Nguyen Mai Hong Tram

I study how AI understands actions, time, and interaction.

My research has evolved from computer vision and medical visual question answering toward multimodal and video understanding, with recent work on temporally grounded action understanding in egocentric video.

Multimodal learning / Video understanding / Action and interaction

From perception to action

Five connected questions trace the evolution of my research interests. The last is a future direction, not a completed project.

How can models reliably extract structured information from complex visual scenes?

Early work in Vietnamese scene-text detection and recognition established my interest in robust visual representations.

Can explicit object relationships improve surgical visual question answering?

I developed GNN-SurgVQA from problem formulation through model development, evaluation, and manuscript writing.

How can models interpret gestures that evolve through time rather than static scenes?

At TMA Solutions, I investigated continuous sign-language recognition and text-to-3D-avatar generation. This shifted my focus toward sequential actions.

How should vision-language models represent actions that unfold over time?

At VinRobotics, I developed and evaluated multi-stage VLM pipelines for temporally grounded action captioning in egocentric video.

How can multimodal models understand action and interaction well enough to support intelligent agents?

My current direction of interest extends this path toward action-centric and embodied intelligence, including VLM and VLA research.

Visual perception

How can models reliably extract structured information from complex visual scenes?

From recognizing what is visible to understanding what is happening.

Questions I want to keep asking

These interests connect my past work to the research I hope to pursue next.

Multimodal learning

Connecting visual information with language and reasoning across modalities.

Video and temporal understanding

Studying action, temporal order, and grounding in egocentric video.

Action-centric intelligence

Exploring models that can understand interactions and support embodied agents.

Selected research

Research questions, my contribution, and the evidence or learning each project produced.

Object-centric multimodal reasoningICDAM 2026 · Best Paper · paper accepted, not yet public

GNN-SurgVQA

Can explicit object relationships improve surgical visual question answering?

My contribution

I carried this research from problem formulation through data preparation, model development, evaluation, and manuscript writing.

Evidence & learning

Using SSG-VQA, a dataset of about 960K question–answer pairs, the model achieved 85.9% accuracy and 85.8% weighted F1 on a 77,198-question visual-oracle test.

Visual oracle uses bounding boxes supplied with SSG-VQA rather than predictions from the fine-tuned YOLOv8 detector.

Explore public code
Egocentric video at VinRoboticsResearch internship · VinRobotics

Temporally grounded action understanding

How should vision-language models represent actions that unfold over time?

My contribution

I developed and evaluated multi-stage vision-language pipelines, and designed controlled experiments on frame sampling, temporal grounding, prompting, and model adaptation.

Evidence & learning

I analyzed directional ambiguity, dataset shift, and overfitting, and investigated recent VLM/VLA methods through literature review and large-scale experimentation using SGLang.

Research internship, AI Platform VLA Team, July to August 2026.

A parallel line of inquiry

Mathematical decision-making under real-world constraints.

Emergency blood transportation

How can emergency logistics balance cost and hospital waiting time under operational constraints?

I formulated a multi-objective electric vehicle routing problem with battery-swapping stations and time windows, then developed hybrid NSGA-II and SPEA2 methods with Savings, 2-opt, and 3-opt heuristics.

Validated on standard EVRP-TW benchmarks and a Ho Chi Minh City case study involving 36 hospitals and 6 battery-swapping stations.

UEH Young Researchers Award, Prize B · Eureka 2025 Semifinalist

Explore public code

Publications

Published and accepted work, with public links where available.

2026

GNN-SurgVQA: Object-Centric Graph Reasoning for Visual Question Answering in Laparoscopic Scene Understanding

International Conference on Data Analytics and Management (ICDAM)

Best Paper · accepted, paper not yet public

Public code
2025

Comprehensive Approach to Vietnamese Scene Text Recognition: Challenges, Models, and Framework Development

ICDAM 2025 · Lecture Notes in Networks and Systems, vol. 1601 · Springer

Published

Read paper

Experience

Research and engineering work across vision, language, and sequential understanding.

Jul–Aug 2026

AI Research Intern

VinRobotics · AI Platform, VLA Team

Temporally grounded action captioning in egocentric video, with controlled experiments in sampling, grounding, prompting, and model adaptation.

Mar–May 2026

AI Research Engineer Intern

TMA Solutions

Bidirectional sign-language communication research and browser inference using PyTorch and ONNX Runtime Web.

Mar 2025–Jan 2026

Research Collaborator

University of Economics Ho Chi Minh City

Surgical VQA, Vietnamese scene-text recognition, and multi-objective optimization across the research lifecycle.

Other projects

Applied systems work alongside the research portfolio.

GradeMind

AI-assisted mathematics grading

Vingroup × VinUniversity Applied AI Talent Program capstone

A teacher-in-the-loop system for Vietnamese secondary-school mathematics, with evidence-backed feedback, class-level error analysis, confidence-based routing, and PII safeguards.

View live demo

Education & programs

B.Sc. in Data Science

University of Economics Ho Chi Minh City

Sep 2022–Mar 2026

3.92 / 4.00GPA

Rank #2 in Data ScienceTop 5% of the School of Business Information Technology

  • Vingroup × VinUniversity Applied AI Talent Program, May–Aug 2026
  • AI Vietnam AIO2026, Jun 2026–present