How can models reliably extract structured information from complex visual scenes?
Early work in Vietnamese scene-text detection and recognition established my interest in robust visual representations.
Concept illustration, not research data
Nguyen Mai Hong Tram
My research has evolved from computer vision and medical visual question answering toward multimodal and video understanding, with recent work on temporally grounded action understanding in egocentric video.
Multimodal learning / Video understanding / Action and interaction
Five connected questions trace the evolution of my research interests. The last is a future direction, not a completed project.
Early work in Vietnamese scene-text detection and recognition established my interest in robust visual representations.
Concept illustration, not research data
I developed GNN-SurgVQA from problem formulation through model development, evaluation, and manuscript writing.
Concept illustration, not research data
At TMA Solutions, I investigated continuous sign-language recognition and text-to-3D-avatar generation. This shifted my focus toward sequential actions.
Concept illustration, not research data
At VinRobotics, I developed and evaluated multi-stage VLM pipelines for temporally grounded action captioning in egocentric video.
Concept illustration, not research data
My current direction of interest extends this path toward action-centric and embodied intelligence, including VLM and VLA research.
Concept illustration, not research data
Visual perception
Concept illustration, not research data
How can models reliably extract structured information from complex visual scenes?
Multimodal reasoning
Concept illustration, not research data
Can explicit object relationships improve surgical visual question answering?
Sequential human actions
Concept illustration, not research data
How can models interpret gestures that evolve through time rather than static scenes?
Temporal grounding
Concept illustration, not research data
How should vision-language models represent actions that unfold over time?
Future direction
Concept illustration, not research data
How can multimodal models understand action and interaction well enough to support intelligent agents?
From recognizing what is visible to understanding what is happening.
These interests connect my past work to the research I hope to pursue next.
Connecting visual information with language and reasoning across modalities.
Studying action, temporal order, and grounding in egocentric video.
Exploring models that can understand interactions and support embodied agents.
Research questions, my contribution, and the evidence or learning each project produced.
Can explicit object relationships improve surgical visual question answering?
I carried this research from problem formulation through data preparation, model development, evaluation, and manuscript writing.
Using SSG-VQA, a dataset of about 960K question–answer pairs, the model achieved 85.9% accuracy and 85.8% weighted F1 on a 77,198-question visual-oracle test.
Visual oracle uses bounding boxes supplied with SSG-VQA rather than predictions from the fine-tuned YOLOv8 detector.
Explore public codeHow should vision-language models represent actions that unfold over time?
I developed and evaluated multi-stage vision-language pipelines, and designed controlled experiments on frame sampling, temporal grounding, prompting, and model adaptation.
I analyzed directional ambiguity, dataset shift, and overfitting, and investigated recent VLM/VLA methods through literature review and large-scale experimentation using SGLang.
Research internship, AI Platform VLA Team, July to August 2026.
Mathematical decision-making under real-world constraints.
How can emergency logistics balance cost and hospital waiting time under operational constraints?
I formulated a multi-objective electric vehicle routing problem with battery-swapping stations and time windows, then developed hybrid NSGA-II and SPEA2 methods with Savings, 2-opt, and 3-opt heuristics.
Validated on standard EVRP-TW benchmarks and a Ho Chi Minh City case study involving 36 hospitals and 6 battery-swapping stations.
UEH Young Researchers Award, Prize B · Eureka 2025 Semifinalist
Explore public codePublished and accepted work, with public links where available.
International Conference on Data Analytics and Management (ICDAM)
Best Paper · accepted, paper not yet public
ICDAM 2025 · Lecture Notes in Networks and Systems, vol. 1601 · Springer
Published
Research and engineering work across vision, language, and sequential understanding.
Jul–Aug 2026
VinRobotics · AI Platform, VLA Team
Temporally grounded action captioning in egocentric video, with controlled experiments in sampling, grounding, prompting, and model adaptation.
Mar–May 2026
TMA Solutions
Bidirectional sign-language communication research and browser inference using PyTorch and ONNX Runtime Web.
Mar 2025–Jan 2026
University of Economics Ho Chi Minh City
Surgical VQA, Vietnamese scene-text recognition, and multi-objective optimization across the research lifecycle.
Applied systems work alongside the research portfolio.
AI-assisted mathematics grading
Vingroup × VinUniversity Applied AI Talent Program capstone
A teacher-in-the-loop system for Vietnamese secondary-school mathematics, with evidence-backed feedback, class-level error analysis, confidence-based routing, and PII safeguards.
View live demoUniversity of Economics Ho Chi Minh City
Sep 2022–Mar 2026
3.92 / 4.00GPA
Rank #2 in Data ScienceTop 5% of the School of Business Information Technology