Nguyen Mai Hong Tram
I build multimodal AI systems for real workflows.
I work across OCR, language models, retrieval, visual reasoning, and browser inference—from implementation through evaluation and delivery.
Selected builds
Systems I built, the decisions behind them, and what can be inspected publicly.
Problem
Help teachers review Vietnamese secondary-school mathematics work with evidence-backed draft feedback and class-level error analysis.
What I built
I personally built the OCR flow, LLM API integration, RAG pipeline, and frontend, and handled project management and documentation.
Engineering approach
A teacher-in-the-loop workflow combines multimodal grading, source-grounded retrieval, confidence-based routing, citation checks, PII safeguards, and human review.
Evidence
The public capstone demo shows the product workflow; the system also uses a golden set for evaluation.
Sign-language inference in the browser
TMA Solutions · AI Engineer internship
Problem
Make continuous sign-language recognition usable in a real-time browser setting.
What I built
I researched continuous sign-language recognition and text-to-3D-avatar generation, then built an end-to-end inference path from PyTorch models to the browser.
Engineering approach
Converted model inference to ONNX Runtime Web and integrated it with a browser-facing pipeline.
Evidence
The pipeline was evaluated in a real-time video-chat setting.
Problem
Can explicit object relationships improve surgical visual question answering?
What I built
I carried this research from problem formulation through data preparation, model development, evaluation, and manuscript writing.
Engineering approach
Combined visual representations, language processing, and graph-based scene reasoning in an end-to-end VQA pipeline.
Evidence
Using SSG-VQA, a dataset of about 960K question–answer pairs, the model achieved 85.9% accuracy and 85.8% weighted F1 on a 77,198-question visual-oracle test.
Air Quality Forecasting Pipeline
Forecasting prototype · UEH course group project
View public code ↗Problem
Forecast air quality from OpenAQ measurements and make predictions available through an API and monitoring dashboard.
What I built
I implemented the technical pipeline end to end: OpenAQ ingestion and preprocessing, RNN/LSTM/GRU training, FastAPI inference, a Streamlit dashboard, and Docker packaging.
Engineering approach
Compared recurrent forecasting models and connected model inference to a dashboard through a containerized API.
Evidence
The public repository contains the implementation, setup steps, and model evaluation. The hosted demo is no longer running.
Read engineering CV ↗Engineering experience
Applied work across multimodal pipelines, video understanding, and browser inference.
Jul–Aug 2026
AI Research Intern
VinRobotics · AI Platform, VLA Team
Temporally grounded action captioning in egocentric video, with controlled experiments in sampling, grounding, prompting, and model adaptation.
Mar–May 2026
AI Research Engineer Intern
TMA Solutions
Bidirectional sign-language communication research and browser inference using PyTorch and ONNX Runtime Web.
How I work
Tools appear here in the context of the systems they helped deliver.
Applied AI
OCR, LLM APIs, RAG, and multimodal workflows in GradeMind.
Model development
PyTorch, vision-language pipelines, graph reasoning, and controlled evaluation.
Inference and delivery
ONNX Runtime Web, SGLang, FastAPI, PostgreSQL/pgvector, and Docker.
Research foundations
Published work and recognition that inform my engineering practice.
2026
GNN-SurgVQA: Object-Centric Graph Reasoning for Visual Question Answering in Laparoscopic Scene Understanding
International Conference on Data Analytics and Management (ICDAM) · Best Paper · accepted, paper not yet public
Public code ↗2025
Comprehensive Approach to Vietnamese Scene Text Recognition: Challenges, Models, and Framework Development
ICDAM 2025 · Lecture Notes in Networks and Systems, vol. 1601 · Springer · Published
Read paper ↗