Nguyen Mai Hong Tram

I build multimodal AI systems for real workflows.

I work across OCR, language models, retrieval, visual reasoning, and browser inference—from implementation through evaluation and delivery.

OCR / RAG / VLM / Browser inference

Selected builds

Systems I built, the decisions behind them, and what can be inspected publicly.

GradeMind

AI-assisted mathematics grading

Open public demo

Problem

Help teachers review Vietnamese secondary-school mathematics work with evidence-backed draft feedback and class-level error analysis.

What I built

I personally built the OCR flow, LLM API integration, RAG pipeline, and frontend, and handled project management and documentation.

Engineering approach

A teacher-in-the-loop workflow combines multimodal grading, source-grounded retrieval, confidence-based routing, citation checks, PII safeguards, and human review.

Evidence

The public capstone demo shows the product workflow; the system also uses a golden set for evaluation.

  • OCR
  • LLM API
  • RAG
  • FastAPI
  • PostgreSQL/pgvector
  • React/TypeScript

Sign-language inference in the browser

TMA Solutions · AI Engineer internship

Problem

Make continuous sign-language recognition usable in a real-time browser setting.

What I built

I researched continuous sign-language recognition and text-to-3D-avatar generation, then built an end-to-end inference path from PyTorch models to the browser.

Engineering approach

Converted model inference to ONNX Runtime Web and integrated it with a browser-facing pipeline.

Evidence

The pipeline was evaluated in a real-time video-chat setting.

  • PyTorch
  • ONNX Runtime Web
  • Browser inference

GNN-SurgVQA

Multimodal surgical visual question answering

View public code

Problem

Can explicit object relationships improve surgical visual question answering?

What I built

I carried this research from problem formulation through data preparation, model development, evaluation, and manuscript writing.

Engineering approach

Combined visual representations, language processing, and graph-based scene reasoning in an end-to-end VQA pipeline.

Evidence

Using SSG-VQA, a dataset of about 960K question–answer pairs, the model achieved 85.9% accuracy and 85.8% weighted F1 on a 77,198-question visual-oracle test.

  • PyTorch
  • PyTorch Geometric
  • Model evaluation

Air Quality Forecasting Pipeline

Forecasting prototype · UEH course group project

View public code

Problem

Forecast air quality from OpenAQ measurements and make predictions available through an API and monitoring dashboard.

What I built

I implemented the technical pipeline end to end: OpenAQ ingestion and preprocessing, RNN/LSTM/GRU training, FastAPI inference, a Streamlit dashboard, and Docker packaging.

Engineering approach

Compared recurrent forecasting models and connected model inference to a dashboard through a containerized API.

Evidence

The public repository contains the implementation, setup steps, and model evaluation. The hosted demo is no longer running.

  • OpenAQ
  • PyTorch
  • RNN/LSTM/GRU
  • FastAPI
  • Streamlit
  • Docker
Read engineering CV

Engineering experience

Applied work across multimodal pipelines, video understanding, and browser inference.

Jul–Aug 2026

AI Research Intern

VinRobotics · AI Platform, VLA Team

Temporally grounded action captioning in egocentric video, with controlled experiments in sampling, grounding, prompting, and model adaptation.

Mar–May 2026

AI Research Engineer Intern

TMA Solutions

Bidirectional sign-language communication research and browser inference using PyTorch and ONNX Runtime Web.

How I work

Tools appear here in the context of the systems they helped deliver.

Applied AI

OCR, LLM APIs, RAG, and multimodal workflows in GradeMind.

Model development

PyTorch, vision-language pipelines, graph reasoning, and controlled evaluation.

Inference and delivery

ONNX Runtime Web, SGLang, FastAPI, PostgreSQL/pgvector, and Docker.

Research foundations

Published work and recognition that inform my engineering practice.

2026

GNN-SurgVQA: Object-Centric Graph Reasoning for Visual Question Answering in Laparoscopic Scene Understanding

International Conference on Data Analytics and Management (ICDAM) · Best Paper · accepted, paper not yet public

Public code

2025

Comprehensive Approach to Vietnamese Scene Text Recognition: Challenges, Models, and Framework Development

ICDAM 2025 · Lecture Notes in Networks and Systems, vol. 1601 · Springer · Published

Read paper