About
Research-trained. Systems-minded. Production accountable.
I'm Sang Truong (@sangtrx), a senior AI engineer and applied AI lead based in Ho Chi Minh City, Vietnam. My work spans model research, LLM/agent systems, computer vision and video intelligence, quantitative ML, distributed backends, product engineering, cloud/edge deployment, and the operational boundaries that decide whether an AI system is actually trustworthy.
Operating model
Own the path from research to acceptance.
I am most useful when a system crosses several disciplines at once: research assumptions, data contracts, model/tool behavior, backend state, user experience, infrastructure, and stakeholder acceptance. I do not treat the model as the whole product.
My current Head of AI role covers clinical AI, education AI, production computer vision/video intelligence, and AI platform engineering. Earlier work spans enterprise conversational agents, multilingual speech systems, environmental forecasting, quantitative research, multimodal/video research, and embedded computer vision.
I prefer explicit boundaries over architecture theater: which component owns facts, what data was available at decision time, what can fail, what is recoverable, what has actually been validated, and what the system deliberately does not claim.
Domain depth
Specialized depth without amputating the rest of the career.
The résumé architecture intentionally keeps one canonical chronology while expanding evidence by domain. These are the four recurring technical centers.
Applied AI · Agents · Voice · RAG
Production AI systems where retrieval, tools, memory, voice/real-time interaction, authority, safety, evaluation, and distributed application state matter as much as the model call.
- Ho Chi Minh City Traditional Medicine Hospital clinical AI: one semantic owner for request/reference/source/tool/scope, bounded approved-corpus research, durable Evidence Workspace, deterministic clinical authority and exact citations
- FPT AI4U: Azure OpenAI, LangGraph/LangChain, Qdrant RAG, tool/model routing, code execution, web search, memory controls and guardrails
- Voice and tool workflows: multilingual Azure Speech transcription/evaluation, Vietnamese ASR/TTS, WebSockets/SSE, Playwright-based browser verification and failure handling
FastAPI · LangGraph · LangChain · Azure OpenAI · Azure Speech · Qdrant · Milvus · pgvector · Playwright · Celery · RabbitMQ · Redis · PostgreSQL · Docker
Computer Vision · Video · Edge AI
Physical-world AI spanning camera/media ingest, perception, tracking, temporal events, evidence capture, edge inference, and recovery under unreliable real-world conditions.
- Production multi-camera video intelligence: YOLO11, Vietnamese ALPR, ByteTrack, line crossing, event video, identity/freshness checks, watchdog recovery
- University of Arkansas: temporal action understanding, vision-language modeling, industrial CV, CarcassFormer, YOLOv8 Jetson deployment
- 5D Agriculture: autonomous braking, face recognition, Intel RealSense D435 RGB-D livestock measurement, embedded AI
PyTorch · TensorFlow · OpenCV · YOLO11/YOLOv8 · Fast-ALPR · ByteTrack · TensorRT · CUDA · ONNX Runtime · Jetson · FFmpeg · MediaMTX
Quantitative Research · Trading Systems
A research-to-production stack built around point-in-time evidence, reusable causal computation, leakage/multiplicity control, durable signal/risk state, execution/reconciliation, and public verification boundaries.
- Curren research: Rust/Python causal core, Arrow/Parquet PIT evidence, shared timeframes/primitives/events, versioned event store, hypothesis views, global OOF and append-only Alpha History
- Curren production path: normalized external alpha-source data, fail-closed ML quality gate, restart-safe lifecycle/risk, guarded execution, reconciliation and research↔streaming parity
- Confidential Fund + Bluebelt: equity/crypto/FX quantitative research, ML ensembles, sentiment-derived signals, AWS execution and MLflow experimentation
Rust · Python · Arrow · PyArrow/Parquet · Polars · DuckDB · LightGBM · CatBoost · XGBoost · SciPy · Statsmodels · Optuna · NautilusTrader
Research · Multimodal & Temporal ML
Peer-reviewed research across temporal video understanding, vision-language learning, medical time-series representation learning, and industrial computer vision.
- ABN → AEI → AOE-Net: action boundaries and actor/object/environment interaction modeling for long untrimmed video
- VLCAP → VLTinT: contrastive vision-language learning and coherent video paragraph captioning; VLTinT was an AAAI 2023 Oral
- sCL-ST + CarcassFormer: medical time-series contrastive learning and industrial localization/segmentation/classification
Transformers · Contrastive Learning · PyTorch · TensorFlow · Detectron2 · MATLAB · NumPy · SciPy · Scikit-learn · Weights & Biases
How I think
Good AI systems are contracts, not demos.
Give models room to reason where useful while keeping authoritative facts, permissions, tools, and failure behavior explicit.
In research and live systems, respect what information existed at decision time. Point-in-time discipline and provenance are engineering requirements.
Separate fluent output from the source, tool result, experiment, runtime observation, or acceptance gate that supports it.
State, retries, idempotency, restart behavior, monitoring, recovery, access control, deployment and handoff are part of the product.
For physical-world AI, camera/media reliability and target-device inference behavior matter more than isolated benchmark numbers.
Implemented, tested, deployed, accepted, and production-ready are different claims. I try to keep them different.
Technical stack
Tools organized by system responsibility.
The stack is broad because the systems cross model, data, product, infrastructure, and physical-world boundaries.
Applied AI / LLM / agents
LLM/RAG/agents · LangGraph/LangChain orchestration · tool/model routing · code/tool execution · web-search workflows · memory/context management · grounded generation · guardrails · human-review boundaries · evaluation · authority separation
Retrieval & knowledge systems
Qdrant · Milvus · Pinecone · pgvector · embeddings · chunking/metadata strategy · governed ingestion · metadata/permission filtering · exact citations · immutable provenance · retrieval evaluation · retrieval failure handling
Voice / real-time AI
Azure Speech · multilingual ASR/transcription · Vietnamese ASR/TTS · speech evaluation · WebSockets/SSE · real-time conversation state · AsyncIO · background processing
Computer vision / video / edge
YOLO11/YOLOv8 · PyTorch · TensorFlow · OpenCV · Detectron2 · Fast-ALPR · ByteTrack · TensorRT · ONNX Runtime · NVIDIA Jetson · CUDA · RGB-D · MediaMTX · RTSP/HLS · FFmpeg
Quantitative ML & research data
Rust/Python causal core · Point-in-time data · Arrow/PyArrow/Parquet · versioned event stores · global OOF · purge/embargo · multiple-testing controls · Alpha History · LightGBM · CatBoost · XGBoost · Scikit-learn · Polars · DuckDB · SciPy · Statsmodels · Optuna · NautilusTrader
Backend & distributed systems
Python · FastAPI · AsyncIO · REST/SSE · WebSockets · Celery · RabbitMQ · Redis · PostgreSQL · SQLite · MongoDB · MinIO/S3 · retries · idempotency · durable state
Product, browser automation, cloud & delivery
Next.js · React · TypeScript/JavaScript · Open edX · Playwright · browser/tool verification · Docker · Kubernetes · Linux · Windows · systemd · Nginx · AWS · Azure
Languages
English — professional working proficiency · Vietnamese — native
Research trajectory
From temporal video reasoning to applied AI systems.
Temporal action understanding
ABN → AEI → AOE-Net: action-boundary and actor/object/environment interaction modeling for long untrimmed video, culminating in IJCV.
Vision-language video understanding
VLCAP → VLTinT: contrastive and Transformer-based modeling for coherent video paragraph captioning; VLTinT was selected as an AAAI 2023 Oral.
Medical & industrial ML
sCL-ST for multi-lead ECG representation learning and CarcassFormer for poultry defect localization, segmentation, and classification, alongside Jetson/TensorRT deployment work.
Education
Research-trained engineering.
Master of Engineering (MEng) in Computer Engineering · University of Arkansas
GPA 4.0/4.0 · thesis: “Towards Multi-modal Interpretable Video Understanding” · advisor: Prof. Ngan Le · fully funded Ph.D. admission in 2021 followed by endowed graduate scholarships in 2022 and 2023.
Bachelor of Science in Automation and Control Engineering · International University — VNU HCMC
GPA 3.5/4.0 · Top 1% · half-tuition scholarship recipient.
Selected research
Peer-reviewed work across multiple AI domains.
Simultaneous localization, segmentation, and classification of poultry carcass defects Poultry Science 2023 AOE-Net
Entity-interaction modeling with adaptive attention for temporal action proposal generation International Journal of Computer Vision 2023 VLTinT
Visual-linguistic Transformer-in-Transformer for coherent video paragraph captioning · AAAI Oral AAAI 2023 sCL-ST
Supervised contrastive learning with semantic transformations for multi-lead ECG arrhythmia classification IEEE JBHI 2022 VLCAP
Vision-language contrastive learning for coherent video paragraph captioning IEEE ICIP 2021 AEI
Actor-environment interaction with adaptive attention for temporal action proposal generation BMVC 2021 ABN
Agent-aware boundary networks for temporal action proposal generation IEEE Access 2021 Multi-module RCNN + Transformer
Multi-module recurrent convolutional neural network with Transformer encoder for ECG arrhythmia classification IEEE BHI Full Google Scholar profile
Explore