Wentao (Tony) Ma

@ Luma AI

Multi-Modal LLM, Unified Physical Model

Toronto & San Francisco
Email: tonyyyma [at] gmail [dot] com

Google Scholar GitHub LinkedIn

Portrait of Wentao Ma

Introduction

The research area I'm focusing on is multi-modal models. I enjoy improving and exploring the abilities of MLLMs and applying them to other fields like robotics.

Currently, I'm doing research on physical foundation models at LumaAI, where I mainly work on pre/mid-training. Before that, I worked at @BosonAI with Alex Smola and Mu Li on post-training for audio foundation models like Higgs-tts-2, -2.5, and Higgs-realtime.

Earlier, I got my Master's degree from the University of Toronto, advised by Dr. Zhijing Jin. In the meantime, I worked closely with Wenhu Chen on video understanding. I also studied at Imperial College London, supervised by Edward Johns, where we validated and improved the multi-modal pattern learning ability of VLMs and applied them to robotics.

I like photography and I'm a member of Toronto Photo Walks (ToPW). I'm also interested in all kinds of sports, including snowboarding and tennis.

News

Works

Higgs Audio

Higgs Audio-LLM Family

Core Contributor, Boson AI Team, 2025-2026

[Repo] [HF] [STT-v3 Blog] [TTS-v2.5 Blog] [Realtime Blog]

IHBench

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

Ahmad Salimi, Wentao Ma, Yuzhi Tang, Dongming Shen, Mu Li, Alex Smola

Preprint

[blog] [paper] [HF] [Github]

Instruct-FD

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

Yuzhi Tang, Wentao Ma, Xiling Zhao, Ahmad Salimi, Sepehr Harfi Moridani, Dongming Shen, and others

Conference on Language Modeling (COLM), 2026

[paper]

WildASR

Back to Basics: Revisiting ASR in the Age of Voice Agents

Geeyang Tay*, Wentao Ma*, Jaewon Lee, Yuzhi Tang, and others

Conference on Language Modeling (COLM), 2026

[paper] [HF] [Github]

Online data reweighting

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods

Wanru Zhao, Yihong Chen, Yuzhi Tang*, Wentao Ma*, and others

International Conference on Learning Representations (ICLR), 2026

[paper]

VideoScore2

VideoScore2: Think before You Score in Generative Video Evaluation

Xuan He*, Dongfu Jiang*, Ping Nie, Minghao Liu, Wentao Ma, Junru Lin, and others

Transactions on Machine Learning Research (TMLR), 2026

[paper] [website]

StructEval

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Jialin Yang*, Dongfu Jiang*, Lipeng He, Sherman Siu, Wentao Ma, Zhiheng Lyu, and others

Transactions on Machine Learning Research (TMLR), 2025

[paper] [website] [benchmark]

VideoEval-Pro

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Wentao Ma*, Weiming Ren*, Yiming Jia, Zhuofeng Li, Ping Nie, Ge Zhang, Wenhu Chen

Transactions on Machine Learning Research (TMLR), 2026

[paper] [website] [benchmark] [Leaderboard]

ProT-GFDM

ProT-GFDM: A Generative Fractional Diffusion Model for Protein Generation

Xiao Liang*, Wentao Ma*, Eric Paquet, Herna Lydia Viktor, Wojtek Michalowski

Computational and Structural Biotechnology Journal (CSBJ), 2025

[paper]

Vamba

Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

Weiming Ren, Wentao Ma, Huan Yang, Cong Wei, Ge Zhang, Wenhu Chen

International Conference on Computer Vision (ICCV), 2025

[paper] [website]

Paint2Plan

Paint2Plan: Image Painting Enables Imitation Learning with VLMs

Tony Ma, Teyun Kwon, Edward Johns

Preprint, 2024

[paper] [website]

LLM Echo Chamber

LLM Echo Chamber: personalized and automated disinformation

Tony Ma, Yves-Alexandre de Montjoye

Machine Learning and Cyber Security Symposium (MLCSS), Imperial, 2024

[paper] [code] [video]

Adversarial patches

Boosting Transferability of Adversarial Patches with Visual Relations

Tony Ma, Songze Li, Yisong Xiao, Shunchang Liu

Conference on Computer Vision and Pattern Recognition (CVPR), AdvVision Workshop, 2023

[paper]

Experience

Luma AI logo

Luma AI

Research Scientist / Engineer

Working on pre/mid-training for physical foundation models

Aug.2026 - Present [website]

Boson AI logo

Boson AI

Machine Learning Engineer

Post-training for Audio Understanding and Generation models

May.2025 - July.2026 [website]

Vector Institute logo

Vector Institute

Machine Learning Associate

Designed a Geo-filtering RAG system with Global Spatial Technology Solutions (GSTS)

Jan.2025 - Apr.2025 [website]

Sony logo

SONY

Edge AI Engineer Intern

Video Object Tracking / Model Quantization / Edge Computing

Sep.2022 - Feb.2023 [website] [Project]

TikTok logo

TikTok

Software Engineer Intern

iOS development for TikTok Pay

May.2022 - Aug.2022 [website]

Selected Certifications and Awards

UofT MScAC Graduate Spotlight --- 2026
AWS Certified Solutions Architect (Associate) --- 2026
Mitacs Research Funding --- 2025-2026
Distinction @ Imperial College London --- 2024
Outstanding Graduates --- 2023
Scholarship for Academic Excellence --- 2020/2021/2022
Scholarship for Discipline Competitions --- 2020/2021/2022
Excellent Student Leader --- 2020

Community Service

Reviewer: AAAI, COLM, RA-L, ICLR workshops, CVPR workshops

© Wentao Ma | Template from Dr. Yueming Jin | Last updated: Sep 2026