Efficient AI Researcher

Yuexiao Ma (马跃萧)

Research Fellow at MMLab@NTU, S-Lab · Ph.D. at Xiamen University

I make large generative and vision models faster, smaller, and more practical. My work spans low-bit quantization, model compression, and efficient video generation. I always welcome research collaborations and academic exchanges. Please feel free to email me with any questions.

Yuexiao Ma smiling outdoors at a city viewpoint

I am a Research Fellow at MMLab@NTU, S-Lab, working with Prof. Chen Change Loy. I received my Ph.D. in Intelligent Science and Technology from Xiamen University, advised by Prof. Rongrong Ji and Assoc. Prof. Xiawu Zheng.

Previously, I spent more than three years with ByteDance's quantization team, translating compression research into low-cost, low-latency inference for generative and foundation models.

Making intelligence efficient by design.

I study the algorithms and systems that remove unnecessary computation without sacrificing model capability.

R1

Low-bit Quantization

Post-training and mixed-precision methods for LLMs, vision transformers, diffusion models, and production inference.

R2

Efficient Video Generation

Temporal and flow-aware caching strategies that accelerate autoregressive generation while preserving visual quality.

R3

Model Compression

Theory-guided pruning, outlier handling, and hardware-aware optimization for deployable neural networks.

Publications

* Equal contribution · Corresponding author

Google Scholar
Overview of Flow Caching
ICLR 2026Video generation

Flow Caching for Autoregressive Video Generation

Yuexiao Ma*, Xuzhe Zheng*, Jing Xu*, Xiwei Xu, Feng Ling, Xiawu Zheng, Huafeng Kuang, Huixia Li, Xing Wang, Xuefeng Xiao, Fei Chao, Rongrong Ji

Flow-aware feature reuse accelerates autoregressive video generation while preserving temporal quality.

MotionCache frame importance map visualization
ICML 2026Video generation

Motion-Aware Caching for Efficient Autoregressive Video Generation

Jing Xu*, Yuexiao Ma*, Xuzhe Zheng, Xing Wang, Shiwei Liu, Chenqian Yan, Xiawu Zheng, Rongrong Ji, Fei Chao, Songwei Liu

A coarse-to-fine cache policy uses inter-frame motion to allocate computation where it matters.

Overview of training-free multimodal orchestration
ICML 2026Multimodal LLMs

Training-Free Multimodal Large Language Model Orchestration

Tianyu Xie, Yuexiao Ma, Yuhang Wu, Wang Chen, Jiayi Ji, Tat-Seng Chua, Xiawu Zheng, Rongrong Ji

A controller composes specialized modality experts into a unified training-free multimodal system.

A2RBench generation pipeline overview
ICML 2026Reasoning benchmark

A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

Qingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie, Xiawu Zheng, Rongrong Ji

An automated pipeline generates abstract reasoning tasks with programmatic verification.

Overview of polybasic speculative decoding
ICML 2025Speculative decoding

Polybasic Speculative Decoding Through a Theoretical Perspective

Ruilin Wang, Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao, Xuefeng Xiao, Rongrong Ji

A theoretical framework generalizes draft-and-verify decoding to multiple interconnected draft models.

ITPruner framework overview
IJCV 2025Network pruning

An Information Theory-inspired Strategy for Automatic Network Pruning

Xiawu Zheng*, Yuexiao Ma*, Teng Xi, Gang Zhang, Errui Ding, Yuchao Li, Jie Chen, Yonghong Tian, Rongrong Ji

Information-theoretic layer importance enables one-pass, search-free automatic pruning.

Relationship between reconstruction granularity, quantization loss, and model performance
ICML 2024Vision quantization

Outlier-aware Slicing for Post-Training Quantization in Vision Transformer

Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, Xuefeng Xiao, Rui Wang, Shilei Wen, Fei Chao, Rongrong Ji

Outlier-aware channel slicing substantially improves accurate INT4 deployment of vision transformers.

AffineQuant transformation overview
ICLR 2024LLM quantization

AffineQuant: Affine Transformation Quantization for Large Language Models

Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, Xuefeng Xiao, Rui Wang, Shilei Wen, Fei Chao, Rongrong Ji

Learned affine transformations reduce activation outliers and post-training quantization error.

MRECG method overview
CVPR 2023Quantization theory

Solving Oscillation Problem in Post-Training Quantization Through a Theoretical Perspective

Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao, Rui Wang, Shilei Wen, Xin Pan, Fei Chao, Rongrong Ji

A theoretical diagnosis of PTQ oscillation leads to mixed reconstruction granularity optimization.

OMPQ framework overview
AAAI 2023Mixed precision

OMPQ: Orthogonal Mixed Precision Quantization

Yuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang, Huixia Li, Yongjian Wu, Yunsheng Wu, Guannan Jiang, Wei Zhang, Rongrong Ji

Orthogonality-guided sensitivity estimates allocate mixed precision without iterative search.

Boundary value problem paper overview
BVP 2019Applied mathematics

Positive Solutions of Fourth-order Problems with Dependence on All Derivatives in Nonlinearity under Stieltjes Integral Boundary Conditions

Yuexiao Ma, Chenyang Yin, Guowei Zhang

Existence results for positive solutions of fourth-order boundary value problems under Stieltjes conditions.

Present

Research

Research Fellow

MMLab@NTU · S-Lab

Working with Prof. Chen Change Loy

2022 — 2025

Industry

Research Intern · Quantization Team

ByteDance Data · Beijing

2022 — 2026

Education

Ph.D. · Intelligent Science and Technology

Xiamen University · Department of Artificial Intelligence

Advisors: Prof. Rongrong Ji and Assoc. Prof. Xiawu Zheng

2020 — 2022

Education

M.S. · Intelligent Science and Technology

Xiamen University · Department of Computer Science

2016 — 2020

Education

B.S. · Mathematics and Applied Mathematics

Northeastern University · Shenyang

Journal reviewer

IEEE TIP · IEEE TMM

Conference reviewer

ECCV · NeurIPS · ICML · ICLR · AAAI · CVPR

Recognition

Top Reviewer

Research · Collaboration · Ideas

Let's make ambitious models practical.

bobma@stu.xmu.edu.cn