Di Chai

Assistant Professor

School of Computing and Artificial Intelligence
Shanghai University of Finance and Economics

Di Chai
Shanghai, China

My research focuses on machine learning systems, with a particular interest in LLM systems. I aim to build high-performance, trustworthy systems with real-world impact. I received my Ph.D. from HKUST, advised by Prof. Kai Chen and Prof. Qiang Yang.

I have also collaborated closely with Prof. Leye Wang and Prof. Junxue Zhang.

Recruiting studentsI am looking for master's and PhD students interested in machine learning systems, particularly LLM systems. Email: chaidi@mail.sufe.edu.cn

To Prospective Students

I am looking for master's and PhD students with strong computer science fundamentals and a genuine passion for building systems. You should enjoy understanding how systems work, turning ideas into working implementations, and investigating why something fails or performs poorly.

I value learning ability and curiosity more than extensive prior research experience. What matters is your willingness to tackle unfamiliar problems: learning the concepts and tools you need, asking thoughtful questions, and testing ideas through careful reasoning and experimentation. I appreciate students who take initiative and revise their assumptions when the evidence calls for it.

Please email chaidi@mail.sufe.edu.cn with a brief introduction to your background and interests. You are welcome to share a project you have built or a technical problem you enjoyed solving—from coursework, personal exploration, open-source contributions, or research. I would like to hear how you approached it and what you learned.

Highlighted Projects

KVMem

arXiv 2026 · Open source

Virtualizes million-token agent workspaces on a consumer GPU, keeping history as reusable KV blocks and retrieving only the context needed for each step.

Paper ↗ Code ↗ Tutorial video ↗
Research insight

The paper separates the persistent KV workspace from the model's active context window. Query-conditioned retrieval selects relevant historical blocks and assembles a bounded execution view, reusing previously computed KV state instead of repeatedly prefilling recalled text.

Centrifuge

ICLR 2026 · Open source

Algorithm–system co-design that turns token filtering into practical training speedups while preserving its model utility benefits.

Paper ↗ Code ↗
Research insight

Token filtering alone does not guarantee faster training. Centrifuge filters activations in attention backward and transforms sparse GEMM into dimension-reduced dense GEMM, enabling efficient execution with standard machine learning libraries.

Up to 34.7% less end-to-end training time. At 50% token filtering; maximum across evaluated settings on 1.1B–40B models.

Conference abstract ↗

Publications

2026

Canopy: Tree-Aware Rollout Scheduling for Agent Reinforcement Learning

Feiyuan Zhang, LI Pengbo, Ziniu Li, Yuhao Jiang, Di Chai, Han Tian, Junhao Wang, Taiqiang Wu, Guanhua Huang, Chaoliang Zeng, Yihao Liu, Kai Chen

2026

PrefixFlow: Training-Time KV Caching for Schedule-Level Prefix Reuse in LLM RL Training

Pengbo Li, Feiyuan Zhang, Guangming Sheng, Guangxin He, Di Chai, Ziniu Li, Taiqiang Wu, Han Tian, Wenyu Mao, Binhang Yuan, Kai Chen

All publicationsShow selected publications

Click image or background to close · Esc