Yuanhao Ban
banyh2000 at gmail.com
I’m a Ph.D. student in the UCLA Department of Computer Science, advised by Professor Cho-Jui Hsieh.
I am broadly interested in rubric rewards, post-training and agentic systems.
During the third year of my Ph.D., I worked on text-to-image reward modeling and post-training at Arena Inc. with I-Hung Hsu and Wei-Lin Chiang.
During the first two years of my Ph.D., I was fortunate to work with Ruochen Wang, Prof. Boqing Gong, Prof. Minhao Cheng, and Prof. Tianyi Zhou.
I graduated from Tsinghua University in 2023 with a degree in Electrical Engineering. At Tsinghua, I was fortunate to conduct research at TSAIL Lab, where I worked with Dr. Yinpeng Dong and Prof. Jun Zhu.
Here is my CV. This site was last updated in October 2026.
news
| Oct 02, 2026 | We released our work on post-training frontier text-to-image models by combining human preference and rubric rewards. Our recipe improves FLUX.2-dev’s Arena Elo score by 69 points, and our post-trained Ideogram-4 ranks above every open-source model in the September 4, 2026 Arena leaderboard snapshot. |
|---|---|
| Sep 24, 2026 | Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist was accepted to NeurIPS 2026 as a Spotlight! |
| Sep 11, 2026 | We released ReCAST, a method for assigning reward credit across diffusion timesteps. |
| Aug 20, 2026 | We released Stream4D, which uses 4D reconstruction rewards and motion priors to improve scene consistency and preserve natural motion in streaming video generation. |
publications
Selected publications highlight a subset of my research, including Self-Forcing++.
- Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric RewardsarXiv preprint arXiv:2610.02967, 2026
-