Yuanhao Ban

prof_pic.jpg

banyh2000 at gmail.com

I’m a Ph.D. student in the UCLA Department of Computer Science, advised by Professor Cho-Jui Hsieh.

I am broadly interested in rubric rewards, post-training and agentic systems.

During the third year of my Ph.D., I worked on text-to-image reward modeling and post-training at Arena Inc. with I-Hung Hsu and Wei-Lin Chiang.

During the first two years of my Ph.D., I was fortunate to work with Ruochen Wang, Prof. Boqing Gong, Prof. Minhao Cheng, and Prof. Tianyi Zhou.

I graduated from Tsinghua University in 2023 with a degree in Electrical Engineering. At Tsinghua, I was fortunate to conduct research at TSAIL Lab, where I worked with Dr. Yinpeng Dong and Prof. Jun Zhu.

Here is my CV. This site was last updated in October 2026.

news

Oct 02, 2026 We released our work on post-training frontier text-to-image models by combining human preference and rubric rewards. Our recipe improves FLUX.2-dev’s Arena Elo score by 69 points, and our post-trained Ideogram-4 ranks above every open-source model in the September 4, 2026 Arena leaderboard snapshot.
Sep 24, 2026 Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist was accepted to NeurIPS 2026 as a Spotlight!
Sep 11, 2026 We released ReCAST, a method for assigning reward credit across diffusion timesteps.
Aug 20, 2026 We released Stream4D, which uses 4D reconstruction rewards and motion priors to improve scene consistency and preserve natural motion in streaming video generation.

publications

Selected publications highlight a subset of my research, including Self-Forcing++.

  1. Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
    Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, and 3 more authors
    arXiv preprint arXiv:2610.02967, 2026
  2. ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
    Yihang Chen, Yuanhao Ban, and Cho-Jui Hsieh
    arXiv preprint arXiv:2609.35433, 2026
  3. SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference
    Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, and 7 more authors
    arXiv preprint arXiv:2609.28064, 2026
  4. ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement
    Yihang Chen*, Yuanhao Ban*, Kuei-Chun Kao, and 1 more author
    arXiv preprint arXiv:2609.13425, 2026
    * Equal contribution.
  5. Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
    Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, and 3 more authors
    arXiv preprint arXiv:2608.19556, 2026
  6. Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
    Yuanhao Ban, Tong Xie, Sohyun An, and 6 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2026
    Spotlight
  7. A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design
    Tong Xie, Yuanhao Ban, Yunqi Hong, and 3 more authors
    arXiv preprint arXiv:2606.11189, 2026
  8. One-Forcing: Towards Stable One-Step Autoregressive Video Generation
    Jiaqi Feng, Justin Cui, Yuanhao Ban, and 1 more author
    arXiv preprint arXiv:2605.23458, 2026
  9. AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
    Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban, and 1 more author
    arXiv preprint arXiv:2605.17602, 2026
  10. LoL: Longer than Longer, Scaling Video Generation to Hour
    Justin Cui, Jie Wu, Ming Li, and 6 more authors
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  11. Reward-Forcing: Autoregressive Video Generation with Reward Feedback
    Jingran Zhang, Ning Li, Yuanhao Ban, and 2 more authors
    In ICLR Second Workshop on World Models, 2026
  12. Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
    Justin Cui, Jie Wu, Ming Li, and 6 more authors
    In International Conference on Learning Representations (ICLR), 2026
  13. Fourier Neural Filter as Generic Vision Backbone
    Chenheng Xu, Yuanhao Ban, Dan Wu, and 3 more authors
    Preprint, 2026
  14. When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
    Tong Xie, Andrew Bai, Yuanhao Ban, and 3 more authors
    In International Conference on Machine Learning (ICML), 2026
  15. IRIS: Intrinsic Reward Image Synthesis
    Yihang Chen*, Yuanhao Ban*, Yunqi Hong, and 1 more author
    arXiv preprint arXiv:2509.25562, 2025
    * Equal contribution.
  16. The Crystal Ball Hypothesis in Diffusion Models: Anticipating Object Positions from Initial Noise
    Yuanhao Ban, Ruochen Wang, Tianyi Zhou, and 3 more authors
    In International Conference on Learning Representations (ICLR), 2025
  17. Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
    Yuanhao Ban, Ruochen Wang, Tianyi Zhou, and 3 more authors
    In European Conference on Computer Vision (ECCV), 2024
  18. Pre-trained Adversarial Perturbations
    Yuanhao Ban, and Yinpeng Dong
    In Advances in Neural Information Processing Systems (NeurIPS), 2022