Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Core Contributor · Xiaomi Robotics Team
Scaling robotic learning with over 100K hours of real-world manipulation trajectories.
Towards scalable and physically grounded robot intelligence.
Advanced Researcher, Robot Foundation Model Team · Xiaomi Robotics
I am an Advanced Researcher at the Robot Foundation Model Team, Xiaomi Robotics, where I was selected for the Top Talent Program. My research focuses on scalability in embodied learning, aiming to build robotic systems that generalize robustly across diverse tasks and environments.
Specifically, I am currently exploring two directions to advance scalable robot learning: (1) leveraging large-scale human data to unlock generalizable manipulation policies, and (2) developing unified visual-tactile representations to enable transferable physical grounding. I have the privilege of working closely with and learning from Wenxuan Song, Yifan Xie, and Zhuorong Li as we pursue these directions.
Previously, I was a Visiting Researcher at Harvard University and MIT-IBM Watson AI Lab, under the guidance of Prof. Chuang Gan and Prof. Yilun Du, and in close collaboration with Haoyu Zhen. Earlier, I was a Research Intern at Shanghai AI Laboratory. I received my M.S. in Electrical and Computer Engineering from Fudan University. Before that, I received a Bachelor of Engineering and a Bachelor of Management from Tianjin University.
Proud to be a core contributor to Xiaomi-Robotics-1, scaling VLA learning with over 100K hours of real-world trajectories. Visit the official website.
Representing Xiaomi Robotics at ICRA 2026, with visits to ETH Zurich and Xiaomi Europe R&D Center.
Action Images: End-to-End Policy Learning via Multiview Video Generation is now on arXiv.
Xiaomi-Robotics-0, an open-sourced VLA model with real-time execution.
TesserAct: Learning 4D Embodied World Models was accepted at ICCV 2025.
Paper, code, model, and website are publicly available.
Served as an on-site volunteer at COLING 2025 in Abu Dhabi.
A first-authored paper was accepted at COLING 2025.
Served as a reviewer for ACL ARR.
Presented MiniConGTS at EMNLP 2024 in Miami.
For a complete list of publications and projects, see my Google Scholar profile.
Core Contributor · Xiaomi Robotics Team
Scaling robotic learning with over 100K hours of real-world manipulation trajectories.
Haoyu Zhen, Zixian Gao, Qiao Sun, Yilin Zhao, Yuncong Yang, Yilun Du, Pengsheng Guo, Tsun-Hsuan Wang, Yi-Ling Qiao, Chuang Gan
Turning pixel-grounded multiview video generation into an end-to-end robot policy without a separate action module.
Core Contributor · Xiaomi Robotics Team
Delivering fast, smooth real-time VLA execution on real robots with a consumer-grade GPU.
Qiao Sun, Liujia Yang, Wei Tang, Wei Huang, Kaixin Xu, Yongchao Chen, Mingyu Liu, Jiange Yang, Haoyi Zhu, Yating Wang, Tong He, Yilun Chen, Xili Dai, Nanyang Ye, Qinying Gu
Composing short-horizon motion primitives into scalable, data-efficient world models for long-horizon robotic tasks.
Accepted at ICCV 2025
Haoyu Zhen*, Qiao Sun*, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, Chuang Gan
Predicting spatially and temporally consistent 4D scenes to support novel-view synthesis and stronger robot policies.
* Denotes Equal Contributions
Selected Reviewing Service: Reviewer for NeurIPS 2025-2026, ICML 2026, CVPR 2026, ICLR 2026, IEEE/ACM TASLP, and ACL ARR (2024-2025).
Conference Service: ACL SRW 2025, COLING 2025, and EMNLP 2024.