Qi Xu 徐淇

I study how intelligent agents can form grounded representations of the physical world, reason about how it evolves, and use that understanding to act.

I am a PhD student at Westlake University, advised by Prof. Peidong Liu. My research connects 3D scene modeling, long-horizon video understanding, and multimodal post-training—building the perceptual and reasoning foundations for spatial world models that support embodied interaction.

Education & Experience

ByteDance logo

Bytedance AML

Research Intern at AML-Doubao Team

TikTok logo

ByteDance TikTok

Research Intern at TikTok Global Live

Mentor: Xiangtai Li

Wuhan University logo

Wuhan University

Bachelor's Degree

Research direction

Perceive. Understand. Interact.

My research aims to enable intelligent systems to understand, reason about, and interact with the physical world.

Observe

Perception

Sensing structure, motion, and change.

Interpret

Understanding

Forming grounded models of how the world works.

Engage

Interaction

Acting, observing feedback, and adapting.

News

Gave an invited talk on post-training multimodal large language models at Hangzhou Dianzi University.

VideoZeroBench was released with code and data for evidence-grounded long-video evaluation.

Any 3D Scene is Worth 1K Tokens was released on arXiv.

Two advised-student papers were accepted to ECCV 2026.

Watch, Remember, Reason was released on arXiv.

HiCI was accepted to ICML 2026.

Towards One-to-Many Temporal Grounding was accepted to ICML 2026.

SIU3R received a NeurIPS 2025 Spotlight.

Selected Publications

Qi Xu is underlined. * denotes equal contribution.

SIU3R teaser
Perception NeurIPS 2025 Spotlight

SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment

Qi Xu*, Dongxu Wei*, Lingzhe Zhao, Wenpu Li, Zhangchi Huang, Shunping Ji, Peidong Liu

Alignment-free framework for state-of-the-art 3D reconstruction and scene understanding from unposed images via a shared pixel-aligned 3D representation.

Perception arXiv 2026

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

Dongxu Wei*, Qi Xu*, Zhiqi Li, Hangning Zhou, Cong Qiu, Hailong Qin, Mu Yang, Zhaopeng Cui, Peidong Liu

Performs 3D scene generation directly within an implicit 3D latent space using 3D Diffusion Transformers.

HiCI teaser
Understanding ICML 2026

HiCI: Hierarchical Construction-Integration for Long-Context Attention

Xiangyu Zeng, Qi Xu, Yunke Wang, Chang Xu

Efficient long-context modeling with hierarchical attention, extending LLaMA-2 to 100K tokens with only 5.5% additional parameters.

Watch, Remember, Reason teaser
Understanding arXiv 2026

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Xiangtai Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang

A human-view survey organizing video MLLM research around watching, remembering, and reasoning for long, multimodal, and knowledge-intensive video understanding.

E-MoFlow teaser
Perception NeurIPS 2025

E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit Regularization

Wenpu Li, Bangyan Liao, Yi Zhou, Qi Xu, Pian Wan, Peidong Liu

Unsupervised framework for joint ego-motion and optical flow estimation using implicit neural representations and geometric constraints.

Publications with Advised Students

ECCV 2026

SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3DGS

Jin Liu, Haodong Li, Jiagang Chen, Dabin leng, Jiguang Li, Zhao Huang, Xiaoshuai Zhang, Qi Xu, Zhiwen Zheng, Xingru Huang

Reconstructs high-resolution 3D thermal scenes from low-resolution inputs via physics-informed degradation modeling.

ECCV 2026

Physically Grounded Dual-Opacity Gaussian Splatting for Joint RGB-TIR Reconstruction

Jin Liu, Dabin leng, Jiagang Chen, Haodong Li, Jiguang Li, Zhao Huang, Xiaoshuai Zhang, Zhiwen Zheng, Xingru Huang, Qi Xu

A unified RGB-TIR reconstruction framework resolving cross-spectral visibility conflicts through thermal field modeling.

Invited Talks

Post-Training of Multimodal Large Language Models

Data, Optimization and Evaluation · Hangzhou Dianzi University

Slides

Contact

I am open to research conversations and collaborations on spatial world models, embodied intelligence, 3D scene representations, and multimodal systems that connect perception with prediction and action.