Westlake University
PhD Student in Artificial Intelligence, supervised by Prof. Peidong Liu.
Research: Embodied AI, Multimodal Large Language Models.
PhD Student in Artificial Intelligence
I am a PhD student at Westlake University, advised by Prof. Peidong Liu. My research focuses on multimodal large language models, especially video reasoning, spatio-temporal grounding, and post-training.
I am interested in building grounded multimodal systems that can find and verify visual evidence, reason over long videos, and understand the spatial structure of physical environments. My earlier work spans 3D reconstruction and scene understanding. I received my master's and bachelor's degrees from Wuhan University.
PhD Student in Artificial Intelligence, supervised by Prof. Peidong Liu.
Research: Embodied AI, Multimodal Large Language Models.
Research Intern, LLM Strategy Department. Mentored by Xiangtai Li.
Master in Pattern Recognition and Intelligent Systems, supervised by Prof. Shunping Ji.
Research: Computer Vision, 3D Reconstruction, Multimodal Learning.
B.S. in Spatial Information and Digital Technology.
VideoZeroBench was released with code and data for evidence-grounded long-video evaluation.
Any 3D Scene is Worth 1K Tokens was released on arXiv.
Two advised-student papers were accepted to ECCV 2026.
Watch, Remember, Reason was released on arXiv.
HiCI was accepted to ICML 2026.
Towards One-to-Many Temporal Grounding was accepted to ICML 2026.
SIU3R received a NeurIPS 2025 Spotlight.
SFT and reinforcement learning for grounded multimodal reasoning, with verifiable task-specific rewards.
Long-video understanding, temporal localization, evidence verification, benchmark construction, and video agents.
Scene reconstruction, understanding, generation, and spatial representations grounded in physical environments.
Qi Xu is underlined. * denotes equal contribution.
NeurIPS 2025 Spotlight
Alignment-free framework for state-of-the-art 3D reconstruction and scene understanding from unposed images via a shared pixel-aligned 3D representation.
My contributions: Project lead, core idea contributor, and primary code implementer; led benchmarking, ablations, and paper writing.
ICML 2026
Introduces the first systematic solution for one-to-many temporal grounding with a 56K-sample dataset and SFT/RL training using temporal and CoT-based caption rewards.
My contributions: Initiated the project and contributed the core idea; led data construction, SFT/RL pipelines, reward design, model training, benchmarking, and paper writing.
arXiv 2026
Performs 3D scene generation directly within an implicit 3D latent space using 3D Diffusion Transformers.
My contributions: Built large-scale datasets, conducted ablations and evaluation, and contributed to paper writing and revision.
arXiv 2026
A human-view survey organizing video MLLM research around watching, remembering, and reasoning for long, multimodal, and knowledge-intensive video understanding.
My contributions: Organized literature for the memory section, created figures and tables, and contributed to writing and revision.
arXiv 2026
A hierarchical long-video QA benchmark that verifies answers with temporal intervals and spatial boxes under a five-level evidence protocol.
My contributions: Co-designed the five-level evaluation protocol, implemented the annotation tool and guidelines, and constructed about 200 complete QA-evidence samples.
NeurIPS 2025
Unsupervised framework for joint ego-motion and optical flow estimation using implicit neural representations and geometric constraints.
My contributions: Contributed to method design and paper revision.
ECCV 2026
Reconstructs high-resolution 3D thermal scenes from low-resolution inputs via physics-informed degradation modeling.
ECCV 2026
A unified RGB-TIR reconstruction framework resolving cross-spectral visibility conflicts through thermal field modeling.
I am open to research conversations and collaborations on embodied AI, multimodal reasoning, 3D Vision, and vision-language systems that connect perception with action.