Research
Research profile
I develop the NVIDIA Nemotron family, with research on foundation models and systems that reason, act, and learn over time.
NVIDIA Nemotron
Model development across data, training, capability development, and evaluation.
Agentic Foundation Models
Developing models that plan, use tools, and learn through interaction to complete open-ended tasks.
Omni-Modal Models and Systems
Co-designing generalist models and specialized components to connect omni-modal perception and generation with decision-making and action.
Long-Horizon Intelligence
Building memory, reasoning, and continual adaptation across extended context and experience.
Academic Service
Reviewer for NeurIPS, ICLR, ICML, CVPR, ICCV, ECCV, AAAI, ACM MM, IJCAI, TPAMI, IJCV, CVIU, and PR.
Featured Work
Representative projects that summarize my current research direction and past trajectory.

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
Guo Chen, Lidong Lu, Yicheng Liu, Liangrui Dong, Lidong Zou, Jixin Lv, Zhenquan Li, Xinyi Mao, Baoqi Pei, Shihao Wang, Zhiqi Li, Karan Sapra, Fuxiao Liu, Yin-Dong Zheng, Yifei Huang, Limin Wang, Zhiding Yu, Andrew Tao, Guilin Liu, Tong Lu
- MM-Lifelong introduces a 181.1-hour benchmark for multimodal lifelong understanding and ReMA, a recursive multimodal agent for sparse long-horizon visual reasoning.

Eagle 2.5: Boosting long-context post-training for frontier vision-language models
Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu.
- Long-context VLM post-training for videos and high-resolution visual inputs, contributing to NVIDIA’s broader embodied and reasoning model ecosystem.

CG-Bench: Clue-grounded question answering benchmark for long video understanding
Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, Limin Wang
- A long-video benchmark centered on clue-grounded reasoning, designed to expose gaps between open-source and commercial multimodal models.

InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, Jifeng Dai
- A general-purpose vision-language model that scales visual foundation models and aligns them with large language models for broad visual-linguistic tasks.

EgoVideo: Exploring egocentric foundation model and downstream adaptation
Guo Chen, Baoqi Pei, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, Yifei Huang, Yali Wang, Tong Lu, Limin Wang, Yu Qiao
- Foundation-model adaptation for egocentric video understanding, connected to champion solutions across Ego4D and EgoVis challenges.