We're hiring. Full-time researchers / engineers and research interns in agentic models, omni-modal models and systems, and long-horizon intelligence. Get in touch.
NVIDIAResearch Scientist
Guo Chen
I am a Research Scientist at NVIDIA ADLR. I earned my Ph.D. in Computer Science from Nanjing University, where I was advised by Prof. Limin Wang and Prof. Tong Lu.
Previously, I contributed to the Intern series at Shanghai AI Lab and to several NVIDIA efforts, including EAGLE, Cosmos, and GR00T. My earlier research spanned long-form video understanding, egocentric perception, and multimodal dataset development. I now develop Nemotron foundation models, focusing on agentic and omni-modal systems for long-horizon intelligence.
Research
What I work on
My work connects foundation-model development with a broader question: how can models and systems reason, act, and learn over time?
Model development
NVIDIA Nemotron
At NVIDIA ADLR, I work on the Nemotron family across data, training, capability development, and evaluation, with recent work on Nemotron 3 Nano Omni.
Research
Current directions
Agentic Foundation Models
Developing models that plan, use tools, and learn through interaction to complete open-ended tasks.
Omni-Modal Models and Systems
Co-designing generalist models and specialized components to connect omni-modal perception and generation with decision-making and action.
Long-Horizon Intelligence
Building memory, reasoning, and continual adaptation across extended context and experience.
Publications
Selected Work
Representative papers. * denotes equal contribution. For the full list, see Google Scholar.
Foundation Models
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
NVIDIA Nemotron Team · Core contributor: Guo Chen
Eagle 2.5: Boosting long-context post-training for frontier vision-language models
Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen*, Shilong Liu, Shihao Wang, Vibashan VS, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Tong Lu, Bryan Catanzaro, Jan Kautz, Andrew Tao, Guilin Liu, Zhiding Yu
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Jilan Xu, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, Limin Wang
InternVL: Scaling Up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, Jifeng Dai
EAGLE Series
Frontier vision-language models with data-centric strategies
-
arXivContributorCurrent
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
-
NeurIPSLead authorCurrent
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
-
arXivCo-first authorCurrent
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
-
arXivSeries originCurrent
Eagle: Exploring the Design Space for Multimodal LLMs with Mixture of Encoders
Nemotron Series
Efficient open foundation models for language and multimodal intelligence
-
arXivCore contributorCurrent
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
-
arXivContributorCurrent
NVIDIA Nemotron Nano V2 VL
-
arXivContributorCurrent
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
Intern Series
General vision and video foundation models built through scaling
-
ECCVContributorCurrent
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
-
ICLR SpotlightContributorCurrent
InternVid: A Large-Scale Video-Text Dataset for Multimodal Understanding and Generation
-
arXivContributorCurrent
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
-
arXivLead authorCurrent
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Datasets & Benchmarks
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Lidong Lu*, Guo Chen*, Zhiqi Li, Yicheng Liu, Tong Lu
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
Guo Chen, Lidong Lu, Yicheng Liu, Liangrui Dong, Lidong Zou, Jixin Lv, Zhenquan Li, Xinyi Mao, Baoqi Pei, Shihao Wang, Zhiqi Li, Karan Sapra, Fuxiao Liu, Yin-Dong Zheng, Yifei Huang, Limin Wang, Zhiding Yu, Andrew Tao, Guilin Liu, Tong Lu
CG-Bench: Clue-grounded question answering benchmark for long video understanding
Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, Limin Wang
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View
Yifei Huang*, Guo Chen*, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, Yu Qiao
CG Series
Clue-grounded benchmarks for long-horizon multimodal understanding
-
CVPRCo-first authorCurrent
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
-
arXivLead authorCurrent
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
-
ICLRLead authorCurrent
CG-Bench: Clue-Grounded Question Answering Benchmark for Long Video Understanding
EgoExo Series
Cross-view learning between first- and third-person video
-
NeurIPSContributorCurrent
EgoExoBench: A Benchmark for First- and Third-Person View Video Understanding in MLLMs
-
ICLRContributorCurrent
EgoExo-Gen: Egocentric Video Prediction by Watching Exocentric Videos
-
CVPRCo-first authorCurrent
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-Centric View
Video Understanding
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Jiahao Wang, Zhe Chen, Zhiqi Li, Kunchang Li, Tong Lu, Limin Wang
Egocentric Object-Interaction Anticipation with Retentive and Predictive Learning
Guo Chen, Yifei Huang, Yin-Dong Zheng, Yicheng Liu, Jiahao Wang, Tong Lu
Memory-and-Anticipation Transformer for Online Action Understanding
Jiahao Wang*, Guo Chen*, Yifei Huang, Limin Wang, Tong Lu
BasicTAD: An Astounding RGB-Only Baseline for Temporal Action Detection
Min Yang*, Guo Chen*, Yin-Dong Zheng, Tong Lu, Limin Wang
DCAN: Improving Temporal Action Detection via Dual Context Aggregation
Guo Chen, Yin-Dong Zheng, Limin Wang, Tong Lu
EgoVideo Series
Egocentric models for perception, reasoning, and anticipation
-
NeurIPSContributorCurrent
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
-
arXivContributorCurrent
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
-
CVPRContributorCurrent
Retrieval-Augmented Egocentric Video Captioning (EgoInstructor)
News
News archive
Research updates, paper acceptances, benchmark releases, challenge results, and honors.
20265 updates
[career] I received my Ph.D. in Computer Science from Nanjing University and joined NVIDIA Applied Deep Learning Research (ADLR) as a full-time Research Scientist, driving the development of the Nemotron family of foundation models.
[publication & release] We release Nemotron 3 Nano Omni, an efficient, open multimodal model for text, images, video, and audio.
[publication & release] 2 CVPR papers are accepted: VideoITG (spotlight) and AV-Reasoner.
[publication & release] We release MM-Lifelong, a 181.1-hour benchmark and agentic baseline for multimodal lifelong understanding.
[publication & release] 2 IJCV papers are accepted: Video-Mamba-Suite and EgoExoSurvey.
202511 updates
[honors&awards] Nanjing University, PhD National Scholarship.
[publication & release] We release NVIDIA Nemotron Nano V2 VL for document understanding, long-video comprehension, and multimodal reasoning.
[publication & release] 3 NeurIPS papers are accepted: Eagle2.5, EgoExoBench, and EgoThinker.
[publication & release] Vinci has been accepted by IMWUT 2025.
[publication & release] We present CG-AV-Counting and AV-Reasoner. The data and code are released.
[honors&awards] EgoExoLearn has been selected for the EgoVis 2023/2024 Distinguished Paper Award.
[publication & release] We present Eagle2.5, boosting long-context vision-language capabilities.
[publication & release] Eagle2 has been adopted by NVIDIA GEAR Team to develop GR00T N1.
[publication & release] 3 ICLR papers are accepted: CG-Bench, EgoHOD, and EgoExo-Gen.
[publication & release] We present Eagle2, with model weights released on Hugging Face.
[publication & release] CG-Bench has been integrated into VLMEvalKit.
20247 updates
[publication & release] We present Vinci, a real-time embodied smart assistant based on egocentric VLM.
[publication & release] We present the clue-grounded long-video understanding benchmark CG-Bench and release basic evaluation code.
[honors&awards] Our team wins Top-1 rankings in 7 tracks of the 1st EgoVis ECCV2024 Challenge.
[publication & release] InternVideo2 has been accepted by ECCV 2024.
[publication & release] We present InternVideo2, with code integrated into InternVideo.
[publication & release] We present Video-Mamba-Suite and release the code.
[publication & release] 4 CVPR papers are accepted: InternVL, MVBench, EgoInstructor, and EgoExoLearn.
2021-20236 updates
[publication & release] We present InternVL and release the code.
[honors&awards] In the first Perception Test challenge, we obtain the best performance in Temporal Sound Localisation and runner-up in Temporal Action Localisation.
[publication & release] MAT is accepted by ICCV.
[honors&awards] WSDM Cup 2023 Toloka VQA Challenge, WSDM 2023, Top-1 Ranking.
[honors&awards] Our team wins Top-1 rankings in 7 tracks of the Ego4D ECCV2022 Challenge.
[publication & release] DCAN is accepted by AAAI.
20173 updates
[honors&awards] CCPC Final Contest, Bronze Medal.
[honors&awards] ACM-ICPC Asia Regional Contest, Silver Medal.
[honors&awards] CCPC Regional Contest, Bronze Medal.