Publications

Selected publications

Selected representative publications. * denotes equal contribution. For a full and up-to-date list, please see my Google Scholar.

arXiv 2026
MM-Lifelong dataset comparison

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

Guo Chen, Lidong Lu, Yicheng Liu, Liangrui Dong, Lidong Zou, Jixin Lv, Zhenquan Li, Xinyi Mao, Baoqi Pei, Shihao Wang, Zhiqi Li, Karan Sapra, Fuxiao Liu, Yin-Dong Zheng, Yifei Huang, Limin Wang, Zhiding Yu, Andrew Tao, Guilin Liu, Tong Lu

PDF dataset

  • MM-Lifelong introduces a 181.1-hour benchmark for multimodal lifelong understanding and ReMA, a recursive multimodal agent for sparse long-horizon visual reasoning.
NeurIPS 2025
Eagle 2.5

Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu.

PDF code

  • Eagle 2.5 is a generalist long-context VLM for videos and high-resolution images, combining post-training data strategy with long-video capability.
arXiv
Eagle 2

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Zhiqi Li, Guo Chen*, Shilong Liu, Shihao Wang, Vibashan VS, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, Nadine Chang, Karan Sapra, Amala Sanjay Deshmukh, Tuomas Rintamaki, Matthieu Le, Ilia Karmanov, Lukas Voegtle, Philipp Fischer, De-An Huang, Timo Roman, Tong Lu, Jose M Alvarez, Bryan Catanzaro, Jan Kautz, Andrew Tao, Guilin Liu, Zhiding Yu

PDF code

  • This work focuses on developing open-source vision-language models by emphasizing data strategy in post-training.
ICLR 2025
CG-Bench

CG-Bench: Clue-grounded question answering benchmark for long video understanding

Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, Limin Wang

PDF code

  • CG-Bench tests multimodal models on long videos with clue-based QA and exposes gaps in long-video reasoning.
CVPR 2024
EgoExoLearn

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

Yifei Huang*, Guo Chen*, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, Yu Qiao

PDF code

  • EgoExoLearn is a dataset with egocentric and demonstration videos, gaze data, and multimodal annotations for cross-view learning.
ICCV 2023
MAT

Memory-and-Anticipation Transformer for Online Action Understanding

Jiahao Wang*, Guo Chen*, Yifei Huang, Limin Wang, Tong Lu

Homepage

  • This work presents a memory-anticipation-based method for online action understanding.