Sign Up

1111 Engineering Drive, Boulder, CO 80309

View map

Abstract: Computer vision has made remarkable progress in recognition, restoration, generation, and embodied perception. Yet many real-world applications require more than a single model trained for a fixed benchmark: they require systems that can diagnose visual conditions, select appropriate tools, reason over intermediate results, and collaborate with other agents in open environments. In this talk, I will present Agentic Visual Intelligence, a system-level paradigm that integrates multimodal reasoning engines with specialized computer vision models for perception, restoration, assessment, planning, and communication.

The talk will cover three directions. First, I will discuss foundation models for human- and machine-centric perception, including efficient backbones, language-guided restoration, and image restoration methods that improve downstream perception under adverse weather and low-light conditions. Second, I will present agentic vision systems for digital media, focusing on 4KAgent, a universal image upscaling system, and Agent Banana, a similarly developed image editing agent. Third, I will extend this idea to physical systems through collaborative autonomous driving, including LangCoop and AirV2X, where language and multi-agent communication enable scalable, interpretable coordination among vehicles, infrastructure, and aerial agents. I will conclude with challenges in optimizing generative vision systems for quality, efficiency, and trustworthiness.

Bio: Dr. Zhengzhong Tu is an Assistant Professor of Computer Science and Engineering at Texas A&M University. He received his Ph.D. degree from the University of Texas at Austin in 2022, advised by Professor Alan Bovik. Dr. Tu has published more than 50 papers in IEEE TPAMI, IEEE TIP, NeurIPS, ICLR, CVPR, ECCV, ICCV, ICRA, IROS, CoRL, TMLR, among others. He is an Associate Editor of IEEE TIP, Action Editor of TMLR, and Co-Chair of the SOGAI special group in VQEG. He served as Area Chairs for multiple AI/ML/CV venues like CVPR/ICLR/NeurIPS/ICCV/ECCV/WACV. He is a recipient of the CVPR 2022 Best Paper Finalist, CVPR 2025 MEIS Workshop Best Paper Award, Google Research Scholar Award 2025, NVIDIA Research Grant Program Award 2025, TensorBlock Researcher Access Program Award, Amazon Research Award, CVPR 2025/2023 highlights, and ICLR 2023 highlight. He leads a team that wins the gold prize (5th/3450 globally) at the 3rd AI Mathematical Olympiad Challenge in 2026. His work has been featured in media outlets such as Google Research Annual Blog, Google I/O, Nature News, Forbes, and WIRED.

  • Dan Hawkley

1 person is interested in this event

User Activity

No recent activity