From Perception to Simulation: The Emergence of World Models in Multi-modal Reasoning
ECCV 2026 Tutorial
Date and time: TBD
ECCV 2026, Malmo, Sweden
Introduction
World models are rapidly reshaping artificial intelligence, evolving from systems that passively perceive the world into engines capable of simulating, reasoning, and planning within it. This tutorial examines how recent advances in generative modeling, self-supervised learning, and multimodal architectures are enabling machines to move beyond recognition and prediction toward mental simulation, counterfactual reasoning, and decision making. We will explore the foundations of world models, approaches for learning dynamics from visual and multimodal data, and the integration of planning and reasoning. The tutorial highlights connections between video generation, diffusion models, discrete representations, and embodied AI, while addressing key challenges such as grounding, causality, physical consistency, and evaluation. Designed for researchers, practitioners, and students, this session provides both conceptual insights and practical perspectives on building AI systems that reason about environments rather than merely interpreting them. This ECCV 2026 tutorial will bring together perspectives on representation learning, generative simulation, multimodal reasoning, and embodied decision making for the next generation of world models.
World Model Tutorial Schedule
| Time | Session |
|---|---|
| TBD |
Detailed ECCV 2026 tutorial schedule will be announced. |
Organizers
Yujun Cai
Staff Research Scientist at Ant Group, Lecturer at University of Queensland
Yujun Cai is a Lecturer in the University of Queensland in Australia and a Staff Research Scientist in Ant Group USA. Before that, she was a Senior Research Scientist in Meta Reality Lab. Her research lies in multi-modal human perception, vision-language models, and natural language processing. She obtained her PhD. degree from Nanyang Technological University in Singapore.
Siyuan Yang
Wallenberg-NTU Presidential Postdoctoral Fellow at KTH Royal Institute of Technology
Siyuan Yang is a Wallenberg-NTU Presidential Postdoctoral Fellow at the Division of Robotics, Perception and Learning, KTH Royal Institute of Technology. He obtained his Ph.D. from Nanyang Technological University in 2024. His research focuses on video and skeleton-based action recognition, 3D human motion understanding, and human pose estimation.
Jun Liu
Professor at Lancaster University
Dr. Jun Liu is Professor and Chair in Digital Health at the School of Computing and Communications, Lancaster University. He earned his PhD from Nanyang Technological University in 2019, subsequently serving as faculty at Singapore University of Technology and Design from 2019 to 2024. Prior to his academic career, he worked at Tencent from 2014 to 2015.
Yiwei Wang
Assistant Professor at University of California, Merced
Dr. Yiwei Wang was an Applied Scientist in Amazon (Seattle) in 2023 and a Postdoc in UCLA NLP Group in 2024. He obtained his Ph.D. degree from National University of Singapore in 2023. Currently, he leads the UC Merced NLP Lab, where his team explores cutting-edge approaches to diffusion llms, reasoning multi-modal llms, and their applications in medicine, advertising, risk detection, signal processing, etc.
Junsong Yuan
Professor at University at Buffalo, SUNY
Dr. Junsong Yuan is a Professor at University at Buffalo, specializing in computer vision, pattern recognition, and multimedia analysis. His research encompasses video understanding, human activity recognition, and multimodal learning with applications in surveillance, healthcare, and autonomous systems.