From Perception to Simulation: The Emergence of World Models in Multi-modal Reasoning

ECCV 2026 Tutorial

Date and time: TBD
ECCV 2026, Malmo, Sweden

Introduction

World models are rapidly reshaping artificial intelligence, evolving from systems that passively perceive the world into engines capable of simulating, reasoning, and planning within it. This tutorial examines how recent advances in generative modeling, self-supervised learning, and multimodal architectures are enabling machines to move beyond recognition and prediction toward mental simulation, counterfactual reasoning, and decision making. We will explore the foundations of world models, approaches for learning dynamics from visual and multimodal data, and the integration of planning and reasoning. The tutorial highlights connections between video generation, diffusion models, discrete representations, and embodied AI, while addressing key challenges such as grounding, causality, physical consistency, and evaluation. Designed for researchers, practitioners, and students, this session provides both conceptual insights and practical perspectives on building AI systems that reason about environments rather than merely interpreting them. This ECCV 2026 tutorial will bring together perspectives on representation learning, generative simulation, multimodal reasoning, and embodied decision making for the next generation of world models.

World Model Tutorial Schedule

Time Session
TBD

Detailed ECCV 2026 tutorial schedule will be announced.

Organizers

Yujun Cai
Yujun Cai

Staff Research Scientist at Ant Group, Lecturer at University of Queensland

Yujun Cai is a Lecturer in the University of Queensland in Australia and a Staff Research Scientist in Ant Group USA. Before that, she was a Senior Research Scientist in Meta Reality Lab. Her research lies in multi-modal human perception, vision-language models, and natural language processing. She obtained her PhD. degree from Nanyang Technological University in Singapore.

Siyuan Yang
Siyuan Yang

Wallenberg-NTU Presidential Postdoctoral Fellow at KTH Royal Institute of Technology

Siyuan Yang is a Wallenberg-NTU Presidential Postdoctoral Fellow at the Division of Robotics, Perception and Learning, KTH Royal Institute of Technology. He obtained his Ph.D. from Nanyang Technological University in 2024. His research focuses on video and skeleton-based action recognition, 3D human motion understanding, and human pose estimation.

Jun Liu
Jun Liu

Professor at Lancaster University

Dr. Jun Liu is Professor and Chair in Digital Health at the School of Computing and Communications, Lancaster University. He earned his PhD from Nanyang Technological University in 2019, subsequently serving as faculty at Singapore University of Technology and Design from 2019 to 2024. Prior to his academic career, he worked at Tencent from 2014 to 2015.

Yiwei Wang
Yiwei Wang

Assistant Professor at University of California, Merced

Dr. Yiwei Wang was an Applied Scientist in Amazon (Seattle) in 2023 and a Postdoc in UCLA NLP Group in 2024. He obtained his Ph.D. degree from National University of Singapore in 2023. Currently, he leads the UC Merced NLP Lab, where his team explores cutting-edge approaches to diffusion llms, reasoning multi-modal llms, and their applications in medicine, advertising, risk detection, signal processing, etc.

Junsong Yuan
Junsong Yuan

Professor at University at Buffalo, SUNY

Dr. Junsong Yuan is a Professor at University at Buffalo, specializing in computer vision, pattern recognition, and multimedia analysis. His research encompasses video understanding, human activity recognition, and multimodal learning with applications in surveillance, healthcare, and autonomous systems.

Invited Speakers