From Perception to Simulation: The Emergence of World Models in Multi-modal Reasoning
ECCV 2026 Tutorial
Date and time: September 8, 2026, 13:30 (Malmo, Sweden time / CEST)
ECCV 2026, Malmo, Sweden
Introduction
World models are rapidly reshaping artificial intelligence, evolving from systems that passively perceive the world into engines capable of simulating, reasoning, and planning within it. This tutorial examines how recent advances in generative modeling, self-supervised learning, and multimodal architectures are enabling machines to move beyond recognition and prediction toward mental simulation, counterfactual reasoning, and decision making. We will explore the foundations of world models, approaches for learning dynamics from visual and multimodal data, and the integration of planning and reasoning. The tutorial highlights connections between video generation, diffusion models, discrete representations, and embodied AI, while addressing key challenges such as grounding, causality, physical consistency, and evaluation. Designed for researchers, practitioners, and students, this session provides both conceptual insights and practical perspectives on building AI systems that reason about environments rather than merely interpreting them. This ECCV 2026 tutorial will bring together perspectives on representation learning, generative simulation, multimodal reasoning, and embodied decision making for the next generation of world models.
World Model Tutorial Schedule
| Time | Session | Speaker |
|---|---|---|
| 13:30 - 13:40 |
Opening Remark: From Perception to Simulation: Setting the Stage for World Models [Abstract]
Abstract: This opening remark sets the stage for the tutorial by highlighting the shift from perception-oriented vision systems toward world models that can predict, simulate, and support action. We briefly discuss the key capabilities required for useful world models, including dynamics, causality, physical consistency, counterfactual reasoning, planning, and evaluation. The session will then introduce the tutorial structure and invited speakers, whose talks provide complementary perspectives on building and understanding multimodal world models.
|
|
| 13:40 - 14:20 |
Invited Talk: From Perception to Action: Unifying Modalities and Tasks for Building Digital Worlds [Abstract]
Abstract: Multimodal foundation models are evolving from isolated perception systems into unified architectures that understand and generate content across images, video, audio, and language. However, unifying modalities and tasks does not automatically make model outputs actionable. A model may understand a design request and generate a visually compelling interface, while still returning only flat pixels: objects cannot be selected, assets cannot be independently edited, and the result cannot be directly executed or verified by an agent. In this talk, we present a three-stage progression from perception to action. First, modality unification establishes a shared perceptual substrate while preserving modality-specific expertise through cross-modal routing and collaborative training. Second, task unification connects understanding with generation, enabling capabilities such as generative segmentation, fine-grained editing, and multimodal content creation. Third, we introduce an actionable interface stack that extends unified models into real digital workflows: Generative UI supports visual exploration, editable visual structures expose semantic RGBA layers and compositional relationships, and Design-grounded Coding translates visual intent, assets, and design constraints into executable interfaces. Using the Ming series as connected system examples, we discuss how intent can be preserved from multimodal perception to visual proposals, structured editing, and code execution. We conclude with evaluation challenges, failure propagation across the pipeline, and open questions for multimodal systems that move beyond understanding and generation toward creating, editing, and acting within digital worlds.
|
Dandan Zheng |
| 14:20 - 15:00 |
Invited Talk: Reasoning Beyond Perception: Toward Reliable Multimodal World Models [Abstract]
Abstract: Multimodal world models are moving beyond passive perception toward systems that can reason about, simulate, and interact with dynamic environments. However, as models increasingly rely on visual evidence, temporal structure, prior knowledge, and multi-step inference, reliability becomes a central challenge. This talk presents a reliability-oriented perspective on multimodal reasoning and world models. I will discuss several key questions: whether the model perceives the required evidence, grounds its reasoning in the correct visual cues, preserves structural and temporal consistency, maintains appropriate knowledge as the environment evolves, and verifies its conclusions before acting. Recent advances in grounded reasoning, continual adaptation, knowledge control, and multimodal unlearning will be discussed, together with examples from our recent work. The broader goal is to move from models that merely generate plausible predictions to systems that maintain coherent, controllable, and trustworthy internal representations of the world.
|
|
| 15:00 - 15:40 |
Invited Talk: Generative Neural Priors for 3D Motion Perception [Abstract]
Abstract: Generative modeling has rapidly evolved from a tool for creating synthetic data to a foundational paradigm for solving diverse real-world problems. In this talk, I will present a series of works where generative priors are designed not merely for pose generation, but as robust, geometry-aware motion models that enable reconstruction, recovery, and reasoning under uncertainty. Starting from pose generation via Neural Riemannian Distance Fields (NRDF), I will show how articulated pose distributions (of humans or hands) can be faithfully captured on product manifolds of rotations, yielding powerful priors for inverse kinematics, image-based pose estimation, and denoising. I will then transition to motion models such as HuMoR, which introduces an expressive, dynamics-aware generative model of human motion via conditional VAEs, enabling robust 3D motion recovery even under noise, occlusions, or missing observations. Finally, Neural Riemannian Motion Fields (NRMF) extend these ideas to higher-order motion dynamics, explicitly modeling pose, velocity, and acceleration on motion manifolds through neural distance fields. Together, these models demonstrate how generative formulations serve as flexible priors that power tasks as diverse as denoising, temporal interpolation, spatiotemporal infilling, and recovery of motion from highly incomplete or corrupted data, as evidenced by Dyn-HaMR. By embedding geometry, dynamics, and stochasticity into generative frameworks, the talk will move towards a unified view where generation, perception, and recovery are all facets of the same modeling principle.
|
|
| 15:40 - 16:20 |
Invited Talk: Disentangling Appearance, Geometry, and Motion for Video Modeling and Generation [Abstract]
Abstract: Video is beyond a sequence of images. The unique spatio-temporal pattern makes video modeling and generation different from representing a single image. Instead of treating videos as a 3D volume of pixels, we propose to learn disentangled representation that separates the appearance, motion, geometry, and semantic information of a video. Utilizing these disentangled representations in an integrated manner can enhance the quality, realism, and interpretability of generated videos and improve the robustness of video modeling. In this talk, we present methods that explicitly utilize disentangled representations, including appearance, motion, geometry, and semantics, to improve performance in video generation and modeling tasks.
|
|
| 16:20 - 17:00 |
Invited Talk: From Imagined Futures to Structured Representations: World Models for Robots [Abstract]
Abstract: World models - internal representations that let a robot predict, plan, and act - are central to progress across robotics, yet turning general-purpose video-generation and vision-language foundation models into something a controller can use remains an open problem. This tutorial presents cases of world models for robots across different robot embodiments, showing how imagined futures and generalist visual-language understanding can be converted into structured, action-ready representations. Through case studies and open discussion, I hope attendees will leave with practical thoughts on building world models across robot embodiments.
|
Organizers
Yujun Cai
Staff Research Scientist at Ant Group, Lecturer at University of Queensland
Yujun Cai is a Lecturer in the University of Queensland in Australia and a Staff Research Scientist in Ant Group USA. Before that, she was a Senior Research Scientist in Meta Reality Lab. Her research lies in multi-modal human perception, vision-language models, and natural language processing. She obtained her PhD. degree from Nanyang Technological University in Singapore.
Siyuan Yang
Wallenberg-NTU Presidential Postdoctoral Fellow at KTH Royal Institute of Technology
Siyuan Yang is a Wallenberg-NTU Presidential Postdoctoral Fellow at the Division of Robotics, Perception and Learning, KTH Royal Institute of Technology. He obtained his Ph.D. from Nanyang Technological University in 2024. His research focuses on video and skeleton-based action recognition, 3D human motion understanding, and human pose estimation.
Jun Liu
Professor at Lancaster University
Dr. Jun Liu is Professor and Chair in Digital Health at the School of Computing and Communications, Lancaster University. He earned his PhD from Nanyang Technological University in 2019, subsequently serving as faculty at Singapore University of Technology and Design from 2019 to 2024. Prior to his academic career, he worked at Tencent from 2014 to 2015.
Yiwei Wang
Assistant Professor at University of California, Merced
Dr. Yiwei Wang was an Applied Scientist in Amazon (Seattle) in 2023 and a Postdoc in UCLA NLP Group in 2024. He obtained his Ph.D. degree from National University of Singapore in 2023. Currently, he leads the UC Merced NLP Lab, where his team explores cutting-edge approaches to diffusion llms, reasoning multi-modal llms, and their applications in medicine, advertising, risk detection, signal processing, etc.
Junsong Yuan
Professor and Interim Chair of Computer Science and Engineering, State University of New York at Buffalo (SUNY)
Dr. Junsong Yuan is Professor and Interim Chair of Computer Science and Engineering, State University of New York at Buffalo (SUNY). Before joining SUNY Buffalo, he was Associate Professor at Nanyang Technological University (NTU), Singapore. He obtained his Ph.D. from Northwestern University, M.Eng. from National University of Singapore, and B.Eng. from Huazhong University of Science Technology. He received Chancellor's Award for Excellence in Scholarship and Creative Activities from SUNY, Faculty Innovation Award from SONY, Nanyang Assistant Professorship from NTU, Outstanding EECS Ph.D. Thesis award from Northwestern University, and Distinguished Lecturer from IEEE Signal Processing Society. He is currently Editor-in-Chief of Journal of Visual Communication and Image Representation (JVCI) and Associate Editor of IEEE Trans. on Pattern Analysis and Machine Intelligence (T-PAMI) etc. He also serves as General/Program Co-chair of ICME/ICASSP/ICIP and Area Chair of a few other conferences. He is a Fellow of IEEE and IAPR, and Distinguished Member of ACM.
Invited Speakers
Dandan Zheng
Ant Group
Dandan Zheng received the M.S. degree in Computer Applications from Beijing University of Posts and Telecommunications, Beijing, China, in 2007, and is currently pursuing the Ph.D. degree in Cybersecurity at Huazhong University of Science and Technology, Wuhan, China. From 2007 to 2014, she was a Research Engineer at IBM China Software Development Lab, focusing on text mining and natural language processing. Since 2014, she has been with Ant Group, where she conducted research on face recognition and visual security from 2014 to 2022, and has been working on visual generation and unified multimodal modeling since 2022. Her research interests include multimodal understanding and generation, diffusion models, autoregressive visual generation, and vision-language unification.
Xiatian Zhu
Reader (Associate Professor) at University of Surrey
Dr. Xiatian Zhu is a Reader (Associate Professor) at the Surrey Institute for People-Centred Artificial Intelligence and the Centre for Vision, Speech, and Signal Processing (CVSSP) at the University of Surrey in Guildford, UK. He leads the Universal Perception (UP) Lab, dedicated to advancing physical spacetime AI - building world models that perceive, generate, and reason about the physical world across space, time, and physical law - with applications spanning creative media, fashion, healthcare, climate science, and cybersecurity. Dr. Zhu earned his PhD from Queen Mary University of London and received the 2016 Sullivan Doctoral Thesis Prize from the British Machine Vision Association. His contributions include the development and commercialization of multi-camera object association systems for industry. During his time as a research scientist at the Samsung AI Centre in Cambridge, Dr. Zhu pioneered sustainable AI algorithms for understanding visual content in images and videos. His work has garnered several best paper awards, and he has been recognized as one of the UK's and the world's best rising stars in science. Dr. Zhu's extensive research output includes over 200 articles in top-tier conferences and journals, with more than 27,000 citations and an H-index of 67. Additionally, Dr. Zhu holds five US patents in the fields of computer vision and AI. He is an IEEE Senior Member. He serves as an Associate Editor of the IEEE Transactions on Multimedia (TMM) and an Action Editor for Transactions on Machine Learning Research (TMLR). He also serves/served as an Area Chair for top conferences including CVPR, ICCV, NeurIPS, and ICLR.
Yirui Wu
Associate Professor at United Arab Emirates University
Yirui Wu is an Associate Professor at United Arab Emirates University. His research interests include trustworthy multimodal AI, continual and adaptive learning, multimodal reasoning, and edge intelligence. His recent work focuses on reliable representation learning, knowledge adaptation and forgetting, and multimodal unlearning, with the broader goal of building intelligent systems that can continuously learn, reason, and adapt while maintaining reliability.
Tolga Birdal
Associate Professor at Imperial College London
Tolga Birdal is a UKRI Future Leaders Fellow and an associate professor in the Department of Computing of Imperial College London. Previously, he was a Postdoctoral Research Fellow at Stanford University within the Geometric Computing Group of Prof. Leonidas Guibas. Tolga has defended his masters and Ph.D. theses at the Computer Vision Group under Chair for Computer Aided Medical Procedures, Technical University of Munich led by Prof. Nassir Navab. He was also a Doktorand at Siemens AG under supervision of Dr. Slobodan Ilic working on "Geometric Methods for 3D Reconstruction from Large Point Clouds". His current foci of interest involve topological / geometric machine learning and 3D computer vision. More theoretical work is aimed at investigating and interrogating limits in geometric computing and non-Euclidean inference as well as principles of deep learning. Tolga is an AC for major vision conferences such as CVPR, ICCV, ECCV and 3DV and he has several publications at the well-respected venues such as NeurIPS, CVPR, ICCV, ECCV, ICLR, ICML, T-PAMI, IJCV, ICRA, IROS, ICASSP and 3DV. Aside from his academic life, Tolga has co-founded multiple companies including Befunky, a widely used web-based image editing platform.