About the project
Objective
This project will develop efficient representation and learning methods for system dynamics that allow a robot to compensate for the inference delay of large robot action models, e.g., vision-language-action (VLA) models. The robot acquires and refines a predictive modelthrough its interaction with the environment. Through rolling out the learned system dynamics and generating the corresponding observations, the policy can be conditioned on ananticipated observation rather than an outdated one.
Background
VLA models allow robots to accomplish a wide range of tasks and have attracted intense interest in both academia and industry. Their capability comes at the expense of long inference time. By the moment an action or action chunk is computed, the robot state and the world have already evolved, and the robot acts on an observation of the past. Closing this gap requires anticipating the state at execution time and the observation that corresponds to it. For the robot’s own state, the prediction is straightforward. In dynamic scenes where objects roll, slide or deform in respond to robot actions, it becomes considerably harder.
About the Digital Futures Postdoc Fellow
Yunfan Gao completed her Ph.D. in 2026 at the University of Freiburg under the supervision of Prof. Moritz Diehl, and simultaneously, until September 2025 she was an industrial Ph.D.student at Bosch Center for AI, under the industrial supervision of Dr. Niels van Duijkeren. Previously, she obtained her master’s degree in Robotics, Systems and Control from ETH Zurich in 2022 and her bachelor’s degree in Electronic Engineering from Fudan University in 2019.
Main supervisor
Danica Kragic Jensfelt, KTH Royal Institute of Technology.
Co-supervisor
Florian Pokorny, KTH Royal Institute of Technology.

