Reinforcement learning and learning from demonstrations under distribution shift
I work on reinforcement learning and learning from demonstrations, with a focus on robustness under distribution shift, reward learning, and curriculum learning.
ETH Zürich PhD · Reinforcement learning and sequential decision-making
My research spans imitation and reward learning, curriculum learning, environment design, and robust learning under distribution shift. I am also interested in active task selection and experimental design as tools for selecting informative training tasks and environments.
01
Imitation and reward learning
How can demonstrations specify useful policies and objectives?
My PhD work studied causal invariance and optimal-transport objectives for learning rewards and policies from limited demonstrations under distribution shift.
Imitation learningInverse RLCausal invarianceOptimal transport
02
Curriculum learning and environment design
Where should the learner train next?
I use policy performance and observed failure modes to define targeted evaluation and curriculum scenarios across terrain, sensing, and dynamics shifts.
What task or experiment should the learner encounter next?
I am exploring how active learning and sequential experimental design can select informative tasks and environments based on current uncertainty and downstream goals.
At ANYbotics, my recent research has focused on how training experience should be allocated in reinforcement learning. I evaluate these methods at scale in simulation and on physical robotic systems. This work is developed and validated in industrial quadruped locomotion, spanning robustness and failure analysis, production training, and sim-to-real deployment.
My ETH Zürich PhD focused on reinforcement learning from demonstrations, including causal invariance for reward learning and optimal-transport objectives for occupancy-matching imitation learning.
More broadly, I am interested in how learning systems acquire robust behavior under limited or imperfect supervision—particularly through active task selection, curriculum design, and experimental design: choosing the environments, tasks, and interventions that are most informative for improving a learner.
My work combines reinforcement learning, imitation and reward learning with large-scale experimentation using PyTorch, JAX, GPU simulation, and HPC infrastructure.