Reinforcement learning and learning from demonstrations under distribution shift
I work on reinforcement learning and learning from demonstrations, with a focus on robustness under distribution shift, reward learning, and curriculum learning.
Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD
My research spans imitation and reward learning, curriculum learning and environment design, and robust robot learning. I am also interested in active task selection and experimental design as tools for selecting informative training tasks and environments.
01
Imitation and reward learning
How can demonstrations specify useful policies and objectives?
My PhD work studied causal invariance and optimal-transport objectives for learning rewards and policies from limited demonstrations under distribution shift.
Imitation learningInverse RLCausal invarianceOptimal transport
02
Curriculum learning and environment design
Where should the learner train next?
I use policy performance and observed failure modes to define targeted evaluation and curriculum scenarios across terrain, sensing, and dynamics shifts.
What task or experiment should the learner encounter next?
I am exploring how active learning and sequential experimental design can select informative tasks and environments based on current uncertainty and downstream goals.
At ANYbotics, I develop reinforcement-learning-based locomotion systems for industrial quadruped robots, spanning policy training, robustness and failure analysis, and sim-to-real deployment.
My ETH Zürich PhD focused on reinforcement learning from demonstrations, including causal invariance for reward learning and optimal-transport objectives for occupancy-matching imitation learning.
More broadly, I am interested in how learning systems acquire robust behavior under limited or imperfect supervision—particularly through active task selection, curriculum design, and experimental design: choosing the environments, tasks, and interventions that are most informative for improving a learner.
My work combines reinforcement learning, imitation and reward learning with large-scale experimentation using PyTorch, JAX, GPU simulation, and HPC infrastructure.