Learning and intervention design for sequential decision-making
I study how structural assumptions and the design of training distributions and objectives constrain policy learning and behavioral generalization in sequential decision-making systems.
Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD
Across these projects, I study how limited training information constrains what a sequential decision-making system can learn, and how additional structure or actively selected experience can resolve the remaining ambiguity. When data are fixed, this motivates structural priors and invariances; when the data-generating process can be controlled, it motivates adapting environments, tasks, or experiments to the current learner.
01
Structure from limited data
What structure should learning assume?
Learning policies or objectives from demonstrations is underdetermined. My PhD work studied causal invariance and distributional geometry as inductive structure for learning rewards and policies that remain useful under distribution shift.
Imitation learningInverse RLCausal invarianceOptimal transport
02
Adaptive training environments
Where should the learner train next?
In deployed robot learning, the environment distribution becomes a controllable part of training. I use policy performance and failure signals to identify weak regions and direct training effort toward behavior that matters under terrain, sensing, and dynamics shift.
A broader direction I am exploring is how to select the next environment, task, or experiment based on what is currently uncertain and what matters for downstream decisions. This connects curriculum and environment design in RL with active data acquisition and goal-oriented experimental design.
Agent post-training. I am exploring how ideas from adaptive environment design transfer to task distributions, generated environments, and verifier-guided learning.
At ANYbotics, I develop and evaluate reinforcement-learning-based locomotion policies for industrial quadruped robots, with a focus on robustness, failure analysis, and sim-to-real iteration. My ETH Zürich PhD studied reinforcement learning from demonstrations: how reward and policy learning can exploit causal structure and distributional geometry when demonstrations provide incomplete information about the intended behavior.
Across both settings, I am increasingly interested in the active version of the same problem: how to choose the environments, tasks, or interventions from which a sequential system should learn next.
I build experiments with PyTorch, JAX, GPU simulation, reproducible evaluation, and HPC workflows.