Ivan Ovinnikov

Ivan Ovinnikov

Reinforcement learning and learning from demonstrations under distribution shift

I work on reinforcement learning and learning from demonstrations, with a focus on robustness under distribution shift, reward learning, and curriculum learning.

Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD

Research interests

My research spans imitation and reward learning, curriculum learning and environment design, and robust robot learning. I am also interested in active task selection and experimental design as tools for selecting informative training tasks and environments.

  1. Imitation and reward learning

    How can demonstrations specify useful policies and objectives?

    My PhD work studied causal invariance and optimal-transport objectives for learning rewards and policies from limited demonstrations under distribution shift.

    Imitation learningInverse RLCausal invarianceOptimal transport
  2. Curriculum learning and environment design

    Where should the learner train next?

    I use policy performance and observed failure modes to define targeted evaluation and curriculum scenarios across terrain, sensing, and dynamics shifts.

    Curriculum learningEnvironment designReinforcement learning
  3. Active task selection and experimental design

    Emerging direction

    What task or experiment should the learner encounter next?

    I am exploring how active learning and sequential experimental design can select informative tasks and environments based on current uncertainty and downstream goals.

    Active learningTask selectionExperimental design

Publications

All publications

Background

At ANYbotics, I develop reinforcement-learning-based locomotion systems for industrial quadruped robots, spanning policy training, robustness and failure analysis, and sim-to-real deployment.

My ETH Zürich PhD focused on reinforcement learning from demonstrations, including causal invariance for reward learning and optimal-transport objectives for occupancy-matching imitation learning.

More broadly, I am interested in how learning systems acquire robust behavior under limited or imperfect supervision—particularly through active task selection, curriculum design, and experimental design: choosing the environments, tasks, and interventions that are most informative for improving a learner.

My work combines reinforcement learning, imitation and reward learning with large-scale experimentation using PyTorch, JAX, GPU simulation, and HPC infrastructure.

Experience and education