Ivan Ovinnikov

Ivan Ovinnikov

Reinforcement learning and learning from demonstrations under distribution shift

I work on reinforcement learning and learning from demonstrations, with a focus on robustness under distribution shift, reward learning, and curriculum learning.

Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD

Research interests

My research spans imitation and reward learning, curriculum learning and environment design, and robust robot learning. I am also interested in active task selection and experimental design as tools for selecting informative training tasks and environments.

  1. Imitation and reward learning

    How can demonstrations specify useful policies and objectives?

    My PhD work studied causal invariance and optimal-transport objectives for learning rewards and policies from limited demonstrations under distribution shift.

    Imitation learningInverse RLCausal invarianceOptimal transport
  2. Curriculum learning and environment design

    Where should the learner train next?

    I use policy performance and observed failure modes to define targeted evaluation and curriculum scenarios across terrain, sensing, and dynamics shifts.

    Curriculum learningEnvironment designReinforcement learning
  3. Active task selection and experimental design

    Emerging direction

    What task or experiment should the learner encounter next?

    I am exploring how active learning and sequential experimental design can select informative tasks and environments based on current uncertainty and downstream goals.

    Active learningTask selectionExperimental design

Publications

All publications

Background

At ANYbotics, I develop and evaluate reinforcement-learning-based locomotion policies for industrial quadruped robots, with a focus on robustness, failure analysis, and sim-to-real iteration. My ETH Zürich PhD studied reinforcement learning from demonstrations, including causal invariance for reward learning and sliced Wasserstein objectives for occupancy-matching imitation learning.

I am also interested in active task selection and experimental design: how to choose informative environments, tasks, or experiments for a learner.

I build experiments with PyTorch, JAX, GPU simulation, reproducible evaluation, and HPC workflows.

Experience and education