Ivan Ovinnikov

Ivan Ovinnikov

Learning and intervention design for sequential decision-making

I study how structural assumptions and the design of training distributions and objectives constrain policy learning and behavioral generalization in sequential decision-making systems.

Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD

Research direction

Across these projects, I study how limited training information constrains what a sequential decision-making system can learn, and how additional structure or actively selected experience can resolve the remaining ambiguity. When data are fixed, this motivates structural priors and invariances; when the data-generating process can be controlled, it motivates adapting environments, tasks, or experiments to the current learner.

  1. Structure from limited data

    What structure should learning assume?

    Learning policies or objectives from demonstrations is underdetermined. My PhD work studied causal invariance and distributional geometry as inductive structure for learning rewards and policies that remain useful under distribution shift.

    Imitation learningInverse RLCausal invarianceOptimal transport
  2. Adaptive training environments

    Where should the learner train next?

    In deployed robot learning, the environment distribution becomes a controllable part of training. I use policy performance and failure signals to identify weak regions and direct training effort toward behavior that matters under terrain, sensing, and dynamics shift.

    Reinforcement learningCurriculum learningEnvironment designSim-to-real
  3. Adaptive intervention design

    Emerging direction

    What intervention should the system run next?

    A broader direction I am exploring is how to select the next environment, task, or experiment based on what is currently uncertain and what matters for downstream decisions. This connects curriculum and environment design in RL with active data acquisition and goal-oriented experimental design.

    Agent post-training. I am exploring how ideas from adaptive environment design transfer to task distributions, generated environments, and verifier-guided learning.

    Adaptive acquisitionTask selectionExperimental design

Publications

All publications

Background

At ANYbotics, I develop and evaluate reinforcement-learning-based locomotion policies for industrial quadruped robots, with a focus on robustness, failure analysis, and sim-to-real iteration. My ETH Zürich PhD studied reinforcement learning from demonstrations: how reward and policy learning can exploit causal structure and distributional geometry when demonstrations provide incomplete information about the intended behavior.

Across both settings, I am increasingly interested in the active version of the same problem: how to choose the environments, tasks, or interventions from which a sequential system should learn next.

I build experiments with PyTorch, JAX, GPU simulation, reproducible evaluation, and HPC workflows.

Experience and education