Ivan Ovinnikov

Reinforcement Learning · Reward Learning · Adaptive Training

Ivan Ovinnikov

Sequential decision-making under distribution shift.

Research on reinforcement learning, learning from demonstrations, and adaptive training distributions for robust, data-efficient decision-making.

ANYbotics · Zürich, Switzerland

About

Researching how objectives and experience shape learning.

I am a machine learning researcher working on sequential decision-making under distribution shift. My research spans reinforcement learning, imitation and inverse reinforcement learning, reward learning, curriculum design, and risk-sensitive objectives.

A central theme of my work is that learning performance depends not only on the objective and model architecture, but also on the distribution of experience used for training. In interactive systems, this distribution is shaped by the current policy, the environment, the curriculum, and the failures selected for further learning. I develop methods for controlling these training and visitation distributions to improve robustness, data efficiency, and rare-event performance.

My work combines methodological research with evaluation in complex simulated and physical systems, including surgical digital twins and deployed quadruped locomotion.

Research areas

Adaptive training distributions

Curriculum, sampling, and replay methods that shape the learner’s experience using capability, uncertainty, learning progress, and failure risk.

Learning from demonstrations

Imitation, inverse reinforcement learning, and reward inference for problems where desired behaviour is easier to demonstrate than to specify directly.

Robust and risk-sensitive learning

Objectives and evaluation methods for distribution shift, reward misspecification, and low-probability but consequential failures.

Interactive model post-training

Applying reinforcement learning, reward learning, and adaptive experience generation to multimodal models and interactive agents.

Selected publications

Papers and preprints.

Work spanning reward learning, reinforcement learning, causal invariance, and optimal transport.

All publications