Ivan Ovinnikov

Ivan Ovinnikov

Reinforcement learning for reliable agents

I develop reinforcement-learning methods for robust behavior under imperfect objectives and distribution shift—from deployed robotics to foundation-model post-training.

Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD

Research themes

I study how supervision, objectives, and training distributions determine behavior after deployment. The work connects reward and imitation learning, failure-focused robot training, and diagnostics for RL post-training.

Robust reinforcement learning

Methods and evaluation for distribution shift, rare failures, curriculum design, sim-to-real transfer, and policies deployed on physical robots.

Learning from imperfect objectives

Reward and imitation-learning methods that address distribution matching, misspecification, and behavioral identifiability.

RL post-training and agent learning

Diagnostics for credit assignment, reward and verifier structure, optimization dynamics, and multimodal or embodied agents.

Publications

All publications

Background

At ANYbotics, I develop and evaluate locomotion policies for industrial quadruped robots. My ETH Zürich PhD focused on reinforcement learning from demonstrations, reward inference, and distribution shift in surgical digital twins.

I build experiments with PyTorch, JAX, GPU simulation, reproducible evaluation, and HPC workflows.

Experience and education