Ivan Ovinnikov

Ivan Ovinnikov

Reinforcement learning for reliable agents

I develop reinforcement-learning methods for robust behavior under imperfect objectives and distribution shift—from deployed robotics to foundation-model post-training.

Reinforcement Learning Engineer, ANYbotics · ETH Zürich PhD

Research themes

I study how supervision, objectives, and training distributions determine behavior after deployment. The work connects reward and imitation learning, failure-focused robot training, and diagnostics for RL post-training.

Robust reinforcement learning

Methods and evaluation for distribution shift, rare failures, curriculum design, sim-to-real transfer, and policies deployed on physical robots.

Learning from imperfect objectives

Reward and imitation-learning methods that address distribution matching, misspecification, and behavioral identifiability.

RL post-training and agent learning

Diagnostics for credit assignment, reward and verifier structure, optimization dynamics, and multimodal or embodied agents.

Publications

All publications

Background

Methods from learning signal to deployment

At ANYbotics, I develop and evaluate locomotion policies for industrial quadruped robots. My ETH Zürich PhD focused on reinforcement learning from demonstrations, reward inference, and distribution shift in surgical digital twins.

I build experiments with PyTorch, JAX, GPU simulation, reproducible evaluation, and HPC workflows.

Experience and education