Adaptive training distributions
Curriculum, sampling, and replay methods that shape the learner’s experience using capability, uncertainty, learning progress, and failure risk.
Reinforcement Learning · Reward Learning · Adaptive Training
Sequential decision-making under distribution shift.
Research on reinforcement learning, learning from demonstrations, and adaptive training distributions for robust, data-efficient decision-making.
ANYbotics · Zürich, Switzerland
About
I am a machine learning researcher working on sequential decision-making under distribution shift. My research spans reinforcement learning, imitation and inverse reinforcement learning, reward learning, curriculum design, and risk-sensitive objectives.
A central theme of my work is that learning performance depends not only on the objective and model architecture, but also on the distribution of experience used for training. In interactive systems, this distribution is shaped by the current policy, the environment, the curriculum, and the failures selected for further learning. I develop methods for controlling these training and visitation distributions to improve robustness, data efficiency, and rare-event performance.
My work combines methodological research with evaluation in complex simulated and physical systems, including surgical digital twins and deployed quadruped locomotion.
Research areas
Curriculum, sampling, and replay methods that shape the learner’s experience using capability, uncertainty, learning progress, and failure risk.
Imitation, inverse reinforcement learning, and reward inference for problems where desired behaviour is easier to demonstrate than to specify directly.
Objectives and evaluation methods for distribution shift, reward misspecification, and low-probability but consequential failures.
Applying reinforcement learning, reward learning, and adaptive experience generation to multimodal models and interactive agents.
Selected work
Adaptive training, learning from demonstrations, reward design, and evaluation in simulated and physical systems.
Internal deployment work on RL locomotion controllers, sim-to-real training workflows, and safety-critical evaluation for industrial quadrupeds.
Reward-model objective design from diverse demonstrations, with an invariance constraint intended to remain useful when the learned objective is optimized under distribution shift.
Peer-reviewed work on reinforcement-learning benchmarks and assistance policies in surgical digital-twin environments.
Turning optimal-transport distances between expert and policy behavior into practical reinforcement-learning rewards.
ArXiv preprint on generative modeling with Wasserstein autoencoders in hyperbolic latent spaces.
Selected publications
Work spanning reward learning, reinforcement learning, causal invariance, and optimal transport.
Preprint / manuscript · Ivan Ovinnikov, Eugene Bykovets, Joachim M. Buhmann
*International Journal of Computer Assisted Radiology and Surgery* · Ivan Ovinnikov, Ami Beuret, Flavia Cavaliere, Joachim M Buhmann
Research manuscript · Ivan Ovinnikov, Alexander Terenin, Joachim M. Buhmann
Preprint / manuscript · Ivan Ovinnikov, Joachim M. Buhmann
Preprint / manuscript · Ivan Ovinnikov
Writing
A work-in-progress measurement plan for locating where the learning signal disappears between reward, advantage estimation, and policy updates.
A short technical note on connecting reward learning, simulation-based training, and safety-critical evaluation for physical AI systems.
Successfully passed my PhD examination