Topic
Turning optimal-transport distances between expert and policy behavior into practical reinforcement-learning rewards.