Topic
ArXiv preprint on reward learning from diverse demonstrations that remain stable under environment shifts.