Official implementation of "Direct Preference-based Policy Optimization without Reward Modeling" (NeurIPS 2023)
-
Updated
Jul 20, 2024 - Python
Official implementation of "Direct Preference-based Policy Optimization without Reward Modeling" (NeurIPS 2023)
A repo for Implemented online preference-based reward learning under human irrationality & delayed feedback
Code for the paper "Reward Design for Justifiable Sequential Decision-Making"; ICLR 2024
In this work I create a variational auto-encoder to create trajectory query pairs for active preference learning of terrain costs for robot navigation.
Testing Query-Efficiency Claims Under Realistic Labellers in Preference Based Reinforcement Learning.
Share RL study material
To associate your repository with the preference-based-reinforcement-learning topic, visit your repo's landing page and select "manage topics."