DeepMind and OpenAI demonstrate reinforcement learning from human preferences
DeepMind and OpenAI researchers showed that reinforcement-learning agents could learn complex behavior from sparse human preference comparisons, using roughly 900 bits of feedback for a simulated backflip and reaching superhuman Atari performance. The work documented reward-hacking limits while establishing a practical technique later deployed in instruction-following language models.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The primary experiment demonstrated a scalable way to transmit human preferences to agents without hand-writing a reward function, and it explicitly surfaced reward-hacking and feedback-quality limitations. Its practical significance is corroborated by OpenAI's later deployment of the same human-feedback pathway in InstructGPT, making this a concrete control technique rather than an untested proposal.
Assessment history
-
R1
Away 56 · confidence 98
Initial inclusion from a dated primary result with separately verified practical deployment impact.
02 Sept 2026
Share this page
-
DoomBench assesses “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” as evidence moving away from doom, with magnitude 56 and confidence 98 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” is based on reporting from Google DeepMind and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” as follows: DeepMind and OpenAI researchers showed that reinforcement-learning agents could learn complex behavior from sparse human...
https://www.doombench.com/news/deepmind-and-openai-demonstrate-reinforcement-learning-from-human-preferences-2017-06-12