Safety and alignment

DeepMind and OpenAI demonstrate reinforcement learning from human preferences

DeepMind and OpenAI researchers showed that reinforcement-learning agents could learn complex behavior from sparse human preference comparisons, using roughly 900 bits of feedback for a simulated backflip and reaching superhuman Atari performance. The work documented reward-hacking limits while establishing a practical technique later deployed in instruction-following language models.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM56confidence 98/100

Why it moved the index

The primary experiment demonstrated a scalable way to transmit human preferences to agents without hand-writing a reward function, and it explicitly surfaced reward-hacking and feedback-quality limitations. Its practical significance is corroborated by OpenAI's later deployment of the same human-feedback pathway in InstructGPT, making this a concrete control technique rather than an untested proposal.

AUDIT TRAIL

Assessment history

  1. R1
    Away 56 · confidence 98

    Initial inclusion from a dated primary result with separately verified practical deployment impact.

    02 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for DeepMind and OpenAI demonstrate reinforcement learning from human preferences.
  1. DoomBench assesses “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” as evidence moving away from doom, with magnitude 56 and confidence 98 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” is based on reporting from Google DeepMind and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “DeepMind and OpenAI demonstrate reinforcement learning from human preferences” as follows: DeepMind and OpenAI researchers showed that reinforcement-learning agents could learn complex behavior from sparse human...

    https://www.doombench.com/news/deepmind-and-openai-demonstrate-reinforcement-learning-from-human-preferences-2017-06-12