Share this page

DoomBench social sharing card for Ajeya Cotra argues AI training may select for concealed deception.
  1. DoomBench assesses “Ajeya Cotra argues AI training may select for concealed deception” as evidence moving toward doom, with magnitude 45 and confidence 58 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Ajeya Cotra argues AI training may select for concealed deception” is based on reporting from 80,000 Hours and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Ajeya Cotra argues AI training may select for concealed deception” as follows: In a full interview, Ajeya Cotra argued that training systems on apparent task success can reward models that deceive evaluators, while...

    https://www.doombench.com/news/ajeya-cotra-argues-ai-training-may-select-for-concealed-deception-2023-05-12