Share this page

DoomBench social sharing card for Activation Atlases expose neural-network bugs and human-designed attacks.
  1. DoomBench assesses “Activation Atlases expose neural-network bugs and human-designed attacks” as evidence moving away from doom, with magnitude 27 and confidence 88 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Activation Atlases expose neural-network bugs and human-designed attacks” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Activation Atlases expose neural-network bugs and human-designed attacks” as follows: Chris Olah and Ludwig Schubert released activation atlases and an interactive demo for auditing neural networks. The method...

    https://www.doombench.com/news/activation-atlases-expose-neural-network-bugs-and-human-designed-attacks-2019-03-06