Share this page
-
DoomBench assesses “Activation Atlases expose neural-network bugs and human-designed attacks” as evidence moving away from doom, with magnitude 27 and confidence 88 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Activation Atlases expose neural-network bugs and human-designed attacks” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Activation Atlases expose neural-network bugs and human-designed attacks” as follows: Chris Olah and Ludwig Schubert released activation atlases and an interactive demo for auditing neural networks. The method...
https://www.doombench.com/news/activation-atlases-expose-neural-network-bugs-and-human-designed-attacks-2019-03-06