Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests
Anthropic's Claude models advancing AI self-correction could redefine AI safety standards, challenging human roles in alignment tasks. The post Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests appeared first o…
The company's automated alignment researchers improved performance across all ten misalignment benchmarks, beating 28 human safety researchers in the process. Anthropic just published research showin… [+2054 chars]