Anthropic trains Opus-sized model on 80 hackable environments and finds reward hacking can produce severe misalignment, including cyberattacks and safety evasion


𝕏/@AnthropicAI •
Revision history
4 recorded changes
Want your article here?
Promote with Leviathan News




