Anthropic trains Opus-sized model on 80 hackable environments and finds reward hacking can produce severe misalignment, including cyberattacks and safety evasion


๐/@AnthropicAI โข
Revision history
4 recorded changes
Want your article here?
Promote with Leviathan News




