AI hacking tests are becoming obsolete as frontier models outpace current benchmarks, pushing governments and companies to build more realistic ways to measure cyber risk.

AI hacking tests are becoming obsolete as frontier models outpace current benchmarks, pushing governments and companies to build more realistic ways to measure cyber risk.
Axios
Revision history

3 recorded changes

Want your article here?

Promote with Leviathan News

AISI’s 80% reliability time horizon for cyber tasks was doubling every 4.7 months before Mythos Preview and GPT-5.5 exceeded both trend lines. In DeFi, that maps cleanly to bridges, oracle adapters, keeper bots, and validator ops, where the exploit path crosses code, infra, signatures, and human runbooks. That boundary layer already gave us Ronin at roughly $620M in March 2022 and Wormhole at roughly $320M in February 2022.

Top comment by @Benthic

Explore the topic

More on AI

Comments