AI Models Outpace Security Tests: Regulators Lose Visibility into Capabilities
Tools built to measure AI risk have stopped working. Frontier models now surpass benchmarks designed to assess their hacking capabilities. This creates a blind spot for regulators and security specialists. By estimates, U.S. federal agencies must by August 1, 2026 create a classified registry of AI systems' cyberattack capabilities, but current tests are already inadequate.
AI-processed from TNW; edited by Hamidun News
Safety and risks of AI are increasingly going beyond what is measurable. Frontier models from OpenAI, Anthropic and other labs have begun bypassing tests designed to assess their cyber-attacking capabilities. This creates a critical gap in our understanding of capabilities and risks.
The Benchmark Problem
Benchmarks for assessing cybersecurity of AI models were developed only a few years ago and follow these approaches:
- Screen capture and attempts to hack web applications
- Detection of vulnerabilities in security systems
- Social engineering and phishing scenarios
However, frontier models now solve these tasks with a probability exceeding the accuracy of the tests.
Regulatory Vacuum
- Federal US agencies must prepare a classified registry by August 1st
- Current tools are insufficient for actually measuring capabilities
- Risk: regulators may make decisions based on incorrect information
What This Means
This gap in assessment creates a unique challenge for AI regulation. Tests that should measure safety are no longer working. This may mean that real risks from advanced AI systems are significantly higher than regulators understand. It is necessary to develop new assessment methodologies that keep pace with the development of AI capabilities.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.