TNW→ original

AI Models Outpace Security Tests: Regulators Lose Visibility into Capabilities

Tools built to measure AI risk have stopped working. Frontier models now surpass benchmarks designed to assess their hacking capabilities. This creates a blind spot for regulators and security specialists. By estimates, U.S. federal agencies must by August 1, 2026 create a classified registry of AI systems' cyberattack capabilities, but current tests are already inadequate.

AI-processed from TNW; edited by Hamidun News
AI Models Outpace Security Tests: Regulators Lose Visibility into Capabilities
Source: TNW. Collage: Hamidun News.
◐ Listen to article

Safety and risks of AI are increasingly going beyond what is measurable. Frontier models from OpenAI, Anthropic and other labs have begun bypassing tests designed to assess their cyber-attacking capabilities. This creates a critical gap in our understanding of capabilities and risks.

The Benchmark Problem

Benchmarks for assessing cybersecurity of AI models were developed only a few years ago and follow these approaches:

  • Screen capture and attempts to hack web applications
  • Detection of vulnerabilities in security systems
  • Social engineering and phishing scenarios

However, frontier models now solve these tasks with a probability exceeding the accuracy of the tests.

Regulatory Vacuum

  • Federal US agencies must prepare a classified registry by August 1st
  • Current tools are insufficient for actually measuring capabilities
  • Risk: regulators may make decisions based on incorrect information

What This Means

This gap in assessment creates a unique challenge for AI regulation. Tests that should measure safety are no longer working. This may mean that real risks from advanced AI systems are significantly higher than regulators understand. It is necessary to develop new assessment methodologies that keep pace with the development of AI capabilities.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…