New findings expose a critical disconnect between lab performance and production reality in AI agent security
Highlights:
· AI agents can pass security benchmarks yet still fail in the real world. Advai's research shows current tests often miss what matters most: speed, cost, reliability and resilience.
· As AI agents move into finance, government, insurance and other high-stakes sectors, failures become business, regulatory and reputational risks.
· Advai is calling for AI security testing that reflects real-world performance, not just success in controlled lab environments.
· These tests address a critical gap not covered by foundation model evaluations, focusing on how AI systems perform in real-world deployment rather than in controlled model assessments.
Advai LTD ("Advai" or the "Company"), a leader in safe and secure AI adoption, is calling out a systemic blind spot in how AI agent security is measured. Benchmarks dominating the field today tell organisations whether an agent completed a task, but say almost nothing about whether that agent is fit for the real world.
The finding comes from Advai's systematic evaluation of CaMeL, a state-of-the-art agent architecture from Google DeepMind and ETH Zurich, conducted as part of the Laboratory for AI Security Research (LASR)* Opportunity Call on Agentic Security, supported by Plexal and Cisco. When Advai moved beyond standard "task completed" metrics to test how agents perform under real conditions, measuring time-to-completion, execution failure rates, cost and resilience under attack, the results were stark.
The benchmark problem is bigger than any single agent
The AI security community has made genuine progress in assessing whether agents resist prompt injection or follow instructions correctly. But evaluations widely used today share a common flaw: they are designed for the lab, not for production. They don't ask how long a task takes, how often execution fails, what the compute bill looks like, or whether an attacker could simply grind an agent to a halt.
Advai's evaluation of CaMeL exposed all four of these gaps. Despite its strong security credentials, the architecture ran up to 120x slower than a standard tool-calling agent on equivalent tasks, produced frequent code execution errors that would be unacceptable in live environments, and offered no effective resistance to Denial-of-Service attacks, with some safeguards actively increasing exposure. None of these findings would have surfaced using conventional benchmarks.
The research also demonstrated that closing the gap between security and operational reality does not require starting over. One straightforward change, limiting automatic retries, reduced deployment costs by more than 50% with negligible impact on results.
David Sully, CEO and Co-Founder of Advai, commented:
" Businesses shouldn't have to choose between secure AI and usable AI.. The AI security field is currently measuring the wrong things. Knowing an agent passed a lab test is not the same as knowing it is safe, cost-effective and resilient in production. The good news is that real-world improvements are achievable right now, without sacrificing security for performance."
Advai has made practical evaluation methods and tooling available to help organisations quantify real-world cost-benefit trade-offs and select security architectures that genuinely fit their risk profile, reinforcing the UK's AI security ecosystem and building justified trust in AI through standards that reflect reality.
Contact
Advai LTD | Email: advai@bursonbuchanan.com |
Media Enquiries) |
|
About Advai
Advai is an independent AI assurance company, helping organisations unlock the potential of artificial intelligence with confidence. Our platform delivers fast, scalable and context-aware testing across multiple AI systems - measuring performance, security, robustness and bias in real-world conditions.
We combine deep technical testing with expert human judgment, giving decision-makers objective, clear insights they can trust. From procurement to post-deployment monitoring, Advai reduces AI model selection from months to weeks, enabling faster, safer adoption at enterprise scale.
Headquartered in the UK with global reach, Advai works with leading organisations in finance, defence, insurance and the public sector. By aligning with international standards and regulatory frameworks, we help enterprises deploy AI securely and responsibly - turning risk into competitive advantage.
*This work was supported by the Laboratory for AI Security Research (LASR). The views expressed in this paper are those of the authors and do not necessarily reflect the position of LASR or His Majesty's Government.
