Claude Sonnet 5 Became First Model to Pass All GitLab Duo Agent Platform Tests
Claude Sonnet 5 from Anthropic is now available on GitLab Duo Agent Platform on all tiers and all deployment variants via AI Gateway. The model is the first in GitLab's evaluation suite to pass all benchmark tasks; its predecessor Sonnet 4.6 closed 93.8% of tasks. GitLab also reports an 8.8% increase in resolved issues.
AI-processed from GitLab Blog; edited by Hamidun News
Claude Sonnet 5 became the first model to pass all tests on GitLab Duo Agent Platform
Anthropic and GitLab announced the launch of Claude Sonnet 5 on the GitLab Duo Agent Platform: it is available on all tiers and in all deployment options through GitLab's AI Gateway and became the first model in GitLab's evaluation set to pass all benchmark tasks.
How much better is Sonnet 5 than its predecessor
GitLab tests models on its own set of tasks that reflect the real work of AI agents in development: multi-step tasks, code generation that withstands code review, and execution of workflows with acceptable costs. Claude Sonnet 5 became the first model to complete all tasks in this set; its predecessor, Sonnet 4.6, covered 93.8% of tasks from the same set.
- Claude Sonnet 5 — first model in GitLab's evaluation set to pass 100% of benchmark tasks
- Sonnet 4.6, the previous model, covered 93.8% of tasks from the same set
- Number of solved issues increased by 8.8%
- Model is available on all tiers and in all deployment options of GitLab Duo through AI Gateway
What changes for development teams
For GitLab Duo Agentic Chat users this changes the usual cycle "prompt — wait — evaluate result". Multi-file refactoring now produces a result suitable for review instead of a task cut off midway. Test generation returns coverage that can actually be used.
Security investigations trace repository history deeper. Basic Duo agents handle most of the assigned work without human intervention, so team time goes to reviewing results rather than restarting failed tasks. GitLab gives a concrete example of such an agent scenario: in the GitLab Orbit tool the model is already used to investigate pipeline failures over the last two months — that is, to analyze the history of build failures and find the cause right in dialogue with an agent.
"Claude Sonnet 5 handled the full range of code writing tasks we tested, and solves more problems.
This is a tangible improvement in both quality and efficiency. We made it available in GitLab Duo Agent Platform today, on all tiers and in all deployment options," — Manav Khurana, Director of Product and Marketing at GitLab.
What this means
The costliest failure of an agent is the one that stops halfway: costs accumulate not only from lost work but also from diagnostics, reprompting, and verifying that something did come back. A model capable of passing the full set of GitLab benchmark tasks without failure turns an agent from a tool that needs constant monitoring into a tool you can truly delegate work to. This is the threshold — reliability, not just the quality of an individual answer — where GitLab and Anthropic propose to evaluate the progress of agent models in software development.
Frequently asked questions
How does
Claude Sonnet 5 differ from Sonnet 4.6 on GitLab tests?
Claude Sonnet 5 became the first model to pass all benchmark tasks in GitLab's evaluation set, whereas Sonnet 4.6 covered 93.8% of tasks from the same set, and the number of solved issues increased by 8.8%.
On which GitLab tiers is Claude Sonnet 5 available?
The model is available on all tiers and in all deployment options of GitLab Duo Agent Platform through GitLab's AI Gateway — no special increased tier was required for access to Sonnet 5.
Where is Claude Sonnet 5 already used in practice at GitLab?
GitLab gives an example of the GitLab Orbit tool, where the model is used to investigate pipeline failures over the last two months — that is, to analyze the history of build failures right in dialogue with an agent.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.