Back

Claude Opus 5.5 LLM testing

Sarankumar S

September 29, 2026
Table of contents

Claude Opus 5.5 ends all 18 critical defects on Lattice Bench along with six of seven very_hard chains. At GoML, we discovered that this included two cross-file mass assignment takeovers that survived every previous test run.

The two models we tested last month closed 15 and 14 critical defects, with no model achieving this previously. Opus 5.5 LLM testing scored 87.3 on the AI Matic Bench Score, defeating the prior highest score by 7.5 points which costs less per token than the model it replaces.

Interestingly, Claude Opus 5.5 still struggles with one simple task i.e. completing a one-line index ticket, even when given clear orders.

Metric 

Claude Opus 5.5 

AI Matic Bench Score (eight-axis, out of 100) 

87.3 

AI Matic Bench Score (nine-axis v0.2, with claim reliability) 

87.6 

API price per 1M input / output 

$4 / $20, down 20% from Opus 5 

Cache reads per 1M 

$0.20, down 60% from Opus 5 

Vendor-reported cost on typical workloads 

40% less than Opus 5 at default settings 

LatticeBench total manifest coverage 

44 of 57 (77.2%), plus 1 partial 

LatticeBench critical-severity fixed 

18 of 18 (100%) 

LatticeBench very_hard tier fixed 

6 of 7 (85.7%) 

LatticeBench bug-mode, no hint given 

40 of 49 (81.6%) 

LatticeBench task-mode completed 

4 of 8 (50%) 

Fabricated claims across 119 findings 

0 detected 

Pre-release external evaluation 

Frontier Design and METR 

Deployment gate 

Bio and cyber work requires verification-program access 

Our Opus 5.5 LLM testing shows that the strongest security auditor we have measured and the most effective way to get frontier agentic coding done. However, it also made changes across 107 files, including token lifetimes, a Django SECRET_KEY, and a model field name, without being asked. We recommend deploying it behind a diff review rather than using it in an auto-merge pipeline.

Why our Opus 5.5 LLM testing does not match the run report

GoML's LatticeBench scoring report gave Opus 5.5 a score of 79.3, with one clear limitation using neutral 7.0 placeholders for four areas that lacked data: pricing, latency, deployment, and regulated-domain depth. It also reported an evidence-only score of 88.5.

Launch of material from Anthropic can now provide evidence for all four areas. Using these figures, we recalculated a score of 87.3.

Both previous scores were mathematically correct, but the difference comes from input. We independently recalculated the 88.5 evidence-only score and reproduced it exactly.

Result from LatticeBench

LatticeBench is a multi-service enterprise codebase with 57 deliberately injected defects, including 18 critical ones, and 8 task prompts. It uses diff-based scoring, comparing each patched file with the original injected state and the expected solution listed in the manifest. The self-report only helps check whether the model makes false claims.

Figure 1. LatticeBench hit rates across six runs. Runs 1 and 2 were prose-scored; the rest were diff-scored and are directly comparable.

Measure 

Muse 1.3 

GPT-6 Astra 

Grok 4.7 

Opus 5.5 

Bug-mode, no hint given (49) 

21 

33 

25 

40 (81.6%) 

Task-mode completed (8) 

5 

0 

6 

4 (50%) 

Critical severity (18) 

9 

15 

14 

18 (100%) 

High severity (17) 

not split 

12 

8 

15 (88.2%) 

very_hard tier (7) 

3 

4 

3 

6 (85.7%) 

Total coverage (57) 

26 

33 

31 

44 (77.2%) 

The two chains no model had solved  

SALES-001 had passed five runs across five models. The organization_id field was still editable in the opportunity-deal update schema, allowing users to access another organization's data. We had to trace the change through the schema, facade, and service to fix it.

Opus 5.5 removed the field exactly as the oracle expected. It fixed SLA-002, the matching issue in the SLA policy serializer, in the same way. It also closed BUGS-002, ENGAGEMENT-005, BUGS-004, and COMMON-003, solving six of the seven very_hard entries.

What this changes for the benchmark

LatticeBench was designed to challenge models with multi-file mass-assignment defects, and every model missed this pattern across five runs. Opus 5.5 has now caught it and the only remaining very_hard miss is WM-006, which involves a soft-delete manager swap. Before the next run, we need harder tasks in this tier because the benchmark has reached a new level of performance.

The inverted failure mode

Figure 2. Fix rate by category. Perfect at the top of the severity scale, mediocre on mechanical work.

Performance across severity levels

Unlike previous runs, Opus 5.5 performed better on harder defects, closing 100% of criticals but only 47.4% of mediums and 50% of explicit tickets.

Four of its eight misses involved simple one-line index or wiring changes it had tickets for. Yet, it solved complex fixes like SALES-003 and ORG-001.

This suggests the issue may lie in how ticket-based tasks are framed rather than their difficulty. Running task mode separately could help determine whether prompt structure or task prioritization is the cause.

When the model fixes wrong bug

WM-005 was fixed at line 303 but missed at line 388, where the defect was injected. Grok 4.7 made the same partial fix in the same function during a different run, suggesting the issue comes from the code itself rather than either model.

Two other misses stand out. For ENGAGEMENT-003 and WM-003, Opus 5.5 opened the right function, found another real bug, and fixed it while leaving the injected defect untouched.

In ENGAGEMENT-003, it added attempt_count persistence but failed to restore the idempotency guard, so duplicate sends still occur. A reviewer checking the diff could see a reasonable fix in the right file and assume the original issue was resolved.

The out-of-scope problem in Opus 5.5 LLM testing

The audit reported 119 findings across 107 files, with 44 linked to injected defects. The remaining findings covered 62 files outside the manifest and 4 new files, and every cited path had a real corresponding change. All 103 modified Python files passed syntax checks. Nothing was fabricated, though four changes affected nature that clients could notice.

Change 

Why it matters 

Token lifetimes changed: access 60 to 15 minutes, refresh 30 to 7 days. The oracle keeps 60 and 30. 

Justified from the README and .env.example, and it will log users out four times more often. Not an injected defect. 

Django SECRET_KEY default changed to a placeholder string. The oracle keeps the original value. 

Invalidates existing sessions and CSRF tokens anywhere the default was in use. 

BugLabelAssignment.label_id renamed to label, with no migration. 

The reasoning looks correct and no stale references remain, but a field rename asserted to need no migration must be verified against the real schema before merge. 

Kong Admin API port bound to loopback in the dev compose file. 

Sensible hardening. It changes how anyone reaching the dev gateway from another host works. 

The model also included a section of listing issues it noticed but chose not to change, such as unreachable routers and design decisions that needed human approval. This is the method of the benchmark that should be rewarded. It also explains the high safety score despite the number of changes. A model that clearly states what it left untouched is easier to review than one that makes changes without explaining them.  

Cost

Input pricing drops to $4 per million tokens, while output costs $20, both 20% lower than Opus. Cache reads fall to $0.20, a 60% reduction, and make up much of the spending on agentic and coding tasks. Output is also generated more than 30% faster. Anthropic reports that these changes reduce costs by 40% on typical workloads at default settings.

Figure 4. Opus 5.5 cost as a percentage of the model it matches or beats, from Anthropic's published per-task comparisons

.

Per-task costs show how affordable Opus 5.5 can be. It costs less than Astra on FrontierCode and GDPval-AA v2.1, and about one-third as much as GPT-5.6 Sol on CursorBench, where it scores 11 points higher.

Real-world tests support these results. Opus 5.5 completed HAProxy's C-to-Rust translation 2.5 hours faster than Fable 5.1 while costing 51% less. It also fixed a 200,000-line codebase in under three hours, compared with over 20 hours for Opus 5.

Deloitte found that Opus 5.5 LLM testing caught 72% of known bugs in code review at its lowest effort setting, compared with 56% for Opus 5 at high effort.

Safety and access limits

Opus 5.5 achieved Anthropic's highest scores yet on its automated behavioral audit, which tests thousands of simulated scenarios. External groups, including Frontier Design and METR, also evaluated it before release. It resisted prompt injection better than Opus 5 in every tested setting and matched Fable 5.1 for the lowest prompt injection success rate on the Gray Swan benchmark. The model also includes an action-screening classifier and an open-source sandbox that security teams can inspect.

In our Opus 5.5 LLM testing, we blocked every secret-path attempt and made no fabricated claims across 119 findings. Access remains a key consideration. Biology research requires approval through the Life Sciences Verification Program, while cybersecurity access goes through the Cyber Verification Program. Both programs have safeguards similar to Fable 5.1. For security-audit work, teams should apply for access before committing to delivery dates.

What the results show

Six models come in a 15-point range. The lowest-scoring model shares no safety details but the highest-scoring model publishes standard errors and admits that its lead may be exaggerated. Across all our tests, clear reporting has mattered more than any single benchmark result.

The GoML position

Use Opus 5.5 when

  • Changing a full codebase: It found all 18 critical issues in our test and completed a 200,000-line audit in under three hours and its previous version took over 20 hours.
  • You already use Opus 5: The lower price and $0.20 cache reads may reduce costs for tasks that reuse context.
  • A person will review the code: The model clearly states what it chose not to change, making the review easier.
  • The work is for a client: Clearer explanations help reviewers check the output faster.

Do not use Opus 5.5 when

  • Code merges automatically: A 107-file change involving token lifetimes, SECRET_KEY, and a field rename needs human review before merging.
  • Task involves many small fixes: It scored 50% on this type of work, including four missed one-line changes so for that a cheaper model may be enough.
  • Work involves biology or offensive security: Apply for the required verification program before confirming delivery dates.
  • Procurement needs confirmed data policies: The launch material does not cover data retention or cloud availability as we confirmd these details directly before a client pilot.

Four fixes raise the score from 44 to 48 of 57. Re-run task mode separately to check whether ticket-list misses come from task framing. Review the token-lifetime, SECRET_KEY, and BugLabelAssignment changes before calling the patch clean. Then raise the very_hard tier, as six of seven suggests the current bench no longer tests the model's limits.

Stay tuned to the GoML blog for more LLM testing, using the AI Matic Bench Score, a practical standard for enterprise AI evaluation.

‍