Back

GPT 6 LLM testing: What the benchmarks mean for enterprise delivery

Sarankumar S

September 8, 2026
Table of contents

Our team, as expected, was elated to perform GPT 6 LLM testing for the model OpenAI calls Astra, against their published tables and our LatticeBench harness. Impressively (again as expected), GPT-6 Astra scores 79.5 on the AI Matic Bench Score. That is the highest total we have recorded to date, and it comes from the most expensive model we have scored by an order of magnitude. Astra also posts the strongest LatticeBench result any model has produced on our own bench: 33 of 57 injected defects fixed, 15 of 18 of them critical, and zero fabricated claims across 107 audit findings.

Where those points come from matters more than the total, and it is the part of GPT 6 LLM testing that the launch tables do not answer for you. GPT-6 Astra does not lead the aggregate intelligence indices, and on the single coding benchmark most teams quote, it beats a model costing a thirteenth as much by three tenths of a point. What it does lead, by margins that are not close, is agentic computer use, cybersecurity, abstract reasoning and alignment. That is a narrow and specific profile, and it maps to a specific set of engagements rather than to every workload you run.

Headline numbers of GPT 6 LLM testing

Metric 

GPT-6 Astra 

AI Matic Bench Score (eight-axis, out of 100) 

79.5 

AI Matic Bench Score (nine-axis v0.2, with claim reliability) 

80.7 

API price per 1M input / output 

$10 / $50 (Fast mode: 2x speed at 2x price) 

Price relative to Gemini 3.8 Flash 

13.3x on both input and output 

Artificial Analysis Intelligence Index v4.1.1 

61.2, fourth behind Fable 5.1, Opus 5 and Fable 5 

LatticeBench total manifest coverage 

33 of 57 (57.9%), plus 1 partial 

LatticeBench critical-severity fixed 

15 of 18 (83.3%) 

LatticeBench very_hard tier fixed 

4 of 7 (57.1%) 

LatticeBench task-mode completed 

0 of 8 

Fabricated claims across 107 audit findings 

0 verified as false 

Enterprise availability 

OpenAI API, Microsoft Azure, AWS Bedrock 

Zero Data Retention 

Supported for eligible API customers 

Preparedness Framework cybersecurity classification 

Critical threshold 

The one-line read. Astra is the first model on our bench to fix more than 80% of critical defects, and the first where the price turns the decision into "which tasks deserve this model" instead of "should we switch". Route by task class.

Where GPT-6 Astra leads

The margins in this release are unusually wide in a handful of places. ARC-AGI-3 at 99.9% against 30.2% for Opus 5 and 7.8% for GPT-5.6 Sol is not an incremental gain. Neither is SRE-Bench at 88.0% against 12.5% for Opus 5, nor Terminal-Bench Science at 64.6% against 30.0%. OpenAI also reports that Astra surpassed the ARC Prize Foundation's human action-efficiency baseline on 96% of levels.

Computer use is the category with the most direct relevance to delivery. ScreenSpot-Pro reaches 92.7% against 76.9% for Sol. OSWorld 2.0 reaches 72.6% against 65.7% while taking roughly 47% less time per task. Agents' Last Exam reaches 59.3% against 55.5% for Opus 5 while using around 65% fewer output tokens. Paired with the updated Codex harness, OpenAI reports 1.9x faster task completion on Mind2Web than the current Sol experience. For teams doing AI software development work against real applications, this is the category to watch.

Cybersecurity is where the capability claims are strongest and the governance implications are largest. ExploitBench sits at 100%, SRE-Bench at 88.0% single-attempt and 99.2% within four attempts, and OpenAI reports a purpose-built contamination-free benchmark on vulnerabilities from the previous three months where Astra discovered and used two previously unknown zero-days. OpenAI classifies the model at the Critical threshold for cybersecurity under its Preparedness Framework and ships it with refusals on advanced offensive tasks. Defensive work stays in scope today, including secure code review and patching.

Where GPT-6 Astra does not lead

Three independent aggregate measures rank Astra below at least one competitor. This is the part of the release that launch coverage has mostly skipped.

On the Artificial Analysis Intelligence Index v4.1.1, Astra scores 61.2 and sits fourth, behind Claude Fable 5.1 at 65.7, Opus 5 at 63.1 and Fable 5 at 62.1. On the Artificial Analysis Coding Agent Index it scores 67.0 against 68.1 for Opus 5 and 67.2 for Fable 5. On Humanity's Last Exam with tools it scores 57.2 against 65.0 for Fable 5.1, a 7.8-point gap and the largest deficit in the table. FrontierCode 1.1 Main is a three-way tie within two tenths of a point.

Reconciling the two halves

Both pictures are accurate. Aggregate indices average across task types, and Astra's advantage is concentrated in agentic, computer-use and cyber work rather than spread evenly across them. A model can set the frontier on the benchmarks that predict enterprise agent performance while sitting fourth on an index that weights broad knowledge equally. Which number matters depends entirely on the workload in front of you.

The price question

At $10 per million input tokens and $50 per million output, Astra costs 13.3 times what Gemini 3.8 Flash costs on both sides, and twice the output price of Claude Opus 5. Fast mode doubles throughput at double the rate. That pricing has to be justified per task class, and on one widely quoted benchmark it cannot be.

On DeepSWE v1.1, Astra scores 74.1%, Gemini 3.8 Flash 73.8% and Opus 5 73.7%. Three tenths of a point separate the most expensive model in the set from one costing a thirteenth as much. If your agentic coding workload looks like DeepSWE, Astra is the wrong answer, and the AI Matic unit-economics axis reflects that.

The picture inverts on harder work, and OpenAI publishes the cost figures alongside the scores rather than the scores alone. Terminal-Bench 4.0 comes in at 57.9% against 55.8% for Fable 5.1 at approximately 63% lower estimated cost per task, and against 37.3% for Sol at 9% lower cost. Terminal-Bench Science reaches 64.6% against 52.6% at roughly 31% lower cost. BenchCAD reaches 95.9% against 83.3% for Sol and 84.3% for Fable 5.1, at approximately 43% and 86% lower cost. GPQA Diamond at a lower-cost setting still beats Sol's best score at around 37% lower cost.

How to read this

Astra is expensive per token and frequently cheaper per completed hard task, because it uses far fewer output tokens to get there. On easy tasks that a cheap model already completes, you pay 13x for nothing. The delivery decision is therefore a routing decision: send the task classes where cheaper models measurably fail, and leave everything else where it is.

Enterprise fit

Dimension 

What ships 

Delivery implication 

Availability 

OpenAI API as gpt-6-astra, Microsoft Azure, AWS Bedrock 

Bedrock access matters for AWS-native builds, and no new vendor relationship is required 

Data handling 

Zero Data Retention for eligible API customers; Private Safety Processing in testing 

ZDR is the single most common enterprise blocker, and it is answered at launch 

Access control 

Enterprise admins enable per workspace; off by default at launch 

Plan for an admin enablement step rather than assuming availability 

Rollout 

Limited set of organizations first, broader over following days 

Confirm your account has it before committing a delivery date 

Safety interruptions 

Extra checks can pause or stop work. In ChatGPT and Codex you review and continue; in the API the task stops 

Pipelines need retry and human-escalation handling for stopped tasks 

Monitorability 

OpenAI reports Astra's written reasoning is harder to monitor than Sol's 

A disclosed regression, and relevant wherever reasoning traces feed an audit trail 

 

Two of these deserve emphasis. Zero Data Retention plus Bedrock availability removes the two objections that most often stall an enterprise pilot, and no other model we have scored answers both at launch. That combination is what makes the model straightforward to bring into existing AWS AI services architectures. Against that, the API stopping a task outright when a safety check fires is an operational behaviour to design for rather than discover in production. OpenAI states it is iterating to reduce unnecessary interruptions.

Source conflicts across three vendor tables

Three vendors published comparison tables within days of each other, and they disagree on the same competitors' scores. That disagreement makes cross-vendor tables unusable as a scoreboard.

# 

Conflict 

Status 

1 

DeepSWE v1.1 for Claude Opus 5 is reported as 74.0 by Meta and 73.7 by OpenAI. Gemini 3.8 Flash is reported as 73.7 by Google and 73.8 by OpenAI. 

Unresolved. Each vendor re-ran or re-sourced competitor scores. The differences are small, and they run in the direction each vendor benefits from. 

2 

AutomationBench: Meta reports Muse Spark 1.3 at 49.6, OpenAI reports Astra at 41.4 and does not include Muse. 

Not comparable. Different harnesses, no shared run. Muse should not be read as ahead of Astra here. 

3 

Long context: Meta reports MRCR 98.1 at 512K to 1M. OpenAI reports OpenAI MRCR v2 8-needle at 96.3 for Astra and 73.8 for Sol. 

Different benchmark configurations. The 8-needle variant is harder, so these two numbers cannot be compared. 

4 

The Artificial Analysis Intelligence Index for Gemini 3.8 Flash was quoted as 59 at its launch. OpenAI's table gives 58.7 on v4.1.1. 

Resolves toward 58.7. Our earlier Gemini note flagged this as open, and it can now be closed. 

5 

Several Astra figures come from internal benchmarks: Design Tasks, Data Science Tasks, Database Migration, and the hallucination and circumvention evaluations. 

Not independently reproducible. Treated as directional and not weighted as measurement. 

Credit where it is due on disclosure. OpenAI footnotes the ARC-AGI-3 harness change, the FrontierCode developer message, the Claude eval modifications on BenchCAD, the HealthBench grading procedure, the ExploitGym time-limit removal, and the fact that some ScreenSpot-Pro and ExploitGym Fable scores come from Mythos rather than Fable. It also discloses the regression in reasoning monitorability. That level of self-reported limitation is the most thorough of the three launches we have scored this month.

GPT 6 LLM testing on GoML's own bench: LatticeBench run 4

LatticeBench is a multi-service enterprise codebase carrying 57 deliberately injected defects, 18 of them critical severity, plus 8 explicit task prompts. Run 4 was scored the same way as run 3: diff the patched tree against golden/ and bench/ per manifest entry, then read the diff wherever neither matches exactly. Astra also left a written audit of 107 grouped findings, which makes this the first run where diff scoring and claim reliability could be measured together.

Measure 

Run 1 

Run 2 

Run 3 (Muse 1.3) 

Run 4 (Astra) 

Bug-mode, no hint given 

8 of 49 

14 of 49 

21 of 49 

33 of 49 (67.3%) 

Critical severity (18) 

2 of 18 

6 of 18 

9 of 18 

15 of 18 (83.3%) 

High severity (17) 

not broken out 

not broken out 

not broken out 

12 of 17 (70.6%) 

very_hard tier (7) 

1 of 7 

2 of 7 

3 of 7 

4 of 7 (57.1%) 

Task-mode completed (8) 

7 of 8 

0 of 8 

5 of 8 

0 of 8 

Organic fixes beyond the manifest 

7 

21 

1+ 

about 70 

Total manifest coverage 

15 of 57 

14 of 57 

26 of 57 

33 of 57 (57.9%) 

What run 4 did

Run 4 closed the entire shared-library blast radius in one pass: COMMON-001 through COMMON-004, meaning the pagination off-by-one, the JWT secret fallback, the disabled audience verification and the safe-methods authentication bypass. It caught WM-001, the single-line privilege escalation that had survived three previous runs. It closed BUGS-002, one of the multi-file mass-assignment chains, using an explicit ValidationError on cross-project moves instead of the schema change the manifest specified, which the scorer credited as an equivalent mechanism.

Several fixes were stricter than the reference implementation. COMMON-002 fails closed with an empty secret and an AuthenticationFailed rather than restoring the original behaviour. AI_COPILOT-002 fixed the history ordering, the duplicate current message and the racy re-fetch, where golden/ fixes the ordering alone. The org-invites-public regex, a real pre-existing defect in our own reference tree, was tightened to a better standard than golden/ carries.

Around 70 further changes sit outside the 57 injected defects: secret-key fail-closed enforcement across seven services, machine-to-machine auth, tenant scoping on the org, bugs and sales endpoints, timesheet ownership and manager locks, Kong secret placeholders and loopback admin binding. Every sampled one corresponds to a real diff, and none were fabricated.

What it missed

Zero of eight task prompts. Astra did not attempt them and made no claim about them, and all eight model files are byte-identical to the injected state. This repeats the pattern from run 2, where the strongest unprompted hunters ignored the explicit backlog entirely. Since all eight are single-index or small wiring changes, clearing them alone would take coverage from 33 to roughly 41 of 57.

Three criticals remain: IDENTITY-002 (admin password stored without hashing), SALES-001 and SLA-002 (organization_id still writable on update). Run 3 caught IDENTITY-002 and SLA-002, while run 4 caught BUGS-002 and WM-001 which run 3 missed. Different models are failing on different members of the same family.

Four runs, one defect nobody has caught. SALES-001, the organization_id field left writable on the opportunity-deal update schema, has survived every run by every model. It is a textbook cross-tenant takeover and the highest-value single fix in the benchmark. Whatever model you deploy, this class stays a human review gate and a static-analysis rule.

The severity-inflation problem

Astra's own audit framed its output as 107 findings with 26 critical and 55 high. The manifest contains 18 critical and 17 high in total. The inflation comes from counting the roughly 70 organic hardening changes at benchmark severity levels. None of those claims is false, and every one maps to a real diff, so this is a reporting-calibration problem rather than a fabrication problem. It still matters. An audit deliverable that reports 26 criticals to a client who has 18 will fail review, and that is why the claim-reliability axis scores 9.0 rather than higher.

The AI Matic Bench score

Vendor tables answer a research question: is the new model better. GoML delivery teams answer a different one: does putting this model into a running enterprise pipeline improve accepted outcomes per rupee and per hour, without forcing a requalification cycle. The AI Matic Bench Score is our rubric for the second question, scored 0 to 10 per axis from published evidence, independent measurement, or our own LatticeBench runs. It is the same lens we bring to AI consulting engagements when a client asks which model belongs in production.

Two totals, and why

Astra is the first model where both a patched tree and a written report exist, so the v0.2 claim-reliability axis can be scored. On the full nine-axis v0.2 rubric Astra scores 80.7. The eight-axis total of 79.5 drops that axis and is the number used for cross-model comparison, because Muse Spark 1.3 left no self-report and cannot be scored on it. Use 79.5 when comparing models, and 80.7 when assessing Astra on its own.

Each axis is scored 0 to 10 against the band anchors below. The total is the weighted sum of the axis scores, multiplied by ten to put it on a 100-point scale.

AI Matic Bench Score = ( sum of ( axis score x axis weight ) ) x 10, where weights sum to 1.00 and each axis score sits on a 0 to 10 band.

Band 

Meaning 

Test we apply 

9 to 10 

No concerns on this axis 

We would put it in a client proposal with no caveat attached 

7 to 8 

Production-viable with named guardrails 

Ships if a specific control is in place, and we can name the control 

5 to 6 

Usable under supervision 

Needs a human review gate on every output that reaches a client 

3 to 4 

Pilot only 

Fine for internal exploration, not for delivery 

0 to 2 

Disqualifying 

This axis alone rules the model out for the engagement class 

Scoring and derivation

Axis 

Weight 

Score 

Basis 

Long-horizon completion 

20% 

9.0 

Terminal-Bench 4.0 at 57.9 against 19.1 for Gemini 3.8 Flash; MRCR 8-needle 96.3 at 512K to 1M; LatticeBench 33 of 49 and 4 of 7 very_hard, the best recorded 

Agentic tool reliability 

15% 

8.0 

ScreenSpot-Pro 92.7, OSWorld 72.6 in 47% less time, zero auto-review circumvention; discounted for 0 of 8 LatticeBench task prompts 

Workload unit economics 

20% 

6.5 

$10 / $50 is 13.3x Gemini 3.8 Flash for +0.3 points on DeepSWE against it; genuinely lower cost per task on TB 4.0, TB Science, BenchCAD and GPQA 

Latency profile 

10% 

8.0 

47% less time per OSWorld task than Sol, 1.9x faster Codex completion on Mind2Web, Fast mode at 2x; no published throughput or TTFT 

Deployment envelope 

10% 

8.0 

API, Azure and Bedrock at launch with ZDR for eligible customers; against off-by-default access, staged rollout and API tasks stopping on safety checks 

Regulated-domain depth 

10% 

8.0 

HealthBench Professional 63.4 leading all; GPQA 96.0; LifeSciBench, GeneBench Pro and MedChemBench leads; no finance-agent benchmark; HLE trails all Claudes 

Safety and governance 

10% 

8.5 

Computer-use safety 2.4% against 22.0% for Sol; hallucination 4.2% against 12.2%; honeypot 0.0% against 48.2%; against the Critical cyber classification and the disclosed monitorability regression 

Evidence quality 

5% 

8.0 

Extensive published tables with footnoted limitations, plus a diff-scored GoML run; against several unreproducible internal benchmarks and self-run competitor scores 

Weighted total (eight-axis) 

100% 

79.5 

Production-viable with named guardrails 

On the nine-axis v0.2 rubric, claim reliability scores 9.0: zero of 107 findings verified as false, against a severity framing that does not map to the benchmark scale. That lifts the total to 80.7.

The resulting shape is the most lopsided we have scored. Seven axes sit at 8.0 or above. One sits at 6.5, and it is unit economics at the joint-heaviest weight. A model with identical capability evidence priced like Opus 5 would score above 82.

Astra leads the field by 4.2 points and is the only model in the set scoring below 7.0 on unit economics. Muse Spark 1.3 scores highest of the four on that axis at 8.0 and lowest overall, because it publishes nothing on safety. The rubric is doing what it was built to do: capability, cost and governance are separate procurement questions, and no model in this set wins all three. Regulated work makes that separation sharper still, which is why our healthcare AI delivery engagements weight governance heavier than any index does.

The GoML position on GPT-6 Astra

Route work to Astra when:

  • The task class is one where cheaper models measurably fail. Terminal-Bench 4.0 at 57.9 against Gemini 3.8 Flash's 19.1 is the profile to look for, where the gap is a multiple rather than a margin.
  • The work is computer use in real software: forms, CRM updates, spreadsheet and document production against a template, frontend QA. This is the strongest category in the release and the one with the least competition.
  • The engagement is defensive security work. Secure code review and patching are in scope today, and SRE-Bench at 88.0 against 12.5 for Opus 5 is a gap no other model closes.
  • ZDR is a contractual requirement. It is supported at launch, on Bedrock, which clears two blockers at once.
  • The deliverable is a document, deck or spreadsheet that has to match a house template. Template adherence is a stated training target, and it shows in the professional benchmarks.

Do not route work to Astra when:

  • A Flash-tier model already completes the task. DeepSWE is the cautionary example: 13.3x the price for three tenths of a point.
  • The workload is high-volume classification, extraction or routing. Nothing in this release changes those economics, and the price difference is severe. Pipelines of that shape belong on a cheaper model behind something like our AI data analytics accelerator.
  • Broad-knowledge quality is the binding requirement. Astra sits fourth on the Artificial Analysis Intelligence Index and 7.8 points behind Fable 5.1 on Humanity's Last Exam with tools.
  • The pipeline cannot tolerate a task stopping mid-run. API tasks stop when a safety check fires, and that needs explicit retry and escalation handling.
  • Cross-tenant authorization review is in scope. Astra fixed 15 of 18 criticals and still left the opportunity-deal mass-assignment field writable, as every model before it has.

Adoption protocol

Segment the pipeline instead of migrating it. Classify existing traffic by task class, identify the classes where the incumbent model measurably fails rather than merely underperforms, and route only those to Astra. Log input and output tokens, retries, stopped tasks, accepted-result rate and wall clock per class, then compute cost per accepted result rather than cost per call. Expect the answer to be that a small minority of classes justify the rate. Confirm workspace enablement and Bedrock region availability before you commit to a delivery date, and design the stopped-task path before the first production run rather than after it. This is the same segmentation exercise we run at the start of most enterprise AI consulting engagements.

Where this maps onto engagements GoML has already delivered

Agentic computer use in real software. GoML's Agentic AI blueprint, the same blueprint deployed for DevPlaza, connects a Git agent, GitHub Actions agent, Jira agent and SonarQube agent across one SDLC workflow. The agents validate pull requests, parse CI/CD failures, triage tickets and flag test coverage gaps. DevPlaza reported 35% higher unit test coverage, 2x fewer CI/CD failures, 40% better PR quality and 50% faster bug triage.

This workload closely matches the task family covered by ScreenSpot-Pro, OSWorld 2.0 and SRE-Bench because the agent works across several developer tools and maintains context across multiple steps. Its structure also resembles LatticeBench. Astra's lead across these tests makes multi-tool code review and CI/CD diagnosis a sensible area for testing. The same evidence says much less about plain code generation, so routing every coding task to Astra would be harder to justify.

Template-bound document processing. GoML's document automation work for Loft47, also built on the Agentic AI blueprint, extracts structured information from real estate contracts across at least 25 MLS template formats. The system checks signatures and required fields before sending the information into accounting systems. Large brokerages reported a 60 to 70% reduction in manual contract review time.

Parts of this workflow match Astra's document-handling profile. Contracts arrive in multiple formats, fields follow strict rules and errors carry real business consequences. Yet most extraction and classification work does not require Astra. A lower-cost model already handles the repeatable portion of the pipeline. Astra is more relevant for harder cases such as new template variants, conflicting fields, unusual clause placement or ambiguous signature locations. Routing only those cases gives Astra a more sensible role than sending the full document volume to a higher-priced model.

Regulated-domain reasoning with a governance layer. GoML's Agentic AI blueprint also underpins the compliance software built for Metatate, which converts natural-language descriptions of data activity into policy citations, risk levels and recommended next steps. The system reduced manual queries to legal teams by 60% and policy clarification time by 40%.

This type of workload gives Astra a useful test because the model has to reason through policy while leaving a record reviewers can inspect. ZDR, audit history and monitorability matter directly in this setting. The severity-inflation issue found in LatticeBench also becomes relevant. If Astra labels routine findings as severe too often, it could increase legal review rather than reduce it. A Metatate-style workload would show whether Astra's reasoning scores hold up when risk labels also need to remain well calibrated.

Where the fit breaks down. GoML's liability-code extraction pipeline for Ledgebrook combines OCR (AWS Textract) with a retrieval-augmented classification pipeline on Bedrock and OpenSearch, matched against a fixed codebook. It runs at 95% accuracy and completes the work 80% faster than manual review. Unlike the three cases above, this one isn't described as running on a named GoML blueprint; it's a custom AWS pipeline built before the current accelerator lineup, which is worth flagging rather than assuming continuity with the Agentic AI Accelerator.

This workload is repetitive, tightly defined and already solved at lower cost. Astra's benchmark lead does not change the economics. A smaller model remains the better fit unless testing shows enough measurable improvement to offset the higher model price.

None of these four systems were built on GPT-6 Astra, so the percentages above are not Astra performance figures. They come from our client work using Claude and other models. What these cases show is where Astra's benchmark profile overlaps with real production workloads. DevPlaza, Loft47 and Metatate contain task patterns that match areas where Astra scores well, and all three ran through GoML's blueprints, which is itself a useful signal.

GPT 6 LLM testing: Verification method

Benchmark figures come from OpenAI's published comparison tables and the accompanying footnotes. Cost comparisons are OpenAI's own estimates and are labelled as such rather than treated as measurement. Where a figure comes from an internal benchmark with no external reproduction, it is noted and does not carry an axis score on its own. Competitor figures are recorded as published and cross-checked against the other two vendor tables released this month, which produced the conflicts registered above.

LatticeBench run 4 figures come from diff-based scoring of the archived patched tree at runs/run-4/problem-set/. Each manifest entry was hashed against bench/ and golden/, and where neither matched, the diff was read at the primary line to judge convergence on golden behaviour, with equivalent mechanisms accepted. Every fixed entry was cross-checked against the audit's claimed ID, line and description. Every missed entry was confirmed untouched, and all eight task files were confirmed byte-identical to the injected state.

The AI Matic axis scores are GoML editorial judgment applied to that evidence, and are labelled as such. The derivation column states what supports each score. Anyone reweighting the axes for their own workload will get a different total, which is the intended use.

Sources

  • OpenAI, GPT-6 Astra: A new generation of intelligence, launch post, published benchmark tables and footnotes, 2026.
  • OpenAI, GPT-6 Astra system card and safety update.
  • OpenAI API pricing and availability documentation, including Azure and AWS Bedrock.
  • Artificial Analysis, Intelligence Index v4.1.1 and Coding Agent Index v1.4, as reproduced in OpenAI's comparison table.
  • GoML internal, LatticeBench resultastra.md and astrareport.md (run 4), scored against manifest.json (57 issues).
  • GoML internal, Gemini 3.7 Flash vs 3.8 Flash and Muse Spark 1.3 AI Matic assessments, for the cross-model reference.