Back

LLM testing of Gemini 3.8 Flash: Scored on the AI Matic Bench

Sarankumar S

September 4, 2026
Table of contents

After doing LLM testing of Gemini 3.8 Flash against Gemini 3.7 Flash, we found that the newer model gives enterprise teams a stronger case for migration. Google kept the pricing, context window, supported modalities, and model ID pattern unchanged, so benchmark performance becomes the main point of comparison.

We evaluated both models using GoML’s AI Matic Bench Score, which measures enterprise delivery readiness using published benchmark results and GoML’s LatticeBench evaluation.

Gemini 3.8 Flash scored 2.7 points higher overall. Our LLM testing shows where those gains came from and what LatticeBench found beyond the publicly reported benchmarks.

Metric 

3.7 Flash 

3.8 Flash 

AImatic Bench Score (weighted, out of 100) 

75.2 

77.9 

Google-published benchmark rows improved 

baseline 

14 of 14 

LatticeBench bug-mode discovery, unprompted 

8 of 49 (16%) 

14 of 49 (29%) 

LatticeBench critical-severity issues found 

2 of 18 (11%) 

6 of 18 (33%) 

LatticeBench very_hard tier found 

1 of 7 

2 of 7 

LatticeBench organic defects surfaced in golden/ 

7 

21 

List price, per 1M input / output tokens 

$0.75 / $3.75 

$0.75 / $3.75 

Measured cost per Artificial Analysis index task 

$0.40 

$0.58 (+45%) 

Output tokens per DeepSWE v1.1 run 

107,000 

143,000 (+34%) 

Agent steps per DeepSWE v1.1 run 

125 

166 (+33%) 

Independent labs corroborating the coding gain 

n/a 

2 (Artificial Analysis, DeepSWE leaderboard) 

Unreconciled conflicts across published sources 

n/a 

4 (registered below) 

Fabricated or unverifiable claims in this note 

0 

0 

The one-line read

Gemini 3.8 Flash improves long-horizon performance, but it also costs more to run in practice even though the listed price stays the same. That trade-off explains why its weighted score rises by 2.7 points rather than showing a much larger jump.

What LLM testing of Gemini 3.8 Flash shows

Gemini 3.8 Flash keeps most of the same specifications as 3.7 Flash. It supports 1,048,576 input tokens and 65,536 output tokens. Its knowledge cut-off reaches March 2026, while some areas only include information up to January 2025. The model accepts text, images, and PDFs as input and returns text output.

Google positions Gemini 3.8 Flash as an update built on 3.7 Flash. The 3.7 model card still carries details on architecture, training data, and hardware. This points to a post-training and agent-loop update rather than a new base model.

Two changes matter in LLM testing of Gemini 3.8 Flash. Google removed the MINIMAL thinking level, so the model now supports LOW, MEDIUM, and HIGH thinking levels only.

The second change is pricing. The introductory rates stay in place until December 31, 2026. After that, input pricing moves from $0.75 to $1.50 and output pricing from $3.75 to $7.50. Context caching pricing also doubles.

LLM testing of Gemini 3.8 Flash published benchmark scores

Google published the evaluation results as an image inside a PDF, so many launch articles quote only a few benchmark numbers instead of the full table.

For this LLM testing of Gemini 3.8 Flash, we transcribed the complete table from Google’s evaluation document and checked it against two independent transcriptions. All three versions matched row by row.

Figure 1. Google-reported scores across the 13 percentage-scored rows. GDPVal-AA v2 is Elo and is excluded here

Benchmark 

What it measures 

3.7 

3.8 

Delta 

DeepSWE v1.1 

Long-horizon software engineering 

65.3% 

73.7% 

+8.4 

GDPVal-AA v2 

Knowledge work (Elo) 

1,482 

1,545 

+63 

Vals Finance Agent v2 

Financial analyst tasks 

59.0% 

61.4% 

+2.4 

Harvey Legal Agent 

Complex legal workflows 

8.8% 

10.0% 

+1.2 

Terminal-Bench 2.1 

Agentic terminal coding 

85.8% 

89.4% 

+3.6 

Terminal-Bench 4.0 

General agent capability 

11.2% 

19.1% 

+7.9 

GDP.PDF 

Expert PDF comprehension 

34.0% 

35.0% 

+1.0 

CharXiv Reasoning 

Synthesis from complex charts 

84.5% 

86.2% 

+1.7 

LVBench 

Long video understanding 

85.4% 

87.8% 

+2.4 

HLE-Verified 

Multidisciplinary expert reasoning 

53.6% 

54.9% 

+1.3 

OSWorld-2.0 

Agentic computer use 

50.6% 

59.0% 

+8.4 

BioMysteryBench (solvable) 

Bioinformatics research 

87.1% 

88.8% 

+1.7 

BioMysteryBench (difficult) 

Bioinformatics research 

43.5% 

56.5% 

+13.0 

LABBench2 

Real-world biology tasks 

82.1% 

86.2% 

+4.1 

The sweep is real. It is also thin in most places. Nine of the fourteen rows move by under 4.2 points, which is inside the range a harness change or a grader revision can produce on its own. Four rows carry the release.

Figure 2. Sorted deltas. The four solid bars are the release; the rest is noise-adjacent movement.

The strongest gains appear on longer tasks that require the model to keep working across many steps. Repository-scale engineering, computer use, and the difficult research split improve the most. Bounded single-pass reasoning changes much less.

For model selection, absolute scores matter more than relative gains. Terminal-Bench 4.0 rises from 11.2% to 19.1%, while Claude Opus 5 reaches 51.8% on the same chart. OSWorld-2.0 reaches 59.0% for Gemini 3.8 Flash and 75.4% for Opus 5. A large percentage gain from a low starting point does not always change the model choice for open-ended computer-use workloads.

Source conflicts we could not reconcile

We found four numeric conflicts across Google’s own pages and launch coverage. We keep both reported figures instead of averaging them, because an average would mix different runs or configurations.

# 

Conflict 

Status 

1 

Terminal-Bench 2.1 reads 89.4% vs 85.8% in the evaluation chart and 90.8% vs 81.6% in the Gemini Enterprise developer guide. 

Unresolved. Different published surfaces or configurations. At least one widely syndicated write-up carries the second pair without noting the first. 

2 

DeepSWE v1.1 for 3.8 Flash is reported as 71.0% across several aggregators; the evaluation PDF says 73.7%. 

The PDF figure is the primary source. The distinction matters: 73.7% against Opus 5's 74.0% is a tie, 71.0% is not. 

3 

3.7 Flash baselines were restated between its own launch package and the new side-by-side chart: GDPVal-AA v2 1,525 to 1,482, OSWorld 47.9% to 50.6%. 

Unresolved. We use the new paired chart for deltas so both columns come from one run set. 

4 

Artificial Analysis Intelligence Index for 3.8 Flash is quoted at 59 (v4.1.1) and at 58.7 elsewhere; separately, a developer-docs Humanity's Last Exam pair (45.4% vs 45.7%) is being reported as if it were the HLE-Verified row (54.9% vs 53.6%). 

Two different benchmarks and two index snapshots. Treat cross-source quotes on either as directional. 

Google explains more of its evaluation setup than many model vendors, but the cross-vendor numbers still need care. The OSWorld runs predate the benchmark team’s 08.08 patch, and the Opus 5 number came from a competitor blog rather than an in-house run. LVBench used 1,024 frames for Gemini and GPT models but only 300 for Claude because of API limits. Some expert-reasoning questions were also blocked by content filters for competitor models. These differences do not affect the Gemini 3.7 Flash versus 3.8 Flash comparison, but they make the Opus 5 and GPT columns useful as context rather than a direct ranking.

LLM testing of Gemini 3.8 Flash workload cost

Google designed Gemini 3.8 Flash to spend more effort on hard tasks. It takes more reasoning steps, makes more tool calls, and produces more output tokens at higher effort levels. Those extra tokens are billable because output pricing includes thinking tokens.

Figure 3. Same list price, measurably different workload cost. Left: dollar figures from the DeepSWE leaderboard and Artificial Analysis. Right: the effort increase behind them.

The public DeepSWE v1.1 leaderboard runs each model through the same mini-swe-agent harness. At high effort, Gemini 3.8 Flash scores 74% plus or minus 1, compared with 65% plus or minus 2 for 3.7 Flash. This supports Google’s reported +8.4-point gain. It also shows the extra work behind the score: 143,000 output tokens and 166 agent steps per run for 3.8 Flash, compared with 107,000 tokens and 125 steps for 3.7 Flash. Average run cost rises from $2.18 to $2.36, while the pass rate rises from about 65% to 74%. Cost per successful run therefore falls from about $3.35 to $3.19. On DeepSWE, the extra effort pays off.

Artificial Analysis shows a different cost pattern across a broader task mix. Gemini 3.8 Flash gains three Intelligence Index points, from 56 to 59 on v4.1.1, while cost per task rises about 45%, from $0.40 to $0.58. Wall-clock time moves from 2.2 to 2.5 minutes. Throughput improves from about 279 to 305 output tokens per second, but time to first token stays above 13 seconds. This profile fits batch and agent workloads better than latency-sensitive chat.

What the cost means for delivery

Cost per accepted result gives the clearest answer. Cost per token makes the two models look the same. Cost per benchmark task makes 3.8 Flash look 45% more expensive. Cost per successful DeepSWE run makes it about 5% cheaper. For an enterprise pipeline, the last measure is the one closest to the outcome being purchased.

GoML's own bench: LatticeBench

Public benchmarks show how the model performs on published test sets. LatticeBench adds GoML’s own enterprise codebase to the comparison. It contains identity, sales, SLA, engagement, organization, file store, work management, gateway, and AI copilot services, with 57 injected defects, including 18 critical-severity issues, plus 8 task prompts. We match each model claim against manifest.json by file path, line number, and description. Claims outside the manifest are checked against the untouched golden/ tree to confirm that the defect already existed. The test focuses on failure types that matter in enterprise software, including cross-tenant mass assignment, fail-open authorization, and shared-module blast radius.

LatticeBench attribution notes

Neither RESULTS file names the model under test. We map run 1 to Gemini 3.7 Flash and run 2 to Gemini 3.8 Flash from the run order. The table below keeps the labels as run 1 and run 2, so a corrected model mapping would change the labels rather than the scores. This mapping should be confirmed before external publication.

Figure 4. LatticeBench hit rates by run. Run 2 nearly doubles unprompted discovery and triples critical finds, and scores zero on the explicit task list.

Measure 

Run 1 

Run 2 

Claims submitted 

22 

35 

Bug-mode issues found, no hint given 

8 of 49 (16%) 

14 of 49 (29%) 

Critical-severity found 

2 of 18 (11%) 

6 of 18 (33%) 

very_hard tier found 

1 of 7 (BUGS-004) 

2 of 7 (BUGS-004, COMMON-003) 

Task-mode completed 

7 of 8 (88%) 

0 of 8 (not addressed) 

Total manifest coverage 

15 of 57 (26%) 

14 of 57 (25%) 

Organic defects surfaced outside the injected set 

7 

21 

Fabricated or hallucinated claims 

0 of 22 

0 of 35 

What improved in LatticeBench

Run 2 finds every issue found by run 1 plus six more, so the extra findings form a strict superset rather than a different sample. Critical-severity discovery rises from 11% to 33%. Run 2 also catches the SQL injection in the handoff service, disabled JWT audience verification, the revoked-share bypass, and hardcoded AWS credentials that run 1 missed. Unprompted discovery rises by 12.6 points on the same fixed corpus, which supports the DeepSWE direction on code Google has not seen.

Claim reliability stays high in both runs. Across 57 claims, every reported issue was verified as a real defect. Run 1 produced zero fabrications across 22 claims, and run 2 produced zero across 35. For audit work, avoiding false positives matters because every incorrect finding creates review time.

What did not improve

Both runs still miss five of the seven very_hard issues. The remaining misses include multi-file mass-assignment chains, fail-open authorization checks, and shared common/ permission issues that spread across services. These cases require the model to trace logic across two or three files instead of finding a local defect in one place.

The main LatticeBench finding

LLM testing of Gemini 3.8 Flash shows a clear limit in multi-file reasoning. The model improves on long-horizon work, but both runs still miss every cross-tenant mass-assignment chain in the test set. These tasks require the model to hold a schema change in memory, follow it through a facade into a service, and then connect an unfiltered setattr loop to a tenant-boundary failure. This is a common failure pattern in multi-tenant SaaS systems, so human review should remain in this part of the workflow.

Two open questions limit the comparison. Run 2 does not address the 8 explicit task prompts, which is why total manifest coverage falls slightly even though bug discovery improves. The RESULTS2 note does not confirm whether those prompts were withheld or ignored, so the drop from 87.5% to 0% in task mode should not be treated as a model regression yet. Neither run also supplied a patched tree, so scoring uses prose claims plus source verification rather than a diff against golden/ for every manifest entry.

The oracle limitation

Together, the two runs found 27 real defects in golden/ that were not part of the injected set. They include a cross-tenant IDOR on team detail, sales analytics queries without organization filters, audit and activity endpoints missing tenant checks on one branch, timesheet approval protected only by IsAuthenticated, and an import that breaks epic enrichment. We verified each issue in the code and confirmed that it existed before the injected defects.

This creates a scoring issue for LatticeBench rather than a model issue. golden/ cannot act as a clean oracle while it still contains 27 untracked defects. Future runs could receive inconsistent credit for finding real bugs that have no manifest entry. Before the next run, GoML should patch those confirmed issues in golden/, bench/, and problem-set/ so the same baseline applies everywhere. Until then, treat the bug-mode hit rates as a lower bound.

LLM testing of Gemini 3.8 Flash using the AI Matic Bench Score

Vendor benchmark tables answer whether a new model scores higher. GoML needs a different answer: does changing the model improve accepted outcomes per rupee and per hour in a running enterprise pipeline without forcing a full requalification cycle? The AI Matic Bench Score v0.2 measures that question across nine axes. Each axis receives a 0 to 10 score from published evidence, independent testing, or GoML’s LatticeBench runs.

Axis 

Weight 

What it scores 

Long-horizon completion 

18% 

Does the model finish multi-step work end to end without a human unblocking it 

Agentic tool reliability 

13% 

Tool calls, recovery from failed paths, computer and terminal control 

Workload unit economics 

18% 

Cost per accepted result, not cost per token 

Claim reliability under audit load 

12% 

Fabrication rate when the model reports findings a human will act on 

Latency profile 

9% 

Time to first token, throughput, wall clock per task, fit to interactive vs batch 

Deployment envelope 

9% 

Context, modalities, API surface stability, migration cost from the incumbent 

Regulated-domain depth 

9% 

Finance, legal and scientific task performance where review is mandatory 

Safety and governance 

7% 

Published safety deltas, refusal behaviour, multilingual behaviour, auditability 

Evidence quality 

5% 

How much of the claim set is independently reproduced rather than vendor reported 

How much of the claim set is independently reproduced rather than vendor reported  

The weighting puts the most emphasis on long-horizon completion and workload unit economics, which together account for 36% of the score. Claim reliability carries 12% because a fabricated audit finding creates review work. Evidence quality carries 5% because it changes confidence in the other scores rather than measuring workload performance on its own.

LatticeBench supplies much of the evidence for three axes: all of claim reliability, part of long-horizon completion alongside DeepSWE, and part of agentic tool reliability alongside terminal and computer-use benchmarks. In total, 43% of the weighted score uses evidence generated by GoML as at least one input.

Figure 5. Axis scores across the nine-axis v0.2 rubric. Five axes improve, four regress.

How the score is calculated

Each axis is scored 0 to 10 against the band anchors below. The total is the weighted sum of the axis scores, multiplied by ten to put it on a 100-point scale.

Formula

AImatic Bench Score  =  ( sum of ( axis score x axis weight ) ) x 10,  where weights sum to 1.00 and each axis score is on a 0 to 10 band.

Band 

Meaning 

Test we apply 

9 to 10 

No concerns on this axis 

We would put it in a client proposal with no caveat attached 

7 to 8 

Production-viable with named guardrails 

Ships if a specific control is in place, and we can name the control 

5 to 6 

Usable under supervision 

Needs a human review gate on every output that reaches a client 

3 to 4 

Pilot only 

Fine for internal exploration, not for delivery 

0 to 2 

Disqualifying 

This axis alone rules the model out for the engagement class 

This axis alone rules the model out for the engagement class  

For long-horizon completion, Gemini 3.8 Flash has three supported gains: DeepSWE improves by 8.4 percentage points and is independently reproduced at 74% versus 65%, BioMysteryBench difficult improves by 13.0 points, and LatticeBench unprompted discovery rises from 16% to 29%. The limit remains the hardest multi-file tier, where the two runs move only from 1 of 7 to 2 of 7 issues found. That places 3.8 Flash at 8.0 for this axis, compared with 6.5 for 3.7 Flash. Human review should still cover multi-file authorization paths.

The score bands still need tighter calibration. They describe the level of confidence but do not yet define a fully repeatable scoring procedure. Two engineers could read the same evidence and land about one point apart on an axis. Until GoML runs a blind inter-rater check, treat the total as accurate to about plus or minus one point.

Scoring and derivation

Axis 

Wt 

3.7 

3.8 

Basis for the delta 

Long-horizon completion 

18% 

6.5 

8.0 

DeepSWE +8.4pp reproduced; LatticeBench discovery +12.6pp; very_hard tier only 1/7 to 2/7 

Agentic tool reliability 

13% 

5.0 

6.5 

TB 4.0 and OSWorld both +8pp but far behind Opus 5; LatticeBench critical finds 11% to 33% 

Workload unit economics 

18% 

8.5 

7.5 

Same list price, +45% cost per index task; cost per DeepSWE pass slightly better 

Claim reliability under audit load 

12% 

9.0 

9.5 

LatticeBench: zero fabrications in 22 claims and zero in 35; 3.8 held it at 59% more claims 

Latency profile 

9% 

8.0 

7.0 

Throughput up 9%, wall clock up 14%, TTFT above 13s 

Deployment envelope 

9% 

9.0 

8.0 

Identical specs and stable ID, but MINIMAL thinking level removed 

Regulated-domain depth 

9% 

7.0 

8.0 

Finance +2.4pp, legal +1.2pp, LABBench2 +4.1pp; legal absolute still 10.0% 

Safety and governance 

7% 

7.5 

7.0 

Text-to-text +0.4pp, tone +0.2pp, refusals worse by 1.1pp, multilingual worse by 5.4pp 

Evidence quality 

5% 

8.0 

9.0 

Two independent labs plus a goML-run harness; four source conflicts remain open 

Weighted total 

100% 

75.2 

77.9 

Net +2.7 on a 100-point scale 

The total score matters less than the shape of the scorecard. Five axes improve and four fall. Long-horizon completion gains 1.5 points while workload unit economics loses 1.0, and both carry an 18% weight, so much of the difference cancels out. Smaller gains in claim reliability, regulated-domain depth, and evidence quality move the result to 77.9 from 75.2. The net change is +2.7 points.

What the AI Matic Bench Score does not cover

  • LatticeBench scores prose claims rather than patched code. No diff-based rescore against golden/ has been run for each manifest entry, and golden/ still contains 27 confirmed defects outside the injected set.
  • The LatticeBench run-to-model mapping comes from run order rather than an explicit model label in the results files. Run 2 also has an unresolved task-mode configuration question.
  • We did not run multilingual or Indic-language tests. Google’s own multilingual score falls by 5.4 points, so teams using non-English workloads should test them separately.
  • We did not run retrieval-grounded tests. None of the published benchmark rows isolates RAG performance.
  • The score bands are descriptive and have not yet gone through a blind inter-rater check. Treat each axis score as about plus or minus one point.
  • The weights reflect GoML’s delivery priorities. A latency-sensitive consumer product could place more weight on latency, which would favour Gemini 3.7 Flash.

The GoML position

Move to Gemini 3.8 Flash when

  • Your workload uses repository-scale coding, multi-step terminal automation, or agent loops that often stop before finishing the task.
  • You measure success by accepted results and can absorb about a third more tokens and turns when the completion rate improves.
  • Your pipeline runs in batch or asynchronously, so a time to first token above 13 seconds has little effect on the end user.
  • You are starting a new Gemini workload. The model ID remains stable and the current list price matches 3.7 Flash.
  • You keep human review for multi-file authorization paths. Neither run catches the cross-tenant mass-assignment cases in LatticeBench.

Stay on Gemini 3.7 Flash when

  • Your pipeline is already validated, runs at high volume, and depends on low latency. Classification, extraction, and routing show little reason to pay for the extra effort.
  • Output-token or tool-call limits are fixed.
  • You depend on the MINIMAL thinking level, which Gemini 3.8 Flash no longer accepts.
  • Non-English performance matters to the contract and you have not retested against the reported 5.4-point multilingual drop.

Migration steps after LLM testing of Gemini 3.8 Flash

Do not change the model and the traffic split in the same release. Shadow Gemini 3.8 Flash against production 3.7 Flash traffic by task class. Log reasoning tokens, visible output tokens, tool calls, retries, accepted-result rate, and p95 latency. Promote only the task classes where the higher accepted-result rate covers the extra token use. Start at LOW or MEDIUM thinking and reserve HIGH for tasks that show a clear need for it. Use the January 2027 pricing in annual cost models rather than the introductory rate.

The larger lesson from this LLM testing of Gemini 3.8 Flash is that model choice needs regular retesting. Google shipped three Flash releases in 43 days, with each one improving most published scores. Teams should keep the model behind an interface, keep one evaluation harness, and rerun the comparison as the Flash tier changes.

How GoML verified LLM testing of Gemini 3.8 Flash

Every number in this comparison traces to a named source. We checked the 14-row Google table against two independent transcriptions of the evaluation PDF, and all three copies matched row by row. Independent figures come from the public DeepSWE v1.1 leaderboard and Artificial Analysis Intelligence Index v4.1.1, both of which identify their harness and version. When sources disagree, we show both figures instead of averaging them.

LatticeBench numbers come from the two scored run reports. Each unmatched claim is checked against the current golden/ file at the cited location to confirm that the defect is real and appears in both golden/ and bench/. Where a run report corrected an earlier count, we use the corrected number. The critical-severity total is 18, not 16. Neither run supplied patched code, so both are scored from prose claims plus source checks.

The AI Matic platform axis scores are GoML’s editorial scoring of the evidence, not direct measurements. The derivation column shows why each score moved. Teams that apply different weights to the nine axes will get a different total, which is expected.

Frequently asked questions

1. Did task-mode completion really fall from 88% to 0%?

Not necessarily. Run 2 did not address the eight task prompts, so the 0% score does not prove a real regression.

2. Do better bug-finding scores mean it is ready for security-critical review?

Only partly. Bug discovery improved and critical findings tripled, but both runs still missed five of the seven hardest multi-file issues. Human review is still needed.

3. Why do benchmark scores differ across sources?

Different test setups are often reported together. GoML uses Google’s main evaluation PDF as the main reference and treats conflicting figures as directional.