Back

LLM testing of Grok 4.6: A cost-curve event, not a capability

Sarankumar S

August 14, 2026
Table of contents

Our team ran the LLM testing of Grok 4.6, and in this article we show just how much SpaceXAI has improved performance without raising its token prices. On the Artificial Analysis Intelligence Index, Claude Opus 5 and Fable 5 still hold the top two spots, while Grok 4.6 scores 61, matching GPT-5.6 Sol Max and sitting one point below Fable 5 Max.

Grok 4.6 gains five points over Grok 4.5 while keeping the same $2/$6 per million token pricing. Claude Opus 5 costs $5/$25, while GPT-5.6 Sol costs $5/$30. Artificial Analysis reports an average Grok cost of $0.84 per completed task.

For companies, the bigger question is where Grok 4.6 can cut costs without giving up too much performance. That opportunity exists, but it appears limited to a narrower set of workloads than early launch claims suggested.

The GoML position on Grok 4.6

Do not think of Grok 4.6 as a platform decision, but rather as a step in the process of routing. We recommend incorporating Grok into one of the tiers of an existing model-abstraction scheme and execute tests with it on your workload and cost-per-completed-task as metrics for comparison. Keep terminal-native and safety-critical agentic execution on your incumbent platform until there is evidence to the contrary.

What launched

Attribute 

Detail 

Release date 

12 August 2026, five weeks after Grok 4.5 

Model string 

grok-4.6 

Context window 

500,000 tokens — unchanged from Grok 4.5 

Standard pricing 

$2.00 / 1M input · $0.50 / 1M cached input · $6.00 / 1M output 

Long-context pricing 

At or above a 200K-token prompt, rates double to $4 / $1 / $12 and apply to every token in the request 

Fast variant 

Approximately twice the standard rate 

Measured throughput 

~67.6 output tokens/sec; time-to-first-token ~42s (Artificial Analysis, high effort) 

Availability 

xAI API, Grok Build, Cursor, OpenRouter, Vercel, Cloudflare 

Training approach 

Extended supplemental pre-training, SFT trajectories regenerated by Grok 4.5, agentic RL across coding, knowledge work and domain environments 

Third-party reporting describes Grok 4.6 as a ~1.5-trillion-parameter model on the same underlying foundation as Grok 4.5. SpaceXAI has not confirmed this; treat it as unverified and do not build capacity assumptions on it.

LLM testing of Grok 4.6 benchmark table

The reported results point to a mixed outcome. Grok 4.6 performs well on knowledge-heavy work and longer business tasks, but it falls behind on advanced software engineering and terminal-based execution. That makes its fit more dependent on the type of work being assigned rather than its overall score alone.

Evaluation 

Grok 4.6 High 

Grok 4.5 High 

GPT-5.6 Sol Max 

Fable 5 Max 

AA Intelligence Index 

61 

56 

61 

62 

GDPVal-AA v2 

1753 

1526 

1728 

1741 

CursorBench v3.2 

69.9% 

66.7% 

67.2% 

70.5% 

DeepSWE v1.1 

65.9% 

54% 

73% 

70% 

FrontierCode v1.1 (Ext.) 

61.3% 

56.6% 

60.6% 

64.9% 

APEX-Agents 

57.5% 

47.1% 

56.7% 

59.2% 

Terminal-Bench v3.0 

26% 

15.7% 

34.6% 

34.1% 

APEX-SWE 

56.4% 

53.6% 

 

58.8% 

AA-Briefcase 

1577 

1313 

1502 

1574 

Harvey LAB (Vals) 

15.8% 

12.9% 

2.5% 

11.3% 

Source: SpaceXAI launch table. Competitor figures are a mix of self-reported and publicly available results — a mixed-provenance table, not a controlled head-to-head.

Four readings that matter

  • Knowledge work is the strength. Grok 4.6 posts the highest published figures on GDPVal-AA v2 (1753), AA-Briefcase (1577) and Harvey LAB (15.8%). These are long-horizon, document-heavy, business-analysis tasks precisely the shape of most enterprise knowledge work.
  • Terminal execution is the weakness. Terminal-Bench v3.0 at 26% against 34.6% and 34.1% is a ~25–30% relative deficit. If your agents live in a shell  CI remediation, infrastructure operations, migration tooling this gap will show up as retries, and retries are where the price advantage evaporates.
  • The generational delta is the real signal. Terminal-Bench moved 15.7% → 26%. DeepSWE moved 54% → 65.9%. APEX-Agents moved 47.1% → 57.5%. Our LLM testing of Grok 4.5 recorded the baseline these numbers are measured against, and Grok 4.6 improves on every row SpaceXAI reports. Position matters less than slope when you are planning a two-year roadmap.
  • Legal remains unsolved. Every model on this table is failing Harvey LAB. 15.8% is the best score and it is still a failing grade. Regulated professional-services work is not solved by any frontier model, and any vendor claim to the contrary should be treated as marketing.

The FrontierCode v1.1 (Extended) score for Fable 5 Max is listed as 64.9% in the launch table graphic but 63.6% in the accompanying text. The difference is small, but anyone citing the figure externally should verify it against the latest published version.

There is also some confusion around the 1753 score. Some coverage has linked it to LMSYS Chatbot Arena, but the figure belongs to GDPVal-AA v2. Keep those two benchmarks separate when citing the result.

The number the matters

The cost per million tokens is a metric used for procurement purposes while the cost per completed task is a metric used for engineering purposes. These two metrics are in constant opposition, and the extent of this opposition is the scale of the budget variance.

According to estimates from Artificial Analysis, Grok 4.6 is priced at $0.84 per task but this pricing is based on both the main price and the token efficiency, i.e., the number of tokens required for task execution. A model that comes up with long-winded explanations or retries three times cannot be considered less expensive than the other models.

Metric 

What Finance Sees 

What Engineering Must Measure 

Unit price 

$2 / $6 per million tokens 

Effective blended rate after long-context banding and cache hit rate 

Volume 

Tokens consumed per month 

Tokens consumed per successfully completed task, including retries 

Comparison 

60%+ below Opus 5 and GPT-5.6 Sol sticker price 

Delta in task success rate; a 10-point accuracy gap can erase a 60% price gap 

Latency 

Not on the invoice 

~42s TTFT and ~68 tok/s — disqualifying for interactive UX, irrelevant for batch 

Three things the launch post does not foreground

The long-context billing cliff

The $2 / $6 rate applies below a 200,000-token prompt. At or above that threshold, rates double to $4 / $12 and the higher rate applies to the entire request, not the excess. A 500K-token context window is available; using it costs double across the board.

The implication is architectural, not commercial: the win with Grok 4.6 is not stuffing 500K tokens into every prompt. It is running a stronger long-horizon model with disciplined retrieval, aggressive prompt caching at $0.50 per million, and periodic context compaction. Teams that treat the large window as an excuse to skip retrieval engineering will pay for it twice in tokens and in accuracy

Latency is not a footnote

At around 42 seconds under high reasoning effort, Grok 4.6's time-to-first-token is astronomical compared to similar models. In the case of asynchronous agents, night-time batch-processing and prolonged studies, this should not pose a problem. However, for most human-facing uses, including copilots, chat interfaces and in-product assistants, times this high are unacceptable without a backup plan. Optimize your workloads before optimizing your budgets.

Self-generated training data

SFT trajectories for Grok 4.6 were regenerated using Grok 4.5 and filtered with model-based checks. This is now standard practice across frontier labs and it demonstrably works. It also concentrates a model family’s failure modes: errors the teacher makes systematically are the errors the filter is least likely to catch. For enterprises running multi-model ensembles specifically to get error decorrelation, this is a reason to keep providers genuinely diverse rather than assuming two strong models fail independently.

How enterprises should read about LLM testing of Grok 4.6

  1. Do not migrate. Route. Single-model standardisation is now the expensive choice. Put a model-abstraction layer between your applications and any provider, then route by task class. Grok 4.6 becomes a tier in that router, not a replacement for it. The abstraction layer is the durable asset; the model behind it has a five-week half-life, as this release demonstrates.
  1. Re-run your own evals, on your own workload. The published table mixes self-reported and public competitor numbers. It tells you nothing about your document formats, your domain vocabulary, or your tool schemas. The only defensible input to a routing decision is one real task, one harness, held constant, measured on cost per completed job. Budget one engineering week per candidate model and make this a standing capability, not a one-off.
  1. Segment by latency tolerance before you segment by price. Batch and asynchronous workloads, such as nightly document processing, research agents, report generation, and bulk enrichment, are where the $0.84-per-task cost can translate into meaningful savings. Interactive applications still need a faster model tier, regardless of intelligence scores. Separate these workload types first, since that tells you how much of your AI spending can realistically be reduced.
  1. Keep terminal-native and safety-critical execution on the incumbent. A 26% Terminal-Bench score is a real constraint for autonomous infrastructure work. Until your own evaluation contradicts it, keep shell-executing, permission-holding agents on the model with the stronger execution profile. The cost differential on that slice of workload is small; the blast radius of a failed remediation is not.
  1. Engineer context, do not buy it. The 200K billing threshold rewards retrieval discipline directly. Invest in chunking quality, reranking, prompt caching and context compaction. These reduce cost on every provider simultaneously and survive every model swap which is a better return than any single-vendor pricing advantage.
  1. Renegotiate, and structure for optionality. Three credible providers within one index point of each other is the strongest procurement position enterprise buyers have had. Use it. Avoid multi-year single-vendor commitments; prefer contracts that permit reallocation of committed spend across model tiers as the frontier moves.

Where LLM testing of Grok 4.6 fits in a multi-model stack

Workload Class 

Recommendation 

Rationale 

Long-horizon knowledge work, research, document analysis 

Strong candidate 

Leads on GDPVal-AA, AA-Briefcase and Harvey LAB; latency is absorbed by asynchronous execution 

Bulk batch processing and enrichment 

Strong candidate 

Price advantage compounds directly with volume; TTFT is irrelevant off the critical path 

IDE-assisted and human-in-the-loop coding 

Test in parallel 

CursorBench 69.9% is competitive; validate against your repository and language mix 

Autonomous SWE and terminal-native agents 

Hold on incumbent 

DeepSWE and Terminal-Bench gaps are material; retry cost erodes the price advantage 

Interactive copilots and chat surfaces 

Not recommended 

~42s TTFT at high effort; requires a lower-latency tier 

Regulated professional-services workflows 

Human-in-the-loop mandatory 

No frontier model clears a usable bar on Harvey LAB; design for review, not autonomy 

What GoML would test before committing

  • Run one representative task from each workload group through the same test harness for every model. Compare cost per completed task and the number of retries required.
  • Test tool-calling accuracy against your real schemas. This is often where benchmark results and production behavior start to differ.
  • Measure effective token cost using your actual prompt-size distribution and cache hit rate. This shows when the 200K-token threshold begins to affect cost.
  • Measure P50 and P95 end-to-end latency under concurrent load rather than relying on single-request time-to-first-token results.
  • Compare failure patterns with your current model. If both models fail in the same way, an ensemble may add less protection than expected.
  • Check data location, storage periods, and provider policies before comparing benchmark scores. For many enterprises, these requirements narrow the model shortlist early.

The bottom line

Grok 4.6 represents a trustworthy first-tier product, focusing on competitiveness primarily based on economic factors instead of technological advantages. This approach is indeed valid and could be more relevant than ranking-based comparisons based on index scores for commercial buyers.

The more important signal comes from the market rather than from any specific vendor. Among three firms competing in the same sector, within one quarter the difference between prices is about 2.5 times, and the data might change every five weeks. Any technology setup that leans on a single vendor carries real risk.

The companies that benefit most won't be the ones that pick the single best model. They will be the ones that build the abstraction layers and evaluation systems that make the right choice repeatable, launch after launch.

Talk to GoML

GoML builds agentic systems, RAG architectures, and multi-model routing on AWS using our propreitory AI Matic platform. It gives teams a robust working foundation for routing instead of forcing them to build the entire layer from scratch.

For terminal-based execution and safety-sensitive Claude workloads, our Claude agent development team handles migration, orchestration, testing and production hardening.

Frequently asked questions

How do you build a model-routing layer instead of choosing one model?

Place an abstraction layer between your application and LLM providers. Add a classification step that assigns each request to the right model tier, then define fallback rules for failures, timeouts, or poor responses. This requires an architectural change rather than a prompt adjustment. A reusable evaluation harness also makes future model changes easier, since you can update routing rules instead of rebuilding the application around a new provider.

How much should you trust a benchmark table that mixes vendor-reported numbers?

Use it as a reference, not as the final basis for a purchasing decision. The launch material itself contains inconsistencies, including two different FrontierCode scores for Fable 5 Max and confusion around the GDPVal-AA figure in outside coverage. Those discrepancies make independent testing more useful. Before committing budget, reproduce the results with your own tasks, prompts, tools, and traffic patterns.

If every leading model still performs poorly on Harvey LAB, is AI ready for legal or compliance work?

A 15.8% pass rate suggests these models are not ready to handle regulated legal work without human review. They can still assist with drafting, research, document review, and early analysis. The safer setup keeps a qualified reviewer between the model’s output and any decision or regulated action.