The Frontier Is a Three-Way Tie — With Very Different Winners
As of August 2026, the gap between the three leading proprietary models has effectively closed. On Artificial Analysis’s Intelligence Index — a composite of nine independent benchmarks — Claude Fable 5 scores 62, while Grok 4.6 and GPT-5.6 Sol are tied at 61. That one-point difference is within noise. What’s not noise is how sharply the three diverge once you move past the headline number into specific tasks.
Grok 4.6 shipped on August 12, 2026, and xAI’s positioning was straightforward: frontier intelligence at cost-efficiency that neither OpenAI nor Anthropic can match. That claim holds in some areas and breaks down in others. Here’s what the independent benchmarks actually show.
Benchmark Results: Where Each Model Wins
The most useful frame is not “which model is best” but “best at what.” These three models have genuinely different strengths, and the benchmarks reveal a clean separation.
Coding and Software Engineering
GPT-5.6 Sol leads on pure software engineering tasks. On Artificial Analysis’s Coding Agent Index, Sol scores 80 — 2.8 points ahead of Claude Fable 5’s 77.2. Grok 4.6 loses both on terminal-based coding (Terminal-Bench) and deep SWE tasks (DeepSWE). On CursorBench v3.2, the ranking flips slightly for IDE-integrated work: Fable 5 leads at 70.5%, Grok 4.6 is close behind at 69.9%, and Sol trails at 67.2%. The pattern suggests Sol is the better choice for pure CLI/agent coding pipelines, while Fable 5 edges it in editor-integrated workflows.
Knowledge Work and Legal Reasoning
Grok 4.6 wins the knowledge-work cluster: GDPval-AA, AA-Briefcase, and the Harvey legal benchmark all go to xAI’s new model. The Harvey result is particularly notable — legal reasoning has historically been an area where models trained on broad instruction-following outperform those tuned for code. If your use case spans contract review, policy analysis, or research synthesis, Grok 4.6 is currently the strongest option at any price.
Agentic Task Completion
Composio’s agentic evaluation — 47 multi-step tasks requiring tool use, memory, and error recovery — shows Claude Fable 5 completing all 47. GPT-5.6 Sol completed 45. Grok 4.6’s score on this evaluation was not independently published at time of writing, but xAI’s own agentic evals place it ahead of Grok 4.5 on every measured dimension. Fable 5’s 100% completion rate is consistent with Anthropic’s stated focus on reliability over raw speed in long-horizon tasks.
The Comparison Matrix
| Benchmark | Claude Fable 5 | GPT-5.6 Sol | Grok 4.6 |
|---|---|---|---|
| AA Intelligence Index | 62 | 61 | 61 |
| Coding Agent Index (AA) | 77.2 | 80.0 | n/a (loses SWE) |
| CursorBench v3.2 | 70.5% | 67.2% | 69.9% |
| Harvey Legal | — | — | Wins |
| Composio Agentic (47 tasks) | 47/47 | 45/47 | not published |
| Input price (per 1M tokens) | $10 | $5 | $2 |
| Output price (per 1M tokens) | $50 | $30 | $6 |
| Context window | ~1M tokens | ~1M tokens | 500K tokens |
Grok 4.6 pricing doubles for prompts over 200K tokens. All scores from Artificial Analysis as of August 2026.
Pricing: Grok 4.6 Is Not a Close Race
Grok 4.6 costs $2 per million input tokens and $6 per million output tokens — for prompts under 200,000 tokens. At that rate it is 2.5x cheaper than Sol on inputs and 5x cheaper on outputs, and against Fable 5 the gap is 5x on input and over 8x on output. For reasoning-heavy workloads where output tokens dominate cost, the difference compounds fast.
The catch: above 200K tokens, xAI’s pricing doubles across the board to $4 input / $12 output. That still undercuts both competitors, but closes the gap meaningfully for long-context document work. Claude Fable 5 and GPT-5.6 Sol both offer roughly 1M-token context windows with stable pricing; Grok 4.6’s 500K context ceiling is also a real constraint for the longest document tasks.
For teams running high-volume inference — product search, document classification, research pipelines — Grok 4.6’s price point is a structural advantage that benchmark scores alone won’t offset.
Who Should Use What
Claude Fable 5 is the choice for long-horizon agentic work and editor-integrated coding where reliability matters more than raw throughput. Its 100% Composio completion rate and CursorBench lead make it the defensible pick for engineering teams where a dropped task costs more than the compute price difference.
GPT-5.6 Sol is the best pure SWE model in this group. If your workflow is terminal-heavy, CLI agents, or automated PR pipelines, Sol’s 80-point Coding Agent Index score is the relevant number. It’s cheaper than Fable 5 and still meaningfully stronger on SWE-bench-class tasks than Grok 4.6.
Grok 4.6 wins on cost and knowledge-work benchmarks. Teams running legal document analysis, research synthesis, or any workload where the Harvey-type evals matter should test it seriously. At under $10/million combined token cost for typical prompts, the price-to-intelligence ratio is currently unmatched at this capability tier.
The real takeaway from August 2026 is that “best frontier model” is no longer a meaningful question. The Artificial Analysis Intelligence Index spread across these three models is 1 point. The only question that matters for your budget is which benchmark cluster maps to your actual workload.
Further Reading
- Grok 4.6 returns to the intelligence frontier — Artificial Analysis — the independent benchmark breakdown behind xAI’s launch claims.
- GPT-5.6 Sol vs Claude Sonnet 5: July 2026 Benchmarks — vortx.ch — how Sol performed one tier down before this comparison.
- Claude Fable 5 and Mythos 5: The First US AI Export Ban — vortx.ch — context on how Fable 5 fits into Anthropic’s model line and export restrictions.

