TL;DR
- Google launched Gemini 4 Argon (Sep 30) at introductory $2/$10 per MTok — one day after OpenAI’s GPT-6.1 Sol launched at identical rates. Anthropic’s Claude Sonnet 5.5 (Sep 28) also prices at $2/$10. Three labs now publish mid-tier frontier models at the same entry point.
- Argon leads DeepSWE v1.1 at 77.9% (vendor-reported by Google); it trails on Terminal-Bench 4.0 (57.4% vs Opus 5.5’s 66.4%) and Harvey Legal Agent (19.6%). No independent scores yet.
- H Company’s Holo4-27B claims 85.2% on OSWorld but scored only 49.3% on a held-out test set — 480 of 600 public tasks were in its training split.
Who should care: Engineering teams choosing a mid-tier frontier model for coding pipelines, agentic workflows, or computer-use automation.
Verdict: Test Gemini 4 Argon on your own coding tasks before the introductory price ends. Claude Sonnet 5.5 is a safe drop-in upgrade for existing Anthropic users today.
The $2/$10 Moment
Three separate announcements this week landed at the same price: OpenAI’s GPT-6.1 Sol on September 29, Google’s Gemini 4 Argon on September 30, and Anthropic’s Sonnet 5.5 on September 28 — all at $2 per million input tokens and $10 per million output tokens. That is not a coincidence. It is the frontier mid-tier settling at a commodity price.
The real competition is now in the benchmarks and the fine print. Argon’s introductory rate will rise to $4/$20 at an unannounced date. Argon’s context window reaches 2 million tokens with up to 1 million output tokens. Sonnet 5.5 caps output at 128K–300K. GPT-6.1 Sol sits between the two on context. The three models are not interchangeable, despite matching invoices.
This week also saw H Company’s Holo4 computer-use models, Perplexity’s first open-source decision model, and Claude Opus 5.5 (launched Sep 22, covered here Tuesday) consolidating its position above the $2/$10 tier.
This Week Ranked
| Rank | Model / Tool | Vendor | What changed | Key sourced metric | Verdict | Best for |
|---|---|---|---|---|---|---|
| 1 | Gemini 4 Argon | New launch; 2M context, 1M output; intro $2/$10 → $4/$20 | DeepSWE v1.1: 77.9% (vendor); CWE-bench: 68% (vendor) | Test it | Long-horizon coding; security scanning; tasks needing >1M context | |
| 2 | Claude Sonnet 5.5 | Anthropic | New launch; 30%+ faster, 30% lower cost vs Sonnet 5 (vendor); $2/$10 | Terminal-Bench 4.0: 70.6% (vendor); OSWorld 2.1: 80.1% (vendor) | Use it | Drop-in Sonnet upgrade; agentic coding pipelines; mixed coding + document tasks |
| 3 | GPT-6.1 Sol | OpenAI | Updated Sol; $2/$10 from launch; 52 Intelligence Index (Artificial Analysis) | AA Intelligence Index: 52/100 (highest in OpenAI lineup per AA) | Use it | OpenAI-native stacks; reasoning-heavy agentic tasks; Codex users |
| 4 | Holo4-27B + 35B-A3B | H Company | New launch; computer-use specialist; 35B Apache 2.0 | OSWorld: 85.2% (27B); AutomationBench held-out: 49.3% | Test it | Desktop/web/Android automation only; evaluate on your own task set first |
| 5 | Perplexity Decider v1 27B | Perplexity | First open-source model from Perplexity; Apache 2.0; routing focus | No independent benchmarks published at launch | Watch | Agent decision routing; teams wanting Apache 2.0 decision models |
1. Gemini 4 Argon
Google released Argon on September 30 with a split performance profile. On DeepSWE v1.1 — a benchmark measuring multi-step software engineering tasks — it scored 77.9%, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%). These figures come from Google; no independent evaluator has published a matching score yet. On Terminal-Bench 4.0, Argon trails: 57.4% against Opus 5.5’s 66.4% (per the same shattered.io source) and Sonnet 5.5’s 70.6% (Anthropic, via The Neuron). Its Harvey Legal Agent score of 19.6% is a clear weak spot for legal automation.
The headline specification is the 2 million token context window with up to 1 million output tokens — both market-leading for a mid-tier model. The introductory price of $2/$10 will increase to $4/$20 at an unannounced date. Teams building on Argon for long-context work should factor in the step-up in their cost models. Verdict: Test it — evaluate on your actual coding tasks before the intro period closes.
2. Claude Sonnet 5.5
Anthropic released Sonnet 5.5 on September 28 at $2/$10 per MTok with a 1 million token context window and up to 128K–300K output. Anthropic’s own measurements show 30%+ faster response and 30% lower task cost than Sonnet 5. Terminal-Bench 4.0 (which the Neuron digest attributes to Anthropic’s announcement) came in at 70.6%, the highest of the three $2/$10 models on that specific benchmark. OSWorld 2.1 score of 80.1% shows competitive computer-use capability, though below Holo4-27B’s claimed headline number.
All performance figures are vendor-reported. For teams already running Sonnet 5 in production, the drop-in upgrade case is strong: same pricing tier, same API endpoint pattern, measurably faster. Verdict: Use it — the safest swap this week for existing Anthropic users.
3. GPT-6.1 Sol
OpenAI updated Sol on September 29, launching GPT-6.1 Sol at the now-standard $2/$10 rate — one day before Google did the same with Argon, which the Yahoo Finance analysis described as a structural signal rather than a coincidence. On Artificial Analysis’s Intelligence Index, GPT-6.1 Sol scores 52 — the highest in OpenAI’s current published lineup on that index. Concrete task-level benchmarks comparing 6.1 Sol to the original Sol were not available at time of writing; this is a gap in OpenAI’s release documentation.
For teams running Codex or OpenAI-native agentic pipelines, the update is a straightforward model swap. Verdict: Use it — if you are already on Sol, this is an automatic upgrade with no price change.
4. Holo4-27B and Holo4-35B-A3B
H Company released two computer-use models on September 28. The 27B dense model scored 85.2% on OSWorld — the GUI automation benchmark — against Sonnet 5.5’s 80.1% and prior leaders below 80%. The 35B mixture-of-experts variant scored 80.8%. H Company used the Apache 2.0 license for the 35B model (CC BY-NC 4.0 for the 27B).
The critical caveat: the held-out AutomationBench score for the 27B dropped to 49.3%. H Company disclosed that 480 of the 600 public AutomationBench tasks were included in its training split, meaning the OSWorld number reflects in-distribution performance. Whether 85.2% holds on genuinely unseen tasks is unanswered. API pricing was listed as $0.40/$3 per MTok for the 27B and $0.30/$2 for the 35B-A3B, per the Neuron digest. Verdict: Test it — build your own held-out task set before committing.
5. Perplexity Decider v1 27B
Perplexity released its first open-source model on October 1 under Apache 2.0. Decider v1 27B is designed for structured decision returns in agent pipelines — routing, classification, and conditional branching. No independent benchmarks were published at launch. Perplexity is positioning this as a lightweight routing layer, not a general-purpose LLM. Verdict: Watch — interesting licensing for EU/enterprise compliance, but no performance data to act on yet.
Cost Per PR Review: The $2/$10 Math
A medium-complexity pull request review — approximately 200,000 input tokens (diff, context files, instructions) and 50,000 output tokens (structured comments) — costs the same across three models this week at $2/$10 pricing.
| Model | Input cost (200K tok) | Output cost (50K tok) | Total per review |
|---|---|---|---|
| Gemini 4 Argon (intro $2/$10) | $0.40 | $0.50 | $0.90 |
| Gemini 4 Argon (standard $4/$20) | $0.80 | $1.00 | $1.80 |
| Claude Sonnet 5.5 ($2/$10) | $0.40 | $0.50 | $0.90 |
| GPT-6.1 Sol ($2/$10) | $0.40 | $0.50 | $0.90 |
| GPT-6 Luna ($0.10/$0.50) | $0.02 | $0.025 | $0.045 |
The math shows the $2/$10 tier at $0.90 per review. At 200 PR reviews per month — a reasonable load for a mid-sized engineering team — that is $180/month per pipeline. GPT-6 Luna at $0.10/$0.50 brings the same task to $9/month, a 20x reduction. BenchLM ranks Luna #37 of 211 models with a composite score of 65/100 — a lower general capability than Sol and the $2/$10 tier. You accept a quality trade-off. The question for each team: is Luna’s capability floor sufficient for your review complexity?
Pricing sources: BenchLM (Luna); shattered.io (Argon); LLM Stats (Sonnet 5.5). For general guidance on building an eval before switching models, see our model evaluation guide.
For Swiss & EU teams
Gemini 4 Argon is available through Google Cloud Vertex AI, which offers EU data-residency regions (europe-west1 through europe-west9). Google’s Data Processing Addendum covers GDPR, and Swiss FADP compliance can be addressed through the same DPA for organisations treating FADP alignment as equivalent to GDPR. Claude Sonnet 5.5 is available through Anthropic’s API; EU-residency options are limited to what AWS Bedrock and GCP Vertex AI offer for Anthropic models in EU regions — teams with strict FADP or EU AI Act data requirements should confirm region availability before migrating. H Company’s Holo4-35B-A3B is Apache 2.0 and can be self-hosted on-prem, which remains the cleanest option for Swiss/EU teams requiring data sovereignty on computer-use tasks.
What to Watch Next Week
Gemini 4 Argon’s standard pricing has no announced start date — Google should clarify this before teams build on introductory rates. Independent evaluations of all three $2/$10 models on the same task sets would resolve the benchmark ambiguity. H Company has not published a timeline for independent AutomationBench testing; if it does not, the OSWorld claim carries a training-contamination asterisk. Perplexity Decider’s first community benchmarks will determine whether it is a useful open-source router or simply a headline release.
For a full history of this week’s model context, see Tuesday’s article on GPT-6 Sol and Luna.
Further Reading
- Gemini 4 Argon: $2/$10 Rates, Benchmark Split — the most detailed independent breakdown of Argon’s benchmark profile, with caveats on vendor-only scores.
- Google’s Gemini 4 Argon Closes the Pricing Triangle — structural analysis of what simultaneous $2/$10 launches mean for the competitive landscape.
- H Company Announces Holo4 Computer-Use Agents — the training-data contamination disclosure that makes Holo4’s OSWorld score hard to read at face value.
Your turn: Are you running Gemini models, Claude, or OpenAI in your coding pipeline this week — and has the $2/$10 convergence actually changed which model you default to? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.
How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

