TL;DR
- Claude Haiku 5.5 cuts input price 90% to $0.10/MTok (≤100K context) — same rate as GPT-6 Luna — while more than doubling every agentic benchmark score versus Haiku 4.5 (Anthropic, vendor-run evals, Oct 7).
- Reflection AI announces Beam, a 501B open-weight MoE with 1M-token context; weights are still in final testing and not yet publicly available.
- Qwen3.8-Omni-Flash ($0.15/$0.47 per MTok) is the week’s cheapest option that handles text, image, audio and video in a single API call.
Who should care: Engineering teams running Claude Haiku 4.5 in CI or agentic pipelines; developers building voice or video agents on a tight token budget; open-source shops watching the Western open-weight space.
Verdict: Use it — migrate from Haiku 4.5 to Haiku 5.5 this week. The price cut alone justifies the switch; the benchmark jump makes it urgent.
The Week’s Defining Move: Claude Haiku 5.5
Anthropic released Claude Haiku 5.5 on October 7, cutting the standard input price from $1.00 to $0.10 per million tokens for prompts up to 100,000 tokens — a 90% reduction on rate. Longer prompts cost $0.50/MTok input and $2.50/MTok output, still far below the Haiku 4.5 baseline of $1.00 and $5.00. Anthropic estimates actual workload costs fall roughly 75% on average, because its updated tokenizer counts approximately 30% more tokens for the same text: the headline 90% rate reduction is real, but effective savings per prompt sit closer to 75–87% depending on content type.
The benchmark jump is just as striking. On OSWorld 2.1 (offline subset), Haiku 5.5 scores 72.4% against Haiku 4.5’s 15.7% — a 4.6× improvement (Anthropic, vendor-run). Terminal-Bench 4.0 at maximum effort moves from 0.0% to 39.2%. FrontierCode 1.1 reaches 46.4%, ahead of GPT-6 Luna’s 42.4% on the same evaluation. The agentic scores are more striking: GDPval-AA v2.1 climbs from 735 to 1,620 (versus 1,437 for GPT-6 Luna), and AA-Briefcase v1.1 from 614 to 1,578. All figures are Anthropic vendor-run; treat absolute values with standard caution, but the direction is consistent across every reported metric.
Haiku 5.5 adds beta computer-use support — browser and desktop — through the Python and TypeScript SDKs. Anthropic positions it as a subagent model, not a replacement for demanding single-step coding tasks. That framing matches the pricing: at roughly $0.011 per 100K-input round trip, agentic loops that were prohibitively expensive just became routine. Customer-reported figures from Anthropic’s launch post: Box reports an 11-point task-completion improvement and roughly half the latency; HubSpot reports 92.8% average across three CRM evaluation runs. These are company-supplied comparisons, not controlled studies.
One caveat: Haiku 5.5 ships with tighter cybersecurity restrictions than its predecessor. Standard safeguards block penetration-testing use cases. Broader cyber or biology access requires Anthropic’s verification programmes.
This Week, Ranked
| Rank | Model | Vendor | What changed | Key number (source) | Verdict | Best for |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 5.5 | Anthropic | 90% price cut to $0.10/MTok input; agentic benchmarks more than double | OSWorld 72.4% vs 15.7% for Haiku 4.5 (Anthropic, vendor-run, Oct 7) | Use it | Agentic subagents, CI coding loops, cost-sensitive pipelines |
| 2 | Reflection Beam | Reflection AI | First open-weight MoE from a Western AI startup; 501B total / 23B active per inference | SWE-bench Verified 80.9% (Reflection, vendor-run, Oct 5 — not independently verified) | Watch | Open-source teams; revisit when weights land in October |
| 3 | Qwen3.8-Omni-Flash | Alibaba | Omnimodal (text/image/audio/video), 1M-token context, $0.47/MTok output | GPQA Diamond 91.0%, LiveCodeBench v6 92.6% (Alibaba, vendor-run, Sep 18) | Test it | Voice and video agents at low per-token cost |
| 4 | Mistral Large 4 | Mistral AI | 1T open-weight multimodal in preview; final weights due end of October | In preview; no public independent scores yet | Watch | EU teams evaluating open-weight frontier; see full analysis |
Cost Math: One Code Review, Three Models
A typical CI code-review pass on a 200-file PR might consume 100,000 input tokens and produce 2,000 output tokens. Here is what each model charges for that task at published API rates (Anthropic pricing via VentureBeat, Oct 8; Qwen pricing via OpenRouter listing):
| Model | Input (100K tokens) | Output (2K tokens) | Per review | 10,000 reviews/month |
|---|---|---|---|---|
| Claude Haiku 5.5 | $0.010 | $0.001 | $0.011 | $110 |
| Qwen3.8-Omni-Flash | $0.015 | <$0.001 | $0.016 | $160 |
| Claude Haiku 4.5 | $0.100 | $0.010 | $0.110 | $1,100 |
Math: Haiku 5.5 at (100,000 ÷ 1,000,000) × $0.10 + (2,000 ÷ 1,000,000) × $0.50 = $0.010 + $0.001 = $0.011. Qwen3.8-Omni-Flash at (100,000 ÷ 1,000,000) × $0.15 + (2,000 ÷ 1,000,000) × $0.47 = $0.015 + $0.00094 ≈ $0.016. The tokenizer change means Haiku 5.5 counts slightly more tokens for the same source text than Haiku 4.5; the table uses round figures consistent with Anthropic’s ∼75% average saving estimate.
Qwen’s input rate ($0.15/MTok) is 50% higher than Haiku 5.5’s, so the gap widens on input-heavy tasks. Its output rate ($0.47/MTok) is nearly identical to Haiku 5.5’s ($0.50/MTok), so the two models converge on generation-heavy jobs. The key differentiator is modality: Qwen handles audio and video in the same call; Haiku 5.5 does not. For text-only or image pipelines, Haiku 5.5 wins on price. GPT-6 Luna matches Haiku 5.5’s short-context rates exactly — for a full cost comparison across the cheap tier, see our GPT-6 Luna cost breakdown.
Model by Model
1. Claude Haiku 5.5 — Use it
The two-tier pricing is worth understanding: prompts up to 100,000 tokens hit the $0.10/$0.50 rate; anything longer jumps to $0.50/$2.50. Most subagent calls stay well under 100K, so the majority of agentic workloads land on the cheaper tier. Cache reads cost $0.01/MTok (short-context), which makes prompt-caching patterns even more attractive than before. The beta computer-use support is not production-ready by Anthropic’s own framing, but it is available to test. For teams running Haiku 4.5 today: there is no benchmark regression, the price fell 90% on rate, and the model ships now. Upgrade.
2. Reflection AI Beam — Watch
Announced October 5, Beam is a text-only sparse MoE: 501B total parameters, 23B active per inference, trained on 23.8 trillion tokens. Reflection followed pretraining with a four-week reinforcement-learning run on 10,500 NVIDIA GB300 GPUs, generating over 100 million rollouts. Stated benchmarks: 80.9% on SWE-bench Verified and 77.2% on SWE-bench Pro v2-Hard (Reflection, vendor-run, no independent verification as of publication). Context window is 1 million tokens. Reflection says Beam uses 3–4× less inference compute than rival Western open-weight models, and matches GLM-5.2 on selected tests while falling short of Kimi K3 on raw capability. The intended licence is Apache 2.0; weights, a technical report and a model card are due later in October. Early access is open via signup at Reflection’s site. Nothing to benchmark in production yet.
3. Qwen3.8-Omni-Flash — Test it
Released September 18 and newly worth evaluating for teams that haven’t looked yet: a single API call accepts text, image, audio and video, served through an OpenAI-compatible endpoint. Alibaba reports GPQA Diamond 91.0%, LiveCodeBench v6 92.6%, and OSWorld 87.1% (all vendor-run, no error bars disclosed). BenchLM ranks it #32 of 125 on instruction following — its strongest tracked category — with coding at 54.1 and knowledge at 54.0 (unranked composites). Context window is 1M tokens; weights are not published. A separate realtime WebSocket variant handles sub-second voice; the benchmarks above apply only to the standard model. If your pipeline currently calls separate APIs for voice transcription, image analysis and text generation, a single-model consolidation test is worth two hours of engineering time.
4. Mistral Large 4 — Watch
We published a full analysis of Mistral Large 4 on Tuesday. Short version: 1T-parameter multimodal model in preview, targeting coding and cyberdefence, final weights due end of October, no public independent scores yet. The open-weight angle and Mistral’s French origin remain the EU-relevant points; neither matters until weights are available for self-hosting evaluation.
For Swiss & EU teams
Reflection Beam’s stated Apache 2.0 licence is the headline for data-residency-conscious teams: if the weights land as promised, self-hosting in EU or Swiss infrastructure is straightforward with no API calls leaving the region. Qwen3.8-Omni-Flash runs through Alibaba Cloud Model Studio; teams with strict FADP or EU GDPR requirements should review Alibaba’s data processing terms before routing production traffic there. On pricing: Haiku 5.5 now matches GPT-6 Luna’s short-context rates, removing the cost argument against Anthropic for EU teams already covered by Anthropic’s EU data processing agreement.
What’s Next
The open-weight queue is filling fast: Reflection Beam weights and Mistral Large 4’s final release are both expected before November. Independent benchmark results on either will reset this comparison table. For now, the week’s verdict is clear — there is no good reason to keep Haiku 4.5 running in any pipeline. October will tell us whether a Western open-weight model can match Kimi K3 on raw coding ability, and whether that matters more than the $0 API bill.
Further Reading
- VentureBeat: Anthropic launches Claude Haiku 5.5 — full pricing breakdown, benchmark table and customer-reported figures
- Let’s Data Science: Reflection AI Beam — architecture details and Reflection’s stated benchmark claims with sourcing notes
- BenchLM: Qwen3.8-Omni-Flash — independent benchmark tracking; composite ranking not yet assigned
Your turn: If you switched from Claude Haiku 4.5 to Haiku 5.5 this week — or decided to hold off — what was the deciding factor? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.
How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

