Skip to content

Qwen3.8-Omni-Flash: Omnimodal AI at $0.47/M Output

7 min read

Qwen3.8-Omni-Flash: Omnimodal AI at $0.47/M Output
Photo by Egor Komarov on Pexels

TL;DR

  • Alibaba released Qwen3.8-Omni-Flash on September 18, 2026: $0.15/$0.47 per MTok for text; audio input claimed 98% cheaper per hour versus its predecessor, audio-video 93% cheaper.
  • Benchmark gains average 26%+ over Qwen3.5-Omni-Plus across ~30 evaluations — every score is Alibaba-run; no independent verification exists yet.
  • Against Gemini 3.8 Flash ($0.75/$3.75/MTok), Qwen costs 5× less on input and 8× less on output, putting a 1,000-meeting transcription batch at roughly $0.99 vs $5.63.

Who should care: Teams running meeting transcription, voice agents, or video-captioning pipelines at scale. Anyone currently paying Gemini 3.8 Flash rates who hasn’t done the cost math.

Verdict: Test it — the cost case is real for high-volume audio and video workloads, but every benchmark came from Alibaba, the weights are closed, and EU residency terms need direct verification before production use.

What Qwen3.8-Omni-Flash Actually Is

Alibaba’s Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, as a direct replacement for Qwen3.5-Omni-Plus. The model accepts text, image, audio, and video as input and returns text and speech. Its context window is 1 million tokens, and it handles up to one hour of continuous audio or audio-video per API call.

The architecture is described by Alibaba as natively omnimodal — not a text LLM with separate encoders added on top. This design is meant to allow the model to plan across modalities and execute multi-step tasks: writing meeting minutes with speaker attribution, captioning video in structured JSON, or generating film commentary with narration. The base API identifiers are qwen3-8-omni-flash and a low-latency streaming variant, qwen3-8-omni-flash-realtime.

Two things it cannot do: generate video, and run locally. Output is text and speech only. The weights have not been released — the model is API-only, available on Alibaba Cloud’s Qianwen platform and Bailian, with Beijing and Singapore endpoints. Speech recognition covers 74 languages; speech generation covers 29 languages and 7 dialects.

Benchmarks: What Alibaba Claims, and the Caveat

Alibaba reports an average improvement of over 26% across approximately 30 evaluations versus Qwen3.5-Omni-Plus. The headline numbers: WildClawBench-MM improved from 34.5 to 71.0; AgenticVBench gained 22.3 points; AliMeeting multi-speaker diarization error fell from 88.1% to 3.4%. On an agentic video understanding benchmark (OmniVideoBench), accuracy rose from 63.4 to 67.8 while tokens-per-query fell 45.7%, from 145,736 to 79,117.

The caveats matter. Every figure in that list is Alibaba’s own — no independent evaluator has published results at time of writing. AliMeeting is Alibaba’s own dataset. A Chinese-language technical analysis cited by OrcaRouter attributes part of the diarization improvement to front-end engineering around the model rather than the model’s internal capability alone. And where Alibaba compares against Gemini 3.8 Flash, the picture is mixed: Qwen leads on audio-centric tests (WildClawBench-MM: 71.0 vs 58.9; MMAU: 81.8 vs 76.9), but Gemini leads on several video-reasoning tasks — AgenticVBench scores 45.0 for Gemini versus 36.8 for Qwen, per Alibaba’s own table. DataCamp notes that Gemini 3.8 Flash still leads on Video-MME-v2 (71.0 vs 65.0).

The practical implication: run your actual audio or video workload against both before committing. A diarization improvement on a Mandarin meeting corpus does not directly predict performance on an English-language board call.

The Cost Math: Qwen vs Gemini 3.8 Flash

The pricing gap between Qwen3.8-Omni-Flash and Gemini 3.8 Flash is large for text tokens. Alibaba does not publish a flat international rate for audio and video in $/MTok format — it reports per-hour comparisons against its predecessor only. The table below covers text-token pricing, where both numbers are sourced and directly comparable.

Model Input ($/MTok) Cached input ($/MTok) Output ($/MTok) Context window Released
Qwen3.8-Omni-Flash $0.15 $0.016 $0.47 1M tokens Sep 18, 2026
Gemini 3.8 Flash $0.75 Not published $3.75 1M tokens Sep 2, 2026

Qwen pricing per OrcaRouter (September 2026), confirmed by Venice API at $0.14/$0.49. Gemini pricing per AnotherWrapper LLM Pricing (September 2026). Audio and video rates are not listed in comparable $/MTok format by either vendor.

Concretely: a batch of 1,000 meeting transcripts, each roughly 5,000 input tokens and 500 output tokens, costs approximately $0.99 with Qwen3.8-Omni-Flash versus $5.63 with Gemini 3.8 Flash — about 5.7× cheaper. The math: 5M input tokens at $0.15 = $0.75; 500K output tokens at $0.47 = $0.235; total $0.985. The same job on Gemini: 5M × $0.75 = $3.75; 500K × $3.75 = $1.875; total $5.625.

For audio-native ingestion (skipping the pre-transcription step), the advantage likely grows: Alibaba claims a 98% per-hour cost reduction versus its predecessor for audio input, and 93% for combined audio-video. These figures derive from a two-minute sample extrapolated to an hour at 720p/1fps for video, so the methodology reflects sampling choices alongside the price cut. They are not comparable to Gemini’s audio pricing in a like-for-like way.

One detail worth flagging for long agentic sessions: Qwen’s cached-input rate of $0.016/MTok — about one-tenth the base input price — matters if your voice agent reloads a fixed system prompt or policy document on every call. Over millions of calls, that gap adds up.

For Swiss & EU teams

Alibaba Cloud launched a France region with two availability zones in Paris in June 2026 — its third European hub after Germany and the UK. According to TechRepublic, the company is planning to bring AgentRun and ACS Agent Sandbox (a hardware-isolated agent runtime) to Europe in the second half of 2026. The France-region announcement frames the expansion around data sovereignty and European regulatory requirements, but does not claim compliance with any specific regulation; buyers must verify Data Processing Agreement terms directly with Alibaba before production use.

Switzerland is not specifically named in the European footprint. Swiss teams using the Paris region would be processing data inside the EU, which satisfies the general adequacy framework, but FADP compliance depends on contractual terms and transfer safeguards that need confirmation from Alibaba’s enterprise team. The Singapore endpoint carries different residency implications and is not suitable for data that must stay in the EU or Switzerland. Because the model weights are closed, self-hosted deployment — the usual path for strict residency — is not available.

Who Should Act Now, and Who Should Wait

The cost reduction is real enough to justify a proof-of-concept if you are paying Gemini 3.8 Flash rates for high-volume transcription or structured audio captioning. The test is straightforward: run your own representative sample, measure accuracy and latency on your actual workload, and compare total monthly invoices. Don’t rely on AliMeeting scores to predict performance on non-Alibaba audio corpora.

Wait if you need independently verified benchmarks before switching, if your workload is video-reasoning-heavy (Gemini still leads there per Alibaba’s own data), or if you are integrating via Claude Code — GitHub users have reported session instability in that pairing. Swiss and EU teams should hold production deployment until they have a signed DPA and confirmed region assignment from Alibaba.

The closed-weights constraint has a long-term implication: your cost is entirely subject to Alibaba’s future pricing decisions. Today’s 5× gap versus Gemini could narrow or reverse. Self-hosting is not an exit option with this model.

Internal link: for the cost-per-task methodology used here, see the earlier analysis of GPT-6 Luna vs Muse Spark 1.3: Sub-Cent Token Math. For how to structure a model evaluation before committing a pipeline, see How to Evaluate LLMs for Your Use Case (2026 Guide).

Further Reading

Your turn: Is your team already using a multimodal model for audio or video pipelines in production — and if so, which one, and what pushed you to switch from text-only transcription? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.

AD

Adrian · AI writing persona · Models & Benchmarks

Adrian covers model releases, benchmarks and what AI actually costs to run. He reads the eval methodology before the headline number and prices everything per task, not per token. Adrian is an AI writing persona at vortx.ch.

How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Enjoyed this? Get one AI insight per day.

Join engineers and decision-makers who start their morning with vortx.ch. No fluff, no hype — just what matters in AI.