Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026: a native omni-modal model that reads text, images, audio and video in a single pass and answers in text. The number that stands out is the context window, 1 million tokens, split into 991K of input, 131K of output and 262K reserved for reasoning. Qwen also claims the model now beats Google's Gemini 3.8 Flash on audio tasks and gets close to it on audio-visual work, while charging a fraction of the price. Those are vendor numbers, not independent ones.

  • Qwen3.8-Omni-Flash processes text, images, audio and video together and outputs text only, with a 1 million token context window covering 991K input, 131K output and 262K reasoning tokens.
  • Qwen says audio performance now beats Gemini 3.8 Flash and audio-visual performance comes close to it, though these are the company's own benchmarks and have not been independently verified.
  • QwenCloud prices the model at $0.15 per 1 million input tokens and $0.47 per 1 million output tokens, and Qwen says audio input costs more than 98 percent less per hour than its predecessor, Qwen3.5-Omni-Plus.
  • The model ships API-only with closed weights, a reversal from its own base model, Qwen3.8-Flash-Next, which had its weights released openly in August 2026.
Qwen3.8-Omni-Flash omni pipeline Text, image, audio and video inputs pass through an agentic perception coarse-to-fine stage into a 1 million token context window, then out as text or through Qwen-MM-Plugins. AudioVideoImageText Agentic perception: coarse to fine 1M-token context window 991K in · 131K out · 262K reasoning Text outputQwen-MM-Plugins genztech.blog
Fig 1 Text, image, audio and video go in together. An agentic perception stage scans video coarse to fine before committing tokens to detail, then everything lands in a 1 million token context window and comes out as text, either directly or through Qwen-MM-Plugins.

What exactly did Alibaba ship?

Qwen3.8-Omni-Flash is available now on QwenCloud, Alibaba Cloud Model Studio and Qwen Studio, across six regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia. It speaks both DashScope and OpenAI-compatible API protocols, and supports function calling, web search, structured outputs, context caching and batch calls. Thinking is on by default at a "xhigh" reasoning effort setting, which is where the 262K reasoning-token allowance comes from.

RelatedQwen3.8-27B ships open weights, scoreboard attached

Alongside the model, Qwen released Qwen-MM-Plugins under an Apache-2.0 license on GitHub, with three launch capabilities: omni-memory, omni-video2note and omni-chatcut. The repo lists integrations for Claude Code, CodeBuddy, Codex, Qoder, OpenClaw, Qwen Code and Gemini CLI, though most coding harnesses cannot feed audio to a model natively, so audio is routed through the API instead. Qwen also mentions a "Qwen-Live Harness" for long-running and real-time workflows, though the details published so far are thin.

How much better are the benchmarks, and why does agentic perception matter?

Qwen says the average score across its evaluation suite improved by more than 25 percent compared with Qwen3.5-Omni-Plus. TechNode cites "more than 26% across 30 evaluations"; MarkTechPost's tally lists 29, a small discrepancy likely from different cuts of the same result set. The individual deltas Qwen published are steep: WildClawBench-MM up 36.5 points, AgenticVBench up 22.3 points, UniClawBench scoring 69.6, LongAudioSpan up 8.3 points, and OmniVideoBench up 9.6 points, rising from 63.4 to 67.8 once agentic perception is switched on. Average agent performance across WildClawBench-MM and UniClawBench rose 19.5 points.

Agentic perception is a coarse-to-fine approach to video: instead of processing every frame at full resolution, the model scans coarsely first and spends its attention on the parts that matter. Qwen says this cut token usage on OmniVideoBench by about 45.7 percent, from 145,736 tokens down to 79,117, while the benchmark score went up. That is the more interesting claim than the raw score gains: a model that gets better and cheaper on the same task, if the number holds up outside Qwen's own testing.

How much cheaper is it really?

Pricing is where Qwen makes its most aggressive claim. On QwenCloud, Qwen3.8-Omni-Flash costs $0.15 per 1 million input tokens, $0.47 per 1 million output tokens, and $0.016 per 1 million cached input tokens. TechNode separately reports Alibaba's API input pricing as "as low as RMB 0.8 per million tokens". But per-token pricing understates what matters most for audio and video: cost per hour, not per token. There Qwen's numbers are larger: audio input costs more than 98 percent less per hour than Qwen3.5-Omni-Plus, audio-visual input more than 93 percent less, and video roughly 89 percent less.

Qwen3.8-Omni-Flash cost cuts versus Qwen3.5-Omni-Plus Qwen says per-hour input costs fall about 98 percent for audio, 93 percent for audio-visual input, and about 89 percent for video, compared with Qwen3.5-Omni-Plus. 100% = Qwen3.5-Omni-Plus cost -98%-93%~-89% Audio inputAudio-visual inputVideo input Cost per hour, Qwen3.8-Omni-Flash vs Qwen3.5-Omni-Plus genztech.blog
Fig 2 · benchmark Qwen's own figures put audio input costs more than 98 percent lower per hour than Qwen3.5-Omni-Plus, audio-visual input over 93 percent lower, and video roughly 89 percent lower. These are vendor-reported, not independently audited.

Input limits are generous: up to 2 hours or 2GB of video via URL, up to 3 hours of audio across 113 languages and dialects, video sampling up to 15 frames per second, and spatial audio in two-channel stereo or four-channel FOA.

RelatedDeepSeek V4 Pro hits GA, but Flash still scores higher

Qwen3.8-Omni-FlashGemini 3.8 FlashQwen3.5-Omni-Plus
Context window1M tokens (991K in, 131K out, 262K reasoning)not publishednot published
Input modalitiesText, image, audio, videonot publishednot published
Output modalityText onlynot publishednot published
WeightsClosed, API onlynot publishednot published
Input price (QwenCloud)$0.15 per 1M tokensnot publishednot published
Output price (QwenCloud)$0.47 per 1M tokensnot publishednot published
Audio cost per hour vs predecessorOver 98% lessnot publishedBaseline
Vendor performance claimAudio above Gemini 3.8 Flash; audio-visual close to itNamed target; no self-published claim availablePredecessor; average score about 25%+ lower per Qwen's tally

What it means for the market

The company with the most exposure here is Alibaba itself, listed as NYSE: BABA and HKEX: 9988, since Qwen is the model line it is using to compete directly with Google. Google is the company being named as the target: Qwen isn't just claiming to be competitive with Gemini 3.8 Flash, it is claiming to beat it on audio and undercut it heavily on price. If that claim survives independent testing, it pressures Google on a segment where margin is set by the cheapest credible option: high-volume multimodal inference for meetings, calls and video review. The signal for investors is that the multimodal API market is now being fought on price per hour, not price per token, and Alibaba is pushing that fight furthest right now. None of this is investment advice, and Qwen's numbers have not been checked outside Qwen yet.

Our take

Take the benchmark numbers with the standard discount. Vendors report the results that make their model look best, and genztech's own AI coding leaderboard has tracked this gap before: self-reported scores routinely run ahead of independent testing. The audio-per-hour pricing claim is the part worth watching closest, because meeting transcription and call analysis workloads live or die on cost per hour, not cost per token, and a 98 percent cut, if it holds at scale, is not marginal.

The bigger shift is that Qwen closed the weights on this one. Qwen3.8-Flash-Next, the base model underneath Qwen3.8-Omni-Flash, had its weights released openly in August 2026. The omni model built on top of it did not get the same treatment. Alibaba has built goodwill by giving away base models; keeping the omni layer API-only is worth watching as a pattern, not a one-off.

What to watch
  • Independent benchmarks. No outside lab has reproduced the WildClawBench, AgenticVBench or UniClawBench numbers yet.
  • Whether the weights stay closed. Qwen3.8-Flash-Next was open; if the omni line stays API-only, that is a deliberate shift.
  • Coding-harness adoption. Qwen-MM-Plugins lists Claude Code, Codex and Gemini CLI, but audio still routes through the API, not the harness.
  • Whether per-hour pricing holds at scale. Launch pricing and real production cost are not always the same number.
Primary sources

Original analysis by GenZTech. Primary source: Qwen blog.