Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Qwen3.8-Omni-Flash

Budget

by Alibaba

Qwen3.8-Omni-Flash is Alibaba's first omni-modal model built around agentic work, released September 18, 2026 as qwen3.8-omni-flash. It takes text, image, audio, and video on one endpoint and is built natively on the Qwen3.8-Flash foundation rather than bolting perception encoders onto a text model. QwenCloud lists it at $0.15 per million input tokens and $0.47 output, with implicit cache reads at $0.016, a 1 million token context window and up to 131K output tokens. The headline is not the per-token rate, which matches Qwen3.8-Flash exactly, but what that rate does to media: Alibaba says an hour of audio input costs 98 percent less than on Qwen3.5-Omni-Plus and an hour of combined audio and video costs more than 93 percent less, which is the difference between transcribing a meeting archive as an experiment and doing it as a standing job. It handles up to an hour of continuous audio or audio-video per call, recognizes speech in 74 languages, and Alibaba reports an average gain of more than 25 percent across 29 evaluations against Qwen3.5-Omni-Plus, all vendor-reported. One limitation deserves to be read before you design around it: despite the Omni name and a predecessor that generated speech, this model outputs text only, so anything voice-facing still needs a separate text-to-speech stage. Weights have not been published. Available through Qwen Chat, QwenCloud, and the Model Studio API on an OpenAI-compatible endpoint.

Input Price

$0.15

per 1M tokens

Output Price

$0.47

per 1M tokens

Context Window

1M

tokens

Released

2026-09

API access

Capabilities

textvisionaudiovideotool-usecodereasoning

Key Strengths

  • ✓Audio input roughly 98 percent cheaper per hour than Qwen3.5-Omni-Plus, vendor-reported
  • ✓Text, image, audio, and video input on a single endpoint
  • ✓$0.15/$0.47 per 1M with implicit cache reads at $0.016
  • ✓1M token context window with 131K max output
  • ✓Up to an hour of continuous audio or audio-video per call
  • ✓Speech recognition across 74 languages

Best For

  • ▸Meeting, call, and long-video summarization at volume
  • ▸Audio-video agents that also need tool calling
  • ▸Multimodal research over recorded material
  • ▸Pipelines that were priced out by per-hour audio rates

Pricing Details

Input tokens

$0.15

per 1M tokens

Output tokens

$0.47

per 1M tokens

Estimated cost per 1K requests

$0.38

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide