Alibaba Qwen3.8 adds 1M-context Omni analysis and 60-language live translation

18-09-2026

Alibaba Cloud has added two Qwen3.8 multimodal endpoints. Omni-Flash combines a 1M context window with audio, video, reasoning and tools, while LiveTranslate targets real-time multilingual translation with speech output.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

alibaba qwen3.8 adds 1m context omni analysis and 60 language live translation
Sign up for our Newsletter

Publication date: 18 September 2026

Alibaba Cloud has added two Qwen3.8 multimodal endpoints to Model Studio:

  • Qwen3.8-Omni-Flash for long-context audio and video analysis,
  • Qwen3.8-LiveTranslate-Flash-Realtime for live multilingual translation.

The endpoints appeared in the public catalog on 17 September 2026, with the Omni documentation updated again on 18 September. The practical change is workload separation. Omni-Flash handles text, images, audio and video with 1M context and agent tools. LiveTranslate is a narrower WebSocket model that translates live audio and visual context into text and speech.

Qwen3.8-Omni-Flash turns multimodal files into agent context

Qwen3.8-Omni-Flash accepts text, images, audio and video, but returns text only. It supports custom function calling, built-in web search, automatic context caching and adjustable reasoning effort. That makes it an analysis model for multimodal agents rather than a speech-to-speech assistant.

Alibaba lists a 1M-token context window, with maximum input of 991,808 tokens without reasoning and 983,616 with reasoning. Maximum output is 131,072 tokens. Audio input covers 113 languages and dialects, while multichannel audio can preserve spatial information.

The biggest jump is scale and agent functionality

SpecificationQwen3.8-Omni-FlashQwen3.5-Omni-Flash
Context1M tokens131,072 tokens
OutputTextText and optional audio
Function callingSupportedNot supported in non-realtime API
ReasoningAdjustableNo comparable control documented

Qwen3.8-Omni-Flash is better suited to long recordings, large video collections and tool-using analysis. Teams needing generated speech should keep a realtime or Qwen3.5 Omni endpoint in the stack.

Alibaba lists up to 30,000 requests per minute for Qwen3.8-Omni-Flash in several regions, with token limits varying by region. Availability includes Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia.

The official pricing documentation was checked, but its searchable table does not yet expose a stable Qwen3.8-Omni-Flash row across all regions. Teams should confirm the console price for their deployment region before forecasting production cost.

LiveTranslate keeps 60 input languages and adds a new Qwen3.8 endpoint

Qwen3.8-LiveTranslate-Flash-Realtime uses WebSocket and accepts audio plus images. It understands 60 languages and can produce translated speech in 29 languages. Context is 53,248 tokens, with 49,152 maximum input and 4,096 maximum output.

Qwen3.5-LiveTranslate-Flash-Realtime also supports 60 input languages, so this is not a language-count upgrade. The material change is a new Qwen3.8 production endpoint with updated API events and image context while preserving broad language coverage.

In Singapore, Alibaba lists audio input at $7.50 per million tokens, image input at $0.55, text output at $20 and audio output at $30. The published limit is 10 requests and 100,000 tokens per minute. High-volume voice operations therefore need quota planning.

What media and localization teams should test first

Media-analysis teams should give Omni-Flash one long recording or video corpus that currently requires chunking. Measure retrieval accuracy across the full context, tool-call correctness, latency and actual token cost. A 1M window matters only if the model reliably finds the relevant moment.

Localization teams should test LiveTranslate on noisy and overlapping speech in their highest-value language pairs. Measure translation accuracy, timing and speech latency separately. Alibaba has not published mature independent evaluations for either endpoint, so buyers need their own production evidence.

What this means

Teams building multimodal research, media intelligence or long-form video agents should test Qwen3.8-Omni-Flash now. Start with a large audio or video job that currently needs manual chunking or a separate transcription stage.

Live translation teams should test Qwen3.8-LiveTranslate-Flash-Realtime when broad language input and spoken output matter. Neither release justifies an automatic switch: Omni pricing needs region-level confirmation, and both models need independent quality evidence.

For more information, visit the official announcement of Qwen3.8-Omni-Flash on the Alibaba Cloud website.

Add DataNorth AI to your Google favorites