Urgent.News

What's breaking now, across thousands of outlets.

AI

Qwen3.8-Omni-Flash Launches With 1M Context and Advanced Audio-Visual Agents

Qwen has launched Qwen3.8-Omni-Flash, its next-generation native omnimodal model designed to handle text, images, audio, and video while also planning … Read More The post Qwen3.8-Omni-Flash Launches With 1M Context and Advanced Audio-Visual Agents appeared first on ProPakistani .

Qwen3.8-Omni-Flash Launches With 1M Context and Advanced Audio-Visual Agents

Qwen has unveiled Qwen3.8-Omni-Flash, its latest native omnimodal model capable of processing text, images, audio, and video. This advanced model features a 1-million-token context window and can plan tasks, utilize tools, and execute multi-step operations. Designed to enhance workflows such as video editing, music video creation, film commentary, audio-visual summarization, and real-time conversations, Qwen3.8-Omni-Flash also boasts significant improvements over its predecessor.

With an average score surpassing Qwen3.5-Omni-Plus by more than 25% across 29 evaluations, Qwen claims a 98% reduction in audio input costs and a 93% decrease in audio-visual input costs under its pricing structure. The model's audio-visual performance closely mirrors Gemini 3.8 Flash while outperforming it in overall audio capabilities, according to Qwen's own testing.

Notably, Qwen3.8-Omni-Flash excels at analyzing specific elements within recordings, such as characters, camera shots, lighting, and sound, based on user requests. This capability, combined with its ability to process up to one hour of audio-visual meeting input, enables the model to identify speakers, transcribe discussions, generate minutes, extract action items, and analyze project risks.

When integrated with external tools, the model can perform tasks like sending emails, organizing tasks, or initiating coding based on meeting requirements.

Additionally, Qwen3.8-Omni-Flash-Realtime allows for real-time audio and video processing, enabling the model to determine the origin of sounds in a given environment and assist with spatial navigation. Speech recognition is supported in 74 languages, including Urdu and Punjabi, while speech generation covers 29 languages. Furthermore, Qwen has expanded Qwen-MM-Plugins and made Qwen-Live Harness available as open-source resources for multimodal agent workflows, long-term memory, task delegation, and real-time interaction.

Written by urgent.news from ProPakistani's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at propakistani.pk →

More in AI

Build and Test an AI Agent Skill with SKILL.md and Python

AI agents are good at interpreting intent, but they are not a reliable place to hide every rule in a workflow. If an agent must count characters, parse a file, or refuse to overwrite an existing…

  • Create skill structure with SKILL.md and scripts directory
  • Implement deterministic script to validate Conventional Commit messages
  • Test script thoroughly covering various message scenarios

The More Powerful the AI, the More the Architecture Matters

The boundaries I designed, the gaps I haven't solved, and why the difference matters Part 14 findings of an experiment: building an LLM-powered support agent with deterministic boundaries.

  • AI architecture crucial for effectiveness, as shown in experiment
  • Deterministic boundaries, allowlist, and gate prevent unexpected dependencies
  • Gate ensures low-risk tasks handled by AI, high-risk tasks require human oversight

More from Saturday 19 September →