Cloudflare's Clef-omni: Open-Source Multimodal Model Processes Audio and Video Natively
Cloudflare Unveils Clef-omni: A Multimodal Leap for Decision Models
Cloudflare has officially released Clef-omni, a new open-weight decision model that natively supports text, images, and—for the first time—audio and full video inputs. The model's weights are now fully open-sourced on Hugging Face, marking a significant upgrade from the existing Clef series.
From Fragmented Pipelines to Unified Processing
Previously, handling video with Clef meant splitting it into static frames along a timeline. That's clunky. Clef-omni changes the game by directly supporting mainstream formats like wav, mp3, mp4, and webm. Developers no longer need to build separate speech transcription and audio-video splitting pipelines. Instead, a single API call can process text, images, audio, and video together. It's a streamlined approach that could save countless hours of integration work.

Built for Speed and Structured Decisions
Under the hood, Clef-omni is based on the Qwen3-Omni-30B-A3B-Instruct architecture, retaining its core understanding capabilities. But this isn't a model designed to write essays—it's a specialized tool for structured decision-making tasks. Official benchmarks show a median response time of about 130 milliseconds for pure text requests and around 150 milliseconds for image requests. Processing a 21-second video with sound? Just 1.5 seconds to complete scoring. That's fast enough for real-time applications.
Price Cuts and Context Adjustments
Alongside the new model, Cloudflare has adjusted pricing for other product lines. The input price for Clef-flash dropped by roughly 58%, from $0.09 per million input tokens to $0.038. However, the context window for the hosted version was reduced from 64k to 24k. For developers, this means lower costs but a tighter context—a trade-off worth noting.
What This Means for Developers
With Clef-omni's open-source release and native multimodal capabilities, Cloudflare is offering a more efficient and cost-effective foundation for building complex intelligent workflows and automated decision systems. Whether you're working on video analysis, audio transcription, or cross-modal tasks, this model could simplify your stack. The open weights also invite community experimentation and fine-tuning.
Key Points
- Native multimodal support: Clef-omni handles text, images, audio, and video directly, eliminating separate pipelines.
- Fast performance: ~130ms for text, ~150ms for images, and ~1.5s for a 21-second video.
- Open-source weights: Available on Hugging Face for developers to use and modify.
- Price reduction: Clef-flash input price cut by 58%, though context window shrinks to 24k.
- Built on Qwen3-Omni: Retains core understanding while focusing on structured decision tasks.