Cloudflare's Clef-omni: Now Your AI Can Watch and Listen
Cloudflare's Clef-omni: One Model to Hear, See, and Decide
Cloudflare has unveiled Clef-omni, an open-weight decision model that now supports audio and full video inputs. Announced on October 9, the model lets developers process multimodal content with a single API call—no more juggling separate transcription and segmentation tools.

From Text and Images to Sound and Motion
Previously, the Clef family handled text, images, and static frames pulled at intervals. Clef-omni changes that: it natively ingests wav, mp3, mp4, and webm files. That means you can throw a podcast episode or a video clip at it and get structured decisions back, all without pre-processing headaches.
Under the hood, Clef-omni is built on Qwen3-Omni-30B-A3B-Instruct and keeps its core understanding chops. But don't expect it to write essays—it's designed for decision-making tasks, not text generation. Speed? Median response for pure text requests is about 130 milliseconds; images clock in around 150 milliseconds. A 21-second video with sound takes roughly 1.5 seconds to evaluate. That's quick enough for near-real-time applications.
Cheaper, Leaner, and Ready for Developers
Alongside the new model, Cloudflare slashed the price of Clef-flash. The cost per million input tokens dropped from $0.09 to $0.038—a 58% cut. The hosted version's context window shrank from 64k to 24k, a trade-off that aligns with typical decision tasks. Overall, the tiered pricing now caters to different budgets and performance needs.
So what's the catch? For one, Clef-omni isn't a general-purpose chatbot. It's a specialized tool for structured outputs, which means it won't replace your favorite LLM for creative writing. But if you're building a system that needs to react to audio or video on the fly—think content moderation, live captioning, or interactive agents—this could be a game-changer.
Key Points
- Multimodal support: Clef-omni natively handles audio (wav, mp3) and video (mp4, webm) without extra preprocessing.
- Built for decisions: Based on Qwen3-Omni-30B-A3B-Instruct, it focuses on structured decision tasks, not text generation.
- Fast performance: ~130 ms for text, ~150 ms for images, and ~1.5 seconds for a 21-second video with audio.
- Price drop: Clef-flash now costs $0.038 per million input tokens, down 58% from $0.09.
- Context window: Hosted version reduced from 64k to 24k tokens, balancing cost and capability.
For developers tired of stitching together multiple models, Clef-omni offers a streamlined path. It's not perfect for every use case, but it's a solid step toward unified multimodal processing. Keep an eye on how the community puts it to work.