Skip to main content

Cloudflare's Clef-omni Now Hears and Sees: New Open Model Tackles Audio and Video

Cloudflare's Clef-omni: A Decision Model That Actually Listens and Watches

On October 9, Cloudflare quietly rolled out something that could make developers' lives a whole lot easier: Clef-omni, an open-weight decision model that now supports audio and video inputs. No more juggling separate APIs for speech-to-text and video frame extraction—this one handles it all in a single call.

From Text and Images to Full Multimodal

The original Clef model was already handy for text, images, and static frames pulled at intervals. But Clef-omni takes it further: it natively supports wav, mp3, mp4, and webm formats. That means you can feed it a podcast, a video clip, or a mix of media without the usual preprocessing headaches. Developers used to spend hours stitching together transcription and segmentation tools; now it's one model, one request.

Image

Under the Hood: Built on Qwen3-Omni

Clef-omni isn't starting from scratch. It's built on Qwen3-Omni-30B-A3B-Instruct, which gives it a solid foundation for understanding multimodal content. But here's the twist: it's not designed to write essays or chat. Instead, it focuses on structured decision-making tasks—think classification, routing, or any scenario where you need a quick, reliable judgment.

Speed? The median response time for pure text requests is around 130 milliseconds, and for images, about 150 milliseconds. For a 21-second video with sound, it takes roughly 1.5 seconds to evaluate. That's fast enough for real-time-ish applications, though you wouldn't want to throw a feature-length film at it.

Price Cut for Clef-flash

Alongside the new model, Cloudflare trimmed the price of Clef-flash. The cost per million input tokens dropped from $0.09 to $0.038—a 58% reduction. The hosted version's context window shrank from 64k to 24k, which might matter for some use cases, but for many decision tasks, it's a fair trade-off. Overall, the tiered pricing now covers a broader range of budgets and performance needs.

What This Means for Developers

If you're building apps that need to process audio or video without sending everything to a heavy LLM, Clef-omni could be a sweet spot. It's open-weight, so you can self-host if you want. And with the price drop on Clef-flash, experimenting just got cheaper. The catch? It's not a general-purpose model—don't expect it to write your blog posts. But for decision-making, it's a specialized tool that might save you a lot of plumbing.

Key Points

  • Clef-omni supports audio and video natively, eliminating separate transcription and segmentation steps.
  • Built on Qwen3-Omni-30B-A3B-Instruct, it targets structured decision tasks, not text generation.
  • Response times: ~130ms for text, ~150ms for images, ~1.5s for a 21-second video.
  • Clef-flash price cut by 58% to $0.038 per million input tokens; context window reduced to 24k.
  • Open-weight model allows self-hosting and customization.