Cloudflare's Clef-omni: Now with Native Audio & Video
Cloudflare's Clef-omni: One Model to Rule Them All?
Remember when handling video meant chopping it into a million static frames? Cloudflare just said goodbye to that headache. Their new Clef-omni model, now fully open-sourced on Hugging Face, natively understands audio and video—no pre-processing gymnastics required.
From Clunky to Seamless
Previously, Cloudflare's Clef series could only handle text and images. Video? You'd have to split it into frames, transcribe audio separately, then stitch everything back together. Clef-omni throws that pipeline out the window. It directly accepts wav, mp3, mp4, and webm files. Developers can now make a single API call to process text, images, audio, and video together. That's a massive simplification for anyone building multimodal apps.

Built for Speed, Not Shakespeare
Under the hood, Clef-omni runs on the Qwen3-Omni-30B-A3B-Instruct architecture. But don't expect it to write poetry—it's a decision-making specialist. It outputs structured results, not essays. And it's fast: median response time for text is about 130 milliseconds, images around 150 milliseconds, and a 21-second video with sound? Just 1.5 seconds to score it. That's quicker than most people can blink twice.
Price Cuts and Context Tweaks
Cloudflare didn't stop at the model. They also slashed prices on Clef-flash: input costs dropped 58%, from $0.09 to $0.038 per million tokens. The hosted version's context window shrank from 64k to 24k—a trade-off for efficiency. For developers watching their cloud bills, this is welcome news.
What This Means for You
If you're building intelligent workflows or automated decision systems, Clef-omni offers a leaner, cheaper foundation. No more duct-taping together speech-to-text and video splitters. One model, one call, native multimodality. The open-source release means you can tinker, deploy, and scale without vendor lock-in.
So, ready to ditch the pipeline spaghetti? Clef-omni might just be your new favorite tool.
Key Points
- Native audio & video: Clef-omni processes wav, mp3, mp4, and webm directly—no pre-splitting needed.
- Single API call: Handles text, images, audio, and video in one go.
- Blazing fast: ~130ms for text, ~150ms for images, 1.5s for a 21-second video.
- Open-source: Weights available on Hugging Face.
- Price drop: Clef-flash input cost cut by 58% to $0.038 per million tokens.
- Context window: Hosted version reduced from 64k to 24k for efficiency.