Cloudflare debuts Clef-omni, supporting audio and video input alongside text and image, makes Clef up to 2x faster, and dramatically cuts Clef-flash pricing (Cloudflare)
Following last week's release of Clef and Clef-flash, Cloudflare's open-weight decision models, we decided to bring forth more gifts.
Cloudflare has unveiled Clef-omni, an enhanced version of its open-weight decision models, capable of accepting audio, video, image, and text inputs. This update makes Clef up to 2 times faster and reduces the price of Clef-flash to be cheaper than Jev. Clef-omni is a result of a model trained over the weekend and launched within days, showcasing Cloudflare's ability to innovate quickly.
By eliminating the need for separate transcription or captioning processes, Clef-omni streamlines multimodal workflows. It achieves this by scoring all modalities and parameter options simultaneously, bypassing the overhead of token generation. The model uses a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation, while discarding non-essential text-to-speech components.
Clef-omni executes prefill passes, scoring all modalities and valid options simultaneously to facilitate fast, schema-constrained scoring. Despite a reduction in the context window from 64k to 24k, Clef-flash remains a cost-effective option for many workflows.
Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.