Cloudflare has released Clef-omni, a decision model that can assess text, images, audio and video in a single API call. The company also says it has lowered the price of Clef-flash and made Clef faster.
Clef-omni accepts WAV or MP3 audio and MP4 or WebM video alongside text and images. Developers can give it a set of possible answers to score, such as whether a machine sounds normal or whether its fan is running. Cloudflare says this lets one model assess the inputs together without first transcribing audio or separating video into other inputs.
Clef-omni scores inputs together
Cloudflare built Clef-omni on Qwen3-Omni-30B-A3B-Instruct. According to the company, it uses that model’s ability to process several input types but omits its text-to-speech output components. Clef-omni scores defined options instead of generating a written response.
Cloudflare reports median response times of about 130 milliseconds for text-only decisions and 150 milliseconds for image inputs. It says audio clips take a few hundred milliseconds, while a 21-second video with sound takes about 1.5 seconds. Those are company-reported figures.

Cloudflare cuts Clef-flash pricing
Cloudflare says Clef-flash now costs less than Jev, a decision model from TypeSafe. The company directs developers to its documentation for current rates; it does not give the new dollar price in the announcement. Cloudflare also says Clef is faster, without providing a speed figure for that update.
Clef-omni is available through Cloudflare’s API, and Cloudflare has released its open weights on Hugging Face.



