Generate, then decide: using Cloudflare Clef as a decision layer
Most LLM pipelines have one model do two jobs: write the answer, then judge whether it's good. Those are different problems. Writing is open-ended. Judging is a bounded choice: supported or not, safe or not, ship or not. Clef is Cloudflare's model for the second job. You don't prompt it for prose. You give it a state (text, JSON, images) and a schema of typed questions , and it returns a…
Cloudflare Clef serves as a decision layer in LLM pipelines, separating the tasks of generating an answer and judging its quality. Unlike standard pipelines where a single model handles both jobs, Clef focuses solely on the judgment aspect. Given a state, schema of typed questions, and evidence, Clef returns probabilities for each allowed option. This enables structured decision-making where code can validate the outcome, rather than parsing sentences.
A proof of concept, Evidence Lab, demonstrates Clef's efficacy in verifying whether a Retrieval-Augmented Generation (RAG) answer is supported by its source documents. For instance, a policy stating members can return unopened items cannot be validated by merely checking citation accuracy; the answer's meaning must also be correct. Clef excels in this scenario, providing a separate classifier that reads both the evidence and claim to assign a definitive label.
A Clef call comprises three components: state, questions, and answers. State represents the material to be evaluated, questions are up to 64 typed queries—each with an ID, and answers provide probabilities per question. There are three question types: yes/no (noul), multiple-choice (choice), and ordered scale (score).
For example, consider a triage scenario where a system failure query must be routed to the appropriate team. The state includes the failure description, questions include "Which team should handle this request?" with options for billing, technical, and sales teams, and answers provide probabilities for each choice. The code validates the schema, citation IDs, and quoted text, ensuring the draft's validity before a Clef call is made. This approach prevents the model from "chatting back" or causing unnecessary retries.
In Evidence Lab, evidence retrieval is fixed to ensure consistency. The generator produces answer blocks, each associated with citation IDs. These blocks are then checked by Clef for support and adherence to global criteria—such as staying within task scope, internal consistency, and absence of contradictory evidence. If an answer fails these checks, a single repair is allowed, iterating through the same validation process. If repair fails, the system returns an abstention, preventing publication.
Technical failures are handled by stopping the run on errors like invalid data, missing checks, hash mismatches, timeouts, or budget exhaustion. To ensure reliability, all verified questions must be answered exactly once, and global checks must pass before the system releases a response.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.