Urgent.News

What's breaking now, across thousands of outlets.

AI

Generate, then decide: using Cloudflare Clef as a decision layer

Most LLM pipelines have one model do two jobs: write the answer, then judge whether it's good. Those are different problems. Writing is open-ended. Judging is a bounded choice: supported or not, safe or not, ship or not. Clef is Cloudflare's model for the second job. You don't prompt it for prose. You give it a state (text, JSON, images) and a schema of typed questions , and it returns a…

Cloudflare Clef serves as a decision layer in LLM pipelines, separating the tasks of generating an answer and judging its quality. Unlike standard pipelines where a single model handles both jobs, Clef focuses solely on the judgment aspect. Given a state, schema of typed questions, and evidence, Clef returns probabilities for each allowed option. This enables structured decision-making where code can validate the outcome, rather than parsing sentences.

A proof of concept, Evidence Lab, demonstrates Clef's efficacy in verifying whether a Retrieval-Augmented Generation (RAG) answer is supported by its source documents. For instance, a policy stating members can return unopened items cannot be validated by merely checking citation accuracy; the answer's meaning must also be correct. Clef excels in this scenario, providing a separate classifier that reads both the evidence and claim to assign a definitive label.

A Clef call comprises three components: state, questions, and answers. State represents the material to be evaluated, questions are up to 64 typed queries—each with an ID, and answers provide probabilities per question. There are three question types: yes/no (noul), multiple-choice (choice), and ordered scale (score).

For example, consider a triage scenario where a system failure query must be routed to the appropriate team. The state includes the failure description, questions include "Which team should handle this request?" with options for billing, technical, and sales teams, and answers provide probabilities for each choice. The code validates the schema, citation IDs, and quoted text, ensuring the draft's validity before a Clef call is made. This approach prevents the model from "chatting back" or causing unnecessary retries.

In Evidence Lab, evidence retrieval is fixed to ensure consistency. The generator produces answer blocks, each associated with citation IDs. These blocks are then checked by Clef for support and adherence to global criteria—such as staying within task scope, internal consistency, and absence of contradictory evidence. If an answer fails these checks, a single repair is allowed, iterating through the same validation process. If repair fails, the system returns an abstention, preventing publication.

Technical failures are handled by stopping the run on errors like invalid data, missing checks, hash mismatches, timeouts, or budget exhaustion. To ensure reliability, all verified questions must be answered exactly once, and global checks must pass before the system releases a response.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

How I Built an n8n AI Voice Outreach Workflow for Roofing Leads

I recently built an automation system that connects lead discovery, AI personalization, voice outreach, and lead tracking into a single n8n workflow.

  • Author built n8n AI workflow for roofing leads
  • Workflow integrates lead discovery, AI personalization, voice outreach, tracking
  • n8n chosen for its service specialization and predictability

More from Tuesday 6 October →