SLATEMOTH / DECISION MODELS
Clef and Jev: probability outputs still need action policy
Cloudflare's Clef puts typed decision models into more workflows. Useful adoption depends on clear choices, local calibration and a separate policy for taking action.
A decision can have a smaller interface
On October 1, Cloudflare announced Clef and Clef-flash, decision models hosted on Workers AI with Apache-2.0 weights. The announcement explicitly acknowledges Jev's influence and claims API compatibility. That is evidence of competition and interface convergence; it does not establish copying of proprietary work. [1]
Clef's model card describes a state and a schema of typed questions as input, with probabilities for the allowed options as output rather than free-form prose. This makes classification easier to connect to code. It also moves an important design decision upstream: someone must decide which answers the system is allowed to consider. [2]
Define the question before measuring the answer
Imagine a support team sorting requests into billing, technical and sales queues. A message about a charge caused by an outage could reasonably fit two categories. A well-formed response cannot repair overlapping definitions. Before replacing the model, clarify whether the question concerns the root cause, the team that owns resolution or the next operational step.
For a choice question, Clef's published interface scores the named options. If none of those descriptions fits the incoming case, choosing the highest-scoring option is still a choice within the supplied set. An explicit review or unknown category may help, but its usefulness also needs testing. It is a policy design, not an automatic guarantee of out-of-distribution detection. [2]
A confidence number is not a success rate
A hypothetical confidence of 0.9 should not be interpreted as nine correct decisions in every ten future cases without evidence from the relevant workload. Evaluate groups of predictions against labeled outcomes, and examine whether rare languages, ambiguous requests or missing context behave differently. A confident result can still be wrong when the deployment data changes.
Cloudflare describes training objectives intended to improve calibration, including a Brier loss. That is a meaningful design choice, but the training method alone does not establish calibration on your organization's inputs. The published model card also reports mixed results across benchmarks; a favorable aggregate comparison is not a universal ranking for every workflow. [1] [2]
Compatible APIs still require local thresholds
A compatible request and response shape can reduce migration work. It does not mean a threshold selected for Jev can be carried unchanged to Clef. The same number may select a different set of cases, and those cases may have different error costs. Compare the decisions the threshold admits, not just whether the client parses the response.
The actual input path matters too. Workers AI documents truncation of long text state to fit the token limit, while the local model-card example has its own configurable input bound. Do not compare hosted and local results as if their retained evidence were necessarily identical. Put decisive context in a documented input contract and test boundary cases. [2] [3]
Classification and authorization are separate steps
Sorting a ticket into a queue and authorizing a refund are different operations. For content teams, identifying a likely topic and publishing it to a brand account are different too. The model's answer can inform an action policy; it cannot supply the missing permission, budget or evidence that the policy requires.
This does not imply that every classification needs human approval. Low-cost, reversible labels may reasonably be automated with an appropriate evaluation and a correction path. Actions with more expensive consequences need tighter admission criteria or review. Match oversight to the consequence, rather than adding a manual gate everywhere or removing it because the output has a type.
Test the policy around the model
Start with representative cases and clear reference decisions. Keep a separate evaluation set that was not used to select the schema or threshold. Include ambiguous cases, insufficient evidence and cases where a wrong positive decision is especially costly. Measure what the chosen action policy gets wrong alongside how often it hands work to a person.
A shadow run can compare a proposed model with the existing process without executing its suggested actions. Record the input and schema versions, the model, its scores and the eventual outcome. Observe whether disagreement comes from the classification, a changed business rule or missing input. Those explanations determine whether to change the model, the threshold or the workflow.
The strongest objection is operational cost: not every small label deserves a large evaluation project. A narrow, reversible task can start with a modest comparison and clear correction mechanism. The essential requirement is proportionality and a record of the remaining uncertainty, not an elaborate benchmark that the team cannot maintain.
Specialization makes the surrounding decisions clearer
Decision models point toward workflows that assign different jobs to different components: classification for a bounded choice, generation for an explanation or draft, and policy for permission to act. This is a plausible direction for system design, not proof that one architecture will replace general-purpose models or complex planning.
An open model and a familiar interface lower the barrier to experimentation. They also make a useful purchasing question more concrete: which component is being replaced, and what observable behavior must stay intact? Access to weights is valuable, but it does not remove hosting, version management or workload evaluation.
For teams considering Clef, the first milestone should be a clearly defined decision that performs acceptably on their own cases. The next should be an action policy that handles mistakes and uncertainty. A fast, typed answer becomes useful intelligence when the organization knows what it means and what it is allowed to trigger.