Decision Models Give Local AI Systems a Narrower, Testable Contract

llama.cpp now exposes decision models through a System One-compatible endpoint. Their constrained outputs can simplify routing and validation, but probabilities still require task-specific calibration and fallbacks.

Source artwork for New in llama.cpp: Decision Models
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Use decision models for bounded choices, then validate option wording, calibrate thresholds on held-out workflow data, and provide an explicit abstention path.

The llama.cpp server now supports decision models through a /v1/systemone endpoint. Instead of generating a free-form answer token by token, these models score options supplied by the caller and return structured probabilities. The October 2 announcement describes three question forms: selecting among choices, estimating a level on an ordered scale, and assigning a yes probability to a binary question. A request can contain multiple questions about the same text, JSON state, or—when the chosen model supports it—image.

That interface is more than a different response format. It gives an application a deliberately narrow contract. A routing component can ask for one of several named queues; a workflow checker can estimate whether a step succeeded; a moderation layer can score a defined policy question. The result is easier to validate mechanically than prose, although it is not automatically correct or calibrated.

Constrained answers change the integration problem

A generative model asked to classify a support ticket may return a label, an explanation, extra punctuation, or an unexpected category. The application then needs parsing rules and must decide what malformed output means. A decision model receives the permitted alternatives as part of the request, so the caller can reason about a fixed result space.

Consider a help desk with billing, delivery, and account-access queues. Each option should include a short operational description, not just a label. The source reports an example in which adding descriptions changed a small model's routing of a duplicate-charge message from shipping to billing. That observation should not be generalized into an accuracy guarantee, but it illustrates that option design is part of the model input. Renaming or rewriting criteria is therefore a behavior change that deserves regression testing.

The probability vector also enables policies that a single label cannot express. An application might accept a clear result, request a second check in a middle band, and send an ambiguous case to a person. Those boundaries are product decisions. A value that works for one model and dataset cannot safely be copied to another.

Build the evaluation around decisions, not demos

Start with examples drawn from the actual workflow, including ordinary cases, costly mistakes, vague inputs, multilingual text where relevant, and states that belong to none of the listed options. Label them independently of the model. Then evaluate the action the system would take at each candidate threshold.

For routing, a confusion matrix shows which queues are being mixed up. Precision and recall per queue are more useful than accuracy alone when one destination is rare. For a yes/no safeguard, examine false approvals and false rejections separately because their costs may differ. For an ordered score, measure both the size and direction of errors; confusing adjacent urgency levels is not equivalent to choosing opposite ends of the scale.

Probability quality needs its own check. Group predictions into ranges and compare stated confidence with observed success on held-out cases. If predictions near 0.8 are correct much less often than that, the number should not be treated as an eight-in-ten guarantee. Calibration can also drift when the model, quantization, criteria, input population, or prompt template changes. Preserve those details alongside every evaluation result.

Design an explicit abstention path

A constrained option list can hide a specification gap. If a ticket concerns legal correspondence but the only choices are billing, shipping, and technical support, the model must still distribute probability among unsuitable answers. Add an other or needs_review route when the workflow permits it, and test inputs that should reach it.

Low confidence is another useful abstention signal, but it should not be the only one. Deterministic rules can catch missing required fields, unsupported file types, or policy categories that always require human review. A robust sequence is: validate the request, obtain model scores, apply task-specific thresholds, and log the decision plus the model and criteria versions. Do not log sensitive source content unless the application's retention policy allows it.

Fallback behavior must be concrete. If the local model cannot load, the image projector is unavailable, or the endpoint times out, decide whether work waits, follows a conservative default, or enters a manual queue. Silent fallback to an arbitrary top option defeats the value of the structured interface.

Choose models and deployment settings empirically

The announced collection spans small text encoders through larger language-model-derived systems, and the listed capabilities and licenses differ. OpenJev is identified as image-capable at announcement time and carries a non-commercial license, while the listed text-only models use Apache 2.0. Teams should verify the model card and current license for the exact artifact before deployment rather than assuming every item in the collection has the same terms.

Model size, quantization, latency, memory use, language coverage, and error cost all belong in the selection record. The source includes speed measurements on one specified GPU; those figures are useful as release context, not predictions for another machine. Benchmark the final GGUF artifact on the target hardware with realistic state lengths and question counts. Router mode can load different models on demand, but that convenience introduces cold-load behavior and capacity considerations that should appear in operational tests.

Decision models are most attractive where the possible actions are known in advance and errors can be measured. They do not replace open-ended generation, nor do their probabilities eliminate judgment. Their practical advantage is a smaller, auditable boundary between model inference and application logic—provided teams define that boundary carefully, evaluate it on their own cases, and keep a safe route for uncertainty.

Source: New in llama.cpp: Decision Models ↗. How we write

← Back to all articles