Stop parsing LLM text for decisions
Decisions need types and probabilities, not prose. Why parsing a model's sentence fails quietly, and what a typed, calibrated decision gives you instead.
LLM structured decisions are answers you get back as a declared type (a label, a yes or no, a level on a scale, a number) with a probability attached, instead of a sentence you parse. Routing a ticket, flagging an email or checking whether an answer is grounded are decisions with a fixed set of outcomes. Asking a text generator to write them in prose and then reading them back is where the bugs hide, because a parser fails quietly. This post spells out how prompt-and-parse breaks, and what a typed decision with a checked probability gives you instead.
Count the LLM calls in a production codebase. Some produce text a person reads. Many produce text that a function immediately turns back into a label, a boolean or a number. The second group is the one to fix.
The prompt-and-parse pattern
Here is a pattern most teams have written, in some form:
reply = llm("Which team should handle this ticket? Reply with one of: billing, technical, sales.\n\n" + ticket)
team = reply.strip().lower().rstrip(".")
if team not in {"billing", "technical", "sales"}:
team = "technical"Each line patches a failure someone saw once: trailing punctuation, capital letters, "Billing team", an explanation before the answer. The last line is the dangerous one. When the reply doesn't parse, the code picks an answer and moves on. That choice never shows up in a metric.
The fix is not a better regex. It is not asking for text at all.
Four ways it fails quietly
None of these raise an exception. They show up weeks later as a skew in your data.
Why LLM structured decisions need types
A typed decision declares its outcomes up front, and the answer always comes back as one of them. In Curva you declare questions such as Choice, Score, Noul (yes or no), Multi, and Text, Number and Integer for extraction. Answers are only ever mapped onto the labels you declared, never parsed from free text, so there is nothing to validate on your side.
The three-line version, from the docs:
import curva
d = curva.decide("I was charged twice, please refund me",
{"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total) # billing True None`d.team` is always one of your options. `d.refund` is `True` when P(yes) is at least 0.5. `d.total` is a number, or `None` when the text has none. And `d["team"]` holds the full answer, with `confidence` and a probability for every option.
Types also give the model two honest exits, which free text never has.
Both are routing decisions in their own right. And the person's answer is the most reliable label you will get, so send it back.
Probabilities you can check
"Billing" and "billing, but only a little more likely than technical" are different answers. The first gets routed. The second deserves a second look. A label without a probability forces you to treat every answer as equally sure.
The probability has to come from the model, not from text it writes. Curva reads it from token log-probabilities when the model returns them (`logprobs` mode), or asks for JSON limited to your labels with a probability for each (`verbal` mode). It also asks every question in two option orders and averages them, which cancels position bias.
Even a probability read from logprobs is not yet a probability that is right. The only way to know is to compare stated confidence with outcomes on your own data, and adjust. That is calibration. Send the true answer when you learn it; after 30 labels for a question, Curva fits a small calibrator, keeps it only when it improves on the raw probabilities on held-out labels, and reports before and after. Once calibrated, a 0.9 means right about 90% of the time on your data. Before that, it is the model's raw estimate.
Measure the model too, not just the system around it. Two small findings from Curva's own runs show why:
The second result is why "confident" and "right" need separate evidence.
A record for every decision
When a decision is wrong, someone will ask why it was made, by which model, under which prompt. Every Curva decision is written to an audit log with its id, project, model, mode, config, answers, latency and cost, and a salted hash of the state instead of the state itself. You can match a record to an input you still have without storing the input.
Two habits make the record useful:
What to keep in code
The other half of the argument is what not to ask. Language models are poor at counting, arithmetic and date comparison. If a decision depends on "more than three open tickets" or "older than 30 days", compute that in code and put the result in the state.
If the rule is fixed, don't call a model at all. Curva's `rules` answer a question on the spot when a condition matches, with no model call and no cost, and say which rule fired. Save the model for what needs reading between the lines.
A checklist for LLM structured decisions
Whatever tool you use:
Text generation is a good interface for people. For decisions, ask for the type you need, and for evidence that the numbers mean something.
Next steps
Read [what is a typed decision](/blog/what-is-a-typed-decision/) for the building blocks, and [LLM calibration explained](/blog/llm-calibration-explained/) for how probabilities get checked. For the "not sure" exit, see [LLM confidence threshold and abstain](/blog/llm-confidence-threshold-abstain/). To weigh this against fine-tuning and other options, read [LLM classification approaches compared](/blog/llm-classification-approaches-compared/). The docs start at [questions and answers](https://itsmohitrohilla.github.io/curva-docs/concepts/questions/). Install with `pip install curva-ai`.