Detect refund requests, not the word refund
A yes/no question that separates "I was charged twice" from "please refund me": the recipe's wording, the probability it returns and how to calibrate it.
Refund request detection works when you ask whether the customer asks for money back, not whether the word "refund" appears. "I was charged twice" is a complaint about a charge. "Please refund the duplicate" is a request. A keyword rule flags both. A yes/no question worded as "the customer explicitly asks for money back" separates them, and Curva returns the probability of yes rather than a word to parse. This post gives the wording from Curva's `support-triage` recipe, shows how to put a threshold on the probability, and how agent labels calibrate it.
Why keyword matching over-fires
A refund keyword list looks fine on day one. Then the misses arrive:
Each fix adds a rule, and each rule creates a new edge case. The real question is about intent, which is what a language model reads well. The problem is getting a clean, checkable answer out of it, not a sentence.
The Noul wording that works
Curva's `support-triage` recipe asks it as a Noul, its yes/no type:
"refund_requested": {
"type": "noul",
"instructions": "The customer explicitly asks for money back (a refund, chargeback or credit). Complaining about a charge without asking for money back is no."
},Three parts do the work. "Explicitly asks" rules out guesses about what the customer might want later. "(a refund, chargeback or credit)" lists the forms money back can take, so "put it back on my card" counts. And the last sentence names the main false positive and says it is a no. Spelling out the "no" case is the single most useful habit in yes/no wording.
For a quick script, the Python shorthand works too: a string ending in `?` is a Noul.
import curva
d = curva.decide("I was charged twice, please refund me",
{"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total) # billing True None`d.refund` is `True` when P(yes) is at least 0.5. For production, use the recipe's fuller wording, because the shorthand can't carry the "no" case.
P(yes) and a threshold for refund request detection
A Noul's answer is one number, `noul`, the probability of yes. Internally it is asked as a normalised two-option choice, so P(yes) and P(no) stay consistent with each other.
from curva import Curva, Noul
client = Curva()
q = {"refund_requested": Noul("The customer explicitly asks for money back (a refund, chargeback "
"or credit). Complaining about a charge without asking for money "
"back is no.")}
d = client.decide({"ticket": ticket_text}, q, project="support")
p = d["refund_requested"].noul
if p >= 0.9:
open_refund_case(ticket_id, decision_id=d.id)
elif p >= 0.5:
tag_for_agent(ticket_id, "possible refund")Two thresholds give you three outcomes: act, flag for a person, or do nothing. Set them by cost. Opening a refund case for a customer who didn't ask wastes an agent's time and can confuse the customer. Missing a real request costs goodwill and, sometimes, a chargeback. Most teams want a high bar for automatic action and a lower bar for a flag.
If you need a promise rather than a threshold, set `coverage` on the Noul. Once the question has 30 labels, the answer carries a `set` of `"true"` and `"false"` values that contains the right answer at least that often, as long as new tickets look like the labeled ones. A set with one value can be automated; a set with both goes to a person.
Calibrate on agent labels
A raw 0.9 from a model is not a measured 90%. Raw LLM confidence is often too high. Calibration fixes that with your own labels.
Your agents already decide whether each ticket was a refund request: they either process a refund or they don't. Keep the decision `id` with the ticket and send that outcome as feedback:
client.feedback(decision_id, "refund_requested", True) # the agent processed a refund request client.feedback(other_id, "refund_requested", False) # a complaint, no request
After 30 labels for this exact question in the project, Curva fits a Platt scaling calibrator, a logistic fit on P(yes). It is applied only when it makes the probabilities more accurate on held-out labels. If the model is already well calibrated, answers stay raw with `calibrated: false`. The docs' tests show the kind of fix it makes: a yes/no model that always says 1.0 moves to its true 70% as labels accumulate.
Label the easy tickets as well as the hard ones. If only the flagged tickets get labels, calibration learns nothing about the confident end. And settle the wording first: calibration belongs to the exact question, so editing the instruction starts it over.
Check the result with the calibration report:
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])Pass `"refund_requested"` instead of `"team"` for this question. `after` is held out, and `accuracy_when_automated` tells you how often answers at 0.9 or above were right, and how many reached it.
Refund and team in one call
A refund flag rarely travels alone. You also need to know which team handles the ticket and how upset the customer is. Ask everything in one request; all questions are answered together:
curva recipe show support-triage > questions.json
The recipe asks `team` (a Choice that abstains below 0.8), `urgency` and `frustration` (Scores), and `refund_requested` and `churn_risk` (Nouls). The `team` instruction carries the same idea as the refund wording: "Judge by what the customer needs done, not by the words they happen to use."
Load it in Python with `json.load` and pass it as the questions. Edit the wording for your domain, but do it before you start sending feedback.
Next steps
Read the [recipes guide](https://itsmohitrohilla.github.io/curva-docs/guides/recipes/) and the [feedback guide](https://itsmohitrohilla.github.io/curva-docs/guides/feedback/) in the docs. Install with `pip install curva-ai`. For yes/no questions in general, see [LLM yes/no probability](/blog/llm-yes-no-probability/). For the full ticket pipeline, read [support ticket triage with AI](/blog/support-ticket-triage-ai/), and for the cancellation flag, [churn risk in support tickets](/blog/churn-risk-support-tickets/). More ideas are in [LLM classification use cases](/blog/llm-classification-use-cases/).