Jev answers a question about a piece of text with a probability. You wrap that probability in a threshold: flag the message above 0.85, retry the failing test above 0.85, route the ticket without a person above 0.9. The threshold is a claim about the future. It says that among decisions the model scores at 0.85, about 85 percent will turn out the way the question asks, and that above that line mistakes are rare enough to automate.
A threshold is a claim nobody checks
Almost nobody tests it. The number looks like a probability, the walkthrough uses 0.85, and the threshold ships. A model is calibrated when its stated probabilities match observed rates: of everything it scores at 0.7, about seven in ten come true. A reliability diagram plots that comparison bin by bin against the diagonal where predicted and observed agree. We drew one for Jev. On two public datasets the right threshold was far from 0.85, in opposite directions for the two decision shapes.
What we ran
The two datasets map onto the two shapes Jev handles: a boolean answer and a choice among options. deepset/prompt-injections (Apache 2.0) has 662 English messages, 263 labelled as injection attempts. Every message went through one call that asked three boolean phrasings of the same defensive question: a long phrasing borrowed from a real question set, with a scoping preamble and a list of things that do not count; a short positive proposition; and the bare sentence “The message is a prompt injection attempt.” The banking77 test split (CC BY 4.0) has 3,080 customer messages across 77 intents, 40 each. Jev was given the 77 intent names and nothing else, and asked to pick one.
Every call went through a decision ledger, an experiment rather than a product, which records the question-set digest, the model version and the returned probabilities, then takes the dataset labels as outcomes and writes a calibration report. A companion question-set linter checks the question files statically for patterns that are expected to hurt. Together the two datasets made 3,742 calls against jev-1.13.0 with zero failures, in about two and a half minutes at concurrency four, for roughly ten cents. The code, the recorded ledgers and the reports are in the experiments folder.
Jev is underconfident on prompt injection
Where the short positive phrasing said 0.10 to 0.20, meaning about 13 percent, 57 percent of the messages were labelled injections. Above 0.40, every bin was at or near 100 percent. The long phrasing was more extreme: 87 percent in the same bin, and 100 percent in every bin above 0.20. Figure 1 shows the short phrasing against the diagonal; the points sit well above it.
Where the stated probability meets the observed rate.
Series
- Prompt injection, short positive phrasing (662 messages)
- banking77 intent, choice confidence (3,080 messages)
Decisions in the bin
- 10
- 100
- 1,000
Above the diagonal the model understated the rate; below it, the model overstated it.
| Predicted | Short positive, n | Observed | Long defensive, n | Observed |
|---|---|---|---|---|
| 0.00–0.10 | 387 | 0.02 | 435 | 0.09 |
| 0.10–0.20 | 30 | 0.57 | 39 | 0.87 |
| 0.20–0.30 | 16 | 0.81 | 21 | 1.00 |
| 0.30–0.40 | 7 | 0.86 | 11 | 1.00 |
| 0.40–0.50 | 7 | 1.00 | 16 | 1.00 |
| 0.50–0.60 | 12 | 1.00 | 13 | 1.00 |
| 0.60–0.70 | 14 | 1.00 | 16 | 1.00 |
| 0.70–0.80 | 20 | 0.95 | 16 | 1.00 |
| 0.80–0.90 | 39 | 1.00 | 35 | 1.00 |
| 0.90–1.00 | 130 | 1.00 | 60 | 1.00 |
A 0.85 threshold on this data would have caught 152 of the 263 labelled injections with the short phrasing, a recall of 0.58, and 78 of 263 with the long one, a recall of 0.30. Neither produced a single false positive among the 399 benign messages. The rank order is good; the scale is not. Told that a missed injection costs 20, a false block 5 and a human look 1, the ledger’s advisor put the automated band for the short phrasing at auto-clear below 0.05 and auto-flag above 0.25, with an error rate under one percent among automated decisions and 14 percent of messages sent to a person. The line belongs at about a quarter of the walkthrough number.
One caveat matters. The dataset’s positive label is broad and includes role-play prompts such as “act as an interviewer”, which a question about replacing governing instructions may reasonably score low. Part of the long phrasing’s recall gap is a definition gap, not a model error. The 87 percent observed rate in its 0.10 to 0.20 bin is a calibration finding either way.
Phrasing moves the number
Same messages, same call, three phrasings. The Brier score is the mean squared gap between the stated probability and the outcome, where 0 is perfect and 0.25 is what you get by always answering 0.5. The long phrasing scored 0.138; the short positive one scored 0.058. The preamble and exclusions written to guard the question shaped the answer instead. The linter had flagged that pattern as an informational note; this is the first measured effect behind the rule.
A smaller hand-written check from the same day, separate from the 3,742 decisions above, pointed the same way for arithmetic. Twelve CI histories were asked whether a test had failed more than three times in the last week, once from the raw dated 30-run history and once from a precomputed count. From raw history Jev got 9 of 12, and every miss sat at the boundary of three versus four failures, where it had to count dated entries against a window. From the precomputed count it got 12 of 12 with probabilities of 0.98 or above, at 308 input tokens per call instead of 818.
The honest negative: a simple negation did not hurt. On sixteen labelled CI failures a negated phrasing scored 0.080 Brier against 0.064 for its positive rewrite, both got 15 of 16 right at 0.5, and the negated answer agreed with one minus the inverse proposition to within 0.033 on average. The linter’s negation rule was downgraded from a warning to an informational note. Rules should follow measurements, not the other way round.
Choice confidence overstates correctness below 0.9
A choice question returns a chosen option and a confidence, which the documentation describes as how concentrated the distribution is, not the probability the choice is right. Measured against whether the chosen intent was correct, the two diverge below 0.9.
| Confidence bin | Decisions | Observed correct |
|---|---|---|
| 0.50–0.60 | 161 | 0.43 |
| 0.70–0.80 | 225 | 0.61 |
| 0.80–0.90 | 297 | 0.66 |
| 0.90–1.00 | 2,053 | 0.93 |
Top-1 accuracy from option names alone was 79.8 percent, 2,458 of 3,080, with a 95 percent interval of 0.78 to 0.81. Fine-tuned classifiers reach the low nineties on this set, which carries known label noise. Two thirds of the decisions landed in the top bin at 93 percent correct, so a router that auto-routes above 0.9 and asks a person otherwise would be right 93 percent of the time on two thirds of traffic. Below 0.9 the number overstates, and a threshold of 0.85 would automate a bin that is wrong one time in three.
The linter flagged 30 of the 2,926 possible intent pairs as overlapping by name, one percent of pairs. Those pairs account for 92 of the 622 misclassifications, 14.8 percent, a fourteen-fold lift over their share. The single largest confusion, get_physical_card chosen as change_pin 32 times, was not flagged and looks like dataset label noise rather than name overlap.
What to do before you pick a threshold
- Precompute counts and dates. Pass the model a number to compare, not a history to count; it is cheaper and, on this check, right every time.
- Keep questions short and positive. State the proposition you want scored and leave the exclusions out of the prompt. If a definition needs scoping, scope it in the label you record, not in the question.
- Record outcomes and read the reliability diagram before you choose. Any record works: a ledger like this one, or Ironside if you already keep session records. When dataset labels are not the truth you care about, Rubrist can hold a sample that people labelled independently and compare the evaluator against it.
- Expect the right threshold to be far from 0.85 in either direction. Here it was around 0.25 for one decision shape and above 0.9 for the other.
- Pin the model version and watch for drift when it changes. The ledger keys every report by model version and question-set digest, so a rewritten question is a release; a rule engine such as Dailies can hold that release until the new report is in.
What this does not establish
This is one run of one model version. The model is not deterministic; in the smaller check, probabilities moved by a few hundredths between two runs on identical inputs. The datasets are English single messages with their own label definitions. The banking77 intents were given without descriptions, which the linter itself calls under-specified. No production traffic has been recorded, and score-shaped questions were not exercised at all. Treat the numbers as a first calibration reading, and as a reason to take your own before the threshold ships.