Most programs contain judgment calls that nobody automated because the judgment was too slow or too expensive to make often. Which queue does this ticket belong in. Is this tool call the one the user asked for. Does this sentence make a claim without a source. A general language model can answer each of those, but at several seconds and a few cents per answer, so you ask rarely, cache aggressively, and keep the model out of anything that loops.

On 15 September 2026, TypeSafe AI released Jev, the first of what it calls System One models. It does not write text. It takes a piece of program state and a set of typed questions, and returns a typed answer for each, with a probability distribution. TypeSafe’s own figures put a decision at well under a second and a fraction of a cent. If those figures hold on your data, the wall moves: a judgment stops being something you ration and becomes something you can make per item, per step, and per tick.

This is a working note on where that matters: a few possible ideas beyond the obvious use, with the vendor’s numbers kept separate from anything I would treat as settled.

Answers instead of text

A request to Jev has two parts. The state is any text or JSON you want judged. The questions are a map of names to one of three question types:

  • Choice picks one option from a set you define, up to 255 of them. It returns the chosen key, a probability per option, and a confidence value.
  • Score places the state on an ordered rubric of two to ten levels described in words. It returns a continuous position, which can land between levels, plus the distribution and a confidence value.
  • Noul is a yes-or-no statement. It returns one number from 0 to 1: the probability the statement is true. There is no separate confidence; the probability is the signal.

Every question in a request is evaluated in parallel and in isolation, so adding a tenth question costs tokens but barely any time. The answer always fits the schema you gave; a Choice cannot return an option you did not list. That is a structural guarantee, not an accuracy claim. An answer can be well typed and still misread the evidence.

The example from TypeSafe’s quickstart, adapted to the JavaScript SDK, shows the shape:

import { TypeSafeClient, choice, score, noul } from '@typesafe-ai/sdk';

const client = new TypeSafeClient(); // reads TYPESAFE_API_KEY

const { answers } = await client.systemOne({
  state: { ticket: ticketText },
  questions: {
    team: choice('Which team should handle this?', {
      billing: 'Payment or subscription issues',
      technical: 'Bugs or integration problems',
      other: 'Anything else',
    }),
    frustration: score('How frustrated does the customer sound?', [
      'Calm, stating facts',
      'Frustrated but civil',
      'Very angry, strong language',
    ]),
    urgent: noul('The message asks for help right now.'),
  },
});

answers.team.choice;        // 'technical'   (illustrative values)
answers.team.confidence;    // 0.91
answers.frustration.score;  // 1.2
answers.urgent.noul;        // 0.97

The types of answers are inferred from the questions, so a typo in an option name is a compile error rather than a runtime surprise.

A smart if statement

The mental model that makes Jev useful is not “a cheaper language model”. It is a function call: unstructured state in, a typed probabilistic decision out. Your code keeps the control flow. The model never decides what happens next; it answers a narrow question, and an if you wrote does the rest.

Figure 01

The model answers. Your code keeps the control flow.

  1. 1
    Code assembles the state

    Only the fields the questions need: the ticket text and the order it references. Nothing the model could be distracted by.

  2. 2
    Jev answers every question at once

    Three questions travel in one request and are evaluated in parallel. Each answer comes back typed, with a probability distribution.

  3. 3
    Code decides what to do

    Thresholds you chose turn probabilities into actions. Uncertain cases go to a person instead of a guess.

team · Choicetechnicalconfidence 0.91
frustration · Score1.2 of 2confidence 0.74
urgent · Noul0.97probability the statement is true
confidence ≥ 0.85Route automatically

to the technical queue

0.5 to 0.85Route, but flag

a person confirms before reply

below 0.5Send to a person

the model does not act

A Noul carries no separate confidence: the probability itself is the signal, and 0.5 means genuinely unsure rather than “medium”.

Illustrative request for a support ticket; the numbers are invented, not a recorded run. The thresholds are one example policy. TypeSafe’s documentation recommends choosing them per action from the cost of a wrong decision.

Two habits follow from this. Decompose broad judgments into atomic questions, because each atomic answer is something code can inspect, threshold, and combine. And treat confidence as a routing input: a high-confidence answer acts, a middling one asks for confirmation, a low one goes to a person. TypeSafe trained the model to make those probabilities calibrated in aggregate, meaning answers given 0.9 should be right about nine times in ten across many cases. That is a property to verify on your own data, not to assume.

Six places it changes the design

The standard use, ticket routing, is real but leaves the architecture where it was. The interesting uses each break an assumption that only held because judgments were expensive. Ranked from most to least buildable today.

1. Verify every step, not just the final answer

Agent pipelines check outputs rarely because a checker costs as much as the worker. When the check costs a hundredth as much, you can check everything: each tool call before it executes (“this call matches what the user asked for”), each extracted field (“this value appears in the source”), each cited passage (“the passage supports the claim”). Jev cannot generate the fix, but it can stop the wrong thing from happening and hand the case to a person or a stronger model. This is the pattern TypeSafe documents as guardrails and citation checking, and the one LangChain’s harness uses to gate risky tool calls in auto mode.

The cost is that a verifier at 67.8% accuracy on the vendor’s workflows is a filter, not a guarantee, and Jev does not treat user-supplied text as hostile. Content inside the state can move the answer. Use it to catch ordinary mistakes; do not use it as your only defence against injection.

2. Store the probabilities, decide at read time

Classification is usually a label written once and kept. That makes every policy change a migration. If the judgment itself is cheap and the answers are numbers, invert it: persist the full distribution and the model version, and derive the decision with thresholds in code when you read the record.

// Persist the judgment, not the verdict.
await db.insert('ticket_judgments', {
  ticketId,
  model: response.model,     // 'jev-1.13.0'
  answers: response.answers, // distributions and confidence
});

// Decide at read time. Change the policy and every past record follows.
function action(a) {
  if (a.team.confidence < 0.5) return 'human';
  if (a.urgent.noul > 0.9 && a.frustration.score > 1.5) return 'page';
  return 'queue';
}

Every decision becomes reversible: tighten a threshold and the old cases re-sort without another request. The audit trail is a set of probabilities rather than a written rationale, which is a real loss in regulated work and a gain everywhere the rationale would have been decorative. The cost is discipline around model versions; pin one, and re-judge when you move.

3. Judge the sentence, not the document

A document-level verdict hides where the problem is. Per-line judgments were unaffordable; at Jev’s prices a thousand lines of “does this sentence assert a fact without a source” costs well under a dollar. That turns fuzzy standards into linters: writing guidelines, code review conventions, contract clauses, moderation rules applied line by line with a location attached. TypeSafe’s cookbooks call this line-by-line search and semantic linting.

What it costs is request volume, and the model’s literal reading. Each line needs enough surrounding context to be judged, and each question needs to say exactly what you mean; negations and scoping words are read at face value. A cheap test: plant twenty known violations in a document, run the linter, and see how many it recovers at the threshold you would ship.

4. Ask thirty questions and keep five

With a language model you ask only what you need, because each question is a round trip. With parallel evaluation you can fetch every judgment a branch might need in one request, then branch in code: intent, complexity, risk, language, sentiment, and whether it is spam, all at once. TypeSafe calls this speculative fan-out and reports that batching thirteen questions into one call was roughly ten times faster and twelve times cheaper than asking them in sequence. Those are their measurements, but the mechanism is plain: marginal latency per question is close to zero.

The trade is that the questions become a schema you own and maintain, and a poorly worded question in the batch is answered with the same confidence as a good one.

5. Turn the answers into features for your own model

A single Choice is the vendor’s judgment about your problem. A vector of twenty Nouls is a set of features you can fit to your own outcomes. Instead of asking “is this spam”, ask twenty narrow things: requests credentials, sender name conflicts with domain, announces an unexpected reward. Each Noul is a column. A logistic regression on those columns, trained on your labelled history, gives you a classifier tuned to your data with an interpretable weight per signal, and no fine-tuning.

signals = {
    "asks_for_credentials": Noul(
        instructions="The message asks the reader to enter or confirm a password or login."
    ),
    "sender_mismatch": Noul(
        instructions="The display name claims an organization that the email domain does not match."
    ),
    "unexpected_reward": Noul(
        instructions="The message announces a prize, refund, or reward the reader did not request."
    ),
    # ... fifteen more, each a column
}
# Fit a plain classifier on the resulting columns against your own labels.

This needs labelled outcomes and some care: the docs warn that questions asking similar things in different formats have no guaranteed relationship to each other, so validate the features rather than reasoning about them. The experiment is cheap: on five hundred labelled rows, compare the classifier’s area under the curve against the single-question baseline.

6. Judgment inside the loop

The least proven and most interesting: a judgment that runs at tick rate. Intent detection while someone types. A non-player character that reassesses the situation a few times a second. A browser agent that scores every candidate element on the page before clicking. Launch-week demos reported by TypeSafe and in the coverage around it include a drone advisory loop and a market-making bot; I have not run them, and I would read them as existence proofs rather than benchmarks.

The numbers do not quite reach a render loop. Seventy to five hundred milliseconds is a network round trip, not a frame, and every input is a possible adversary. Keep the control loop in code with the model advisory, exactly as the demo authors did.

What to keep away from it

TypeSafe publishes a candid list of where the current model is jagged. Read it before designing questions. The short version:

  • Anything generative. No prose, code, summaries, or explanations. If you need a value the model must invent, extract candidates in code or with a language model first and let Jev pick among them.
  • Counting and arithmetic. Error grows with the size of the list. Ask one Noul per item and add in code.
  • Dates and numbers as quantities. Dates are read as text, so “which came first” is unreliable. Parse the date in code and ask about what it means.
  • Long, unrelated context. Accuracy falls as irrelevant state accumulates. Retrieve and filter in code, and send only the fields the question needs.
  • Double negatives and contradictions. The model answers the question you wrote. Keep instructions and criteria pointing the same way, and test negations explicitly.
  • Open answer spaces or one-off hard reasoning. If the options are not enumerable, or the decision needs a written justification, use a reasoning model.
  • Images, audio, video. Text only for now. Transcribe or caption first.

Try it in an hour

The test that matters is not the vendor’s benchmark but agreement with judgments you already trust.

  1. Create a key in the TypeSafe console and install a client: npm install @typesafe-ai/sdk or pip install typesafe-sdk. Both read TYPESAFE_API_KEY from the environment.
  2. Take one hundred to two hundred records you have already labelled by hand: routed tickets, moderated posts, screened resumes. Write the label as a Choice with an explicit “other”, and add two or three atomic Nouls for the signals a reviewer would look at.
  3. Run one request per record with the version pinned, for example model: 'jev-1.13.0', and keep the full response. The alias jev-latest can change without notice, which matters once you start tuning thresholds.
  4. Measure two things. Agreement with your labels overall. Then agreement bucketed by confidence: if accuracy does not rise with confidence on your data, the thresholds are decoration and you should not route on them.
  5. Put it behind an if with a human path for the low bucket, and watch the disagreements for a week.

If you already record runs, that labelled set is the same evidence an evaluator needs. Coeval is built for exactly this comparison between an automated judgment and independent human review.

What I would still check

The performance figures in this note come from TypeSafe. The company reports 67.8% accuracy on a four-workflow benchmark it designed, level with a mid-tier reasoning model and behind the frontier ones, at roughly a two-hundredth of the cost and a fiftieth of the latency. The consensus labels were built by averaging two frontier models. Nobody has reproduced this on a neutral harness that I could find, and the input price of 4.2 cents per million tokens with free output tokens may or may not be sustainable.

Two limitations are structural rather than temporary. The model gives probabilities but no rationale, so any process that needs a written reason for a decision has to produce it elsewhere. And calibration is a claim about aggregates: a single 0.9 tells you nothing about that one case.

None of that changes the design argument. If a typed judgment costs almost nothing and takes a fraction of a second, the interesting question is not whether it can replace a language model. It is which parts of your program have been waiting for a cheap, consistent answer to a narrow question, and never got one.

Sources: TypeSafe’s launch post, documentation, and Jev 1.13 jaggedness notes; the LangChain harness post; and the practical guide by Valyu. Read on 20 September 2026.