When you change a prompt, model, or agent workflow, you need a way to check whether its answers improved. If another model checks them, you also need to know whether its judgments make sense.

Rubrist helps you develop that check. You bring recorded examples, define one thing that matters, run an evaluator, and inspect where its results need human attention. For more consequential decisions, you can validate the evaluator against separately reviewed examples.

Start with one quality question

“Was this good?” leaves too much open to interpretation. A useful check asks something narrow enough that another person can understand what passing means. Rubrist calls that question a criterion; the evaluator is the versioned mechanism that answers it.

Figure 01

One example. Two independent judgments.

Same example + same criterion

Was the success claim supported?

Independent human reviewFail

The recorded evidence does not justify the claim.

Evaluator under testPass

The evaluator accepts the answer.

Compare after reviewA false pass

The evaluator accepted an example the reference judgment rejected.

Illustrative comparison with an independently reviewed reference label. Reviewers do not see the evaluator’s answer while labeling. Across many examples, these comparisons reveal false passes, false fails, coverage, and uncertainty.

Start with examples you already have. You can use a small dataset without connecting production tracing, or import from Ironside, LangSmith, or Langfuse. Rubrist also has an analysis workflow for reviewing a sample of traces and naming recurring failure modes.

Run a check on your examples

The Guided view uses three terms: a Run is recorded work, a Check evaluates one quality question, and a Result is what the check concluded. Start by running a check on an example you understand.

Review the reasoning, especially where you disagree. Was the instruction unclear? Did the evaluator overlook evidence? Does your definition of a passing result need to change? Those observations help you improve the evaluator.

A first result is not proof of accuracy. A label you supply during setup, or a correction made after seeing the evaluator’s answer, is useful development feedback. It is different from an independent assessment collected without that influence.

Check the evaluator against people

For independent validation, reviewers receive the example and the criterion without seeing the evaluator’s answer. Their labels and disagreements stay in the record. Resolving a disagreement does not erase the original judgments.

Rubrist can compare a binary evaluator against protected, independently reviewed examples. This is calibration: measuring how the evaluator agrees with the reference judgments, including the direction of its mistakes, how much evidence was classified, and the uncertainty in the measurements.

A false pass and a false fail can have different consequences. In the deployment example, overlooking an unsupported success claim is a different problem from rejecting a correctly verified deployment. Rubrist keeps those errors visible.

The current calibration runtime executes one binary trial. Repeated-trial execution and calibration of scalar or categorical evaluators are not currently implemented. A dataset used to tune an evaluator cannot simply be relabeled as protected validation data.

Keep evidence through changes

As you revise the criterion, prompt, or model, Rubrist records evaluator versions and their evidence. Regression examples from known failures help you check whether the revised evaluator still handles cases you’ve already reviewed.

You can group several focused evaluators into a suite and keep their results separate. Rubrist can export assessment evidence that another system verifies. Dailies can then apply your release rules; Rubrist itself does not decide whether your application or skill should ship.

Try Rubrist

For help setting up the tools around a specific question, build a local stack setup prompt for your coding agent.

Use the current quickstart to set up a workspace. Local development requires Node.js 24 or newer, pnpm, and Postgres; real model judging also needs a supported provider key. The in-memory demo is useful for exploring screens, but independent review and protected calibration require a persistent, authenticated workspace.

  1. Choose one failure you want your AI system to avoid.
  2. Add a few recorded examples that show what happened.
  3. Define a check and inspect its first results.
  4. Improve the evaluator from development feedback before collecting separate validation evidence.

To retain the original session records, see Ironside.