Where would you like to start?

Choose how much guidance you’d like. You can switch paths at any time.

Start with made-up data. Before using real examples, confirm you’re allowed to use them. Calls to an external model can send their contents off your machine.

Before using real data

Try your first agent check

An evaluation, or “eval,” checks an agent’s work against a clear expectation. Let’s start with one: did the agent show that it ran tests before saying the work was done?

A made-up session
Agent says
“Fixed the bug. Tests pass.”
Recorded actions
Edited a file. No test result is recorded.

The record doesn’t show that tests passed, so the claim needs a closer look. Missing results don’t prove that no tests ran.

Copy this prompt into your coding agent. It uses an Overclock skill, a set of instructions your agent follows during setup. Your agent will explain the example and unfamiliar terms, then help you choose the next step. You won’t need your own session data.

I’m new to evaluations and use Codex. Help me try one simple check: did an agent show evidence that it ran tests before saying the work was done?

Use the local-eval-stack skill from https://github.com/luka-zivkovic/overclock. If it is missing, explain how to install it for my agent, with its supporting files, and stop there.

Start with a clearly made-up session. Show me the agent’s claim and recorded actions, then walk me through checking them. Explain unfamiliar terms as they come up and guide me through one decision at a time. Don’t ask me to repeat answers I’ve already given. Missing test results leave the claim unsupported by this record; they don’t prove that no tests ran.

Then explain which local tools would help me inspect the example and check what’s installed. Agree on the next step with me before changing anything. Don’t require an AI model to judge the example just so I can understand it.

Use no real session logs, customer data, private code, or credentials in the example. Keep ongoing capture and automatic judging off. Keep secrets out of commands and output. Before credential writes or external model calls, explain what would be written or sent, to whom, and at what cost, and ask for any authorization still missing. This walkthrough is for learning, not proof that an agent is reliable.
What will I have at the end?

You’ll know how to compare an agent’s claim with its recorded actions and spot what the record doesn’t tell you. Your agent can then help you view the example in Ironside, which stores recorded sessions. Once you know what you want to check, you can add automated checks.

When to use it

Use local-eval-stack when you want to inspect and evaluate coding-agent work on your own machine. It brings together Ironside for traces, Rubrist for evaluator development and review, and Casefile for inspecting the skills under test.

How it works

Figure 01

Use recorded agent work as evaluation material.

  1. Capture

    Import supported session records into Ironside.

  2. Select

    Inspect real runs and choose relevant examples.

  3. Evaluate

    Develop a focused check and review results in Rubrist.

ResultLocal traces and evaluator evidence
This shows the session-to-evaluation path. Casefile separately inspects the skills under test. Import scheduling is optional setup, not enabled by installation.

The prompts tell your coding agent where you’re starting and how much guidance you’d like. They ask it to reuse the choices you’ve already made.

The skill checks the local services and guides their connections. Bundled importers read Claude Code and Codex session logs; a pi extension supports live tracing. Reads of SKILL.md can become skill tags on traces, helping relate recorded behavior to actual skill use.

From those records, you can select examples, develop a focused judge, and review its results. The setup keeps ingest credentials scoped to writing traces and pauses at the points that require secrets or model-provider spending.

An example request

Ask your agent

“Set up a local evaluation workflow for my coding sessions. Start by checking which services are already running, then show how to preview one Codex session import.”

This illustrates a request you can adapt; it is not a transcript of a completed run.

A local path from recorded work to inspectable examples and evaluator results, with service checks and explicit setup steps.

Install and use

The skill’s identifier is local-eval-stack. It ships in the eval-stack plugin. In Claude Code, first add the Overclock marketplace, then install the package:

/plugin install eval-stack@overclock

For a standalone installation, keep the skill directory and its bundled resources together. The repository also includes per-skill Codex metadata. Read the full skill instructions for its exact workflow and supporting files.

What to keep in mind

The log importers run after records exist. Installing the plugin does not enable continuous capture; hook and scheduling recipes are optional. Redaction and truncation do not guarantee that confidential content has been removed.

This workflow targets one developer’s machine. Production hosting, multi-user operation, and CI gating against an existing Rubrist installation need separate setup.

  • Lessons learned — Turn a correction into a reusable lesson, and keep repeated lessons up to date.
  • Setup advisor — Choose the smallest useful set of Overclock plugins for your work.