Where would you like to start?
Choose how much guidance you’d like. You can switch paths at any time.
Start with made-up data. Before using real examples, confirm you’re allowed to use them. Calls to an external model can send their contents off your machine.
Before using real data
Try your first agent check
An evaluation, or “eval,” checks an agent’s work against a clear expectation. Let’s start with one: did the agent show that it ran tests before saying the work was done?
- Agent says
- “Fixed the bug. Tests pass.”
- Recorded actions
- Edited a file. No test result is recorded.
The record doesn’t show that tests passed, so the claim needs a closer look. Missing results don’t prove that no tests ran.
Copy this prompt into your coding agent. It uses an Overclock skill, a set of instructions your agent follows during setup. Your agent will explain the example and unfamiliar terms, then help you choose the next step. You won’t need your own session data.
I’m new to evaluations and use Codex. Help me try one simple check: did an agent show evidence that it ran tests before saying the work was done?
Use the local-eval-stack skill from https://github.com/luka-zivkovic/overclock. If it is missing, explain how to install it for my agent, with its supporting files, and stop there.
Start with a clearly made-up session. Show me the agent’s claim and recorded actions, then walk me through checking them. Explain unfamiliar terms as they come up and guide me through one decision at a time. Don’t ask me to repeat answers I’ve already given. Missing test results leave the claim unsupported by this record; they don’t prove that no tests ran.
Then explain which local tools would help me inspect the example and check what’s installed. Agree on the next step with me before changing anything. Don’t require an AI model to judge the example just so I can understand it.
Use no real session logs, customer data, private code, or credentials in the example. Keep ongoing capture and automatic judging off. Keep secrets out of commands and output. Before credential writes or external model calls, explain what would be written or sent, to whom, and at what cost, and ask for any authorization still missing. This walkthrough is for learning, not proof that an agent is reliable.What will I have at the end?
You’ll know how to compare an agent’s claim with its recorded actions and spot what the record doesn’t tell you. Your agent can then help you view the example in Ironside, which stores recorded sessions. Once you know what you want to check, you can add automated checks.
Write a question above to generate your setup prompts.
Your setup prompts
Use Ironside for recorded work and Rubrist for a focused evaluator. Add Casefile only if we are evaluating an agent skill or plugin.
The prompts use Overclock’s local-eval-stack skill. Install it with its supporting files, then work through the steps in order.
Check prerequisites
A short plan showing what is installed, what is missing, and which services the goal needs.
Use Overclock’s local-eval-stack skill for one developer’s machine. My coding agent is Codex. Read the installed skill and the reference for the requested phase. If the skill is unavailable, show the installation steps for this agent and stop.
Goal: Check task completion.
Question: When the agent reports that a task succeeded, does the session contain evidence that it verified the result?
Use Ironside for recorded work and Rubrist for a focused evaluator. Add Casefile only if we are evaluating an agent skill or plugin.
Data: Use a small synthetic example with no real customer content, private code, or credentials. Do not search or read my existing session logs.
Keep automatic capture, scheduled imports, and auto-judging off. Keep secrets out of prompts, command arguments, logs, and Git. Ask before writing credentials or sending examples to an external model provider, and make the proposed data, model, and spending limit clear first.
For this step, only inspect the installed tools and local service configuration without exposing secrets. Do not read session logs, install packages, start services, or write files. Reuse my choices and ask only for information that is still missing. Show the setup plan and the checks you will use to verify it.Start the local services
Verified local service addresses, with any missing setup reported clearly.
Use Overclock’s local-eval-stack skill for one developer’s machine. My coding agent is Codex. Read the installed skill and the reference for the requested phase. If the skill is unavailable, show the installation steps for this agent and stop.
Goal: Check task completion.
Question: When the agent reports that a task succeeded, does the session contain evidence that it verified the result?
Use Ironside for recorded work and Rubrist for a focused evaluator. Add Casefile only if we are evaluating an agent skill or plugin.
Data: Use a small synthetic example with no real customer content, private code, or credentials. Do not search or read my existing session logs.
Keep automatic capture, scheduled imports, and auto-judging off. Keep secrets out of prompts, command arguments, logs, and Git. Ask before writing credentials or sending examples to an external model provider, and make the proposed data, model, and spending limit clear first.
Set up only the services needed for this goal. Check existing installations before changing them. Follow the current phase references, including Ironside owner setup before seed, and verify versions, ports, and authentication against the running services. Keep services local and ingest credentials write-scoped. Do not read or capture real sessions yet. Show the verified addresses and any unfinished steps.Preview the first examples
A small, reviewed sample and a clear account of what was imported or left out.
Use Overclock’s local-eval-stack skill for one developer’s machine. My coding agent is Codex. Read the installed skill and the reference for the requested phase. If the skill is unavailable, show the installation steps for this agent and stop.
Goal: Check task completion.
Question: When the agent reports that a task succeeded, does the session contain evidence that it verified the result?
Use Ironside for recorded work and Rubrist for a focused evaluator. Add Casefile only if we are evaluating an agent skill or plugin.
Data: Use a small synthetic example with no real customer content, private code, or credentials. Do not search or read my existing session logs.
Keep automatic capture, scheduled imports, and auto-judging off. Keep secrets out of prompts, command arguments, logs, and Git. Ask before writing credentials or sending examples to an external model provider, and make the proposed data, model, and spending limit clear first.
Prepare the selected data for inspection. For Claude Code or Codex session logs, use the bundled importer’s dry-run mode on the exact chosen file; pi tracing is a separate opt-in action, not permission to capture an ongoing session. If the format is unsupported, explain that rather than inventing a converter or substituting another source. Show what the preview will include and omit, and wait for confirmation before importing real data. Verify a sample trace in Ironside after an authorized import. Do not call an external judge in this step.Try the evaluation
A proposed criterion and a small set of results for human review, with uncertainty and missing evidence visible.
Use Overclock’s local-eval-stack skill for one developer’s machine. My coding agent is Codex. Read the installed skill and the reference for the requested phase. If the skill is unavailable, show the installation steps for this agent and stop.
Goal: Check task completion.
Question: When the agent reports that a task succeeded, does the session contain evidence that it verified the result?
Use Ironside for recorded work and Rubrist for a focused evaluator. Add Casefile only if we are evaluating an agent skill or plugin.
Data: Use a small synthetic example with no real customer content, private code, or credentials. Do not search or read my existing session logs.
Keep automatic capture, scheduled imports, and auto-judging off. Keep secrets out of prompts, command arguments, logs, and Git. Ask before writing credentials or sending examples to an external model provider, and make the proposed data, model, and spending limit clear first.
Turn my question into one observable criterion and let me review what counts as a pass, a failure, or insufficient evidence. Inspect the current Rubrist workflow before configuring it. Agree on the criterion and the small batch before running paid evaluation. Keep development feedback separate from independent human validation. Review disagreements and unscorable examples with me; missing evidence must not silently become a pass. A successful wiring test or agreement on a few examples does not establish evaluator accuracy or readiness for a release gate.Inspecting a few results helps you develop a check. Independent human review is still needed before relying on its judgments.
Continue your existing workflow
Choose the step you need. Each prompt tells your agent to work with your current setup and ask only for what’s missing.
Check the local setup
Inspect what is already running and identify what the next step needs.
Use Overclock’s local-eval-stack skill with Codex to continue my existing evaluation workflow. The skill is at https://github.com/luka-zivkovic/overclock. Read the reference for the requested step. If the skill is missing, show how to install it with its supporting files and stop. Keep explanations concise; reuse my existing choices, services, and configuration. Ask only for missing information needed for this step. Keep the work on one developer’s machine and do not expand it into a fresh full-stack setup.
Use only explicitly selected, authorized examples. Establish their destination, access, and retention before reading real data; redaction is not proof that data is safe. Keep ongoing capture, scheduling, and auto-judging off unless separately authorized. Keep secrets out of commands, output, and Git. Before credential writes or external model calls, make the proposed data, provider/model, and budget clear and obtain any authorization still missing.
Only check prerequisites and the existing local services relevant to my next step. Ask what I want to do next if I have not said. Verify versions, ports, and authentication without exposing credentials. Do not install, start, seed, or write anything yet. Report what works, what is unverified, and the smallest proposed change. Preserve existing data and configuration; when a fresh Ironside instance is needed, owner setup comes before seed.Import selected examples
Preview a chosen file or batch and verify what reaches the trace store.
Use Overclock’s local-eval-stack skill with Codex to continue my existing evaluation workflow. The skill is at https://github.com/luka-zivkovic/overclock. Read the reference for the requested step. If the skill is missing, show how to install it with its supporting files and stop. Keep explanations concise; reuse my existing choices, services, and configuration. Ask only for missing information needed for this step. Keep the work on one developer’s machine and do not expand it into a fresh full-stack setup.
Use only explicitly selected, authorized examples. Establish their destination, access, and retention before reading real data; redaction is not proof that data is safe. Keep ongoing capture, scheduling, and auto-judging off unless separately authorized. Keep secrets out of commands, output, and Git. Before credential writes or external model calls, make the proposed data, provider/model, and budget clear and obtain any authorization still missing.
Ask for the exact example files, their source format, and the intended existing Ironside instance if I have not supplied them. For Claude Code or Codex logs, use the bundled importer’s dry-run on only those files. Keep preview output within the authorized data scope. Explain what will be included or omitted and obtain any missing import authorization before ingest. Do not choose the latest session, sweep directories, or turn on pi live capture. If the format is unsupported, stop and explain. Verify a sample trace after import; require a skill tag only when the source records that skill use. Do not call a judge or install other services.Compare evaluation results
Compare a baseline and a change using the same examples and criterion.
Use Overclock’s local-eval-stack skill with Codex to continue my existing evaluation workflow. The skill is at https://github.com/luka-zivkovic/overclock. Read the reference for the requested step. If the skill is missing, show how to install it with its supporting files and stop. Keep explanations concise; reuse my existing choices, services, and configuration. Ask only for missing information needed for this step. Keep the work on one developer’s machine and do not expand it into a fresh full-stack setup.
Use only explicitly selected, authorized examples. Establish their destination, access, and retention before reading real data; redaction is not proof that data is safe. Keep ongoing capture, scheduling, and auto-judging off unless separately authorized. Keep secrets out of commands, output, and Git. Before credential writes or external model calls, make the proposed data, provider/model, and budget clear and obtain any authorization still missing.
Use my existing evaluation workflow. Establish only the missing baseline and candidate versions, selected examples, criterion, and human-review records. Prefer inspecting existing results before proposing new model calls. Compare the same tasks and criterion under comparable conditions, report improvements and regressions separately, and mark missing evidence as unresolved. If using Rubrist, inspect its current workflow before changing configuration. Keep development feedback separate from independent human validation. A small comparison or agreement on a few examples does not prove evaluator accuracy or release readiness.Need the skill itself? Go to installation instructions.
When to use it
Use local-eval-stack when you want to inspect and evaluate coding-agent work on your own machine. It brings together Ironside for traces, Rubrist for evaluator development and review, and Casefile for inspecting the skills under test.
How it works
Use recorded agent work as evaluation material.
- 1Capture
Import supported session records into Ironside.
- 2Select
Inspect real runs and choose relevant examples.
- 3Evaluate
Develop a focused check and review results in Rubrist.
The prompts tell your coding agent where you’re starting and how much guidance you’d like. They ask it to reuse the choices you’ve already made.
The skill checks the local services and guides their connections. Bundled importers read Claude Code and Codex session logs; a pi extension supports live tracing. Reads of SKILL.md can become skill tags on traces, helping relate recorded behavior to actual skill use.
From those records, you can select examples, develop a focused judge, and review its results. The setup keeps ingest credentials scoped to writing traces and pauses at the points that require secrets or model-provider spending.
An example request
“Set up a local evaluation workflow for my coding sessions. Start by checking which services are already running, then show how to preview one Codex session import.”
A local path from recorded work to inspectable examples and evaluator results, with service checks and explicit setup steps.
Install and use
The skill’s identifier is local-eval-stack. It ships in the eval-stack plugin. In Claude Code, first add the Overclock marketplace, then install the package:
/plugin install eval-stack@overclockFor a standalone installation, keep the skill directory and its bundled resources together. The repository also includes per-skill Codex metadata. Read the full skill instructions for its exact workflow and supporting files.
What to keep in mind
The log importers run after records exist. Installing the plugin does not enable continuous capture; hook and scheduling recipes are optional. Redaction and truncation do not guarantee that confidential content has been removed.
This workflow targets one developer’s machine. Production hosting, multi-user operation, and CI gating against an existing Rubrist installation need separate setup.
Related skills
- Lessons learned — Turn a correction into a reusable lesson, and keep repeated lessons up to date.
- Setup advisor — Choose the smallest useful set of Overclock plugins for your work.