← Products

evalctl

Local-first evals for agents, not just prompts.

pip install evalctl

evalctl is an evaluation harness for AI agents. It runs on your own machine. It scores what an agent did — the files it wrote, the diffs it made, and the commands it ran — not only the text it returned. You need no gateway, no dashboard, and no SaaS account.

Why agent evals need a different tool

Most eval tools test a prompt. They send a prompt, read the completion, and grade the text. That works for a chatbot. It does not work for a coding agent. A coding agent changes a workspace. The result that matters is the change: did it edit the right files, leave the other files alone, and exit cleanly?

evalctl grades that result. Each case starts from a fixture folder. The agent runs against a copy of it. evalctl then scores the workspace with deterministic scorers:

How it works

evalctl keeps the model simple. Eval cases are files. Runners are shell commands. Results are folders on disk.

  1. Author a suite. evalctl init, suite add, case add, and scorer add build the suite for you. An agent can author a suite without hand-editing JSON.
  2. Plan the run. evalctl plan shows every action a run would take. It creates nothing and runs nothing.
  3. Run it. evalctl run executes the cases with bounded parallelism and scores each one.
  4. Read the report. evalctl report reads the run folder. Another person, or another agent, can read the same folder later without your shell history.
pip install evalctl
export EVALCTL_ACKNOWLEDGE_UNSANDBOXED_RUNNER=1
evalctl init --json
evalctl suite add demo --runner-argv "python3 $EVALCTL_WORKSPACE/r.py" --json
evalctl case add demo --task "do X" --workspace fixtures/x --expect-json '{"exact":"ok"}' --json
evalctl scorer add demo --name exact --required --json
evalctl run demo --json
evalctl report <run-id> --format json

Runners and command scorers run local code without a sandbox. evalctl asks you to acknowledge this before a run.

Runs that survive a crash

Long agent evals get killed: a laptop sleeps, a session ends, a process dies. evalctl writes run state to disk before it starts each case. After a crash, evalctl run --resume <run-id> skips the finished cases and runs only the rest. evalctl doctor reports the state of runs, reservations, and optional integrations.

To re-test only the failures, use evalctl replay --failed <run-id>. It selects the failed cases and runs them again as a new, linked run.

Built for agents to operate

evalctl is agent-first. Every command has a --json output with a stable shape. Errors have fixed codes. evalctl capabilities publishes the command contract, so an agent can check what the installed version supports before it acts.

Compared with promptfoo

promptfoo is the established local eval CLI. It is built around prompts and completions. evalctl is built around agent runs. If you test prompt text, promptfoo is a good fit. If you test what an agent does to a codebase, use evalctl.

Works with inferctl and spoolctl

evalctl runs on its own. It can also use the other Ozhiaki tools:

See one command, three tools for a full example.

Requirements and status

evalctl needs Python 3.11 or later. It has no runtime dependencies beyond the standard library. It is pre-1.0 and licensed under Apache 2.0. See the evalctl documentation for the full command reference.