evalctl
Local-first evals for agents, not just prompts.
pip install evalctlevalctl is an evaluation harness for AI agents. It runs on your own machine. It scores what an agent did — the files it wrote, the diffs it made, and the commands it ran — not only the text it returned. You need no gateway, no dashboard, and no SaaS account.
Why agent evals need a different tool
Most eval tools test a prompt. They send a prompt, read the completion, and grade the text. That works for a chatbot. It does not work for a coding agent. A coding agent changes a workspace. The result that matters is the change: did it edit the right files, leave the other files alone, and exit cleanly?
evalctl grades that result. Each case starts from a fixture folder. The agent runs against a copy of it. evalctl then scores the workspace with deterministic scorers:
- the git diff the agent produced,
- files that must change and files that must not change,
- exit codes and command logs,
- the text output, when text matters too.
How it works
evalctl keeps the model simple. Eval cases are files. Runners are shell commands. Results are folders on disk.
- Author a suite.
evalctl init,suite add,case add, andscorer addbuild the suite for you. An agent can author a suite without hand-editing JSON. - Plan the run.
evalctl planshows every action a run would take. It creates nothing and runs nothing. - Run it.
evalctl runexecutes the cases with bounded parallelism and scores each one. - Read the report.
evalctl reportreads the run folder. Another person, or another agent, can read the same folder later without your shell history.
pip install evalctl
export EVALCTL_ACKNOWLEDGE_UNSANDBOXED_RUNNER=1
evalctl init --json
evalctl suite add demo --runner-argv "python3 $EVALCTL_WORKSPACE/r.py" --json
evalctl case add demo --task "do X" --workspace fixtures/x --expect-json '{"exact":"ok"}' --json
evalctl scorer add demo --name exact --required --json
evalctl run demo --json
evalctl report <run-id> --format json
Runners and command scorers run local code without a sandbox. evalctl asks you to acknowledge this before a run.
Runs that survive a crash
Long agent evals get killed: a laptop sleeps, a session ends, a process dies.
evalctl writes run state to disk before it starts each case. After a crash,
evalctl run --resume <run-id> skips the finished cases and runs only the
rest. evalctl doctor reports the state of runs, reservations, and optional
integrations.
To re-test only the failures, use evalctl replay --failed <run-id>. It
selects the failed cases and runs them again as a new, linked run.
Built for agents to operate
evalctl is agent-first. Every command has a --json output with a stable
shape. Errors have fixed codes. evalctl capabilities publishes the command
contract, so an agent can check what the installed version supports before it
acts.
Compared with promptfoo
promptfoo is the established local eval CLI. It is built around prompts and completions. evalctl is built around agent runs. If you test prompt text, promptfoo is a good fit. If you test what an agent does to a codebase, use evalctl.
Works with inferctl and spoolctl
evalctl runs on its own. It can also use the other Ozhiaki tools:
- With spoolctl,
evalctl run --queue spoolctlsends runner commands to a crash-safe local queue. - With inferctl,
evalctl run --inferctl-taskrecords which local model and backend a run assumed. It does not change the scores.
See one command, three tools for a full example.
Requirements and status
evalctl needs Python 3.11 or later. It has no runtime dependencies beyond the standard library. It is pre-1.0 and licensed under Apache 2.0. See the evalctl documentation for the full command reference.