← Blog

One command, three tools: inspect, evaluate, queue

Run a complete agent eval on your laptop with all three Ozhiaki tools at once, in one command. inferctl inspects your local model stack and advises the route, evalctl runs and scores the suite, and spoolctl runs the cases as a queue. You need no API key, no Docker, and no login.

Each tool does one job:

Install the tools

evalctl and spoolctl are on PyPI; inferctl installs from its public module path.

pip install evalctl "spoolctl>=0.4.11"
go install github.com/inferctl/inferctl/cmd/inferctl@latest

The install fetches the packages. After that, the run needs no network — it inspects and executes on your machine.

Scaffold a suite

evalctl init writes a ready-to-run suite. It ships one passing case and one failing case, so you see both verdicts.

evalctl init

Run the suite

Run the suite, delegate execution to spoolctl, and capture inferctl’s preflight for the run:

export EVALCTL_ACKNOWLEDGE_UNSANDBOXED_RUNNER=1

evalctl run code-review --queue spoolctl --slots 4 --inferctl-task code --json

The acknowledgment is required because runner commands execute local code. evalctl scores what the code did, so it runs that code; it is not a sandbox.

What each tool reported

inferctl advised the route. Before any case ran, inferctl inspected the stack and reported what a code task would use. On the machine that wrote this post, the advice was an on-device MLX model:

{ "selected_model": "mlx-community/Qwen2.5-0.5B-Instruct-4bit",
  "selected_backend": "mlx",
  "runnable": true,
  "reason": "route satisfies preflight policy" }

Your route will differ, because it reflects your own stack. If nothing suitable is installed, inferctl says so; that is still the answer you need.

evalctl wrote that advice into the run folder, one pair of files per case:

evals/runs/<run-id>/cases/cr-pass/inferctl-preflight.json
evals/runs/<run-id>/cases/cr-pass/inferctl-provenance.json

spoolctl ran the queue. Under --queue spoolctl, the runner commands ran through spoolctl’s local queue instead of in-process. evalctl kept everything else.

evalctl scored the run. The suite finished with one pass and one fail, as scaffolded:

{ "case_count": 2, "status_counts": { "pass": 1, "fail": 1, "error": 0 } }

The run’s meta recorded that inferctl was present and captured, not absent or degraded:

"inferctl": { "requested_mode": "preflight", "actual_mode": "preflight",
              "available": true, "degraded": false }

What each tool owns

The division of labor is deliberate. Even under --queue spoolctl, evalctl still prepares each case’s workspace, normalizes output, captures diffs, scores, and writes the run folder. spoolctl runs only the runner command. inferctl only inspects and advises; its output is provenance, not a change to the run.

That separation is why the run is reproducible: the scoring and the record stay evalctl-owned, whatever ran the runner and whatever the model stack looked like.

Resume after a crash

If a run is interrupted — a worker killed, the machine shut down — resume it with evalctl run --resume <run-id>. spoolctl keeps the queue state on disk, in a per-run .spoolctl.db, so a resumed run can continue the remaining cases instead of starting over. The run folder and its scores stay evalctl-owned throughout.

What you get

One command, three tools, and a run folder you can hand to someone else to replay — with the model-stack assumptions recorded next to the scores. It runs on your laptop, the grade is deterministic, and each tool does one job: inferctl tells you what the run would use, evalctl grades what the run did, and spoolctl runs the work.