One command, three tools: inspect, evaluate, queue
Run a complete agent eval on your laptop with all three Ozhiaki tools at once, in one command. inferctl inspects your local model stack and advises the route, evalctl runs and scores the suite, and spoolctl runs the cases as a queue. You need no API key, no Docker, and no login.
Each tool does one job:
- inferctl — inspect and advise. It reads your local LLM stack and reports what a run would use, before anything runs. It does not run inference; that is not its job. Its output is a record of what the run assumed.
- evalctl — evaluate. It runs the cases, scores them by what they did to the workspace, and owns the run folder you can replay later.
- spoolctl — queue. It runs the runner commands as a local queue, so the work is bounded and survives a killed worker.
Install the tools
evalctl and spoolctl are on PyPI; inferctl installs from its public module path.
pip install evalctl "spoolctl>=0.4.11"
go install github.com/inferctl/inferctl/cmd/inferctl@latest
The install fetches the packages. After that, the run needs no network — it inspects and executes on your machine.
Scaffold a suite
evalctl init writes a ready-to-run suite. It ships one passing case and one
failing case, so you see both verdicts.
evalctl init
Run the suite
Run the suite, delegate execution to spoolctl, and capture inferctl’s preflight for the run:
export EVALCTL_ACKNOWLEDGE_UNSANDBOXED_RUNNER=1
evalctl run code-review --queue spoolctl --slots 4 --inferctl-task code --json
The acknowledgment is required because runner commands execute local code. evalctl scores what the code did, so it runs that code; it is not a sandbox.
What each tool reported
inferctl advised the route. Before any case ran, inferctl inspected the stack
and reported what a code task would use. On the machine that wrote this post,
the advice was an on-device MLX model:
{ "selected_model": "mlx-community/Qwen2.5-0.5B-Instruct-4bit",
"selected_backend": "mlx",
"runnable": true,
"reason": "route satisfies preflight policy" }
Your route will differ, because it reflects your own stack. If nothing suitable is installed, inferctl says so; that is still the answer you need.
evalctl wrote that advice into the run folder, one pair of files per case:
evals/runs/<run-id>/cases/cr-pass/inferctl-preflight.json
evals/runs/<run-id>/cases/cr-pass/inferctl-provenance.json
spoolctl ran the queue. Under --queue spoolctl, the runner commands ran
through spoolctl’s local queue instead of in-process. evalctl kept everything
else.
evalctl scored the run. The suite finished with one pass and one fail, as scaffolded:
{ "case_count": 2, "status_counts": { "pass": 1, "fail": 1, "error": 0 } }
The run’s meta recorded that inferctl was present and captured, not absent or
degraded:
"inferctl": { "requested_mode": "preflight", "actual_mode": "preflight",
"available": true, "degraded": false }
What each tool owns
The division of labor is deliberate. Even under --queue spoolctl, evalctl still
prepares each case’s workspace, normalizes output, captures diffs, scores, and
writes the run folder. spoolctl runs only the runner command. inferctl only
inspects and advises; its output is provenance, not a change to the run.
That separation is why the run is reproducible: the scoring and the record stay evalctl-owned, whatever ran the runner and whatever the model stack looked like.
Resume after a crash
If a run is interrupted — a worker killed, the machine shut down — resume it with
evalctl run --resume <run-id>. spoolctl keeps the queue state on disk, in a
per-run .spoolctl.db, so a resumed run can continue the remaining cases instead
of starting over. The run folder and its scores stay evalctl-owned throughout.
What you get
One command, three tools, and a run folder you can hand to someone else to replay — with the model-stack assumptions recorded next to the scores. It runs on your laptop, the grade is deterministic, and each tool does one job: inferctl tells you what the run would use, evalctl grades what the run did, and spoolctl runs the work.