inferctl
Inspect and diagnose your local LLM stack from the command line.
go install github.com/inferctl/inferctl/cmd/inferctl@latestinferctl explains your local LLM stack. It finds your inference backends, lists the models they serve, checks their health, and tells you which backend and model a task would use — and why. It reports only. It does not run inference.
The problem it solves
A local LLM setup often has more than one backend: Ollama on one port, llama.cpp on another, LM Studio in the background, MLX on a Mac. When a job fails, the cause is often simple. A server is down. A model is not loaded. A config points at the wrong port. It takes time to find out which.
An agent has the same problem, but it cannot look around. It needs a clear, machine-readable answer before it starts work. inferctl gives that answer.
What it does
- Diagnose health.
inferctl doctorchecks each configured backend and reports what is wrong. - List backends and models. It reads what each backend serves without loading or running a model.
- Explain routing.
inferctl route <task>names the backend and model a task would use, and gives the reason for the choice. - Check readiness.
inferctl preflight <task>tells an automation whether it can start now, as JSON or Markdown. - Track changes.
snapshot,diff, andstatuscapture the state of the stack and show what changed since the last check.
Supported backends
inferctl has read-only adapters for:
- Ollama
- llama.cpp
- LM Studio
- MLX
- any server with an OpenAI-compatible
/v1/modelsendpoint
Each adapter has a recorded, redacted test run against a real local server.
Built for agents and tool builders
Every command accepts --json and returns a stable envelope. Errors have fixed
codes. Most outputs include the next command to run. inferctl capabilities
publishes the command contract, so a caller can check compatibility first.
Tool builders can use inferctl as a control plane. Ask inferctl which backend and model to use, then call that backend directly:
inferctl route code --json
inferctl config show --json
The first command returns the selected backend and model. The second maps the backend name to its connection details, such as its base URL.
Install
inferctl is written in Go. Install it from its public module path on macOS, Linux, or Windows:
go install github.com/inferctl/inferctl/cmd/inferctl@latest
inferctl config explain
inferctl config explain describes the TOML config format. Define at least one
backend, then run inferctl doctor.
Works with evalctl
evalctl can call inferctl preflight before an eval run.
It stores the result with each case, so you know which model setup the run
assumed. See one command, three tools for a full
example.
Status
inferctl is pre-1.0 and licensed under Apache 2.0. It does not warm up models, manage locks, measure latency, or run inference. See the inferctl documentation for the verb and error reference.