← Products

inferctl

Inspect and diagnose your local LLM stack from the command line.

go install github.com/inferctl/inferctl/cmd/inferctl@latest

inferctl explains your local LLM stack. It finds your inference backends, lists the models they serve, checks their health, and tells you which backend and model a task would use — and why. It reports only. It does not run inference.

The problem it solves

A local LLM setup often has more than one backend: Ollama on one port, llama.cpp on another, LM Studio in the background, MLX on a Mac. When a job fails, the cause is often simple. A server is down. A model is not loaded. A config points at the wrong port. It takes time to find out which.

An agent has the same problem, but it cannot look around. It needs a clear, machine-readable answer before it starts work. inferctl gives that answer.

What it does

Supported backends

inferctl has read-only adapters for:

Each adapter has a recorded, redacted test run against a real local server.

Built for agents and tool builders

Every command accepts --json and returns a stable envelope. Errors have fixed codes. Most outputs include the next command to run. inferctl capabilities publishes the command contract, so a caller can check compatibility first.

Tool builders can use inferctl as a control plane. Ask inferctl which backend and model to use, then call that backend directly:

inferctl route code --json
inferctl config show --json

The first command returns the selected backend and model. The second maps the backend name to its connection details, such as its base URL.

Install

inferctl is written in Go. Install it from its public module path on macOS, Linux, or Windows:

go install github.com/inferctl/inferctl/cmd/inferctl@latest
inferctl config explain

inferctl config explain describes the TOML config format. Define at least one backend, then run inferctl doctor.

Works with evalctl

evalctl can call inferctl preflight before an eval run. It stores the result with each case, so you know which model setup the run assumed. See one command, three tools for a full example.

Status

inferctl is pre-1.0 and licensed under Apache 2.0. It does not warm up models, manage locks, measure latency, or run inference. See the inferctl documentation for the verb and error reference.