Measure
Your GPU's spec sheet does not say what an answer costs.
This benchmark does. It runs a fixed set of tasks against the inference engine you already have, reads the GPU's energy counter around every item, and reports joules per correct answer, the number that a model cannot inflate by writing more. Run it on your own machine, or read what it found on ours.
Measure your own machine
One clone, one command, and roughly half an hour on an NVIDIA box. It uses the GPU's own power reading, so no wall meter is needed for this tier. Nothing leaves the box unless you opt in.
Runs on Linux and on Apple Silicon. The item counts are identical everywhere — that is what makes two results comparable — so a slower machine runs the same suite for longer rather than a smaller one: a Mac takes closer to 80 minutes, and the controller budgets for that on its own.
git clone https://github.com/boringbots/async-energy-controller.git
cd async-energy-controller
python3 -m venv .venv && source .venv/bin/activate
pip install -e .
ollama serve & # the suite attaches to an engine you
ollama pull qwen3.5:9b-q4_K_M # already run; it never pulls for you
async-energy-controller bench quick # no smart plug needed
async-energy-controller bench opt-in # optional: contribute the run (CC BY 4.0)
Measuring needs no account. Contributing does: opting in sends the
bundle under your API key, so the leaderboard can count distinct
machines and you can see your own rows.
Sign up,
mint a key on the
keys page,
and put it in the controller's .env before you opt in.
HM_ASYNC_API_URL=https://api.async.energy
HM_ASYNC_API_KEY=... # minted once at /app/keys/ -- a key, never your password The run writes a bundle to disk: energy, duration, throughput and accuracy for each task, with the honesty fields the methodology describes. Opting in shares the bundle and a hardware class (GPU model, VRAM, driver, CPU model, RAM), never prompts, commands, serials or hostnames. Contributed rows grow the corpus everyone below reads, and they give the scheduler a measured cost for your box instead of a nameplate guess. The full consent text is in the controller's README.
Or read what we measured
Three consumer GPUs, twenty-two models, four benchmarks, a thousand runs. Four views of the same data:
-
Leaderboard
Every measured configuration on one chart: accuracy across, joules per correct answer up. Filter by benchmark, GPU, model source and size; hover a dot for its numbers and grades.
-
Model picker
Pick the objective that matters to you: energy, speed, accuracy or balance. Every measured configuration is ranked on it, on the same chart, with the numbers behind each rank one hover away.
-
Energy data
Every variable the corpus has swept, as trajectories: thinking on and off, power caps, clocks, engines, quantization, output caps, and the same cell on different cards.
-
Methodology
What a number here means: three independent energy readings per run, the honesty fields that ride along, and the known limits. Read before quoting a figure.
For agents: the same data is JSON at the endpoints listed in llms.txt, including a recommendation endpoint that takes a workload shape and an MCP server. No key, no auth.
The research
The findings behind these pages are being written up as a paper, What a Correct Answer Costs. Its subtitle says what it covers: Joules per correct answer on consumer GPUs, at the wall, and which knobs move it without moving the answer. It is in preparation; the link will appear here when the preprint is posted. Until then, the introductory post carries the measured results with their sources.
Then move the work you can move
Once you know what a job costs, the remaining lever is when it runs. The same controller schedules deferrable work, such as agent runs, batch inference, indexing and fine-tunes, into the cheapest electricity hours of your day, and reports the measured energy of every run back to you.