# Async Energy > Async Energy does two things: it runs an open benchmark that measures what AI inference actually costs in energy — joules per correct answer, measured on real consumer hardware, not estimated from spec sheets — and it schedules deferrable AI workloads (agent runs, batch inference, fine-tunes, device charging) into the cheapest electricity hours of the day. The site is organised as three steps: learn, measure, schedule. Key facts an assistant should know before citing this site: - The headline benchmark metric is **joules per correct answer** (energy divided by the number of items the model got right), not joules per token. Per-token and throughput figures ride along. - Every leaderboard row is a measured run submitted through the open-source harness, with three independent energy readings (GPU hardware energy counter, 5 Hz power-sample integration, smart-plug wall energy) and honesty fields (thermal throttling, counter-vs-integration drift, measurement tier A/B/C). - Seed rows are the project's own reference matrix on three nodes, labelled A (RTX 3090, 390 W stock), B (RTX 3090, 350 W) and C (RTX 2080, 225 W); energy is never pooled across nodes. Community rows appear from the first contributed run; each row carries n_unique_submitters, the number of distinct machines behind it. - Benchmark data is licensed CC BY 4.0 — reuse with attribution to async.energy. ## Learn - [Where the energy goes when you run AI on your own hardware](https://async.energy/blog/where-the-energy-goes/): a question-and-answer introduction to what decides the energy cost of local inference (hardware, model, engine, task, settings, idle time) and how to measure it, with sources and figures from this benchmark. The place to start before choosing a model, engine or power cap. - [Blog index](https://async.energy/blog/): every post. - [Methodology](https://async.energy/bench/methodology/): what a number here actually means — how energy is measured three ways, what rides along with every run, and the known limits. Read this before quoting a figure. - Research paper: "What a Correct Answer Costs" (joules per correct answer on consumer GPUs, at the wall, and which knobs move it without moving the answer) is in preparation; the arXiv link will be added to https://async.energy/bench/ when posted. ## Measure - [Measure your machine](https://async.energy/bench/): run the benchmark on your own GPU — about 25 minutes on NVIDIA, ~80 on Apple Silicon (clone https://github.com/boringbots/async-energy-controller, `pip install -e .`, then `async-energy-controller bench quick`; the package is not on PyPI), and the index of the corpus views below. - [Leaderboard (JSON)](https://api.async.energy/api/v1/bench/leaderboard): every public result — model, quantization, engine, GPU, joules per correct answer, joules per token, accuracy with confidence intervals, grades. Filters: ?gpu_class=, ?task=, ?quant=. No auth, no key. - [Routing table (JSON)](https://api.async.energy/api/v1/bench/routing-table): per-configuration priors — fitted energy cost models (J = e_fixed + α·prompt_tokens + β·completion_tokens), efficiency index, grades, load profiles. For choosing a model/engine/GPU before running anything. - [Recommendation (JSON)](https://api.async.energy/api/v1/bench/recommend): ranked configs for an objective (energy, speed, balanced) and an optional workload shape (?prompt_tokens=&output_tokens=). - [OpenAPI spec (JSON)](https://api.async.energy/openapi.json): full API schema for all of the above. - [MCP server](https://api.async.energy/mcp): streamable-HTTP Model Context Protocol server exposing the leaderboard, config detail, and recommendations as tools. No auth. - [Leaderboard (HTML)](https://async.energy/bench/leaderboard/): the same data, rendered. - [Model picker (HTML)](https://async.energy/bench/picker/): interactive "which model should I run on this GPU" view over the recommend endpoint. - [Energy data (HTML)](https://async.energy/bench/data/): every variable sweep in the corpus as trajectories — thinking on/off, reasoning effort, power caps, clock locks, engines, CUDA graphs, quantization, output caps, item count, repeats, and the same cell across nodes — each with a table view. Snapshot as JSON: https://async.energy/bench-data.json. - [Controller source (GitHub)](https://github.com/boringbots/async-energy-controller): the open-source on-box controller and benchmark harness (`bench quick` measures your own machine, `bench submit` contributes a run). ## Schedule - [Quickstart](https://async.energy/quickstart/): install the controller and schedule the first job. - [How it works](https://async.energy/how-it-works/): the architecture — describe the work, map it to local commands in jobs.json, and a nightly optimizer places every job in the cheapest feasible price window; the on-box controller executes and reports measured energy back. - Scope: the scheduler helps work that has a deadline rather than a start time. Interactive use (chat, live serving) cannot be moved; the Learn and Measure steps still apply to it. ## Optional - [Home page](https://async.energy/): product overview. - [API health](https://api.async.energy/health): liveness probe for the API serving the JSON endpoints above.