async.energy

Measure · energy data

Every knob the benchmark has turned, on one plane.

Capability across, energy up. Accuracy on the x-axis, joules per correct answer on a log y-axis. Every run is a point. Every time one variable changed while the rest held, an arrow joins the runs, pointing from before to after. Panels are split by benchmark because a correct answer costs 100× more on math500 than on mmlu. Snapshot of 1,207 runs, merged 2026-09-21, on the three nodes below. Licensed CC BY 4.0; the JSON behind this page is one file.

Reading the plane

Down and to the right is better: more answers right, fewer joules for each. A vertical arrow means the variable moved energy and not accuracy, which is the pattern that dominates this corpus. A horizontal arrow means it moved accuracy at constant cost, which almost never happens.

What an arrow holds fixed

Model, task, node, engine, quantization, output cap, thinking mode, CUDA-graph mode, item count and repeat index: all except the one variable the figure is about. Hover any point for the full configuration. Every figure has a table view, so nothing is reachable only by hovering.

The three nodes

Energy is never pooled across nodes: two identical cards differ by 22% in joules for identical work (section 6). Every energy figure on this page names its node.

NodeGPUStock power limitBehaviour under load
ARTX 3090 24 GB390 WPinned at its power cap for most of a long run; never thermally throttles.
BRTX 3090 24 GB350 WThermally throttles on long runs and settles near 300 W. Same die as A, different cooling and host.
CRTX 2080 8 GB225 WSmall-model tier only; no bf16; vLLM cannot serve a 7B at the sanctioned caps.

Which modelCapability & speed vs energy

One dot per configuration, (model, quantization, engine, node, settings), on a single benchmark, filtered by GPU and board. The first four modes plot a metric against energy per correct answer (down is cheaper); TTFT vs accuracy and TTFT vs tok/s plot responsiveness against capability or throughput, with energy moved into the dot colour. Axis labels name which direction is better; the Pareto front (amber dots joined by a stepped hairline, or ink-ringed dots in the coloured modes) is nothing-measured-is-better-on-both-axes. A hollow dot is a quantized checkpoint.

Loading the corpus…

Why one benchmark at a time. J/correct-answer is only comparable within a task: the same model needs ~54 J for an mmlu-redux answer and ~6,600 J for a math-500 one, because the questions are not the same work. Pooling tasks would draw a chart whose real axis is "which benchmark", and a Pareto front across it would be an artifact.

Chance floor. On a four-choice task a model that guesses still books ~25 correct answers per 100 for almost no energy, so its J/correct is tiny and it sorts to the top. Since 2026-09-09 a cell may only rank when its accuracy interval's lower bound clears the floor (0.25 on mmlu and gpqa). Unticked, sub-floor cells are drawn with a red ring and kept off the front.

Tier A only. Unticked, the panel shows every configuration in the corpus (410 cells), pooling eager-mode, uncapped and old-stop rows and warm-cache repeats; ticked, only settings-matched cells (170), one run each.

What is measuredThe four benchmarks

Every task is scored through the same generative path: the model writes its answer as text, and the answer is extracted and checked. The standard alternative for multiple choice, comparing token likelihoods without generating, would leave joules-per-token undefined, so it is not used. Two tasks are prefill-shaped (long prompt, tiny answer) and two are decode-shaped (short prompt, long answer); a decode token costs 40–60× a prefill token, which is why the shape matters to energy.

gsm8k_platinum

madrylab/gsm8k-platinum · decode · 5-shot · 100 items · cap 8,192

Tests multi-step arithmetic on grade-school word problems, using the re-labelled "platinum" test set so known label errors don't count against a model.

How: five worked examples, then the question; the model reasons in prose and ends with #### <number>. Exact numeric match on the final number.

Read with care: saturated. At 100 items the strong models sit at 0.90–0.99 and 16 of 29 cells tie the leader. It separates models on energy, not accuracy.

mmlu_redux

edinburgh-dawg/mmlu-redux-2.0 · prefill · 5-shot · 100 items (500 in the engine wave) · cap 8,192

Tests breadth of knowledge: four-way multiple choice across 57 subjects from law and medicine to college mathematics, on the human re-annotated subset with the broken questions removed.

How: five examples, then the question; the model emits a letter A–D. Letter extraction.

Read with care: the median answer is two tokens, so per-request overhead dominates the energy, which is exactly why Ollama looks so expensive here. Accuracy is not leaderboard-comparable (generative protocol, our prompt).

gpqa_diamond

Idavidrein/gpqa (gated) · prefill · zero-shot · 50 items · cap 16,384

Tests graduate-level reasoning in biology, physics and chemistry. Written to be "Google-proof": PhDs outside the question's subfield reach about 34% with open web access. The one task current models do not saturate.

How: zero-shot, matching the paper; four shuffled options; the model emits a letter. Reasoning models think at length first.

Read with care: at 50 items the interval is ±0.13, so the ranking is informative but unresolved. A truncated answer scores at chance (0.25), not zero. The only task where turning thinking on raised accuracy.

math500

HuggingFaceH4/MATH-500 · decode · 5-shot · 50 items · cap 16,384

Tests competition mathematics: algebra, number theory, geometry, counting, with long derivations.

How: five algebra examples from the original MATH train set, then the problem; the model derives and ends with \boxed{answer}. String equality after light LaTeX normalisation, not a computer-algebra check, so \frac{1}{2} and 0.5 do not match.

Read with care: accuracy is a lower bound, so energy per correct is an upper bound. Inherently fat-tailed: the five longest items are 32–50% of the tokens on every engine. Qwen3.5 models often never emit \boxed at all (a format failure, not a maths failure).

What is testedThe models, by family

Size, architecture, and the three features that move energy the most: whether the model has a thinking mode and what its template does when nothing is set, whether it is a mixture of experts (only a fraction of the weights run per token), and whether it carries a vision tower (extra weights that these text-only benchmarks never exercise). "Built for" is the publisher's own positioning; "what we found" is from this corpus.

ModelSizeArchitectureFeaturesBuilt for (publisher)What we found
Alibaba (Qwen)
Qwen3-0.6B0.6Bdense transformerthinking · default ONEdge and on-device assistant; the same dual-mode recipe as the larger Qwen3s. Most-downloaded text model on the hub.Rambles: 20× LFM2.5-1.2B's energy for lower accuracy. The clearest case of size not predicting cost.
Qwen3-8B + AWQ · FP8 · Q4_K_M · Q8_08Bdense transformerthinking · default ONGeneral-purpose assistant with switchable step-by-step reasoning, 40k context. Third most-downloaded on the hub.In the tied top group with thinking on, at 8–10× gpt-oss's energy. With thinking off it drops to 0.75 pooled and the AWQ build does that for 161 kJ; the FP8, Q8_0 and Q4_K_M repacks land between.
Qwen3.5-0.8B / 2B / 4B / 9B0.8–9Bhybrid: Gated-DeltaNet linear attention + attentionthinking · 0.8B/2B OFF, 4B/9B ONvisionNatively multimodal, 262k-token context, efficient long-context assistants; the switchable thinking mode.The reference model (9B). Same family, opposite thinking defaults by size. 9B fp16 gets only a 4k window on 24 GB. math500 is a format failure (never emits \boxed).
Qwen3.5-9B-AWQ · Qwen3.8-27B-AWQ9B · 27Bsame hybrid, 4-bitthinking · default ONvision (9B)Quantized repacks; the 27B is the family's larger reasoning tier.AWQ gets the 9B a 32k window where fp16 gets 4k. The 27B cannot run at tier A on 24 GB (OOM in CUDA-graph capture).
Qwen2.5-7B-Instruct + AWQ · GPTQ-Int4 · GPTQ-Int8 · Q4_K_M7Bdense transformer, GQAno thinkingThe 2024 general assistant, pre-reasoning era; long-text generation, structured output, multilingual.The continuity anchor and the workhorse: the whole quantization ladder, every power and clock ladder, both engine waves.
Qwen3-Coder-30B-A3B-AWQ + Ollama tag30B · 3B activeMoE, 128 expertsMoEno thinkingAgentic coding: repository-scale code generation and tool use, 262k context. Not a general assistant.Scored on four non-coding tasks on purpose. No measurable accuracy deficit against granite-4.2-8b at 46× less energy; 52× less than Qwen3.5-9B. On SGLang, 251 J per correct answer for the whole suite. Ollama's Q4_K_M tag scores 0.81 pooled for 180 kJ, the cheapest configuration in the top statistical group.
OpenAI
gpt-oss-20b + Ollama tag21B · 3.6B activeMoE, MXFP4-native weightsMoEreasoning effort low / medium / highOpen-weight reasoning and agentic model sized to run on 16 GB; the "harmony" format with a graded reasoning dial rather than an on/off switch.The accuracy anchor of the frontier: 0.857 pooled for 580 kJ, with every dense thinking model near it costing 5.8–17× more. Since the Ollama parity wave, two MoE Q4_K_M tags (Qwen3-Coder, gemma-4-26B) share its statistical group at 1.8–3.2× less energy.
IBM
granite-4.2-8b8Bdense transformer, GQAthinking · default ONEnterprise assistant: tool calling, RAG, long documents; Apache 2.0 with a documented training corpus.In the tied top group on accuracy, at 13–17× gpt-oss's energy. Its thinking switch is load-bearing (trained in with RL).
Ornith AI
Ornith-1.5-9B9Bdense transformerthinking · default ONReasoning and coding specialist, continued-pretrained on Qwen3.5 and Gemma 4 with heavy RL self-improvement; MIT.Accuracy point-leader (0.873), tied with three others; 5.8× gpt-oss's energy for +0.02.
HuggingFace
SmolLM3-3B3Bdense transformerthinking · default ONSmall multilingual long-context model with a dual reasoning mode, fully open training recipe.The model whose unpinned thinking scored 0% on every multiple-choice task and forced the read-the-template rule. No tier-A rows yet.
Liquid AI
LFM2.5-1.2B-Instruct1.2Bhybrid: gated short convolutions + GQAno thinkingOn-device: fast CPU and GPU inference at low memory, built for phones and laptops rather than servers.19 J per correct mmlu answer at 0.37 accuracy, the fourth-cheapest tier-A cell in the corpus. Three of four tasks at tier A, so no roster row. Only measured on node C.
LFM2.5-2.6B2.6Bsame hybridreasoning always on (forced)The family's reasoning tier.Force-opens a thinking block on every answer with no way to turn it off, so it is not comparable to its 1.2B sibling on size alone.
PrismML
Ternary-Bonsai-1.7B / 4B / 8B1.7–8Bdense, Qwen3 bases, ternary weights (2.13 bits)thinking · pinned OFFPost-training quantization of Qwen3 checkpoints to ternary weights; accuracy per gigabyte is the publisher's headline and phones are the target.The 8B at 2.13 bits matches its own F16 render to the digit on three of four tasks for 3.4–8× less energy. The 4B on node C is the cheapest tier-A cell in the corpus: 11 J per correct mmlu answer.
Ternary-Bonsai-2-27B · Bonsai-27B27Bsame, ternary (PQ2_0 / PTQ1_0) and one-bit (Q1_0)thinking · pinned OFFThe family's largest tier: a 27B in 7–8 GB.At 1.58 bits a 27B fits an 8 GB RTX 2080 and scores 0.80 pooled. Below 2 bits the saving reverses: PTQ1_0 matches PQ2_0 on accuracy with 13% less VRAM but 1.2–1.6× more energy on Ampere. The one-bit Bonsai-27B holds 0.77.
Google
gemma-4-26B-A4B-AWQ · gemma-4-31B-AWQ · gemma-4-e2b-it26B (4B active) · 31B · 2BMoE (26B) · dense (31B, e2b)MoE (26B)thinking · default OFFvisionOpen multimodal generalists in three sizes; the 26B is the efficient MoE tier, the 31B the quality tier.The 31B scored the highest of anything measured (1.00 gsm8k, 0.96 mmlu) but only in tier B, and cannot run at tier A on 24 GB. The 26B AWQ build does not fit gpqa/math500 at the sanctioned caps on vLLM; Ollama's Q4_K_M tag of it does, at 0.83 pooled for 330 kJ, in the top statistical group.
Mistral AI
Mistral-7B-Instruct-v0.2 + AWQ · Q4_K_M7Bdense, GQA, sliding-window attentionno thinkingThe 2023 open 7B that set the small-model bar; general chat.Pre-reasoning-era control. 0.41 pooled accuracy at 231 kJ: cheap, and behind everything modern.
Microsoft
Phi-3-mini-4k-instruct3.8Bdense, no GQAno thinkingSmall model trained on curated "textbook-quality" data; reasoning-heavy for its size, 4k context.Its 4k window cannot hold the sanctioned caps: does-not-fit on two tasks. No tier-A rows.
TII (Technology Innovation Institute)
Falcon-H1R-7B7BMamba-2 / attention hybridreasoning always onState-space hybrid with a reasoning post-train; long-context efficiency.In the roster, never run. The only Mamba hybrid.
Meta
Llama-3.1-8B · Llama-3.2-1B8B · 1.2Bdense transformerno thinkingThe deployment-reality anchors: the most-run open models in the world.The FP16 build stays gated, but Ollama's Q4_K_M tag of the 8B ran on node A on 2026-09-21: 0.89 on gsm8k, 0.24 on math500, tier A on all four tasks.

Quantization variants (AWQ, GPTQ-Int4/Int8, Q4_K_M, MXFP4, and PrismML's PQ2_0 / PTQ1_0 / Q1_0 below two bits) are the same weights re-packed at fewer bits; each is its own row in the index because two packings of "the same" model differ measurably. Ollama tags are Ollama's own GGUF conversions of the named model.

1 · RosterEvery model at its reference configuration, per node

Tier A, all four tasks pooled at the reference item counts (100 / 100 / 50 / 50), Wilson 95% intervals, one configuration per (model, quantization, engine, node, thinking mode), so a model that ran thinking on and off appears twice. The shaded band is every configuration whose interval overlaps the leader's: one statistical group, spread more than an order of magnitude on energy. This is the only figure that pools tasks.

Suite energy against pooled accuracy

configurations with all four tasks at tier A · suite kJ on a log scale · whiskers are the 95% CI on accuracy · red ring where the gpqa cell cannot beat guessing

2 · Model sizeWalking up a family's size ladder

Arrows step from the smallest to the largest build within one family on one node, tier A only, and within one thinking mode, so a Qwen3.5 ladder with thinking off is a different line from the same ladder with thinking on. Where an arrow goes up without going right, size bought tokens, not answers.

Size ladders, per family and node

arrow = one size step · hollow point = smallest build

3 · ThinkingThinking off → on, ten models, three nodes

Every pair in the corpus where only the thinking setting changed: the same model, task, node, cap, graph mode and items. Arrow colour is the verdict on the Wilson intervals: green where thinking raised accuracy, red where it lowered it, grey where the intervals overlap. Pooled over the tier-A pairs, thinking costs 9× and helps in 6 of 48.

Thinking off → on, per pair

colour is the verdict · hollow = thinking off · label carries the cost ratio · only tier-A pairs are drawn

3b · Reasoning effortgpt-oss at low → medium → high

The only graded reasoning dial in the roster, run on all four tasks at tier A with gpqa at the full 198 items. Every arrow is short: accuracy intervals overlap on every task, token counts sit within 4% of each other, and energy moves by less than the run-to-run floor. The dial never reached the model. The low and high reasoning traces are byte-identical on 142 of 198 gpqa items, while each differs from the medium row on 185: vLLM 0.25.1 takes harmony reasoning effort from the top-level reasoning_effort request field, and the harness sent it inside chat_template_kwargs. All three rungs ran at medium, so what the arrows show is boot-to-boot drift on an MoE, not a dose-response. Kept here because it is the kind of mistake anyone measuring this dial can make.

Effort ladder, per task

hollow = low · arrows step to medium then high · gpqa at n=198 · every rung actually ran at medium

3c · Item countgpqa at 50 → 198 items

The accuracy interval halves in width from 50 to 198 items, and the point estimates regress toward each other: Qwen3.5-4B fell from 0.74 to 0.68, gpt-oss rose from 0.59 to 0.61–0.65. They still overlap. What 198 items did resolve is the bottom of the frontier: the no-thinking coder's 0.42 is now clearly below both, at 886 J per correct answer against 14–64 kJ.

gpqa_diamond, 50 → 198 items, same configuration

hollow = 50 items · arrow ends at 198 · hover for the intervals

4 · Power limitEvery power ladder, every node

Nine ladders across three cards and three models. The knee sits near 225 W on both 3090s. Accuracy is flat on every ladder, so the plane view would be nine vertical lines; these keep the cap on the x-axis so the knee is visible.

Energy per correct answer against power cap

per node · colour is task · label is the model · repeats averaged

Throughput against power cap

same ladders

5 · Clock lockEvery clock ladder

Six ladders: three tasks on node A, two on node B, one on node C. Node A's gsm8k minimum is at 1350 MHz. The second panel is the finding that a "smoothness" grade, as first defined, rewards a slower card: power CV rises with clock even with no cap in play.

Energy per correct answer against SM clock

per node · colour is task · label is the model

Power CV against SM clock

same ladders

6 · NodeThe same cell on two or three cards

139 cells ran the same configuration on more than one node. Accuracy lands in the same column; energy does not. Node A at 390 W sits above node B at 350 W on almost every cell, and node C's 2080 is cheaper still on the small models it can hold.

Same configuration, different node

colour is node · a line joins one cell across nodes

7 · CUDA graphsEager → graphs on

26 pairs where only the graph flag changed. On decode-heavy tasks the arrow points down (graphs save energy); on the two-token multiple-choice tasks it can point up. Accuracy never moves.

Eager → CUDA graphs, per pair

hollow = eager · arrow ends at graphs on

8 · EngineThe same weights on different engines

34 cells from both engine waves. Within a cell the weights are byte-identical (vLLM ↔ SGLang share AWQ; llama.cpp ↔ Ollama share Q4_K_M; gpt-oss is MXFP4 on both). Accuracy ties everywhere; the vertical spread is the engine.

Engines, per cell

colour is engine · a line joins one cell across engines

9 · RepeatsThe same cell run again

22 clusters of three or more runs. Dense models land on the same point every boot; mixtures-of-experts scatter 2–4% in energy; and the earliest repeat studies, run before prefix caching was disabled, walk downward as the cache warms.

Repeat clusters

colour is architecture · points in one cluster are joined in run order

10 · Output capRaising the cap, up to uncapped

123 sequences: the deliberate cap ladders plus every capped/uncapped pair in the reference wave. On gsm8k the arrows stop moving after 800 tokens. On math500 they keep climbing. Uncapped points at the top are the runaway items measuring the context window.

Cap sequences, smallest cap → uncapped

hollow = smallest cap · one line per model and configuration

11 · Quantizationfp16 → Int8 → Int4 → AWQ

Every precision sequence on a shared configuration, one panel per model and node, the benchmarks as the coloured series. The full four-step ladder exists on Qwen2.5-7B on node B; two-step fp16 → AWQ pairs exist for Qwen2.5-7B and Qwen3.5-9B on node A. Arrows go down, rarely left. Each line states its tier; only sequences with the same tier are comparable to each other.

Precision sequences, per model and node

colour is benchmark · hollow = fp16 · label is the tier

CoverageWhere each model has been swept

Distinct levels per axis for every build with runs. ⊘ marks a does-not-fit verdict somewhere in the fleet.

Model × axis coverage

39 builds