Measure · energy data
Every knob the benchmark has turned, on one plane.
Capability across, energy up. Accuracy on the x-axis, joules per correct answer on a log y-axis. Every run is a point. Every time one variable changed while the rest held, an arrow joins the runs, pointing from before to after. Panels are split by benchmark because a correct answer costs 100× more on math500 than on mmlu. Snapshot of 1,207 runs, merged 2026-09-21, on the three nodes below. Licensed CC BY 4.0; the JSON behind this page is one file.
Reading the plane
Down and to the right is better: more answers right, fewer joules for each. A vertical arrow means the variable moved energy and not accuracy, which is the pattern that dominates this corpus. A horizontal arrow means it moved accuracy at constant cost, which almost never happens.
What an arrow holds fixed
Model, task, node, engine, quantization, output cap, thinking mode, CUDA-graph mode, item count and repeat index: all except the one variable the figure is about. Hover any point for the full configuration. Every figure has a table view, so nothing is reachable only by hovering.
The three nodes
Energy is never pooled across nodes: two identical cards differ by 22% in joules for identical work (section 6). Every energy figure on this page names its node.
| Node | GPU | Stock power limit | Behaviour under load |
|---|---|---|---|
| A | RTX 3090 24 GB | 390 W | Pinned at its power cap for most of a long run; never thermally throttles. |
| B | RTX 3090 24 GB | 350 W | Thermally throttles on long runs and settles near 300 W. Same die as A, different cooling and host. |
| C | RTX 2080 8 GB | 225 W | Small-model tier only; no bf16; vLLM cannot serve a 7B at the sanctioned caps. |
Which modelCapability & speed vs energy
One dot per configuration, (model, quantization, engine, node, settings), on a single benchmark, filtered by GPU and board. The first four modes plot a metric against energy per correct answer (down is cheaper); TTFT vs accuracy and TTFT vs tok/s plot responsiveness against capability or throughput, with energy moved into the dot colour. Axis labels name which direction is better; the Pareto front (amber dots joined by a stepped hairline, or ink-ringed dots in the coloured modes) is nothing-measured-is-better-on-both-axes. A hollow dot is a quantized checkpoint.
Loading the corpus…
Why one benchmark at a time. J/correct-answer is only comparable within a task: the same model needs ~54 J for an mmlu-redux answer and ~6,600 J for a math-500 one, because the questions are not the same work. Pooling tasks would draw a chart whose real axis is "which benchmark", and a Pareto front across it would be an artifact.
Chance floor. On a four-choice task a model that guesses still books ~25 correct answers per 100 for almost no energy, so its J/correct is tiny and it sorts to the top. Since 2026-09-09 a cell may only rank when its accuracy interval's lower bound clears the floor (0.25 on mmlu and gpqa). Unticked, sub-floor cells are drawn with a red ring and kept off the front.
Tier A only. Unticked, the panel shows every configuration in the corpus (410 cells), pooling eager-mode, uncapped and old-stop rows and warm-cache repeats; ticked, only settings-matched cells (170), one run each.
What is measuredThe four benchmarks
Every task is scored through the same generative path: the model writes its answer as text, and the answer is extracted and checked. The standard alternative for multiple choice, comparing token likelihoods without generating, would leave joules-per-token undefined, so it is not used. Two tasks are prefill-shaped (long prompt, tiny answer) and two are decode-shaped (short prompt, long answer); a decode token costs 40–60× a prefill token, which is why the shape matters to energy.
gsm8k_platinum
Tests multi-step arithmetic on grade-school word problems, using the re-labelled "platinum" test set so known label errors don't count against a model.
How: five worked examples, then the question; the model reasons in prose and ends with #### <number>. Exact numeric match on the final number.
Read with care: saturated. At 100 items the strong models sit at 0.90–0.99 and 16 of 29 cells tie the leader. It separates models on energy, not accuracy.
mmlu_redux
Tests breadth of knowledge: four-way multiple choice across 57 subjects from law and medicine to college mathematics, on the human re-annotated subset with the broken questions removed.
How: five examples, then the question; the model emits a letter A–D. Letter extraction.
Read with care: the median answer is two tokens, so per-request overhead dominates the energy, which is exactly why Ollama looks so expensive here. Accuracy is not leaderboard-comparable (generative protocol, our prompt).
gpqa_diamond
Tests graduate-level reasoning in biology, physics and chemistry. Written to be "Google-proof": PhDs outside the question's subfield reach about 34% with open web access. The one task current models do not saturate.
How: zero-shot, matching the paper; four shuffled options; the model emits a letter. Reasoning models think at length first.
Read with care: at 50 items the interval is ±0.13, so the ranking is informative but unresolved. A truncated answer scores at chance (0.25), not zero. The only task where turning thinking on raised accuracy.
math500
Tests competition mathematics: algebra, number theory, geometry, counting, with long derivations.
How: five algebra examples from the original MATH train set, then the problem; the model derives and ends with \boxed{answer}. String equality after light LaTeX normalisation, not a computer-algebra check, so \frac{1}{2} and 0.5 do not match.
Read with care: accuracy is a lower bound, so energy per correct is an upper bound. Inherently fat-tailed: the five longest items are 32–50% of the tokens on every engine. Qwen3.5 models often never emit \boxed at all (a format failure, not a maths failure).
What is testedThe models, by family
Size, architecture, and the three features that move energy the most: whether the model has a thinking mode and what its template does when nothing is set, whether it is a mixture of experts (only a fraction of the weights run per token), and whether it carries a vision tower (extra weights that these text-only benchmarks never exercise). "Built for" is the publisher's own positioning; "what we found" is from this corpus.
| Model | Size | Architecture | Features | Built for (publisher) | What we found |
|---|---|---|---|---|---|
| Alibaba (Qwen) | |||||
| Qwen3-0.6B | 0.6B | dense transformer | thinking · default ON | Edge and on-device assistant; the same dual-mode recipe as the larger Qwen3s. Most-downloaded text model on the hub. | Rambles: 20× LFM2.5-1.2B's energy for lower accuracy. The clearest case of size not predicting cost. |
| Qwen3-8B + AWQ · FP8 · Q4_K_M · Q8_0 | 8B | dense transformer | thinking · default ON | General-purpose assistant with switchable step-by-step reasoning, 40k context. Third most-downloaded on the hub. | In the tied top group with thinking on, at 8–10× gpt-oss's energy. With thinking off it drops to 0.75 pooled and the AWQ build does that for 161 kJ; the FP8, Q8_0 and Q4_K_M repacks land between. |
| Qwen3.5-0.8B / 2B / 4B / 9B | 0.8–9B | hybrid: Gated-DeltaNet linear attention + attention | thinking · 0.8B/2B OFF, 4B/9B ONvision | Natively multimodal, 262k-token context, efficient long-context assistants; the switchable thinking mode. | The reference model (9B). Same family, opposite thinking defaults by size. 9B fp16 gets only a 4k window on 24 GB. math500 is a format failure (never emits \boxed). |
| Qwen3.5-9B-AWQ · Qwen3.8-27B-AWQ | 9B · 27B | same hybrid, 4-bit | thinking · default ONvision (9B) | Quantized repacks; the 27B is the family's larger reasoning tier. | AWQ gets the 9B a 32k window where fp16 gets 4k. The 27B cannot run at tier A on 24 GB (OOM in CUDA-graph capture). |
| Qwen2.5-7B-Instruct + AWQ · GPTQ-Int4 · GPTQ-Int8 · Q4_K_M | 7B | dense transformer, GQA | no thinking | The 2024 general assistant, pre-reasoning era; long-text generation, structured output, multilingual. | The continuity anchor and the workhorse: the whole quantization ladder, every power and clock ladder, both engine waves. |
| Qwen3-Coder-30B-A3B-AWQ + Ollama tag | 30B · 3B active | MoE, 128 experts | MoEno thinking | Agentic coding: repository-scale code generation and tool use, 262k context. Not a general assistant. | Scored on four non-coding tasks on purpose. No measurable accuracy deficit against granite-4.2-8b at 46× less energy; 52× less than Qwen3.5-9B. On SGLang, 251 J per correct answer for the whole suite. Ollama's Q4_K_M tag scores 0.81 pooled for 180 kJ, the cheapest configuration in the top statistical group. |
| OpenAI | |||||
| gpt-oss-20b + Ollama tag | 21B · 3.6B active | MoE, MXFP4-native weights | MoEreasoning effort low / medium / high | Open-weight reasoning and agentic model sized to run on 16 GB; the "harmony" format with a graded reasoning dial rather than an on/off switch. | The accuracy anchor of the frontier: 0.857 pooled for 580 kJ, with every dense thinking model near it costing 5.8–17× more. Since the Ollama parity wave, two MoE Q4_K_M tags (Qwen3-Coder, gemma-4-26B) share its statistical group at 1.8–3.2× less energy. |
| IBM | |||||
| granite-4.2-8b | 8B | dense transformer, GQA | thinking · default ON | Enterprise assistant: tool calling, RAG, long documents; Apache 2.0 with a documented training corpus. | In the tied top group on accuracy, at 13–17× gpt-oss's energy. Its thinking switch is load-bearing (trained in with RL). |
| Ornith AI | |||||
| Ornith-1.5-9B | 9B | dense transformer | thinking · default ON | Reasoning and coding specialist, continued-pretrained on Qwen3.5 and Gemma 4 with heavy RL self-improvement; MIT. | Accuracy point-leader (0.873), tied with three others; 5.8× gpt-oss's energy for +0.02. |
| HuggingFace | |||||
| SmolLM3-3B | 3B | dense transformer | thinking · default ON | Small multilingual long-context model with a dual reasoning mode, fully open training recipe. | The model whose unpinned thinking scored 0% on every multiple-choice task and forced the read-the-template rule. No tier-A rows yet. |
| Liquid AI | |||||
| LFM2.5-1.2B-Instruct | 1.2B | hybrid: gated short convolutions + GQA | no thinking | On-device: fast CPU and GPU inference at low memory, built for phones and laptops rather than servers. | 19 J per correct mmlu answer at 0.37 accuracy, the fourth-cheapest tier-A cell in the corpus. Three of four tasks at tier A, so no roster row. Only measured on node C. |
| LFM2.5-2.6B | 2.6B | same hybrid | reasoning always on (forced) | The family's reasoning tier. | Force-opens a thinking block on every answer with no way to turn it off, so it is not comparable to its 1.2B sibling on size alone. |
| PrismML | |||||
| Ternary-Bonsai-1.7B / 4B / 8B | 1.7–8B | dense, Qwen3 bases, ternary weights (2.13 bits) | thinking · pinned OFF | Post-training quantization of Qwen3 checkpoints to ternary weights; accuracy per gigabyte is the publisher's headline and phones are the target. | The 8B at 2.13 bits matches its own F16 render to the digit on three of four tasks for 3.4–8× less energy. The 4B on node C is the cheapest tier-A cell in the corpus: 11 J per correct mmlu answer. |
| Ternary-Bonsai-2-27B · Bonsai-27B | 27B | same, ternary (PQ2_0 / PTQ1_0) and one-bit (Q1_0) | thinking · pinned OFF | The family's largest tier: a 27B in 7–8 GB. | At 1.58 bits a 27B fits an 8 GB RTX 2080 and scores 0.80 pooled. Below 2 bits the saving reverses: PTQ1_0 matches PQ2_0 on accuracy with 13% less VRAM but 1.2–1.6× more energy on Ampere. The one-bit Bonsai-27B holds 0.77. |
| gemma-4-26B-A4B-AWQ · gemma-4-31B-AWQ · gemma-4-e2b-it | 26B (4B active) · 31B · 2B | MoE (26B) · dense (31B, e2b) | MoE (26B)thinking · default OFFvision | Open multimodal generalists in three sizes; the 26B is the efficient MoE tier, the 31B the quality tier. | The 31B scored the highest of anything measured (1.00 gsm8k, 0.96 mmlu) but only in tier B, and cannot run at tier A on 24 GB. The 26B AWQ build does not fit gpqa/math500 at the sanctioned caps on vLLM; Ollama's Q4_K_M tag of it does, at 0.83 pooled for 330 kJ, in the top statistical group. |
| Mistral AI | |||||
| Mistral-7B-Instruct-v0.2 + AWQ · Q4_K_M | 7B | dense, GQA, sliding-window attention | no thinking | The 2023 open 7B that set the small-model bar; general chat. | Pre-reasoning-era control. 0.41 pooled accuracy at 231 kJ: cheap, and behind everything modern. |
| Microsoft | |||||
| Phi-3-mini-4k-instruct | 3.8B | dense, no GQA | no thinking | Small model trained on curated "textbook-quality" data; reasoning-heavy for its size, 4k context. | Its 4k window cannot hold the sanctioned caps: does-not-fit on two tasks. No tier-A rows. |
| TII (Technology Innovation Institute) | |||||
| Falcon-H1R-7B | 7B | Mamba-2 / attention hybrid | reasoning always on | State-space hybrid with a reasoning post-train; long-context efficiency. | In the roster, never run. The only Mamba hybrid. |
| Meta | |||||
| Llama-3.1-8B · Llama-3.2-1B | 8B · 1.2B | dense transformer | no thinking | The deployment-reality anchors: the most-run open models in the world. | The FP16 build stays gated, but Ollama's Q4_K_M tag of the 8B ran on node A on 2026-09-21: 0.89 on gsm8k, 0.24 on math500, tier A on all four tasks. |
Quantization variants (AWQ, GPTQ-Int4/Int8, Q4_K_M, MXFP4, and PrismML's PQ2_0 / PTQ1_0 / Q1_0 below two bits) are the same weights re-packed at fewer bits; each is its own row in the index because two packings of "the same" model differ measurably. Ollama tags are Ollama's own GGUF conversions of the named model.
1 · RosterEvery model at its reference configuration, per node
Tier A, all four tasks pooled at the reference item counts (100 / 100 / 50 / 50), Wilson 95% intervals, one configuration per (model, quantization, engine, node, thinking mode), so a model that ran thinking on and off appears twice. The shaded band is every configuration whose interval overlaps the leader's: one statistical group, spread more than an order of magnitude on energy. This is the only figure that pools tasks.
Suite energy against pooled accuracy
configurations with all four tasks at tier A · suite kJ on a log scale · whiskers are the 95% CI on accuracy · red ring where the gpqa cell cannot beat guessing
2 · Model sizeWalking up a family's size ladder
Arrows step from the smallest to the largest build within one family on one node, tier A only, and within one thinking mode, so a Qwen3.5 ladder with thinking off is a different line from the same ladder with thinking on. Where an arrow goes up without going right, size bought tokens, not answers.
Size ladders, per family and node
arrow = one size step · hollow point = smallest build
3 · ThinkingThinking off → on, ten models, three nodes
Every pair in the corpus where only the thinking setting changed: the same model, task, node, cap, graph mode and items. Arrow colour is the verdict on the Wilson intervals: green where thinking raised accuracy, red where it lowered it, grey where the intervals overlap. Pooled over the tier-A pairs, thinking costs 9× and helps in 6 of 48.
Thinking off → on, per pair
colour is the verdict · hollow = thinking off · label carries the cost ratio · only tier-A pairs are drawn
3b · Reasoning effortgpt-oss at low → medium → high
The only graded reasoning dial in the roster, run on all four tasks at tier A with gpqa at the full 198 items. Every arrow is short: accuracy intervals overlap on every task, token counts sit within 4% of each other, and energy moves by less than the run-to-run floor. The dial never reached the model. The low and high reasoning traces are byte-identical on 142 of 198 gpqa items, while each differs from the medium row on 185: vLLM 0.25.1 takes harmony reasoning effort from the top-level reasoning_effort request field, and the harness sent it inside chat_template_kwargs. All three rungs ran at medium, so what the arrows show is boot-to-boot drift on an MoE, not a dose-response. Kept here because it is the kind of mistake anyone measuring this dial can make.
Effort ladder, per task
hollow = low · arrows step to medium then high · gpqa at n=198 · every rung actually ran at medium
3c · Item countgpqa at 50 → 198 items
The accuracy interval halves in width from 50 to 198 items, and the point estimates regress toward each other: Qwen3.5-4B fell from 0.74 to 0.68, gpt-oss rose from 0.59 to 0.61–0.65. They still overlap. What 198 items did resolve is the bottom of the frontier: the no-thinking coder's 0.42 is now clearly below both, at 886 J per correct answer against 14–64 kJ.
gpqa_diamond, 50 → 198 items, same configuration
hollow = 50 items · arrow ends at 198 · hover for the intervals
4 · Power limitEvery power ladder, every node
Nine ladders across three cards and three models. The knee sits near 225 W on both 3090s. Accuracy is flat on every ladder, so the plane view would be nine vertical lines; these keep the cap on the x-axis so the knee is visible.
Energy per correct answer against power cap
per node · colour is task · label is the model · repeats averaged
Throughput against power cap
same ladders
5 · Clock lockEvery clock ladder
Six ladders: three tasks on node A, two on node B, one on node C. Node A's gsm8k minimum is at 1350 MHz. The second panel is the finding that a "smoothness" grade, as first defined, rewards a slower card: power CV rises with clock even with no cap in play.
Energy per correct answer against SM clock
per node · colour is task · label is the model
Power CV against SM clock
same ladders
6 · NodeThe same cell on two or three cards
139 cells ran the same configuration on more than one node. Accuracy lands in the same column; energy does not. Node A at 390 W sits above node B at 350 W on almost every cell, and node C's 2080 is cheaper still on the small models it can hold.
Same configuration, different node
colour is node · a line joins one cell across nodes
7 · CUDA graphsEager → graphs on
26 pairs where only the graph flag changed. On decode-heavy tasks the arrow points down (graphs save energy); on the two-token multiple-choice tasks it can point up. Accuracy never moves.
Eager → CUDA graphs, per pair
hollow = eager · arrow ends at graphs on
8 · EngineThe same weights on different engines
34 cells from both engine waves. Within a cell the weights are byte-identical (vLLM ↔ SGLang share AWQ; llama.cpp ↔ Ollama share Q4_K_M; gpt-oss is MXFP4 on both). Accuracy ties everywhere; the vertical spread is the engine.
Engines, per cell
colour is engine · a line joins one cell across engines
9 · RepeatsThe same cell run again
22 clusters of three or more runs. Dense models land on the same point every boot; mixtures-of-experts scatter 2–4% in energy; and the earliest repeat studies, run before prefix caching was disabled, walk downward as the cache warms.
Repeat clusters
colour is architecture · points in one cluster are joined in run order
10 · Output capRaising the cap, up to uncapped
123 sequences: the deliberate cap ladders plus every capped/uncapped pair in the reference wave. On gsm8k the arrows stop moving after 800 tokens. On math500 they keep climbing. Uncapped points at the top are the runaway items measuring the context window.
Cap sequences, smallest cap → uncapped
hollow = smallest cap · one line per model and configuration
11 · Quantizationfp16 → Int8 → Int4 → AWQ
Every precision sequence on a shared configuration, one panel per model and node, the benchmarks as the coloured series. The full four-step ladder exists on Qwen2.5-7B on node B; two-step fp16 → AWQ pairs exist for Qwen2.5-7B and Qwen3.5-9B on node A. Arrows go down, rarely left. Each line states its tier; only sequences with the same tier are comparable to each other.
Precision sequences, per model and node
colour is benchmark · hollow = fp16 · label is the tier
CoverageWhere each model has been swept
Distinct levels per axis for every build with runs. ⊘ marks a does-not-fit verdict somewhere in the fleet.
Model × axis coverage
39 builds