async.energy

Blog

Where the energy goes when you run AI on your own hardware

A question-and-answer introduction to what decides the energy cost of local inference: hardware, model, engine, task, settings, and idle time. With sources, and numbers measured on this site's benchmark.

· 19 min read · local-ai, energy, measurement, agents

The question. Say you run local AI models on consumer hardware, as agents or for everyday tasks such as a personal assistant. What decides how much energy that uses? The hardware, the model, the engine, the task? The settings, such as thinking on or off, context length, the type of model? What else?

This post is the answer. It is written for a person setting up a local assistant and for an agent deciding how to spend its own compute. Where a claim has a number behind it, the number comes from a paper linked inline or from runs on this site’s own benchmark, which measures joules per correct answer on consumer GPUs. The methodology page says how those runs were measured and where their limits are.

Start with power times time

Energy is power multiplied by time, summed over everything the machine does, including the hours it does nothing. Every factor below changes one or both terms: how many watts the system draws, and for how long. For a personal assistant the number to manage is energy per completed task, and energy per day for the whole setup. Tokens per second and peak watts are inputs to that number, not the number itself.

Prefill and decode

Inference has two phases with different energy profiles.

Prefill processes the input: the system prompt, tool definitions, documents, conversation history, tool outputs. It is compute-bound and parallel, so a GPU gets through thousands of input tokens per second. Energy per input token is low, but the attention cost of each new token grows with the context already in place.

Decode produces the output one token at a time. Each step reads the model’s active weights and the whole KV cache out of memory, so the phase is bound by memory bandwidth and the processor spends much of its time waiting (Pope et al. 2022). Energy per output token is several times energy per input token.

Most of what follows makes sense once you ask which phase a factor touches. A task that reads a long document and answers in a sentence is prefill. A task that writes a page from a short prompt is decode.

The two phases even look different on a power meter. Prefill is a burst; decode is a plateau.

GPU power over time for a prefill-shaped task and a decode-shaped task, same model, same card Two 5 Hz GPU power traces from Qwen2.5-7B-Instruct on node A. The multiple-choice task draws a spiky burst for 17 seconds with a power CV of 0.24; the math task draws a flat plateau for 323 seconds with a power CV of 0.05. Prefill-shaped: mmlu_redux, a long prompt and a two-token answer mean 285 W · σ 67 W · peak 387 W · 17 s · power CV 0.24 100 W 200 W 300 W 400 W 0 s 5 s 10 s 15 s mmlu_redux: 86 samples at 5 Hz over 17.4 s, mean 285.3 W, peak 387.3 W, CV 0.236 mean 285 W Decode-shaped: gsm8k_platinum, a short prompt and a long worked answer mean 379 W · σ 19 W · peak 390 W · 323 s · power CV 0.05 100 W 200 W 300 W 400 W 0 s 60 s 120 s 180 s 240 s 300 s gsm8k_platinum: 1575 samples at 5 Hz over 322.9 s, mean 379.1 W, peak 390.1 W, CV 0.051 mean 379 W
Prefill vs decode: is prefill really lumpier? Two runs of Qwen2.5-7B-Instruct on the same RTX 3090 (node A), GPU power sampled at 5 Hz, drawn from the samples with the ±σ band shaded and the mean in amber. The prefill-shaped task whipsaws between 200 W and the 390 W cap and is over in 17 seconds; the decode-shaped task holds a plateau near the cap for five minutes. The coefficient of variation is the ratio σ ÷ mean; every model and engine on this benchmark shows the same 3 to 5× split between the two shapes. Plotted lines are per-bin means of the 5 Hz samples; the statistics are computed from all samples.
RunShapeDurationMean powerPeakPower CV
mmlu_redux, 100 itemsprefill17 s285 W387 W0.24
gsm8k_platinum, 100 itemsdecode323 s379 W390 W0.05

On this benchmark the split shows up as a caching effect. On a multiple-choice task where the prompt dominates, a warm prefix cache cut a run from 17.2 s and 4,902 J to 9.6 s and 2,712 J, with identical answers. On a math task where the model writes long solutions, the same cache moved total energy by a quarter of one percent.

Hardware

Memory bandwidth sets decode speed. That is why unified-memory Macs, high-end GPUs with GDDR or HBM, and the newer wide-bus “AI PC” chips decode well, and why a CPU on dual-channel DDR5 is slow. The spec sheets tell the story: an RTX 3090 moves about 936 GB/s, an M4 Max about 546 GB/s, and dual-channel DDR5-5600 about 90 GB/s. Slow generation keeps the machine at high power for longer.

Wall power differs by an order of magnitude between machines that both work. The RTX 3090 on this site’s main test node sits pinned at its 390 W GPU power cap for 89% of a long run, before counting the host. Apple lists the Mac Studio with M4 Max at 6 W idle and 145 W maximum, and the M3 Ultra at 9 W and 270 W (Apple). The Mac generates more slowly. Which one spends fewer joules per task depends on the model and the workload, so measure rather than assume.

Fit the model in fast memory. If weights spill from VRAM into system RAM, or layers are offloaded to the CPU, the spilled part decodes at system-RAM bandwidth. The bandwidth ratio above predicts about a tenfold slowdown for those layers, while power stays high. That combination is usually the worst case for energy. The 24 GB cards on this fleet hold about 19 GB of weights once the engine has its working memory; larger models here run in a 4-bit format or not at all.

Cap the GPU’s power. GPUs ship tuned for peak speed. On both RTX 3090s in this fleet, a power cap near 225 W (58 to 64% of stock) cut energy per correct answer by 24 to 34%, cost 2 to 21% of throughput depending on the task, and left accuracy identical at every setting. Above that knee the card buys clock speed that memory-bound decode cannot use. Two of these cards doing identical work, one at 390 W and one settling near 300 W: the 390 W card finished 0.2% sooner and spent 22% more energy per token. Capping the clock does the same job: 1,350 MHz on a 3090 gave 20% less energy for 11% less throughput. Samsi et al. 2023 saw the same shape on datacenter cards. On NVIDIA the cap is one command, and it resets at reboot:

nvidia-smi -q -d POWER          # note the default limit before you change it
sudo nvidia-smi -pl 225         # cap this card at 225 W; restore the default when done

Idle draw is where the kilowatt-hours hide on an always-on machine. The GPUs alone on this fleet idle at 19 to 42 W with nothing loaded, measured at the card. This site has not metered a whole desktop at idle, and a host adds its own tens of watts on top. At 80 W, a machine that never sleeps uses about 1.9 kWh a day and 700 kWh a year, more than the inference itself for a lightly used assistant. A machine that idles in single digits of watts, as Apple’s figures above show a Mac Studio does, changes that sum by an order of magnitude.

The rest of the box counts. PSU efficiency drops at low load, which is where an idle desktop lives. CPU, RAM, fans, storage, and any NPU running small side models all draw from the same plug. Only a wall meter sees them.

The model

Active parameters drive decode cost. Every generated token reads the active weights. A dense 70B costs far more per token than a dense 8B.

Mixture-of-experts separates memory from work. An MoE model with 30B total and 3B active parameters needs the memory of a 30B model but decodes close to a 3B one, because each token routes through a few experts (Fedus et al. 2022). On this benchmark, a 30B MoE with 3B active used 52 times less energy than a dense 9B on the same suite, and a 20B MoE beat a dense 8B by 13 to 17 times. This is one of the biggest levers a local setup has, provided the whole model fits in fast memory.

Parameter count alone predicts nothing. The smallest model in this corpus, a 0.6B, burned 20 times the energy of a 1.2B because it thinks by default and writes long. What predicts energy is output length, which is a property of the model, its thinking default, and the engine together.

Quantization buys energy, not accuracy. Smaller weights mean less data per token, so faster decode and fewer joules, and larger models fit (GPTQ, AWQ). On a 7B model here, 8-bit reproduced the fp16 accuracy exactly at 1.8× less energy, and 4-bit gave up four points of accuracy for 4.3× less. Which 4-bit format was cheapest changed by task. Push quantization too far and quality drops, and lost quality has an energy cost of its own, covered under tasks below.

Architecture decides the cost of long context. Grouped-query attention (Ainslie et al. 2023), sliding-window attention, multi-head latent attention (DeepSeek-V2), and hybrid designs with linear-attention or state-space layers all shrink the KV cache, which is what an agent’s growing history is made of.

Speculative decoding lets a small draft model, or a built-in multi-token head, propose several tokens that the large model verifies in one pass (Leviathan et al. 2023). When the drafts are usually accepted it cuts decode time by a large fraction. When they are not, it adds work.

The engine

llama.cpp, Ollama, MLX, vLLM, SGLang, and ExLlama can differ by a lot on the same weights and hardware. Where the differences come from:

  • Kernels. Flash attention and quantized matrix multiplies tuned for your hardware. On this benchmark, SGLang was cheaper per token than vLLM in 8 of 8 cells, by 8 to 24%. On a MoE that became 40% less total energy; on a dense 7B, about none. Accuracy never separated any pair of engines. Engine choice changes what an answer costs, not whether it is right.
  • Per-request overhead. Ollama’s cost over llama.cpp on byte-identical weights was per request, not per token: 370% more energy on a task with one-to-twenty-token answers, 15% more on long solutions. Short replies are where an assistant lives.
  • Prefix caching. For agents this may be the largest software factor. If the system prompt, tool definitions and history are cached, each turn prefills only the new part. Without it, every step re-reads the whole growing context, and total prefill work over a session grows with roughly the square of its length. vLLM and SGLang cache by default and llama.cpp reuses the last prompt. Check what yours does, and whether tool calls or a rotating system prompt keep breaking the prefix.
  • KV cache quantization reduces memory pressure for long contexts.
  • Batching. One user at batch size 1 leaves most of the GPU idle, and most of the energy goes into moving weights. Serving several requests at once amortizes that, which is why throughput per joule in a datacenter is many times what a single-user box gets (Pope et al. 2022; Kwon et al. 2023). Locally it matters when an agent runs parallel subtasks or several services share one model.
  • Engine flags. Not every flag is a free win. On this benchmark, running vLLM in eager mode instead of with CUDA graphs changed energy by anywhere from −22% to +78%, depending on model and task shape. Graphs lower per-token cost and raise time to first token, so decode-heavy work wins with them and prefill-heavy work can lose.
  • Load and unload. Keep-alive decides whether the model stays in VRAM (fast responses, higher idle) or unloads (lower idle, plus the time and energy of reloading from disk on every cold start).

The task and the settings

Energy per successful outcome is the number that matters. A small model that fails, loops, calls the wrong tool, or needs three tries can use more total energy than a larger model that gets it right once. Size, quantization and quality trade off here, and the trade sharpens in agentic work, where errors compound over many steps. On this benchmark’s math set, two models both scored 0.99 and one used 11.4 times less energy; a third scored 0.94 at 24 times the cheaper one’s cost. The accuracy column separates them by six points, the energy column by a factor of 24. Energy per correct answer as a metric has prior art at datacenter scale (Intelligence per Watt); the reason to prefer it over joules per token is exactly this ranking.

Thinking mode multiplies output tokens. Across 48 paired runs on nine models, thinking cost 9 times the energy on average. It helped in 6 pairs, hurt in 10, and made no measurable difference in 32. Every gain that cleared the noise cost at least 100 times the energy. On small models thinking was strictly worse: lower accuracy at 10 to 386 times the cost, with most items running out of room before an answer. For “turn on the lights” or “summarize this email” it is wasted energy. For a hard multi-step problem it can pay for itself by preventing a failed run. Route by difficulty, and pin the setting rather than let the model choose.

Input-heavy and output-heavy tasks behave differently. Summarizing a long document is mostly prefill: cheap per token, large in volume. Writing long content is mostly decode. Classification and extraction with short outputs are very cheap.

Context length is paid on every turn. It raises attention cost and KV cache size. Stuffing retrieved documents in “just in case”, or never pruning agent history, costs energy each time the model runs. A related lever: give the model a large window but cap output per task. With the same model, items and seed, letting four runaway items run to a 64k wall made them 89% of all tokens generated and more than tripled energy per correct answer, with accuracy unchanged.

Agent loop design. How many steps, how verbose the tool outputs are (a tool returning 20,000 tokens of raw HTML is expensive to prefill), how often the agent reflects, and whether it polls or reacts to events. A single runaway item carried 40% of a run’s energy in this corpus. A loop with no ceiling will produce such items.

Output verbosity. Asking for concise answers reduces decode energy in direct proportion.

Routing and cascading. Use a tiny model for intent classification, simple commands and embeddings, and escalate to a large one only when needed. At the bottom of this leaderboard a 1.2B model answers for 432 J per correct answer at 49% accuracy; the frontier MoE answers for 2,257 J at 86%. A router that sends only the hard questions upward gets most of the accuracy for a fraction of the energy.

The whole pipeline, not just the LLM

A personal assistant is rarely one model. There may be wake-word detection running all the time, speech-to-text, text-to-speech, an embedding model for memory or search, a vision encoder for images and screenshots (each image becomes hundreds or thousands of prefill tokens), a vector database, and whatever the tools themselves do. An always-listening pipeline has a baseline draw independent of how often you use it.

Beyond the machine

Duty cycle usually dominates. Total energy is roughly idle power times idle hours plus active power times active hours. For most personal use, active inference adds up to minutes or an hour a day, so idle behavior and the ability to sleep and wake matter more than inference efficiency.

Heat. Nearly every watt ends up as heat in the room. In summer it adds cooling load. In a heated home in winter it offsets some heating.

Grid and timing. The carbon intensity of your electricity changes by region and by hour (Electricity Maps), and so does its price. Deferring batch work such as indexing, summarizing, or overnight agent jobs to cheap or low-carbon hours changes both the bill and the emissions. That scheduling is what this site’s controller does.

Embodied energy. Manufacturing a GPU or a new computer carries a large carbon footprint of its own (Gupta et al. 2021). Buying hardware for local AI has an upfront cost that takes a long time to recover through efficiency. Reusing hardware you already own has none.

Local versus cloud. Datacenters batch heavily, run custom accelerators, and keep utilization high, so cloud inference is often more efficient per token. Google reports a median text prompt to Gemini at 0.24 Wh, counting idle machines, host CPU and RAM, and datacenter overhead (Elsworth et al. 2025). The industry-average overhead ratio, PUE, is about 1.56 against 1.1 at the best operators (Uptime Institute 2024), and datacenters drew about 415 TWh in 2024, 1.5% of world electricity (IEA 2025). Against that, the cloud adds network transfer and often a far larger model than a task needs. A small local model on efficient hardware for simple tasks competes well. A large local model at batch size 1 on a power-hungry desktop usually does not. Privacy, latency and availability are separate reasons to run local, and they are yours to weigh.

Measure it

Estimates are unreliable enough that measuring is worth the trouble.

  • A plug-in wall meter, around $20, captures the whole system including idle and PSU loss. Its counter is coarse. The plugs on this fleet tick in 0.01 kWh steps, so short runs need the GPU counter instead.
  • Component readings. nvidia-smi reports GPU power, and NVML exposes a cumulative energy counter on cards since Volta; powermetrics reports package power on macOS; RAPL covers Intel and AMD CPUs. They miss PSU loss and the rest of the box. On this benchmark, integrating power samples at 5 Hz overestimated the hardware energy counter by 2.6 to 5.9%, so prefer the counter when you have one.
  • Anything else on the GPU corrupts the reading. One stray job sharing the card here added 15.8% to energy per correct answer.
  • Know the noise floor. Energy repeats to about 1% on a dense model between boots, and 2 to 4.5% on a MoE, which emits a different token count every boot even at temperature zero. Accuracy is far noisier. On a 300-item suite the 95% interval is about ±5 points; at 50 items it swallows most of the differences between models, and energy per correct answer inherits that width. Spend GPU time on more items before spending it on more repeats.

The numbers to keep: joules per token, split by prefill and decode; watt-hours per typical task; and kWh per day for the whole setup with its idle included. Task-level energy is the one that transfers between setups (Luccioni et al. 2024), and the ML.ENERGY benchmark is the reference point for datacenter GPUs.

Where the biggest gains usually are

For a typical always-on assistant, in rough order of how much each one moved the number here:

  1. Hardware that idles low, and a sleep policy for the hours nobody is asking.
  2. A model that fits in fast memory, with nothing spilling.
  3. An MoE or a right-sized model at 8-bit or 4-bit.
  4. Prompt caching that survives your agent’s turn structure.
  5. Thinking pinned off, and switched on by a router only for tasks that earn it.
  6. Concise outputs and lean tool results.
  7. A power cap near the knee on any discrete GPU.

Which of these you pull is your call. The order is a starting point, and your own wall meter has the final say.

What to do next

The measured rows behind the claims above are public, and the two tools that produced them are free to run.

  1. Measure your own machine. The controller’s bench quick runs this benchmark against the engine you already have in about 25 minutes, no wall meter needed, and writes the joules per correct answer for your box. Opt in and the run joins the leaderboard under CC BY 4.0. The Measure page has the one-line install.
  2. Read the corpus. The leaderboard and model picker for people; the energy data page for every sweep in this post as a chart with a table view. For agents, the same data is JSON, plus a recommendation endpoint that takes a workload shape and an MCP server, all listed in llms.txt.
  3. Move the work you can move. Anything with a deadline rather than a start time, such as agent runs, indexing, batch inference and fine-tunes, can run in the cheapest electricity hours of your day. The same controller schedules it and reports the measured energy of every run. Interactive use cannot be moved, which is why the first two steps matter more for a personal assistant.

The methodology page says what each number means and where the limits are.

Sources

Benchmark figures in this post come from the energy-bench reference document (findings F1 to F10 and the sweep record, runs through 2026-09-09), measured on RTX 3090 and RTX 2080 nodes at temperature 0 with the GPU’s own energy counter. The methodology page describes the instrument; the rows are on the leaderboard.

  • Pope, R. et al. (2022). Efficiently Scaling Transformer Inference. arXiv:2211.05102.
  • Kwon, W. et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180.
  • Leviathan, Y., Kalman, M., Matias, Y. (2023). Fast Inference from Transformers via Speculative Decoding. arXiv:2211.17192.
  • Ainslie, J. et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245.
  • Fedus, W., Zoph, B., Shazeer, N. (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv:2101.03961.
  • DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434.
  • Frantar, E. et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323.
  • Lin, J. et al. (2023). AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. arXiv:2306.00978.
  • Samsi, S. et al. (2023). From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference. arXiv:2310.03003.
  • Chung, J.-W. et al. (2025). The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization. arXiv:2505.06371. Leaderboard at ml.energy.
  • Luccioni, S., Jernite, Y., Strubell, E. (2024). Power Hungry Processing: Watts Driving the Cost of AI Deployment? FAccT 2024. arXiv:2311.16863.
  • Stanford Hazy Research (2025). Intelligence per Watt. arXiv:2511.07885.
  • Elsworth, C. et al. (2025). Measuring the environmental impact of delivering AI at Google Scale. arXiv:2508.15734.
  • Uptime Institute (2024). Global Data Center Survey 2024.
  • International Energy Agency (2025). Energy and AI.
  • Gupta, U. et al. (2021). Chasing Carbon: The Elusive Environmental Footprint of Computing. HPCA 2021. arXiv:2011.02839.
  • Apple. Mac Studio power consumption and thermal output.
  • Electricity Maps. Live carbon intensity by region.