Where the energy goes when you run AI on your own hardware
A question-and-answer introduction to what decides the energy cost of local inference: hardware, model, engine, task, settings, and idle time. With sources, and numbers measured on this site's benchmark.
· 19 min read · local-ai, energy, measurement, agents
The question. Say you run local AI models on consumer hardware, as agents or for everyday tasks such as a personal assistant. What decides how much energy that uses? The hardware, the model, the engine, the task? The settings, such as thinking on or off, context length, the type of model? What else?
This post is the answer. It is written for a person setting up a local assistant and for an agent deciding how to spend its own compute. Where a claim has a number behind it, the number comes from a paper linked inline or from runs on this site’s own benchmark, which measures joules per correct answer on consumer GPUs. The methodology page says how those runs were measured and where their limits are.
Start with power times time
Energy is power multiplied by time, summed over everything the machine does, including the hours it does nothing. Every factor below changes one or both terms: how many watts the system draws, and for how long. For a personal assistant the number to manage is energy per completed task, and energy per day for the whole setup. Tokens per second and peak watts are inputs to that number, not the number itself.
Prefill and decode
Inference has two phases with different energy profiles.
Prefill processes the input: the system prompt, tool definitions, documents, conversation history, tool outputs. It is compute-bound and parallel, so a GPU gets through thousands of input tokens per second. Energy per input token is low, but the attention cost of each new token grows with the context already in place.
Decode produces the output one token at a time. Each step reads the model’s active weights and the whole KV cache out of memory, so the phase is bound by memory bandwidth and the processor spends much of its time waiting (Pope et al. 2022). Energy per output token is several times energy per input token.
Most of what follows makes sense once you ask which phase a factor touches. A task that reads a long document and answers in a sentence is prefill. A task that writes a page from a short prompt is decode.
The two phases even look different on a power meter. Prefill is a burst; decode is a plateau.
| Run | Shape | Duration | Mean power | Peak | Power CV |
|---|---|---|---|---|---|
| mmlu_redux, 100 items | prefill | 17 s | 285 W | 387 W | 0.24 |
| gsm8k_platinum, 100 items | decode | 323 s | 379 W | 390 W | 0.05 |
On this benchmark the split shows up as a caching effect. On a multiple-choice task where the prompt dominates, a warm prefix cache cut a run from 17.2 s and 4,902 J to 9.6 s and 2,712 J, with identical answers. On a math task where the model writes long solutions, the same cache moved total energy by a quarter of one percent.
Hardware
Memory bandwidth sets decode speed. That is why unified-memory Macs, high-end GPUs with GDDR or HBM, and the newer wide-bus “AI PC” chips decode well, and why a CPU on dual-channel DDR5 is slow. The spec sheets tell the story: an RTX 3090 moves about 936 GB/s, an M4 Max about 546 GB/s, and dual-channel DDR5-5600 about 90 GB/s. Slow generation keeps the machine at high power for longer.
Wall power differs by an order of magnitude between machines that both work. The RTX 3090 on this site’s main test node sits pinned at its 390 W GPU power cap for 89% of a long run, before counting the host. Apple lists the Mac Studio with M4 Max at 6 W idle and 145 W maximum, and the M3 Ultra at 9 W and 270 W (Apple). The Mac generates more slowly. Which one spends fewer joules per task depends on the model and the workload, so measure rather than assume.
Fit the model in fast memory. If weights spill from VRAM into system RAM, or layers are offloaded to the CPU, the spilled part decodes at system-RAM bandwidth. The bandwidth ratio above predicts about a tenfold slowdown for those layers, while power stays high. That combination is usually the worst case for energy. The 24 GB cards on this fleet hold about 19 GB of weights once the engine has its working memory; larger models here run in a 4-bit format or not at all.
Cap the GPU’s power. GPUs ship tuned for peak speed. On both RTX 3090s in this fleet, a power cap near 225 W (58 to 64% of stock) cut energy per correct answer by 24 to 34%, cost 2 to 21% of throughput depending on the task, and left accuracy identical at every setting. Above that knee the card buys clock speed that memory-bound decode cannot use. Two of these cards doing identical work, one at 390 W and one settling near 300 W: the 390 W card finished 0.2% sooner and spent 22% more energy per token. Capping the clock does the same job: 1,350 MHz on a 3090 gave 20% less energy for 11% less throughput. Samsi et al. 2023 saw the same shape on datacenter cards. On NVIDIA the cap is one command, and it resets at reboot:
nvidia-smi -q -d POWER # note the default limit before you change it
sudo nvidia-smi -pl 225 # cap this card at 225 W; restore the default when done
Idle draw is where the kilowatt-hours hide on an always-on machine. The GPUs alone on this fleet idle at 19 to 42 W with nothing loaded, measured at the card. This site has not metered a whole desktop at idle, and a host adds its own tens of watts on top. At 80 W, a machine that never sleeps uses about 1.9 kWh a day and 700 kWh a year, more than the inference itself for a lightly used assistant. A machine that idles in single digits of watts, as Apple’s figures above show a Mac Studio does, changes that sum by an order of magnitude.
The rest of the box counts. PSU efficiency drops at low load, which is where an idle desktop lives. CPU, RAM, fans, storage, and any NPU running small side models all draw from the same plug. Only a wall meter sees them.
The model
Active parameters drive decode cost. Every generated token reads the active weights. A dense 70B costs far more per token than a dense 8B.
Mixture-of-experts separates memory from work. An MoE model with 30B total and 3B active parameters needs the memory of a 30B model but decodes close to a 3B one, because each token routes through a few experts (Fedus et al. 2022). On this benchmark, a 30B MoE with 3B active used 52 times less energy than a dense 9B on the same suite, and a 20B MoE beat a dense 8B by 13 to 17 times. This is one of the biggest levers a local setup has, provided the whole model fits in fast memory.
Parameter count alone predicts nothing. The smallest model in this corpus, a 0.6B, burned 20 times the energy of a 1.2B because it thinks by default and writes long. What predicts energy is output length, which is a property of the model, its thinking default, and the engine together.
Quantization buys energy, not accuracy. Smaller weights mean less data per token, so faster decode and fewer joules, and larger models fit (GPTQ, AWQ). On a 7B model here, 8-bit reproduced the fp16 accuracy exactly at 1.8× less energy, and 4-bit gave up four points of accuracy for 4.3× less. Which 4-bit format was cheapest changed by task. Push quantization too far and quality drops, and lost quality has an energy cost of its own, covered under tasks below.
Architecture decides the cost of long context. Grouped-query attention (Ainslie et al. 2023), sliding-window attention, multi-head latent attention (DeepSeek-V2), and hybrid designs with linear-attention or state-space layers all shrink the KV cache, which is what an agent’s growing history is made of.
Speculative decoding lets a small draft model, or a built-in multi-token head, propose several tokens that the large model verifies in one pass (Leviathan et al. 2023). When the drafts are usually accepted it cuts decode time by a large fraction. When they are not, it adds work.
The engine
llama.cpp, Ollama, MLX, vLLM, SGLang, and ExLlama can differ by a lot on the same weights and hardware. Where the differences come from:
- Kernels. Flash attention and quantized matrix multiplies tuned for your hardware. On this benchmark, SGLang was cheaper per token than vLLM in 8 of 8 cells, by 8 to 24%. On a MoE that became 40% less total energy; on a dense 7B, about none. Accuracy never separated any pair of engines. Engine choice changes what an answer costs, not whether it is right.
- Per-request overhead. Ollama’s cost over llama.cpp on byte-identical weights was per request, not per token: 370% more energy on a task with one-to-twenty-token answers, 15% more on long solutions. Short replies are where an assistant lives.
- Prefix caching. For agents this may be the largest software factor. If the system prompt, tool definitions and history are cached, each turn prefills only the new part. Without it, every step re-reads the whole growing context, and total prefill work over a session grows with roughly the square of its length. vLLM and SGLang cache by default and llama.cpp reuses the last prompt. Check what yours does, and whether tool calls or a rotating system prompt keep breaking the prefix.
- KV cache quantization reduces memory pressure for long contexts.
- Batching. One user at batch size 1 leaves most of the GPU idle, and most of the energy goes into moving weights. Serving several requests at once amortizes that, which is why throughput per joule in a datacenter is many times what a single-user box gets (Pope et al. 2022; Kwon et al. 2023). Locally it matters when an agent runs parallel subtasks or several services share one model.
- Engine flags. Not every flag is a free win. On this benchmark, running vLLM in eager mode instead of with CUDA graphs changed energy by anywhere from −22% to +78%, depending on model and task shape. Graphs lower per-token cost and raise time to first token, so decode-heavy work wins with them and prefill-heavy work can lose.
- Load and unload. Keep-alive decides whether the model stays in VRAM (fast responses, higher idle) or unloads (lower idle, plus the time and energy of reloading from disk on every cold start).
The task and the settings
Energy per successful outcome is the number that matters. A small model that fails, loops, calls the wrong tool, or needs three tries can use more total energy than a larger model that gets it right once. Size, quantization and quality trade off here, and the trade sharpens in agentic work, where errors compound over many steps. On this benchmark’s math set, two models both scored 0.99 and one used 11.4 times less energy; a third scored 0.94 at 24 times the cheaper one’s cost. The accuracy column separates them by six points, the energy column by a factor of 24. Energy per correct answer as a metric has prior art at datacenter scale (Intelligence per Watt); the reason to prefer it over joules per token is exactly this ranking.
Thinking mode multiplies output tokens. Across 48 paired runs on nine models, thinking cost 9 times the energy on average. It helped in 6 pairs, hurt in 10, and made no measurable difference in 32. Every gain that cleared the noise cost at least 100 times the energy. On small models thinking was strictly worse: lower accuracy at 10 to 386 times the cost, with most items running out of room before an answer. For “turn on the lights” or “summarize this email” it is wasted energy. For a hard multi-step problem it can pay for itself by preventing a failed run. Route by difficulty, and pin the setting rather than let the model choose.
Input-heavy and output-heavy tasks behave differently. Summarizing a long document is mostly prefill: cheap per token, large in volume. Writing long content is mostly decode. Classification and extraction with short outputs are very cheap.
Context length is paid on every turn. It raises attention cost and KV cache size. Stuffing retrieved documents in “just in case”, or never pruning agent history, costs energy each time the model runs. A related lever: give the model a large window but cap output per task. With the same model, items and seed, letting four runaway items run to a 64k wall made them 89% of all tokens generated and more than tripled energy per correct answer, with accuracy unchanged.
Agent loop design. How many steps, how verbose the tool outputs are (a tool returning 20,000 tokens of raw HTML is expensive to prefill), how often the agent reflects, and whether it polls or reacts to events. A single runaway item carried 40% of a run’s energy in this corpus. A loop with no ceiling will produce such items.
Output verbosity. Asking for concise answers reduces decode energy in direct proportion.
Routing and cascading. Use a tiny model for intent classification, simple commands and embeddings, and escalate to a large one only when needed. At the bottom of this leaderboard a 1.2B model answers for 432 J per correct answer at 49% accuracy; the frontier MoE answers for 2,257 J at 86%. A router that sends only the hard questions upward gets most of the accuracy for a fraction of the energy.
The whole pipeline, not just the LLM
A personal assistant is rarely one model. There may be wake-word detection running all the time, speech-to-text, text-to-speech, an embedding model for memory or search, a vision encoder for images and screenshots (each image becomes hundreds or thousands of prefill tokens), a vector database, and whatever the tools themselves do. An always-listening pipeline has a baseline draw independent of how often you use it.
Beyond the machine
Duty cycle usually dominates. Total energy is roughly idle power times idle hours plus active power times active hours. For most personal use, active inference adds up to minutes or an hour a day, so idle behavior and the ability to sleep and wake matter more than inference efficiency.
Heat. Nearly every watt ends up as heat in the room. In summer it adds cooling load. In a heated home in winter it offsets some heating.
Grid and timing. The carbon intensity of your electricity changes by region and by hour (Electricity Maps), and so does its price. Deferring batch work such as indexing, summarizing, or overnight agent jobs to cheap or low-carbon hours changes both the bill and the emissions. That scheduling is what this site’s controller does.
Embodied energy. Manufacturing a GPU or a new computer carries a large carbon footprint of its own (Gupta et al. 2021). Buying hardware for local AI has an upfront cost that takes a long time to recover through efficiency. Reusing hardware you already own has none.
Local versus cloud. Datacenters batch heavily, run custom accelerators, and keep utilization high, so cloud inference is often more efficient per token. Google reports a median text prompt to Gemini at 0.24 Wh, counting idle machines, host CPU and RAM, and datacenter overhead (Elsworth et al. 2025). The industry-average overhead ratio, PUE, is about 1.56 against 1.1 at the best operators (Uptime Institute 2024), and datacenters drew about 415 TWh in 2024, 1.5% of world electricity (IEA 2025). Against that, the cloud adds network transfer and often a far larger model than a task needs. A small local model on efficient hardware for simple tasks competes well. A large local model at batch size 1 on a power-hungry desktop usually does not. Privacy, latency and availability are separate reasons to run local, and they are yours to weigh.
Measure it
Estimates are unreliable enough that measuring is worth the trouble.
- A plug-in wall meter, around $20, captures the whole system including idle and PSU loss. Its counter is coarse. The plugs on this fleet tick in 0.01 kWh steps, so short runs need the GPU counter instead.
- Component readings.
nvidia-smireports GPU power, and NVML exposes a cumulative energy counter on cards since Volta;powermetricsreports package power on macOS; RAPL covers Intel and AMD CPUs. They miss PSU loss and the rest of the box. On this benchmark, integrating power samples at 5 Hz overestimated the hardware energy counter by 2.6 to 5.9%, so prefer the counter when you have one. - Anything else on the GPU corrupts the reading. One stray job sharing the card here added 15.8% to energy per correct answer.
- Know the noise floor. Energy repeats to about 1% on a dense model between boots, and 2 to 4.5% on a MoE, which emits a different token count every boot even at temperature zero. Accuracy is far noisier. On a 300-item suite the 95% interval is about ±5 points; at 50 items it swallows most of the differences between models, and energy per correct answer inherits that width. Spend GPU time on more items before spending it on more repeats.
The numbers to keep: joules per token, split by prefill and decode; watt-hours per typical task; and kWh per day for the whole setup with its idle included. Task-level energy is the one that transfers between setups (Luccioni et al. 2024), and the ML.ENERGY benchmark is the reference point for datacenter GPUs.
Where the biggest gains usually are
For a typical always-on assistant, in rough order of how much each one moved the number here:
- Hardware that idles low, and a sleep policy for the hours nobody is asking.
- A model that fits in fast memory, with nothing spilling.
- An MoE or a right-sized model at 8-bit or 4-bit.
- Prompt caching that survives your agent’s turn structure.
- Thinking pinned off, and switched on by a router only for tasks that earn it.
- Concise outputs and lean tool results.
- A power cap near the knee on any discrete GPU.
Which of these you pull is your call. The order is a starting point, and your own wall meter has the final say.
What to do next
The measured rows behind the claims above are public, and the two tools that produced them are free to run.
- Measure your own machine. The controller’s
bench quickruns this benchmark against the engine you already have in about 25 minutes, no wall meter needed, and writes the joules per correct answer for your box. Opt in and the run joins the leaderboard under CC BY 4.0. The Measure page has the one-line install. - Read the corpus. The leaderboard and model picker for people; the energy data page for every sweep in this post as a chart with a table view. For agents, the same data is JSON, plus a recommendation endpoint that takes a workload shape and an MCP server, all listed in llms.txt.
- Move the work you can move. Anything with a deadline rather than a start time, such as agent runs, indexing, batch inference and fine-tunes, can run in the cheapest electricity hours of your day. The same controller schedules it and reports the measured energy of every run. Interactive use cannot be moved, which is why the first two steps matter more for a personal assistant.
The methodology page says what each number means and where the limits are.
Sources
Benchmark figures in this post come from the energy-bench reference document (findings F1 to F10 and the sweep record, runs through 2026-09-09), measured on RTX 3090 and RTX 2080 nodes at temperature 0 with the GPU’s own energy counter. The methodology page describes the instrument; the rows are on the leaderboard.
- Pope, R. et al. (2022). Efficiently Scaling Transformer Inference. arXiv:2211.05102.
- Kwon, W. et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180.
- Leviathan, Y., Kalman, M., Matias, Y. (2023). Fast Inference from Transformers via Speculative Decoding. arXiv:2211.17192.
- Ainslie, J. et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245.
- Fedus, W., Zoph, B., Shazeer, N. (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv:2101.03961.
- DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434.
- Frantar, E. et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323.
- Lin, J. et al. (2023). AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. arXiv:2306.00978.
- Samsi, S. et al. (2023). From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference. arXiv:2310.03003.
- Chung, J.-W. et al. (2025). The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization. arXiv:2505.06371. Leaderboard at ml.energy.
- Luccioni, S., Jernite, Y., Strubell, E. (2024). Power Hungry Processing: Watts Driving the Cost of AI Deployment? FAccT 2024. arXiv:2311.16863.
- Stanford Hazy Research (2025). Intelligence per Watt. arXiv:2511.07885.
- Elsworth, C. et al. (2025). Measuring the environmental impact of delivering AI at Google Scale. arXiv:2508.15734.
- Uptime Institute (2024). Global Data Center Survey 2024.
- International Energy Agency (2025). Energy and AI.
- Gupta, U. et al. (2021). Chasing Carbon: The Elusive Environmental Footprint of Computing. HPCA 2021. arXiv:2011.02839.
- Apple. Mac Studio power consumption and thermal output.
- Electricity Maps. Live carbon intensity by region.