Ian O’Byrne
Writing / Overstory

How Much Energy Does an AI Question Use? It Depends on What You Count

Why an AI query has no universal energy cost—and what local models can teach us about the infrastructure hidden behind every answer.

Posted
Aug 31, 2026
Author
Ian O’Byrne
Read
10 min
Topics
ai

Whenever I’m writing or teaching about artificial intelligence (AI), one question always comes up. How much energy is being used? It ties into the recent pushback about data centers, but it also opens up something more valuable: a chance to practice critical thinking and human agency as we weigh the effects of these technologies on our planet.

This summer, that question was front and center during a week-long professional development session on integrating AI into K–8 education. I was focusing on local, private AI by having teachers build and use models on a Raspberry Pi. We compared the output from that small computer to local AI running on the high-priced desktops in our lab, and to what we could do with the phones in our pockets. We talked about the AI hallucinating and getting things wrong. We talked about how fast it responded. We unpacked the size of the files and the parameters each model used.

Just before lunch, a participant sent me down a rabbit hole with a simple question: How much energy does it take to ask an AI a question?

The question landed because, in the workshop, participants watched power and battery consumption directly. They examined what the CPU and GPU were doing and asked where all of that “work” happens when they use an online AI model. Running a model locally, they could see the power cable going into the computer, hear the fans kick on, and feel the heat coming off the sides. With an online, hosted service, the answer arrived through a browser tab and left no similar evidence in the room.

As we broke for lunch and I thought about it that night, I realized I’d answered badly.

Two years ago, I wrote about running AI models on my own machine and argued that cloud computing doesn’t make the material cost disappear. It moves that cost somewhere the user can’t see. In my own work and outreach, I’ve been exploring more use and development of local, private AI. In this post, I want to outline what “cost” actually means in this conversation, and the decisions we face as we use (or don’t use) AI in our work.

Here’s the topline for my research. The trouble isn’t that we can’t agree on a number. It’s that the word query hides most of what we’d need to know to produce one. To get anywhere, it helps to think about AI the way we think about a car trip, and to keep that comparison in mind all the way through.

Not all queries are the same

When I started thinking about how much energy an AI chat uses, I started with the individual query. A query is a question or request for information. It sounds like a unit of measurement, but it isn’t.

A watt-hour is a unit. A joule is a unit. A query is simply something we ask a computer to do. In my earlier work, a query usually meant typing something into Google. Here, I’m using it to mean a request sent to an AI model.

The problem is that one query can be wildly different from another.

Think about saying a car uses half a gallon of gas “per trip.” That number tells us almost nothing unless we know whether the trip was two miles or two hundred. Both count as one trip. AI queries work the same way.

“What is the capital of France?” is one query. Asking a model to read a long document, compare several arguments, and write a detailed analysis is also one query. The second request may demand vastly more computation and energy, even if they both get counted as one.

Tokens give us a slightly better way to describe the distance traveled. Models break text into small pieces called tokens as they process a prompt and generate a response. A short question and answer might involve a few hundred tokens. A long document and detailed response might involve tens of thousands.

But tokens aren’t units of energy either. Processing 1,000 tokens with a small model isn’t equivalent to processing 1,000 tokens with a much larger one. The hardware matters. So does the length of the context, the way the model is configured, and whether the machine is serving one request or many at the same time.

So here’s the frame I keep coming back to, and it’s the one I’ll use for the rest of this post:

The query tells us how many trips we took. The tokens tell us roughly how far we traveled. The model and hardware tell us what we were driving.

One AI query separates into the number of trips, the token distance, and the model and hardware used as the vehicle.

Figure 1. A query counts the trip, not the energy. Tokens, model, hardware, and configuration describe the work more clearly.

This is why there’s no universal energy cost for an AI query. The disagreement isn’t just that researchers can’t settle on a figure. The word query collapses the trip, the distance, and the vehicle into one word.

It helps to separate power from energy, too. Watts describe how quickly a device uses electricity. Watt-hours describe how much it uses over time. A powerful computer can draw more watts but still use less total energy if it finishes the task quickly. A smaller machine running much longer may consume more in total. Fast and thirsty isn’t the same as slow and frugal. Neither maps cleanly onto “big” or “small.”

That’s only part of the story.

The range of the trip

Once I stopped treating a query as a standard unit, the problem made a bit more sense. It also needed more context.

A Google study of its production Gemini service reported 0.24 watt-hours for the median text prompt when the researchers counted the serving system broadly. This included the accelerators, host computers and memory, provisioned idle capacity, and data-center overhead. In plain English, this meant the AI chips doing the work, the computers supporting them, the machines kept ready for the next request, and the electricity needed to cool and run the building.

But even that number varies enormously, because the amount of work behind a prompt varies enormously.

The ML.ENERGY benchmark shows how strongly energy use tracks the workload. Long answers generally require more computation than short ones. Larger models usually require more work than smaller ones. And the same model can consume different amounts of energy depending on the hardware running it and how efficiently it’s used.

There are less visible differences too. Models can run at different levels of numerical precision, trading some computational detail for speed and efficiency. And in a data center, a single powerful AI chip may process requests from many users at once, spreading its energy across all of them.

So asking a small model for a short answer is a fundamentally different task from asking a large reasoning model to work through a complex problem and generate several thousand words. Researchers aren’t necessarily measuring the same trip.

To make this concrete, the next step is to measure what a single machine actually draws at the outlet. Software can report what the GPU is using, but that leaves out the CPU, memory, storage, power-supply losses, and the rest of the system. A useful measurement needs an idle baseline, the power draw while the model is processing a prompt and generating a response, and the time spent doing that work. An hour-long conversation does not mean an hour of continuous inference. The machine works harder while the model responds, then largely idles while I read and type. Until I measure that complete path, attaching a per-query number to my desktop would create the same false precision I’m arguing against.

Once again, that is one vehicle making one kind of trip. When I use a hosted frontier model such as Claude, ChatGPT, or Gemini, I can still see the electricity my own computer uses. What I cannot see is the energy consumed in the data center. That raises another problem: which parts of that larger system belong in the count?

What questions are we actually asking?

Even if we agreed on a consistent unit of measurement, a second problem would still be waiting: where we draw the boundary. When someone says an AI prompt “used” a certain amount of energy, they may be answering at least three different questions.

The first is marginal: How much additional electricity flowed because I submitted this request? The second is allocated: What share of the entire serving system should be assigned to an average request, including machines held ready for traffic that may never arrive? The third is amortized: Should some portion of the energy used to train the model earlier be divided across every request it will ever serve?

All three are legitimate accounting questions. They are not interchangeable.

Three nested accounting boundaries show marginal, allocated, and amortized energy estimates.

Figure 2. The answer changes as the accounting boundary widens from the additional work caused by a request to the serving system and model training.

Training shows the problem clearly. The more requests you divide its energy cost across, the smaller the cost assigned to each one. The electricity used to train the model hasn’t changed—only the way we divide it. Calling that number “the energy my prompt used” makes an estimate sound like a direct measurement.

For an individual user, the most useful question is how much extra work their request caused. Which model ran, how much material it processed, how long the response was, and whether they asked it to try again. At a larger scale, millions of requests can create demand for more machines and data centers. All of these points matter.

Owning the truck or taking the Uber

I love using analogies to make a point and hopefully it helps to think of using AI like we think about transportation. I think this captures the local-versus-cloud choice better than any single number.

Running an AI model on hardware you own is like owning a truck. You get independence, privacy, and no per-trip tolls. You also buy the vehicle and pay for its fuel and maintenance. In my office, those costs show up as hardware, electricity, and heat. Every cost is mine, and every cost is visible. Using a hosted model is like taking an Uber. You skip the maintenance and the fixed costs, and you scale instantly. The challenge is that you pay a recurring fare, you depend on the provider, and someone else is burning the fuel on your behalf, out of your sight.

A local model exposes its hardware, electricity, heat, and meter while a hosted model moves the provider's meter out of sight.

Figure 3. Local and hosted models both use material infrastructure. What changes is who owns it and who can see the meter.

When I use a local AI model on my machine, I can hear the work happening. The fans spin up, the room gets warmer, my electric bill goes up and words appear on the screen. When I ask a hosted model the same kind of question, my computer barely notices.

The difference is not that my local machine uses energy and the data center does not. The difference is that I can see one of them working. A hosted service may process my request more efficiently by sharing its hardware across many users. It may also send my request to a much larger model that uses more energy. From the browser, I cannot tell. I see the answer, not the meter.

That brings me back to the question I was asked before lunch. I wanted to give that teacher a number. A more honest answer is that a number means very little until we know what kind of request was made, which model handled it, what hardware did the work, and which parts of the system were counted.

This uncertainty does not make the question useless. It gives us better questions to ask. Providers can be clearer about what they measure and what their estimates include. As users, we can consider whether a smaller model is enough, whether we need another response, or whether the task requires AI at all. As educators, we can help learners recognize that a clean browser window still connects to physical infrastructure somewhere else.

I still cannot say exactly how much energy that teacher’s question used. No universal per-query number can provide that. What I can do is look at the trip, the distance, the vehicle, and who had access to the fuel gauge. The cost does not disappear when it leaves the room. Our ability to see it does.