Something slightly odd is happening with AI at the moment. The models are getting better, and at the same time they are getting slower and more expensive.
That sounds like a complaint. It isn’t. I think there’s an opportunity hiding in it, and I don’t think many of us have spotted it yet.
The Bill Is Going the Wrong Way
Nobody really warned us about this bit. As models got smarter, they started thinking for longer. A question that used to come back in three seconds now takes two minutes, because the thing is reasoning its way through the problem before it answers. And a proper agentic task, the kind that searches and writes and checks its own work, can chew through a million tokens or more before it hands you anything at all.
Frontier reasoning models sit somewhere around $5 to $10 per million input tokens, and output tokens cost three to five times more than input. Run that maths on a big research job and you’re comfortably into $10 or $20 for a single question.
I’ve done exactly this. Fired off a chunky research task, watched it grind away, and then thought: hang on, did I actually need that in the next 10 minutes?
Almost always, no. I needed it by tomorrow.
Charging the Car Overnight
Which got me thinking about electric cars.
Nobody charges an EV at 5pm if they can help it. You plug it in, it sits there overnight when demand is low and power is cheap, and in the morning it’s full. You gave up speed, which you didn’t need, and got a lower bill, which you did.
So why are we running every single AI task like it’s an emergency?
Here’s the version I keep coming back to. You spend the day collecting up all the things you want done. Research, drafting, analysis, the tedious stuff. You don’t run any of it. You just stack it up. Then at 6pm, before you shut the laptop, you hand the whole pile over and go home. It runs overnight. It’s waiting for you in the morning.
I Was Wrong About Why This Works
Now, my first instinct on why this would be cheaper was completely wrong, and it’s worth saying so.
I assumed overnight work would cost less because it’d run on fewer GPUs, quietly chugging along on spare hardware. It turns out it’s almost exactly backwards.
The cost of a token is basically the hourly cost of the GPU divided by how many tokens it managed to produce in that hour. A GPU handling one request at a time is sitting there mostly idle, waiting to load model weights, running at under 5% utilisation. Pack hundreds of requests onto it at once and the same hardware, at the same hourly cost, produces vastly more tokens. Cheaper tokens come from a busier GPU, not a smaller one.
And this isn’t theoretical. Anthropic, OpenAI and Google all run batch APIs, and all three charge exactly half price if you’re willing to accept a 24 hour turnaround. Half. Same model, same quality, no catch other than patience. In practice most batches come back in a few hours anyway.
The electricity angle turns out to be real too, just smaller. Overnight power can run at roughly half the peak daytime rate, and in some markets with tiered pricing the gap between peak and the deepest overnight tier is enormous. It’s a lever, it’s just not the main one.
There’s a third option too. Instead of paying frontier prices for an instant answer, let a cheaper, smaller model think about it for much longer. More reasoning, less money, more time. You’d never accept that trade in a live chat. Overnight, who cares?
Two Kinds of Work
Once you accept all that, the interesting bit isn’t the pricing. It’s what it does to how you organise your day.
Because suddenly there are two types of work. There’s the stuff you need now, where you pay frontier prices and you’re essentially buying speed. And there’s the stuff that just needs to exist by tomorrow morning, where speed is worth nothing to you at all.
I’d bet most of what we currently run at full tilt is secretly the second kind. We run it now because now is the only option the interface offers us.
And if speed stops being the constraint, the question changes. It stops being “how fast can I get this” and becomes “how much value am I actually getting for this spend”. That’s a much better question.
The Queue
Here’s where I start speculating properly.
If cost is the thing you’re optimising for instead of speed, you’d want to pool the work rather than fire it off one job at a time. Collect everything up through the day, sort it by what’s actually worth paying for, and push the winners through overnight.
You wouldn’t get an instant answer back. You’d get a delivery date. Which is much closer to briefing a freelancer than using a search bar, and I suspect it would make people think a lot harder about what they’re actually asking for. Right now the cost of asking is invisible, so we ask for everything.
The Catch
Obviously this only holds up if the models are good enough to be left alone. Right now a lot of agentic work still needs a human poking at it, and a batch job that quietly goes off the rails at 2am has wasted your entire night rather than four minutes of your afternoon.
But that’s a temporary problem. The models get less needy every few months. And our ability to brief them gets better as well.
The permanent shift is the mindset one. We’ve been trained to think faster is better, because for the whole history of computing it was. With this stuff, speed is just one thing you can buy, and quite often it’s the thing you need least.
Sometimes the smartest thing you can do is go home and let it run.

Leave a Reply