Your AI Bill Is Full of Work a Laptop Can Now Do
Token prices keep falling and AI bills keep climbing. A model released on August 14 runs on a good laptop, and it changes which work is worth paying premium rates for.
Your AI bill went up this year while the price of AI went down. Both of those are true at once, and the space between them is where a lot of marketing money is quietly going.
We’ve written some version of this story seven times. Prices fell and budgets didn’t. The people who build the computers started arguing prices might stop falling. Uber burned through its AI budget in four months. Canva cut its growth forecast by a third to pay for its own features. Each time the recommended fix was discipline: watch the meter, set a cap, prove the return. That advice is fine and it isn’t working, because it treats the bill as a spending problem when it’s a buying problem.
Here’s the number that made me rethink it. Alex Atallah, co-founder and CEO of OpenRouter, sat down with Harry Stebbings on August 10. OpenRouter sits between companies and the model vendors, so it sees what everyone is spending. Stebbings asked whether falling prices hurt his business. Atallah walked through what happened to one model on his own platform:
“OpenAI cut prices by 5x and then in coordination with us by another 2x. So in total, the price of Luna has dropped 10x on OpenRouter over the last 2 weeks. And guess how much usage has grown? 13x.”
Ten times cheaper, thirteen times more used. The bill went up. Atallah names the theory behind it, the Jevons paradox, where a resource getting cheaper causes people to consume so much more of it that total spending rises. He’s also honest that the theory is under-tested: “no one has really done a great job modeling it,” he says, before offering his own platform’s data as one clean case.
This is the part that should bother a marketing team. If your AI spend rises when prices fall, then no amount of waiting for the next price cut fixes your budget. You aren’t buying a thing whose cost declines. You’re buying a thing whose cost per unit declines while you buy far more units, and nobody on the team is deciding which units are worth buying.
What changed on August 14
On August 14, Alibaba’s Tongyi Lab released Qwen3.8-27B under an Apache 2.0 license, which means anyone can download it, run it, and use it commercially without asking. Open weights are the model’s parameters published as a file, as opposed to a model you can only reach by calling a vendor’s API and paying per use.
The specification that matters here is the size, more than any benchmark score. At 27 billion parameters, quantized to 4-bit (a compression step that trades a little accuracy for a lot less memory), the model needs somewhere around 17GB to run. That fits comfortably on a Mac with 32GB of unified memory, and on a 24GB machine with very little room to spare. It’s listed in Ollama’s library, which is a free tool that downloads and runs a model locally in about one command.
So the machine your content lead is reading this on can probably run a capable model, for nothing per use, with the data never leaving the building.
We covered the previous version of this in July, when Moonshot’s Kimi K3 put frontier-grade capability into open weights. That story was real but abstract for most teams, because running K3 meant provisioning serious hardware. The distance between “you could self-host this” and “this runs on the laptop in your bag” is the whole difference between an interesting development and a budget decision.
The strongest argument against me
Self-hosting has a long record of costing more than it saves. Standing up a production model service means GPU capacity you pay for whether or not it’s busy, engineers who know how to keep it healthy, and a maintenance job that never ends. At low utilization the real cost per unit can exceed what you’d have paid a vendor outright. Teams talk themselves into the free model and buy an expensive problem.
Atallah is the sharpest version of this counter, and he has an obvious interest in it, since his company exists at the API layer. His view is that the answer to cost chaos is better routing, sending each request to the right model automatically. He thinks it eventually pushes down to individuals, and told Stebbings that choosing models well becomes part of the job:
“Really, your employees all cost totally dynamic different amounts now.”
He may be right about large engineering organizations. The counter doesn’t survive contact with a ten-person marketing team, because it’s answering a question nobody here is asking. Nothing above requires a server, a GPU contract, or an ops hire. It requires a laptop somebody already expensed and a free tool. The failure mode being warned about is running a model service for your product. This is running a model for your own back-office work, which is a different activity with a different bill.
And the tiering idea isn’t hypothetical, because a company your team already pays has done it. When Canva’s AI features cost too much to roll out, it didn’t ration them. It rebuilt them to run up to 30 times cheaper and kept shipping. The vendor solved its own margin problem by moving work to cheaper compute, then charged you the same subscription. You can make the same move.
What to do about it
Sort your AI work into two piles, by whether the output faces a customer or faces you.
- Anything routine, high volume, and internal goes local. Summarizing calls, cleaning list data, drafting first-pass briefs, tagging tickets, reformatting exports. This work is repetitive, tolerant of a good-not-perfect answer, and you do a great deal of it. It’s most of what a marketing team runs through a model in a week.
- Anything touching customer data goes local by default. Not for cost. Because the answer to where the data went becomes nowhere, and that ends a conversation with legal rather than starting one.
- Buy frontier tokens for judgment. Hard strategy work, original writing, and anything where being wrong is expensive. Pay top rates there without flinching, because that’s where the extra capability shows up in the output.
- Put one number in front of the team. Not total AI spend. Spend on work that could have run locally. That’s the figure that tells you whether you have a buying problem, and no dashboard reports it today.
The mistake I’d been making, and I think the industry is making, is arguing about how much AI to use. The more useful question is where in the stack it runs. Model capability has been getting cheaper for three years and marketing keeps spending the savings on the same tier of work. On August 14 the floor dropped out of the cheap tier, and almost nobody in marketing noticed, because it arrived as a model release rather than a pricing announcement.
By June 30, 2027, I expect at least one major martech platform to publicly offer a local or open-weight processing tier and to name it as a way to cut your AI costs, not as a privacy feature. The margin pressure that pushed Canva to rebuild is on every vendor in the category, and the first one to pass the savings on gets to say so in a sales cycle. If that hasn’t happened by then, my read on where this pressure lands was wrong, and the tiering stays a thing individual teams do quietly rather than something they can buy.
Quoted in this story
- Alex Atallah, Co-founder and CEO, OpenRouter (source)
Want your perspective in coverage like this? Get quoted.
Sources
This story is part of our running coverage: the full picture →
Get Net Effect.
The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.


