Home Data-Driven Thinking Marketers, The Free Token Lunch Is Over. Is Your Stack Ready?

Marketers, The Free Token Lunch Is Over. Is Your Stack Ready?

SHARE:
Jon Morra, chief AI officer, Zefr

Anthropic just dropped Claude Sonnet 5, reigniting a broader debate about how AI should be priced. Some argue that published token rates no longer tell the full economic story because changes in token consumption and model behavior, such as the use of reasoning tokens, can materially affect real-world costs.

Meanwhile, other alternatives like Fable 5 are priced at twice Anthropic’s prior top-tier model and no longer treated as unlimited plan features. Regardless of where you land in that debate, it’s becoming clear that the economics of AI are changing. Not to mention the actual, ever-changing pricing structures.

Most companies running AI-powered marketing tools are not ready for what this means. They built their systems assuming both tokens would stay cheap forever and the number of tokens needed to achieve an outcome would stay flat.

For years, AI builders and buyers have been rewarded for a pretty simple strategy: use the biggest model you can get your hands on. That worked while the cost of the hardware needed to run the models stayed within budget. It was easy to assume that more intelligence was always the right answer, but when you have to consider the cost of that intelligence, that is not always the correct, economical choice.

Now marketing needs to pay attention because this changes the conversation and the questions that marketers should be asking about their AI tools. 

Matching model to task

Think about the kinds of decisions AI is making across advertising today. Is this video suitable for a family-friendly campaign? Does this creator align with a brand’s values? Is this creative likely to violate a platform policy? Important questions, but they aren’t the same as asking an AI model to solve a complex legal problem or prove cutting-edge math problems.

Ramp, which tracks AI spending across tens of thousands of businesses, found that although the average cost per million tokens has fallen dramatically over the past year, AI spending continues to climb for two reasons: breadth and depth. On the breadth side, companies are using AI to solve far more problems than they were a year ago.

On the depth side, most modern AI models now ship with reasoning enabled by default. Reasoning uses many times more tokens to “think” about answers to problems before generating the actual output tokens that respond to the prompt. Cheaper tokens didn’t lower AI budgets; they encouraged organizations to consume many more of them, a classic case of Jevons paradox.

We are shifting back to unit economics, and marketers need to start looking at right-sizing models to the problems they actually need to solve.

Researchers are also beginning to evaluate AI in a different way. Instead of asking how much a token costs, they’re asking how much it costs to get the right answer. Researchers recently proposed a framework they call “cost-of-pass,” which measures the expected cost of producing a correct result. They concluded that different kinds of models are economically optimal for different kinds of work. Lightweight models excel on simpler tasks, larger models on knowledge-intensive ones and reasoning models earn their keep only when the problem is complex enough to justify the additional cost. My colleagues and I also studied this concept under the brand-safety paradigm, asking when to trust model outputs or escalate decisions to more expensive human review.

That’s exactly where I think enterprise AI is heading. We’re relearning that solving billions of repetitive decisions efficiently is a very different engineering challenge from solving one extraordinarily difficult problem. For engineers like myself, it can be challenging to overcome our shiny object system. But when a technology needs to work at scale, we must.

Per-token price for frontier models may go up or down, but the cost to solve complex problems is going to increase. Anthropic won’t be the last lab to make that move. The question every marketer should be asking their AI vendors this quarter isn’t “Which model do you use?” It’s “Why does every task get the same model?” If the answer is habit rather than architecture, that’s where the next invoice surprise is coming from.

“Data-Driven Thinking” is written by members of the media community and contains fresh ideas on the digital revolution in media.

Follow Zefr and AdExchanger on LinkedIn.

Must Read

Apple Has Far-Reaching Plans To Block Hundreds Of Programmatic Data Companies From iOS

Apple’s WebKit crackdown appears to extend well beyond The Trade Desk, putting hundreds of ad tech, data and identity vendors on a mysterious, dynamically updated block list.

Josh Reed, Zoom's VP of brand and content, speaking at AdExchanger's Programmatic IO event in New York City (September 28, 2006)

Zoom’s Marketing Challenge Is That It’s Too Well Known For Its Own Good

Zoom has 99% unaided brand awareness, which sounds great on paper. But there’s a catch: Most people still think it’s just a video-call app.

Why Agencies Think They Shouldn’t Own Agentic AI Tools Or The Data Used To Build Them

Agencies are differentiating their tech stacks by building custom agentic AI tools for their clients. And they’re rethinking owning those AI tools – particularly since licensing them creates new revenue streams.

Privacy! Commerce! Connected TV! Read all about it. Subscribe to AdExchanger Newsletters

Programmatic IO: Insurers Are Building Ad Tech’s AI Accountability Layer

Agencies and marketers discussed the future of AI governance at AdExchanger’s Programmatic IO NYC this week. The main takeaway? Expect insurers to play an increasingly important role in managing AI compliance.

Apple’s Latest Operating System Blocks The Trade Desk From Serving Ads On Safari

The Trade Desk is unable to serve ads to the Safari browser for Apple device owners that have downloaded iOS 27. Apple has been investigating the issue since last week.

Who Will Stand Up For The Open Web?

The open web is done, stick a fork in it. Banner blindness is near universal, search traffic has run dry and publishers are struggling for oxygen. But what if that’s … not true?