The Business Engineer

The Business Engineer

Inference Engineering: The Tokenomics of AI

Gennaro Cuofano's avatar
Gennaro Cuofano
Oct 03, 2026
∙ Paid

For almost four years, the AI Supercycle was dominated by one question: how do we build more capable models?

First came pretraining. Scale the Transformer with more data, more parameters, and more compute, and increasingly powerful capabilities emerged. Then came post-training. Instruction tuning, RLHF, reinforcement learning, verifiable rewards, and eventually reasoning techniques turned those pretrained models into systems that could follow instructions, solve harder problems, and interact with people in useful ways.

That was the first phase of the cycle. Most of the industry’s attention, capital, and technical prestige went into creating intelligence.

We are now moving into a different phase: the inference phase.

Training does not become less important. Frontier models will continue to require enormous amounts of compute. But training happens in large, concentrated runs. Inference happens every time the model is used. Every question, search, document review, coding task, reasoning step, tool call, and agent action requires the machine to run again. Google reported processing more than 3.2 quadrillion tokens per month as read in 2026, around seven times the level of a year earlier. At the same time, reasoning models are producing longer outputs, context windows have expanded dramatically, and agents can invoke models dozens of times to complete a single piece of work.

My expectation is that over the next three to five years, enterprise inference becomes at least as important economically as the pretraining layer, and potentially much larger. The reason is straightforward: once intelligence becomes useful enough, the economic problem shifts from producing the model to operating it continuously across millions of workflows.

And this is where something important changes.

For years, inference looked like plumbing. You picked a model, called an API, paid for the tokens, and focused on the application above it. But at enterprise scale, that abstraction starts to break down. The price of the model is only one part of the cost. The economics increasingly depend on how the request is served.

A model request actually contains two very different jobs.

Prefill is reading. The model processes the prompt, the instructions, the documents, the conversation history, and everything else in the context. Those input tokens can largely be processed in parallel, which makes prefill relatively compute-intensive and strongly influences how long the user waits before the first token appears.

Decode is writing. Once the model starts answering, it generates one token after another because each new token depends on what came before. That makes generation far more sensitive to memory bandwidth and to how many conversations share the hardware at the same time.

That distinction sounds technical, but it determines the economics of enterprise AI.

A system that reads a 100-page contract and returns one classification has a very different cost structure from a coding agent generating thousands of tokens from a short instruction. An agent that repeatedly rereads 50,000 tokens of context over thirty steps has a different profile again. The same model, at the same token price, can produce radically different costs depending on what it is asked to read, generate, repeat, and remember.

This is why inference engineering is becoming a discipline of its own.

It is about batching requests so expensive hardware stays productive without making users wait too long. It is about managing the KV cache, the working memory of each conversation, so long contexts do not consume the entire machine. It is about reusing prefixes instead of repeatedly paying to read the same instructions and documents. It is about deciding which work needs real-time capacity, which can run in batch, when dedicated infrastructure makes sense, and when renting tokens remains cheaper.

Above all, it is about changing the unit of optimization.

The objective is not the cheapest token.

It is the lowest cost per accepted outcome at the latency the workload actually requires.

That becomes even more important as reasoning and agents spread. Reasoning models can consume more computation to solve harder problems. Agents multiply inference calls and repeatedly reread growing contexts. The cost per token may keep collapsing while the number of tokens required to perform useful work rises much faster. That is how intelligence can become dramatically cheaper while the total inference market becomes dramatically larger.

So this piece goes beneath the API and into the machinery that increasingly determines the real economics of AI: prefill and decode, working memory, batching, caching, serving modes, latency, utilization, hardware, and the software engines connecting them all.

The concepts are technical, but the logic is surprisingly simple.

Serve to the target. Keep the cache warm. Pay for busy, not idle.

That is the operating system of the inference phase.


Preview the book & course module here!

Subscribe To Premium To Gain Access!


For the last three years, I’ve been rebuilding the Business Engineer’s curriculum from the ground up. That curriculum has now become the foundation of a new discipline, with the entire series taking shape around it.

Subscribe To Premium To Gain Access!

If you’re already a paid member, simply reply to this email, and we’ll send it your way.

Preview the book & course module here!

Subscribe To Premium To Gain Access!

User's avatar

Continue reading this post for free, courtesy of Gennaro Cuofano.

Or purchase a paid subscription.
© 2026 Gennaro Cuofano · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture