This is a resource that any CTO should read. Why?
For almost four years, much of the AI conversation has centered on how intelligence gets built. Pretraining gave models broad capabilities by learning from enormous datasets. Post-training shaped those capabilities into more useful behavior. Reinforcement learning from human feedback, or RLHF, helped turn language models into assistants that could follow instructions and interact with people naturally. These were not separate technologies replacing one another, but successive layers of improvement around the model.
We have now entered what I would call the inference phase of the AI Supercycle. Not because training has stopped mattering, but because the economic question is expanding from How do we build a capable model? to How do we operate that capability repeatedly, across millions of tasks, at a cost the customer can justify? Reasoning models make this transition particularly important: additional computation during inference can itself improve performance. The model’s capabilities depend not only on how it was trained, but also on how much work the system performs when answering.
My expectation is that over the next three to five years, production inference could become a larger spending pool than pretraining alone, with enterprise deployment an important driver. Evidence already points in that direction within a defined market: Gartner forecasts inference to account for 55% of AI-optimized cloud infrastructure spending in 2026 and 59% in 2027, exceeding training spending in that segment. This isn't an industry-wide accounting of every AI investment, but it shows why inference deserves to move from a technical consideration to the center of enterprise strategy (see Gartner).
The mechanism is straightforward. Training costs recur as models are developed and improved. Inference costs recur whenever those models are used. And an agentic workflow can require much more than a single answer. Imagine a system investigating a customer issue: it reads the account history, checks a contract, chooses a tool, processes the result, revises its plan, and prepares an action. Each model call adds work. Multiply that pattern across departments, customers, and continuous operations, and serving intelligence becomes a substantial production system.
But inference is not one uniform workload, and that is where the nuances become fundamental. In a typical autoregressive language model, two phases do different jobs:
Prefill processes the input. The model works through the prompt, retrieved documents, instructions, and other context, building intermediate attention state commonly held in a key-value, or KV, cache. Much of this processing can happen in parallel. For sufficiently long inputs, available computation is often the main constraint, and prefill contributes substantially to the wait before generation begins.
Decode generates the output. The model produces subsequent tokens sequentially, using the context and previously generated tokens. Especially at smaller batch sizes, performance is often constrained by how quickly weights and cached state can move through memory rather than by arithmetic capacity alone. This affects how quickly the response progresses after generation begins.
That distinction has direct business consequences. A system reading a large contract pack to return a short classification has a different workload profile from one generating a long software change from a short instruction. An agent repeatedly using the same context introduces another consideration: whether the serving system can reuse previous computation instead of paying to process the same material again. The same model can have very different economics depending on what the application asks it to read, generate, and repeat.
This is why serving architecture matters. Prefill and decode can sometimes benefit from separate resource pools, different scheduling, and different capacity allocations. But separation isn't automatically better: transferring cached state adds overhead, and short or heavily cached requests may be more efficient without it. The right architecture follows the workload, not a universal rule about which chip, cloud, or model is best.
For an enterprise, this turns apparently technical choices into operating decisions. How much context should each request receive? Which repeated inputs can be cached? Which tasks can wait for batch processing? Which require immediate responses? Where does a more capable model reduce retries and human review enough to justify its price? The objective is not simply cheaper tokens. It is lower cost per accepted outcome, with quality, latency, and reliability held to the required standard.
That is also why this phase requires a map of the whole stack. Inference depends on the physical infrastructure underneath the model, but useful enterprise work depends equally on the context, routing, tools, permissions, memory, and evaluation around it. The company must decide not only where intelligence runs, but where the knowledge and authority that make it useful accumulate. Those are the decisions that determine whether it can change providers without rebuilding its operating capability.
The first phase made us ask how powerful the model could become. The inference phase forces us to understand the system that turns that capability into recurring work. This piece maps that system, from power and silicon to serving, agents, and enterprise control, and asks which parts should be rented and which must remain under the company’s control.
Rent the layers. Own the seams. Price the exit.
Preview the book & course module here!
For the last three years, I’ve been rebuilding the Business Engineer’s curriculum from the ground up. That curriculum has now become the foundation of a new discipline, with the entire series taking shape around it.
If you’re already a paid member, simply reply to this email, and we’ll send it your way.




