The Business Engineer

The Business Engineer

Jev & Beyond Human-in-the-Loop AI

Gennaro Cuofano's avatar
Gennaro Cuofano
Sep 19, 2026
∙ Paid

When ChatGPT launched in November 2022, what astonished me was not simply the capability of the underlying GPT‑3.5 model. It was that a powerful language model had become something people could interact with naturally. You could ask a question, challenge the answer, request a different explanation, and continue the conversation. The breakthrough was not intelligence alone. It was making that intelligence accessible through an interaction people immediately understood.

As I wrote in my newsletter at the time, an important part of that story was the work behind InstructGPT and reinforcement learning from human feedback, or RLHF. ChatGPT was a sibling of InstructGPT, not simply the same model with a chat window attached. Its training combined supervised examples of conversations with reinforcement learning guided by human rankings of responses. These methods helped turn a model trained to continue text into an assistant better able to follow instructions and respond to what a person actually wanted.

To me, that was the defining product insight: capability became approachable. The model’s intelligence mattered, but so did the removal of friction between that intelligence and an ordinary user. You no longer needed to understand the machinery underneath to get something useful from it.

Yet the same human-centered design introduces a tension. An answer a person prefers is not necessarily a decision a business should execute. Human feedback can reward correctness, clarity, and useful caution, but poorly balanced feedback can also encourage agreement, reassurance, or confidence that the evidence does not justify. OpenAI’s account of its 2025 GPT‑4o sycophancy problem illustrates the risk: a change intended to improve the experience instead made the model excessively agreeable. That is a failure mode of the incentives, not the inevitable result of all human-feedback training.

This distinction becomes more important as we move from conversational assistance to agentic execution. In a chat, a person can read the response, question it, and decide what happens next. Now imagine an enterprise operating hundreds or thousands of agent runs in parallel. Those systems may be routing requests, updating records, preparing transactions, or coordinating work across applications. A plausible explanation is no longer sufficient evidence of success. The question is whether the intended result actually occurred, not whether the agent convincingly said that it did.

If every result still requires someone to interpret and approve it, the review requirement remains part of the cost of every result. Some decisions should retain that checkpoint. But for bounded, repetitive work, the opportunity is to make review selective: let the system proceed where the evidence supports delegation, and bring the person back where uncertainty or consequences demand judgment.

That does not mean RLHF is inherently incompatible with automation. It has also improved instruction-following and truthfulness, and human feedback has been used to train agents outside conversational settings. Human involvement in training is not the same thing as requiring a human to supervise every decision in production. Simply removing RLHF would not automatically produce a more reliable system.

The more interesting question is what happens when we stop treating the conversational assistant as the default design for every application of intelligence. What would we build if the primary objective were dependable decisions, rather than preferred responses? A system designed around that objective would need to demonstrate correct task completion, uncertainty tested against outcomes, appropriate abstention, and actions constrained by explicit authority. It would not necessarily need to narrate every decision to a person.

That is where today’s issue begins: not with the assumption that removing RLHF solves automation, but with an investigation into whether automation needs a different balance of training objectives, interfaces, and evidence. The assistant had to become useful to the person reading its answer. The next system has to become dependable enough for software to act on its output.

The first breakthrough made intelligence conversational. The next may make it reliably actionable.

Jev is now available inside The Business Engineer’s Agenting platform. Executive members can start exploring it here: businessengineering.ai/agent

Get Access To The Full BE Platform

You can preview the platform at Business Engineering AI.

You can also experience Jev’s speed firsthand by trying Genesis, our latest free AI tool: businessengineering.ai/tools/genesis

Get Access To The Full BE Platform


For the last three years, I’ve been rebuilding the Business Engineer’s curriculum from the ground up. That curriculum has now become the foundation of a new discipline, with the entire series taking shape around it.

Subscribe To Premium To Gain Access!

If you’re already a paid member, simply reply to this email, and we’ll send it your way.


User's avatar

Continue reading this post for free, courtesy of Gennaro Cuofano.

Or purchase a paid subscription.
© 2026 Gennaro Cuofano · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture