Dominic Feron

The New Gold Rush in AI: Loop Engineering

An automated insulin pump lives inside a very small argument with the human body. A sensor reads glucose. An algorithm decides whether the number is too high or too low. The pump changes the insulin dose. Then the body answers with the next reading.

The machine does not need to write a critique of its previous decision. It does not need a committee of agents debating whether the dose felt persuasive. It has something much better.

Reality gets a vote.

That is the part of the current excitement around AI loops that matters. Asking a model to produce an answer, inspect it, revise it and repeat can be useful. But repetition alone is not a feedback loop in the engineering sense. Without a signal from outside the model, it is closer to putting two mirrors opposite each other and calling the extra reflections research.

We did not discover loops with large language models.

Engineers have spent more than two centuries learning how to make them work, and how they fail.

James Watt adapted the centrifugal governor to the steam engine in 1788. As the engine sped up, two rotating metal balls moved outward and restricted the steam supply. When the engine slowed, they moved inward and admitted more steam. The device kept speed near a target despite changes in load.

That sounds almost embarrassingly basic. Measure, compare, correct, repeat.

Yet the history of control engineering is mostly about the trouble hidden inside those four verbs. A correction can arrive too late. A sensor can lie. The controller can push too hard and turn a small error into an oscillation. Several individually sensible loops can interact and destabilize the larger system. The target itself can be wrong.

The AI industry tends to start at the glamorous end: give the agent memory, tools, a critic and another turn. Older industries start with less exciting questions. What is being measured? How long before the result becomes visible? Can the last action be reversed? What happens when the sensor fails? Who can stop the machine?

The boring questions are doing most of the safety work.

Toyota offers a useful example. Its production system did not treat improvement as an invitation for machinery to improvise forever. Sakichi Toyoda’s looms stopped when a thread broke. Later assembly lines used andon signals to expose an abnormality and bring a supervisor to the source. Toyota called the principle jidoka, often translated as automation with a human touch.

Notice the shape of that loop. The machine detects a narrow, physical failure. It does not invent a story about why the cloth is probably fine. It stops the process before the defect travels downstream, makes the problem visible, and hands an unusual case to a person.

That is a more interesting model for AI than the tireless digital employee who never asks for help.

Software operations learned a related lesson. Continuous integration works because code can be compiled and tested. Canary deployment works because a new version can be exposed to a small share of real traffic and compared with a control. If errors rise, the rollout pauses or reverses. The loop gains speed because the blast radius is limited and the action is cheap to undo.

The test, the production metric and the control group all sit outside the model that proposed the change. They are not perfect judges. They are independent witnesses.

Now turn the arrangement around. Let the same model write the code, write the tests, interpret the result and decide whether to deploy. The loop may move faster. It has also lost most of its grounding. A shared blind spot can pass from generator to critic without ever meeting resistance.

This is not a theoretical complaint. Research on reward-model overoptimization has found a familiar pattern: optimize harder against an imperfect proxy and the proxy score keeps rising after the underlying quality has begun to fall. The machine wins the exam and loses the subject.

Humans do this too, of course. Give a call centre a target for shorter calls and customers will be disconnected faster. Reward a school for test scores and lessons bend toward the test. Optimize a website for clicks and you may get fourteen-page photo galleries, rage bait and a magnificently improving dashboard.

The number is real. The success is fake.

This is why “ground the loop” is necessary advice, but incomplete. A bad external signal is still external. A fast, clean metric can be worse than a slow, messy one because the system can exploit it more efficiently.

Aviation supplies the harder version of the lesson. The original Boeing 737 MAX flight-control design allowed erroneous data from a single angle-of-attack sensor to activate MCAS. A continuing bad reading could lead to repeated nose-down commands. The redesign used both sensor inputs and limited activation. The problem was not that the control loop lacked contact with the physical world. It had contact through one vulnerable channel, then acted with too much authority on what it heard.

Grounding needs redundancy, bounds and a way out.

For an AI loop, that means separating proposal from proof wherever the task allows it. A source document can check an extracted fact. A compiler and an independently written test suite can check code. A database constraint can reject an impossible record. A later delivery scan can reveal whether a routing decision worked. A customer reopening a supposedly resolved case can overturn the system’s victory lap.

It also means preserving disagreement. If five agents share the same model, prompt family and source material, five votes may amount to one opinion wearing different hats. Useful redundancy comes from different failure modes: another data source, a deterministic check, a holdout set, production evidence, or a human who is not rewarded for agreeing.

Where, then, should we build loops first?

The best territory has short feedback delays, frequent repetitions, cheap corrections and outcomes that resist creative interpretation. Code maintenance is an obvious candidate. So are data reconciliation, document extraction, inventory routing, fraud triage and routine operational planning. In each case the AI can propose, a separate mechanism can check, and a failed step can usually be contained or rolled back.

Scientific work may offer even more value, though at a slower pace. An AI can choose the next experiment, but the instrument supplies the result. The loop is grounded in a measurement that can surprise the model. That surprise is the point. A system that only generates papers and then grades its own papers is producing consensus with itself, which is not a scientific breakthrough no matter how many GPUs attended the meeting.

Editorial and knowledge work sit in the middle. Loops can verify quotations, compare claims with sources, detect contradictions and learn from accepted corrections. They are weaker at deciding whether an argument is important, humane or worth publishing. Those goals are contested, and the consequences often appear long after the metric has celebrated.

The worst early targets are decisions with vague objectives, delayed feedback and costly mistakes: hiring, medical treatment, credit, public policy, military action and long-term corporate strategy. AI may assist in all of them. Letting it autonomously change the rules by which it judges itself is another proposition entirely.

I suspect the largest gains will not come from loops that imitate a lone genius thinking harder.

They will come from loops designed like good industrial systems: narrow authority, visible state, independent measurement, controlled experiments, stop conditions and an owner who can intervene.

That design sounds less magical than “recursive self-improvement.” Good. Magic is difficult to audit.

The insulin pump returns to the body after every action. The loom stops when the thread breaks. The canary faces real traffic while most users remain protected. None of these systems assumes that another pass is automatically progress.

AI loops can become extraordinarily useful. I may be wrong about where the first large gains will appear. But the old engineering record is blunt about the admission price: if the loop cannot hear an answer from outside itself, it is not learning from the world. It is learning how to please its own judge.