Written in October 2026. The opening scene set in 2027 is hypothetical.
Fast-forward to 2027. Models have gotten much stronger. Given a requirement, they can read the repository, change the code, add tests, submit changes, and handle deployment. In the demo video, the programmer steps away for a cup of milk tea and returns to a finished feature.
Put the same model into a project, though, and QA stops the release. The task is to change billing for customer plans. The code implements the ticket, and all the tests pass. But a meeting last week established that existing customers would keep their current benefits. That decision stayed in the meeting and a few messages. It never made it into the ticket or the tests. If nobody remembered it, those customers would be migrated along with everyone else.
At this point, saying “LLMs still haven't solved coding” seems fair enough. Engineers still have to catch omissions and decide whether to release, and someone has to take responsibility when things go wrong. But what did this engineer do? Did they exercise judgment the model doesn't yet have, or supply a fact it never received?
Those two situations call for different improvements. A vague claim that “humans understand the project” can easily end the discussion: code generation has improved, but understanding, verification, and maintenance still need people. In my work on Context Infra, I need to break that work down and see which parts a system could take on next.
01|Code is cheaper. The other costs remain.
For an agreed project scope and maintenance period, we can roughly break down the total cost of software delivery as:
When AI speeds up code generation, implementation is the first cost it reduces. Understanding the goal, checking whether acceptance criteria miss a constraint, and making sure a release won't break existing functionality still take time. Failures still cause losses. Even if the code for that migration were written instantly, someone would still need to check the existing customers' benefits.
I agree with some of the criticism of AI coding: a patch that passes tests is still a long way from a project you can maintain over time. Faster code generation doesn't make those costs disappear. Benchmarks are starting to assess this capability separately. SWE-EVO collects 48 software evolution tasks from seven mature Python projects, asking Agents to make changes across files and multiple steps based on higher-level requirements. The May 2026 version of the paper reports that GPT-5.4 paired with OpenHands completed 25% under that experimental setup. That number can't predict the next generation of models, but local fixes and long-term software evolution do need separate evaluations.
What I don't quite buy is the inference that these remaining costs prove some work will always be beyond technology's reach. An engineer spends half a day finding the latest requirements, comparing two documents, and recalling a decision from a meeting. We can call that half-day “understanding the project.” But that work includes judgment and tradeoffs as well as searching and checking versions. Unless we separate those activities, we can't discuss what else could be saved.
People can remain accountable even as the evidence they need, the time they spend, and the way they do the work change. The fact that someone has to remember the meeting today doesn't prove that every future release must depend on one person remembering every agreement.
To judge how much further these costs can fall, we need to look at how a system detects and corrects errors.
02|Put the model back into the workflow
If we reduce programming to “requirements in, code out,” it's easy to judge a model solely by how much it gets right in one shot. Actual development involves repeated attempts: run the program, notice a problem, check the logs, change the implementation, and run it again. One wrong result doesn't necessarily mean the task is beyond reach, and one successful run may not be enough to ship.
This process involves observation, action, and feedback. We can borrow the language of control theory to describe it:
is the true state of the work, is the system's action, and represents external changes. The system receives an observation , which may contain noise or error . In a software project, the true state includes code, deployed versions, and data, as well as current requirements, customer constraints, and other changes in progress. The model receives files, tool outputs, and event records. From those, it must form a state estimate and choose an action toward the goal :
If the model in the opening scene were to migrate the existing customers too, the problem might lie in that state estimate. Its ticket doesn't reflect the latest decision. Following the ticket perfectly would still miss the real goal. The system might also choose the wrong action despite having all the necessary information. An evaluation needs to check both possibilities.
In control theory, observability concerns whether internal state can be recovered from measurements of inputs and outputs. An observer uses those measurements to keep updating the state estimate, which the controller then acts on. An enterprise project can't simply be plugged into a linear control model, but the distinction between “knowing what's happening now” and “deciding what to do about it” is useful. Which facts the system has read, which have become outdated, and what happened after an action all affect the result.
A model's probabilistic nature alone doesn't rule out reliable software, because the system can verify what it generates. Candidate generation can involve randomness; the system can then run, check, and select those candidates. AlphaEvolve combines model-generated programs, automated evaluation, and evolutionary search, using measurable results to choose what to improve next. The prerequisite is an automated evaluation that can meaningfully distinguish between candidates.
The plan migration may not meet that condition. If the tests don't cover the agreement with existing customers, another run won't reveal the omission. Fixing existing problems can also introduce regressions, or the system may treat a lucky success as a lesson learned. We need to account for the errors each round corrects and the new ones it may introduce.
03|The law of closed-loop convergence for software
Let be the expected risk-weighted deviation from the true goal after the system completes round of “execute, verify, correct.” This measures actual deviation. Neither the number of failed tests nor the model's self-reported confidence is a direct substitute. An omission like the agreement with existing customers remains a deviation even when every test passes.
Suppose each round removes an average fraction of the existing deviation while introducing an average amount of new deviation. Regressions caused by fixes, incorrect state updates, and external changes can all contribute to . Then:
The next round's expected deviation is what remains uncorrected from this round plus the newly introduced deviation. This is an expectation, not a guarantee that any particular run will improve. The model assumes and stay approximately constant within the scope being studied. I call the relationship that follows under these conditions the “law of closed-loop convergence for software”:
When the effective correction rate and newly introduced deviation are approximately stable, the expected deviation above the current steady state decays geometrically. The steady-state level depends on the ratio of newly introduced deviation to the effective correction rate.
This “law” is a first-order recurrence model. Unlike Moore's law, which rests on long-term empirical data, it gives us a mathematical conclusion under stated assumptions. Whether a project meets those assumptions requires separate measurement.
The effective correction rate can be decomposed further:
is the fraction of deviation effectively detected, and is the conditional probability of a successful fix once it has been detected. Both use consistent risk weights. This conditional decomposition does not require detection and repair to be independent.
In the plan migration, the system has to know that existing customers' benefits are affected before it can correct the implementation. Better repair capabilities won't solve a problem that never gets detected. Finding a pile of problems without being able to fix them won't get the software shipped either. Models, context, tests, and tools can all affect and . We can't simply assign “detection” to infrastructure and “repair” to the model.
First, find where it settles
If the system reaches a steady state , the next round's expected deviation no longer changes:
Rearranging gives:
When each round removes as much old deviation as it introduces new deviation, the expected deviation stays at this level. Subtracting the steady state from both sides of the recurrence shows how the system approaches it:
Expanding repeatedly gives:
As long as the assumptions hold, the system approaches . When the initial deviation is above the steady state, each additional round reduces the expected deviation by a smaller amount.
Then, work out how long it takes
Let . The equivalent number of rounds needed to halve the deviation above the steady state is:
In practice, we need to round up to a whole number of rounds. If each round takes approximately , the corresponding half-life in time is . It tells us how long the system needs to halve the deviation above its current steady state, accounting for both correction effectiveness and elapsed time.
That metric is more useful than just counting how many tool calls an Agent makes per minute. Skipping checks can make a round faster, but it may lower and raise . A faster loop doesn't necessarily finish the task sooner.
The horizontal line can move too
is only the steady state for the current parameters. Increasing the effective correction rate or reducing the deviation introduced per round lowers that steady state.

Figure 1 is a mathematical illustration, not measured data. All curves start with a deviation of 100. A has a steady state of 20. Increasing the correction rate lowers B's steady state to 10; further reducing newly introduced deviation lowers C's to 2.5. With an acceptance threshold of 8, C meets the requirement in round 6. A and B cannot meet it under these fixed parameters.
With unchanged parameters, more rounds can only approach the existing steady state. Improving observation, verification, repair, or execution mechanisms may change those parameters. As long as can fall, residual deviation can fall too. The current result does not establish a permanent floor.
If we can establish only an inequality:
then the derivation gives an upper bound on deviation. We cannot treat it as a lower bound that can never be crossed. An upper bound, a steady state, and a permanent floor are three different things.
But even if deviation does fall, we still haven't shown that delivery gets cheaper. Every additional round of execution and verification costs something.
04|Does another round cost anything?
Let the project's acceptance threshold be , with initial deviation . Under the equality model above, meeting the threshold in a finite number of rounds requires:
If the steady state is above the acceptance threshold, another hundred attempts won't meet the requirement under these fixed parameters. The system needs to change how it works or bring in a person. When the condition is met, the minimum number of rounds is:
A better loop may finish feasible tasks in fewer rounds. It may also make automated delivery possible for tasks that previously couldn't meet the standard. Whether this saves money depends on what those improvements cost.
Let include initial implementation, environment setup, and allocated infrastructure costs. Let be the cost of execution and verification per round, and use the local linear approximation to estimate the expected loss from residual deviation. Holding the acceptance standard fixed, choose the lowest total cost among all round counts that meet it:
The cost change from one more round is:
is the added expense. The subtracted term is the reduction in expected loss from that round. Once the task meets the standard, there is no economic reason to keep running if the benefit is smaller than the expense, even if the system can still correct errors.

Figure 2 is a parametric illustration, not measured data. Both curves use the same acceptance threshold. Environment 2 makes deviations easier to detect and introduces fewer new deviations per round, while incurring a higher fixed cost. Under these parameters, its optimal total cost is still lower. The horizontal axis is the conditional probability of a successful fix, not years or general intelligence.
A family of curves better describes software cost, since both model capability and the conditions for observation and verification affect it. The calculation also has to include any extra deployment, computing, and maintenance costs from infrastructure. Further reductions in error do not imply that cost will reach zero. Nor do today's remaining costs imply a permanent positive floor.
This simplified model is not enough to handle irreversible operations, tail losses, or hard safety requirements. Those need separate constraints; averages alone won't do. The plan migration also leaves a question unanswered: can the system detect the missing agreement with those customers?
05|More thinking can't supply a missing fact
Add one constraint to the opening scenario. Imagine two worlds. World A requires existing customers to keep their old benefits. World B requires all customers to move to the new benefits. The code, tickets, tests, and every record the model is allowed to query are identical in both worlds. The fact that determines how to migrate has never entered any accessible information channel.
If the two worlds are equally likely and the model chooses to preserve the old benefits with probability , its average accuracy is:
No matter how much internal reasoning it does, the same evidence still corresponds to two opposite correct answers. The model has no information to distinguish them. It can ask a follow-up question, seek a new information source, or hold off on execution. Those actions help because they acquire new evidence or change the decision process.
If the available mechanisms can never detect a certain class of deviation, we cannot assume it has a positive effective correction rate in the earlier model. Rerunning the same tests won't supply the missing customer agreement. Models may get better and better at finding, understanding, and using facts. But the fact that determines the answer still has to be available somewhere.
In the opening project, the decision still exists in the meeting and messages; the ticket and tests simply haven't caught up. People on projects often fill these gaps: pointing out an outdated document, flagging a customer exception, or explaining why a change still cannot ship. Information that is entirely missing needs a different response from information already recorded but scattered across applications and time. In the latter case, we can try to build a system that retrieves those facts and checks whether they still apply.
Once the system finds the information, it still has to use it correctly. Fixing this migration only solves the immediate problem. Whether the system repeats fewer mistakes on the next task depends on what it retains from this experience.
06|RSI also has to explain how it knows it has improved
Recursive self-improvement, or RSI, faces the same question. An Agent that changes code in response to errors until the tests pass is already using feedback. But if the model, tools, and methods remain unchanged after the task, it may only have fixed the current patch. It will have to start over when a similar problem appears.
We need to distinguish improvements to the current result, the system doing the work, and the method that produces improvements. The earlier loop with fixed describes correction within a task. If a system updates its strategy, tools, memory, or model from experience, we can write the update as:
Here, is the current system, is experience, and is the update mechanism. If the update mechanism can itself be modified:
the system starts changing how it finds and verifies improvements. That is the recursive step in RSI. Existing research targets different parts of this process.
In Absolute Zero, from 2025, the model generates tasks and a code executor verifies the tasks and answers, providing reinforcement learning feedback. The system helps construct its own training curriculum. It still uses a pretrained model. “Zero Data” refers to dependence on external task data at that stage; it does not mean there is no prior knowledge, nor does it eliminate the execution and verification environment.
Darwin Gödel Machine attempts to modify the Agent's own program and then selects versions through task evaluations. It explores improvements to code editing tools, context management, and working methods. But parts of the paper's open-ended exploration process remain fixed. It does not establish that the entire research and development mechanism can already rewrite itself.
Hyperagents, from March 2026, lets the system modify the meta-level improvement process: the task Agent and the meta-Agent responsible for modifying the system sit in the same editable program. The paper reports improved task performance and transfer of some improvement mechanisms across tasks. Those results remain limited to its research setting. They do not establish unlimited, general self-acceleration.
In terms of the earlier parameters, these studies are trying to give subsequent tasks a higher , a lower , or a lower correction cost. What one task leaves behind starts to shape how the next task is done. That makes verification trickier. If an error affects only the current answer, its damage is bounded. Once it gets written into memory, tools, or an update method, it may repeatedly affect later tasks.
“On the Fragility of Self-Improving Agents” reevaluates two memory-based self-improvement methods. The authors find substantial variation across runs. Self-improvement loops may amplify noise, and task order affects the results. More detailed task criteria and environmental feedback can mitigate some of the degradation, but do not eliminate all the problems.
So when I hear that “the system verifies itself,” I still want to know what it verifies against, who maintains the criteria, and whether it can improve its score by relaxing them. A system can try generating tests, finding counterexamples, and designing experiments. It can also improve those methods, but progress still needs evidence independent of the system's own assessment. Otherwise, better optimization may just help it find scoring loopholes faster.
The plan migration is missing an acceptance criterion; RSI may misremember why an attempt succeeded. Both need feedback that distinguishes real improvement from lucky success and mistaken attribution. Watching the loop run alone cannot tell us which of these it is accumulating.
07|What work should Context Infra take over?
My work on Context Infra focuses on how systems acquire facts and form an understanding of the current state. For the migration in the opening scene, the system at least needs to connect the meeting decision to the plan requirements, identify which customers it applies to, notice that the ticket and tests still follow the old requirements, and expose the conflict before execution.
Putting all the material into context doesn't get us there. An obsolete document and a newly confirmed decision can appear in the same window. The system still has to decide which one to act on. The same requirement may be discussed in messages, revised in documents, acted on through tickets, and implemented in code. After finding those records, the system must preserve their relationships and distinguish suggestions, decisions, and things that actually happened. “The code is fixed,” “it met the acceptance criteria,” and “it's live” can't all be recorded as “done.”
By Context Infra, I mean a system that continuously collects structure and events from software environments, within the limits of what it is authorized and able to observe. It organizes them into a working representation that retains sources, timestamps, permissions, and validity status, then updates that representation based on action outcomes.
This requires observable evidence from interfaces, application structures, and event streams, as well as models to organize information by meaning, identify tasks, and infer state. Direct observations, model inferences, and things that remain unconfirmed need to stay distinct. Opening a document cannot be recorded as approving the proposal inside it.
Models can also help improve context management itself. The September 2026 Context Language Models preprint treats context as an editable file, trains models to maintain it, and studies the corresponding training and inference mechanisms. The model can learn to organize context, so fixed external rules don't always have to control it.
Still, organizing existing material and continuously acquiring new facts across applications are separate problems. The system needs a way to learn that a customer has just changed a requirement, a previous decision has been revoked, or an operation did not actually succeed. Otherwise, it may organize its old material beautifully and still act on an outdated state.
Having these facts doesn't remove the other difficulties of programming. Even with all the necessary information, a model may misunderstand constraints, make design mistakes, or fail to fix the code. Reasoning, planning, and execution still need to improve. Verification and execution environments also need work: business constraints must enter acceptance criteria, runtime results must be observable, and failed checks must provide enough evidence for the next revision. High-risk operations need permission controls, staged rollouts, and rollback. Candidate generation and final acceptance also need an appropriate degree of separation.
The full process includes:
For an experience to help with the next task, a record of “what we did” isn't enough. Its usefulness may depend on the evidence available before the action, the judgment based on that evidence, how the result was verified, and why it was later corrected. Without those details, a system may preserve a complete record of an operation without knowing why it succeeded. It may also turn a mistaken attribution into a rule and keep applying it.
Context Infra can supply facts and state for this process. We still have to evaluate whether a lesson holds and whether an update improves the system. Familiarity with one project also doesn't directly count as improved general capability. I want it to start by reducing how often people have to explain known facts, how long they spend reconstructing the state of their work, and how much rework follows code written on false premises.
08|We still have to run the experiments
We need experiments to verify these gains, and the results may invalidate the recurrence model itself. First, we need to check whether and can remain approximately stable within the scope under study. If remaining errors become harder to detect, fixes are strongly coupled, or goals keep changing, an exponential curve may no longer explain the behavior. Then we should change the model, not force a fit to protect the “law.”
To evaluate Context Infra, hold the model, task scope, and acceptance standard fixed. Compare humans supplying background information, Agents retrieving it themselves, ordinary document indexes, and systems that continuously maintain state. Beyond task completion rates, record how often people supply missing information, when omitted constraints are discovered, the number of regressions and the amount of rework, and changes in delivery risk. Include all acquisition, computing, deployment, and maintenance costs in the total.
The experiments also need multiple runs, shuffled task order, and evaluation tasks held out from the improvement process. Research on self-improvement reminds us that a single run and a fixed task order can hide important instabilities.
We can start with the plan migration. Can the system find the decision to preserve existing benefits before execution, confirm that it still applies, and incorporate it into acceptance criteria? After this task is complete, can the next task use the relevant state correctly?
In my work on Context Infra, I need to know how much less time engineers spend explaining background, checking state, and handling rework under the same delivery requirements. Once we include the system's own costs, how much does the total cost fall?
References
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
https://arxiv.org/abs/2512.18470 - Feedback Systems — Chapter 8: Output Feedback
https://fbswiki.org/wiki/index.php/Output_Feedback - AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ - Absolute Zero: Reinforced Self-play Reasoning with Zero Data
https://arxiv.org/abs/2505.03335 - Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
https://arxiv.org/html/2505.22954v3 - Hyperagents
https://arxiv.org/abs/2603.19461 - On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
https://arxiv.org/abs/2608.18066 - Context Language Models
https://arxiv.org/abs/2609.37725
