Software Delivery
Model Independence
Recursive self-improvement needs a better measure
September 29, 2026 - 4 minute read
Software Delivery
Model Independence
September 29, 2026 - 4 minute read
If your team is doing more work with AI, it can be difficult to tell which improvements will last. More generated code, more model calls, and a more ambitious demonstration are visible. Less supervision, fewer repeated mistakes, and useful software reaching users take longer to establish. Recursive self-improvement adds another claim to that conversation, often without making the measure of progress clear.
In his interview with Sabrina Halper, Matan Grinberg separates a system's ability to improve itself from the speed and usefulness of that improvement. He also discusses the move from encouraging AI adoption to evaluating its return. For engineering teams, that means measuring what reaches an accepted result and how much human work it takes to get there.
The interview offers a useful distinction between a capability and its observed effect. A system can contribute to its own improvement without demonstrating rapid, unbounded progress. A team can also spend more on AI without producing more valuable work.
For an engineering team, the immediate task is concrete. Choose a recurring workflow, establish its current result, and inspect what changes after a proposed improvement. That gives you something testable without requiring a prediction about the future pace of AI research.
Grinberg argues that a model's ability to contribute to its own improvement says little about the rate of progress. He allows for improvement that is slow or less effective than work assisted by humans.
The distinction helps avoid combining separate claims. Producing a useful experiment, completing a training run, and demonstrating a better model on independent evaluation are different outcomes. Evidence for one does not automatically establish the others. A repeatable improvement claim needs to identify the changed system and the evaluation used to accept the change.
This also matters when engineering teams use the same phrase for their own tools. Updating repository instructions after a failed task can improve a later run. It does not establish that the underlying model trained itself or changed its weights. A better test environment can make an unchanged model more effective by giving it clearer feedback.
When you assess a proposed improvement, preserve the earlier configuration and the tasks used to evaluate it. Include examples that were not used to shape the change. If the procedure only succeeds on the failure that inspired it, you have evidence of a specific repair, not yet a general improvement.
At 23:30, Grinberg describes the shift from encouraging AI adoption to measuring its return. Getting people to use AI was an early goal. The next question for a business is whether that spending produces useful work.
A workflow that investigates a difficult defect can have a different cost profile from one that summarizes a short document. Set budgets around the work and its quality requirements, with room to investigate why a task needed more attempts.
Factory Router handles model selection with quality, latency, cost, and cache state in view. To evaluate routing for a workflow, track both usage and the work needed to reach an accepted result.
For a practical evaluation, record the full path to an accepted result. Include failed attempts, extra instructions, time spent preparing the environment, and the maintainer's review. A cheaper model request may require more attempts. A larger bill may reflect more useful work, or simply more repetition. Keep those explanations separate before deciding whether to expand the workflow.
At 28:18, Grinberg argues for preserving alternatives to a single model provider. For a software team, that means keeping the workflow usable when an approved provider, model, or endpoint changes.
Factory's Custom Models documentation describes provider keys, compatible endpoints, and locally hosted models in the Droid CLI and desktop app. Custom configurations are not available in the hosted web or mobile products. Compatibility with an API format also does not establish equivalent behavior across models.
For a team with an approved endpoint, evaluate the actual configuration against its tasks and tools. Check whether it follows repository instructions, uses the required checks, and stops when a requirement is unclear. A familiar model name or an available connection should not substitute for that evaluation.
Keep privacy review separate as well. Factory's data-flow documentation distinguishes local file operations from the context sent in model requests. Operating a machine inside your network does not, by itself, establish that inference or connected-tool traffic remains there. Model choice and data-flow control are related decisions with different evidence requirements.
For day-to-day software work, start with a feedback loop you can inspect. Factory's software factory approach combines continuous feedback around software delivery with model independence, sovereign intelligence, and continual learning. Teams can use the record of shipped work and failures to improve how the next task runs.
When operating a software factory on Factory, a team can standardize task inputs and tooling, measure output, and preserve enough context to replay work. A useful improvement might be a clearer task specification, a reusable procedure, or a check that catches a previously missed failure.
Factory's Agent Readiness overview gives teams a way to inspect repository conditions that affect agent work. Use a discovered gap to choose a bounded preparation task. For example, if a test command is undocumented, making it reproducible creates an improvement you can check without changing the model.
Then rerun representative work and inspect the difference. Keep the acceptance criteria stable, record the human intervention required, and retain failures. If the revised instructions cause the agent to skip an important check, a smoother-looking run is not an improvement. The result should support a specific statement about the workflow you changed.
That bounded approach keeps responsibility clear. Engineers choose which changes to the process are acceptable and where they apply. A procedure that helps with maintenance does not automatically justify broader permissions or production access. Expand it when the evidence supports the next use, rather than treating improvement as permission to remove oversight.
Grinberg puts the distinction plainly: “improve themselves is not synonymous with like hyperbolic growth.” On enterprise adoption, he says, “Now, we need to actually measure the ROI”.
Both observations point back to measurement. Track the accepted outcome, the cost of reaching it, and the human intervention required. Those records show whether a change to the workflow is worth keeping.
Start building