AI Coding Agents
Software Delivery
The multi-agent architecture behind reliable missions
September 29, 2026 - 4 minute read
AI Coding Agents
Software Delivery
September 29, 2026 - 4 minute read
If several coding agents are working on your project, keeping them aligned can become a job of its own. Each may produce a plausible change while leaving you to reconcile conflicting edits, recover context, and decide whether the result works. A useful multi-agent architecture needs to make that coordination visible rather than pass it back to you.
In The Multi-Agent Architecture That Actually Ships, Factory's Luke Alvoeiro describes how Missions separates planning, implementation, and validation. Workers receive bounded assignments, while independent checks test the result against completion conditions agreed before coding begins.
Watch the full AI Engineer talk.
Alvoeiro's central concern is the attention required to supervise extended work. The architecture he describes assigns planning, implementation, and verification to distinct roles, with shared state and explicit handoffs between them.
Follow a requirement from the plan through the worker's assignment to its acceptance check. If it disappears at a handoff, the final change can look complete while missing part of the requested behavior.
At 4:36, Alvoeiro describes orchestrators, workers, and validators. The orchestrator clarifies the goal and produces a plan with features, milestones, and completion conditions. Workers implement assigned features with fresh context. Validators check the result separately.
Factory's current Missions overview describes the corresponding product workflow. You collaborate on the scope and success criteria, approve a plan, and let the orchestration layer carry the project forward. This leaves an important responsibility with the person requesting the work. The system needs an agreed outcome before it can organize execution toward it.
For a first project, make one boundary especially clear. State what the worker may change and what must remain compatible. A feature that adds a new input field might also affect a stored record, an API response, and an older client. An instruction to update the form alone can leave those relationships implicit.
A fresh worker context is useful only if the assignment contains enough information to recover those obligations. Review the specification as if a teammate who missed the planning conversation will implement it. It should identify the relevant interfaces, the acceptance evidence, and the conditions under which the worker should stop and ask for clarification.
Decide who will respond when the project reaches one of those stopping points. A task that needs a policy decision should return to its owner rather than be repeatedly reassigned as an implementation failure. Keep the decision and any resulting scope change with the plan, so the next worker does not have to reconstruct it from a separate conversation.
At 6:34, the talk explains the validation contract. It is written during planning and defines correctness independently of implementation. Alvoeiro's concern is that checks written around an existing solution can end up confirming its choices without testing the original requirement.
Preserve the expected behavior from planning as you add tests during implementation. A check should be able to reject a plausible solution that misses the requirement, including when a later worker takes a different approach.
At 7:21, Alvoeiro describes the user-testing validator launching the application and exercising functional flows. Factory's Missions planning guide makes the preparation concrete: the application needs a reliable way to start and to receive input or be simulated. That setup belongs in the project plan.
For a settings change, save a value, reload the page, and verify that the saved value remains. Then submit invalid input and confirm that existing data is preserved. Use isolated test data so verification does not create unintended changes in a real account.
At 8:23, Alvoeiro explains what workers preserve in a structured handoff. It includes completed and unfinished work, commands and their exit codes, discovered issues, and adherence to the assigned procedure. Milestone checks use that record to identify corrective work.
That is more useful to a reviewer than a completion message alone. A command may have failed because a dependency was missing. A browser check may have reached the wrong application instance. An implementation may be complete while an acceptance condition remains untested. Keeping these outcomes distinct helps the next worker address the actual problem.
The same record makes the limits of a result visible. If a worker could not exercise a payment callback, a later reviewer should not have to infer that gap from a transcript. Record the missing check, why it was unavailable, and what would be required to run it safely.
Treat the handoff as evidence to inspect rather than a substitute for the evidence itself. A reported exit code should correspond to an actual command result. A screenshot should show the state being claimed. When a corrective task changes the implementation, rerun the relevant acceptance checks rather than carrying forward success from an earlier revision.
At 9:31, Alvoeiro describes problems Factory encountered with concurrent implementation, including conflicting changes and inconsistent architectural decisions. In the design presented, features execute serially. Read-only research and review can run in parallel inside a feature or validation step.
Before splitting work, check whether the tasks compete for the same mutable state. Researching separate APIs can proceed independently. Changes to a shared interface need coordination across its callers and tests.
At 11:26, Alvoeiro connects model choice to the role. Planning, implementation, and validation place different demands on a model. That is a reason to evaluate a model in its assigned workflow, rather than assume the strongest result on one activity transfers to every other activity.
The talk later describes orchestration built largely from prompts and skills, with a thin deterministic layer handling operational safeguards. Factory's skills documentation explains how reusable procedures are discovered and loaded when relevant. The practical implication is to keep your procedures easy to inspect and test as models change. Evaluate a revised procedure against the same acceptance conditions before widening its authority.
The validation contract and the orchestration logic serve different purposes. At 6:34, Alvoeiro says the contract “defines correctness independently of implementation”.
At 14:48, Alvoeiro says that “almost all of the orchestration logic is defined in prompts and skills”.
The validation contract keeps the expected behavior stable, while prompts and skills let the orchestration procedure evolve as the team learns from failed runs.
Start building