Why We Built S.A.M.I. Around Multi-Agent Consensus
Single-model coding tools make a silent assumption: one AI is enough. We disagree. This post explains the engineering reasoning that led us to build a consensus-first orchestration system and what we learned along the way.
The SAMI Team
Engineering
Every AI coding assistant available today — Copilot, Cursor, Windsurf, and the rest — routes each prompt to a single model and returns the answer that model produces. This is fast, cheap, and works well for the majority of autocomplete tasks. But it has a structural flaw that becomes visible as soon as you ask something non-trivial.
Single-model tools inherit all the blind spots of their underlying model. Different frontier models miss different edge cases — one may struggle with concurrent code, another over-sanitises outputs in security-sensitive contexts, a third has its own systematic tendencies. No model is wrong in the same way, and that difference matters.
The Case for Disagreement
In software engineering, code review exists precisely because the author of a function is the worst person to find its bugs. They already know what they intended — their brain autocorrects mistakes before they notice them. A second pair of eyes, especially one that does not share the author's mental model, catches what the author missed.
Multi-agent consensus applies this principle to AI: we run the same task through multiple models and force them to critique and verify each other's output before we synthesise a final answer. The value is not blind agreement, but structured disagreement under a shared checkpoint protocol.
This is not an original idea in academia. Ensemble methods in machine learning have existed for decades, and multi-agent debate has been studied as a technique to improve reasoning (see Du et al., "Improving Factuality and Reasoning in Language Models through Multiagent Debate", 2023). What we built is a practical implementation of that idea on top of real production LLM APIs.
How the Orchestrator Works
The S.A.M.I. orchestrator runs server-side, parallelising work across all council members. When a task arrives it goes through five stages:
- Understand & Contract. An interactive clarification loop — if the task is ambiguous, the council asks targeted questions before proceeding. In the same stage the council fixes the binding contract: interfaces, data flow, module boundaries and file layout. All members vote, and the coordinating member records the result.
- Plan. All council members pressure-test the approach and vote on weak spots; the coordinating member synthesises one plan for the complete structure.
- Build. Only one council member writes the binding code. Every other member critiques direction, sketches alternatives in prose, and files change requests — they do not write code at this stage. Right after the council agrees, the project's own build and tests run, if it has any.
- Harden. All council members challenge the build for defects, correctness and security gaps, checking it against sources, repo context, and prior stage artifacts. The validated fixes are applied and the project is run again.
- Acceptance. A final audit and one last run of the project's checks before the result is delivered. A blocking vote sends the result back to the council to resolve the objection.
What We Got Wrong First
Our initial design ran all agents sequentially. The task goes to Model A, the result goes to Model B for review, then to Model C for a security pass, and so on. This was conceptually clean but slow — nothing reached the user until the whole chain had finished.
We moved to parallel work inside each stage and immediately hit a different problem: models producing structurally incompatible outputs. If two models solve the same problem with different class hierarchies, merging their outputs is not trivial. So every stage now produces its result against a declared artifact schema, and the binding contract fixed in the first stage tells the later stages what structure to follow — instead of leaving it to be discovered when outputs are merged.
Consistent Review, Not Selective Shortcuts
A common design temptation is to skip the council on "small" tasks to save latency and cost. S.A.M.I. deliberately does not: every task runs through the full council. Blind spots do not announce themselves in advance — a one-line change can still touch a security boundary or an API contract — so the review that catches them has to be consistent, not conditional.
We do not keep cost down by shortening the council; we keep it visible. The desktop app shows a live cost meter while a run streams, so you see what the review costs as it happens. What you control is the autonomy level and your credits — not whether the work gets reviewed.
Where It Runs
The orchestrator is a server-side, highly concurrent runtime, and it is the production path behind every run — the engine behind every agent run in the S.A.M.I. desktop app, from the Free mini-council up. The desktop app is not publicly released yet; if you want to try the consensus system when it is, join the waitlist on the download page.