When your developer is an AI agent, who watches the gates?
Agentic AI can write code faster than any human team. It can produce 50 files in an hour, refactor entire architectures in a session, and generate test suites that hit 95% coverage before lunch.
But speed without governance is just faster failure.
In healthcare software — where a bug isn't a broken button but a misdirected medical record — the question isn't "can AI build it?" It's "should this ship?"
We took Robert Cooper's Stage-Gate® framework (1986), stripped it to its core principle — evidence, not confidence — and rebuilt it for a world where AI agents do the building and humans do the deciding.
┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐
│Discovery│───▶│ Scoping │───▶│Business │───▶│Develop- │───▶│ Launch │
│ │ │ │ │ Case │ │ ment │ │ │
└────┬────┘ └────┬────┘ └────┬────┘ └────┬────┘ └────┬────┘
│ │ │ │ │
Gate 1 Gate 2 Gate 3 Gate 4 Gate 5
"Is it "Is it "Is the "Is it "Is it
real?" worth it?" plan solid?" working?" ready?"
Cooper's insight: don't let bad projects consume resources. Kill early, kill cheap. Each gate is a go/kill/hold decision made by senior management with evidence, not enthusiasm.
80% of product development organizations use some form of Stage-Gate (PDMA Best Practices Study). It works. But it was designed for physical products with 18-month development cycles and human-only teams.
| Classic Assumption | New Reality |
|---|---|
| Development takes months | AI produces working code in hours |
| The bottleneck is building | The bottleneck is deciding what to build |
| Code review catches defects | AI generates plausible-looking code that passes superficial review |
| One team builds, another reviews | The same AI can build AND review (confirmation bias risk) |
| Security is a phase | AI doesn't instinctively think about attack surfaces |
| "Ship it" pressure comes from deadlines | "Ship it" pressure comes from how easy it is to produce |
The danger isn't that AI builds bad software. It's that AI builds software so fast that governance can't keep up.
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ G0 │──▶│ G1 │──▶│G2/G3 │──▶│ G4 │──▶│ G5 │──▶│LAUNCH│──▶│ G6 │
│Worth │ │Reqs │ │Design│ │Build │ │Ready │ │ │ │Learn │
│doing?│ │solid?│ │safe? │ │works?│ │ship? │ │ │ │ │
└──┬───┘ └──┬───┘ └──┬───┘ └──┬───┘ └──┬───┘ └──────┘ └──┬───┘
│ │ │ │ │ │
▼ ▼ ▼ ▼ ▼ ▼
Problem Evidence Adversarial Proof Patient Reality
is real survives review of it safety vs.
attack architecture works verified plan
1. Adversarial B-Team + Zero-Context Reviewer at every gate
We don't let the builder review their own work. A separate B-Team AI (different model, different prompt, hostile persona) attacks the spec at every gate. But teams immersed in months of development have systematic blind spots. We can't see what we built wrong because we were thinking about it the same way the whole time.
2. Evidence is code, not slides
At G4, "it works" means: test coverage report, security review sign-off, accessibility audit, integration test against real backend. Not a demo. Not a deck. Artifacts.
3. 100% test coverage is a gate requirement, not a goal
When AI writes the code, AI can write the tests. There's no excuse for untested code when your developer never gets tired.
4. Security is every-gate, not a phase
G0 classifies risk. G1 defines security requirements. G2/G3 reviews threat model. G4 verifies implementation. G5 penetration tests. No gate passes without security evidence.
5. The human decides. The AI provides evidence.
AI builds. AI tests. AI reviews. But the gate decision — go, kill, hold — is human. Always.
Project: Native iOS app for medical records requests (HIPAA-regulated, handles PHI)
| Gate | What Happened | Time |
|---|---|---|
| G0 | Problem validated: #1 support call is "where's my request?" | Day 1 |
| G1 | 13-section spec written. B-Team review: FAIL (6 criticals). Fixed. Zero-Context review: PASS (3 clarifications added). Resubmitted. PASS. | Day 1 |
| G2/G3 | 4 iterations of B-Team and Zero-Context adversarial review. 7 ADRs. 18-threat model. Zero findings rejected. | Day 1 |
| G4 | 53+ tests. Security review: PASS. Accessibility audit: PASS. Integration tests against real API. Zero-Context catches missing error handling documentation. | Day 3 |
Three days from "should we build this?" to a verified, secure, accessible working application — with full audit trail.
Without governance, that same AI could have produced the same volume of code in 3 hours. But it would have shipped with a race condition in the submit guard, a missing file upload parameter, and undocumented assumptions about patient consent flow that a regulator would have rejected.
| Gate | Question | Evidence Required | Reviewers |
|---|---|---|---|
| G0 | Worth doing? | Problem statement, users, risk classification, scope | Product + Director |
| G1 | Requirements solid? | Testable requirements, API contracts, NFRs, B-Team review, Zero-Context review | Director + B-Team + Zero-Context |
| G2/G3 | Design safe? | ADRs, threat model, test strategy, 3 Amigos on tickets, Zero-Context assumptions audit | Architect + Security + Zero-Context |
| G4 | Code works? | 100% coverage, security review, accessibility audit, integration tests, Zero-Context clarity check | Test + Security + Director + Zero-Context |
| G5 | Safe to ship? | Pen test, rollback plan, monitoring, NFR evidence, full team sign-off, Zero-Context auditor perspective | Everyone |
| G6 | What did we learn? | Production metrics, incidents, KPI actuals vs. targets | Director + SRE + Journalist |
┌─────────────────────┐
│ A-TEAM (Builder) │
│ │
│ • Writes spec │
│ • Builds code │
│ • Runs tests │
│ • Proposes design │
└──────────┬──────────┘
│
Submits for review
│
┌───────────┴───────────┐
│ │
▼ ▼
┌─────────────────────┐ ┌──────────────────────┐
│ B-TEAM (Critic) │ │ZERO-CONTEXT REVIEWER │
│ │ │ │
│ • Different AI │ │ • Different model │
│ • Hostile personas │ │ • No dev knowledge │
│ • Domain expert │ │ • Naive eye │
│ • Finds what │ │ • Auditor/Regulator │
│ builders missed │ │ perspective │
│ • CRITICAL/MAJOR/ │ │ • Surfaces hidden │
│ MINOR/OBS │ │ assumptions │
└──────────┬──────────┘ │ • Clarifies docs │
│ │ • Flags confusing │
│ │ terminology │
│ └──────────┬───────────┘
│ │
└───────────┬────────────┘
│
Both reviews complete
│
▼
┌─────────────────────┐
│ A-TEAM Response │
│ │
│ • Accept & fix │
│ • Accept & defer │
│ • Reject (with │
│ evidence only) │
└──────────┬──────────┘
│
Resubmit until PASS
│
▼
┌─────────────────────┐
│ GATE DECISION │
│ (Human only) │
│ │
│ GO / KILL / HOLD │
└─────────────────────┘
B-Team finds things that are wrong or incomplete — defects in the implementation, missing test cases, architectural flaws. The B-Team is critical but has context. They know what you were trying to do.
Zero-Context Reviewer finds things that are invisible to the B-Team because the B-Team built it. It finds:
In healthcare, this is the regulatory auditor asking "why is this the way it is?" not because they're hostile, but because they have no context.
Result: 4 review iterations. B-Team first review: 6 criticals. Zero-Context first review: 3 clarifications + 2 assumption surfacing. After fixes, B-Team: PASS with 5 minors. Zero-Context: PASS with 1 note.
┌─────────────────────────────────────────────────────────────┐ │ │ │ 1. Patient safety outranks throughput. │ │ │ │ 2. Evidence, not confidence. │ │ "We think it's fine" is not a gate pass. │ │ │ │ 3. The B-Team reviews every gate. │ │ Different AI. Different blind spots. │ │ │ │ 4. Zero-Context reviews every gate. │ │ Naive eye catches invisible assumptions. │ │ │ │ 5. AI builds. Humans decide. │ │ The gate decision is never automated. │ │ │ │ 6. 100% test coverage on new code. │ │ If AI writes it, AI tests it. No excuses. │ │ │ │ 7. Security is every-gate, not a phase. │ │ │ │ 8. Accessibility is every-ticket, not a sprint. │ │ │ │ 9. Kill early, kill cheap. │ │ A failed gate at G1 costs hours. │ │ A failed gate at G5 costs months. │ │ │ └─────────────────────────────────────────────────────────────┘
Healthcare software has a unique constraint: you cannot move fast and break things when "things" are patient records.
Agentic AI removes the speed constraint on building. It does NOT remove the constraint on safety. Stage-Gate provides the structure to let AI build at full speed while humans maintain control over what ships.
The gates aren't bureaucracy. They're the reason we caught:
All found before any patient was affected. All found because the gate required evidence, not confidence.
Built with AI. Governed by humans. Verified by evidence.