Module T3-01Neill's Vibe · Tier 3 · The Factory~8 min

Three other AIs found the bug mine swore wasn't there

The model that wrote the code said it was clean. Three rival models, asked separately, found two ways it would break in production. This is how a non-coder safely runs ~30 apps.

01 · the scene

"Looks good. Tests pass. Ship it." It would have shipped a bug.

One model wrote a change (one commit — a single saved edit to the code), reviewed its own work, and gave it a clean bill of health. The test suite was green. The only reason a real bug never reached users is that three different AIs were asked the same question — separately — and disagreed, each turning up a P0: a ship-blocking bug, the highest severity, the kind that must be fixed before anything goes out to real people.

cross-provider review of commit 42bf412 — 2026-05-15
# reenactment of the real scar — "review" here is shorthand for the runnable cross-review.mjs you'll download (real call: node cross-review.mjs 42bf412)
$ review 42bf412 --providers claude,codex,coderabbit # the 3-way first pass
claude → looks clean. tests green. no blockers.
codex → no blockers.
coderabbit→ no blockers.
$ review 42bf412 --providers gemini # the independent 4th model
gemini → P0 ×2 — unhandled error path + API-contract mismatch # P0 = ship-blocker
# re-run after Gemini's catch, with the others looking harder...
coderabbit→ P0 ×3 more — secondary writes silently skipped
Neill51% · the terrain

"I can't read the code. I'm trusting an AI that the thing it just built is safe. If the same AI that wrote it is also the one telling me it's fine, I have one opinion wearing two hats. Across ~30 apps, one opinion is how a quiet bug ships to real people."

Claude49% · the execution

"I was confidently wrong, and so were two of the others on the first pass. That's not a character flaw — it's what a probabilistic component is. The fix isn't a smarter single reviewer. It's a second, third, and fourth set of eyes from different models, asked separately, and you wait for all of them before you ship."

02 · the why

An AI is a confident guesser. So you don't ask just one.

Here's the one idea: the AI that wrote the code is the worst judge of it. It doesn't know your code is correct — it predicts the most likely next words, and "this looks fine" is often the most likely answer whether or not it's true. When it reviews its own work, it grades its own homework with the same blind spots it built in.

So you don't ask just one. Cross-provider review means sending the same change to different AI vendors — built by different labs, on different training — and asking each to review it separately. They don't share blind spots, so where one is confidently wrong, another flags it. The disagreement is the value.

And tests don't rescue you here. A test only checks the cases someone thought to write, often against fake stand-in data — so it can flash green while a whole untested path (a way the code can fail that nobody handled) slips through.

◎

Confident guesser

probabilistic, not sure

The model predicts plausible text. "Looks fine" can be the likely answer even when it's wrong. Self-review repeats its own blind spots.

≡

Different blind spots

cross-provider

Separate vendors miss different things. Where one is confidently wrong, an independent one catches it. The disagreement is the value.

✓

Tests aren't enough

green ≠ correct

A test checks only the cases written, often on fake data. A whole unhandled path can stay green. Review reads what tests never run.

The model: one engineer who reviews their own pull request (the formal "please merge my change" request) is theater — they share every assumption that produced the bug. A real shop sends the change to several reviewers from different teams and ships only when all of them sign off. Cross-provider review is exactly that, with AI reviewers: many independent eyes, no single confident voice trusted alone. This is the load-bearing reason a non-coder can run a fleet — the machine, not the person, polices the code.

03 · predict it

Tests green, the author-AI says "clean." Will a second AI find anything?

The model that wrote commit 42bf412 reviewed itself and found nothing. The tests passed. Before you read what happened — what's the most likely outcome when three different AIs review the same change separately?

Not it. Self-review repeats the same blind spots that wrote the bug, and a green test only proves the cases someone thought to write. Think about what a different model, built by a different lab, sees that the author can't.
Exactly. On 42bf412 the 3-way first pass missed two P0s (ship-blocking bugs) that Gemini caught — a way the code could fail that nobody had handled, and a place where one part of the code expected one shape of data and another part sent a different shape — and on the re-run CodeRabbit caught three more. Different models, different blind spots. The disagreement is the whole point.
04 · the walkthrough

How to put a change in front of four different judges.

grab the diff
step 1Take the exact change — the git diff of what's about to ship (a diff is the exact lines your change added or removed). You review the diff, not vibes, so every reviewer sees the same thing.
fan out to vendors
step 2Send that same diff to several different AI vendors separately — Claude, Codex, CodeRabbit, Gemini — each asked to find blocking bugs.
wait for all of them
step 3Wait for every reviewer, even the slow one. Dropping a reviewer because it's slow is how the P0 it would've caught ships.
verify before fixing
step 4For each flagged P0, check it against the real code before acting — a reviewer with no file access can raise a confident false alarm. Fix the real ones; ship. The tool you'll download does this gate for you: it exits non-zero on any P0, so a failed run is the one-click "do not ship" check you now own.
05 · prove it · gate

Teach it back to unlock the payoff.

Same rule as the whole course: answer it right, or the module isn't done.

Why is cross-provider review load-bearing — not overhead — even when tests pass and the author-AI says the code is clean?
// one correct answer unlocks Module T3-01
Not the load-bearing reason. Green tests and a self-review don't prove correctness — they repeat the author's assumptions. Think about why different vendors, with different blind spots, catch what one voice can't.
That's it. The model is a confident guesser; the author grades its own homework; tests only run the cases someone wrote. Independent vendors don't share blind spots, so the disagreement surfaces the real bug — and you ship only after all of them have looked. On 42bf412 that turned "clean" into five caught P0s. Module unlocked. ↓
the durable idea

Never bet the build on one set of eyes. A confident guesser that grades its own homework will ship a bug a green test never noticed — so you ask several different models, separately, and wait for all of them.

06 · factory payoff
⊘ locked — pass the teach-back to claim this piece of the factory

The tool that fans one change out to four reviewers — and waits.

A real, runnable cross-review.mjs you download and run on your own machine: it takes the git diff of your change, sends that same diff to several review providers at once and waits for all of them to finish even if some fail (that's a Promise.allSettled — run several checks at once and wait for ALL to finish, even the ones that error), then prints the union of their findings with a clear ship / don't-ship verdict (every P0 blocks the ship until you verify it against the real code). The provider calls are a stub — a placeholder you replace with the real call — behind one tiny function you swap for whichever models you use; nothing is hidden. Plus the one-page rule it enforces: the piece of the machine that lets a non-coder trust a fleet they can't read line by line.

Don't take the tool's word for it — make it catch a real one. Download it, wire your providers into the one stubbed function, then plant an obvious bug in a change and run it against that diff. If the wall holds, it refuses to ship. Here's exactly what that looks like on a diff with one planted P0 (a reenactment — you run this in your own terminal, not on this page):

cross-review.mjs — planted P0, providers wired
$ node cross-review.mjs main # review everything since main
Reviewed by 4 providers (claude, codex, coderabbit, gemini).
✗ 1 P0(s) found — verify each against the real code before fixing:
- [gemini] unhandled error path: API response not checked before use
$ echo $? # the exit code
1 # non-zero = do NOT ship — this is the wall a CI merge step reads

Read that 1 out loud — that's the wall. A green test never saw this path; one independent model did, and the tool turned "looks clean" into a hard stop before it reached real people. That non-zero exit is the machine policing the fleet for you.

✓ Module T3-01 complete