Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.
Reviewed 30 items for 2026-W40, policy review-routing-2026-09-27-v1. Kept 2 at Watch: a Sonnet 5.5 model-watch entry already moving through the existing auto-eval pipeline (no new artifact needed until the scorecard lands), and the deferred Three Wise Men PEP-636 proposal for tax-scanner-api's rule dispatcher, held at Watch because its code claims are explicitly self-flagged as assumed/unverified pending a read of the real repo. Dropped the remaining 28: seven Gemini blog posts are 4-month-stale catalog churn already superseded by later releases; several Claude Fable/Opus items describe a model stack already adopted or already deprecated per current policy; a cluster of Sebastian Raschka/HuggingFace/LlamaIndex/Lovable/Latent Space posts are Tier 2-3 generic practitioner content with no named product decision or artifact change; one model-watch entry (gpt-6.1-sol-pro) is pure catalog noise with no incumbent or product list attached.
Model-watch auto-detected anthropic/claude-sonnet-5.5, same publisher lineage as the incumbent Sonnet tier used across quincy, my-day, lengua, newton, cicero, euclid, ledger-essentials, financial-planning-agent, floodstream; auto-eval already queued by existing infra (cap 5/run) and will produce a scorecard row.
Why this verdict: Tier 1 (direct model-registry detection), medium novelty (named successor to current incumbent, pricing not yet compared), but the artifact work (auto-eval, scorecard) is already running automatically — no new Factory artifact is needed today, so it doesn't clear Factory Candidate; gain is real but not yet measured, so holding at Watch until the scorecard exists avoids acting on an unverified frontier claim.