Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.
30 items judged, 7 kept / 23 dropped. Two real signals: Mistral OCR 4 (Product Input — confidence-scored OCR directly upgrades Floodstream/tax doc ingestion, needs a 5-doc accuracy test before adoption) and Vercel Konsistent (Factory Candidate — structural lint parity for agent- and human-written code, worth a single-package pilot; effort reassessed Low since it is a single config rule, no migration). Five Watches held pending more data or a named pilot: DeepSeek V4 Flash (scorecard pending), two Cursor changelog items (no current Cursor usage), Mistral's agentic test-writing pattern, and Eugene Yan's summary-eval ladder (both reusable for Factory's own judge/verification gates but not yet tied to a specific retrofit). The 17 model-watch batch-SKU entries are almost entirely noise: deprecated Opus 4.1/4.5/4.6/4.7 tiers, superseded legacy models (GPT-3.5, o3-mini), and duplicate batch variants of already-judged frontier events (Sonnet 5, Haiku 4.5, Fable 5, GPT-5.1) — none cleared the novelty or incumbent-policy gates. Remaining cloudflare/HF blog items were generic tag pages or patterns with no named product decision attached.
New OCR API version adds bounding boxes, block classification, and per-block confidence scores at a stated batch price, which lets an ingestion pipeline route low-confidence blocks to human review instead of silently trusting flat-text OCR output.
Why this verdict: Tier 1 vendor announcement with concrete callable API and pricing; novelty is real (confidence scores + block classification are new vs typical OCR passthrough); expected gain is Medium (fewer silent mis-extractions in money/claims documents) for Medium effort (swap-in test, no pipeline rewrite) — gain plausibly exceeds effort, but pricing/accuracy claims are vendor-stated and unverified.
Konsistent enforces custom structural lint rules (file-shape, export contracts, type contracts) identically for AI-agent-written and human-written code in a monorepo, catching drift a normal linter/style-guide can't.
Why this verdict: Tier 1 vendor changelog for a shipped feature; novel enforcement mechanism (agent+human parity) not currently present in any repo; expected gain Medium (catches agent-code drift before it compounds) for Low effort (one rule, one package, no migration, feature is config-only) — gain exceeds effort for a bounded pilot; full-repo rollout would be higher effort and is NOT what's being proposed here.
New-generation DeepSeek V4 Flash at very low per-token pricing with 1M context; if it lands on the price/quality frontier it could displace a current low-cost rotation-tier model once the auto-eval scorecard exists.
Why this verdict: Tier 1 model registry detection with real pricing; novelty is genuine (new major version, V3→V4) and it was deferred by cap last week rather than resolved, so it's not yet actioned; expected gain unknown pending data, so it sits at Watch rather than Candidate until the scorecard exists — promoting now would be gain-unproven, not gain>effort.