Factory Review · 2026-W39

What the factory watched this week.

Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.

30 items in, 11 survived, 19 dropped. Three near-duplicate Cursor SDK posts collapsed to one Product Input for shepherd/ai-practice-watch. Five llamaindex liteparse/llamaparse posts collapsed to one Watch; Mistral OCR 4 (carried over from last week's cap) promoted to Product Input for floodstream/tax-return-tools given a concrete callable API and price. Eight Sentry blog posts collapsed to one low-priority Watch on the Claude-skill authoring pattern — rest were the same content campaign with no named product decision. Five model-watch entries (Opus 5.5, DeepSeek V4.1 Flash, Fable 5.1, GPT-6 Astra Pro) held at Watch since auto-eval is already queued and will produce the actual scorecard; two other model-watch entries dropped as untrusted-publisher or batch-pricing catalog variants. The internal Three Wise Men PEP636 report flags a plausible tax-scanner-api billing dispatch bug but is explicitly built on an assumed (unverified) code pattern and is out of scope for this session's repo — logged as Product Input requiring verification against the real dispatcher before any change. Nothing here clears the bar for Factory Candidate or Immediate Factory Upgrade this week.

The headline calls

Product Inputmodel catalog

The Three Wise Men Debate Report

Proposes a decorator-based Rule Registry (O(1) dispatch) plus Pydantic-validated pattern matching for tax-scanner-api's rule dispatcher and Claude-response parser, explicitly to close a silent fall-through path that could mis-bill.

Why this verdict: Internally generated proposal (Tier 3 — self-authored analysis, explicitly speculative: 'Current pattern (assumed)') naming a plausible billing-correctness defect, which is material if real. Held at Product Input, not higher, because the premise is unverified against the actual repo, and this session is scoped to a different repo (one-repo-per-window) so it cannot verify or act on it here.

claims to check before believing:
  • Assumed if/elif dispatcher pattern and 'silent fall-through' billing bug in tax-scanner-api not confirmed against actual source

The rest of the week (10)

Product InputIntroducing Mistral OCR 4
Product Inputsdk release
WatchNew model: `anthropic/claude-opus-5.5`
Watchliteparse server self hostable document parsing
WatchValidating agentic behavior when “correct” isn’t deterministic
WatchNew model: `deepseek/deepseek-v4.1-flash`
WatchNew model: `anthropic/claude-fable-5.1`
Watch100 things we announced at I/O 2026
Watchai bug reproduction
WatchNew model: `openai/gpt-6-astra-pro`