Factory Review · 2026-W36

What the factory watched this week.

Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.

30 items judged, 7 kept / 23 dropped. Two real signals: Mistral OCR 4 (Product Input — confidence-scored OCR directly upgrades Floodstream/tax doc ingestion, needs a 5-doc accuracy test before adoption) and Vercel Konsistent (Factory Candidate — structural lint parity for agent- and human-written code, worth a single-package pilot; effort reassessed Low since it is a single config rule, no migration). Five Watches held pending more data or a named pilot: DeepSeek V4 Flash (scorecard pending), two Cursor changelog items (no current Cursor usage), Mistral's agentic test-writing pattern, and Eugene Yan's summary-eval ladder (both reusable for Factory's own judge/verification gates but not yet tied to a specific retrofit). The 17 model-watch batch-SKU entries are almost entirely noise: deprecated Opus 4.1/4.5/4.6/4.7 tiers, superseded legacy models (GPT-3.5, o3-mini), and duplicate batch variants of already-judged frontier events (Sonnet 5, Haiku 4.5, Fable 5, GPT-5.1) — none cleared the novelty or incumbent-policy gates. Remaining cloudflare/HF blog items were generic tag pages or patterns with no named product decision attached.

The headline calls

Product Inputmistral-news

Introducing Mistral OCR 4

New OCR API version adds bounding boxes, block classification, and per-block confidence scores at a stated batch price, which lets an ingestion pipeline route low-confidence blocks to human review instead of silently trusting flat-text OCR output.

Why this verdict: Tier 1 vendor announcement with concrete callable API and pricing; novelty is real (confidence scores + block classification are new vs typical OCR passthrough); expected gain is Medium (fewer silent mis-extractions in money/claims documents) for Medium effort (swap-in test, no pipeline rewrite) — gain plausibly exceeds effort, but pricing/accuracy claims are vendor-stated and unverified.

claims to check before believing:
  • Batch pricing of $2 per 1,000 pages
  • Confidence-score accuracy/calibration claims not independently benchmarked
Factory Candidatevercel-changelog

enforce consistent code for agents and humans with konsistent

Konsistent enforces custom structural lint rules (file-shape, export contracts, type contracts) identically for AI-agent-written and human-written code in a monorepo, catching drift a normal linter/style-guide can't.

Why this verdict: Tier 1 vendor changelog for a shipped feature; novel enforcement mechanism (agent+human parity) not currently present in any repo; expected gain Medium (catches agent-code drift before it compounds) for Low effort (one rule, one package, no migration, feature is config-only) — gain exceeds effort for a bounded pilot; full-repo rollout would be higher effort and is NOT what's being proposed here.

claims to check before believing:
  • Exact rule syntax/capabilities of Konsistent beyond the changelog description
Watchmodel catalog

New model: `deepseek/deepseek-v4-flash-0731`

New-generation DeepSeek V4 Flash at very low per-token pricing with 1M context; if it lands on the price/quality frontier it could displace a current low-cost rotation-tier model once the auto-eval scorecard exists.

Why this verdict: Tier 1 model registry detection with real pricing; novelty is genuine (new major version, V3→V4) and it was deferred by cap last week rather than resolved, so it's not yet actioned; expected gain unknown pending data, so it sits at Watch rather than Candidate until the scorecard exists — promoting now would be gain-unproven, not gain>effort.

claims to check before believing:
  • Stated per-token pricing and 1,048,576 token context not independently verified

The rest of the week (4)

WatchEvaluation & Hallucination Detection for Abstractive Summaries
WatchRails testing on autopilot: Building an agent that writes what developers won't
Watch06 18 26
Watchcloud in agents window