Quality checklist
Scaled pages only rank if each one is useful. These rules are enforced in code; the monthly checks are done by a person.
Enforced in code
| Rule | Where |
|---|---|
| Text built only from real facts; missing fact ⇒ sentence dropped | packages/shared/src/describe.ts (tested) |
| Browse page needs ≥ 3 published works, else real 404 | MIN_TAXON_WORKS, API /v1/taxa/*, worker |
| Only famous-first publishing; nothing published without an image and an English title | ingest + publisher |
| Human Wikidata labels beat machine translation; Wikipedia disambiguation stripped | translateTitles(), ingest clean() |
Canonical, hreflang, 404 status, noindex on personal pages | apps/web/worker/seo.ts (tested) |
| Painter died before 1926 ⇒ public domain everywhere (life + 100) | ingest query |
Monthly (≈ 1 hour)
- Translation spot-check: 50 random titles published that month per language (es, fr, de, ja, vi). Fix bad ones in D1; if > 5 % are wrong in one language, adjust the prompt in
publisher.ts. - Search Console: indexed vs. discovered, Core Web Vitals, structured-data errors (Product, BreadcrumbList).
- Sample 20 artwork pages in a browser per language: image loads, facts correct, order form works.
- Decide
PUBLISH_PER_DAYfor next month (strategy §3).
Known weak spots
- Machine-translated Vietnamese titles are the weakest; e.g. "Đại họa chân dung …" instead of "Chân dung …". Spot-checks matter most there.
- Seeded (curated) works have hand-written descriptions; ingested works use generated ones. Curated text for the top 500 works is roadmap item F-SEO-10.