Content pipeline — how new pages go live every day
Wikidata ──(ingest, daily if queue low)──▶ D1: status='queued' ──(publish cron 01:00 UTC)──▶ status='published' ──▶ sitemaps + IndexNow
│ │
│ translate missing titles (LLM)
│ publish artists that now have works
│ recount artist & browse-page sizes
└──────────────────────────────────────────────1. Ingest (fill the queue)
- Script:
scripts/ingest-wikidata.mjs - Workflow:
.github/workflows/ingest-catalog.yml— 00:15 UTC daily; ingests only when fewer thanQUEUE_TARGET(15,000) artworks are queued. Manual run: Actions → ingest-catalog → Run workflow (input: number of painters,force). - Order: most prolific public-domain painters first (≥ 5 paintings with an image, died before 1926), 20 per run.
- What it pulls per painting: labels in 6 languages, image, year, height/width (converted to cm), museum, movement, genre, up to 8 subjects, popularity (Wikipedia language links).
- Idempotent: keyed on Wikidata IDs (
INSERT OR IGNORE), so re-running never duplicates. The 31 hand-curated works are linked to their Wikidata IDs, so ingestion never creates a second page for them. - Nothing becomes public here — rows land as
queued.
Local trial:
bash
node scripts/ingest-wikidata.mjs --artists 2 --out /tmp/batch.sql
cd workers/api && npx wrangler d1 execute artlove365-platform --local --file /tmp/batch.sql2. Publish (daily, best first)
- Code:
workers/api/src/modules/seo/publisher.ts, cron*/10 * * * *(runPublishTick): publishes storied works as soon as they are ready, capped atPUBLISH_PER_DAYper UTC day. Browse/artist counts are refreshed after every publish and at least hourly. - Knob:
PUBLISH_PER_DAYinworkers/api/wrangler.jsoncvars(starts at 300). Change it, push tomain, done. Ramp plan: strategy §3. - Steps each run:
- take the next N
queuedartworks that already have a story (stories) — public artists below 10 works first, new artists only as a block of 10; - translate missing title languages — 10 titles per LLM call (
@cf/meta/llama-3.3-70b-instruct-fp8-fast), human Wikidata labels always win, disambiguation like "(Rubens, Prado)" is stripped; - set
status='published',published_at=now; - publish their artists;
- recount
artists.artwork_countandtaxa.artwork_count(decides which browse pages exist); - send all new URLs × 6 languages to IndexNow (Bing, Yandex, Seznam, Naver); Google picks them up from the sitemaps;
- write a row in
publish_log.
- take the next N
- Measured locally: 40 artworks incl. translation in ~12 s → 1,000/day ≈ 5 min.
3. Watch and steer
Staff endpoints (header Authorization: Bearer $ADMIN_TOKEN):
| Endpoint | Use |
|---|---|
GET https://api.artlove365.com/v1/admin/seo/stats | queue vs published, indexable browse pages, last 30 runs |
POST https://api.artlove365.com/v1/admin/seo/publish {"count": 500} | publish now (e.g. launch day) |
POST https://api.artlove365.com/v1/admin/seo/refresh-counts | recount after manual data fixes |
Weekly: check Google Search Console → Pages (indexed vs. discovered) and Sitemaps; raise or hold PUBLISH_PER_DAY per the strategy rule.