Skip to content

Content pipeline — how new pages go live every day ​

Wikidata ──(ingest, daily if queue low)──▶ D1: status='queued' ──(publish cron 01:00 UTC)──▶ status='published' ──▶ sitemaps + IndexNow
                                                        │                     │
                                                        │          translate missing titles (LLM)
                                                        │          publish artists that now have works
                                                        │          recount artist & browse-page sizes
                                                        └──────────────────────────────────────────────

1. Ingest (fill the queue) ​

  • Script: scripts/ingest-wikidata.mjs
  • Workflow: .github/workflows/ingest-catalog.yml — 00:15 UTC daily; ingests only when fewer than QUEUE_TARGET (15,000) artworks are queued. Manual run: Actions → ingest-catalog → Run workflow (input: number of painters, force).
  • Order: most prolific public-domain painters first (≥ 5 paintings with an image, died before 1926), 20 per run.
  • What it pulls per painting: labels in 6 languages, image, year, height/width (converted to cm), museum, movement, genre, up to 8 subjects, popularity (Wikipedia language links).
  • Idempotent: keyed on Wikidata IDs (INSERT OR IGNORE), so re-running never duplicates. The 31 hand-curated works are linked to their Wikidata IDs, so ingestion never creates a second page for them.
  • Nothing becomes public here — rows land as queued.

Local trial:

bash
node scripts/ingest-wikidata.mjs --artists 2 --out /tmp/batch.sql
cd workers/api && npx wrangler d1 execute artlove365-platform --local --file /tmp/batch.sql

2. Publish (daily, best first) ​

  • Code: workers/api/src/modules/seo/publisher.ts, cron */10 * * * * (runPublishTick): publishes storied works as soon as they are ready, capped at PUBLISH_PER_DAY per UTC day. Browse/artist counts are refreshed after every publish and at least hourly.
  • Knob: PUBLISH_PER_DAY in workers/api/wrangler.jsonc vars (starts at 300). Change it, push to main, done. Ramp plan: strategy §3.
  • Steps each run:
    1. take the next N queued artworks that already have a story (stories) — public artists below 10 works first, new artists only as a block of 10;
    2. translate missing title languages — 10 titles per LLM call (@cf/meta/llama-3.3-70b-instruct-fp8-fast), human Wikidata labels always win, disambiguation like "(Rubens, Prado)" is stripped;
    3. set status='published', published_at=now;
    4. publish their artists;
    5. recount artists.artwork_count and taxa.artwork_count (decides which browse pages exist);
    6. send all new URLs × 6 languages to IndexNow (Bing, Yandex, Seznam, Naver); Google picks them up from the sitemaps;
    7. write a row in publish_log.
  • Measured locally: 40 artworks incl. translation in ~12 s → 1,000/day ≈ 5 min.

3. Watch and steer ​

Staff endpoints (header Authorization: Bearer $ADMIN_TOKEN):

EndpointUse
GET https://api.artlove365.com/v1/admin/seo/statsqueue vs published, indexable browse pages, last 30 runs
POST https://api.artlove365.com/v1/admin/seo/publish {"count": 500}publish now (e.g. launch day)
POST https://api.artlove365.com/v1/admin/seo/refresh-countsrecount after manual data fixes

Weekly: check Google Search Console → Pages (indexed vs. discovered) and Sitemaps; raise or hold PUBLISH_PER_DAY per the strategy rule.