SEO strategy — where the millions of URLs come from
Goal: ~1.8 million indexable URLs within 12 months, published gradually every day, each page carrying real, page-specific information. Numbers below were measured on Wikidata on 2026-10-04 (not estimates), except where marked est.
1. The multiplier
Every page exists in 6 languages (/en /es /fr /de /ja /vi), linked with hreflang. So 1 entity = 6 URLs. The question is which entities, and how many.
2. URL areas and their size
| # | Area | URL pattern | Source | Entities | × 6 locales | Phase |
|---|---|---|---|---|---|---|
| A | Artwork pages (product pages) | /{l}/artworks/{slug} | Wikidata paintings with an image, painter died before 1926 | 276,775 | 1,660,650 | 1 · live |
| B | Artist pages | /{l}/artists/{slug} | Painters with ≥ 5 such paintings | 8,689 | 52,134 | 1 · live |
| C | Subject pages ("paintings of horses") | /{l}/subjects/{slug} | Wikidata depicts (P180), ≥ 5 paintings | 4,731 | 28,386 | 1 · live |
| D | Museum pages | /{l}/museums/{slug} | Wikidata collection (P195), ≥ 5 paintings | 2,463 | 14,778 | 1 · live |
| E | Genre pages (landscape, portrait…) | /{l}/genres/{slug} | Wikidata genre (P136) | ~300 est. | ~1,800 | 1 · live |
| F | Movement pages (Impressionism…) | /{l}/movements/{slug} | Wikidata movement (P135), ≥ 5 paintings | 91 | 546 | 1 · live |
| G | Listing pagination | ?page=N on B–F, gallery | derived | — | ~60,000 est. | 1 · live |
| H | Combination pages (artist × subject, movement × subject, museum × artist) | /{l}/artists/{a}/{subject} … | derived, only where ≥ 8 works | ~25,000 est. | ~150,000 | 2 |
| I | Custom-portrait landing pages (occasion, pet breed, style) | /{l}/portraits/{occasion}, /{l}/portraits/pets/{breed} | curated lists | ~600 est. | ~3,600 | 2 |
| J | Journal (art history articles) | /{l}/journal/{slug} | written + reviewed | ~1,000/yr est. | ~6,000 | 3 |
| K | Customer gallery (finished commissions, with consent) | /{l}/made/{order} | real orders | grows with sales | — | 3 |
Phase 1 total ≈ 1.76 M URLs (A–F, plus pagination). Phase 2–3 add ~0.2 M high-intent pages.
Where the traffic is expected to come from:
- A (artworks) — long tail: "[painting] oil reproduction", "buy [painting] hand painted". ~94 % of URLs, low competition each.
- B, C, D, F — mid tail: "Monet paintings", "paintings of horses", "Prado museum paintings".
- I (portraits) — highest commercial intent: "custom dog oil portrait". Few pages, most revenue per page.
3. Daily publishing ramp
Pages are not published all at once. A daily cron publishes the next PUBLISH_PER_DAY artworks, most famous first (ranked by Wikipedia language links), and their artist and browse pages appear as soon as they qualify. See content pipeline.
| Period | PUBLISH_PER_DAY | Artworks published (cumulative) | Indexable URLs (cumulative, all areas) |
|---|---|---|---|
| Month 1 | 300 | 9,000 | ~60,000 |
| Month 2 | 600 | 27,000 | ~180,000 |
| Month 3 | 1,000 | 57,000 | ~380,000 |
| Months 4–6 | 1,000 | 147,000 | ~950,000 |
| Months 7–9 | 1,000 | 237,000 | ~1,500,000 |
| Month 10–11 | 1,000 | 276,775 (catalog complete) | ~1,760,000 |
Rule for raising the rate: only increase PUBLISH_PER_DAY when Google Search Console shows indexed / submitted ≥ 50 % for the previous month. Publishing faster than Google crawls a young domain only creates "Discovered – currently not indexed" pages.
4. Quality rules (non-negotiable)
Google treats mass-produced, low-value pages as scaled content abuse. Every page must be worth visiting on its own:
- Facts, not filler. Page text is built only from the entity's real facts (year, size, museum, movement, subjects) — a missing fact removes its sentence. Two pages never share the same body unless they share the same facts. →
packages/shared/src/describe.ts - Thin-page guard. Browse pages (C–F) exist only with ≥
MIN_TAXON_WORKS(3) published works; below that the URL is a real 404, not an empty page. Combination pages (H) will need ≥ 8. - Real translations. Titles use Wikidata's human labels when they exist and an LLM with art context otherwise (machine translation by m2m100 was tested and rejected). Spot-check 50 titles per language every month → quality.
- Real 404s. Unknown slugs return status 404 — no soft-404s.
- One canonical per page, always
https://artlove365.com/{locale}/…. - Personal pages are
noindex: order tracking, checkout, admin.
5. Technical foundations (all live)
| Piece | Where | Doc |
|---|---|---|
| Server-rendered HTML for crawlers (title, meta, canonical, hreflang, Open Graph, JSON-LD, main content) | apps/web/worker/seo.ts | Rendering |
| Sharded sitemaps with hreflang + lastmod, ≤ 50k URLs per file | apps/web/worker/index.ts, workers/api/src/modules/seo/routes.ts | Sitemaps |
| Daily publisher: translate → publish → recount → IndexNow | workers/api/src/modules/seo/publisher.ts | Content pipeline |
| Wikidata ingest into the queue | scripts/ingest-wikidata.mjs, .github/workflows/ingest-catalog.yml | Content pipeline |
| Edge cache: one render serves crawlers for an hour | apps/web/worker/index.ts | Rendering |
6. Open risks
- Domain age. A new domain gets a small crawl budget; the ramp above is an upper bound, not a promise.
- Images hot-linked from Wikimedia. Fine to start; must move to R2 + an image CDN before traffic grows (Roadmap F-SEO-07).
- D1 size. 276k artworks + ~2.5M taxonomy links fits comfortably in D1's 10 GB; search needs FTS5 past ~50k works (F-SEO-12).