Skip to content

SEO strategy — where the millions of URLs come from ​

Goal: ~1.8 million indexable URLs within 12 months, published gradually every day, each page carrying real, page-specific information. Numbers below were measured on Wikidata on 2026-10-04 (not estimates), except where marked est.

1. The multiplier ​

Every page exists in 6 languages (/en /es /fr /de /ja /vi), linked with hreflang. So 1 entity = 6 URLs. The question is which entities, and how many.

2. URL areas and their size ​

#AreaURL patternSourceEntities× 6 localesPhase
AArtwork pages (product pages)/{l}/artworks/{slug}Wikidata paintings with an image, painter died before 1926276,7751,660,6501 · live
BArtist pages/{l}/artists/{slug}Painters with ≥ 5 such paintings8,68952,1341 · live
CSubject pages ("paintings of horses")/{l}/subjects/{slug}Wikidata depicts (P180), ≥ 5 paintings4,73128,3861 · live
DMuseum pages/{l}/museums/{slug}Wikidata collection (P195), ≥ 5 paintings2,46314,7781 · live
EGenre pages (landscape, portrait…)/{l}/genres/{slug}Wikidata genre (P136)~300 est.~1,8001 · live
FMovement pages (Impressionism…)/{l}/movements/{slug}Wikidata movement (P135), ≥ 5 paintings915461 · live
GListing pagination?page=N on B–F, galleryderived—~60,000 est.1 · live
HCombination pages (artist × subject, movement × subject, museum × artist)/{l}/artists/{a}/{subject} …derived, only where ≥ 8 works~25,000 est.~150,0002
ICustom-portrait landing pages (occasion, pet breed, style)/{l}/portraits/{occasion}, /{l}/portraits/pets/{breed}curated lists~600 est.~3,6002
JJournal (art history articles)/{l}/journal/{slug}written + reviewed~1,000/yr est.~6,0003
KCustomer gallery (finished commissions, with consent)/{l}/made/{order}real ordersgrows with sales—3

Phase 1 total ≈ 1.76 M URLs (A–F, plus pagination). Phase 2–3 add ~0.2 M high-intent pages.

Where the traffic is expected to come from:

  • A (artworks) — long tail: "[painting] oil reproduction", "buy [painting] hand painted". ~94 % of URLs, low competition each.
  • B, C, D, F — mid tail: "Monet paintings", "paintings of horses", "Prado museum paintings".
  • I (portraits) — highest commercial intent: "custom dog oil portrait". Few pages, most revenue per page.

3. Daily publishing ramp ​

Pages are not published all at once. A daily cron publishes the next PUBLISH_PER_DAY artworks, most famous first (ranked by Wikipedia language links), and their artist and browse pages appear as soon as they qualify. See content pipeline.

PeriodPUBLISH_PER_DAYArtworks published (cumulative)Indexable URLs (cumulative, all areas)
Month 13009,000~60,000
Month 260027,000~180,000
Month 31,00057,000~380,000
Months 4–61,000147,000~950,000
Months 7–91,000237,000~1,500,000
Month 10–111,000276,775 (catalog complete)~1,760,000

Rule for raising the rate: only increase PUBLISH_PER_DAY when Google Search Console shows indexed / submitted ≥ 50 % for the previous month. Publishing faster than Google crawls a young domain only creates "Discovered – currently not indexed" pages.

4. Quality rules (non-negotiable) ​

Google treats mass-produced, low-value pages as scaled content abuse. Every page must be worth visiting on its own:

  1. Facts, not filler. Page text is built only from the entity's real facts (year, size, museum, movement, subjects) — a missing fact removes its sentence. Two pages never share the same body unless they share the same facts. → packages/shared/src/describe.ts
  2. Thin-page guard. Browse pages (C–F) exist only with ≥ MIN_TAXON_WORKS (3) published works; below that the URL is a real 404, not an empty page. Combination pages (H) will need ≥ 8.
  3. Real translations. Titles use Wikidata's human labels when they exist and an LLM with art context otherwise (machine translation by m2m100 was tested and rejected). Spot-check 50 titles per language every month → quality.
  4. Real 404s. Unknown slugs return status 404 — no soft-404s.
  5. One canonical per page, always https://artlove365.com/{locale}/….
  6. Personal pages are noindex: order tracking, checkout, admin.

5. Technical foundations (all live) ​

PieceWhereDoc
Server-rendered HTML for crawlers (title, meta, canonical, hreflang, Open Graph, JSON-LD, main content)apps/web/worker/seo.tsRendering
Sharded sitemaps with hreflang + lastmod, ≤ 50k URLs per fileapps/web/worker/index.ts, workers/api/src/modules/seo/routes.tsSitemaps
Daily publisher: translate → publish → recount → IndexNowworkers/api/src/modules/seo/publisher.tsContent pipeline
Wikidata ingest into the queuescripts/ingest-wikidata.mjs, .github/workflows/ingest-catalog.ymlContent pipeline
Edge cache: one render serves crawlers for an hourapps/web/worker/index.tsRendering

6. Open risks ​

  • Domain age. A new domain gets a small crawl budget; the ramp above is an upper bound, not a promise.
  • Images hot-linked from Wikimedia. Fine to start; must move to R2 + an image CDN before traffic grows (Roadmap F-SEO-07).
  • D1 size. 276k artworks + ~2.5M taxonomy links fits comfortably in D1's 10 GB; search needs FTS5 past ~50k works (F-SEO-12).