Back Office · office.temerarii.xyz
SEO-FRAMEWORK.md

← all docs

SEO + Site Speed — the framework

The Chairman asked "what are we doing for SEO and site speed". The answer, as data and stages rather

than a slide: free tier now (Search Console · GA4 organic → doors · CrUX · PageSpeed Insights · our

own crawler) and DataForSEO live behind credentials (he funds a $50 deposit and pastes the login;

until then the stage prints what it would cost and refuses). Every number on the office's

/reports/seo page comes from a row a pull READ from a named system, on a fiscal-week key

(2026W40); a missing source is an empty cell with its reason. Nothing is typed by hand.

Governing rule (shared with docs/strategy/REPORTING-SPEC.md): never fabricate a number.

No search volumes without DataForSEO, no "estimated traffic", no keyword difficulty scores from a

model. A draft keyword list has terms and sources, never volumes.


Sources by tier — what each answers, what it costs

TierSourceStageAnswersCostNeeds
ownour crawler (crawl_site)engine/stages/crawl_site.pyWhat a non-rendering bot receives from every public page: title / description / canonical / og:image / viewport / h1 / alt / robots meta, redirect chains, broken links, sitemap vs realityfreenothing — public pages
ownthe built pages on diskengine/gates/verify_seo.pyThe same contract BEFORE deploy (apps/lab/index.html + apps/lab/public/**/index.html, the blog included); sitemap.xml == built pages; robots.txtfreenothing
freeGoogle Search Consoleengine/stages/pull_metrics_gsc.pyClicks · impressions · CTR · position per day, per property and per key url; the queries people typed (sidecar); sitemap status; URL Inspection = is the key url indexedfreethe Google OAuth token (python scripts/_google_oauth.py, expires ~weekly while the app is in testing mode)
freeGA4 organic → doorsengine/stages/pull_metrics_ga4.py (extension)Organic-search sessions, and which land on /go/<segment>freethe same token; the doors on a GA4-tagged domain (the vercel.app host carries no tag) — declared not_connected until both
freeChrome UX Report (CrUX)engine/stages/pull_metrics_crux.pyReal-user p75 LCP / INP / CLS per origin and per key url, trailing 28 days — the field truth Google ranks withfreePSI_API_KEY (a Google Cloud API key with the CrUX + PSI APIs enabled; no billing)
freePageSpeed Insights (Lighthouse)engine/stages/pull_metrics_psi.pyLab performance / accessibility / SEO / best-practices 0–100 + lab LCP / CLS / TBT per key url, mobile + desktop, and the failed audits (what to fix)free (~25 s per url per strategy)the same key
freeLighthouse CI.github/workflows/lighthouse.ymlThe same lab run on every push to master + weekly, held to governance/seo/budgets.yaml, reports as artifactsfree (GitHub minutes)the credential purge first — CI cannot be pushed on the present history
freeGoogle Autocompleteengine/stages/derive_keyword_themes.pyWhat people actually type after a seed term (public suggest endpoint, rate-limited) — phrasing, not volumefreenothing
paidDataForSEO — Google Ads search volumeengine/stages/pull_metrics_dataforseo.py --volumeMonthly volume · CPC · competition for every term in keywords.yaml (≤ 1,000 per task)≈ $0.075 per task of 1,000 terms — verify on dataforseo.com/pricingDATAFORSEO_LOGIN + DATAFORSEO_PASSWORD, a prepaid balance (min. $50)
paidDataForSEO — SERP organic (standard queue)… --serpOur Google position per term (US / en, depth 10) + the top-10 domains≈ $0.0006 per term — verifysame
paidDataForSEO — backlinks summary… --backlinksBacklinks · referring domains · domain rank for our two domains + the competitors the Chairman lists≈ $0.02 per domain — verifysame

Order of magnitude for the paid tier with today's draft (~190 terms, 2 domains): **well under $1 per

run**; weekly for a year ≈ $30–50. The stage prints the exact estimate before any call and stops

without --yes. Prices are constants in the stage marked verify — they change.


The loop — how a search query becomes a calendar cell and a result column

GSC queries (pull_metrics_gsc → seo-gsc.json)
      │  what people typed to reach us · impressions with no clicks = demand we do not answer yet
      ▼
door themes (derive_keyword_themes → governance/seo/keywords.yaml, DRAFT)
      │  per icp.segments[] door: pain theme + solution theme, seeds from the segment's OWN copy,
      │  GSC queries that overlap, Autocomplete expansions; the Chairman edits, deletes, adds
      ▼
volumes + positions (pull_metrics_dataforseo, paid, cost printed first)
      │  which themes carry demand · where we rank today · who ranks above us
      ▼
calendar cells (the content machine — blog posts on temerarii.xyz/blog, door copy on /go/<segment>)
      │  a cell names its target theme; render_static_blog writes the page WITH its contract
      │  (title ≤ 60 aligned to h1 · description · self-canonical · og:image · JSON-LD)
      ▼
verify_seo (pre-deploy) → crawl_site (post-deploy) → the page is in sitemap.xml, indexable, fast
      ▼
result column (/reports/seo, fiscal weeks TY / LY / YoY / LW / WoW)
      │  Search Clicks · Impressions · CTR · Average Position (lower is better) · Indexed Pages ·
      │  CWV p75 LCP / INP / CLS · Lighthouse per key url · Keyword Positions · Referring Domains ·
      │  Crawl Findings — per property and per key url (the doors are rooms)
      └──▶ back to the top: the queries the new page now earns feed the next derivation

The join is structural, not statistical: a door is a room (go-publishers), a blog post is a page in

GSC's page dimension, and the fiscal week is the same key the content index uses — so "we shipped

the publishers post in W41; what did search do in W42–W45" is a lookup, not an attribution model.


Run order

python -m engine.stages.crawl_site                       # free, now — writes seo-crawl.json (+ rows once the channel is declared)
python engine/gates/verify_seo.py                        # the gate; also in engine/goal.py
python -m engine.stages.pull_metrics_gsc                 # needs the OAuth token; 16 months first, incremental after
python -m engine.stages.derive_keyword_themes            # draft keywords.yaml (better AFTER gsc: it reads the queries)
python -m engine.stages.pull_metrics_crux                # needs PSI_API_KEY
python -m engine.stages.pull_metrics_psi                 # needs PSI_API_KEY; ~15 min for the full set — nightly, not per build
python -m engine.stages.pull_metrics_dataforseo --yes    # needs DATAFORSEO_*; prints the estimate first; --collect for queued SERPs
python apps/office/render_reports.py                     # /reports/seo from the fold — no API at render

Cadence: crawl + gate on every build · GSC daily (it lags 2–3 days) · CrUX weekly (it is a 28-day

window) · PSI nightly · DataForSEO weekly (positions) and monthly (volumes, backlinks).

Registries: governance/seo/urls.yaml (the key urls = rooms; the 13 doors are generated from

gtm.yaml icp.segments[]), governance/seo/keywords.yaml (draft, Chairman edits),

governance/seo/budgets.yaml (the thresholds — one file for the gate AND Lighthouse CI),

governance/seo/reporting-seo-block.json (the seo channel for governance/reporting.json).


What needs the Chairman

ItemWhyHow
Re-run the Google OAuththe token is expired (invalid_grant, verified 2026-08-24); GSC and GA4 refuse until thenpython scripts/_google_oauth.py (the app is in testing mode → ~weekly; publishing the OAuth app ends that)
A PageSpeed / CrUX API keyCrUX requires a key; PSI is rate-limited without oneGoogle Cloud console → enable PageSpeed Insights API + Chrome UX Report API → create an API key → PSI_API_KEY= in .env (free)
DataForSEO login + $50the paid tier; nothing is called without itapp.dataforseo.com → deposit → API access → DATAFORSEO_LOGIN= / DATAFORSEO_PASSWORD= in .env; the stage prints the estimate and needs --yes
Edit governance/seo/keywords.yamlit is a draft from the GTM's own words + autocomplete; the Chairman knows which phrases buyers usedelete / add terms; list competitor_domains; a chairman: block survives regeneration
Move the doors to a real domain/go/<segment> lives on temerarii-office-app.vercel.app (no GA4 tag, not a GSC property, no robots.txt, no sitemap, no canonical, no og:image) — search rows for the doors cannot exist until they are on a propertycut apps/web over to the final domain; add it in Search Console; add the GA4 tag; then organic_door_sessions can be connected
The credential purgeCI (gates.yml, lighthouse.yml) cannot be pushed on a history that carries live keysNovember: purge + rotate (memory: open-source-december-2026)
temerarii.com (Duda) fixesthe crawl's findings on www.temerarii.com are outside the codebase (six h1s, no sitemap line, no robots meta)in Duda; the gate reports them as review, not failure

What we deliberately do not do

  • No scraping of Google result pages (positions, People-Also-Ask, "related searches"). It is against

Google's terms and it is the thing the consent-layer GTM argues against. Positions come from

Search Console (our own data) and, paid, from DataForSEO's SERP API — a licensed source. PAA is

available through the same API if ever wanted; keywords.yaml says so instead of pretending.

  • No bought keyword or backlink lists, no "estimated traffic" from a model, no keyword-difficulty

score without a declared source. A term in keywords.yaml carries source: gtm.yaml:<field>,

source: gsc or source: autocomplete — nothing else.

  • No browser automation for measurement (memory: no-browser-automation). Lighthouse runs on

Google's PSI infrastructure and in GitHub CI; CrUX is an API; the crawler is urllib with a browser

UA (r2.dev and some hosts 403 a bare Python-urllib UA — engine/lib/http_probe.py).

  • No JS execution in the crawler. What it sees is what a non-rendering bot sees; that the Lab's

root ships no <h1> to such a bot is a finding, not a crawler limitation.

  • No third-party report hosting from CI (temporaryPublicStorage: false).
  • No numbers from a source that is not connected. A KPI whose source is missing is declared

not_connected with what it needs, and renders as such — never 0.


Current state (2026-08-24, the first real run)

  • The crawl ran against all three origins; verify_seo.py reads it. The findings are in the gate's

output and the report — the doors host is the loud one (no canonical / og:image / robots / sitemap,

one title on 13 pages); temerarii.xyz is close (root has no h1 in the served HTML, two pages lack

og:image, one description is over 160, /blog/author/… canonical and sitemap disagree on the slash).

  • GSC · CrUX · PSI · DataForSEO all refuse cleanly today with the exact next step; nothing is written.
  • governance/reporting.json does not yet carry the seo channel; every stage checks and refuses to

write rows until it does (the crawl JSON is still written — it is the gate's input, not a metric).