Back Office · office.temerarii.xyz
PERCEPTION-STACK.md

← all docs

PERCEPTION STACK — total sight + hearing for every asset (every frame, every second)

Goal (operator, 2026-06-16): give the agent COMPLETE sight + hearing — full comprehension of every

video, GIF, blog, email, SMS — at 30 fps (108,000 frames/hr) fidelity, every second of audio, nothing missed.

The principle: COVERAGE ≠ COMPREHENSION

Running a VLM on every 30 fps frame across the calendar = tens of millions of calls — intractable, and an

8 GB GPU can't host a frame-rate VLM. So don't. Total perception = two layers:

  1. COVERAGE (100%, cheap, deterministic): a per-frame + per-second signal pass over EVERY frame/second —

nothing is unseen/unheard at the signal level. This is what makes "every single frame" literally true.

  1. COMPREHENSION (semantic, expensive): a VLM/LLM reads the SALIENT subset (scene-change keyframes) +

escalates to 100% on any frame the coverage layer flags anomalous. Understanding where it matters,

guaranteed wherever something's wrong.

This is how a human editor perceives 108k frames without consciously studying each: reflex catches the

wrong frame; attention studies the meaningful ones. The agent gets both.

SIGHT — video & GIF (3 tiers)

  • Tier 0 · EVERY-FRAME signal (30 fps, deterministic, local, cheap): ffmpeg decodes all frames →

per-frame: blackdetect · freezedetect · signalstats (luma/contrast/saturation) · dominant-color +

red-% (colorway/§03) · pHash (dup/stutter) · zone occupancy (logo corner · caption band · astronaut

zone). Output = a per-frame timeline. Catches ANY black/frozen/blank/chrome-missing/color-drift frame,

anywhere. (Upgrade certify._walk_frames from 2 fps → 30 fps for this layer.)

  • Tier 1 · Salience (keyframes): ffmpeg scene-cut (select=gt(scene,0.3)) + I-frames + 1/scene-beat →

the semantically-distinct frame set (~1 per second of real change).

  • Tier 2 · Comprehension (VLM — keyframes + every flagged frame): Claude-vision (the Max-authed

critic.py judge, in-session) reads each keyframe + any Tier-0 anomaly → on-brand? matches caption/script?

craft? defect? Escalates to 100% on anomalies. Batch option: routed cloud VLM (fal/replicate qwen-VL).

  • On-screen TEXT: OCR at ≥1 fps → the full text-over-time track (banned words / typed wordmark / NaN /

placeholder anywhere, not just midpoints).

HEARING — audio (full)

  • Whisper full transcript (word-level timestamps) → every spoken word heard.
  • Per-second audio signal: LUFS-over-time · silencedetect (dead air) · clipping · VO-vs-music duck.
  • Cross-check (one-track): transcript == caption == script per scene; VO covers the duration (no dead air);

the spoken brand = "The Big T-M", 0 banned words — verified against what's actually HEARD, not just planned.

READ — blogs / email / text assets

  • Full text parse (NOT sampled) → voice · 0-banned · grounding (claims trace to a real source) · SEO/AEO ·

readability (FK) + an LLM coherence/accuracy read.

  • Render → screenshot → VLM so SIGHT also covers the rendered layout (not just the source).

EMAIL / SMS / BLOG — every asset = CREATIVE component (SIGHT) + COPY (READ), both in the matrix

Each email/SMS/blog carries TWO things and BOTH must be reviewed + linked in the one matrix:

  • The creative component (its Hero GIF / featured motion-graphic) → the GIF is a kind:gif asset

in the index (email-gif-W##-<day>, sms-gif-W##-<day>) → it gets the full 30fps SIGHT walk in

perceive AND a certify entry. VERIFIED on 2026-W27: all 9 GIFs (7 email + 2 sms) frame-walked + certified.

  • The copy (block-stack / subject / preheader / body / SMS text / blog article) → the READ tier:

certify_email (blocks · lockup · banned · fabricated, gated on a produced out/email/W##-<day>.html),

certify_sms (≤160 · banned · mms_gif), certify_blog (live · JSON-LD · canonical · banned).

  • Email: HTML → headless screenshot → VLM (sight) + text parse (read). SMS: text + its GIF → the video stack.

KNOWN COVERAGE GAPS (2026-W27, to close so the review is complete per asset)

  1. Email COPY not produced for 2026-W27 → certify_email skips (no out/email/W27-<day>.html), so each

email is represented ONLY by its GIF. FIX = render the 2026-W27 email block-stack HTML (task #129) →

certify_email then covers subject/preheader/blocks/banned. (2026-W27 needs these drafts anyway.)

  1. certify_blog hardcoded to the 2026-W23 live URL + a blog-W23 gate → blog-W27 reviews the 2026-W23

article, not 2026-W27's. FIX = resolve the blog URL/HTML per week from the index; SKIP-not-FAIL the

live check when that week's blog isn't published yet (2026-W27 holds until launch).

  1. SMS copy ✓ (W27-Tue/Fri.txt produced + certified). The GIF creative components ✓ (all SIGHT-walked).

The rule: NO asset is "reviewed" until BOTH its creative component (SIGHT/30fps) AND its copy (READ) are

in comp-id-review.json, linked. Video/social/thread/GIF = ✓ today; email-copy + blog-per-week = the close-out.

COST / LOCAL REALITY (honest)

  • Tier 0 + OCR + Whisper + per-second audio = local, GPU-light, 100% coverage, cheap (ffmpeg decodes

30 fps fast; per-frame numpy + pHash are cheap).

  • Tier 2 VLM = the only real cost → bounded to keyframes + flagged frames (NOT 108k); Claude-Max

in-session for the smart sample, cloud VLM for batch. The tiering is necessary, not a compromise.

  • Net: **every frame SEEN (signal) · every second HEARD (audio) · salient frames UNDERSTOOD (VLM) · every

anomaly escalated to full comprehension.**

THE 100% PRINCIPLE — comprehension · best-content-by-intention · zero waste (operator-locked 2026-06-16)

The sim is not just a CHECKER, it is the CONSTRUCTOR + OPTIMIZER. Three guarantees, each by design:

(1) 100% COMPREHENSION = COVERAGE (every frame seen @30fps + every second heard, perceive) + CONFORMANCE (render==spec parity: the deterministic render provably matches the declared ground/logo/caption/shape/motion) + GROUNDED-TRUTH (every claim cites the registry). Full per-frame VLM is intractable + unnecessary; comprehension ESCALATES to anomalies/keyframes. So: every frame examined, every defect pinned, every claim sourced — effective 100% comprehension, honest about the escalation.

(2) BEST CONTENT BY INTENTION = each cell carries an intention (what it must land, _intent_for). The planner BUILDS the spec from intention (content-aware best-fit comp + grounded copy from grounding.yaml); the craft-judge scores it; memory.learn_fit → planner-fit.json teaches the planner which (beat→comp) lands each intention best → it CONVERGES on best-content-per-intention, not random. The optimize loop (engine.sim optimize/loop) drives the worklist until the bar. Intention → grounded+brand-locked spec → judged → learned → better next time.

(3) ZERO WASTE — the core anti-waste IS the determinism: the spec sim is CHEAP (CPU, run freely) and the render is DETERMINISTIC, so we KNOW the exact pixels BEFORE spending GPU. Therefore: simulate cheaply → fix at the spec layer → render ONCE (GPU), only when the spec is green (verify_render_ready gates the pour) → no re-pours. Plus: coverage≠comprehension (no VLM-every-frame), produced-only (no walking/shipping stale), cert --workers (no OOM), the unified matrix (no duplicate stores). Render is the only spend and it follows a passing spec.

THE TIGHTEN-UP to make (1)+(2) literally 100% (build after the 2026-W27 unify lands):

  • cite-or-fail grounding — authoring may only assert facts that cite a grounding.yaml row (source:); the gate FAILs any number/tool/capability claim without a citation → a false claim is impossible to author (the registry, operator-curated, is the only truth-bound).
  • render==spec parity as a BLOCKING certify dim — extend the existing pixel checks (perceive/run_pixel) into a hard render_matches_spec gate covering ground·logo·caption·astronaut·shape·motion, so craft-conformance can't ship below 100%.
  • one-time golden + planner-fit — the operator signs the golden once (the taste bar); planner-fit reproduces it deterministically after. No per-asset taste re-litigation.

FIDELITY CONTRACT — the harness reflects what is ACTUALLY PUBLISHED (operator Q 2026-06-16: "what ensures what's reflected in the harness for every asset is what's actually published?")

The identity chain that must hold per asset: plan (content-index.json) == produced (out/final, certified+perceived) == published (CDN/posted) == measured. State today:

  • plan == produced: HARD. certify/perceive read the ACTUAL out/final/<aid>-<tag>.mp4 bytes (full decode + 30fps walk), not the plan; ONE source (content-index) drives harness + renderer + office; verify_content_alignment gates drift. So the matrix verdict is on the real produced file.
  • produced == published: PARTIAL (display-guarded, not cert-gated). The office scene_video_src serves the CDN copy ONLY when cdn.epoch >= RENDER_EPOCH (render_post_view.py:167), else falls back to the still — so the office never displays a stale publish. BUT there is no cert DIMENSION that FAILS an asset whose published copy ≠ the certified master.
  • THE CLOSE-OUT (makes it a hard gate): stamp a per-asset content fingerprint (sha of the final mp4 + the content-index record hash) at RENDER → carry into the cdn-map at UPLOAD (engine/stages/upload_media.py writes cdn.fp + cdn.epoch) → NEW certify dim published_current (= cdn.fp == sha(out/final master) AND cdn.epoch >= RENDER_EPOCH); FAIL ⇒ the published copy is NOT what the harness certified ⇒ it cannot ship. _content_fingerprint (static.py:213) is the primitive to reuse. Then plan==produced==published is fully ASSERTED, not display-hoped. Send gate stays: nothing publishes until the 2026-W27 GO; #262 produced-only model ships only current finals.

GRANULARITY CONTRACT — comp_id × dimension × frame/second, ONE bidirectional matrix (operator-locked 2026-06-16)

The grain is non-negotiable: every check records at the SCENE/comp_id level, per dimension, and (for video/audio) per frame/second — not rolled up to the asset. Verified state of the sim today:

  • comp-id-review.json ALREADY keys per-asset AND per-scene {n, beat, comp_id, tier, flags} → every flag traces to its exact comp_id. KEEP + extend this as the canonical matrix.
  • certify-<wk>.json today rolls the per-scene/comp_id static detail into 3 asset-level flags (static_clean · copy_specificity · craft_grade) + 10 render dims. GAP → fix: the cert must carry the full comp_id × dimension × timestamp breakdown forward (or store the join key) so a failed chrome_caption/ocr_clean/craft_grade links to the exact scene·comp_id·frame, not just the asset.
  • Bidirectional, ONE matrix: perceive.py writes its per-scene/per-frame perception record back INTO comp-id-review.json (the harness's canonical store) so harness == office read the same granular truth. The back-edge to the planner is ALREADY wired + verified — critic/operator verdict → memory.learn_fit() → planner-fit.json → treatment_planner.py reads it on the next pick; tune_from_reviews() re-weights the scorer. So: sim → comp-id-review (forward, granular) ⇄ verdict → memory → planner-fit → planner (back). The office surfaces it per comp_id × dimension × timestamp; nothing is asset-rolled-up that could be comp_id-grained.

BUILD PATH (on the existing sim — additive)

  • engine/sim/certify.py: _walk_frames step 2 fps → 30 fps for the Tier-0 signal pass (decode-all +

per-frame numpy signals + pHash); keep VLM/OCR on keyframes + flagged. **Record each per-frame signal under

its scene·comp_id** (not asset-global) so the timeline is comp_id-attributed.

  • NEW engine/sim/perceive.py: orchestrate sight+hearing → a per-scene/comp_id perception report

(seen/heard/understood + anomaly list with timestamps) merged INTO comp-id-review.json (the one matrix

the harness + office both read) AND summarized up into the cert matrix.

  • Hearing: a Whisper transcript stage + per-second loudness/silence, cross-checked to caption/script, per scene.
  • engine/sim/critic.py: feed the keyframe set + flagged frames (already Claude-as-judge); verdicts → memory → planner-fit (the back-edge).
  • Blog/email: screenshot→VLM + full-text read tiers, keyed per asset (no scenes) but per dimension.
  • Verify: the matrix shows, per comp_id, 30 fps coverage + full transcript + the VLM verdicts + every anomaly

frame's timestamp; the office links each cert dimension → its exact comp_id/scene/frame; 2026-W27 = the proof week.

Today vs target (honest delta)

  • Today: certify does a real full-DECODE pass (corruption, every frame) + full blackdetect + a 2 fps

visual walk + OCR at scene-midpoints/10 s + audio stream/duration/clipping. 2026-W27 = 54/54 at that fidelity.

  • Target (this doc): 30 fps signal coverage + Whisper full-hearing + keyframe/flagged VLM comprehension +

blog/email screenshot-VLM → genuine "every frame, every second, full comprehension."