Back Office · office.temerarii.xyz
One asset, all the way in — composition, the wireframe + storyboard, the output format stack, and the template, all read from the SAME content-index record. The expected output matches what /media surfaces for this post.
post thread-the-stack-W34-Tuekind threadweek W34date 2026-08-25campaign thread · Tuepillar socialbeat Tueasset videoduration 37.7sground blackscenes 6

Checklist the per-video bar — engine/sim

86.8/100
plain languagevo coverageno dead airuniquenesscaption fitcompletenesscleanliness
quantitative quality · weights learn from your reviews (engine.sim.memory review thread-the-stack-W34-Tue good|bad)
⚠ 5 flag(s) — not yet ship-ready: low_vo_coveragedead_airgeneric_scene · see docs/strategy/VIDEO-CHECKLIST.md

Composition comp · template family · expected output

composition SceneReelfamily / template SchematicCard
9:16 Reelrendered1:1 Squarepending16:9 Widepending9:16 4Kpending1:1 4Kpending16:9 4KpendingGIF (SMS)pending
▶ open rendered mp4
expected output: 1/7 rendered — same matrix the /media preview surfaces for this asset.

Composition layer × scene 6 scenes · 37.7s · comp_id + rendered still + tier + the script

#Layer (comp_id · still · tier)BeatTimecodeMotionLogoAudioVO / on-screen / caption
1s1
matches intent
shared field
signature-3d
hook0–6.9skinetic-buildicon·wireframe♪ node_lock
The tool is Whisper, uncanny transcription. The limit: on a long file it drifts and the timestamps wander off.
on-screen: Whisper drifts on long audio
expected on screen: black ground · torus hero in the shared Signal Field · Nuntius leads · mark · kinetic-build · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id shared Signal Fieldvisual markshape torusground blacktreatment wireframemotion kinetic-buildpower laser-lockinstrument laser-trace→Whisper drifts on long audio
2s2
matches intent
ChecklistCard
template
teach6.9–13.600000000000001skinetic-buildicon·wireframe♪ node_lock
So we chunk the audio first, then run WhisperX to force word-level alignment back onto the real waveform.
on-screen: Chunk + WhisperX align
expected on screen: black ground · a ChecklistCard panel over a dimmed Signal Field · Nuntius leads · node-graph · kinetic-build · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id ChecklistCardvisual node-graphshape torusground blacktreatment wireframemotion kinetic-buildpower morphinstrument morph+laser→Chunk + WhisperX aligncurate items, nodes
3s3
matches intent
DiffCard
template
build13.6–20.9stype-onicon·wireframe♪ node_lock
Now captions snap to each spoken word, so auto-generated subtitles look hand-tuned and your edits land on the syllable.
on-screen: Caption that snaps to the word
expected on screen: black ground · a DiffCard panel over a dimmed Signal Field · Nuntius leads · code · type-on · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id DiffCardvisual codeshape torusground blacktreatment wireframemotion type-onpower summoninstrument draw-on→Caption that snaps to the wordcurate codeLines, lines
4s4
matches intent
LegacyCard
template
proof20.9–27.7sreceipts-counticon·wireframe♪ node_lock
So take one long recording, chunk it, and run WhisperX. Compare the timestamps. The fix sells itself.
on-screen: Run WhisperX on a clip
expected on screen: black ground · a LegacyCard panel over a dimmed Signal Field · Nuntius leads · receipts · receipts-count · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id LegacyCardvisual receiptsshape torusground blacktreatment wireframemotion receipts-countpower receiptsinstrument spotlight→Run WhisperX on a clipcurate statsLabels
5s5
first render · fix pending
shared field
signature-3ddead_airgeneric_scene
futurist27.7–32.7sspatial-parallaxicon·wireframe♪ bed
The Stack
expected on screen: black ground · torus hero in the shared Signal Field · Nuntius leads · spatial-parallax · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id shared Signal Fieldshape torusground blacktreatment wireframemotion spatial-parallaxpower spatial-parallaxinstrument laser-fire→The Stack
6s6
first render · fix pending
shared field
signature-3ddead_airgeneric_scene
resolve32.7–37.7scoalescenceicon·wireframe♪ bed_out
The Stack
expected on screen: black ground · torus hero in the shared Signal Field · Nuntius leads · coalescence · coalescence · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id shared Signal Fieldvisual coalescenceshape torusground blacktreatment wireframemotion coalescencepower coalescenceinstrument coalescence→The Stack

Format stack 3 aspects · same scenes[], re-cropped

9:16
1080×1920
Stories · TikTok · YouTube Shorts · Reels
1:1
1080×1080
LinkedIn · Facebook · Instagram
16:9
1920×1080
X/Twitter · YouTube · LinkedIn video

Channels 9 destinations

LinkedInX/TwitterYouTubeInstagramFacebookThreadsTikTokPinterestBluesky

Social captions supplemental published copy · per channel (comp_id level)

tiktokThe tool: Whisper, for uncanny transcription. The limit: on a long file it drifts and the timestamps wander off. The fix: chunk the audio, then run WhisperX to force word-level alignment back onto the waveform. Captions snap to each word. (AI-assisted)
instagramThe Stack: Whisper. The limit: it drifts on long audio. The fix: Chunk the audio. Run WhisperX for word-level alignment. Captions snap to each spoken word. #whisper #transcription #captions #ai #buildinpublic
linkedinThe tool is Whisper, and the transcription is uncanny. The limit nobody mentions: on a long file it drifts, and the timestamps wander off the actual audio. The fix: chunk the audio first, then run WhisperX to force word-level alignment back onto the real waveform. Now captions snap to each spoken word, so auto-generated subtitles look hand-tuned and your edits land on the syllable. Take one long recording, chunk it, run WhisperX, and compare the timestamps. The fix sells itself.
xTool: Whisper. Uncanny transcription. Limit: on long audio it drifts and timestamps wander. Fix: chunk the audio, run WhisperX for word-level alignment. Captions snap to each word.
facebookThe tool is Whisper, and the transcription is uncanny. The limit is that on a long file it drifts and the timestamps wander off the audio. The fix: chunk the audio first, then run WhisperX to force word-level alignment back onto the real waveform. Now captions snap to each spoken word and look hand-tuned. Try it on one long recording and compare the timestamps.
threadsThe tool is Whisper, uncanny transcription. The limit: on a long file it drifts and the timestamps wander off. The fix: chunk the audio, then run WhisperX to force word-level alignment onto the real waveform. Captions snap to each spoken word. Try it on one clip and compare.
pinterestFixing Whisper transcription drift on long audio with WhisperX. The method: chunk the audio first, then run WhisperX to force word-level alignment back onto the real waveform so captions snap to each spoken word. A plain guide to accurate auto-generated subtitles and word-level timestamps.
blueskyTool: Whisper. Uncanny transcription. Limit: on long audio it drifts and timestamps wander. Fix: chunk the audio, run WhisperX for word-level alignment. Captions snap to each word.
youtubeTitle: The Stack: Fix Whisper's Drift on Long Audio with WhisperX The tool is Whisper, and the transcription is uncanny. The limit nobody mentions: on a long file it drifts and the timestamps wander off the actual audio. The fix: chunk the audio first, then run WhisperX to force word-level alignment back onto the real waveform. Now captions snap to each spoken word, so auto-generated subtitles look hand-tuned and your edits land on the syllable. Take one long recording, chunk it, run WhisperX, and compare the timestamps.

Cross-links every lens is a view on this one record