Back Office · office.temerarii.xyz
One asset, all the way in — composition, the wireframe + storyboard, the output format stack, and the template, all read from the SAME content-index record. The expected output matches what /media surfaces for this post.
post longform-W44-Thukind longformweek W44date 2026-11-05campaign longform-youtubepillar emerging_techbeat asset videoduration 149.4sground blackscenes 8

Checklist the per-video bar — engine/sim

100.0/100
plain languagevo coverageno dead airuniquenesscaption fitcompletenesscleanliness
quantitative quality · weights learn from your reviews (engine.sim.memory review longform-W44-Thu good|bad)
✓ all static checks pass — one-focal/scene · tier-by-beat · one-track caption · colorway · cast+shape correct · no banned/fabricated. (audio + visual tiers verify on the rendered finals — Phase 2)

Composition comp · template family · expected output

composition LongFormChaptersfamily / template LongFormChapters
9:16 Reelpending1:1 Squarepending16:9 Widepending9:16 4Kpending1:1 4Kpending16:9 4KpendingGIF (SMS)pending
render pending — silent master not yet on disk
expected output: 0/7 rendered — same matrix the /media preview surfaces for this asset.

Composition layer × scene 8 scenes · 149.4s · comp_id + rendered still + tier + the script

#Layer (comp_id · still · tier)BeatTimecodeMotionLogoAudioVO / on-screen / caption
1s1
matches intent
shared field
signature-3d
open0–18.4sspatial-parallaxicon·liquid-chrome♪ bed_in
Today is content creation again, but the moving kind: video. Building presence in public means short clips, not just text, and people think video needs a crew. It does not. Here is how we turn one written idea into a finished social video with code, on a laptop, no studio.
on-screen: Turn a post into a video
expected on screen: black ground · cone hero in the shared Signal Field · Augur leads · node-graph · spatial-parallax · icon·liquid-chrome logo · caption bottom-left
spec (the prompt): comp_id shared Signal Fieldvisual node-graphshape coneground blacktreatment liquid-chromemotion spatial-parallaxpower summoninstrument summon→Turn a post into a video
2s2
matches intent
NumberedList
template
teach18.4–37.2skinetic-buildicon·ember-fill♪ node_lock
First step. The video does not start from scratch. It starts from the same script the agent already wrote for this idea, pulled from the source file. One idea, one script, then we point the video pipeline at it. The words you would have posted become the words the video says.
on-screen: Start from the same script
expected on screen: black ground · a NumberedList panel over a dimmed Signal Field · Augur leads · node-graph · kinetic-build · icon·ember-fill logo · caption bottom-left
spec (the prompt): comp_id NumberedListvisual node-graphshape coneground blacktreatment ember-fillmotion kinetic-buildpower morphinstrument morph+laser→Start from the same scriptcurate items, nodes
3s3
matches intent
ChecklistCard
template
teach37.2–55.6skinetic-buildicon·wireframe♪ node_lock
Second step. We make the voiceover with a text-to-speech tool, ElevenLabs, called through an MCP server. We hand it the script and one locked voice, and it returns a clean audio file. No microphone, no booth, no retakes. The same voice every day, so the channel sounds like one person.
on-screen: Make the voice with a model
expected on screen: black ground · a ChecklistCard panel over a dimmed Signal Field · Augur leads · node-graph · kinetic-build · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id ChecklistCardvisual node-graphshape coneground blacktreatment wireframemotion kinetic-buildpower morphinstrument morph+laser→Make the voice with a modelcurate items, nodes
4s4
matches intent
PipelineMap
template
teach55.6–75.8skinetic-buildicon·white-knockout♪ node_lock
Third step. The picture is code, not a camera. We use Remotion, which lets you build video out of plain web code, and render the frames to match the audio. Captions, motion, the brand look, all written once as a template and reused for every clip. Change the template, every future video changes with it.
on-screen: Render the frames with code
expected on screen: black ground · a PipelineMap panel over a dimmed Signal Field · Augur leads · node-graph · kinetic-build · icon·white-knockout logo · caption bottom-left
spec (the prompt): comp_id PipelineMapvisual node-graphshape coneground blacktreatment white-knockoutmotion kinetic-buildpower morphinstrument morph+laser→Render the frames with codecurate nodes, stages
5s5
matches intent
LogStream
template
teach75.8–96.3skinetic-buildicon·liquid-chrome♪ node_lock
Fourth step. Last, we stitch the voice and the picture into one file with ffmpeg, the free command-line tool that handles video. One command lays the audio under the frames and writes the finished clip. The whole chain, script to voice to frames to final file, runs from the terminal without a single timeline to drag.
on-screen: Mux audio and video together
expected on screen: black ground · a LogStream panel over a dimmed Signal Field · Augur leads · node-graph · kinetic-build · icon·liquid-chrome logo · caption bottom-left
spec (the prompt): comp_id LogStreamvisual node-graphshape coneground blacktreatment liquid-chromemotion kinetic-buildpower morphinstrument morph+laser→Mux audio and video togethercurate nodes, rows
6s6
matches intent
AnnotatedDiagram
template
teach96.3–115.8skinetic-buildicon·ember-fill♪ node_lock
The honest proof. This video, the one playing now, was built with exactly that chain. A script from a file, a model voice, code for frames, ffmpeg to finish. No camera was switched on. You can watch the other clips it made, all from the same template, at office dot temerarii dot xyz.
on-screen: This clip was built that way
expected on screen: black ground · a AnnotatedDiagram panel over a dimmed Signal Field · Augur leads · node-graph · kinetic-build · icon·ember-fill logo · caption bottom-left
spec (the prompt): comp_id AnnotatedDiagramvisual node-graphshape coneground blacktreatment ember-fillmotion kinetic-buildpower morphinstrument morph+laser→This clip was built that waycurate callouts, nodes
7s7
matches intent
KpiGrid
template
proof115.8–132.4sreceipts-counticon·wireframe♪ node_lock
Fifth step, and the real lesson. The win is not one video. It is the template. Build the look in code once, and every future clip is cheap. Script in, video out, same style, every day. That is how presence stays consistent without a crew.
on-screen: Code the template, reuse forever
expected on screen: black ground · a KpiGrid panel over a dimmed Signal Field · Augur leads · receipts · receipts-count · icon·wireframe logo · caption bottom-left
spec (the prompt): comp_id KpiGridvisual receiptsshape coneground blacktreatment wireframemotion receipts-countpower receiptsinstrument spotlight→Code the template, reuse forevercurate kpis, statsLabels
8s8
matches intent
shared field
signature-3d
resolve132.4–149.4scoalescenceicon·liquid-chrome♪ bed_out
The takeaway. Video is not a budget problem anymore. It is a script, a voice model, a code template, and one stitching command. Build the template once and feed it ideas forever. Go make your first clip, and see ours at office dot temerarii dot xyz.
on-screen: Make the moving proof yourself
expected on screen: black ground · cone hero in the shared Signal Field · Augur leads · coalescence · coalescence · icon·liquid-chrome logo · caption bottom-left
spec (the prompt): comp_id shared Signal Fieldvisual coalescenceshape coneground blacktreatment liquid-chromemotion coalescencepower coalescenceinstrument coalescence→Make the moving proof yourself

Format stack 1 aspects · same scenes[], re-cropped

16:9
1920×1080
X/Twitter · YouTube · LinkedIn video

Channels 2 destinations

YouTubeBlog

Social captions supplemental published copy · per channel (comp_id level)

youtubeTurn One Written Post Into a Finished Social Video With Code, No Camera People think video needs a crew. It does not. In this video we show how we turn one written idea into a finished social video using code on a laptop, no studio, no microphone. The chain, step by step: - Start from the same script. The video does not start from scratch. It starts from the script the agent already wrote for this idea, pulled from our source file. The words you would have posted become the words the video says. - Make the voice with a model. We make the voiceover with a text-to-speech tool, ElevenLabs, called through an MCP server. We hand it the script and one locked voice and get back a clean audio file. Same voice every day, so the channel sounds like one person. - Render the frames with code. The picture is code, not a camera. We use Remotion, which lets you build video out of plain web code, and render frames to match the audio. Captions, motion, and the brand look are written once as a template and reused. - Mux it together. We stitch the voice and the picture into one file with ffmpeg, the free command-line video tool. One command lays the audio under the frames and writes the finished clip. The honest proof: this exact video was built with that chain. A script from a file, a model voice, code for frames, ffmpeg to finish. No camera was switched on. You can watch the other clips it made at office.temerarii.xyz. The real lesson is the template. Build the look in code once, and every future clip is cheap. Script in, video out, same style, every day. That is how presence stays consistent without a crew. Keywords: Remotion, ElevenLabs, ffmpeg, text to speech, code video, content automation.

Cross-links every lens is a view on this one record