Back Office · office.temerarii.xyz
Where the encoder lives, who owns it, and what happens when it dies

The show page says what the format is. The stream guide says what the settings are. This says where the machine lives, who owns it, and what happens when it dies.

Why the encoder is not in the room

 Number
Measured local upload5.50 Mbps
Safe streaming ceiling (~80% of measured upload)~3,500–4,000 kbps
Target bitrate6,000 kbps
That gap does not close with settings. There is no encoder preset, codec or tuning option that sends more data than the connection carries. Moving the encoder into a datacenter is arithmetic, not preference — and every other decision on this page follows from it.

The second reason is bigger than the first. With the encoder in a datacenter, the broadcast survives the operator's machine. A local crash, a power cut, a hotspot dropping — none of it reaches air. The show keeps running and whoever was driving reconnects.

The signal path

PARTICIPANTS                 CLOUD MASTER                    DESTINATIONS
(anywhere, any device)       (datacenter, GPU)               (6, two aspects)

  browser  ──────────┐
  browser  ──────────┼────▶  ingest ─▶ compose ─▶ encode ─┬─▶ YouTube   16:9
  phone    ──────────┤          │         │              ├─▶ Twitch    16:9
  browser  ──────────┘          │         │              ├─▶ Kick      16:9
                                │         │              ├─▶ X         16:9
  OPERATOR ─── remote ctrl ─────┘         │              ├─▶ LinkedIn  16:9
  PRODUCER ─── remote ctrl ───────────────┘              └─▶ TikTok/IG 9:16
                                                                             
                                └─▶ local record (crash-safe) ─▶ clip pipeline
Read the operator and producer lines carefully. Neither of them is in the video path. Your machine is a remote control, not the encoder — which is why a 5.5 Mbps connection is irrelevant once this is running, and why two people in two countries can drive the same show.

Three control surfaces, not one

Controlling the programme and controlling the participants are different problems. Being honest about which are built is how a Producer knows what they actually have on show day.

SurfaceCoversStatus
ProgrammeScenes, layouts, overlays, sound effectsBuilt, verified against a live OBS
ParticipantsTheir camera, microphone, resolution and bitrate — through the guest platform’s directorNot yet built. Guest links are generated; nothing controls them
Their own co-streamThe OBS on their desk pushing to their own channelUncontrollable by design — coaching only
Remote control reaches a guest’s browser feed. It does not reach their lighting, their room, or a laptop that is thermally throttling. Those are settled by the per-device setup sheets before air. A Producer burning ten live minutes trying to fix a dark room from another continent is a preventable failure, and knowing the boundary is what prevents it.

Transport — what travels how

These boundaries get relitigated constantly, so they are written down once here.

NeedTransportWhy, and the trap
Participants → serverWebRTC (guest links)Crosses the internet; guests install nothing
Producer → broadcast controlHTTPS to the relayA few hundred bytes per command, so it survives a bad network
Producer → GUI workParsecLayout, installs and licences only. It streams video of a screen, so it degrades exactly when the network is worst — never operate through it
Source to source on the boxNDI (optional)Local network only. NDI never crosses the internet
Server → platformsRTMP / SRT simulcastSix destinations, two aspects
NDI on the box. WebRTC across the world. HTTP for control.

Scene layouts are solved, not drawn

The landscape collection is 37 scenes. The multi-person layouts in it were computed, not positioned by hand. Roughly fifteen of the sixty-five weekday episodes carry a big cast and the rest are small, so authoring a layout per headcount produces twenty-odd near-identical scenes that drift apart the first time somebody nudges a gutter. Given a headcount the grid is arithmetic, so it is computed — and the design effort goes where judgement genuinely beats maths: a solo host, a screen share, the hot seat.

The solver exists to prevent one specific thing: twenty-four faces on a 1920×1080 canvas. That fails three ways at once, and only the first is visible in a preview window.

CeilingWhere it bitesWhy no setting closes it
LegibilityA nameplate needs a 560px cell to carry name and role, 380px for a name alone, 260px for a short label. Below 260px nothing honest fits, and a 24-up grid lands at 273px.Text needs roughly 16px to survive streaming compression, and 16px inside a 300px cell is about six characters. Shipping a plate nobody can read is worse than shipping none.
Bitrate per faceThe encoder divides the 6,000 kbps target from the top of this page across every region that is moving. Four faces get ~1,500 kbps each; twenty-four get ~250, and a moving face visibly falls apart below ~500.The constraint is division. There is no preset that gives twenty-four faces the bitrate of four. Twenty-four blocky faces is a worse picture than four clean ones.
Decode countEvery visible feed is a separate video decode, on the same box doing the encode. Past roughly eight, the encoder starves.OBS reports it as “skipped frames (encoding lag)”, which reads as a CPU problem and is not one. So the natural response — lowering the encoder preset — achieves nothing, and the ten minutes spent trying it are live minutes.
The solver returns the geometry and what it costs. It will still hand back a 24-up grid, because the “everyone is here” beat is occasionally worth ten seconds — but it says which ceilings that crosses, and it says when to take a composited room view as one source or cut to pods instead. A layout that is past a ceiling is a decision, not an accident.

Sixteen faces is a beat, not a layout

Sixteen is the hard cap on faces in one frame. Sixteen fit geometrically at 352px cells; past that the grid is thumbnails. But sixteen is already twice the decode ceiling and under the bitrate floor, which is the point — it is the ceiling for a brief roll-call taken as one composited room source, not for a sustained layout of sixteen individual feeds.

A big cast is shown in pods of four to six, grouped by track. Everyone stays legible, the decode count stays inside the ceiling, and the grouping carries meaning rather than being an arbitrary slice of a roster. The wall of twenty-four is a transition, not a place to sit.

Screens set the layout; faces take what is left

A face at 350px is recognisable. Code at 350px is not readable at any font size, and an audience watching somebody build has come for the screen. So the screen is placed first and the faces take the remainder — never the other way round.

The arithmetic is short. A 1920-wide desktop rendered into a fraction of the frame is scaled by that fraction, so a 14px editor font arrives on stream at 14 × the fraction. Streaming compression needs about 12px to stay legible:

required source font = 12 / (screen’s fraction of the frame)
ModeScreen, as a fraction of frameFont the sharer must set
Corners — screen near full frame, faces as picture-in-picture at the edges80%15px
Rail-right — screen left, up to three faces stacked down the side64%19px
Rail-bottom — screen on top, faces across the bottom as a reacting panel53–61%20–23px
Dual — two screens side by side, two people racing the same build44%28px
Which means the fix is not a layout change. A default 14px editor needs the screen at 86% of frame — wider than any layout that also holds people. At every multi-person shape the participant has to raise their own font. So this is a hard setup requirement, not advice: editor at 20px or more, terminal at 18px or more, and share a 1920×1080 region rather than a 4K desktop. A producer cannot recover an unreadable screen from the gallery, and the audience reads it as our fault.

The virtual set — production value around the people

On a remote show you cannot fix the cameras. Twelve to twenty-four unknown webcams, in unknown rooms, on three continents, and no amount of layout work improves a poor sensor in poor light. So the production value has to come from around the people rather than from the people. That is what a physical studio’s set does, and the virtual set is the remote equivalent of it.

Mechanically it is one browser source per layout. It sits above the cameras and is opaque everywhere except where a camera should show through, with a hole cut for each one. Everything a physical set would do — the surround, the vignette, the texture — happens in the region outside those holes, because a texture crossing a camera edge is what makes a composite look assembled.

The holes and the camera positions are generated from one slot declaration and checked against each other at build time. Two sources of truth would guarantee a set whose windows do not line up with the faces in them, and that is the kind of fault nobody notices in a preview and everybody notices on air.

Light is the primary colourway

Black script, red accent, white ground. Two reasons, and the first is a constraint rather than a preference.

The mark only composites correctly on a light ground. The supplied transparency was cut with a global white key, which removed the background and the white fill inside the tail along with it. On white the ground shows through and the mark reads correctly; on a dark field the tail loses its fill and the black tail text sits on a dark field. Dark scenes therefore carry no mark at all until a dark-ground lockup exists — a missing logo is invisible, a black script on a black field is a defect people screenshot.

The second reason is editorial. This is a hiring programme, not a gaming stream. A light ground reads as daytime and as editorial rather than as a late-night broadcast, and it is a more welcoming frame for people who are being judged on their own webcams in their own houses. The dark ground stays available as a second colourway; it is not the default.

A colourway is not only a stylesheet, and this one bit. Every overlay is a browser source and picks its ground up from generated tokens — but a colour source is a native input holding a literal value. So when the colourway flipped to light, the background source stayed black, and every scene without a set rendered light panels on a black field. It was invisible in the scenes that have a set, because the set covers it. Anything holding a literal colour has to be regenerated with the tokens, not alongside them.

What renders what

Three tools, and choosing the wrong one is expensive in licence fees, in specialist time, or in something else that can fail while live.

LayerToolBecause
Overlay graphics — lower third, ticker, scoreboard, HUD. They coexist with the pictureHTML browser sourcesData-driven: they read the system’s own JSON. Free, version-controlled, styled from brand tokens. Two-dimensional only
Motion graphics — titles, stings, transitions, reveals, credits. They replace the pictureRemotion, pre-renderedRendered once and played as a media source, so they cost nothing at runtime. Code-native and on the existing render path
Real-time 3D bound to live dataWASP3D — only if genuinely requiredCosts a licence, Windows, a specialist and another machine that can fail live. Its scenes are authored in a GUI, so they do not version-control like everything else here
HTML is for information; motion graphics are for spectacle. This show’s premise is that the record is read from the system — that is information, so HTML is the correct primary tool rather than a compromise.

Graphics fire from the system, not from a person

The Producer cuts cameras. Everything else triggers itself: the sting from the scene change, the segment card from the calendar, the gate verdict from the gate emitting a result.

A show whose premise is that the record is read from the system must not have a person manually typing the record onto the screen. That would quietly contradict the format. It also keeps the Producer’s attention on twenty-four people across three continents, which is where it belongs.

HUD for everyone; AR only at the host position

HUD is two-dimensional, locked to the screen, and shows state. It makes no claim to exist in the room — which is the honest and correct instrument for a dozen unknown webcams with no tracking, no depth data and no controlled lighting.

AR means graphics that appear to occupy the physical space of the shot. It is viable at the host position only, because that is the one controlled machine: locked camera, known lighting, a capable GPU.

What makes it workDetail
A locked camera removes the need for trackingOne-time perspective calibration does the same job. Tape the tripod — the illusion dies the instant the camera shifts
Occlusion comes from layering, not the rendererAR objects above the host, environment below, and background removal cuts the host out between them
It renders in the browserSame toolchain as every overlay, reading the same data, version-controlled
A green screen beats automatic segmentation by a wide margin on hair and movement, and is the single highest-leverage physical purchase if AR is taken seriously.

Who owns what

ThingHeld byWhy
The master machine and its accountExecutive ChairmanOwner-level, for the whole season
Platform credentials and stream keysExecutive ChairmanOwner-level on every destination
Operator access to the broadcastProducerEverything needed to run the show, nothing that ends it permanently
Scene layouts, encoder profiles, overlaysProducer, in version controlReviewable, and recoverable without them
The person who can end the broadcast must not be one of the people competing in it. That is not distrust — it is the same principle as a Producer never self-approving an episode. Every access decision on this page follows from it, including why credentials are never held during selection.

What is automated, and what is not

Worth being precise, because this is routinely overstated in both directions.

AutomatableHuman only
Scene switching, source creation, text updates, filter toggles, start/stop stream and record, reading stats — all of it over the WebSocket control APIInstallers. Nothing clicks through a setup wizard for you.
Overlay rendering, telemetry, publishing, the clip pipeline, firewall and proxy config, the whole rebuild scriptLicence activation and platform logins
Encoder profiles applied from version-controlled filesVisual scene layout — deciding where a camera sits on screen
The control API is built into the broadcast software already. No plugin, no licence, documented protocol. That is what makes a remote producer and an automated overlay layer possible at zero cost — and it is the first thing built, because nothing built on it is wasted regardless of what else gets chosen.

When it breaks: replace, never repair

The server is disposable by design. Everything that makes it the broadcast machine lives in version control, and a single script rebuilds it from nothing.

The drill is part of pre-season, not a theory. Terminate the instance deliberately, redeploy, run only the script, and confirm the show comes back. If it needs one manual step, the script is incomplete — and you will find that out on a Tuesday otherwise.

Telemetry — the broadcast reports into this office

Stream state, bitrate, dropped and missed frames, encoder load, per-participant connection health, what is on air, and which episode it is. The broadcast machine pushes that out; nothing here reaches into it.

RuleWhy
Push, never pullThe office is gated, but its build feeds a public sandbox. A pull would create a path from a public artefact toward a live production machine.
Honest empty statesNo report means the page says no data and when it last heard anything. A stale number rendered as current is worse than a blank, because someone acts on it.
Why it goes here rather than a separate dashboard. The Producer's live status view, the incident log that pay decisions are settled from, and the elimination question “were you the bottleneck” all need the same facts. Machine timestamps beat memory, and putting them anywhere else means a second place to look.

A pilot broadcast comes before season one

A pilot is produced before Episode 1 — the season dates are on the show page and are not repeated here. The pilot is not a rehearsal of the format. It exists to find the three things this system does not yet have, at a point where the cost of finding them is one broadcast rather than an episode of the season. The pilot date is not yet set.

Not yet builtWhat the pilot has to establish
Audio architecture for many open microphonesThe soundboard exists and the per-source filter chain is documented in the stream guide. What does not exist is a bus design for twelve to twenty-four microphones open at once: routing, who hears whom, how the public/private split behaves under load, and what the mix does when eight people talk together. None of that is yet measured.
A confidence feedThe producer drives the programme through the control relay, and uses Parsec for layout. Neither gives a low-latency picture of what is actually going to air. A producer cutting a show they cannot see is cutting blind, and the transport table above is the reason a screen-sharing tool is not the answer.
Guest controlNamed as not built under the three control surfaces above: guest links are generated and nothing controls the camera, microphone, resolution or bitrate behind them. The pilot establishes how much of that can be settled by the setup sheets alone, and what genuinely needs a control path.
Do not read the scene collection as readiness. Thirty-seven scenes verified against a live OBS says the programme surface works. It says nothing about twenty-four live microphones, nothing about a producer’s ability to see their own output, and nothing about controlling a guest’s device. Those three are the pilot’s whole purpose, and claiming otherwise before it has run would be the kind of overstatement that costs an episode.

Decisions still open

Open3D graphics engine, and therefore the operating system.
One graphics product would force Windows and a licence cost. Its free tier is claimed but not verified. Until it is, the plan builds the control layer and HTML-based overlays first — those run anywhere, cost nothing, and cover the large majority of what is wanted. Roughly two graphics out of thirty-three genuinely need a 3D engine.
OpenReal cost of the instance.
GPU hourly rate plus operating-system licensing billed on top, priced against the exact card rather than an estimate. This changes whether the schedule is affordable, so it is answered before anything is provisioned.
OpenParticipant telemetry is deliberately NOT in scope.
Typing speed, process monitoring and biometric readouts would need software installed on twenty-four personal machines and broadcast publicly. The signed release covers likeness and screen share; it does not cover keystroke or biometric capture. The same on-air tension is available from work product — commits, gates, renders, publishes — which we already emit, from machines we own, with no consent question attached.

Rig settings, audio chain and the pre-flight gate · the Producer's operational runbook · the show