Tooling is not the hard part; licensing is. Split captions, voiceover and motion into seven stages and every stage has a mature, self-hostable option. What actually kills projects is the render layer's licence, dependencies widely mistaken for MIT, and word-level alignment.
This piece lays out the seven stages with acceptance assertions and fallbacks, grounds caption readability in first-party standards (the BBC and Netflix source values), and designs the cache keys that decide your monthly bill.
The most useful finding is counter-intuitive: for the motion render layer, the cleanest licence with 50k stars behind it is HeyGen's HyperFrames (Apache-2.0) — not Remotion.
Conclusion: the tooling is not the hard part. Licensing is.
Once you split "captions + speech + motion" into seven stages, every stage has a mature, free, self-hostable option, and assembling them is not difficult. What actually kills a project six months in is three things: the licence on the motion rendering layer (Remotion charges companies of 4 or more, and bills automation per render), several dependencies everyone assumes are MIT but are really GPL/LGPL, and word-level time alignment, the stage that is most often underestimated.
And the biggest payoff of this research was a counter-intuitive finding: at the motion rendering layer, the cleanest licence — validated by 50,000 stars — belongs to HeyGen's HyperFrames (Apache-2.0), not Remotion.
How to read the labels: first-party = official pages, repositories and registries I fetched myself; inference = my own judgment; not verified = explicitly not done. This article is a design proposal, not a test report — I did not run the whole pipeline on this machine (reasons at the end). The tool comparison extends the previous post, and the two pieces reference each other.
HyperFrames vs. Remotion: the licence gap is an order of magnitude
Read this section first. It decides whether you can turn a product into a business.
| Dimension | HeyGen HyperFrames | Remotion |
|---|---|---|
| Licence | Apache-2.0 | source-available (GitHub classifies it as NOASSERTION; the vendor itself says it is not OSI open source) |
| Free tier threshold | No commercial threshold | Individuals / for-profit organisations of ≤3 people; 4 or more must buy a company licence |
| How headcount is counted | — | Vendor's wording: partner and contractor headcount is counted together toward the 4-person threshold |
| Automation / batch billing | No per-render fee | $0.01 / render, $100/month minimum |
| Seat price | — | $25/seat/month (Creators) |
| Enterprise tier | — | from $500/month |
| Multi-tenant SaaS | Permitted (Apache-2.0) | Permitted with conditions; you may not let users upload their own Remotion project to be rendered |
| Community size | 50,496★ / 4,607 forks | 59,000+★ |
| Determinism promise | Vendor's wording: “same input, same frames, same output” | Frame-by-frame rendering; determinism is good |
I verified both projects first-hand (first-party, 2026-09-16):
- HyperFrames: repository
heygen-com/hyperframes, Apache-2.0, 50,496 stars / 4,607 forks, created 2026-03-10 and still receiving pushes that day; its self-description is “Write HTML. Render video. Built for agents.” - Remotion: the official pricing page (marked as updated 2026-09-15) spells out the free tier, the Automators tier at
$0.01/render + $100/mo, the Creators tier at$25/mo/seat, and Enterprise from$500/mo.
Four dependencies that bite
This part is independent of your stack choice. Anyone can step on it.
| Component | Common misconception | Reality | Impact |
|---|---|---|---|
piper1-gpl (TTS) |
Piper is MIT | The new piper1-gpl is GPL-3.0; the original rhasspy/piper is archived |
GPL copyleft risk; must be handled before any closed-source distribution |
edge-tts |
It is MIT | The main body is LGPLv3 | Safe to call as a separate process or service; static linking into a closed-source artifact needs attention |
| Coqui XTTS (TTS) | Open source, commercially usable | Coqui Public Model License; its terms page returned 404 when tested, so it cannot be verified | Unverifiable terms = cannot be relied on |
| FFmpeg | It is free, so use it however | LGPLv2.1+ by default; once built with --enable-gpl and libx264/libx265 it becomes GPL; --enable-nonfree builds are not redistributable |
The risk is not money, it is copyleft and the H.264/AAC patent pools |
Incidentally, the previous post's conclusion holds here too: Hypit itself is modified Apache-2.0 that forbids multi-tenant services and commercial redistribution, while the HyperFrames underneath it is clean Apache-2.0. If you are building a product, go upstream.
Seven stages, each with a verifiable exit
I designed this chain as seven stages, each with explicit inputs, artifacts, and acceptance assertions. The point of the assertions is that when something breaks you can localise it to a stage instead of re-running the whole thing.
| Stage | Input | Artifact | Acceptance assertion | Failure fallback |
|---|---|---|---|---|
| 2 Speech | Script segments | Audio file (wav 48k) | Loudness inside the target range; no clipping (true peak < −1 dBTP) | Retry with a different voice; if the whole pass fails, split and retry |
| 3 Align | Audio + known text | Word-level JSON | Word count equals the input word count; monotonically increasing with no overlaps; per-word confidence above threshold | Fall back to a sentence-level timeline and raise an alarm (never emit wrong timings silently) |
| 4 Captions | Word-level JSON | ASS / SRT | CPS ≤ the ceiling; ≤ N characters per line; ≤ N lines; inside the safe area | Automatic re-wrapping / splitting; still over the limit means an error |
| 5 Motion | Word-level JSON + structure | A frame-evaluable composition | The same frame number must yield the same picture (determinism) | Switch off non-essential motion; keep the captions |
| 6 Render | Composition + audio | mp4 | Duration matches; no black or dropped frames; resolution and aspect ratio correct | Lower concurrency; re-render at a lower resolution |
| 7 Accept | mp4 + assertion set | Pass/fail + a report | All automated assertions pass; spot-check the first and last frames and the captions by hand | Localise to the stage and re-run, not the whole video |
Caption readability: use first-party standards, not community folklore
Every number in this section was verified against first-party documentation, with the original wording kept.
BBC Subtitle Guidelines (portrait values are official)
First-party — I pulled the official BBC page and checked the raw HTML line by line. The changelog for that version explicitly notes newly added "size and position guidance for 9:16 portrait".
| Item | 16:9 landscape | 9:16 portrait |
|---|---|---|
| Safe area | Central 90% vertically, central 75% horizontally | Central 75% vertically, central 90% horizontally (swapped) |
| Line width | 68% of frame width | 90% of frame width |
| Characters per line (conversion guide) | 37 characters | ≈25 characters |
| Maximum lines | 2 lines | 3 lines |
| Line height (type size) | 7%–8% of frame height | 3.9%–4.5% of frame height |
Two key sentences from the BBC, quoted verbatim so my paraphrase cannot distort them:
"As a guide, the equivalent to 37 characters in a 75% width region of a
16:9 (landscape) video is 25 characters in a 90% width region of a
9:16 (vertical) video."
"For 9:16 video in portrait or vertical mode, this is reversed: subtitles
should not be placed outside the central 75% vertically and the central
90% horizontally."
On reading speed, the BBC gives 160–180 words per minute (that is 0.33–0.375 seconds per word).
Engineering consequence: converted for 1080×1920, a 90% horizontal safe area gives x ∈ [54, 1026] and a 75% vertical safe area gives y ∈ [240, 1680]. Those two pixel values go straight into your QC assertions (inference, derived from the BBC percentages).
Netflix Timed Text Style Guide (Simplified Chinese has hard constraints)
First-party — I fetched the Simplified Chinese specification itself. The rules most relevant to Chinese speech:
- 16 characters per line, at most 2 lines (SDH may stretch to 18 characters).
- Reading speed capped at 9 characters/second (7 for children; 11 for adult SDH).
- "Do not use commas or periods. Use one single space instead." — the original wording, and it runs against Chinese writing instincts, so it is easy to trip over.
- Ellipses use U+2026; U+22EF is not supported. No italics. Numerals half-width, and spell 1–10 as Chinese characters where possible. Weekdays must not use Arabic numerals:
星期2is wrong,星期二is correct. - Line shape prefers an inverted pyramid — short on top, long below — avoiding a top line with only one or two words.
A boundary note: Netflix's rules are written for delivery to Netflix, not as a universal platform standard. Using them as thresholds is "borrowing a stricter industry standard" — safe, but be clear it was not designed for short-form video. inference
Word-level alignment: text first, then synthesize the speech
Of the seven stages, this is the one most often underestimated and most capable of ruining how the finished video feels.
The core design decision: make alignment happen against known text, instead of asking ASR to guess the text.
| Route | Approach | Problem |
|---|---|---|
| ASR-first | Audio first → Whisper transcription + word-level timestamps | Whisper's word-level timestamps are heuristic (derived from token times and attention alignment) and are not guaranteed to match phoneme boundaries; it also rewrites text, drops words, and merges numbers with their units |
| script-first | Script first → TTS synthesis → forced alignment against the known text | Needs a real forced-alignment tool, one extra step; but the text is 100% correct and the time boundaries are reliable |
Usable forced-alignment tools: WhisperX (wav2vec2 phoneme alignment on top of Whisper; the community default), Montreal Forced Aligner (the academic standard; accurate, but you must supply a pronunciation dictionary and acoustic model), and ctc-forced-aligner (lightweight). Third-party sources — the positioning above comes from their official repository descriptions, and I did not run an accuracy comparison.
For the output contract, pin down a structure with few but strict fields:
{
"words": [
{ "w": "rewrite", "start": 1.240, "end": 1.605, "conf": 0.94 },
{ "w": "the", "start": 1.605, "end": 1.720, "conf": 0.88 }
],
"audio": { "path": "vo.wav", "sr": 48000, "lufs": -14.2, "truePeak": -1.4 },
"scriptHash": "sha256:..."
}
Motion must be driven by frame number, not by wall-clock time
This is the line between "can you make a video" and "can you make a reproducible video".
Video rendering is, at heart: evaluate frame n once and get one image. So the animation function must be a pure function: frame → picture.
// Correct: driven by frame number, same frame always gives the same picture
const t = frame / fps;
const y = interpolate(t, [0, 0.6], [40, 0], { easing: easeOut });
// Wrong: depends on wall-clock time, so the same frame differs on fast and slow machines
const t = (Date.now() - startTime) / 1000;
const y = spring(t); // not reproducible, and it can drop frames
Why requestAnimationFrame is wrong in video rendering: it is bound to the real clock, so if one frame takes too long during rendering the animation state jumps ahead, and the same source produces different pictures on different machines. That makes visual regression testing impossible and makes the diff from "change one line of dialogue" untrustworthy. inference (this is the general engineering consensus on render determinism)
The two animation paradigms each have their place:
- Timeline-driven: right for title cards, transitions, decorative motion — it only cares how much time has passed.
- Event-driven (word anchoring): right for caption highlighting, emphasis, B-roll cuts — it cares which word is being spoken. This is the highest-yield category for talking-head video, because it means "change the line" no longer requires "re-time everything".
The Hypit teardown in the previous post turns exactly that second category into a language-level primitive (the attachment polarity of @name / @name!). You can borrow the idea without its whole framework: define anchors over your word-level JSON, and have motion subscribe to anchors rather than to seconds.
Cache keys: what decides your monthly bill
This section is the economics of the whole pipeline. There is one goal: when you change a line of dialogue, re-run only the stages that must re-run.
The method is content addressing plus per-stage keys. Each stage's cache key is made of every effective input to that stage:
| Stage | What the cache key must include | Invalidated when the script changes? |
|---|---|---|
| 2 Speech | Text + voice id + rate/emotion parameters + model version + audio format | Yes (necessarily regenerated) |
| 3 Align | Audio hash + text hash + aligner version + language | Yes |
| 4 Captions | Word-level JSON hash + caption style version + platform spec | Yes |
| 5 Motion | Structural source hash + word-level JSON + asset versions | Partially (only the affected ranges recompute) |
| 6 Render | Composition definition hash + encoding parameters + resolution | Partially (only the affected frame ranges re-render) |
Key design rule: put the cheap stages before the expensive ones, and keep the expensive stage's inputs as few as possible. Stages 4 and 5 are pure local computation, so changing them a hundred times costs nothing; stage 2 is where the money goes. So:
- Do not put style parameters into stage 2's key — changing a caption color should not regenerate the voiceover. This is exactly the value of separating
.svs(appearance) from.svml(content) in the previous post. - The model version must go into the key — a vendor silently upgrading its model makes identical input produce a different voice, and that is the most insidious source of non-reproducibility.
- Rendering must support frame ranges (like Hypit's
start-frame/end-frame-exclusive), so a review pass can render only seconds 8–12.
build-record and the whole video regenerates. If your pipeline also uses explicit reuse, make it the default step in your revision workflow rather than something someone has to remember. Better still is automatic dirty propagation: let content hashes determine the invalidation scope, with no declarations from the author.
Six failure modes you will actually hit
| Failure | How to detect it | Fallback strategy |
|---|---|---|
| Audio and captions offset throughout | Cross-check the first word's time against the audio's first frame; measure against a known audio clip | Apply a global shift; if the offset is too large, fall back to a sentence-level timeline and raise an alarm |
| Chinese renders as tofu boxes | Check before rendering that the font resolves (fc-list / font files exist) |
Bundle a CJK font with the artifact instead of depending on system fonts |
| Chromium version drift causes visual regressions | Pin the browser version and diff against baseline frames in CI | Lock the version; treat the browser as a build dependency, not a system dependency |
| TTS is non-deterministic | Request the same input twice and compare audio hashes | Persist the first result to disk and reuse it; record it in the cache key |
| Captions hidden behind platform UI | Assert coordinate bounds against the BBC portrait safe area | Shift the whole block up into the safe position in the lower third |
| Loudness off target / true-peak clipping | Measure integrated loudness and true peak | Normalise loudness to the target value (platform targets like -14 LUFS) |
One caution on loudness targets: different platforms use different values, and I did not verify each platform's official wording one by one, so no specific numbers are given here. Go by the official documentation of the platform you actually publish to, and make it a configurable setting rather than a hard-coded constant. not verified
The smallest viable stack, and what it costs
| Stage | Recommendation | Licence | Why not something else |
|---|---|---|---|
| 2 Speech | A commercial TTS API (pay as you go) or edge-tts (isolated) | LGPLv3 (isolated) | Among open-source local options, piper1-gpl's GPL-3.0 complicates closed-source distribution |
| 3 Align | WhisperX / faster-whisper | BSD-2 / MIT | The most mature in the community; word-level accuracy is sufficient; MFA is more accurate but heavier to deploy |
| 4 Captions | A layout engine you write yourself, emitting ASS/SRT | Your own | ASS supports per-word highlighting; you implement the layout rules yourself against the standards above |
| 5+6 Motion and render | HeyGen HyperFrames | Apache-2.0 | No commercial threshold, no per-render fee, and an explicit determinism promise |
| 6 Encoding | FFmpeg (LGPL build) | LGPLv2.1+ | Avoid --enable-gpl so GPL copyleft does not spread |
Cost structure (inference, dependent on your volume): software licensing can be $0 (an all Apache-2.0/LGPL/BSD combination). The real costs are ① metered TTS, ② your own machine or cloud compute, and ③ if you pick Remotion and your team is ≥4 people, a continuing monthly or per-render fee.
I did not test this pipeline
- This machine has no ffmpeg, no TTS and no alignment library. I tried to install them but the sandbox does not permit writing to system directories, so I ran none of the stages. This article is a design, not a test report. not verified
- I did not benchmark WhisperX / MFA / ctc-forced-aligner against each other; I only describe where each sits.
- Loudness targets were not verified per platform, and I deliberately omit specific numbers.
- Remotion's multi-tenant details are quoted from the v5.0 Terms that the vendor marks as "Upcoming"; I could not fetch the text of the version currently in force, so confirm your use-case boundary with the vendor in writing. not verified
- Coqui XTTS's licence page returned 404 when tested, so the terms cannot be verified, which is why I do not recommend depending on it.
- All prices in this article are a 2026-09-16 snapshot. Licence policies change; check the official pages again before you commit.
The method here is "first-party evidence where it can be verified + clearly labelled inference + an honest list of gaps". Treat this design as a starting point, not a conclusion.