Back to Writing

Build Notes

An Automated Video Editing Pipeline: Captions, Voiceover and Motion

Seven acceptance-gated stages for captions, voiceover and motion, grounded in first-party numbers, plus the licence gap that makes HyperFrames safer than Remotion.

Tooling is not the hard part; licensing is. Split captions, voiceover and motion into seven stages and every stage has a mature, self-hostable option. What actually kills projects is the render layer's licence, dependencies widely mistaken for MIT, and word-level alignment.

This piece lays out the seven stages with acceptance assertions and fallbacks, grounds caption readability in first-party standards (the BBC and Netflix source values), and designs the cache keys that decide your monthly bill.

The most useful finding is counter-intuitive: for the motion render layer, the cleanest licence with 50k stars behind it is HeyGen's HyperFrames (Apache-2.0) — not Remotion.

Conclusion: the tooling is not the hard part. Licensing is.

Once you split "captions + speech + motion" into seven stages, every stage has a mature, free, self-hostable option, and assembling them is not difficult. What actually kills a project six months in is three things: the licence on the motion rendering layer (Remotion charges companies of 4 or more, and bills automation per render), several dependencies everyone assumes are MIT but are really GPL/LGPL, and word-level time alignment, the stage that is most often underestimated.

And the biggest payoff of this research was a counter-intuitive finding: at the motion rendering layer, the cleanest licence — validated by 50,000 stars — belongs to HeyGen's HyperFrames (Apache-2.0), not Remotion.

How to read the labels: first-party = official pages, repositories and registries I fetched myself; inference = my own judgment; not verified = explicitly not done. This article is a design proposal, not a test report — I did not run the whole pipeline on this machine (reasons at the end). The tool comparison extends the previous post, and the two pieces reference each other.

01 / The motion rendering layer

HyperFrames vs. Remotion: the licence gap is an order of magnitude

Read this section first. It decides whether you can turn a product into a business.

DimensionHeyGen HyperFramesRemotion
LicenceApache-2.0source-available (GitHub classifies it as NOASSERTION; the vendor itself says it is not OSI open source)
Free tier thresholdNo commercial thresholdIndividuals / for-profit organisations of ≤3 people; 4 or more must buy a company licence
How headcount is counted—Vendor's wording: partner and contractor headcount is counted together toward the 4-person threshold
Automation / batch billingNo per-render fee$0.01 / render, $100/month minimum
Seat price—$25/seat/month (Creators)
Enterprise tier—from $500/month
Multi-tenant SaaSPermitted (Apache-2.0)Permitted with conditions; you may not let users upload their own Remotion project to be rendered
Community size50,496★ / 4,607 forks59,000+★
Determinism promiseVendor's wording: “same input, same frames, same output”Frame-by-frame rendering; determinism is good

I verified both projects first-hand (first-party, 2026-09-16):

  • HyperFrames: repository heygen-com/hyperframes, Apache-2.0, 50,496 stars / 4,607 forks, created 2026-03-10 and still receiving pushes that day; its self-description is “Write HTML. Render video. Built for agents.”
  • Remotion: the official pricing page (marked as updated 2026-09-15) spells out the free tier, the Automators tier at $0.01/render + $100/mo, the Creators tier at $25/mo/seat, and Enterprise from $500/mo.
Monthly software licence cost for a 4-person team at 10,000 renders per month Bar chart: HeyGen HyperFrames costs $0 (Apache-2.0, no commercial threshold); a Remotion company licence costs $100/month (the Automators tier, $0.01 per render with a $100 monthly minimum, and 10,000 renders lands exactly on that minimum); Remotion on the seat model would be 4 × $25 = $100/month. 4-person team · 10,000 renders/month · monthly software licence cost (USD) HyperFrames Remotion (Automators) Remotion (4 seats) $0 $100/mo $100/mo The threshold gap is 3 vs 4 people: solo devs get both free; once you cross, Remotion is a recurring monthly cost. Vendor-published prices, fetched 2026-09-16.
The question is not whether $100 is expensive. It is that cost scales linearly with volume while revenue may not. For an automation product billed per render, $0.01/render goes straight onto gross margin.
The counter-argument, which has to be stated. Remotion's ecosystem is more mature: thicker documentation, Lambda cloud rendering, the Studio editor, and more complete component libraries and templates, so answers are easier to find when you get stuck. HyperFrames is younger (created 2026-03). If you are a one- or two-person team and your product is not billed per render, Remotion's free tier is entirely sufficient, and its engineering maturity is worth the cost. The reason to choose HyperFrames is licence freedom and predictable cost — not better features.
02 / Licence traps

Four dependencies that bite

This part is independent of your stack choice. Anyone can step on it.

ComponentCommon misconceptionRealityImpact
piper1-gpl (TTS) Piper is MIT The new piper1-gpl is GPL-3.0; the original rhasspy/piper is archived GPL copyleft risk; must be handled before any closed-source distribution
edge-tts It is MIT The main body is LGPLv3 Safe to call as a separate process or service; static linking into a closed-source artifact needs attention
Coqui XTTS (TTS) Open source, commercially usable Coqui Public Model License; its terms page returned 404 when tested, so it cannot be verified Unverifiable terms = cannot be relied on
FFmpeg It is free, so use it however LGPLv2.1+ by default; once built with --enable-gpl and libx264/libx265 it becomes GPL; --enable-nonfree builds are not redistributable The risk is not money, it is copyleft and the H.264/AAC patent pools
What to do in practice: treat both TTS and FFmpeg as standalone executables invoked through a subprocess, rather than linking them into your artifact. That is the most common and least painful way to isolate copyleft. But note: this is engineering convention, not legal advice. If you really intend to distribute closed-source commercially, have a lawyer read it.

Incidentally, the previous post's conclusion holds here too: Hypit itself is modified Apache-2.0 that forbids multi-tenant services and commercial redistribution, while the HyperFrames underneath it is clean Apache-2.0. If you are building a product, go upstream.

03 / Pipeline

Seven stages, each with a verifiable exit

I designed this chain as seven stages, each with explicit inputs, artifacts, and acceptance assertions. The point of the assertions is that when something breaks you can localise it to a stage instead of re-running the whole thing.

The seven-stage automated video editing pipeline Flowchart: 1 script (prose/JSON) → 2 spoken TTS → 3 forced word-level alignment → 4 caption layout to ASS/SRT → 5 motion and composition (HTML/JS driven by frame number) → 6 render and audio mux → 7 quality acceptance. The word-level timings from stage 3 feed both the captions in stage 4 and the motion in stage 5. 1 Script Prose / JSON 2 Voice TTS Audio + LUFS 3 Forced align Word JSON (key) 4 Captions ASS / SRT + CPS 5 Motion Frame-driven 6 Render + mux Encode + mix 7 QC gate Asserts + samples Stage 3's word-level timing is the hinge of the whole chain: captions and motion both consume it. Hold that stage's error down and the next two stages can be stable.
Of the seven stages, only 2 and 6 are obviously compute-hungry; stage 3 is accuracy-hungry. Stages 4 and 5 are pure computation, so you can change them a hundred times for free — which is why the design should push changes toward stages 4 and 5 whenever possible.
StageInputArtifactAcceptance assertionFailure fallback
2 SpeechScript segmentsAudio file (wav 48k) Loudness inside the target range; no clipping (true peak < −1 dBTP) Retry with a different voice; if the whole pass fails, split and retry
3 AlignAudio + known textWord-level JSON Word count equals the input word count; monotonically increasing with no overlaps; per-word confidence above threshold Fall back to a sentence-level timeline and raise an alarm (never emit wrong timings silently)
4 CaptionsWord-level JSONASS / SRT CPS ≤ the ceiling; ≤ N characters per line; ≤ N lines; inside the safe area Automatic re-wrapping / splitting; still over the limit means an error
5 MotionWord-level JSON + structureA frame-evaluable composition The same frame number must yield the same picture (determinism) Switch off non-essential motion; keep the captions
6 RenderComposition + audiomp4 Duration matches; no black or dropped frames; resolution and aspect ratio correct Lower concurrency; re-render at a lower resolution
7 Acceptmp4 + assertion setPass/fail + a report All automated assertions pass; spot-check the first and last frames and the captions by hand Localise to the stage and re-run, not the whole video
04 / Captions

Caption readability: use first-party standards, not community folklore

Every number in this section was verified against first-party documentation, with the original wording kept.

BBC Subtitle Guidelines (portrait values are official)

First-party — I pulled the official BBC page and checked the raw HTML line by line. The changelog for that version explicitly notes newly added "size and position guidance for 9:16 portrait".

Item16:9 landscape9:16 portrait
Safe areaCentral 90% vertically, central 75% horizontallyCentral 75% vertically, central 90% horizontally (swapped)
Line width68% of frame width90% of frame width
Characters per line (conversion guide)37 characters≈25 characters
Maximum lines2 lines3 lines
Line height (type size)7%–8% of frame height3.9%–4.5% of frame height

Two key sentences from the BBC, quoted verbatim so my paraphrase cannot distort them:

"As a guide, the equivalent to 37 characters in a 75% width region of a
16:9 (landscape) video is 25 characters in a 90% width region of a
9:16 (vertical) video."

"For 9:16 video in portrait or vertical mode, this is reversed: subtitles
should not be placed outside the central 75% vertically and the central
90% horizontally."

On reading speed, the BBC gives 160–180 words per minute (that is 0.33–0.375 seconds per word).

Engineering consequence: converted for 1080×1920, a 90% horizontal safe area gives x ∈ [54, 1026] and a 75% vertical safe area gives y ∈ [240, 1680]. Those two pixel values go straight into your QC assertions (inference, derived from the BBC percentages).

Netflix Timed Text Style Guide (Simplified Chinese has hard constraints)

First-party — I fetched the Simplified Chinese specification itself. The rules most relevant to Chinese speech:

  • 16 characters per line, at most 2 lines (SDH may stretch to 18 characters).
  • Reading speed capped at 9 characters/second (7 for children; 11 for adult SDH).
  • "Do not use commas or periods. Use one single space instead." — the original wording, and it runs against Chinese writing instincts, so it is easy to trip over.
  • Ellipses use U+2026; U+22EF is not supported. No italics. Numerals half-width, and spell 1–10 as Chinese characters where possible. Weekdays must not use Arabic numerals: 星期2 is wrong, 星期二 is correct.
  • Line shape prefers an inverted pyramid — short on top, long below — avoiding a top line with only one or two words.
A counter-intuitive but important corollary (inference): typical Chinese speech runs at roughly 5 characters/second, below the 9 CPS ceiling, so your CPS check will often be all green. The real bottleneck is the capacity limit of "16 characters × 2 lines = 32 characters" — long sentences must be split. And word-by-word highlighting eats additional width, so the practical usable count is often only 12–14 characters. Therefore: write the CPS assertion, but stress-test the "32 characters per cue + line overflow" case.

A boundary note: Netflix's rules are written for delivery to Netflix, not as a universal platform standard. Using them as thresholds is "borrowing a stricter industry standard" — safe, but be clear it was not designed for short-form video. inference

05 / Alignment

Word-level alignment: text first, then synthesize the speech

Of the seven stages, this is the one most often underestimated and most capable of ruining how the finished video feels.

The core design decision: make alignment happen against known text, instead of asking ASR to guess the text.

RouteApproachProblem
ASR-first Audio first → Whisper transcription + word-level timestamps Whisper's word-level timestamps are heuristic (derived from token times and attention alignment) and are not guaranteed to match phoneme boundaries; it also rewrites text, drops words, and merges numbers with their units
script-first Script first → TTS synthesis → forced alignment against the known text Needs a real forced-alignment tool, one extra step; but the text is 100% correct and the time boundaries are reliable

Usable forced-alignment tools: WhisperX (wav2vec2 phoneme alignment on top of Whisper; the community default), Montreal Forced Aligner (the academic standard; accurate, but you must supply a pronunciation dictionary and acoustic model), and ctc-forced-aligner (lightweight). Third-party sources — the positioning above comes from their official repository descriptions, and I did not run an accuracy comparison.

For the output contract, pin down a structure with few but strict fields:

{
  "words": [
    { "w": "rewrite", "start": 1.240, "end": 1.605, "conf": 0.94 },
    { "w": "the",     "start": 1.605, "end": 1.720, "conf": 0.88 }
  ],
  "audio": { "path": "vo.wav", "sr": 48000, "lufs": -14.2, "truePeak": -1.4 },
  "scriptHash": "sha256:..."
}
Three invariants you must check: ① the word count is strictly equal to the input word count (any mismatch means the aligner rewrote the text — raise an error rather than continuing); ② times are monotonically increasing and non-overlapping; ③ any word whose confidence falls below threshold raises an alarm. Turn those three into assertions and you will catch the vast majority of "why are the captions drifting?" problems.
06 / Motion

Motion must be driven by frame number, not by wall-clock time

This is the line between "can you make a video" and "can you make a reproducible video".

Video rendering is, at heart: evaluate frame n once and get one image. So the animation function must be a pure function: frame → picture.

// Correct: driven by frame number, same frame always gives the same picture
const t = frame / fps;
const y = interpolate(t, [0, 0.6], [40, 0], { easing: easeOut });

// Wrong: depends on wall-clock time, so the same frame differs on fast and slow machines
const t = (Date.now() - startTime) / 1000;
const y = spring(t);   // not reproducible, and it can drop frames

Why requestAnimationFrame is wrong in video rendering: it is bound to the real clock, so if one frame takes too long during rendering the animation state jumps ahead, and the same source produces different pictures on different machines. That makes visual regression testing impossible and makes the diff from "change one line of dialogue" untrustworthy. inference (this is the general engineering consensus on render determinism)

The two animation paradigms each have their place:

  • Timeline-driven: right for title cards, transitions, decorative motion — it only cares how much time has passed.
  • Event-driven (word anchoring): right for caption highlighting, emphasis, B-roll cuts — it cares which word is being spoken. This is the highest-yield category for talking-head video, because it means "change the line" no longer requires "re-time everything".

The Hypit teardown in the previous post turns exactly that second category into a language-level primitive (the attachment polarity of @name / @name!). You can borrow the idea without its whole framework: define anchors over your word-level JSON, and have motion subscribe to anchors rather than to seconds.

07 / Incremental re-runs

Cache keys: what decides your monthly bill

This section is the economics of the whole pipeline. There is one goal: when you change a line of dialogue, re-run only the stages that must re-run.

The method is content addressing plus per-stage keys. Each stage's cache key is made of every effective input to that stage:

StageWhat the cache key must includeInvalidated when the script changes?
2 SpeechText + voice id + rate/emotion parameters + model version + audio formatYes (necessarily regenerated)
3 AlignAudio hash + text hash + aligner version + languageYes
4 CaptionsWord-level JSON hash + caption style version + platform specYes
5 MotionStructural source hash + word-level JSON + asset versionsPartially (only the affected ranges recompute)
6 RenderComposition definition hash + encoding parameters + resolutionPartially (only the affected frame ranges re-render)

Key design rule: put the cheap stages before the expensive ones, and keep the expensive stage's inputs as few as possible. Stages 4 and 5 are pure local computation, so changing them a hundred times costs nothing; stage 2 is where the money goes. So:

  • Do not put style parameters into stage 2's key — changing a caption color should not regenerate the voiceover. This is exactly the value of separating .svs (appearance) from .svml (content) in the previous post.
  • The model version must go into the key — a vendor silently upgrading its model makes identical input produce a different voice, and that is the most insidious source of non-reproducibility.
  • Rendering must support frame ranges (like Hypit's start-frame / end-frame-exclusive), so a review pass can render only seconds 8–12.
A warning carried over from the previous post. Hypit's reuse is explicit: leave out build-record and the whole video regenerates. If your pipeline also uses explicit reuse, make it the default step in your revision workflow rather than something someone has to remember. Better still is automatic dirty propagation: let content hashes determine the invalidation scope, with no declarations from the author.
08 / Failures

Six failure modes you will actually hit

FailureHow to detect itFallback strategy
Audio and captions offset throughout Cross-check the first word's time against the audio's first frame; measure against a known audio clip Apply a global shift; if the offset is too large, fall back to a sentence-level timeline and raise an alarm
Chinese renders as tofu boxes Check before rendering that the font resolves (fc-list / font files exist) Bundle a CJK font with the artifact instead of depending on system fonts
Chromium version drift causes visual regressions Pin the browser version and diff against baseline frames in CI Lock the version; treat the browser as a build dependency, not a system dependency
TTS is non-deterministic Request the same input twice and compare audio hashes Persist the first result to disk and reuse it; record it in the cache key
Captions hidden behind platform UI Assert coordinate bounds against the BBC portrait safe area Shift the whole block up into the safe position in the lower third
Loudness off target / true-peak clipping Measure integrated loudness and true peak Normalise loudness to the target value (platform targets like -14 LUFS)

One caution on loudness targets: different platforms use different values, and I did not verify each platform's official wording one by one, so no specific numbers are given here. Go by the official documentation of the platform you actually publish to, and make it a configurable setting rather than a hard-coded constant. not verified

09 / Stack

The smallest viable stack, and what it costs

StageRecommendationLicenceWhy not something else
2 SpeechA commercial TTS API (pay as you go) or edge-tts (isolated)LGPLv3 (isolated)Among open-source local options, piper1-gpl's GPL-3.0 complicates closed-source distribution
3 AlignWhisperX / faster-whisperBSD-2 / MITThe most mature in the community; word-level accuracy is sufficient; MFA is more accurate but heavier to deploy
4 CaptionsA layout engine you write yourself, emitting ASS/SRTYour ownASS supports per-word highlighting; you implement the layout rules yourself against the standards above
5+6 Motion and renderHeyGen HyperFramesApache-2.0No commercial threshold, no per-render fee, and an explicit determinism promise
6 EncodingFFmpeg (LGPL build)LGPLv2.1+Avoid --enable-gpl so GPL copyleft does not spread

Cost structure (inference, dependent on your volume): software licensing can be $0 (an all Apache-2.0/LGPL/BSD combination). The real costs are ① metered TTS, ② your own machine or cloud compute, and ③ if you pick Remotion and your team is ≥4 people, a continuing monthly or per-render fee.

The selection advice, in one line. A one- or two-person team not billing per render → Remotion's free tier, and enjoy the more mature ecosystem. A team of ≥4 people, or a multi-tenant / per-render automation product → HyperFrames: clean licence, predictable cost.
10 / Limits

I did not test this pipeline

  • This machine has no ffmpeg, no TTS and no alignment library. I tried to install them but the sandbox does not permit writing to system directories, so I ran none of the stages. This article is a design, not a test report. not verified
  • I did not benchmark WhisperX / MFA / ctc-forced-aligner against each other; I only describe where each sits.
  • Loudness targets were not verified per platform, and I deliberately omit specific numbers.
  • Remotion's multi-tenant details are quoted from the v5.0 Terms that the vendor marks as "Upcoming"; I could not fetch the text of the version currently in force, so confirm your use-case boundary with the vendor in writing. not verified
  • Coqui XTTS's licence page returned 404 when tested, so the terms cannot be verified, which is why I do not recommend depending on it.
  • All prices in this article are a 2026-09-16 snapshot. Licence policies change; check the official pages again before you commit.

The method here is "first-party evidence where it can be verified + clearly labelled inference + an honest list of gaps". Treat this design as a starting point, not a conclusion.