Hypit does not generate video; it compiles it. Sources are planned into an immutable build definition, a Runtime satisfies each need, and Providers do the rendering. Once that clicks, the whole command surface and directory layout become coherent.
This teardown covers the six stages, how the three source kinds divide responsibility, the one command that spends money, and explicit reuse via build-record — including the trap that there is no implicit cache.
It ends on a surprise: Hypit's local renderer is not its own. It depends directly on HeyGen's open-source HyperFrames engine, pinned to a stale version.
In one line: Hypit's pipeline does not "generate a video" — it compiles one.
It treats video as a build artifact with a dependency graph. Source files (.svml / .svs / .svrun) are first planned into an immutable build definition; the Runtime then satisfies each "need" in that definition, one at a time; a Provider finally renders and produces the output. Once that clicks, the whole command set, the directory layout, and the ability to "change one line of dialogue and re-run only part of it" all become self-consistent.
Every claim below carries a source label: source means a repository file backs it up (path given), docs means the project's own documentation, and hands-on means I ran or re-checked it myself. This is the sequel to the earlier teardown. That piece covered what Hypit is; this one covers how it runs.
Six stages, each owned by a different party
The most common misreading: the coding agent does not render, and does not generate footage. It writes source, decides which services to use, and submits work. The real execution happens in the Runtime and the Provider.
| Stage | Owner | Input | Output | Written to |
|---|---|---|---|---|
| 1 Analyze reference | agent + Skill | Reference video / one-line brief | A read on the reference: shot and pacing judgments | Project notes |
| 2 Write Source | agent | That understanding + assets | .svml / .svs / .svrun | Project directory |
| 3 Plan | Core | Source closure | Immutable BuildDefinition (with its Need list) | Memory / Runtime |
| 4 Satisfy Need | Runtime + Provider | A single Need | Asset / timing / component instance | .hypit/runtimes/local |
| 5 Render + composite | Local Provider | Composition + timeline | Frame sequence → encoded video | Same |
| 6 Export | CLI | Build id + output name | The finished video file | The path you specify |
The output of stages 3–6 is all recorded in the Build Result, which lands inside the project at .hypit/results (source: packages/video-cli/src/distribution.ts:19); runtime data lives at .hypit/runtimes/local (same file, :28). hands-on
Why split a video into three files
Not for looks. It is so that changing one thing does not drag everything else along with it.
| File | What it declares | What editing it costs you |
|---|---|---|
.svml | Content and structure: Script dialogue, visual components, tracks, the composition | Drags along components that depend on it, but assets already produced stay reusable |
.svs | Appearance parameter tables: named Recipes (layout, anchors, font sizes, motion) | Appearance only, content untouched, re-renderable on its own |
.svrun | What this run should produce, and what it reuses | Does not change the work itself, only the scope of this execution |
A real .svrun is three lines long — hands-on (examples/interview/swap-host.svrun):
<?svml using="@hypit/run-markup@1"?>
<svrun version="1">
<author source="./swap-host.svml"/>
<target output="final.video"/>
</svrun>
And .svs is a key-value table of named recipes: the caption box's position and alignment are first-class parameters — hands-on (examples/interview/recipes.svs):
caption.boy {
stack-order: 70;
x: 0.5; y: 0.5;
width: 0.88; height: 0.22;
anchor-x: center; anchor-y: center;
align: center;
}
Command-level flow: only build spends money
The official docs draw the line bluntly (docs/quickstart/run.md): “Only build submits work.” Everything else is either read-only or diagnostic. docs
hypit runtime use hypit.runtime.json # pick the execution environment (once)
hypit plan build.svrun # show the work that would run, without running it
hypit build build.svrun --follow # submit the build (this is where money is spent)
hypit get <build-id> --output final.video --to output/final.mp4
The helper commands around those four steps (packages/cli/src/output.ts:983-1022, source):
checkverifies that a single Source is self-consistent;pricingreads the current Provider's price list (read-only, so you can do the math before spending).builds/history/status --watch/logs/inspect: find things among existing Results, watch progress, read failure evidence.activity/cancel: list active builds and withdraw one.doctor: diagnose the external environment (I ran it; see the end of this post).
One design detail matters: submitting and observing are decoupled. The docs state plainly that "the terminal closing, or observation being interrupted, is a different thing from a Provider request inside the Build failing" (skills/hypit/references/production/builds.md) — docs. So a dropped --follow does not mean the build failed; hypit builds can still recover the result. That is genuinely useful for long renders.
The cheapest part: Target, Candidate, and build-record
Of the whole pipeline, this is the section I rate highest on engineering value.
The Run source distinguishes two concepts — docs (skills/hypit/references/production/runs.md):
- Target: the output this run must deliver (for example
final.video). - Candidate: an output needed along the way that can also be supplied from outside.
The key move is that a Candidate can come from a previous Build result:
<author source="./production.svml"/>
<target output="final.video"/>
<build-record id="kept-take" build="bld_..." output="opening-semantic.take"/>
<satisfy output="opening-semantic.take" candidate="kept-take"/>
The official description: “changing a title can keep the performance and its timing while producing a new final video” — retitle the video and the already-recorded performance, along with its timing, survives; only a new final cut gets produced. docs
One easy misjudgment to flag: Hypit has no implicit cache. In the source, resource ids look like res_<uuid>, and generated objects are only canonicalized — they do not participate in automatic cross-build matching. If you want reuse, the author has to write build-record explicitly in the Run, and a Forward does not copy bytes. source
The implication is direct: the money-saving switch is manual. Forget to write build-record and changing a single title regenerates the entire video from scratch. That is the habit this design most demands you build, and it is a real weakness compared with automatic dirty-tracking.
Rendering: evaluate a frame range, not "export the whole video"
Rendering is sliceable too — docs (skills/hypit/references/production/rendering.md):
<render:Video id="review" composition={main.composition} timeline={speech.timeline}
start-frame="240" end-frame-exclusive="360"/>
Both bounds point at the original program's frame clock and must land on frame boundaries: 240 to 360 exclusive is 120 frames, which at 30fps is seconds 8–12. When you render only one range, the cost of frame captures and source-frame extraction both shrink — which is exactly what you want for "just take a look at second 8" review passes.
Concurrency is not hard-coded. Within its own reserved ceiling, the Provider adjusts capture concurrency to the current job (the docs' wording: “the Provider adjusts capture concurrency within its reserved ceiling using the current job's work”). The ceiling formula is at packages/provider-hyperframes-local/src/concurrency.ts:9-14: min(CPU-2, memory/2/1.5GiB) — source (in the previous post I disproved the "64 Chromium instances" claim; that number was simply this formula evaluated on the author's large machine).
How captions, speech, and motion each flow
Captions: word anchors → timing → layout → presentation
The chain runs: word-level anchors in the Script (@name / @name! and friends) → alignment produces times → the caption package handles layout → the screen presents it. The relevant packages are layered by responsibility: script (parses dialogue and anchors), temporal / timeline-author (the timeline), caption and caption-fine (standard and fine-grained captions), whisperx (word-level alignment), fonts-open (fonts). source
Caption presentation parameters travel through Recipes rather than being inlined into the content — the caption.boy block above is the evidence: the caption box's normalized coordinates and alignment are reusable, theme-swappable data.
Speech: A-roll supplies the performance and the timing
The docs define A-roll as "the material supplying a spoken passage and its local timing" (“A-roll is the performance supplying a spoken passage and its local timing”), and stress that its picture may be cut away or shrunk into picture-in-picture while the audio and its semantic role persist — docs (skills/hypit/SKILL.md). Speech generation goes through a Provider (the built-in vendor mapping includes fishaudio, elevenlabs, and others).
Motion: triggered by word-anchor events, not by seconds
This is the real consumer side of "word anchoring." In the previous post I verified it line by line: a single @manifest! anchor is consumed in three places — a sound effect, an icon, and a flash — and none of the three contains a second count. Rewrite the dialogue, change the language, and all three follow the word. hands-on
Motion itself is driven by a Recipe's appearance / motion parameters (for example recipes.motion.card). It is declarative animation, not scripted animation. source
Its local renderer is really someone else's engine
This was the most surprising finding after mapping the pipeline, and it corrects the intuition that Hyperframes is an internal Hypit component.
Hypit ships four packages with "hyperframes" in the name, which makes it easy to assume the whole stack is in-house. In fact it is three in-house layers plus one upstream engine:
| Package | Role | Ownership |
|---|---|---|
@hypit/hyperframes | Compiles a Composition into a Hyperframes document | In-house at Hypit |
@hypit/render-hyperframes | Surface / capability layer (self-described as "is not a renderer") | In-house at Hypit |
@hypit/provider-hyperframes-local | Local execution layer | In-house at Hypit |
hyperframes · @hyperframes/engine · @hyperframes/producer | The engine that actually does the work | Open source from HeyGen |
The evidence is hard: Hypit's local Provider depends directly on external packages (packages/provider-hyperframes-local/package.json, hands-on):
"hyperframes": "0.7.101"
"@hyperframes/engine": "0.7.101"
"@hyperframes/producer": "0.7.101"
And the source imports the capture session straight from the upstream engine: import { getCdpSession } from "@hyperframes/engine" (src/opaque-capture.ts:2, hands-on). In other words, Hypit's stage-5 rendering capability is built on top of HeyGen's HyperFrames.
Who is upstream
I verified this project first-hand (hands-on, 2026-09-16):
- Repository
heygen-com/hyperframes, Apache-2.0 licensed, 50,496 stars / 4,607 forks, created 2026-03-10 and still receiving pushes that day. - Its self-description is a single line: “Write HTML. Render video. Built for agents.”
- The npm package
hyperframesis at 0.8.41, maintained byvance@heygen.com(official HeyGen).
0.7.101 while upstream has reached 0.8.41 — if you want to play with this renderer yourself, going straight at HeyGen's engine is the better deal. ② The licence gap is much larger: HyperFrames is clean Apache-2.0, whereas Hypit's modified Apache licence forbids multi-tenant use and commercial redistribution. If "HTML renders to video" is also your goal, upstream carries none of those restrictions. The next post will work that comparison out properly.
What this design buys, and what it costs
| Design choice | Benefit | Cost |
|---|---|---|
| Video as a build artifact | Sliceable, reusable, verifiable | Many concepts (Run/Target/Candidate/Need/Build); steep to learn |
| Content / appearance / run scope split three ways | Restyling does not regenerate assets | You maintain three source types plus Recipe tables |
| Word-anchored timing | Rewrite the script without re-timing | Semantic anchors cannot express purely decorative motion |
| Provider abstraction | Swap models / run local / bring your own key | Expensive to integrate; no *_API_KEY convention |
| Component packages installed at exact versions | Reproducible | First-time setup is one install after another, and the errors do not tell you which one |
This is a design that pays for maintainability up front: it charges a higher conceptual cost in exchange for "change one thing without redoing the entire video." If you only ever make one-off videos, the abstraction is a burden. If you need to produce hundreds or thousands of variants, it is the most serious answer I have seen.
Incidentally, its Studio can only step in after a build: hypit studio --run <build.svrun> — hands-on (hypit --help). So what you preview is a build that already exists, not a project you are editing. That is the inevitable consequence of a compiler-style architecture.
What I could not verify
- I never ran a real build. On this machine
hypit doctorreports only0 diagnostics, but that is conditional on selecting a Runtime Profile first; ffmpeg and paid models are missing, so for stages 3–6 I have only source and documentation evidence — no end-to-end acceptance. hands-on doctorhas a small trap: with no Profile selected it still reports "No problems found," because it only checks the Profile already selected. Do not read it as proof that the environment is complete. hands-on- For the full list of "Need" types I confirmed existence and responsibility from the docs and package structure, but did not enumerate them exhaustively.
- I did not check the exact fields of the Build Result manifest file one by one.
In the next post I will pull those three tracks apart and set them beside tools like Remotion and Hyperframes, to work out a general, deployable automated editing pipeline.