Back to Writing

Build Notes

Inside Hypit's Production Flow: From a Reference Video to a Finished Cut

How Hypit works end to end: six stages, three source kinds, the one command that spends money, and explicit reuse via build-record — plus the HyperFrames surprise.

Hypit does not generate video; it compiles it. Sources are planned into an immutable build definition, a Runtime satisfies each need, and Providers do the rendering. Once that clicks, the whole command surface and directory layout become coherent.

This teardown covers the six stages, how the three source kinds divide responsibility, the one command that spends money, and explicit reuse via build-record — including the trap that there is no implicit cache.

It ends on a surprise: Hypit's local renderer is not its own. It depends directly on HeyGen's open-source HyperFrames engine, pinned to a stale version.

In one line: Hypit's pipeline does not "generate a video" — it compiles one.

It treats video as a build artifact with a dependency graph. Source files (.svml / .svs / .svrun) are first planned into an immutable build definition; the Runtime then satisfies each "need" in that definition, one at a time; a Provider finally renders and produces the output. Once that clicks, the whole command set, the directory layout, and the ability to "change one line of dialogue and re-run only part of it" all become self-consistent.

Every claim below carries a source label: source means a repository file backs it up (path given), docs means the project's own documentation, and hands-on means I ran or re-checked it myself. This is the sequel to the earlier teardown. That piece covered what Hypit is; this one covers how it runs.

01 / Full pipeline

Six stages, each owned by a different party

The most common misreading: the coding agent does not render, and does not generate footage. It writes source, decides which services to use, and submits work. The real execution happens in the Runtime and the Provider.

Hypit's six production stages and who owns each Flowchart: 1 analyze the reference (agent + Skill) → 2 write Source (agent) → 3 plan the BuildDefinition (Core compiler) → 4 satisfy Need (Runtime + Provider) → 5 render and composite (local Provider renderer) → 6 export (the get command). 1 Analyze ref agent + Skill 2 Write Source agent .svml/.svs/.svrun 3 Plan Core compiler 4 Satisfy Need Runtime + Provider 5 Render Local Provider 6 Export hypit get Build Result (.hypit/results) — reusable by a later Run's <build-record> reference, for incremental reuse
Stages 3–5 are fully automatic. A human only shows up at stages 1 and 2, and at the "are we spending money?" decision point. The dashed arrow means stage 6's output flows back in as input to stage 3 — that is the single biggest cost saver in the design.
StageOwnerInputOutputWritten to
1 Analyze referenceagent + SkillReference video / one-line briefA read on the reference: shot and pacing judgmentsProject notes
2 Write SourceagentThat understanding + assets.svml / .svs / .svrunProject directory
3 PlanCoreSource closureImmutable BuildDefinition (with its Need list)Memory / Runtime
4 Satisfy NeedRuntime + ProviderA single NeedAsset / timing / component instance.hypit/runtimes/local
5 Render + compositeLocal ProviderComposition + timelineFrame sequence → encoded videoSame
6 ExportCLIBuild id + output nameThe finished video fileThe path you specify

The output of stages 3–6 is all recorded in the Build Result, which lands inside the project at .hypit/results (source: packages/video-cli/src/distribution.ts:19); runtime data lives at .hypit/runtimes/local (same file, :28). hands-on

02 / Three source types

Why split a video into three files

Not for looks. It is so that changing one thing does not drag everything else along with it.

FileWhat it declaresWhat editing it costs you
.svmlContent and structure: Script dialogue, visual components, tracks, the compositionDrags along components that depend on it, but assets already produced stay reusable
.svsAppearance parameter tables: named Recipes (layout, anchors, font sizes, motion)Appearance only, content untouched, re-renderable on its own
.svrunWhat this run should produce, and what it reusesDoes not change the work itself, only the scope of this execution

A real .svrun is three lines long — hands-on (examples/interview/swap-host.svrun):

<?svml using="@hypit/run-markup@1"?>
<svrun version="1">
  <author source="./swap-host.svml"/>
  <target output="final.video"/>
</svrun>

And .svs is a key-value table of named recipes: the caption box's position and alignment are first-class parameters — hands-on (examples/interview/recipes.svs):

caption.boy {
  stack-order: 70;
  x: 0.5;  y: 0.5;
  width: 0.88;  height: 0.22;
  anchor-x: center;  anchor-y: center;
  align: center;
}
This split is worth copying. Separating "content," "appearance," and "run scope" into three kinds of source file means high-frequency edits — a new font, a new palette, a new caption position — never touch the content at all, and therefore never require regenerating any paid asset. Most video tools mash all three into one project file, so changing a color forces a full re-render.
03 / Commands

Command-level flow: only build spends money

The official docs draw the line bluntly (docs/quickstart/run.md): “Only build submits work.” Everything else is either read-only or diagnostic. docs

hypit runtime use hypit.runtime.json   # pick the execution environment (once)
hypit plan  build.svrun                # show the work that would run, without running it
hypit build build.svrun --follow       # submit the build (this is where money is spent)
hypit get <build-id> --output final.video --to output/final.mp4

The helper commands around those four steps (packages/cli/src/output.ts:983-1022, source):

  • check verifies that a single Source is self-consistent; pricing reads the current Provider's price list (read-only, so you can do the math before spending).
  • builds / history / status --watch / logs / inspect: find things among existing Results, watch progress, read failure evidence.
  • activity / cancel: list active builds and withdraw one.
  • doctor: diagnose the external environment (I ran it; see the end of this post).

One design detail matters: submitting and observing are decoupled. The docs state plainly that "the terminal closing, or observation being interrupted, is a different thing from a Provider request inside the Build failing" (skills/hypit/references/production/builds.md) — docs. So a dropped --follow does not mean the build failed; hypit builds can still recover the result. That is genuinely useful for long renders.

04 / Incremental reuse

The cheapest part: Target, Candidate, and build-record

Of the whole pipeline, this is the section I rate highest on engineering value.

The Run source distinguishes two concepts — docs (skills/hypit/references/production/runs.md):

  • Target: the output this run must deliver (for example final.video).
  • Candidate: an output needed along the way that can also be supplied from outside.

The key move is that a Candidate can come from a previous Build result:

<author source="./production.svml"/>
<target output="final.video"/>
<build-record id="kept-take" build="bld_..." output="opening-semantic.take"/>
<satisfy output="opening-semantic.take" candidate="kept-take"/>

The official description: “changing a title can keep the performance and its timing while producing a new final video” — retitle the video and the already-recorded performance, along with its timing, survives; only a new final cut gets produced. docs

Why this beats a "cache," and why it is also more trouble. An ordinary cache keys on file hashes, so changing one character invalidates everything. What gets reused here is a semantic object (SemanticTake — "this performance and its word-level timing"). The docs are explicit: as long as the Script's identity still matches, both the media and its timing survive, which lets downstream captions and motion recompute instead of regenerate. The paid asset calls are skipped while the free layout math re-runs — the tradeoff points the right way.

One easy misjudgment to flag: Hypit has no implicit cache. In the source, resource ids look like res_<uuid>, and generated objects are only canonicalized — they do not participate in automatic cross-build matching. If you want reuse, the author has to write build-record explicitly in the Run, and a Forward does not copy bytes. source

The implication is direct: the money-saving switch is manual. Forget to write build-record and changing a single title regenerates the entire video from scratch. That is the habit this design most demands you build, and it is a real weakness compared with automatic dirty-tracking.

05 / Rendering

Rendering: evaluate a frame range, not "export the whole video"

Rendering is sliceable too — docs (skills/hypit/references/production/rendering.md):

<render:Video id="review" composition={main.composition} timeline={speech.timeline}
  start-frame="240" end-frame-exclusive="360"/>

Both bounds point at the original program's frame clock and must land on frame boundaries: 240 to 360 exclusive is 120 frames, which at 30fps is seconds 8–12. When you render only one range, the cost of frame captures and source-frame extraction both shrink — which is exactly what you want for "just take a look at second 8" review passes.

Concurrency is not hard-coded. Within its own reserved ceiling, the Provider adjusts capture concurrency to the current job (the docs' wording: “the Provider adjusts capture concurrency within its reserved ceiling using the current job's work”). The ceiling formula is at packages/provider-hyperframes-local/src/concurrency.ts:9-14: min(CPU-2, memory/2/1.5GiB) — source (in the previous post I disproved the "64 Chromium instances" claim; that number was simply this formula evaluated on the author's large machine).

06 / Three tracks

How captions, speech, and motion each flow

Captions: word anchors → timing → layout → presentation

The chain runs: word-level anchors in the Script (@name / @name! and friends) → alignment produces times → the caption package handles layout → the screen presents it. The relevant packages are layered by responsibility: script (parses dialogue and anchors), temporal / timeline-author (the timeline), caption and caption-fine (standard and fine-grained captions), whisperx (word-level alignment), fonts-open (fonts). source

Caption presentation parameters travel through Recipes rather than being inlined into the content — the caption.boy block above is the evidence: the caption box's normalized coordinates and alignment are reusable, theme-swappable data.

Speech: A-roll supplies the performance and the timing

The docs define A-roll as "the material supplying a spoken passage and its local timing" (“A-roll is the performance supplying a spoken passage and its local timing”), and stress that its picture may be cut away or shrunk into picture-in-picture while the audio and its semantic role persist — docs (skills/hypit/SKILL.md). Speech generation goes through a Provider (the built-in vendor mapping includes fishaudio, elevenlabs, and others).

Motion: triggered by word-anchor events, not by seconds

This is the real consumer side of "word anchoring." In the previous post I verified it line by line: a single @manifest! anchor is consumed in three places — a sound effect, an icon, and a flash — and none of the three contains a second count. Rewrite the dialogue, change the language, and all three follow the word. hands-on

Motion itself is driven by a Recipe's appearance / motion parameters (for example recipes.motion.card). It is declarative animation, not scripted animation. source

07 / Implementation truth

Its local renderer is really someone else's engine

This was the most surprising finding after mapping the pipeline, and it corrects the intuition that Hyperframes is an internal Hypit component.

Hypit ships four packages with "hyperframes" in the name, which makes it easy to assume the whole stack is in-house. In fact it is three in-house layers plus one upstream engine:

PackageRoleOwnership
@hypit/hyperframesCompiles a Composition into a Hyperframes documentIn-house at Hypit
@hypit/render-hyperframesSurface / capability layer (self-described as "is not a renderer")In-house at Hypit
@hypit/provider-hyperframes-localLocal execution layerIn-house at Hypit
hyperframes · @hyperframes/engine · @hyperframes/producerThe engine that actually does the workOpen source from HeyGen

The evidence is hard: Hypit's local Provider depends directly on external packages (packages/provider-hyperframes-local/package.json, hands-on):

"hyperframes": "0.7.101"
"@hyperframes/engine": "0.7.101"
"@hyperframes/producer": "0.7.101"

And the source imports the capture session straight from the upstream engine: import { getCdpSession } from "@hyperframes/engine" (src/opaque-capture.ts:2, hands-on). In other words, Hypit's stage-5 rendering capability is built on top of HeyGen's HyperFrames.

Who is upstream

I verified this project first-hand (hands-on, 2026-09-16):

  • Repository heygen-com/hyperframes, Apache-2.0 licensed, 50,496 stars / 4,607 forks, created 2026-03-10 and still receiving pushes that day.
  • Its self-description is a single line: “Write HTML. Render video. Built for agents.”
  • The npm package hyperframes is at 0.8.41, maintained by vance@heygen.com (official HeyGen).
Two actionable conclusions. ① A version gap: Hypit is pinned to 0.7.101 while upstream has reached 0.8.41 — if you want to play with this renderer yourself, going straight at HeyGen's engine is the better deal. ② The licence gap is much larger: HyperFrames is clean Apache-2.0, whereas Hypit's modified Apache licence forbids multi-tenant use and commercial redistribution. If "HTML renders to video" is also your goal, upstream carries none of those restrictions. The next post will work that comparison out properly.
08 / Tradeoffs

What this design buys, and what it costs

Design choiceBenefitCost
Video as a build artifactSliceable, reusable, verifiableMany concepts (Run/Target/Candidate/Need/Build); steep to learn
Content / appearance / run scope split three waysRestyling does not regenerate assetsYou maintain three source types plus Recipe tables
Word-anchored timingRewrite the script without re-timingSemantic anchors cannot express purely decorative motion
Provider abstractionSwap models / run local / bring your own keyExpensive to integrate; no *_API_KEY convention
Component packages installed at exact versionsReproducibleFirst-time setup is one install after another, and the errors do not tell you which one

This is a design that pays for maintainability up front: it charges a higher conceptual cost in exchange for "change one thing without redoing the entire video." If you only ever make one-off videos, the abstraction is a burden. If you need to produce hundreds or thousands of variants, it is the most serious answer I have seen.

Incidentally, its Studio can only step in after a build: hypit studio --run <build.svrun> — hands-on (hypit --help). So what you preview is a build that already exists, not a project you are editing. That is the inevitable consequence of a compiler-style architecture.

09 / Limits

What I could not verify

  • I never ran a real build. On this machine hypit doctor reports only 0 diagnostics, but that is conditional on selecting a Runtime Profile first; ffmpeg and paid models are missing, so for stages 3–6 I have only source and documentation evidence — no end-to-end acceptance. hands-on
  • doctor has a small trap: with no Profile selected it still reports "No problems found," because it only checks the Profile already selected. Do not read it as proof that the environment is complete. hands-on
  • For the full list of "Need" types I confirmed existence and responsibility from the docs and package structure, but did not enumerate them exhaustively.
  • I did not check the exact fields of the Build Result manifest file one by one.

In the next post I will pull those three tracks apart and set them beside tools like Remotion and Hyperframes, to work out a general, deployable automated editing pipeline.