Skip to content

perf(producer): chunks read remote videos from their link instead of copying them - #5449

Closed
robinebers wants to merge 1 commit into
heygen-com:mainfrom
robinebers:feat/remote-video-range-reads
Closed

robinebers wants to merge 1 commit into
heygen-com:mainfrom
robinebers:feat/remote-video-range-reads

Conversation

@robinebers

@robinebers robinebers commented Oct 11, 2026 •

Copy link
Copy Markdown

What

When a render is split into chunks (the AWS Lambda path), a <video> given as an HTTPS link is no longer downloaded by the planner and copied to every chunk. Each chunk reads only the part of the video it shows, from the link.

Local renders on one machine are unchanged.

Why

Today the planner downloads the whole video, and every chunk that shows it gets its own copy, even when the chunk only shows a few seconds. With long recordings, that copying is most of the render's time and cost.

On a 20-minute 4K edit made from 50 minutes of screen and camera recordings, rendered on AWS Lambda:

main This PR
Time 21.5 min 11.2 min
Cost $7.00 $3.38
Picture Same

Related work

Refs #5290. Its notes leave this as a follow-up: "Reading only the byte ranges a chunk needs would cut that and is left for a follow-up."

How

  1. The planner checks each video link once. It asks for the first 512 bytes. The link is read in place only if all of these hold:

    • It's public HTTPS and passes the downloader's existing safety rules.
    • It answers with just those bytes (status 206), without redirecting.
    • It has a strong ETag (a version tag for the file).
    • It's an MP4, MOV, Matroska or WebM file. A playlist (HLS, DASH) is downloaded instead, because FFmpeg would fetch its segments itself, outside the relay.

    Anything else is downloaded exactly as it is today.

  2. The plan stores the link and its ETag instead of a file path, so nothing is copied to the chunks.

  3. FFmpeg never reads the link itself. Each process (the planner and each chunk) starts a small relay on 127.0.0.1, and FFmpeg reads the video from there. The relay fetches the real link and:

    • checks every redirect with the downloader's rules (HTTPS only, no private addresses);
    • sends the ETag as If-Match on every request, so a file replaced mid-render fails instead of mixing old and new frames;
    • if the link fails partway through (a 503, or a 412 because the file changed), fails the FFmpeg run that was reading it. FFmpeg alone would exit cleanly with the frames it had;
    • asks the link for the bytes FFmpeg wants in growing pieces (1 MB, 2 MB, 4 MB, then 8 MB at a time), so when FFmpeg stops reading, the transfer stops within one piece.
  4. The chunk's browser is told not to fetch those exact links. The chunk supplies the frames itself, as it already does for extracted videos.

  5. The link is checked once per render. Compile, probe, frame extraction and audio all share one answer, so they always agree on which links are read in place.

Review

  • Code review: GPT-6 Astra. It found four problems in the first version, all fixed here:
    1. FFmpeg could follow a later redirect to a plain-HTTP or private address. Now FFmpeg only talks to the relay, which checks every hop.
    2. If the file changed, one fallback path quietly read the new file without If-Match. Now every read goes through the relay with If-Match, so it fails instead.
    3. Audio and the duration probe read the link without If-Match, so old video could be paired with new audio. Same fix.
    4. The chunk's browser blocked a link with any query string, which could also block a different signed link to the same file. Now it blocks exactly that link, query included (checked in headless Chrome).
  • Second code review. It found two more problems and a test bug, all fixed:
    1. An HLS playlist passed the link check, and FFmpeg would then fetch its segments directly, around the relay's checks. Now only MP4, MOV, Matroska and WebM files are read in place.
    2. A link that failed partway through a read (a 503, or a 412) gave a render that "succeeded" with missing frames: 128 of 288. Now the run fails.
    3. Two test fixtures advertised bytes past the end of the file. They hid problem 2.
  • Human review, with a short AWS Lambda test of the version before the second review (results below).

Questions you might have

Will this break anything that works today?
No. A link that fails any check is downloaded as before. A render on one machine never runs the check.

Why a relay instead of handing the link to FFmpeg?
FFmpeg follows redirects with no way to check where they go. And the first version only added If-Match on some reads, which is how findings 2 and 3 slipped in. The relay is one place that does both, for every read: frames, the duration probe and audio.

Is the relay reachable from outside?
No. It listens on 127.0.0.1 only, serves only the links this render chose to read in place (each under a random path), and only answers GET.

What if the video changes during the render?
Every read carries the ETag, so a changed file fails the render with a clear error instead of producing a mix. Weak ETags (W/...) don't qualify, because If-Match never matches them; those links are downloaded.

What about signed links that expire?
This is the one new requirement. A signed link now has to stay valid until the last chunk finishes, not just until planning ends. If it expires mid-render, the chunk fails with an error; it doesn't render wrong frames.

Does it put more load on the media host?
Far fewer bytes, a few more requests. Each chunk now makes a handful of range requests instead of the planner downloading the whole file.

Why read in growing pieces?
FFmpeg asks for "everything from here on" and hangs up once it has enough. Node's fetch slows down when nobody reads, but Bun's reads the whole rest of the file anyway (the Cloud Run image runs on Bun). Reading 3 seconds from the middle of a 73 MB file:

FFmpeg reading the link directly Through the relay
Node 7.4 MB 9.8 MB
Bun, one request for the rest of the file 7.6 MB 127 MB
Bun, growing pieces (this PR) 7.6 MB 23 MB

What about the frame cache?
Frames read from a link skip the local frame cache, because that cache is keyed by local files. Downloaded sources still use it.

How big is it?
About 590 lines of code (the relay file is 300 of them) and 550 lines of tests, across 27 files, plus one docs sentence in docs/deploy/aws-lambda.mdx.

Test plan

  • Unit tests added/updated
  • Manual testing performed
  • Documentation updated (if applicable)
  • Comments follow CONTRIBUTING.md "Comments": they say why, not what, and a bug fix says what the code must do and how to reproduce the bug

Same render on main and on this branch. A local HTTPS server with range and ETag support served a 73 MB video. The render was a 12-second clip in 4 chunks, using the plan v2 path, run under Bun (the worst case for the relay, see above).

main This PR
Planner reads from the link 146.7 MB (the whole file, twice) 53.4 MB of partial reads
Video copied into the plan Yes, to all 4 chunks No
Plan size 70 MB 0.4 MB
Each chunk reads from the link 0 MB (it has the copy) 14 to 30 MB
Chunk browsers fetch the video No No
Final video Byte-identical to main

On AWS Lambda. A 1,000-frame 4K render (40 seconds, two 50-minute HEVC recordings read from signed S3 links) of the version before the second review. The second review's fixes were checked locally: same final video as main. It also carried the 25 fps support from #5423, which that footage needs.

First version of this PR Before the second review
Cost $0.1186 $0.1190
Plan size 1 MB, no video copied
Picture and sound Identical (same decoded MD5)

The full 20-minute edit in the table at the top was rendered with the first version of this change.

New unit tests.

  • The link check reads a link in place when it should. It downloads instead when the link ignores ranges, has no ETag, has only a weak ETag, redirects, or returns a text page or a playlist. Matroska and WebM are read in place. A local path or a blocked address never gets a request, and each link is asked about only once.
  • The relay forwards the range and If-Match, passes a 412 through, refuses redirects to a private address or to plain HTTP (without fetching them), reads in growing pieces, cuts the connection when the link sends less than it promised, and serves only links it was given. A link that fails partway is reported to the run that read it; FFmpeg hanging up normally is not.
  • Frames read through the relay match frames read from the local file, and a changed file fails.
  • A link that answers 503 or 412 partway through a read fails the extraction instead of losing frames.
  • A deferred video whose file changed after the check fails instead of being read (review finding 2).
  • The plan accepts a link with its ETag, and rejects a link without an ETag or a link that isn't public HTTPS.
  • A chunk reads a deferred video through its relay at the planner's version, matching the planner's frames, and its page never fetches that link.
  • Compile keeps only qualifying <video> links and downloads every other source. The probe doesn't load the videos read in place.

Existing suites. Typecheck, lint, format and the pre-commit checks pass. On my machine, 7 frame-sampling and colour tests and 1 audio-level test fail, and they fail the same way on main (most likely my local FFmpeg build).

@robinebers
robinebers force-pushed the feat/remote-video-range-reads branch from d5780fb to 415cfd9 Compare October 11, 2026 05:56
@robinebers

Copy link
Copy Markdown
Author

What changed since the first version

Code-reviewed by GPT-6 Astra and human-reviewed with a short AWS Lambda test. The review found four problems; all are fixed:

  1. Redirects. FFmpeg could follow a later redirect to plain HTTP or a private address. FFmpeg now reads from a small relay on 127.0.0.1, which checks every redirect hop with the downloader's rules.
  2. Changed file, fallback path. One fallback quietly read a replaced file. Every read now goes through the relay with If-Match, so it fails instead.
  3. Audio and duration probe. These read the link without If-Match. Same fix.
  4. Blocking in the browser. It ignored the query string, so it could block a different signed link to the same file. It now blocks the exact link (checked in headless Chrome).

One more thing came up while testing: Bun's fetch keeps downloading after FFmpeg hangs up, and Cloud Run runs on Bun. So the relay asks for the bytes in growing pieces (1 MB up to 8 MB), and a hang-up stops the transfer within one piece.

Lambda test of this exact commit: 1,000 frames of 4K from two 50-minute HEVC recordings on signed S3 links. Cost $0.119, a 1 MB plan with no video copied, and the same decoded picture and sound as the first version.

The description is updated with the details.

…d of shipping them

A distributed plan that defers extraction now leaves a remote <video> source at its URL.
Each chunk reads only the byte ranges of the frames it renders, and its page never fetches
the video. The planner no longer downloads the source, and no copy of it is shipped to any
chunk.

A URL is read in place only when it is public HTTPS, answers a range request without
redirecting, carries a strong entity tag and holds an MP4, MOV, Matroska or WebM file, so
never a playlist whose segments FFmpeg would fetch itself. The plan asks once
per URL, so compile, probe, extraction and audio agree on the answer. Anything else is
downloaded exactly as before. Local renders are unchanged.

FFmpeg never reads the URL itself. Each process serves the sources it reads in place from
a relay on 127.0.0.1, which fetches them with the downloader's checks on every redirect
hop and with If-Match on every request, so a redirect cannot reach a private or plain-HTTP
address, and a source replaced mid-render fails instead of mixing two versions. A relayed read that
fails partway fails the FFmpeg or FFprobe run that made it, even when FFmpeg exits cleanly
with what it had. The relay
asks for growing windows of the range FFmpeg wants, so a read FFmpeg stops early stops the
transfer within one window.

Co-authored-by: Cursor <cursoragent@cursor.com>
@robinebers
robinebers force-pushed the feat/remote-video-range-reads branch from 415cfd9 to 320023b Compare October 11, 2026 06:28
@robinebers

Copy link
Copy Markdown
Author

Minor follow-up from a second review, pushed as 320023b:

  • Playlists are no longer read in place. An HLS playlist passed the link check, and FFmpeg would then fetch its segments itself, around the relay's checks. Now only MP4, MOV, Matroska and WebM files are read in place. Everything else is downloaded as before.
  • A link that fails partway now fails the render. With a 503 (or a 412) after the first piece, extraction used to "succeed" with 128 of 288 frames. The relay now records the failure, and the FFmpeg or FFprobe run that was reading fails.
  • Two test fixtures advertised bytes past the end of the file, which hid the second problem. Fixed.

Checked locally: the same final video as main, and the existing suites show the same results as main. Lambda wasn't rerun for this follow-up.

@robinebers

Copy link
Copy Markdown
Author

this has ballooned a little too much. closing this, reworking, and opening a new one.

@robinebers robinebers closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant