A faceless YouTube pipeline on a free local stack

3 min read YouTubeffmpegedge-ttsAutomation

Turn a script JSON into a captioned long-form video plus three derivative Shorts with no subscriptions, using edge-tts, headless Edge, and ffmpeg.

The symptom

I needed a repeatable way to turn a written script into a captioned YouTube video and its three derivative Shorts, without recording a face or paying a monthly SaaS fee. Every tool I found either cost money or handed control to a third-party renderer I could not inspect, and I wanted something I could run locally, understand completely, and extend when something went wrong.

The first attempt failed silently in three places: headless Chrome wrote no PNG, ffmpeg shook on zoomed clips, and the captions were synced to sentences instead of words, so the karaoke effect was useless.

What was happening

The headless browser was attaching to the running instance instead of rendering. Edge and Chrome need --user-data-dir pointed at an isolated, non-shared directory when run headlessly from a subprocess. Pass the path to your normal profile, or omit the flag, and the process attaches to whatever is already running, fires no render, and exits cleanly with code 0. The PNG never appears and there is no error.

ffmpeg’s zoompan was shaking on every scene. The filter expects input frames at its output resolution; if the source frame is smaller or larger, the crop window shakes because ffmpeg rounds pixel coordinates against the wrong grid. Scaling the input to a supersample width (4x or 6x the output width, for short versus long scenes) before piping into zoompan fixes it. Without that scale step the Ken Burns motion stutters at every whole-pixel boundary.

edge-tts was emitting sentence timings instead of word timings. Version 7.x changed the default boundary type to SentenceBoundary, and word-synced karaoke captions need WordBoundary events, which you have to request explicitly:

communicate = edge_tts.Communicate(text, voice, rate=rate)
async for event in communicate.stream():
    if event["type"] == "WordBoundary":
        timings.append((event["offset"], event["duration"], event["text"]))
    elif event["type"] == "audio":
        audio_chunks.append(event["data"])

Without that filter each timing event covers a full sentence, and the caption highlight snaps once per sentence instead of once per word.

The fix

The production pipeline, now packaged as the Faceless YouTube Automation Kit, resolves all three:

scene JSON
  -> fill HTML template per scene
  -> headless Edge (isolated --user-data-dir, device-scale-factor 2)
       => 2160x3840 PNG (4x headroom for zoompan)
  -> edge-tts with WordBoundary -> per-word timing list
       => karaoke ASS subtitle file
  -> ffmpeg:
       scale to supersample -> zoompan Ken Burns (ZOOM_RATE = 0.009/s)
       -> burn captions -> whoosh/pop SFX in silent tail -> fade
       -> concat all scenes -> final encode
  -> long-form mp4 (prints chapter timestamps for YouTube description)
  -> 3 Short mp4s from the same scene JSON

The b-roll is fetched per buyer from Pexels or Pixabay with the buyer’s own API key, which keeps the licence clean (you hold the download right) and keeps the kit distributable. A remotion-broll/ module ships for animated data overlays (FocusCurve, StatCard, CompareBars) if you want motion graphics rather than footage; it is optional and Node-only, and the core pipeline is Python with no Node requirement.

The music bed is CC-BY, Kevin MacLeod tracks fetched at setup time. The attribution is written to ATTRIBUTION.txt automatically and the long-form description template includes the credit line.

The lesson

A headless render that exits 0 and writes nothing means the process attached to the wrong target. Verify the output file exists before treating the subprocess as done.

For zoompan, scale first and zoom second. The filter maths breaks if the input is not already at the resolution the filter expects.

For word-synced captions with edge-tts, request WordBoundary explicitly. The library changed its default and will not warn you when it switches to sentence boundaries.


The kit is available now with launch code LAUNCH for EUR 19 (regular price EUR 29). The code is use-capped and will be removed once the intro period closes.

Related fixes

Discussion

Powered by GitHub. Sign in to leave a comment.