2026-09-29 · 7 min read
ffmpeg burn captions: when the CLI works and when to use an API
Every developer who needs captions burned into a video starts in the same place: ffmpeg. It is already installed, it is free, and the subtitles filter does exactly what it says. For a one-off render, that is the right answer. This post is the honest map of where the ffmpeg path ends and where an API starts paying for itself, with working examples of both.
The ffmpeg way, actually working
Burning an existing subtitle file into a video is one filter:
ffmpeg -i input.mp4 -vf "subtitles=captions.srt" -c:a copy output.mp4
# styled, via an ASS file with fonts, colors and positioning:
ffmpeg -i input.mp4 -vf "ass=captions.ass" -c:a copy output.mp4The catch is that ffmpeg only renders the captions you hand it. It does not transcribe audio, it does not produce word-level timing, and it does not invent the karaoke-style active-word highlight that short-form feeds run on. For that look, every word needs its own timing and its own style event in the ASS file. That is a generator you now have to write and maintain.
Where the CLI path breaks at volume
- Transcription: you need speech-to-text with word timestamps before ffmpeg has anything to render. That is a second vendor, a second job queue, and a joining step.
- Styling: ASS styling is powerful but arcane. Fonts must exist on every worker that renders, and a small change to the look means regenerating every file.
- Operations: render jobs need retries, timeouts, progress tracking and storage. One video is a command; ten thousand is a distributed system.
- Fonts on servers: the karaoke captions you see on social use custom typefaces. Shipping font files to every render worker and keeping fontconfig happy is its own quiet tax.
None of this means ffmpeg is the wrong tool. It means the real system you are building is transcription plus timing plus styling plus rendering plus job management, and ffmpeg only covers one of those five.
The API version of the same job
caption.sh collapses the whole chain into one async job. POST a video URL, poll the job, download a captioned MP4:
curl -X POST https://api.caption.sh/v1/videos \
-H "Authorization: Bearer $CAPTION_API_KEY" \
-H "Content-Type: application/json" \
-d '{"source_url": "https://example.com/raw.mp4"}'
# -> 202 { "job_id": "..." }
# poll GET /v1/jobs/{id} until status is "done"The style layer is JSON instead of ASS: words per group, typeface, colors, the active-word animation, position. Change the look by changing data, not by regenerating subtitle files, and fonts are handled server-side. Transcription with word-level timing is included, so there is nothing to join.
When to stay on ffmpeg
If you caption a handful of videos a month with a fixed subtitle file you already have, stay on ffmpeg. It is free, deterministic, and scriptable, and an API would add nothing. The crossover point is when captions become a pipeline: transcription needed, styles changing per template or per customer, volume high enough that you are operating render jobs instead of running commands. That is the moment the API stops being a cost and starts being the smaller system.
The playground on the caption.sh homepage runs a real render in the browser, no account needed. If you want the spec first, the OpenAPI document is public at api.caption.sh/openapi.json.