2026-09-30 · 6 min read

Subtitle burn-in API: hardcode styled subtitles into video with one POST

Burning subtitles into a video means the text becomes part of the pixels. No player toggle, no sidecar file, no font support roulette on the viewer's device. For anything headed to a social feed, that is the point: the captions render identically everywhere because they are the video. A subtitle burn-in API is the service version of that idea - you POST a video, you get back a video with the captions already inside it.

Hardcoded vs sidecar: pick based on the destination

  • Sidecar (SRT/VTT): a separate text file the player renders. Good for accessibility archives and platforms with native caption support like YouTube. Breaks the moment the video is reposted without its file.
  • Hardcoded (burned-in): the captions are baked into the frames. They survive downloads, reposts and re-uploads, and the styling you chose is the styling everyone sees. This is what TikTok, Reels and Shorts workflows actually ship.

The social case won the argument. When a clip moves between platforms and devices, the only caption format guaranteed to survive is the one inside the pixels. The trade-off is that burned-in captions are expensive to produce well: someone has to transcribe the audio, time every word, style the text, and render the result. That used to mean an editor or an ffmpeg pipeline. An API collapses it into a request.

What the one-call version looks like

curl -X POST https://api.caption.sh/v1/captions \
  -H "Authorization: Bearer $CAPTION_API_KEY" \
  -F "video=@clip.mp4" \
  -F 'preset=karaoke'

caption.sh transcribes the audio, produces word-level timing, applies the preset's styling, and burns the result into the video. The response is a job you poll or a webhook you receive, and the finished file is a render-ready MP4 with the captions inside. Presets cover the looks short-form feeds run on - karaoke-style active-word highlight, clean single-line, boxed - and custom styles are fields on the same request.

Why word-level timing is the whole product

The difference between captions that look professional and captions that look generated is timing granularity. Sentence-level timing gives you blocks of text that linger. Word-level timing gives you the karaoke effect, correct line breaks, and text that changes exactly when the speaker does. Any burn-in API worth integrating exposes word timing from its transcription step rather than asking you to supply it.

When to build the pipeline yourself

If you render a handful of videos a month and already have transcripts, ffmpeg with an ASS file is free and fine. The API starts paying for itself the moment you need transcription, volume, retries, or a look that has to stay consistent across hundreds of renders. At $0.10 per minute of video, the build-versus-buy math turns quickly - the pipeline you are not maintaining is the expensive part.

caption.sh is a subtitle burn-in API built for exactly this: one POST in, a captioned MP4 out. There is a free trial on the homepage, and the same engine is available to AI agents over MCP.

Back to all articles