Automatic TikTok-Style Captions

The style of viral videos: a few words at a time, in the center of the picture, with the spoken word highlighted in the color of your choice. MP4 export in 9:16.

100% local processing, nothing is uploaded

How does it work?

Drop your video: it is transcribed in your browser, each word with its timestamp. The “Social media” style shows 1 to 3 words at a time, in large type in the center of the picture, and highlights the spoken word in the color of your choice, in time with the voice.

The 9:16 format (1080 × 1920) is applied, with the picture cropped to fill the screen; you can switch back to the original format. Choose the font (Anton by default), size, colors, outline and text height, with an instant preview.

Correct misheard words before exporting: the highlighting automatically follows the corrected text. Export to MP4, ready for TikTok, Instagram Reels and YouTube Shorts.

Examples

A 45-second Reel

45 seconds of speech, about 110 words, shown two at a time: about 55 groups on screen, each with its highlighted word. The 1080p export weighs about 46 MB.

Word by word for a fast pace

With “1 word at a time” and a size of 120 px, each word is shown alone in the center of the screen: the most dynamic effect, suited to videos under 30 seconds.

Frequently asked questions

How is the spoken word synchronized?
The transcription timestamps every word (not every sentence) thanks to Whisper’s alignment. If you correct a word, it keeps the timing of the word it replaced; added words are spread between their neighbors. Accuracy is about a tenth of a second.
Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
Is there a length limit?
Transcription accepts files of 30 minutes at most: beyond that, split the video into parts. The video with burned-in subtitles can last up to 10 minutes, because the MP4 file is created in the browser’s memory. SRT, VTT and text files have no limit.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.