Auto Subtitle Generator
Animated captions like the ones on TikTok and Shorts — generated automatically, then editable word by word. No account, no watermark, no upload.
The AI runs on your device — your video is never uploadedDrop your video here
or click to browse — a vertical clip for Reels, Shorts or TikTok works best
MP4 · MOV · WebM · MP3 · WAV · M4A
Built for short-form clips — under about ten minutes. Transcription runs on your own processor, so keep this tab open while it works.
Three steps
No account, no watermark, and the video never leaves your computer.
Drop the clip in
A vertical video for Reels, Shorts or TikTok is the usual case, but any video or audio file works.
Generate, then fix
The AI writes the captions with punctuation and proper capitals. Correct any misheard word directly — the timing follows what you type.
Style and export
Pick a look, adjust the typeface and colours, then download a subtitle file or the finished video with captions burned in.
Who this is for
Anyone posting short-form video
Most of it is watched with the sound off. Captions are not a nicety on Reels and TikTok — they are the difference between being watched and being scrolled past.
Podcasters cutting clips
Turn a good two minutes of a long episode into something that works on a feed, with the words on screen.
Students and journalists
An interview or a lecture transcribed without paying per minute — and without handing the recording to a company you have never heard of.
Anyone tired of per-minute pricing
Every captioning service charges by the minute because transcription costs them GPU time. Running it here costs nothing, so it is free and unlimited.
Why captions decide whether your video gets watched
The majority of short-form video is watched without sound. People scroll in bed next to someone sleeping, on a bus, in a waiting room, at a desk. If your first sentence only exists in the audio, most of the feed never hears it — they just keep scrolling.
That is why every serious creator burns captions in. Not the platform's own auto-captions, which are unstyled, often wrong, and sometimes not shown at all, but captions that are part of the picture.
What makes captions readable
An outline, always. White text over a bright sky disappears. A dark outline around every letter keeps the words legible over anything, which is why the outline control here defaults to on rather than off.
Few words at a time. Two or three words per caption is the short-form convention. It forces the eye to move with the speech instead of reading ahead, and it keeps the text large enough to read on a phone held at arm's length.
A highlight on the spoken word. Colouring or enlarging the word currently being said is the single detail that makes captions feel professional rather than pasted on. It also genuinely helps comprehension — the eye knows where it is.
Room from the edges. Platforms overlay their own interface at the bottom and along the right side. Keep captions clear of the bottom fifth or the username and buttons will sit on top of them.
Why the automatic version still needs you
Whisper, the model this tool runs, is good — it punctuates, it capitalises, and it handles accents better than anything freely available a few years ago. It is not perfect. It will mishear proper nouns, brand names, technical terms and anything said over noise.
Which is exactly why every caption here is an editable box rather than a finished result. Fix the two words it got wrong, and the timing of the surrounding words is recalculated so the highlight stays in step. A tool that gives you an uneditable transcript is a tool that wastes your time the moment it makes a mistake.
SRT, VTT, or burned in?
Burn them in for TikTok, Reels and Shorts. Styled captions that are part of the picture always look better than the platform's own, and they cannot be switched off or rendered in some default font.
Use an SRT file for YouTube, where uploading subtitles separately means viewers can turn them on or off, and — importantly — Google can read them. Subtitles are indexed, so an SRT makes the video findable by what is said in it.
Use VTT if you are embedding the video on your own website with an HTML video player.
Nothing stops you doing both, and plenty of people do: burned-in captions for the visual, plus an SRT for the search engine.
About the one thing that is downloaded
This site's promise is that your files never leave your device, and that holds here: your video and its audio are never transmitted anywhere.
There is one honest asterisk worth stating plainly. The speech recognition model itself — around 80 MB — is downloaded to your browser the first time you use the tool, and cached afterwards. That is data coming in, in exactly the same category as a font or a JavaScript library. Your recording does not go out. Once the model is cached you can disconnect from the internet entirely and the tool keeps working, which is the simplest proof of the difference.
Every commercial captioning service works the other way round: your recording goes to their servers, gets processed on their hardware, and you are billed by the minute. That is why they cost money and this does not.
Common questions
Is my video uploaded anywhere?
No. The speech model runs inside your browser on your own processor, and your video and audio are never transmitted. The one thing that travels is the model itself — downloaded once from Hugging Face and then cached — which is data coming in, not your recording going out. After that first download the tool works offline.
Can I fix words the AI got wrong?
Yes, and you should. Every caption is an editable text box. Correct a name, fix the spelling, rewrite a whole line — the word timing is recalculated as you type so the highlight stays in sync with the speech.
Can I change the font, colour and size?
Four presets to start from, then full control: typeface, size, text colour, highlight colour, outline thickness, background box, vertical position, capitals, and how many words appear at once.
What do I get at the end?
Either a subtitle file — SRT for YouTube and video editors, VTT for websites — or a finished video with the captions burned in and no watermark on it.
How long does it take?
The model downloads once, which takes a moment the first time and is instant afterwards. Transcription then runs at roughly real time on a modern laptop, so a one-minute clip takes about a minute. Editing and restyling after that are instant.
Which languages does it handle?
Whisper is multilingual and covers around a hundred languages, including Spanish, Portuguese, French, German, Italian and Japanese. Leave the setting on automatic and it detects the language, or choose one to remove any doubt.
Why does exporting the video take as long as the video?
Because a browser can only record video as it plays — a two minute clip needs two minutes. We tried the faster-looking route, encoding frames directly with WebCodecs, and measured it: on ordinary hardware that encoder runs at about 47 milliseconds a frame, which works out slower than simply recording in real time. So real time it is. You can leave the tab in the background while it runs.
Generating the subtitles is the part where the model size matters. Burning them in is bound by the browser, not by the model.
What format is the exported video?
WebM. TikTok, Instagram, YouTube and every modern player accept it on upload. Browsers claim to be able to record MP4 from a canvas and then produce an empty file, so WebM is what we ship. If you specifically need an MP4, run the result through our video compressor.
More tools
Every one of them runs on your device. No uploads, no accounts.