Audio and video, worked on where it sits
Transcribe a recording, write and fix the subtitles, convert between the two formats that look alike and aren't, or pull the sound out of a video. Every one of them runs in this browser: your media is never uploaded, and there is no account or key in the loop.
Video & Audio Transcriber
Drop in a recording — or capture a tab that is playing one — and get a transcript with timings, an editor beside the player, and SRT, VTT, TXT or JSON out the other end. The speech model runs on your own machine, downloaded once and cached after that.
Open the transcriber- Your fileOr a tab's audio
- On this deviceNothing uploaded
- SubtitlesEdited against the player
- SRT · VTT · TXTYours to keep
The four tools
Each does one part of the job, and hands over to the next when it is done.
Start from what you have
- A video, and no subtitles yetThe transcriber makes them, then the editor tidies them.Transcribe it
- A subtitle file the player rejectsNine times in ten it is SRT wearing a .vtt name. The converter writes a real one.Convert it
- Subtitles that are out of syncShift every cue by a fixed amount, or correct a frame-rate mismatch.Shift the timing
- Subtitles that read badlyWrap long lines and enforce a minimum duration and a gap between cues.Open the editor
- A video you only want the sound fromExtract it as a WAV, trimmed to the part you need.Extract the audio
Every one of these reads your file in this browser. Nothing is uploaded, so nothing has to be deleted afterwards.
Private by design
Every tool in this category is built on the same principle as the rest of Toolsfully: your files are processed locally in your browser, not uploaded to a server. The transcriber uses an on-device AI speech model (downloaded once and cached), and the subtitle engine that powers it is shared across these tools so behaviour stays consistent. We describe exactly how each tool works rather than making absolute security claims.
What actually runs, and where
Transcription here uses OpenAI's Whisper model running inside your browser through transformers.js, on WebGPU where your device offers it and WebAssembly otherwise. Your recording is decoded, resampled to the 16 kHz the model expects, and processed on your own machine. It is never uploaded, and there is no account or API key in the loop.
The cost of that is the model itself, which has to be fetched once before anything can run. There are three sizes: roughly 40 MB for the fastest, 80 MB for the balanced one, and 240 MB for the most accurate. Your browser caches it afterwards, so the wait happens on the first transcription and not again. A larger model is slower and better, and on a long recording the difference in both directions is substantial — it is worth trying the middle option first rather than assuming you need the largest.
SRT and WebVTT are nearly the same file
The two subtitle formats you will meet are close enough to look interchangeable and different enough to break a player. SRT writes its timestamps with a comma before the milliseconds — 00:01:23,480 — while WebVTT uses a full stop, 00:01:23.480, and requires the literal word WEBVTT on the first line. Rename an .srt to .vtt and a browser will usually reject the file outright, because neither of those two things is true of it.
That is nearly the whole of the difference for ordinary subtitles, which is why converting between them is quick and lossless in both directions. WebVTT can additionally carry positioning and styling that SRT has no way to express, so a VTT using those features loses them on the way to SRT. Plain dialogue does not.
What makes a subtitle readable
Subtitling conventions exist because a viewer is reading and watching at the same time. Two lines is the practical ceiling, and around 42 characters a line is the widely used limit — wide enough for normal phrasing, narrow enough to be taken in at a glance without crowding the picture. Longer lines do not fail; they just stop being read in time.
Timing matters as much as wrapping. A cue that appears for a fraction of a second cannot be read at all, one that lingers invites the viewer to re-read it, and two cues that overlap will either fight for the same space or flicker, depending on the player. Enforcing a minimum duration, a maximum duration and a small gap between cues fixes all three, and it is worth doing after any edit that moved timings around rather than trusting that they still line up.
Where automatic transcription goes wrong
It is worth knowing the failure modes rather than discovering them in a finished file. Overlapping speakers are the hardest case — the model transcribes a single stream of speech and has no concept of who is talking, so a crosstalk section tends to come out as one speaker's words with the other's missing. Background music and room noise degrade accuracy steadily rather than obviously. Proper nouns, technical vocabulary and names are guessed phonetically and are the first thing to check.
The distinctive one is repetition. Whisper occasionally falls into a loop and emits the same phrase over and over, usually across a passage of silence or noise where there is nothing to transcribe. It is a known behaviour of the model rather than a fault in the audio, it is obvious once you know to look for it, and the editor here detects and strips those runs.
None of this makes automatic transcription unreliable so much as unfinished. Treat the output as a first draft that saves you the typing, not as a transcript that is ready to publish.