The sound out of a video, without the video

Pull the soundtrack out as a WAV — for a podcast edit, a transcript, or just to keep the audio. Your browser decodes the file itself, so nothing is uploaded and nothing is queued behind anyone else's job.

Home › Audio & Video › Extract Audio
1

The file to take it from

Anything your browser can play: MP4, WebM, MOV, M4A, MP3.

Files stay on this device

Drop a video or audio file, or click to choose

Decoded here in your browser — nothing is uploaded

Why WAV, and what that means for size

Your browser can decode almost any audio it can play, but it has no built-in encoder for compressed formats — so writing an MP3 would mean shipping a large encoder library to your device. Rather than do that, this tool writes WAV, which the browser can produce natively and which every editor, transcription tool and DAW accepts.

The trade-off is size: WAV is uncompressed, roughly 10 MB per minute in stereo at CD quality. Two options cut that sharply when you don't need fidelity — mono halves it, and dropping the sample rate to 16 kHz cuts it further while remaining perfectly good for speech. That combination is what transcription software actually wants, and it produces files about a tenth the size. If you need MP3 for distribution, convert the WAV afterwards in an audio editor.

What it can and can't open

Extraction works with whatever your browser can decode — MP4/AAC, WebM/Opus, MOV, MP3, M4A and WAV all work in current browsers. Unusual containers or codecs may fail to decode, and when that happens the tool tells you plainly rather than silently sending your video to a server. Very long files are decoded into memory, so a feature-length video may be limited by your device's RAM; trimming to a section keeps that manageable.

Sample rate, channels, and what to keep

Two numbers determine most of the size of the result. The sample rate is how many times a second the waveform was measured — 44.1 kHz is the CD rate and 48 kHz is the video and broadcast rate, which is why audio pulled from video is usually the latter. The channel count is one for mono and two for stereo, and it doubles the file.

If the destination is transcription, both can come down considerably. Speech recognition models typically resample to 16 kHz internally and treat the audio as mono, so a stereo 48 kHz export is carrying roughly six times the data the model will actually use. For a podcast edit or anything a person will listen to, keep the original rate and the stereo image; the file is larger and nothing has been thrown away that you might want back.

Every re-encode costs something

The audio inside a video file has almost always been compressed already, usually with AAC. Extracting it to WAV decodes that compression and stores the result uncompressed — which is why the WAV is so much larger than the video's audio track, and why it is not any better: the detail discarded during the original encode is gone, and a large file cannot bring it back.

What it does do is stop the loss accumulating. Compressing an already-compressed track again, to MP3 or to another AAC, applies a second round of the same lossy process to material that has already been approximated, and the artefacts compound. So WAV is the right intermediate for editing or processing, and the right time to compress is once, at the end, from the working file — not at every step along the way.