Sign up free

Free Podcast Transcript Generator

Drop an audio file, get your transcript. Free and unlimited, right in your browser: download it as VTT, SRT, or plain text, ready to publish.

🔒100% private, 100% in your browser. The AI model runs on your device, so your audio is never uploaded anywhere. No cloud, no account, no limits.

🎙️
Drop your audio here
or click to choose a file  •  MP3, M4A, WAV, FLAC, even MP4/MOV video
The first run downloads the AI model to your browser, cached for next time: about 120 MB (Fast), 210 MB (Balanced), 590 MB (Best), or 1.6 GB (Max). For non-English audio, Best or Max gives noticeably better results. If the transcript comes out in the wrong language, pick yours explicitly.
You can switch to another tab while it works: the progress shows in the tab title, and your transcript will be waiting right here.
cancel
Transcript

Frequently asked questions

Is it really free?

Yes, completely. No account, no watermark, no length limits, no trial that expires. It is a free tool from RSS.com for the podcasting community. And if you would rather never drop a file at all, all RSS.com paid plans generate transcripts automatically for every episode you publish; if your show lives elsewhere, you can move it to RSS.com and get 3 full months free on any paid plan.

How does this work without uploading my audio?

This page downloads an open-source speech recognition model (OpenAI's Whisper) straight into your browser, and the transcription runs on your own device. Your audio never leaves your computer: nothing is uploaded, there is no server, and you can even disconnect from the internet once the model has loaded.

Which accuracy level should I choose?

Each level is a different size of the same AI model. Bigger models understand speech better, especially in languages other than English, but they are heavier: the download is larger, they take more space on your device, and they need more computing power. Fast (~120 MB) and Balanced (~210 MB) run well almost anywhere. Best (~590 MB) is noticeably more accurate and fine on most laptops. Max (~1.6 GB) produces professional-grade transcripts but needs a computer with a modern graphics chip (GPU); if yours does not have one, pick a smaller model.

What is a VTT file?

WebVTT is the standard format for podcast transcripts and captions, and the recommended format for the Podcasting 2.0 transcript tag. Add it to your episode and podcast apps can show your transcript alongside the audio. You also get SRT (the classic subtitle format) and plain text.

Does it work on my phone?

Yes, in any modern browser on Windows, Mac, and Linux, and on recent smartphones with one caveat: on phones, Fast may be the only accuracy level that works, because bigger models often need more memory than a phone allows (if a run fails or the page suddenly reloads, that is why). Shorter episodes work best on mobile. For full episodes and higher accuracy, a laptop or desktop is the right tool.

Can I drop a video file?

Yes. Drop an MP4, MOV, or WebM and the audio track is extracted and transcribed; the video itself is never read. Handy for video podcasts and YouTube exports. Keep in mind the whole file is loaded into memory first, so for very large video files it is lighter to extract the audio before dropping it.

Do I have to download the AI model every time?

No. The first time you use an accuracy level, your browser downloads that model once and stores (caches) a copy in its own storage, on your device. On your next visit the model loads from that copy instantly, with no new download. The cache belongs to your browser, and your browser manages it: it is emptied if you clear your browsing data, it is not kept in private or incognito windows, and it may be quietly removed if your device runs low on space. Nothing breaks either way: the page simply downloads the model again next time.

How long does it take?

It depends on your device and the accuracy level. As a rule of thumb, a one-hour episode takes a few minutes on a modern laptop with the Balanced model, and five minutes or more on slower machines (like Chromebooks) or phones, where transcription runs on the CPU. Long files also look quiet for the first minute or so while the audio is prepared, before the progress bar starts moving. Keep the tab open while it works; the progress bar and time estimate will tell you where you are.