Drop an audio file, get your transcript. Free and unlimited, right in your browser: download it as VTT, SRT, or plain text, ready to publish.
🔒100% private, 100% in your browser. The AI model runs on your device, so your audio is never uploaded anywhere. No cloud, no account, no limits.
Yes, completely. No account, no watermark, no length limits, no trial that expires. It is a free tool from RSS.com for the podcasting community. And if you would rather never drop a file at all, all RSS.com paid plans generate transcripts automatically for every episode you publish; if your show lives elsewhere, you can move it to RSS.com and get 3 full months free on any paid plan.
This page downloads an open-source speech recognition model (OpenAI's Whisper) straight into your browser, and the transcription runs on your own device. Your audio never leaves your computer: nothing is uploaded, there is no server, and you can even disconnect from the internet once the model has loaded.
Each level is a different size of the same AI model. Bigger models understand speech better, especially in languages other than English, but they are heavier: the download is larger, they take more space on your device, and they need more computing power. Fast (~120 MB) and Balanced (~210 MB) run well almost anywhere. Best (~590 MB) is noticeably more accurate and fine on most laptops. Max (~1.6 GB) produces professional-grade transcripts but needs a computer with a modern graphics chip (GPU); if yours does not have one, pick a smaller model.
WebVTT is the standard format for podcast transcripts and captions, and the recommended format for the Podcasting 2.0 transcript tag. Add it to your episode and podcast apps can show your transcript alongside the audio. You also get SRT (the classic subtitle format) and plain text.
Yes, in any modern browser on Windows, Mac, and Linux, and on recent smartphones with one caveat: on phones, Fast may be the only accuracy level that works, because bigger models often need more memory than a phone allows (if a run fails or the page suddenly reloads, that is why). Shorter episodes work best on mobile. For full episodes and higher accuracy, a laptop or desktop is the right tool.
No. The first time you use an accuracy level, your browser downloads that model once and stores (caches) a copy in its own storage, on your device. On your next visit the model loads from that copy instantly, with no new download. The cache belongs to your browser, and your browser manages it: it is emptied if you clear your browsing data, it is not kept in private or incognito windows, and it may be quietly removed if your device runs low on space. Nothing breaks either way: the page simply downloads the model again next time.
It depends on your device and the accuracy level. On a modern laptop, expect a few minutes for a one-hour episode on the Balanced model. Keep the tab open while it works.