How to use
- Add a recording.
- Choose the model (Tiny is fastest, Small most accurate), the language or automatic detection, and whether to translate to English.
- Download the model (first time only), transcribe, review, and copy or download TXT, SRT, VTT or JSON.
Worked example
The 6-second spoken sample ('Welcome to the audio tools…') is transcribed by the Tiny model as text with start and end times, ready to save as SRT or VTT subtitles.
Supported formats and limits
| Input | MP3, WAV, M4A, Ogg, FLAC, WebM, MP4, MOV |
|---|---|
| Output | TXT, SRT, VTT, JSON with segments |
| Limits | First use downloads 44 MB (Tiny, English), 80 MB (Base) or 252 MB (Small) from Hugging Face, more on WebGPU; cached afterwards. Up to about 2 hours, processed in sections of up to 30 seconds. |
| Engine | Whisper (MIT) via transformers.js in a Web Worker: WebGPU when available, otherwise WebAssembly; cuts at quiet points, silent sections skipped |
Limitations
- Speaker names are not identified.
- Background music and crosstalk reduce accuracy.
Questions
Which model should I pick?
Tiny (about 44 MB, English only) is the fastest. Base (about 80 MB) and Small (about 252 MB) handle other languages and are more accurate but slower. WebGPU is much faster when your browser supports it.
How accurate is the transcript?
It depends on the recording. Music, overlapping speakers and heavy accents lower accuracy, and names or numbers can be wrong, so check the text before relying on it.
Guides
Privacy
On-device model. Processing runs in this browser. The open-source model files are downloaded once from the model host (Hugging Face) and cached; your content is not uploaded.
- Hugging Face: Downloads of the open-source Whisper model files when you ask for them. No audio or text is sent.
See the privacy policy for how toolsdocks handles data.