Skip to content
toolsdocks

Convert text to speech

Turn text into natural speech and download it as WAV or MP3.

On-device model
Loading tool…

How to use

  1. Paste text.
  2. Pick a voice and speed, then download the model (first time only) and generate. Browser voices play instantly but cannot be downloaded.
  3. Listen, then download WAV or MP3.

Worked example

The 40-word sample text with the Heart voice at 1.0× becomes 16.5 seconds of speech, a 771 KB WAV.

Supported formats and limits

InputText up to 5,000 characters
OutputWAV (24 kHz mono), MP3, Browser voices: playback only
LimitsFirst use downloads about 93 MB of model files from Hugging Face (cached afterwards). English voices (US and UK).
EngineKokoro-82M (Apache-2.0, 8-bit ONNX) via kokoro-js in a Web Worker, sentence by sentence; or the browser's speechSynthesis voices

Limitations

  • Unusual names and abbreviations may be mispronounced; spell them phonetically.
  • Only English is supported by the bundled voices.

Questions

Can I download the audio?

Yes with Kokoro: the speech is generated as a 24 kHz mono WAV, or MP3. Browser voices play instantly but are playback only, because browsers do not hand that audio to the page.

What is downloaded the first time?

About 93 MB of Kokoro model files from Hugging Face, plus a 0.5 MB style file per voice, cached for later visits. Your text is not sent.

Privacy

On-device model. Processing runs in this browser. The open-source model files are downloaded once from the model host (Hugging Face) and cached; your content is not uploaded.

  • Hugging Face: Downloads of the open-source model and voice files when you ask for them. No text or audio is sent.

See the privacy policy for how toolsdocks handles data.