What is Supertonic-3?
Supertonic 3 is a lightweight text-to-speech system from Supertone that runs entirely on your device through ONNX Runtime, with no cloud call for synthesis. The third release expands the open weights from 5 to 31 languages, improves reading stability with fewer repeat and skip failures, raises speaker similarity over Supertonic 2, and adds simple expression tags such as laugh, breath and sigh.
| Specification | Supertonic-3 |
|---|---|
| Parameters | About 99M across the public ONNX assets |
| Modality | Text to speech |
| Languages | 31, including English, Korean, Japanese, German, Spanish, French, Hindi, Russian |
| Runtime | ONNX Runtime, on-device |
| Voices | Fixed presets included; custom styles via Supertone Voice Builder |
| Release date | May 2026 |
| License | OpenRAIL-M, sample code MIT |
What Supertonic 3 is good at
The design target is practical on-device inference. Supertone reports the model runs fast on CPU, even against larger baselines measured on an A100 GPU, and uses substantially less memory; no GPU is required, which makes local, browser and edge deployment straightforward. On reading accuracy, the vendor's measurements keep it within a competitive WER and CER range against much larger open TTS systems such as VoxCPM2.
At about 99M parameters it is far smaller than the 0.7B to 2B class of open TTS models, which shows up as smaller downloads and faster startup. The whole ONNX asset set is roughly 400 MB: a text encoder, duration predictor, vector estimator and vocoder.
Supertonic 3 hardware requirements
Any machine that runs Python. The public assets total about 400 MB, the runtime is CPU-first, and Supertone highlights browser and edge deployment as intended targets. There is no VRAM requirement to plan around.
How to run Supertonic 3 locally
Install the Python SDK and the model downloads itself from Hugging Face on first run:
- pip install supertonic
- Pick a preset voice style and call synthesize with your text and language code.
- Save the wav. Expression tags such as laugh and sigh go inline in the text.
Custom voice styles can be built from reference audio with Supertone's Voice Builder; purchased styles include embeddings for both Supertonic 2 and 3. For chat models to pair the voice with, see the full catalog of local models.
Supertonic-3 license
The model is released under the OpenRAIL-M license and the sample code under MIT. OpenRAIL-M is free to use with usage-based restrictions listed in the license file, so read it before commercial deployment.
