Supertonic-3

Updated
24.08.2026
Audio
Multilingual

Supertonic 3 is a 99M-parameter on-device TTS from Supertone that speaks 31 languages and runs fast on CPU, with about 400 MB of ONNX assets.

At a glance

  • License: OpenRAIL-M for the model, MIT for the sample code
  • Parameters: about 99M across the public ONNX assets
  • Languages: 31, expanded from 5 in Supertonic 2
  • Minimum hardware: CPU only, roughly 400 MB of ONNX assets
  • Extras: expression tags for laugh, breath and sigh; preset and custom voice styles

What is Supertonic-3?

Supertonic 3 is a lightweight text-to-speech system from Supertone that runs entirely on your device through ONNX Runtime, with no cloud call for synthesis. The third release expands the open weights from 5 to 31 languages, improves reading stability with fewer repeat and skip failures, raises speaker similarity over Supertonic 2, and adds simple expression tags such as laugh, breath and sigh.

SpecificationSupertonic-3
ParametersAbout 99M across the public ONNX assets
ModalityText to speech
Languages31, including English, Korean, Japanese, German, Spanish, French, Hindi, Russian
RuntimeONNX Runtime, on-device
VoicesFixed presets included; custom styles via Supertone Voice Builder
Release dateMay 2026
LicenseOpenRAIL-M, sample code MIT

What Supertonic 3 is good at

The design target is practical on-device inference. Supertone reports the model runs fast on CPU, even against larger baselines measured on an A100 GPU, and uses substantially less memory; no GPU is required, which makes local, browser and edge deployment straightforward. On reading accuracy, the vendor's measurements keep it within a competitive WER and CER range against much larger open TTS systems such as VoxCPM2.

At about 99M parameters it is far smaller than the 0.7B to 2B class of open TTS models, which shows up as smaller downloads and faster startup. The whole ONNX asset set is roughly 400 MB: a text encoder, duration predictor, vector estimator and vocoder.

Supertonic 3 hardware requirements

Any machine that runs Python. The public assets total about 400 MB, the runtime is CPU-first, and Supertone highlights browser and edge deployment as intended targets. There is no VRAM requirement to plan around.

How to run Supertonic 3 locally

Install the Python SDK and the model downloads itself from Hugging Face on first run:

  1. pip install supertonic
  2. Pick a preset voice style and call synthesize with your text and language code.
  3. Save the wav. Expression tags such as laugh and sigh go inline in the text.

Custom voice styles can be built from reference audio with Supertone's Voice Builder; purchased styles include embeddings for both Supertonic 2 and 3. For chat models to pair the voice with, see the full catalog of local models.

Supertonic-3 license

The model is released under the OpenRAIL-M license and the sample code under MIT. OpenRAIL-M is free to use with usage-based restrictions listed in the license file, so read it before commercial deployment.

Get the weights from Hugging Face

pip install supertonic
# the SDK downloads the model assets from Hugging Face on first run
from supertonic import TTS

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")

text = "A gentle breeze moved through the open window while everyone listened to the story."
wav, duration = tts.synthesize(text, voice_style=style, lang="en")

tts.save_audio(wav, "output.wav")
print(f"Generated {duration:.2f}s of audio")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Supertonic 3 is a lightweight open-weight text-to-speech system from Supertone that runs entirely on your device through ONNX Runtime, with no cloud call for synthesis. It expands the open release from 5 to 31 languages, improves reading stability, and supports expression tags such as laugh, breath and sigh.

No. Supertone reports it runs fast on CPU, even compared with larger baselines measured on an A100 GPU, while using substantially less memory. At about 99M parameters and roughly 400 MB of ONNX assets, it is built for local, browser and edge deployment.

The model weights are released under the OpenRAIL-M license and the sample code under MIT. OpenRAIL-M permits free use with usage-based restrictions listed in the license file, so read it before commercial deployment.

Install the Python SDK with pip install supertonic, and the SDK downloads the model assets from Hugging Face on first run. Pick a preset voice style, call synthesize with your text and language code, and save the wav. Custom voice styles can be built from reference audio with Supertone's Voice Builder.

Language coverage grew from 5 to 31 languages, repeat and skip failures dropped, and speaker similarity improved across the shared-language set. Supertone also reports the model stays within a competitive WER and CER range against much larger open TTS systems such as VoxCPM2.