Quick takeaways: Kokoro-82M is an open-weight 82M-parameter text-to-speech model with an Apache-2.0 license, an official Hugging Face model card, a demo Space, and a Python inference library. The safest first test is the official Python workflow: install
kokoro,soundfile, andespeak-ng, create aKPipeline, choose a language code and voice, then save WAV output. This guide does not claim independent quality benchmarks.
What is Kokoro-82M?
Kokoro-82M is the hexgrad/Kokoro-82M text-to-speech model on Hugging Face. Its model card describes it as an open-weight TTS model with 82 million parameters and an Apache-2.0 license.
The project is tutorial-friendly because it has an official model card, a GitHub inference library, a Hugging Face demo Space, and voice documentation.

When should you use it?
Use Kokoro when you need a lightweight TTS experiment for demos, personal projects, prototypes, or a Python-based speech generation workflow.
Do not use it to impersonate real people, create deceptive voice content, or attach a generated voice to someone else’s identity without permission.
Prerequisites
The official examples use Python, kokoro, soundfile, and espeak-ng. On Windows, the GitHub README points users to the espeak-ng releases page and an MSI installer.
For Apple Silicon, the README documents this environment variable for MPS fallback:
PYTORCH_ENABLE_MPS_FALLBACK=1 python run-your-kokoro-script.pyQuick Python/Colab test
The official setup cell is:
!pip install -q kokoro>=0.9.4 soundfile
!apt-get -qq -y install espeak-ng > /dev/null 2>&1Then import the pipeline and WAV writer:
from kokoro import KPipeline
from IPython.display import display, Audio
import soundfile as sf
import torchCreate an American English pipeline:
pipeline = KPipeline(lang_code='a')Generate and save audio:
text = 'Kokoro is an open-weight text-to-speech model.'
generator = pipeline(text, voice='af_heart')
for i, (gs, ps, audio) in enumerate(generator):
print(i, gs, ps)
display(Audio(data=audio, rate=24000, autoplay=i==0))
sf.write(f'{i}.wav', audio, 24000)Choosing languages and voices
The advanced README lists language codes including a for American English, b for British English, e for Spanish, f for French, h for Hindi, i for Italian, j for Japanese, p for Brazilian Portuguese, and z for Mandarin Chinese.
Japanese requires misaki[ja]; Mandarin Chinese requires misaki[zh]. The VOICES.md file also warns that non-English support can be thin due to G2P and training-data limitations.
Verification
A successful first run should import the libraries, display playable audio in the notebook, and write WAV files such as 0.wav. If phonemes print but audio fails, check soundfile and the audio playback environment.
For long text, VOICES.md warns that voices may rush above roughly 400 tokens. Split long scripts into shorter chunks.
Common issues
Missing espeak-ng is common outside Colab. Install it for your operating system.
A language-code and voice mismatch can produce poor results. Match the lang_code to the voice family.
Non-English quality varies. Test your target language and voice before using it in a public workflow.
Safety and legal notes
The model card warns about fake Kokoro-branded websites that are not affiliated with the model author. Use the official Hugging Face and GitHub sources.
The model license does not make every generated use acceptable. Avoid impersonation, fraud, and unauthorized voice cloning.
FAQ
Is Kokoro-82M open source?
The model card lists Apache-2.0 licensing for the model. Always verify the license on the official page before commercial deployment.
Does Kokoro support Vietnamese?
In the checked VOICES.md file, Vietnamese is not listed among the official voice groups. Do not assume Vietnamese support.
Do I need a GPU?
The official quickstart can be tried in a notebook without a GPU requirement, but performance depends on your runtime.








