Shisa TTS
Shisa TTS turns text into lifelike speech. It produces natural-sounding Japanese and English audio with genuine intonation and prosody — voices vetted by native Japanese speakers rather than English voices stretched to fit Japanese. Send text, pick a voice, and get back binary audio in the format you ask for.
POST https://api.shisa.ai/tts
Why Shisa TTS
- Natural Japanese voices — highly realistic Japanese and English speech with natural intonation and prosody, not translated-sounding output.
- Voice catalogue — choose from active voices with published language, format, sample-rate, and streaming capabilities.
- Real-time streaming — low-latency streaming for interactive applications and real-time conversations.
- Multiple languages and formats — Japanese, English, and more with native pronunciation quality, delivered as
mp3,wav,ogg,pcm, orflac. - Enterprise security — see the Privacy Policy on the Shisa platform for how your data is handled.
Use cases
- Virtual assistants — power AI assistants, chatbots, and voice interfaces with natural-sounding speech for engaging interactions, including phone agents, smart-home assistants, and interactive voice response.
- Audiobooks and content — create audiobooks, podcasts, e-learning courses, and video voiceovers with professional-quality narration at scale.
- Accessibility — make content accessible with screen readers, news-article audio, document narration, and navigation assistance.
Next steps
- Quickstart — generate your first audio with curl, Python, or JavaScript.
- Voices — browse the voice catalogue and the
GET /tts/voicesendpoint. - API reference — full endpoint, parameter, response, and error documentation.
- Pricing — how TTS usage is billed.