Skip to main content

Shisa TTS

Shisa TTS turns text into lifelike speech. It produces natural-sounding Japanese and English audio with genuine intonation and prosody — voices vetted by native Japanese speakers rather than English voices stretched to fit Japanese. Send text, pick a voice, and get back binary audio in the format you ask for.

POST https://api.shisa.ai/tts

Why Shisa TTS

  • Natural Japanese voices — highly realistic Japanese and English speech with natural intonation and prosody, not translated-sounding output.
  • Emotional tone and expressiveness — express joy, concern, excitement, or professionalism through emotional voice modulation to match your content's mood.
  • Custom voice cloning — clone your brand voice or create custom voices for consistent audio branding.
  • Real-time streaming — low-latency streaming for interactive applications and real-time conversations.
  • Multiple languages and formats — Japanese, English, and more with native pronunciation quality, delivered as mp3, wav, ogg, pcm, or flac.
  • Enterprise-grade — secure processing and dedicated support for enterprise needs.

Use cases

  • Virtual assistants — power AI assistants, chatbots, and voice interfaces with natural-sounding speech for engaging interactions, including phone agents, smart-home assistants, and interactive voice response.
  • Audiobooks and content — create audiobooks, podcasts, e-learning courses, and video voiceovers with professional-quality narration at scale.
  • Accessibility — make content accessible with screen readers, news-article audio, document narration, and navigation assistance.

Next steps

  • Quickstart — generate your first audio with curl, Python, or JavaScript.
  • Voices — browse the voice catalogue and the GET /tts/voices endpoint.
  • API reference — full endpoint, parameter, response, and error documentation.
  • Pricing — how TTS usage is billed.