Back to library
AI
prompt

Direct AI Narrators with Natural Language Voice Controls

Direct AI voices using natural language instead of complex code.

Use natural language prompts to control the style, accent, pace, and tone of generated audio using Gemini-TTS for more human-like narration.

Google Gemini

The Scenario

You are producing a podcast or an audiobook where the narration needs to convey specific emotions or switch between multiple distinct character voices naturally.

Before & after

The old way

Standard TTS tools often require fiddling with SSML tags or manual speed/pitch sliders, taking 15–20 minutes to get the 'vibe' right through trial and error.

With AI

Use natural language prompts with Gemini-TTS to specify the exact accent, speed, and emotional tone in seconds. The AI handles the synthesis in under a minute without specialized audio engineering.

The Prompt

Generate audio for the following text. Use a [TONE: e.g., cheerful and high-energy] tone with a [PACING: e.g., slow and deliberate] pace. For the second speaker, use a [STYLE: e.g., skeptical and deep-voiced] profile. Text: [INSERT_TEXT_HERE]

Unlike traditional TTS which uses rigid settings, Gemini TTS is 'steerable' via natural language, allowing for nuanced control over multi-speaker dialogues and emotional expression.

Source

Release notes  |  Gemini API  |  Google AI for Developers
"TTS through the Gemini API is tailored for scenarios that require exact text recitation with fine-grained control over style and sound."