Back to library
AI
best practice

Control Audio Style and Tone Using Natural Language Descriptions

Adjust speed and emotion using plain English instead of technical settings.

Direct the 'style, accent, pace, and tone' of your audio by including descriptive natural language instructions in your TTS request.

Google Gemini

The Scenario

You need a voiceover for a marketing video that sounds energetic and urgent, or an audiobook narration that needs to be slow and soothing.

Before & after

The old way

Typically, you had to manually adjust pitch, speed, and equalization in post-production software, which could take 15–20 minutes per clip.

With AI

Simply specify the desired emotional tone or speed in the prompt to Gemini-TTS, and receive matching audio in approximately 30 seconds.

The Prompt

Convert the following text to speech: [PASTE TEXT HERE]. Delivery style: [SPECIFY STYLE, E.G., ENERGETIC, WHISPERED, OR FAST-PACED]. Use a [SPECIFY ACCENT] accent.

The 'controllable' nature of Gemini-TTS means you don't need technical sliders; you can describe the delivery style (e.g., 'excited', 'slow', 'British accent') in plain English.

Source

Release notes  |  Gemini API  |  Google AI for Developers
"Text-to-speech (TTS) generation is controllable... you can use natural language to structure interactions and guide the style, accent, pace, and tone."