Back to library
AI
tutorial

Refine AI Videos Iteratively Using Conversational Editing

Use conversational editing to refine AI video clips without starting over.

Leverage the 'Interactions API' with Gemini Omni Flash to perform multi-turn video editing. You can build on previous generations to change costumes, settings, or timing while maintaining character consistency.

Google Gemini

The Scenario

You need to create a short marketing video or social media clip, but the initial AI generation is slightly off. Instead of regenerating everything, you want to tweak specific elements like costume or lighting.

Before & after

The old way

Previously, you would have to manually edit video clips in software like Premiere Pro or restart the entire AI generation with a new prompt, costing 30–60 minutes of trial and error.

With AI

By using the Gemini Omni Flash API, you can generate a 10-second 720p base clip and then refine it through conversational commands like "make the lighting warmer." This process takes roughly 3–5 minutes per iteration.

The Prompt

Initial generation: "Generate a 10-second video of a [CHARACTER/SUBJECT] in a [SETTING] with [ANIMATION STYLE]."

Follow-up edit: "Keep the character consistency but change the background to [NEW SETTING] and make the camera movement [SLOWER/FASTER]."

Gemini Omni Flash allows for iterative, conversational adjustments to video output, meaning you don't have to start from scratch if one detail is wrong. This is powered by the Interactions API.

Source

Release notes  |  Gemini API  |  Google AI for Developers
"Every prompt builds on the last. Change a costume, retime an action, or swap a setting on the same shot."