Refine AI Videos Iteratively Using Conversational Editing
Use conversational editing to refine AI video clips without starting over.
Leverage the 'Interactions API' with Gemini Omni Flash to perform multi-turn video editing. You can build on previous generations to change costumes, settings, or timing while maintaining character consistency.
The Scenario
You need to create a short marketing video or social media clip, but the initial AI generation is slightly off. Instead of regenerating everything, you want to tweak specific elements like costume or lighting.
Before & after
Previously, you would have to manually edit video clips in software like Premiere Pro or restart the entire AI generation with a new prompt, costing 30–60 minutes of trial and error.
By using the Gemini Omni Flash API, you can generate a 10-second 720p base clip and then refine it through conversational commands like "make the lighting warmer." This process takes roughly 3–5 minutes per iteration.
The Prompt
Initial generation: "Generate a 10-second video of a [CHARACTER/SUBJECT] in a [SETTING] with [ANIMATION STYLE]." Follow-up edit: "Keep the character consistency but change the background to [NEW SETTING] and make the camera movement [SLOWER/FASTER]."
Gemini Omni Flash allows for iterative, conversational adjustments to video output, meaning you don't have to start from scratch if one detail is wrong. This is powered by the Interactions API.
Source
Release notes | Gemini API | Google AI for Developers"Every prompt builds on the last. Change a costume, retime an action, or swap a setting on the same shot."
