One or two directed speakers
Render narration or a two-person scene in one output. Give each speaker a voice, profile, accent, style, and pace, then assign every turn.
Create AI video drafts directly in Img2Vid. Choose a compatible model, review price and plan eligibility, then generate in the same workspace.
Gemini 3.1 Flash TTS is Google's preview model for controllable narration and dialogue. On Img2Vid, define one or two speakers, choose from 30 voices, direct style, pace, accent, and character, then write an ordered performance with audio tags. The result is structured speech generation rather than a generic prompt box.
Every example contains audio generated by Gemini 3.1 Flash TTS through the configured KIE task API. Img2Vid packages the original speech inside a 16:9 waveform video for consistent comparison across narration, bilingual delivery, distinct speakers, character acting, and product guidance.
Direction
Zephyr voice, American accent, Promo/Hype style, and natural pace. A premium launch read that builds from composed confidence to an energetic invitation.
Core Capabilities
Gemini 3.1 Flash TTS combines exact recitation with natural-language direction. Google documents 30 voices, more than 70 languages, up to two speakers, and audio tags. Img2Vid presents those capabilities as explicit controls.
Render narration or a two-person scene in one output. Give each speaker a voice, profile, accent, style, and pace, then assign every turn.
Gemini 3.1 Flash TTS exposes Google's 30 preset voices, including Zephyr, Puck, Kore, Fenrir, Aoede, and Charon. Preview them because names alone cannot describe a performance.
Generate speech across more than 70 languages with automatic language detection. Native reviewers should still check names, abbreviations, and code switching.
Use directions such as [whispers], [laughs], or [gasp] inside dialogue. Gemini 3.1 Flash TTS can change one moment without changing the global profile.
Describe setting, relationships, emotional arc, and tone separately from the exact words. Scene context complements per-speaker controls.
Google states that Gemini 3.1 Flash TTS audio includes SynthID. Creators still need consent, rights review, disclosure, and editorial checks.
Production Workflow
Strong speech starts with casting and direction. Img2Vid keeps those decisions visible so each revision can target one variable.
Choose one or two voices, then define role, energy, age impression, and delivery without imitating a real person.
Select accent, style, pace, and temperature. Keep scene direction compatible with the dialogue instead of giving contradictory instructions.
Assign each turn to a speaker and write the exact text. Img2Vid allows eight turns of 250 characters as a product guardrail, not a Google limit.
Review pronunciation, speaker separation, timing, and emotion. Change one Gemini 3.1 Flash TTS control at a time to isolate improvements.
Practical Uses
Gemini 3.1 Flash TTS suits controlled recitation rather than open microphone conversation. These uses benefit from a known script and explicit roles.
Create scripted intros, recaps, explainers, or two-host prototypes. Separate profiles keep the conversation intentional.
Dialogue · Editorial
Direction
Two studio hosts, distinct energy, natural timing, no announcer exaggeration.
Direct a narrator or short scene with atmosphere, pacing, and tags. Gemini 3.1 Flash TTS helps test prose before full production.
Narration · Story
Direction
Restrained narrator, intimate distance, slower pace, controlled suspense.
Turn onboarding and help scripts into clear guidance. Keep a stable global tone while tags emphasize key actions.
Onboarding · Training
Direction
Empathetic guide, concise phrasing, pause between each action.
Test game dialogue and pitch scenes with contrasting speakers. Final casting still requires creative and rights review.
Games · Previsualization
Direction
Weary guardian and urgent traveler, clear tension, final whispered reveal.
Create multilingual review drafts or code-switching lessons. Native speakers should approve pronunciation and cultural tone.
Localization · Learning
Direction
Patient teacher, clear Mandarin phrase, natural English response.
Generate scripted weather, accessibility, or in-app updates. Verify dynamic facts before they reach Gemini 3.1 Flash TTS.
Product · Accessibility
Direction
Calm neutral delivery, factual language, no dramatic emphasis.
Model Positioning
Compare workflow fit, not unsupported absolute rankings. Gemini 3.1 Flash TTS emphasizes Google-native control and multilingual scripts; other platforms may prioritize cloning or streaming.
| Decision | Gemini 3.1 Flash TTS | ElevenLabs | PlayHT |
|---|---|---|---|
| Primary fit | Directed scripted narration and two-speaker scenes | Broad voice platform and voice design workflows | Streaming and conversational speech workflows |
| Current Img2Vid controls | 30 voices, profiles, accent, style, pace, tags | Not configured on this page | Not configured on this page |
| Speaker structure | One or two speakers in a single ordered request | Varies by selected product and model | Varies by selected model and endpoint |
| Voice cloning here | Not exposed by this integration | Available in its broader platform | Available in its broader platform |
FAQ
Practical answers about the model and the exact Gemini 3.1 Flash TTS configuration currently available on Img2Vid.
Create on Img2Vid
Define each voice, write the dialogue, and use Gemini 3.1 Flash TTS to turn a script into expressive audio.