Transparent pricing with no hidden fees. Pay as you go.
| Model | Spec | PoYo price | Official price | You save |
|---|---|---|---|---|
Text to speech | $0.120 24 credits per 1000 characters | $0.150 Reference | 20% |
* Actual fees are based on the final output.
Complete guide to using Affordable Gemini 3.1 Flash TTS API for Precise Audio-Tag Voice Control
Affordable Gemini 3.1 Flash TTS API for Precise Audio-Tag Voice Control
Available Gemini 3.1 Flash TTS API Models on PoYo
Why Gemini 3.1 Flash TTS Is Different
Gemini 3.1 Flash TTS is built for speech that follows direction, switches delivery, and supports multilingual product workflows without a complex voice pipeline.
01
Granular Audio Tags
- Control emotion, pacing, and non-verbal delivery cues.
- Use tags for stories, product videos, podcasts, and character lines.
- Keep expressive speech direction close to the script.

02
Two-Speaker Dialogue
- Create host and guest podcast segments.
- Prototype character dialogue and assistant conversations.
- Configure speaker aliases and voices for each dialogue role.

What You Can Build
Podcasts And Narration
Generate expressive host reads, intro segments, audiobook narration, and editorial audio with guided delivery.
Learning Content
Turn course scripts, onboarding lessons, and training material into clear multilingual spoken audio.
Character Dialogue
Prototype game lines, roleplay flows, story scenes, and two-speaker conversations with distinct voices.
Localized Product Voice
Create app text inputs, support messages, product demos, and accessibility speech in multiple languages.
Gemini 3.1 Flash TTS Pricing: PoYo vs fal.ai
| Feature | PoYo | Fal.ai |
|---|---|---|
| Public model ID | gemini-3-1-flash-tts | Not exposed |
| Fal.ai | $0.12 per 1000 characters | $0.15 per 1000 characters |
| PoYo credits | 24 credits per 1000 characters | N/A |
| Core controls | text, style_instructions, voice, language_code, speakers, temperature, output_format | Text input with style, voice, language, speaker, temperature, and format controls |
| Workflow | Submit task, poll standard status, download audio files | Reference queue workflow |
PoYo pricing uses 24 credits per 1000 characters. One credit is priced at $0.005.
How to Use Gemini 3.1 Flash TTS on PoYo
- Get your API key: Create a PoYo account and generate an API key from the dashboard. Get API Key ->
- Write the text: Enter the speech text and add audio tags such as [laughing], [sigh], [whispering], or [short pause] when you need expressive delivery.
- Choose voice settings: Select a Gemini voice, optional language_code, style_instructions, temperature, output_format, or configure exactly two speakers with speaker_id and voice.
- Submit and retrieve: Submit through PoYo's generate API, then poll the standard task status endpoint for audio output. View API Docs ->
Gemini 3.1 Flash TTS FAQ
What is Gemini 3.1 Flash TTS?
What should I put in the text field?
How do style_instructions work?
Does Gemini 3.1 Flash TTS support audio tags?
Which voices and languages are supported?
Can I create multi-speaker dialogue?
What output formats can I request?
How is billing calculated and how do I get results?
Why Use PoYo
Clear Per-Character Pricing
Use Gemini 3.1 Flash TTS at $0.12 per 1000 characters with predictable credit billing.
Unified Model Access
Use the same PoYo account, billing, status polling, and callback workflow across text, image, video, and audio models.
Production Workflow
Submit async jobs, poll stable task status, and retrieve finished speech files from one API surface.
Expressive Controls
Expose practical production controls for text, style, voice, language, speakers, temperature, and output format.