GPT Image 2.5 API is live. Create and edit images →
GPT Image 2.5 API is live. Create and edit images →
PoYoPoYo
Market
Pricing
Enterprise
Ctrl+K
poyo.ai

One API. Every model.

Built for developers

All systems operational

Resources

  • Alternative Hub
  • Comparison Hub
  • Blog
  • Providers
  • Affiliate Program
  • Features
  • Sitemap

© 2025 PoYo API. All rights reserved.

Popular Modelshot

  • Seedance 2.5
  • MiniMax H3 / Hailuo 03
  • FLUX 3
  • Nano Banana Pro
  • Seedance 2
  • GPT Image 2

New Releasesnew

  • GPT-6.1 Sol
  • Claude Sonnet 5.5
  • MiniMax H3 Max
  • MiniMax H3 Max Turbo
  • Qwen Image 2.1
  • GPT-6 Luna

Coming Soonsoon

  • Wan Animate 2

All Providers

Google

  • Nano Banana Pro
  • Gemini Omni 1.1 Flash
  • Gemini 3.7 Flash
  • Gemini 3.8 Flash
  • Gemini 3.1 Flash TTS
  • Nano Banana
  • Veo 3.1
  • Veo 3.1 Official
  • Omni Flash
  • Nano Banana 2
  • Gemini 3.5 Flash
  • Nano Banana 2 Lite
  • Gemini 3 Series

Kling

  • Kling 3.0 Motion Control
  • Kling 3.0
  • Kling O3
  • Kling 1.6
  • Kling 2.1
  • Kling 2.6
  • Kling Avatar 2.0
  • Kling 2.5 Turbo Pro
  • Kling 3.0 Turbo
  • Kling 2.6 Motion Control
  • Kling O3 Image
  • Kling O1 Image

Anthropic

  • Claude Sonnet 5
  • Claude Fable 5.1
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 4.7
  • Claude 4.6 API
  • Claude 4.5 Series

MiniMax

  • MiniMax H3 / Hailuo 03
  • MiniMax H3 Max
  • MiniMax H3 Max Turbo
  • MiniMax Music 2.6
  • Hailuo 02
  • Hailuo 2.3

Tripo3D

  • Tripo3D H3.1
  • Tripo3D P1

Moonshot AI

  • Kimi K3

OpenAI

  • GPT Image 2
  • Sora 2 Official
  • GPT-6 Astra
  • GPT-6 Sol
  • GPT Image 2.5
  • GPT-5.6
  • GPT-6 Luna
  • GPT-6.1 Sol
  • GPT Image 1.5
  • GPT-4o Image
  • GPT-5.5
  • GPT-5.4
  • GPT-5.2

Seedream

  • Seedream 5.0 Pro
  • Seedream 4
  • Seedream 4.5
  • Seedream 5.0 Lite

Seedance

  • Seedance 2.5
  • Seedance 2
  • Seedance 2.0 Mini
  • Seedance 1.0 Pro
  • Seedance 1.5 Pro

DeepSeek

  • DeepSeek V4 Flash
  • DeepSeek V4 Pro
  • DeepSeek V4.1 Flash

Hunyuan

  • Hunyuan 3D v3.1

ElevenLabs

  • ElevenLabs Music
  • ElevenLabs TTS Turbo v2.5
  • ElevenLabs V3 TTS

Black Forest Labs

  • FLUX 3
  • FLUX.2
  • Flux Kontext
  • Flux Dev
  • Flux Schnell

Alibaba

  • Qwen Image 3.0
  • Happy Horse 1.1
  • Wan 2.7 Video
  • Wan 3.0
  • Wan 3.0 Video Prime
  • Qwen Image 2.1
  • Wan Animate 2
  • Wan 2.2 Fast
  • Z-Image
  • Wan 2.5
  • Wan 2.6
  • Wan 2.7 Image
  • Wan Animate
  • Happy Horse

xAI

  • Grok 4.6
  • Grok 4.7
  • Grok Imagine Image 2.0
  • xAI TTS 1
  • Grok Imagine
  • Grok Imagine Video 1.5
  • Grok Imagine Image Quality

Meshy

  • Meshy 6 3D

Runway

  • Runway Gen-4.5

PoYo

  • AI Video Background Removal
  • AI Video Upscaler
  • Image Translator
  • Video Translator
  • Suno Music
GitHubXInstagramYouTubeDiscordTelegramEmail
Customer SupportDocumentation
ImageVideo3DAudioAvatarTools
Loading playground...
Audio Generator Guide

Complete guide to using A Complete Audio Generator for Speech and Music

A Complete Audio Generator for Speech and Music

PoYo Audio Generator brings leading speech, voice, and music generation models into one focused workspace. Instead of switching between separate provider tools, you can choose ElevenLabs, Gemini, xAI, MiniMax, and other executable audio models directly inside the Input panel.

Each model keeps its native controls, so language, voice, emotion tags, speed, lyrics, instrumental mode, output format, preview, JSON response, and download workflows appear only when the selected model supports them.

Start Creating

Key Features of PoYo Audio Generator

Generate speech, voiceovers, and music, then preview, download, and move stable workflows into your API pipeline.

01

Expressive multilingual text-to-speech

Turn scripts, product copy, narration, and dialogue into natural audio. Supported models expose voice selection, language control, emotion tags, text normalization, speed, and timestamps when available.

  • Generate voiceovers from text prompts and scripts
  • Use language, voice, and delivery controls where supported
  • Create speech for apps, videos, courses, and agents
Expressive multilingual text-to-speech workflow

02

Prompt-to-music and song creation

Create music from prompts, lyrics, instrumental settings, and structured composition ideas. Use text-to-music models for demos, hooks, backing tracks, and creative direction before moving into production.

  • Generate songs, instrumentals, and music concepts
  • Control mood, genre, arrangement, lyrics, and vocals
  • Use model-specific composition fields when available
Prompt-to-music generation workflow

03

Model-specific quality, latency, and format controls

Switch between audio model families without losing their native controls. Fast TTS, expressive voice models, music generators, sample rate, bitrate, output format, and quality options stay attached to the active model.

  • Compare audio providers from one selector
  • Expose only parameters supported by the active model
  • Choose MP3, WAV, OGG, PCM, or provider-specific formats
Audio model controls and format options

04

Preview, download, and API handoff

Listen to generated audio directly in the browser, inspect structured JSON and timestamps when returned, download the final file, and move validated settings into API or agent workflows.

  • Preview audio outputs before downloading
  • Download production files for editing and publishing
  • Validate prompts and parameters before API integration
Audio preview and API handoff workflow

What Can You Create with PoYo Audio Generator?

01

Video voiceovers and dubbing

Create narration, product walkthrough audio, multilingual dubbing drafts, and social video voice tracks from scripts.

02

Podcasts, audiobooks, and course narration

Generate spoken lessons, chapter previews, podcast segments, and long-form narration tests before recording or mastering.

03

Multilingual app voices and accessibility

Prototype app voices, screen-reader style audio, onboarding narration, and localized speech experiences.

04

Music demos, songs, and instrumental tracks

Create song ideas, instrumental beds, hooks, demos, and backing tracks from prompts, lyrics, and style direction.

05

Social, game, and creative audio concepts

Explore character voices, game UI sounds, creator assets, and audio concepts for interactive media.

06

Developer, API, and agent workflows

Validate prompts, voices, formats, previews, and JSON responses before adding audio generation to an API product.

How to Use Audio Generator on PoYo

1. Choose an audio model in the Input panel. Pick the provider and executable model that matches your speech or music workflow.

2. Add text, lyrics, or music direction. Write the script, prompt, language, voice style, mood, instrumentation, or lyrics required by the selected model.

3. Adjust model-specific controls. Set only the options shown by the active model, such as voice, emotion tags, speed, sample rate, bitrate, output format, lyrics optimizer, or instrumental mode.

4. Run, preview, and download. Submit the generation, listen to the audio preview, inspect JSON if needed, and download the generated file.

Frequently Asked Questions about Audio Generator

What can Audio Generator create?

PoYo Audio Generator can create text-to-speech, voiceover audio, narration, dialogue, music demos, songs, and instrumental tracks depending on the selected model. The Input panel only shows the fields supported by the active audio model.

Which audio model should I choose first?

Start with ElevenLabs V3 TTS for expressive voiceovers, ElevenLabs TTS Turbo v2.5 or xAI TTS 1 for faster speech tests, Gemini 3.1 Flash TTS for style-instructed speech, ElevenLabs Music for structured music generation, MiniMax Music 2.6 for prompt and lyrics workflows, and Generate Music for broad music creation.

How realistic or natural is the generated audio?

Quality depends on the model, prompt, language, voice, and output settings. Modern TTS models can sound highly natural, but production use still benefits from listening checks, pronunciation edits, pacing review, and final mastering.

Can I control language, emotion, voice, speed, or music style?

Yes, when the selected model supports those controls. Some speech models expose language, voice, speed, stability, style, emotion tags, or timestamps. Music models may expose genre, mood, lyrics, instrumental mode, section structure, and output format.

What audio formats can I preview or download?

The gallery previews common audio files in the browser. Depending on the provider response, downloads can include MP3, WAV, OGG, Opus, PCM, mu-law, A-law, timestamps, or supporting files.

How do I improve audio quality?

Use clean scripts, clear pronunciation hints, specific voice direction, realistic pacing, and concise emotion tags. For music, describe genre, mood, tempo, instruments, vocals, structure, lyrics, and what to avoid.

How long does generation take and how is pricing calculated?

Generation time and cost vary by model, text length, duration, format, music length, quality settings, and provider. Check the active model cost in the Input footer before running longer narration or music generations.

Can I use generated audio commercially?

Commercial use depends on the selected model, provider terms, your account plan, and the rights attached to scripts, lyrics, voices, references, and uploaded material. Review the relevant terms and avoid content you do not have permission to use.

Why Use PoYo for Audio Generation?

01

One workspace for audio models

Compare speech, voice, and music generators without rebuilding prompts across provider tools.

02

Native audio parameters

Each model keeps its own voice, language, style, lyrics, format, and result workflow.

03

Browser audio previews

Listen to generated audio directly in the output gallery before downloading files.

04

Playground and API workflow

Test scripts, prompts, voices, formats, and JSON responses before moving into API-powered products.