Text to Speech API

AI APIs for developers

Text to Speech API

Add AI voice generation to your product with the PoYo Text to Speech API. Compare TTS models, voice controls, languages, audio formats, timestamps and pricing for narration workflows.

Text to Speech Model APIs - Pricing and Model Fit

Compare providers, supported inputs, output formats and prices before choosing a model.

Exact task match

Models appear here only when their catalog metadata includes Text to Speech.

Provider comparison

Review model families across providers without leaving the task directory.

API-ready paths

Each model card links to PoYo pricing, examples, playground controls and API details.

ModelProviderTask typesPrice
ElevenLabs TTS Turbo v2.5

elevenlabs-tts-turbo-2-5

ElevenLabsText to Speech$0.04/1,000 characters
ElevenLabs V3 TTS

elevenlabs-v3-tts

ElevenLabsText to Speech$0.08/1,000 characters
Gemini 3.1 Flash TTS

gemini-3-1-flash-tts

GoogleText to Speech$0.12/1,000 characters
xAI TTS 1

xai-tts-1

xAIText to Speech$0.012/1,000 characters

Frequently asked questions

What is a text to speech API?

+

A text to speech API generates spoken audio from written text. It enables narration and voice features in an application without recording every script manually. The selected model determines available voices, languages, speaking controls and output formats.

How do I choose a TTS API model?

+

Test your actual scripts, including names, numbers and long sentences. Compare pronunciation, naturalness, expressive control, voice consistency and cost. Confirm the required language and audio format, and check the documented request flow if response time is important to your product.

Can I choose the voice, speaking speed and emotion?

+

Available controls vary by model. An endpoint may offer voice presets, speed, stability, style instructions or audio tags. Use documented fields and supported values; a model without a speed or emotion control should not be expected to interpret an arbitrary parameter.

Which languages does the text to speech API support?

+

Language coverage and pronunciation quality depend on the model and voice. Check the documented language options or detection behavior, then test native text and mixed-language passages. A multilingual label does not guarantee equal quality for every accent or proper name.

How can I fix mispronounced names, abbreviations or numbers?

+

Rewrite the passage for spoken language, expand ambiguous abbreviations and test a short sentence. Use documented pronunciation or normalization controls where available. Keep a reviewed text version for narration rather than assuming a written display string is always the best speech input.

Can I generate long-form narration?

+

Check the endpoint text limit and supported duration behavior. Divide long scripts at natural paragraph or sentence boundaries, reuse voice settings and review transitions after assembly. Preserve context fields where the model supports them to help continuity across separate requests.

Does a TTS API automatically clone a voice?

+

No. Selecting a synthetic voice or preset is different from creating a clone from a recording. Voice cloning requires a specifically supported workflow and appropriate permission. Check the available endpoint before designing a product around a custom voice identity.

Can I get word timestamps or generate dialogue?

+

Some models provide timestamps or multiple-speaker controls, while others do not. Confirm the output fields and request options on the model page. Validate alignment on your own script before using timestamps for subtitles or synchronized text highlighting.

Which audio formats can I request?

+

Use the formats and quality options documented for the chosen TTS model. Playback, telephony and editing can require different codecs or sample rates. Inspect the completed output and convert it in a separate media step when your application needs an unsupported format.

How do I retrieve speech generated through PoYo?

+

Use the selected model generation endpoint, save task_id and retrieve the completed audio through status queries or a supported callback_url. Do not infer a streaming endpoint from the phrase low-latency TTS; check the exact PoYo interface documented for that model.

How is text to speech API pricing calculated?

+

TTS pricing can use characters, tokens, generation or another model-specific unit. Check how the chosen tier counts usage and how optional settings affect it. Estimate the complete script and expected retries, rather than comparing raw prices with different billing units.

How can I connect TTS with a talking avatar API?

+

Generate and review the speech first, then confirm that its format and length meet the avatar endpoint requirements. Supply the audio with a suitable portrait and track the avatar task separately. Keep narration text and voice settings so revisions can reproduce the intended delivery.