PoYo.ai AI APIs for production teams

PoYo.ai AI API Platform

Build image, video, music, chat, and 3D generation into your product with one async API. Access premium models, use polling or webhooks, and pay only for successful generations.

StablePolling + WebhookFailed tasks are not charged
ChatDeepSeek

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

From$0.30Per 1M InputTry it
Text to ImageOpenAI

GPT Image 2.5

OpenAI image generation and editing with high-fidelity reference images, precise inpainting, multi-turn consistency, and transparent backgrounds. Explore Flare and Sunburst and compare GPT Image 2.5 with GPT Image 2 on PoYo.

From$0.007Per ReqTry it
ChatOpenAI

GPT-6 Astra

GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

From$8.00Per 1M Input20% OFFTry it
ChatAnthropic

Claude Fable 5.1

Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

From$8.00Per 1M Input20% OFFTry it
ChatGoogle

Gemini 3.8 Flash

Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

From$0.06Per 1M Input20% OFFTry it
Text to VideoGoogle

Gemini Omni 1.1 Flash

Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

From$0.225Per ReqTry it
Text to VideoAlibaba

Wan 3.0

Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

From$0.05Per SecTry it
Text to VideoAlibaba

Wan 3.0 Video Prime

Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

From$0.068Per SecTry it
Text to VideoSeedance

Seedance 2.5

ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

From$0.085Per Sec36% OFFTry it
Image to VideoAlibaba

Wan Animate 2

Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

Coming soonTry it
ChatxAI

Grok 4.6

xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

From$1.60Per 1M Input20% OFFTry it
ChatGoogle

Gemini 3.7 Flash

Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

From$0.06Per 1M Input20% OFFTry it
Text to ImagexAI

Grok Imagine Image 2.0

xAI image generation and editing with up to three input images, five supported aspect ratios, 1K or 2K resolution, selectable quality, and up to four outputs.

From$0.04Per ReqTry it
Text to VideoBlack Forest Labs

FLUX 3

Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

From$0.17Per SecTry it
Text to VideoMiniMax

MiniMax H3 / Hailuo 03

MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.

From$0.105Per Sec20% OFFTry it
ChatAnthropic

Claude Opus 5

Anthropic's flagship agentic model for advanced coding, long-horizon execution, computer use, deep research, and professional knowledge work.

From$2.00Per 1M Input60% OFFTry it
Text to ImageAlibaba

Qwen Image 3.0

Alibaba's third-generation image model for long prompts, information-dense layouts, fine text rendering, multilingual visuals, and knowledge-rich image generation.

From$0.024Per Req40% OFFTry it
ChatMoonshot AI

Kimi K3

Moonshot AI's 2.8T-parameter flagship multimodal model with a 1M-token context window, built for long-horizon coding, deep reasoning, knowledge work, and tool-using agents.

From$2.28Per 1M Input24% OFFTry it
ChatOpenAI

GPT-5.6

GPT-5.6 API access for Sol, Terra, and Luna through the Responses API, from high-throughput automation to frontier coding and agent workflows.

From$0.056Per 1M Input72% OFFTry it
Text to ImageSeedream

Seedream 5.0 Pro

ByteDance's flagship image model for precise prompt following, dense layouts, multilingual text rendering, and region-precise editing.

From$0.075Per Req20% OFFTry it
Text to ImageGoogle

Nano Banana 2 Lite

Google's fastest and most cost-efficient Gemini Image model, built for rapid ideation, high-throughput image generation, low-latency creative workflows, and scaled production pipelines.

From$0.025Per Req26% OFFTry it
ChatAnthropic

Claude Sonnet 5

Claude Sonnet 5 API access for agentic coding, long-running agents, browser and computer use, professional workflows, 1M context, and up to 128K output.

From$0.85Per 1M Input57% OFFTry it
Text to VideoSeedance

Seedance 2.0 Mini

ByteDance Seedance 2.0 Mini video generation for lower-cost text-to-video, image-to-video, and multimodal reference workflows with 480p and 720p output.

From$0.03Per Sec20% OFFTry it
Text to VideoAlibaba

Happy Horse 1.1

Alibaba Happy Horse 1.1 video generation with text-to-video, image-to-video, and reference-to-video workflows.

From$0.11Per Sec20% OFFTry it