PoYo.ai AI APIs for production teams

PoYo.ai AI API Platform

Build image, video, music, chat, and 3D generation into your product with one async API. Access premium models, use polling or webhooks, and pay only for successful generations.

StablePolling + WebhookFailed tasks are not charged
ChatOpenAI

GPT-6.1 Sol

GPT-6.1 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

From$1.60Per 1M Input20% OFFTry it
ChatAnthropic

Claude Sonnet 5.5

Claude Sonnet 5.5 API on PoYo. Use Chat Completions or Claude Messages at $1.6 input and $8 output per million tokens, 20% below official standard rates.

From$1.60Per 1M Input20% OFFTry it
Text to VideoMiniMax

MiniMax H3 Max

MiniMax H3 Max generates 5–15 second videos from text, first/last frames, or image, video and audio references at 480p, 768p and 1080p.

From$0.05Per SecTry it
Text to VideoMiniMax

MiniMax H3 Max Turbo

MiniMax H3 Max Turbo generates 5–15 second videos from text or first/last frames at 480p, 768p and 1080p.

From$0.025Per SecTry it
Text to ImageAlibaba

Qwen Image 2.1

Alibaba's 7B open-weight unified text-to-image and editing model with native RGBA transparency, up to 10 reference images, and 2K resolution.

From$0.03Per ReqTry it
ChatOpenAI

GPT-6 Luna

GPT-6 Luna API on PoYo. Use Chat Completions or Responses at $0.08 input and $0.4 output per million tokens, 20% below official standard rates.

From$0.08Per 1M Input20% OFFTry it
ChatOpenAI

GPT-6 Sol

GPT-6 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

From$1.60Per 1M Input20% OFFTry it
ChatAnthropic

Claude Opus 5.5

Claude Opus 5.5 API on PoYo. Use Chat Completions or Claude Messages at $3.2 input and $16 output per million tokens, 20% below official standard rates.

From$3.20Per 1M Input20% OFFTry it
ChatxAI

Grok 4.7

Grok 4.7 API on PoYo. Use Chat Completions at $1.6 input and $4.8 output per million tokens, 20% below official standard rates.

From$1.60Per 1M Input20% OFFTry it
ChatDeepSeek

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

From$0.30Per 1M InputTry it
Text to ImageOpenAI

GPT Image 2.5

OpenAI image generation and editing with high-fidelity reference images, precise inpainting, multi-turn consistency, and transparent backgrounds. Explore Flare and Sunburst and compare GPT Image 2.5 with GPT Image 2 on PoYo.

From$0.007Per ReqTry it
ChatOpenAI

GPT-6 Astra

GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

From$8.00Per 1M Input20% OFFTry it
ChatAnthropic

Claude Fable 5.1

Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

From$8.00Per 1M Input20% OFFTry it
ChatGoogle

Gemini 3.8 Flash

Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

From$0.06Per 1M Input20% OFFTry it
Text to VideoGoogle

Gemini Omni 1.1 Flash

Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

From$0.225Per ReqTry it
Text to VideoAlibaba

Wan 3.0

Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

From$0.05Per SecTry it
Text to VideoAlibaba

Wan 3.0 Video Prime

Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

From$0.068Per SecTry it
Text to VideoSeedance

Seedance 2.5

ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

From$0.085Per Sec36% OFFTry it
Image to VideoAlibaba

Wan Animate 2

Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

Coming soonTry it
ChatxAI

Grok 4.6

xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

From$1.60Per 1M Input20% OFFTry it
ChatGoogle

Gemini 3.7 Flash

Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

From$0.06Per 1M Input20% OFFTry it
Text to ImagexAI

Grok Imagine Image 2.0

xAI image generation and editing with up to three input images, five supported aspect ratios, 1K or 2K resolution, selectable quality, and up to four outputs.

From$0.04Per ReqTry it
Text to VideoBlack Forest Labs

FLUX 3

Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

From$0.17Per SecTry it
Text to VideoMiniMax

MiniMax H3 / Hailuo 03

MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.

From$0.105Per Sec20% OFFTry it