PoYo.ai AI API Platform
Build image, video, music, chat, and 3D generation into your product with one async API. Access premium models, use polling or webhooks, and pay only for successful generations.

GPT-6.1 Sol
GPT-6.1 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

Claude Sonnet 5.5
Claude Sonnet 5.5 API on PoYo. Use Chat Completions or Claude Messages at $1.6 input and $8 output per million tokens, 20% below official standard rates.

MiniMax H3 Max
MiniMax H3 Max generates 5–15 second videos from text, first/last frames, or image, video and audio references at 480p, 768p and 1080p.

MiniMax H3 Max Turbo
MiniMax H3 Max Turbo generates 5–15 second videos from text or first/last frames at 480p, 768p and 1080p.

Qwen Image 2.1
Alibaba's 7B open-weight unified text-to-image and editing model with native RGBA transparency, up to 10 reference images, and 2K resolution.

GPT-6 Luna
GPT-6 Luna API on PoYo. Use Chat Completions or Responses at $0.08 input and $0.4 output per million tokens, 20% below official standard rates.

GPT-6 Sol
GPT-6 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

Claude Opus 5.5
Claude Opus 5.5 API on PoYo. Use Chat Completions or Claude Messages at $3.2 input and $16 output per million tokens, 20% below official standard rates.

Grok 4.7
Grok 4.7 API on PoYo. Use Chat Completions at $1.6 input and $4.8 output per million tokens, 20% below official standard rates.

DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

GPT Image 2.5
OpenAI image generation and editing with high-fidelity reference images, precise inpainting, multi-turn consistency, and transparent backgrounds. Explore Flare and Sunburst and compare GPT Image 2.5 with GPT Image 2 on PoYo.

GPT-6 Astra
GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

Claude Fable 5.1
Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

Gemini 3.8 Flash
Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

AlibabaWan 3.0
Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

AlibabaWan 3.0 Video Prime
Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

Seedance 2.5
ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

AlibabaWan Animate 2
Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

Grok 4.6
xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

Gemini 3.7 Flash
Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

Grok Imagine Image 2.0
xAI image generation and editing with up to three input images, five supported aspect ratios, 1K or 2K resolution, selectable quality, and up to four outputs.

FLUX 3
Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

MiniMax H3 / Hailuo 03
MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.
