PoYo.ai AI API Platform
Build image, video, music, chat, and 3D generation into your product with one async API. Access premium models, use polling or webhooks, and pay only for successful generations.

DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

GPT Image 2.5
OpenAI image generation and editing with high-fidelity reference images, precise inpainting, multi-turn consistency, and transparent backgrounds. Explore Flare and Sunburst and compare GPT Image 2.5 with GPT Image 2 on PoYo.

GPT-6 Astra
GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

Claude Fable 5.1
Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

Gemini 3.8 Flash
Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

AlibabaWan 3.0
Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

AlibabaWan 3.0 Video Prime
Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

Seedance 2.5
ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

AlibabaWan Animate 2
Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

Grok 4.6
xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

Gemini 3.7 Flash
Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

Grok Imagine Image 2.0
xAI image generation and editing with up to three input images, five supported aspect ratios, 1K or 2K resolution, selectable quality, and up to four outputs.

FLUX 3
Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

MiniMax H3 / Hailuo 03
MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.

Claude Opus 5
Anthropic's flagship agentic model for advanced coding, long-horizon execution, computer use, deep research, and professional knowledge work.

Qwen Image 3.0
Alibaba's third-generation image model for long prompts, information-dense layouts, fine text rendering, multilingual visuals, and knowledge-rich image generation.

Kimi K3
Moonshot AI's 2.8T-parameter flagship multimodal model with a 1M-token context window, built for long-horizon coding, deep reasoning, knowledge work, and tool-using agents.

GPT-5.6
GPT-5.6 API access for Sol, Terra, and Luna through the Responses API, from high-throughput automation to frontier coding and agent workflows.

Seedream 5.0 Pro
ByteDance's flagship image model for precise prompt following, dense layouts, multilingual text rendering, and region-precise editing.

Nano Banana 2 Lite
Google's fastest and most cost-efficient Gemini Image model, built for rapid ideation, high-throughput image generation, low-latency creative workflows, and scaled production pipelines.

Claude Sonnet 5
Claude Sonnet 5 API access for agentic coding, long-running agents, browser and computer use, professional workflows, 1M context, and up to 128K output.

Seedance 2.0 Mini
ByteDance Seedance 2.0 Mini video generation for lower-cost text-to-video, image-to-video, and multimodal reference workflows with 480p and 720p output.

Happy Horse 1.1
Alibaba Happy Horse 1.1 video generation with text-to-video, image-to-video, and reference-to-video workflows.
