announcements

Poyo.ai August 2026 Update: 8 New Models for Image, Video, Chat, and Music

Poyo.ai Team
9 min read
Share:

A dark AI studio with holographic screens for video, image, and music generation

August 2026 added eight production models to Poyo.ai, covering image, video, and chat in one month. The same month also upgraded Suno Music into a single workspace and expanded Grok Imagine Video 1.5 with additional workflows and 1080p output.

This monthly update covers only models that are live on Poyo.ai today. Prices below match the public model pages as of September 2, 2026. Confirm the current rate on the model page before you ship a production workflow.

How to choose

Start with the job, then open the matching model page.

If you needStart withWhy
Lower-cost video, 2–30 secondsWan 3.0Text, image, and reference workflows at 480p, 720p, or 1080p
Faster Wan 3.0 turnaroundWan 3.0 Video PrimeSame workflow family, priced for speed
Longer clips and many referencesSeedance 2.54–30 seconds and up to 50 reference files
First/last frames, extension, keyframes, native audioFLUX 3Five video workflows in one family
Dense layouts and multilingual text in imagesQwen Image 3Long prompts, fine text, 1K and 2K output
Text-to-image plus light reference editingGrok Imagine Image 2.0Up to 3 input images and 4 outputs per request
Long-running agents, coding, and knowledge workGrok 4.6Chat API at 20% below official token rates
Multimodal agents, tools, and long contextGemini 3.7 FlashChat API with cached-input pricing
Songs, sound effects, voices, and mashupsSuno Music24 music tools in one workspace

Video: the largest August change

Seedance 2.5

Seedance 2.5 went live on August 8 for longer, reference-heavy video. Use it when a clip needs more than a prompt: first and last frames, images, video, and audio can all travel in one request.

  • Model ID: seedance-2.5
  • Duration: 4–30 seconds
  • Output: 480p, 720p, or 1080p
  • References: up to 30 images, 10 videos, and 10 audio files, with 50 files maximum per request
  • Optional synchronized audio through generate_audio
OutputWith reference videoWithout reference video
480p17 credits / $0.085 per billed second28 credits / $0.14 per second
720p38 credits / $0.19 per billed second63 credits / $0.315 per second
1080p68.5 credits / $0.3425 per billed second114 credits / $0.57 per second

When a reference video is present, billed duration equals output duration plus reference-video duration. See the Seedance 2.5 API docs for request fields. For a side-by-side with MiniMax H3, read MiniMax H3 vs Seedance 2.5.

Wan 3.0 and Wan 3.0 Video Prime

Wan 3.0 and Wan 3.0 Video Prime went live on August 25. Both support text-to-video, image-to-video, and multimodal reference-to-video. Choose standard Wan 3.0 when cost matters more; choose Prime when generation time matters more.

Shared capabilities:

  • 2–30 second output at 480p, 720p, or 1080p
  • One or two images as the first frame and optional last frame
  • Up to 10 reference images, 5 reference videos, 5 reference audio clips, and one document or public webpage
  • Optional generated audio on all three workflows
ResolutionWan 3.0Wan 3.0 Video Prime
480p10 credits / $0.05 per second13.6 credits / $0.068 per second
720p20 credits / $0.10 per second28 credits / $0.14 per second
1080p40 credits / $0.20 per second56 credits / $0.28 per second

Model IDs:

  • Wan 3.0: wan3.0-text-to-video, wan3.0-image-to-video, wan3.0-reference-to-video
  • Wan 3.0 Video Prime: wan3.0-prime-text-to-video, wan3.0-prime-image-to-video, wan3.0-prime-reference-to-video

FLUX 3

FLUX 3 became a production video family on August 6, with five workflows instead of a single text-to-video endpoint.

  • flux-3/text-to-video
  • flux-3/image-to-video
  • flux-3/first-last-frame-to-video
  • flux-3/extend-video
  • flux-3/keyframes-to-video

Standard workflows are 34 credits / $0.17 per second at 720p and 58 credits / $0.29 per second at 1080p. Video extension is 82 credits / $0.41 per second at 720p and 106 credits / $0.53 per second at 1080p. Native audio is part of the family.

Use FLUX 3 when you need shot control, continuation, or placed keyframes rather than a one-shot generation. Deeper walkthroughs: FLUX 3: Everything You Need to Know and FLUX 3 vs FLUX 2.

Grok Imagine Video 1.5 upgrade

This is an upgrade, not a new catalog entry. Grok Imagine Video 1.5 now covers text-to-video, single-frame image-to-video, and reference-to-video with 1–7 images. Output includes 480p, 720p, and 1080p.

ItemPrice
480p14.5 credits / $0.0725 per second
720p25 credits / $0.125 per second
1080p text or image mode45 credits / $0.225 per second
Input image2 credits / $0.01 per image

Image: two complementary options

Qwen Image 3 and Qwen Image 3 Pro

Qwen Image 3 moved from a preview page to a live generator on August 6. It is the better fit for long prompts, information-dense layouts, small text, and multilingual graphics.

  • Model IDs: qwen-image-3, qwen-image-3-pro
  • Output: 1K or 2K
  • Tasks: text-to-image and image-to-image
Variant1K2K
Qwen Image 34.8 credits / $0.0244.8 credits / $0.024
Qwen Image 3 Pro6.4 credits / $0.03212 credits / $0.06

Input images add 0.5 credits / $0.0025 each. Full background: Qwen Image 3.0: Everything You Need to Know.

Grok Imagine Image 2.0

Grok Imagine Image 2.0 went live on August 15. One model ID handles text-to-image and editing with up to three input images.

  • Model ID: grok-imagine-image-2.0
  • Resolution: 1K or 2K
  • Quality: low or medium
  • Up to 4 images per request
  • Output: JPEG, PNG, or WebP
Quality / resolutionPrice per image
Low / 1K8 credits / $0.04
Medium / 1K12 credits / $0.06
Low / 2K12 credits / $0.06
Medium / 2K16 credits / $0.08
Input image add-on2 credits / $0.01 per input image

Chat: two new production models

Grok 4.6

Grok 4.6 went live on August 15 for long-running agents, coding, engineering, and knowledge work. Call it through the OpenAI-compatible Chat Completions API.

  • Model ID: grok-4.6
  • Input: 320 credits / $1.60 per 1M tokens (official $2.00)
  • Output, including reasoning tokens: 960 credits / $4.80 per 1M tokens (official $6.00)

That is 20% below the official token rates. Docs: OpenAI-compatible chat.

Gemini 3.7 Flash

Gemini 3.7 Flash also went live on August 15. Use it for agentic coding, multimodal reasoning, tool use, and long-context work when you want Google's GA Flash model rather than a larger flagship.

  • Model ID: gemini-3.7-flash
  • Input: 120 credits / $0.60 per 1M tokens (official $0.75)
  • Cached input: 12 credits / $0.06 per 1M tokens (official $0.075)
  • Output, including thinking tokens: 600 credits / $3.00 per 1M tokens (official $3.75)

Docs: Gemini native format.

Music: Suno workspace upgrade

On August 28, Suno Music became one workspace with 24 tools. You can switch between generation, editing, mashups, reusable voices, stem splitting, and format conversion without leaving the page. Each tool shows its model ID, price, and API docs.

New and upgraded capabilities include:

  • Generate Mashup: blend two audio tracks with style, vocal, model version, and duration controls. 12 credits / $0.06 per generation.
  • Generate Sounds: sound effects and ambience with loop, tempo, and key controls. 2.5 credits / $0.0125 per generation.
  • Suno Voice: validate a voice, create a reusable Voice ID, regenerate the verification phrase, and check availability. Voice operations are free.
  • Generate MIDI: 1 credit / $0.005 per generation.
  • Music generation and extension now support V5, V5.5, Style Persona, Voice Persona, and duration up to 360 seconds.
  • Editing covers, extensions, section replacement, added vocals, added instrumentals, vocal separation, and 12-stem splits.

API docs: Generate Music.

Platform changes you can use now

These are the August product changes that affect how you work on Poyo.ai:

  • A dashboard overview that keeps history and usage in one place.
  • Generation history that shows result previews again.
  • A clearer model market, with live discount pricing on model cards when a promotion is active.
  • New developer entry points: MCP, CLI, Skill, plus n8n and ComfyUI guides.
  • Vietnamese and Indonesian site locales.

August 2026 model cheat sheet

Live dateModelTypeModel IDsPage
Aug 6Qwen Image 3 / ProImageqwen-image-3, qwen-image-3-pro/models/qwen-image-3
Aug 6FLUX 3Videoflux-3/text-to-video and four related endpoints/models/flux-3
Aug 8Seedance 2.5Videoseedance-2.5/models/seedance-2-5
Aug 15Grok 4.6Chatgrok-4.6/models/grok-4-6
Aug 15Gemini 3.7 FlashChatgemini-3.7-flash/models/gemini-3-7-flash
Aug 15Grok Imagine Image 2.0Imagegrok-imagine-image-2.0/models/grok-imagine-image-2
Aug 25Wan 3.0Videowan3.0-text-to-video, wan3.0-image-to-video, wan3.0-reference-to-video/models/wan-3-0
Aug 25Wan 3.0 Video PrimeVideowan3.0-prime-text-to-video, wan3.0-prime-image-to-video, wan3.0-prime-reference-to-video/models/wan-3-0-video-prime

August upgrades, not counted in the eight launches: Grok Imagine Video 1.5 and Suno Music.

How to start

  1. Open the model page and run one test generation with the live tool.
  2. Copy the model ID into your existing Poyo workflow: async submit and status polling for media, or Chat Completions / Gemini native format for chat.
  3. If you are wiring this into an internal tool, start from MCP, CLI, n8n, or ComfyUI.

After August: early September launches

Two more models went live just after this August update window:

They are not part of the August eight. Use the links above if you are already evaluating the newest video and chat options.

Share: