图像生成视频 AI 模型

图像生成视频模型

图像生成视频 AI 模型

把产品图、人物图、角色设定或首帧画面动画化为视频,让运动围绕已有视觉素材展开。

匹配模型

38 个模型

MiniMax H3 / Hailuo 03
Text to Video

$0.105 / second

hailuo-03

MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.

Seedance 2.5
Text to Video

$0.085 / second

seedance-2.5

ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

FLUX 3
Text to Video

$0.17 / second

flux-3/text-to-video

Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

Seedance 2
Text to Video

$0.045 / second

seedance-2

ByteDance Seedance 2 video generation with text-to-video, image-to-video, multimodal reference-to-video, video-to-video, and audio-to-video workflows using image, video, and audio references with optional native audio.

Gemini Omni 1.1 Flash
Text to Video

$0.225 / generation

gemini-omni-1.1-flash

Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

Seedance 2.0 Mini
Text to Video

$0.03 / second

seedance-2-mini

ByteDance Seedance 2.0 Mini video generation for lower-cost text-to-video, image-to-video, and multimodal reference workflows with 480p and 720p output.

Kling 3.0 Motion Control
Motion Control

$0.045 / second

kling-3.0-motion-control

Kling 3.0 Motion Control is a reference-driven motion transfer model that combines one character image and one source video with transparent per-second pricing.

Happy Horse 1.1
Text to Video

$0.11 / second

happy-horse-1.1

Alibaba Happy Horse 1.1 video generation with text-to-video, image-to-video, and reference-to-video workflows.

Kling 3.0
Text to Video

$0.135 / second

kling-3.0/standard

Kuaishou's advanced Kling 3.0 video model family with Standard, Pro, and native 4K generation, multi-shot storyboarding, synchronized audio, and element-based consistency.

Kling O3
Text to Video

$0.05 / second

kling-o3/standard

Kling O3 video generation on PoYo with Standard, Pro, and 4K variants for text-to-video, image-to-video, and reference-to-video workflows.

Sora 2 Official
Text to Video

$0.20 / generation

sora-2-official

OpenAI Sora 2 on PoYo with synced audio, improved physics, optional reference image input, and fixed 4-second, 8-second, 12-second, 16-second, and 20-second tiers.

Wan 2.7 Video
Text to Video

$0.06 / second

wan2.7-text-to-video

Wan 2.7 Video models on PoYo for text-to-video, image-to-video, reference-to-video, and edit-video workflows.

Wan 3.0
Text to Video

$0.05 / second

wan3.0-text-to-video

Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

Wan 3.0 Video Prime
Text to Video

$0.068 / second

wan3.0-prime-text-to-video

Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

Wan Animate 2
Image to Video

Coming soon

Wan Animate 2

Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

MiniMax H3 Max
Text to Video

$0.05 / second

h3-max

MiniMax H3 Max generates 5–15 second videos from text, first/last frames, or image, video and audio references at 480p, 768p and 1080p.

MiniMax H3 Max Turbo
Text to Video

$0.025 / second

h3-max-turbo

MiniMax H3 Max Turbo generates 5–15 second videos from text or first/last frames at 480p, 768p and 1080p.

Grok Imagine
Text to Image

$0.03 / generation

grok-imagine-image

xAI's Aurora-powered visual AI for image generation and video creation with Fun, Normal, and Spicy creative modes.

Hailuo 02
Text to Video

$0.035 / second

hailuo-02

MiniMax's #2 globally-ranked video model with NCR architecture, ultra-realistic physics, and 1080p cinematic output.

Veo 3.1
Text to Video

$0.10 / generation

veo3.1-lite

Google DeepMind's VEO 3.1 video family with text-to-video, image-to-video, 720p, 1080p, and 4K output options.

Veo 3.1 Official
Text to Video

$0.018 / second

veo3.1-fast-official

Official VEO 3.1 video generation with text, image, first/last-frame and reference workflows across fast, lite and quality tiers.

Wan 2.2 Fast
Text to Video

$0.03 / generation

wan2.2-image-to-video-fast

Wan 2.2 Fast provides fast text-to-video and image-to-video generation with low-cost 480p and 720p tiers for quick iteration.

Grok Imagine Video 1.5
Image to Video

$0.072 / second

grok-imagine-video-1.5

Grok Imagine Video 1.5 API for text-to-video, single-frame image-to-video, and 1–7 image reference-to-video generation with up to 1080p output.

Hailuo 2.3
Text to Video

$0.175 / generation

hailuo-2.3

MiniMax's Hailuo 2.3 video model for realistic human motion, expressive characters, and text-to-video or first-frame guided generation at 768p and 1080p.

Omni Flash
Text to Video

$0.60 / generation

omni-flash

Omni Flash API access for text-to-video, image-to-video, three-image reference fusion, and video-input generation workflows.

Seedance 1.0 Pro
Text to Video

$0.105 / generation

seedance-1.0-pro

ByteDance's #1 ranked video model with multi-shot storytelling, cinema-grade motion, and bilingual text-to-video generation.

Seedance 1.5 Pro
Text to Video

$0.045 / generation

seedance-1.5-pro

ByteDance's latest video model with synchronized audio generation, flexible aspect ratios, and enhanced motion control.

Wan 2.5
Text to Video

$0.15 / generation

wan2.5-image-to-video

Wan 2.5 combines text-to-video and image-to-video generation with 5-second and 10-second output, synchronized audio support, and multiple size and resolution tiers.

Kling 1.6
Text to Video

$0.045 / second

kling-1.6/standard

Kling 1.6 Standard and Pro video generation on PoYo with text, first/last frame, and multi-image element reference workflows.

Kling 2.1
Image to Video

$0.15 / generation

kling-2.1/standard

Kling 2.1 on PoYo provides Standard and Pro image-to-video modes with 5-second and 10-second clips, start-frame control, and optional end-frame control in Pro.

Kling 2.6
Text to Video

$0.325 / generation

kling-2.6

Kuaishou's revolutionary video model that simultaneously generates visuals with synchronized dialogue, sound effects, and ambient audio in one pass.

Kling Avatar 2.0
Image to Video

$0.035 / second

kling-avatar-2.0/standard

Kling Avatar 2.0 creates audio-driven talking avatar videos from one reference image and one driving audio file, with Standard and Pro quality modes.

Wan 2.6
Text to Video

$0.40 / generation

wan2.6-text-to-video

Alibaba's Wan 2.6 video generation family for text-to-video, image-to-video, and video-to-video with multi-shot 1080p output.

Kling 2.5 Turbo Pro
Text to Video

$0.21 / generation

kling-2.5-turbo-pro

Kling 2.5 Turbo Pro is a flexible short-form video model with text-to-video, optional frame guidance, smooth motion, cinematic depth, and fixed 5-second and 10-second tiers.

Kling 3.0 Turbo
Text to Video

$0.085 / second

kling-3.0-turbo/standard

Kling 3.0 Turbo is the speed-optimized Kling 3.0 variant for high-volume text-to-video and image-to-video generation with multi-shot storyboarding. Standard (720p) and Pro (1080p) models available.

Wan Animate
Image to Video

$0.035 / second

wan-animate-move

Alibaba's 14B-parameter character animation model that transfers motion from reference videos to static characters with exceptional identity preservation.

Kling 2.6 Motion Control
Image to Video

$0.04 / second

kling-2.6-motion-control

Kuaishou's motion control model that transfers motion from reference videos to character images while maintaining identity and adapting environments.

Happy Horse
Text to Video

$0.08 / second

happy-horse

Alibaba Happy Horse 1.0 video generation and editing with text-to-video, image-to-video, reference-to-video, and video-edit workflows.

图像生成视频模型 API - 价格和模型适配

通过 PoYo.ai 调用 图像生成视频模型 API,查看透明的 API 价格、模型页面和统一的异步生成流程。

精确任务匹配

只有模型目录元数据包含 Image to Video 的模型才会出现在这里。

供应商对比

无需离开任务目录,即可查看不同供应商的模型系列。

API 就绪路径

每张模型卡都会跳转到 PoYo 的价格、示例、Playground 控件和 API 详情。

模型供应商任务类型价格
MiniMax H3 / Hailuo 03

hailuo-03

MiniMaxText to Video, Image to Video$0.105次生成
Seedance 2.5

seedance-2.5

SeedanceText to Video, Image to Video, Video to Video$0.085次生成
FLUX 3

flux-3/text-to-video

Black Forest LabsText to Video, Image to Video$0.17次生成
Seedance 2

seedance-2

SeedanceText to Video, Image to Video, Video to Video$0.045次生成
Gemini Omni 1.1 Flash

gemini-omni-1.1-flash

GoogleText to Video, Image to Video, Video to Video$0.225次生成
Seedance 2.0 Mini

seedance-2-mini

SeedanceText to Video, Image to Video, Video to Video$0.03次生成
Kling 3.0 Motion Control

kling-3.0-motion-control

KlingMotion Control, Image to Video$0.045次生成
Happy Horse 1.1

happy-horse-1.1

AlibabaText to Video, Image to Video, Uncensored$0.11次生成
Kling 3.0

kling-3.0/standard

KlingText to Video, Image to Video$0.135次生成
Kling O3

kling-o3/standard

KlingText to Video, Image to Video$0.05次生成
Sora 2 Official

sora-2-official

OpenAIText to Video, Image to Video$0.20次生成
Wan 2.7 Video

wan2.7-text-to-video

AlibabaText to Video, Image to Video, Uncensored$0.06次生成
Wan 3.0

wan3.0-text-to-video

AlibabaText to Video, Image to Video, Uncensored$0.05次生成
Wan 3.0 Video Prime

wan3.0-prime-text-to-video

AlibabaText to Video, Image to Video, Uncensored$0.068次生成
Wan Animate 2

wan-animate-2

AlibabaImage to Video$0.00次生成
MiniMax H3 Max

h3-max

MiniMaxText to Video, Image to Video$0.05次生成
MiniMax H3 Max Turbo

h3-max-turbo

MiniMaxText to Video, Image to Video$0.025次生成
Grok Imagine

grok-imagine-image

xAIText to Image, Image to Video, Uncensored$0.03次生成
Hailuo 02

hailuo-02

MiniMaxText to Video, Image to Video$0.035次生成
Veo 3.1

veo3.1-lite

GoogleText to Video, Image to Video$0.10次生成
Veo 3.1 Official

veo3.1-fast-official

GoogleText to Video, Image to Video$0.018次生成
Wan 2.2 Fast

wan2.2-image-to-video-fast

AlibabaText to Video, Image to Video$0.03次生成
Grok Imagine Video 1.5

grok-imagine-video-1.5

xAIImage to Video$0.072次生成
Hailuo 2.3

hailuo-2.3

MiniMaxText to Video, Image to Video$0.175次生成
Omni Flash

omni-flash

GoogleText to Video, Image to Video, Video to Video$0.60次生成
Seedance 1.0 Pro

seedance-1.0-pro

SeedanceText to Video, Image to Video$0.105次生成
Seedance 1.5 Pro

seedance-1.5-pro

SeedanceText to Video, Image to Video$0.045次生成
Wan 2.5

wan2.5-image-to-video

AlibabaText to Video, Image to Video$0.15次生成
Kling 1.6

kling-1.6/standard

KlingText to Video, Image to Video$0.045次生成
Kling 2.1

kling-2.1/standard

KlingImage to Video$0.15次生成
Kling 2.6

kling-2.6

KlingText to Video, Image to Video$0.325次生成
Kling Avatar 2.0

kling-avatar-2.0/standard

KlingImage to Video, Audio to Video, Talking Avatar$0.035次生成
Wan 2.6

wan2.6-text-to-video

AlibabaText to Video, Image to Video, Video to Video$0.40次生成
Kling 2.5 Turbo Pro

kling-2.5-turbo-pro

KlingText to Video, Image to Video$0.21次生成
Kling 3.0 Turbo

kling-3.0-turbo/standard

KlingText to Video, Image to Video$0.085次生成
Wan Animate

wan-animate-move

AlibabaImage to Video$0.035次生成
Kling 2.6 Motion Control

kling-2.6-motion-control

KlingImage to Video, Motion Control$0.04次生成
Happy Horse

happy-horse

AlibabaText to Video, Image to Video, Uncensored$0.08次生成

常见问题

PoYo.ai 上的图像生成视频页面包含什么?

+

这个页面汇总任务类型包含 Image to Video 的模型,覆盖首帧生成、参考图引导视频、产品图动画和角色动画等工作流。

图像生成视频模型更适合哪些场景?

+

它更适合已有视觉素材的场景,例如产品动效、角色动作、头像视频、海报动画、首帧镜头延展和需要保持主体一致的视频生成。

接入图像生成视频 API 前应该先看什么?

+

建议先确认模型支持的输入图片类型、视频时长、分辨率、运动控制、价格和异步任务流程,再用同一张参考图做效果对比。