
Modelos de imagen a vídeo
Modelos de IA de imagen a vídeo
Anime tomas de producto, retratos, diseños de personajes o primeros fotogramas en vídeo para que el movimiento siga una referencia visual existente.
Modelos compatibles
Modelos 38

$0.105 / second
hailuo-03
MiniMax H3 (Hailuo 03) generates native 2K video from text, first/last frames, or image, video, and audio references.

$0.085 / second
seedance-2.5
ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

$0.17 / second
flux-3/text-to-video
Black Forest Labs' FLUX 3 video family for text, images, first and last frames, source-video extension, positioned keyframes, and native audio generation.

$0.045 / second
seedance-2
ByteDance Seedance 2 video generation with text-to-video, image-to-video, multimodal reference-to-video, video-to-video, and audio-to-video workflows using image, video, and audio references with optional native audio.

$0.225 / generation
gemini-omni-1.1-flash
Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

$0.03 / second
seedance-2-mini
ByteDance Seedance 2.0 Mini video generation for lower-cost text-to-video, image-to-video, and multimodal reference workflows with 480p and 720p output.

$0.045 / second
kling-3.0-motion-control
Kling 3.0 Motion Control is a reference-driven motion transfer model that combines one character image and one source video with transparent per-second pricing.

$0.11 / second
happy-horse-1.1
Alibaba Happy Horse 1.1 video generation with text-to-video, image-to-video, and reference-to-video workflows.

$0.135 / second
kling-3.0/standard
Kuaishou's advanced Kling 3.0 video model family with Standard, Pro, and native 4K generation, multi-shot storyboarding, synchronized audio, and element-based consistency.

$0.05 / second
kling-o3/standard
Kling O3 video generation on PoYo with Standard, Pro, and 4K variants for text-to-video, image-to-video, and reference-to-video workflows.

$0.20 / generation
sora-2-official
OpenAI Sora 2 on PoYo with synced audio, improved physics, optional reference image input, and fixed 4-second, 8-second, 12-second, 16-second, and 20-second tiers.

$0.06 / second
wan2.7-text-to-video
Wan 2.7 Video models on PoYo for text-to-video, image-to-video, reference-to-video, and edit-video workflows.

$0.05 / second
wan3.0-text-to-video
Generate videos up to 30 seconds with Wan 3.0 using text, first and last frames, or multimodal references including images, video, audio, documents, and links.

$0.068 / second
wan3.0-prime-text-to-video
Wan 3.0 Video Prime is Alibaba's Wan 3.0 tier optimized for faster turnaround, with multimodal image, video, audio, document, and public webpage references, videos up to 30 seconds, and optional generated audio.

Coming soon
Wan Animate 2
Alibaba's end-to-end character animation model for direct motion transfer, identity preservation, and text-driven viewpoint control. Coming soon to PoYo.

$0.05 / second
h3-max
MiniMax H3 Max generates 5–15 second videos from text, first/last frames, or image, video and audio references at 480p, 768p and 1080p.

$0.025 / second
h3-max-turbo
MiniMax H3 Max Turbo generates 5–15 second videos from text or first/last frames at 480p, 768p and 1080p.

$0.03 / generation
grok-imagine-image
xAI's Aurora-powered visual AI for image generation and video creation with Fun, Normal, and Spicy creative modes.

$0.035 / second
hailuo-02
MiniMax's #2 globally-ranked video model with NCR architecture, ultra-realistic physics, and 1080p cinematic output.

$0.10 / generation
veo3.1-lite
Google DeepMind's VEO 3.1 video family with text-to-video, image-to-video, 720p, 1080p, and 4K output options.

$0.018 / second
veo3.1-fast-official
Official VEO 3.1 video generation with text, image, first/last-frame and reference workflows across fast, lite and quality tiers.

$0.03 / generation
wan2.2-image-to-video-fast
Wan 2.2 Fast provides fast text-to-video and image-to-video generation with low-cost 480p and 720p tiers for quick iteration.

$0.072 / second
grok-imagine-video-1.5
Grok Imagine Video 1.5 API for text-to-video, single-frame image-to-video, and 1–7 image reference-to-video generation with up to 1080p output.

$0.175 / generation
hailuo-2.3
MiniMax's Hailuo 2.3 video model for realistic human motion, expressive characters, and text-to-video or first-frame guided generation at 768p and 1080p.

$0.60 / generation
omni-flash
Omni Flash API access for text-to-video, image-to-video, three-image reference fusion, and video-input generation workflows.

$0.105 / generation
seedance-1.0-pro
ByteDance's #1 ranked video model with multi-shot storytelling, cinema-grade motion, and bilingual text-to-video generation.

$0.045 / generation
seedance-1.5-pro
ByteDance's latest video model with synchronized audio generation, flexible aspect ratios, and enhanced motion control.

$0.15 / generation
wan2.5-image-to-video
Wan 2.5 combines text-to-video and image-to-video generation with 5-second and 10-second output, synchronized audio support, and multiple size and resolution tiers.

$0.045 / second
kling-1.6/standard
Kling 1.6 Standard and Pro video generation on PoYo with text, first/last frame, and multi-image element reference workflows.

$0.15 / generation
kling-2.1/standard
Kling 2.1 on PoYo provides Standard and Pro image-to-video modes with 5-second and 10-second clips, start-frame control, and optional end-frame control in Pro.

$0.325 / generation
kling-2.6
Kuaishou's revolutionary video model that simultaneously generates visuals with synchronized dialogue, sound effects, and ambient audio in one pass.
$0.035 / second
kling-avatar-2.0/standard
Kling Avatar 2.0 creates audio-driven talking avatar videos from one reference image and one driving audio file, with Standard and Pro quality modes.

$0.40 / generation
wan2.6-text-to-video
Alibaba's Wan 2.6 video generation family for text-to-video, image-to-video, and video-to-video with multi-shot 1080p output.

$0.21 / generation
kling-2.5-turbo-pro
Kling 2.5 Turbo Pro is a flexible short-form video model with text-to-video, optional frame guidance, smooth motion, cinematic depth, and fixed 5-second and 10-second tiers.

$0.085 / second
kling-3.0-turbo/standard
Kling 3.0 Turbo is the speed-optimized Kling 3.0 variant for high-volume text-to-video and image-to-video generation with multi-shot storyboarding. Standard (720p) and Pro (1080p) models available.

$0.035 / second
wan-animate-move
Alibaba's 14B-parameter character animation model that transfers motion from reference videos to static characters with exceptional identity preservation.

$0.04 / second
kling-2.6-motion-control
Kuaishou's motion control model that transfers motion from reference videos to character images while maintaining identity and adapting environments.

$0.08 / second
happy-horse
Alibaba Happy Horse 1.0 video generation and editing with text-to-video, image-to-video, reference-to-video, and video-edit workflows.
APIs de modelos de Imagen a vídeo: precios y ajuste del modelo
Llame a APIs de modelos de Imagen a vídeo a través de PoYo.ai para revisar precios transparentes, páginas de modelos y un flujo de generación asíncrono unificado.
Coincidencia exacta de tareas
Los modelos aparecen aquí solo cuando los metadatos de su catálogo incluyen Image to Video.
Comparación de proveedores
Revise las familias de modelos de todos los proveedores sin salir del directorio de tareas.
Rutas listas para API
Cada tarjeta de modelo enlaza con los precios de PoYo, ejemplos, controles del área de juegos y detalles de API.
| Modelo | Proveedor | Tipos de tareas | Precio |
|---|---|---|---|
| MiniMax H3 / Hailuo 03 hailuo-03 | MiniMax | Text to Video, Image to Video | $0.105generación |
| Seedance 2.5 seedance-2.5 | Seedance | Text to Video, Image to Video, Video to Video | $0.085generación |
| FLUX 3 flux-3/text-to-video | Black Forest Labs | Text to Video, Image to Video | $0.17generación |
| Seedance 2 seedance-2 | Seedance | Text to Video, Image to Video, Video to Video | $0.045generación |
| Gemini Omni 1.1 Flash gemini-omni-1.1-flash | Text to Video, Image to Video, Video to Video | $0.225generación | |
| Seedance 2.0 Mini seedance-2-mini | Seedance | Text to Video, Image to Video, Video to Video | $0.03generación |
| Kling 3.0 Motion Control kling-3.0-motion-control | Kling | Motion Control, Image to Video | $0.045generación |
| Happy Horse 1.1 happy-horse-1.1 | Alibaba | Text to Video, Image to Video, Uncensored | $0.11generación |
| Kling 3.0 kling-3.0/standard | Kling | Text to Video, Image to Video | $0.135generación |
| Kling O3 kling-o3/standard | Kling | Text to Video, Image to Video | $0.05generación |
| Sora 2 Official sora-2-official | OpenAI | Text to Video, Image to Video | $0.20generación |
| Wan 2.7 Video wan2.7-text-to-video | Alibaba | Text to Video, Image to Video, Uncensored | $0.06generación |
| Wan 3.0 wan3.0-text-to-video | Alibaba | Text to Video, Image to Video, Uncensored | $0.05generación |
| Wan 3.0 Video Prime wan3.0-prime-text-to-video | Alibaba | Text to Video, Image to Video, Uncensored | $0.068generación |
| Wan Animate 2 wan-animate-2 | Alibaba | Image to Video | $0.00generación |
| MiniMax H3 Max h3-max | MiniMax | Text to Video, Image to Video | $0.05generación |
| MiniMax H3 Max Turbo h3-max-turbo | MiniMax | Text to Video, Image to Video | $0.025generación |
| Grok Imagine grok-imagine-image | xAI | Text to Image, Image to Video, Uncensored | $0.03generación |
| Hailuo 02 hailuo-02 | MiniMax | Text to Video, Image to Video | $0.035generación |
| Veo 3.1 veo3.1-lite | Text to Video, Image to Video | $0.10generación | |
| Veo 3.1 Official veo3.1-fast-official | Text to Video, Image to Video | $0.018generación | |
| Wan 2.2 Fast wan2.2-image-to-video-fast | Alibaba | Text to Video, Image to Video | $0.03generación |
| Grok Imagine Video 1.5 grok-imagine-video-1.5 | xAI | Image to Video | $0.072generación |
| Hailuo 2.3 hailuo-2.3 | MiniMax | Text to Video, Image to Video | $0.175generación |
| Omni Flash omni-flash | Text to Video, Image to Video, Video to Video | $0.60generación | |
| Seedance 1.0 Pro seedance-1.0-pro | Seedance | Text to Video, Image to Video | $0.105generación |
| Seedance 1.5 Pro seedance-1.5-pro | Seedance | Text to Video, Image to Video | $0.045generación |
| Wan 2.5 wan2.5-image-to-video | Alibaba | Text to Video, Image to Video | $0.15generación |
| Kling 1.6 kling-1.6/standard | Kling | Text to Video, Image to Video | $0.045generación |
| Kling 2.1 kling-2.1/standard | Kling | Image to Video | $0.15generación |
| Kling 2.6 kling-2.6 | Kling | Text to Video, Image to Video | $0.325generación |
| Kling Avatar 2.0 kling-avatar-2.0/standard | Kling | Image to Video, Audio to Video, Talking Avatar | $0.035generación |
| Wan 2.6 wan2.6-text-to-video | Alibaba | Text to Video, Image to Video, Video to Video | $0.40generación |
| Kling 2.5 Turbo Pro kling-2.5-turbo-pro | Kling | Text to Video, Image to Video | $0.21generación |
| Kling 3.0 Turbo kling-3.0-turbo/standard | Kling | Text to Video, Image to Video | $0.085generación |
| Wan Animate wan-animate-move | Alibaba | Image to Video | $0.035generación |
| Kling 2.6 Motion Control kling-2.6-motion-control | Kling | Image to Video, Motion Control | $0.04generación |
| Happy Horse happy-horse | Alibaba | Text to Video, Image to Video, Uncensored | $0.08generación |
Preguntas frecuentes
¿Qué incluye la página de imagen a vídeo de PoYo.ai?
+
Esta página enumera los modelos cuyos metadatos de tareas incluyen imagen a vídeo, con generación desde primer fotograma, vídeo guiado por referencia, animación de productos y flujos de animación de personajes.
¿Para qué flujos de trabajo son mejores los modelos de imagen a vídeo?
+
Son mejores cuando ya tiene material visual de origen para animar productos, personajes, avatares, carteles, extensiones desde primer fotograma o vídeos con sujeto consistente.
¿Qué debo comprobar antes de integrar una API de imagen a vídeo?
+
Confirme las entradas de imagen admitidas, duración, resolución, controles de movimiento, precios y flujo de tareas asíncronas. Después compare las salidas con la misma imagen de referencia.
Explorar todos los modelos de IA
Abra el mercado de modelos completo para comparar modelos de imágenes, videos, audio, chat, 3D y herramientas.
Ver todos los modelosCrear con la API
Utilice la documentación de la API de PoYo cuando esté listo para enviar trabajos, consultar el estado o recibir callbacks.
Documentos API