
AI APIs for developers
AI Video to Video API
Build reference-driven video workflows with the PoYo AI Video to Video API. Compare video-input models, transformation controls, output quality and pricing for creative applications.
Matching models
6 models

$0.085 /billed second
seedance-2.5
ByteDance Seedance 2.5 video generation with 4-30 second output, first and last frame control, up to 50 image, video, and audio references, optional synchronized audio, and 480p, 720p, or 1080p delivery.

$0.045 /output + reference video second
seedance-2
ByteDance Seedance 2 video generation with text-to-video, image-to-video, multimodal reference-to-video, video-to-video, and audio-to-video workflows using image, video, and audio references with optional native audio.

$0.225 /generation
gemini-omni-1.1-flash
Gemini Omni 1.1 Flash API for text-to-video, first and last frame interpolation, reference image generation, and reference-video workflows.

$0.03 /output + reference video second
seedance-2-mini
ByteDance Seedance 2.0 Mini video generation for lower-cost text-to-video, image-to-video, and multimodal reference workflows with 480p and 720p output.

$0.60 /generation
omni-flash
Omni Flash API access for text-to-video, image-to-video, three-image reference fusion, and video-input generation workflows.

$0.40 /video
wan2.6-text-to-video
Alibaba's Wan 2.6 video generation family for text-to-video, image-to-video, and video-to-video with multi-shot 1080p output.
Video to Video Model APIs - Pricing and Model Fit
Compare providers, supported inputs, output formats and prices before choosing a model.
Exact task match
Models appear here only when their catalog metadata includes Video to Video.
Provider comparison
Review model families across providers without leaving the task directory.
API-ready paths
Each model card links to PoYo pricing, examples, playground controls and API details.
| Model | Provider | Task types | Price |
|---|---|---|---|
| Seedance 2.5 seedance-2.5 | Seedance | Text to Video, Image to Video, Video to Video | $0.085/billed second |
| Seedance 2 seedance-2 | Seedance | Text to Video, Image to Video, Video to Video | $0.045/output + reference video second |
| Gemini Omni 1.1 Flash gemini-omni-1.1-flash | Text to Video, Image to Video, Video to Video | $0.225/generation | |
| Seedance 2.0 Mini seedance-2-mini | Seedance | Text to Video, Image to Video, Video to Video | $0.03/output + reference video second |
| Omni Flash omni-flash | Text to Video, Image to Video, Video to Video | $0.60/generation | |
| Wan 2.6 wan2.6-text-to-video | Alibaba | Text to Video, Image to Video, Video to Video | $0.40/video |
Frequently asked questions
What is an AI video to video API?
+
An AI video to video API accepts footage to guide a generated video or supported transformation. The source can influence motion, appearance or scene structure, depending on the endpoint. It is a generative workflow, so you should inspect the output rather than expect unchanged source frames.
Does Video to Video mean a full video editing API?
+
Not necessarily. Video references, generative edits, timeline rendering and video conversion are different capabilities. Check the chosen model for the exact operation and controls. If your task is trimming, joining clips or transcoding, use the appropriate media-processing step alongside generation.
What source videos can I submit?
+
Check the endpoint for accepted formats, duration, file size, resolution and reference count. Use a reachable URL where the API requires one. Choose a short clip containing the relevant action or composition rather than sending unrelated scenes that make the intended reference harder to follow.
Can I preserve the original motion while changing the appearance?
+
Some models can follow movement or composition from a source video, but preservation is not exact across every model or request. State what should remain and test a simple transformation first. If the primary task is transferring a performance to a character image, compare Motion Control endpoints.
Can a video to video API replace only one object?
+
Only assume localized editing when the selected endpoint documents the necessary controls. A prompt-based transformation may also alter other parts of the scene. Check for supported object references, masks or editing modes and evaluate surrounding detail before adopting the workflow.
Why does the generated video flicker or change between frames?
+
Ambiguous instructions, complex movement, occlusion and conflicting references can make temporal consistency harder. Try a shorter, clearer source clip and a narrower change. Compare output stability across models, and avoid simultaneously changing the subject, setting and motion when diagnosing an issue.
Does the output keep the source video audio?
+
Audio behavior depends on the model: it may generate new audio, accept audio as a reference or return a silent result. Do not assume automatic preservation. If the original soundtrack must remain exact, plan a separate audio extraction and remux step in your media workflow.
Is video upscaling the same as video to video generation?
+
Dedicated upscaling focuses on increasing resolution or enhancing existing footage. Generative video to video models can reinterpret content and change visual details. Use a dedicated video upscaler when the main requirement is enhancement, and use reference-driven generation when a creative transformation is intended.
How do I track a video to video API job?
+
Submit the chosen model and its video inputs to the documented PoYo generation endpoint. Save task_id and use status queries or a supported webhook to retrieve the completed result. Keep the source clip and settings with the job so you can compare results and avoid duplicate submissions.
What affects video to video API pricing?
+
The model and selected settings determine the billing unit. Some tiers charge output duration, while others also account for reference-video duration or quality options. Read the current pricing description carefully and estimate both input and output processing where applicable.
How do I choose between Video to Video and Motion Control?
+
Choose Video to Video when existing footage should guide a scene or supported transformation. Choose Motion Control when a reference performance should drive a separate character image. Compare required inputs and subject preservation with the same representative assets before deciding.
What should I do when a video input is rejected?
+
Check the returned error against the documented URL, format, duration and size requirements. Confirm the source can be fetched without your browser login. Correct the input before retrying, and query the original task when the problem was a network timeout rather than a confirmed rejection.