Model icon
omni-flash
Text to Video
Image to Video
Video to Video
Model:
Omni Flash API supports text-to-video, image-to-video, three-image reference fusion, and video-input generation with 720p, 1080p, duration, and aspect-ratio controls.
Input
168/20000
Added 1/3

Drop files here, paste files, or add a URL. Use no image for text-to-video, one image for image-to-video, or three images for reference fusion.

Image Urls 1

Drop files here, paste files, or add a URL. video/mp4,video/quicktime,video/webm,video/x-matroska · 100 MiB

Output

Example Output

This is sample data. Generate a new result to see real output.

Task ID:OMNIFLASHPREVIEWCreated:2026-06-03T00:00:00
Status:finishedProgress:100%

Example video

Pricing details

Transparent pricing with no hidden fees. Pay as you go.

Googleomni-flash
720p/1080p - no video input - 4s
PoYo price
$0.600/gen
120 credits
Official price
-
You save
-
Googleomni-flash
720p/1080p - no video input - 6s
PoYo price
$0.750/gen
150 credits
Official price
-
You save
-
Googleomni-flash
720p/1080p - no video input - 8s
PoYo price
$1.00/gen
200 credits
Official price
-
You save
-
Googleomni-flash
720p/1080p - no video input - 10s
PoYo price
$1.10/gen
220 credits
Official price
-
You save
-
Googleomni-flash
4K - no video input - 4s
PoYo price
$1.25/gen
250 credits
Official price
-
You save
-
Googleomni-flash
4K - no video input - 6s
PoYo price
$1.50/gen
300 credits
Official price
-
You save
-
Googleomni-flash
4K - no video input - 8s
PoYo price
$1.75/gen
350 credits
Official price
-
You save
-
Googleomni-flash
4K - no video input - 10s
PoYo price
$2.25/gen
450 credits
Official price
-
You save
-
Googleomni-flash
720p/1080p / with video input
PoYo price
$1.50/gen
300 credits
Official price
-
You save
-
Googleomni-flash
4K / with video input
PoYo price
$2.00/gen
400 credits
Official price
-
You save
-

* Actual fees are based on the final output.

Introduction

Complete guide to using Omni Flash API for Multimodal Video Generation

Omni Flash API for Multimodal Video Generation

Omni Flash on PoYo accepts a prompt alone, a prompt with one source image, three reference images, or a short video reference. It is built for short video workflows where teams need practical controls for resolution, aspect ratio, and duration while keeping the public API request compact.

Key Features of Omni Flash API

The page examples use PoYo-hosted CDN media and are matched to the supported API modes: text-to-video, image-to-video, and video-input generation.

01

Generate motion from a single prompt

Use text-only generation when you want fast concept clips without preparing source media. Describe scene physics, camera movement, material behavior, and atmosphere in the prompt, then choose 720p or 1080p output.

  • Text-to-video with no image or video input
  • 4s, 6s, 8s, or 10s duration controls
  • Aspect ratios for widescreen, vertical, square, and portrait layouts

02

Animate one image while preserving the subject

For image-to-video, upload one source image and explain how it should move. This official Google DeepMind dandelion drawing example keeps the simple line art as the visual anchor while turning it into a short animated scene.

  • One image URL for image-to-video
  • Prompt-led motion and environment details
  • Useful for sketches, product frames, and visual concepts

03

Use a video reference for motion and style transfer

When a source video is provided, Omni Flash switches to a video-input workflow. Use it to transfer motion cues, timing, or scene structure while optionally adding reference imagery through the same request.

  • At most one video URL per request
  • Duration is omitted for video-input jobs
  • Supports 720p and 1080p resolution tiers

What Teams Build With Omni Flash API

01

Concept video drafts

Generate short clips from plain prompts for pitches, storyboards, and creative direction tests.

02

Image-to-video assets

Animate product photos, illustrations, and still frames into short social or campaign clips.

03

Reference image fusion

Submit three source images when a scene needs multiple visual references in one generation request.

04

Video-input variations

Use a video reference to guide motion while adjusting the final look and output resolution.

Omni Flash vs Veo 3.1 vs Seedance 2.0 - Video Model Comparison

FeatureOmni FlashVeo 3.1Seedance 2.0
ProviderGoogle Gemini OmniGoogle DeepMind VeoByteDance Seedance
Best ForConversational video editing, prompt-to-video, image-to-video, and reference-guided iterationCinematic generation, native audio, image-guided shots, and production previsualizationMulti-shot stories, native audio-video generation, lip-sync, and complex multimodal references
Official PositioningCreate and edit video from any input, with world understanding and conversational refinementLeading video generation model for filmmakers and storytellers, with realism, prompt adherence, and audioNative multimodal audio-video generation model for text, image, audio, and video inputs
Input Modes on PoYoPrompt only, 1 image, 3 images, or 1 video URLPrompt, image guidance, first/last frames, reference images, or video extension depending on tierText, image, video, and audio reference workflows through the Seedance 2 API page
Output and AudioVideo output on PoYo; official model card describes high-resolution video with audioVideo with native audio or silent output depending on model settingsNative audio-video joint generation with dialogue, SFX, music, and ambience
Resolution on PoYo720p, 1080p720p or 1080p depending on tier480p, 720p, 1080p depending on model
Duration on PoYo4s, 6s, 8s, 10s for no-video-input jobs; video input uses its own modeUsually 4s, 6s, or 8s generation; extension supports source-video continuation4-15 seconds, billed per second
Published Evaluation StatusGoogle says detailed T2VA, I2VA, R2VA, editing, and image generation evaluations will be shared when developer and enterprise APIs roll outGoogle publishes Veo 3.1 preference, alignment, visual quality, audio-video alignment, and physics claimsSeedance 2.0 model card reports broad improvements and leading-level performance in expert and public user tests
Choose This WhenYou need Google Omni Flash search coverage, conversational editing language, and compact PoYo API controlsYou need a mature cinematic model with native audio options and stronger public benchmark evidenceYou need longer multi-shot clips, audio/lip-sync workflows, and richer multimodal references

Model capabilities change quickly. The table combines public model information and PoYo's current API surface as of June 2026; verify current request fields and pricing in the PoYo docs before production integration.

How to Use Omni Flash API on PoYo

Step 1: Create your API key
Generate a PoYo API key and submit requests to the unified asynchronous generation endpoint.
Create API Key

Step 2: Choose an input mode
Send only prompt for text-to-video, add one image_urls item for image-to-video, send three images for reference fusion, or add one video_urls item for video-input generation.
Add Credits

Step 3: Poll status or use webhook
Use the returned task ID to query progress, or provide callback_url to receive the finished video asynchronously.
Open API Documentation

Omni Flash API FAQ

What is Omni Flash in Google Flow?

In Google Flow, Omni Flash refers to Gemini Omni Flash, Google's conversational video model for creating and editing video from different kinds of input. Google describes it as a model that can create anything from any input, starting with video, and says Flow users can blend real-world inspiration with generated content and iterate conversationally. Google Flow announcement

What is Omni Flash?

Omni Flash is commonly used to refer to Google's Gemini Omni Flash. The Google DeepMind model card describes it as a native multimodal model for text, image, audio, and video inputs that creates high-quality, high-resolution video with audio, supports conversational editing, and is designed around world understanding and multimodality. On PoYo, the Omni Flash API exposes prompt, image, and video-reference workflows with resolution, duration, and aspect-ratio controls. Google DeepMind model card

How much better is Omni Flash better than Veo 3.1?

There is no public, apples-to-apples score that proves Omni Flash is a fixed percentage better than Veo 3.1. Google says Omni Flash evaluation details for T2VA, I2VA, R2VA, video editing, and image generation will be shared when developer and enterprise APIs roll out. Veo 3.1 already has public Google benchmark claims for preference, prompt alignment, visual quality, audio-video alignment, and realistic physics. In practice, choose Omni Flash for conversational editing and any-input iteration; choose Veo 3.1 when you want a more established cinematic video model with published benchmark evidence. Veo 3.1 overview

How to use Google Omni Flash?

In Google products, use Gemini Omni Flash through Gemini, Google Flow, YouTube, or Flow Music where available to your plan and region. Start with a prompt, image, video, or other reference, then refine the result through follow-up instructions. On PoYo, create an API key, send model: "omni-flash" with a prompt and optional image_urls or video_urls, then poll the status endpoint or use a webhook to retrieve the video.

What input modes does Omni Flash support on PoYo?

PoYo supports text-to-video with no source media, image-to-video with one image, reference fusion with three images, and video-input generation with one video URL.

Can I upload two images?

No. The public API accepts 0, 1, or 3 images for this model. Requests with two images are rejected so the input mode stays unambiguous.

Can I send duration with video_urls?

No. When video_urls is provided, omit duration. Video-input jobs use their own fixed billing and generation mode.

Which resolutions and aspect ratios are available?

Resolution supports 720p and 1080p. Aspect ratio supports 16:9, 9:16, 1:1, 4:3, and 3:4.

Does PoYo store internal upstream payload fields?

No. PoYo stores only user-facing request fields such as prompt, image_urls, video_urls, resolution, duration, aspect_ratio, model, and callback_url when provided.

Why Use PoYo for Omni Flash API

01

Unified async API

Submit tasks through the same PoYo generation endpoint and retrieve final files with the standard task status flow.

02

Clear controls

Resolution, duration, aspect ratio, image count, and video input are explicit request fields.

03

Transparent credits

The model card and pricing table expose duration-based and video-input credit tiers before generation.

04

Clean stored input

Task input records keep user-facing fields only, which simplifies debugging and customer support.