Audio + image → video

Create with MiniMax H3, Gemini Omni, Veo 3.1, Seedance 2.0, or Seedance 2 Mini in Flow AI Video. Start from text or supported reference media, track every generation, and keep finished videos ready to download or reuse.
Total credits = video rate x charged time + additional image credits. Input audio is free.
Made with Flow AI Video
Explore cinematic scenes, advertising, interfaces, dynamic posters, image-to-video animation, fine detail, and synchronized audio across the creative workflows available in Flow AI Video.
From prompt to video
Combine prompts with images, video, or audio when the selected model supports them. Flow AI Video shows the reference modes, durations, resolutions, and creative controls available for each model.



Why Flow AI Video
Flow AI Video keeps model selection, prompting, reference uploads, generation tracking, and asset reuse in one workspace. Each model keeps its own supported inputs and settings clear.
In Flow AI Video, compare supported inputs, durations, resolutions, aspect ratios, and audio options before you create.
Write a prompt, then add images, video, audio, or first and last frames when the selected model supports them.
Describe the subject, action, camera, style, pacing, and reference relationships in one prompt instead of splitting direction across separate tools.
See the selected model, output settings, and required credits before starting a generation.
Follow processing tasks from the queue while you continue working elsewhere in the site.
Successful videos are saved to the Library for review, download, and reuse as references. Failed generations automatically return the credits charged for that attempt.
How it works
Choose a model, describe the result, add supported references, and review the output settings and credit cost. Flow AI Video tracks the task and saves successful results to your Library.
Pick the model in Flow AI Video that matches your workflow, then describe the subject, action, camera movement, style, atmosphere, and sound you want.

Upload supported images, video, or audio to guide the subject, motion, timing, or sound. You can also choose compatible media already saved in your Library.

Review the duration, aspect ratio, resolution, audio option, and credit cost available for the selected model. Follow progress in the queue, then find successful results in your Library.

Built for production work
Use Flow AI Video models and reference workflows for campaign production, product storytelling, previsualization, motion design, and fast creative iteration.
Explore a hero film, paid-social variation, or localized cutdown from the same set of product, motion, voice, and brand references.
Use typography, product color, camera language, voice, and sound references to direct a short branded sequence.
Turn packshots and product references into demonstrations, macro reveals, seasonal scenes, and vertical storefront videos.
Use campaign art as the design reference, then describe the movement, pacing, sound, and text details the animated version should follow.
Explore lenses, blocking, camera movement, transitions, and title sequences before crews, locations, and post-production are booked.
Develop cinematic tests, animated UI, environmental motion, and character moments from one world-building reference set.
Show a proposed object in context, animate its behavior, compare finishes, and communicate an interaction before the physical prototype is ready.
Move from a static interface frame to a presentable interaction concept with transitions, sound, and product storytelling included.
Model selection guide
Each Flow AI Video model supports a different mix of reference media, duration, resolution, aspect ratios, and output controls. Start with the requirements that matter most to your brief.
| Criterion | MiniMax H3 | Gemini Omni | Google Veo 3.1 | Seedance 2.0 |
|---|---|---|---|---|
| Generation inputs and references | Prompt, up to nine images, three videos, and three audio files | Prompt with image references or one source video | Prompt with one to three reference images, or first and last frames | Prompt, up to nine images and three videos, or first and last frames |
| Generation workflows | Text-to-video, reference-led generation, first/last frame, and video editing | Text-to-video, image-guided generation, and video editing | Text-to-video, reference images, and first/last frame transitions | Text-to-video, reference-led generation, first/last frame, and video editing |
| Available duration | 4 to 15 seconds | 4, 6, 8, or 10 seconds | 4, 6, or 8 seconds | 4, 5, 6, 8, 10, or 15 seconds |
| Available resolution | 768P or 2K | 720p, 1080p, or 4K | 720p, 1080p, or 4K | 480p, 720p, or 1080p |
| Available aspect ratios | Six formats from 21:9 landscape to 9:16 portrait | 16:9 or 9:16 | 16:9 or 9:16 | Seven fixed formats plus adaptive framing |
| Distinctive controls | Image, video, and audio references with precise media mentions in the prompt | One source video up to two minutes for video-guided workflows | Dedicated reference-image and first/last-frame modes | Optional generated audio and adaptive aspect ratio |
How creators use the workspace
Teams can match a model to each deliverable while keeping prompts, references, costs, task status, and completed work in one consistent production flow.
Some briefs need a fast proof of concept; others need careful control over motion and detail. I can switch models without changing how the team reviews the work.
Keeping the prompt, reference frames, and source footage together saves me from hunting through old project folders. When a direction works, I can pick it up again without rebuilding the setup.
The cost preview makes planning a batch of campaign variants much more predictable. We can explore broadly at the draft stage, then spend more only on the concepts that earn a final pass.
I usually start with a loose motion study, compare a few model outputs, and carry the strongest one into the next round. The workspace makes that experimentation feel like one project instead of scattered tests.
Our producers can see what is running, what is ready, and which result belongs to each brief. That visibility has made handoffs between the creative and campaign teams noticeably calmer.
A launch film and its vertical cutdowns rarely need the exact same treatment. I can choose the right model and format for each deliverable, then keep every approved version easy to find.
Learn how model selection, reference media, credits, usage rights, and the Library fit into the workflow.

Start from text or supported reference media, review the settings and credit cost, then follow the result from the generation queue to your Library.