Text-to-Video or Image-to-Video?
Choose the right starting point for exploration, consistency, and motion control.
OmniVideo supports two starting points: a written prompt, or a prompt paired with a reference image. The best option depends on what you already know about the shot.
Start with text when you want to explore
Text-to-video is useful when the visual direction is still open. It lets you test a subject, setting, mood, and camera idea without preparing artwork first.
Use it for:
- early concepts and visual brainstorming
- scenes where exact character or product continuity is not essential
- quickly comparing landscape and portrait directions
Add an image when the frame already matters
Image-to-video is a better fit when you already have a composition, product shot, illustration, or approved key visual. The reference supplies the visual starting point; the prompt can concentrate on motion.

Use it for:
- animating a product or campaign still
- keeping a recognizable composition
- testing several motion directions from the same frame
A practical workflow
Begin with text-to-video to discover the broad direction. Once you find a composition you like, use an approved still as a reference and focus your next prompt on camera movement, environmental motion, and pacing.
Generated results can vary, particularly while the model is in preview. Keep source files and prompt versions so you can recreate the intent even when a specific output cannot be reproduced exactly.