WORKFLOW
Text to Video AI
Text to video AI is prompt-to-video: you describe the scene, pick a compatible model in VIBE, and generate a clip on iPhone or Android.
Text to video AI is the shortest path from an idea to a moving shot. You write the camera, subject, setting, and mood. VIBE sends that prompt to the AI video model you selected.
Generate video from text
Example prompt: “Cinematic drone shot flying over a futuristic city at night, neon reflections, slow camera movement.” Put motion in the sentence. Vague nouns make vague clips.
Models that run from a prompt
Most models in VIBE support text-to-video, including Sora 2, Veo 3.1, WAN 3.0, and Seedance 2.5. The app shows what the selected model accepts.
Related
Already have a still? Use image to video AI. Want the category overview? Start at AI video generator.
How to write a prompt that actually moves
Name the camera (drone, handheld, slow push-in). Name the subject. Name the place. Name the weather or time of day. If you want sound, mention it. “City at night” is a postcard. “Cinematic drone shot flying over a futuristic city at night, neon reflections, slow camera movement” is a shot.
Reuse the same prompt on more than one model
The point of VIBE is that text to video AI is not locked to a single engine. Run the line on Veo 3.1, Sora 2, WAN 3.0 or Seedance 2.5 and keep the take that matches the post. See the model comparison hub and the long comparison article.
When not to use text to video
If you already have a product photo, a character sheet, or artwork you need to preserve, start with image to video AI instead. Mixing a new prompt with a new still and a new model at the same time teaches you nothing.