Guide · Text to video
From text to video: how to turn a prompt into a usable shot
Checked: September 7, 2026, using supplied research; product access and plan terms can change.
Choose the kind of video you need#
“Text to video” covers three different products. Generative short-clip tools such as Bing Video Creator and Runway create individual shots from visual prompts; choose these for an original scene or product motion. Script-to-finished-video tools such as Renderforest and Fliki build narrated videos with stock footage or assembled scenes; choose these for an explainer whose script carries the message. Talking-avatar tools such as Synthesia put a presenter on screen; choose these for a spoken introduction or training segment.
The workflow below focuses on generative shots. For narrated or avatar output, start with the spoken script and review scene timing, pronunciation, and captions instead.
Text-to-video works best when you ask for one clear shot at a time. A prompt can create a few seconds of a product reveal, an establishing shot, or a character action. Asking the same generation to produce a complete 30-second advertisement with several locations, edits, and spoken lines usually leads to missing actions and inconsistent details.
The practical approach is simple: plan the sequence, prompt each shot, keep the good takes, and assemble them in an editor.
Plan the shot before you type#
Write down what the viewer should see during the clip. Name one subject, one action, one setting, and one camera behavior. If the idea contains “then,” it may need another shot.
“A baker removes bread from an oven, then serves a customer outside” asks for a location change and a time jump. Split it into a close-up of the bread leaving the oven and a separate exterior handoff. Each prompt now has a visible goal that can fit into a short generation.
Choose the target format at this stage. A 16:9 frame suits most landscape players, 9:16 suits vertical short-form platforms, and 1:1 can work for square feeds. Composition words help: “centered with space above for a headline” or “subject in the left third with empty space on the right.” Leave titles and exact brand copy for editing because generated lettering is often unstable.
Use a concrete prompt formula#
A reliable starting structure is:
Shot type + subject + visible action + setting + light or time + camera movement + motion pace + constraints
Here is a landscape example:
Wide establishing shot of a small fishing boat crossing a calm lake at dawn, mist drifting above the water, soft cool light, slow camera pan from left to right, quiet natural movement, no text.
For a vertical product clip:
9:16 close-up of a red running shoe on a dark studio plinth, thin light sweeps across the sole, slow clockwise camera orbit, crisp controlled movement, product remains centered, no extra objects or lettering.
Use nouns and physical verbs. “The curtain moves gently toward the open window” gives the model a clearer task than “beautiful cinematic energy.” A few visual details are useful, but stacked style labels can conflict. Start plain, generate, and add one change at a time.
If people appear, describe actions that can be seen: “raises the cup,” “takes two steps,” or “turns slightly toward camera.” Dialogue, motivation, and backstory do not tell the system how pixels should move. Generate speech separately only when the chosen product explicitly supports it, then check lip movement closely.
Build a sequence shot by shot#
For a short promotion, make a compact shot list before using credits:
- A wide shot establishes the place.
- A medium shot introduces the person or product.
- A close-up shows the important action or detail.
- A clean end frame leaves room for text added in editing.
Repeat stable descriptors across prompts when continuity matters. If a character is “a middle-aged cyclist in a yellow rain jacket with a matte black helmet,” keep that phrase unchanged. Text-only generation may still alter the face, clothing, or props between shots. If the tool supports reference images, try one to guide appearance, but still inspect every shot for changes.
Generate short drafts first. Lower-cost or draft modes help compare camera direction and composition before spending higher-quality credits. Check the displayed queue status rather than assuming a fixed completion time. Do not submit duplicates until you know the first job has stopped.
Diagnose a weak result#
Too many actions disappear. Break the prompt into separate shots and put the essential action near the beginning.
The subject changes shape. Reduce action speed, shorten the duration, simplify clothing or props, and avoid full rotations that reveal unseen surfaces.
The camera behaves unpredictably. Ask for one move: locked camera, slow push-in, pan, tilt, tracking shot, or orbit. Do not combine a crane, zoom, orbit, and handheld shake in the same short clip.
The frame looks busy. Remove background characters, weather effects, reflections, or style terms until the main action works. Add complexity in later attempts.
Generated words are unreadable. Ask for a blank sign or packaging area, then overlay the exact words in an editor.
Motion flickers or jumps. Reduce rapid light changes and small repeating patterns. Try a slower action, simpler environment, or shorter clip.
The tool rejects the prompt. Read the displayed reason and the service’s content policy, then revise the part that caused the rejection. If the error is unclear, try a simpler, unambiguous description.
Export and assemble the final video#
Download successful shots promptly. Keep filenames tied to the shot list, such as 02-product-closeup-take3.mp4. Do not rely on a generation history without checking its retention rules.
In an editor, trim unstable opening and closing frames, arrange the shots, add text, record or license audio, and balance color between clips. Do not stretch a weak two-second moment into a long sequence. A clean cut to another angle usually looks better.
Choose MP4 if available and accepted by your destination, then open the downloaded file to confirm playback. Export at the aspect ratio used during generation. Reframing landscape footage into vertical can remove subjects or captions. Confirm the service’s available resolution before planning delivery; free tiers may restrict output quality, credits, storage, or watermark removal. A larger export does not cure inconsistent frames.
Protect rights and sensitive material#
Read both the generation terms and the rules for every reference asset. Commercial use may depend on the plan. Free output can carry a watermark or be limited to personal projects, and stock music, logos, celebrity likenesses, and reference images have separate rights.
Text prompts reveal less than personal uploads, but they can still contain confidential campaign names, product details, or client information. Avoid entering secrets in a public demo. If you add an image, video, or audio reference, check retention, deletion, and training controls first. No account does not automatically mean no storage.
Save the prompt, source licenses, export date, and a copy of the applicable terms for paid client work. Free plans change, so check the live limits and license again when the project moves from experimentation to publication.
Sources#
- Bing Video Creator launch announcement (June 2025) (external site)
- Runway Free plan details (external site)
- Runway usage rights (external site)
- Renderforest AI video generator (external site)
- Renderforest commercial-use help (external site)
- Fliki pricing (external site)
- Synthesia pricing (external site)
- Synthesia watermark help (external site)