Script free with an account·render in about a minute·no watermark on any plan
Text to Video Generator
Paste any block of text and get it back as a narrated, captioned vertical video with a frame per scene.
Start free — no credit card, and the script and the shot list cost no credits.
01Paste the text. A paragraph, a note, a slack message you wrote too well to waste. No formatting requirements.
02It becomes scenes. The text is rewritten for the ear and split so that each scene carries exactly one idea, with a frame to match it.
03Voiced, captioned, cut. Narration is generated, captions are timed to it, music sits underneath, and it renders as a single file.
Voice:
Adam
Captions:
Bold
Music:
Lo-fi
Format:
Vertical 9:16
What comes out
The settings this page ships with, and the shape they produce
Every example on this page is labelled with the exact configuration behind it, including the one that says it is a drawing.
Layout mock-up
We spent six months building a feature nobody used. Here i…
Style: Cinematic
Voice: Adam
Captions: Bold
Ratio: 9:16
A mock-up of the frame and caption layout at this page’s settings — TikTok, Reels and Shorts, lo-fi music bed. It is a drawing, not a rendered video: we have no published render library to show you yet, and a stock clip dressed up as our output would be worth less than saying so.
Your input
“We spent six months building a feature nobody used. Here is what we should have done instead: talked to twenty customers before writing a line of code, shipped the ugliest possible version in a week, and watched what people did with it rather than what they said about it.”
Plus this page’s brief, verbatim
Turn the text below into a spoken short. Keep the writer’s meaning and their strongest phrases, but rewrite for the ear: shorter sentences, one idea per scene, and the sharpest line moved to the front.
That string is not marketing copy — it is the instruction the script writer is given before it reads a word of yours, and it is the only thing separating this page from the other forty-seven.
The shape that comes back
Scene 1 — the hook
One line that works with no context, because the viewer arrived mid-scroll.
Scenes 2–6 — the body
One idea each, one or two spoken sentences, each with its own image prompt and its own duration.
Final scene — the close
A line that lands the point rather than asking for a follow.
Per scene — the shot
The exact image prompt the model would receive, with this page’s style suffix already appended.
This is the structure, not a saved example. The box at the top of the page writes the real one — with your wording, in about ten seconds, free once you sign in.
5–8 scenes
each with its own line, its own frame and its own length
Under 60s
cut to the narration, not to a fixed per-scene timer
Word-timed
captions burned from the voice model’s own alignment
1080-class
vertical, square or landscape — no watermark on any plan
01
Writing for the eye and writing for the ear are different jobs
Text that reads beautifully often sounds terrible. Subordinate clauses that a reader can re-scan are lost on a listener, who gets one pass at the sentence and no way back. Semicolons vanish. Parentheses become confusion.
So this does not simply read your paragraph aloud. It rewrites it for speech — shorter sentences, one idea per scene, the strongest phrase moved to the front where a scroll-stopper has to live — while keeping your meaning and, wherever they survive the move, your actual words.
02
One idea per scene, because the picture has to agree with the line
Every scene gets a generated frame built from that scene’s own line. That only works if the line is about one thing. A sentence that covers a statistic, a metaphor and a caveat has no single image, and the frame behind it ends up generic.
The split is therefore semantic rather than mechanical. It is not chopping your text every fifteen words; it is finding where the ideas change and cutting there, which is also where a viewer expects a cut.
03
What comes back, and what you can change
You get a title, five to eight scenes, the spoken line for each, the image prompt behind each, and an estimated runtime — the whole plan of the video, before a single asset has been generated.
In the app, everything on that plan is editable. Change a line and regenerate that scene; change the image prompt and leave the narration alone; swap the voice and keep the script. Nothing forces a full re-render, and a scene regeneration is free.
Who this is for
Four people this was built for
Newsletter writers
One section of this week’s issue, posted as a short, is the cheapest audience growth available to you.
Support and docs teams
The answer you have already written twenty times becomes a 40-second video you can link instead.
Course creators
A module’s key idea, cut down to the one claim that makes people want the module.
Anyone with a notes app full of good sentences
The writing is done. This is the part where it becomes something people actually see.
Compared
What this replaces
Six things somebody has to arrange to make one short, and who arranges them.
Input
The usual way: Rewrite the text as a script, then storyboard it, then find the footage.
With Veedio: Paste the text. The rewrite, the split and the shot list happen in one pass.
Visuals
The usual way: Search stock for a clip that vaguely matches paragraph three.
With Veedio: A frame generated from the line it sits behind, so it never vaguely matches.
Voiceover
The usual way: Book a VO artist, or record yourself and cut the breaths out.
With Veedio: Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
The usual way: Burn in auto-captions, then fix the six words it heard wrong.
With Veedio: Captions come from the script you approved, timed to the voice that read it.
Turnaround
The usual way: Days, and most of it waiting on someone else.
With Veedio: About a minute, and you watch it fill in scene by scene.
Changing it
The usual way: Re-record the line, re-cut the timeline, re-export, re-upload.
With Veedio: Edit the line and regenerate that one scene. Regenerating a scene is free.
The usual way
With Veedio
Input
Rewrite the text as a script, then storyboard it, then find the footage.
Paste the text. The rewrite, the split and the shot list happen in one pass.
Visuals
Search stock for a clip that vaguely matches paragraph three.
A frame generated from the line it sits behind, so it never vaguely matches.
Voiceover
Book a VO artist, or record yourself and cut the breaths out.
Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
Burn in auto-captions, then fix the six words it heard wrong.
Captions come from the script you approved, timed to the voice that read it.
Turnaround
Days, and most of it waiting on someone else.
About a minute, and you watch it fill in scene by scene.
Changing it
Re-record the line, re-cut the timeline, re-export, re-upload.
Edit the line and regenerate that one scene. Regenerating a scene is free.
Related tools
Same engine, different preset
Twenty more pages over the same pipeline. Each one is a different brief, a different look and a different page — and each one prints its brief the way this one does.
Roughly 1,500 characters into the free box on this page, and 5,000 in the app. Beyond that you are describing a video longer than a short, and the scriptwriter will cut it down rather than run past a minute.
Will it keep my exact words?
Some of them. It keeps your meaning and your best phrasing but rewrites for the ear. If you want your wording preserved as closely as possible, use the script-to-video tool instead — that preset is explicitly told to keep the original.
Can I paste something with headings and bullet points?
Yes. Structure is a help rather than a hindrance: headings usually mark exactly where the scene breaks belong, and bullets often become one scene each.
What if the text is about something the model does not know?
Then keep it descriptive rather than referential. The scriptwriter works from what you paste, so facts, numbers and names in the text will survive; a passing reference to something you have not explained may not.
Does it add facts I did not write?
It should not, and the prompt tells it not to, but it is a language model and it can still reach for a connective claim. Read the script before you render — that is exactly why the script comes back first.
Can I choose how many scenes it makes?
Not directly. Scene count follows the length you choose — five to eight scenes in a minute, roughly one per seven or eight seconds of speech — so picking 90 seconds or 3 minutes in the composer is how you get more of them. Pasting less text makes a shorter video.
What voice suits pasted text?
This page defaults to Adam, a deeper documentary read, because pasted prose tends to be explanatory. All five voices are one dropdown away, and switching voices later re-records the narration only.
Does the music get in the way of the words?
No. The bed is ducked underneath the narration and the whole master is loudness-normalised, so it sits at the level everything else in the feed does. You can also set music to none.
Do I have to sign up to try this?
Yes, a free one. Signup takes a moment and needs no card. Once you are signed in, the box at the top writes the script and the shot list for no credits — that is the part that decides whether the video is worth making — and rendering it into a file uses the 60 free credits the account starts with, about four videos.
What does it cost after the free credits?
Starter is $19 a month, Pro $49, Studio $149. A credit is a unit of pipeline cost rather than a video, so a five-scene short costs less than a nine-scene one, and you see the estimate before you spend anything.
Do I own the videos, and is there a watermark?
You own them, on every plan including the free credits, and there is no watermark on any plan, no resolution cap and no platform withheld. Four things do follow the plan, because they are what a video costs to make: AI motion and sound effects need Starter, a presenter avatar and the choice of video model need Pro, and a free account tops out at a minute where paid plans go to three. Everything else — every voice, every language, every caption style, every aspect ratio — is the same on all of them.