Script free with an account·render in about a minute·no watermark on any plan
Tutorial Video Generator
A how-to in five steps, each step its own scene, its own frame and its own length.
Start free — no credit card, and the script and the shot list cost no credits.
01Say what you are teaching. The outcome, and who it is for. A tutorial written for a beginner is a different script.
02Five steps, five scenes. Each step is one instruction with one concrete detail, and gets its own frame and duration.
03Narrated and subtitled. A warm, clear read with small bottom captions, so the picture stays visible while you follow.
Voice:
Rachel
Captions:
Subtle
Music:
Uplifting
Format:
Vertical 9:16
What comes out
The settings this page ships with, and the shape they produce
Every example on this page is labelled with the exact configuration behind it, including the one that says it is a drawing.
Layout mock-up
How to season a cast iron pan properly
Style: Documentary
Voice: Rachel
Captions: Subtle
Ratio: 9:16
A mock-up of the frame and caption layout at this page’s settings — TikTok, Reels and Shorts, uplifting music bed. It is a drawing, not a rendered video: we have no published render library to show you yet, and a stock clip dressed up as our output would be worth less than saying so.
Your input
“How to season a cast iron pan properly”
Plus this page’s brief, verbatim
Write this as a step-by-step tutorial. Open by naming the outcome and how long it takes, not by introducing yourself. Then five steps in order, one per scene, each an instruction in the imperative with one specific detail — a setting, a measurement, a duration. The last scene says how to tell it worked. No filler, no recap.
That string is not marketing copy — it is the instruction the script writer is given before it reads a word of yours, and it is the only thing separating this page from the other forty-seven.
The shape that comes back
Scene 1 — the hook
One line that works with no context, because the viewer arrived mid-scroll.
Scenes 2–6 — the body
One idea each, one or two spoken sentences, each with its own image prompt and its own duration.
Final scene — the close
A line that lands the point rather than asking for a follow.
Per scene — the shot
The exact image prompt the model would receive, with this page’s style suffix already appended.
This is the structure, not a saved example. The box at the top of the page writes the real one — with your wording, in about ten seconds, free once you sign in.
5–8 scenes
each with its own line, its own frame and its own length
Under 60s
cut to the narration, not to a fixed per-scene timer
Word-timed
captions burned from the voice model’s own alignment
1080-class
vertical, square or landscape — no watermark on any plan
01
A tutorial short is not a shortened tutorial
A ten-minute how-to can afford to explain why. A forty-five-second one cannot, and the videos that try are the ones people leave at step two. What holds a short-form tutorial is the outcome stated in the first line and then instructions with no throat-clearing between them.
So this brief opens on the result and the time it takes, then goes straight into the imperative: do this, then this. Every step is required to carry one specific thing — a temperature, a setting, a number of minutes — because a step with no specifics is the step people already knew.
02
Why the captions are the small ones
This is the one preset where the default caption style is subtle rather than bold. Big centre-screen captions are a retention device for entertainment, and on a tutorial they cover the middle of a frame that is supposed to be showing you the thing.
Small bottom captions keep the video watchable on mute — which is most of the feed — without hiding the picture. It is a dropdown, so change it if your audience is used to the louder style; but the default is set this way on purpose.
03
What the frames can and cannot show
Every frame is generated, which means the visuals illustrate the step rather than record it. For a physical process — cooking, cleaning, a repair — that reads well: a pan, a hand, a workbench, in a consistent documentary style.
For anything on a screen it does not. A generated image of a settings menu will be a convincing picture of a menu that does not exist, and the wrong thing to put behind an instruction. We cannot take a screen recording — nothing can be uploaded — so for software tutorials use this to write and voice the script, then put your own capture behind it in an editor you already have.
Who this is for
Four people this was built for
Teachers and trainers
The same process explained the same way every time, without recording it again.
Home and hobby channels
Physical processes illustrate well, and this is the format the whole category runs on.
Support and success teams
The five answers you type out weekly, as videos you can send instead.
Anyone with expertise and no camera
You know the steps. This turns them into something watchable in about a minute.
Compared
What this replaces
Six things somebody has to arrange to make one short, and who arranges them.
Input
The usual way: Script it, set up a camera over the workspace, film it, cut it down.
With Veedio: The outcome in one line. The steps come back split and timed.
Visuals
The usual way: Overhead footage you have to shoot, or stock that shows a different process.
With Veedio: A documentary-style frame per step, generated from that step’s instruction.
Voiceover
The usual way: Book a VO artist, or record yourself and cut the breaths out.
With Veedio: Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
The usual way: Auto-captions covering the middle of the frame you are trying to show.
With Veedio: Small bottom captions by default, word-timed, so the picture stays visible.
Turnaround
The usual way: Days, and most of it waiting on someone else.
With Veedio: About a minute, and you watch it fill in scene by scene.
Changing it
The usual way: Re-record the line, re-cut the timeline, re-export, re-upload.
With Veedio: Edit the line and regenerate that one scene. Regenerating a scene is free.
The usual way
With Veedio
Input
Script it, set up a camera over the workspace, film it, cut it down.
The outcome in one line. The steps come back split and timed.
Visuals
Overhead footage you have to shoot, or stock that shows a different process.
A documentary-style frame per step, generated from that step’s instruction.
Voiceover
Book a VO artist, or record yourself and cut the breaths out.
Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
Auto-captions covering the middle of the frame you are trying to show.
Small bottom captions by default, word-timed, so the picture stays visible.
Turnaround
Days, and most of it waiting on someone else.
About a minute, and you watch it fill in scene by scene.
Changing it
Re-record the line, re-cut the timeline, re-export, re-upload.
Edit the line and regenerate that one scene. Regenerating a scene is free.
Related tools
Same engine, different preset
Twenty more pages over the same pipeline. Each one is a different brief, a different look and a different page — and each one prints its brief the way this one does.
No. Nothing is captured and no video goes in — a screenshot could only be the frame the video opens on — so for a software tutorial the generated frames would be pictures of interfaces that do not exist. Use this for the script and the voiceover and put your own screen capture behind it.
How many steps will I get?
Five by default, which with a hook and a closing check comes in under a minute. Ask for fewer in the box if the process is genuinely three steps — padding a tutorial is worse than a short one.
Is the advice accurate?
It comes from a language model, so treat it the way you would treat a knowledgeable stranger: fine on settled processes, unreliable on anything specific to a version or a product released recently. Read the script before you render it — that is free.
Why are the captions small here?
Because big centred captions cover the part of the frame a tutorial is meant to be showing. It is a dropdown, so switch to bold if your channel uses that style.
Can I write the steps myself?
Yes, and for anything precise you should. Paste your steps into the box and the script writer keeps them, adds the hook and the close, and splits them into scenes.
Do I need an account?
Yes, a free one. The script and the shot list cost no credits; rendering is what spends them, and signup includes 60 with no card.
Do I have to sign up to try this?
Yes, a free one. Signup takes a moment and needs no card. Once you are signed in, the box at the top writes the script and the shot list for no credits — that is the part that decides whether the video is worth making — and rendering it into a file uses the 60 free credits the account starts with, about four videos.
What does it cost after the free credits?
Starter is $19 a month, Pro $49, Studio $149. A credit is a unit of pipeline cost rather than a video, so a five-scene short costs less than a nine-scene one, and you see the estimate before you spend anything.
Do I own the videos, and is there a watermark?
You own them, on every plan including the free credits, and there is no watermark on any plan, no resolution cap and no platform withheld. Four things do follow the plan, because they are what a video costs to make: AI motion and sound effects need Starter, a presenter avatar and the choice of video model need Pro, and a free account tops out at a minute where paid plans go to three. Everything else — every voice, every language, every caption style, every aspect ratio — is the same on all of them.