Script free with an account·render in about a minute·no watermark on any plan
Recipe Video Generator
A recipe as five steps, one generated frame each, narrated — not filmed in your kitchen.
Start free — no credit card, and the script and the shot list cost no credits.
01Name the dish. Plus time and servings. Constraints — one pan, fifteen minutes — make a much better video.
02Ingredients, then four steps. Each step in the imperative with a real quantity, temperature or duration attached.
03Narrated with highlight captions. A warm read with word-by-word captions, which is how a quantity actually lands.
Voice:
Rachel
Captions:
Highlight
Music:
Uplifting
Format:
Vertical 9:16
What comes out
The settings this page ships with, and the shape they produce
Every example on this page is labelled with the exact configuration behind it, including the one that says it is a drawing.
Layout mock-up
A one-pan chicken traybake with potatoes and lemon, for fo…
Style: Documentary
Voice: Rachel
Captions: Highlight
Ratio: 9:16
A mock-up of the frame and caption layout at this page’s settings — TikTok, Reels and Shorts, uplifting music bed. It is a drawing, not a rendered video: we have no published render library to show you yet, and a stock clip dressed up as our output would be worth less than saying so.
Your input
“A one-pan chicken traybake with potatoes and lemon, for four people”
Plus this page’s brief, verbatim
Write this as a recipe video. Open by naming the dish, the time and the number of servings. Then the ingredients in one scene, and four steps in the imperative, each with a real quantity, temperature or duration. Close with how to tell it is done. No stories about the origin of the dish and no preamble.
That string is not marketing copy — it is the instruction the script writer is given before it reads a word of yours, and it is the only thing separating this page from the other forty-seven.
The shape that comes back
Scene 1 — the hook
One line that works with no context, because the viewer arrived mid-scroll.
Scenes 2–6 — the body
One idea each, one or two spoken sentences, each with its own image prompt and its own duration.
Final scene — the close
A line that lands the point rather than asking for a follow.
Per scene — the shot
The exact image prompt the model would receive, with this page’s style suffix already appended.
This is the structure, not a saved example. The box at the top of the page writes the real one — with your wording, in about ten seconds, free once you sign in.
5–8 scenes
each with its own line, its own frame and its own length
Under 60s
cut to the narration, not to a fixed per-scene timer
Word-timed
captions burned from the voice model’s own alignment
1080-class
vertical, square or landscape — no watermark on any plan
01
Read this first: the frames are generated, not filmed
No video or audio goes in — no clip of your hands, no recording of the kitchen — and a photograph of the finished dish can only be the opening frame, uploaded in the composer, not the scenes. Every scene is generated from the script, which means what this produces is an illustrated recipe rather than a cooking video.
If the point is your food, this is the wrong tool and any phone plus a free timeline editor is the right one. If the point is the recipe — the quantities, the order, the timing — this does that in about a minute and does it well.
02
Why quantities belong in the captions
The default caption style here is the word-by-word highlight, and that is a functional choice rather than an aesthetic one. A viewer cooking from a short video is pausing on the quantities, and highlighted captions put “180°C for 40 minutes” on screen in a form that survives a screenshot.
It also matters that the captions are burned from the voice model’s own word timings rather than estimated. A number that appears half a second after it is spoken is a number somebody writes down wrong.
03
The food-image problem, stated plainly
Image models are good at food and not good at process. A generated frame of a finished traybake looks excellent; a generated frame of “fold the mixture until it just comes together” looks like a plausible bowl of something that is not doing that.
So the useful pattern is: let the frames set the scene and let the narration and captions carry the instructions. And do not present a generated image of the finished dish as a photograph of the food you made — it is not, and in this niche the audience is unusually good at telling.
Who this is for
Four people this was built for
Recipe writers
A recipe you already wrote, as a video, without cooking it again on camera.
Meal-plan and budget accounts
Where the constraint is the content and the food is illustration.
Restaurants and food brands
A weekly recipe post without booking a food shoot for it.
Anyone with a family recipe
Written down properly, said out loud once, and kept.
Compared
What this replaces
Six things somebody has to arrange to make one short, and who arranges them.
Input
The usual way: Cook it twice, film it overhead, edit for an afternoon.
With Veedio: The recipe in a line. The steps, the timings and the frames come back.
Visuals
The usual way: Your own footage, if you have the kitchen light for it.
With Veedio: Generated documentary-style frames — good at food, poor at process, and honest about which.
Voiceover
The usual way: Book a VO artist, or record yourself and cut the breaths out.
With Veedio: Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
The usual way: Auto-captions that put the oven temperature on screen a beat late.
With Veedio: Word-timed highlight captions, so a quantity lands exactly when it is said.
Turnaround
The usual way: Days, and most of it waiting on someone else.
With Veedio: About a minute, and you watch it fill in scene by scene.
Changing it
The usual way: Re-record the line, re-cut the timeline, re-export, re-upload.
With Veedio: Edit the line and regenerate that one scene. Regenerating a scene is free.
The usual way
With Veedio
Input
Cook it twice, film it overhead, edit for an afternoon.
The recipe in a line. The steps, the timings and the frames come back.
Visuals
Your own footage, if you have the kitchen light for it.
Generated documentary-style frames — good at food, poor at process, and honest about which.
Voiceover
Book a VO artist, or record yourself and cut the breaths out.
Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
Auto-captions that put the oven temperature on screen a beat late.
Word-timed highlight captions, so a quantity lands exactly when it is said.
Turnaround
Days, and most of it waiting on someone else.
About a minute, and you watch it fill in scene by scene.
Changing it
Re-record the line, re-cut the timeline, re-export, re-upload.
Edit the line and regenerate that one scene. Regenerating a scene is free.
Related tools
Same engine, different preset
Twenty more pages over the same pipeline. Each one is a different brief, a different look and a different page — and each one prints its brief the way this one does.
No. Video and audio cannot be uploaded at all, and a photo can only be the opening frame. Every scene is generated, so this makes an illustrated recipe rather than a cooking video. If your own food is the point, use a normal editor.
Will the images show the actual dish?
They will show a generated image of that kind of dish. It is not a photograph of anything, and it should not be presented as one — food audiences notice.
Are the quantities reliable?
They come from a language model, so treat them the way you would treat a recipe from a stranger: plausible, occasionally wrong, and worth reading before you cook. The script is free to read.
Why the highlight captions?
Because people screenshot quantities. Word-by-word highlighting puts a temperature or a weight on screen in a form that survives a pause, and the timings come from the voice model rather than an estimate.
Can I write the recipe myself?
Yes, and for a recipe you actually use you should. Paste it into the box; the script writer keeps your quantities and builds the scenes around them.
Do I need an account?
Yes, a free one. The script and the shot list cost no credits once you are signed in; rendering spends from the 60 free credits, and signup needs no card.
Do I have to sign up to try this?
Yes, a free one. Signup takes a moment and needs no card. Once you are signed in, the box at the top writes the script and the shot list for no credits — that is the part that decides whether the video is worth making — and rendering it into a file uses the 60 free credits the account starts with, about four videos.
What does it cost after the free credits?
Starter is $19 a month, Pro $49, Studio $149. A credit is a unit of pipeline cost rather than a video, so a five-scene short costs less than a nine-scene one, and you see the estimate before you spend anything.
Do I own the videos, and is there a watermark?
You own them, on every plan including the free credits, and there is no watermark on any plan, no resolution cap and no platform withheld. Four things do follow the plan, because they are what a video costs to make: AI motion and sound effects need Starter, a presenter avatar and the choice of video model need Pro, and a free account tops out at a minute where paid plans go to three. Everything else — every voice, every language, every caption style, every aspect ratio — is the same on all of them.