Script free with an account·render in about a minute·no watermark on any plan
Voxel Video Generator
A blocky voxel world — textured cubes, low sun, saturated colour — built from a line of text.
Start free — no credit card, and the script and the shot list cost no credits.
01Name a process. How something forms, moves or spreads. Terrain and systems are what this look is for.
02A cube world per scene. Every line becomes a wide, high-angle voxel landscape in one consistent palette and one low sun.
03Bright read, lo-fi bed. An energetic voice over a calm track, with bold captions timed to the words.
Voice:
Bella
Captions:
Bold
Music:
Lo-fi
Format:
Vertical 9:16
What comes out
The settings this page ships with, and the shape they produce
Every example on this page is labelled with the exact configuration behind it, including the one that says it is a drawing.
Layout mock-up
How a river decides where to bend
Style: Voxel world
Voice: Bella
Captions: Bold
Ratio: 9:16
A mock-up of the frame and caption layout at this page’s settings — TikTok, Reels and Shorts, lo-fi music bed. It is a drawing, not a rendered video: we have no published render library to show you yet, and a stock clip dressed up as our output would be worth less than saying so.
Your input
“How a river decides where to bend”
Plus this page’s brief, verbatim
Write this for a voxel-styled short about how something works. Favour processes that happen to a landscape or a system over time — erosion, growth, routing, settlement — because each scene becomes a wide terrain shot and terrain is what this look renders best. One mechanism per scene, stated plainly, and no scene that depends on a close-up of a person’s face.
That string is not marketing copy — it is the instruction the script writer is given before it reads a word of yours, and it is the only thing separating this page from the other forty-seven.
The shape that comes back
Scene 1 — the hook
One line that works with no context, because the viewer arrived mid-scroll.
Scenes 2–6 — the body
One idea each, one or two spoken sentences, each with its own image prompt and its own duration.
Final scene — the close
A line that lands the point rather than asking for a follow.
Per scene — the shot
The exact image prompt the model would receive, with this page’s style suffix already appended.
This is the structure, not a saved example. The box at the top of the page writes the real one — with your wording, in about ten seconds, free once you sign in.
5–8 scenes
each with its own line, its own frame and its own length
Under 60s
cut to the narration, not to a fixed per-scene timer
Word-timed
captions burned from the voice model’s own alignment
1080-class
vertical, square or landscape — no watermark on any plan
01
The only look in the catalogue that renders a process as terrain
A voxel world is made of identical cubes, which means a hillside, a river bend and a road are all built from the same unit. That is unusually good for an explanation: when the frame shows a valley made of visible blocks, a viewer can see the quantity of a thing — how much was moved, how far it spread — in a way a photograph of the same valley hides.
So the brief points the script at processes rather than at events. Erosion, settlement, routing, growth, drainage: things that happen to a landscape over time, one mechanism per scene. That is the pairing this style has an actual advantage in, and everything else it does is decoration.
02
An audience that reads the style before it reads the words
The cube-world aesthetic carries an association for anyone under about thirty-five, and the association is not a brand — it is a way of looking at terrain, built up over a decade of watching worlds get dug out one block at a time. A viewer who recognises it grants the video a few seconds of attention that a stock landscape photograph would not get.
It is worth being deliberate about that. Used on a geography or an infrastructure explainer, the style is doing real work on retention. Used on a subject with no landscape in it at all, it is a costume, and the same audience is unusually quick to notice when a look has been put on for no reason.
03
What this is not: it is not gameplay, and it is not a named game
There is no screen recording anywhere in this pipeline. We do not capture a game, we do not clip one, and nothing here is footage of anybody playing anything — the frames are generated images in a blocky style, and describing them as gameplay would be a straightforward lie about where the pixels came from.
It also will not reproduce a specific game’s world, characters, interface or logo, and the preset will not be argued into it. What you get is the general voxel look: textured cubes, a low sun, soft shadows and saturated colour, applied to whatever you typed in the box above.
Who this is for
Four people this was built for
Geography and earth-science channels
Landscape processes shown as terrain rather than described over a photograph.
Infrastructure and city accounts
Roads, grids and growth, at a scale the viewer can see all at once.
Faceless explainer operators
A recognisable house look that needs no footage, no face and no licence.
Teachers and STEM creators
A frame where quantity is visible, which is most of what an explanation is trying to show.
Compared
What this replaces
Six things somebody has to arrange to make one short, and who arranges them.
Input
The usual way: Model the terrain, texture it, light it, and render an animation of it.
With Veedio: A sentence about a process. Every frame is already a cube world.
Visuals
The usual way: Record a game and hope nobody minds, or commission a voxel artist per shot.
With Veedio: A generated voxel landscape per scene, in one palette and one light, from that scene’s own line.
Voiceover
The usual way: Book a VO artist, or record yourself and cut the breaths out.
With Veedio: Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
The usual way: Auto-captions that drift, then an hour nudging keyframes.
With Veedio: Burned in from the voice model’s own word alignment, so they land on the syllable.
Turnaround
The usual way: Days, and most of it waiting on someone else.
With Veedio: About a minute, and you watch it fill in scene by scene.
Changing it
The usual way: Re-render the sequence and re-time everything downstream of it.
With Veedio: Rewrite the line, regenerate that one landscape, and nothing else in the video moves.
The usual way
With Veedio
Input
Model the terrain, texture it, light it, and render an animation of it.
A sentence about a process. Every frame is already a cube world.
Visuals
Record a game and hope nobody minds, or commission a voxel artist per shot.
A generated voxel landscape per scene, in one palette and one light, from that scene’s own line.
Voiceover
Book a VO artist, or record yourself and cut the breaths out.
Five ElevenLabs voices, re-recorded free whenever the line changes.
Captions
Auto-captions that drift, then an hour nudging keyframes.
Burned in from the voice model’s own word alignment, so they land on the syllable.
Turnaround
Days, and most of it waiting on someone else.
About a minute, and you watch it fill in scene by scene.
Changing it
Re-render the sequence and re-time everything downstream of it.
Rewrite the line, regenerate that one landscape, and nothing else in the video moves.
Related tools
Same engine, different preset
Twenty more pages over the same pipeline. Each one is a different brief, a different look and a different page — and each one prints its brief the way this one does.
No. Nothing in this pipeline records or clips a game. Every frame is a generated image in a blocky voxel style, which is why it can be about a river rather than about a level.
Can it look like a specific voxel game?
No, and it should not — that game’s world, characters and interface belong to somebody. What the preset gives you is the general cube-world look applied to your own subject.
Why is the camera always wide and high?
Because terrain is what this style renders best. Close-ups of faces come back as blocky approximations, so the brief keeps every scene at landscape scale, where the cubes are a feature rather than a limitation.
Can it show something moving, like water flowing?
Only as a sequence of states. Each scene is a still with a slow move on it, so a process is shown as before, during and after rather than as continuous motion. That is usually clearer anyway for an explanation.
Will the blocks spell out a label?
No. Generated text is invented text, and an invented label on a diagram is worse than no label at all. Keep the naming in the narration and in the caption layer.
Do I need to pay to see what it writes?
No. A free account with no card gets you the script, the scene split and the image prompts for zero credits. Only rendering spends, and the account starts with 60 free credits.
Do I have to sign up to try this?
Yes, a free one. Signup takes a moment and needs no card. Once you are signed in, the box at the top writes the script and the shot list for no credits — that is the part that decides whether the video is worth making — and rendering it into a file uses the 60 free credits the account starts with, about four videos.
What does it cost after the free credits?
Starter is $19 a month, Pro $49, Studio $149. A credit is a unit of pipeline cost rather than a video, so a five-scene short costs less than a nine-scene one, and you see the estimate before you spend anything.
Do I own the videos, and is there a watermark?
You own them, on every plan including the free credits, and there is no watermark on any plan, no resolution cap and no platform withheld. Four things do follow the plan, because they are what a video costs to make: AI motion and sound effects need Starter, a presenter avatar and the choice of video model need Pro, and a free account tops out at a minute where paid plans go to three. Everything else — every voice, every language, every caption style, every aspect ratio — is the same on all of them.