Prompt builder

A reference-to-video prompt is not a sentence you write, it is a structure you fill. Set the frame, say what each reference is for, then spend the seconds deliberately — the two parts people skip are the one that stops the product drifting and the one that decides how long the payoff gets. Fill it in, or load the worked example and take it apart.

0 of 6 parts · 0 beats

Your references

Numbered per media type

Declare what you are attaching before you write, because the prompt addresses them by number and the numbering runs separately for each type. The first clip is [Video1] even if three stills came before it.

No references yet. A prompt with none is still valid — it just has nothing holding the product still.

01Register — what kind of film is this

The first clause sets the whole grade. "Bright and colourful commercial style" and "moody cinematic short" produce different lighting, different lenses and different cutting from identical instructions further down. Say it first and say it plainly; a model that has to infer the register infers it late, after it has already committed to a look.

02Subject and variants — what is the star

03Reference bindings — what each reference is FOR

04Arrangement — how the frame is organised

05Timeline — how the seconds are spent

beats total 0s

06On-screen text — exact words, exact positionoptional

07Music bed — or an explicit refusaloptional

The prompt

0 characters · 0 references
Fill a part above, or load the worked example.

What the form doesn't tell you

Bind references inline, not in a list at the end

A reference mentioned where it applies governs that clause. The same reference listed in a block at the bottom governs everything and therefore nothing — the model averages it into the general look instead of applying it to the shot you meant.

Tokens are numbered per media type, not overall

The first still is [Image1] and the first clip is [Video1] even if three images were uploaded before it. Get this wrong and the prompt points at a file that is not there, which fails silently — you get a plausible video built on the wrong reference.

Quote the words you want rendered

Text inside quotation marks is treated as literal on-screen copy. "Describe the tagline" gets you invented letterforms; 'the tagline reads "Simple. Done right."' gets you the tagline.

Physics beats adjectives

"Crumbs scatter and the filling bursts open" is executable. "Looks delicious" is not. Describe what moves, what breaks and what the speed does, and the model has something to simulate.

State the register first and the sound last

The opening clause sets the grade for everything after it; the audio cue reads best as a closing instruction because it applies across the whole take rather than to any one beat.

One take, one climax

Two payoff moments in a 15-second cut means neither lands. If a second idea is worth having, it is worth its own render — which at 480p costs about a fifth of a finished one.

@Image1 in the playground, [Image1] through the API

The Dreamina playground writes references as @Image1 because it resolves them from an attachment picker as you type. Through the API — which is what this studio uses — they are square-bracketed, [Image1]. The structure is identical; only the sigil changes. Paste a playground prompt straight into an API call with the @ intact and the tokens are read as literal text.