One scene, three AI video prompts: LTX, MiniMax H3 and Kling

LTX prompting guideLightricks' own guide to writing prompts for LTX video models

Three neon consoles in orange, violet and cyan, each with a small screen, joined by glowing lines to one shared screen above them

Three AI video models now generate picture and sound together, cut between shots inside one clip, and put words in a character's mouth. Each publishes a guide to writing for it, and the three guides disagree about the form a scene should take. Lightricks says it plainly: a prompt written for another model loses its tag syntax and shot-list formatting on the way into LTX, and tends to underperform. What follows is one scene written three ways, from those guides.

The scene: a woman waits at a rainy bus stop at night; she checks her phone and says "He's late"; a bus pulls in.

LTX: write it as prose

Lightricks wants a single flowing paragraph in the present tense — roughly four to eight sentences for a single take — with dialogue in quotation marks and emotion shown through physical cues rather than named. For LTX-2.5's multi-shot scenes the cut goes into the prose too: name the transition, re-establish the new framing, say whether the sound carries across. It recommends two to four shots per generation.

A wide shot frames a bus stop at night, rain streaking through a streetlight's cone. A woman in a grey coat stands under the shelter, glancing at her phone; its light catches her face. She exhales and says quietly, "He's late." A hard cut moves to a close-up of the phone screen as rain ticks on the shelter roof; the ambience continues across the cut. Headlights sweep across the glass as a bus pulls in.

One more line in the guide changes how a writer paces a scene: LTX-2.5's duration predictor times the clip to the action as written, and will not add a pause you did not prompt for. If the beat matters, write "a beat of silence" into the scene.

MiniMax H3: a tagged script

MiniMax's prompt guide for H3 is the most like a screenplay, and the strictest. The prompt has three named fields: integrated_multimodal_description for everything seen and heard in order, overall_soundscape for ambience and physical sound, and non_diegetic_music for score the characters cannot hear. Shots are numbered; every shot after the first opens with its cut time; a speaker gets a stable ID and their line goes inside a dialogue tag with the language named.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a wide shot frames a woman in a grey coat under a bus shelter at night as rain falls past a streetlight. She glances at her phone, and the tired woman with a low, quiet voice (S1) says: <d>[English] He's late.</d> [Shot 2] At 00:04.000, the camera cuts to a close-up of the phone screen as headlights sweep across the shelter glass.

overall_soundscape: Steady rain drums on the shelter roof and runs off the gutter. A bus engine approaches and its brakes hiss as it stops. non_diegetic_music: N/A

The guide also asks for camera moves written as sentences with an amplitude and a speed — "the camera pushes in with small amplitude at slow speed" — rather than labels stacked at the end, and keeps on-screen text in double quotation marks exactly as it should appear.

Kling: a timeline and keyframes

Kling 4.0 takes up to ten keyframes and prompts of up to 8,000 tokens, and its keyframes guide lays a prompt out as a timeline: 0–5s:, 5–10s: and so on, each segment saying what happens. The keyframes fix the important moments; the prompt, in Kling's words, explains how the video moves between them. For this scene that is two keyframes — the woman at the stop, the bus arriving — and a timeline that carries her from one to the other.

0–4s: wide shot, a woman in a grey coat waits under a bus shelter at night in heavy rain; she checks her phone and says quietly, "He's late." 4–8s: headlights sweep across the shelter glass as a bus pulls in and stops beside her.

Which to write first

Write the scene once as prose — it is the form closest to how a writer thinks, and it is what LTX wants anyway. Converting prose into H3's fields is mechanical: split it at the cuts, give each speaker an ID, move the rain into the soundscape. Converting it into Kling's timeline means deciding which moments are keyframes, which is a directing decision rather than a formatting one.

One scene, three AI video prompts: LTX, MiniMax H3 and Kling: questions

Can I use the same prompt for LTX, MiniMax H3 and Kling?

Not unchanged. Lightricks' guide says prompts written for other models, Kling's included, lose their tag syntax and shot-list formatting in LTX and tend to underperform; MiniMax H3 expects its own three-field format, and Kling 4.0 works best with a timeline tied to keyframes.

How do I write dialogue in an AI video prompt?

For LTX, put the spoken line in quotation marks. For MiniMax H3, give the speaker an ID such as (S1) and put the line inside a dialogue tag with the language named: <d>[English] …</d>.

How many shots can one LTX-2.5 prompt hold?

Lightricks recommends two to four shots per generation, with each cut named in the prose.

Source: Lightricks, MiniMax and Kling prompting guides checked against the source

More AI video news