LTX 2.5 Is Here: Native Multi-Shot Video, DFR, and Better Prompt Control

Aug 13, 2026

LTX 2.5 has arrived, and it is a much bigger update than a routine checkpoint refresh. Lightricks describes it as an open-weights world model for video generation, with changes aimed at the problems creators encounter after the first impressive frame: maintaining continuity across cuts, preserving detail through motion, following complex direction, and producing assets that fit professional post-production workflows.

The headline additions are native multi-shot generation, Diffusion Fidelity Rendering (DFR), a new diffusion video decoder, a fine-tuned Gemma 4 12B text encoder, automatic duration selection, native 4K HDR and RAW workflows, and a more capable distilled model.

Ready to create instead of configure? Generate a video with LTX 2.5 from a text prompt →. No local model download or GPU setup is required.

LTX 2.5 AI video generation

What Is LTX 2.5?

LTX 2.5 is the latest generation of Lightricks' joint audio-video foundation model. It can serve as the engine for text-to-video, image-to-video, audio-to-video, keyframe interpolation, video retakes, video transformation, HDR output, and fine-tuned production pipelines.

Unlike a closed web-only model, LTX 2.5 is distributed as downloadable weights. The official LTX 2.5 model page, LTX-2 GitHub repository, and Hugging Face model repository provide the model components and reference pipelines for teams that want to run or customize it on their own infrastructure.

For creators who do not want to manage a large local stack, this site provides browser-based LTX 2.5 workflows for text-to-video, image-to-video, and audio-to-video. Available generation controls vary by workflow, so model-level capabilities such as RAW/EXR export should not be confused with the options exposed by every online tool.

LTX 2.5 vs LTX 2.3: What Changed?

AreaLTX 2.3LTX 2.5
Shot structurePrimarily single-shot generationNative connected multi-shot scenes in one generation
Detail pipelineConventional multistage upscalingDFR with keyframes, spatial detailing, and optional temporal refinement
Video decodingConvolutional VAE workflowNew diffusion decoder for higher visual fidelity, plus a lighter convolutional option
Prompt understandingGemma 3-based text stackFine-tuned Gemma 4 12B encoder with a custom prompt enhancer
DurationCreator selects a supported lengthOptional duration head predicts an appropriate clip length from the prompt
Professional outputHigh-resolution audio-video generationNative 4K HDR and RAW/EXR-oriented pipelines
Fast workflowDistilled generation for rapid iterationUpdated distilled model trained for better quality and efficiency

The practical shift is from generating a visually attractive isolated clip to building a more controllable sequence. LTX 2.5 gives filmmakers and creative teams more ways to preserve intent between shots and carry the result deeper into an editing or finishing pipeline.

1. Native Multi-Shot Generation

Multi-shot is the most immediately visible LTX 2.5 feature. A single generation can contain connected cuts while preserving important scene properties such as character identity, environment, lighting, visual style, and voice.

That matters because assembling independently generated clips usually creates continuity errors: a jacket changes color, the room layout shifts, the key light moves, or a voice sounds different after the cut. Native multi-shot generation lets the model reason about the connected sequence instead of treating every shot as a separate request.

The best prompts describe a short sequence chronologically. For example:

A woman in a red raincoat waits under a neon shop sign at night. Wide establishing shot in heavy rain. Cut to a close-up as she hears a bicycle bell and turns left. Cut to an over-the-shoulder shot as a cyclist stops beside her. Preserve her face, red coat, blue-magenta lighting, rain intensity, and natural city ambience across every shot.

Keep the number of cuts appropriate for the requested duration. Two or three purposeful shots generally give each action more screen time than a dense edit list.

2. Diffusion Fidelity Rendering (DFR)

LTX 2.5 introduces Diffusion Fidelity Rendering, a rendering approach that assigns compute according to scene complexity. The official DFR pipeline generates keyframes, performs a spatial detailing pass, and can add temporal refinement rounds with a temporal upscaler.

LTX 2.5 Diffusion Fidelity Rendering

Why does that matter? Not every part of a video needs the same attention. A static wall is comparatively easy; a moving face, reflective material, fast camera move, or readable sign is much harder. DFR is designed to spend more rendering effort where visual errors are most noticeable, producing detail that holds together more convincingly through motion and refinement.

DFR is especially relevant to local and custom pipelines. The browser tools on this site focus on a streamlined generation experience rather than exposing every low-level DFR stage.

3. Cleaner Motion Through a New Video Decoder

The LTX 2.5 video VAE now offers a diffusion decoder designed for improved output quality. According to the official repository, it trades additional decode time and memory for higher fidelity, while a lighter convolutional decoder remains available for less demanding setups.

In practice, the new decoder targets the familiar failure modes that become obvious once a clip starts moving: unstable facial detail, crawling textures, malformed text, material flicker, and fine edges that dissolve between frames. The goal is not simply a sharper still image—it is detail that survives across time.

4. Stronger Prompt Understanding With Gemma 4

LTX 2.5 replaces the earlier text stack with a fine-tuned Gemma 4 12B encoder and a custom prompt enhancer. This improves the model's ability to translate concise creative direction into ordered actions, camera behavior, visual continuity, and audio cues.

You should still write prompts like a director, not like a bag of keywords. A reliable structure is:

  1. Start with the subject and primary action.
  2. Describe actions in chronological order.
  3. Specify framing and camera movement.
  4. Lock identity, wardrobe, environment, and lighting details.
  5. Add dialogue, ambience, sound effects, and music cues when relevant.
  6. State what must remain consistent across cuts.

For fast iteration, begin with the shortest prompt that captures the shot. Add constraints only when the output shows a specific ambiguity. This makes it easier to understand which instruction caused a change.

Try a prompt now: Open the LTX 2.5 text-to-video generator →, choose Fast for exploration or Pro when you want to prioritize final fidelity.

5. Auto Duration and More Efficient Iteration

An optional duration head can predict the clip length from the requested action, so a local pipeline does not always need a manually fixed frame count. A brief reaction can stay brief, while a more involved action can receive enough time to complete naturally.

LTX 2.5 also ships with an improved distilled model. The official reference workflow uses a small predefined set of diffusion steps for faster results, making it suitable for prompt exploration and high-volume iteration. Full and guided pipelines remain available when quality and control take priority over turnaround time.

The same fast-to-final mindset applies online: explore motion and composition with the Fast option where available, then move to Pro for the version you intend to keep.

6. Native 4K HDR, RAW, and EXR Workflows

For professional pipelines, LTX 2.5 adds native 4K HDR generation and support for workflows built around linear image data. Official local pipelines can accept EXR stills or frame folders, work in supported HDR color spaces, and output EXR frames plus an HDR master suitable for tonemapping and finishing.

These features are important for VFX and color workflows because they preserve a broader range of scene information than a delivery-ready SDR video. They also make LTX 2.5 more useful as one stage in a production pipeline rather than only as a final social-media clip generator.

Again, this is a model and local-pipeline capability. Individual online endpoints may expose a smaller set of resolution, format, frame-rate, or duration choices.

Three Ways to Create With LTX 2.5 Online

Text to Video

Use LTX 2.5 Text to Video when you are starting with an idea, script beat, product concept, or shot description. It is the quickest route for testing composition, camera language, atmosphere, and action.

Image to Video

Use LTX 2.5 Image to Video when art direction is already established. Upload a reference image, describe the motion and camera behavior, and optionally provide an end frame when you need more control over where the shot finishes.

Audio to Video

Use LTX 2.5 Audio to Video when timing begins with sound. Upload a supported audio clip, add a visual prompt, and optionally guide the result with an image. This workflow is useful for dialogue-led shots, performance concepts, music visuals, and sound-driven motion.

Should You Run LTX 2.5 Locally or Use an Online Tool?

Run locally when you need deep pipeline control, custom inference code, LoRA training, DFR tuning, EXR/HDR finishing, or private infrastructure. The official quick start notes that a representative set of full local components is roughly 66 GiB, before accounting for generated assets, caches, and additional checkpoints. GPU memory requirements also depend heavily on the selected pipeline, decoder, quantization, offloading, resolution, and duration.

Use an online tool when you want to move from prompt to result without downloading gated weights, installing CUDA dependencies, or tuning memory settings. It is also the faster way to validate whether a concept works before investing in a custom local setup.

Final Takeaway

LTX 2.5 is focused on the parts of AI video that matter after the first demo: continuity between shots, persistent detail, better interpretation of direction, efficient iteration, and production-oriented output. Native multi-shot generation and DFR are the defining upgrades, while Gemma 4 prompt understanding, auto duration, the new decoder, and the improved distilled workflow make the model more practical across both creative exploration and technical production.

Create your first LTX 2.5 video: Start from a text prompt →, animate an image →, or generate visuals from audio →.