New #1 open-source AI video generator is here!
Quick Overview
Lightricks released the LTX-2 open-source audio-video generative model, which offers significantly improved performance and control over previous models, featuring native 4K 50 FPS generation, 20-second clips, and extensive LoRA support for style and camera motion control, making it a new standard for AI video generation.
Key Points: LTX-2 is a new open-source audio-video generative model released by Lightricks, representing a new standard for AI video generation. The model supports native 4K resolution at 50 frames per second (FPS) and can generate clips up to 20 seconds long. LTX-2 is natively multimodal, supporting text-to-video, image-to-video, and audio-conditioned generation. It features extensive LoRA support, including specialized Camera LoRAs for precise control over dolly, jib, and pan movements. The model is optimized for NVIDIA RTX GPUs, with the reviewer testing it successfully on an RTX 4090. The base model weights and accompanying LoRAs are available for local download and use via ComfyUI. Performance benchmarks show the distilled model generating a 5-second clip in 53 seconds on an RTX 4090, significantly faster than previous models.
Context: The video introduces LTX-2, the latest AI video generation model released by Lightricks, which they claim sets a new standard for the field. The presenter details the model's capabilities, including high frame rates, longer clip generation, and enhanced control mechanisms like LoRAs for camera motion. The presenter runs demonstrations using the ComfyUI interface, comparing the performance of the full model versus the distilled model on his local RTX 4090 workstation.
Detailed Analysis
The video announces the release of LTX-2, Lightricks' new open-source audio-video generative model, emphasizing its superior capabilities over previous iterations. The model achieves native 4K resolution at 50 FPS and supports up to 20-second video clips. It is natively multimodal, handling text-to-video, image-to-video, and audio conditioning. A key feature is the extensive support for LoRAs, including specialized Camera LoRAs (like Dolly-Left, Dolly-Right, Jib-Down) that allow precise control over camera movement, structure, and style, which is demonstrated through the ComfyUI node graph. The presenter confirms he is running the model locally on an NVIDIA RTX 4090, noting that the distilled version is much faster (53 seconds for a 5-second clip) than the full model (2 minutes 27 seconds). The video concludes by showing the model's ability to animate static images, such as Edvard Munch's 'The Scream,' and provides the link to the Hugging Face repository for users to download the weights and test the model themselves.