ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

LTX-2.5 开源世界模型完全指南:模型家族、下载安装与图像转视频实战

LTX-2.5 开源世界模型完全指南:模型家族、下载安装与图像转视频实战 大模型媒体生成音视频多模态本地部署【免费下载链接】LTX-2.5项目地址https://ai.gitcode.com/hf_mirrors/Lightricks/LTX-2.5点击查看免费下载LTX-2.5 是 Lightricks 推出的开源权重世界模型定位本地执行与微调以文本、图像与视频输入生成同步的高保真视频与音频。本文以仓库 README.md 为骨架结合仓库内全部实际发布的模型文件系统梳理 LTX-2.5 的新特性、完整模型家族清单、ltx-pipelines与 Diffusers 两种主流调用方式的实战配置并给出帧数/分辨率约束、低显存技巧与微调适配建议帮助你直接在本机跑通文生视频与图生视频管线。模型概览LTX-2.5 是什么根据 README.md 的定义LTX-2.5 是一个权重开放的开放式世界模型open world model with open weights其设计初衷就是本地执行local execution与微调fine-tuning。它的既定用途是从文本、图像和视频输入生成同步synchronized、高保真high-fidelity的视频与音频而对于机器人、物理 AI 等新兴领域的适用性仍在持续探索之中属于发展中的能力而非已承诺的卖点。从生态定位看LTX-2.5 强调完全控制与定制——在你的自有基础设施上自托管self-host on your own infrastructure没有按次计费的 API 账单、没有按席位锁定、没有强制的 API 依赖。这一原则也直接决定了本仓库的交付形态——一份按组件拆分的 .safetensors 文件集合而不是一个单一巨型权重文件。值得注意的授权边界详见 README 中的商业条款摘要营收按整个实体包括共同控制下的子公司与关联公司合并计算年营收低于 1000 万美元Under $10M annual revenue可在 LTX-2.x Community License 下免费进行商业与生产使用但微调模型的转移transfer of fine-tunes可能需要付费许可年营收超过 1000 万美元Over $10M annual revenue则需签署付费商业使用协议Commercial Use Agreement可获得完整权重、工程支持、LoRA 与灵活部署选项。LTX-2.5 的新特性Whats newREADME 明确列出了六项核心升级理解它们是读懂后续模型文件划分的前提原生多镜头生成Native multishot generation——单次生成即可产出多个相连镜头multiple shots在镜头切换间保持角色身份、环境、光照、声音与视觉风格的一致性。这是相对前代的重大跃迁早期版本只能产出单个连续镜头a single continuous shot。扩散保真渲染Diffusion fidelity rendering——不再把每个场景锁定到单一压缩率而是根据场景复杂度与预算动态分配算力在关键处渲染无瑕细节在其余部分保持高效。全新的扩散视频解码器New diffusion video decoder——取代了原先的 VAE 重建阶段带来更锐利的面部、纹理与屏幕文字更好的运动表现以及高要求场景下更少的伪影。定制的 Gemma 4 12B 文本编码器Custom Gemma 4 12B text encoder——能够hold together复杂提示词多角色、运镜、光照、动作不再在长提示序列中丢失细节。这解释了为何仓库中文本编码器单文件就达到 12B 量级。提示词增强器Prompt enhancer——以极低额外算力将简短提示扩展为更丰富的电影化指令。时长预测器Duration predictor可选——一个可选节点根据提示词预测片段长度并自动设置帧数取代固定时长参数。这对应仓库中的 duration-head 补丁文件。显著改进的蒸馏模型Substantially improved distilled model——在更小、更快的检查点上保留更多完整模型的视觉质量、提示遵循度与运动一致性。模型家族与检查点清单LTX-2.5 以**拆分、Comfy 对齐的包split, Comfy-aligned pack**形式发布每个组件一个 .safetensors 文件而不是单一整体。使用时只需把各 CLI 参数 / 加载器指向对应文件。下面按 README 的分类逐项核对仓库实际存在的文件。扩散模型Transformers / DiT文件说明diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors蒸馏版 DiTbf16。固定 8 步采样调度CFG1适合追求速度的默认推理。diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors完整 / 可训练版 DiTbf16用于微调与高保真推理。diffusion_models/ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors蒸馏 DiT 的 Comfy int8 convrot 量化版。仅 ComfyUI 可用不适用于ltx-pipelines/ PyTorch 路径。diffusion_models/ltx-2.5-22b-dev-transformer-comfy-int8-convrot.safetensors完整 DiT 的 Comfy int8 convrot 量化版。仅 ComfyUI 可用。diffusion_models/ltx-2.5-22b-distilled-transformer-nvfp4.safetensors蒸馏 DiT 的 NVFP4 量化版。可用于 ComfyUI或ltx-pipelines配合--quantization nvfp4-prequant需要 Blackwell 硬件 /ltx-kernels。关键区分点bf16 检查点供ltx-pipelinesPyTorch路径使用*-comfy-int8-convrot.safetensors文件是 ComfyUI 专属不会被 PyTorch 路径加载务必按运行环境选择对应文件。其他组件文件说明text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensorsGemma4 文本编码器 投影层bf16。text_encoders/gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors同款文本编码器Comfy int8 版——仅 ComfyUI。vae/ltx-2.5-video-vae-bf16.safetensors视频 DiffVAE——质量更高、更重。vae/ltx-2.5-video-vae-conv-bf16.safetensorsConv VAE——更快、更轻。vae/ltx-2.5-audio-vae-bf16.safetensors音频 VAE 声码器vocoder。loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors蒸馏版 LoRA用于 dev-transformer 工作流。model_patches/ltx-2.5-duration-head-bf16.safetensors时长预测头省略--num-frames时自动决定时长。latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensorsx2 空间上采样器多阶段管线必需。latent_upscale_models/ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensorsx2 时间上采样器。这套拆分包结构是理解 LTX-2.5 一切用法的钥匙文本编码、视频生成、音频生成、时长决策、分辨率提升分别由独立文件承载你可以按需只下载/替换其中一部分也可以自由混搭例如用轻量 Conv VAE 换速度。使用方式一Pythonltx-pipelinesLTX-2.5 的权重是**拆分split**的官方 LTX-2 仓库的ltx-pipelines包通过--transformer-path、--text-encoder-path等参数逐组件加载。环境安装README 给出的推荐环境要求Python 3.12、CUDA 12.7、PyTorch ~ 2.7。安装步骤如下git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync source .venv/bin/activate关于注意力后端attention backends与可选扩展请参阅官方仓库的 README。使用uv管理虚拟环境也是后续所有推理命令能直接以uv run执行的前提。下载权重由于本仓库采用拆分交付下载时按需指定组件文件。官方路径使用 Hugging Face CLI 的hf命令先登录再下载hf auth login # LTX-2.5 distilled split pack hf download Lightricks/LTX-2.5 \ diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \ text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ vae/ltx-2.5-video-vae-bf16.safetensors \ vae/ltx-2.5-audio-vae-bf16.safetensors \ model_patches/ltx-2.5-duration-head-bf16.safetensors \ latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \ --local-dir models/ltx-2.5下载完成后models/ltx-2.5下的目录结构即与仓库根目录保持一致diffusion_models/、text_encoders/、vae/、model_patches/、latent_upscale_models/。蒸馏版文生视频蒸馏版distilled采用固定 8 步采样、CFG1的调度是最快的入口。完整命令如下uv run python -m ltx_pipelines.distilled \ --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \ --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \ --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \ --duration-head-path models/ltx-2.5/model_patches/ltx-2.5-duration-head-bf16.safetensors \ --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \ --prompt A golden retriever running through a sunny meadow, cinematic lighting \ --seed 42 \ --output-path output_distilled.mp4参数要点--duration-head-path指定 时长预测头。省略--num-frames时模型会根据提示词自行预测片段长度LTX-2.5 的新能力也可以显式指定例如--num-frames 121。--spatial-upsampler-path多阶段multi-stage管线必需的空间上采样器对应 latent_upscale_models 下的 x2 空间上采样文件。帧数与分辨率约束--num-frames必须满足frames % 8 1即 1、9、17、…、121、…宽高必须能被 32 整除。这两个约束是 LTX-2.5 的硬性输入规则详见下文Constraints小节。图生视频在文生视频的基础上只需追加一个或多个--image PATH FRAME_IDX STRENGTH标志即可帧索引 0 首帧条件控制uv run python -m ltx_pipelines.distilled \ --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \ --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \ --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \ --duration-head-path models/ltx-2.5/model_patches/ltx-2.5-duration-head-bf16.safetensors \ --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \ --image path/to/first_frame.jpg 0 1.0 \ --prompt The camera slowly dollies out as wind moves through the grass \ --seed 42 \ --output-path output_i2v.mp4--image的第三个参数STRENGTH为条件强度1.0表示完全遵循首帧。支持传入多个--image以做多帧条件控制。低显存技巧如果 VRAM 紧张README 给出的方案是动态降精度downcast到 fp8 CPU 卸载# Downcast bf16 transformer on the fly CPU offload ...existing flags... \ --quantization fp8-cast \ --offload cpu再次强调PyTorch 路径必须使用 bf16 检查点*-comfy-int8-convrot.safetensors是 ComfyUI 专属格式不会也不应被ltx-pipelines加载。若你使用 NVFP4 量化文件则需在ltx-pipelines中传--quantization nvfp4-prequant且要求 Blackwell 架构 GPU 与ltx-kernels支持。Python API同一套拆分路径除 CLI 外ltx-pipelines还暴露了 Python 接口路径参数完全一致from ltx_pipelines.distilled import DistilledPipeline from ltx_pipelines.utils.model_paths import ModelPaths model_paths ModelPaths.from_split( transformer_pathmodels/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors, text_encoder_pathmodels/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors, video_vae_pathmodels/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors, audio_vae_pathmodels/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors, duration_head_pathmodels/ltx-2.5/model_patches/ltx-2.5-duration-head-bf16.safetensors, ) pipe DistilledPipeline( model_pathsmodel_paths, spatial_upsampler_pathmodels/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors, ) # See packages/ltx-pipelines for __call__ args (prompt, seed, num_frames, images, ...).这里可以看到拆分路径的映射关系ModelPaths.from_split(...)显式接收 transformer、text encoder、两个 VAE 与 duration head 的路径DistilledPipeline再单独接收空间上采样器路径。CLI 的所有参数prompt、seed、num_frames、images等在__call__中均有对应实现具体签名以packages/ltx-pipelines源码为准。查看完整 CLI 参数列表uv run python -m ltx_pipelines.distilled --help使用方式二ComfyUI官方 LTX-2.5 工作流模板随 ComfyUI 一起发布README 称之为 official LTX-2.5 workflow templates。ComfyUI 用户应选用仓库中带comfy标记的 int8 convrot 量化文件Transformer 与文本编码器各一份例如 蒸馏版 Comfy int8 Transformer 与 Comfy int8 文本编码器。这两类文件与 PyTorch 路径互不通用选择运行环境时务必区分。使用方式三Diffusers两阶段图生视频README 提供了一个 Diffusers 兼容包Lightricks/LTX-2.5-Diffusers同一模型、Diffusers 友好的打包形式。安装LTX-2.5 支持尚未进入正式的diffusersrelease需要从 main 分支安装pip install githttps://github.com/huggingface/diffusers两阶段图生视频完整示例这是 README 中最完整的代码示例核心思想是两阶段two-stage阶段 1 在较低分辨率生成阶段 2 经潜空间上采样后在 2 倍分辨率精修。代码要点已用注释标出import torch from diffusers import LTX2ImageToVideoPipeline, LTX2LatentUpsamplePipeline from diffusers.pipelines.ltx2.latent_upsampler import LTX2LatentUpsamplerModel from diffusers.pipelines.ltx2.utils import ( DEFAULT_NEGATIVE_PROMPT, DISTILLED_SIGMA_VALUES, STAGE_2_DISTILLED_SIGMA_VALUES, ) from diffusers.utils import encode_video, load_image MODEL_ID Lightricks/LTX-2.5-Diffusers # Stage 1 resolution; stage 2 runs at 2x this. HEIGHT, WIDTH, NUM_FRAMES, FRAME_RATE 544, 960, 121, 24.0 pipe LTX2ImageToVideoPipeline.from_pretrained(MODEL_ID, dtypetorch.bfloat16) pipe.enable_model_cpu_offload() pipe.vae.enable_tiling() # stage 2 decodes at 2x latent_upsampler LTX2LatentUpsamplerModel.from_pretrained( MODEL_ID, subfolderlatent_upsampler, dtypetorch.bfloat16 ).to(cuda) upsample_pipe LTX2LatentUpsamplePipeline(vaepipe.vae, latent_upsamplerlatent_upsampler) generator torch.Generator(cuda).manual_seed(42) shared dict( imageload_image(path/to/first_frame.jpg), promptThe camera slowly dollies out as wind moves through the grass, negative_promptDEFAULT_NEGATIVE_PROMPT, frame_rateFRAME_RATE, guidance_scale1.0, audio_guidance_scale1.0, stg_scale0.0, audio_stg_scale0.0, modality_scale1.0, audio_modality_scale1.0, generatorgenerator, return_dictFalse, ) stage_1_latents, audio_latents pipe( heightHEIGHT, widthWIDTH, num_framesNUM_FRAMES, sigmasDISTILLED_SIGMA_VALUES, output_typelatent, **shared, ) upsampled_latents upsample_pipe( latentsstage_1_latents, output_typelatent, return_dictFalse )[0] # Stage 2 takes its size from the upsampled latents, so pass no height/width. video, audio pipe( num_framesNUM_FRAMES, sigmasSTAGE_2_DISTILLED_SIGMA_VALUES, latentsupsampled_latents, audio_latentsaudio_latents, noise_scaleSTAGE_2_DISTILLED_SIGMA_VALUES[0], output_typenp, **shared, ) encode_video( video[0], fpsint(FRAME_RATE), output_pathoutput_i2v_two_stage.mp4, audioaudio[0].float().cpu(), audio_sample_ratepipe.vocoder.config.output_sampling_rate, )值得注意的细节分辨率策略阶段 1 为544 x 960121 帧、24 fps阶段 2 通过LTX2LatentUpsamplePipeline对潜变量做 2 倍上采样因此阶段 2 不再传 height/width尺寸自动取自上采样后的潜变量最终输出约 1088 x 1920。这与仓库中 空间上采样器 的 x2 定位一一对应。采样调度常量蒸馏模型使用DISTILLED_SIGMA_VALUES阶段 2 精修使用STAGE_2_DISTILLED_SIGMA_VALUES同时以该常量的首个值作为noise_scale注入噪声。音频链路管线同时产出audio潜变量阶段 2 返回音频后由声码器vocoder以pipe.vocoder.config.output_sampling_rate采样率编码进最终 MP4——这就是 README 强调同步视频与音频的 Diffusers 侧实现。默认负向提示词与尺度参数DEFAULT_NEGATIVE_PROMPT、guidance_scale1.0、audio_guidance_scale1.0、stg_scale0.0、audio_stg_scale0.0、modality_scale1.0、audio_modality_scale1.0均来自 diffusers.pipelines.ltx2.utils对应 LTX-2.5 蒸馏调度与多模态尺度设计。输入约束与提示词实践README 明确列出的硬性约束帧数num_frames % 8 1即 1、9、17、…、121、…宽高必须能被 32 整除两条约束在三条使用路径CLI 的--num-frames、Diffusers 的NUM_FRAMES/HEIGHT/WIDTH、ComfyUI 工作流中同样适用是排查参数非法类报错的第一检查项。提示词PromptingREADME 明确指出结构良好、详细well-structured, detailed的提示词能实质性提升生成质量。多镜头multishot提示写法与完整指南可参考官方如何提示 LTX-2文档结合本文开头的新特性多镜头场景下要充分利用 Gemma 4 12B 编码器兜住长提示的能力把角色、运镜、光照、动作逐项写全而不是依赖模型自行脑补。训练与微调dev Transformer 完全可训练fully trainable用于复现官方发布的 LoRA 与 IC-LoRA 需配合 LTX-2 Trainer 使用。基于 Lightricks 自身测试绝大多数在 LTX-2.3 上训练的 LoRA 与 IC-LoRA 可在 LTX-2.5 上无改动直接运行但存在少量例外——生产环境使用前务必对适配器进行验证。仓库中的 蒸馏版 LoRA 文件distilled LoRA用于 dev-transformer 工作流即是这一生态的落地示例。已知限制LimitationsREADME 以模型卡形式如实列出局限使用前应充分认知本模型不打算也没有能力提供事实性信息not intended or able to provide factual information。作为统计模型检查点可能放大现存的社会偏见。提示遵循度受提示风格显著影响prompt following is heavily influenced by prompting style。模型可能无法生成与提示词完全匹配的视频。模型可能生成不当或冒犯性的内容。引用与进一步阅读README 给出了 LTX-2 论文的 BibTeX 引用article{hacohen2025ltx2, title{LTX-2: Efficient Joint Audio-Visual Foundation Model}, author{HaCohen, Yoav and Brazowski, Benny and Chiprut, Nisan and Bitterman, Yaki and Kvochko, Andrew and Berkowitz, Avishai and Shalem, Daniel and Lifschitz, Daphna and Moshe, Dudu and Porat, Eitan and Richardson, Eitan and Guy Shiran and Itay Chachy and Jonathan Chetboun and Michael Finkelson and Michael Kupchick and Nir Zabari and Nitzan Guetta and Noa Kotler and Ofir Bibi and Ori Gordon and Poriya Panet and Roi Benita and Shahar Armon and Victor Kulikov and Yaron Inger and Yonatan Shiftan and Zeev Melumian and Zeev Farbman}, journal{arXiv preprint arXiv:2601.03233}, year{2026} }如需继续深入可直接在本仓库内核对各组件文件模型权重清单见 diffusion_models/、text_encoders/、vae/、loras/、model_patches/、latent_upscale_models/ 各目录模型卡片正文与授权条款见 README.md。本仓库为只读镜像所有查看、下载与推理均应在本地执行并按上文路径参数自行组织models/ltx-2.5目录。赞分享大模型媒体生成音视频多模态本地部署【免费下载链接】LTX-2.5项目地址https://ai.gitcode.com/hf_mirrors/Lightricks/LTX-2.5点击查看免费下载相关推荐AudioCraft 开源音频生成库完全指南模型家族、安装配置与训练实践AudioCraft 开源音频生成库完全指南模型家族、安装配置与训练实践 AudioCraft 是一个基于 PyTorch 的深度学习音频生成研究库集成了训人工智能深度学习音频媒体生成音乐生成d3-annotation常见问题解答从安装到部署的全方位解决方案d3 annotation常见问题解答从安装到部署的全方位解决方案 你是否在使用D3.js创建数据可视化时想要为图表添加专业的标注和注释却不知道从何入手突破视频修复瓶颈SeedVR-3B模型家族全版本选型与实战指南突破视频修复瓶颈SeedVR 3B模型家族全版本选型与实战指南 你是否还在为模糊视频修复效果不佳而困扰面对AIGC视频的细节丢失束手无策普通修复模型在处理计算机视觉深度学习大模型创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表