【ComfyUI】LTX-2.5の使い方|T2V・I2V・FLF2Vの公式ワークフロー

LTX-2.5がComfyUIで公式にサポートされたので、テンプレートのworkflowをそのまま動かしてみました。

テキストからの生成に加えて、画像1枚から動かすi2v、最初と最後のコマをつなぐflf2vの3種類が用意されています。

この記事ではアクセス申請からモデルの配置、3つのworkflowの中身までをまとめました。

LTX-2.5とは?

LTX-2.5はLightricksが公開している動画生成モデルで、映像と音声を同時に生成できます。

1回の生成でカットをまたいだ複数ショットを作れるようになったのが2.5の目玉とのこと。

ComfyUI向けにはint8+convrotで量子化されたファイルが配布されていて、こちらはComfyUI専用と書かれています。

公式のモデルカードに挙げられている主な特徴は以下の通り。

  • Pixel Diffusion:高精細なキーフレームを先に作ってから映像を組み立てる方式
  • Diffusion Video Decoder:顔や文字のにじみ、速い動きでの崩れを抑えるデコーダー
  • ネイティブマルチショット:1回の生成でカットをまたいでキャラや照明、声を保つ
  • Gemma 4 12Bベースのテキストエンコーダー:長いプロンプトの要素を取りこぼしにくい
  • Prompt Enhancer:短いプロンプトを映像向けの指示に膨らませる軽量モデル
  • Auto Duration:書かれた動作の内容からクリップの長さを推定
  • ネイティブ4Kと音声同期は2.3から引き継ぎ

モデルカードとファイル一式はHugging Faceで公開されています。

Lightricks/LTX-2.5 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

LTX-2.5のライセンス

LTX-2.5は従来と同じくLTX-2.xのコミュニティライセンスで配布されています。

年間売上が1000万ドル未満であれば、商用や制作物への利用も無償。

これを超える規模の場合は、別途有償の商用利用契約が必要になります。

ファインチューンしたモデルの配布には別の条件が付く場合があると記載。

詳細はモデルカードのライセンス欄で確認できます。

LTX-2/LICENSE.md at main · Lightricks/LTX-2

動作環境

私の動作環境は以下の通りです。

項目内容
GPUGeForce RTX 4060 Ti(VRAM 16GB)
メインメモリ64GB
ComfyUIStability Matrix経由で導入

公式のモデルカードに必要VRAMの明記は見当たらず、VRAMが少ない場合は量子化とCPUオフロードを使う案内が出ている程度でした。

ダウンロードするモデルは合計で約46GBあるので、ディスクの空きも先に見ておく必要があります。

ComfyUIのインストール

ComfyUIのインストールがまだの方はこちらを参考にしてみてください。

【ComfyUI】の使い方・始め方まとめ|インストールから基本操作・カスタムノードまで
ComfyUIを始めたいけど、インストール方法がいくつかあってどれを選べばいいか迷う、という方は多いと思います。この記事では、ComfyUIの導入方法の違いと選び方、そして基本的な使い方の入口をまとめて紹介します。各トピックの詳しい手順は、…

LTX-2.5はComfyUI 0.32.0から対応しています。

アップデートがまだの方はアップデートも必要です。

モデルのダウンロード

Hugging Faceでアクセス申請する

LTX-2.5のリポジトリはgated(アクセス制限付き)です。

ログインした状態でモデルページを開き、ライセンスに同意して申請を出せばアクセスできるようになります。

Hugging FaceのLTX-2.5ページに出る、連絡先の共有に同意してアクセスを申請するボタン
Lightricks/LTX-2.5 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

ダウンロードするファイルと置き場所

公式workflowで使うファイルは6つあり、置き場所は以下の通り。

ファイル置き場所サイズの目安
ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensorsmodels\diffusion_models約20GB
gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensorsmodels\text_encoders約14GB
gemma4_e2b_it_bf16.safetensorsmodels\text_encoders約9.5GB
ltx-2.5-video-vae-bf16.safetensorsmodels\vae約1.4GB
ltx-2.5-audio-vae-bf16.safetensorsmodels\vae約0.3GB
ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensorsmodels\latent_upscale_models約0.9GB

Stability Matrixの場合はData\Models側に置いても認識されないため、ComfyUIパッケージ内のmodelsフォルダへ直接入れてください。

ComfyUI\models\latent_upscale_models

workflowの読み込み

ComfyUIのメニューからテンプレートを開き、Videoのカテゴリ内でLTX-2.5を選びます。

用意されているのはt2v、i2v、flf2vの3つ。

ComfyUIのテンプレート一覧に並ぶLTX-2.5のImage to Video・FLF2V・Text to Videoのworkflow

同じworkflowの解説はComfyUI公式のチュートリアルにも用意されています。

LTX-2.5: ComfyUI ワークフロー例 – ComfyUI
ComfyUI で LTX-2.5 を使用する方法を説明します。テキストから動画へ、画像から動画へ、先頭・最終フレームから動画生成のためのネイティブワークフローを備え、ピクセルディフュージョン品質、ネイティブマルチショットシーン、同期オーデ…

LTX-2.5の使い方

3つのworkflowはどれも中身がサブグラフ1個にまとまっています。

表に出ているのはプロンプトとサイズ、長さ、モデル指定くらいでパラメータもほぼ共通です。

項目内容
promptシーン、動き、照明、カメラ、音などを書くプロンプト
prompt_enhancePrompt Enhancerのオンオフ
durationクリップの長さ(秒)
width / height出力サイズ(初期値は1280×720)
seedシード値
frame_rateフレームレート(初期値は24)
unet_name本体のdistilled transformer
video_vae / audio_vae映像用と音声用のVAE
clip_nameGemma 4 12Bのテキストエンコーダー
upscale_modelデコード前に使うlatent upscaler
prompt_enhance_modelPrompt Enhancerが使うモデル

Text to Video(t2v)

テキストから動画を作るworkflowです。

Text to Video (LTX-2.5)というサブグラフにまとまっています。

Text to Video (LTX-2.5)ノードのプロンプト欄とprompt_enhance・duration・モデル指定の項目

基本はプロンプト書いて、モデルセットするくらいです。

あと長さを変えたい場合はduration、prompt_enhanceのオン・オフ、解像度はResolution Selectorで指定可能です。

デフォルトでは1280×720になっています。

アスペクト比とmegapixelsから解像度を決めるResolution Selectorノード

Text to Video(t2v)の実行結果

1280×720(5秒)の動画で約2分19秒(139.54 seconds)でした。

prompt_enhanceもオンの状態で実行したのでかなり生成速度は速いと思います。

A cinematic medium close-up of an elderly watchmaker seated at a cluttered wooden workbench in a small repair shop at night. A single brass desk lamp throws warm amber light across scattered gears and a half-open pocket watch, while the rest of the room falls into deep green-brown shadow. He is in his seventies, thin silver hair combed back, a worn grey cardigan over a collared shirt, a magnifying loupe clipped to his glasses. He lowers the loupe to his eye, lifts the pocket watch with steel tweezers, and his jaw tightens as the second hand finally begins to move; a slow smile spreads across his face. The camera pushes in gently from his hands to his face, holding on his eyes as they catch the lamplight. Ambient audio carries the layered ticking of dozens of clocks, faint rain against the window, and the soft metallic click of tweezers on brass. He murmurs in a low, gravelly English voice, “There you are, old friend.”

Image to Video(i2v)

Load First Frameに画像を入れるくらいで他はt2vと一緒です。

サブグラフの中にはText to Videoへ切り替えるスイッチもあり、画像なしの生成もそのまま試せます。

Load First Frameにエルフの画像を読み込んだImage to Video (LTX-2.5)のworkflow

Image to Video(i2v)の実行結果

生成時間は約3分7秒(187.73 seconds)でした。

prompt_enhanceオンでやったのですが、あまり元のプロンプトの内容は反映できていません。

Use the provided start image as the first frame. A cinematic medium shot of a blonde elf archer in 2D anime style pressed against a moss-covered tree in a sunlit forest, dappled green light shifting across her cloak. A branch snaps off-screen right and her head whips toward the sound, eyes widening as she shoves off the trunk and breaks into a sprint down the shadowed trail, cloak snapping out behind her and hair streaming back. The camera whip-pans with her and settles into a fast handheld tracking shot from behind her shoulder, ferns and trunks rushing past the lens in motion blur while shafts of sunlight strobe across the frame. Behind her a massive dark shape bursts through the undergrowth, a hulking horned beast with glowing amber eyes, shouldering trees aside as it gives chase. She glances back over her shoulder mid-stride, then rips an arrow from the quiver as she runs. Ambient audio carries pounding footsteps on wet earth, cracking wood, her ragged breathing, and a deep guttural roar closing in from behind.

prompt_enhanceオフでやると生成時間は約2分(120.12 seconds)。

若干内容はズレてますが、prompt_enhanceオンのときよりはちゃんと動きが再現できました。

First & Last Frame to Video(flf2v)

flf2vは最初と最後のフレームを指定して動画生成できるものです。

Load First FrameとLoad Last Frameで2枚指定できます。

Load First FrameとLoad Last Frameに2枚の画像を読み込んだFirst & Last Frame to Video (LTX-2.5)のworkflow

First & Last Frame to Video(flf2v)の実行結果

生成時間はprompt_enhanceありで約3分17秒(197.57 seconds)でした。

ただ途中トランジションみたいな切り替えが入ってしまい、MiniMaxH3のflf2vと比較すると一貫性などは劣る印象です。

Use the provided start image as the first frame and the provided end image as the final frame. A continuous studio portrait take of a young woman with shoulder-length dark brown hair and blunt bangs, gold hoop earrings and a thin gold pendant necklace, wearing a beige off-shoulder ribbed knit dress, in front of a softly lit white paneled wall with dried pampas grass in a glass vase at frame left. She begins seated in a navy velvet chair in medium close-up, smiling gently at the lens, then plants a hand on the armrest and rises smoothly to her feet, her hair swaying and the dress falling into place as she straightens. The camera pulls back and tilts down in one steady move, widening from her face to a full-body shot as she steps forward and settles her weight on one hip. She places both hands on her waist, squares her shoulders toward the lens and holds a warm closed-lip smile, chin slightly lowered. Soft diffused daylight stays constant across the whole move, with gentle shadows on the wall behind her. Ambient audio carries a quiet room tone, the faint rustle of knit fabric, and a soft breath as she stands.

Prompt Enhancerとプロンプトの書き方

prompt_enhanceをオンにすると、短いプロンプトをGemma 4ベースの軽量モデルが映像向けに書き直してくれます。

そのために必要なのがgemma4_e2b_it_bf16.safetensorsで、これだけで約9.5GB使います。

自分で細かく書く場合はオフにすれば、生成速度も向上し、入力した文章がそのまま使われます。

書き方の指針は、LTX-2.5向けに公開されている公式のプロンプトガイドがあるので、こちらを参考にしてください。

LTX-2.5 Prompt Guide: Prompting Tips For LTX-2.5 | LTX Blog
Learn how to write effective LTX-2.5 video prompts — shot design, multi-shot cuts, Dub-It dialogue, and Video Editing IC…

うまく動かないとき

おそらく詰まりやすいのはこのあたりです。

  • テンプレートにLTX-2.5が出てこない、ノードが赤い:ComfyUI本体が古い
  • ダウンロードで権限エラーになる:Hugging Faceのアクセス申請がまだ承認されていない
  • モデルが選択肢に出てこない:フォルダ名か置き場所が違う(latent_upscale_modelsは特に注意)
  • 途中でVRAMが足りなくなる:解像度とdurationを下げる、prompt_enhanceをオフにする

特にgatedの承認とlatent_upscale_modelsの置き場所は環境によって迷いやすいので、確認してみてください。

まとめ

MiniMaxH3と比較するとこんな印象でした。

LTX-2.5:低スペックでも高解像度が生成しやすい、生成速度が速い、画質も綺麗

MiniMaxH3:LTX-2.5より処理速度は劣るが、プロンプトの再現性、人物の一貫性などは高い

生成速度や解像度とか考えるとLTX-2.5の方が生成しやすいと思いますが、自分の脳内の動画をちゃんと再現するのであれば、MiniMaxH3の方が強いかなと感じました。

あくまでデフォルトのモデルを使用して感じたことなので、今後いろいろ便利なLoRAや高速化が出たら評価は変わるかもしれません。

以上、ComfyUIでLTX-2.5を動かす方法を紹介しました。

参考になれば幸いです。

タイトルとURLをコピーしました