Qwen-Image-2512をComfyUIで使う方法|GGUF・4ステップLoRA比較

Qwen-Image旧版と2512の生成結果を並べた比較画像

Qwen-Image-2512は、人物の質感や画像内の文字表現が改善されたQwen-Imageの更新モデルです。

ComfyUIでは公式テンプレートが用意されており、通常版に加えてGGUFや4ステップLoRAも利用できます。

今回はQwen-Image-2512をComfyUIで使う方法と、旧版・2512・4ステップLoRAの比較結果を紹介します。

この記事で分かること
  • Qwen-Image-2512で必要なモデルと保存先
  • ComfyUI公式ワークフローの読み込み方と設定
  • 旧版・2512・4ステップLoRAの生成時間と画質差

Qwen-Image-2512とは?

人物や動物、料理、文字表現をまとめたQwen-Image-2512の公式作例

Qwen-Image-2512は、2025年12月31日に公開されたQwen-Imageの画像生成モデルです。

公式モデルカードでは、従来版から人物のリアリズム、自然物の細部、テキストレンダリングが改善されてます。

主な特徴

  • AI特有の質感を抑えた人物表現
  • 風景や動物の毛並みなど、自然物の細かな描写
  • 画像内の文字とレイアウト精度の向上
  • FP8・BF16・GGUFなど、環境に合わせてモデル形式を選択可能
Qwen/Qwen-Image-2512 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

ライセンス

Qwen-Image-2512のモデルカードに記載されているライセンスはApache License 2.0です。

ただし、生成物に含まれる人物・ロゴ・キャラクターなどの権利まで一律に保証するものではないため、公開や商用利用では内容を個別に確認してください。

必要なもの・動作環境

公式テンプレートを使うため、先にComfyUIを最新状態へ更新してください。

テンプレートや必要なコアノードが見つからない場合は、ComfyUIのバージョンが古い可能性があります。

ComfyUIをまだ導入していない場合はこちらを参考にしてください。

【ComfyUI】の使い方・始め方まとめ|インストールから基本操作・カスタムノードまで
ComfyUIを始めたいけど、インストール方法がいくつかあってどれを選べばいいか迷う、という方は多いと思います。この記事では、ComfyUIの導入方法の違いと選び方、そして基本的な使い方の入口をまとめて紹介します。各トピックの詳しい手順は、…

GGUF版を使う場合は、標準のLoad Diffusion ModelではなくComfyUI-GGUFのUnet Loader (GGUF)が必要になります。

GitHub – city96/ComfyUI-GGUF: GGUF Quantization support for native ComfyUI models
GGUF Quantization support for native ComfyUI models – city96/ComfyUI-GGUF

ComfyUI Managerでインストールするかカスタムノードフォルダに手動でクローンしてください。

git clone https://github.com/city96/ComfyUI-GGUF.git

モデルのダウンロードと配置

通常版では、Diffusion Model、Text Encoder、VAEの3ファイルが必要です。

4ステップで生成したい場合は、Lightning LoRAも追加してください。

種類ファイル保存先
Diffusion Modelqwen_image_2512_fp8_e4m3fn.safetensorsComfyUI\models\diffusion_models
Text Encoderqwen_2.5_vl_7b_fp8_scaled.safetensorsComfyUI\models\text_encoders
VAEqwen_image_vae.safetensorsComfyUI\models\vae
4ステップLoRAQwen-Image-2512-Lightning-4steps-V1.0-bf16.safetensorsComfyUI\models\loras

GGUF版はこちら。

unsloth/Qwen-Image-2512-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

公式ワークフローの読み込み

ComfyUIのメニューから「Workflow」→「Browse Templates」を開き、Qwen-ImageのText to Imageテンプレートを選択してください。

ComfyUIのテンプレート一覧に表示されたQwen-Image Text to Image

公式ワークフローには通常の50ステップ生成と、Lightning LoRAを使う4ステップ生成が含まれています。

Qwen-Image-2512 ComfyUI Native Workflow Example – ComfyUI
Qwen-Image-2512 is the December update of Qwen-Image's text-to-image foundational model, featuring enhanced human realis…

ノードの設定と生成

モデルを選択する

Load Diffusion Model

qwen_image_2512_fp8_e4m3fn.safetensorsを選択します。

Load Diffusion ModelでQwen-Image-2512のFP8モデルを選択した設定

GGUFの場合はUnet Loader (GGUF)ノードに置き換えてください。

Unet Loader (GGUF)をLoRAローダーへ接続したComfyUIのノード構成

Load CLIP

qwen_2.5_vl_7b_fp8_scaled.safetensorsを選び、typeはqwen_imageにします。

Load CLIPでQwen 2.5 VLのText Encoderとqwen_imageを選択した設定

Load VAE

qwen_image_vae.safetensorsを指定します。

Load VAEでqwen_image_vae.safetensorsを選択した設定

EmptySD3LatentImage

公式が案内している1:1の解像度は1328×1328です。

ほかにも1664×928(16:9)、928×1664(9:16)、1472×1104(4:3)などに対応しています。

EmptySD3LatentImageを1328×1328、バッチサイズ1にした設定

CLIP Text Encode (Positive Prompt)

生成したい内容を入力します。

CLIP Text Encodeへドラゴンの英語プロンプトを入力した例

LoraLoaderModelOnly

4ステップLoRAを使う場合はLoraLoaderModelOnlyのバイパスを解除し、ダウンロードしたLoRAを選びます。

LoraLoaderModelOnlyでQwen-Image-2512 Lightning LoRAを選択した設定

KSampler

私の環境では、Euler・CFG 1の組み合わせで格子状のアーティファクトが出ることがありました。

そのため比較時はstepsを8、CFGを1、sampler_nameをres_multistep、schedulerをbetaに変更しています。

KSamplerを8ステップ、CFG 1、res_multistep、betaにした設定

従来のモデル・2512・4ステップLoRAを比較

生成時間は従来のモデルと2512は8ステップで約40秒前後、4ステップLoRAは約30秒でした。

(VRAM16GB)

同じプロンプトで、上から従来のQwen-Image、Qwen-Image-2512、Qwen-Image-2512+4ステップLoRAで比較してます。

4ステップLoRAは劣化しているようにも感じました。

従来版、Qwen-Image-2512、4ステップLoRAの生成結果比較

プロンプトはAIで作成したものを使っています。

1. ファンタジーもの
An epic, high-fantasy wide shot of a majestic ancient dragon perched atop a floating crystalline citadel. The dragon’s scales are iridescent, shimmering with metallic hues of cobalt and violet, each scale exhibiting individual reflections and razor-sharp edges. Its wings are partially translucent, showing complex vein structures under the glow of a twin-moon eclipse. The citadel features intricate gothic architecture carved from white marble and floating obsidian fragments, surrounded by swirling mana currents and glowing magical particles. The atmosphere is filled with a soft luminescence; ethereal light rays pierce through swirling nebulae in the background. Cinematic composition, hyper-detailed textures, volumetric lighting, ethereal color palette of deep purples, teals, and gold. 8k resolution, Unreal Engine 5 render style.

2. リアルな女性
A hyper-realistic candid portrait of a woman in her late 20s, captured in the soft, directional light of the golden hour. The focus is razor-sharp on her eyes, revealing intricate amber patterns in the iris and individual wet reflections. Her skin is a study of natural realism: visible pores, faint freckles across the bridge of the nose, fine peach fuzz along the jawline, and subtle natural oils. A few stray strands of chestnut hair are backlit by the sun, creating a glowing halo effect. She wears a coarse-knit linen sweater, with every fiber and stray thread rendered in tactile detail. The background is a soft-focus autumn meadow with creamy bokeh. Naturalistic color grading, warm earthy tones, shallow depth of field, 85mm lens aesthetic, extreme macro detail.


3. テキストを再現するプロンプト
A sleek, modern minimalist storefront at night, featuring a prominent high-contrast LED glass display. Centered perfectly on the glass is the text “Innovation and Excellence: Redefining the Boundaries of Technology in 2026” written in a clean, sophisticated sans-serif typography. The letters are crisp, glowing with a steady cool white light, reflecting clearly onto the polished dark pavement below. The surrounding environment is a high-tech urban setting with blurred city lights and rain-slicked surfaces. The focus is locked on the clarity of the characters, ensuring sharp edges and consistent spacing. Cinematic night photography, cold blue and charcoal tones, sharp focus on text, 8k resolution, hyper-realistic glass reflections.


4. 中世ヨーロッパ的なもの
A gritty, atmospheric scene of a narrow cobblestone street in a 14th-century European town during a light drizzle. The timber-framed houses lean inward, their weathered wood grain and crumbling plaster showing centuries of decay. Wet stones reflect the flickering orange glow of iron braziers and distant torches. Details include thick mud in the gutters, rusted iron signage hanging from wooden beams, and the rough texture of wool tunics on distant figures. The air is heavy with mist and woodsmoke, creating a deep sense of layered atmosphere. The palette is dominated by muted browns, slate grays, and warm flickers of firelight. Historical realism, cinematic chiaroscuro lighting, high-contrast shadows, ultra-fine architectural textures.

まとめ

従来のモデルより画像そのもののクオリティはかなり上がってるように感じました。

生成内容にもよりますが、AI特有の安っぽさが消えてる印象です。

あとZIT程ではないものの、LoRA使わなくても40秒前後なので生成しやすいと思います。

以上Qwen-Image-2512の使い方を紹介しました。

参考になれば幸いです。

タイトルとURLをコピーしました