SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[Commercial Use Allowed] Z-Image_clear_photoreal Released | A Photorealistic Model Combining High Speed and High Quality [Image Generation AI]

Introduction

Hello, I'm Kimama / Easygoing.

This time, I am introducing the definitive version of an image generation AI that is fast, high-quality, and allows for commercial use: Z-Image_clear_photoreal.


Z-Image_clear_photorealで生成された、高精細でリアルな女性のポートレート。鮮やかな光の表現と肌の質感が特徴的。
Z-Image_clear_photoreal

What is Z-Image?

Z-Image is an image generation AI released by Alibaba Group in China on January 27, 2026, under the Apache-2.0 license, which also permits commercial use.

Z-Image と Flux.2 [klein] の登場時期を示したガントチャート

The Z-Image series gained popularity as an image generation AI that runs on mid-range PCs, following the release of the high-speed distilled version, Z-Image-Turbo, in November 2025.

While Z-Image-Turbo was a distilled model specialized for high-speed generation, the Z-Image (Base) released this time is the core model before distillation, which is essentially the raw brain of the trained AI.

Distilled models have less variety

Distilled models, including Z-Image-Turbo, undergo adjustments to improve usability.

Characteristics of distilled models

  • Improved stability and speed

  • Degraded image quality

  • Reduced variety

Distilled models undergo adjustments to increase stability in order to reduce failures in illustration generation.

This allows illustrations to converge faster, enabling high-speed generation with low steps, but in exchange, vivid color expression is lost, the overall tone becomes brownish, and residual noise tends to remain in the finished illustration.

Z-Image_clear_photorealで生成された、高精細で青い服を着た女性とケーキとワインのポートレート。
Z-Image_clear_photoreal

Furthermore, increasing stability makes it easier to generate similar illustrations, leading to a lack of variety in faces and compositions.

Mixing the core model based on the distilled model

For this Z-Image_clear_photoreal, we aimed to Z-Image-Turbo_clear, the distillation model introduced previously, as a base, and by mixing in the original Z-Image (Base) at specific layers, we aimed for improved image quality and variety.

By doing this, we aim for further improvements in image quality while maintaining the high speed without using CFG of Z-Image-Turbo.

Actual illustrations!

Now, let's compare the actual illustrations.

On the left is the newly adjusted Z-Image_clear_photoreal model, and on the right is the output from the original Z-Image-Turbo_clear.

Fireplace room

「暖炉のある部屋」のプロンプトによる比較。左のphotorealモデルは右のTurboモデルよりライティングが自然。
Z-Image_clear_photoreal | Z-Image-Turbo_clear

Late-night radio

「深夜ラジオのスタジオ」の生成比較。マージモデルでは細部のガジェット感や空気感が向上している。

Magnified view

人物の肌のクローズアップ比較。右のTurbo版に見られる微細な粒状ノイズが、左のphotoreal版では滑らかに改善されている。

When comparing image quality, it is easy to tell by looking at the skin texture; while the Z-Image-Turbo_clear on the right retains the granular noise characteristic of high-speed distillation models in the finished illustration, you can see that this has been improved and looks cleaner in the Z-Image_clear_photoreal on the left.

Comparing variety!

Next, let's compare the variety of the illustrations.

I will now generate eight consecutive illustrations while keeping the same prompt but changing the seed value.

Prompt (Woman in traditional attire)

realistic, photorealistic, a female wears a purple and gold traditional dress and jewelry, standing in front of a snowy village at sunset or sunrise, surrounded by snow - covered houses and a warm orange and pink sky with orange and pink hues., dynamic angle, dutch angle, upper body, close up, face close up, happy, smile, laugh, peaceful, wind, rouge, alizarin, burgundy, maroon, indigo, royal blue, deep blue, deep purple, royal purple, stylish, elegant, turn around



Z-Image_clear_photoreal (This merged model)

Z-Image_clear_photorealを使用し、同じプロンプトでseed値を変えて生成した8枚の画像のうち前半の4枚。顔立ちや構図が多様。
Z-Image_clear_photorealを使用し、同じプロンプトでseed値を変えて生成した8枚の画像のうち後半の4枚。顔立ちや構図が多様。
Rich variety

Z-Image-Turbo_clear (Distillation model)

Z-Image-Turbo_clearで生成した8枚の画像のうち前半の4枚。顔のパーツや基本的な構図がほぼ同じで、変化が少ない様子。
Z-Image-Turbo_clearで生成した8枚の画像のうち後半の4枚。顔のパーツや基本的な構図がほぼ同じで、変化が少ない様子。
Little change in faces or composition

Comparing the two models, you can see that the newly adjusted Z-Image_clear_photoreal shows rich variety in the character's face and composition, whereas the original Z-Image-Turbo_clear has little change in faces or composition, resulting in similar illustrations.

Merge Recipe for Z-Image_clear_photoreal

Now, I will introduce the merge recipe for the Z-Image_clear_photoreal I created this time.

ComfyUI上の「Model Merge Z-Image」ノードの設定画面。BaseとTurboのモデルを層ごとにマージする設定値が表示されている。

The Model Merge Z-Image node on the right is a node for mixing two Z-Image models, where the 0 setting value merges layers from Z-Image (Base), and the 1 setting value merges layers from Z-Image-Turbo_clear.

In other words, the initial layers use the diverse training data of Z-Image (Base) to increase variation, while the middle to later layers use the distilled Z-Image-Turbo_clear to maintain stability and generation speed.

Model Merge Z-Image Custom Node

Let's use Z-Image_clear_vae for the VAE!

For this Z-Image_clear_photoreal model, I recommend using it in conjunction with Z-Image_clear_vae, a custom model I created.

  • Z-Image_natural_vae: Natural expression with suppressed saturation

  • Z-Image_clear_vae: Vivid expression with increased saturation

Since I have prepared several Z-Image_clear_vae models with different color expressions, please try using them according to your preference.

Next time: Introduction to Z-Image_clear_anime!

This time, when adjusting Z-Image, the direction of adjustment was completely different between photorealistic illustrations and anime illustrations, so I created Z-Image_clear_photoreal as a dedicated model for photorealism.

Z-Image_clear_animeのティーザー画像。鮮やかな色使いのアニメ調キャラクターイラスト。
Next time: Introduction to Z-Image_clear_anime!

Next time, I plan to introduce the Z-Image_clear_anime model for anime illustrations.

Summary: Let's try using Z-Image_clear_photoreal

  • Z-Image (Base) has abundant variations

  • Z-Image-Turbo offers stability and high-speed generation

  • Z-Image_clear_photoreal takes the best of both worlds

This time, I introduced the Z-Image_clear_photoreal model.

When actually performing adjustments, I realized that Z-Image has almost no redundant structure, and it is a model that Alibaba has deeply researched Qwen-Image to achieve the best optimization and lightweighting.

店の飲料の冷蔵棚の前の女性のポートレート。Z-Image_clear_photorealモデルのコンセプトイメージグラフィック。
Z-Image offers an exquisite balance of speed and image quality.

With the release of the core Z-Image (Base) model, it has become easier than ever to develop the Z-Image series, and I believe that excellent improved versions from the community will emerge from here on.

Why not take this opportunity to try out Z-Image's high-quality, high-speed generation for yourself?

Thank you for reading until the end!


English Article


いいなと思ったら応援しよう!