見出し画像

ComfyUIでFP8/GGUF版のOvis-Imageを試す(VRAM 8GB以上)

Last update 12-5-2025
※ ワークフローはモデルの選択が異なる3種類があるので、1-2.を参考に検討してください。
※ 関連記事は、ComfyUIの記事一覧にまとめてあります。




■ 0. 概要

▼ 0-0. はじめに

 本記事では、ComfyUIで「Ovis-Image」を利用するための手順を説明します。モデルはFP8版とGGUF版を利用します。

 ComfyUIのインストール方法は、下記の記事をご覧ください。


▼ 0-1. Ovis-Imageについて

 Ovis-Imageは、Alibaba International Digital Commerce Group(AIDC)のAIチームであるAIDC-AIが11-29-2025に公開しました。AIDCはAlibaba.comやAliExpress等のeコマース事業を行っている部門で、Qwen、Wan、Z-Image等を手がけているところとは別です。

 Ovis-Imageは画像生成アーキテクチャで、7Bのパラメーター数という小型の規模で、画像内のテキストレンダリングに長けているのが特徴とのことです。


▼ 0-2. 参考リンク

 公式のリポジトリと、テクニカルレポートです。

 ComfyUIに関連するZ-Image-Turboの情報です。 


▼ 0-3. 関連リンク1

 無料で生成のおためしができます。

 生成AIプレイグラウンドのfalも利用可能です。執筆時点では$0.012/MP です。


▼ 0-4. 関連リンク2

 公式のモデルです。

 本記事で利用するモデルです。VAEはZ-Image-Turboと同様、FLUX.1と共通です。



■ 1. 留意事項と使用リソース量

▼ 1-1. 留意事項

 ComfyUIを更新していないと、必要なノードや機能が存在しない場合があります。バグフィックスを含むこともあるので、なるべく最新版にしてください。

 次回の生成を高速化するため、VRAMに読み込んだモデルが自動的にメインRAMへ待避(オフロード)されることがあります。そのため、メインRAMは32GB以上が望ましいです。


▼ 1-2. 必要なリソースと生成時間

 VRAM使用量の目安です。ただし、実行中のリソース使用量は環境によって大きく異なります。VRAMが12~16GBの場合はGGUF版1を、8GBの場合はGGUF版2をおすすめします。

  • FP8版
    おおむね10GB程度。

  • GGUF版1(FP8版のModel、GGUF版のText Encoder (IQ4_NL) 、GGUF版のVAEを使用)
    おおむね10GB程度。実行時間はGGUF版1とほぼ同じ。

  • GGUF版2(GGUF版のModel (IQ4_NL) 、GGUF版のText Encoder (IQ4_NL) 、GGUF版のVAEを使用)
    おおむね8GB未満。Modelの量子化タイプにより上下する。

 筆者の環境(GeForce RTX 5060Ti 16GB)での生成時間は下記のとおりです。生成画像の大きさは1MP(1024x1024等)としました。ディスクからのモデルの読み込みや、プロンプトの処理の時間は含まれていません。その他の条件は上記と同じです。

  • FP8版とGGUF版1
    54秒程度(20Steps)

  • GGUF版2
    70秒程度(20Steps)

 Z-Image-Turboと比べて時間がかかる理由は、(単純比較はできませんが)Stepsが9から20に上がっている点と、CFGが1.0ではない点が考えられます。



■ 2. モデルの設置

▼ 2-1. はじめに

 FP8版を利用する場合は2-2.を、GGUF版(2パターンあり)の場合は2-3.を参照してください。

  • FP8版 → 2-2.

  • GGUF版 → 2-3.


▼ 2-2. FP8版

 今回はFP8版のModelをダウンロードするのではなく、BF16形式のModelをFP8形式でロードして利用します。

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。なお、VAEの設置場所のみ「flux1」フォルダにしておきました。

  • Comfy-Org/Ovis-Image
    https://huggingface.co/Comfy-Org/Ovis-Image 

    • split_files/diffusion_models/
      ovis_image_bf16.safetensors (13.7GB)
      models\diffusion_models\ovis-image\

    • split_files/text_encoders/
      ovis_2.5.safetensors (4.78GB)
      models\text_encoders\ovis-image\


▼ 2-3. GGUF版

 GGUF形式のモデルを利用するにあたり、「Text Encoderのみで利用する」「ModelとText Encoderで利用する」の2パターンを提案し、前者をおすすめします。

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。

 BF16形式のModelを利用(FP8形式でロード)したい場合は、下記をダウンロードします。3.に掲載しているワークフローは「GGUF版1」を選択してください。

 GGUF形式のModelを利用したい場合は、代わりに下記をダウンロードします。3.に掲載しているワークフローは「GGUF版2」を選択してください。

 次に、GGUF形式のText EncoderとVAEをダウンロードします。VAEの設置場所のみ「flux1」フォルダにしておきました。



■ 3. 画像の生成

▼ 3-1. はじめに

 FP8版やGGUF版のワークフローを用いて画像を生成します。FP8版の場合は3-3.へ進んでください。

  • 必要な拡張機能について(GGUF版のみ) → 3-2.

  • ワークフローについて → 3-3.

  • 実行例 → 3-4.


▼ 3-2. 必要な拡張機能について(GGUF版のみ)

 本記事に掲載したGGUF版のワークフローを読み込んだとき、拡張機能が不足していると「Some Nodes Are Missing」の表示が出ます。その場合は、ComfyUI Managerの機能を利用してインストールすることができます。

 拡張機能をインストールする方法は、下記の記事の「自動インストールの手順 」または「手動インストールの手順」を参照してください。

 本記事のワークフローでは下記の拡張機能を利用します。


▼ 3-3. ワークフローについて

 生成用のワークフローを提供しますので、改造、再公開を含め自由にご利用ください。

 FP8版、GGUF版1(Text EncoderとVAEがGGUF形式)、GGUF版2(ModelとText EncoderとVAEがGGUF形式)の3パターンがあり、おすすめはGGUF版1です(VRAM 8GBの場合はGGUF版2)。

 なお、Negative Promptは元となったワークフローの内容をそのまま拝借しています。

(Last update 12-5-2025)

FP8版

GGUF版1(Text EncoderとVAEがGGUF形式)
GGUF版2(ModelとText EncoderとVAEがGGUF形式)


▼ 3-4. 実行例(画像の生成)

 設置したモデル(Model、Text Encoder、VAE)に合わせて設定を行い、解像度(初期値は1024x1024)を好みで書き換えて、プロンプトを入力して「Run」をクリックすると生成が始まります。

実行例(解像度を1280x640に変更)

 実行例は本記事の表紙画像です。GGUF版1を使用して、解像度は1280x640、Seedは271429449280069です。なぜか「text rendering」を「text rending」としてしまう傾向があり、正しく記述できませんでした。

A super-deformed Japanese-anime style illustration of a chibi schoolgirl proudly pointing at a chalkboard where the English text is written in noticeably large, bold chalk letters: “Ovis-Image is a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints.” The letters appear big and clear, surrounded by her playful doodles. She has round sparkling amber eyes, a tiny cheerful mouth, and fluffy chestnut hair with a small ribbon clip. Wearing a simplified sailor uniform with an oversized red neckerchief and short navy skirt, she stands in an energetic pose with one hand on her hip and the other pointing confidently at the enlarged chalk text. Background: a cozy, lightly stylized classroom with soft sunlight, simple wooden desks, gentle chalk dust, pastel tones, and clean deformed anime linework for an overall cute, warm atmosphere.



■ 4. おまけ

▼ 4-1. Ovis-Imageの生成画像

 Ovis-Imageで生成した画像を掲載します。プロンプトはZ-Image-Turboの記事と同じものを使用します。出力時の解像度も同じく1536x864としています。

 最近の他のアーキテクチャと比べると見劣りする内容ですが(特に、Z-Image-Turboと比べるとかなり)、初出でありながら様々なスタイルのイメージを生成する能力が確実に備わっていると思います。

A heartwarming puppet-theater style illustration with felt-fabric textures, soft woolly details and visible marionette strings, chibi proportions: a delighted purple-haired girl in a flowing red felt hooded cape and white cotton dress sits happily on a giant fluffy cloud made of stuffed cotton, floating above a dreamy flower meadow stage at twilight. She gently hugs and holds hands with an adorable oversized plush baby red panda that has sparkling button big eyes, embroidered rosy cheeks, tiny stitched paws and super-soft fur texture, wearing a tiny red satin ribbon bow that flutters in the breeze. In the foreground, endless fields of glowing fabric cherry blossoms and felt dandelions sway softly on wires, clusters of friendly hand-puppet forest animals like bunnies and fawns gather curiously around them, tiny LED fireflies dance like lantern props, delicate silk petals drift in the warm evening air, all lit by gentle paper-moon light and painted sunset backdrop for a perfectly balanced, magical, heart-melting wholesome puppet-show scene.
A timid 19-year-old girl drawn in a warm and fluffy modern shoujo manga style for today’s readers, with extremely long flowing pale-lavender hair styled in soft loose pigtails tied low with delicate ribbons, airy bangs sweetly hiding one eye, large sparkling deep-crimson eyes brimming with quiet wonder and gazing gently off to the side, fair porcelain skin with a faint rosy blush, wearing a classic white long-sleeve sailor uniform with navy collar, bright red scarf neatly tied, and a gently swaying navy pleated skirt, white knee-high socks with tiny ribbon charms, and shiny chestnut loafers. She steps shyly toward the viewer, filling the frame, one foot forward, body leaning in cutely, both hands carefully offering a small glass jar full of twinkling golden fireflies toward the camera with lovely foreshortening, the softest happiest smile ever. The scene is wrapped in deeper twilight darkness so the fireflies and jar glow even brighter, countless golden fireflies drifting like tiny lanterns in the tall swaying emerald grass, a velvet star-filled sky with gentle pastel shooting stars. Palette of rich midnight indigo, soft lavender, and warm golden highlights, delicate halftone shading, abundant tiny sparkles, and a tender glowing rim light that beautifully outlines her hair and clothes. High-resolution modern shoujo illustration, soft dreamy bokeh, vibrant yet heart-warming romantic mood, the kind of art that’s super popular in current kawaii communities.
Horizontal wide manga page layout with 3 equal panels side by side. Depict a cute chibi girl with long chestnut hair tied with navy ribbon wearing navy blue school swimsuit at bright indoor swimming pool. Left panel: inside the pool facility hallway or locker area doing side stretch warm-up with text "Stretch♪" and only walls and doors visible in background, middle panel: standing on starting block giving big peace sign and wink with text "Ready!" and pool water and lane ropes behind her, right panel: mid-air diving forward into the pool with huge excited smile and text "Dive!!" showing sparkling blue water below and splash starting. No speech balloons, dynamic motion lines and droplets, thin strong black outlines, some effects breaking borders. Style: soft warm deformed chibi style, full-color manga with delicate pale watercolor texture, gentle healing atmosphere, bright indoor lighting, high saturation pastel colors, fluffy round character design, very cute and heartwarming, soft watercolor wash effect, delicate shojo-manga detail.
Documentary-style portrait of a Japanese elementary school-aged girl sitting on a gently sloping lush green meadow with her knees drawn up to her chest and arms wrapped around her legs in a gentle embrace, photographed from a diagonal front angle at close range so her figure fills most of the frame. She looks toward the camera with a soft, warm smile, her face prominently visible and occupying significant space in the composition. Medium-length brunette hair in twin braids frames her face, light gray eyes sparkling. Wearing pale blue lolita-style dress with elaborate lace and ribbon details over matching culotte-style short pants in the same pale blue fabric, matching headpiece, white ankle socks, and black patent shoes with feet positioned close together. Her seated posture is compact and natural, knees held close to chest, arms gently circling her shins, creating an intimate and peaceful composition. Refreshing natural daylight, vast open grassland softly blurred behind, shallow depth of field, crisp 85mm lens perspective, bright and clear close-up from diagonal front view emphasizing facial expression and relaxed seated pose.
A cozy traditional Japanese izakaya interior at night, warm lantern lighting, wooden table in the foreground with various izakaya dishes like yakitori skewers, edamame, sashimi platter, tempura, grilled fish, several glasses of cold draft beer with foam overflowing, chopsticks and small plates scattered naturally, across the table sits a 50-year-old Japanese salaryman in a dark business suit, white shirt, slightly loosened tie, short graying hair, gentle tired smile, relaxed posture leaning back a little, realistic detailed photography style, cinematic warm atmosphere, sharp focus, high resolution
Photorealistic wide-angle view of the iconic Kaminarimon Gate in Asakusa, Tokyo, on a bright sunny day, massive red chochin lantern with "雷門" lettering hanging prominently, fierce wind and thunder god statues on both sides, dense crowd of excited tourists from all over the world taking selfies, raising smartphones and cameras, groups wearing yukata and kimono, street vendors with colorful carts, traditional rickshaws passing by, red temple gate framing the bustling Nakamise-dori shopping street in the background, vivid colors, sharp details, lively atmosphere, cinematic composition, 8k
Ultra-detailed 3D MMORPG combat screenshot in medieval fantasy style, dynamic battlefield at dusk with dramatic orange-purple sky and volumetric god rays piercing through thick battle smoke, foreground shows a heavily armored human warrior in massive plate armor with glowing blue runes on the shoulder plates swinging a gigantic two-handed flaming greatsword mid-attack creating bright fire trails and sparks, massive impact shockwave distorting the air, directly in front of him a huge dark orc warlord in spiked black armor blocking with a blood-stained tower shield while being pushed back, cracked ground and flying debris around them, in the midground a female elven mage in flowing emerald robes casting a massive swirling ice storm spell with glowing magic circles and frost particles exploding outward, behind her a rogue in dark leather dual-wielding poisoned daggers leaping from the shadows, background filled with chaotic player vs player combat, arrows raining down, fireballs exploding, warriors clashing, destroyed siege weapons burning, detailed UI elements with health bars, mana bars, floating damage numbers in red and gold, party frames on the left, minimap top right, glowing skill icons on hotbar with cooldown timers, high fantasy atmosphere, ultra realistic textures, cinematic depth of field, 8k resolution, unreal engine rendering, extremely detailed environment and character models, epic large-scale battle vibe
Retro MS-DOS era visual novel scene, EGA 16-color graphics, pixel art style, cute anime heroine with long hair and school uniform standing in classroom during sunset, warm orange light through window, classic dialogue box at bottom with cyan background and white pixel font text saying "Will you go home with me today?", heartfelt innocent atmosphere, 1990s Japanese galge aesthetic, crisp dithering, nostalgic CRT monitor glow, high detail
A highly detailed whimsical anthropomorphic vegetable garden scene at golden hour, ultra-tall muscular carrot with thick green leafy hair and powerful root legs standing proudly in the soil like a giant guardian, mischievous plump red tomato with skinny dangling arms and cheeky grin lazily leaning against a wooden fence, determined brown potato with freshly sprouted strong human-like arms and legs energetically swinging a tiny gardening hoe, slender green cucumber with long elegant arms thoughtfully climbing a wooden trellis while looking into the distance, various expressive vegetables peeking from the lush foliage, vibrant colorful raised garden beds, dewdrops on leaves, soft warm golden sunlight filtering through, long playful shadows, rich earthy tones mixed with bright vegetable colors, cute storybook illustration style with cinematic lighting, highly detailed, sharp focus, whimsical and heartwarming atmosphere



■ 5. その他

 私が書いた他の記事は、メニューよりたどってください。

 ComfyUIに限定した記事の一覧もあります。

 記事に関することで何かありましたら、Xの@riddi0908までお願いします。

いいなと思ったら応援しよう!