SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Kakudai V1 EX with ollama

As I mentioned in my previous article, apart from the upscaler I posted last time, I am a big fan of Kakudai V1, which was developed by the Japanese AI development group 'Maverick'.

However, four months have passed since the original Kakudai V1 was released, and the world of generative AI has continued to evolve at a rapid pace during that time, so there are parts that can be improved by incorporating those latest technologies.

In particular, the original version used an OpenAI API key to access GPT-4 for the prompt generation part that complements faithful image upscaling, so it was a system that incurred API usage fees every time it was used.

Naturally, that part can be substituted with WD14Tagger, which is familiar in A1111, but it goes without saying that its image analysis and text generation capabilities are far inferior to GPT-4.

However, with the development of various extensions that allow ComfyUI to access LLMs built in a local environment using the ollama application, an environment has been prepared where this part can be completely self-contained locally.

This time, I have created and am distributing an improved version of the original Kakudai V1 where the part that accessed GPT-4 to analyze the original image and generate prompts has been replaced with ollama.

As with the last time, the underlying workflow itself is distributed for free, and since I only made minor changes to it, I will distribute it for free as well.

To be honest, I would like to develop a node that I feel comfortable charging for soon.

However, not only the prompt generation part but also the model and sampler parts reflect various advancements made over the past four months, so the settings are different from the original version.

いいなと思ったら応援しよう!