SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

I want everyone to know about the greatness of the new image generation AI model called Flux.1. I'll even show you how to try it out!

* All generated images in this article are unedited and were created using Flux.1 Schnell

So, the story is that a new image generation model called Flux.1 has been released, and it's quite interesting!

For example, you can easily create images like this

However!
You probably won't be able to create the kind of 'beautiful girl' AI illustrations you might be satisfied with!

But it's interesting! Why not try generating something other than beautiful girls for a change?

If you're going for something realistic, it might turn out pretty well

What is Flux.1? How is it different from Stable Diffusion?

I don't really know, but it's probably a relative of Stable Diffusion.
As of August 7, 2024, it cannot be used in webui.

If you want to know more, you should read articles like this one!

It can currently generate top-tier images, especially for photorealistic photos and art-style images.

And like Stable Diffusion, you can download the model and run it on your own computer!(*)

In other words, to put it very loosely, it's like being able to use Midjourney for free. Is that too loose?

* Some computers may not be able to run it

It broke...

What's so great about it?

It's quite beautiful, and it will respond to your wild requests to a certain extent! Probably!

The official announcement page lists the following features!

The highest performance in the image generation AI world as of July 2024

Midjourney-v6.0, Stable Diffusion 3 Ultra, DALL-E 3, and the lesser-known but super powerful Ideogram—the official site claims it can generate images that are equal to or better than these.

This is likely a graph of scores from a showdown where people were shown images generated with the same prompt and asked,
'Which model's generated image do you think is better?'
A comparison by field. A graph showing that it is the best in everything except for Ideogram's text reproduction.

The 2 images above are cited from https://blackforestlabs.ai/announcing-black-forest-labs/

Actually, I haven't done a serious comparison, so I don't know for sure.
But I feel it has enough power to make you think that might be the case.

Bow down!

It seems like there are three types, but what are these?

There are three types of FLUX.1: pro, dev, and schnell!

pro is for professionals, so you can't use it without paying! The model is also not public!

dev is the most standard one. The
model is public, and you can use it virtually without limits just for generating images!

schnell has slightly lower quality but is fast at generating! (About 5 times faster than dev).
The model is also public.
If you are familiar with Stable Diffusion, you can think of it as something like LCM or Turbo/Lightning.
This one can be used with almost no restrictions, even more so than dev!

But it must be expensive, right?

It's free to try a little on the web or use on your own computer!
Basically, anyway!

The lion is happy too.

Let's try it out!

As I said earlier, you can try it on the web or on your own computer.

Try it on your own computer!
The PC specs required for that are 16GB or more of memory and a graphics card with 12GB or more of VRAM!

For those who don't know if their PC meets this!
...I think it's a bit difficult, so you should probably look at trying it on the web instead.

Actually, I think you can run it with specs lower than this, but it takes time to generate, so I don't really recommend it.

Also, I think it's better to make your prompts natural English sentences!
Let's ask ChatGPT!

Try it on your own computer

I think writing explanations is a hassle and reading them is a hassle too, so I'll keep it as short as possible!

We will use ComfyUI for generation! How to install it... please look it up!
From here on, I assume the installation of ComfyUI is finished!
I'm not slacking off!

First, download the model file. There are
dev and schnell, but schnell is easier, so I will explain using that.

Open the following page and download the file from the download link near the center of the screen.
(By the way, this is the fp8 version, which has slightly lower performance but half the file size.)

Once the download is finished, please move flux1-schnell-fp8.safetensors into the models folder inside the folder where ComfyUI is located, and then into the checkpoints folder.

Then, launch ComfyUI, download the workflow file below, and load it.

You will then see this incredibly easy-to-understand workflow.

There is no need to touch anything other than the prompt and the image size!

Don't see the koala in the bottom right? Press the Queue button!
Depending on your computer's performance, koala sushi will pop up in a few to several dozen seconds.

Now that you're all set, go ahead and change the prompt to whatever you like and generate!
It seems you can stably generate images in various aspect ratios as long as the image size is between 0.1 megapixels (about 320x320) and 2 megapixels (about 1920x1080)!

The penguin is also giving a thumbs up.

Try it on the web

I don't have a computer like that!
Don't worry. There are ways to use it online too!

The easiest way is the official Hugging Face Space.

Just enter a prompt and wait a little, and an image should be generated! It's completely free!

It's bean sprouts, this.

It's available in many other places too, so try looking for them (you can also do it on Civitai).

Our Discord server also runs a generation bot, so you can try it out there.
(However, you might be forced to see nonsensical images, which could be unpleasant.)

Includes an automatic prompt enhancer.

Licenses... or whatever they're called, they're probably complicated, right!

Not really!
dev
and schnell both have absolutely no restrictions on the images you generate, I think.

To be precise,
dev is FLUX.1 [dev] Non-Commercial License , and it has restrictions on commercial use of the model itself, and restrictions against using the model's output to train competing models (like Stable Diffusion), etc.
However, commercial use here refers to things like image generation services using the model, and it is explicitly stated that you can use the output images themselves for various purposes, including commercial use, so there's no need to worry!

schnell is Apache License 2.0, and to put it very loosely, it's basically like having no restrictions at all!

In other words! You don't need to worry about anything when playing around with image generation!

Yay!

Summary

Did you understand?

Whether you understood or not, I would be happy if you could give it a try for now.

This is the end
or so you might think, but there is a main part after this, so if you have some free time, please take a look!

Aside

The company that created this FLUX.1
is apparently planning to release a text-to-video generation model next!
(Whether it will be released, or whether we will be able to run it if it is, is another matter)

The announcement pageis cool!

Below is just a large collection of images generated by Flux.1

This is the main part
I wrote everything up to this point just to create a justification for posting a large number of images.

What lies ahead from here
is a collection of cool images generated by someone in our Discord server, including myself, using the generation bot.
(The names of the creators have been omitted out of consideration for privacy.)

I hope you can experience the wonder of Flux.1 by looking at these.
Thank you.

Note: Some images may be a bit vulgar.

A cat wearing sunglasses with its hands on the steering wheel of a car, CCTV video quality
An owl pole dancing on a stage
An owl showing off its miraculous biceps
Slippers made of human noses
Raw steak meat with arms and legs running while kicking up dust, motion blur
A potato that looks just like Dwayne Johnson
A carrot shaped like a person sitting with their knees hugged, photorealistic
A painting of the Mona Lisa speeding down a highway at night, drifting, cornering, motion blur, photograph
A Google Street View image capturing a man dressed as a rhinoceros beetle
A Clockwork Orange
A person with an alarm clock for a head and face
A Showa-era family portrait of a family whose heads have become Poke Balls. Monochrome. There is a father, mother, daughter, two sons, a grandmother, and a grandfather.
A creature that is a fusion of a surveillance camera and a crow
A photo of a skeleton drinking beer from a mug, but the beer is spilling everywhere through the gaps in its bones
A CCTV surveillance camera photo capturing a crazed clown koala wandering the hallway of an office late at night. Grainy image
A tilt-shift photo taken with a toy camera, steampunk-style cityscape
A wide-shot tilt-shift photo taken with a toy camera, landscape of an abandoned factory
Photo realistic, Castlevania, 8K ultra-high resolution, high-pass filter. A large number of translucent Nicolas Cage faces in the background
A trophy giving the middle finger
A tank with legs wearing fishnet stockings
A muscular eggplant wearing a diaper
Express the scene of poop being run over and squashed by a car in a 4-panel manga
An X-ray of 💩
Before and after image, with a high-class dinner and the word "BEFORE" on the left, and poop and the word "AFTER" on the right
An old man buried in a rice paddy with only his head sticking out, with grass growing from his head

These are just a few examples, so if you want to see more, come join the Discord server!


いいなと思ったら応援しよう!