SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

What I Learned from Creating 300,000 Images with AI


I started my journey with generative AI around August 2022. It began with Midjourney (early version), then I moved to NAI, and after Midjourney updated and I played around with Niji-Journey a bit, I bought an RTX 4090 and shifted my environment to a local Stable Diffusion setup.

From August 2022 to February 2023, that's about one year and 200 days (approximately 565 days).
The total number of generated images is over 320,000. About 30,000 were created using online generation services (Midjourney + NAI). Since moving to a local environment, I have generated 290,000 images.

The early Midjourney version period. Around August 2022. Back then, I was just going 'ooh-aah' over this.
The NAI period, where I was going 'ooh-aah' because 'I can make anime characters!'
A manga (unpublished) created with Midjourney + Niji + NAI, themed around the AI's unique ability to cross art styles (dimensions).
The Stable Diffusion period. I'm still going 'ooh-aah' every day.

It all started because I thought, 'AI images are creepy, so they might be useful for horror expressions,' and I tried out Midjourney. As everyone knows, it reached a level of practical utility in the short span of just a year and a half.


The Narrative of a Work

When looking at AI, there is a large gap between the creator and the viewer.
The creator uploads with great excitement, saying, 'This is amazing, it looks like a pro did it, it looks like a photo,' but the reaction from the viewer side is generally lukewarm.
The creator wonders, 'Why?'
Not knowing the answer, they upload more or change direction, but the response remains poor. So, they try making it into a video or adding music.
Even then, the reaction doesn't change much. Why?
The fundamental cause, as I see it, is what is most lacking in AI generation.
The answer I've arrived at is that there is 'no narrative.'

'No narrative' means, well, 'no story,' and also
'no direction,' 'no intent,' 'no sense of purpose,'
and 'no message.'

Broadly speaking, what is a picture? It is a 'means of communication,' a technique for conveying the author's intent and worldview to others.
From professional illustrations to amateur hand-drawn sketches, everything is a message. It is something constructed by turning meaning and intent into a two-dimensional form.

Then what about AI-generated art?
The motivation for the picture is completely absent. The creator inputs a 'vague prompt,' selects the best-looking result from the AI's output, and publishes it.
The reason the motivation is absent is that current AI technology does not allow for that level of complete control.

To convey motivation through a picture, a massive number of elements are required.
First, you decide on the message you want to convey, find the optimal composition, and then determine the necessary colors, character placement and poses, gaze, hand positions, costumes, lighting, props, and background.
When humans draw, regardless of the difference in skill, everyone does this. While drawing, they repeat the tuning process, thinking, 'Wouldn't it convey more if I did it this way?'
A picture is built upon precise code.

Since that is impossible in AI generation (at present), one outputs hundreds of images and selects the one closest to 'what I probably wanted.'
During that selection process, the AI creator is actually searching for a 'narrative.' Even while doing the fruitless task of checking for missing limbs, they are intuitively selecting images that feel 'good' due to slight differences in character placement or gaze. They are finding the tiny narratives within the random AI output, like searching for gold dust. But even that is a narrative as small as gold dust. It is too small.

A visual without a message is barren. No matter how beautiful or high-definition it may be, it is boring, and the eyes just slide over the image.
What stops that gaze is the narrative of the picture. The meaning and intent contained within the picture are what capture the viewer's interest.

You cannot generate a narrative in a work. This is the missing piece of current AI illustration.

Inoffensive images


The author's narrative

What is the major current difference between human illustrators, who are physical laborers, and AI creators, who are futuristic exploitative creators?
It is the lack of a narrative for the work, as well as the lack of a narrative for the artist.
Incredibly, an illustrator and their art style are inseparable.
This is surprising to AI creators. AI creators, who have achieved the democratization of painting techniques, can share every art style and use all visuals as a paintbrush.
However, that was the trap.
In the world of illustration, the art and the artist are one. Behind the art is the artist, and the artist has their art as their history.

However, for AI generators, that sense of unity does not exist.
An AI artist with the strongest, invincible power, capable of using magic of all attributes, mastering the special moves of all jobs, and possessing all items,
The strongest, me, losing to a human's hand-drawn picture!?
Unfortunately, at present, they cannot win.
People actually care about 'who is drawing it'.

A person's art style changes as slowly as a tortoise.
Changes in art style are only possible through long years and tireless persistence.
It happens that a manga artist's art looks completely different at the end of a long serialization compared to the beginning, but in reality, it only changes to 'that extent'.
Unless a person tries very hard to force a change, the range of change is very small.
In contrast, AI generators lack that sense of unity.
The AI used, CKP, LORA, prompts—because the technology is democratized, anyone can use it. Anyone is replaceable. Among the countless AI generators on the internet, how many have their art and themselves as the artist connected?
To gain recognition as an artist, one must abandon the possibilities of AI and limit their art style or theme.
To limit the art style or theme, one needs a motive and must find oneself as an artist.
One must make the work and the artist inseparable.

There is no narrative to connect the work and the artist.
However, this is a deficiency that can be solved in the future, as humans with a certain kind of perversion will start engaging in creative activities combined with AI.

An unobtrusive image that does not interfere with the text


What is 'creation'?

That's a big subject.
Actually, I am a hand-drawing illustrator. This is not a lie; I have reported it as such to the tax office and pay taxes. If you think it's a lie, please ask the tax office.
As a result of someone like me getting obsessed with AI generation for about a year and a half, I was forced to think about what it means to draw and create.
To put it briefly,
I have come to the conclusion that it means 'to possess the world'.

Everything in life is borrowed. The Earth, countries, houses are rented, things are property, money is banknotes, and even life is a borrowed thing that will eventually be asked to be returned.
The whole world belongs to others, even oneself is uncertain, and in such a situation,
I can say for sure that 'only creative works are absolutely my own'.
From paintings to scribbles in textbooks, if I draw it myself, it's my own.
No one can deny that.
In the vast universe, a tiny, 0.00000000000000~01% area becomes my own.
The pleasure of a part of the world becoming one's own.
I have come to think that this is the pleasure of creation.

(There is spatial occupation, but created characters, dramas, and worldviews also become one's own. There is also the pleasure of being able to become a god in that place.)

I made about 300,000 images with AI generation, but I didn't feel that pleasure.
The moment of 'this is mine alone' never came.
Hand-drawn illustrations are 100% me. I control everything. If it's bad, it's my fault. If it's good, it just happened to go well. That's the kind of world it is.
However, in AI generation, the 'me' part is probably less than 20%. It might even be in the single digits.
Even after generating 300,000 times, I didn't have the feeling that I had acquired a part of the world.


The picture between the sentences

Conclusion

So, which one are you? Are you for it or against it?
The answer is neither. I will just keep doing what I want to do. I intend to continue until I find something or until I get completely bored.
Plus, there is the realization that 'I won't get any better at drawing than this,' and there is also the aspect of continuing it as a skill to survive in old age.
However, it seems I am a pervert who 'gets pleasure when I press a button and a good picture comes out,' so I clicked away and made 300,000 images because it felt good.
Pleasure is the source of action.


Technological innovation in generative AI will continue in the future.
It is the human role to treat it as a 'paintbrush'.

いいなと思ったら応援しよう!