SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The Blurring Line Between 'Real and Illustration' Melted by AI Image Generation: A Meta-World Created with ChatGPT

I believe the greatest strength of the image generation AI currently equipped in ChatGPT lies in its prompt precision. Whether it's people, backgrounds, composition, clothing, color schemes, poses, art styles, lighting, or text, even if you pack in a long list of numerous elements, it generally reproduces them for you.
In addition, it is capable of generating high-quality, realistic images as well as anime-style illustrations.

Furthermore, by combining these features, you can create images that
look like anime-style characters fused into a realistic world.

For example, it looks something like the example below.

This image has a composition where an anime-style character appears to be jumping out of a realistic smartphone, doesn't it?

In this way, being able to naturally fuse realism with anime-style illustrations is what I consider to be the hidden brilliance of image generation with ChatGPT.

So, this time, I will introduce several actual examples of such images and touch upon why ChatGPT is able to generate images like these.




Introduction of actual generated image examples

This time, I have created the following 'chibi character in a lab coat' using ChatGPT.
I will introduce examples where this is used as a base and combined with realistic images.

A scene where she is puzzled in front of a suspicious liquid, wondering, 'Why didn't it explode?'


Standing on a notebook on a desk


Sitting on the rim of a coffee cup


Hugging a teddy bear


Eyes sparkling in front of a pudding


In this way, by instructing the realistic background and illustration-style character as prompts, I think images where they coexist well are being generated.

While the high precision of ChatGPT's prompts (cross-attention) that makes this distinction possible is amazing, I also think that the seamless feel between reality and illustration is actually a point worth noting.


Normally, when creating an image like this, you would need to composite an illustrated character onto a realistic image.
And if you composite them, the illustration will end up looking out of place if you don't process it.

This is because if left as is, there will be a sense of discomfort in the light and shadows, and there will be no sense of unity in contrast or color.
That is why you need to put in the effort to blend the two well.

Nevertheless, in this image, there is no such discomfort, and it feels like there is a sense of unity.

The way the color and contrast are applied, as well as how the shadows fall and the light hits, all appear natural.
Furthermore, the fact that it reflects the weight of the character based on the indentation of the sofa also enhances the seamless feel.


Thinking about it that way, I find it interesting that it can be generated by fusing them so naturally.


Why is it possible to generate a seamless image?

There seem to be two main reasons why ChatGPT can generate such images.

1. Simultaneous global optimization derived from diffusion models

The current mainstream of AI image generation uses 'diffusion models' to gradually generate images from noise according to prompts, and in doing so, it generates while optimizing the entire image simultaneously.

Each element is depicted while interacting with the others, and since it won't converge if consistency isn't maintained along the way, in other words, the converged image is one where the overall consistency is maintained.

Although it has not been publicly disclosed what kind of model ChatGPT uses, as far as it is estimated from literature, it is said to be a hybrid of 'diffusion model + autoregressive model'.

Therefore, even if different textures like reality and illustration coexist, by generating while optimizing the whole, it is thought that the contrast and color are adjusted so that the image as a whole has a sense of unity.


2. Acquisition of 'physical realism' through learning

Generative AI does not understand physical phenomena, but it can reproduce the appearance of physical phenomena from a large amount of training data.

For example, it can draw how light naturally hits materials with different textures, or how shadows naturally fall.

Therefore, even if reality and illustration coexist, I think it creates a natural and fused impression because it enables 'plausible and realistic' shadows and lighting that follow each material.


Generative AI surely has no 'wall of meaning'

Thinking about other factors, I feel that 'AI does not semantically distinguish between real and illustration' is one of them.

For us humans, photorealistic images are continuous with reality and signify the 'real'.
On the other hand, illustrations are clearly creative works and carry a 'fictional' meaning.

In other words, humans have a clear conscious barrier between the real and the fictional, so we hesitate to blend them.
When we seek meaning, such as asking 'What is going on in this part?', a sense of discomfort inevitably arises.


However, generative AI does not have this consciousness, and real and illustration are merely differences in expression style to it.

Just as Impressionism, Rococo, and modern art have different expression styles, real and illustration are also just differences in style.

Therefore, generative AI can draw both without hesitation or discomfort is an aspect of it, I think.


Summary

This time, I have covered the blending of realism and illustration.
This meta-feeling found in the gap between reality and fiction is interesting and gives off a futuristic vibe, doesn't it?
It's also relatively easy in terms of prompting, as you just need to give separate instructions for the background and the character, so please try it if you'd like.

If you found this helpful, I'd be happy if you could 'like' it.
See you next time.


〇Related Links

いいなと思ったら応援しよう!

Alpaka もし気に入っていただけたら、チップで応援して頂けると嬉しいです。その気持ちが次の創作の大きな励みになります!

この記事が参加している募集