[For AI Beginners] No GPU Required! Turn Pose Illustrations into Realistic Photos with ChatGPT! How to Create Your Ideal Character by Freely Changing Clothes, Face, Background, and Body Type | Compared with Gemini | SFW

In this article, I will introduce how to turn pose illustrations into realistic photos using ChatGPT.
Are you interested in AI image generation but don't have a high-performance PC with a GPU?
Even if that is the case, you will be fine.
As long as you have an environment where you can use ChatGPT, you can create realistic images based on pose illustrations and build your ideal character from your PC or smartphone.
Furthermore, you can freely change the face, clothing, background, and body type of the character you have created.
In this article, I will explain the steps using actual examples. I also tried it with Gemini, so I will share those results as well.

Click here for the video explanation
Why use pose illustrations?
In AI image generation, you usually specify the pose using a prompt (instruction text).
However, when you actually use it,
the pose doesn't turn out as expected
the position of the limbs is wrong
the composition changes significantly
These things happen often.
In a sense, there is a 'gacha' element involved, and sometimes you need to repeat the generation many times.
Therefore, to reduce that gacha element as much as possible, I tried a method of preparing a pose illustration in advance and creating a realistic image by using that pose as a reference.
Preparing a pose illustration
This time, I used pose materials from the free-to-use 'Random Pose Maker'.
I have tried various poses, but this time I will explain using a breakdancing pose as an example.

Trying to make it realistic with ChatGPT
First, I uploaded the pose illustration to ChatGPT and entered the following prompt.
Using the pose in this image, create a photo-quality image of a young Japanese woman breakdancing, in 16:9 aspect ratio, smiling, with a white background

The generated image accurately reflected the characteristics of the pose.

Although it is not the exact same pose, it has a dynamic movement typical of breakdancing, and the result was close enough to the image I had in mind.
Also, despite specifying a white background, there was an interesting phenomenon where an image that looked like a transparent background was generated.
Changing the outfit
Next, I will try changing the outfit.
I added the following prompt.
In a white tank top and white mini skirt
Then, I was able to change it to a refreshing summer outfit while keeping the same pose.

The background also became a clean white background.
Changing the face
Furthermore, I prepared another person's image and
replaced it with the face of the woman in this image
I changed the face with that instruction.


However, since the smile disappeared at this point,
with a smile
I gave an additional instruction.
As a result, I was able to maintain the smile while changing to the desired face.

Changing body type too

You can also do it all at once
Prepare a pose illustration and a face photo,
Using the pose in this image, create a photo-quality image of a young Japanese woman boxing, in 16:9, with a smile, on a white background
I gave the instruction.


Realistic conversion with ChatGPT (Drop the pose illustration and face into ChatGPT and prompt)

My thoughts after actually trying it
What I felt after trying this time is that by using pose illustrations, it becomes easier to create a composition close to the ideal from the start.
Especially,
Dancing
Sports
Action scenes
Model poses
etc. I felt that this is highly effective for poses that are difficult to express with prompts alone.
It is easier to get the intended results than by specifying poses only with prompts.
Summary
By using pose illustrations, you can create your ideal character relatively easily with ChatGPT.
Moreover, you don't need an expensive computer equipped with a GPU.
Anyone can try this as long as they have a computer or smartphone that can access ChatGPT.
"I can't get the pose I want"
If you have that kind of trouble, please try image generation using pose illustrations at least once.
Creating your ideal character might become much easier.
I also tested it with Gemini under the same conditions
This time, I also tried turning pose illustrations into realistic photos using Google Gemini under the same conditions as ChatGPT.
I used the same pose illustration images and compared them with the same prompts.
As a result, in this verification, I got the impression that ChatGPT has higher pose reproducibility.
Gemini was also able to generate images, but there were relatively many cases where the image would break down when the pose became complex.
For example,
it looks like there are three legs
The arm is unnaturally piercing through the neck (I self-censored it because it was creepy)
The shoe on one foot disappears, leaving it barefoot




These kinds of phenomena occurred.
Of course, these are the results of this verification, and there is a possibility that they will be improved by future model updates.
However, at least within the scope of what I tried this time, regarding the purpose of creating realistic images by referencing pose illustrations, ChatGPT produced consistently higher accuracy.
