Is the ChatGPT Turbo updated today a savior for image generation? - Blog 2023/11/07
On the 6th (US time), OpenAI held its developer conference, "OpenAI DevDay", in San Francisco. The keynote speech is available on YouTube.
The next-generation model "GPT-4 Turbo" was announced, and a massive amount of posts were flowing on social media since early morning.
GPT-4 Turbo can now handle up to 128,000 tokens, which is equivalent to over 300 pages in a typical book.
As for the knowledge base of GPT-4 Turbo, it has information up to April 2023 (GPT-4 was up to September 2021). However, for news information, Bing runs in the background, so it searches and displays the latest information. It can also display today's news.
Since "OpenAI DevDay" is an event for developers, I will leave the details to IT news media and focus on topics that affect creative work.
Reference:
Executing prompt generation and image generation simultaneously
When you open ChatGPT, an update pop-up is displayed.


Just to be sure, I will ask about the specifications.
Me:
For example, can I upload an image, generate a prompt to reproduce that image, and then proceed to execute that prompt to generate an image?
ChatGPT:
Yes, that is possible. First, I will analyze the uploaded image and create a detailed prompt describing its content. Next, I can use that prompt to generate a new image using DALL·E. If you upload an image, I will start that process.

Now, let's try it out.
First, I dragged an image file generated in Midjourney into the ChatGPT window to upload it, had it generate a prompt to reproduce it, and simultaneously generated an image.
Until now, after generating a prompt, I had to switch to DALL·E 3 in the menu to generate an image, but it is now possible to execute them simultaneously (what was called All Tools).
Playback time: 25 seconds
Next, I will assign roles to ChatGPT and have two experts work together.
Upload a Midjourney generated image. Use this image as a reference.
Expert A will devise a prompt and generate an image at the same time.
Expert B will point out areas for improvement in the generated image.
Expert A will incorporate the feedback into the prompt and generate the image again.
Although the interaction is in English, an image is generated based on the prompt Expert A thought of, Expert B evaluates that image, and Expert A revises the prompt considering the improvements to generate the image again.
Playback time: 20 seconds
This is... It has the impact to change the image generation workflow.
The time spent thinking about prompts will be simplified, and roles like art director will become important.
Have it read a project proposal and generate images based on the content.
I will upload the story draft (text file) for "Creating a Music Video with Generative AI Live" to ChatGPT and have it generate images based on what is written. First, I will confirm whether it understood the content of the story draft as follows.
Please understand the content of Act 1, Act 2, and Act 3 of the "MV 'Never Your Friend (tentative)' Story Draft" written in this text file. You do not need to write it out.
If you understand, please answer "I understand".

I will have it generate images up to Act 3 based on the written content.
As an expert in generative AI for images, please generate images for the content of Act 1, Act 2, and Act 3 respectively. Please also write the prompts used in both English and Japanese.
It generated all the images.
Until now, the number that could be output at once was two, but six were generated.
Since the specifications change every few weeks, making a paper manual is tough...

As a test, I will try generating the Act 1 prompt created by ChatGPT using Midjourney.
A picturesque coastal town where childhood bonds shine. 18-year-old Meg, with a passion for photography, captures the everyday life of her hometown with her camera. Her neighbor and childhood friend Jake dreams of becoming a professional surfer. Their friendship is strong as they share their dreams and spend their days together. The scene shows a sunny coastal town with Meg taking pictures and Jake surfing in the background, reflecting their strong friendship and shared dreams.
18-year-old Meg, who loves photography, captures the everyday life of her hometown with her camera. Her neighbor and childhood friend Jake dreams of becoming a professional surfer. Their friendship is strong, and they spend their days sharing their dreams. This scene shows a sunny coastal town where Meg is taking pictures and Jake is surfing in the background.
The quality of the generated images is wonderful, but DALL·E 3 is the one that is faithful to the prompt. Midjourney completely fails to express the scene of "a sunny coastal town where Meg is taking pictures and Jake is surfing in the background."
Its interpretation of prompts is inferior to DALL·E 3 and Adobe Firefly, so it is difficult to get it close to the intended image.

Since DALL·E 3 faithfully expresses the prompt, it can reduce trial and error compared to Midjourney, but photorealistic expressions (realistic expressions indistinguishable from photographs) are often blocked.
It displays a message like the following.
I was unable to generate images based on the latest prompt due to content policy restrictions.
I was unable to generate images based on the latest prompt due to content policy restrictions.
As shown below, there is no perfect AI image generation service yet (each has its pros and cons), so you have no choice but to combine multiple generative AIs.
Midjourney:
Excellent at photorealistic expressions indistinguishable from photos, but often fails to reflect the content of the promptDALL·E 3:
Faithfully expresses the content of the prompt, but photorealistic expressions are often blocked
Place the generated images into an A4 document and export as a PDF
The feature integration formerly known as ALL Tools is now available, so I tried having it automatically process multiple tasks. The tasks are as follows:
Generate two images based on a theme
Place the generated images into an A4 size document
Export as a PDF
I have turned the work record into a 16-second video, so please take a look.
Playback time: 16 seconds
If you try to execute all processes at once, it may ignore part of the prompt, so it works better if you split it into two tasks.
Image generation task:
As an expert in generative AI, please generate images under the following conditions.
#Conditions
・Generate two photos (A and B) of young Japanese men and women in Shibuya, Tokyo
・Set the aspect ratio to 16:9

Task for image placement and PDF creation:
Please place the generated images (A and B) into a PDF under the following conditions.
#Conditions
・The PDF should be A4 size in portrait orientation, and place the images (A and B) vertically. Never change the aspect ratio of the images.
・Set the spacing between images to 2cm
・Include the prompt used for image generation in English in the margins
Errors occur frequently, but ChatGPT repeats trial and error on its own and eventually displays a download link for the PDF.
However, it is a bit unstable. Sometimes it places the same image twice or fails to include the prompt. It is a bit troublesome because I have to correct it with a prompt each time.


Expectations for Google Bard & Adobe Firefly
On May 11, Adobe announced a partnership between Firefly and Google Bard.
I don't know when it will be released yet, but it might become possible to generate images based on dialogue in Bard as well.
Me:
Can Bard generate images?
Bard:
As of November 8, 2023, Bard cannot generate images. However, Google has announced a partnership with Adobe's generative AI, "Adobe Firefly," and it is planned that image generation will become possible in Bard in the future.

Me:
Once Adobe Firefly is implemented, will it be possible to give instructions for corrections to the generated images?
Bard:
Yes, it is expected that once Adobe Firefly is implemented, it will be possible to give instructions for corrections to the generated images.
Adobe Firefly is an AI that can generate high-quality images based on inputs such as text, code, and images. Therefore, it is thought that even if you give instructions to correct the generated image, the AI will be able to understand and modify the image according to those instructions.
Specifically, the following can be mentioned:
・Adding or deleting specific objects or people
・Changing the color or shape of objects or people
・Changing the background of the image
I also look forward to the integration of Google Bard and Adobe Firefly!
I believe this ChatGPT update has advanced the image generation workflow and accelerated the move away from prompt engineering.
As I wrote in the article below, isn't this the most effective interface for breaking through the limitations of trial and error with text prompts?
In the design industry, I am convinced that the mainstream approach will become giving instructions to ChatGPT (a capable assistant) as an art director to get closer to the intended image.
ChatGPT will handle things by running the most suitable AI technology in the background as needed (we don't need to know about cutting-edge AI technology).
Why chat-based image generation will become mainstream
-
Improved usability:
Users spend less time worrying about word choice and expression in prompts, allowing them to intuitively get closer to the desired image.
With improved ease of use, it is expected that more people will use these image generation AIs.
-
Easy fine-tuning:
Users can add or modify details sequentially from the initial request, making it possible to generate ideal images by repeating more specific instructions and corrections.
-
Reduced educational costs:
Since you can use image generation AI while chatting, the barrier to entry for image generation will be lowered.
There are more errors during image generation...
It seems that heavy users from all over the world are flocking to it, and errors during image generation have increased. Let's take a break for now...
Errors are more likely to occur from the evening to late at night.

We are experiencing heavy server load. To ensure the best experience for everyone, we have rate limits in place. Please wait for 7 minutes before generating more images.
The server is under a heavy load. To ensure the best experience for everyone, we have set rate limits. Please wait for 7 minutes before generating more images.
That's all for now
I will continue to verify the updated ChatGPT.
Updated: Wednesday, November 8, 2023 / Published: Tuesday, November 7, 2023
