SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Understanding AI Drawing as Magic

On AI image generation, which becomes more like sorcery the more you do it.

The following is a rough guide to logic and control that might make experts in the field very angry.

For the record,this is know-how that works across major services (DALL-E 2, MidJourney, Stable Diffusion, Disco Diffusion, crayon, dall-e mini, etc.).


A rough understanding of how image AI works

For conversational AI,spellspromptsare merely things that determine thedirectionvectorof image generation.

For example, the following is an example of an image generated with "I love apple." Somehow, something vague comes out.

I Love Apple

This is because the directionvector of "Apple" simultaneously holds multiple possibilities, such as "apple (fruit)," "green apple," "Apple Computer (old rainbow logo)," and "Apple Computer (new logo)."

In other words, "Apple" is a very ambiguous directionvector for AI.

It doesn't even know if it should be a photo or an illustration. Also, it's combined with a "heart." Such short incantations become uncontrollable in many cases.

AI will generate pictures that satisfy ambiguous commands simultaneously (Note: faces splitting is a separate phenomenon).

Excellent AI manipulation requires a design that bundles directionvector through incantations.

The following is an incantation with the constraint "apple as a fruit". It improved, but red apples and green apples still exist at the same time.

I love Apple (fruit)

As you can see, even to produce one apple, a certain amount of directionvector incantation is required.

Actually, for more stable incantations, you need to change your way of thinking.

"A single red apple on a simple background, high quality, for an e-commerce site"

If you do that, an apple will be generated in one go.

good quality photography of a singular red apple on simple background for e-commerce site --stop 75

"Product photos for e-commerce" have high reproducibility in background and layout. Once you realize this, you can refine better spellsprompts.
However, AI vectors extend into thousands of dimensions. If you just list words randomly, multiple words will end up canceling each other's vectors out.

Therefore, to effectively command an AI, you need two understandings: "vocabulary of powerful wordsrunes that strongly define vector direction" and "constructing spells with reproducibility."

Trivia
When colors are in a "state where red, blue, and green seem mixed together," it is likely that the vector is too ambiguous. That is not the MidJourney art style. It is a state that requires more vector limitation.


Seek out powerful words

Powerful wordsrunes are words with clear directionstrong vectors. You could also call them "pure words without ambiguity."

This stems from the AI's learning mechanism. AI learns by pairing images found on the internet with their explanatory text. From learning countless pairs, it forms a vector space that could be called an Akashic Recordpossibility space.

Because of that mechanism, a strong power is born in "pairs of images and text that appear repeatedly in learning without the content wavering."

If you place runes like "Unreal Engine" or "Playstation5" in the final section of your spell chant, image quality tends to improve. This is likely because the AI has learned content such as "{Unreal Engine (game creation tool name) has amazing image quality}" from pairs of articles and images.

Similarly, adding "beautiful concept art of" to the first section is also powerful. This is also because "{images called average concept art are well-drawn}."

Comparison diagrams are shown below.

landscape
landscape by unrealengine
beautiful concept art of landscape

Personally, I recommend camera settings as strong runes.

If you add things like "shot with a Canon EOS 5D Mark 4 and SIGMA Art Lens 35mm F1.4 DG HSM, F-stop 2.4, ISO 200, shutter speed 2000" to the end of your chant, the quality will increase significantly.

You can control not just image quality, but also composition, lighting, bokeh, and more overall.

landscape taken by Caanon EOS 5D Mark4 and SIGMA Art Lens 35mm F1.4 DG HSM, F1.4, ISO 200 Shutter Speed 2000

This is because the AI has learned that "images with such notations are professional photos, official manufacturer sample photos, or works from photo sites taken by high-level amateurs." When you want to aim for foreground or background bokeh, they function as strong runes.

When in doubt, use a camera!

First, beginners should find a repertoire of strong runes while playing around.


Find powerful syntax

The order of word selection during incantation is also important.

Below is the difference in display between "landscape with robot" and "robot with landscape".

robot with landscape
landscape with robot

Even with commands that look the same at first glance, you can see that the handling of the robot changes significantly. In this way, "what you say and in what order?" greatly influences the quality of image alchemy.

As a general rule, it loosely has characteristics such as "the earlier the sentence, the stronger it is, and the later the word, the weaker it is," "the last sentence is moderately strong," and "words before prepositions or relative pronouns appear stronger, while words after them become weaker."

Basically, it is best to create a spellprompt| in a "style close to captions in art museums, photo sites, or art books."

Among advanced mages who have been doing AI alchemy since ancient times, the following structure seems to have become a standard.

<Overall format> <Subject> <Subject supplement> <Artist> <Overall supplement> <Flavor>

Recommended incantation rhythm

<Overall format> Detailing oil painting of
<Subject> The great white castle on deep forest landscape<Heroic Spirit> by CASPAR DAVID FRIEDRICH and CLAUDE LORRAIN,
<Overall supplement> perfect lighting, golden hour,<Flavor> taken with Canon 5D Mk4

Detailing oil painting of The great white castle on deep forest landscape by CASPAR DAVID FRIEDRICH and CLAUDE LORRAIN,perfect lighting, golden hour, taken with Canon 5D Mk4 --ar 16:9



Summon the great ones of old

In AI magic incantations, including the name of a great one of old in the second or third syllable stabilizes accuracy. It is the name of a top Heroic Spirit like Da Vinci or Rembrandt, or one of the ancients not yet forgotten like Caravaggio.

Basically, when drawing landscapes, weaving the name of a landscape painter into the incantation, and when drawing portraits, weaving the name of a portrait painter into it, will increase accuracy.

Chanting the names of masters of portraiture stabilizes face generation and lighting.

However, it is better to think that what is summoned here is not Da Vinci himself, but rather "something that embodies the Da Vinci-esque".

All the AI has learned is "a collection of images associated with the word Da Vinci." This is because it also contains elements of Da Vinci's disciples, works influenced by him, his family home, and works compared to his.

If you chant "portrait by Dalí," Dalí himself will show up!

portrait by Salvador Dalí

This is because the vector space for "Salvador Dalí" contains many photos of Dalí himself performing with gusto, as well as his self-portraits. In this way, we are not talking to the AI at the limit. It is important to be aware that we are indirectly passing vector parameters to the AI through language.

By the way, heroic spirits and spirits are subject to regional and popularity modifiers. Even if you chant the names of Japanese heroic spirits, they are difficult to control under the Western magical foundation (American AI).In terms of popularity modifiers, it is hard to stabilize unless it is a great heroic spirit of the Hokusai class.

Even with the same chant of "landscape painting of forests," changing the author makes this much of a difference.

landscape painting of forests by J. M. W. Turner --ar 16:9
landscape painting of forests by Hieronymus Bosch --ar 16:9

Personally, I recommend enshrining multiple heroic spirits together. So-called composite heroic spirits. If you mix heroic spirits with the same attributes, the expression will stabilize. Conversely, if you mix heroic spirits with different attributes, it will become unstable, but the expressions will blend.

The figure below is a composite of the Bosch painting used as an example in the figure above with another artist. You can see that while maintaining the lumpy object feel unique to Bosch, it has become a completely different art style.

A synthesis of Bosch's vector with another artist's vector


I think it is better to avoid using the names of contemporary artists alone as much as possible. Basically, when including a painter's name, I think it is better to include multiple names, or use the name of someone who has been dead for 70 years as the main one, and not create prompts that rely too much on the art style of a specific individual.

You might get involved in emotional or moral troubles rather than legal ones. (As a general rule, copyright law is granted to individual works themselves, and styles, ideas, and concepts are not subject to copyright protection on their own).

Those who prioritize risk avoidance should summon entire eras or movements, such as "Renaissance" or "Rococo."

I feel like this will probably become a matter of etiquette in the near future.


Never include the names of real actors, politicians, or celebrities!

One piece of know-how for stabilizing human images is to "include the names of real celebrities."However, you absolutely, absolutely should not do this.

It is better to think of this as a forbidden spell.
I once witnessed a terrifying, instant-death-level incident on a Discord timeline.

Someone was trying to create an image by combining phrases like "Hollywood actress" and "photo of a captured princess knight." Perhaps the person themselves had no malicious intent and just wanted to improve the quality of their "captured princess knight-like something" work. I think they just casually included the name of a "Hollywood actress" for that purpose.

However, the AI suddenly outputted a "photo-realistic, nipple-slip princess knight-like something, and the face was that of a Hollywood actress"!!

It was likely an accident, but it was a pretty dangerous image.

The person who accidentally created it deleted the image, or the management deleted it after receiving a report, so no harm was done. I think it was probably an unfortunate accident without malice.

But if this image had been released to the world... it would have become a full-fledged fake pornographic image. If it had spread, it would have been a catastrophe with no way to avoid massive damages.

Human transmutation using famous people as a base has that kind of instant-death risk.With the current behavior of AI, NSFW images of real people can be suddenly transmuted even if you don't intend it.I think it is better not to lightly create prompt text that includes the names of famous people.

The police will come knocking.

In any case, it is best to absolutely never create "prompts containing real names that could form the face of a real person." Since such accidents and incidents are bound to occur frequently from now on, I am issuing a strong warning in advance.


Understand the roots of AI

Below is a more in-depth technical discussion.

Image generation methods like MidJourney, DALL-E, and Stable Diffusion are based on a technology called Clip Guided. If you understand this mechanism, it is easier to visualize the vectors of text.

Clip is an AI that converts images and text into "mutually comparable vectors." To put it more simply, it can also be called an AI that measures "how much does this pair of image and text match?" with a score.

Current image generation uses this Clip to generate images so that the distance between "text" and "picture" becomes small.

AI is created by training on hundreds of millions of "paired images and text." In other words, it is trained on "images on the internet and the captions attached to them."

For this reason, Clip-type image generation tries to generate "an image that receives a caption from the user and is likely to be paired with that caption." If you understand this basic rule, the accuracy of image transmutation will increase significantly.

To get the image you want, you just need to imagine, "If the image I want were on an art museum site, an e-commerce site, or a photo site, what kind of caption would it have?" and then think of text that matches that.

The aforementioned "basic syntax" and the like are built with a structure very similar to that.

<Overall Format><Subject><Subject Supplement><Author><Overall Supplement><Flavor>

Oil painting 'Landscape with Sunflowers' (by Van Gogh). Bright red sun, passionate colors, strong brushstrokes... Writing in this style is very similar to descriptions found in art museums or art auction catalogs.

As you get more used to it, there are even more specialized notations (directly injecting vector weights), but that is advanced magic, so that is a story for another time.

Example of direct vector injection notation
oil painting::1 landscape with sunflower::1 by Vincent van Gogh::0.25 water color ::-0.25


There are various other tricks like using noise or original images for guidance, inpainting (correction and completion), upscaling, downscaling, composite AI, and repurposing for Photoshop or Blender materials... but that would take too long, so I will omit them.

Now everyone, please follow the instructions and enjoy your magical life!

High quality concept art of a mage standing at the center of a magic circle and exercising great magic. Center of the legendary chapel, many crowds. High fantasy. Golden. XXXXXXXX composition. XXXXX light. The glow of magic. XXXXXXXX XXXXXXXX. art station trending, octane render, 8k, by XXXXXXXX and XXXXXXXX and XXXXXXXX, golden, crazy detailing


Bonus: Keep the mystery hidden...?

Occultism is written as the study of hidden things, but magic can be thought of as a technical system that 'conceals the process and only reveals the output.' In other words, from the perspective of magic, such know-how only gains value and power when kept to oneself. It is the idea that '

mysteries must be kept secret
.'

However, this is the 21st century. In the internet age, the spread of information has become irreversible and rapid. Someone will scatter any hidden mystery for free on the net within a few days. It is no longer realistic to monopolize information.

If that is the case, I think we should periodically release information and bundle knowledge together to reach the deep roots of magic.

Regarding prompts, there is a DALL-E Prompt Book cheat sheet below. If you read this, I think you will quickly understand the necessary basic vocabulary.


Note that memorizing the quirks of individual AIs or spellsprompts is not very meaningful. Because you can just look at the cheat sheet.

The know-how you should truly learn is 'When you encounter a new black-box technology, how can you quickly form a hypothesis about its behavior and tentatively bring it under control?' I think this reproducibility of technique is what is important.




Like a sketch of a mecha



For those who want to understand it technically or try touching the code

Those who want to study the mechanisms thoroughly should read the following.


If you want to read the underlying code, look up Latent Diffusion Model or Glide X3.



While I was spending about three days refining this article and trying hard to make it consistent with the Type-Moon world, @shi3z published their article first. I'm so frustrated.


いいなと思ったら応援しよう!

深津 貴之 (fladdict) いただいたサポートは、コロナでオフィスいけてないので、コロナあけにnoteチームにピザおごったり、サービス設計の参考書籍代にします。

この記事が参加している募集