SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

I prepared training images using a batch generation tool and actually created a character LoRA with Anima

Previously, I created a GUI tool that allows for the batch generation of images with different poses and expressions while keeping the character consistent.

With this tool, once you set the common prompt for the character, you can read differences such as face direction, composition, expression, and pose sequentially from a txt file to generate images.

In my previous article, I focused on the mechanism and usage of the tool.

This time, as a practical example,

I will batch generate images for training a character LoRA and use those materials to actually create a character LoRA in Anima.

Rather than just finishing after making the tool,

I will verify whether it is truly useful for preparing LoRA training images,

including the process from training to the final generation results.




What I used this time

The batch generation tool I used this time and the 35 original prompts are available in the following article.

GUI tool for batch generating different poses with a fixed character

This is a tool that manages common character prompts and differential prompts for poses and expressions separately, generating multiple images in sequence.

I have released versions for Stable Diffusion AUTOMATIC1111 and ComfyUI. I am using ComfyUI this time.

Summary of 35 prompts for LoRA training images

This is an article that summarizes the 35 prompts for training images I used when I previously created the character LoRA for Shion Amatsumi.

This time, I based the composition and pose distribution on those used in this article, while replacing only the character portion with the features of a new character.


The character I am creating this time

The character I am creating a LoRA for this time is based on white hair and pink.

I set the main features as follows:

・White hair
・Pink eyes
・Long ponytail
・Pink inner color
・Pink ribbon
・Pink jacket
・White shirt
・Pink skirt
・Sharp teeth

For the LoRA trigger word, I will use the following word based on the character's name.

momose_ririka


STEP 1: Change the existing 35 types of prompts for the new character

In the prompts I used previously, I had prepared multiple compositions for LoRA training, such as close-ups of the face, upper body shots, and full-body shots.

This time, I will change only the character-specific parts to the new features while keeping the compositions and poses as much as possible.

The txt file used this time

The 35 types of prompts used this time are summarized in the attached file below.

The common elements of the character are summarized under [CHARACTER] in the txt file.

Based on that, I have listed the parts I want to change for each image, such as face close-ups, upper body, full body, expressions, poses, and backgrounds, separated from [POSE1] to [POSE35].

Common elements that become unnecessary depending on the composition can be adjusted using DEL:, and costumes or accessories you want to add can be adjusted using ADD:.

For example, if you want to delete features that are only visible from the front when generating a shot from behind, you can write it as follows.

DEL:pink eyes
DEL:white shirt
DEL:sharp teeth

Conversely, if you want to add costumes or accessories only for a specific pose, you can specify them with ADD:.

If you rewrite the character part, I think you can reuse the same compositions and poses for other characters as well.


STEP 2: Batch generate a total of 140 images from 35 types of prompts

This time, I prepared 35 types of differential prompts and generated 4 images for each.

Calculating this,

35 types × 4 images = 140 images in total

is the result.

In the txt file, if you write it as follows, you can generate 4 images of that pose.

[POSE1 4]

Even with the same prompt, the expression, fine composition, and state of the hands and feet will change slightly depending on the Seed.

Therefore, instead of generating only one image for each composition, I decided to generate four candidates each time and choose the one most suitable for a training image from among them.

I will not be using all 140 images for training.

Ultimately, I will choose one image from each prompt, resulting in

35 types × 1 selected image = 35 training images

as the final configuration.


Testing one by one before the actual batch generation

Although I will eventually generate four images for each of the 35 types of prompts, I recommend testing each prompt one by one first, rather than generating all 140 images at once.

Batch generation is convenient, but if there is a mistake in the prompt, multiple images containing that mistake will be generated at once.

Even in the txt file I prepared this time, there was one place where the pose description was missing.

Simple descriptive omissions like this can be noticed before the actual batch generation if you check the generation results one by one in advance.

Also, even if there are no mistakes in the description itself, elements included in the common prompt may interfere with the specified composition.

A failure example where a looking-back composition is created instead of a back view

For example, if you want to generate a back view but 'pink eyes' remains in the common prompt, the AI may try to draw the eyes, resulting in an incomplete back view.

Sometimes the face may be turned slightly to the side, or it may result in a composition where the character is looking back at you, as shown in the image.

Elements like 'white shirt' that are not visible from the back of the clothes may also be more stable if removed depending on the composition.

Therefore, for the back view prompts, I individually removed features that are not visible or elements that might interfere with the composition.

DEL:pink eyes
DEL:white shirt
DEL:sharp teeth

It is reassuring to check the following points before the actual generation.

・Are the descriptions of poses and compositions missing?
・Are the specifications for close-ups, upper body, and full-body shots reflected?
・Is the face visible even though it is a back view?
・Are the elements added with ADD: reflected?
・Do the elements specified with DEL: remain?
・Are the common prompt features interfering with the composition?

After correcting the problematic areas, I will change the number of generated images to 4 and start the actual batch generation.

Although it takes a little more effort, it prevents generating 140 images with configuration errors, which ultimately reduces the need for redoing the work.


Parts that became especially easier with batch generation

Previously, when performing the same task manually, the following operations were required:

Rewrite the prompt

Press the generate button

Wait until generation is finished

Change to the next prompt

Generate again

This needs to be repeated for 35 types.

Each operation is not that difficult.

However, when generating dozens of types while changing compositions and expressions, it takes quite a lot of effort.

Since you need to check when the generation finishes to change to the next prompt, you also end up staying in front of the PC for a long time.

With this tool, once you load the txt file and start the generation, it will process the registered prompts in order from the top.

Since you can basically do other work while generating, you don't need to manually change prompts many times.

For tasks like this one,

generating 4 images each for 35 types of prompts

I was able to really feel the convenience.

However, what the tool does automatically is only generating images in order.

The generated images also include the following:

・Faces or hands are distorted
・Character features are weak
・Not in the specified pose
・Similar compositions
・Contains elements you do not want to train

Even if generation can be automated, final image confirmation and selection are necessary.

Even so, because the time spent repeating generation operations has decreased, compared to the work of creating images,

You will be able to focus on the task of "selecting which images to use for training."

STEP 3: Select 35 images to use for training from the 140 available


Once generation is complete, select 35 images from the 140 to actually use for LoRA training.

Screenshot of the folder when created in batches of 4

This time, I selected one image from each prompt.

I compared the 4 images generated with the same prompt and checked the following points:

・Are the face or hands significantly distorted? ・Do the character's basic features appear? ・Is the expression or pose as intended? ・Are there any unwanted accessories or features included? ・Is the composition too similar to other training images?

Also, when looking at the training images as a whole, I ensured that the ratio of close-ups, upper-body shots, and full-body shots was not too skewed.



I also selected images to ensure that expressions, body orientation, and backgrounds were not all the same.

[List of the 35 selected images]

These are the 35 images selected from the 140 generated results, one for each prompt. I chose them to ensure a good balance of close-ups, upper-body, and full-body shots, as well as expressions and poses.

STEP 4: Create a LoRA using Anima LoRA Factory


Once the selection of training images is complete, the next step is to actually create the LoRA.

This time, I used "Anima LoRA Factory," which makes it relatively easy to create LoRAs for Anima.

This is an article by Fukashigi-san👇

Anima LoRA Factory provides features for automatic captioning and bulk tag editing.

Instead of manually entering all captions for each image, you can automatically tag them first and then organize unnecessary tags in bulk.

I think it is a tool that makes it relatively easy to grasp the workflow, even if you are creating an Anima LoRA for the first time.

Note that this article does not explain all the features of Anima LoRA Factory in detail, but focuses on the flow of creating a character LoRA as I did this time.

End of process.


Gather the 35 selected images into a folder

First, gather the 35 images selected from the 140 into a single folder.

Next, specify the folder where you saved the 35 images in the image folder path under "Dataset Preparation" in Anima LoRA Factory.

Screen where the image folder is specified in Anima LoRA Factory

Execute automatic tagging

Once the image folder is specified, execute the automatic tagging indicated by the red circle.

The content of the images will be analyzed, and caption tags will be automatically added to each image.

Features recognized from the images, such as clothing, facial expressions, poses, backgrounds, and composition, are written out as tags.

Caption screen after automatic tagging

Automatic tagging is convenient, but instead of using the output tags as they are, organize them to suit the LoRA you want to create.

Organize the caption tags

For character LoRAs, there is a basic concept of:

deleting from the caption the features you want to be automatically reproduced when the trigger word is entered

that approach.

If you leave features present in the image in the caption, the model will more easily learn those features as elements specified by individual tags rather than as the character itself.

Conversely, if you delete them from the caption, those features will be more easily learned as being tied to the trigger word.

Features you want reproduced just by the trigger word

The following types of features are candidates for deletion from the caption:

・Unique hair color
・Unique eye color
・Character-specific hairstyle
・Distinctive accessories
・Unique outfit
・Character-specific ears or horns

For example, if you want to reproduce white hair or pink eyes just by entering the trigger word, delete "white hair" and "pink eyes" from the caption.

Features you want to change or control during generation

On the other hand, you should keep elements you want to change via prompts during generation in the caption.

・Facial expressions
・Poses
・Backgrounds
・Camera distance
・Face and body orientation
・States such as sitting or standing
・Outfits and hairstyles you want to change during generation

For example, by keeping tags like smile, sitting, from side, or outdoors, it becomes easier to train the model while separating facial expressions, poses, and composition from the character's unique features.

It is not necessarily correct to delete all tags related to the character.

If you delete everything including facial expressions, poses, backgrounds, and composition, those elements might become tied to the trigger word.

As a result, the same facial expression or similar composition may appear every time you use the trigger word, which reduces the flexibility of the LoRA.

Which features to delete and which to keep depends on how you want to use the finished LoRA.


Tags deleted this time

For this character, I deleted the following tags from the caption.

pink eyes
white hair
sharp teeth

You can organize tags for 35 images using bulk addition and deletion

These three are elements I want reproduced as the character's basic features when the trigger word is entered.

On the other hand, I did not delete tags related to hairstyle and clothing, but left them in the caption.

This is because I wanted to leave room to change the hairstyle and outfit to some extent with the finished LoRA.

For example, suppose the caption immediately after automatic tagging looks like this:

1girl, white hair, pink eyes, ponytail, sharp teeth, pink jacket, smile, upper body, indoors

After organizing, it looks like this:

Example of a caption with three tags removed

In this way, you only delete the features you want to tie to the character, while keeping facial expressions, composition, and outfits you want to be able to change.


Adding trigger words to all captions

Once the tag organization is complete, add the LoRA trigger word to all captions.

A trigger word is a word that acts as a signal to call up that character during generation.

Character names are often used, but it is safer to avoid names that overlap with existing characters or common English words.

If it overlaps with a word that already has a meaning, it may interfere with concepts the base model already knows.

This time, I used the following trigger word:

momose_ririka

By using a unique name that includes an underscore, I have made it less likely to overlap with common words.

The captions after adding the trigger word will look like this:

Screen after bulk-adding the trigger word

Note that I do not necessarily consider the caption settings used here to be the best.

The reproducibility and flexibility of the completed LoRA will change depending on which tags are deleted or kept. Also, the appropriate captions will differ depending on the content of the training images and which features you want to be able to change after completion.

Therefore, I am not presenting these settings as the correct answer, but rather as one example I actually tried.


Configuring training settings

Once the caption preparation is finished, move to the training settings screen.

Here, specify the storage locations for the Anima base model, VAE, and text encoder to be used.

At the same time, set the folder where the completed LoRA will be saved and the file name for the output LoRA.

There is generally no problem with changing the LoRA file name after training.

You can also change the number of training steps and the learning rate, but this time I created it based on the initial settings.

If you are creating one for the first time, I think it is fine to start by training with settings close to the defaults and then adjust after checking the generation results.

However, the optimal settings will vary depending on the number of images, the content of the images, the base model used, and the desired level of reproduction.

If you need to make strict decisions regarding training settings, we recommend checking the official Anima LoRA Factory information or consulting with an expert knowledgeable in LoRA training.

Once the settings are complete, press "Start LoRA Training" and wait for the training to finish.

When training is complete, the LoRA file will be output to the specified save destination.


Trying to generate with the completed LoRA

Once training was complete, I loaded the created LoRA and actually tried generating the character.

I was able to successfully generate the character and even change their outfit.

Four generation results using the completed LoRA

The purpose of this article is not to perform a detailed comparison of the performance of the completed LoRA.

The main goal is to confirm whether a character LoRA can actually be created from training images prepared using a batch generation script.

is the primary objective.

Judging by the generation results, I was able to create a LoRA with the character's basic features even from materials generated in a batch.

On the other hand, there may be compositions where the character's features become weak, or conditions where it is difficult to change outfits or hairstyles.

I would like to verify the detailed reproduction levels and differences between settings in another article if necessary.


What I felt after actually using it

This time, I used my own batch generation script to create training images for the LoRA all at once.

What was particularly convenient was being able to generate a total of 140 candidate images at once without having to manually swap out 35 types of prompts.

Unlike before, there is no longer a need to repeatedly perform the operations of:

changing the prompt
starting generation
waiting for it to finish
changing to the next prompt

over and over again.

On the other hand, not all generated images can be used for training as they are.

It is necessary to check for images with weak character features, distorted limbs, or compositions that did not match the specifications, and ultimately select them by human eye.

Also, while automatic captioning is convenient, it was necessary to organize the tags while considering which features to associate with the character, rather than using the automatically assigned tags as they are.

In other words, what could be automated with this batch is mainly the image generation process.

It is not possible to completely automate everything, including image selection and caption organization.

Even so, by being able to prepare candidates for training images in bulk with fewer operations, I was able to significantly simplify the material preparation process, which is particularly time-consuming in LoRA creation.


Summary

This time, I used a custom batch generation script to create training images for a character LoRA.

I generated a total of 140 images, 4 each from 35 types of prompts, and selected 35 of them to use for training.

Furthermore, I automatically captioned the selected images, organized the necessary tags, and then actually created the character LoRA using Anima LoRA Factory.

By using the batch generation script, I eliminated the need to repeatedly change prompts and perform generation operations, which streamlined the process of preparing training materials.

However, since image selection and caption organization affect the reproducibility and flexibility of the completed LoRA, human verification is ultimately necessary.

Although it is not possible to automate everything,

automating the task of creating images and having humans focus on the task of selecting images

is a way of using it that I felt was quite compatible.

I hope to further verify the detailed generation results of the LoRA created this time, as well as the differences when changing captions and training settings in the future.

Thank you for reading until the end.

[Usage Environment]
- Comfy UI
- Model: anima-base

* In this generation, LoRA is applied on the ComfyUI workflow side. Therefore, the prompts in the distributed txt file alone may not result in the same style or outcome as the posted images.


いいなと思ったら応援しよう!