SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Taking on impossible challenges with the new model "Tsubaki". | PIXAI | Next-gen model |

"Tsubaki" was officially released on July 10, 2025. I was able to use it during the αTEST and βTEST phases. This is my report on the user experience.
First of all, "Tsubaki" is the latest model in the Stable Diffusion (SD) series developed independently by PIXAI, and it is only available within the site. Web service providers create models exclusive to their own sites for differentiation. (After all, it costs hundreds of millions in development.) Also, please understand that Tsubaki is not a photorealistic model.
At the same time, the prompt helper(automatic prompt conversion) has also been updated to the latest version, and its performance is high-quality and compatible with Tsubaki.
"Tsubaki" is an SD-based model, but it demonstrates performance that transcends dimensions in the following areas. SD-based models generally evolve significantly about once every three months.
Six months ago is ancient history. (Illustrious 2.0 and Animagin 1.0 were probably the previous peaks.)
I've gone on for a while, but regarding Tsubaki,to conclude, it's the best.
I've looked at various models, and usually it's just "Oh, it's new!" or "It can draw XX beautifully!", but with this Tsubaki, even at the αTEST stage,I was shocked enough to make all XL models look outdated.(It's a major change.)
・Natural language understanding:This is the result of technology also used in GPT. Language comprehension has improved abnormally (also thanks to the synergistic effect of automatic prompt conversion). It can also accept input in various languages, and natural sentences are fine. (*1 case available) It is close to DALL-E 3, and in some areas, its understanding exceeds it. You could say the SD series has caught up.
・Prompt understanding:Similar to the above, but it's the ability to turn that into a concrete picture.
Screen composition, positional relationships of objects, and for people, gestures and expressions are wonderful 👍.
・Understanding relationships between multiple objects: This is also related to positional relationships, but as the name implies, there are multiple subjects. And while it is difficult for generative AI to draw multiple distinct characters (all AI generation), it can draw 4-5 distinct characters. And you can also specify the position of those characters. (*2 cases available)

My evaluation criteria are as follows:Since there are things that are hard to quantify,
I judge based on the following abilities. It is strongly linked to the actual user experience.

(High-level language) + (Language reproduction/Screen composition ability) + (Image quality) + (Existing knowledge) + LoRA

It's a bit of a condescending attitude, like "If you can't do at least this much, you're not an AI" (lol).
LoRA is not yet implemented as of the time of writing this article, but it is said to be released soon. Since it is a new architecture, past LoRAs cannot be used.
I'm looking forward to it because LoRA creation can also be done within PIXAI.


Below is my user experience based on the provider's boasts. I evaluate the provider's points of pride.
Character consistency engine= Unconfirmed.
Story comprehension= Confirmed. (Rating: ⭐️⭐️⭐️) It's crazy! (*3 cases available)
Innovation in composition expression= Confirmed. (Rating: ⭐️⭐️⭐️) I need to improve my user-side control techniques. Composition-based instructions work. Confirmed.)
Advanced direction of light and shadow
= (Rating: ⭐️⭐️⭐️) I need to improve my user-side control techniques... Good things come out even if I stay silent, but whether the user can give instructions is unconfirmed. (It's a problem with my skills.)

*I rarely call up existing content (copyrighted material), so I will leave it as unconfirmed because the number of trials is low, but since it's PIXAI, it will probably be able to do it.
Also, there is a detailed explanation from the provider, so please read that as well.

Now, to the main topic. I tried using "Tsubaki".

🔸Case 1 ⚠️ All images are from α or β, not the current version.
Since the image quality wasn't good back then, I used what I drew with Tsubaki as a reference image and
finished it with other models. Everything posted here is Tsubaki before finishing. How much can it draw with short Japanese? (I did translate it into modern language, as expected.)
The caption is the prompt as is.

  The sound of a frog jumping into an old pond              Even when coughing, I am alone
shopping, lots of spresent box, holding present box, stacked present box, girl walk in GINZA, beautiful details, fine illustration, , head up, looking at viewer, white hair, necklace, 1 girl, cute, bob cut, thick lips, (((from side, full body,white background ))),
*Walking while holding stacked boxes is difficult even with prompts. It pulled it off easily.
A woman with long black hair is playing the shamisen in a kimono, the woman next to her has a blonde short bob and is wearing street-punk clothing, and is playing air guitar (a pose of just playing without holding a guitar) *Unfortunately, it seems it cannot understand air guitar.

🔸Case 2 ⚠️ All images are from α or β, not the current version.
I am drawing by specifying the order. Just in case, I have also added the reverse order.

I will draw a fantasy dream world. Animals are stacked like the Town Musicians of Bremen (fairy tale). From the top, a rooster, a black cat, a dog, and an alpaca are stacked on each other's backs in that order. (From the bottom, alpaca, dog, black cat, rooster order) But this is just for fun. They look happy. Small birds in the sky, and a pair of hedgehogs at their feet are watching the scene and laughing. Rural countryside, on a hill, a sunny spring day.

No, but there are a lot of typos and omissions in the prompt. But it is drawing in this state.









🔸Case 32girls, multiple girls, animal ear.A two-shot of a cat-eared girl in a green adventure outfit and a cat-eared girl with red hairThe red cat-eared girl (smiling, red adventure outfit) puts her arm around the green cat-eared adventurer, acting familiar.The green cat-eared girl looks a little bewildered

■ Guild Registration
The inside was lively.

■ Chance Encounter (Kagura)At that moment.

"Ah, that's the all-you-can-eat one, right?"

The one who thrust her face in from the side was a red-haired, energetic-looking girl.She smiles, showing her beast-man ears and large fangs.

"Can I come too? I was just getting hungry!"

Mina flinched involuntarily.

"Eh, um, well...?"

The receptionist nodded easily."She is Kagura-san, from the famous fighting tribe 'Agna-ryu'. Her skills are solid.""Hehe. Taking down a rat is a piece of cake!"Saying that, Kagura slapped Mina on the shoulder.

"Nice to meet you, Mina!"

I suddenly wrote a novel, but this novel is the prompt itself.
I color-coded the characters in the novel so you can see if they are being drawn distinctly.
By the way, it can handle about 4000 characters. However, since it only draws one image, there's no point in it being too long.
Let's cut out the scene part.

Well, this has gotten long, so I'll stop here. Man, even reading it back now, there are a lot of typos in the prompt. Still, I'm grateful it draws it for me. Also, I think this is the case for others too, but what I've pasted here is the best or most understandable one.
If you're into AI generation, you probably know this, but the AI will think for itself and draw parts you didn't instruct it to, so there are good results and ones you might not like. If you have a clear image in mind, make sure to be specific in your prompt.

I wonder when LoRA will be available. If it were, I could compose it with my original characters. (Probably)


いいなと思ったら応援しよう!

picko よろしければ応援お願いします! いただいたチップは 『誰かにチップを配ります』パカ