Original video in Chinese.
Key Takeaway
- Fine-tuning Flux LoRA is a valuable skill that lets AI generate people or products with a specified appearance, and multiple LoRAs can be stacked together.
- The Flux model produces highly realistic images; ComfyUI solves the controllability problem in image generation, and LoRA solves the model experience problem.
- Making a LoRA requires preparing source images (20 as a starting point) and training through a fine-tuning tool (such as a project on Replicate).
- A trained LoRA can be used in the cloud or locally, and multiple LoRAs can be stacked to achieve more complex effects.
- The article emphasizes the business potential of AI image generation in e-commerce, IP operations, and other fields.
Full Content
If you want to make money, this is a skill you should really learn: fine-tuning Flux LoRA.
Once you’ve mastered it, you’ll be able to make AI generate people or products with the exact look you want. You can even stack two LoRAs together. For example, you can have your virtual model hold a specific product.
If you do e-commerce or run an IP, you’ll definitely know how much trouble real-life shoots for people and products can be. By fine-tuning Flux LoRA, you can save a lot of money and a ton of time.
Hello everyone, welcome to my channel. I’m one of the few creators in China who can explain both the Why and the How of AI clearly. Remember to follow me. If you want to connect with me, come to the newtype community.
In the last video, I introduced how the Flux model plus the ComfyUI workflow works, and I also demonstrated the effect of a celebrity-style LoRA. If you haven’t watched it yet, be sure to go find it on my homepage. After watching, you’ll get the idea. I can briefly and bluntly summarize it like this:
For the Flux model, just remember that the realism of the images it generates is higher than DALLE and SD. Only when the realism problem is solved is commercial use possible.
For ComfyUI, just remember that it solves the control problem, allowing the model to generate according to the steps you set and within the scope you define. Only when the controllability problem is solved is large-scale commercial use possible.
As for the LoRA we’re focusing on today, just remember that it solves the model experience problem.
For example, the model knows what a beautiful woman looks like, but it doesn’t know what a beautiful influencer looks like. But we can’t train a model from scratch—the cost is too high and it’s not realistic. So we give the model a skill pack and tell it: this is what a beautiful influencer looks like. Next time someone makes a request, you generate according to this.
That skill pack is LoRA. Making this skill pack is very simple. Compared with fine-tuning a large language model, it’s basically so easy it’s “a piece of cake.” Let me give you a quick demo.
The young lady you’re seeing now is a teacher I especially like, so I’ll use her as the example. For instance, I want AI to generate images according to her appearance. How do I do that?
First of course, prepare the learning materials—that is, images of the teacher from different angles and with different expressions. That way AI can generate based on them, right? You don’t need too many images; 20 is the starting point. I’ve prepared 25 here.
Next, we need a fine-tuning tool. Many platforms at home and abroad already provide ready-made ones. I’m using a project on Replicate.
Go to the Replicate website. Search this string of keywords. Then find this project.
There are two steps here that you absolutely have to do:
First, upload a ZIP archive. I tried it, and the RAR format doesn’t seem to work. Remember to compress it in ZIP format.
Second, create a new model. That way, after training is complete, it will exist on the platform as a model.
The other settings can be adjusted or left alone. For example, if you want the training results to be synced to Hugging Face, enter your ID and Token.
Once all of that is configured, you can start training. Because the GPU used is H100, it’s very fast—basically done in about 20 minutes.
The trained LoRA can be used directly on Replicate, or downloaded.
Let’s test the effect directly on Replicate. Remember, the prompt must include the trigger keyword set earlier.
I’m pretty satisfied with this result. If there were more photos, with richer expressions and angles, the trained effect would definitely be better.
You can also run it locally. Put the downloaded file into the loras folder. Then go into ComfyUI and select it in the LoRA loader. The default strength is 0.8; different values will produce different results. I tested several cases, and I’ll show them to you.
As mentioned earlier, multiple LoRAs can be stacked together. So I also trained a product LoRA, using Ray-Ban sunglasses images. I wanted to test whether it’s feasible for a virtual model to use a specific product. Because this would be used for placements or sales.
Stacking multiple LoRAs is simple: hold down the Alt key to duplicate a LoRA loader, then connect the lines and select the file. Be sure to include all the trigger keywords in the prompt.
I won’t adjust the strength value here; let’s just generate it directly.
This stacking worked pretty well. Our teacher successfully put on Ray-Ban sunglasses. So everyone can let their imagination run: a virtual model plus a product, or virtual model A plus virtual model B—this should all be possible. If you have a feel for doing business, you’ll understand how much room there is here.
OK, that’s it for this episode. If you want to discuss it further, come to the newtype community. See you next time!