Original video in Chinese.
Key Takeaway
- When the Flux model is combined with ComfyUI workflows and influencer LoRAs, it can generate highly realistic AI images that can even pass for the real thing.
- The Flux model was developed by the core team behind Stable Diffusion. It produces highly realistic images and also enables precise control.
- Through a node-based workflow, ComfyUI solves the traditional problem of AI image generation being hard to control precisely, making fine-grained output possible.
- As a “skill pack,” LoRA lets the model generate images in a specific style, and multiple LoRAs can be stacked together.
- AI image generation has entered the implementation phase, with commercial potential in fields like e-commerce, and ComfyUI workflows can be shared.
Full Content
Friends who like looking at beauties on Xiaohongshu, please pay attention:
What you’re seeing now is very likely all AI-generated.
Don’t talk to me about how platforms can detect it. You don’t realize how realistic the latest technology can make images.
For example, look at this image. Do you think it’s real or fake?
Actually, this image was generated by me using AI. More specifically, I used the Flux model plus a simple ComfyUI workflow. There are two key points here:
First, the prompt part, meaning the text description of the image, was generated by Claude for me. I gave it an existing image, asked it to describe it in detail in English, and then used that.
Second, the reason the girl in the image looks so familiar to everyone is that I added an influencer LoRA. You can think of it simply as a small plugin that lets the model generate in a specific style.
Using such a simple method, you can make a fake look real. In fact, if you wanted to be even more ruthless, you could go straight from image to image. For example, find an image on Xiaohongshu that suits people’s tastes, then have AI generate based on that. It’s very easy to make the pose, body, and background basically the same, while only the face is different.
Older models weren’t good at local details, for example fingers would often end up with an extra one. But today’s models have improved a lot. These domestic platforms can’t detect it. So the people making accounts and selling accounts are relying on the Flux model I just demonstrated, together with ComfyUI.
Let’s start with the Flux model.
Over the past month or so, this model has been especially hot in the community. Many companies and teams are already using it in practice, for example in e-commerce.
So where did this insanely powerful model come from?
I’m sure everyone has heard of Stable Diffusion. Flux was created by the core team behind SD. They founded a new company called Black Forest Labs.
On August 1, Black Forest Labs officially released the Flux model, with three versions: schnell, the fast version with lower hardware requirements; the dev version, which has higher quality but also higher hardware requirements, ideally a 4090 GPU; and the Pro version, a closed-source version that can only be accessed through an API.
After the official version came out, the entire community also gave it strong support. For example, a GGUF version was released to make Flux easier to use for people with insufficient VRAM.
Once you have the model, the next question is how to run it. Right now, the best method is through ComfyUI.
Traditional AI image generation uses a long string of prompts, commonly called “spells.” That creates a very headache-inducing problem:
You can’t precisely control what AI generates.
After you send over a string of text, you have no idea how AI handles the rest of the process. And if you’re not satisfied with the result, you can only make tweaks at the text level. Often, this approach isn’t precise enough, and it’s also very inefficient.
So ComfyUI came along. It builds a workflow out of individual nodes. This node-based interface makes it very clear to users exactly how AI is generating the image, and if there’s a problem, where it gets stuck. Users can control the output in very fine detail.
A simple example: suppose you do e-commerce and don’t have the money to hire that many models for photo shoots, then just do face swapping. You or one of the girls under you first put on the sample clothes and take the photos, then put them into a ComfyUI workflow and create a mask specifically for the face area. In this way, AI only generates the face. It will generate a new face according to that outline, and then place it back into the original position.
Using this method, you get a virtual model. Doesn’t that feel a bit like skin-changing magic? Thinking about it like that is actually kind of scary.
If you feel the generated image looks too obviously AI-made, too greasy, too perfect, you can add a LoRA. For example, some expert made one that simulates amateur photography, making the image look like it was shot by a beginner, which makes it much more realistic. The influencer-style LoRA I used in my demo was also made by another expert. After I downloaded it and put it into a specific folder, I could select it in the workflow.
So you see, with ComfyUI, what used to be a huge lump of work gets broken down into individual steps and nodes, making it much simpler, much clearer, and much more controllable.
Even better, these workflows can also be shared. Once you get the workflow’s JSON file, you just drag it onto the canvas and it loads automatically. So whether in China or overseas, many people are making highly professional workflows. This is already a ready-made business.
You’ve definitely seen content like this on short-video platforms: first they show off how amazing the generated images are, then they display the extremely complex workflow they built, and finally they tell you to add them on WeChat if you want it.
If your machine can’t handle it, that’s fine too. Almost all compute rental platforms have partnerships with creators and provide ready-made image sets that users can directly use.
I bought an integrated package made by someone else, and it cost me a total of 1,500. They packaged everything up for me. It was over 100GB to download, and I didn’t even need to install it, which saved me a huge amount of time.
The benefit of paying for a finished product is that a lot of the basics don’t need to be fiddled with again; you just need to understand them. For example, besides the model, what exactly are Clip and VAE for? What files go in the key folders?
Practice and disassembly are the real focus. Once you fully understand other people’s stuff, you end up building your own. That’s my talent, and I know it very well. So spend the money when it should be spent, and you’ll definitely earn it back doubled.
This wave of Flux signals that AI image generation has entered the implementation phase. The fast movers have already started harvesting the fruits. That’s also why I waited more than a year and only started researching it now. I recommend that whether or not you want to use this technology to make some money, you should at least understand it. Think about how much our lives would change when what we see is not necessarily what’s real.
OK, that’s it for this episode. If you want to find me, come to the newtype community. See you next time!