Original video in Chinese.
Key Takeaway
- The iPhone 17 Pro’s performance is a pleasant surprise: it can run an 8B model locally and efficiently (such as Qwen3-8B at 12 tokens/s), reaching desktop-level performance and exceeding expectations; I used PocketPal to download and test it, and the inference is smooth.
- My main reasons for upgrading: the Hong Kong 512G model (11,750 yuan), for the convenience of eSIM activation (Red Tea eSIM, with network options taking priority) and for strong video shooting capability (with the FX30 as a supplement, trying new styles).
- My view on on-device AI: for AI to become widespread, it needs to land on-device; the iPhone hardware is ready. I don’t have high hopes for Apple Intelligence, and I want to fine-tune 4B/8B models, hoping for good iOS clients that support both local running and APIs.
iPhone 17 Pro can actually run an 8B large language model! Can you believe that?
Look at this: this is Qwen3-8B. I downloaded it from Hugging Face with PocketPal, in Q4 precision. This is the test I ran last night, and the speed reached 12 tokens per second.
Let me show you the actual inference speed. I asked it a simple question: why is the sky blue? See that? It’s pretty smooth, right?
What does it mean to be able to run an 8B model on a phone?
You have to know that before this, consumer-grade devices on mobile usually ran 1.5B or 3B models; on desktop, 7B and 8B models were the starting point, and being able to run up to 70B was already very impressive.
Now, the iPhone 17 Pro has actually brushed up against the tail end of desktop-class performance. That’s something I didn’t expect before I got the new phone.
The one I have is the 512G Pro, Hong Kong version. I found a reseller on Taobao and paid 11,750. I bought it that day, and it was sent out by SF Express that same evening.
Actually, I had already been using this iPhone 12 mini for many years. There were mainly two reasons I decided to upgrade this time.
The first was for eSIM. Just like I said on the newtype Knowledge Planet, if some eSIM services can be activated and used directly in mainland China, then that gives you one more option on the network side — and I care a lot about that. I firmly believe:
Over a ten-year horizon, network, payments, and identity are very fundamental elements. Their priority is very high, and they absolutely need to be handled well.
In the end, I was very satisfied with the results of this Hong Kong version.
I’m using Red Tea’s eSIM. Once I turn on data roaming, it works directly. Activation is also very simple. I posted screenshots from that time on the newtype Knowledge Planet. Scanning the QR code with the camera automatically jumps to the activation page, and it’s done in a minute.
I’m not convenient to go into detail on this topic. Its value, those who know will know.
The second reason was for shooting video.
My main setup now is the FX30 I’m currently using, paired with three Sigma lenses. The one I’m using right now is the 10-18 F2.8. Although the FX30 is only an entry-level cinema camera, its image quality is already more than enough.
That said, it’s fine for fixed-position shooting. If I want to use it to shoot other kinds of scenes, it’s always a bit inconvenient. Just at this time, the iPhone 17 Pro’s video capability is extremely strong. As long as it’s not a professional camera, then in terms of image quality, it’s the best right now. It’s very suitable for me to try more content styles. Honestly, after sitting here chatting nonsense for a hundred episodes, I was already sick of it.
And being able to run an 8B large language model is an unexpected bonus.
I still stand by the view I had more than a year ago: if AI is going to become widespread, it absolutely needs to land on-device. You can’t possibly meet such broad demand by relying on cloud compute alone.
But unfortunately, I’ve been waiting for almost two years, and on-device AI still isn’t very compelling. It has clearly fallen behind the broader pace of development.
The performance of the chip in the iPhone 17 Pro is a very good sign, because it shows that hardware is fully ready. Running large language models locally is no longer just at the level of “it can run”; it can run efficiently and stably.
I suddenly got the urge to fine-tune models. I want to create a few custom 4B or 8B models based on my own needs, and both the iPhone and Mac can run them locally.
What’s more, the training datasets can be continuously iterated. Over the long run, this accumulation will be very valuable, just like notes.
Now that the hardware is no longer a problem, it all comes down to software and the system. I’m not counting on Apple Intelligence. What I want now is for someone to hurry up and release a better client that supports both local execution and API calls.
Once I see an iOS client I’m satisfied with, I’ll make another video to introduce it.
OK, that’s all for this episode. If you want to learn about AI, want to become a super-individual, and want to find like-minded people, come join our newtype community. See you next time!