ArticleAugust 25, 2026Free to read

What Hardware Do You Need to Run Qwen 3.8 27B Locally?

Understand memory capacity and bandwidth, then compare the tradeoffs of Apple Silicon, RTX 5090, and RTX PRO 6000 for local model inference.

Originally published . English translation: . Read the Chinese original.

Original video in Chinese.

Key Takeaways

  • Capacity determines the model size and precision you can run; bandwidth affects generation speed. In the configurations discussed here, a 27B model at 4-bit precision starts at roughly 24GB of memory. Large capacity does not necessarily mean high bandwidth.
  • Apple Silicon offers relatively affordable unified memory with lower bandwidth. An RTX 5090 offers 32GB of fast memory for a single Qwen 3.8 27B instance, or two independent instances for concurrent agents. The 96GB RTX PRO 6000 suits long contexts and multiple resident models but costs much more.
  • Qwen 3.8 27B represents what I call a DeepSeek moment for open models: strong performance in a compact model that individuals can deploy as a practical productivity tool.

Qwen 3.8 27B feels like a DeepSeek moment for open models. It performs well, has a compact size, and can deliver useful results on hardware owned by an individual.

Over several days, I researched local deployment options and posted the process in the newtype community. Here is my overview, starting with two important measures.

Memory capacity

The model’s weights must fit into GPU memory or unified memory. If they do not, the model either cannot run or must place some weights in system memory and access them through the CPU or PCIe, greatly reducing speed.

Capacity determines how large a model you can run and at what precision.

For a 27B model, BF16 weights alone require about 54GB. Quantization to 8-bit brings that to around 30GB; 4-bit brings it to roughly 17–20GB.

The operating system, inference framework, and context also consume memory. A 4-bit model therefore needs more than 17GB: approximately 24GB is a starting point for barely fitting it in this example.

Memory bandwidth

Once the model fits, bandwidth becomes the next key measure. Capacity determines whether it can run; bandwidth affects how quickly it generates.

NVIDIA’s DGX Spark has 128GB of memory, more than the RTX PRO 6000’s 96GB. That can look like good value. But the bandwidth figures discussed here are 273GB/s for DGX Spark and 1,792GB/s for RTX PRO 6000.

Lower bandwidth means that although DGX Spark can fit a large model, generation with a large dense model may be slow. It is better suited to MoE models, which activate only a subset of their total parameters for each token.

One community member bought two units to run DeepSeek V4 Flash and reported a speed above 30 tokens per second, enough for their needs.

Capacity and bandwidth are the most important starting points, though actual performance also depends on GPU compute, architecture, and the inference framework.

Option one: Mac

Apple Silicon uses unified memory: the CPU and GPU share a memory pool. This lets you obtain more memory for a relatively lower price.

Its drawback is bandwidth, which may be half that of a 5090 or 6000, or lower. Even with equal or greater capacity, a Mac may generate more slowly on dense models. That is a tradeoff to consider.

Option two: RTX 5090

The 5090 configuration discussed here has 32GB of GDDR7 memory and 1,792GB/s bandwidth, making it well suited to a model such as Qwen 3.8 27B.

If your sole objective is to run Qwen 3.8 27B locally, I consider a single 5090 generally more suitable than a Mac Studio.

For concurrent agents, you can use two 5090s. The 5090 has no NVLink, so communication across cards goes through PCIe and does not behave like a single native 64GB card.

The purpose of two cards here is to run a complete model on each: one card serves some agents, and the other serves the rest. This particularly suits high concurrency with short or medium contexts.

The capacity limit is 32GB. A 27B model may fit, but a 70B model, very long context, or several resident models can exceed it.

Option three: RTX PRO 6000

This is the most complete local option discussed here, with 96GB of ECC GDDR7 memory and 1,792GB/s bandwidth.

The combination of capacity and bandwidth is well suited to long contexts, multiple resident models, and production agent services. ECC and professional drivers also support long-running use.

The major drawback is the price. These three options span from tens of thousands of yuan to well over a hundred thousand or around two hundred thousand. No hardware wins on every dimension; you must make tradeoffs.

What hardware are you using for local models, and how well does it work? Share your experience in the comments.

That is all for this episode. If you want to understand AI, reclaim personal sovereignty, and meet people with similar interests, join the newtype community. See you next time.