Large Unified Memory PCs for Local AI: What the Extra Memory Buys
Desktops with large unified memory let the GPU address most of system RAM, so they can load models far larger than a consumer graphics card's memory allows. That makes large open-weight models and long contexts practical locally. The trade-off is speed: memory bandwidth limits tokens per second, so these machines suit capacity-bound work more than high-throughput serving.
What unified memory changes
On a typical PC, a model must fit in the graphics card's dedicated memory to run quickly. Consumer cards carry a limited amount, which caps model size, and anything that spills into system RAM runs much more slowly.
Machines with unified memory share one large pool between CPU and GPU. Apple Silicon popularised this for local AI, and x86 desktops built on AMD's Strix Halo platform, such as the Framework Desktop, brought large unified memory configurations to PCs. Newer configurations push memory capacity higher still. The practical effect is that models and contexts which would never fit on a consumer GPU can be loaded entirely into fast-enough memory.
Capacity versus bandwidth
Two numbers decide local model performance. Capacity decides what fits: the model's weights at your chosen quantisation, plus memory for the context. Bandwidth decides speed: generating each token requires reading the active weights from memory, so tokens per second are roughly bounded by bandwidth divided by the bytes read per token.
Large unified memory systems are strong on capacity and moderate on bandwidth compared with high-end discrete GPUs. They run big models, but a dense model that uses most of the memory will generate text noticeably more slowly than a smaller model on a fast GPU.
Models that suit the hardware
Mixture-of-experts models are a good match. They have a large total parameter count, which needs capacity, but activate only a fraction of parameters for each token, which reduces the bytes read and therefore the bandwidth bottleneck. That combination plays to exactly what unified memory machines offer.
Dense models in the mid-size range also run comfortably, with room left for long contexts or for keeping several models loaded at once, such as a coding model, a general model, and an embedding model.
Who these machines suit
- Developers experimenting with large open-weight models who need capacity more than speed.
- Privacy-sensitive work where data should not leave the machine.
- Agent development that benefits from several local models running together.
- Long-context tasks such as working over large documents or codebases.
They suit high-volume serving to many users less well, where discrete GPUs or hosted inference deliver far more throughput per unit of cost.
Getting more from the hardware
On bandwidth-limited machines, every token of prompt and context costs time to process. Keeping prompts short pays off directly. Retrieval helps: instead of loading whole documents into a long context, index them and give the model only the relevant passages. That shortens prompt processing, leaves memory for the model, and often improves answer quality because the model is not searching a huge context for the relevant part.
Also test quantisation levels: a slightly more aggressive quantisation can raise speed noticeably with modest quality loss, depending on the model and task.
Software support
Hardware is only useful if the software stack supports it well. Check that the runtimes you plan to use, such as llama.cpp-based tools or other local inference engines, support the platform's GPU and memory model, and how mature that support is. Driver and runtime updates can change performance substantially, so look at recent measurements rather than launch-day reviews. Also confirm that tools you rely on, such as embedding models and agent frameworks, run on the same machine without extra friction.
Before buying
Write down the specific models and context lengths you intend to run, estimate their memory needs, and check measured tokens per second from people running those models on the hardware you are considering. Compare the cost with hosted inference for your expected usage. A machine sized for what you actually run beats one sized for headlines.
Frequently asked questions
- Why is unified memory good for local AI?
- Because the GPU can use most of the system's memory, so models far larger than a consumer graphics card's dedicated memory can be loaded and run without spilling into slow system RAM. That makes large open-weight models and long contexts practical on a single desktop.
- Are unified memory machines faster than GPUs for LLMs?
- Usually not for models that fit on a high-end GPU. Discrete GPUs typically have higher memory bandwidth, which drives token generation speed. Unified memory machines win on capacity, running models that would not fit on a GPU at all, while generating more slowly than a fast GPU on smaller models.
- What models run well on large unified memory?
- Mixture-of-experts models suit them well, because their large total size needs capacity while only a fraction of parameters is used per token, easing bandwidth limits. Mid-size dense models also run comfortably with room for long contexts or several models loaded at once.
- How much memory do I need to run a 70B model locally?
- It depends on quantisation and context length. As a rough guide, weights need about the parameter count multiplied by the bytes per parameter, so a 70 billion parameter model at 4-bit quantisation needs on the order of 35 to 40 GB before context memory. Long contexts add several more gigabytes, so leave generous headroom.
- Is a unified memory desktop better than a Mac for local AI?
- They target a similar niche: large memory pools usable by an integrated GPU. Differences come down to memory bandwidth, maximum memory configurations, software support for your chosen runtimes, operating system preference, and price. Compare measured tokens per second for the specific models you plan to run rather than headline specifications.