
TL;DR: A 128GB AI PC shares one pool of memory between the processor and the GPU, so an AI model never gets the full 128GB. Software sets the GPU’s share, through an operating-system rule, a chip maker’s setting, or both. Ordinary Windows has long capped it at half, AMD’s Windows setting reserves up to 96GB, and Microsoft has raised the limit for its new Surface Laptop Ultra without saying how far. OpenAI’s gpt-oss-120b needs about 65GB, right at the old half-of-memory line and well inside AMD’s 96GB.
Microsoft opened pre-orders for the Surface Laptop Ultra on October 7, built on NVIDIA’s RTX Spark chip. The laptop ships from October 16, with the 128GB model priced at $5,899.99. The fine print on the laptop’s product page reads: “128 GB total system memory in a unified memory architecture shared between CPU and GPU. Maximum amount addressable by the GPU depends on system configuration and workload and is less than the total.”
So how much of the 128GB can the GPU actually use, and will a 65GB model such as OpenAI’s gpt-oss-120b fit? AMD’s Ryzen AI Max+ 395 mini PCs, NVIDIA’s DGX Spark and Apple’s Macs raise the same question, and the local AI mini PC guide compares them.
Where a 128GB AI PC’s memory goes
A local AI memory budget has three parts, and all three come out of the same 128GB.
The first is what the operating system and every other program keep. Whatever the GPU is given is taken from them. AMD says so plainly: memory assigned to its GPU “will no longer be available for use as system RAM.”
The second is the model itself. A model’s size is given in parameters, the numbers it learned in training, so “120B” means 120 billions of them. Models run locally are usually stored at around 4 bits, about half a byte, per parameter. On that rough rule a 120B model comes to about 60GB and a 70B model to about 35GB. Real downloads run somewhat larger: the 4-bit Llama 3.3 70B file in the Q4_K_M format is 42.5GB.
The third is the memory a model uses while it works. The largest piece is the KV cache, the model’s running record of the conversation so far. It grows with every token the model reads or writes, and a token is roughly three-quarters of a word by OpenAI’s rule of thumb. A short chat adds little. A long document pasted into the prompt can add nearly as much as the model weighs. The software running the model also needs working space. For gpt-oss-120b, OpenAI’s downloadable model of about 117 billion parameters, llama.cpp’s guide lists 2.7GB (llama.cpp is a popular free program for running models locally).
The GPU’s share has to cover the second and third parts together.
How much each platform lets the GPU use
| Machine | GPU memory limit on a 128GB machine | What kind of limit it is |
|---|---|---|
| Ordinary Windows PC with shared graphics | Half, so 64GB | Windows cap on shared memory; recent Intel chips allow an override |
| Surface Laptop Ultra (NVIDIA RTX Spark chip) | More than half, less than 128GB; not published | Raised Windows cap |
| AMD Ryzen AI Max+ 395 on Windows | 96GB reserved, up to 112GB in total | Reserved in AMD’s app or BIOS, plus shared memory; needs a restart |
| AMD Ryzen AI Max+ 395 on Linux | About half by default; AMD’s example raises it to 100GB | Driver limit, changed with an AMD tool |
| NVIDIA DGX Spark | No fixed figure; whatever the processor isn’t using | Handed out on demand |
| Mac with 128GB | Not published by Apple | macOS limit; raised with one terminal command |
Because these limits work differently, the numbers compare loosely. Windows caps shared memory at any moment, AMD sets memory aside in advance, and DGX Spark hands it out as programs ask.
The half rule is a long-standing Windows policy, and Intel’s current Windows 11 support page still lists shared graphics memory as “limited by OS to one-half of System Memory.” For RTX Spark machines like the Surface, Microsoft announced “a new higher, smarter limit on total system memory accessible by the GPU” in May. It has not said what the new limit is.
AMD’s Variable Graphics Memory setting reserves up to 96GB for the GPU, which can borrow more from the shared pool to reach 112GB. AMD adds that “the highest performance can be unlocked by limiting workloads to 96GB.”
What gpt-oss-120b and Llama 3.3 70B need
Both models are shown in their common 4-bit versions. For scale, 32K tokens is roughly 25,000 words and 131K tokens roughly 98,000.
| Model and conversation length | Model weights | Total memory needed |
|---|---|---|
| gpt-oss-120b, 32K tokens | 61.0GB | 64.9GB |
| gpt-oss-120b, 131K tokens (its maximum) | 61.0GB | 68.5GB |
| Llama 3.3 70B, 32K tokens | 42.5GB | About 53GB (estimated), plus working space* |
| Llama 3.3 70B, 131K tokens (its maximum) | 42.5GB | About 85GB (estimated), plus working space* |
*gpt-oss figures from llama.cpp’s guide. Llama figures are estimates that assume a 16-bit conversation record (8-bit roughly halves it) and exclude working space.
The smaller model ends up needing more memory at full length because of how each was built. Llama 3.3 70B keeps a detailed record of every token in each of its 80 layers, the stacked stages a model passes each word through. gpt-oss-120b has 36 layers, keeps a smaller record in each, and in half of them remembers only the last 128 tokens. Working from each model’s published design, that comes to about 330KB per token for Llama and about 37KB for gpt-oss. Over a 131K-token conversation, the difference is about 43GB of conversation memory against under 5GB.
Under AMD’s 96GB setting, everything in the table fits, though Llama 3.3 70B at full length only just does, at about 85GB before its working space.
Where the 120B claim stands
Microsoft claims the new Surface can run “up to 120B parameter models locally,” and that line carries a footnote about computing speed, citing NVIDIA’s theoretical FP4 figure measured with sparsity. Memory is a separate question. gpt-oss-120b needs about 65GB with a 32K-token conversation and about 68.5GB at its longest. Half of a 128GB machine is about 64GB, so the model sits right on the line ordinary Windows has used for years, with nothing to spare for a longer document or a second program. llama.cpp doesn’t say which kind of gigabyte it counts in, and that difference alone could tip the model to either side.
So the claim leans on Microsoft’s higher limit for RTX Spark machines, the one figure it has not published. AMD’s 96GB is a useful comparison, though it describes AMD’s design rather than Microsoft’s. Microsoft sells the $5,899.99 configuration partly on running large models locally, but buyers can’t check the memory arithmetic until that figure is known.
Loading a model versus running it
Speed is the next test, and gpt-oss-120b has an advantage here, because it is a mixture-of-experts model that uses about 5.1 billion of its roughly 117 billion parameters for each token. A dense model such as Llama 3.3 70B uses all of its parameters for every token, so it has far more to read from memory each time.
On a DGX Spark, one tester’s late-2025 llama.cpp benchmarks of gpt-oss-120b showed 46 to 61 tokens per second on a fresh conversation, depending on the software version, and 34 to 41 once 32K tokens were already in it. Testers on Framework’s community forum reported about 33 tokens per second in August 2025 for the same model on a Framework Desktop with AMD’s chip and 128GB.
Those figures are for one person. Every extra user needs a separate conversation record in memory. On Llama 3.3 70B, each additional 32K-token conversation adds about 11GB, so a box that comfortably serves one developer fills up quickly with several. How fast each user’s answers arrive then depends on how the software schedules the work. The GPU memory hierarchy explainer covers the same pressure on larger servers.
Before buying one
Ask where the GPU’s share is set and whether changing it needs a restart. On AMD machines it does. Once the machine arrives, load the model the team actually plans to use and set its context length to the longest conversation or document it will realistically see, since programs such as llama.cpp reserve the conversation memory for that length when the model loads. Then watch the peak memory use, since the download size leaves out the conversation and the working space.
On Windows, the GPU memory figure in Task Manager shows how close a model is to the GPU’s limit. On graphics cards, llama.cpp’s guide warns that going past it means “slow swapping to RAM and very bad performance.” How RTX Spark machines behave at their limit has not been measured yet. On a DGX Spark, NVIDIA warns that its usual memory check “may be smaller than the actual allocatable memory,” so loading the model is the better test.
The hardware being replaced belongs in the budget too. A team moving its local AI work off a workstation with a 24GB or 32GB graphics card can sell the used GPU toward the new machine, and the workstation’s memory can go with it. Server memory is still tight, and memory makers expect the memory shortage to last through 2028 or longer, although DDR4 prices have started to soften from their peak. That makes the RAM in a retired workstation, and in any servers retired with it, worth a resale quote before it goes into storage.
Whether the Surface Laptop Ultra clears the 65GB that gpt-oss-120b needs will show in the first reviews after October 16. The cleanest test is Llama 3.3 70B loaded at three context lengths: 32K tokens (about 53GB), 64K (about 64GB) and its full 131K (about 85GB), all estimates before working space. If speed collapses at 64K, the old half-of-memory rule is still in charge. If it holds at 131K, the raised limit is above 85GB. A collapse anywhere in between shows where Microsoft set it.