Used EPYC Servers for CPU-Only MoE Inference in 2026
A used 8-channel EPYC server with 512GB of registered DDR4 will run a 300-460GB mixture-of-experts model that no consumer GPU can touch. Reported speeds sit at 4-6 tokens per second. The pitch you will read elsewhere is that used server RAM escaped the 2026 memory shortage. It did not escape; it escaped less. DDR4 ECC currently tracks at a median of $8.66 per gigabyte, so the 512GB that makes this build interesting costs about $4,400 on its own.
Bottom Line
- It works, and it is slow. Reported CPU-only speeds on 8-channel EPYC: 4.2 tok/s on DeepSeek-R1 Q5_K_S (461.81GB), ~5.2 tok/s on GLM-5 Q3_K_XL (309GB). Usable for batch work, painful for chat.
- Memory channels beat cores, and it is not close. Bandwidth is the constraint, not core count. Buy DIMMs before you buy cores.
- CORRECTED 2026-08-25. This page originally read two cross-machine reports — a dual 96-core 7K62 at 2.9 tok/s against a single 48-core 7K62 at 4.2 tok/s — as proof that dual socket is slower. Those are two different machines, so the comparison cannot isolate the socket count. A same-machine A/B shows the second socket helping: 1.83x on a dense 70B, and 1.02x on DeepSeek R1. One socket is still the right buy for MoE work, for a better reason. See dual-socket vs single-socket EPYC for LLM inference.
- The cheap-RAM premise is half true. DDR4 ECC RDIMM tracked at a $8.66/GB median against $36.20/GB for DDR5 when checked on 24 August 2026. That is a 4.2x relative advantage, but 512GB still costs roughly $2,300 at the cheapest listing and about $4,400 at the median.
- The CPU is the cheap part. Used EPYC 7402 listings around $89; a used EPYC 7513 32-core around $379. The RAM costs five to fifty times the processor.
- This only makes sense for models that fit nowhere else. If it fits in 24GB, a used RTX 3090 at $1,000-1,300 is cheaper and roughly six times faster.
- Prompt processing is the hidden cost. Generation at 5 tok/s is tolerable. CPU prompt processing on a long context is not, and no build guide mentions it.
- Plan for noise and a power bill. Server chassis fans and 200W-class CPUs are not home-office equipment.
Why Anyone Considers This
The models that matter most in 2026 stopped fitting on consumer hardware. As we set out in open weights are not local any more, frontier open-weight releases are now mixture-of-experts models in the 300GB to 700GB range at usable quantisation. No consumer GPU holds that. A 96GB workstation card does not hold that.
But MoE models have a property that rescues CPU inference: only a fraction of the parameters activate per token. A 671B model with roughly 37B active parameters does 37B-worth of arithmetic per token while needing all 671B resident in memory. That is a memory-capacity problem, not a compute problem, and system RAM is the cheapest capacity you can buy per gigabyte.
Hence the pitch: a used server with eight memory channels and 512GB of registered DDR4. We cover the small end of this idea in running a local LLM on 128GB of RAM with no GPU. This page is the large end, with the honest arithmetic.
The Measured Speeds
These figures come from user reports in the llama.cpp CPU-inference discussion. They are self-reported by individual owners, not a controlled benchmark, so read them as an order of magnitude rather than a spec sheet.
| System | Memory | Model | Speed |
|---|---|---|---|
| EPYC 7K62, 48 cores | 8×64GB (8-ch) | DeepSeek-R1 Q5_K_S, 461.81GB | 4.2 tok/s |
| Dual EPYC 7K62, 96 cores | 16×64GB | DeepSeek-R1 Q5_K_S, 461.81GB | 2.9 tok/s |
| Threadripper Pro 3955WX, 16 cores | 8×64GB | DeepSeek-R1 Q5_K_S, 461.81GB | 2.8 tok/s |
| EPYC 7552, 48 cores | 512GB DDR4-2666 (8-ch) | GLM-5 Q3_K_XL, 309GB | ~5.2 tok/s (24 threads) |
| EPYC 9654 (DDR5-4800) | 12-ch DDR5 | 671B Q8 | 6.2 tok/s |
Three things in that table deserve attention.
Row two carries a correction, added 2026-08-25. The dual-socket machine has twice the cores and twice the DIMMs, and it reports 31% slower than the single-socket machine on the identical workload. We originally concluded from that pair that dual socket is slower. That conclusion does not survive a controlled test. These are two different machines owned by two different people, configured differently, so the comparison cannot isolate the socket count. When a single machine is measured with one socket and then with both, the second socket is faster every time — although on DeepSeek R1 it is faster by only 2.2%. The full numbers, the mechanism, and the NUMA placement fix that recovers about 80% on a badly configured dual-socket box are in dual-socket vs single-socket EPYC for LLM inference. Buy one socket for MoE work — but buy it because the second socket returns almost nothing on MoE, not because it makes the machine slower.
Row three shows cores are not the lever either. A 16-core Threadripper Pro with the same 8 channels lands at 2.8 tok/s against 4.2 for a 48-core EPYC. There is some core scaling, but the EPYC 7552 report used only 24 of its 48 threads and still hit 5.2 tok/s. Past a point, adding threads adds contention, not speed.
Row five is the ceiling. Twelve channels of DDR5-4800 buys 6.2 tok/s. Even the modern, expensive version of this machine is not fast. Do not expect a newer platform to change the category.
The Real Cost of the RAM
This is where the topic usually gets oversold, so here are the numbers as checked on 24 August 2026.
A live server-memory tracker showed DDR4 ECC RDIMM at a $8.66/GB median and DDR5 at a $36.20/GB median, with the cheapest DDR4 listing at $4.53/GB for a 16GB DDR4-2133 RDIMM. Separately, refurbished-server-parts coverage of the 2026 market reports that refurbished DDR4 has itself risen 30-50%, while a new 64GB DDR5 RDIMM went from about $255 to over $900 in a year.
Applied to the capacity that makes this build worth doing:
| Capacity | At cheapest listing ($4.53/GB) | At median ($8.66/GB) | Same capacity in DDR5 ($36.20/GB) |
|---|---|---|---|
| 256GB | ~$1,160 | ~$2,220 | ~$9,270 |
| 512GB | ~$2,320 | ~$4,430 | ~$18,530 |
| 1TB | ~$4,640 | ~$8,870 | ~$37,070 |
Say the honest version out loud: the DDR5 column is why people call DDR4 cheap. Against DDR5 it looks like a bargain. Against a $1,150 used RTX 3090 it does not. Buying 512GB at the median price costs roughly what four used 3090s cost, and four 3090s give you 96GB of memory running at more than ten times the bandwidth.
Two practical notes on sourcing. The prices above come from a scraped tracker and a parts-reseller write-up, which is the weakest source class we use, so treat the ranges as indicative and price your actual DIMM configuration before committing. And the cheapest listing is a 16GB module: filling 8 channels to 512GB at $4.53/GB means 32 sticks, which most single-socket boards cannot hold. In practice you buy 64GB modules and pay closer to the median.
What the Rest of the Machine Costs
The processor is genuinely cheap, which is the part of the pitch that survives scrutiny. Recent listings show a used EPYC 7402 24-core around $89 and a used EPYC 7513 32-core around $379. Supermicro H12SSL-I boards for the 7002/7003 generation list around $1,178 new in Canadian dollars, with used and bundled options below that.
So a rough single-socket build, with every figure treated as indicative:
- CPU: $89-379 used
- Motherboard: several hundred used, roughly $900 US equivalent new
- 512GB registered DDR4: $2,300-4,400
- Chassis, PSU, cooling, storage: several hundred
The RAM is 60-80% of the bill. Every optimisation that matters is a memory-buying decision, not a compute decision.
When To Build This, and When Not To
Build it when all of these are true:
- The model you need is genuinely 200GB or larger at a quantisation you accept.
- It is a mixture-of-experts model with a small active-parameter fraction. A dense model of that size will be several times slower again.
- Your workload is batch or asynchronous. Overnight document processing, agent runs you do not watch, evaluation sweeps.
- You can put a loud machine somewhere that is not your office.
Do not build it when:
- The model fits in 24GB or 48GB. Buy a card. It is cheaper and faster. Our 48GB setup guide covers that tier.
- You want interactive chat. At 5 tok/s a 500-token answer takes over 90 seconds, and that ignores prompt processing.
- Your context is long. CPU prompt processing scales badly and is the part these reports mostly do not measure. Assume it is worse than you hope, and test it before you buy.
- You were sold on “cheap RAM.” Re-read the table above. It is cheaper than DDR5 and it is not cheap.
Before you spend $3,000-5,000 on an EPYC build, check whether your model fits in 24GB. A used RTX 3090 costs $1,000-1,300 as of August 2026 and runs a 27B model at roughly 30 tok/s — about six times the speed of the server, for a quarter to a third of the price. The server wins only on capacity. We do not stock affiliate links for used EPYC parts, and we are not going to invent one; source those from eBay or a server-parts reseller directly.
Amazon affiliate link — we earn a small commission at no cost to you.
Making It As Fast As It Can Be
If you build it, these are the levers that actually move the number:
- One socket, all channels filled. Eight DIMMs on an 8-channel board. A half-populated board halves your bandwidth and therefore halves your speed.
- Fastest DIMM speed the platform supports. DDR4-3200 over DDR4-2666 where the CPU allows it. This is straight bandwidth.
- Fewer threads than you expect. The EPYC 7552 report used 24 of 48 threads. Sweep the thread count; the best value is often around half the physical cores.
- Use the MoE offload flags. Even a small GPU can hold the attention layers and the KV cache while system RAM holds the experts. Our guide to llama.cpp MoE offload flags covers the exact arguments, and running a 160GB MoE on 8GB of VRAM shows how far that goes.
- Quantise harder than feels comfortable. Q3 versus Q5 is the difference between fitting in 256GB and needing 512GB, and at these prices that is thousands of dollars.
See Also
- Dual-Socket vs Single-Socket EPYC for LLM Inference — the controlled A/B that corrects this page, and the NUMA fix
- Can I run a local LLM on 128GB of RAM with no GPU? — the same idea at consumer scale
- Running a 160GB MoE model on 8GB of VRAM — the hybrid alternative to going CPU-only
- Open weights are not local any more — why these models stopped fitting on a GPU
- llama.cpp MoE offload flags explained — the exact arguments for expert offload
- Dense 70B vs MoE 120B for local use — why active parameters decide the speed
- Should I buy RAM now in the 2026 shortage? — the market this build is priced into
- What your local AI rig will be worth in two years — server parts liquidate very differently from GPUs
Sources
- Inference LLM DeepSeek-V3 671B on CPU only (llama.cpp discussion #11765) — all reported tok/s figures; self-reported by individual users
- Server RAM prices, DDR4 ECC live $/GB (DatacenterDisk) — $8.66/GB DDR4 median, $36.20/GB DDR5 median, $4.53/GB cheapest listing, read 2026-08-24; scraped tracker, indicative
- Refurbished DDR4 server memory and the 2026 RAM price crisis (PCSP) — reseller write-up, indicative
- EPYC processor listings (eBay) — used CPU and board prices, checked August 2026
Used RTX 3090 pricing is our own August 2026 price reference, compiled from eBay-derived trackers. The 30 tok/s comparison figure is the measured 27B result from the RTX 3090 power-limit sweep, not a like-for-like test against the server models above; it is there to show the order-of-magnitude difference, not to benchmark the same workload.
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session