DeepSeek V4 Flash is a 284.33B-parameter MoE.
๐คฏ This shouldn't work ๐ DeepSeek V4 Flash on 7.7GB RAM. No GPU, and instead using an increasingly popular and more affordable technique.
DeepSeek V4 Flash is a 284.33B-parameter MoE.
The GGUF used here is 78.62 GiB.
The machine?
๐พ 8 GB RAM
๐ฎ ZERO GPU / CUDA
๐ง CPU-only
๐ฟ NVMe SSD
๐ฆ 78.62 GiB model
๐งฎ 284.33B parameters
They got the model
to execute.
Even crazier โก Controlled cold first-token diagnostic: 5.33 sec ๐ง Maximum RSS: ~5.9 GiB ๐ Process swaps: 0 โ ๏ธ No tps reported How? The whole 79GB model never goes into RAM.
The GGUF stays memory-mapped on the NVMe SSD, and Linux demand-pages the pieces of the model into RAM as theyโre needed.
๐ฌ NVMe becomes backing storage for model weights โ RAM becomes a moving working set. This is NOT fast.
The developer explicitly says they are not claiming competitive sustained decode speeds, and the 5.33-second result is a controlled first-token diagnostic, not normal chat latency.
๐ฏ So the fact that the model doesn't fit in RAM is more of a performance problem rather than an NVRAM barrier.
If smart paging + MoE expert loading + faster NVMe keeps improving, the SSD may become another tier in the Local AI memory hierarchy:
VRAM โ RAM โ NVMe