DeepSeek V4 Flash is a 284.33B-parameter MoE.

avatar

๐Ÿคฏ This shouldn't work ๐Ÿ‘‰ DeepSeek V4 Flash on 7.7GB RAM. No GPU, and instead using an increasingly popular and more affordable technique.

DeepSeek V4 Flash is a 284.33B-parameter MoE.

The GGUF used here is 78.62 GiB.

The machine?

๐Ÿ’พ 8 GB RAM

๐ŸŽฎ ZERO GPU / CUDA

๐Ÿง  CPU-only

๐Ÿ’ฟ NVMe SSD

๐Ÿ“ฆ 78.62 GiB model

๐Ÿงฎ 284.33B parameters

They got the model

to execute.

Even crazier โšก Controlled cold first-token diagnostic: 5.33 sec ๐Ÿง  Maximum RSS: ~5.9 GiB ๐Ÿ”„ Process swaps: 0 โš ๏ธ No tps reported How? The whole 79GB model never goes into RAM.

The GGUF stays memory-mapped on the NVMe SSD, and Linux demand-pages the pieces of the model into RAM as theyโ€™re needed.

๐Ÿ’ฌ NVMe becomes backing storage for model weights โ†’ RAM becomes a moving working set. This is NOT fast.

The developer explicitly says they are not claiming competitive sustained decode speeds, and the 5.33-second result is a controlled first-token diagnostic, not normal chat latency.

๐ŸŽฏ So the fact that the model doesn't fit in RAM is more of a performance problem rather than an NVRAM barrier.

If smart paging + MoE expert loading + faster NVMe keeps improving, the SSD may become another tier in the Local AI memory hierarchy:

VRAM โ†’ RAM โ†’ NVMe



0
0
0.000
0 comments