问HN:HotPin – 在24GB内存(CPU,50行代码)上进行无损120B MoE推理
我是一名机电一体化设计师,拥有控制系统、机器人技术、PCB设计和嵌入式硬件的背景。我设计物理系统:电机、传感器、微控制器和实时控制回路。
我将这种设计思维应用于大型语言模型(LLM)的内存管理——结果非常成功。
HotPin 是一组针对 llama.cpp 的补丁,可以在比其磁盘占用更少的 RAM 上运行 30B–120B 的专家混合(MoE)模型,并且输出与原始模型完全相同(无损)。
在 AMD Ryzen AI 9 HX 370(Zen5,AVX512)、23.6GB LPDDR5X、NVMe >1GB/s 的 CPU 环境下进行了测试。
结果如下:
| 模型 | 磁盘 | 最小内存 | 节省 | 速度(tok/s) |
|-------|------|---------|---------|-------|
| gpt-oss:120b | 58.5GB | 19.1GB | -67% | 3.84 |
| qwen3:30b-a3b | 18.0GB | 10.4GB | -42% | 19.7 |
| gemma4:26b-a4b | 16.2GB | 10.6GB | -35% | 11.5 |
| GLM-4.7-Flash | 19.0GB | 13.3GB | -30% | 12.4 |
输出与全内存运行的 SHA-256 位完全相同,已验证。
工作原理(在 llama.cpp 中约 50 行 C++ 代码):
1. 记录 MoE 专家路由频率。
2. 从磁盘映射整个模型。
3. 仅将最活跃的专家锁定到物理 RAM 中。
4. 使用 posix_fadvise / 预取冷专家到 NVMe 中,确保在需要之前加载。
边界条件:如果磁盘 > RAM,锁定可以带来 +45% 的速度提升(gpt-oss: 2.64 → 3.84 tok/s)。如果模型可以完全放入 RAM,锁定则不会增加任何开销。
在以下环境中进行了测试:Linux 原生、WSL、Windows 原生(VirtualLock)。
代码库地址:https://github.com/LozzKappa/hotpin-llm
论文(PDF + LaTeX)在代码库中,arXiv 提交正在等待 cs.LG 的认可。
我正在寻找:
1. 一位 arXiv 认可者(cs.LG)——如果您是一名研究人员,对这项工作感兴趣,请与我联系。
2. 希望从任何在其硬件上测试此技术的人那里获得反馈。
该技术简单、无损,并且今天就能使用。请测试并告诉我您的基准结果。
感谢您的阅读。
查看原文
I'm a mechatronics designer with a background in control systems, robotics, PCB design, and embedded hardware. I design physical systems: motors, sensors, microcontrollers, and real-time control loops.<p>I applied this design thinking to LLM memory management – and it worked.<p>HotPin is a set of patches for llama.cpp that runs 30B–120B Mixture of Experts (MoE) models on far less RAM than their disk footprint, with bit-identical (lossless) output.<p>Tested on an AMD Ryzen AI 9 HX 370 (Zen5, AVX512), 23.6GB LPDDR5X, NVMe >1GB/s, CPU-only.<p>Results:
| Model | Disk | Min RAM | Savings | tok/s |
|-------|------|---------|---------|-------|
| gpt-oss:120b | 58.5GB | 19.1GB | -67% | 3.84 |
| qwen3:30b-a3b | 18.0GB | 10.4GB | -42% | 19.7 |
| gemma4:26b-a4b | 16.2GB | 10.6GB | -35% | 11.5 |
| GLM-4.7-Flash | 19.0GB | 13.3GB | -30% | 12.4 |<p>Output is SHA-256 bit-identical to full-RAM runs. Verified.<p>How it works (~50 lines of C++ in llama.cpp):
1. Profile MoE expert routing frequencies.
2. mmap the entire model from disk.
3. mlock only the hottest experts into physical RAM.
4. posix_fadvise / prefetch cold experts from NVMe before they're needed.<p>Boundary condition: if Disk > RAM, pinning gives +45% speedup (gpt-oss: 2.64 → 3.84 tok/s). If model fits in RAM, pinning adds zero overhead.<p>Tested on: Linux native, WSL, Windows native (VirtualLock).<p>Repo: https://github.com/LozzKappa/hotpin-llm
Paper (PDF + LaTeX) in the repo – arXiv submission pending endorsement in cs.LG.<p>I'm looking for:
1. An arXiv endorser (cs.LG) – if you're a researcher and this work interests you, please reach out.
2. Feedback from anyone who tests it on their hardware.<p>The technique is simple, lossless, and works today. Test it and tell me your benchmarks.<p>Thanks for reading.