展示HN:通过打破DDR4时序规则在DRAM中运行PrismML的Bonsai
围绕PrismML的1位/三元Bonsai模型的兴奋情绪使得业界密切关注智能手机巨头,尤其是苹果公司,如何在边缘设备上实现大语言模型(LLMs)。<p>将人工智能移至设备端是一项聪明且必要的策略。这不仅确保了用户隐私符合欧盟法规,还从根本上改变了高昂的云推理经济模式,并为用户寻求真正具备人工智能能力的硅芯片铺平了道路,开启了显著的硬件升级超级周期。<p>为了创建一个智能的设备端“语义路由器”,模型需要达到超过270亿参数的规模。在手机上实现这一点需要极端量化,例如PrismML的三元权重。<p>然而,软件界常常忽视的一个关键硬件现实是:将权重放入RAM并不等同于将其移动。 在标准LPDDR上运行一个270亿的三元模型会遇到显著的内存带宽限制。每次生成令牌时在系统芯片(SoC)总线之间传输数GB的数据可能导致神经处理单元(NPU)过热降频和过度耗电。<p>这引发了一个重要问题:我们为什么还在将数据传输到计算单元?为什么不在内存中本地执行人工智能推理?<p>由于对忽视裸金属物理的学术PIM模拟感到沮丧,我开发了CaSA,这是一种通过电荷共享直接在商用DRAM中执行三元LLM推理的架构,完全绕过了内存总线。<p>软件量化是一个很好的初步步骤,而CaSA提供了完成这一桥梁所需的物理硬件基础:<a href="https://github.com/pcdeni/CaSA" rel="nofollow">https://github.com/pcdeni/CaSA</a>
查看原文
The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices.<p>Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon.<p>To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a phone requires extreme quantization, such as PrismML’s ternary weights.<p>However, a critical hardware reality often overlooked by the software world is that fitting the weights in RAM is not equivalent to moving them. Running a 27B ternary model on standard LPDDR encounters a significant memory bandwidth limitation. Transferring gigabytes of data across the SoC bus for each token generation can lead to thermal throttling of the NPU and excessive battery drain.<p>This raises an important question: why are we still transferring data to the compute? Why not execute AI inference natively within the memory?<p>Frustrated with academic PIM simulations that overlook bare-metal physics, I developed CaSA, an architecture that performs ternary LLM inference directly inside COTS DRAM through charge-sharing, completely bypassing the memory bus.<p>Software quantization is a great initial step, and CaSA provides the physical hardware substrate needed to complete the bridge: <a href="https://github.com/pcdeni/CaSA" rel="nofollow">https://github.com/pcdeni/CaSA</a>