DGX Spark的新推理服务器:大模型 C4: 55-90 令牌/秒,无规格解码

3 分•作者: medicis123•3 个月前•原帖
大家好,我们非常高兴地分享我们新构建的推理服务器的性能数据和基准报告,该服务器专门用于在DGX Spark集群上运行多模型智能工作流。我们在一个由2个DGX Spark节点组成的集群上进行了LlamaBench测试以及我们自己的模拟流量测试,得到了非常不错的结果。以下是详细信息。完整的详细报告可以在我们的WoolyAI网站上找到。 WoolyAI私有多智能体推理堆栈为DGX Spark而构建,旨在使企业内部的团队能够建立自己的私有、低成本的推理堆栈,以支持业务智能工作流应用。它是基于前瞻性的愿景构建的,认为公司将需要自己的私有、低成本的推理设置,而单模型推理不足以满足复杂企业工作流智能应用的需求。这些工作流需要多个不同专业(因此大小不同)的模型来处理工作流中的不同步骤。为每个模型配备专用的多GPU推理堆栈成本非常高昂。 第一次基准测试:我们首先在一个由2个DGX Spark节点组成的集群上对我们的推理服务器进行了LlamaBenchy测试,涉及3个模型,没有量化和推测解码。使用推测解码将会得到更高的性能数据。 模型 加载-预填充-系统-令牌/解码-系统-令牌/每请求解码-令牌 - DeepSeek V4 Flash C1 1,518.91 21.15 21.15 - DeepSeek V4 Flash C4 1,533.15 55.99 14.00 - Gemma 4 26B A4B C1 4,579.73 30.22 30.22 - Gemma 4 26B A4B C4 4,702.16 63.75 15.94 - Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42 - Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71 第二次基准测试:一个端点,三种不同模型控制的激活。调度程序对每次突发进行批处理,协调两个排名,并仅在安全边界处更改驻留模型。 模型 预填充-令牌/解码-令牌/模型-激活-等待 - DeepSeek V4 Flash 4,154.34 49.30 16秒 - Gemma 4 26B A4B 18 4,781.44 64.67 6秒 - Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2秒 我们认为通过进一步优化可以将这些数据提高20%。请分享您的反馈。https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/
查看原文
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.<p>WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.<p>First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.<p>MODEL LOAD-PREFILL-SYSTEM-TOK&#x2F;S DECODE-SYSTEM-TOK&#x2F;S DECODE-TOK&#x2F;S-PER-REQUEST<p>DeepSeek V4 Flash C1 1,518.91 21.15 21.15<p>DeepSeek V4 Flash C4 1,533.15 55.99 14.00<p>Gemma 4 26B A4B C1 4,579.73 30.22 30.22<p>Gemma 4 26B A4B C4 4,702.16 63.75 15.94<p>Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42<p>Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71<p>Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary<p>MODEL PREFILL-TOK&#x2F;S DECODE-TOK&#x2F;S Model-Activation-Wait<p>DeepSeek V4 Flash 4,154.34 49.30 16s<p>Gemma 4 26B A4B 18 4,781.44 64.67 6s<p>Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s<p>We think we can improve these numbers by 20% with more optimization. Please share your feedback. https:&#x2F;&#x2F;woolyai.com&#x2F;ai-compute-software&#x2F;dgx-spark-inference-stack&#x2F;