启动 HN:Speko(YC S26)——语音人工智能的开放路由器
嗨,HN!我是Bek,Speko的创始人。Speko是一个平台,可以在所有公开基准选项中,根据您的约束条件找到最佳的语音转文本(STT)、大型语言模型(LLM)和文本转语音(TTS)模型组合,并告诉您原因。
<p>演示:<a href="https://youtu.be/no2LY2gRh-c" rel="nofollow">https://youtu.be/no2LY2gRh-c</a>
<p>典型的生产语音代理由三个模型组成:STT、LLM和TTS。
<p>每一层都有十多个可信的供应商,每个月市场上都会出现新的模型。几乎每个人只评估一次,选择自己喜欢的堆栈,然后再也不检查,因为从一个供应商切换到另一个供应商涉及额外的集成和关于数据的争论。
<p>结果是,您使用的语音代理运行的是上个季度的模型,而更好、更便宜的选项已经可用。
<p>在创办Speko之前,我花了四年时间作为联合创始人和首席技术官,为亚洲的企业构建语音代理,支持10多种语言。每当有新的语音模型出现时,我们都会重复同样的流程:雇佣母语评估员,将其与我们现有的堆栈进行基准测试,如果有改进就更新生产。Speko将这个过程转变为一个API。一个每天处理数千个电话的团队告诉我们:“我们可以直接去这个仪表盘,切换模型,它会为我们完成。”
<p>工作原理:您发送一个请求,包含您的优化标准(准确性、延迟、成本或平衡)、语言和地区。路由器会过滤出我们针对给定约束组合测量过的模型,对其进行基准测试,选择获胜者,并返回包含供应商、模型名称和分数的响应头。网关会预取已签名的会话计划,因此新的会话可以直接从内存中拨打供应商的电话;在呼叫者等待时不会有控制平面往返。
<p>故障切换仅在连接设置阶段发生:如果供应商拒绝连接请求,我们会开始连接备选供应商。
<p>一些客户故事:一位创始人来找我们时完全不知道该选择什么:他给我们提供了他的用例,现在所有内容都通过这个平台进行路由。一家物业管理AI在Python中运行LiveKit,自启动以来没有更新STT或TTS:他们不知道自己的STT在通话中错误率很高,更好的选项存在,而切换看起来总像是一个研发项目。一个团队不知道该选择哪个西班牙语模型。一个医疗团队不知道哪个STT最适合处理医学词汇。在每种情况下,我们都帮助他们从基准中找到合适的堆栈,现在他们通过我们进行路由。
<p>测量部分是公开的:我们在同一区域对每个模型进行不同日期的相同输入测试,并发布这些结果,包括我们选择的表现不如替代方案的情况。发布的演示回答了哪个30秒的片段听起来更好;生产环境则关注哪个模型能在第八分钟存活,因此我们测试自发语音、金钱和日期、十分钟的录音,排名会发生变化。我们为TTS自然度训练了一个自动评分器,在我们的盲听投票中;对于它从未见过投票的供应商,它选择的获胜者与我们的评估员一致的频率与评估员之间的共识相当。
<p>我们不自己训练或销售模型,这正是我们保持排名公正的原因。
<p>我们还开源了网关,供希望避免音频路径上额外网络跳转并且不想与我们的云共享密钥的团队使用(<a href="https://github.com/SpekoAI/gateway" rel="nofollow">https://github.com/SpekoAI/gateway</a>, MIT):一个Go二进制文件,作为您代理容器中的侧车运行,通过Unix套接字使用一种本地协议,与提供商主机固定并附加您的密钥。在BYOK模式下,它根本不与我们通信。
<p>请注意,匿名的、无内容的遥测默认启用,您只需设置一个环境变量即可禁用。
<p>费用:网关和BYOK设置将永远免费,我们对托管路由器和管理密钥收取费用,并提供合并账单。自从我们在6月底启动批量以来,外部使用量平均每周增长约25%,在发布周前期集中。
<p>我非常希望得到社区的反馈:您现在是如何选择语音模型的?是什么让您信任第三方基准?
查看原文
Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public benchmarked options, and tells you why.<p>Demo: <a href="https://youtu.be/no2LY2gRh-c" rel="nofollow">https://youtu.be/no2LY2gRh-c</a><p>Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS.<p>Each of those layers offers a dozen credible vendors, and each month there are new models on the market. Almost everyone evaluates once, picks a stack of their choice, and never rechecks because switching from a vendor to another involves yet another integration and arguments about the numbers.<p>The result is that you use voice agents running last quarter's models while better and cheaper options are available.<p>Before founding Speko, I spent four years as cofounder and CTO building voice agents for enterprises across Asia in 10+ languages. Each time a new speech model would arrive, we repeated the same ritual: hire native-speaking raters, benchmark it against our existing stack, and update production if it improved. Speko turns this process into an API. A team running thousands of calls a day told us: "we can literally go to this dashboard, switch the model, and it will do it for us."<p>How it works: you send a request with your optimization criteria (accuracy, latency, cost or balanced), language and region. The router filters to models which we measured for the given combination of constraints, benchmarks them, selects the winner, and returns a response with headers containing provider, model names, and the scores. The gateway prefetches signed session plans, so a new session dials the provider straight from memory; no control-plane round trip while a caller waits.<p>Failover happens only during connection setup stage: if the provider refuses the connection attempt, we start connecting to the runners-up.<p>Some of the customer stories: one founder came to us not knowing what to pick at all: he gave us his use case and now routes everything through the platform. A property management AI runs LiveKit in Python and had not updated STT or TTS since launch: they did not know their STT had high error rates on their calls, better options existed, and swapping always looked like an R&D project. One team did not know which models to pick for Spanish. A medical team did not know which STT handles medical vocabulary best. In every case we helped find the right stack from the benchmarks, and now they route through us.<p>The measuring part is public: we pass the same inputs to every model in one region in different dated runs and we publish the boards, including those where our selections perform worse than alternatives. A launch demo answers which 30-second clip sounds better; production asks which model survives minute eight, so we test spontaneous speech, money and dates, ten-minute takes, and the rankings change. We trained an automatic scorer for TTS naturalness on our blind head-to-head listening votes; on providers it has never seen a vote for, it picks the same winner our raters do about as often as raters agree with each other.<p>We don't train or sell models ourselves, that's precisely how we keep our rankings impartial.<p>We also open sourced the gateway for teams who want to avoid an extra network hop on the audio path and don't want to share keys with our cloud (<a href="https://github.com/SpekoAI/gateway" rel="nofollow">https://github.com/SpekoAI/gateway</a>, MIT): one Go binary, which is running as a sidecar in your agent's container, speaks one local protocol over Unix socket, pins provider hosts and attaches your keys. In BYOK mode it doesn't communicate with us at all.<p>Notice that the anonymous, content-free telemetry is enabled by default, and one env var disables it.<p>Cost: the gateway and BYOK setup will be free forever, we charge for the hosted router and managed keys with consolidated billing. Since we started the batch in late June, external usage has grown about 25 percent per week on average, front-loaded toward the launch weeks.<p>I would love feedback from the community: how do you pick speech models now, and what makes you trust the third-party benchmark?