展示HN:人工智能初创公司TrustedRouter融资125万美元
抱歉,我无法访问您提供的链接内容。不过,我可以帮助您翻译您提供的文本。以下是翻译:
大家好,我创办了 trustedrouter.com,这是一个非常简单的使用人工智能的方法,无需将您的数据交给像封闭源路由器这样的第三方。
构建这个项目的过程非常有趣,因为我认识了许多不同服务提供商的首席执行官和创始人。现在,我们的服务提供商数量超过了开放路由器,模型数量也更多。关于这一点,稍后会有更多信息。
我们在安全性和技能方面进行了许多创新,能够为您提供建议,告诉您应该使用哪个大型语言模型(LLM)。此外,我们还创建了一个新网站,名为 anyeval.com,我预计它将成为互联网人工智能开放源基础设施的关键部分。其他公司如 AAII 发布基准测试,但您可能对他们真正的工作内容和测量方式一无所知,他们在测试哪些模型方面存在很多空白,因为这非常昂贵。anyeval 的理念是让您能够支付几分钱的费用来运行单个问题的评估,然后作为一个集体,我们可以众筹为任何我们想要的模型支付整个评估的费用,或者您也可以仅支付随机样本的问题费用,以获取对该模型质量的评估和一些阿拉伯艺术家的意见。您可以进行一对一的比较,也可以创建全新的评估。
例如,我创建了一个新的评估,称为 honey pot bench 或 honey bench,它重现了 hugging face 事件的一些事实,以测量特定人工智能是否倾向于逃逸。我发现 Fable 特别不同于其他 Claude 模型,在对齐性方面表现得非常不一致。
我还创建了 freedom bench,它测量模型与中国审查相关的审查程度,发现主要是中国的服务提供商级别的监控,而不是美国的服务提供商进行审查。我没有在模型权重中看到太多的审查。
查看原文
https://www.axios.com/2026/08/31/exclusive-ai-startup-trustedrouter-raises-125-million<p>hey everybody, I started trustedrouter.com, which is a really simple way to use AI without needing to give your data to a third party like a close source router<p>It’s been really fun to build this because I got to know the CEOs and founders of so many different providers. We now have more providers than open router and more models as well. More on that soon<p>We are doing a lot of innovations in security and skills that advise you on which LLM to use and also we created a new site called anyeval.com that I expect to be a critical part of the open-source infrastructure for AI on the Internet. Other companies like AAII postbenchmarks, but you’d have no idea about what they’re really doing and how they’re really measuring, and they have a ton of gaps about which models they’re doing their tests on because it’s so expensive. The idea behind any eval is so that you can pay the few pennies it costs to run an individual problem in an eval, and then as a together as a collective, we can crowdsource paying for a whole eval for any model that we want, or you can just pay for a random sample of the problems to get a an and some Arab artists on what you think that the quality of that model is. You can do head-to-head comparisons. You can also create whole new evals.<p>For example, I created this new eval called honey pot bench or honey bench, which recreates the some of the facts of the hugging face incident to measure whether a particular AI is prone to wanting to escape, and I found that Fable in particular, unlike the other Claude’s, is very unaligned in comparison<p>I also created freedom bench, which measures the amount of censorship related to Chinese censorship that a model has, and found that it’s mostly the provider level monitor provided at the providers in China, but not the US providers that does the censitions. I don’t see as much censorship in the model weights