技能双倍,108次运行:优化效能与代币效率

1 分•作者: darvh•大约 1 个月前•原帖
马尾辫与信号 基准测试虽然有运行,但并不系统化。模型展示了对答案的记忆,这在如今是有些预期的。 比起技能,能够让基准测试顺利进行更有趣 :). 如果关于基准测试的进行或构建的前言有问题,请告诉我。 博客文章:darvh.com/posts/when-coding-agents-raced-through-108-bugs/ 基准测试:github.com/darvh/bench 信号:github.com/darvh/signal
查看原文
Ponytail vs Signal<p>Bench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days.<p>More than the skills, it was fun getting the benchmark running :).<p>Let me know if the preamble on how the bench was conducted or constructed is flawed.<p>Blog Post: darvh.com&#x2F;posts&#x2F;when-coding-agents-raced-through-108-bugs&#x2F;<p>Bench: github.com&#x2F;darvh&#x2F;bench<p>Signal: github.com&#x2F;darvh&#x2F;signal