技能双倍,108次运行:优化效能与代币效率
马尾辫与信号
基准测试虽然有运行,但并不系统化。模型展示了对答案的记忆,这在如今是有些预期的。
比起技能,能够让基准测试顺利进行更有趣 :).
如果关于基准测试的进行或构建的前言有问题,请告诉我。
博客文章:darvh.com/posts/when-coding-agents-raced-through-108-bugs/
基准测试:github.com/darvh/bench
信号:github.com/darvh/signal
查看原文
Ponytail vs Signal<p>Bench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days.<p>More than the skills, it was fun getting the benchmark running :).<p>Let me know if the preamble on how the bench was conducted or constructed is flawed.<p>Blog Post: darvh.com/posts/when-coding-agents-raced-through-108-bugs/<p>Bench: github.com/darvh/bench<p>Signal: github.com/darvh/signal