说明性人工智能的构思
我听到了一些关于前沿模型在<i>解释数学</i>方面表现不佳的投诉,因此我在考虑一个可能有帮助的强化学习环境:<p>- 选择一个非常困难的数学问题,并确保有可验证的答案<p>- 让前沿模型向一个小型模型(参数在0.5-1亿之间,并且在该问题上表现不佳)解释如何解决,但不提供具体的答案,并对前沿模型给予奖励,鼓励其提供能够帮助小型模型解决问题的提示/解释<p>显然,需要一定程度的人类监督,以防止前沿模型提供过多的信息。
查看原文
I've heard some complaints about the frontier models still be bad at <i>explaining math</i> and was thinking of an RL environment that would help might be to:<p>-Take very hard math problem with a verifiable answer<p>-Have frontier model explain to a tiny model like (0.5-1B params and provably bad score on the problem) how to solve but not the solution, and reward the frontier model for prompts/explanations that helped the tiny model solve the problem<p>Obviously some amount of human supervision is needed to weed out it giving too much information