陌生人使用HF PRs和定时任务预训练了一个语言模型。

1作者: somevyn大约 16 小时前原帖
上周,我们几个人从零开始预训练了一个1500万参数的语言模型——超出了其Chinchilla最佳令牌预算——没有集群、没有服务器,也没有预算。训练协调员是一个GitHub Actions的定时任务。梯度则是通过拉取请求提交的。最后,有一个我们从未见过的人加入并训练了2300万个令牌。 github.com/commonsense-ai/coop — 模型 + 卡片 · huggingface.co/commonsense-ai/tinystories-15m — 梯度收件箱 — 排行榜
查看原文
Last week a few of us pretrained a 15M-parameter language model from scratch — past its Chinchilla-optimal token budget — with no cluster, no server, and no budget. The training coordinator is a GitHub Actions cron job. The gradients are pull requests. By the end, someone we'd never met had joined and trained 23M tokens of it. github.com/commonsense-ai/coop — Model + card · huggingface.co/commonsense-ai/tinystories-15m — Gradient inbox — Leaderboard