训练盒子 Poppy 第一部分:起步 (Poppy the Training Box, Part 1: The Beginnings)
文章背景与核心概要
作者 Giles Thomas 厌倦了长期占用主力日常电脑(perry)进行数日之久的本地大语言模型(LLM)训练,于是决定复活一台名为 poppy 的老旧小型(SFF)PC。这台最初为旅行和游戏于 2020 年组装的电脑,换上了宽敞的新机箱、大功率 1600W 电源以及二手 RTX 3090 显卡。在经历了一次长达 11 天的意外训练、CPU 风扇损坏导致的 115°C 高温飙升以及随后的硬件更换后,poppy 成功被改造为一台专注且功能齐全的本地 LLM 专用训练主机。
[作者有一段时间一直计划组装一台专门用于本地大语言模型(LLM)训练的独立机器。在此之前,训练工作都是在日常主力桌面电脑 perry(配备了一张 RTX 3090)上进行的。尽管取得了成功——例如在 JAX 中训练了一个 1.63 亿参数的 GPT-2 small 风格 LLM——但这种设置也带来了几个缺点:]
For a while, the plan has been to put together a separate machine for local LLM training. Until now, training has been done on the daily driver desktop,
perry(equipped with an RTX 3090). While successful—such as training a 163M-parameter GPT-2 small style LLM in JAX—this setup posed several drawbacks:
[ 性能迟钝: CPU 和 GPU 的高负载让日常任务变得迟钝。] [ 无法游戏: 电脑被训练任务占用意味着无法进行游戏。] [ 无法并行实验:* 当 GPU 忙于训练时,探索下一步计划或运行其他实验是不可能的。]
- Sluggish Performance: CPU and GPU heavy loads make everyday tasks sluggish.
- No Gaming: Tying up the PC for days rules out gaming.
- No Parallel Experiments: While the GPU is busy training, scoping out next steps or running other experiments is impossible.
[此外,训练运行时间往往被武断地限制在两天以内,原因仅仅这是 perry 可承受的最大停机时间。更长时间的训练可能会带来令人着迷的结果。]
Additionally, training runs were arbitrarily capped at two days simply because that was the maximum acceptable downtime for
perry. Longer training runs could yield fascinating results.
[长远的目标包括组装一台多 GPU 机器,以便在本地测试大规模云端并行训练,而无需支付昂贵的每小时云租赁费用。最后,作者一直想搭建一套定制水冷循环系统。]
Longer-term goals include building a multi-GPU box to test large-scale cloud parallel training locally without paying high hourly cloud rental fees. Finally, the author has always wanted to build a custom water-cooling loop.
[这篇文章涵盖了基础工作:改造旧 PC、安装从 eBay 购买的二手 RTX 3090、意外地进行了一场长达 11 天的 LLM 训练,以及差点把 CPU 烧毁的经历。]
This post covers the baseline: repurposing an old PC, installing a second-hand eBay RTX 3090, accidentally training an LLM for 11 days, and nearly cooking a CPU.
[在全职搬到里斯本之前,一台名为 poppy 的小型表单因数(SFF)PC 于 2020 年组装完成,它有着特定的限制条件:]
* 适合随身携带: 小到可以塞进登机箱,方便在伦敦和葡萄牙之间旅行时携带。
* 便携: 招待客人时可以轻松在公寓内搬动。
* 具备游戏能力: 足够强大,能够运行诸如《刺客信条:奥德赛》之类的游戏。
Before moving to Lisbon full-time, a small form-factor PC named
poppywas built in 2020 with specific constraints: * Carry-on friendly: Small enough to fit in a carry-on bag for travel between London and Portugal. * Portable: Easy to move around the flat when hosting guests. * Gaming capable: Powerful enough to run games like Assassin's Creed Odyssey.
[### 原配置清单]
Original Component List
[ CPU: AMD Ryzen 5 3600(3.6GHz,6 核)] [ CPU 散热器: Noctua NH-L9a-AM4] [ 主板: Gigabyte X570 I AORUS PRO WIFI Mini ITX] [ 内存: 32 GiB Corsair Vengeance DDR4] [ 存储: 2x Samsung 970 Evo 500 GB NVMe SSD] [ GPU: Zotac GTX 1660 Super 6 GiB] [ 机箱: Lian Li PC-TU100 Mini ITX] [ 电源: Corsair SF450 450W SFF]
- CPU: AMD Ryzen 5 3600 (3.6GHz, 6-Core)
- CPU Cooler: Noctua NH-L9a-AM4
- Motherboard: Gigabyte X570 I AORUS PRO WIFI Mini ITX
- RAM: 32 GiB Corsair Vengeance DDR4
- Storage: 2x Samsung 970 Evo 500 GB NVMe SSDs
- GPU: Zotac GTX 1660 Super 6 GiB
- Case: Lian Li PC-TU100 Mini ITX
- PSU: Corsair SF450 450W SFF
[一旦 perry 在里斯本成为主力机器,poppy 就被闲置在书房的角落里。是时候进行升级了。]
Once
perrybecame the main daily driver in Lisbon,poppysat unused in the corner of the study. It was time for an upgrade.
[初步排查显示 poppy 无法开机,这指向了一个有故障的电源。考虑到未来需要支持多张显卡的需求,购置了新组件:]
* 电源: ASRock Phantom Gaming PG-1600G 1600W(能够带动最多三张 RTX 3090 和一个 CPU)。
* 机箱: Fractal Design North XL Mesh ATX 全塔机箱(为多块 GPU 和水冷提供了充足的空间)。
Initial troubleshooting revealed
poppywouldn't power on, pointing to a faulty PSU. Anticipating the need to support multiple graphics cards eventually, new components were acquired: * PSU: ASRock Phantom Gaming PG-1600G 1600W (capable of handling up to three RTX 3090s and a CPU). * Case: Fractal Design North XL Mesh ATX Full Tower (offering plenty of space for multiple GPUs and water cooling).
[将旧的 mini-ITX 主板和新电源安装到 North XL 机箱后,系统成功开机。随后抹除了 Arch Linux 操作系统并进行了全新安装和配置。]
After installing the old mini-ITX motherboard and new PSU into the North XL case, the system powered on successfully. The Arch Linux OS was wiped and reinstalled with a fresh configuration.
[作为压力测试,使用 PyTorch 环境训练了一个精简版的 GPT-2 small 模型:] * 词表大小: 50,257 * 上下文长度: 512(从 1024 缩减) * 嵌入维度: 512(从 768 缩减) * 注意力头 / 层数: 8 / 8(从 12 / 12 缩减) * 参数量: 约 7690 万(需要进行约 15 亿 Token 的训练运行)
As a burn-in test, a cut-down version of GPT-2 small was trained using a PyTorch setup: * Vocab Size: 50,257 * Context Length: 512 (down from 1024) * Embedding Dimensions: 512 (down from 768) * Heads / Layers: 8 / 8 (down from 12 / 12) * Parameters: ~76.9 million (requiring a ~1.5B token training run)
[### 结果]
Results
[ Perry(RTX 3090): 耗时约 9 小时完成了基准测试,功率为 368W。] [ Poppy(GTX 1660 Super): 报告显示使用率为 100%,但实际上在 53% 的利用率(67W 功耗)下严重卡顿,耗时 267.57 小时(约 11 天)。] [ 电费成本: Poppy 消耗了近 18kWh 的电量,而 Perry 仅为 3.3kWh。买一张 RTX 3090,拯救地球!*]
- Perry (RTX 3090): Completed the baseline run in ~9 hours drawing 368W.
- Poppy (GTX 1660 Super): Ran at 100% reported usage, but effectively choked at 53% utilization (67W draw), taking 267.57 hours (~11 days).
- Energy Cost: Poppy consumed nearly 18kWh compared to Perry's 3.3kWh. Buy an RTX 3090, save the planet!
[尽管存在速度瓶颈,但评估测试显示出了很有前景的文本生成效果,并且损失(Loss)分数与 Perry 的模型相当。]
Despite the speed bottleneck, evaluation tests showed promising text generation and comparable loss scores to Perry's models.
[从保加利亚通过 eBay 购得了一张价格实惠且靠谱的 RTX 3090。安装好后,显卡亮起了充满活力、带有水晶纹理的 RGB 迪斯科灯效——这让带有网状侧板的机箱相比玻璃侧板成为了一个更受欢迎的选择。]
An affordable, trustworthy RTX 3090 was sourced from Bulgaria via eBay. Upon installation, the card powered up with a vibrant, crystal-textured RGB disco display—making the mesh-sided case a welcome choice over glass.
[在使用新 3090 进行第一次完整训练测试期间,poppy 在十分钟后突然关机。]
During the first full training test with the new 3090,
poppyabruptly shut down after ten minutes.
[调查发现 Noctua CPU 散热器风扇没有旋转。InfluxDB 监控显示了历史数据:CPU 在超过一个多月的时间里一直处于 70°C+ 的高温待机状态,并在最终关机时飙升至 115°C 的紧急热保护关机点。]
Investigation revealed the Noctua CPU cooler fan wasn't spinning. InfluxDB monitoring showed historical data: the CPU had been idling at a scorching 70°C+ for over a month, culminating in an emergency thermal shutdown spike at 115°C.
[通过亚马逊次日达服务订购了一个替换风扇(Noctua NF-A9x14 PWM)。安装后,待机温度稳定在健康的 35.5°C。]
A replacement fan (Noctua NF-A9x14 PWM) was ordered via Amazon next-day delivery. Upon installation, idle temperatures stabilized at a healthy 35.5°C.
[散热恢复后,启动了一次标准长度的完整 LLM 训练:]
* 完成时间: 约 40 小时(与 perry 相当)。
* 处理的 Token 数: 32.6 亿
* 训练损失 / 测试损失: 约 3.530 / 约 3.548(与 Perry 的基准测试几乎完全相同)。
With cooling restored, a full-length standard LLM training run was initiated: * Completion Time: ~40 hours (comparable to
perry). * Tokens Seen: 3.26 Billion * Train Loss / Test Loss: ~3.530 / ~3.548 (virtually identical to Perry's benchmarks).
[模型的输出("Every effort moves you and your customers...")证明了这台机器已经完全可以正常工作,随时可以投入实战。]
The model's output ("Every effort moves you and your customers...") proved the rig was fully operational and ready for action.
[poppy 现在是一台配置齐全的本地训练主机,拥有单张 RTX 3090、稳定的 CPU 以及充足的扩展空间。]
poppyis now a fully configured local training box featuring a single RTX 3090, a stable CPU, and ample room for expansion.
[下一步计划:] 1. 过渡到定制水冷。 2. 由于支持多块 GPU 最终需要更换主板和 CPU,因此跳过对当前 CPU 的水冷改造。 3. 相反,将首先在 GPU 上安装水冷头,构建一个单组件的水冷循环——并希望顺便去掉 RGB 灯效。
Next Steps: 1. Transition to custom water cooling. 2. Because supporting multiple GPUs will eventually require a new motherboard and CPU, water-cooling the current CPU is skipped. 3. Instead, a water block will be installed on the GPU first to build a single-component loop—and hopefully ditch the RGB lighting along the way.