Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
Provides a detailed, reproducible methodology for training capable small language models on consumer hardware, directly addressing the cost barrier in academic/open-source research.
AI Summary
Researchers present an open, low-cost recipe that trains a 2B-parameter model from scratch for under $6.9K, approaching Qwen2.5-1.5B performance.
Excerpt
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we pre
