A training stack rebuilt around sparse models
Ai2's Olmo-core 3 keeps experts resident on GPUs instead of regathering weights, and reports 2.7 times the throughput of its predecessor.
2 minOpen-Source AI ModelsFresh · 1 Oct
Ai2 has released Olmo-core 3, a rebuild of the framework behind its Olmo models around mixture-of-experts training. The promise of an MoE is that a model can hold many more parameters without every input using all of them. The catch is that the whole model still has to live across GPU memory and be updated, and routing each token to the right experts across a cluster costs communication and coordination that grows with the model.
The central change is where the experts sit. The earlier implementation used fully sharded data parallelism, gathering and resharding weights for each small batch. Olmo-core 3 moves to distributed data parallelism and keeps experts resident on their GPUs, routing the data to them instead of moving the weights back and forth.
The benchmark Ai2 leads with isolates the effect. Growing the expert pool from 8 to 128 while still selecting four experts per token held active parameters at roughly 3.2B and raised total capacity from 4.6B to 47B, with training throughput falling by less than 5%. On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU against 19,400 with the earlier implementation, about 2.7 times the throughput. The same infrastructure has been benchmarked at over one trillion total parameters.
Three techniques carry the scaling: expert parallelism spreads experts across GPUs, pipeline parallelism splits layers across groups of GPUs, and a distributed optimiser spreads optimiser state rather than copying it everywhere. On top of those, rowwise expert parallelism places routed data straight into expert input buffers, and routing metadata stays on the GPUs so the CPU does not wait for it to be copied back.
Ai2 frames the release as a cost argument rather than a capability one. Training large models puts development out of reach for academic groups and smaller labs, and the framework behind each Olmo release is published for the same reason the weights are.
Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.