Olmo-core 3 Reframes MoE Scaling as a Data-Movement Problem

Ai2's redesigned open training stack keeps experts resident on GPUs, adds several forms of parallelism, and exposes the practical tradeoffs behind scaling sparse mixture-of-experts models.

Source artwork for Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Olmo-core 3 makes sparse-model scaling more inspectable, but teams still need workload-matched tests of routing balance, communication, precision, and recovery before treating headline throughput as transferable.

Mixture-of-experts models promise a useful bargain: increase a model's total capacity without activating every parameter for every token. Realizing that bargain, however, depends on much more than selecting a small number of experts. The system still has to store those experts, move tokens to the right devices, synchronize updates, and prevent communication from swallowing the compute saved by sparsity.

Ai2's newly released Olmo-core 3 focuses on that systems problem. The open framework replaces its earlier fully sharded approach to MoE training with a distributed-data-parallel design that keeps expert weights resident on GPUs and sends routed data to them. Ai2 reports that an eight-B300 test of a 47-billion-parameter model reached 52,000 tokens per second per GPU, versus 19,400 for its prior implementation. That result is specific to the published configuration; it should not be read as a universal speedup for arbitrary clusters or models.

Why resident experts change the bottleneck

A sharded training system can conserve memory by repeatedly assembling the weights needed for a unit of work and then redistributing them. That is attractive for dense models, where most weights participate in each pass. For an MoE, it can become an awkward match: only a subset of experts handles a token, yet the machinery may repeatedly move large weight tensors.

Keeping experts resident reverses that choice. Weights stay put, while comparatively smaller token representations travel to the devices that own the selected experts. This does not eliminate communication. It makes communication explicit and gives the runtime a better opportunity to organize it. The useful question is no longer simply whether sparse activation reduces arithmetic, but whether token dispatch, expert computation, and result collection can be scheduled efficiently together.

Consider a simplified 64-expert layer spread across eight GPUs. If each token selects four experts, every GPU need not execute all 64. But a token batch originating on one GPU may need experts located on several others. An efficient stack must group destinations, exchange data, run many uneven expert workloads, and return the outputs without leaving devices idle. The model architecture and the cluster topology therefore become a coupled design problem.

Parallelism is a portfolio, not one switch

Olmo-core 3 combines expert parallelism, pipeline parallelism, and distributed optimizer state. Each addresses a different memory pressure. Expert parallelism partitions the expert pool; pipeline parallelism assigns different layer ranges to device groups; optimizer sharding avoids replicating all training-state data everywhere.

These dimensions should be evaluated together. More pipeline stages can lower per-device model memory, but introduce boundaries where activations must move and where idle pipeline gaps may appear. Wider expert parallelism can fit a larger expert pool, but can also increase all-to-all traffic. Sharding optimizer state helps memory capacity, while adding synchronization requirements. A configuration that fits is not necessarily one that uses the hardware well.

The new stack also includes GPU-resident routing metadata, rowwise expert parallelism, and grouped matrix multiplication. The common theme is reducing orchestration overhead: avoid unnecessary trips through the CPU, place routed values closer to their final buffers, and turn many small expert operations into work shaped more efficiently for accelerators. These optimizations matter because MoE workloads can fragment otherwise powerful GPUs into small, irregular jobs.

Lower precision still needs end-to-end accounting

Support for MXFP8 offers another lever. Ai2 measured roughly 21% higher end-to-end throughput than BF16 in a controlled four-B300 benchmark with uniformly distributed work, while peak active memory decreased from 103 GiB to 95 GiB. The setup matters: uniform routing removes an important source of imbalance, and conversion overhead can offset lower-precision savings in other workloads.

Teams evaluating the feature should separate at least four measurements: arithmetic throughput, inter-device transfer volume, conversion cost, and training behavior. Faster kernels alone do not establish a faster training run. Likewise, lower memory use may be valuable even when wall-clock gains are modest, because it can enable a larger batch, more experts, or a less complicated partitioning plan. Numerical quality also remains a model-training question rather than a conclusion that follows from a systems benchmark.

What the trillion-parameter result does and does not show

Ai2 also reports a 1.2-trillion-parameter configuration across 512 B300 GPUs, with 58.36 billion parameters active per token and a peak observation of 858 TFLOP/s per GPU. A separate short-capacity experiment using DeepEP v2 reached 2.38 trillion total parameters. These are infrastructure scale demonstrations, not evidence that a model of that size was trained to convergence or achieved a particular downstream quality level. The 1.2-trillion test used random routing, so it isolates system behavior rather than learned routing quality.

That boundary is important. Capacity, sustained utilization, convergence, and model usefulness are distinct milestones. A stack may prove that tensors fit and operations execute before researchers know whether routing remains balanced over a long run, whether failures are recoverable at scale, or whether the additional experts improve capability enough to justify their cost.

A practical evaluation checklist

Before adopting the stack, a training team can build a small matrix of representative tests. Vary expert count and top-k routing while holding active parameters approximately constant. Repeat measurements with realistic token distributions rather than only uniform or random assignments. Track tokens per second alongside peak memory, network utilization, expert load variance, and time lost to stragglers. Use identical input values when comparing kernels, since data values themselves can affect timing.

Finally, test the failure cases that headline throughput omits: checkpoint duration, restart behavior, reproducibility, and sensitivity to topology. Olmo-core 3 is valuable because it opens more of this machinery for inspection and adaptation. Its strongest contribution may be making MoE training choices testable as a complete system, rather than presenting sparsity as an automatic efficiency win.

Source: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs ↗. How we write

← Back to all articles