FR
live
AI

Olmo-core 3 opens MoE training to the trillion-parameter scale

Ai2 ships Olmo-core 3, an open mixture-of-experts training stack that holds a 1.2-trillion-parameter model across 512 GPUs. The FSDP-to-DDP switch and MXFP8 precision change the compute economics for labs that do not have Megatron-Core.

An aisle of identical dark GPU racks in a dim data center, one single amber status LED lit on one rack.

October 1, 2026. Ai2 — the Allen Institute for AI — publishes Olmo-core 3, a redesign of its open language-model training infrastructure centered on mixture-of-experts (MoE). October 1, 2026. The stack has been benchmarked on a 1.2-trillion-parameter model with 58.36 billion active per token, spread across 512 NVIDIA B300 GPUs. October 1, 2026. Switching to the low-precision MXFP8 format adds roughly 21% throughput over BF16 while cutting peak active memory from 103 GiB to 95 GiB. Why it matters: this is a credible open-source alternative to Megatron-Core for training large sparse models, and it attacks exactly the costs that eroded the MoE advantage at scale.

The problem MoEs never quite solved

MoEs have promised the same thing for years: carry far more learned parameters without every input having to use them all. A dense model activates almost all of its capacity on every token; an MoE activates only a handful of experts — the specialized components of the model — per token. The idea is to pay less compute for the same capacity.

Except the full model still has to live in GPU memory and be updated during training. And routing each input to the right expert, across a cluster, creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the saving gained by using only a fraction of the model for each input.

Olmo-core 3 is built to close that gap. In one benchmark, the team grew the expert pool from 8 to 128 while still selecting only four experts per token. Active parameters per token stayed roughly fixed at about 3.2 billion, while total model capacity grew from 4.6 to 47 billion parameters — with a training-throughput drop of less than 5%. In other words, capacity increased tenfold without nearly touching speed.

FSDP to DDP, or how to stop reloading weights

The key structural change is abandoning fully sharded data parallelism (FSDP) in favor of distributed data parallelism (DDP). The earlier MoE implementation in Olmo-core relied on FSDP, configured to gather and re-shard model weights on every small batch of training data. That back-and-forth gets expensive as the model grows.

Olmo-core 3 keeps experts resident on GPUs and routes the relevant data to them, avoiding the repeated weight gathering. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, versus 19,400 with the old one — about 2.7× the throughput.

The comparison is explicitly with Megatron-Core, NVIDIA’s established option for training large MoEs. Olmo-core 3 brings an integrated MoE stack to the framework behind Olmo, with a redesign that beats the throughput of the earlier FSDP implementation. It is a clear positioning: an open path that does not depend on the vendor’s proprietary stack.

Three parallelisms, three routing optimizations

How the model and its training state are split across hardware rests on three techniques. Expert parallelism spreads the experts across GPUs, so each stores only part of the expert pool. Pipeline parallelism splits the model’s layers — the successive stages that transform an input — across groups of GPUs, reducing how much of the model each must hold in memory. A distributed optimizer spreads optimizer state across GPUs instead of keeping a full copy on each. Together, these let an MoE grow without requiring every GPU to keep the entire model and its training state in memory.

Routing gets three targeted optimizations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing rearrangement work. GPU-resident routing keeps routing metadata on the GPUs so the CPU can queue work without waiting for a copy back. And grouped GEMM combines many small expert computations so GPUs execute them more efficiently.

On top of that, MXFP8 is a lower-precision number format that represents some values in fewer bits. In a controlled benchmark on four NVIDIA B300 GPUs, enabling it where it helped most delivered about 21% higher training throughput than BF16, with peak active memory falling from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts, not attention alone.

The trade-offs are the point

None of these techniques work in isolation, and the team is explicit about it. Speeding up one part of training can create costs elsewhere: faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving researchers using the open stack control over how the pieces fit together. An interactive walkthrough ships alongside the release, showing how data, expert, and pipeline parallelism combine to scale MoE training from a single GPU to many.

Who this is actually for

The deeper stake is not the throughput record but access. Training large language models takes enormous compute, driving up cost and energy use and putting advanced development out of reach for many academic researchers and smaller labs. An open MoE stack that holds a trillion parameters without throughput collapse lowers the entry ticket.

The nuance is in the measurement method. The benchmarks — including the 1.2 trillion parameters at 512 GPUs, with a peak observed throughput of 858 TFLOP/s per GPU — use random routing to measure system performance, not the quality of a trained model. In other words, Olmo-core 3 proves the infrastructure can carry the load; the quality of the models that come out of it remains to be shown on the next generations of Olmo.

The project sits in a trajectory. Ai2’s work on sparse models goes back to OlmoE, an MoE with 64 routed experts; Olmo 3, by contrast, was dense, and its training stack was built around that density. Olmo-core 3 reconciles the two lines by rebuilding the training tool for far larger MoEs, with a promise to open up the infrastructure behind every future model. The open-source angle matters here because it decouples MoE training from a single vendor’s roadmap: researchers can inspect, modify, and rerun the stack, which is exactly what a reproducibility-conscious lab wants from its training infrastructure.

Verdict

If you train — or plan to train — large sparse models without wanting to depend on Megatron-Core, Olmo-core 3 is the open alternative to watch closely: the FSDP-to-DDP switch and MXFP8 are concrete, quantified levers, not promises. If you only do inference or light fine-tuning, this release does not concern you directly — it is training infrastructure, not a model to download. If you want to measure for yourself, the technical report, code, and interactive demo are public, and the throughput numbers are reproducible on B300 GPUs.

Références

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

OpenAI disrupts a reasoning extraction campaign tied to Moonshot AI associates

OpenAI says it neutralized a coordinated distillation campaign that extracted protected reasoning from its models, attributed to individuals associated with Moonshot AI. An August 2026 study shows the encrypted traces of Claude, Gemini, and GPT are interchangeable across sessions, enabling a scalable decryption jailbreak.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss