Olmo-core 3 opens MoE training to the trillion-parameter scale
Ai2 ships Olmo-core 3, an open mixture-of-experts training stack that holds a 1.2-trillion-parameter model across 512 GPUs. The FSDP-to-DDP switch and MXFP8 precision change the compute economics for labs that do not have Megatron-Core.