FR
live
tag

#training

Olmo-core 3 opens MoE training to the trillion-parameter scale

Ai2 ships Olmo-core 3, an open mixture-of-experts training stack that holds a 1.2-trillion-parameter model across 512 GPUs. The FSDP-to-DDP switch and MXFP8 precision change the compute economics for labs that do not have Megatron-Core.

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss