FR
live
tag

#moe

Olmo-core 3 opens MoE training to the trillion-parameter scale

Ai2 ships Olmo-core 3, an open mixture-of-experts training stack that holds a 1.2-trillion-parameter model across 512 GPUs. The FSDP-to-DDP switch and MXFP8 precision change the compute economics for labs that do not have Megatron-Core.

Qwen3.8-Max ships open weights, but not under Apache 2.0

On August 12, 2026, Alibaba published the weights of Qwen3.8-Max, a 2.4-trillion-parameter MoE model, under a custom license with revenue thresholds rather than Apache 2.0. Before you deploy, read the clauses: the checkpoint is text-only and resale above a threshold becomes paid.

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss