Ai2 Releases Olmo-Core 3 Framework for Efficient Mixture-of-Experts AI Training
Tech

Ai2 Releases Olmo-Core 3 Framework for Efficient Mixture-of-Experts AI Training

TechNews Editorial
TechNews EditorialOct 2, 2026 · 3 min read
Share

Why it matters

The framework makes training trillion-parameter mixture-of-experts models more accessible and affordable by significantly reducing memory overhead and hardware requirements.

The facts

  • Ai2 released the Olmo-core 3 framework to improve training efficiency for large mixture-of-experts language models.
  • The system achieved 52,000 tokens per second on Nvidia B3000 GPUs, offering 2.7 times the throughput of Megatron-core.
  • The open-source project is available on GitHub to help researchers scale models to over 1 trillion parameters.

The Allen Institute for AI announced a development framework for large language models on Thursday. The new system significantly improves how mixture-of-experts large language models are trained. Named Olmo-core 3, the framework allows mixture-of-experts training to reach the trillion-parameter scale while keeping costs low by preserving computational efficiency.

Mixture-of-experts models divide work across specialist parts

Mixture-of-experts models operate differently from dense artificial intelligence models. They split computation across specialist portions of the model each time a token is generated. Dense models activate the entire model for every token.

A mixture-of-experts model may contain far more total parameters, which are individual dials that tune model behavior. However, it computes with only a small number of them as it generates each part of an answer. Training an expert model allows activating only portions of it while learning each token. A token is a small piece of text, often a word or a part of a word, that an artificial intelligence reads and generates.

This lowers overall compute costs. However, the full model still must be stored across graphics processing unit memory. Coordinating networking between experts during training still generates additional costs.

A distributed training cluster exchanges data through network switches while four processing paths work and memory remains occupied across the cluster.
Illustration: AI & Tech News

Olmo-core 3 scales expert pools and processes more tokens

The Allen Institute for AI stated that Olmo-core 3 bridges the gap between dense models and mixture-of-experts models. It allows the expert pool to grow from eight to 128 while still selecting only four experts per token. Using the same infrastructure, large language models can scale to more than 1 trillion parameters.

In benchmarks, the company reported that Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 graphics processing units for a 47 billion-parameter model. This compared with Nvidia Corporation's Megatron-core training architecture, an established option for training large mixture-of-experts models that topped out at about 19,400 tokens per second. The result represented a jump of about 2.7 times the throughput.

Read nextMicrosoft AI Releases New Transcription and Text-to-Speech Models

New architecture uses advanced parallelism to reduce memory overhead

In a white paper published about the project, the Allen Institute for AI noted that the new architecture uses expert parallelism to spread experts across multiple graphics processing units. This allows each card to store only part of the full expert pool. It also splits the model layers, which are successive stages generating inputs, across groups of graphics processing units to reduce how much of the model each card needs to keep in memory.

A distributed optimizer spreads the optimizer state across multiple graphics processing units instead of storing full copies on every card. The optimizer state is additional data used to calculate and apply updates during training. Altogether, this reduces memory overhead as models scale because the entire model and its training state do not need to be stored in memory all at once.

Additionally, the creator supports MXFP8, a number format for large language models that represents some values with fewer bits. This format can reduce computation and the amount of data moved between graphics processing units.

The new training architecture and its higher efficiency form part of the company vision to give researchers tools to build and train larger models. Trillion-parameter artificial intelligence models are often beyond the reach of those without access to state and enterprise infrastructure. The Allen Institute for AI added that Olmo-core 3 will also allow researchers to adapt mixture-of-experts training to different hardware and experiment with routing, parallelism, and other parts of the system to build out a new ecosystem.

The project and related systems are currently available for developers and the open-source community on GitHub.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading