2 min read

Olmo-core 3: Scalable Infrastructure for Trillion-Parameter MoEs

machine learningdeep learningMoEscalable AIdistributed systemsGPU computing

Executive Summary

Olmo-core 3, an advancement from Hugging Face, enhances the framework for developing large language models by effectively scaling Mixture-of-Experts (MoEs) training to trillions of parameters. This release targets scalability and computational efficiency challenges of large-scale AI, empowering researchers with open and adaptable infrastructure.

The Architecture / Core Concept

Olmo-core 3 focuses on managing and scaling complex MoE architectures efficiently. Unlike dense models that activate all parameters for every input, MoEs select only a subset of experts for processing, facilitating greater parameter efficiency. These experts are specialized sub-models, each handling specific data aspects.

Traditionally, the challenge with MoEs is directing inputs, or tokens, to the relevant experts, especially when expanding into the trillion-parameter range. Olmo-core 3 mitigates these challenges through distributed data parallelism—keeping expert weights resident on GPUs to lower data movement, unlike prior approaches that reshuffled weights for each mini-batch. This optimizes not only the processing but also the network and memory usage during training.

Implementation Details

Key techniques underpinning Olmo-core 3 include expert parallelism, pipeline parallelism, and the use of a distributed optimizer:

# Pseudo code illustrating distributed data parallelism
for token in input_tokens:
    available_experts = select_experts(token)
    for expert in available_experts:
        dispatch_to_gpu(expert, token)

These methods ensure GPUs maintain only portions of the overall training state, minimizing memory overheads.

Moreover, advanced techniques like rowwise expert parallelism, GPU-resident routing, and grouped GEMM optimize computational loads. The introduction of MXFP8 further enhances throughput by utilizing lower-precision calculations without sacrificing model accuracy.

Engineering Implications

Implementing Olmo-core 3 implicates several trade-offs:

  • Scalability: By efficiently distributing model components across GPUs, Olmo-core 3 significantly enhances the feasible size of models, enabling trillion-parameter configurations.
  • Latency and Throughput: Innovations in routing and computation support notable improvements in throughput, as seen in benchmark comparisons—achieving up to 2.7× the throughput compared to earlier implementations.
  • Cost and Complexity: The reduction in data transfer and lower precision calculations reduces the operational and energy costs. However, the complexity of implementation increases with these optimizations, requiring meticulous tuning and real-time monitoring.

My Take

Olmo-core 3 represents a transformative update in the MoE domain by making sophisticated AI model training more accessible and less resource-intensive, even at unprecedented scales. However, the crux of its success, especially in diverse real-world applications, will depend on the community's engagement with the open-source framework. As Olmo-core harnesses modular techniques, its adaptability and support potentially lead to new breakthroughs—balancing innovations with practical usability will define its broader impact.

Share this article

J

Written by James Geng

Software engineer passionate about building great products and sharing what I learn along the way.