Test training scale before expanding GPU capacity
In a report published on October 5, 2026, Aleph Alpha describes scaling a 30B-A3B MoE model from 16 to 512 B200 GPUs, reporting near-linear scalability and 35.3% MFU.
Source: Escalando o pré-treinamento de 16 para 512 GPUs B200 (aleph-alpha.com). Text prepared with AI from this source.
What happened and what to do
In a report published on October 5, 2026, Aleph Alpha says it scaled pretraining of a 30B-A3B MoE model from 16 to 512 B200 GPUs. The reported result was near-linear scalability and 35.3% mean GPU utilization (MFU). The approach narrows the search space at small scale before increasing the GPU count.
A company planning model training can apply this approach in a test environment: compare configurations with limited resources, record runtime, utilization, and cost, then set criteria for expanding infrastructure. Wendelmaques can help structure data collection and monitoring and automate the evaluation of configurations.
How the consultancy can help
Wendelmaques can diagnose the opportunity and infrastructure constraints, define a testing scope, and implement and operate monitoring and evaluation pipelines to inform training scale-up.
Next step
Send a short description of your case to receive a scoped proposal for diagnosis and implementation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal