18 August 2026

Nvidia releases efficient model with fewer active parameters

  • Nvidia released Nemotron 3.5 Lightning, a model designed to run efficiently by activating only 3 billion of its 30 billion total parameters at any given time.
  • The model can predict multiple tokens simultaneously, reducing the number of computational steps needed to generate text.
  • Nvidia published research on training methods that prevent efficiency loss between training and real-world deployment for large models with this architecture.

How it was covered

Latent Spaceswyx & Alessio

Nemotron 3.5 Lightning exemplifies a shift toward architecture-level efficiency with its 30B MoE design using 3B active parameters, multi-token prediction support, and research on RL for large MoEs showing zero train-infer mismatch.