DiffusionGemma attacks the slowest LLM habit: one token at a time
Google DeepMind introduced DiffusionGemma, an experimental open text generation model that the model page says can produce up to 4x to 5x faster output on NVIDIA GPUs and exceed 1,000 tokens per second on an H100. The bigger point is architectural: text models are starting to attack the sequential bottleneck.