FireTofu
Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

Technology · en

Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

The Decoder · Aug 9, 2026, 10:01 AM UTC

Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second.…