[{"data":1,"prerenderedAt":29},["ShallowReactive",2],{"nr-en-google-diffusiongemma-training-budget":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":21,"trackerLabel":22,"headlineStat":23,"image":24,"ogImage":25,"imageAlt":5,"csv":10,"minutes":26,"words":27,"html":28},"google-diffusiongemma-training-budget","Google DeepMind: DiffusionGemma with 90% Less Training Budget","Google converted Gemma 4 into a diffusion model without training from scratch – using less than 10 percent of the original budget. The result: 1,500 tokens per second instead of one at a time.","2026-08-09","12:28","2026-08-09T12:28:00+02:00","","August 9, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20],"Large Language Models","Diffusion Models","Training Efficiency","Google DeepMind","\u002Fstand-der-ki","AI Progress","90% less training budget","\u002Fnewsroom\u002Fimg\u002Fgoogle-diffusiongemma-training-budget.webp","\u002Fog-nr\u002Fgoogle-diffusiongemma-training-budget.en.png",3,575,"\u003Cp>Google DeepMind has found a new way to build text diffusion models without training them from the ground up. The company converted the finished \u003Cstrong>Gemma 4\u003C\u002Fstrong> into a diffusion model using less than 10 percent of the original training budget. The model is called \u003Cstrong>DiffusionGemma\u003C\u002Fstrong> and was released in June 2026; now the technical report has followed.\u003C\u002Fp>\n\u003Ch2>The essentials\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>DiffusionGemma\u003C\u002Fstrong> generates \u003Cstrong>256 tokens in parallel\u003C\u002Fstrong> instead of sequentially, achieving about \u003Cstrong>1,500 tokens per second\u003C\u002Fstrong> on an Nvidia H100\u003C\u002Fli>\n\u003Cli>Converting existing models requires only \u003Cstrong>&lt;10% of the original training budget\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>Bidirectional reasoning: the model can correct errors during generation, not just afterward\u003C\u002Fli>\n\u003Cli>Quality lags behind the original, especially on reasoning tasks – but speed is several times higher\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>How does the conversion work?\u003C\u002Fh2>\n\u003Cp>Instead of training a new model, Google uses two training phases. In the first step, the model learns to reconstruct noisy text blocks from example data – similar to how image AI models generate a picture from noise. Then comes a combined phase of \u003Cstrong>reinforcement learning and sampler distillation\u003C\u002Fstrong>, which Google calls \u003Cstrong>SD·RL\u003C\u002Fstrong>. Reinforcement learning improves answer quality, while sampler distillation lets the model work with fewer compute steps.\u003C\u002Fp>\n\u003Cp>This combined approach raises quality on reasoning benchmarks by an average of \u003Cstrong>10 points\u003C\u002Fstrong> while nearly \u003Cstrong>quadrupling\u003C\u002Fstrong> tokens per compute step. Side effect: DiffusionGemma&#39;s answers are about 50 percent shorter – which further boosts speed.\u003C\u002Fp>\n\u003Ch2>Self-correction instead of linear thinking\u003C\u002Fh2>\n\u003Cp>Autoregressive models like Gemma 4 must commit to the first digit of an answer before they&#39;ve worked through the reasoning. Google&#39;s report shows an example: Gemma 4 writes &quot;-1&quot; first, realizes during the derivation that &quot;-25&quot; is correct, and adds a correction afterward.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>DiffusionGemma develops answer and reasoning in parallel.\u003C\u002Fstrong> The model can fix mistakes before the output is finalized. With Sudoku puzzles, the advantage is clear: after minimal fine-tuning, DiffusionGemma solves nearly \u003Cstrong>85 percent of puzzles correctly\u003C\u002Fstrong>, while the base model fails at the task entirely. Structured outputs like JSON or code repairs finish after just \u003Cstrong>2–3 refinement steps\u003C\u002Fstrong>, because the input already determines most tokens.\u003C\u002Fp>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Aspect\u003C\u002Fth>\n\u003Cth>DiffusionGemma\u003C\u002Fth>\n\u003Cth>Gemma 4 (autoregressive)\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>\u003Cstrong>Throughput\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>~1,500 tokens\u002Fs\u003C\u002Ftd>\n\u003Ctd>Significantly slower\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>Quality (benchmarks)\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>Behind original\u003C\u002Ftd>\n\u003Ctd>Higher\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>Reasoning ability\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>Bidirectional, self-correcting\u003C\u002Ftd>\n\u003Ctd>Linear, correctable only afterward\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>Training budget\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>&lt;10% of original\u003C\u002Ftd>\n\u003Ctd>100%\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Ch2>Where are the limits?\u003C\u002Fh2>\n\u003Cp>Absolute performance falls short of the autoregressive base model – especially on demanding reasoning tasks. Google cites several reasons: DiffusionGemma was not trained as a diffusion model from scratch, but converted from an existing model. This has consequences for quality. Multi-user scenarios are also still limited.\u003C\u002Fp>\n\u003Cp>One advantage remains: DiffusionGemma can switch between both modes – users can choose between parallel and sequential generation depending on the task.\u003C\u002Fp>\n\u003Ch2>What does this mean for businesses?\u003C\u002Fh2>\n\u003Cp>The approach significantly lowers the barrier to entry for text diffusion models. If you already have a base model, you can convert it into a faster model with 90 percent less compute – relevant for organizations with limited GPU budgets. For applications that prioritize speed over maximum quality (real-time systems, mass processing), DiffusionGemma could be interesting. However, it remains unclear whether this approach works for specialized models – developers would need to test it themselves.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fgoogles-diffusiongemma-proves-you-dont-need-to-train-from-scratch-to-build-a-text-diffusion-model\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fgoogle-diffusiongemma-technischer-bericht-erklaert-das-schnelle-text-diffusionsmodell\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1786280106321]