RUN THIS LLM
Search local LLM hardware requirements
DiffusionGemma 26B MoE
Google · 25B MoE · Vision
Google's diffusion language model built on Gemma 4 26B MoE. Generates 256-token blocks in parallel via iterative denoising — 1,100+ tok/s on an H100. Accepts text, image, and video input.
VRAM Requirements
| Quantization | VRAM |
|---|---|
| Q4_K_M (smallest) | 15 GB |
| Q8_0 (balanced) | 27.5 GB |
| FP16 (full quality) | 50 GB |
Specifications
- Parameters: 25B MoE (3.8B active per token)
- Category: Vision
- Max context: 256K tokens
- System RAM: 24 GB minimum
- HuggingFace: google/diffusiongemma-26B-A4B-it
Benchmarks
- HumanEval: 69 — Code generation (pass@1)
- MATH: 69 — Competition-level math reasoning
Loading interactive analysis...