LLM Comparison
DeepSeek V4 Flash 0731 vs DiffusionGemma
Side-by-side specs, pricing & capabilities · Updated August 2026
Price vs Intelligence
Add to comparison
2/6 modelsSame tier:
| Organization | ||
| OpenTools Score | 46 219 | 36 |
| Family | DeepSeek V4 Flash | Gemma |
| Status | Current | Current |
| Release Date | Jul 2026 | Jun 2026 |
| Context Window | 1.0M tokens | 256K tokens |
| Input Price | $0.14/M tokens | Free |
| Output Price | $0.28/M tokens | Free |
| Pricing Notes | DeepSeek API pricing lists $0.0028 cache-hit input, $0.14 cache-miss input, and $0.28 output per 1M tokens for deepseek-v4-flash; peak pricing may change after a future official announcement. | Open-weights model under Apache 2.0; API pricing depends on the host or local infrastructure used. |
| Capabilities | reasoningcodingtool-usestructured-outputlong-contextresponses-api | textvisioncodereasoninglocal-inference |
| Max Output | 393K tokens | 256 tokens |
| API Identifier | deepseek-v4-flash | google/diffusiongemma-26b-a4b-it |
| Benchmarks | ||
| Artificial Analysis Intelligence Index v4.1 | 50artificial-analysis | — |
| Artificial Analysis Intelligence Index v4.1.1 | 52artificial-analysis | — |
| MMLU Pro | — | 77.6official-google-model-card |
| GPQA Diamond | — | 73.2official-google-model-card |
| LiveCodeBench v6 | — | 69.1official-google-model-card |
| MMMLU | — | 81.5official-google-model-card |
| HLE no tools | — | 11official-google-model-card |
| View DeepSeek V4 Flash 0731 | View DiffusionGemma | |
Cost Calculator
Enter your expected monthly token usage to compare costs.
| Model | Input | Output | Total / mo | vs Best |
|---|---|---|---|---|
| DiffusionGemmaCheapest | $0.00 | $0.00 | $0.00 | — |
| DeepSeek V4 Flash 0731 | $0.14 | $0.14 | $0.28 | +0% |
DeepSeek
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is an open-weight, cost-efficient reasoning model for long-context coding, agent tasks, and API workflows through DeepSeek.
DiffusionGemma
DiffusionGemma is Google DeepMind’s experimental open-weights text-diffusion model based on Gemma 4 26B A4B. It uses discrete diffusion and parallel canvas denoising to trade some benchmark quality for much faster local generation on dedicated GPUs.
More Comparisons
Looking for more AI models?
Browse All LLMs