LLM Comparison
DeepSeek V3.2 Speciale vs Gemini 3.5 Flash-Lite
Side-by-side specs, pricing & capabilities · Updated August 2026
Price vs Intelligence
Add to comparison
2/6 models| Organization | ||
| OpenTools Score | 63 78.1 | 13 9.0 |
| Family | DeepSeek | Gemini |
| Status | Current | Current |
| Release Date | Dec 2025 | Jul 2026 |
| Context Window | 164K tokens | 1.0M tokens |
| Input Price | $0.40/M tokens | $0.30/M tokens |
| Output Price | $1.20/M tokens | $2.50/M tokens |
| Pricing Notes | Cache read: $0.2000/M tokens | DeepMind model page lists $0.30 input and $2.50 output per 1M tokens with no caching in the performance table. The ai.google.dev API docs were attempted through ScrapingBee and returned HTTP 400, so exact billing details beyond the official DeepMind price table were not used. |
| Capabilities | textcode | textvisionaudio-inputvideo-inputpdf-inputreasoningcodingtool-usesearch-groundingcomputer-usestructured-output |
| Max Output | 164K tokens | 64K tokens |
| API Identifier | deepseek/deepseek-v3.2-speciale | gemini-3.5-flash-lite |
| Benchmarks | ||
| MMLU-Pro | 85deepseek | — |
| AIME 2025 | 96deepseek | — |
| LiveCodeBench | 88.7deepseek | — |
| HLE | 30.6deepseek | — |
| HMMT 2025 | 99deepseek | — |
| SWE-bench Verified | 73deepseek | — |
| Terminal-Bench 2.0 | 46deepseek | — |
| Artificial Analysis Intelligence Index | — | 36artificial-analysis |
| SWE-Bench Pro | — | 54.2google-deepmind |
| Terminal-Bench 2.1 | — | 54google-deepmind |
| GDPVal-AA v2 | — | 1140google-deepmind |
| View DeepSeek V3.2 Speciale | View Gemini 3.5 Flash-Lite | |
Cost Calculator
Enter your expected monthly token usage to compare costs.
| Model | Input | Output | Total / mo | vs Best |
|---|---|---|---|---|
| DeepSeek V3.2 SpecialeCheapest | $0.40 | $0.60 | $1.00 | — |
| Gemini 3.5 Flash-Lite | $0.30 | $1.25 | $1.55 | +55% |
DeepSeek
DeepSeek V3.2 Speciale
DeepSeek V3.2 Speciale is a large language model from DeepSeek. Supports up to 163,840 token context window. Achieves 87.1% on MMLU. Available from $0.40/M input tokens.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is Google DeepMind general availability low-latency model for high-volume agentic systems, coding, document parsing, translation, and structured automation. Official DeepMind evidence lists 1M input tokens, 64k output tokens, text, image, video, audio, and PDF inputs, text output, tool use, Gemini API availability, and $0.30 input / $2.50 output pricing per million tokens.
More Comparisons
Looking for more AI models?
Browse All LLMs