LLM Comparison
Gemini 3.5 Flash-Lite vs Nemotron 3.5 Lightning 30B A3B
Side-by-side specs, pricing & capabilities · Updated August 2026
Price vs Intelligence
Add to comparison
2/6 models| Organization | ||
| OpenTools Score | 14 9.8 | 6 12.0 |
| Family | Gemini | Nemotron |
| Status | Current | Current |
| Release Date | Jul 2026 | Aug 2026 |
| Context Window | 1.0M tokens | 1.0M tokens |
| Input Price | $0.30/M tokens | $0.10/M tokens |
| Output Price | $2.50/M tokens | $0.95/M tokens |
| Pricing Notes | DeepMind model page lists $0.30 input and $2.50 output per 1M tokens with no caching in the performance table. The ai.google.dev API docs were attempted through ScrapingBee and returned HTTP 400, so exact billing details beyond the official DeepMind price table were not used. | NVIDIA build lists serverless NIM pricing at $0.10 input and $0.95 output per million tokens. Self-hosted weights are also available; infrastructure costs vary. |
| Capabilities | textvisionaudio-inputvideo-inputpdf-inputreasoningcodingtool-usesearch-groundingcomputer-usestructured-output | textreasoningcodingtool-usefunction-callinglong-contextagenticlocal-deploymentopen-weights |
| Training Cutoff | — | May 2026 post-training; September 2025 pre-training |
| Max Output | 64K tokens | 33K tokens |
| API Identifier | gemini-3.5-flash-lite | nvidia/nemotron-3.5-lightning-30b-a3b |
| Benchmarks | ||
| Artificial Analysis Intelligence Index | 36artificial-analysis | — |
| SWE-Bench Pro | 54.2google-deepmind | — |
| Terminal-Bench 2.1 | 54google-deepmind | — |
| GDPVal-AA v2 | 1140google-deepmind | — |
| Artificial Analysis Intelligence Index v4.1.1 | — | 24artificial-analysis |
| View Gemini 3.5 Flash-Lite | View Nemotron 3.5 Lightning 30B A3B | |
Cost Calculator
Enter your expected monthly token usage to compare costs.
| Model | Input | Output | Total / mo | vs Best |
|---|---|---|---|---|
| Nemotron 3.5 Lightning 30B A3BCheapest | $0.10 | $0.48 | $0.58 | — |
| Gemini 3.5 Flash-Lite | $0.30 | $1.25 | $1.55 | +170% |
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is Google DeepMind general availability low-latency model for high-volume agentic systems, coding, document parsing, translation, and structured automation. Official DeepMind evidence lists 1M input tokens, 64k output tokens, text, image, video, audio, and PDF inputs, text output, tool use, Gemini API availability, and $0.30 input / $2.50 output pricing per million tokens.
NVIDIA
Nemotron 3.5 Lightning 30B A3B
Nemotron 3.5 Lightning is NVIDIA open-weight 30B-A3B reasoning model for fast, long-running agents. Its hybrid Mamba-2, mixture-of-experts, and attention architecture supports function calling, coding, tool use, long context, and efficient local or serverless deployment.
More Comparisons
Looking for more AI models?
Browse All LLMs