Google Frontier Models

Gemini & Gemma
Pricing & Specs Dashboard

Compare model sizes, specifications, and modalities. Calculate production API costs in real-time with caching and batch optimizations.

Cost Calculator

Model custom prompt sizing & optimizations

Context Caching
90% off input cached
Batch API
50% asynchronous discount
1,000
Estimated Cost / Request
$0.0000
Input Tokens (250,000) $0.0000
Output Tokens (50,000) $0.0000
Daily Running Bill $0.00
Monthly Projected Expense $0.00

Cost Visualizer

Gemini Family pricing comparison

Comparing relative cost for processed workload: Input: 250,000 tokens | Output: 50,000 tokens

Ultra Economical (< 15%)
Balanced
Frontier Flagship (> 75%)

Model Specification Matrix

Explore model constraints, modalities, parameters, and license terms.

Model Name & Capability Context Length Modalities Supported API Cost per 1M Deployment & License

Understanding Context Caching

Google AI Studio and Vertex AI feature native context caching. Developers can store large, persistent documents, videos, source code, or system instructions in the model memory, and read from it across multiple API requests.

  • Massive Savings: Cached input tokens are billed at a 90% discount compared to base inputs.
  • Large Context Support: Excellent for codebases, long user histories, PDF libraries, or full movies.
  • Fast Response: Speeds up model latency significantly since large contexts do not need to be processed again.
📦

Leveraging Google's Batch API

For workloads that are not time-sensitive, developers can queue requests and receive responses asynchronously within 24 hours. This is perfect for high-volume offline classification, backtesting, translation, or batch analytics.

  • 50% Cost Cut: Directly slashes the rate for input and output tokens by half.
  • Huge Limits: Higher rate limits compared to the standard, live online pay-as-you-go endpoint.
  • Simple Pipeline: Ideal for processing massive backlogs of company documents or training files.