Day 0 Gemini 3.7 Flash release #36799
kaylaberriai
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Claude Opus 5 is now live on LiteLLM
Faster with better quality - Quicker responses than 3.6 Flash and higher scores on complex multi-step agentic, coding, and reasoning benchmarks.
Thinking levels via reasoning_effort - LiteLLM maps OpenAI reasoning_effort to Gemini’s thinkingLevel, so the same request shape you use for other reasoning models works here. The minimal level is not yet supported by the Gemini API at launch, and Google plans it as a fast follow.
1M context, 65K output - a 1,048,576-token context window and 65,536-token max output, with full multimodal input (text, image, audio, video) and function calling with thought signatures.
50% launch pricing through December 31, 2027 - $0.75 / MTok input and $3.75 / MTok output (standard is $1.50 / $7.50). Cache reads, batch, flex, and priority tiers are discounted proportionally, and LiteLLM tracks cost at the promotional rate.
Get started, no redeploy needed:
Gemini 3.7 Flash is already in the LiteLLM model cost map. Reload your pricing in the UI under Models + Endpoints -> Price Data -> Reload Price Data (or POST /reload/model_cost_map as an admin) and add the following model entry:
Running with LITELLM_LOCAL_MODEL_COST_MAP=true ghcr.io/berriai/litellm:v1.98.0-dev.2 will be released today, and that version (or any later version) adds Gemini 3.7 Flash support.
Read the full guide → Day 0 Support: Gemini 3.7 Flash
Mateo Wang
LiteLLM AI Team
All reactions