You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reasoning and coding close to larger frontier models - Google reports gains on long-horizon software engineering, stronger multi-step reasoning in finance and legal domains, and 54.9% on HLE-Verified.
Half price through December 31, 2026 - $0.75 / MTok input and $3.75 / MTok output, going to $1.50 / $7.50 on January 1, 2027. Cache reads are $0.075 / MTok, batch and flex run at half the standard rate, and LiteLLM tracks cost at the promotional rate.
1M context, 65K output - a 1,048,576-token context window and 65,536-token max output, with full multimodal input (text, image, audio, video), function calling, prompt caching from 4,096 tokens, and web search.
Get started, no redeploy needed:
Gemini 3.8 Flash is in the LiteLLM model cost map as of PR #39340. Reload your pricing in the UI under Models + Endpoints -> Price Data -> Reload Price Data (or POST /reload/model_cost_map as an admin) and add the following model entry.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Gemini 3.8 Flash is now live on LiteLLM
What's new:
Reasoning and coding close to larger frontier models - Google reports gains on long-horizon software engineering, stronger multi-step reasoning in finance and legal domains, and 54.9% on HLE-Verified.
Half price through December 31, 2026 - $0.75 / MTok input and $3.75 / MTok output, going to $1.50 / $7.50 on January 1, 2027. Cache reads are $0.075 / MTok, batch and flex run at half the standard rate, and LiteLLM tracks cost at the promotional rate.
1M context, 65K output - a 1,048,576-token context window and 65,536-token max output, with full multimodal input (text, image, audio, video), function calling, prompt caching from 4,096 tokens, and web search.
Get started, no redeploy needed:
Gemini 3.8 Flash is in the LiteLLM model cost map as of PR #39340. Reload your pricing in the UI under Models + Endpoints -> Price Data -> Reload Price Data (or POST /reload/model_cost_map as an admin) and add the following model entry.
Read the full guide → Day 0 Support: Gemini 3.8 Flash
Misbah Syed
DevRel @ LiteLLM
All reactions