Core package badges:
Quality and tooling:
Project/community:
Docs:
Estimate GenAI prompt costs from a unified, auto-updated pricing table. This repo provides a small usage-based cost estimator plus parsers for LiteLLM JSON and markdown pricing tables.
- Parses a local LiteLLM pricing snapshot when present, otherwise falls back to
genai_pricing.PRICING_URL - Computes costs from prompt, completion, cache-creation, and cache-read token usage
- Includes internal helpers for OpenAI/Gemini-style usage extraction and fallback token counting
- Python 3.8+
- Packages:
- tiktoken
pip install tiktokenThe included example shows how to estimate cost from an args-like object and token usage dictionary using genai_pricing.estimate_costs. See example.py.
# Minimal example
import os
from openai import OpenAI
from types import SimpleNamespace
from genai_pricing import estimate_costs
"""Estimate the cost of an OpenAI prompt using genai_pricing."""
api_key = os.environ.get("OPENAI_API_KEY")
client = OpenAI(api_key=api_key)
model = "gpt-5.6-sol"
prompt = "Why is the sky blue?"
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_completion_tokens=50,
)
answer = resp.choices[0].message.content
usage = {
"prompt_tokens": resp.usage.prompt_tokens,
"completion_tokens": resp.usage.completion_tokens,
"cache_read_input_tokens": resp.usage.prompt_tokens_details.cached_tokens,
}
args = SimpleNamespace(model=model)
estimate = estimate_costs(args, usage) # <- use this line in your project
print("Cost (USD):", estimate["total_cost"])Run the example:
python example.pyPrices are looked up by model name in the pricing table, then applied to token counts. prompt_tokens is the provider-reported total input count; cache creation and cache read counts are subsets charged in place of regular input:
-
$t_\text{in}$ : prompt tokens -
$t_\text{create}$ :cache_creation_input_tokens, when reported -
$t_\text{read}$ :cache_read_input_tokens, when reported -
$t_\text{out}$ : completion tokens -
$p$ : the corresponding USD price per 1M tokens
The result includes prompt_cost, cache_creation_cost, cache_read_cost, and completion_cost when applicable, plus total_cost. If a model has no cache-specific rate, cached tokens fall back to its regular input rate.
Provide token counts from your model provider when available. Internal helpers can extract OpenAI- and Gemini-style usage metadata and fall back to tiktoken or a lightweight heuristic when needed.
By default, prices are resolved in this order:
model_prices_and_context_window_backup.jsonin the current working directorymodel_prices_and_context_window_backup.jsonnear the package/repository locationgenai_pricing.PRICING_URL, the remote LiteLLM JSON source
The parsed pricing table is cached. Call clear_pricing_cache() after changing the local snapshot or when you want the remote source fetched again.
The project uses Python’s built-in unittest.
- Run all tests (discovery):
python -m unittest discover -s test -p "*_test.py" -vgenai_pricing.estimate_costs- Computes a dict with prompt/completion costs and
total_costfrom an object with a.modelattribute and a usage dictionary
- Computes a dict with prompt/completion costs and
genai_pricing.clear_pricing_cache- Clears cached pricing data so the configured source is fetched or read again
Key constant:
genai_pricing.PRICING_URL— remote table to fetch by default
MIT © 2025 Roberto Rossi