Repository navigation
Cost Reporting
What reuse saved, in seconds and in money, and why the money stays blank until you supply a price.
| Status | ⚠ Works, and ships inert by design: the money line stays blank until you supply a price |
| Verified | against the shipped library: with no price set, no dollar figure is reported and the reason is named |
| Function | merlin_get_savings |
Reuse saves GPU time. Turning that into a number your finance team will accept needs four things, and each is a different kind of claim:
- how many cache hits: counted exactly
- how many tokens those avoided: counted exactly
- how many GPU-seconds that is: measured on your hardware
- how many dollars: needs your price, and a statement of whether saved seconds become fewer billed seconds
⭐ Layers 1 to 3 report themselves. Layer 4 stays empty until you supply the price and say what your situation is.
⭐ Empty, not zero. A number that cannot be backed up is absent, never guessed.
Your platform team reports "we saved 40% of prefill work this month": a counted fact. Turning it into "we saved $8,000" needs your GPU price and a statement about whether your cluster was saturated. Until you provide both, the money line is blank and says why.
How it works
- Galahad counts every time saved memory was reused, and the GPU time that saved.
- With your GPU price, it turns that time into money.
- If a price or setting is missing, it shows no money figure and says what is missing. It never guesses.
How to use it
- Set your GPU price per hour, where the price comes from, and its date.
- Say how busy your GPUs are (the utilisation settings below).
- Read the result with
merlin_get_savings, or from the metrics when you use vLLM.
merlin_savings s;
memset(&s, 0, sizeof s);
s.struct_size = sizeof s;
merlin_status st = merlin_get_savings(avoided_seconds, added_seconds, &s);-
avoided_seconds: the GPU seconds that the reused prefills would have cost. You measure this on your own hardware. -
added_seconds: the GPU seconds Galahad itself spent. Pass the real value, not 0.
When st is MERLIN_OK:
| Field | Meaning |
|---|---|
net_gpu_seconds |
avoided seconds minus Galahad's own seconds |
dollars |
the money saved. Can be negative when Galahad's own cost is higher than the saving |
confidence |
from 0 to 1. Higher when more inputs are measured |
reason |
empty |
The refusals, and what they mean:
| Status | reason |
What to do |
|---|---|---|
MERLIN_ERR_NOT_FOUND |
price_and_utilisation_unset |
set the price settings and the utilisation mode |
MERLIN_ERR_NOT_FOUND |
price_unset |
set the GPU hour price, source and date |
MERLIN_ERR_NOT_FOUND |
utilisation_unset |
declare your regime |
MERLIN_ERR_INVALID_ARG |
config_error |
a setting is not valid; merlin_last_error() names it |
MERLIN_ERR_ABI_MISMATCH |
set struct_size = sizeof
|
|
MERLIN_ERR_INVALID_ARG |
out was null |
MERLIN_ERR_NOT_FOUND here is the normal state of an unconfigured deployment,
not an error to alert on.
⭐ out is zeroed before any early return, so a caller that ignores the
status reads 0.00 and an empty reason, never an old dollar figure from a
previous call.
export GALAHAD_FINOPS_GPU_HOUR_USD=2.40
export GALAHAD_FINOPS_PRICE_SOURCE="AWS p4d on-demand, eu-north-1"
export GALAHAD_FINOPS_PRICE_AS_OF=2026-09-17Plus a utilisation regime, below.
⭐ The price carries its source and its date, so a cost figure can be audited later. A price without a source or a date is refused.
⚠ A corrected price shows on the next scrape; no restart needed.
The last step turns avoided GPU-seconds into billed seconds. That fraction, R, depends on whether your cluster is saturated. If your GPUs are idle anyway, saving seconds does not reduce the bill.
R is a value you declare, not a measurement. Until you declare it, no dollar figure is given.
| Your cluster | Set | Meaning |
|---|---|---|
| GPUs are the bottleneck |
GALAHAD_FINOPS_REGIME=saturated, or GALAHAD_FINOPS_UTILISATION_MODE=measured with a measured GALAHAD_FINOPS_R
|
avoided seconds become billed seconds |
| GPUs have headroom | GALAHAD_FINOPS_REGIME=unsaturated |
avoided seconds free capacity, not spend |
| the cluster scales with load | GALAHAD_FINOPS_REGIME=autoscaled |
GPUs are added and removed with the load |
All the settings:
| Setting | Values | Meaning |
|---|---|---|
GALAHAD_FINOPS_GPU_HOUR_USD |
a number, 0 or more | your price for one GPU hour, in US dollars |
GALAHAD_FINOPS_PRICE_SOURCE |
text | where the price comes from. Required when a price is set |
GALAHAD_FINOPS_PRICE_AS_OF |
a date | when the price was valid. Required when a price is set |
GALAHAD_FINOPS_UTILISATION_MODE |
measured, derived or declared
|
how you know your utilisation. Required for a money figure |
GALAHAD_FINOPS_REGIME |
saturated, unsaturated or autoscaled
|
how busy your GPUs are |
GALAHAD_FINOPS_R |
a number from 0 to 1 | the share of saved GPU time that lowers your bill. Required when the mode is measured. When set, it is used instead of the regime |
GALAHAD_FINOPS_INDUCED_DEMAND |
a number from 0 to 1 | optional: a correction if you know how much extra work cheaper reuse created |
Values are not case sensitive. A value that is not valid is refused, never rounded or replaced.
⭐ Every figure says where it came from, so a reader can see which numbers are counted and which are declared. The seconds are measured; the money is your declared price multiplied by your declared regime.
With the vLLM Connector, the savings metrics appear on every metrics scrape:
| Metric | Meaning |
|---|---|
merlin_hit_tokens_total |
tokens served from memory |
merlin_prefills_avoided_total |
cache hits that avoided a prefill |
merlin_gpu_seconds_added_total |
GPU seconds Galahad itself spent |
merlin_savings_model_info |
the price, regime and source behind the figure. Its unavailable label names a missing input |
merlin_net_compute_dollars_saved |
money saved. Only present when every input is known |
merlin_savings_confidence |
from 0 to 1. Only present with the money figure |
⚠ An unconfigured deployment still shows no dollar figure, by design. You get the counted layers plus an info metric naming the missing input. Set the price settings above to see money.
⭐ merlin_stats.tokens_served_from_cache is the single source for tokens
served from memory, so every report quotes the same number.
| Symptom | Cause | Fix |
|---|---|---|
price_unset |
no GPU hour price | set the three price settings |
utilisation_unset |
no regime declared | say whether your cluster is saturated |
config_error |
a setting is not valid, or a price has no source or date | read merlin_last_error() and fix the setting |
| dollars read 0.00 | you ignored the status | check st and reason; zero is the safe default, not an answer |
dollars are 0.00 with MERLIN_OK
|
regime unsaturated
|
expected: saved time on spare GPUs does not lower the bill |
MERLIN_ERR_ABI_MISMATCH |
struct_size unset |
s.struct_size = sizeof s |
| cost metrics missing from Prometheus | a malformed GALAHAD_FINOPS_* value |
the startup log names it |
- API Reference
- Running on Kubernetes: where the other metrics live
- vLLM Connector
- Install