Peter/model cache pricing - #47
Merged
Merged
Conversation
…column fix: show cache read multiplier in catalog
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Pull request overview
Updates the auto-generated Inference API model catalog markdown to present prompt-caching “cache read” pricing as a multiplier (instead of a per-1M token price), aligning the index page display with the underlying cache pricing model data.
Changes:
- Adjust the models index table “Cache read” column to display
cacheReadMultiplier(e.g.,0.1×) and use a visual fallback for unsupported models. - Update the prompt-caching callout text to describe the new multiplier-based meaning of the column.
- Update/extend unit tests to assert the new header formatting, alignment row, and multiplier/fallback rendering.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| src/lib/model-artifacts.ts | Changes the generated models index markdown table to show cache-read multipliers and updates accompanying explanatory text. |
| src/lib/model-artifacts.test.ts | Updates expectations to validate the new cache-read column header/alignment and the multiplier/unsupported rendering. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
226
to
+229
| The table below outlines the models we have available and their pricing per 1M tokens. If you are interested in understanding pricing for a model not listed below or if you'd like to request a new model - please reach out to support@doubleword.ai. | ||
|
|
||
| :::info{title="Prompt caching"} | ||
| Prompt-caching availability and rates are model-specific. The **Cache read** column shows the current cache-read price for supported models. See the [prompt caching guide](/inference-api/prompt-caching) for setup, TTLs, and write pricing. | ||
| Prompt-caching availability and rates are model-specific. The **Cache read** column shows the current multiplier on the model's standard input price. See the [prompt caching guide](/inference-api/prompt-caching) for setup, TTLs, and write pricing. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.