Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
3377fa6
Add live GitHub star counter to docs header
May 19, 2026
efe7b89
docs(evaluation): add Set up guardrails how-to guide
abhijaisrivastava15 Jul 9, 2026
4974826
docs(evaluation): add screenshot placeholders for guardrails guide
abhijaisrivastava15 Jul 9, 2026
f59f1c7
docs(evaluation): add annotated screenshots for guardrails guide
abhijaisrivastava15 Jul 9, 2026
0b44ee0
docs(evaluation): fix dead Dive deeper link (evaluate -> guides/runni…
abhijaisrivastava15 Jul 10, 2026
c31d166
docs(evaluation): add SDK method and full monitoring tabs to guardrai…
abhijaisrivastava15 Jul 15, 2026
5f04515
docs(evaluation): Explore playground overview
Jul 15, 2026
5b8bddd
docs(evaluation): add configure-flow GIF to guardrails guide
abhijaisrivastava15 Jul 15, 2026
5f4801c
docs(evaluation): rewrite custom models guide to house style with rea…
abhijaisrivastava15 Jul 15, 2026
15eccf1
docs(evaluation): annotate Create custom model callout on custom mode…
abhijaisrivastava15 Jul 15, 2026
a80d27e
docs(evaluation): annotate remaining guardrails screenshots
abhijaisrivastava15 Jul 15, 2026
96ff272
docs(evaluation): annotate remaining custom models screenshots
abhijaisrivastava15 Jul 15, 2026
0192180
docs(evaluation): Test an eval
Jul 15, 2026
e6c5ecc
docs(evaluation): address review on Test an eval
Jul 15, 2026
74f4d08
docs(evaluation): recapture Test an eval screenshots at uniform size
Jul 15, 2026
c337a16
docs(evaluation): Usage & analytics
Jul 15, 2026
b624486
docs(evaluation): address review on Usage & analytics
Jul 15, 2026
a8d16e2
docs(evaluation): recapture Usage tab screenshot at matching size
Jul 15, 2026
26d80d4
Merge branch 'docs/evaluation-revamp' into docs/eval-explore-playgrou…
khushalsonawat Jul 16, 2026
f021e82
docs(evaluation): fix Output Type and tag scope on Explore playground…
khushalsonawat Jul 16, 2026
fe7329b
Fixing sidebar also
khushalsonawat Jul 16, 2026
4c07efe
Merge branch 'docs/evaluation-revamp' into docs/eval-explore-playgrou…
khushalsonawat Jul 16, 2026
c0c94f6
docs(evaluation): rework Explore playground overview from reader feed…
khushalsonawat Jul 16, 2026
a0d6906
Merge pull request #751 from future-agi/docs/eval-explore-playground-…
khushalsonawat Jul 16, 2026
dae7083
Merge pull request #753 from future-agi/docs/eval-explore-playground-…
khushalsonawat Jul 16, 2026
f3bd610
docs(evaluation): add the missing select-the-eval step to Test an eval
khushalsonawat Jul 16, 2026
99f7eab
Merge branch 'docs/evaluation-revamp' into docs/eval-explore-playgrou…
khushalsonawat Jul 16, 2026
90d8e90
Fixing the analytics page
khushalsonawat Jul 16, 2026
bfca7d6
Merge pull request #754 from future-agi/docs/eval-explore-playground-…
khushalsonawat Jul 16, 2026
1dd97f4
docs(evaluation): add Collect feedback guide
SuhaniNagpal7 Jul 16, 2026
890a685
Merge branch 'docs/evaluation-revamp' into docs/collect-feedback
suhani-725 Jul 16, 2026
6d5c316
Update navigation.ts
khushalsonawat Jul 16, 2026
0d2a1ba
docs(evaluation): clarify feedback eligibility and relabel re-scoring…
khushalsonawat Jul 16, 2026
e67c6e3
Merge pull request #767 from future-agi/docs/collect-feedback
khushalsonawat Jul 16, 2026
8612388
Merge branch 'docs/evaluation-revamp' into docs/eval-guide-guardrails
khushalsonawat Jul 16, 2026
5e4a498
docs(evaluation): rework custom models guide from review feedback
khushalsonawat Jul 16, 2026
a7949e5
docs(evaluation): center Set up guardrails on the Protect SDK
khushalsonawat Jul 16, 2026
9810075
Merge pull request #752 from future-agi/docs/eval-guide-custom-models
khushalsonawat Jul 16, 2026
2afe06b
Merge pull request #729 from future-agi/docs/eval-guide-guardrails
khushalsonawat Jul 16, 2026
cf2165d
Merge pull request #764 from future-agi/docs/evaluation-revamp
khushalsonawat Jul 16, 2026
bf285d0
Merge pull request #658 from future-agi/fix/star-on-github-counter
khushalsonawat Jul 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file removed public/images/custom-model/1.png
Binary file not shown.
Binary file removed public/images/custom-model/2.png
Binary file not shown.
Binary file removed public/images/custom-model/3.png
Binary file not shown.
Binary file removed public/images/custom-model/4.png
Binary file not shown.
Binary file removed public/images/custom-model/5.png
Binary file not shown.
Binary file removed public/images/custom-model/6.png
Binary file not shown.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
55 changes: 50 additions & 5 deletions src/components/Header.astro
Original file line number Diff line number Diff line change
Expand Up @@ -86,16 +86,24 @@ const currentPath = Astro.url.pathname;

<!-- Right Actions -->
<div class="flex items-center gap-3 flex-shrink-0">
<!-- Star on GitHub -->
<a
href="https://github.com/future-agi"
href="https://github.com/future-agi/future-agi"
target="_blank"
rel="noopener noreferrer"
aria-label="GitHub"
class="hidden sm:inline-flex items-center p-1.5 text-[var(--color-text-secondary)] hover:text-[var(--color-text-primary)] transition-colors"
class="hidden sm:inline-flex items-center gap-1.5 px-2.5 py-1.5 text-sm text-[var(--color-text-secondary)] hover:text-[var(--color-text-primary)] transition-colors group"
aria-label="Star on GitHub"
>
<svg class="w-5 h-5" fill="currentColor" viewBox="0 0 24 24" aria-hidden="true">
<path fill-rule="evenodd" clip-rule="evenodd" d="M12 2C6.477 2 2 6.484 2 12.017c0 4.425 2.865 8.18 6.839 9.504.5.092.682-.217.682-.483 0-.237-.008-.868-.013-1.703-2.782.605-3.369-1.343-3.369-1.343-.454-1.158-1.11-1.466-1.11-1.466-.908-.62.069-.608.069-.608 1.003.07 1.531 1.032 1.531 1.032.892 1.53 2.341 1.088 2.91.832.092-.647.35-1.088.636-1.338-2.22-.253-4.555-1.113-4.555-4.951 0-1.093.39-1.988 1.029-2.688-.103-.253-.446-1.272.098-2.65 0 0 .84-.27 2.75 1.026A9.564 9.564 0 0112 6.844c.85.004 1.705.115 2.504.337 1.909-1.296 2.747-1.027 2.747-1.027.546 1.379.202 2.398.1 2.651.64.7 1.028 1.595 1.028 2.688 0 3.848-2.339 4.695-4.566 4.943.359.309.678.92.678 1.855 0 1.338-.012 2.419-.012 2.747 0 .268.18.58.688.482A10.02 10.02 0 0022 12.017C22 6.484 17.522 2 12 2z"/>
<span>Star on</span>
<svg class="w-[18px] h-[18px]" fill="currentColor" viewBox="0 0 24 24">
<path fill-rule="evenodd" clip-rule="evenodd" d="M12 2C6.477 2 2 6.484 2 12.017c0 4.425 2.865 8.18 6.839 9.504.5.092.682-.217.682-.483 0-.237-.008-.868-.013-1.703-2.782.605-3.369-1.343-3.369-1.343-.454-1.158-1.11-1.466-1.11-1.466-.908-.62.069-.608.069-.608 1.003.07 1.531 1.032 1.531 1.032.892 1.53 2.341 1.088 2.91.832.092-.647.35-1.088.636-1.338-2.22-.253-4.555-1.113-4.555-4.951 0-1.093.39-1.988 1.029-2.688-.103-.253-.446-1.272.098-2.65 0 0 .84-.27 2.75 1.026A9.564 9.564 0 0112 6.844c.85.004 1.705.115 2.504.337 1.909-1.296 2.747-1.027 2.747-1.027.546 1.379.202 2.398.1 2.651.64.7 1.028 1.595 1.028 2.688 0 3.848-2.339 4.695-4.566 4.943.359.309.678.92.678 1.855 0 1.338-.012 2.419-.012 2.747 0 .268.18.58.688.482A10.019 10.019 0 0022 12.017C22 6.484 17.522 2 12 2z" />
</svg>
<span class="inline-flex items-center gap-1 px-1.5 py-0.5 rounded-md bg-[var(--color-bg-tertiary)] border border-[var(--color-border-subtle)] text-[11px] text-[var(--color-text-secondary)] group-hover:border-[var(--color-border-accent)] transition-colors min-w-[28px] justify-center">
<svg class="w-2.5 h-2.5" style="color: #fbbf24;" fill="currentColor" viewBox="0 0 24 24">
<path d="M12 2l3.09 6.26L22 9.27l-5 4.87 1.18 6.88L12 17.77l-6.18 3.25L7 14.14 2 9.27l6.91-1.01L12 2z"/>
</svg>
<span data-github-stars>—</span>
</span>
</a>

<a
Expand Down Expand Up @@ -257,5 +265,42 @@ const currentPath = Astro.url.pathname;

setupHeader();
document.addEventListener('astro:page-load', setupHeader);

// GitHub star counter - fetches live count, caches for 10 min
function formatStars(count) {
if (count >= 1000) {
return (count / 1000).toFixed(count >= 10000 ? 0 : 1).replace(/\.0$/, '') + 'k';
}
return String(count);
}

async function loadGitHubStars() {
const targets = document.querySelectorAll('[data-github-stars]');
if (!targets.length) return;
const CACHE_KEY = 'gh-stars-future-agi';
const CACHE_TTL = 10 * 60 * 1000;
try {
const cached = localStorage.getItem(CACHE_KEY);
if (cached) {
const { count, ts } = JSON.parse(cached);
if (Date.now() - ts < CACHE_TTL && typeof count === 'number') {
targets.forEach(el => { el.textContent = formatStars(count); });
return;
}
}
} catch {}
try {
const res = await fetch('https://api.github.com/repos/future-agi/future-agi');
if (!res.ok) throw new Error('GitHub API error');
const data = await res.json();
const count = data.stargazers_count || 0;
targets.forEach(el => { el.textContent = formatStars(count); });
try { localStorage.setItem(CACHE_KEY, JSON.stringify({ count, ts: Date.now() })); } catch {}
} catch {
targets.forEach(el => { el.textContent = '★'; });
}
}
loadGitHubStars();
document.addEventListener('astro:page-load', loadGitHubStars);
})();
</script>
10 changes: 10 additions & 0 deletions src/lib/navigation.ts
Original file line number Diff line number Diff line change
Expand Up @@ -324,9 +324,19 @@ export const tabNavigation: NavTab[] = [
title: 'Guides',
items: [
{ title: 'Running Evaluations', href: '/docs/evaluation/guides/running-evaluations' },
{
title: 'Explore playground',
items: [
{ title: 'The Evaluations page', href: '/docs/evaluation/guides/explore-playground' },
{ title: 'Test an eval', href: '/docs/evaluation/guides/explore-playground/test-an-eval' },
{ title: 'Usage & analytics', href: '/docs/evaluation/guides/explore-playground/usage-analytics' },
]
},
{ title: 'Create a custom eval', href: '/docs/evaluation/guides/custom-evals' },
{ title: 'Build a composite evals', href: '/docs/evaluation/guides/composite-evals' },
{ title: 'Set up guardrails', href: '/docs/evaluation/guides/guardrails' },
{ title: 'Add ground truth', href: '/docs/evaluation/guides/ground-truth' },
{ title: 'Collect feedback', href: '/docs/evaluation/guides/collect-feedback' },
{ title: 'Use custom models', href: '/docs/evaluation/guides/custom-models' },
{ title: 'Evaluate in CI/CD', href: '/docs/evaluation/guides/cicd' },
{ title: 'Advanced usage', href: '/docs/evaluation/guides/advanced-usage' },
Expand Down
132 changes: 132 additions & 0 deletions src/pages/docs/evaluation/guides/collect-feedback.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,132 @@
---
title: "Collect feedback"
description: "Teach an evaluator your standard by correcting the results it gets wrong"
---

[Feedback](/docs/evaluation/concepts/feedback) is a correction you record on an eval result, which every later run retrieves and shows to the [evaluator model](/docs/evaluation/concepts/evaluator-models) as an example before it scores. This guide corrects a result on a dataset, chooses what gets re-scored, and reviews everything your team has corrected.

You usually work in batches: correct the results an eval got wrong across a run, then re-run and watch it converge on your team's judgment. A single correction nudges the evaluator; a batch is what moves it, and most evals settle in a few rounds.

<Note>
Feedback isn't available for [code evals](/docs/evaluation/concepts/eval-types) (no evaluator model to steer), [composite evals](/docs/evaluation/concepts/composite-evals) (a roll-up of child scores), raw-number metrics (a computed value with no pass or fail to correct, unlike a Score eval's gradable 0-100), or results in an error state (no score to correct). Code and composite evals also have no Feedback tab.
</Note>

## Collect feedback on a dataset

This example corrects an eval result on a [dataset](/docs/dataset), then chooses what gets re-scored.

### Open the drawer

- Hover any result in an eval column to see the reason the eval gave
- Click **Add feedback** in the popover under that reason

You can also open a row and click **Add Feedback** on the eval in the datapoint drawer. Either way the drawer names the eval and repeats its explanation, so you're correcting against what it actually said rather than from memory.

<img src="/images/docs/evaluation/guides/collect-feedback/dataset-add-feedback-hover.png" alt="Hovering an eval result in a dataset grid, with the reason in a popover and an Add feedback button beneath it" style={{ borderRadius: '5px' }} />
*Hover a result to see the reason the eval gave, and the button that opens the drawer*

### Correct the result

The first field takes the shape of the eval's [output type](/docs/evaluation/reference/output-types):

| Output type | Label | What you enter |
|---|---|---|
| Pass/Fail, or a single choice | **Select a right value** | The verdict it should have returned |
| Multiple choices | **Select the right value(s)** | Every label that should've applied |
| Score | **Write a right value** | A number between 0 and 100 |
| Reason | **Write a right value** | The corrected text |

Then write the explanation the eval should have given. This is the field that teaches it, so name the rule you're applying instead of restating the verdict. The value and the explanation are both required.

<img src="/images/docs/evaluation/guides/collect-feedback/dataset-correct-result.png" alt="The Add feedback drawer for a score eval, with the eval's own explanation on top, a Write a right value number field, and an explanation box" style={{ borderRadius: '5px' }} />
*The drawer repeats the eval's explanation above the fields, so you correct against what it said*

### Choose what gets re-scored

Every option stores your correction against the eval. What they differ on is how much gets re-scored:

| Option | What it does |
|---|---|
| **Re-tune** | Stores the correction. Nothing is re-scored, and later runs pick it up |
| **Re-calculate for this row** | Stores it, then re-runs the eval on this row |
| **Re-tune and re-calculate for this dataset** | Stores it, then re-runs the eval on every run in the dataset |

The last two re-score existing results, so they take a while on a big dataset; the eval column updates in place as each result finishes, so you can watch it there. Reach for **Re-tune** when you're labelling a batch of corrections and only want them counting from the next run onward.

Whichever you pick, the eval's own criteria stay as they are: your correction is stored and pulled into later runs as an example.

<img src="/images/docs/evaluation/guides/collect-feedback/dataset-rescore-options.png" alt="The filled Add feedback drawer showing the three re-scoring options with Re-calculate for this row selected, and Submit feedback" style={{ borderRadius: '5px' }} />
*Pick what gets re-scored, then submit*

## Collect feedback in the eval playground

You can also correct a result straight from the eval's own page, without opening a dataset. The steps match the dataset flow, apart from three things: the field labels differ, there are two re-scoring options instead of three, and the row gets a thumb once you submit.

### Open the drawer from the Usage tab

- From **Evals**, open an eval and go to its **Usage** tab
- Click a row to open its panel
- Click **Add Feedback**, or **Edit Feedback** if the row already carries one

<img src="/images/docs/evaluation/guides/collect-feedback/playground-add-feedback.png" alt="An eval's Usage tab with a result's panel open, showing its score and reason and an Add Feedback button" style={{ borderRadius: '5px' }} />
*Open a result's panel on the Usage tab, then click Add Feedback*

### Enter your correction

The drawer here is titled **Feedbacks for Auto Learning**. It has the same two fields as a dataset, under different labels:

- Pick the verdict under **Choose a right value**, or **Write a right value** for a score or text eval
- Fill in **What would you like to improve** with why the result was wrong

<img src="/images/docs/evaluation/guides/collect-feedback/playground-correct-result.png" alt="The Feedbacks for Auto Learning drawer with a Choose a right value Passed or Failed field and a What would you like to improve box" style={{ borderRadius: '5px' }} />
*The playground drawer, with the same two fields under different labels*

### Pick a re-scoring option

Both store the correction; the difference is whether past runs get re-scored too:

| Option | What it does |
|---|---|
| **Re-tune** | Stores the correction for later runs |
| **Re-calculate and re-tune** | Stores it, then re-scores every past run of this eval |

Submit, and the row gets a thumb on the Usage tab: a green thumbs-up where you answered passed, a red thumbs-down where you answered failed. That's the quickest way to see which results you've already been through.

<img src="/images/docs/evaluation/guides/collect-feedback/playground-rescore-options.png" alt="The filled playground drawer with Failed chosen, an improvement note written, and the Re-tune and Re-calculate and re-tune options" style={{ borderRadius: '5px' }} />
*Pick one of the two options, then submit feedback*

## Review your corrections

Every correction on an eval collects on its **Feedback** tab, whichever surface it came from. That tab is the record of what your team has taught the evaluator, so it's where you work from when you want to know whether it's converging. Open an eval and click **Feedback**.

**Feedback History** lists one row per correction:

| Column | What it shows |
|---|---|
| **Feedback** | The verdict you gave, as a **Correct** or **Incorrect** chip |
| **Improvement Note** | The explanation you wrote |
| **Action** | **Re-tune** or **Re-calculate**, whichever you picked |
| **Source** | Where it came from, **Dataset** or **Playground** |
| **By** | Who submitted it |
| **Date** | When |

A colored bar on the left edge of each row repeats the verdict at a glance. Before anyone has corrected the eval, the tab reads "No feedback submitted yet".

Click a row to open the whole entry beside the list: the improvement note in full, the log ID it came from, and a read-only **Raw Data** view of the stored record. Step through entries with **j** and **k**, close with **Escape**, and click **Edit Feedback** to change one.

<img src="/images/docs/evaluation/guides/collect-feedback/feedback-history.png" alt="The Feedback tab's history table with a correction's full entry open on the right, showing its improvement note, action, source, and raw data" style={{ borderRadius: '5px' }} />
*The Feedback tab lists every correction, with the full entry open on the right*

## Dive deeper

<CardGroup cols={3}>
<Card title="Eval correction loop" icon="arrows-rotate" href="/docs/cookbook/evaluation/eval-correction-loop">
A worked example of turning corrections into a better eval
</Card>
<Card title="Create a custom eval" icon="wand-magic-sparkles" href="/docs/evaluation/guides/custom-evals">
Encode your corrections as rules the evaluator follows
</Card>
<Card title="Ground truth" icon="database" href="/docs/evaluation/concepts/ground-truth">
Calibrate an eval against reference rows instead of corrections
</Card>
</CardGroup>
Loading
Loading