Skip to content

Robertomonteromiguel/llm validation gate - #12409

Merged
robertomonteromiguel merged 3 commits into
robertomonteromiguel/dd-apm-sdk-review-core-overridesfrom
robertomonteromiguel/llm-validation-gate
Sep 4, 2026
Merged

Robertomonteromiguel/llm validation gate#12409
robertomonteromiguel merged 3 commits into
robertomonteromiguel/dd-apm-sdk-review-core-overridesfrom
robertomonteromiguel/llm-validation-gate

Conversation

@robertomonteromiguel

Copy link
Copy Markdown
Contributor

What Does This Do

Motivation

Additional Notes

Contributor Checklist

Jira ticket: [PROJ-IDENT]

Replace the promptfoo suite with .llm-validation cases, wire the reusable GitLab job, and document local Docker runs against the published llmval image.

Co-authored-by: Cursor <cursoragent@cursor.com>
@robertomonteromiguel
robertomonteromiguel changed the base branch from master to robertomonteromiguel/dd-apm-sdk-review-core-overrides September 4, 2026 12:50
…into robertomonteromiguel/llm-validation-gate
@datadog-datadog-prod-us1-2

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 13.88 s 13.91 s [-0.9%; +0.4%] (no difference)
startup:insecure-bank:tracing:Agent 12.90 s 12.96 s [-1.3%; +0.3%] (no difference)

Commit: e73b497a · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@pr-commenter

pr-commenter Bot commented Sep 4, 2026

Copy link
Copy Markdown

LLM Validation

LLM Validation Gate — dd-apm-sdk-review

✅ PASS

  • Overall quality improved by 8.7 points with no blocking-case regressions.

Analysis

Changed instruction file(s): .agents/skills/dd-apm-sdk-review/SKILL.md, .agents/skills/dd-apm-sdk-review/reviewers/_common.md, .agents/skills/dd-apm-sdk-review/reviewers/correctness.md, .agents/skills/dd-apm-sdk-review/reviewers/performance.md, .agents/skills/dd-apm-sdk-review/reviewers/report-template.md, .agents/skills/dd-apm-sdk-review/reviewers/coherence.md, .agents/skills/dd-apm-sdk-review/reviewers/security.md, .agents/skills/dd-apm-sdk-review/reviewers/design.md, .agents/skills/dd-apm-sdk-review/reviewers/maintainability.md, .agents/skills/dd-apm-sdk-review/reviewers/conventions.md, .agents/skills/dd-apm-sdk-review/reviewers/cross-sdk.md, .agents/dd-apm-sdk-review-overrides/repo-context.md, .agents/dd-apm-sdk-review-overrides/reviewers/security.md, .agents/dd-apm-sdk-review-overrides/reviewers/performance.md, .agents/dd-apm-sdk-review-overrides/reviewers/design.md, .agents/dd-apm-sdk-review-overrides/reviewers/conventions.md, .agents/dd-apm-sdk-review-overrides/reviewers/maintainability.md.

No safety or blocking-case regressions across 8 case(s). Overall pairwise win-rate 70% [63%–78%], quality +8.7 — see the verdict above for whether that clears the noise band.

Results

  • Pairwise win-rate: 70% [63%–78%] — candidate's share of blind comparisons (90% CI; spanning 50% = no clear difference)
  • Overall quality: 78.4 → 87.1 (/100, +8.7)
  • Bad signals introduced (advisory): 0
  • Candidate criteria coverage (advisory): 33/33 (100%) — expected_criteria the candidate met; does not affect the gate
  • Blocking-case regressions: 0

Cases

Case Mode Quality Δ Win-rate (90% CI) Safety
java-perf-lens-wrong-collection-001 block +7.9 88% [67%–100%] ok
java-perf-pipeline-full-review-002 block +29.6 100% [100%–100%] ok
java-security-crash-handler-before-trust block +7.9 62% [42%–83%] ok
java-correctness-capture-before-send block -0.8 62% [42%–83%] ok
java-correctness-sqs-queue-name-incomplete block +5.0 50% [50%–50%] ok
java-maintainability-resource-leak-streams block +5.8 75% [51%–99%] ok
java-correctness-span-events-list-only block +11.7 75% [51%–99%] ok
java-correctness-mapper-state-leak block +2.5 50% [50%–50%] ok

Per-dimension scores, token usage, latency, and estimated cost are in the CI job logs.

@robertomonteromiguel
robertomonteromiguel marked this pull request as ready for review September 4, 2026 14:46
@robertomonteromiguel
robertomonteromiguel requested review from a team as code owners September 4, 2026 14:46
@robertomonteromiguel
robertomonteromiguel requested review from bric3 and randomanderson and removed request for a team September 4, 2026 14:46
@dd-octo-sts dd-octo-sts Bot added the tag: ai generated Largely based on code generated by an AI or LLM label Sep 4, 2026
@robertomonteromiguel
robertomonteromiguel merged commit 11fadad into robertomonteromiguel/dd-apm-sdk-review-core-overrides Sep 4, 2026
589 of 595 checks passed
@robertomonteromiguel
robertomonteromiguel deleted the robertomonteromiguel/llm-validation-gate branch September 4, 2026 14:47

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bb027dfa05

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +111 to +113
- "Flags that the JVM handler flag (`OnError`) is set to the target script/path BEFORE `CrashUploaderScriptInitializer.initialize(...)` — the ownership/trust validation — has run and returned a result. It is not enough to note that a validation step exists; the report must call out the ordering problem — if `initialize(...)` later rejects the path, the flag has already been committed to the untrusted path and the JVM crash handler will still execute it on a crash."
- "Classifies this as a security finding (P0 or P1, not merely a style/maintainability nit) — a security control that can be silently bypassed."
- "Recommends gating the flag-setter on the validator's result (e.g. `if (initialize(...)) { flags.setValue(...) }`) or clearing/restoring the flag on validation failure — not just \"add more logging\" or \"add a comment\"."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Stop grading public reports on private security details

When this security case runs, these criteria require the consolidated report to disclose the ordering flaw, execution consequence, and concrete remedy, while .agents/skills/dd-apm-sdk-review/SKILL.md:199-201 requires security findings in that report to omit the location, failure mode, and reproduction. A candidate that correctly strengthens private routing can therefore score worse or trip the “misses the ordering problem” bad signal; grade the public report on its private-routing placeholder and verify the detailed diagnosis only in a confidential artifact.

AGENTS.md reference: AGENTS.md:L79-L79

Useful? React with 👍 / 👎.

Comment thread .llm-validation/README.md
Comment on lines +112 to +114
| `minimum` | **1** (`java-perf-lens-wrong-collection-001`) | 3 | First smoke |
| `gate` (default) | **8** listed in `config.yaml` | 5 | CI-shaped |
| `full` | **every** case in `suites/` | 3 | Broader pass |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the documented preset run counts

The new configuration sets the minimum, gate, and full presets to two runs, but this table still documents 3, 5, and 3 respectively. Anyone using the README to reproduce the CI-shaped evaluation will run a different number of samples or incorrectly estimate its cost and variance; update all three table entries to match .llm-validation/config.yaml.

Useful? React with 👍 / 👎.

Comment on lines +135 to +138
+ // PR #12207: processCaptureExpressions() runs for every hit, regardless of
+ // whether the probe's condition/sampler has already decided this hit will
+ // not be sent. logStatus.shouldSend() reflects that effective send decision
+ // and is already computed by the caller before this method runs.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Remove the diagnosis from the capture-expression input

When this gate case evaluates a weakened correctness or performance reviewer, the model is already told that expressions run for rejected hits, that shouldSend() is the effective decision, and that it was computed before this method. The response can therefore restate the prompt and satisfy the criteria without deriving the defect from the reviewer rules, so the case cannot reliably detect a regression in those rules; retain only neutral call-path facts and leave the missing send gate for the reviewer to identify.

Useful? React with 👍 / 👎.

Comment thread .llm-validation/README.md

| Level | Cases | Default runs | Use |
|---|---|---|---|
| `minimum` | **1** (`java-perf-lens-wrong-collection-001`) | 3 | First smoke |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the minimum smoke exercise SKILL.md changes

The recommended minimum run selects only java-perf-lens-wrong-collection-001, whose files list contains _common.md, performance.md, and its override but not SKILL.md. When the change being evaluated is an orchestration-only edit to SKILL.md—one of the suite's stated monitored targets—baseline and candidate for this smoke have no changed instruction file, so it can pass without exercising the edit; select the full-pipeline case for minimum or add and invoke SKILL.md in the current smoke case.

Useful? React with 👍 / 👎.

- java-maintainability-resource-leak-streams
- java-correctness-span-events-list-only
- java-correctness-mapper-state-leak
runs: 2

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve enough gate repeats to reach the CI threshold

With the gate reduced to two repeats, even a blocking criterion that fails in both runs does not produce an upper binomial confidence bound below the configured blocking_fail_ci_upper: 0.55 (the upper bound remains above 0.55 with the platform's confidence calculation). Consequently, the blocking-failure confidence condition cannot be met and a consistently broken candidate can be downgraded to a warning instead of failing CI; keep at least five gate repeats or recalibrate the confidence policy for the smaller sample.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tag: ai generated Largely based on code generated by an AI or LLM

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant