Skip to content

ci: exclude release-proposal wfs from green_ci and exclude benchmarks on release* branches - #2206

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 5 commits into
mainfrom
igor/versioning/green_ci
Jul 16, 2026
Merged

ci: exclude release-proposal wfs from green_ci and exclude benchmarks on release* branches#2206
gh-worker-dd-mergequeue-cf854d[bot] merged 5 commits into
mainfrom
igor/versioning/green_ci

Conversation

@iunanua

@iunanua iunanua commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

What does this PR do?

  • Exclude from Green CI release-proposal-dispatch.yml and release-proposal-test.yml
  • Exclude benchmarks on release, release-proposal, release-* branches

@iunanua iunanua changed the title ci: exclude release-proposal wfs from green_ci and exclude benchmars on release* branches ci: exclude release-proposal wfs from green_ci and exclude benchmarks on release* branches Jul 7, 2026
@iunanua
iunanua marked this pull request as ready for review July 7, 2026 15:38
@iunanua
iunanua requested a review from a team as a code owner July 7, 2026 15:38

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2b545d85a3

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .gitlab/benchmarks.yml Outdated
Comment thread .github/workflows/release-proposal-dispatch.yml
@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

Tests

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
Patch Coverage: 100.00%
Overall Coverage: 74.73% (-0.00%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 0bdafcf | Docs | Datadog PR Page | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

Artifact Size Benchmark Report

aarch64-alpine-linux-musl
Artifact Baseline Commit Change
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.so 7.88 MB 7.88 MB 0% (0 B) 👌
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.a 85.91 MB 85.91 MB 0% (0 B) 👌
aarch64-unknown-linux-gnu
Artifact Baseline Commit Change
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.so 10.61 MB 10.61 MB 0% (0 B) 👌
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.a 97.11 MB 97.11 MB 0% (0 B) 👌
libdatadog-x64-windows
Artifact Baseline Commit Change
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.dll 25.46 MB 25.46 MB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.lib 88.44 KB 88.44 KB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.pdb 184.60 MB 184.60 MB 0% (0 B) 👌
/libdatadog-x64-windows/debug/static/datadog_profiling_ffi.lib 946.40 MB 946.40 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.dll 8.32 MB 8.32 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.lib 88.44 KB 88.44 KB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.pdb 24.62 MB 24.62 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/static/datadog_profiling_ffi.lib 49.04 MB 49.04 MB 0% (0 B) 👌
libdatadog-x86-windows
Artifact Baseline Commit Change
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.dll 22.06 MB 22.06 MB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.lib 89.82 KB 89.82 KB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.pdb 188.61 MB 188.62 MB +0% (+16.00 KB) 👌
/libdatadog-x86-windows/debug/static/datadog_profiling_ffi.lib 935.37 MB 935.37 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.dll 6.43 MB 6.43 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.lib 89.82 KB 89.82 KB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.pdb 26.43 MB 26.43 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/static/datadog_profiling_ffi.lib 46.65 MB 46.65 MB 0% (0 B) 👌
x86_64-alpine-linux-musl
Artifact Baseline Commit Change
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.a 76.59 MB 76.59 MB 0% (0 B) 👌
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.so 8.78 MB 8.78 MB 0% (0 B) 👌
x86_64-unknown-linux-gnu
Artifact Baseline Commit Change
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.a 92.11 MB 92.11 MB 0% (0 B) 👌
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.so 10.69 MB 10.69 MB 0% (0 B) 👌

iunanua and others added 3 commits July 8, 2026 09:56
Replace `npm install -g @datadog/datadog-ci` in the Green CI exclusion
steps with a download of the pinned v5.21.0 release binary verified
against its SHA-256 checksum, matching the pattern in test.yml. This
avoids running unpinned third-party npm install lifecycle hooks with the
DATADOG_API_KEY secret in their environment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment on lines +53 to +71
# Download the pinned, checksum-verified datadog-ci binary instead of
# `npm install -g`, so no third-party install-time code runs with the
# Datadog API key in its environment.
URL="https://github.com/DataDog/datadog-ci/releases/download/v5.21.0/datadog-ci_linux-x64"
OUTPUT="datadog-ci"
EXPECTED_CHECKSUM="be4a6473fc451fec967ff277df3856060814b9a54d707d055a9c1542ae2869f0"

echo "Downloading datadog-ci from $URL"
curl -L --fail --retry 3 -o "$OUTPUT" "$URL"
chmod +x "$OUTPUT"

ACTUAL_CHECKSUM=$(sha256sum "$OUTPUT" | cut -d' ' -f1)
if [ "$ACTUAL_CHECKSUM" != "$EXPECTED_CHECKSUM" ]; then
echo "Checksum verification failed! expected=$EXPECTED_CHECKSUM actual=$ACTUAL_CHECKSUM"
exit 1
fi
echo "Checksum verification passed"

./"$OUTPUT" tag --level pipeline --tags green_ci.excluded:true

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this code is run in two workflows, Wouldn't it be better to put this code in scripts and run it from the workflow?

@iunanua
iunanua force-pushed the igor/versioning/green_ci branch from cdd6207 to 3ec993c Compare July 16, 2026 14:03
The identical datadog-ci download + pipeline tagging step was duplicated in
release-proposal-test.yml and release-proposal-dispatch.yml. Move it into
scripts/exclude-from-green-ci.sh and call it from both. The step now runs
after checkout (so the script is on disk) instead of before.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@iunanua
iunanua force-pushed the igor/versioning/green_ci branch from 3ec993c to 0bdafcf Compare July 16, 2026 14:04
@iunanua

iunanua commented Jul 16, 2026

Copy link
Copy Markdown
Contributor Author

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Jul 16, 2026

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-07-16 14:05:57 UTC ℹ️ Start processing command /merge


2026-07-16 14:06:04 UTC ℹ️ MergeQueue: waiting for PR to be ready

This pull request is not mergeable according to GitHub. Common reasons include pending required checks, missing approvals, or merge conflicts — but it could also be blocked by other repository rules or settings.
It will be added to the queue as soon as checks pass and/or get approvals. View in MergeQueue UI.
Note: if you pushed new commits since the last approval, you may need additional approval.
You can remove it from the waiting list with /remove command.


2026-07-16 15:30:16 UTC ℹ️ MergeQueue: merge request added to the queue

The expected merge time in main is approximately 1h (p90).


2026-07-16 16:13:55 UTC ℹ️ MergeQueue: This merge request was merged

@pr-commenter

pr-commenter Bot commented Jul 16, 2026

Copy link
Copy Markdown

Benchmarks

Comparison

Benchmark execution time: 2026-07-16 14:39:37

Comparing candidate commit 0bdafcf in PR branch igor/versioning/green_ci with baseline commit 4a914bb in branch main.

Found 10 performance improvements and 18 performance regressions! Performance is the same for 120 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:credit_card/is_card_number/x371413321323331

  • 🟥 execution_time [+1.138µs; +1.141µs] or [+19.945%; +19.989%]
  • 🟥 throughput [-29203171.687op/s; -29135607.882op/s] or [-16.663%; -16.625%]

scenario:credit_card/is_card_number_no_luhn/ 3782-8224-6310-005

  • 🟥 execution_time [+4.993µs; +5.117µs] or [+8.199%; +8.403%]
  • 🟥 throughput [-1276193.648op/s; -1242917.655op/s] or [-7.771%; -7.568%]

scenario:credit_card/is_card_number_no_luhn/x371413321323331

  • 🟥 execution_time [+1.139µs; +1.142µs] or [+19.972%; +20.014%]
  • 🟥 throughput [-29236161.840op/s; -29174648.640op/s] or [-16.680%; -16.645%]

scenario:msgpack_decoder::v05/high_sharing/200

  • 🟩 execution_time [-6.953µs; -6.846µs] or [-4.127%; -4.064%]
  • 🟩 throughput [+50303.728op/s; +51097.456op/s] or [+4.237%; +4.304%]

scenario:msgpack_decoder::v05/low_sharing/2000

  • 🟩 execution_time [-156.210µs; -155.206µs] or [-4.268%; -4.240%]
  • 🟩 throughput [+24199.844op/s; +24358.109op/s] or [+4.429%; +4.458%]

scenario:profile_serialize_compressed_pprof_timestamped_x1000

  • 🟩 execution_time [-280.556µs; -277.017µs] or [-22.836%; -22.548%]

scenario:sql/obfuscate_sql_string

  • 🟩 execution_time [-13.799µs; -13.565µs] or [-4.436%; -4.360%]

scenario:vec_map/as_deduped_map/already_deduped/16

  • 🟥 execution_time [+7.832ns; +7.889ns] or [+33.287%; +33.531%]

scenario:vec_map/as_deduped_map/already_deduped/8

  • 🟥 execution_time [+0.939ns; +0.954ns] or [+6.763%; +6.874%]

scenario:vec_map/contains_key/16

  • 🟥 execution_time [+11.840ns; +12.130ns] or [+4.728%; +4.844%]
  • 🟥 throughput [-2954358.445op/s; -2883112.075op/s] or [-4.624%; -4.512%]

scenario:vec_map/contains_key/8

  • 🟥 execution_time [+10.796ns; +11.086ns] or [+14.890%; +15.290%]
  • 🟥 throughput [-14630624.017op/s; -14274161.025op/s] or [-13.260%; -12.937%]

scenario:vec_map/get_mut/128

  • 🟩 execution_time [-1.628µs; -1.535µs] or [-10.411%; -9.817%]
  • 🟩 throughput [+893912.906op/s; +948928.641op/s] or [+10.920%; +11.592%]

scenario:vec_map/get_mut/64

  • 🟩 execution_time [-383.621ns; -344.708ns] or [-9.049%; -8.131%]
  • 🟩 throughput [+1344799.949op/s; +1500976.073op/s] or [+8.904%; +9.938%]

scenario:vec_map/iter/128

  • 🟥 execution_time [+7.062ns; +7.132ns] or [+6.795%; +6.862%]
  • 🟥 throughput [-79083675.159op/s; -78342360.720op/s] or [-6.422%; -6.361%]

scenario:vec_map/iter/16

  • 🟥 execution_time [+0.578ns; +0.590ns] or [+4.420%; +4.509%]
  • 🟥 throughput [-52805705.985op/s; -51730408.902op/s] or [-4.318%; -4.230%]

scenario:vec_map/iter/8

  • 🟥 execution_time [+0.567ns; +0.572ns] or [+8.500%; +8.585%]
  • 🟥 throughput [-94942017.133op/s; -93946561.510op/s] or [-7.912%; -7.829%]

Benchmark execution time: 2026-07-16 14:43:56

Comparing candidate commit 0bdafcf in PR branch igor/versioning/green_ci with baseline commit 4a914bb in branch main.

Found 13 performance improvements and 6 performance regressions! Performance is the same for 104 metrics, 10 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:ddsketch_add/add_with_count/clustered_near_zero

  • 🟥 execution_time [+5.937µs; +6.469µs] or [+4.967%; +5.412%]
  • 🟥 throughput [-1753752.704op/s; -1614959.274op/s] or [-5.118%; -4.713%]

scenario:glob_matcher/ascii_case_insensitive_match/wall_time

  • 🟩 execution_time [-2.661ns; -2.581ns] or [-9.087%; -8.814%]

scenario:glob_matcher/ascii_exact_match/wall_time

  • 🟩 execution_time [-2.344ns; -2.263ns] or [-7.999%; -7.723%]

scenario:glob_matcher/ascii_exact_miss/wall_time

  • 🟩 execution_time [-3.286ns; -3.250ns] or [-21.978%; -21.737%]

scenario:normalization/normalize_name/normalize_name/Too-Long-.Too-Long-.Too-Long-.Too-Long-.Too-Long-.Too-Lo...

  • 🟩 execution_time [-18.837µs; -18.679µs] or [-9.184%; -9.107%]
  • 🟩 throughput [+488718.451op/s; +492805.896op/s] or [+10.024%; +10.108%]

scenario:normalization/normalize_name/normalize_name/bad-name

  • 🟩 execution_time [-1.257µs; -1.217µs] or [-6.725%; -6.512%]
  • 🟩 throughput [+3730463.147op/s; +3855809.230op/s] or [+6.973%; +7.207%]

scenario:normalization/normalize_name/normalize_name/good

  • 🟩 execution_time [-1.043µs; -1.010µs] or [-9.733%; -9.428%]
  • 🟩 throughput [+9732443.039op/s; +10025268.014op/s] or [+10.430%; +10.743%]

scenario:normalization/normalize_service/normalize_service/A0000000000000000000000000000000000000000000000000...

  • 🟩 execution_time [-37.231µs; -36.946µs] or [-6.940%; -6.887%]
  • 🟩 throughput [+137893.841op/s; +138971.811op/s] or [+7.398%; +7.456%]

scenario:normalization/normalize_service/normalize_service/Test Conversion 0f Weird !@#$%^&**() Characters

  • 🟩 execution_time [-25.187µs; -25.082µs] or [-12.870%; -12.816%]
  • 🟩 throughput [+751436.287op/s; +754550.854op/s] or [+14.705%; +14.766%]

scenario:normalization/normalize_service/normalize_service/test_ASCII

  • 🟥 execution_time [+1.983µs; +2.022µs] or [+4.274%; +4.358%]
  • 🟥 throughput [-900377.357op/s; -883043.082op/s] or [-4.177%; -4.097%]

scenario:trace_buffer/4_senders/no_delay

  • 🟥 execution_time [+112.881µs; +141.499µs] or [+4.850%; +6.079%]
  • 🟥 throughput [-91261.760op/s; -72450.437op/s] or [-5.892%; -4.677%]

Candidate

Omitted due to size.

Baseline

Omitted due to size.

@iunanua

iunanua commented Jul 16, 2026

Copy link
Copy Markdown
Contributor Author

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Jul 16, 2026

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-07-16 15:28:45 UTC ℹ️ Start processing command /merge


2026-07-16 15:28:48 UTC ❌ MergeQueue

PR already in the queue with status waiting

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants