Skip to content

[New Integration] GDACS - #19290

Merged
efd6 merged 20 commits into
elastic:mainfrom
tehbooom:gdacs
Aug 9, 2026
Merged

[New Integration] GDACS#19290
efd6 merged 20 commits into
elastic:mainfrom
tehbooom:gdacs

Conversation

@tehbooom

@tehbooom tehbooom commented May 29, 2026

Copy link
Copy Markdown
Contributor

Proposed commit message

This PR adds the Global Disaster Awareness and Coordination System (GDACS) integration to collect natural disaster events events.

Checklist

  • I have reviewed tips for building integrations and this pull request is aligned with them.
  • I have verified that all data streams collect metrics or logs.
  • I have added an entry to my package's changelog.yml file.
  • I have verified that Kibana version constraints are current according to guidelines.
  • I have verified that any added dashboard complies with Kibana's Dashboard good practices

Author's Checklist

  • [ ]

How to test this PR locally

Simply install the integration and run against the real API. No keys or additional information is needed.

Related issues

Screenshots

gdacs_events

Additional Info

Ive left codeowners empty as unsure which team would own this integration.

@tehbooom
tehbooom requested a review from a team as a code owner May 29, 2026 13:01
@github-actions

github-actions Bot commented May 29, 2026

Copy link
Copy Markdown
Contributor

Elastic Docs Style Checker (Vale)

Summary: 7 warnings, 2 suggestions found

⚠️ Warnings (7): Fix when the suggestion improves clarity or correctness.
File Line Rule Message
packages/gdacs/data_stream/events/fields/fields.yml 47 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'.
packages/gdacs/data_stream/events/fields/fields.yml 58 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'.
packages/gdacs/data_stream/events/fields/fields.yml 85 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'.
packages/gdacs/data_stream/events/fields/fields.yml 108 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'and so on' instead of 'etc'.
packages/gdacs/data_stream/events/fields/fields.yml 152 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'.
packages/gdacs/data_stream/events/fields/fields.yml 152 Elastic.QuotesPunctuation Place punctuation inside closing quotation marks.
packages/gdacs/data_stream/events/manifest.yml 46 Elastic.Latinisms Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'.
💡 Suggestions (2): Optional style improvements. Apply when helpful.
File Line Rule Message
packages/gdacs/changelog.yml 1 Elastic.Versions Use 'later versions' instead of 'newer versions' when referring to versions.
packages/gdacs/data_stream/events/manifest.yml 60 Elastic.WordChoice Consider using 'deactivated, deselected, hidden, turned off, unavailable' instead of 'disabled', unless the term is in the UI.

The Vale linter checks documentation changes against the Elastic Docs style guide. To use Vale locally or report issues, refer to Elastic style guide for Vale.

@github-actions

This comment has been minimized.

@github-actions

This comment has been minimized.

@github-actions

Copy link
Copy Markdown
Contributor

TL;DR

Buildkite failed in Check go sources because packages/gdacs/manifest.yml declares owner.github: elastic/integrations, but that owner is not represented in .github/CODEOWNERS for this package. Update the manifest owner to match the CODEOWNERS team for gdacs.

Remediation

  • In packages/gdacs/manifest.yml (line 35), change owner.github from elastic/integrations to the team that owns the package in CODEOWNERS (elastic/integration-experience), or alternatively align CODEOWNERS and manifest to the same exact owner value.
  • Re-run .buildkite/scripts/check_sources.sh (or CI) to confirm github.com/elastic/integrations/dev/codeowners.Check passes.
Investigation details

Root Cause

dev/codeowners.Check validates that each package manifest owner exists in .github/CODEOWNERS.

  • packages/gdacs/manifest.yml:35 has github: elastic/integrations.
  • PR head .github/CODEOWNERS:290 has /packages/gdacs @elastic/integration-experience``.

Those values do not match, so the check fails.

Evidence

Error: error validating packages in directory 'packages': error checking manifest 'packages/gdacs': owner "elastic/integrations" defined in "packages/gdacs/manifest.yml" is not in ".github/CODEOWNERS"

Verification

  • Not run locally in this environment (read-only detective workflow).

Follow-up

If you intended the package owner to be elastic/integrations, add a matching CODEOWNERS owner entry used by this repository’s ownership checks; otherwise prefer updating the manifest owner to the existing gdacs CODEOWNERS team.

Note

🔒 Integrity filter blocked 3 items

The following items were blocked because they don't meet the GitHub integrity level.

  • [New Integration] GDACS #19290 pull_request_read: has lower integrity than agent requires. The agent cannot read data with integrity below "approved".
  • #19290 pull_request_read: has lower integrity than agent requires. The agent cannot read data with integrity below "approved".
  • [New Integration] GDACS #19290 issue_read: has lower integrity than agent requires. The agent cannot read data with integrity below "approved".

To allow these resources, lower min-integrity in your GitHub frontmatter:

tools:
  github:
    min-integrity: approved  # merged | approved | unapproved | none

What is this? | From workflow: PR Buildkite Detective

Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.

@elastic-vault-github-plugin-prod

Copy link
Copy Markdown
Contributor

🚀 Benchmarks report

To see the full report comment with /test benchmark fullreport

@elasticmachine

Copy link
Copy Markdown

💚 Build Succeeded

History

@tehbooom tehbooom added enhancement New feature or request New Integration Issue or pull request for creating a new integration package. labels May 29, 2026
@andrewkroh andrewkroh added the documentation Improvements or additions to documentation. Applied to PRs that modify *.md files. label May 29, 2026
@botelastic

botelastic Bot commented Jun 28, 2026

Copy link
Copy Markdown

Hi! We just realized that we haven't looked into this PR in a while. We're sorry! We're labeling this issue as Stale to make it hit our filters and make sure we get back to it as soon as possible. In the meantime, it'd be extremely helpful if you could take a look at it as well and confirm its relevance. A simple comment with a nice emoji will be enough :+1. Thank you for your contribution!

@botelastic botelastic Bot added the Stalled label Jun 28, 2026
@jamiehynds jamiehynds added the Team:Security-Service Integrations Security Service Integrations team [elastic/security-service-integrations] label Jul 24, 2026
@botelastic botelastic Bot removed the Stalled label Jul 24, 2026
@infra-vault-gh-plugin-prod

Copy link
Copy Markdown

Pinging @elastic/security-service-integrations (Team:Security-Service Integrations)

@ShourieG

Copy link
Copy Markdown
Contributor

@vera-review-bot review

- remove:
tag: remove_event_original
field: event.original
if: "!ctx.tags.contains('preserve_original_event')"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:23

The remove_event_original condition calls .contains() on ctx.tags without a null guard, so every document fails the pipeline when the user clears the Tags variable. Add a null check to the condition.

Details

The condition is !ctx.tags.contains('preserve_original_event'). The tags variable in the data stream manifest has a default of [forwarded, gdacs-events], but it is not marked required, so a user can clear it in Fleet. When tags is absent the CEL template emits no tags: key, ctx.tags is null, and null.contains(...) throws a NullPointerException. The pipeline-level on_failure then catches it, so every single document is indexed as event.kind: pipeline_error instead of a parsed GDACS event. The pipeline test fixture cannot catch this because _dev/test/pipeline/test-common-config.yml always injects tags. Separately, note that the pipeline review checklist treats a remove of event.original keyed on the absence of the preserve_original_event tag as a deprecated pattern; if you keep it, it must at minimum be null-safe.

Recommendation:

Guard the condition against a missing tags array:

  - remove:
      tag: remove_event_original
      field: event.original
      if: 'ctx.tags == null || !(ctx.tags.contains("preserve_original_event"))'
      ignore_missing: true

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

value: gdacs
- set:
tag: set_event_category
field: event.category

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:437

event.category and event.type are ECS array fields but are written with the set processor, which overwrites anything already present. Use append with allow_duplicates: false instead.

Details

set_event_category (lines 435-439) and set_event_type (lines 440-444) both use set with a list value. The pipeline review checklist requires append for event.category and event.type because they are ECS array fields: set replaces the whole array, so any value contributed by a routing pipeline, a Logstash hop, or a future branch in this pipeline is silently discarded. append with allow_duplicates: false is idempotent and additive.

Recommendation:

Replace both set processors with append:

  - append:
      tag: append_event_category
      field: event.category
      value: threat
      allow_duplicates: false
  - append:
      tag: append_event_type
      field: event.type
      value: indicator
      allow_duplicates: false

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- gdacs.geometry_id
- gdacs.geometry_role
ignore_missing: true
on_failure:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: medium path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:506

The pipeline-level on_failure block deviates from the required structure: the steps are in the wrong order, the preserve_original_event tag append is missing, and the mustache placeholders use double braces which HTML-escape the error text. Rewrite it to the standard three-step form.

Details

Three deviations from the standard on_failure block for a new package. (1) Order: set event.kind runs before append error.message; the required order is error.message, then event.kind, then the tags append. (2) The third step is missing entirely: append tags: preserve_original_event is what keeps event.original on the document when the pipeline fails, so a failing document currently loses its raw payload to the remove_event_original processor and becomes undebuggable. (3) The template uses double-brace mustache ({{ _ingest.on_failure_message }}); double braces HTML-escape the substituted value, so an error message containing quotes or angle brackets is mangled. Triple braces are required. Note _ingest.on_failure_pipeline is also not part of the standard template.

Recommendation:

Replace the on_failure block with the standard three-step form:

on_failure:
  - append:
      field: error.message
      value: >-
        Processor '{{{ _ingest.on_failure_processor_type }}}'
        {{{#_ingest.on_failure_processor_tag}}}with tag '{{{ _ingest.on_failure_processor_tag }}}'
        {{{/_ingest.on_failure_processor_tag}}}failed with message '{{{ _ingest.on_failure_message }}}'
  - set:
      field: event.kind
      tag: set_pipeline_error_to_event_kind
      value: pipeline_error
  - append:
      field: tags
      value: preserve_original_event
      allow_duplicates: false

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- name: data_stream.namespace
type: constant_keyword
description: Data stream namespace.
- name: input.type

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/fields/base-fields.yml:10

base-fields.yml does not match the required six-entry shape: event.module and event.dataset are missing, input.type is not part of it, and no entry uses external: ecs. Replace the file with the standard six entries.

Details

The required base-fields.yml contains exactly six entries — data_stream.type, data_stream.dataset, data_stream.namespace, event.module, event.dataset, @​timestamp — all declared with external: ecs, with event.module and event.dataset overridden to constant_keyword carrying their literal values. This file instead has five entries, declares input.type (which belongs in beats.yml, and is not needed for a CEL stream at all), inlines descriptions rather than inheriting them from ECS, and omits event.module/event.dataset. Those two are currently declared as plain external: ecs keywords in fields/ecs.yml, so they are mapped as regular keywords and indexed per document instead of being folded into a zero-cost constant_keyword. Remove the duplicate event.dataset and event.module entries from fields/ecs.yml when you make this change.

Recommendation:

Replace the contents of base-fields.yml:

- name: data_stream.type
  external: ecs
- name: data_stream.dataset
  external: ecs
- name: data_stream.namespace
  external: ecs
- name: '@​timestamp'
  external: ecs
- name: event.module
  type: constant_keyword
  external: ecs
  value: gdacs
- name: event.dataset
  type: constant_keyword
  external: ecs
  value: gdacs.events

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- name: geo.country_name
type: keyword
description: Name of the primary affected country.
- name: geo.location

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: medium path: packages/gdacs/data_stream/events/fields/fields.yml:10

The geo_point and the rest of the geo fieldset are declared at the document root (geo.location, geo.name, geo.country_name) instead of nested under a parent entity. Move them under the gdacs namespace.

Details

The field review checklist requires that geo fields, and geo_point fields in particular, are never declared at root level — the ECS geo fieldset is reusable and is expected under a parent entity (source.geo, destination.geo, host.geo, observer.geo, ...), never as a bare top-level geo.*. This package declares geo.location (geo_point), geo.country_iso_code, geo.country_name and geo.name at the root, and the pipeline writes them there (the extract_geometry and flatten_affected_countries scripts build ctx.geo, and set_geo_name copies into geo.name). Since a GDACS disaster event has no ECS entity that owns the location, the natural home is the package namespace alongside the other custom fields. This touches the pipeline scripts, fields.yml, the map layers in the dashboard, and the README validation steps, so it is best done now rather than after 1.0.0.

Recommendation:

Nest the fieldset under gdacs and update the pipeline scripts and the dashboard geoField references to match:

- name: gdacs.geo.location
  type: geo_point
  description: Centroid coordinates of the event or affected area.
- name: gdacs.geo.country_iso_code
  type: keyword
  description: ISO 3166-1 alpha-2 country code of the primary affected country.
- name: gdacs.geo.country_name
  type: keyword
  description: Name of the primary affected country.
- name: gdacs.geo.name
  type: keyword
  description: Human-readable name of the event location.

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- name: geo.location
type: geo_point
description: Centroid coordinates of the event or affected area.
- name: geo.location.coordinates

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: medium path: packages/gdacs/data_stream/events/fields/fields.yml:13

Four subfields are declared underneath the geo_point and geo_shape leaves (geo.location.coordinates/.type and gdacs.affected_area.coordinates/.type) but the pipeline never writes them. Delete the four declarations.

Details

geo.location is declared geo_point (line 10) and gdacs.affected_area is declared geo_shape (line 142), yet both also get .coordinates (object/double) and .type (keyword) children declared beneath them at lines 13-19 and 146-152. The pipeline writes geo.location as a {lon, lat} map (extract_geometry, line 307) and gdacs.affected_area as a WKT string (line 317), which the pipeline test fixture confirms — test-events-ndjson.log-expected.json line 29 shows "affected_area": "POLYGON ((...))" and lines 82-85 show "location": {"lat": ..., "lon": ...}. The GeoJSON {type, coordinates} shape seen in sample_event.json is Elasticsearch reconstructing the geo fields from synthetic source at read time, not a document shape the pipeline produces. Declaring them as real mapped subfields of a leaf field type is an orphan declaration: nothing writes them, and they surface as phantom entries in the generated field reference in the README.

Recommendation:

Drop the four child declarations and keep only the geo leaves:

- name: geo.location
  type: geo_point
  description: Centroid coordinates of the event or affected area.

and under the gdacs group:

    - name: affected_area
      type: geo_shape
      description: >-
        WKT polygon, multipolygon or line representing the affected area of the disaster.

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Elasticsearch exposes those subfields via synthetic source at read time, so sample_event.json contains them and they need to be declared

- set:
tag: set_event_kind
field: event.kind
value: alert

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: low path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:426

The ECS categorization models natural disaster alerts as threat intelligence (event.category threat, event.type indicator) which also conflicts with event.kind alert. Reconsider the triple, most likely event.kind alert with no threat categorization.

Details

The pipeline sets event.kind: alert, event.category: [threat] and event.type: [indicator]. Both values are in the ECS allowed list, so this is a semantic rather than a syntactic issue, but the combination is inconsistent: in ECS, event.type: indicator denotes a threat-intelligence indicator record and pairs with event.kind: enrichment, not with event.kind: alert. GDACS earthquakes, floods and cyclones are not adversarial threat indicators, and mapping them into the threat category will make them show up in threat-intel-scoped queries and dashboards alongside real IoCs. event.kind: alert on its own is a defensible fit for a disaster alert feed; there is no ECS event.category value that describes a natural hazard, and leaving the category unset is preferable to a misleading one.

Recommendation:

Keep event.kind and drop the threat-intel categorization:

  - set:
      tag: set_event_kind
      field: event.kind
      value: alert

If you want a categorization value for filtering, prefer one that matches the semantics of the feed rather than threat, and document the choice in the README.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

[{"message": f.encode_json()}]
).flatten(),
"cursor": {
"page_number": int(state.?cursor.page_number.orValue(1)) + 1,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: medium path: packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs:119

When the last page of a run is partial the cursor keeps the incremented page_number and the frozen date window, so the next poll spends a whole interval fetching an empty page. Reset the cursor when want_more is false.

Details

The success branch always writes page_number: <current> + 1 together with the pinned from_date/to_date, regardless of whether want_more is true. The cursor is only cleared in the empty-result branch (lines 128-134). So a normal run ends like this: the final page returns fewer than page_size features, want_more becomes false, and the cursor is left pointing at page N+1 of an exhausted, frozen date window. The next scheduled poll therefore requests that empty page, gets zero features, and only then resets the cursor — meaning fresh events are picked up on every second poll and the effective collection interval is double the configured one. With the default interval: 1h that is an hour of avoidable latency on a disaster alert feed.

Recommendation:

Reset paging state on the terminal page instead of carrying it forward:

        "cursor": size(body.features) >= state.page_size ? {
          "page_number": int(state.?cursor.page_number.orValue(1)) + 1,
          "from_date": state.?cursor.from_date.orValue(
            (now - duration(string(state.lookback_hours) + "h")).format("01/02/2006")
          ),
          "to_date": state.?cursor.to_date.orValue(now.format("01/02/2006")),
        } : {
          "last_poll_date": now.format("01/02/2006"),
        },
        "want_more": size(body.features) >= state.page_size,

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

copy_from: gdacs.name
ignore_empty_value: true
- remove:
tag: remove_raw_fields

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:489

The CEL program computes geometry_doc, geometry_id and geometry_role for tropical cyclone child documents, the pipeline renames them into gdacs.*, and then remove_raw_fields deletes them all. Keep gdacs.geometry_role or stop producing it.

Details

remove_raw_fields deletes gdacs.geometry_type, gdacs.geometry_doc, gdacs.geometry_id and gdacs.geometry_role after three dedicated rename processors (lines 125-139) moved them into the gdacs namespace. The expected-output fixture confirms none of them survive. The only surviving consumer is event.id, which folds geometry_id into the fingerprint. geometry_role in particular classifies each TC child document as wind, cone or track — exactly the distinction the dashboard's TC map layers currently have to reconstruct with gdacs.class: Line_Line_* OR gdacs.class: Poly_Cones kuery filters. Either declare and keep it, or drop the renames and the CEL fields that feed them.

Recommendation:

Keep the role field, declare it in fields.yml, and narrow the cleanup list:

  - remove:
      tag: remove_raw_fields
      field:
        - properties
        - geometry
        - bbox
        - type
        - polygon_geometry
        - polygon_class
        - polygon_label
        - gdacs.geometry_type
        - gdacs.geometry_doc
        - gdacs.geometry_id
      ignore_missing: true

with the matching declaration:

    - name: geometry_role
      type: keyword
      description: Role of a tropical cyclone child geometry document - wind, cone or track.

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

Comment thread packages/gdacs/manifest.yml Outdated
type: image/png
icons:
- src: /img/gdacs-logo.svg
title: Sample logo

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/manifest.yml:23

The package icon still carries the scaffold placeholder title "Sample logo" and the screenshot title has a doubled space. Give both descriptive titles.

Details

icons[0].title is Sample logo, which is the value the scaffolder emits and is meant to be replaced. It is user-visible in the Fleet integration listing. Line 18 also has title: Dashboard Overview with two consecutive spaces between the words.

Recommendation:

Give both assets real titles:

screenshots:
  - src: /img/gdacs_events.png
    title: GDACS Events dashboard overview
    size: 600x600
    type: image/png
icons:
  - src: /img/gdacs-logo.svg
    title: GDACS logo
    size: 32x32
    type: image/svg+xml

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

Comment thread .github/CODEOWNERS Outdated
Comment thread .github/CODEOWNERS Outdated
Comment thread packages/gdacs/_dev/build/build.yml Outdated
Comment thread packages/gdacs/data_stream/events/fields/ecs.yml
Comment thread packages/gdacs/data_stream/events/fields/fields.yml Outdated
Comment thread packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml Outdated
Comment thread packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs
"events": [],
"want_more": false,
"cursor": {
"last_poll_date": now.format("01/02/2006"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The cursor.last_poll_date is not used anywhere inside the request. It looks like we want to set next request's fromDate based on this value?

Comment thread packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs
target_field: event.original
ignore_missing: true
if: ctx.event?.original == null
- json:

@kcreddy kcreddy Jul 28, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The processor remove_agentless_tags isn't applicable anymore. This package doesn't have agentless.

The rest of the review bot's suggestion can be done.

if (polyType == "Polygon" || polyType == "MultiPolygon" || polyType == "LineString" || polyType == "MultiLineString") {
if (ctx.gdacs == null) { ctx.gdacs = new HashMap(); }
ctx.gdacs.affected_area = gdacsShapeToWkt(polyGeom);
ctx.gdacs.geometry_type = polyType;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:318

The extract_geometry script computes gdacs.geometry_type and remove_raw_fields deletes it a few processors later, so the value never reaches the index. Either keep the field (drop it from remove_raw_fields and declare it in fields.yml) or stop computing it.

Details

Line 318 sets ctx.gdacs.geometry_type = polyType alongside ctx.gdacs.affected_area, but remove_raw_fields (line 501) lists gdacs.geometry_type among the fields to delete. The sibling affected_area survives while geometry_type is always dropped, so the assignment is dead code and the field is absent from both fields.yml and the pipeline-test expected output.

This matters beyond tidiness: geometry_type is the only value that distinguishes a Polygon/MultiPolygon wind band from a LineString cyclone track once affected_area has been flattened to WKT. Because it is discarded, the dashboard map layer has to reconstruct that distinction from the label string with a wildcard query ("gdacs.event_type: TC AND (gdacs.class: Line_Line_* OR gdacs.class: Poly_Cones)"), which breaks as soon as GDACS renames a Class value.

Recommendation:

Keep the field: remove the gdacs.geometry_type entry from remove_raw_fields and declare it.

  - remove:
      tag: remove_raw_fields
      field:
        - properties
        - geometry
        - bbox
        - type
        - polygon_geometry
        - polygon_class
        - polygon_label
        - geometry_doc
        - geometry_id
        - geometry_role
        - gdacs.geometry_doc
        - gdacs.geometry_id
        - gdacs.geometry_role
      ignore_missing: true

and in fields/fields.yml, under the gdacs group:

    - name: geometry_type
      type: keyword
      description: >-
        GeoJSON geometry type of gdacs.affected_area
        (Polygon, MultiPolygon, LineString, MultiLineString).

Otherwise, if the field is genuinely not wanted, delete the assignment instead so the pipeline does not compute a value it discards:

            ctx.gdacs.affected_area = gdacsShapeToWkt(polyGeom);

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- set:
tag: set_message
field: message
value: "{{{gdacs.description}}}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: medium path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:481

set_message builds message from a mustache template instead of copy_from, so a document without gdacs.description gets message set to an empty string. Switch to copy_from with ignore_empty_value.

Details

The set_message processor writes message from the mustache template "{{{gdacs.description}}}". When gdacs.description is absent the template renders to the empty string and the set processor stores message: "" — ignore_failure does not prevent this, because nothing failed. The result is an indexed empty message field rather than no field at all.

Every other value-copying set processor in this pipeline (set_event_url, set_event_reference, set_geo_name, set_timestamp_modified, set_timestamp_start) already uses copy_from together with ignore_empty_value, which skips the write when the source is missing or empty. set_message is the only one that does not follow that pattern.

Recommendation:

Use copy_from so the processor is skipped when the source field is absent:

  - set:
      tag: set_message
      field: message
      copy_from: gdacs.description
      ignore_empty_value: true

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- name: class
type: keyword
description: >-
GeoJSON feature class: Point_Centroid (epicenter) or Poly_area

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/data_stream/events/fields/fields.yml:92

The descriptions for gdacs.class and gdacs.polygon_label list values the integration never produces ("Poly_area", "Affected Area") and omit the ones it does. Update both descriptions to the actual value sets.

Details

gdacs.class is documented as "Point_Centroid (epicenter) or Poly_area (affected area polygon)". Poly_area is never emitted. The extract_geometry script overwrites gdacs.class with the enrichment's polygon_class, and the values actually produced — visible in the committed test fixture and expected output — are Point_Centroid, Poly_Circle, Poly_Green, Poly_Orange, Poly_Red, Poly_Cones and Line_Line_. The dashboard queries three of these (Line_Line_*, Poly_Cones) that the description does not mention.

gdacs.polygon_label has the same problem: it is documented as "(Centroid, Affected Area)", but the fixtures carry "Centroid", radius labels such as "100km", and tropical-cyclone forecast timestamps such as "01/06 06:00" — which the dashboard also uses as a categorical colour stop.

These descriptions are rendered into the published field reference by the README's {{ fields "events" }} directive, so the inaccuracy reaches users.

Recommendation:

Describe the values the pipeline actually emits:

    - name: class
      type: keyword
      description: >-
        GDACS GeoJSON feature class. Point_Centroid for the event epicentre, or
        the enrichment polygon class: Poly_Circle, Poly_Green, Poly_Orange,
        Poly_Red, Poly_Cones, or Line_Line_<n> for tropical cyclone track lines.
    - name: polygon_label
      type: keyword
      description: >-
        Label for the geometry, e.g. "Centroid", an impact radius such as
        "100km", or a tropical cyclone forecast timestamp such as "01/06 06:00".

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- set:
tag: set_ecs_version
field: ecs.version
value: "8.11.0"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:453

The pipeline pins ecs.version to 8.11.0 but _dev/build/build.yml pins the ECS reference to v9.4.0; set the pipeline value to 9.4.0 so the two agree.

Details

_dev/build/build.yml declares dependencies.ecs.reference: "git@​v9.4.0", so the package's field definitions are generated from ECS 9.4.0. The pipeline, however, stamps every document with ecs.version: "8.11.0" (this value is also baked into the committed pipeline test expectations). The two must agree: consumers and downstream rules read ecs.version to decide which ECS schema a document conforms to, and here it advertises a schema three minor versions older than the one the mappings were actually built from. This is not a "could be newer" nit — it is a concrete mismatch between two files in this PR.

Recommendation:

Set the pipeline ecs.version to match the build.yml ECS pin:

  - set:
      tag: set_ecs_version
      field: ecs.version
      value: "9.4.0"

Then regenerate the pipeline test expectations (elastic-package test pipeline --generate) so test-events-ndjson.log-expected.json and sample_event.json carry the updated value.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

if: ctx.event?.original != null
ignore_missing: true
- remove:
tag: remove_event_original

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:26

The pipeline deletes event.original based on the absence of the preserve_original_event tag, which is a prohibited pattern; remove this processor entirely and let the agent control preservation.

Details

Deleting event.original from inside the ingest pipeline based on ctx.tags not containing preserve_original_event is a deprecated pattern that current integrations must not use. Preservation is already controlled at the agent level by the preserve_original_event variable in the data stream manifest, which adds the tag to the CEL input's tags list; the pipeline should only rename message to event.original and drop message.

It also defeats this pipeline's own error handling. The pipeline-level on_failure block ends with append: tags / value: preserve_original_event (line 563), whose entire purpose is to retain the raw document for debugging when a processor fails. Because this remove runs as processor #​5, event.original is already gone by the time any later processor fails, so failed documents land with the preserve_original_event tag set but no original payload to inspect. With the default preserve_original_event: false, that is every failure in this pipeline.

For reference, packages/neon_cyber/data_stream/events/elasticsearch/ingest_pipeline/default.yml shows the current shape: rename, remove message, and no removal of event.original anywhere.

Recommendation:

Delete the processor entirely, keeping only the rename and the message cleanup:

  - rename:
      tag: rename_message_to_event_original
      field: message
      target_field: event.original
      ignore_missing: true
      if: ctx.event?.original == null
  - json:
      tag: json_parse_event_original
      field: event.original
      add_to_root: true
      add_to_root_conflict_strategy: replace
      if: ctx.event?.original != null
  - remove:
      tag: remove_message
      field: message
      if: ctx.event?.original != null
      ignore_missing: true

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

pf.properties.Class != "Point_Centroid" &&
has(pf.geometry) && has(pf.geometry.type) &&
(pf.geometry.type == "Polygon" || pf.geometry.type == "MultiPolygon")
).as(polys, size(polys) > 0 ?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs:110

The CEL program nests .as() bindings six levels deep, exceeding the hard cap of five; hoist the cursor/date pre-bindings out of the HTTP chain to flatten it.

Details

The program stacks six nested .as() bindings inside a single expression: countryParams (line 27) -> resp (line 40) -> body (line 67) -> polyResp (line 75) -> polyBody (line 76) -> polys (line 110). The documented hard cap is five levels measured from state.with() inward, with the HTTP core expected to be two levels (do_request().as(resp, decode_json().as(body, {...}))).

A contributing cause is that window bounds and cursor defaults are computed inline inside the HTTP chain instead of being bound up front, and they are duplicated verbatim: the from_date default expression (now - duration(string(state.lookback_hours) + "h")).format("01/02/2006") appears at lines 35-37 and again at lines 131-133, and now.format("01/02/2006") is repeated at lines 38, 45, 47, 134, 136 and 147. Hoisting those into pre-bindings before the request both removes the duplication and drops nesting levels.

Recommendation:

Bind the window bounds and the country params before the request, so the HTTP chain stays shallow:

  state.?cursor.from_date.orValue(
    (now - duration(string(state.lookback_hours) + "h")).format("01/02/2006")
  ).as(fromDate,
  state.?cursor.to_date.orValue(now.format("01/02/2006")).as(toDate,
  now.format("01/02/2006").as(today,
    state.with(
      request("GET", state.url.trim_right("/") + "/events/geteventlist/SEARCH?" + {
        "eventlist": [state.event_types],
        "alertlevel": [state.alert_levels],
        "pageSize": [string(state.page_size)],
        "pageNumber": [string(state.?cursor.page_number.orValue(1))],
        "fromDate": [fromDate],
        "toDate": [toDate],
      }.with(
        state.?country.orValue("") != "" ? {"country": [state.country]} : {}
      ).format_query()).do_request().as(resp,
        resp.Body.decode_json().as(body, {
          // ...
        })
      )
    )
  )))

Extracting the per-feature geometry enrichment into its own top-level binding (rather than nesting polyResp/polyBody/polys inside the map()) will bring the remaining depth under the cap.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

type: geo_shape
description: >-
WKT polygon, multipolygon or line representing the affected area of the disaster.
- name: affected_area.coordinates

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: medium path: packages/gdacs/data_stream/events/fields/fields.yml:132

fields.yml declares .coordinates and .type sub-paths beneath the geo_shape and geo_point leaves gdacs.affected_area and gdacs.geo.location; delete these four entries as the pipeline never produces that shape.

Details

gdacs.affected_area is declared geo_shape (line 128) and gdacs.geo.location is declared geo_point (line 139), but the file then declares sub-paths beneath both leaves: affected_area.coordinates (132), affected_area.type (136), geo.location.coordinates (142) and geo.location.type (146). A field cannot be both a concrete geo_shape/geo_point leaf and an object with properties — Elasticsearch rejects properties on those types, so this risks failing index-template installation for the data stream.

Separately, these are orphan declarations regardless of the mapping question: the pipeline never writes that shape. The extract_geometry script sets ctx.gdacs.affected_area to a WKT string ("POLYGON ((...))", default.yml line 332) and ctx.gdacs.geo.location to a ['lon': ..., 'lat': ...] map (line 322). The committed expectations confirm it — test-events-ndjson.log-expected.json shows "location": {"lat": -10.5397, "lon": 162.4478} and affected_area as a WKT string, with no coordinates/type sub-keys anywhere.

This pattern appears in no other package: a grep for *.coordinates/*.type sub-declarations under geo fields across all packages/*/data_stream/*/fields/*.yml matches only this file.

Recommendation:

Drop the four sub-path declarations and keep only the two leaves:

    - name: affected_area
      type: geo_shape
      description: >-
        WKT polygon, multipolygon or line representing the affected area of the disaster.
    - name: geo.location
      type: geo_point
      description: Centroid coordinates of the event or affected area.

Both types natively accept the values this pipeline produces (WKT strings for geo_shape, lat/lon maps for geo_point), so no sub-field declarations are needed. Run elastic-package build && elastic-package check afterwards to confirm the index template installs.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

description: Polls the GDACS Search API for natural disaster alerts and events.
owner:
github: elastic/security-service-integrations
type: community

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: medium path: packages/gdacs/manifest.yml:36

owner.type is set to community while CODEOWNERS assigns the package to the Elastic team elastic/security-service-integrations; set it to elastic to match.

Details

The manifest declares owner.github: elastic/security-service-integrations with owner.type: community. Per the package spec, community means "built and maintained by non-Elastic community members", while elastic means built and maintained by Elastic. This PR's CODEOWNERS change assigns /packages/gdacs to @​elastic/security-service-integrations @​elastic/sit-crest-contractors — the same ownership line used by that team's other packages, all of which declare owner.type: elastic (for example packages/github/manifest.yml line 128, same GitHub team). owner.type drives the support-level badge Fleet shows users, so community here understates the support commitment implied by the CODEOWNERS entry.

Recommendation:

Align the declared owner type with the assigned CODEOWNERS team:

owner:
  github: elastic/security-service-integrations
  type: elastic

If the package is genuinely intended to ship as community-supported, leave type: community and drop the Elastic team from the CODEOWNERS entry instead, so the two files agree.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

description: "Ingest Global Disaster Alert and Coordination System (GDACS) events — earthquakes, tropical cyclones, floods, volcanoes, droughts, and wildfires."
type: integration
categories:
- custom

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: medium path: packages/gdacs/manifest.yml:10

categories is set to the generic custom bucket, which hides the package in Fleet's browser; pick descriptive categories such as security and observability.

Details

custom is a valid enum value in the package spec, so this will not fail validation, but it is the catch-all bucket for user-defined packages and no other package in the repository uses it (a grep for - custom across every packages/*/manifest.yml returns no matches). Fleet groups the integration browser by category, so a first-party package categorised as custom is effectively undiscoverable by anyone browsing for disaster/situational-awareness data. The package's own README pitches it for "situational awareness dashboards" and "correlation of disaster events with infrastructure or supply chain data", which maps onto existing categories.

Recommendation:

Choose categories that describe what the package actually collects:

categories:
  - security
  - observability

Any combination from the spec's enum works — threat_intel is another reasonable fit given the alert-feed shape — but custom should not be the only category on a package shipped in this repository.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.


1. In Kibana, go to **Discover** and select the `logs-gdacs.events-*` data stream.
2. You should see documents with `event.kind: alert` and `event.module: gdacs`.
3. Check for `gdacs.event_type`, `gdacs.alert_level`, and `geo.location` fields.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/_dev/build/docs/README.md:63

The README validation step tells users to look for a geo.location field, but the pipeline writes gdacs.geo.location; correct the field name so the verification step works.

Details

Step 3 of the "Validation" section instructs users to check for gdacs.event_type, gdacs.alert_level, and geo.location. There is no geo.location field in this package — the pipeline's extract_geometry script writes the centroid to ctx.gdacs.geo.location (default.yml line 322) and fields/fields.yml declares it as gdacs.geo.location (line 139). A user following this step verbatim will find nothing and conclude the integration is broken.

Line 22 of the same file is also inaccurate: it states events are deduplicated "using a fingerprint of {event_id}-{episode_id}", but the set_event_id script (default.yml lines 470-482) also appends gdacs.geometry_id for tropical-cyclone child geometry documents, so those fingerprints are three-part.

Recommendation:

Use the field name the pipeline actually writes:

3. Check for `gdacs.event_type`, `gdacs.alert_level`, and `gdacs.geo.location` fields.

And correct the deduplication description on line 22:

- Deduplicates events using a fingerprint of `{event_id}-{episode_id}`, extended with `{geometry_id}` for tropical cyclone child geometry documents

Remember to regenerate packages/gdacs/docs/README.md (elastic-package build) so the published copy picks up both edits.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

Comment thread packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs Outdated
Comment thread packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs Outdated
- set:
tag: set_ecs_version
field: ecs.version
value: "8.11.0"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml:453

The pipeline pins ecs.version to 8.11.0 while _dev/build/build.yml imports ECS v9.4.0; set ecs.version to 9.4.0 so the declared version matches the imported field definitions.

Details

packages/gdacs/_dev/build/build.yml pins the ECS dependency to git@​v9.4.0, but the set_ecs_version processor stamps every document with ecs.version: "8.11.0". The mappings generated for this data stream come from ECS 9.4.0, so the value written to ecs.version misreports the schema the documents were normalized against. sample_event.json and the pipeline test expectations both carry the incorrect 8.11.0 value as well, so they will need to be regenerated after the fix.

Recommendation:

Set the pipeline's ecs.version to the version pinned in build.yml:

  - set:
      tag: set_ecs_version
      field: ecs.version
      value: "9.4.0"

Then regenerate the pipeline test expectations and sample_event.json so they carry the same value.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

"url": "https://www.gdacs.org/report.aspx?eventid=1476137&episodeid=1632460&eventtype=EQ"
},
"gdacs": {
"affected_area": {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: high path: packages/gdacs/data_stream/events/sample_event.json:40

sample_event.json does not match what the pipeline actually produces (GeoJSON objects and scalars instead of WKT strings and arrays); regenerate it from a system test run.

Details

This sample event was not produced by the current pipeline. The extract_geometry script sets gdacs.affected_area to a WKT string (gdacsShapeToWkt returns "POLYGON ((...))") and gdacs.geo.location to ['lon': ..., 'lat': ...], and flatten_affected_countries builds affected_country_names/_iso2/_iso3 as ArrayLists. The generated pipeline expectations in _dev/test/pipeline/test-events-ndjson.log-expected.json confirm this: they contain "affected_area": "POLYGON ((...", "location": {"lat": -10.5397, "lon": 162.4478}, "affected_countries": [ { ... } ] and "affected_country_names": ["Solomon Is."].

sample_event.json instead shows gdacs.affected_area as a GeoJSON object with coordinates/type, gdacs.geo.location as {"coordinates": [...], "type": "Point"}, affected_countries as a single object, and affected_country_names/_iso2/_iso3 as bare strings. Because docs/README.md embeds this file via {{ event "events" }}, the published documentation currently shows a document shape users will never see.

Recommendation:

Delete the hand-written file and regenerate it from a system test run so it reflects the real pipeline output, e.g.:

# packages/gdacs/data_stream/events/sample_event.json (regenerated)
{
    "gdacs": {
        "affected_area": "POLYGON ((163.364 -10.54, ...))",
        "affected_country_names": [
            "Solomon Is."
        ],
        "geo": {
            "location": {
                "lat": -10.5397,
                "lon": 162.4478
            }
        }
    }
}

Run elastic-package test system -g for the events data stream and commit the generated file, then rebuild the docs.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

type: geo_shape
description: >-
WKT polygon, multipolygon or line representing the affected area of the disaster.
- name: affected_area.coordinates

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟠 High confidence: medium path: packages/gdacs/data_stream/events/fields/fields.yml:132

fields.yml declares object subfields underneath the geo_shape and geo_point leaves (affected_area.coordinates/.type, geo.location.coordinates/.type); remove them, as they conflict with the leaf mapping and are never produced.

Details

gdacs.affected_area is declared as geo_shape on line 129 and gdacs.geo.location as geo_point on line 139, but lines 132-138 and 142-148 then declare affected_area.coordinates, affected_area.type, geo.location.coordinates and geo.location.type as separate object/keyword fields. A geo_shape/geo_point mapping is a leaf and cannot also carry properties, so these declarations describe a mapping that cannot coexist with the leaf type.

They are also unreachable: the extract_geometry script writes a WKT string to gdacs.affected_area and a {lon, lat} map to gdacs.geo.location, so no document ever contains affected_area.coordinates, affected_area.type, geo.location.coordinates or geo.location.type. The generated expectations in test-events-ndjson.log-expected.json confirm neither subfield is emitted. These four entries appear to be left over from the GeoJSON-object shape still shown in sample_event.json (see finding 2).

Recommendation:

Drop the four subfield declarations and keep only the leaf definitions:

    - name: affected_area
      type: geo_shape
      description: >-
        WKT polygon, multipolygon or line representing the affected area of the disaster.
    - name: geo.location
      type: geo_point
      description: Centroid coordinates of the event or affected area.

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

@@ -0,0 +1,157 @@
- name: event.modified

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: high path: packages/gdacs/data_stream/events/fields/fields.yml:1

event.modified is not an ECS field, so this adds a custom field to the ECS-reserved event.* namespace; move it under gdacs.*.

Details

event.modified is declared locally as a custom date field rather than with external: ecs. It is not part of ECS — the ECS v9.4.0 field list (which build.yml pins via git@​v9.4.0) defines event.created, event.start, event.end and event.ingested under event.*, but no event.modified. Declaring a package-specific field inside the ECS-owned event.* namespace risks colliding with a future ECS field of the same name and breaks the convention that non-ECS data lives under the package namespace. The date_datemodified processor in default.yml (line 219) targets this field.

Recommendation:

Move the field into the package namespace and retarget the date processor:

# fields/fields.yml
    - name: date_modified
      type: date
      description: Timestamp when the GDACS event was last modified.
# elasticsearch/ingest_pipeline/default.yml
  - date:
      tag: date_datemodified
      field: properties.datemodified
      target_field: gdacs.date_modified
      formats:
        - "yyyy-MM-dd'T'HH:mm:ss"
        - ISO8601

Update the set_timestamp_modified / set_timestamp_start processors to copy from gdacs.date_modified accordingly.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

Comment thread packages/gdacs/manifest.yml
(now - duration(string(state.lookback_hours) + "h")).format("01/02/2006")
),
],
"toDate": [state.?cursor.to_date.orValue(now.format("01/02/2006"))],

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🟡 Medium confidence: medium path: packages/gdacs/data_stream/events/agent/stream/cel.yml.hbs:34

The request's toDate is read from the stored cursor instead of the current time, so each poll queries a window that ends at the previous poll's date; compute toDate from now.

Details

Line 34 builds toDate as state.?cursor.to_date.orValue(now.format("01/02/2006")). On the terminal branches (lines 141-154 and 159-169) the program writes cursor.to_date = now, so from the second poll onward cursor.to_date is always populated and the orValue(now) fallback never applies. The query window therefore ends at the timestamp recorded by the previous poll rather than the present time.

Because the format is day-granularity (01/02/2006 is MM/DD/YYYY), the concrete effect is at each day rollover: the first poll after midnight still sends toDate = the previous day, so events from the new day are excluded until the following poll rewrites the cursor. The fromDate side is correct — it must come from the cursor to advance the window — but toDate should track the current time.

Recommendation:

Always end the window at the current time:

      "fromDate": [
        state.?cursor.from_date.orValue(
          (now - duration(string(state.lookback_hours) + "h")).format("01/02/2006")
        ),
      ],
      "toDate": [now.format("01/02/2006")],

Keep writing cursor.to_date on the terminal branches so the next fromDate picks up where this run ended; only the pagination branch (lines 141-148) needs to pin to_date for the duration of a paged run.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.


1. In Kibana, go to **Discover** and select the `logs-gdacs.events-*` data stream.
2. You should see documents with `event.kind: alert` and `event.module: gdacs`.
3. Check for `gdacs.event_type`, `gdacs.alert_level`, and `geo.location` fields.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/_dev/build/docs/README.md:63

The validation step tells users to look for a geo.location field, but the package writes gdacs.geo.location; correct the field name.

Details

Step 3 of the Validation section instructs users to check for gdacs.event_type, gdacs.alert_level, and geo.location. The pipeline never writes a top-level ECS geo.* object — extract_geometry sets ctx.gdacs.geo.location, and fields.yml declares the field as gdacs.geo.location. Following the documented step returns no field. This text is also carried into the generated packages/gdacs/docs/README.md.

Recommendation:

Use the field name the package actually produces:

3. Check for `gdacs.event_type`, `gdacs.alert_level`, and `gdacs.geo.location` fields.

Rebuild the docs with elastic-package build so docs/README.md picks up the change.


🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

- Extracts centroid `geo_point` and affected-area `geo_shape` from GeoJSON geometry
- Maps GDACS alert levels (Red/Orange/Green) to numeric `event.severity` and `event.risk_score`
- Flattens affected country arrays into searchable keyword fields
- Deduplicates events using a fingerprint of `{event_id}-{episode_id}`

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/_dev/build/docs/README.md:22

The How it works section says events are deduplicated on {event_id}-{episode_id}, but the fingerprint also includes the geometry id for tropical-cyclone child documents; update the description.

Details

The set_event_id script in default.yml (lines 470-482) joins gdacs.event_id, gdacs.episode_id and gdacs.geometry_id, and fingerprint_event_id hashes that value into _id. For tropical cyclone events the CEL program emits one child document per wind band, cone and track line, each with a distinct geometry_id; without it in the key those child documents would all collapse onto a single _id. The README's {event_id}-{episode_id} description omits this and understates how the child geometry documents are kept distinct.

Recommendation:

Describe the actual key:

- Deduplicates events using a fingerprint of `{event_id}-{episode_id}`, extended with the GDACS geometry id for tropical cyclone wind band, cone, and track child documents

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

},
"showApplySelections": false
},
"description": "",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: 🔵 Low confidence: high path: packages/gdacs/kibana/dashboard/gdacs-998e30fb-8dff-4724-8186-fd9410478f8e.json:104

The [GDACS] Events Overview dashboard ships with an empty description; add one describing what the dashboard shows.

Details

The dashboard saved object sets "description": "". The description is shown in the Kibana dashboard listing and in the integration's asset list, so an empty value leaves users without any indication of what the dashboard covers. The saved search GDACS Raw Events in kibana/search/ has the same empty description.

Recommendation:

Populate the description in the exported saved object:

        "description": "Overview of GDACS natural disaster alerts: alert level and event type breakdowns, affected countries, and a map of event centroids and affected-area geometry.",

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

@vera-review-bot

Copy link
Copy Markdown

Review summary

Issues found across the latest commits 25229f4 — 3 high, 3 medium, 3 low
  • 🟠 The pipeline pins ecs.version to 8.11.0 while _dev/build/build.yml imports ECS v9.4.0 (link) (Unresolved)
  • 🟠 sample_event.json does not match what the pipeline actually produces (GeoJSON objects and scalars instead of WKT strings and arrays) (link) (Unresolved)
  • 🟠 fields.yml declares object subfields underneath the geo_shape and geo_point leaves (affected_area.coordinates/.type, geo.location.coordinates/.type) (link) (Unresolved)
  • 🟡 event.modified is not an ECS field, so this adds a custom field to the ECS-reserved event.* namespace (link) (Unresolved)
  • 🟡 owner.type is set to community even though the package is owned by the Elastic team elastic/security-service-integrations (link) (Unresolved)
  • 🟡 The request's toDate is read from the stored cursor instead of the current time, so each poll queries a window that ends at the previous poll's date (link) (Unresolved)
  • 🔵 The validation step tells users to look for a geo.location field, but the package writes gdacs.geo.location (link) (Unresolved)
  • 🔵 The How it works section says events are deduplicated on {event_id}-{episode_id}, but the fingerprint also includes the geometry id for tropical-cyclone child documents (link) (Unresolved)
  • 🔵 The [GDACS] Events Overview dashboard ships with an empty description (link) (Unresolved)
Issues found across earlier commits de2f39302605ed (250 commits) — 4 high, 1 medium, 2 low
  • 🟠 The pipeline pins ecs.version to 8.11.0 but _dev/build/build.yml pins the ECS reference to v9.4.0 (link) (Unresolved)
  • 🟠 The pipeline deletes event.original based on the absence of the preserve_original_event tag, which is a prohibited pattern (link) (Unresolved)
  • 🟠 The CEL program nests .as() bindings six levels deep, exceeding the hard cap of five (link) (Unresolved)
  • 🟠 fields.yml declares .coordinates and .type sub-paths beneath the geo_shape and geo_point leaves gdacs.affected_area and gdacs.geo.location (link) (Unresolved)
  • 🟡 owner.type is set to community while CODEOWNERS assigns the package to the Elastic team elastic/security-service-integrations (link) (Unresolved)
  • 🔵 categories is set to the generic custom bucket, which hides the package in Fleet's browser (link) (Unresolved)
  • 🔵 The README validation step tells users to look for a geo.location field, but the pipeline writes gdacs.geo.location (link) (Unresolved)
Issues found across earlier commits 15f50b1, 4e6a6d6 — 1 medium, 2 low
  • 🟡 The extract_geometry script computes gdacs.geometry_type and remove_raw_fields deletes it a few processors later, so the value never reaches the index. Either keep the field (drop it from remove_raw_fields and declare it in fields.yml) or stop computing it. (link) (Unresolved)
  • 🔵 set_message builds message from a mustache template instead of copy_from, so a document without gdacs.description gets message set to an empty string. Switch to copy_from with ignore_empty_value. (link) (Unresolved)
  • 🔵 The descriptions for gdacs.class and gdacs.polygon_label list values the integration never produces ("Poly_area", "Affected Area") and omit the ones it does. Update both descriptions to the actual value sets. (link) (Unresolved)
Issues found across earlier commits b44ccd0 — 5 high, 4 medium, 2 low
  • 🟠 The remove_event_original condition calls .contains() on ctx.tags without a null guard, so every document fails the pipeline when the user clears the Tags variable. Add a null check to the condition. (link) (Unresolved)
  • 🟠 event.category and event.type are ECS array fields but are written with the set processor, which overwrites anything already present. Use append with allow_duplicates: false instead. (link) (Unresolved)
  • 🟠 The pipeline-level on_failure block deviates from the required structure: the steps are in the wrong order, the preserve_original_event tag append is missing, and the mustache placeholders use double braces which HTML-escape the error text. Rewrite it to the standard three-step form. (link) (Unresolved)
  • 🟠 base-fields.yml does not match the required six-entry shape: event.module and event.dataset are missing, input.type is not part of it, and no entry uses external: ecs. Replace the file with the standard six entries. (link) (Unresolved)
  • 🟠 The geo_point and the rest of the geo fieldset are declared at the document root (geo.location, geo.name, geo.country_name) instead of nested under a parent entity. Move them under the gdacs namespace. (link) (Unresolved)
  • 🟡 The CEL-only opening processors are missing, so a collector error document (which has error.message but no message) reaches the json processor and throws. Add the agentless remove and the terminate guard, and make the json processor conditional. (link) (Unresolved)
  • 🟡 Four subfields are declared underneath the geo_point and geo_shape leaves (geo.location.coordinates/.type and gdacs.affected_area.coordinates/.type) but the pipeline never writes them. Delete the four declarations. (link) (Unresolved)
  • 🟡 The ECS categorization models natural disaster alerts as threat intelligence (event.category threat, event.type indicator) which also conflicts with event.kind alert. Reconsider the triple, most likely event.kind alert with no threat categorization. (link) (Unresolved)
  • 🟡 When the last page of a run is partial the cursor keeps the incremented page_number and the frozen date window, so the next poll spends a whole interval fetching an empty page. Reset the cursor when want_more is false. (link) (Unresolved)
  • 🔵 The CEL program computes geometry_doc, geometry_id and geometry_role for tropical cyclone child documents, the pipeline renames them into gdacs.*, and then remove_raw_fields deletes them all. Keep gdacs.geometry_role or stop producing it. (link) (Unresolved)
  • 🔵 The package icon still carries the scaffold placeholder title "Sample logo" and the screenshot title has a doubled space. Give both descriptive titles. (link) (Unresolved)

A new commit triggers another review — at most once every 15 minutes. I skip the PR while it's approved or has merge conflicts.

🤖 AI-Generated Review | Vera Review Bot | 📚 Knowledge base: integration-skills

⚠️ Automated review — verify suggestions before applying.

@kcreddy kcreddy left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor nit on ecs.version. LGTM otherwise for my comments.

Please wait for @efd6 approval.

Comment thread packages/gdacs/data_stream/events/elasticsearch/ingest_pipeline/default.yml Outdated
Comment thread packages/gdacs/manifest.yml
@mergify

mergify Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request

@efd6 efd6 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, then lgtm

Comment thread packages/gdacs/img/gdacs_events.png
@elastic-vault-github-plugin-prod

Copy link
Copy Markdown
Contributor

✅ All changelog entries have the correct PR link.

@infra-vault-gh-plugin-prod

Copy link
Copy Markdown

💚 Build Succeeded

History

@efd6
efd6 merged commit b3b8435 into elastic:main Aug 9, 2026
13 checks passed
@elastic-vault-github-plugin-prod

Copy link
Copy Markdown
Contributor

Package gdacs - 0.1.0 containing this change is available at https://epr.elastic.co/package/gdacs/0.1.0/

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation. Applied to PRs that modify *.md files. enhancement New feature or request New Integration Issue or pull request for creating a new integration package. Team:Security-Service Integrations Security Service Integrations team [elastic/security-service-integrations]

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants