You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat/integration-tests (#119)
* build(charts): update desmos image tag to v2.0.18
* add it tests
* fix(it): unblock desmos startup in the k3s harness
Two independent problems kept the desmos pod in Init:0/1:
- dome-ied was deployed even though nothing asks for it. Helm enables a
dependency when the path in its `condition` does not exist, and no values
file defines `dome-ied.enabled`. Its templates set no namespace, so it
landed in `default`, where k3s-maven-plugin waited on it until timeout.
- the dlt-adapter service was published on port 80, the upstream chart
default. `charts/access-node/values.yaml` does set `service.port: 8080`,
but under the key `dlt-adapter`, which reaches neither of the aliased
subcharts. desmos calls the adapter on 8080, so its check-external-services
initContainer polled /health forever.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(it): ignore empty-vs-null collections when comparing replicas
Replication works, but the NGSI-LD round trip normalises an absent
collection into an empty one, so the consumer's copy came back with
productSpecificationRelationship: [] where the provider held null. Same
entity, failed assertEquals.
Flatten empty collections to null on both sides instead of comparing a
hand-picked subset of fields, so the whole entity is still compared. Only
empty collections are flattened: if replication ever dropped the contents
of a populated list, the assertion still fails.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* build(actions): downgrade action versions for consistency
* ci(it): give the k3s harness room to pull its images
The CI job timed out in pre-integration-test waiting for the infra
dome-contract job, 300s being the k3s plugin default. The job is not
stuck: locally it completes in 92s, of which ~74s is image pulling --
62s of that the 808MB quay.io/wi_stefan/dome-contract image alone, on
top of the init container's apt-get. A 2-vCPU runner does not fit that
into 300s, so raise the wait to 900s for every execution of the plugin.
Also dump the cluster when the job fails: the plugin only reports which
resource it gave up on, never why, which left the previous failures
impossible to tell apart from a pod that never came up.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(it): drop the init container's runtime package install
The dome-contract job never completed on CI: its init container ran
apt-get to fetch curl and jq, external DNS does not resolve from inside
the pod on the runners ("Temporary failure resolving security.ubuntu.com"),
so the container exited non-zero and OnFailure restarted it forever. The
earlier timeout bump only made the build take 17 minutes to reach the
same dead end, so revert it.
Note the asymmetry the cluster dump showed: image pulls all succeeded in
under 4s, because containerd resolves on the node, while apt-get resolves
from the pod network. Taking the install out removes the pod's only
external dependency -- it now just talks to blockchain-testnode over
in-cluster DNS, which CoreDNS answers without upstream forwarding.
Verified by running the rendered init.sh as the pod does, /bin/sh -ec
against a geth dev node: both accounts funded, exit 0, and a second run
still exits 0 with the imports reporting "account already exists", so the
retry path the backoffLimit relies on stays intact.
The image is alpine, hence sh over bash, and is pinned to :alpine rather
than :latest so the implicit Always pull policy does not hit Docker Hub
on every CI run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(it): wait for the dev node before funding the test accounts
Taking the apt-get out of the init container removed the ~12s it
incidentally spent installing curl and jq, which had been giving the geth
dev node time to come up. Without it the container gets there first,
eth_accounts comes back empty, and because curl exits 0 even when the RPC
answers with an error the script funded nobody and still reported success.
The contract deployment then failed with INSUFFICIENT_FUNDS, and since
OnFailure only restarts the main container, never the init one, the pod
stayed in CrashLoopBackOff with both accounts on a zero balance.
So poll for the dev account rather than assume it, and check each balance
after funding it: an unfunded account otherwise only surfaces at contract
deployment, too late for the init container to retry.
Verified by starting the script 20s ahead of the node: it waits, imports
and funds both accounts, and exits 0 with each holding 0xfffffffffffffff0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(it): run the init container on glibc, not musl
CI got past the package install and then sat in the init container for
five minutes without a restart, so the poll loop was spinning: the pod
could not reach the dev node over its in-cluster name. The alpine tag I
picked is musl based, and musl walks the search list differently from
glibc under the ndots:5 that Kubernetes writes into every pod, a mismatch
that bites on some hosts and not others -- it resolved fine locally. The
rest of this harness was built around a glibc image, so use the ubuntu tag
of the same image: still curl and jq preinstalled, still no apt-get, and
the tag pins the distro that plain "ubuntu" no longer does.
Bound the loop while here. Untimed curls against an unreachable endpoint
took 400s to give up, longer than the 300s the k3s plugin waits for this
job, so the plugin killed the container before it could report anything.
With per-attempt timeouts and a lower cap it now gives up in 121s and
says so, and logs each attempt instead of failing mute.
Also name the init container when dumping logs: --all-containers skips
init containers, which is why the last dump came back with nothing but
"waiting to start: PodInitializing" for the one container that mattered.
Probe DNS and the RPC from a throwaway pod too, to tell a broken pod
network apart from a slow one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci(it): pin the k3s image and probe each network layer
These tests passed on CI until they were taken out of the workflows, and
the chart reached the dev node by the same in-cluster name then that it
does now, so cluster networking used to work on the runners. Nothing in
the harness explains the change -- but the k3s image is pulled by the
"latest" tag, which the plugin itself warns about, so the cluster under
the suite has been drifting with nothing in this repo changing. Pin it to
the line these tests were written against.
The evidence so far only says pods cannot resolve service names, with
resolv.conf correct and CoreDNS healthy and quiet, meaning the queries
never arrive. That points below DNS, but reaching the node by name walks
CNI, kube-proxy and CoreDNS together, so it cannot say which. Probe the
pod IP, the service IP and the name in turn: the first to fail names the
layer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci(it): dump what kube-proxy did, and read the probe from the pod log
The suite dies because pods cannot reach a ClusterIP: the service IP of
the dev node does not answer, and the name does not resolve because
CoreDNS sits behind a ClusterIP too. CoreDNS itself is healthy and quiet,
so the queries never arrive. That is below DNS, but nothing so far says
whether kube-proxy failed to program the service VIPs or whether pod
traffic never leaves the pod at all.
So dump the nat table and the k3s container log, where kube-proxy runs
and would complain, plus the netfilter modules it depends on. Add a
reachability check against the CoreDNS pod IP, which separates CNI from
kube-proxy without going through a service.
The previous probe lost its first lines, including the pod-IP check that
would have settled this: kubectl run --attach connects after the
container has started. Read the log once the pod finishes instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(ci): load br_netfilter before starting the k3s cluster
The suite never got past the infra job because pods could not reach any
ClusterIP. kube-proxy was not at fault -- it programmed the KUBE-SVC
chains correctly -- but without br_netfilter the pod traffic crossing the
CNI bridge never traverses iptables, so the DNAT for a service VIP is
simply skipped. Probing each layer in turn showed exactly that shape: pod
to pod answered, pod to ClusterIP did not, and name resolution failed
with it because CoreDNS is reached through a ClusterIP. That last part is
what surfaced first, months of symptoms ago, as apt-get in the init
container failing to resolve anything.
The runners no longer come with the module loaded, which is also why this
used to pass with no change on our side: k3s logs "Failed to set sysctl:
open /proc/sys/net/bridge/bridge-nf-call-iptables: no such file or
directory" there and nothing of the sort on a machine where it is loaded.
Load it in both workflows, since both run the integration tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Update helm documentation
* ci(it): run the integration tests through a single lifecycle
Naming four phases made Maven walk the lifecycle up to each one, so a run
generated the twelve OpenAPI clients three times, templated the charts
three times, compiled the 620 sources three times, and executed the
Cucumber scenario twice. The second scenario run is not just wasted time:
it creates another product offering and replicates it on top of whatever
the first one left behind. "clean verify" reaches the same phases once --
measured locally at 5:49 against 8:35, with every plugin down to a single
execution and the tests still green.
pre-release and release were also stopping at integration-test. Failsafe
records results there and only fails the build from its verify goal,
bound to the verify phase, so a broken test would not have failed either
job -- and both gate a deploy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: stop the pre-release deploy from rewriting the PR branch
The deploy job commits the chart version bump and the regenerated helm
docs back to the PR branch, which is how the PRE version reaches main and
has to keep working. Its side effects are what needed fixing.
On a pull_request event checkout lands on the merge ref rather than the
branch tip, so pushing HEAD back carried a merge of main onto the branch:
593379e, 9c8c96e and 3cbc9d6 all arrived that way, one per run that got
as far as deploying. Checking out github.head_ref pushes only what the
job actually changed, and fetch-depth 0 covers what the manual unshallow
was for -- that command fails on a complete clone.
The commit also becomes the PR head, and runs attributed to
github-actions[bot] land in action_required, which is why the PR showed
no checks at all after the suite first went green. Marking the bot's
commit so GitHub skips CI stops those runs from being queued at all.
The marker is spelled out in the workflow rather than here, since GitHub
reads it from any commit message in the push -- including this one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(ci): only rebuild the chart index when something was released
The deploy job broke with "cr: command not found" as soon as a second
push reached it. chart-releaser installs the cr binary lazily, on the
path where it has charts to package; once the PRE tag for this PR
existed, it found no chart changes, returned early and never installed
anything, while the next step called cr regardless.
Installing cr by hand would only have moved the error: cr index builds
the index from .cr-release-packages, which stays empty when nothing is
packaged, so the index would come out unchanged and the commit would
fail on having nothing to commit. With no chart released there is no
index to rebuild, so gate the step on the action's own changed_charts
output, which it documents as empty in exactly that case.
This was latent rather than new: deploy had never run twice on one pull
request until the integration tests started passing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: run the integration suite from one reusable workflow
The it job was copied into three workflows, so a pull request built two
full k3s clusters to check the same thing -- 6m37s and 6m41s on the last
push -- and any fix had to be applied three times. It already had been
missed once: release.yaml never got the br_netfilter step, so merging to
main would have failed the way pull requests did until now.
Move the job to a workflow_call workflow and have pre-release and release
call it. A pull request now runs the suite once, and there is a single
place to change it.
Note for whoever merges: the Test check no longer exists on its own, the
job reports as Pre-Release / it. If main's branch protection requires
Test by name, it needs updating first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: make the failure dump generic instead of bug-specific
The dump was written while chasing one particular failure, so it named
dome-contract and blockchain-testnode outright and carried a layered
network probe, an iptables nat dump and lsmod. That question is answered
-- the runners had stopped loading br_netfilter -- and none of it helps
with whatever breaks next.
Keep the parts that would help with any failure and let them find their
own targets: describe and dump every pod that is neither Completed nor
Running with all containers ready, which covers one stuck initialising,
one crash looping and one whose sidecar never came up. Naming each
container individually stays, since --all-containers omits init
containers and that gap cost a whole cycle of blind debugging. The k3s
container log stays too: a cluster that never came up explains itself
there and nowhere else.
Output is grouped so a run with many unhealthy pods stays readable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* build(actions): update the actions and pin the third-party ones
The workflows were spread across checkout v1, v2 and v3 and setup-java
v1, all deprecated and already being forced onto Node 24 with a warning.
This reverses 35ed8be, which had downgraded them for consistency.
Bring the GitHub-owned ones up to current and give setup-java the
distribution it has required since v2, plus the maven cache -- the logs
show every dependency being fetched again on each run. setup-go lost its
implicit version default, so name one rather than take whatever the
runner ships. The duplicate azure/setup-helm goes: the second was a
no-op.
Pin the third-party actions by commit. tj-actions/changed-files was on
v14.6, from before the March 2025 compromise in which its tags were
rewritten to leak secrets, and a moving tag would have carried that in;
ad-m/github-push-action and helm/chart-releaser-action were tracking
master and main outright. The readable tag stays beside each SHA.
Checked before pinning: changed-files still exposes
all_changed_and_modified_files under v47, which the label and release
checks read.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: report the integration test results on the pull request
Until now the outcome was only visible by opening the run and reading
the Maven output, which is a poor way to find out why a check went red.
Failsafe already writes JUnit reports naming the feature and the
scenario, so publish those: the action posts a comment on the pull
request and keeps editing it on later pushes rather than adding a new
one, writes the job summary, and raises a check run.
Reporting must never decide whether the build passes, so the step is
continue-on-error and runs on always() -- a failed suite is exactly when
its output matters. The token is read-only on pull requests from forks
and the comment would be rejected there.
The scopes it needs are declared by the callers: a reusable workflow
cannot grant itself more than the workflow calling it already has.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* build(deps): let dependabot keep the pinned actions moving
Pinning the third-party actions by commit stops a rewritten tag from
reaching us, but it also freezes them: with nothing bumping the pins they
rot in place and the pin turns from a defence into debt. Dependabot reads
the tag written beside each SHA and moves both together, so it is what
makes that decision sustainable. The repository had no configuration at
all until now.
Two settings carry their weight here. A cooldown holds back releases
published in the last week, because pinning keeps a rewritten tag out but
does nothing about adopting a compromised release on the day it ships --
which is how tj-actions/changed-files would have reached us. And grouping
the minor and patch bumps into one pull request matters more than usual,
since every pull request builds a k3s cluster and runs the whole suite;
majors stay out of the group so they land alone, which is where the
breakage tends to be.
Covers the two ecosystems that exist: the workflows, and the pom under
it/. Helm values are not something Dependabot reads, so the chart image
tags stay manual.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>