Skip to content

v4.4.21

Choose a tag to compare

@anionDev anionDev released this 11 Sep 16:58
· 88 commits to main since this release
v4.4.21
ab0abd3

Release notes

Changes

  • SCTaskRunnerServer transfers the archives of a job chunk-wise between the network-connection and the disk instead of
    holding them in memory as a whole. Before this change the server read the complete payload-archive (the whole repository
    of the client) into memory when a job was submitted, and read the complete result-archive (the whole codeunit-folder
    including its build-output) into memory when the result was fetched. Both are regularly hundreds of megabytes up to
    several gigabytes, which a runner that runs in a container with a memory-limit can not hold in addition to the build it
    executes. __start_job now receives the request-body as a stream, and the result-archive is sent from its file.
  • SCTaskRunnerServer no longer deletes the workspace of a job while the result-archive of that same job is still being
    packed. A client deletes a job as soon as it stops waiting for the result, which also happens when it gives up on a
    still-running result-request - for example because a reverse-proxy in front of the runner ran into its read-timeout while
    the (large) archive was being created, since no data flows during that time. The deletion then removed the workspace
    underneath the running transfer, which failed with a FileNotFoundError naming an arbitrary file of the build-output and
    could leave a partially deleted workspace behind. A per-job transfer-lock now makes the deletion wait for a running
    transfer; because the job is removed from the job-list first, no new transfer can start in the meantime.
  • SCTaskRunnerServer logs its start and the result of every job it runs. Before this change a runner produced no output
    at all: everything it logs is logged as Information, while the default-loglevel of a ScriptCollectionCore is
    Warning, so all of it was discarded - neither the start of a runner nor the jobs it ran were visible in the
    console-output of the process, which is the only log a permanently running server has. A server which creates its own
    ScriptCollectionCore now raises that loglevel (a caller which passes one keeps the loglevel it configured there), the
    start is logged before the server begins to listen (so a start which fails on its port or its certificate is visible
    too), and every job is logged when it is accepted and when it ended, with its state and its exitcode.
  • TFCPS_RemoteBuild ends a delegated build with a defined statement when the connection to the runner breaks. Before
    this change every connection-error (for example a reset connection while the client polled the state of the job)
    propagated as the raw error of urllib and ended the build with a traceback which stated nothing about the delegated
    build. Such an error is now reported as RunnerNotReachableError, naming the runner and the request; an error-status
    answered by the runner is reported as RunnerRequestFailedError. Because a job runs on the runner independently of the
    connection to it, requests which concern an already existing job are repeated for at most five minutes while the runner
    is not reachable, instead of giving up a build which is still producing a result; submitting a job is not repeated,
    because that would start a second build. A runner which does not know a job any more (which is what a restart of a
    runner while a job was running looks like from the client, since a runner holds its jobs in memory only) is reported as
    exactly that instead of as an unspecific 404.
  • SCTaskRunnerServer answers the submit of a job as soon as the transferred archive has arrived and extracts it in the
    thread of the job. Before this change the extraction happened in the request which submitted the job, which takes
    minutes for the repository of a large codeunit and transports nothing while it runs: the client and a reverse-proxy
    between it and the runner ran into their timeouts although the job was being prepared correctly. A failing extraction
    is a failed job now (with a state and an exitcode like every other job) instead of a failing request.
  • TFCPS_RemoteBuild waits an hour instead of five minutes for the result-archive of a job, because the runner packs the
    complete build-output of the codeunit before it sends the first byte of it. Every other request keeps the previous
    timeout, so a connection which transports nothing for five minutes still counts as interrupted.
  • SCTaskRunnerServer separates an error of the runner itself from an error of what it was asked to build. An
    unexpected error while a request is handled is answered with a 500 which names an error-id and contains no details -
    the details (stacktraces and paths inside the runner) go to the log of the runner under that error-id. Before this
    change such an error left the request without any answer at all, which a client sees as a broken connection. A job
    which fails because of the runner states that and the error-id in its log, again without the details. The output of
    the program of a job is unchanged: it is what the client asked for and is transferred in full, so a build which does
    not compile states why. A request which does not contain the headers which state what is to be run is answered with a
    400 which names them.
  • SCTaskRunnerServer logs what a job is and what it does: the program with its arguments, its folder, its codeunit and
    the amount of transferred bytes when a job arrives; which repository a job builds (taken from the .git-configuration
    of the transferred repository, so the client does not have to send anything in addition, and with credentials removed
    from that address); every 60 seconds that a job is still running, since how long and with how much output so far; the
    duration of a job when it ends; and the last 25 lines of the output of a job which failed, so that the log states why
    a build failed even when no client fetched that output.
  • A remote build transfers back only the folder in which the delegated build-step produces its result instead of the
    whole codeunit-folder, and the client no longer deletes anything for it. Before this change the local
    codeunit-folder was removed and replaced by the one of the runner, which destroyed everything of that codeunit which
    was not part of the answer of the runner - including files which are not in git and can therefore not be restored -
    and which on Windows can not even complete, because the build-script of the codeunit runs in a folder below the
    codeunit and a folder a process runs in can not be removed there: the deletion failed in the middle and left the
    codeunit incomplete. Besides that, the workspace of a runner contains state of that machine after a build (its
    toolchain-paths in local.properties, its absolute paths in .dart_tool, caches), which has no business in the
    repository of a developer. The client states the folder as the new header X-Result-Folder of the submit, and
    run_program_on_remote_runner has a new parameter result_folder for it; for a flutter-codeunit these are the folders
    the artifacts were copied from anyway (build/windows/x64/runner/Release, build/macos/Build/Products/Release,
    build/ios/iphoneos, build/app/outputs/bundle/release) and for a C/C++-codeunit it is the sourcecode-folder, in
    which its build-commands produce their output. A result-folder outside of the workspace of the job is refused, and a
    job which produced no result-folder is answered with a 409 naming it.
  • New specification openspec/specs/testcases/spec.md: a testcase does not send a request to a server and does not
    verify that a connection to one can be established - neither to a server it starts itself, nor to one in the network,
    nor to one in the internet. A testcase states something about the code of this product, while such a result would also
    state something about what is reachable, about a free port and about what a third party currently answers.
  • The android-target of a flutter-codeunit publishes its artifacts again: the app-bundle which came back from the runner
    as Other/Artifacts/BuildResult_AAB/<codeunit>.aab and the universal apk generated from it with Google's bundletool
    as Other/Artifacts/BuildResult_APK/<codeunit>.apk. Both were switched off with the note that the apk-generation is
    not platform-independent, which was true: it called the java of the machine, which is not part of every machine
    which builds a codeunit. It uses the pinned JRE from the global cache now - the same one the plantuml-rendering
    already uses - so the step needs nothing which a build does not need anyway. Two defects of that code were fixed on
    the way: the arguments of bundletool are passed as an array instead of as one string (they contain absolute paths, and
    a path containing a space was split into several arguments), and the parameter which states whether the cached
    bundletool may be used was passed inverted.
  • TFCPS_RemoteBuild states that a runner could not be reached when it selects the runner for an operating-system.
    Before this change a runner which did not answer was reported as No configured remote-build-runner provides the required operating-system, although the runner was configured and it was merely unknown whether it provides that
    operating-system, and the unreachability itself was logged with a complete python-traceback.