24.10.0
We are proud to announce the release of:
✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨
Ryax 24.10.0
✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨ ✨
Multi-site full power!
New features
- A new service called Ryax Worker can now be used to attached any Slurm or Kubernetes cluster resources
- Ryax can now run any action on SLURM and Kubernetes seamlessly
- Action are now scheduled according to user defined constraints and objectives
- Add the possibility to pin Ryax services to a dedicated resources (nodeSelector)
- Enhance Ryax documentation with updated content (doc)
- New Jupyter Notebook action with GPU support in default actions
- Action builds now can be canceled
- Kubernetes addon now support injection of service
Bug fixes and Improvements
- Fix volume permission for NFS based storage volumes (defaults to 1200 now)
- Fix fail properly when a pip install fails during builds
Upgrade to this version
This is a major release of Ryax which implies some extra step for the upgrade.
Update configuration
This release introduce a new service, the Worker. In order to define the nodes that will be used by your actions, the Worker requires a site configuration. Please, add a configuration in your Ryax installation configuration file using the following example: in your local cluster has a node pool named default with a label my.provider.com/pool-name: default on each node, it has 4 CPU and 8G of memory per node.
worker:
values:
config:
site:
name: local
spec:
nodePools:
- cpu: 4
memory: 8G
name: default
selector:
my.provider.com/pool-name: defaultSee the Worker configuration documentation for more details.
Update DNS
If you use public IP with TLS enabled, you will need to create a new DNS entry to support all subdomain for your cluster. This is used for example for an external Worker to access the internal container repository.
Please add an entry in your DNS using star notation:
*.<clusterName>.<domainName>
See installation doc for more details.
Add HPC site
The users of HPC actions have to install a Worker dedicated to each cluster following this documentation.
Apply and clean
Once configured, you can apply the configuration with ryax-adm as usual.
The log capture service, Loki, was moved into the ryaxns namespace. Thus, the old Loki deployment can be removed.
After applying, we have to remove the old deployment:
helm uninstall -n ryaxns-monitoring loki
kubectl delete pvc -n ryaxns-monitoring storage-loki-0The Worker is now handling deployment. So, to avoid dangling actions and failing deployment, you have to clean the Runner state.
Be aware that, this will reset the execution history and stop all running workflows.
ryax-adm clean runner worker