Skip to content

v0.5.0

Choose a tag to compare

@kbushick kbushick released this 30 Sep 22:31

Orchestrator Version 0.5 Release Notes - 2025-08-15

Overview

Orchestrator is a modular, extensible Python framework designed to streamline the end-to-end workflow for building, training, testing, running, and analyzing Interatomic Potentials (IAPs) and large-scale molecular dynamics (MD) simulations. It provides a uniform API for integrating diverse simulation codes and tools, reducing human effort in complex scientific workflows.

Design Features

1. Modular Architecture

  • Abstract Base Classes: Core functionality is defined via abstract classes, enabling drop-in replacement of concrete implementations via the uniform API.
  • Extensible Factories/Builders: Uniform instantiation of modules via factory and builder patterns.

2. Supported Workflows

  • IAP Development: Build, train, validate, and deploy interatomic potentials (empirical, ML-based).
  • Simulation Management: Run and analyze MD simulations, including property calculations (melting point, elastic constants, etc.).
  • Ground Truth Calculation: Generate, run, and parse DFT or other ground truth calculations for training data generation, cataloguing outputs and settings for consistency.
  • Active Learning/Pruning: Dataset augmentation and reduction using scoring and selection modules.

3. Integration with Existing Code Infrastructure

  • KIM Suite Integration: Seamless use of KIM-API, KIMkit, KIM Tests, KliFF, and Colabfit for model management and simulation.
  • AiiDA Support: Automated provenance tracking, error handling, and job management for DFT codes (VASP, Quantum Espresso).
  • ASE Atoms: ASE Atoms are used as the internal representation for atomic-scale configurations.

4. Data Management

  • Flexible Storage: Local (filesystem) and Colabfit (PostgreSQL) storage backends for datasets, supporting ASE Atoms as the core data structure. KIMkit provides similar functionalities for IAPs.
  • Metadata & Versioning: Built-in metadata tracking, version control, and property mapping for datasets and potentials.

5. Testing & Validation

  • Comprehensive Test Suite: Unit tests for all modules, with curated reference outputs and pytest integration.
  • Semi-Automated Checking: Scripts for setup, execution, and validation of tests across modules.

6. Job Execution & Workflow Management

  • Local & HPC Execution: Support for local execution, Slurm, LSF, and hybrid Slurm-to-LSF workflows.
  • Asynchronous/Synchronous Modes: Flexible job submission and blocking/waiting mechanisms.
  • Checkpointing & Restart: Robust checkpointing for long-running or multi-step workflows.

7. Analysis & Scoring

  • Score Modules: Quantify uncertainty, diversity, efficiency, and importance using information-theoretic and UQ metrics (LTAU, QUESTS, FIM).
  • Augmentor: Advanced dataset pruning, novelty detection, and subcell extraction for active learning.

Module Organization

Module Type Description
Turn-key Application-style execution (Executor, under development)
Coordinating Modules which leverage one or more "atomic" modules in simple to complex coordination for their operation. Include: Augmentor, TargetProperty
Atomic Modules which serve as the building blocks of core functionality. Simpler "input" --> "output" usage. Include: Descriptor, Oracle, Potential, Score, Simulator, Trainer
Utility Backend support for module, data, and file management. Include: Factory/Builder, Restart, Storage, Workflow

More Information

Full docs can be found at https://orchestrator-docs.readthedocs.io/en/latest/index.html

Bugs, feature requests, or other comments can be addressed via Issues
or messages sent to orchestrator-help@llnl.gov

Full Changelog: https://github.com/LLNL/orchestrator/commits/v0.5.0