Skip to content

Multi-AutoML Interface v5.6.0

Latest

Choose a tag to compare

@github-actions github-actions released this 01 Oct 03:56

Multi-AutoML Interface 5.6.0

Fixed

  • An H2O run no longer leaves a Java cluster running for the rest of the process. The client
    keeps one connection per process, so train_h2o_model used to hand the live H2OAutoML to the
    session and never shut the cluster down: that is what let post-training prediction work, and it
    also meant 2-4 GB of heap parked in a shared, multi-session server, plus the risk of one run's
    cleanup_h2o() killing the cluster another run was still training on. h2o_cluster() now holds
    a lock, starts a private cluster on a free loopback port, and shuts it down on the way out -
    including when the body raises or initialize_h2o() itself fails, which used to strand the lock.
    The free port matters: without one the client adopts whatever cluster is listening on 54321,
    which is how a second session ended up shutting down the first one's JVM.
    Training returns H2OSessionModel(run_id); predict_with_h2o reopens a cluster, reloads the
    model with fetch_h2o_model, and materialises the result as numpy before releasing. Verified
    live on the 3.11 interpreter: after train, after a leaderboard read and after prediction, jps -l
    lists no H2OApp.
  • The H2O model reload path had never worked. fetch_h2o_model only matched a *.zip
    artifact, but h2o.save_model writes an archive named after the model id with no extension,
    so every reload raised H2O model not found in artifacts. It was invisible because training had
    always handed the live object to the caller. Both layouts are accepted now, .zip first.
  • The H2O Inspector reads its leaderboard from the run's artifacts. The first version of the
    reloaded-cluster Inspector called model.leaderboard on the model h2o.load_model returns; that
    attribute does not exist. Measured on h2o 3.46 with a real run: hasattr is False for
    leaderboard, leader, best_model and all_models, and reading it raises
    AttributeError: type object 'ModelBase' has no attribute 'leaderboard' - a reloaded model is one
    estimator, not the AutoML object. h2o_run_leaderboard(run_id) now parses the
    h2o_leaderboard_<run>.csv the run logs (writer and reader share one constant, and a test fails
    if the training path stops using it), which also means opening that expander costs no JVM. A run
    that trained no model says so instead of showing an empty table.

Fixed

  • Two FLAML runs in one process used to kill each other. flaml/tune/tune.py keeps its trial
    runner in a module global (_runner, line 45), so a second search replaced the first one's
    runner and the first died with AttributeError: 'NoneType' object has no attribute 'stop_trial'.
    Reproduced with the current code: three fits started in threads straight against
    _train_flaml_model -> one failed with exactly the message the UI had shown; the same three
    through train_flaml_model, which now queues behind _EXPERIMENT_LOCK the way PyCaret does, all
    finished. A run queued behind another can still be cancelled (StopIteration), and two runs
    started from the page now both complete.
  • FLAML could not train with a validation holdout and cross-validation together - which is what
    the split section produces: fit answered AssertionError: eval_method must be 'auto' or 'holdout' for custom validation data. _apply_evaluation_settings now chooses holdout when a
    validation frame arrived and cv only when it did not, in both the single-target and the
    multi-target branches. A run started from the interface with Cross-Validation selected finished
    in 2m 20s where the same form had failed in five seconds.
  • Computer Vision Multi-Label Classification could not be trained from the interface at all.
    Driving the real UI (upload a ZIP + annotations.csv, pick the row, train, predict) found three
    faults behind the engine-level tests, all of them in the path between the page and the engine:
    • app.py split every multi-label selection into one experiment per column. For Tabular that is
      what the engines want; for Computer Vision it handed train_model a single label column, so
      the run died on its own "needs at least two label columns" guard and the page showed two
      failures for one dataset. The decision now lives in label_run_plan()
      (src/task_catalog.py), which keeps the tabular fan-out and gives the CV row one run - and the
      test that mirrors it also checks that app.py still calls it.
    • After both predictors were fitted and saved, the reporting step called
      MultiModalPredictor.evaluate(), which computes ROC AUC - undefined when the holdout holds
      one class, which a random 10% split of a small dataset does routinely. The whole run was lost
      over a metric. _evaluate_with_single_class_fallback() now scores accuracy in that case and
      says so in the log; the live run recorded circle_roc_auc = 1.0 and red_accuracy = 0.6
      instead of failing.
    • The Pipeline Inspector asked the predictor for a leaderboard(). MultiLabelAutoGluonPredictor
      and MultiModalPredictor have evaluate(), not a leaderboard, so a completed run's Inspector
      showed AttributeError. It now reads the leaderboard.csv the run logged
      (read_run_leaderboard), the same source the metrics came from.
  • The Prediction section did not appear after clicking 🔮 Predict. The button lives inside the
    5-second dashboard fragment and only sets session_state; the section below it is outside the
    fragment, so it re-rendered only on the next unrelated interaction (switching pages made
    "Active model: autogluon" and the batch uploader appear). Loading the model now reruns the app.

Added requirements.txt

gains autogluon.tabular/core/features/common==1.6.3 plus the six packages its closure needs
(boto3, botocore, s3transfer, jmespath, networkx, psutil). Measured on a fresh
Python 3.12 install of the new lock: 16 packages added and no pin moved - AutoGluon's own
caps (numpy<2.6, scipy<1.19, pandas<2.4, scikit-learn<1.10, Pillow<13) are already
satisfied by what the project pins - 41 MB of extra site-packages, pip-audit -r requirements.txt --strict still exits 0, tests/test_engine_matrix.py passes the five Tabular
rows there, and the six Text/Vision/Multimodal rows skip.

  • AutoGluon multimodal stays out, and the reason is now measured rather than assumed.
    autogluon/multimodal/data/templates.py does import pkg_resources at module scope (no
    try, no lazy branch: with setuptools 84, import autogluon.multimodal raises
    ModuleNotFoundError: No module named 'pkg_resources'), and pkg_resources is shipped by
    setuptools only through 81.0.0 - it is gone from 82.0.0, which is still inside the range
    GHSA-h35f-9h28-mq5c / PYSEC-2026-3447 flag, fixed only in 83.0.0. pip install autogluon.multimodal into the shipped closure resolves setuptools down to 81.0.0 (dry-run
    output: - setuptools==84.0.0 / + setuptools==81.0.0) and adds torch, transformers, ray and
    scikit-image, about 1.7 GB. So the installer keeps the audit gate clean and the catalog keeps
    hiding those rows until the engine is really importable.

  • tests/test_h2o_cluster_lifecycle.py - 16 tests against a fake h2o module, so the whole
    lifecycle is covered without Java: one cluster per nested operation, release when the body
    raises, release when the cluster never started, exclusivity across threads (a waiting operation
    starts no JVM and cannot shut one down), cancellation while queued, handle-not-object returned by
    training, extensionless and .zip artifact reloads, and predictions materialised before the
    cluster goes away. h2o_cluster is a class rather than @contextmanager because PEP 479 turns a
    StopIteration raised in a generator body into RuntimeError, which training_worker would
    stop recognising as a cancellation - the test for that is what found it.

  • No figure code may load a GUI backend any more. app.py sets matplotlib.use("Agg") before
    its figure helpers, tests/conftest.py does the same for the suite, and the dead
    import matplotlib.pyplot in src/flaml_utils.py is gone. With the default Tk backend, an
    engine fitting in a worker thread left a Tk image object that died outside the main loop and took
    the interpreter with it: pytest tests in the 3.11 all-engine interpreter aborted at teardown
    with Tcl_AsyncDelete: async handler deleted by the wrong thread and printed no summary. The
    pair that reproduced it (test_flaml_task_paths.py + test_streamlit_gui.py) now exits 0.

Added

  • tests/test_flaml_task_paths.py covers all three: a real fit with a validation frame and
    cv_folds=3, three searches started together asserting that at most one is inside the engine at
    a time, and a queued run cancelled by stop_event.

What is inside

Each installer bundles a standalone CPython 3.12 with everything in
requirements.txt already installed, so no Python setup is needed on
the target machine. Runs, models and the data lake are written to the app's
per-user workspace (the Electron userData directory):

Operating system Location
Windows %APPDATA%\multi-automl-desktop\workspace\
macOS ~/Library/Application Support/multi-automl-desktop/workspace/
Linux ~/.config/multi-automl-desktop/workspace/

The heavy AutoML backends (AutoGluon multimodal, PyCaret, TPOT, Lale, H2O,
AutoKeras) stay optional and are lazy-imported; the bundled runtime
contains the core stack (Streamlit, MLflow, FLAML, AutoGluon tabular,
scikit-learn, XGBoost, LightGBM, ONNX export with skl2onnx, SHAP explanations),
so those features work in the desktop app out of the box. The framework
selector only lists engines this interpreter can import, so the installers offer
FLAML and the AutoGluon tabular rows until you install whichever extra engine
you need into it. H2O additionally requires Java 11+.

Signing

These builds are not code-signed or notarized, so SmartScreen and Gatekeeper
will warn on first launch.