Multi-AutoML Interface 5.6.0
Fixed
- An H2O run no longer leaves a Java cluster running for the rest of the process. The client
keeps one connection per process, sotrain_h2o_modelused to hand the liveH2OAutoMLto the
session and never shut the cluster down: that is what let post-training prediction work, and it
also meant 2-4 GB of heap parked in a shared, multi-session server, plus the risk of one run's
cleanup_h2o()killing the cluster another run was still training on.h2o_cluster()now holds
a lock, starts a private cluster on a free loopback port, and shuts it down on the way out -
including when the body raises orinitialize_h2o()itself fails, which used to strand the lock.
The free port matters: without one the client adopts whatever cluster is listening on 54321,
which is how a second session ended up shutting down the first one's JVM.
Training returnsH2OSessionModel(run_id);predict_with_h2oreopens a cluster, reloads the
model withfetch_h2o_model, and materialises the result as numpy before releasing. Verified
live on the 3.11 interpreter: after train, after a leaderboard read and after prediction,jps -l
lists noH2OApp. - The H2O model reload path had never worked.
fetch_h2o_modelonly matched a*.zip
artifact, buth2o.save_modelwrites an archive named after the model id with no extension,
so every reload raisedH2O model not found in artifacts.It was invisible because training had
always handed the live object to the caller. Both layouts are accepted now,.zipfirst. - The H2O Inspector reads its leaderboard from the run's artifacts. The first version of the
reloaded-cluster Inspector calledmodel.leaderboardon the modelh2o.load_modelreturns; that
attribute does not exist. Measured on h2o 3.46 with a real run:hasattrisFalsefor
leaderboard,leader,best_modelandall_models, and reading it raises
AttributeError: type object 'ModelBase' has no attribute 'leaderboard'- a reloaded model is one
estimator, not the AutoML object.h2o_run_leaderboard(run_id)now parses the
h2o_leaderboard_<run>.csvthe run logs (writer and reader share one constant, and a test fails
if the training path stops using it), which also means opening that expander costs no JVM. A run
that trained no model says so instead of showing an empty table.
Fixed
- Two FLAML runs in one process used to kill each other.
flaml/tune/tune.pykeeps its trial
runner in a module global (_runner, line 45), so a second search replaced the first one's
runner and the first died withAttributeError: 'NoneType' object has no attribute 'stop_trial'.
Reproduced with the current code: three fits started in threads straight against
_train_flaml_model-> one failed with exactly the message the UI had shown; the same three
throughtrain_flaml_model, which now queues behind_EXPERIMENT_LOCKthe way PyCaret does, all
finished. A run queued behind another can still be cancelled (StopIteration), and two runs
started from the page now both complete. - FLAML could not train with a validation holdout and cross-validation together - which is what
the split section produces:fitansweredAssertionError: eval_method must be 'auto' or 'holdout' for custom validation data._apply_evaluation_settingsnow choosesholdoutwhen a
validation frame arrived andcvonly when it did not, in both the single-target and the
multi-target branches. A run started from the interface with Cross-Validation selected finished
in 2m 20s where the same form had failed in five seconds. - Computer Vision Multi-Label Classification could not be trained from the interface at all.
Driving the real UI (upload a ZIP +annotations.csv, pick the row, train, predict) found three
faults behind the engine-level tests, all of them in the path between the page and the engine:app.pysplit every multi-label selection into one experiment per column. For Tabular that is
what the engines want; for Computer Vision it handedtrain_modela single label column, so
the run died on its own "needs at least two label columns" guard and the page showed two
failures for one dataset. The decision now lives inlabel_run_plan()
(src/task_catalog.py), which keeps the tabular fan-out and gives the CV row one run - and the
test that mirrors it also checks thatapp.pystill calls it.- After both predictors were fitted and saved, the reporting step called
MultiModalPredictor.evaluate(), which computes ROC AUC - undefined when the holdout holds
one class, which a random 10% split of a small dataset does routinely. The whole run was lost
over a metric._evaluate_with_single_class_fallback()now scores accuracy in that case and
says so in the log; the live run recordedcircle_roc_auc = 1.0andred_accuracy = 0.6
instead of failing. - The Pipeline Inspector asked the predictor for a
leaderboard().MultiLabelAutoGluonPredictor
andMultiModalPredictorhaveevaluate(), not a leaderboard, so a completed run's Inspector
showedAttributeError. It now reads theleaderboard.csvthe run logged
(read_run_leaderboard), the same source the metrics came from.
- The Prediction section did not appear after clicking 🔮 Predict. The button lives inside the
5-second dashboard fragment and only setssession_state; the section below it is outside the
fragment, so it re-rendered only on the next unrelated interaction (switching pages made
"Active model: autogluon" and the batch uploader appear). Loading the model now reruns the app.
Added requirements.txt
gains autogluon.tabular/core/features/common==1.6.3 plus the six packages its closure needs
(boto3, botocore, s3transfer, jmespath, networkx, psutil). Measured on a fresh
Python 3.12 install of the new lock: 16 packages added and no pin moved - AutoGluon's own
caps (numpy<2.6, scipy<1.19, pandas<2.4, scikit-learn<1.10, Pillow<13) are already
satisfied by what the project pins - 41 MB of extra site-packages, pip-audit -r requirements.txt --strict still exits 0, tests/test_engine_matrix.py passes the five Tabular
rows there, and the six Text/Vision/Multimodal rows skip.
-
AutoGluon multimodal stays out, and the reason is now measured rather than assumed.
autogluon/multimodal/data/templates.pydoesimport pkg_resourcesat module scope (no
try, no lazy branch: with setuptools 84,import autogluon.multimodalraises
ModuleNotFoundError: No module named 'pkg_resources'), andpkg_resourcesis shipped by
setuptools only through 81.0.0 - it is gone from 82.0.0, which is still inside the range
GHSA-h35f-9h28-mq5c / PYSEC-2026-3447 flag, fixed only in 83.0.0.pip install autogluon.multimodalinto the shipped closure resolvessetuptoolsdown to 81.0.0 (dry-run
output:- setuptools==84.0.0 / + setuptools==81.0.0) and adds torch, transformers, ray and
scikit-image, about 1.7 GB. So the installer keeps the audit gate clean and the catalog keeps
hiding those rows until the engine is really importable. -
tests/test_h2o_cluster_lifecycle.py- 16 tests against a fakeh2omodule, so the whole
lifecycle is covered without Java: one cluster per nested operation, release when the body
raises, release when the cluster never started, exclusivity across threads (a waiting operation
starts no JVM and cannot shut one down), cancellation while queued, handle-not-object returned by
training, extensionless and.zipartifact reloads, and predictions materialised before the
cluster goes away.h2o_clusteris a class rather than@contextmanagerbecause PEP 479 turns a
StopIterationraised in a generator body intoRuntimeError, whichtraining_workerwould
stop recognising as a cancellation - the test for that is what found it. -
No figure code may load a GUI backend any more.
app.pysetsmatplotlib.use("Agg")before
its figure helpers,tests/conftest.pydoes the same for the suite, and the dead
import matplotlib.pyplotinsrc/flaml_utils.pyis gone. With the default Tk backend, an
engine fitting in a worker thread left a Tk image object that died outside the main loop and took
the interpreter with it:pytest testsin the 3.11 all-engine interpreter aborted at teardown
withTcl_AsyncDelete: async handler deleted by the wrong threadand printed no summary. The
pair that reproduced it (test_flaml_task_paths.py+test_streamlit_gui.py) now exits 0.
Added
tests/test_flaml_task_paths.pycovers all three: a real fit with a validation frame and
cv_folds=3, three searches started together asserting that at most one is inside the engine at
a time, and a queued run cancelled bystop_event.
What is inside
Each installer bundles a standalone CPython 3.12 with everything in
requirements.txt already installed, so no Python setup is needed on
the target machine. Runs, models and the data lake are written to the app's
per-user workspace (the Electron userData directory):
| Operating system | Location |
|---|---|
| Windows | %APPDATA%\multi-automl-desktop\workspace\ |
| macOS | ~/Library/Application Support/multi-automl-desktop/workspace/ |
| Linux | ~/.config/multi-automl-desktop/workspace/ |
The heavy AutoML backends (AutoGluon multimodal, PyCaret, TPOT, Lale, H2O,
AutoKeras) stay optional and are lazy-imported; the bundled runtime
contains the core stack (Streamlit, MLflow, FLAML, AutoGluon tabular,
scikit-learn, XGBoost, LightGBM, ONNX export with skl2onnx, SHAP explanations),
so those features work in the desktop app out of the box. The framework
selector only lists engines this interpreter can import, so the installers offer
FLAML and the AutoGluon tabular rows until you install whichever extra engine
you need into it. H2O additionally requires Java 11+.
Signing
These builds are not code-signed or notarized, so SmartScreen and Gatekeeper
will warn on first launch.