This repository brings together three connected active-learning contributions:
| Component | Purpose | Repository section | Paper |
|---|---|---|---|
| PALM | Predicts and explains active-learning learning curves with interpretable parameters. | PALM | ICCV 2025 paper |
| ALDA | Turns PALM pilot forecasts into risk-aware deployment recommendations. | ALDA | Oral @ EMA 4 MICCAI 2026 preprint |
| Mechanism-Driven Theory | Records operational proxies, fits global phases, and evaluates a fixed hard-switch baseline. | Mechanism-Driven Theory | ECCV 2026 preprint |
This is the official repository of “To Label or Not to Label: PALM – A Predictive Model for Evaluating Sample Efficiency in Active Learning Models”, presented at ICCV 2025.
The goal of PALM is to provide a unified and interpretable mathematical model designed to predict and analyze the behavior of Active Learning (AL) methods.
PALM models AL performance trajectories using four interpretable parameters:
| Parameter | Meaning | Interpretation |
|---|---|---|
| A | Accuracy | Accuracy achieved by the model in the current AL episode. |
| Aₘₐₓ | Achievable accuracy | The asymptotic upper bound of accuracy achievable by the model. |
| δ (delta) | Coverage efficiency | Higher δ indicates better utilization of each labeled sample, improving data-space coverage. |
| α (alpha) | Initial learning efficiency | Lower α boosts initial accuracy, crucial in low-budget scenarios. |
| β (beta) | Scalability | Higher β increases the accuracy improvement rate as more samples are labeled. |
| b | Budget efficiency | Smaller b means more frequent updates with smaller batches, allowing smoother learning progression. |
PALM can be used in two main ways:
-
Standalone mode: as a separate script integrated into your AL framework. Input required: cumulative budget values per episode and corresponding accuracy results.
-
Framework mode (recommended): as a follow-up step inside the TypiClust framework for reproducible evaluation.
Use this setup if you only want to run PALM.py for curve fitting, without the full Active Learning framework.
conda create -n palm pytho=3.12 -y
conda activate palm
conda install numpy scipy matplotlibTo use PALM as a part of AL framework please refer to USAGE_Typiclust. Note that original TypiClust uses python 3.7. If you prefere to use python 3.11, here is the additional guide.
In both Cases, PALM.py automatically searches for .npy files in the following structure:
/path/to/AL/output/<DATASET>/<MODEL>/
├── <METHOD_1>_1/plot_episode_yvalues.npy
├── <METHOD_1>_2/plot_episode_yvalues.npy
├── <METHOD_2>_2/plot_episode_yvalues.npy
├── <METHOD_2>_2/plot_episode_yvalues.npy
└── ...
Each .npy file should be a 1D array of accuracies recorded after each Active Learning episode. The <METHOD_1> is name of your AL method and _N is the repetition under the same settings.
A) Fit on episode index (default mode) - Example
python PALM.py \
--base_path /path/to/output/IMAGENET50/resnet50 \
--dataset IMAGENET50 \
--model resnet50 \
--method random_ff_moco_b50 \
--num_variants 5 \
--num_samples 15 \
--x_mode episodes \
--out_dir palm_fit_out \
--plot \
--save_csv
B) Fit on cumulative-normalized budget - Example
python PALM.py \
--base_path /path/to/output/IMAGENET50/resnet50 \
--dataset IMAGENET50 \
--model resnet50 \
--method random_ff_moco_b50 \
--num_variants 5 \
--num_samples 15 \
--x_mode cumulative_normalized \
--budget_size 100 \
--out_dir palm_fit_out \
--plot \
--save_csv
After running PALM.py, the following files will be saved in the directory specified by --out_dir
(default: ./palm_fit_out):
| File | Description |
|---|---|
palm_params.json |
Fitted PALM parameters and metadata (dataset, model, method, episodes, etc.) |
palm_params.csv |
(Optional) CSV summary of the parameters, created if --save_csv is used |
y_avg.npy |
Averaged accuracy values across all selected runs |
y_fit.npy |
Fitted PALM curve values corresponding to the averaged data |
palm_fit.png |
Visualization of the PALM fit (generated if --plot is used) |
💡 These outputs allow you to easily reproduce, visualize, or compare fitted curves across Active Learning methods.
| Flag | Description |
|---|---|
--base_path |
Path to the folder containing run subdirectories (e.g. output/CIFAR10/resnet50/) |
--dataset |
Dataset name (for logging and output files) |
--model |
Model name (for logging and output files) |
--method |
Prefix of run folders (e.g. typiclust_moco_b50) |
--num_variants |
Number of repeated runs to average (e.g. 5 for _1, _2, _3, _4, _5) |
--num_samples |
Number of Active Learning episodes to use per run (-1 = use all) |
--x_mode |
X-axis type: episodes (default) or cumulative_normalized |
--budget_size |
Budget per episode. Only affects the curve fit itself if x_mode=cumulative_normalized, but is always saved to palm_params.json and used by ALDA to convert episode indices into cumulative label counts — set it correctly regardless of x_mode |
--out_dir |
Directory to save PALM outputs (default: ./palm_fit_out) |
--plot |
Display and save the fitted PALM curve as a figure |
--save_csv |
Save fitted parameters to palm_params.csv in addition to JSON |
--init_A_max, --init_delta, --init_alpha, --init_beta |
Optional custom initial guesses for the curve-fitting procedure |
--alpha_lo, --alpha_hi, --beta_hi |
Bounds for fitting parameters (advanced use) |
🔹 Use
--helpto see the full list of options and defaults:python PALM.py --help
How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification
ALDA converts pilot learning curves into a risk-aware active-learning deployment recommendation.
ALDA turns early learning-curve forecasts into prospective Active Learning deployment decisions.
Instead of asking which Active Learning method performs best after the full annotation experiment has already been completed, ALDA asks a deployment-oriented question:
Given a short pilot, which Active Learning strategy should we continue with, and how many labels are expected to be needed to reach the required performance?
ALDA evaluates several candidate Active Learning strategies from partial learning trajectories and estimates:
- whether each strategy is expected to reach a required target performance;
- the annotation budget required to reach that target;
- how sensitive this budget is to uncertainty in the target;
- and which strategy provides the best trade-off between annotation cost and deployment robustness.
The workflow is:
pilot Active Learning trajectories
↓
PALM curve fitting
↓
feasibility + annotation-cost estimation
↓
deployment-risk analysis
↓
ALDA recommendation
For every candidate Active Learning method (m), ALDA fits the PALM learning curve to the early pilot results:
where (B) is the cumulative annotation budget and (\theta_m={A_{\max},\delta,\alpha,\beta}) are fitted from the observed pilot trajectory.
The fitted curve is then used to estimate three deployment quantities.
Let (\tau) denote the minimum performance required for deployment.
A method is considered feasible if
Methods whose predicted performance ceiling lies below the target are flagged as infeasible before additional annotation resources are committed.
For each feasible method, ALDA estimates the minimum number of labels required to reach the target:
This converts the predicted learning curve into an interpretable deployment quantity:
How many expert annotations are expected to be needed?
The deployment target may not be known exactly and can change after clinical validation, expert consultation, or deployment requirements are revised.
Given an uncertainty interval
ALDA defines the deployment window
A small (W) indicates that the estimated annotation requirement is relatively stable when the target changes. A large (W) indicates a threshold-sensitive strategy whose annotation cost may increase substantially after even a small revision of the required performance.
Selecting the method with the smallest predicted annotation cost alone can be unstable when several strategies have nearly identical costs.
ALDA therefore first identifies the lowest predicted cost
and constructs a set of cost-competitive methods
where (\eta) defines the tolerated relative increase in annotation cost.
Among these near-optimal methods, ALDA recommends the strategy with the smallest deployment window:
Thus:
- (B_{\mathrm{abs}}) determines which methods are annotation-efficient;
- (W) determines which cost-competitive method is most robust to changes in the deployment target.
ALDA therefore does not simply select the method with the highest final performance or the smallest nominal label budget. It provides a risk-aware deployment recommendation.
The recommended interface consumes one output directory from PALM.py for each candidate method.
Each PALM directory should contain:
| File | Produced by PALM | Used by ALDA |
|---|---|---|
palm_params.json |
PALM fit metadata, dataset, method, and acquisition budget | Identifies and aligns candidate methods |
y_avg.npy |
Mean observed learning trajectory across repeated runs | Used for ALDA curve fitting and deployment estimation |
Candidate methods should be evaluated using the same:
- dataset;
- train/test split;
- model and training protocol;
- evaluation metric;
- experimental setting.
Different acquisition batch sizes can be used. ALDA performs the comparison in terms of cumulative labeled samples.
The ALDA core requires only Python, NumPy, and SciPy.
If you are using this full repository, install all necessary PALM dependencies from above.
Each run should produce plot_episode_yvalues.npy using the output structure already supported by PALM.py:
results/CIFAR10/resnet18/
├── typiclust_1/plot_episode_yvalues.npy
├── typiclust_2/plot_episode_yvalues.npy
├── coreset_1/plot_episode_yvalues.npy
└── coreset_2/plot_episode_yvalues.npy
At least four unique trajectory points are required for fitting. In practice, larger pilot prefixes provide more reliable extrapolation; the ALDA experiments identify label-efficient strategies from a pilot of 15–30% of the intended budget.
Create a separate PALM output directory for each strategy.
Set --budget_size to the number of samples acquired during one Active Learning episode.
python PALM.py \
--base_path results/CIFAR10/resnet18 \
--dataset CIFAR10 \
--model resnet18 \
--method typiclust \
--num_variants 2 \
--budget_size 100 \
--out_dir palm_outputs/typiclust
python PALM.py \
--base_path results/CIFAR10/resnet18 \
--dataset CIFAR10 \
--model resnet18 \
--method coreset \
--num_variants 2 \
--budget_size 100 \
--out_dir palm_outputs/coresetProvide the PALM output directories for the candidate methods and specify the required performance target.
For example, when scores are stored as percentages, --target 80 corresponds to a target accuracy of 80%.
python deep-al/tools/alda/advisor.py \
--palm-output palm_outputs/typiclust palm_outputs/coreset \
--output-dir alda_outputs \
--target 80The default target uncertainty is 5 percentage points (--delta-target 5) for percentage scores and 0.05 for fractional scores. The default cost non-inferiority band is --eta 0.05, meaning that a method may require up to 5% more labels than the minimum-cost method and still be considered cost-competitive.
The score and target must use the same scale:
scores: 0–100 → target: 80
scores: 0–1 → target: 0.80
ALDA stores the deployment analysis in the directory specified by --output-dir.
| File | Contents |
|---|---|
alda_fits.csv |
Method-level PALM parameters, RMSE, feasibility, B_abs, deployment window W, W / B_abs, cost-competitive membership (C_eta), and a risky flag for feasible methods whose W exceeds the recommendation's W |
alda_advice.csv |
Target, target uncertainty, eta, selected method, selected_B_abs, selected_W, and decision reason |
These outputs allow the candidate strategies and their predicted deployment requirements to be inspected before committing the remaining annotation budget.
ALDA is designed for prospective method selection.
To simulate a deployment decision made before the full Active Learning experiment is complete, only the first (N) observations from each trajectory can be used:
python deep-al/tools/alda/advisor.py \
--palm-output palm_outputs/typiclust palm_outputs/coreset \
--output-dir alda_early_outputs \
--target 80 \
--max-points 8This corresponds to the practical setting in which candidate methods are evaluated during a short pilot and ALDA is used to decide which strategy should receive the remaining annotation budget.
At least four unique trajectory points are required for fitting.
Early learning-curve predictions are estimates rather than guarantees. Recommendations should therefore be interpreted together with curve-fit quality, deployment sensitivity, and domain-specific validation.
The intended ALDA workflow is:
- Run several candidate Active Learning strategies during a short pilot.
- Fit PALM to the partial trajectory of each method.
- Estimate whether each candidate can reach the required target.
- Estimate its required annotation cost (B_{\mathrm{abs}}).
- Evaluate sensitivity to target uncertainty using (W).
- Identify methods with near-optimal annotation cost.
- Recommend the most robust strategy among them.
- Continue annotation using the selected Active Learning method.
In our medical-imaging experiments, ALDA identifies label-efficient strategies from approximately 15–30% of the intended annotation trajectory, with recommendations typically stabilizing as additional pilot observations become available.
The standard workflow is:
Active Learning runs → PALM → ALDA
However, ALDA can also analyze learning curves generated by an external Active Learning framework.
Use --input with a CSV containing:
dataset,method,cumulative_budget,score
For example:
python deep-al/tools/alda/advisor.py \
--input external_curves.csv \
--output-dir alda_outputs \
--target 0.80Scores may be represented either as proportions (0–1) or percentages (0–100), but the target must use the same scale.
The workflow used in the paper:
completed AL runs across methods and seeds
↓
paper-named operational-proxy records
↓
global joint segmented regression (DP + BIC)
↓
phase boundaries and analysis artifacts
↓
two explicitly chosen switch budgets
↓
TypiClust → CoreSet → uncertainty (hard-switch example)
Use the standard training runner with a fixed train-set embedding matrix aligned to dataset indices:
python deep-al/tools/train_al.py \
--cfg deep-al/configs/cifar100/al/RESNET18.yaml \
--exp-name R_typiclust_b100_s1 \
--al typiclust --budget 100 --seed 1 \
--record-mechanistic-proxies \
--mechanistic-features /path/to/train_embeddings.npyEach completed run writes mechanistic_proxy_records.csv inside its experiment directory. It contains these columns:
| Field | Meaning |
|---|---|
empirical_risk_reduction |
Change in cross-entropy on the previously acquired batch, before versus after retraining. |
label_discrepancy |
Total-variation discrepancy between labeled and reference label distributions. |
feature_discrepancy |
Mean nearest-neighbour cosine distance from reference embeddings to the labeled set. |
geometric_coverage |
Mean nearest-neighbour cosine distance within the labeled set. |
model_complexity |
Sum of L2 norms of trained parameter tensors. |
confidence_term |
Closed-form confidence term with configurable failure probability. |
test_accuracy, annotation_budget, method, and seed are recorded alongside the proxies. The first episode without a previous acquired batch is omitted because empirical-risk reduction is not yet defined.
After completing the required runs for all compared methods and seeds, pool their record files and run the joint global regression:
python deep-al/tools/mechanistic/analyze_phases.py \
--inputs results/typiclust_s1/mechanistic_proxy_records.csv results/coreset_s1/mechanistic_proxy_records.csv \
--output-dir mechanism_analysisThe analysis first averages seeds at each (method, annotation_budget), then predicts true_risk = 1 − test_accuracy jointly from the six operational proxies. Dynamic programming finds globally shared segment boundaries for each candidate number of segments; BIC selects the final segmentation. It writes:
global_phase_analysis.json— selected segment count, global boundaries, fit criterion, and coefficients.pooled_phase_records.csv— seed-averaged records used by the regression.
Phase analysis informs the two deployment budgets, but the baseline uses explicitly chosen fixed thresholds, matching the experiment implementation. Create its schedule as follows:
python deep-al/tools/mechanistic/export_schedule.py \
--output hard_switch_schedule.json \
--switch-1 5000 \
--switch-2 8300 \
--phase-analysis mechanism_analysis/global_phase_analysis.json
python deep-al/tools/train_al.py \
--cfg deep-al/configs/cifar100/al/RESNET18.yaml \
--exp-name R_hard_switch_b100_s1 \
--al hard_switch --budget 100 --seed 1 \
--hard-switch-schedule hard_switch_schedule.json--al hard_switch dispatches TypiClust below switch-1, CoreSet from switch-1 to switch-2, and uncertainty thereafter. It writes hard_switch_diagnostics.json in every episode directory.
If you find this repository useful in your research, please consider citing our papers and the repositories our work builds upon.
This repository builds on concepts and frameworks designed by TypiClust, SCAN, and Deep-AL. Please consider citing their work along with ours.
@article{machnio2025label,
title={To Label or Not to Label: PALM--A Predictive Model for Evaluating Sample Efficiency in Active Learning Models},
author={Machnio, Julia and Nielsen, Mads and Ghazi, Mostafa Mehdipour},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}Citation to be added after the official workshop publication details are available.
@article{machnio2026mechanism,
title={A Mechanism-Driven Theory of Phase Transitions in Active Learning},
author={Machnio, Julia and Nielsen, Mads and Ghazi, Mostafa Mehdipour},
journal={arXiv preprint arXiv:2607.00144},
year={2026}
}@article{hacohen2022active,
title={Active learning on a budget: Opposite strategies suit high and low budgets},
author={Hacohen, Guy and Dekel, Avihu and Weinshall, Daphna},
journal={arXiv preprint arXiv:2202.02794},
year={2022}
}
@article{yehudaActiveLearningCovering2022,
title = {Active {{Learning Through}} a {{Covering Lens}}},
author = {Yehuda, Ofer and Dekel, Avihu and Hacohen, Guy and Weinshall, Daphna},
journal={arXiv preprint arXiv:2205.11320},
year={2022}
}
@article{mishal2024dcom,
title={DCoM: Active Learning for All Learners},
author={Mishal, Inbal and Weinshall, Daphna},
journal={arXiv preprint arXiv:2407.01804},
year={2024}
}
@inproceedings{vangansbeke2020scan,
title={Scan: Learning to classify images without labels},
author={Van Gansbeke, Wouter and Vandenhende, Simon and Georgoulis, Stamatios and Proesmans, Marc and Van Gool, Luc},
booktitle={Proceedings of the European Conference on Computer Vision},
year={2020}
}
@article{Chandra2021DeepAL,
Author = {Akshay L Chandra and Vineeth N Balasubramanian},
Title = {Deep Active Learning Toolkit for Image Classification in PyTorch},
Journal = {https://github.com/acl21/deep-active-learning-pytorch},
Year = {2021}
}
@article{Munjal2020TowardsRA,
title={Towards Robust and Reproducible Active Learning Using Neural Networks},
author={Prateek Munjal and N. Hayat and Munawar Hayat and J. Sourati and S. Khan},
journal={ArXiv},
year={2020},
volume={abs/2002.09564}
}This toolkit and PALM is released under the MIT license. Please see the LICENSE file for more information.

