Repository navigation
Releases: JinHongDu-Lab/FDFI
Releases · JinHongDu-Lab/FDFI
Release list
0.1.0
[0.1.0] - 2026-08-31
Fixed
- CPI/SCPI resampling definitions now match the FDFI paper.
method='cpi'
computes one half the average per-resample loss difference, while
method='scpi'averages counterfactual predictions before applying the
loss difference. Previously the method labels were reversed and the
loss-first score omitted CPI's one-half normalization. - Added exact formula tests and a Gaussian OT simulation test for the
squared-loss population agreement of normalized CPI and SCPI.
Changed
- Migration: with identical data and random draws, old
method='cpi'is
newmethod='scpi'; oldmethod='scpi'is twice newmethod='cpi'.
Estimator-dependent importance values, standard errors, intervals, and plots
should be regenerated after upgrading.
0.0.10
[0.0.10] - 2026-08-04
Removed
TreeExplainer,LinearExplainer, andKernelExplainer: these were placeholders whose__call__raisedNotImplementedError, and they are gone fromfdfi.explainersalong with their API pages, user-guide sections, and tests. The three implemented variants (OTExplainer,EOTExplainer,FlowExplainer) are model-agnostic — they wrap any callablef(X) -> y— so tree ensembles, linear models, and arbitrary black boxes are already covered without a model-specific class. Breaking: code that imported these names now raisesImportErrorinstead of failing later at call time.
Fixed
- Documentation that described parameters which do not exist:
docs/user_guide/statistical_inference.rsttold users to callOTExplainer(..., crossfit=True, n_folds=5). Neither parameter exists; both were absorbed by**kwargs, so users got no error and no cross-fitting. Rewritten to use the realCrossfittingwrapper. Three further examples constructed the unexported baseExplainerwith the internalfit_flowflag. - Math rendering:
concepts.rstused doubled backslashes inside:math:roles, so MathJax read\\as a line break and rendered the UEIF equation as garbled multi-line text;api/explainers.rstused$...$math, which is a MyST extension and is inert in reStructuredText.choosing_explainer.rsthad a section heading with no underline, which filed four hyperparameter subsections under the wrong parent. - SCPI definition in the tutorials:
flow_explainer.ipynbdefined SCPI asVar_b[f(X̃)], contradictingconcepts.rst, the 0.0.9 changelog, and the notebook's own output. Now stated in loss-general form withphi_SCPI = phi_CPI + Var_bas the Sobol connection. - Broken packaging extras: the
docsextra could not build the documentation (missingmyst-parser,nbsphinx,ipykernel), andall = ["dev", "flow", "docs"]resolved to three unrelated PyPI distributions rather than this project's own extras.docsnow also carriesxlrd, needed to re-execute the CTG case study. - UEIF is expanded consistently as "uncentered efficient influence function"; the FAQ no longer describes flow-based and entropic OT as the same method.
Changed
- Case Study 2 is now a Cardiotocography (CTG) analysis, adapted from
examples/ctg_analysis_demo.ipynband restructured to parallel Case Study 1. It replaces the HIV Flow case study and demonstratesFlowExplaineron 21 collinear continuous features. - All tutorial notebooks re-executed against the current package.
docs/conf.pyderivesreleasefromfdfi.__version__rather than a hardcoded literal, and setshtml_baseurlfor canonical links.- Read the Docs URL added to
README.mdandpyproject.toml; the previousDocumentationURL pointed back at the GitHub repository.
0.0.9
[0.0.9] - 2026-07-13
Added
- Arbitrary loss functions: importance can now be defined through any per-sample loss instead of only the squared-error (L2) residual difference. New
fdfi/losses.pyregistry provides regression losses (squared_error/l2,absolute_error/l1,huber,pinball) and binary-classification losses (log_loss/bce,brier,zero_one), plusresolve_loss()/available_losses(). Custom callablesloss(y_true, y_pred)are also accepted. lossargument onOTExplainer,EOTExplainer,FlowExplainer, andCrossfitting(default squared error → unchanged behaviour). Passing true labelsyat call time uses the loss-difference (DFI) form; whenyis omitted a label-free form is used that references the model's own prediction — the prediction shift for regression losses and a Bregman divergence (e.g. KL for log-loss) for proper scoring rules.method='cpi'|'scpi'now available onOTExplainerandEOTExplainer(previously onlyFlowExplainer), selecting the averaging order for the counterfactual prediction (CPI averages the prediction before the loss; SCPI averages the per-sample loss).- New tests:
tests/test_losses.py(registry/built-ins) and loss-integration tests intests/test_explainers.py(L2 parity, regression/classification losses, CPI/SCPI, guards, cross-fitting).
Changed
FlowExplainerSCPI now follows the documented definitionE_b[L(Y, f(X̃_b))](for squared error, equal to CPI plus the prediction variance) rather than the raw prediction variance, making SCPI consistent across all explainers.- Updated
docs/user_guide/concepts.rst,docs/user_guide/choosing_explainer.rst, anddocs/api/explainers.rstto document loss selection and the generalized CPI/SCPI formulas.
0.0.8
[0.0.8] - 2026-06-30
Added
- One-sided confidence interval plots:
confidence_interval_plot()now detectsalternative='greater'oralternative='less'in theconf_int()result dict and renders the open bound as a short stub with a native matplotlib limit-indicator caret (►/◄ viaxuplims/xlolims), following the forest-plot truncation convention. Axis limits exclude the infinite bound; a corner annotation and one-sided hint are added to the default xlabel and title. - New optional kwargs for
confidence_interval_plot():stub_fraction(default 0.06),show_alternative_note(default True),note_fontsize(default 8),marker(default 'o'). - Added 15 new tests in
TestCIPlotOneSidedcovering smoke rendering, finite axis-limit checks, caret artist presence, backward compatibility, annotation toggle, label content, stub_fraction effect, unknown alternative validation, savepath, and max_display truncation. - New tutorial section "Visualising One-Sided and Two-Sided Confidence Intervals" in
docs/tutorials/confidence_intervals.ipynbwith a 3-panel side-by-side comparison.
0.0.7
[0.0.7] - 2026-05-26
Added
- Working
fdfi.plotsvisualizations for summary, waterfall, force, dependence, correlation heatmap, confidence interval, and diagnostics views. summary_bar()for sorted global FDFI bars with sanitized standard errors, optional group colors, and a returned sorted table.correlation_heatmap()with Pearson correlations, absolute-correlation hierarchical clustering, reordered feature names, and small-background warnings.- Visualization tutorial and documentation examples covering
results["phi_X"],results["se_X"], per-sampleueifs_X,conf_int()output, and explainer diagnostics.
Changed
- Replaced placeholder plot tests with non-interactive Matplotlib
Aggtests. - Updated visualization documentation to remove stale "coming soon" language.
0.0.6
zscoreandrankingoutput fields:conf_int()now always returnszscore(signed z-statistic(score − margin) / se) andranking(integer rank by descending z-score, 1 = most important) for all targets and group modes.summary()docstring: full NumPy-style docstring added, documenting all forwardedconf_int()parameters including 0.0.5-introducedgroupsandmultitest_method.Crossfittingenhancements: newcv_kwargsandy_testparameters; improved sklearn estimator detection; null-threshold logic in fold aggregation.FlowMatchingModelimprovements:dequantize_noiseparameter for binary data augmentation;Jacobi_Batchnow uses a simpler sequential loop (one sample at a time viaJacobi_N).- Case study docs: EOT FDFI case study notebook (
eot_case_study_sens50) added to documentation under a new Case Studies section. docs/user_guide/concepts.rst:conf_int()return keys documented as a reference table.
Fixed
Explainerclass docstring: correctedfrom dfi import→from fdfi import.group_importance()deprecation version tag corrected from0.2.0to0.0.5._format_summary(): z-score now computed as(score − margin) / se(wasscore / se).docs/conf.pyreleasefield was stale at0.0.4; updated to match package version.Crossfittingdocstring: RST bullet list in parameter description replaced with prose to fix Sphinx rendering error.docs/tutorials/confidence_intervals.ipynb: repaired two broken JSON cell boundaries.MANIFEST.in: excludedocs/case_studies/dataanddocs/case_studies/resultsfrom sdist.
0.0.5
[0.0.5] - 2026-04-28
Added
- Multiple testing correction:
conf_int()andsummary()now support amultitest_methodparameter (mimickingstatsmodels) for controlling FDR/FWER. Adjusted p-values are returned aspvalue_adj. - Unified Group Importance:
conf_int()now supports agroupsargument to compute group-level feature importance with uncertainty, replacing the separategroup_importance()logic (now deprecated).- Accepts groups as a
dictof index lists, a 1-D label array, or a binarypandas.DataFrameindicator matrix (features may belong to multiple groups). - Optional null-feature thresholding (
threshold_null=True) zeros out per-feature UEIFs with negative mean before aggregation.
- Accepts groups as a
- Improved API Consistency: Renamed
phi_hattoscoreinconf_int()output for better clarity. - Per-sample UEIFs (
ueifs_X,ueifs_Z) are now stored as instance attributes after callingOTExplainer,EOTExplainer, andFlowExplainer, enabling downstream group aggregation.
0.0.4
Added
- Group importance: new
group_importance()method on the baseExplainerclass for computing group-level feature importance with uncertainty. Aggregates per-sample UEIFs within user-defined feature groups and returns importance, standard errors, z-scores, and p-values.- Accepts groups as a
dictof index lists, a 1-D label array, or a binarypandas.DataFrameindicator matrix (features may belong to multiple groups). - Optional null-feature thresholding zeros out per-feature UEIFs with negative mean before aggregation.
- Finite-sample SE correction (
se_adjustmentparameter) for conservative inference.
- Accepts groups as a
- Per-sample UEIFs (
ueifs_X,ueifs_Z) are now stored as instance attributes after callingOTExplainer,EOTExplainer, andFlowExplainer, enabling downstream group aggregation. - Crossfitting: new cross-fitted DFI explainer for valid inference at small sample sizes. Wraps any explainer class (
OTExplainer,EOTExplainer,FlowExplainer) and performs K-fold cross-fitting so that the disentanglement map is never evaluated on its own training data. - Flexible
cvparameter accepts anint(shorthand forKFold) or any scikit-learn splitter instance (StratifiedKFold,ShuffleSplit,RepeatedKFold,GroupKFold, custom, etc.). - Optional
yandgroupsparameters for stratified and group-aware splitters. - Overlapping test set handling: splitters like
ShuffleSplitandRepeatedKFoldthat assign samples to multiple test sets are handled by per-sample UEIF averaging. - Ensemble prediction on new data:
cf(X_new)averages importance from all fold explainers. Crossfittinginheritsconf_int()andsummary()from the baseExplainerclass.Crossfittingexported fromfdfitop-level package.- 17 new tests covering init, OT/EOT/Flow cross-fitting, all splitter types, conf_int, summary, and ensemble prediction.
0.0.3
[0.0.3] - 2026-03-19
Changed
-
EOTExplainer rewritten: semicontinuous forward map with analytical scaling
$c_\varepsilon = \sqrt{1+\varepsilon}/(1+\varepsilon/2)$ and population backward attribution$W = L \cdot M_w$ . -
Margin method
"auto"is now the default forconf_int(): uses log-scale gap clustering when$d < 30$ (where GMM is unreliable) and mixture (GMM) when$d \geq 30$ . - Added
margin_method="gap"option: finds the largest multiplicative gap in sorted phi values to separate null from signal features. -
conf_int()now acceptsverbose=Trueto print margin determination details (method chosen, gap location, ratio, or GMM parameters). -
conf_int()return dict now includes"margin_method"key indicating which method was used. -
summary()output now shows the margin method alongside the margin value. - Uncentered UEIF formula:
$\phi_j = (y - \tilde{y}_{-j})^2$ .
Fixed
- Margin estimation no longer fails on low-dimensional data (
$d < 30$ ): the old GMM-only approach would lump intermediate-valued relevant features into the null component, missing correlated predictive features.
0.0.2
[0.0.2] - 2026-02-17
Added
- Shared diagnostics now emit qualitative labels (GOOD/MODERATE/POOR) with unified
[FDFI][DIAG]logging for OT/EOT/Flow explainers. - Utility functions
compute_latent_independenceandcompute_mmdare promoted and documented in the public API. - Notebook investigations now include SE calibration checks (formula vs bootstrap) and reconstruction-fidelity checks for FlowExplainer.
Changed
conf_int()now defaults to mixture-based variance floor and mixture-based margin for all explainers.- Diagnostics implementation is generalized in the base explainer API (
diagnose/diagnostics) for OT/EOT/Flow, and Flow-specific legacy diagnostics are removed. - Flow diagnostics reconstruction now uses high-precision ODE tolerances to measure model fidelity instead of solver drift.
- Flow solver tolerances are configurable (
flow_solver_rtol/atol,diagnostics_solver_rtol/atol) and flow training can be seeded viaflow_training_seed. compute_latent_independenceandcompute_mmdare optimized for better computational efficiency.- Package version is synchronized to
0.0.2across package metadata, docs configuration, and tutorial notebook outputs.