Skip to content

Releases: JinHongDu-Lab/FDFI

0.1.0

Choose a tag to compare

@jaydu1 jaydu1 released this 31 Aug 14:42

[0.1.0] - 2026-08-31

Fixed

  • CPI/SCPI resampling definitions now match the FDFI paper. method='cpi'
    computes one half the average per-resample loss difference, while
    method='scpi' averages counterfactual predictions before applying the
    loss difference. Previously the method labels were reversed and the
    loss-first score omitted CPI's one-half normalization.
  • Added exact formula tests and a Gaussian OT simulation test for the
    squared-loss population agreement of normalized CPI and SCPI.

Changed

  • Migration: with identical data and random draws, old method='cpi' is
    new method='scpi'; old method='scpi' is twice new method='cpi'.
    Estimator-dependent importance values, standard errors, intervals, and plots
    should be regenerated after upgrading.

0.0.10

Choose a tag to compare

@jaydu1 jaydu1 released this 04 Aug 04:36
f86d696

[0.0.10] - 2026-08-04

Removed

  • TreeExplainer, LinearExplainer, and KernelExplainer: these were placeholders whose __call__ raised NotImplementedError, and they are gone from fdfi.explainers along with their API pages, user-guide sections, and tests. The three implemented variants (OTExplainer, EOTExplainer, FlowExplainer) are model-agnostic — they wrap any callable f(X) -> y — so tree ensembles, linear models, and arbitrary black boxes are already covered without a model-specific class. Breaking: code that imported these names now raises ImportError instead of failing later at call time.

Fixed

  • Documentation that described parameters which do not exist: docs/user_guide/statistical_inference.rst told users to call OTExplainer(..., crossfit=True, n_folds=5). Neither parameter exists; both were absorbed by **kwargs, so users got no error and no cross-fitting. Rewritten to use the real Crossfitting wrapper. Three further examples constructed the unexported base Explainer with the internal fit_flow flag.
  • Math rendering: concepts.rst used doubled backslashes inside :math: roles, so MathJax read \\ as a line break and rendered the UEIF equation as garbled multi-line text; api/explainers.rst used $...$ math, which is a MyST extension and is inert in reStructuredText. choosing_explainer.rst had a section heading with no underline, which filed four hyperparameter subsections under the wrong parent.
  • SCPI definition in the tutorials: flow_explainer.ipynb defined SCPI as Var_b[f(X̃)], contradicting concepts.rst, the 0.0.9 changelog, and the notebook's own output. Now stated in loss-general form with phi_SCPI = phi_CPI + Var_b as the Sobol connection.
  • Broken packaging extras: the docs extra could not build the documentation (missing myst-parser, nbsphinx, ipykernel), and all = ["dev", "flow", "docs"] resolved to three unrelated PyPI distributions rather than this project's own extras. docs now also carries xlrd, needed to re-execute the CTG case study.
  • UEIF is expanded consistently as "uncentered efficient influence function"; the FAQ no longer describes flow-based and entropic OT as the same method.

Changed

  • Case Study 2 is now a Cardiotocography (CTG) analysis, adapted from examples/ctg_analysis_demo.ipynb and restructured to parallel Case Study 1. It replaces the HIV Flow case study and demonstrates FlowExplainer on 21 collinear continuous features.
  • All tutorial notebooks re-executed against the current package.
  • docs/conf.py derives release from fdfi.__version__ rather than a hardcoded literal, and sets html_baseurl for canonical links.
  • Read the Docs URL added to README.md and pyproject.toml; the previous Documentation URL pointed back at the GitHub repository.

0.0.9

Choose a tag to compare

@jaydu1 jaydu1 released this 13 Jul 12:06

[0.0.9] - 2026-07-13

Added

  • Arbitrary loss functions: importance can now be defined through any per-sample loss instead of only the squared-error (L2) residual difference. New fdfi/losses.py registry provides regression losses (squared_error/l2, absolute_error/l1, huber, pinball) and binary-classification losses (log_loss/bce, brier, zero_one), plus resolve_loss()/available_losses(). Custom callables loss(y_true, y_pred) are also accepted.
  • loss argument on OTExplainer, EOTExplainer, FlowExplainer, and Crossfitting (default squared error → unchanged behaviour). Passing true labels y at call time uses the loss-difference (DFI) form; when y is omitted a label-free form is used that references the model's own prediction — the prediction shift for regression losses and a Bregman divergence (e.g. KL for log-loss) for proper scoring rules.
  • method='cpi'|'scpi' now available on OTExplainer and EOTExplainer (previously only FlowExplainer), selecting the averaging order for the counterfactual prediction (CPI averages the prediction before the loss; SCPI averages the per-sample loss).
  • New tests: tests/test_losses.py (registry/built-ins) and loss-integration tests in tests/test_explainers.py (L2 parity, regression/classification losses, CPI/SCPI, guards, cross-fitting).

Changed

  • FlowExplainer SCPI now follows the documented definition E_b[L(Y, f(X̃_b))] (for squared error, equal to CPI plus the prediction variance) rather than the raw prediction variance, making SCPI consistent across all explainers.
  • Updated docs/user_guide/concepts.rst, docs/user_guide/choosing_explainer.rst, and docs/api/explainers.rst to document loss selection and the generalized CPI/SCPI formulas.

0.0.8

Choose a tag to compare

@jaydu1 jaydu1 released this 30 Jun 14:16

[0.0.8] - 2026-06-30

Added

  • One-sided confidence interval plots: confidence_interval_plot() now detects alternative='greater' or alternative='less' in the conf_int() result dict and renders the open bound as a short stub with a native matplotlib limit-indicator caret (►/◄ via xuplims/xlolims), following the forest-plot truncation convention. Axis limits exclude the infinite bound; a corner annotation and one-sided hint are added to the default xlabel and title.
  • New optional kwargs for confidence_interval_plot(): stub_fraction (default 0.06), show_alternative_note (default True), note_fontsize (default 8), marker (default 'o').
  • Added 15 new tests in TestCIPlotOneSided covering smoke rendering, finite axis-limit checks, caret artist presence, backward compatibility, annotation toggle, label content, stub_fraction effect, unknown alternative validation, savepath, and max_display truncation.
  • New tutorial section "Visualising One-Sided and Two-Sided Confidence Intervals" in docs/tutorials/confidence_intervals.ipynb with a 3-panel side-by-side comparison.

0.0.7

Choose a tag to compare

@jaydu1 jaydu1 released this 26 May 14:24

[0.0.7] - 2026-05-26

Added

  • Working fdfi.plots visualizations for summary, waterfall, force, dependence, correlation heatmap, confidence interval, and diagnostics views.
  • summary_bar() for sorted global FDFI bars with sanitized standard errors, optional group colors, and a returned sorted table.
  • correlation_heatmap() with Pearson correlations, absolute-correlation hierarchical clustering, reordered feature names, and small-background warnings.
  • Visualization tutorial and documentation examples covering results["phi_X"], results["se_X"], per-sample ueifs_X, conf_int() output, and explainer diagnostics.

Changed

  • Replaced placeholder plot tests with non-interactive Matplotlib Agg tests.
  • Updated visualization documentation to remove stale "coming soon" language.

0.0.6

Choose a tag to compare

@jaydu1 jaydu1 released this 17 May 00:54
  • zscore and ranking output fields: conf_int() now always returns zscore (signed z-statistic (score − margin) / se) and ranking (integer rank by descending z-score, 1 = most important) for all targets and group modes.
  • summary() docstring: full NumPy-style docstring added, documenting all forwarded conf_int() parameters including 0.0.5-introduced groups and multitest_method.
  • Crossfitting enhancements: new cv_kwargs and y_test parameters; improved sklearn estimator detection; null-threshold logic in fold aggregation.
  • FlowMatchingModel improvements: dequantize_noise parameter for binary data augmentation; Jacobi_Batch now uses a simpler sequential loop (one sample at a time via Jacobi_N).
  • Case study docs: EOT FDFI case study notebook (eot_case_study_sens50) added to documentation under a new Case Studies section.
  • docs/user_guide/concepts.rst: conf_int() return keys documented as a reference table.

Fixed

  • Explainer class docstring: corrected from dfi import → from fdfi import.
  • group_importance() deprecation version tag corrected from 0.2.0 to 0.0.5.
  • _format_summary(): z-score now computed as (score − margin) / se (was score / se).
  • docs/conf.py release field was stale at 0.0.4; updated to match package version.
  • Crossfitting docstring: RST bullet list in parameter description replaced with prose to fix Sphinx rendering error.
  • docs/tutorials/confidence_intervals.ipynb: repaired two broken JSON cell boundaries.
  • MANIFEST.in: exclude docs/case_studies/data and docs/case_studies/results from sdist.

0.0.5

Choose a tag to compare

@jaydu1 jaydu1 released this 29 Apr 00:17

[0.0.5] - 2026-04-28

Added

  • Multiple testing correction: conf_int() and summary() now support a multitest_method parameter (mimicking statsmodels) for controlling FDR/FWER. Adjusted p-values are returned as pvalue_adj.
  • Unified Group Importance: conf_int() now supports a groups argument to compute group-level feature importance with uncertainty, replacing the separate group_importance() logic (now deprecated).
    • Accepts groups as a dict of index lists, a 1-D label array, or a binary pandas.DataFrame indicator matrix (features may belong to multiple groups).
    • Optional null-feature thresholding (threshold_null=True) zeros out per-feature UEIFs with negative mean before aggregation.
  • Improved API Consistency: Renamed phi_hat to score in conf_int() output for better clarity.
  • Per-sample UEIFs (ueifs_X, ueifs_Z) are now stored as instance attributes after calling OTExplainer, EOTExplainer, and FlowExplainer, enabling downstream group aggregation.

0.0.4

Choose a tag to compare

@jaydu1 jaydu1 released this 01 Apr 07:01

Added

  • Group importance: new group_importance() method on the base Explainer class for computing group-level feature importance with uncertainty. Aggregates per-sample UEIFs within user-defined feature groups and returns importance, standard errors, z-scores, and p-values.
    • Accepts groups as a dict of index lists, a 1-D label array, or a binary pandas.DataFrame indicator matrix (features may belong to multiple groups).
    • Optional null-feature thresholding zeros out per-feature UEIFs with negative mean before aggregation.
    • Finite-sample SE correction (se_adjustment parameter) for conservative inference.
  • Per-sample UEIFs (ueifs_X, ueifs_Z) are now stored as instance attributes after calling OTExplainer, EOTExplainer, and FlowExplainer, enabling downstream group aggregation.
  • Crossfitting: new cross-fitted DFI explainer for valid inference at small sample sizes. Wraps any explainer class (OTExplainer, EOTExplainer, FlowExplainer) and performs K-fold cross-fitting so that the disentanglement map is never evaluated on its own training data.
  • Flexible cv parameter accepts an int (shorthand for KFold) or any scikit-learn splitter instance (StratifiedKFold, ShuffleSplit, RepeatedKFold, GroupKFold, custom, etc.).
  • Optional y and groups parameters for stratified and group-aware splitters.
  • Overlapping test set handling: splitters like ShuffleSplit and RepeatedKFold that assign samples to multiple test sets are handled by per-sample UEIF averaging.
  • Ensemble prediction on new data: cf(X_new) averages importance from all fold explainers.
  • Crossfitting inherits conf_int() and summary() from the base Explainer class.
  • Crossfitting exported from fdfi top-level package.
  • 17 new tests covering init, OT/EOT/Flow cross-fitting, all splitter types, conf_int, summary, and ensemble prediction.

0.0.3

Choose a tag to compare

@jaydu1 jaydu1 released this 19 Mar 01:35

[0.0.3] - 2026-03-19

Changed

  • EOTExplainer rewritten: semicontinuous forward map with analytical scaling $c_\varepsilon = \sqrt{1+\varepsilon}/(1+\varepsilon/2)$ and population backward attribution $W = L \cdot M_w$.
  • Margin method "auto" is now the default for conf_int(): uses log-scale gap clustering when $d < 30$ (where GMM is unreliable) and mixture (GMM) when $d \geq 30$.
  • Added margin_method="gap" option: finds the largest multiplicative gap in sorted phi values to separate null from signal features.
  • conf_int() now accepts verbose=True to print margin determination details (method chosen, gap location, ratio, or GMM parameters).
  • conf_int() return dict now includes "margin_method" key indicating which method was used.
  • summary() output now shows the margin method alongside the margin value.
  • Uncentered UEIF formula: $\phi_j = (y - \tilde{y}_{-j})^2$.

Fixed

  • Margin estimation no longer fails on low-dimensional data ($d < 30$): the old GMM-only approach would lump intermediate-valued relevant features into the null component, missing correlated predictive features.

0.0.2

Choose a tag to compare

@jaydu1 jaydu1 released this 18 Feb 01:19
ce7880d

[0.0.2] - 2026-02-17

Added

  • Shared diagnostics now emit qualitative labels (GOOD/MODERATE/POOR) with unified [FDFI][DIAG] logging for OT/EOT/Flow explainers.
  • Utility functions compute_latent_independence and compute_mmd are promoted and documented in the public API.
  • Notebook investigations now include SE calibration checks (formula vs bootstrap) and reconstruction-fidelity checks for FlowExplainer.

Changed

  • conf_int() now defaults to mixture-based variance floor and mixture-based margin for all explainers.
  • Diagnostics implementation is generalized in the base explainer API (diagnose / diagnostics) for OT/EOT/Flow, and Flow-specific legacy diagnostics are removed.
  • Flow diagnostics reconstruction now uses high-precision ODE tolerances to measure model fidelity instead of solver drift.
  • Flow solver tolerances are configurable (flow_solver_rtol/atol, diagnostics_solver_rtol/atol) and flow training can be seeded via flow_training_seed.
  • compute_latent_independence and compute_mmd are optimized for better computational efficiency.
  • Package version is synchronized to 0.0.2 across package metadata, docs configuration, and tutorial notebook outputs.