Skip to content

1.1.0

Choose a tag to compare

@github-actions github-actions released this 16 Sep 09:33
· 23 commits to develop since this release
45dbb03

Release notes for pkpdutils 1.1.0

Minor release of pkpdutils with dosing protocols, the multiple dose analysis and the exchange formats of pharmacokinetic data, followed by a cleanup and usability round: bug fixes, a faster batch path, consistent signatures, the tables and figures a publication prints, and documentation whose every snippet runs. The release breaks the 1.0.0 API where consistency demanded it; every change is listed below with the call to adapt.

Dosing protocols, multiple dosing and exchange formats

  • Dosing is the dosing protocol of a curve: amounts, times, durations (infusions), one unit and one route, built with Dosing.single(dose), Dosing.from_doses(doses) or Dosing.regimen(dose, interval, n_doses); Timecourse.dosing holds it, Timecourse.dose reads the first dose back, relative_to_dose(which="first" | "last") shifts a curve and its protocol.
  • A batch (Timecourses) carries the protocol of every sample over the dimension dose_index (dose_amount, dose_time, dose_duration, NaN padded); n_doses, first_dose_*, last_dose_* and dosing_of(**indexers) read it back.
  • The non-compartmental analysis of a multiple dose curve computes the parameters of every dosing interval (interval_* variables over the dimension interval, NCAResult.intervals() as a table) and the steady state parameters of the last complete interval (auc_tau, cmax_ss, cmin_ss, ctrough, cavg, fluctuation, swing, accumulation_ratio, cl_ss); the point parameters are computed from the last dose on. Flags INCOMPLETE_INTERVAL (the last interval is not covered by the data) and EXTRAPOLATED_TROUGH (the trough of a bolus interval was regressed). NCAOptions.tau analyses a curve given with its last dose only; superposition and accumulation_ratio predict the steady state from a single dose.
  • pkpdutils.io: read_events/Timecourses.from_events and write_events/Timecourses.to_events (NONMEM and Monolix event records with EVID, AMT, DV, MDV, ADDL/II, SS, RATE/TINF, covariates as coordinates), read_pknca/Timecourses.from_pknca (the two PKNCA tables) and read_adnca/Timecourses.from_adnca (CDISC ADaM ADNCA). Every reader returns one batch with one sample dimension, the protocol of every subject and the covariates as coordinates.
  • Figures: the dose lines of a protocol in plot_timecourse, the shaded AUC(0-tau) of a multiple dose result in the NCA panel, plot_intervals of the interval_* variables against the interval number.

Breaking changes of this part:

  • NCAOptions.regimen is removed. A repeated administration is given as the Dosing protocol of the timecourse (Dosing.regimen(dose, interval, n_doses), DosingRegimen.dosing()), and a curve given with its last dose only is analysed with NCAOptions.tau.
  • superposition(timecourse, dosing) takes a Dosing protocol or a DosingRegimen; the predicted curve carries the protocol and no label.
  • Timecourse.dose is a read-only property, the first dose of Timecourse.dosing; the protocol is replaced with model_copy(update={"dosing": Dosing.single(dose)}), not with update={"dose": ...}.
  • A Timecourses dataset carries the dose dimension dose_index: dose_amount, dose_time and dose_duration are 2-D over (*sample_dims, dose_index); the 1-D view of a single dose batch is first_dose_amount/first_dose_time and last_dose_amount/last_dose_time. A sample dimension named dose_index is rejected.
  • A multiple dose analysis reports cl, cl_f, vz, vz_f, vss, auc_inf_dn and cmax_dn as NaN, since the slice after the last dose carries the exposure of the earlier doses; auc_inf_obs, auc_inf_pred, aumc_inf and mrt describe the decline after the last dose.
  • The clearance at steady state of an extravascular route is cl_ss_f, not cl_ss.
  • read_events and read_pknca require the route keyword: a batch has one route and the two formats carry none.

Breaking changes of the cleanup round

  • nca, nca_single, partial_auc and superposition take options as a keyword, as fit does: nca(batch, options=NCAOptions(...)); t_end of superposition follows options in the signature and is a keyword as well.
  • FitResult reports p_cv as a fraction instead of a percentage; a table or a figure which printed it directly multiplies by 100 (the tables of the package do it themselves).
  • Every plot_* function takes ax/axes and style as keywords, and a logarithmic axis is log_x/log_y instead of log; plot_dose_proportionality takes the ProportionalityResult which proportionality_test now returns instead of its former tuple.
  • read_events, read_pknca and read_adnca name their column keywords *_col (id_col, time_col, dv_col, amt_col, ...); groups of read_pknca is covariates.
  • Timecourses.n_workers and NCAOptions.n_workers: None is automatic (the pool only above the size threshold) and 1 is serial; pass a number to force that many workers.
  • stats.effects_from_arrays uses EffectKind.HEDGES_G by default, like the other entry points; pass kind= for another one.
  • The readers and Timecourses.from_dataframe raise ValueError naming the sample instead of a pydantic ValidationError: a sample with fewer than two time points, a NaN time, duplicate times or a dose which is not a valid protocol is caught by the batch itself, which no longer builds one Timecourse per sample.
  • Timecourses.from_dataframe rejects a value which is neither missing nor a number with a ValueError naming the sample and the column; a non-numeric time column raised a TypeError from pandas before, and a non-numeric dose column with dose_time was read as a missing dose.
  • ParameterResult.summarize no longer reports _sd, _se, _ci_low, _ci_high, _geomean and _geocv for the discrete parameters (tmax, tlast, tau, the counts and the diagnostics of the terminal regression); they keep x, x_median, x_q25, x_q75 and x_n.
  • plot_timecourse no longer takes ax positionally; plot_forest takes exp before ax and style; draw_nca_panel(timecourse, values, flags, *, log_y, title, ax, style) takes its data first, ax as a keyword, and returns the Axes instead of None.
  • plot_goodness_of_fit: log is replaced by log_x and log_y, which scale and mask the two axes on their own. plot_bland_altman: log is renamed log_ratio, since it selects the statistic (the log ratio against the log mean) and not only the scale of an axis.
  • plot_nca, plot_nca_grid and plot_fit take axes, the panels to draw into; the log of plot_nca_grid is log_y, and so is the log of plot_parameters.
  • plot_dose_proportionality and plot_bland_altman take ax.
  • write_events names its column keywords *_col like the readers, so id no longer shadows the builtin.
  • read_adnca and read_pknca gain covariates; the ADNCA column keywords subject, param, value, value_unit, time_first, time_ref, dose, dtype and lloq are subject_col, param_col, value_col, value_unit_col, time_first_col, time_ref_col, dose_col, dtype_col and lloq_col; the PKNCA subject is subject_col.
  • stats.multiple_comparison(p_values, *, method=...) and stats.effects_from_arrays(estimates, variances, *, labels=..., kind=..., ci_level=...) take their options by keyword.
  • stats.substrate_sensitivity(auc_ratio, *, thresholds=None) takes the thresholds as an optional keyword (the FDA ones by default) instead of a required positional argument.
  • compare_models(models, x, y): y defaults to None and x also takes a Timecourse or a Timecourses, in which case y, sd, dims, coords, x_unit and y_unit must not be given.
  • The result of the NCA carries the variables lambda_z_t_last and lambda_z_span and the flag NCAFlag.SPAN_LOW (1024); pkpdutils.nca.terminal.TerminalFit gains the field t_last after t_first, so a positional construction of it has to be adapted.
  • ParameterResult.summarize reports x_min and x_max for every parameter and x_cv for every non-discrete one, and reduces the point variables of summarized_point_variables (the interval_* parameters of a multiple dose analysis, which it dropped before) over the sample dimension, keeping their interval dimension.
  • pkpdutils.result.SUMMARY_SUFFIXES holds _min and _max, so a fit model whose parameter or derived name ends in one of them (e_max, c_min) is rejected by the reserved suffix guard of pkpdutils.fit.engine and has to be spelled emax, cmin.
  • plot_timecourse(by=...) colors the curves by group and writes one legend entry per group, where it used to give every sample its own color and entry with the group value as its label; a figure of more than max_legend (12) entries gets no legend at all. The function gains facet and, with it, axes.
  • plot_nca draws the legend on the linear panel only, plot_nca_grid once for the figure (into the first panel when the caller supplies axes); its panel titles are dose = 50 mg, individual = s1 instead of 50.0|s1, and its ncols is clamped to the number of samples, so a batch of one sample no longer produces a figure of three panels. draw_nca_panel gains legend.
  • plot_ratio and plot_forest annotate the rows with estimate [low, high] (and the weight of a study) by default and widen the x axis to hold the column; annotate=False restores the bare figure. plot_ratio gains labels for the row names.
  • A curve with an infusion is drawn with the window of the infusion (a shaded span from the dose time to the end of the dose) in plot_timecourse and draw_nca_panel; the NCA panel marks the dose it analyses alone, since it starts at that dose.
  • The figures label an axis whose unit is the canonical long form of pint with the short symbols instead (trough [mg/l], not trough [milligram / liter]), and leave the unit out of the label of a dimensionless variable; a unit the data spells itself (hr, ng/ml) is unchanged.
  • plot_ratio draws no tick at unity when the interaction thresholds are drawn, whose 0.8 and 1.25 crowd it.
  • fit_timecourse, fit_timecourses and fit_table store attrs["x_name"] and attrs["y_name"] on the result; plot_fit and plot_dose_proportionality label their axes with them instead of x and y.
  • The n of a batch is one count per sample or, when a count varies over the curve, one count per sample and time point: Timecourses.mean keeps the count of every time point, and from_timecourses and from_dataframe keep a varying count instead of reducing it to its maximum with a warning. Timecourses.n_subjects is the number of subjects of a sample in either layout and is what the analyses read.

Features

  • Batch ergonomics: Timecourses.select(**indexers) cuts a batch by labels, lists or slices of a sample dimension or of a coordinate along one, groupby(coord) walks its groups, mean(dim, *, spread, min_n) reduces a sample dimension to the group curve with sd, se and the count of every time point, dose_normalized(reference=None) divides the values by the dose, relative_to_dose(which) shifts every sample by its own dose time, and Timecourse.to_batch(dim, label) wraps a single curve as a batch.
  • Publication tables: summary_table(result, dim, *, by, parameters, stats, digits, units, layout) (also a method of every result) writes one row per parameter and group with the statistics n, mean, sd, cv, geomean, geocv, median, min, max and range as formatted strings; stats.ratio_table, stats.ddi_table and fit.proportionality_table do the same for the ratio, the interaction and the dose proportionality tables.
  • ParameterResult.summarize(dim) adds x_cv, x_min and x_max and reduces the interval_* parameters of a multiple dose analysis over the sample dimension as well.
  • Quality of the terminal phase: the analysis reports the window it used as lambda_z_t_first, lambda_z_t_last and lambda_z_span (the half-lives it covers) and flags a window below two half-lives with NCAFlag.SPAN_LOW.
  • Figures: plot_mean_timecourse (the mean of every group with its spread band, the individuals faint behind it, on a linear and a semi-logarithmic panel), plot_troughs (the trough of every dosing interval, the figure of steady state), plot_timecourse with by, facet and max_legend, a legend drawn once per figure, panel titles with names and units, annotated rows in plot_ratio and plot_forest, and axis labels from the names of the fitted variables.
  • Every public function of pkpdutils.stats takes its options as the enumeration member or as its string (compare(a, b, scale="log")), and so does plot_parameters(scale=...); an unknown string is rejected instead of falling through to a default.
  • Exchange formats: all three readers take covariates, write_events writes SD, SE and N columns when the batch carries them and read_events reads them back (sd_col, se_col, n_col), and route strings are coerced everywhere.
  • pkpdutils exports what a script configures its analyses with: the enumerations of the options, the model library of the fit, summary_table and pkpdutils.plot as a module, so a script imports from pkpdutils and pkpdutils.plot only.
  • pkpdutils.parallel holds the worker pools of the package: one lazily created, atexit-closed executor per kind, shared by the non-compartmental analysis (threads) and the fit (processes).

Performance

  • read_events of 100 000 event rows is 21 times faster (vectorized with pandas instead of scalar loops).
  • Timecourses.from_dataframe of 10 000 subjects is 232 times faster (the batch is validated once instead of once per subject).
  • Iterating a batch of 10 000 curves is 15 times faster (the curves are built from the numpy arrays of the batch).
  • The interval analysis of a multiple dose batch is 3.3 times faster (the interval block is gathered once per row).
  • nca of 100 000 rows is 2.4 times faster (chunked rows over the shared thread pool).
  • The first pooled fit takes 0.08 s instead of 1.18 s, because the process pool is created once and reused.
  • The bootstrap of 1 000 curves with 1 000 replicates peaks at 485 MB instead of 891 MB, because the replicates are materialized per chunk.

Fixes

  • The formulas of the documentation are rendered again after a navigation between pages: the instant navigation of the site swapped the page content without typesetting, so only a directly loaded page showed its formulas (#76).
  • pkpdutils.stats: a string argument is coerced to its enumeration at the entry point instead of silently taking the default branch.
  • A degenerate sample (no finite value, one value, zero variance) gives NaN statistics or a clear ValueError instead of a ZeroDivisionError or an escaping RuntimeWarning.
  • The delta method reports x_geocv as the geometric CV over subjects, consistent with the bootstrap.
  • A batch which mixes single dose and multiple dose subjects analyses every row by its own protocol.
  • TerminalMethod.LAST_N honours exclude_cmax.
  • A zero dose gives NaN for the dose dependent parameters instead of an infinity.
  • unit="" is rejected with a message naming "dimensionless".
  • tissue survives the round trips of a batch.
  • write_events and read_events carry SD, SE and N.
  • A ragged batch round-trips through to_dataframe and from_dataframe, samples in their original order.
  • IU doses are accepted.
  • plot_fit(log_x=True) masks the non-positive values instead of drawing an empty panel.
  • A figure which draws into an ax of the caller leaves the layout engine of that figure untouched.

Documentation

  • docs/workflows.md: five walk-throughs from the data of a study to its table and figure (a study table, bioequivalence, drug-drug interaction, steady state, dose proportionality), every one of them runnable.
  • The quickstart of the start page is the first of those workflows in six lines, on the event table docs/data/study.csv.
  • docs/gallery.md: one card per example of the repository with its figure, one sentence, the core snippet and the links.
  • Six mermaid diagrams: the data model, the NCA pipeline, the paths of the uncertainty, the input paths of the statistics, the engine of the fit and the formats.
  • Every page carries one self-contained runnable snippet with its output, and the figures of the examples are rendered into docs/images/ by scripts/render_examples.py and committed.
  • tests/docs/test_snippets.py runs every python fence of every page in order, with warnings as errors, so the documentation cannot drift from the library.

Your pkpdutils team