Repository navigation
1.1.0
Release notes for pkpdutils 1.1.0
Minor release of pkpdutils with dosing protocols, the multiple dose analysis and the exchange formats of pharmacokinetic data, followed by a cleanup and usability round: bug fixes, a faster batch path, consistent signatures, the tables and figures a publication prints, and documentation whose every snippet runs. The release breaks the 1.0.0 API where consistency demanded it; every change is listed below with the call to adapt.
Dosing protocols, multiple dosing and exchange formats
Dosingis the dosing protocol of a curve:amounts,times,durations(infusions), oneunitand oneroute, built withDosing.single(dose),Dosing.from_doses(doses)orDosing.regimen(dose, interval, n_doses);Timecourse.dosingholds it,Timecourse.dosereads the first dose back,relative_to_dose(which="first" | "last")shifts a curve and its protocol.- A batch (
Timecourses) carries the protocol of every sample over the dimensiondose_index(dose_amount,dose_time,dose_duration,NaNpadded);n_doses,first_dose_*,last_dose_*anddosing_of(**indexers)read it back. - The non-compartmental analysis of a multiple dose curve computes the parameters of every dosing interval (
interval_*variables over the dimensioninterval,NCAResult.intervals()as a table) and the steady state parameters of the last complete interval (auc_tau,cmax_ss,cmin_ss,ctrough,cavg,fluctuation,swing,accumulation_ratio,cl_ss); the point parameters are computed from the last dose on. FlagsINCOMPLETE_INTERVAL(the last interval is not covered by the data) andEXTRAPOLATED_TROUGH(the trough of a bolus interval was regressed).NCAOptions.tauanalyses a curve given with its last dose only;superpositionandaccumulation_ratiopredict the steady state from a single dose. pkpdutils.io:read_events/Timecourses.from_eventsandwrite_events/Timecourses.to_events(NONMEM and Monolix event records withEVID,AMT,DV,MDV,ADDL/II,SS,RATE/TINF, covariates as coordinates),read_pknca/Timecourses.from_pknca(the two PKNCA tables) andread_adnca/Timecourses.from_adnca(CDISC ADaM ADNCA). Every reader returns one batch with one sample dimension, the protocol of every subject and the covariates as coordinates.- Figures: the dose lines of a protocol in
plot_timecourse, the shadedAUC(0-tau)of a multiple dose result in the NCA panel,plot_intervalsof theinterval_*variables against the interval number.
Breaking changes of this part:
NCAOptions.regimenis removed. A repeated administration is given as theDosingprotocol of the timecourse (Dosing.regimen(dose, interval, n_doses),DosingRegimen.dosing()), and a curve given with its last dose only is analysed withNCAOptions.tau.superposition(timecourse, dosing)takes aDosingprotocol or aDosingRegimen; the predicted curve carries the protocol and no label.Timecourse.doseis a read-only property, the first dose ofTimecourse.dosing; the protocol is replaced withmodel_copy(update={"dosing": Dosing.single(dose)}), not withupdate={"dose": ...}.- A
Timecoursesdataset carries the dose dimensiondose_index:dose_amount,dose_timeanddose_durationare 2-D over(*sample_dims, dose_index); the 1-D view of a single dose batch isfirst_dose_amount/first_dose_timeandlast_dose_amount/last_dose_time. A sample dimension nameddose_indexis rejected. - A multiple dose analysis reports
cl,cl_f,vz,vz_f,vss,auc_inf_dnandcmax_dnasNaN, since the slice after the last dose carries the exposure of the earlier doses;auc_inf_obs,auc_inf_pred,aumc_infandmrtdescribe the decline after the last dose. - The clearance at steady state of an extravascular route is
cl_ss_f, notcl_ss. read_eventsandread_pkncarequire theroutekeyword: a batch has one route and the two formats carry none.
Breaking changes of the cleanup round
nca,nca_single,partial_aucandsuperpositiontakeoptionsas a keyword, asfitdoes:nca(batch, options=NCAOptions(...));t_endofsuperpositionfollowsoptionsin the signature and is a keyword as well.FitResultreportsp_cvas a fraction instead of a percentage; a table or a figure which printed it directly multiplies by 100 (the tables of the package do it themselves).- Every
plot_*function takesax/axesandstyleas keywords, and a logarithmic axis islog_x/log_yinstead oflog;plot_dose_proportionalitytakes theProportionalityResultwhichproportionality_testnow returns instead of its former tuple. read_events,read_pkncaandread_adncaname their column keywords*_col(id_col,time_col,dv_col,amt_col, ...);groupsofread_pkncaiscovariates.Timecourses.n_workersandNCAOptions.n_workers:Noneis automatic (the pool only above the size threshold) and1is serial; pass a number to force that many workers.stats.effects_from_arraysusesEffectKind.HEDGES_Gby default, like the other entry points; passkind=for another one.- The readers and
Timecourses.from_dataframeraiseValueErrornaming the sample instead of a pydanticValidationError: a sample with fewer than two time points, aNaNtime, duplicate times or a dose which is not a valid protocol is caught by the batch itself, which no longer builds oneTimecourseper sample. Timecourses.from_dataframerejects a value which is neither missing nor a number with aValueErrornaming the sample and the column; a non-numerictimecolumn raised aTypeErrorfrom pandas before, and a non-numeric dose column withdose_timewas read as a missing dose.ParameterResult.summarizeno longer reports_sd,_se,_ci_low,_ci_high,_geomeanand_geocvfor the discrete parameters (tmax,tlast,tau, the counts and the diagnostics of the terminal regression); they keepx,x_median,x_q25,x_q75andx_n.plot_timecourseno longer takesaxpositionally;plot_foresttakesexpbeforeaxandstyle;draw_nca_panel(timecourse, values, flags, *, log_y, title, ax, style)takes its data first,axas a keyword, and returns theAxesinstead ofNone.plot_goodness_of_fit:logis replaced bylog_xandlog_y, which scale and mask the two axes on their own.plot_bland_altman:logis renamedlog_ratio, since it selects the statistic (the log ratio against the log mean) and not only the scale of an axis.plot_nca,plot_nca_gridandplot_fittakeaxes, the panels to draw into; thelogofplot_nca_gridislog_y, and so is thelogofplot_parameters.plot_dose_proportionalityandplot_bland_altmantakeax.write_eventsnames its column keywords*_collike the readers, soidno longer shadows the builtin.read_adncaandread_pkncagaincovariates; the ADNCA column keywordssubject,param,value,value_unit,time_first,time_ref,dose,dtypeandlloqaresubject_col,param_col,value_col,value_unit_col,time_first_col,time_ref_col,dose_col,dtype_colandlloq_col; the PKNCAsubjectissubject_col.stats.multiple_comparison(p_values, *, method=...)andstats.effects_from_arrays(estimates, variances, *, labels=..., kind=..., ci_level=...)take their options by keyword.stats.substrate_sensitivity(auc_ratio, *, thresholds=None)takes the thresholds as an optional keyword (the FDA ones by default) instead of a required positional argument.compare_models(models, x, y):ydefaults toNoneandxalso takes aTimecourseor aTimecourses, in which casey,sd,dims,coords,x_unitandy_unitmust not be given.- The result of the NCA carries the variables
lambda_z_t_lastandlambda_z_spanand the flagNCAFlag.SPAN_LOW(1024);pkpdutils.nca.terminal.TerminalFitgains the fieldt_lastaftert_first, so a positional construction of it has to be adapted. ParameterResult.summarizereportsx_minandx_maxfor every parameter andx_cvfor every non-discrete one, and reduces the point variables ofsummarized_point_variables(theinterval_*parameters of a multiple dose analysis, which it dropped before) over the sample dimension, keeping theirintervaldimension.pkpdutils.result.SUMMARY_SUFFIXESholds_minand_max, so a fit model whose parameter or derived name ends in one of them (e_max,c_min) is rejected by the reserved suffix guard ofpkpdutils.fit.engineand has to be spelledemax,cmin.plot_timecourse(by=...)colors the curves by group and writes one legend entry per group, where it used to give every sample its own color and entry with the group value as its label; a figure of more thanmax_legend(12) entries gets no legend at all. The function gainsfacetand, with it,axes.plot_ncadraws the legend on the linear panel only,plot_nca_gridonce for the figure (into the first panel when the caller suppliesaxes); its panel titles aredose = 50 mg, individual = s1instead of50.0|s1, and itsncolsis clamped to the number of samples, so a batch of one sample no longer produces a figure of three panels.draw_nca_panelgainslegend.plot_ratioandplot_forestannotate the rows withestimate [low, high](and the weight of a study) by default and widen the x axis to hold the column;annotate=Falserestores the bare figure.plot_ratiogainslabelsfor the row names.- A curve with an infusion is drawn with the window of the infusion (a shaded span from the dose time to the end of the dose) in
plot_timecourseanddraw_nca_panel; the NCA panel marks the dose it analyses alone, since it starts at that dose. - The figures label an axis whose unit is the canonical long form of pint with the short symbols instead (
trough [mg/l], nottrough [milligram / liter]), and leave the unit out of the label of a dimensionless variable; a unit the data spells itself (hr,ng/ml) is unchanged. plot_ratiodraws no tick at unity when the interaction thresholds are drawn, whose 0.8 and 1.25 crowd it.fit_timecourse,fit_timecoursesandfit_tablestoreattrs["x_name"]andattrs["y_name"]on the result;plot_fitandplot_dose_proportionalitylabel their axes with them instead ofxandy.- The
nof a batch is one count per sample or, when a count varies over the curve, one count per sample and time point:Timecourses.meankeeps the count of every time point, andfrom_timecoursesandfrom_dataframekeep a varying count instead of reducing it to its maximum with a warning.Timecourses.n_subjectsis the number of subjects of a sample in either layout and is what the analyses read.
Features
- Batch ergonomics:
Timecourses.select(**indexers)cuts a batch by labels, lists or slices of a sample dimension or of a coordinate along one,groupby(coord)walks its groups,mean(dim, *, spread, min_n)reduces a sample dimension to the group curve withsd,seand the count of every time point,dose_normalized(reference=None)divides the values by the dose,relative_to_dose(which)shifts every sample by its own dose time, andTimecourse.to_batch(dim, label)wraps a single curve as a batch. - Publication tables:
summary_table(result, dim, *, by, parameters, stats, digits, units, layout)(also a method of every result) writes one row per parameter and group with the statisticsn,mean,sd,cv,geomean,geocv,median,min,maxandrangeas formatted strings;stats.ratio_table,stats.ddi_tableandfit.proportionality_tabledo the same for the ratio, the interaction and the dose proportionality tables. ParameterResult.summarize(dim)addsx_cv,x_minandx_maxand reduces theinterval_*parameters of a multiple dose analysis over the sample dimension as well.- Quality of the terminal phase: the analysis reports the window it used as
lambda_z_t_first,lambda_z_t_lastandlambda_z_span(the half-lives it covers) and flags a window below two half-lives withNCAFlag.SPAN_LOW. - Figures:
plot_mean_timecourse(the mean of every group with its spread band, the individuals faint behind it, on a linear and a semi-logarithmic panel),plot_troughs(the trough of every dosing interval, the figure of steady state),plot_timecoursewithby,facetandmax_legend, a legend drawn once per figure, panel titles with names and units, annotated rows inplot_ratioandplot_forest, and axis labels from the names of the fitted variables. - Every public function of
pkpdutils.statstakes its options as the enumeration member or as its string (compare(a, b, scale="log")), and so doesplot_parameters(scale=...); an unknown string is rejected instead of falling through to a default. - Exchange formats: all three readers take
covariates,write_eventswritesSD,SEandNcolumns when the batch carries them andread_eventsreads them back (sd_col,se_col,n_col), androutestrings are coerced everywhere. pkpdutilsexports what a script configures its analyses with: the enumerations of the options, the model library of the fit,summary_tableandpkpdutils.plotas a module, so a script imports frompkpdutilsandpkpdutils.plotonly.pkpdutils.parallelholds the worker pools of the package: one lazily created,atexit-closed executor per kind, shared by the non-compartmental analysis (threads) and the fit (processes).
Performance
read_eventsof 100 000 event rows is 21 times faster (vectorized with pandas instead of scalar loops).Timecourses.from_dataframeof 10 000 subjects is 232 times faster (the batch is validated once instead of once per subject).- Iterating a batch of 10 000 curves is 15 times faster (the curves are built from the numpy arrays of the batch).
- The interval analysis of a multiple dose batch is 3.3 times faster (the interval block is gathered once per row).
ncaof 100 000 rows is 2.4 times faster (chunked rows over the shared thread pool).- The first pooled fit takes 0.08 s instead of 1.18 s, because the process pool is created once and reused.
- The bootstrap of 1 000 curves with 1 000 replicates peaks at 485 MB instead of 891 MB, because the replicates are materialized per chunk.
Fixes
- The formulas of the documentation are rendered again after a navigation between pages: the instant navigation of the site swapped the page content without typesetting, so only a directly loaded page showed its formulas (#76).
pkpdutils.stats: a string argument is coerced to its enumeration at the entry point instead of silently taking the default branch.- A degenerate sample (no finite value, one value, zero variance) gives
NaNstatistics or a clearValueErrorinstead of aZeroDivisionErroror an escapingRuntimeWarning. - The delta method reports
x_geocvas the geometric CV over subjects, consistent with the bootstrap. - A batch which mixes single dose and multiple dose subjects analyses every row by its own protocol.
TerminalMethod.LAST_Nhonoursexclude_cmax.- A zero dose gives
NaNfor the dose dependent parameters instead of an infinity. unit=""is rejected with a message naming"dimensionless".tissuesurvives the round trips of a batch.write_eventsandread_eventscarrySD,SEandN.- A ragged batch round-trips through
to_dataframeandfrom_dataframe, samples in their original order. IUdoses are accepted.plot_fit(log_x=True)masks the non-positive values instead of drawing an empty panel.- A figure which draws into an
axof the caller leaves the layout engine of that figure untouched.
Documentation
docs/workflows.md: five walk-throughs from the data of a study to its table and figure (a study table, bioequivalence, drug-drug interaction, steady state, dose proportionality), every one of them runnable.- The quickstart of the start page is the first of those workflows in six lines, on the event table
docs/data/study.csv. docs/gallery.md: one card per example of the repository with its figure, one sentence, the core snippet and the links.- Six mermaid diagrams: the data model, the NCA pipeline, the paths of the uncertainty, the input paths of the statistics, the engine of the fit and the formats.
- Every page carries one self-contained runnable snippet with its output, and the figures of the examples are rendered into
docs/images/byscripts/render_examples.pyand committed. tests/docs/test_snippets.pyruns every python fence of every page in order, with warnings as errors, so the documentation cannot drift from the library.
Your pkpdutils team