Bayesian own-price elasticity estimation with PyMC for long-format panels (one row per SKU–time observation).
This is currently an MVP for log-demand panels (per-SKU intercept + per-SKU elasticity) with:
- optional shared control regressors (
control_* columns) - optional hierarchical / partial pooling across group columns
- optional cross-price elasticities (
CrossElasticitySpec) - simple diagnostics + plotting helpers
- posterior predictive simulation for counterfactual price scenarios
- per-SKU revenue price optimization (
optimize_prices)
pip install pypricingOptional model-graph rendering (also needs the system Graphviz binaries):
pip install 'pypricing[graphviz]'Docs: pypricing.readthedocs.io
From a clone of this repository:
uv sync --extra dev
uv run pytestFor local docs builds: uv sync --extra docs.
import numpy as np
from pypricing import LogLogDemandModel, generate_mock_data
# 1) Example data (replace with your own panel)
df = generate_mock_data(
n_periods=20,
n_skus=5,
n_controls=2, # creates control_1, control_2
random_state=0,
)
# 2) Fit
model = LogLogDemandModel()
idata = model.fit(
df,
draws=1000,
tune=1000,
chains=4,
random_seed=42,
)
print(model.run_diagnostics())
print(model.fit_summary().head())
# 3) Counterfactual prediction (quantity column not required)
df_scenario = df.drop(columns=["quantity"]).copy()
df_scenario["price"] = df_scenario["price"] * 1.05 # +5% price scenario
pred = model.predict(df_scenario, hdi_prob=0.9, random_seed=123)
print(pred[["sku", "period", "price", "quantity_mean", "quantity_hdi_lower", "quantity_hdi_upper"]].head())
# 4) Plots
_ = model.plot_elasticity_posterior(hdi_prob=0.9)
sku0 = model.sku_levels_[0]
price_grid = np.linspace(df["price"].min(), df["price"].max(), 25)
controls = {c: float(df.loc[df["sku"] == sku0, c].iloc[0]) for c in model.control_names_}
_ = model.plot_response_curve(sku=sku0, price_grid=price_grid, controls=controls, hdi_prob=0.9)LogLogDemandModel.fit(df) expects a pandas DataFrame with level-scale columns:
- Required
sku(configurable viasku_col)price(configurable viaprice_col) — must be strictly positivequantity(configurable viaquantity_col) — must be non-negative
- Optional
control_* columns (or pass an explicitcontrol_columns=(...)viaPanelColumns) — must be numeric, no NaNs- hierarchy columns (e.g.
category_1,category_2) viaPanelColumns(group_columns=...) period(and oftenregion) forfit_train_test()/ cross-elasticity market cells
Internally the model works on logs:
log_price = log(price)log_quantity = log(max(quantity, quantity_floor))
The default quantity_floor=1.0 allows zero quantities without \log(0).
At a high level this package fits log-demand with Gaussian noise:
\log Q \sim \mathcal{N}(\mu, \sigma)
Where \mu is a per-SKU demand curve plus optional global control effects.
Pick the model class directly:
LogLogDemandModel(default/simple)QuadraticLogDemandModelSigmoidSaturationDemandModel
Constant elasticity log-log:
- Interpretation: \epsilon_{\text{sku}} is own-price elasticity.
- Example: \epsilon=-1.5 implies a 1% price increase → ~1.5% quantity decrease (locally / in expectation).
- Best when: you want a simple constant-elasticity approximation.
Allows elasticity to vary with price (curvature in log-price):
This parameterization enforces that the elasticity at a per-SKU midpoint price equals elasticity_sku.
The midpoint is computed from the training data as the median log-price per SKU.
- Best when: elasticity changes with price level (e.g., premium vs discount regimes).
Saturating response curve in level price (softplus / log-sigmoid form):
\mu = \alpha_{\text{sku}} - \mathrm{softplus}(b_{\text{sku}}(P - P_{\text{center,sku}})) + X\beta
The per-SKU center P_{\text{center,sku}} = \exp(\texttt{log\_price\_center\_sku}) is a learned parameter.
Its prior is centered at the empirical per-SKU median log-price from training data (log_price_midpoint_sku_), with default sigma=0.5.
The curve is parameterized so the elasticity at P_{\text{center,sku}} equals elasticity_sku (via b_sku = -2 * elasticity_sku / price_center_sku).
- Best when: response “flattens out” at extreme prices (a simple saturation behavior).
If your frame contains control_* columns (or you pass control_columns=(...)), the model includes a shared linear term X\beta:
- one global coefficient vector
beta_controlshared across all SKUs - controls must also be provided at prediction time
You can override priors by passing model_config to each model class constructor.
Each entry uses:
{
"dist": <PyMC distribution constructor>,
"kwargs": { ... },
}Example:
import pymc as pm
from pypricing import QuadraticLogDemandModel
model = QuadraticLogDemandModel(
model_config={
"alpha_sku": {"dist": pm.Normal, "kwargs": {"mu": 0.0, "sigma": 1.0}},
"elasticity_sku": {"dist": pm.Normal, "kwargs": {"mu": -1.0, "sigma": 0.7}},
"sigma": {"dist": pm.HalfNormal, "kwargs": {"sigma": 0.3}},
# "curvature_sku": ... (only used by QuadraticLogDemandModel)
# "beta_control": ... (only used when you have controls)
}
)Defaults today:
alpha_sku ~ Normal(mu=6, sigma=2)elasticity_sku ~ Normal(mu=-1, sigma=2)sigma ~ HalfNormal(sigma=0.5)beta_control ~ Normal(mu=0, sigma=0.5)(if controls exist)curvature_sku ~ Normal(mu=0, sigma=0.2)(only forquadratic)log_price_center_sku ~ Normal(mu=log_price_midpoint_sku_, sigma=0.5)(only forsigmoid)
sample_posterior_predictive(df)returns anxarray.Datasetwith posterior draws forlog_quantityandquantity.predict(df)returns a DataFrame with:quantity_meanquantity_hdi_lowerquantity_hdi_upper
Prediction requires:
skuandprice- all control columns used during fit (if any)
quantityis not required
Unknown SKUs at prediction time raise an error (no cold-start handling yet).
fit_train_test(df, period_col="period", test_size=0.2, ...):
- holds out the last fraction of unique periods (to avoid time leakage)
- fits on train, predicts on test
- returns:
rmseonquantity_meanhdi_coverage: fraction of true quantities inside the predicted HDItest_predictions: the prediction frame
<ModelClass>.save(path) saves the model InferenceData to NetCDF (.nc) with model attrs.
<ModelClass>.load(path) restores the model from that NetCDF artifact and rebuilds the PyMC model.
pypricing.posterior.summarize_quantity_multiplier_one_sku(...) converts posterior elasticity draws into a posterior over relative quantity change for a price multiplier m via m^{\epsilon}.
For all SKUs at once, use the model method:
model.quantity_multiplier_summary(price_multiplier=..., hdi_prob=...)- returns one row per SKU with mean and HDI bounds of the quantity multiplier
- optionally,
model.quantity_multiplier_summary(..., return_draws=True)returns(summary_df, draws_da)wheredraws_dahas dims("chain", "draw", "sku")
This is only valid for LogLogDemandModel (constant elasticity).
optimize_prices (also available as model.optimize_prices(...)) chooses one price per SKU that maximizes the posterior mean of revenue price * exp(μ), where μ is the demand-curve mean log-quantity (same convention as prediction; residual σ is not folded into exp(μ)).
from pypricing import LogLogDemandModel, generate_mock_data
df = generate_mock_data(n_periods=20, n_skus=5, n_controls=1, random_state=0)
model = LogLogDemandModel()
model.fit(df, draws=500, tune=500, chains=2, random_seed=0)
price_bounds = {
sku: (float(g["price"].min() * 0.8), float(g["price"].max() * 1.3))
for sku, g in df.groupby("sku")
}
controls_df = (
df.groupby("sku", as_index=False)[model.control_names_].median()
)
opt = model.optimize_prices(price_bounds=price_bounds, controls_df=controls_df)
print(opt.head())Caveats:
- Optimization is per-SKU and independent (1D SciPy search on log-price within bounds).
- Not supported when
cross_elasticityis enabled (revenue then depends on the full market cell). - For
LogLogDemandModel, expected revenue is proportional top^(1+ε)draw-wise; if elasticity is roughly constant and|ε| ≠ 1, the optimum often sits on a bound. - If the model was fit with controls, pass
controls_dfwith one row per SKU.
See docs/source/notebooks/quickstart.ipynb for plots (plot_revenue_vs_price, plot_optimization_summary).
Pass column mapping via PanelColumns:
from pypricing import CrossElasticitySpec, LogLogDemandModel, PanelColumns
model = LogLogDemandModel(
panel_columns=PanelColumns(group_columns=("category_1", "category_2")),
cross_elasticity=CrossElasticitySpec(mode="within_group", group_level=0),
)Omit cross_elasticity for own-price only. Use mode="all" for every directed SKU pair (no group_level).
- Cold-start prediction for unseen SKUs
- Joint / cross-aware price optimization (use counterfactual prediction on a full market cell instead)
Hosted docs (API + notebooks): https://pypricing.readthedocs.io
To build HTML locally (Sphinx + notebooks, no notebook re-execution):
uv sync --extra docs
cd docs && uv run make html
# open docs/build/html/index.htmluv sync --extra dev
uv run pytest # unit + integration (default)
uv run pytest -m unit # fast only
uv run pytest -m integration # short MCMC smoke
uv run pytest -m recovery # parameter recovery (slower)
uv run ruff check src tests