Skip to content

Tuning POUNCE per problem

John Kitchin edited this page Aug 24, 2026 · 1 revision

Tuning POUNCE per problem

POUNCE registers 441 solver options. This page is a measured answer to "which ones should I actually set?", based on a study that ran ten non-default options against all 47 problems of the Mittelmann ampl-nlp suite — 595 recorded solves, with every claimed win re-timed and independently re-verified.

Read this first

There is no globally good non-default setting. Every one of the ten options tested was net-negative or inert at the median across the suite. The best single global candidate improved the median by about 1%, which was inside the measurement noise of the machine it ran on.

Most problems get nothing. 26 of 47 improved beyond measurement noise. The other 21 were best left alone.

But per-problem tuning is worth real money when it lands. Across the suite, picking the best option per problem was worth 1.9×, and individual problems reached 14×, 10×, and 9×.

So: don't go looking for a flag to turn on globally. Do consider tuning a specific model you solve repeatedly — especially a slow one.

All numbers below come from one machine, one build, and one benchmark suite. They are a guide to what to try, not a promise about your model. Measure.

The method: diagnose, then choose

Guessing options is expensive and mostly fails. Diagnosing first works better.

1. Find out where the time goes

pounce yourmodel.nl \
  timing_statistics=yes print_timing_statistics=yes \
  --json-output run.json --json-detail full

The timing block breaks wall clock down by phase. The two questions that matter:

  • Is LinearSystemFactorization dominating? Then each iteration is expensive, and the lever is linear algebra.
  • Are there far more iterations than you'd expect? Then the trajectory is the problem, and the lever is the barrier strategy or the line search.

These take disjoint sets of options. Treating an iteration-count problem with an ordering change is the most common way to waste an afternoon.

2. Match the symptom to a lever

Symptom in the timing/iteration data Try
LinearSystemFactorization is an extreme share (>90%) feral_ordering=amd, then metis
Hessian assembly/factorization is a large share, model is large hessian_approximation=limited-memory
Many iterations; μ barely moves between them mu_strategy=adaptive
Many iterations; model is mostly linear/equality constraints mehrotra_algorithm=yes
Objective gradients are badly scaled, or scaling looks harmful nlp_scaling_method=none
Model has obvious redundancy before the solver sees it presolve=yes

3. Check the answer, not just the clock

A faster run that lands somewhere else is not automatically wrong — nonconvex problems have many local optima, and a different feasible KKT point is a legitimate answer. But you should know which happened:

pounce verify yourmodel.nl yourmodel.sol   # exit 0 = feasible

Compare objective values between runs deliberately. In the study, most winners landed on essentially the same objective, and among those that moved, more found a better point than a worse one — but a handful moved by more than 1%.

What each option actually did

Across 42–47 problems each. "Cand" = beat the default by more than that problem's measurement noise; "Invalid" = failed to solve.

Option Cand Neutral Loss Invalid Median Character
feral_fma=yes 7 33 2 0 1.01× the only never-harmful one
linear_system_scaling=ruiz 0 41 1 0 1.00× inert on this suite
presolve=yes 1 39 1 1 1.00× free; rarely helps, once enormously
nlp_scaling_method=none 3 23 9 7 1.00× mixed
feral_ordering=amd 3 29 10 0 0.99× mostly inert, occasionally 3×
mu_strategy=adaptive 8 10 16 8 0.93× high variance
mehrotra_algorithm=yes 4 5 11 22 0.92× transformative or broken
feral_ordering=metis 3 14 25 0 0.89× usually a loss
feral_ordering=auto_race 1 14 27 0 0.82× see note below
hessian_approximation=limited-memory 4 2 29 7 0.78× sharpest instrument here

Two readings worth taking away:

hessian_approximation=limited-memory is high-risk, high-reward. It lost on 29 of 42 problems — and it produced the two largest wins in the entire study (14.4× and 10.2×). If your model is large and a lot of your time is going into the exact Hessian, it is worth one experiment. Do not enable it blindly.

mehrotra_algorithm=yes failed outright on 22 of 47. It is not a general-purpose NLP setting; Ipopt's own documentation says as much. But on a model whose constraints are mostly linear it can be dramatic — one problem went from 194 iterations to 16.

A note on feral_ordering=auto_race

The options documentation currently recommends auto_race as the safe choice when the per-problem winner is uncertain. On this suite it was the worst of the ten options tested: 27 losses, one win, median 0.82×.

The reason is visible in the reported factor statistics — it frequently returns the same ordering auto picks unaided, after paying roughly four symbolic factorizations to find it. Where ordering genuinely paid, a pinned choice won: the most factorization-dominated problem in the suite (96% of wall clock) went 3.08× faster with feral_ordering=amd, and auto_race was the weakest of the three explicit choices even there.

Tracking issue: #768.

Worked examples

Each of these was predicted from the diagnosis before being measured.

LP-like model → predictor-corrector. corkscrw is 86% equality constraints and only 14% nonlinear. mehrotra_algorithm=yes took it from 194 iterations to 16, and 7.7 s to 0.85 s — 9.0×.

Stalled barrier → predictor-corrector. clnlbeam ran 2465 iterations across just 6 μ levels, with μ frozen for 2448 consecutive iterations. Same option: 164 iterations, 98 s → 12 s — 8.1×.

Extreme factorization share → pin the ordering. qssp180 spent 96% of its wall clock in factorization. feral_ordering=amd: 45 s → 14.6 s — 3.1×, at an unchanged iteration count.

Expensive exact Hessian → limited memory. henon120 and lane_emden120 are large, factorization-dominated models with dense Jacobian rows. hessian_approximation=limited-memory gave 14.4× and 10.2× — in one case while increasing the iteration count, because each iteration got roughly thirteen times cheaper.

When the diagnosis will mislead you

A cautionary result, because it is the most useful thing the study found.

dirichlet120 is structurally almost identical to henon120 and lane_emden120 — same constraint count, same dense-row profile, same factorization-dominated timing. The same option that gave those two 10–14× made it 2.2× slower: limited-memory failed to converge at all, burning 3000 iterations before a fallback solve rescued it.

It was not a sparsity effect. Limited-memory shrank its KKT factor more than either sibling's (12.6M → 233k nonzeros) and it still lost.

The lesson: structural similarity does not predict whether a limited-memory approximation will capture your model's curvature. There is no substitute for running the experiment on your model.

Practical recipe

  1. Time the default run with print_timing_statistics=yes. Keep it as your baseline.
  2. Repeat that baseline once. The difference between the two runs is your noise floor — on the study's machine it was about 1%, but yours will differ. Any "improvement" smaller than it is not real.
  3. Pick two or three options from the symptom table. Run each once.
  4. Take anything that beats the baseline by clearly more than your noise floor, run it two more times, and pounce verify the solution.
  5. Record what you chose and why, next to the model. Future-you will not remember.

If nothing clears the noise floor, that is a normal outcome — it happened on 21 of 47 problems here. Leave the defaults alone.

Caveats

  • One host, one build, one benchmark suite. Your model may behave differently, and the study's own results include a problem that behaved opposite to two near-identical siblings.
  • Single options only. Combinations were not systematically explored.
  • These are not proposed defaults. Each is a trajectory change, and POUNCE requires a fixture sweep before any default moves. Nothing here has had one.

Source

Full data, per-problem write-ups and the exact commands to reproduce every run are summarised in dev-notes/research/mittelmann-option-tuning.md in the repository, which also records the defects and documentation issues the study turned up.