Repository navigation
Tuning POUNCE per problem
POUNCE registers 441 solver options. This page is a measured answer to "which
ones should I actually set?", based on a study that ran ten non-default options
against all 47 problems of the Mittelmann ampl-nlp suite — 595 recorded
solves, with every claimed win re-timed and independently re-verified.
There is no globally good non-default setting. Every one of the ten options tested was net-negative or inert at the median across the suite. The best single global candidate improved the median by about 1%, which was inside the measurement noise of the machine it ran on.
Most problems get nothing. 26 of 47 improved beyond measurement noise. The other 21 were best left alone.
But per-problem tuning is worth real money when it lands. Across the suite, picking the best option per problem was worth 1.9×, and individual problems reached 14×, 10×, and 9×.
So: don't go looking for a flag to turn on globally. Do consider tuning a specific model you solve repeatedly — especially a slow one.
All numbers below come from one machine, one build, and one benchmark suite. They are a guide to what to try, not a promise about your model. Measure.
Guessing options is expensive and mostly fails. Diagnosing first works better.
pounce yourmodel.nl \
timing_statistics=yes print_timing_statistics=yes \
--json-output run.json --json-detail fullThe timing block breaks wall clock down by phase. The two questions that matter:
-
Is
LinearSystemFactorizationdominating? Then each iteration is expensive, and the lever is linear algebra. - Are there far more iterations than you'd expect? Then the trajectory is the problem, and the lever is the barrier strategy or the line search.
These take disjoint sets of options. Treating an iteration-count problem with an ordering change is the most common way to waste an afternoon.
| Symptom in the timing/iteration data | Try |
|---|---|
LinearSystemFactorization is an extreme share (>90%) |
feral_ordering=amd, then metis
|
| Hessian assembly/factorization is a large share, model is large | hessian_approximation=limited-memory |
| Many iterations; μ barely moves between them | mu_strategy=adaptive |
| Many iterations; model is mostly linear/equality constraints | mehrotra_algorithm=yes |
| Objective gradients are badly scaled, or scaling looks harmful | nlp_scaling_method=none |
| Model has obvious redundancy before the solver sees it | presolve=yes |
A faster run that lands somewhere else is not automatically wrong — nonconvex problems have many local optima, and a different feasible KKT point is a legitimate answer. But you should know which happened:
pounce verify yourmodel.nl yourmodel.sol # exit 0 = feasibleCompare objective values between runs deliberately. In the study, most winners landed on essentially the same objective, and among those that moved, more found a better point than a worse one — but a handful moved by more than 1%.
Across 42–47 problems each. "Cand" = beat the default by more than that problem's measurement noise; "Invalid" = failed to solve.
| Option | Cand | Neutral | Loss | Invalid | Median | Character |
|---|---|---|---|---|---|---|
feral_fma=yes |
7 | 33 | 2 | 0 | 1.01× | the only never-harmful one |
linear_system_scaling=ruiz |
0 | 41 | 1 | 0 | 1.00× | inert on this suite |
presolve=yes |
1 | 39 | 1 | 1 | 1.00× | free; rarely helps, once enormously |
nlp_scaling_method=none |
3 | 23 | 9 | 7 | 1.00× | mixed |
feral_ordering=amd |
3 | 29 | 10 | 0 | 0.99× | mostly inert, occasionally 3× |
mu_strategy=adaptive |
8 | 10 | 16 | 8 | 0.93× | high variance |
mehrotra_algorithm=yes |
4 | 5 | 11 | 22 | 0.92× | transformative or broken |
feral_ordering=metis |
3 | 14 | 25 | 0 | 0.89× | usually a loss |
feral_ordering=auto_race |
1 | 14 | 27 | 0 | 0.82× | see note below |
hessian_approximation=limited-memory |
4 | 2 | 29 | 7 | 0.78× | sharpest instrument here |
Two readings worth taking away:
hessian_approximation=limited-memory is high-risk, high-reward. It lost on
29 of 42 problems — and it produced the two largest wins in the entire study
(14.4× and 10.2×). If your model is large and a lot of your time is going into
the exact Hessian, it is worth one experiment. Do not enable it blindly.
mehrotra_algorithm=yes failed outright on 22 of 47. It is not a general-purpose
NLP setting; Ipopt's own documentation says as much. But on a model whose
constraints are mostly linear it can be dramatic — one problem went from 194
iterations to 16.
The options documentation currently recommends auto_race as the safe choice
when the per-problem winner is uncertain. On this suite it was the worst of
the ten options tested: 27 losses, one win, median 0.82×.
The reason is visible in the reported factor statistics — it frequently returns
the same ordering auto picks unaided, after paying roughly four symbolic
factorizations to find it. Where ordering genuinely paid, a pinned choice won:
the most factorization-dominated problem in the suite (96% of wall clock) went
3.08× faster with feral_ordering=amd, and auto_race was the weakest of the
three explicit choices even there.
Tracking issue: #768.
Each of these was predicted from the diagnosis before being measured.
LP-like model → predictor-corrector. corkscrw is 86% equality constraints
and only 14% nonlinear. mehrotra_algorithm=yes took it from 194 iterations to
16, and 7.7 s to 0.85 s — 9.0×.
Stalled barrier → predictor-corrector. clnlbeam ran 2465 iterations across
just 6 μ levels, with μ frozen for 2448 consecutive iterations. Same option:
164 iterations, 98 s → 12 s — 8.1×.
Extreme factorization share → pin the ordering. qssp180 spent 96% of its
wall clock in factorization. feral_ordering=amd: 45 s → 14.6 s — 3.1×, at
an unchanged iteration count.
Expensive exact Hessian → limited memory. henon120 and lane_emden120 are
large, factorization-dominated models with dense Jacobian rows.
hessian_approximation=limited-memory gave 14.4× and 10.2× — in one case
while increasing the iteration count, because each iteration got roughly
thirteen times cheaper.
A cautionary result, because it is the most useful thing the study found.
dirichlet120 is structurally almost identical to henon120 and
lane_emden120 — same constraint count, same dense-row profile, same
factorization-dominated timing. The same option that gave those two 10–14×
made it 2.2× slower: limited-memory failed to converge at all, burning 3000
iterations before a fallback solve rescued it.
It was not a sparsity effect. Limited-memory shrank its KKT factor more than either sibling's (12.6M → 233k nonzeros) and it still lost.
The lesson: structural similarity does not predict whether a limited-memory approximation will capture your model's curvature. There is no substitute for running the experiment on your model.
- Time the default run with
print_timing_statistics=yes. Keep it as your baseline. - Repeat that baseline once. The difference between the two runs is your noise floor — on the study's machine it was about 1%, but yours will differ. Any "improvement" smaller than it is not real.
- Pick two or three options from the symptom table. Run each once.
- Take anything that beats the baseline by clearly more than your noise floor,
run it two more times, and
pounce verifythe solution. - Record what you chose and why, next to the model. Future-you will not remember.
If nothing clears the noise floor, that is a normal outcome — it happened on 21 of 47 problems here. Leave the defaults alone.
- One host, one build, one benchmark suite. Your model may behave differently, and the study's own results include a problem that behaved opposite to two near-identical siblings.
- Single options only. Combinations were not systematically explored.
- These are not proposed defaults. Each is a trajectory change, and POUNCE requires a fixture sweep before any default moves. Nothing here has had one.
Full data, per-problem write-ups and the exact commands to reproduce every run
are summarised in dev-notes/research/mittelmann-option-tuning.md in the
repository, which also records the defects and documentation issues the study
turned up.