Skip to content

External Sampler Requirements

Jared O'Neal edited this page Jul 9, 2026 · 1 revision

Intent

This is a collaborative document for developers to work in to build up a common understanding of what we want to achieve as we alter surmise so that it can use bilby samplers under the hood.

At some point, this document will have served its purpose and we will likely neither finish nor maintain it. We will endeavor to update this section when that occurs.

User stories

As a surmise user, I would like

  1. To be able to use surmise with as many different emulators, calibrators and samplers (ECSs) as possible so that I have the flexibility and freedom to choose the best options for my problem.
  2. To know that all the ECSs are well-tested within the package so that my trust in the package improves.
  3. clear examples of how to specify which ECSs to use including any differences in the package interface that might vary depending on my choices so that I can learn about this mechanism and trust my ability to use it correctly.
  4. to programmatically access the current set of supported ECSs so that my surmise applications can dynamically assess which tools are available and which are not with my current installation.
  5. it to be made clear and explicit if surmise provides different implementations of the same ECSs (e.g., a surmise native implementation and an external implementation) so that I know that I have options and how to choose between implementations.
  6. To be able to use the same sampler one or more times to execute independent sampling using different functions or on the same set of functions.
  7. To be able to run more than one sampler in parallel and all independently as part of a single calibration process for the purposes of checking the quality of the sampler's results. There are several possiblities:
    • Allow users to run multiple independent calibrations in parallel with each calibrator using one and only one sampler. This puts the burden of running and executing the verification fully on the user, but has a simpler design with a small interface.
    • Allow each calibrator to access and run in parallel multiple samplers. This reduces the burden on the user, but the calibrators must determine what to do with the multiple different sampler results, how to use these internally, and how to make their samples available to the user. In addition, the calibrator interface will need to be expanded.
    • Allow each sampler to be calibrated so that it can run many different sampler executions independently and in parallel internally with the sampler determining how to combine all results into a single set of results for the calling calibrator. In this case, each calibrator still performs a single calibration with a single calibrator to produce a single result. The calibrator does not make any decisions about how to manage multiple results, but its interface would have to allow for relaying verification information from the sampler back to the user. We will already have a potential case of this type if users want to use emcee via bilby.
  8. To be able to export an instance of the chosen ECS option trained with one set of data so that a separate user may use the trained ECS on a different machine assuming compatible surmise versions.
  9. To be able to access intermediate computational results if desired, either for debugging or scientific understanding purposes. This should include bundling intermediate results produced by user and external methods with all results produced natively by surmise so that file management and recordkeeping should be easy and sufficient.
    • MC: store all intermediate outputs in a log. If code is run in debug mode(?), save into a local file and provide location. Otherwise, if runs are successful, delete the intermediate outputs. In the case where failures occur, save the log.
  10. To be able to access samples and relevant information during and after sampling, such that corner plots, convergence analyses, and trace plots of posterior samples can be produced.

As a surmise power user, I would like

  1. to be able to add my own ECSs to the list of available implementations of a current surmise installation without having to alter that installation so that I can use surmise to its fullest on my challenging problems without altering/breaking pre-existing Python virtual environments.
  2. to be able to add my own ECSs to the list of available implementations of a current surmise installation without having to understand deeply the inner design of surmise so that I can use surmise to its fullest on my challenging problems without becoming a surmise developer or having to change my implementation to adjust to low-level implementation changes in surmise.
  3. To be able to select external ECSs (e.g., bilby), user-provided ECSs, and surmise-internal ECSs in a simple and similar way so that the interface treats all possible options in the same way regardless of their source.
  4. To be able to emulate simultaneously a function and its gradient so that surmise methods can make use of the gradients when my computational model can provide these.
    • Some of the calibration methods are written in such a way that it appears that a user could provide gradients not through the emulator but rather by providing a lpdf_grad member function in their prior. Are users really allowed this pathway or should the grad only be made available via the emulator? If they can provide their own, perhaps we need a new user story including specification of which should be taken first.
    • MC: The gradient of the likelihood is sometimes provided by a user to benefit the MLE problem in fitting an emulator. Similarly, gradients of the log-likelihood in the calibrator method can be supplied as the loglik_grad, for example in directbayeswoodbury.

As a surmise developer, I would like

  1. To couple in external and user-provided ECSs with some checks to ensure that they are adhering to the design interface specified by surmise so that surmise provides extra sanity checks for users and can detect interface changes/mismatches as quickly as possible.
  2. to make the user aware that the likelihood of the emulator is determined by the choice of the emulator and the emulator is fitted via maximum likelihood estimation. (See ch.5 of Gramacy (2020) for discussion.)
  3. to keep the surmise design relatively simple by decoupling as much as possible emulation and calibration. In the case where a fully Bayesian analysis is desired, meaning that the hyperparameters in the emulator are also sampled, surmise is not the suitable choice of implementation, since sampling packages like bilby or emcee can handle this use case.
  4. To allow each concrete calibration method to determine the final configuration of the sampler that it will use based on a given user-provided sampler, given user-provided sampler configuration values, and its calibration-specific knowledge.
  5. To inform users of technical details regarding the likelihood imposed by their choice of calibration method, but design the software so that calling code has no need to directly access the likelihood. Ditto for the emulator. This is motivated by the desire to simplify the public interface of the software as much as possible.
  6. To design the public interface of the software so that users have no direct access to samplers in the interest of making the interface as simple as possible. However, the full set of results returned from samplers should be made available to users so that users can judge for themselves if the performance is acceptable.
  7. To allow users to use in principle any possible combination of sampler and calibration methods as part of normal use of the high-level interface while allowing developers to prohibit with clear communication to user when an unsupported or invalid pairing is requested.
    • If some combinations are prohibited, does this mean that we should provide a list of prohibited combinations to satisfy User Req 6?
  8. To give each calibration method the ability to also provide to samplers the derivative of the loglikelihood when gradient of the log of the joint prior are available so that the samplers can also make use of the derivative of the log joint posterior.
  9. To allow each calibration method to be able to construct their log joint posterior function and optionally the derivative of the log joint posterior from not only a user-provided joint prior and the internally defined log likelihood (and potentially its derivative), but also extra method-specific terms so that the methods have the flexibility to construct method-specific posteriors.
    • Ideally, the construction of the log joint prior and potentially its derivative would be generic and could always be done the same way for each calibaration method, which needs to just supply the log joint prior, log likelihood, and derivatives. In such a case, we could create a construct_log_joint_posterior() function in the sampler portion of the private interface.
    • A use case for this story appears to be mlwoodbury, which can use the additional function logprobacce to build its log joint posterior. Could this method provide arguments to the sampler so that the sampler could build the correct log joint posterior using construct_log_joint_posterior()? There are presently no other use cases.
      • MC: It seems like the logprobacce and its gradient are general finite difference code to estimate the gradient.
  10. To allow all log densities to be provided as unnormalized quantities but insist that a value of -numpy.inf be provided for all theta values where the density is zero so that writing and maintaining log density codes can be as simple as possible by leveraging the use of log densitities in the different methods while imposing a common standard. Note that this standard must also be communicated to users since they must provide log prior densities. TRUE?
  11. I would like to develop new methods in the repository in accord with the project's git workflow with small feature branches that are merged into main frequently despite the fact that the new method is not yet ready for official use. Ideally, the package would have a scheme whereby inclusion of methods in the package does not mean that users have access to the new methods. In addition, it would be nice if a power user, such as the developer, can force the use of such unofficial methods in the package. Ideally this can be done without moving files in the git repo so that the commit history of the files is easy to access.

As a surmise maintainer, I would like

  1. To have the surmise package designed so that the sets of available, externally-provided ECSs can grow, shrink, or change dynamically without having to adjust the surmise package so that surmise and the external providers are decoupled as much as possible and so that users are immediately able to access all possible ECSs at installation without making requests to maintainers.
  2. To have the GH actions run automatically every time a new version of a package that provides external ECSs is published so that we can discover changes/breakages as soon as possible and confirm compatibility otherwise.
  3. To be able to run each surmise test case with all random number generation under my control so that I can test exact reproduction of key results when refactoring code in such a way that results should not change.

Clone this wiki locally