Skip to content

YAMLInputSpecs

Vichayanun Wachirapusitanand edited this page May 8, 2025 · 3 revisions

YAML input specifications

Refer to specfile_test_2022.yaml and plotfile_2022.yaml for examples of YAML file structures for make_histograms.py and plot_histograms.py.

make_histograms.py

The YAML input for make_histograms.py must contain details about your input ROOT files and the configuation you would like to have. Example input file for this script is specfile_test_2022.yaml.

Required options

  • year: Analysis year.

  • lumi: Integrated luminosity.

  • lumiunit: Unit of integrated luminosity.

  • genweight: Generator weight definitions. Can be written as an expression.

  • treename: Tree name to look for in the input ROOT files. The tree name must be the same for all ROOT files.

  • processes: List of processes to be added into tagging categories. The structure must be as follows:

    • One data key is required. This specifies the input ROOT files for data. Another key nominal_files must be present in data, and must contain a list of input file paths.
    • Any number of processes, each as one key. Inside it must contain the keys nominal_files that specify the file paths for nominal event distributions, and unc_files that specify the file paths for event distributions for shape-based uncertainties, such as JES or JER.

    Multiple file paths for the same process can be added.

  • basecut: Base cut to be applied to all events.

  • categories: Tagging categories to be added to the datacard. (This is equivalent to processes in HiggsCombine.) Each key represents a tagging category, and must contain the list of processes (under processes key) and further cuts for that category (under cut key).

  • tagger: Details on the designated tagger.

    • name: Name of the tagger to be used in files. Required.
    • varname: Name of the discrimnator as seen in the ntuple files.
    • cut: Discriminator cut value for the tagger. Passing and failing events are defined as events with discriminator values higher and lower than this cut respectively.
    • cutrule: Rule for tagger cuts. Events that pass this tagger cut rule will be put in passing distribution, while events that fail this will be put in failing distribution. Overrides varname and cut keys.
  • distribution: Details on the final distribution for scale factor measurement. Must have the following two keys:

    • variables: A list of three-key objects as follows:

      • variable Target variable to be populated in the distribution, such as jet mass.
      • range: Range of the target variable.
      • bins: Number of bins in the distribution. Uneven bins are not yet supported.

      This design allows one YAML input file to quickly generate multiple distributions to inspect multiple variables at the same time.

    • event_categories: Definitions for event categories. Every event should be categorised into one of the event categories (provided that they are orthogonal), and then further classified into passing and failing categories based on the tagger. (This is equivalent to bins in HiggsCombine.)

      Each category must contain the following keys:

      • name: Name of the event category. The name used here will be used as output file names.
      • rule: Rule of the event category. This rule will be plugged directly into TTree.Project method, so the rules defined here should be in the compatible ROOT format.
  • uncertainties: Details on uncertainties to be added into the datacard. Each key is a uncertainty name, and should contain mode key inside. Currently mode supports lnN, factor, and file, and has different behaviours as follows:

    • lnN: Log-normal uncertainty. If no category key is present, this uncertainty will be applied to all tagging categories with the specified size value. Specify tagging categories to apply this uncertainty using category key.
    • factor: Shape uncertainty calculated from nominal input files with the designated ROOT expression. Must contain keys up and down.
    • file: Shape uncertainty calculated from files specified in unc_files under processes key. In this case, the uncertainty name must be the same as specified in processes key.

Optional

  • analysisname: Name of the output folder to be stored. Default is . or the current directory.

  • perfileweights: Adds a new branch with a constant value to ROOT files. New branches will be added directly to the TTree named in treename in specified ROOT files before any histograms are made, which is ideal for updating the cross section for certain files. This option should contain a list of dictionaries following this pattern:

    • name: Name for a new branch of the target TTree (named in treename)
    • value: Value of a new branch. Must be constant number. Values calculated based on other columns are not supported.
    • files: List of ROOT file paths for the new branch to be added.

    Optional:

    • lowess: Apply LOWESS smoothing algorithm to the uncertainty shape. Two options can be added here:
      • limit: Maximum relative size of the uncertainty. Default is 0.5.
      • symmetrize: Force the up/down uncertainty shape to be symmetric. Default is true.

    See specfile_test_2022.yaml for examples on how to use this option.

plot_histograms.py

The YAML input for plot_histograms.py has a different structure, aimed at plotting the prefit and postfit histograms. Example input file for this script is plotfile_test_2022.yaml.

Required Options

  • lumi: Integrated luminosity.
  • xlabel: Label on the x-axis. The range of x-axis is automatically determined by the histogram bins.
  • savedir: Save directory.
  • categories: Tagging categories as defined in the YAML input file for make_histograms.py. However, each key (name of the tagging category) must contain the following sub-keys:
    • color: Histogram colour, used in both prefit and postfit plots. Must be in the matplotlib-compatible format, such as RGB hex code (#0033a0).
    • propername: Name to be shown in the legend.
  • eventcats: List of event categories to be plotted.
    • name: Event category name as defined in event_categories option in the input YAML file for make_histograms.py.
    • variable: Target variable used to populate the histograms.
    • propername: Text to be added in the legend of the plot. If left unspecified, the name in the legend will be the variable name used in variable key.
    • prefitfile: Path to prefit file, or the output ROOT file from make_histograms.py.
    • postfitfile: Optional. Path to postfit file, or the output ROOT file from FitDiagnostics method of Higgs Combine.

Optional

  • com: Centre-of-mass energy, in TeV. Default is 13.