Skip to content
Vichayanun Wachirapusitanand edited this page May 6, 2025 · 4 revisions

Usage

This framework contains two main Python scripts:

# Make distributions for scale factor fitting.
python make_histogram.py YAML_FILE [--diagnosis]

# Make prefit and postfit plots.
python plot_histograms.py PLOT_YAML_FILE

make_histogram.py is the main script that generates the final 1D distributions containing events passing and failing the designated tagger. It requires a YAML input file (YAML_FILE) detailing everything regarding the setup, such as input ROOT file location, processes and tagging categories involved, and uncertainty definitions.

On the other hand, plot_histograms.py is an additional script that generates the prefit and postfit plots of the scale factor fitting. With this Python script, you can choose to either generate only prefit plots (to check if the prefit distributions are good for measurement) or both prefit and postfit plots (after the actual fitting has been done). As with the first script, it requires another YAML input file (PLOT_YAML_FILE) which has a different structure.

Workflow in a nutshell

  1. Prepare a set of ROOT ntuple files as inputs. The file should have the same tree name, which should be easy if all the ntuple files are made with the same set of analysis code.
  2. Prepare an input YAML file for make_histograms.py which defines event categories, tagging categories, and a rule for passing/failing the tagger.
  3. Run make_histograms.py with the input YAML file. Output files can then be used for scale factor fitting in Higgs Combine.
  4. Use Higgs Combine to calculate the scale factor fitting.
  5. Prepare another input YAML file for plot_histograms.py, specifying where the prefit and postfit distributions are, which variables to be plotted, which colour should be used for each process, etc.
  6. Run plot_histograms.py with the second input YAML file.

make_histograms.py

make_histograms.py will classify events into different event categories, tagging categories, and passing/failing distributions. Here are the definitions for each term:

  • Each MC ntuple file can be assigned into different tagging categories, such as events containing top-tagged jets, W-tagged jets, etc. This is equivalent to Higgs combine processes.
  • For each MC file, an event may be classified into event categories, such as $p_T \in [300, 400]$.
  • The events in each event category are then further classified depending on whether or not the event passes or fails the tagger cut, such as BDT cut greater than 0.5.

The end results are, per event category, two 1D distributions for events passing tagger cut and events failing tagger cut. Each distribution will contain multiple tagging categories in the same sense as the normal event distribution containing different MC processes. Finally, data events are classified in the same way and assigned into these distributions, but are not associated with any tagging categories.

The output files from this script are, per one event category,

  • ROOT file containing 1D distributions
  • an accompanying combine datacard for that event category
  • if --diagnosis option is present, diagnosis ROOT file containing all distributions created from each input ROOT file for MC, arranged by the file name order

This script will also create one helpful bash script invoking text2workspace program, which can be used on machines with HiggsCombine set up. You will need the output ROOT file, the combine datacard, and the script for Higgs Combine.

Caveats for make_histograms.py

  • This script will only generate 1D distribution and no intermediate 2D histogram templates. Normally, to save time, other frameworks may generate the intermediate 2D histogram templates (containing jet pT versus jet mass distribution, for example) for fast datacard generation in case the user wants to adjust the jet pT range. Unfortunately this may lead to bugs since the 2D histogram may not always have the exact pT ranges encoded.

  • To offer more flexibility in event categories which may not entirely rely on one pT variable only (such as scale factor measurements for two or more variables, where the categories do not have to follow in the grid fashion), instead of only defining pT ranges, you can (and must) define your own event categories. This means, for each event category, you must include all the variables needed in the rule associated with the category.

What's inside the output ROOT file

The output ROOT file contains all 1D histograms including passing and failing distributions. The naming convention is as follows:

  • For data histograms, data_{VARIABLE_NAME}_{EVENT_CATEGORY}_pass or data_{VARIABLE_NAME}_{EVENT_CATEGORY}_fail
  • For MC histograms,
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_pass_nominal
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_pass_{UNCERTAINTY}Up
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_pass_{UNCERTAINTY}Down
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_fail_nominal
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_fail_{UNCERTAINTY}Up
    • {TAGGING_CATEGORY}_{VARIABLE_NAME}_{EVENT_CATEGORY}_fail_{UNCERTAINTY}Down

If --diagnosis option is turned on for make_histograms.py, another ROOT file, per event category, will be created with the name diagnosis_{EVENT_CATEGORY}.root The naming convention in this file is {PROCESS}_{VARIABLE_NAME}_{UNCERTAINTY}_{up|down}_{FILE_INDEX}_{EVENT_CATEGORY}_{TAGGING_CATEGORY}_{pass|fail} where FILE_INDEX represents the order of the input file specified in the YAML file.

plot_histogram.py

plot_histogram.py will generate the prefit and postfit plots for the fitting. Prefit plots can be based on the output of make_histograms.py, while postfit plots must be based on the Higgs combine output using FitDiagnostics method. Specifically, the postfit file this script uses must be the output of

combine -M FitDiagnostics <WORKSPACE FILE> --redefineSignalPOIs SF_topmatched --saveShapes --saveWithUncertainties

which gives the output file named fitDiagnostics.root containing postfit histograms required. The structure of this file is well known, and is hardcoded in this script. See HiggsCombine example commands for further details.

You may choose to plot only prefit plots to inspect the distributions before fitting. To do this, omit the postfit section of the YAML input file.

The output plot files will be located in {SAVE_DIRECTORY}/{VARIABLE_NAME}/ with the names {prefit|postfit}_{pass|fail}_{EVENT_CATEGORY}, and are saved in both PNG and PDF formats.