Skip to content

operon_rules

github-actions[bot] edited this page Jul 21, 2026 · 6 revisions

Operon Rules

Efesto applies operon-context filters to remove spurious single-gene hits and enforce co-occurrence requirements before reporting clusters. Two rule engines are available: a default FeGenie-exact port, and a configurable JSON rule engine.


Default filtering (no operon_rules.json)

Without operon_rules.json in --hmm_dir, Efesto uses an exact reimplementation of FeGenie's per-category operon rules. Each cluster is routed to one handler based on which special genes are present. For iron acquisition categories, a per-ORF pass/break pattern is used: a gene that fails co-occurrence is silently skipped, but other genes in the same cluster are still reported.

Use --all_results to bypass all filtering entirely.


JSON rule engine

If operon_rules.json exists in --hmm_dir, the JSON engine replaces the default. The file must be in the same directory as the HMM files.

Top-level keys

{
  "report_all_categories": ["metal_resistance-*", "iron_storage"],
  "rules": [ ... ]
}

report_all_categories — glob patterns. Clusters whose entire category set matches bypass all rules and are reported unconditionally. MetHMMDB categories (metal_resistance-*) and iron_storage bypass by default. iron_sulfur_assembly and iron_stress are also in this list — single-gene hits are informative.

Rule object schema

{
  "name":            "RULE_NAME",
  "categories":      ["category_1", "category_2"],
  "genes":           ["GeneA", "GeneB"],
  "rule":            "require_n_of",
  "min_genes":       2,
  "on_fail":         "drop",
  "canonical_size":  5,
  "max_bp_gap":      2000,
  "neutral_energizer_stems": ["PF03544_TonB_C", "PF13103_TonB_2"]
}
Field Required Description
name yes Rule identifier (shown in debug output)
categories yes Categories this rule governs
genes rule-dependent Gene names from FeGenie-map.txt
rule yes Rule type (see below)
min_genes rule-dependent Minimum count threshold
on_fail yes Action when rule fails
canonical_size no Expected gene count for completeness scoring
max_bp_gap no Override global --max_bp_gap for this rule's genes
neutral_energizer_stems no Stems that are substrate-agnostic (see below)

Rule types

Rule Behaviour
require_n_of Cluster must contain ≥ min_genes unique members of genes list
require_anchor A specific anchor gene must be present in the cluster
require_n_cat Cluster must contain ≥ min_genes distinct HMMs from categories
require_n_cat_or_lone_trusted As require_n_cat, or lone hit in trusted_lone list
mtr_disambiguation Re-assigns iron_oxidation vs iron_reduction based on MtrA/MtoA co-presence

On-fail actions

Action Behaviour
passthrough_non_members Keep genes NOT in the rule's genes list; drop genes that are in it
drop Remove the entire cluster from output
keep_all Keep everything regardless of rule outcome

Per-rule bp-gap override (max_bp_gap)

When max_bp_gap is set in a rule, Efesto applies the more restrictive of the two adjacent genes' rule-specific gaps during clustering. This makes tight operons (e.g. MtrMto: 2 000 bp, FoxABC: 1 000 bp) stricter than the global window without penalising loosely encoded systems (e.g. SIDERO_SYNTH: 5 000 bp).


The TonB energizer guard (neutral_energizer_stems)

Background — why TonB/ExbBD are substrate-agnostic

TonB, ExbB, and ExbD form the Ton motor complex — an inner membrane machinery that harnesses the proton motive force (PMF) and delivers it to TonB-dependent outer membrane transporters (TBDTs). The structural basis is well established: ExbB forms a pentamer with an ExbD dimer inside its transmembrane pore; this complex converts the electrochemical gradient into mechanical energy; TonB spans the periplasm and cyclically contacts TBDTs via a conserved five-residue TonB box, triggering plug movement and substrate release into the periplasm (Celia et al. 2016, Nature; Silale & van den Berg 2023, Annu Rev Microbiol; Braun 2024, Mol Microbiol).

Critically, the Ton motor is substrate-agnostic. The same TonB-ExbBD complex energises ALL TBDTs regardless of their cargo. In Bacteroidetes, TBDTs (SusC-like proteins) are the invariant component of every polysaccharide utilisation locus (PUL), and TonB-ExbBD powers their carbohydrate import — not iron uptake (Bolam & Koropatkin 2012, Curr Opin Struct Biol; Pollet et al. 2021, Mol Microbiol). During phytoplankton blooms, TonB-dependent transporters were the most highly expressed protein class in marine bacterioplankton (~16.7% of all detected proteins), with the majority predicted to target polysaccharides (Francis et al. 2021, ISME J). In Xanthomonas plant pathogens, CAZyme clusters are organised in CUT systems (Carbohydrate Utilisation with TonB-dependent transporters), directly analogous to Bacteroidetes PULs (Giuseppe et al. 2023, Essays Biochem).

The problem

Efesto models the Ton motor with five HMMs:

Stem Pfam Role
PF03544_TonB_C PF03544 TonB C-terminal periplasmic domain
PF13103_TonB_2 PF13103 TonB domain 2
PF16031_TonB_N PF16031 TonB N-terminal anchor
PF01618-MotA_TolQ_ExbB_proton_channel_family PF01618 ExbB / MotA / TolQ proton channel
PF02472-ExbD PF02472 ExbD periplasmic signalling

All five are in iron_acquisition-siderophore_transport_potential. The SIDERO_TRANSPORT rule requires min_genes: 2 from siderophore_transport_potential. Therefore:

PF03544_TonB_C  +  PF01618_ExbB  →  2 hits  →  passes SIDERO_TRANSPORT

This is a false positive in Bacteroidetes-rich metagenomes: a Ton motor hit without a siderophore receptor is evidence of carbohydrate transport, not iron acquisition.

The fix — neutral_energizer_stems

The SIDERO_TRANSPORT rule accepts an optional neutral_energizer_stems list. If ALL hits in a cluster come exclusively from this list (no substrate-specific receptor), the rule treats it as a fail → drop.

{
  "name":       "SIDERO_TRANSPORT",
  "categories": ["iron_acquisition-siderophore_transport_potential",
                 "iron_acquisition-heme_transport",
                 "iron_acquisition-siderophore_transport"],
  "rule":       "require_n_cat_or_lone_trusted",
  "min_genes":  2,
  "neutral_energizer_stems": [
    "PF03544_TonB_C",
    "PF13103_TonB_2",
    "PF16031_TonB_N",
    "PF01618-MotA_TolQ_ExbB_proton_channel_family",
    "PF02472-ExbD",
    "HasB-TonB_paralog_heme_uptake-rep"
  ],
  "trusted_lone": [ ... ],
  "on_fail": "drop",
  "max_bp_gap": 2000
}

Logic (proposed implementation in operon.py):

def passes_energizer_guard(cluster_hits, neutral_stems):
    """Return False (fail) if all hits are neutral energizers."""
    non_neutral = [h for h in cluster_hits if h["stem"] not in neutral_stems]
    return len(non_neutral) > 0

A cluster passes the guard when at least one hit is a substrate-specific gene (receptor, periplasmic binding protein, ABC permease, etc.).

Clusters the guard catches

Composition Guard outcome Reason
TonB_C + ExbB Drop Ton motor only; substrate unknown
TonB_C + ExbD + FutA1 Pass FutA1 is an iron-specific ABC binding protein
TonB_C + PF00593-TonB_dependent_receptor Pass TonB-dep receptor is substrate-specific
ExbB + FepC-ATPase Pass FepC is siderophore-specific
HasB + HxuA Pass HasB (heme-specific TonB paralog) + HxuA is heme-specific

HasB — the heme-specific TonB paralog

HasB (HasB-TonB_paralog_heme_uptake-rep) is a TonB paralog in Serratia marcescens and related organisms. Unlike canonical TonB it is dedicated to the heme acquisition system (HasASB), acting specifically with the HasR outer membrane heme receptor. Its ExbB-interaction surface has a unique long periplasmic extension that discriminates it from canonical TonB (Biou et al. 2022, Commun Biol). HasB is therefore substrate-specific (heme) — it is listed in neutral_energizer_stems only when appearing WITHOUT HasR; when co-clustered with HasR or other heme acquisition genes, it is a valid heme-transport signal.

Status: The neutral_energizer_stems guard is the proposed implementation. The operon.py rule engine requires a ~20-line update to support it. See the open issue on the development tracker.


Current rules (active operon_rules.json)

FLEET

Iron oxidation via the electron shuttle complex in Sideroxydans-type organisms. Requires ≥ 5 of 8 subunits (EetA/B, Ndh2, FmnB/A, DmkA/B, PplA). canonical_size: 8, max_bp_gap: 2000. On fail: keep non-FLEET genes.

MAM

Magnetosome island. Requires ≥ 5 of 10 MAM genes (MamA/B/E/K/P/M/Q/I/L/O). canonical_size: 10, max_bp_gap: 3000. On fail: keep non-MAM genes.

FOXABC

Acidithiobacillus FoxABC iron oxidation complex. Requires ≥ 2 of 3 subunits. canonical_size: 3, max_bp_gap: 1000.

FOXEYZ

FoxEYZ cytochrome bc1-type complex. FoxE is the mandatory anchor gene. canonical_size: 3, max_bp_gap: 1000.

DFE1 / DFE2

Desulfosporosinus iron reduction operons 1 and 2. Require ≥ 3 of 4 / 3 of 5 subunits respectively. max_bp_gap: 1000.

MtrMto

Mtr/Mto disambiguation. Re-assigns iron_oxidation vs iron_reduction based on co-presence of MtrA, MtrB, MtrC (reduction) vs MtoA (oxidation) and CymA. canonical_size: 3, max_bp_gap: 2000. On fail: keep all.

SIDERO_TRANSPORT

Siderophore/heme transport. Requires ≥ 2 distinct transport HMMs, or a lone trusted receptor. max_bp_gap: 2000. On fail: drop. Trusted lone genes: FutA1, FutA2, FutC, LbtU/LbtB variants, IroC.

SIDERO_SYNTH

Siderophore biosynthesis. Requires ≥ 3 distinct synthesis HMMs. max_bp_gap: 5000 (biosynthesis operons can be large). On fail: drop.

IRON_TRANSPORT

Iron and heme transport. Requires ≥ 2 distinct HMMs from iron_acquisition-iron_transport and iron_acquisition-heme_oxygenase. max_bp_gap: 2000. On fail: drop.

TAD_PILUS

Tad/Flp pilus (Chen et al. 2025, Nature, MISO — iron oxide respiration coupled to sulfide oxidation). Requires both Flp_Fap (Pfam PF04964 pilin subunit) and cpaB (NCBIfam TIGR03177 assembly platform) — they are obligate operon partners, and either gene alone is a widespread, non-metal-specific adhesion/conjugation signal found across unrelated bacteria (the paper itself notes the organism's other type IV pilin, pilA, is constitutively expressed and not a useful iron-reduction marker). canonical_size: 2, max_bp_gap: 2000. On fail: passthrough_non_members (a lone hit of either gene is dropped from the reported cluster).

SUF_OPERON

SUF Fe-S cluster assembly operon (sufABCDSE). sufA, sufD, and sufE all have near-zero TC–NC calibration margin (confirmed empirically: the sufE model cross-hits its non-operon paralog csdE at a comparable bitscore to true sufE — see docs/hmm_library_curation.md) and are marked needs_confirmation in the registry. Requires ≥ 3 of the 6 SUF genes co-occurring — in practice this means co-occurrence with the well-calibrated anchors sufB/sufC/sufS, since a false-positive hit for sufA/sufD/sufE alone has no reason to coincidentally cluster with two others. canonical_size: 6, max_bp_gap: 2000. On fail: passthrough_non_members (a lone weak hit unsupported by anchors is dropped).


--catalog_mode

Use when input is a deduplicated (non-redundant) gene catalog — e.g. an mmseqs2/CD-HIT clustered set of representative ORFs pooled across many samples/assemblies — where genomic context is unavailable.

Without --catalog_mode, and without --gff_dir/--fna_dir supplying real coordinates, efesto falls back to grouping ORFs by parsing the trailing _<int> off each header and treating IDs numerically close together (within --max_gap) as genomic neighbors. That assumption holds for a single genome/MAG's native Prodigal output (contig_1, contig_2, contig_3… really are adjacent), but not for a gene catalog: representative sequence IDs reflect catalog/cluster bookkeeping, not genomic position. Two unrelated genes from different samples/contigs can end up with adjacent numeric IDs purely by chance of catalog construction, and would otherwise be merged into a fake "operon cluster."

--catalog_mode turns off clustering entirely — every ORF is treated as its own singleton cluster. As a consequence it bypasses:

  • co-occurrence count rules (FLEET ≥ 5, siderophore ≥ 2/3, iron_transport ≥ 2, SUF_OPERON ≥ 3, etc.) — these need ≥2 genes in one cluster to fire at all
  • Mtr/Mto co-occurrence disambiguation — hits keep their per-HMM default category from the registry instead

It keeps: bitscore cutoffs and Cyc2 logic, both of which are evaluated per-ORF and don't depend on cluster membership.

Before trusting any operon-context/co-occurrence feature on a catalog input (i.e. before running without --catalog_mode on catalog data — don't do this, but if you're unsure whether your catalog's IDs happen to preserve real genomic order), check whether your header IDs encode real coordinates. Many gene-calling tools (Prodigal, pyrodigal) write full positional metadata directly in the FASTA defline, e.g.:

>SAMPLE_CONTIGID_5 # 91309 # 92004 # 1 # ID=...;partial=00;start_type=ATG;...

If your catalog retains this format, pick one recurring SAMPLE_CONTIGID stem with many entries and check whether its start coordinate increases monotonically with the trailing index:

grep "^>SAMPLE_CONTIGID_" your_catalog.faa | \
  sed 's/^>//' | awk '{print $1, $3, $5}' | sort -t_ -k3 -n

Monotonically increasing, non-overlapping positions across the whole stem means that stem is a real single contig (e.g. a complete long-read assembly), and coordinate-based clustering would be biologically meaningful for it. Positions that jump backward or repeat in blocks mean the ID has been reused/relabeled across unrelated original contigs during catalog construction — in that case there is no genuine adjacency to recover for that stem, and --catalog_mode's singleton-cluster behavior is the only safe option regardless of your header format.

Clone this wiki locally