A set of tools to analyze and categorized genome neighborhoods in microbial genomes, and to generate publication-friendly diagrams.
Particular focuses: investigation and classification of genome neighborhoods that do not resemble natural product pathways, integration of hypothetical proteins into analyses, and compatibility with IMG datasets.
- Introductory note
- Components
of
prettyClusters, and how you use them - Development information, including updates (most recent: 20250813) and a to-do list
There are a number of other excellent tools for looking at genome
neighborhoods in various ways, including
antiSMASH,
EFI-GNT,
BiG-SCAPE and
BiG-SCAPE-CORASON,
CAGECAT (a server wrapper for
cblaster and
clinker),
zol/fai, and
socialGene. However, you may want to
use prettyClusters if:
- you’re starting out with a gene/protein of interest and don’t know much about the pathways it’s part of yet.
- you are working on pathways that don’t look like the type of well-studied pathways found in MiBIG and detected by antiSMASH.
- you are working on pathways that have a lot of hypothetical proteins, or that have a lot of proteins with vague/blanket annotations.
- you want to visualize both genome neighborhood content and relationships, and potentially to correlate it to sequence similarity for your protein of interest.
- you use a lot of datasets from the JGI’s IMG database.
Additionally: this is a work in progress, and I’m a chemical biologist: there will be bugs and inefficient code, but I do my best to fix them as they’re identified.
The Wiki entries for each function contain a detailed description of theuse of specific functions.
generateNeighbors. Sets up the list of neighboring genes for your gene of interest.prepNeighbors. QC, and provisional assignment of hypothetical protein families.analyzeNeighborsQuantification of common types of neighboring genes, genome neighborhood classification.prettyClusterDiagramsVisualization of genome neighborhoods.
gbToIMG. Processing GenBank files for aprettyClustersworkflow.incorpIprScan. Incorporate InterProScan annotations (primarily needed when using GenBank input.)repnodeTrim. When integrating EFI-EST SSNs, trim the active dataset to representative nodes.identifySubgroups. Assign provisional subgroup annotations within a large group of related proteins.trimFasta. Provide a trimmed multiFASTA file for a subset of proteins.
- A basic installation guide is the best starting point. Includes cross-platform installation instructions for everything needed to make this package work.
- For users working with IMG data, I’ve got a walkthrough for a standard run.
- For users working with data from other sources, I’ve got a rough
workflow
for getting data from non-IMG sources, and another
workflow
for standardizing it for use in
prettyClusters. - Also check out the troubleshooting list for cryptic-sounding errors that I’ve seen pop up enough times to be worth noting, as well as some additional limitations.
Illustrating the output of some of the components (or, in the case of
the cluster diagrams themselves, just under 10% of the output for this
example):
Notably, sequence similarity and genome neighborhood similarity are not
always tightly coupled. The analyses in prettyClusters make it
possible to investigate a protein family along both axes.
Until I get a proper paper out, you can use citation() to generate a
basic citation for the version of prettyClusters that you are
currently using.
citation("prettyClusters")
#> To cite package 'prettyClusters' in publications use:
#>
#> Kenney G (2024). _prettyClusters: Exploring and Classifying Genomic
#> Neighborhoods Using IMG-Like Data_. R package version 0.3.0.
#>
#> A BibTeX entry for LaTeX users is
#>
#> @Manual{,
#> title = {prettyClusters: Exploring and Classifying Genomic Neighborhoods Using IMG-Like Data},
#> author = {G. E. Kenney},
#> year = {2024},
#> note = {R package version 0.3.0},
#> }Additionally, this package makes use of other tools, including:
Camacho C., Coulouris G., Avagyan V., Ma N., Papadopoulos J., Bealer K., Madden T.L. BMC Bioinformatics (2008) 10:421
Nakamura T., Yamada K.D., Tomii K., Katoh K. Bioinformatics (2018) 34:2490–2492
HMMER - note that citing the website is preferred.
HMMER 3.4 (Aug 2023); http://hmmer.org/
Eddy S. R. PLOS Comp. Biol. (2011) 7:e1002195
citation("gggenes")
#> To cite package 'gggenes' in publications use:
#>
#> Wilkins D (2023). _gggenes: Draw Gene Arrow Maps in 'ggplot2'_. R
#> package version 0.5.1, <https://CRAN.R-project.org/package=gggenes>.
#>
#> A BibTeX entry for LaTeX users is
#>
#> @Manual{,
#> title = {gggenes: Draw Gene Arrow Maps in 'ggplot2'},
#> author = {David Wilkins},
#> year = {2023},
#> note = {R package version 0.5.1},
#> url = {https://CRAN.R-project.org/package=gggenes},
#> }citation("tidygraph")
#> To cite package 'tidygraph' in publications use:
#>
#> Pedersen T (2024). _tidygraph: A Tidy API for Graph Manipulation_. R
#> package version 1.3.1,
#> <https://CRAN.R-project.org/package=tidygraph>.
#>
#> A BibTeX entry for LaTeX users is
#>
#> @Manual{,
#> title = {tidygraph: A Tidy API for Graph Manipulation},
#> author = {Thomas Lin Pedersen},
#> year = {2024},
#> note = {R package version 1.3.1},
#> url = {https://CRAN.R-project.org/package=tidygraph},
#> }citation("gggenomes")
#> To cite package 'gggenomes' in publications use:
#>
#> Hackl T, Ankenbrand M, van Adrichem B (2024). _gggenomes: A Grammar
#> of Graphics for Comparative Genomics_. R package version 1.0.1,
#> <https://CRAN.R-project.org/package=gggenomes>.
#>
#> A BibTeX entry for LaTeX users is
#>
#> @Manual{,
#> title = {gggenomes: A Grammar of Graphics for Comparative Genomics},
#> author = {Thomas Hackl and Markus J. Ankenbrand and Bart {van Adrichem}},
#> year = {2024},
#> note = {R package version 1.0.1},
#> url = {https://CRAN.R-project.org/package=gggenomes},
#> }- Version 0.3.0 See the release notes.
- Switched GenBank import systems to handle
genbankrdeprecation. Check changes to function and the workflow. - Improved handling of coloring edge cases in
prettyClusterDiagrams - Minor tweaks in peptide handling and NA handling in
prepNeighborsandidentifySubgroups(and in metadata/graph export for the latter) incorpIprScandeals gracefully with some InterProScan updates- An unhelpful and non-fatal blast error is provisionally silenced in
prepNeighborsandidentifySubgroups
Under consideration, but no guarantees about order or timeframe.
- Looking into setting up a shiny and potentially interactive GUI, if not full online deployment. The activation energy for getting a new user up and running on R is unfortunately real!
- A way to handle non-gene things as neighborhood “anchors” - regulator binding sites, riboswitches, etc.? Annotated elements like tRNAs are more straightforward; user-supplied coordinates and pseudo-genes (analogous to approaches used for unannotated peptides) might be the way to go.
- A gene co-occurrence tool to handle datasets where there is no one
gene of interest to “anchor” the cluster (i.e. a cluster has 2+ of a
larger set of gene families within a constrained genomic region, or
within the genome.) Basically loosening some of the
analyzeNeighborsrequirements. Will need a modified approach for diagram generation without a core GoI.coGenesor something similar. - May try to auto-generate a “typical” genome neighborhood, illustrating
order and abundance of neighbors in a given gene cluster family? Still
working on how to automate this well. Would be
averageCluster. - Similarly, possibly a third sort of per-cluster diagram with
highlighting of %ID between homologs - more along the lines of
clinker, but without the GenBank
input. Or like gggenomes, but
protein rather than nucleotide similarity. (Possibly employing one of
those tools if I can figure out a way to easily do so!) Likely to be
compareCluster.
- For use when working with really large families, tweak
repnodeTrimand add an option for using the EFI-EST toolset before even generating neighborhoods, as a way of getting a more manageable dataset. “If you need more than 128 GB of RAM to open the SSN, you may find this helpful…” - Generation of HMMs for hypothetical protein families identified in
prepNeighborsandidentifySubgroups(and with it the ability to turn on and off MSA and HMM generation in both tools.) - User-supplied HMMs for annotation of predefined custom protein (sub)families as a standalone subfunction.
- Switch to MMseqs2 for clustering
- Options to let the user specify distance and clustering methods in
prepNeighborsandanalyzeNeighbors. (Perhaps you prefer Jaccard to Euclidean distance, for example?)