The AlloPipe tool is a computational workflow designed to compute, given a pair of annotated human genomic datasets:
- directional amino acid mismatches, and
- the related candidate minor histocompatibility antigens.
The product is provided free of charge, and, therefore, on an "as is" basis, without warranty of any kind.
AlloPipe is also available as a web application.
A. Dhuyser, P. Delaugère, P. Laville, et al., AlloPipe and Its Web Server Allogenomics: From Genomic Data to Candidate Minor Histocompatibility Antigens, HLA 107, no. 2 (2026): e70590.
Full text: https://doi.org/10.1111/tan.70590.
The AlloPipe tool is divided into two sequential modules: Allo-Count, then Allo-Affinity.
After reformating relevant data from the variant-annotated .VCF file(s), Allo-count performs a stringent data cleaning and computes the directional comparison of the genomic sequences from sample 1 and sample 2.
Allo-Count returns:
- a quantitative output called the Allogenomic Mismatch Score (AMS): a discrete quantitative variable that is counting the number of directional amino acid mismatches.
- a qualitative output stored in the AMS (mismatch) table: providing information about the polymorphisms contributing to the AMS.
Directional comparison
The sample comparison is directional and accounts for either polymorphisms that are present in the donor but absent in the recipient (donor-to-recipient) or that are present in the recipient but absent in the donor (recipient-to-donor).
- Donor-to-recipient accounts for polymorphisms present by the donor but absent by the recipient, i.e. triggerring the recipient's immune system after solid organ transplantation.
- Recipient-to-donor accounts for polymorphisms present by the recipient but absent by the donor, i.e. triggerring the donor's immune system after allogeneic haematopoietic cell transplantation.
Allo-Affinity reconstructs peptides of requested length around the polymorphisms present in the mismatches tables.
The affinity of those peptides towards the HLA molecules can then be assessed using third party tools such as NetMHCpan or MixMHCpred softwares, to retrieve the candidate mHAgs.
Please read the terms of use for the NetMHCpan and MixMHCpred softwares. 4-digit HLA typing must be provided by the user for the HLA molecules of interest, including the alpha/beta chain combination for HLA-DR and HLA-DQ molecules.
There are two modes of operation for each module: as single pair or as multiple pairs
-
Single pair: Run as 'single pair mode' if you aim to compute Allo-Count and/or Allo-Affinity for one pair at a time.
You need to provide one variant-annotated.VCFfile per individual. -
Cohort (Multiple pairs): Run through the Nextflow cohort workflow if you aim to compute Allo-Count and/or Allo-Affinity for more than one pair at a time.
You need to provide one unique variant-annotated.VCFfile containing the genotypes of all individuals you want to analyse - i.e. a joint.VCFfile - and the.csvformatted list of the pairs you want to process. For cohort mode, the pair list can also carry a per-pairhlacolumn with a comma-separated HLA typing string.
For installing AlloPipe you will specifically require the following softwares:
-
Nextflow 24.10.9 installed to run the AlloPipe workflow.
-
Conda installed for your operating system. The Nextflow workflow uses Conda to create its execution environment from
recipes/conda/allopipe.yml. AlloPipe is currently developed on Python v3.10. -
To run Allo-Affinity, you need to assess the affinity of the reconstructed peptides towards the HLA molecules. We recommend two groups of software suites for that (only NetMHCpan is supported in command line for now):
- NetMHCpan and NetMHCIIpan, which should be downloaded as command line tools (be careful with version numbers).
- MixMHCpred and MixMHC2pred (support development in progress).
-
To predict proteasomal cleavage on the proteins of the donor or the recipient, you will also need the NetChop tool installed as a standalone version.
Make sure you use each software in accordance with its user license.
-
Clone the repository from git
You might be requested to create a token for you to log in. See the GitHub tutorial -
Run AlloPipe through the Nextflow workflow
The following command lines will clone the repository and show the workflow help:
git clone https://github.com/huguesrichard/Allopipe.git
cd Allopipe
nextflow run main.nf --help
- Remember that to run prediction of affinity for the peptides you will also need NetMHCpan installed (NetChop to account for proteasomal cleavage).
AlloPipe input file(s) must be variant-annotated .VCF file(s). We highly recommend performing the variant annotation with the most recent version of VEP using the command line installation and all the arguments specified below.
Any variant annotator could be used at this step, but keep in mind that AlloPipe has been developed with
.VCFfiles in version 4.2 annotated with VEP command line installation for versions older than 103.
To install the VEP command line tool, follow the installation tutorial available here.
During the installation, you will be asked if you want to download cache files, FASTA files and plugins.
- We recommend downloading the cache files for the assembly of your
.VCFfiles to be able to run VEP offline.
Download the VEP cache files which correspond to your Ensembl VEP installation and genome reference! - We recommend downloading the FASTA files for the assembly of your
.VCFfiles to be able to run VEP offline.
Download the FASTA files which correspond to your Ensembl VEP installation and genome reference! - In default mode, we do not recommend downloading any plugin, except the Frameshift plugin when frameshift neoantigen handling is required. See the Frameshift plugin section.
We then recommend adding VEP to your PATH by adding the following line to your ~/.profile or ~/.bash_profile:
export PATH=%%PATH/TO/VEP%%:${PATH}
Run the following command to annotate your .VCF file(s) with VEP.
All specified options are mandatory, with the exception of the assembly if you only downloaded one cache file.
vep --fork 4 --cache --assembly <GRChXX> --offline --af_gnomade -i <FILE-TO-ANNOTATE>.vcf -o <ANNOTATED-FILE>.vcf --coding_only --pick_allele --use_given_ref --vcf
Where:
<GRChXX>is the version of the genome used to align the sequences.<FILE-TO-ANNOTATE>.vcfis the path to your file to annotate.<ANNOTATED-FILE>.vcfis the path to the output annotated file.
This command line works for individual .VCF files or joint .VCF files, whether compressed (.vcf.gz) or not (.vcf).
Run this command for every file you want to input in AlloPipe.
When using the Nextflow workflow with automatic VEP annotation enabled, AlloPipe selects the VEP assembly from --ensembl_path. The path must contain GRCh37 or GRCh38; for example, --ensembl_path data/Ensembl/GRCh38 uses GRCh38 for VEP annotation.
Once the variant-annotation of your file(s) is(are) complete, you are now ready to run your first AlloPipe run!
If you want to take into account frameshift neoantigens peptide generation in the af-AMS, you need to add the Frameshift plugin from the pVACtools software to your VEP installation.
mv Frameshift.pm ~/.vep/Plugins
You should then add these options to the VEP command:
--plugin Frameshift --dir_plugins <PLUGIN-PATH>
with <PLUGIN-PATH> being the path to your VEP plugins directory.
The Nextflow workflow runs the two AlloPipe modules in sequence: Allo-Count, then Allo-Affinity. It can process either one donor/recipient pair (--mode pair) or several pairs from a cohort/joint VCF (--mode cohort).
Run commands from the root of the AlloPipe directory.
By default, each new run must use a unique output destination. The directory formed by --output_dir and --run_name (<output_dir>/runs/<run_name>) must not already exist; if it does, AlloPipe stops before launching the workflow. To continue an interrupted run, launch the same command with Nextflow -resume. To start over instead, use a different --run_name or --output_dir, or pass --force_overwrite to remove and replace that run directory explicitly.
VEP-annotated VCFs are published under <output_dir>/runs/<run_name>/vcf_vep. In cohort mode, extracted VCFs and their .tbi indexes are published under <output_dir>/runs/<run_name>/vcf_indiv. Each completed Allo-Count and Allo-Affinity task also publishes its files progressively into the same run directory, so completed upstream and per-pair results remain available if a downstream task fails or the workflow is interrupted.
Use pair mode when you have one VCF file for the donor and one VCF file for the recipient. Because pair is the default value of --mode, passing --mode pair is optional.
| Argument | Required | Description |
|---|---|---|
--mode pair |
no | Selects single-pair mode. Default: pair. |
--donor <VCF> |
yes | Donor VCF file (.vcf or .vcf.gz). |
--recipient <VCF> |
yes | Recipient VCF file (.vcf or .vcf.gz). |
--run_name <NAME> |
yes | Name used for the output run directory. An existing directory is accepted with Nextflow -resume, or replaced when --force_overwrite is passed. |
--orientation dr or --orientation rd |
yes | Direction of the mismatch comparison: dr for donor-to-recipient, rd for recipient-to-donor. |
--imputation imputation or --imputation no-imputation |
yes | Missing genotype handling mode. Individual VCFs usually use imputation. |
--ensembl_path <DIR> |
yes | Ensembl data directory used by Allo-Affinity. The path must contain GRCh37 or GRCh38 so the workflow can automatically select the VEP assembly. |
--hla_typing <HLA_LIST> |
yes | Comma-separated HLA typing used by Allo-Affinity. |
Example:
nextflow run main.nf -profile conda \
--mode pair \
--donor tutorial/HG002-VEPannotated.vcf \
--recipient tutorial/HG007-VEPannotated.vcf \
--run_name test_pair \
--orientation dr \
--imputation imputation \
--ensembl_path data/Ensembl/GRCh38 \
--hla_typing "HLA-A*01:01,HLA-A*02:01,HLA-B*08:01,HLA-B*27:05,HLA-C*01:02,HLA-C*07:01" \
--skip_vep_annotation true
Use --mode cohort when several donor/recipient pairs must be extracted from one joint multi-sample VCF.
| Argument | Required | Description |
|---|---|---|
--mode cohort |
yes | Selects cohort mode. Required when running a cohort, because the default mode is pair. |
--multi_vcf <VCF> |
yes | Joint multi-sample VCF containing all donors and recipients. |
--pairs <CSV> |
yes | CSV file describing the donor/recipient pairs and per-pair HLA typing. |
--run_name <NAME> |
yes | Name used for the output run directory. An existing directory is accepted with Nextflow -resume, or replaced when --force_overwrite is passed. |
--orientation dr or --orientation rd |
yes | Direction of the mismatch comparison for all pairs in the run. |
--imputation imputation or --imputation no-imputation |
yes | Missing genotype handling mode. Joint VCF cohort runs usually use no-imputation. |
--ensembl_path <DIR> |
yes | Ensembl data directory used by Allo-Affinity. The path must contain GRCh37 or GRCh38 so the workflow can select the VEP assembly. |
The cohort CSV must contain at least the columns donor, recipient, and hla. The hla column is mandatory and must be the last column because HLA values are comma-separated.
Pairs are identified as P1, P01, P001, etc., with zero-padding adapted to the cohort size, and displayed in process tags as PAIR 1, PAIR 01, PAIR 001, etc.
Example:
nextflow run main.nf -profile conda \
--mode cohort \
--multi_vcf tutorial/HG002-HG007-VEPannotated.vcf \
--pairs tutorial/example.csv \
--run_name test_cohort \
--orientation dr \
--imputation no-imputation \
--ensembl_path data/Ensembl/GRCh38 \
--skip_vep_annotation true
| Argument | Description |
|---|---|
--output_dir <DIR> |
Base output directory. Default: output. |
--force_overwrite |
Remove <output_dir>/runs/<run_name> before launching when it already exists. When combined with -resume, the published run directory is removed while Nextflow's cached tasks remain available. |
--skip_vep_annotation true |
Skip the built-in VEP annotation step when the input VCFs are already VEP-annotated. |
--vep_cache <DIR> |
VEP cache path inside the execution environment. Default: /cache. |
--vep_version <TAG> |
VEP container tag used by the VEP process. Default: release_113.4. To change the workflow default, edit params.vep_version in nextflow.config. |
--frameshift true |
Enable frameshift neoantigen handling. Requires VEP Frameshift plugin annotation. |
--frameshift_plugin_path <PATH> |
Path to the VEP Frameshift.pm plugin or plugin directory. |
--allo_count_opts "<OPTIONS>" |
Extra options passed to the Allo-Count step, for example filtering thresholds. |
--allo_affinity_opts "<OPTIONS>" |
Extra options passed to the Allo-Affinity step, for example --length, --el_rank, --class_type, --cleavage, or --dry_run to skip the long NetMHCpan computation. |
Which parameters Allo-Count considers?
From variant-annotated .VCF file(s), data are first reformatted to obtain one data frame per individual.
Those data frames are then filtered considering a set of quality metrics (defaults values):
- minimal depth per position (20x)
- maximal depth per position (400x)
- minimal allelic depth (5x)
- homozygosity threshold (0.2)
- GnomADe allele frequency threshold (0.01)
- genotype quality threshold (0: you might adjust this value according to your sequencing platform)
- maximal length for insertions or deletions (indels, 3)
The curated data frames are then queried to assess the directional mismatches between samples.
Directional comparison
The sample comparison is directional and accounts for either polymorphisms that are present in the donor but absent in the recipient (donor-to-recipient) or that are present in the recipient but absent in the donor (recipient-to-donor).
- Donor-to-recipient accounts for polymorphisms present by the donor but absent by the recipient, i.e. triggering the recipient's immune system after solid organ transplantation.
- Recipient-to-donor accounts for polymorphisms present by the recipient but absent by the donor, i.e. triggering the donor's immune system after allogenic hematopoietic cell transplantation.
How does AlloPipe handle missing data?
We provide the possibility to impute genotype missing data as being ref/ref (e.g. 0/0 or homozygous on the nucleotide of reference.
-
If you are using individual
.VCFfiles as input ('single pair mode'), you most probably want to run with theimputationargument as ref/ref variants are omitted in those files. -
If you are using a joint
.VCFthrough the Nextflow cohort workflow, running with theno-imputationargument will only keep variants sequenced in the two datasets of each pair.
Once variant annotation is complete, run Allo-Count through the Nextflow pair workflow from the root of the AlloPipe directory. The workflow also launches the Allo-Affinity step, so --hla_typing and --ensembl_path are required.
nextflow run main.nf -profile conda \
--mode pair \
--donor tutorial/HG002-VEPannotated.vcf \
--recipient tutorial/HG007-VEPannotated.vcf \
--run_name test_pair \
--orientation dr \
--imputation imputation \
--ensembl_path data/Ensembl/GRCh38 \
--hla_typing "HLA-A*01:01,HLA-A*02:01,HLA-B*08:01,HLA-B*27:05,HLA-C*01:02,HLA-C*07:01" \
--skip_vep_annotation true
More detailed help can be obtained with nextflow run main.nf --help.
For instance, you can generate frameshift neoantigen candidates by adding --frameshift true to the Nextflow command. Note: you need to have installed the VEP Frameshift plugin.
Multiple-pair execution is handled by the Nextflow cohort workflow. Use an annotated or unannotated joint .VCF file containing the genomic data of interest and provide a CSV file specifying donor/recipient pairs.
Only one directional comparison is accepted within the same command line.
nextflow run main.nf -profile conda \
--mode cohort \
--multi_vcf tutorial/HG002-HG007-VEPannotated.vcf \
--pairs tutorial/example.csv \
--run_name test_cohort \
--orientation dr \
--imputation no-imputation \
--ensembl_path data/Ensembl/GRCh38 \
--skip_vep_annotation true
Again, more detailed help can be obtained with nextflow run main.nf --help.
After the run is complete, have a look at the output/runs/NAME-RUN/ directory that was created.
The directory is structured as followed :
- the
AMS/subdirectory contains the AMS value(s) - the
plots/subdirectory contains visual output - the
run_tables/subdirectory contains the tables created during the run.
In the run_tables/ directory, you can find:
1) D0-TABLE and R0-TABLE:
The D0/R0 tables are tab delimited files that summarize the genotype information contained in the .VCF file(s), whether individual or joint.
They can be used to navigate through this data in a more simple way, by opening them with a spreadsheet software.
2) The MISMATCHES-TABLE:
This table gives you information on the mismatched positions. For each type of information (VCF, Sample, VEP, AlloPipe results), the columns names are the following (types given in parenthesis)
- VCF information:
- CHROM (str): Chromosome of the variant
- POS (int): Position on the chromosome
- ID_{x, y} (str): Reference SNP cluster ID for the donor (x) or recipient (y)
- REF, ALT (str): REF and ALT alleles at the given position
- QUAL_{x, y} (float: Phred-scaled quality score for the assertion made in ALT
- FILTER_{x, y} (str): PASS if this position has passed all filters
- FORMAT_{x, y} (list): Format of the sample column post AlloPipe processing
- Sample_{x, y} (str): Sample information regarding the position. Note that the column name is the one provided in the original
.VCF- In the case of transplantation,
Sample_xis the donor andSample_yis the recipient
- In the case of transplantation,
- Sample information:
- GT_{x, y} (str): Predicted genotype of the sample
- GQ_{x, y} (float): Score of quality of the predicted genotype
- AD_{x, y} (str): Allelic depth
- FT_{x, y} (str): Sample genotype filter indicating if this genotype was “called”
- phased_{x, y} (str): Predicted genotype containing phased information (if provided in the sample column)
- DP_{x, y} (int): Sequencing Depth at position
- TYPE_{x, y} (str): type of genotype (homozygous, heterozygous)
- VEP information:
- consequences_{x, y} (int): Count of each consequence type (i.e. frameshift indel, missense variant, ...)
- transcripts_{x, y} (str): Transcripts recorded for the variant
- genes_{x, y} (str): Genes recorded for the variant
- aa_REF, aa_ALT (str): Amino-acid for REF and ALT alleles for the variant
- gnomADe_AF_{x, y} (float): Frequency of existing variant in gnomAD exomes combined population
- Frameshift_sequence_{x, y} (str): Frameshift sequences annotated by VEP (if
--frameshiftis activated, empty otherwise) - aa_ref_indiv_{x, y}, aa_alt_indiv_{x, y} (str): REF and ALT amino acids recorded for the sample (x and y)
- aa_indiv_{x, y} (str): REF and ALT amino acids combined in one column
- AlloPipe information:
- diff (str): difference between the amino acids of both samples
- mismatch (int): number of mismatches in the diff field
- mismatch_type (str): type of mismatch (homozygous, heterozygous)
3) The TRANSCRIPTS-TABLE:
This table contains mandatory data to perform the reconstruct peptides in the second step
What does the Allo-Affinity tool do?
From previously generated files that are the MISMATCHES-TABLE and the TRANSCRIPTS-TABLE, Allo-Affinity reconstructs the set of peptides that are different between the donor and the recipient. All the peptides of a given length (defined by the user) are generated around mismatch position using the principle of a sliding window.
The directionality of the mismatch is kept, meaning that:
- if Allo-Count has been run in the donor-to-recipient direction (
dr), only peptides exhibiting a polymorphism present by the donor but absent from the recipient will be reconstructed. - In the same way, if Allo-Count has been run within the recipient-to-donor direction (
rd), only peptides exhibiting a polymorphism present by the recipient but absent from the donor will be reconstructed.
To reconstruct the peptides, you will need the following files (replace XXX by the version of Ensembl used by VEP in the following links):
Homo_sapiens.<REFERENCE-GENOME>.cdna.all.fa.gz: https://ftp.ensembl.org/pub/release-XXX/fasta/homo_sapiens/cdnaHomo_sapiens.<REFERENCE-GENOME>.pep.fa.gz: https://ftp.ensembl.org/pub/release-XXX/fasta/homo_sapiens/pepHomo_sapiens.<REFERENCE-GENOME>.<VEP-VERSION>.refseq.tsv.gz: https://ftp.ensembl.org/pub/release-XXX/tsv/homo_sapiens/Please be aware the number of the Ensembl release has to be the same as the one used by the VEP tool version that generated the annotated VCF.
Do not forget to select the reference genome used to perform the alignment. We provide the v111 of those files for GRCh37 and GRCh38 here.
Before running a Nextflow command that launches Allo-Affinity, unzip the files corresponding to your assembly (GRCh37 or GRCh38):
gzip -d data/Ensembl/GRCh38/*
Allo-Affinity output those peptides in a fasta file that can be processed by the following third party softwares:
Each of these tool imputes the affinity of the reconstructed peptides towards the HLA peptide grooves, therefore outputs candidate minor histocompatibility antigens (mHAgs). Please note that the HLA typing has to be known before running the command line, as the AlloPipe tool does not impute the HLA typing from genomic data.
Hint: You can use nfcore-HLAtyping to assess the HLA class I from exome data.
Allo-Affinity is run automatically after Allo-Count by the Nextflow workflow. Some parameters are handled directly by Nextflow, while others are still passed to the Python Allo-Affinity step through --allo_affinity_opts.
nextflow run main.nf -profile conda \
--mode pair \
--donor tutorial/HG002-VEPannotated.vcf \
--recipient tutorial/HG007-VEPannotated.vcf \
--run_name test_pair_affinity \
--orientation dr \
--imputation imputation \
--ensembl_path data/Ensembl/GRCh38 \
--hla_typing "HLA-A*01:01,HLA-A*02:01,HLA-B*08:01,HLA-B*27:05,HLA-C*01:02,HLA-C*07:01" \
--skip_vep_annotation true \
Nextflow arguments:
--run_nameis the unique name of the run--ensembl_pathis the path to the Ensembl files previously downloaded:.cdna.fa,.pep.faand.refseq.tsv--hla_typingis the HLA typing, e.g.HLA-A*01:01,HLA-A*02:01,HLA-B*08:01,HLA-B*27:05,HLA-C*01:02,HLA-C*07:01
Python Allo-Affinity arguments passed through --allo_affinity_opts:
--lengthis the length of peptides to be imputed (default: 9 for class I)--el_rankis the elution threshold (default: 2 for class I)--dry_runskips the long NetMHCpan computation
For example:
--allo_affinity_opts "--length --el_rank"
Note on providing HLA typing
Allo-Affinity lets you be flexible in providing the HLA alleles used for typing, as long as they can be set as parameters for the affinity prediction program (NetMHCpan for the moment). In pair mode, pass them with
--hla_typing. In cohort mode, provide them per pair in the mandatoryhlacolumn of the cohort CSV. In most scenarios you should provide both HLA alleles to compute affinity, but it is perfectly possible to provide for instance only one allele (as could be the case for bone marrow transplantation).
Multiple-pair Allo-Affinity execution is also handled by the Nextflow cohort workflow. Provide the same pair list as the Allo-Count step, with a per-pair hla column containing the HLA typing to use for each donor/recipient pair.
This second step of AlloPipe uses the AMS information of the first step.
You will find 3 new subdirectories in the output/runs/<run_name>/ directory:
- the
AAMS/directory contains a subdirectory created for these run parameters specifically, the AAMS value contained in a.csvfile. - the
netMHCpan_out/subdirectory contains all tables generated during the NetMHCpan step. - the
aams_run_tables/subdirectory contains all the other tables created during the run
If you want more in-depth information on the mismatches contributing to the AAMS, you will find a mismatches table in the aams_run_tables/ directory.
It contains the mismatches information from the AMS run along with information provided by NetMHCpan :
- NetMHCpan information
- hla_peptides (str): Potential ligand peptide built from VEP information and Ensembl information
- Gene_id (str): Ensembl Gene ID
- NB (int): Number of Weak Binding/Strong Binding peptides accross given HLA
- EL-score (float): Raw prediction score
- EL_Rank (float): Rank of the predicted EL-score compared to a set of random natural peptides
- BA-score (float): Binding-Affinity score
- BA_Rank (float): Rank of the predicted BA-score
- HLA (str): Specified MHC molecule / Allele name
- Transcript_id (str): Ensembl Transcript ID
- Peptide_id (str): Ensembl Peptide ID
AlloPipe can also run the NetChop tool to annotate the potential proteasomal cleavage sites on the proteins that contain mismatch. This then give you a reduced set of candidate peptides that you can compare with their affinity values.
The cleaved sites are predicted on a protein sequence which depends of the directionality of the run:
drdirection: Proteins reconstructed from the genotype of the donor.rddirection: Proteins reconstructed from the genotype of the recipient.
- Cleaved peptide information
More information about the cleaved peptide is available in the netChop/ directory in the netchop_table.csv file, which contains the following information.
- CHROM (str): Chromosome of the variant
- POS (int): Position on the chromosome
- Protein_position (str): Position on the protein
- Gene_id (str): Ensembl Gene ID
- Transcript_id (str): Ensembl Transcript ID
- Peptide_id (str): Ensembl Peptide ID
- Sequence_aa (str): Amino acid sequence of the peptide
- aa_REF (str): Amino acid for REF
- aa_ALT (str): Amino acid for ALT
- peptide_ALT (str): Amino acid sequence of the peptide with mutation(s)
Each row of the table corresponds to a cleaved peptide on a protein that contributes to a mismatch in the AMS.
We provide a couple of example data in /tutorial, i.e. tutorial/donor_to_annotate.vcf and tutorial/recipient_to_annotate.vcf (those files correspond to human chr6).
To test your VEP installation (v111 in this tutorial), run the following commands:
vep --fork 4 --cache --assembly GRCh38 --offline --af_gnomade -i tutorial/donor_to_annotate.vcf -o tutorial/donor_annotated_vep111.vcf --coding_only --pick_allele --use_given_ref --vcf
vep --fork 4 --cache --assembly GRCh38 --offline --af_gnomade -i tutorial/recipient_to_annotate.vcf -o tutorial/recipient_annotated_vep111.vcf --coding_only --pick_allele --use_given_ref --vcf
Once the VEP annotation is complete, go to the root of the AlloPipe directory to run the workflow:
nextflow run main.nf -profile conda \
--mode pair \
--donor tutorial/HG002-VEPannotated.vcf \
--recipient tutorial/HG007-VEPannotated.vcf \
--run_name test-run \
--orientation rd \
--imputation no-imputation \
--ensembl_path data/Ensembl/GRCh38 \
--hla_typing "HLA-A*01:01,HLA-A*02:01,HLA-B*08:01,HLA-B*27:05,HLA-C*01:02,HLA-C*07:01" \
--skip_vep_annotation true
The expected AMS are:
| Orientation | Imputation | No imputation |
|---|---|---|
HSCT = rd |
2812 | 42 |
SOT = dr |
1155 | 34 |
The same Nextflow command also runs Allo-Affinity and writes the af-AMS and related tables in the output run directory.
If you want to run the cleaved peptide prediction, add --allo_affinity_opts="--cleavage" to the Nextflow command.
You can now enjoy AlloPipe. If you have any feedback, please get in touch, we will be happy to help!
