Repository navigation
Designing assays
primer-finder find reports regions. primer-finder design turns them into assays: it runs Primer3
on the regions of a finished run, works out what would make each assay selective, and writes the primer
files that insilicoPCR reads — optionally running it as well, to
check the assays against both groups of genomes.
primer-finder design results/ -o assays/ -t 16
primer-finder design results/ -o assays/ -t 16 --insilico-pcr /path/to/insilicoPCR-linux-x64The genomes come from what the run recorded in run_info.json, so the two folders need not be given again;
-i and -e override that if they have moved.
A region is inclusion-specific as a whole: every kmer of it is absent from the exclusion genomes. An assay designed inside it is not specific by itself. Only two things make it so, and the command tells them apart:
- Absence. The amplicon does not exist in any exclusion genome, so there is nothing to amplify. This is the strongest case and it is how the published assays this tool is validated against work — their oligos cover no differing base at all; the stretch they sit in is simply missing from the other subspecies.
-
Differences. The amplicon does exist there, but an oligo sits on bases that differ — the lower-case
bases of
final_kmers.fasta. The assay then depends on a mismatch stopping the reaction.
An assay that is neither — amplicon present, no oligo on a difference — would amplify both groups. Those are
listed in assays.tsv with specific_by=nothing and go no further.
Three kinds of request go in for each region, and the results are pooled and de-duplicated:
| Request | What it asks for | Why |
|---|---|---|
| plain | the best assay anywhere in the region | the region may be absent from the exclusion genomes, in which case any assay in it works |
SEQUENCE_FORCE_LEFT_END / SEQUENCE_FORCE_RIGHT_END
|
a primer whose 3' end sits on the last (or first) difference of a target | allele-specific PCR: a mismatch at the 3' end hinders extension, and the rest of the target falls under the primer |
SEQUENCE_INTERNAL_OVERLAP_JUNCTION_LIST |
a probe straddling the middle of a target | a probe that cannot bind the exclusion template |
A target is a stretch of at most 25 bases — the widest primer Primer3 may return — holding as many differences as possible. The four richest, non-overlapping ones are aimed at. See below for why that is the unit rather than a run of consecutive differences.
The forced requests often come back empty, which is not an error: no oligo of the required length and melting temperature may fit there. A forced request cannot produce a bad oligo — the constraints below are part of every request, forced or not, apart from the GC clamp (which cannot be, see below), and Primer3 returns nothing rather than something outside them.
This is the point the whole pipeline turns on, so it is worth stating plainly. primer-finder find keeps a
region when its differences from an exclusion genome could sit in one oligo — a run of them, or two of
them fewer than 21 matching bases apart, 21 being about a primer's length
(Methods). A region is selected for that and nothing else; the whole
reason it is a candidate is that an oligo can be placed to carry two or more mismatches at once.
Drawn — a region with two differences seven bases apart, which is not a run but is well within one oligo:
region ··········G······T······················
target └──────┘ 8 bases, 2 differences
left primer ▶ its 3' end on the SECOND difference, so both sit under the primer
right primer ◀ its 3' end on the FIRST, since it reads the other way
probe ╳ straddling the middle of the two
Up to 1.3.0 there was no target here at all: two differences seven apart are two runs of one, so a request
pinned a 3' end on G or on T, and whichever primer came back carried a single mismatch.
So the design step has to actually do it, and two things make sure of it:
- The forced requests aim at a window, not at a run. Two differences seven bases apart are not consecutive, so a run-based target would see two runs of one and pin a primer's 3' end on a single base, wasting the pair the region was kept for. On the Xylella validation set that is the usual case: 19 of the 20 regions designed on hold such a pair, and 8 hold nothing else. Aiming at any 25-base window instead raised the share of difference-based assays whose best single oligo covers two or more differences from 43% to 77%, on the same regions with everything else unchanged.
-
--min-primer-differences, 2 by default, sets aside the rest. An assay that rests on differences and spends only one of them is the weak case thefindrule exists to avoid — a single mismatch, even at a 3' end, often does not stop amplification (Lefever et al. 2019). Those assays are still listed, withbest_primer_variantssaying how many the better primer carries, but they are not written to the insilicoPCR files and not carried further.--min-primer-differences 1keeps them, which is a weaker assay rather than no assay.
An assay that is specific by absence is exempt: it does not rest on a difference at all, so how many sit under its primers is beside the point.
--min-primer-differences looks at the primers only, however many differences the probe covers. That is
deliberate, and it is the one place where the reasoning behind this filter does not carry over from a primer
to an oligo in general.
A probe here is longer than a primer — 18 to 27 bases against 18 to 25, 22 optimal against 20 — and its melting temperature is higher, 62–72 °C against 58–63 °C, because the extra length is how that temperature is reached. In a reaction annealing near 60 °C it therefore starts with far more binding energy in hand, and two mismatches spread over a long duplex need not stop it hybridising or being cleaved. Three findings say so more precisely than the argument does:
- A conventional TaqMan probe still gave a detectable signal through five mismatches, and under standard conditions neither it nor an MGB probe was sequence-specific (Yao, Nellåker & Karlsson 2006). Whatever a small number of mismatches under a probe is evidence of, it is not evidence of discrimination.
- A probe discriminates by being short, not by carrying more mismatches. A 12-base MGB probe has the same melting temperature, 65 °C, as an unmodified 27-base one (Kutyavin et al. 2000, Nucleic Acids Res 28:655–661, PMC102528). Shortening the duplex is what makes one mismatch a large fraction of its stability, which is why SNP genotyping uses short MGB or LNA probes with the difference near the middle.
- In a 5'-nuclease assay the discrimination lives in the primers. A single mismatch in a primer's 3' region shifts the quantification cycle by anything from under 1.5 to over 7 cycles depending on which base pair it is, and by up to sevenfold between master mixes (Stadhouders et al. 2010; see also Lefever et al. 2013, Clin Chem 59:1470–1480). The probe reports the amplification; it does not gate it.
So an assay whose only differences sit under its probe is set aside like any other that no primer can tell
apart. Its probe differences are still counted in probe_variants, and they still break ties in the
ranking (step 8), but they cannot qualify it.
If the probe is meant to do the discriminating, primer-finder cannot give you what that needs: order it as a shortened MGB or LNA probe with the difference near its centre, and raise the annealing temperature. What the design step can do is stop presenting a 27-mer as though it were specific.
- Yao et al. is one study, one target, from 2006, under "standard conditions". "A detectable signal" is not the same as "indistinguishable": a probe mismatch that delays the cycle by several cycles may still discriminate if you set a threshold for it and validate that threshold.
- Every number above is condition-dependent. Stadhouders' sevenfold spread between master mixes is the point rather than a footnote: the same oligo behaves differently in a different mix, at a different annealing temperature, with a different polymerase.
- In silico PCR cannot settle it either way, in either direction. It decides whether an oligo binds by counting mismatches, so it will report a probe with two mismatches as not binding — which is exactly the assumption in question. In the Xylella sweep every one of the 58 probe-only assays came out "selective" in qPCR mode and none of them in PCR mode; that gap was the model trusting the probe, not evidence about it.
So this is a reason to stop treating two mismatches under a probe as qualification, not a demonstration that
such a probe never discriminates. --min-primer-differences 1 accepts those assays if you would rather
judge them yourself.
On the Xylella set the two changes together take the 511 assays Primer3 proposed to 363 carried forward (195 by absence, 168 by difference), setting 141 aside and dropping the 7 that would amplify both groups. The ranking is unchanged: it still prefers more differences in total, and does not take a view on whether two under one primer beat one under each.
These are hard constraints, not score terms: Primer3 returns nothing rather than an oligo outside them, so a doomed assay is never proposed and never reaches the ranking. The ranking below only ever orders assays that already satisfy all of this.
| Primer | Probe | |
|---|---|---|
| Length | 18–25 nt, 20 optimal | 18–27 nt, 22 optimal |
| Melting temperature | 58–63 °C, 60 optimal | 62–72 °C, 68 optimal |
| GC | 30–70% | 30–80% |
| Ambiguous bases | none | none |
| G or C at the 3' end |
--gc-clamp, 1 by default, except in a request that pins an end |
— (Primer3 has no clamp for the probe) |
| Hairpin | melts below --max-hairpin-tm, 47 °C by default |
same |
| Self-dimer, whole oligo and 3' end | below --max-dimer-tm, 47 °C by default |
same |
| Dimer with the other primer, whole oligo and 3' end | below --max-dimer-tm
|
— |
A G or a C at the 3'-most base holds the primer down where extension starts, so one is asked for by default
(PRIMER_GC_CLAMP=1); --gc-clamp 2 asks for two, --gc-clamp 0 for none.
The clamp is dropped from the requests that pin a primer's 3' end on a differing base. Those are the
SEQUENCE_FORCE_LEFT_END / SEQUENCE_FORCE_RIGHT_END requests, where the 3' base is whatever the genomes
made it — asking for a G or a C there is asking for a difference that may not exist. On an A/T-rich target
the two requirements are almost never satisfiable at once: with the clamp left on, seven of eight forced-end
requests on the Xylella set returned nothing.
PRIMER_GC_CLAMP is a single tag for a whole request, with no per-side or per-oligo variant, so dropping it
drops it from both primers of such a request — the partner, whose end Primer3 was free to choose, loses
it too — and --gc-clamp therefore reaches only the requests that pin nothing and the one that pins the
probe. There is no way to ask Primer3 for "a clamp on the partner but not on the pinned primer". The clamp
never applies to the probe either: Primer3 has no PRIMER_INTERNAL_GC_CLAMP.
An oligo that folds back on itself, pairs with a copy of itself, or pairs with its partner is spent before it
ever reaches the template — and a 3'-end dimer is worse than an internal one, because the polymerase can
extend it. Primer3 models all of these thermodynamically (PRIMER_MAX_HAIRPIN_TH, PRIMER_MAX_SELF_ANY_TH
and _SELF_END_TH, PRIMER_PAIR_MAX_COMPL_ANY_TH and _COMPL_END_TH, and the PRIMER_INTERNAL_*
equivalents for the probe), and the design step sets every one of them, for the probe as well as the primers.
Nothing extra has to be installed: this is Primer3's own thermodynamic alignment, so there is one melting
temperature model for the oligo and for the structures it might form.
The default limit of 47 °C is Primer3's own, which is roughly 10 °C below the annealing temperature these
oligos are designed for: a structure that has melted by the time the reaction anneals does not compete with
the template. Lower it to be stricter (--max-hairpin-tm 40), raise it on a target where nothing else will
fit — at the cost of candidates that may fold.
What Primer3 predicted is also reported, so a reaction run under other conditions can be judged:
pair_dimer_tm and pair_dimer_end_tm per assay, and forward_hairpin_tm, forward_self_dimer_tm and the
reverse_ and probe_ equivalents per oligo. A 0.0 means no structure was predicted, not that none was
looked for.
A melting temperature is only meaningful for a given reaction: salt stabilises a duplex, magnesium more so per mole, dNTPs chelate magnesium away, and oligo concentration shifts the equilibrium. Primer3 is therefore told what reaction to predict for, and the defaults are an ordinary TaqMan qPCR:
| Option | Default | Primer3 tag |
|---|---|---|
--monovalent |
50 mM | PRIMER_SALT_MONOVALENT |
--divalent |
3 mM Mg²⁺ | PRIMER_SALT_DIVALENT |
--dntp |
0.8 mM total (0.2 mM each) | PRIMER_DNTP_CONC |
--primer-conc |
250 nM | PRIMER_DNA_CONC |
--probe-conc |
200 nM | PRIMER_INTERNAL_DNA_CONC |
The probe gets its own concentration, since it is normally used below the primers. These are not cosmetic: they change both the temperatures reported and which oligos come back at all. Measured on the first assay of the Xylella run, the same oligos read:
| bare Primer3 defaults | the reaction above | its floor here | |
|---|---|---|---|
| its forward primer | 56.1 °C | 60.0 °C | 58 °C |
| its probe | 54.6 °C | 64.0 °C | 62 °C |
Either oligo is rejected under the bare defaults and accepted under the stated reaction, so this is not a cosmetic difference in a reported number: it decides what Primer3 returns.
Set them to your own master mix if it differs; the values used are recorded in design_info.json under
parameters.conditions, so a table of assays can always be traced back to the reaction it was designed for.
Changing them does not change the oligos' behaviour at the bench, only which ones this step proposes.
The order is a heuristic for which assays to look at first, not a prediction of what will work at the bench. Nothing here has been tested in a laboratory. Read it as "these are the ones worth trying", and let in silico PCR and then the bench decide.
Assays are ordered by this key, each step breaking the ties of the one before:
- Absence beats everything. An assay whose amplicon no exclusion genome holds has nothing to amplify there. This is the one step that does not depend on how a reaction behaves.
- How many differences the primers cover, in total.
- What those differences replace, weighted: a position where the exclusion genomes have a G or a C counts 1, an A or a T counts 0.5, and one whose base is not known counts 0.75. G and C pair with three hydrogen bonds against two, so a mismatch there costs more to form. This only ever separates assays that cover the same number of differences.
- The longest run of differences ending at a primer's 3' end.
- Differences within five bases of a 3' end, run or not.
- Primer3's pair penalty, by band (<1, <2, <4, worse), so that a probe covering more differences never wins over an assay that is clearly better made.
- Copies of the amplicon in the inclusion genomes, since a repeated target usually gives a better limit of detection. See below.
- Differences under the probe, then the exact penalty, then the names, so that a run is reproducible.
Nine assays, all with the same Primer3 penalty so that only the differences separate them. Each row is a
20-base forward primer written 5' to 3', with ▶ for the 3' end where the polymerase extends:
· a base the exclusion genomes share with the region — nothing for an oligo to discriminate on
G a base where they differ, written as what THEY have there: the base the primer would mismatch
▶ the 3' end of the primer ◀ the 3' end of the reverse primer, which reads the other way
In the order the ranking puts them. The two marked no are set aside before the order matters, so
assays.tsv writes them at the end of the file whatever their rank:
| the primer | diffs | weight | 3' run | kept? | |
|---|---|---|---|---|---|
| 1 | the amplicon is missing from every exclusion genome | — | — | — | yes |
| 2 | ·················GCG▶ |
3 | 3.0 | 3 | yes |
| 3 | ··················GC▶ |
2 | 2.0 | 2 | yes |
| 4 |
···················G▶ and ◀C···················
|
2 | 2.0 | 1 | no |
| 5 | ················G·C·▶ |
2 | 2.0 | 0 | yes |
| 6 | ·····G······C·······▶ |
2 | 2.0 | 0 | yes |
| 7 | ··················AT▶ |
2 | 1.0 | 2 | yes |
| 8 | ···················G▶ |
1 | 1.0 | 1 | no |
| 9 |
no difference under either primer; two under the probe: ··········GC··········
|
0 | 0 | 0 | no |
Reading the ladder:
- 1 beats everything. Nothing to amplify is better than something hard to amplify, and it is the only step that does not depend on how a reaction behaves.
- 2 beats 3 on the count alone: three differences under one primer rather than two.
- 3 beats 7 — the one your eye should go to. Both are a run of two at the 3' end; 3's are over a G and a C, 7's over an A and a T. Breaking G:C costs three hydrogen bonds against two, so at an equal count the G/C pair is worth more (step 3 of the key).
- 5 and 6 also beat 7, which is the step order showing its hand: what the differences replace (step 3) is weighed before where they sit (step 4). Two G/C differences in the middle of a primer therefore outrank two A/T differences at its 3' end. 5 beats 6 because its differences are within five bases of the 3' end even though neither reaches it (step 5).
- 9 comes last, and is set aside: its primers cannot tell the groups apart at all. Two differences under a probe are not evidence that the probe will not bind — see why the probe does not count.
Rows 4, 8 and 9 are listed in assays.tsv but carried no further — not written to the
insilicoPCR files, not among the candidates — because of
--min-primer-differences, 2 by default:
- 8 rests on a single difference. One mismatch, even at the 3' end, often does not stop amplification.
- 4 is the subtle one. It covers two differences and its rank key is high — higher than 5, 6 and 7 — but they are one each on two different primers, so neither primer carries two. The region was kept because two differences could sit in one oligo, and this assay does not do that. The filter is applied before the order matters, so a high rank does not save it.
- 9 has two differences under one oligo, but that oligo is the probe, which does not count: here is why.
All three are kept in the table with best_primer_variants saying what they rest on, and
--min-primer-differences 1 accepts them — a weaker assay rather than no assay.
Row 1 is exempt from the filter: an assay specific by absence rests on no difference at all, so how many sit under its primers is beside the point.
Steps 2 to 4 follow common practice in allele-specific design rather than any measurement made here:
- a single mismatch, even at the 3' end, is often not enough on its own. Allele-specific PCR is known for "low discriminating power", and a common remedy is to introduce a second, artificial mismatch in the primer — the basis of double-mismatch allele-specific qPCR (Lefever et al. 2019). More differences under a primer is the same idea, arrived at without having to engineer one;
- where a mismatch sits, and which bases are involved, changes its effect (Sharma et al. 2022 summarise this for PCR while characterising it for RPA). Hence two tie-breaks: the 3' end is the position that matters most for extension, and a difference replacing a G or a C is worth more than one replacing an A or a T, since G:C holds with three hydrogen bonds and A:T with two.
The bases are read from the exclusion genomes' own alignments of each amplicon, so what is weighed is what
an oligo would actually have to mismatch, genome by genome, rather than what the region happens to carry.
assays.tsv reports the weight (primer_variant_weight) and the count of G/C positions per primer
(forward_strong_variants, reverse_strong_variants), so the order can be checked.
How much a given mismatch costs depends on the assay, not only on the sequence: annealing temperature, polymerase, magnesium, cycling and template concentration all change it. A design that discriminates under one set of conditions may not under another, and optimisation at the bench can recover an assay this ranking puts low — or ruin one it puts high.
Running the designed assays through in silico PCR gives numbers that look like evidence for the order above. They are weaker than they look, and are reported here only so that nobody mistakes them for more. On 287 difference-based assays from Xylella fastidiosa subsp. multiplex regions, against 17 exclusion genomes, the share that amplified only the inclusion group rises with the number of differences under the primers:
| differences under the primers | 0 (probe only) | 1 | 2 | 3 | 4 or more |
|---|---|---|---|---|---|
| qPCR mode, selective | 43% | 36% | 60% | 64% | 94% |
| PCR mode, selective | 0% | 7% | 39% | 40% | 91% |
Three reasons not to read that as a result about PCR:
-
It is close to circular. in silico PCR decides whether a primer binds by counting mismatches against
its own tolerance (
--mismatches). More differences therefore means fewer reported hits almost by construction. The numbers describe the model's rule as much as the biology. - The model cannot see a difference at the very 3' end — the position the ranking cares most about. See the three zones below. Assays whose only differences are one or two bases at a 3' end are reported selective in none of the 65 cases seen, at any tolerance, because those bases are not counted. That is a property of the alignment, and it is the one number from an earlier version of this page that has had to be withdrawn as evidence: it says nothing either way about whether such an assay would discriminate at the bench.
- It is one organism, one dataset, one set of parameters. 25 genomes, 20 regions, one in silico tool.
So: the ranking is a sensible order to work through, supported by what others have published about
mismatches, and the in silico numbers are a consistency check on one dataset. They are not a measurement of
how these assays behave in a tube. The full sweep is in
validation/results/2026-10-08_v1.3.0/SUMMARY.md.
One thing in that record is not circular, and is worth acting on: of the 178 assays that are specific by
absence — no designed mismatch anywhere, their amplicon simply missing from the exclusion genomes — 30
amplify something else in an exclusion genome at the default tolerance, and 50 do at --mismatches 2.
Off-target amplification is a real specificity risk that has nothing to do with the ranking, and it is the
main reason to run the check at all.
insilicoPCR is a separate program and stays one. primer-finder design writes what it reads, and will run
it for you:
primer-finder design results/ -o assays/ --insilico-pcr /path/to/insilicoPCR-linux-x64--insilico-pcr takes the folder of an extracted portable release (which brings its own Java), its jar, or
a launcher script. Without it, the same thing can be run later with the run_insilico_pcr.sh the command
writes.
Two files, two runs. insilicoPCR cannot report both kinds of assay at once: one probe in a primer file
puts the whole report in qPCR mode, where an assay without a probe is never positive. So every usable assay
is written twice — assays_qpcr.fasta with its probe, assays_pcr.fasta with the primers alone — and each
file is run against both groups. That gives two answers per assay, and they are worth having separately:
-
pcr_selectivesays whether the primer pair discriminates on its own; -
qpcr_selectivesays whether the whole assay reports positive, with the probe having to bind too.
An assay is selective in a mode when it amplifies every inclusion genome and no exclusion genome.
insilicoPCR takes a mismatch tolerance, and primer-finder passes --mismatches 1 by default. What that
governs was measured rather than assumed — a primer pair in a synthetic template, the template mutated one
base at a time at a known distance from the primer's 3' end, run at tolerances from 0 to 10
(mismatch_zones.py).
A primer turns out to have three zones, and the tolerance only controls one:
| Where the mismatch is | What insilicoPCR does |
|---|---|
| the last 2 bases |
free. Not counted as a mismatch at any tolerance, including 0. blast trims an unmatched base off the end of its alignment, and insilicoPCR reports the trim as a negative EndMismatch offset and calls the primer bound. |
| 3 or 4 bases from the 3' end | never amplifies, at any tolerance — 10 was tested. A trim that long is rejected, and keeping the mismatch scores worse for blast than trimming. |
| 5 or more bases in |
what --mismatches decides, counted per primer: at --mismatches 1 each primer may carry one, so an assay may carry two. |
This is a property of the alignment, not of PCR, and there is no setting that changes it. Two consequences: a difference at the very 3' end — what the forced-end requests aim for — cannot be rewarded by this check, and a difference 3 or 4 bases in is treated as fatal when it may not be.
Why 1 and not 0, 2 or 3. --mismatches 0 assumes any mismatch five or more bases from the 3' end stops
a primer, which is the optimistic end of what is known: a single internal mismatch frequently does not stop
amplification, which is why double-mismatch allele-specific designs exist
(Lefever et al. 2019). It also hid 12 of the 30 off-target
amplifications above. --mismatches 2 assumes two mismatches per primer still bind, which is more than
that literature supports — two deliberate mismatches are what an allele-specific design uses to
discriminate — so it belongs in a second, stricter pass rather than in the default. --mismatches 3
measured nothing that 2 did not: on the Xylella set the two give identical verdicts in PCR mode, one assay
apart in qPCR mode.
primer-finder design results/ -o assays/ --insilico-pcr <dir> # the default, -m 1
primer-finder design results/ -o assays/ --insilico-pcr <dir> --mismatches 2 # the stricter checkAn assay that still passes at --mismatches 2 rests on something the model cannot explain away. Raising the
tolerance only ever finds more exclusion amplification, never less, so it can only take assays away.
Because the last two bases are free, an assay can be reported non-selective for a reason that is about the
alignment rather than about the assay. assays.tsv therefore carries qpcr_exclusion_terminal_only and
pcr_exclusion_terminal_only: of the exclusion genomes an assay amplified, how many did so only where
a difference sits in the last two bases of a primer.
Two rules decide that, and between them they make the count mean one thing:
-
only the amplicons that needed no counted mismatch are looked at. Those are reported at every
tolerance, so they are the ones the check could not have refused. An amplicon that bound through a
mismatch
--mismatchesallowed is the tolerance talking, not these bases, and is left out. This is what makes the answer the same whatever tolerance was used: on the Xylella set the same 256 exclusion amplifications are unrefusable at--mismatches0, 1, 2 and 3; - a genome counts only when none of those amplicons is an exact match. One exact match is a real amplification, and no number of trimmed ones beside it changes that.
When the count equals *_exclusion_amplified, every exclusion amplification behind the verdict is one
insilicoPCR could not have refused, and the design is one in silico PCR cannot judge in either direction.
That gets its own label rather than being lumped in with a cross-reaction: *_selective reads
undecided (3' end), and those assays are ranked above the ones this check refused on evidence it can
defend and below the ones it cleared. On the Xylella set:
| PCR mode, 465 usable assays | pcr_selective |
|
|---|---|---|
| amplifies every inclusion genome, no exclusion genome | yes |
220 |
| every exclusion amplification rests on an uncounted 3'-end difference | undecided (3' end) |
96 |
| some of them do | no |
16 |
| amplifies exclusion genomes on evidence the check can defend | no |
133 |
Of those 96, 73 rest on a difference the ranking deliberately put at a 3' end: 63 of one base and 10 of two. Cut the other way, 80 are specific by a difference and 16 by absence — in the second case the amplification is an off-target amplicon somewhere else in an exclusion genome, which itself only binds through an ignored terminal mismatch, so a weak off-target rather than a clean one. Either way: the column tells you which verdict to argue with, and 149 rather than 245 is the number of assays this check rejects on its own evidence.
The inclusion side needs no such column. Of 3,720 assay-and-genome amplifications in the inclusion group, none depended on a trimmed primer end: every inclusion genome that counted had at least one clean amplicon, which is what should happen when the assays were designed on sequence all of them share.
-p/--min-inclusion lets find keep a region that some inclusion genomes lack. An assay designed on such a
region cannot amplify genomes that do not hold it, so judging it against all of them would be unfair. The
design step reads the -p of the run out of its run_info.json and applies the same bar:
qpcr_selective / pcr_selective
|
What it means |
|---|---|
yes |
amplifies every inclusion genome and no exclusion genome |
partial (88%) |
amplifies 88% of the inclusion genomes — at or above the run's -p — and no exclusion genome |
undecided (3' end) |
covers the inclusion group, but it amplifies exclusion genomes only where a difference sits in the last two bases of a primer, which this check cannot see. undecided (3' end, 88%) when the inclusion coverage is partial as well. What that means
|
no |
below the threshold, or it amplifies an exclusion genome on evidence the check can defend |
The percentage is also in qpcr_inclusion_percent and pcr_inclusion_percent, and is rounded down so that
it never overstates the coverage. Amplifying an exclusion genome is never excused, whatever -p was. Assays
that amplify every inclusion genome are listed before those that merely meet the threshold.
A target present several times in each genome gives more template per cell, which usually improves the limit
of detection. assays.tsv reports inclusion_copies_min and inclusion_copies_max, counted by blasting
each amplicon against the inclusion genomes, and the ranking prefers more copies.
The default -d 1 of the find command discards repeated regions before any of this. -d caps how
often a kmer may occur per inclusion genome, and KMC counts over the whole group: with N genomes and
-d 1, a kmer occurring twice in each genome has a count of 2N, above the -cx N limit, and is dropped.
Verified with two genomes holding a region twice: -d 1 returns only the single-copy kmers, -d 2 returns
both. So if a multi-copy target is what you want:
primer-finder -i inclusion/ -e exclusion/ -o results/ -d 2 # or more
primer-finder design results/ -o assays/and look at inclusion_copies_min in the table. The trade-off is that -d above 1 also lets through kmers
that occur several times inside a single genome, which is why it is not the default.
| File | What it holds |
|---|---|
assays.tsv |
every assay, the usable ones first, with the numbers behind the order (see Outputs) |
assays_qpcr.fasta |
the assays that have a probe, for insilicoPCR in qPCR mode |
assays_pcr.fasta |
every assay, primers only, for insilicoPCR in standard PCR mode |
run_insilico_pcr.sh |
the in silico PCR of both files against both groups, to run later |
design_info.json |
the parameters, the programs and their versions, and the counts |
primer3/ |
exactly what was sent to Primer3 and what came back |
amplicons/ |
the blast databases and hits of the amplicon checks |
insilico_pcr/ |
what insilicoPCR wrote, when it was run |
- The ranking is a prediction, not a result. It orders assays by what makes them likely to work; in silico PCR is the check, and the wet lab is the answer.
- The numbers behind the ranking come from one dataset. They are from the Xylella validation set, with one organism, 25 genomes and 20 regions. They are consistent with what is known about 3' mismatches, but they are not a universal law — and in silico PCR cannot see a difference in the last two bases of a primer at all, so it can never confirm the step of the ranking that cares most about it.
-
Only the first 50 regions are designed on by default (
--max-regions), because the regions are already ordered most-promising-first and in silico PCR of thousands of assays takes a while.--max-regions 0does all of them. - The structure predictions are predictions. Primer3's thermodynamic model is the same one behind its melting temperatures; it is good enough to throw out the obvious failures, not a substitute for looking at a candidate before ordering it.
- Nothing here checks the assay against anything but these genomes: no cross-reaction with what is not in the two folders, and no check against a transcriptome or a host.