Skip to content

pipeline_detection

SlimaniF edited this page Aug 14, 2026 · 1 revision

In this section is described the detection.py script.

Algorithm functionnement

Three algorithm are exploited in the process of detecting single molecules, a detection algorithm, a deconvolution algorithm and a clustering algorithm. Here is a completed breakdown.

Detection

Description

Fluorescent single molecules appears on FISH images as bright, most often tiny, spots, the idea of this algorithm is to detect variation of intensity in regions of diameters equal to spot size parameter.

  1. A gaussian blur is first applied, filtering out intensity variation due to noise and focusing dection at a spatial scale matching the spot size.
  2. Secondly, we use a Laplacian (intensity second derivative) to detect intensity maxima.
  3. Thirdly a maximum filter finds local maxima according to the spot_size.
  4. Automatic or user selected threshold is applied on filtered image to select local maxima considered as single molecules.

For each single molecule pixel coordinates of local maxima are passed down to the software as the single molecules positions.

Selecting a threshold :
Threshold is aplied only on the image resulting from LoG filter meaning that the threshold value should NOT be based on original intensity values.*

This algorithm can have really good performance given that you are willing to set a threshold. However regions with a high density of spot such as transcription site or foci will yield only one or two spots due to the max filter. To adress this issue a second algorithm is ran and is described below.

Parameters

  • spot size : Expected size of single molecule in nanometers (zyx). This parameter is used to compute Gaussian blur and maximal filter kernel sizes. For diffraction limited spots, the same value as the voxel size can be used. Note that it is possible to use a value lower than the voxel size but will make the algorithm much more sensible to noise, in case the the spot size is too low, the software will return an error.

  • threshold : Threshold value used AFTER application of detection filters (see above). If no threshold is set (blank space) an automatic threshold will be inferred (see above).

Automatic threshold

An automatic threshold can be used to avoid user bias but the quality of this automatic threshold is strongly impacted by the quality of the signal to noise ration of your image. The automatic threshold will be used for each genes whith a blanck in threshold columns.

To compute the automatic threshold the software will plot the number of single molecules detected versus the threshold used (see below).

Correct signal to noise ratio Bad signal to noise ratio
good_elbow bad_elbow

Logically the curve is always decreasing and should hit a critical threshold where all background points are already filtered leading to an inflexion point. The automatic threshold is set by detecting the curve's inflexion point. This point can be more or less well defined depending on the quality of your signal and the presence of different spot population (i.e. : single molecules and clusters). In the case where the user wants to use automatic threshold but is not satisfied with the result the threshold penalty parameter can be used and will act as a multiplicator of the automatic threshold.

Inflexion point is computed using a maximum distance rule. Let us consider the elbow curve going from A to B, assuming it does follow an elbow shape (L shape) the inflexion is the point farther from the AB segment.

elbow_plot_max_distance

Dense region deconvolution

If your image has dense bright regions the spot dection algorithm might yield an underestimated number of spots because of the size of the maximum filter (see section 1). In such a case you can use this algorithm to compute a more accurate estimation of single molecule numbers in dense regions. This algorithm is based on the detected spots and 3 user defined parameters : alpha, beta, gamma.

Algorithm

The aim of this algorithm is to get a more accurate estimation of single molecule number by fitting Gaussian spots into the intensity signal of the image.

First of all a Gaussian blur is applied to the image according to the gamma parameter removing noise and setting spatial scale accordingly.

To get the fitting parameters a reference spot is computed as the alpha-percentile of the spot distribution by default the median spot is taken as the reference spot (i.e. : alpha = 0.5)

reference_spots

Then bright regions are detected as regions with at least 4 connected pixel above an intensity threshold. This threshold is defined with :

Threshold = beta x median_intensity With median_intesity being the actual median intensity independently of alpha value.

The algorithm will then try to compute more accurate estimation of single number in those bright regions by reconstructing the original signal using only the reference spot. At each iteration the algorithm will add one spot randomly in the bright region and compute the residual sum square (RSS) between the reconstructed and the original signal. As soon as the RSS starts increasing the last added spot will be removed and all the new spots are kept for the software quantification.

Parameters

  • alpha : Percentile for reference spot selection, float between 0 and 1.

  • beta : Penalty for bright region detection, is applied as a multiplicative factor of the median spot intensity.

  • gamma : Multiplicative factor use to compute the gaussian kernel size:

    $kernel\space size = \frac{\gamma.spot\space radius}{voxel\space size}$

    We perform a large gaussian filter with such scale to estimate image background and remove it from original image. A large gamma increases the scale of the gaussian filter and smooth the estimated background. To decompose very large bright areas, a larger gamma should be set. .

  • deconvolution kernel : Standard deviation used for the gaussian kernel (one for each dimension), in pixel. If it’s a scalar, the same standard deviation is applied to every dimensions. If None, we estimate the kernel size from ‘spot_radius’, ‘voxel_size’ and ‘gamma’

Clustering

The cluster detection is based on well described Density Based Spatial Clustering of Applications with Noise. (DBSCAN) (add citations) and uses the sklearn package implementation. The user needs to give 2 parameters for this algorithm : cluster size and min spot number.

Algorithm

The idea for this algorithm is that each point having more than min spot number spots in its surrounding (cluster_size) is labelled as a core point and its surrounding spots as directly reachable points. All directly reachable points are checked to see if they fit the neighbouring spot condition too. In such a case they become core points allowing the cluster to grow. Spots that are not assigned to clusters are called outliers.

In Small Fish all core and directly reachable points are considered as clustered spots and outliers as free spots. If you are using the clustering feature is it recommended to use the dense region deconvolution algorithm as well.

DBSCAN_illut

Clustering parameters

  • cluster_size : Expected size of a cluster in nanometers.
  • min spot number : Minimum number of spot that must be found in neighborhood for a spot to be considered a core point

Script logic

All the algorithms described above form the detection process. It is ran for each cycle and each channel including washout cycles. All the parameters used, the cycle index as well as the color (channel) index are saved in the Detection table. Each detection is referenced with a detection_id.

Concurency

To reduce computing time the software performs several detections simultaneously using concurence. This can be memory intensive, the number of detection executed in paralell can be set using the detection_MAX_WORKERS.

Exclude slices at stack extremities

Some experimental material can yield noisy signal in first or last slices. To adress this issue the DETECTION_SLICE_TO_REMOVE can be used. It can contain two integers, the first is the numbler of slices to exclude from the bottom of the stack and the second from the top of the stack. If you don't want to exclude any leave a blank, for example to exclude 5 slices on the bottom of the stack but none on the top set the parameter to : 5,.

Advice on threshold settings

#TODO

About LoG+Threshold algorithm
Unfortunately, in a large majority of cases the perfect threshold does not exist and you will end up or missing true spots or detecting false spots. To my experience, though, this is hardly an issue as this error is negligible and will not impact your statistics.
The LoG filter + threshold algorithm is very robust to noise and allows quantification as long as the signal is decent. I have tried different deep learning models to achieve this task and up to this date have not find anything more satisfying, one the one hand you do have to set user parameters which is both time consuming and dependend on user but on the other hand the detection process is very clear and we keep the possibility to reach another outcome by modifying spot size parameter without having to retrain a detection model.
Finally, due to training logic a deep learning based detection cannot perform better than the human eye which trained it or if does we cannot assess it on un-labeled data. Using a deterministic algorithm we know the logic that seperate an undetected spot from a detected spot.

Clone this wiki locally