Skip to content

Distributing tasks to fully use computing power

etiaum edited this page Aug 5, 2026 · 10 revisions

By Etienne Aumont, with the guidance of Gleb Bezgin
With this guide, You'll learn how to make tasks run exponentially faster.

image

Table of contents

  1. Connecting to other computers (beginner-friendly)
  2. Using all computer cores (Parallel)
  3. Distributing tasks to Burnside computers (SLURM)
    3.1 Preparing your jobs
    3.2 Running and following-up on jobs

1. Connecting to other computers

The ssh command switches to a different machine on Linux systems. It is useful for our TNL lab operations because most of the better machines are located outside the lab. If option '-Y' is specified, it would allow graphics forwarding, so that programs with graphical user interface (like RStudio, MATLAB, Display, etc) will look the same way as on the local machine.

Example:

  ssh -Y anton  

After that, start your RStudio, MATLAB, etc from that terminal, the program will then run on the specified machine - in this case, the machine called anton.

List of fast Burnside machines that can be used for heavier computations:

adie, agraval, anton, aran, bastian, beck, beevor, bender

2. Using all computer cores (Parallel)

Learning for loops is highly recommended before learning parallelization (see the BASH tutorial). Command lines are generally run on a single computer core. However, the lab's computers all have 40+ cores. Parallel solves that by allocating tasks to different cores. It is ideal for moderate to heavy computations that are repeated at least a couple of times. It makes tasks run MUCH faster on the powerful Burnside computers. For example, I wanted to convert 1104 MRIs from .mnc to .nii, something that takes ~1 second per image without parallel (18 mins total). With parallel on Anton, it only took 30 seconds.

 for i in *.mnc; do echo mnc2nii $i ; done > imageconversion.txt  
   #First, transform the for loop into a temporary text file with 1 line per image to transform  
 cat imageconversion.txt | parallel  
   #Second, run the lines in parallel  

More complex tasks such as ASHS may also use parallel with an in-built option. (see the -P option in the ASHS guide). However, they only make individual jobs run more quickly. If you need to run several tasks, the paralleled loop will be more efficient.

For longer tasks, I recommend using nohup to let it run without needing the terminal to be open and to store the logs in a file

  nohup bash -c 'your command' > output.log &  

3. Distributing tasks to Burnside computers (SLURM)

3.2 Preparing your jobs

For very heavy computing, you can use an even more powerful method that distributes tasks across the Burnside cluster. For example, after converting my 1104 MRIs, I wanted to run the ASHS segmentation. To do this, I first build the folder structure with shell/, logs/ and errors/ folders.

 mkdir shell  
 mkdir logs  
 mkdir errors  

Then, I have to populate the shell folder with shell scripts

 for i in *.nii ;  
 do (echo '#!/bin/bash' && echo hostname &&      #mandatory part, do not modify  
 echo $ASHS_ROOT/bin/ashs_main.sh -P -I ${i:0:16} -a ../ashsT1_atlas_upennpmc_wm_07172023 -g $i -f $i -w ${i:0:16};  
 echo rm ${i:0:16}/bootstrap -r ;                
 echo rm ${i:0:16}/multiatlas -r) > shell/${i:0:16}.sh
 ; done   

Here, the loop does 5 things:

  1. run through each image file (participant is identified by the 16 first characters)
  2. identify the compiler (bash) and specify who we are
  3. run ASHS
  4. clear superfluous files to avoid encumbering the disk
  5. put all of this in a shell script

3.2 Running and following-up on jobs

Then, we run all of the scripts in a loop using SLURM:

 for i in shell/* ; 
 do sbatch --array=1-N%150 --cpus-per-task=4 --mem=12G --time=8:00:00 --exclude="" -o logs/${i:6:16}_log.out -e errors/${i:6:16}_error.out $i
 ; done

For lighter tasks, you can also run them with less CPU per task, but since ASHS runs with parallel built-in, it's better with more than 1 CPU.

You can follow-up on what jobs are run by what computer with:

 squeue -u etienne             #to show all of my jobs
 squeue -u etienne | wc -l     #to show how many lines remain to do
 scancel -u etienne            #to interrupt your jobs

Clone this wiki locally