<div style="border: 2px solid #8A9AD0; margin: 1em 0.2em; padding: 0.5em;">

# Advanced CLI in Galaxy

by [The Carpentries](https://training.galaxyproject.org/hall-of-fame/carpentries/), [Helena Rasche](https://training.galaxyproject.org/hall-of-fame/hexylena/), [Bazante Sanders](https://training.galaxyproject.org/hall-of-fame/bazante1/), [Erasmus+ Programme](https://training.galaxyproject.org/hall-of-fame/erasmusplus/), [Avans Hogeschool](https://training.galaxyproject.org/hall-of-fame/avans-atgm/)

CC-BY licensed content from the [Galaxy Training Network](https://training.galaxyproject.org/)

**Objectives**

- How can I combine existing commands to do new things?
- How can I perform the same actions on many different files?
- How can I find files?
- How can I find things in files?

**Objectives**

- Redirect a command's output to a file.
- Process a file instead of keyboard input using redirection.
- Construct command pipelines with two or more stages.
- Explain what usually happens if a program or pipeline isn't given any input to process.
- Explain Unix's 'small pieces, loosely joined' philosophy.
- Write a loop that applies one or more commands separately to each file in a set of files.
- Trace the values taken on by a loop variable during execution of the loop.
- Explain the difference between a variable's name and its value.
- Explain why spaces and some punctuation characters shouldn't be used in file names.
- Demonstrate how to see what commands have recently been executed.
- Re-run recently executed commands without retyping them.
- Use <code>grep</code> to select lines from text files that match simple patterns.
- Use <code>find</code> to find files and directories whose names match simple patterns.
- Use the output of one command as the command-line argument(s) to another command.
- Explain what is meant by 'text' and 'binary' files, and why many common tools don't handle the latter well.

**Time Estimation: 2H**
</div>


<p>This tutorial will walk you through the basics of how to use the Unix command line.</p>
<blockquote class="comment" style="border: 2px solid #ffecc1; margin: 1em 0.2em">
<h3 id="-icon-comment--comment">💬 Comment</h3>
<p>This tutorial is <strong>significantly</strong> based on <a href="https://carpentries.org">the Carpentries</a> <a href="https://swcarpentry.github.io/shell-novice/">“The Unix Shell”</a> lesson, which is licensed CC-BY 4.0. Adaptations have been made to make this work better in a GTN/Galaxy environment.</p>
</blockquote>
<blockquote class="agenda" style="border: 2px solid #86D486;display: none; margin: 1em 0.2em">
<h3 id="agenda">Agenda</h3>
<p>In this tutorial, we will cover:</p>
<ol id="markdown-toc">
<li><a href="#pipes-and-filtering" id="markdown-toc-pipes-and-filtering">Pipes and Filtering</a></li>
</ol>
</blockquote>
<h1 id="pipes-and-filtering">Pipes and Filtering</h1>
<p>Now that we know a few basic commands,
we can finally look at the shell’s most powerful feature:
the ease with which it lets us combine existing programs in new ways.
We’ll start with the directory called <code>shell-lesson-data/molecules</code>
that contains six files describing some simple organic molecules.
The <code>.pdb</code> extension indicates that these files are in Protein Data Bank format,
a simple text format that specifies the type and position of each atom in the molecule.</p>


In [None]:
cd ~/Desktop/shell-lesson-data/
ls molecules

<p>Let’s go into that directory with <code>cd</code> and run an example command <code>wc cubane.pdb</code>:</p>


In [None]:
cd molecules
wc cubane

<p><code class="language-plaintext highlighter-rouge">wc</code> is the ‘word count’ command:
it counts the number of lines, words, and characters in files (from left to right, in that order).</p>
<p>If we run the command <code>wc *.pdb</code>, the <code>*</code> in <code>*.pdb</code> matches zero or more characters,
so the shell turns <code>*.pdb</code> into a list of all <code>.pdb</code> files in the current directory:</p>


In [None]:
wc *.pdb

<p>Note that <code>wc *.pdb</code> also shows the total number of all lines in the last line of the output.</p>
<p>If we run <code>wc -l</code> instead of just <code>wc</code>,
the output shows only the number of lines per file:</p>


In [None]:
wc -l *.pdb

<p>The <code>-m</code> and <code>-w</code> options can also be used with the <code>wc</code> command, to show
only the number of characters or the number of words in the files.</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--why-isnt-it-doing-anything">💡 Why Isn’t It Doing Anything?</h3>
<p>What happens if a command is supposed to process a file, but we
don’t give it a filename? For example, what if we type:</p>
<div class="language-bash highlighter-rouge"><div><pre style="color: inherit; background: white"><code><span class="nv">&#36; </span><span class="nb">wc</span> <span class="nt">-l</span>
</code></pre></div>  </div>
<p>but don’t type <code>*.pdb</code> (or anything else) after the command?
Since it doesn’t have any filenames, <code>wc</code> assumes it is supposed to
process input given at the command prompt, so it just sits there and waits for us to give
it some data interactively. From the outside, though, all we see is it
sitting there: the command doesn’t appear to do anything.</p>
<p>If you make this kind of mistake, you can escape out of this state by holding down
the control key (<kbd>Ctrl</kbd>) and typing the letter <kbd>C</kbd> once and
letting go of the <kbd>Ctrl</kbd> key.
<kbd>Ctrl</kbd>+<kbd>C</kbd></p>
</blockquote>
<h2 id="capturing-output-from-commands">Capturing output from commands</h2>
<p>Which of these files contains the fewest lines?
It’s an easy question to answer when there are only six files,
but what if there were 6000?
Our first step toward a solution is to run the command:</p>


In [None]:
wc -l *.pdb > lengths.txt

<p>The greater than symbol, <code>&gt;</code>, tells the shell to <strong>redirect</strong> the command’s output
to a file instead of printing it to the screen. (This is why there is no screen output:
everything that <code>wc</code> would have printed has gone into the
file <code>lengths.txt</code> instead.)  The shell will create
the file if it doesn’t exist. If the file exists, it will be
silently overwritten, which may lead to data loss and thus requires
some caution.
<code>ls lengths.txt</code> confirms that the file exists:</p>


In [None]:
ls lengths.txt

<p>We can now send the content of <code>lengths.txt</code> to the screen using <code>cat lengths.txt</code>.
The <code>cat</code> command gets its name from ‘concatenate’ i.e. join together,
and it prints the contents of files one after another.
There’s only one file in this case,
so <code>cat</code> just shows us what it contains:</p>


In [None]:
cat lengths.txt

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--output-page-by-page">💡 Output Page by Page</h3>
<p>We’ll continue to use <code>cat</code> in this lesson, for convenience and consistency,
but it has the disadvantage that it always dumps the whole file onto your screen.
More useful in practice is the command <code>less</code>,
which you use with <code>less lengths.txt</code>.
This displays a screenful of the file, and then stops.
You can go forward one screenful by pressing the spacebar,
or back one by pressing <code>b</code>.  Press <code>q</code> to quit.</p>
</blockquote>
<h2 id="filtering-output">Filtering output</h2>
<p>Next we’ll use the <code>sort</code> command to sort the contents of the <code>lengths.txt</code> file.
But first we’ll use an exercise to learn a little about the sort command:</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--what-does-sort--n-do">❓ What Does <code>sort -n</code> Do?</h3>
<p>The file <code>shell-lesson-data/numbers.txt</code>
contains the following lines:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>10
2
19
22
6
</code></pre></div>  </div>
<p>If we run <code>sort</code> on this file, the output is:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>10
19
2
22
6
</code></pre></div>  </div>
<p>If we run <code>sort -n</code> on the same file, we get this instead:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2
6
10
19
22
</code></pre></div>  </div>
<p>Explain why <code>-n</code> has this effect.</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>The <code>-n</code> option specifies a numerical rather than an alphanumerical sort.</p>
</blockquote>
</blockquote>
<p>We will also use the <code>-n</code> option to specify that the sort is
numerical instead of alphanumerical.
This does <em>not</em> change the file;
instead, it sends the sorted result to the screen:</p>


In [None]:
sort -n lengths.txt

<p>We can put the sorted list of lines in another temporary file called <code>sorted-lengths.txt</code>
by putting <code>&gt; sorted-lengths.txt</code> after the command,
just as we used <code>&gt; lengths.txt</code> to put the output of <code>wc</code> into <code>lengths.txt</code>.
Once we’ve done that,
we can run another command called <code>head</code> to get the first few lines in <code>sorted-lengths.txt</code>:</p>


In [None]:
sort -n lengths.txt > sorted-lengths.txt

<p>Using <code>-n 1</code> with <code>head</code> tells it that
we only want the first line of the file;
<code>-n 20</code> would get the first 20,
and so on.
Since <code>sorted-lengths.txt</code> contains the lengths of our files ordered from least to greatest,
the output of <code>head</code> must be the file with the fewest lines.</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--redirecting-to-the-same-file">💡 Redirecting to the same file</h3>
<p>It’s a very bad idea to try redirecting
the output of a command that operates on a file
to the same file. For example:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; sort -n lengths.txt &gt; lengths.txt
</code></pre></div>  </div>
<p>Doing something like this may give you
incorrect results and/or delete
the contents of <code>lengths.txt</code>.</p>
</blockquote>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--what-does--mean">❓ What Does <code>&gt;&gt;</code> Mean?</h3>
<p>We have seen the use of <code>&gt;</code>, but there is a similar operator <code>&gt;&gt;</code>
which works slightly differently.
We’ll learn about the differences between these two operators by printing some strings.
We can use the <code>echo</code> command to print strings e.g.</p>
<blockquote class="code-in" style="border: 2px solid #86D486; margin: 1em 0.2em">
<h3 id="-icon-code-in--input-bash">⌨️ Input: Bash</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; echo The echo command prints text
</code></pre></div>    </div>
</blockquote>
<blockquote class="code-out" style="border: 2px solid #fb99d0; margin: 1em 0.2em">
<h3 id="-icon-code-out--output">🖥 Output</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>The echo command prints text
</code></pre></div>    </div>
</blockquote>
<p>Now test the commands below to reveal the difference between the two operators:</p>
<blockquote class="code-in" style="border: 2px solid #86D486; margin: 1em 0.2em">
<h3 id="-icon-code-in--input-bash-1">⌨️ Input: Bash</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; echo hello &gt; testfile01.txt
</code></pre></div>    </div>
</blockquote>
<p>and:</p>
<blockquote class="code-in" style="border: 2px solid #86D486; margin: 1em 0.2em">
<h3 id="-icon-code-in--input-bash-2">⌨️ Input: Bash</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; echo hello &gt;&gt; testfile02.txt
</code></pre></div>    </div>
</blockquote>
<p><strong>Hint</strong>: Try executing each command twice in a row and then examining the output files.</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<h2 id="solution">Solution</h2>
<p>In the first example with <code>&gt;</code>, the string ‘hello’ is written to <code>testfile01.txt</code>,
but the file gets overwritten each time we run the command.</p>
<p>We see from the second example that the <code>&gt;&gt;</code> operator also writes ‘hello’ to a file
(in this case<code class="language-plaintext highlighter-rouge">testfile02.txt</code>),
but appends the string to the file if it already exists
(i.e. when we run it for the second time).</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--appending-data">❓ Appending Data</h3>
<p>We have already met the <code>head</code> command, which prints lines from the start of a file.
<code>tail</code> is similar, but prints lines from the end of a file instead.</p>
<p>Consider the file <code>shell-lesson-data/data/animals.txt</code>.
After these commands, select the answer that
corresponds to the file <code>animals-subset.txt</code>:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; head -n 3 animals.txt &gt; animals-subset.txt
&#36; tail -n 2 animals.txt &gt;&gt; animals-subset.txt
</code></pre></div>  </div>
<ol>
<li>The first three lines of <code>animals.txt</code></li>
<li>The last two lines of <code>animals.txt</code></li>
<li>The first three lines and the last two lines of <code>animals.txt</code></li>
<li>The second and third lines of <code>animals.txt</code></li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>Option 3 is correct.
For option 1 to be correct we would only run the <code>head</code> command.
For option 2 to be correct we would only run the <code>tail</code> command.
For option 4 to be correct we would have to pipe the output of <code>head</code> into <code>tail -n 2</code>
by doing <code>head -n 3 animals.txt | tail -n 2 &gt; animals-subset.txt</code></p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<h2 id="passing-output-to-another-command">Passing output to another command</h2>
<p>In our example of finding the file with the fewest lines,
we are using two intermediate files <code>lengths.txt</code> and <code>sorted-lengths.txt</code> to store output.
This is a confusing way to work because
even once you understand what <code>wc</code>, <code>sort</code>, and <code>head</code> do,
those intermediate files make it hard to follow what’s going on.
We can make it easier to understand by running <code>sort</code> and <code>head</code> together:</p>


In [None]:
sort -n lengths.txt | head -n 1

<p>The vertical bar, <code>|</code>, between the two commands is called a <strong>pipe</strong>.
It tells the shell that we want to use
the output of the command on the left
as the input to the command on the right.</p>
<p>This has removed the need for the <code>sorted-lengths.txt</code> file.</p>
<h2 id="combining-multiple-commands">Combining multiple commands</h2>
<p>Nothing prevents us from chaining pipes consecutively.
We can for example send the output of <code>wc</code> directly to <code>sort</code>,
and then the resulting output to <code>head</code>.
This removes the need for any intermediate files.</p>
<p>We’ll start by using a pipe to send the output of <code>wc</code> to <code>sort</code>:</p>


In [None]:
wc -l *.pdb | sort -n

<p>We can then send that output through another pipe, to <code>head</code>, so that the full pipeline becomes:</p>


In [None]:
wc -l *.pdb | sort -n | head -n 1

<p>This is exactly like a mathematician nesting functions like <em>log(3x)</em>
and saying ‘the log of three times <em>x</em>’.
In our case,
the calculation is ‘head of sort of line count of <code>*.pdb</code>’.</p>
<p>The redirection and pipes used in the last few commands are illustrated below:</p>
<p><img src="https://training.galaxyproject.org/training-material/topics/data-science/tutorials/cli-advanced/../../images/carpentries-cli/redirects-and-pipes.svg" alt="Redirects and Pipes of different commands" /></p>
<p><code>wc -l *.pdb</code> will direct the output to the shell. <code>wc -l *.pdb &gt; lengths</code> will
direct output to the file lengths. <code>wc -l *.pdb | sort -n | head -n 1</code> will
build a pipeline where the output of the wc command is the input to the sort
command, the output of the sort command is the input to the head command and
the output of the head command is directed to the shell</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--piping-commands-together">❓ Piping Commands Together</h3>
<p>In our current directory, we want to find the 3 files which have the least number of
lines. Which command listed below would work?</p>
<ol>
<li><code>wc -l * &gt; sort -n &gt; head -n 3</code></li>
<li><code>wc -l * | sort -n | head -n 1-3</code></li>
<li><code>wc -l * | head -n 3 | sort -n</code></li>
<li><code>wc -l * | sort -n | head -n 3</code></li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>Option 4 is the solution.
The pipe character <code>|</code> is used to connect the output from one command to
the input of another.
<code>&gt;</code> is used to redirect standard output to a file.
Try it in the <code>shell-lesson-data/molecules</code> directory!</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<h2 id="tools-designed-to-work-together">Tools designed to work together</h2>
<p>This idea of linking programs together is why Unix has been so successful.
Instead of creating enormous programs that try to do many different things,
Unix programmers focus on creating lots of simple tools that each do one job well,
and that work well with each other.
This programming model is called ‘pipes and filters’.
We’ve already seen pipes;
a <strong>filter</strong> is a program like <code>wc</code> or <code>sort</code>
that transforms a stream of input into a stream of output.
Almost all of the standard Unix tools can work this way:
unless told to do otherwise,
they read from standard input,
do something with what they’ve read,
and write to standard output.</p>
<p>The key is that any program that reads lines of text from standard input
and writes lines of text to standard output
can be combined with every other program that behaves this way as well.
You can <em>and should</em> write your programs this way
so that you and other people can put those programs into pipes to multiply their power.</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--pipe-reading-comprehension">❓ Pipe Reading Comprehension</h3>
<p>A file called <code>animals.txt</code> (in the <code>shell-lesson-data/data</code> folder) contains the following data:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2012-11-05,deer
2012-11-05,rabbit
2012-11-05,raccoon
2012-11-06,rabbit
2012-11-06,deer
2012-11-06,fox
2012-11-07,rabbit
2012-11-07,bear
</code></pre></div>  </div>
<p>What text passes through each of the pipes and the final redirect in the pipeline below?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; cat animals.txt | head -n 5 | tail -n 3 | sort -r &gt; final.txt
</code></pre></div>  </div>
<p>Hint: build the pipeline up one command at a time to test your understanding</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>The <code>head</code> command extracts the first 5 lines from <code>animals.txt</code>.
Then, the last 3 lines are extracted from the previous 5 by using the <code>tail</code> command.
With the <code>sort -r</code> command those 3 lines are sorted in reverse order and finally,
the output is redirected to a file <code>final.txt</code>.
The content of this file can be checked by executing <code>cat final.txt</code>.
The file should contain the following lines:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2012-11-06,rabbit
2012-11-06,deer
2012-11-05,raccoon
</code></pre></div>    </div>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--pipe-construction">❓ Pipe Construction</h3>
<p>For the file <code>animals.txt</code> from the previous exercise, consider the following command:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; cut -d , -f 2 animals.txt
</code></pre></div>  </div>
<p>The <code>cut</code> command is used to remove or ‘cut out’ certain sections of each line in the file,
and <code>cut</code> expects the lines to be separated into columns by a <kbd>Tab</kbd> character.
A character used in this way is a called a <strong>delimiter</strong>.
In the example above we use the <code>-d</code> option to specify the comma as our delimiter character.
We have also used the <code>-f</code> option to specify that we want to extract the second field (column).
This gives the following output:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>deer
rabbit
raccoon
rabbit
deer
fox
rabbit
bear
</code></pre></div>  </div>
<p>The <code>uniq</code> command filters out adjacent matching lines in a file.
How could you extend this pipeline (using <code>uniq</code> and another command) to find
out what animals the file contains (without any duplicates in their
names)?</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; cut -d , -f 2 animals.txt | sort | uniq
</code></pre></div>    </div>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--which-pipe">❓ Which Pipe?</h3>
<p>The file <code>animals.txt</code> contains 8 lines of data formatted as follows:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2012-11-05,deer
2012-11-05,rabbit
2012-11-05,raccoon
2012-11-06,rabbit
...
</code></pre></div>  </div>
<p>The <code>uniq</code> command has a <code>-c</code> option which gives a count of the
number of times a line occurs in its input.  Assuming your current
directory is <code>shell-lesson-data/data/</code>, what command would you use to produce
a table that shows the total count of each type of animal in the file?</p>
<ol>
<li><code>sort animals.txt | uniq -c</code></li>
<li><code>sort -t, -k2,2 animals.txt | uniq -c</code></li>
<li><code>cut -d, -f 2 animals.txt | uniq -c</code></li>
<li><code>cut -d, -f 2 animals.txt | sort | uniq -c</code></li>
<li><code>cut -d, -f 2 animals.txt | sort | uniq -c | wc -l</code></li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>Option 4. is the correct answer.
If you have difficulty understanding why, try running the commands, or sub-sections of
the pipelines (make sure you are in the <code>shell-lesson-data/data</code> directory).</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<h2 id="nelles-pipeline-checking-files">Nelle’s Pipeline: Checking Files</h2>
<p>Nelle has run her samples through the assay machines
and created 17 files in the <code>north-pacific-gyre/2012-07-03</code> directory described earlier.
As a quick check, starting from her home directory, Nelle types:</p>


In [None]:
cd ~/Desktop/shell-lesson-data/north-pacific-gyre/2012-07-03
wc -l *.txt

<p>The output is 18 lines that look like this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>300 NENE01729A.txt
300 NENE01729B.txt
300 NENE01736A.txt
300 NENE01751A.txt
300 NENE01751B.txt
300 NENE01812A.txt
... ...
</code></pre></div></div>
<p>Now she types this:</p>


In [None]:
wc -l *.txt | sort -n | head -n 5

<p>Whoops: one of the files is 60 lines shorter than the others.
When she goes back and checks it,
she sees that she did that assay at 8:00 on a Monday morning — someone
was probably in using the machine on the weekend,
and she forgot to reset it.
Before re-running that sample,
she checks to see if any files have too much data:</p>


In [None]:
wc -l *.txt | sort -n | tail -n 5

<p>Those numbers look good — but what’s that ‘Z’ doing there in the third-to-last line?
All of her samples should be marked ‘A’ or ‘B’;
by convention,
her lab uses ‘Z’ to indicate samples with missing information.
To find others like it, she does this:</p>


In [None]:
ls *Z.txt

<p>Sure enough,
when she checks the log on her laptop,
there’s no depth recorded for either of those samples.
Since it’s too late to get the information any other way,
she must exclude those two files from her analysis.
She could delete them using <code>rm</code>,
but there are actually some analyses she might do later where depth doesn’t matter,
so instead, she’ll have to be careful later on to select files using the wildcard expressions
<code class="language-plaintext highlighter-rouge">NENE*A.txt NENE*B.txt</code>.</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--removing-unneeded-files">❓ Removing Unneeded Files</h3>
<p>Suppose you want to delete your processed data files, and only keep
your raw files and processing script to save storage.
The raw files end in <code>.dat</code> and the processed files end in <code>.txt</code>.
Which of the following would remove all the processed data files,
and <em>only</em> the processed data files?</p>
<ol>
<li><code>rm ?.txt</code></li>
<li><code>rm *.txt</code></li>
<li><code>rm * .txt</code></li>
<li><code>rm *.*</code></li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<ol>
<li>This would remove <code>.txt</code> files with one-character names</li>
<li>This is correct answer</li>
<li>The shell would expand <code>*</code> to match everything in the current directory,
so the command would try to remove all matched files and an additional
file called <code>.txt</code></li>
<li>The shell would expand <code>*.*</code> to match all files with any extension,
so this command would delete all files</li>
</ol>
</blockquote>
</blockquote>
<h1 id="loops">Loops</h1>
<p><strong>Loops</strong> are a programming construct which allow us to repeat a command or set of commands
for each item in a list.
As such they are key to productivity improvements through automation.
Similar to wildcards and tab completion, using loops also reduces the
amount of typing required (and hence reduces the number of typing mistakes).</p>
<p>Suppose we have several hundred genome data files named <code>basilisk.dat</code>, <code>minotaur.dat</code>, and
<code class="language-plaintext highlighter-rouge">unicorn.dat</code>.
For this example, we’ll use the <code>creatures</code> directory which only has three example files,
but the principles can be applied to many many more files at once. First, go
into the creatures directory.</p>


In [None]:
# Change directories here!


<p>The structure of these files is the same: the common name, classification, and updated date are
presented on the first three lines, with DNA sequences on the following lines.
Let’s look at the files:</p>


In [None]:
head -n 5 basilisk.dat minotaur.dat unicorn.dat

<p>We would like to print out the classification for each species, which is given on the second
line of each file.
For each file, we would need to execute the command <code>head -n 2</code> and pipe this to <code>tail -n 1</code>.
We’ll use a loop to solve this problem, but first let’s look at the general form of a loop:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for thing in list_of_things
do
operation_using &#36;thing    # Indentation within the loop is not required, but aids legibility
done
</code></pre></div></div>
<p>and we can apply this to our example like this:</p>


In [None]:
for filename in basilisk.dat minotaur.dat unicorn.dat
do
head -n 2 $filename | tail -n 1
done

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--follow-the-prompt">💡 Follow the Prompt</h3>
<p>The shell prompt changes from <code>&#36;</code> to <code>&gt;</code> and back again as we were
typing in our loop. The second prompt, <code>&gt;</code>, is different to remind
us that we haven’t finished typing a complete command yet. A semicolon, <code>;</code>,
can be used to separate two commands written on a single line.</p>
</blockquote>
<p>When the shell sees the keyword <code>for</code>,
it knows to repeat a command (or group of commands) once for each item in a list.
Each time the loop runs (called an iteration), an item in the list is assigned in sequence to
the <strong>variable</strong>, and the commands inside the loop are executed, before moving on to
the next item in the list.
Inside the loop,
we call for the variable’s value by putting <code>&#36;</code> in front of it.
The <code>&#36;</code> tells the shell interpreter to treat
the variable as a variable name and substitute its value in its place,
rather than treat it as text or an external command.</p>
<p>In this example, the list is three filenames: <code>basilisk.dat</code>, <code>minotaur.dat</code>, and <code>unicorn.dat</code>.
Each time the loop iterates, it will assign a file name to the variable <code>filename</code>
and run the <code>head</code> command.
The first time through the loop,
<code>&#36;filename</code> is <code>basilisk.dat</code>.
The interpreter runs the command <code>head</code> on <code>basilisk.dat</code>
and pipes the first two lines to the <code>tail</code> command,
which then prints the second line of <code>basilisk.dat</code>.
For the second iteration, <code>&#36;filename</code> becomes
<code class="language-plaintext highlighter-rouge">minotaur.dat</code>. This time, the shell runs <code>head</code> on <code>minotaur.dat</code>
and pipes the first two lines to the <code>tail</code> command,
which then prints the second line of <code>minotaur.dat</code>.
For the third iteration, <code>&#36;filename</code> becomes
<code class="language-plaintext highlighter-rouge">unicorn.dat</code>, so the shell runs the <code>head</code> command on that file,
and <code>tail</code> on the output of that.
Since the list was only three items, the shell exits the <code>for</code> loop.</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--same-symbols-different-meanings">💡 Same Symbols, Different Meanings</h3>
<p>Here we see <code>&gt;</code> being used as a shell prompt, whereas <code>&gt;</code> is also
used to redirect output.
Similarly, <code>&#36;</code> is used as a shell prompt, but, as we saw earlier,
it is also used to ask the shell to get the value of a variable.</p>
<p>If the <em>shell</em> prints <code>&gt;</code> or <code>&#36;</code> then it expects you to type something,
and the symbol is a prompt.</p>
<p>If <em>you</em> type <code>&gt;</code> or <code>&#36;</code> yourself, it is an instruction from you that
the shell should redirect output or get the value of a variable.</p>
</blockquote>
<p>When using variables it is also
possible to put the names into curly braces to clearly delimit the variable
name: <code>&#36;filename</code> is equivalent to <code>&#36;{filename}</code>, but is different from
<code class="language-plaintext highlighter-rouge">&#36;{file}name</code>. You may find this notation in other people’s programs.</p>
<p>We have called the variable in this loop <code>filename</code>
in order to make its purpose clearer to human readers.
The shell itself doesn’t care what the variable is called;
if we wrote this loop as:</p>


In [None]:
for x in basilisk.dat minotaur.dat unicorn.dat
do
head -n 2 $x | tail -n 1
done

<p>or:</p>


In [None]:
for temperature in basilisk.dat minotaur.dat unicorn.dat
do
head -n 2 $temperature | tail -n 1
done

<p>it would work exactly the same way.</p>
<p><strong>Don’t do this.</strong></p>
<p>Programs are only useful if people can understand them,
so meaningless names (like <code>x</code>) or misleading names (like <code>temperature</code>)
increase the odds that the program won’t do what its readers think it does.</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--variables-in-loops">❓ Variables in Loops</h3>
<p>This exercise refers to the <code>shell-lesson-data/molecules</code> directory.
<code>ls</code> gives the following output:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
</code></pre></div>  </div>
<p>What is the output of the following code?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in *.pdb
do
    ls *.pdb
done
</code></pre></div>  </div>
<p>Now, what is the output of the following code?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in *.pdb
do
   ls &#36;datafile
done
</code></pre></div>  </div>
<p>Why do these two loops give different outputs?</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>The first code block gives the same output on each iteration through
the loop.
Bash expands the wildcard <code>*.pdb</code> within the loop body (as well as
before the loop starts) to match all files ending in <code>.pdb</code>
and then lists them using <code>ls</code>.
The expanded loop would look like this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; for datafile in cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
&gt; do
&gt;     ls cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
&gt; done
</code></pre></div>    </div>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
cubane.pdb  ethane.pdb  methane.pdb  octane.pdb  pentane.pdb  propane.pdb
</code></pre></div>    </div>
<p>The second code block lists a different file on each loop iteration.
The value of the <code>datafile</code> variable is evaluated using <code>&#36;datafile</code>,
and then listed using <code>ls</code>.</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cubane.pdb
ethane.pdb
methane.pdb
octane.pdb
pentane.pdb
propane.pdb
</code></pre></div>    </div>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--limiting-sets-of-files">❓ Limiting Sets of Files</h3>
<p>What would be the output of running the following loop in thei
<code>shell-lesson-data/molecules</code> directory?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for filename in c*
do
    ls &#36;filename
done
</code></pre></div>  </div>
<ol>
<li>No files are listed.</li>
<li>All files are listed.</li>
<li>Only <code>cubane.pdb</code>, <code>octane.pdb</code> and <code>pentane.pdb</code> are listed.</li>
<li>Only <code>cubane.pdb</code> is listed.</li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>4 is the correct answer. <code>*</code> matches zero or more characters, so any file name starting with
the letter c, followed by zero or more other characters will be matched.</p>
</blockquote>
<p>How would the output differ from using this command instead?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for filename in *c*
do
    ls &#36;filename
done
</code></pre></div>  </div>
<ol>
<li>The same files would be listed.</li>
<li>All the files are listed this time.</li>
<li>No files are listed this time.</li>
<li>The files <code>cubane.pdb</code> and <code>octane.pdb</code> will be listed.</li>
<li>Only the file <code>octane.pdb</code> will be listed.</li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<h3 id="-icon-solution--solution-1">👁 Solution</h3>
<p>4 is the correct answer. <code>*</code> matches zero or more characters, so a file name with zero or more
characters before a letter c and zero or more characters after the letter c will be matched.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--saving-to-a-file-in-a-loop---part-one">❓ Saving to a File in a Loop - Part One</h3>
<p>In the <code>shell-lesson-data/molecules</code> directory, what is the effect of this loop?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for alkanes in *.pdb
do
    echo &#36;alkanes
    cat &#36;alkanes &gt; alkanes.pdb
done
</code></pre></div>  </div>
<ol>
<li>Prints <code>cubane.pdb</code>, <code>ethane.pdb</code>, <code>methane.pdb</code>, <code>octane.pdb</code>, <code>pentane.pdb</code> and
<code>propane.pdb</code>, and the text from <code>propane.pdb</code> will be saved to a file called <code>alkanes.pdb</code>.</li>
<li>Prints <code>cubane.pdb</code>, <code>ethane.pdb</code>, and <code>methane.pdb</code>, and the text from all three files
would be concatenated and saved to a file called <code>alkanes.pdb</code>.</li>
<li>Prints <code>cubane.pdb</code>, <code>ethane.pdb</code>, <code>methane.pdb</code>, <code>octane.pdb</code>, and <code>pentane.pdb</code>,
and the text from <code>propane.pdb</code> will be saved to a file called <code>alkanes.pdb</code>.</li>
<li>None of the above.</li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>1 is correct. The text from each file in turn gets written to the <code>alkanes.pdb</code> file.
However, the file gets overwritten on each loop iteration, so the final content of <code>alkanes.pdb</code>
is the text from the <code>propane.pdb</code> file.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--saving-to-a-file-in-a-loop---part-two">❓ Saving to a File in a Loop - Part Two</h3>
<p>Also in the <code>shell-lesson-data/molecules</code> directory,
what would be the output of the following loop?</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in *.pdb
do
    cat &#36;datafile &gt;&gt; all.pdb
done
</code></pre></div>  </div>
<ol>
<li>All of the text from <code>cubane.pdb</code>, <code>ethane.pdb</code>, <code>methane.pdb</code>, <code>octane.pdb</code>, and
<code>pentane.pdb</code> would be concatenated and saved to a file called <code>all.pdb</code>.</li>
<li>The text from <code>ethane.pdb</code> will be saved to a file called <code>all.pdb</code>.</li>
<li>All of the text from <code>cubane.pdb</code>, <code>ethane.pdb</code>, <code>methane.pdb</code>, <code>octane.pdb</code>, <code>pentane.pdb</code>
and <code>propane.pdb</code> would be concatenated and saved to a file called <code>all.pdb</code>.</li>
<li>All of the text from <code>cubane.pdb</code>, <code>ethane.pdb</code>, <code>methane.pdb</code>, <code>octane.pdb</code>, <code>pentane.pdb</code>
and <code>propane.pdb</code> would be printed to the screen and saved to a file called <code>all.pdb</code>.</li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>3 is the correct answer. <code>&gt;&gt;</code> appends to a file, rather than overwriting it with the redirected
output from a command.
Given the output from the <code>cat</code> command has been redirected, nothing is printed to the screen.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<p>Let’s continue with our example in the <code>shell-lesson-data/creatures</code> directory.
Here’s a slightly more complicated loop:</p>


In [None]:
for filename in *.dat
do
echo $filename
head -n 100 $filename | tail -n 20
done

<p>The shell starts by expanding <code>*.dat</code> to create the list of files it will process.
The <strong>loop body</strong>
then executes two commands for each of those files.
The first command, <code>echo</code>, prints its command-line arguments to standard output.
For example:</p>


In [None]:
echo hello there

<p>prints:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>hello there
</code></pre></div></div>
<p>In this case,
since the shell expands <code>&#36;filename</code> to be the name of a file,
<code>echo &#36;filename</code> prints the name of the file.
Note that we can’t write this as:</p>


In [None]:
for filename in *.dat
do
$filename
head -n 100 $filename | tail -n 20
done

<p>because then the first time through the loop,
when <code>&#36;filename</code> expanded to <code>basilisk.dat</code>, the shell would try to run <code>basilisk.dat</code> as a program.
Finally,
the <code>head</code> and <code>tail</code> combination selects lines 81-100
from whatever file is being processed
(assuming the file has at least 100 lines).</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--spaces-in-names">💡 Spaces in Names</h3>
<p>Spaces are used to separate the elements of the list
that we are going to loop over. If one of those elements
contains a space character, we need to surround it with
quotes, and do the same thing to our loop variable.
Suppose our data files are named:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>red dragon.dat
purple unicorn.dat
</code></pre></div>  </div>
<p>To loop over these files, we would need to add double quotes like so:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; for filename in "red dragon.dat" "purple unicorn.dat"
&gt; do
&gt;     head -n 100 "&#36;filename" | tail -n 20
&gt; done
</code></pre></div>  </div>
<p>It is simpler to avoid using spaces (or other special characters) in filenames.</p>
<p>The files above don’t exist, so if we run the above code, the <code>head</code> command will be unable
to find them, however the error message returned will show the name of the files it is
expecting:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>head: cannot open ‘red dragon.dat’ for reading: No such file or directory
head: cannot open ‘purple unicorn.dat’ for reading: No such file or directory
</code></pre></div>  </div>
<p>Try removing the quotes around <code>&#36;filename</code> in the loop above to see the effect of the quote
marks on spaces. Note that we get a result from the loop command for unicorn.dat
when we run this code in the <code>creatures</code> directory:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>head: cannot open ‘red’ for reading: No such file or directory
head: cannot open ‘dragon.dat’ for reading: No such file or directory
head: cannot open ‘purple’ for reading: No such file or directory
CGGTACCGAA
AAGGGTCGCG
CAAGTGTTCC
...
</code></pre></div>  </div>
</blockquote>
<p>We would like to modify each of the files in <code>shell-lesson-data/creatures</code>, but also save a version
of the original files, naming the copies <code>original-basilisk.dat</code> and <code>original-unicorn.dat</code>.
We can’t use:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cp *.dat original-*.dat
</code></pre></div></div>
<p>because that would expand to:</p>


In [None]:
cp basilisk.dat minotaur.dat unicorn.dat original-*.dat

<p>This wouldn’t back up our files, instead we get an error.</p>
<p>This problem arises when <code>cp</code> receives more than two inputs. When this happens, it
expects the last input to be a directory where it can copy all the files it was passed.
Since there is no directory named <code>original-*.dat</code> in the <code>creatures</code> directory we get an
error.</p>
<p>Instead, we can use a loop:</p>


In [None]:
for filename in *.dat
do
cp $filename original-$filename
done

<p>This loop runs the <code>cp</code> command once for each filename.
The first time,
when <code>&#36;filename</code> expands to <code>basilisk.dat</code>,
the shell executes:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cp basilisk.dat original-basilisk.dat
</code></pre></div></div>
<p>The second time, the command is:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cp minotaur.dat original-minotaur.dat
</code></pre></div></div>
<p>The third and last time, the command is:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cp unicorn.dat original-unicorn.dat
</code></pre></div></div>
<p>Since the <code>cp</code> command does not normally produce any output, it’s hard to check
that the loop is doing the correct thing.
However, we learned earlier how to print strings using <code>echo</code>, and we can modify the loop
to use <code>echo</code> to print our commands without actually executing them.
As such we can check what commands <em>would be</em> run in the unmodified loop.</p>
<p>The following diagram
shows what happens when the modified loop is executed, and demonstrates how the
judicious use of <code>echo</code> is a good debugging technique.</p>
<p><img src="https://training.galaxyproject.org/training-material/topics/data-science/tutorials/cli-advanced/../../images/carpentries-cli/shell_script_for_loop_flow_chart.svg" alt="The for loop 'for filename in *.dat; do echo cp &#36;filename original-&#36;filename; done' will successively assign the names of all '*.dat' files in your current directory to the variable '&#36;filename' and then execute the command. With the files 'basilisk.dat', 'minotaur.dat' and 'unicorn.dat' in the current directory the loop will successively call the echo command three times and print three lines: 'cp basislisk.dat original-basilisk.dat', then 'cp minotaur.dat original-minotaur.dat' and finally 'cp unicorn.dat original-unicorn.dat'" /></p>
<h2 id="nelles-pipeline-processing-files">Nelle’s Pipeline: Processing Files</h2>
<p>Nelle is now ready to process her data files using <code>goostats.sh</code> —
a shell script written by her supervisor.
This calculates some statistics from a protein sample file, and takes two arguments:</p>
<ol>
<li>an input file (containing the raw data)</li>
<li>an output file (to store the calculated statistics)</li>
</ol>
<p>Since she’s still learning how to use the shell,
she decides to build up the required commands in stages.
Her first step is to make sure that she can select the right input files — remember,
these are ones whose names end in ‘A’ or ‘B’, rather than ‘Z’.
Starting from her home directory, Nelle types:</p>


In [None]:
cd ~/Desktop/shell-lesson-data/north-pacific-gyre/2012-07-03
for datafile in NENE*A.txt NENE*B.txt
do
echo $datafile
done

<p>Her next step is to decide
what to call the files that the <code>goostats.sh</code> analysis program will create.
Prefixing each input file’s name with ‘stats’ seems simple,
so she modifies her loop to do that:</p>


In [None]:
for datafile in NENE*A.txt NENE*B.txt
do
echo $datafile stats-$datafile
done

<p>She hasn’t actually run <code>goostats.sh</code> yet,
but now she’s sure she can select the right files and generate the right output filenames.</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--top-terminal-tip-re-running-previous-commands">💡 Top Terminal Tip: Re-running previous commands</h3>
<p>Typing in commands over and over again is becoming tedious,
though,
and Nelle is worried about making mistakes,
so instead of re-entering her loop,
she presses <kbd>↑</kbd>.
In response,
the shell redisplays the whole loop on one line
(using semi-colons to separate the pieces):</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in NENE*A.txt NENE*B.txt; do echo &#36;datafile stats-&#36;datafile; done
</code></pre></div>  </div>
</blockquote>
<p>Using the left arrow key,
Nelle backs up and changes the command <code>echo</code> to <code>bash goostats.sh</code>:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in NENE*A.txt NENE*B.txt; do bash goostats.sh &#36;datafile stats-&#36;datafile; done
</code></pre></div></div>
<p>When she presses <kbd>Enter</kbd>,
the shell runs the modified command.
However, nothing appears to happen — there is no output.
After a moment, Nelle realizes that since her script doesn’t print anything to the screen
any longer, she has no idea whether it is running, much less how quickly.
She kills the running command by typing <kbd>Ctrl</kbd>+<kbd>C</kbd>,
uses <kbd>↑</kbd> to repeat the command,
and edits it to read:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in NENE*A.txt NENE*B.txt; do echo &#36;datafile; bash goostats.sh &#36;datafile stats-&#36;datafile; done
</code></pre></div></div>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h2 id="beginning-and-end">Beginning and End</h2>
<p>We can move to the beginning of a line in the shell by typing <kbd>Ctrl</kbd>+<kbd>A</kbd>
and to the end using <kbd>Ctrl</kbd>+<kbd>E</kbd>.</p>
</blockquote>
<p>When she runs her program now,
it produces one line of output every five seconds or so:</p>
<p>1518 times 5 seconds,
divided by 60,
tells her that her script will take about two hours to run.
As a final check,
she opens another terminal window,
goes into <code>north-pacific-gyre/2012-07-03</code>,
and uses <code>cat stats-NENE01729B.txt</code>
to examine one of the output files.
It looks good,
so she decides to get some coffee and catch up on her reading.</p>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--those-who-know-history-can-choose-to-repeat-it">💡 Those Who Know History Can Choose to Repeat It</h3>
<p>Another way to repeat previous work is to use the <code>history</code> command to
get a list of the last few hundred commands that have been executed, and
then to use <code>!123</code> (where ‘123’ is replaced by the command number) to
repeat one of those commands. For example, if Nelle types this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; history | tail -n 5
</code></pre></div>  </div>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>  456  ls -l NENE0*.txt
  457  rm stats-NENE01729B.txt.txt
  458  bash goostats.sh NENE01729B.txt stats-NENE01729B.txt
  459  ls -l NENE0*.txt
  460  history
</code></pre></div>  </div>
<p>then she can re-run <code>goostats.sh</code> on <code>NENE01729B.txt</code> simply by typing
<code>!458</code>. This number will be different for you, you should check your history before running it!</p>
</blockquote>
<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--other-history-commands">💡 Other History Commands</h3>
<p>There are a number of other shortcut commands for getting at the history.</p>
<ul>
<li><kbd>Ctrl</kbd>+<kbd>R</kbd> enters a history search mode ‘reverse-i-search’ and finds the
most recent command in your history that matches the text you enter next.
Press <kbd>Ctrl</kbd>+<kbd>R</kbd> one or more additional times to search for earlier matches.
You can then use the left and right arrow keys to choose that line and edit
it then hit <kbd>Return</kbd> to run the command.</li>
<li><code>!!</code> retrieves the immediately preceding command
(you may or may not find this more convenient than using <kbd>↑</kbd>)</li>
<li><code>!&#36;</code> retrieves the last word of the last command.
That’s useful more often than you might expect: after
<code>bash goostats.sh NENE01729B.txt stats-NENE01729B.txt</code>, you can type
<code>less !&#36;</code> to look at the file <code>stats-NENE01729B.txt</code>, which is
quicker than doing <kbd>↑</kbd> and editing the command-line.</li>
</ul>
</blockquote>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--doing-a-dry-run">❓ Doing a Dry Run</h3>
<p>A loop is a way to do many things at once — or to make many mistakes at
once if it does the wrong thing. One way to check what a loop <em>would</em> do
is to <code>echo</code> the commands it would run instead of actually running them.</p>
<p>Suppose we want to preview the commands the following loop will execute
without actually running those commands:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cd ~/Desktop/shell-lesson-data/pdb/
for datafile in *.pdb
do
    cat &#36;datafile &gt;&gt; all.pdb
done
</code></pre></div>  </div>
<p>What is the difference between the two loops below, and which one would we
want to run?</p>
<blockquote class="code-in" style="border: 2px solid #86D486; margin: 1em 0.2em">
<h3 id="-icon-code-in--version-1">⌨️ Version 1</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in *.pdb
do
    echo cat &#36;datafile &gt;&gt; all.pdb
done
</code></pre></div>    </div>
</blockquote>
<blockquote class="code-in" style="border: 2px solid #86D486; margin: 1em 0.2em">
<h3 id="-icon-code-in--version-2">⌨️ Version 2</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for datafile in *.pdb
do
    echo "cat &#36;datafile &gt;&gt; all.pdb"
done
</code></pre></div>    </div>
</blockquote>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<h3 id="-icon-tip--solution">💡 Solution</h3>
<p>The second version is the one we want to run.
This prints to screen everything enclosed in the quote marks, expanding the
loop variable name because we have prefixed it with a dollar sign.</p>
<p>The first version appends the output from the command <code>echo cat &#36;datafile</code>
to the file, <code>all.pdb</code>. This file will just contain the list;
<code>cat cubane.pdb</code>, <code>cat ethane.pdb</code>, <code>cat methane.pdb</code> etc.</p>
<p>Try both versions for yourself to see the output! Be sure to open the
<code>all.pdb</code> file to view its contents.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--nested-loops">❓ Nested Loops</h3>
<p>Suppose we want to set up a directory structure to organize
some experiments measuring reaction rate constants with different compounds
<em>and</em> different temperatures.  What would be the
result of the following code:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for species in cubane ethane methane
do
    for temperature in 25 30 37 40
    do
        mkdir &#36;species-&#36;temperature
    done
done
</code></pre></div>  </div>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>We have a nested loop, i.e. contained within another loop, so for each species
in the outer loop, the inner loop (the nested loop) iterates over the list of
temperatures, and creates a new directory for each combination.</p>
<p>Try running the code for yourself to see which directories are created!</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<h1 id="finding-things">Finding Things</h1>
<p>In the same way that many of us now use ‘Google’ as a
verb meaning ‘to find’, Unix programmers often use the
word ‘grep’.
‘grep’ is a contraction of ‘global/regular expression/print’,
a common sequence of operations in early Unix text editors.
It is also the name of a very useful command-line program.</p>
<p><code>grep</code> finds and prints lines in files that match a pattern.
For our examples,
we will use a file that contains three haiku taken from a
1998 competition in <em>Salon</em> magazine. For this set of examples,
we’re going to be working in the writing subdirectory:</p>


In [None]:
cd
cd Desktop/shell-lesson-data/writing
cat haiku.txt

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--forever-or-five-years">💡 Forever, or Five Years</h3>
<p>We haven’t linked to the original haiku because
they don’t appear to be on <em>Salon</em>’s site any longer.
As <a href="https://www.clir.org/wp-content/uploads/sites/6/ensuring.pdf">Jeff Rothenberg said</a>,
‘Digital information lasts forever — or five years, whichever comes first.’
Luckily, popular content often <a href="http://wiki.c2.com/?ComputerErrorHaiku">has backups</a>.</p>
</blockquote>
<p>Let’s find lines that contain the word ‘not’:</p>


In [None]:
grep not haiku.txt

<p>Here, <code>not</code> is the pattern we’re searching for.
The grep command searches through the file, looking for matches to the pattern specified.
To use it type <code>grep</code>, then the pattern we’re searching for and finally
the name of the file (or files) we’re searching in.</p>
<p>The output is the three lines in the file that contain the letters ‘not’.</p>
<p>By default, grep searches for a pattern in a case-sensitive way.
In addition, the search pattern we have selected does not have to form a complete word,
as we will see in the next example.</p>
<p>Let’s search for the pattern: ‘The’.</p>


In [None]:
grep The haiku.txt

<p>This time, two lines that include the letters ‘The’ are outputted,
one of which contained our search pattern within a larger word, ‘Thesis’.</p>
<p>To restrict matches to lines containing the word ‘The’ on its own,
we can give <code>grep</code> with the <code>-w</code> option.
This will limit matches to word boundaries.</p>
<p>Later in this lesson, we will also see how we can change the search behavior of grep
with respect to its case sensitivity.</p>


In [None]:
grep -w The haiku.txt

<p>Note that a ‘word boundary’ includes the start and end of a line, so not
just letters surrounded by spaces.
Sometimes we don’t
want to search for a single word, but a phrase. This is also easy to do with
<code>grep</code> by putting the phrase in quotes.</p>


In [None]:
grep -w "is not" haiku.txt

<p>We’ve now seen that you don’t have to have quotes around single words,
but it is useful to use quotes when searching for multiple words.
It also helps to make it easier to distinguish between the search term or phrase
and the file being searched.
We will use quotes in the remaining examples.</p>
<p>Another useful option is <code>-n</code>, which numbers the lines that match:</p>


In [None]:
grep -n "it" haiku.txt

<p>Here, we can see that lines 5, 9, and 10 contain the letters ‘it’.</p>
<p>We can combine options (i.e. flags) as we do with other Unix commands.
For example, let’s find the lines that contain the word ‘the’.
We can combine the option <code>-w</code> to find the lines that contain the word ‘the’
and <code>-n</code> to number the lines that match:</p>


In [None]:
grep -n -w "the" haiku.txt

<p>Now we want to use the option <code>-i</code> to make our search case-insensitive:</p>


In [None]:
grep -n -w -i "the" haiku.txt

<p>Now, we want to use the option <code>-v</code> to invert our search, i.e., we want to output
the lines that do not contain the word ‘the’.</p>


In [None]:
grep -n -w -v "the" haiku.txt

<p>If we use the <code>-r</code> (recursive) option,
<code>grep</code> can search for a pattern recursively through a set of files in subdirectories.</p>
<p>Let’s search recursively for <code>Yesterday</code> in the <code>shell-lesson-data/writing</code> directory:</p>


In [None]:
grep -r Yesterday .

<p><code class="language-plaintext highlighter-rouge">grep</code> has lots of other options. To find out what they are, we can type:</p>


In [None]:
grep --help

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--using-grep">❓ Using <code>grep</code></h3>
<p>Which command would result in the following output:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>and the presence of absence:
</code></pre></div>  </div>
<ol>
<li><code>grep "of" haiku.txt</code></li>
<li><code>grep -E "of" haiku.txt</code></li>
<li><code>grep -w "of" haiku.txt</code></li>
<li><code>grep -i "of" haiku.txt</code></li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>The correct answer is 3, because the <code>-w</code> option looks only for whole-word matches.
The other options will also match ‘of’ when part of another word.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--wildcards">💡 Wildcards</h3>
<p><code>grep</code>’s real power doesn’t come from its options, though; it comes from
the fact that patterns can include wildcards. (The technical name for
these is <strong>regular expressions</strong>, which
is what the ‘re’ in ‘grep’ stands for.) Regular expressions are both complex
and powerful; if you want to do complex searches, please look at the lesson
on <a href="http://v4.software-carpentry.org/regexp/index.html">our website</a>. As a taster, we can
find lines that have an ‘o’ in the second position like this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; grep -E "^.o" haiku.txt
</code></pre></div>  </div>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>You bring fresh toner.
Today it is not working
Software is like that.
</code></pre></div>  </div>
<p>We use the <code>-E</code> option and put the pattern in quotes to prevent the shell
from trying to interpret it. (If the pattern contained a <code>*</code>, for
example, the shell would try to expand it before running <code>grep</code>.) The
<code>^</code> in the pattern anchors the match to the start of the line. The <code>.</code>
matches a single character (just like <code>?</code> in the shell), while the <code>o</code>
matches an actual ‘o’.</p>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--tracking-a-species">❓ Tracking a Species</h3>
<p>Leah has several hundred
data files saved in one directory, each of which is formatted like this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2013-11-05,deer,5
2013-11-05,rabbit,22
2013-11-05,raccoon,7
2013-11-06,rabbit,19
2013-11-06,deer,2
</code></pre></div>  </div>
<p>She wants to write a shell script that takes a species as the first command-line argument
and a directory as the second argument. The script should return one file called <code>species.txt</code>
containing a list of dates and the number of that species seen on each date.
For example using the data shown above, <code>rabbit.txt</code> would contain:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>2013-11-05,22
2013-11-06,19
</code></pre></div>  </div>
<p>Put these commands and pipes in the right order to achieve this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>cut -d : -f 2
&gt;
|
grep -w &#36;1 -r &#36;2
|
&#36;1.txt
cut -d , -f 1,3
</code></pre></div>  </div>
<p>Hint: use <code>man grep</code> to look for how to grep text recursively in a directory
and <code>man cut</code> to select more than one field in a line.</p>
<p>An example of such a file is provided in <code>shell-lesson-data/data/animal-counts/animals.txt</code></p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>grep -w &#36;1 -r &#36;2 | cut -d : -f 2 | cut -d , -f 1,3 &gt; &#36;1.txt
</code></pre></div>    </div>
<p>Actually, you can swap the order of the two cut commands and it still works. At the
command line, try changing the order of the cut commands, and have a look at the output
from each step to see why this is the case.</p>
<p>You would call the script above like this:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>&#36; bash count-species.sh bear .
</code></pre></div>    </div>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--little-women">❓ Little Women</h3>
<p>You and your friend, having just finished reading <em>Little Women</em> by
Louisa May Alcott, are in an argument.  Of the four sisters in the
book, Jo, Meg, Beth, and Amy, your friend thinks that Jo was the
most mentioned.  You, however, are certain it was Amy.  Luckily, you
have a file <code>LittleWomen.txt</code> containing the full text of the novel
(<code class="language-plaintext highlighter-rouge">shell-lesson-data/writing/data/LittleWomen.txt</code>).
Using a <code>for</code> loop, how would you tabulate the number of times each
of the four sisters is mentioned?</p>
<p>Hint: one solution might employ
the commands <code>grep</code> and <code>wc</code> and a <code>|</code>, while another might utilize
<code>grep</code> options.
There is often more than one way to solve a programming task, so a
particular solution is usually chosen based on a combination of
yielding the correct result, elegance, readability, and speed.</p>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<h3 id="-icon-solution--solutions">👁 Solutions</h3>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for sis in Jo Meg Beth Amy
do
	echo &#36;sis:
	grep -ow &#36;sis LittleWomen.txt | wc -l
done
</code></pre></div>    </div>
<p>Alternative, slightly inferior solution:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>for sis in Jo Meg Beth Amy
do
	echo &#36;sis:
	grep -ocw &#36;sis LittleWomen.txt
done
</code></pre></div>    </div>
<p>This solution is inferior because <code>grep -c</code> only reports the number of lines matched.
The total number of matches reported by this method will be lower if there is more
than one match per line.</p>
<p>Perceptive observers may have noticed that character names sometimes appear in all-uppercase
in chapter titles (e.g. ‘MEG GOES TO VANITY FAIR’).
If you wanted to count these as well, you could add the <code>-i</code> option for case-insensitivity
(though in this case, it doesn’t affect the answer to which sister is mentioned
most frequently).</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<p>While <code>grep</code> finds lines in files,
the <code>find</code> command finds files themselves.
Again,
it has a lot of options;
to show how the simplest ones work, we’ll use the directory tree shown below.</p>
<p><img src="https://training.galaxyproject.org/training-material/topics/data-science/tutorials/cli-advanced/../../images/carpentries-cli/find-file-tree.svg" alt="A file tree under the directory 'writing' contians several sub-directories and files such that 'writing' contains directories 'data', 'thesis', 'tools' and a file 'haiku.txt'; 'writing/data' contains the files 'Little Women.txt', 'one.txt' and 'two.txt'; 'writing/thesis' contains the file 'empty-draft.md'; 'writing/tools' contains the directory 'old' and the files 'format' and 'stats'; and 'writing/tools/old' contains a file 'oldtool'" /></p>
<p>Nelle’s <code>writing</code> directory contains one file called <code>haiku.txt</code> and three subdirectories:
<code>thesis</code> (which contains a sadly empty file, <code>empty-draft.md</code>);
<code>data</code> (which contains three files <code>LittleWomen.txt</code>, <code>one.txt</code> and <code>two.txt</code>);
and a <code>tools</code> directory that contains the programs <code>format</code> and <code>stats</code>,
and a subdirectory called <code>old</code>, with a file <code>oldtool</code>.</p>
<p>For our first command,
let’s run <code>find .</code> (remember to run this command from the <code>shell-lesson-data/writing</code> folder).</p>


In [None]:
find .

<p>As always,
the <code>.</code> on its own means the current working directory,
which is where we want our search to start.
<code class="language-plaintext highlighter-rouge">find</code>’s output is the names of every file <strong>and</strong> directory
under the current working directory.
This can seem useless at first but <code>find</code> has many options
to filter the output and in this lesson we will discover some
of them.</p>
<p>The first option in our list is
<code>-type d</code> that means ‘things that are directories’.
Sure enough,
<code class="language-plaintext highlighter-rouge">find</code>’s output is the names of the five directories in our little tree
(including <code>.</code>):</p>


In [None]:
find . -type d

<p>Notice that the objects <code>find</code> finds are not listed in any particular order.
If we change <code>-type d</code> to <code>-type f</code>,
we get a listing of all the files instead:</p>


In [None]:
find . -type f

<p>Now let’s try matching by name:</p>


In [None]:
find . -name *.txt

<p>We expected it to find all the text files,
but it only prints out <code>./haiku.txt</code>.
The problem is that the shell expands wildcard characters like <code>*</code> <em>before</em> commands run.
Since <code>*.txt</code> in the current directory expands to <code>haiku.txt</code>,
the command we actually ran was:</p>


In [None]:
find . -name haiku.txt

<p><code class="language-plaintext highlighter-rouge">find</code> did what we asked; we just asked for the wrong thing.</p>
<p>To get what we want,
let’s do what we did with <code>grep</code>:
put <code>*.txt</code> in quotes to prevent the shell from expanding the <code>*</code> wildcard.
This way,
<code>find</code> actually gets the pattern <code>*.txt</code>, not the expanded filename <code>haiku.txt</code>:</p>


In [None]:
find . -name "*.txt"

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--listing-vs-finding">💡 Listing vs. Finding</h3>
<p><code>ls</code> and <code>find</code> can be made to do similar things given the right options,
but under normal circumstances,
<code>ls</code> lists everything it can,
while <code>find</code> searches for things with certain properties and shows them.</p>
</blockquote>
<p>As we said earlier,
the command line’s power lies in combining tools.
We’ve seen how to do that with pipes;
let’s look at another technique.
As we just saw,
<code>find . -name "*.txt"</code> gives us a list of all text files in or below the current directory.
How can we combine that with <code>wc -l</code> to count the lines in all those files?</p>
<p>The simplest way is to put the <code>find</code> command inside <code>&#36;()</code>:</p>


In [None]:
wc -l $(find . -name "*.txt")

<p>When the shell executes this command,
the first thing it does is run whatever is inside the <code>&#36;()</code>.
It then replaces the <code>&#36;()</code> expression with that command’s output.
Since the output of <code>find</code> is the four filenames <code>./data/one.txt</code>, <code>./data/LittleWomen.txt</code>,
<code class="language-plaintext highlighter-rouge">./data/two.txt</code>, and <code>./haiku.txt</code>, the shell constructs the command:</p>


In [None]:
wc -l ./data/one.txt ./data/LittleWomen.txt ./data/two.txt ./haiku.txt

<p>which is what we wanted.
This expansion is exactly what the shell does when it expands wildcards like <code>*</code> and <code>?</code>,
but lets us use any command we want as our own ‘wildcard’.</p>
<p>It’s very common to use <code>find</code> and <code>grep</code> together.
The first finds files that match a pattern;
the second looks for lines inside those files that match another pattern.
Here, for example, we can find PDB files that contain iron atoms
by looking for the string ‘FE’ in all the <code>.pdb</code> files above the current directory:</p>


In [None]:
grep "FE" $(find .. -name "*.pdb")

<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--matching-and-subtracting">❓ Matching and Subtracting</h3>
<p>The <code>-v</code> option to <code>grep</code> inverts pattern matching, so that only lines
which do <em>not</em> match the pattern are printed. Given that, which of
the following commands will find all files in <code>/data</code> whose names
end in <code>s.txt</code> but whose names also do <em>not</em> contain the string <code>net</code>?
(For example, <code>animals.txt</code> or <code>amino-acids.txt</code> but not <code>planets.txt</code>.)
Once you have thought about your answer, you can test the commands in the <code>shell-lesson-data</code>
directory.</p>
<ol>
<li><code>find data -name "*s.txt" | grep -v net</code></li>
<li><code>find data -name *s.txt | grep -v net</code></li>
<li><code>grep -v "net" &#36;(find data -name "*s.txt")</code></li>
<li>None of the above.</li>
</ol>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<p>The correct answer is 1. Putting the match expression in quotes prevents the shell
expanding it, so it gets passed to the <code>find</code> command.</p>
<p>Option 2 is incorrect because the shell expands <code>*s.txt</code> instead of passing the wildcard
expression to <code>find</code>.</p>
<p>Option 3 is incorrect because it searches the contents of the files for lines which
do not match ‘net’, rather than searching the file names.</p>
</blockquote>
</blockquote>


In [None]:
# Explore the possible solutions here!

<blockquote class="tip" style="border: 2px solid #FFE19E; margin: 1em 0.2em">
<h3 id="-icon-tip--binary-files">💡 Binary Files</h3>
<p>We have focused exclusively on finding patterns in text files. What if
your data is stored as images, in databases, or in some other format?</p>
<p>A handful of tools extend <code>grep</code> to handle a few non text formats. But a
more generalizable approach is to convert the data to text, or
extract the text-like elements from the data. On the one hand, it makes simple
things easy to do. On the other hand, complex things are usually impossible. For
example, it’s easy enough to write a program that will extract X and Y
dimensions from image files for <code>grep</code> to play with, but how would you
write something to find values in a spreadsheet whose cells contained
formulas?</p>
<p>A last option is to recognize that the shell and text processing have
their limits, and to use another programming language.
When the time comes to do this, don’t be too hard on the shell: many
modern programming languages have borrowed a lot of
ideas from it, and imitation is also the sincerest form of praise.</p>
</blockquote>
<p>The Unix shell is older than most of the people who use it. It has
survived so long because it is one of the most productive programming
environments ever created — maybe even <em>the</em> most productive. Its syntax
may be cryptic, but people who have mastered it can experiment with
different commands interactively, then use what they have learned to
automate their work. Graphical user interfaces may be easier to use at
first, but once learned, the productivity in the shell is unbeatable.
And as Alfred North Whitehead wrote in 1911, ‘Civilization advances by
extending the number of important operations which we can perform
without thinking about them.’</p>
<blockquote class="question" style="border: 2px solid #8A9AD0; margin: 1em 0.2em">
<h3 id="-icon-question--find-pipeline-reading-comprehension">❓ <code>find</code> Pipeline Reading Comprehension</h3>
<p>Write a short explanatory comment for the following shell script:</p>
<div class="language-plaintext highlighter-rouge"><div><pre style="color: inherit; background: white"><code>wc -l &#36;(find . -name "*.dat") | sort -n
</code></pre></div>  </div>
<blockquote class="solution" style="border: 2px solid #B8C3EA;color: white; margin: 1em 0.2em">
<div style="color: #555; font-size: 95%;">Hint: Select the text with your mouse to see the answer</div><h3 id="-icon-solution--solution">👁 Solution</h3>
<ol>
<li>Find all files with a <code>.dat</code> extension recursively from the current directory</li>
<li>Count the number of lines each of these files contains</li>
<li>Sort the output from step 2. numerically</li>
</ol>
</blockquote>
</blockquote>
<h1 id="final-notes">Final Notes</h1>
<p>All of the commands you have run up until now were ad-hoc, interactive commands.</p>


# Key Points

- `wc` counts lines, words, and characters in its inputs.
- `cat` displays the contents of its inputs.
- `sort` sorts its inputs.
- `head` displays the first 10 lines of its input.
- `tail` displays the last 10 lines of its input.
- `command > [file]` redirects a command's output to a file (overwriting any existing content).
- `command >> [file]` appends a command's output to a file.
- `[first] | [second]` is a pipeline: the output of the first command is used as the input to the second.
- The best way to use the shell is to use pipes to combine simple single-purpose programs (filters).
- A `for` loop repeats commands once for every thing in a list.
- Every `for` loop needs a variable to refer to the thing it is currently operating on.
- Use `$name` to expand a variable (i.e., get its value). `${name}` can also be used.
- Do not use spaces, quotes, or wildcard characters such as '*' or '?' in filenames, as it complicates variable expansion.
- Give files consistent names that are easy to match with wildcard patterns to make it easy to select them for looping.
- Use the up-arrow key to scroll up through previous commands to edit and repeat them.
- Use <kbd>Ctrl</kbd>+<kbd>R</kbd> to search through the previously entered commands.
- Use `history` to display recent commands, and `![number]` to repeat a command by number.
- `find` finds files with specific properties that match patterns.
- `grep` selects lines in files that match patterns.
- `--help` is an option supported by many bash commands, and programs that can be run from within Bash, to display more information on how to use these commands or programs.
- `man [command]` displays the manual page for a given command.
- `$([command])` inserts a command's output in place.

# Congratulations on successfully completing this tutorial!

Please [fill out the feedback on the GTN website](https://training.galaxyproject.org/training-material/topics/data-science/tutorials/cli-advanced/tutorial.html#feedback) and check there for further resources!
