Command-line skills

This post is part of a three-part series: #1, #2, #3.

The rest of my notes from Efficient Linux at the Command Line (2022) by Daniel J. Barrett, focused on a handful of simple tricks that make a real difference.

Combining commands with pipes

less

tree .git/objects/ | less

wc

Count lines in each file:

wc -l *

How many files/directories?

ls -1 | wc -l

Why ls -1? Because ls changes its behavior when redirected:

  • unlike other commands ls is aware of whether stdout is the screen or whether it's been redirected (to a pipe or otherwise):
    • when stdout is the screen, ls arranges its output in multiple columns (the reason is user friendliness)
    • when ls is redirected, it produces a single column
  • ls -1 forces a single column output
    • it is not strictly necessary when redirected but it's more explicit

head prints the first lines of a file.

Efficient for large files.

Can be used to reduce the output from another command:

ls -1 /bin | head -n5

cut

cut prints one or more column from a file.

cut by field (-f) when the input consist of strings (fields) separated by tab characters:

cut -f2 animals.txt
cut -f2,3 animals.txt  # Multiple fields.
cut -f2-4 animals.txt  # By numeric range.

Or cut by character position (-c):

cut -c1-3 animals.txt

Use a custom delimiter (-d) to cut a CSV…

cut -d";" -f2 data.csv

…and re-pipe to cut to get FirstName from LastName FirstName:

cut -d";" -f5 data.csv | cut -d" " -f2 | head -n10

sort

sort sorts alphabetically (default) or numerically (-n) in ascending (default) or descending (-r) order:

sort animals.txt
sort -r animals.txt

cut -f3 animals.txt | sort -n
cut -f3 animals.txt | sort -rn

Get min and max:

cut -f3 animals.txt | sort -n | head -n1   # min.
cut -f3 animals.txt | sort -rn | head -n1  # max.

uniq

uniq detect repeated "adjacent" lines in a file. It does not scan the entire file for duplicates unless the duplicates are next to each other.

By default, it removes the repeats if any:

uniq animals.txt
uniq -c animals.txt

# grades.txt contains alphabetical grades (A, B, C…)
cut -f1 grades.txt | sort | uniq -c | sort -nr

Testing tools

Combining commands with pipes often requires trial and error.

Tools and techniques to help:

  • ls or echo to test destructive commands (e.g., echo instead of rm)
  • tee to view intermediate results:
    • command1 | command2 | command3 | tee outfile | command4 | command5
    • tee saves the output from command3 in outfile
    • while piping the same output to command4

Producing text

date

date
date +%Y-%m-%d
date +%H:%M:%S
date +%Y-%m-%d\ %H:%M:%S
date +"I cannot believe it's already %A"

seq (sequence of numbers)

seq 1 5
seq 1 2 10        # Middle number = increment.
seq 3 -1 0
seq 1.1 0.1 2
seq -s/ 1 5       # --separator
seq -w 8 10       # --fixed-width

BSD alternative: jot.

Brace expansion

Or curly brace expansion.

{x..y..z} generates the values x through y incrementing by z:

echo {1..10}
echo {10..1}
echo {01..10}

echo {A..Z}
echo {A..Z} | tr -d " "  # Delete spaces.

find

List files in a directory recursively.

The default action is -print (it can be omitted it in the following commands):

find /etc -print    # List all of /etc recursively.

find . -type f -print
find . -type d -print

find . -type f -name "*.py" -print

find . -iname "*txt" -print

Execute a command for each file path with -exec:

  1. construct a find command
  2. append -exec
    • followed by the command to execute
    • {} indicates where the file path should appear in the command
  3. end with a quoted or escaped semicolon (";" or \;)

E.g.

# Print an @ symbol on either side of the file path.
find . -exec echo @ {} @ ";"

-exec is often used to ls or rm files:

# Delete files with names ending in a tilde within the directory and its subdirectories.
# Use `echo` for safety.
find /tmp -type f -name "*~" -exec echo rm {} ";"
# Delete for real.
find /tmp -type f -name "*~" -exec rm {} ";"

But some built-in actions are more efficient than -exec ls or -exec rm:

find . -type f -name "*.txt" -ls

find /tmp -type f -name "*~" -delete

yes

Prints the same string until terminated:

yes woof!

The main use for our purpose is printing a string a specific number of times:

yes woof! | head -n3

Isolating text

grep

# Whole words only.
grep --color -w enough cite.txt

# Only nonblank lines with `-v` (--invert-match).
grep --invert-match "^$" cite.txt

# Literal matches only. Ignore regular expressions with `-F` ("fixed").
grep -F --color "e." cite.txt

tail

tail -n3 cite.txt    # Last 3 lines.

tail -n+90 cite.txt  # From line 90 to the end.

Combine tail and head to print the fourth line only:

head -n4 cite.txt | tail -n1

awk '{print}'

awk extracts columns in a way cut cannot.

Key concepts:

  • $1, $2, etc. → columns
  • $NF → last column ("number of fields")
  • $0 → entire line
  • field separator → any run of spaces/tabs/newlines (vs a single character in cut)

awk does not print whitespaces values by default, use commas to add whitespace:

df | awk '{print $1 $3}'    # No whitespace.

df | awk '{print $1, $3}'   # Whitespace.

Combining text

cat

cat (concatenate) prints the output of multiple files to stdout:

cat poem1 poem2 poem3

tac

tac:

  • cat spelled backward
  • reverses a file line by line
  • great for processing data that is already in chronological order but not reversible with sort -r

E.g. reverse a web-server log file to process its lines from newest to oldest:

tac access.log

paste

paste:

  • combines files side by side in columns separated by a single tab character
  • good partner to cut (which extracts columns from a tab separated file)
paste words1.txt words2.txt | cut -f2

paste -d, words1.txt words2.txt         # Change separator.

paste -s words1.txt words2.txt          # Produce rows instead of columns.

paste -d"\n" words1.txt words2.txt      # Interleave data with a new line as separator.

diff

diff can be used as text processor that interleaves lines from two files.

Many users don't think of diff in this way but it can be used to solve certain problems.

Isolate differing lines:

diff file1.txt file2.txt

diff file1.txt file2.txt | grep '^[<>]'

diff file1.txt file2.txt | grep '^[<>]' | cut -c3-

Transforming text

tr

tr translates one set of characters into another:

echo efficient | tr a-z A-Z

echo Efficient Linux | tr " " "\n"      # Convert spaces into new lines.

echo Efficient Linux | tr -d " \t"      # Delete spaces and tabs.

rev

rev reverses the character of each line of input:

echo Efficient Linux | rev

Beyond the obvious entertainment value, rev can be used to extract the final word of each line:

rev file1.txt | cut -d' ' -f1 | rev

sed essentials

sed uses a sed script (a sequence of instructions) to transform text from stdin into any other text.

The most common type of script is a substitution script:

echo Efficient Linux | sed s/Linux/macOS/

# The forward slashes may be replaced by any other character.
echo Efficient Linux | sed s@Linux@macOS@
echo Efficient Linux | sed s+Linux+macOS+

Another type of script is a deletion script:

seq 10 20 | sed 4d               # Remove the 4th line.

seq 101 200 | sed '/[13579]$/d'  # Delete lines ending with an odd digit.

Delete carriage returns and newlines:

sed -e s/$'\r'/,/g a.txt > b.txt

Practical CLI tips

Launching browser from the CLI

open <url>  # Default browser.

open -a Safari
open -a Safari "https://marcarea.com"

open -a "Google Chrome"
open -a "Google Chrome" "https://marcarea.com"
open -na "Google Chrome" --args -incognito "https://marcarea.com"

Retrieving rendered web content with a text-based browser

Use a text-based brower such as lynx or links to download a rendered page with the -dump option:

lynx -dump https://marcarea.com > tmpfile
  • useful for checking out a suspicous-looking link
  • they can't promise complete security, so use your best judgement

Clipboard control

On macOS, pbcopy and pbpaste connect selections to stdin and stdout:

pbcopy < 00-intro.txt

pbpaste | wc -w

pbcopy < ~/.ssh/id_rsa.pub
pbpaste > main.go
pbpaste >> main.go

# Encode clipboard in base64.
pbpaste | base64 | pbcopy

# Copy external IP address to clipboard.
curl -Ss https://icanhazip.com | pbcopy

# Copy private IP address to clipboard.
ipconfig getifaddr en0 | pbcopy

Editing files that contains a given string

Open files containing a string in an editor:

vi $(grep -l string *)

mate $(grep -l string *)

Recursively with the -r option and beginning in the current directory (the dot):

mate $(grep -lr string .)

For faster search of large directory trees:

mate $(find . -type f -print0 | xargs -0 grep -l string)

Processing a file one line at a time

cat the file into a while read loop:

cat myfile | while read line; do
    echo "$line" | wc -c
done

Generating test files

  1. select a random number of words from a file using $RANDOM (a positive integer between 0 and 32,767):

    gshuf -n $RANDOM /usr/share/dict/words
    
  2. create a directory for files:

    mkdir -p ~/Desktop/randomfiles && cd ~/Desktop/randomfiles
    
  3. generate a file with a random filename containing a random number of random words:

    # Generate a random string.
    openssl rand -hex 5
    
    
    gshuf -n $RANDOM -o $(openssl rand -hex 5).txt /usr/share/dict/words
    
  4. run the solution multiple times:

    • with a for loop:

      for i in {1..10}; do
          gshuf -n $RANDOM -o $(openssl rand -hex 5).txt /usr/share/dict/words
      done
      
    • or a one-liner (e.g., run the command 10 times using yes):

      yes 'gshuf -n $RANDOM -o $(openssl rand -hex 5).txt /usr/share/dict/words' | head -n 10 | bash
      

Longer learning

  • cron
  • crontab
  • at (schedules commands to run once)
  • rsync
  • make

Avant Ways to run a shell command Après Awk

A Kemar Joint