Awk

Notes from The AWK Programming Language (2024) by Alfred V. Aho, Brian W. Kernighan, and Peter J. Weinberger.

I don't use Awk often, so this is my way of holding onto the essentials for later.

What is Awk?

Awk is a scripting language designed for text processing.

An Awk program:

  • is a sequence of pattern-actions statements
    • patterns can involve numeric and strings comparisons
    • actions can include computation and formatted printing
  • every input line is tested against the patterns in order
  • when a pattern matches, it performs the corresponding action

Awk:

  • reads through the input files automatically
  • splits each input line into fields
  • provides a number of built-in variables and functions
  • lets us define our own variables and functions

With this combination of features, many useful computations can be expressed by short programs because the details that would be needed in another language are handled implicitely in Awk.

Awk tutorial

Suppose you have a file called emp.data with three fields:

  • name
  • pay rate in dollars (per hour)
  • number of hours worked

The fields are separated by spaces or tabs:

Beth    21      0
Dan     19      0
Kathy   15.50   10
Mark    25      20
Mary    22.50   22
Susie   17      18

To print names of employees who did not work:

awk '$3 == 0 { print $1 }' emp.data

The part inside the quotes is a complete Awk program.

Structure of an Awk program

An Awk program is a sequence of one or more pattern-action statements:

pattern1 { action1 }
pattern2 { action2 }
…

Awk processes input like this:

  • reads input line by line
  • from any number of files
  • checks each line against all patterns
  • when a pattern matches, it runs the corresponding action

Either the pattern or the action (but not both) may be omitted:

  1. pattern only ($3 == 0): prints every matching line
  2. action only ({ print $1 }): the action is performed for every line

Blank lines are ignored.

Running an Awk program

  1. on files:

    awk '$3 == 0 { print $1 }' file1 file2
    
  2. from terminal input (until an end-of-file signal - ctrl-D):

    awk '$3 == 0 { print $1 }'
    foo foo 0
    foo
    bar bar 1
    baz baz 0
    baz
    ^D
    
  3. By loading its program from a file:

    awk -f progfile optional list of input files
    

Errors

If there's a mistake in an Awk program:

  • "Syntax error" → there's a grammatical mistake (shown between >>> <<<)
  • "Bailing out" → Awk stopped, no recovery was attempted

Simple output

Awk works with only two types of data:

  1. numbers
  2. strings of characters

A file like emp.data is typical: each line contains words and numbers separated by spaces or tabs.

Printing every line

  1. print by itself prints the current input line:

    awk '{ print }' emp.data
    
  2. $0 is the whole line, so this is the same thing:

    awk '{ print $0 }' emp.data
    

Printing specific fields

Fields are referenced as $1, $2, $3, etc.

  1. Print specific fields:

    awk '{ print $1 $3 }' emp.data
    
  2. Awk does not print whitespaces values by default, use commas to add whitespace:

    awk '{ print $1, $3 }' emp.data
    

Default behavior:

  • expressions separated by a comma in a print statement are separated by a space
  • each line produced by print ends with a newline character

NF - Number of fields

awk '{ print NF, $1, NF }' emp.data

Computing and printing

awk '{ print $1, $2 * $3 }' emp.data

NR - Line numbers

awk '{ print NR, $0 }' emp.data

Putting text in the output

awk '{ print "Total pay for", $1, "is", $2 * $3 }' emp.data

Formatted output with printf

The printf statement lets you control exactly how output looks:

printf(format, value_1, value_2, … , value_n)
  • format: a string containing specifications that defines how values are displayed
  • a specification is a % followed by a few characters that control the format of a value
    • the 1st specification tells how value_1 is to be printed
    • the 2nd specification tells how value_2 is to be printed
    • etc.

Examples:

  • %s → string
  • %.2f → number with 2 decimal places
  • %-8s → left-align text in 8-character width
  • %6.2f → number in 6-character width, 2 decimals

Important difference from print

With printf, no spaces or newlines are produced automatically, so don't forget the \n:

awk '{ printf("Total pay for %s is $%.2f\n", $1, $2 * $3) }' emp.data

awk '{ printf("%-8s $%6.2f\n", $1, $2 * $3) }' emp.data

Sorting the output

The easiest way to sort in order of increasing pay is to prefix the total pay and run the output through a sorting program:

awk '{ printf("%6.2f %s\n", $2 * $3, $0) }' emp.data | sort

Selection

A pattern by itself automatically prints all matching lines.

That's why many Awk programs are just a single pattern.

Selection by comparison

awk '$2 >= 20' emp.data

Selection by computation

awk '$2 * $3 > 200 { printf("$%.2f for %s\n", $2 * $3, $1) }' emp.data

Selection by text content

Exact match:

awk '$1 == "Susie"' emp.data

All lines that contain Susie anywhere (regular expression):

awk '/Susie/' emp.data

Combinations of patterns

Patterns can be combined with parentheses and the logical operators &&, ||, and !:

awk '$2 >= 20 || $3 >= 20' emp.data

Data validation

Real-world data often contains errors.

Awk is useful for checking that data is in the right format.

Examples:

awk 'NF != 3 { print $0, "number of fields is not equal to 3" }' emp.data

awk '$2 < 15 { print $0, "rate is too low" }' emp.data

awk '$2 > 25 { print $0, "rate exceeds $25 per hour"}' emp.data

awk '$3 < 0 { print $0, "negative hours worked" }' emp.data

awk '$3 > 60 { print $0, "too many hours worked" }' emp.data

BEGIN and END special patterns

BEGIN runs before any input is processed.

END runs after all input is processed.

awk 'BEGIN { print "NAME    RATE    HOURS"; print "" } { print }' emp.data

Notice that:

  • print "" prints a blank line
  • print prints the current input line

Computing with Awk

An action can contain multiple statements.

Statements can be separated by:

  • new lines

    $3 > 15 { emp = emp + 1 }
    END { print emp, "employes worked more than 15 hours" }
    
  • or semicolons

    $3 > 15 { emp = emp + 1 } ; END { print emp, "employes worked more than 15 hours" }
    

Counting

Counts employees who worked more than 15 hours using an emp variable:

awk '$3 > 15 { emp = emp + 1 } ; END { print emp, "employes worked more than 15 hours" }' emp.data

Shorter version with the increment operator ++:

awk '$3 > 15 { emp++ } ; END { print emp, "employes worked more than 15 hours" }' emp.data

Awk variables for numbers:

  • begin with the value 0
  • aren't explicitely initialized

Computing sums and averages

Count total employees using the built-in NR variable:

awk 'END { print NR, "employees" }' emp.data

To compute the average pay:

awk '{ pay = pay + $2 * $3 } ; END { print NR, "employees" ; print "total pay is", pay ; print "average pay is", pay / NR }' emp.data

Handling text

Awk variables can hold strings of characters.

Find the employee who is paid the most per hour:

awk '
    $2 > maxrate { maxrate = $2; maxemp = $1 }
    END { print "highest hourly rate:", maxrate, "for", maxemp }
' emp.data

String concatenation

Collect all the employee names into a string by appending each name and a space to the previous value in the variable names:

awk '{ names = names $1 " " } ; END { print names }' emp.data

For each line:

  • the first statement concatenates 3 strings:

    1. the previous value of names
    2. the first field
    3. a space
  • and assigns the resulting string to names

The string concatenation operator:

  • is represented by writing string values one after the other
  • there is no explicit concatenation operator
  • in hindsight, this design might not be ideal because it can lead to hard-to-spot errors

Awk variables for strings:

  • begin life with the value null string (a string containing no characters)
  • aren't explicitely initialized

Printing the last input line

Built-in variables like NR and $0 retain their value in an END action.

So one way to print the last line is:

awk 'END { print $0 }' emp.data

Built-in functions

Awk has built-in functions for computing values.

To compute the length of each person's name:

awk '{ print $1, length($1) }' emp.data

Counting lines, words, and characters

Use length, NF, and NR to count the number of lines, words, and characters:

awk '{ nc += length($0) + 1 ; nw += NF } END { print NR, "lines,", nw, "words,", nc, "characters" }' emp.data

We added 1 to nc to count the newline character at the end of each input line.

Control-flow statements

Control-flow statements are used inside actions to control how your program runs.

If-else statement

Use an if to defend against any potential division by zero in computing the average pay:

awk '$2 > 20 { n++; pay += $2 * $3 } END {
  if (n > 0) {
    print n, "high-pay employees Total pay is", pay, "Average pay is", pay/n
  } else {
    print "No employees are paid more than $30/hour"
  }
}' emp.data

While statement

We write an Awk program in a file named interest1.awk:

  • it shows how the value of an amount of money invested at a particular interest rate grows over a number of years
  • using the formula value = amount(1 + rate)^years
# interest1.awk
# Compute compound interest
#   input:  amount  rate  years
#   output: compounded value at the end of each year

{   i = 1
    while (i <= $3) {
        printf("\t%.2f\n", $1 * (1 + $2) ^ i)
        i++
    }
}

In this program:

  • the \t in the printf specification string stands for a tab character
  • the ^ is the exponentiation operator

To use it:

awk -f programs/interest1.awk
1000 .05 5

1000 .15 10

For statement

Here is the previous interest computation with a for:

# interest2.awk
# Compute compound interest
#   input:  amount  rate  years
#   output: compounded value at the end of each year

{ for (i = 1; i <= $3; i++)
    printf("\t%.2f\n", $1 * (1 + $2) ^ i)
}

FizzBuzz

Print the numbers from 1 to 100:

  • if the number is divisible by 3, print "fizz"
  • if the number is divisible by 5, print "buzz"
  • if the number is divisible by both, print "fizzbuzz"
awk 'BEGIN {
  for (i = 1; i <= 100; i++) {
    if (i%15 == 0)  # divisible by both 3 and 5
      print i, "fizzbuzz" 
    else if (i%5 == 0)
      print i, "buzz"
    else if (i%3 == 0)
      print i, "fizz"
    else
      print i
  }
}'

Arrays

Print the input in reverse order:

awk '
    { line[NR] = $0 }  # Remember each input line.
    END {
        for (i = NR; i > 0; i--) {
            print line[i]
        }
    }' emp.data

The first action stores input lines in successive elements of the array line:

  • the first line goes into line[1]
  • the second line goes into line[2]
  • and so on

The expression inside [] is called the subscript.

The subscripts in this example are numeric but they can be arbitrary strings of characters.

Useful one-liners

Description Command
Print the total number of input lines awk 'END { print NR }' emp.data
Print the first 10 input lines awk 'NR <= 10' emp.data
Print the tenth input line awk 'NR == 10' emp.data
Print every tenth input line, starting with line 1 awk 'NR % 10 == 1' emp.data
Print the last field of every input line awk '{ print $NF }' emp.data
Print the last field of the last input line awk 'END { print $NF }' emp.data
Print every input line with more than four fields awk 'NF > 4' emp.data
Print every input line that does not have exactly four fields awk 'NF != 4' emp.data
Print every input line in which the last field is greater than 4 awk '$NF > 4' emp.data
Print the total number of fields in all input lines awk '{ nf += NF } END { print nf }' emp.data
Print the total number of lines that contain `Beth` awk '/Beth/ { nlines++ } END { print nlines }' emp.data
Print the largest field and the line that contains it awk '$1 > max { max = $1; maxline = $0 } END { print max, maxline }' emp.data
Print every line that has at least one field awk 'NF > 0' emp.data
Print every line longer than 80 characters awk 'length($0) > 80' emp.data
Print the number of fields in every line followed by the line itself awk '{ print NF, $0 }' emp.data
Print the first two fields, in opposite order, of every line awk '{ print $2, $1 }' emp.data
Interchange the first two fields of every line and then print the line awk '{ temp = $1; $1 = $2; $2 = temp; print }' emp.data
Print every line preceded by its line number awk '{ print NR, $0 }' emp.data
Print every line with the first field replaced by the line number awk '{ $1 = NR; print }' emp.data
Print every line after erasing the second field awk '{ $2 = ""; print }' emp.data
Print in reverse order the fields of every line
awk '{
        for (i = NF; i > 0; i--) printf("%s ", $i)
        printf("\n")
    }' emp.data
Print the sums of the fields of every line
awk '{
        sum = 0
        for (i = 1; i <= NF; i++) sum = sum + $i
        print sum
    }' emp.data
Add up all fields in all lines and print the sum
awk '
        { for (i = 1; i <= NF; i++) sum = sum + $i }
    END { print sum }
    ' emp.data
Print every line after replacing each field by its absolute value
awk '{
        for (i = 1; i <= NF; i++) if ($i < 0) $i = -$i
        print
    }' emp.data

Passing parameters to Awk (ARGC and ARGV)

Built-ins variables:

  • ARGC → number of arguments
  • ARGV → array of arguments
    • ARGV[1] → first argument
    • ARGV[ARGC-1] → last argument
    • ARGV[0] → name of the program (usually awk)
awk 'BEGIN { argv1 = ARGV[1] ; print argv1 }' foo

Using arguments in a script file

  • a program stored in a file ($* is the shell notation for all the parameters to the script or function):

    # progfile
    
    
    awk 'BEGIN {
        argv0 = ARGV[0]
        argv1 = ARGV[1]
        argv2 = ARGV[2]
        print argv0, argv1, argv2
    }' $*
    
  • make the file executable:

    $ chmod +x progfile
    
  • then call the program:

    $ ./progfile foo bar
    awk foo bar
    

Awk for exploratory data analysis (EDA)

The purpose of exploratory data is to get a sense of what the data is, looking for both patterns and anomalies.

Be approximately right rather than exactly wrong. — John W. Tukey

Awk is great for quickly inspecting data.

Check data format

Validate field structure, each line should have 5 fields and the total field should be correct:

awk 'NF != 5 || $3 != $4 + $5' programs/titanic.tsv

Ensure all records have the same number of fields:

awk --csv '{ print NF }' programs/passengers.csv | sort | uniq -c | sort -nr

CSV dataset

CSV is not rigorously defined, but commonly:

  • fields containing commas or quotes must be enclosed in double quotes
  • any field may be surrounded by quotes (whether it contains commas and quotes or not)
  • an empty field is just ""
  • quotes inside a field are doubled: """,""" represents ","
  • fields may contain newline characters

Awk ≥ 2023 support the --csv argument which causes input lines to be split into fields according to this rule:

awk --csv 'NR > 1 { print $2 }' programs/passengers.csv

Generating CSV

Double each quote and surround the result with quotes:

# to_csv - convert to proper "..."

function to_csv(s) {
    gsub(/"/, "\"\"", s)
    return "\"" s "\""
}

This function can be used within a loop to insert commas between fields:

awk '

    function to_csv(s) {
        gsub(/"/, "\"\"", s)
        return "\"" s "\""
    }

    function record_to_csv(     s, i) {
        for (i = 1; i < NF; i++) {
            s = s to_csv($i) ","
        }
        s = s to_csv($NF)  # No comma after the last field `$NF`.
        return s
    }

    {
        print record_to_csv()
    }

' data/emp.data

Or for an array:

function array_to_csv(arr,       s, i, n) {
    n = length(arr)
    for (i = 1; i <= n; i++) {
        s = s to_csv(arr[i]) ","
    }
    return substr(s, 1, length(s) - 1)  # Remove trailing comma.
}

Generating TSV

awk 'NR > 1 { OFS="\t" ; print $1, $2, $3}' data/emp.data > ~/Desktop/foo.tsv

Associative arrays

Awk supports associative arrays:

  • the subscripts (or indices) of Awk arrays can be arbitrary strings:
    • arrays that allow arbitrary strings as subscripts are called associative arrays
    • other languages provide the same facility with dictionary, map or hashmap
  • Awk has a special form of the for statement for iterating over the indices of an associative array:
    • for (i in array) { statements }
    • the element of the array are visited in an unspecified order
    • you can't count on any particular order

How many people are there in each category?

awk '
    NR > 1 { types[$1] += $3 ; classes[$2] += $3 }
    END {
        for (type in types) {
            print type, types[type]
        }
        print ""
        for (class in classes) {
            print class, classes[class]
        }
    }
' programs/titanic.tsv

Compute the survival rate for each category and pipe the result to sort:

awk 'NR > 1 { printf("%6s %6s %6.1f%%\n", $1, $2, 100 * $4 / $3) }' programs/titanic.tsv | sort -k3 -nr

Beer reviews dataset

Dataset source: kaggle.

File size:

wc data/beer_reviews.csv  # Number of lines, words, and bytes.

Awk equivalent:

awk '
    { nc += length($0) + 1 ; words_num += NF }
    END {
        print NR,  # Lines.
        words_num,
        nc,
        FILENAME
    }
' data/beer_reviews.csv

Fields of the CSV:

  1. brewery_id
  2. brewery_name
  3. review_time
  4. review_overall
  5. review_aroma
  6. review_appearance
  7. review_profilename
  8. beer_style
  9. review_palate
  10. review_taste
  11. beer_name
  12. beer_abv (alcohol content, percentage of alcohol by volume or ABV)
  13. beer_beerid

Find the strongest beer:

awk --csv '
    NR > 1 && $12 > max_abv { max_abv = $12 ; brewery = $2 ; name = $11 }
    END { print max_abv, brewery, name }
' data/beer_reviews.csv

The result is surprisingly high.

That raises a follow-up question: is this value an outlier or the tip of a substantial alcoholic iceberg:

awk --csv 'NR > 1 && $12 >= 10 { print $2, $11, $12 }' data/beer_reviews.csv

What about low-alcohol beer?

awk --csv 'NR > 1 && $12 <= 0.5 && $12 > 0 { print $2, $11, $12 }' data/beer_reviews.csv

What ratings are associated with high and low alcohol?

# Rating of high alcohol.
awk --csv '
    $12 >= 10 { rate += $4 ; nrate++ }
    END { print rate / nrate, nrate }
' data/beer_reviews.csv

# Rating of low alcohol.
awk --csv '
    $12 <= 0.5 && $12 > 0 { rate += $4 ; nrate++ }
    END { print rate / nrate, nrate }
' data/beer_reviews.csv

# Average rating.
awk --csv '
    $12 > 0 { rate += $4 ; nrate++ }
    END { print rate / nrate, nrate }
' data/beer_reviews.csv

We use $12 > 0 because some fields don't have a ABV.

When exploring data, always check:

  • how many fields are empty?
  • how many fields have an explicitely non-useful value like "N/A"?
  • what is the range of values in a column (min/max)?
  • what are the distinct values?

Automating these checks with small scripts can save a lot of time.

Avant Command-line skills Après English tenses

A Kemar Joint