Notes from The AWK Programming Language (2024) by Alfred V. Aho, Brian W. Kernighan, and Peter J. Weinberger.
I don't use Awk often, so this is my way of holding onto the essentials for later.
What is Awk?
Awk is a scripting language designed for text processing.
An Awk program:
- is a sequence of pattern-actions statements
- patterns can involve numeric and strings comparisons
- actions can include computation and formatted printing
- every input line is tested against the patterns in order
- when a pattern matches, it performs the corresponding action
Awk:
- reads through the input files automatically
- splits each input line into fields
- provides a number of built-in variables and functions
- lets us define our own variables and functions
With this combination of features, many useful computations can be expressed by short programs because the details that would be needed in another language are handled implicitely in Awk.
Awk tutorial
Suppose you have a file called emp.data with three fields:
- name
- pay rate in dollars (per hour)
- number of hours worked
The fields are separated by spaces or tabs:
Beth 21 0
Dan 19 0
Kathy 15.50 10
Mark 25 20
Mary 22.50 22
Susie 17 18
To print names of employees who did not work:
awk '$3 == 0 { print $1 }' emp.data
The part inside the quotes is a complete Awk program.
Structure of an Awk program
An Awk program is a sequence of one or more pattern-action statements:
pattern1 { action1 }
pattern2 { action2 }
…
Awk processes input like this:
- reads input line by line
- from any number of files
- checks each line against all patterns
- when a pattern matches, it runs the corresponding action
Either the pattern or the action (but not both) may be omitted:
- pattern only (
$3 == 0): prints every matching line - action only (
{ print $1 }): the action is performed for every line
Blank lines are ignored.
Running an Awk program
on files:
awk '$3 == 0 { print $1 }' file1 file2from terminal input (until an end-of-file signal -
ctrl-D):awk '$3 == 0 { print $1 }' foo foo 0 foo bar bar 1 baz baz 0 baz ^DBy loading its program from a file:
awk -f progfile optional list of input files
Errors
If there's a mistake in an Awk program:
- "Syntax error" → there's a grammatical mistake (shown between
>>> <<<) - "Bailing out" → Awk stopped, no recovery was attempted
Simple output
Awk works with only two types of data:
- numbers
- strings of characters
A file like emp.data is typical: each line contains words and numbers separated by spaces or tabs.
Printing every line
printby itself prints the current input line:awk '{ print }' emp.data$0is the whole line, so this is the same thing:awk '{ print $0 }' emp.data
Printing specific fields
Fields are referenced as $1, $2, $3, etc.
Print specific fields:
awk '{ print $1 $3 }' emp.dataAwk does not print whitespaces values by default, use commas to add whitespace:
awk '{ print $1, $3 }' emp.data
Default behavior:
- expressions separated by a comma in a
printstatement are separated by a space - each line produced by
printends with a newline character
NF - Number of fields
awk '{ print NF, $1, NF }' emp.data
Computing and printing
awk '{ print $1, $2 * $3 }' emp.data
NR - Line numbers
awk '{ print NR, $0 }' emp.data
Putting text in the output
awk '{ print "Total pay for", $1, "is", $2 * $3 }' emp.data
Formatted output with printf
The printf statement lets you control exactly how output looks:
printf(format, value_1, value_2, … , value_n)
format: a string containing specifications that defines how values are displayed- a specification is a
%followed by a few characters that control the format of a value- the 1st specification tells how
value_1is to be printed - the 2nd specification tells how
value_2is to be printed - etc.
- the 1st specification tells how
Examples:
%s→ string%.2f→ number with 2 decimal places%-8s→ left-align text in 8-character width%6.2f→ number in 6-character width, 2 decimals
Important difference from print
With printf, no spaces or newlines are produced automatically, so don't forget the \n:
awk '{ printf("Total pay for %s is $%.2f\n", $1, $2 * $3) }' emp.data
awk '{ printf("%-8s $%6.2f\n", $1, $2 * $3) }' emp.data
Sorting the output
The easiest way to sort in order of increasing pay is to prefix the total pay and run the output through a sorting program:
awk '{ printf("%6.2f %s\n", $2 * $3, $0) }' emp.data | sort
Selection
A pattern by itself automatically prints all matching lines.
That's why many Awk programs are just a single pattern.
Selection by comparison
awk '$2 >= 20' emp.data
Selection by computation
awk '$2 * $3 > 200 { printf("$%.2f for %s\n", $2 * $3, $1) }' emp.data
Selection by text content
Exact match:
awk '$1 == "Susie"' emp.data
All lines that contain Susie anywhere (regular expression):
awk '/Susie/' emp.data
Combinations of patterns
Patterns can be combined with parentheses and the logical operators &&, ||, and !:
awk '$2 >= 20 || $3 >= 20' emp.data
Data validation
Real-world data often contains errors.
Awk is useful for checking that data is in the right format.
Examples:
awk 'NF != 3 { print $0, "number of fields is not equal to 3" }' emp.data
awk '$2 < 15 { print $0, "rate is too low" }' emp.data
awk '$2 > 25 { print $0, "rate exceeds $25 per hour"}' emp.data
awk '$3 < 0 { print $0, "negative hours worked" }' emp.data
awk '$3 > 60 { print $0, "too many hours worked" }' emp.data
BEGIN and END special patterns
BEGIN runs before any input is processed.
END runs after all input is processed.
awk 'BEGIN { print "NAME RATE HOURS"; print "" } { print }' emp.data
Notice that:
print ""prints a blank lineprintprints the current input line
Computing with Awk
An action can contain multiple statements.
Statements can be separated by:
new lines
$3 > 15 { emp = emp + 1 } END { print emp, "employes worked more than 15 hours" }or semicolons
$3 > 15 { emp = emp + 1 } ; END { print emp, "employes worked more than 15 hours" }
Counting
Counts employees who worked more than 15 hours using an emp variable:
awk '$3 > 15 { emp = emp + 1 } ; END { print emp, "employes worked more than 15 hours" }' emp.data
Shorter version with the increment operator ++:
awk '$3 > 15 { emp++ } ; END { print emp, "employes worked more than 15 hours" }' emp.data
Awk variables for numbers:
- begin with the value
0 - aren't explicitely initialized
Computing sums and averages
Count total employees using the built-in NR variable:
awk 'END { print NR, "employees" }' emp.data
To compute the average pay:
awk '{ pay = pay + $2 * $3 } ; END { print NR, "employees" ; print "total pay is", pay ; print "average pay is", pay / NR }' emp.data
Handling text
Awk variables can hold strings of characters.
Find the employee who is paid the most per hour:
awk '
$2 > maxrate { maxrate = $2; maxemp = $1 }
END { print "highest hourly rate:", maxrate, "for", maxemp }
' emp.data
String concatenation
Collect all the employee names into a string by appending each name and a space to the previous value in the variable names:
awk '{ names = names $1 " " } ; END { print names }' emp.data
For each line:
the first statement concatenates 3 strings:
- the previous value of
names - the first field
- a space
- the previous value of
and assigns the resulting string to
names
The string concatenation operator:
- is represented by writing string values one after the other
- there is no explicit concatenation operator
- in hindsight, this design might not be ideal because it can lead to hard-to-spot errors
Awk variables for strings:
- begin life with the value
null string(a string containing no characters) - aren't explicitely initialized
Printing the last input line
Built-in variables like NR and $0 retain their value in an END action.
So one way to print the last line is:
awk 'END { print $0 }' emp.data
Built-in functions
Awk has built-in functions for computing values.
To compute the length of each person's name:
awk '{ print $1, length($1) }' emp.data
Counting lines, words, and characters
Use length, NF, and NR to count the number of lines, words, and characters:
awk '{ nc += length($0) + 1 ; nw += NF } END { print NR, "lines,", nw, "words,", nc, "characters" }' emp.data
We added 1 to nc to count the newline character at the end of each input line.
Control-flow statements
Control-flow statements are used inside actions to control how your program runs.
If-else statement
Use an if to defend against any potential division by zero in computing the average pay:
awk '$2 > 20 { n++; pay += $2 * $3 } END {
if (n > 0) {
print n, "high-pay employees Total pay is", pay, "Average pay is", pay/n
} else {
print "No employees are paid more than $30/hour"
}
}' emp.data
While statement
We write an Awk program in a file named interest1.awk:
- it shows how the value of an amount of money invested at a particular interest rate grows over a number of years
- using the formula
value = amount(1 + rate)^years
# interest1.awk
# Compute compound interest
# input: amount rate years
# output: compounded value at the end of each year
{ i = 1
while (i <= $3) {
printf("\t%.2f\n", $1 * (1 + $2) ^ i)
i++
}
}
In this program:
- the
\tin theprintfspecification string stands for a tab character - the
^is the exponentiation operator
To use it:
awk -f programs/interest1.awk
1000 .05 5
1000 .15 10
For statement
Here is the previous interest computation with a for:
# interest2.awk
# Compute compound interest
# input: amount rate years
# output: compounded value at the end of each year
{ for (i = 1; i <= $3; i++)
printf("\t%.2f\n", $1 * (1 + $2) ^ i)
}
FizzBuzz
Print the numbers from 1 to 100:
- if the number is divisible by 3, print "fizz"
- if the number is divisible by 5, print "buzz"
- if the number is divisible by both, print "fizzbuzz"
awk 'BEGIN {
for (i = 1; i <= 100; i++) {
if (i%15 == 0) # divisible by both 3 and 5
print i, "fizzbuzz"
else if (i%5 == 0)
print i, "buzz"
else if (i%3 == 0)
print i, "fizz"
else
print i
}
}'
Arrays
Print the input in reverse order:
awk '
{ line[NR] = $0 } # Remember each input line.
END {
for (i = NR; i > 0; i--) {
print line[i]
}
}' emp.data
The first action stores input lines in successive elements of the array line:
- the first line goes into
line[1] - the second line goes into
line[2] - and so on
The expression inside [] is called the subscript.
The subscripts in this example are numeric but they can be arbitrary strings of characters.
Useful one-liners
| Description | Command |
|---|---|
| Print the total number of input lines | awk 'END { print NR }' emp.data |
| Print the first 10 input lines | awk 'NR <= 10' emp.data |
| Print the tenth input line | awk 'NR == 10' emp.data |
| Print every tenth input line, starting with line 1 | awk 'NR % 10 == 1' emp.data |
| Print the last field of every input line | awk '{ print $NF }' emp.data |
| Print the last field of the last input line | awk 'END { print $NF }' emp.data |
| Print every input line with more than four fields | awk 'NF > 4' emp.data |
| Print every input line that does not have exactly four fields | awk 'NF != 4' emp.data |
| Print every input line in which the last field is greater than 4 | awk '$NF > 4' emp.data |
| Print the total number of fields in all input lines | awk '{ nf += NF } END { print nf }' emp.data |
| Print the total number of lines that contain `Beth` | awk '/Beth/ { nlines++ } END { print nlines }' emp.data |
| Print the largest field and the line that contains it | awk '$1 > max { max = $1; maxline = $0 } END { print max, maxline }' emp.data |
| Print every line that has at least one field | awk 'NF > 0' emp.data |
| Print every line longer than 80 characters | awk 'length($0) > 80' emp.data |
| Print the number of fields in every line followed by the line itself | awk '{ print NF, $0 }' emp.data |
| Print the first two fields, in opposite order, of every line | awk '{ print $2, $1 }' emp.data |
| Interchange the first two fields of every line and then print the line | awk '{ temp = $1; $1 = $2; $2 = temp; print }' emp.data |
| Print every line preceded by its line number | awk '{ print NR, $0 }' emp.data |
| Print every line with the first field replaced by the line number | awk '{ $1 = NR; print }' emp.data |
| Print every line after erasing the second field | awk '{ $2 = ""; print }' emp.data |
| Print in reverse order the fields of every line |
awk '{
for (i = NF; i > 0; i--) printf("%s ", $i)
printf("\n")
}' emp.data
|
| Print the sums of the fields of every line |
awk '{
sum = 0
for (i = 1; i <= NF; i++) sum = sum + $i
print sum
}' emp.data
|
| Add up all fields in all lines and print the sum |
awk '
{ for (i = 1; i <= NF; i++) sum = sum + $i }
END { print sum }
' emp.data
|
| Print every line after replacing each field by its absolute value |
awk '{
for (i = 1; i <= NF; i++) if ($i < 0) $i = -$i
print
}' emp.data
|
Passing parameters to Awk (ARGC and ARGV)
Built-ins variables:
ARGC→ number of argumentsARGV→ array of argumentsARGV[1]→ first argumentARGV[ARGC-1]→ last argumentARGV[0]→ name of the program (usuallyawk)
awk 'BEGIN { argv1 = ARGV[1] ; print argv1 }' foo
Using arguments in a script file
a program stored in a file (
$*is the shell notation for all the parameters to the script or function):# progfile awk 'BEGIN { argv0 = ARGV[0] argv1 = ARGV[1] argv2 = ARGV[2] print argv0, argv1, argv2 }' $*make the file executable:
$ chmod +x progfilethen call the program:
$ ./progfile foo bar awk foo bar
Awk for exploratory data analysis (EDA)
The purpose of exploratory data is to get a sense of what the data is, looking for both patterns and anomalies.
Be approximately right rather than exactly wrong. — John W. Tukey
Awk is great for quickly inspecting data.
Check data format
Validate field structure, each line should have 5 fields and the total field should be correct:
awk 'NF != 5 || $3 != $4 + $5' programs/titanic.tsv
Ensure all records have the same number of fields:
awk --csv '{ print NF }' programs/passengers.csv | sort | uniq -c | sort -nr
CSV dataset
CSV is not rigorously defined, but commonly:
- fields containing commas or quotes must be enclosed in double quotes
- any field may be surrounded by quotes (whether it contains commas and quotes or not)
- an empty field is just
"" - quotes inside a field are doubled:
""","""represents"," - fields may contain newline characters
Awk ≥ 2023 support the --csv argument which causes input lines to be split into fields according to this rule:
awk --csv 'NR > 1 { print $2 }' programs/passengers.csv
Generating CSV
Double each quote and surround the result with quotes:
# to_csv - convert to proper "..."
function to_csv(s) {
gsub(/"/, "\"\"", s)
return "\"" s "\""
}
This function can be used within a loop to insert commas between fields:
awk '
function to_csv(s) {
gsub(/"/, "\"\"", s)
return "\"" s "\""
}
function record_to_csv( s, i) {
for (i = 1; i < NF; i++) {
s = s to_csv($i) ","
}
s = s to_csv($NF) # No comma after the last field `$NF`.
return s
}
{
print record_to_csv()
}
' data/emp.data
Or for an array:
function array_to_csv(arr, s, i, n) {
n = length(arr)
for (i = 1; i <= n; i++) {
s = s to_csv(arr[i]) ","
}
return substr(s, 1, length(s) - 1) # Remove trailing comma.
}
Generating TSV
awk 'NR > 1 { OFS="\t" ; print $1, $2, $3}' data/emp.data > ~/Desktop/foo.tsv
Associative arrays
Awk supports associative arrays:
- the subscripts (or indices) of Awk arrays can be arbitrary strings:
- arrays that allow arbitrary strings as subscripts are called associative arrays
- other languages provide the same facility with dictionary, map or hashmap
- Awk has a special form of the
forstatement for iterating over the indices of an associative array:for (i in array) { statements }- the element of the array are visited in an unspecified order
- you can't count on any particular order
How many people are there in each category?
awk '
NR > 1 { types[$1] += $3 ; classes[$2] += $3 }
END {
for (type in types) {
print type, types[type]
}
print ""
for (class in classes) {
print class, classes[class]
}
}
' programs/titanic.tsv
Compute the survival rate for each category and pipe the result to sort:
awk 'NR > 1 { printf("%6s %6s %6.1f%%\n", $1, $2, 100 * $4 / $3) }' programs/titanic.tsv | sort -k3 -nr
Beer reviews dataset
Dataset source: kaggle.
File size:
wc data/beer_reviews.csv # Number of lines, words, and bytes.
Awk equivalent:
awk '
{ nc += length($0) + 1 ; words_num += NF }
END {
print NR, # Lines.
words_num,
nc,
FILENAME
}
' data/beer_reviews.csv
Fields of the CSV:
- brewery_id
- brewery_name
- review_time
- review_overall
- review_aroma
- review_appearance
- review_profilename
- beer_style
- review_palate
- review_taste
- beer_name
- beer_abv (alcohol content, percentage of alcohol by volume or ABV)
- beer_beerid
Find the strongest beer:
awk --csv '
NR > 1 && $12 > max_abv { max_abv = $12 ; brewery = $2 ; name = $11 }
END { print max_abv, brewery, name }
' data/beer_reviews.csv
The result is surprisingly high.
That raises a follow-up question: is this value an outlier or the tip of a substantial alcoholic iceberg:
awk --csv 'NR > 1 && $12 >= 10 { print $2, $11, $12 }' data/beer_reviews.csv
What about low-alcohol beer?
awk --csv 'NR > 1 && $12 <= 0.5 && $12 > 0 { print $2, $11, $12 }' data/beer_reviews.csv
What ratings are associated with high and low alcohol?
# Rating of high alcohol.
awk --csv '
$12 >= 10 { rate += $4 ; nrate++ }
END { print rate / nrate, nrate }
' data/beer_reviews.csv
# Rating of low alcohol.
awk --csv '
$12 <= 0.5 && $12 > 0 { rate += $4 ; nrate++ }
END { print rate / nrate, nrate }
' data/beer_reviews.csv
# Average rating.
awk --csv '
$12 > 0 { rate += $4 ; nrate++ }
END { print rate / nrate, nrate }
' data/beer_reviews.csv
We use $12 > 0 because some fields don't have a ABV.
When exploring data, always check:
- how many fields are empty?
- how many fields have an explicitely non-useful value like "N/A"?
- what is the range of values in a column (min/max)?
- what are the distinct values?
Automating these checks with small scripts can save a lot of time.