Text Processing
grep — Search
grep prints lines that match a pattern. It is the first tool to reach for on logs and config files.
Useful Options
grep "error" app.log # matching lines
grep -i "error" app.log # ignore case
grep -c "error" app.log # count matches
grep -n "error" app.log # show line numbers
grep -v "debug" app.log # invert: lines that do NOT match
grep -r "TODO" src/ # recurse through directories
grep -l "error" *.log # print only filenames with matches
grep -o "[0-9]\+" file # print only the matched part
grep -q "ready" file # quiet: status only, for 'if'
Basic vs Extended Regex
grep 'abc\{3\}' f # BRE: braces escaped
grep -E 'abc{3}' f # ERE: -E, readable
grep -F 'a.b*c' f # fixed string: no regex, fastest
grep -E '^[0-9]{3}-[0-9]{4}$' phones
Always quote the pattern. An unquoted
* or ? is expanded by the shell before grep ever sees it.
sed — Stream Editor
Substitution
sed edits a stream line by line. Substitution is its most common job.
sed 's/foo/bar/' file # first occurrence per line
sed 's/foo/bar/g' file # every occurrence (global)
sed 's/foo/bar/2' file # second occurrence only
sed -n '3,7p' file # print only lines 3..7
sed '/^#/d' file # delete comment lines
sed -E 's/[0-9]+/N/g' file # extended regex
In-Place Editing
sed -i 's/old/new/g' file # GNU sed: edit the file, no output
sed -i.bak 's/old/new/g' file # GNU sed: keep file.bak as backup
sed -i '' 's/old/new/g' file # BSD/macOS sed requires an argument
GNU and BSD
sed -i differ. For portable scripts write to a temp file and move it: sed 's/a/b/g' f > f.tmp && mv f.tmp f.
Other useful commands: -e chains several expressions, and a leading address restricts the command to selected lines.
sed -e '/^$/d' -e 's/[[:space:]]\+$//' file # drop blanks, strip trailing space
sed 's/^/ /' file # indent every line
awk — Fields & Records
Fields and Records
awk splits each line into fields. $1, $2 … are the fields; $0 is the whole line; NF is the field count.
awk '{ print $1 }' file # first column
awk -F: '{ print $1, $3 }' /etc/passwd # custom delimiter (colon)
awk '{ print NR, $0 }' file # number every line
awk 'NF > 0 { print }' file # drop blank lines
Patterns and Blocks
awk '$3 > 100 { print $1 }' data # condition on a field
awk '/ERROR/ { c++ } END { print c }' log # count matches
awk 'NR==1 { print "header: " $0 }' file
Practical Aggregation
# Sum the third column and print the average
awk '{ sum += $3; n++ } END { printf "%.2f\n", sum/n }' data
# Join fields with a separator
awk -F: '{ print $1 ": " $7 }' /etc/passwd
# Group and total by first column (like a mini GROUP BY)
awk '{ total[$1] += $2 } END { for (k in total) print k, total[k] }' sales
Unlike shell loops, awk reads a whole file in one process — often an order of magnitude faster than piping line by line.
tr & cut
tr 'a-z' 'A-Z' < file # uppercase
tr -d '\r' < dos.txt > unix.txt # delete carriage returns
tr -s ' ' < file # squeeze repeated spaces
cut -d: -f1 /etc/passwd # first field
cut -c1-10 file # first ten characters
Regular Expressions
Basic vs Extended
| Token | Means | BRE | ERE |
|---|---|---|---|
. | any character | yes | yes |
* | zero or more | yes | yes |
+ | one or more | \+ | yes |
? | zero or one | \? | yes |
| | alternation | \| | yes |
( ) | grouping | \( \) | yes |
{n} | exactly n | \{n\} | yes |
Anchors and Classes
grep -E '^[A-Za-z_][A-Za-z0-9_]*$' ids # a valid identifier
grep -E '^[0-9]{4}-[0-9]{2}-[0-9]{2}$' dates # ISO date
grep -E '[[:alnum:]]+@[[:alnum:].]+' mail # crude e-mail match
Pitfalls
- Unquoted patterns are globbed by the shell before the tool sees them.
- BRE and ERE disagree on
+ ? | ( ) { }— pass-Efor readability. sed -isyntax differs between GNU and BSD — use a temp file for portability.- Regular expressions are not a parser. For structured data (JSON, XML) use a real tool.