Showing posts with label command-line tool. Show all posts
Showing posts with label command-line tool. Show all posts

How to extract or remove lines with grep?

 


In this post, we go through some examples of extracting or removing lines of a file with the grep command line tool.

Suppose we have a file text.txt with the following content.


sed command stands for stream editor
sed allows us to substitute, insert or delete, find and replace strings in a file
sed is quite useful and sed is awesome!

How to filter lines having "useful"? 

Command:

$grep useful text.txt
Output:

sed is quite useful and sed is awesome!

How to filter those lines with case insensitive? 

For example, filtering lines with sed or SED

Suppose we have an updated file text.txt with the following content.


SED command stands for stream editor
SED allows us to substitute, insert or delete, find and replace strings in a file
SED is quite useful and sed is awesome!
Command (with -i option):

$grep -i sed
Ouput:

SED command stands for stream editor
SED allows us to substitute, insert or delete, find and replace strings in a file
SED is quite useful and sed is awesome!

How to filter lines with regex patterns? 

For example, lines starting with SED and containing substitute 


Command (using -E option):

$grep -E '^SED.*substitute.*'
Ouput:

SED allows us to substitute, insert or delete, find and replace strings in a file

Bash scripting - How to delete lines of a file?

 



In this tutorial, we play with deleting some lines of a file using Bash.

Content
  • Create a simple file for the tutorial
  • How to delete n-th line?
  • How to delete several lines within a range?
  • How to delete lines with patterns?

Create a simple file for the tutorial. 


First, let's create a simple file for this tutorial using the following command with seq and tee.

$seq 6 | tee delete.txt

1
2
3
4
5
6


How to delete n-th line? with sed

$sed '5d' delete.txt	# delete 5th line
Output:

1
2
3
4
6

$sed '$d' delete.txt	# delte the last line, $ inidicates the last line in sed
Output:

1
2
3
4
5


How to delete several lines within a range?

$sed '2,4d' delete.txt
Output:

1
5
6


How to delete lines with patterns?

$sed '/2/d' delete.txt	# /pattern/ to delete
Output:

1
3
4
5
6

Bash scripting - How to filter lines of a file?



In this tutorial, we play with filtering some lines of a file using Bash.

Content
  • Create a simple file for the tutorial
  • How to filter the first/last n lines?
  • How to remove n lines?
  • How to filter specific lines?

Create a simple file for the tutorial. 


First, let's create a simple file for this tutorial using the following command with seq and tee.

$seq 6 | tee filtering.txt

1
2
3
4
5
6

How to filter the first n lines?


Given the dummy file - filtering.txt -we've just created, we now move on to several ways of filtering n=3 lines as an example.

Using head

$< filtering.txt head -n 3 
Using sed

$< filtering.txt sed -n '1,3p'
Using awk (NR refers to Number of Records)

$< filtering.txt awk 'NR<=3'
Filtering the last 3 lines is straighforward with tail

$< filtering.txt tail -n 3
How to filter the last n lines?

$< lines tail -n 3

How to remove n lines? (e.g., n=3)


Using tail

$< filtering.txt tail -n +4
Using sed

$< filtering.txt sed ‘1,3d’
How to remove the last n lines?

$< lines head -n -3

How to filter specific lines? For example, 4-6 lines


Using sed

$ < filtering.txt sed -n '4,6p'
Using awk

$ < filtering.txt awk '(NR>=4)&&(NR<=6)'
Using head and tail

$ < filtering.txt head -n 6 | tail -n 3  

How to filter odd or even lines?

 
Using sed

$< filtering.txt sed -n '1~2p'	# '0~2p' for even 
Using awk

$< filtering.txt awk 'NR%2'	# '(NR+1)%2' for even

How to count the number of occorrences of each word in each line of a file?



Suppose we have a file - words.txt - with its content where each line is a word and we would like to count each unique word and sort them in ascending order. 

More specifically, we have words.txt as follows:

  1 foo
  2 bar
  3 foo
  4 foo
  5 bar
And we would like to get the result as follows:

  word,count
  foo,3
  bar,2

We use the following steps to approach this problem:
  1. Read and sort the content to count unique words and their counts
  2. Sort in reversing order by comparing their count values
  3. Print the count and word informatin in the "<count>,<word>" format
  4. Finally, append the column names "word,count" and print out





1. Read and sort the content to count unique words and their counts

$cat words.txt | sort | uniq -c
Note you can always check options available for a command, e.g., uniq, with --help option

$uniq --help
Usage: uniq [OPTION]... [INPUT [OUTPUT]]
Filter adjacent matching lines from INPUT (or standard input),
writing to OUTPUT (or standard output).

With no options, matching lines are merged to the first occurrence.

Mandatory arguments to long options are mandatory for short options too.
  -c, --count           prefix lines by the number of occurrences
This will give us the following results so far:

2 bar
3 foo

2. Sort in reversing order by comparing their count values

$cat words.txt | sort | uniq -c | sort -rn
where -rn indicate comparing numerical values and sort in reversed/descending order. And we get the results so far as:

3 foo
2 bar

3. Print the count and word informatin in the "<count>,<word>" format using awk command (awk is abbreviated from the names of the developers – Aho, Weinberger, and Kernighan)

cat words.txt | sort | uniq -c | sort -nr | awk '{print $2","$1}'
Now, we are almost done with the output as

foo,3
bar,2
and just need to append headers on top.

4. Finally, append the column names "word,count" and print out using header command line tool from dsutils

$cat words.txt | sort | uniq -c | sort -nr | awk '{print $2","$1}' | header -a word,count

Putting everythign together, we have the final solution as follows:

$cat words.txt | sort | uniq -c | sort -nr | awk '{print $2","$1}' | header -a word,count
with our desired output

  word,count
  foo,3
  bar,2