Linux 正则表达式基础 & grep 匹配正则表达式

Linux 正则表达式

  • 正则表达式作为一个 pattern,将 pattern 与要搜索的字符串进行匹配,以便查找一个或多个字符串。
  • 正则表达式,自成体系,由普通字符(例如字符 a 到 z)和元字符组成的文字模式。
    • 普通字符:没有显式指定为元字符的所有可打印和不可打印字符字符,包括所有大写和小写字母、所有数字、所有标点符号和其他一些符号。
    • 元字符:出了普通字符之外的字符。
  • 正则表达式,工具(vim、grep、less等)和程序语言(Perl、Python、C等)都使用正则表达式。

正则表达式分类:

  • 普通正则表达式
  • 扩展正则表示,支持更多的元字符。

环境准备

[laoma@shell ~]$ vim words
cat
category
acat
concatenate
cbt
c1t
cCt
c-t
c.t
dog

普通字符

[laoma@shell ~]$ grep 'cat' words
cat
category
acat
concatenate

字符集

.

匹配除换行符(\n\r)之外的任何单个字符,相等于\[^\n\r]

[laoma@shell ~]$ grep 'c.t' words
cat
category
acat
concatenate
cbt
c1t
cCt
c-t
c.t

[…]

匹配 [...] 中的任意一个字符。

[laoma@shell ~]$ grep 'c[ab]t' words
cat
category
acat
concatenate
cbt

[a-z] [A-Z] [0-9]

  • [a-z],匹配所有小写字母。
  • [A-Z],匹配所有大写字母。
  • [0-9],匹配所有数字。
[laoma@shell ~]$ grep 'c[a-z]t' words
cat
category
acat
concatenate
cbt
[laoma@shell ~]$ grep 'c[A-Z]t' words
cCt	
[laoma@shell ~]$ grep 'c[0-9]t' words
c1t
[laoma@shell ~]$ grep 'c[a-z0-9]t' words
cat
category
acat
concatenate
cbt
c1t
[laoma@shell ~]$ grep 'c[a-zA-Z0-9]t' words
cat
category
acat
concatenate
cbt
c1t
cCt
# 要想匹配-符号,将改符号写在第一个位置
[laoma@shell ~]$ grep 'c[-a-zA-Z0-9]t' words
cat
category
acat
concatenate
cbt
c1t
cCt
c-t

[^…]

匹配除了 [...] 中字符的所有字符。

[laoma@shell ~]$ grep 'c[^ab]t' words
c1t
cCt
c-t
c.t
# ^放中间会被当做普通字符
[laoma@shell ~]$ grep 'c[a^b]t' words
cat
category
acat
concatenate
cbt

\

将下一个字符标记为或特殊字符、或原义字符、或向后引用、或八进制转义符。

例如, ‘n’ 匹配字符 ‘n’。\n 匹配换行符。序列 \\ 匹配 \,而 \( 则匹配 (

[laoma@shell ~]$ grep 'c\.t' words
c.t
# 匹配普通字符,虽然可以匹配,但强烈建议不要在前面加\
[laoma@shell ~]$ grep 'c\at' words
cat
category
acat
concatenate

|

| 符号是扩展表达式中元字符,指明两项之间的一个选择。要匹配 |,请使用 \|

# 使用egrep或者grep -E 匹配
[laoma@shell ~]$ egrep 'cat|dog' words
cat
category
acat
concatenate
dog

# 或者
[laoma@shell ~]$ grep -E 'cat|dog' words
cat
category
acat
concatenate
dog
选项描述
[[:digit:]]数字: 0 1 2 3 4 5 6 7 8 9 等同于[0-9]
[[:xdigit:]]十六进制数字: 0 1 2 3 4 5 6 7 8 9 A B C D E F a b c d e f等同于[0-9a-fA-F]
[[:lower:]]小写字母:在 C 语言环境和ASCII字符编码中,对应于[a-z]
[[:upper:]]大写字母:在 C 语言环境和ASCII字符编码中,对应于[A-Z]
[[:alpha:]]字母字符:[[:lower:]和[[:upper:]];在C语言环境和ASCII字符编码中,等同于**[A-Za-z]**
[[:alnum:]]字母数字字符:[:alpha:]和[:digit:];在C语言环境和ASCII字符编码中,等同于**[0-9A-Za-z]**
[[:blank:]]或者[[:space:]]空白字符:在 C 语言环境中,它对应于制表符、换行符、垂直制表符、换页符、回车符和空格。
[[:punct:]]标点符号:在C语言环境和ASCII字符编码中,它对应于!" # $ % &'()*+,-./:;<=>?@[]^_`{|}~
[[:print:]]或者 [[:graph:]]可打印字符: [[:alnum:]]、[[:punct:]]。
[[:cntrl:]]控制字符。在 ASCII中, 这些字符对应八进制代码000到037和 177 (DEL)。

非打印字符

终端中不显示的字符,例如换行符。

字符描述
\cx匹配由x指明的控制字符。例如, \cM 匹配一个 Control-M 或回车符。x 的值必须为 A-Z 或 a-z 之一。否则,将 c 视为一个原义的 ‘c’ 字符。
\f匹配一个换页符。等价于 \x0c\cL
\n匹配一个换行符。等价于 \x0a\cJ
\r匹配一个回车符。等价于 \x0d\cM
\s匹配任何空白字符,包括空格、制表符、换页符等等。等价于 [ \f\n\r\t\v]。注意 Unicode 正则表达式会匹配全角空格符。
\S匹配任何非空白字符。等价于 [^ \f\n\r\t\v]
\w匹配字母、数字、下划线。等价于 [A-Za-z0-9_]
\W匹配任何非单词字符。等价于[^A-Za-z0-9_]
\t匹配一个制表符。等价于 \x09\cI
\v匹配一个垂直制表符。等价于 \x0b\cK

grep 命令支持\w\W\s\S

定位符

^

匹配行首位置。

[laoma@shell ~]$ grep '^cat' words
cat
category

$

匹配行末位置。

[laoma@shell ~]$ grep 'cat$' words
cat
acat

示例:

# 查看 /var/log/message Aug 19 14:01 到 Aug 19 14:06 时间段发生的事件
[laoma@shell ~]$ sudo cat /var/log/messages | egrep '^Aug 19 14:0[1-6]# 只包含cat的行
[laoma@shell ~]$ cat words | grep '^cat$'
cat

# 排除/etc/profile文件中以#开头的行
[laoma@shell ~]$ cat /etc/profile | egrep  '^[^#]'

# 查询/etc/profile文件中有效行
[laoma@shell ~]$ cat /etc/profile | egrep -v '^#|^$'
# -v 取反,不显示匹配的内容

[root@server ~]# egrep -v '^\s*#|^$' /etc/profile

# 查看系统中有哪些仓库
[laoma@shell ~]$ yum-config-manager | grep '^\[' | grep -v main
[base]
[epel]
[extras]
[updates]

# 查看 /var/log/message Aug 19 14:01 到 Aug 19 14:06 时间段发生的事件
[laoma@shell ~]$ sudo cat /var/log/messages | egrep '^Aug 19 14:0[1-6]

\b

匹配一个单词边界。

[laoma@shell ~]$ echo hello cat kitty >> words 
[laoma@shell ~]$ grep '\bcat' words
cat
category
hello cat kitty
[laoma@shell ~]$ grep 'cat\b' words
cat
acat
hello cat kitty
[laoma@shell ~]$ grep '\bcat\b' words
cat
hello cat kitty

\B

非单词边界匹配。

[laoma@shell ~]$ grep '\Bcat\B' words
concatenate

\< 和 \>

  • \< ,匹配一个单词左边界。
  • \>,匹配一个单词右边界。
[laoma@shell ~]$ grep '\<cat' words
cat
category
hello cat kitty

[laoma@shell ~]$ grep 'cat\>' words
cat
acat
hello cat kitty

限定次数

*

匹配前面的子表达式任意次数

[laoma@shell ~]$ echo dg >> words 
[laoma@shell ~]$ echo doog >> words 
[laoma@shell ~]$ grep 'do*g' words
dog
dg
doog

+

+ 是扩展表达式元字符,匹配前面的子表达式一次以上次数

[laoma@shell ~]$ egrep 'do+g' words
dog
doog

?

? 是扩展表达式元字符,匹配前面的子表达式一次以下次数

[laoma@shell ~]$ egrep 'do?g' words
dog
dg

{n}

{} 是扩展表达式元字符,用于匹配特定次数。例如:{n},配置n次。

[laoma@shell ~]$ egrep 'do{2}g' words
doog

{m,n}

{m,n},是扩展表达式元字符,用于匹配次数介于m-n之间。

[laoma@shell ~]$ echo dooog >> words
[laoma@shell ~]$ echo doooog >> words 

[laoma@shell ~]$ egrep 'do{2,3}g' words
doog
dooog

{m,}

{m,},是扩展表达式元字符,匹配前面的子表达式m次以上次数

[laoma@shell ~]$ egrep 'do{2,}g' words
doog
dooog
doooog

{,n}

{,n},是扩展表达式元字符,匹配前面的子表达式n次以下次数

[laoma@shell ~]$ egrep 'do{,3}g' words
dog
doog
dg
dooog

()

标记一个子表达式。

[laoma@shell ~]$ echo dogdog >> words 
[laoma@shell ~]$ echo dogdogdog >> words 
[laoma@shell ~]$ echo dogdogdogdog >> words 

[laoma@shell ~]$ egrep '(dog){2,3}' words
dogdog
dogdogdog
dogdogdogdog

[laoma@shell ~]$ egrep '(dog){2,}' words
dogdog
dogdogdog
dogdogdogdog

综合案例

如何过滤出以下内容中所有有效IPv4地址?

0.0.0.0
1.1.1.1
11.11.11.111
111.111.111.111
999.9.9.9
01.1.1.1
10.0.0.0
0.1.1.1
266.1.1.1
248.1.1.1
256.1.1.1

参考答案

\b(([1-9]?[0-9])|(1[0-9]{2})|(2[0-4][0-9])|(25[0-4]))(\.(([1-9]?[0-9])|(1[0-9]{2})|(2[0-4][0-9])|(25[0-4]))){3}\b

grep -Eo ‘\b((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?).){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\b’ 文件名

1. 命令各部分作用
  • grep:文本搜索工具

  • -E:启用扩展正则表达式(无需对 | () 等符号转义)

  • -o:仅输出匹配到的内容(而非整行)

  • 正则部分:

    \b((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\b
    
    • \b:单词边界,避免匹配到 192.168.1.1000 中的 192.168.1.100 这类片段

    • (25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)
      

      :匹配 0-255 的数字段:

      • 25[0-5] → 250-255
      • 2[0-4][0-9] → 200-249
      • [01]?[0-9][0-9]? → 0-199([01]? 匹配可选的 0/1,后接 1-2 位数字)
    • \.:匹配点号(需转义,否则 . 代表任意字符)

    • {3}:重复前一段(数字段+.)3 次,即匹配前 3 段(如 192.168.1.

    • 最后一段:单独匹配第 4 个数字段(无末尾点号)

2. 使用示例

假设文件 ips.txt 内容:

有效IP:192.168.1.1、10.0.0.255、255.255.255.255
无效IP:256.0.0.1、192.168.1.、abc.123.45.67
混合内容:服务器IP是172.16.0.0,端口8080

执行命令:

grep -Eo '\b((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\b' ips.txt

输出结果:

192.168.1.1
10.0.0.255
255.255.255.255
172.16.0.0
3. 注意事项
  • 匹配范围:仅匹配标准 IPv4 地址,不支持带掩码(如 192.168.1.1/24)或端口(如 192.168.1.1:80)的格式
  • 空值 / 前导零:会匹配 0.0.0.0010.01.001.1 这类合法(但非常规)的 IP
  • 性能:正则较复杂,处理超大文件时效率略低,可结合 grep -F 先过滤粗匹配再精准校验

反向引用

对一个正则表达式模式或部分模式两边添加圆括号将导致相关匹配存储到一个临时缓冲区中,所捕获的每个子匹配都按照在正则表达式模式中从左到右出现的顺序存储。缓冲区编号从 1 开始,最多可存储 99 个捕获的子表达式。

每个缓冲区都可以使用 \N 访问,其中 N 为一个标识特定缓冲区的一位或两位十进制数。

\N 这用引用方式称之为反向引用。

[laoma@shell ~]$ echo 'laoma laoniu laohu laoma laoniu laohu' | \
> egrep -o '(laoma) (laoniu).*\1'

# 过滤结果如下
laoma laoniu laohu laoma

[laoma@shell ~]$ echo 'Is is the cost of of gasoline going up up?' | \
> egrep -o '\b([a-z]+) \1\b' 
# 过滤结果如下
of of
up up

[laoma@shell ~]$ echo 'Is is the cost of of of gasoline going up up?' | egrep -o '(\b[a-z]+\b\s+)\1{1,}'
# 过滤结果如下
of of of

[laoma@shell ~]$ echo 'Is is the cost of  of of gasoline going up up?' | egrep -o '(\b[a-z]+\b\s+)\1{1,}'
# 过滤结果如下
of of
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值