Python正则表达式匹配HTML页面编码
html页面一般都会指定一个编码,如何获取到是处理html页面的第一步,因为错误的编码必然带来后面处理的问题。这里我用python的正则表达式写了个:
importre a=["<metahttp-equiv="Content-Type"content="text/html;charset=utf-8"/>", '<metahttp-equiv=Content-Typecontent="text/html;charset=gb2312">', '<metahttp-equiv="Content-Type"content="text/html;charset=iso-8859-1">', '<metahttp-equiv="Content-Type"content="text/html;charset=gb2312"/>', '<metahttp-equiv="content-type"content="text/html;charset=utf-8"/>', '<metahttp-equiv="Content-Type"content="text/html;charset=gb2312"/>', '<metahttp-equiv="Content-Type"content="text/html;charset=gb2312"/>' ] b="<meta[]+http-equiv=["']?content-type["']?[]+content=["']?text/html;[]*charset=([0-9-a-zA-Z]+)["']?" B=re.compile(b,re.IGNORECASE) foraxina: r1=B.search(ax) ifr1: printr1.group() printr1.group(1),len(r1.group()) else: print'notmatch'
热门推荐
10 朋友新年祝福语大全 简短
11 新年祝福语简短大方兔年
12 搬新家礼物祝福语简短
13 同学见面花束祝福语简短
14 五一假期祝福语幽默简短
15 离职欢送敬酒祝福语简短
16 对学弟的祝福语简短
17 考老师辞职祝福语简短
18 祝福语驱散霉运的话简短