2.6 字符串忽略大�写的�索替�¶

问题¶

你需�以忽略大�写的方��索与替�文本字符串

解决方案¶

为了在文本�作时忽略大�写,你需�在使用 re 模�的时候给这些�作�供 re.IGNORECASE 标志�数。比如:

>>> text = 'UPPER PYTHON, lower python, Mixed Python'
>>> re.findall('python', text, flags=re.IGNORECASE)
['PYTHON', 'python', 'Python']
>>> re.sub('python', 'snake', text, flags=re.IGNORECASE)
'UPPER snake, lower snake, Mixed snake'
>>>

最�的那个例��示了一个�缺陷,替�字符串并�会自动跟被匹�字符串的大�写��一致。 为了修�这个,你�能需�一个辅助函数,就�下�的这样:

def matchcase(word):
    def replace(m):
        text = m.group()
        if text.isupper():
            return word.upper()
        elif text.islower():
            return word.lower()
        elif text[0].isupper():
            return word.capitalize()
        else:
            return word
    return replace

下�是使用上述函数的方法:

>>> re.sub('python', matchcase('snake'), text, flags=re.IGNORECASE)
'UPPER SNAKE, lower snake, Mixed Snake'
>>>

译者注: matchcase('snake') 返回了一个回调函数(�数必须是 match 对象),��一节�到过, sub() 函数除了接�替�字符串外,还能接�一个回调函数。

讨论¶

对于一般的忽略大�写的匹��作,简�的传递一个 re.IGNORECASE 标志�数就已�足够了。 但是需�注�的是,这个对于�些需�大�写转�的Unicode匹��能还�够, �考2.10�节了解更多细节。