6.7 利用命�空间解�XML文档¶
问题¶
ä½ æƒ³è§£æž�æŸ�个XML文档,文档ä¸ä½¿ç”¨äº†XML命å��空间。
解决方案¶
考虑下�这个使用了命�空间的文档:
<?xml version="1.0" encoding="utf-8"?>
<top>
<author>David Beazley</author>
<content>
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<title>Hello World</title>
</head>
<body>
<h1>Hello World!</h1>
</body>
</html>
</content>
</top>
å¦‚æžœä½ è§£æž�è¿™ä¸ªæ–‡æ¡£å¹¶æ‰§è¡Œæ™®é€šçš„æŸ¥è¯¢ï¼Œä½ ä¼šå�‘现这个并ä¸�æ˜¯é‚£ä¹ˆå®¹æ˜“ï¼Œå› ä¸ºæ‰€æœ‰æ¥éª¤éƒ½å�˜å¾—相当的ç¹�ç��。
>>> # Some queries that work
>>> doc.findtext('author')
'David Beazley'
>>> doc.find('content')
<Element 'content' at 0x100776ec0>
>>> # A query involving a namespace (doesn't work)
>>> doc.find('content/html')
>>> # Works if fully qualified
>>> doc.find('content/{http://www.w3.org/1999/xhtml}html')
<Element '{http://www.w3.org/1999/xhtml}html' at 0x1007767e0>
>>> # Doesn't work
>>> doc.findtext('content/{http://www.w3.org/1999/xhtml}html/head/title')
>>> # Fully qualified
>>> doc.findtext('content/{http://www.w3.org/1999/xhtml}html/'
... '{http://www.w3.org/1999/xhtml}head/{http://www.w3.org/1999/xhtml}title')
'Hello World'
>>>
ä½ å�¯ä»¥é€šè¿‡å°†å‘½å��空间处ç�†é€»è¾‘包装为一个工具类æ�¥ç®€åŒ–这个过程:
class XMLNamespaces:
def __init__(self, **kwargs):
self.namespaces = {}
for name, uri in kwargs.items():
self.register(name, uri)
def register(self, name, uri):
self.namespaces[name] = '{'+uri+'}'
def __call__(self, path):
return path.format_map(self.namespaces)
通过下�的方�使用这个类:
>>> ns = XMLNamespaces(html='http://www.w3.org/1999/xhtml')
>>> doc.find(ns('content/{html}html'))
<Element '{http://www.w3.org/1999/xhtml}html' at 0x1007767e0>
>>> doc.findtext(ns('content/{html}html/{html}head/{html}title'))
'Hello World'
>>>
讨论¶
解��有命�空间的XML文档会比较��。
上é�¢çš„ XMLNamespaces 仅仅是å…�è®¸ä½ ä½¿ç”¨ç¼©ç•¥å��代替完整的URI将其å�˜å¾—ç¨�微简æ´�一点。
很ä¸�幸的是,在基本的 ElementTree è§£æž�䏿²¡æœ‰ä»»ä½•途径获å�–命å��空间的信æ�¯ã€‚
ä½†æ˜¯ï¼Œå¦‚æžœä½ ä½¿ç”¨ iterparse() 函数的è¯�å°±å�¯ä»¥èŽ·å�–更多关于命å��空间处ç�†èŒƒå›´çš„ä¿¡æ�¯ã€‚例如:
>>> from xml.etree.ElementTree import iterparse
>>> for evt, elem in iterparse('ns2.xml', ('end', 'start-ns', 'end-ns')):
... print(evt, elem)
...
end <Element 'author' at 0x10110de10>
start-ns ('', 'http://www.w3.org/1999/xhtml')
end <Element '{http://www.w3.org/1999/xhtml}title' at 0x1011131b0>
end <Element '{http://www.w3.org/1999/xhtml}head' at 0x1011130a8>
end <Element '{http://www.w3.org/1999/xhtml}h1' at 0x101113310>
end <Element '{http://www.w3.org/1999/xhtml}body' at 0x101113260>
end <Element '{http://www.w3.org/1999/xhtml}html' at 0x10110df70>
end-ns None
end <Element 'content' at 0x10110de68>
end <Element 'top' at 0x10110dd60>
>>> elem # This is the topmost element
<Element 'top' at 0x10110dd60>
>>>
最å�Žä¸€ç‚¹ï¼Œå¦‚æžœä½ è¦�处ç�†çš„XML文本除了è¦�使用到其他高级XML特性外,还è¦�使用到命å��空间,
å»ºè®®ä½ æœ€å¥½æ˜¯ä½¿ç”¨ lxml 函数库æ�¥ä»£æ›¿ ElementTree 。
例如,lxml 对利用DTD验è¯�文档ã€�更好的XPath支æŒ�和一些其他高级XML特性ç‰éƒ½æ��供了更好的支æŒ�。
这一å°�节其实å�ªæ˜¯æ•™ä½ 如何让XMLè§£æž�ç¨�微简å�•一点。