6.3 解�简�的XML数�¶

问题¶

你想从一个简�的XML文档中��数�。

解决方案¶

�以使用 xml.etree.ElementTree 模�从简�的XML文档中��数�。 为了演示,�设你想解�Planet Python上的RSS�。下�是相应的代�:

from urllib.request import urlopen
from xml.etree.ElementTree import parse

# Download the RSS feed and parse it
u = urlopen('http://planet.python.org/rss20.xml')
doc = parse(u)

# Extract and output tags of interest
for item in doc.iterfind('channel/item'):
    title = item.findtext('title')
    date = item.findtext('pubDate')
    link = item.findtext('link')

    print(title)
    print(date)
    print(link)
    print()

�行上�的代�,输出结果类似这样:

Steve Holden: Python for Data Analysis
Mon, 19 Nov 2012 02:13:51 +0000
http://holdenweb.blogspot.com/2012/11/python-for-data-analysis.html

Vasudev Ram: The Python Data model (for v2 and v3)
Sun, 18 Nov 2012 22:06:47 +0000
http://jugad2.blogspot.com/2012/11/the-python-data-model.html

Python Diary: Been playing around with Object Databases
Sun, 18 Nov 2012 20:40:29 +0000
http://www.pythondiary.com/blog/Nov.18,2012/been-...-object-databases.html

Vasudev Ram: Wakari, Scientific Python in the cloud
Sun, 18 Nov 2012 20:19:41 +0000
http://jugad2.blogspot.com/2012/11/wakari-scientific-python-in-cloud.html

Jesse Jiryu Davis: Toro: synchronization primitives for Tornado coroutines
Sun, 18 Nov 2012 20:17:49 +0000
http://feedproxy.google.com/~r/EmptysquarePython/~3/_DOZT2Kd0hQ/

很显然,如果你想�进一步的处�,你需�替� print() 语��完�其他有趣的事。

讨论¶

在很多应用程�中处�XML编�格�的数�是很常�的。 �仅因为XML在Internet上�已�被广泛应用于数�交�, �时它也是一�存储应用程�数�的常用格�(比如字处�,音�库等)。 接下�的讨论会先�定读者已�对XML基础比较熟悉了。

在很多情况下,当使用XML�仅仅存储数�的时候,对应的文档结构�常紧凑并且直观。 例如,上�例�中的RSS订阅�类似于下�的格�:

<?xml version="1.0"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/">
    <channel>
        <title>Planet Python</title>
        <link>http://planet.python.org/</link>
        <language>en</language>
        <description>Planet Python - http://planet.python.org/</description>
        <item>
            <title>Steve Holden: Python for Data Analysis</title>
            <guid>http://holdenweb.blogspot.com/...-data-analysis.html</guid>
            <link>http://holdenweb.blogspot.com/...-data-analysis.html</link>
            <description>...</description>
            <pubDate>Mon, 19 Nov 2012 02:13:51 +0000</pubDate>
        </item>
        <item>
            <title>Vasudev Ram: The Python Data model (for v2 and v3)</title>
            <guid>http://jugad2.blogspot.com/...-data-model.html</guid>
            <link>http://jugad2.blogspot.com/...-data-model.html</link>
            <description>...</description>
            <pubDate>Sun, 18 Nov 2012 22:06:47 +0000</pubDate>
        </item>
        <item>
            <title>Python Diary: Been playing around with Object Databases</title>
            <guid>http://www.pythondiary.com/...-object-databases.html</guid>
            <link>http://www.pythondiary.com/...-object-databases.html</link>
            <description>...</description>
            <pubDate>Sun, 18 Nov 2012 20:40:29 +0000</pubDate>
        </item>
        ...
    </channel>
</rss>

xml.etree.ElementTree.parse() 函数解�整个XML文档并将其转��一个文档对象。 然�,你就能使用 find() �iterfind() 和 findtext() 等方法��索特定的XML元素了。 这些函数的�数就是�个指定的标签�,例如 channel/item 或 title 。

�次指定�个标签时,你需��历整个文档结构。�次�索�作会从一个起始元素开始进行。 �样,�次�作所指定的标签�也是起始元素的相对路径。 例如,执行 doc.iterfind('channel/item') ��索所有在 channel 元素下�的 item 元素。 doc 代表文档的最顶层(也就是第一级的 rss 元素)。 然�接下�的调用 item.findtext() 会从已找到的 item 元素�置开始�索。

ElementTree 模�中的�个元素有一些��的属性和方法,在解�的时候�常有用。 tag 属性包�了标签的�字,text 属性包�了内部的文本,而 get() 方法能获�属性值。例如:

>>> doc
<xml.etree.ElementTree.ElementTree object at 0x101339510>
>>> e = doc.find('channel/title')
>>> e
<Element 'title' at 0x10135b310>
>>> e.tag
'title'
>>> e.text
'Planet Python'
>>> e.get('some_attribute')
>>>

有一点�强调的是 xml.etree.ElementTree 并�是XML解�的唯一方法。 对于更高级的应用程�,你需�考虑使用 lxml 。 它使用了和ElementTree�样的编程接�,因此上�的例��样也适用于lxml。 你�需�将刚开始的import语��� from lxml.etree import parse 就行了。 lxml 完全�循XML标准,并且速度也�常快,�时还支�验�,XSLT,和XPath等特性。