如何解决将XML文件转换为CSV文件的有效方法?
我正在尝试找到一种使用Python将xml文件转换为csv文件的方法。我想这样做,以便脚本将解析每个警报的xml文件(请参阅下面的xml片段)。
因此,它将创建一个xls文件,其中包含eventType
,probableCause
,description
和severities
的列,类似于此格式:
我拥有的代码不起作用,它仅更新列名。
XML示例:
<?xml version="1.0" encoding="UTF-8"?>
<faults version="1" xmlns="urn:nortel:namespaces:mcp:faults" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="urn:nortel:namespaces:mcp:faults NortelFaultSchema.xsd ">
<family longName="1OffMsgr" shortName="OOM"/>
<family longName="ACTAGENT" shortName="ACAT">
<logs>
<log>
<eventType>RES</eventType>
<number>1</number>
<severity>INFO</severity>
<descTemplate>
<msg>Accounting is enabled upon this NE.</msg>
</descTemplate>
<note>This log is generated when setting a Session Manager's AM from <none> to a valid AM.</note>
<om>On all instances of this Session Manager,the <NE_Inst>:<AM>:STD:acct OM row in the StdRecordStream group will appear and start counting the recording units sent to the configured AM.
On the configured AM,the <NE_inst>:acct OM rows in RECSTRMCOLL group will appear and start counting the recording units received from this Session Manager's instances.
</om>
</log>
<log>
<eventType>RES</eventType>
<number>2</number>
<severity>ALERT</severity>
<descTemplate>
<msg>Accounting is disabled upon this NE.</msg>
</descTemplate>
<note>This log is generated when setting a Session Manager's AM from a valid AM to <none>.</note>
<action>If you do not intend for the Session Manager to produce accounting records,then no action is required. If you do intend for the Session Manager to produce accounting records,then you should set the Session Manager's AM to a valid AM.</action>
<om>On all instances of this Session Manager,the <NE_Inst>:<AM>:STD:acct OM row in the StdRecordStream group that matched the previous datafilled AM will disappear.
On the previously configured AM,the <NE_inst>:acct OM rows in RECSTRMCOLL group will disappear.
</om>
</log>
</logs>
</family>
<family longName="ACODE" shortName="AC">
<alarms>
<alarm>
<eventType>ADMIN</eventType>
<number>1</number>
<probableCause>INFORMATION_MODIFICATION_DETECTED</probableCause>
<descTemplate>
<msg>Configured data for audiocode server updated: $1</msg>
<param>
<num>1</num>
<description>AudioCode configuration data got updated</description>
<exampleValue>acgwy1</exampleValue>
</param>
</descTemplate>
<manualClearable></manualClearable>
<correctiveAction>None. Acknowledge/Clear alarm and deploy the audiocode server if appropriate.</correctiveAction>
<alarmName>Audiocode Server Updated</alarmName>
<severities>
<severity>MINOR</severity>
</severities>
</alarm>
<alarm>
<eventType>ADMIN</eventType>
<number>2</number>
<probableCause>CONFIG_OR_CUSTOMIZATION_ERROR</probableCause>
<descTemplate>
<msg>Deployment for audiocode server failed: $1. Reason: $2.</msg>
<param>
<num>1</num>
<description>AudioCode Name</description>
<exampleValue>audcod</exampleValue>
</param>
<param>
<num>2</num>
<description>AudioCode Deployment failed reason</description>
<exampleValue>Failed to parse audiocode configuration data</exampleValue>
</param>
</descTemplate>
<manualClearable></manualClearable>
<correctiveAction>Check the configuration of audiocode server. Acknowledge/Clear alarm and deploy the audiocode server if appropriate.</correctiveAction>
<alarmName>Audiocode Server Deploy Failed</alarmName>
<severities>
<severity>MINOR</severity>
</severities>
</alarm>
</alarms>
</family>
</faults>
我尝试过的方法(小样本):
from logging import root
from xml.etree import ElementTree
import os
import csv
tree = ElementTree.parse('Fault.xml')
sitescope_data = open('Out.csv','w',newline='',encoding='utf-8')
csvwriter = csv.writer(sitescope_data)
col_names = ['eventType','probableCause','description']
csvwriter.writerow(col_names)
root = tree.getroot()
for eventData in root.findall('alarms'):
event_data = []
event = eventData.find('alarm')
event_id = event.find('eventType')
if event_id != None :
event_id = event_id.text
event_data.append(event_id)
csvwriter.writerow(event_data)
sitescope_data.close()
解决方法
root = tree.getroot()
def get_uri(elem):
if elem.tag[0] == "{":
uri,ignore,tag = elem.tag[1:].partition("}")
return f"{{{uri}}}"
return ""
uri = get_uri(root)
def recurse(root):
for child in root:
recurse(child)
print(child.tag)
for event in root.findall(f'{uri}alarm'):
event_data = []
event_id = event.find(f'{uri}eventType')
if event_id != None :
event_id = event_id.text
event_data.append(event_id)
probableCause = event.find(f'{uri}probableCause')
if probableCause != None:
probableCause = probableCause.text
event_data.append(probableCause)
severities = event.find(f'{uir}severities')
if severities:
severity_data = ','.join([sv.text for sv in severities.findall('f{uri}severity')])
event_data.append(severity_data)
else:
event_data.append("")
csvwriter.writerow(event_data)
recurse(root)
要注意的事情:
- 使用递归遍历XML
- 打印语句将向您显示,每个标签都有{urn:nortel:namespaces:mcp:faults}来自根目录中的xmlns属性,这可能是导致您绊倒的大部分原因。 我添加了一个函数来获取“ uri”文本并将其添加到每个标签之前。
- 您每次写入csv时都希望添加多个列
版权声明:本文内容由互联网用户自发贡献,该文观点与技术仅代表作者本人。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如发现本站有涉嫌侵权/违法违规的内容, 请发送邮件至 dio@foxmail.com 举报,一经查实,本站将立刻删除。