Why convert XML to JSON
Plenty of systems still speak XML — SOAP services, bank and payment files, sitemaps, RSS, Office documents, Android resources — while the code consuming them prefers JSON. Converting lets you inspect the structure, write a JSONPath or jq query against it, or paste it into a JavaScript test. Because XML has features JSON lacks (attributes, mixed content, namespaces), every converter needs a convention. This one is documented below so you can rely on it.
The mapping, rule by rule
- The root element becomes the single top-level key:
<order>…</order>becomes{"order": …}. - Attributes become keys prefixed with
@:id="ord_8f2k1"becomes"@id": "ord_8f2k1". - Child elements become keys named after the element.
- Repeated child elements with the same name become an array, in document order. Elements with the same name are grouped even when other elements sit between them.
- An element with only text and no attributes becomes a plain string:
<city>Paris</city>becomes"city": "Paris". - When an element has attributes or children as well as text, the text goes under
"#text"; CDATA sections go under"#cdata". - Empty elements (
<note/>or<note></note>) become an empty string. - Namespace prefixes are kept in names (
"dc:title","@xmlns:dc"); namespaces are not resolved. - The five built-in entities and character references such as
 are decoded. Entities declared in a DTD, like , are left as written and reported with a warning.
The top-level JSON is always an object, and the output uses two-space indentation.
Detect numbers and booleans
By default every text value and attribute stays a string, because XML itself has no types and a ZIP code or account number must not lose its leading zeros.
Turn on Detect numbers and booleans to convert values that are exactly true, false or a canonical number. The rule is strict: text becomes a number only if writing it back produces the same text. So 42 and -3.5 become numbers, while 007, 1.50, 1e3 and integers too large for a double stay strings. The option applies to attribute values as well as element text.
What XML loses on the way
- Comments, processing instructions and the DOCTYPE are dropped, with an info message counting them.
- Mixed content such as
<p>Hello <b>big</b> world</p>keeps its text, but the text pieces are trimmed and joined ("Hello world") and their position relative to<b>is lost. JSON is a poor fit for prose markup. - Single versus repeated: one
<item>gives an object, two give an array. Code reading the JSON should accept both shapes. - Order between different elements that are grouped into arrays is not recorded.
Converting back with JSON to XML uses the same convention, so data-oriented XML round-trips well; document-oriented XML does not.
Tips
Run messy XML through the XML formatter first to spot structural problems, and use XML to YAML when you want the same mapping in a more readable form. Parsing is strict XML 1.0, and everything happens on your device.
Examples
Order with attributes and CDATA
The two item elements become an array of objects with @sku, @qty and #text, and the CDATA note is kept under #cdata.
<?xml version="1.0" encoding="UTF-8"?>
<order id="ord_8f2k1" status="paid">
<customer vip="true">Aisha Tan</customer>
<item sku="KB-104" qty="1">Mechanical keyboard</item>
<item sku="MS-220" qty="2">Wireless mouse</item>
<note><![CDATA[Leave at <front> door]]></note>
</order>
{
"order": {
"@id": "ord_8f2k1",
"@status": "paid",
"customer": {
"@vip": "true",
"#text": "Aisha Tan"
},
"item": [
{
"@sku": "KB-104",
"@qty": "1",
"#text": "Mechanical keyboard"
},
{
"@sku": "MS-220",
"@qty": "2",
"#text": "Wireless mouse"
}
],
"note": {
"#cdata": "Leave at <front> door"
}
}
}
Typed values in a feed
With detection on, stock becomes 42 and active becomes true, while the zero-padded SKU and 129.90 stay strings because rewriting them would change the text.
<products updated="2026-09-14">
<product>
<sku>00123</sku>
<price>129.90</price>
<stock>42</stock>
<active>true</active>
</product>
</products>
{
"products": {
"@updated": "2026-09-14",
"product": {
"sku": "00123",
"price": "129.90",
"stock": 42,
"active": true
}
}
}
Namespaced RSS item
The dc: prefix stays part of the key and the namespace declaration becomes an @xmlns:dc attribute key.
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
<title>Release notes</title>
<item>
<title>PasteKit 2.0</title>
<dc:creator>Ben Okafor</dc:creator>
<pubDate>Mon, 14 Sep 2026 08:21:05 GMT</pubDate>
</item>
</channel>
</rss>
{
"rss": {
"@version": "2.0",
"@xmlns:dc": "http://purl.org/dc/elements/1.1/",
"channel": {
"title": "Release notes",
"item": {
"title": "PasteKit 2.0",
"dc:creator": "Ben Okafor",
"pubDate": "Mon, 14 Sep 2026 08:21:05 GMT"
}
}
}
}
Common errors and how to fix them
| Error | Cause | Fix |
|---|---|---|
Expected </b> to close <b> opened at line 1, found </c>Explained | A closing tag does not match the most recently opened element. | Rename the closing tag or close the inner element first. |
Unescaped '&' — it must start an entity reference like & or &Explained | A bare ampersand appears in text or an attribute, common in URLs with query strings. | Write & instead of & or wrap the text in a CDATA section. |
Only one root element is allowed — <b> follows the root element <a> that closed at line 1Explained | The input contains several top-level elements, for example two XML fragments pasted together. | Wrap them in a single parent element. |
Text is not allowed before the root elementExplained | Something other than whitespace, a declaration or a comment precedes the first tag, often a byte of log output. | Delete everything before the <?xml declaration or the root element. |
Entity is not declared (only & < > " ' are built in)Explained | A warning: an HTML entity is used in XML without a DTD declaring it. | Use the character reference or the character itself. |
Frequently asked questions
How are XML attributes represented in JSON?
As keys starting with @ inside the element’s object, for example “@id”. Element text then moves to “#text” so both can coexist.
Why is a single child an object but two children an array?
XML does not say whether an element is meant to repeat, so the converter only creates an array when it actually sees repetition. Code reading the JSON should handle both shapes.
Are numbers converted automatically?
Only with Detect numbers and booleans switched on, and then only when the text reads back identically. Values like 007 or 1.50 always stay strings.
Are comments kept?
No. Comments, processing instructions and the DOCTYPE have no JSON equivalent and are dropped, with a message saying how many.
Is the XML uploaded to a server?
No, the parser runs in your browser, so internal SOAP payloads and bank files stay on your machine.