Common causes
1. Extra text before the document
A status line, a log prefix or response headers captured with the body put text in front of the XML. Pass the parser only the body.
HTTP/1.1 200 OK
<invoice number="2026-091"/><invoice number="2026-091"/>2. Whitespace or a blank line before <?xml
The XML declaration must be the very first thing in the file. Even a newline before it makes Java report that a processing instruction named “xml” is not allowed. Trim leading whitespace when generating the file.
<?xml version="1.0" encoding="UTF-8"?>
<invoice number="2026-091"/><?xml version="1.0" encoding="UTF-8"?>
<invoice number="2026-091"/>3. A byte-order mark read with the wrong encoding
A UTF-8 BOM decoded as Windows-1252 becomes the visible text “” before the declaration. Let the parser read the raw bytes (an InputStream rather than a Reader) so it can detect the encoding itself.
<?xml version="1.0" encoding="UTF-8"?>
<invoice number="2026-091"/><?xml version="1.0" encoding="UTF-8"?>
<invoice number="2026-091"/>4. The parser was given a file name or JSON instead of XML
Passing a path to a method that expects XML text (for example new StringReader("invoice.xml")) makes the parser read the path itself. A JSON error body from an API produces the same message.
{"error": "invoice not found"}<error>invoice not found</error>Frequently asked questions
What exactly is the XML prolog?
Everything before the root element: the optional XML declaration, which must come first, followed by any comments, processing instructions and a DOCTYPE. Only markup and whitespace may appear there.
Why does the file look fine in my editor?
A BOM and some whitespace characters are invisible. Open the file in a hex viewer, or paste it into PasteKit, which reports the exact character and position before the root element.
How do I avoid this in Java?
Pass DocumentBuilder.parse a File or an InputStream rather than a String or Reader, so the parser handles the BOM and encoding declaration itself, and log the first bytes of any input that fails.