Python ships a solid JSON library in the standard library. Most of its surprises come from a handful of defaults: ASCII escaping, NaN output, missing datetime support and float parsing. This guide covers the functions, those defaults, and the fixes.
The four functions
The names follow one rule: a trailing s means string.
import json
data = json.loads('{"id": 7, "tags": ["a", "b"]}') # str/bytes -> Python
text = json.dumps(data) # Python -> str
with open("order.json", encoding="utf-8") as f:
order = json.load(f) # file -> Python
with open("out.json", "w", encoding="utf-8") as f:
json.dump(order, f, indent=2, ensure_ascii=False) # Python -> file
Always pass encoding="utf-8" when opening files. Without it, Python uses the platform’s locale encoding, which on older Windows setups is a legacy code page, and non-ASCII text is garbled or rejected. json.loads also accepts bytes directly and detects UTF-8, UTF-16 and UTF-32.
Types map as you would expect: objects become dict, arrays list, strings str, integers int, other numbers float, true/false become True/False and null becomes None. Serialising also accepts tuples, which become arrays; sets do not, and raise TypeError.
Output options: indent, separators, ensure_ascii, sort_keys
json.dumps({"name": "café"})
# '{"name": "caf\u00e9"}'
json.dumps({"name": "café"}, ensure_ascii=False)
# '{"name": "café"}'
json.dumps({"a": [1, 2]}, separators=(",", ":"))
# '{"a":[1,2]}'
json.dumps({"b": 1, "a": 2}, indent=2, sort_keys=True)
- ensure_ascii defaults to
True, so every non-ASCII character is written as a\uXXXXescape, and characters outside the Basic Multilingual Plane, such as emoji, become surrogate pairs like\ud83d\ude00. The output is valid JSON and parses back identically, but it is larger and unreadable for non-English text. Set it toFalseand write the file as UTF-8. - indent pretty-prints. When it is set, the item separator becomes a plain comma, so lines do not end with trailing spaces.
- separators controls the punctuation;
(",", ":")gives the most compact output. - sort_keys produces stable output for diffs, caches and signatures.
For a quick look at a file from the shell, python -m json.tool data.json validates and pretty-prints it.
NaN and Infinity: Python writes invalid JSON by default
JSON has no representation for NaN or infinities, but Python’s allow_nan defaults to True:
json.dumps(float("nan")) # 'NaN' (not valid JSON)
json.dumps(float("nan"), allow_nan=False) # ValueError: Out of range float values are not JSON compliant: nan
json.loads("[NaN, Infinity]") # [nan, inf]
JavaScript’s JSON.parse and most other parsers reject NaN, so a pandas result with missing values can produce a file that only Python can read. Pass allow_nan=False in services so the problem surfaces where it starts, and convert missing values to None (which becomes null) before dumping. The NaN and Infinity explainer shows what other languages do.
datetime, Decimal and other types: default=
Serialising an unsupported type raises TypeError: Object of type datetime is not JSON serializable. The default parameter is a function that is called for any object the encoder does not understand; it returns something serialisable or raises TypeError itself:
from datetime import datetime, timezone
from decimal import Decimal
def to_json(o):
if isinstance(o, datetime):
return o.isoformat()
if isinstance(o, Decimal):
return str(o)
if isinstance(o, set):
return sorted(o)
raise TypeError(f"Object of type {type(o).__name__} is not JSON serializable")
json.dumps({"at": datetime(2026, 10, 6, 9, 0, tzinfo=timezone.utc), "price": Decimal("19.90")}, default=to_json)
# '{"at": "2026-10-06T09:00:00+00:00", "price": "19.90"}'
default=str is a popular shortcut, but it hides bugs: any unexpected object is silently turned into its str() form, and a datetime becomes "2026-10-06 09:00:00+00:00" with a space rather than T. Raising for unknown types is safer. For classes you reuse, subclass json.JSONEncoder, override its default method and pass cls=YourEncoder.
Note the choice for Decimal: returning float(o) would round values like 0.1000000000000000000001 to 0.1, so a string keeps money and measurements exact. The standard library cannot write an unquoted number with more precision than a float; the third-party simplejson package can, with use_decimal=True. For dates and times, the JSON date format guide explains which string format to choose.
Parsing numbers: parse_float, parse_int and big integers
By default JSON numbers with a fraction or exponent become float, with the usual binary rounding: 19.90 is read as 19.9, and summing such values gives results like 0.30000000000000004. Hand the text to Decimal instead:
json.loads('{"price": 19.90, "qty": 2}', parse_float=Decimal)
# {'price': Decimal('19.90'), 'qty': 2}
Integers are a strength: Python’s int has arbitrary precision, so 9007199254740993, which JavaScript rounds, is read exactly. One limit exists since Python 3.11: converting a string of more than 4300 digits to an int raises ValueError: Exceeds the limit (4300 digits) for integer string conversion, a guard against denial-of-service. sys.set_int_max_str_digits() raises the limit if you really need gigantic numbers.
parse_int and parse_constant exist too; the latter lets you reject NaN and Infinity while parsing.
Large files and JSON Lines
json.load reads the whole document into memory before returning, and the resulting Python objects take several times more memory than the file itself. For logs and exports, prefer JSON Lines, one document per line, which you can process as a stream:
with open("events.jsonl", encoding="utf-8") as f:
for line in f:
if line.strip():
event = json.loads(line)
For one huge array that cannot be split, a streaming parser such as the third-party ijson package yields items one at a time. When speed matters more than the standard library’s flexibility, orjson is a popular drop-in alternative; note that its dumps returns bytes rather than str, and that it serialises datetime and dataclasses itself. The JSON Lines guide compares the two layouts.
Errors, duplicate keys and other details
Invalid input raises json.JSONDecodeError, a subclass of ValueError with msg, lineno, colno and pos attributes:
try:
json.loads('{"a": 1,}')
except json.JSONDecodeError as e:
print(e) # Expecting property name enclosed in double quotes: line 1 column 9 (char 8)
Expecting value: line 1 column 1 (char 0) almost always means the input was empty or was not JSON at all, typically an HTML error page or an empty HTTP body; the explainer for that message lists the usual causes. Single-quoted keys, the result of printing a dict with str() instead of json.dumps, produce the “enclosed in double quotes” error.
Duplicate keys are accepted silently and the last one wins. To reject them, use object_pairs_hook, which receives the raw list of pairs:
def unique(pairs):
keys = [k for k, _ in pairs]
if len(keys) != len(set(keys)):
raise ValueError("duplicate keys")
return dict(pairs)
json.loads('{"a": 1, "a": 2}', object_pairs_hook=unique) # ValueError
Dictionary keys must be strings when serialising, but int, float, bool and None keys are converted to strings for you, so {1: "x"} comes back from a round trip as {"1": "x"}. Tuple keys raise TypeError. Extremely deep nesting, in the thousands of levels, exceeds the interpreter’s recursion limit and raises RecursionError rather than JSONDecodeError, so catch both when parsing untrusted data. To check a document by eye, paste it into the JSON formatter, which reports the line and column of the first error.
Frequently asked questions
What is the difference between json.load and json.loads?
json.loads parses a string or bytes object; json.load reads from a file-like object. Likewise json.dumps returns a string and json.dump writes to a file.
How do I stop Python escaping accented characters and emoji?
Pass ensure_ascii=False to json.dumps or json.dump, and write the result with UTF-8 encoding.
How do I serialise a datetime to JSON in Python?
Convert it to an ISO 8601 string with isoformat(), either before dumping or in a default= function. Make it timezone-aware first, otherwise the string has no offset.
Why does json.loads turn 19.90 into 19.9?
Numbers with a fraction become binary floats, which drop the trailing zero and cannot store most decimals exactly. Pass parse_float=decimal.Decimal to keep the exact text.
Does Python lose precision on large integers in JSON?
No. Integers are parsed with arbitrary precision. Only strings longer than 4300 digits are refused by default since Python 3.11.
Is the output of json.dumps always valid JSON?
Not by default: NaN and Infinity are written as bare words that other parsers reject. Use allow_nan=False to raise an error instead.