What is in the library
The datasets fall into six groups:
- Web and HTTP: every permanently registered HTTP status code with its RFC 9110 reason phrase, the 148 CSS named colors with hex and RGB values, and common file extensions mapped to their registered MIME types.
- Geography and language: ISO 3166-1 country codes (alpha-2, alpha-3, numeric), ISO 4217 currencies with their minor units, ISO 639 language codes, major IANA time zones and the 50 US states with abbreviations and capitals.
- Science and reference: the NATO spelling alphabet, the first twenty chemical elements and the eight planets.
- Mock data: fictional users, products, orders, analytics events and comments.
- Config files: package.json, tsconfig.json, a legacy .eslintrc.json, a Compose file, a Kubernetes Deployment and a GitHub Actions workflow.
- Parser edge cases: documents built to break naive JSON handling.
Each dataset names its source. Reference data comes from the standard that defines it, and facts that change often (moon counts, exchange rates, daylight-saving offsets) are deliberately left out rather than going stale.
Formats and what each one is good for
JSON is the master copy. Flat lists of records, such as status codes or currencies, can also be viewed as CSV for spreadsheets and bulk imports. YAML and XML views are produced by the same converters as the JSON to YAML and JSON to XML pages, so what you see is exactly what those tools produce. Nested datasets like the orders are not offered as CSV, because flattening them would hide structure; use JSON to CSV when you want to choose how.
The three deployment configs open as YAML, since that is how Compose, Kubernetes and GitHub Actions files normally live. Codes with leading zeros, such as ISO numeric codes “036” for Australia, are strings so that no tool turns them into 36.
Mock data you can rely on in tests
The mock users, products, orders, events and comments are generated by code from a fixed seed, so they are identical on every visit and in every build. Snapshot tests and screenshots stay stable, and two people following the same tutorial see the same rows.
Names come from a small built-in list of invented first and last names, every email address ends in example.com (a domain reserved for documentation by RFC 2606), and dates fall in 2025 and 2026. The datasets reference each other: order line items use real SKUs from the products, and user ids in orders and events fall within the mock users’ range, so joins and lookups work. Money is computed in whole cents and only then divided, so order totals equal the sum of their lines exactly.
Parser edge cases
These documents are valid JSON under RFC 8259 but behave differently from parser to parser:
- Big integers beyond 2^53, where JavaScript rounds 9007199254740993 down by one and 64-bit IDs lose digits. The big integers guide explains the fixes.
- Unusual numbers:
-0,1e400(which becomes Infinity as a double),1e-400(which becomes 0) and decimals with more digits than a double can hold. - Strings and Unicode: escaped and raw surrogate pairs, a lone surrogate, the null character, the U+2028 line separator, and combining accents that look identical to precomposed letters.
- Keys: empty keys,
__proto__, dots and slashes that need escaping in JSONPath and JSON Pointer, and two duplicate keys, one of which only appears after unescaping\u00e9. Duplicates are flagged in the JSON validator. - Deep nesting: 512 nested arrays, enough to find recursion limits.
- Empty values:
{},[],"",null,falseand0, which converters often drop or merge.
They are offered as JSON only, exactly as written, because converting them would already change them.
Using a dataset elsewhere
Copy and Download give you the text in the selected format, with a matching file name. “Open in formatter” and “Open converter” carry the text to that tool page in this browser tab, so you can reshape the data, generate TypeScript types from it, or turn a CSV into SQL inserts. The data comes from this site’s own files; nothing you do here is sent anywhere.
Examples
NATO phonetic alphabet as CSV
A two-column table that converts cleanly to CSV. Note the official spellings Alfa and Juliett.
nato-phonetic-alphabetCSVletter,codeWord
A,Alfa
B,Bravo
C,Charlie
D,Delta
E,Echo
F,Foxtrot
G,Golf
H,Hotel
I,India
J,Juliett
K,Kilo
L,Lima
M,Mike
N,November
O,Oscar
P,Papa
Q,Quebec
R,Romeo
S,Sierra
T,Tango
U,Uniform
V,Victor
W,Whiskey
X,X-ray
Y,Yankee
Z,Zulu
HTTP status codes as YAML
The full registry as a YAML list, ready to paste into an OpenAPI description or a fixture file.
http-status-codesYAMLBig integers that JavaScript rounds
Paste this into any JSON tool built on JSON.parse and compare the last digits of maxSafePlusTwo and int64Max.
edge-big-integersJSONKubernetes Deployment manifest
A complete apps/v1 Deployment shown as YAML; switch to JSON for the equivalent form, which kubectl apply accepts too.
config-kubernetes-deploymentYAMLCommon errors and how to fix them
| Error | Cause | Fix |
|---|---|---|
My CSV import turned 036 into 36 | Spreadsheet apps guess column types and treat numeric-looking codes as numbers, dropping leading zeros. | Import the column as text, or use the JSON file, where the codes are strings. |
Duplicate key "duplicate" — most parsers keep only the last valueExplained | The edge-keys dataset contains duplicate keys on purpose, to test how your parser reacts. | Nothing to fix; pick another dataset if you need clean data. |
Unexpected token or depth exceeded on the deep nesting sample | Some recursive parsers and printers limit nesting depth or run out of stack. | That is the point of the sample: make sure your code reports the problem instead of crashing. |
Frequently asked questions
Can I use these datasets in my own projects?
Yes. Reference data such as status codes and ISO codes are public standards, and the mock data is invented for this site. Check the original standard when you need its complete, authoritative list.
Why are some countries or currencies missing?
The country, currency and language lists are curated selections of widely used entries, not the full ISO tables, so every row could be checked. Each dataset’s source line says so.
Why is 418 listed as "(Unused)" instead of "I'm a teapot"?
The teapot code comes from RFC 2324, an April Fools’ joke. It was deployed often enough that RFC 9110 reserved 418 so it is never assigned, and the IANA registry lists it as unused. This dataset follows the registry.
Will the mock data change between visits?
No. It is generated from fixed seeds, so the same rows appear every time, which keeps snapshot tests and tutorials reproducible.
Why do the time zones show only a standard offset?
Daylight saving time changes the real offset for half the year in many zones, and governments change rules. Store the zone name, such as Europe/Berlin, and let a current time zone library work out the offset for each date.