Why the same data costs different amounts
Language models read text as tokens, chunks of a few characters each, and both pricing and context limits are counted in them. JSON spends many tokens on punctuation: every key is wrapped in quotes and followed by a colon, every value is separated by a comma, and pretty-printing adds a run of spaces before each line. None of that carries meaning for the model.
The table shows the same data written five ways, each counted with the real tokenizer:
- As pasted: your text exactly, including its indentation.
- Minified JSON: all optional whitespace removed. Usually the safest saving, since the data and the syntax stay identical.
- Pretty JSON (2 spaces): one value per line, the way most tools print it.
- YAML: no braces and few quotes, indentation instead. Often cheaper than pretty JSON and close to minified.
- CSV: offered when the data is an array of flat records. Field names are written once in a header row instead of once per record, which is frequently the biggest saving of all.
The row with the fewest tokens is marked, and the headline says how much smaller it is than what you pasted.
Choosing an encoding
OpenAI models share a handful of published byte-pair encodings, and the counts here are exact for them:
- o200k_base, used by GPT-4o, GPT-4.1, GPT-5 and the o-series reasoning models (o1, o3, o4-mini). Its 200,000-entry vocabulary handles non-English text and code more compactly.
- cl100k_base, used by GPT-4, GPT-4 Turbo, GPT-3.5 Turbo and the text-embedding-3 models.
Text that looks like a special token, such as <|endoftext|>, is counted as ordinary characters, the way the APIs treat user content. Chat requests add a few tokens per message for roles and separators, so a prompt costs slightly more than the bare text shown here.
Claude, Gemini and other models
Anthropic and Google do not publish their tokenizers, and open models such as Llama and Mistral each use their own vocabulary. The same JSON produces a different count on each. Use the numbers on this page as an estimate for those models, and rely on the provider’s own token-counting endpoint or usage report when you need the exact bill. The ranking between representations usually carries over to other tokenizers, because the savings come from dropping the same punctuation and whitespace.
Things that make JSON expensive
- Long key names repeated in every record. In an array of a thousand objects each name is paid for a thousand times. The JSON Size Analyzer shows which keys cost the most bytes.
- Indentation. Deeply nested, pretty-printed documents can spend a third of their tokens on spaces.
- Escaped Unicode. A character written as
\u00e9takes several tokens; the same character written directly usually takes one or two. - Embedded blobs. Base64 images, hashes and random IDs tokenize poorly: random Base64 averages fewer than two characters per token, against about four for English prose.
- Data the model does not need. Dropping unused fields with a jq or JSONPath query beats any formatting trick.
Accuracy and privacy
Counting uses the MIT-licensed gpt-tokenizer library, a JavaScript port of OpenAI’s tiktoken that our tests check against published reference counts. Each encoding’s vocabulary (about 1.2 MB for cl100k_base and 2.4 MB for o200k_base, before compression) downloads from this site the first time you use it, and counting a megabyte of JSON in all its forms typically takes under a second, in a background worker. The YAML and CSV forms come from the same converters as the JSON to YAML and JSON to CSV tools, and big numbers keep their exact digits in all of them. No text leaves your browser.
Examples
A product list: CSV wins
Five flat records. The header row names each field once, so CSV needs far fewer tokens than any JSON layout; YAML and minified JSON land in between.
[
{ "sku": "KB-104", "name": "Mechanical keyboard", "category": "peripherals", "price": 129.9, "in_stock": true },
{ "sku": "MS-220", "name": "Wireless mouse", "category": "peripherals", "price": 39.5, "in_stock": true },
{ "sku": "MN-270", "name": "27-inch monitor", "category": "displays", "price": 289.0, "in_stock": false },
{ "sku": "HD-310", "name": "USB-C hub", "category": "accessories", "price": 49.99, "in_stock": true },
{ "sku": "WC-090", "name": "HD webcam", "category": "peripherals", "price": 64.0, "in_stock": true }
]As pasted: 194 tokens
Minified JSON: 139 tokens
Pretty JSON (2 spaces): 229 tokens
YAML: 159 tokens
CSV: 86 tokens
Fewest tokens: CSVA nested order: minify it
Nested objects cannot become CSV, so the choice is between JSON layouts and YAML. Minifying removes the indentation tokens without changing a single value.
{
"order": {
"id": "ord_8f2k1",
"placed_at": "2026-09-14T08:21:05Z",
"customer": { "name": "Aisha Tan", "tier": "gold" },
"lines": [
{ "sku": "KB-104", "qty": 1, "unit_price": 129.9 },
{ "sku": "MS-220", "qty": 2, "unit_price": 39.5 }
],
"shipping": { "method": "express", "address": { "city": "Singapore", "postcode": "018956" } }
}
}As pasted: 149 tokens
Minified JSON: 99 tokens
Pretty JSON (2 spaces): 168 tokens
YAML: 119 tokens
Fewest tokens: Minified JSONNon-English text and emoji
Japanese, Korean and emoji take more tokens per character than English. Compare the two encodings: o200k_base is noticeably cheaper for non-Latin scripts.
{
"greetings": {
"en": "Thank you for your order!",
"de": "Vielen Dank für Ihre Bestellung!",
"ja": "ご注文ありがとうございます!",
"ko": "주문해 주셔서 감사합니다!",
"emoji": "🎉🙏📦"
}
}As pasted: 66 tokens
Minified JSON: 48 tokens
Pretty JSON (2 spaces): 66 tokens
YAML: 47 tokens
Fewest tokens: YAMLCommon errors and how to fix them
| Error | Cause | Fix |
|---|---|---|
Not valid JSON, so only the text as pasted was counted | The text has a syntax error, so it cannot be minified or converted. The count for the raw text is still exact. | Use the Go to the error link, or repair the document with the JSON formatter, to unlock the other representations. |
No CSV row in the table | CSV is only offered for an array of objects whose values are plain values or one level of nested objects. Arrays inside records do not fit in a cell cleanly. | Flatten the data first with the JSON flatten tool, or pick the fields you need with a jq query. |
The count differs from my API bill | Chat requests add tokens for message roles and formatting, tool definitions and system prompts count too, and non-OpenAI models use other tokenizers. | Compare like with like: the bare text here is the content part of one message. |
Frequently asked questions
Which models are these counts exact for?
o200k_base counts are exact for GPT-4o, GPT-4.1, GPT-5 and the o1, o3 and o4-mini models; cl100k_base counts are exact for GPT-4, GPT-4 Turbo, GPT-3.5 Turbo and the text-embedding-3 models.
Can I count tokens for Claude or Gemini here?
Only approximately. Anthropic and Google have not published their tokenizers, so no browser tool can count their tokens exactly. Their APIs offer token-counting endpoints for exact numbers.
Is minified JSON always cheaper?
It is never more expensive than the same JSON with indentation, because it is the same text minus whitespace. Whether YAML or CSV beats it depends on the data, which is what the table shows.
Does the model understand CSV and YAML as well as JSON?
Current models read all three well. CSV loses types, so numbers and strings that look alike become indistinguishable, and nested data does not fit. Keep JSON when the exact structure matters, for example for tool-call arguments.
Is my data sent to OpenAI or anywhere else?
No. Tokenization runs in your browser with a bundled copy of the encoding, so the page works offline after loading and nothing you paste is transmitted.
Can I count plain text that is not JSON?
Yes. Any text gets an exact count for the selected encoding; the extra representations appear only when the text is valid JSON.