A recent analysis compared the token usage of seven data formats when used in LLM prompts, finding that pretty-printed JSON consumes approximately three times more tokens than CSV. This difference is attributed to repeated keys, punctuation, and unnecessary indentation, which add to the token count without benefiting the model. The study suggests using CSV or TSV for simple tabular data, minified JSON for structured output, and Markdown for human-readable tables, while advising against pretty-printed JSON and XML in prompts to optimize costs. AI
IMPACT Optimizing prompt formatting can significantly reduce LLM operational costs and improve efficiency.
RANK_REASON Analysis of data format token efficiency for LLM prompts.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →