Developer tools

RAG Chunking Calculator

Estimate RAG chunk counts, token totals, and overlap overhead from document length, chunk size, and overlap locally in your browser.

Enter the document length in characters, the average characters per token for your tokenizer, and the chunk size and overlap in tokens. Overlap must be smaller than the chunk size. The estimate runs locally in your browser; your values are not uploaded or stored.

Your RAG chunking estimate will appear here.

Split a document into retrieval chunks before you embed it.

Enter the document length in characters, the average characters per token for your tokenizer, and the chunk size and overlap in tokens. The calculator estimates total tokens, the number of chunks, and the extra tokens repeated by overlap. Storage and recall depend on your tokenizer, splitter, and vector store, so confirm the plan against your pipeline. Everything runs locally in your browser, so your values are not uploaded or stored.

Frequently Asked Questions

Everything you need to know about this tool, how it works, and privacy.

What does the RAG chunking calculator estimate?

Enter the document length in characters, the average characters per token for your tokenizer, and the chunk size and overlap in tokens. The calculator estimates total tokens as characters divided by characters per token, then counts fixed-size chunks with a sliding window: one chunk when the document fits, otherwise 1 plus the remaining tokens divided by the step of chunk size minus overlap, rounded up. For example, 50000 characters at 4 chars per token with 512-token chunks and 50-token overlap give about 12500 tokens in 27 chunks with 1300 repeated overlap tokens.

What should I enter for characters per token?

Characters per token converts your character count into tokens for the chunking math. English text with common subword tokenizers averages about 4 characters per token; code, dense symbols, or other languages can differ. Enter a value from 1 to 10 that matches your tokenizer, or measure it on a sample of your own documents.

How does chunk overlap change the result?

Overlap repeats tokens between neighbouring chunks to preserve context across boundaries. It must be a whole number from 0 up to one less than the chunk size. Larger overlap adds repeated tokens on top of the document total, which the result reports as overlap overhead with its percentage.

Do my values stay private?

Yes. The calculation and copy action run locally in your browser. CodeASystem does not upload or store the values you enter.