BounceGrip

RAG chunk budget planner

Plan a retrieval-augmented generation chunk budget around a selected model’s context window, system prompt, output reserve, and retrieved chunk count.

How to use this tool

  1. Choose the generation model that will receive retrieved context.
  2. Enter system-prompt and output-token reservations.
  3. Choose how many retrieved chunks to include and use the resulting per-chunk budget.

Worked example: Example: with a 200,000-token context window, a 5,000-token system prompt, a 5,000-token output reserve, and 10 chunks, each chunk can be about 19,000 tokens.

Retrieval budget397,000 tokens
Maximum per chunk49,625 tokens
Total document capacity397,000 tokens

Get your team on one AI workspace.