Compare context windows
Choose a model with enough room for retrieval.
Plan a retrieval-augmented generation chunk budget around a selected model’s context window, system prompt, output reserve, and retrieved chunk count.
Worked example: Example: with a 200,000-token context window, a 5,000-token system prompt, a 5,000-token output reserve, and 10 chunks, each chunk can be about 19,000 tokens.