Key idea: You pay for every token in and out. Caching makes a repeated prompt start much cheaper, if you build prompts for it.
- Ask two questions in a row. The second request reads most of its prompt from the cache and costs less.
- Turn off Cache-friendly prompt and ask again. A timestamp at the top breaks the cache every time.
- Turn off the document. The prompt drops below ~1,024 tokens and can't be cached at all.
- Switch the answer length to Detailed. Output tokens cost about 4x input, so watch the Answer slice grow.
- Compare Small and Large models on the same question, and check the cost per 1,000 requests.
Try a prompt