How an LLM Answers a Prompt

Key idea: An LLM predicts one token at a time inside a fixed context window.

  1. Send the tagline prompt at temperature 0 three times, then at 1. Compare how much the answers vary.
  2. Set Max output tokens to 40 and ask for the list of 10 fruits. Watch it get cut off.
  3. Have a few turns, switch the context window to Tiny, then ask what you asked first.
  4. Type in the box and watch the token chips update. Try emoji, code, or a non-English word.
Try a prompt
Enter to send · Shift+Enter for a new line

Change how it works, then send again

0.7
probabilitiesrepeatPromptyour textTokenizertext → token idsContext windowwhat the model seesChat historyearlier turnsSystem promptinstructionsModelscores next tokensministral-8bSamplerpicks one tokenOutputstreamed reply
Send a message to watch it run

Tokens · last prompt

Start typing to see how text is split into tokens.

Approximate (OpenAI cl100k tokenizer). Each model family splits text a little differently.

Behind the scenes

Send a message and every step the system takes will show up here.