Glossary
AI terms for product managers
The 61 ideas the lessons teach, in plain words. Each one links to the lesson where you can try it on a real model and see what it changes.
A
Ablation
Removing one part of a prompt or system at a time to see how much it really contributes. It shows which ingredients matter and which are just habit.
Try it inWriting Good Prompts →
Abstaining
Letting the model say "I don't know" instead of guessing. Explicit permission cuts made-up answers, at the risk of being too cautious.
Try it inHallucination & Grounding →
Agent Loop
The cycle an agent repeats: the model decides on the next step, calls a tool, reads the result and goes again until it can answer. Each lap is another model call.
Try it inHow an AI Agent Works →
Agents
A model that decides its own next steps and tools until a goal is met. Flexible with unexpected requests, but costs more and can wander.
Try it inWorkflows vs. Agents →
Approvals
Requiring a person to confirm an agent's action before it happens, typically for money, deletions or anything that can't be undone.
Try it inHuman-in-the-Loop →
Autonomy
How much an agent may do without asking. More autonomy is faster and cheaper to run; less is safer.
Try it inHuman-in-the-Loop →
C
Cascades
Trying a cheap model first and escalating to a stronger one only when the answer looks unsure or fails a check.
Try it inModel Routing →
Context Window
The most tokens a model can take in at once: instructions, conversation history, documents and the reply together. What doesn't fit is cut, so the model can forget early turns.
Try it inHow an LLM Answers a Prompt →
Cost
What a run costs in model calls: the tokens in and out on every call, times the price per token. Agents, reasoning and big models multiply it.
E
Embeddings
Lists of numbers that represent the meaning of a piece of text, so texts with similar meanings get similar numbers. They power semantic search and RAG.
Try it inEmbeddings & Semantic Search →Retrieval-Augmented Generation →
F
Few-Shot Examples
Worked examples included in the prompt, showing an input and the output you want. They teach format and judgement faster than describing them.
Try it inWriting Good Prompts →
Fine-tuning
Further training a model on your own examples so it learns a style or pattern. Good for consistent behaviour; poor for facts, which go stale and are hard to update.
Try it inPrompting vs. RAG vs. Fine-tuning →
G
Grounding
Giving the model the facts it needs (documents, data, search results) and asking it to answer only from them, ideally citing where each answer came from.
Try it inHallucination & Grounding →
Guardrails
Checks around a model that block bad inputs or outputs, such as attack detection, content filters, format checks or limits on what actions are allowed.
Try it inHow an AI Agent Works →Prompt Injection & Guardrails →
H
Hallucination
A confident answer the model made up. It sounds the same as a correct one, so it has to be caught with grounding, checks or evals.
Try it inHallucination & Grounding →
Handoffs
Passing work and context from one agent to another. Each handoff adds time and cost, and can lose details on the way.
Try it inMulti-Agent Systems →
Human Review
People checking an AI's work, either before an action (approval) or afterwards on a sample, to catch errors and improve the system.
Try it inHuman-in-the-Loop →
J
JSON Schema
A description of the exact fields, types and allowed values a JSON reply must have. Providers can enforce it, so replies always parse.
Try it inStructured Outputs →
K
Keyword Search
Classic search that matches the words you typed. Precise for names and codes, but misses results that say the same thing in other words.
Try it inEmbeddings & Semantic Search →
L
Latency
How long a user waits for an answer. It grows with output length, model size, reasoning and the number of calls in a chain.
Try it inReasoning Models →
Least Privilege
Giving an agent only the tools and permissions the task needs, so a successful attack or mistake can do little harm.
Try it inPrompt Injection & Guardrails →
Live Traffic
What real users actually send, which always includes cases your test set didn't imagine.
Try it inProduction Monitoring →
LLM-as-Judge
Using a second model with a rubric to grade answers that have no single right wording, such as tone or helpfulness. Cheaper than people, but it must be checked against human grades.
Try it inEvaluating AI Changes →
M
Max Tokens
A cap on how long the reply can be. The model stops when it hits the cap, even mid-sentence, which keeps cost and latency predictable.
Try it inHow an LLM Answers a Prompt →
MCP
Model Context Protocol: an open standard for connecting AI assistants to tools and data. A server describes its tools once and any MCP client (Claude, ChatGPT, Cursor) can use them.
Try it inTool Design & MCP →
Memory
What an agent carries from step to step or session to session: the conversation so far, notes it saved, or facts stored for later. It all costs context and tokens.
Try it inHow an AI Agent Works →
Model Choice
Picking which model handles a task. Bigger models are better at hard cases and cost more; most traffic doesn't need them.
Try it inModel Routing →
Monitoring
Watching a live AI feature: sampling real conversations, grading them and tracking failures, cost and latency, then turning failures into new test cases.
Try it inProduction Monitoring →
O
Orchestration
Coordinating several model calls or agents: who does what, in which order, and how results are combined.
Try it inMulti-Agent Systems →
Output Format
Telling the model exactly what shape to reply in, such as one word, a list or JSON. It's often the single change that most improves results a program has to read.
Try it inWriting Good Prompts →
P
Parsing
Turning the model's reply into data your program can use. Free text often fails to parse; enforced schemas make it dependable.
Try it inStructured Outputs →
Pass Rate
The share of test cases that pass every grader. A single number to compare prompt or model versions.
Try it inEvaluating AI Changes →
Pricing
Model APIs charge per token, with separate prices for input and output tokens; output usually costs several times more. Bigger models cost more per token.
Try it inPrompt Caching & Cost →
Prompt Caching
The provider remembers the start of a prompt it has just seen and charges much less to reuse it. Put the fixed parts (instructions, documents) first so they can be cached.
Try it inPrompt Caching & Cost →
Prompt Design
Choosing what goes into a prompt: the task, the format you want back, definitions of your categories, examples and edge cases. Each part can be tested to see whether it helps.
Try it inWriting Good Prompts →
Prompt Injection
An attack that hides instructions in content the model reads, such as an email, web page or document, to make it do something its user didn't ask for.
Try it inPrompt Injection & Guardrails →
Prompting
Changing what you tell the model to change what it does. The cheapest, fastest fix, and the right one for rules you can write down.
Try it inPrompting vs. RAG vs. Fine-tuning →
R
RAG
Retrieval-augmented generation: search your own data for the passages relevant to a question, then give them to the model to answer from. It supplies facts the model was never trained on.
Try it inRetrieval-Augmented Generation →Prompting vs. RAG vs. Fine-tuning →
Reasoning
Models that think step by step before answering, spending hidden "reasoning tokens". It helps multi-step problems and is wasted on easy ones.
Try it inReasoning Models →
Regressions
Things that used to work and broke after a change. Evals exist mostly to catch these before users do.
Reliability
How often a system does the right thing across many runs, not just in a demo. Models vary from run to run, so it has to be measured.
Reranking
A second, more careful pass that reorders search results by how well they actually answer the question, before the model sees them.
Try it inRetrieval-Augmented Generation →
Risk Policy
Written rules for which actions an agent may take on its own and which need a person, kept in its instructions so fewer cases reach a reviewer.
Try it inHuman-in-the-Loop →
Routing
Sending each request to the model that fits it: easy ones to a cheap model, hard ones to a strong one. The risk is hard requests that look easy.
Try it inModel Routing →
S
Semantic Search
Searching by meaning instead of exact words, using embeddings. "Can't log in" finds a ticket about a locked account even though no words match.
Try it inEmbeddings & Semantic Search →
Similarity
How close two embeddings are, usually measured as cosine similarity. Higher means closer in meaning; search returns the most similar items.
Try it inEmbeddings & Semantic Search →
Specialists
Agents set up for one job each, with their own instructions, tools or data. Useful when each brings something the others lack.
Try it inMulti-Agent Systems →
Structured Output
Asking the model to reply in a fixed data shape (usually JSON matching a schema) so your code can read every reply reliably.
Try it inStructured Outputs →
T
Temperature
A setting for how adventurous the model's word choice is. At 0 it almost always picks the most likely next token, so answers repeat; higher values vary more, for ideas, and also make more mistakes.
Try it inHow an LLM Answers a Prompt →Hallucination & Grounding →
Test Sets
A fixed collection of inputs with known good answers, run after every change to see whether things got better or worse.
Thinking Effort
A setting for how much a reasoning model thinks before answering. More effort costs time and tokens; past a point it stops helping.
Try it inReasoning Models →
Token Cost
What a part of the request costs in tokens. Every connected tool's description is sent with every request, so unused tools still cost money.
Try it inTool Design & MCP →
Tokens
The pieces of text a model reads and writes: whole words, parts of words or punctuation. Models count length, limits and price in tokens, roughly 3 to 4 characters of English each.
Try it inHow an LLM Answers a Prompt →Prompt Caching & Cost →
Tool Descriptions
The name and text that tell a model what a tool does and when to use it. The model chooses tools only from these, so vague descriptions get the wrong tool called.
Try it inTool Design & MCP →
Tool Use
Letting a model call functions you provide, such as search, a calculator or your API. The model asks for a call; your code runs it and returns the result.
Top-K
How many search results retrieval passes to the model. Too few misses the answer; too many adds noise and cost.
Try it inRetrieval-Augmented Generation →
Trade-offs
Every AI design choice buys something and costs something: quality against cost, speed against care, flexibility against predictability. Product sense is knowing which side matters for this feature.
Try it inPrompting vs. RAG vs. Fine-tuning →Multi-Agent Systems →
U
Unit Economics
What one request, conversation or user costs you in model calls, compared with what it's worth. It decides whether a feature can scale.
Untrusted Data
Any text that didn't come from you or your user, such as tool results, web pages, emails or uploads. It can contain instructions, so treat it as data, never as commands.
Try it inPrompt Injection & Guardrails →
V
Vector Search
Finding the stored embeddings closest to a query's embedding, usually in a vector database. It's the retrieval step of most RAG systems.
Try it inRetrieval-Augmented Generation →
W
Workflows
A fixed sequence of steps, some using a model, decided in code. Cheap and predictable, but it mishandles anything it wasn't designed for.
Try it inWorkflows vs. Agents →