Skip to content
Nerdy beta

AI Engineering Glossary

Half the arguments about AI engineering are two people using the same word for different things. Here's the working vocabulary, grouped by what the terms do together rather than alphabetized, because almost none of them mean much alone.

Core Building Blocks

Model (Foundation Model, LLM) The trained neural network itself: a function that takes a sequence of tokens and predicts the ones that come next. It has no memory, no tools, and no ability to act. Every other term in this glossary names something built around it to make it useful.

Token The unit of text a model reads and writes, usually a word fragment of three or four characters of English. Tokens are the currency of AI engineering. Context limits, pricing, and latency are all denominated in them.

Inference Running a model to produce output, as opposed to training it. Inference cost and latency are the constraints your architecture trades against every day.

Context Window The maximum number of tokens a model can attend to in one request, covering the system prompt, conversation history, tool results, and retrieved documents together. Treat it as working memory. Nothing survives past the end of a request unless the surrounding software deliberately carries it forward.

System Prompt Instructions placed at the start of the context that set the model’s role, its constraints, the tools it can call, and the rules it operates under. In an agent it does the work of a job description and an operating manual at once.

Agents and Their Harness

Agent A model running in a loop. It decides what to do next, takes an action (usually a tool call), reads the result, and continues until the task is done or it gives up. The defining property is that the model directs its own control flow. Contrast a workflow, where code fixes the sequence and the model fills in individual steps. Anthropic’s Building effective agents is the reference statement of that split.

Harness The software that makes the agent loop run: parsing and executing tool calls, feeding results back, managing the context window, enforcing permissions, handling retries, and defining the system prompt and tool set. The model emits tokens. The harness does everything else. Claude Code, OpenHands, and any bespoke loop built on a messages API are all harnesses. An agent is what you get when a capable model runs inside a good harness, and neither one alone produces it.

Workflow A sequence of steps that code orchestrates, with LLM calls filling specific stages. Classify this, summarize that, draft the other. Workflows are cheaper and more predictable than agents, and agents handle open-ended tasks a workflow can’t anticipate. Start as a workflow and make the model earn its way to agency.

Orchestration The layer that sequences model calls, tools, and data flows, whether in a workflow or an agent. LangGraph, the Claude Agent SDK, and systems built on Temporal are opinionated harness kits sold as frameworks.

Tool Use (Function Calling) How a model asks for something to happen in the outside world. It emits a structured call, a name plus arguments, the harness executes it, and the result goes back into context. Tool design does more for agent quality than almost anything else and gets the least attention: naming, descriptions, granularity, and the wording of error messages.

MCP (Model Context Protocol) An open standard for connecting models to external tools and data. An MCP server exposes tools in a common format, and any compliant harness can consume them. MCP does for tool integration what USB did for peripherals: one connector, many devices.

Subagent An agent that another agent spawns to handle a scoped subtask, usually with a fresh context window of its own. This is a context-management technique first. The parent delegates a deep dive, the subagent burns its own context doing the reading, and only the distilled answer comes back.

Multi-Agent System Several agents with distinct roles working one task, coordinated by an orchestrator or through shared artifacts. It pays off on genuinely parallel work. Applied to a task that one well-harnessed agent could finish, it buys you coordination bugs and a larger bill.

Skill A packaged bundle of instructions, scripts, and reference material that a harness loads when the situation calls for it, teaching the model how to do a specific kind of work: generate a Word document, follow your team’s PR conventions, run a domain process. Skills are progressive disclosure for capability. The model sees a one-line description up front and pulls in the full playbook only when the task needs it, which makes them the practical place to put organizational know-how that a prompt is too temporary to hold and fine-tuning is too heavy to justify.

Computer Use An agent driving a graphical interface the way a person does, by looking at screenshots, moving a cursor, clicking, and typing. It extends agency to software that has no API, and it costs you speed and reliability against native tool use.

Trajectory The full record of one agent run: every model turn, tool call, observation, and decision from the opening prompt to the outcome. Trajectories are the raw material for debugging and evals, and increasingly for training. Labs improve agentic models by learning from the runs that worked.

Checkpointing Saving an agent’s progress at defined points so a long task can resume after a failure, an interruption, or a context reset instead of starting over. With compaction, it’s what makes a multi-hour agent run practical.

Human-in-the-Loop (HITL) Points where a person reviews, approves, or redirects an agent’s actions before they take effect. This is the primary governance control for an agentic system, and it’s a harness feature, not a model property.

Getting Knowledge Into the Model

Prompt Engineering Writing the instructions and examples that shape a model’s output. Prompt engineering is the entry-level literacy, and named prompt frameworks like CO-STAR and STOKE are the checklists that stop you leaving out the context and success criteria the task needed.

Few-Shot and Zero-Shot Prompting Zero-shot asks the model to do a task from instructions alone. Few-shot includes worked examples so the model infers the pattern. Examples remain the cheapest reliable lever on output quality and format compliance.

Chain of Thought (CoT) Prompting or training a model to reason step by step before it answers, named by Wei and colleagues in 2022. CoT is the prompt-level ancestor of reasoning models, and still the technique you reach for with a model that has no extended thinking built in.

Structured Output Constraining output to a schema, usually JSON, so downstream code can parse it. Modern APIs enforce this during decoding. Before that, engineers ran on prompting and prayer, and a whole category of production bugs lived in the gap.

Context Engineering Context engineering is the discipline of curating everything in the window: which instructions, which documents, which tool results, which history, in what order, at what level of compression. As tasks run longer and windows grow, it displaces prompt engineering as the core skill. The question stops being “what do I say” and becomes “what does the model need in view right now.”

RAG (Retrieval-Augmented Generation) Fetching relevant documents at request time and putting them in context so the model answers from them rather than from training data. Lewis and colleagues coined the term in 2020 for a much smaller setup than the one you’re building; the pattern outlived the architecture. It’s the standard way to ground a model in private, current, or bulky information.

Embeddings Numeric vectors that place semantically similar text near each other in vector space. They’re the machinery under semantic search and RAG retrieval.

Vector Database A datastore built for similarity search over embeddings: Pinecone, pgvector, Weaviate, and kin. It’s the retrieval backend RAG grew up on, though agentic search over plain files and grep has been eating into its territory.

Chunking Splitting documents into pieces sized for embedding and retrieval. Chunk boundaries quietly decide what a RAG system can and cannot find. When someone tells you the retrieval doesn’t work, look at the chunking before you look at the model.

Reranking A second-pass model that reorders retrieval results by relevance before they enter the context window. Cheap insurance on retrieval quality: cast a wide net, then let the reranker pick what actually matters.

Grounding Tying output to source material a reader can check, meaning retrieved documents, tool results, or structured data, instead of the model’s parametric memory. It’s the main defense against hallucination.

Hallucination (Confabulation) Fluent, confident output that is wrong or invented. This is a statistical property of next-token prediction, not a bug waiting for a patch. Grounding, tool use, and verification steps reduce it. Nothing eliminates it.

Fine-Tuning Further training a model on domain examples to shift its behavior or style. Teams reach for it early and regret it. The enterprise problem that looks like a fine-tuning problem is usually a context problem: cheaper to solve, faster to iterate, and it doesn’t freeze your knowledge at training time.

Compaction (Context Compression) Summarizing or pruning conversation history and tool results so a long task stays inside the window. Anthropic calls it the first lever in context engineering. Any agent working longer than a few minutes needs it, and it’s the harness’s job to do it.

Context Rot Performance degrading as the window fills with stale, redundant, or irrelevant material: old tool results, superseded plans, accumulated noise. Chroma tested 18 models and found that performance shifts with input length even on trivial tasks, and that how the context is arranged matters more than how much of it fits. This is the failure mode compaction and context engineering exist to prevent. A long context is a capacity, never a free lunch.

Prompt Caching Reusing the computed state of a stable prompt prefix (system prompt, tool definitions, reference documents) across requests. On the Claude API a cache read costs a tenth of the base input token price, so cache-aware prompt structure is the first place to look when agentic spend surprises you.

Memory Whatever persists information across sessions: files the agent writes to itself, retrieval over past conversations, structured stores the harness injects. The model is stateless, so all memory is a harness feature.

Training and Alignment

Pre-training The initial large-scale training run where a model learns language, facts, and reasoning patterns from a broad corpus. It sets raw capability.

Post-training (RLHF, RLAIF, RL) The phase that shapes a pre-trained model into a useful assistant: instruction following, tool-use conventions, refusals, style. RLHF learns from human preference labels, as in OpenAI’s InstructGPT. RLAIF puts a model in the labeler’s seat, as in Anthropic’s Constitutional AI. Labs now run reinforcement learning against their own harnesses, so a model learns the conventions of specific software. That’s one reason a model can shine in its native harness and underperform in a generic one.

Alignment The work of making model behavior match human intent and values, spanning training technique, evaluation, and deployment safeguards.

Distillation Training a smaller, cheaper model to imitate a larger one’s behavior on a task distribution. Hinton, Vinyals, and Dean described it in 2015, well before anyone needed it to cut an inference bill. It’s how labs produce fast mid-tier models, and how a sharp team gets frontier-quality behavior on a narrow task at small-model prices.

Quantization Reducing the numeric precision of a model’s weights to shrink its memory footprint and speed up inference, usually at modest quality cost. It’s the workhorse behind running capable models on constrained or local hardware.

Reasoning Model (Extended Thinking) A model trained to emit intermediate reasoning tokens before it answers, trading latency and cost for reliability on hard problems. The trace is also a debugging artifact. Reading how a model decomposed a problem is a shortcut to the metacognitive work of naming the steps in your own.

Quality, Safety, and Operations

Evals (Evaluations) Systematic tests of model or agent behavior against criteria you set before you looked at the output. Evals are the test suite of AI engineering, and whether a team has them is the strongest available predictor of whether its AI feature gets better over time. Vibes-based quality assessment stops scaling almost immediately.

Benchmark A standardized public eval for comparing models. SWE-bench scores agents on 2,294 problems taken from real GitHub issues and pull requests across 12 Python repositories; MMLU covers 57 subjects from mathematics to law. Both are a rough capability signal and a poor substitute for evals built on your own task distribution, because nobody’s product is 12 Python repositories.

Guardrails Runtime checks on inputs and outputs: content filters, schema validators, policy classifiers, spend limits. Like HITL, guardrails live in the harness and the system around it. That’s why a governance conversation turns out to be a harness conversation.

Sandboxing Isolating an agent’s execution environment with containers, restricted filesystems, and network allowlists so its actions stay bounded whatever it decides to do. Defense in depth for the gap between what a model should do and what it might.

Prompt Injection An attack where instructions hidden in data the model processes, such as a web page, an email, or an attached document, hijack its behavior. Simon Willison has catalogued the attacks and the attempted defenses since 2022. It’s the defining security problem of agentic systems, because an agent’s whole job is acting on content it reads.

Red Teaming Adversarial testing of an AI system: jailbreaks, injections, data exfiltration, and misuse scenarios, run before an attacker or a customer finds them. Evals ask whether it works. Red teaming asks how it breaks.

Observability (Tracing) Instrumentation that records every model call, tool invocation, and decision in a run so engineers can debug failures and audit behavior. Traces are to agents what logs are to services.

Determinism and Temperature Temperature controls output randomness. Even at zero, an LLM system is only approximately deterministic. Engineering around that means designing for verification and retry instead of assuming a repeat run returns the same answer.

The Latency-Cost-Quality Triangle The standing trade-off in AI engineering: bigger models and longer reasoning buy quality with speed and spend. You don’t escape the triangle. You route around it.

Model Routing Choosing which model handles a request based on difficulty, cost ceiling, or the capability the task actually needs. It’s a harness or gateway concern, and invisible to the user when it works.

AI Gateway A proxy between your applications and model providers that centralizes authentication, routing, rate limiting, caching, and spend tracking. Several of the levers in the AI Cost-Savings Scorecard land here. It’s the control point where AI spend stops being a surprise and becomes a number someone owns.

Resources