GLOSSARY
One term, one clear explanation.
Search the English name, Chinese name, acronym, and common alternative wording together. You can also return to the map to see how the ideas connect.
Artificial Intelligence (AI)
Technology that lets computers recognize, generate, judge, and act on information; the result still depends on software, data, tools, and human review.
AI Application
The website, desktop app, or software surface where a person uses a model with a specific interface, context, files, and tools.
Model
A trained system of parameters that processes input and produces an output; usually called by an application.
Agent
An AI system that works toward a goal, observes results, and decides what to do next, often with tools and a runtime environment.
Prompt
The task instructions given to an AI system, including the goal, background, constraints, and requested output.
Context
The information available to the model for the current task, such as instructions, conversation, retrieved material, and tool results.
Memory
Information an application stores and may provide again in later interactions.
Token
A unit of text or data used by a model for processing and measurement; not equivalent to one Chinese character or one English word.
Tool
A specific capability an AI system can invoke, such as searching, reading a file, calculating, generating a file, or changing an external system.
Skill
A reusable package of instructions, rules, and supporting material for a class of tasks, used by systems that support skills.
Model Context Protocol (MCP)
An open protocol that lets AI applications connect to external tools, resources, and services through a shared communication model.
Workflow
A task organized into ordered steps, with possible branches, repetition, checks, and stopping conditions.
Permission
The scope of data and actions that an application, service, account, or environment allows.
Hallucination
Content that appears plausible but is false, unsupported, or lacks reliable evidence.
Follow curiosity one level deeper.
Training, architecture, algorithms, and context mechanisms are available when you need them.
Model Training (Training)
Model training uses data and objectives to repeatedly adjust model parameters so that it performs better on a class of tasks.
Model Inference (Inference)
Model inference uses trained parameters to compute on the current input and produce an output.
Parameters (Parameters)
Parameters are a large number of values adjusted during training and used in computation; they collectively affect the output.
Training Data (Training Data)
Training data is the sample used for model learning; what the data covers and its quality affect what the model learns.
Pretraining (Pretraining)
Pretraining allows models to first acquire basic capabilities from large amounts of data that can be used for subsequent tasks.
Next-token Prediction (Next-token Prediction)
Next-token prediction estimates the probability of each subsequent token based on existing content.
Loss Function (Loss Function)
Loss functions convert the gap between predictions and training targets into a number for the training program to optimize.
Backpropagation (Backpropagation)
Backpropagation computes derivatives backwards through the computation to determine how each parameter affects the loss change.
Gradient (Gradient)
The gradient describes how the loss changes when parameters are slightly adjusted in each direction at the current position.
Optimizer (Optimizer)
The optimizer combines the gradient with the update rule to determine how the parameters are updated in this step.
Stochastic Gradient Descent (SGD)
SGD estimates the gradient using sampled examples, then updates parameters in the direction that reduces the loss.
Adaptive Moment Estimation (Adam)
Adam maintains historical statistics of gradients and adjusts the update scale for different parameters, making it a common gradient optimization algorithm.
Post-training
Post-training continues to adjust the behavior of base models, making them more suitable for following instructions, answering questions, or completing specific tasks.
Supervised Fine-Tuning (SFT)
Supervised fine-tuning uses demonstrations of inputs and ideal outputs to help existing models learn response patterns that better align with task requirements.
Reinforcement Learning from Human Feedback (RLHF)
RLHF uses human feedback on outputs to shape model behavior; a common pipeline converts preferences into rewards and then retrains the model.
Proximal Policy Optimization (PPO)
PPO is a class of policy optimization methods in reinforcement learning that encourages beneficial behaviors while discouraging overly large policy changes.
Direct Preference Optimization (DPO)
DPO directly optimizes models using data about which answers are preferred, without requiring a separate reward model in the standard approach.
Group Relative Policy Optimization (GRPO)
GRPO generates a set of answers for the same problem and uses the relative differences in rewards within the group to guide policy updates.
Low-Rank Adaptation (LoRA)
LoRA freezes the original weights and changes the model's behavior on target tasks by training a small number of low-rank adaptation parameters.
Model Evaluation (Evaluation)
Model evaluation uses clear tasks and standards to check results and determine if capabilities meet the goals.
Transformer Architecture (Transformer)
Transformer is a neural network architecture that uses operations like attention to allow different positions in the input to exchange information.
Attention Mechanism (Attention)
Attention allows one position to gather information from other allowed positions according to computed weights.
Context Window (Context Window)
The context window is the range of tokens that a model can accommodate in a single computation; it limits how much input and generated content can be processed in that instance.
Output Length and Budget (Output Budget)
Maximum output length is the maximum number of tokens allowed for generation at once; it is a related but different constraint from the context window.
Truncation and Compression (Context Management)
Context management is about selecting what information to retain, how to compress it, and when to retrieve data when capacity is limited.
Retrieval-Augmented Generation (RAG)
RAG first finds content relevant to the question from external sources, then provides it to the model to assist in generating answers.
Key-Value Cache (KV Cache)
KV cache stores attention keys and values of processed tokens, reusing them in subsequent generation to reduce redundant computation.
MCP Connection Type (Transport)
How MCP client and service exchange messages; common methods include local stdio and HTTP-based remote connections.
MCP Capabilities (Capabilities)
MCP services can provide tools, resources, and prompt templates according to conventions; what specifically can be used depends on the service implementation.
MCP Lifecycle
A typical MCP usage goes through initialization, capability discovery, specific calls, result return, and termination or continuation.