AI Fundamentals Glossary
What This AI Fundamentals Glossary Covers
- All key terms from BotServ’s foundational articles.
- Beginner-friendly explanations with examples.
- Cross-references to articles where terms are explained in detail.
- Organized by topic: Local AI, AI Agents, AI Hardware.
Introduction: AI Fundamentals Glossary
Learning about local AI, AI agents, and AI hardware means encountering many technical terms. This glossary brings together all the central concepts from BotServ’s foundational articles in one place. Each term gets a brief, accessible explanation, often with an example. When you need more detail, you’ll find a link to the corresponding foundational article.
Local AI Terms
LLM
LLM stands for Large Language Model. An LLM is an AI that can process and generate text. Examples include Llama 3.1, Qwen 2.5, and Mistral. An LLM consists of billions of parameters and was trained on vast amounts of text. Learn more in What is Local AI?.
Inference
Inference is the process where a trained model takes input and produces output. When you start Ollama, enter a prompt, and get a response, that’s inference. The opposite is training, which creates the model in the first place. Learn more in Running LLMs Locally.
Training
Training is the process of creating a model. A model learns from large datasets so it can understand and generate text. Training demands enormous computing power and takes weeks or months. Inference is using the finished model. Learn more in What is Local AI?.
Token
A token is the smallest unit a language model processes. It can be a word, part of a word, or a character. A single word typically equals about 1.3 to 1.5 tokens in German. A model’s context length is measured in tokens. Learn more in Context Length.
Context Length
Context length, also called context window, is the maximum number of tokens a model can process at once. A model with 8,000 tokens of context can read roughly 6,000 words simultaneously. Longer texts must be split or shortened. Learn more in Context Length.
Context Window
Synonym for context length. Both terms describe the maximum input length a model can handle, measured in tokens.
KV-Cache
The KV-Cache is a buffer that builds up during inference. It stores intermediate values for every token the model has already processed. With long contexts, the KV-Cache grows and consumes additional memory. This is why longer contexts require more VRAM or RAM. Learn more in Context Length.
Prompt
A prompt is the input you give to a language model. It can be a question, an instruction, or text for the model to work with. The prompt plus the model’s response must together fit within the context window.
Completion
Completion is the response the model generates after a prompt. In chat models, the completion is the message the model returns.
System Prompt
A system prompt is an instruction the model receives at the start of a conversation. It defines the model’s role, behavior, and boundaries. Example: “You are a helpful assistant for local AI. Answer in English.”
Ollama
Ollama is a runtime environment for local language models. It downloads models, launches them, and provides an API. Ollama is the simplest way to get started with local AI. Learn more in Running LLMs Locally and Ollama.
LM Studio
LM Studio is a graphical application for local language models. It provides a user interface for loading and testing models. LM Studio is an alternative to Ollama for users who prefer clicking to typing. Learn more in Running LLMs Locally.
GGUF
GGUF is a file format for quantized language models. It’s used by llama.cpp and Ollama. GGUF files store the model in compressed form so it runs with less memory. Learn more in Quantization.
Quantization
Quantization reduces the precision of model weights to save memory. A model originally trained with 16-bit values gets reduced to 4, 5, or 8 bits. This significantly lowers memory requirements. The trade-off: response quality may decline slightly. Learn more in Quantization.
FP16
FP16 stands for 16-bit floating-point numbers. This is the standard format in which models are trained. An FP16 model is large and memory-hungry. For local AI, it’s usually quantized.
INT8
INT8 is an 8-bit integer representation. A model in INT8 uses roughly half the memory of FP16. Quality typically stays close to the original.
INT4
INT4 is a 4-bit representation. A model in INT4 uses about a quarter of the memory of FP16. INT4 is the most common quantization level for local AI on consumer hardware.
Q4, Q5, Q6, Q8
These abbreviations denote quantization levels of 4, 5, 6, or 8 bits. Q4 saves the most memory, Q8 offers the best quality. Learn more in Quantization.
K-Quantization
K-Quantization is an improved mixed quantization approach. It stores sensitive parts of the model with more bits than less sensitive ones. Q4_K_M is better than Q4_0 because it gives more precision to important weights. Learn more in Quantization.
RAG
RAG stands for Retrieval-Augmented Generation. The model searches for relevant text passages from a database and uses them as context for its answer. RAG lets you use a model with your own documents without retraining it. Learn more in Local RAG.
Embedding
An embedding is a numerical representation of text. The text is converted into a vector of numbers that captures its meaning. Embeddings are used for RAG search. Learn more in Embedding Models.
Vector Database
A vector database stores embeddings and enables searching for similar texts. Examples include Chroma and Qdrant. Vector databases are central to RAG systems. Learn more in Vector Databases.
Chunking
Chunking is the process of splitting documents into smaller pieces before loading them into a vector database. Each chunk becomes an embedding. Good chunking is essential for good RAG results. Learn more in Chunking.
Self-Hosting
Self-hosting means running software on your own hardware instead of using it as a cloud service. Local AI is a form of self-hosting. Learn more in the article What is local AI?.
Cloud AI
Cloud AI means the model runs on a provider’s servers. You send your input via API or browser, and the response comes back. Examples include ChatGPT, Claude, and Gemini. Learn more in the article Local AI vs. Cloud AI.
API
API stands for Application Programming Interface. An API is an interface through which programs communicate with each other. Ollama, for example, provides an API at http://localhost:11434 that other programs can use to interact with the model.
Hybrid
A hybrid setup combines local AI and cloud AI. Simple or non-sensitive tasks run in the cloud, while sensitive or frequent tasks run locally. Learn more in the article Local AI vs. Cloud AI.
GDPR
The GDPR (General Data Protection Regulation) governs how personal data is handled in the EU. Local AI helps with compliance because data never leaves your own network. Learn more in the article Local AI vs. Cloud AI.
Modelfile
A Modelfile is a configuration file for Ollama. It defines which model is loaded, which system prompt is used, and which parameters apply. Think of it as a Dockerfile, but for AI models.
Parameters
Parameters are the weights of a model. They are learned during training and determine how the model processes inputs. Models are often named after their parameter count, for example 7B for 7 billion parameters.
AI Agent Terminology
Agent
An agent is an AI system that pursues a goal and takes action independently. It observes its environment, plans steps, and uses tools to reach its objective. Learn more in the article What is an AI Agent?.
AI Agent
Synonym for agent. An AI agent uses a language model as its brain and tools as its hands to complete tasks.
Chatbot
A chatbot is a dialogue system that responds to input with text. It reacts to queries but doesn’t plan independently and doesn’t use tools. Learn more in the article Chatbot vs. AI Agent.
Tool
A tool is an external function that an agent can call. Examples include web search, file access, database queries, or email sending. Tools are what make an agent truly useful.
Tool-Calling
Tool-calling is the ability of a language model to invoke external functions. The model doesn’t just produce text, but generates a structured message with a function name and parameters. Your program executes the function and returns the result. Learn more in the article Tool-Calling.
Function Calling
Synonym for tool-calling. Both terms refer to the same capability.
Schema
A schema describes a tool for the model. It contains the name, a description, and the parameters with their types. The model uses the schema to decide whether and how to call the tool.
Memory
Memory is an agent’s storage. Short-term memory holds current steps and intermediate results. Long-term memory preserves knowledge across multiple sessions.
Mnemonic Storage
English term for memory. See above.
Planning
Planning is breaking down a goal into individual steps. The agent decides which action to take next and adapts the plan based on intermediate results.
Autonomy
Autonomy describes how independently an agent makes decisions. Greater autonomy means more possibilities but also more risk. Autonomous agents should run in sandboxed environments.
Multi-Agent
Multi-agent means multiple agents with different roles work together. One agent writes code, another reviews it, a third documents it. Examples include CrewAI and AutoGen. Learn more in the Frameworks overview.
Workflow
A workflow is a sequence of steps that an agent executes. Workflows can be linear or branched. A research workflow might look like this: search, read, summarize, store.
MCP
MCP stands for Model Context Protocol. It’s a standard for tools and context that agents can use. MCP makes it easier to share tools between different frameworks.
Reasoning
Reasoning is the ability of a model to think logically and draw conclusions. Models with strong reasoning capabilities are better suited for agents because they plan steps more effectively.
ReAct
ReAct is an agent architecture that combines reasoning and acting. The agent thinks first (Reason), then acts (Act), then observes the result (Observe), and starts again. Learn more in the article What is an AI Agent?.
Plan-and-Execute
Plan-and-execute is an agent architecture in which the agent creates a complete plan first and then executes it step by step. Unlike ReAct, the plan is created upfront rather than being adapted at each step.
Sandbox
A sandbox is an isolated environment where an agent runs. It prevents the agent from performing unexpected actions on the actual system. Sandboxing is critical for safety with autonomous agents.
Human Approval
Human approval (human-in-the-loop) means critical actions by an agent must be confirmed by a person before execution. This increases safety for autonomous workflows.
Framework
A framework is a software library that helps you build agents. Frameworks handle the loop of observing, planning, and executing. Examples include LangGraph, CrewAI, and AutoGen. Learn more in the Frameworks overview.
AI Hardware Terminology
RAM
RAM (Random Access Memory) is your computer’s working memory. It’s available to the CPU and all programs. During CPU inference, the model resides in RAM. Learn more in the article RAM vs. VRAM.
VRAM
VRAM (Video RAM) is the memory on a dedicated graphics card. It’s exclusively available to the GPU and is significantly faster than RAM for AI computations. During GPU inference, the model resides in VRAM. Learn more in the article RAM vs. VRAM.
GPU
GPU stands for Graphics Processing Unit, the processor on a graphics card. GPUs have many parallel cores and are significantly faster than CPUs for AI inference. Learn more in the article RAM vs. VRAM.
CPU
CPU stands for Central Processing Unit, your computer’s main processor. CPUs have few but flexible cores. AI inference on a CPU is possible but slower than on a GPU.
GPU-Offloading
GPU-offloading moves portions of a model to the GPU when the entire model doesn’t fit in VRAM, while the remainder stays in RAM. This approach is slower than full GPU inference but faster than pure CPU inference. See RAM vs. VRAM for more details.
Memory Bandwidth
Memory bandwidth is the rate at which data transfers between storage and processor. High bandwidth matters for AI because models move large amounts of data constantly. VRAM typically offers higher bandwidth than system RAM.
Unified Memory
Unified Memory allows the CPU and GPU to access the same memory pool. A Mac with 32 GB of RAM can use all 32 GB for GPU computation. By contrast, a discrete graphics card has separate VRAM. See RAM vs. VRAM for details.
Shared Memory
Shared Memory is another term for Unified Memory. Both refer to a single memory space accessible by both CPU and GPU.
APU
APU stands for Accelerated Processing Unit. An APU integrates the CPU and GPU on a single chip. Apple Silicon and AMD Ryzen AI Max are common examples. The advantage is that both processors can access the same memory pool. Learn more in AI Hardware Basics.
Apple Silicon
Apple Silicon refers to Apple’s custom chips: M1, M2, M3, and M4. They combine CPU and GPU on one die and use Unified Memory. Apple Silicon is popular for local AI because it handles large memory capacity efficiently. See RAM vs. VRAM for more.
TOPS
TOPS stands for Tera Operations Per Second. It measures AI compute performance. A higher TOPS value indicates a chip can potentially execute more AI operations per second. However, TOPS alone doesn’t tell the full story; memory bandwidth and architecture matter equally. Learn more in AI Hardware Basics.
NTOPS
NTOPS stands for NPU TOPS, measuring the AI compute performance of an NPU (Neural Processing Unit). NPUs are specialized AI accelerators found in modern chips. NTOPS measure only NPU performance, not GPU performance.
SSD
SSD stands for Solid State Drive, a type of flash storage. An SSD loads models significantly faster than a traditional hard drive (HDD). NVMe SSDs are even faster than SATA SSDs.
NVMe
NVMe is an interface standard for SSDs. NVMe drives deliver substantially faster speeds than SATA SSDs and can load large models in seconds rather than minutes. For local AI, an NVMe SSD is recommended.
Memory Requirement
Memory requirement is the amount of RAM or VRAM a model needs during inference. It depends on model size, quantization, and the KV-cache. See RAM and VRAM Requirements for more.
Further Reading and Glossary Resources
- Local AI Basics
- AI Agents Basics
- AI Hardware Basics
- What is Local AI?
- What is an AI Agent?
- RAM vs. VRAM
FAQ - Glossary Questions
What’s the difference between training and inference?
Training creates a model using large datasets. Inference uses the finished model to generate answers. Local AI relies on inference.
What’s the difference between RAM and VRAM?
RAM is your computer’s main memory; VRAM is a graphics card’s memory. CPU inference loads the model into RAM, GPU inference into VRAM. GPU inference is significantly faster. Learn more in RAM vs. VRAM.
What’s the difference between Tool-Calling and Function Calling?
No difference. Both terms describe the same capability: a language model that can invoke external functions. See Tool-Calling for details.
What’s the difference between context length and context window?
No difference. Both refer to the maximum number of tokens a model can process at once. Learn more in Context Length.
What’s the difference between quantization and compression?
Quantization is a specific type of compression. It reduces the precision of model weights, for example from 16-bit to 4-bit. Compression is the broader term.
What’s the difference between an agent and a chatbot?
A chatbot responds to inputs with text. An agent plans multiple steps, uses tools, and pursues a goal. See Chatbot vs. AI Agent for more.
What’s the difference between Unified Memory and Shared Memory?
No difference. Both terms describe a single memory space shared by CPU and GPU, as in Apple Silicon or Ryzen AI Max.
What’s the difference between Q4_0 and Q4_K_M?
Q4_0 quantizes all weights uniformly. Q4_K_M uses mixed quantization, storing more sensitive model components with higher precision. Q4_K_M generally delivers better results. Learn more in Quantization.
What does GGUF mean?
GGUF is a file format for quantized language models. It’s used by llama.cpp and Ollama. GGUF files contain the model in compressed form.
What’s the difference between RAG and Fine-Tuning?
RAG retrieves relevant passages from a database and feeds them to the model as context. Fine-tuning further trains the model on new data. RAG is simpler and more flexible; fine-tuning requires more effort but integrates deeper.
What’s the difference between TOPS and TFLOPS?
TOPS measures integer operations (INT), while TFLOPS measures floating-point operations (FP). For AI inference with quantized models, TOPS is more relevant. For training, TFLOPS matters more.
What’s the difference between MCP and Tool-Calling?
Tool-Calling is the model’s ability to invoke functions. MCP is a standard that defines how tools and context are structured and exposed. MCP uses Tool-Calling as its underlying mechanism.


