Skip to content
BotServBotServ
GlossaryAI TermsLLMRAGEmbeddingAgent

AI Glossary: Key Terms Explained

Essential AI glossary covering LLM, RAG, Embedding, Agent, Tool-Calling and more with clear explanations.

S

schutzgeist

5 min read
AI Glossary: Key Terms Explained

Glossary: AI Concepts Explained

What this article covers

  • The most important AI concepts explained clearly.
  • From LLM to RAG, Embedding, and Agent.
  • When each term matters and what it means.
  • Straightforward explanations without jargon.
  • Links to detailed articles for deeper learning.

Introduction: Understanding AI Concepts

The AI world is filled with technical terms: LLM, RAG, Embedding, Agent, Tool-Calling, Fine-Tuning, Prompt Injection. This glossary breaks down the most important concepts in plain language, whether you’re just starting out or already experienced.

This article is for anyone who wants to understand AI terminology. For fundamentals, check out What is local AI?

Why do you need a glossary?

Imagine reading an article about “RAG with embeddings and tool-calling.” What does that actually mean? This glossary explains it: RAG means searching documents plus generating answers. Embedding means converting text to numbers. Tool-Calling means the model uses available tools. Clear and simple.

Essential AI concepts

A

Agent (AI Agent) An AI system that acts autonomously: it plans, uses tools, and takes action. Not just answering questions, but solving tasks. See AI Agents.

API (Application Programming Interface) An interface for programs to communicate. For AI: how you interact with models (Ollama API, OpenAI API).

B

Batch Processing Processing many requests simultaneously. For performance: vLLM uses batch processing to handle parallel queries.

Benchmark A test to measure performance. For models: how good is it? MMLU, HumanEval, etc.

C

Chunking Breaking text into smaller pieces. For RAG: splitting documents into chunks so they fit in the context window. See Chunking.

Cloud AI AI services hosted in the cloud: OpenAI, Anthropic, Google. Your data goes to third parties. The opposite is local AI.

Context Window How much text the model can process. 4K, 8K, 32K, 128K tokens. See Context Length.

D

Data Protection Safeguarding personal data. For AI: processing locally instead of in the cloud. See Data Protection.

Docker A containerization platform. For AI: running Ollama, Qdrant, and agents in containers. See Docker.

E

Embedding Converting text into numbers (vectors). For semantic search: similar texts produce similar vectors. See Embeddings.

Embedding Model A model that creates embeddings: nomic-embed-text, multilingual-e5, bge-m3. See Embedding Models.

F

Fine-Tuning Training a model on specific data. For specialized tasks: medicine, law, your own documents.

Function-Calling A model calls functions. For agents: the LLM decides which tool to use. See Function-Calling.

G

GGUF A file format for quantized models. For Ollama and LM Studio: load models in GGUF format.

GPU (Graphics Processing Unit) A graphics card for LLM inference. For AI: VRAM determines model capacity. See Hardware.

H

Hallucination The model invents facts. For RAG: citations help prevent hallucinations. See Hallucinations.

Hardware Machines for AI: GPU, VRAM, RAM, CPU. See Hardware Requirements.

I

Inference The model generates answers. For AI: LLM inference on GPU.

L

LLM (Large Language Model) A large language model: llama3.1, mistral, qwen. For text generation, chat, analysis. See Text Models.

Local AI AI on your own computer: Ollama, LM Studio. No data sent to third parties. See Local AI.

M

MCP (Model Context Protocol) A protocol for tool integration. For agents: connecting to file systems, databases, APIs. See MCP.

Model An AI model: LLM, Vision, Embedding. For different tasks: text, images, search.

O

Ollama A local model server for LLMs. For local AI: load models and use the API. See Ollama.

OCR (Optical Character Recognition) Extracting text from images. For documents: Tesseract or vision models. See OCR.

P

Prompt An instruction for the model. For AI: what should the model do? See Prompting.

Prompt Injection An attack that manipulates the model into performing harmful actions. For security: input validation, permissions. See Prompt Injection.

Q

Quantization Compressing a model to use less memory. Q4, Q8: smaller models, less VRAM. For local AI: Q4 is the standard.

Query Engine A search engine for RAG. For LlamaIndex: finds relevant documents and generates answers.

R

RAG (Retrieval-Augmented Generation) Search first, answer second: search documents (Retrieval), then generate an answer (Generation). For Q&A systems. See Local RAG.

Reasoning Model A model that “thinks” before answering: DeepSeek-R1, QwQ. For complex problems: math, logic, debugging. See Reasoning Models.

S

Sandboxing Isolation for critical tools. For security: run code in containers without system access. See Sandboxing.

Semantic Search Searching by meaning, not keywords. For RAG: embeddings find similar texts.

System Prompt Instructions that define the model’s role and behavior. For agents: “You are a helpful assistant.”

T

Token The smallest unit for LLMs. 1 token ≈ 0.75 words. For costs: cloud APIs charge per token.

Tool-Calling The model calls tools. For agents: the LLM decides which tool to use. See Tool-Calling.

V

Vector Database A database for embeddings: Qdrant, ChromaDB, Weaviate. For RAG: stores embeddings. See Vector Databases.

Vision Model A model that understands images: LLaVA, MiniCPM-V. For image analysis, OCR, diagrams. See Vision Models.

VRAM (Video RAM) GPU memory for LLMs. For AI: how much VRAM does the model need? See VRAM Calculator.

W

Workflow An automated process: trigger → processing → action. For n8n, Node-RED: build workflows.

Further resources

FAQ

What is an LLM?

Large Language Model: a large language model like llama3.1, GPT-4, or Claude. It understands and generates text for chat, analysis, code, and more.

What is RAG?

Retrieval-Augmented Generation: first search documents (Retrieval), then generate an answer (Generation). Used for Q&A systems and knowledge bases.

What is an embedding?

Converting text into numbers (vectors). Similar texts produce similar vectors. Used for semantic search and RAG.

What is an AI agent?

An AI system that acts autonomously: it plans, uses tools, and takes action. Not just answering questions, but solving tasks.

What is MCP?

Model Context Protocol: a protocol for tool integration. It connects agents to file systems, databases, APIs, and other tools.

What is VRAM?

Video RAM: GPU memory for LLMs. The larger the model, the more VRAM it needs. 8B model: about 5 GB. 70B: about 40 GB.

What is Ollama?

A local model server for LLMs: load models, use the API, everything stays local. Great for privacy and no cloud costs.

What is prompt injection?

An attack that manipulates the model into performing harmful actions. Prevented through input validation, permissions, and sandboxing.

References and further reading

Back to Blog
Share:

Related Posts