Concepts: Foundations Behind Local AI, Agents, and Self-Hosting
What This Article Covers
- Core ideas behind local AI, agents, and self-hosting.
- Links to relevant articles for each concept.
- How these concepts connect.
Introduction
Tutorials teach you how. This section explains what and why. Understanding these concepts helps you choose better tools and debug faster. The topics below are the building blocks that make up most of what BotServ covers.
Local AI: Why Run Locally Instead of Cloud?
Local AI runs on your device rather than on a provider’s servers. The trade-offs:
- Privacy: Prompts and documents stay on your hardware. Essential for sensitive data.
- Cost: One-time hardware investment instead of recurring API charges.
- Independence: No vendor lock-in, no surprise price increases, no outages from third parties.
- Control: You choose the model, version, and parameters.
Get started: What is Local AI? and Local AI vs Cloud AI.
Language Models and Quantization
A language model (LLM) is a neural network that predicts text. Size is measured in parameters (7B = 7 billion). To fit the model into memory, it’s quantized: the precision of its weights is reduced, for example from 16-bit to 4-bit.
Get started: Quantization, RAM vs VRAM, and Model Size and Memory Requirements.
RAG: Load Knowledge Without Retraining
Retrieval-Augmented Generation works like this: instead of answering questions only from its training data, the model fetches relevant documents from a database first. This lets a local AI answer questions about your own documents without them being part of its training.
The process:
- Documents are split into small chunks.
- Each chunk is converted to an embedding (a vector of numbers).
- When you ask a question, the question’s embedding is compared against document embeddings.
- The most similar chunks are passed to the model as context.
Get started: Local RAG, Embedding Models, and Vector Databases.
AI Agents: Models That Take Action
An AI agent goes beyond a chatbot: it can invoke tools, make decisions, and work through tasks step by step. The most popular pattern is ReAct: Reason (think) → Act (call a tool) → Observe (read the result) → repeat.
Get started: What is an AI Agent?, Tool-Calling, and Multi-Agent Systems.
Self-Hosting: Running Services on Your Own Hardware
Self-hosting means running software like Nextcloud, n8n, or Open WebUI on your own hardware or a rented server instead of using finished cloud services. Benefits: privacy, control, and learning. Costs: maintenance and responsibility.
Get started: Self-Hosting Basics, Docker Basics, and Homelab.
Security: Prompt Injection and Access Control
Local AI isn’t automatically secure. Prompt injection, unsafe tool calls, and open ports are real risks. The Secure Operation section covers this systematically: from access control to secrets management to agent security.
Further Reading
- Glossary - Key terms explained clearly.
- Literature - Books, papers, and blogs.
- Local AI - The main section for self-hosting.
- AI Agents - Agents, frameworks, and workflows.
FAQ - Frequently Asked Questions
In what order should I learn these concepts? Start with local AI fundamentals, then quantization and model selection. Next, RAG, then agents. Self-hosting can run in parallel if you want to run tools like Ollama on your own server.
Are these concepts only for advanced users? No. All concepts are explained for beginners. The links lead to articles that need no prior knowledge.
What’s the difference between a concept and a tutorial? Concepts explain ideas and relationships. Tutorials show step-by-step how to install or configure something. You’ll find concepts here, and tutorials in the relevant sections.

