AnythingLLM: RAG Desktop App for Local AI
What this article covers
- What AnythingLLM is and how it lets you chat with your documents without writing a single line of code
- How to install AnythingLLM as a desktop app or via Docker, and connect it to local models
- How Workspaces function and why they make managing different document collections easier
- Which features AnythingLLM offers, from Agent Mode to pinning answers
- Common pitfalls and how to avoid typical issues
Introduction: Understanding AnythingLLM
If you work with local AI, you quickly run into a familiar desire: you want to chat with your own documents. A PDF manual, a collection of notes, a contract, all that shouldn’t just sit on your hard drive, but be queryable. You ask a question and get an answer based on your content.
That’s exactly where AnythingLLM shines. It’s a desktop app that makes RAG possible without programming. You don’t need Python, API keys, or the command line. You install the app, upload documents, and chat about them. Behind the scenes, AnythingLLM handles embeddings, vector databases, and context management.
In this article, you’ll learn what AnythingLLM can do, how to install it, and how to use it in your daily workflow. If you’re new to the topic, first read our article What is local AI?.
Why do you need AnythingLLM?
Imagine you have a folder full of PDFs: manuals, research reports, invoices, contracts. You want to ask targeted questions about them without opening and searching through each document individually.
Without AnythingLLM, it looks like this:
- You search for a Python RAG framework and read through the documentation
- You write code for embeddings, vector databases, and retrieval
- You set up an API or configure Ollama manually
- You debug errors until it finally works
With AnythingLLM, you open the app, drag your PDFs into a Workspace, and ask your question. Done. No code, no configuration, no API setup.
You might want to do any of the following:
- Upload documents and ask targeted questions about them without programming
- Keep different topics clean and separate instead of mixing everything together
- Use local models via Ollama or LM Studio without sending data to the cloud
- Use an agent that combines web search and RAG
- Have a desktop app that works like a normal program
AnythingLLM solves all of these. For more background, check out our overview of Local AI Software.
AnythingLLM at a glance
AnythingLLM is a desktop application for RAG using local and cloud-based models. It manages your documents in Workspaces, automatically creates embeddings, and lets you ask questions about your content via chat. You don’t need programming skills.
Its key characteristics:
- Desktop app: Runs as a normal program on Windows, macOS, and Linux
- RAG without code: Upload documents, ask questions, done
- Workspaces: Each Workspace has its own documents and its own context
- Multi-provider: Works with Ollama, LM Studio, OpenAI, Anthropic, and others
- Agent Mode: Built-in agents for web search, RAG, and tools
- Docker option: Available for server deployment or multi-user setups
If you’re unfamiliar with Ollama, start with our article on Ollama.
Who is AnythingLLM for?
AnythingLLM is designed for several audiences:
- Individual users who want to chat with their documents without programming
- Knowledge workers who want to make research, notes, and manuals queryable
- Small teams building a shared knowledge base
- Developers who want to experiment with RAG quickly before building their own solution
- Self-hosters who want to run AnythingLLM via Docker on a server
You don’t need prior knowledge. If you use the desktop app, installation is as straightforward as any other program. For RAG fundamentals, see our article RAG Basics.
Key terms related to AnythingLLM
| Term | Explanation |
|---|---|
| AnythingLLM | Desktop app for RAG with local and cloud models, without programming |
| Workspace | A separate area with its own documents, context, and configuration |
| RAG | Retrieval-Augmented Generation; asking questions about your own documents |
| Embedding | A numerical vector that represents the content of a text chunk for search purposes |
| Vector DB | Vector database where embeddings are stored and searched |
| Document | A file you upload to a Workspace, such as PDF or DOCX |
| Ollama | Tool that loads and runs local models, often used as a backend for AnythingLLM |
| LLM Provider | A service or tool that supplies the language model, either locally or in the cloud |
| Agent | A mode in which the model uses tools, such as web search or RAG |
| Pin | Keep an answer or text snippet so it stays in the context |
Installation
AnythingLLM offers two installation paths: the desktop app for individual users and Docker for server deployment.
Desktop app
The desktop app is the simplest option. You download it from the official website and install it like any other program.
- Windows: Download the installer, run it, and follow the wizard
- macOS: Download the
.dmgfile, open it, and drag AnythingLLM to your Applications folder - Linux: Download the appropriate package,
.debfor Ubuntu and Debian or.AppImagefor other distributions
After installation, launch AnythingLLM like any other application. On first startup, a setup wizard guides you through selecting your LLM provider.
Docker
For server deployment or if you prefer Docker, there’s an official Docker image.
docker run -d -p 3001:3001 \
-v anythingllm:/app/server/storage \
-e STORAGE_DIR=/app/server/storage \
--name anythingllm \
--restart always \
mintplexlabs/anythingllm
Here’s what each part does:
-d: Container runs in the background-p 3001:3001: Port 3001 on your machine maps to port 3001 in the container-v anythingllm:/app/server/storage: Documents and configuration persist--restart always: Container automatically starts after a reboot
Then open http://localhost:3001 in your browser. The initial setup walks you through choosing your LLM provider and embedding source.
Setting up LLM providers
AnythingLLM supports various providers, both local and cloud-based. You decide where your requests are processed.
Local providers
- Ollama: The most common choice for local models. Install Ollama, pull a model with
ollama pull llama3.1, and select “Ollama” as the provider in AnythingLLM. More details in Ollama. - LM Studio: A desktop app that loads models and provides an OpenAI-compatible API. Enter the LM Studio API address in AnythingLLM. Details in LM Studio.
Local providers have the advantage that your data never leaves your machine. However, you need sufficient hardware, especially for larger models. Learn more in Running LLMs locally.
Cloud Providers
- OpenAI: You’ll need an API key. Requests go through OpenAI’s servers, and your documents are sent there for processing.
- Anthropic: Same approach as OpenAI. You need an API key and send requests to Anthropic’s servers.
Cloud providers are simpler to set up and require no local hardware. The trade-off is that your data leaves your machine. You decide which compromise works for you.
Embedding Providers
Beyond choosing an LLM provider, you also need to select an embedding provider. AnythingLLM uses Ollama’s built-in embedding engine when you choose Ollama as your provider. You can also use OpenAI embeddings or AnythingLLM’s built-in embedding model. For fully local operation, Ollama or the built-in model are your best options.
Creating Workspaces
Workspaces are AnythingLLM’s core concept. A workspace is an isolated area with its own documents, context, and configuration. You can create as many workspaces as you need.
Examples of workspaces:
- Manuals: All PDF guides for your equipment
- Contracts: Invoices, contracts, and legal documents
- Notes: Your own notes and Markdown files
- Project: Documents related to a specific project
Each workspace includes:
- Its own document collection
- Its own chat history
- Its own system prompt configuration
- Its own context, meaning the model only sees documents in that workspace
This matters: when you ask a question in the “Manuals” workspace, the model searches only those manuals, not your contracts. It keeps answers focused and prevents confusion.
To create a workspace, use the sidebar. Click “New Workspace”, enter a name, and you’re ready to go. Then upload documents and start chatting.
Learn more about RAG in our article Local RAG.
Uploading Documents and Chatting
Once you’ve created a workspace, upload your documents. AnythingLLM supports these formats:
- PDF: The most common case, like manuals and reports
- DOCX: Word documents
- TXT: Plain text files
- MD: Markdown files
- CSV: Tabular data
- HTML: Web pages saved as files
The process:
- Open the workspace where you want to load documents.
- Click “Upload Documents” or drag and drop files into the window.
- AnythingLLM processes the files automatically: text is extracted, split into chunks, converted to embeddings, and stored in the vector database.
- Once upload completes, you can ask questions about the documents in chat.
You don’t need to worry about embeddings, chunking, or the vector database. AnythingLLM handles that in the background. You upload, wait briefly, and chat.
A typical workflow: upload a 50-page PDF to the “Manuals” workspace, ask “How do I reset the clock?”, and get an answer with references to the relevant part of the document.
Agent Mode
AnythingLLM includes an agent mode that goes beyond simple chat. In agent mode, the model can use tools to complete tasks.
Available tools include:
- Web Search: The model searches the internet for current information
- RAG: The model searches documents in the current workspace
- File Creation: The model can generate files and offer them for download
- Summarization: Automatically summarize long documents
- Lists and Tables: Generate structured output
You enable agent mode in the chat window. Select the mode and the model decides which tool to use for your question. For example, if you ask “What do my documents say about data protection and what does the current GDPR say about it?”, the agent combines RAG with web search.
Agent mode is especially useful when you want to pull information from multiple sources. It works with local and cloud models, though not every model supports all tools equally well.
AnythingLLM vs. Open WebUI vs. Custom RAG
These three approaches are often compared, but they suit different scenarios:
| Feature | AnythingLLM | Open WebUI | Custom RAG |
|---|---|---|---|
| Type | Desktop App | Web Interface | Own Project |
| RAG without Code | Yes | Yes, limited | No |
| Workspaces | Yes, core feature | No | Self-built |
| Agent Mode | Yes, built-in | Limited | Self-built |
| Desktop App | Yes | No | No |
| Docker Support | Yes | Yes | Self-built |
| Multi-User | Yes, in Docker version | Yes | Self-built |
| Target Audience | Individual users, knowledge workers | Teams, families | Developers |
In short: AnythingLLM is the best choice if you want RAG without coding and need workspaces. Open WebUI is better if you manage multiple users and prefer a pure web interface. Custom RAG only makes sense if you need full control and can code. Learn more about Open WebUI in our article Open WebUI.
Example: Searching PDFs with AnythingLLM
Here’s a complete workflow from installation to your first question:
- Install AnythingLLM: Download the desktop app and install it. Launch the app.
- Run the setup wizard: Choose Ollama as your LLM provider and Ollama as your embedding provider. If Ollama isn’t running yet, install it first and pull a model with
ollama pull llama3.1. - Create a workspace: Click “New Workspace” and name it, for example, “Manuals”.
- Upload documents: Click “Upload Documents” and select your PDFs. Wait until processing finishes.
- Ask a question: Type a question in the chat window, for example, “How do I clean the filter?”.
- Check the answer: AnythingLLM searches your PDFs, finds the relevant section, and bases its answer on that.
- Follow up: Ask follow-up questions to clarify details. The context is preserved.
The entire workflow takes about 10 to 15 minutes, depending on the volume of documents and your hardware speed.
Common Pitfalls with AnythingLLM
AnythingLLM is straightforward to use, but a few things regularly prompt questions:
- Ollama can’t be found: If Ollama isn’t running or listening on a different port, AnythingLLM can’t connect. Check that Ollama is running and the address in settings is correct.
- Embeddings are slow: When uploading large documents for the first time, embedding creation takes a while, especially on CPU. This is normal and happens only once per document.
- Answers are inaccurate: If a workspace contains too many unrelated documents, the context becomes unwieldy. Split content across multiple workspaces.
- Documents aren’t found: AnythingLLM searches only the active workspace. If you ask in the wrong workspace, the model finds nothing. Check which workspace you’re in.
- Storage fills up: Embeddings and documents require space. With many large PDFs, storage needs grow quickly. Monitor your storage, especially in Docker deployment.
- Cloud providers send data outside: If you choose OpenAI or Anthropic as your provider, your documents are sent to the cloud service for processing. For local-only operation, use Ollama or LM Studio.
- Agent mode delivers unexpected results: Not every model supports all tools well. If the agent produces unusable answers, switch models or disable agent mode.
- Data loss after update: Without volume mapping in Docker, all workspaces disappear after a restart. The
-vparameter in your Docker command is important.
Hardware, Costs, and Security with AnythingLLM
Hardware: AnythingLLM itself demands minimal resources. The real load comes from the language model and embedding calculations. For smaller models up to 8 billion parameters, a standard machine with 16 GB RAM suffices. Larger models benefit from a GPU with adequate VRAM. As a rough guide: 8 GB VRAM for models up to 8 billion parameters, more for anything larger. Embedding computation runs fine on the CPU, though it will be slower there.
Costs: AnythingLLM is open source and free. You pay nothing for the software itself. Costs arise only from hardware upgrades if needed, and electricity. If you use cloud providers, API call expenses apply.
Security: Using local providers keeps your data on your machine. That’s the biggest advantage over cloud services. Still, keep these points in mind:
- Use cloud providers only if you accept that your documents leave your machine
- Don’t expose AnythingLLM in Docker mode to the open internet without authentication
- Keep the software up to date to patch security vulnerabilities
- Maintain clean workspace separation, especially with multiple users
Further Resources and Information on AnythingLLM
- AnythingLLM on GitHub
- Official Website
- AnythingLLM Documentation
- Ollama at BotServ.de
- LM Studio at BotServ.de
- Open WebUI at BotServ.de
- Local RAG at BotServ.de
- RAG Fundamentals at BotServ.de
FAQ: AnythingLLM - Common Questions
Is AnythingLLM free? Yes, AnythingLLM is open source and free. You pay nothing for the software itself, only for the hardware running it. Cloud providers may add extra costs.
Do I need Ollama for AnythingLLM? Not necessarily. AnythingLLM also supports LM Studio, OpenAI, Anthropic, and other providers. Ollama is the most common and easiest choice for local models.
Can I install AnythingLLM without Docker? Yes, the desktop app installs like any other program. Docker is only needed if you want to run AnythingLLM on a server.
Does AnythingLLM work on macOS?
Yes, there’s a native macOS version. Download the .dmg file and install it as you normally would.
Can I use multiple workspaces at once? Yes, you can create as many workspaces as you want and switch between them. Each workspace has its own documents and context.
What file formats does AnythingLLM support? PDF, DOCX, TXT, MD, CSV, and HTML are among the supported formats. Check the documentation for the complete list.
Do I need a GPU? For smaller models, the CPU is enough. For smooth work with larger models and faster embeddings, a GPU is recommended.
Can I connect AnythingLLM to cloud services? Yes, you can integrate OpenAI, Anthropic, and other cloud providers. Requests then go through the cloud service, and your data leaves your machine.
How do I update AnythingLLM?
The desktop app has an update function built in. With Docker, pull the latest image using docker pull mintplexlabs/anythingllm and restart the container. Your volume retains your data.
What’s the difference between workspaces and regular chats? Regular chats have no connection to your documents. Workspaces link chats to a specific document collection, so the model accesses them selectively.
Can I copy documents from one workspace to another? Yes, you can assign documents to multiple workspaces. This is useful when certain files are relevant in several contexts.
Is agent mode available with every model? Agent mode works with most models, but not all support every tool equally well. Local models with 8 billion parameters or more generally produce usable results.
Sources and Further Reading
- AnythingLLM Documentation: https://docs.anythingllm.com/
- AnythingLLM GitHub Repository: https://github.com/Mintplex-Labs/anything-llm
- Ollama Documentation: https://ollama.com/
- More articles at BotServ.de: Local AI Software, What is Local AI?, Local RAG


