Milvus: Vector Database for Large-Scale Operations
What this article covers
- What Milvus is and why it’s ideal for handling large datasets
- How to install and run Milvus locally using Docker
- Which index types are available and how to choose the right one for your use case
- How to store, search, and filter data with pymilvus
- How Milvus compares to Chroma and Qdrant, and when it’s the right choice
Introduction
You’ve discovered local RAG and may have already experimented with Chroma or Qdrant. Now you’re facing a new challenge: what happens when your document collection keeps growing? When you’re dealing with millions of vectors instead of thousands?
That’s where Milvus comes in. Milvus is an open-source vector database engineered specifically for large-scale operations and high performance. It supports multiple index types, can be run in a distributed setup, and handles data volumes that push other databases to their limits.
This article walks you through installation, index selection, data storage, and search queries. If you haven’t read up on RAG fundamentals yet, start there. Familiarity with embedding models is also recommended.
Why do you need Milvus?
Imagine running a RAG application for a large enterprise. Your knowledge base contains hundreds of thousands of documents, with new ones arriving daily. Your vector database must not only store all these vectors but also find the most similar ones in milliseconds.
Chroma hits a wall with datasets at this scale. Qdrant goes significantly further, though it’s primarily optimized for single-node setups. Milvus, by contrast, is built from the ground up for distributed architecture. You can work with multiple nodes, distribute data across machines, and still get fast search results.
Another advantage is the range of index types available. Milvus supports IVF, HNSW, DiskANN, and several others. Each index type has its own strengths: some build faster, others use less memory, and some are optimized for very large datasets. You choose the index that fits your use case rather than being locked into a one-size-fits-all solution.
For local setups, you don’t necessarily need to build a distributed cluster. Milvus Lite and Milvus Standalone work perfectly for most projects. But if your project grows, you can upgrade to Milvus Distributed without switching databases.
Milvus at a glance
Milvus stores vectors in collections, similar to tables in a relational database. Each collection has a schema that defines the fields and their data types. One of these fields is the primary vector used for similarity search.
When you run a search query, Milvus converts the query into a vector and searches through the collection’s index. The index is a data structure that speeds up search by avoiding direct comparison of every vector, instead using intelligent grouping and pre-filtering. Your choice of index affects search speed, memory usage, and result accuracy.
Milvus operates in three modes: Milvus Lite for testing directly in Python, Milvus Standalone as a single Docker container, and Milvus Distributed as a multi-node cluster. For local RAG projects, Milvus Standalone is the practical choice because it’s simple to set up while offering all essential features.
Who this article is for
This article is aimed at developers who already have basic familiarity with RAG and vector databases. You should understand what local AI is and how RAG works. If you’ve tried Chroma or Qdrant and are looking for a solution that scales to larger datasets, this is your next step. You’ll need basic Python knowledge and should be comfortable with Docker.
Key terms
| Term | Definition |
|---|---|
| Collection | A container for vectors and metadata, similar to a table. |
| Schema | The definition of fields, data types, and vector parameters for a collection. |
| Index | A data structure that accelerates vector search, such as IVF or HNSW. |
| IVF | Inverted File Index, groups vectors into clusters to speed up search. |
| HNSW | Hierarchical Navigable Small World, a graph-based index for very fast search. |
| DiskANN | An index that keeps data on disk and is optimized for very large datasets. |
| Metric Type | The distance metric used for similarity search, such as Cosine, L2, or IP. |
| Partition | A division of a collection to restrict searches to specific subsets. |
| Segment | The physical storage unit in Milvus where data is organized. |
Installation: Milvus Standalone with Docker
Milvus Standalone is the best choice for local projects. It runs in a single Docker container and requires etcd and MinIO as dependencies. The simplest approach is to use the official docker-compose.yml file from Milvus.
Download the file and start the service:
wget https://github.com/milvus-io/milvus/releases/download/v2.4.0/milvus-standalone-docker-compose.yml -O docker-compose.yml
docker compose up -d
After startup, verify that all containers are running:
docker compose ps
You should see three running containers: milvus, etcd, and minio. Milvus is accessible on port 19530. If you haven’t installed Docker yet, check out the Docker fundamentals.
Alternatively, you can use Milvus Lite, which runs directly in Python without Docker. This is ideal for testing but not suitable for production:
pip install pymilvus
With Milvus Lite, you don’t need a server. Instead, you use Milvus as an embedded library. For this article, we’ll use Milvus Standalone with Docker because it more closely resembles a real production system.
Install the Python client
The Python client for Milvus is pymilvus. Install it with pip:
pip install pymilvus
Then connect to your local Milvus instance in Python:
from pymilvus import MilvusClient
client = MilvusClient(uri="http://localhost:19530")
print(client.list_collections()) # Lists all existing collections
If you’re using Milvus Lite, provide a file path instead of a URI:
client = MilvusClient(uri="./milvus.db")
Create a collection and define a schema
Before storing data, you create a collection. In Milvus, you explicitly define the schema fields. Each collection needs at least a primary key field and a vector field.
from pymilvus import MilvusClient, DataType
client.create_collection(
collection_name="articles",
fields=[
{"name": "id", "dtype": DataType.INT64, "is_primary": True},
{"name": "title", "dtype": DataType.VARCHAR, "max_length": 200},
{"name": "content", "dtype": DataType.VARCHAR, "max_length": 2000},
{"name": "source", "dtype": DataType.VARCHAR, "max_length": 100},
{"name": "vector", "dtype": DataType.FLOAT_VECTOR, "dim": 768}
]
)
The vector field vector has a dimension of 768, which corresponds to the output of intfloat/multilingual-e5-large. If you use a different embedding model, adjust the dimension accordingly.
Creating an Index
Once you’ve created a collection, set up an index on the vector field. The index is crucial for search performance. Without one, Milvus would scan every vector individually, which becomes impractical with large datasets.
client.create_index(
collection_name="articles",
field_name="vector",
index_type="IVF_FLAT",
metric_type="COSINE",
params={"nlist": 128}
)
The nlist parameter controls how many clusters Milvus groups vectors into. A higher value produces finer clusters and more accurate results, but requires more memory. A good starting point is 128 to 1024, depending on your data size.
Comparing Index Types
Milvus offers several index types. Here’s an overview of the most important ones:
IVF_FLAT: Groups vectors into clusters and searches only the relevant ones. Easy to set up and suitable for medium-sized datasets. Memory usage is moderate, and accuracy is high.
IVF_SQ8: Similar to IVF_FLAT, but uses quantized vectors. Reduces memory footprint to roughly one quarter with slightly lower accuracy. Choose this when storage space is limited.
HNSW: A graph-based index with multiple hierarchy levels. Delivers very fast search and high accuracy, but requires more memory and takes longer to build. Ideal when search speed matters more than storage efficiency.
DiskANN: Stores the index on disk rather than in RAM. Perfect for very large datasets that don’t fit in memory. Search is slightly slower than HNSW, but storage efficiency is unbeatable.
For most local RAG projects, HNSW is the best choice because it delivers the fastest search performance and datasets usually fit in RAM. As your data grows beyond available memory, switch to DiskANN.
Inserting Data
Now add data to the collection. You pass vectors together with metadata fields.
data = [
{
"id": 1,
"title": "Was ist lokale KI?",
"content": "Lokale KI läuft auf dem eigenen Rechner ohne Cloud-Dienste.",
"source": "blog",
"vector": [0.12, 0.34, 0.56] + [0.0] * 765 # Padded to 768 dimensions
},
{
"id": 2,
"title": "RAG erklärt",
"content": "RAG verbindet Dokumentensuche mit Sprachmodellen.",
"source": "blog",
"vector": [0.23, 0.45, 0.67] + [0.0] * 765
}
]
client.insert(collection_name="articles", data=data)
In practice, you generate vectors using a real embedding model. The vectors shown above are placeholders. What matters is that each vector has exactly the dimension defined in your schema.
Running Searches
After storing data and creating an index, you can query it. Milvus first loads the collection into memory:
client.load_collection(collection_name="articles")
results = client.search(
collection_name="articles",
data=[[0.15, 0.40, 0.60] + [0.0] * 765], # Search vector
limit=3,
output_fields=["title", "content", "source"]
)
for hits in results:
for hit in hits:
print(hit["entity"]["title"], hit["distance"])
The limit parameter specifies how many results to return. output_fields controls which fields appear in results. The distance shows how similar the retrieved vector is to your query vector.
Filtering
Milvus supports filter expressions to narrow search results. This is essential when you want to search only specific sources or categories.
results = client.search(
collection_name="articles",
data=[[0.15, 0.40, 0.60] + [0.0] * 765],
limit=3,
filter='source == "blog"',
output_fields=["title", "content", "source"]
)
for hits in results:
for hit in hits:
print(hit["entity"]["title"], hit["entity"]["source"])
Filter expressions in Milvus use SQL-like syntax. You can combine comparisons, logical operators, and range queries. This is more flexible than simple metadata filters in other databases.
Milvus vs. Chroma vs. Qdrant
| Feature | Milvus | Chroma | Qdrant |
|---|---|---|---|
| Scalability | Very high, distributed possible | Small to medium | High, single node |
| Index types | IVF, HNSW, DiskANN and more | HNSW (internal) | HNSW |
| Architecture | Distributed, with etcd and MinIO | Embedded or server | Server, single process |
| Resource overhead | Medium to high | Low | Low to medium |
| Ease of entry | Medium | Very low | Low |
| Filtering | SQL-like expressions | Simple | Highly expressive |
| Performance | Very high with large datasets | Good with small datasets | Very high |
| Setup effort | Higher, multiple containers | Minimal | Low |
Milvus shines when managing large datasets and you need a database that scales with you. Chroma is ideal for quick starts, and Qdrant is a strong all-rounder. Milvus pays off when you plan to work with millions of vectors long-term.
Common Pitfalls
- Forgetting to load the collection: Milvus requires you to load a collection into memory (
load_collection) before searching. Skip this step and you’ll get an error. After inserting new data, you don’t need to reload, but you do after index changes. - Mismatched vector dimensions: The schema dimension must match your embedding model exactly. A wrong value causes insertion errors. Check your model’s dimension before creating the schema.
- Creating the index before data: Build the index before inserting large amounts of data. If you insert first and index later, index construction takes much longer.
- Mixing up metric types: If your embedding model is optimized for cosine distance but you use L2, you’ll get poor search results. Verify which metric your model recommends.
- Missing etcd and MinIO: Milvus Standalone needs etcd and MinIO to function. Starting just the Milvus container won’t work. Always use the official
docker-compose.ymlfile. - Skipping persistence configuration: By default, Milvus stores data in volumes defined in
docker-compose.yml. If you customize this file, make sure volumes persist. - Requesting too many results: On large collections, high
limitvalues degrade performance. Start small and increase only when needed. - Ignoring partitions: Very large collections benefit from partitions, which scope searches to subsets. This significantly improves performance but is often overlooked.
Hardware, Costs, and Security
Milvus Standalone needs at least 8 GB RAM to run reasonably. For larger datasets, 16 GB or more is recommended. Disk storage depends on data size and index type. HNSW requires more space than IVF_SQ8, and DiskANN uses disk more efficiently than RAM-based indexes.
Milvus is open source and free. The community edition covers all features you need for local RAG projects. There’s also a commercial cloud version called Zilliz Cloud, which isn’t relevant for local setups.
Security works the same as with other local databases. As long as Milvus is only accessible within your network, your data is protected. If you expose it over the network, use a reverse proxy with TLS and enable authentication. Since all data stays local, that’s a major advantage over cloud-based vector databases, as explained in the article on local AI vs. cloud AI.
Further Reading
- Vector Database Overview - All vector database articles
- Local RAG - Overview of all RAG articles
- RAG Fundamentals - How RAG works
- Embedding Models - Generating vectors from text
- Hybrid Search - Combining vector and keyword search
- Chroma - Simple vector database
- Qdrant - Scalable vector database
- Docker Fundamentals - Understanding container technology
FAQ
Do I need to build a distributed cluster to use Milvus?
No. Milvus Standalone is sufficient for most local projects. Distributed mode is only necessary if your data volume grows so large that a single node can’t handle it.
What’s the difference between Milvus Lite and Milvus Standalone?
Milvus Lite runs directly in Python without Docker and works well for testing. Milvus Standalone runs as a Docker setup with etcd and MinIO, bringing you closer to a production environment. For serious projects, use Milvus Standalone.
Is Milvus free?
Yes, the open-source version is free. There’s a commercial cloud offering called Zilliz Cloud, which isn’t needed for local setups.
How many vectors can Milvus handle locally?
That depends on your RAM and disk space. Milvus is designed to work with millions of vectors. With DiskANN, you can even process datasets that don’t fit in RAM.
Can I connect Milvus with Ollama?
Yes. Ollama provides the language model, Milvus the vector database. Your Python script bridges both: Milvus retrieves relevant documents, Ollama generates the response.
Which index type should I choose?
For most local RAG projects, HNSW is the best choice because it delivers the fastest search results. If RAM is limited, switch to DiskANN. For small datasets, IVF_FLAT is sufficient.
Do I need Docker for Milvus?
For Milvus Standalone, yes, since etcd and MinIO run as additional containers. For testing, you can use Milvus Lite without Docker.
Does Milvus support hybrid search?
Yes, starting from version 2.4, Milvus supports hybrid search with sparse vectors. You can combine dense and sparse vectors to unite semantic and keyword-based search. Learn more in the Hybrid Search article.
Can I run Milvus on a NAS?
Yes, with Docker. Ensure you have sufficient RAM and fast storage. Milvus benefits greatly from SSD storage, especially with HNSW indices.
What’s the difference between Milvus and Qdrant?
Qdrant is leaner and simpler to set up, ideal for medium-sized datasets. Milvus is more complex to configure but better suited for very large datasets and distributed architectures.
How much RAM do I need for Milvus?
At least 8 GB for Milvus Standalone. For larger datasets with HNSW, 16 GB or more is recommended. With DiskANN, you can reduce RAM requirements because the index lives on disk.
Sources
- Milvus official website: milvus.io
- Milvus documentation: milvus.io/docs
- pymilvus on GitHub: github.com/milvus-io/pymilvus
- Milvus GitHub repository: github.com/milvus-io/milvus
- Article on Vector Databases in this series
- Article on Hybrid Search in this series


