FAISS: Meta’s Vector Search Library for Research
What This Article Covers
- What FAISS is and why it’s a library, not a database.
- Which index types FAISS offers and how to choose the right one.
- How to install and use FAISS with Python.
- How GPU support accelerates vector search.
- How FAISS compares to Chroma and Qdrant.
Introduction
FAISS stands for Facebook AI Similarity Search and is Meta’s library for efficient vector search. Unlike Chroma or Qdrant, FAISS is not a database but a C++ library with Python bindings. It stores no metadata, offers no REST API, and doesn’t handle persistence automatically. In exchange, FAISS is blazingly fast and scales to billions of vectors. This article shows you how to install FAISS, choose index types, and use it for RAG.
Why Do You Need FAISS?
When you need to search through massive numbers of vectors and performance is critical, FAISS ranks among the fastest available solutions. Meta develops FAISS for use in its own products and continuously optimizes it. GPU support is where FAISS excels: it can index and search vectors on the GPU, dramatically reducing query times.
FAISS suits research, prototyping, and applications where you want to manage infrastructure yourself. If you’re looking for a ready-made database with persistence, filtering, and an API, Chroma or Qdrant are better choices. But if you need maximum control and speed, FAISS is the right tool.
FAISS Explained
FAISS works with indexes that store and search vectors. An index is an object you populate with vectors and then query. FAISS offers different index types that make different tradeoffs between speed, accuracy, and memory. The main ones are:
- IndexFlatL2: Brute-force search using L2 distance. Exact but slow on large datasets.
- IndexFlatIP: Brute-force search using inner product. Exact, also slow on large datasets.
- IndexIVFFlat: Cluster-based search. Faster than Flat, slight accuracy loss.
- IndexIVFPQ: Cluster-based with product quantization. Very memory efficient, larger accuracy loss.
- IndexHNSWFlat: Graph-based search. Very fast and accurate, higher memory overhead.
You can also combine indexes, for example IVF with PQ for memory-efficient approximate search.
Who Is This Article For?
This article is for developers and researchers who need maximum performance from vector search and are willing to do more manual work. You should have Python experience and understand what embeddings are. If you’re new to vector databases, start with the vector databases overview and the article on embedding models.
Key Concepts
| Term | Meaning |
|---|---|
| Index | FAISS object that stores and searches vectors |
| Flat | Brute-force search without approximation |
| IVF | Inverted File, cluster-based approximate search |
| PQ | Product Quantization, vector compression for memory savings |
| HNSW | Hierarchical Navigable Small World, graph-based approximate search |
| GPU Index | Index that runs on the graphics card |
| Recall | Proportion of correctly found neighbors among all actual neighbors |
| nlist | Number of clusters in IVF |
| nprobe | Number of clusters searched during a query |
Installation
FAISS for CPU
pip install faiss-cpu
FAISS for GPU
For GPU support, you need an NVIDIA graphics card with CUDA:
pip install faiss-gpu
If you want to compile FAISS from source, check the FAISS GitHub Repository for instructions. For most use cases, the pip installation is sufficient.
Getting Started
Creating a Simple Index
import faiss
import numpy as np
dimension = 768
# Random vectors for the example
vectors = np.random.rand(1000, dimension).astype('float32')
# Flat index with L2 distance
index = faiss.IndexFlatL2(dimension)
index.add(vectors)
print(f"Vectors in index: {index.ntotal}")
# Search query
query = np.random.rand(1, dimension).astype('float32')
distances, indices = index.search(query, k=5)
print("Found indices:", indices)
print("Distances:", distances)
FAISS expects vectors as numpy arrays of type float32. Make sure your embeddings are in this format.
IVF Index with Training
IVF indexes must be trained before adding vectors. Training determines the cluster centers.
import faiss
import numpy as np
dimension = 768
nlist = 100 # Number of clusters
vectors = np.random.rand(10000, dimension).astype('float32')
# Quantizer as the base for IVF
quantizer = faiss.IndexFlatL2(dimension)
# Create IVF index
index = faiss.IndexIVFFlat(quantizer, dimension, nlist)
# Train with vectors
index.train(vectors)
# Add vectors
index.add(vectors)
# Set search parameters
index.nprobe = 10 # Number of clusters to search
# Search query
query = np.random.rand(1, dimension).astype('float32')
distances, indices = index.search(query, k=5)
print("Found indices:", indices)
The nprobe parameter controls the tradeoff between speed and accuracy. Higher nprobe means more accurate results but slower search.
HNSW Index
import faiss
import numpy as np
dimension = 768
vectors = np.random.rand(10000, dimension).astype('float32')
# HNSW index
index = faiss.IndexHNSWFlat(dimension, 32)
index.hnsw.efConstruction = 40
index.hnsw.efSearch = 16
index.add(vectors)
query = np.random.rand(1, dimension).astype('float32')
distances, indices = index.search(query, k=5)
print("Found indices:", indices)
HNSW requires no training. The parameter M (32 here) controls graph connections, efConstruction controls build quality, and efSearch controls search accuracy.
Using GPU
import faiss
import numpy as np
dimension = 768
vectors = np.random.rand(100000, dimension).astype('float32')
# Get GPU resources
res = faiss.StandardGpuResources()
# Create CPU index and move to GPU
cpu_index = faiss.IndexFlatL2(dimension)
gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index)
gpu_index.add(vectors)
query = np.random.rand(1, dimension).astype('float32')
distances, indices = gpu_index.search(query, k=5)
print("Found indices:", indices)
The number 0 in index_cpu_to_gpu is the GPU ID. With multiple graphics cards, you can choose which one to use.
Saving and Loading Indexes
# Save
faiss.write_index(index, "my_index.faiss")
# Load
index = faiss.read_index("my_index.faiss")
FAISS saves only the vectors and index structure. You must manage metadata like text or source information separately, for example in a JSON file or SQLite database.
Index Types Compared
| Index Type | Training | Accuracy | Speed | Memory | GPU |
|---|---|---|---|---|---|
| IndexFlatL2 | No | Exact | Slow | High | Yes |
| IndexFlatIP | No | Exact | Slow | High | Yes |
| IndexIVFFlat | Yes | High | Medium | Medium | Yes |
| IndexIVFPQ | Yes | Medium | Fast | Low | Yes |
| IndexHNSWFlat | No | High | Fast | High | No |
For small datasets up to about 10,000 vectors, IndexFlatL2 is sufficient. For medium-sized datasets, IVFFlat offers a solid balance. When dealing with large datasets and memory constraints, IVFPQ helps. HNSW is your best bet for fast and accurate search when memory is available.
Comparison with Chroma and Qdrant
| Property | FAISS | Chroma | Qdrant |
|---|---|---|---|
| Type | Library | Vector database | Vector database |
| Persistence | Manual (write_index) | Automatic | Automatic |
| Metadata | No | Yes | Yes |
| API | Python/C++ | Python | REST, gRPC, Python |
| Filtering | No | Simple | Advanced |
| GPU Support | Yes | No | No |
| Scalability | Very high | Small to medium | High |
| Learning curve | High | Low | Medium |
FAISS is the fastest solution for pure vector search, especially with GPU acceleration. However, it lacks features that dedicated databases provide: persistence, metadata, filtering, and an API. In practice, you’ll often wrap FAISS as an engine underneath your own layer that adds these capabilities. If you don’t want to build that abstraction, Chroma and Qdrant are better choices. For more on combining different search methods, see Hybrid Search.
Common Pitfalls
- Wrong data type: FAISS expects
float32arrays. Passingfloat64orintcauses errors or silent failures. Convert with.astype('float32'). - Skipping training on IVF: IVF indices must be trained before adding vectors. If you skip this step,
add()will fail. - Insufficient training data: IVF needs enough vectors for training. A good rule of thumb is at least 39 times as many vectors as
nlist. - nprobe too low: A low
nprobespeeds up search but reduces accuracy. Test different values and measure recall. - No persistence: FAISS doesn’t save indices automatically. If your process terminates, data is lost unless you call
write_index. - Managing metadata separately: FAISS returns only index positions. You must maintain a mapping table linking positions to text or sources.
- GPU memory overflow: Very large indices may exceed GPU memory. Either partition the index or use IVFPQ for compression.
- Index out of sync with data changes: New vectors are automatically integrated into HNSW. With IVF, quality degrades if you add many new vectors without retraining.
- Dimension mismatch: The index is created with a fixed dimensionality. Vectors with different dimensions cause errors.
Hardware, Costs, and Security
FAISS runs on any CPU-based machine. Small to medium datasets work fine on a standard desktop. For large datasets and GPU support, you’ll need an NVIDIA graphics card with sufficient VRAM (at least 8 GB, ideally 24 GB or more). GPU speedup can be 10x to 100x depending on index type and dataset size.
Costs: FAISS is open source under the MIT license and free. Hardware expenses, particularly GPUs, account for the main costs.
Security: FAISS itself has no network interface and no authentication. It runs as a library in your process. Security depends on how you protect the environment running FAISS. Since all data stays local, there’s no risk of exposure to external services. If you expose FAISS through a service, secure that service like any other application.
Further Reading
- FAISS GitHub Repository
- FAISS Documentation
- FAISS Tutorial
- Vector Database Overview
- RAG Fundamentals
- Local RAG
- Embedding Models
- Setting Up Chroma
- Setting Up Qdrant
- Hybrid Search
- Docker Basics
FAQ
Is FAISS a database?
No. FAISS is a library for vector search. It provides no persistence, metadata support, or API. You must add these features yourself.
Do I need a GPU for FAISS?
No. FAISS works on CPU. A GPU accelerates search on large datasets significantly but is not required.
Which graphics cards are supported?
FAISS supports NVIDIA graphics cards with CUDA. AMD cards are not officially supported.
How do I persist a FAISS index?
Use faiss.write_index(index, filepath) to save the index to a file. Load it with faiss.read_index(filepath).
Can I store metadata in FAISS?
No. FAISS stores only vectors. You must manage metadata in a separate database or file and link it via index positions.
Which index type is best?
It depends on your requirements. For small datasets, IndexFlatL2 is sufficient. For large datasets with speed requirements, HNSW or IVFFlat are recommended. For memory constraints, IVFPQ helps.
How do I choose nlist and nprobe for IVF?
A good starting point for nlist is the square root of the vector count. nprobe controls the speed-accuracy trade-off. Start with nprobe equal to nlist divided by 10 and experiment.
Can I combine FAISS with Ollama?
Yes. Ollama provides the language model, FAISS handles vector search. Your Python script glues them together and manages metadata.
Is FAISS free?
Yes, FAISS is open source under the MIT license and free. You pay only for hardware.
How do I measure index accuracy?
Compare results from an approximate index against an exact index like IndexFlatL2. The proportion of matching results is the recall.
Can I use FAISS in Docker?
Yes. Official Docker images with GPU support are available. You need to set up Docker with NVIDIA Container Toolkit. See Docker Basics for details.
What’s the difference between FAISS and Chroma?
Chroma is a complete vector database with persistence and metadata. FAISS is a pure search library that’s faster but requires more manual work.
Sources
- FAISS GitHub Repository
- FAISS Documentation
- FAISS Wiki with Tutorials
- HNSW Paper by Malkov and Yashunin
- Product Quantization Paper by Jegou et al.


