Fix ModuleNotFoundError: No module named 'faiss', FAISS_THROW_IF_NOT(assign_index->ntotal != k), and Index Out of Bounds Errors
ModuleNotFoundError: No module named 'faiss': runpython -m pip install faiss-cpuwith the same Python that runs your code. There is no PyPI package calledfaiss;pip install faissends inNo matching distribution found for faiss.Number of training points (50) should be at least as large as number of clusters (100): givetrain()at leastnlistvectors (39 × nlist to silence the warning), or lowernlist. PQ and RQ codebooks with 8 bits need at least 256 training vectors on their own.assign_index->ntotal == Kfailed: in faiss 1.15.1 this check sits in the residual quantizer (RQ) encoder, not in plain IVF training. It means an assignment index was handed in already holding a different number of vectors than the codebook size K. Pass an empty index, or let FAISS build its own.-1in the results:kis larger thanindex.ntotal. Capkand filter-1, becausedocs[-1]on a list silently returns the wrong document.
FAISS reports failures from C++ through the Python bindings with little context, and several of the worst mistakes here raise nothing at all. I reran every case in this post on the setup above; the outputs quoted below are copied from those runs, not typed from memory.
Reproduced
| What I ran | Exact result | Fix that worked |
|---|---|---|
import faiss in a fresh venv | ModuleNotFoundError: No module named 'faiss' | pip install faiss-cpu in that venv |
pip install faiss (pip 24.0 and 26.2.1) | ERROR: No matching distribution found for faiss | Install faiss-cpu; import name stays faiss |
python3 -m pip install faiss-cpu on system Python | error: externally-managed-environment | Create a venv, install there |
IndexIVFFlat, nlist=100, 50 training vectors | 'nx >= static_cast<idx_t>(k)' failed: Number of training points (50) should be at least as large as number of clusters (100) | 100 or more vectors trains; 3,900+ removes the warning |
beam_search_encode_step with an assign index holding 10 vectors, K=16 | Error: 'assign_index->ntotal == static_cast<idx_t>(K)' failed | Pass an assign index whose ntotal is 0 or K |
add() on an untrained IndexIVFFlat | Error: 'is_trained' failed | index.train(x) first |
search(k=10) on 5 vectors | No error. Indices [0, 2, 3, 1, 4, -1, -1, -1, -1, -1] | k = min(k, index.ntotal), drop -1 |
| 64-dim query on a 128-dim index | Bare AssertionError at assert d == self.d | Embed with the model that built the index |
1D query of shape (128,) | ValueError: not enough values to unpack (expected 2, got 1) | query.reshape(1, -1) |
| GPU index on GTX 1070, faiss-gpu 1.15.1 and faiss-gpu-cu12 1.14.1.post1 | CUDA error 209 no kernel image is available for execution on the device, process aborted | Use faiss-cpu on this card |
ModuleNotFoundError: No module named 'faiss'
This is the full error from a fresh Python 3.12 venv:
Traceback (most recent call last):
File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'faiss'
The package name on PyPI and the module name differ. You import faiss, but you install faiss-cpu. Installing the name you import does not work:
$ pip install faiss
ERROR: Could not find a version that satisfies the requirement faiss (from versions: none)
ERROR: No matching distribution found for faiss
The install that worked:
python -m venv .venv
.venv/bin/python -m pip install faiss-cpu numpy
.venv/bin/python -c "import faiss, numpy; print(faiss.__version__, numpy.__version__)"
# 1.15.1 2.5.3
If the package is installed and you still get the error, you are running a different Python than the one pip installed into. The cases I hit on this machine:
- Installed in a venv, ran with system Python.
faiss-cpuwas in/tmp/faisstest/base, andpython3 -c "import faiss"(which is/usr/bin/python3) still raisedModuleNotFoundError. Activate the venv or call its interpreter directly. - Tried to install into system Python instead. On Ubuntu 24.04,
python3 -m pip install faiss-cpustops witherror: externally-managed-environment. Do not reach for--break-system-packages; use a venv. - Jupyter kernel or IDE interpreter (not reproduced here, same cause): the kernel is another Python. Install from inside the notebook with
%pip install faiss-cpu, which targets the running kernel.
To see which interpreter is which, compare these two:
python -c "import sys; print(sys.executable)"
python -m pip show faiss-cpu # "Location:" must be under that interpreter's prefix
Always prefer python -m pip over a bare pip; it guarantees the package lands in the interpreter you just named. For GPU builds, see Error 4 before installing: on the GPU I tested, both pip GPU packages installed cleanly and then crashed at the first search.
Error 1: -1 Indices and IndexError from a k Larger Than ntotal
An index with 5 vectors, searched for 10 neighbors:
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
vectors = np.random.rand(5, d).astype('float32')
index.add(vectors)
query = np.random.rand(1, d).astype('float32')
distances, indices = index.search(query, k=10)
FAISS does not raise. On faiss 1.15.1 it returned:
indices [0, 2, 3, 1, 4, -1, -1, -1, -1, -1]
distances ... 3.4028234663852886e+38, 3.4028234663852886e+38
You asked for k=10 neighbors but the index holds 5 vectors, so the unfillable slots get -1 and the largest float32 value as distance. What happens next depends on your lookup code, and the worst case is the quiet one:
docs = ["doc0", ..., "doc4"] # Python list
[docs[i] for i in indices[0]]
# ['doc0', 'doc2', 'doc3', 'doc1', 'doc4', 'doc4', 'doc4', 'doc4', 'doc4', 'doc4']
id_to_doc = {0: "doc0", ..., 4: "doc4"} # dict
[id_to_doc[int(i)] for i in indices[0]]
# KeyError: -1
Negative indexing turns -1 into "the last document", so a list or NumPy array lookup returns wrong results with no error. A dict raises KeyError: -1, which at least tells you something is off.
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
vectors = np.random.rand(5, d).astype('float32')
index.add(vectors)
query = np.random.rand(1, d).astype('float32')
# Always cap k at the number of indexed vectors
k = min(10, index.ntotal)
distances, indices = index.search(query, k=k)
# Filter out -1 sentinels defensively even after capping
valid_mask = indices[0] != -1
valid_indices = indices[0][valid_mask]
valid_distances = distances[0][valid_mask]
print(f"Found {len(valid_indices)} neighbors: {valid_indices}")
Use index.ntotal to query how many vectors are in the index at any time. Cap k against it before every search, and still filter -1: an IVF index with a small nprobe can return fewer than k hits even when ntotal is large. With nlist=100, 1,000 vectors, nprobe=1 and k=50, four of five test queries got between 7 and 39 -1 slots.
Error 2: Dimension Mismatch Between Index and Query Vector
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
vectors = np.random.rand(10, d).astype('float32')
index.add(vectors)
# Query vector has wrong dimension
query = np.random.rand(1, 64).astype('float32')
distances, indices = index.search(query, k=3)
On faiss 1.15.1 the message is empty. The last lines of the traceback:
assert d == self.d
^^^^^^^^^^^
AssertionError
The index was created with d=128 and the query has 64 dimensions. The check is a plain Python assert in the FAISS wrapper, so it prints no numbers; that is why the explicit assert with a message below is worth adding. index.add() with the wrong width fails the same way. A 1D query fails differently, before the dimension check:
ValueError: not enough values to unpack (expected 2, got 1)
Common causes:
- You switched embedding models mid-pipeline (e.g., from
all-MiniLM-L6-v2at 384d totext-embedding-ada-002at 1536d) but reused an old index - You transposed a matrix:
shape (128, 1)instead of(1, 128) - You passed a 1D array instead of 2D:
shape (128,)instead of(1, 128)
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
vectors = np.random.rand(10, d).astype('float32')
index.add(vectors)
# Wrong: 1D array
raw_query = np.random.rand(d).astype('float32')
# Fix 1: reshape 1D to 2D
query = raw_query.reshape(1, -1)
# Fix 2: assert before searching to get a clear error message
assert query.shape[1] == index.d, (
f"Query dimension {query.shape[1]} does not match index dimension {index.d}. "
f"Re-embed your query with the same model used to build the index."
)
distances, indices = index.search(query, k=3)
print(indices)
To diagnose silently mismatched pipelines, store the embedding model name alongside the index:
import json, faiss
def save_index(index, path, embedding_model: str):
faiss.write_index(index, path + ".faiss")
with open(path + ".meta.json", "w") as f:
json.dump({"d": index.d, "embedding_model": embedding_model}, f)
def load_index(path, expected_model: str):
with open(path + ".meta.json") as f:
meta = json.load(f)
if meta["embedding_model"] != expected_model:
raise ValueError(
f"Index was built with '{meta['embedding_model']}' "
f"but you are using '{expected_model}'. Rebuild the index."
)
return faiss.read_index(path + ".faiss")
Error 3: Searching an Empty Index
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
# Forgot to add vectors
query = np.random.rand(1, d).astype('float32')
distances, indices = index.search(query, k=5)
print(indices) # [[-1 -1 -1 -1 -1]]
print(distances) # [[3.4028235e+38 3.4028235e+38 ...]] (float32 max)
There is no exception. FAISS returns all -1 indices, and downstream code fails later (or, with list lookups, returns the last document five times) with no obvious link to the empty index.
Searching an empty FAISS index is not an error at the FAISS level. index.ntotal == 0 and there are simply no candidates to return. This happens when:
- Ingestion failed silently upstream and no vectors were ever added
- You saved an index before anything was added to it; it reloads fine with
ntotal == 0 - You created a fresh index but the embedding step threw an exception that was caught and swallowed
import faiss
import numpy as np
d = 128
index = faiss.IndexFlatL2(d)
# Simulate population step (e.g., from a database or file)
vectors = np.random.rand(100, d).astype('float32')
index.add(vectors)
def safe_search(index, query: np.ndarray, k: int):
"""Search with guards against empty index and wrong k."""
if index.ntotal == 0:
raise RuntimeError(
"FAISS index is empty. Add vectors before searching. "
"Check your ingestion pipeline for silent failures."
)
if query.ndim == 1:
query = query.reshape(1, -1)
if query.dtype != np.float32:
query = query.astype('float32')
k = min(k, index.ntotal)
distances, indices = index.search(query, k)
# Remove sentinel -1 rows
mask = indices[0] != -1
return distances[0][mask], indices[0][mask]
query = np.random.rand(d).astype('float32')
dists, idxs = safe_search(index, query, k=5)
print(f"Top-{len(idxs)} results: {idxs}")
The same check belongs right after loading a persisted index from disk. A damaged file does raise: a file cut in half gave Error: 'ret == (size)' failed: read error in trunc.faiss: 25577 != 51200, and a 0-byte file gave 'ret == (1)' failed. An index that was saved empty loads without any error, though, so only the ntotal check catches it:
import faiss
index = faiss.read_index("my_index.faiss")
if index.ntotal == 0:
raise RuntimeError("Loaded index is empty; rebuild from source data.")
print(f"Loaded index with {index.ntotal} vectors of dimension {index.d}")
Error 4: faiss-gpu Installs, Then Crashes on the GPU
import faiss
import numpy as np
# Attempting GPU index without checking availability
res = faiss.StandardGpuResources()
d = 128
cpu_index = faiss.IndexFlatL2(d)
gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index)
With faiss-cpu installed, the first line fails with:
AttributeError: module 'faiss' has no attribute 'StandardGpuResources'
The CPU package has no GPU classes. For GPU you need a GPU build, and on pip there are two today. Both installed cleanly on Python 3.12:
pip install faiss-gpu # faiss-gpu 1.15.1 + nvidia-cublas-cu12 12.9, cuda-runtime 12.9
pip install faiss-gpu-cu12 # faiss-gpu-cu12 1.14.1.post1, same CUDA 12.9 runtime wheels
Install one or the other, never both: they ship the same faiss module and overwrite each other, and neither should share a venv with faiss-cpu. With either one, hasattr(faiss, "StandardGpuResources") was True and faiss.get_num_gpus() returned 1 for the GTX 1070. The first GPU search then killed the process:
Faiss assertion 'err__ == cudaSuccess' failed in void faiss::gpu::runL2Norm(...)
at /project/faiss/gpu/impl/L2Norm.cu:257; details: CUDA error 209 no kernel image
is available for execution on the device
Aborted (core dumped)
Shell exit code 134 (SIGABRT), the same for both packages. The wheels are built against CUDA 12.9 and ship no kernels for the GTX 1070's Pascal architecture (compute capability 6.1), so the card is detected but cannot run anything. Two consequences:
get_num_gpus() > 0does not prove the GPU works. It only proves CUDA sees a device.- This is a C++ abort, not a Python exception.
try/exceptnever runs, so an in-process fallback cannot catch it.
On a Pascal card, use faiss-cpu (or a build compiled for your architecture). On a newer card, run a one-query smoke test in a subprocess before trusting the GPU path:
import subprocess, sys
SMOKE = (
"import faiss, numpy as np;"
"x = np.random.rand(100, 16).astype('float32');"
"i = faiss.index_cpu_to_gpu(faiss.StandardGpuResources(), 0, faiss.IndexFlatL2(16));"
"i.add(x); i.search(x[:1], 1)"
)
def faiss_gpu_works() -> bool:
r = subprocess.run([sys.executable, "-c", SMOKE], capture_output=True)
return r.returncode == 0 # -6 (SIGABRT) on the GTX 1070 run above
The fallback below then only takes the GPU path when faiss_gpu_works() is true:
import faiss
import numpy as np
d = 128
def build_index(vectors: np.ndarray, use_gpu: bool = False):
vectors = vectors.astype('float32')
d = vectors.shape[1]
cpu_index = faiss.IndexFlatL2(d)
cpu_index.add(vectors)
if not use_gpu:
return cpu_index
# GPU package installed?
if not hasattr(faiss, 'StandardGpuResources'):
print("faiss-gpu not installed, falling back to CPU.")
return cpu_index
ngpus = faiss.get_num_gpus()
if ngpus == 0 or not faiss_gpu_works():
print("No usable CUDA GPU, falling back to CPU.")
return cpu_index
res = faiss.StandardGpuResources()
gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index)
print(f"Index moved to GPU 0 ({ngpus} GPU(s) available)")
return gpu_index
# Usage
vectors = np.random.rand(1000, d).astype('float32')
index = build_index(vectors, use_gpu=True) # CPU index if the GPU check fails
query = np.random.rand(1, d).astype('float32')
distances, indices = index.search(query, k=5)
print(indices)
To move a GPU index back to CPU (e.g., for persistence, since FAISS cannot save GPU indexes directly):
import faiss
# gpu_index is a GpuIndex object
cpu_index = faiss.index_gpu_to_cpu(gpu_index)
faiss.write_index(cpu_index, "index.faiss")
# On reload, move back to GPU
loaded = faiss.read_index("index.faiss")
res = faiss.StandardGpuResources()
gpu_index = faiss.index_cpu_to_gpu(res, 0, loaded)
Error 5: FAISS_THROW_IF_NOT(assign_index->ntotal != k) and Training Errors
People search for this assertion as assign_index->ntotal != k. The source actually checks assign_index->ntotal == K and throws when it does not hold (faiss-cpu 1.8.0 spells it assign_index->ntotal == K, 1.15.1 spells it assign_index->ntotal == static_cast<idx_t>(K)). An earlier version of this post blamed it on too few training vectors for an IVF index. That was wrong for current FAISS: too few training vectors raises a different, readable error, shown next.
Too few training vectors for nlist
import faiss
import numpy as np
d = 128
nlist = 100
quantizer = faiss.IndexFlatL2(d)
index = faiss.IndexIVFFlat(quantizer, d, nlist)
# Fewer training vectors than clusters
training_vectors = np.random.rand(50, d).astype('float32')
index.train(training_vectors)
RuntimeError: Error in void faiss::Clustering::train_encoded(...) at
/project/faiss/Clustering.cpp:66: Error: 'nx >= static_cast<idx_t>(k)' failed:
Number of training points (50) should be at least as large as number of clusters (100)
Training an IVF index runs k-means with k = nlist centroids, and k-means cannot place 100 centroids on 50 points. The hard limit is n >= nlist: with exactly 100 points the index trained without error and is_trained was True. Below 39 × nlist you only get a warning:
WARNING clustering 10240 points to 1024 centroids: please provide at least 39936 training points
Watch for the same error where you do not expect k-means. IVF16,PQ4x8 with 200 training vectors failed with Number of training points (200) should be at least as large as number of clusters (256): the 16 IVF lists were fine, but each 8-bit PQ sub-quantizer has 256 centroids. RQ2x8 and PQ4x8 on 100 vectors failed the same way.
Where assign_index->ntotal == K actually fires
In faiss-cpu 1.15.1 the check is in faiss::beam_search_encode_step (impl/residual_quantizer_encode_steps.cpp), the beam search used by the residual quantizer (ResidualQuantizer, RQ index factory strings, IVF with RQ). The function can take an "assign index" used to find the nearest codebook entries. If that index is empty, FAISS adds the K codebook vectors to it. If it already holds vectors, FAISS assumes those are the codebook and requires exactly K of them. I triggered it by calling the function directly with an assign index holding 10 vectors while the codebook had K=16:
RuntimeError: Error in void faiss::beam_search_encode_step(...) at
/project/faiss/impl/residual_quantizer_encode_steps.cpp:74:
Error: 'assign_index->ntotal == static_cast<idx_t>(K)' failed
Ordinary index_factory use did not hit it: RQ1x4_1x6, RQ1x6_1x4 and RQ1x4_1x6_1x8 (codebooks of different sizes per stage) all trained and encoded on 5,000 vectors. So if you see this assertion, look for code that supplies its own assign index or assign-index factory to a residual quantizer, or reuses one across codebooks of different sizes. The fix is to hand FAISS an empty index (assign_index.reset()) or one holding exactly the K centroids for that step.
Adding to an untrained index
This one raises, contrary to what this post used to say:
RuntimeError: Error in virtual void faiss::IndexIVFFlat::add_core(...) at
/project/faiss/IndexIVFFlat.cpp:66: Error: 'is_trained' failed
A training guard that checks the real limits
import faiss
import numpy as np
RECOMMENDED_MULTIPLIER = 39 # below this FAISS warns; below 1 it raises
def safe_train(index, training_vectors: np.ndarray, nlist: int):
n = training_vectors.shape[0]
if n < nlist:
raise ValueError(
f"Only {n} training vectors for nlist={nlist}. "
f"FAISS needs at least {nlist}; lower nlist or gather more data."
)
if n < nlist * RECOMMENDED_MULTIPLIER:
print(f"warning: {n} < {nlist * RECOMMENDED_MULTIPLIER} recommended training vectors")
index.train(training_vectors)
assert index.is_trained
d, nlist = 128, 100
index = faiss.IndexIVFFlat(faiss.IndexFlatL2(d), d, nlist)
safe_train(index, np.random.rand(5000, d).astype('float32'), nlist)
index.add(np.random.rand(5000, d).astype('float32'))
Does the 39x rule matter for recall? nprobe matters more
I measured it on 200,000 synthetic 128-dim vectors (Gaussian blobs around 500 centers), IndexIVFFlat with nlist=1024, 1,000 queries, recall@10 against IndexFlatL2, 4 CPU threads:
training vectors nprobe=1 nprobe=8 nprobe=32 nprobe=128
1,024 (1x nlist) 0.661 0.996 1.000 1.000
10,240 (10x nlist) 0.855 1.000 1.000 1.000
39,936 (39x nlist) 0.794 1.000 1.000 1.000
On this data the training-set size moved recall at nprobe=1, not in a clean direction, and stopped mattering by nprobe=8. Treat the 39x figure as FAISS's safe default, not a threshold where results break. Timing was noisy on this shared machine: the flat search took about 300 ms for 1,000 queries, nprobe=8 took 48 to 128 ms, and nprobe=128 took 680 to 1,090 ms, which is slower than brute force. Synthetic blobs are easier than real embeddings, so tune nprobe on your own data against a flat baseline. Setting nprobe above nlist is not an error: nprobe=64 on an nlist=16 index searched normally and scanned every list.
Quick Diagnostic Checklist
Print these before a failing search():
import faiss
import numpy as np
def diagnose_index(index, query: np.ndarray, k: int):
print(f"Index type: {type(index).__name__}")
print(f"Index d: {index.d}")
print(f"Index ntotal: {index.ntotal}")
print(f"Query shape: {query.shape}")
print(f"Query dtype: {query.dtype}")
print(f"k requested: {k}")
print(f"k effective: {min(k, index.ntotal)}")
if index.ntotal == 0:
print("WARNING: index is empty")
if query.ndim == 1:
print("WARNING: query is 1D, needs reshape(1, -1)")
if query.dtype != np.float32:
print(f"NOTE: query dtype is {query.dtype}; the wrapper converts it to float32 on every call")
if query.shape[-1] != index.d:
print(f"ERROR: dimension mismatch, query {query.shape[-1]} vs index {index.d}")
One wrapper fixes most of this
Three of the cases above give you nothing useful to debug: k above ntotal and an empty index raise nothing, and a dimension mismatch raises an AssertionError with no message. Call diagnose_index() (or at least check ntotal, d and the query shape) before search(), and the silent ones turn into readable messages.