Embedding Quantization for Production Vector Search in October 2026: Int8, Binary, and Matryoshka Truncation Without Losing Recall
Vector search gets expensive in a quiet way. The first million chunks fit in memory without anyone noticing. Then the corpus grows, you add a second embedding model for a new language, and suddenly the vector index is the largest line item on your retrieval bill. Before you buy a bigger box, look at how you store the vectors themselves. Embedding quantization lets you shrink the index by 4x to 32x, and with a rescoring step most teams keep retrieval quality close to where it was.
This guide explains the three techniques that matter in practice (scalar int8 quantization, binary quantization, and Matryoshka-style dimension truncation), how to combine them with rescoring, and how to measure whether you actually kept your recall.
Why the index gets big
Most embedding models output float32 vectors. Each dimension costs 4 bytes. The math is simple and unforgiving:
- 1 million vectors at 1,024 dimensions in float32 is about 4.1 GB of raw vector data, before any graph or index overhead.
- The same vectors in int8 take about 1 GB.
- The same vectors as binary (1 bit per dimension) take 128 bytes each, or about 128 MB total.
Approximate nearest neighbor indexes such as HNSW want vectors in RAM to stay fast. So the storage format directly decides how many machines you need and how much a cold start costs.
Option 1: Scalar (int8) quantization
Scalar quantization maps each float dimension into a small integer range, usually 256 buckets. You pick a minimum and maximum per dimension from a calibration sample of your real embeddings, then linearly map values into that range.
import numpy as np
def fit_int8(calibration: np.ndarray):
# calibration: (n, d) float32 embeddings from your real corpus
lo = np.percentile(calibration, 0.5, axis=0)
hi = np.percentile(calibration, 99.5, axis=0)
return lo, hi
def to_int8(x: np.ndarray, lo, hi):
scaled = (np.clip(x, lo, hi) - lo) / (hi - lo + 1e-12)
return (scaled * 255 - 128).round().astype(np.int8)
Using percentiles instead of the absolute min and max keeps a few outlier values from wasting most of the range. Calibrate on a few thousand to tens of thousands of embeddings drawn from the same distribution you will index, not from a toy sample.
When to use it: int8 is the safe default. It cuts memory by 4x, it is supported by most vector databases and libraries, and the quality loss is usually small. If you only make one change, make this one.
Option 2: Binary quantization
Binary quantization keeps only the sign of each dimension: 1 if the value is above zero, 0 otherwise. Distance becomes Hamming distance, which a CPU computes with XOR and a population count. That is extremely fast and extremely compact.

def to_binary(x: np.ndarray) -> np.ndarray:
# (n, d) float32 -> (n, d/8) uint8, one bit per dimension
return np.packbits(x > 0, axis=1)
def hamming(query_bits: np.ndarray, doc_bits: np.ndarray) -> np.ndarray:
xor = np.bitwise_xor(doc_bits, query_bits)
return np.unpackbits(xor, axis=1).sum(axis=1)
The catch is that binary quantization throws away a lot of information, so it works much better with some embedding models than others. Models with higher dimensions and values centered around zero tend to survive it better. Do not assume; test your specific model.
When to use it: large corpora where memory or latency is the main constraint, and only together with rescoring (below). Binary search alone is a first-pass filter, not a final ranking.
Option 3: Matryoshka truncation
Some embedding models are trained with Matryoshka Representation Learning, which packs the most important information into the first dimensions. With those models you can keep, for example, the first 256 of 1,024 dimensions and still get useful similarity scores. That is a 4x reduction before any quantization.
def truncate(x: np.ndarray, dims: int) -> np.ndarray:
t = x[:, :dims]
return t / np.linalg.norm(t, axis=1, keepdims=True)
Two rules matter here. First, re-normalize after truncating if you use cosine or dot-product similarity. Second, only truncate models whose documentation says they support it. Chopping dimensions off a model that was not trained this way usually degrades results badly.
Truncation and quantization stack. A Matryoshka model truncated to 512 dimensions and stored as int8 is 8x smaller than the original float32 index.
The step that makes it work: rescoring
Quantized vectors are good at finding the right neighborhood and worse at ordering the final few results. The standard fix is a two-stage search:
- Oversample. If the application needs the top 10, retrieve the top 40 to 100 candidates from the quantized index.
- Rescore. Recompute similarity for just those candidates using higher-precision vectors, then keep the best 10.
The higher-precision vectors do not need to live in RAM. Keep the float32 (or float16) copies on local SSD or in object storage keyed by document ID, and fetch only the candidates. A cheaper variant is to compare the float32 query against int8 document vectors, which improves ordering without any extra fetch.
def search(query_f32, bin_index, f32_store, k=10, oversample=8):
q_bits = to_binary(query_f32[None, :])
dist = hamming(q_bits, bin_index)
cand = np.argsort(dist)[: k * oversample]
full = f32_store[cand] # fetch only candidates
scores = full @ query_f32
return cand[np.argsort(-scores)[:k]]
Many vector stores implement this pattern for you. Qdrant offers scalar, binary, and product quantization with optional rescoring from original vectors. FAISS provides scalar-quantizer and binary index types. pgvector supports half-precision and bit vector types, so you can index binary vectors and rescore against the full vectors in the same SQL query. Check the docs for the version you run, since options and defaults change.
How to measure what you lost
Never ship quantization on vibes. Build a small, honest benchmark:
- Collect real queries. Pull a few hundred to a few thousand queries from production logs, with personal data removed.
- Compute ground truth. Run exact (brute-force) float32 search for each query over a representative slice of the corpus. Save the true top-k.
- Score each configuration. For int8, binary plus rescoring, truncated dimensions, and combinations, measure recall@k against the ground truth, plus p50 and p99 latency and memory.
- Check end-to-end quality. Retrieval recall is a proxy. Run your existing RAG evaluation set through the pipeline too, because a small recall drop on easy queries can hide a bigger drop on the hard ones that matter.
Then pick the cheapest configuration that stays inside your quality budget. Write that budget down before you run the numbers, for example "recall@10 must stay within two points of float32," so the decision is not argued after the fact.
A practical rollout plan
- Start with int8 plus rescoring. It is the lowest-risk change and often enough on its own.
- Try binary only when memory still hurts. Pair it with generous oversampling and measure again.
- Use truncation only with Matryoshka-trained models. Test several cutoffs; the right one depends on your data.
- Keep the full-precision vectors. They make rescoring possible and let you rebuild the index if you change quantization settings later.
- Re-calibrate when the model changes. Int8 ranges are fitted to one model's output distribution. A new embedding model means new calibration and a new benchmark.
- Shadow before switching. Run the quantized index alongside the current one, compare result overlap on live traffic, then cut over.
Common mistakes
- Calibrating on the wrong data. Fitting int8 ranges on generic text and indexing legal contracts gives poor bucket usage.
- Skipping rescoring with binary. Hamming distance alone ties many candidates and orders them poorly.
- Forgetting to quantize the query the same way. The query and documents must use the same transform, or distances are meaningless.
- Measuring only average recall. Look at per-query results; quantization tends to hurt long-tail, specific queries first.
Bottom line
Embedding quantization is one of the few retrieval optimizations that saves real money without changing your model or your prompts. Start with int8, add rescoring, measure recall against exact search on real queries, and move to binary or truncation only when the numbers say you can afford it. Done carefully, a vector index that needed several large machines can often fit comfortably on one.
Comments
Post a Comment