BM25 Sparse Retrieval
Sparse Retrieval is based on exact keyword matching (bag of words). BM25 improves TF-IDF with two additions:
- term frequency saturation — each subsequent occurrence gives a smaller and smaller increase (parameter k1);
- document length normalization (parameter b).
Rare words weigh more (IDF).
Strength — exact names, technical codes, terms. Weakness — doesn’t understand synonyms (“kitty” won’t find “cat”).
Trainable variants (SPLADE, BGE-M3 sparse) add neural network weights to terms.
Related: Dense Embedding, Hybrid Search