BM25 Sparse Retrieval

Sparse Retrieval is based on exact keyword matching (bag of words). BM25 improves TF-IDF with two additions:

  • term frequency saturation — each subsequent occurrence gives a smaller and smaller increase (parameter k1);
  • document length normalization (parameter b).

Rare words weigh more (IDF).

Strength — exact names, technical codes, terms. Weakness — doesn’t understand synonyms (“kitty” won’t find “cat”).

Trainable variants (SPLADE, BGE-M3 sparse) add neural network weights to terms.

Related: Dense Embedding, Hybrid Search