Skip to article
The Evidence Path
Reader Systems Research Notebooks Python PDF

RAG glossary

17 of 18 · 3,106 words

Terms are defined in the sense used by this repository. Similar names can have different meanings across papers and products; the surrounding corpus, metric, and implementation still need to be specified.

A

Abstention — Choosing not to provide a factual answer when evidence, confidence, authority, or policy is insufficient. Measure error among answered queries and missed useful answers separately.

ACL (access-control list) — Principals or groups allowed to access an object. RAG must propagate ACLs to chunks, embeddings, graph edges, caches, and logs and enforce them before content crosses a component boundary.

Adaptive RAG — A system that routes among retrieval strategies, sources, or budgets based on the request/state rather than always using one fixed path.

Agentic RAG — A stateful policy that plans and executes multiple retrieval, read, tool, verify, or answer actions under a stopping rule and hard budget.

ANN (approximate nearest-neighbor search) — Retrieves likely nearest vectors without exhaustive comparison, trading recall for latency/memory.

ANN recall — Overlap between approximate and exact nearest-neighbor top-k. It is an index diagnostic, not evidence relevance or answer correctness.

Answer relevance — How directly an answer addresses the question. A relevant answer can still be false or unsupported.

Answerability — Whether available evidence is sufficient to answer a question under the declared policy.

ASQA — Ambiguous/long-form QA benchmark often used for attributed generation and citation evaluation.

Attribution — Mapping generated claims to evidence actually supporting them.

Authority — Source appropriateness or standing for a claim, distinct from semantic relevance and entailment.

B

BGE — Family of general/multilingual embedding and reranking models. Exact model, revision, instructions, dimension, and benchmark slice must be stated.

Bitemporal data — Stores valid/event time and system/observed time, enabling both historical truth and “what the system knew then” queries.

BM25 — Probabilistic lexical ranking with inverse document frequency, term-frequency saturation, and length normalization.

BM25F — Fielded BM25 that combines title/body/other fields with independent weights and length normalization.

Block-Max WAND — Exact/approximate dynamic-pruning algorithm using score upper bounds for postings blocks to skip documents unable to enter top-k.

BRIGHT — Reasoning-intensive retrieval benchmark where ordinary embedding leaderboards transfer poorly.

C

Calibration — Mapping a score/confidence to an empirically reliable probability/risk under a specified distribution.

Candidate generation — Fast high-recall first-stage retrieval before an expensive reranker or selector.

Canonical document — Faithful structured representation of source bytes with alignment/provenance; separate from retrieval-specific normalized views.

CDC (change-data capture) — Incremental stream of inserts, updates, and deletes used to keep indexes synchronized with a source.

Citation completeness/recall — Fraction of externally verifiable claims that have sufficient cited support.

Citation correctness/precision — Fraction of citations that actually support their attached claims.

Claim — Atomic externally verifiable proposition extracted from an answer or source for support evaluation.

Claim recall — Fraction of required/reference claims covered by retrieved evidence or generated output, depending on stage.

Chunk — A retrievable unit derived from a source document with stable source identity and offsets.

Chunking — Constructing retrieval units; includes fixed, structural, semantic, proposition, parent-child, hierarchical, visual, and other policies.

ColBERT — Multi-vector late-interaction retriever using query-token to document-token MaxSim.

ColPali — Visual document retriever using text-query/image-patch late interaction over rendered pages.

CombMNZ — Score fusion that multiplies summed normalized scores by number of retrievers contributing a positive score.

CombSUM — Sum of normalized scores across retrieval runs.

Compression — Reducing evidence tokens/vectors using extractive, abstractive, token-level, or latent methods. Must evaluate support preservation.

Conformal risk/control — Uses calibration data to provide statistical risk bounds under explicit assumptions.

Context precision — Fraction of retrieved/selected context that is relevant or supports required claims, depending on definition.

Context utilization — Degree to which a generator uses relevant provided evidence; not identical to faithfulness.

Contriever — Unsupervised contrastively trained dense retriever.

Corrective RAG (CRAG method) — Grades initial retrieval and refines or switches search when evidence appears poor. Distinguish from CRAG benchmark.

CRAG benchmark — Dynamic/long-tail/complex QA benchmark with web and KG/API resources, unrelated to the corrective method acronym collision.

Cross-encoder — Jointly encodes query and document for high-interaction ranking; documents cannot be independently pre-indexed by that score.

D

Data plane — Runtime path that processes user requests and evidence, as opposed to ingestion, control, and evaluation planes.

DeepCT — Learned contextual term-impact weighting stored in an inverted index.

DeepImpact — Learned per-term impact retrieval combined with document expansion.

Dense retrieval — Independently maps query/document to dense vectors and ranks by vector similarity.

Dense X Retrieval — Work advocating proposition-level retrievable units, not a generic synonym for dense search.

Denoised negative — Hard-negative candidate filtered/soft-labeled by a stronger teacher to reduce false negatives.

Direct preference optimization (DPO) — Optimizes preference pairs relative to a reference policy without an explicit learned reward model.

DiskANN — SSD-oriented graph ANN system using Vamana-style navigation and compressed/in-memory routing structures.

Distillation — Training a smaller/retrieval model to match teacher scores, distributions, attention, margins, or downstream utility.

Document expansion — Adds predicted query/term text to an index-only document representation to improve lexical matching.

DPR — Dense Passage Retrieval; dual BERT encoders trained with positives, in-batch negatives, and lexical hard negatives for open QA.

Drift — Change in query, corpus, model/index, or evaluation distribution that can invalidate quality and calibration.

E

E5 — Embedding family trained from broad weakly supervised text pairs with contrastive learning and query/passage prefixes.

Embedding — Numeric representation of query/document/unit; not anonymized data and not evidence by itself.

EMDR² — End-to-end training approach for multi-document reader and retriever using an EM-like objective.

Entailment — Whether evidence logically supports a claim. Semantic similarity or answer mention does not imply entailment.

Entity linking — Mapping textual mentions to canonical entity IDs; a central graph retrieval bottleneck.

Evidence set — One or more units collectively sufficient to support an answer or required claims.

Exact match (EM) — Strict normalized string answer metric; insensitive to many semantic and citation properties.

Exact vector search — Scores every vector, used as an oracle to measure ANN loss on feasible corpora/samples.

F

Factual correctness — Truth relative to reference/world/time, distinct from faithfulness to provided context.

FAISS — Library/research system for efficient exact and approximate vector search, including IVF/PQ and GPU algorithms.

Faithfulness/groundedness — Degree to which generated claims follow from provided evidence; exact formulation varies across evaluators.

False negative (retrieval training) — Unlabeled document treated as negative even though it is relevant/supporting.

Federated retrieval — Queries several independently owned or implemented stores and fuses authorized results.

FiD (Fusion-in-Decoder) — Independently encodes question-passage pairs and lets one decoder attend over their combined states.

FiD-KD — Distills reader signals from FiD into a retriever.

Filter-aware ANN — Integrates metadata/ACL predicates into vector search rather than discarding unauthorized/nonmatching results only afterward.

FLARE — Active retrieval method that triggers retrieval from uncertain tentative generation.

FLOPS regularization — Learned-sparse penalty based on average vocabulary activation, intended to reduce inverted-index work.

Freshness — Operational property of capturing, indexing, selecting, and answering from correct current/as-of evidence.

G

GaRAGe — Human-annotated benchmark for relevance-aware factuality, attribution, and insufficient-evidence deflection over web/private data.

Generator — Model producing answer text/structure from request and selected evidence.

Global search — Corpus-level synthesis over themes/communities/summaries, contrasted with local fact retrieval.

GPL — Domain adaptation using generated queries, mined negatives, and cross-encoder teacher margins.

GraphRAG — Ambiguous family using graphs for construction, retrieval, organization, or generation. Microsoft GraphRAG specifically uses extracted graphs, communities, and reports for global/local search.

GRIP — 2026 framework representing retrieval and intermediate/answer/stop states as structured tokens inside generation.

GritLM — Unified model trained for both embeddings and generation.

Ground truth/qrels — Human or constructed relevance/support labels used to evaluate retrieval; often incomplete and version-specific.

H

Hallucination — Generated content unsupported or false under a declared definition. “Unsupported by context” and “factually false” are not identical.

Hard negative — Highly ranked or semantically similar nonpositive used for training; may be an unlabeled false negative.

Hierarchical RAG — Retrieves across structural or generated summary levels, including parent-child and RAPTOR-style trees.

HiPRAG — 2026 method adding hierarchical process rewards for search/no-search decisions.

HippoRAG — Associative graph-memory approach using entities and Personalized PageRank; HippoRAG 2 adds passage nodes/contextual relationships.

HNSW — Hierarchical navigable small-world graph ANN index with parameters such as M, efConstruction, and efSearch.

Hybrid retrieval — Combines complementary retrieval families, commonly sparse and dense.

HyDE — Generates a hypothetical answer document and embeds it for zero-shot dense retrieval.

I

ICT (inverse cloze task) — Self-supervised retriever pretraining using a sentence as query and surrounding context as positive.

IDF — Inverse document frequency; higher weight for rarer terms under a specific corpus/smoothing formula.

Index generation — Immutable, complete version of corpus representations and index configuration served atomically.

Index staleness — Mismatch between source/current model and served index; also used during retriever training when document embeddings lag parameters.

Indirect prompt injection — Malicious instructions embedded in retrieved data/tool outputs that attempt to override policy or control actions.

Information gain — Reduction in uncertainty or uncovered needs after a retrieval step; generally estimated by a proxy.

Ingestion lag — Time from upstream source change to availability in the served index.

INSTRUCTOR — Instruction-conditioned embedding approach producing task- specific representations from natural-language instructions.

InfoNCE — Contrastive softmax objective distinguishing a positive from sampled negatives, controlled by temperature and negative distribution.

Inverted index — Maps terms to compressed postings of documents, positions, frequencies, or learned impacts.

IRCoT — Interleaves chain-of-thought-style intermediate steps and retrieval for multi-step questions.

ITER-RETGEN — Iterates answer generation and retrieval so generated output guides later search.

IVF (inverted file vector index) — Partitions dense vectors by coarse centroids and searches selected lists controlled by nprobe.

J–K

Joint training — Optimizes retriever and generator/control together; creates index refresh and evidence credit-assignment challenges.

Judge — Human, classifier, NLI model, or LLM evaluating quality. Judge error, bias, prompt/version, and calibration must be audited.

kNN-LM — Interpolates a language model’s next-token distribution with tokens from nearest hidden-state neighbors in an external datastore.

Knowledge graph (KG) — Typed entities and relations with provenance/time; may be curated or extracted.

KILT — Shared Wikipedia snapshot benchmark combining downstream task performance with provenance-aware retrieval metrics.

L

Late chunking — Encodes longer document context before pooling token spans into chunk vectors.

Late interaction — Independently precomputes token/patch vectors but performs fine-grained query-unit interaction at scoring time.

Latency percentiles — p50/p95/p99 response times; tail latency is critical for fan-out and sequential agent pipelines.

LDAR — 2026 distraction-aware retrieval learning focused on downstream utility rather than monotonically increasing top-k.

Learned sparse retrieval — Neural term weighting/expansion retained as sparse vocabulary vectors for inverted-index serving.

Lineage — Reverse/forward mapping among source versions, derived units, embeddings, summaries, graph edges, outputs, caches, and training artifacts.

Listwise ranking — Learns or infers ordering using an entire candidate list, not independent items or pairs.

LLMLingua — Prompt compression family using learned token selection/budget control; LongLLMLingua is query-aware for long contexts.

Long context — Model input capacity large enough for full/large sources; does not guarantee effective evidence use or remove retrieval economics.

LongMemEval — Benchmark for long-term conversational memory extraction, updates, temporal reasoning, multi-session use, and abstention.

Lost in the Middle — Finding that long-context models often underuse evidence at middle positions relative to beginning/end.

M

MAP (mean average precision) — Mean of average precision over queries; requires relevance judgments and accounts for ranks of multiple relevant items.

Matryoshka representation — Embedding trained so prefixes of several dimensions remain useful, enabling adaptive storage/search/reranking.

MaxSim — Sum over query tokens of maximum similarity to any document token, used by ColBERT/ColPali.

Memory RAG — Persistent external state with explicit write, retrieve, consolidate, update, forget, delete, and use policies.

Metadata filter — Exact predicate over fields such as tenant, time, language, jurisdiction, or product version.

MIPS — Maximum inner-product search, common for unnormalized dense scores.

MMR (maximal marginal relevance) — Greedy selection balancing query relevance and redundancy against already selected evidence.

MTEB/MMTEB — Broad embedding evaluation suites; global averages do not identify the best retriever for one RAG workload.

Multi-hop retrieval — Finds evidence connected across two or more reasoning steps, often requiring decomposition, iteration, or graphs.

Multi-vector retrieval — Stores several token/region/patch vectors per unit rather than one pooled embedding.

MuSiQue — Multi-hop QA benchmark designed to reduce shortcut reasoning.

N–O

nDCG@k — Discounted cumulative gain normalized by ideal ranking, supporting graded relevance.

Negative mining — Selecting nonpositive training documents, including random, in-batch, lexical, ANN, teacher-denoised, or adversarial negatives.

NoMIRACL — Multilingual benchmark reporting hallucination when evidence is absent and miss when evidence exists; it does not measure full answer accuracy.

Nougat — Vision-to-markup model for parsing scientific document pages.

Oracle context — Gold supporting evidence given directly to the generator, used to isolate generator capability.

Oracle candidate recall — Maximum possible retrieval/selection performance given the first-stage candidate pool.

ORQA — Latent dense retrieval trained from answer strings, initialized with ICT, preceding REALM/DPR.

Outcome reward — RL reward based on final answer; vulnerable to spurious search/evidence credit.

Over-search/under-search — Retrieving unnecessarily versus failing to retrieve when needed; report both.

P

Parent-child retrieval — Indexes small child units for matching and expands to a larger parent for coherent generation.

Passage — Contiguous text retrieval unit; length/boundary definition must be specified.

Personalized PageRank — Random-walk ranking biased to query/seed nodes, used for associative graph retrieval.

PoisonedRAG — Targeted knowledge-poisoning attack that crafts retrieval and generation payloads; not a defensive method.

Precision@k — Relevant retrieved items divided by k (or returned count under a declared convention).

Process reward — RL signal for intermediate retrieval/search behavior rather than only final answer.

Prompt RAG — Serializes retrieved evidence into the generator input without architectural integration.

Proposition retrieval — Indexes atomic self-contained claims with parent/ source alignment.

Provenance — Authentic source identity, version, location, transformation, and custody for evidence.

Pseudo-relevance feedback (PRF) — Assumes top initial results are relevant and derives expansion/reweighting; can cause query drift.

PQ/OPQ — Product quantization encodes vector subvectors with codebooks; optimized PQ learns a rotation to reduce distortion.

Q

Q-RAG — 2026 value-based RL system that learns evidence selection/STOP with a frozen LLM.

QAMPARI — Long-form/list QA benchmark often used for attributed generation.

Query2Doc — LLM query expansion that appends a generated pseudo-document while retaining original query terms.

Query drift — Transformation/feedback moves search away from the user’s actual information need.

Query likelihood — Ranks documents by smoothed probability of generating the query terms.

Query routing — Chooses retriever/source/tool/no-search/long-context path.

Qrels — Query-document relevance judgments; incomplete qrels bias evaluation and hard-negative training.

R

RAG-Sequence — Original RAG variant marginalizing one latent document for the entire generated sequence.

RAG-Token — Original RAG variant marginalizing latent documents separately at each output token.

RAGAS — Automated reference-light evaluator framework; metrics/judges/APIs must be versioned and calibrated.

RAGChecker — Claim-oriented framework diagnosing retriever and generator contributions/failures.

RAGTruth — Human span-level hallucination corpus for generated outputs under provided references; not a retrieval benchmark.

RankRAG — Instruction-tunes one LLM for candidate ranking and answer generation; first-stage recall remains a ceiling.

RankGPT — LLM listwise/permutation reranking, often applied with sliding windows.

RAPTOR — Recursive clustering and abstractive summarization into a retrieval tree.

RaLMSpec — Speculative retrieval/batched verification for faster iterative retrieval-augmented LM serving.

Recall@k — Fraction of relevant items found in top k; definition may require one item, all supporting facts, or complete alternative provenance sets.

RECOMP — Trained extractive/abstractive evidence compression with selective empty augmentation.

Relevance — Query-document usefulness under a judgment guideline; topical similarity, support, and authority are different.

RemoteRAG/PRAG — Research directions for privacy-preserving remote/ distributed similarity retrieval under explicit assumptions.

Reranker regret — Relevant evidence present in candidates but pushed below the retained cutoff by reranking.

RETRO — Retrieval-enhanced language model using chunk-neighbor cross-attention during pretraining/inference over a massive datastore.

Retriever — Component returning ranked external units for a query/state.

RRF (reciprocal-rank fusion) — Sums inverse shifted ranks across runs; robust to incomparable raw score scales.

Risk-coverage curve — Error/risk among answered examples as answer coverage varies by confidence threshold.

RM3 — Relevance-model pseudo-feedback interpolation widely used in lexical retrieval.

S

SafeRAG — Security benchmark covering noise, conflicts, advertisements, and denial-of-service attacks; not a universal safety guarantee.

ScaNN — ANN system combining partitioning, anisotropic quantization, and reordering for inner-product search.

Score calibration — Transforms retrieval/ranking scores into comparable probabilities or stable fusion inputs using labeled data.

Self-RAG — Trains reflection tokens for retrieve/relevance/support/utility control during generation.

Semantic cache — Reuses results/answers for similar queries; high privacy, intent, time, and invalidation risk.

Selector — Chooses the evidence set under relevance, coverage, redundancy, authority, freshness, risk, and token constraints.

Set selection — Optimizes collective evidence coverage rather than independent top-k ranking.

SFT (supervised fine-tuning) — Likelihood training on demonstrations such as answers, citations, or search trajectories.

SPLADE — Learned sparse vocabulary expansion/weighting with ranking and sparsity objectives.

SPANN — Hybrid memory/disk ANN using in-memory centroids and disk posting lists with closure augmentation/query pruning.

Sparse retrieval — Scores nonzero vocabulary/feature overlap, usually via an inverted index; includes lexical and learned sparse methods.

Source diversity — Independence/coverage across source families; duplicates do not count as corroboration.

Source of truth — Authoritative upstream state used for reconciliation, permissions, corrections, and deletion.

Stale answer rate — Fraction of answers using superseded/temporally incompatible evidence under the declared task.

Sufficient context — Evidence contains enough information to answer, beyond individual passage relevance.

T

Table RAG — Retrieval and reasoning over structured tables, often combining schema/table/row/cell retrieval with SQL/calculation.

Temperature — Contrastive softmax/decode parameter controlling distribution sharpness; distinct uses should not be conflated.

Temporal RAG — Uses valid/event/system time, versions, and temporal constraints to answer current or historical questions.

Tombstone — Deletion marker preventing stale events/replicas from resurrecting removed content.

Top-k truncation — Only highest retrieved candidates enter a latent sum or context, imposing a gradient and recall ceiling.

TREC RAG — Annual NIST shared evaluation for retrieval, generation, nuggets, and sentence-level citations; pin year and judgments.

Trust domain — Set of data/components/principals sharing security assumptions and policy boundary.

TTFT (time to first token) — Delay before streaming output begins; includes retrieval and model prefill in end-to-end measurement.

U–V

Unanswerable query — No sufficient authorized evidence exists; distinct from technical retrieval failure.

Utility-aware retrieval — Optimizes downstream answer/claim performance rather than relevance alone.

Valid time — Interval when a fact is true in the represented world.

Vamana — Pruned navigable graph underlying DiskANN-style search.

Vector database — Storage/search system for embeddings plus metadata; does not by itself define chunking, retriever quality, generation, or evaluation.

Verifier — Component checking claims, citations, schema, calculations, policy, or safety. It is fallible and requires calibration/audit.

Visual RAG — Retrieves and generates from page/image/region/patch evidence, not merely OCR text.

VisRAG — Direct visual document retrieval plus VLM generation approach.

W–Z

WAND — Dynamic-pruning top-k algorithm using term score upper bounds.

Weak supervision — Labels derived from answer strings, citations, clicks, teachers, or synthetic generation rather than direct expert relevance.

Working memory — Temporary state for the current task, distinct from persistent user/agent memory.

xRAG/latent compression — Compresses retrieved text into learned embedding tokens for a generator, improving token efficiency at an auditability cost.

Zero-shot retrieval — Evaluation on tasks/domains without task-specific relevance fine-tuning; contamination and broad pretraining still matter.

← Chronological index of RAG and its technical substrateRAG coverage and verification matrix →
Typesetting mathematics…
The Evidence Path · evidence cutoff 9 August 2026