What is retrieval-augmented generation?
RAG retrieves material from an external source before the model answers and puts what came back in its context. A document store may split files into chunks and search their embeddings; a structured source can query rows with SQL. Vector embeddings aren’t required for the retrieval step.
It exists because a model’s weights only hold what it was trained on, and a context window is finite, so an agent can’t carry its whole knowledge base into every call. RAG pulls the relevant slice from internal docs, past tickets, or a product catalog, and you can update that source without retraining anything.
If retrieval misses the source or brings back an old one, the model can still write a fluent answer. It has not seen the material the request needed.