all answers

General agent concepts

Retrieval augmented generation

Retrieval-augmented generation, RAG, is the pattern of fetching relevant content at request time and putting it in the model's context before it answers. Documents are split into chunks, each chunk is embedded into a vector, and the chunks sit in an index. At request time the question is embedded the same way, the index returns the closest chunks, and the model generates its answer from what came back.

It helps an agent because a model's weights hold only what it was trained on, and context windows are finite, so an agent can't carry its whole knowledge base into every request. RAG hands it exactly the slice a request needs, from internal docs, past tickets, or a product catalog, without retraining anything, and the corpus can be updated at any time.

RAG is bad at questions whose answers aren't localized in a few passages. Retrieval returns a top handful of chunks, so anything that needs aggregating across many documents, counting, or summarizing a whole corpus won't work. It's also weak when the question's wording doesn't resemble the document's wording, since similarity search leans on phrasing. And when the corpus holds stale or contradictory content, the model will faithfully answer from the wrong version.

11 questions

Answered, plainly.

Does RAG stop a model from hallucinating?No. RAG reduces hallucination by giving the model real content to answer from, but it can't stop the model from ignoring or misreading it.answer →How do I tell a stale RAG index apart from a truncated context window?A stale index returns an old source; truncation loses a retrieved chunk before the model sees it. Compare retrieval with the prompt to tell them apart.answer →What can retrieval grading over traces not tell me?Trace grading catches noise and gaps in what came back, but it can't see the relevant chunk that existed and was never retrieved at all.answer →What is retrieval-augmented generation?RAG pulls relevant information from an external source at request time and puts it in a model's context before the model answers.answer →Why can't I just increase the context window instead of using RAG?You can, up to a point, but a bigger window doesn't shrink a corpus to fit, and stuffing it full raises cost and hurts retrieval inside the prompt.answer →Why did my RAG answer miss the exception to the rule?The chunk boundary split the rule from its exception, so only the rule came back and the model answered faithfully from half of it.answer →Why does RAG fail when my question doesn't match the wording of the source document?Retrieval runs on embedding similarity, which tracks meaning loosely, so a chunk phrased differently than the question can rank low or never surface.answer →Should you use RAG or fine-tuning to teach a model new information?RAG for facts that change or need a citation; fine-tuning for a fixed skill, tone, or format baked permanently into the model.answer →What's the difference between grounding and RAG?Grounding is the goal, tying an answer to a checkable source; RAG is one way to reach it, retrieving relevant text and putting it in the model's context first.answer →What's the difference between RAG and GraphRAG?Standard RAG retrieves the chunks closest to a question; GraphRAG first builds a knowledge graph of the corpus and summarizes it by cluster, so it can also answer aggregation questions.answer →What's the difference between RAG and prompt engineering?RAG fetches new content into the model's context at request time; prompt engineering only changes how the model is told to use the context it has.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y