Why can't I just increase the context window instead of using RAG?

A longer context window helps, but it doesn’t solve the problem RAG solves. Even a million-token window doesn’t hold a company’s full documentation, ticket history, and product catalog at once, and every token you paste in gets billed and processed on every call, whether that document was needed or not.

Stuffing a window full of documents also degrades retrieval within the prompt itself. Research on long-context models found they’re measurably worse at using information buried in the middle of a long context than information near the start or end (Liu et al., “Lost in the Middle”). A bigger window filled with mostly irrelevant text can make a model perform worse than a smaller window filled with exactly the chunks that matter.

The two aren’t opposites. A bigger window makes each retrieved chunk cheaper to include, but it doesn’t remove the need to find the right ones first.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y