all answers

General agent concepts

Hallucinations

A hallucination is an agent stating something its sources don't support. The term covers three distinct properties of an answer, and each one can be checked on its own.

Faithfulness asks whether the answer accurately represents the sources it cites. An agent can retrieve the right document and still misquote it, invert a number, or attribute a claim the document doesn't make.

Groundedness asks whether the answer comes from the retrieved context or from the model's own priors, with the context left unused. An ungrounded answer can be correct whenever the prior happens to be right, which makes it hard to notice, and it fails when the prior is wrong.

Context presence asks whether the information needed to answer was retrieved at all. When it wasn't, the failure is in the retrieval step, before generation begins.

Hallucination is a problem because a made-up answer looks the same as a right one. The person asking usually can't verify it themselves, and if the answer feeds a decision or another agent, the made-up fact gets acted on.

11 questions

Answered, plainly.

How can an answer be faithful to its context and still be wrong?Faithfulness only checks the answer against what was retrieved, so an answer built entirely from a partial or unrepresentative slice of the source passes even though it's wrong.answer →What's the difference between faithfulness, groundedness, and context presence?Faithfulness vs groundedness: one checks the answer against its source, the other whether the retrieved context was used. Context presence, whether it arrived.answer →How do I catch agent hallucinations when my error rate stays flat?Run a content check against the retrieved context on production traffic, not a test set: a hallucination returns a clean status, so only what the answer says can reveal it.answer →What is a hallucination in an AI agent?An agent stating something its sources don't support, whether it misquotes a document it retrieved, ignores that document entirely, or never retrieved anything to begin with.answer →Why do AI agents hallucinate?A language model generates the most probable next word, not a verified one, so it produces fluent text whether or not anything backs it, unless something forces it to check.answer →Why does a wrong answer often look just as confident as a right one?Nothing about how an answer reads reveals where it came from; a number pulled from the wrong source is written in the same fluent, certain sentence as one pulled from the right one.answer →Does setting temperature to 0 stop an agent from hallucinating?No. Temperature controls how randomly the model samples tokens, not whether its most likely answer is actually true, so a deterministic model still hallucinates.answer →Do bigger models hallucinate less?No, not reliably. On Vectara's leaderboard a 7-billion-parameter model beats GPT-4 on hallucination rate; training predicts it better than size.answer →Why does an AI model make up an answer instead of saying it doesn't know?Because the tests it's graded on score a confident guess the same as a right answer, and score 'I don't know' as a guaranteed loss, so guessing wins.answer →How do you reduce hallucinations in an AI agent?Ground every claim in something the agent actually retrieved or looked up, not the model's own memory, then check what it wrote for claims that still don't trace back.answer →Does checking whether an answer is grounded require the model that generated it?No. A 2026 paper found a separate 'observer' model reading hidden states can flag ungrounded answers about as well as reading the generator's own activations.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y