Do bigger models hallucinate less?
No, not reliably. On Vectara’s HHEM leaderboard, which scores models on how often they state something a source document doesn’t support, Intel’s Neural Chat, a 7-billion-parameter model, hallucinates at 2.8%, lower than GPT-4’s 3%, a model estimated at roughly 1.8 trillion parameters. Bigger models tend to hallucinate less on average, but the correlation is loose enough that a much smaller model, tuned well for staying faithful to a source, beats a much larger general-purpose one. What predicts a model’s hallucination rate better than raw size is how it was trained and aligned for the specific job of representing its context accurately. A small model built for that one job can outperform a large one optimized for something broader. Judge a model on a hallucination benchmark that matches your actual use case, summarization, RAG, tool use, not on its parameter count.