There's a reflex, when a grounded system makes something up, to go rewrite the prompt. Add a stern instruction. Add a few-shot example. Sometimes it helps. More often the problem is upstream: the model was handed four sources that disagreed, or one source that didn't actually contain the answer, and it did the reasonable thing with bad inputs.
Thin evidence is the main culprit
If the retrieval step returns nothing relevant, a well-behaved model should say so. Most will instead synthesize something plausible from parametric memory and cite the nearest source, which is worse than an error because it looks like an answer.
The mitigation is a relevance threshold applied before the model sees anything. If the top passage scores below a cutoff, return no context and let the system report that it couldn't find an answer. Teams resist this because the refusal rate goes up. The alternative is a confident wrong answer with a citation attached, which is the failure mode that actually costs you trust.
Disagreement needs to survive to the model
When sources genuinely conflict — two outlets reporting different numbers — collapsing them into a single passage destroys the signal. Keeping them separate, with dates attached, lets the model notice and say so.
Our Answer endpoint declines when evidence is thin rather than filling the gap. It scores worse on naive benchmarks and better on anything you'd put in front of a customer.