CITATION QUALITY SCORECARD · REVIEWED AUGUST 2026
How to Evaluate RAG Citation Quality in 2026
Tests for citation correctness, completeness and source authority so a grounded answer does more than display links.
A retrieval system can cite a real document that does not support the sentence, omit the decisive source or quote an outdated version. Link presence is not evidence quality.
This guide is designed for RAG builders, knowledge managers and support teams. It turns the topic into a reviewable sequence rather than asking readers to trust a provider label, a detector score or a fluent model answer.
Practical recommendation: Evaluate citations at claim level: check entailment, coverage, freshness, authority and whether users can open the supporting passage.
Before you start
Write down the exact task, accountable owner, approved data, affected people and the result that would be unacceptable. Use safe representative examples during the first pass. Where health, legal, employment, financial, safety or regulatory obligations may apply, involve a qualified professional and follow the rules that govern your organization.
1. Create claim-level examples
Include factual, procedural, exception and unanswerable questions. Mark the authoritative source and passages that should support each claim.
Document the decision made during “Create claim-level examples”, the evidence consulted and the person responsible for the next action. That short record helps RAG builders, knowledge managers and support teams distinguish a repeatable control from an informal habit.
2. Measure citation correctness
For every cited claim, verify that the linked passage actually supports it without relying on unstated assumptions or nearby unrelated text.
Test “Measure citation correctness” with a normal case and a deliberately difficult case. Record what passed, what required correction and which condition should trigger a human review for RAG builders, knowledge managers and support teams.
3. Measure citation completeness
Identify material claims with no supporting source. A response with one correct citation can still contain several unsupported conclusions.
Assign an owner and completion criterion for “Measure citation completeness”. If the evidence is missing or contradictory, pause the workflow instead of allowing speed or model confidence to become the approval rule.
4. Test authority and freshness
Prefer current approved policy, product documentation or controlled knowledge over an old duplicate. Encode effective dates and document status.
Keep the input, output version and reviewer note associated with “Test authority and freshness” where policy permits. This makes later corrections traceable without retaining unnecessary sensitive data.
5. Test the reader path
Confirm links open with the user's permission, anchors reach the relevant passage and the interface distinguishes quoted fact from model synthesis.
Review this step after material changes to the model, provider, prompt, data source or connected system. A control that worked in one configuration should not be assumed to cover the next one.
Common failure modes and controls
The following table is a pre-launch challenge list. Teams should adapt it to the systems, people and permissions in their real deployment.
| Failure mode | Practical control |
|---|---|
| Citation points to right document, wrong passage | Store passage identifiers and test entailment. |
| Outdated duplicate ranks higher | Filter by status and effective date. |
| User lacks source permission | Apply retrieval authorization before generation. |
| System answers when evidence is absent | Include abstention and escalation cases. |
What to measure
Do not optimize a single headline number. Measure useful outcomes together with correction effort, critical failures and the human work needed to make the result acceptable.
- citation correctnessDefine the numerator, denominator, owner and review period for citation correctness; compare like-for-like workflow versions.
- claim coverageTrack claim coverage beside correction effort and serious exceptions so a faster result does not hide weaker quality.
- outdated-source rateSample outdated-source rate by risk level and user group; investigate material changes instead of relying on one aggregate percentage.
- appropriate abstention rateSet a baseline for appropriate abstention rate, record the intervention and review whether the change remained useful after human verification.
Final review checklist
- Gold passages are documented
- Claims are evaluated separately
- Source dates are checked
- Permissions are enforced
- Links open to evidence
- No-evidence cases abstain
Frequently asked questions
Is a high retrieval score enough?
No. Similarity does not prove that the passage supports the generated claim.
Should every sentence have a citation?
Material factual claims should be traceable; purely connective or clearly labelled analysis may not require one.
How are conflicting sources handled?
Apply source authority and effective-date rules, and show the conflict when no approved resolution exists.
Primary and official sources
- OpenAI guide to working with evaluations (checked August 13, 2026)
- Anthropic guidance for reducing hallucinations (checked August 13, 2026)
- NIST AI Risk Management Framework and Generative AI Profile (checked August 13, 2026)
This independent guide was reviewed against the linked primary or official materials on August 13, 2026. It provides an operational framework, not legal, medical, financial or security certification. Product features, terms and policies can change, so verify time-sensitive details at the source.
Continue your comparison
Use AI Tools Galaxy to compare access models and read the detailed editorial profiles available for selected tools. Keep tests small, protect sensitive data and verify important output before acting on it.
Browse AI tools