The benchmark and answer key are fixed, but no tool result is published on this page until a real fresh-session run and retained evidence exist.
What this benchmark tests
Three fictional project documents contain six objectively checkable questions. The pack includes known facts, two conflicting launch dates, an unstated vendor, a material accessibility gate and a sample-size detail that can be confused with total capacity. Because the facts are synthetic, the ground truth is controlled by AI Tools Galaxy rather than by changing web pages.
This is not a universal intelligence score. It tests one narrow document-grounding workflow under recorded conditions.
Objective dimensions
- Fact extraction: does the assistant return the exact budget, learner-seat and demo-station facts?
- Contradiction detection: does it report both October dates instead of silently choosing one?
- Unknown handling: does it state that no contracted vendor is named rather than inventing one?
- Qualifier preservation: does it preserve the 30 September accessibility sign-off gate and staff-only consequence?
- Source traceability: do the cited document sections support the nearby answer?
- Correction effort: how many user corrections are needed after the first complete answer?
Download the exact test pack
Use these files unchanged. The answer key is intentionally not linked from this public page because a tested assistant should not receive it before the run.
Dataset license
The public DOC-GROUNDING-V1 benchmark files are distributed under the AI Tools Galaxy Benchmark Dataset License v1.0. You may use the files to run, reproduce, discuss or critique benchmark experiments with attribution. Republishing the benchmark pack as a substitute download, selling the dataset itself, removing the source notice or presenting it as your own original benchmark requires permission.
Run conditions that must be recorded
Before a result can become HANDS-ON TESTED, record the product, visible model or mode, plan/account type if known, platform, test date, and whether web search, connectors, memory or other tools are active. Preserve the first complete response before correcting the assistant.
A result with missing conditions or missing first-response evidence stays awaiting evidence.
Initial comparison set
The first planned comparison uses the same pack with ChatGPT, Claude, Gemini, Perplexity and Microsoft Copilot. Results will be shown by dimension and test condition, not collapsed into a fabricated universal winner.
Why no result is shown yet
The build process can create protocols and validation rules, but it cannot honestly invent a product run. The current AI Tools Galaxy development conversation has already inspected the hidden answer key, so it is deliberately excluded as a valid ChatGPT test session. The first ChatGPT result must come from a fresh conversation that has not seen the answer key.
