ORIGINAL TESTING SYSTEM · PHASE 4

AI Tools Galaxy Test Lab

The Test Lab separates provider-source research from observed product behavior. A protocol can be public before any tool has been run, but a page cannot claim hands-on testing until the run conditions and evidence are recorded.

Current evidence state0 completed hands-on runs are published. 6 planned run records are waiting for real evidence.

Evidence states

SOURCE-CHECKED

Claims are tied to suitable first-party sources; no site-run result is implied.

PROTOCOL READY

Inputs and measurement rules are fixed, but observed results are still blank.

HANDS-ON TESTED

Allowed only after a real run has complete conditions and retained evidence.

RETEST NEEDED

A material product change makes an older evidenced run insufficient for current claims.

Why these benchmarks are harder to fake

The controlled document and writing packs use fictional facts created for the benchmark. That gives the site its own answer key and makes unsupported additions, missed qualifiers and contradiction handling directly checkable. The image protocol uses countable constraints instead of an arbitrary beauty score.

Results are recorded by dimension. AI Tools Galaxy does not turn unrelated dimensions into a single universal winner or hide plan/model conditions behind a star rating.

Published protocols

PROTOCOL-READY

Controlled document grounding and contradiction test

Protocol ID: DOC-GROUNDING-V1

Measure whether an assistant extracts known facts, preserves a material qualifier, identifies a deliberate contradiction and refuses to invent a fact that the source pack never states.

Initial targets: ChatGPT, Claude, Gemini, Perplexity AI, Microsoft Copilot

Measured dimensions

  • fact-extraction: Count exact answers to known facts.
  • contradiction-detection: Record whether the answer reports both conflicting dates instead of silently choosing one.
  • unknown-handling: Record whether the assistant says the vendor is not stated instead of inventing one.
  • qualifier-preservation: Record whether the accessibility sign-off condition is preserved.
  • source-traceability: Check whether cited document/section references support the nearby claim.
  • correction-effort: Count user corrections needed before all scored items are accurate.

Limits

  • One synthetic document pack cannot represent every real document workflow.
  • Model and product behavior can change after the recorded test date.
  • Account-specific retrieval, file limits and connectors must be recorded with the result.

Open benchmark instructions

PROTOCOL-READY

Constrained factual rewrite test

Protocol ID: WRITE-CONSTRAINT-V1

Measure instruction following and factual preservation during a short rewrite without judging subjective writing style as a universal score.

Initial targets: ChatGPT, Claude, Gemini, Perplexity AI, Microsoft Copilot

Measured dimensions

  • required-facts: Check every required fact against the source brief.
  • invented-facts: Count unsupported names, numbers, dates or causal claims.
  • format-constraints: Check headline, paragraph count, prohibited claims and requested callout.
  • length-constraint: Record whether the response stays inside the specified word range.
  • correction-effort: Count revisions required to satisfy all objective constraints.

Limits

  • The protocol evaluates constraint following, not literary quality or creativity.
  • Language performance may differ outside English and must be tested separately.

Open benchmark instructions

PROTOCOL-READY

Countable image constraint and revision test

Protocol ID: IMAGE-CONSTRAINT-V1

Measure whether an image system follows countable layout, object, text and revision constraints without turning aesthetics into an unsupported universal score.

Initial targets: Adobe Firefly

Measured dimensions

  • object-count: Check whether exactly three blue cubes and one yellow circle appear.
  • text-accuracy: Check whether the required text ATG TEST is legible and exact.
  • layout-constraint: Check required upper-right text placement and white background.
  • prohibited-elements: Record people or extra objects not requested.
  • revision-effort: Count generations/edits required to satisfy the objective brief.
  • export-condition: Record actual aspect ratio/dimensions and any upscaling used.

Limits

  • The protocol does not assign a universal beauty or artistic-quality score.
  • Different models inside the same product must be recorded as separate conditions.

Open benchmark instructions

First controlled benchmark

DOC-GROUNDING-V1: AI Document Grounding Benchmark publishes the exact synthetic test pack, objective dimensions and validity rules without publishing a fabricated result. The first ChatGPT run must be completed in a fresh conversation that has not seen the answer key.

Evidence provenance gate

A run cannot become HANDS-ON TESTED merely because a result field was filled in. The build now requires recorded run conditions, the preserved first-response hash, an evidence-bundle hash, scoring version, scorer identity, scoring time and a resolvable evidence reference.

Download the evidence provenance policy. Missing evidence stays blank rather than becoming a zero score.

Raw data policy

The authoritative run records live in one structured dataset. Public JSON and CSV exports are generated from that source so tables and future comparison pages cannot silently drift apart. Blank result fields mean the test has not happened; they are not treated as zero scores.

Download JSON run records · Download CSV run records

Product change tracking

0 independently verified material provider change records are currently published. A record is added only after a dated primary source is checked and the user-facing impact is written independently; an empty log is preferred to guessed history.

Download the current change log

What still requires a human run

Accounts, model selectors, file uploads, image generation and product-specific interfaces can require login or paid access. When that happens, the site owner must provide the real output or sanitized evidence. Until then, the corresponding run remains awaiting evidence.

See the review methodology for publication, correction and disclosure rules.