AI Search Evidence Lab
Turn repeated AI-answer captures into evidence you can audit: exact denominators, citation persistence, topic-specific source patterns, prompt coverage, fact conflicts, URL failures, and bounded intervention results.
Example data — fabricated observations that demonstrate repeated citations, a broken URL, a fact conflict, and a correction.
Bundled methodology and source registries last verified Jul 28, 2026.
Runs entirely in your browser — nothing you paste is uploaded, and data is saved locally only when you explicitly choose the browser-save action. Clear saved removes that browser copy without changing the text currently in the editor. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
Site passport Local context for this saved site
Local data
Saved targets, named lists, and recent check summaries remain only in this browser.
Rate this tool
What this adds to AI-search measurement
- Observation provenance: exact prompt version, surface, model/checkpoint, retrieval mode, sample group, date, and extraction method.
- Passage-level evidence: distinguish a known cited URL from a citation whose supporting passage was actually captured.
- Repeatability: persistence and 95% Wilson intervals instead of a single-run visibility label.
- Source influence: repeated citations grouped by normalized domain, with prompt and surface breadth shown separately.
- Correction fixtures: human overrides remain attached to the original observation and classifier version.
- Intervention discipline: before/after panels remain directional unless compatible control evidence exists.
Packet structure
The top-level object accepts observations, corrections, and interventions. Historical records should be append-only. Change a prompt by creating a new version; change a classifier or denominator by creating a new methodology version. Do not rewrite earlier captures.
Required observation fields
id, capturedAt, promptId, promptVersion, sampleGroup, sampleOrdinal, source, provider, surface, retrievalMode, evaluationState, extractorVersion, methodologyVersion, evidenceState
How to use it
- Capture repeated answers under the same prompt version, surface, retrieval mode, locale, and methodology.
- Include citations and their passages when the source exposes them. Use a passage hash when retention rules prevent storing the text.
- Paste or upload the normalized packet and review excluded observations before reading rates.
- Investigate fact conflicts and broken URLs manually. A conflict shows disagreement, not which value is correct.
- When evaluating an edit, preserve the deployment and recrawl dates and add a comparable control panel when possible.
Limitations
The lab analyzes supplied evidence; it cannot see undisclosed retrieval, internal ranking, cached context, or every answer shown to every user. Citation frequency is not rank, authority, causal influence, traffic, or conversion. A verified crawler request proves that request, not later use in an answer. Human corrections are counted but never silently rewrite raw evidence.
Frequently asked questions
Does this query AI systems for me?
No. It analyzes observation packets you already captured or exported. Keeping acquisition separate prevents an API execution from being mislabeled as a consumer-product result.
What counts in the denominator?
Only evaluated observations with usable evidence. Refusals, provider failures, and not-evaluated records remain visible but do not become zeroes.
How many repeats do I need?
The ai-search-evidence-v1 contract requires at least three compatible observations before applying a stability label and prefers five. Even then, the result is a sample, not a universal rank.
How does the topic-specific source map work?
Add version-matched prompt records with topic labels. The report groups observed citation domains by those supplied labels and preserves citation appearances, distinct prompt breadth, surface breadth, and captured-passage counts. It does not infer topical authority.
Does a before-and-after increase prove my edit caused it?
No. An uncontrolled change is directional evidence. A compatible control panel makes the inference stronger, but the tool still avoids causal language.
Is pasted data uploaded or stored?
Analysis runs in your browser. Nothing is saved unless you explicitly choose Save in this browser; you can clear that local copy at any time.
Feature requests for Ai Search Evidence Lab
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
➕ Request a feature
New requests are reviewed before they appear here.
ツールについて
繰り返し取得したAI回答を分析し、引用の安定性、トピック別のソース傾向、プロンプトの網羅性、事実の競合、URL失敗を確認します。
制限付きの介入結果を記録し、観測された回答と検証不能な推測を分けて比較します。
機能
- 反復AI回答の引用安定性
- トピック別ソースとプロンプト網羅性
- 事実競合とURL失敗の検出
- 制限付き介入の前後比較
仕組み
回答、引用URL、プロンプト、取得時刻を正規化し、同じトピックでの引用の出現、欠落、変動を集計します。事実の競合と取得失敗を分け、介入後の結果を元のキャプチャと比較します。
制限事項
- 結果は提供されたキャプチャとプロンプトに限り、AIモデルの全体的な挙動、将来の回答、順位、トラフィックを保証しません。
- 繰り返し取得でもモデル、時刻、地域、プロバイダーの変化を完全には再現できません。
よくある質問
引用が一度出れば安定していますか?
いいえ。複数回のキャプチャで出現率と変動を確認してください。モデルやプロンプトの変化で結果は変わります。
このラボはAIの内部理由を見せますか?
いいえ。回答と出典の観測結果を比較します。内部推論や非公開のランキング要因は推測しません。
介入は本番コンテンツを変更しますか?
いいえ。入力されたキャプチャと制限付きの比較結果を扱います。公開変更は別途承認・検証してください。