Home › Methodology

Methodology

No black box. Here is exactly how the diagnostic works.

Every number in a Citewise report is produced by a fixed scoring rule, not by an AI's judgment — the AI writes prose, never scores. A human sits on top of that as a QC step: they can flag or throw out a false match (a common-word firm name that hit by accident), but they never hand-tune a score up or down. This page documents each stage, including the quality gates that stop a bad report from ever reaching you.

1 · Question generation

We generate 12 prospect-intent questions for your firm — category-level questions a real buyer would type, phrased for your specialty, fee model, and locality ("best fee-only advisor for physicians in Austin", never "is [your firm] good"). You can review them in the report; the raw phrasing ships with your results.

Why category-level questions?

Asking an engine about your firm by name tests brand recall. Asking the question your prospect actually asks tests whether you make the shortlist — that is the moment revenue is won or lost.

2 · Four-engine capture

Each question runs fresh against ChatGPT, Claude, Gemini and Perplexity — 48 answers total, captured in a single measurement window so the results are comparable. We keep the complete raw answers and include them in your report: you read exactly what your prospects read.

Exactly which surface we capture — so the claim is checkable.

We measure the free, logged-out consumer surface a first-time prospect lands on, on the engine's default model at the time of your run (not a pinned older version, and not a paid power-user tier): ChatGPT — the free web app at chat.openai.com, default model; Claude — the free claude.ai web app, default model; Gemini — the Gemini web app (gemini.google.com) free tier, default model — this is the conversational app, not Google Search's AI Overviews, which is a separate surface; Perplexity — the free default (Quick) answer at perplexity.ai, not Pro search. Every report records the model label and date each answer was captured under, so you can re-ask on the same free surface and compare. If you want a specific paid tier or the AI-Overviews surface measured instead, say so in your intake and we'll note the swap in the report.

How we control for personalization.

Consumer chat apps personalize on account history, saved memory and location, and that is the largest source of run-to-run variance. So each question is asked in a fresh, logged-out session with memory and chat history off, at the engine's default sampling settings (we don't override temperature — we want the answer a real prospect gets, not a lab setting), and we record the locale used so a geo-sensitive query ("near me") is captured for your market, not ours. This gets us close to a clean, unpersonalized baseline — but it is a baseline, not a claim that every prospect sees exactly this. A logged-in prospect with their own history may see something different, which is one more reason we measure a window and re-measure on the Monitor rather than treating a single run as ground truth.

How many times we ask each question — and the honest variance that creates.

In the standard diagnostic, each of the 12 questions is asked once per engine inside a single measurement window — that's how you get 48 answers. We're upfront about the trade-off: these models are stochastic, so even in a clean logged-out session an answer can name a different set of firms between two runs, and a single unlucky (or lucky) generation can flip one mention. That's a real limit of a one-shot capture. Two things blunt it. First, the score spans 48 answers across four engines, so no single generation dominates the grade — one flipped mention moves it by roughly a point, not a letter. Second, borderline or surprising results are exactly what a human reviewer re-asks before the report ships. Where a specific answer looks run-sensitive, we say so in prose next to that raw answer rather than presenting it as settled — and the Monitor's monthly re-run is the real fix, because a stable signal is one that survives repeated measurement, not one lucky pass. If you want a question sampled multiple times in one window for a tighter read, ask in your intake.

3 · Mechanical scoring

Each answer is scored by rules, not judgment:

Your score out of 100 is the share of possible mentions and citations you actually hold across the window.

Why a mention and a citation count the same — and where that's a simplification.

In the score, one mention and one citation each count as one point. We weight them equally on purpose: engines differ in how they surface sources (Perplexity attaches links; ChatGPT and Claude often name a firm without one), so a fixed "a citation is worth 2× a mention" rule would just encode our guess about which engine your prospects use, and we'd rather not pretend to know that. It is a simplification — commercially, a Perplexity citation with a live link can be worth more than a bare mention with no way to click through. Rather than bury that in one weighted number, the report keeps mention and citation as separate columns and ships the raw answer, so you can see whether you were linked or merely named and judge the commercial value yourself. A future weighted score is on the roadmap; when there's evidence for a specific weighting we'll publish it here before it changes any number.

4 · Quality gates

A report only ships if it passes mechanical QC: at least 70% of engine calls must succeed and at least two distinct engines must respond — a half-failed capture never reaches a paying client. Failed runs route to a human, and your 48-hour clock keeps running on us, not you.

Free scan vs. paid report — what runs where.

The paid diagnostic runs the full pipeline above: 12 questions × 4 engines = 48 answers, mechanical scoring, and this QC gate before it ships. The free scan is a smaller, human-run preview — a person runs 2 of your prospects' real questions across the same engines by hand and emails you the exact answers, usually within one business day. It's a real look at the same engines, not the full scored pipeline; treat it as a taste of the method, and the paid report as the measured version with the scoring and QC applied.

5 · The prescriptive layer

Only after the measured results exist does AI write anything: a strategist model converts your specific losses (which questions, which engines, which competitors) into five prioritized actions — directory completeness, citation surfaces, comparison content — in the order that moves answers in your category.

6 · The fix pack

Two paste-ready assets are generated for your firm and mechanically validated before delivery:

We're candid about limits: llms.txt adoption is ahead of AI-crawler consumption today (a 2026 million-site crawl found most llms.txt files get few crawler requests yet), which is why the action plan weights third-party surfaces — directories, comparisons, communities — that engines demonstrably cite now. The fix pack makes your own site ready; the plan makes the rest of the web vouch for you.

7 · The monthly delta (Monitor)

On the Monitor plan, the identical pipeline re-runs every month, automatically, on your renewal. Your report leads with the delta: score movement, questions gained and lost, competitor changes — and a fresh action set. AI answers drift with model updates; the delta is how you catch it before your pipeline feels it.

Honest limitations

See the methodology run on your firm — free.

We run 2 of your prospects' real questions across the major AI engines and email you the exact answers, usually within one business day. No payment, no call.

Get my free scan See a sample report