Skip to content

Methodology

Every finding AlphaCitation publishes is measured, not guessed.

This page explains exactly how, so you can trust a score without taking our word for it.

What we query
Buyer-intent questions put to ChatGPT, Claude, Gemini and Perplexity in your market frame. How the questions are built
What we record
Every answer, the entities named in it, the order they appear, and the sources cited. How answers are scored
How it is checked
Against a human-labelled gold set, with an interval on every score rather than a bare number. How we verify
What it cannot establish
It measures answers, not prescribing, sales or patient outcomes. Where the measurement stops
Last updated
24 September 2026
Sections
10
Stated boundaries
4

Operating stance

Foundational accuracy and neutrality are the product. AlphaCitation measures what AI systems tell people, where those answers come from, and whether the documented record supports them. The method is the same for every client in every sector, and no client can change it: the questions are set before measurement begins, the engines and the decision rule are fixed, and every finding traces back to the verbatim answers behind it and, where a claim is checked, to the primary source it was checked against.

We report what we measure, including results a client would rather not see. We help authoritative information become findable and verifiable, and we never claim to control what any AI system says. We do not seed content, fabricate grassroots activity, produce synthetic media, or profile or target individuals. We work for companies, institutions and campaigns on these terms and no others.

What we measure

Each scan runs a battery of real buyer-intent questions (the same kind of question a real prospect would actually ask an AI assistant, not a synthetic “does X exist” prompt) against ChatGPT, Claude, Gemini and Perplexity. We record whether your brand is named, where it ranks relative to competitors, and how it's framed (sentiment, objections, pricing perception, and more).

Sequential sampling, not a fixed sample size

We sample prompts in rounds across every engine and stop as soon as your brand's mention-rate estimate reaches a target precision (a 95% Wilson confidence interval of a set half-width), or the prompt pool is exhausted. A clear-cut brand (named almost always, or almost never) converges quickly; a contested one keeps sampling toward the cap. This means the sample size follows the precision the reading needs, rather than the precision following however many prompts happened to run. Where the pool is exhausted before the target is met, the interval is simply wider — and it is published at that width, rather than presented as though the target had been reached.

Confidence intervals, not point estimates

Every score ships with a real 95% confidence interval, not a bare number, built by combining several independently measured sources of real-world uncertainty the statistically correct way, not stacked naively and not invented. When a metric's uncertainty genuinely can't be pinned down, the interval says so, rather than rounding to look more precise than the data actually supports.

Deterministic scoring

Given the same answers, the scoring engine always produces the same score. There's no hidden randomness in how we grade an AI's response once it's been collected. The uncertainty in your interval comes entirely from real, independently measured sources, not from an unstable scoring formula.

What this doesn't do

A score is a point-in-time snapshot. AI answers shift as the underlying models and their sources update, so a re-scan next month can move for reasons that have nothing to do with anything you did. That's why trends across scans matter more than any single number. We're also honest about the limits of automation: our own reading of an AI's answer is measured, not assumed to be perfect, which is exactly why it's folded into your confidence interval instead of ignored.

Data sources

Every scan queries ChatGPT, Claude, Gemini and Perplexity directly: real API calls to real models, never scraped. Brand and competitor questions are put to the models on every scan. The shared category questions — the ones that are identical whichever brand is being measured — may be answered from a cache for up to 24 hours. Growth-tier plans and above add real sentiment measured from Reddit, X, YouTube, GitHub, Hacker News and Stack Overflow, kept as a separate, clearly labeled signal rather than blended into the AI-engine numbers.

Findings are also checked against live official records rather than against our own database: FDA Structured Product Labeling via openFDA and EMA EPAR for medicines, and the authoritative public register for each other sector covered — NHTSA safety ratings, USDA FoodData Central, the NIH Dietary Supplement Label Database, EU Ecolabel and EU eAmbrosia among them. A register that is read and returns nothing is recorded as absence of confirmation, never as evidence of absence.

Controlled measurement environment

Measurement runs in a controlled, repeatable query environment against each provider's model interface. That is what makes like-for-like comparison possible: the same battery, the same parser and the same decision rule, cycle after cycle, so a change in the reading is a change in the provider rather than a change in how we asked.

An individual consumer session can differ. Personalisation, interface-specific retrieval, search behaviour, model routing, account configuration and product-level system instructions all vary between users and between apps. Each measured response is therefore an observation within a defined test environment, not a claim that every result reproduces every possible consumer session. What the method is built to do is measure provider behaviour, and movement in it, repeatably over time.

Verifying what we found

For regulated industries, every finding is anchored to an independent regulatory source and sealed in a tamper-evident, independently verifiable evidence chain, so a claim's history is provable without trusting our database. See Label Ledger.

Where the measurement stops

Where the measurement stops

A reading is only trustworthy if its edges are stated. Four boundaries describe where this measurement stops, and every capability a client surface links to is filed beneath the one it belongs to.

  1. Where a reading needs history

    One cycle shows where something stands; showing movement takes the same thing observed on more than one occasion.

    Some readings need the same thing observed on more than one occasion, or need a source's own revision history alongside what the answers say. A single cycle establishes where something stands, not how it moved. Where movement is reported, it is movement in the answers, measured between cycles this programme ran.

    • Clinical guideline revisions

      What is measured is whether a revision has reached the answers: a guideline change becomes visible here when engines begin to reflect it, which is the lag this programme exists to observe. The guideline's own publication history is a record about the guideline, and this feed reads the answers.

    • HTA decision changes

      Health technology assessment determinations are held as current state rather than as a dated sequence, so a movement between two determinations is not raised as an event here. The determinations themselves are reported on the Access surface.

    • Prescribing restriction changes

      Prescribing restrictions are read as they stand at each cycle, so this feed reports what the answers now describe. Dating the move itself would mean reading the restriction's own history rather than the answers.

    • Cross-engine claim propagation

      Following one claim between engines would require the same claim to be recognised as the same claim across separate providers and separate cycles. Each cycle is measured on its own terms, so the passport reports where a claim stands rather than the path it travelled.

    • Label revision timeline

      Each claim is checked against the approved label as it reads at the time of the scan, and the passport reports that comparison rather than a sequence of them. A timeline is a record about the label's own history, which is a different subject from the claim.

    • Per-theme trend across scans

      A trend needs the same theme observed on at least two cycles, with something having happened in between. Until a claim has been re-observed after a logged intervention there is no movement to report, and a single reading drawn as a line would suggest a direction the data does not contain.

    • Temporal currency

      Scoring how current the evidence behind an answer is would require dating each source an engine drew on. Publication dates are not read from cited sources, so this radar reports what the evidence says rather than how recent it is.

    • Longitudinal discourse shift

      A trend needs the same scope observed on more than one occasion. This layer reports the discourse as measured in a single collection, so it describes where the discussion stands rather than how it has moved.

  2. Where a reading needs a consequence

    We observe AI answers and the evidence behind them, never prescribing, dispensing, patient outcomes or a regulator's conclusion.

    This programme observes what AI says and whether the evidence supports it. It observes no prescribing, no dispensing and no patient outcome, and it makes no regulatory determination. What an answer caused, and what a regulator would conclude about it, sit outside what any scan of answers can establish.

    • Commercial impact of a claim

      This programme measures what AI says and whether the evidence supports it. It observes no prescribing, dispensing or patient outcome, so what a wrong claim costs cannot be established from any scan — however many cycles are run. A figure here would be a model of consequence, not a measurement of one.

    • Regulatory indication

      A claim theme is a therapeutic subject an answer discusses, not a specific approved indication. Binding one to the other is a regulatory determination, and reading it off the language of an answer would assert a precision the answer does not carry.

    • Why an engine names the molecule instead of the brand

      Brand attribution records whether an answer that chooses a medicine's molecule also names the medicine's brand, engine by engine. It does not establish why an engine writes the molecule or the dose instead of the brand: that depends on the sources an engine retrieves and how it composes an answer, which no reading of the answers themselves can determine.

  3. Where a reading needs a record we do not read

    A comparison that needs a record outside AI answers is made only where that record is public.

    Some comparisons need a record that lives outside AI answers: a payer's own determination, a manufacturer's communication, a fact about your organisation. Where such a record is public, the comparison is made and reported on the surface that holds it. Where none is public, no accuracy reading is claimed rather than estimated.

    • Safety communication changes

      This feed reports movement in what AI engines say from one cycle to the next. It does not receive safety communications from manufacturers or regulators directly, so a new communication appears here only once it has changed what the engines answer — not at the moment it is issued.

    • Reimbursement decision changes

      A reimbursement decision recorded against this brand is a fact entered for the propagation clock, not a change this programme detected. Reporting it as a detected change would misstate how it was learned.

    • Payer formulary changes

      Movement in the access tier an AI engine places this brand at is measured and does appear on this feed. A real payer's own formulary movement is a different fact, read as current state rather than as a tracked history, and the two are never merged.

    • Claim ownership

      Themes are observed in answers, not assigned to people. Who owns a theme internally is a fact about your organisation rather than about the answers, and this product does not hold it.

    • Coverage accuracy

      This radar measures how AI represents cost and access, never what a payer decided; the two readings are kept apart by design. Whether an AI coverage claim matches a real determination is a comparison against public payer records, and it is reported on the Access surface.

    • Eligibility accuracy

      Same boundary as coverage accuracy: an eligibility claim is checked against a matched coverage determination's own criteria text on the Access surface, not scored by this radar, which reads only how AI presents eligibility rather than whether a real determination agrees.

    • Restriction accuracy

      Whether AI describes a plan's real restrictions correctly is measured where a plan publishes its own plan-level files — ACA Marketplace coverage, reported on the Access surface. Medicare, Medicaid and employer coverage publish no plan-level source, so no accuracy reading is claimed there.

    • Clinician-payer disagreement

      Comparing what clinicians conclude against what payers require needs two separate datasets, from two different partners, scoped to the same market and period, with topic vocabularies never designed to line up. A divergence number computed across two mismatched populations would look identical to a real finding.

    • Plan-level access reality

      What an individual plan decides about a drug is not published by any source this layer can reach, so no claim here is checked against a real determination. The payer axes compare emphasis between populations, never coverage outcomes.

  4. Where a reading is made in aggregate

    Some readings are taken across a whole question battery or at the level of a claim, rather than item by item.

    Several readings are taken across a whole question battery, or at the level of a claim, rather than item by item. A finer cut of the same data is a different judgement rather than a smaller one, so it is named here instead of implied.

    • Claim prevalence changes

      How widely a claim has spread, and how firmly it has settled, is measured against a registry of approved facts, on compliance cases. Themes observed in open AI answers have no approved record to be counted against, so this feed reports what was said rather than how far it has travelled.

    • Per-question drill-down

      Readings on this surface are aggregated across the whole question battery. Reporting them question by question, or market by market within a question, would need each answer's claims and sources held against that individual question rather than against the cycle.

    • Evidence relationship tagging

      Whether a cited source supports, qualifies or contradicts a claim is judged at the level of the claim against the registry, not source by source. Tagging each citation individually is a judgement about a source's contents, which is a different reading from the one this makes.

    • Explicit versus inferred source attribution

      An answer does not reliably reveal whether the model quoted a source or drew on it without saying so, so what is recorded is whether the answer cited a source at all.

    • Scientific qualification

      This radar reports whether what was said holds up against the registry, not how confidently it was said. How carefully an answer hedges a medical claim is a reading about the language of the answer rather than about its accuracy.

    • Position-level agreement

      This layer compares how much emphasis each population places on a topic, never whether their conclusions agree. Comparing conclusions needs a structured vocabulary of positions, and inferring one from a summary sentence would be a fabricated comparison.

    • Isolated persona effects

      Audience readings compare answers to questions framed for each audience. A framed question changes both who is asking and how the question is worded, so a difference between audiences reflects the framing as a whole. Separating the audience from the wording needs the same question asked with only the audience changed, which is a finer reading than the one reported.

AlphaCitation is a digital product owned and operated by Akunudo LLC-FZ, a free-zone company registered in the United Arab Emirates. Questions about this methodology? analyst@alphacitation.com.