Insight 05
We Tested an Advantage a Marketed Product Did Not Have
AlphaCitation ran a counterfactual on four engines to test what an approved cardiovascular indication would do to a weight-management medicine's standing in AI treatment decisions. Favourable decisions rose 8.8 points, in the same direction on every engine. The brief named four events that would reopen it; ten days later, a competing medicine was approved for a cardiovascular claim on the same attribute the counterfactual had tested.
An indication the medicine does not hold was worth 8.8 percentage points in AI treatment decisions. On 18 August 2026 we assumed that a weight-management medicine in the incretin family carried an approved indication for cardiovascular risk reduction in the Gulf market we were measuring — it does not — held every other condition identical, and put the same 46 pre-registered decision questions to four AI engines in both worlds.
Favourable decisions rose from 83 of 147 matched pairs to 96: 56.5 per cent to 65.3 per cent, 8.8 percentage points absolute and a 15.7 per cent relative increase, resolved at an exact two-sided p of 0.041 and positive in every one of the four providers tested. The recommended posture was Monitor. Under “act now”, the brief said: nothing.
The more useful half of that result is the half that can expire. The brief closed by naming, in advance, four events that should reopen it. The fourth read: a competing product gains, loses or strengthens a claim on the same attribute the counterfactual assumes. Ten days later that condition fired: a competing incretin medicine received a United States indication to reduce the risk of major adverse cardiovascular events, approved on 27 August and announced the following day. The market met a written expiry condition inside a fortnight.
A frozen two-arm design is what makes the number refutable
The subject is withheld deliberately. This was our own run on our own instrument, not a commissioned engagement, and naming a medicine beside a regulatory state it does not hold is a sentence we will not publish. The method is published in full, because the method is what lets a reader judge the number.
Two arms ran against one design. One carried a single assumed change; the other carried nothing. Question population, provider set, order, parser, endpoint, threshold and inferential method were identical, so the assumption was the only difference between the arms. The design was frozen before the first paid call, which is what stops a choice made after seeing a result from shaping the result.
Thirteen net decisions moved, and the interval excludes zero
Most decisions did not move, which is the ordinary condition of a decision environment. Of the 35 pairs whose outcome changed at all, 24 moved toward the medicine and 11 moved away: a net of 13 across 147 pairs. The inference rests on those two discordant cells, not on the 112 pairs that agreed with themselves.
Where 147 matched decisions went when one claim was added
Of 147 matched decision pairs: 72 favourable in both arms, 24 moved toward the medicine, 11 moved away, 40 unfavourable in both arms. Net movement 13 pairs, 8.8 percentage points.
| Transition | Pairs | What it means |
|---|---|---|
| Favourable in both arms | 72 pairs | The decision was already favourable and stayed there. |
| Moved toward the medicine | 24 pairs | Unfavourable in the real arm, favourable once the indication was assumed. |
| Moved away from it | 11 pairs | Favourable in the real arm, unfavourable in the scenario arm. |
| Unfavourable in both arms | 40 pairs | The stronger proposition did not reach these decisions at all. |
Favourable in both arms
- Pairs
- 72 pairs
- What it means
- The decision was already favourable and stayed there.
Moved toward the medicine
- Pairs
- 24 pairs
- What it means
- Unfavourable in the real arm, favourable once the indication was assumed.
Moved away from it
- Pairs
- 11 pairs
- What it means
- Favourable in the real arm, unfavourable in the scenario arm.
Unfavourable in both arms
- Pairs
- 40 pairs
- What it means
- The stronger proposition did not reach these decisions at all.
Twenty-four decisions that were unfavourable in the real world turned favourable on a single assumed claim. Eleven went the other way.
- Source
- AlphaCitation Scenario Lab, run confirmatory-2026-08-18. Transition matrix read from the frozen run record.
- Observed
- 18 August 2026
- Scope
- Base: 147 complete-valid matched pairs of 184 launched, across 46 pre-registered questions and four providers. Each pair is one decision, not a patient and not a prescription.
The 95 per cent interval runs from +1 to +15 percentage points. Excluding zero is what resolves the finding; reaching 15 is what puts a major movement inside the range the data admits, although the observed estimate of +8.8 does not reach it. That threshold is a magnitude classification and nothing more — not a finding about commercial importance, which this experiment does not measure.
The direction held everywhere it was tested — Anthropic +15, Perplexity +11.4, OpenAI +6.5, Google +6.5 percentage points — so no single engine carries the pooled figure. Those four magnitudes are descriptive: the strata are smaller, none was individually tested, and the differences between them were not measured.
The resistance that remained was about access, not efficacy
Forty decisions stayed unfavourable with the stronger proposition in hand. Where a decision resisted, the reasoning sometimes said why, and efficacy was never the reason it gave.
What the reasoning named when the indication was not enough
Of 40 decisions that resisted: supply and availability stated in 7 across 3 providers; safety and tolerability in 3 within 1 provider; a competitor's established position in 2 across 2 providers; reimbursement and coverage in 2 across 2 providers; and 28 decisions gave no reason attributable to the medicine.
| Stated in the reasoning | Decisions | Evidence state |
|---|---|---|
| Supply and availability | 7 decisions, 3 providers | Stated more than once and in more than one decision context. |
| Safety and tolerability | 3 decisions, 1 provider | Stated more than once, within a single provider. |
| A competitor's established position | 2 decisions, 2 providers | Stated more than once and in more than one decision context. |
| Reimbursement and coverage | 2 decisions, 2 providers | Stated more than once and in more than one decision context. |
| No reason attributable to the medicine | 28 decisions | The largest group, and larger than every stated reason combined. |
Supply and availability
- Decisions
- 7 decisions, 3 providers
- Evidence state
- Stated more than once and in more than one decision context.
Safety and tolerability
- Decisions
- 3 decisions, 1 provider
- Evidence state
- Stated more than once, within a single provider.
A competitor's established position
- Decisions
- 2 decisions, 2 providers
- Evidence state
- Stated more than once and in more than one decision context.
Reimbursement and coverage
- Decisions
- 2 decisions, 2 providers
- Evidence state
- Stated more than once and in more than one decision context.
No reason attributable to the medicine
- Decisions
- 28 decisions
- Evidence state
- The largest group, and larger than every stated reason combined.
Supply and availability was stated most often and across the most providers — a claim about access, which no indication resolves.
- Source
- AlphaCitation Scenario Lab, run confirmatory-2026-08-18. Reasons classified as trace-supported only where stated more than once and in more than one decision context.
- Observed
- 18 August 2026
- Scope
- Base: the 40 decisions that were unfavourable in both arms. This is evidence about what the environment says while resisting — not evidence that the constraint caused the resistance, and not evidence that removing it would change a decision.
The scenario moved the decision. It did not remove the resistance.
That boundary is what a brand team can act on. A regulatory or clinical win is measurable in the decision environment, and this one measured at roughly nine points; it does not buy the access argument, because the access argument is made of different material. The brief’s instruction followed from that: test supply and availability next, then reimbursement and coverage. Not the indication, which had already been shown to work.
A pre-registered reopen condition fired ten days after the run
The event is public and specific. On 27 August the FDA approved a supplement to the Mounjaro (tirzepatide) application, and Eli Lilly announced it on 28 August. The prescribing information revised that month carries the indication “to reduce the risk of major adverse cardiovascular (CV) events (CV death, non-fatal myocardial infarction, non-fatal stroke) in adults with type 2 diabetes mellitus who are at high risk for these events”, on the strength of SURPASS-CVOT.
That is not the claim we assumed. Our counterfactual assumed cardiovascular risk reduction in adults with established cardiovascular disease, alongside an existing weight-management indication; the approved claim is event reduction in adults with type 2 diabetes at high cardiovascular risk. Different company, different product, different market, different population. It confirms nothing we measured, and we do not present it as confirmation.
It is the condition we wrote down. The counterfactual isolated one attribute — a cardiovascular claim attached to an incretin medicine — and priced what that attribute was worth in a decision environment where nobody in the set held it. Ten days later a competitor held a cardiovascular claim on that attribute, in a reference market. The measured 8.8 points do not transfer to another product in another market, and nobody should read them that way. What changes is the question: the attribute is no longer unclaimed, so the baseline the scenario was measured against is no longer the environment that exists.
A forecast would have aged silently through that fortnight. This one published the condition that retires it before the event that met it, which is the difference between a scenario and an opinion with a number attached.
The capability is pricing the event that has not happened
Brand teams are routinely asked to plan for things that have not happened: an indication, a competitor’s label change, a guideline revision, a supply constraint. The usual answer is a workshop and a judgement. The question underneath it is measurable, and it is not “will this happen”.
It is: if it happened, would the decision environment move, by how much, in which engines, and where would it still refuse. That is a controlled experiment — two arms, one difference, a frozen design, a pre-registered question set and an inference chosen before the data exists. It answers a strategic question with a measurement instead of a conviction, and it states the conditions under which its own answer stops being current.
Scenario Lab measures how a controlled AI decision environment responds to an assumed change. It does not measure, predict or estimate prescriber, payer or market behaviour, and no result it produces is evidence that the assumed change has occurred.
AlphaCitation is a digital product owned and operated by Akunudo LLC-FZ.