How ARA Measures Trust
What trust means when the audience is a machine, why counting recommendations is not the same thing, and what we do about the fact that two of the four models we audit cannot read the live web.
Brand trust has always meant the same thing: the confidence someone has that you can and will deliver on what you promise. It is measured, conventionally, after the fact with repeat purchase, retention, referral, Net Promoter. All of it lagging, all of it human.
There is now a second audience making that judgement, and it makes it before your customer ever sees you. When somebody asks an AI model what to buy, the model decides whether to put your brand forward. That decision is a trust judgement, made on your customer's behalf, in a channel where you have no visibility at all.
You cannot survey it. There is no panel and no unaided recall. So the question becomes how you measure trust in an audience that will not answer a questionnaire.
The answer is not to count how often it recommends you.
The drivers do not change. The audience does.
Marketing already knows what builds brand trust: a consistent voice, corroboration from people who are not you, evident competence, and transparency about who you are and what you do. Those drivers are not superseded by AI. They are what the machine reads.
Our five dimensions are those same drivers, measured in the machine population rather than a survey panel.
Conventional trust driver · What ARA scores
Consistent brand voice — Identity & Voice
Public perception — ratings, reviews, word of mouth — External Sentiment
Competence, reliability, "can deliver on promises" — Semantic Clarity
Transparency, findability, integrity of the record — Structural Readiness
The trust judgement itself — Recommendation
The last row is the one that has moved. Conventionally the trust judgement is the customer's, and you observe it afterwards through retention, repeat purchase and referral. Increasingly it is made first by a machine, on the customer's behalf, before your brand is ever seen and it is observable at the moment it happens rather than a quarter later.
Recommendation is presence. Trust is presence plus accuracy.
Counting appearances is the easy half, and it is what most AI-visibility work stops at. It tells you a machine said your name. It cannot tell you whether the machine was right.
Trust needs both halves, and each fails differently without the other:
Accurate but not chosen. The machine has a correct, complete picture of your brand and still reaches for someone else. Your problem is not comprehension, and every dollar you spend on content and structured data is funding a battle you have already won.
Chosen but not accurate. Worse, and much less visible. You are being recommended on residual familiarity rather than on anything the machine currently understands about you. That position reprices the moment the models refresh. In our Q3 beverage study, Pepsi held the third-highest recommendation score in the category and the lowest semantic clarity — named often, understood least.
A count cannot tell those two apart. They look identical on a mentions dashboard and they require opposite interventions.
The independent record is what makes accuracy measurable
To score whether a machine is right about a brand, you need something to be right against.
Before any scoring happens, we build a dossier on the brand — researched, dated and sourced, deliberately not the AI's opinion of itself. Every near-term claim carries a date and a source URL. Product lines are named rather than described. The brand's own first-person language is captured in direct quotes from its own channels. Gaps are recorded as gaps.
That document is the reference standard. The scorer is handed it as ground truth, and the score is the distance between what the brand actually is and what the machines say it is.
This is the part that cannot be done by looking at model output alone. A system with no independent record can only observe that a model said something. It cannot know that the thing was false. In the Q3 Beer study one model asserted that a Czech brewery was owned by the wrong multinational — confidently, in the register of fact. Only an independent record makes that an error rather than a data point.
A mirror shows you what the machines say. It cannot tell a true reflection from a distorted one.
Five dimensions: four describe the picture, one describes the act
Each scored 0–20, summed to 100.
Structural Readiness — can it retrieve you. The infrastructure layer, and the only dimension that is not an inference. Alongside the model responses we run a live scan of the brand's own web estate — crawler permissions, robots.txt, structured data, redirects, response times — and record which checks pass and fail, per URL. A brand can block the crawler that decides whether it gets recommended, and most brands do not know whether they do.
Semantic Clarity — is the picture accurate and distinct. Findable is not understood. This measures what the machine knows once it arrives, and whether that knowledge separates you from the brand beside you. Failure has a signature: a brand that exists in every dataset ever assembled and comes back described as "the challenger" — accurate, and not a reason to choose anything.
External Sentiment — does anyone but you corroborate it. The social-proof layer. What the machine retrieves about you from voices that are not yours: consumer review, cultural memory, institutional endorsement. High scores return a chorus. Low scores return specs and press releases with no human voice behind them.
Identity and Voice — is the picture consistent enough to reproduce. Not whether you have a voice. Whether it is coherent enough that a machine can read it, reproduce it, and carry your register intact into an answer. A brand whose personality lives in imagery and sixty years of televised association has very little for a system that reads rather than watches.
Recommendation — will it act on the picture. The behaviour. We ask each model the questions a real customer asks — what to buy, what to bring, what to order, what is best for this occasion — and record who is named, and where.
The first four are the machine's account of you, scored for accuracy against the record. The fifth is what it does with that account. Trust is both, which is why the Trust Tier is a total rather than a single number.
Why the fifth dimension is scored differently
Being named first is not the same as being named fifth, so the score cannot be a count.
Recommendation is scored on reciprocal rank: first position carries its full value, second half, third a third, averaged across the models that answered and summed across the questions the brand is eligible for. That signal runs through `round(20 × tanh(signal × 0.8))`.
The curve is deliberate. Early presence moves the score a great deal; further presence moves it less. Going from never-named to occasionally-named is worth more than going from frequent to dominant — which matches how recommendation actually converts. The expensive gap is between zero and one.
And zero is a real number. It does not mean ranked low. It means that across every buying question, in every framing, the brand was never returned. Nineteen brands in our consumer corpus sit there. One of them had been named by an investment bank that same quarter as the biggest World Cup beneficiary among global brewers — the business case was being made in public while the machines returned a blank.
What the corpus shows
Across 132 consumer brands in 12 categories — beer, beverages, banking, beauty, oral care, deodorant, QSR, travel, luxury fashion, non-alcoholic spirits, feed additives and hospitality consultancies — the machines' picture of a brand runs at a median of 16 out of 20. Their willingness to act on it runs at 9.5.
103 of those 132 brands are described more accurately than they are recommended. That is 78% of every consumer brand we have scored: the machine has the picture and will not act on it.
Every one of the 12 categories shows it. What varies is how far it has been closed.
Banking has all but closed it, which matters more than any of the wide gaps: it proves the distance is closable rather than inherent. Non-alcoholic spirits is understood well enough at 13 and acted on at 2.5 — a category that spent a decade explaining itself to people and has not yet explained itself to the machines.
A note on what is not in this. We have measured 53 sports properties across the NFL and the Premier League, and they are excluded here deliberately. A club is not recommended the way a product is: fandom is inherited, geographic and usually lifelong, so "will AI put this forward to someone choosing" is not the question a supporter is asking. Both sat at the extreme end of the distance, and including them would have flattered the finding by measuring purchase intent in a category that does not have any.
The Trust Tiers
Trust Tier / Score / What it means
AWESOME 83–100
Category authority. AI carries the brand's intended frame and acts on it across the demand moments that matter.
STRONG 70–82
Accurate picture, reliably acted on in primary recommendation contexts. Not yet category-defining.
AVERAGE 56–69
Understood, inconsistently chosen.
WEAK 40–55
Known and largely passed over.
MISSING 0–39
Present in the data, absent from the answer.
Not how much a machine knows about you. How far it will go on your behalf, and whether it is right when it does.
The instrument notes we would rather you heard from us
Two of the four models do not read the live web. Gemini fails a live-fetch control every time and invents a plausible answer instead — measured against our own dossiers, 13.3% of its checkable specifics contradict the record, against 5.2% for the control model on identical prompts. GPT-4o declines the fetch outright, which is quieter and arguably worse: it answers from its training corpus, so for any brand that recently rebranded, was acquired or launched something, it is measuring corpus age rather than the brand.
We report it because it is a property of the subject, not a defect in the instrument. What a model invents about your brand is your brand inside that model, and the customer asking it gets the invention. The scoring handles it correctly — accuracy is graded against the record, so a confident falsehood scores down rather than up.
Position and movement are different measurements. Position is scored on the full current question battery. Movement is scored only on the questions carried forward word-for-word from the prior quarter, because that is the only basis on which two quarters are the same measurement. A brand can gain on one and lose on the other, and both numbers are true.
A brand can only move against a number it already had. Brands entering a cohort carry a position and no delta. Zero and absent are different claims and we do not print one for the other.
Same inputs, same score. Every rubric is versioned and every change is logged. Re-scoring an old study under a new rubric is a deliberate act with an audit trail, not a silent drift. A measurement you cannot reproduce is an opinion with a number attached.
Your customers' trust you already measure. This is the audience you don't.
ARA scores five dimensions across four frontier models plus a live agent-readiness scan, against an independently researched record of what each brand actually is. Trust Tier definitions and full methodology: araco.ai/methodology

