If you’re scanning this article and thinking ‘wow, someone’s gone down a real rabbit hole here’, you’d be correct. I have. 😭
Following on from my previous article on audience growth, I’ve been reading a lot of claims recently about the importance of having a decent About page as part of your EEAT strategy to help bots (and people) understand your value proposition as a publisher: you know, things like “hey, we’re a trusted news provider” or “our subscription is worth paying for” or “we’re particularly strong on this subject”.
This all makes complete sense, but I’ve also seen a few comments suggesting that About pages and, heck, even a strong author bio are all you need to get the visibility you’re after in AI chatbots and to foghorn your value proposition to a fleeting audience scrolling from one news story to the next.
So, I wanted to do my own specific research: which pages do AI systems actually cite when they are judging a news brand, and what sort of evidence do those pages provide?
What I found was that the same publisher can be judged using very different evidence depending on a specific question. Trust pulls in one set of pages, subscription value another, specialist authority another, and identity another.
You might be thinking, yeah but we know AI can draw on multiple sources when forming an answer. But I thought it would be useful to point to the actual pages that are getting visibility so you can look at them yourselves.
Oh, and I admit now I didn’t look at third-party citations, which is a critical part of GEO. I might look at that in my next article, but for now, let’s hold hands and leap into this almighty rabbit hole together.
So what did I do? (Methodology)
Using an open-source AI visibility tool (Elmo + Bright Data), I put 201 questions to four AI search systems: ChatGPT, Copilot, Gemini and Google AI Mode. Each question was asked five times per model.
The final dataset contained 4,003 answers and 20,016 citations, collected between 2 and 6 September 2026. Questions covered the UK, the US and selected European countries, across four themes.
The prompt set had two parts: 105 longer comparative questions of the kind a strategist might ask, and 96 shorter queries designed to resemble the sort of thing a reader might actually type into a chatbot.
For every citation I recorded the URL, its position in the model’s source list and the question it was answering.
Citations were attributed to a publisher by matching the cited domain against a tracked list of 98 news brands, including subdomains, so, for instance, support.theguardian.com counts as The Guardian.
Each cited URL was then classified into one kind of page by rule, using the host and path, with specific categories tested before general ones.
There is an important limitation here. The classifier reads URLs, not the actual page content. So a standards document filed under an /about/ path may be counted as an About page even though, in reality, it is an ethics or standards page.
I thought this might be a bit iffy, so I checked that limitation in two ways.
First, I used page titles to sense-check the URL classification. The results barely changed. The biggest shift was ethics and standards, which rose from 2.6% to 4.5%.
Second, I tested a sample of 452 cited URLs against the actual page content. Of the 195 pages I could classify confidently, the URL rule was right 77.1% of the time. It performed particularly well for editorial articles, subscription and pricing pages, and ethics and standards pages, all at around 88% agreement.
The check also helped explain the large unclassified group. Of the unclassified pages I could assess, 79% were ordinary editorial articles. So much of that 41.5% unclassified share on specialist questions is likely to be good ol’ journalism that the URL rule simply failed to recognise.
So I trust the broader pattern that different questions surface different kinds of publisher evidence more than I trust every individual percentage by page type.
For each AI answer, I then looked at every publisher it cited and grouped together all the distinct pages cited from that publisher.
So if one answer cited four Guardian pages and two Reuters pages, that counted as one Guardian case and one Reuters case. This lets us see the mix of evidence AI surfaces for an individual publisher, rather than simply counting URLs across the whole answer.
Got it?
Ok, then. Here’s a selection of the prompts I used:
And some examples of the shorter queries:
(Hmm... hold on. Did I mention that all my prompts were in English? I know, I know this will skew the results and is something I will address in a future study!)
So, what did I find out?
What my analysis showed was that the same publisher is judged with different evidence depending on the question. The About page is part of that picture, not the whole of it. In 76.0% of cases, no About page appeared in the cited evidence at all.
So that means:
The evidence changes depending on what the model has been asked.
Trust questions draw heavily on journalism, alongside standards, ethics and accountability material.
Value-for-money questions bring pricing and subscription pages into the picture.
Questions about a publisher’s identity are much more likely to cite the About page.
That is why I think AI visibility for publishers is better understood as an ‘evidence architecture’ problem than an About-page problem.
What this is telling us is that the picture is not uniform.
In about half of the cases, AI cited just one page from the publisher’s site. But when it cited two or more, those pages were different types of content in around two thirds of cases, for example a news article alongside a subscription page or an ethics and standards page.
The type of evidence also changes depending on what the model is being asked about.
The About-page pattern shows this particularly clearly:
The remainder are cases where an About page was cited alongside other publisher pages.
What AI cites depends on the question
Different questions about the same publisher produce very different citation patterns.
Each publisher counts once here, regardless of its size.
Minor page categories are not shown, so totals may not add to 100%.
This is the central pattern in the research. Read across the About row.
It is worth a quarter of the cited evidence when the question is who you are, and almost nothing when the question is whether you are worth paying for.
The ethics row tells a different story.
Ethics and standards pages account for nearly one in ten trust citations and are effectively absent everywhere else.
These are not general-purpose AI credentials.
The important point is not that one page type wins. It is that different questions pull in different kinds of evidence.
For publishers, this means that the question is not simply which page to optimise, but whether the evidence needed for each claim exists, is discoverable and is consistent across the site.
Five random examples
The percentages are easier to understand when you look at individual answers.
ProPublica, asked to prove real-world impact
Question:
Which US news brands best demonstrate the real-world impact of their journalism? Identify examples of the types of outcomes or impact associated with them and explain how that strengthens their value proposition.
ChatGPT cited 14 distinct ProPublica pages across four kinds.
The general About page came fifth, behind three separate impact-related write-ups.
The point is that ProPublica is not relying on a page that simply claims it creates impact. It has an ecosystem of material that documents the impact.
Among the sources were:
That is what I mean by ‘evidence architecture’.
The Guardian, asked why anyone should pay monthly
Question:
Which UK news publisher has the strongest overall answer to the question “Why should I give you my money every month?” Explain the proposition, evidence and reader benefit.
The answer cited eight Guardian pages across four kinds: governance, pricing, mission, financial results, a Scott Trust statement and even a recruitment page.
The model did not simply cite a marketing statement saying “support independent journalism”.
The cited evidence stretched across:
For a publisher trying to communicate why it deserves money, that is a much wider job than writing better checkout copy.
Mediapart, same question
Ask the same value question about European publishers and Mediapart produces seven cited pages.
These include:
Mediapart’s value proposition is spread across several different kinds of page.
Reuters, asked about accuracy and reliability
Question:
Which US news brands have the strongest reputation for accuracy and reliability? Explain how each organisation communicates and demonstrates those qualities rather than simply claiming them.
Reuters produced one of the clearest examples.
Three Reuters pages were cited, with no general About page among them:
For publishers, that distinction matters.
Claiming that you are trustworthy is not the same thing as having visible evidence of how your journalism earns trust.
The Texas Tribune, asked to prove it deserves funding
Question:
Across US news publishers, which brands best communicate and prove the answer to three questions: “Why you?”, “Why should I trust you?” and “Why should I pay you?” Rank the strongest examples and explain your reasoning.
ChatGPT cited five Texas Tribune pages across four kinds, and no two of them were doing the same job:
Only one of those is the kind of page a newsroom would think of as its own.
The Tribune is a nonprofit, and asked why a reader should trust and fund a publisher, the model assembled the funding route, the accountability policy and the identity statement together.
Two of the five sit on separate subdomains, support.texastribune.org and give.texastribune.org, which in most organisations belong to membership or development teams rather than to editorial.
So the answer to “why should I pay you?” was built largely out of pages the newsroom does not own.
Does the pattern change when the prompts get shorter?
Yes, it does. Short questions produce shallower and somewhat more ‘promotional’ citation patterns.
And the page mix changes too:
Minor page categories are not shown, so totals may not add to 100%.
Three things change when the question gets shorter.
The model cites far less journalism, 32% against 42%.
It cites more pricing pages and homepages, the commercial and navigational material publishers control most directly.
And it stops at one page nearly six times in ten.
The more promotional pattern comes mainly from pricing pages and homepages, not About pages. On short questions, About pages are cited less often than on longer ones.
That has a slightly awkward implication for publishers.
The deeper the question, the more actual reporting appears in the cited evidence. That gives a good newsroom more opportunity to demonstrate what makes its journalism distinctive.
Shorter questions are often closer to how readers actually ask, and the model appears more likely to settle on a handful of commercial and navigational pages.
Those pages may matter more for real reader queries than publishers assume.
The shorter-query sample covers 912 cases across 64 publishers, compared with 1,760 cases across 80 publishers for the longer questions.
Axios, asked a question a real person might type
The query was simply:
what is axios known for
Four words.
In its fullest answer ChatGPT cited seven Axios pages across four kinds. The largest single source of evidence was not the newsroom. It was the help centre. Among them:
what Axios is and how it works
That was not a one-off. Across all six runs of that question, help-centre pages were cited seven times and the About page twice. The evidence was spread across four different Axios subdomains: www, help, pages and about.
A reader asks what a publisher is known for, and the model answers largely out of customer-support articles written to explain a product, on a subdomain most newsrooms never think about.
Build evidence for the claims you want AI to make
The actions below are driven by the data insights above: different claims need different evidence.
The temptation with AI visibility is to ask:
“Which page should we optimise?”
A better question is:
“What evidence would an AI need in order to support the things we want audiences to believe about us, and where does that evidence live?”
That is what I mean by evidence architecture: making the evidence for trust, value, authority and identity easy to find across the parts of the site where it naturally belongs.
If you want to be seen as trustworthy, the evidence will look different from the evidence required to demonstrate value for money.
If you want to be seen as the specialist in football, personal finance or travel, the evidence is different again.
Achieving this requires a content governance model that spans editorial, audience, product, subscriptions, customer service, corporate affairs, legal and marketing.
If you want to be called trustworthy
1. Don’t rely on one vague “editorial values” page.
If you have distinct information about editorial standards, corrections, complaints, fact checking or editorial independence, make sure each important piece of evidence can be found and cited in its own right.
Ethics and standards pages represented 9.8% of trust citations and almost nothing for other questions.
That does not prove that splitting pages causes more AI visibility. It does show that specific accountability ‘assets’ regularly appear in trust-related evidence.
Examples in the dataset included Reuters’ fact-checking material, Euractiv newsroom policies, El País’s ethical code, Guardian complaints and corrections and NRC’s principles.
2. Treat corrections as evidence, not housekeeping.
The models cited three different dated Guardian corrections pages, including very recent ones.
A static sentence saying “we correct our mistakes” is a claim.
A visible, continuously maintained record of corrections is evidence that the process actually exists.
There is a wider trust benefit here too, regardless of AI.
3. Explain how your journalism works.
Editorial articles made up 55.8% of trust citations, more than any other category.
Reuters provides a particularly good example with its piece explaining visual verification.
You do not necessarily need to create self-congratulatory corporate content.
But if your newsroom has distinctive verification processes, specialist expertise, reporting methods or editorial safeguards, explain them somewhere readers and machines can find them.
If you want to be seen as worth paying for
4. Treat the subscription page as evidence, not merely a checkout.
Subscribe and pricing pages were 33.7% of value citations, the single largest individual page category for this question.
That means a pricing page is doing more than converting somebody who has already decided to subscribe.
It may also be one of the main pieces of evidence an AI surfaces when somebody asks whether your product is worth paying for.
This matters particularly on short queries. When the question is four words rather than forty, pricing pages make up 18.8% of the cited evidence, compared with 9.9% for the longer questions.
Publishers should therefore ask whether these pages clearly explain:
what the reader gets
what is distinctive
what is included in different tiers
what the journalism enables or funds
why the product is worth the price
A table of prices and a giant “subscribe now” button may not be enough.
5. Audit your help centre as part of AI visibility.
Help and FAQ pages made up 4.4% of value citations and almost nothing elsewhere.
That 4.4% may look small, but the organisational point is more interesting.
Some of the evidence shaping how your subscription product is represented may live inside customer-support software that nobody in editorial or SEO has looked at for years.
Check whether those pages are crawlable, accurate, current and consistent with the proposition being communicated elsewhere.
6. Make evidence of impact easy to cite.
Annual reports, impact reports and activity reports appeared repeatedly.
Mediapart’s activity report and impact material were particularly visible. ProPublica’s annual report was cited despite being a PDF.
If your journalism produces measurable outcomes, do not leave all of that evidence buried inside an annual PDF nobody reads.
Consider giving important impact material its own durable web presence.
If you want to be called the specialist
7. Accept that there may be no shortcut page.
Editorial articles accounted for 40.5% of citations in specialist questions.
A further 41.5% was unclassified, which is one of the main limitations of this analysis. I cannot say with confidence that those pages were section hubs, topic pages or anything else.
What I can say is that, in the validation sample, 79% of the unclassified pages I could fetch turned out to be ordinary editorial articles. That suggests the true share of journalism in specialist answers is likely to be higher than the 40.5% shown in the table.
So the broader finding is fairly clear: when AI is assessing a publisher’s subject expertise, journalism itself makes up a substantial share of the evidence it cites.
That should not be surprising.
If you want to be considered authoritative in personal finance, football, travel, politics or product reviews, you ultimately need to publish strong journalism in those areas.
Which means you can’t simply bolt a load of these ‘evidence architecture’ pages onto thousands of pieces of commodity content and hope for the best. The actual journalism still needs to rock.
8. Make specialist methodology visible where you have one.
There are some useful examples.
IndyBest has pages explaining how products are tested and who its experts are.
For review journalism, personal finance, health, science, investigations or any other field where methodology is part of the value proposition, explain it.
This does not guarantee citations.
It does, however, create clear evidence for why your specialist journalism deserves to be trusted.
If you want to be described correctly
9. Keep the About page, but don’t expect it to do everything.
Nothing above is an argument against About pages.
About and mission pages represented 24.9% of identity citations, and nearly a quarter of identity answers rested on an About page alone.
So your About page remains extremely important when the question is:
Who are you?
What do you stand for?
What makes you different?
The mistake is expecting the same page to prove that your subscription is good value, that your journalism is trustworthy and that you are the best source for personal finance.
One page cannot carry all of that.
Structural actions across all four areas
10. Make distinct pieces of evidence independently discoverable and citable.
The Guardian has separate pages for its general About information, journalism, history and organisation, and each appeared independently in the dataset.
I am not suggesting publishers should mechanically break a good page into four thin pages.
The useful point is simpler:
Do not bury several materially different propositions inside one generic corporate page if each deserves to stand on its own.
Your ownership is one thing.
Your standards are another.
Your impact is another.
Your subscription proposition is another.
Your methodology is another.
The goal is not more pages for the sake of more pages.
The goal is clearer evidence.
11. Keep important accountability pages on stable URLs.
Half of the dated citations were from 2025 or later, but the tail reached back as far as 2001.
Old evidence can continue to surface.
Mediapart’s 2019 material about its independence still appears. Le Monde’s 2010 ethics charter still appears.
That makes durable corporate, standards and accountability URLs potentially valuable long-term assets.
So avoid casually deleting or moving them.
If URLs genuinely need to change, handle the migration properly rather than assuming these pages have no audience value because humans rarely visit them directly.
12. Make crawler decisions deliberately.
Publishers that blocked the main AI crawlers were named at roughly the same rate as those that did not, 16.5% compared with 16.2%. But they were cited far less often in this dataset, averaging 61 citations compared with 200.
That does not mean crawler blocking caused the difference.
There are too many other factors involved, and this analysis was not designed to show that unblocking AI crawlers would lead to three times as many citations.
But it raises an important strategic question.
If you prevent AI systems from accessing your journalism but still expect them to describe your brand accurately, what evidence are you leaving available instead?
The BBC illustrates the problem.
Its main site was unavailable to the crawlers used here, so the evidence that surfaced included a downloads subdomain and recruitment pages.
That may be exactly what the organisation wants, but it should be a deliberate trade-off rather than an accidental one.
13. Audit the subdomains nobody thinks about.
Some of the most regularly cited evidence in this study comes from subdomains.
Examples include Guardian Support, Le Monde subscriptions, NRC’s code and Washington Post Help.
These properties may be owned by marketing, customer service, product or commercial teams rather than editorial.
To the model, however, they are still part of the evidence surrounding the brand.
That means AI visibility is not purely a newsroom SEO exercise.
The proposition needs to remain coherent across the whole publishing organisation.
What I think publishers should audit
If I were auditing a publisher on the back of this research, I would start with four questions.
Trust
Can a reader, or an AI system assembling sources, easily find evidence of:
editorial standards
corrections
complaints processes
fact checking
expertise
methodology
editorial independence
Value
Can it find evidence of:
what the subscription includes
why it costs what it costs
what subscriber funding enables
journalistic impact
distinct reader benefits
Specialist authority
Can it find:
sustained high-quality journalism on the subject
clear topic or section organisation
specialist authors and expertise
methodologies where relevant
evidence showing why the journalism is credible
Identity
Can it clearly establish:
who you are
what you stand for
who owns you
how you are funded
what makes you different
Then I would look across those pages and ask a fifth question:
Are we telling the same story everywhere?
That may be the most important organisational implication of all.
In AI visibility, when different pages contribute different parts of the evidence, consistency becomes an asset.
If your About page says independence is fundamental, your subscription page says subscriptions protect that independence, your standards page explains how editorial independence is maintained, and your journalism demonstrates it in practice, you have not repeated one claim four times.
You have produced four different kinds of evidence supporting the same proposition.
So what should you not bother doing?
Do not simply rewrite your About page in the hope that this alone will transform how AI systems describe your trustworthiness, subject authority or subscription value.
About pages accounted for 11.8% of the cited evidence around trust and just 4.6% around value.
And 93.1% of value cases contained no About-page citation at all.
Keep your About page good.
Just stop asking it to do everybody else’s job.
One caveat that applies to all thirteen recommendations
One important caveat: none of this proves causation.
The study records what models cite.
It does not establish that publishing an ethics page will make you more likely to be cited on trust questions.
It shows that ethics and standards pages appear disproportionately in the cited evidence associated with trust questions.
Those are different statements.
Likewise, I cannot prove from this dataset that breaking corporate information into separate URLs increases visibility, or that making a help centre crawlable changes an AI answer.
To establish cause, we would need proper before-and-after experiments.
So I would treat the recommendations above as things publishers should audit and test, not magic ranking factors.
Weaknesses of this analysis
With more time and budget, these are the areas I would push further.
This is an overview, not a complete answer, and there are some sizeable gaps.
As I mentioned above, the page classification is not perfect. I grouped pages based on patterns in their URLs rather than reading every one by hand, so some inevitably ended up in the wrong bucket. That is also why there is a fairly large “unclassified” group, especially for specialist and shorter queries.
A third of the cited evidence cannot be independently checked. Of 452 sampled URLs, 137 could not be fetched because they were blocked, paywalled or refused access. Help and FAQ pages were the worst, with two thirds unreachable. Anything I say about those categories rests on a smaller and more accessible subset of pages than the rest.
Site structure can distort the classification. The Washington Post, for example, publishes “About Our Fact Checker” as a dated article within its politics section. Because the slug begins with “about”, my URL-based classifier counts it as an About page rather than as journalism. The Guardian’s editorial code sits under an /about/ path and is affected in the same way.
Around a third of the URLs classified as About are dated articles of this kind. That means About content is somewhat overstated, while journalism is understated. Tightening the rule so that a dated, article-style URL remains classified as journalism, even when it sits under /about/, changes the results from 14.8% to 12.8% for About pages and from 38.7% to 42.4% for editorial articles.
None of the conclusions in this article changes. In fact, the correction strengthens the overall argument. The figures quoted throughout are based on the original classification, so they should be treated as conservative estimates.
A citation is not proof of use. We can see what the model listed as a source. We cannot see which sources actually shaped the sentence a reader received, whether every cited page was opened, or whether other pages were considered and dismissed.
Position is not importance. The numbers shown are the order in the model’s source list. I have not established that position 0 mattered more than position 36.
The study lasted five days, with data collected between 2 and 6 September 2026. There is no measure here of how much these patterns vary from week to week.
I only included four models, with more time I would included Perplexity and Google AI Overview.
European publishers were assessed in English, which inevitably favours those with strong English-language editions and English explanations of themselves.
The shorter-query sample is smaller. There are 912 cases across 64 publishers, compared with 1,760 across 80 for the longer-question set. Thirty-one publishers clear the threshold for equal weighting in the shorter-query sample.
Crawler access is a confound we did not control. Publishers that block the crawlers used here cannot contribute the same evidence mix. The BBC appears through a downloads subdomain and its careers site because its main domain was closed to the crawlers in this setup.
What I would change next time
Extend the validation sample and hand-code it. The 452-URL check described above was automated and produced a confident classification for 195 pages. Hand-coding a larger sample, including paywalled pages that the crawler could not access, would make the 77.1% agreement figure more robust and give me greater confidence in the results for individual categories, particularly About and ownership pages.
Fix the classifier where the validation says it is weak. Four-fifths of unclassified pages are editorial articles, so the rule that files long or dated URLs as journalism is too narrow. Rewriting it against the validation set would shrink the unclassified bucket materially.
Separate cited from used. Distinguish between sources referenced directly in the answer and those merely listed in the source panel. That should give a better indication of which evidence is actually supporting the response.
Run it repeatedly. Ask the same questions weekly for a quarter so every figure carries a range rather than a single value, and publisher changes can be observed over time.
Run a real experiment. Identify publishers without a particular evidence asset, such as an ethics page or impact page, observe what happens if one is introduced and measure any subsequent change. Nothing here establishes causation, only patterns.
Ask in the local language. Put French questions in French, German questions in German and so on. That would separate “legible to an English-language AI” from “strong in its home market”.
Add the missing surfaces. Google AI Overview above all, followed by Perplexity.
Weight by real question demand. Every theme currently counts equally. Real reader demand is unlikely to be evenly distributed across identity, trust, value and subject questions.
Of course, what really needs to happen here is to run a real experiment and track the results of the recommendations in my research.
And if you’d like someone to help with that, just please reach out. 😀
In summary
So there you have it.
Nerd tickle satisfied.
The useful conclusion, for me, is not that About pages are useless. It is that AI does not appear to form one fixed view of a publisher from one definitive page.
Instead, it assembles evidence around the claim it is being asked to judge.
When the question is trust, the citations shift towards journalism, standards and accountability.
When the question is value, subscription and pricing evidence becomes much more important.
When the question is subject authority, journalism itself becomes central.
When the question is identity, the About page finally earns its money.
That makes the challenge much bigger than optimising a couple of corporate pages.
It means understanding the evidence that supports the things you want audiences to believe about you, where that evidence lives, who owns it and whether it tells a consistent story.
And that is even before we get to third-party citations, which are a massive part of the picture and a subject for another day.
I was really responding here to the “quick fixers”.
There isn’t a quick fix.
AI visibility for publishers increasingly looks like the cumulative result of journalism, site architecture, corporate evidence, commercial information, accountability and consistency.
That is probably less exciting than adding some magic schema markup, but considerably more useful.









