How AI Assistants Decide Which University to Recommend
In brief: AI assistants retrieve against searches a student never typed, and only 12% of the URLs they cite actually rank in Google’s top 10 for the same query — so a first-page ranking buys you less inside the assistants than your dashboard suggests. None of twenty Canadian university robots.txt files block the AI crawlers, so if you’re not being cited, it isn’t because you shut the door. This piece covers what’s documented, what’s inference, and three things you can change without a new budget line.
Your best program page took four months. New photography, outcomes data, four rounds of faculty review. Then a 17-year-old asks ChatGPT which schools in the province run that program, and it names three schools, a licensing body and a forum thread. You’re not in it.
How these systems choose what to cite is part documented, part guesswork.
How do AI assistants choose which schools to name?
Nobody outside the labs can see the ranking function, and anyone claiming to have reverse-engineered it is selling something. What is documented is the retrieval step.
Google’s Search Central documentation says AI Overviews and AI Mode “may use a ‘query fan-out’ technique — issuing multiple related searches across subtopics and data sources.” So one typed question becomes a set of searches the student never typed: cost, admission average, licensing outcome, what people say about the co-op placement.
There is also a gap between what gets cited and what ranks. Ahrefs ran 15,000 queries through ChatGPT, Gemini, Copilot and Perplexity and published the results in August 2025: only 12% of the URLs the assistants cited ranked in Google’s top 10 for the same query. Perplexity was the outlier at 28.6%, the rest sat near 8%, and in an earlier Ahrefs study, Google’s own AI Overviews overlapped with the top 10 at 76%.
AI Overviews behave roughly like search. The standalone assistants do not, so a first-page ranking buys you less there than your dashboard implies. My read is that they retrieve against the sub-questions rather than the one the student typed, which is inference from the fan-out documentation, not something any vendor has confirmed.
Why doesn’t my program page show up in ChatGPT?
Because most program pages contain no sentence that can be lifted as an answer.
Here’s the genre: “Our graduates leave prepared to thrive in a rapidly changing profession.” No number, no date, no regulator, no credential name. A system hunting for a passage that answers “how long is the diploma and does it lead to licensure” will take the licensing body’s page instead, or a provincial database, or a thread where four graduates argue about it.
The sentence that would get cited by AI — the one carrying the averages, the placement rates, the specific data — is usually the sentence that gets softened in the copy review and editing process. Committee editing is very good at removing the exact material a machine needs. That is my read, not a study finding.
Forums win the trust-heavy questions for a separate set of reasons, which I went through in Why Reddit Outranks Your Admissions Office.
Does adding llms.txt or schema markup get you into AI answers?
Not on the 2026 evidence, and this is where a lot of budget is about to go.
Google’s documentation is blunt: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features,” and “There’s also no special schema.org structured data that you need to add.”
llms.txt has a study behind it now. Ahrefs looked at 137,210 domains in May 2026 and found roughly 38,000 with a valid file. 97% of those files were never requested once that month. Structured data still earns its keep for rich results in ordinary search, and ordinary search is what feeds AI Overviews at that 76% overlap, so keep it. Just don’t buy it a second time under a new name.
What can a university change to get cited by AI?
Three things, in the order I’d do them.
Start with what you are blocking, because the answer is probably nothing. On 24 August 2026 I pulled the robots.txt of twenty public Canadian universities and colleges. Sixteen responded, and not one of them disallows GPTBot or OAI-SearchBot. Most do not mention AI crawlers at all. The blanket AI blocks that went up across publishers in 2023 and 2024 did not happen here. If your pages are not being cited, it is not because you shut the door. Confirm your own file anyway, and know which bot is which: GPTBot crawls content that may be used to train models, while OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features.” Blocking the first does not keep you out of the second.
Publish the answer with the number in it. For each program that matters, one page should state the credential, the length, the cost, the admission range you really use, and what the graduate is eligible to do, in sentences that survive being quoted out of context.
Then look outside your own site. Directories, provincial databases, licensing bodies, Wikipedia and forums are all in the corpus, and where they disagree with you the model picks whichever it trusts more. Correcting a stale database entry moves more than a homepage refresh.
How do I check which sources AI trusts about my programs?
Take the ten questions your inquiry inbox gets most, run them through three AI assistants in fresh sessions, and write down every domain that comes back, including the answers where you were never mentioned. That list is what the machine is reading when someone asks about your programs.
One caution: the Ahrefs overlap figures were measured in 2025 and retrieval changes without announcement, so re-run the check every couple of months rather than buying a permanent fix. What ChatGPT Tells a 17-Year-Old About Your University has the question list I start from, and the AI Search Readiness Audit itself sits under services.
Finding out what AI assistants already say about your programs, and fixing what’s citable, is part of the AI Search Readiness Audit.
Book a discovery call →