By Sara Khan · Last updated August 2026.

Somewhere in Lahore this week, a marketing manager is signing off an invoice for an llms.txt file on the promise that it will help the brand show up inside ChatGPT.

The belief is reasonable on its face, because the file sounds like exactly the kind of standard the modern web runs on, modeled as it is on robots.txt. The problem is that the analogy is the entire argument, and the analogy is wrong. John Mueller, a search advocate at Google, has stated publicly that no AI system currently uses llms.txt and that the absence is obvious to anyone who inspects their own server logs. The consumer chatbots that Pakistani brands actually want traffic from, ChatGPT, Perplexity, Claude, and Google AI Overviews, fetch individual pages when a person asks a question, and none of them fetch the llms.txt file first. The deliverable being invoiced has, in the most literal sense, no documented reader. A file with no reader is a courtesy, not a strategy.

It is like handing a neatly printed price list to a shopkeeper at Liberty Market who has already decided what to charge you; the list changes nothing, because nobody consults it before the decision gets made. The file exists, the file gets crawled, and the file gets indexed, and not one of those facts is evidence that the file does anything at all.

The four proofs that prove nothing

The case for llms.txt rests on four observations that practitioners repeat inside proposals and decks: the bots crawled it, Google indexed it, a large language model repeated it, and ChatGPT endorsed it. Each observation is real, and each is also exactly what you would expect to see if the file did nothing. That is the trap, and the trap is the whole product.

Search Engine Journal demonstrated the problem cleanly by inventing cats.txt, a plain-text standard for declaring office cats, their job titles, their breeds, and a mandatory purr metric scored out of ten. The publication wrote a specification, seeded it with earnest-sounding text, and ran it through the same four tests the industry uses to certify llms.txt. Cats.txt was crawled by the AI bots, indexed by Google, repeated by a language model, and confirmed at length by ChatGPT as helpful for ranking, a file about a fictional Tuxedo cat named Odd that exists nowhere but the file. It cleared the entire evidentiary bar. The underlying mechanic is not that llms.txt works or fails; the underlying mechanic is that the standard of proof being applied would certify a joke.

A crawler fetching a file and a model relying on that file are different events separated by an entire inference the industry has decided to skip. Conflating the two is the intellectual equivalent of concluding that your umbrella causes the rain to stop, because every time you put the umbrella away the rain does eventually stop. Pakistani brands are not paying for a result; they are paying for a set of observations that would occur whether or not any result existed, and the agency invoicing them cannot tell the difference because the tests it runs cannot tell the difference.

Infographic: Infographic-style horizontal bar chart titled 'What correlates with AI visibility' showing branded web mentions 0.68 and

What noindex could not do, robots.txt only partly does

The instinct, once a brand suspects the file is theater, is to reach for stronger controls, to noindex the pages it does not want surfaced, or to block the bots outright in robots.txt. Seer Interactive’s audit work shows why that instinct also runs into a wall, and the wall is structural rather than tactical.

A crawler has to fetch a page before it can read a noindex tag, which means that by the time the bot sees the instruction it has already ingested the content. Noindex was built for cooperative search engines that agreed to honor the signal, and large language models have not made the same commitment, so the tag functions as a legacy mechanism without a widely adopted replacement. Robots.txt operates one step upstream, so a well-behaved crawler that checks the file will never read the page at all, which is meaningfully better than noindex. The qualifier is doing all the work in that sentence, and the qualifier is well-behaved.

Training crawlers and AI search crawlers will generally stay out when you add explicit Disallow directives, and that is low-effort and worth doing, particularly if paid landing pages live on a subdomain you can scope without touching the main site. User-initiated retrieval is a different category entirely. When a real person in Karachi asks ChatGPT a question and the bot fetches live content in that moment, robots.txt will not reliably stop it; OpenAI removed ChatGPT-User compliance language from its documentation, and Perplexity has described Perplexity-User as an agent and therefore exempt. The only reliable control for that tier is server-level blocking at the WAF or Cloudflare layer, a heavier lift that most Pakistani SMEs do not need and should not be sold. The practical decision is narrower than the pitch: block the training crawlers if you care, ignore the rest unless monitoring shows a real problem, and stop treating a text file as a strategy.

Infographic: Infographic-style checklist titled 'The four proofs that prove nothing' with four checked rows (Crawled, Indexed, LLM re

The citation data that explains why a file cannot save you

Ready to improve your marketing results?

Book a free strategy call - we'll audit your current setup and identify the highest-impact fixes.

Book Free Call

Even if llms.txt worked, it could not solve the problem Pakistani brands are paying it to solve, because that problem is not a file problem at all. Kevin Indig’s first-half 2026 research found that roughly ninety-one percent of AI citations appear in only one of ChatGPT, Perplexity, or AI Overviews, and almost never in more than one. A brand that optimizes for a single answer in a single engine is, by construction, invisible to the other engines for the same query, and no text file changes that geometry. Google’s own Search Console data is reportedly about seventy-five percent incomplete for this landscape, which means even careful practitioners are measuring from a partial picture and then paying for a file to improve a number they cannot fully see. The consequence is plain: a Lahore brand can be cited inside ChatGPT for its category and entirely absent from Perplexity for the identical question, and a file does nothing to reconcile the two.

Model choice has also become a genuine business risk rather than a preference, which makes single-engine thinking more fragile, not less. ChatGPT’s share of agent usage slid from seventy-eight percent in July 2025 to fifty-six percent a year later, while Gemini climbed from fifteen to thirty and Claude grew from two to ten. A Pakistani brand that “won” ChatGPT in 2025 may have lost close to half its audience share by 2026 without anything changing on its own website. Buyers now check an average of 2.4 platforms before validating a purchase, which is a concrete, surveyable proxy for the influence that should replace the ranking on the dashboard. Visibility has to be built across a panel of models rather than chased inside one, and a file cannot populate a panel.

The pattern repeats wherever you look. The unit of measurement has moved from the single ranking to the panel of prompts, and the deliverable being sold has not moved with it. Pakistani businesses that pay for a file are buying a 2020 artifact to solve a 2026 problem, and the gap between the two is where the budget leaks.

Where the budget is actually going

The spending tells the story of an industry that has decided to act before it understands, and the numbers are starker than the rhetoric. Fractl’s survey of marketers found that teams are routing about twenty-four percent of search or content budgets to AI visibility, that eighty-two percent have allocated at least something, and that eighteen percent have allocated nothing at all. Digiday’s reporting adds the harder figure: sixty-seven percent of brands say they appear less frequently in AI-generated answers than they would like, which is the quiet admission that all of this spend is not yet producing the outcome being sold. Ahrefs ran the file question across one hundred thousand domains and found that llms.txt is, in practice, largely ignored by the very crawlers it is meant to court, a finding echoed by other large studies showing no measurable citation advantage for sites that add one.

What this means for a Pakistani SME is that the money routed toward file implementation is money not spent on the two activities the same research shows actually correlate with visibility. Branded web mentions and YouTube impressions correlate with AI visibility in the 0.50 to 0.74 range, while backlink count and ad spend sit below 0.30, a reallocation signal away from the tactics being invoiced and toward earned coverage and original video. The honest reading of the data is that brands should spend less on files and more on the original research, the proprietary data, and the named expertise that an AI system cannot generate for itself.

There is a useful counterweight to the file economy in the broader research on brand equity. A WARC study conducted by the agency Charlie Oscar estimated that sixty-three percent of a brand’s visibility was attributable to long-term brand equity and only twenty-six percent to current marketing activity. The implication for a Pakistani brand is that the durable lever is the brand itself, the reputation, the named people, the original work, and not a text file bolted onto the root domain on a Tuesday afternoon.

The honest alternative to a text file

If a file is not the answer, the answer is measurement, and specifically the kind of measurement that treats AI visibility the way public relations has treated share of voice for a decade. AMEC, the body behind the Barcelona Principles, released seven GEO principles that organize the work into upstream reputation, search and content readiness, and downstream output tracking across presence, framing, citations, and accuracy. That is a dashboard, not a file, and it is what allows a Pakistani brand to see whether it is being mentioned, recommended, or ignored across ChatGPT, Perplexity, and AI Overviews in the same week rather than once a quarter. Search Engine Journal’s AEO reporting adds the tactical spine: YouTube is the single most cited source across major language models, the average AI prompt runs twenty-three words rather than three or four, and the models cite Reddit only when no brand has produced a better answer. Those are inputs a brand can act on. A text file is not.

WeProms Digital approaches this the way an analytics team would, not the way a file vendor would, and its AI search visibility monitoring and citation tracking work is built to measure real mentions across a panel of models rather than to invoice for a standard no engine reads. The difference matters because a citation panel tells a brand exactly which prompt it lost and to whom, while a file tells it nothing at all. Pakistani brands that want to know where they actually appear should read the field note on where Pakistani brands show up in AI search and the deeper look at how content volume feeds citations before spending on another file.

A standard worth nothing is still priced like a service

See this in action

How we helped a Pakistani business achieve measurable results.

Read case study

The principle worth holding onto is simple, and it is worth stating without softening. An activity that produces no measurable change in a single answer inside a single engine is not a strategy; it is a line item. Pakistani brands are not failing at AI search because they lack an llms.txt file; they are failing because they are paying for artifacts instead of evidence, and because the evidence, real citations, real mentions, real recommendations across a panel of models, is harder to produce than a text file and therefore less often sold. Treat the file the way Google’s own advocate treats it: a free, harmless, optional courtesy that deserves five minutes of effort and zero rupees of budget. Spend the rest on the original research, the YouTube presence, and the citation monitoring that actually moves where a brand appears when a Pakistani buyer asks a question.

Read next: AI visibility tools that waste Pakistani SME budgets and how competitors win the ChatGPT answers Pakistani brands lose.

If your agency handed you an llms.txt file as a deliverable and could not show you where you appear across ChatGPT, Perplexity, and Google AI Overviews this month, that is the signal to change the conversation. WeProms Digital builds the measurement layer most Pakistani brands are missing, a citation panel, a freshness process, and a content plan tied to the prompts your buyers actually ask. Talk to the team at weproms.com/contact-us, email hello@weproms.com, or message WhatsApp +92 300 0133399.

Sources & References

  1. Search Engine Journal — How cats.txt showed llms.txt evidence is GEO astrology — August 7, 2026
  2. Search Engine Journal — AI’s impact is outrunning measurement: the trust and attribution gap — August 8, 2026
  3. Seer Interactive — Do LLMs respect robots.txt? Where robots.txt can and can’t block AI — August 7, 2026
  4. Search Engine Land — What six perspectives reveal about demand generation in AI search — August 7, 2026
  5. Digiday — By the numbers: How marketers are building the infrastructure for AI search — August 7, 2026
  6. Search Engine Journal — The AEO Playbook: How to get cited and stay visible — August 7, 2026
  7. Digiday — Clients are hungry for AI visibility aids, but buyers are skeptical — August 7, 2026
  8. Search Engine Land — How business context changes AI recommendations — August 7, 2026

Additional reading from industry feeds: