Skip to content

Research

Can AI read India's best-known brands?.

We checked 75 well-known Indian brand websites for what ChatGPT, Perplexity, Gemini and Google AI need to read and understand them. None block AI crawlers. Many make it hard anyway.

Updated 6 October 2026 · 5 min read

Key findings

On 6 October 2026 we ran our free AI visibility checker over the homepages of 75 well-known Indian brands: 15 each in fintech, B2B SaaS, D2C, consumer internet, and health and education. Here is what we found.

  • None of the 64 robots.txt files we could read block any AI crawler. Not ChatGPT's, Perplexity's, Google's, Anthropic's or Common Crawl's. Indian brands have chosen to be open to AI, unlike many news publishers.
  • One site in six hides its content from most AI crawlers. 9 of the 56 homepages we could read had fewer than 150 words of text before JavaScript runs, and 7 had fewer than 50. Most AI crawlers don't run JavaScript, so they see a nearly blank page.
  • One site in five has no structured data at all. 12 of 56 homepages had no schema.org markup, and 18 had nothing that describes the business as an organization.
  • Almost half already publish llms.txt. 26 of 56 sites had one, led by B2B SaaS (9 of 13) and fintech (7 of 14).
  • A quarter of the sites turned our checker away. 19 of 75 returned an error or didn't respond, mostly in D2C (8 of 15) and consumer internet (7 of 15). That doesn't prove they block AI crawlers, which often get special treatment, but it's worth checking.
  • Fintech and B2B SaaS lead. Among the sites we could read, median scores were 95 for B2B SaaS and 91 for fintech, against 88 for consumer internet and 90 for D2C, where several sites scored below 60.

What we checked

For each brand we read four public files, the same ones any crawler can fetch: the homepage, robots.txt, sitemap.xml and llms.txt. We didn't use any AI service, log in, or load any page other than these. The check identifies itself as RecvisBot and ran from India.

Each site gets a score out of 100. Access counts for 55 points: whether the homepage loads over HTTPS, whether robots.txt lets AI crawlers in, and how much text exists before JavaScript runs. Understanding counts for 45: title, meta description, main heading, structured data, linked official profiles, sitemap, llms.txt and link previews.

This measures whether AI systems can read and understand a website, not whether they recommend the brand. A site can score 100 and still be missing from AI answers, and a famous brand can be recommended despite a low score. It's one input among many.

Nobody is blocking AI crawlers

We checked robots.txt rules for ten crawlers. Five fetch pages for live answers: Googlebot (AI Overviews and AI Mode), Bingbot (Copilot), OAI-SearchBot and ChatGPT-User (ChatGPT), and PerplexityBot. Five collect training data: GPTBot, Google-Extended, ClaudeBot, CCBot and Applebot-Extended.

Across all 64 robots.txt files we could read, not one blocks any of them. Most don't mention AI crawlers at all, which leaves them allowed by default. That's a deliberate or default choice to be visible to AI, and it's the right one for brands that want to be recommended.

It also means access isn't the problem for these brands. The gaps are in what crawlers find once they arrive.

Content AI crawlers can't see

Google renders JavaScript before indexing a page, so a site built entirely in the browser can still rank in search. Most AI crawlers don't. They read the HTML the server sends, and if the content only appears after scripts run, they see a shell.

9 of the 56 homepages we could read had fewer than 150 words of text in that first HTML, and 7 had fewer than 50. Three of the nine were D2C stores and three were consumer apps, both sectors that tend to build their websites as single-page apps. For an AI assistant asked about these brands, the homepage says almost nothing about what they sell.

The fix: render key pages on the server, or pre-render them, so the text is in the HTML. Most modern frameworks do this with a setting rather than a rebuild.

Telling AI who you are

Once a crawler can read the page, it needs clear signals about who the business is. Results across the 56 readable homepages:

SignalSites with itShare
Sitemap53 of 5695%
Main heading (h1)47 of 5684%
Organization structured data38 of 5668%
Open Graph title and image38 of 5668%
Page title of 10–70 characters37 of 5666%
Official profiles linked in structured data (sameAs)35 of 5663%
Meta description of 50–170 characters35 of 5663%
llms.txt26 of 5646%

Structured data is the biggest gap. Organization markup with sameAs links tells an AI system, in a format it doesn't have to guess at, that this website, this LinkedIn page and this Wikipedia article are the same company. A third of these brands don't provide it.

llms.txt is a newer, optional standard: a plain-text summary of a business for AI tools. It's unproven as a ranking signal, but it's cheap to add, and almost half these brands already have one.

When bot protection turns everyone away

19 of the 75 websites didn't let our checker read the homepage: 9 refused it outright (HTTP 403), 5 said it was making too many requests (HTTP 429) even when we waited between attempts, and 5 never responded. We re-checked each one slowly before counting it.

We can't tell from this whether these sites also block AI crawlers. Firewall services often recognise and allow the big AI companies' crawlers while blocking unknown ones, and results changed with the software making the request. So we excluded these sites from the page-level figures above instead of scoring them zero.

It's still worth checking. Bot protection is usually set up to stop scrapers, and it's easy for the same rules to catch AI crawlers by accident. If you use Cloudflare, Akamai or a similar service, check that verified AI crawlers are allowed.

Scores by sector

SectorSitesReadableMedian scoreAverage score
B2B SaaS15139593
Fintech15149190
Health & education15149187
D2C brands1579077
Consumer internet1588877
All75569186

Scores cover only the sites we could read. B2B SaaS and fintech companies, many of which market to developers and other businesses, have the most complete technical setup. D2C and consumer internet have the widest spread: some perfect scores, some JavaScript-only homepages, and the most sites behind strict bot protection. 11 sites scored a perfect 100.

What to do about it

  1. Keep AI crawlers allowed in robots.txt, as every brand here does.
  2. Make sure your key pages have their text in the HTML, not only after JavaScript runs.
  3. Add Organization structured data with your logo and sameAs links to your official profiles.
  4. Write a title and meta description that say plainly what you sell and who it's for.
  5. Check that your bot protection allows verified AI crawlers.
  6. Consider adding an llms.txt summary.

You can run the same check on your own site in a few seconds with our free AI visibility checker. Whether AI assistants actually recommend you is a different question, and it's what our audit measures.

Limits of this study

We checked homepages only, once, from one location in India, on 6 October 2026. Websites change often, so any single result may already be out of date. Brands were chosen to cover five sectors with well-known names, not as a random sample, so the figures describe these 75 brands rather than Indian websites in general. We didn't contact the brands before publishing.

Read nextAI Search GuideHow AI assistants like ChatGPT, Gemini, Perplexity and Google AI Overviews find, judge and recommend brands, and how to measure where you stand.

Where does your brand stand?

Tell us your website. A founder will reply within one working day.

Talk to us