Log File Analysis for Indian Ecommerce: What Googlebot Actually Crawls
Most SEO decisions on large Indian ecommerce sites are made on assumptions about crawling. The team believes Googlebot is reaching the category pages, believes the parameter URLs are being ignored, believes the new collection launched last month has been discovered. Log files turn all of that from belief into fact, and the facts are usually unpleasant.
A server log records every request: which user agent, which URL, what status code, when. Filter it to verified Googlebot and you have an exact record of where your crawl budget went.
What the first look usually shows
Three findings come up again and again on Indian catalogue sites.
Most of the budget is going to URLs you do not want indexed. Faceted filters, sort parameters, session IDs, pagination beyond page three, internal search results. It is common to find 60-75% of Googlebot's requests landing on parameter combinations, while the actual category pages get crawled once a fortnight. Every one of those wasted requests is a request not spent on a product page.
Deep products are effectively invisible. Group the crawl data by click depth from the homepage. On a typical large catalogue, pages at depth two get crawled weekly; pages at depth five get crawled every few months, if at all. If a product sits four clicks deep and the logs show no Googlebot hit in ninety days, it is not underperforming – it has not been looked at.
Response codes tell a story nobody was tracking. A steady trickle of 5xx errors during peak traffic hours, redirect chains three and four hops long left over from a migration, and a set of 404s that are still being crawled months after the URLs were removed because something still links to them.
How to run it
Get thirty days of raw access logs from the server or CDN. Verify Googlebot properly by reverse DNS – a meaningful share of traffic claiming to be Googlebot is not. Then join the log data against a crawl of the site so every URL carries its depth, its status, its template type and whether it appears in the sitemap.
Four questions to answer first. What percentage of crawl requests hit indexable pages versus parameters and junk? What is the crawl frequency by template – homepage, category, product, blog? Which indexable URLs received zero Googlebot requests in the period? And which non-indexable URLs received the most?
That last query, sorted descending, is your fix list. It is usually short and the top five entries usually account for most of the waste.
The fixes, in the order they pay
Block the parameter patterns that produce no unique content in robots.txt – and block the pattern, not the individual URLs, because the combinatorics are endless. Set canonical tags on filtered views pointing to the clean category URL, then check the logs a fortnight later to confirm the crawl actually moved.
Flatten the depth of anything that matters. A product four clicks from the homepage should be reachable in two, which usually means better category cross-linking and genuinely useful hub pages rather than a bigger footer menu.
Collapse redirect chains to single hops. Every extra hop is a wasted request and a small loss of signal.
Fix the 5xx errors. Nothing suppresses crawl rate faster than a server that intermittently fails – Googlebot backs off, and it does not come back quickly.
Then measure it again
The point of log analysis is that it is repeatable. Run it again after six weeks and compare the same four numbers. The percentage of requests hitting indexable pages should be up. The zero-crawl list should be shorter. If neither moved, the fix did not take, and you know that in six weeks rather than in a year.
This is also the diagnostic that explains most cases of “we published it and Google ignored it”. Usually Google did not ignore it. Google never reached it.
Teams that want the log pull, the crawl join and the fix list produced as one piece of work can get it done by Deep Bhardwaj, who will normally ask for the CDN logs before anything else. The indexing problems on a large catalogue are almost always upstream of the content.