Friday, August 3, 2007

Clean Code Counts

This piece is taken from an interview by SearchEngineWatch with Googler Dan Crow who runs the Crawl Infrastructure Group. Google can't index the entire web so they are forced to make decisions on which pages they crawl and indexe. Certianly one limiting factor is Google's own infrastructure and crawled site inefficiencies further reduce the overall amount of the web that Google can crawl. So, does Google factor in how the actual code of your website is constructed in determining whether they will crawl it or crawl deeper pages inside the site? It appears to be the case:
What can we do to get more pages indexed? I've always suspected that streamlining HTML code is a good way to facilitate indexing. Reducing code bloat helps pages load faster and use less bandwidth. I asked if it would help to move JavaScript and CSS definitions to external files, and clean up tag soup. Dan's answer was refreshingly clear. "Those would be very good ideas," he said.

1 comment:

Anonymous said...

Good for people to know.