@lightpohl They want a copy of all the data. For orgs where the goal is to give anyone all the data they might want for any reason, usually there will be a bulk download file (big cached, compressed file). Orgs don't offer these, or if they do, the crawlers ignore them and just point the same stupid spider doing a naive crawl. Stackoverflow at one time served the google crawler a different version of the site for similar reasons.
Post
July 30, 2024