web.archive.org is also explicitly setup to thwart page harvesters.
I suppose archive.org don't like to be drowned in page-request and starts to say no if it detects that.... so set Max connections/second to say 0.1 or even slower and set it to Pause after downloading 2000000 bytes, and of course set the Spider to care about no robots.txt rules, and set another Browser ID should help, or?