Web scraping is an incredibly powerful technique for extracting data from websites at scale. Whether you‘re a marketer analyzing competitor pricing, a data scientist gathering training data, or a hobbyist building your own product database, the ability to automatically collect structured web data opens up a world of possibilities.
However, many websites are understandably protective of their data and go to great lengths to detect and block web scrapers. Getting blocked while scraping is incredibly frustrating, ruining hours of work and potentially resulting in permanent IP bans.
In this guide, I‘ll share proven tips and techniques to help you scrape websites more effectively while avoiding detection. With the right approach, you can gather the web data you need without triggering anti-bot systems or overwhelming servers. Let‘s dive in!
Why Websites Block Scrapers
Before we get into scraping tactics, it‘s important to understand why websites attempt to block scrapers in the first place. The two main reasons are:
- To prevent servers from being overloaded with requests
- To protect valuable data and content
When a scraper indiscriminately rapid-fires requests at a website, it can dramatically increase the load on that site‘s servers. Too many concurrent requests can overwhelm a server, slowing it to a crawl or even knocking it offline entirely. No website owner wants their site taken down by an overly aggressive scraper.
Additionally, the data on many websites is proprietary and considered a competitive advantage. E-commerce sites, for example, don‘t want competitors easily scraping their product info and undercutting their prices. Blocking suspicious traffic helps keep valuable data from falling into the wrong hands.
So while it may be frustrating to face anti-scraping measures, it‘s important to respect a website‘s desire to protect itself. The key is finding ways to collect the data you need without adversely impacting the site or accessing info you shouldn‘t. Here‘s how.
Slow Down Your Scraping
One of the biggest giveaways that you‘re a bot and not a human is scraping a website too quickly. A normal user might view a few pages per minute, but an unthrottled scraper can easily make hundreds or even thousands of requests in the same timeframe.
Anti-bot systems look for this kind of abnormally fast traffic and flag it as suspicious. The solution is to drastically slow down your scraping rate to more closely mimic human behavior. Here are a few tips:
- Add a delay of a few seconds between each request
- Randomize the delay a bit so the timing isn‘t too consistent
- Limit concurrent requests to 1-2 at most
- Throttle your scraping speed to a few requests per second, max
Yes, this means your scraper will run much more slowly. But it‘s better to scrape slowly than not at all! If a site detects an abnormal spike in traffic from your IP, you risk having your access cut off completely.
Use Proxies to Rotate IP Addresses
Another major red flag is when a site sees a large number of requests all coming from a single IP address. Even if you slow your scraping rate to mimic human behavior, making thousands of requests from one IP will still get you noticed and likely banned.
The solution is to spread your requests across many different IP addresses using proxies. A proxy server acts as an intermediary, routing your requests through a different IP address to mask your true location.
Of course, making all your requests through a single proxy is little better than using your real IP. You need a large pool of proxies to choose from and rotate through them, ideally at random, for each new request you make.
Some popular proxy services for web scraping include:
- Bright Data (formerly Luminati) – P2P network with over 72 million IPs
- IPRoyal – Ethically-sourced datacenter and residential proxies
- Proxy-Seller – Affordable P2P residential proxies
- SOAX – Flexible proxy service with millions of IPs
- Smartproxy – High-quality residential proxies
- Proxy-Cheap – Bulk IPv4 proxies at rock-bottom prices
- HydraProxy – Easy-to-use proxy API and browser
Vary Your Scraping Patterns
Humans are highly unpredictable when browsing the web. We might hop from one page to another, use the back button, open multiple tabs, leave a site entirely and return later, etc.
Scrapers, on the other hand, tend to follow rigid, repetitive patterns. They might crawl every page in sequential order or make the same requests over and over.
Varying these patterns is key to avoiding detection. Some techniques to appear more human:
- Randomize the pages and order you scrape
- Adjust your crawl depth and breadth for different sites
- Throw in occasional random actions like pausing or reloading a page
- Mimic human actions by clicking buttons or filling out forms
- Mix in requests to static resources like CSS/JS/image files
Of course, you don‘t want to randomize things so much that you break the actual scraping logic. But adding some unpredictability to your scraper can go a long way in making it seem more human.
Rotate User-Agents & Headers
When a normal web browser makes a request, it includes information in the request headers about the client, OS, and more in a string called the user-agent.
The user-agent allows sites to serve content tailored to a specific device or browser. But it‘s also an easy way for sites to spot scrapers—if a site sees thousands of requests all with identical user-agents, that‘s a dead giveaway it‘s being scraped.
Fortunately, it‘s trivial to spoof your user-agent and pretend to be a different browser. Even better, you can rotate through a list of user-agents for each request, which makes you look even more like normal traffic from a variety of visitors.
There are big lists of common user agents you can download and use in your scraper. You can also grab real user agents from live web traffic using a tool like fake-useragent.
In addition to rotating user agents, you should spoof other request headers as well. Things like Accept-Language, Referer, and DNT can all help make your requests seem more authentic.
Watch Out for Honeypot Traps
Some sneaky websites will create invisible links or honeypot pages that aren‘t visible to normal users but that a scraper might accidentally crawl. If you follow one of these honeypot links, the jig is up—the site knows you‘re a bot.
To avoid falling for honeypots, be very precise about the links you choose to crawl. Inspect the page source and look for any suspicious links that might be honeypots. You can also check the CSS to see if links are being hidden with display: none or visibility: hidden.
Configure your link extractor to only follow URLs matching a whitelist pattern, and avoid crawling any URL that looks machine-generated or otherwise fishy. A bit of caution when crawling links goes a long way.
Use Headless Browsers for Dynamic Pages
So far we‘ve focused on scraping relatively simple static web pages. But modern websites are increasingly dynamic, using JavaScript to load content, react to user input, and more.
Scrapers that simply make HTTP requests and parse the HTML responses will have trouble with these dynamic sites, since much of the content is loaded asynchronously by JavaScript after the initial page load.
One solution is to use a real web browser like Chrome or Firefox to load pages, allowing the JavaScript to run before scraping. You can automate these browsers with tools like Selenium or Puppeteer.
For production scraping, I recommend using a headless browser instead of a full GUI browser. Headless browsers are stripped-down browsers with no graphical interface, perfect for automated scraping.
The two most popular headless browsers are PhantomJS (now deprecated) and Puppeteer/Headless Chrome. With these tools, you can fully render pages, interact with UI elements, and scrape content from even the most complex JavaScript-driven sites.
Monitor for CAPTCHAs, Rate Limits & Bans
Even with all these techniques, your scrapers will sometimes get caught by anti-bot defenses like CAPTCHAs, rate limiting, and IP bans. It‘s important to monitor your scrapers for these issues and have fallback mechanisms in place.
Some sites will throw up a CAPTCHA if they detect suspicious traffic. You can try to solve these with automated CAPTCHA-solving services, but it‘s often better to just back off and try again later with more stealth measures in place.
Rate limiting is another common anti-bot measure where a site will start timing out requests or serving 429 Too Many Requests errors if you make too many requests too quickly. Again, the solution is to slow down and spread your requests out more.
Finally, some websites will ban offending IPs outright. If you find yourself blocked from a site, stop scraping immediately. Give it a rest, then try again later with a new IP and more careful techniques. Pushing a banned scraper will just get you blocked more severely.
Tips for Ethical & Effective Scraping
We‘ve covered a lot of ground in this guide to stealthy web scraping! As you put these techniques into practice, here are a few final tips to keep in mind:
- Respect robots.txt files and terms of service. Don‘t scrape any data a site has explicitly asked you not to.
- Don‘t overwhelm sites with too many requests. Use delays, rate limiting, and proxies to lighten the load.
- Spread scraping tasks across multiple machines for greater concurrency without increasing the load on any one server.
- Use pre-built scraping tools and frameworks when possible to save development time. Only roll your own code when absolutely necessary.
- Consider the legal implications of scraping. In general, scraping publicly available data for personal use is ok, but scraping private data or using it commercially may be a grey area.
With the right techniques and some common sense, you can gather valuable web data without negatively impacting the sites you scrape. Go forth and scrape responsibly!