Web scraping has become an increasingly vital tool for businesses looking to extract valuable public data from websites. However, many site owners employ various anti-scraping techniques to prevent bots from harvesting their data. As a web scraper, it‘s important to understand these common roadblocks and how to work around them legally and ethically.
In this article, we‘ll dive into five of the most prevalent anti-scraping measures you‘re likely to come across and share strategies for dealing with each of them effectively. From IP blocking to CAPTCHAs to dynamic content, we‘ve got you covered. Let‘s get started!
1. IP Address Blocking
One of the simplest and most common ways websites detect and block web scraping activity is by monitoring IP addresses. Site owners can easily see which IP addresses are making unusually frequent requests to their servers. Too many requests too quickly is a dead giveaway of automated scraping behavior.
If a particular IP hits a predetermined request limit, the website will typically block that address, either temporarily or permanently. Sometimes entire IP ranges get banned if they‘re associated with cloud hosting providers known to be used by web scrapers.
To avoid having your IP address blocked while scraping, try the following techniques:
Slow down your request speed. Introduce delays between requests to better simulate human browsing behavior. A few seconds is usually sufficient. Randomizing the length of the delay can help avoid tripping rate limit alarms.
Rotate your IP addresses frequently. Don‘t make too many requests from a single IP address. Proxy rotation services allow you spread requests across a pool of IP addresses. Rotating IPs from different geographical regions can make your traffic appear more organic.
Choose your proxies carefully. Shared proxies are often blocked sooner because multiple scrapers may be using the same IP addresses simultaneously. Semi-dedicated or private proxies reduce this risk. Residential proxies tend to be more trusted by websites than data center IPs.
Slowing down your bots and rotating IP addresses will go a long way towards evading IP-based blocking. However, websites are constantly coming up with more sophisticated techniques for sniffing out suspicious traffic, as we‘ll see.
2. CAPTCHAs – Completely Automated Public Turing tests to tell Computers and Humans Apart
You‘ve likely encountered CAPTCHAs countless times while browsing the web. Those annoying little tests asking you to identify crosswalks or fire hydrants in a grid of images. Or to decipher wobbly, distorted text snippets. Their goal is to filter out bots by forcing "human-like" cognitive tasks.
While CAPTCHAs can often be triggered by suspicious traffic patterns, some sites implement them by default on certain pages or actions. For example, logging into an account or submitting a web form may require passing a CAPTCHA challenge.
Dealing with CAPTCHAs while scraping can be tricky, but you have a few options:
Avoid them entirely. Carefully control your request volume and rate to avoid raising red flags that trigger CAPTCHAs. Use IP rotation, user agent spoofing, and other techniques to keep a low profile.
Outsource to CAPTCHA farms. Services like 2captcha and DeathByCaptcha employ low-wage workers to manually solve CAPTCHAs on behalf of their customers. For a small fee per challenge ($0.50-$3), you can pipe your CAPTCHAs to these services which return the solution to your scraper.
Automate solving with ML. While certainly not easy, it‘s possible to leverage machine learning to solve certain types of CAPTCHAs with a reasonably high accuracy. Open source libraries exist for solving simple text-based CAPTCHAs. Image classification tools like TensorFlow can theoretically be trained to identify common CAPTCHA elements. However, implementation is not for the faint of heart.
In general, your goal should be avoiding CAPTCHAs altogether by flying under the radar. Triggering them means your other stealth techniques are failing.
3. Login Walls – Blocking Access to Unauthenticated Users
Some websites make scraping extra difficult by putting valuable content behind login walls. News sites with paywalls are a classic example. So are social networks or forums that only allow logged-in members to view certain pages.
Fortunately, login walls can often be bypassed if you‘re willing to jump through a few hoops. The most common approaches include:
Simulating the login process. Most login forms simply send a POST request with an email/username and password to an authentication endpoint, which returns a session token. You can automate this process by collecting valid login credentials (either your own or via an account creation script) and passing the right headers and parameters to simulate a real user logging in.
Reusing session tokens. Once you‘ve automated the login process, you can usually reuse the same authenticated session for a period of time (from minutes to days, depending on the site). Scrapers can serialize and reuse these session tokens to avoid re-logging in on every run. Be sure to monitor for session timeouts and re-authenticate as needed.
Using browser profiles. For tricker sites that perform extensive fingerprinting and validation on login requests, you may need to use a fully fledged browser environment. Tools like Puppeteer or Selenium allow you to automate interactions with real browser sessions. You can log into the site manually and save the authenticated browser profile to disk. Your scraper can then load the profile and reuse the logged in session.
While it‘s often possible to bypass login walls with these techniques, be aware that you may still hit a "paywall" that limits how much content you can scrape. Many sites only allow logged-in users to access a certain number of pages or resources before requiring a paid subscription. At that point, your only option is to pony up for an account or look elsewhere for the data you need.
4. User Agent Sniffing – Examining Browser Fingerprints
Besides tracking IP addresses and sessions, many websites examine the "user agent" string sent by your scraper as an additional authentication factor. The user agent provides details about the client application making the request – typically things like the browser name, operating system, and version numbers.
Servers use this information to tailor content to specific browsers and collect analytics on their traffic sources. The default user agents sent by scraping frameworks like Python‘s requests library often look very different from standard web browsers. For example:
python-requests/2.22.0
Seeing an unfamiliar user agent is a big hint to website operators that the traffic is coming from a bot rather than a real user. That‘s why it‘s crucial to spoof your user agent to mimic a real web browser. For example:
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.74 Safari/537.36 Edg/79.0.309.43
There are countless possible user agent combinations, depending on the browser, rendering engine, operating system, and other variables. It‘s a good idea to choose a relatively common one to avoid standing out.
However, spoofing your user agent once isn‘t enough. Anti-scraping scripts can profile traffic coming from a particular user agent over time. If the traffic patterns don‘t match those expected from human users of that user agent, red flags will be raised.
So in addition to spoofing legitimate user agent strings, it‘s important to rotate them regularly. Don‘t make too many requests from the same user agent fingerprint. Vary things up by choosing different combinations of browsers, operating systems, and version numbers.
While examining user agents alone is not a very robust anti-scraping strategy, it‘s often used as part of a multi-pronged approach in combination with IP tracking, rate limiting, and other signals. Varying your user agent is one more way to keep your scraper from being easily identified.
5. Dynamic Content Loading – AJAX and XHR
As web applications have grown richer and more complex, there‘s been a significant shift away from rendering HTML server-side to client-side rendering with JavaScript. Rather than getting a simple HTML page on the initial request, many modern websites load a minimal page template and then fetch content dynamically from API endpoints using techniques like AJAX (Asynchronous JavaScript and XML).
This presents some unique challenges for web scrapers, since the data they‘re after is often not present in the initial HTML response. Instead, that data gets loaded asynchronously by JavaScript after the page loads.
To scrape these modern "single page apps", you‘ll need to level up your scraping toolkit. Some strategies to consider:
Use a real browser engine. Tools like Puppeteer and Selenium allow you to automate real browsers like Chrome and Firefox. With a real browser, the JavaScript will run and dynamic content will load just like for human visitor. You can wait for elements to appear on the page before extracting them. The downsides are that browser automation is much slower and more resource intensive than sending simple HTTP requests.
Reverse engineer the API calls. One elegant solution is to open your browser‘s network tab and monitor the HTTP requests that get sent as the page loads. You can usually identify the API endpoints that return the data you‘re after. By replicating those requests directly from your scraper, you can access the raw structured data while bypassing the HTML entirely!
Identify URL patterns for dynamic parameters. While modern apps tend to update elements on the page without a full refresh, they often push updates to the URL string for bookmarking and navigation purposes. By understanding the patterns for these dynamically loaded URL paths and query parameters, you can sometimes uncover ways to jump directly to certain pages or states.
While scraping JavaScript-heavy websites is undoubtedly harder than old school static HTML, it‘s still very much possible. You may need to get creative and combine multiple approaches depending on the quirks of each particular target site.
Other Anti-Scraping Measures to Look Out For
Beyond the major anti-scraping approaches covered above, there are countless other smaller techniques that websites use to put up speed bumps for scrapers. While we won‘t dive into them deeply, here are a few more to have on your radar:
Honeypot links. Websites sometimes include hidden links that are invisible to normal website visitors. However, naïve scrapers may still find and follow them. These links act as traps, signaling to the site that the visitor is a bot.
Rate limiting. Many websites implement global rate limits for their pages and API endpoints. Too many requests from a single client in a short time window will get blocked. Use caching, proxy rotation, and respect rate limits to avoid hogging server resources and drawing unwanted attention to your scrapers.
Bot detection scripts. There are numerous commercial "bot detection" scripts like Distil and Imperva that combine many of the techniques discussed here into a single service. These tools fingerprint scraper traffic based on IP, user agent, and behavioral signals. Some can even run JavaScript challenges and examine mouse movements or other signs of "human-like" behavior.
The Web Scraping Arms Race
Web scraping is a constant cat-and-mouse game between site operators and data harvesters. As scrapers come up with new techniques to blend in with human traffic and parse even the most convoluted web apps, websites respond with ever more sophisticated ways to detect and block bots.
To stay ahead of the curve, web scrapers must constantly iterate on their stealth techniques. Using a combination of the strategies outlined in this article will go a long way towards bypassing the most common anti-scraping defenses. However, some websites are so heavily fortified that it becomes impractical to scrape them sustainably without getting blocked sooner or later.
As with any web scraping project, it‘s critical to be ethical and respect the website owner‘s wishes. Scrape only public data, never overwhelm a site with requests, and obey the robots.txt file. With the right techniques and tools – and a healthy dose of finesse – you can usually find a way to get the data you‘re after while staying on the right side of the web scraping arms race.