Is Web Crawling Legal? Well, It Depends

Web crawling, also known as web scraping, has become an increasingly popular technique for extracting data from websites. By using automated scripts or bots, web crawlers can quickly gather large amounts of information that would be impractical to collect manually. This has enabled a wide range of innovative uses across industries:

  • E-commerce businesses use web scraping to monitor competitor prices, optimize product listings, and analyze customer sentiment.
  • Marketing firms leverage scraped data for lead generation, brand monitoring, and market research.
  • Financial analysts use web crawling to extract data for investment models, economic forecasting, and risk assessment.
  • Researchers and academics scrape data to study social trends, online behaviors, and public health patterns.

However, the rising adoption of web crawling has also brought its legality into question. Is it legal to scrape data from websites without the owner‘s explicit permission? The short answer is – it depends. The legality of web crawling is a complex issue that is still evolving and can vary based on several key factors.

In this article, we‘ll take a deep dive into the murky legal waters surrounding web scraping. We‘ll examine what makes crawling legal or not, discuss relevant laws and legal precedent, and provide some best practices for staying on the right side of the law. Let‘s get started!

The Growing Popularity and Risks of Web Scraping

The use of web scraping has exploded in recent years as the amount of valuable data on the web continues to grow exponentially. According to a 2020 study by Openprise, 24% of companies now use web scraping for lead generation, up from just 12% in 2018.

Another survey by Oxylabs found that 52% of companies use web scraping for market research, 49% for competitor analysis, and 35% for pricing optimization. The same study revealed that 79% of senior managers believe web scraping has given their business a competitive edge.

However, this rising adoption has also led to increased legal scrutiny and risk. A report by Osterman Research found that 22% of companies have received cease-and-desist letters due to their web scraping activities. 33% have had their IP addresses blocked, and 25% have even faced legal action.

The financial risks can be significant – the average cost to defend against a scraping-related lawsuit is $500,000, not including any settlements or judgments. Some high-profile cases have resulted in millions of dollars in damages, such as the $60 million judgment against Bidder‘s Edge in 2000 for scraping eBay‘s auction listings.

So what determines whether a particular web crawling operation is likely to be considered legal or not? It boils down to three main factors:

1. How the crawling is performed

One of the most important considerations is what data is being accessed and how the crawler goes about collecting it. Generally speaking, scraping publicly available information is much less likely to raise legal issues than extracting data from private accounts or pages that require a login.

As attorney Aaron Rubin explains, "Courts have consistently held that scraping data that is publicly accessible and not protected by any kind of authentication is not a violation of the Computer Fraud and Abuse Act (CFAA)."

A 2017 case involving LinkedIn and hiQ Labs set a major precedent in this area. hiQ, an analytics company, was scraping publicly available LinkedIn profiles to offer services such as identifying employees at risk of being recruited away. LinkedIn sent a cease-and-desist letter and attempted to block hiQ‘s access, claiming it violated the CFAA.

However, the judge ruled in hiQ‘s favor, stating that the data was public and that barring the company would pose an existential threat to its business model. The case established that scraping public data likely does not violate the CFAA.

On the other hand, bypassing authentication or hacking into private accounts to scrape data would almost certainly be illegal under the CFAA and other laws. Even for public data, the way the crawling is carried out makes a difference.

Another key factor is the crawling rate and politeness toward the target website. Overly aggressive crawlers that slam servers with a high volume of requests may be viewed as a denial-of-service attack or trespassing on computing resources.

For example, in the case of eBay v. Bidder‘s Edge, the court ruled that the auction aggregator‘s web crawling activities constituted trespass to chattels because they caused harm to eBay‘s computer systems. Bidder‘s Edge was hitting eBay with 100,000 requests per day, about 1.5% of the site‘s total traffic.

Best practice is to insert delays between requests and limit the overall speed to avoid negatively impacting the website‘s performance. The Web Robots Pages (robots.txt) is a voluntary protocol that allows site owners to communicate their preferences for which parts of their site may be crawled and at what rate. However, violating the robots.txt alone is not necessarily illegal in the U.S., according to experts.

"There is no law that grants legal authority to a robots.txt file," says attorney Bennet Kelley. "While violating the robots.txt would be a factor in assessing whether access is authorized, it is not legally binding on its own."

Following a website‘s terms of service can also help keep crawling above board, but again, violating terms of service alone is not necessarily illegal. It can still result in getting banned and may be a factor in legal cases though.

2. How the scraped data is used

Another critical legal consideration is the purpose for collecting the data and what is ultimately done with it. Scraping for personal, non-commercial, or research uses is generally safer under fair use doctrines. However, using extracted data in a commercial product or service starts to push into riskier territory.

Fair use is a legal doctrine that permits limited use of copyrighted material without permission for purposes such as criticism, commentary, news reporting, teaching, scholarship, or research. To determine if something qualifies as fair use, courts weigh four main factors:

  1. The purpose and character of the use, including whether it is commercial or non-profit educational
  2. The nature of the copyrighted work (factual works get less protection than creative works)
  3. The amount and substantiality of the portion used in relation to the work as a whole
  4. The effect of the use on the potential market for or value of the original work

So while scraping data for personal use or academic study is more likely to be considered fair use, scraping a large portion of copyrighted content to resell or build a competitive commercial service is much riskier.

Publishing scraped data or using it to compete directly with the source website is much more likely to provoke a legal response. For example, in the Associated Press v. Meltwater case, the court ruled that a news clipping service that scraped AP articles was not protected by fair use. The service was effectively competing with AP by drawing away its subscribers and potential market.

Copyright law also comes into play if the extracted content itself is copyrighted. Facts alone cannot be copyrighted in most cases, but scraping content like articles, photos, or videos can infringe if used without a license or permission.

Attorney Nathaniel Fairfield from the Electronic Frontier Foundation explains, "Copyright protects creative expression. It does not protect facts that are arranged in a creative way. The classic example is that you can‘t copyright the phone book, even though you can copyright a creative arrangement of facts."

However, it‘s not always clear-cut what constitutes a "fact" vs. an "expression." In the infamous case of HiQ Labs. v. LinkedIn, one key question was whether publicly available user profiles were LinkedIn‘s copyrighted content or simply facts that could be freely scraped. The court ultimately decided scraping the public profiles did not violate LinkedIn‘s copyrights.

3. Applicable laws and precedent

Finally, the specific laws and legal precedent that apply based on jurisdiction are major determining factors for legality. In the United States, several key laws are commonly cited in web crawling cases:

The Computer Fraud and Abuse Act (CFAA) makes it illegal to intentionally access a computer without authorization. Originally designed as an anti-hacking law, it has been used in civil cases against web scrapers. However, courts have ruled both ways on whether scraping violates the CFAA, largely depending on the specific circumstances.

Two notable U.S. Circuit Court cases in 2020 came down on opposite sides on this question. In hiQ Labs v. LinkedIn, the Ninth Circuit reaffirmed that scraping public data likely does not violate the CFAA. Meanwhile, the Eleventh Circuit ruled in Compulife Software v. Newman that scraping may violate the CFAA even if the data is public.

This split between circuits creates uncertainty on how the CFAA applies to web crawling. The Supreme Court recently narrowed the scope of the CFAA in the case of Van Buren v. United States, but some experts believe the law is still unclear and open to abuse in civil cases against scrapers.

Other key laws include:

  • Copyright law and the Digital Millennium Copyright Act (DMCA) protect copyrighted content from being scraped and reused without permission.
  • Trespass to Chattels is a common law tort sometimes applied to web crawling, claiming unauthorized use of computer systems.
  • Breach of Contract can apply if crawling violates a website‘s terms of service that prohibit scraping.

Privacy regulations, anti-hacking statutes, trade secrets law, and even insider trading rules could also come into play depending on the specifics of the crawling activity and data use.

In the European Union, the recently enacted General Data Protection Regulation (GDPR) adds another layer of legal considerations. Scraping any personally identifiable information on EU citizens requires a valid legal basis under the GDPR such as consent from the individual. The GDPR also grants data subjects certain rights like accessing the data collected on them.

According to a survey by Bright Data (formerly Luminati Networks), GDPR compliance is now the top concern for companies that collect web data, with 68% ranking it as a high priority. Fines for non-compliance can be severe – up to €20 million or 4% of a company‘s annual global revenue.

Beyond the core factors of how scraping is done, how data is used, and what laws apply, some additional issues can come into play:

  • If the scraper is gathering data from websites in different countries, it may need to comply with varying data protection laws beyond the GDPR. Even within the U.S., states like California have their own strict privacy regulations such as the California Consumer Privacy Act (CCPA).

  • Using proxy servers or VPNs to hide the scraper‘s identity and location may violate a site‘s terms of service. Although this alone may not be illegal, it could still lead to getting blocked or even legal action in some cases. The anonymity provided by proxies could also make it easier to perform clearly illegal activities.

  • Scraped data could potentially be used for insider trading if it provides non-public information that gives an unfair advantage in the stock market before it becomes available to other investors. While the law is still unclear, regulators are starting to scrutinize web scraping used for financial gain more closely.

  • Jurisdiction and venue can impact legal outcomes significantly. For example, many lawsuits in the U.S. are filed in the Northern District of California, which may be more familiar with web scraping technology and business models. Some legal experts believe this court is more friendly to defendants in scraping cases than other jurisdictions.

So with all the potential legal pitfalls, how can you take advantage of web scraping while minimizing legal risks? Here are some expert tips:

  1. Carefully review and respect the target website‘s terms of service and robots.txt file. Get permission from the site owner if possible. "If a website doesn‘t want you to scrape it, don‘t do it unless you have an exceptionally good reason," advises attorney Zach Slobin.

  2. Use reasonable crawling rates and don‘t adversely impact the site‘s performance or availability. Set delays between requests and obey rate limits. However, keep in mind that "there is no hard and fast rule on what crawl rate is acceptable," notes data expert Sanaea Daruwalla. It depends on the specific site infrastructure.

  3. Monitor the site for technical countermeasures against scraping like CAPTCHAs, user agent blocking, etc. "Unusual traffic patterns or excessive requests from a single IP are common red flags that can get you banned," explains web scraping consultant Isaiah Hull. Be prepared to stop crawling if asked by the site owner.

  4. Avoid scraping content that is access-restricted, copyrighted, or contains personal information, especially under GDPR. Focus on factual, publicly available data. "The more ‘creative‘ the data is, the more likely it is to be protected by copyright, requiring a license or permission to use," states attorney Ruth Carter.

  5. Consider how scraped data will be used and if it falls under fair use. Limit commercial applications that compete with the source site. "Scraping data to provide a service that the website already provides is much riskier than scraping to provide a new service or insight," says Bennet Kelley.

  6. Consult with a lawyer specializing in technology and intellectual property law if you have legal concerns or questions about a specific use case. "It‘s crucial to do a risk assessment upfront before starting any major scraping project," recommends attorney James Bikoff.

  7. Use an established, reputable web scraping tool or outsource to a service provider with experience in keeping web crawling legal and ethical. Sami Saleh from web scraping service OxyLabs suggests "partnering with a company that maintains its own proxy servers in-house and has strict compliance processes can offload a lot of potential headaches."

The Bottom Line

The legality of web crawling and scraping remains a complex issue that depends heavily on the specifics of each case. With the proper precautions and an awareness of the key risk factors, it is possible to legally leverage this powerful technology.

However, the laws in this area continue to evolve and can vary widely by jurisdiction. Just because a use case seems low-risk now doesn‘t guarantee it will always be in the clear. It‘s important to continually re-evaluate the landscape and adjust policies as needed.

When in doubt, it‘s advisable to get professional legal guidance tailored to your situation. Engaging a lawyer well-versed in these matters can help devise a compliant crawling strategy and limit your liability.

Enlisting a trusted web scraping service provider can also reduce risk by shifting some of the burden of legal compliance to the vendor. Choose an experienced partner with a commitment to white hat best practices.

As the Bright Data survey revealed, 79% of businesses believe web scraping will be even more important in the coming years. But 48% are at least somewhat worried about potential legal action over their data collection practices.

By being proactive about legal issues, companies can harness the power of web data while minimizing dangers. With a thoughtful approach, web scraping can be an invaluable tool for driving smarter business decisions in an increasingly data-driven world.

Leave a Reply

Your email address will not be published. Required fields are marked *