Web crawling is the process of systematically browsing and indexing websites by following links from page to page. Search engines like Google use complex web crawlers to discover and catalog the trillions of pages on the internet. But you can also create your own small-scale web crawler in PHP to extract specific data and content from a set of websites.
There are many reasons you might want to build a web crawler:
- To scrape product data from ecommerce sites for market research
- To aggregate news articles and blog posts on certain topics
- To monitor your brand mentions and online reputation
- To generate sales leads from business directories
- To automate website testing and monitoring
- And many more!
While there are pre-built web scraping tools available, coding your own crawler in PHP gives you complete flexibility and control. With some basic programming knowledge, it‘s not too difficult to create a simple PHP web crawler. Let‘s walk through the process step-by-step.
Step 1: Set Up the PHP Crawler Script
To start, create a new PHP file for your crawler script. At the top, include an HTML form with an input field where you can paste the URL of the page you want to crawl. You‘ll also need a submit button to initiate the crawl.
<form method="post">
<input type="text" name="url" placeholder="Enter URL">
<input type="submit" name="submit" value="Crawl">
</form>Below that, open up your PHP tags to begin the crawler script:
<?php
if (isset($_POST[‘submit‘])) {
$url = $_POST[‘url‘];
// crawler code will go here
}
?>This checks if the form has been submitted, and if so, stores the entered URL in a variable.
Step 2: Use Regular Expressions to Extract Data
With the target URL in hand, you can now fetch the page‘s HTML using PHP‘s built-in file_get_contents() function. However, the raw HTML will contain a lot of irrelevant code and content. To extract just the data you want, you‘ll need to parse the HTML.
Regular expressions (regex) allow you to search for specific patterns within a block of text. You can use regex to locate and extract elements like titles, links, images, and paragraphs.
Here‘s a handy custom function that searches a string between a starting and ending regex pattern:
function extractStringRegex($str, $start, $end) {
$temp = preg_split($start, $str);
$result = preg_split($end, $temp[1]);
return $result[0];
}For example, to extract a page title wrapped in <title> tags:
$html = file_get_contents($url);
$title = extractStringRegex($html, ‘/<title>/‘, ‘/<\/title>/‘);
echo $title;You can modify the start and end patterns to match any HTML element. Regex can get quite complex for heavily nested elements. There are many online tools to help you build and test regular expressions.
Step 3: Use String Functions for Simple Parsing
For simpler parsing tasks, regular expressions may be overkill. In those cases, you can use PHP‘s string functions like explode() and strpos() to split strings on a delimiter.
For instance, here‘s a function that extracts the substring between two strings:
function extractStringBetween($str, $start, $end) {
$temp = explode($start, $str, 2);
$result = explode($end, $temp[1], 2);
return $result[0];
}So to extract an image URL that appears after src=" and before the closing ":
$imgurl = extractStringBetween($html, ‘src="‘, ‘"‘);The string manipulation route tends to be more straightforward than regex, but may not always work if the HTML structure is inconsistent between pages.
Step 4: Add Functions to Save Content and Images
Extracting data is one thing, but you‘ll likely want to save it for later use. You can write the scraped text content to a local file using fopen() and fwrite().
It‘s wise to include some simple logging to save the raw HTML in case you need to troubleshoot the extracted content later:
function saveLog($content) {
$file = fopen(‘log.txt‘, ‘w‘) or die("Unable to open file!");
fwrite($file, $content);
fclose($file);
}Saving images is a bit more involved. You‘ll need to detect the file type from the URL extension, generate a unique filename, create the local directory path, fetch the binary image data, and write it to a file.
Here‘s a function to handle the image download process:
function downloadImage($url, $dirName) {
$defaultName = basename($url);
$ext = pathinfo($url, PATHINFO_EXTENSION);
$allowed = [‘jpg‘, ‘jpeg‘, ‘gif‘, ‘png‘, ‘webp‘];
if (!in_array($ext, $allowed)) {
return false;
}
$filename = time() . rand(0, 9999) . ‘.‘ . $ext;
$dirPath = $dirName . ‘/‘ . date(‘Y/m/d‘) . ‘/‘;
if (!is_dir($dirPath)) {
mkdir($dirPath, 0777, true);
}
$binary = @file_get_contents($url);
if ($binary) {
file_put_contents($dirPath . $filename, $binary);
return $dirPath . $filename;
}
return false;
}The function returns the final path of the saved image on success. It‘s a good idea to organize the downloaded files by date to avoid overwriting files.
Step 5: Put It All Together
Now let‘s combine the previous steps into a working PHP crawler that extracts the title, description, and main image from a product page on Amazon:
<?php
if (isset($_POST[‘submit‘])) {
$url = $_POST[‘url‘];
$html = file_get_contents($url);
saveLog($html);
$title = extractStringRegex($html, ‘/<span id="productTitle" class="a-size-large product-title-word-break">/‘, ‘/<\/span>/‘);
echo "<strong>Title:</strong> $title<br>";
$desc = extractStringRegex($html, ‘/<div id="productDescription" class="a-section a-spacing-small">/‘, ‘/<\/div>/‘);
echo "<strong>Description:</strong> $desc<br>";
$imgurl = extractStringBetween($html, ‘<img id="landingImage" class="a-dynamic-image a-stretch-vertical" src="‘, ‘"‘);
$imgpath = downloadImage($imgurl, ‘img‘);
echo ‘<img src="‘ . $imgpath . ‘">‘;
}
?>Of course, this just scratches the surface of what‘s possible with web crawling. You‘ll likely run into complications like inconsistent HTML templates, dynamically loaded content, pagination, rate limiting, CAPTCHAs, and more.
There are strategies to overcome these obstacles, such as detecting and adapting to different page templates, using headless browsers to render JavaScript, crawling with delays and retries, and utilizing proxies and spoofing techniques. But those are more advanced topics.
When to Use a Pre-Built Web Scraping Tool
Coding your own web crawler in PHP certainly has advantages in terms of flexibility and control. You can build it to your exact specifications.
However, there‘s also an argument for using established web scraping tools and frameworks, especially for non-programmers or those who want to quickly set up data extraction on a larger scale.
Pre-built scrapers offer plug-and-play setup with a visual interface to map data fields. They handle a lot of the complexities like AJAX rendering, pagination, and formatting out of the box.
They also make it easy to schedule scraping jobs, export data, and integrate with other apps. Some even provide web-based APIs and hosted services so you don‘t have to run the crawler on your own server. The tradeoff is less ability to customize.
Popular web scraping platforms include:
- Octoparse
- ParseHub
- Mozenda
- Scrapy Cloud
- Apify
If you do opt to build your own PHP crawler, there are also many open source libraries that can help with the heavy lifting, such as:
- Goutte – a simple PHP web scraper
- HTTPFul – RESTful PHP library
- PHP cURL – PHP client URL library
- QueryPath – a PHP library for parsing HTML
- Symfony DomCrawler – web crawler component
Scaling and Automating PHP Web Crawlers
A basic PHP crawler running on your local machine is fine for scraping a handful of pages. But if you want to crawl hundreds or thousands of URLs on a regular basis, you‘ll need to think about scalability and automation.
Some tips for scaling up your PHP crawlers:
- Use a queuing system like RabbitMQ to manage crawl jobs
- Store crawl data in a database like MySQL instead of flat files
- Run multiple crawler instances in parallel on different servers
- Leverage PHP‘s multi-curl functions to fetch pages concurrently
- Cache frequently crawled pages to reduce server load
- Continuously monitor and optimize crawler performance
To automate your crawlers, you can set up cron jobs to trigger PHP scripts on a schedule. For example, to scrape a list of URLs every 6 hours.
Managing large-scale web scraping projects also requires planning around politeness (crawl rate), stealth (spoofing), and reliability (error handling). There are risks in aggressive crawling, from overloading servers to getting blocked or banned.
Legal Considerations for Web Crawling
Is it legal to scrape data from websites using PHP or any other method? The short answer is: it depends.
Web scraping itself is not illegal, but there are some legal pitfalls to be aware of:
Copyright – you should only crawl and reproduce publicly available data, not content behind logins or paywalls. Use scraped content within fair use guidelines.
Terms of Service – many websites prohibit scraping in their ToS. While these are not always legally binding, they could suspend your account or IP for violating them.
Trespass to Chattels – scraping that intentionally overloads servers could be considered a denial of service attack in extreme cases.
GDPR and CCPA – if you scrape personal data, you are subject to data privacy laws around notice, consent, and individual rights.
To stay on the right side of the law, follow these best practices:
- Respect robots.txt rules that indicate which pages are off limits to crawlers
- Identify your crawler with a custom user agent string and provide a way for sites to contact you
- Slow down your crawl speed to a reasonable rate
- Don‘t crawl sites with "no scraping" language in their ToS
- Don‘t republish scraped content or personal data without permission
Wrap-up
Web crawling is a powerful technique for extracting data from websites at scale. With a language like PHP, you can build your own customized web crawler to scrape specific content for all kinds of research, aggregation, and analysis.
The basic process involves fetching pages, parsing HTML, and saving the extracted data. From there, you can expand your crawler‘s capabilities with more advanced techniques to overcome scraping obstacles.
Before you start scraping, be sure to understand the legal implications and follow crawling best practices to avoid trouble. Consider the ongoing maintenance and infrastructure required to keep your crawlers running smoothly.
If you don‘t have the technical expertise or resources to build and manage web crawlers in-house, a pre-built web scraping tool may be the way to go. In any case, there are many options available to automate data extraction from the web.