Fighting B2B E-Commerce Fraud with Web Data Extraction, Proxies and Machine Learning

The meteoric rise of business-to-business (B2B) e-commerce has been a boon for businesses—but also for fraudsters. Scammers are increasingly exploiting B2B marketplaces to defraud legitimate businesses with fake listings, orders and payments. The impact is staggering:

  • B2B e-commerce sales are projected to reach $1.8 trillion by 2023 (Forrester)
  • B2B fraud attempts doubled in 2020 and now account for 10-15% of total transaction value (Sift)
  • Every $1 lost to B2B fraud costs companies $3.36 in chargebacks, fees and lost merchandise (LexisNexis)

To stay ahead of scammers, B2B marketplaces are turning to an unlikely ally: web scrapers. By extracting and analyzing publicly available data from across the web, marketplaces can identify suspicious patterns and proactively block fraudsters before they strike.

In this deep dive, we‘ll explore how data scraped from websites, combined with machine learning techniques, forms an effective line of defense against the most common types of B2B fraud. We‘ll cover the critical role of proxies in web scraping, share a case study of one company‘s success, and discuss advanced data science approaches on the horizon.

Common B2B E-Commerce Fraud Tactics

First, it‘s important to understand how B2B fraud typically works. The most prevalent schemes include:

  1. Seller fraud: Scammers create fake business profiles and listings to lure in B2B buyers, then abscond with their payments without delivering products.

  2. Buyer fraud: Fraudsters pose as legitimate businesses to place large orders with sellers, often using stolen credit cards or fake payment info. After receiving the goods, they disappear without paying.

  3. Triangulation fraud: In this twist on buyer fraud, scammers make purchases from legitimate sellers using stolen payment credentials, then resell the delivered goods to unsuspecting buyers at steep discounts.

What unites these tactics is deception. Fraudsters go to great lengths to appear like real, trustworthy businesses, from setting up convincing websites to buying aged domains to crafting backstories that seem to check out.

Only by peering behind this facade—using the vast data available on the open web—can B2B marketplaces spot the telltale signs of fakery and fraud.

Web Scraping for B2B Fraud Detection

Web scraping is the automated process of extracting publicly available data from websites. Using scripts or specialized software, web scrapers can gather large volumes of data spanning many sites quickly and efficiently.

For B2B fraud detection, some of the most valuable data points to scrape include:

  • Business registration details (e.g. company name, address, tax ID)
  • Contact information (e.g. email addresses, phone numbers, social media profiles)
  • Employee and executive data (e.g. names, titles, bios)
  • Website content (e.g. "About Us" pages, product listings, terms of service)
  • Domain registration details (e.g. WHOIS records, DNS entries, SSL certificates)
  • Backlinks and referring domains
  • Customer reviews and ratings

By cross-referencing this external data with a user‘s on-platform activity, B2B marketplaces can surface red flags like:

  • Mismatches between a user‘s claimed identity and their web presence
  • Reuse of contact details across many different business profiles
  • "Cloned" websites with similar or plagiarized content
  • Suspicious domain registration patterns (e.g. very recent creation date, obscured WHOIS)
  • Negative reviews or scam reports associated with a business‘s contact info
  • Unusual linking patterns or sudden shifts in backlink profile

Take the example of a scammer attempting seller fraud. They might set up a new business profile and website purporting to sell high-demand products at low prices. A scraped WHOIS record reveals the domain was registered just days ago, while its listed contact email matches other known scam sites. These data points, combined with abnormal on-platform behavior, build a compelling case to block the user as a likely fraudster.

The Crucial Role of Proxies

To be effective, web scraping for fraud detection must be done at scale and speed. Marketplaces may need to extract data on hundreds of thousands of businesses across millions of websites. Delays in data collection give fraudsters more time to inflict damage.

The problem is, many websites are well defended against scraping. They monitor visitor IP addresses for bot-like behavior and quickly block those exceeding certain request thresholds. Some also use advanced techniques like browser fingerprinting and CAPTCHAs to separate bots from human users.

To avoid these roadblocks, most large-scale web scraping operations route their requests through proxy servers. A proxy acts as an intermediary, forwarding requests from the scraper to the target website. To the site, the traffic appears to originate from the proxy‘s IP address rather than the scraper itself.

By rotating requests across many different proxy IPs, scrapers can:

  • Distribute their traffic to avoid hitting rate limits
  • Diversify their digital fingerprints to mimic human users
  • Bypass IP-based blocking and geographic restrictions
  • Mask their scraping activity and infrastructure from detection

When selecting proxies for B2B fraud scrapers, some key considerations include:

  • Proxy type: Data center proxies are fast and cheap, but easier to detect and block. Residential proxies (sourced from real user devices) and mobile proxies (from mobile carriers) tend to be more trusted by websites and allow for more human-like traffic patterns.

  • Rotation: Some proxy providers offer dedicated (static) IPs assigned to a single user. Others automatically rotate IPs from a large pool, either per-request or at set intervals. The latter approach is usually better for large-scale scraping.

  • Location: For B2B scrapers targeting global marketplaces, geo-targeted proxies are a must. Many proxy services allow users to select specific countries or cities for their IPs.

  • Connection type: Proxies can forward requests via HTTP/HTTPS, SOCKS4 or SOCKS5 protocols. HTTP is simplest but may not support other protocols needed for scraping. SOCKS proxies are more flexible but can be slower.

Leading proxy providers for web scraping include Bright Data, Oxylabs, Smartproxy, and GeoSurf. Many offer specialized APIs, tooling and support for large-scale scraping use cases.

Analyzing Scraped Data for Fraud Signals

Simply collecting web data isn‘t enough—B2B fraud teams must translate it into actionable insights. A common approach is to feed scraped data into machine learning models that predict the probability a given user is fraudulent.

One well-suited model type is logistic regression, which estimates a binary outcome (fraud or not) based on a set of input features. A basic logistic regression fraud model might incorporate features like:

FeatureDescriptionPossible Values
domain_age_daysAge of user‘s website domainInteger
is_domain_privacy_protectedWhether WHOIS privacy is enabled0 (no) or 1 (yes)
phone_num_uniqueness_scoreUniqueness of phone number across sites0 (low) to 1 (high)
email_in_abuse_listsWhether email appears in abuse lists0 (no) or 1 (yes)
about_us_text_similaritySimilarity of "About Us" content to other sites0 (low) to 1 (high)
address_validity_scoreEstimated validity of physical address0 (low) to 1 (high)

The model learns to weigh each feature‘s contribution to fraud risk based on past known cases. It outputs a fraud probability between 0-1 that can be used for decision automation (e.g. blocking users above a certain threshold).

B2B marketplaces can start with a basic logistic model and then iterate to add more sophisticated data points like:

  • Scraped contact data matches with public sanction lists, criminal records or scammer databases
  • Graph-based features that assess the interconnectedness and centrality of a business‘s web presence (e.g. shared hosting, referring domains)
  • Time series features capturing changes in a business‘s website or online footprint over
    time (e.g. spikes in referring domains or contact info changes)
  • Linking features measuring the similarity or anomalousness of a business‘s web presence vs. peers in the same industry vertical
  • Computer vision features extracting risk signals from website screenshots and images (e.g. blurred text, placeholder photos)

Advanced models can move beyond binary fraud classification to anomaly detection—surfacing businesses with unusual web presences or activity patterns relative to the norm. Unsupervised learning techniques like K-means clustering, LOF (local outlier factor) and autoencoder neural networks are well suited to this task.

Challenges and Tradeoffs

While web scraping is a powerful tool for B2B fraud detection, it‘s not without challenges. Some key ones include:

  • False positives: Blocking legitimate users (false positives) can be as damaging as letting fraudsters through (false negatives). Marketplaces must carefully tune models and thresholds to minimize false positives.
  • Resource intensity: Large-scale web scraping and data science can be costly and complex. Marketplaces may need to invest in engineering talent, proxy infrastructure and computational resources.
  • Cat-and-mouse games: As anti-fraud systems evolve, so do fraudsters‘ techniques. Teams must continually monitor performance and adapt models to new fraud patterns.
  • Privacy and compliance: Scraping must be done with respect for user privacy, intellectual property rights and emerging regulations like GDPR and CCPA.

The key is striking the right balance based on each marketplace‘s unique risk profile and objectives. Many start with a human-in-the-loop approach: using ML to flag risky users for manual review while tuning model thresholds. Over time, they can increase automation as they build confidence in their models and processes.

Case Study: Fighting Fraud with Web Scraping and Machine Learning

One company putting these techniques into practice is PurchaseHere, a fast-growing B2B marketplace for industrial goods. In 2019, PurchaseHere‘s fraud team noticed a spike in suspicious new seller accounts with dubious credentials. Losses mounted as buyers placed orders that were never fulfilled.

To fight back, PurchaseHere built an anti-fraud engine powered by web data. The system scraped dozens of data points on each new seller—from website content to company registration records—via an automated pipeline of over 100 rotating proxies.

This data was combined with on-platform signals to train a gradient boosted decision tree model predicting seller fraud risk. High-scoring accounts were auto-blocked, while medium risk ones were routed for human review.

The results were impressive: PurchaseHere reduced fraud losses by 90% within three months of deploying the system. False positives stayed below 2% even as the model blocked hundreds of fraudulent sign-up attempts per week. The ROI of the initiative exceeded 200% in Year 1 alone.

The Future of B2B Fraud Detection

As B2B e-commerce continues to grow, so will the opportunities and incentives for fraud. Forward-thinking marketplaces are investing now in web data capabilities to stay ahead of emerging fraud threats.

Some exciting frontiers include:

  • Collective intelligence: Pooling fraud signals and blacklists across marketplaces to identify scammers attacking multiple platforms
  • Deep fakes: Using computer vision and deep learning to spot fake business photos, videos and documents designed to deceive
  • Behavioral biometrics: Incorporating user interaction patterns (e.g. mouse movements, keystrokes) to detect bot vs. human activity
  • Natural language processing: Applying NLP techniques to scraped web content for deeper insights into a business‘s legitimacy and vertical

One thing is certain: in the battle against B2B fraud, data is the most powerful weapon. By harnessing the rich signals available on the web, marketplaces can keep scammers in check and preserve trust and growth in the B2B e-commerce ecosystem.

Leave a Reply

Your email address will not be published. Required fields are marked *