Scraping & Visualizing YouTube Comments to Uncover Fan Sentiment During the 2018 World Cup

The 2018 FIFA World Cup, hosted by Russia, was a tournament for the ages. Over the course of a thrilling month of action, a total of 32 teams representing countries from around the globe competed for the most prestigious trophy in football. In the end, France emerged victorious, defeating Croatia 4-2 in a memorable final to claim their second World Cup title.

But while the action on the pitch captivated billions of viewers, just as fascinating was the global conversation happening online as fans reacted to every twist and turn of the tournament. And there‘s no larger platform for that conversation than YouTube.

As the world‘s second most visited website with over 2 billion monthly active users, YouTube is a goldmine of user engagement data. During the World Cup, football fans flocked to the platform to watch highlights, interviews, recaps, and more – generating millions of views, likes, and comments in the process.

For data analysts and decision-makers, this trove of unstructured data presents an incredible opportunity to gain insight into fan behavior, preferences, and sentiment. But collecting, cleaning, and making sense of YouTube data at scale is no easy feat.

In this post, we‘ll share how we used the web scraping tool Octoparse and proxy services to extract and analyze a dataset of over 600 videos and 6,000 comments related to the 2018 World Cup. We‘ll dive into the challenges of scraping YouTube data, the process of conducting sentiment analysis and data visualization on the comments, and the insights we uncovered as a result. Let‘s kick things off!

The Challenges of Scraping YouTube Data

While YouTube offers an incredible volume and variety of user-generated content to analyze, it also presents some unique technical challenges for web scrapers.

Unlike static web pages with predictable structures, many aspects of the YouTube interface are dynamically generated and loaded via JavaScript. This can make it difficult to locate and extract specific page elements using traditional parsing methods.

YouTube is also highly protective of its platform and users, employing various techniques to detect and block suspicious scraping activity. This includes rate limiting IP addresses that make too many requests in a short period of time, as well as more sophisticated defenses like browser fingerprinting and user behavior analysis.

To successfully scrape YouTube at scale, scrapers must find creative workarounds to these challenges. Some best practices include:

  • Using residential proxy services that provide IP addresses associated with real human users, as opposed to data center IPs which are more easily flagged
  • Employing IP rotation to distribute requests across a wide pool of IPs and avoid hitting rate limits
  • Customizing request headers (user agent, referrer, etc.) to mimic organic traffic
  • Inserting random delays between requests to avoid appearing bot-like
  • Leveraging headless browsers or real browser engines to execute JavaScript and render dynamic page content
  • Utilizing captcha solving services to automatically bypass challenges

For our World Cup YouTube scraping project, we employed several of these techniques – most notably using rotating residential IPs from providers like Smartproxy and Proxy-Cheap to avoid detection and scale up our operation. More on our exact process in the next section.

Scraping YouTube Data with Octoparse

To collect our World Cup dataset, we used Octoparse, a powerful and user-friendly visual web scraping tool. Octoparse allows users to build scrapers by simply pointing and clicking on the desired data fields in their web browser. It then generates a scraping workflow that can be run on a schedule or on demand.

We set up two separate Octoparse crawlers for this project:

  1. A crawler to scrape video metadata for 600+ YouTube search results for the query "World Cup 2018"
  2. A crawler to scrape comments from each of the individual video pages found in the first crawler

For the video crawler, we extracted the following fields:

  • Video title
  • View count
  • Like count
  • Dislike count
  • Comment count
  • Video duration
  • Channel name
  • Video description
  • Publish date

And for the comments crawler, we targeted:

  • Commenter name
  • Comment text
  • Comment like count
  • Replies count
  • Comment date

To ensure we collected as much data as possible, we configured the video crawler to automatically scroll to the bottom of the search results page and click the "Load more" button until all results were loaded. And for the comments crawler, we set up a loop to paginate through all the comments on each video.

After running our crawlers and waiting a few hours, we had our structured World Cup dataset ready for analysis. In total, we collected:

  • Metadata for 612 videos from 200+ channels
  • 6,327 comments and associated metadata

Now came the fun part – slicing and dicing the data to see what insights we could glean from it!

Analyzing Video Popularity & Engagement

First, we wanted to get a high level view of which World Cup YouTube videos captured the most attention and engagement from fans. So we analyzed our dataset to identify the top videos by a few key metrics.

Videos with the most views

Rank | Video Title | View Count (millions)
1 | France v Croatia – 2018 FIFA World Cup™ FINAL – HIGHLIGHTS | 51.8
2 | Mário Bros Reaction to France Winning the World Cup 2018 | 33.9
3 | 2018 FIFA World Cup Russia™ Official Song | 31.7
4 | Brasil v Belgium – 2018 FIFA World Cup™ MATCH 58 | 29.6
5 | Hyundai Goal of the Tournament 2018 FIFA World Cup Russia™ | 21.4

Interestingly, while the official match recap videos dominated the top of the list as expected, a few of the most viewed World Cup videos were actually music videos for the tournament‘s official songs. This shows the power of music to drive viral engagement and reach a massive global audience.

Videos with the highest like to dislike ratio

To gauge which videos were the most well-received by fans, we calculated the ratio of likes to dislikes for each video. A higher ratio indicates a video with a large proportion of positive reactions.

Rank | Video Title | Like/Dislike Ratio
1 | 2018 FIFA World Cup Russia™ Review | 98.3
2 | Lionel Messi Reaction after Croatia vs Argentina (3-0) | 97.8
3 | Luka Modric accepts his Golden Ball award | 88.2
4 | Diego Maradona Interview about Lionel Messi | 85.3

Some of the most positively received videos were player interviews, tournament recaps, and clips of stars like Messi and Modric. This suggests that fans on YouTube value exclusive behind-the-scenes content that offers a more intimate perspective on the tournament and their favorite players.

Videos with the most comments

Finally, we looked at which videos generated the most discussion in terms of raw comment volume.

Rank | Video Title | Comment Count
1 | Brasil v Belgium – 2018 FIFA World Cup™ MATCH 58 | 15,862
2 | France v Croatia – 2018 FIFA World Cup™ FINAL – HIGHLIGHTS | 13,289
3 | Portugal v Spain – 2018 FIFA World Cup™ MATCH 3 | 7,590

Unsurprisingly, the videos with the most comments tended to be recaps of the most high-profile and dramatic matches of the tournament, especially those featuring global superstars like Cristiano Ronaldo and Neymar. The final match between France and Croatia also generated a massive amount of discussion.

Analyzing Comment Sentiment

While comment volume is one way to measure fan engagement, to truly take the pulse of fan sentiment we needed to analyze the actual content of the comments.

One useful technique for gauging the emotional tenor of large volumes of text is sentiment analysis. Sentiment analysis uses natural language processing algorithms to classify pieces of text as positive, negative, or neutral in sentiment.

We ran the comments for several of the most popular World Cup matches through the VADER sentiment analysis model and calculated the proportion of positive, negative, and neutral comments for each match. Here‘s what we found:

France vs Croatia (Final)

  • Positive comments: 63%
  • Negative comments: 12%
  • Neutral comments: 25%

England vs Croatia (Semi-final)

  • Positive comments: 45%
  • Negative comments: 35%
  • Neutral comments: 20%

Brazil vs Belgium (Quarter-final)

  • Positive comments: 37%
  • Negative comments: 52%
  • Neutral comments: 11%

As you might expect, the final match generated the highest proportion of positively classified comments, as fans were generally happy with the result and the quality of play. The dramatic semifinal between England and Croatia, which Croatia won in extra time, was more polarizing, with a nearly even split between positive and negative reactions.

And in the shocking quarterfinal that saw tournament favorites Brazil eliminated by Belgium, negative comments significantly outweighed positive ones as Brazilian fans expressed their disappointment and frustration.

While sentiment analysis provides a helpful way to quantify the overall emotional response to different matches, we also wanted to go a level deeper and surface some of the specific topics and talking points that fans were buzzing about.

To do this, we generated word clouds from the comments on several key match recap videos. Word clouds provide a visual representation of how frequently different words appear in a text corpus – the larger the word in the cloud, the more often it was mentioned.

Here are a few of the most insightful word clouds we created:

France vs Croatia (Final)

France vs Croatia Word Cloud

For the final, the most frequent terms in the comments were the names of the two teams and their key players, like "Mbappe", "Modric", "Pogba" and "Perisic". There was also a high prevalence of positive words like "congratulations", "deserved", "best", reflecting the general sentiment that France was a worthy champion.

England vs Colombia (Round of 16)

England vs Colombia Word Cloud

England‘s dramatic penalty shootout victory over Colombia was a major talking point. "Penalties", "penalty" and "shootout" were among the top terms, along with the names of England players like "Kane", "Trippier" and "Pickford" and frequent mentions of "IKEA" in reference to Colombian keeper David Ospina.

Germany vs South Korea (Group Stage)

Germany vs South Korea Word Cloud

The match that sealed reigning champion Germany‘s shocking early exit from the tournament generated a deluge of colorful fan commentary. Frequent terms included "karma", "schadenfreude", "defending champs", "curse", and "shocked", painting a vivid picture of the dismay and amusement of fans witnessing a massive upset.

Interesting Findings & Applications

Sentiment analysis of YouTube comments is still an emerging field, and our findings really only scratched the surface of what‘s possible. As we continue refining our web scraping and natural language processing pipelines, here are a few other potential insights we could glean from this data:

  • Identifying the most positively and negatively perceived players based on mentions and sentiment of comments related to individual players
  • Analyzing sentiment broken down by language or country of origin to understand how fans from different nations reacted to key moments
  • Using emoji analysis as a proxy for fan emotion, looking at the most frequently used emojis in comments

We‘re also excited about some of the potential real-world applications of this YouTube comment data, such as:

  • Sports betting & odds making: Sentiment analysis of comments could potentially be used as a signal to predict match outcomes or adjust betting odds in real-time
  • Targeted advertising: Comments reveal valuable info about viewer demographics, preferences, and purchase intent that could inform YouTube ad targeting
  • Content recommendations: Identifying the most popular types of videos for each team or player could help inform what other content to create or suggest to fans
  • Fan engagement: Teams and leagues could use comment data to understand how fans are reacting to key moments and identify opportunities to connect with their audience

Of course, any applications of this data would need to be pursued thoughtfully and with respect for fan privacy. YouTube comments are a public forum, but analyzing and deriving insights from that data at scale enters an ethical gray area that warrants careful consideration.

Conclusion

Scraping and analyzing YouTube comments from the 2018 World Cup proved to be a fascinating data-driven lens into the thoughts and feelings of football fans around the world.

The scale of the data – tens of thousands of comments across hundreds of videos – required some creative technical solutions to extract and wrangle into an analyzable format. By leveraging tools like Octoparse and proxy services to gather the data, and natural language processing techniques like sentiment analysis and word clouds to process it, we were able to quantify the fan zeitgeist in a novel way.

From identifying the most popular and polarizing matches by volume and sentiment of comments, to uncovering the key topics and terms on fans‘ minds, this analysis gave us a deeper understanding of how one of the world‘s largest sporting events was perceived and discussed online.

Of course, YouTube comments are just one of many data sources that can be mined for insights into fan behavior and sentiment. Social media platforms like Twitter and Facebook, online forums like Reddit, and even traditional channels like TV and radio all offer rich troves of sports fan data ripe for exploration.

The challenge for data analysts and decision-makers is figuring out how to efficiently collect, clean, and extract meaning from these disparate datasets at scale. With the right tools and techniques, there are endless opportunities to gain a competitive edge.

At the end of the day, major sporting events like the World Cup are as much about the fans as they are the players on the pitch. They offer a unique window into human sentiment and behavior at a massive scale. And by harnessing the power of web scraped data, there‘s never been more potential to understand and connect with sports fans in a meaningful way.

*This post was written by the team at ScrapeOps. We provide powerful tools and managed web scraping services to help organizations turn messy web data into actionable insights.

If you enjoyed this analysis, be sure to check out our blog for more posts on web scraping, data science, and interesting findings! You can also follow us on Twitter for the latest updates.*

Leave a Reply

Your email address will not be published. Required fields are marked *