Content Aggregators: The Content Publishers of the Future?

The rise of the internet and digital publishing has radically transformed how we discover and consume content. Gone are the days when our news and entertainment primarily came from a select few newspapers, magazines, and TV channels. We now have access to a virtually unlimited supply of content on any topic imaginable, available instantly at our fingertips.

While this content renaissance has been a boon for consumers, it has created challenges for traditional publishers. Faced with falling print circulation and fierce online competition for eyeballs and advertising dollars, many have struggled to adapt their business models for the digital age.

Into this tumultuous publishing landscape have stepped content aggregators – digital platforms and apps that collect content from a wide range of sources and serve it up in a convenient, personalized format for users. By making it easy to access exactly the content you want all in one place, aggregators like Flipboard, Reddit, and Apple News have quickly gained massive popularity. But are these platforms simply useful tools for navigating today‘s vast content ecosystem, or do they represent the future of publishing itself?

The Age of Aggregation

To understand the disruptive impact of content aggregators, it‘s helpful to compare how we consumed media in the past to how we do it today. Just a few decades ago, if you wanted to stay informed, you would pick up a newspaper on your way to work or tune into the nightly news on TV. Your content diet was limited to what a small group of editors and producers chose to publish or broadcast.

Fast forward to the present, and the average person is now bombarded with content 24/7 across a dizzying array of channels and devices. Over 4 million blog posts are published every day, alongside half a million tweets and 300 hours of new YouTube video uploaded every minute. No human could possibly keep up with even a fraction of all this content.

Enter content aggregators. These platforms use a combination of human curation and machine learning algorithms to sift through the deluge of available content and surface the most relevant, interesting pieces for each individual user based on their interests and reading habits. Instead of having to visit dozens of different websites or apps, users can access a personalized content feed all in one place. It‘s like having a super-smart, hyper-productive personal assistant to read the entire internet for you and highlight only the best parts.

Some of the most popular aggregator platforms and apps include:

  • Flipboard – collects articles and photos based on your selected interests
  • Apple News – comes pre-installed on iOS devices and pulls in stories from major news outlets
  • Google News – algorithmically aggregates headlines from thousands of publishers
  • Reddit – user-submitted content is up-voted or down-voted by the community
  • Pocket – lets you save articles from any publication in a central reading list
  • Feedly – a customizable RSS feed reader for staying on top of your favorite sites

The use of content aggregators has exploded in recent years as our content consumption has shifted to mobile devices. The average American now spends over 3 hours per day consuming digital media on mobile, according to eMarketer. And the majority of that time is spent in just a handful of aggregator apps.

AggregatorMonthly Active Users
Flipboard145 million
Reddit430 million
Apple News125 million
Pocket30 million
Feedly14 million

Sources: Flipboard, Reddit, Apple, Pocket, Feedly

The Secret Sauce: Web Scraping and Data Extraction

So how do content aggregators manage to wrangle the web‘s infinite sea of content into tidy personalized packages? The key is web scraping – the process of programmatically extracting data and content from websites. By unleashing fleets of web crawler bots to systematically browse the internet and parse the HTML of web pages, aggregators can collect and structure massive amounts of data to power their content recommendation engines.

Web scraping allows content aggregators to ingest, categorize and analyze content from virtually any online source at scale. The scraped data can include things like:

  • Headlines and body text
  • Topic tags and keywords
  • Author and publisher info
  • Engagement metrics (views, likes, shares, etc.)
  • Sentiment analysis
  • Named entities and facts
  • Quotes and attributions
  • Images, videos and other multimedia

Aggregator engineering teams build sophisticated content ingestion pipelines to clean, normalize and store all of this disparate data in structured databases. They can then apply machine learning models to the data to extract insights and power the hyper-personalized content recommendations that keep users engaged.

For example, Topic, a fast-growing news aggregator focused on business and tech content, uses web scraping to collect over 100,000 articles per day from across the web. Their NLP engine scans and tags each article for over 300 different topics, sentiment, and keywords. This rich metadata is the foundation of the app‘s content matching system. As Topic CEO Jordan Jacobs explained in an interview with Digiday:

"Our core IP is really around content ingestion and classification. Being able to take in a huge volume of content and understand what it‘s about, who it‘s for, and match that content to the right users is really the heart of what we do."

Of course, scraping content from thousands of websites every day is no easy feat. Content aggregators have to deal with a host of technical challenges, from bot detection and CAPTCHAs to dynamic page loading and inconsistent site structures. One of the most common obstacles is IP blocking. When a website detects a suspicious amount of traffic coming from a single IP address, it will often block that IP to prevent abuse.

To get around this, web scrapers have to distribute their scraping activity across many different IP addresses using proxies. By rotating their IP on each request, or parallelizing the workload across multiple IPs, scrapers can avoid triggering blocking mechanisms and ensure high availability for their data pipelines.

The top content aggregators will often employ a combination of data center proxies and residential proxies from trusted networks like Bright Data, Oxylabs, and Smartproxy to power their scraping infrastructure. Some, like Flipboard, have even built custom proxy management solutions to intelligently route different scraping jobs across their diverse proxy pools based on the target site and request type.

According to Shad Baxter of scraping tool provider ScrapeOwl:

"As the web has become more dynamic and sites are better at detecting scraping traffic, it‘s becoming crucial to use rotating proxies when collecting content data at scale. The top aggregators all do this in some form. Using a combination of data center and residential IPs with smart proxy management helps ensure high success rates and keep costs down."

The Battle for Attention

As content aggregators have gained traction, they‘ve begun to change user expectations and behaviors around content consumption. A 2020 study by the Reuters Institute found that across age groups, people are now more likely to access news through "side door" channels like aggregators and social media rather than going directly to publishers.

18-2425-3435-4445-5455+
Direct to Publisher28%36%46%50%55%
Search, Social, Aggregators72%64%54%50%45%

Source: Reuters Institute Digital News Report 2020

For publishers, this shift presents both an opportunity and a threat. On one hand, aggregators can be a powerful source of new readers, particularly for niche or lesser-known publications that may struggle to attract direct traffic. Referrals from aggregators expose publishers‘ content to highly targeted potential subscribers.

However, having their content live in someone else‘s platform means publishers have far less control over the user experience and the direct relationship with readers. Aggregators often display the full text of articles in their own apps, which can make the publisher‘s own website feel irrelevant. And since most aggregators monetize through their own ads or subscriptions, not through revenue shares, publishers may see little direct financial benefit even if an aggregator sends them substantial traffic.

There are also concerns that content aggregators — and the recommendation algorithms that power them — can exacerbate ideological echo chambers and make it harder for people to encounter diverse perspectives. If an aggregator only shows you content that matches your existing interests and beliefs, it may reinforce your prejudices and make you less open to differing viewpoints.

The Future of Aggregation

So what does the future hold for content aggregation? All signs point to continued growth and evolution. Deloitte predicts that the percentage of consumers who use news aggregators will grow to 75% by 2025. And as media becomes more fragmented across new digital channels and formats, aggregators are well-positioned to help users navigate the noise.

Already, we‘re seeing a new wave of aggregation startups emerge to tackle different verticals and content types beyond traditional text articles. Some examples:

  • Odysee – Aggregates and categorizes the best podcast episodes into personalized playlists
  • Artifact – A "TikTok for news" that uses machine learning to curate short-form video news stories
  • Upnext – Compiles the top performing content across social media (Instagram, TikTok, YouTube, etc.) into a single feed
  • Workweek – Aggregates business content (newsletters, podcasts, tweets, etc.) aimed at professionals

Upnext app
Upnext aggregates top social media content. Source: Upnext.com

As these apps proliferate, publishers and content creators will have to continue adapting their strategies to reach audiences. Some may embrace aggregators as a key distribution channel, optimizing their content to rank highly in recommendation feeds. Others may focus on building more direct relationships with readers to reduce their reliance on aggregators. The most successful will likely take a balanced "barbell" approach, using aggregators strategically to grow reach while doubling down on a unique value proposition and user experience on their owned channels.

It‘s clear that the days of one-size-fits-all, unified content destinations are over. The future of media is personalized, algorithm-driven and aggregated. Users increasingly want easy access to diverse, high-quality content that matches their unique interests. Content aggregators are at the forefront of delivering on that demand today, leveraging web scraping, machine learning and massive troves of user data to bring order to the chaos of the modern internet.

Does this mean the death of the traditional publication as we know it? Not necessarily. But it does mean that all publishers will have to get more creative and user-centric to stay relevant in an aggregator-driven world. Those who do will find no shortage of opportunities to connect with engaged audiences. Those who don‘t may find themselves increasingly invisible. Welcome to the age of aggregation.

Leave a Reply

Your email address will not be published. Required fields are marked *