The world is awash in data. According to a report from DOMO, over 2.5 quintillion bytes of data are created every single day. By 2025, it‘s estimated that 463 exabytes of data will be created each day globally – the equivalent of over 212 million DVDs per day!
However, a significant share of this data is generated by private companies and organizations, locked away in proprietary databases and internal systems. For researchers, analysts, and data enthusiasts, getting access to interesting datasets can be a major challenge.
Thankfully, the open data movement has been growing steadily in recent years. Governments, universities, non-profits, and even some private companies are embracing the idea that certain data should be freely available to everyone. As of 2021, over 1 million datasets are published on governmental open data portals alone.
In this post, we‘ll curate 70 of the best free and open datasets available online. We‘ll also share some tips and best practices for finding and working with open data, including using web scraping and proxies to collect data yourself. Let‘s dive in!
Categories of Open Data Sources
Government and Census Data
Government agencies are some of the largest producers of data, covering topics like population demographics, economic indicators, energy usage, transportation, and more. Much of this data is made available through public data portals and websites. Here are a few of the top government data sources:
Data.gov – The US government‘s open data portal, with over 250,000 datasets from 14 different federal agencies.
- Example: Farmers Markets Geographic Data – locations and details on over 8,500 farmer‘s markets in the US.
European Union Open Data Portal – A single point of access for data from institutions and agencies of the European Union.
- Example: Eurostat Regional Data – statistics on population, employment, health, education, and more for EU regions.
UK Government Web Archive – An archive of over 3 billion web pages from UK central government websites, including data and documents.
- Example: National Public Transport Access Nodes (NaPTAN) dataset – a database of all public transport access points in the UK.
Using government data sources, you can conduct in-depth policy analysis, examine demographic trends, or look for economic insights. Many of these portals include visualization and analysis tools directly in the browser. For example, Data.gov includes embedded previews of the datasets using charts and maps.
Some other government data sources to check out:
- World Bank Open Data
- China Statistical Yearbooks
- Kenya Open Data Portal
- Singapore Department of Statistics
Health and Drug Data
Advances in medical research generate vast amounts of data on health, medicine, and biology. Many governments and organizations make this data open to facilitate scientific collaboration and enhance public health initiatives. Some of the best open health data sources include:
NCBI Datasets – An extensive collection of biomedical and genomic datasets from the National Center for Biotechnology Information (NCBI), including gene sequences, clinical studies, and more.
- Example: ClinVar – a public archive of interpretations of clinically relevant genetic variants.
World Health Organization Global Health Observatory – A gateway to health-related statistics for more than 1000 indicators for UN member states.
- Example: Mortality and Burden of Disease – data on life expectancy, death rates, and disease burdens globally.
Medicare.gov Datasets – Detailed, anonymized data on Medicare services, providers, and payments in the US.
- Example: Inpatient Prospective Payment System Provider Summary – hospital-level data on Medicare payments and services.
By using open health data, researchers can study disease trends, evaluate public health interventions, or investigate correlations between health and other factors like environment or socioeconomic status. Many open health datasets are quite large, containing millions of records, so tools like BigQuery or Spark can be helpful for analysis.
More health data sources worth exploring:
- Project Tycho – Global health data from the 20th century
- Broad Bioimage Benchmark Collection – biological images for researching algorithms
- MIMIC Critical Care Database – anonymized health data from intensive care units
Financial and Economic Data
Financial markets move fast, and having access to accurate, up-to-date data is essential for making informed decisions. Multiple organizations provide open data on stocks, trade, employment, and other economic indicators. Some top financial data sources are:
Quandl – Over 35 million financial and economic time series datasets from over 500 publishers.
- Example: London Stock Exchange (LSE) End of Day Data – daily stock prices and trading volume from LSE.
OpenCorporates – The world‘s largest open database of companies, with over 180 million records.
- Example: UK Companies House Data – all businesses registered with Companies House in UK, updated daily.
Federal Reserve Economic Data (FRED) – Over 500,000 economic time series from 87 different sources, including the World Bank, OECD, and BLS.
- Example: Effective Federal Funds Rate – historical data on the interest rate that banks charge each other for overnight loans.
Financial data can be analyzed to understand market trends, assess investment opportunities, or build predictive models. APIs provide a way to access this data in real-time and incorporate it into applications. However, be aware that some financial data sources may have restrictions on commercial use.
Other financial data sources to explore:
- IEX Cloud – Real-time and historical stock prices and market data
- World Inequality Database – Global data on income and wealth inequality
- Global Financial Data – Long-term historical financial and economic data series
Social Media and Web Data
Some of the most interesting and timely data today comes from the web and social media. User posts, product reviews, web traffic patterns, and other online activities can provide valuable insights into human behavior and consumer trends. Here are some of the top web and social data sources:
Common Crawl – An open repository of web crawl data, collected over 8 years of web crawling. Contains raw web page data, extracted metadata and text, and web graph information.
- Example: Common Crawl 2018-47 Index – 2.72 billion web pages crawled in November 2018.
Twitter API – Access Twitter data including tweets, user profiles, mentions, followers, and more. Both streaming and historical data is available.
- Example: COVID-19 Stream – Real-time stream of tweets related to COVID-19.
Pushshift Reddit Datasets – A collection of submissions and comments from Reddit, dating back to 2005. Includes both current data and historical archives.
- Example: Reddit Comments – Archive of all publicly available Reddit comments.
Analyzing web and social data can uncover trends and insights not available in conventional datasets. Data scientists use NLP techniques to understand the sentiment and topics in unstructured text, or network analysis to map the relationships between people and entities online.
Some challenges with web data include its sheer scale (often measured in terabytes or petabytes), inconsistent structure, and concerns about data ethics and user privacy. When working with personal or potentially sensitive data, be sure to anonymize records and follow relevant regulations like GDPR.
Other web data sources to check out:
- Wikipedia Dumps – Snapshots of all Wikipedia content, in various languages
- Yelp Open Dataset – User reviews and business data from Yelp
- GitHub Archive – A record of all public GitHub activity
Tips for Collecting and Using Open Data
Web Scraping for Data Collection
Not all data you need will be available in ready-to-download datasets. Sometimes, the data you want is spread across multiple web pages, or buried in PDFs, images, or other unstructured formats. In these cases, you may need to resort to web scraping to compile the data yourself.
Web scraping is the process of programmatically extracting data from websites. Using libraries like BeautifulSoup, Scrapy, and Selenium, you can parse and extract specific data fields from web pages. For example, you could scrape product names, prices, and reviews from an e-commerce site, or extract event titles and dates from an online calendar.
Some tips for effective and responsible web scraping:
- Respect websites‘ terms of service and robots.txt files. Some sites prohibit scraping, or ask crawlers to limit request rates.
- Space out your requests to avoid overwhelming servers. Add delays between requests and follow crawl-delay directives.
- Rotating your IP address and user-agent can help prevent blocking. Using a pool of proxy servers will distribute your requests across IPs.
- Cache your data locally to avoid repeated requests. Be sure to re-run your scraper regularly to fetch updated data.
- Anonymize any personal data you collect, and do not release scraped datasets publicly without permission.
There are many open-source web scraping tools available in various languages. Some popular options are:
- BeautifulSoup – A Python library for parsing HTML and XML documents
- Scrapy – A fast and powerful web crawling framework for Python
- Puppeteer – A Node library for controlling a headless Chrome browser
- rvest – An R package for web scraping and parsing HTML
Challenges of Open Data
While open data offers many opportunities, working with it also comes with challenges. Here are a few common issues to watch out for:
- Inconsistent formats – Open datasets come in all sorts of file formats, from CSV and JSON to XML and PDF. Data formatting may be inconsistent across files.
- Missing or inaccurate metadata – Without clear metadata on what each field means, how data was collected, or when it was last updated, an open dataset can be hard to use effectively.
- Changing data access – The availability or location of open data can change over time, especially for government datasets. Data portals may be redesigned or migrate to different URLs.
- Licensing and terms of use – Open data licenses can vary widely. Some datasets are in the public domain, while others have restrictions on commercial use, derivative works, or redistributing the data.
To mitigate these challenges, adopt good data management practices. Use open, standard file formats like CSV and JSON. Document your data collection and processing steps. Regularly check dataset links and update your code to handle changes. Read data licenses carefully and comply with their terms to avoid legal issues down the road.
Data Ethics and Responsible Use
Open data is a powerful tool, but like any technology, it can be misused. It‘s important for data practitioners to consider the ethical implications of their work and use open data responsibly. Some key ethical principles include:
Privacy – Even if a dataset does not contain direct personal identifiers, it may still be possible to re-identify individuals by combining it with other datasets. Practice data minimization and avoid collecting or exposing unnecessary personal information.
Bias and Fairness – Open datasets can contain biases baked in from their collection or production. Using biased data in machine learning models or analysis can reinforce and amplify those biases. Audit your data for potential bias and consider ways to mitigate it.
Transparency and Accountability – Be open about your data sources, methodologies, and limitations. Document your assumptions and provide a way for others to replicate your work. Consider the potential impacts of your project and consult with relevant stakeholders.
Consent and Agency – If you are collecting data directly from people, be transparent about how that data will be used and shared, and give them a way to opt-out. Respect people‘s rights to privacy and control over their personal data.
These ethical principles apply whether you are working with open data or private data. By holding ourselves to high ethical standards, we can unlock data‘s benefits while avoiding its potential harms.
The Importance of Open Data
In our data-driven world, access to high-quality, open datasets is crucial. Open data enhances transparency, enables innovation, and promotes collaboration across sectors. Here are a few reasons why open data matters:
Transparency and Accountability – Open government data allows citizens and watchdog groups to scrutinize government activities, budgets, and performance. This public accountability can help combat corruption and build trust.
Research and Innovation – Open data accelerates scientific research by allowing scientists to access more datasets and build on each other‘s work. In the business world, open data can fuel new products, services, and economic opportunities.
Efficiency and Cost Savings – Sharing data openly breaks down silos between departments and organizations. This can reduce duplicated efforts, speed up decision-making, and save money.
Equity and Inclusion – Open data can help expose and address social inequities. For example, open education data has revealed disparities in student performance and disciplinary outcomes by race and income level.
As data becomes an increasingly important resource, we must continue to advocate for more accessible, usable open data. This means supporting open data initiatives in government, academia, and the private sector. It means contributing to the maintenance and improvement of open data platforms. And it means using open data responsibly and ethically in our own projects.
Conclusion
We‘ve explored 70 open datasets across a range of categories, from government and health data to financial records and web archives. These open data sources are a treasure trove for researchers, analysts, and curious data enthusiasts. By leveraging this data, we can uncover new insights, build innovative applications, and tackle important social challenges.
But we‘ve also seen that working with open data comes with its own challenges and ethical considerations. From inconsistent data formats to concerns over bias and privacy, we must approach open data thoughtfully and responsibly.
If we can‘t find the open data we need, web scraping provides a powerful tool to collect it ourselves. By using web scraping best practices and respecting website owners‘ terms of service, we can expand the universe of open data even further.
Ultimately, the open data movement is about more than just releasing datasets. It‘s about cultivating a culture of transparency, collaboration, and data-driven innovation. It‘s about using data as a tool to solve problems and improve people‘s lives.
As data practitioners, we have an important role to play in this movement. By using open data in our work, advocating for more open data initiatives, and holding ourselves to high ethical standards, we can help unlock data‘s full potential for public good. So let‘s get out there and start exploring the amazing open datasets at our fingertips!