Skip to content
whatismyalternative

Common Crawl

Free, open repository of web crawl data usable by anyone.

Visit Common Crawl

Common Crawl is a 501(c)(3) non-profit founded in 2007 that crawls the web and freely provides its archives and datasets to the public. It maintains an open repository of web crawl data collected since 2008, totaling more than 10 petabytes. Crawls are published approximately once a month, each typically containing more than two billion web pages, hosted on Amazon Web Services through its Open Data Sponsorship Program and downloadable at no cost.

The organization's crawler, CCBot, is based on Apache Nutch and identifies itself via the User-Agent string CCBot, obeying robots.txt directives and rate-limiting its requests to individual servers. Crawl data is provided in WARC (raw HTTP responses), WAT (metadata and link graphs), and WET (plaintext extracted from HTML) formats. Monthly index files are published in the CDXJ Index and URL Index, the latter queryable with Amazon Athena. Common Crawl also publishes host-level and domain-level Web Graphs with computed Harmonic Centrality and PageRank values, and experimental data products on Hugging Face.

Common Crawl is intended for researchers and others who need wholesale extraction, transformation, and analysis of open web data without maintaining their own crawling infrastructure. The dataset is freely available at no cost.

12 alternatives to Common Crawl

Ranked by how well each tool replaces Common Crawl: shared features, audience, price and popularity.

  1. Distributed data warehouse for massive-scale analytics.

    Covers 3 of 15 key features and is open source.

    Open source
    56 out of 100 match—
  2. A platform for analyzing large data sets with a high-level language and infrastructure for

    Covers 3 of 15 key features and is open source.

    Open source
    55 out of 100 match—
  3. Highly extensible, scalable, production-ready web crawler

    Covers 4 of 15 key features and is open source.

    Free planOpen source
    50 out of 100 matchFree
  4. Web data infrastructure for developers and AI

    Covers 4 of 15 key features.

    Free plan
    50 out of 100 match$99/mo
  5. Unified engine for large-scale data analytics.

    Covers 1 of 15 key features and is open source.

    Open source
    48 out of 100 match—
  6. Apache Kylin Overview.

    Covers 1 of 15 key features and is open source.

    Open source
    48 out of 100 match—
  7. The Cost Efficient Data Lake

    Covers 1 of 15 key features.

    Free plan
    48 out of 100 matchUsage-based
  8. A platform purpose-built for high-speed data engineering.

    Covers 3 of 15 key features and is open source.

    Free planOpen source
    47 out of 100 matchFree
  9. Open-source self-hosted web archiving

    Covers 5 of 15 key features and is open source.

    Open source
    47 out of 100 match—
  10. The AI platform for data and analytics

    Covers 4 of 15 key features.

    47 out of 100 matchContact sales
  11. The web data API to search, scrape, and interact with the web at scale, turning websites L

    Covers 1 of 15 key features and is open source.

    Free planOpen source
    47 out of 100 match$16/mo
  12. The Open Graph Viz Platform

    Covers 5 of 15 key features and is open source.

    Free planOpen source
    47 out of 100 matchContact sales