Apache Hive
Distributed data warehouse for massive-scale analytics.
Apache Hive is a distributed, fault-tolerant data warehouse system that enables analytics at massive scale. It lets users read, write, and manage petabytes of data in distributed storage using SQL, and it is built on Apache Hadoop with support for S3, Azure Data Lake, Google Cloud Storage, and other cloud storage systems. The Hive Metastore (HMS) provides a central metadata repository for tables and partitions, serving as a building block for data lake architectures.
Key components include HiveServer2 (HS2), which supports multi-client concurrency, authentication, and JDBC/ODBC access for business intelligence tools; the Hive Metastore Server; and the Beeline client. Hive provides full ACID transactions for ORC tables and insert-only support for other formats, along with query-based and MapReduce-based data compaction. It supports Apache Iceberg tables through Hive StorageHandler, Apache Calcite's cost-based query optimizer, LLAP for low-latency interactive queries, and data replication for backup and disaster recovery. Security features include Kerberos authentication, integration with Apache Ranger for authorization, and Apache Atlas for data lineage and governance.
Apache Hive is released under the Apache License v2 and is managed by the Apache Software Foundation. It is available as open source, with documentation, a Docker quickstart, and community resources including mailing lists and issue tracking.
12 alternatives to Apache Hive
Ranked by how well each tool replaces Apache Hive: shared features, audience, price and popularity.
- 69 out of 100 matchUsage-based
Open-source software for reliable, scalable, distributed computing
Covers 1 of 15 key features and has a free plan.
Free planOpen source66 out of 100 matchFreeA platform for analyzing large data sets with a high-level language and infrastructure for
Covers 4 of 15 key features.
Open source66 out of 100 match—- 64 out of 100 match—
- 63 out of 100 matchUsage-based
A platform purpose-built for high-speed data engineering.
Covers 6 of 15 key features and has a free plan.
Free planOpen source62 out of 100 matchFreeOpen source, native analytic database for open data and table formats.
Covers 1 of 15 key features.
Open source62 out of 100 match—- 62 out of 100 matchFree
A cluster management framework for partitioned and replicated distributed resources
Covers 4 of 15 key features and has a free plan.
Free planOpen source61 out of 100 matchContact salesDistributed real-time computation system for processing unbounded data streams
Covers 3 of 15 key features and has a free plan.
Free planOpen source61 out of 100 matchFreeA unified programming model for defining and executing batch and streaming data processing
Covers 1 of 15 key features.
Open source59 out of 100 match—