Hadoop Hive Bucketing

Question

Why do we use Bucketing in Hive ?

Omkar · Answer 1 · Dec 14, 2018

Partition helps in increasing the efficiency when performing a query on a table. Instead of scanning the whole table, it will only scan for the partitioned set and does not scan or operate on the unpartitioned sets, which helps us to provide results in lesser time and the details will be displayed very quickly because of Hive Partition.

Now, let’s assume a condition that there is a huge dataset. At times, even after partitioning on a particular field or fields, the partitioned file size doesn’t match with the actual expectation and remains huge and we want to manage the partition results into different parts. To overcome this problem of partitioning, Hive provides Bucketing concept, which allows the user to divide table data sets into more manageable parts.

Thus, Bucketing helps the user to maintain parts that are more manageable and the user can set the size of the manageable parts or Buckets too.