Running Mapreduce on compressed data

Question

Hi. I have compressed the output of Map as follows:

conf.set("mapreduce.map.output.compress", true) 
conf.set("mapreduce.output.fileoutputformat.compress", false)

Now I want to run a Mapreduce on this data. How to do this?

score 0 · Answer 1 · Jul 24, 2019

It is very straight forward, no need to implement any custom input format for the same. You can use any input formats with compression. The only step is to add the compression codec to the value in io.compression.codecs

Suppose if you are using LZO then your value would look something like

io.compression.codecs  =  org.apache.hadoop.io.compress.GzipCodec, org.apache.hadoop.io.compress.DefaultCodec, com.hadoop.compression.lzo.LzopCodec

Then configure and run your map reduce jobs as you do normally on uncompressed files. When map wants to process a file and if it is compressed it would check for the io.compression.codecs and use a suitable codec from there to read the file.