What is the purpose of shuffling and sorting phase in the reducer in Map Reduce

Question

Hi Guys,

I am new to Map Reduce. In Map Reduce programming, the reduce phase has shuffling, sorting, and reduction as its sub-parts. Can anyone tell me the purpose of the shuffling and sorting phase in the reducer in Map Reduce Programming?

MD · Answer 1 · Dec 20, 2020

Hi@akhtar,

Shuffle phase in Hadoop transfers the map output from Mapper to a Reducer in MapReduce. Sort phase in MapReduce covers the merging and sorting of map outputs. Data from the mapper are grouped by the key, split among reducers, and sorted by the key. Every reducer obtains all values associated with the same key. Shuffle and sort phase in Hadoop occur simultaneously and are done by the MapReduce framework.

Learn more about Big Data and its applications from the Azure Data Engineer certification.