Which is the easiest way for text analytics with hadoop?

0 votes

 am experimenting with hadoop and the distributions of Hortonwork and cloudera in order to do some simple text analytics. All the examples I have found until now on the web regarding e.g. wordcount deal with only one column. But I have many text files on which wordcount must be applied and the results must be saved in a spreadsheet, each in a separate column. So I was wondering what is the easiest way to do text analytics with hadoop in conjunction with spreadsheets. The functions I need are:

  • transform to lower case
  • filter stopwords
  • transpose results
  • write to excel

Can this be accomplished easily with Pig or Rhadoop or something else?

Nov 22, 2018 in Big Data Hadoop by Neha
• 6,280 points
96 views

1 answer to this question.

0 votes
Apache pig provides CSVExcelStorage class for loading or storing into csv format, it uses CSV conventions of Excel 2007. Apart from that I have also experimented with storing the results from Pig to mongoDB and then reading it into R using rmongodb library.
answered Nov 22, 2018 by Frankie
• 9,810 points

Related Questions In Big Data Hadoop

0 votes
1 answer

Best way of starting & stopping the Hadoop daemons with command line

First way is to use start-all.sh & ...READ MORE

answered Apr 15, 2018 in Big Data Hadoop by Shubham
• 13,300 points
1,888 views
0 votes
1 answer
0 votes
9 answers

Is there any way to check which Hadoop daemons are running?

use jps command, It will show all the running ...READ MORE

answered Dec 27, 2018 in Big Data Hadoop by Rakesh
• 160 points
8,961 views
0 votes
1 answer

Hadoop Mapreduce word count Program

Firstly you need to understand the concept ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,710 points
3,325 views
0 votes
1 answer

hadoop.mapred vs hadoop.mapreduce?

org.apache.hadoop.mapred is the Old API  org.apache.hadoop.mapreduce is the ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,710 points
398 views
0 votes
10 answers

hadoop fs -put command?

put syntax: put <localSrc> <dest> copy syntax: copyFr ...READ MORE

answered Dec 7, 2018 in Big Data Hadoop by Aditya
16,432 views
0 votes
1 answer

Hadoop dfs -ls command?

In your case there is no difference ...READ MORE

answered Mar 16, 2018 in Big Data Hadoop by kurt_cobain
• 9,260 points
1,197 views
0 votes
1 answer

Which is the Real Time Monitoring tool/API for Hadoop?

If you're using Yarn, there's a rest ...READ MORE

answered Sep 4, 2018 in Big Data Hadoop by Frankie
• 9,810 points
150 views