Query regarding a spark split logic

0 votes
I have a csv file containing 10 lines. I need to have a spark code to split even and odd numbers of lines into 2 text files results. Could you help me out?
Feb 9, 2019 in Apache Spark by Joshi
56 views

1 answer to this question.

0 votes

First, import the data in Spark and add IDs to it. Run these commands in scala console:

val df = spark.read.csv("file.csv")
val df1 = df.withColumn("id",monotonicallyIncreasingId)

Then use this logic to split:

val df2 = df1.filter($"id"%2!==0)
val df2 = df1.filter($"id"%2!==0)
answered Feb 9, 2019 by Omkar
• 69,040 points

Related Questions In Apache Spark

0 votes
1 answer

Error : split value is not a member of org.apache.spark.sql.Row

spark.read.csv is used when loading into a ...READ MORE

answered Jul 10, 2019 in Apache Spark by Rishi
2,470 views
0 votes
1 answer

Query regarding Appending " to a string in Scala

You can perform this task in two ...READ MORE

answered Jul 10, 2019 in Apache Spark by Esha
892 views
0 votes
1 answer

Error : split value is not a member of org.apache.spark.sql.Row

spark.read.csv is used when loading into a ...READ MORE

answered Jul 22, 2019 in Apache Spark by Firoz
852 views
0 votes
1 answer

Unable to run select query with selected columns on a temp view registered in spark application

from pyspark.sql.types import FloatType fname = [1.0,2.4,3.6,4.2,45.4] df=spark.createDataFrame(fname, ...READ MORE

answered Mar 28 in Apache Spark by GAURAV
• 140 points
256 views
+1 vote
2 answers
+1 vote
1 answer

Hadoop Mapreduce word count Program

Firstly you need to understand the concept ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,920 points
5,299 views
0 votes
1 answer

hadoop.mapred vs hadoop.mapreduce?

org.apache.hadoop.mapred is the Old API  org.apache.hadoop.mapreduce is the ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,920 points
774 views
+1 vote
11 answers

hadoop fs -put command?

put syntax: put <localSrc> <dest> copy syntax: copyF ...READ MORE

answered Dec 7, 2018 in Big Data Hadoop by Aditya
32,321 views
–1 vote
1 answer

Not able to use sc in spark shell

Seems like master and worker are not ...READ MORE

answered Jan 3, 2019 in Apache Spark by Omkar
• 69,040 points
393 views
0 votes
1 answer

Spark and Scale Auxiliary constructor doubt

println("Slayer") is an anonymous block and gets ...READ MORE

answered Jan 8, 2019 in Apache Spark by Omkar
• 69,040 points
69 views