load/save in spark

0 votes

1) load command will load only parquet 

val a = spark.read.load("employee.parquet")--works fine
val a=spark.read.load("employee.txt")--error

2)val a =spark.read.format("csv").load("employee.txt")--works fine

3)what is the difference between 1 and 2 point loading text as data 

4)except parquet file remaining format are not accepting the below stmt 

val a = spark.read.load("text/csv/json") 

5) what is the difference of this two stmts 

val a =spark.read.json()
val a =spark.read.format("json").load("a.json")
Jul 5 in Apache Spark by Esha
14 views

1 answer to this question.

0 votes

The reason why you are able to load employee.parquet and not employee.txt using load in spark.read.load by default it assumes that data source is in parquet format so it is able to load it but we can use format function which can be used to specify the different format and use the load function to load the data 

spark.read.json()

Spark SQL can automatically infer the schema of a JSON dataset and load it as a Dataset[Row]. This conversion can be done by spark.read.json()

Both spark.read.json() and spark.read.format("json").load("a.json") are same as you can see in the below screenshot,

image

image

answered Jul 5 by Firoz

Related Questions In Apache Spark

0 votes
1 answer

Changing Column position in spark dataframe

Yes, you can reorder the dataframe elements. You need ...READ MORE

answered Apr 19, 2018 in Apache Spark by Ashish
• 2,630 points
3,421 views
0 votes
1 answer

Efficient way to read specific columns from parquet file in spark

As parquet is a column based storage ...READ MORE

answered Apr 20, 2018 in Apache Spark by kurt_cobain
• 9,240 points
993 views
+2 votes
4 answers

use length function in substring in spark

You can use the function expr val data ...READ MORE

answered May 3, 2018 in Apache Spark by kurt_cobain
• 9,240 points
9,722 views
0 votes
1 answer

cache tables in apache spark sql

Caching the tables puts the whole table ...READ MORE

answered May 4, 2018 in Apache Spark by Data_Nerd
• 2,360 points
529 views
0 votes
0 answers
0 votes
1 answer

Hadoop Mapreduce word count Program

Firstly you need to understand the concept ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,110 points
2,048 views
0 votes
1 answer

hadoop.mapred vs hadoop.mapreduce?

org.apache.hadoop.mapred is the Old API  org.apache.hadoop.mapreduce is the ...READ MORE

answered Mar 16, 2018 in Data Analytics by nitinrawat895
• 10,110 points
196 views
0 votes
10 answers

hadoop fs -put command?

copy command can be used to copy files ...READ MORE

answered Dec 7, 2018 in Big Data Hadoop by Sujay
10,478 views
+5 votes
11 answers

Concatenate columns in apache spark dataframe

its late but this how you can ...READ MORE

answered Mar 21 in Apache Spark by anonymous
22,964 views
0 votes
1 answer

How can I write a text file in HDFS not from an RDD, in Spark program?

Yes, you can go ahead and write ...READ MORE

answered May 29, 2018 in Apache Spark by Shubham
• 13,190 points
865 views