Feynman Liang created SPARK-10199:
-------------------------------------
Summary: Avoid using reflections for parquet model save
Key: SPARK-10199
URL: https://issues.apache.org/jira/browse/SPARK-10199
Project: Spark
Issue Type: Improvement
Components: ML
Reporter: Feynman Liang
Multiple model save/load in ML use case classes to infer a schema for the data
frame saved to Parquet. However, inferring a schema from case classes uses
[runtime
reflection|https://github.com/apache/spark/blob/d7b4c095271c36fcc7f9ded267ecf5ec66fac803/sql/core/src/main/scala/org/apache/spark/sql/SQLContext.scala#L361]
which is unnecessary since the types are already known at the time `save` is
called.
It would be better to just specify the schema for the data frame directly using
`sqlContext.createDataFrame(dataRDD, schema)`.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]