[
https://issues.apache.org/jira/browse/FLINK-10684?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Flink Jira Bot updated FLINK-10684:
-----------------------------------
Labels: CSV stale-major (was: CSV)
I am the [Flink Jira Bot|https://github.com/apache/flink-jira-bot/] and I help
the community manage its development. I see this issues has been marked as
Major but is unassigned and neither itself nor its Sub-Tasks have been updated
for 30 days. I have gone ahead and added a "stale-major" to the issue". If this
ticket is a Major, please either assign yourself or give an update. Afterwards,
please remove the label or in 7 days the issue will be deprioritized.
> Improve the CSV reading process
> -------------------------------
>
> Key: FLINK-10684
> URL: https://issues.apache.org/jira/browse/FLINK-10684
> Project: Flink
> Issue Type: Improvement
> Components: API / DataSet
> Reporter: Xingcan Cui
> Priority: Major
> Labels: CSV, stale-major
>
> CSV is one of the most commonly used file formats in data wrangling. To load
> records from CSV files, Flink has provided the basic {{CsvInputFormat}}, as
> well as some variants (e.g., {{RowCsvInputFormat}} and
> {{PojoCsvInputFormat}}). However, it seems that the reading process can be
> improved. For example, we could add a built-in util to automatically infer
> schemas from CSV headers and samples of data. Also, the current bad record
> handling method can be improved by somehow keeping the invalid lines (and
> even the reasons for failed parsing), instead of logging the total number
> only.
> This is an umbrella issue for all the improvements and bug fixes for the CSV
> reading process.
--
This message was sent by Atlassian Jira
(v8.3.4#803005)