[ 
https://issues.apache.org/jira/browse/DERBY-4555?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15305423#comment-15305423
 ] 

Danoja Dias commented on DERBY-4555:
------------------------------------

I think the type of skip should be short. 

When I study the code, I found followings.

org.apache.derby.catalog.SystemProcedures.java contains the 
SYSCS_IMPORT_TABLE function It calls ImportTable function in the 
org.apache.derby.impl.load.Import.java .

Reading of csv file happens at readNextToken method in 
org.apache.derby.impl.load.ImportReadData.java
readNextToken reads one column's value at a time.

In org.apache.derby.impl.load.ImportReadData.java  the method ignoreFirstRow() 
ignores the first row by looking for the record separator. It is done by 
readNextToken() method. 

I think we can use that mechenism for skipping the headerlines.





> Expand SYSCS_IMPORT_TABLE to accept CSV file with header lines
> --------------------------------------------------------------
>
>                 Key: DERBY-4555
>                 URL: https://issues.apache.org/jira/browse/DERBY-4555
>             Project: Derby
>          Issue Type: Improvement
>          Components: Miscellaneous
>            Reporter: Yair Lenga
>            Assignee: Danoja Dias
>         Attachments: petlist.csv, petlist.csv, petlist.csv, repro.java
>
>
> The SYSCS_IMPORT_TABLE (and SYSCS_IMPORT_DATA) function allow import of data 
> from external resources. In general, they can process CSV files that created 
> with various tools - with one exception: the header line.
> While there is no accepted standard, most tools will include a header line in 
> the CSV file with column names. This convention is supported in Excel and 
> many other tools.
> My Request: extend the SYSCS_IMPORT_TABLe and SYSCS_IMPORT_DATA (and other 
> related procedures) to include an extra indicator for the number of header 
> lines to be ignored.
> As an extra bonus it will be help is the SYSCS_IMPORT_DATA will accept column 
> names (instead of column indexes) in the 'COLUMNINDEXES' arguments. E.g., it 
> should be possible to indicate COLUMNINDEXES of '1,3,sales,5,'. This feature 
> will make it significantly easier to handle cases where the external input 
> files is extended to include additional columns.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to