[
https://issues.apache.org/jira/browse/ARROW-5666?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16870952#comment-16870952
]
Robin Kåveland commented on ARROW-5666:
---------------------------------------
That's just a slight modification to the expression above:
{code:java}
all(key.isdigit() or (key.startswith('-') and key[1:].isdigit()) for key in
self.keys){code}
But it's starting to feel like a bad idea to attempt to do this coercing at
all. Maybe it's better to force the user to deal with what type the partition
key should have? If it were to be interpreted as something that looks a bit
like {{pd.Categorical}}, it would be relatively cheap to read it into memory
and deal with it after the user has read the file?
> [Python] Underscores in partition (string) values are dropped when reading
> dataset
> ----------------------------------------------------------------------------------
>
> Key: ARROW-5666
> URL: https://issues.apache.org/jira/browse/ARROW-5666
> Project: Apache Arrow
> Issue Type: Bug
> Components: Python
> Affects Versions: 0.13.0
> Reporter: Julian de Ruiter
> Priority: Major
> Labels: parquet
>
> When reading a partitioned dataset, in which the partition column contains
> string values with underscores, pyarrow seems to be ignoring the underscores
> in the resulting values.
> For example if I write and then read a dataset as follows:
> {code:java}
> import pyarrow as pa
> import pandas as pd
> df = pd.DataFrame({
> "year_week": ["2019_2", "2019_3"],
> "value": [1, 2]
> })
> table = pa.Table.from_pandas(df.head())
> pq.write_to_dataset(table, 'test', partition_cols=["year_week"])
> table2 = pq.ParquetDataset('test').read()
> {code}
> The resulting 'year_week' column in table 2 has lost the underscores:
> {code:java}
> table2[1] # Gives:
> <Column name='year_week' type=DictionaryType(dictionary<values=int64,
> indices=int32, ordered=0>)>
> [
> -- dictionary:
> [
> 20192,
> 20193
> ]
> -- indices:
> [
> 0
> ],
> -- dictionary:
> [
> 20192,
> 20193
> ]
> -- indices:
> [
> 1
> ]
> ]
> {code}
> Is this intentional behaviour or is this a bug in arrow?
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)