jorisvandenbossche commented on PR #51157:
URL: https://github.com/apache/arrow/pull/51157#issuecomment-5909207899

   One remaining behaviour question I have, is what to do about NaN, which we 
explicitly disallow for object dtype when converting to string (or to any type 
other than float, in general), unless you explicitly specify `from_pandas=True`:
   
   ```python
   # converting None to null -> fine
   >>> pa.array(np.array(["a", None], dtype=object))
   <pyarrow.lib.StringArray object at 0x7f813a761300>
   [
     "a",
     null
   ]
   
   # converting NaN -> complains about not being a string or null (error msg 
could be better, though ;))
   >>> pa.array(np.array(["a", np.nan], dtype=object))
   ...
   ArrowTypeError: Expected bytes, got a 'float' object
   
   # unless specifying `from_pandas=True`, and then NaN also gets recognized as 
null
   >>> pa.array(np.array(["a", np.nan], dtype=object), from_pandas=True)
   <pyarrow.lib.StringArray object at 0x7f813bd98ac0>
   [
     "a",
     null
   ]
   ```
   
   So on the one hand, we could argue for consistency that also here NaN should 
only work if `from_pandas=True` is specified. But on the other hand, here it is 
numpy that explicitly recognizes and models it as "missing" (since you have the 
`has_nan_na` vs `has_string_na` on the dtype descr, and the goal of the 
sentinel is for missing values (if not a string sentinel)). So I assume that 
the current behaviour to always treat a non-string sentinel as missing is fine.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to