Hi all,
I went down a bit of a rabbit hole with the aid of an AI model today
related to my ByteStringDType and ongoing work to improve StringDType. It
led me to conclude that there is a missing feature in the DType API. I'd
appreciate it if I could get some feedback about whether I'm on the right
track.
## Problem Statement
Consider this example with np.nan and StringDType:
```python
import numpy as np
from numpy.dtypes import StringDType
text = np.array(["a", "b"], dtype=StringDType(na_object=np.nan))
text[1] = np.nan # stores a missing value
text.tolist() # ['a', nan]
np.where([True, False], text, np.nan) # raises DTypePromotionError
```
Here, `np.where` treats `np.nan` as a float, so it gets promoted as such,
and we end up raising an exception *despite* the `na_object` being set to
`np.nan`.
Other functions lose dtype information or string contents when converting
Python values independently:
```python
words = np.array(["a", "longer"], dtype=StringDType())
np.vectorize(lambda s: s)(words).dtype # dtype('<U6')
np.apply_along_axis(lambda row: row[0], 1,
words[:, None]).tolist() # ['a', 'l']
nul_text = np.array(["x", "x\0"], dtype=StringDType())
np.isin(nul_text, ["x\0"]).tolist() # [True, False]
```
Related problems affect `select`, `append`, `concatenate`, and set
operations.
Let's consider some hypothetical user defined DTypes:
```python
rank = CategoricalDType(["low", "medium", "high"], ordered=True)
levels = np.array(["low", "high"], dtype=rank)
levels[0] # 'low', a Python str
np.sort(levels).tolist() # ['low', 'high']: category order
levels[0] = "medium" # accepted
levels[0] = "urgent" # error: unknown category
```
```python
metres = UnitDType("m")
lengths = np.array([1.0, 2.0], dtype=metres)
lengths[0] = 3.0 # stores three metres
lengths.astype(UnitDType("cm")) # values [300.0, 200.0] in
centimetres
lengths.astype(UnitDType("s")) # error: incompatible units
```
Both the unit and categorical DTypes have a similar issue to the one
StringDType has above: plain python scalars are compatible but rejected
because promotion happens without considering this possibility. For
example, with `np.append` (via `concatenate`)
```python
np.append(levels, ["medium"]) # error: cannot promote rank with Unicode
np.append(lengths, [3.0]) # error: cannot promote metres with float64
```
My reading of NEP 41-43 is that these issues came up but were implicitly
deferred. For example, the discussion about categorical dtypes losing their
dtype when converted back to arrays in NEP 42. I think this mechanism might
be able to help these deferred cases.
## Proposal
I'd like to propose a new DType API hook
`NPY_DT_discover_descr_with_context` and typedef:
```c
int discover_descr_with_context(
npy_intp ndescrs, PyArray_Descr *const descrs[],
PyObject *value, NPY_DTYPE_CONTEXT context, PyArray_Descr **out);
```
For most common things there will only be one input descriptor:
```python
a = np.array(["hello"], dtype=StringDType(na_object=np.nan))
np.where([False], a, np.nan)
```
In this example the one input descriptor is a.dtype. It's also possible to
call `np.vectorize` with an arbitrary number of input arrays, so this takes
an arbitrary number of inputs, but the hook only fires if all inputs have
the same dtype.
The `value` is a Python object that could be a possible scalar for the
dtype instance. In the `np.where` example above, the `value` is `np.nan`.
The NPY_DTYPE_CONTEXT enum tells this function the kind of operation that
is being handled and then `out` is filled in with the appropriate
descriptor for `value`. You can also return 0 to indicate that you have no
special instructions.
The three types of `context` argument I'd like to propose are:
- **OPERAND:** convert one input, then promote it with the other inputs.
- **SEQUENCE_ELEMENT:** require every scalar element to be accepted by the
same DType, combine their settings, then convert the sequence.
- **RESULT:** choose the dtype for a value returned by the user's function.
NumPy calls this hook through the private `_array_converter` methods
`as_arrays(with_context=True)` and `result_type_hint(value)`, retaining
original values until conversion.
StringDType can then recognize strings and its configured missing values
before conversion. On the proposed branch, the earlier examples produce:
```python
np.where([True, False], text, np.nan).tolist() # ['a', nan]
np.vectorize(lambda s: s)(words).dtype # StringDType()
np.apply_along_axis(lambda row: row[0], 1,
words[:, None]).tolist() # ['a', 'longer']
np.isin(nul_text, ["x\0"]).tolist() # [False, True]
```
The hypothetical dtypes could likewise make `append` follow their
assignment rules while continuing to reject implicit conversion of typed
arrays:
```python
# With hooks for the hypothetical dtypes:
np.append(levels, ["medium"]) # ['medium', 'high', 'medium'], dtype=rank
np.append(lengths, [3.0]) # values [3.0, 2.0, 3.0] in metres
```
I've coded (with substantial AI assistance) an [experimental
implementation](
https://github.com/numpy/numpy/compare/main...ngoldbaum:numpy:context-dtype-discover?expand=1)
that includes StringDType support.
_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3//lists/numpy-discussion.python.org
Member address: [email protected]