Hi all,

I went down a bit of a rabbit hole with the aid of an AI model today
related to my ByteStringDType and ongoing work to improve StringDType. It
led me to conclude that there is a missing feature in the DType API. I'd
appreciate it if I could get some feedback about whether I'm on the right
track.

## Problem Statement

Consider this example with np.nan and StringDType:

```python
import numpy as np
from numpy.dtypes import StringDType

text = np.array(["a", "b"], dtype=StringDType(na_object=np.nan))
text[1] = np.nan                       # stores a missing value
text.tolist()                         # ['a', nan]
np.where([True, False], text, np.nan)   # raises DTypePromotionError
```

Here, `np.where` treats `np.nan` as a float, so it gets promoted as such,
and we end up raising an exception *despite* the `na_object` being set to
`np.nan`.

Other functions lose dtype information or string contents when converting
Python values independently:

```python
words = np.array(["a", "longer"], dtype=StringDType())
np.vectorize(lambda s: s)(words).dtype                # dtype('<U6')
np.apply_along_axis(lambda row: row[0], 1,
                    words[:, None]).tolist()        # ['a', 'l']

nul_text = np.array(["x", "x\0"], dtype=StringDType())
np.isin(nul_text, ["x\0"]).tolist()                   # [True, False]
```

Related problems affect `select`, `append`, `concatenate`, and set
operations.

Let's consider some hypothetical user defined DTypes:

```python
rank = CategoricalDType(["low", "medium", "high"], ordered=True)
levels = np.array(["low", "high"], dtype=rank)
levels[0]                             # 'low', a Python str
np.sort(levels).tolist()               # ['low', 'high']: category order
levels[0] = "medium"                   # accepted
levels[0] = "urgent"                   # error: unknown category
```

```python
metres = UnitDType("m")
lengths = np.array([1.0, 2.0], dtype=metres)
lengths[0] = 3.0                       # stores three metres
lengths.astype(UnitDType("cm"))        # values [300.0, 200.0] in
centimetres
lengths.astype(UnitDType("s"))         # error: incompatible units
```

Both the unit and categorical DTypes have a similar issue to the one
StringDType has above: plain python scalars are compatible but rejected
because promotion happens without considering this possibility. For
example, with `np.append` (via `concatenate`)

```python
np.append(levels, ["medium"])  # error: cannot promote rank with Unicode
np.append(lengths, [3.0])      # error: cannot promote metres with float64
```

My reading of NEP 41-43 is that these issues came up but were implicitly
deferred. For example, the discussion about categorical dtypes losing their
dtype when converted back to arrays in NEP 42. I think this mechanism might
be able to help these deferred cases.

## Proposal

I'd like to propose a new DType API hook
`NPY_DT_discover_descr_with_context` and typedef:

```c
int discover_descr_with_context(
    npy_intp ndescrs, PyArray_Descr *const descrs[],
    PyObject *value, NPY_DTYPE_CONTEXT context, PyArray_Descr **out);
```

For most common things there will only be one input descriptor:

```python
a = np.array(["hello"], dtype=StringDType(na_object=np.nan))
np.where([False], a, np.nan)
```

In this example the one input descriptor is a.dtype. It's also possible to
call `np.vectorize` with an arbitrary number of input arrays, so this takes
an arbitrary number of inputs, but the hook only fires if all inputs have
the same dtype.

The `value` is a Python object that could be a possible scalar for the
dtype instance. In the `np.where` example above, the `value` is `np.nan`.

The NPY_DTYPE_CONTEXT enum tells this function the kind of operation that
is being handled and then `out` is filled in with the appropriate
descriptor for `value`. You can also return 0 to indicate that you have no
special instructions.

The three types of `context` argument I'd like to propose are:

- **OPERAND:** convert one input, then promote it with the other inputs.
- **SEQUENCE_ELEMENT:** require every scalar element to be accepted by the
same DType, combine their settings, then convert the sequence.
- **RESULT:** choose the dtype for a value returned by the user's function.

NumPy calls this hook through the private `_array_converter` methods
`as_arrays(with_context=True)` and `result_type_hint(value)`, retaining
original values until conversion.

StringDType can then recognize strings and its configured missing values
before conversion. On the proposed branch, the earlier examples produce:

```python
np.where([True, False], text, np.nan).tolist()         # ['a', nan]
np.vectorize(lambda s: s)(words).dtype                # StringDType()
np.apply_along_axis(lambda row: row[0], 1,
                    words[:, None]).tolist()        # ['a', 'longer']
np.isin(nul_text, ["x\0"]).tolist()                   # [False, True]
```

The hypothetical dtypes could likewise make `append` follow their
assignment rules while continuing to reject implicit conversion of typed
arrays:

```python
# With hooks for the hypothetical dtypes:
np.append(levels, ["medium"])  # ['medium', 'high', 'medium'], dtype=rank
np.append(lengths, [3.0])      # values [3.0, 2.0, 3.0] in metres
```

I've coded (with substantial AI assistance) an [experimental
implementation](
https://github.com/numpy/numpy/compare/main...ngoldbaum:numpy:context-dtype-discover?expand=1)
that includes StringDType support.
_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3//lists/numpy-discussion.python.org
Member address: [email protected]

Reply via email to