On Thu, Sep 24, 2026 at 7:33 PM Marten van Kerkwijk via NumPy-Discussion <
[email protected]> wrote:

> Hi Nathan,
>
> I think it makes sense to let the dtypes influence the inferred dtype
> for objects that do not have an explicit descriptor.  For StringDType
> with a nan object as a missing value, it all does seem a bit tricky,
> since it is not just that it suddenly can handle float, but it must
> actually look at the value.  (I think this is related to the
> difficulties Joren mentioned about typing StringDType; it is not about
> the class anymore but about a literal value.)


Yes it’s definitely related. It would be fine if all scalar get and set
operations used numpy scalars and errored otherwise. However nothing in the
DType API requires that there be a NumPy scalar. I also think at least
Sebastian wanted to keep that as a useful option.


>
> I haven't looked at your PR in detail, but I guess the overall question
> is API design, and in particular whether it can be done more simply.


Maybe? I can try to take another look…

I do like how this approach solves lots of different edge cases in the
current design with one new slot. The enum also makes it possibly
extensible in the future.


>
> All new-style dtypes have a `discover_descr_from_pyobject`, which at
> first glance should help.  But I guess the problem is that this is on
> the DType rather than the descriptor?  Obviously, StringDType cannot
> know what na_object a given instance has...
>
> There's also descr.type -- if for StringDType that returned the vstring
> scalar that you are considering elsewhere, would it be able to recognize
> its own na_object?  Could that fact be used?  I.e., could one have a
> fall-back scheme in dtype discovery where, if things do not work out the
> regular way with `discover_descr_from_pyobj`, a second round is done
> whether it is tried to convert the pyobj using descr.pyobj?


I don’t think so, at least not in a way that makes StringDType a drop-in
replacement for object string arrays. Too much code checks for sentinels
with e.g. `is None`


>
> Anyway, not sure that is better, just not quite feeling the proposed API
> is obviously right either, so hoping that bringing up different
> suggestions can be helpful.
>
> All the best,
>
> Marten
>
> p.s.  Note that I don't understand the unit examples: I'd really want
> any assignment to (or concatenation of) an array a dtype with length
> units to raise if one used a float value!  One should indicate the unit
> of the float, not guess that the unit of the array is meant.
>

Personally I think a unit system that allowed inferring units from the
context is a perfectly valid choice for a unit library. It doesn’t give you
all the safety a unit library could possibly give you but it also allows
avoiding a lot of boilerplate when the units are obvious (to my eyes) from
the context.


> Nathan via NumPy-Discussion <[email protected]> writes:
>
> > Hi all,
> >
> > I went down a bit of a rabbit hole with the aid of an AI model today
> related to my ByteStringDType and ongoing work to
> > improve StringDType. It led me to conclude that there is a missing
> feature in the DType API. I'd appreciate it if I could get
> > some feedback about whether I'm on the right track.
> >
> > ## Problem Statement
> >
> > Consider this example with np.nan and StringDType:
> >
> > ```python
> > import numpy as np
> > from numpy.dtypes import StringDType
> >
> > text = np.array(["a", "b"], dtype=StringDType(na_object=np.nan))
> > text[1] = np.nan                       # stores a missing value
> > text.tolist()                         # ['a', nan]
> > np.where([True, False], text, np.nan)   # raises DTypePromotionError
> > ```
> >
> > Here, `np.where` treats `np.nan` as a float, so it gets promoted as
> such, and we end up raising an exception *despite* the
> > `na_object` being set to `np.nan`.
> >
> > Other functions lose dtype information or string contents when
> converting Python values independently:
> >
> > ```python
> > words = np.array(["a", "longer"], dtype=StringDType())
> > np.vectorize(lambda s: s)(words).dtype                # dtype('<U6')
> > np.apply_along_axis(lambda row: row[0], 1,
> >                     words[:, None]).tolist()        # ['a', 'l']
> >
> > nul_text = np.array(["x", "x\0"], dtype=StringDType())
> > np.isin(nul_text, ["x\0"]).tolist()                   # [True, False]
> > ```
> >
> > Related problems affect `select`, `append`, `concatenate`, and set
> operations.
> >
> > Let's consider some hypothetical user defined DTypes:
> >
> > ```python
> > rank = CategoricalDType(["low", "medium", "high"], ordered=True)
> > levels = np.array(["low", "high"], dtype=rank)
> > levels[0]                             # 'low', a Python str
> > np.sort(levels).tolist()               # ['low', 'high']: category order
> > levels[0] = "medium"                   # accepted
> > levels[0] = "urgent"                   # error: unknown category
> > ```
> >
> > ```python
> > metres = UnitDType("m")
> > lengths = np.array([1.0, 2.0], dtype=metres)
> > lengths[0] = 3.0                       # stores three metres
> > lengths.astype(UnitDType("cm"))        # values [300.0, 200.0] in
> centimetres
> > lengths.astype(UnitDType("s"))         # error: incompatible units
> > ```
> >
> > Both the unit and categorical DTypes have a similar issue to the one
> StringDType has above: plain python scalars are
> > compatible but rejected because promotion happens without considering
> this possibility. For example, with `np.append` (via
> > `concatenate`)
> >
> > ```python
> > np.append(levels, ["medium"])  # error: cannot promote rank with Unicode
> > np.append(lengths, [3.0])      # error: cannot promote metres with
> float64
> > ```
> >
> > My reading of NEP 41-43 is that these issues came up but were implicitly
> deferred. For example, the discussion about
> > categorical dtypes losing their dtype when converted back to arrays in
> NEP 42. I think this mechanism might be able to help
> > these deferred cases.
> >
> > ## Proposal
> >
> > I'd like to propose a new DType API hook
> `NPY_DT_discover_descr_with_context` and typedef:
> >
> > ```c
> > int discover_descr_with_context(
> >     npy_intp ndescrs, PyArray_Descr *const descrs[],
> >     PyObject *value, NPY_DTYPE_CONTEXT context, PyArray_Descr **out);
> > ```
> >
> > For most common things there will only be one input descriptor:
> >
> > ```python
> > a = np.array(["hello"], dtype=StringDType(na_object=np.nan))
> > np.where([False], a, np.nan)
> > ```
> >
> > In this example the one input descriptor is a.dtype. It's also possible
> to call `np.vectorize` with an arbitrary number of input
> > arrays, so this takes an arbitrary number of inputs, but the hook only
> fires if all inputs have the same dtype.
> >
> > The `value` is a Python object that could be a possible scalar for the
> dtype instance. In the `np.where` example above, the
> > `value` is `np.nan`.
> >
> > The NPY_DTYPE_CONTEXT enum tells this function the kind of operation
> that is being handled and then `out` is filled in with
> > the appropriate descriptor for `value`. You can also return 0 to
> indicate that you have no special instructions.
> >
> > The three types of `context` argument I'd like to propose are:
> >
> > - **OPERAND:** convert one input, then promote it with the other inputs.
> > - **SEQUENCE_ELEMENT:** require every scalar element to be accepted by
> the same DType, combine their settings, then
> > convert the sequence.
> > - **RESULT:** choose the dtype for a value returned by the user's
> function.
> >
> > NumPy calls this hook through the private `_array_converter` methods
> `as_arrays(with_context=True)` and `result_type_hint
> > (value)`, retaining original values until conversion.
> >
> > StringDType can then recognize strings and its configured missing values
> before conversion. On the proposed branch, the
> > earlier examples produce:
> >
> > ```python
> > np.where([True, False], text, np.nan).tolist()         # ['a', nan]
> > np.vectorize(lambda s: s)(words).dtype                # StringDType()
> > np.apply_along_axis(lambda row: row[0], 1,
> >                     words[:, None]).tolist()        # ['a', 'longer']
> > np.isin(nul_text, ["x\0"]).tolist()                   # [False, True]
> > ```
> >
> > The hypothetical dtypes could likewise make `append` follow their
> assignment rules while continuing to reject implicit
> > conversion of typed arrays:
> >
> > ```python
> > # With hooks for the hypothetical dtypes:
> > np.append(levels, ["medium"])  # ['medium', 'high', 'medium'], dtype=rank
> > np.append(lengths, [3.0])      # values [3.0, 2.0, 3.0] in metres
> > ```
> >
> > I've coded (with substantial AI assistance) an [experimental
> implementation]
> > (
> https://github.com/numpy/numpy/compare/main...ngoldbaum:numpy:context-dtype-discover?expand=1)
> that includes
> > StringDType support.
> > _______________________________________________
> > NumPy-Discussion mailing list -- [email protected]
> > To unsubscribe send an email to [email protected]
> > https://mail.python.org/mailman3//lists/numpy-discussion.python.org
> > Member address: [email protected]
> _______________________________________________
> NumPy-Discussion mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
> https://mail.python.org/mailman3//lists/numpy-discussion.python.org
> Member address: [email protected]
>
_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3//lists/numpy-discussion.python.org
Member address: [email protected]

Reply via email to