Hi,

On Thu, 2026-09-24 at 19:46 -0600, Nathan via NumPy-Discussion wrote:
> On Thu, Sep 24, 2026 at 7:33 PM Marten van Kerkwijk via NumPy-
> Discussion <
> [email protected]> wrote:
> 
> > Hi Nathan,
> > 
> > I think it makes sense to let the dtypes influence the inferred
> > dtype
> > for objects that do not have an explicit descriptor.  For
> > StringDType
> > with a nan object as a missing value, it all does seem a bit
> > tricky,
> > since it is not just that it suddenly can handle float, but it must
> > actually look at the value.  (I think this is related to the
> > difficulties Joren mentioned about typing StringDType; it is not
> > about
> > the class anymore but about a literal value.)
> 
> 
> Yes it’s definitely related. It would be fine if all scalar get and
> set
> operations used numpy scalars and errored otherwise. However nothing
> in the
> DType API requires that there be a NumPy scalar. I also think at
> least
> Sebastian wanted to keep that as a useful option.

Well, my original plan was a two-pass discovery so that everything can
first check for the right DType and only then discover the precise
dtype instance/descriptor (and in many case you don't have to do the
second pass, because your dtype is non-parametric).
(Mainly, this would be nice for ufuncs, as well)

I guess your approach is somewhat similar? You gather up everything,
then do effectively a second promotion over it. I think one should
worry a bit about the performance impact on simple numerical data, but
not sure.

The second part of what I thought already is possible and might even
solve some of this.  Once you decide for StringDType (which is tricky,
because it is somewhat an invalid promotion, but say you just allow
it), then I had imagined that: StringDType.discover_from_pyobj(nan)`
would just return `StringDType(na_object=nan)` maybe with a "do not
instantiate" flag.  I.e. basically deciding that things are solvable by
carrying the context on weird dtype instances.
(Inconvenient maybe, but can probably do a lot in principle.)

I suppose I am fine with such `info` and all, there is a point that
even in my two-pass scheme one would return the actual dtype instance
if an array was found.
Although, I believe you are basically adding a two-path approach to the
array coercion here, and at that point it feels like a lot of
complexity that does need thought.
And that might also add quite a bit of extra work when it kicks in?

Some of the examples you gave around assignments/Categoricals, look
like they already work (e.g. all of the assignment ones to me seem like
they can, but I am not quite sure).
And I am not immediately sure how much complexity we want to avoid
having to write the `dtype=` here:
```
concatenate([np.nan, arr], dtype=arr.dtype)
# or
concatenate([np.nan, arr], dtype=StringDType)
```

- Sebastian



> 


> > I haven't looked at your PR in detail, but I guess the overall
> > question
> > is API design, and in particular whether it can be done more
> > simply.
> 
> 
> Maybe? I can try to take another look…
> 
> I do like how this approach solves lots of different edge cases in
> the
> current design with one new slot. The enum also makes it possibly
> extensible in the future.
> 
> 
> > 
> > All new-style dtypes have a `discover_descr_from_pyobject`, which
> > at
> > first glance should help.  But I guess the problem is that this is
> > on
> > the DType rather than the descriptor?  Obviously, StringDType
> > cannot
> > know what na_object a given instance has...
> > 
> > There's also descr.type -- if for StringDType that returned the
> > vstring
> > scalar that you are considering elsewhere, would it be able to
> > recognize
> > its own na_object?  Could that fact be used?  I.e., could one have
> > a
> > fall-back scheme in dtype discovery where, if things do not work
> > out the
> > regular way with `discover_descr_from_pyobj`, a second round is
> > done
> > whether it is tried to convert the pyobj using descr.pyobj?
> 
> 
> I don’t think so, at least not in a way that makes StringDType a
> drop-in
> replacement for object string arrays. Too much code checks for
> sentinels
> with e.g. `is None`
> 
> 
> > 
> > Anyway, not sure that is better, just not quite feeling the
> > proposed API
> > is obviously right either, so hoping that bringing up different
> > suggestions can be helpful.
> > 
> > All the best,
> > 
> > Marten
> > 
> > p.s.  Note that I don't understand the unit examples: I'd really
> > want
> > any assignment to (or concatenation of) an array a dtype with
> > length
> > units to raise if one used a float value!  One should indicate the
> > unit
> > of the float, not guess that the unit of the array is meant.
> > 
> 
> Personally I think a unit system that allowed inferring units from
> the
> context is a perfectly valid choice for a unit library. It doesn’t
> give you
> all the safety a unit library could possibly give you but it also
> allows
> avoiding a lot of boilerplate when the units are obvious (to my eyes)
> from
> the context.
> 
> 
> > Nathan via NumPy-Discussion <[email protected]> writes:
> > 
> > > Hi all,
> > > 
> > > I went down a bit of a rabbit hole with the aid of an AI model
> > > today
> > related to my ByteStringDType and ongoing work to
> > > improve StringDType. It led me to conclude that there is a
> > > missing
> > feature in the DType API. I'd appreciate it if I could get
> > > some feedback about whether I'm on the right track.
> > > 
> > > ## Problem Statement
> > > 
> > > Consider this example with np.nan and StringDType:
> > > 
> > > ```python
> > > import numpy as np
> > > from numpy.dtypes import StringDType
> > > 
> > > text = np.array(["a", "b"], dtype=StringDType(na_object=np.nan))
> > > text[1] = np.nan                       # stores a missing value
> > > text.tolist()                         # ['a', nan]
> > > np.where([True, False], text, np.nan)   # raises
> > > DTypePromotionError
> > > ```
> > > 
> > > Here, `np.where` treats `np.nan` as a float, so it gets promoted
> > > as
> > such, and we end up raising an exception *despite* the
> > > `na_object` being set to `np.nan`.
> > > 
> > > Other functions lose dtype information or string contents when
> > converting Python values independently:
> > > 
> > > ```python
> > > words = np.array(["a", "longer"], dtype=StringDType())
> > > np.vectorize(lambda s: s)(words).dtype                #
> > > dtype('<U6')
> > > np.apply_along_axis(lambda row: row[0], 1,
> > >                     words[:, None]).tolist()        # ['a', 'l']
> > > 
> > > nul_text = np.array(["x", "x\0"], dtype=StringDType())
> > > np.isin(nul_text, ["x\0"]).tolist()                   # [True,
> > > False]
> > > ```
> > > 
> > > Related problems affect `select`, `append`, `concatenate`, and
> > > set
> > operations.
> > > 
> > > Let's consider some hypothetical user defined DTypes:
> > > 
> > > ```python
> > > rank = CategoricalDType(["low", "medium", "high"], ordered=True)
> > > levels = np.array(["low", "high"], dtype=rank)
> > > levels[0]                             # 'low', a Python str
> > > np.sort(levels).tolist()               # ['low', 'high']:
> > > category order
> > > levels[0] = "medium"                   # accepted
> > > levels[0] = "urgent"                   # error: unknown category
> > > ```
> > > 
> > > ```python
> > > metres = UnitDType("m")
> > > lengths = np.array([1.0, 2.0], dtype=metres)
> > > lengths[0] = 3.0                       # stores three metres
> > > lengths.astype(UnitDType("cm"))        # values [300.0, 200.0] in
> > centimetres
> > > lengths.astype(UnitDType("s"))         # error: incompatible
> > > units
> > > ```
> > > 
> > > Both the unit and categorical DTypes have a similar issue to the
> > > one
> > StringDType has above: plain python scalars are
> > > compatible but rejected because promotion happens without
> > > considering
> > this possibility. For example, with `np.append` (via
> > > `concatenate`)
> > > 
> > > ```python
> > > np.append(levels, ["medium"])  # error: cannot promote rank with
> > > Unicode
> > > np.append(lengths, [3.0])      # error: cannot promote metres
> > > with
> > float64
> > > ```
> > > 
> > > My reading of NEP 41-43 is that these issues came up but were
> > > implicitly
> > deferred. For example, the discussion about
> > > categorical dtypes losing their dtype when converted back to
> > > arrays in
> > NEP 42. I think this mechanism might be able to help
> > > these deferred cases.
> > > 
> > > ## Proposal
> > > 
> > > I'd like to propose a new DType API hook
> > `NPY_DT_discover_descr_with_context` and typedef:
> > > 
> > > ```c
> > > int discover_descr_with_context(
> > >     npy_intp ndescrs, PyArray_Descr *const descrs[],
> > >     PyObject *value, NPY_DTYPE_CONTEXT context, PyArray_Descr
> > > **out);
> > > ```
> > > 
> > > For most common things there will only be one input descriptor:
> > > 
> > > ```python
> > > a = np.array(["hello"], dtype=StringDType(na_object=np.nan))
> > > np.where([False], a, np.nan)
> > > ```
> > > 
> > > In this example the one input descriptor is a.dtype. It's also
> > > possible
> > to call `np.vectorize` with an arbitrary number of input
> > > arrays, so this takes an arbitrary number of inputs, but the hook
> > > only
> > fires if all inputs have the same dtype.
> > > 
> > > The `value` is a Python object that could be a possible scalar
> > > for the
> > dtype instance. In the `np.where` example above, the
> > > `value` is `np.nan`.
> > > 
> > > The NPY_DTYPE_CONTEXT enum tells this function the kind of
> > > operation
> > that is being handled and then `out` is filled in with
> > > the appropriate descriptor for `value`. You can also return 0 to
> > indicate that you have no special instructions.
> > > 
> > > The three types of `context` argument I'd like to propose are:
> > > 
> > > - **OPERAND:** convert one input, then promote it with the other
> > > inputs.
> > > - **SEQUENCE_ELEMENT:** require every scalar element to be
> > > accepted by
> > the same DType, combine their settings, then
> > > convert the sequence.
> > > - **RESULT:** choose the dtype for a value returned by the user's
> > function.
> > > 
> > > NumPy calls this hook through the private `_array_converter`
> > > methods
> > `as_arrays(with_context=True)` and `result_type_hint
> > > (value)`, retaining original values until conversion.
> > > 
> > > StringDType can then recognize strings and its configured missing
> > > values
> > before conversion. On the proposed branch, the
> > > earlier examples produce:
> > > 
> > > ```python
> > > np.where([True, False], text, np.nan).tolist()         # ['a',
> > > nan]
> > > np.vectorize(lambda s: s)(words).dtype                #
> > > StringDType()
> > > np.apply_along_axis(lambda row: row[0], 1,
> > >                     words[:, None]).tolist()        # ['a',
> > > 'longer']
> > > np.isin(nul_text, ["x\0"]).tolist()                   # [False,
> > > True]
> > > ```
> > > 
> > > The hypothetical dtypes could likewise make `append` follow their
> > assignment rules while continuing to reject implicit
> > > conversion of typed arrays:
> > > 
> > > ```python
> > > # With hooks for the hypothetical dtypes:
> > > np.append(levels, ["medium"])  # ['medium', 'high', 'medium'],
> > > dtype=rank
> > > np.append(lengths, [3.0])      # values [3.0, 2.0, 3.0] in metres
> > > ```
> > > 
> > > I've coded (with substantial AI assistance) an [experimental
> > implementation]
> > > (
> > https://github.com/numpy/numpy/compare/main...ngoldbaum:numpy:context-dtype-discover?expand=1
> > )
> > that includes
> > > StringDType support.
> > > _______________________________________________
> > > NumPy-Discussion mailing list -- [email protected]
> > > To unsubscribe send an email to [email protected]
> > > https://mail.python.org/mailman3//lists/numpy-discussion.python.org
> > > Member address: [email protected]
> > _______________________________________________
> > NumPy-Discussion mailing list -- [email protected]
> > To unsubscribe send an email to [email protected]
> > https://mail.python.org/mailman3//lists/numpy-discussion.python.org
> > Member address: [email protected]
> > 
> _______________________________________________
> NumPy-Discussion mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
> https://mail.python.org/mailman3//lists/numpy-discussion.python.org
> Member address: [email protected]
_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3//lists/numpy-discussion.python.org
Member address: [email protected]

Reply via email to