Hi Marten, On Tue, Aug 25, 2026 at 2:49 PM Marten van Kerkwijk via NumPy-Discussion < [email protected]> wrote:
> Hi Nathan, > > ByteStringDType looks very nice. The broad outline seems eminently > sensible. My only broad comment is that arguably there is a bit too > much implementation detail in the NEP; much of that might be better done > in the documentation. > Thanks, I'll cut down the implementation details in the NEP text and limit it to describing user-observable behavior. > > On the open questions in the NEP: > > > 1. The type character ``'R'`` and the name ``ByteStringDType``. > > Neither ideal but perhaps the best one can do. I wondered about > ByteArrayDType, but since our dtype is not mutable in the same way, that > is clearly worse. > > > 2. Should setitem also accept buffer-protocol objects (``bytearray``, > > ``memoryview``)? I prefer to defer support to a later iteration to > > reduce the complexity of the initial version. > > Seems fine to postpone. > > > 3. The shared-descriptor-struct approach (vestigial ``coerce``) vs. a > > distinct struct. > > Isn't this an implementation detail? I.e., is any of that visible from > outside of numpy? If so, I think we need accessor functions regardless > of the solution taken. > It's visible because we publicly expose the PyArray_StringDTypeObject struct in NumPy's headers. That means we need to make a choice here about how to represent ByteStringDType instances in the C API. My inclination is to add a typedef that makes PyArray_ByteStringDTypeObject an alias to PyArray_StringDTypeObject and also publicly expose it. We could add accessor functions but the struct is already public. We'd need to do that anyway if anyone ever needs to access the struct fields in the opaque PyObject ABI. > > Somewhat relatedly: I think you currently have StringDType and > ByteStringDType both inherit directly from numpy.dtype. Might it make > sense to have an intermediate VariableLengthDType or so? Though > probably this is again an implementation detail, and I noticed the > existing StrDType and BytesDType also inherit directly from numpy.dtype. > > We could do that. I'm not sure what it buys us since you don't usually want to write code that handles bytes and text identically. > > 4. The ``np.strings.encode`` transition for StringDType input: this NEP > > proposes emitting a ``DeprecationWarning`` now and flipping the > default > > result dtype to ByteStringDType in a later release, after the full > > encoding/errors matrix lands. The flip also changes values whose > > encoding ends in NUL bytes (see :ref:`encoding_decoding`). Ratify the > > flip and its timing, or keep the variable-width result opt-in via > > ``dtype=`` indefinitely? > > I would consider S and U, and StringDType and ByteStringDType pairs, and > feel we should go towards that, so am in favour of your suggestion. I > think StringDType is young enough -- and encoding of it rare enough -- > that deprecation can be relatively fast. > Thanks! See also my reply to Joren where I write a bit about adding a scalar type. I think, on the whole, it is worth doing here, and I've been told enough times that the lack of a scalar type on StringDType is annoying that I'd like to take the opportunity to avoid that trouble for ByteStringDType. It also turns out not to be so bad to implement. > > All the best, > > Marten > > Nathan via NumPy-Discussion <[email protected]> writes: > > > Hi all, > > > > I just opened a PR to add NEP 58, which proposes a new variable-width > > dtype, corresponding to the python bytes type. See > > https://github.com/numpy/numpy/pull/32433. > > > > Per NEP 0, I've copied the NEP content up through "Usage and Impact". > > Please keep substantive comments about the NEP as a whole in this > > thread rather than in the github PR for the NEP text. However, please > > direct specific comments about the text to the PR. > > > > Thanks all for your ideas and input on this, > > > > -Nathan > > > > ===================================== > > NEP 58 — A variable-width bytes DType > > ===================================== > > > > :Author: Nathan Goldbaum > > :Status: Draft > > :Type: Standards Track > > :Created: 2026-08-25 > > > > Abstract > > -------- > > > > I propose ``ByteStringDType``, a variable-width bytes data type: the > bytes > > sibling of ``StringDType`` (:ref:`NEP 55 <NEP55>`). It reuses > StringDType's > > arena-backed storage, allocator, and missing-data machinery while > > > > * storing and returning Python :class:`bytes`, > > * supporting embedded and trailing NUL bytes by construction, and > > * exposing only operations meaningful on raw bytes. > > > > Text and bytes never promote or cast implicitly. The only conversion > between > > ``StringDType`` and ``ByteStringDType`` is via the > > ``np.strings.encode``/``np.strings.decode`` pair. These gain a C UFunc > > implementation, along with the ability to add codec-aware variable-width > loops > > mirroring :meth:`str.encode` and :meth:`bytes.decode`. > > > > A working prototype accompanies this NEP as its reference implementation. > > > > Motivation and scope > > -------------------- > > > > NumPy's only bytes data type is the fixed-width ``S`` dtype > > (``np.dtypes.BytesDType``, scalar ``np.bytes_``). Because ``S`` is > > null-padded, it cannot represent trailing NUL bytes:: > > > > >>> np.array([b"x\x00"])[0] > > b'x' > > > > This makes ``S`` unsuitable for generic byte streams such as binary > record > > formats, encoded blobs, and network data. The truncation cannot be fixed > in > > place: existing code relies on it, so reports have been closed as not > planned > > since 2011 (NumPy `issue #2414 < > https://github.com/numpy/numpy/issues/2414>`__; > > more recently `issue #25268 <https://github.com/numpy/numpy/issues/25268 > >`__, > > which loses the final byte of a SHA-256 digest). :ref:`NEP 55 <NEP55>` > called > > this out, explicitly ruled a bytes/arbitrary-encoding dtype out of its > > scope, and > > listed an improved binary dtype as complementary future work. NumPy > `issue > > #27701 <https://github.com/numpy/numpy/issues/27701>`__ is the open > > request for a > > StringDType equivalent for ``bytes``. This NEP proposes that dtype. > > > > Many downstream libraries fall back to object arrays of ``bytes`` > > wherever they carry variable-width binary data, giving up NumPy's flat > > memory layout and loop machinery. PyArrow materializes every Arrow binary > > column as an `object array of boxed bytes objects > > < > https://github.com/apache/arrow/blob/118892700b95fdfa9a6b3e482a6e5399563f5d75/python/pyarrow/src/arrow/python/arrow_to_pandas.cc#L154-L180 > >`_, > > and the reverse conversion from ``S`` has to guess each element's length > > with ``strnlen``, `truncating at the first NUL > > < > https://github.com/apache/arrow/blob/118892700b95fdfa9a6b3e482a6e5399563f5d75/python/pyarrow/src/arrow/python/numpy_to_arrow.cc#L567-L577 > >`_. > > The > > h5py library reads HDF5 variable-length byte strings `as object arrays > > < > https://github.com/h5py/h5py/blob/821e503405b5e26a1333b28f2b6418d1a2f8c88a/h5py/h5t.templ.pyx#L1900-L1903 > >`_. > > Zarr defines a first-class ``variable_length_bytes`` data type and > > `stores it in object arrays > > < > https://github.com/zarr-developers/zarr-python/blob/20ba31e3e1142fae83b178d6e0a29538c2b18725/src/zarr/core/dtype/npy/bytes.py#L938-L963 > >`_ > > for the same reason. Astropy holds FITS variable-length binary columns > > in object arrays and cannot round-trip fixed-width character columns > > without rewriting their padding (`astropy#11341 > > <https://github.com/astropy/astropy/issues/11341>`__). pandas inherits > > the ``S`` truncation for ``bytes`` columns (`pandas#58205 > > <https://github.com/pandas-dev/pandas/issues/58205>`__). > > > > In scope: > > > > * A variable-width, NUL-transparent bytes DType with the same operation > > *set* as fixed-width ``S`` (ASCII case folding and predicates, > > byte-indexed search/slice), the same missing-data support as > > StringDType, and casts to/from ``S``, void, and bool. > > * Explicit, codec-aware ``encode`` and ``decode`` ufuncs as the only > > text-to-bytes path for the variable-width pair. > > * As a structural side effect, StringDType's UTF-8 assumptions are > > identified and confined to a small encoding-specific surface of the > > implementation. > > > > Out of scope: > > > > * Changing how Python ``bytes`` values are inferred (``np.array([b"x"])`` > > stays fixed-width ``S``, as NEP 55 kept ``str`` inference at ``U``). > > * Exposing other encodings (latin-1, utf-16) as array dtypes. The > > encoding-specific surface identified here is a starting point for an > > encoding-parameterized StringDType, but that is a new user-facing > > semantic that would need its own proposal. > > > > Usage and impact > > ---------------- > > > > Because it uses the same representation and arena-backed storage as > > ``StringDType``, the new ``ByteStringDType`` supports embedded and > trailing > > NUL bytes automatically: > > > > .. code-block:: python > > > > >>> import numpy as np > > >>> from numpy.dtypes import ByteStringDType > > > > >>> a = np.array([b"x\x00", b"a\x00b", b"\xff\xfe"], > > dtype=ByteStringDType()) > > >>> a[0] # trailing NULs survive > > b'x\x00' > > >>> a[1] # embedded NULs too > > b'a\x00b' > > >>> np.strings.str_len(a) # lengths are in bytes, length-explicit > > array([2, 3, 2]) > > >>> np.strings.find(a, b"\x00") > > array([ 1, 1, -1]) > > > > This dtype does not support the ``coerce`` argument that ``StringDType`` > > supports, so data that is not bytes will be rejected by ``np.array()``: > > > > .. code-block:: python > > > > >>> np.array(["text"], dtype=ByteStringDType()) > > Traceback (most recent call last): > > ... > > TypeError: ByteStringDType only allows bytes data, got an instance of > > 'str'; convert text to bytes explicitly with str.encode(encoding) > > > > Converting between ``StringDType`` and ``ByteStringDType`` happens > through > > ``np.strings.encode`` and ``np.strings.decode``. While the ``encode`` > > default transitions (see :ref:`backward_compatibility`), the > ByteStringDType > > result is requested explicitly: > > > > >>> s = np.array(["héllo"], dtype=np.dtypes.StringDType()) > > >>> b = np.strings.encode(s, "utf-8", dtype=ByteStringDType()) > > >>> np.strings.decode(b, "utf-8") > > array(['héllo'], dtype=StringDType()) > > > > .. _backward_compatibility: > > > > Backward compatibility > > ---------------------- > > > > There is only one major backward compatibility concern: dealing with > > ``np.strings.encode``. The function already exists and supports > > ``StringDType``, but sub-optimally in a manner that cannot perserve > trailing NUL > > bytes. This presents some awkward backward compatibility concerns for > > this proposal. In the prototype branch, ``np.strings.encode`` emits a > > ``DeprecationWarning`` for StringDType input when the new ``dtype=`` > argument is > > unspecified. Until the default flips to ByteStringDType in a later > release, its > > behavior is otherwise unchanged: fixed-width ``S`` result, full codec and > > error-mode support, and 0-d arrays for 0-d input. See > :ref:`encoding_decoding` > > for more detail and :ref:`open_questions` for whether this deprecation > should > > happen. > > _______________________________________________ > > NumPy-Discussion mailing list -- [email protected] > > To unsubscribe send an email to [email protected] > > https://mail.python.org/mailman3//lists/numpy-discussion.python.org > > Member address: [email protected] > _______________________________________________ > NumPy-Discussion mailing list -- [email protected] > To unsubscribe send an email to [email protected] > https://mail.python.org/mailman3//lists/numpy-discussion.python.org > Member address: [email protected] >
_______________________________________________ NumPy-Discussion mailing list -- [email protected] To unsubscribe send an email to [email protected] https://mail.python.org/mailman3//lists/numpy-discussion.python.org Member address: [email protected]
