Hi all,

I sent in a PR yesterday updating the NEP text and it was merged overnight.
You can view the updated text on the NEP index:
https://numpy.org/neps/nep-0058-bytestring-dtype.html.

Updates since the last round are:

* Specifying that bytestringdtype will include a `np.vbytes` scalar type
that subclasses `bytes` and `np.generic`
* Closing the "open questions"
* Adding a few new "rejected ideas" to respond to the mailing list
discussion.
* Copy-editing and updating to reflect recently merged work in NumPy `main`.

I've also updated the bytestringdtype prototype to reflect recent updates
and rebased it:
https://github.com/numpy/numpy/compare/main...ngoldbaum:numpy:bytestringdtype
.

Please reply if you continue to have any concerns that you'd like to see
addressed before I call for the NEP to be accepted. This is hopefully the
near-final draft.

-Nathan



On Mon, Aug 31, 2026 at 1:39 PM Nathan <[email protected]> wrote:

> Hi Marten,
>
> On Tue, Aug 25, 2026 at 2:49 PM Marten van Kerkwijk via NumPy-Discussion <
> [email protected]> wrote:
>
>> Hi Nathan,
>>
>> ByteStringDType looks very nice.  The broad outline seems eminently
>> sensible.  My only broad comment is that arguably there is a bit too
>> much implementation detail in the NEP; much of that might be better done
>> in the documentation.
>>
>
> Thanks, I'll cut down the implementation details in the NEP text and limit
> it to describing user-observable behavior.
>
>
>>
>> On the open questions in the NEP:
>>
>> > 1. The type character ``'R'`` and the name ``ByteStringDType``.
>>
>> Neither ideal but perhaps the best one can do.  I wondered about
>> ByteArrayDType, but since our dtype is not mutable in the same way, that
>> is clearly worse.
>>
>> > 2. Should setitem also accept buffer-protocol objects (``bytearray``,
>> >    ``memoryview``)? I prefer to defer support to a later iteration to
>> >    reduce the complexity of the initial version.
>>
>> Seems fine to postpone.
>>
>> > 3. The shared-descriptor-struct approach (vestigial ``coerce``) vs. a
>> >    distinct struct.
>>
>> Isn't this an implementation detail? I.e., is any of that visible from
>> outside of numpy?  If so, I think we need accessor functions regardless
>> of the solution taken.
>>
>
> It's visible because we publicly expose the PyArray_StringDTypeObject
> struct in NumPy's headers. That means we need to make a choice here about
> how to represent ByteStringDType instances in the C API. My inclination is
> to add a typedef that makes PyArray_ByteStringDTypeObject an alias to
> PyArray_StringDTypeObject and also publicly expose it.
>
> We could add accessor functions but the struct is already public. We'd
> need to do that anyway if anyone ever needs to access the struct fields in
> the opaque PyObject ABI.
>
>
>>
>> Somewhat relatedly: I think you currently have StringDType and
>> ByteStringDType both inherit directly from numpy.dtype.  Might it make
>> sense to have an intermediate VariableLengthDType or so?  Though
>> probably this is again an implementation detail, and I noticed the
>> existing StrDType and BytesDType also inherit directly from numpy.dtype.
>>
>>
> We could do that. I'm not sure what it buys us since you don't usually
> want to write code that handles bytes and text identically.
>
>
>> > 4. The ``np.strings.encode`` transition for StringDType input: this NEP
>> >    proposes emitting a ``DeprecationWarning`` now and flipping the
>> default
>> >    result dtype to ByteStringDType in a later release, after the full
>> >    encoding/errors matrix lands. The flip also changes values whose
>> >    encoding ends in NUL bytes (see :ref:`encoding_decoding`). Ratify the
>> >    flip and its timing, or keep the variable-width result opt-in via
>> >    ``dtype=`` indefinitely?
>>
>> I would consider S and U, and StringDType and ByteStringDType pairs, and
>> feel we should go towards that, so am in favour of your suggestion.  I
>> think StringDType is young enough -- and encoding of it rare enough --
>> that deprecation can be relatively fast.
>>
>
> Thanks! See also my reply to Joren where I write a bit about adding a
> scalar type. I think, on the whole, it is worth doing here, and I've been
> told enough times that the lack of a scalar type on StringDType is annoying
> that I'd like to take the opportunity to avoid that trouble for
> ByteStringDType. It also turns out not to be so bad to implement.
>
>
>>
>> All the best,
>>
>> Marten
>>
>> Nathan via NumPy-Discussion <[email protected]> writes:
>>
>> > Hi all,
>> >
>> > I just opened a PR to add NEP 58, which proposes a new variable-width
>> > dtype, corresponding to the python bytes type. See
>> > https://github.com/numpy/numpy/pull/32433.
>> >
>> > Per NEP 0, I've copied the NEP content up through "Usage and Impact".
>> > Please keep substantive comments about the NEP as a whole in this
>> > thread rather than in the github PR for the NEP text. However, please
>> > direct specific comments about the text to the PR.
>> >
>> > Thanks all for your ideas and input on this,
>> >
>> > -Nathan
>> >
>> > =====================================
>> > NEP 58 — A variable-width bytes DType
>> > =====================================
>> >
>> > :Author: Nathan Goldbaum
>> > :Status: Draft
>> > :Type: Standards Track
>> > :Created: 2026-08-25
>> >
>> > Abstract
>> > --------
>> >
>> > I propose ``ByteStringDType``, a variable-width bytes data type: the
>> bytes
>> > sibling of ``StringDType`` (:ref:`NEP 55 <NEP55>`). It reuses
>> StringDType's
>> > arena-backed storage, allocator, and missing-data machinery while
>> >
>> > * storing and returning Python :class:`bytes`,
>> > * supporting embedded and trailing NUL bytes by construction, and
>> > * exposing only operations meaningful on raw bytes.
>> >
>> > Text and bytes never promote or cast implicitly. The only conversion
>> between
>> > ``StringDType`` and ``ByteStringDType`` is via the
>> > ``np.strings.encode``/``np.strings.decode`` pair. These gain a C UFunc
>> > implementation, along with the ability to add codec-aware
>> variable-width loops
>> > mirroring :meth:`str.encode` and :meth:`bytes.decode`.
>> >
>> > A working prototype accompanies this NEP as its reference
>> implementation.
>> >
>> > Motivation and scope
>> > --------------------
>> >
>> > NumPy's only bytes data type is the fixed-width ``S`` dtype
>> > (``np.dtypes.BytesDType``, scalar ``np.bytes_``). Because ``S`` is
>> > null-padded, it cannot represent trailing NUL bytes::
>> >
>> >     >>> np.array([b"x\x00"])[0]
>> >     b'x'
>> >
>> > This makes ``S`` unsuitable for generic byte streams such as binary
>> record
>> > formats, encoded blobs, and network data. The truncation cannot be
>> fixed in
>> > place: existing code relies on it, so reports have been closed as not
>> planned
>> > since 2011 (NumPy `issue #2414 <
>> https://github.com/numpy/numpy/issues/2414>`__;
>> > more recently `issue #25268 <
>> https://github.com/numpy/numpy/issues/25268>`__,
>> > which loses the final byte of a SHA-256 digest). :ref:`NEP 55 <NEP55>`
>> called
>> > this out, explicitly ruled a bytes/arbitrary-encoding dtype out of its
>> > scope, and
>> > listed an improved binary dtype as complementary future work.  NumPy
>> `issue
>> > #27701 <https://github.com/numpy/numpy/issues/27701>`__ is the open
>> > request for a
>> > StringDType equivalent for ``bytes``. This NEP proposes that dtype.
>> >
>> > Many downstream libraries fall back to object arrays of ``bytes``
>> > wherever they carry variable-width binary data, giving up NumPy's flat
>> > memory layout and loop machinery. PyArrow materializes every Arrow
>> binary
>> > column as an `object array of boxed bytes objects
>> > <
>> https://github.com/apache/arrow/blob/118892700b95fdfa9a6b3e482a6e5399563f5d75/python/pyarrow/src/arrow/python/arrow_to_pandas.cc#L154-L180
>> >`_,
>> > and the reverse conversion from ``S`` has to guess each element's length
>> > with ``strnlen``, `truncating at the first NUL
>> > <
>> https://github.com/apache/arrow/blob/118892700b95fdfa9a6b3e482a6e5399563f5d75/python/pyarrow/src/arrow/python/numpy_to_arrow.cc#L567-L577
>> >`_.
>> > The
>> > h5py library reads HDF5 variable-length byte strings `as object arrays
>> > <
>> https://github.com/h5py/h5py/blob/821e503405b5e26a1333b28f2b6418d1a2f8c88a/h5py/h5t.templ.pyx#L1900-L1903
>> >`_.
>> > Zarr defines a first-class ``variable_length_bytes`` data type and
>> > `stores it in object arrays
>> > <
>> https://github.com/zarr-developers/zarr-python/blob/20ba31e3e1142fae83b178d6e0a29538c2b18725/src/zarr/core/dtype/npy/bytes.py#L938-L963
>> >`_
>> > for the same reason.  Astropy holds FITS variable-length binary columns
>> > in object arrays and cannot round-trip fixed-width character columns
>> > without rewriting their padding (`astropy#11341
>> > <https://github.com/astropy/astropy/issues/11341>`__). pandas inherits
>> > the ``S`` truncation for ``bytes`` columns (`pandas#58205
>> > <https://github.com/pandas-dev/pandas/issues/58205>`__).
>> >
>> > In scope:
>> >
>> > * A variable-width, NUL-transparent bytes DType with the same operation
>> >   *set* as fixed-width ``S`` (ASCII case folding and predicates,
>> >   byte-indexed search/slice), the same missing-data support as
>> >   StringDType, and casts to/from ``S``, void, and bool.
>> > * Explicit, codec-aware ``encode`` and ``decode`` ufuncs as the only
>> >   text-to-bytes path for the variable-width pair.
>> > * As a structural side effect, StringDType's UTF-8 assumptions are
>> >   identified and confined to a small encoding-specific surface of the
>> >   implementation.
>> >
>> > Out of scope:
>> >
>> > * Changing how Python ``bytes`` values are inferred
>> (``np.array([b"x"])``
>> >   stays fixed-width ``S``, as NEP 55 kept ``str`` inference at ``U``).
>> > * Exposing other encodings (latin-1, utf-16) as array dtypes. The
>> >   encoding-specific surface identified here is a starting point for an
>> >   encoding-parameterized StringDType, but that is a new user-facing
>> >   semantic that would need its own proposal.
>> >
>> > Usage and impact
>> > ----------------
>> >
>> > Because it uses the same representation and arena-backed storage as
>> > ``StringDType``, the new ``ByteStringDType`` supports embedded and
>> trailing
>> > NUL bytes automatically:
>> >
>> > .. code-block:: python
>> >
>> >     >>> import numpy as np
>> >     >>> from numpy.dtypes import ByteStringDType
>> >
>> >     >>> a = np.array([b"x\x00", b"a\x00b", b"\xff\xfe"],
>> > dtype=ByteStringDType())
>> >     >>> a[0]                      # trailing NULs survive
>> >     b'x\x00'
>> >     >>> a[1]                      # embedded NULs too
>> >     b'a\x00b'
>> >     >>> np.strings.str_len(a)     # lengths are in bytes,
>> length-explicit
>> >     array([2, 3, 2])
>> >     >>> np.strings.find(a, b"\x00")
>> >     array([ 1,  1, -1])
>> >
>> > This dtype does not support the ``coerce`` argument that ``StringDType``
>> > supports, so data that is not bytes will be rejected by ``np.array()``:
>> >
>> > .. code-block:: python
>> >
>> >     >>> np.array(["text"], dtype=ByteStringDType())
>> >     Traceback (most recent call last):
>> >         ...
>> >     TypeError: ByteStringDType only allows bytes data, got an instance
>> of
>> >     'str'; convert text to bytes explicitly with str.encode(encoding)
>> >
>> > Converting between ``StringDType`` and ``ByteStringDType`` happens
>> through
>> > ``np.strings.encode`` and ``np.strings.decode``. While the ``encode``
>> > default transitions (see :ref:`backward_compatibility`), the
>> ByteStringDType
>> > result is requested explicitly:
>> >
>> >     >>> s = np.array(["héllo"], dtype=np.dtypes.StringDType())
>> >     >>> b = np.strings.encode(s, "utf-8", dtype=ByteStringDType())
>> >     >>> np.strings.decode(b, "utf-8")
>> >     array(['héllo'], dtype=StringDType())
>> >
>> > .. _backward_compatibility:
>> >
>> > Backward compatibility
>> > ----------------------
>> >
>> > There is only one major backward compatibility concern: dealing with
>> > ``np.strings.encode``. The function already exists and supports
>> > ``StringDType``, but sub-optimally in a manner that cannot perserve
>> trailing NUL
>> > bytes. This presents some awkward backward compatibility concerns for
>> > this proposal. In the prototype branch, ``np.strings.encode`` emits a
>> > ``DeprecationWarning`` for StringDType input when the new ``dtype=``
>> argument is
>> > unspecified. Until the default flips to ByteStringDType in a later
>> release, its
>> > behavior is otherwise unchanged: fixed-width ``S`` result, full codec
>> and
>> > error-mode support, and 0-d arrays for 0-d input. See
>> :ref:`encoding_decoding`
>> > for more detail and :ref:`open_questions` for whether this deprecation
>> should
>> > happen.
>> > _______________________________________________
>> > NumPy-Discussion mailing list -- [email protected]
>> > To unsubscribe send an email to [email protected]
>> > https://mail.python.org/mailman3//lists/numpy-discussion.python.org
>> > Member address: [email protected]
>> _______________________________________________
>> NumPy-Discussion mailing list -- [email protected]
>> To unsubscribe send an email to [email protected]
>> https://mail.python.org/mailman3//lists/numpy-discussion.python.org
>> Member address: [email protected]
>>
>
_______________________________________________
NumPy-Discussion mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3//lists/numpy-discussion.python.org
Member address: [email protected]

Reply via email to