Public bug reported:

[ Impact ]

 * Affected Ubuntu source versions: openblas 0.3.32+ds-5 in resolute (26.04 
LTS) and 0.3.33+ds-3 in stonking
   (26.10 development). This report requests a backport of the fix already 
released in OpenBLAS 0.3.34.

 * On affected arm64 runtime-dispatch paths, `sdot` / `cblas_sdot` can silently 
return incorrect results for
   unit-stride vectors with n >= 64. The ASIMD vector loop uses v0 as an 
accumulator without initializing it,
   so its result depends on the register's previous contents. The affected LP64 
packages are
   libopenblas0-pthread, libopenblas0-openmp and libopenblas0-serial.

 * The attached reproducer uses a length-64 dot product of two all-ones 
vectors. Loading four lanes of
   1,000,000 into v0 immediately before the call produces 4,000,064 instead of 
64 in the affected library.
   None of the call's arguments is a floating-point value, and v0 is 
caller-saved: its entry contents are not
   specified by the calling convention. In the recorded GCC build, three 
consecutive plain-C calls also
   returned 64, 128, 192 because the previous result remained in s0. That 
second symptom depends on the
   compiler-generated code; the explicit-register check is the primary 
reproducer.

 * Upstream found the bug through SciPy. Its single-precision iterative-solver 
tests exposed stagnation in
   `lgmres` and divergence in `gcrotmk` on a well-conditioned Poisson problem, 
traced to incorrect sdot-based
   orthogonalisation (https://github.com/OpenMathLib/OpenBLAS/issues/5917).

 * The observed defect affects the NEOVERSEN1, THUNDERX2T99 and THUNDERX3T110 
sdot kernels. Which kernel a
   process uses depends on OpenBLAS runtime dispatch and the CPU identity and 
features exposed by the system:
   - NEOVERSEN1: Neoverse N1 (AWS Graviton2, Ampere Altra), AmpereOne, Apple 
silicon when the guest's CPU
     identity selects this kernel (as in the tested M5 VM), and Neoverse 
N2/V1/V2 when SVE is not exposed;
   - THUNDERX2T99: Cavium ThunderX2, Broadcom Vulcan;
   - THUNDERX3T110: ThunderX3.

   Normal SVE dispatch on Graviton3/4 and A64FX avoids this unit-stride defect; 
V1/V2 systems can instead reach
   the affected NEOVERSEN1 fallback when SVE is unavailable. The 
Cortex-A53/A57/A72/A76 and generic ARMV8
   kernels were also unaffected by this reproducer. The recorded test called 
all 56 per-core dot kernels in
   the 0.3.32+ds-5 library (14 cores x sdot, ddot, dsdot, sdsdot) with v0 set. 
Only the three sdot kernels above
   changed their result. ddot, dsdot, sdsdot, strided sdot and n < 64 did not 
show this defect in those tests.

 * Cause: `kernel/arm64/dot_kernel_asimd.c` accumulates the vector loop in 
v0..v7. It zeroes d1..d7 and the
   asm output operand, and relies on the compiler placing that operand in v0. 
In 0.3.32 upstream marked the
   result variable `volatile` (3f6e928d34). In the shipped Ubuntu build, GCC 
places the single-precision output
   operand in s31, leaving v0 uninitialized. Disassembly of that 0.3.32+ds-5 
body at 0xa0daa0 shows
   `fmov s31, wzr; fmov d1, xzr ... fmov d7, xzr`, then `fmla v0.4s, v16.4s, 
v24.4s`. In the recorded rebuild
   with Ubuntu's flags but without the `volatile`, the operand was allocated to 
v0 again.

 * The fix is upstream commit dc3aa2cbd9 (#5918, in 0.3.34), applied as a quilt 
patch. It drops the `volatile`,
   zeroes d0 explicitly, and adds d0 and v16-v31 to the asm clobber list. In 
the recorded test build, GCC
   allocated the single-precision result to s15 and saved/restored the 
containing d15 register. The explicit
   initialization makes the accumulator start at zero regardless of the 
compiler's output-register choice.

[ Test Plan ]

 * New arm64-only autopkgtest `sdot-v0`, with one stanza for each LP64 
threading flavour: pthread, openmp and
   serial. It compiles the attached reproducer against the installed library, 
logs the selected alternatives,
   and runs the core selected by runtime dispatch. It then attempts the three 
kernels where the unit-stride
   sdot defect was observed: THUNDERX2T99, NEOVERSEN1 and THUNDERX3T110, using 
`OPENBLAS_CORETYPE`.
   The script checks selected /proc/cpuinfo feature flags before these forced 
runs and reports which runs it
   skips.

   Upstream's own regression test (the utest hunk) is compiled only when 
`ARMV8` is defined. Debian's
   `DYNAMIC_ARCH TARGET=GENERIC` build does not define it, so that build-time 
regression test is not run.
   The new autopkgtest covers the installed library independently of that 
condition.

 * Manual reproduction on an arm64 machine compatible with an affected
kernel:

       sudo apt install libopenblas0-pthread libopenblas-pthread-dev gcc 
libc6-dev
       gcc -O2 sdot_v0_repro.c -o sdot_v0_repro -lopenblas
       ./sdot_v0_repro                                  # default dispatch; 
affected when it selects one of the kernels above
       OPENBLAS_CORETYPE=NEOVERSEN1 ./sdot_v0_repro      # force on a 
compatible ARMv8.2+ CPU

   Recorded output with 0.3.32+ds-5 and the affected kernel:

       check 1: cblas_sdot(64, ones, 1, ones, 1) with v0 = 1e6 x 4: 4000064.0 
(expected 64.0)
       check 2: the same call three times: 64.0 128.0 192.0 (expected 64.0 64.0 
64.0)
       FAIL: sdot depends on v0 (OpenBLAS #5917)

   With the fixed build, every reported result is 64.0 and the last line is 
`PASS` (exit status 0).
   The three-call sequence on a buggy build can differ with the compiler and 
caller's register contents.

 * Verification from -proposed:
   1. Enable resolute-proposed.
   2. Install the fixed packages:

          sudo apt install -t resolute-proposed
libopenblas0-pthread=0.3.32+ds-5ubuntu0.1 libopenblas-pthread-
dev=0.3.32+ds-5ubuntu0.1

   3. Rerun the reproducer on the default core and, on a compatible CPU, with 
`OPENBLAS_CORETYPE=NEOVERSEN1`.
   4. Install the same proposed version of the -openmp and -serial package 
pairs and repeat for each flavour.
      Each run tests the runtime flavour selected by
      `sudo update-alternatives --config libopenblas.so.0-aarch64-linux-gnu`.
      Record the selected library with `readlink -f 
/usr/lib/aarch64-linux-gnu/libopenblas.so.0`.

 * Regression testing:
   - Run the existing `upstream-testsuite` autopkgtests (BLAS level 1-3 tests) 
on arm64.
   - The patch also changes the ddot kernel's code on the same cores. Check 
that ddot and sdot results do not
     change when v0 is zero. They did not change in the recorded local test 
build (see Other Info).
   - Other architectures build nothing from this file.

[ Where problems could occur ]

 * The patch changes the machine code of the arm64 ASIMD dot kernel for sdot 
and ddot on eight core types:
   NEOVERSEN1, NEOVERSEN2, NEOVERSEV1, ARMV8SVE, A64FX, ARMV9SME, THUNDERX2T99 
and THUNDERX3T110. On the SVE
   cores the kernel is reached only for non-unit strides. A mistake would show 
up as wrong `sdot`/`ddot` results
   on those CPUs. It would also show up in callers of these routines: LAPACK 
routines built on ddot/sdot, and
   NumPy or SciPy through the BLAS alternative.

 * In the recorded GCC build, the new clobber list makes GCC allocate the 
result in s15/d15 and save and
   restore d15. An incorrect save or restore could corrupt the caller's 
preserved register state, causing
   wrong floating-point values in calling code after a call to sdot or ddot.

 * The patch removes the `volatile` that 0.3.32 added to the result variable, 
whose stated purpose was to keep
   compilers from optimizing it out. If a compiler did drop the asm's output, 
sdot would return 0 or a stale
   value. With GCC 15 the output is still computed and returned as before; the 
recorded tests cover that build.

 * The multithreaded path (n > 10000 with several threads) calls the same 
kernel from worker threads. A
   regression there would show up only for long vectors. The attached n=64 
reproducer does not test that path.

 * The new autopkgtest could fail for reasons unrelated to the fix:
   - forcing a core on a runner that lacks instructions used by that build 
could crash with SIGILL; the
     /proc/cpuinfo checks reduce that risk and skipped forced runs must be 
recorded;
   - a toolchain change could keep the previous result out of v0 between the 
plain-C calls. That can hide the
     second symptom on an affected build, so the explicit v0 preload is the 
primary check.

[ Other Info ]

 * Upstream: issue https://github.com/OpenMathLib/OpenBLAS/issues/5917, fix
   https://github.com/OpenMathLib/OpenBLAS/pull/5918 (commit dc3aa2cbd9, 
released in 0.3.34, July 2026).
   - A later upstream commit, 75c11cbe95, is not included. It corrects the 
constraints for modified pointer,
     stride and counter operands, addressing a compiler register-allocation 
hazard demonstrated with LTO.
   - The recorded resolute comparison used GCC 15.2 and Ubuntu's package flags 
without LTO. With those flags,
     dc3aa2cbd9 with or without 75c11cbe95 gave byte-identical objects. That 
comparison has not been made with
     stonking's compiler.

 * Development release first: stonking has 0.3.33+ds-3, which contains the 
affected source and has no Ubuntu delta.
   - Proposed for 26.10, which is in Pre-release Freeze: the second debdiff, 
0.3.33+ds-3ubuntu1, a patch-only
     upload with the same patch and test as the SRU, subject to the applicable 
release review.
   - For the next Ubuntu development series, merge the current Debian package 
and retain the new regression
     test. A sync is appropriate once the Ubuntu changes are included in Debian 
or are otherwise no longer
     needed.
   - If 26.10 releases with 0.3.33+ds-3, it needs the same fix through the SRU 
process, with a newly prepared
     0.3.33+ds-3ubuntu0.1 update instead of this pre-release debdiff.

 * Previously recorded local tests, before any PPA build. Both sources were 
built in an Ubuntu 26.04 arm64 VM
   on an Apple M5.
   - Resolute's source with the patch: built with debian/rules' recipe and 
Ubuntu's build flags
     (`dpkg-buildflags`, DYNAMIC_ARCH, TARGET=GENERIC), pthread flavour only.
   - `make tests`: all BLAS tests passed, and utest reported 127/127 and 
1522/1522.
   - The new test script, run with `sh`: exit 1 against the installed 
0.3.32+ds-5, exit 0 against the build (via
     `LD_LIBRARY_PATH`).
   - The reproducer passes with each of the 14 cores forced. The five SVE cores 
ran under qemu-user `-cpu max`.
   - With v0 = 0, all 56 per-core dot kernels give results bit-identical to 
0.3.32+ds-5 on 24 inputs.
   - The stonking source also builds with the patch, and its test script passes.
   - Not yet done: a PPA build of the three threading variants for both LP64 
and ILP64 (six builds), and an
     autopkgtest run on an arm64 runner. The new `sdot-v0` test exercises the 
LP64 variants.

 * Prepared with Claude Opus 5.5 (an AI model): the analysis, the reproducer, 
the debdiffs and the recorded local
   tests. Report wording and source-application review with OpenAI GPT-6 Astra 
Pro.

** Affects: openblas (Ubuntu)
     Importance: Undecided
         Status: New

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2169719

Title:
  sdot on arm64 adds the caller's v0 register to the result (Neoverse
  N1, AmpereOne, Apple silicon, ThunderX2/3)

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/openblas/+bug/2169719/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to