Public bug reported: [ Impact ]
* Affected Ubuntu source versions: openblas 0.3.32+ds-5 in resolute (26.04 LTS) and 0.3.33+ds-3 in stonking (26.10 development). This report requests a backport of the fix already released in OpenBLAS 0.3.34. * On affected arm64 runtime-dispatch paths, `sdot` / `cblas_sdot` can silently return incorrect results for unit-stride vectors with n >= 64. The ASIMD vector loop uses v0 as an accumulator without initializing it, so its result depends on the register's previous contents. The affected LP64 packages are libopenblas0-pthread, libopenblas0-openmp and libopenblas0-serial. * The attached reproducer uses a length-64 dot product of two all-ones vectors. Loading four lanes of 1,000,000 into v0 immediately before the call produces 4,000,064 instead of 64 in the affected library. None of the call's arguments is a floating-point value, and v0 is caller-saved: its entry contents are not specified by the calling convention. In the recorded GCC build, three consecutive plain-C calls also returned 64, 128, 192 because the previous result remained in s0. That second symptom depends on the compiler-generated code; the explicit-register check is the primary reproducer. * Upstream found the bug through SciPy. Its single-precision iterative-solver tests exposed stagnation in `lgmres` and divergence in `gcrotmk` on a well-conditioned Poisson problem, traced to incorrect sdot-based orthogonalisation (https://github.com/OpenMathLib/OpenBLAS/issues/5917). * The observed defect affects the NEOVERSEN1, THUNDERX2T99 and THUNDERX3T110 sdot kernels. Which kernel a process uses depends on OpenBLAS runtime dispatch and the CPU identity and features exposed by the system: - NEOVERSEN1: Neoverse N1 (AWS Graviton2, Ampere Altra), AmpereOne, Apple silicon when the guest's CPU identity selects this kernel (as in the tested M5 VM), and Neoverse N2/V1/V2 when SVE is not exposed; - THUNDERX2T99: Cavium ThunderX2, Broadcom Vulcan; - THUNDERX3T110: ThunderX3. Normal SVE dispatch on Graviton3/4 and A64FX avoids this unit-stride defect; V1/V2 systems can instead reach the affected NEOVERSEN1 fallback when SVE is unavailable. The Cortex-A53/A57/A72/A76 and generic ARMV8 kernels were also unaffected by this reproducer. The recorded test called all 56 per-core dot kernels in the 0.3.32+ds-5 library (14 cores x sdot, ddot, dsdot, sdsdot) with v0 set. Only the three sdot kernels above changed their result. ddot, dsdot, sdsdot, strided sdot and n < 64 did not show this defect in those tests. * Cause: `kernel/arm64/dot_kernel_asimd.c` accumulates the vector loop in v0..v7. It zeroes d1..d7 and the asm output operand, and relies on the compiler placing that operand in v0. In 0.3.32 upstream marked the result variable `volatile` (3f6e928d34). In the shipped Ubuntu build, GCC places the single-precision output operand in s31, leaving v0 uninitialized. Disassembly of that 0.3.32+ds-5 body at 0xa0daa0 shows `fmov s31, wzr; fmov d1, xzr ... fmov d7, xzr`, then `fmla v0.4s, v16.4s, v24.4s`. In the recorded rebuild with Ubuntu's flags but without the `volatile`, the operand was allocated to v0 again. * The fix is upstream commit dc3aa2cbd9 (#5918, in 0.3.34), applied as a quilt patch. It drops the `volatile`, zeroes d0 explicitly, and adds d0 and v16-v31 to the asm clobber list. In the recorded test build, GCC allocated the single-precision result to s15 and saved/restored the containing d15 register. The explicit initialization makes the accumulator start at zero regardless of the compiler's output-register choice. [ Test Plan ] * New arm64-only autopkgtest `sdot-v0`, with one stanza for each LP64 threading flavour: pthread, openmp and serial. It compiles the attached reproducer against the installed library, logs the selected alternatives, and runs the core selected by runtime dispatch. It then attempts the three kernels where the unit-stride sdot defect was observed: THUNDERX2T99, NEOVERSEN1 and THUNDERX3T110, using `OPENBLAS_CORETYPE`. The script checks selected /proc/cpuinfo feature flags before these forced runs and reports which runs it skips. Upstream's own regression test (the utest hunk) is compiled only when `ARMV8` is defined. Debian's `DYNAMIC_ARCH TARGET=GENERIC` build does not define it, so that build-time regression test is not run. The new autopkgtest covers the installed library independently of that condition. * Manual reproduction on an arm64 machine compatible with an affected kernel: sudo apt install libopenblas0-pthread libopenblas-pthread-dev gcc libc6-dev gcc -O2 sdot_v0_repro.c -o sdot_v0_repro -lopenblas ./sdot_v0_repro # default dispatch; affected when it selects one of the kernels above OPENBLAS_CORETYPE=NEOVERSEN1 ./sdot_v0_repro # force on a compatible ARMv8.2+ CPU Recorded output with 0.3.32+ds-5 and the affected kernel: check 1: cblas_sdot(64, ones, 1, ones, 1) with v0 = 1e6 x 4: 4000064.0 (expected 64.0) check 2: the same call three times: 64.0 128.0 192.0 (expected 64.0 64.0 64.0) FAIL: sdot depends on v0 (OpenBLAS #5917) With the fixed build, every reported result is 64.0 and the last line is `PASS` (exit status 0). The three-call sequence on a buggy build can differ with the compiler and caller's register contents. * Verification from -proposed: 1. Enable resolute-proposed. 2. Install the fixed packages: sudo apt install -t resolute-proposed libopenblas0-pthread=0.3.32+ds-5ubuntu0.1 libopenblas-pthread- dev=0.3.32+ds-5ubuntu0.1 3. Rerun the reproducer on the default core and, on a compatible CPU, with `OPENBLAS_CORETYPE=NEOVERSEN1`. 4. Install the same proposed version of the -openmp and -serial package pairs and repeat for each flavour. Each run tests the runtime flavour selected by `sudo update-alternatives --config libopenblas.so.0-aarch64-linux-gnu`. Record the selected library with `readlink -f /usr/lib/aarch64-linux-gnu/libopenblas.so.0`. * Regression testing: - Run the existing `upstream-testsuite` autopkgtests (BLAS level 1-3 tests) on arm64. - The patch also changes the ddot kernel's code on the same cores. Check that ddot and sdot results do not change when v0 is zero. They did not change in the recorded local test build (see Other Info). - Other architectures build nothing from this file. [ Where problems could occur ] * The patch changes the machine code of the arm64 ASIMD dot kernel for sdot and ddot on eight core types: NEOVERSEN1, NEOVERSEN2, NEOVERSEV1, ARMV8SVE, A64FX, ARMV9SME, THUNDERX2T99 and THUNDERX3T110. On the SVE cores the kernel is reached only for non-unit strides. A mistake would show up as wrong `sdot`/`ddot` results on those CPUs. It would also show up in callers of these routines: LAPACK routines built on ddot/sdot, and NumPy or SciPy through the BLAS alternative. * In the recorded GCC build, the new clobber list makes GCC allocate the result in s15/d15 and save and restore d15. An incorrect save or restore could corrupt the caller's preserved register state, causing wrong floating-point values in calling code after a call to sdot or ddot. * The patch removes the `volatile` that 0.3.32 added to the result variable, whose stated purpose was to keep compilers from optimizing it out. If a compiler did drop the asm's output, sdot would return 0 or a stale value. With GCC 15 the output is still computed and returned as before; the recorded tests cover that build. * The multithreaded path (n > 10000 with several threads) calls the same kernel from worker threads. A regression there would show up only for long vectors. The attached n=64 reproducer does not test that path. * The new autopkgtest could fail for reasons unrelated to the fix: - forcing a core on a runner that lacks instructions used by that build could crash with SIGILL; the /proc/cpuinfo checks reduce that risk and skipped forced runs must be recorded; - a toolchain change could keep the previous result out of v0 between the plain-C calls. That can hide the second symptom on an affected build, so the explicit v0 preload is the primary check. [ Other Info ] * Upstream: issue https://github.com/OpenMathLib/OpenBLAS/issues/5917, fix https://github.com/OpenMathLib/OpenBLAS/pull/5918 (commit dc3aa2cbd9, released in 0.3.34, July 2026). - A later upstream commit, 75c11cbe95, is not included. It corrects the constraints for modified pointer, stride and counter operands, addressing a compiler register-allocation hazard demonstrated with LTO. - The recorded resolute comparison used GCC 15.2 and Ubuntu's package flags without LTO. With those flags, dc3aa2cbd9 with or without 75c11cbe95 gave byte-identical objects. That comparison has not been made with stonking's compiler. * Development release first: stonking has 0.3.33+ds-3, which contains the affected source and has no Ubuntu delta. - Proposed for 26.10, which is in Pre-release Freeze: the second debdiff, 0.3.33+ds-3ubuntu1, a patch-only upload with the same patch and test as the SRU, subject to the applicable release review. - For the next Ubuntu development series, merge the current Debian package and retain the new regression test. A sync is appropriate once the Ubuntu changes are included in Debian or are otherwise no longer needed. - If 26.10 releases with 0.3.33+ds-3, it needs the same fix through the SRU process, with a newly prepared 0.3.33+ds-3ubuntu0.1 update instead of this pre-release debdiff. * Previously recorded local tests, before any PPA build. Both sources were built in an Ubuntu 26.04 arm64 VM on an Apple M5. - Resolute's source with the patch: built with debian/rules' recipe and Ubuntu's build flags (`dpkg-buildflags`, DYNAMIC_ARCH, TARGET=GENERIC), pthread flavour only. - `make tests`: all BLAS tests passed, and utest reported 127/127 and 1522/1522. - The new test script, run with `sh`: exit 1 against the installed 0.3.32+ds-5, exit 0 against the build (via `LD_LIBRARY_PATH`). - The reproducer passes with each of the 14 cores forced. The five SVE cores ran under qemu-user `-cpu max`. - With v0 = 0, all 56 per-core dot kernels give results bit-identical to 0.3.32+ds-5 on 24 inputs. - The stonking source also builds with the patch, and its test script passes. - Not yet done: a PPA build of the three threading variants for both LP64 and ILP64 (six builds), and an autopkgtest run on an arm64 runner. The new `sdot-v0` test exercises the LP64 variants. * Prepared with Claude Opus 5.5 (an AI model): the analysis, the reproducer, the debdiffs and the recorded local tests. Report wording and source-application review with OpenAI GPT-6 Astra Pro. ** Affects: openblas (Ubuntu) Importance: Undecided Status: New -- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2169719 Title: sdot on arm64 adds the caller's v0 register to the result (Neoverse N1, AmpereOne, Apple silicon, ThunderX2/3) To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/openblas/+bug/2169719/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
