solrbot opened a new pull request, #4842:
URL: https://github.com/apache/solr/pull/4842

   > ℹ️ **Note**
   > 
   > This PR body was truncated due to platform limits.
   
   This PR contains the following updates:
   
   | Package | Type | Update | Change | Pending |
   |---|---|---|---|---|
   | [org.threeten:threetenbp](https://www.threeten.org/threetenbp) 
([source](https://redirect.github.com/ThreeTen/threetenbp)) | dependencies | 
patch | `1.7.3` → `1.7.4` |  |
   | io.swagger.core.v3.swagger-gradle-plugin | plugin | patch | `2.2.52` → 
`2.2.54` | `2.2.55` |
   | 
[io.swagger.core.v3:swagger-jaxrs2-jakarta](https://redirect.github.com/swagger-api/swagger-core)
 | dependencies | patch | `2.2.52` → `2.2.54` | `2.2.55` |
   | 
[io.swagger.core.v3:swagger-annotations-jakarta](https://redirect.github.com/swagger-api/swagger-core)
 | dependencies | patch | `2.2.52` → `2.2.54` | `2.2.55` |
   | [com.github.spotbugs:spotbugs-annotations](https://spotbugs.github.io/) 
([source](https://redirect.github.com/spotbugs/spotbugs)) | dependencies | 
patch | `4.10.2` → `4.10.4` |  |
   | org.openapi.generator | plugin | minor | `7.23.0` → `7.25.0` |  |
   | 
[com.microsoft.onnxruntime:onnxruntime](https://microsoft.github.io/onnxruntime/)
 ([source](https://redirect.github.com/microsoft/onnxruntime)) | dependencies | 
minor | `1.26.0` → `1.29.0` |  |
   | 
[io.nlopez.compose.rules:ktlint](https://redirect.github.com/mrmans0n/compose-rules)
 | dependencies | patch | `0.6.2` → `0.6.4` |  |
   | 
[no.nav.security:mock-oauth2-server](https://redirect.github.com/navikt/mock-oauth2-server)
 | dependencies | patch | `5.0.1` → `5.0.2` |  |
   | net.ltgt.errorprone | plugin | patch | `5.1.0` → `5.1.1` |  |
   | [dev.logchange](https://redirect.github.com/logchange/logchange) | plugin 
| patch | `1.19.15` → `1.19.16` |  |
   | nl.littlerobots.version-catalog-update | plugin | patch | `1.1.0` → 
`1.1.1` |  |
   | 
[dev.langchain4j:langchain4j-bom](https://redirect.github.com/langchain4j/langchain4j/tree/main/langchain4j-bom)
 
([source](https://redirect.github.com/langchain4j/langchain4j/tree/HEAD/langchain4j-bom))
 | dependencies | minor | `1.17.0` → `1.19.0` |  |
   | [joda-time:joda-time](https://www.joda.org/joda-time/) 
([source](https://redirect.github.com/JodaOrg/joda-time)) | dependencies | 
patch | `2.14.2` → `2.14.3` |  |
   | [org.jctools:jctools-core](https://redirect.github.com/JCTools) 
([source](https://redirect.github.com/JCTools/JCTools)) | dependencies | patch 
| `4.0.6` → `4.0.7` |  |
   | [com.google.guava:guava](https://redirect.github.com/google/guava) | 
dependencies | minor | `33.6.0-jre` → `33.7.1-jre` |  |
   | 
[org.eclipse.jgit:org.eclipse.jgit](https://eclipse.gerrithub.io/admin/repos/eclipse-jgit/jgit)
 | dependencies | patch | `7.7.0.202606012155-r` → `7.7.1.202607240634-r` |  |
   | com.diffplug.spotless | plugin | minor | `8.7.0` → `8.10.0` | `8.10.1` |
   | [com.nvidia.cuvs:cuvs-java](https://rapids.ai) 
([source](https://redirect.github.com/rapidsai/cuvs)) | dependencies | minor | 
`26.06.0` → `26.08.1` |  |
   | 
[commons-codec:commons-codec](https://commons.apache.org/proper/commons-codec/) 
([source](https://redirect.github.com/apache/commons-codec)) | dependencies | 
patch | `1.22.0` → `1.22.1` |  |
   | [org.checkerframework:checker-qual](https://checkerframework.org/) 
([source](https://redirect.github.com/typetools/checker-framework)) | 
dependencies | patch | `4.2.0` → `4.2.2` |  |
   | [com.carrotsearch:hppc](https://redirect.github.com/carrotsearch/hppc) | 
dependencies | minor | `0.10.0` → `0.11.1` |  |
   | 
[org.bouncycastle:bcprov-jdk18on](https://www.bouncycastle.org/download/bouncy-castle-java/)
 ([source](https://redirect.github.com/bcgit/bc-java)) | dependencies | minor | 
`1.84` → `1.85.2` |  |
   | 
[org.bouncycastle:bcpkix-jdk18on](https://www.bouncycastle.org/download/bouncy-castle-java/)
 ([source](https://redirect.github.com/bcgit/bc-java)) | dependencies | minor | 
`1.84` → `1.85` |  |
   | com.github.ben-manes.versions | plugin | minor | `0.54.0` → `0.61.0` |  |
   | [org.apache.tika:tika-core](https://tika.apache.org/) 
([source](https://redirect.github.com/apache/tika)) | dependencies | patch | 
`3.3.1` → `3.3.2` |  |
   | [org.apache.opennlp:opennlp-tools](https://www.apache.org/) 
([source](https://redirect.github.com/apache/opennlp)) | dependencies | patch | 
`2.5.10` → `2.5.11` |  |
   | [org.apache.opennlp:opennlp-dl](https://www.apache.org/) 
([source](https://redirect.github.com/apache/opennlp)) | dependencies | patch | 
`2.5.10` → `2.5.11` |  |
   | 
[org.apache.commons:commons-collections4](https://commons.apache.org/proper/commons-collections/)
 ([source](https://gitbox.apache.org/repos/asf?p=commons-collections.git)) | 
dependencies | minor | `4.5.0` → `4.6.0` |  |
   | 
[com.adobe.testing:s3mock-testcontainers](https://redirect.github.com/adobe/S3Mock)
 | dependencies | minor | `5.1.0` → `5.2.0` |  |
   
   ---
   
   ### Release Notes
   
   <details>
   <summary>ThreeTen/threetenbp (org.threeten:threetenbp)</summary>
   
   ### 
[`v1.7.4`](https://redirect.github.com/ThreeTen/threetenbp/releases/tag/v1.7.4)
   
   See the [change 
notes](https://www.threeten.org/threetenbp/changes-report.html) for more 
information.
   
   </details>
   
   <details>
   <summary>swagger-api/swagger-core 
(io.swagger.core.v3:swagger-jaxrs2-jakarta)</summary>
   
   ### 
[`v2.2.54`](https://redirect.github.com/swagger-api/swagger-core/blob/HEAD/CHANGELOG.md#2254---2026-08-18)
   
   ##### Fixed
   
   - Java 8 date/time types (`OffsetTime`, `Duration`, `LocalTime`) now map by 
default
     to the correct OpenAPI Formats Registry strings (`"time"`, `"duration"`, 
`"time-local"`)
     instead of an unusable expanded object. 
([#&#8203;5172](https://redirect.github.com/swagger-api/swagger-core/issues/5172))
   - `LocalDateTime` deserialization from an existing OpenAPI spec now correctly
     round-trips through the new 
`TimeSchema`/`DurationSchema`/`DateTimeLocalSchema`/
     `TimeLocalSchema` classes instead of falling back to a generic 
`StringSchema`.
   
   ##### Added
   
   - `PrimitiveType.enableJava8Formats()` — opt-in to map `LocalDateTime` to the
     registry-compliant `"date-time-local"` format (default remains 
`"date-time"`
     for backward compatibility).
   
   ##### Deprecated
   
   - `PrimitiveType.enablePartialTime()` — prefer the new default `"time-local"`
     mapping for `LocalTime`; kept for callers who specifically need the
     non-registry `"partial-time"` format.
   
   ### 
[`v2.2.53`](https://redirect.github.com/swagger-api/swagger-core/releases/tag/v2.2.53):
 Swagger-core 2.2.53 released!
   
   - chore: update Jackson to 2.22.1 
([#&#8203;5258](https://redirect.github.com/swagger-api/swagger-core/issues/5258))
   - refactor: replace `writer(new DefaultPrettyPrinter())` with 
`writerWithDefaultPrettyPrinter()` 
([#&#8203;5252](https://redirect.github.com/swagger-api/swagger-core/issues/5252))
   - fix: Stabilize CI Maven and Gradle builds 
([#&#8203;5238](https://redirect.github.com/swagger-api/swagger-core/issues/5238))
   - test: remove system.out.println from tests 
([#&#8203;5236](https://redirect.github.com/swagger-api/swagger-core/issues/5236))
   - chore: bump dependencies 
([#&#8203;5229](https://redirect.github.com/swagger-api/swagger-core/issues/5229))
   - refactor: simplify type handling in ModelDeserializer 
([#&#8203;5227](https://redirect.github.com/swagger-api/swagger-core/issues/5227))
   - Revert "hotfix: temporarily allow Release workflow to skip mvn deploy to 
recover 2.2.52 
([#&#8203;5220](https://redirect.github.com/swagger-api/swagger-core/issues/5220))"
 
([#&#8203;5223](https://redirect.github.com/swagger-api/swagger-core/issues/5223))
   - chore(deps-dev): bump org.codehaus.groovy:groovy from 3.0.23 to 3.0.25 
([#&#8203;5216](https://redirect.github.com/swagger-api/swagger-core/issues/5216))
   - chore(deps): bump commons-cli:commons-cli from 1.9.0 to 1.11.0 
([#&#8203;5209](https://redirect.github.com/swagger-api/swagger-core/issues/5209))
   - chore(deps-dev): bump commons-codec:commons-codec from 1.17.2 to 1.22.0 
([#&#8203;5208](https://redirect.github.com/swagger-api/swagger-core/issues/5208))
   - fix: emit $ref for array items when cycle guard suppresses implementation 
processing 
([#&#8203;5205](https://redirect.github.com/swagger-api/swagger-core/issues/5205))
   - Restore inner property name from map key in handleUnwrapped 
([#&#8203;5193](https://redirect.github.com/swagger-api/swagger-core/issues/5193))
   - Honor PropertyNamingStrategy for get/is-prefixed property names 
([#&#8203;5192](https://redirect.github.com/swagger-api/swagger-core/issues/5192))
   - fix: let explicit 
[@&#8203;Schema](https://redirect.github.com/Schema)(format) override 
type-derived format 
([#&#8203;5185](https://redirect.github.com/swagger-api/swagger-core/issues/5185))
 
([#&#8203;5186](https://redirect.github.com/swagger-api/swagger-core/issues/5186))
   - docs: update format of javadoc to produce a functional link 
([#&#8203;5182](https://redirect.github.com/swagger-api/swagger-core/issues/5182))
   - fix: exclude overridable annotation values when parsing composed 
annotations 
([#&#8203;5179](https://redirect.github.com/swagger-api/swagger-core/issues/5179))
   - fix: negative and positive validation annotations uses relevant OAS 3.1 
syntax ( 
[#&#8203;5170](https://redirect.github.com/swagger-api/swagger-core/issues/5170))
 
([#&#8203;5171](https://redirect.github.com/swagger-api/swagger-core/issues/5171))
   
   </details>
   
   <details>
   <summary>spotbugs/spotbugs 
(com.github.spotbugs:spotbugs-annotations)</summary>
   
   ### 
[`v4.10.4`](https://redirect.github.com/spotbugs/spotbugs/blob/HEAD/CHANGELOG.md#4104---2026-08-19)
   
   [Compare 
Source](https://redirect.github.com/spotbugs/spotbugs/compare/4.10.3...4.10.4)
   
   ##### Fixed
   
   - Fix `NN_NAKED_NOTIFY` false negatives when a field read is stored in a 
local variable before `notify()` or `notifyAll()` 
([#&#8203;3884](https://redirect.github.com/spotbugs/spotbugs/issues/3884))
   - Fix `ASE_ASSERTION_WITH_SIDE_EFFECT` and 
`ASE_ASSERTION_WITH_SIDE_EFFECT_METHOD` false positives in every method 
analysed after a method that reads `$assertionsDisabled` without throwing an 
`AssertionError` 
([#&#8203;3483](https://redirect.github.com/spotbugs/spotbugs/issues/3483))
   - Fix `INT_BAD_COMPARISON_WITH_SIGNED_BYTE` false positive for meaningful 
comparisons of a signed byte with `127` (`b < 127`, `b >= 127`) 
([#&#8203;4201](https://redirect.github.com/spotbugs/spotbugs/pull/4201))
   - Fix `EI_EXPOSE_REP` false negative for public getters in anonymous classes 
([#&#8203;4237](https://redirect.github.com/spotbugs/spotbugs/pull/4237))
   - Fix missing class report for `java.util.Collections$EmptyNavigableSet` and 
`java.util.Collections$EmptyNavigableMap` when the result of 
`Collections.emptySortedSet()`, `emptyNavigableSet()`, `emptySortedMap()` or 
`emptyNavigableMap()` is stored 
([#&#8203;4244](https://redirect.github.com/spotbugs/spotbugs/pull/4244))
   - Fix `URF_UNREAD_FIELD` false negative for unread instance fields declared 
in enums 
([#&#8203;4246](https://redirect.github.com/spotbugs/spotbugs/issues/4246))
   - Stop publishing global dependency-management constraints to consumer POMs. 
([#&#8203;4223](https://redirect.github.com/spotbugs/spotbugs/pull/4223))
   
   ### 
[`v4.10.3`](https://redirect.github.com/spotbugs/spotbugs/blob/HEAD/CHANGELOG.md#4103---2026-07-12)
   
   [Compare 
Source](https://redirect.github.com/spotbugs/spotbugs/compare/4.10.2...4.10.3)
   
   ##### Fixed
   
   - Fix `LI_LAZY_INIT_STATIC` false negative when the null guard is written in 
yoda-style (`null == field`) 
([#&#8203;4144](https://redirect.github.com/spotbugs/spotbugs/pull/4144))
   - Fix `DC_DOUBLECHECK`, `NP_SYNC_AND_NULL_CHECK_FIELD` and 
`SP_SPIN_ON_FIELD` false negatives when the null guard is written in yoda-style 
(`null == field`) 
([#&#8203;4144](https://redirect.github.com/spotbugs/spotbugs/pull/4144))
   - Fix message for `UNS_UNSAFE_CALL` bug pattern
   - Restore CLI plugin loading by fixing DetectorFactoryCollection bootstrap 
ordering 
([#&#8203;4191](https://redirect.github.com/spotbugs/spotbugs/pull/4191))
   - Fix `UWF_NULL_FIELD` false negative for fields initialized with cast null 
values 
([#&#8203;4034](https://redirect.github.com/spotbugs/spotbugs/issues/4034))
   - Fix `UMAC_UNCALLABLE_METHOD_OF_ANONYMOUS_CLASS` false positive for methods 
reached only through method references 
([#&#8203;4059](https://redirect.github.com/spotbugs/spotbugs/pull/4059))
   
   ##### Changed
   
   - Ant `FindBugsViewerTask`: use default look and feel by default. 
([#&#8203;4165](https://redirect.github.com/spotbugs/spotbugs/pull/4165))
   
   ##### Refactor
   
   - Ant `FindBugsViewerTask`: extend `AbstractFindBugsTask` to reduce 
duplicate code. 
([#&#8203;4165](https://redirect.github.com/spotbugs/spotbugs/pull/4165))
   
   </details>
   
   <details>
   <summary>microsoft/onnxruntime 
(com.microsoft.onnxruntime:onnxruntime)</summary>
   
   ### 
[`v1.29.0`](https://redirect.github.com/microsoft/onnxruntime/releases/tag/v1.29.0):
 ONNX Runtime v1.29.0
   
   #### Announcements & Breaking Changes
   
   - onnxruntime-web has announced the deprecation of WebGL and JSEP. The 
native WebGPU EP is the recommended path going forward. See the deprecation and 
migration plans for details 
([#&#8203;29716](https://redirect.github.com/microsoft/onnxruntime/pull/29716), 
[#&#8203;31683](https://redirect.github.com/microsoft/onnxruntime/pull/31683)).
   - POSIX telemetry is now available on Linux, macOS, Android, and iOS when 
ONNX Runtime is built with telemetry enabled. It does not change the public 
ABI, WebAssembly remains telemetry-free, and setting `ORT_DISABLE_TELEMETRY=1` 
before initialization disables non-Windows telemetry for the process 
([#&#8203;27379](https://redirect.github.com/microsoft/onnxruntime/pull/27379), 
[#&#8203;29872](https://redirect.github.com/microsoft/onnxruntime/pull/29872)).
   - The unused internal `onnxruntime/python/tools/tensorrt` dashboard tooling 
was removed. This does not affect the TensorRT Execution Provider APIs 
([#&#8203;29395](https://redirect.github.com/microsoft/onnxruntime/pull/29395)).
   
   #### Security Fixes
   
   ##### Path, bounds, and input validation
   
   - Fixed a path traversal vulnerability in TensorRT and NvTensorRTRTX engine 
refitting by making external-data path validation unconditional 
([#&#8203;29396](https://redirect.github.com/microsoft/onnxruntime/pull/29396)).
   - Validated the CPU MoE `k` attribute against the number of experts and 
fixed a CPU `TensorScatter` security issue 
([#&#8203;29907](https://redirect.github.com/microsoft/onnxruntime/pull/29907), 
[#&#8203;29916](https://redirect.github.com/microsoft/onnxruntime/pull/29916)).
   - Added missing rank, shape, and parameter validation for pooling, LSTM and 
DynamicQuantizeLSTM, Sampling, FeatureVectorizer, SkipLayerNorm, QLinearConv, 
Whisper decoding, RNN activations, GridSample, contrib `Range`, and 
`CropAndResize` 
([#&#8203;29254](https://redirect.github.com/microsoft/onnxruntime/pull/29254), 
[#&#8203;29255](https://redirect.github.com/microsoft/onnxruntime/pull/29255), 
[#&#8203;29265](https://redirect.github.com/microsoft/onnxruntime/pull/29265), 
[#&#8203;29579](https://redirect.github.com/microsoft/onnxruntime/pull/29579), 
[#&#8203;29595](https://redirect.github.com/microsoft/onnxruntime/pull/29595), 
[#&#8203;29605](https://redirect.github.com/microsoft/onnxruntime/pull/29605), 
[#&#8203;29871](https://redirect.github.com/microsoft/onnxruntime/pull/29871), 
[#&#8203;31636](https://redirect.github.com/microsoft/onnxruntime/pull/31636), 
[#&#8203;31671](https://redirect.github.com/microsoft/onnxruntime/pull/31671), 
[#&#8203;31675](https://redirect.github.com/m
 icrosoft/onnxruntime/pull/31675), 
[#&#8203;31676](https://redirect.github.com/microsoft/onnxruntime/pull/31676), 
[#&#8203;31684](https://redirect.github.com/microsoft/onnxruntime/pull/31684)).
   - Hardened CUDA indexing and buffer handling in GridSample, transpose, 
GatherBlockQuantized, InstanceNormalization, LayerNorm/RMSNorm, BeamSearch, 
DeformConv, AveragePool, and MaxPool 
([#&#8203;29581](https://redirect.github.com/microsoft/onnxruntime/pull/29581), 
[#&#8203;29631](https://redirect.github.com/microsoft/onnxruntime/pull/29631), 
[#&#8203;29638](https://redirect.github.com/microsoft/onnxruntime/pull/29638), 
[#&#8203;31640](https://redirect.github.com/microsoft/onnxruntime/pull/31640), 
[#&#8203;31642](https://redirect.github.com/microsoft/onnxruntime/pull/31642), 
[#&#8203;31644](https://redirect.github.com/microsoft/onnxruntime/pull/31644), 
[#&#8203;31645](https://redirect.github.com/microsoft/onnxruntime/pull/31645), 
[#&#8203;31647](https://redirect.github.com/microsoft/onnxruntime/pull/31647), 
[#&#8203;31650](https://redirect.github.com/microsoft/onnxruntime/pull/31650)).
   - Fixed packed sub-byte tensor over-copying in `OrtApi::GetValue` and 
validated DML constant tensor byte sizes 
([#&#8203;29157](https://redirect.github.com/microsoft/onnxruntime/pull/29157), 
[#&#8203;31665](https://redirect.github.com/microsoft/onnxruntime/pull/31665)).
   
   ##### Supply chain and tooling
   
   - Updated npm lockfiles, refreshed the Next.js end-to-end fixture lockfile 
for security advisories, and upgraded `adm-zip` for `onnxruntime-node` 
([#&#8203;29827](https://redirect.github.com/microsoft/onnxruntime/pull/29827), 
[#&#8203;29926](https://redirect.github.com/microsoft/onnxruntime/pull/29926), 
[#&#8203;31192](https://redirect.github.com/microsoft/onnxruntime/pull/31192)).
   
   #### New Features
   
   ##### Core APIs & Runtime
   
   - Default intra-op and inter-op thread-pool sizes can now be set with 
`ORT_INTRA_OP_NUM_THREADS` and `ORT_INTER_OP_NUM_THREADS`. Explicit thread 
settings still take precedence, and `0` preserves machine-sized defaults 
([#&#8203;29688](https://redirect.github.com/microsoft/onnxruntime/pull/29688)).
   - Added weightless-model support for all initializer types, allowed 
zero-input `EpContext` nodes, and wired maximum-shape inference into workspace 
estimation 
([#&#8203;29607](https://redirect.github.com/microsoft/onnxruntime/pull/29607), 
[#&#8203;29799](https://redirect.github.com/microsoft/onnxruntime/pull/29799), 
[#&#8203;31613](https://redirect.github.com/microsoft/onnxruntime/pull/31613)).
   - Added ONNX-domain support for rotary embedding and a fused 
`MRotaryEmbedding` contrib operator for Qwen mRoPE variants 
([#&#8203;29261](https://redirect.github.com/microsoft/onnxruntime/pull/29261), 
[#&#8203;31728](https://redirect.github.com/microsoft/onnxruntime/pull/31728)).
   - Added multi-shape profiling to `onnxruntime_perf_test` through 
`--data_shape`, plus verbose graph-transformer tracing and broader 
inference-session error-path coverage 
([#&#8203;29555](https://redirect.github.com/microsoft/onnxruntime/pull/29555), 
[#&#8203;29558](https://redirect.github.com/microsoft/onnxruntime/pull/29558), 
[#&#8203;29569](https://redirect.github.com/microsoft/onnxruntime/pull/29569), 
[#&#8203;29571](https://redirect.github.com/microsoft/onnxruntime/pull/29571)).
   
   ##### Execution Provider ABI & Plugin EPs
   
   - WebGPU now supports device-free compile-only sessions for offline graph 
transformation 
([#&#8203;29681](https://redirect.github.com/microsoft/onnxruntime/pull/29681)).
   - Expanded CUDA plugin EP packaging and testing, including Windows ARM64 
package and size options, updated package outputs, and aligned architecture 
selections across Python, C API, TensorRT, Node.js, and plugin packages 
([#&#8203;31635](https://redirect.github.com/microsoft/onnxruntime/pull/31635), 
[#&#8203;31722](https://redirect.github.com/microsoft/onnxruntime/pull/31722), 
[#&#8203;31992](https://redirect.github.com/microsoft/onnxruntime/pull/31992)).
   - Improved plugin lifecycle handling by unloading failed EP library loads 
and fixing allocator-deleter lifetime 
([#&#8203;29634](https://redirect.github.com/microsoft/onnxruntime/pull/29634), 
[#&#8203;29770](https://redirect.github.com/microsoft/onnxruntime/pull/29770)).
   
   #### Execution Provider Updates
   
   ##### NVIDIA CUDA EP
   
   ##### Attention and decoding
   
   - Added `PagedAttention` with quantized KV cache, XQA decode, MLA, QK-Norm, 
and head-sink support 
([#&#8203;29912](https://redirect.github.com/microsoft/onnxruntime/pull/29912)).
   - Extended quantized KV-cache support with attention sinks, independent and 
per-channel scales, sliding-window cache support, and a fused K/V 
dequantization launch 
([#&#8203;29900](https://redirect.github.com/microsoft/onnxruntime/pull/29900), 
[#&#8203;29904](https://redirect.github.com/microsoft/onnxruntime/pull/29904), 
[#&#8203;31480](https://redirect.github.com/microsoft/onnxruntime/pull/31480)).
   - Added a cuDNN SDPA decode tier to the standard ONNX `Attention` CUDA 
kernel and enabled cuDNN SDPA for contrib `Attention` 
([#&#8203;29715](https://redirect.github.com/microsoft/onnxruntime/pull/29715), 
[#&#8203;29717](https://redirect.github.com/microsoft/onnxruntime/pull/29717)).
   - Added `attention_bias` support to the GroupQueryAttention unfused path and 
`state_window` support to LinearAttention and CausalConvWithState for MTP 
([#&#8203;29525](https://redirect.github.com/microsoft/onnxruntime/pull/29525), 
[#&#8203;31157](https://redirect.github.com/microsoft/onnxruntime/pull/31157)).
   - Fixed LinearAttention on GPUs with limited shared memory 
([#&#8203;31982](https://redirect.github.com/microsoft/onnxruntime/pull/31982)).
   
   ##### MoE and quantized GEMM
   
   - Added NVFP4 QMoE, including native FP4xFP4 prefill on SM120, faster decode 
GEMV, fused routing/finalization paths, and reduced activation and 
weight-dequantization overhead 
([#&#8203;29697](https://redirect.github.com/microsoft/onnxruntime/pull/29697), 
[#&#8203;29824](https://redirect.github.com/microsoft/onnxruntime/pull/29824), 
[#&#8203;29887](https://redirect.github.com/microsoft/onnxruntime/pull/29887), 
[#&#8203;29919](https://redirect.github.com/microsoft/onnxruntime/pull/29919), 
[#&#8203;31156](https://redirect.github.com/microsoft/onnxruntime/pull/31156), 
[#&#8203;31159](https://redirect.github.com/microsoft/onnxruntime/pull/31159), 
[#&#8203;31349](https://redirect.github.com/microsoft/onnxruntime/pull/31349), 
[#&#8203;31479](https://redirect.github.com/microsoft/onnxruntime/pull/31479)).
   - Added `MatMulBlockQuantizedFp4Weight` and `MatMulBlockQuantizedFp8Weight`, 
plus block-scaled tensor-core/GEMV decode paths, packed FP4 decode, M-tiling, 
and folded W8A8 activation QDQ 
([#&#8203;29818](https://redirect.github.com/microsoft/onnxruntime/pull/29818), 
[#&#8203;29850](https://redirect.github.com/microsoft/onnxruntime/pull/29850), 
[#&#8203;29896](https://redirect.github.com/microsoft/onnxruntime/pull/29896), 
[#&#8203;31155](https://redirect.github.com/microsoft/onnxruntime/pull/31155), 
[#&#8203;31481](https://redirect.github.com/microsoft/onnxruntime/pull/31481)).
   - Improved MatMulNBits and QMoE robustness and efficiency by optimizing 
8-bit dequantization, releasing raw MXFP4 initializers after prepack, and 
fixing subgraph prepacking and mixed FP8/FP4 build failures 
([#&#8203;29852](https://redirect.github.com/microsoft/onnxruntime/pull/29852), 
[#&#8203;31141](https://redirect.github.com/microsoft/onnxruntime/pull/31141), 
[#&#8203;31154](https://redirect.github.com/microsoft/onnxruntime/pull/31154), 
[#&#8203;31350](https://redirect.github.com/microsoft/onnxruntime/pull/31350)).
   
   ##### Operators and collectives
   
   - Added `LinearAttentionGate`, `GatedRMSNorm`, and `GatedAdd` contrib 
operators 
([#&#8203;31158](https://redirect.github.com/microsoft/onnxruntime/pull/31158), 
[#&#8203;31835](https://redirect.github.com/microsoft/onnxruntime/pull/31835)).
   - Added bfloat16 support to `AllReduce`, `AllGather`, and `AllToAll` 
([#&#8203;31571](https://redirect.github.com/microsoft/onnxruntime/pull/31571)).
   - Fixed the default zero point in CUDA `GatherBlockQuantized` 
([#&#8203;31693](https://redirect.github.com/microsoft/onnxruntime/pull/31693)).
   
   ##### WebGPU EP
   
   - Added DFT, HardSwish, Max/Min, Trilu, GRU, PRelu, MatMulBnb4, and 
MRotaryEmbedding support 
([#&#8203;29454](https://redirect.github.com/microsoft/onnxruntime/pull/29454), 
[#&#8203;29587](https://redirect.github.com/microsoft/onnxruntime/pull/29587), 
[#&#8203;29828](https://redirect.github.com/microsoft/onnxruntime/pull/29828), 
[#&#8203;29833](https://redirect.github.com/microsoft/onnxruntime/pull/29833), 
[#&#8203;29840](https://redirect.github.com/microsoft/onnxruntime/pull/29840), 
[#&#8203;29845](https://redirect.github.com/microsoft/onnxruntime/pull/29845), 
[#&#8203;30512](https://redirect.github.com/microsoft/onnxruntime/pull/30512), 
[#&#8203;31976](https://redirect.github.com/microsoft/onnxruntime/pull/31976)).
   - Expanded integer support across Clip, Reshape, Cast, Add, Tile, Concat, 
Expand, Gather, CumSum, Max, and Min 
([#&#8203;29830](https://redirect.github.com/microsoft/onnxruntime/pull/29830), 
[#&#8203;29834](https://redirect.github.com/microsoft/onnxruntime/pull/29834), 
[#&#8203;29835](https://redirect.github.com/microsoft/onnxruntime/pull/29835), 
[#&#8203;29839](https://redirect.github.com/microsoft/onnxruntime/pull/29839), 
[#&#8203;29844](https://redirect.github.com/microsoft/onnxruntime/pull/29844), 
[#&#8203;29847](https://redirect.github.com/microsoft/onnxruntime/pull/29847), 
[#&#8203;29854](https://redirect.github.com/microsoft/onnxruntime/pull/29854), 
[#&#8203;29861](https://redirect.github.com/microsoft/onnxruntime/pull/29861), 
[#&#8203;29897](https://redirect.github.com/microsoft/onnxruntime/pull/29897), 
[#&#8203;29918](https://redirect.github.com/microsoft/onnxruntime/pull/29918), 
[#&#8203;31049](https://redirect.github.com/microsoft/onnxruntime/pull/31049), 
[#&#8203;31702
 ](https://redirect.github.com/microsoft/onnxruntime/pull/31702), 
[#&#8203;31709](https://redirect.github.com/microsoft/onnxruntime/pull/31709)).
   - Added the initial WebGPU PagedAttention implementation and moved Softmax 
and non-flash Attention to online algorithms 
([#&#8203;29694](https://redirect.github.com/microsoft/onnxruntime/pull/29694), 
[#&#8203;29724](https://redirect.github.com/microsoft/onnxruntime/pull/29724), 
[#&#8203;31611](https://redirect.github.com/microsoft/onnxruntime/pull/31611)).
   - Improved MatMulNBits wide-tile accumulation precision 
([#&#8203;29611](https://redirect.github.com/microsoft/onnxruntime/pull/29611)).
   - Added and extended Intel subgroup-matrix MatMul/Gemm kernels, including 
f16, batched-B, and odd-N support 
([#&#8203;29592](https://redirect.github.com/microsoft/onnxruntime/pull/29592), 
[#&#8203;29749](https://redirect.github.com/microsoft/onnxruntime/pull/29749), 
[#&#8203;29813](https://redirect.github.com/microsoft/onnxruntime/pull/29813), 
[#&#8203;29893](https://redirect.github.com/microsoft/onnxruntime/pull/29893)).
   - Reduced cold-start and upload overhead with deferred dispatch and 
staging-buffer improvements; tuned FlashAttention, subgroup Gemm/MatMul, 
Split-K on Panther Lake, and Xe im2col-matmul 
([#&#8203;29271](https://redirect.github.com/microsoft/onnxruntime/pull/29271), 
[#&#8203;29505](https://redirect.github.com/microsoft/onnxruntime/pull/29505), 
[#&#8203;29557](https://redirect.github.com/microsoft/onnxruntime/pull/29557), 
[#&#8203;29586](https://redirect.github.com/microsoft/onnxruntime/pull/29586), 
[#&#8203;29846](https://redirect.github.com/microsoft/onnxruntime/pull/29846), 
[#&#8203;30514](https://redirect.github.com/microsoft/onnxruntime/pull/30514)).
   - Upgraded Dawn and improved reliability by avoiding exceptions in Dawn 
callbacks, fixing a Linux adapter-failure self-deadlock, correcting Windows x86 
transfer callbacks, and fixing TurboQuant batched sequence lengths 
([#&#8203;29389](https://redirect.github.com/microsoft/onnxruntime/pull/29389), 
[#&#8203;29591](https://redirect.github.com/microsoft/onnxruntime/pull/29591), 
[#&#8203;29625](https://redirect.github.com/microsoft/onnxruntime/pull/29625), 
[#&#8203;29752](https://redirect.github.com/microsoft/onnxruntime/pull/29752), 
[#&#8203;31568](https://redirect.github.com/microsoft/onnxruntime/pull/31568)).
   
   ##### WebNN EP
   
   - Added uint8-packed 4-bit `GatherBlockQuantized` and `LpNormalization`, 
reused the shared WASM loader for Blob-backed external data, and fixed per-axis 
QDQ and MatMulNBits edge cases 
([#&#8203;29475](https://redirect.github.com/microsoft/onnxruntime/pull/29475), 
[#&#8203;29801](https://redirect.github.com/microsoft/onnxruntime/pull/29801), 
[#&#8203;31151](https://redirect.github.com/microsoft/onnxruntime/pull/31151), 
[#&#8203;31152](https://redirect.github.com/microsoft/onnxruntime/pull/31152), 
[#&#8203;31197](https://redirect.github.com/microsoft/onnxruntime/pull/31197)).
   
   ##### OpenVINO / QNN / DML / XNNPACK / TensorRT
   
   - OpenVINO fixed float16 constant-output corruption and output-name routing, 
added dot-separated KV-cache names to the stateful transform, and corrected 
raw-data-backed float initializer handling 
([#&#8203;29729](https://redirect.github.com/microsoft/onnxruntime/pull/29729), 
[#&#8203;29882](https://redirect.github.com/microsoft/onnxruntime/pull/29882), 
[#&#8203;29895](https://redirect.github.com/microsoft/onnxruntime/pull/29895), 
[#&#8203;31138](https://redirect.github.com/microsoft/onnxruntime/pull/31138)).
   - QNN added a reshape handler for split-axis reshapes 
([#&#8203;29660](https://redirect.github.com/microsoft/onnxruntime/pull/29660)).
   - DML fixed wide-string handling and made fused graph kernels own their 
model paths 
([#&#8203;31656](https://redirect.github.com/microsoft/onnxruntime/pull/31656), 
[#&#8203;31664](https://redirect.github.com/microsoft/onnxruntime/pull/31664)).
   - XNNPACK now reads dynamic Gemm `M` from the input tensor at compute time 
([#&#8203;31189](https://redirect.github.com/microsoft/onnxruntime/pull/31189)).
   - TensorRT deduplicated context-path handling and added a build option for 
fused-attention cubins 
([#&#8203;29640](https://redirect.github.com/microsoft/onnxruntime/pull/29640), 
[#&#8203;31632](https://redirect.github.com/microsoft/onnxruntime/pull/31632)).
   
   #### CPU & Core Optimizations
   
   ##### MLAS
   
   - Added Arm64 half-precision GEMM and convolution support through KleidiAI, 
including FP16 MatMul/Gemm/Conv paths and asymmetric Q4 and SME2 MatMulNBits 
kernels 
([#&#8203;28786](https://redirect.github.com/microsoft/onnxruntime/pull/28786), 
[#&#8203;29654](https://redirect.github.com/microsoft/onnxruntime/pull/29654), 
[#&#8203;29709](https://redirect.github.com/microsoft/onnxruntime/pull/29709), 
[#&#8203;29898](https://redirect.github.com/microsoft/onnxruntime/pull/29898)).
   - Added a RISC-V RVV QNBitGemm backend, an Arm64 NEON fp32 RoPE kernel, 
portable SVE elementwise kernels with FEXPA exp, and Arm64 UDOT routing for 
S8U8 QGEMM 
([#&#8203;29537](https://redirect.github.com/microsoft/onnxruntime/pull/29537), 
[#&#8203;29787](https://redirect.github.com/microsoft/onnxruntime/pull/29787), 
[#&#8203;29836](https://redirect.github.com/microsoft/onnxruntime/pull/29836), 
[#&#8203;31145](https://redirect.github.com/microsoft/onnxruntime/pull/31145)).
   - Added AVX2/VNNI 2-bit weight kernels and vectorized 2-bit dequantization, 
and improved fp16 MatMulNBits paths by avoiding fp32 temporaries and writing 
fp16 output directly across 2-, 4-, and 8-bit paths 
([#&#8203;29619](https://redirect.github.com/microsoft/onnxruntime/pull/29619), 
[#&#8203;29766](https://redirect.github.com/microsoft/onnxruntime/pull/29766), 
[#&#8203;29791](https://redirect.github.com/microsoft/onnxruntime/pull/29791), 
[#&#8203;29842](https://redirect.github.com/microsoft/onnxruntime/pull/29842), 
[#&#8203;29864](https://redirect.github.com/microsoft/onnxruntime/pull/29864), 
[#&#8203;29901](https://redirect.github.com/microsoft/onnxruntime/pull/29901)).
   
   ##### CPU Attention & Kernels
   
   - Improved masked Attention performance, enabled CPU FlashAttention on Linux 
Arm64 through L2-cache detection, and added FP16 GQA with quantized KV cache 
([#&#8203;29621](https://redirect.github.com/microsoft/onnxruntime/pull/29621), 
[#&#8203;29719](https://redirect.github.com/microsoft/onnxruntime/pull/29719), 
[#&#8203;29825](https://redirect.github.com/microsoft/onnxruntime/pull/29825)).
   - Added double support to CPU `Cos` and int32 support to CPU `Trilu`, and 
fixed int8 QLinearSoftmax saturation and AvgPool 
`ceil_mode`/`count_include_pad` behavior 
([#&#8203;28975](https://redirect.github.com/microsoft/onnxruntime/pull/28975), 
[#&#8203;29476](https://redirect.github.com/microsoft/onnxruntime/pull/29476), 
[#&#8203;29629](https://redirect.github.com/microsoft/onnxruntime/pull/29629), 
[#&#8203;29728](https://redirect.github.com/microsoft/onnxruntime/pull/29728)).
   - Fixed `TfIdfVectorizer` weight indexing and skipped MinLength 
logits-processor construction when `eos_token_id` is negative 
([#&#8203;29604](https://redirect.github.com/microsoft/onnxruntime/pull/29604), 
[#&#8203;31649](https://redirect.github.com/microsoft/onnxruntime/pull/31649)).
   - Tightened K/V and cache-indirection shape contracts in CPU Attention and 
MultiHeadAttention, and fixed LinearAttention output shape inference for 
grouped-query attention 
([#&#8203;29892](https://redirect.github.com/microsoft/onnxruntime/pull/29892), 
[#&#8203;31190](https://redirect.github.com/microsoft/onnxruntime/pull/31190), 
[#&#8203;31634](https://redirect.github.com/microsoft/onnxruntime/pull/31634)).
   
   ##### Graph, Optimizer, and Runtime
   
   - Extended reshape fusion, fixed double recursion in subgraph type/shape 
inference, and made constant-folding output deterministic 
([#&#8203;29027](https://redirect.github.com/microsoft/onnxruntime/pull/29027), 
[#&#8203;29617](https://redirect.github.com/microsoft/onnxruntime/pull/29617), 
[#&#8203;29789](https://redirect.github.com/microsoft/onnxruntime/pull/29789)).
   - Fixed in-memory external initializer loading, memory-pattern allocation 
stream selection, and a leak in `GetOverridableInitializerNames()` 
([#&#8203;29349](https://redirect.github.com/microsoft/onnxruntime/pull/29349), 
[#&#8203;29589](https://redirect.github.com/microsoft/onnxruntime/pull/29589), 
[#&#8203;29616](https://redirect.github.com/microsoft/onnxruntime/pull/29616)).
   - Reduced small MatMul batch allocations and redundant LUT initialization 
([#&#8203;29085](https://redirect.github.com/microsoft/onnxruntime/pull/29085), 
[#&#8203;29690](https://redirect.github.com/microsoft/onnxruntime/pull/29690)).
   - Fixed static-initialization-order crashes when importing ONNX Runtime and 
reduced eager runtime initialization 
([#&#8203;29880](https://redirect.github.com/microsoft/onnxruntime/pull/29880), 
[#&#8203;31964](https://redirect.github.com/microsoft/onnxruntime/pull/31964)).
   - Negative CPU `Split` axes now produce an error instead of being accepted 
([#&#8203;31149](https://redirect.github.com/microsoft/onnxruntime/pull/31149)).
   
   #### Web & JavaScript
   
   - Added on-demand loading of Blob-backed external data in JSPI builds 
([#&#8203;29477](https://redirect.github.com/microsoft/onnxruntime/pull/29477)).
   - Fixed JSEP pooling output shape for `ceil_mode`, allowed DFT to ignore 
excess input data, and fixed a webpack/Terser release-build crash 
([#&#8203;29627](https://redirect.github.com/microsoft/onnxruntime/pull/29627), 
[#&#8203;29680](https://redirect.github.com/microsoft/onnxruntime/pull/29680), 
[#&#8203;31652](https://redirect.github.com/microsoft/onnxruntime/pull/31652)).
   
   #### Build, Packaging & CI
   
   - CUDA package architecture selections are now aligned across plugin EP, 
Python, C API, TensorRT, and Node.js pipelines. Windows arm64 is only available 
in CUDA plugin EP 
([#&#8203;31992](https://redirect.github.com/microsoft/onnxruntime/pull/31992)):
   
     | OS            | CUDA | CUDA architectures (all in `-real` form) |
     | ------------- | ---- | ---------------------------------------- |
     | Linux x64     | 12.8 | 60;70;75;80;86;89;90a;120a               |
     | Linux x64     | 13.x | 75;80;86;89;90a;120a                     |
     | Linux aarch64 | 13.x | 89;90a;120a;121a                         |
     | Windows x64   | 12.8 | 61;75;86;89;120a                         |
     | Windows x64   | 13.x | 75;80;86;89;120a                         |
     | Windows arm64 | 13.x | 120a;121a                                |
   - Reduced CUDA compilation time and memory usage by splitting generated SM80 
MoE, fpA\_intB, and MatMulNBits translation units and adding two-level 
workspace estimation 
([#&#8203;29614](https://redirect.github.com/microsoft/onnxruntime/pull/29614), 
[#&#8203;29699](https://redirect.github.com/microsoft/onnxruntime/pull/29699), 
[#&#8203;29811](https://redirect.github.com/microsoft/onnxruntime/pull/29811), 
[#&#8203;31834](https://redirect.github.com/microsoft/onnxruntime/pull/31834), 
[#&#8203;31837](https://redirect.github.com/microsoft/onnxruntime/pull/31837)).
   - Fixed CUDA 13 plugin and packaging builds on Windows, Linux, and Windows 
ARM64, including MSVC/TMA compatibility and CI memory limits 
([#&#8203;31608](https://redirect.github.com/microsoft/onnxruntime/pull/31608), 
[#&#8203;31609](https://redirect.github.com/microsoft/onnxruntime/pull/31609), 
[#&#8203;31615](https://redirect.github.com/microsoft/onnxruntime/pull/31615), 
[#&#8203;31616](https://redirect.github.com/microsoft/onnxruntime/pull/31616), 
[#&#8203;31617](https://redirect.github.com/microsoft/onnxruntime/pull/31617), 
[#&#8203;31622](https://redirect.github.com/microsoft/onnxruntime/pull/31622), 
[#&#8203;31729](https://redirect.github.com/microsoft/onnxruntime/pull/31729), 
[#&#8203;31748](https://redirect.github.com/microsoft/onnxruntime/pull/31748)).
   - Fixed MLAS AVX2 builds on toolchains without AVX-VNNI assembler support, 
GCC 15 `-Werror` builds, and an MSVC C1001 issue in the W2 AVX-512-VNNI 
dispatch path 
([#&#8203;28767](https://redirect.github.com/microsoft/onnxruntime/pull/28767), 
[#&#8203;29679](https://redirect.github.com/microsoft/onnxruntime/pull/29679), 
[#&#8203;29885](https://redirect.github.com/microsoft/onnxruntime/pull/29885)).
   - Improved Windows compatibility by skipping DXGI discovery when Win32k 
system calls are unavailable and delay-loading `shell32` 
([#&#8203;29755](https://redirect.github.com/microsoft/onnxruntime/pull/29755), 
[#&#8203;30889](https://redirect.github.com/microsoft/onnxruntime/pull/30889)).
   - Fixed Dawn parallel-build races, and GPU discovery in build/test 
environments 
([#&#8203;29858](https://redirect.github.com/microsoft/onnxruntime/pull/29858), 
[#&#8203;29866](https://redirect.github.com/microsoft/onnxruntime/pull/29866)).
   
   #### Contributors
   
   Thanks to our 63 contributors for this release!
   
   [@&#8203;adrastogi](https://redirect.github.com/adrastogi), 
[@&#8203;ahsan-ca](https://redirect.github.com/ahsan-ca), 
[@&#8203;AngelGalindo7](https://redirect.github.com/AngelGalindo7), 
[@&#8203;ankitm3k](https://redirect.github.com/ankitm3k), 
[@&#8203;apsonawane](https://redirect.github.com/apsonawane), 
[@&#8203;blazingphoenix7](https://redirect.github.com/blazingphoenix7), 
[@&#8203;bmehta001](https://redirect.github.com/bmehta001), 
[@&#8203;chilo-ms](https://redirect.github.com/chilo-ms), 
[@&#8203;claude](https://redirect.github.com/claude), 
[@&#8203;daijh](https://redirect.github.com/daijh), 
[@&#8203;ducviet00](https://redirect.github.com/ducviet00), 
[@&#8203;edgchen1](https://redirect.github.com/edgchen1), 
[@&#8203;elwhyjay](https://redirect.github.com/elwhyjay), 
[@&#8203;eserscor](https://redirect.github.com/eserscor), 
[@&#8203;GopalakrishnanN](https://redirect.github.com/GopalakrishnanN), 
[@&#8203;guptaishaan](https://redirect.github.com/guptaishaan), 
[@&#8203;hariharans29](
 https://redirect.github.com/hariharans29), 
[@&#8203;Honry](https://redirect.github.com/Honry), 
[@&#8203;huningxin](https://redirect.github.com/huningxin), 
[@&#8203;jchen10](https://redirect.github.com/jchen10), 
[@&#8203;jiafatom](https://redirect.github.com/jiafatom), 
[@&#8203;jiangzhuo](https://redirect.github.com/jiangzhuo), 
[@&#8203;Jiawei-Shao](https://redirect.github.com/Jiawei-Shao), 
[@&#8203;JonathanC-ARM](https://redirect.github.com/JonathanC-ARM), 
[@&#8203;justinchuby](https://redirect.github.com/justinchuby), 
[@&#8203;kjg0724](https://redirect.github.com/kjg0724), 
[@&#8203;kunal-vaishnavi](https://redirect.github.com/kunal-vaishnavi), 
[@&#8203;kylo5aby](https://redirect.github.com/kylo5aby), 
[@&#8203;Laan33](https://redirect.github.com/Laan33), 
[@&#8203;martin-klacer-arm](https://redirect.github.com/martin-klacer-arm), 
[@&#8203;mastryukov1990](https://redirect.github.com/mastryukov1990), 
[@&#8203;mcollinswisc](https://redirect.github.com/mcollinswisc), 
[@&#8203;miaobin](ht
 tps://redirect.github.com/miaobin), 
[@&#8203;mingmingtasd](https://redirect.github.com/mingmingtasd), 
[@&#8203;mirounga](https://redirect.github.com/mirounga), 
[@&#8203;mustjab](https://redirect.github.com/mustjab), 
[@&#8203;n1harika](https://redirect.github.com/n1harika), 
[@&#8203;namgyu-youn](https://redirect.github.com/namgyu-youn), 
[@&#8203;neilmsft](https://redirect.github.com/neilmsft), 
[@&#8203;nenad1002](https://redirect.github.com/nenad1002), 
[@&#8203;nicholascelestin](https://redirect.github.com/nicholascelestin), 
[@&#8203;OscarFree](https://redirect.github.com/OscarFree), 
[@&#8203;prathikr](https://redirect.github.com/prathikr), 
[@&#8203;qjia7](https://redirect.github.com/qjia7), 
[@&#8203;quic-muchhsu](https://redirect.github.com/quic-muchhsu), 
[@&#8203;Sammy-Dabbas](https://redirect.github.com/Sammy-Dabbas), 
[@&#8203;sanaa-hamel-microsoft](https://redirect.github.com/sanaa-hamel-microsoft),
 [@&#8203;shiyi9801](https://redirect.github.com/shiyi9801), 
[@&#8203;skottmckay](
 https://redirect.github.com/skottmckay), 
[@&#8203;tairenpiao](https://redirect.github.com/tairenpiao), 
[@&#8203;TedThemistokleous](https://redirect.github.com/TedThemistokleous), 
[@&#8203;the0cp](https://redirect.github.com/the0cp), 
[@&#8203;tianleiwu](https://redirect.github.com/tianleiwu), 
[@&#8203;titaiwangms](https://redirect.github.com/titaiwangms), 
[@&#8203;velonica0](https://redirect.github.com/velonica0), 
[@&#8203;wangw-1991](https://redirect.github.com/wangw-1991), 
[@&#8203;wuisabel-gif](https://redirect.github.com/wuisabel-gif), 
[@&#8203;xadupre](https://redirect.github.com/xadupre), 
[@&#8203;xhcao](https://redirect.github.com/xhcao), 
[@&#8203;xiaofeihan1](https://redirect.github.com/xiaofeihan1), 
[@&#8203;xiaoyu-work](https://redirect.github.com/xiaoyu-work), 
[@&#8203;yen-shi](https://redirect.github.com/yen-shi), 
[@&#8203;zlma7001](https://redirect.github.com/zlma7001)
   
   Full Changelog: 
[v1.28.0...v1.29.0](https://redirect.github.com/microsoft/onnxruntime/compare/v1.28.0...v1.29.0)
   
   ### 
[`v1.28.0`](https://redirect.github.com/microsoft/onnxruntime/releases/tag/v1.28.0):
 ONNX Runtime v1.28.0
   
   #### Announcements & Breaking Changes
   
   - Upgraded to **ONNX 1.22.0** and protobuf 6.33.5 
([#&#8203;28754](https://redirect.github.com/microsoft/onnxruntime/pull/28754), 
[#&#8203;29606](https://redirect.github.com/microsoft/onnxruntime/pull/29606), 
[#&#8203;28967](https://redirect.github.com/microsoft/onnxruntime/pull/28967)). 
Graph optimizer opset version checks were updated accordingly 
([#&#8203;28966](https://redirect.github.com/microsoft/onnxruntime/pull/28966)).
   - **cuDNN and cuFFT are now optional at runtime** for the CUDA EP, and 
`nvrtc` is no longer linked, which significantly reduces the required CUDA 
redistributable footprint 
([#&#8203;29252](https://redirect.github.com/microsoft/onnxruntime/pull/29252), 
[#&#8203;29808](https://redirect.github.com/microsoft/onnxruntime/pull/29808), 
[#&#8203;29705](https://redirect.github.com/microsoft/onnxruntime/pull/29705), 
[#&#8203;29620](https://redirect.github.com/microsoft/onnxruntime/pull/29620)).
   - An **experimental C/C++ API surface** was introduced. `OrtModelPackageApi` 
now lives in the experimental C API and may change in future releases 
([#&#8203;28746](https://redirect.github.com/microsoft/onnxruntime/pull/28746), 
[#&#8203;29142](https://redirect.github.com/microsoft/onnxruntime/pull/29142), 
[#&#8203;28990](https://redirect.github.com/microsoft/onnxruntime/pull/28990)).
   - **Deprecated / removed:**
     - SkipLayerNorm strict mode is deprecated 
([#&#8203;29388](https://redirect.github.com/microsoft/onnxruntime/pull/29388)).
     - The TensorRT fused causal attention kernels were removed from the CUDA 
EP 
([#&#8203;29143](https://redirect.github.com/microsoft/onnxruntime/pull/29143)).
     - The dynamic WGSL generator (duktape/Node) path was removed in favor of 
the Python `wgsl-gen` implementation 
([#&#8203;29141](https://redirect.github.com/microsoft/onnxruntime/pull/29141), 
[#&#8203;28355](https://redirect.github.com/microsoft/onnxruntime/pull/28355)).
     - `CUDA_QUANT_PREPROCESS` is off by default 
([#&#8203;29687](https://redirect.github.com/microsoft/onnxruntime/pull/29687)).
   - NPM packages are now published from the CUDA 13 pipeline 
([#&#8203;28773](https://redirect.github.com/microsoft/onnxruntime/pull/28773)).
   - The CUDA 12.8 package architecture list was refreshed for this release 
([#&#8203;29711](https://redirect.github.com/microsoft/onnxruntime/pull/29711)).
   
   #### Security Fixes
   
   ##### Memory safety & input validation
   
   - Hardened the ORT FlatBuffer model loader against malformed buffers, and 
removed now-redundant table offset validation 
([#&#8203;28186](https://redirect.github.com/microsoft/onnxruntime/pull/28186), 
[#&#8203;29068](https://redirect.github.com/microsoft/onnxruntime/pull/29068))
   - Fixed type confusion in raw-pointer `bind_input` causing an out-of-bounds 
write 
([#&#8203;28839](https://redirect.github.com/microsoft/onnxruntime/pull/28839))
   - Fixed out-of-bounds pointer in `TensorAt` for sub-byte packed types 
([#&#8203;28973](https://redirect.github.com/microsoft/onnxruntime/pull/28973))
   - Fixed arbitrary memory read, out-of-bounds dereference, and other OOB 
accesses in kernels 
([#&#8203;28991](https://redirect.github.com/microsoft/onnxruntime/pull/28991), 
[#&#8203;29011](https://redirect.github.com/microsoft/onnxruntime/pull/29011), 
[#&#8203;29012](https://redirect.github.com/microsoft/onnxruntime/pull/29012), 
[#&#8203;29014](https://redirect.github.com/microsoft/onnxruntime/pull/29014))
   - Validated `Col2Im` inputs to prevent heap over-read 
([#&#8203;28706](https://redirect.github.com/microsoft/onnxruntime/pull/28706))
   - Hardened `CropAndResize` against malformed `crop_size` tensors 
([#&#8203;28766](https://redirect.github.com/microsoft/onnxruntime/pull/28766))
   - Validated `BeamSearch` `vocab_size` against logits width 
([#&#8203;28774](https://redirect.github.com/microsoft/onnxruntime/pull/28774))
   - Fixed bounds in `WhisperDecoderSubgraph::CreateInitialFeeds` 
([#&#8203;29239](https://redirect.github.com/microsoft/onnxruntime/pull/29239))
   - Validated `SparseAttention` CSR indices/key lengths and rejected 
zero-dimension `block_row_indices` 
([#&#8203;29015](https://redirect.github.com/microsoft/onnxruntime/pull/29015), 
[#&#8203;29242](https://redirect.github.com/microsoft/onnxruntime/pull/29242))
   - Clamped derived sequence lengths and KV-cache index in CUDA 
GroupQueryAttention, and fixed a CPU GQA out-of-bounds read in the past-KV 
buffer 
([#&#8203;29240](https://redirect.github.com/microsoft/onnxruntime/pull/29240), 
[#&#8203;29447](https://redirect.github.com/microsoft/onnxruntime/pull/29447))
   - Clamped 1D attention `mask_index` to valid bounds 
([#&#8203;29449](https://redirect.github.com/microsoft/onnxruntime/pull/29449))
   - Validated `MaxpoolWithMask` kernel rank against input spatial rank 
([#&#8203;29253](https://redirect.github.com/microsoft/onnxruntime/pull/29253))
   - Rejected CUDA BERT `EmbedLayerNorm`/`SkipLayerNorm` shapes exceeding 
32-bit output indexing 
([#&#8203;29264](https://redirect.github.com/microsoft/onnxruntime/pull/29264))
   - Fixed the optional-output guard in `DecoderAttention`/`MultiHeadAttention` 
shape inference and negative-axis handling in `ExpandDims` shape inference 
([#&#8203;29268](https://redirect.github.com/microsoft/onnxruntime/pull/29268), 
[#&#8203;29448](https://redirect.github.com/microsoft/onnxruntime/pull/29448))
   - Fixed `TreeEnsemble` target id validation and added input validation to 
`LinearClassifier` 
([#&#8203;29293](https://redirect.github.com/microsoft/onnxruntime/pull/29293), 
[#&#8203;29060](https://redirect.github.com/microsoft/onnxruntime/pull/29060))
   - Fixed `DynamicQuantizeLSTM` zero-point/scale validation typos 
([#&#8203;29462](https://redirect.github.com/microsoft/onnxruntime/pull/29462))
   - Handled non-trivially-copyable types in `Loop`/`Scan` output concatenation 
([#&#8203;29397](https://redirect.github.com/microsoft/onnxruntime/pull/29397))
   - Normalized bool tensor `raw_data` to `{0, 1}` on unpack 
([#&#8203;29238](https://redirect.github.com/microsoft/onnxruntime/pull/29238))
   - Addressed hardening gaps in `Resize`, `PadFusion`, and LoRA handling 
([#&#8203;28779](https://redirect.github.com/microsoft/onnxruntime/pull/28779), 
[#&#8203;28780](https://redirect.github.com/microsoft/onnxruntime/pull/28780), 
[#&#8203;28801](https://redirect.github.com/microsoft/onnxruntime/pull/28801))
   - Fixed unbounded lifetime on `WithOutputTensor` in the Rust bindings 
([#&#8203;29251](https://redirect.github.com/microsoft/onnxruntime/pull/29251))
   
   ##### Integer overflow & allocation size
   
   - Guarded `MlasConvPrepare` working-buffer products and `ConvTranspose` pad 
computation with SafeInt 
([#&#8203;29444](https://redirect.github.com/microsoft/onnxruntime/pull/29444), 
[#&#8203;29446](https://redirect.github.com/microsoft/onnxruntime/pull/29446))
   - Fixed signed-int overflow in `SamplingState::Init` that could cause a heap 
buffer overflow 
([#&#8203;29443](https://redirect.github.com/microsoft/onnxruntime/pull/29443))
   - Hardened QMoE against integer overflow and partial K tiles 
([#&#8203;29067](https://redirect.github.com/microsoft/onnxruntime/pull/29067))
   - Validated `B`/scales/zero-points shape in `MatMulNBits::PrePack` 
([#&#8203;29445](https://redirect.github.com/microsoft/onnxruntime/pull/29445))
   - Pre-checked `ConstantOfShape` output size against the input initializer 
before constant folding 
([#&#8203;28751](https://redirect.github.com/microsoft/onnxruntime/pull/28751))
   - Fixed integer overflow in RKNPU implicit bias allocation 
([#&#8203;29249](https://redirect.github.com/microsoft/onnxruntime/pull/29249))
   - Fixed WebGPU out-of-bounds reads in `Pad` (int64/int32 truncation), 
`Slice`, and `GatherBlockQuantized` 
([#&#8203;28721](https://redirect.github.com/microsoft/onnxruntime/pull/28721), 
[#&#8203;28704](https://redirect.github.com/microsoft/onnxruntime/pull/28704), 
[#&#8203;28718](https://redirect.github.com/microsoft/onnxruntime/pull/28718))
   
   ##### Supply chain & tooling
   
   - Updated protobuf to mitigate CVE-2026-0994 and bumped ONNX/protobuf to fix 
additional CVEs 
([#&#8203;28967](https://redirect.github.com/microsoft/onnxruntime/pull/28967), 
[#&#8203;29606](https://redirect.github.com/microsoft/onnxruntime/pull/29606))
   - Avoided shell injection in the training helper and switched Triton compile 
helpers to `subprocess` 
([#&#8203;28776](https://redirect.github.com/microsoft/onnxruntime/pull/28776), 
[#&#8203;28775](https://redirect.github.com/microsoft/onnxruntime/pull/28775))
   - Validated archive extraction paths in the transformers tooling 
([#&#8203;28777](https://redirect.github.com/microsoft/onnxruntime/pull/28777))
   - Validated and inlined external data in node tensor attributes during 
session initialization 
([#&#8203;29250](https://redirect.github.com/microsoft/onnxruntime/pull/29250))
   - Enabled Spectre-mitigated MSVC libraries for BinSkim builds 
([#&#8203;29624](https://redirect.github.com/microsoft/onnxruntime/pull/29624))
   - Bumped npm dependencies: `shell-quote`, `esbuild`, `tmp`, `ws`, 
`protobufjs`, `js-yaml`, `tar`, `markdown-it`, `@babel/core` 
([#&#8203;29022](https://redirect.github.com/microsoft/onnxruntime/pull/29022), 
[#&#8203;29044](https://redirect.github.com/microsoft/onnxruntime/pull/29044), 
[#&#8203;29055](https://redirect.github.com/microsoft/onnxruntime/pull/29055), 
[#&#8203;29057](https://redirect.github.com/microsoft/onnxruntime/pull/29057), 
[#&#8203;29061](https://redirect.github.com/microsoft/onnxruntime/pull/29061), 
[#&#8203;29062](https://redirect.github.com/microsoft/onnxruntime/pull/29062), 
[#&#8203;29063](https://redirect.github.com/microsoft/onnxruntime/pull/29063), 
[#&#8203;29079](https://redirect.github.com/microsoft/onnxruntime/pull/29079), 
[#&#8203;29090](https://redirect.github.com/microsoft/onnxruntime/pull/29090), 
[#&#8203;29156](https://redirect.github.com/microsoft/onnxruntime/pull/29156))
   
   #### New Features
   
   ##### Execution Provider ABI & Plugin EPs
   
   - Model Package support Phase 2, plus authoring tools, schema versioning, 
and folding `external_data` into session options 
([#&#8203;28271](https://redirect.github.com/microsoft/onnxruntime/pull/28271), 
[#&#8203;28989](https://redirect.github.com/microsoft/onnxruntime/pull/28989), 
[#&#8203;29501](https://redirect.github.com/microsoft/onnxruntime/pull/29501))
   - Added an API to select the best compiled-model compatibility info from 
candidate strings 
([#&#8203;28387](https://redirect.github.com/microsoft/onnxruntime/pull/28387))
   - Added crypto support: applications can supply I/O callbacks to an EP, with 
callback and fallback helpers 
([#&#8203;28624](https://redirect.github.com/microsoft/onnxruntime/pull/28624))
   - Implemented name-based partitioning with accompanying documentation 
([#&#8203;28903](https://redirect.github.com/microsoft/onnxruntime/pull/28903))
   - Added Linux NPU discovery through sysfs accel devices 
([#&#8203;28703](https://redirect.github.com/microsoft/onnxruntime/pull/28703))
   - Relaxed `CompileModel` validation to accept zero-input `OrtModel` graphs 
([#&#8203;28771](https://redirect.github.com/microsoft/onnxruntime/pull/28771))
   - CUDA plugin EP: user compute stream with CUDA graph, kernel sync stream 
exposed for scratch allocation, and Windows ARM64 packages 
([#&#8203;29221](https://redirect.github.com/microsoft/onnxruntime/pull/29221), 
[#&#8203;29244](https://redirect.github.com/microsoft/onnxruntime/pull/29244), 
[#&#8203;28896](https://redirect.github.com/microsoft/onnxruntime/pull/28896), 
[#&#8203;28789](https://redirect.github.com/microsoft/onnxruntime/pull/28789))
   - WebGPU plugin EP version bumped to 0.3.0 
([#&#8203;29056](https://redirect.github.com/microsoft/onnxruntime/pull/29056))
   
   ##### Core APIs & Runtime
   
   - Added `OrtErrorCode` documentation, single-sourced the values so 
`StatusCode` stays in sync, and added `OrtErrorCode::ORT_DEVICE_RESET` 
([#&#8203;29018](https://redirect.github.com/microsoft/onnxruntime/pull/29018), 
[#&#8203;29065](https://redirect.github.com/microsoft/onnxruntime/pull/29065), 
[#&#8203;29748](https://redirect.github.com/microsoft/onnxruntime/pull/29748))
   - Added memory statistics to profiling output 
([#&#8203;29058](https://redirect.github.com/microsoft/onnxruntime/pull/29058))
   - Added EP version logging on inference failure, in the `EpDeviceUsage` 
event, and ORT version logging 
([#&#8203;28794](https://redirect.github.com/microsoft/onnxruntime/pull/28794))
   - User-supplied external initializers are now used in place when already on 
the planned device 
([#&#8203;29013](https://redirect.github.com/microsoft/onnxruntime/pull/29013))
   - `model_external_initializers_file_folder_path` is now honored for 
file-path model loads 
([#&#8203;29459](https://redirect.github.com/microsoft/onnxruntime/pull/29459))
   - Added a Python API for `HOST_ACCESSIBLE` `OrtValue` allocation 
([#&#8203;28038](https://redirect.github.com/microsoft/onnxruntime/pull/28038))
   
   ##### Quantization Tooling
   
   - Added `CudaQuantizer` to `onnxruntime.quantization` 
([#&#8203;29509](https://redirect.github.com/microsoft/onnxruntime/pull/29509))
   - Registered `Flatten` as a Direct8Bit op in the Python QDQ static quantizer 
([#&#8203;28340](https://redirect.github.com/microsoft/onnxruntime/pull/28340))
   - Skipped `MaxPool` during FP8 static quantization and fixed the FP8 
(`FLOAT8E4M3FN`) scale reference distribution 
([#&#8203;28488](https://redirect.github.com/microsoft/onnxruntime/pull/28488), 
[#&#8203;29350](https://redirect.github.com/microsoft/onnxruntime/pull/29350))
   - Added Float16/BFloat16/Float8 support in the `TensorArray` custom op 
([#&#8203;28335](https://redirect.github.com/microsoft/onnxruntime/pull/28335))
   - Clarified CPU parameter recommendations in the quantization docs 
([#&#8203;28415](https://redirect.github.com/microsoft/onnxruntime/pull/28415))
   
   #### Execution Provider Updates
   
   ##### NVIDIA CUDA EP
   
   **Attention & LLM decode**
   
   - Enabled XQA by default for FP16/BF16 GroupQueryAttention, and extended XQA 
decode with attention sink, sliding window, and QK-Norm support 
([#&#8203;29046](https://redirect.github.com/microsoft/onnxruntime/pull/29046), 
[#&#8203;29162](https://redirect.github.com/microsoft/onnxruntime/pull/29162), 
[#&#8203;29177](https://redirect.github.com/microsoft/onnxruntime/pull/29177), 
[#&#8203;29186](https://redirect.github.com/microsoft/onnxruntime/pull/29186))
   - Upgraded `cudnn_frontend` to 1.24 and enabled cuDNN SDPA for MHA/GQA 
([#&#8203;28849](https://redirect.github.com/microsoft/onnxruntime/pull/28849))
   - Added decode-optimized LinearAttention (GatedDeltaNet) kernels 
([#&#8203;28985](https://redirect.github.com/microsoft/onnxruntime/pull/28985))
   - Optimized FlashDecode split planning for local-window GQA and fixed 
Flash/Lean attention split heuristics 
([#&#8203;29161](https://redirect.github.com/microsoft/onnxruntime/pull/29161), 
[#&#8203;29554](https://redirect.github.com/microsoft/onnxruntime/pull/29554))
   - Updated the GroupQueryAttention contrib op documentation 
([#&#8203;29173](https://redirect.github.com/microsoft/onnxruntime/pull/29173))
   
   **MoE & quantized GEMM**
   
   - Prepacked int4/int8 QMoE expert weights in the `PrePack` hook, symmetric 
with `MatMulNBits`, and fix
   
   > ✂ **Note**
   > 
   > PR body was truncated to here.
   
   
   </details>
   
   ---
   
   ### Configuration
   
   📅 **Schedule**: (UTC)
   
   - Branch creation
     - "before 9am on the first day of the month"
   - Automerge
     - At any time (no schedule defined)
   
   🚦 **Automerge**: Disabled by config. Please merge this manually once you are 
satisfied.
   
   ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry 
checkbox.
   
   👻 **Immortal**: This PR will be recreated if closed unmerged. Get [config 
help](https://redirect.github.com/renovatebot/renovate/discussions) if that's 
undesired.
   
   ---
   
    - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this 
box
   
   ---
   
   This PR has been generated by [Renovate 
Bot](https://redirect.github.com/solrbot/renovate-github-action)
   
<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4xNzAuMTgiLCJ1cGRhdGVkSW5WZXIiOiI0My4xNzAuMTgiLCJ0YXJnZXRCcmFuY2giOiJtYWluIiwibGFiZWxzIjpbImV4ZW1wdC1zdGFsZSJdfQ==-->
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to