The GitHub Actions job "CI" on tvm-ffi.git/main has succeeded.
Run started by GitHub user cyx-6 (triggered by cyx-6).

Head commit for run:
230ae6c86ababb170baac6c485214e0a5af30702 / Piotr Mazurek <[email protected]>
Include torch build/ABI in torch C-DLPack addon cache key (#644)

## Problem

The prebuilt torch C-DLPack addon is cached under a filename derived
only from torch **major.minor** + a coarse device string:

```python
major, minor = torch.__version__.split(".")[:2]
device = _torch_extension_device(torch)          # "cuda" / "rocm" / "cpu"
libname = f"libtorch_c_dlpack_addon_torch{major}{minor}-{device}{suffix}"
lib_path = cache_dir / libname                   # cache_dir defaults to 
~/.cache/tvm-ffi
if not lib_path.exists():
    ...build...                                  # otherwise reuse whatever is 
there
```

This key omits:
- the torch **patch** version (`2.9.0` vs `2.9.1`),
- the build local-version tag carried in `torch.__version__` (`+cu121`
vs `+cu124`, `+cpu`, …),
- the C++ ABI flag (`torch._C._GLIBCXX_USE_CXX11_ABI`).

Since the addon is a compiled extension linking libtorch's C++ ABI, two
torch installs that share `major.minor` + device but differ in patch /
CUDA toolkit / ABI resolve to the **same** cached `.so`. The addon built
against the first torch is then silently reused by the second — an ABI
mismatch in the DLPack bridge that surfaces as crashes, memory faults,
or **silently wrong tensor data**, not a clean error.

This is easy to hit whenever `~/.cache/tvm-ffi` is shared across
environments — a shared/NFS home, or container images that mount the
host home and see the same cache under different torch builds.

## Reproduce

```bash
# env A
pip install torch==2.9.0+cu121 --index-url 
https://download.pytorch.org/whl/cu121
python -c "import tvm_ffi"     # builds 
~/.cache/tvm-ffi/libtorch_c_dlpack_addon_torch29-cuda.so

# env B: same major.minor, different build/ABI, same cache dir
pip install torch==2.9.1+cu124 --index-url 
https://download.pytorch.org/whl/cu124
python -c "import tvm_ffi"     # REUSES the cu121-built torch29-cuda.so -> ABI 
mismatch
```
The same applies to any two builds that share
`torch{major}{minor}-{device}`, including CPU and ROCm builds.

## Fix

Fold the full torch build identity (`torch.__version__`, which already
carries patch + `+cuXXX`/`+rocmX.Y`/`+cpu`) and the C++ ABI flag into
the cached addon name via a short hash. Incompatible builds now get
distinct cache entries, while same-build reuse is unchanged. The build
subprocess receives `libname` from this call site, so it stays
consistent automatically.

Signed-off-by: Piotr Mazurek <[email protected]>
Co-authored-by: Piotr Mazurek <[email protected]>

Report URL: https://github.com/apache/tvm-ffi/actions/runs/28384085888

With regards,
GitHub Actions via GitBox


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to