================
@@ -0,0 +1,48 @@
+/// Check that the x86 intrinsic headers parse and compile for a SPIR-V device
+/// when the auxiliary host target is x86-64 MSVC, which is what the sse/sse2
+/// device features derived from that host are for.
+
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple 
x86_64-pc-windows-msvc \
+// RUN:   -fsycl-is-device -ffreestanding -emit-llvm -o - %s | FileCheck %s
+
+/// With sse/sse2 disabled, the always_inline intrinsics cannot be inlined into
+/// device code; this is the failure the derived features prevent.
+/// Codegen stops at the first failing function, so only add() is diagnosed.
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple 
x86_64-pc-windows-msvc \
+// RUN:   -fsycl-is-device -ffreestanding -target-feature -sse -target-feature 
-sse2 \
+// RUN:   -emit-llvm -verify=no-sse -o - %s
+
+/// Intrinsics that lower to x86 builtins are rejected for the device.
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple 
x86_64-pc-windows-msvc \
+// RUN:   -fsycl-is-device -ffreestanding -DX86_BUILTIN -emit-llvm 
-verify=x86-builtin \
+// RUN:   -o - %s
+
+#include <immintrin.h>
+
+// CHECK-LABEL: define {{.*}}spir_func noundef float @_Z3addff
+[[clang::sycl_external]] float add(float x, float y) {
+  // no-sse-error@+1 {{always_inline function '_mm_set1_ps' requires target 
feature 'sse', but would be inlined into function 'add' that is compiled 
without support for 'sse'}}
+  __m128 a = _mm_set1_ps(x);
+  // no-sse-error@+1 {{always_inline function '_mm_set1_ps' requires target 
feature 'sse', but would be inlined into function 'add' that is compiled 
without support for 'sse'}}
+  __m128 b = _mm_set1_ps(y);
+  // CHECK: fadd <4 x float>
----------------
tahonermann wrote:

Ah, yes, I missed that. `_mm_set1_ps()` is defined in 
`clang/lib/Headers/xmmintrin.h` as:
```c++
  36 #define __DEFAULT_FN_ATTRS                                                 
    \
  37   __attribute__((__always_inline__, __nodebug__, __target__("sse"),        
    \
  38                  __min_vector_width__(128)))
  39 #define __DEFAULT_FN_ATTRS_SSE2                                            
    \
  40   __attribute__((__always_inline__, __nodebug__, __target__("sse2"),       
    \
  41                  __min_vector_width__(128)))
  42
  43 #if defined(__cplusplus) && (__cplusplus >= 201103L)
  44 #define __DEFAULT_FN_ATTRS_CONSTEXPR __DEFAULT_FN_ATTRS constexpr
  45 #define __DEFAULT_FN_ATTRS_SSE2_CONSTEXPR __DEFAULT_FN_ATTRS_SSE2 constexpr
  46 #else
  47 #define __DEFAULT_FN_ATTRS_CONSTEXPR __DEFAULT_FN_ATTRS
  48 #define __DEFAULT_FN_ATTRS_SSE2_CONSTEXPR __DEFAULT_FN_ATTRS_SSE2
  49 #endif
  ..
1914 static __inline__ __m128 __DEFAULT_FN_ATTRS_CONSTEXPR
1915 _mm_set1_ps(float __w) {
1916   return __extension__ (__m128){ __w, __w, __w, __w };
1917 }
```

The body of the function is just a C99 compound literal, so the only basis we 
have for rejection is the [`__target__` 
attribute](https://clang.llvm.org/docs/AttributeReference.html#target) attached 
to the function. Diagnosing the `arch=`, `cpu=`, `tune=`, and 
`branch-protection=` cases should be straight forward. Diagnosing the subtarget 
feature cases looks like it will require differentiating which features are 
enabled for host compatibility vs which ones are supported for the actual 
target. I don't have a good sense of how difficult that would be to do.

https://github.com/llvm/llvm-project/pull/227665
_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits

Reply via email to