https://github.com/Thyre created 
https://github.com/llvm/llvm-project/pull/210038

For the OpenMP Tools Interface, a number of callbacks should get dispatched for 
the `sections` construct, and the `section` directive. For each 
_sections-begin_ event, a work callback with `ompt_scope_begin` as the endpoint 
and `ompt_work_sections` as the work_type dispatched. For each _section-begin_, 
a dispatch callback with the kind set to `ompt_dispatch_section` is dispatched. 
And finally, on _sections-end_, another work callback with `ompt_scope_end` is 
emitted.

In Clang and Flang, the `sections` construct and `section` directive is 
rewritten during code generation. A user code such as:

```
{
  #pragma omp section
  {}

  #pragma omp section
  {}
}
```

is transformed to:

```
[...]

omp.inner.for.cond:
  %6 = load i32, ptr %.omp.sections.iv., align 4
  %7 = load i32, ptr %.omp.sections.ub., align 4
  %cmp = icmp sle i32 %6, %7
  br i1 %cmp, label %omp.inner.for.body, label %omp.inner.for.end

omp.inner.for.body:
  %8 = load i32, ptr %.omp.sections.iv., align 4
  switch i32 %8, label %.omp.sections.exit [
    i32 0, label %.omp.sections.case
    i32 1, label %.omp.sections.case1
  ]

.omp.sections.case:
  br label %.omp.sections.exit

.omp.sections.case1:
  br label %.omp.sections.exit

.omp.sections.exit:
  br label %omp.inner.for.inc

[...]
```

Clang and Flang both use a static loop, splitting work across threads in the 
team, with a switch-case to select the correct `section` directive. On the call 
to `__kmpc_for_static_init`, Clang passes `OMP_IDENT_WORK_SECTIONS` as part of 
the ident flags. As such, the OpenMP runtime identifies the transformed 
construct as a `sections` construct with one or more `section` directives.

In #207746, it was found that this handling can lead to an incorrect event 
sequence emitted to tools. Since `__kmpc_for_static_init` is only called _once_ 
at the beginning of the loop for each thread in the team, and each thread 
dispatches one work and dispatch callback to an attached tool, regardless of 
the actual number of section directives handles, a tool might therefore 
incorrectly assume that the number of `section` directives matches the number 
of threads.

To improve the dispatched events to an attached tool, use `OMP_IDENT_WORK_LOOP` 
for the `sections` construct and `section` directives as well. With this, the 
behavior of Clang is consistent with the behavior of Flang. A tool is able to 
identify the number of loop iterations done per thread, hence is able to filter 
threads not executing any `section` directive.

Closes #207746

From 223da05d52980521d0053828a2235349746b80d7 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Jan=20Andr=C3=A9=20Reuter?= <[email protected]>
Date: Wed, 15 Jul 2026 16:04:44 +0200
Subject: [PATCH] [Clang][OpenMP] Use loop as ident flag for sections
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit

For the OpenMP Tools Interface, a number of callbacks should get dispatched
for the `sections` construct, and the `section` directive.
For each _sections-begin_ event, a work callback with `ompt_scope_begin` as
the endpoint and `ompt_work_sections` as the work_type dispatched. For each
_section-begin_, a dispatch callback with the kind set to
`ompt_dispatch_section` is dispatched. And finally, on _sections-end_,
another work callback with `ompt_scope_end` is emitted.

In Clang and Flang, the `sections` construct and `section` directive
is rewritten during code generation. A user code such as:

```
{
  #pragma omp section
  {}

  #pragma omp section
  {}
}
```

is transformed to:

```
[...]

omp.inner.for.cond:
  %6 = load i32, ptr %.omp.sections.iv., align 4
  %7 = load i32, ptr %.omp.sections.ub., align 4
  %cmp = icmp sle i32 %6, %7
  br i1 %cmp, label %omp.inner.for.body, label %omp.inner.for.end

omp.inner.for.body:
  %8 = load i32, ptr %.omp.sections.iv., align 4
  switch i32 %8, label %.omp.sections.exit [
    i32 0, label %.omp.sections.case
    i32 1, label %.omp.sections.case1
  ]

.omp.sections.case:
  br label %.omp.sections.exit

.omp.sections.case1:
  br label %.omp.sections.exit

.omp.sections.exit:
  br label %omp.inner.for.inc

[...]
```

Clang and Flang both use a static loop, splitting work across threads in the
team, with a switch-case to select the correct `section` directive.
On the call to `__kmpc_for_static_init`, Clang passes `OMP_IDENT_WORK_SECTIONS`
as part of the ident flags. As such, the OpenMP runtime identifies the
transformed construct as a `sections` construct with one or more `section`
directives.

In #207746, it was found that this handling can lead to an incorrect event
sequence emitted to tools. Since `__kmpc_for_static_init` is only called _once_
at the beginning of the loop for each thread in the team, and each thread
dispatches one work and dispatch callback to an attached tool, regardless of
the actual number of section directives handles, a tool might therefore
incorrectly assume that the number of `section` directives matches the number
of threads.

To improve the dispatched events to an attached tool, use
`OMP_IDENT_WORK_LOOP` for the `sections` construct and `section` directives
as well. With this, the behavior of Clang is consistent with the behavior of
Flang. A tool is able to identify the number of loop iterations done per
thread, hence is able to filter threads not executing any `section` directive.

Closes #207746

Signed-off-by: Jan André Reuter <[email protected]>
---
 clang/lib/CodeGen/CGOpenMPRuntime.cpp              | 10 +++-------
 clang/test/OpenMP/sections_codegen.cpp             |  2 +-
 openmp/runtime/test/ompt/worksharing/sections.c    |  8 ++++----
 .../test/ompt/worksharing/sections_dispatch.c      | 14 ++++++--------
 4 files changed, 14 insertions(+), 20 deletions(-)

diff --git a/clang/lib/CodeGen/CGOpenMPRuntime.cpp 
b/clang/lib/CodeGen/CGOpenMPRuntime.cpp
index eb2f92cdbf972..f68fbb8b50824 100644
--- a/clang/lib/CodeGen/CGOpenMPRuntime.cpp
+++ b/clang/lib/CodeGen/CGOpenMPRuntime.cpp
@@ -2752,10 +2752,8 @@ void CGOpenMPRuntime::emitForStaticInit(CodeGenFunction 
&CGF,
                                Values.Ordered);
   assert((isOpenMPWorksharingDirective(DKind) || (DKind == OMPD_loop)) &&
          "Expected loop-based or sections-based directive.");
-  llvm::Value *UpdatedLocation = emitUpdateLocation(CGF, Loc,
-                                             isOpenMPLoopDirective(DKind)
-                                                 ? OMP_IDENT_WORK_LOOP
-                                                 : OMP_IDENT_WORK_SECTIONS);
+  llvm::Value *UpdatedLocation =
+      emitUpdateLocation(CGF, Loc, OMP_IDENT_WORK_LOOP);
   llvm::Value *ThreadId = getThreadID(CGF, Loc);
   llvm::FunctionCallee StaticInitFunction =
       OMPBuilder.createForStaticInitFunction(Values.IVSize, Values.IVSigned,
@@ -2799,9 +2797,7 @@ void CGOpenMPRuntime::emitForStaticFinish(CodeGenFunction 
&CGF,
                          isOpenMPDistributeDirective(DKind) ||
                                  (DKind == OMPD_target_teams_loop)
                              ? OMP_IDENT_WORK_DISTRIBUTE
-                         : isOpenMPLoopDirective(DKind)
-                             ? OMP_IDENT_WORK_LOOP
-                             : OMP_IDENT_WORK_SECTIONS),
+                             : OMP_IDENT_WORK_LOOP),
       getThreadID(CGF, Loc)};
   auto DL = ApplyDebugLocation::CreateDefaultArtificial(CGF, Loc);
   if (isOpenMPDistributeDirective(DKind) &&
diff --git a/clang/test/OpenMP/sections_codegen.cpp 
b/clang/test/OpenMP/sections_codegen.cpp
index 5a5e32751c677..957833e60daaa 100644
--- a/clang/test/OpenMP/sections_codegen.cpp
+++ b/clang/test/OpenMP/sections_codegen.cpp
@@ -10,7 +10,7 @@
 #ifndef HEADER
 #define HEADER
 // CHECK-DAG: [[IMPLICIT_BARRIER_SECTIONS_LOC:@.+]] = private unnamed_addr 
constant %{{.+}} { i32 0, i32 194, i32 0, i32 {{[0-9]+}}, ptr
-// CHECK-DAG: [[SECTIONS_LOC:@.+]] = private unnamed_addr constant %{{.+}} { 
i32 0, i32 1026, i32 0, i32 {{[0-9]+}}, ptr
+// CHECK-DAG: [[SECTIONS_LOC:@.+]] = private unnamed_addr constant %{{.+}} { 
i32 0, i32 514, i32 0, i32 {{[0-9]+}}, ptr
 // CHECK-LABEL: foo
 void foo() { extern void mayThrow(); mayThrow(); };
 // CHECK-LABEL: bar
diff --git a/openmp/runtime/test/ompt/worksharing/sections.c 
b/openmp/runtime/test/ompt/worksharing/sections.c
index 5cf927def6905..b319a0f1a9a3f 100644
--- a/openmp/runtime/test/ompt/worksharing/sections.c
+++ b/openmp/runtime/test/ompt/worksharing/sections.c
@@ -28,11 +28,11 @@ int main() {
 
   // CHECK: 0: NULL_POINTER=[[NULL:.*$]]
 
-  // CHECK: {{^}}[[MASTER_ID:[0-9]+]]: ompt_event_sections_begin: 
parallel_id=[[PARALLEL_ID:[0-f]+]], task_id=[[TASK_ID:[0-f]+]], 
codeptr_ra=[[SECT_BEGIN:(0x)?[0-f]+]], count=2
-  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_sections_end: 
parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}}, 
codeptr_ra=[[SECT_END:(0x)?[0-f]+]]
+  // CHECK: {{^}}[[MASTER_ID:[0-9]+]]: ompt_event_loop_static_begin: 
parallel_id=[[PARALLEL_ID:[0-f]+]], task_id=[[TASK_ID:[0-f]+]], 
codeptr_ra=[[SECT_BEGIN:(0x)?[0-f]+]], count=2
+  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_loop_static_end: 
parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}}, 
codeptr_ra=[[SECT_END:(0x)?[0-f]+]]
 
-  // CHECK: {{^}}[[THREAD_ID:[0-9]+]]: ompt_event_sections_begin: 
parallel_id=[[PARALLEL_ID]], task_id=[[TASK_ID:[0-f]+]], 
codeptr_ra=[[SECT_BEGIN]], count=2
-  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_sections_end: 
parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}}, codeptr_ra=[[SECT_END]]
+  // CHECK: {{^}}[[THREAD_ID:[0-9]+]]: ompt_event_loop_static_begin: 
parallel_id=[[PARALLEL_ID]], task_id=[[TASK_ID:[0-f]+]], 
codeptr_ra=[[SECT_BEGIN]], count=2
+  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_loop_static_end: 
parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}}, codeptr_ra=[[SECT_END]]
   // clang-format on
 
   return 0;
diff --git a/openmp/runtime/test/ompt/worksharing/sections_dispatch.c 
b/openmp/runtime/test/ompt/worksharing/sections_dispatch.c
index ec3bb5bd7552d..fc573ee5ae971 100644
--- a/openmp/runtime/test/ompt/worksharing/sections_dispatch.c
+++ b/openmp/runtime/test/ompt/worksharing/sections_dispatch.c
@@ -29,24 +29,22 @@ int main() {
 
   // CHECK: 0: NULL_POINTER=[[NULL:.*$]]
 
-  // CHECK: {{^}}[[MASTER_ID:[0-9]+]]: ompt_event_sections_begin:
+  // CHECK: {{^}}[[MASTER_ID:[0-9]+]]: ompt_event_loop_static_begin:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID:[0-f]+]],
   // CHECK-SAME: task_id=[[TASK_ID:[0-f]+]],
   // CHECK-SAME: codeptr_ra=[[SECT_BEGIN:(0x)?[0-f]+]], count=2
-  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_section_begin:
+  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_ws_loop_chunk_begin:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID]], task_id=[[TASK_ID]]
-  // CHECK-SAME: codeptr_ra=[[SECT_BEGIN]]
-  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_sections_end:
+  // CHECK: {{^}}[[MASTER_ID]]: ompt_event_loop_static_end:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}},
   // CHECK-SAME: codeptr_ra=[[SECT_END:(0x)?[0-f]+]]
 
-  // CHECK: {{^}}[[THREAD_ID:[0-9]+]]: ompt_event_sections_begin:
+  // CHECK: {{^}}[[THREAD_ID:[0-9]+]]: ompt_event_loop_static_begin:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID]], task_id=[[TASK_ID:[0-f]+]],
   // CHECK-SAME: codeptr_ra=[[SECT_BEGIN]], count=2
-  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_section_begin:
+  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_ws_loop_chunk_begin:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID]], task_id=[[TASK_ID]]
-  // CHECK-SAME: codeptr_ra=[[SECT_BEGIN]]
-  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_sections_end:
+  // CHECK: {{^}}[[THREAD_ID]]: ompt_event_loop_static_end:
   // CHECK-SAME: parallel_id=[[PARALLEL_ID]], task_id={{[0-f]+}},
   // CHECK-SAME: codeptr_ra=[[SECT_END]]
   // clang-format on

_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits

Reply via email to