================
@@ -3934,6 +3961,55 @@ emitInnerParallelForWhenCombined(CodeGenFunction &CGF,
         HasCancel = D->hasCancel();
     }
     CodeGenFunction::OMPCancelStackRAII CancelRegion(CGF, EKind, HasCancel);
+
+    CodeGenModule &CGM = CGF.CGM;
+    if (CGM.getOpenMPRuntime().canPromoteToNoLoop()) {
+      // Prepare the loop variables and their privatization.
+      emitLoopIterationspaceVars(CGF, S);
+      OMPLoopScope PreInitScope(CGF, S);
+
+      CodeGenFunction::OMPPrivateScope PrivateScope(CGF);
+      CGF.EmitOMPPrivateClause(S, PrivateScope);
----------------
nicebert wrote:

> hm, but why does the existing EmitOMPWorksharingLoop handle it?

Because the existing path isn't only used for SPMD kernels. It's shared by 
every worksharing loop: plain `for`, `parallel for`, host code, and `distribute 
parallel for` inside a generic mode kernel. In those cases either there is no 
distribute level at all or only the main thread ran it, so the loop has to make 
the per-thread copy itself. For SPMD kernels the regular path handling 
firstprivate again just creates a redundant copy, e.g. a firstprivate array 
gets copied once at the distribute level and then again in the parallel region. 
A no-loop kernel is always SPMD, so every thread runs the distribute level and 
the copy made there is already per-thread.

> Regarding the split: I was just thinking about sth like this: #220540 (I 
> already created that for demonstration in an earlier discussion). Without the 
> no-loop work, this split is a bit hard to test, but in the context of your 
> work, it might be useful and would probably simplify things for you?

I don't think it would simplify things here. This PR already factors out the 
pieces it makes sense to share with the existing paths, the iteration space 
setup and the canonical loop creation, so the no-loop path reuses them instead 
of duplicating them. It was also split up to keep the review focused, and 
mixing a bigger restructuring of the loop emission into it would make it harder 
to review, not easier. If we want that split, I think it should be its own PR.

https://github.com/llvm/llvm-project/pull/224041
_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits

Reply via email to