[SWDEV-497665] Blocked cudaMemcpyAsync race condition by synchronizing (#1447)

* Switched calls to `cudaMemcpyAsync` to be `cudaMemcpy` in `ncclTransportP2pSetup` to avoid race condition with `cudaIpcOpenMemHandle` inside p2p `connect`. See `ncclP2pImportShareableBuffer`.

* Moved synchronize outside of the loop, as it isn't necessary to sync between every iteration of the loop.

[ROCm/rccl commit: c158d3a9b4]
This commit is contained in:
corey-derochie-amd
2025-01-03 13:06:47 -07:00
committed by GitHub
parent dd3fc22531
commit 17becdb7f8
+2
View File
@@ -263,6 +263,8 @@ ncclResult_t ncclTransportP2pSetup(struct ncclComm* comm, struct ncclTopoGraph*
}
}
CUDACHECKGOTO(cudaStreamSynchronize(comm->sharedRes->hostStream.cudaStream), ret, fail);
if (timeReported) {
struct timeval now;
gettimeofday(&now, NULL);