P4 to Git Change 1997981 by cpaquot@cpaquot-ocl-lc-lnx on 2019/09/13 11:17:32

SWDEV-203438 - [HIP] AllGather RCCL test issue
	The test tries to launch a kernel on two devices at once and they need to communicate with each other.
	For that, it uses a custom stream for each devices.
	Problem is in getNullStream we used to call syncStreams all the time
	and it was syncing all the streams even the ones on different devices.
	So that made the second kernel launch (on 2n dev) to wait for the first kernel to finish which
	would never occur since the first one was waiting for the second one.
	The fix is to not call syncStreams from getNullStream because we sync already anyway prior in general.

Affected files ...

... //depot/stg/opencl/drivers/opencl/api/hip/hip_context.cpp#21 edit
... //depot/stg/opencl/drivers/opencl/api/hip/hip_event.cpp#16 edit
... //depot/stg/opencl/drivers/opencl/api/hip/hip_internal.hpp#40 edit
... //depot/stg/opencl/drivers/opencl/api/hip/hip_memory.cpp#70 edit
... //depot/stg/opencl/drivers/opencl/api/hip/hip_module.cpp#41 edit
... //depot/stg/opencl/drivers/opencl/api/hip/hip_stream.cpp#24 edit
Cette révision appartient à :
foreman
2019-09-13 11:28:33 -04:00
Parent b0b5c92469
révision 0aa24f6d4d
6 fichiers modifiés avec 30 ajouts et 91 suppressions
+5 -8
Voir le fichier
@@ -218,6 +218,9 @@ hipError_t ihipModuleLaunchKernel(hipFunction_t f,
hipEvent_t startEvent, hipEvent_t stopEvent, uint32_t flags = 0,
uint32_t params = 0)
{
HIP_INIT_API(f, gridDimX, gridDimY, gridDimZ, blockDimX, blockDimY, blockDimZ,
sharedMemBytes, hStream, kernelParams, extra, startEvent, stopEvent, flags, params);
hip::Function* function = hip::Function::asFunction(f);
amd::Kernel* kernel = function->function_;
amd::Device* device = hip::getCurrentContext()->devices()[0];
@@ -226,14 +229,8 @@ hipError_t ihipModuleLaunchKernel(hipFunction_t f,
hip::Event* eStart = reinterpret_cast<hip::Event*>(startEvent);
hip::Event* eStop = reinterpret_cast<hip::Event*>(stopEvent);
amd::HostQueue* queue;
if (hStream == nullptr) {
hip::syncStreams();
queue = hip::getNullStream();
} else {
hip::getNullStream()->finish();
queue = reinterpret_cast<hip::Stream*>(hStream)->asHostQueue();
}
amd::HostQueue* queue = hip::getQueue(hStream);
if ((params & amd::NDRangeKernelCommand::CooperativeGroups) &&
!device->info().cooperativeGroups_) {
return hipErrorLaunchFailure;