P4 to Git Change 1982729 by gandryey@gera-win10 on 2019/08/13 17:40:55
SWDEV-79445 - OCL generic changes and code clean-up
- Use max number of waves per SIMD in the scratch calculation to allow async kernel execution with the scratch buffer
Affected files ...
... //depot/stg/opencl/drivers/opencl/runtime/device/pal/paldevice.cpp#155 edit
... //depot/stg/opencl/drivers/opencl/runtime/device/pal/paldevice.hpp#43 edit
... //depot/stg/opencl/drivers/opencl/runtime/device/pal/palvirtual.cpp#147 edit
[ROCm/clr commit: 34e526d77f]
Šī revīzija ir iekļauta:
@@ -2444,7 +2444,7 @@ bool VirtualGPU::submitKernelInternal(const amd::NDRangeContainer& sizes, const
|
||||
// Use maximum available slots for all dispatches to allow async on the same queue
|
||||
// HW value loaded into SGPR is an offset value calculated as
|
||||
// wave_slot * COMPUTE_TMPRING_SIZE.WAVESIZE
|
||||
dispatchParam.workitemPrivateSegmentSize = scratch->privateMemSize_;
|
||||
dispatchParam.workitemPrivateSegmentSize = std::max(hsaKernel.spillSegSize(), scratch->privateMemSize_);
|
||||
}
|
||||
dispatchParam.pCpuAqlCode = hsaKernel.cpuAqlCode();
|
||||
dispatchParam.hsaQueueVa = hsaQueueMem_->vmAddress();
|
||||
|
||||
Atsaukties uz šo jaunā problēmā
Block a user