TF doesn't reserve all available memory now. If any
client wants to reserve they can explicitly set
HIP_HIDDEN_FREE_MEM env var
Change-Id: Ied3a948b79f49aa7327f6a820e9789e39cec143b
[ROCm/clr commit: d8ca3c632c]
Workaround hipStream deadlock issue as the same lock was used twice SWDEV-236746
Change-Id: Icc60104ce6edf4cfd2a3a889bab78a6caadd50b7
[ROCm/clr commit: f3ee29cdb2]
If a offset of the pointer is passed to free it may release
the mem object but may not release from MemObjMap. Erase the map
by getting the parent pointer.
Change-Id: I06b92548de2d49b4029efe6b511329225007cc55
[ROCm/clr commit: 2b6fea4348]
Support gfx908 as part of the default AMDGPU_TARGETS. MIGraphX requires this change.
Change-Id: I692f87f27829778e04f59c9ca655c6e8cbc00abc
[ROCm/clr commit: 29c7c9b1c2]
Similar to HCC, link with compiler-rt to support __fp16 and _Float16 type conversions in ONNX models. This should resolve SWDEV-238491.
Change-Id: Iad8dcff568831719f501f562a04023326ae8036c
[ROCm/clr commit: d93134e727]
For some reason the test is causing latest rocr to crash. Disabling for now to investigate.
Change-Id: I78996241e6756c36af4b7f5fcb34e915bf33573e
[ROCm/clr commit: ce2c28c745]
Tests are now excluded from the main build, however ocltst can still be built by running `cmake --build . --target ocltst`.
Test modules can be run by invoking the test_${module_name} target, for example `cmake --build . --target test_oclruntime`.
Change-Id: I6f40d3bd318e821e4fbf35ccfa395dc7036673cb
[ROCm/clr commit: 8790099ba9]
Select cpu in terms of the smallest Numa distance for a GPU device.
This will improve performance of hipMemcpy in the mode of
hipMemcpyHostToDevice or hipMemcpyDeviceToHost for small buffer.
`
Change-Id: I2860f1f83b79be0dff7bf5e64cf68ab4448db0a1
[ROCm/clr commit: aedb9590be]
Fix a build error in centos 6.10. When linking, ld cannot find function
clock_getres. This function is from librt (for glibc version before 2.17).
Change-Id: I698f768d0618a054d5d0fc73b9f1916a7ce542fb
[ROCm/clr commit: 9a8bd9e68b]
This only adds source files for ocltst and the following test modules - oclruntime, oclperf, oclgl, ocldx. There's no build files for now.
Change-Id: I0f8d9d074c45d82e92f7d30bf22753102f272f4f
[ROCm/clr commit: 75e6add24d]
This reverts commit 5540e98745.
Reason for revert: HIP_ENABLE_LAZY_KERNEL_LOADING is needed before the runtime is initialized, so this utility cannot be used
Change-Id: I49f8ddb98c9a85b9a77b8fd4b236d06b6b2b0f32
[ROCm/clr commit: ac8d1ba687]
The hipOccupancyMaxPotentialBlockSize API is meant to return the
number of threads for the highest-occupancy workgroup, and the number
of those workgroups. It was previously calculating the number of
maximum-sized workgroups that would fit on a single CU. This is
a mixture of the API we wanted (to calculate max potential block size)
and the MaxBlocksPerMultiprocessor function.
This patch fixes it up so that the internal occupancy calculation
function works for two uses: the traditional function that calculates
the maximum blocks per multiprocessor when a user passes in a fixed
block size (used for hipMaxBlocksPerMultiprocessor style functions)
and a function that calculates the size of a block that would lead
to maximum occupancy, and how many blocks of that size would be
needed to fill the whole GPU (for hipOccupancyMaxPotentialBlockSize
style functions).
This also updates the occupancy calculation function to prepare for
gfx10, which does not have SGPR-based occupancy limits.
Change-Id: Ie007b3f9d5ebc4e166b50a3a051498af35650f35
[ROCm/clr commit: 90453b68d3]