Граф коммитов

14 Коммитов

Автор SHA1 Сообщение Дата
Ben Sander 4dfe77a99b Improve copy testing implementation.
- add tests for (unpinned/pinned) x H2H x D2D.
- Free memory at end of test.


[ROCm/clr commit: 1128610801]
2016-02-12 18:24:08 -06:00
Ben Sander 89e461988e Step1 in staging buffer copy.
- use StagingBuffer class for copies.
- refactor g_device to use array rather than vector.
   (keeps pointers from moving).


[ROCm/clr commit: 90af462b85]
2016-02-12 18:24:08 -06:00
Ben Sander 5978d5f372 Query tracked memory sizes.
Support more accurate hipMemGetInfo.  Add test to hipPointerAttrib.


[ROCm/clr commit: f464cedcf4]
2016-02-12 18:24:08 -06:00
Ben Sander 80d7c867d1 Remove ! USE_PINNED_HOST support
[ROCm/clr commit: f2c1bf3bc0]
2016-02-12 18:24:08 -06:00
Ben Sander 712750e1a5 Use memtracker 'appID' to store deviceID associated with ptr
[ROCm/clr commit: c04b5d3afb]
2016-02-12 18:24:08 -06:00
Ben Sander 2089e549eb Tracker improvements
- add API to add / remove user-pointers from the tracker.
- test for thread-safety with MultiThreadtest_2 - rapid
  insertions/removal.
- add mutex to provide thread-safety.
- rename tracker interface to "memtracker_..." for consistency.
- add am_memtracker_reset, connect to hipDeviceReset.
-


[ROCm/clr commit: 7216727fba]
2016-02-12 18:24:08 -06:00
Ben Sander fe67be1134 Create address tracker for am_alloc.
Tracks device where memory is allocated, pinned-host or device, and
more.

Uses memory-range-based lookups - so pointers that exist anywhere in

the range of hostPtr + size will find the associated AmPointerInfo.

The insertions and lookups use a self-balancing binary tree and
should support O(logN) lookup speed.


[ROCm/clr commit: 721508cc2f]
2016-02-12 18:24:08 -06:00
Ben Sander a50fa0f78e Fix bug in device bounds comparison.
Shows up in multi-GPU.


[ROCm/clr commit: f1bc9af294]
2016-02-12 18:24:08 -06:00
Evgeny Mankov 6add51ef8c Fix typo: maxThreadsPerMultiProcessor -> MaxSharedMemoryPerMultiprocessor
Device property MaxSharedMemoryPerMultiprocessor set equal to totalGlobalMem (HIP path).
Reason: MaxSharedMemoryPerMultiprocessor should be as the same as group memory size. Group memory will not be paged out, so, the physical memory size = total shared memory size = group region size. NVCC path remains untouched: CUDA's device property MaxSharedMemoryPerMultiprocessor is reported.

hipify is updated as well.


[ROCm/clr commit: 460b501cbb]
2016-02-12 01:29:20 +03:00
Evgeny Mankov 735d4738ad Device property maxThreadsPerMultiProcessor set equal to totalGlobalMem (HIP path).
Reason: maxThreadsPerMultiProcessor should be as the same as group memory size. Group memory will not be paged out, so, the physical memory size = total shared memory size = group region size.

NVCC path remains untouched: CUDA's device property maxThreadsPerMultiProcessor is reported.


[ROCm/clr commit: 1025341300]
2016-02-12 00:04:14 +03:00
Evgeny Mankov a8b7647f8b BDFID (BusID/DeviceID/FunctionID) support.
Except FunctionID (or DomainID in CUDA) support, because cudaDeviceProp::pciDomainID is not reported by CUDA.


[ROCm/clr commit: 658e9f0484]
2016-02-11 22:26:01 +03:00
Evgeny Mankov 9f596e0aab Device property concurrentKernels is added to hipDeviceProp_t struct.
For HCC path concurrentKernels is set to true since all ROCR hardware supports this feature.
For NVCC path concurrentKernels is obtained from CUDA's device property cudaDeviceProp::concurrentKernels.


[ROCm/clr commit: 4d4ca3ef3f]
2016-02-09 17:10:35 +03:00
Sam Kolton 136baccbe5 Implementation of hipDeviceGetAttribute()
[ROCm/clr commit: afe45964ae]
2016-02-04 17:39:27 +03:00
Ben Sander 28f87a0428 Initial commit for GPUOpen Launch
[ROCm/clr commit: 304171c1a2]
2016-01-26 20:14:33 -06:00