Ben Sander
784ebcbc86
Fix memcpy for Titan. Add <threads> to common includes
2016-02-22 15:09:23 -06:00
Ben Sander
16b04fc0d3
Merge branch 'memtracker' of https://github.com/AMDComputeLibraries/HIP-privatestaging into memtracker
2016-02-22 08:33:47 -06:00
Ben Sander
28990567fb
Improve async copy implementation.
...
- Add device-side signal waits when transitioning between command classes
(Kernel, H2D copy, D2H copy).
- Support waiting in staged memory copies as well.
- Add several chicken bits to control implementation:
- HIP_DISABLE_ENQ_BARRIER
- HIP_DISABLE_BIDIR_MEMCPY
- HIP_ONESHOT_COPY_DEP
- Refactor signal pool to support efficient deallocation based on
signsequnm.
- Deallocate copy signals on eventSynchronize.
- Improve copy tests, add pingpong.
2016-02-22 23:15:24 -06:00
Ben Sander
d5c777268a
Track last command to a stream.
...
Passing simple tests.
2016-02-20 11:02:07 -06:00
Ben Sander
16ff0757a6
Describe how to update HTML docs
2016-02-19 01:56:17 -06:00
Evgeny Mankov
8aace64dce
Device property memoryClockRate implementation.
...
+ Device property memoryClockRate is added to hipDeviceProp_t struct.
+ Device attribute hipDeviceAttributeMemoryClockRate is added to hipDeviceAttribute_t struct.
+ Tests update.
+ Rename hipDevAttrConcurrentKernels to hipDeviceAttributeConcurrentKernels.
2016-02-18 17:25:28 +03:00
Evgeny Mankov
d4bd94e9a0
Attribute hipDevAttrConcurrentKernels for obtaining Device property concurrentKernels is added.
2016-02-18 14:34:18 +03:00
Ben Sander
d0f9881d60
Merge branch 'memtracker' of https://github.com/AMDComputeLibraries/HIP-privatestaging into memtracker
2016-02-17 23:06:51 -06:00
Ben Sander
9a82d316c3
Support HSA_PATH env, async path tweak
2016-02-17 21:22:07 -06:00
Ben Sander
0cdbe1ff05
more work on async copies
2016-02-17 00:59:12 -06:00
Ben Sander
731a2a58d3
Add comments to tests
2016-02-16 01:58:24 -06:00
Ben Sander
3b2d4acabc
Remove old include path.
2016-02-15 05:40:37 -06:00
Ben Sander
afbe451b0d
Fix tests to account for multi-gpu
2016-02-15 05:19:52 -06:00
Ben Sander
8939b4f0e5
Add multi-threading synchonization on staging buffers and signals.
...
Also pre-allocate a couple signals for copies.
2016-02-13 03:18:01 -06:00
Ben Sander
a002833a89
D2H multi-buffer
2016-02-13 01:15:23 -06:00
Ben Sander
2353cbb028
Improve copy testing
2016-02-12 18:24:08 -06:00
Ben Sander
1128610801
Improve copy testing implementation.
...
- add tests for (unpinned/pinned) x H2H x D2D.
- Free memory at end of test.
2016-02-12 18:24:08 -06:00
Ben Sander
90af462b85
Step1 in staging buffer copy.
...
- use StagingBuffer class for copies.
- refactor g_device to use array rather than vector.
(keeps pointers from moving).
2016-02-12 18:24:08 -06:00
Ben Sander
f464cedcf4
Query tracked memory sizes.
...
Support more accurate hipMemGetInfo. Add test to hipPointerAttrib.
2016-02-12 18:24:08 -06:00
Ben Sander
7216727fba
Tracker improvements
...
- add API to add / remove user-pointers from the tracker.
- test for thread-safety with MultiThreadtest_2 - rapid
insertions/removal.
- add mutex to provide thread-safety.
- rename tracker interface to "memtracker_..." for consistency.
- add am_memtracker_reset, connect to hipDeviceReset.
-
2016-02-12 18:24:08 -06:00
Ben Sander
721508cc2f
Create address tracker for am_alloc.
...
Tracks device where memory is allocated, pinned-host or device, and
more.
Uses memory-range-based lookups - so pointers that exist anywhere in
the range of hostPtr + size will find the associated AmPointerInfo.
The insertions and lookups use a self-balancing binary tree and
should support O(logN) lookup speed.
2016-02-12 18:24:08 -06:00
Evgeny Mankov
460b501cbb
Fix typo: maxThreadsPerMultiProcessor -> MaxSharedMemoryPerMultiprocessor
...
Device property MaxSharedMemoryPerMultiprocessor set equal to totalGlobalMem (HIP path).
Reason: MaxSharedMemoryPerMultiprocessor should be as the same as group memory size. Group memory will not be paged out, so, the physical memory size = total shared memory size = group region size. NVCC path remains untouched: CUDA's device property MaxSharedMemoryPerMultiprocessor is reported.
hipify is updated as well.
2016-02-12 01:29:20 +03:00
Evgeny Mankov
658e9f0484
BDFID (BusID/DeviceID/FunctionID) support.
...
Except FunctionID (or DomainID in CUDA) support, because cudaDeviceProp::pciDomainID is not reported by CUDA.
2016-02-11 22:26:01 +03:00
Maneesh Gupta
ed2d86f3a9
Updated readme for test
2016-02-11 13:06:58 +05:30
streamhsa
90add185fd
Remove test for atomicInc and atomicDec
2016-02-10 21:02:52 +08:00
streamhsa
56f1832e70
Updated readme for test
2016-02-10 20:05:59 +08:00
streamhsa
2f8d56e903
Resolved test issues
2016-02-10 20:01:16 +08:00
streamhsa
310023e273
Rename test hipInfo as hipGetDeviceAttribute
2016-02-09 13:19:32 +08:00
Ben Sander
ce2fc0f7fe
Test fixes:
...
- Remove reference to missing test.
- Add hipMemset back.
- Parse --gpu option to specify default starting GPU.
2016-02-08 22:55:23 -06:00
Ben Sander
c482a3f456
in HIPCHECK, only run command once even if error occurs
2016-02-08 21:45:49 -06:00
Ben Sander
26854bb31c
Fix HIP_PLATFORM detection
2016-02-05 07:15:46 -06:00
Sam Kolton
afe45964ae
Implementation of hipDeviceGetAttribute()
2016-02-04 17:39:27 +03:00
Ben Sander
2faf1dfe6e
Merge branch 'master' into privatestaging
2016-02-03 09:39:19 -06:00
Ben Sander
182296ce59
Remove warning on ballot/any/all and pop/clz.
...
Since these are supported in HIP no reason to emit warnings.
2016-02-02 10:02:48 -06:00
streamhsa
e5a491f3c8
Added test for ballot and removing HIP_FUNCTION from hipSampleAtomicsTest.cpp -sandeep
2016-02-02 14:50:55 +05:30
Maneesh Gupta
405ee35a04
Add double and integer intrinsics to test
2016-02-01 16:00:45 +05:30
Jack Chung
518ef58652
Disable sincosf which has trouble on hcc now.
2016-02-01 17:42:37 +08:00
Maneesh Gupta
97fb876c6a
Split math function tests into several smaller tests
2016-02-01 14:36:50 +05:30
Maneesh Gupta
8d01a1db15
Disable testing of unsupported single precision intrinsics
2016-02-01 14:34:28 +05:30
Ben Sander
304171c1a2
Initial commit for GPUOpen Launch
2016-01-26 20:14:33 -06:00