gilbertlee-amd
c981e76efe
Clique kernel support ( #295 ) ( #15 )
...
* Adding experimental clique-based kernels (opt-in only)
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
[ROCm/rccl commit: 3e62ceddc5 ]
2021-01-28 09:45:01 -07:00
Wenkai Du
7f9c15b843
Use less unroll for clique kernels ( #313 )
...
[ROCm/rccl commit: 41e47a36e7 ]
2021-01-15 17:48:10 -08:00
Wenkai Du
d4382de267
Improve collective trace
...
[ROCm/rccl commit: 2ddbe6646b ]
2021-01-14 19:28:01 -05:00
Wenkai Du
2c49121171
Port alltoall[v]
...
[ROCm/rccl commit: f4d5d3d620 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
41bead5a4e
Do not allow GPU as intermediate
...
[ROCm/rccl commit: 105db19a11 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
34c6013299
Revert "Changes to topology based on XGMI ( #272 )"
...
This reverts commit 0a9adc16f4 .
[ROCm/rccl commit: e055229e56 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
adff98765c
Merge remote-tracking branch 'nccl/master' into no-target-id
...
[ROCm/rccl commit: d469947641 ]
2021-01-14 19:27:53 -05:00
Jonas Zhou
1db601566f
x86: Add CPU detection for Zhaoxin processors
...
Signed-off-by: Jonas Zhou <JonasZhou@zhaoxin.com >
[ROCm/rccl commit: 3996562690 ]
2020-12-17 11:15:18 -08:00
Wenkai Du
4ea285c527
Fix Rome PCIe 2 node topology generation ( #310 )
...
[ROCm/rccl commit: 373a108516 ]
2020-12-15 17:16:17 -08:00
Wenkai Du
b68ff1ebba
Add Rome model and improve search ( #305 )
...
[ROCm/rccl commit: 975b14dffa ]
2020-11-17 14:55:06 -08:00
Sylvain Jeaugey
a8908b34ee
2.8.3-1
...
Optimization for Tree allreduce on A100.
Improve aggregation performance.
Use shared buffers for inter-node send/recv.
Add NVTX profiling hooks.
Accelerate alltoall connections by merging communication for all
channels.
Add support for one hop communication through NVLink, for faster
send/recv communication on cubemesh topologies like DGX-1.
Improve alltoall scheduling to better balance intra/inter node
communication.
Increase send/recv parallelism by 8x, each warp sending or
receiving to a different peer.
Net: move to v4.
Net: make flush operation asynchronous to accelerate alltoall.
Net: define maximum number of requests.
Fix hang when using LL128 protocol after 2^31 steps.
Fix #379 : topology injection failing when using less GPUs than
described in the XML.
Fix #394 : protocol mismatch causing hangs or crashes when using
one GPU per node.
[ROCm/rccl commit: 920dbe5b35 ]
2020-11-17 11:08:52 -08:00
Wenkai Du
f19cbc8e51
Use device's link width and speed if port doesn't report ( #304 )
...
[ROCm/rccl commit: 554729079d ]
2020-11-13 17:58:04 -08:00
Stanley Tsang
f373cd2fdc
Fixing IPC handle leak ( #302 )
...
[ROCm/rccl commit: 2958f7eace ]
2020-11-13 10:32:42 -07:00
gilbertlee-amd
f66d05193a
Adding RCCL_CLIQUE_DEBUG to help debug experimental clique feature ( #300 )
...
[ROCm/rccl commit: c8d08a7c2f ]
2020-11-13 09:07:11 -07:00
Wenkai Du
62d21047b8
Skip unused peer connection in scatter and gather ( #301 )
...
[ROCm/rccl commit: 4e68229c8b ]
2020-11-12 15:47:34 -08:00
gilbertlee-amd
a7ef699687
Clique kernel support ( #295 )
...
* Adding experimental clique-based kernels (opt-in only)
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
[ROCm/rccl commit: 41bcfb8878 ]
2020-11-10 15:44:10 -07:00
Wenkai Du
a0109b2eae
Use ncclSend/ncclRecv for alltoall type of collectives as default ( #297 )
...
[ROCm/rccl commit: 2e8b3a0857 ]
2020-11-09 11:23:17 -08:00
Wenkai Du
a4dd1a9548
Improve GPU direct RDMA handling on Rome ( #294 )
...
[ROCm/rccl commit: 709b7e4880 ]
2020-11-03 14:29:08 -08:00
Wenkai Du
c0c64d970a
Add more Rome models ( #292 )
...
[ROCm/rccl commit: dfa3c41ede ]
2020-10-30 21:26:04 -07:00
xietingwew
0277094dd2
fix proxyArgs for trace log
...
[ROCm/rccl commit: 084207e685 ]
2020-10-21 09:18:40 -07:00
Wenkai Du
1aae6b1344
Fix incorrect pointer checking for scatter and gather ( #285 )
...
[ROCm/rccl commit: dcad0ef7cb ]
2020-10-19 13:27:09 -07:00
Wenkai Du
194135a40c
Merge remote-tracking branch 'nccl/master' into nccl_sync
...
[ROCm/rccl commit: c835d8263a ]
2020-10-15 18:42:38 -04:00
gilbertlee-amd
94437eef28
Revert "Initial support for clique-based kernels ( #276 )" ( #280 )
...
This reverts commit d68a532bc6 .
[ROCm/rccl commit: 84a2541e01 ]
2020-10-15 11:30:18 -07:00
Sylvain Jeaugey
591ffd32fe
Fix affinity move
...
[ROCm/rccl commit: 0e14394c5f ]
2020-10-13 16:58:05 -07:00
Sylvain Jeaugey
5de6b6681d
Make sure proxy threads inherit the CPU affinity.
...
[ROCm/rccl commit: c6dbdb0084 ]
2020-10-13 16:37:52 -07:00
Wenkai Du
8b120c0508
Update Rome single node models ( #277 )
...
[ROCm/rccl commit: 33babcb5e2 ]
2020-10-13 13:33:09 -07:00
gilbertlee-amd
d68a532bc6
Initial support for clique-based kernels ( #276 )
...
* Initial support for clique-based kernels
[ROCm/rccl commit: 2b8184808d ]
2020-10-13 11:22:04 -06:00
Wenkai Du
41260bb948
Rework Rome detection and add multiple network ports models ( #274 )
...
* Rework Rome detection and add multiple network ports models
* Remove unused opCount in p2p transport
[ROCm/rccl commit: ae008fd2db ]
2020-10-07 13:37:36 -07:00
Wenkai Du
dbde26e681
Add Alltoallv RCCL kernel implementation ( #269 )
...
* Add alltoallv API and implementation
* Extend Rome P2P channel limit to multinode and alltoall kernels
* topo_expl: fix compilation and sync up with main
* gtest: use RCCL alltoallv API
* Code review changes
[ROCm/rccl commit: b871ea3c0c ]
2020-09-30 16:25:36 -07:00
Stanley Tsang
67a8d86d78
Updating inline asm to not require explicit L1 cache invalidation ( #270 )
...
[ROCm/rccl commit: acca2ae20a ]
2020-09-25 13:46:26 -06:00
gilbertlee-amd
0a9adc16f4
Changes to topology based on XGMI ( #272 )
...
* Alterations to topology search to improve XGMI-enabled nodes
[ROCm/rccl commit: 01bd2573db ]
2020-09-25 12:20:09 -06:00
Wenkai Du
7ba087e069
Ensure all ranks on same send/receive or alltoall kernel path ( #271 )
...
[ROCm/rccl commit: 44fcde7835 ]
2020-09-24 08:25:04 -07:00
Wenkai Du
37f7eec6b7
Change network plugin name to librccl-net.so ( #266 )
...
[ROCm/rccl commit: d871fceb54 ]
2020-09-18 13:23:30 -07:00
Wenkai Du
f0a303664e
Limit P2P channels on Rome
...
[ROCm/rccl commit: 42955f5f4f ]
2020-09-17 17:20:32 -07:00
Wenkai Du
a3402d6aeb
Merge pull request #262 from wenkaidu/alignment
...
Make data alignment requirements matching ISA manual
[ROCm/rccl commit: 60819dcf8d ]
2020-09-08 10:40:42 -07:00
Wenkai Du
09639a5d54
Fix broken profiling build ( #263 )
...
[ROCm/rccl commit: e2042ccf8a ]
2020-09-02 15:39:52 -07:00
Wenkai Du
cfa1228504
Make data alignment requirements matching ISA manual
...
From https://developer.amd.com/wp-content/resources/Vega_Shader_ISA.pdf
8.1.7. Alignment
For Dword or larger reads or writes, the two LSBs of the byte-address
are ignored, thus forcing Dword alignment.
[ROCm/rccl commit: 4751992231 ]
2020-09-01 21:21:58 +00:00
Wenkai Du
778ab61097
Fix incorrect threads split in sendrecv ( #261 )
...
[ROCm/rccl commit: 4180e6409e ]
2020-08-31 17:33:22 -07:00
Wenkai Du
03bb6bcb54
Increase minimal channels for gfx908 ( #259 )
...
[ROCm/rccl commit: c5cbece6d0 ]
2020-08-26 11:40:11 -07:00
Wenkai Du
0898fea746
Only use software barrier for synchronization ( #258 )
...
[ROCm/rccl commit: b0919dc46c ]
2020-08-25 13:16:34 -07:00
Wenkai Du
5f49a0e088
Add NPS4 support on some models ( #256 )
...
* Add NPS4 support on some models
* Add XML models
[ROCm/rccl commit: 391bbf3f1e ]
2020-08-19 11:03:20 -07:00
Wenkai Du
3d5fb8142e
Add another Rome model ( #249 )
...
* Add another Rome model
* Add gfx908 4P3L models and support
* Revert "Use cached value for detecting GDR support only once"
This reverts commit 0108a1219d .
* Skip using ibverb for GPU direct RDMA detection
* Fine tune one Rome model
[ROCm/rccl commit: a51e4071e3 ]
2020-08-17 10:51:02 -07:00
Wenkai Du
f242a2f0b0
Collect gcnArch and hipDeviceArch_t in XML ( #252 )
...
[ROCm/rccl commit: 7e3d8a31cc ]
2020-08-12 15:48:38 -07:00
Wenkai Du
e5ec2d94d5
Merge pull request #248 from wenkaidu/2.7.8
...
2.7.8
[ROCm/rccl commit: 066223333d ]
2020-08-11 08:20:37 -07:00
Wenkai Du
14ad6ff3b4
Merge remote-tracking branch 'nccl/master' into 2.7.8
...
[ROCm/rccl commit: 7e3f841fab ]
2020-08-10 16:11:00 +00:00
Wenkai Du
c9815aaa36
Add more Rome 4P2H models
...
[ROCm/rccl commit: 09ef75656a ]
2020-08-06 18:20:02 +00:00
Jack Snyder
dca19952fd
Setting type when gpu sub node is discovered
...
[ROCm/rccl commit: de49a77074 ]
2020-08-05 13:39:23 -07:00
Eric Badger
d6a78cb1c7
Don't require NIC devices to have specific PCI class
...
If a PCI node is the parent of a NIC, treat it as such, regardless of
the PCI class code for the device. This allows non-traditional devices
to act as NICs via the net plugin mechanism.
For consistency, treat GPUs similarly.
[ROCm/rccl commit: 700c0e0f24 ]
2020-08-05 12:46:29 -07:00
Wenkai Du
5f96f13e75
Allow setup ring through NCCL_RINGS to facilitate testing
...
[ROCm/rccl commit: 5b03132ace ]
2020-08-04 21:07:00 +00:00
Wenkai Du
22a6211eaf
Improve 4P2H topology on Rome ( #243 )
...
1. Use bi-directional rings
2. GPU search is sorted by PCI device ID to get consistent results
[ROCm/rccl commit: d1e20b4c5e ]
2020-07-28 14:21:44 -07:00