Stanley Tsang
6680e23c63
Message queue refactor to POSIX implementation and leak fix ( #355 )
...
* Fixing message queue leak.
* Using POSIX implementation of Message Queues
* Adding unlink to msgqueue
* MsgQueue update
* Adding timeout check to msgqueue broadcast; tightening up system checks
* Removing unnecessary code
* Removing extra argument from print
* Adding explicit msg queue close call to all other ranks
[ROCm/rccl commit: 70597789d0 ]
2021-04-23 11:33:20 -06:00
Wenkai Du
e28bac31aa
Tune number of channels for gfx90a ( #349 )
...
[ROCm/rccl commit: 415c7cd3d1 ]
2021-04-19 15:27:01 -07:00
Wenkai Du
951d89b12f
Use correct WARP_SIZE for gfx1030 ( #348 )
...
[ROCm/rccl commit: 9c718ce6d6 ]
2021-04-14 14:09:52 -07:00
Wenkai Du
661b1351a3
Limit max channels for ring graph on single node Rome ( #347 )
...
* Limit max channels for ring graph on single node Rome
* Partially revert "Use non-temporal access for streaming data (#341 )"
[ROCm/rccl commit: a79f74082e ]
2021-04-14 10:14:54 -07:00
Wenkai Du
0f4d497edc
Add gfx90a target ( #344 )
...
* Add gfx90a target
* Support gfx90a topology
Co-authored-by: Eiden Yoshida <eiden.yoshida@amd.com >
[ROCm/rccl commit: 1fe031402a ]
2021-04-14 09:29:00 -06:00
Sylvain Jeaugey
20da390b96
2.9.6-1
...
Add support for CUDA graphs.
Fuse BCM Gen4 switches to avoid suboptimal performance on some platforms. Issue #439 .
Fix bootstrap issue caused by connection reordering.
Fix CPU locking block.
Improve CollNet algorithm.
Improve performance on DGX A100 for communicators with only one GPU per node.
[ROCm/rccl commit: a46ea10583 ]
2021-04-12 16:00:46 -07:00
TomSang
6105af2dfc
Add detection of cooperative multi device launch attribute ( #345 )
...
[ROCm/rccl commit: 87f12cbb86 ]
2021-04-11 13:29:24 -07:00
Wenkai Du
6b3389b790
Use non-temporal access for streaming data ( #341 )
...
* Use non-temporal access for streaming data
* Revert to ulong2 after fixing compiling issue
[ROCm/rccl commit: 9dfc2c183e ]
2021-04-07 17:34:35 -07:00
gilbertlee-amd
a7e99734e8
Fixing clique-topology detection ( #342 )
...
* Fixing clique-topology detection
* Fix to enable multi-process clique-based kernels
[ROCm/rccl commit: caba0a63d2 ]
2021-04-07 11:29:44 -06:00
Wenkai Du
b4a7fa7011
Cleanup number of channels calculation ( #340 )
...
[ROCm/rccl commit: e26ad2995e ]
2021-04-05 17:51:56 -07:00
Wenkai Du
8927d8bf17
Fix incorrect net counting ( #339 )
...
* Fix incorrect net counting
* Add comments
[ROCm/rccl commit: 17491c918e ]
2021-04-05 12:21:57 -07:00
Wenkai Du
018a31877c
Rework network port trimming code ( #338 )
...
* Rework network port trimming code
* Move Rome related changes to separate source files
[ROCm/rccl commit: 1d2946ee4b ]
2021-03-31 10:25:59 -07:00
Wenkai Du
4da9c54d4e
Check fine grained memory on peer GPU before enabling P2P ( #337 )
...
[ROCm/rccl commit: 0c78553ee0 ]
2021-03-30 09:06:39 -07:00
Wenkai Du
065bde98d8
collnet: support multiple NICs ( #335 )
...
[ROCm/rccl commit: d87dc7c2e8 ]
2021-03-25 20:59:32 -07:00
Stanley Tsang
016a11f61b
Fixing message queue leak. ( #331 )
...
[ROCm/rccl commit: 289db2a636 ]
2021-03-25 19:11:43 -06:00
Wenkai Du
07548845a6
Remove HDP workaround for ROCm 4.2 HIP ( #334 )
...
[ROCm/rccl commit: 0fbb9510a5 ]
2021-03-23 20:11:37 -07:00
Wenkai Du
287ed0f18a
Enable collnet in RCCL ( #333 )
...
* Enable CollNet and use different number of channels
* topo_expl: enable collnet
[ROCm/rccl commit: 1d6244b18d ]
2021-03-19 12:58:13 -07:00
Wenkai Du
052119c20a
Sort GPUs by HIP device ID ( #329 )
...
* Sort GPUs by HIP device ID
* Remove extra space
[ROCm/rccl commit: b46260260a ]
2021-03-16 16:51:32 -07:00
Wenkai Du
7374c512d9
Add GPU memory usage tracker ( #326 )
...
[ROCm/rccl commit: f60b76c67a ]
2021-03-06 20:32:30 -08:00
Wenkai Du
b7253710ca
Revert "Port alltoall[v]" ( #325 )
...
This reverts commit 2c49121171 .
[ROCm/rccl commit: 8e180cf087 ]
2021-03-06 13:59:31 -08:00
Wenkai Du
bcf4ecb0e3
Enable local sendrecv over network if GDR is available on all GPUs ( #324 )
...
[ROCm/rccl commit: c018edf0f2 ]
2021-03-05 19:59:41 -08:00
gilbertlee-amd
8d90821062
Adding pthread_join / pthread_detach to clean up pthreads to avoid leaks ( #322 )
...
[ROCm/rccl commit: f4a9b9acba ]
2021-02-26 16:29:55 -07:00
Wenkai Du
020ccb6730
Update tuning parameters for XGMI and NET
...
[ROCm/rccl commit: e820a943e9 ]
2021-02-23 21:41:26 +00:00
Wenkai Du
d24df90882
Match NBIO only when GPUs and NICs are directly connected to CPU
...
[ROCm/rccl commit: ec8d89b1dd ]
2021-02-22 18:52:29 -05:00
Stanley Tsang
ce02076d64
Fixing cache deletion for CliqueManager; updating copyright
...
[ROCm/rccl commit: 45f5255f7c ]
2021-02-19 22:22:46 +00:00
Wenkai Du
6c3ccc2192
Add support to another Rome model
...
[ROCm/rccl commit: 95f178324c ]
2021-02-18 02:00:31 +00:00
Wenkai Du
ab71643c99
Merge remote-tracking branch 'nccl/master' into 2.8.3
...
[ROCm/rccl commit: c985358e11 ]
2021-02-15 18:44:47 -05:00
Wenkai Du
34d7735419
Move HDP flush to CPU
...
[ROCm/rccl commit: bf8eb40705 ]
2021-02-12 18:06:19 +00:00
Sylvain Jeaugey
fc7bdb38a5
2.8.4-1
...
Fix hang in corner cases of alltoallv using point to point send/recv.
Harmonize error messages.
Fix missing NVTX section in the license.
Update README.
[ROCm/rccl commit: 911d61f214 ]
2021-02-09 15:36:48 -08:00
Wenkai Du
602fe60428
Fix GDRDMA read and remove unused files
...
[ROCm/rccl commit: 9cc3b56166 ]
2021-02-09 01:34:39 +00:00
Stanley Tsang
f152c8d160
Update MP UT to support arbitrary # of GPUs; multiple bugfixes ( #16 )
...
* Fixing temp file creation/deletion for Clique kernel mode.
* Refactoring of MP unit tests; include bugfixes and general support for any number of GPUs
* GroupCall MP UT properly quits when too many devices specified
* MP UT will programmatically set NCCL_COMM_ID if not specified; updated install script
[ROCm/rccl commit: d00b7d17bd ]
2021-02-05 16:49:25 -08:00
Wenkai Du
ae5779702a
Merge remote-tracking branch 'origin/develop' into 2.8.3
...
[ROCm/rccl commit: ab1e7a0318 ]
2021-02-04 20:02:34 -05:00
gilbertlee-amd
16d625ca27
Tuning some clique-based kernel parameters ( #315 )
...
[ROCm/rccl commit: 1990ffd76a ]
2021-02-03 20:00:08 -07:00
Wenkai Du
57abf599b2
Enable GPU direct RDMA read from GPU
...
[ROCm/rccl commit: 5f97122442 ]
2021-02-03 02:48:30 +00:00
gilbertlee-amd
c981e76efe
Clique kernel support ( #295 ) ( #15 )
...
* Adding experimental clique-based kernels (opt-in only)
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
[ROCm/rccl commit: 3e62ceddc5 ]
2021-01-28 09:45:01 -07:00
Wenkai Du
7f9c15b843
Use less unroll for clique kernels ( #313 )
...
[ROCm/rccl commit: 41e47a36e7 ]
2021-01-15 17:48:10 -08:00
Wenkai Du
d4382de267
Improve collective trace
...
[ROCm/rccl commit: 2ddbe6646b ]
2021-01-14 19:28:01 -05:00
Wenkai Du
2c49121171
Port alltoall[v]
...
[ROCm/rccl commit: f4d5d3d620 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
41bead5a4e
Do not allow GPU as intermediate
...
[ROCm/rccl commit: 105db19a11 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
34c6013299
Revert "Changes to topology based on XGMI ( #272 )"
...
This reverts commit 0a9adc16f4 .
[ROCm/rccl commit: e055229e56 ]
2021-01-14 19:28:01 -05:00
Wenkai Du
adff98765c
Merge remote-tracking branch 'nccl/master' into no-target-id
...
[ROCm/rccl commit: d469947641 ]
2021-01-14 19:27:53 -05:00
Jonas Zhou
1db601566f
x86: Add CPU detection for Zhaoxin processors
...
Signed-off-by: Jonas Zhou <JonasZhou@zhaoxin.com >
[ROCm/rccl commit: 3996562690 ]
2020-12-17 11:15:18 -08:00
Wenkai Du
4ea285c527
Fix Rome PCIe 2 node topology generation ( #310 )
...
[ROCm/rccl commit: 373a108516 ]
2020-12-15 17:16:17 -08:00
Wenkai Du
b68ff1ebba
Add Rome model and improve search ( #305 )
...
[ROCm/rccl commit: 975b14dffa ]
2020-11-17 14:55:06 -08:00
Sylvain Jeaugey
a8908b34ee
2.8.3-1
...
Optimization for Tree allreduce on A100.
Improve aggregation performance.
Use shared buffers for inter-node send/recv.
Add NVTX profiling hooks.
Accelerate alltoall connections by merging communication for all
channels.
Add support for one hop communication through NVLink, for faster
send/recv communication on cubemesh topologies like DGX-1.
Improve alltoall scheduling to better balance intra/inter node
communication.
Increase send/recv parallelism by 8x, each warp sending or
receiving to a different peer.
Net: move to v4.
Net: make flush operation asynchronous to accelerate alltoall.
Net: define maximum number of requests.
Fix hang when using LL128 protocol after 2^31 steps.
Fix #379 : topology injection failing when using less GPUs than
described in the XML.
Fix #394 : protocol mismatch causing hangs or crashes when using
one GPU per node.
[ROCm/rccl commit: 920dbe5b35 ]
2020-11-17 11:08:52 -08:00
Wenkai Du
f19cbc8e51
Use device's link width and speed if port doesn't report ( #304 )
...
[ROCm/rccl commit: 554729079d ]
2020-11-13 17:58:04 -08:00
Stanley Tsang
f373cd2fdc
Fixing IPC handle leak ( #302 )
...
[ROCm/rccl commit: 2958f7eace ]
2020-11-13 10:32:42 -07:00
gilbertlee-amd
f66d05193a
Adding RCCL_CLIQUE_DEBUG to help debug experimental clique feature ( #300 )
...
[ROCm/rccl commit: c8d08a7c2f ]
2020-11-13 09:07:11 -07:00
Wenkai Du
62d21047b8
Skip unused peer connection in scatter and gather ( #301 )
...
[ROCm/rccl commit: 4e68229c8b ]
2020-11-12 15:47:34 -08:00
gilbertlee-amd
a7ef699687
Clique kernel support ( #295 )
...
* Adding experimental clique-based kernels (opt-in only)
Co-authored-by: Stanley Tsang <stanley.tsang@amd.com >
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com >
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com >
[ROCm/rccl commit: 41bcfb8878 ]
2020-11-10 15:44:10 -07:00