Commit Graph

505 Commits

Author SHA1 Message Date
Stanley Tsang 6680e23c63 Message queue refactor to POSIX implementation and leak fix (#355)
* Fixing message queue leak.

* Using POSIX implementation of Message Queues

* Adding unlink to msgqueue

* MsgQueue update

* Adding timeout check to msgqueue broadcast; tightening up system checks

* Removing unnecessary code

* Removing extra argument from print

* Adding explicit msg queue close call to all other ranks

[ROCm/rccl commit: 70597789d0]
2021-04-23 11:33:20 -06:00
Wenkai Du e28bac31aa Tune number of channels for gfx90a (#349)
[ROCm/rccl commit: 415c7cd3d1]
2021-04-19 15:27:01 -07:00
Wenkai Du 951d89b12f Use correct WARP_SIZE for gfx1030 (#348)
[ROCm/rccl commit: 9c718ce6d6]
2021-04-14 14:09:52 -07:00
Wenkai Du 661b1351a3 Limit max channels for ring graph on single node Rome (#347)
* Limit max channels for ring graph on single node Rome
* Partially revert "Use non-temporal access for streaming data (#341)"

[ROCm/rccl commit: a79f74082e]
2021-04-14 10:14:54 -07:00
Wenkai Du 0f4d497edc Add gfx90a target (#344)
* Add gfx90a target

* Support gfx90a topology

Co-authored-by: Eiden Yoshida <eiden.yoshida@amd.com>

[ROCm/rccl commit: 1fe031402a]
2021-04-14 09:29:00 -06:00
Sylvain Jeaugey 20da390b96 2.9.6-1
Add support for CUDA graphs.
Fuse BCM Gen4 switches to avoid suboptimal performance on some platforms. Issue #439.
Fix bootstrap issue caused by connection reordering.
Fix CPU locking block.
Improve CollNet algorithm.
Improve performance on DGX A100 for communicators with only one GPU per node.


[ROCm/rccl commit: a46ea10583]
2021-04-12 16:00:46 -07:00
TomSang 6105af2dfc Add detection of cooperative multi device launch attribute (#345)
[ROCm/rccl commit: 87f12cbb86]
2021-04-11 13:29:24 -07:00
Wenkai Du 6b3389b790 Use non-temporal access for streaming data (#341)
* Use non-temporal access for streaming data

* Revert to ulong2 after fixing compiling issue

[ROCm/rccl commit: 9dfc2c183e]
2021-04-07 17:34:35 -07:00
gilbertlee-amd a7e99734e8 Fixing clique-topology detection (#342)
* Fixing clique-topology detection
* Fix to enable multi-process clique-based kernels

[ROCm/rccl commit: caba0a63d2]
2021-04-07 11:29:44 -06:00
Wenkai Du b4a7fa7011 Cleanup number of channels calculation (#340)
[ROCm/rccl commit: e26ad2995e]
2021-04-05 17:51:56 -07:00
Wenkai Du 8927d8bf17 Fix incorrect net counting (#339)
* Fix incorrect net counting

* Add comments

[ROCm/rccl commit: 17491c918e]
2021-04-05 12:21:57 -07:00
Wenkai Du 018a31877c Rework network port trimming code (#338)
* Rework network port trimming code

* Move Rome related changes to separate source files

[ROCm/rccl commit: 1d2946ee4b]
2021-03-31 10:25:59 -07:00
Wenkai Du 4da9c54d4e Check fine grained memory on peer GPU before enabling P2P (#337)
[ROCm/rccl commit: 0c78553ee0]
2021-03-30 09:06:39 -07:00
Wenkai Du 065bde98d8 collnet: support multiple NICs (#335)
[ROCm/rccl commit: d87dc7c2e8]
2021-03-25 20:59:32 -07:00
Stanley Tsang 016a11f61b Fixing message queue leak. (#331)
[ROCm/rccl commit: 289db2a636]
2021-03-25 19:11:43 -06:00
Wenkai Du 07548845a6 Remove HDP workaround for ROCm 4.2 HIP (#334)
[ROCm/rccl commit: 0fbb9510a5]
2021-03-23 20:11:37 -07:00
Wenkai Du 287ed0f18a Enable collnet in RCCL (#333)
* Enable CollNet and use different number of channels

* topo_expl: enable collnet

[ROCm/rccl commit: 1d6244b18d]
2021-03-19 12:58:13 -07:00
Wenkai Du 052119c20a Sort GPUs by HIP device ID (#329)
* Sort GPUs by HIP device ID

* Remove extra space

[ROCm/rccl commit: b46260260a]
2021-03-16 16:51:32 -07:00
Wenkai Du 7374c512d9 Add GPU memory usage tracker (#326)
[ROCm/rccl commit: f60b76c67a]
2021-03-06 20:32:30 -08:00
Wenkai Du b7253710ca Revert "Port alltoall[v]" (#325)
This reverts commit 2c49121171.

[ROCm/rccl commit: 8e180cf087]
2021-03-06 13:59:31 -08:00
Wenkai Du bcf4ecb0e3 Enable local sendrecv over network if GDR is available on all GPUs (#324)
[ROCm/rccl commit: c018edf0f2]
2021-03-05 19:59:41 -08:00
gilbertlee-amd 8d90821062 Adding pthread_join / pthread_detach to clean up pthreads to avoid leaks (#322)
[ROCm/rccl commit: f4a9b9acba]
2021-02-26 16:29:55 -07:00
Wenkai Du 020ccb6730 Update tuning parameters for XGMI and NET
[ROCm/rccl commit: e820a943e9]
2021-02-23 21:41:26 +00:00
Wenkai Du d24df90882 Match NBIO only when GPUs and NICs are directly connected to CPU
[ROCm/rccl commit: ec8d89b1dd]
2021-02-22 18:52:29 -05:00
Stanley Tsang ce02076d64 Fixing cache deletion for CliqueManager; updating copyright
[ROCm/rccl commit: 45f5255f7c]
2021-02-19 22:22:46 +00:00
Wenkai Du 6c3ccc2192 Add support to another Rome model
[ROCm/rccl commit: 95f178324c]
2021-02-18 02:00:31 +00:00
Wenkai Du ab71643c99 Merge remote-tracking branch 'nccl/master' into 2.8.3
[ROCm/rccl commit: c985358e11]
2021-02-15 18:44:47 -05:00
Wenkai Du 34d7735419 Move HDP flush to CPU
[ROCm/rccl commit: bf8eb40705]
2021-02-12 18:06:19 +00:00
Sylvain Jeaugey fc7bdb38a5 2.8.4-1
Fix hang in corner cases of alltoallv using point to point send/recv.
Harmonize error messages.
Fix missing NVTX section in the license.
Update README.


[ROCm/rccl commit: 911d61f214]
2021-02-09 15:36:48 -08:00
Wenkai Du 602fe60428 Fix GDRDMA read and remove unused files
[ROCm/rccl commit: 9cc3b56166]
2021-02-09 01:34:39 +00:00
Stanley Tsang f152c8d160 Update MP UT to support arbitrary # of GPUs; multiple bugfixes (#16)
* Fixing temp file creation/deletion for Clique kernel mode.

* Refactoring of MP unit tests; include bugfixes and general support for any number of GPUs

* GroupCall MP UT properly quits when too many devices specified

* MP UT will programmatically set NCCL_COMM_ID if not specified; updated install script

[ROCm/rccl commit: d00b7d17bd]
2021-02-05 16:49:25 -08:00
Wenkai Du ae5779702a Merge remote-tracking branch 'origin/develop' into 2.8.3
[ROCm/rccl commit: ab1e7a0318]
2021-02-04 20:02:34 -05:00
gilbertlee-amd 16d625ca27 Tuning some clique-based kernel parameters (#315)
[ROCm/rccl commit: 1990ffd76a]
2021-02-03 20:00:08 -07:00
Wenkai Du 57abf599b2 Enable GPU direct RDMA read from GPU
[ROCm/rccl commit: 5f97122442]
2021-02-03 02:48:30 +00:00
gilbertlee-amd c981e76efe Clique kernel support (#295) (#15)
* Adding experimental clique-based kernels (opt-in only)

Co-authored-by: Stanley Tsang <stanley.tsang@amd.com>
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com>
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com>

Co-authored-by: Stanley Tsang <stanley.tsang@amd.com>
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com>

[ROCm/rccl commit: 3e62ceddc5]
2021-01-28 09:45:01 -07:00
Wenkai Du 7f9c15b843 Use less unroll for clique kernels (#313)
[ROCm/rccl commit: 41e47a36e7]
2021-01-15 17:48:10 -08:00
Wenkai Du d4382de267 Improve collective trace
[ROCm/rccl commit: 2ddbe6646b]
2021-01-14 19:28:01 -05:00
Wenkai Du 2c49121171 Port alltoall[v]
[ROCm/rccl commit: f4d5d3d620]
2021-01-14 19:28:01 -05:00
Wenkai Du 41bead5a4e Do not allow GPU as intermediate
[ROCm/rccl commit: 105db19a11]
2021-01-14 19:28:01 -05:00
Wenkai Du 34c6013299 Revert "Changes to topology based on XGMI (#272)"
This reverts commit 0a9adc16f4.


[ROCm/rccl commit: e055229e56]
2021-01-14 19:28:01 -05:00
Wenkai Du adff98765c Merge remote-tracking branch 'nccl/master' into no-target-id
[ROCm/rccl commit: d469947641]
2021-01-14 19:27:53 -05:00
Jonas Zhou 1db601566f x86: Add CPU detection for Zhaoxin processors
Signed-off-by: Jonas Zhou <JonasZhou@zhaoxin.com>


[ROCm/rccl commit: 3996562690]
2020-12-17 11:15:18 -08:00
Wenkai Du 4ea285c527 Fix Rome PCIe 2 node topology generation (#310)
[ROCm/rccl commit: 373a108516]
2020-12-15 17:16:17 -08:00
Wenkai Du b68ff1ebba Add Rome model and improve search (#305)
[ROCm/rccl commit: 975b14dffa]
2020-11-17 14:55:06 -08:00
Sylvain Jeaugey a8908b34ee 2.8.3-1
Optimization for Tree allreduce on A100.
Improve aggregation performance.
Use shared buffers for inter-node send/recv.
Add NVTX profiling hooks.
Accelerate alltoall connections by merging communication for all
channels.
Add support for one hop communication through NVLink, for faster
send/recv communication on cubemesh topologies like DGX-1.
Improve alltoall scheduling to better balance intra/inter node
communication.
Increase send/recv parallelism by 8x, each warp sending or
receiving to a different peer.
Net: move to v4.
Net: make flush operation asynchronous to accelerate alltoall.
Net: define maximum number of requests.
Fix hang when using LL128 protocol after 2^31 steps.
Fix #379 : topology injection failing when using less GPUs than
described in the XML.
Fix #394 : protocol mismatch causing hangs or crashes when using
one GPU per node.


[ROCm/rccl commit: 920dbe5b35]
2020-11-17 11:08:52 -08:00
Wenkai Du f19cbc8e51 Use device's link width and speed if port doesn't report (#304)
[ROCm/rccl commit: 554729079d]
2020-11-13 17:58:04 -08:00
Stanley Tsang f373cd2fdc Fixing IPC handle leak (#302)
[ROCm/rccl commit: 2958f7eace]
2020-11-13 10:32:42 -07:00
gilbertlee-amd f66d05193a Adding RCCL_CLIQUE_DEBUG to help debug experimental clique feature (#300)
[ROCm/rccl commit: c8d08a7c2f]
2020-11-13 09:07:11 -07:00
Wenkai Du 62d21047b8 Skip unused peer connection in scatter and gather (#301)
[ROCm/rccl commit: 4e68229c8b]
2020-11-12 15:47:34 -08:00
gilbertlee-amd a7ef699687 Clique kernel support (#295)
* Adding experimental clique-based kernels (opt-in only)

Co-authored-by: Stanley Tsang <stanley.tsang@amd.com>
Co-authored-by: Gilbert Lee <gilbert.lee@amd.com>
Co-authored-by: Wenkai Du <43822138+wenkaidu@users.noreply.github.com>

[ROCm/rccl commit: 41bcfb8878]
2020-11-10 15:44:10 -07:00