Wenkai Du
74f9e5db64
Add new GPU model ( #1080 )
2024-02-23 12:19:42 -08:00
Bertan Dogancay
2fb12a9358
Merge pull request #1079 from BertanDogancay/2.19.4-sync
...
2.19.4 Sync
2024-02-16 09:50:11 -07:00
akolliasAMD
bac57421c7
Allow bus id to be null ( #1085 )
...
* Allow bus id to be null
2024-02-15 16:36:51 -07:00
Wenkai Du
d999d9ad21
Merge remote-tracking branch 'rccl/develop' into 2.19.4
2024-02-09 11:31:03 -06:00
Wenkai Du
5669b0d7b6
2.18.5 fix ( #1077 )
...
* Revert "Revert "2.18.5-1""
This reverts commit 767fde8210 .
* Fix initial net device value
2024-02-09 09:18:38 -08:00
Wenkai Du
704c9ef0d1
Doubling P2P channels per peer on single node gfx94x only ( #1074 )
2024-02-07 14:05:57 -08:00
Wenkai Du
1d989f6524
Doubling P2P channels per peer on single node only ( #1069 )
2024-02-02 12:41:00 -08:00
Wenkai Du
1a134b283b
Merge remote-tracking branch 'rccl/develop' into 2.19.4
2024-01-31 11:53:10 -06:00
BertanDogancay
9ff53eeeae
Merge remote-tracking branch 'nccl/master' into develop
2024-01-30 14:43:43 -08:00
Pedram Alizadeh
ccfb35fa6d
modifying the tuning table to improve the performance of allreduce for 8MB and 16MB for single-node MI300X ( #1063 )
2024-01-26 09:05:53 -05:00
Wenkai Du
ffde530af5
Increase P2P channels per peer ( #1060 )
2024-01-25 11:21:58 -08:00
BertanDogancay
81ddf9de89
Merge remote-tracking branch 'nccl/v2.19' into develop
2024-01-24 15:25:33 -08:00
Wenkai Du
3325f96c56
Only use full MAXCHANNELS for gfx94x ( #1050 )
2024-01-17 09:00:49 -08:00
Pedram Alizadeh
b08124c85d
adding rccl tuning parameters for MI300X gfx942 with 8 GPUs single and multi-node ( #1047 )
2024-01-16 13:44:32 -05:00
Wenkai Du
261707d90a
Add option to force enable network transport on single node ( #1046 )
2024-01-16 07:54:18 -08:00
PedramAlizadeh
767fde8210
Revert "2.18.5-1"
...
This reverts commit 559b70f86c .
2024-01-12 16:54:19 +00:00
Hossein Pourreza
735178c1fe
cover more gpu/nic mapping cases ( #1037 )
2024-01-10 08:01:37 -08:00
PedramAlizadeh
0d515f9388
resolved conflicts, fixed the localNetCount/0 bug
2023-12-18 08:11:34 +00:00
Ziyue Yang
655742a3a6
Fully disable MSCCL when machine is not matched ( #1017 )
...
* Disable MSCCL algorithm meta loading when machine is not matched
* fully disable init
* fix potential segfault
2023-12-13 08:36:21 -08:00
Wenkai Du
53d807a5b9
msccl: disable on multi-node ( #1018 )
2023-12-13 07:41:40 -08:00
Wenkai Du
4ba65d1d6a
Increase max channles to 64 ( #993 )
2023-12-01 16:01:11 -08:00
pradeep-ramanna
0b53f79196
Fix GPU to NIC mapping for peertopeer ( #994 )
2023-12-01 08:00:17 -08:00
Ziyue Yang
4bb0b4a380
Move MSCCL algorithm loading to initialization to workaround HIP graph conflict ( #982 )
...
* MSCCL: pre-specify channels and pre-load algorithms
* add mutex
* fix bug
* clean include
* disable all-gathers temporarily
2023-11-30 09:47:20 -08:00
Wenkai Du
50b2dd9fd7
Add special handling of gfx940 ( #976 )
...
* Add special handling of gfx940
* Update ring base
2023-11-22 15:07:36 -08:00
Sylvain Jeaugey
b6d7438d31
Merge remote-tracking branch 'origin/master'
2023-11-20 05:07:23 -08:00
Alexander Grund
cece6415b0
Fix use of CPUID overwriting registers in use.
...
CPUID writes to EAX, EBX, ECX, and EDX so the inline-asm must state that.
Otherwise currently in-use register might get overwritten which may
cause all kinds of failures like segfaults or wrong results.
Alternatively `__cpuid` can be used which avoids this and related issues.
So do that as suggested in the GCC issue https://gcc.gnu.org/bugzilla/show_bug.cgi?id=112513
2023-11-14 12:38:02 +01:00
Sylvain Jeaugey
88d44d777f
2.19.4-1
...
Split transport connect phase into multiple steps to avoid port
exhaustion when connecting alltoall at large scale. Defaults to 128
peers per round.
Fix memory leaks on CUDA graph capture.
Fix alltoallv crash on self-sendrecv.
Make topology detection more deterministic when PCI speeds are not
available (fix issue #1020 ).
Properly close shared memory in NVLS resources.
Revert proxy detach after 5 seconds.
Add option to print progress during transport connect.
Add option to set NCCL_DEBUG to INFO on first WARN.
2023-11-13 10:36:12 -08:00
Wenkai Du
dbcba2923b
Use parallel init of LDS and adjust P2P channels for gfx94x ( #943 )
...
* Use parallel init of LDS and adjust P2P channels for gfx94x
* Move another init to parallel
* Fix NCCL_NCHANNELS_PER_PEER setting
2023-11-03 16:06:49 -07:00
Nilesh M Negi
1e5ca6820b
Fix gcnArchName bug in topology dump ( #937 )
...
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com >
2023-10-28 12:30:36 -05:00
Nilesh M Negi
f22df90e5c
remove gcnArch support ( #920 )
...
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com >
2023-10-26 12:09:15 -05:00
Wenkai Du
c4e65fd382
Add missing gfx942 support ( #927 )
2023-10-23 12:04:37 -07:00
Wenkai Du
4278a9918b
Update rome models ( #922 )
2023-10-18 17:28:01 -07:00
Bertan Dogancay
a6ff4618c7
Revert "Remove 2H4P condition from P2P channels adjustment ( #890 )" ( #904 )
...
This reverts commit 16dd05a58a .
2023-10-04 09:46:11 -06:00
Sylvain Jeaugey
8c6c595185
2.19.3-1
...
H800/H100 fixes and tuning.
Re-enable intra-process direct pointer buffer access when CUMEM is
enabled.
2023-09-26 05:57:15 -07:00
Sylvain Jeaugey
f9c3dc251e
2.19.1-1
...
Add local user buffer registration for NVLink SHARP.
Add tuning plugin support.
Increase net API to v7 to allow for device-side packet reordering;
remove support for v4 plugins.
Add support for RoCE ECE.
Add support for C2C links.
Better detect SHM allocation failures to avoid crash with Bus Error.
Fix missing thread unlocks in bootstrap (Fixes #936 ).
Disable network flush by default on H100.
Move device code from src/collectives/device to src/device.
2023-09-26 05:50:33 -07:00
akolliasAMD
b85d73c02e
changed the form that RCCL_TREE uses ( #888 )
...
* changed the form that RCCL_TREE uses
2023-09-15 15:01:33 -06:00
Wenkai Du
16dd05a58a
Remove 2H4P condition from P2P channels adjustment ( #890 )
2023-09-13 12:54:21 -07:00
Ziyue Yang
c1bfd5f0d8
Add single-node MI300X topology ( #889 )
2023-09-13 11:07:17 -07:00
Audrey MP
e58ec78d35
Gcn arch name ( #886 )
...
We use CMake to determine if we're compiling against a version of ROCm that supports gcnArchName and handles architecture checking appropriately. It includes a few helper functions as drop ins for the functionality we used gcnArch for before; sometimes to enable flags, and sometimes to set frequencies.
2023-09-12 15:34:40 -04:00
Wenkai Du
aeca1af374
Add MSCCL xml files ( #861 )
2023-08-23 14:12:34 -07:00
Sylvain Jeaugey
559b70f86c
2.18.5-1
...
Fix NVLS search (issue #931 ).
Increase max IB NICs to 32.
Fix inconsistent device ordering (issue #820 ).
Try to use different devices for different GPUs in systems with
more than one NIC per GFU.
2023-08-23 06:32:36 -07:00
Wenkai Du
6a0a6a37d9
Use relaxed atomics for LL on GFX11 ( #859 )
2023-08-21 16:28:39 -07:00
akolliasAMD
d33cd5a233
NCCL_TREES variable and rome model fixes ( #856 )
2023-08-21 10:35:37 -06:00
Wenkai Du
7044599575
Add new model support ( #847 )
...
* Add new model support
* Update new rings
2023-08-10 17:14:51 -07:00
Wenkai Du
8e58b65873
gfx11xx: disable LL protocol to workaround mtype issue ( #840 )
2023-08-04 07:53:07 -07:00
Sylvain Jeaugey
8ed014bae9
Fix inter-node NVLS graph search
...
We were passing a net ID instead of a gpu index, which could cause
crashes if those were unrelated (and they usually are).
Issue #931
2023-08-02 07:06:35 -07:00
Wenkai Du
3db371c9a5
Revert "Enable Ll128 on gfx90a ( #823 )" ( #829 )
...
This reverts commit 420f8af6a0 .
Also increase number of parallel jobs for linking
2023-07-27 20:25:18 -07:00
Wenkai Du
a7fcd58a97
Enable gfx94x ( #808 ) ( #816 )
...
(cherry picked from commit 94da229a7788d74685d1591a4e75a8341de64f41)
2023-07-21 07:31:27 -07:00
Wenkai Du
abd0615351
Merge remote-tracking branch 'nccl/master' into develop
2023-06-26 22:51:56 +00:00
Bertan Dogancay
0c77c66221
Disable Colltrace for --fast option ( #778 )
...
* Disable Colltrace for --fast option
* Limit nprocs for CI
2023-06-21 14:16:09 -06:00