Grafik Komit

857 Melakukan

Penulis SHA1 Pesan Tanggal
Wen-Heng (Jack) Chung 293f0fb752 Use a map to host scratch buffers (#1004)
* Use a map to host scratch buffers

* Address review feedbacks. Deliberately keep mscclSetupScratch function.
2023-12-05 13:15:28 -06:00
Nilesh M Negi bc44e3faa7 Fix gcnArch bug in IFC mix build (#998) (#1002)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>
2023-12-04 16:20:22 -06:00
Bertan Dogancay 7c0f49a878 IFC mix build (#998) 2023-12-02 18:49:52 -07:00
Wenkai Du 4ba65d1d6a Increase max channles to 64 (#993) 2023-12-01 16:01:11 -08:00
pradeep-ramanna 0b53f79196 Fix GPU to NIC mapping for peertopeer (#994) 2023-12-01 08:00:17 -08:00
Ziyue Yang e44e112a17 Fix mscclAlgoHandle not initialized issue (#995) 2023-12-01 07:58:01 -08:00
Ziyue Yang 4bb0b4a380 Move MSCCL algorithm loading to initialization to workaround HIP graph conflict (#982)
* MSCCL: pre-specify channels and pre-load algorithms

* add mutex

* fix bug

* clean include

* disable all-gathers temporarily
2023-11-30 09:47:20 -08:00
akolliasAMD 56ce9ef05f recreated pr 914 to work with current develop branch (#979) 2023-11-28 16:33:47 -07:00
Wenkai Du 50b2dd9fd7 Add special handling of gfx940 (#976)
* Add special handling of gfx940

* Update ring base
2023-11-22 15:07:36 -08:00
Sylvain Jeaugey b6d7438d31 Merge remote-tracking branch 'origin/master' 2023-11-20 05:07:23 -08:00
Wenkai Du 569d3f7d59 msccl: allocate scratch as ext-scope fine-grained (#968) 2023-11-16 09:57:25 -06:00
Wenkai Du bc8661f092 Fix kernel command line warnings (#961)
* Fix kernel command line warnings

* Remove while loop
2023-11-15 18:01:12 -08:00
Ziyue Yang 7fc891bc8d Fix MSCCL work FIFO allocation with HIP graph enabled (#967) 2023-11-15 16:43:28 -08:00
Bertan Dogancay 198f14923b Check to support older ROCm versions (#963) 2023-11-15 12:36:31 -07:00
Ziyue Yang 7ae95db5b8 Optimize MSCCL all-gather algorithms for gfx942 (#964) 2023-11-15 08:18:59 -08:00
Ziyue Yang df128879a6 Optimize MSCCL reduce primitive switching for gfx942 (#962)
* Optimize reduce primitive switching for gfx942

* address comment
2023-11-15 08:18:44 -08:00
Alexander Grund cece6415b0 Fix use of CPUID overwriting registers in use.
CPUID writes to EAX, EBX, ECX, and EDX so the inline-asm must state that.
Otherwise currently in-use register might get overwritten which may
cause all kinds of failures like segfaults or wrong results.

Alternatively `__cpuid` can be used which avoids this and related issues.
So do that as suggested in the GCC issue https://gcc.gnu.org/bugzilla/show_bug.cgi?id=112513
2023-11-14 12:38:02 +01:00
Sylvain Jeaugey 88d44d777f 2.19.4-1
Split transport connect phase into multiple steps to avoid port
exhaustion when connecting alltoall at large scale. Defaults to 128
peers per round.
Fix memory leaks on CUDA graph capture.
Fix alltoallv crash on self-sendrecv.
Make topology detection more deterministic when PCI speeds are not
available (fix issue #1020).
Properly close shared memory in NVLS resources.
Revert proxy detach after 5 seconds.
Add option to print progress during transport connect.
Add option to set NCCL_DEBUG to INFO on first WARN.
2023-11-13 10:36:12 -08:00
Wenkai Du 5a800e00cd msccl: enable basic collective trace (#959)
To avoid increasing number of kernels, colltrace is only enabled with
RCCL_MSCCL_FORCE_FULLOPS=1
2023-11-08 20:14:28 -08:00
Wen-Heng (Jack) Chung efc42d9045 Use send instead of sendWithBarrier. (#727) 2023-11-07 13:47:24 -06:00
Nusrat Islam 022735d208 Merge pull request #950 from nusislam/msccl-red2
msccl: remove cases from numReduction switch statement
2023-11-04 02:48:03 -05:00
Wenkai Du dbcba2923b Use parallel init of LDS and adjust P2P channels for gfx94x (#943)
* Use parallel init of LDS and adjust P2P channels for gfx94x

* Move another init to parallel

* Fix NCCL_NCHANNELS_PER_PEER setting
2023-11-03 16:06:49 -07:00
Nusrat Islam f545b94d4b msccl: remove cases from numReduction switch statement 2023-11-03 16:56:51 -05:00
Wenkai Du bb84345943 msccl: use 32-bit LDS access and add RCCL_MSCCL_FORCE_FULLOPS (#953) 2023-11-03 10:38:02 -07:00
akolliasAMD 988efe605a MSCCL stream fix (#948) 2023-11-03 09:10:52 -06:00
Wenkai Du f484ff17b9 msccl: add templated kernel (#945)
* msccl: add templated kernel

* Use defines to improve code readability

* Fix kernel indexing and review feedback
2023-11-02 17:21:53 -07:00
Nusrat Islam 6b80a0d0d4 msccl: remove dereference of reduce args
It can be removed because the msccl kernel will never execute this code
according to the current msccl setup.
2023-11-02 13:20:00 -05:00
Wenkai Du a7400218a2 msccl: use atomic to set dependency flags (#941) 2023-10-31 14:46:57 -07:00
Wenkai Du a497722894 NPkit: misc fixes for MSCCL (#936)
* msccl: add xcc_id to timestamp sync

* NPKit: add timestamp for rrc operator

* NPKit: add timestamp for MSCCL init
2023-10-30 10:00:12 -07:00
Nilesh M Negi 1e5ca6820b Fix gcnArchName bug in topology dump (#937)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>
2023-10-28 12:30:36 -05:00
Ziyue Yang 4c117e5335 Fix MSCCL work FIFO out-of-bound issue (#935) 2023-10-27 11:24:52 -07:00
Nilesh M Negi 96ec3ffe2e SRC/INIT: fix typo for ENABLE_PROFILING (#934)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>
2023-10-26 23:52:46 -05:00
Nilesh M Negi f22df90e5c remove gcnArch support (#920)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>
2023-10-26 12:09:15 -05:00
Wenkai Du fb0eccb57b msccl: reduce debug output when using NCCL_DEBUG=INFO (#932) 2023-10-25 08:05:19 -07:00
Wenkai Du c4e65fd382 Add missing gfx942 support (#927) 2023-10-23 12:04:37 -07:00
Wenkai Du dbb5611a3a Remove LDS based software barriers from MSCCL (#923) 2023-10-19 16:39:41 -05:00
Wenkai Du 4278a9918b Update rome models (#922) 2023-10-18 17:28:01 -07:00
Wenkai Du 39812ce757 NPKit: add xcc_id field (#918) 2023-10-13 15:24:59 -07:00
Wenkai Du 1b80d041cb Fix incorrect arch name parsing (#916) 2023-10-13 10:01:11 -07:00
Wenkai Du 6d0b5c1e89 Port init_once fix from NCCL (#915) 2023-10-13 08:01:12 -07:00
Wen-Heng (Jack) Chung 7ee5c1c28b Change MSCCL kernel signature to allow kernel arguments be preloaded via SGPR (#911)
* Adding a script that will download/compile/run TransferBench/RCCL/UCX/RCCL-tests/RCCL-Unittests/hip-mpi-testsuite (#895)

Co-authored-by: Pedram Alizadeh <pmohamma@banff-pla-r27-05.pla.dcgpu>

* Only build gfx941

* demo

* fine tune malloc

* Fix merge errors

* Fix merge errors

* Disable parallel build

* Adopt --amdgpu-kernarg-preload-count

* Revert "Adding a script that will download/compile/run TransferBench/RCCL/UCX/RCCL-tests/RCCL-Unittests/hip-mpi-testsuite (#895)"

This reverts commit f5e252dddf02a41b4d1bc512f306f45f97166304.

* Revert CMake changes.

* NPKIT changes.

* Remove some license declarations.

* Address code review feedbacks on msccl_kernel_impl.h

* Update CMakeLists.txt

* Add CMake logic to check the existence of --amdgpu-kernarg-preload-count

* Fix NPKIT trace logic.

---------

Co-authored-by: Pedram Alizadeh <pmohamma@amd.com>
Co-authored-by: Pedram Alizadeh <pmohamma@banff-pla-r27-05.pla.dcgpu>
Co-authored-by: Ziyue Yang <ziyyang@microsoft.com>
2023-10-12 20:17:08 -05:00
Bertan Dogancay a6ff4618c7 Revert "Remove 2H4P condition from P2P channels adjustment (#890)" (#904)
This reverts commit 16dd05a58a.
2023-10-04 09:46:11 -06:00
akolliasAMD 28d7fe5629 Dma buf support optin (#905)
* dmaBufSupport Optin added on every part of the code that should invoke it
2023-10-03 03:17:48 -06:00
Bertan Dogancay c1f57a7041 Modify All-To-All doc (#896)
* Modify All-To-All doc

* Update nccl.h.in

* update unit-tests

---------

Co-authored-by: gilbertlee-amd <44450918+gilbertlee-amd@users.noreply.github.com>
2023-09-27 12:45:21 -04:00
Sylvain Jeaugey 8c6c595185 2.19.3-1
H800/H100 fixes and tuning.
Re-enable intra-process direct pointer buffer access when CUMEM is
enabled.
2023-09-26 05:57:15 -07:00
Sylvain Jeaugey 3435178b6c Merge remote-tracking branch 'origin/master' into v2.19 2023-09-26 05:55:56 -07:00
Sylvain Jeaugey f9c3dc251e 2.19.1-1
Add local user buffer registration for NVLink SHARP.
Add tuning plugin support.
Increase net API to v7 to allow for device-side packet reordering;
remove support for v4 plugins.
Add support for RoCE ECE.
Add support for C2C links.
Better detect SHM allocation failures to avoid crash with Bus Error.
Fix missing thread unlocks in bootstrap (Fixes #936).
Disable network flush by default on H100.
Move device code from src/collectives/device to src/device.
2023-09-26 05:50:33 -07:00
Kaiming Ouyang 4365458757 Fix cudaMemcpyAsync bug
We are trying to use the copy result of first cudaMemcpyAsync in the
second cudaMemcpyAsync without sync in between. This patch fixes it
by allocating a CPU side array to cache device side addr so that we
can avoid this consecutive cuda mem copy.

Fixes #957
2023-09-20 05:51:14 -07:00
akolliasAMD b85d73c02e changed the form that RCCL_TREE uses (#888)
* changed the form that RCCL_TREE uses
2023-09-15 15:01:33 -06:00
Wenkai Du 26e982d913 Reduce NPKit latency overhead in MSCCL kernel (#893)
* Reduce NPKit latency overhead in MSCCL kernel

* Fix build error without NPKit enable
2023-09-15 13:28:26 -07:00