Граф коммитов

756 Коммитов

Автор SHA1 Сообщение Дата
Wenkai Du 157cc5f6ba Add new Rome model (#1304)
* Add another rome model and override

* Fix bug

* Fix typo

* Add ring

* Update ring

* Fix model matching

* Clean up

* Clean up

* Reverse rings for NCCL_RINGS input

* Only reverse NCCL_RINGS for ring graph

* Fix mapping issue when using  NCCL_RINGS

* Add NCCL_RINGS_REMAP to handle inconsistant net names

[ROCm/rccl commit: 532b70afb6]
2024-08-23 08:45:43 +08:00
mberenjk 886b576722 adding all nccl apis to api_support to enable rccl tracing by rocprofv3 (#1297)
* adding all nccl apis to api_support to enable rccl tracing by rocprofv3

Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com>
Co-authored-by: Jonathan R. Madsen <jonathanrmadsen@gmail.com>

[ROCm/rccl commit: db840f024e]
2024-08-22 12:36:07 -05:00
Wenkai Du d5a33e7107 Fix gfx940 CPX mode (#1290)
[ROCm/rccl commit: d3171b51b7]
2024-08-16 08:46:06 +08:00
Wenkai Du e770054b7f Fix model matching with PXN enable (#1295)
[ROCm/rccl commit: eff56735b0]
2024-08-16 06:16:00 +08:00
akolliasAMD 38e189bb1e removed hcc mentions (#1291)
[ROCm/rccl commit: d6c317d6ae]
2024-08-14 15:04:13 -06:00
Pedram Alizadeh e40a0939f1 adding new tunning table for very large number of nodes (#1288)
[ROCm/rccl commit: a25ca9bb90]
2024-08-09 10:47:42 -04:00
Tim 9fdecceefb Adding core binding in info (#1212)
Signed-off-by: AtlantaPepsi <timhu102@amd.com>

[ROCm/rccl commit: 4200964202]
2024-08-08 11:36:24 -04:00
Richard Barnes 92d874be50 Remove unused but set variable from all_reduce.h (#1258)
Allows `-Wunused-but-set-variable` to pass

[ROCm/rccl commit: d09b152aa0]
2024-07-29 08:11:24 -07:00
Richard Barnes 3d208c8eb9 Remove unused but set variable from prims_ll128.h (#1257)
Allows `-Wunused-but-set-variable` to pass

[ROCm/rccl commit: 86a4ad6e8b]
2024-07-29 08:11:01 -07:00
Richard Barnes 780324296c Remove unused but set variable from prims_ll.h (#1256)
Allows `-Wunused-but-set-variable` to pass

[ROCm/rccl commit: 7ad432ee23]
2024-07-29 08:10:38 -07:00
akolliasAMD 37c44d531b gfx12 Disable ll protocol (#1268)
[ROCm/rccl commit: c246e25f8e]
2024-07-26 08:59:55 -06:00
corey-derochie-amd 94910b8f80 Fix bug where the first collective call was using MSCCL instead of MSCCL++ (#1260)
[ROCm/rccl commit: 69135976d6]
2024-07-22 15:46:47 -06:00
saurabhAMD 24e1ed5288 Adding performance collection feature in rccl_replayer, and updating MSCCL logging and replayer parsing (#1265)
* Adding performance collection feature in rccl_replayer, and updating MSCCL logging and replayer parsing

* Performance collection feature in rccl_replayer, and updating MSCCL logging and replayer parsing

[ROCm/rccl commit: cf311b71ee]
2024-07-22 10:21:29 -05:00
corey-derochie-amd f2b2372056 Only initialize MSCCL++ when runtime-enabled. (#1266)
[ROCm/rccl commit: b31b4082dd]
2024-07-22 00:41:31 -06:00
Nusrat Islam df63c9772f Enable CPX mode for MI300X (#1259)
* graph: enable cpx mode for MI300X

* graph: tune limits for cpx and cleanup

[ROCm/rccl commit: 6f331b0d43]
2024-07-19 11:30:37 -05:00
Wenkai Du 54e4899607 Template unroll for RCCL kernels (#1250)
* Template unroll for RCCL kernels

* Adding unroll template arg during CMake hipification

* Reduce linking parallel jobs to avoid OOM in CI

* Workaround issues with UT tests

SWDEV-469533: register spill fix is needed for mainline build
LWPCOMMLIBS-369: cannot enable 112 channels with 80 CUs
Use -parallel-jobs=8 for linking

* CI: do not use -j 16 when building

* CI: use -j 8 when building

* Only reduce parallel linking job for CI extended

* Restore original jenkins command. Change parallel linking jobs in cmake

* Disable MSCCLPP

---------

Co-authored-by: gilbertlee-amd <gilbert.lee@amd.com>

[ROCm/rccl commit: 89349f2ce4]
2024-07-19 08:15:59 -07:00
Nilesh M Negi 73e17b3e70 Consistent channel shuffling for MI300X multi-node (#1255)
* Revert "[GRAPH] Use channel shuffling only for IB systems (#1228)"

This reverts commit 8a7dd0e590.

* Revert "Revert "Changing channel stride for MI300X multinode (#1196)" (#1224)"

This reverts commit bb6fab3d8e.

[ROCm/rccl commit: a1ef217b32]
2024-07-18 10:18:09 -05:00
Nilesh M Negi 13134c6c64 [GRAPH] Disable MSCCL override of no. of channels (#1187)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

[ROCm/rccl commit: 67e867271f]
2024-07-15 10:45:21 -05:00
corey-derochie-amd da08e5ed1e Only enable MSCCL++ AllReduce for message sizes that are multiples 32 (#1253)
* Only enable MSCCL++ AllReduce for message sizes that are multiples of 32. MSCCL++ does not handle these other sizes.

* Sanitized MSCCL++ logging.

[ROCm/rccl commit: 9cbb3da224]
2024-07-12 17:04:23 -07:00
corey-derochie-amd b8542c2477 Integrated RCCL with MSCCL++ for small message sizes (#1231)
[ROCm/rccl commit: 6dc47eecd7]
2024-07-12 15:32:58 -06:00
Rahul Vaidya f60367f1c3 Improved version reporting in NCCL_DEBUG=VERSION (#1232)
* Improved version reporting in NCCL_DEBUG=VERSION.

Signed-off-by: rahulvaidya20 <ravaidya@amd.com>

* Version reporting changes

Signed-off-by: rahulvaidya20 <ravaidya@amd.com>

* Versioning changes: Initialized char arrays to null and fixed typo.

---------

Signed-off-by: rahulvaidya20 <ravaidya@amd.com>

[ROCm/rccl commit: c755b9cf93]
2024-07-12 08:14:29 -05:00
akolliasAMD 942555fd21 gfx12 initial enablement (#1219)
[ROCm/rccl commit: 63e4d76e23]
2024-07-10 13:32:09 -06:00
corey-derochie-amd 37bf54b8f8 Enable multi-threading for MSCCL (#1203)
MSCCL can now run in a multi-threaded configuration. To test in the unit tests, added the ENABLE_OPENMP compile definition flag and the --openmp-test-enable flag to the unit test build script. To activate, set the environment variables UT_MULTITHREADED=1 and UT_PROCESS_MASK=1. Set Jenkins to use this mode.

[ROCm/rccl commit: 0c36d571ea]
2024-07-04 09:34:38 -06:00
Wenkai Du b5bc883f61 Checking kernel header files only when missing sysfs entry (#1239)
[ROCm/rccl commit: 45f3fbc52f]
2024-07-03 15:53:15 -07:00
Nilesh M Negi 8a7dd0e590 [GRAPH] Use channel shuffling only for IB systems (#1228)
* [GRAPH] Use channel shuffling only for IB systems

Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

* [GRAPH] Define channels=48 for gfx94 RoCE systems

Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

* [GRAPH] Increase channels for RoCE gfx94 systems

Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

---------

Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

[ROCm/rccl commit: 5be3b713ef]
2024-07-02 12:20:40 -05:00
Wenkai Du b3e9d2d61b NPKit: separate time stamps for GPU access from different blocks (#1229)
To avoid races in memory access in GPU

[ROCm/rccl commit: 9d8f68b4ee]
2024-06-28 08:00:22 -07:00
Nusrat Islam c24b1ff14f graph: fix minNchannels for multi-node overwrite (#1230)
[ROCm/rccl commit: b09ea29d66]
2024-06-26 16:56:10 -05:00
Wenkai Du bb6fab3d8e Revert "Changing channel stride for MI300X multinode (#1196)" (#1224)
This reverts channel stride change in commit
a009e43f3c

[ROCm/rccl commit: ad31d93f3d]
2024-06-25 14:03:30 -07:00
saurabhAMD de7ea612d7 Unit Tests for testing channels (#1222)
[ROCm/rccl commit: e170f41ddd]
2024-06-25 10:10:10 -05:00
Nusrat Islam 3c7bbd3243 Merge pull request #1223 from nusislam/minNchannels-multinode
graph: fix minNchannels for multi-node

[ROCm/rccl commit: 9f2514e5c8]
2024-06-25 10:03:35 -05:00
Wenkai Du 2e6c26a36c Fix DMABUF support (#1218)
* Fix DMABUF support

* Reduce log output by moving dmabuf allocation details to TRACE

* Enable peer memory GDR support if ib_umem_get_peer is in kernel

[ROCm/rccl commit: 5d7078e383]
2024-06-25 08:00:15 -07:00
Nusrat Islam bf164da244 graph: fix minNchannels for multi-node
Multi-node rccl was not correctly setting the minNchannels value. This
PR fixes the bug.


[ROCm/rccl commit: 05df0f8cea]
2024-06-24 16:42:44 -05:00
Paul Emberson f7fb3392fb fix initOnceFunc setting incorrect result code (#1205)
Addresses DMA-BUF support check unexpectedly failing

[ROCm/rccl commit: 435756af02]
2024-06-07 16:47:19 -07:00
Nusrat Islam 7c029672a8 Merge pull request #1200 from nusislam/multi-node-256-fix
graph: fix multi-node channel count

[ROCm/rccl commit: 9660e2e2dc]
2024-06-07 14:34:20 -05:00
Nilesh M Negi 9b88e59ea4 Fix min_nchannels bug for gfx94* nranks=4 (#1202)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

[ROCm/rccl commit: d9661c17e6]
2024-06-07 14:31:28 -05:00
gilbertlee-amd efe851cfff Disabling NUMA maching for model 79 for some VM configs (#1204)
[ROCm/rccl commit: 9b94a1052f]
2024-06-06 17:15:04 -06:00
Nusrat Islam 3cbf67715f graph: restrict maxChannels to 64 for multi-node and RCCL_ENABLE_INTRANET=1
[ROCm/rccl commit: 526cce9bf4]
2024-06-06 10:58:41 -05:00
Nusrat Islam bf35178250 graph: fix multi-node minChannel count
[ROCm/rccl commit: 6ab20a7c6b]
2024-06-06 10:56:39 -05:00
Nusrat Islam 62cbb3bcdd set MIN_NCHANNEL limit to 64 for multi-node
[ROCm/rccl commit: 9746d8ca3f]
2024-06-03 13:05:05 -05:00
Nusrat Islam b34fd115a1 doubling debug buffer size with increased channels
[ROCm/rccl commit: 0634c5c8e1]
2024-06-03 13:05:05 -05:00
Nusrat Islam 48821ad0d7 set MAXCHANNELS to 128
[ROCm/rccl commit: ef442f8f92]
2024-06-03 13:05:05 -05:00
Nusrat Islam 2447161fb4 graph: restrict MAXCHANNELS for certain platforms
[ROCm/rccl commit: 9f654f6cf5]
2024-06-03 13:05:01 -05:00
Nusrat Islam ed2f96bc6a device: update the logic for channelId assignment
[ROCm/rccl commit: 48859a97b1]
2024-06-03 13:03:18 -05:00
Nusrat Islam b0b2aa1166 add 256 channels support
[ROCm/rccl commit: 506f16c506]
2024-06-03 13:03:18 -05:00
gilbertlee-amd a009e43f3c Changing channel stride for MI300X multinode (#1196)
* Shuffling MI300X multi-node channels
* Updating tree channel logic

[ROCm/rccl commit: 0948eecbba]
2024-06-03 10:00:55 -06:00
ClementLinCF 4f56aa5f8c Optimize NCHANNELS and MSCCL config for gfx942 80CUs (#1195)
* Optimize NCHANNELS and MSCCL config for gfx942 80CUs

Set appropriately for different NCCL_MIN_NCHANNELS and MSCCL config,
potentially improving communication perf on the MI300x 80CUs

* Delete tools/msccl-algorithms/allreduce_1step_mccl_8_2_16777216_LL.xml

* Change the factor of gfx94 and update msccl config

[ROCm/rccl commit: cab25f919e]
2024-06-01 07:07:46 -07:00
Nilesh M Negi 8518ef4afb [MSCCL]: Move scratch buffer debug msgs to TRACE (#1189)
Signed-off-by: nileshnegi <Nilesh.Negi@amd.com>

[ROCm/rccl commit: 1249a6c3fd]
2024-05-31 17:54:23 -05:00
gilbertlee-amd b35e2b8c4b Addressing possible out-of-bounds mem access during channel duplication (#1193)
[ROCm/rccl commit: 354e0b29a6]
2024-05-30 14:02:14 -06:00
Wenkai Du a724f1ebb7 Add ring simple chunk size tuning (#1180)
* Add ring simple chunk size tuning

* modifying the tuning table to improve the performance of broadcast for 8MB to 32MB for single-node MI300X after ring simple chunk size tuning

* modifying the tuning table to improve the performance of reduce for 1MB to 4MB for single-node MI300X after ring simple chunk size tuning

---------

Co-authored-by: PedramAlizadeh <pmohamma@amd.com>

[ROCm/rccl commit: 73221b4230]
2024-05-29 07:59:47 -07:00
Wenkai Du ae6c372406 Make WSL detection thread safe (#1178)
* Make WSL detection thread safe

* Change static to beginning

* Switch to use atomics

[ROCm/rccl commit: 8f099b1adb]
2024-05-28 17:23:50 -07:00