rocm-systems

Autor	SHA1	Zpráva	Datum
Nusrat Islam	eb347a0dd3	GDA support for alltoall via rocshmem integration (#2099 ) * ROCSHMEM linking/building to match MSCCL++ style * add rocSHMEM as a submodule * Move rocSHMEM submodule to ext-src/rocSHMEM * Adding submodule support proper, as well as a patch for rocshmem * Cleaning up INCLUDE_DIR vs INCLUDE_DIRS mixup * updating patch file * Pointing rocshmem submodule to edgars fixup patch * Adding IBVERBS link to the submodule build * More IBVERBS patching * pin rocshmem submodule to `b534423` * Adding IPC support in rocSHMEM build * updating rocshmem submodule to resolve CQ errors * Updating submodule to include recent a2a optimizations * invoke rocshmem alltoall from rccl * Updating submodule to CQ error number hang * Updating submodule to include a2a improvements and bug fixes * Updating submodule to point to Yiltan's fork and doorbell ring removal commit * Updating hash to correspond with submodule change * Updating to no-ctx wg call and updating submodule * copy-in/copy-out using multiples CUs * Updating rocSHMEM submodule to include doorbell improvs * updating gitmodule to point to upstream * code cleanup and adjust threashold * guard rocshmem a2a invocation * Only build with rocshmem when specified * code cleanup * address review comments * Removing debugging failure case Signed-off-by: Thomas Huber <thomas.huber@amd.com> * whitespace fix * Adding rocshmem compile guard * Removing unneccesary comment Signed-off-by: Thomas Huber <thomas.huber@amd.com> * remove commented lines * address review comments * cleanup --------- Signed-off-by: Thomas Huber <thomas.huber@amd.com> Co-authored-by: Thomas Huber <thomas.huber@amd.com> Co-authored-by: Nusrat Islam <nusislam@dell300x-ccs-aus-k12-27.cs-aus.dcgpu> Co-authored-by: Nusrat Islam <nusislam@dell300x-ccs-aus-k13-09.cs-aus.dcgpu> Co-authored-by: Islam <nusislam@amd.com> Co-authored-by: Nusrat Islam <nusislam@dell300x-ccs-aus-k13-03.cs-aus.dcgpu> [ROCm/rccl commit: `27648b0900`]	2026-01-09 14:04:54 -06:00
AbandiGa	7f7c8d14f6	Disable Bfloatf16 pipelining for reduction collectives for gfx950 (#2047 ) * disable bf16 reduce_copy pipelining for gfx950 * edit CHANGELOG * Combine unroll and pipeline local arch calculation into single function * fix multi-node error and disbale for gfx950 even if it's not a local build * removed has_gfx950 * disable pipelining for gfx950 in rcclSetPipelining --------- Co-authored-by: Ghadeer Alabandi <galaband@cv350-zts-gtu-h30-08.prov.gtu.zts.cpe.ice.amd.com> Co-authored-by: Ghadeer Alabandi <galaband@cv350-zts-gtu-h30-18.prov.gtu.zts.cpe.ice.amd.com> Co-authored-by: Ghadeer Alabandi <galaband@cv350-zts-gtu-h28a-08.prov.gtu.zts.cpe.ice.amd.com> [ROCm/rccl commit: `277b6e9bac`]	2025-11-13 14:55:09 -06:00
Bertan Dogancay	b955a7df40	[GEN/BUILD] Refactor generator script and reduce build time for old archs. (#2030 ) [ROCm/rccl commit: `b1e680adc0`]	2025-11-07 15:15:25 -05:00
Nilesh M Negi	dd625edf56	Revert "[GEN/BUILD] Refactor generate.py and reduce build time for older archs (#2006 )" (#2021 ) This reverts commit `40f3faead0`. [ROCm/rccl commit: `62ab7a22d7`]	2025-10-31 10:04:12 -05:00
Bertan Dogancay	40f3faead0	[GEN/BUILD] Refactor generate.py and reduce build time for older archs (#2006 ) [ROCm/rccl commit: `bed7cdf863`]	2025-10-30 11:45:53 -04:00
Nilesh M Negi	0aa56fb0a5	Fix ncclDevFuncId for AllReduceWithBias (#1980 ) [ROCm/rccl commit: `c35bc721ad`]	2025-10-17 09:28:57 -05:00
mberenjk	433251272b	fixing the ar_with_bias test issue when running rccl-tests (#1912 ) * fixing the AR_With_Bias issue when running rccl-tests [ROCm/rccl commit: `e738c03e39`]	2025-10-13 13:58:21 -07:00
Mustafa Abduljabbar	f37f290134	[Device] Add dynamic fetch/reduce pipelining for reduction collectives - Simple protocol (#1861 ) * Support pipelining codegen and template specialization * Support ReduceCopy pipelining for AllReduce, ReduceScatter, and Reduce (currently enabled for bfloat16) * Remove need for FUNC_INDEX_TOTAL * Add pipeline field to device function key construction logic * Avoid unneeded codegen for LL/LL64 kernels * Modify conditions and add pipeline dtypes env * Optimize selection for both gfx942 and gfx950 * Increase pipeline bitfield width * Use __forceinline__ for all device functions * Realign reduceCopy with original form * Add opt-out option to enable perf debugs * Remove force-reduce-pipelining option from README * Update CHANGELOG.md --------- Co-authored-by: Jeffrey Novotny <jnovotny@amd.com> [ROCm/rccl commit: `277747c199`]	2025-08-26 15:03:54 -04:00
Nilesh M Negi	fbe014c870	[BUILD] Populate host_table entries only for 1 unroll (#1871 ) [ROCm/rccl commit: `bf6660ee4e`]	2025-08-23 00:15:38 -05:00
Mustafa Abduljabbar	5025a9aab9	Have ncclDevFuncId use 64-Bit keyed map with field packing (#1857 ) - Updated ncclDevFuncId to use a hash-based lookup with std::unordered_map. - Keys are now 64-bit integers, which pack coll, algo, proto, devRedOp, and type fields. - Improved flexibility and maintainability by moving away from row-based indexing. - Added error handling for missing keys in the hash map. - Aligned key generation logic with generate.py and updated generate.py. [ROCm/rccl commit: `c1b3cd8911`]	2025-08-19 16:41:19 -04:00
mberenjk	c76a4492f1	Added useAcc as a template parameter to address the performance regression (#1856 ) * Added useAcc as a template parameter to address the 2% performance regression in allreduceWithBias --------- Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com> [ROCm/rccl commit: `c61152baa4`]	2025-08-14 15:58:54 -05:00
Nilesh M Negi	be810f10f3	[DEVICE] Add unroll=2 for gfx950 multi-node (#1824 ) [ROCm/rccl commit: `bd55f876e9`]	2025-07-31 02:35:26 -05:00
Wenkai Du	caff9764d3	Support fused all reduce and elementwise operations (#1729 ) * Support fused all reduce and elementwise operations Add additional "acc" parameter to RCCL Replayer logs Add flag which indicates availability of new API * Fix Recorder json parsing * Remove unreachable code * Remove extra acc pointer check * . * Revert "[DEVICE] Adding ability to choose unroll factor at runtime (#1734)" This reverts commit `4cadf3597c`. * Use noinline to reduce kernels linking time * Don't use noinline for gfx942 and gfx950 to avoid perf regression --------- Co-authored-by: AtlantaPepsi <timhu102@amd.com> Co-authored-by: BertanDogancay <bertan.dogancay@gmail.com> [ROCm/rccl commit: `9a4213356d`]	2025-07-23 09:04:17 -07:00
Bertan Dogancay	d4aafe31fa	[GEN] Fix typo in IFC code gen (#1796 ) [ROCm/rccl commit: `7158adb57f`]	2025-07-11 09:19:39 -04:00
Nilesh M Negi	ba31e4e846	[INIT] Fix fallback for unsupported user-specified runtime unroll factor (#1780 ) * [INIT] Fix fallback for unsupported user-specified runtime unroll factor * Add CollTrace guard * Move `commSetUnrollFactor()` to rccl_wrap.cc * Modify comments in the device-code generator script [ROCm/rccl commit: `2c099fe29a`]	2025-07-10 10:56:18 -05:00
Bertan Dogancay	471fc6bff2	[DEVICE] Enable PAT algo for RCCL 1ppn (#1756 ) * Enable PAT algo for RCCL 1ppn [ROCm/rccl commit: `e96c8473a1`]	2025-07-04 13:45:18 -04:00
Nilesh M Negi	4cadf3597c	[DEVICE] Adding ability to choose unroll factor at runtime (#1734 ) * Adding runtime unroll factor selection via RCCL_UNROLL_FACTOR * [BUILD] Add support for user-defined UNROLL for debugging * Update CHANGELOG.md * Fix COLLTRACE errors in CI * Add debug statements for unroll and resolve warnings * Incorporate UNROLL into ONLY_FUNCS for debugging --------- Signed-off-by: nileshnegi <Nilesh.Negi@amd.com> Co-authored-by: gilbertlee-amd <44450918+gilbertlee-amd@users.noreply.github.com> Co-authored-by: Jeffrey Novotny <jnovotny@amd.com> [ROCm/rccl commit: `9d72be7b2f`]	2025-06-11 00:07:59 -05:00
Nilesh M Negi	7abc3160e7	[BUILD] Enable LL128 on gfx950 (#1731 ) * [BUILD] Enable LL128 on gfx950 * Modify comment in src/rccl_wrap.cc * Update CHANGELOG Signed-off-by: nileshnegi <Nilesh.Negi@amd.com> Co-authored-by: corey-derochie-amd <161367113+corey-derochie-amd@users.noreply.github.com> [ROCm/rccl commit: `ef5b4ff630`]	2025-06-09 00:25:54 -05:00
Nilesh M Negi	19ed482121	Re-apply unroll=1 and 112 channels for gfx950 (#1706 ) * Reapply "[SRC] Enable unroll=1 for gfx950 (#1602)" (#1667) This reverts commit `a6972c0d09`. * Reapply "[GRAPH] Increase default nChannels to 112 for gfx950 (#1596)" (#1620) This reverts commit `1a2eca1756`. [ROCm/rccl commit: `12517a957e`]	2025-05-28 14:58:10 -05:00
Nilesh M Negi	a6972c0d09	Revert "[SRC] Enable unroll=1 for gfx950 (#1602 )" (#1667 ) * Revert "[SRC] Enable unroll=1 for gfx950 (#1602)" This reverts commit `210f90ae0f`. * Update Changelog --------- Signed-off-by: nileshnegi <Nilesh.Negi@amd.com> [ROCm/rccl commit: `329e13efff`]	2025-04-30 23:33:08 -05:00
BertanDogancay	d045d0ca23	Merge remote-tracking branch 'nccl/master' into develop [ROCm/rccl commit: `a6bf9bfc9e`]	2025-04-23 20:47:43 -07:00
Nilesh M Negi	210f90ae0f	[SRC] Enable unroll=1 for gfx950 (#1602 ) * [SRC] Enable unroll=1 for gfx950 * Fix typo from rebase in generate.py * Support for unroll=1 and gfx90a when building for all GPU targets --------- Signed-off-by: nileshnegi <Nilesh.Negi@amd.com> [ROCm/rccl commit: `307bc10781`]	2025-03-27 18:21:35 -05:00
Wenkai Du	0bd40f5a87	Enable LL128 on gfx942 (#1549 ) [ROCm/rccl commit: `245c2de909`]	2025-03-16 15:10:05 -07:00
Nilesh M Negi	a3f44b4a10	[BUILD] Fix generate.py for gfx950 (#1575 ) * [BUILD] Fix generate.py for gfx950 Signed-off-by: nilenegi <Nilesh.Negi@amd.com> * [BUILD] Cleaner way to check for gfx targets Signed-off-by: nilenegi <Nilesh.Negi@amd.com> --------- Signed-off-by: nilenegi <Nilesh.Negi@amd.com> [ROCm/rccl commit: `02b8e504f0`]	2025-02-26 18:43:37 -06:00
Bertan Dogancay	e171f59719	[BUILD] Fix unsupported arguments in generator (#1519 ) * Fix unsupported arguments in generator * Get ROCM_PATH as env variable [ROCm/rccl commit: `5804603632`]	2025-02-03 14:51:55 -05:00
Sylvain Jeaugey	db3bfd118f	2.24.3-1 Network user buffer support for collectives * Leverage user buffer registration to achieve zero-copy inter-node communications for Ring, NVLS and Collnet Add RAS subsystem * Create a RAS thread keeping track of all NCCL communicators. * Add a ncclras tool contacting the RAS thread and getting a report. Add fp8 support * Add support for e5m2 and e4m3 8-bit floating point operations. * Use Tree/PAT algorithms when possible for better numerical stability. Add NIC fusion * Add a NET API to ask the network plugin to fuse a set of interfaces together. * Fuse multiple NICs under the same PCI switch as a single, larger NIC. Socket connection failure retry * Retry in case of socket connection failure (unreachable host) * Avoid "Software caused connection abort" errors on retries QP connection failure retry * Retry in case of IB QP connection failure during ibv_modify_qp. NET API improvements * Allow plugins to force a flush in case data and completion ordering is not guaranteed. * Indicate when completion is not needed (e.g. for the LL128 protocol), allowing plugins to skip generating a completion. * Allow for full offload of allgather operations when using one GPU per node. NCCL_ALGO/NCCL_PROTO strict enforcement * Extend NCCL_ALGO/NCCL_PROTO syntax to be able to specify ALGO/PROTO filters for each collective operation. * Strictly enforce the ALGO/PROTO filters, no longer fall back on the ring algorithm when the filtering leaves no option and error out instead. Enable CUMEM host allocations * Use cumem functions for host memory allocation by default. Improved profiler plugin API * Avoid dependencies with NCCL includes. * Add information on whether the buffer is registered or not Adjust PAT tuning * Improve transition between PAT and ring at scale. Fix hangs when running with different CPU architectures * Detect when we use a mix of GPU architectures * Ensure Algo/Proto decisions are made based on that unified state. Fix FD leak in UDS * Fix a leak when mapping buffers intra-node with cumem IPCs. Fix crash when mixing buffer registration and graph buffer registration. * Separate local and graph registration to avoid crashes when we free buffers. Fix user buffer registration with dmabuf * Make ncclSend/ncclRecv communication with buffer registration functional on network plugins relying on dmabuf for buffer registration. Fix crash in IB code caused by uninitialized fields. Fix non-blocking ncclSend/ncclRecv * Fix case where ncclSend/ncclRecv would return ncclSuccess in non-blocking mode even though the operation was not enqueued onto the stream. * Issue #1495 Various compiler tweaks and fixes * PR #758 Fix typo in ncclTopoPrintGraph * Issue #1468 [ROCm/rccl commit: `6aae379278`]	2025-01-07 02:01:15 -08:00
Bertan Dogancay	df51937f09	Fix typo in ncclGetKernelIndex macro (#1424 ) [ROCm/rccl commit: `dfe4a3ed81`]	2024-11-18 10:40:05 -05:00
Bertan Dogancay	9041572e23	Template generic kernel for unroll factor (#1419 ) * Template generic kernel for unroll factor [ROCm/rccl commit: `cb175fb0b3`]	2024-11-12 18:27:29 -05:00
Bertan Dogancay	57710c1183	Dynamically select unroll factor to build for when targeting local arch (#1371 ) * Dynamically select unroll factor to build for when targeting local arch only [ROCm/rccl commit: `373f113524`]	2024-10-21 10:53:11 -04:00
Bertan Dogancay	974c13cd62	[BUILD] Move code generation to python from CMake (#1360 ) * Use generate.py for func generation * Convert AddUnroll.cmake to bash [ROCm/rccl commit: `2dd10c8f17`]	2024-10-03 10:21:19 -04:00
Sylvain Jeaugey	60240fec77	2.23.4-1 Add scalable init API * Add new ncclCommInitRankScalable to allow for passing multiple unique IDs to the init function. * Spreads the load onto multiple bootstrap roots, allowing for constant bootstrap time. * Requires multiple ranks to create a unique ID, and the CPU-side ID exchange code to call allgather[v] instead of broadcast. Accelerate init bootstrap operations * Reduce the number of calls to allgather. * Allow roots to reply early to ranks when information is already available. * Add an option to use ncclNet instead of sockets to perform bootstrap allgather operations. Add PAT algorithms for Allgather and ReduceScatter * Parallel Aggregated Trees, variation of Bruck algorithm. * Logarithmic number of network steps for small sizes at scale. * Only supports one rank per node at the moment. Add support for registered buffers for intra-node communication. * Allow registered user buffers to be accessed directly intra-node * Avoids extra copies in algorithms which permit it, saving memory bandwidth and helping with compute overlap. Add profiler plugin API * New plugin API for profiling * Supports various levels of profiling, with a hierarchy. Asynchronous graph allocation * Make calls to cudaMalloc and cudaMemcpy during graph allocation asynchronous. * Significantly speeds up graph capture. Use fatal IB asynchronous events to stop network operation * Avoids many other error messages * Only fatal errors are affected; potentially transient errors (e.g. port down) do not cause an immediate stop. Set P2P level to PXB on AMD CPUs when using more than 2 GPUs per node * P2P would cause a significant performance degradation when using many GPUs, and therefore many interleaved data flows. * Disable P2P through the CPU when we have 3+ GPUs per node; keep it enabled when we only have 2 GPUs. Improve the init logs to report the real NCCL function. * Make the log report ncclCommInitRank or ncclCommSplit, rather than the generic ncclCommInitRankFunc. Add a parameter to set the location of the user configuration file. * Add NCCL_CONF_FILE environment variable to set where the user's configuration file resides. Increase default IB timeout * Increase IB timeout value from 18 to 20. * Should help avoid fatal errors on large RoCE systems. Add new check for nvidia peermem * On linux kernels 6.6+, /sys/kernel/mm/memory_peers is no longer present; check for /sys/module/nvidia_peermem/version instead. Fix old performance regression when mixing small and large operations. * Improves distribution of work on channels. Fix crash when NUMA IDs are equal to -1. * Can happen when a NIC is a virtual NIC, or when linux doesn't know which NUMA node a device is attached to * Issue NVIDIA/nccl-tests#233 Fix tree graph search when NCCL_CROSS_NIC is set to 1. * Would force NCCL to use the balanced_tree pattern, thereby disabling LL128 on platforms with 1 GPU+1 NIC per PCI switch. * Would also try to use alternate rings even though it was not needed. Compiler tweaks and fixes * PR #1177 * PR #1228 Fix stack smash * PR #1325 Fixes for multi-node NVLink + IB operation Coverity fixes and comments. [ROCm/rccl commit: `68b542363f`]	2024-09-16 23:41:17 -07:00
Sylvain Jeaugey	5ca1b6c160	2.22.3-1 Rework core for NVIDIA Trusted Computing * Compress work structs so that they are shared between channels * Utilize the full amount of kernel argument space permitted (4k) before resorting to work fifo. * Rework the task preprocessing phase. * Use a separate abortDevFlag which is kept in sync with abortFlag using cudaMemcpy operations. * Rename src/include/align.h to src/include/bitops.h Add lazy connection establishment for collective operations * Move buffer allocation and connection establishment to the first collective operation using that algorithm. * Accelerate init time and reduce memory usage. * Avoid allocating NVLS buffers if all calls are registered. * Compute algo/proto in ncclLaunchCollTasksInfo early on. * Connect peers in ncclCollPreconnectFunc if not connected already. * Also move shared buffer creation to the first send/recv call. Accelerate intra-node NVLink detection * Make each rank only detect NVLinks attached to its GPU. * Fuse XMLs to reconstruct the full NVLink topology Add init profiling to report time spend in different init phases. * Report timings of bootstrap, allgather, search, connect, etc. * Add new "PROFILE" category for NCCL_DEBUG_SUBSYS. Add support for PCI p2p on split PCI switches * Detect split PCI switches through a kernel module exposing switch information. * Update the topology XML and graph to add those inter-switch connections. Add cost estimation API * Add a new ncclGroupEndSimulate primitive to return the estimated time a group would take. Net/IB: Add separate traffic class for fifo messages * Add NCCL_IB_FIFO_TC to control the traffic class of fifo messages independently from NCCL_IB_TC. Merges PR #1194 Net/IB: Add support for IB router * Use flid instead of lid if subnets do not match * Warn if flid is 0 Optimizations and fixes for device network offload (unpack) * Double the default number of channels * Cache netDeviceType * Fix save/increment head logic to enable Tree support. Support ncclGroupStart/End for ncclCommAbort/Destroy * Allow Abort/Destroy to be called within a group when managing multiple GPUs with a single process. Improve Tuner API * Provide to the plugin the original cost table so that the plugin can leave unknown or disabled algo/proto combinations untouched. * Remove nvlsSupport and collnetSupport. Do not print version to stdout when using a debug file * Also print version from all processes with INFO debug level. Fixes issue #1271 Fix clang warnings in NVTX headers * Update NVTX headers to the latest version Fixes issue #1270 Disable port fusion in heterogeneous systems * Do not fuse ports if a mix of multi-port and single port are detected. Fix NVLS graphs search for dual NICs. * Fix NVLS graph search when we have more than one NIC per GPU. Fix crash with collnetDirect * Add separate graph search for collnetDirect, testing alltoall paths and working similarly to the NVLS search. Fix hang when nodes have different CPU types * Add the CPU type to the rank peer info. * Align all ranks on the CPU type after the first allgather. * Only use the aligned CPU type for all tuning operations. Fixes issue #1136 Fixes issue #1184 Fix performance of registered send/recv operations * Allow for single full size operations * Add INFO to confirm the registration of send/recv buffers. Move all sync ops to finalize stage * Ensure ncclCommDestroy is non-blocking if ncclCommFinalize has been called. Improve error reporting during SHM segment creation Improve support of various compilers Merges PR #1177 Merges PR #1228 Allow net and tuner plugins to be statically linked * Search for ncclNet or ncclTuner symbols in the main binary. Merges PR #979 Plugin examples includes cleanup * Harmonize err.h and common.h usage. * Add mixed plugin with both net and tuner. [ROCm/rccl commit: `178b6b7590`]	2024-06-19 01:57:16 -07:00
Sylvain Jeaugey	2ab8a3a750	2.20.3-1 Add support for alternating rings, allow for cross-nic rings without cross-rail communication. Add support for user buffer registration for network send/recv. Optimize aggregated operations to better utilize all channels. Add flattening for BCM PCI gen5 switches. Add support for inter-node NVLink communication Add support for port fusion in NET/IB. Add support for ReduceScatter and AllGather using Collnet. Update net API to v8. Fix hang during A2A connection. [ROCm/rccl commit: `b6475625fb`]	2024-02-13 04:22:38 -08:00
Sylvain Jeaugey	506d6c332c	2.19.1-1 Add local user buffer registration for NVLink SHARP. Add tuning plugin support. Increase net API to v7 to allow for device-side packet reordering; remove support for v4 plugins. Add support for RoCE ECE. Add support for C2C links. Better detect SHM allocation failures to avoid crash with Bus Error. Fix missing thread unlocks in bootstrap (Fixes #936). Disable network flush by default on H100. Move device code from src/collectives/device to src/device. [ROCm/rccl commit: `f9c3dc251e`]	2023-09-26 05:50:33 -07:00

34 Commity