* [rocm_regression] Return errors when HSA_NO_SCRATCH_RECLAIM=1 even for rocm >= 6.4.0
* [rocm_regression] Check firmware version
* [rocm_regression] Resolve review comments
* [rocm_regression] Move hsa env checking into init once func
* [rocm_regression] Prevent hot fix version in firmware
* [rocm_regression] Improve unit tests
* Support pipelining codegen and template specialization
* Support ReduceCopy pipelining for AllReduce, ReduceScatter, and Reduce (currently enabled for bfloat16)
* Remove need for FUNC_INDEX_TOTAL
* Add pipeline field to device function key construction logic
* Avoid unneeded codegen for LL/LL64 kernels
* Modify conditions and add pipeline dtypes env
* Optimize selection for both gfx942 and gfx950
* Increase pipeline bitfield width
* Use __forceinline__ for all device functions
* Realign reduceCopy with original form
* Add opt-out option to enable perf debugs
* Remove force-reduce-pipelining option from README
* Update CHANGELOG.md
---------
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* add direct allgather algorithm
* minor fix
* add debug print for memory allocation tracker
* add message size threshold for direct allgather
* scatter transfers across ranks
* update changelog
* minor fix
* Update CHANGELOG.md
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* enable direct AG when pxn is ON on MI300X or MI350
---------
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
- Updated ncclDevFuncId to use a hash-based lookup with std::unordered_map.
- Keys are now 64-bit integers, which pack coll, algo, proto, devRedOp, and type fields.
- Improved flexibility and maintainability by moving away from row-based indexing.
- Added error handling for missing keys in the hash map.
- Aligned key generation logic with generate.py and updated generate.py.
* change gfx950 algo/proto selection for multinode allreduce, allgather, reduceScatter
* gfx950 tuning: enable tuning for broadcast, allreduce starts LL128 earlier and switches to ring earlier, change LL128 start for allgather and reduceScatter
* lower LL128 threshold
* update reduceScatter LL128 min to match LL max for consistency
* enable multinode PXN and increase chunksize for gfx950
* change LL128 start to 128KB, adjust ring-start according to node-count
* disable code-path for fused-AR on LL128 for gfx950
* use LL128 starting from 1KB for multinode allgather on gfx950
* start LL128 earlier for multinode reduceScatter on gfx950
* start LL128 earlier for multinode broadcast on gfx950
* set multinode allreduce to start simple on 64MB for gfx950
* start LL128 from 1KB for multinode broadcast on gfx950
* setting multinode AR to use tree instead of ring at 16MB, 64MB, 128MB
* set multinode broadcast to use LL for up to 256KB depending on node-count for gfx950
* adjust algo for 32MB multinode allreduce on gfx950
* make 32MB tree LL128 for multinode AR on gfx950
* make sure ring is not picked on 2N allreduce on small sizes
* Added useAcc as a template parameter to address the 2% performance regression in allreduceWithBias
---------
Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com>
* Support fused all reduce and elementwise operations
Add additional "acc" parameter to RCCL Replayer logs
Add flag which indicates availability of new API
* Fix Recorder json parsing
* Remove unreachable code
* Remove extra acc pointer check
* .
* Revert "[DEVICE] Adding ability to choose unroll factor at runtime (#1734)"
This reverts commit 9d72be7b2f.
* Use noinline to reduce kernels linking time
* Don't use noinline for gfx942 and gfx950 to avoid perf regression
---------
Co-authored-by: AtlantaPepsi <timhu102@amd.com>
Co-authored-by: BertanDogancay <bertan.dogancay@gmail.com>
Boosts single node bfloat16 allreduce performance by up to 20% for some data sizes and provides gating with the RCCL_GFX942_CHEAP_FENCE_OFF environment variable
* Add support for extended fine-grained system memory pool
* Use hipHostRegisterUncached
* Add "sc0 sc1" flags for LL store on gfx950
* Update after HIP flag is changed to hipExtHostRegisterUncached
This feature tracks the proxy events and status of each send/recv op. ProxyTrace keeps a fixed number of active ops in host mem and dumps the status of each op when the program crashes or hangs.
* First version of new replayer, with comments on future TODOs
* plus minor fixes for UT
* Updated format of recorder, especially in binary department, according to replayer's need
* Reapply "[AG and RS channel tuning] Add thread work threshold to tuning models and precompute reg index in LL128 (#1641)"
This reverts commit 943ad6f7820739385a0b54e81f823d0df1dbf71c.
* Decreasing NCCL_LL128_SHMEM_ELEMS_PER_THREAD from 16 to 8
Symmetric memory API and symmetric kernels
* Redesign from the ground up, enabling major latency and bandwidth
improvements.
* Add new API calls to register user-allocated memory among communicator
ranks into a NCCL window: ncclCommWindowRegister() and
ncclCommWindowDeregister(). The calls currently support symmetric
registration for P2P and NVLS, and require VMM memory buffers (i.e.,
CUMEM must be operational).
* Implement specialized kernels taking advantage of symmetrically
registered memory, with performance gains expected particularly for
small to medium message sizes.
* The kernels support 32 bit floating point types and smaller, and sum as
the reduction operator, with no more than one collective operation per
group.
* Floating point summation is always done in fp32 accumulators (with the
exception of fp8 on NVLS, where it uses fp16 inside the switch). Thus,
the accuracy with fp8 and fp16 data types should be much improved.
* This initial implementation supports non-network communicators only (P2P
and NVLS transports).
* To explore this functionality users need to use the new memory
registration API calls with the NCCL_WIN_COLL_SYMMETRIC flag and all
ranks of a communicator must pass buffers at the same offset in the same
registration when invoking a collective NCCL operation.
Add support for DGX Spark.
Add support for DirectNIC (CX8) to the internal IB plugin.
Add a new ncclCommShrink() API call
* It is a non-collective call similar to ncclCommSplit(), which makes it
possible to exclude some (possibly unresponsive) ranks from the parent
communicator.
Add support for loading multiple network plugins
* This enables the creation of generic containers that can work across a
range of providers.
* Allow NCCL_NET_PLUGIN to accept a comma-separated list of plugins to
load.
NVLink SHARP (NVLS) improvements
* Implement NVLS+IB SHARP support for AllGather and ReduceScatter with
user buffer registration. This improves performance and reduces the
number of CTAs needed to achieve peak bandwidth.
* Gracefully fall back by default to other transports if NVLS
initialization fails (the old behavior of returning an error code from a
NCCL call can be preserved by setting NCCL_NVLS_ENABLE=1).
* Decrease the NVLS channel count to 24 on Blackwell systems with multiple
NVLink domains per communicator.
* Enable fine-tuning of NCCL behavior per communicator using new
"ncclConfig_t" members "collnetEnable", "CTAPolicy", and "nvlsCTAs".
Profiler improvements
* Extend the init function by adding communicator name, comm id (hash),
rank, number of ranks, number of nodes, and the NCCL log function to the
argument list. This makes the name and the comm id available to all
events in the communicator without explicitly passing them to each
individual event. Add the communicator id and rank to the profiler trace
filename. Now, the communicator name can be set via a new "ncclConfig_t"
member "commName".
* Improve the accuracy of the GPU kernel events by providing GPU-generated
timestamps for the start and stop of every NCCL operation.
* Harmonize proxy events, removing overlaps between ProxyOp and ProxyStep
states.
* Add support for network-defined event updates (through
"recordEventState").
* Report the correct number of channels used by every collective/p2p
operation (used to be set to nMaxChannels for collectives and absent for
p2ps).
* Fix the logic on proxyCtrl Idle/Active events (Issue #1162).
* Fix an issue where the network proxy profiler could lose track of an
event identifier (Issue #1682).
* Improve the backward compatibility with plugins older than v4.
* Ensure that the work counters are 0-initialized.
* Fix a potential race condition in the network profiler that could result
in an event being linked to a wrong parent.
MNNVL improvements
* Increase to 16 the number of NICs used to communicate between MNNVL
domains on GB200 systems, to optimize the performance of collective
operations.
* Add support for more complex MNNVL topologies with up to 32 NICs per
node.
* If the MNNVL fabric initialization was unsuccessful, NCCL will now fail
by default, so as to avoid inadvertently falling back to a potentially
much slower network transport. Such failures are typically due to a
misconfigured IMEX support on the system. To continue without MNNVL,
restart the job with NCCL_MNNVL_ENABLE=0.
* Fix a potential hang in alltoall-like communication patterns at a scale
of over 80 ranks.
* Make NCCL_P2P_DISABLE=1 imply NCCL_MNNVL_ENABLE=0 (so the latter no
longer needs to be specified on MNNVL systems).
* Fix an initialization failure when NCCL_TOPO_FILE is used on MNNVL
systems.
* Fix the graph search to exclude non-local NICs.
* Fix the SHM transport to use fabric handles on MNNVL systems.
NIC Fusion improvements
* Disable the creation of fused NICs for physical devices that haven't
been merged.
* Flatten multiple ports to a single PCI device within the internal IB
plugin and reparent dual-port NICs under the first PCI parent. If the
parent is not a PCI switch, PCI devices for fused NICs won't be
duplicated.
* Route traffic on GB200-CX8 systems through DirectNIC, not the host
interface.
Improve support for platforms with C2C connectivity (e.g., GB200)
* Enable GPUDirect RDMA for the NICs by default.
* Add support for P2C (PXN over C2C) and the LL128 protocol.
Extend NCCL fault tolerance in multithreaded scenarios
* Support the creation of multiple nonblocking communicators within a
single group and polling in parallel for the completion using multiple
threads (one per communicator).
Enable ncclImplicitOrderLaunch for CUDA 12.9+
* This can potentially speed up NCCL_IMPLICIT_LAUNCH_ORDER.
Improve the netSocket transport latency and control
* Provide finer control over the size of the socket send/receive buffers,
the task size, and the number of sockets that a single peer can open.
* Add support for the inlining of small messages behind the header when
using multiple sockets per connection.
Improve the readability of the CPU affinity in the debug output
* Print it as a range string rather than a bitmask.
Fix a potential race condition in graph execution
* A contention could arise when mixing graph and non-graph execution.
Improve PXN connection code
* Avoid duplicate and unused connections.
RAS fixes
* Fix a memory corruption at job termination time in case of a previously
failed initialization of a RAS socket connection.
* Fix a race condition leading to a crash when generating a RAS report
during communicator initialization (Issues #1669, #1718).
* Fix a potential race condition when gathering data for a RAS status
report.
Fix a potential memory corruption in ncclCommSplit()
* Memory could get corrupted when resource sharing was in use and the size
of the NVLink domain in the new communicator was smaller than in the old
one.
Fix asynchronous graph upload
* Fix a small memory leak.
* Fix oversychronization.
Add a check for out-of-memory conditions in ncclMemAlloc()
Clean up the NCCL socket code
* accept() will retry also if just reading the magic failed (Issue #1613).
* connect() will retry also if poll() did not return a POLLOUT event
(Issue #1618).
* Add error checking in a few instances (Issue #1539).
* Fix the loop condition in ncclFindInterfaceMatchSubnet() (Issue #1574).
* Clean up the debug output, downgrading WARN messages to INFO in
non-critical cases, and printing the peer's address where relevant.
Switch NCCL_DEBUG_FILE to line buffering
* This should help avoid mixed-up partial output lines in multithreaded
cases.
Other minor fixes
* Improve the checks for buffer overflows in the graph code (Issue #1585).
* Extend logging and state clearing to all four events in the internal IB
plugin (Issue #1650).
* Fix the error path in case IB communication is not ready (Issue #1489).
* Add ECE logging for IB fabric.
* Fix various minor issues in the graph module (Issue #1635).
* Clean up the debug output in the graph code, downgrading WARN messages
to INFO in non-critical cases.
* Add a missing argument to a directSend() call (Issue #1628).
* Remove duplicate code in sendProxySetup() (Issue #1420).
* Fix the order of arguments of cudaDeviceCanAccessPeer() (Issue #1507).
* Fix compiler warnings with GCC 14.
* Fix a typo in a comment (Issue #1236).
* Revert "Revert "replacing rccl_float8 with hip_fp8 and address compatibility …"
This reverts commit 824b81c034.
* [UT] Modify max stack size to 496
* adding a check for OCP type and replacing ROCM_VERSION with HIP_VERSION
* addressing the ci failure
* Adding the device tag
---------
Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com>
* Update LL128 elems per thread
* Precompute ix[g] in LL128 prim
* Make Threadthreshold part of tuning models
* Ignore channel tuning when channels are env controlled
* Tune LL128 max limit for AG
* Tune LL128 max limit for RS
* Retune AR LL128 limits due to changes
* Update CHANGELOG.md
---------
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
This commit handles DMABUF initialization and call appropriate handling function. This fixes crash in OS with no peermem support and relying on only DMABUF.
* Initial test commit
* Handling Dmabuf_fd opening and closing
* Cleanup
* Use DMABuff or Peermem as needed
* Using user input for ibDmaBufSupportInitOnce
* Revert all changes to rocmwrap.cc
* Revert all changes to rocmwrap.cc
* Changing to func definition braces
* Reverting line removal in utils.h
* useDmaBuf to calculate flushEnabled
* Internal RCCL/NCCL functionality exposed when RCCL_EXPOSE_STATIC is enabled
* Algo/protocol/max channels can be obtained with the new RCCL API
* Introduce rccl_static and rccl_static_inline macros to work around invisible functions in core source files like enqueue.cc
* Add usage example in topo-explorer tool