Add new API for creating a reduction operation which multiplies the input by a rank-specific scalar before doing an inter-rank summation (see: ncclRedOpCreatePreMulSum).
Improve CollNet (SHARP) performance of ncclAllReduce when captured in a CUDA Graph via user buffer registration.
Add environment variable NCCL_NET_PLUGIN="<suffix>" to allow user to choose among multiple NCCL net plugins by substituting into "libnccl-net-<suffix>.so".
Fix memory leak of NVB connections.
Fix topology detection of IB Virtual Functions (SR-IOV).


[ROCm/rccl commit: e11238b302]
Этот коммит содержится в:
Ke Wen
2021-09-08 13:56:25 -07:00
родитель 04553b802a
Коммит e51f2a83da
41 изменённых файлов: 1532 добавлений и 599 удалений
+2
Просмотреть файл
@@ -9,6 +9,7 @@
#include "nccl.h"
#include "devcomm.h"
#include "collectives.h"
typedef enum {
ncclPatternRing,
@@ -38,6 +39,7 @@ struct ncclInfo {
int chunkSteps;
int sliceSteps;
// Computed later
ncclDevRedOpFull opFull;
int algorithm;
int protocol;
ncclPattern_t pattern;