Modifying ReduceOrCopyMulti to accept number of preOp source to support clique-based kernels

Αυτή η υποβολή περιλαμβάνεται σε:
Gilbert Lee
2021-08-11 11:00:34 -05:00
γονέας a4929465c5
υποβολή ae13d2a354
4 αρχεία άλλαξαν με 21 προσθήκες και 15 διαγραφές
@@ -66,7 +66,7 @@ __device__ void AllReduceCliqueSplitKernel(struct ncclWorkElem* args)
// Perform the reduction
#define ALL_REDUCE_CLIQUE_UNROLL 1
ReduceOrCopyMulti<ALL_REDUCE_CLIQUE_UNROLL, FUNC, T, NUM_RANKS, NUM_RANKS, NUM_RANKS, NUM_RANKS>(
threadIdx.x, blockDim.x, redOp, true, true, NUM_RANKS, srcs, NUM_RANKS, dsts, blockN);
threadIdx.x, blockDim.x, redOp, NUM_RANKS, true, NUM_RANKS, srcs, NUM_RANKS, dsts, blockN);
}
// Even if there was nothing for this GPU to do, it must participate in a barrier