Add AllGather LL128 multi-node tuning and include LL cutoff points in tuning models (#1618)

* Enable LL/LL128 cutoff points in tuning models

* Initializing ll/ll128 model cutoffs for MI300

* Use RCCL_LL_LIMITS_UNDEFINED

---------

Co-authored-by: PedramAlizadeh <pmohamma@amd.com>
Bu işleme şunda yer alıyor:
Mustafa Abduljabbar
2025-04-02 16:26:23 -04:00
işlemeyi yapan: GitHub
ebeveyn aace4e27f8
işleme 4be06f04d8
4 değiştirilmiş dosya ile 68 ekleme ve 23 silme
+2 -1
Dosyayı Görüntüle
@@ -517,6 +517,7 @@ struct ncclComm {
float bandwidths[NCCL_NUM_FUNCTIONS][NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
float ringbdw[NCCL_NUM_FUNCTIONS][NCCL_NUM_PROTOCOLS];
int maxThreads[NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
uint64_t minMaxLLRange[NCCL_NUM_FUNCTIONS][NCCL_NUM_PROTOCOLS - 1][2];
/* This attribute can indicate the states of communicators and return code of
* asynchronous NCCL operations. */
@@ -599,7 +600,7 @@ struct ncclComm {
struct ncclKernelPlanner planner;
hipStream_t sideStream; // [RCCL] Cached non-captured stream
cudaMemPool_t memPool;
// Queue of events and associated callbacks for cleaning up asynchronous work.
// Using this is preferable to using CUDA host callbacks because host callbacks