rocm-systems

작성자	SHA1	메시지	날짜
Nusrat Islam	ef442f8f92	set MAXCHANNELS to 128	2024-06-03 13:05:05 -05:00
Nusrat Islam	506f16c506	add 256 channels support	2024-06-03 13:03:18 -05:00
Wenkai Du	73221b4230	Add ring simple chunk size tuning (#1180 ) * Add ring simple chunk size tuning * modifying the tuning table to improve the performance of broadcast for 8MB to 32MB for single-node MI300X after ring simple chunk size tuning * modifying the tuning table to improve the performance of reduce for 1MB to 4MB for single-node MI300X after ring simple chunk size tuning --------- Co-authored-by: PedramAlizadeh <pmohamma@amd.com>	2024-05-29 07:59:47 -07:00
Wenkai Du	eeea3b693b	Report error when collective is not enabled in build (#1177 ) * Report error when collective is not enabled in build * Fix typo	2024-05-16 10:11:12 -07:00
Wenkai Du	ecafc1969c	Support WSL2 (#1173 )	2024-05-10 07:31:12 -07:00
Wenkai Du	cd6e840e0b	Add back tree simple chunk size tuning (#1157 )	2024-04-28 19:48:53 -07:00
Wenkai Du	9e0c9b4ed8	Replace __HIP_PLATFORM_HCC__ with __HIP_PLATFORM_AMD__ (#1154 )	2024-04-25 07:19:18 -07:00
BertanDogancay	e1a835910e	Merge remote-tracking branch 'nccl/master' into develop	2024-04-23 13:34:00 -07:00
mberenjk	428837ffe4	replacing rccl_bfloat16 with hip_bfloat16 (#1126 ) Co-authored-by: mberenjk <mberenjk@amd.com>	2024-04-11 11:30:37 -05:00
Andy li	6777e65c1d	Enable fp8 support (#1101 ) * initial checkin * resolve cr comments * resolve the build issue * fix the data correctless issue * update fp8 header file and update the unit test for fp8 support * remove fp16 from fp8 headers * fix ut issue and catch up the latest code from develop * udate according to cr comments * update ut according to cr comments * update num floats for each SumPostDiv from 4 to 6 * update fp8 header file name * fix the typo	2024-03-08 15:17:53 -08:00
Bertan Dogancay	b275ed0b56	LL128 check if all XGMI (#1089 )	2024-02-21 09:41:40 -07:00
Sylvain Jeaugey	b6475625fb	2.20.3-1 Add support for alternating rings, allow for cross-nic rings without cross-rail communication. Add support for user buffer registration for network send/recv. Optimize aggregated operations to better utilize all channels. Add flattening for BCM PCI gen5 switches. Add support for inter-node NVLink communication Add support for port fusion in NET/IB. Add support for ReduceScatter and AllGather using Collnet. Update net API to v8. Fix hang during A2A connection.	2024-02-13 04:22:38 -08:00
BertanDogancay	00fdb1ef51	Clean up	2024-01-31 17:27:15 -08:00
Wenkai Du	1a134b283b	Merge remote-tracking branch 'rccl/develop' into 2.19.4	2024-01-31 11:53:10 -06:00
BertanDogancay	9ff53eeeae	Merge remote-tracking branch 'nccl/master' into develop	2024-01-30 14:43:43 -08:00
Bertan Dogancay	01b359027b	Include common.h in enqueue.cc instead (#1067 )	2024-01-30 08:24:22 -08:00
BertanDogancay	81ddf9de89	Merge remote-tracking branch 'nccl/v2.19' into develop	2024-01-24 15:25:33 -08:00
Wenkai Du	7e25d5bc55	Use new HIP graph API compatible with CUDA 11030 (#991 ) * Use new HIP graph API compatible with CUDA 11030 * Update dependency to ROCm 6.1 * Fix single stream use case	2024-01-21 19:00:50 -08:00
Bertan Dogancay	28d9b170c9	[DEV] Configure functions in RCCL (#986 ) * configure functions in rccl	2024-01-18 15:07:16 -07:00
Sylvain Jeaugey	88d44d777f	2.19.4-1 Split transport connect phase into multiple steps to avoid port exhaustion when connecting alltoall at large scale. Defaults to 128 peers per round. Fix memory leaks on CUDA graph capture. Fix alltoallv crash on self-sendrecv. Make topology detection more deterministic when PCI speeds are not available (fix issue #1020). Properly close shared memory in NVLS resources. Revert proxy detach after 5 seconds. Add option to print progress during transport connect. Add option to set NCCL_DEBUG to INFO on first WARN.	2023-11-13 10:36:12 -08:00
Sylvain Jeaugey	8c6c595185	2.19.3-1 H800/H100 fixes and tuning. Re-enable intra-process direct pointer buffer access when CUMEM is enabled.	2023-09-26 05:57:15 -07:00
Sylvain Jeaugey	f9c3dc251e	2.19.1-1 Add local user buffer registration for NVLink SHARP. Add tuning plugin support. Increase net API to v7 to allow for device-side packet reordering; remove support for v4 plugins. Add support for RoCE ECE. Add support for C2C links. Better detect SHM allocation failures to avoid crash with Bus Error. Fix missing thread unlocks in bootstrap (Fixes #936). Disable network flush by default on H100. Move device code from src/collectives/device to src/device.	2023-09-26 05:50:33 -07:00
Wenkai Du	6a0a6a37d9	Use relaxed atomics for LL on GFX11 (#859 )	2023-08-21 16:28:39 -07:00
akolliasAMD	d33cd5a233	NCCL_TREES variable and rome model fixes (#856 )	2023-08-21 10:35:37 -06:00
Wenkai Du	f70e3e569b	gfx11: don't use LL for sendrecv (#853 ) * gfx11: don't use LL for sendrecv * Use builtin instead of inline asm	2023-08-17 08:50:51 -07:00
Bertan Dogancay	8bab4f04b7	Implement RCCL Replayer (#817 ) * Implement RCCL Replayer	2023-07-24 16:26:22 -06:00
Wenkai Du	abd0615351	Merge remote-tracking branch 'nccl/master' into develop	2023-06-26 22:51:56 +00:00
Bertan Dogancay	0c77c66221	Disable Colltrace for --fast option (#778 ) * Disable Colltrace for --fast option * Limit nprocs for CI	2023-06-21 14:16:09 -06:00
Sylvain Jeaugey	ea38312273	2.18.3-1 Fix data corruption with Tree/LL128 on systems with 1GPU:1NIC. Fix hang with Collnet on bfloat16 on systems with less than one NIC per GPU. Fix long initialization time. Fix data corruption with Collnet when mixing multi-process and multi-GPU per process. Fix crash when shared memory creation fails. Fix Avg operation with Collnet/Chain. Fix performance of alltoall at scale with more than one NIC per GPU. Fix performance for DGX H800. Fix race condition in connection progress causing a crash. Fix network flush with Collnet. Fix performance of aggregated allGather/reduceScatter operations. Fix PXN operation when CUDA_VISIBLE_DEVICES is set. Fix NVTX3 compilation issues on Debian 10.	2023-06-14 01:29:17 -07:00
Wenkai Du	53a1f91857	Merge remote-tracking branch 'nccl/master' into develop	2023-04-25 15:38:32 -07:00
Sylvain Jeaugey	d97a32fac8	2.18.1-1 Add support for IB SHARP to NVLS (NVLink SHARP algorithm). Add NVLS+Tree algorithm. Add support for memory management using cuMem* functions. Use all NICs for Send/Receive operations on systems with more than one NIC per GPU (#804). Add ncclCommSplit primitive, with resource sharing option in config. Fix alltoallv hang (#788) Increase number of channels on H100 when we're not limited by NVLink. Improve error reporting in case of IB failure, printing local and remote ID (#779). Add build option to allow compilation against RDMA includes instead of dynamically loading IB verbs symbols (#802). Fix context creation for progress thread (#803). NET/IB: add option to use multiple QPs in round-robin mode. Fix tree performance issue when NVB is disabled on HCM topologies.	2023-04-18 03:58:25 -07:00
Wenkai Du	b02fd04165	Fix unit test HIP graph error (#712 )	2023-03-20 15:34:09 -07:00
Sylvain Jeaugey	5d3ab08b69	2.17.1-1 Add new NVLS algorithm for allreduce using NVLink SHARP (intra-node only). Add new config options: cgaClusterSize, minCTAs, maxCTAs, netName. Enable LL128 when we use PXN to close rings. NVTX3 includes update. Fix crash when one CollNet (SHARP) rail fails to initialize.	2023-03-01 00:39:04 -08:00
Wenkai Du	86e7b71234	Fix P2P scheduling (#690 )	2023-02-21 07:49:54 -08:00
Wenkai Du	e1cb45ff22	Merge remote-tracking branch 'nccl/master' into HEAD	2023-02-04 01:44:43 +00:00
Sylvain Jeaugey	28189e2df8	2.16.2-1 Add support for CUDA 12.0, drop Kepler (sm_35). Support for H100 features. Make socket code more robust and protected. Solves #555. Improve performance on large CUDA graphs, reducing dependencies. Reduce inter-socket bandwidth on AMD CPUs to favor better paths. Various fixes to ncclCommAbort. Make service thread polling resistant to EINTR. Compile with profiling API by default. Extend NVTX instrumentation with call arguments.	2022-11-30 02:31:59 -08:00
Wenkai Du	562dd87036	Move hipify to cmake stage Add minimal ROCm/HIP version requirements for Graph support	2022-11-14 18:10:45 +00:00
Wenkai Du	9a077e6947	Merge remote-tracking branch 'nccl/master' into develop	2022-11-03 21:17:42 +00:00
Wenkai Du	72ef100050	Fix P2P scheduling	2022-10-31 08:54:34 -07:00
Sylvain Jeaugey	cb111f764a	2.15.5-1 Fix crash with CollnetChain on some node topologies Fix hang when interleaving the capture of different graphs Fix hang during init in multi-threaded mode Fix potential data corruption with LL128 protocol on unaligned buffers. Fix CPU usage during preconnect Fixes double-free in the error path for ncclCommInitAll Workaround hang on H100 with Ring/LL128 on 2 GPUs.	2022-10-25 00:55:55 -07:00
Wenkai Du	4f0e223db4	Merge remote-tracking branch 'nccl/master' into develop	2022-10-20 15:41:29 +00:00
Edgar Gabriel	e645b02cd8	introduce a hw topology aware bintree for hayabusa architecture.	2022-10-03 15:26:21 +00:00
Wenkai Du	021932b3c8	Add LL128 tuning (#630 )	2022-09-27 09:39:09 -07:00
Sylvain Jeaugey	da8152e57a	2.15.1-1 Add support for H100 (sm90). Make sure NCCL kernel honor user stream priorities.	2022-09-27 02:31:13 -07:00
Edgar Gabriel	8f3219dbd4	make binary tree work on 2.13.4	2022-09-15 00:01:54 +00:00
Edgar Gabriel	be935d7ce7	Merge branch 'develop' into 2.13.4	2022-09-13 17:19:04 -05:00
Edgar Gabriel	65e2ae20e5	add binary tree In addition, introduce the ability to have 2 trees at the same time. Only for allreduce at the moment.	2022-09-13 20:52:32 +00:00
Wenkai Du	a79d9e3586	Merge remote-tracking branch 'nccl/master' into develop	2022-09-09 16:05:38 +00:00
Wenkai Du	7bbce085cc	Enable LL128 protocol support (#605 ) * Enable LL128 protocol support * Use shared memory object directly when possible	2022-09-08 14:45:27 -07:00
gilbertlee-amd	47b2fc3a30	Adding opt-in hipGraph support for RCCL via RCCL_ENABLE_HIPGRAPH (#608 ) Adding opt-in hipGraph support via RCCL_ENABLE_HIPGRAPH	2022-09-06 10:29:46 -06:00

1 2 3

116 커밋