Nilesh M Negi
6f1b11ad49
Merge remote-tracking branch 'nccl-tests/master' into develop
2025-08-16 16:10:04 -04:00
Kajsa Arnold
a7809b3243
Standardize output formats ( #140 )
...
* remove spaces from csv
* consistently set redop to none when applicable
* write output file after test finishes
2025-07-30 17:28:04 -05:00
David Addison
fae7cb4727
Merge pull request #316 from martin-belanger/print-program-name
...
Print the name of the program being executed before and after test output
2025-07-24 14:58:54 -07:00
Bertan Dogancay
645be0eb45
[Common] Use NCCL API to allocate/free memory ( #144 )
2025-07-24 11:14:49 -04:00
David Addison
6edafa0a9c
Add extra reserved space during maxBytes calculation
...
Also, don't allow minBytes > maxBytes
2025-07-23 16:19:37 -07:00
BertanDogancay
50a26637fb
Merge remote-tracking branch 'nccl-tests/master' into develop
2025-07-23 14:23:22 -05:00
David Addison
e7c8825b0b
Wrap ncclCommWindowRegister() calls within ncclGroup
2025-06-03 10:36:53 -07:00
Martin Belanger
dafb70408d
Print the name of the program being executed
...
One thing missing from the stdout of each performance test is
the name of the test that is actually being run.
This patch adds 2 new messages to the stdout. At the beginning
of the execution of a test (e.g. sendrecv_perf) we will now
see this message:
Collective test starting: sendrecv_perf
And at the end, we will now see this:
Collective test concluded: sendrecv_perf
This is needed when running several tests consecutively and we're
trying to parse the stdout to collect the results.
For example, using a Python script to parse the stdout, one could
retrieve the results for each test and plot them on a graph. This
patch makes it easier to implement such a script.
Signed-off-by: Martin Belanger <martin.belanger@dell.com >
2025-06-03 11:43:02 -04:00
David Addison
a5c539e68b
Add support for Symmetric Memory Registration
...
From NCCL 2.27.x we can now use the Symmetric Memory APIs (-R 2)
2025-05-30 17:31:34 -07:00
mberenjk
4b2b635766
Switched to using the hip_fp8 header instead of rccl_float8, resolving compatibility issues.( #109 )
...
* addressing hip_fp8 support compatibility issue
* skipping mulsum and avg test for fp8, using hip_fp8 for product
* syncing with nccl-tests
removing the fp8 filter for pre-hopper gpus and resolving the merge conflict
---------
Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com >
2025-05-14 15:30:07 -05:00
Wenkai Du
cac33a8c2f
Automatically set in-place option from out-of-place ( #123 )
2025-05-09 16:48:42 -05:00
Grant Pinkert
f611dbd49a
Fix message size logging ( #115 )
...
Previously, the logger was logging the number of expected bytes a node was to recieve.
This differs from the stdout logging, where the reported message size is the total size of a message.
Signed-off-by: Grant Pinkert <gpinkert@amd.com >
2025-04-25 11:05:21 -05:00
nileshnegi
5625599dda
Merge remote-tracking branch 'nccl-tests/master' into develop
2025-04-21 19:46:10 -05:00
David Addison
501a149d57
Add support for FP8 datatypes
...
Added new datatypes: f8e4m3, f8e5m2
Only supported on H100+ architectures and NCCL versions >= 2.24.0
2025-04-18 19:20:59 -07:00
David Addison
b4300cc79d
Add PCI domain and device ID for GPU device BDF display
2025-02-28 13:25:51 -08:00
Junyu Ma
a89cf07fe8
Perftests: Introduce NCCL_TESTS_SPLIT env
...
`NCCL_TESTS_SPLIT` serves as new way of computing the color for splitting communicators.
Will be overrided by `NCCL_TESTS_SPLIT_MASK`.
Examples:
NCCL_TESTS_SPLIT_MASK="0x7" # color = rank & 0x7. What we do today to run on a DGX with one GPU per node.
NCCL_TESTS_SPLIT="AND 0x7" # color = rank & 0x7. New way to run on one GPU per node on a DGX, equivalent to NCCL_TESTS_SPLIT_MASK=0x7
NCCL_TESTS_SPLIT="MOD 72" # color = rank % 72. One GPU per NVLink domain on an NVL72 system.
NCCL_TESTS_SPLIT="DIV 72" # color = rank / 72. Intra NVLink domain on NVL72.
You can also use: "%" "&" "|" "/" for short.
Extra spaces in the middle will be automatically ignored.
Not case sensitive.
The followings are all equivalent:
NCCL_TESTS_SPLIT="%0x7"
NCCL_TESTS_SPLIT="%0b111"
NCCL_TESTS_SPLIT="AND 7"
NCCL_TESTS_SPLIT="and 0x7"
2025-02-04 15:18:09 -08:00
David Sidler
959cc19920
Add option to output results to a file ( #93 )
...
* Use find_package for MPI
* Add functionality to output results to file
* fix compilation
* report num gpus
* Revert "Use find_package for MPI"
This reverts commit c8fa253724ef4d0beac0d9c72f968062fbc6908e.
* Change inplace key
* remove dependency on json library
* Print "ranks, ranksPerNode, gpusPerRank"
* Add "nodes" field
---------
Co-authored-by: nileshnegi <Nilesh.Negi@amd.com >
2025-01-13 17:28:29 -06:00
saurabhAMD
fc9917e0da
Updating to use hipDeviceMallocUncached ( #95 )
...
Use hipDeviceMallocUncached instead of hipDeviceMallocFinegrained on newer ROCm versions.
2025-01-11 23:25:24 -06:00
Mustafa Abduljabbar
f2a48983ae
Memset to fix inflated performance when GPU is reset ( #94 )
...
* Memset to fix inflated performance when GPU is reset
* use hipMemset for both memsets
2025-01-11 23:25:24 -06:00
Tim
f7a5df7fc4
hot fixing ncclMemFree for mscclpp ( #100 )
2025-01-09 12:03:52 -05:00
John Bachan
29f4114f02
Fixes to all tests that divide buffers by nranks so that they trim buffer sizes to be multiples of 16 bytes.
...
This ensures non-pow2 ranks have buffer addresses aligned suitably for performance.
2024-12-18 11:20:28 -08:00
AtlantaPepsi
afd5ca10ae
Merge -R option for memory allocation
...
Signed-off-by: AtlantaPepsi <timhu102@amd.com >
2024-07-31 14:57:20 +00:00
David Addison
9d26b8422b
Merge pull request #226 from netgroup/master
...
improve parsing of stepbytes (increment size) argument
2024-07-30 14:58:54 -07:00
David Addison
d2d40cc824
Added -N,--run_cycles option
2024-07-25 22:00:23 -07:00
Rahul Vaidya
c5cae38bb8
Fix --root all issue. ( #83 )
...
Signed-off-by: rahulvaidya20 <ravaidya@amd.com >
2024-06-14 11:46:08 -05:00
Stefano Salsano
746549b28d
improve parsing of stepbytes (increment size) argument
2024-06-14 11:28:55 +02:00
Kaiming Ouyang
d028efcf35
Change ncclCommRegister size to maxBytes in serial comm init
2024-06-06 06:54:48 -07:00
saurabhAMD
36a2c372ac
Rotating tensor -R (default:off)
2024-06-04 11:35:39 -05:00
Giuseppe Congiu
a1efb427e7
Add -R option to register user buffers
2024-06-03 01:04:58 -07:00
saurabhAMD
74c4177f58
updating cache flush on functionality
2024-05-10 08:46:13 -07:00
saurabhAMD
699478dadf
Enable cache flush after every -F iteration. Default : 0 (No cache flush)
2024-05-07 11:32:30 -05:00
saurabhAMD
3c0728e8eb
Cache flush
2024-05-07 11:09:32 -05:00
Wenkai Du
16dfeaf89b
Fix incorrect device ordinal with limited device visibility ( #74 )
2024-05-02 11:14:57 -07:00
corey-derochie-amd
f74c04b686
Fixed spelling
2024-05-02 09:18:25 -06:00
Corey Derochie
0c762d210c
Wrapped the warmup iters in captures when doing graph mode to do a proper warmup.
2024-05-01 20:41:12 -05:00
mberenjk
eb65dadfc5
replacing rccl_bfloat16 with hip_bfloat16 ( #70 )
...
Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com >
2024-04-23 17:00:20 -05:00
mberenjk
3f7f7859bf
adding git version to rccl-tests ( #69 )
...
Co-authored-by: mberenjk <mberenjk@amd.com >
2024-03-28 14:03:59 -05:00
akolliasAMD
91609be0ef
Revert "adding git version to rccl-test ( #66 )"
...
This reverts commit a31679775c .
2024-03-22 10:21:37 -06:00
mberenjk
a31679775c
adding git version to rccl-test ( #66 )
...
* adding git version to rccl-test
---------
Co-authored-by: mberenjk <mberenjk@banff-cyxtera-s74-2.ctr.dcgpu >
2024-03-20 10:04:12 -05:00
Andy li
e447c17382
update the fp8 header file name ( #65 )
...
* update the fp8 header name
2024-03-08 10:02:40 -08:00
Andy li
21e59fb283
Enable fp8 support ( #63 )
...
* initial checkin
* rename the fp8 datatype name
* update based on cr comments
* resolve the build issue
* resolve fp8 campability issue
* fix minior bug and catch up to reflex latest develop branch change
* add fp8 + operatior support
* update fp8 header file
* resolve merge issue from develop branch
2024-03-07 16:54:41 -08:00
Bertan Dogancay
88cf7dbf45
Add hipify steps prior to build ( #62 )
...
* Add hipify steps prior to build
2024-03-05 09:47:18 -07:00
Wenkai Du
621dde544d
Merge remote-tracking branch 'nccl-tests/master' into HEAD
2024-03-01 18:34:44 +00:00
Wenkai Du
7715a0cf1f
Fix typo in rank assignment ( #59 )
2024-02-15 12:04:38 -08:00
David Addison
c6afef0b6f
Added missing MPI_Comm_free() call before MPI_Finalize()
2024-02-05 08:53:54 -08:00
Nusrat Islam
a2bec5d2f6
Add option to disable out-of-place
2024-01-04 16:43:50 -06:00
Wenkai Du
5ee7a08994
Warm up both out-of-place and in-place collectives ( #51 )
2023-10-16 12:13:50 -07:00
David Addison
1292b25553
Added an MPI_Barrier() call after MPI_Bcast() for HCOLL issue
2023-10-12 16:53:32 -07:00
David Addison
6c46206a47
Make the -c option be a datacheck iteration count parameter
...
Default is 1
2023-09-13 14:03:38 -07:00
Wenkai Du
652a24d38d
Fix merge error
2023-06-14 20:26:33 +00:00