Commit Graph

124 Commits

Author SHA1 Message Date
amd-jiali d5e8f372dc Fix Out of Memory issue when allocating bias buffer (#160)
* Add argument to select performance test with bias or not; if with bias, the maximum memory usage should be re-calculated and reduce the data size to avoid the Out of Memory issue; if without bias, no need to allocate buffers for bias

* Remove argument option for bias; memory calculation and buffer allocation are determined by the exec name.

---------

Co-authored-by: Li <jialili@ctr2-alola-ctrl-01.amd.com>

[ROCm/rccl-tests commit: 5272cd16ef]
2025-12-11 14:00:29 -08:00
Wenkai Du 75a69211a0 Add all_reduce_bias_perf to support All Reduce with Bias (#130)
Use dynamic symbol loading of ncclAllReduceWithBias

Co-authored-by: mberenjk <146776561+mberenjk@users.noreply.github.com>

[ROCm/rccl-tests commit: db6ea5a594]
2025-10-13 16:09:10 -05:00
Mustafa Abduljabbar cb4b286d2b Enable viewing algo/proto/channels used in rccl-tests output (#151)
* Enable algo/proto/channel viewing 

* Use dynamic symbol loading to avoid build/runtime issues with non-compatible RCCL versions

* Reduce code duplication

[ROCm/rccl-tests commit: 0c94d4d2b3]
2025-09-26 18:09:01 -04:00
Nilesh M Negi 15bf0f5fd1 Merge remote-tracking branch 'nccl-tests/master' into develop
[ROCm/rccl-tests commit: 6f1b11ad49]
2025-08-16 16:10:04 -04:00
Kajsa Arnold aed68678a4 Standardize output formats (#140)
* remove spaces from csv
* consistently set redop to none when applicable
* write output file after test finishes

[ROCm/rccl-tests commit: a7809b3243]
2025-07-30 17:28:04 -05:00
David Addison 33b74ad124 Merge pull request #316 from martin-belanger/print-program-name
Print the name of the program being executed before and after test output

[ROCm/rccl-tests commit: fae7cb4727]
2025-07-24 14:58:54 -07:00
Bertan Dogancay 7111d2dd99 [Common] Use NCCL API to allocate/free memory (#144)
[ROCm/rccl-tests commit: 645be0eb45]
2025-07-24 11:14:49 -04:00
David Addison 146ecc2212 Add extra reserved space during maxBytes calculation
Also, don't allow minBytes > maxBytes


[ROCm/rccl-tests commit: 6edafa0a9c]
2025-07-23 16:19:37 -07:00
BertanDogancay 0010193b64 Merge remote-tracking branch 'nccl-tests/master' into develop
[ROCm/rccl-tests commit: 50a26637fb]
2025-07-23 14:23:22 -05:00
David Addison 0ae7c8cbf4 Wrap ncclCommWindowRegister() calls within ncclGroup
[ROCm/rccl-tests commit: e7c8825b0b]
2025-06-03 10:36:53 -07:00
Martin Belanger ce1a83a0e8 Print the name of the program being executed
One thing missing from the stdout of each performance test is
the name of the test that is actually being run.

This patch adds 2 new messages to the stdout. At the beginning
of the execution of a test (e.g. sendrecv_perf) we will now
see this message:

  Collective test starting: sendrecv_perf

And at the end, we will now see this:

  Collective test concluded: sendrecv_perf

This is needed when running several tests consecutively and we're
trying to parse the stdout to collect the results.

For example, using a Python script to parse the stdout, one could
retrieve the results for each test and plot them on a graph. This
patch makes it easier to implement such a script.

Signed-off-by: Martin Belanger <martin.belanger@dell.com>


[ROCm/rccl-tests commit: dafb70408d]
2025-06-03 11:43:02 -04:00
David Addison 46e09f18c8 Add support for Symmetric Memory Registration
From NCCL 2.27.x we can now use the Symmetric Memory APIs (-R 2)


[ROCm/rccl-tests commit: a5c539e68b]
2025-05-30 17:31:34 -07:00
mberenjk ed6ebb12a7 Switched to using the hip_fp8 header instead of rccl_float8, resolving compatibility issues.(#109)
* addressing hip_fp8 support compatibility issue

* skipping mulsum and avg test for fp8, using hip_fp8 for product

* syncing with nccl-tests

removing the fp8 filter for pre-hopper gpus and resolving the merge conflict

---------

Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com>

[ROCm/rccl-tests commit: 4b2b635766]
2025-05-14 15:30:07 -05:00
Wenkai Du fe47d3dd77 Automatically set in-place option from out-of-place (#123)
[ROCm/rccl-tests commit: cac33a8c2f]
2025-05-09 16:48:42 -05:00
Grant Pinkert 3f962f5d58 Fix message size logging (#115)
Previously, the logger was logging the number of expected bytes a node was to recieve.
This differs from the stdout logging, where the reported message size is the total size of a message.

Signed-off-by: Grant Pinkert <gpinkert@amd.com>

[ROCm/rccl-tests commit: f611dbd49a]
2025-04-25 11:05:21 -05:00
nileshnegi 8d887aad0d Merge remote-tracking branch 'nccl-tests/master' into develop
[ROCm/rccl-tests commit: 5625599dda]
2025-04-21 19:46:10 -05:00
David Addison 8d71063e05 Add support for FP8 datatypes
Added new datatypes: f8e4m3, f8e5m2

Only supported on H100+ architectures and NCCL versions >= 2.24.0


[ROCm/rccl-tests commit: 501a149d57]
2025-04-18 19:20:59 -07:00
David Addison d516392fac Add PCI domain and device ID for GPU device BDF display
[ROCm/rccl-tests commit: b4300cc79d]
2025-02-28 13:25:51 -08:00
Junyu Ma e2a9cbb362 Perftests: Introduce NCCL_TESTS_SPLIT env
`NCCL_TESTS_SPLIT` serves as new way of computing the color for splitting communicators.

Will be overrided by `NCCL_TESTS_SPLIT_MASK`.

Examples:

NCCL_TESTS_SPLIT_MASK="0x7" # color = rank & 0x7. What we do today to run on a DGX with one GPU per node.
NCCL_TESTS_SPLIT="AND 0x7"  # color = rank & 0x7. New way to run on one GPU per node on a DGX, equivalent to NCCL_TESTS_SPLIT_MASK=0x7
NCCL_TESTS_SPLIT="MOD 72"   # color = rank % 72.  One GPU per NVLink domain on an NVL72 system.
NCCL_TESTS_SPLIT="DIV 72"   # color = rank / 72.  Intra NVLink domain on NVL72.

You can also use: "%" "&" "|" "/" for short.
Extra spaces in the middle will be automatically ignored.
Not case sensitive.

The followings are all equivalent:

NCCL_TESTS_SPLIT="%0x7"
NCCL_TESTS_SPLIT="%0b111"
NCCL_TESTS_SPLIT="AND 7"
NCCL_TESTS_SPLIT="and 0x7"


[ROCm/rccl-tests commit: a89cf07fe8]
2025-02-04 15:18:09 -08:00
David Sidler 93977b8866 Add option to output results to a file (#93)
* Use find_package for MPI
* Add functionality to output results to file
* fix compilation
* report num gpus
* Revert "Use find_package for MPI"
This reverts commit c8fa253724ef4d0beac0d9c72f968062fbc6908e.
* Change inplace key
* remove dependency on json library
* Print "ranks, ranksPerNode, gpusPerRank"
* Add "nodes" field
---------

Co-authored-by: nileshnegi <Nilesh.Negi@amd.com>

[ROCm/rccl-tests commit: 959cc19920]
2025-01-13 17:28:29 -06:00
saurabhAMD a2a3515e33 Updating to use hipDeviceMallocUncached (#95)
Use hipDeviceMallocUncached instead of hipDeviceMallocFinegrained on newer ROCm versions.

[ROCm/rccl-tests commit: fc9917e0da]
2025-01-11 23:25:24 -06:00
Mustafa Abduljabbar 7fbad10633 Memset to fix inflated performance when GPU is reset (#94)
* Memset to fix inflated performance when GPU is reset
* use hipMemset for both memsets

[ROCm/rccl-tests commit: f2a48983ae]
2025-01-11 23:25:24 -06:00
Tim 549f58f2a1 hot fixing ncclMemFree for mscclpp (#100)
[ROCm/rccl-tests commit: f7a5df7fc4]
2025-01-09 12:03:52 -05:00
John Bachan 69b9a05e71 Fixes to all tests that divide buffers by nranks so that they trim buffer sizes to be multiples of 16 bytes.
This ensures non-pow2 ranks have buffer addresses aligned suitably for performance.


[ROCm/rccl-tests commit: 29f4114f02]
2024-12-18 11:20:28 -08:00
AtlantaPepsi ccf92dcea2 Merge -R option for memory allocation
Signed-off-by: AtlantaPepsi <timhu102@amd.com>


[ROCm/rccl-tests commit: afd5ca10ae]
2024-07-31 14:57:20 +00:00
David Addison 58349a42b2 Merge pull request #226 from netgroup/master
improve parsing of stepbytes (increment size) argument

[ROCm/rccl-tests commit: 9d26b8422b]
2024-07-30 14:58:54 -07:00
David Addison cf3ffb2f5f Added -N,--run_cycles option
[ROCm/rccl-tests commit: d2d40cc824]
2024-07-25 22:00:23 -07:00
Rahul Vaidya 77524bffd0 Fix --root all issue. (#83)
Signed-off-by: rahulvaidya20 <ravaidya@amd.com>

[ROCm/rccl-tests commit: c5cae38bb8]
2024-06-14 11:46:08 -05:00
Stefano Salsano a0e06f2133 improve parsing of stepbytes (increment size) argument
[ROCm/rccl-tests commit: 746549b28d]
2024-06-14 11:28:55 +02:00
Kaiming Ouyang 1922bd71cb Change ncclCommRegister size to maxBytes in serial comm init
[ROCm/rccl-tests commit: d028efcf35]
2024-06-06 06:54:48 -07:00
saurabhAMD 351a047627 Rotating tensor -R (default:off)
[ROCm/rccl-tests commit: 36a2c372ac]
2024-06-04 11:35:39 -05:00
Giuseppe Congiu 411a7912f8 Add -R option to register user buffers
[ROCm/rccl-tests commit: a1efb427e7]
2024-06-03 01:04:58 -07:00
saurabhAMD 5700751b7e updating cache flush on functionality
[ROCm/rccl-tests commit: 74c4177f58]
2024-05-10 08:46:13 -07:00
saurabhAMD ce8e61cc3b Enable cache flush after every -F iteration. Default : 0 (No cache flush)
[ROCm/rccl-tests commit: 699478dadf]
2024-05-07 11:32:30 -05:00
saurabhAMD e1a3c5a6dc Cache flush
[ROCm/rccl-tests commit: 3c0728e8eb]
2024-05-07 11:09:32 -05:00
Wenkai Du baf9242e07 Fix incorrect device ordinal with limited device visibility (#74)
[ROCm/rccl-tests commit: 16dfeaf89b]
2024-05-02 11:14:57 -07:00
corey-derochie-amd fe151f517b Fixed spelling
[ROCm/rccl-tests commit: f74c04b686]
2024-05-02 09:18:25 -06:00
Corey Derochie 85bdda3812 Wrapped the warmup iters in captures when doing graph mode to do a proper warmup.
[ROCm/rccl-tests commit: 0c762d210c]
2024-05-01 20:41:12 -05:00
mberenjk 9cfb0745c0 replacing rccl_bfloat16 with hip_bfloat16 (#70)
Co-authored-by: Marzieh Berenjkoub <mberenjk@amd.com>

[ROCm/rccl-tests commit: eb65dadfc5]
2024-04-23 17:00:20 -05:00
mberenjk ca4ba933a3 adding git version to rccl-tests (#69)
Co-authored-by: mberenjk <mberenjk@amd.com>

[ROCm/rccl-tests commit: 3f7f7859bf]
2024-03-28 14:03:59 -05:00
akolliasAMD f8c62f64c3 Revert "adding git version to rccl-test (#66)"
This reverts commit 82c71f1838.


[ROCm/rccl-tests commit: 91609be0ef]
2024-03-22 10:21:37 -06:00
mberenjk 82c71f1838 adding git version to rccl-test (#66)
* adding git version to rccl-test

---------

Co-authored-by: mberenjk <mberenjk@banff-cyxtera-s74-2.ctr.dcgpu>

[ROCm/rccl-tests commit: a31679775c]
2024-03-20 10:04:12 -05:00
Andy li aaf1e27af2 update the fp8 header file name (#65)
* update the fp8 header name

[ROCm/rccl-tests commit: e447c17382]
2024-03-08 10:02:40 -08:00
Andy li c128f0422d Enable fp8 support (#63)
* initial checkin

* rename the fp8 datatype name

* update based on cr comments

* resolve the build issue

* resolve fp8 campability issue

* fix minior bug and catch up to reflex latest develop branch change

* add fp8 + operatior support

* update fp8 header file

* resolve merge issue from develop branch

[ROCm/rccl-tests commit: 21e59fb283]
2024-03-07 16:54:41 -08:00
Bertan Dogancay 882a96f5cb Add hipify steps prior to build (#62)
* Add hipify steps prior to build

[ROCm/rccl-tests commit: 88cf7dbf45]
2024-03-05 09:47:18 -07:00
Wenkai Du b49f6da1ec Merge remote-tracking branch 'nccl-tests/master' into HEAD
[ROCm/rccl-tests commit: 621dde544d]
2024-03-01 18:34:44 +00:00
Wenkai Du ff97af6529 Fix typo in rank assignment (#59)
[ROCm/rccl-tests commit: 7715a0cf1f]
2024-02-15 12:04:38 -08:00
David Addison 5d52f0285c Added missing MPI_Comm_free() call before MPI_Finalize()
[ROCm/rccl-tests commit: c6afef0b6f]
2024-02-05 08:53:54 -08:00
Nusrat Islam 26b1b0b822 Add option to disable out-of-place
[ROCm/rccl-tests commit: a2bec5d2f6]
2024-01-04 16:43:50 -06:00
Wenkai Du ccad358bc9 Warm up both out-of-place and in-place collectives (#51)
[ROCm/rccl-tests commit: 5ee7a08994]
2023-10-16 12:13:50 -07:00