In this commit it disabled by default and can be enabled via
`RCCL_ENABLE_CONTEXT_TRACKING=1` for both (CDNA, RDNA)
Original PR https://github.com/ROCm/rccl/pull/1927
[ROCm/rccl commit: 00a42c80f3]
In this commit it disabled by default and can be enabled via
`RCCL_ENABLE_CONTEXT_TRACKING=1` for both (CDNA, RDNA)
Original PR https://github.com/ROCm/rccl/pull/1927
* Try outputting LastTest.log
* Update if condition for outputting log
* Another attempt
* Only run Ubuntu Noble on MI355 in push/PR
* Try exclude matrix
* Move conditional statement in matrix exclusion
* Create ci-matrix.yml file
* Add needs parameter to ubuntu job
* Fix typo in matrix output variable
* Add back pull_request_template.md
* Add back pull_request_template.md
* Add OMPT to ROCpd
* Use correct category
* Added wrapper functions for future control
* Formatting
* Fix naming
* Comment change
* Remove ompt_get_cb_args
* Switched to using region_sample for OMPT
* Remove relic function
* Remove get_use_rocpd that was used in this pr (one still remains)
* Rename ompt_get_args_string and reuse in tool_tracing_callback_stop
* Make lock init and destroy cb instant
* [Prototype] ROCPD Name fix
* [Prototype] ROCPD Name fix P1
* [Prototype] ROCPD Name fix P2
* ROCPD Name fix
* Var name changes
* Rewrite cb overwrite to single function
* [Important] Use parallel_data as key for parallel callback map
* Fix workflow failure
* Make cpp USE_ROCM consistent with hpp and use default constructor if USE_ROCM = 0
* Add missing ROCPROFILER_VERSION check
* Improve readability
* Make ompt storage maps thread local
* Part 1: Variable name fix, memory cleanup, and fixed asserts
* Part 2: Add comments
* Part 3: Add CI_THROW
* Part 4: Formatting
* Part 5: Move #include to cpp
* add support for compiling all backends
also include the logic to select backends either based on user requests
or through some heuristics
* checkpoint for compiling all backends
* final checkpoint
all tests seem to pass when compiling all three backends simultaneasly
and forcing to use any of the three Backends.
* update PR to new envvar system
[ROCm/rocshmem commit: a1269e3db5]
* add support for compiling all backends
also include the logic to select backends either based on user requests
or through some heuristics
* checkpoint for compiling all backends
* final checkpoint
all tests seem to pass when compiling all three backends simultaneasly
and forcing to use any of the three Backends.
* update PR to new envvar system
Some sections were being displayed multiple times in the web GUI.
Code to append the section was nested inside the subsection loop,
so each time a new subsction was appened to the section,
the entire section was appended.
* Revamp findibverbs to find ionic again
* gda ionic: rename ionic_sq_buf ionic_cq_buf
Avoid duplicating member names used by mlx5 gda.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda: move spin lock to util.hpp
Move spin lock out of ionic gda to util.hpp.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: assume latest fwabi changes
There is no firmware abi compatibility in this ionic gda code yet, so
assume we are using the latest firmware abi as of now.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: allow doorbell with incomplete wqes
Use spin lock to ensure doorbell is only written with an increasing
producer index. Ring the doorbell after this wave has initialized its
wqes. Wqes of other waves might not be fully initialized, but firmware
will not process them until the phase/color flag is updated in the
respecitve wqes.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: poll cq for additional completions
Keep polling the cq for more than just the minimum number of completions
for this wave of threads to make progress, as long as the cq is not
empty. A part of wave-optimized cq polling, at the expense of one wave
polling additional completions, it was observed that nearly all other
waves avoid taking the cq lock at all.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda: max_rd_atomic in rts transition
In modify_qp(RTS), specify max_rd_atomic, not max_dest_rd_atomic.
By not speicfying max_rd_atomic (rather, max_rd_atomic=zero), the local
nic may get stuck transmitting the first read or atomic request. One
read or atomic request is greater than the initiator depth of zero.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: allow specifying traffic class
Allow specifying a traffic class. The network might have a specific
traffic class configured as no-drop, for example.
Co-authored-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: tweak uxdma assignment
The ideal arrangement will have an equal number of QPs active on each
uxdma pipeline.
Pre-rebase, the better arrangement for rocshmem funcitonal test
benchmarks was [0, 1], [1, 0], [0, 1], [1, 0], ...
Now, following changes that add 'ROCSHMEM_GDA_ALTERNATE_QP_PORTS=1' by
default, the better arrangement is [0, 1], [0, 1], [0, 1], [0, 1], ...
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
---------
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
Co-authored-by: Aurelien Bouteiller <abouteil@amd.com>
Co-authored-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
[ROCm/rocshmem commit: c84bbc250b]
* Revamp findibverbs to find ionic again
* gda ionic: rename ionic_sq_buf ionic_cq_buf
Avoid duplicating member names used by mlx5 gda.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda: move spin lock to util.hpp
Move spin lock out of ionic gda to util.hpp.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: assume latest fwabi changes
There is no firmware abi compatibility in this ionic gda code yet, so
assume we are using the latest firmware abi as of now.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: allow doorbell with incomplete wqes
Use spin lock to ensure doorbell is only written with an increasing
producer index. Ring the doorbell after this wave has initialized its
wqes. Wqes of other waves might not be fully initialized, but firmware
will not process them until the phase/color flag is updated in the
respecitve wqes.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: poll cq for additional completions
Keep polling the cq for more than just the minimum number of completions
for this wave of threads to make progress, as long as the cq is not
empty. A part of wave-optimized cq polling, at the expense of one wave
polling additional completions, it was observed that nearly all other
waves avoid taking the cq lock at all.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda: max_rd_atomic in rts transition
In modify_qp(RTS), specify max_rd_atomic, not max_dest_rd_atomic.
By not speicfying max_rd_atomic (rather, max_rd_atomic=zero), the local
nic may get stuck transmitting the first read or atomic request. One
read or atomic request is greater than the initiator depth of zero.
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: allow specifying traffic class
Allow specifying a traffic class. The network might have a specific
traffic class configured as no-drop, for example.
Co-authored-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
* gda ionic: tweak uxdma assignment
The ideal arrangement will have an equal number of QPs active on each
uxdma pipeline.
Pre-rebase, the better arrangement for rocshmem funcitonal test
benchmarks was [0, 1], [1, 0], [0, 1], [1, 0], ...
Now, following changes that add 'ROCSHMEM_GDA_ALTERNATE_QP_PORTS=1' by
default, the better arrangement is [0, 1], [0, 1], [0, 1], [0, 1], ...
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
---------
Signed-off-by: Allen Hubbe <allen.hubbe@amd.com>
Co-authored-by: Aurelien Bouteiller <abouteil@amd.com>
Co-authored-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
* Trigger CI run on pull request
* Enabling CI run on different PR types
---------
Co-authored-by: arravikum <arravikum@amd.com>
[ROCm/rccl commit: 1858a31c41]
Added `GPU LINK PORT STATUS` table to `amd-smi xgmi` command
The `amd-smi xgmi -s` or `amd-smi xgmi --source-status` will show `GPU LINK PORT STATUS` table.
---------
Signed-off-by: Bindhiya Kanangot Balakrishnan <Bindhiya.KanangotBalakrishnan@amd.com>
Added `GPU LINK PORT STATUS` table to `amd-smi xgmi` command
The `amd-smi xgmi -s` or `amd-smi xgmi --source-status` will show `GPU LINK PORT STATUS` table.
---------
Signed-off-by: Bindhiya Kanangot Balakrishnan <Bindhiya.KanangotBalakrishnan@amd.com>
[ROCm/amdsmi commit: 7ddd91653e]