SWDEV-362046 - Report HIP_OPS activities using the ROCr driver_node_id instead of the device's index

The ROCclr assigns zero-based IDs to GPUs in the order they are
discovered. That zero-based ID is what is used to identify the GPU
on which the HIP_OPS activity took place.

When multiple ranks are used, each rank's first logical device always
has GPU ID 0, regardless of which physical device is selected with
CUDA_VISIBLE_DEVICES. Because of this, when merging trace files from
multiple ranks, GPU IDs from different processes may overlap.

The long term solution is to use the KFD's gpu_id which is stable
across APIs and processes. Unfortunately the gpu_id is not yet exposed
by the ROCr, so for now use the driver's node id.

Change-Id: Ib78854527d600d175bb76e2df0747c33f898c615
This commit is contained in:
Laurent Morichetti
2022-10-18 19:51:02 -07:00
committed by Laurent Morichetti
szülő dacd55f3d7
commit 9a82118c85
3 fájl változott, egészen pontosan 11 új sor hozzáadva és 2 régi sor törölve
+2 -2
Fájl megtekintése
@@ -73,8 +73,8 @@ void ReportActivity(const amd::Command& command) {
command.profilingInfo().start_, // begin timestamp, ns
command.profilingInfo().end_, // end timestamp, ns
{{
static_cast<int>(queue->device().index()), // device id
queue->vdev()->index() // queue id
static_cast<int>(queue->device().info().driverNodeId_), // device id
queue->vdev()->index() // queue id
}},
{} // copied data size for memcpy, or kernel name for dispatch
};