[rocprof-compute] update yamls for docs (#1887)

Этот коммит содержится в:
xuchen-amd
2025-11-19 10:46:02 -05:00
коммит произвёл GitHub
родитель eddd4c3601
Коммит c778acdb70
83 изменённых файлов: 708 добавлений и 836 удалений
@@ -200,37 +200,37 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -263,7 +263,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -296,7 +296,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -170,15 +170,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -266,7 +266,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+13 -13
Просмотреть файл
@@ -134,48 +134,48 @@ Panel Config:
/ 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
from MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -196,22 +196,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -141,6 +141,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -168,7 +168,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -140,7 +140,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -436,7 +436,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
@@ -218,37 +218,37 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -281,7 +281,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -314,7 +314,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -170,15 +170,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -266,7 +266,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+13 -13
Просмотреть файл
@@ -132,48 +132,48 @@ Panel Config:
/ 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
from MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -194,22 +194,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -141,6 +141,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -168,7 +168,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -268,7 +268,7 @@ Panel Config:
floating-point operands issued to the VALU per normalization unit.
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
on 64-bit floating-point operands issued to the VALU per normalization unit.
Conversion: |-
Conversion: >-
The total number of type conversion instructions (such as converting
data to or from F32\u2194F64) issued to the VALU per normalization unit.
Global/Generic Instr: The total number of global & generic memory instructions
@@ -237,37 +237,37 @@ Panel Config:
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
unit: (OPs + $normUnit)
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (INT8): |-
MFMA IOPs (INT8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -140,7 +140,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -436,7 +436,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
@@ -227,12 +227,12 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
@@ -242,27 +242,27 @@ Panel Config:
from VALU instructions. This is also presented as a percent of the peak theoretical
F8 MFMA operations achievable on the specific accelerator. It is supported on
AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -295,7 +295,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -328,7 +328,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -162,15 +162,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -252,7 +252,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+13 -13
Просмотреть файл
@@ -140,17 +140,17 @@ Panel Config:
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
@@ -160,33 +160,33 @@ Panel Config:
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
on the specific accelerator is displayed alongside for comparison. It is supported
on AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -207,22 +207,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -141,6 +141,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -168,7 +168,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -273,7 +273,7 @@ Panel Config:
floating-point operands issued to the VALU per normalization unit.
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
on 64-bit floating-point operands issued to the VALU per normalization unit.
Conversion: |-
Conversion: >-
The total number of type conversion instructions (such as converting
data to or from F32\u2194F64) issued to the VALU per normalization unit.
Global/Generic Instr: The total number of global & generic memory instructions
@@ -251,37 +251,37 @@ Panel Config:
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
unit: (OPs + $normUnit)
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (INT8): |-
MFMA IOPs (INT8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -140,7 +140,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -398,7 +398,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
@@ -227,12 +227,12 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
@@ -242,27 +242,27 @@ Panel Config:
from VALU instructions. This is also presented as a percent of the peak theoretical
F8 MFMA operations achievable on the specific accelerator. It is supported on
AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -295,7 +295,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -328,7 +328,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -162,15 +162,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -252,7 +252,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+13 -13
Просмотреть файл
@@ -140,17 +140,17 @@ Panel Config:
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
@@ -160,33 +160,33 @@ Panel Config:
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
on the specific accelerator is displayed alongside for comparison. It is supported
on AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -207,22 +207,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -141,6 +141,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -168,7 +168,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -273,7 +273,7 @@ Panel Config:
floating-point operands issued to the VALU per normalization unit.
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
on 64-bit floating-point operands issued to the VALU per normalization unit.
Conversion: |-
Conversion: >-
The total number of type conversion instructions (such as converting
data to or from F32\u2194F64) issued to the VALU per normalization unit.
Global/Generic Instr: The total number of global & generic memory instructions
@@ -251,37 +251,37 @@ Panel Config:
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
unit: (OPs + $normUnit)
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (INT8): |-
MFMA IOPs (INT8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -140,7 +140,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -398,7 +398,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
@@ -227,12 +227,12 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
@@ -242,27 +242,27 @@ Panel Config:
from VALU instructions. This is also presented as a percent of the peak theoretical
F8 MFMA operations achievable on the specific accelerator. It is supported on
AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -295,7 +295,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -328,7 +328,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -162,15 +162,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -252,7 +252,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+13 -13
Просмотреть файл
@@ -140,17 +140,17 @@ Panel Config:
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
@@ -160,33 +160,33 @@ Panel Config:
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
on the specific accelerator is displayed alongside for comparison. It is supported
on AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -207,22 +207,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -141,6 +141,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -168,7 +168,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -273,7 +273,7 @@ Panel Config:
floating-point operands issued to the VALU per normalization unit.
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
on 64-bit floating-point operands issued to the VALU per normalization unit.
Conversion: |-
Conversion: >-
The total number of type conversion instructions (such as converting
data to or from F32\u2194F64) issued to the VALU per normalization unit.
Global/Generic Instr: The total number of global & generic memory instructions
@@ -251,37 +251,37 @@ Panel Config:
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
unit: (OPs + $normUnit)
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (INT8): |-
MFMA IOPs (INT8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -140,7 +140,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -398,7 +398,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
@@ -233,12 +233,12 @@ Panel Config:
pop: None
coll_level: SQ_IFETCH_LEVEL
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
@@ -248,27 +248,27 @@ Panel Config:
from VALU instructions. This is also presented as a percent of the peak theoretical
F8 MFMA operations achievable on the specific accelerator. It is supported on
AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -301,7 +301,7 @@ Panel Config:
IPC: The ratio of the total number of instructions executed on the CU over the
total active CU cycles. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
@@ -334,7 +334,7 @@ Panel Config:
if only a single value is requested in a cache line, the data movement will
still be counted as a full cache line. This is also presented as a percent of
the peak theoretical bandwidth achievable on the specific accelerator.
L2-Fabric Read BW: |-
L2-Fabric Read BW: >-
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
per unit time. This is also presented as a percent of the peak theoretical
bandwidth achievable on the specific accelerator.
@@ -172,15 +172,15 @@ Panel Config:
Active CUs: Total number of active compute units (CUs) on the accelerator during
the kernel execution.
Num CUs: Total number of compute units (CUs) on the accelerator.
VGPR: |-
VGPR: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
SGPR: |-
SGPR: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -268,7 +268,7 @@ Panel Config:
or data (atomic with return value) was returned to the L2.
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
of data from the accelerator's local HBM, per normalization unit.
HBM Wr: |-
HBM Wr: >-
The total number of L2 requests to Infinity Fabric to write or atomically
update 32B or 64B of data in the accelerator's local HBM, per normalization
unit.
+14 -14
Просмотреть файл
@@ -148,17 +148,17 @@ Panel Config:
Start_Timestamp) / 1e9) ) / 1e9
unit: GFLOP/s
metrics_description:
VALU FLOPs (F16): |-
VALU FLOPs (F16): >-
The total 16-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F16 FLOPs achievable
on the specific accelerator. Note: this does not include any F16 operations
from MFMA instructions.
VALU FLOPs (F32): |-
VALU FLOPs (F32): >-
The total 32-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F32 FLOPs achievable
on the specific accelerator. Note: this does not include any F32 operations
from MFMA instructions.
VALU FLOPs (F64): |-
VALU FLOPs (F64): >-
The total 64-bit floating-point operations executed per second on the VALU.
This is presented with the value of the peak empirical F64 FLOPs achievable
on the specific accelerator. Note: this does not include any F64 operations
@@ -168,39 +168,39 @@ Panel Config:
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
on the specific accelerator is displayed alongside for comparison. It is supported
on AMD Instinct MI300 series and later only.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point
operations from VALU instructions. The peak empirically measured BF16 MFMA
operations achievable on the specific accelerator is displayed alongside
for comparison.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. The peak empirically measured F16 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. The peak empirically measured F32 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. The peak empirically measured F64 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
MFMA FLOPs (F6F4): |-
MFMA FLOPs (F6F4): >-
The total number of 4-bit and 6-bit floating point MFMA operations executed
per second. Note: this does not include any floating point operations from
VALU instructions. The peak empirically measured F6F4 MFMA operations
achievable on the specific accelerator is displayed alongside for comparison.
It is supported on AMD Instinct MI350 series (gfx950) and later only.
MFMA IOPs (Int8): |-
MFMA IOPs (Int8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
The peak empirically measured INT8 MFMA operations achievable on the specific
accelerator is displayed alongside for comparison.
HBM Bandwidth: |-
HBM Bandwidth: >-
The total number of bytes read from and written to High-Bandwidth
Memory (HBM) per second. The peak empirically measured bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
@@ -221,22 +221,22 @@ Panel Config:
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
example for more detail). The peak empirically measured LDS bandwidth achievable
on the specific accelerator is displayed alongside for comparison.
AI L1: |-
AI L1: >-
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L1 cache and the processing units. This value is used as the x-coordinate
for the L1 roofline.
AI L2: |-
AI L2: >-
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
of total floating-point operations (FLOPs) to total bytes transferred between
the L2 cache and the L1 cache. This value is used as the x-coordinate for
the L2 roofline.
AI HBM: |-
AI HBM: >-
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
It is the ratio of total floating-point operations (FLOPs) to total bytes
transferred between HBM and the L2 cache. This value is used as the x-coordinate
for the HBM roofline.
Performance (GFLOPs): |-
Performance (GFLOPs): >-
The overall achieved performance, measured in GigaFLOPs
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
operations divided by the total execution time. This value is used as the y-coordinate
@@ -162,6 +162,6 @@ Panel Config:
the CPC-L2 interface was active doing any work.
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
translation
CPC-UTCL2 Utilization: |-
CPC-UTCL2 Utilization: >-
Percent of total cycles counted by the CPC's L2 address translation
interface where the CPC was busy doing address translation work.
@@ -204,7 +204,7 @@ Panel Config:
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
resources.
Not-scheduled Rate (Scheduler-Pipe): |-
Not-scheduled Rate (Scheduler-Pipe): >-
The percent of total scheduler-pipe cycles in the kernel where a workgroup
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
rather than a lack of a CU or SIMD with sufficient resources.
+6 -6
Просмотреть файл
@@ -121,26 +121,26 @@ Panel Config:
Workgroup Size: The total number of work-items (or, threads) in each workgroup
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
to the total block size.
Total Wavefronts: |-
Total Wavefronts: >-
The total number of wavefronts launched as part of the kernel dispatch.
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
size is always 64 work-items. Thus, the total number of wavefronts should
be equivalent to the ceiling of grid size divided by 64.
Saved Wavefronts: The total number of wavefronts saved at a context-save.
Restored Wavefronts: The total number of wavefronts restored from a context-save.
VGPRs: |-
VGPRs: >-
The number of architected vector general-purpose registers allocated
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
requested by the compiler due to allocation granularity.
AGPRs: |-
AGPRs: >-
The number of accumulation vector general-purpose registers allocated
for the kernel, see AGPRs. Note: this may not exactly match the number of
AGPRs requested by the compiler due to allocation granularity.
SGPRs: |-
SGPRs: >-
The number of scalar general-purpose registers allocated for the kernel,
see SALU. Note: this may not exactly match the number of SGPRs requested by
the compiler due to allocation granularity.
LDS Allocation: |-
LDS Allocation: >-
The number of bytes of LDS memory (or, shared memory) allocated for
this kernel. Note: This may also be larger than what was requested at compile
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
@@ -173,7 +173,7 @@ Panel Config:
rather than identification of a precise limiter. The sum of this metric, Issue
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
metric.
Wavefront Occupancy: |-
Wavefront Occupancy: >-
The time-averaged number of wavefronts resident on the accelerator over
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
kernels (less than 1ms).
@@ -283,7 +283,7 @@ Panel Config:
floating-point operands issued to the VALU per normalization unit.
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
on 64-bit floating-point operands issued to the VALU per normalization unit.
Conversion: |-
Conversion: >-
The total number of type conversion instructions (such as converting
data to or from F32\u2194F64) issued to the VALU per normalization unit.
Global/Generic Instr: The total number of global & generic memory instructions
@@ -267,37 +267,37 @@ Panel Config:
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
unit: (OPs + $normUnit)
metrics_description:
VALU FLOPs: |-
VALU FLOPs: >-
The total floating-point operations executed per second on the VALU.
This is also presented as a percent of the peak theoretical FLOPs achievable
on the specific accelerator. Note: this does not include any floating-point
operations from MFMA instructions.
VALU IOPs: |-
VALU IOPs: >-
The total integer operations executed per second on the VALU. This is
also presented as a percent of the peak theoretical IOPs achievable on the
specific accelerator. Note: this does not include any integer operations from
MFMA instructions.
MFMA FLOPs (BF16): |-
MFMA FLOPs (BF16): >-
The total number of 16-bit brain floating point MFMA operations executed
per second. Note: this does not include any 16-bit brain floating point operations
from VALU instructions. This is also presented as a percent of the peak theoretical
BF16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F16): |-
MFMA FLOPs (F16): >-
The total number of 16-bit floating point MFMA operations executed per
second. Note: this does not include any 16-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F16 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F32): |-
MFMA FLOPs (F32): >-
The total number of 32-bit floating point MFMA operations executed per
second. Note: this does not include any 32-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F32 MFMA operations achievable on the specific accelerator.
MFMA FLOPs (F64): |-
MFMA FLOPs (F64): >-
The total number of 64-bit floating point MFMA operations executed per
second. Note: this does not include any 64-bit floating point operations from
VALU instructions. This is also presented as a percent of the peak theoretical
F64 MFMA operations achievable on the specific accelerator.
MFMA IOPs (INT8): |-
MFMA IOPs (INT8): >-
The total number of 8-bit integer MFMA operations executed per second.
Note: this does not include any 8-bit integer operations from VALU instructions.
This is also presented as a percent of the peak theoretical INT8 MFMA operations
@@ -180,7 +180,7 @@ Panel Config:
unit.
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
stalls from non-dword aligned addresses per normalization unit.
Mem Violations: |-
Mem Violations: >-
The total number of out-of-bounds accesses made to the LDS, per normalization
unit. This is unused and expected to be zero in most configurations for
modern CDNA\u2122 accelerators.
@@ -92,7 +92,7 @@ Panel Config:
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
loaded line the cache. Calculated as the ratio of the number of L1I requests
that hit over the number of all L1I requests.
L1I-L2 Bandwidth Utilization: |-
L1I-L2 Bandwidth Utilization: >-
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
achieved. Calculated as the ratio of the total number of requests from the
L1I to the L2 cache over the total L1I-L2 interface cycles.
@@ -154,7 +154,7 @@ Panel Config:
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
bandwidth acheived. Calculated as total number of bytes read from, written to,
or atomically updated across the sL1D - L2 interface.
sL1D-L2 BW: |-
sL1D-L2 BW: >-
The total number of bytes read from, written to, or atomically updated
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
writes and atomics are typically unused on current CDNA accelerators, so
@@ -164,7 +164,7 @@ Panel Config:
unit.
Hits: The total number of sL1D requests that hit on a previously loaded cache
line, per normalization unit.
Misses - Non Duplicated: |-
Misses - Non Duplicated: >-
The total number of sL1D requests that missed on a cache line that was
not already pending due to another request, per normalization unit.
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
@@ -187,6 +187,6 @@ Panel Config:
unit.
Write Req: The total number of write requests from sL1D to the L2, per normalization
unit. Typically unused on current CDNA accelerators.
Stall Cycles: |-
Stall Cycles: >-
The total number of cycles the sL1D\u2194L2 interface was stalled, per
normalization unit.
@@ -501,7 +501,7 @@ Panel Config:
per normalization unit.
Translation Misses: The total number of translation requests that missed in the
UTCL1 due to translation not being present in the cache, per normalization unit.
Permission Misses: |-
Permission Misses: >-
The total number of translation requests that missed in the UTCL1 due
to a permission error, per normalization unit. This is unused and expected
to be zero in most configurations for modern CDNA\u2122 accelerators.
+1 -1
Просмотреть файл
@@ -706,7 +706,7 @@ Panel Config:
requests are only considered atomic by Infinity Fabric if they are targeted
at non-write-cacheable memory, such as fine-grained memory allocations or uncached
memory allocations on the MI2XX.
Read Stall: |-
Read Stall: >-
The ratio of the total number of cycles the L2-Fabric interface was
stalled on a read request to any destination (local HBM, remote PCIe\xAE
connected accelerator or CPU, or remote Infinity Fabric connected accelerator