[rocprof-compute] update yamls for docs (#1887)
Этот коммит содержится в:
+9
-9
@@ -200,37 +200,37 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -263,7 +263,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -296,7 +296,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -170,15 +170,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -266,7 +266,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+13
-13
@@ -134,48 +134,48 @@ Panel Config:
|
||||
/ 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
from MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -196,22 +196,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -141,6 +141,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -168,7 +168,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -140,7 +140,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -436,7 +436,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+9
-9
@@ -218,37 +218,37 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -281,7 +281,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -314,7 +314,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -170,15 +170,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -266,7 +266,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+13
-13
@@ -132,48 +132,48 @@ Panel Config:
|
||||
/ 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
from MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -194,22 +194,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -141,6 +141,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -168,7 +168,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -268,7 +268,7 @@ Panel Config:
|
||||
floating-point operands issued to the VALU per normalization unit.
|
||||
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
|
||||
on 64-bit floating-point operands issued to the VALU per normalization unit.
|
||||
Conversion: |-
|
||||
Conversion: >-
|
||||
The total number of type conversion instructions (such as converting
|
||||
data to or from F32\u2194F64) issued to the VALU per normalization unit.
|
||||
Global/Generic Instr: The total number of global & generic memory instructions
|
||||
|
||||
+7
-7
@@ -237,37 +237,37 @@ Panel Config:
|
||||
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
|
||||
unit: (OPs + $normUnit)
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (INT8): |-
|
||||
MFMA IOPs (INT8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
|
||||
+1
-1
@@ -140,7 +140,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -436,7 +436,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+9
-9
@@ -227,12 +227,12 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
@@ -242,27 +242,27 @@ Panel Config:
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F8 MFMA operations achievable on the specific accelerator. It is supported on
|
||||
AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -295,7 +295,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -328,7 +328,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -162,15 +162,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -252,7 +252,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+13
-13
@@ -140,17 +140,17 @@ Panel Config:
|
||||
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
@@ -160,33 +160,33 @@ Panel Config:
|
||||
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
|
||||
on the specific accelerator is displayed alongside for comparison. It is supported
|
||||
on AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -207,22 +207,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -141,6 +141,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -168,7 +168,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -273,7 +273,7 @@ Panel Config:
|
||||
floating-point operands issued to the VALU per normalization unit.
|
||||
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
|
||||
on 64-bit floating-point operands issued to the VALU per normalization unit.
|
||||
Conversion: |-
|
||||
Conversion: >-
|
||||
The total number of type conversion instructions (such as converting
|
||||
data to or from F32\u2194F64) issued to the VALU per normalization unit.
|
||||
Global/Generic Instr: The total number of global & generic memory instructions
|
||||
|
||||
+7
-7
@@ -251,37 +251,37 @@ Panel Config:
|
||||
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
|
||||
unit: (OPs + $normUnit)
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (INT8): |-
|
||||
MFMA IOPs (INT8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
|
||||
+1
-1
@@ -140,7 +140,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -398,7 +398,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+9
-9
@@ -227,12 +227,12 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
@@ -242,27 +242,27 @@ Panel Config:
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F8 MFMA operations achievable on the specific accelerator. It is supported on
|
||||
AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -295,7 +295,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -328,7 +328,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -162,15 +162,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -252,7 +252,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+13
-13
@@ -140,17 +140,17 @@ Panel Config:
|
||||
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
@@ -160,33 +160,33 @@ Panel Config:
|
||||
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
|
||||
on the specific accelerator is displayed alongside for comparison. It is supported
|
||||
on AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -207,22 +207,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -141,6 +141,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -168,7 +168,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -273,7 +273,7 @@ Panel Config:
|
||||
floating-point operands issued to the VALU per normalization unit.
|
||||
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
|
||||
on 64-bit floating-point operands issued to the VALU per normalization unit.
|
||||
Conversion: |-
|
||||
Conversion: >-
|
||||
The total number of type conversion instructions (such as converting
|
||||
data to or from F32\u2194F64) issued to the VALU per normalization unit.
|
||||
Global/Generic Instr: The total number of global & generic memory instructions
|
||||
|
||||
+7
-7
@@ -251,37 +251,37 @@ Panel Config:
|
||||
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
|
||||
unit: (OPs + $normUnit)
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (INT8): |-
|
||||
MFMA IOPs (INT8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
|
||||
+1
-1
@@ -140,7 +140,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -398,7 +398,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+9
-9
@@ -227,12 +227,12 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
@@ -242,27 +242,27 @@ Panel Config:
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F8 MFMA operations achievable on the specific accelerator. It is supported on
|
||||
AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -295,7 +295,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -328,7 +328,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -162,15 +162,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -252,7 +252,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+13
-13
@@ -140,17 +140,17 @@ Panel Config:
|
||||
* 512) ) / (SUM(End_Timestamp - Start_Timestamp) / 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
@@ -160,33 +160,33 @@ Panel Config:
|
||||
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
|
||||
on the specific accelerator is displayed alongside for comparison. It is supported
|
||||
on AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -207,22 +207,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -141,6 +141,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -168,7 +168,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -273,7 +273,7 @@ Panel Config:
|
||||
floating-point operands issued to the VALU per normalization unit.
|
||||
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
|
||||
on 64-bit floating-point operands issued to the VALU per normalization unit.
|
||||
Conversion: |-
|
||||
Conversion: >-
|
||||
The total number of type conversion instructions (such as converting
|
||||
data to or from F32\u2194F64) issued to the VALU per normalization unit.
|
||||
Global/Generic Instr: The total number of global & generic memory instructions
|
||||
|
||||
+7
-7
@@ -251,37 +251,37 @@ Panel Config:
|
||||
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
|
||||
unit: (OPs + $normUnit)
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (INT8): |-
|
||||
MFMA IOPs (INT8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
|
||||
+1
-1
@@ -140,7 +140,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -398,7 +398,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+9
-9
@@ -233,12 +233,12 @@ Panel Config:
|
||||
pop: None
|
||||
coll_level: SQ_IFETCH_LEVEL
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
@@ -248,27 +248,27 @@ Panel Config:
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F8 MFMA operations achievable on the specific accelerator. It is supported on
|
||||
AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
@@ -301,7 +301,7 @@ Panel Config:
|
||||
IPC: The ratio of the total number of instructions executed on the CU over the
|
||||
total active CU cycles. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms). This is also presented as a percent of the peak theoretical
|
||||
@@ -334,7 +334,7 @@ Panel Config:
|
||||
if only a single value is requested in a cache line, the data movement will
|
||||
still be counted as a full cache line. This is also presented as a percent of
|
||||
the peak theoretical bandwidth achievable on the specific accelerator.
|
||||
L2-Fabric Read BW: |-
|
||||
L2-Fabric Read BW: >-
|
||||
The number of bytes read by the L2 over the Infinity Fabric\u2122 interface
|
||||
per unit time. This is also presented as a percent of the peak theoretical
|
||||
bandwidth achievable on the specific accelerator.
|
||||
|
||||
+4
-4
@@ -172,15 +172,15 @@ Panel Config:
|
||||
Active CUs: Total number of active compute units (CUs) on the accelerator during
|
||||
the kernel execution.
|
||||
Num CUs: Total number of compute units (CUs) on the accelerator.
|
||||
VGPR: |-
|
||||
VGPR: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
SGPR: |-
|
||||
SGPR: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -268,7 +268,7 @@ Panel Config:
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
HBM Rd: The total number of L2 requests to Infinity Fabric to read 32B or 64B
|
||||
of data from the accelerator's local HBM, per normalization unit.
|
||||
HBM Wr: |-
|
||||
HBM Wr: >-
|
||||
The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per normalization
|
||||
unit.
|
||||
|
||||
+14
-14
@@ -148,17 +148,17 @@ Panel Config:
|
||||
Start_Timestamp) / 1e9) ) / 1e9
|
||||
unit: GFLOP/s
|
||||
metrics_description:
|
||||
VALU FLOPs (F16): |-
|
||||
VALU FLOPs (F16): >-
|
||||
The total 16-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F16 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F16 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F32): |-
|
||||
VALU FLOPs (F32): >-
|
||||
The total 32-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F32 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F32 operations
|
||||
from MFMA instructions.
|
||||
VALU FLOPs (F64): |-
|
||||
VALU FLOPs (F64): >-
|
||||
The total 64-bit floating-point operations executed per second on the VALU.
|
||||
This is presented with the value of the peak empirical F64 FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any F64 operations
|
||||
@@ -168,39 +168,39 @@ Panel Config:
|
||||
from VALU instructions. The peak empirically measured F8 MFMA operations achievable
|
||||
on the specific accelerator is displayed alongside for comparison. It is supported
|
||||
on AMD Instinct MI300 series and later only.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point
|
||||
operations from VALU instructions. The peak empirically measured BF16 MFMA
|
||||
operations achievable on the specific accelerator is displayed alongside
|
||||
for comparison.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F16 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F32 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. The peak empirically measured F64 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
MFMA FLOPs (F6F4): |-
|
||||
MFMA FLOPs (F6F4): >-
|
||||
The total number of 4-bit and 6-bit floating point MFMA operations executed
|
||||
per second. Note: this does not include any floating point operations from
|
||||
VALU instructions. The peak empirically measured F6F4 MFMA operations
|
||||
achievable on the specific accelerator is displayed alongside for comparison.
|
||||
It is supported on AMD Instinct MI350 series (gfx950) and later only.
|
||||
MFMA IOPs (Int8): |-
|
||||
MFMA IOPs (Int8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
The peak empirically measured INT8 MFMA operations achievable on the specific
|
||||
accelerator is displayed alongside for comparison.
|
||||
HBM Bandwidth: |-
|
||||
HBM Bandwidth: >-
|
||||
The total number of bytes read from and written to High-Bandwidth
|
||||
Memory (HBM) per second. The peak empirically measured bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
@@ -221,22 +221,22 @@ Panel Config:
|
||||
from, stored to, or atomically updated in the LDS per unit time (see LDS Bandwidth
|
||||
example for more detail). The peak empirically measured LDS bandwidth achievable
|
||||
on the specific accelerator is displayed alongside for comparison.
|
||||
AI L1: |-
|
||||
AI L1: >-
|
||||
The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L1 cache and the processing units. This value is used as the x-coordinate
|
||||
for the L1 roofline.
|
||||
AI L2: |-
|
||||
AI L2: >-
|
||||
The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio
|
||||
of total floating-point operations (FLOPs) to total bytes transferred between
|
||||
the L2 cache and the L1 cache. This value is used as the x-coordinate for
|
||||
the L2 roofline.
|
||||
AI HBM: |-
|
||||
AI HBM: >-
|
||||
The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM).
|
||||
It is the ratio of total floating-point operations (FLOPs) to total bytes
|
||||
transferred between HBM and the L2 cache. This value is used as the x-coordinate
|
||||
for the HBM roofline.
|
||||
Performance (GFLOPs): |-
|
||||
Performance (GFLOPs): >-
|
||||
The overall achieved performance, measured in GigaFLOPs
|
||||
per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point
|
||||
operations divided by the total execution time. This value is used as the y-coordinate
|
||||
|
||||
+1
-1
@@ -162,6 +162,6 @@ Panel Config:
|
||||
the CPC-L2 interface was active doing any work.
|
||||
CPC-UTCL1 Stall: Percent of CPC busy cycles where the CPC was stalled by address
|
||||
translation
|
||||
CPC-UTCL2 Utilization: |-
|
||||
CPC-UTCL2 Utilization: >-
|
||||
Percent of total cycles counted by the CPC's L2 address translation
|
||||
interface where the CPC was busy doing address translation work.
|
||||
|
||||
+1
-1
@@ -204,7 +204,7 @@ Panel Config:
|
||||
in the kernel where a workgroup could not be scheduled to a CU due to a bottleneck
|
||||
within the workgroup manager rather than a lack of a CU or SIMD with sufficient
|
||||
resources.
|
||||
Not-scheduled Rate (Scheduler-Pipe): |-
|
||||
Not-scheduled Rate (Scheduler-Pipe): >-
|
||||
The percent of total scheduler-pipe cycles in the kernel where a workgroup
|
||||
could not be scheduled to a CU due to a bottleneck within the scheduler-pipes
|
||||
rather than a lack of a CU or SIMD with sufficient resources.
|
||||
|
||||
+6
-6
@@ -121,26 +121,26 @@ Panel Config:
|
||||
Workgroup Size: The total number of work-items (or, threads) in each workgroup
|
||||
(or, block) launched as part of the kernel dispatch. In HIP, this is equivalent
|
||||
to the total block size.
|
||||
Total Wavefronts: |-
|
||||
Total Wavefronts: >-
|
||||
The total number of wavefronts launched as part of the kernel dispatch.
|
||||
On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront
|
||||
size is always 64 work-items. Thus, the total number of wavefronts should
|
||||
be equivalent to the ceiling of grid size divided by 64.
|
||||
Saved Wavefronts: The total number of wavefronts saved at a context-save.
|
||||
Restored Wavefronts: The total number of wavefronts restored from a context-save.
|
||||
VGPRs: |-
|
||||
VGPRs: >-
|
||||
The number of architected vector general-purpose registers allocated
|
||||
for the kernel, see VALU. Note: this may not exactly match the number of VGPRs
|
||||
requested by the compiler due to allocation granularity.
|
||||
AGPRs: |-
|
||||
AGPRs: >-
|
||||
The number of accumulation vector general-purpose registers allocated
|
||||
for the kernel, see AGPRs. Note: this may not exactly match the number of
|
||||
AGPRs requested by the compiler due to allocation granularity.
|
||||
SGPRs: |-
|
||||
SGPRs: >-
|
||||
The number of scalar general-purpose registers allocated for the kernel,
|
||||
see SALU. Note: this may not exactly match the number of SGPRs requested by
|
||||
the compiler due to allocation granularity.
|
||||
LDS Allocation: |-
|
||||
LDS Allocation: >-
|
||||
The number of bytes of LDS memory (or, shared memory) allocated for
|
||||
this kernel. Note: This may also be larger than what was requested at compile
|
||||
time due to both allocation granularity and dynamic per-dispatch LDS allocations.
|
||||
@@ -173,7 +173,7 @@ Panel Config:
|
||||
rather than identification of a precise limiter. The sum of this metric, Issue
|
||||
Wait Cycles and Active Wait Cycles should be equal to the total Wave Cycles
|
||||
metric.
|
||||
Wavefront Occupancy: |-
|
||||
Wavefront Occupancy: >-
|
||||
The time-averaged number of wavefronts resident on the accelerator over
|
||||
the lifetime of the kernel. Note: this metric may be inaccurate for short-running
|
||||
kernels (less than 1ms).
|
||||
|
||||
+1
-1
@@ -283,7 +283,7 @@ Panel Config:
|
||||
floating-point operands issued to the VALU per normalization unit.
|
||||
F64-Trans: The total number of transcendental instructions (such as sqrt) operating
|
||||
on 64-bit floating-point operands issued to the VALU per normalization unit.
|
||||
Conversion: |-
|
||||
Conversion: >-
|
||||
The total number of type conversion instructions (such as converting
|
||||
data to or from F32\u2194F64) issued to the VALU per normalization unit.
|
||||
Global/Generic Instr: The total number of global & generic memory instructions
|
||||
|
||||
+7
-7
@@ -267,37 +267,37 @@ Panel Config:
|
||||
max: MAX(((SQ_INSTS_VALU_MFMA_MOPS_I8 * 512) / $denom))
|
||||
unit: (OPs + $normUnit)
|
||||
metrics_description:
|
||||
VALU FLOPs: |-
|
||||
VALU FLOPs: >-
|
||||
The total floating-point operations executed per second on the VALU.
|
||||
This is also presented as a percent of the peak theoretical FLOPs achievable
|
||||
on the specific accelerator. Note: this does not include any floating-point
|
||||
operations from MFMA instructions.
|
||||
VALU IOPs: |-
|
||||
VALU IOPs: >-
|
||||
The total integer operations executed per second on the VALU. This is
|
||||
also presented as a percent of the peak theoretical IOPs achievable on the
|
||||
specific accelerator. Note: this does not include any integer operations from
|
||||
MFMA instructions.
|
||||
MFMA FLOPs (BF16): |-
|
||||
MFMA FLOPs (BF16): >-
|
||||
The total number of 16-bit brain floating point MFMA operations executed
|
||||
per second. Note: this does not include any 16-bit brain floating point operations
|
||||
from VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
BF16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F16): |-
|
||||
MFMA FLOPs (F16): >-
|
||||
The total number of 16-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 16-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F16 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F32): |-
|
||||
MFMA FLOPs (F32): >-
|
||||
The total number of 32-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 32-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F32 MFMA operations achievable on the specific accelerator.
|
||||
MFMA FLOPs (F64): |-
|
||||
MFMA FLOPs (F64): >-
|
||||
The total number of 64-bit floating point MFMA operations executed per
|
||||
second. Note: this does not include any 64-bit floating point operations from
|
||||
VALU instructions. This is also presented as a percent of the peak theoretical
|
||||
F64 MFMA operations achievable on the specific accelerator.
|
||||
MFMA IOPs (INT8): |-
|
||||
MFMA IOPs (INT8): >-
|
||||
The total number of 8-bit integer MFMA operations executed per second.
|
||||
Note: this does not include any 8-bit integer operations from VALU instructions.
|
||||
This is also presented as a percent of the peak theoretical INT8 MFMA operations
|
||||
|
||||
+1
-1
@@ -180,7 +180,7 @@ Panel Config:
|
||||
unit.
|
||||
Unaligned Stall: The total number of cycles spent in the LDS scheduler due to
|
||||
stalls from non-dword aligned addresses per normalization unit.
|
||||
Mem Violations: |-
|
||||
Mem Violations: >-
|
||||
The total number of out-of-bounds accesses made to the LDS, per normalization
|
||||
unit. This is unused and expected to be zero in most configurations for
|
||||
modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Panel Config:
|
||||
Cache Hit Rate: The percent of L1I requests that hit [#l1i-cache]_ on a previously
|
||||
loaded line the cache. Calculated as the ratio of the number of L1I requests
|
||||
that hit over the number of all L1I requests.
|
||||
L1I-L2 Bandwidth Utilization: |-
|
||||
L1I-L2 Bandwidth Utilization: >-
|
||||
The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth
|
||||
achieved. Calculated as the ratio of the total number of requests from the
|
||||
L1I to the L2 cache over the total L1I-L2 interface cycles.
|
||||
|
||||
+3
-3
@@ -154,7 +154,7 @@ Panel Config:
|
||||
sL1D-L2 BW Utilization: The percentage of the peak theoretical sL1D - L2 interface
|
||||
bandwidth acheived. Calculated as total number of bytes read from, written to,
|
||||
or atomically updated across the sL1D - L2 interface.
|
||||
sL1D-L2 BW: |-
|
||||
sL1D-L2 BW: >-
|
||||
The total number of bytes read from, written to, or atomically updated
|
||||
across the sL1D\u2194L2 interface, divided by total duration. Note that sL1D
|
||||
writes and atomics are typically unused on current CDNA accelerators, so
|
||||
@@ -164,7 +164,7 @@ Panel Config:
|
||||
unit.
|
||||
Hits: The total number of sL1D requests that hit on a previously loaded cache
|
||||
line, per normalization unit.
|
||||
Misses - Non Duplicated: |-
|
||||
Misses - Non Duplicated: >-
|
||||
The total number of sL1D requests that missed on a cache line that was
|
||||
not already pending due to another request, per normalization unit.
|
||||
Misses- Duplicated: The total number of sL1D requests that missed on a cache line
|
||||
@@ -187,6 +187,6 @@ Panel Config:
|
||||
unit.
|
||||
Write Req: The total number of write requests from sL1D to the L2, per normalization
|
||||
unit. Typically unused on current CDNA accelerators.
|
||||
Stall Cycles: |-
|
||||
Stall Cycles: >-
|
||||
The total number of cycles the sL1D\u2194L2 interface was stalled, per
|
||||
normalization unit.
|
||||
|
||||
+1
-1
@@ -501,7 +501,7 @@ Panel Config:
|
||||
per normalization unit.
|
||||
Translation Misses: The total number of translation requests that missed in the
|
||||
UTCL1 due to translation not being present in the cache, per normalization unit.
|
||||
Permission Misses: |-
|
||||
Permission Misses: >-
|
||||
The total number of translation requests that missed in the UTCL1 due
|
||||
to a permission error, per normalization unit. This is unused and expected
|
||||
to be zero in most configurations for modern CDNA\u2122 accelerators.
|
||||
|
||||
+1
-1
@@ -706,7 +706,7 @@ Panel Config:
|
||||
requests are only considered atomic by Infinity Fabric if they are targeted
|
||||
at non-write-cacheable memory, such as fine-grained memory allocations or uncached
|
||||
memory allocations on the MI2XX.
|
||||
Read Stall: |-
|
||||
Read Stall: >-
|
||||
The ratio of the total number of cycles the L2-Fabric interface was
|
||||
stalled on a read request to any destination (local HBM, remote PCIe\xAE
|
||||
connected accelerator or CPU, or remote Infinity Fabric connected accelerator
|
||||
|
||||
Ссылка в новой задаче
Block a user