Fixing some accumulate metrics (#1089)

* Fixing some accumulate metrics

* Fixing some more accumulate metrics

---------

Co-authored-by: Benjamin Welton <bewelton@amd.com>
This commit is contained in:
Giovanni Lenzi Baraldi
2024-10-03 13:44:31 -03:00
کامیت شده توسط GitHub
والد 1d6a0b5c80
کامیت fa7fc7ec5d
2فایلهای تغییر یافته به همراه18 افزوده شده و 18 حذف شده
+2
مشاهده پرونده
@@ -108,4 +108,6 @@ Full documentation for ROCprofiler-SDK is available at [Click Here](source/docs/
- Creation of subdirection when rocprofv3 `--output-file` contains a folder path - Creation of subdirection when rocprofv3 `--output-file` contains a folder path
- Fix misaligned stores (undefined behavior) for buffer records - Fix misaligned stores (undefined behavior) for buffer records
- Fix crash when only scratch reporting is enabled - Fix crash when only scratch reporting is enabled
- Fixed MeanOccupancy* metrics
- Fix aborted-app validation test to properly check for hipExtHostAlloc command now that it is supported - Fix aborted-app validation test to properly check for hipExtHostAlloc command now that it is supported
@@ -461,8 +461,8 @@ GpuUtil:
description: 'Unit: percent' description: 'Unit: percent'
InstrFetchLatency: InstrFetchLatency:
architectures: architectures:
gfx90a: gfx942/gfx941/gfx940/gfx90a:
expression: SQ_ACCUM_PREV_HIRES/SQ_IFETCH expression: accumulate(SQ_IFETCH_LEVEL, HIGH_RES)/SQ_IFETCH
description: 'Unit: cycles' description: 'Unit: cycles'
L1iCacheHitRate: L1iCacheHitRate:
architectures: architectures:
@@ -508,8 +508,8 @@ LdsBankConflict:
description: 'Unit: conflicts/access' description: 'Unit: conflicts/access'
LdsLatency: LdsLatency:
architectures: architectures:
gfx90a: gfx942/gfx941/gfx940/gfx90a/gfx10/gfx1010/gfx1030/gfx1031/gfx1032/gfx11/gfx1100/gfx1101/gfx1102:
expression: SQ_ACCUM_PREV_HIRES/SQ_INSTS_LDS expression: accumulate(SQ_INST_LEVEL_LDS, HIGH_RES)/SQ_INSTS_LDS
description: 'Unit: cycles' description: 'Unit: cycles'
LdsPipeIssueUtil: LdsPipeIssueUtil:
architectures: architectures:
@@ -528,19 +528,17 @@ MAX_WAVE_SIZE:
description: Max wave size constant description: Max wave size constant
MeanOccupancyPerActiveCU: MeanOccupancyPerActiveCU:
architectures: architectures:
gfx10/gfx1010/gfx1030/gfx1031/gfx1032:
expression: GRBM_COUNT*0+SQ_LEVEL_WAVES*0+SQ_ACCUM_PREV*4/SQ_BUSY_CYCLES/CU_NUM
gfx942/gfx941/gfx940/gfx90a: gfx942/gfx941/gfx940/gfx90a:
expression: SQ_LEVEL_WAVES*0+SQ_ACCUM_PREV_HIRES*4/SQ_BUSY_CYCLES/CU_NUM expression: accumulate(SQ_LEVEL_WAVES, LOW_RES)/SQ_BUSY_CU_CYCLES
gfx11/gfx1100/gfx1101/gfx1102:
expression: SQ_WAVE_CYCLES/SQ_BUSY_CYCLES
description: Mean occupancy per active compute unit. description: Mean occupancy per active compute unit.
MeanOccupancyPerCU: MeanOccupancyPerCU:
architectures: architectures:
gfx10/gfx1010/gfx1030/gfx1031/gfx1032: gfx10/gfx1010/gfx1030/gfx1031/gfx1032/gfx90a/gfx942/gfx941/gfx940:
expression: GRBM_COUNT*0+SQ_LEVEL_WAVES*0+SQ_ACCUM_PREV/GRBM_GUI_ACTIVE/CU_NUM expression: accumulate(SQ_LEVEL_WAVES, HIGH_RES)/reduce(GRBM_GUI_ACTIVE,max)/CU_NUM
gfx90a: gfx11/gfx1100/gfx1101/gfx1102:
expression: SQ_LEVEL_WAVES*0+SQ_ACCUM_PREV_HIRES/GRBM_GUI_ACTIVE/CU_NUM expression: SQ_WAVE_CYCLES/GRBM_GUI_ACTIVE/CU_NUM
gfx942/gfx941/gfx940:
expression: reduce(SQ_LEVEL_WAVES,sum)*0+reduce(SQ_ACCUM_PREV_HIRES,sum)/reduce(GRBM_GUI_ACTIVE,sum)/CU_NUM
description: Mean occupancy per compute unit. description: Mean occupancy per compute unit.
MemUnitBusy: MemUnitBusy:
architectures: architectures:
@@ -1063,7 +1061,7 @@ SQ_BUSY_CU_CYCLES:
with units in quad-cycles(4 cycles). with units in quad-cycles(4 cycles).
SQ_BUSY_CYCLES: SQ_BUSY_CYCLES:
architectures: architectures:
gfx942/gfx941/gfx10/gfx1010/gfx1030/gfx1031/gfx11/gfx1032/gfx1102/gfx1100/gfx1101/gfx940/gfx90a: gfx942/gfx941/gfx940/gfx90a/gfx10/gfx1010/gfx1030/gfx1031/gfx11/gfx1032/gfx1102/gfx1100/gfx1101:
block: SQ block: SQ
event: 3 event: 3
description: Number of clock cycles there are active waves in a shader engine (as reported by the distributed description: Number of clock cycles there are active waves in a shader engine (as reported by the distributed
@@ -2061,8 +2059,8 @@ ScaPipeIssueUtil:
description: 'Unit: percent' description: 'Unit: percent'
SmemLatency: SmemLatency:
architectures: architectures:
gfx90a: gfx942/gfx941/gfx940/gfx90a:
expression: SQ_ACCUM_PREV_HIRES/SQ_INSTS_SMEM_NORM expression: accumulate(SQ_INST_LEVEL_SMEM, HIGH_RES)/SQ_INSTS_SMEM_NORM
description: 'Unit: cycles' description: 'Unit: cycles'
SpiUtil: SpiUtil:
architectures: architectures:
@@ -4008,8 +4006,8 @@ ValuPipeIssueUtil:
description: 'Unit: percent' description: 'Unit: percent'
VmemLatency: VmemLatency:
architectures: architectures:
gfx90a: gfx942/gfx941/gfx940/gfx90a:
expression: SQ_ACCUM_PREV_HIRES/SQ_INSTS_VMEM expression: accumulate(SQ_INST_LEVEL_VMEM, HIGH_RES)/SQ_INSTS_VMEM
description: 'Unit: cycles' description: 'Unit: cycles'
VmemPipeIssueUtil: VmemPipeIssueUtil:
architectures: architectures: