f9aa7be97c
* Add MI 350 hardware information
* Refactor MI GPU YAML file and corresponding interface
* Add SoC file for gfx950 architecture
* Add analysis report configs for MI 350 containing existing metrics
* Add placeholder None valued metrics for previous architectures to make
baseline comparison work
* Enable testing on MI 350
* Analysis config metric changes
- SPI changes
- Update metric formula for default SPI pipe counter
- Use efficiently collected pipe wise SPI counters
- Add SPI Wave Occupancy
- Add Scheduler-Pipe Wave Utilization
- Update formula for VGPR Writes
- Add Scheduler-Pipe FIFO Full Rate
- CPC changes
- Add CPC SYNC FIFO Full Rate
- Add CPC CANE Stall Rate
- Add CPC ADC Utilization
- SQ changes
- Add VALU co-issue efficiency
- Add F6F4 datatype metrics
- Update formula for total FLOPs by adding F6F4 counters
- Add LDS STORE / LOAD / ATOMIC metrics
- Add LDS STORE / LOAD / ATOMIC bandwidth
- Add LDS FIFO and TA ADDR / CMD / DATA FIFO full rates
* Collect TCP_TCP_LATENCY_sum only for gfx950 (MI 350)
* Do not inject SQ_ACCUM_PREV_HIRES unnecesarily
* Do not hardcode memory and shader clock speeds
* Write num_hbm_channels to sysinfo.csv instead of hbm_bw while profiling
* Move generate sysinfo.csv to pre processing step of profiling
* Add warnings to use --specs-correction for missing sysinfo.csv values during analysis phase
* Update CHANGELOG
* Analysis phase warning to use --specs-correction when needed
1.3 KiB
1.3 KiB
| 1 | Dispatch_ID | GPU_ID | Queue_ID | PID | TID | Grid_Size | Workgroup_Size | LDS_Per_Workgroup | Scratch_Per_Workitem | Arch_VGPR | Accum_VGPR | SGPR | Wave_Size | Kernel_Name | Start_Timestamp | End_Timestamp | Correlation_ID | Kernel_ID | SPI_RA_VGPR_SIMD_FULL_CSN | SPI_VWC1_VDATA_VALID_WR | SQC_DCACHE_REQ_READ_2 | SQ_ACTIVE_INST_ANY | SQ_ACTIVE_INST_MISC | SQ_INSTS_BRANCH | SQ_INSTS_LDS_STORE_BANDWIDTH | SQ_INSTS_VALU_CVT | SQ_LDS_ATOMIC_RETURN | SQ_LDS_MEM_VIOLATIONS | TCC_MISS_sum | TCC_PROBE_sum | TCC_REQ_sum | TCC_WRITEBACK_sum | TCP_TCC_CC_WRITE_REQ_sum | TCP_TOTAL_ACCESSES_sum | TCP_TOTAL_ATOMIC_WITHOUT_RET_sum | TCP_TOTAL_READ_sum |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 0 | 0 | 1 | 15793 | 15793 | 1048576 | 256 | 0 | 0 | 4 | 0 | 16 | 64 | vecCopy(double*, double*, double*, int, int) | 1248228890145284 | 1248228890166603 | 1 | 13 | 0.0 | 32768.0 | 32768.0 | 278528.0 | 32768.0 | 16384.0 | 0.0 | 0.0 | 0.0 | 0.0 | 131138.0 | 0.0 | 198002.0 | 65596.0 | 0.0 | 2097152.0 | 0.0 | 1048576.0 |
| 3 | 1 | 0 | 1 | 15793 | 15793 | 1048576 | 256 | 0 | 0 | 4 | 0 | 16 | 64 | vecCopy(double*, double*, double*, int, int) | 1248228890710662 | 1248228890727781 | 2 | 13 | 0.0 | 32768.0 | 32768.0 | 278528.0 | 32768.0 | 16384.0 | 0.0 | 0.0 | 0.0 | 0.0 | 131120.0 | 0.0 | 197984.0 | 65596.0 | 0.0 | 2097152.0 | 0.0 | 1048576.0 |
| 4 | 2 | 0 | 1 | 15793 | 15793 | 1048576 | 256 | 0 | 0 | 4 | 0 | 16 | 64 | vecCopy(double*, double*, double*, int, int) | 1248228891192003 | 1248228891208963 | 3 | 13 | 0.0 | 32768.0 | 32768.0 | 278528.0 | 32768.0 | 16384.0 | 0.0 | 0.0 | 0.0 | 0.0 | 131112.0 | 0.0 | 197976.0 | 65596.0 | 0.0 | 2097152.0 | 0.0 | 1048576.0 |