Files
rocm-systems/tests/validate-causal-json.py
T
Jonathan R. Madsen 9618ddefba Causal profiling (#229)
* Addition of basic structure

* Reworked categories

* More causal integration additions

* Causal implementation

* Update examples

* delete virtual_speedup files

* Update perfetto submodule to v31.0

* Update dyninst submodule

* Update timemory submodule

* ElfUtils build for libdw

* OMNITRACE_LIKELY and OMNITRACE_UNLIKELY

* Update common lib join

* Examples updates for causal profiling

* config updates with causal options

- OMNITRACE_CAUSAL_FIXED_LINE
- OMNITRACE_CAUSAL_FIXED_SPEEDUP
- OMNITRACE_CAUSAL_FILE
- OMNITRACE_CAUSAL_BINARY_SCOPE
- OMNITRACE_CAUSAL_SOURCE_SCOPE
- version info in banner
- support increments in parse_numeric_range
- fix occasional deadlock in first call to get_config

* PTL general task group

* Always include PID in debug/verbose messages

* Add blocking/unblocking gotchas to runtime init bundle

* CausalState

* thread_data updates

- generic component_bundle_cache

* Improve handling of causal in category_region

* components updates

- backtrace_causal component
- backtrace::get_data member func
- decrease ignore_depth in backtrace::sample(int)
- handle "omnitrace_main" in backtrace::filter_and_patch(...)
- tweak internal thread state scope for pthread_mutex_gotcha wrappers

* simplify tracing get_instrumentation_bundles usage

* sampling updates

- include backtrace_causal component
- disable backtrace_metrics if using causal and not using perfetto
- disable backtrace and backtrace_timestamp when using causal
- post_process_causal

* causal updates

- more checks in blocking_gotcha and unblocking_gotcha start/stop
- miscellaneous overhaul of data
- experiment update

* Remove virtual speedup

* libomnitrace code_object

* causal-profiling test

* libomnitrace library.cpp updates

- handle causal profiling
- fini_bundle

* Disable causal profiling by default

* Updated causal code and example

- example: three execution variants: cpu + rng, cpu, rng
- example: three instrumentation variants: none, omni, coz
- fix blocking gotcha credit
- rework perform_experiment_impl
- get_eligible_address_ranges
- compute_eligible_lines
- support fixed lines/speedups/functions
- update selected_entry to support function mode
- fix causal::delay
- experiment updates

* omnitrace_progress / omnitrace_user_progress

- with accompanying omnitrace_annotated_progress / omnitrace_user_annotated_progress

* Update timemory submodule

* CausalMode

- mode indicated whether causal predictions source be at line-level or function-level

* code_object, config, runtime, sampling, thread_data

- code_object: address_range
- code_object: basic::line_info serialize(), name(), hash()
- config updates
- two signals for causal sampling
- thread_data init fixes

* pthread updates

- pthread_create_gotcha processes delays
- pthread_mutex_gotcha does not wrap pthread_join in causal mode

* backtrace_causal update

- dynamic delay period stats

* main wrapper uses basename of argv[0]

* update elfio submodule

* perf support (currently unused)

* Fix experiment JSON serialization

- static_vector.hpp (unused)

* causal executable + config options updates

- omnitrace-causal exe simplifies running multiple causal configs
- changed the causal config option names

* Support both throughput and latency points

* process-causal-json.py script

- will be used later for testing

* stable_vector

* Rework thread_data

* Improve omnitrace-causal exe

- better verbosity handling
- correct diagnosis of status for child process
- execvpe when only one iteration (debugging)

* Update timemory submodule

* exe --version

- omnitrace, omnitrace-avail, and omnitrace-sample all support --version on command-line

* OMNITRACE_INTERNAL_API + OMNITRACE_{LIKELY,UNLIKELY}

* omnitrace-causal cmake format

* omnitrace config update

- OMNITRACE_CAUSAL_FILE_CLOBBER

* custom exception

- wraps STL exception and gets stacktrace during construction

* exit_gotcha supports _Exit

* use global construct_on_init + max threads

- add some safety when exceeding max # of threads

* update code_object binary filter

- exclude dyninst and tbbmalloc library

* containers: c_array, static_vector, stable_vector

- moved utility::c_array to container::c_array
- created static_vector: std::vector bound to std::array
- created stable_vector: vector with stable references

* grow thread_data when new thread created

* causal updates

- data: improve compute_eligible_lines to ignore lambdas
- data: use new thread_data
- delay: use new thread_data
- experiment: properly support latency points
- experiment: support file clobber
- experiment: ensure non-zero experiment time
- progress_point: use new thread_data
- backtrace_causal: use new thread_data

* Update causal-profiling tests

* fix omnitrace-causal backslash escaping

* process-causal-json script

* restructure causal implementation

- update verbose messages for omnitrace-causal diagnose_status
- migrated causal implementation in sampling.cpp to causal/sampling.cpp
- OMNITRACE_USE_CAUSAL does not require OMNITRACE_USE_SAMPLING
- added Mode::Causal
- causal sampling uses same signals as regular sampling
- moved tracing::thread_init to implementation file
- combined tracing::thread_init and tracing::thread_init_sampling
- added causal/components folder
- pthread_create_gotcha::wrapper_config
- omnitrace_preload checks OMNITRACE_USE_CAUSAL
  - updates mode accordingly

* update timemory submodule

* update timemory submodule

* causal example updates

- causal for lulesh

* perf code + utility - helpers

- relocated causal perf code
- placement new when generating unique ptr trait for potentially allocating during sampling
- additions to utility header
- removed previously added helpers.hpp

* update timemory submodule

* Default env variables for omnitrace-causal

- activate OMNITRACE_USE_KOKKOSP, etc.

* update stable_vector and static_vector

- static vector can use atomic for size tracking for thread-safe situations

* update causal example header

- CAUSAL_PROGRESS_NAMED
- use CAUSAL_ prefix for some macros

* Tweak lulesh example

- use CAUSAL_PROGRESS instead of CAUSAL_BEGIN and CAUSAL_END

* omnitrace-sample support for causal mode

- set OMNITRACE_USE_SAMPLING to off when OMNITRACE_MODE=causal

* refactor and cleanup code_object

- scope filter
- fixes to address_range

* overhaul causal data + causal config options

- full support for function and line mode
- support static vector of instruction pointers
- improve line info mapping resolution
- remove thread-locality from miscellanous functions where unnecessary
- causal options for {binary,source,function,fileline} exclusion

* causal experiment, sampling, and backtrace updates

- is_selected + unwind address array
- experiment warning about progress points
- increased buffer size for backtrace_casual sampler
- backtrace_causal only stores IP addresses instead of full unwind info

* category_region updates

- minor refactor
- local_category_region::mark

* Update causal tests

* Bump version to 1.8.0

* omnitrace-causal args + CLOBBER -> RESET

- renamed OMNITRACE_CAUSAL_FILE_CLOBBER to OMNITRACE_CAUSAL_FILE_RESET
- updated omnitrace-causal exe to support recently added configuration options
- other miscellaneous tweaks to data.cpp, experiment.cpp, and sampling.cpp

* Refactor causal and code_object

- code_object.hpp and code_object.cpp moved into binary folder
- causal components namespaced into omnitrace::causal::component
- moved sample_data out of backtrace_causal and into own file
- renamed backtrace_causal to causal::component::backtrace

* preload omnitrace_init + OMNITRACE_DEBUG_MARK

- env OMNITRACE_DEBUG_MARK
- fix omnitrace_init call when LD_PRELOAD-ing omnitrace

* Fix fileline support + line-info output names + experiment log

- line-info log files are prefixed with experiment name
- don't print experiment duration when E2E
- account for fileline scope in analysis

* KokkosP: OMNITRACE_KOKKOSP_NAME_LENGTH_MAX

- config option to limit the name of kokkos tool callbacks
- remove [kokkos] from KokkosP names

* Update causal example

- minor tweaks to decrease probability of overlapping regions in binary

* omnitrace-causal update

- prefix N / Ntot in environment printout

* Miscellaneous updates

- causal::finish_experimenting()
- OMNITRACE_CAUSAL_RANDOM_SEED
- KokkosP causal updates
  - exclude some callbacks, make some callbacks unique, etc.
- address_range::operator+=(address_range)
- combine contiguous ranges in binary/analysis.cpp when file, func, line is same and address range is contiguous
- bfd_line_info reads inline info
- wait for perform_experiment_impl to complete
- causal::delay updates
  - delay::process checks if experiment is active
  - uses threading::get_id()
- experiment scales duration up for larger speedup experiments
- line info samples includes excluded lines
- sampler uses CLOCK_REALTIME
- blocking_gotcha updates
  - is no longer fully static
  - adds audit routine which sets the postblock value to zero if try/timed routine fails
- category::host was added to causal_throughput_categories_t
- pthread_create_gotcha sets new threads local parent delay
  - was using internal value, now uses sequent value

* Causal improvements to KokkosP

* Updates to experiment time scaling

- use stats instead of just max

* binary/link_map.{hpp,cpp}

* update process-causal-json.py

* Folded fileline scope into source scope

* Update documentation

- Add documentation for causal profiling
- Replace 'Omnitrace' with 'OmniTrace' everywhere

* Update causal-helpers.cmake + omnitrace-testing.cmake

- split tests/CMakeLists.txt partially into omnitrace-testing.cmake

* omnitrace/causal.h

- OMNITRACE_CAUSAL_PROGRESS
- OMNITRACE_CAUSAL_PROGRESS_NAMED
- OMNITRACE_CAUSAL_BEGIN
- OMNITRACE_CAUSAL_END

* selected_entry + remove default filters for lambdas and operator()

- selected entry stores range and binary load address

* update process-causal-json.py

* format examples/lulesh/CMakeLists.txt

* causal-helpers find_package(Threads)

* OMNITRACE_KOKKOSP_KERNEL_LOGGER

- was OMNITRACE_KOKKOS_KERNEL_LOGGER

* quiet find of coz-profiler

* Fix rocm_smi exception handling

* Update timemory submodule (binutils)

- fix binutls compile error on some systems
- bump binutils to v2.40

* Fix miscellaneous tests

* OMNITRACE_KOKKOSP_PREFIX

* revert rocm_smi handling

* ElfUtils updates

- default to download version 0.188
- add -Wno-error=null-dereference due to GCC 12 compiler error

* Update causal example

* Remove OMNITRACE_VERBOSE from global workflow envs

* Reliable causal test

* disable compilation of causal perf files

* Remove set_current_selection with unwind stack

* update timemory submodule

* fix for segfault on bionic

- locking in TLS dtor was causing segfault

* remove experiment::is_selected(unwind_stack_t)

* update default init of selected_entry

* Fix for when IP is not offset by load address

* Update CMakeLists.txt

* Miscellaneous updates

- OMNITRACE_WARNING_OR_CI_THROW
- OMNITRACE_REQUIRE
- OMNITRACE_PREFER
- fixed issues with no ASLR
-  added load address variable and ipaddr() func to basic/bfd line info
- removed get_basic() from dwarf_line_info
- TIMEMORY_PREFER -> OMNITRACE_PREFER
- removed previously added binary_address and range variables from selected_entry

* Removed superfluous CausalState

* Additional causal tests (lulesh + kokkos)

* filter, prefer, analysis ASLR handling

- removed default filter on cold functions
- fixed OMNITRACE_PREFER
- fixed analysis ASLR handling

* Tweak line-info output

* Removed some superfluous code

- causal/delay
- causal/selected_entry

* Exclude main.cold in function mode

* Update validate-perfetto-proto.py

- account for occasional http errors

* Add sampling test disabling tmp files

* argparser for process-causal-json

- support validation
- support filtering

* Avoid pthread_{lock,unlock} in sampling offload

- use homemade atomic_mutex/atomic_lock since contention will be low and using pthread tools might trigger our wrappers

* Rename process-causal-json.py

- validate-causal-json.py

* rework omnitrace_add_causal_test

- capable of performing validation
- added validation tests

* Fix kokkosp_begin_deep_copy + causal

* Tweak address range in bfd_line_info::read_pc

* Tweak analysis and data IP handling

- look for gaps

* Disable scaling experiment time by speedup

* Revert change in max threads during CI

* binary updates

- significant overhaul of binary analysis implementation
- removed "basic_line_info" and "bfd_line_info" in lieu of "symbol" class
  - symbol class has basic BFD info + vector of inlines + vector of dwarf info

* Updated causal to use new binary analysis

- Fix symbol.cpp includes

* Updated formatting target

- include *.cmake files

* Updated causal tests

- causal tests should be stable now

* Update timemory and dyninst submodules

- TPLs are stripped + built w/o debug info

* Increase tolerance for causal validation speedups

- higher speedups have more variance (increased to +/- 5 from 3)

* Support causal output for MPI

- i.e. tag with MPI rank

* omnitrace-causal launcher argument

* improve experiment sampling output

* causal data updates

- call compute lines once
- fixed filtered cached binary info
- debugging info when experiment fails to start

* Tweaked causal validation tests

* dwarf_entry ranges

* CI updates

- increase max threads to 64

* Tweak causal E2E validation tests

- more threads
- shorter thread runtime
- more iterations

* Fix shadowed variable

* fix symbol read_bfd last PC calculation

* fix maybe-uninitialized warning

* omnitrace-causal launcher update

- only inject "omnitrace-causal --" once
- throw error if no matches found

* Update causal profiling docs for launcher

* fix address range boundaries
2023-01-24 18:53:23 -06:00

404 řádky
12 KiB
Python
Spustitelný soubor

#!/usr/bin/env python3
import os
import re
import sys
import json
import argparse
from collections import OrderedDict
num_stddev = 1
def mean(_data):
return sum(_data) / float(len(_data)) if len(_data) > 0 else 0.0
def stddev(_data):
if len(_data) == 0:
return 0.0
_mean = mean(_data)
_variance = sum([((x - _mean) ** 2) for x in _data]) / float(len(_data))
return _variance**0.5
class validation(object):
def __init__(self, _exp_re, _pp_re, _virt, _expected, _tolerance):
self.experiment_filter = re.compile(_exp_re)
self.progress_pt_filter = re.compile(_pp_re)
self.virtual_speedup = int(_virt)
self.program_speedup = float(_expected)
self.tolerance = float(_tolerance)
def validate(self, _exp_name, _pp_name, _virt_speedup, _prog_speedup):
if (
not re.search(self.experiment_filter, _exp_name)
or not re.search(self.progress_pt_filter, _pp_name)
or _virt_speedup != self.virtual_speedup
):
return None
return _prog_speedup >= (
self.program_speedup - self.tolerance
) and _prog_speedup <= (self.program_speedup + self.tolerance)
class experiment_data(object):
def __init__(self, _speedup):
self.speedup = _speedup
self.duration = []
def __iadd__(self, _val):
self.duration += [float(_val)]
def __len__(self):
return len(self.duration)
def __eq__(self, rhs):
return self.speedup == rhs.speedup
def __neq__(self, rhs):
return not self == rhs
def __lt__(self, rhs):
return self.speedup < rhs.speedup
def mean(self):
return mean(self.duration)
def stddev(self):
return stddev(self.duration)
class line_speedup(object):
def __init__(self, _name="", _prog="", _exp_data=None, _exp_base=None):
self.name = _name
self.prog = _prog
self.data = _exp_data
self.base = _exp_base
def virtual_speedup(self):
if self.data is None or self.base is None:
return 0.0
return self.data.speedup
def compute_speedup(self):
if self.data is None or self.base is None:
return 0.0
return ((self.base.mean() - self.data.mean()) / self.base.mean()) * 100
def compute_speedup_stddev(self):
if self.data is None or self.base is None:
return 0.0
_data = []
_base = self.base.mean()
for ditr in self.data.duration:
_data += [((_base - ditr) / _base) * 100]
return stddev(_data)
def get_name(self):
return ":".join(
[
os.path.basename(x) if os.path.isfile(x) else x
for x in self.name.split(":")
]
)
def __str__(self):
if self.data is None or self.base is None:
return f"{self.name}"
_line_speedup = self.compute_speedup()
_line_stddev = (
float(num_stddev) * self.compute_speedup_stddev()
) # 3 stddev == 99.87%
_name = self.get_name()
return f"[{_name}][{self.prog}][{self.data.speedup:3}] speedup: {_line_speedup:6.1f} +/- {_line_stddev:6.2f} %"
def __eq__(self, rhs):
return (
self.name == rhs.name
and self.prog == rhs.prog
and self.data == rhs.data
and self.base == rhs.base
)
def __neq__(self, rhs):
return not self == rhs
def __lt__(self, rhs):
if self.name != rhs.name:
return self.name < rhs.name
elif self.prog != rhs.prog:
return self.prog < rhs.prog
elif self.data != rhs.data:
return self.data < rhs.data
elif self.base != rhs.base:
return self.base < rhs.base
return False
class experiment_progress(object):
def __init__(self, _data):
self.data = _data
def get_impact(self):
"""
speedup_c = [x.compute_speedup() for x in self.data]
speedup_v = [x.virtual_speedup() for x in self.data]
impact = []
for i in range(len(self.data) - 1):
x = speedup_v[i + 1] - speedup_v[i]
y_low = speedup_c[i]
y_upp = speedup_c[i + 1]
a_low = x * min([y_low, y_upp])
a_high = 0.5 * x * (max([y_low, y_upp]) - min([y_low, y_upp]))
impact += [a_low + a_high]
"""
impact = [x.compute_speedup() for x in self.data]
return [sum(impact), mean(impact), stddev(impact)]
def __len__(self):
return len(self.data)
def __str__(self):
_impact_v = self.get_impact()
_name = self.data[0].get_name()
_prog = self.data[0].prog
_impact = [
f"[{_name}][{_prog}][sum] impact: {_impact_v[0]:6.1f}",
f"[{_name}][{_prog}][avg] impact: {_impact_v[1]:6.1f} +/- {_impact_v[2]:6.2f}",
]
return "\n".join([f"{x}" for x in self.data] + _impact)
def __lt__(self, rhs):
self.data.sort()
return self.get_impact()[0] < rhs.get_impact()[0]
def find_or_insert(_data, _value):
if _value not in _data:
_data[_value] = experiment_data(_value)
return _data[_value]
def process_data(data, _data, args):
if not _data:
return data
_selection_filter = re.compile(args.experiments)
_progresspt_filter = re.compile(args.progress_points)
for record in _data["omnitrace"]["causal"]["records"]:
for exp in record["experiments"]:
_speedup = exp["virtual_speedup"]
_duration = exp["duration"]
_file = exp["selection"]["info"]["file"]
_line = exp["selection"]["info"]["line"]
_func = exp["selection"]["info"]["dfunc"]
_sym_addr = exp["selection"]["symbol_address"]
_selected = ":".join([_file, f"{_line}"]) if _sym_addr == 0 else _func
if not re.search(_selection_filter, _selected):
continue
if _selected not in data:
data[_selected] = {}
for pts in exp["progress_points"]:
_name = pts["name"]
if not re.search(_progresspt_filter, _name):
continue
if _name not in data[_selected]:
data[_selected][_name] = {}
if "delta" in pts:
_delt = pts["delta"]
if _delt > 0:
itr = find_or_insert(data[_selected][_name], _speedup)
itr += float(_duration) / float(_delt)
else:
_diff = pts["arrival"] - pts["departure"] + 1
_rate = pts["arrival"] / float(_duration)
if _rate > 0:
itr = find_or_insert(data[_selected][_name], _speedup)
itr += float(_diff) / float(_rate)
else:
_delt = pts["laps"]
if _delt > 0:
itr = find_or_insert(data[_selected][_name], _speedup)
itr += float(_duration) / float(_delt)
return data
def compute_speedups(_data, args):
data = {}
for selected, pitr in _data.items():
if selected not in data:
data[selected] = {}
for progpt, ditr in pitr.items():
data[selected][progpt] = OrderedDict(sorted(ditr.items()))
from os.path import dirname
ret = []
for selected, pitr in _data.items():
for progpt, ditr in pitr.items():
if 0 not in ditr.keys():
# print(f"missing baseline data for {progpt} in {selected}...")
continue
_baseline = ditr[0].mean()
for speedup, itr in ditr.items():
if len(args.speedups) > 0 and speedup not in args.speedups:
continue
if speedup != itr.speedup:
raise ValueError(f"in {selected}: {speedup} != {itr.speedup}")
_val = line_speedup(selected, progpt, itr, ditr[0])
ret.append(_val)
ret.sort()
_last_name = None
_last_prog = None
result = []
for itr in ret:
if itr.name != _last_name or itr.prog != _last_prog:
result.append([])
result[-1].append(itr)
_last_name = itr.name
_last_prog = itr.prog
_data = []
for itr in result:
_data.append(experiment_progress(itr))
_data.sort()
return _data
def get_validations(args):
data = []
_len = len(args.validate)
if _len == 0:
return data
elif _len % 5 != 0:
raise ValueError(
"validation requires format: {experiment regex} {progress-point regex} {virtual-speedup} {expected-speedup} {tolerance} (i.e. 5 args per validation. There are {} extra/missing arguments".format(
_len % 5
)
)
v = args.validate
for i in range(int(_len / 5)):
off = 5 * i
data.append(
validation(v[off + 0], v[off + 1], v[off + 2], v[off + 3], v[off + 4])
)
return data
def main():
import argparse
parser = argparse.ArgumentParser()
parser.add_argument(
"-e", "--experiments", type=str, help="Regex for experiments", default=".*"
)
parser.add_argument(
"-p",
"--progress-points",
type=str,
help="Regex for progress points",
default=".*",
)
parser.add_argument(
"-n", "--num-points", type=int, help="Minimum number of data points", default=5
)
parser.add_argument(
"-i", "--input", type=str, nargs="*", help="Input file(s)", required=True
)
parser.add_argument(
"-s",
"--speedups",
type=int,
help="List of speedup values to report",
nargs="*",
default=[],
)
parser.add_argument(
"-d",
"--stddev",
type=int,
help="Number of standard deviations to report",
default=1,
)
parser.add_argument(
"-v",
"--validate",
type=str,
nargs="*",
help="Validate speedup: {experiment regex} {progress-point regex} {virtual-speedup} {expected-speedup} {tolerance}",
default=[],
)
args = parser.parse_args()
num_stddev = args.stddev
num_speedups = len(args.speedups)
if num_speedups > 0 and args.num_points > num_speedups:
args.num_points = num_speedups
data = {}
for inp in args.input:
with open(inp, "r") as f:
inp_data = json.load(f)
data = process_data(data, inp_data, args)
results = compute_speedups(data, args)
for itr in results:
if len(itr) < args.num_points:
continue
print("")
print(f"{itr}")
validations = get_validations(args)
expected_validations = len(validations)
correct_validations = 0
if expected_validations > 0:
print(f"\nPerforming {expected_validations} validations...\n")
for eitr in results:
_experiment = eitr.data[0].get_name()
_progresspt = eitr.data[0].prog
for ditr in eitr.data:
_virt_speedup = ditr.virtual_speedup()
_prog_speedup = ditr.compute_speedup()
for vitr in validations:
_v = vitr.validate(
_experiment, _progresspt, _virt_speedup, _prog_speedup
)
if _v is None:
continue
if _v is True:
correct_validations += 1
else:
sys.stderr.write(
f" [{_experiment}][{_progresspt}][{_virt_speedup}] failed validation: {_prog_speedup:8.3f} != {vitr.program_speedup} +/- {vitr.tolerance}\n"
)
if expected_validations != correct_validations:
sys.stderr.flush()
sys.stderr.write(
f"\nCausal profiling predictions not validated. Expected {expected_validations}, found {correct_validations}\n"
)
sys.stderr.flush()
sys.exit(-1)
elif expected_validations > 0:
print(f"Causal profiling predictions validated: {expected_validations}")
if __name__ == "__main__":
main()