Support for tracing mutex locking (#52)

* Parallel overhead example with locks

* Support tracing mutex locking + more

- support wrapping pthread_mutex_lock
- support wrapping pthread_mutex_unlock
- support wrapping pthread_mutex_trylock
- get_perfetto_combined_traces setting
- OMNITRACE_TRACE_THREAD_LOCKS option
- ThreadState
- critical trace includes queue id
- enabled/disabled settings in timemory
- fix OMNITRACE_TIMEMORY_COMPONENTS
- fix reading config
- fix setting categories
- applied ThreadState::Internal in various places
- utility::get_filled_array
- utility::get_reserved_vector
- utility::get_thread_index
- fork_gotcha messages about forks
- split out some pthread_gotcha functionality into pthread_create_gotcha
- handle queue id in roctracer callbacks

* Update timemory and PTL submodules

* Misc CMake updates

- Includes fix to omnitrace-static-lib{gcc,stdcxx}

* Misc cleanup to pthread_mutex_gotcha and backtrace

* Fix to duplicate field in module_function json

* Improvement to debug messages

* omnitrace-dl and common improvements

- tweak to delimit
- common::ignore message
- common::join quoting of strings
- omnitrace_set_env ignores if inited and active
- omnitrace_set_mpi ignores if inited and active

* nsync for transpose example

* Fix to thread_deleter<void> functor invoke

* Fix thread state and HIP stream enums
This commit is contained in:
Jonathan R. Madsen
2022-05-08 04:40:10 -05:00
zatwierdzone przez GitHub
rodzic bab90baf0b
commit b208047741
50 zmienionych plików z 1736 dodań i 483 usunięć
@@ -1,14 +1,23 @@
#include <atomic>
#include <cstdio>
#include <cstdlib>
#include <string>
#include <thread>
#include <vector>
#if defined(USE_LOCKS)
# include <mutex>
using auto_lock_t = std::unique_lock<std::mutex>;
long total = 0;
std::mutex mtx{};
#else
# include <atomic>
std::atomic<long> total{ 0 };
#endif
long
fib(long n) __attribute__((noinline));
void
run(size_t nitr, long) __attribute__((noinline));
@@ -21,10 +30,19 @@ fib(long n)
void
run(size_t nitr, long n)
{
#if defined(USE_LOCKS)
for(size_t i = 0; i < nitr; ++i)
{
auto _v = fib(n);
auto_lock_t _lk{ mtx };
total += _v;
}
#else
long local = 0;
for(size_t i = 0; i < nitr; ++i)
local += fib(n);
total += local;
#endif
}
int
@@ -42,7 +60,7 @@ main(int argc, char** argv)
if(argc > 2) nthread = atol(argv[2]);
if(argc > 3) nitr = atol(argv[3]);
printf("[%s] Threads: %zu\n[%s] Iterations: %zu\n[%s] fibonacci(%li)...\n",
printf("\n[%s] Threads: %zu\n[%s] Iterations: %zu\n[%s] fibonacci(%li)...\n",
_name.c_str(), nthread, _name.c_str(), nitr, _name.c_str(), nfib);
std::vector<std::thread> threads{};
@@ -53,13 +71,16 @@ main(int argc, char** argv)
threads.emplace_back(&run, _nitr, nfib);
}
#if !defined(USE_LOCKS)
auto _nitr = std::max<size_t>(nitr - 0.25 * nitr, 1);
run(_nitr, nfib - 0.1 * nfib);
#endif
for(auto& itr : threads)
itr.join();
printf("[%s] fibonacci(%li) x %lu = %li\n", _name.c_str(), nfib, nthread,
total.load());
static_cast<long>(total));
return 0;
}