Update branding to ROCm Systems Profiler in documentation (#2)
* Update branding in docs
* Rename image used in documentation
* Update names of code samples.
In the code snippets, the "-" is not valid. ex., rocprof-sys_ --> rocprofsys_
* Update ASCII art
* update Doxyfile strip_from_path
* Add a "Formerly known as" message.
* Fixed typo in product name
ROCm Systems Profiler, not ROCm Profiler System
* Add "Omnitrace" back to the metadata keywords
* Update "install via package manager" section
* Update paths to user API files
* Rename configuration and environment settings
* Update Doxyfiles
Update publisher name & ID to "AMD".
Update bundle ID to "rocprofiler-systems"
* Update docs/what-is-rocprof-sys.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/conceptual/data-collection-modes.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/tutorials/video-tutorials.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/conceptual/rocprof-sys-feature-set.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/configuring-runtime-options.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/configuring-validating-environment.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/general-tips-using-rocprof-sys.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/reference/rocprof-sys-glossary.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/reference/development-guide.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/instrumenting-rewriting-binary-application.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/install/quick-start.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Note that videos were recorded using the "Omnitrace" name.
* Rebase and update some file paths
* Update paths to doc images
* Update Omnitrace references in code snippets
* Rename examples still using the "omni" prefix.
* Update docs/how-to/performing-causal-profiling.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/profiling-python-scripts.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/sampling-call-stack.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/how-to/understanding-rocprof-sys-output.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
* Update docs/install/install.rst
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
---------
Co-authored-by: Peter Park <peter.park@amd.com>
Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
[ROCm/rocprofiler-systems commit: 032d39f15c]
This commit is contained in:
zatwierdzone przez
GitHub
rodzic
181a782835
commit
d13617cf91
@@ -1,122 +1,122 @@
|
||||
.. meta::
|
||||
:description: Omnitrace documentation and reference
|
||||
:keywords: Omnitrace, ROCm, profiler, tracking, visualization, tool, Instinct, accelerator, AMD
|
||||
:description: ROCm Systems Profiler development documentation and reference
|
||||
:keywords: rocprof-sys, rocprofiler-systems, Omnitrace, ROCm, development, developers guide, profiler, tracking, visualization, tool, Instinct, accelerator, AMD
|
||||
|
||||
****************************************************
|
||||
Development guide
|
||||
****************************************************
|
||||
|
||||
This guide discusses the `Omnitrace <https://github.com/ROCm/omnitrace>`_ design.
|
||||
It includes a list of the executables and libraries, along with a discussion of the application's
|
||||
This guide discusses the `ROCm Systems Profiler <https://github.com/ROCm/rocprofiler-systems>`_ design.
|
||||
It includes a list of the executables and libraries, along with a discussion of the application's
|
||||
memory, sampling, and time-window constraint models.
|
||||
|
||||
Executables
|
||||
========================================
|
||||
|
||||
This section lists the Omnitrace executables.
|
||||
This section lists the ROCm Systems Profiler executables.
|
||||
|
||||
omnitrace-avail: `source/bin/omnitrace-avail <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/bin/omnitrace-avail>`_
|
||||
rocprof-sys-avail: `source/bin/rocprof-sys-avail <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/bin/rocprof-sys-avail>`_
|
||||
-------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
The ``main`` routine of ``omnitrace-avail`` has three important sections:
|
||||
The ``main`` routine of ``rocprof-sys-avail`` has three important sections:
|
||||
|
||||
* Printing components
|
||||
* Printing options
|
||||
* Printing hardware counters
|
||||
|
||||
omnitrace-sample: `source/bin/omnitrace-sample <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/bin/omnitrace-sample>`_
|
||||
rocprof-sys-sample: `source/bin/rocprof-sys-sample <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/bin/rocprof-sys-sample>`_
|
||||
----------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* Requires a command-line format of ``omnitrace-sample <options> -- <command> <command-args>``
|
||||
* Requires a command-line format of ``rocprof-sys-sample <options> -- <command> <command-args>``
|
||||
* Translates command-line options into environment variables
|
||||
* Adds ``libomnitrace-dl.so`` to ``LD_PRELOAD``
|
||||
* Adds ``librocprof-sys-dl.so`` to ``LD_PRELOAD``
|
||||
* Is launched by using ``execvpe`` with ``<command> <command-args>`` and a modified environment
|
||||
|
||||
omnitrace-casual: `source/bin/omnitrace-causal <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/bin/omnitrace-causal>`_
|
||||
rocprof-sys-casual: `source/bin/rocprof-sys-causal <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/bin/rocprof-sys-causal>`_
|
||||
----------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
When there is exactly one causal profiling configuration variant (which enables debugging),
|
||||
``omnitrace-casual`` has a nearly identical design to ``omnitrace-sample``
|
||||
``rocprof-sys-casual`` has a nearly identical design to ``rocprof-sys-sample``
|
||||
|
||||
When the command-line options produce more than one causal profiling configuration variant,
|
||||
the following actions take place for each variant:
|
||||
|
||||
* ``omnitrace-causal`` calls ``fork()``
|
||||
* ``rocprof-sys-causal`` calls ``fork()``
|
||||
* the child process launches ``<command> <command-args>`` using ``execvpe``, which modifies the environment for the variant
|
||||
* the parent process waits for the child process to finish
|
||||
|
||||
omnitrace-instrument: `source/bin/omnitrace-instrument <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/bin/omnitrace-instrument>`_
|
||||
rocprof-sys-instrument: `source/bin/rocprof-sys-instrument <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/bin/rocprof-sys-instrument>`_
|
||||
----------------------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* Requires a command-line format of ``omnitrace-instrument <options> -- <command> <command-args>``
|
||||
* Allows the user to provide options specifying whether to perform runtime instrumentation, use binary rewrite, or
|
||||
* Requires a command-line format of ``rocprof-sys-instrument <options> -- <command> <command-args>``
|
||||
* Allows the user to provide options specifying whether to perform runtime instrumentation, use binary rewrite, or
|
||||
attach to process
|
||||
* Either opens the instrumentation target (for binary rewrite), launches the target and stops it
|
||||
before it starts executing ``main``, or attaches to a running executable and pauses it
|
||||
* Finds all functions in the targets
|
||||
* Finds ``libomnitrace-dl`` and locates the functions
|
||||
* Iterates over and instruments all the functions, provided they satisfy the
|
||||
* Finds ``librocprof-sys-dl`` and locates the functions
|
||||
* Iterates over and instruments all the functions, provided they satisfy the
|
||||
defined criteria (such as a minimum number of instructions)
|
||||
|
||||
* See the ``module_function`` class
|
||||
|
||||
* Until this point, the workflow has been the same for the different options,
|
||||
* Until this point, the workflow has been the same for the different options,
|
||||
but it diverges after instrumentation is complete:
|
||||
|
||||
* For a binary rewrite: it produces a new instrumented binary and exits
|
||||
* For runtime instrumentation or attaching to a process: it instructs the application
|
||||
* For runtime instrumentation or attaching to a process: it instructs the application
|
||||
to resume and then waits for it to exit
|
||||
|
||||
Libraries
|
||||
========================================
|
||||
|
||||
Common library: `source/lib/common <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/common>`_
|
||||
Common library: `source/lib/common <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/common>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* General header-only functionality used in multiple executables and/or libraries.
|
||||
* General header-only functionality used in multiple executables and/or libraries.
|
||||
* Not installed or exported outside of the build tree.
|
||||
|
||||
Core library: `source/lib/core <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/core>`_
|
||||
Core library: `source/lib/core <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/core>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* Static PIC library with functionality that does not depend on any components.
|
||||
* Static PIC library with functionality that does not depend on any components.
|
||||
* Not installed or exported outside of the build tree.
|
||||
|
||||
Binary library: `source/lib/binary <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/binary>`_
|
||||
Binary library: `source/lib/binary <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/binary>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* Static PIC library with functionality for reading/analyzing binary info.
|
||||
* Mostly used by the causal profiling sections of ``libomnitrace``.
|
||||
* Mostly used by the causal profiling sections of ``librocprof-sys``.
|
||||
* Not installed or exported outside of the build tree.
|
||||
|
||||
libomnitrace: `source/lib/omnitrace <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/omnitrace>`_
|
||||
librocprof-sys: `source/lib/rocprof-sys <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/rocprof-sys>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
This is the main library encapsulating all the capabilities.
|
||||
|
||||
libomnitrace-dl: `source/lib/omnitrace-dl <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/omnitrace-dl>`_
|
||||
librocprof-sys-dl: `source/lib/rocprof-sys-dl <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/rocprof-sys-dl>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
This is a lightweight, front-end library for ``libomnitrace`` which serves three primary purposes:
|
||||
This is a lightweight, front-end library for ``librocprof-sys`` which serves three primary purposes:
|
||||
|
||||
* Dramatically speeds up instrumentation time compared to using ``libomnitrace`` directly because
|
||||
Dyninst must parse the entire library in order to find the instrumentation functions
|
||||
(a ``dlopen`` call is made on ``libomnitrace`` when the instrumentation functions get called)
|
||||
* Prevents re-entry if ``libomnitrace`` calls an instrumented function internally
|
||||
* Coordinates communication between ``libomnitrace-user`` and ``libomnitrace``
|
||||
* Dramatically speeds up instrumentation time compared to using ``librocprof-sys`` directly because
|
||||
Dyninst must parse the entire library in order to find the instrumentation functions
|
||||
(a ``dlopen`` call is made on ``librocprof-sys`` when the instrumentation functions get called)
|
||||
* Prevents re-entry if ``librocprof-sys`` calls an instrumented function internally
|
||||
* Coordinates communication between ``librocprof-sys-user`` and ``librocprof-sys``
|
||||
|
||||
libomnitrace-user: `source/lib/omnitrace-user <https://github.com/ROCm/omnitrace/tree/amd-mainline/source/lib/omnitrace-user>`_
|
||||
librocprof-sys-user: `source/lib/rocprof-sys-user <https://github.com/ROCm/rocprofiler-systems/tree/amd-mainline/source/lib/rocprof-sys-user>`_
|
||||
--------------------------------------------------------------------------------------------------------------------------------
|
||||
|
||||
* Provides a set of functions and types for the users to add to their code, for example,
|
||||
disabling data collection globally or on a specific thread or
|
||||
user-defined region
|
||||
* If ``libomnitrace-dl`` is not loaded, the user API is effectively a set of no-op function calls.
|
||||
* If ``librocprof-sys-dl`` is not loaded, the user API is effectively a set of no-op function calls.
|
||||
|
||||
Testing tools
|
||||
========================================
|
||||
|
||||
* `CDash Testing Dashboard <https://my.cdash.org/index.php?project=Omnitrace>`_ (requires a login)
|
||||
* `CDash Testing Dashboard <https://my.cdash.org/index.php?project=rocprofiler-systems>`_ (requires a login)
|
||||
|
||||
Components
|
||||
========================================
|
||||
@@ -124,34 +124,34 @@ Components
|
||||
Most measurements and capabilities are encapsulated into a "component" with the following definitions:
|
||||
|
||||
Measurement
|
||||
A recording of some data relevant to performance, for instance, the current call-stack,
|
||||
A recording of some data relevant to performance, for instance, the current call-stack,
|
||||
hardware counter values, current memory usage, or timestamp
|
||||
|
||||
Capability
|
||||
Handles the implementation or orchestration of some feature which is used
|
||||
to collect measurements, for example, a component which handles setting up function wrappers
|
||||
Handles the implementation or orchestration of some feature which is used
|
||||
to collect measurements, for example, a component which handles setting up function wrappers
|
||||
around various functions such as ``pthread_create`` or ``MPI_Init``.
|
||||
|
||||
Components are designed to either hold no data at all or only the data for both an instantaneous
|
||||
Components are designed to either hold no data at all or only the data for both an instantaneous
|
||||
measurement and a phase measurement.
|
||||
|
||||
Components which store data typically implement a static ``record()`` function
|
||||
Components which store data typically implement a static ``record()`` function
|
||||
for getting a record of the measurement,
|
||||
``start()`` and ``stop()`` member functions for calculating a phase measurement,
|
||||
``start()`` and ``stop()`` member functions for calculating a phase measurement,
|
||||
and a ``sample()`` member function for storing an
|
||||
instantaneous measurement. In reality, there are several more "standard" functions
|
||||
instantaneous measurement. In reality, there are several more "standard" functions
|
||||
but these are the most commonly-used ones.
|
||||
|
||||
Components which do not store data might also have ``start()``, ``stop()``, and ``sample()``
|
||||
Components which do not store data might also have ``start()``, ``stop()``, and ``sample()``
|
||||
functions. However, components which
|
||||
implement function wrappers typically provide a call operator or ``audit(...)``
|
||||
implement function wrappers typically provide a call operator or ``audit(...)``
|
||||
functions. These are invoked with the
|
||||
wrapped function's arguments before the wrapped function gets called and with the return value
|
||||
wrapped function's arguments before the wrapped function gets called and with the return value
|
||||
after the wrapped function gets called.
|
||||
|
||||
.. note::
|
||||
|
||||
The goal of this design is to provide relatively small and resuable lightweight objects
|
||||
The goal of this design is to provide relatively small and resuable lightweight objects
|
||||
for recording measurements and implementing capabilities.
|
||||
|
||||
Wall-clock component example
|
||||
@@ -195,7 +195,7 @@ A component for computing the elapsed wall-clock time looks like this:
|
||||
Function wrapper component example
|
||||
--------------------------------------
|
||||
|
||||
A component which implements wrappers around ``fork()`` and ``exit(int)`` (and stores no data)
|
||||
A component which implements wrappers around ``fork()`` and ``exit(int)`` (and stores no data)
|
||||
could look like this:
|
||||
|
||||
.. code-block:: cpp
|
||||
@@ -219,7 +219,7 @@ could look like this:
|
||||
void operator()(const gotcha_data&, void (*real_exit)(int), int _exit_code)
|
||||
{
|
||||
// catch the call to exit and finalize before truly exiting
|
||||
omnitrace_finalize();
|
||||
rocprofsys_finalize();
|
||||
|
||||
real_exit(_exit_code);
|
||||
}
|
||||
@@ -298,22 +298,22 @@ Collected data is generally handled in one of the three following ways:
|
||||
* It is managed implicitly by Timemory and accessed as needed
|
||||
* As thread-local data
|
||||
|
||||
In general, only instrumentation for relatively simple data is directly passed to
|
||||
In general, only instrumentation for relatively simple data is directly passed to
|
||||
Perfetto and/or Timemory during runtime.
|
||||
For example, the callbacks from binary instrumentation, user API instrumentation,
|
||||
For example, the callbacks from binary instrumentation, user API instrumentation,
|
||||
and roctracer directly invoke
|
||||
calls to Perfetto or Timemory's storage model. Otherwise, the data is stored
|
||||
by Omnitrace in the thread-data model
|
||||
calls to Perfetto or Timemory's storage model. Otherwise, the data is stored
|
||||
by ROCm Systems Profiler in the thread-data model
|
||||
which is more persistent than simply using ``thread_local`` static data, which gets deleted
|
||||
when the thread stops.
|
||||
|
||||
Thread identification
|
||||
--------------------------------------
|
||||
|
||||
Each CPU thread is assigned two integral identifiers. One identifier, the ``internal_value``, is
|
||||
Each CPU thread is assigned two integral identifiers. One identifier, the ``internal_value``, is
|
||||
atomically incremented every time a new thread is created.
|
||||
The other identifier, known as the ``sequent_value``, tries to account for the fact that Omnitrace, Perfetto, ROCm, and other applications
|
||||
start background threads. When a thread is created as a by-product of Omnitrace,
|
||||
The other identifier, known as the ``sequent_value``, tries to account for the fact that ROCm Systems Profiler, Perfetto, ROCm, and other applications
|
||||
start background threads. When a thread is created as a by-product of ROCm Systems Profiler,
|
||||
the index is offset by a large value. This serves
|
||||
two purposes:
|
||||
|
||||
@@ -325,88 +325,88 @@ The ``sequent_value`` identifier is typically used to access the thread-data.
|
||||
Thread-data class
|
||||
--------------------------------------
|
||||
|
||||
Currently, most thread data is effectively stored in a static
|
||||
``std::array<std::unique_ptr<T>, OMNITRACE_MAX_THREADS>`` instance.
|
||||
``OMNITRACE_MAX_THREADS`` is a value defined a compile-time and set to ``2048``
|
||||
Currently, most thread data is effectively stored in a static
|
||||
``std::array<std::unique_ptr<T>, ROCPROFSYS_MAX_THREADS>`` instance.
|
||||
``ROCPROFSYS_MAX_THREADS`` is a value defined a compile-time and set to ``2048``
|
||||
for release builds. During finalization,
|
||||
Omnitrace iterates through the thread-data and transforms that data
|
||||
ROCm Systems Profiler iterates through the thread-data and transforms that data
|
||||
into something that can be passed along to Perfetto and/or Timemory.
|
||||
The downside of the current model is that if the user exceeds ``OMNITRACE_MAX_THREADS``,
|
||||
The downside of the current model is that if the user exceeds ``ROCPROFSYS_MAX_THREADS``,
|
||||
a segmentation fault occurs. To fix this issue,
|
||||
a new model is being adopted which has all the benefits of this model
|
||||
a new model is being adopted which has all the benefits of this model
|
||||
but permits dynamic expansion.
|
||||
|
||||
Sampling model
|
||||
========================================
|
||||
|
||||
The general structure for the sampling is within Timemory (``source/timemory/sampling``).
|
||||
The general structure for the sampling is within Timemory (``source/timemory/sampling``).
|
||||
Currently, all sampling is done per-thread
|
||||
via POSIX timers. Omnitrace supports both a real-time timer and a CPU-time timer.
|
||||
via POSIX timers. ROCm Systems Profiler supports both a real-time timer and a CPU-time timer.
|
||||
Both have adjustable frequencies, delays, and durations.
|
||||
By default, only CPU-time sampling is enabled. Initial settings are inherited from
|
||||
the settings starting with ``OMNITRACE_SAMPLING_``.
|
||||
By default, only CPU-time sampling is enabled. Initial settings are inherited from
|
||||
the settings starting with ``ROCPROFSYS_SAMPLING_``.
|
||||
|
||||
For each type of timer, timer-specific settings can be used to
|
||||
override the common and inherited timer settings.
|
||||
These settings begin with ``OMNITRACE_SAMPLING_CPUTIME`` for the CPU-time sampler
|
||||
and ``OMNITRACE_SAMPLING_REALTIME`` for
|
||||
the real-time sampler. For example, ``OMNITRACE_SAMPLING_FREQ=500`` initially sets the
|
||||
sampling frequency to 500 interrupts per second. Adding the setting ``OMNITRACE_SAMPLING_REALTIME_FREQ=10``
|
||||
For each type of timer, timer-specific settings can be used to
|
||||
override the common and inherited timer settings.
|
||||
These settings begin with ``ROCPROFSYS_SAMPLING_CPUTIME`` for the CPU-time sampler
|
||||
and ``ROCPROFSYS_SAMPLING_REALTIME`` for
|
||||
the real-time sampler. For example, ``ROCPROFSYS_SAMPLING_FREQ=500`` initially sets the
|
||||
sampling frequency to 500 interrupts per second. Adding the setting ``ROCPROFSYS_SAMPLING_REALTIME_FREQ=10``
|
||||
lowers the sampling frequency for the real-time sampler
|
||||
to 10 interrupts per second of real-time.
|
||||
|
||||
The Omnitrace-specific implementation can be found in
|
||||
`source/lib/omnitrace/library/sampling.cpp <https://github.com/ROCm/omnitrace/blob/main/source/lib/omnitrace/library/sampling.cpp>`_.
|
||||
Within `sampling.cpp <https://github.com/ROCm/omnitrace/blob/main/source/lib/omnitrace/library/sampling.cpp>`_,
|
||||
The ROCm Systems Profiler-specific implementation can be found in
|
||||
`source/lib/rocprof-sys/library/sampling.cpp <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/rocprof-sys/library/sampling.cpp>`_.
|
||||
Within `sampling.cpp <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/rocprof-sys/library/sampling.cpp>`_,
|
||||
there is a bundle of three sampling components:
|
||||
|
||||
* `backtrace_timestamp <https://github.com/ROCm/omnitrace/blob/main/source/lib/omnitrace/library/components/backtrace_timestamp.hpp>`_ simply
|
||||
* `backtrace_timestamp <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/rocprof-sys/library/components/backtrace_timestamp.hpp>`_ simply
|
||||
records the wall-clock time of the sample.
|
||||
* `backtrace <https://github.com/ROCm/omnitrace/blob/main/source/lib/omnitrace/library/components/backtrace.hpp>`_
|
||||
* `backtrace <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/rocprof-sys/library/components/backtrace.hpp>`_
|
||||
records the call-stack via libunwind.
|
||||
* `backtrace_metrics <https://github.com/ROCm/omnitrace/blob/main/source/lib/omnitrace/library/components/backtrace_metrics.hpp>`_
|
||||
* `backtrace_metrics <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/rocprof-sys/library/components/backtrace_metrics.hpp>`_
|
||||
records the sample metrics, such as peak RSS and the hardware counters.
|
||||
|
||||
These three components are bundled together in
|
||||
These three components are bundled together in
|
||||
a tuple-like ``struct`` (``tuple<backtrace_timestamp, backtrace, backtrace_metrics>``).
|
||||
A buffer of at least 1024 instances of this tuple is mapped using ``mmap``
|
||||
per-thread. When this buffer is full,
|
||||
A buffer of at least 1024 instances of this tuple is mapped using ``mmap``
|
||||
per-thread. When this buffer is full,
|
||||
the sampler hands the buffer off to its allocator thread and maps a new buffer with ``mmap``
|
||||
before taking the next sample. The allocator thread takes this data
|
||||
and either dynamically stores it in memory or writes it to a file depending on the
|
||||
value of ``OMNITRACE_USE_TEMPORARY_FILES``.
|
||||
This schema avoids all allocations in the signal handler, lets the data grow
|
||||
dynamically, avoids potentially slow I/O within the signal handler, and also enables
|
||||
before taking the next sample. The allocator thread takes this data
|
||||
and either dynamically stores it in memory or writes it to a file depending on the
|
||||
value of ``ROCPROFSYS_USE_TEMPORARY_FILES``.
|
||||
This schema avoids all allocations in the signal handler, lets the data grow
|
||||
dynamically, avoids potentially slow I/O within the signal handler, and also enables
|
||||
the capability of avoiding I/O altogether.
|
||||
The maximum number of samplers handled by each allocator is governed by the
|
||||
``OMNITRACE_SAMPLING_ALLOCATOR_SIZE`` setting (the default is eight). Whenever an allocator
|
||||
The maximum number of samplers handled by each allocator is governed by the
|
||||
``ROCPROFSYS_SAMPLING_ALLOCATOR_SIZE`` setting (the default is eight). Whenever an allocator
|
||||
has reached its limit,
|
||||
a new internal thread is created to handle the new samplers.
|
||||
|
||||
Time-window constraint model
|
||||
========================================
|
||||
|
||||
With the recent introduction of tracing delay and duration, the
|
||||
`constraint namespace <https://github.com/ROCm/omnitrace/blob/main/source/lib/core/constraint.hpp>`_
|
||||
was introduced to improve the management of delays and duration limits for
|
||||
With the recent introduction of tracing delay and duration, the
|
||||
`constraint namespace <https://github.com/ROCm/rocprofiler-systems/blob/main/source/lib/core/constraint.hpp>`_
|
||||
was introduced to improve the management of delays and duration limits for
|
||||
data collection. The ``spec`` class accepts a clock identifier, a delay value, a duration value, and an
|
||||
integer indicating how many times to repeat the delay and duration cycle. It is therefore
|
||||
integer indicating how many times to repeat the delay and duration cycle. It is therefore
|
||||
possible to perform tasks such as periodically enabling tracing for brief periods
|
||||
of time in between long periods without data collection while the application runs. The
|
||||
syntax follows the format ``clock_identifier:delay:capture_duration:cycles``, so a value of
|
||||
syntax follows the format ``clock_identifier:delay:capture_duration:cycles``, so a value of
|
||||
``10:1:3`` for the last three parameters represents the following sequence of operations:
|
||||
|
||||
* Ten seconds where no data is collected, then one second where it is
|
||||
* Ten seconds where no data is collected, then one second where it is
|
||||
* Ten seconds where no data is collected, then one second where it is
|
||||
* Ten seconds where no data is collected, then one second where it is
|
||||
* Ten seconds where no data is collected, then one second where it is
|
||||
* Stop
|
||||
|
||||
As another example, ``OMNITRACE_TRACE_PERIODS = realtime:10:1:5 process_cputime:10:2:20`` translates
|
||||
As another example, ``ROCPROFSYS_TRACE_PERIODS = realtime:10:1:5 process_cputime:10:2:20`` translates
|
||||
to this sequence:
|
||||
|
||||
* Five cycles of: no data collection for ten seconds of real-time followed by one second of data collection
|
||||
* Twenty cycles of: no data collection for ten seconds of process CPU time followed by two CPU-time seconds of data collection
|
||||
|
||||
Eventually, the goal is to migrate all subsets of data collection which currently support
|
||||
Eventually, the goal is to migrate all subsets of data collection which currently support
|
||||
more rudimentary models of time window constraints, such as process sampling and causal profiling,
|
||||
to this model.
|
||||
|
||||
+39
-39
@@ -1,40 +1,40 @@
|
||||
.. meta::
|
||||
:description: Omnitrace documentation and reference
|
||||
:keywords: Omnitrace, ROCm, profiler, tracking, visualization, tool, Instinct, accelerator, AMD
|
||||
:description: ROCm Systems Profiler glossary and reference
|
||||
:keywords: rocprof-sys, rocprofiler-systems, Omnitrace, ROCm, glossary, terminology, profiler, tracking, visualization, tool, Instinct, accelerator, AMD
|
||||
|
||||
*******************
|
||||
Omnitrace Glossary
|
||||
ROCm Systems Profiler Glossary
|
||||
*******************
|
||||
|
||||
This topic explains the terminology necessary to use Omnitrace.
|
||||
The list below provides a basic glossary for those who
|
||||
are new to binary instrumentation. It also clarifies ambiguities
|
||||
when certain terms have different
|
||||
contextual meanings, for example, the Omnitrace meaning of the term "module"
|
||||
This topic explains the terminology necessary to use ROCm Systems Profiler.
|
||||
The list below provides a basic glossary for those who
|
||||
are new to binary instrumentation. It also clarifies ambiguities
|
||||
when certain terms have different
|
||||
contextual meanings, for example, the ROCm Systems Profiler meaning of the term "module"
|
||||
when instrumenting Python.
|
||||
|
||||
**Binary**
|
||||
A file written in the Executable and Linkable Format (ELF). This is the standard file
|
||||
A file written in the Executable and Linkable Format (ELF). This is the standard file
|
||||
format for executable files, shared libraries, etc.
|
||||
|
||||
**Binary instrumentation**
|
||||
Inserting callbacks to instrumentation into an existing binary. This can be performed
|
||||
Inserting callbacks to instrumentation into an existing binary. This can be performed
|
||||
statically or dynamically.
|
||||
|
||||
**Static binary instrumentation**
|
||||
Loads an existing binary, determines instrumentation points, and generates a new binary
|
||||
with instrumentation directly embedded. It is applicable to executables and libraries but
|
||||
Loads an existing binary, determines instrumentation points, and generates a new binary
|
||||
with instrumentation directly embedded. It is applicable to executables and libraries but
|
||||
limited to only the functions defined in the binary. This is also known as **Binary rewrite**.
|
||||
|
||||
**Dynamic binary instrumentation**
|
||||
Loads an existing binary into memory, inserts instrumentation, and runs the binary.
|
||||
It is limited to executables but is capable of instrumenting linked libraries.
|
||||
Loads an existing binary into memory, inserts instrumentation, and runs the binary.
|
||||
It is limited to executables but is capable of instrumenting linked libraries.
|
||||
This is also known as **Runtime instrumentation**.
|
||||
|
||||
**Statistical sampling**
|
||||
At periodic intervals, the application is paused and the current call-stack of the CPU
|
||||
is recorded along with various other metrics. It uses timers that measure either
|
||||
(A) real clock time or (B) the CPU time used by the current thread and the CPU time
|
||||
**Statistical sampling**
|
||||
At periodic intervals, the application is paused and the current call-stack of the CPU
|
||||
is recorded along with various other metrics. It uses timers that measure either
|
||||
(A) real clock time or (B) the CPU time used by the current thread and the CPU time
|
||||
expended on behalf of the thread by the system. This is also known as simply **sampling**.
|
||||
|
||||
**Sampling rate**
|
||||
@@ -45,12 +45,12 @@ when instrumenting Python.
|
||||
* How long to wait before (A) and (B) begin triggering at their designated rate
|
||||
|
||||
**Sampling duration**
|
||||
* The amount of time (in real-time) after the start of the application to record samples.
|
||||
* The amount of time (in real-time) after the start of the application to record samples.
|
||||
* After this time limit has been reached, no more samples are recorded.
|
||||
|
||||
**Process sampling**
|
||||
At periodic (real-time) intervals, a background thread records global metrics without
|
||||
interrupting the current process. These metrics include, but are not limited to:
|
||||
At periodic (real-time) intervals, a background thread records global metrics without
|
||||
interrupting the current process. These metrics include, but are not limited to:
|
||||
CPU frequency, CPU memory high-water mark (i.e. peak memory usage), GPU temperature,
|
||||
and GPU power usage.
|
||||
|
||||
@@ -62,41 +62,41 @@ when instrumenting Python.
|
||||
* How long to wait (in real-time) before recording samples
|
||||
|
||||
**Sampling duration**
|
||||
* The amount of time (in real-time) after the start of the application to record samples.
|
||||
* The amount of time (in real-time) after the start of the application to record samples.
|
||||
* After this time limit has been reached, no more samples are recorded.
|
||||
|
||||
**Module**
|
||||
With respect to binary instrumentation, a module is defined as either the filename
|
||||
(such as ``foo.c``) or library name (``libfoo.so``) which contains the definition
|
||||
With respect to binary instrumentation, a module is defined as either the filename
|
||||
(such as ``foo.c``) or library name (``libfoo.so``) which contains the definition
|
||||
of one or more functions.
|
||||
|
||||
With respect to Python instrumentation, a module is defined as the **file** which contains
|
||||
the definition of one or more functions. The full path to this file typically contains the
|
||||
With respect to Python instrumentation, a module is defined as the **file** which contains
|
||||
the definition of one or more functions. The full path to this file typically contains the
|
||||
name of the "Python module".
|
||||
|
||||
**Basic block**
|
||||
A straight-line code sequence with no branches in (except for the entry) and
|
||||
A straight-line code sequence with no branches in (except for the entry) and
|
||||
no branches out (except for the exit).
|
||||
|
||||
**Address range**
|
||||
The instructions for a function in a binary start at certain address with the ELF file
|
||||
The instructions for a function in a binary start at certain address with the ELF file
|
||||
and end at a certain address. The range is ``end - start``.
|
||||
|
||||
The address range is a decent approximation for the "cost" of a function.
|
||||
The address range is a decent approximation for the "cost" of a function.
|
||||
For example, a larger address range approximately equates to more instructions.
|
||||
|
||||
**Instrumentation traps**
|
||||
On the x86 architecture, because instructions are of variable size, an instruction
|
||||
might be too small for Dyninst to replace it with the normal code sequence
|
||||
used to call instrumentation. When instrumentation is placed at points other
|
||||
than subroutine entry, exit, or call points, traps may be used to ensure
|
||||
the instrumentation fits. (By default, ``omnitrace-instrument`` avoids instrumentation
|
||||
On the x86 architecture, because instructions are of variable size, an instruction
|
||||
might be too small for Dyninst to replace it with the normal code sequence
|
||||
used to call instrumentation. When instrumentation is placed at points other
|
||||
than subroutine entry, exit, or call points, traps may be used to ensure
|
||||
the instrumentation fits. (By default, ``rocprof-sys-instrument`` avoids instrumentation
|
||||
which requires a trap.)
|
||||
|
||||
**Overlapping functions**
|
||||
Due to language constructs or compiler optimizations, it might be possible for
|
||||
multiple functions to overlap (that is, share part of the same function body)
|
||||
or for a single function to have multiple entry points. In practice, it's
|
||||
impossible to determine the difference between multiple overlapping functions
|
||||
and a single function with multiple entry points. (By default, ``omnitrace-instrument``
|
||||
Due to language constructs or compiler optimizations, it might be possible for
|
||||
multiple functions to overlap (that is, share part of the same function body)
|
||||
or for a single function to have multiple entry points. In practice, it's
|
||||
impossible to determine the difference between multiple overlapping functions
|
||||
and a single function with multiple entry points. (By default, ``rocprof-sys-instrument``
|
||||
avoids instrumenting overlapping functions.)
|
||||
Reference in New Issue
Block a user