Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions .github/workflows/signaloid-python.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,24 @@ jobs:
echo "github.event.release.tag_name:" ${{ github.event.release.tag_name }}
echo "contains(github.event.pull_request.labels.*.name, 'ci:upload-binaries'):" ${{ contains(github.event.pull_request.labels.*.name, 'ci:upload-binaries') }}

- name: Check submodules are up to date
# Independent of the matrix, so run it once. The fetch below reuses the
# credentials actions/checkout persisted into each submodule, so this
# works whether the assets remote is private or public.
if: ${{ matrix.args.host == 'ubuntu-22.04' && matrix.python-version == '3.10' }}
run: |
git submodule foreach --quiet --recursive '
branch=$(git config -f "$toplevel/.gitmodules" "submodule.$name.branch" || echo HEAD)
git fetch --quiet origin "$branch"
pinned=$(git rev-parse HEAD)
latest=$(git rev-parse FETCH_HEAD)
if [ "$pinned" != "$latest" ]; then
echo "::error::Submodule $sm_path is behind origin/$branch. Pinned $pinned, latest $latest. Update it with: git submodule update --remote $sm_path && git add $sm_path"
exit 1
fi
echo "Submodule $sm_path is up to date with origin/$branch ($pinned)"
'

- name: Set up Python
uses: actions/setup-python@v6
with:
Expand Down
39 changes: 33 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,15 +45,17 @@ benchmarks for UxHw Core microarchitectures Athens and Jupiter for precisions 8,
python -m signaloid.benchmarking.automation \
--path-to-application ./my-uxhw-app \
--path-to-uxhw-sdk ~/project-uxhw-sdk \
--path-to-pin ~/pin-external-4.2 \
-u Athens Jupiter \
-s 8 16 32 \
-c Disabled Autocorrelation \
-r Mean
```

The tool needs access to the Signaloid UxHw SDK to build the applications for
UxHw, and access to the Intel Pin tool for accurate benchmarking. Arguments
UxHw. The Intel Pin tool is optional and off by default. Export `PIN_ROOT` and
pass `--measure-dynamic-instructions` to also measure the dynamic instruction
count. Without that flag the run never uses Pin, even when `PIN_ROOT` is set,
and reports the count as missing. Arguments
`-u/--representation-types`, `-s/--representation-sizes`,
`-c/--uncertainty-correlation_types`, `-r/--reporting-methods` can also be
supplied using a YAML file with `--config <file>`.
Expand Down Expand Up @@ -81,17 +83,42 @@ dist_value = DistributionalValue.parse(ux_binary_buffer)
### Create Distribution Plots
Create plots to visualize distributional information by using the
[`plot` function](./src/signaloid/distributional_information_plotting/plot_wrapper.py)
with a `DistributionalValue` object containing Ux Data. The `plot` function is a
wrapper function for the `PlotHistogramDiracDeltas` class for plotting a
distributional value as a histogram with variable bin widths.
with a `PlotData` object built from a `DistributionalValue` containing Ux Data.
The `plot` function is a wrapper function for the `PlotHistogramDiracDeltas`
class for plotting a distributional value as a histogram with variable bin
widths.

```python
from signaloid.distributional_information_plotting.plot_histogram_dirac_deltas import PlotData
from signaloid.distributional_information_plotting.plot_wrapper import plot

# Intermediate code which writes to ux_string
# ...

# Create distributional value object from Ux String
dist_value = DistributionalValue.parse(ux_string)
plot(dist_value)
plot(PlotData(dist_value))
```

For plotting from raw samples, saving to a file, and the other `plot` options,
see the package [README.md](src/signaloid/distributional_information_plotting/README.md).

### Sample from Ux Data
Draw random samples from a distributional value with the
[`sample_generator` function](./src/signaloid/distributional_information_plotting/sample_generator.py).
Samples of the finite part of the distribution are drawn by inverse transform
sampling of the binned distribution. Distributions that also carry non-finite
mass (`NaN`, `-Inf`, `+Inf`) are sampled as a mixture, with each sample drawn
from the finite or the non-finite part in proportion to their masses.

```python
from signaloid.distributional_information_plotting.sample_generator import sample_generator

# Intermediate code which writes to ux_string
# ...

samples = sample_generator(ux_string, n_samples=1000)
```

To sample from a `DistributionalValue` that is already parsed, use
`sample_from_distributional_value` from the same module.
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ requires = ["poetry-core>=1.9.1", "poetry-dynamic-versioning>=1.0.0,<2.0.0"]
build-backend = "poetry_dynamic_versioning.backend"

[tool.poetry.dependencies]
python = ">=3.10"
python = ">=3.10,<4"
numpy = [
{ version = ">=2.0.0", python = ">=3.10,<3.13" },
{ version = ">=2.1.0", python = ">=3.10,<3.14" },
Expand Down
55 changes: 37 additions & 18 deletions src/signaloid/benchmarking/automation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,11 +18,10 @@ third-party dependencies.
### System dependencies
- gcc or g++
- GNU Scientific Library (on Ubuntu: libgsl-dev).
- Python (3.10+; see root-level pyproject.toml)
- Python (3.10+, see root-level pyproject.toml)
- GNU Make (make)
- bash
- lscpu
- hyperfine

C and C++ files are compiled separately and linked with `c++`. Applications that
link against GSL need `-lgsl -lgslcblas -lm`.
Expand All @@ -34,6 +33,8 @@ virtual environment, then install the package with pip:
python -m venv .venv
source .venv/bin/activate
pip install .
# Or, to also get the optional Google Sheets upload:
pip install ".[sheets]"
```
For development, you can install in editable mode with `pip install -e .`
instead.
Expand Down Expand Up @@ -87,11 +88,21 @@ See the [Signaloid documentation for more details about using GitHub repositorie
with UxHw](https://docs.signaloid.io/docs/api/guides/builds/builds-repository/).


### Intel PIN
### Intel PIN (Optional)

Intel PIN is necessary for the dynamic instruction count on any timing run. The
tool uses it via command-line argument `--path-to-pin` or environment variable
`PIN_ROOT`. A timing run fails if neither is set.
Intel PIN adds a dynamic instruction count to a timing run. It is optional and
off by default. The tool turns it on only when you pass
`--measure-dynamic-instructions`. The kit location comes from the `PIN_ROOT`
environment variable.

That flag is the single switch. A run without it never uses PIN, even when
`PIN_ROOT` is set in your shell. So you do not need to unset anything to turn
the measurement off.

Passing the flag is an explicit opt-in, so an unset `PIN_ROOT`, or a kit that
is missing or not built, is a hard error. Without the flag a timing run still
completes. It logs that it skipped the count and reports `pinDynInstCount` as
missing.

The dynamic instruction count (`pinDynInstCount`) is from the `inscount0` tool.

Expand All @@ -100,8 +111,11 @@ Basic installation instructions:
[Pin binary-instrumentation tool downloads](https://www.intel.com/content/www/us/en/developer/articles/tool/pin-a-binary-instrumentation-tool-downloads.html)
(validated against PIN 4.2).
2. Extract it to a directory `<dir>` of your choice.
3. Set parameter `--path-to-pin <dir>` (or `export PIN_ROOT=<dir>`).
3. Build the `inscount0` counter once (the kit does not ship it prebuilt):
3. Build the `inscount0` counter once. The kit does not ship it prebuilt.
```
make -C <dir>/source/tools/ManualExamples obj-intel64/inscount0.so
```
4. `export PIN_ROOT=<dir>`, then run with `--measure-dynamic-instructions`.


### Google Sheets (Optional)
Expand All @@ -116,7 +130,9 @@ To enable it you need:
pip install ".[sheets]"
```
A plain `pip install .` does **not** pull in the Sheets stack (`gspread`,
`google-api-python-client`, `oauth2client`).
`google-api-python-client`, `oauth2client`). Passing `--write-sheets`
without it raises `RuntimeError` at startup (before the pipeline runs),
naming the modules it could not find.
2. A Google Cloud service-account credentials JSON file, supplied via
`--google-credentials <path>` or the `GOOGLE_APPLICATION_CREDENTIALS`
environment variable (whichever resolves to an existing file). There is no
Expand Down Expand Up @@ -188,7 +204,7 @@ signaloid-benchmarking [OPTIONS]
| Flag | Default | Description |
|---|---|---|
| `--path-to-uxhw-sdk` | `~/project-uxhw-sdk` | Path to the UxHw SDK. |
| `--path-to-pin` | `None` (uses `PIN_ROOT` env) | Path to the Intel PIN kit, exported as `PIN_ROOT` for the timing script's dynamic instruction count. Omit to keep any existing `PIN_ROOT`. A timing run fails clearly if neither `--path-to-pin` nor `PIN_ROOT` is set. |
| `--measure-dynamic-instructions` | `False` | Also measure dynamic instruction counts with Intel PIN, using the kit that `PIN_ROOT` points at. This flag is the only way to switch the measurement on, so a run without it never invokes PIN even when `PIN_ROOT` is set. Passing it without a usable `PIN_ROOT` is an error. |
| `--demo-cli-args` | `""` | Extra command-line arguments passed to both native-MC and UxHw executions. Use this when the demo application requires additional flags (e.g., `--demo-cli-args "--asc-file inputs/blink.asc"`). |
| `--ground-truth-size` | `1` | Number of Monte Carlo samples for ground truth generation. |
| `--ground-truth-type` | `MonteCarlo` | Type of ground truth: `MonteCarlo` or `WeightedSamples`. |
Expand Down Expand Up @@ -306,12 +322,12 @@ output):
4. **UxHw Tracing Database** — Runs the application through the UxHw tracing
pipeline to produce UxHw distributional outputs for each representation
type/size/correlation combination. If a tracing database already exists, the
script will prompt before overwriting. The tracing build uses `-O0` (so the
script will prompt before overwriting. The tracing build uses `-O0`, so the
`addDistValueTrace` `file:line` directives resolve against unoptimised debug
info); to guard against optimisation changing the traced values, each config
is also built at `-O2` and its Ux strings are checked (byte-for-byte) against
the `-O0` ones. Any difference is reported as a warning and does not stop the
run. If a config's `-O2` build or run fails, that config is skipped and
info. Optimisation could change the traced values. To guard against that,
each config is also built at `-O2` and its Ux strings are checked
(byte-for-byte) against the `-O0` ones. Any difference is reported as a
warning and does not stop the run. If a config's `-O2` build or run fails, that config is skipped and
reported as failing in an `-O2 ux-string verification FAILED` summary (the
run still continues). Set `TRACING_VERIFY_OPTFLAGS` to compare against a
different level.
Expand All @@ -325,8 +341,8 @@ output):
validation.
9. **Equivalent Monte Carlo** — Computes the true EMCC by comparing adversary MC
distances to UxHw distances.
10. **UxHw Timings** — Measures execution time and dynamic instruction counts
for each UxHw configuration.
10. **UxHw Timings**. Measures execution time for each UxHw configuration. It
also measures dynamic instruction counts when Intel PIN is enabled.
11. **Native MC Timings** — Measures execution time for native MC at each EMCC
sample size.
12. **Load Measurements** — Loads all measurement data from the timing file.
Expand Down Expand Up @@ -405,7 +421,8 @@ underlying bash scripts used by the pipeline:
(not executed) by the Python tool with pre-set environment variables. It
handles UxHw compilation, native MC benchmarking, UxHw tracing, and timing
collection (the UxHw cores are compiled and timed via UxHw. The dynamic
instruction count comes from Intel PIN). Compilation warnings are redirected
instruction count comes from Intel PIN when PIN is enabled, and is left
unset otherwise). Compilation warnings are redirected
to log files (`uxhw-build.log` and `native-mc-build.log` in the `logs/`
directory). Source files are discovered recursively (excluding `build/`
directories), and C++ files (`.cc`, `.cpp`) are automatically included when
Expand Down Expand Up @@ -476,6 +493,8 @@ Notes on the schema:
- Missing numeric fields (e.g. `dbTime` on native runs) are serialised as
`null`.
- `dbDynInstCount` for UxHw rows is `0`.
- `pinDynInstCount` is `null` when the run had no Intel PIN kit configured.
Intel PIN is optional and off by default, so this is the common case.
- `uxhwTargetRepetitions` is the UxHw-loop target at session start. The actual
rep count used per measurement can differ — native-MC rows use
`NATIVE_MC_REPETITION` (dynamically computed per precision), and the UxHw
Expand Down
16 changes: 8 additions & 8 deletions src/signaloid/benchmarking/automation/arguments.py
Original file line number Diff line number Diff line change
Expand Up @@ -273,15 +273,15 @@ def create_argument_parser() -> ArgumentParser:
)

parser.add_argument(
"--path-to-pin",
dest="path_to_pin",
type=str,
default=None,
"--measure-dynamic-instructions",
dest="measure_dynamic_instructions",
action="store_true",
help=(
"Path to the Intel PIN kit (sets PIN_ROOT for the timing "
"script, which uses it to count dynamic instructions). When "
"omitted, an inherited PIN_ROOT is used. If neither is set "
"the timing run errors with 'Intel Pin not found'."
"Also measure dynamic instruction counts, using the Intel PIN "
"kit that PIN_ROOT points at. Off by default. This flag is the "
"only way to switch the measurement on, so a run without it "
"never invokes PIN even when PIN_ROOT is set. Passing it "
"without a usable PIN_ROOT is an error."
),
)

Expand Down
6 changes: 3 additions & 3 deletions src/signaloid/benchmarking/automation/benchmark.py
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,7 @@ def __init__(
self,
path_to_application: str,
path_to_uxhw_sdk: str = DEFAULT_UXHW_SDK_PATH,
path_to_pin: str | None = None,
measure_dynamic_instructions: bool = False,
has_analytic_ground_truth: bool = False,
path_to_ground_truth_file: str = "",
ground_truth_size: int = DEFAULT_GROUND_TRUTH_SIZE,
Expand Down Expand Up @@ -116,7 +116,7 @@ def __init__(
"""
self.path_to_application = os.path.expanduser(path_to_application)
self.path_to_uxhw_sdk = os.path.expanduser(path_to_uxhw_sdk)
self.path_to_pin = os.path.expanduser(path_to_pin) if path_to_pin else None
self.measure_dynamic_instructions = measure_dynamic_instructions
self.has_analytic_ground_truth = has_analytic_ground_truth
if self.has_analytic_ground_truth:
self.path_to_ground_truth_file = os.path.expanduser(
Expand Down Expand Up @@ -410,7 +410,7 @@ def get_application_info(self) -> None:
# Export common timing environment variables
self.tracing_db_path = export_timing_env(
path_to_uxhw_sdk=self.path_to_uxhw_sdk,
path_to_pin=self.path_to_pin,
measure_dynamic_instructions=self.measure_dynamic_instructions,
path_to_application=self.path_to_application,
application_name=self.application_name,
application_version=self.application_version,
Expand Down
30 changes: 26 additions & 4 deletions src/signaloid/benchmarking/automation/benchmark_application.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
# DEALINGS IN THE SOFTWARE.

import datetime
import importlib.util
import os
import sys
import traceback
Expand Down Expand Up @@ -62,6 +63,11 @@
TOTAL_STEPS = 15
LOG_FILE_PREFIX = "benchmarking_automation_error"

# Top-level modules the Google Sheets upload imports (see report_writer's
# write_results_to_spreadsheet and _get_credentials). They ship in the
# optional `sheets` extra, so a plain `pip install .` leaves them absent.
SHEETS_MODULE_NAMES = ("gspread", "googleapiclient", "oauth2client")


def _use_color() -> bool:
if os.environ.get("NO_COLOR"):
Expand Down Expand Up @@ -91,12 +97,28 @@ def _resolve_sheets_credentials(args: Namespace) -> str | None:
``None`` when the user opted out.

Raises:
RuntimeError: If ``--write-sheets`` is set but no credentials file
resolves on disk, or the Drive folder / Sheets template IDs are
unset.
RuntimeError: If ``--write-sheets`` is set but the ``sheets`` extra
is not installed, no credentials file resolves on disk, or the
Drive folder / Sheets template IDs are unset.
"""
if not args.write_sheets:
return None
# The upload's dependencies are imported lazily inside
# write_results_to_spreadsheet, so without this check a missing `sheets`
# extra only surfaces at step 15, once the whole pipeline has run.
# find_spec locates the modules without importing them, keeping the
# opted-out path free of the extra's import cost.
missing_module_names = [
module_name
for module_name in SHEETS_MODULE_NAMES
if importlib.util.find_spec(module_name) is None
]
if missing_module_names:
raise RuntimeError(
"--write-sheets requires the `sheets` extra, but "
f"{', '.join(missing_module_names)} could not be found. "
'Install it with: pip install ".[sheets]"'
)
credentials_path = resolve_google_credentials_path(
google_credentials=args.google_credentials,
)
Expand Down Expand Up @@ -248,7 +270,7 @@ def _run_pipeline(args: Namespace, *, credentials_path: str | None) -> None:
benchmark = Benchmark(
path_to_application=args.path_to_application,
path_to_uxhw_sdk=args.path_to_uxhw_sdk,
path_to_pin=args.path_to_pin,
measure_dynamic_instructions=(args.measure_dynamic_instructions),
has_analytic_ground_truth=(args.has_analytic_ground_truth),
path_to_ground_truth_file=(args.path_to_ground_truth_file),
ground_truth_type=args.ground_truth_type,
Expand Down
Loading
Loading