Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 25 additions & 6 deletions .github/workflows/python-app.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,16 +24,16 @@ jobs:
with:
python-version: "3.10"
cache: "pip"
cache-dependency-path: pyproject.toml

- name: Install Python dependencies
run: |
python -m pip install --upgrade pip
pip install pytest
if [ -f requirements-ci.txt ]; then pip install -r requirements-ci.txt; fi
pip install -e ".[dev]"

- name: Test with pytest
run: |
PYTHONPATH=`pwd` pytest
pytest

test-csl:
runs-on: ubuntu-latest
Expand All @@ -46,12 +46,12 @@ jobs:
with:
python-version: "3.10"
cache: "pip"
cache-dependency-path: pyproject.toml

- name: Install Python dependencies
run: |
python -m pip install --upgrade pip
pip install pytest
if [ -f requirements-ci.txt ]; then pip install -r requirements-ci.txt; fi
pip install -e ".[dev]"

- name: Install CSL dependencies
run: |
Expand Down Expand Up @@ -91,6 +91,25 @@ jobs:

- name: Test CSL with simulator
run: |
pip install --no-deps -e .
export PATH=$PATH:`pwd`/cerebras-sdk
./tests/csl_runtime/run_tests.sh

build-package:
runs-on: ubuntu-latest

steps:
- uses: actions/checkout@v4

- name: Set up Python 3.10
uses: actions/setup-python@v6
with:
python-version: "3.10"
cache: "pip"
cache-dependency-path: pyproject.toml

- name: Build sdist and wheel
run: |
python -m pip install --upgrade pip
pip install build twine
python -m build
twine check dist/*
37 changes: 21 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
# SPADA — A Spatial Dataflow Architecture Programming Language
# SpaDA — A Spatial Dataflow Architecture Programming Language

SPADA is a programming language and compiler for spatial dataflow architectures such as the [Cerebras Wafer-Scale Engine](https://www.cerebras.net/). It provides precise control over data placement, communication streams, and asynchronous execution while abstracting architecture-specific routing details. SPADA also serves as a compiler intermediate representation (IR) for domain-specific languages; this repository includes a complete end-to-end compilation pipeline from the [GT4Py](https://github.com/GridTools/gt4py) stencil DSL (used in production weather forecasting at CSCS/MeteoSwiss) to Cerebras CSL.
SpaDA is a programming language and compiler for spatial dataflow architectures such as the [Cerebras Wafer-Scale Engine](https://www.cerebras.net/). It provides precise control over data placement, communication streams, and asynchronous execution while abstracting architecture-specific routing details. SpaDA also serves as a compiler intermediate representation (IR) for domain-specific languages; this repository includes a complete end-to-end compilation pipeline from the [GT4Py](https://github.com/GridTools/gt4py) stencil DSL (used in production weather forecasting at CSCS/MeteoSwiss) to Cerebras CSL.

Spatial dataflow architectures achieve exceptional throughput through disaggregated memory: each processing element (PE) holds only fast local SRAM, eliminating cache hierarchies and shared-memory contention. However, programming these architectures demands explicit orchestration of data movement over a circuit-switched network-on-chip (NoC), with limited concurrent communication channels and asynchronous, data-triggered task execution. SPADA addresses this by offering high-level constructs—`place`, `dataflow`, and `compute` blocks; `async`/`await`; `foreach` and `map` loops—alongside a formal dataflow semantics that defines routing correctness, data races, and deadlocks at compile time.
Spatial dataflow architectures achieve exceptional throughput through disaggregated memory: each processing element (PE) holds only fast local SRAM, eliminating cache hierarchies and shared-memory contention. However, programming these architectures demands explicit orchestration of data movement over a circuit-switched network-on-chip (NoC), with limited concurrent communication channels and asynchronous, data-triggered task execution. SpaDA addresses this by offering high-level constructs—`place`, `dataflow`, and `compute` blocks; `async`/`await`; `foreach` and `map` loops—alongside a formal dataflow semantics that defines routing correctness, data races, and deadlocks at compile time.

Key capabilities:
- **Explicit placement and dataflow**: Declare where data lives and how it moves between PEs.
- **Automatic routing assignment**: A checkerboard decomposition algorithm guarantees conflict-free channel allocation by construction, eliminating manual reasoning about hardware routing.
- **Multi-level compilation**: GT4Py stencils → Stencil IR → SPADA IR → Cerebras CSL, with automatic vectorization via Data Structure Descriptors (DSDs) and task fusion.
- **Compact code**: Hand-written SPADA kernels require 6–8× fewer lines than equivalent CSL; GT4Py stencils compile with up to 700× code reduction.
- **Multi-level compilation**: GT4Py stencils → Stencil IR → SpaDA IR → Cerebras CSL, with automatic vectorization via Data Structure Descriptors (DSDs) and task fusion.
- **Compact code**: Hand-written SpaDA kernels require 6–8× fewer lines than equivalent CSL; GT4Py stencils compile with up to 700× code reduction.
- **Near-ideal weak scaling**: Compiler-generated stencil kernels achieve >150 TFlop/s on the WSE-2 with near-ideal weak scaling across three orders of magnitude.

For full details, see the paper:
Expand All @@ -21,28 +21,33 @@ For full details, see the paper:

### Prerequisites

- Python ≥ 3.8
- Python ≥ 3.10
- [Cerebras SDK](https://sdk.cerebras.net/) (required to compile and run generated CSL code on WSE hardware; optional for compiler development)

### Installation

Clone the repository and install the package:

```bash
git clone https://github.com/glukas/spada.git
git clone https://github.com/spcl/spada.git
cd spada
pip install -e .
```

To install with development dependencies:
To install with development dependencies (test suite, formatters, and the
standalone placement research code):

```bash
pip install -e ".[dev]"
```

### Compiling a SPADA Program
Other optional dependency groups: `docs` (MkDocs site in `irspec/`), `placement`
(igraph/matplotlib/hilbertcurve), `render` (pycairo, needs system cairo), and
`all`.

The `sptlc` command-line tool compiles a SPADA Spatial IR (`.sptl`) file to Cerebras CSL:
### Compiling a SpaDA Program

The `sptlc` command-line tool compiles a SpaDA Spatial IR (`.sptl`) file to Cerebras CSL:

```bash
sptlc samples/benchmarks/laplacian_128_128_80.sptl output/ --param I=128 --param J=128
Expand All @@ -61,7 +66,7 @@ Key options:

### Compiling from GT4Py

To compile a GT4Py stencil file to SPADA IR (`.spst` and `.sptl`):
To compile a GT4Py stencil file to SpaDA IR (`.spst` and `.sptl`):

```bash
python -m spada.cli.gt4py_to_spatial samples/stencils.py 128,128,80 output/ --function-name laplacian
Expand Down Expand Up @@ -94,7 +99,7 @@ The runtime reads `metadata.json` generated by `sptlc` to determine the PE grid

### Example Kernels

Sample SPADA programs are in `samples/`:
Sample SpaDA programs are in `samples/`:

| Directory | Contents |
|---|---|
Expand Down Expand Up @@ -125,7 +130,7 @@ pytest tests/ --ignore=tests/csl_runtime

### CSL Runtime Tests (Singularity / Cerebras SDK)

End-to-end tests in `tests/csl_runtime/` compile and simulate SPADA programs using the Cerebras SDK and simulator. The Cerebras SDK ships as a Singularity Image File (`.sif`) and requires Singularity/Apptainer and an x86_64 Linux environment. Follow the Cerebras installation guide for full details: [Installation and Setup](https://sdk.cerebras.net/installation-guide).
End-to-end tests in `tests/csl_runtime/` compile and simulate SpaDA programs using the Cerebras SDK and simulator. The Cerebras SDK ships as a Singularity Image File (`.sif`) and requires Singularity/Apptainer and an x86_64 Linux environment. Follow the Cerebras installation guide for full details: [Installation and Setup](https://sdk.cerebras.net/installation-guide).

**Linux or x86_64 VM setup**

Expand All @@ -141,7 +146,7 @@ This saves the tarball to `tests/csl_runtime/cerebras-sdk.tar.gz` and extracts i
3. Install Python dependencies for the compiler:

```bash
python3 -m pip install -r requirements-ci.txt
python3 -m pip install -e ".[dev]"
```

4. Verify the toolchain:
Expand Down Expand Up @@ -217,7 +222,7 @@ make -C tests/csl_runtime clean-sdk # also remove the downloaded SDK

Questions, discussions, and feedback are welcome via GitHub Issues:

- **Bug reports and feature requests**: [GitHub Issues](https://github.com/glukas/spada/issues)
- **Bug reports and feature requests**: [GitHub Issues](https://github.com/spcl/spada/issues)

---

Expand All @@ -241,6 +246,6 @@ For significant changes (new language constructs, compiler passes, or architectu

## Release

SPADA is released under BSD-3-Clause License, see [LICENSE](LICENSE) for details.
SpaDA is released under BSD-3-Clause License, see [LICENSE](LICENSE) for details.

LLNL-CODE-2000963
97 changes: 97 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
[build-system]
requires = ["setuptools>=77"]
build-backend = "setuptools.build_meta"

[project]
name = "SpaDA"
description = "SpaDA: a spatial dataflow architecture programming language and compiler"
readme = "README.md"
requires-python = ">=3.10"
license = "BSD-3-Clause"
license-files = ["LICENSE", "NOTICE"]
authors = [
{ name = "Lukas Gianinazzi" },
{ name = "Tal Ben-Nun" },
{ name = "Torsten Hoefler" },
]
keywords = [
"stencil",
"compiler",
"spatial",
"dataflow",
"high-performance-computing",
"csl",
"cerebras",
]
classifiers = [
"Development Status :: 3 - Alpha",
"Intended Audience :: Developers",
"Intended Audience :: Science/Research",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Topic :: Scientific/Engineering",
"Topic :: Software Development :: Compilers",
]
dependencies = [
"click>=8.0",
"lark>=1.3.1",
"networkx>=2.8",
"numpy>=1.24",
]
dynamic = ["version"]

[project.urls]
Homepage = "https://github.com/spcl/spada"
Repository = "https://github.com/spcl/spada"
Paper = "https://arxiv.org/abs/2511.09447"

[project.scripts]
sptlc = "spada.cli.compiler:compile_spatial_ir"
sptlc-appliance = "spada.cli.appliance_compiler:compile_spatial_ir_appliance"
spada-wse-launcher = "spada.runtime.appliance_launcher:launch"

[project.optional-dependencies]
# Standalone placement research code (spada/placement, tests/placement).
placement = [
"hilbertcurve>=2.0.5",
"igraph>=0.11.4",
"matplotlib>=3.7",
]
# Graph rendering through igraph; needs the system cairo libraries.
render = ["pycairo>=1.26"]
docs = [
"mkdocs>=1.5.3",
"mkdocs-material>=9.2.7",
"mkdocs-material-extensions>=1.2",
"pymdown-extensions>=10.2.1",
]
dev = [
"SpaDA[placement]",
"black",
"flake8",
"isort",
# Backs networkx's nx_pydot dot export, used only by commented-out debug
# dumps of the completion DAG (spada/syntax/csl/tasks.py).
"pydot",
"pytest",
"pytest-cov",
"yapf",
]
all = ["SpaDA[dev,docs,placement,render]"]

[tool.setuptools.dynamic]
version = { attr = "spada.__version__" }

[tool.setuptools.packages.find]
include = ["spada*"]

[tool.setuptools.package-data]
spada = ["**/*.lark", "assets/csl/sync/*.csl"]

[tool.black]
line-length = 120
target-version = ["py310"]

[tool.isort]
profile = "black"
line_length = 120
14 changes: 0 additions & 14 deletions requirements-ci.txt

This file was deleted.

15 changes: 0 additions & 15 deletions requirements.txt

This file was deleted.

Loading
Loading