Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,28 @@
## 0.8.0 - 2026-09-27

### Added

* Opt-in ring band-limited far-field evaluation for ordinary fields, both
embedded normalizations, and isolated-element characterization. Select
`farFieldEvaluator: "ring"` at model or array-solver creation, or override
individual requests with `evaluator: "ring"` or `"exact"`.
* A shared C++ evaluator supports free space and perfect ground, with
deterministic sparse sampling, interpolation, and direct fallback. Worker
pools schedule complete rings and preserve the existing output layout.
* `fieldEvaluation` metadata identifies the approximation, execution mode,
direct-evaluation counts, analytical truncation bound, and fallback reason.
Packed characterization handoffs carry this provenance alongside the buffers.
* Reproducible native/WASM profiling, accuracy, and integration benchmarks in
`docs/ring-far-field-evaluation.md`. The large-array stateful WASM cases improve
by 6.76–8.91×; all 154 comparisons meet the 1e-7 relative-error target.

### Compatibility

* Exact evaluation remains the default. Existing exact ABI entry points,
numerical goldens, normalization, packed NECQ/NECF formats, and retained
factorization behavior are preserved. NEC2++ remains `2.5.0`, and WASM ABI
remains `1`; ring entry points are additive.

## 0.7.1 - 2026-09-23

### Improved
Expand Down
33 changes: 33 additions & 0 deletions bench/ring-far-field/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# Ring far-field bench

See [the report](../../docs/ring-far-field-evaluation.md) for the derivation,
measured results, limitations, and exact build/test/run commands.

The harness is outside the production package; it now exercises the shared
production kernel. The original prototype and its measurements are preserved in
commit `970d1ed` and the original `evidence/` files.

- `ring.hpp`: thin adapter to the shared production C++ kernel.
- `bench.cpp`: full NEC solves, phase timings, direct/ring comparisons and gates.
- `test.cpp`: deterministic interpolation and fallback tests, native and WASM.
- `build.py`: separate binaries with phase probes in a temporary source copy.
- `run.py`: serial native/WASM measurements and environment metadata.
- `summarize.py`: report tables from raw observations.
- `profile-package.mjs`: initial public-package relevance probe.
- `package-production.mjs`: packaged worker timings and embedded/packed verification.
- `evidence/`: original observations; `evidence/production/` records integration measurements.

The production comparison requires Emscripten **4.0.7**, matching the package
build. Run timing commands sequentially on an otherwise idle machine. The
embedded package benchmark uses the same seven geometries on a compact
3 × 181 grid to keep the 256-port basis arrays practical; ordinary fields and
worker timings use each geometry's full visualizer grid.

Standalone tests now link the unchanged direct kernel explicitly:

```sh
g++ -std=c++17 -O3 -Isrc -Isrc/eigen -I/tmp/ring-production-native \
bench/ring-far-field/test.cpp src/nec_field_evaluator_wasm.cpp \
-o /tmp/ring-test
/tmp/ring-test
```
149 changes: 149 additions & 0 deletions bench/ring-far-field/bench.cpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
// Analysis-only harness; not linked into any production target.
#include "nec_stateful_model.h"
#include "electromag.h"
#include "ring.hpp"
#include <algorithm>
#include <chrono>
#include <cmath>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <vector>
double ring_fill_ms=0, ring_factor_ms=0;
using Clock=std::chrono::steady_clock;
double ms(Clock::time_point t){return std::chrono::duration<double,std::milli>(Clock::now()-t).count();}
constexpr double pi_b=3.14159265358979323846;
int main(int argc,char**argv){
const std::string only=argc>1?argv[1]:"regular4";
const int rounds=argc>2?std::stoi(argv[2]):3;
const int side=only=="regular16"||only=="sunflower256"?16:only=="regular8"?8:4;
const bool ground=only!="free4";
const double lambda=em::get_wavelength(300e6);
for(int round=0;round<rounds;++round){
nec_stateful_model model;
std::vector<nec_port_definition> ports;
std::vector<double> xs,ys;
for(int i=0;i<side*side;++i){
double x=.5*(i%side-(side-1)/2.), y=.5*(i/side-(side-1)/2.);
if(only=="sunflower256") {double r=.5*std::sqrt(i/pi_b), phi=i*pi_b*(3-std::sqrt(5.)); x=r*std::cos(phi);y=r*std::sin(phi);}
// Nonzero horizontal centroid exercises removal and restoration of phase.
x+=.37;y-=.21;
double z=.5+(only=="heights4"?.11*std::sin(i*1.7):0);
double ux=0,uy=0,uz=1;
if(only=="tilted4"){ux=std::cos(i*1.3)*.8;uy=std::sin(i*1.3)*.8;uz=.6;}
model.add_wire({i+1,11,(x-.235*ux)*lambda,(y-.235*uy)*lambda,(z-.235*uz)*lambda,(x+.235*ux)*lambda,(y+.235*uy)*lambda,(z+.235*uz)*lambda,.001*lambda});
ports.push_back({i+1,6}); xs.push_back(x);ys.push_back(y);
}
model.complete_geometry(); model.define_ports(ports);
model.set_ground({ground?nec_ground_kind::perfect:nec_ground_kind::free_space});
const double d=std::hypot(*std::max_element(xs.begin(),xs.end())-*std::min_element(xs.begin(),xs.end()),*std::max_element(ys.begin(),ys.end())-*std::min_element(ys.begin(),ys.end()))+.47;
const double step=std::min(pi_b/36,.886/(d*8));
const int nt=std::ceil(pi_b/2/step)+1,np=std::ceil(2*pi_b/step);
const nec_far_field_grid grid{1,0,nt,90./(nt-1),0,np,360./np};
auto t=Clock::now();model.prepare(300);const double prepare=ms(t);
for(int steer=0;steer<2;++steer){
std::vector<nec_complex> voltages;
for(int i=0;i<side*side;++i)voltages.push_back(std::polar(1.,-2*pi_b*(xs[i]*(.12+steer*.03)+ys[i]*.19)));
t=Clock::now(); const auto solution=model.solve_port_voltages_detailed(voltages);const double solve=ms(t);
if(model.factorization_generation()!=1)throw std::runtime_error("LU not retained");
t=Clock::now();const auto& exact=model.compute_far_field(grid);const double direct=ms(t);
const auto snapshot=model.capture_far_field_snapshot();
t=Clock::now(); const auto tile=ring_bench::direct(snapshot,grid);const double tileMs=ms(t);
t=Clock::now(); const auto ring=ring_bench::evaluate(snapshot,grid);const double ringMs=ms(t);
double peakT=0,peakP=0,peak=0,errT=0,errP=0,interpolation=0,oracle=0,pDirect=0,pRing=0;int worstT=0,worstP=0;
for(size_t i=0;i<exact.e_theta.size();++i){
const int it=i%nt;
const double dt=std::abs(exact.e_theta[i]-ring.field.theta(i)),dp=std::abs(exact.e_phi[i]-ring.field.phi(i));
if(dt>errT){errT=dt;worstT=it;}if(dp>errP){errP=dp;worstP=it;}
peakT=std::max(peakT,std::abs(exact.e_theta[i]));peakP=std::max(peakP,std::abs(exact.e_phi[i]));
peak=std::max(peak,std::sqrt(std::norm(exact.e_theta[i])+std::norm(exact.e_phi[i])));
interpolation=std::max({interpolation,std::abs(tile.theta(i)-ring.field.theta(i)),std::abs(tile.phi(i)-ring.field.phi(i))});
oracle=std::max({oracle,std::abs(exact.e_theta[i]-tile.theta(i)),std::abs(exact.e_phi[i]-tile.phi(i))});
const double weight=std::sin(it*grid.theta_step_deg*pi_b/180)*(it==0||it==nt-1?.5:1.);
pDirect+=weight*(std::norm(exact.e_theta[i])+std::norm(exact.e_phi[i]));
pRing+=weight*(std::norm(ring.field.theta(i))+std::norm(ring.field.phi(i)));
}
const double scale=grid.theta_step_deg*pi_b/180*2*pi_b/np/(2*std::sqrt(4*pi_b*1e-7/8.854e-12));
// free4 has identical vertical elements at one height: lower-hemisphere
// intensity is its upper-hemisphere reflection, so double its integral.
pDirect*=scale*(ground?1:2);pRing*=scale*(ground?1:2);
const double closureDelta=(pRing-pDirect)/solution.power_budget.radiated_power_w;
if(!std::isfinite(errT)||!std::isfinite(errP)||errT>1e-7*peak||errP>1e-7*peak||std::abs(closureDelta)>1e-7)throw std::runtime_error("accuracy gate failed");
const auto contributions=exact.diagnostics.segment_direction_contributions;
auto ringGrid=grid;ringGrid.ring=true;
t=Clock::now();const auto& integrated=model.compute_far_field(ringGrid);const double integratedMs=ms(t);
for(size_t i=0;i<integrated.sample_count();++i)
if(std::abs(integrated.e_theta[i]-ring.field.theta(i))>1e-7*peak ||
std::abs(integrated.e_phi[i]-ring.field.phi(i))>1e-7*peak)throw std::runtime_error("integrated ring gate failed");
std::cout
<< std::setprecision(17)
<< "{\"case\":\""
<< only
<< "\",\"round\":"
<< round
<< ",\"steer\":"
<< steer
<< ",\"segments\":"
<< side*side*11
<< ",\"nt\":"
<< nt
<< ",\"np\":"
<< np
<< ",\"contributions\":"
<< contributions
<< ",\"prepareMs\":"
<< prepare
<< ",\"fillMs\":"
<< ring_fill_ms
<< ",\"factorMs\":"
<< ring_factor_ms
<< ",\"solveMs\":"
<< solve
<< ",\"directMs\":"
<< direct
<< ",\"tileMs\":"
<< tileMs
<< ",\"ringMs\":"
<< ringMs
<< ",\"integratedRingMs\":"
<< integratedMs
<< ",\"ringDirections\":"
<< ring.directions
<< ",\"fallbackRings\":"
<< ring.fallback
<< ",\"maxL\":"
<< ring.maxL
<< ",\"maxSegmentExtra\":"
<< ring.maxSegmentExtra
<< ",\"errorTheta\":"
<< errT/peak
<< ",\"errorPhi\":"
<< errP/peak
<< ",\"componentErrorTheta\":"
<< (peakT?errT/peakT:0)
<< ",\"componentErrorPhi\":"
<< (peakP?errP/peakP:0)
<< ",\"worstThetaRingDeg\":"
<< worstT*grid.theta_step_deg
<< ",\"worstPhiRingDeg\":"
<< worstP*grid.theta_step_deg
<< ",\"tileOracleError\":"
<< oracle/peak
<< ",\"interpolationError\":"
<< interpolation/peak
<< ",\"closureDirect\":"
<< pDirect/solution.power_budget.radiated_power_w
<< ",\"closureRing\":"
<< pRing/solution.power_budget.radiated_power_w
<< ",\"closureDelta\":"
<< closureDelta
<< ",\"tailBoundRelative\":"
<< ring.bound/peak
<< ",\"sourceCondition\":"
<< ring.sourceNorm/peak
<< "}"
<< std::endl;
}
}
}
63 changes: 63 additions & 0 deletions bench/ring-far-field/build.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
#!/usr/bin/env python3
"""Build an isolated harness. Production sources and artifacts are never written."""
import argparse
import concurrent.futures
import pathlib
import re
import subprocess

parser = argparse.ArgumentParser()
parser.add_argument('--compiler', default='g++')
parser.add_argument('--out', required=True)
parser.add_argument('--wasm', action='store_true')
parser.add_argument('--jobs', type=int, default=2)
args = parser.parse_args()
root = pathlib.Path(__file__).resolve().parents[2]
out = pathlib.Path(args.out).resolve()
out.mkdir(parents=True, exist_ok=True)
(out / 'config.h').write_text(
'#define NECPP_VERSION "2.5.0"\n'
'#define NECPP_BUILD_DATE "ring-bench"\n')
source = (root / 'src/nec_context.cpp').read_text()
# Time exactly the two operations without changing the public interface or sources.
for statement, phase in [
('cmset( neq, cm, rkh );', 'fill'),
('factrs(m_output, npeq, neq, cm, ip );', 'factor'),
]:
assert source.count(statement) == 1
source = source.replace(statement,
'{ const auto start = std::chrono::steady_clock::now(); ' + statement
+ ' ring_' + phase + '_ms = std::chrono::duration<double, std::milli>('
'std::chrono::steady_clock::now()-start).count(); }')
(out / 'nec_context.cpp').write_text(
'#include <chrono>\nextern double ring_fill_ms, ring_factor_ms;\n' + source)
names = re.search(r'set\(NECPP_LIB_SRCS(.*?)\)',
(root / 'src/CMakeLists.txt').read_text(), re.S)[1].split()
sources = [out / name if name == 'nec_context.cpp' else root / 'src' / name
for name in names]
sources.append(root / 'bench/ring-far-field/bench.cpp')
# No fast math, SIMD, native-arch, or parallel solver. Match scalar release flags.
flags = [
'-std=c++17', '-O3', '-DNDEBUG', '-fexceptions', '-flto',
'-DNECPP_FAR_FIELD_CACHE_DIRECTIONS=1', '-DNECPP_FAR_FIELD_REUSE_OUTPUTS=1',
'-I' + str(root / 'src'), '-I' + str(root / 'src/eigen'), '-I' + str(out),
]


def compile_source(path):
obj = out / (path.stem + '.o')
# Always compile: avoiding stale header/flag dependencies matters more than
# incremental-build speed in a reproducibility harness.
subprocess.run([args.compiler, *flags, '-c', str(path), '-o', str(obj)],
check=True)
return str(obj)


with concurrent.futures.ThreadPoolExecutor(max_workers=args.jobs) as pool:
objects = list(pool.map(compile_source, sources))
link = [
'-sSTACK_SIZE=4194304', '-sALLOW_MEMORY_GROWTH=1', '-sENVIRONMENT=node',
'-sDISABLE_EXCEPTION_CATCHING=0',
] if args.wasm else []
subprocess.run([args.compiler, *flags, *objects, *link, '-o',
str(out / ('bench.js' if args.wasm else 'bench'))], check=True)
14 changes: 14 additions & 0 deletions bench/ring-far-field/evidence/metadata.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"date": "2026-09-26T23:41:03+0200",
"commit": "09e4d413d835285a6631216baf8b2172f08651e5",
"platform": "Linux-7.0.0-31-generic-x86_64-with-glibc2.43",
"cpu": "13th Gen Intel(R) Core(TM) i7-1355U",
"node": "v24.20.0",
"g++": "g++ (Ubuntu 15.2.0-16ubuntu1) 15.2.0",
"rounds": 3,
"nativeSha256": "b4349498d658a6b18ecdbe96f83136a6317d15d587261437f301ef7848f7e3d7",
"benchWasmSha256": "0ecce8df6570e5e42de0d61276f6da0a4dc75c26f9f4522e88a3385372f517bb",
"packageWasmSha256": "745631fc178b541c5eca44688d95fee72fdafac306227931f7f4669335106b90",
"wasmCompiler": "Emscripten 4.0.10; see build.py for exact flags",
"protocol": "3 fresh models; one cold + one retained voltage RHS per model; scalar; explicit geometry; serial execution; module/process startup excluded"
}
Loading
Loading