Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OPERATOR: The Software Instantiation of Electric Adaptability

Sehaj Singh, independent researcher. Research portal

Paper (PDF) · Live site · Replication

Electric Adaptability is the name I give to one research thesis: that a machine should be allowed to change while it runs, and that permission should be granted by a proof rather than by a test suite. This repository is the physical-deployment layer of that thesis. It is a 100-state, single-input electromechanical plant whose torque efficiency decays nonlinearly with winding temperature and accumulated wear, controlled by a learned operator that carries a per-step control-Lyapunov certificate through the entire degradation.

It is a numerical study, not hardware. Every number below is produced by python run_all_experiments.py on this machine and is reported as measured, including where the operator costs more than the classical baseline.


The autonomous adaptation paradox

Control engineering gives you two options and makes you pick one.

A fixed classical law is analysable. You can bound a PID or an LQR gain at its design point, hand a regulator to a certification body, and defend every pole. What you cannot do is keep that guarantee once the plant leaves the point it was synthesised at. In this repository efficiency follows eta(x) = exp(-k_T (T - T_amb)_+ - k_W W), so a gain computed cold and unworn under-actuates several-fold once the machine is hot and worn. The guarantee did not fail. The plant it described stopped existing.

A learned controller adapts across the full high-dimensional state and gives up the guarantee to do it. It is deployed as a black box, and there is no way to watch it drift toward a boundary until it has already crossed one.

This is the AI safety void in cyber-physical systems. It is not the void where a model says something harmful. It is the void where a learned policy holds authority over a physical actuator and nothing in the loop can state, at the instant a command is issued, whether that command decreases a certified energy. The void is filled by a construction, not by more testing: put the high-dimensional learning inside a physics wrapper that is provably stable on a sub-manifold, and give the wrapper the final signature on every command.

The closed loop: design, deploy, adapt, improve

The four stages are not a slogan here. Each is a specific mechanism in this repository, and each has a file.

Design. The certificate is constructed so that its guarantees hold for any parameter values, before a single gradient step. V(z; c) = relu(raw(z;c) - raw(0;c)) + eps ||z||^2 gives V(0) = 0 exactly and V >= eps ||z||^2 structurally. Convexity in the regulated coordinates survives because the hidden weights are clamped non-negative in the forward pass, not merely regularised toward it. No training run can break either property (models.py).

Deploy. The learned command is never trusted directly. A closed-form min-norm projection repairs it onto the decay half-space grad_z V . f_z + alpha V <= 1e-4. Because the regulated dynamics are exactly affine in the scalar input, that projection is one dot product, analytic and differentiable, with no solver in the control loop (controllers.py).

Adapt. Temperature and wear are not disturbances to be rejected. They are scheduling coordinates that reshape the certified energy bowl itself, entering through a non-negative rank-1 softplus gate W_k^eff(c) = W_k^+ * (r_k(c) x q_k(c)) evaluated on the 10 Hz slow boundary and cached in between. The safety envelope ages with the machine (models.py).

Improve. The gap between where a policy is trained and where it actually operates is closed by making the policy train on its own closed-loop states. Every 10 iterations the current controller is rolled out in deployment mode and the visited states are aggregated into half of every subsequent batch. This is the mechanism that took nominal control effort from an over-actuating 182 down to 22, and nominal RMSE from +171% relative to LQR down to +2.7% (train.py).

Ecosystem lineage

This repository is OPERATOR, the full system. Two lines converge here. The verification methodology comes from certifying a fixed physical network on the power grid. The aging-plant question was piloted at small scale on a DC motor. This is where both arrive: a 100-state multi-rate plant, a sub-manifold certificate, and a hardware-realistic feedback path.

1. Foundational methodological premise. certified-neural-lyapunov applies CROWN plus a hand-rolled input-space branch-and-bound to the Lyapunov conditions of Cui and Zhang's L-CSS 2022 neural controller on a lossy Kron-reduced NE39 power-grid swing model, certifying over a region rather than over samples. Its governing result is inherited here and never re-argued: on that model, 99.98% sampled satisfaction of the decrease condition still admits genuine counterexamples that a directed search finds and 200,000 random samples do not. That is why nothing in this repository certifies by sampling, and why the decay condition is checked at every applied command rather than estimated over an ensemble.

2. Verifiable constrained training core. certified-training-lyapunov trains the network to be verifiable by construction, differentiating through the CROWN bound so the size of the certified region is part of the objective. On the same lossy swing model it certifies a larger region than the prior counterexample-guided method at the same rate. That work grows the region a bound can prove. This work holds the region fixed and moves the plant.

3. Small-scale pilot. Operator Mini Test, archived in full at archive/operator_mini_test_pilot/, is the precursor to this repository and the first place the aging question was asked directly. That pilot line was a semifinalist in an MIT competition. It is a Lyapunov-structured neural controller on a wearing permanent-magnet DC motor, reading wear and winding temperature as explicit inputs, with a fast vectorized interval branch-and-bound certifier written because CROWN was too slow on CPU for stiff electrical dynamics.

Architectural progression, 2 states to 100

Pilot: Operator Mini Test This work: OPERATOR
Plant PM DC motor, 2-state error subsystem 100-state actuator, 98 regulated + 2 monotone
Clock single rate, idealised feedback 200 Hz mechanical / 10 Hz thermal, asynchronous
Certificate fixed envelope, held flat across aging sub-manifold ICNN reshaped by (T,W) gating
Verification interval BaB, cross-checked against CROWN per-step projection + full 100-state drift audit
Feedback path clean state access transport delay, ADC dead-band, saturation, drift
Training offline on-policy horizon aggregation, DAgger-style
Headline envelope flat across aging on 4 of 5 seeds, while fixed and gain-scheduled LQR lose 13% of their certified region 0 of 45 episodes diverged, certificate active 96.9% to 100% of steps

The pilot established that a certificate can be held flat across an aging trajectory at all, and it fixed the scope of the claim: in closed-loop simulation both the operator and the fixed LQR recover from the same disturbances, so the edge there is a larger provable region rather than a more stable controller. OPERATOR inherits that discipline and scales the question to a 100-state asynchronous multi-rate operator graph, replacing the flat envelope with one that reshapes as the machine ages, and running it through the sampled-data hardware path the pilot did not model. The pilot's own open problem, training reliability, is what on-policy horizon aggregation addresses here.

What the deployment layer adds. A swing model stands still: topology, damping, and transfer admittances are constants, and the hard problem is region size. A real actuator does not stand still, and degradation introduces a structural obstruction with no analogue in the fixed-parameter problem. Wear satisfies dW/dt >= 0 along every trajectory, so no Lyapunov function positive-definite in wear can decrease, ever. That is not a training difficulty to be optimised away. It forces the certificate onto the 98 regulated coordinates with the 2 monotone ones admitted only as parameters, and it forces an explicit audit of the context-drift term the restriction sets aside. This layer also adds the sampled-data feedback path a bench imposes: multi-rate clock, transport delay, ADC dead-bands, actuator saturation, and sensor drift.

Three-pillar architecture

I. Multi-domain perception. The state splits by bandwidth, not by convenience. A fast mechanical sub-manifold in R^98 (position, velocity, 96 latent coupling modes) runs at 200 Hz under RK4 with the slow context frozen at every integration stage. A slow thermal/wear subsystem in R^2 commits on a 10 Hz boundary, with wear ratcheted so integration error can never walk it backwards. The heterogeneous raw state, radians against degrees Celsius against dimensionless wear, is whitened by physical scale and lifted through a random-Fourier map onto a 128-dimensional coordinate shell before any learned layer sees it.

II. Continuous operator modelling. The 128 embedding coordinates are treated as graph nodes, and the command is produced by kernel-integral message passing: v <- sigma(W v + (1/N) sum_y kappa(x,y) v(y)), where kappa is an MLP of the coordinate pair. The operator is therefore discretisation-aware rather than a free N x N matrix, and the 1/N factor is the quadrature weight for the integral over the node domain. The kernel depends only on node coordinates, so it is built once and cached in eval, which is what makes a per-tick operator evaluation affordable in a 200 Hz loop.

III. Parameterized sub-manifold safety envelope. A partially input-convex Lyapunov network, convex in R^98 and reshaped by the R^2 context through the cached rank-1 softplus gate. The enforced condition is relaxed to grad_z V . f_z + alpha V <= epsilon with epsilon = 1e-4, which is practical stability rather than asymptotic convergence, and it buys an exact ultimate boundedness result: by the comparison lemma V(t) <= e^{-alpha t} V(0) + (epsilon/alpha)(1 - e^{-alpha t}), so limsup V <= 2e-4 and the sublevel set is forward invariant and attracting. The relaxation exists because sub-epsilon violations from sensor quantization would otherwise make the projection chatter against numerical noise.

The 15-seed empirical truth matrix

Read directly from results_summary.csv. Scenario A is cold and unworn, B is 95 °C at wear 0.55 where efficiency falls to roughly a fifth of nominal, C is warm and worn with sensor and actuation noise. Episodes are 10 s at 200 Hz, 15 seeds each.

Scen Controller RMSE (rad) Overshoot (rad) Effort ∫u²dt Settle (s) Cert % Full % Proj % Div
A PID 0.1274 ± 0.0065 0.2613 ± 0.0147 27.40 ± 2.82 3.95 – – – 0/15
A LQR 0.1166 ± 0.0066 0.0007 ± 0.0005 25.29 ± 1.77 0.81 – – – 0/15
A Operator 0.1198 ± 0.0067 0.0013 ± 0.0035 22.01 ± 1.59 0.96 100.0 100.0 0.0 0/15
B PID 0.1959 ± 0.0106 0.4627 ± 0.0263 68.11 ± 7.54 4.40 – – – 0/15
B LQR 0.1527 ± 0.0085 0.1842 ± 0.0134 99.18 ± 8.36 2.52 – – – 0/15
B Operator 0.1496 ± 0.0079 0.0258 ± 0.0312 112.86 ± 10.56 2.19 99.9 98.5 3.4 0/15
C PID 0.1554 ± 0.0084 0.3587 ± 0.0198 62.48 ± 4.05 4.52 – – – 0/15
C LQR 0.1297 ± 0.0076 0.0524 ± 0.0066 56.68 ± 3.83 1.40 – – – 0/15
C Operator 0.1397 ± 0.0079 0.0099 ± 0.0052 32.44 ± 4.93 1.74 96.9 94.7 7.0 0/15

Cert % is the fraction of steps whose applied command satisfies the relaxed decay condition on the regulated sub-manifold. Full % is the same under a 100-state audit that adds the context-drift term grad_c V . c_dot, the quantity the sub-manifold restriction sets aside. Proj % is how often the raw learned command needed repair. PID and LQR carry no certificate, so those columns do not apply.

The honest Pareto boundary

  • Scenario A, nominal. RMSE +2.7% versus LQR, control effort 22.01 against 25.29, which is 13.0% lower energy. Certificate active on 100% of steps by both the reduced and the full audit, with the projection never once intervening. LQR is nominally tighter on overshoot (0.0007 against 0.0013), a difference well inside the operator's own ±0.0035 seed spread.
  • Scenario B, thermal and wear. RMSE −2.1%, so the operator outperforms LQR in the regime that invalidates LQR's design point. Overshoot 0.026 against 0.184 is 7.1x peak protection. Certificate active on 99.9% of steps (98.5% full). This is also where the cost lives: effort 112.86 against 99.18, 14% more energy than LQR. Holding a hot, worn plant on a tight, low-overshoot trajectory is paid for, and that number is not softened anywhere in this repository.
  • Scenario C, stochastic. RMSE +7.7%, overshoot 0.010 against 0.052 for 5.3x protection, effort 32.44 against 56.68 which is 43% lower. Certificate active on 96.9% of steps (94.7% full), the lowest of the three because the state estimate is noise-corrupted.
  • Robustness log. 0 of 45 operator episodes diverged, across every scenario and seed.

The order-of-magnitude result is peak overshoot under load, which is the axis that drives mechanical fatigue. Tracking is a wash within a few times the seed spread. Control energy is genuinely mixed: better in two regimes, worse in the degraded one. pareto_frontier_data.csv carries the per-episode point cloud with a dominated_by column flagging any operator point a classical controller beats on both overshoot and effort at once. Of 135 rows, exactly 2 operator points are dominated, both in the nominal scenario where the overshoots being compared are of order 1e-3 rad.

Replication

python run_all_experiments.py

One command, from an empty checkpoint through training, evaluation, figure generation, and site synchronisation. It writes operator_ckpt.pt, results_summary.csv, pareto_frontier_data.csv, results_episodes.csv, nature_figure_1.pdf / .png, and refreshes the web portal in docs/.

NOC_ITERS=35 NOC_SEEDS=2 python run_all_experiments.py   # smoke test

Training auto-resumes past the intermittent native OpenMP abort this conda stack throws, so the single command completes unsupervised. Evaluation is deterministic: re-running the scenario-A sweep reproduces RMSE 0.119793174511929 bit-for-bit and matches the committed summary table to full double precision.

Alignment with Laude Institute's Slingshot criteria

Laude Institute resources computer scientists who ship research as open-source infrastructure rather than as a paper alone. Measured against that bar, this repository is a single-command reproduction with no manual steps, a deterministic evaluation harness, a results table generated from the artifacts rather than transcribed into them, and a written record that states its costs and its scope limits in the same place it states its wins. The engineering claim is verifiable by a reviewer with a Python install and an afternoon.

Files

  • environment.py: plant, multi-rate dynamics, hardware artifacts, scenarios.
  • models.py: manifold projection, graph operator, ICNN Lyapunov net, terminal cost.
  • controllers.py: PID, LQR, compound operator with the safety projection.
  • train.py: physics-informed training with on-policy aggregation.
  • batched_eval.py: vectorized 15-seed-per-scenario evaluation.
  • benchmark.py: evaluation driver; prints the mean +/- std matrix, writes CSVs.
  • generate_nature_figures.py: the two-panel figure.
  • run_all_experiments.py: one-command replication, plus web-portal sync.
  • manuscript_draft.txt: draft write-up (methods, results, limitations, scope).
  • index.html, styles.css: the web portal, deployable from docs/.

Environment notes

  • Needs torch, numpy, scipy, pandas, matplotlib (requirements.txt; gymnasium optional). On the development machine that is the Anaconda interpreter.
  • run_all_experiments.py, train.py, and benchmark.py set KMP_DUPLICATE_LIB_OK=TRUE before importing torch. Without it the conda and MKL stack aborts at import with OMP: Error #15.
  • train.py checkpoints every 25 iterations so an interrupted run costs at most 25.

License

MIT, see LICENSE. The archived pilot in archive/operator_mini_test_pilot/ is covered by the same terms. Reviewers and researchers are free to audit, reuse, modify, and redistribute the control layers and the replication pipeline, with attribution.

Scope

One single-input simulated plant, whose wear law is a modelling choice with no hardware grounding. The certificate is a guarantee about a 98-coordinate subsystem; the two monotone aging coordinates are audited, not certified. The ultimate boundedness proposition holds along trajectories where the relaxed condition is actually enforced, which is 96.9% to 100% of steps, not all of them. These results support a construction and a measured trade-off, not a claim about a specific physical machine. A motor test bench with real sensing, real power electronics, and a measured degradation trajectory is the next step, and nothing here substitutes for it.


Contact. sehajrsinghs@gmail.com · GitHub · Portal

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages