Skip to content

Repository files navigation

PPO Robot Controller

CI licence: MIT language: Python 3.10+ tests: 39 status: 94.4% held-out success (5 seeds)

A Proximal Policy Optimization implementation (PyTorch, no RL library) for a PyBullet reaching task, with a hand-derived GAE bootstrap tested against an independent reference, a tanh-squashed Gaussian policy verified against the PPO ratio identity, multi-episode rollout collection with minibatched updates, and a workspace bound verified by sabotage.

Result

On the full random-target task (target sampled uniformly from [-1, 1] x [-1, 1] each episode), performance is measured across 5 independent training runs (--seed 0 through --seed 4, 3000 episodes each, otherwise identical hyperparameters) rather than a single run, each evaluated on the same held-out protocol (deterministic policy, frozen observation-normalizer statistics, 50 episodes, seed 1000):

seed held-out success mean reward
0 46/50 (92%) -30.01 +/- 37.45
1 47/50 (94%) -27.73 +/- 34.67
2 47/50 (94%) -28.40 +/- 37.17
3 46/50 (92%) -34.08 +/- 52.14
4 50/50 (100%) -22.08 +/- 19.63

Mean 94.4% +/- 3.3% held-out success across the 5 seeds (pooled: 236/250 episodes), with every individual seed at 92% or higher, the result holds consistently, not just on one lucky run. Raw per-seed numbers are in results/seed_sweep_summary.csv.

The training curve backs this up: chunked into ten 300-episode blocks, mean reward for a representative run climbs from -172 in the first block to a stable -20 to -30 range by the back half, with success rate rising alongside it. A linear fit over the ten chunks gives a positive slope (r=0.649); the other four runs show the same shape.

Training reward on the full random-target task: the 50-episode rolling mean climbs from about -400 to about -25 within the first 200 episodes and stays there for the remaining 3000 episodes

Real training run, full random-target task, 3000 episodes. See docs/design.md for the full numbers.

RobotReachEnv rendered in PyBullet: the r2d2 stand-in robot and a red target marker on a checkered ground plane

PyBullet render of a real reset of RobotReachEnv. The target is drawn as an enlarged red marker for visibility, the actual sphere_small.urdf collision/visual target is a few millimetres across at this scale and would not read as a dot in a screenshot.

How it works

RobotReachEnv (PyBullet, Gymnasium API) rewards the negative planar distance to a random target, with a bonus for reaching it. The "robot" is a fixed-base r2d2.urdf mesh whose base position/velocity is set directly rather than driven by joint torques or forces, see Limitations below for what that does and doesn't mean for the results. The target is a fixed-base body too, so it is a static landmark rather than something the robot can physically shove out of place. The robot's planar position is clamped to a fixed workspace bound so that a bad policy's reward is capped rather than growing unboundedly worse the longer it drifts.

PPOAgent implements PPO from the ground up: a shared-trunk actor-critic network (PolicyValueNetwork), a tanh-squashed diagonal Gaussian policy, clipped surrogate policy loss, GAE advantage estimation, and entropy regularization. Sampling, squashing into the environment's action bounds, and log-probability scoring (including the change-of-variables Jacobian correction) are computed consistently for the exact action returned to the caller, so the PPO importance ratio exp(log_pi_new(a) - log_pi_old(a)) is provably 1.0 for every stored transition before any optimizer step, checked directly by a dedicated test, including near-saturated action bounds where a naive implementation would disagree.

Updates run over a rollout buffer that can span several episodes (--rollout-steps, default 2048), shuffled into minibatches (--minibatch-size) across --ppo-epochs epochs; the GAE recursion is cut at every episode boundary inside the buffer so one episode's advantage never leaks into another's. Checkpoints save network weights, optimizer state, configuration, training progress, and observation-normalizer statistics together, so evaluation can restore and freeze the exact statistics used during training; evaluation is deterministic by default (--stochastic is explicit). No RL library (Stable-Baselines3, RLlib, ...) is used for the algorithm itself, only PyTorch, Gymnasium (for the environment API), and PyBullet (for the physics).

Before spending a full training budget on the random-target task, a cheap fixed-target overfit diagnostic (one static goal, a few hundred episodes) is a useful sanity check that the whole pipeline can learn at all: on this implementation it converges cleanly, reward climbing from around -350 to around -15 over the first ~230 of 500 episodes, and a held-out evaluation on that fixed goal resolves 50/50.

Fixed-target overfit diagnostic: episode reward climbing smoothly from about -350 to about -15 over 230 episodes on a single static goal, then staying converged for the rest of the 500-episode run

The cheaper diagnostic run: a single fixed goal, 500 episodes, 50/50 held-out at convergence.

Installation

git clone https://github.com/andrealo20/ppo-robot-controller.git
cd ppo-robot-controller
pip install -r requirements.txt

Quick start

Run everything as a module, from the repository root (python src/train.py does not work, running a script directly puts src/ itself on sys.path, not the repository root, so from src.x import y has nothing to resolve against):

python -m src.train --num-episodes 3000 --lr 1e-4 --rollout-steps 2048 --minibatch-size 64 --seed 42
python -m src.evaluate --model-path experiments/checkpoints/best_model.pt --episodes 50 --seed 1000
python -m pytest tests -q

# Cheap overfit diagnostic on one fixed goal before a full run
python -m src.train --num-episodes 500 --fixed-target 0.6 0.4 --output-dir experiments/fixed_target

--seed seeds NumPy, PyTorch, and the environment for a reproducible run; omit it to train unseeded. src.evaluate's own --seed is separate from the training seed: it controls which held-out targets are sampled and seeds the evaluation process's NumPy and PyTorch generators (so that --stochastic runs repeat exactly), independent of how the model being evaluated was trained.

Repository layout

src/
├── environment/    reaching_env.py, the PyBullet/Gymnasium environment
├── agent/          ppo.py, the PPO agent (tanh-squashed Gaussian policy, multi-episode rollout buffer, GAE, minibatched clipped update)
├── network/        policy_value.py, shared-trunk actor-critic network
├── utils/          running_normalizer.py, running observation normalization
├── train.py        training entry point + collect_rollout (python -m src.train)
└── evaluate.py      evaluation entry point (python -m src.evaluate)

tests/              39 tests: GAE reference + multi-episode leak + sabotage
                    checks, PPO ratio-identity invariant, rollout collection
                    bookkeeping, network initialization, agent smoke tests,
                    environment terminated/truncated contract + workspace
                    bound, a hand-coded oracle solvability check, running
                    normalizer
docs/design.md      architecture and design rationale
conftest.py         empty; exists only so pytest puts the repo root on
                    sys.path (see docs/design.md)
experiments/        checkpoints and TensorBoard logs land here at runtime
results/            seed_sweep_summary.csv, per-seed held-out results
                    backing the table above
assets/             reaching_env.png, training_reward_fixed_target.png,
                    training_reward_random_target.png,
                    seed_sweep_success_rate.png, real renders and real
                    training data, not mockups

Limitations

  • The task is not solved outright. 94.4% +/- 3.3% held-out success across 5 independent 3000-episode runs is a strong, checked result, not a guarantee, the worst seed still fails 4/50 held-out episodes, and 3000 episodes is still a modest budget by continuous-control standards.
  • The reaching task is kinematic, not dynamic. The robot's base position/velocity is set directly each step; gravity and rigid-body dynamics play no role. This is a 2D point-reaching problem with a robot mesh on top, not a locomotion or manipulation task, nothing here should be read as evidence about either.
  • No parallel environments. Rollouts are collected from a single environment instance, serially.

References

Licence

MIT, see LICENSE.

About

From-scratch PPO (PyTorch) on a PyBullet reaching task - GAE bootstrap verified against a hand-derived reference

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages