Source for nstarkman.space — an Astro site, a CV database, and the tooling that renders one into the other.
Everything comes out of data/. One JSON file per item, and the same database
renders the website, the profile README, the CV PDFs, and
nstarkman_publications.bib. Adding a paper or
a conference is one small file.
Adding something? Read AGENTS.md.
data/ one item per file, data/<date.start>-<id>.json — the source of truth
data/lists/ bare enumerations (peer review) as plain arrays
config/ things that are configuration rather than items: the CV presets,
the person, and the generated contributions list
schema/ the item JSON Schema, plus invalid fixtures the tests assert on
src/ the Astro site; src/styles/ is the stylesheet, split by page
cv/ the Typst CV template — cv.typ, lib/ (the palette and the
styles), and the portrait and QR it draws
tests/ vitest suites over src/lib and the generated data
scripts/ test, build and generation scripts
public/ served verbatim: CNAME, files/starkman_cv.pdf, and the fonts
the CV is compiled from
npm install
npm run dev # local server
npm run build # static build into dist/
npm run build:all # the site, then the four CV PDFs (needs typst)
npm test # everything below, in order; CI runs it on every pull requestnpm test is the whole gate, not a subset — test:unit (which includes the schema checks),
test:bibtex, a build, test:a11y, test:links. Each runs on its own too:
npm run test:unit # vitest over src/lib and the generated data
npm run test:schema # every record against schema/, including the invalid fixtures
npm run test:bibtex # escaping, names, case protection, maths in abstracts
npm run test:a11y # every link and heading on every built page has a name
npm run test:links # every internal link, anchor, published record and PDF resolves
npm run validate # the schema check alone, against data/*.jsontest:a11y and test:links read dist/, so they need a build first — which is
why npm test builds in the middle rather than at the end.
Five files, each imported by what renders it, so a page ships only the rules it
can use: the home page carries no CV and no map, and the CV carries no map from
/research/.
| file | imported by | what is in it |
|---|---|---|
src/styles/global.css |
src/layouts/Base.astro, so every page |
the site: the palette and the type stack, the page frame, nav and footer, the buttons and marks every page uses, the Embed Builder's pill rows, an item as the site renders it, the cards, and the CV pieces other pages reuse (below) |
src/styles/map.css |
CollabMap.astro, ConfMap.astro |
the two world maps on /research/ |
src/styles/cv/page.css |
CvView.astro |
the CV as /cv/ and /cv/<preset>/ render it — the ruled headings, the sticky bar, the length slider, the source marks |
src/styles/cv/minimap.css |
CvView.astro |
the collaborator and conference maps in the CV's right margin, and the tint on a row that answers what was picked in one — a paper written with that person, a talk given in that place |
src/styles/cv/builder.css |
src/pages/tools/cv.astro |
what /tools/cv/ adds: a tick box on every entry and every elaboration line, the per-section selectors, the compile controls |
Order is the cascade: Base.astro loads global.css before any page or
component adds its own, so the site comes first and the CV layers on top of it.
The test for whether a rule belongs in cv/ is whether anything outside /cv/
renders it, not whether the CV happens to use it. A dated row — .tl — stays in
global.css even though the CV is full of them, because the home page's
Background section is built from the same rows; so do its date gutter and
detail links, the contents rail (which /publications/ reuses for years), and
the mini software cards (/software/'s long tail). Only the variants the CV
alone has (.tl--pub, .tl--unpub, .tl--pick) sit under cv/.
The four pre-built PDFs and the in-browser builder at /tools/cv/ compile the
same template from the same render model, so the PDF and the site cannot
disagree about what a preset contains. Under CI the Typst CLI reads cv.json
off disk; in the browser typst.ts is handed the same filename through its
virtual filesystem. Nothing forks.
| file | what is in it |
|---|---|
cv/cv.typ |
the document — the header, the shape of a section, and the loop over cv.json. Names no colour, no font and no glyph. |
cv/lib/theme.typ |
the palette, and the two figure styles that are not the document's old-style default |
cv/lib/styles.typ |
the icon fonts, the mark table, and one unit per style |
cv/ is the compilation root in both runtimes, so cv/lib/styles.typ is
/lib/styles.typ to typst.ts and the relative imports resolve unchanged either
way. The builder globs /cv/**/*.typ into that virtual filesystem, so a new
module needs no change there.
The builder adds a style menu beside Compile PDF:
| style | what it is |
|---|---|
| Default | what the pre-built PDFs are. Font Awesome and Academicons marks on links, section headings and the contact block. |
| Default - 🎨 | the same CV with none of the marks. Where a glyph stood alone the word takes its place — a resource trail reads code, docs rather than two icons — and the PDF embeds no icon fonts at all. |
| adrn | a close mimic of adrn/cv. Lato at 12pt rather than a serif, steel-blue headings over a rule of the same blue, US Letter, no portrait or QR, dates inline instead of in a gutter, and a running head and "Last updated" foot. Values are taken from that repo's apw-cv.cls, not eyeballed. |
A style is one key on cv.json:
{ "style": "plain" }read by cv.typ as cv.at("style", default: "default"). Only the builder
sets it. The CLI never writes it, so every pre-built PDF is the default and
the one- and two-page contracts cannot be affected by a style. plain holds at
one and two pages too, though that was not a given — words are wider than
glyphs. adrn does not, and cannot: US
Letter with 1in margins is about a third less text area than this A4 setup, so
the one-page preset runs to two. That is a property of the design being copied,
not something spacing can recover.
A style is a self-contained unit rather than a branch: one entry in STYLES
in cv/lib/styles.typ, answering three questions and
nothing else.
| the unit answers | asked by |
|---|---|
glyph(name) |
a mark drawn beside words — a section heading, a contact line. none draws nothing. |
solo(name, word) |
a mark standing alone, given the word it was standing in for — a bare arXiv or paper link on a publication. Without a deliberate answer here a style compiles an empty link. |
trail(links) |
a run of marks — the code, docs, data trail at the end of an entry. |
marked(), a mark with words beside it, is derived from glyph rather than
restated, so a style cannot answer those two inconsistently. Everything else is
shared — the same type, the same spacing, the same sections, off the same
model — which is why adding a style is that entry plus an <option> in
src/pages/tools/cv.astro, touches cv.typ not at all, and never becomes a
second template.
Most of data/ is written by hand. One file is not:
config/contributions.json — the open-source repositories with merged pull
requests that are not mine and have no Software card of their own.
node scripts/collect-contributions.mjs # rewrite the list
node scripts/collect-contributions.mjs --check # exit 1 if anything is newIt walks merged pull requests oldest-first through fixed date windows — the
GitHub search API caps a query at 1000 results — so a repository's first
appearance really is the first contribution. Forks, anything already rendered
as a Software card, and anything under nstarman/ are dropped automatically;
config/contributions-exclude.json holds only the judgement calls the script
cannot make, such as papers and workshops that happen to be repositories.
.github/workflows/refresh-contributions.yml runs it monthly and opens a pull
request when something new appears — but only after the schema, build and a11y
gates pass on the regenerated file.
config/collaborators.json is the other generated file: where each co-author
has worked, and when, from their public ORCID employment records.
node scripts/collect-collaborators.mjs # rebuild
node scripts/collect-collaborators.mjs --check # exit 1 if it is out of dateOnly people already named as authors in data/ are looked up, and only their
ORCID is sent. ORCID is self-reported: 27 of 32 list any employment at all, and
an empty history is recorded as empty rather than guessed at.
Two sources answer "where were they when we wrote this", and they are not the same claim:
authors[].affiliationon a publication is what that paper printed.scripts/fetch-affiliations.mjsfills it from Crossref, only where it is missing and only where Crossref names one.- the employment history answers for everyone else, through
affiliationAt(orcid, date)insrc/lib/data.js, which returnsnullrather than a guess when nothing covers the date — careers have gaps, and papering over one would invent a fact.
Between them, 35 of 56 co-author entries can say where that person was.
.github/workflows/refresh-collaborators.yml rebuilds from ORCID every
1 September and opens an issue when the file would change. An issue rather
than a pull request, unlike the contributions refresh: that one only adds
repositories, while this one can rewrite where a person is recorded as having
worked, which deserves a human reading the diff.
GitHub Pages via Actions (.github/workflows/deploy.yml), not a branch deploy:
Astro builds to dist/, which a branch deploy would mean committing — and
would publish the repo source instead of the built site.
public/CNAME ships the custom domain in the artifact so it survives every
deploy.
Every pull request also gets a full preview — the preview job in ci.yml
builds the site and the four PDFs and deploys them to Cloudflare Pages, then
posts the URL as a sticky comment. GitHub Pages serves one deployment per repository and that one
is production, which is why previews live elsewhere. The job skips rather than
fails when the Cloudflare secrets are absent, so a fork does not see a red
tick. Cloudflare refuses any file over 25 MiB, which is why the 27 MiB Typst
compiler wasm ships gzipped (scripts/gzip-compiler.mjs) and the CV builder
inflates it in the browser.
Jekyll never runs. Source "GitHub Actions" serves the uploaded artifact
verbatim, so Astro's underscore-prefixed _astro/ is untouched — the live site
serves it while no .nojekyll exists anywhere, which is the proof. A
.nojekyll would also be unshippable: upload-pages-artifact strips every
dotfile from the tarball.
That leaves one thing holding the guarantee up, and it lives in the web UI
where nothing in the repo can see it change. So deploy.yml asserts
build_type == workflow against the API before it builds, and fails with the
setting to correct rather than publishing a styleless site.
This repo previously held an academicpages Jekyll site — 211 files under
_config.yml, _posts/, _publications/ and the rest. All of it went in the
rebuild and survives only in the git history. files/starkman_cv.pdf kept its
original URL.