Independent LLM-evaluation researcher. Sydney, Australia.
My work audits large language models the way clinical psychology audits human subjects, through validated measurement instruments, population ground truth, effect-size estimation, and fully released receipts rather than binary pass/fail benchmarks. The recurring finding across these studies is consistent: aggregate metrics conceal fine-grained distortion. A bias will show its direction plainly while its distribution is lost entirely. Most studies run solo and unfunded, without institutional compute or supervision, and this is deliberate. Independent evaluation grounded in its own data is the argument, and every repository here is built to let you check the numbers yourself.
Selected work
- Plausible Patients, Impossible Populations: language models simulate psychiatric patients who are individually plausible but form populations that diverge from epidemiological ground truth.
- The Granularity Gap: an LLM judge's own pass/fail verdict explains under a third of the variance in the severity scores that same judge assigns.
- Private Documents, Local Models (with R. H. Tai): given the same retrieved passages as hosted frontier models, small local models on one consumer GPU trail them by a small but measurable amount, and local judges agree with the hosted judge on which side leads.
Papers on arXiv
- arXiv:2606.05183 The Granularity Gap
- arXiv:2604.17359 Plausible Patients, Impossible Populations
Research repositories
| Repository | What it holds |
|---|---|
| The-Granularity-Gap | Data, code and verification gates for a continuous-scale audit of sycophancy across 8 Gemini variants and 8,830 responses, with 10,792 per-vote judge logs. |
| plausible-patients | Four cost-tier models generate 28,800 simulated psychiatric patients across 120 demographic cohorts, scored against survey-weighted NHANES population norms. |
| psych_scope | A counterfactual audit of socioeconomic-cue bias in LLM mental-health assessment across three model families, with a sparse-autoencoder test of whether interpretability can explain it. |
| deeper-than-the-guardrails | A ten-pair base/abliterated audit of demographic bias in clinical screening, 95,920 observations graded against author-derived NHANES 2005-2018 norms. |
| telling-more-than-they-can-know | A four-model factorial audit of whether model self-explanations reveal or conceal the demographic drivers of their psychiatric-instrument scoring, across 18 intersectional cohorts. |
| the-listening-gap | Six LLMs translate semantic audio descriptors into parametric EQ curves, graded against SAFE-DB settings from real audio engineers. No LLM judge. |
| leave-the-image-alone | 30 CORD-v2 receipts across 4 preprocessing conditions, 4 vision models and 3 repeats, 1,440 transcriptions. Raw input beats binarize, upscale and downscale. |
Systems
Mimesis_Voice_Clone is an offline MCP server for drafting in a target author's voice, using SQLite FTS5 hybrid search, stylometric profiling and local ONNX embeddings. claude_mind is a self-maintaining personal memory system: vector and graph indexes over an Obsidian vault, kept current by autonomous agent loops.
The full index of studies, findings, and receipts is in the research collection.
I also write fiction and philosophical essays, collected at my portfolio.
Open to research collaboration and funded work in LLM evaluation and safety. · pskeough@gmail.com
