Comparing sequential forecasters via confidence sequences & e-processes
-
Updated
Oct 24, 2023 - Jupyter Notebook
Comparing sequential forecasters via confidence sequences & e-processes
Taming False Positives in Out-of-Distribution Detection with Human Feedback (AISTATS '24)
让 prompt 从"感觉好了"变成"真好了" · Where "feels better" becomes "measurably better" for your AI prompts.
Compute confidence Intervals (CLT, Chebyshev, Hoeffding) and Confidence Sequences (in a sequential setting if needed) in Julia
A GxP-purpose-built LLM evaluation framework: anytime-valid validation, judge qualification, ALCOA+ audit, human plus AI workflow validation, drift monitoring, listener hooks, and a high-level facade on Inspect AI. Enables responsible AI use under risk-based assurance.
Always-valid sequential A/B testing engine: mSPRT confidence sequences make peeking safe by construction, CUPED cuts variance up to 49%, SRM gates bad data. Built-in adversarial peeking harness proves the claim: naive daily peeking hit 27.5% false positives in 2,000 simulations; this engine held 1.7%. All numbers reproducible from committed seeds.
Certified Sparse Attention: runtime-verified fidelity for sparse-attention LLM serving. Label-free detection, anytime-valid confidence sequences, elastic probe scheduling. Smoke-scale H1-H4 results on Qwen2.5-0.5B/1.5B (AUC 0.92, Spearman -0.90).
Add a description, image, and links to the confidence-sequences topic page so that developers can more easily learn about it.
To associate your repository with the confidence-sequences topic, visit your repo's landing page and select "manage topics."