Building toward applied AI/ML, with a focus on evaluation, data quality, and reliable AI tooling.
I like projects where the hard part is proving what a system actually does: checking leakage, choosing honest baselines, testing failure paths, and making limitations visible next to the result.
An ML pipeline for forecasting 2030 EV charging demand across Prague grid zones, classifying charger types, and replaying flexible charging against transformer capacity.
- 24.03 kWh MAE, 38.8% below a population baseline and 17.0% below ridge on the same 55 features.
- Found and removed validation leakage from early stopping.
- Reverse-engineered 99.75% of the synthetic charger label rule and narrowed the product claim accordingly.
- Built an auditable Streamlit dashboard over committed outputs rather than mock data.
Team project at the Czech AI Olympiad 2026; my work covered the two modelling tracks and the post-competition evaluation rewrite.
Five independent plugins for multi-model review, structured ideation, API and repository discovery, supply-chain inspection, and agent guardrails.
- A three-stage model council with anonymized peer review and 52 deterministic failure-path tests.
- A local catalogue of 1,679 public APIs plus repository vetting.
- Explicit privacy and security boundaries: failed checks stay unknown, never silently “clean.”
- Third-party code and prior art are attributed down to the component and approximate share.
- Earcon — a zero-runtime-dependency Claude Code plugin that plays only when input is needed or a turn finishes; 49 behavioral tests across macOS and Linux.
- BBWA — a WhatsApp client for the Android 4.3 runtime on BlackBerry 10, with a hardened Node backend and a reproducible APK release.
ML evaluation and data provenance · classical ML baselines · agent reliability and evals · Python, Polars, scikit-learn, LightGBM, Node.js, shell, and GitHub Actions.
The code, case studies, and limitations are public here. That is the portfolio.

