I'm a senior software engineer based in Berlin with more than ten years of experience across SaaS, ad tech and gaming.
I focus on AI support reliability particularly how support systems retrieve sources, assemble evidence, generate answers and expose enough internal state to diagnose failures.
My work combines assessments of real support assistants with retrieval evaluation, evidence review workflows and reliability-focused engineering, especially around billing, payments, subscriptions and account-access workflows.
- Evaluating high-risk AI support workflows against documented source material
- Measuring retrieval and ranking quality at aggregate and question level
- Inspecting context construction, evidence coverage, groundedness and correctness
- Building evaluation datasets, review queues, dashboards and observability tooling
- Turning repeated failures into inspectable findings and regression cases
A retrieval evaluation workbench for testing whether an AI support system finds the correct source material before answer generation.
It compares keyword, BM25, embedding and hybrid retrieval strategies using Hit@k, Recall@k and MRR, with grouped diagnostics and question-level inspection.
Demonstrates: retrieval evaluation, source coverage analysis, metric interpretation and failure investigation.
A human-review workflow for support messages involving payment, subscription, access and revenue risk.
It turns structured AI output into reviewable findings rather than treating model output as an automatic decision.
Demonstrates: structured outputs, human review, billing-risk workflows and inspectable AI-assisted decisions.
A developer-facing tool that analyzes screenshots and DOM context to produce structured frontend QA findings.
The system separates observable evidence from inferred conclusions and keeps findings reviewable before action.
Demonstrates: multimodal AI workflows, structured evaluation and evidence-aware product interfaces.
My implementation work spans React, TypeScript, Node.js, APIs, backend integrations, internal tools, billing systems and complex product workflows.
I am particularly interested in systems where AI output needs to be evaluated, reviewed and made easier to inspect.



