Senior AI Engineer — GenAI, LLM and agentic systems PhD candidate — AI safety & alignment · Dubai / Abu Dhabi, UAE
I build LLM systems and then measure whether they actually work. Most of what follows is the second half of that sentence.
Currently researching whether safety guardrails survive translation across language and modality — and whether the benchmarks we use to check are measuring refusal or just reading comprehension.
A set of repositories, each built to answer one measured question rather than to demonstrate a technology. Every number in them is generated by a committed script, and CI fails if a result drifts from the code that produced it.
| The question | |
|---|---|
| attention-kv-cache-from-scratch | What did grouped-query attention actually buy, in concurrent sequences? |
| llm-serving-benchmark | Which serving config meets a stated SLO at the lowest cost per million tokens? |
| prefill-decode-roofline | Where does measured bandwidth fall short of theoretical peak, and why? |
| quantization-accuracy-curves | How much does the calibration set choice change measured degradation? |
| speculative-decoding-when-it-pays | At what batch size does speculative decoding stop paying for itself? |
| moe-vs-dense-serving-profile | Why do active parameters mislead capacity planning for a sparse model? |
| context-extension-and-long-context-eval | How far short of advertised context does effective context fall? |
| The question | |
|---|---|
| llm-eval-harness | How far can an LLM judge be trusted, measured against human labels? |
| guardrail-transfer-study | How much of the apparent cross-modal safety gap is a reading-ability artifact rather than an alignment gap? |
| rag-eval-retrieval-vs-generation | Would a single end-to-end RAG score have hidden a real regression? |
| agent-eval-trajectory-vs-outcome | How often does an agent reach the right answer through a wrong process? |
| The question | |
|---|---|
| preference-optimization-landscape | Which preference-optimisation method for which situation, and what does each give up? |
| rl-environment-for-rubric-graded-tasks | What reward hack did the agent find, and what environment change closed it? |
| sft-loss-masking-and-packing | How much does packing without attention-mask correction actually cost? |
| peft-lora-memory-and-quality | At what rank does quality plateau, and does the memory math match reality? |
| synthetic-data-pipeline | What fraction of a generated set is contaminated against the eval set? |
| The question | |
|---|---|
| mcp-server-and-agent | What is each agent topology's failure rate, and does supervisor actually beat single-agent? |
| retrieval-internals-and-tuning | At what latency budget does re-ranking stop being worth it? |
| doc-extraction-ocr-vs-vlm | Which extraction approach per document class — and does constrained decoding beat a retry loop? |
| llm-client-kit | Where does the wall-clock actually go when you fan out LLM calls? |
| inference-gateway | At what monthly request volume does self-hosting beat the API? |
| voice-loop-latency-budget | Which hop dominates perceived voice latency? |
Bold = flagship. Several are early — each README states plainly what is built and what is not, because a status table beats a guess.
- Now — Senior AI engineer on healthcare AI: multimodal document pipelines, agentic orchestration, human-in-the-loop safety layers for automated clinical-financial decisioning, self-hosted inference on vLLM.
- PhD (in progress) — Attenuation of Safety Alignment Across Language and Modality in Multimodal LLMs. Two preprints in preparation: MCP multi-agent orchestration, and cross-lingual/cross-modal guardrail transfer.
- Before — travel and booking platforms at scale (React/TypeScript, FastAPI, C++ microservices); autonomous-vehicle perception (YOLO, ROS2, sensor fusion, SLAM, TensorRT edge inference).
- MSc Computer Science (distinction) · BTech Robotics & Automation
Core Python · TypeScript · C++ · SQL LLM LangGraph · MCP · vLLM · Hugging Face (Transformers, TRL, PEFT) · PyTorch Data PostgreSQL/pgvector · Redis · FAISS Infra Docker · Kubernetes · FastAPI · OpenTelemetry · GitHub Actions
- moclamzw@gmail.com
- Portfolio site — coming soon
This account dates from July 2026, after my previous GitHub account was compromised. Work here begins from that rebuild.