ML / Data Engineer — building RAG systems, data pipelines, and poking at what's inside LLMs
MS CS @ UT Dallas (May 2026) · Dallas, TX
- 🔍 Digging into mechanistic interpretability — trained an SAE on GPT-2-small and built a live safety monitor from it
- 🛠️ Shipped a RAG evaluation tool at an OpenAI hackathon (Build Week)
- 💼 Past experience across ML (IoT/CV at MIT-WPU × Capgemini), data engineering (DRDO — LiDAR/camera perception pipelines), and TA'ing DSA at UT Dallas
- 🎓 MS in Computer Science from UT Dallas, open to full-time ML / Data Engineering roles
|
Downshift · live demo · PyPI Cuts LLM costs per PR. Finds every LLM call in a Python repo, tests cheaper models against evals written for each call site, recommends the safe downgrades, and posts the projected monthly cost change on every pull request. On the demo app, 3 of 8 call sites downgraded for an 18% projected saving with pass rate held steady.
|
|
|
ActivationLens Mechanistic interpretability on GPT-2-small — trained a sparse autoencoder on the layer-6 residual stream, cut dead-feature collapse from ~80% to 0.77%, then built a live per-token safety monitor (0.758 AUROC) with a one-pass kernel that trimmed monitoring overhead 6x.
|
OpenAI Build Week GPT-5.6-powered tool that scores, diagnoses, and auto-tunes RAG pipeline outputs — built with Codex during OpenAI's Build Week hackathon.
|
|
DocumentSync AI RAG pipeline over 30 ArXiv papers across 4 chunking strategies, evaluated with RAGAS across 120 LLM-as-judge calls — best config improved context precision by 59% and recall 3x over baseline.
|
Real-Time Crypto Streaming Pipeline Streams live market data via Kafka + Spark Structured Streaming into a Redshift star schema, orchestrated with Airflow — cut Athena bytes scanned by 89%.
|
