I build AI systems that hold up under real production pressure β at the intersection of distributed systems and agentic AI. 3+ years across Fortune 500 energy infrastructure (Shell PLC) and research-scale AI (NYU).
- π¬ Currently building GeneCart, an AI-assisted genomics discovery platform at NYU's Center for Genomics & Systems Biology β owning the agent/LLM layer with LangGraph + PyTorch.
- β‘ Cut P99 RAG latency by 78% (450ms β <100ms) on a Multi-Agent research engine serving 3,000+ RPS at 99.9% uptime.
- πͺ¨ Pushed LLM inference to 15ms on Snapdragon NPUs (10Γ faster than cloud) via QLoRA + AWQ quantization β won the Qualcomm Edge AI Hackathon.
- π’οΈ Kept Shell's maritime telemetry alive at 115GB/day across 200+ offshore stations, zero data loss.
- π M.S. Computer Science, NYU (GPA 3.8) Β· TA for Algorithms & ML for Bioinformatics.
- π‘ I like chasing unconventional ideas, and I care about engineering that serves a real human need.
Curious how I built this profile? My portfolio runs Avocado AI, a streaming RAG agent you can actually talk to β jayaremala.com
Languages
AI, ML & Agents
Systems & Cloud
Frameworks & Databases
π₯ jayaremala.com β Avocado AI Β· Live
Production portfolio fronted by a streaming RAG agentic chatbot with a 4-stage hybrid retrieval pipeline: query expansion β batched dense search (ChromaDB / all-MiniLM-L6-v2 ONNX) β BM25 lexical search β Reciprocal Rank Fusion (k=60) β all before Gemini 2.5 Flash sees the question. Multi-provider fallback (Gemini / Groq / OpenRouter), incrementally auto-syncing knowledge base, blue-green deploys on AWS Lightsail.
Next.js 16 FastAPI ChromaDB BM25 RRF fastembed ONNX Docker GitHub Actions
π SnapLog β Edge AI Security Engine Β· Qualcomm Hackathon Winner
15ms token latency on Snapdragon NPUs β 10Γ over cloud inference β by fine-tuning Llama 3.2 3B on security logs with QLoRA and deploying via 4-bit AWQ quantization through ONNX Runtime on-device. Offline-first SQLite buffer guarantees zero data loss during network partitions.
QLoRA AWQ ONNX Runtime Llama 3.2 Snapdragon NPU FastAPI
AI-powered platform making genomic data exploration actionable for researchers. Owning the agent/LLM layer with LangGraph + PyTorch and the full-stack delivery path on AWS + Kubernetes.
LangGraph PyTorch FastAPI React PostgreSQL AWS
Production LangGraph + Llama 3.1 70B system semantically mapping global researcher collaboration networks across millions of Elsevier papers. 78% P99 latency cut via write-through Redis caching; 99.9% uptime at 3,000+ RPS on AWS ECS.
LangGraph Llama 3.1 70B BGE-M3 Redis AWS Kubernetes
Conflict-free multi-user editing with Yjs (CRDTs) + WebSockets, horizontally scaled behind Nginx. 65% better AI auto-complete context via an AST-chunking, Voyage-Code-2 embedding agent.
CRDTs WebSockets Claude 3.5 Voyage AI Redis Docker
π gradeVITian Β· GitHub
Academic grade-forecasting PWA β 17K+ monthly active users, 20K+ accounts, #2 Google Search ranking. Ran 6+ years in production, recently rebuilt on Next.js + FastAPI.
Next.js React FastAPI PWA SEO
| Where | Role | Impact |
|---|---|---|
| NYU CAS β Genomics & Systems Biology | Software Engineer (Jun 2025 β Present) | Building GeneCart's agent/LLM layer (LangGraph + PyTorch) |
| NYU IT β High-Speed Research Network | Software Engineer, Research Infra | 78% P99 latency cut Β· 99.9% uptime @ 3K+ RPS Β· 65% faster deploys |
| Wipro (Client: Shell PLC) | Software Engineer | Zero-data-loss Kafka pipeline Β· 115GB/day Β· 200+ offshore stations |
| VIT University | Full Stack Developer | gradeVITian β 17K+ MAU, #2 on Google |
Open to roles in Healthcare, Finance, and Consumer (Entertainment & Retail) where AI infrastructure has to actually work at scale.
