I build the layer that makes language models trustworthy: orchestration that verifies itself, retrieval you can measure, and guardrails that hold. Every project ships with a real eval harness, because "it felt better" is not a metric.
|
A meta-agent plans a task DAG, routes steps to web and typed-API agents, runs a plan-verify-iterate loop, and synthesizes new sub-agents at runtime.
|
BM25 vs dense vs hybrid (RRF) vs hybrid + rerank over a hand-labelled corpus. Local models, no API key required.
|
|
Local-first LLM guardrails: layered PII, injection, secrets, and schema validators with a built-in eval harness.
|
Recommends the cheapest refinement level (prompt vs RAG vs agent vs fine-tune) that clears a quality bar, via cost-performance Pareto benchmarking.
|
|
Open-source, local "System One" decision layer for LLM apps: typed decisions with provably calibrated confidence. No API key, no waitlist.
|
Image to validated, personalized food-product intelligence: multi-agent extraction, a misleading front-of-pack claim detector, and a data flywheel that lifts attribute F1 across rounds.
|
Based in Mumbai, India · open to conversations about hard LLM reliability problems.



