Software Engineer (AI) at Red Hat Β· Researcher (LLM evaluation & CoT process monitoring) Β· 2x EMNLP Main Β· Open Source
I am a Software Engineer (AI) at Red Hat. I work on agentic systems in production: MCP servers, diagnostics agents, and support automation that customers actually use.
I work on LLM evaluation and chain-of-thought monitoring. I care about when a score or a trace stops being a trustworthy signal. I look at whether two models can share the same accuracy and still fail in different ways, and whether a chain that still looks clean is already getting riskier with depth.
I belong in the overlap of shipping agents and measuring them. I maintain Provena. I contribute to PyTorch and vLLM. I mentor students who are starting the same kind of work.
Patents
-
System and Method for Utilizing Digital Footprints of a User to Generate Conversational AI Thereof
202421048008Β· Filed Jun 2024 Β· Granted (20-year term) -
Recommendation and Intent Reconciliation in a Virtual Leader Framework
202521025963Β· Filed Apr 2025 Β· Under review
Publications
-
When Does Reasoning Age? Survival Analysis of Step-Level Error Hazard in LLM Chains Β· EMNLP 2026 Main Β· first author Β· CoT process monitoring
-
The Correlation Mirage: Benchmark Dependence Collapses for Top-Performing LLMs Β· EMNLP 2026 Main Β· first author Β· science of evaluations
-
GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon Β· ICCUBEA 2026 (IEEE Xplore) Β· energy as an eval axis
-
A Survey on Advanced Recommendation Systems: Content-Based Filtering, Collaborative Filtering, Hybrid and Opinion Mining Approaches Β· ICTIS 2025 / Springer LNNS Β· 2025
-
Proposed Model of Hindi Book Review Sentiment Analysis Β· IJERT 2023 Β· earlier NLP
PyTorch β Deep Learning Framework
Active contributor working on core framework improvements:
- Input validation & safety β Adding proper bounds checking to prevent silent failures in
max_pool3d,channel_shuffle,RNN cells, and convolution ops - Optimizer improvements β
maximizeparameter for LBFGS, integer step tensor support in foreach optimizers - Numerical stability β Fixing NaN propagation in
lp_pool,hardtanhbackward pass corrections - API enhancements β
keepdimforcosine_similarity,NanDetectModefor forward-pass diagnostics,dtypecontext manager
vLLM β LLM Inference Engine
Contributing to the high-throughput LLM serving engine:
- Responses API β Namespace tools support for harmony/GPT-OSS models
| Project | Description | Stack |
|---|---|---|
| survival-llm-reasoning | Code for When Does Reasoning Age? β CoT error hazard, 18,969 chains | Python, survival analysis |
| correlation-mirage-benchmarks | Code for The Correlation Mirage β copula tail dependence of LLM benches | Python, copulas |
| provena | Context governance for agentic AI β tamper-evident audit trails, provenance validation, EU AI Act compliance | Python, PostgreSQL, MCP, Policy Engine |
| sumo-logic-mcp | MCP server for Sumo Logic with 48 tools β log search, monitors, alerts, dashboards, metrics | Python, MCP Protocol |
| repo-time-machine | Agentic RAG for codebases β ask questions answered by code, git history, issues & PRs | Python, FAISS, Ollama |
Languages
ML & AI
Infrastructure & DevOps
Databases & Tools
Currently measuring when LLM evals and CoT traces stop being trustworthy β and shipping agents at Red Hat.


