Senior AI Engineer with 10 years of software engineering experience building backend and distributed systems with Python, TypeScript, Go, and AWS.
I focus on production agentic AI systems: agents that can retrieve enterprise knowledge, query governed data, call real tools, execute workflows safely, and be evaluated and observed in production.
My current focus is the intersection of:
- 🤖 Agentic AI — Amazon Bedrock · AgentCore · Strands · tool calling · MCP
- 🧠 Enterprise AI — RAG · semantic layers · NL-to-SQL · governed data access
- 🧪 LLMOps — golden datasets · agent evals · regression testing · CI quality gates
- 🔐 Safe execution — authorization · human-in-the-loop · policy-controlled tools
- ☁️ Production AWS — CDK · CloudWatch · OpenTelemetry · Step Functions · Lambda
- 🏗️ Backend & distributed systems — APIs · event-driven systems · reliability · scaling
Currently shipping production services at Amazon Web Services / Amazon Connect and working on production-oriented agentic AI systems.
📍 Based in Mexico and open to remote contractor opportunities with US and international teams.
Rather than building generic AI infrastructure from scratch, I focus on applying managed AI and cloud capabilities to concrete production problems.
An agent that correlates CloudWatch logs, metrics, deployment history, and engineering runbooks to assist with incident investigation and produce evidence-backed failure hypotheses.
Focus areas
- tool selection
- operational reasoning
- evidence grounding
- agent observability
- incident evaluation
- safe access to production telemetry
Stack: Amazon Bedrock · Strands · CloudWatch · OpenTelemetry
An enterprise analytics agent that translates natural-language questions into safe queries through a semantic layer and policy-controlled data tools.
The interesting problem is not generating SQL — it is making sure the agent understands business semantics, requests clarification when necessary, and never accesses data the user is not authorized to see.
Focus areas
- semantic layers
- NL-to-SQL
- tool authorization
- row/data access boundaries
- adversarial evaluation
- regression testing
Stack: Amazon Bedrock · Strands · AgentCore · Python · SQL · AWS
A production-oriented knowledge agent for internal engineering documentation with hybrid retrieval, metadata-aware access control, reranking, citations, and retrieval evaluation.
Focus areas
- RAG
- hybrid retrieval
- metadata filtering
- reranking
- citation grounding
- Recall@K / MRR
- golden datasets
Stack: Amazon Bedrock · Bedrock Knowledge Bases · Python · RAG
Agent workflows where reasoning is probabilistic but side effects remain deterministic and controlled.
Sensitive actions require explicit approval before invoking operational tools or APIs.
Focus areas
- human-in-the-loop
- approval gates
- idempotency
- retries
- tool policies
- safe side effects
- workflow evaluation
Stack: Strands · Amazon Bedrock · AWS Step Functions · Lambda · AgentCore
For me, an agent is not production-ready because it works in a demo.
Production AI needs a repeatable quality loop:
Build
↓
Trace
↓
Evaluate
↓
Analyze failures
↓
Improve
↓
Regression test
↓
Deploy
↓
Observe
Areas I'm particularly interested in:
- versioned golden datasets
- agent and retrieval evaluation
- LLM-as-judge with human alignment
- tool-selection evaluation
- adversarial testing
- prompt/model regression testing
- CI quality gates
- OpenTelemetry tracing
- CloudWatch observability
- latency and cost monitoring
I enjoy the infrastructure side of AI systems, especially the boundary between AI Engineering, LLMOps, and platform engineering.
At AWS, I've worked with deployment safety and reliability practices including AWS CDK, pre-production validation, GameDay chaos testing, dependency failure scenarios, on-call alerting, and runbooks.
I like treating AI systems the same way as any serious distributed system:
systems fail, dependencies timeout, models change, tools misbehave, and quality regresses — production engineering is about detecting and containing those failures.
Agents · RAG · Tool Calling · MCP · NL-to-SQL · Semantic Layers · Golden Datasets · LLM-as-Judge · Agent Evaluations · CI Quality Gates
Python · TypeScript · Go · Ruby · FastAPI · Rails · Node.js · REST APIs · SQL
Lambda · Step Functions · EventBridge · SQS · SNS · IAM · KMS · CloudWatch · S3 · RDS · DynamoDB · EKS · AWS CDK
CloudWatch · OpenTelemetry · SLOs · CI/CD · Infrastructure as Code · AWS CDK · GameDays · Chaos Testing · Load Testing · TDD
I'm particularly interested in problems around:
- production AI agents
- agent reliability and evaluation
- governed enterprise tool access
- semantic data layers
- retrieval engineering
- AI observability
- LLMOps
- Infrastructure as Code for AI workloads
- distributed systems
- safe human-in-the-loop automation
I'm open to remote Senior AI Engineer / AI Platform Engineer / GenAI Engineer contractor roles, particularly around:
- Amazon Bedrock / AgentCore / Strands
- enterprise AI agents
- RAG and knowledge systems
- NL-to-SQL and governed enterprise data access
- AI evaluation and LLMOps
- agent observability and production reliability
- AWS architecture and Infrastructure as Code
If your team is moving an AI system from prototype to production, that's the kind of problem I enjoy working on.
⭐️ Building AI agents that can safely interact with real enterprise systems.



