Skip to content
View vnponce's full-sized avatar

Block or report vnponce

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
vnponce/README.md

Hi there 👋, I'm Abel Ponce

Senior AI Engineer · Building AI agents that safely interact with real enterprise systems

LinkedIn Email Location Status


🚀 About Me

Senior AI Engineer with 10 years of software engineering experience building backend and distributed systems with Python, TypeScript, Go, and AWS.

I focus on production agentic AI systems: agents that can retrieve enterprise knowledge, query governed data, call real tools, execute workflows safely, and be evaluated and observed in production.

My current focus is the intersection of:

  • 🤖 Agentic AI — Amazon Bedrock · AgentCore · Strands · tool calling · MCP
  • 🧠 Enterprise AI — RAG · semantic layers · NL-to-SQL · governed data access
  • 🧪 LLMOps — golden datasets · agent evals · regression testing · CI quality gates
  • 🔐 Safe execution — authorization · human-in-the-loop · policy-controlled tools
  • ☁️ Production AWS — CDK · CloudWatch · OpenTelemetry · Step Functions · Lambda
  • 🏗️ Backend & distributed systems — APIs · event-driven systems · reliability · scaling

Currently shipping production services at Amazon Web Services / Amazon Connect and working on production-oriented agentic AI systems.

📍 Based in Mexico and open to remote contractor opportunities with US and international teams.


🤖 What I'm Building

Rather than building generic AI infrastructure from scratch, I focus on applying managed AI and cloud capabilities to concrete production problems.

🔎 CloudOps Investigation Agent

An agent that correlates CloudWatch logs, metrics, deployment history, and engineering runbooks to assist with incident investigation and produce evidence-backed failure hypotheses.

Focus areas

  • tool selection
  • operational reasoning
  • evidence grounding
  • agent observability
  • incident evaluation
  • safe access to production telemetry

Stack: Amazon Bedrock · Strands · CloudWatch · OpenTelemetry


📊 Governed Natural-Language Analytics Agent

An enterprise analytics agent that translates natural-language questions into safe queries through a semantic layer and policy-controlled data tools.

The interesting problem is not generating SQL — it is making sure the agent understands business semantics, requests clarification when necessary, and never accesses data the user is not authorized to see.

Focus areas

  • semantic layers
  • NL-to-SQL
  • tool authorization
  • row/data access boundaries
  • adversarial evaluation
  • regression testing

Stack: Amazon Bedrock · Strands · AgentCore · Python · SQL · AWS


📚 Enterprise Knowledge Agent

A production-oriented knowledge agent for internal engineering documentation with hybrid retrieval, metadata-aware access control, reranking, citations, and retrieval evaluation.

Focus areas

  • RAG
  • hybrid retrieval
  • metadata filtering
  • reranking
  • citation grounding
  • Recall@K / MRR
  • golden datasets

Stack: Amazon Bedrock · Bedrock Knowledge Bases · Python · RAG


🛡️ Approval-Gated Agent Workflows

Agent workflows where reasoning is probabilistic but side effects remain deterministic and controlled.

Sensitive actions require explicit approval before invoking operational tools or APIs.

Focus areas

  • human-in-the-loop
  • approval gates
  • idempotency
  • retries
  • tool policies
  • safe side effects
  • workflow evaluation

Stack: Strands · Amazon Bedrock · AWS Step Functions · Lambda · AgentCore


🧪 LLMOps & Production AI

For me, an agent is not production-ready because it works in a demo.

Production AI needs a repeatable quality loop:

Build
  ↓
Trace
  ↓
Evaluate
  ↓
Analyze failures
  ↓
Improve
  ↓
Regression test
  ↓
Deploy
  ↓
Observe

Areas I'm particularly interested in:

  • versioned golden datasets
  • agent and retrieval evaluation
  • LLM-as-judge with human alignment
  • tool-selection evaluation
  • adversarial testing
  • prompt/model regression testing
  • CI quality gates
  • OpenTelemetry tracing
  • CloudWatch observability
  • latency and cost monitoring

☁️ Infrastructure as Code & Reliability

I enjoy the infrastructure side of AI systems, especially the boundary between AI Engineering, LLMOps, and platform engineering.

At AWS, I've worked with deployment safety and reliability practices including AWS CDK, pre-production validation, GameDay chaos testing, dependency failure scenarios, on-call alerting, and runbooks.

I like treating AI systems the same way as any serious distributed system:

systems fail, dependencies timeout, models change, tools misbehave, and quality regresses — production engineering is about detecting and containing those failures.


🛠️ Tech Stack

Agentic AI & LLMOps

AWS Bedrock AgentCore Strands MCP OpenAI Anthropic

Agents · RAG · Tool Calling · MCP · NL-to-SQL · Semantic Layers · Golden Datasets · LLM-as-Judge · Agent Evaluations · CI Quality Gates

Languages & Backend

Python TypeScript Go Ruby FastAPI Rails

Python · TypeScript · Go · Ruby · FastAPI · Rails · Node.js · REST APIs · SQL

AWS & Cloud

AWS Docker Kubernetes PostgreSQL Redis Kafka

Lambda · Step Functions · EventBridge · SQS · SNS · IAM · KMS · CloudWatch · S3 · RDS · DynamoDB · EKS · AWS CDK

Reliability & Observability

OpenTelemetry pytest Cypress

CloudWatch · OpenTelemetry · SLOs · CI/CD · Infrastructure as Code · AWS CDK · GameDays · Chaos Testing · Load Testing · TDD


🧭 Engineering Interests

I'm particularly interested in problems around:

  • production AI agents
  • agent reliability and evaluation
  • governed enterprise tool access
  • semantic data layers
  • retrieval engineering
  • AI observability
  • LLMOps
  • Infrastructure as Code for AI workloads
  • distributed systems
  • safe human-in-the-loop automation

🤝 Open to Contractor Opportunities

I'm open to remote Senior AI Engineer / AI Platform Engineer / GenAI Engineer contractor roles, particularly around:

  • Amazon Bedrock / AgentCore / Strands
  • enterprise AI agents
  • RAG and knowledge systems
  • NL-to-SQL and governed enterprise data access
  • AI evaluation and LLMOps
  • agent observability and production reliability
  • AWS architecture and Infrastructure as Code

If your team is moving an AI system from prototype to production, that's the kind of problem I enjoy working on.

📫 vnponce8@gmail.com


📊 GitHub Stats

GitHub Stats

GitHub Streak

Top Languages


📫 Connect with me

LinkedIn Email GitHub


⭐️ Building AI agents that can safely interact with real enterprise systems.

Pinned Loading

  1. healthcare-intake-coordination-system healthcare-intake-coordination-system Public

    A small multi-agent AI system for healthcare intake coordination, built to learn multi-agent orchestration patterns end-to-end without an SDK

    Python 1

  2. football-intelligence-agent football-intelligence-agent Public

    AI-powered serverless agent that answers football questions using AWS Lambda and Claude AI

    JavaScript

  3. mundial-ticket-agent mundial-ticket-agent Public

    small learning project where I explored integrating AWS Bedrock with Lambda to create a virtual assistant for World Cup ticket registrations.

    TypeScript

  4. SeederMX SeederMX Public

    This project shows Laravel model factories to generate 'Estados' 'Municipios' from México.

    PHP 8 9

  5. deliveries-exercise deliveries-exercise Public

    Deliveries exercise with react using TDD

    JavaScript

  6. apartadoenlinea apartadoenlinea Public

    Main idea is to have a sort of ecom but schedulling the orders to specific day, time and store with its owb dashboard to track orders.

    PHP