Skip to content

feat: Agente de kubernetes + alertas + acciones - #50

Closed
nicolas344 wants to merge 2 commits into
mainfrom
feature/agente-kubernetes-implementado
Closed

nicolas344 wants to merge 2 commits into
mainfrom
feature/agente-kubernetes-implementado

Conversation

@nicolas344

Copy link
Copy Markdown
Owner

No description provided.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Este PR agrega soporte de Kubernetes al flujo de incidentes de Sentinel (ingesta de alertas → enrutamiento a agente → propuesta/ejecución de acciones → verificación), integrando alertas basadas en kube-state-metrics, un KubernetesAgent con tools de observación, y la capacidad de ejecutar acciones vía kubectl.

Changes:

  • Se incorporan reglas de alertas Kubernetes en Prometheus y soporte de container_runtime: kubernetes a nivel backend/frontend.
  • Se añade un KubernetesAgent (prompt + tools) y runbooks seed para investigación asistida.
  • Se habilita ejecución/verificación de acciones Kubernetes y snapshots de métricas para pods.

Reviewed changes

Copilot reviewed 17 out of 18 changed files in this pull request and generated 13 comments.

Show a summary per file
File Description
prometheus/alert_rules.yml Nuevas alertas Kubernetes (kube-state-metrics).
Frontend/src/pages/Dashboard.jsx Badge de runtime para kubernetes.
docker-compose.yml Variables/volúmenes para kubeconfig y proxy en entorno local.
demo_kubernetes.sh Script de demo para disparar incidente Kubernetes end-to-end.
Backend/services/verification.py Verificación post-acción para recursos Kubernetes.
Backend/services/prometheus.py Snapshot de métricas para pods Kubernetes.
Backend/services/alert_processor.py Enrutamiento/runtime Kubernetes y snapshot de métricas.
Backend/services/alert_processor.py Dedupe por container_runtime y ajuste de logs/Loki.
Backend/services/agents/supervisor.py Propuesta de acciones kubectl para incidentes Kubernetes.
Backend/services/agents/kubernetes/tools.py Tools read-only del KubernetesAgent (SDK k8s).
Backend/services/agents/kubernetes/prompt.md Prompt y formato de respuesta del KubernetesAgent.
Backend/services/agents/kubernetes/agent.py Implementación del KubernetesAgent (loop ReAct).
Backend/services/agents/kubernetes/init.py Inicialización del módulo KubernetesAgent.
Backend/services/agents/guardrails.py Whitelist de comandos kubectl permitidos.
Backend/services/agents/init.py Auto-registro del KubernetesAgent.
Backend/scripts/seed_kubernetes_runbooks.py Seed de runbooks Kubernetes en ChromaDB.
Backend/routers/actions.py Validación/ejecución de comandos kubectl bajo aprobación.
Backend/requirements.txt Dependencia kubernetes (Python client).
Backend/db/migrations/004_kubernetes_runtime.sql Constraint DB para permitir container_runtime='kubernetes'.
Comments suppressed due to low confidence (1)

Backend/services/agents/supervisor.py:180

  • La acción propuesta para Kubernetes usa el mismo clean tanto para borrar pods como para reiniciar deployments. Para incidentes de pod (NotReady/CrashLoop) o de nodo (NodeNotReady), kubectl rollout restart deployment/<pod-o-nodo> probablemente fallará o actuará sobre un recurso equivocado. Conviene elegir el recurso desde labels (pod vs deployment vs node) y devolver None (sin acción automática) para alertas de nodo u otros casos no seguros.
        # Crash/OOM/restart → delete the pod so the ReplicaSet recreates it fresh
        if incident_type in {"app_crash", "oom", "restart_loop", "dependency_failure", "config_error"}:
            return f"kubectl delete pod {clean} -n {namespace}"
        return f"kubectl rollout restart deployment/{clean} -n {namespace}"


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread demo_kubernetes.sh Outdated
Comment on lines +6 to +9
BACKEND="http://localhost:8000"
ANON_KEY="eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImV1c3p1cGRlY2l0cXpqdXp0cWVvIiwicm9sZSI6ImFub24iLCJpYXQiOjE3NzE0NDM3OTcsImV4cCI6MjA4NzAxOTc5N30.mgpAaQviU-dMWohAPwggO3mOJrcWrUR7WlbwQAcLmIk"
SERVICE_KEY="eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImV1c3p1cGRlY2l0cXpqdXp0cWVvIiwicm9sZSI6InNlcnZpY2Vfcm9sZSIsImlhdCI6MTc3MTQ0Mzc5NywiZXhwIjoyMDg3MDE5Nzk3fQ.6cPnJBvaEctDVJ1NhUahfbf7xWYK69NlyoT43HkR3MI"
SUPABASE_URL="https://euszupdecitqzjuztqeo.supabase.co"
Comment thread demo_kubernetes.sh Outdated
Comment on lines +47 to +50
TOKEN=$(curl -s -X POST "$SUPABASE_URL/auth/v1/token?grant_type=password" \
-H "apikey: $ANON_KEY" \
-H "Content-Type: application/json" \
-d '{"email":"ricomontesinonicolas@gmail.com","password":"Motoraton_147"}' \
Comment thread docker-compose.yml Outdated
Comment on lines +26 to +31
# K8S_PROXY_URL: URL del kubectl proxy para entornos Docker Desktop (macOS/Windows).
# En Docker Desktop el API server (127.0.0.1:6443) no es alcanzable desde contenedores.
# Solución: antes del demo ejecutar en el host:
# kubectl proxy --port=8555 --address=<LAN_IP> --accept-hosts='.*' &
# Y ajustar esta variable con la IP LAN del host (ver: ifconfig | grep "inet 192")
K8S_PROXY_URL: "http://192.168.1.2:8555"
Comment thread docker-compose.yml Outdated
Comment on lines +42 to +43
# En macOS/Linux: ~/.kube/config montado como read-only.
- ${KUBECONFIG:-${HOME}/.kube}:/root/.kube:ro
Comment on lines +299 to +303
labels:
severity: critical
container_runtime: kubernetes
annotations:
summary: >-
Comment thread Backend/routers/actions.py Outdated

_CONTAINER_NAME_RE = re.compile(r"^[a-zA-Z0-9][a-zA-Z0-9_.-]{0,127}$")
_PG_DATNAME_RE = re.compile(r"^[a-zA-Z0-9][a-zA-Z0-9_-]{0,62}$")
_K8S_NAME_RE = re.compile(r"^[a-zA-Z0-9][a-zA-Z0-9-]{0,62}$")
Comment thread Backend/services/verification.py Outdated
dep = apps_v1.read_namespaced_deployment(name=resource_name, namespace=namespace)
desired = dep.spec.replicas or 0
available = dep.status.available_replicas or 0
healthy = available >= desired and desired > 0
Comment on lines +252 to +253
result = []
for e in sorted(events.items, key=lambda x: x.last_timestamp or x.event_time or "", reverse=True)[:20]:
Comment on lines +109 to +122
for i in range(_MAX_TOOL_ITERATIONS + 1):
response: AIMessage = llm.invoke(messages)
messages.append(response)

pending = getattr(response, "tool_calls", None) or []
if not pending:
return (response.content or "").strip(), recorded

if i == _MAX_TOOL_ITERATIONS:
messages.append(HumanMessage(
content="Límite de tool calls alcanzado. Redacta el análisis "
"final ahora con la información que tienes."
))
continue
Comment on lines 212 to 219
_runtime_hint = labels.get("container_runtime", "")
if source_type != "container":
container_runtime = None
elif _runtime_hint in {"kubernetes", "k8s", "containerd"}:
container_runtime = "kubernetes"
elif _runtime_hint in {"podman"}:
container_runtime = "podman"
else:
- Remove demo scripts from git tracking (contain hardcoded credentials)
- Add demo_*.sh to .gitignore; add *.example templates for teammates
- docker-compose: K8S_PROXY_URL empty by default, fix kubeconfig volume
- alert_processor: build k8s target from pod/deployment/node labels
- supervisor: parse resource type explicitly, return None for nodes
- agent._react_loop: final LLM call without tools after max iterations
- tools.get_pod_events: use epoch datetime fallback to avoid TypeError
- verification: treat desired==0 (scale-to-zero) as healthy
- actions: fix _K8S_NAME_RE to DNS-1123 lowercase
- ruff: fix E402/F841/E501 in kubernetes files
@nicolas344 nicolas344 closed this May 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants