This guide covers deployment options for Cyber-AutoAgent in various environments.
The React terminal uses one of three execution profiles. The setup wizard uses friendly display names, while the
configuration and --deployment-mode option use the canonical values shown below.
React UI
├── Python / Local CLI (local-cli)
│ └── Local Python agent process
├── Single Container (single-container)
│ └── Docker agent container
└── Full Stack (full-stack)
└── Docker Compose: agent + supporting services
The React UI starts the Python agent directly on the host through the local Python execution service. The host Python environment supplies the agent runtime, installed dependencies, and host-local security tools. Provider credentials, model configuration, and output paths are passed to that process, and no Docker services are started.
The Python process emits structured execution events that the React UI consumes for progress, tool activity, metrics, reports, and terminal completion. This mode has the smallest footprint and is suited to development or environments where the required Python tools are already installed locally.
The React UI uses the Docker execution service to start one isolated agent container. The container supplies the Python runtime and bundled assessment tools, while the UI passes the selected configuration, provider credentials, target information, and output mounts into the container.
The container runs the assessment and exits when the operation completes. This mode does not start the persistent observability and evaluation service stack from Docker Compose. If observability is configured for this mode, it must use an externally available service.
The React UI starts the Docker Compose deployment and communicates with the agent service over the Compose network. Alongside the agent, the stack provides the supporting services required for the complete platform experience, including observability, evaluation, service networking, databases, caching, and object storage.
The supporting services remain available across agent processes and provide the infrastructure used for Langfuse tracing, evaluation workflows, and persisted service data. This mode has the largest disk, memory, startup-time, and operational requirements, but is the recommended profile when the complete platform is needed.
Cyber-AutoAgent supports 4 invocation methods, each with different use cases:
Best for: Automation, scripting, CI/CD pipelines
# Configure via environment variables
export AZURE_API_KEY="your_key"
export AZURE_API_BASE="https://your-endpoint.openai.azure.com/"
export AZURE_API_VERSION="2024-12-01-preview"
export CYBER_AGENT_LLM_MODEL="azure/gpt-5"
export CYBER_AGENT_EMBEDDING_MODEL="azure/text-embedding-3-large"
export REASONING_EFFORT="medium"
export KMP_DUPLICATE_LIB_OK="TRUE"
# Run with uv (recommended)
uv run python src/cyberautoagent.py \
--target "https://example.com" \
--objective "Bug bounty assessment" \
--max-duration 120 \
--provider litellmBest for: Repeated testing with saved config, development
# Uses ~/.cyber-autoagent/config.json for settings
cd src/modules/interfaces/react
npm start -- --auto-run \
--target "https://example.com" \
--objective "Security assessment" \
--max-duration 60Configure via ~/.cyber-autoagent/config.json:
{
"modelProvider": "litellm",
"modelId": "azure/gpt-5",
"embeddingModel": "azure/text-embedding-3-large",
"azureApiKey": "your_key",
"azureApiBase": "https://your-endpoint.openai.azure.com/",
"azureApiVersion": "2024-12-01-preview",
"reasoningEffort": "medium"
}Best for: Isolated environments, clean tooling, reproducibility
With Interactive React Terminal:
docker run -it --rm \
-e AZURE_API_KEY=your_key \
-e AZURE_API_BASE=https://your-endpoint.openai.azure.com/ \
-e CYBER_AGENT_LLM_MODEL=azure/gpt-5 \
-v $(pwd)/outputs:/app/outputs \
cyber-autoagent:latestDirect Python Execution (Override Entrypoint):
docker run --rm --entrypoint python \
-e AZURE_API_KEY=your_key \
-e AZURE_API_BASE=https://your-endpoint.openai.azure.com/ \
-e AZURE_API_VERSION=2024-12-01-preview \
-e CYBER_AGENT_LLM_MODEL=azure/gpt-5 \
-e CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-large \
-e REASONING_EFFORT=medium \
-v $(pwd)/outputs:/app/outputs \
cyber-autoagent:latest \
src/cyberautoagent.py \
--target https://example.com \
--objective "Security assessment" \
--max-duration 60 \
--provider litellmBest for: Observability, team deployments, production monitoring
# Uses docker/.env for configuration
docker compose -f docker/docker-compose.yml up -dCyber-AutoAgent supports 300+ LLM providers via LiteLLM. Examples:
Azure OpenAI:
-e AZURE_API_KEY=your_key
-e AZURE_API_BASE=https://your-endpoint.openai.azure.com/
-e AZURE_API_VERSION=2024-12-01-preview
-e CYBER_AGENT_LLM_MODEL=azure/gpt-5
-e CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-largeAWS Bedrock:
-e AWS_ACCESS_KEY_ID=your_key
-e AWS_SECRET_ACCESS_KEY=your_secret
-e CYBER_AGENT_LLM_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0
-e CYBER_AGENT_EMBEDDING_MODEL=amazon.titan-embed-text-v2:0OpenRouter:
-e OPENROUTER_API_KEY=your_key
-e CYBER_AGENT_LLM_MODEL=openrouter/openrouter/polaris-alpha
-e CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-largeMoonshot AI:
-e MOONSHOT_API_KEY=your_key
-e CYBER_AGENT_LLM_MODEL=moonshot/kimi-k2-thinking
-e CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-large
-e AZURE_API_KEY=azure_key # Required for Azure embeddings
-e AZURE_API_BASE=https://your-endpoint.openai.azure.com/
-e AZURE_API_VERSION=2024-12-01-previewMixed Providers: You can combine any LLM with any embedding model!
# Clone the repository
git clone https://github.com/double16/Cyber-AutoAgent-ng.git
cd cyber-autoagent
# Build and run with Docker Compose (includes observability)
cd docker
docker compose -f docker-compose.yml up -d
# Run a penetration test
docker run --rm \
--network cyber-autoagent_default \
-e AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID} \
-e AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY} \
-e LANGFUSE_HOST=http://langfuse-web:3000 \
-e LANGFUSE_PUBLIC_KEY=cyber-public \
-e LANGFUSE_SECRET_KEY=cyber-secret \
-v $(pwd)/outputs:/app/outputs \
cyber-autoagent \
--target "example.com" \
--objective "Web application security assessment"For just the agent without observability:
# Build the image
# (optional) docker build --pull --platform linux/amd64,linux/arm64 -f docker/Dockerfile.tools -t public.ecr.aws/bramblethorn/cyber-autoagent-ng/tools:latest .
docker build -t cyber-autoagent -f docker/Dockerfile .
# Run with AWS Bedrock
docker run --rm \
-e AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID} \
-e AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY} \
-e AWS_REGION=${AWS_REGION:-us-east-1} \
-v $(pwd)/outputs:/app/outputs \
cyber-autoagent \
--target "192.168.1.100" \
--objective "Network security assessment" \
--provider bedrock
# Run with Ollama (local)
docker run --rm \
-e OLLAMA_HOST=http://host.docker.internal:11434 \
-e OLLAMA_CONTEXT_LENGTH=32768 \
-v $(pwd)/outputs:/app/outputs \
cyber-autoagent \
--target "testsite.local" \
--objective "Basic security scan" \
--provider ollama \
--model qwen3.6:27b- Network Isolation: Deploy in an isolated network segment
- Resource Limits: Set memory and CPU limits in docker-compose.yml
- Secure Keys: Generate proper encryption keys for Langfuse:
# Generate secure keys openssl rand -hex 32 # For ENCRYPTION_KEY openssl rand -base64 32 # For SALT openssl rand -base64 32 # For NEXTAUTH_SECRET
Cyber-AutoAgent uses a modular, three-tier configuration system with automatic model detection and safe token limit allocation.
Configuration Modules:
config/
├── manager.py # Core orchestration (ConfigManager)
├── types.py # Type definitions and dataclasses
├── models/ # Model creation and capabilities
├── system/ # Environment, logging, validation
└── providers/ # Provider-specific helpers
See src/modules/config/README.md for complete module documentation.
Settings are applied in this priority order:
1. CLI/API Arguments (Highest)
└─ Flags: --provider, --model, --max-duration, --max-tokens, --max-cost
└─ Direct parameters to create_agent()
2. Environment Variables (Override)
└─ CYBER_AGENT_LLM_MODEL
└─ CYBER_AGENT_EMBEDDING_MODEL
└─ REASONING_EFFORT
└─ Provider-specific: AZURE_API_KEY, AWS_REGION, etc.
3. Provider Defaults (Fallback)
└─ Safe defaults for all providers
└─ Automatically selected based on provider
Example:
# Default: temperature=0.5 (from provider defaults)
# Override via environment: CYBER_LLM_TEMPERATURE=0.8
# Override via CLI: create_agent(..., temperature=0.7)
# Result: Uses 0.7 (CLI has highest priority)Token limits are automatically detected using the models.dev API with resilient fallback:
Three-Tier Fallback:
- Disk cache (
~/.cache/cyber-autoagent/models.json, 24h TTL) - Live API (
https://models.dev/api.json) - Embedded snapshot (
models_snapshot.json, 432KB bundled)
Benefits:
- Accurate limits for 1,100+ models across 58 providers
- Works offline (embedded snapshot)
- Safe token allocation (50% of actual limit by default)
- Automatic capability detection (reasoning, tools, attachments)
Safe Token Limits:
# Specialist tools use 50% of model's output limit for reliability
safe_max = model_output_limit * 0.5
# Example: Bedrock Claude 3.5 Sonnet v2
# Actual limit: 8,192 tokens
# Safe allocation: 4,096 tokensContext-window limits use the existing provider-aware precedence:
- Explicit override -
CYBER_PROMPT_LIMIT_FORCEenvironment variable - Ollama configuration or runtime detection -
OLLAMA_CONTEXT_LENGTH, model metadata, or the loaded model - Models.dev/static model registry - Authoritative known-model data
- LiteLLM provider detection - Remote-provider model metadata
- Context window maximum -
CYBER_CONTEXT_LIMIT, used as a clamp or configured fallback - Provider defaults - Conservative final resolution
The resolved value is written to the Strands model's context_window_limit and reused for prompt budgeting and
conversation compression. Ollama also receives the same value as num_ctx. Values such as 48,000 are configuration or
detection results, not application defaults. Startup fails when no positive context window can be resolved.
Artifact pages use the same resolved input context window. Each page is limited to 5% of the window at four UTF-8 bytes per token, clamped from 8 KiB through 64 KiB. For example, a 48,000-token context permits 9,600 bytes. Oversized pages are rejected before their content is materialized.
Example fallback configuration:
export CYBER_CONTEXT_WINDOW_FALLBACKS='[
{"azure/gpt-5": ["azure/gpt-4o", "azure/gpt-4"]},
{"anthropic/claude-opus": ["anthropic/claude-sonnet-4-5"]}
]'| Variable | Description | Required |
|---|---|---|
CYBER_AGENT_PROVIDER |
Provider choice (bedrock/ollama/litellm/gemini) | No (default: bedrock) |
CYBER_AGENT_LLM_MODEL |
Main LLM model ID | Yes |
CYBER_AGENT_EMBEDDING_MODEL |
Embedding model ID | No (provider default) |
REASONING_EFFORT |
Reasoning effort (low/medium/high) | No (default: medium) |
MAX_TOKENS |
Override LLM max (output) tokens | No (models.dev default) |
CYBER_AGENT_SWARM_MODEL |
Swarm tool LLM model ID | No |
CYBER_AGENT_SWARM_MAX_TOKENS |
Override specialist max tokens | No (models.dev default) |
MAX_TOKENS_LIMIT |
Override LLM output token upper bound | No (12,000 default) |
MAX_TOKENS_REASONING_LIMIT |
Override LLM output token upper bound (reasoning) | No (32,000 default) |
CYBER_CONTEXT_LIMIT |
Limit detected prompt tokens | No (auto-detected) |
CYBER_PROMPT_LIMIT_FORCE |
Force prompt token limit | No (auto-detected) |
CYBER_SDK_CONTEXT_MANAGER |
Strands context facade (auto, agentic, false) |
No (default: false) |
CYBER_WORKFLOW_PLAN_REFINEMENT_ITERATIONS |
Maximum initial plan critic reviews; 0 disables critique |
No (default: 7) |
CYBER_WORKFLOW_TASK_PROMPT_REFINEMENT_ITERATIONS |
Maximum task prompt critic reviews; 0 disables critique |
No (default: 3) |
CYBER_WORKFLOW_TASK_EXECUTION_CYCLES |
Maximum normal executor passes per task | No (default: 3, minimum 1) |
CYBER_TASK_EVALUATOR_MAX_CORRECTIONS |
Extra executor passes for actionable evaluator feedback | No (default: 1, minimum 0) |
CYBER_TASK_EVALUATOR_ARTIFACT_PAGES_PER_FILE |
Successful pages per authorized evaluator evidence artifact | No (default: 4, minimum 1; 200 lines per page) |
CYBER_SECLISTS_DIR |
Absolute SecLists root for wordlist-consuming tools | No (common locations; container default: /usr/share/seclists) |
CYBER_REPORT_REFINEMENT_CYCLES |
Critic-guided revision cycles per generated report section | No (default: 2; 0 disables) |
CYBER_REPORT_EVIDENCE_GROUPING |
Merge corroborating evidence into canonical report findings | No (default: false) |
CYBER_TAXONOMY_CACHE_DIR |
Local cache directory for CWE and ATT&CK catalogs | No (default: user cache directory) |
CYBER_TAXONOMY_REFRESH_DAYS |
Days before cached taxonomy catalogs are refreshed | No (default: 30) |
CYBER_TAXONOMY_REFRESH |
Fetch current catalogs when the cache is stale | No (default: true) |
CYBER_TAXONOMY_OFFLINE |
Disable taxonomy catalog network refreshes | No (default: false) |
CYBER_TAXONOMY_CATALOG_URL |
Optional normalized taxonomy catalog mirror | No |
CYBER_TOOL_RECOVERY_MAX_POLICY_VIOLATIONS |
Repeated blocked recovery calls before stopping execution | No (default: 2, minimum 1) |
CYBER_TOOL_RECOVERY_MAX_CORRECTIONS |
Changed retries allowed for one failed task invocation | No (default: 2, minimum 1) |
CYBER_TASK_CREATOR_MAX_CORRECTIONS |
Retained correction turns after rejected task creation | No (default: 6, minimum 0) |
CYBER_TASK_ACCEPTANCE_MAX_CORRECTIONS |
Retained correction turns after rejected task acceptance | No (default: 2, minimum 0) |
AWS_ACCESS_KEY_ID |
AWS credentials for Bedrock | For Bedrock provider |
AWS_SECRET_ACCESS_KEY |
AWS credentials for Bedrock | For Bedrock provider |
AWS_REGION |
AWS region (default: us-east-1) | For Bedrock provider |
OLLAMA_HOST |
Ollama API endpoint | For Ollama provider |
OLLAMA_CONTEXT_LENGTH |
Ollama model context length | No, Ollama default |
OLLAMA_TIMEOUT |
Ollama API timeout in seconds | No (default: 120) |
OLLAMA_KEEP_ALIVE |
Ollama model keep alive | No (default: 30m) |
AZURE_API_KEY |
Azure OpenAI API key | For Azure/LiteLLM |
AZURE_API_BASE |
Azure endpoint URL | For Azure/LiteLLM |
AZURE_API_VERSION |
Azure API version | For Azure/LiteLLM |
CYBER_MEMORY_MODE |
Memory query scope: operation or shared |
No (default: operation) |
QDRANT_URL |
Qdrant service endpoint; unset uses outputs/qdrant |
No |
QDRANT_API_KEY |
Optional Qdrant service API key | Only for authenticated services |
QDRANT_COLLECTION |
Qdrant semantic-memory collection | No (default: cyber_autoagent_memories) |
LANGFUSE_HOST |
Langfuse observability endpoint | For observability |
LANGFUSE_PUBLIC_KEY |
Langfuse API public key | For observability |
LANGFUSE_SECRET_KEY |
Langfuse API secret key | For observability |
ENABLE_AUTO_EVALUATION |
Enable automatic Ragas evaluation | For evaluation |
CYBER_RATE_LIMIT_REQ_PER_MIN |
Limit model requests per minute | No (no limit) |
CYBER_RATE_LIMIT_TOKENS_PER_MIN |
Limit model tokens per minute | No (no limit) |
CYBER_RATE_LIMIT_MAX_CONCURRENT |
Limit model concurrent requests | No (Ollama defaults to 1) |
CYBER_HEAP_MONITOR_AUTOSTART |
Auto-start heap monitor thread (0 disables) |
No (default: 1) |
CYBER_AGENT_PRICING_INPUT |
Model price per 1M input tokens | No (defaults to models.dev) |
CYBER_AGENT_PRICING_OUTPUT |
Model price per 1M output tokens | No (defaults to models.dev) |
CYBER_AGENT_PRICING_CACHE_READ |
Model price per 1M cache read tokens | No (defaults to models.dev) |
CYBER_AGENT_PRICING_CACHE_WRITE |
Model price per 1M cache write tokens | No (defaults to models.dev) |
CYBER_SDK_ENABLE_STREAMING |
Enable/disable model streaming. | No (defaults to false) |
Example deployment manifest:
apiVersion: apps/v1
kind: Deployment
metadata:
name: cyber-autoagent
spec:
replicas: 1
selector:
matchLabels:
app: cyber-autoagent
template:
metadata:
labels:
app: cyber-autoagent
spec:
containers:
- name: cyber-autoagent
image: cyber-autoagent:latest
env:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: aws-credentials
key: access-key-id
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: aws-credentials
key: secret-access-key
volumeMounts:
- name: outputs
mountPath: /app/outputs
volumes:
- name: outputs
persistentVolumeClaim:
claimName: outputs-pvc- Access Langfuse UI at http://localhost:3000
- Default credentials: admin@cyber-autoagent.com / changeme
- View real-time traces of agent operations
- Export results for reporting
The React terminal interface provides interactive configuration and real-time monitoring:
# Install and build
cd src/modules/interfaces/react
npm install
npm run build
# Start the interface
npm start
# The interface will guide you through:
# 1. Docker environment setup
# 2. Deployment mode selection (local-cli, single-container, full-stack)
# 3. Model provider configuration (Bedrock, Ollama, LiteLLM, Gemini)
# 4. First assessment executionAccess the interface at http://localhost:3000 when using full-stack deployment with observability.
Qdrant is the semantic-memory backend. With no service variables it uses filesystem storage at outputs/qdrant.
Set QDRANT_URL and, when required, QDRANT_API_KEY to connect to a service. Every point is tagged with exact target
values and operation ID. CYBER_MEMORY_MODE=operation queries both fields; shared omits only the operation criterion.
See the memory guide for the complete model.
export AZURE_API_KEY=your_key
export AZURE_API_BASE=https://your-endpoint.openai.azure.com/
export AZURE_API_VERSION=2024-12-01-preview
export CYBER_AGENT_LLM_MODEL=azure/gpt-5
export CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-large
export REASONING_EFFORT=high
export MAX_TOKENS=8000 # Optional: Override defaultexport AWS_REGION=us-east-1
export CYBER_AGENT_LLM_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0
export CYBER_AGENT_EMBEDDING_MODEL=amazon.titan-embed-text-v2:0
export QDRANT_URL=http://localhost:6333 # Optional; omit for filesystem storage
export REASONING_EFFORT=mediumexport MOONSHOT_API_KEY=your_key
export CYBER_AGENT_LLM_MODEL=moonshot/kimi-k2-thinking
export CYBER_AGENT_EMBEDDING_MODEL=azure/text-embedding-3-large
export AZURE_API_KEY=your_azure_key # For embeddings
export AZURE_API_BASE=https://your-endpoint.openai.azure.com/
export AZURE_API_VERSION=2024-12-01-preview
export QDRANT_URL=http://localhost:6333 # Optional memory serviceexport OLLAMA_HOST=http://localhost:11434
export OLLAMA_CONTEXT_LENGTH=32768
export CYBER_AGENT_LLM_MODEL=qwen3.6:27b
export CYBER_AGENT_EMBEDDING_MODEL=nomic-embed-text:latest
export CYBER_CONTEXT_WINDOW_FALLBACKS='[
{"qwen3-coder:30b": ["qwen3-coder:14b", "llama3.2:3b"]}
]'Common deployment issues:
- Container fails to start: Check Docker logs with
docker logs cyber-autoagent - AWS credentials error: Ensure IAM role has Bedrock access and correct region
- Ollama connection failed: Verify Ollama is running and accessible at specified host
- Out of memory: Increase Docker memory limits or reduce token-heavy workloads
- React interface issues: Run
npm run buildafter any code changes - Memory backend errors: Verify environment variables and network connectivity
- Model not found: Check model ID format (use
provider/modelfor LiteLLM) - Token limit errors: Verify models.dev snapshot exists at
src/modules/config/models/models_snapshot.json - Specialist failures: Check swarm max_tokens configuration (should be >100 tokens)