diff --git a/modules/ROOT/nav.adoc b/modules/ROOT/nav.adoc index 39ec1cab6..78e27f949 100644 --- a/modules/ROOT/nav.adoc +++ b/modules/ROOT/nav.adoc @@ -23,6 +23,7 @@ *** xref:model-proxy-request.adoc[] *** xref:model-proxy-semantic-service.adoc[] *** xref:model-proxy-try-out.adoc[] + *** xref:exp-detect-and-contain-rogue-agents.adoc[] ** xref:exp-scanners-add-from-providers.adoc[] *** xref:exp-scanners-prerequisites-reference.adoc[] *** xref:exp-scanners-view-details.adoc[] @@ -30,6 +31,8 @@ *** xref:exp-providers-manage.adoc[] ** xref:exp-governance-view-cost-and-token-usage.adoc[] *** xref:model-proxy-token-reports.adoc[] + *** xref:exp-models-manage-costs.adoc[] + *** xref:exp-model-wallets-manage.adoc[] ** xref:exp-akamai-risk-correlation.adoc[] ** xref:exp-governance-work-with-strategies.adoc[] *** xref:exp-governance-create-strategy.adoc[] @@ -48,6 +51,10 @@ * xref:learning-map-mulesoft-ai.adoc[] * xref:agent-fabric-overview.adoc[Agent Fabric] ** xref:learning-map-agent-fabric.adoc[Get Started with Agent Fabric] + ** xref:agent-fabric-use-cases.adoc[] + *** xref:af-use-case-mcp-bridge.adoc[] + *** xref:af-use-case-cost-management.adoc[] + *** xref:af-use-case-orchestration.adoc[] ** xref:agent-fabric-release-notes.adoc[] ** xref:agent-networks-get-started.adoc[] * xref:learning-map-api-management.adoc[API Management] diff --git a/modules/ROOT/pages/af-use-case-cost-management.adoc b/modules/ROOT/pages/af-use-case-cost-management.adoc new file mode 100644 index 000000000..afc64d19d --- /dev/null +++ b/modules/ROOT/pages/af-use-case-cost-management.adoc @@ -0,0 +1,220 @@ += Control LLM Costs in Agent Fabric + +Agent Fabric gives you full visibility into token consumption and spend across every model call, and the controls to enforce limits before usage compounds. Use Model Wallets to set hard limits in tokens or dollars that agents can't exceed, and attribute every token and dollar to the team, application, or agent that generated them. + +Key benefits include: + +* Stop budget overruns before they happen: Model Wallets block requests when a budget limit is reached, so an unattended agent can't run spend past the cap. The limit is enforced, not just flagged. +* Cut cost without cutting quality: Semantic routing sends routine queries to a cheaper, faster model and reserves higher-cost models for complex work, lowering average cost per query while holding response quality. +* Turn spend into decisions: See token consumption and cost by application, agent, user, and model in one place, so you can act on what's driving cost instead of guessing. +* End cost surprises: Trace every token and dollar to the team, application, or agent that drove it, so overruns have an owner and a root cause. +* Govern once, everywhere: Apply the same policies, prompt protection, and automatic fallback across every provider, so adding a model or team doesn't mean rebuilding controls. * Govern once, everywhere: Apply the same policies, prompt protection, and automatic fallback across every provider, so adding a model or team doesn't mean rebuilding controls. + +== The Problem + +As agentic systems scale, LLM costs can quickly spiral out of control without proper visibility and enforcement: + +* Unpredictable costs: Token usage varies widely based on agent behavior and queries. +* No visibility: Costs can't be attributed to specific applications, agents, or users. +* Inefficient model usage: Expensive models handle simple queries that cheaper models could address. +* Budget overruns: Nothing enforces the budget, so a number in the plan isn't a real limit. +* No optimization path: Without data on usage patterns, optimization is guesswork. + +Gain visibility into what's driving costs and the controls to optimize spending without sacrificing quality. + +== The Solution + +Agent Fabric provides comprehensive cost management for agentic AI through the Model Proxy: + +* Usage visibility: Tracks token consumption by application, agent, user, and model. +* Semantic routing: Routes each query to the appropriate model based on complexity, automatically. +* Budget enforcement: Caps token or spend usage with Model Wallets that block requests when a budget limit is reached. +* Model optimization: Provides data-driven insight into which models are used for which tasks. +* Cost attribution: Shows exactly what's driving your spend, down to the team, app, or agent. + +== How Semantic Routing Works + +Not all queries require the most powerful (and most expensive) AI models. Semantic routing analyzes query complexity and routes each request appropriately: + +[source,text] +---- +Simple query: "What's the status of order #12345?" +→ Route to: Fast, cost-effective model + +Complex query: "Analyze this quarter's sales trends, identify anomalies, and recommend strategic adjustments" +→ Route to: Powerful, higher-cost model +---- + +The Model Proxy makes these routing decisions automatically based on configurable rules, maintaining quality while minimizing cost. + +== How Model Wallets Enforce Usage Limits + +Budget alerts tell you when you've overrun a limit. They don't stop it. By the time an alert fires at 90 percent of budget, an agent running unattended can push token or spend usage far past the limit before anyone reads the notification. + +A Model Wallet closes that gap by turning a budget from a number in a plan into a financial control that agents can't overrun: + +* Hard limits: Set enforced token or spend limits per provider so that when a budget reaches its limit, requests are blocked rather than flagged. +* Periodic caps: Set a daily, weekly, or monthly limit that resets automatically on a schedule, so token or spend usage stays bounded each period. +* Flexible metrics: Set limits in dollars (USD) or tokens, so budgets reflect the metric that matters most to your team. +* Identity-scoped controls: Tie usage controls to a Model Proxy's client identity so that each application or agent routes through its own governed access point with dedicated budget limits. + +Model Wallets build on the Model Proxy. The proxy provides one governed access point for every provider, making token consumption and spend visible; Model Wallets add the enforcement layer on top, so the usage you can see is also usage you can cap. + +== Who This Is For + +These cost management capabilities are ideal for: + +* FinOps teams managing cloud and AI spending +* Engineering leaders optimizing infrastructure costs +* Product teams building cost-effective agentic features +* Enterprises scaling agent deployments beyond proof of concept +* Any organization where LLM costs are a significant budget item + +== Example Scenarios + +These scenarios demonstrate how different organizations use Agent Fabric to bring LLM costs under control. + +=== Scenario 1: Enterprise-Wide Cost Visibility + +*Challenge:* A large organization has agents deployed across multiple business units but no visibility into which ones are driving LLM costs. + +With Agent Fabric you get complete cost visibility: + +. Route all LLM traffic through the Model Proxy. +. Tag requests by business unit, application, and agent. +. Generate cost dashboards showing usage patterns. +. Identify high-cost agents and opportunities for optimization. + +*Result:* Every dollar of LLM spend maps to the business unit, application, and agent that drove it. Leaders can see which deployments cost the most and target optimization where it pays off, instead of splitting an unattributed bill across teams. + +=== Scenario 2: Semantic Routing for Customer Support + +*Challenge:* A customer support agent system uses expensive models for all queries, including simple lookups. + +With semantic routing you cut cost without sacrificing quality: + +. Configure semantic routing rules: + * Simple queries (order status, account lookups) route to a cost-effective model. + * Complex queries (troubleshooting, technical issues) route to an advanced model. +. The Model Proxy analyzes each request and routes it appropriately. +. Monitor quality metrics to ensure the customer experience isn't impacted. + +*Result:* Routine lookups, which are often the majority of support traffic, move to a model that can cost roughly an order of magnitude less per query, while complex cases still reach a more advanced, higher-cost model. Average cost per query drops with no measurable change in resolution quality. + +=== Scenario 3: Budget Controls for Development Teams + +*Challenge:* Development teams experimenting with agents have no spending limits, leading to budget surprises. + +With Model Wallets you get budgets agents can't overrun: + +. Create a separate Model Proxy for each team, with budget limits configured per provider. +. Set a daily, weekly, or monthly limit in dollars (USD) or tokens depending on how the team tracks spend. +. Enable hard limits so requests are blocked when a budget is reached rather than allowed to overrun. +. Give teams visibility into budget consumption through the Model Proxy dashboard. + +*Result:* A runaway experiment stops at its cap instead of surfacing as a month-end invoice. Finance gets predictable per-team spend, and teams keep the autonomy to experiment within a limit they can see. + +=== Scenario 4: Multi-Model Optimization + +*Challenge:* An organization uses multiple LLM providers but doesn't know which models are most cost-effective for its use cases. + +With Agent Fabric you get evidence-based model selection: + +. Deploy agents with flexible model configuration. +. Agent Fabric tracks cost and quality metrics per model. +. Analyze which models deliver the best cost/quality ratio for each use case. +. Adjust routing rules based on the data. + +*Result:* Model selection becomes an evidence-based decision instead of a default. Each use case runs on the provider that delivers the best cost and quality ratio for it, and routing keeps pace as providers change pricing or release new models. + +== Implementation Steps + +Follow these steps to set up cost management in Agent Fabric. + +=== Step 1: Deploy the Model Proxy + +. Set up the Model Proxy as part of your Omni Gateway deployment. +. Configure integration with your LLM providers. +. Route agent traffic through the Model Proxy. + +=== Step 2: Implement Usage Tracking + +. Tag all requests with relevant metadata: + * Business unit or cost center + * Application name + * Agent identifier + * User or session ID +. Configure logging and metrics collection. +. Set up dashboards for cost visibility. + +=== Step 3: Establish Baseline Costs + +. Monitor usage patterns for one to two weeks without optimization. +. Identify current cost drivers. +. Categorize query types by complexity and frequency. +. Calculate baseline cost metrics. + +=== Step 4: Configure Semantic Routing + +. Define query complexity categories. +. Assign an appropriate model to each category. +. Create routing rules in the Model Proxy. +. Test routing with sample queries to ensure quality. + +=== Step 5: Set Budget Controls with Model Wallets + +. Create a Model Proxy per team, application, or agent to scope budget controls to a client identity. +. Add a budget for each provider the Model Proxy routes to, choosing a daily, weekly, or monthly period. +. Set the limit in dollars (USD) or tokens, and enable hard limits where overruns are unacceptable so requests are blocked when a limit is reached. +. Share the Model Proxy dashboard with teams so they can track budget consumption per provider. + +=== Step 6: Monitor and Optimize + +. Track cost trends after optimization. +. Monitor quality metrics to ensure no degradation. +. Refine routing rules based on results. +. Share cost savings and optimization wins. + +== Metrics to Track + +Track these metrics to measure and improve the impact of cost control. + +=== Cost Metrics + +* Total token usage by time period +* Cost per application, agent, and user +* Cost per model and provider +* Average cost per query + +=== Usage Metrics + +* Queries per time period +* Model distribution (which models handle what percentage of queries) +* Token usage distribution by query complexity +* Peak usage times and patterns + +=== Quality Metrics + +* Response quality scores +* User satisfaction ratings +* Error rates by model +* Average response time by model + +=== Optimization Metrics + +* Cost savings from semantic routing +* Percentage of queries routed to cost-effective models +* Budget adherence by team and application +* ROI of optimization efforts + +== Related Documentation + +* xref:agent-fabric-use-cases.adoc[] + +== Next Steps + +With cost controls in place, you're ready to: + +* xref:af-use-case-policy-enforcement.adoc[Add policy enforcement for governance] +* xref:use-case-orchestration.adoc[Build complex workflows with cost awareness] +* xref:use-case-identity.adoc[Move to production with identity and cost controls] diff --git a/modules/ROOT/pages/af-use-case-mcp-bridge.adoc b/modules/ROOT/pages/af-use-case-mcp-bridge.adoc new file mode 100644 index 000000000..12fe2ce44 --- /dev/null +++ b/modules/ROOT/pages/af-use-case-mcp-bridge.adoc @@ -0,0 +1,324 @@ += Make APIs Agent-Ready with MCP Bridge + +Agent Fabric MCP Bridge transforms your existing API instances into agent-ready tools without touching code. Choose which operations to expose as tools, limit agents to read operations where needed, and make your Mule or API management investment immediately available to agents with the same governance policies you already use. + +Key benefits of MCP Bridge include: + +* Connect agents in minutes, not weeks: Turn an existing API into agent-ready tools through configuration alone, skipping the custom integration each agent platform would otherwise require. +* Never expose an operation by accident: Agents can call only the operations you explicitly map as tools, so write and delete stay off-limits unless you choose to expose them. +* One controlled interface for every agent: All agents reach your APIs the same way, eliminating the inconsistent, per-platform access that's hard to audit. +* Get more from what you already built: Reuse APIs managed through Anypoint Platform without rebuilding them for agentic use. +* Govern agent API access from one place: A single point of control means access, policy, and monitoring don't fragment as you add agents. + +== The Problem + +You've invested heavily in API infrastructure, but adapting these APIs for agentic use is challenging: + +* API complexity: Existing APIs aren't designed for agent consumption. +* Security concerns: APIs often include write and delete operations that are too risky for agents. +* Integration overhead: Building custom integrations for each agent platform is time-consuming. +* Inconsistent access: Different agents access the same APIs in different ways. +* Governance gaps: It's hard to track and control how agents use existing APIs. + +Bridge your existing API infrastructure to agents safely and efficiently with MCP Bridge. + +== The Solution + +Agent Fabric Model Context Protocol (MCP) Bridge enables safe, efficient API integration for agents: + +* Rapid integration: Connects existing API instances to agents through simple configuration, with no rewriting required. +* Selective exposure: Exposes only the API operations you choose as discrete MCP tools that agents can discover and call. Limit agent access to read-only operations when needed, without modifying the underlying API. +* Centralized management: Provides a single point of control for API access, with governance policies applied from the same portfolio. +* Agent-facing abstraction: Exposes existing REST API operations as MCP tools so agents can call them without needing API-specific details. +* Reuse existing assets: Works with your existing MuleSoft API infrastructure without rebuilding assets. + +== How MCP Bridge Works + +MCP Bridge operates at the Omni Gateway layer as a set of automated policies. You configure it through API Manager, and it deploys to the same gateway infrastructure you already use to manage API traffic. + +Each agent request follows this path: + +. The AI agent makes a request using MCP. +. The request flows through Omni Gateway, where MCP Bridge policies execute. +. MCP Bridge maps the request to the underlying API operation. +. Only operations you expose as MCP tools are reachable. +. The call reaches the backend API. +. Responses flow back through the gateway to the agent. + +== Who This Is For + +MCP Bridge is ideal for: + +* Organizations with existing API infrastructure +* Security teams concerned about agent access to sensitive operations +* Integration teams looking to enable agents quickly without custom development +* Architects who are designing safe agent access to enterprise systems +* Teams who want to build on existing investments in API management and governance + +== Example Scenarios + +These scenarios demonstrate how different organizations use MCP Bridge to solve specific agent access challenges. + +=== Scenario 1: Customer Service Agent with Read-Only Access + +*Challenge:* Customer service agents need to answer questions about customer accounts and orders but can't modify data. Without MCP Bridge, you build custom integrations for each agent platform. This approach risks exposing write operations and creates inconsistent behavior across platforms. + +With MCP Bridge, you get rapid deployment with built-in safety: + +. In API Manager, create an MCP Bridge instance for your customer API. +. Map only GET operations to MCP tools. +. Deploy the instance (policies execute on Omni Gateway). +. Connect all agent platforms to the MCP server endpoint. + +*Result:* Agents answer account and order questions with zero write access exposed, from a single configuration. There is no per-platform integration to build or maintain, and no path for an agent to modify customer data. + +=== Scenario 2: Inventory Check Agent + +*Challenge:* Your inventory API includes operations to check stock levels, reserve items, adjust quantities, and process transfers. Sales agents, warehouse agents, and customer service agents all need to check current stock, but only warehouse systems should adjust it. + +Without MCP Bridge, you either expose the full API (risking accidental adjustments) or build filtered endpoints for each agent type. + +With MCP Bridge, you give all agents read-only inventory access from a single configuration: + +. In API Manager, create an MCP Bridge instance from the inventory API. +. Map only GET operations for stock queries to MCP tools. +. Connect all agent types to the MCP server endpoint. +. Monitor agent usage in the enhanced MuleSoft experience. + +*Result:* Sales, warehouse, and customer service agents all read live stock levels from one read-only tool surface, while stock adjustments stay restricted to warehouse systems, without building or maintaining a filtered endpoint per agent type. + +=== Scenario 3: Financial Data Agent with Layered Tool Selection + +*Challenge:* Different agents need different levels of access to financial APIs. + +Without MCP Bridge, you build separate custom integrations for each agent access level. This duplication increases the risk that a high-privilege operation reaches the wrong agent. + +With MCP Bridge, you get fine-grained access control with safety guarantees: + +. In API Manager, create multiple MCP Bridge instances from your financial API, each with different operation mappings: + * Basic server: Map only GET operations for account balances. + * Analytics server: Map GET operations for transaction history and reports. + * Approval server: Map GET operations plus POST for creating approval requests. +. Connect each agent to the appropriate MCP server endpoint based on access requirements. +. In API Manager, apply policies to each MCP server instance. + +*Result:* Each agent gets exactly the access its role requires. Agents balances only, analytics or approval requests from one financial API, with no duplicated integrations and no path for a high-privilege operation to reach the wrong agent. + +=== Scenario 4: Mule Integration Platform + +*Challenge:* Your organization has extensive Mule infrastructure and wants to enable agentic access without rebuilding integrations. + +Without MCP Bridge, you would need to build agent-specific adapters on top of each Mule API, duplicating logic and bypassing existing Anypoint Platform governance. + +With MCP Bridge, you extend your existing Mule investment to agentic use cases: + +. In API Manager, create MCP Bridge instances for key Mule APIs. +. Map the API operations you want to expose to MCP tools. +. Deploy the instances (policies execute on your existing Omni Gateway infrastructure). +. Monitor through the enhanced MuleSoft experience and apply policies in API Manager. + +*Result:* Existing Mule APIs become agent-callable without new adapters or duplicated logic, and agent traffic inherits the same Anypoint Platform governance already protecting those APIs. + +== Implementation Steps + +=== Step 1: Identify APIs for Agent Access + +. List APIs that agents access. +. Document the current API operations (GET, POST, PUT, DELETE). +. Identify which operations are safe for agents. +. Prioritize APIs by business value and risk. + +=== Step 2: Plan Your Tool Selection + +For each API instance, decide: + +* Which operations can agents access? +* Are read-only operations sufficient? +* If agents require write access, which specific write operations do they use? +* What Omni Gateway policies apply to the resulting MCP server? + +=== Step 3: Create and Deploy an MCP Bridge Instance + +. In API Manager, create an MCP Bridge instance for your target API. +. Configure downstream settings (base path, protocol, port) and upstream settings (route label, upstream URL). +. Map specific API operations to MCP tools by selecting the HTTP method and resource for each tool. +. Define tool names, descriptions, and input schemas for each mapped operation. +. Test the MCP server endpoint with sample requests. + +For more information, see xref:api-manager::create-instance-task-mcp-bridge.adoc[]. + +=== Step 4: Apply Omni Gateway Policies + +MCP Bridge controls which API operations are available as tools. Omni Gateway policies control how agents can use them at runtime. After deploying MCP Bridge, apply policies in API Manager to the MCP server instance to enforce authentication, rate limiting, audit logging, and alerting. + +See xref:use-case-policy-enforcement.adoc[Policy Enforcement] for implementation steps. + +=== Step 5: Connect Agents + +. Update agents to use MCP Bridge. +. Test with realistic scenarios. +. Verify that tool selection works (agents can only call the operations you expose). +. Monitor initial usage. + +=== Step 6: Monitor and Refine + +. In the enhanced MuleSoft experience, review latency, error rates, and request volume for your MCP server. +. Analyze which tools agents call and identify missing operations. +. Check gateway logs for attempts to call unexposed operations. +. To change tool mappings, create a new MCP Bridge instance with the updated configuration in API Manager. MCP Bridge instances are immutable after deployment. + +For information about monitoring MCP servers and other services, see xref:exp-services-monitoring.adoc[]. + + +== Tool Selection Strategies + +Choose which API operations to expose as discrete MCP tools that agents can discover and call. During configuration, you map specific operations (HTTP method and resource) to tools. Expose all operations, only read operations, or any subset you choose, without modifying the underlying API. + +Consider a Customer API with both read and write operations: + +[source,text] +---- +Customer API: +- GET /customers/{id} (read) +- POST /customers (create) +- PUT /customers/{id} (update) +- DELETE /customers/{id} (delete) +- GET /customers/{id}/orders (read) +- POST /customers/{id}/orders (create) +---- + +After you configure MCP Bridge to expose only read operations, agents see a reduced tool surface: + +[source,text] +---- +Customer API (Agent Tool Surface): +- GET /customers/{id} (read - exposed as an MCP tool) +- GET /customers/{id}/orders (read - exposed as an MCP tool) +---- + +The agent can read customer and order data but can't create, modify, or delete anything. The write operations exist in the underlying API but aren't exposed as MCP tools. + +=== Read-Only Operation Mapping + +Map only GET/read operations to tools first. Explicitly map write operations only when a use case requires them. + +=== Operation Allowlisting + +Select specific operations to expose rather than trying to identify operations to block. Allowlisting is safer and easier to maintain. + +=== Data Scoping + +Combine tool selection with data filters (for example, expose only operations that access certain data categories). + +For runtime controls such as rate limiting and time-based access restrictions, see xref:use-case-policy-enforcement.adoc[Policy Enforcement]. + +== Common Tool Selection Patterns + +These patterns represent typical approaches to selecting which operations to expose as tools based on your security requirements and use cases. + +=== Read-Only Tool Surface + +Use this pattern when agents only retrieve data. Expose all GET/read operations as tools and exclude all write (POST, PUT, DELETE) operations. + +[source,text] +---- +(exposed) GET /resource/{id} +(exposed) GET /resource +(not exposed) POST /resource +(not exposed) PUT /resource/{id} +(not exposed) DELETE /resource/{id} +---- + +=== Read and Create Tool Surface + +Use this pattern when agents look up records and create new ones but don't modify or delete existing data. + +[source,text] +---- +(exposed) GET /resource/{id} +(exposed) POST /resource +(not exposed) PUT /resource/{id} +(not exposed) DELETE /resource/{id} +---- + +=== Read and Request-Only Tool Surface + +Use this pattern when agents initiate a workflow (for example, submitting an approval request) without directly modifying records. + +[source,text] +---- +(exposed) GET /resource/{id} +(exposed) POST /resource/request +(not exposed) PUT /resource/{id} +(not exposed) DELETE /resource/{id} +---- + +=== Granular Tool Surface + +Use this pattern for fine-grained control, exposing specific safe operations and blocking others based on individual business risk assessments. For example, expose GET /orders/{id} so agents can view order details, but block PUT /orders/{id}/status to prevent agents from changing order status. This pattern gives agents access to the information they need while protecting sensitive operations. + +[source,text] +---- +(exposed) GET /orders/{id} +(exposed) POST /orders/{id}/notes (add note - safe) +(not exposed) PUT /orders/{id}/status (change status - risky) +(not exposed) DELETE /orders/{id} (delete order - risky) +---- + +== Security Considerations + +When deploying MCP Bridge, follow security best practices to minimize risk and maintain control over agent access to your APIs. + +=== Default Deny + +Start with no access and explicitly enable only what agents use. + +Apply Omni Gateway policies to enable audit logging and monitor MCP server usage through Anypoint Monitoring. See xref:use-case-policy-enforcement.adoc[Policy Enforcement]. + +=== Regular Reviews + +Periodically review: + +* Which operations are in use +* Whether your tool selection is still appropriate +* Whether to remove additional operations from the agent tool list + +== Integration with Existing Tools + +MCP Bridge integrates with your existing MuleSoft infrastructure, identity systems, and gateway deployment to build on investments you've already made. + +=== Mule Integration + +MCP Bridge works directly with Mule APIs managed in Anypoint Platform, so you can expose existing Mule API instances as MCP servers without rebuilding them. + +* Use MCP Bridge with existing Mule APIs. +* Use Anypoint Platform governance. +* Extend Mule monitoring to agent access. +* Reuse existing API policies. + +=== Omni Gateway Integration + +MCP Bridge runs on Omni Gateway infrastructure. MCP server instances are managed through API Manager and protected by Omni Gateway policies, just like API instances. + +* Create and deploy MCP Bridge instances on the same Omni Gateway that manages your APIs. +* Apply policies to MCP servers the same way you apply them to APIs. +* Use unified monitoring and alerting across APIs and MCP servers. + +=== Identity Systems + +MCP Bridge can work with existing identity infrastructure so that agent access to APIs respects the same authentication and authorization rules already in place. + +* Integrate with existing authentication. +* Use existing authorization rules where applicable. +* Extend identity-based access control to agents. + +== Related Documentation + +* xref:agent-fabric-use-cases.adoc[] + +== Next Steps + +With MCP Bridge providing safe API access, you're ready to: + +* xref:af-use-case-policy-enforcement.adoc[Apply Omni Gateway policies to broker endpoints] +* xref:af-use-case-orchestration.adoc[Build agent networks] diff --git a/modules/ROOT/pages/af-use-case-orchestration.adoc b/modules/ROOT/pages/af-use-case-orchestration.adoc new file mode 100644 index 000000000..48974d33e --- /dev/null +++ b/modules/ROOT/pages/af-use-case-orchestration.adoc @@ -0,0 +1,204 @@ += Orchestrate Multi-Agent Processes with Agent Broker + +Agent brokers and agent networks coordinate task delegation across specialized agents with guided determinism. This coordination produces reliable, predictable outcomes from complex business processes. You define broker routing logic in Agent Script and configure the network composition in `agent-network.yaml`. This file brings agents, LLMs, and MCP servers together in a single coordinated system. Use this approach when a single agent can't handle the process and your business process requires multiple specialized agents working in sequence or in parallel. + +Key benefits of agent networks and brokers include: + +* Reliable outcomes from complex processes: Graph-based routing runs operations in the correct order with defined paths for both expected outcomes and error conditions, so multi-step processes behave predictably instead of drifting. +* Better results by using the right specialist: Route each task to the agent best suited to handle it, instead of stretching one agent to do everything and accepting weaker output on every task. +* Build once, reuse across teams: Publish brokers and agent networks to your portfolio so other teams compose them into their own networks rather than rebuilding the same process. +* Know exactly where a process failed: Monitor broker routing decisions, agent performance, and request flows with Agent Visualizer, so a failure points to the agent and step that caused it. +* Start small, scale when ready: Begin with a simple agent network and add brokers as complexity grows. You don't need a broker to get started, so there's no upfront coordination cost. + +== The Problem + +A single agent can handle straightforward tasks, but complex business processes expose the limits of working with one agent in isolation: + +* Specialization gaps: One agent can't excel at research, financial analysis, regulatory review, and customer communication simultaneously. +* No intelligent routing: Without a broker, there's no component to match an incoming request to the right specialist agent based on context. +* Process reliability: Multi-step processes that involve multiple agents and systems are difficult to make reliable and predictable without explicit coordination. +* Opaque execution: When something goes wrong in a multi-agent process, it's hard to know which agent failed, why, and what state the process was in. +* Duplication: Teams that build similar multi-agent processes independently create inconsistent behavior and duplicated effort. + +Making a single agent smarter doesn't solve these challenges. They require a coordination layer—an agent network with a broker. + +== The Solution + +Agent Fabric's agent brokers and agent networks provide a coordination layer for multi-agent processes: + +* Agent networks: A YAML-configured composition of agents, brokers, LLMs, and MCP servers that defines the structure of your agentic solution. +* Agent brokers: Intelligent routing services, defined in Agent Script, that delegate tasks to the right A2A-compliant agent based on context. +* Guided determinism: Graph-based broker logic makes sure that tasks follow defined paths, handling both expected outcomes and error conditions. +* Anypoint Exchange integration: Published agent networks and brokers appear in Exchange for discovery and reuse across your organization. +* Integrated observability: Agent Visualizer displays the network topology, real-time request flows, and performance metrics for the entire network. + +== How Agent Brokers Work + +When a request arrives at a broker, it follows this path: + +. A request arrives at the broker, either from an external caller or from another broker higher in the network. +. The broker evaluates the request against its routing logic, defined as a graph of nodes in Agent Script. +. The broker generates a context ID and task ID to track state across the interaction. +. The broker delegates the task to the A2A-compliant agent or sub-broker best suited to handle it. +. The assigned agent processes the task, calling LLMs and MCP servers as needed. +. Results return to the broker, which continues routing through the graph until the process is complete. +. The broker returns the final result to the original caller. + +== Who This Is For + +Agent brokers and agent networks are for: + +* Teams building multi-agent solutions where tasks are routed across specialized agents based on context +* Companies that need reliable, predictable outcomes from complex processes spanning multiple agents and systems +* Architects designing reusable agentic components that other teams can discover and compose +* Teams that need end-to-end visibility into how multi-agent processes execute and where failures occur + +== Example Scenarios + +=== Scenario 1: Order Processing + +*Challenge:* Processing customer orders involves checking inventory, routing to fulfillment, handling payment, and sending confirmation. Each step is handled by a specialized agent. Without coordination, the process is fragile and hard to monitor. + +Agent brokers and agent networks provide reliable coordination with clear routing logic: + +. Define an agent network in `agent-network.yaml` that registers a broker, an inventory agent, a payment agent, and a fulfillment agent. +. Define the broker in Agent Script with nodes for each step and explicit routing logic for inventory shortfalls, payment failures, and partial fulfillment. +. Deploy the agent network to CloudHub 2.0. +. Monitor execution in Agent Visualizer to trace each step and identify where issues occur. + +*Result:* Orders move through inventory, payment, and fulfillment on defined paths, with explicit handling for shortfalls and payment failures instead of silent breakage. When a step fails, Agent Visualizer shows which agent and which state, so recovery is fast rather than a guessing game. + +=== Scenario 2: Multi-Discipline Research + +*Challenge:* Answering a strategic research question requires market analysis, competitive intelligence, technical feasibility assessment, and financial modeling — capabilities that belong in separate specialized agents. + +With an agent network, a broker decomposes the request and coordinates the specialists: + +. Define an agent network that registers market, competitive, technical, and financial analysis agents alongside a synthesis agent. +. Define a broker in Agent Script that routes each research sub-task to the appropriate specialist agent, then routes all results to the synthesis agent. +. The synthesis agent produces the final report from the aggregated specialist outputs. +. Agent Visualizer shows the routing path, which agents were called, and where time was spent. + +*Result:* A strategic question is answered by four specialists working in parallel and synthesized into one report. The report is faster and higher-quality than a single general-purpose agent, with a visible trail of which analysis contributed what. + +=== Scenario 3: Document Processing Pipeline + +*Challenge:* Processing incoming documents requires classification, data extraction, quality validation, and routing for human review — in sequence, with defined rules for what happens when quality thresholds aren't met. + +With guided determinism in Agent Script: + +. Define an agent network that includes a classification agent, an extraction agent, and a review routing agent. +. Define a broker with nodes for each processing stage and explicit paths for low-confidence extractions that require human review. +. Deploy the network and monitor processing quality and throughput in Agent Visualizer. + +*Result:* Documents flow through classification, extraction, and validation automatically, and only low-confidence cases route to a human. Staff review the exceptions instead of every document, and quality thresholds are enforced by the graph rather than left to chance. + +=== Scenario 4: Customer Support Triage + +*Challenge:* Incoming support requests range from simple account questions to complex technical issues and billing disputes. Each type requires a different specialized agent, and misrouting wastes time. + +With a broker handling intelligent triage: + +. Define a broker in Agent Script that classifies the incoming request and routes it to the appropriate specialist: account agent, technical agent, or billing agent. +. Each specialist agent calls the relevant MCP servers to retrieve customer data, account history, or billing records. +. The broker consolidates the specialist's response and returns it to the caller. +. Publish the agent network to Exchange so other teams can reuse the triage broker. + +*Result:* Requests reach the right specialist on the first hop, so account, technical, and billing issues stop bouncing between the wrong agents. Publishing the triage broker to Exchange lets other teams adopt the same routing instead of rebuilding it. + +== Implementation Steps + +=== Step 1: Map the Process + +. List the steps in the process and identify which require specialized capabilities. +. Document decision points such as where the process branches based on outcomes or data. +. Identify which steps can run in parallel and which must be sequential. +. Determine what backend systems each agent needs to call, and whether MCP servers already expose those systems. + +=== Step 2: Design the Agent Network + +. Identify which agents exist and which to build. +. Determine whether a broker is needed. If the process involves multiple specialists that are routed dynamically, add a broker. +. Sketch the broker graph: nodes represent steps, edges represent routing logic between them. +. Identify which LLMs agents use for reasoning and which MCP servers provide backend access. + +=== Step 3: Configure agent-network.yaml + +. In Anypoint Code Builder, create a new agent network project. +. Configure `agent-network.yaml` to declare the registry (agents and MCP servers), context (connections), and brokers. +. For each broker, create an Agent Script file (`.agent`) that defines the routing graph and node behavior. + +For more information, see xref:agent-network::af-define-your-agent-network-specification.adoc[Define Your Agent Network Specification]. + +=== Step 4: Deploy and Test + +. Deploy the agent network to CloudHub 2.0. +. Send test requests that exercise each routing path in the broker graph. +. Verify that each agent receives the tasks it is responsible for. +. Test failure paths. + +=== Step 5: Monitor with Agent Visualizer + +. Open Agent Visualizer to view the network topology and confirm that agents, brokers, and MCP servers are connected as expected. +. Send live requests and watch routing decisions in real time. +. Use performance metrics to identify latency, error rates, and bottlenecks. +. Review historical analysis to identify patterns and optimize routing logic. + +For more information, see xref:agent-visualizer::index.adoc[Agent Visualizer]. + +=== Step 6: Publish for Reuse + +. After validating the agent network, publish it to Exchange. +. Other teams can discover the published broker and reuse it as a component in their own agent networks. +. Apply Omni Gateway policies to the broker endpoint to enforce authentication, rate limiting, and audit logging. + +See xref:use-case-policy-enforcement.adoc[Policy Enforcement] for implementation steps. + +== Broker Design Patterns + +These patterns represent common approaches to structuring broker routing logic in Agent Script. + +[cols="1,2,2"] +|=== +| Pattern | Description | When to Use + +| Sequential Routing +| Route a request through a fixed sequence of specialist agents, where each agent's output is the input for the next. +| Steps must run in order and each step depends on the previous result. + +| Parallel Dispatch +| Dispatch a request to multiple specialist agents simultaneously and aggregate the results. +| Sub-tasks are independent and can run concurrently for faster completion. + +| Conditional Routing +| Route a request to different agents based on the content or classification of the request. +| The right specialist depends on context that isn't known until the request arrives. + +| Hierarchical Delegation +| A broker delegates to another broker, which in turn delegates to specialist agents. +| Complex processes can be decomposed into distinct sub-processes, each managed by its own broker. +|=== + +== Security Considerations + +When deploying agent networks and brokers in production: + +* Apply Omni Gateway policies to broker endpoints to enforce authentication, rate limiting, and audit logging. See xref:use-case-policy-enforcement.adoc[Policy Enforcement]. +* Make sure that each specialist agent's MCP server connections use appropriate authentication. +* Review which agents have access to which backend systems and apply the principle of least privilege to MCP tool selection. +* Use Agent Visualizer to monitor for unexpected routing paths or anomalous agent behavior. + +== Related Documentation + +* xref:agent-network::af-agent-networks.adoc[Building Agent Networks for Agent Fabric] +* xref:agent-network::af-define-your-agent-network-specification.adoc[Define Your Agent Network Specification] +* xref:agent-visualizer::index.adoc[Agent Visualizer] +* xref:anypoint-code-builder::index.adoc[Anypoint Code Builder] + +== Next Steps + +With agent networks and brokers coordinating your multi-agent processes, you're ready to: + +* xref:af-use-case-policy-enforcement.adoc[Enforce consistent policies across all agents in the network] +* xref:use-case-identity.adoc[Add user identity context to agent interactions] +* xref:use-case-cost-control.adoc[Optimize LLM costs across the network with semantic routing] diff --git a/modules/ROOT/pages/agent-fabric-use-cases.adoc b/modules/ROOT/pages/agent-fabric-use-cases.adoc new file mode 100644 index 000000000..5f20abe51 --- /dev/null +++ b/modules/ROOT/pages/agent-fabric-use-cases.adoc @@ -0,0 +1,85 @@ += Agent Fabric Use Cases + +== Why Agent Fabric + +Enterprises don't struggle to build one agent. They struggle to run many — across Agentforce, Bedrock, custom Mule implementations, and vendor platforms — without losing control of cost, security, and consistency. The usual response is to solve each problem per platform: custom API adapters here, a spend dashboard there, guardrails reimplemented for every team. That work duplicates effort, drifts out of sync, and leaves gaps no one owns. + +Agent Fabric replaces that per-platform sprawl with one governed layer that every agent, API, model, and provider routes through. Expose APIs as agent tools without touching code, enforce the same policies and budgets everywhere, attribute every token to the team that spent it, and contain a rogue agent before it causes damage — all from a single point of control. The result is faster time-to-agent, predictable spend, and governance that holds as you scale from proof of concept to production. + +Each use case below maps a specific business problem to the Agent Fabric capabilities that solve it, so you can start where the pain is and build incrementally. Use the quick-reference table to match your immediate goal to the right capability. + +== Find Your Use Case + +[cols="1,2,1"] +|=== +|If you want to... |Use case |Go to + +|Catalog and manage agents across multiple platforms +|Discover and register agents from a centralized catalog. Know what exists and where before you apply governance. +|//xref:af-use-case-registry.adoc[] + +|Make existing APIs agent-ready without modifying code +|Select which API operations to expose as MCP tools, apply read-only filtering, and control agent access to your APIs. +|xref:af-use-case-mcp-bridge.adoc[] + +|Apply consistent business rules and guardrails across all agents +|Enforce PII policies, rate limits, and compliance rules across all agents through a single gateway rather than reimplementing them per platform. +|//xref:af-use-case-policy-enforcement.adoc[Policy Enforcement] + +|Control token usage and optimize model selection +|Monitor token usage across agents and applications, then route queries to cost-effective models based on complexity. +|xref:af-use-case-cost-management.adoc[Cost Control] + +|Build complex, deterministic workflows with error handling +|Handle multi-step business logic, error scenarios, and unhappy paths that simple agent loops can't reliably address. +|xref:af-use-case-orchestration.adoc[Agentic Orchestration] + +|=== + +== Discover and Govern Agents Across Your Organization + +Use a centralized registry and automated scanners to discover, catalog, and manage agents across your organization. Know what exists and where before you apply governance. For example, a business process outsourcer managing agents across multiple client platforms, or an enterprise security team asked to audit every agent in production, can use Registry and Scanners to get a complete, current picture without manually tracking deployments. + +//For more information, see xref:af-use-case-registry.adoc[]. + +== Make APIs Agent-Ready with MCP Bridge + +Transform your existing API instances into agent-ready tools without modifying code. Select which operations to expose, apply read-only filtering where needed, and control agent access to your APIs. For example, an organization with hundreds of existing Mule APIs can make them callable by agents in minutes through configuration alone, without touching the underlying implementations. + +For more information, see xref:af-use-case-mcp-bridge.adoc[]. + +== Enforce Consistent Policies Across Agent Platforms + +Consistently enforce business rules, PII policies, and guardrails across all agents regardless of platform. Apply policies one time through a gateway rather than reimplementing them for each platform. For example, a regulated industry deploying agents across Agentforce, Bedrock, and custom Mule implementations can enforce the same data privacy and rate limiting rules everywhere through a single Omni Gateway policy. + +//For more information, see xref:af-use-case-policy-enforcement.adoc[]. + +== Monitor Costs and Optimize Model Selection + +Monitor token usage across applications and agents, then optimize costs by routing simple queries to cost-effective models while reserving powerful models for complex tasks. For example, a team running hundreds of daily agent interactions can route routine lookups to a smaller, cheaper model and reserve a frontier model only for tasks that require complex reasoning. + +For more information, see xref:af-use-case-cost-management.adoc[]. + +== Coordinate Multi-Agent Processes with Brokers + +Use agent brokers and agent networks to coordinate task delegation across A2A-compliant agents with guided determinism. Define broker routing logic in Agent Script to handle complex business processes, error scenarios, and unhappy paths that require multiple specialized agents working in sequence. For example, an order management process that involves a research agent, a pricing agent, and an approval agent can be wired together in a single agent network with a broker that routes each step to the right specialist. + +For more information, see xref:af-use-case-orchestration.adoc[]. + + + +== Get Started + +New to Agent Fabric? Start with the xref:agent-fabric-overview.adoc[Agent Fabric Overview] and xref:learning-map-agent-fabric.adoc[Agent Fabric learning map] to understand the fundamentals. + +After you understand the basics, choose a use case to implement. + +== See Also + +* xref:agent-fabric-overview.adoc[] +* xref:learning-map-agent-fabric.adoc[] +* xref:af-use-case-cost-management.adoc[] +* xref:af-use-case-mcp-bridge.adoc[] +* xref:af-use-case-orchestration.adoc[] +* //xref:af-use-case-policy-enforcement.adoc[] +* //xref:af-use-case-registry.adoc[] diff --git a/modules/ROOT/pages/exp-detect-and-contain-rogue-agents.adoc b/modules/ROOT/pages/exp-detect-and-contain-rogue-agents.adoc new file mode 100644 index 000000000..5a853099f --- /dev/null +++ b/modules/ROOT/pages/exp-detect-and-contain-rogue-agents.adoc @@ -0,0 +1,130 @@ += Detect and Contain Rogue Agents +:keywords: quarantine agent, contain rogue agent, agent behavioral drift, tool-call abuse detection, token spend anomaly, flag history, reactivate quarantined agent, unauthorized agent actions, agent security monitoring, model proxy policies, mulesoft + +As AI agents take on more autonomous work across your organization, a single compromised, misconfigured, or malfunctioning agent can quickly leak sensitive data, take unauthorized actions, or run up costs before anyone notices. Monitor agent traffic from a single location in the enhanced MuleSoft experience to detect anomalous behavior and contain a rogue agent before it causes damage. + +Key benefits: + +* Detect anomalous or unauthorized behavior in real time, using built-in detectors and your own custom rules. +* Get notified the moment the platform flags an agent, through email or Slack. +* Review flagged activity before you act. Nothing quarantines automatically without your review. +* Quarantine a rogue agent instantly to stop data leakage, harmful actions, or runaway costs. Quarantine is fully reversible, so you can undo a false positive in seconds. +* Meet compliance requirements. The platform logs every action with full attribution. + +== Before You Begin + +Confirm these prerequisites are in place. + +* You need these API Manager permissions: +** API Creator: Create instances +** View APIs Configuration: View instances +** Edit APIs Configuration: Edit and manage instances +** View API Alerts: View API alerts in a specific environment +** Manage API Alerts: Manage API alerts in a specific environment + +* Register each agent you want to monitor with a unique instance name and ID, and link it to its model proxies. This lets flags, logs, and containment actions target the correct agent. + +* Set up your agents as service identities in your enterprise IdP, and configure the IdP to include each agent's identifier and team claims in the tokens it issues. + +* Apply these policies to your agent tools and model proxy instances: +** xref:gateway::policies-included-jwt-validation.adoc[JWT Token Policy]: validates the agent's identity. +** xref:gateway::policies-outbound-oauth-obo.adoc[OBO (On-Behalf-Of) Token Validation Policy]: identifies both the agent and the end user it's acting for. + +== Enable Agent Monitoring for a Model Proxy + +Enable monitoring on each model proxy whose agent traffic you want to track. + +. Log in to the MuleSoft enhanced experience with an account that has the required permissions. +. In *Portfolio*, select *Model Proxies*. +. Select the model proxy you want to monitor. +. Select the *Policies* tab and then select *+Apply Policy*. +. Select the *Rogue Agent Detection* policy and click *Next*. +. Configure the policy settings by selecting the criteria monitor: ++ +* *Token-spend anomaly* ++ +Flags when an LLM virtual key's token-spend rate over a rolling window is statistically anomalous against the baseline. +* *Tool-call loop* ++ +Flags when an agent repeats near-identical tool calls inside a short window past the configured threshold. +* *Behavioral drift* ++ +Flags when an agent's behavioral fingerprint (tool selection, response shape) drifts persistently from its baseline. +* *Custom detection criteria* ++ +Flags activity that matches a rule you define in plain language, for example, `flag any agent that asks a user for a password`. +. Select *Apply Policy*. + +Repeat these steps for each model proxy to protect. + +== Set Up Notifications for Flagged Agents + +Configure alerts so you're notified as soon as an agent is flagged. + +. Navigate to *Notifications* and select *New Notification*. +. Under *Select Alert Target*, complete the fields: +.. *Environment* ++ +Select the environment to monitor, for example, *Production* or *Sandbox*. +.. *Service Type* ++ +Select *Model Proxy*. +.. *Target Service* ++ +Search for and select the model proxy to receive alerts about. +. Under *Specify Alert Configuration*, select the *Alert Metric* for quarantine activity _(label to be confirmed)_, then set the comparison operator, threshold value, and time window. +.. *Alert Metric* ++ +Select *Policy Violation* and then select the policy. +.. *Alert When* ++ +Select the operator to use in the alert condition, for example, the policy violation occurs greater than 10 times in 5 minutes. +. Under *Set Alert Delivery*: +.. For *Severity*, select *Critical*, *Warning*, or *Info*. +.. In *Alert Name*, enter a descriptive name (four or more characters). +.. In *Delivery Channels*, select *Email*, *Slack*, or both. +. Select *Create Alert*. + +== Review and Quarantine a Flagged Agent + +When an agent is flagged, review its activity and quarantine it if needed. Quarantine stops the agent from acting. + +. Navigate to *Security*. +. On the *Needs Review* tab, find the flagged agent. ++ +The *Security* page shows flagged and reviewed agents from the last 90 days. +. Select *Review* to open the agent's flag history. +. Review the timeline of detection events, then select an action: ++ +* *Quarantine agent* ++ Stops the agent from acting. Confirm at the prompt to complete this action. +* *Clear flags* ++ +Dismisses the flags if the activity was legitimate. Confirm at the prompt to complete this action. +* *Cancel* ++ +Closes the view with no changes. The agent stays in the review queue. + +The audit log records the action with the agent instance ID, the user who performed it, and details such as the timestamp and environment. + +== Reactivate a Quarantined Agent + +After you confirm a quarantined agent is safe to return to service, reactivate it. + +. Navigate to *Security*. +. On the *Reviewed* tab, find the quarantined agent. +. Select *Review* to review the agent's details. +. Select *Restore Model Access*. +. In the confirmation prompt, select *Restore Model Access* to return the agent to service. +. Select *Close* to close the view with no changes. ++ +After reactivation the audit log records the action. + +== See Also + +* xref:exp-alerts-configure-notifications.adoc[] +* xref:exp-instances-add.adoc[] +* xref:exp-home-start.adoc[] +* xref:exp-services-monitoring.adoc[] +* xref:gateway::policies-included-jwt-validation.adoc[] +* xref:gateway::policies-outbound-oauth-obo.adoc[] \ No newline at end of file diff --git a/modules/ROOT/pages/exp-model-wallets-manage.adoc b/modules/ROOT/pages/exp-model-wallets-manage.adoc new file mode 100644 index 000000000..06a577b28 --- /dev/null +++ b/modules/ROOT/pages/exp-model-wallets-manage.adoc @@ -0,0 +1,117 @@ += Managing Model Wallets +:keywords: model wallets, model wallet, budgets, token budget, spend limit, model proxies, access control, jwt claims, cost management, anypoint platform + +A model wallet controls who can use your model proxies and how much they can spend. It combines the authentication credentials that callers present with the budget limits that you set on that access. Use model wallets to grant controlled access to a team, user, or agent, and to cap their token or dollar spend against a provider over a set time window. + +Each wallet has a system-generated client ID and required JWT claims that you define. When a caller connects to a model proxy, the proxy checks the client ID and JWT claims in the request against the wallet's configuration. If the claims match, the proxy authorizes the request and counts the request against the wallet's budgets. + +[[before-you-begin]] +== Before You Begin + +To manage model wallets, you need: + +* An Anypoint Platform account +* At least one configured model proxy. See xref:model-proxy-create-model-proxy.adoc[]. +* A configured identity provider (IdP) that issues JWTs for callers +* A xref:gateway::policies-included-jwt-validation.adoc[JWT Validation] policy applied to each model proxy that the wallet authorizes. Apply this policy on the model proxy in the enhanced experience. See <>. +* These permissions: ++ +-- +** API Manager: API Creator +** API Manager: View APIs Configuration +** API Manager: Edit APIs Configuration +** API Manager: Manage Policies, to apply the JWT Validation policy +-- ++ +For more information, see xref:exp-home-start.adoc#permissions[Enhanced Experience Permissions]. + +Organizations without an IdP configured can't use model wallets. Those organizations continue to use existing access methods. + +[[apply-jwt-validation-to-a-model-proxy]] +== Apply a JWT Validation Policy to a Model Proxy + +A model wallet matches the client ID and JWT claims in a request. The JWT Validation policy on the model proxy validates the token from your IdP first. Apply the policy on each model proxy that callers reach through the wallet. You apply it from the model proxy, not from *Model Wallets*. + +. From *Portfolio* > *Model Proxies*, open the model proxy. +. Select the *Policies* tab, then select *+Apply Policy*. +. Select *JWT Validation* and select *Next*. +. Configure the policy to validate tokens from your IdP, including the JWT origin and the JWKS URL or signing key. For configuration parameters, see xref:gateway::policies-included-jwt-validation.adoc[]. +. Select *Apply Policy*. + +Repeat these steps for each model proxy the wallet authorizes. + +[[create-a-model-wallet]] +== Create a Model Wallet + +. From *Portfolio* > *Model Proxies* > *Model Wallets*, select *New Model Wallet*. +. In *Name*, enter a name for the wallet, for example, `Finance Analytics Bot`. +. (Optional) Enter a *Description* explaining what the wallet is used for. +. Under *Authentication*, copy the system-generated *Client ID*. ++ +This value is read-only. Callers send it as the `X-Client-Id` request header to select the wallet. +. Select *Add Claim* and enter at least one required claim: ++ +-- +** In *Required Claims*, enter one or more claim keys, for example, `group`. +** Enter one or more comma-separated values. +** To add more claims, select *Add Claim* again. +-- ++ +The required claims must match claims in the JWT after the JWT Validation policy on the model proxy validates the token. +. (Optional) Add one or more budgets to cap spend or token usage. See <>. +. Select *Create Model Wallet*. + +[[view-model-wallets]] +== View Model Wallets + +From *Portfolio* > *Model Proxies*, select *Model Wallets*. The wallet list shows: + +* *Name*: The wallet's display name. +* *Description*: What the wallet is used for. +* *Budgets*: How many budgets are attached to the wallet. +* *Last Updated*: When the wallet was last changed. + +To find a specific wallet, use the search box to filter by name or description. + +[[edit-a-model-wallet]] +== Edit a Model Wallet + +. From *Portfolio* > *Model Proxies* > *Model Wallets*, open the wallet and select *Edit*. +. Change the *Name*, *Description*, or *Required Claims*. ++ +The *Client ID* can't be changed. +. Select *Save*. + +Manage budgets separately from the wallet's core details. + +[[add-a-budget-to-a-wallet]] +== Add a Budget to a Wallet + +A budget caps usage for a wallet against a provider over a recurring period. Add a budget when you create or edit a model wallet, or add one from the budgets view in *Governance* > *Cost Management*. + +Add multiple budget limits to a wallet. Assign each model to only one budget limit. + +. In the *Budgets* section, select *Add Budget*. +. Enter the budget details: ++ +-- +** *Provider*: The provider the budget tracks. +** *Period*: *Daily*, *Weekly*, or *Monthly*. +** *Resets on Day of Month*: The day each period the usage counter returns to zero. +** *Metric*: Whether the limit is measured in dollars (*Spend*) or *Tokens*. +** *Spend Limit (USD)*: The limit amount for the selected metric. +-- +. Select *Add Budget*. + +Budgets depend on model costs. For a budget to track usage accurately, configure the cost per token for each model on the *Models* page. Models without configured costs default to zero and don't count against the budget. See xref:exp-models-manage-costs.adoc[]. + +When a wallet reaches its budget, the proxy blocks requests to that provider's models. If a fallback route exists, the request routes to the next model in the proxy's route. + +Budgets are approximate rather than a real-time hard cutoff. The system tracks total token consumption across routes as a governance and cost-awareness tool. Usage can briefly exceed a limit before the system blocks requests. + +== See Also + +* xref:model-proxy-create-model-proxy.adoc[] +* xref:gateway::policies-included-jwt-validation.adoc[] +* xref:exp-governance-view-cost-and-token-usage.adoc[] +* xref:exp-services-register-manually.adoc[] \ No newline at end of file diff --git a/modules/ROOT/pages/exp-models-manage-costs.adoc b/modules/ROOT/pages/exp-models-manage-costs.adoc new file mode 100644 index 000000000..e60fe55e1 --- /dev/null +++ b/modules/ROOT/pages/exp-models-manage-costs.adoc @@ -0,0 +1,46 @@ += Managing Models and Model Costs +:keywords: models, model costs, cost per token, input tokens, output tokens, spend tracking, budgets, model proxies, cost management, anypoint platform + + +The *Models* page lists the AI models available in your organization. Set the cost per token for each model so that token usage translates into accurate spend and budget tracking across model proxies and wallet budgets. + +== Before You Begin + +To manage model costs, you need: + +* An Anypoint Platform account. +* Any of these permissions: ++ +-- +** API Manager: API Creator to create instances +** API Manager: View APIs Configuration to view instances +** API Manager: Edit APIs Configuration to edit instances +-- ++ +For more information, see xref:exp-home-start.adoc#permissions[Enhanced Experience Permissions]. + +== View Models + +. From *Portfolio* > *Model Proxies* > *Model Wallets*, select *Models*. +. Review the list of available models. +. To see details for a model, select its name. ++ +The model detail page shows the model proxies that surface the model and its configured costs. + +[[edit-model-costs]] +== Edit Model Costs + +Model costs drive spend and budget calculations. Configure them for each model you use. + +. On the *Models* page, select a model — for example, `Claude 3.5 Haiku`. +. Select *Edit model costs*. +. Enter the *Cost per 1K input tokens* — for example, `0.001`. +. Enter the *Cost per 1K output tokens* — for example, `0.005`. +. Select *Save changes*. + +Models without configured costs default to zero and don't factor into spend tracking or budget calculations. If a model's costs aren't configured, wallet budgets don't reflect actual usage for that model. See xref:exp-model-wallets-manage.adoc[]. + +== See Also + +* xref:exp-model-wallets-manage.adoc[] +* xref:exp-governance-view-cost-and-token-usage.adoc[] \ No newline at end of file