Skip to content
This repository was archived by the owner on Sep 28, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/CI.yml
Original file line number Diff line number Diff line change
Expand Up @@ -211,9 +211,9 @@ jobs:
location=${{ env.AZURE_LOCATION }} \
azureAiServiceLocation=${{ env.AZURE_LOCATION }} \
deploymentType="GlobalStandard" \
gptModelName="gpt-4.1-mini" \
gptModelName="gpt-5-mini" \
gptDeploymentCapacity=${{ env.GPT_CAPACITY }} \
gptModelVersion="2025-04-14" \
gptModelVersion="2025-08-07" \
embeddingModelName="text-embedding-3-large" \
embeddingDeploymentCapacity=${{ env.TEXT_EMBEDDING_CAPACITY }} \
embeddingModelVersion="1" \
Expand Down
2 changes: 1 addition & 1 deletion Deployment/checkquota.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ Write-Host "✅ Azure subscription set successfully."

# Define models and their minimum required capacities
$MIN_CAPACITY = @{
"OpenAI.GlobalStandard.gpt4.1-mini" = $GPT_MIN_CAPACITY
"OpenAI.GlobalStandard.gpt-5-mini" = $GPT_MIN_CAPACITY
"OpenAI.GlobalStandard.text-embedding-3-large" = $TEXT_EMBEDDING_MIN_CAPACITY
}

Expand Down
2 changes: 1 addition & 1 deletion Deployment/quota_check_params.sh
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ log_verbose() {
}

# Default Models and Capacities (Comma-separated in "model:capacity" format)
DEFAULT_MODEL_CAPACITY="gpt4.1-mini:150,text-embedding-3-large:100"
DEFAULT_MODEL_CAPACITY="gpt-5-mini:150,text-embedding-3-large:100"

# Convert the comma-separated string into an array
IFS=',' read -r -a MODEL_CAPACITY_PAIRS <<< "$DEFAULT_MODEL_CAPACITY"
Expand Down
2 changes: 1 addition & 1 deletion docs/AVMPostDeploymentGuide.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,7 @@ Upon successful completion, you'll see a success message with important informat

| Model Name | Recommended TPM | Minimum TPM |
|------------------------|----------------|-------------|
| gpt-4.1-mini | 100K TPM | 10K TPM |
| gpt-5-mini | 100K TPM | 10K TPM |
| text-embedding-3-large | 200K TPM | 50K TPM |

> **⚠️ Warning**: Insufficient quota will cause failures during document upload and processing. Ensure adequate capacity before proceeding.
Expand Down
4 changes: 2 additions & 2 deletions docs/CustomizingAzdParameters.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,9 @@ By default this template will use the environment name as the prefix to prevent
| `AZURE_LOCATION` | string | `<User selects during deployment>` | Location of the Azure resources. Controls where the infrastructure will be deployed. |
| `AZURE_ENV_AI_SERVICE_LOCATION` | string | `<User selects during deployment>` | Location for Azure OpenAI resources. Can be different from AZURE_LOCATION for optimized AI service placement. |
| `AZURE_ENV_MODEL_DEPLOYMENT_TYPE` | string | `GlobalStandard` | Defines the deployment type for the AI model (e.g., Standard, GlobalStandard). |
| `AZURE_ENV_GPT_MODEL_NAME` | string | `gpt-4.1-mini` | Specifies the name of the GPT model to be deployed. |
| `AZURE_ENV_GPT_MODEL_NAME` | string | `gpt-5-mini` | Specifies the name of the GPT model to be deployed. |
| `AZURE_ENV_GPT_MODEL_CAPACITY` | int | `100` | Sets the GPT model capacity (in thousands of tokens per minute). |
| `AZURE_ENV_GPT_MODEL_VERSION` | string | `2025-04-14` | Version of the GPT model to be used for deployment. |
| `AZURE_ENV_GPT_MODEL_VERSION` | string | `2025-08-07` | Version of the GPT model to be used for deployment. |
| `AZURE_ENV_EMBEDDING_MODEL_NAME` | string | `text-embedding-3-large` | Sets the name of the embedding model to use. |
| `AZURE_ENV_EMBEDDING_MODEL_VERSION` | string | `1` | Version of the embedding model to be used for deployment. |
| `AZURE_ENV_EMBEDDING_DEPLOYMENT_CAPACITY` | int | `100` | Capacity for embedding model deployment (in thousands of tokens per minute). |
Expand Down
10 changes: 5 additions & 5 deletions docs/QuotaCheck.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
## Check Quota Availability Before Deployment

Before deploying the accelerator, **ensure sufficient quota availability** for the required model.
> **For Global Standard | gpt4.1-mini - increase the capacity to at least 150K tokens for optimal performance.**
> **For Global Standard | gpt-5-mini - increase the capacity to at least 150K tokens for optimal performance.**

### Login if you have not done so already
```
Expand All @@ -11,7 +11,7 @@ azd auth login

### 📌 Default Models & Capacities:
```
gpt4.1-mini:150, text-embedding-3-large:100
gpt-5-mini:150, text-embedding-3-large:100
```
### 📌 Default Regions:
```
Expand All @@ -37,19 +37,19 @@ eastus, uksouth, eastus2, northcentralus, swedencentral, westus, westus2, southc
```
✔️ Check specific model(s) in default regions:
```
./quota_check_params.sh --models gpt4.1-mini:150,text-embedding-3-large:100
./quota_check_params.sh --models gpt-5-mini:150,text-embedding-3-large:100
```
✔️ Check default models in specific region(s):
```
./quota_check_params.sh --regions eastus,westus
```
✔️ Passing Both models and regions:
```
./quota_check_params.sh --models gpt4.1-mini:150 --regions eastus,westus2
./quota_check_params.sh --models gpt-5-mini:150 --regions eastus,westus2
```
✔️ All parameters combined:
```
./quota_check_params.sh --models gpt4.1-mini:150,text-embedding-3-large:100 --regions eastus,westus --verbose
./quota_check_params.sh --models gpt-5-mini:150,text-embedding-3-large:100 --regions eastus,westus --verbose
```

### **Sample Output**
Expand Down
11 changes: 7 additions & 4 deletions infra/main.bicep
Original file line number Diff line number Diff line change
Expand Up @@ -34,12 +34,12 @@ param deploymentType string = 'GlobalStandard'
@minLength(1)
@description('Optional. Name of the GPT model to deploy:')
@allowed([
'gpt-4.1-mini'
'gpt-5-mini'
])
param gptModelName string = 'gpt-4.1-mini'
param gptModelName string = 'gpt-5-mini'
Comment thread
Priyanka2-Microsoft marked this conversation as resolved.

@description('Optional. Version of the GPT model to deploy.')
param gptModelVersion string = '2025-04-14'
param gptModelVersion string = '2025-08-07'

@description('Optional. Capacity of the GPT model deployment:')
@minValue(10)
Expand Down Expand Up @@ -95,7 +95,7 @@ param enableScalability bool = false
azd: {
type: 'location'
usageName: [
'OpenAI.GlobalStandard.gpt4.1-mini,150'
'OpenAI.GlobalStandard.gpt-5-mini,150'
'OpenAI.GlobalStandard.text-embedding-3-large,100'
]
}
Expand Down Expand Up @@ -984,6 +984,9 @@ module managedCluster 'br/public:avm/res/container-service/managed-cluster:0.13.
minCount: 1
maxCount: 2

// Disable zonal placement; not all regions/subscriptions expose availability zones for the agent pool SKU
availabilityZones: []

// WAF aligned configuration for Private Networking
enableAutoScaling: true
scaleSetEvictionPolicy: 'Delete'
Expand Down
13 changes: 7 additions & 6 deletions infra/main.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
"_generator": {
"name": "bicep",
"version": "0.43.8.12551",
"templateHash": "9930092765515882543"
"templateHash": "5797490851931362105"
}
},
"parameters": {
Expand Down Expand Up @@ -48,9 +48,9 @@
},
"gptModelName": {
"type": "string",
"defaultValue": "gpt-4.1-mini",
"defaultValue": "gpt-5-mini",
"allowedValues": [
"gpt-4.1-mini"
"gpt-5-mini"
],
"minLength": 1,
"metadata": {
Expand All @@ -59,7 +59,7 @@
},
"gptModelVersion": {
"type": "string",
"defaultValue": "2025-04-14",
"defaultValue": "2025-08-07",
"metadata": {
"description": "Optional. Version of the GPT model to deploy."
}
Expand Down Expand Up @@ -177,7 +177,7 @@
"azd": {
"type": "location",
"usageName": [
"OpenAI.GlobalStandard.gpt4.1-mini,150",
"OpenAI.GlobalStandard.gpt-5-mini,150",
"OpenAI.GlobalStandard.text-embedding-3-large,100"
]
},
Expand Down Expand Up @@ -43817,8 +43817,8 @@
}
},
"dependsOn": [
"[format('avmPrivateDnsZones[{0}]', variables('dnsZoneIndex').storageBlob)]",
"[format('avmPrivateDnsZones[{0}]', variables('dnsZoneIndex').storageQueue)]",
"[format('avmPrivateDnsZones[{0}]', variables('dnsZoneIndex').storageBlob)]",
"userAssignedIdentity",
"virtualNetwork"
]
Expand Down Expand Up @@ -52360,6 +52360,7 @@
"type": "VirtualMachineScaleSets",
"minCount": 1,
"maxCount": 2,
"availabilityZones": [],
"enableAutoScaling": true,
"scaleSetEvictionPolicy": "Delete",
"scaleSetPriority": "Regular",
Expand Down
Loading