Skip to content

test: add Azure Principal Architect agent eval suite - #326

Open
759132989-crypto wants to merge 3 commits into
Azure:mainfrom
759132989-crypto:759132989-crypto-patch-1
Open

759132989-crypto wants to merge 3 commits into
Azure:mainfrom
759132989-crypto:759132989-crypto-patch-1

Conversation

@759132989-crypto

Copy link
Copy Markdown

test: add Azure Principal Architect agent eval suite

Addresses #95.

Changes

  • Add an auto-discovered eval suite under .github/evals/agents/azure-principal-architect/; no manifest changes.
  • Mirror the current canonical azure-principal-architect.agent.md without changing the production agent.
  • Add two positive tasks: a production architecture review across the five WAF pillars, and an Azure Functions Consumption/Premium trade-off analysis.
  • Add two negative tasks: unrelated marketing content, and a request to deploy or modify live Azure resources.
  • Use continuation-session prompt graders for response quality and scope handling.
  • Suppress the implicit VS Code/SDK cross-taxonomy tool allowlist following the repository's guidance. The deployment-negative task separately rejects the SDK tool names bash, edit, create, sql, and task in recorded calls.

Validation

  • git diff --check HEAD^ HEAD: passed.
  • Canonical/mirrored agent byte comparison: passed.
  • node scripts/validate-structure.js: passed, 0 errors and 6 warnings in existing, unchanged source files. Provisioned gray-matter@4.0.3 locally for this check; no dependency manifests or lockfiles changed. Website build not run.
  • Waza v0.37.0 mock smoke run on the revised suite: all 4 tasks loaded and ran, 0 runtime errors; 0 succeeded, 4 failed, exit code 1. The prompt graders returned score 0 with mock responses and no grade callbacks. This is a loading/execution check, not a passing agent-quality evaluation.
  • The mock run used a temporary copy with executor: mock and one trial per task. The submitted suite retains copilot-sdk and two trials per task.

Review boundaries

The actual model evaluation has not been run. Maintainer CI/review is needed to establish response quality, refusal behavior, and documentation-grounded correctness.

The negative tool grader checks only the listed SDK tool names in the available transcript. It is not a sandbox or proof that every possible external side effect is prevented. Mock results contain no real tool-use behavior.

This contribution was prepared with AI assistance. No claim of production adoption or measured business impact is made.

@759132989-crypto

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant