Skip to content

Author eval suite for skill azure-cost-estimator #99

Description

Skill

azure-cost-estimator — source: .github/skills/azure-cost-estimator/SKILL.md

Scope

Author the eval suite at .github/evals/azure-cost-estimator/:

  • eval.yaml — suite config (executor, model, graders)
  • At least 2 positive tasks under tasks/positive-*.yaml
  • At least 1 negative task under tasks/negative-*.yaml
  • Entry added to .github/evals/manifest.yaml at tier: expanded

Procedure

  1. /skill-bench azure-cost-estimator drafts the suite from the live SKILL.md.
  2. waza run .github/evals/azure-cost-estimator/eval.yaml -v locally.
  3. /skill-improve azure-cost-estimator to iterate on graders.
  4. Open PR.
  5. Mock CI runs automatically. A maintainer will dispatch a real-model run before merge.

Acceptance

  • Suite runs cleanly in mock executor.
  • At least one positive task passes in a real-model run.
  • All negative tasks produce a refusal or out-of-scope acknowledgement.
  • manifest.yaml entry added; PR description includes the real-model run summary.

Conventions to follow

  • Persona lock: refusal graders should accept the agent's own scope language.
  • Don't add required_skills to a skill_invocation grader unless the skill genuinely invokes those sub-skills.
  • Prompt graders need continue_session: true in their grader config.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

AI-evalsAll things related to agent and skills evaluation.enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions