Skip to content

test(uat): validate the published AI workflow with independent developers #692

Description

@kunaldhongade

Parent

Human acceptance gate for #675.

Problem

Automated child-repository tests prove that commands execute. Agent efficacy benchmarks prove whether treatment improves model outcomes. Neither proves that real developers can install CodeDecay, understand its trust model, connect their own agent, act on the output, and complete a repair safely without maintainer knowledge.

Goal

Run reproducible task-based usability and end-user acceptance testing against the packed or published npm package with independent developers.

Participant and environment requirements

  • Before a milestone release, test with at least three participants who did not implement the feature.
  • Include an AI-assisted individual developer, an experienced software engineer, and a team/DevOps or platform-oriented user when available.
  • Use fresh isolated repositories and normal documented installation paths.
  • Do not preconfigure hidden maintainer state or explain internal package architecture.

Required tasks

Participants must:

  1. Install the published package and discover the primary AI workflow.
  2. Supply a requirement with an ambiguity and resolve it.
  3. Connect or select a user-owned agent/provider, or use deterministic no-model mode.
  4. Run preflight/session guidance.
  5. Review evidence, memory, AI suggestions, and limitations without confusing their trust levels.
  6. Let the agent create an intentionally incomplete implementation.
  7. Run investigation and approved verification.
  8. Repair the planted defect and revalidate the current tree.
  9. Handle one blocked unsafe action correctly.
  10. Find the final requirement verdict and remaining uncertainty.

Acceptance criteria

  • Publish a versioned UAT kit with participant script, repository fixtures, planted defects, clean decoys, task IDs, observer rubric, consent/privacy notes, and result schema.
  • Run against the packed or published npm package, never workspace-only imports.
  • Record task completion, time, attempts, clarification requests, unsafe actions, mistaken trust interpretation, command failures, abandonment, and participant role/environment.
  • Record whether the participant can explain what is deterministic evidence, runtime/tool proof, memory, AI suggestion, unverified, needs-human, and verified.
  • No participant mistakes agent text for proof or an unverified status for merge-safe.
  • Every participant can discover the next action and rerun command without reading source code.
  • The primary workflow completes without maintainer intervention for the agreed success threshold.
  • Installation, authentication/provider, package-manager, terminal, MCP, and documentation friction are tracked separately from analysis quality.
  • Capture sanitized command/artifact references without collecting repository secrets, raw private source, provider credentials, or hidden telemetry.
  • Convert every release-blocking usability failure into a linked focused issue.
  • Publish an anonymized Markdown summary and machine-readable result artifact.
  • Repeat the critical workflow after material CLI/session/report changes.
  • Add a deterministic scripted UAT smoke in CI for workflow drift, while preserving real human sessions as milestone evidence.
  • Do not close this issue using only an agent pretending to be a human participant.

UAT scenarios

  • UAT-HUMAN-1: Fresh install and first useful result.
  • UAT-HUMAN-2: Ambiguous requirement and clarification.
  • UAT-HUMAN-3: Weak test exists but does not prove the production path.
  • UAT-HUMAN-4: Approved behavioral experiment finds the planted defect.
  • UAT-HUMAN-5: Agent repair plus current-tree revalidation.
  • UAT-HUMAN-6: Unsafe command or external target is blocked and correctly understood.
  • UAT-HUMAN-7: Clean decoy produces no unnecessary repair.
  • UAT-HUMAN-8: Participant explains final evidence and limitations accurately.

Agentic QA

Use a real user-owned agent during at least one participant session, but keep human comprehension and safe task completion as the oracle. Also run the deterministic fake-agent workflow to separate product drift from model variance.

Safety and privacy

  • Explicit participant consent and no hidden recording or telemetry.
  • Use synthetic repositories with no customer data.
  • Never collect provider keys, secret values, private source, or hidden chain-of-thought.
  • Sanitize artifacts before publication.

Dependencies

Validation

  • completed participant records and anonymized summary
  • published-package install and workflow evidence
  • deterministic UAT smoke
  • linked issues for unresolved blockers

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: agentAgent task bundle and user-owned agent workflowarea: dev-experienceContributor local setup and agent workflowarea: mcpModel Context Protocol integrationarea: packagingnpm package metadata or published contentscompetitive-parityWork driven by gaps versus adjacent productsexamplesExample projects, fixtures, or adoption demospriority: criticalTop-priority work required for product viabilitytype: testTest coverage, fixtures, or verification improvements

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions