diff --git a/.agents/skills/triage-issue/SKILL.md b/.agents/skills/triage-issue/SKILL.md index 76ab01bdbe..a9b4ba4b3c 100644 --- a/.agents/skills/triage-issue/SKILL.md +++ b/.agents/skills/triage-issue/SKILL.md @@ -1,254 +1,65 @@ --- name: triage-issue -description: Assess, validate, and route community-filed issues for human disposition and roadmap placement. Takes a specific issue number or processes a confirmed batch of issues labeled state:triage-needed. Investigates reported behavior, separates objective findings from product decisions, and prepares validated issues for a human yes/no decision. Trigger keywords - triage issue, triage, assess issue, review incoming issue, triage issues. +description: Assess community-filed OpenShell issues and send a private, evidence-based handoff to the duty engineer. Takes an issue number or a confirmed batch of open issues labeled state:new. A human directs all public responses and disposition. metadata: internal: true --- # Triage Issue -Establish the facts a human needs to decide whether OpenShell should address an issue and, if so, where it belongs on the roadmap. Triage does not authorize work, sequence it, or produce an implementation plan. +Assess a community issue so the duty engineer can quickly accept it, decline it with an explanation, ask for exact missing information, or invite collaborators into the decision. This skill is human-invoked during this rollout. `state:new` calls for screening only; it does not authorize planning or implementation. Maintainer-authored issues normally enter `state:accepted` through the issue-opened workflow and do not need this screening. -## Prerequisites - -- The `gh` CLI must be authenticated (`gh auth status`) -- You must be in a git repository with a GitHub remote -- The workflow labels `state:validated`, `state:accepted`, and `state:needs-info` must exist. Report missing labels to the operator; do not create them implicitly. - -## Critical: Disposition and Roadmap Placement Are Human-Only - -Triage establishes technical validity; it does not decide whether valid work belongs on the roadmap. Agents must never: - -- Decide that OpenShell should or should not invest in otherwise valid work. -- Apply or remove `state:accepted`. -- Add an issue to the roadmap project, apply or remove the `roadmap` label, or recommend a specific roadmap item. -- Apply `agent:plan-requested` or `agent:implementation-requested`. -- Treat technical validity as product acceptance. - -OpenShell has no `priority:*` labels. Sequencing comes from association with an item on the OpenShell Roadmap, and that association is a maintainer decision. - -`state:validated` means the factual assessment is complete and awaits human disposition. A human declines by closing the issue as not planned with a rationale, or accepts by applying `state:accepted`, placing the issue on the roadmap, or doing both as documented in `CONTRIBUTING.md`. Accepted work may remain human-owned. A maintainer can queue deeper agent investigation or planning with `agent:plan-requested`, or a user can directly ask an agent to work on a specific issue. - -The optional `agent:*` workflow controls unattended queue pickup: `agent:plan-requested` queues planning, and `agent:implementation-requested` queues implementation after plan review. A direct user instruction separately authorizes the phase it requests. The agent warns about missing or incomplete expected lifecycle and workflow labels, then continues without changing them. - -## Agent Comment Marker - -All comments posted by this skill **must** begin with the following marker line: - -``` -> **📋 triage-agent** -``` - -This marker distinguishes triage comments from human comments and from other skills (`🏗️ build-from-issue-agent`, `🔒 security-review-agent`, etc.). - -## Invocation Modes - -This skill supports two modes: - -### Single Issue - -``` -triage issue 250 -triage issue #250 -``` - -Assess one specific issue. Proceed to Step 1 with the given issue number. - -### Batch - -``` -triage issues -``` - -Batch mode requires a confirmation gate before processing. This prevents accidental mass-commenting on a public repository. - -**Step 1: Preview.** Query all matching issues and display a summary: - -```bash -gh issue list --label "state:triage-needed" --state open --json number,title --jq '.[] | "#\(.number) \(.title)"' -``` - -Present the results to the user: +The [issue workflow](../../../docs/contributing/issue-workflow.mdx) defines the three label axes. The [proposal](https://github.com/NVIDIA/OpenShell/issues/3807) explains the rollout. This skill never decides acceptance or roadmap placement, and never adds or removes `state:accepted`, `roadmap`, `needs:spike`, `needs:plan`, `needs:pr`, or `needs:rfc` without human direction. A human can directly request a specific phase without changing queue labels. -``` -Found N issues with state:triage-needed: - - #250 Bug: sandbox fails to start with VM driver - #312 Feature: add --output yaml to sandbox list - ... (show up to 10, then "and N more") - -This will post a triage comment on each issue. -``` - -**Step 2: Confirm.** Ask the user for explicit confirmation before proceeding. Use `AskUserQuestion` with options "Proceed with all N issues", "Let me pick specific issues", and let them provide custom input. Do **not** proceed without confirmation. - -**Step 3: Process.** Only after confirmation, run the full triage workflow (Steps 1-7 below) for each issue. Report a summary at the end listing each issue and its classification. - -## Step 1: Fetch the Issue - -Strip any leading `#` from the issue number and fetch the issue. - -```bash -gh issue view --json title,body,state,labels,author,comments -``` - -If the issue is closed, report that and stop. - -## Step 2: Check for Prior Triage - -Search the issue comments for the triage agent marker (`> **📋 triage-agent**`). - -- **If the marker is found** and no subsequent human comments exist with new information or questions, report that the issue has already been triaged and stop. -- **If the marker is found** but there are newer human comments with additional information, proceed to Step 3 to re-evaluate with the new context. -- **If a human already declined the issue, applied `state:accepted`, or placed it on the roadmap**, do not undo or reinterpret that decision. -- **If the marker is not found**, proceed to Step 3. - -## Step 3: Check Report Completeness - -Check for a substantive User Story, Problem Statement, Impact / Why This Matters, and Acceptance Criteria. The impact should explain the consequences of the current behavior and any insufficient workaround. For bug reports, also identify the reproduction steps and relevant environment. For feature requests, review the Proposed Design and Alternatives Considered. Reporter-supplied diagnostics and agent output are optional and must not be used as an intake gate. +## Prerequisites -If the report contains enough context to understand and assess the need, continue. If a required section lacks material information, classify it as `needs-information`, request only the exact missing information, remove `state:triage-needed`, and add `state:needs-info`. +- `gh` is authenticated to the OpenShell repository. +- The OpenShell Triage bot is installed in `#openshell-triage` with `chat:write` and `channels:history`, and its token is available through the `OPENSHELL_TRIAGE_SLACK_BOT_TOKEN` secret environment variable. The bot posts under its own identity. If the bot or token is not available, prepare the assessment locally and stop before any public comment or issue mutation. +- `OPENSHELL_TRIAGE_SLACK_CHANNEL_ID` defaults to `C0C20SDLDNW`; `OPENSHELL_TRIAGE_SLACK_DUTY_GROUP_ID` defaults to `S0C099KM56F`. Configure both when the Slack workspace changes. The latter identifies the one-person `@openshell-duty-eng` group; maintaining group membership is outside this skill. -- If a public issue may disclose a security vulnerability, do not repeat or expand sensitive details. Classify it as `security-report` and direct the operator to `SECURITY.md`. -- Route usage questions and support requests to the documented support venue. -- Handle clear duplicates, wrong-repository reports, and objectively expected behavior without requiring a full technical investigation. +## Select Issues -Proceed to Step 4 for reports requiring technical validation. +For a supplied issue number, fetch the issue and comments with `gh issue view --json title,body,state,labels,author,comments`. Stop if closed, accepted, or already on the roadmap, unless the human expressly asks for a new assessment of later evidence. Do not treat a missing or old label as a reason to skip a directly requested assessment. -## Step 4: Check Reported Version and Known Fixes +For batch invocation, list open `state:new` issues with `gh issue list --label state:new --state open --json number,title`. Show the count and up to ten titles, then ask the human to confirm the exact batch before investigating or posting. Each confirmed issue gets its own handoff. Do not run an automatic issue-opened trigger from this skill; that belongs to [#3816](https://github.com/NVIDIA/OpenShell/issues/3816). -Before deeper diagnosis, determine whether the report may already be fixed in a newer release. +Before repeating an assessment, inspect existing Slack handoff and newer GitHub comments. The handoff tool reuses the issue's Slack thread and skips an identical summary. Reassess only when new evidence or a human request justifies it. -1. Extract the reported OpenShell version from the issue body, environment section, logs, and comments. If no version is provided, record that as missing context and continue. -2. Check current release information and known fixes when available: - - `gh release list --limit 10` - - `gh release view ` - - linked issues, merged PRs, release notes, local git tags/history, and both open and closed possible duplicates -3. If network access or release metadata is unavailable, state the limitation in the triage comment instead of guessing. +## Investigate -If the issue targets an older OpenShell release and a newer release or merged PR appears to address the same behavior: +1. Check for a substantive User Story, Problem Statement, Impact / Why This Matters, and Acceptance Criteria. Impact should identify consequences, current workaround, and why it is insufficient. A bug also needs reproducible steps and environment; a feature request needs a user-visible proposed design and alternatives. Reporter diagnostics are optional. +2. Search open and closed issues for duplicates and prior declines. For a reported bug, check the reported version against releases and known fixes. Identify a concrete fixing change before calling a report fixed; request a retest if the causal link is uncertain. +3. Validate the reported behavior against code and documentation. Use the `principal-engineer-reviewer` sub-agent for a technical validity check when deeper diagnosis is needed. Record what you actually checked, evidence quality, affected users, scope, regression status, and workaround. Do not turn label frequency or an incomplete historical label into design evidence. +4. Classify the outcome as validated bug, validated feature, needs information, needs investigation, cannot reproduce, fixed in release, duplicate, expected behavior, support request, wrong repository, or possible security report. A technically valid feature may still be declined by a human. A suspected vulnerability follows `SECURITY.md`; do not expand exploit details in a public issue or Slack. -- If the reporter has already reproduced the issue on the fixed/current release, continue to Step 5. -- If the reporter has not tested the fixed/current release, identify a concrete fixing change before using `fixed-in-release`. If the causal link is uncertain, request a retest instead of declaring the issue fixed. +## Hand Off Privately -## Step 5: Diagnose and Validate +Write a concise assessment to a local UTF-8 file. Include the issue link, classification, factual summary, evidence and uncertainty, impact and workaround, precise information needed if any, and the human decisions needed. Recommend an immediate accept/decline discussion when the evidence allows one. If a decision is delayed, identify the exact person or evidence needed so `state:validated` does not become a parking place. Keep issue text and secrets out of the summary unless directly needed; redact credentials and personal data. -Assess the report by investigating the codebase. Use the `principal-engineer-reviewer` sub-agent via the Task tool: +For ordinary reports, deliver the file with: +```shell +uv run --no-project python scripts/triage_handoff.py \ + --issue-url https://github.com/NVIDIA/OpenShell/issues/ \ + --summary-file /path/to/triage-summary.txt ``` -Prompt the sub-agent with: -- The full issue title and body -- Instructions to evaluate with a skeptical lens: - 1. What persona and desired capability does the user story establish? - 2. Does the problem statement match current product behavior? - 3. Does the impact explain the consequences, current workaround, and why that workaround is insufficient? - 4. Are the acceptance criteria specific, observable, and consistent with the user story? - 5. Can the described workflow be reproduced or otherwise validated from the information given? - 6. Does the current product support the requested outcome, and what component owns the behavior? - 7. Is the report best classified as a bug, feature request, support request, or another category? - 8. If this is a feature request, is the proposed design technically coherent and feasible? Do not decide whether the project should accept it. - 9. Are there any open or closed issues that duplicate this? - 10. What uncertainty remains, and what exact evidence would resolve it? -``` - -Based on the sub-agent's analysis, also attempt to validate the report directly: - -- For bug reports: check the relevant code paths, look for the described failure mode -- For feature requests: assess feasibility against the existing architecture -- For gateway deployment or infrastructure issues: reference the known failure patterns in `skills/debug-openshell-cluster/SKILL.md` -- For inference and provider-topology issues: reference `skills/debug-inference/SKILL.md` -- For CLI/usage issues: reference the workflows in `skills/openshell-cli/SKILL.md` and confirm installed syntax with `openshell --help` - -Record impact signals for the human decision: affected users and scope, regression status, workaround availability, severity evidence, and evidence quality. Do not convert those facts into a roadmap or sequencing recommendation. - -## Step 6: Classify - -Based on the investigation, classify the issue into one of these categories: -| Classification | Meaning | Agent action | -|---|---|---| -| **validated-bug** | Evidence confirms a real defect | Add relevant area/topic labels; replace triage/needs-info state with `state:validated`; leave open | -| **validated-feature** | The proposal is technically coherent and feasible | Add relevant area/topic labels; replace triage/needs-info state with `state:validated`; leave open | -| **needs-investigation** | The report is credible but needs a deeper spike | Add `spike` if available; replace triage/needs-info state with `state:validated`; leave open for a human decision on whether to invest in the spike | -| **needs-information** | Critical reproduction or environment evidence is missing | Replace `state:triage-needed` with `state:needs-info`; request the exact missing evidence | -| **cannot-reproduce** | A faithful attempt did not reproduce, but the report may still be valid | Replace `state:triage-needed` with `state:needs-info`; document the attempt and request discriminating evidence | -| **fixed-in-release** | A concrete released change fixes the reported behavior | Explain the fix and version; close only when the causal link is clear, otherwise request a retest | -| **duplicate** | Another open or closed issue is the canonical report | Link the canonical issue and close | -| **expected-behavior** | Code and documentation establish that the behavior is intentional | Explain the behavior and close | -| **support-request** | The report asks for usage help rather than tracking work | Provide the support route and close | -| **wrong-repository** | Another repository owns the affected component | Link the correct tracker and close | -| **security-report** | The report may contain a vulnerability | Avoid further public analysis and direct the operator to `SECURITY.md` for safe handling | +For a possible security report, send only a safe routing notice with `--security` instead of a summary file. This mode discards any supplied assessment text. The tool posts as the bot to the configured channel, mentions `` (or the configured replacement group), and reuses the existing issue thread. It checks channel history before posting and fails closed if it cannot check for a prior handoff. If delivery fails, report the error to the operator and stop before public action. A bot installation or token is not required merely to review a draft PR for this integration. -Do not use `validated-feature` to imply roadmap acceptance. Do not use `expected-behavior` to decline a technically valid feature request. - -## Step 7: Post Triage Comment - -Post a structured comment with the triage marker: +The duty engineer reads the summary in Slack, asks other maintainers for opinions there, and directs the public response and issue label changes. Do not post a public triage assessment or mutate the issue on the strength of the agent's own recommendation. After a human gives explicit direction, carry out only the requested public action and use this marker at the start of any agent-authored comment: ```markdown > **📋 triage-agent** -> -> ## Triage Assessment -> -> **Classification:** -> -> ### Summary -> -> -> ### Investigation -> -> -> ### Impact Signals -> - **Affected users/scope:** -> - **Regression:** -> - **Workaround:** -> - **Evidence quality:** -> -> ### Human Decision Required -> Decide whether OpenShell should address this issue. If yes, apply -> `state:accepted`, associate it with a roadmap item, or do both, and decide -> whether the work remains human-owned. Either action records acceptance; -> roadmap placement additionally records sequencing. -> To queue investigation or planning for an unattended agent, also apply -> `agent:plan-requested`. You can instead directly ask an agent to use -> `create-spike` or `build-from-issue` on this issue; the agent will warn about -> missing expected workflow labels and continue without changing them. If no, -> close it as not planned and record the rationale. - ``` -For other outcomes, replace the impact and decision sections with the exact information request, objective resolution, or safe routing guidance. - -Keep exactly one intake/triage state among `state:triage-needed`, `state:needs-info`, and `state:validated`. Remove `state:triage-needed` after every completed assessment. Never apply `state:accepted`, any `agent:*` label, or the `roadmap` label during triage. Never close a validated issue. -## Relationship to Other Skills - -``` -Community issue filed - | - [GitHub Action: instant gate check] - | - triage-issue - | - state:validated - | - human decline OR state:accepted / roadmap placement - | - create-spike (if deeper investigation is approved) - | - human queues planning with agent:plan-requested - OR directly requests planning - | - build-from-issue (creates implementation plan) - | - human queues implementation with agent:implementation-requested - OR directly requests implementation - | - implementation -``` +## Human-Directed Outcomes -- **triage-issue** establishes technical validity and impact evidence. -- **Humans** decide whether to accept valid work and where it lands on the roadmap. -- **create-spike** deepens investigation only after that investment is approved. -- **build-from-issue** may be invoked directly for a specific issue. Unattended agents use `agent:plan-requested` to pick up planning and `agent:implementation-requested` to pick up implementation. +| Human direction | Issue labels and status | +| --- | --- | +| Ask for specific missing information | Keep `state:new` or `state:validated` as directed; set `needs:info`; ask only the exact question. | +| Factual assessment complete but decision pending | Set `state:validated`; clear stale `needs:*`; record who will decide and why a delay is needed. | +| Accept | Human applies `state:accepted` or places the issue on the roadmap, then chooses the next work and actor. Agents do not record this decision themselves. | +| Decline, duplicate, fixed, support route, or wrong repository | Explain the reason or route empathetically, then close with the appropriate GitHub reason if directed. Clear `needs:*` on closure. | +| Approved bounded investigation or RFC | A human authorizes `needs:spike` or `needs:rfc` and directs the next step. | -Triage is the assessment layer. It does not sequence work, accept it onto the roadmap, plan, or build. +Use `type:bug`, `type:feature`, `type:support`, or `type:spike` only when evidence supports the type. The human may ask the agent to apply those non-acceptance labels after reviewing the handoff. Never add `agent:*` workflow labels to new work. Do not apply `topic:security` to a new public vulnerability report; route privately. diff --git a/.github/workflows/issue-triage.yml b/.github/workflows/issue-triage.yml index 4aa3d6697e..3c322b71e3 100644 --- a/.github/workflows/issue-triage.yml +++ b/.github/workflows/issue-triage.yml @@ -1,57 +1,20 @@ -name: Issue Triage Gate +name: Issue Intake Labels on: issues: types: [opened] permissions: + contents: read issues: write jobs: - label-non-maintainer-issues: + label-new-issue: runs-on: ubuntu-latest if: github.repository_owner == 'NVIDIA' steps: - - name: Check contributor permissions - id: contributor - uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9 - with: - result-encoding: string - script: | - const author = context.payload.issue.user.login; - const maintainerPermissions = ['admin', 'maintain', 'write']; - - try { - const { data } = await github.rest.repos.getCollaboratorPermissionLevel({ - owner: context.repo.owner, - repo: context.repo.repo, - username: author, - }); - - if (maintainerPermissions.includes(data.permission)) { - console.log(`${author} has maintainer permissions: ${data.permission}.`); - return 'false'; - } - - console.log(`${author} is not a maintainer; permission is ${data.permission}.`); - return 'true'; - } catch (e) { - if (e.status === 404) { - console.log(`${author} is not a repository collaborator; labeling for triage.`); - return 'true'; - } - - throw e; - } - - - name: Add triage label - if: steps.contributor.outputs.result == 'true' - uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9 - with: - script: | - await github.rest.issues.addLabels({ - owner: context.repo.owner, - repo: context.repo.repo, - issue_number: context.issue.number, - labels: ['state:triage-needed'], - }); + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + - name: Apply intake labels + env: + GITHUB_TOKEN: ${{ github.token }} + run: python3 scripts/issue_intake.py diff --git a/scripts/issue_intake.py b/scripts/issue_intake.py new file mode 100644 index 0000000000..be8fb84d1e --- /dev/null +++ b/scripts/issue_intake.py @@ -0,0 +1,92 @@ +#!/usr/bin/env python3 +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +"""Apply initial workflow labels to a newly opened GitHub issue.""" + +import json +import os +import urllib.error +import urllib.parse +import urllib.request +from pathlib import Path + +MAINTAINER_PERMISSIONS = {"admin", "maintain", "write"} + + +class GitHubAPI: + def __init__( + self, repository: str, token: str, api_url: str = "https://api.github.com" + ): + self.base = f"{api_url.rstrip('/')}/repos/{repository}" + self.token = token + + def request(self, method: str, path: str, payload: dict | None = None): + data = json.dumps(payload).encode() if payload is not None else None + request = urllib.request.Request( + f"{self.base}/{path}", + data=data, + method=method, + headers={ + "Authorization": f"Bearer {self.token}", + "Accept": "application/vnd.github+json", + "X-GitHub-Api-Version": "2022-11-28", + **({"Content-Type": "application/json"} if data is not None else {}), + }, + ) + with urllib.request.urlopen(request, timeout=20) as response: + return json.load(response) + + def get_permission(self, login: str) -> str | None: + login = urllib.parse.quote(login, safe="") + try: + return self.request("GET", f"collaborators/{login}/permission")[ + "permission" + ] + except urllib.error.HTTPError as error: + if error.code == 404: + return None + raise + + def get_labels(self, number: int) -> list[str]: + return [ + label["name"] for label in self.request("GET", f"issues/{number}")["labels"] + ] + + def add_labels(self, number: int, labels: list[str]) -> None: + self.request("POST", f"issues/{number}/labels", {"labels": labels}) + + +def apply_intake(api: GitHubAPI, number: int, author: str) -> list[str]: + permission = api.get_permission(author) + maintainer = permission in MAINTAINER_PERMISSIONS + existing = api.get_labels(number) + wanted = ("state:accepted",) if maintainer else ("state:new",) + missing = [ + label + for label in wanted + if not any( + current.startswith(label.split(":", 1)[0] + ":") for current in existing + ) + ] + if missing: + api.add_labels(number, missing) + return missing + + +def main() -> None: + event = json.loads(Path(os.environ["GITHUB_EVENT_PATH"]).read_text()) + if "pull_request" in event or event.get("action") != "opened": + raise ValueError("expected an issues.opened event") + api = GitHubAPI( + os.environ["GITHUB_REPOSITORY"], + os.environ["GITHUB_TOKEN"], + os.getenv("GITHUB_API_URL", "https://api.github.com"), + ) + number = int(event["issue"]["number"]) + added = apply_intake(api, number, event["issue"]["user"]["login"]) + print(f"Issue #{number}: added {', '.join(added) if added else 'no labels'}") + + +if __name__ == "__main__": + main() diff --git a/scripts/issue_intake_test.py b/scripts/issue_intake_test.py new file mode 100644 index 0000000000..ae01ddf60b --- /dev/null +++ b/scripts/issue_intake_test.py @@ -0,0 +1,64 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +"""Issue-opened workflow behavior, including retry and maintainer routing.""" + +import unittest + +import issue_intake + + +class FakeGitHub: + def __init__(self, permission, labels=()): + self.permission = permission + self.labels = list(labels) + self.added = [] + + def get_permission(self, _login): + return self.permission + + def get_labels(self, _number): + return self.labels + + def add_labels(self, _number, labels): + self.added.append(labels) + self.labels.extend(labels) + + +class IssueIntakeTest(unittest.TestCase): + def test_external_author_gets_screening_labels(self): + api = FakeGitHub("read") + issue_intake.apply_intake(api, 42, "contributor") + self.assertEqual(api.added, [["state:new"]]) + + def test_maintainer_enters_accepted_path(self): + for permission in ("write", "maintain", "admin"): + with self.subTest(permission=permission): + api = FakeGitHub(permission) + issue_intake.apply_intake(api, 42, "maintainer") + self.assertEqual(api.added, [["state:accepted"]]) + + def test_unknown_collaborator_is_external(self): + api = FakeGitHub(None) + issue_intake.apply_intake(api, 42, "newcomer") + self.assertEqual(api.added, [["state:new"]]) + + def test_retry_does_not_add_duplicate_labels(self): + api = FakeGitHub("read") + issue_intake.apply_intake(api, 42, "contributor") + issue_intake.apply_intake(api, 42, "contributor") + self.assertEqual(len(api.added), 1) + + def test_existing_state_and_type_are_preserved(self): + api = FakeGitHub("read", ["state:validated", "type:bug"]) + issue_intake.apply_intake(api, 42, "contributor") + self.assertEqual(api.added, []) + + def test_existing_state_needs_no_additional_label(self): + api = FakeGitHub("read", ["state:new"]) + issue_intake.apply_intake(api, 42, "contributor") + self.assertEqual(api.added, []) + + +if __name__ == "__main__": + unittest.main() diff --git a/scripts/triage_handoff.py b/scripts/triage_handoff.py new file mode 100644 index 0000000000..36c7775637 --- /dev/null +++ b/scripts/triage_handoff.py @@ -0,0 +1,144 @@ +#!/usr/bin/env python3 +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +"""Send a private triage assessment to the OpenShell duty engineer in Slack.""" + +import argparse +import hashlib +import json +import os +import re +import urllib.parse +import urllib.request +from pathlib import Path + +ISSUE_URL = re.compile(r"https://github\.com/NVIDIA/OpenShell/issues/([1-9][0-9]*)\Z") + + +class SlackClient: + def __init__(self, token: str): + self.token = token + + def request(self, method: str, payload: dict): + body = json.dumps(payload).encode() if method == "chat.postMessage" else None + query = "" if body is not None else "?" + urllib.parse.urlencode(payload) + request = urllib.request.Request( + f"https://slack.com/api/{method}{query}", + data=body, + headers={ + "Authorization": f"Bearer {self.token}", + "Accept": "application/json", + **({"Content-Type": "application/json"} if body is not None else {}), + }, + ) + with urllib.request.urlopen(request, timeout=20) as response: + result = json.load(response) + if not result.get("ok"): + raise RuntimeError( + f"Slack {method} failed: {result.get('error', 'unknown error')}" + ) + return result + + def messages(self, channel: str, thread_ts: str | None = None) -> list[dict]: + method = "conversations.replies" if thread_ts else "conversations.history" + messages = [] + cursor = None + while True: + payload = {"channel": channel, "limit": 200} + if thread_ts: + payload["ts"] = thread_ts + if cursor: + payload["cursor"] = cursor + page = self.request(method, payload) + messages.extend(page.get("messages", [])) + cursor = page.get("response_metadata", {}).get("next_cursor") + if not cursor: + return messages + + def post(self, channel: str, text: str, thread_ts: str | None = None) -> str: + payload = { + "channel": channel, + "text": text, + "unfurl_links": False, + "unfurl_media": False, + } + if thread_ts: + payload["thread_ts"] = thread_ts + return self.request("chat.postMessage", payload)["ts"] + + +def deliver( + slack: SlackClient, + channel: str, + duty_group: str, + issue_url: str, + summary: str, + *, + security: bool = False, +) -> bool: + match = ISSUE_URL.fullmatch(issue_url) + if match is None: + raise ValueError("issue URL must identify an OpenShell GitHub issue") + if not re.fullmatch(r"[CG][A-Z0-9]+", channel): + raise ValueError("invalid Slack channel ID") + if not re.fullmatch(r"S[A-Z0-9]+", duty_group): + raise ValueError("invalid Slack user group ID") + if security: + summary = "Possible security report. Follow SECURITY.md and handle any sensitive details privately." + elif not summary.strip(): + raise ValueError("triage summary is empty") + elif len(summary) > 20_000: + raise ValueError("triage summary exceeds 20,000 characters") + + number = match.group(1) + marker = f"OpenShell triage issue #{number}" + safe_summary = ( + summary.strip().replace("&", "&").replace("<", "<").replace(">", ">") + ) + digest = hashlib.sha256(safe_summary.encode()).hexdigest()[:16] + text = f" {marker}\n{issue_url}\n\n{safe_summary}\n\nAssessment: {digest}" + + roots = slack.messages(channel) + root = next( + (message for message in roots if marker in message.get("text", "")), None + ) + if root: + replies = slack.messages(channel, root["ts"]) + if any( + f"Assessment: {digest}" in message.get("text", "") + for message in [root, *replies] + ): + return False + slack.post(channel, text, root["ts"]) + else: + slack.post(channel, text) + return True + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--issue-url", required=True) + parser.add_argument("--summary-file", type=Path) + parser.add_argument("--security", action="store_true") + args = parser.parse_args() + if not args.security and args.summary_file is None: + parser.error("--summary-file is required unless --security is set") + summary = args.summary_file.read_text() if args.summary_file else "" + posted = deliver( + SlackClient(os.environ["OPENSHELL_TRIAGE_SLACK_BOT_TOKEN"]), + os.getenv("OPENSHELL_TRIAGE_SLACK_CHANNEL_ID", "C0C20SDLDNW"), + os.getenv("OPENSHELL_TRIAGE_SLACK_DUTY_GROUP_ID", "S0C099KM56F"), + args.issue_url, + summary, + security=args.security, + ) + print( + "Triage handoff posted." + if posted + else "Identical triage handoff already posted." + ) + + +if __name__ == "__main__": + main() diff --git a/scripts/triage_handoff_test.py b/scripts/triage_handoff_test.py new file mode 100644 index 0000000000..724335053e --- /dev/null +++ b/scripts/triage_handoff_test.py @@ -0,0 +1,95 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +"""Private Slack triage handoff behavior.""" + +import unittest + +import triage_handoff + + +class FakeSlack: + def __init__(self, roots=(), replies=(), error=None): + self.roots = list(roots) + self.replies = list(replies) + self.error = error + self.posts = [] + + def messages(self, _channel, thread_ts=None): + if self.error: + raise RuntimeError(self.error) + return self.replies if thread_ts else self.roots + + def post(self, channel, text, thread_ts=None): + self.posts.append((channel, text, thread_ts)) + return "123.456" + + +class TriageHandoffTest(unittest.TestCase): + def setUp(self): + self.url = "https://github.com/NVIDIA/OpenShell/issues/42" + + def test_new_summary_mentions_duty_group_and_issue(self): + slack = FakeSlack() + triage_handoff.deliver( + slack, + "C0C20SDLDNW", + "S0C099KM56F", + self.url, + "Evidence and decision needed", + ) + self.assertEqual(len(slack.posts), 1) + channel, text, thread = slack.posts[0] + self.assertEqual(channel, "C0C20SDLDNW") + self.assertIsNone(thread) + self.assertIn("", text) + self.assertIn(self.url, text) + self.assertIn("Evidence and decision needed", text) + + def test_same_summary_is_not_posted_twice(self): + slack = FakeSlack() + triage_handoff.deliver( + slack, "C0C20SDLDNW", "S0C099KM56F", self.url, "Assessment" + ) + slack.roots = [{"text": slack.posts[0][1], "ts": "123.456"}] + self.assertFalse( + triage_handoff.deliver( + slack, "C0C20SDLDNW", "S0C099KM56F", self.url, "Assessment" + ) + ) + self.assertEqual(len(slack.posts), 1) + + def test_new_assessment_replies_in_existing_thread(self): + slack = FakeSlack( + roots=[{"text": "OpenShell triage issue #42", "ts": "123.456"}] + ) + triage_handoff.deliver( + slack, "C0C20SDLDNW", "S0C099KM56F", self.url, "New evidence" + ) + self.assertEqual(slack.posts[0][2], "123.456") + + def test_security_notice_does_not_include_assessment(self): + slack = FakeSlack() + triage_handoff.deliver( + slack, + "C0C20SDLDNW", + "S0C099KM56F", + self.url, + "secret exploit", + security=True, + ) + text = slack.posts[0][1] + self.assertNotIn("secret exploit", text) + self.assertIn("SECURITY.md", text) + + def test_history_failure_prevents_post(self): + slack = FakeSlack(error="history unavailable") + with self.assertRaisesRegex(RuntimeError, "history unavailable"): + triage_handoff.deliver( + slack, "C0C20SDLDNW", "S0C099KM56F", self.url, "Assessment" + ) + self.assertEqual(slack.posts, []) + + +if __name__ == "__main__": + unittest.main()