Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
# Archive Report: hackathon-analysis

**Change**: hackathon-analysis (Roadmap change 3 of 4)
**Archived**: 2026-09-29
**Archived to**: `openspec/changes/archive/2026-09-29-hackathon-analysis/`
**Mode**: openspec

## Verdict

PASS WITH WARNINGS. Archive successful with all operator tasks verified and resolved.

## Tasks Completion Status

All 11 phases complete:
- Phases 1-10: Implemented and merged (PRs #20-#37).
- Phase 11: Operator rollout manual tasks verified and complete:
- 11.1: Queue created on Cloudflare (2026-09-28)
- 11.2: Workers AI and Browser Rendering bindings confirmed working in production (2026-09-29)
- 11.3: Bot promoted to group ADMINISTRATOR with "Pin Messages" right. Verified pinning works (topic 262, pinned_message_id: 271)
- 11.4: Production models verified and set (Free-plan: glm-4.7-flash / qwen3-30b-a3b-fp8)
- 11.5: Smoke test passed on production. General chat: bnb-ai-hack and tokenized-stocks. Topic: /hackathon bnb-hack-tokenized-stocks-edition with successful link and pin.

## Specs Synced to Main

Three new spec domains created from delta specs:

| Domain | Action | Location |
|--------|--------|----------|
| hackathon-analysis | Created | openspec/specs/hackathon-analysis/spec.md |
| page-fetch | Created | openspec/specs/page-fetch/spec.md |
| llm-extraction | Created | openspec/specs/llm-extraction/spec.md |

All three specs are complete specifications (not deltas) and have been merged into the main specs directory. No existing specs were modified.

## Archive Contents

Located at: `openspec/changes/archive/2026-09-29-hackathon-analysis/`

- explore.md — Exploration and approach decisions
- proposal.md — Scope, capabilities, risks, dependencies
- design.md — Technical architecture, interfaces, error taxonomy
- apply-progress.md — Phase 1 apply batch evidence (PRs #20-#26 from first apply batch; phases 2-11 remain in later apply batches)
- tasks.md — All 11 phases marked complete with verification notes
- verify-report.md — Verification report with post-verify update confirming resolution of blockers 11.3 and 11.5
- specs/hackathon-analysis/spec.md — Full spec for admin-only fresh analysis, slug management, topic linking, daily cap
- specs/page-fetch/spec.md — Full spec for safe fetch with SSRF guards, browser fallback, quota handling
- specs/llm-extraction/spec.md — Full spec for Workers AI extraction with strict schema, null-over-guess, snippet validation

## Tests and Verification

- **Unit tests**: 70 files, 803 tests passing
- **Typecheck**: Clean, no errors
- **Production smoke test**: Passed on 2026-09-29 with bot promoted to admin for pinning
- **Production-faithful harness**: Passed for tokenized-stocks after PRs #33-#37

## Non-Blocking Follow-ups

The verify-report identifies these non-blocking follow-ups for future work:
- R3-001: an all-null extraction counts as usable.
- R3-002: the text/plain static path is not normalized with `normalizePageText`.
- R3-003: escape test gap.
- 405 is not in the bot-wall set.
- The harness always runs the rendered fetch.
- A claimed-job transient retry keeps the claim.
- **New**: When the pin fails on a fresh run through the queue consumer, the "Pinning failed" note is dropped, so the user is not told. Consider surfacing pin failure in the linked message or via a separate notification.

## Artifact Summary

All SDD artifacts for hackathon-analysis have been archived:
- Full change history and decision rationale preserved
- Delta specs successfully merged into main specs
- All task phases marked complete with verification evidence
- Operator readiness confirmed on production

## SDD Cycle Closure

The hackathon-analysis change is now complete and closed:
- Proposed, specified, designed, implemented, verified, and archived
- Ready for roadmap change 4 (natural language layer) to depend on the /hackathon domain APIs
- All operator steps completed and tested on production

---

**Archive Date**: 2026-09-29
**Archived By**: SDD Archive Executor
**Engram Topic Key**: sdd/hackathon-analysis/archive-report
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,7 @@ index.queue ─ buildHackathonConsumer(env) ─ runHackathonJob(msg)

### Worker egress reality

Verified with the harness (`npm run harness`, real modules on Cloudflare with real AI and BROWSER bindings) against https://www.bnbchain.org/en/hackathons/tokenized-stocks: from Cloudflare egress the static fetch gets HTTP 403, so production always takes the rendered (browser) path, whose text is 2.5k-3.4k chars. Rendered `innerText` separates table cells with TAB characters, and Qwen copies snippets containing raw TABs into JSON string literals, which `JSON.parse` rejects ("Bad control character in string literal"). Consequences in the design: (1) both fetch paths pass their text through `normalizePageText` (every ASCII control character except `
` becomes a space, space runs collapse per line, 3+ newlines collapse to a blank line; the 22,000 cap still applies), and (2) the extractor repairs control characters in string literals as a second line of defense. GLM returns fields it cannot find as `{ "value": null, "snippet": null, "confidence": 0 }`, handled by the per-field validation above.
Verified with the harness (`npm run harness`, real modules on Cloudflare with real AI and BROWSER bindings) against https://www.bnbchain.org/en/hackathons/tokenized-stocks: from Cloudflare egress the static fetch gets HTTP 403, so production always takes the rendered (browser) path, whose text is 2.5k-3.4k chars. Rendered `innerText` separates table cells with TAB characters, and Qwen copies snippets containing raw TABs into JSON string literals, which `JSON.parse` rejects ("Bad control character in string literal"). Consequences in the design: (1) both fetch paths pass their text through `normalizePageText` (every ASCII control character except `\n` becomes a space, space runs collapse per line, 3+ newlines collapse to a blank line; the 22,000 cap still applies), and (2) the extractor repairs control characters in string literals as a second line of defense. GLM returns fields it cannot find as `{ "value": null, "snippet": null, "confidence": 0 }`, handled by the per-field validation above.

The rendered fetcher calls `page.setRequestInterception(true)`. Each request goes through the pure `browserRequestPolicy`: http(s) only, the same host guard, images, fonts, media and stylesheets aborted, and at most 100 requests. After `goto`, it re-checks `page.url()`, reads `innerText` (capped), and calls `browser.close()` in `finally`. A 429 raises `BrowserQuotaExceededError`.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,7 @@ Chain strategy: stacked-to-main
## Phase 11: Operator Rollout (Manual — Not Performed by Apply)

- [x] 11.1 Run `npx wrangler queues create hackathon-analysis`. (Queue created in Cloudflare on 2026-09-28.)
- [ ] 11.2 Enable the Workers AI and Browser Rendering bindings for the Worker.
- [ ] 11.3 Give the bot the "can pin messages" right in the target chat(s).
- [x] 11.2 Enable the Workers AI and Browser Rendering bindings for the Worker. Confirmed in production (Workers AI and Browser Rendering bindings working; production-faithful harness run on Cloudflare).
- [x] 11.3 Give the bot the "can pin messages" right in the target chat(s). The bot must be a group ADMINISTRATOR with "Pin Messages". The member permission toggle is not enough, because the Bot API requires admin `can_pin_messages`. Verified on 2026-09-29: the first topic run had pinned_message_id null while the bot was only a member; after promotion, pinned_message_id is 271 in topic 262.
- [x] 11.4 Verify the exact `@cf/...` catalog IDs and context windows (>=10k tokens) for GLM-5.3-Flash and DeepSeek V4 Flash, and set them in `vars.HACKATHON_MODEL_PRIMARY`/`HACKATHON_MODEL_FALLBACK`. Verified 2026-09-28 with `wrangler ai models schema`: primary `@cf/zai-org/glm-5.3-flash` (context 1,310,720), fallback `@cf/deepseek-ai/deepseek-v4-flash-0731` (context 1,310,720). Both declare an OpenAI-style `choices[].message.content` output (string or null), not `{ response }`; `parseModelOutput` unwraps it. Revised after the production smoke test (2026-09-29): GLM-5.3-Flash and DeepSeek V4 need Workers Paid (403 AiError 5035), so the vars are now the Free-plan models primary `@cf/zai-org/glm-4.7-flash` (context 131,072) and fallback `@cf/qwen/qwen3-30b-a3b-fp8` (context 32,768).
- [ ] 11.5 Apply the D1 migration remotely, deploy, smoke-test `/hackathon <url>` in general chat then in a topic.
- [x] 11.5 Apply the D1 migration remotely, deploy, smoke-test `/hackathon <url>` in general chat then in a topic. Smoke test passed on 2026-09-29. General chat: bnb-ai-hack and tokenized-stocks. Topic: link and pin via `/hackathon bnb-hack-tokenized-stocks-edition`. After the production fixes PR #33–#37, which are documented in apply-progress.md.
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# Verify Report: hackathon-analysis

Base: `main` @ bf43602 (PRs #20-#37 merged). Mode: openspec, Strict TDD.

## Verdict: PASS WITH WARNINGS
0 CRITICAL, 3 WARNING, 6 SUGGESTION (non-blocking follow-ups). Archive is blocked only by operator items (see "Blocks archive").

## Evidence
- `npm test`: 70 files passed, 803 tests passed, 0 failed (~48 s).
- `npm run typecheck` (`tsc --noEmit`): clean, no errors.

## tasks.md
- Phases 1-10: done.
- 11.1 done. 11.4 done (models revised to Free-plan `glm-4.7-flash` / `qwen3-30b-a3b-fp8`).
- 11.2 unchecked in tasks.md but confirmed by evidence (AI + Browser bindings work in production and on the Cloudflare harness). Checkbox should be ticked.
- 11.3 pending (operator: bot "can pin messages" right).
- 11.5 unchecked: passed in production for bnb-ai-hack; tokenized-stocks passed on the production-faithful harness after PR #37; final Telegram confirmation pending Browser Rendering daily-quota reset.

## Design/spec consistency (grep spot-checks)
- wrangler.jsonc: `HACKATHON_MODEL_PRIMARY=@cf/zai-org/glm-4.7-flash`, `HACKATHON_MODEL_FALLBACK=@cf/qwen/qwen3-30b-a3b-fp8`, `max_retries: 5` - OK.
- workers-ai-extractor.ts: chat `messages` input, `max_tokens` 2500 (comment; passed as `MAX_TOKENS`) - OK.
- analyze-hackathon.ts: `BOT_WALL_STATUSES = {401,403,429,503}` - OK.
- extraction.ts: per-field `wrong-shape` rejection - OK.
- `normalizePageText` in html-to-text.ts, used by the static HTML path and rendered-fetcher.ts - OK.

## Coverage table (keyword mapping; file hits in test/)
| Spec requirement / scenario | Covering tests | Status |
|---|---|---|
| Admin-only fresh analysis (admin / non-admin) | request-hackathon-analysis, command-outcome, hackathon-command-e2e | Covered |
| Member re-shows by slug (member / admin in topic) | show-analysis, argument, e2e | Covered |
| No-argument (linked topic / nothing linked) | show-topic-analysis, argument | Covered |
| Slug vs URL classification | argument.test.ts, url.test.ts | Covered |
| Slug generation + collision suffix | slug.test.ts, hackathon-analysis-repo | Covered |
| Same-URL refresh keeps slug | run-hackathon-job / repo (refresh) | Covered |
| Failed re-analysis keeps prior result | no explicit keyword hit | WARNING (verify by inspection; likely structural: persist only on success) |
| One analysis per topic / conflicts move link | link-analysis-to-topic | Covered |
| Pin failure falls back to unpinned | pin hits in commands, run-hackathon-job, link-analysis | Covered |
| Daily cap | analysis-quota, request-hackathon-analysis | Covered |
| Job safety: running / enqueue failure / duplicate / retries exhausted | request-hackathon-analysis, run-hackathon-job (line 406), queue tests | Covered |
| Listing read-only + truncated | list-analyses, text-limit | Covered |
| Plain text replies | format.test.ts, commands | Covered |
| Strict schema (well-formed / malformed / per-field / null-not-found) | extraction.test.ts, workers-ai-extractor.test.ts (incl. GLM null fields, line 363) | Covered |
| Null over guess | extraction.test.ts | Covered |
| Snippet per non-null field (present / flattened ws / absent / null) | extraction.test.ts | Covered |
| Untrusted page framing | prompt.test.ts | Covered |
| Distinct schema-failure error | analyze-hackathon.test.ts (line 290), run-hackathon-job | Covered |
| Workers AI quota exhaustion non-retrying | workers-ai-extractor, run-hackathon-job | Covered |
| Scheme/destination guard (static + browser + redirect) | safe-fetcher, rendered-fetcher, url.test.ts | Covered |
| Size cap / time cap | safe-fetcher.test.ts lines 152, 171 | Covered |
| Thin text triggers browser fallback / sufficient skips | analyze-hackathon.test.ts | Covered |
| Bot-wall triggers fallback / 404 does not / browser fails after bot-wall | analyze-hackathon.test.ts (line 165) | Covered |
| Browser 429 with usable / insufficient static text | analyze-hackathon.test.ts | Covered |
| No raw page stored or logged (bounded output; safe fetch-error log) | run-hackathon-job.test.ts 481-561 (SECRET/example.com not logged), safe-logger | Covered (no-storage half by inspection: WARNING-level evidence) |

Uncovered by keyword: "Failed re-analysis keeps prior result" (no explicit test found), "successful analysis stores only bounded output" (no explicit assertion found).

## Warnings
1. No explicit test for "Failed re-analysis keeps the prior result".
2. No explicit test asserting only bounded output is stored (raw page never persisted).
3. tasks.md checkbox 11.2 not ticked despite confirmed evidence.

## Non-blocking follow-ups
- R3-001: an all-null extraction counts as usable.
- R3-002: the text/plain static path is not normalized with `normalizePageText`.
- R3-003: escape test gap.
- 405 is not in the bot-wall set.
- The harness always runs the rendered fetch.
- A claimed-job transient retry keeps the claim.

## Blocks archive
- 11.3: operator must grant the bot the "can pin messages" right.
- 11.5 final confirmation: tokenized-stocks Telegram smoke test after the Browser Rendering daily quota resets (production-faithful harness already passed).
- (Housekeeping) tick 11.2 in tasks.md.

## Post-verify update

Archive blockers 11.3 and 11.5 are now resolved:

- **11.3 resolved** (2026-09-29): The bot was promoted to group ADMINISTRATOR with "Pin Messages" right. Verified: first topic run (topic 262) had `pinned_message_id: null` while the bot was only a member; after promotion, `pinned_message_id: 271`.
- **11.5 resolved** (2026-09-29): Smoke test passed on production. General chat runs: bnb-ai-hack and tokenized-stocks. Topic run: link and pin via `/hackathon bnb-hack-tokenized-stocks-edition`. All production fixes from PRs #33–#37 documented in apply-progress.md.

**Verdict remains**: PASS WITH WARNINGS with 0 CRITICAL blockers. Archive may proceed.

**Additional non-blocking follow-up**:
- When the pin fails on a fresh run through the queue consumer (`postAnalysisAndLinkTopic` in src/domain/usecases/run-hackathon-job.ts), the "Pinning failed" note is dropped, so the user is not told. Consider surfacing pin failure in the linked message or via a separate notification.
Loading
Loading