add: [wip] transformation servos feature (custom opensearch ingestion… - #388
Merged
Merged
Conversation
OpenAPI changesAdded endpoints
|
Reject attributes a servo discarded. A `drop` processor makes OpenSearch answer the index request with HTTP 200 and `result: noop` (verified on 3.4), which create_attribute never checked -- so the API returned 201 for an attribute that was not in the index, and queued correlation, reactor and notification work for a document that does not exist. The result is now verified before any of that happens, raising AttributeNotIndexedError, which POST /attributes/ maps to a 422 and feed loops count as a failed row. The existing attribute tests had to start saying what OpenSearch returned, which is the check doing its job. Surface servo failures. The on_failure handler already wrote to expanded.servo_errors, but nothing read it back, so a servo erroring on every single document still showed a green "enabled" badge. GET /tech-lab/servos/errors aggregates counts and distinct messages per servo, rendered as a badge with the messages in its tooltip. Let servos be reordered. Order matters as soon as one servo reads a field another produced. The arrows in the new order column post the whole order to /tech-lab/servos/reorder, so the chain is rebuilt once rather than once per servo and is never half-sorted in between. A servo containing a `drop`, including one nested in foreach or on_failure, is flagged with a "discards" badge and a warning in the editor rather than looking like any other servo. Dropping stays allowed -- deduplication is a legitimate use -- it just stops being invisible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ApiTester's autouse cleanup fixture deletes every user, organisation, feed, server, tag, taxonomy and galaxy in Postgres and every event, attribute, object, object reference and analyst-data document in OpenSearch. It does that unconditionally, before the first test body runs, against whatever the environment points at -- so the documented `docker compose exec api poetry run pytest` silently wipes a dev stack, and a single ApiTester-based file wipes just as much as the whole suite. require_test_environment() now stops that unless ENVIRONMENT=test, exiting 1 with the target it was about to destroy and the command to proceed. It is called from the db fixture, so the run ends before anything connects, and from teardown_db and _cleanup_opensearch so no other caller can reach them. Gating on ENVIRONMENT rather than the database name because CI and the dev stack both use a database called `misp`; the name cannot tell them apart. CI already sets ENVIRONMENT=test, so it is unaffected. Docs updated where the old command was published: CLAUDE.md, the development guide and the api README. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
create_attribute now verifies the index response, so the bare MagicMock in _create_two_attributes read as a dropped document and raised AttributeNotIndexedError. Same fix as the one already applied to the repository tests: the client reports a real create. Caught by CI, not locally -- the previous change was only exercised against the servo suites and the attribute repository tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #388 +/- ##
==========================================
+ Coverage 83.95% 84.43% +0.48%
==========================================
Files 201 208 +7
Lines 18276 19138 +862
==========================================
+ Hits 15343 16159 +816
- Misses 2933 2979 +46 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Searching a bare indicator value is the common case, and it matches attributes rather than events -- an event does not carry the value in its own fields. The view opened on "Events 0" next to "Attributes 3", which makes the platform look like it found nothing. Two watchers already tried to cover this, but only between events and attributes, and each read the other store's total at whatever moment its own result landed. The three searches are dispatched together, so an early responder compared against the previous search's totals and the swap never happened. Correlations was not handled at all. Replaced with one watcher that waits for all three to settle, then picks the first tab in display order with results. A tab that has results is never taken away, the choice is armed per search rather than on every store update -- so deliberately opening an empty tab to reach its filters still works -- and when nothing matched anywhere the tab is left alone rather than shuffling through three empty panels. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /attributes/` adds `must_not: exists: object_uuid` when no object is given, so the tab lists only attributes that are not inside an object. The badge showed `event.attribute_count`, which counts every attribute on the event. The Emotet fixture event advertises 9 and lists 2. The badge now counts what the tab lists: fetched once on mount, because the panel is mounted lazily and the number has to be right before anyone opens the tab, then tracked from the tab's own page total once it is. AttributesIndex is the only writer of that store, so creates, deletes and edits keep it current without new plumbing. The title block still reads "9 attributes in 3 objects" -- that is the whole-event total, and it is what says where the other 7 are. The graph spec asserted the old 9, so it now asserts 2 alongside the Objects count of 3. Its screenshot still shows the old badge and needs regenerating on an instance seeded with docs fixtures only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "Event results" branch matched on `index_target === 'events'` with no hunt-type guard. A cpe hunt targets events, so it was caught there and never reached its own renderer a hundred lines below -- producing a uuid / info / org / date table with one row per CVE and every cell empty, 246 of them for the Exchange fixture hunt. Guarded it the way the attributes branch already was, with an explicit `opensearch` / `mitre-attack-pattern` check, so any hunt type that has its own renderer falls through to it rather than being intercepted by a generic index_target match. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both quickstarts told the reader to fork before running but not that the fork begins with nothing imported. Kernels are keyed per (user, notebook), so a fork is a new notebook and therefore a new kernel; running any cell before the imports in section 1 fails with `NameError: name 'pd' is not defined`. The note names that error so searching for it finds the answer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`seed-docs-fixtures` covers the events, attributes, objects, hunts and audit rows behind the documentation screenshots, but only 5 of the 15 screenshot specs run on real data. The other 10 stub every call with Playwright, so their content lives in TypeScript fixtures and never reaches a database -- which is exactly the part a demo needs. seed-demo builds on the docs fixtures and adds servos, reactor scripts, analyst notes and opinions, feed definitions, library notebooks, per-hunt run history and notifications, plus a second set of events whose indicators deliberately overlap the first so the correlation engine has something real to find. Additive by design: everything is keyed by pinned uuid, exact name, or a payload marker, so re-running refreshes the demo in place and leaves the rest of the instance alone. --reset removes only rows the seeder created. Details worth knowing: - Hunt history is generated from a fixed RNG seed, so the heatmap is identical on every machine -- a demo that looks different each run cannot be rehearsed. It is capped at one run per day because get_hunt_history returns the *oldest* 90 rows, so a busier hunt would draw the wrong end of the window. The seeder clears the Redis history and results keys as it goes, or the endpoint serves the previous seed from cache. - Each hunt is then executed once, because the results panel reads a Redis key only a real run writes. Without it the chart is populated above an empty table. The cpe and rulezet hunts reach external services and are allowed to fail so the demo still comes up offline. - Servos reference the shipped templates by slug rather than copying their processors, so the demo cannot drift from them. - Feeds are seeded disabled. A fetch reaches the network and returns whatever is live that day, which is the opposite of what fixture data is for. - Reactor scripts are seeded paused. An active one tags events as a side effect of seeding, which silently mutates existing data and changes the docs screenshots. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A route through the app for people seeing it for the first time, following one indicator -- the IP 185.220.101.42 -- from the search box, through the events that share it, to the correlation that ties them together. Six stops with what to click, what to say, and what should appear, plus timings; the first four are the core if there are only six minutes. Written against a seeded instance rather than from the code, so the numbers are what the dataset actually produces and the gotchas are real ones: the MITRE hunt legitimately returns only 2 because it matches a galaxy tag rather than doing text search, notebooks need the lab worker before anything is run, and fetching a feed live pulls whatever is on the internet that day and buries the curated data mid-demo. Ends with a troubleshooting table mapping each likely failure to its cause. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The demo had no written intelligence in it -- every seeded artefact was machine-readable, which leaves out the half of the product where an analyst actually writes something down. The report is a short incident write-up: summary, delivery chain, an indicator table, actions taken and open questions. It leans on the rest of the seed rather than standing alone, pointing at the shared C2 address and at the opinion attached to the Cobalt Strike event, so it reads as part of one investigation. Written to the index with a pinned uuid rather than through reports_repository.create_event_report, which generates its own and would stack a fresh copy on every seed. Reports hang off the event by uuid and survive a docs re-seed -- _clear_event_children drops attributes, objects and object references, not reports. Verified rendering: five headings, the indicator table and the blockquote all come out as Markdown, with no raw syntax leaking into the page. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The event had a written report but nothing showing the reasoning behind it, and the seed carried no relationship at all -- so one of the three analyst data types never appeared anywhere in a demo. Added a note on how the timeline was reconstructed, an opinion scoring the TA542 attribution at 40/100, a note on the loader URL recording that the host is a compromised third party rather than attacker-owned, and a related-to relationship from the Emotet event to the Cobalt Strike one. They pick up the thread the report leaves open, so the low-confidence attribution and the shared C2 address are argued in the same place an analyst would argue them. _seed_demo_analyst_data could not recognise a relationship on a re-run: entries are matched by content because analyst data generates its own uuid, and the matcher only looked at note and comment text, which a relationship has neither of -- so every seed would have stacked another copy. It now keys a relationship on its type and target, and checks the relationships bucket alongside notes and opinions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The demo had no server connections, so the sync section was empty and there was nothing to show anyone asking whether the platform talks to their existing MISP. Two are seeded: a pull-only partner with tag-based pull rules, and an internal hub that pushes and pulls with rules excluding tlp:red and tlp:amber. Between them they cover most of the configuration surface -- galaxy cluster sync, sighting push, self-signed certificates, proxy skipping, priority. Both point at hostnames on the reserved `.invalid` TLD, which by RFC 2606 can never resolve, so pressing Pull on a demo instance fails at DNS with "Remote MISP instance not reachable" rather than reaching somebody's real server -- verified by actually pressing it. The auth keys are placeholders that say so in the value itself; there is no credential in the fixture. Left with pull and push enabled so the list looks like a real deployment. Nothing syncs unattended: a scheduled pull is a redbeat entry a user creates, not a static beat schedule. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude Desktop installs MCP servers as .mcpb bundles, and there was no way to connect one to an instance short of hand-editing config. Two constraints decided the shape. A .mcpb cannot declare a remote HTTP server -- all four server types run locally -- and Claude Desktop's native remote path, Settings -> Connectors, expects OAuth, which leaves nowhere to put the scoped bearer token the MCP endpoint actually uses. So the bundle ships a local stdio bridge that forwards to the instance over HTTP. The bridge is written out rather than delegating to `npx mcp-remote`: the bundle then fetches nothing from npm at launch, puts nothing third-party in the path of a credential, and avoids the documented bug where Claude Desktop on Windows mangles an argument containing a space -- which is exactly the shape of "Authorization: Bearer ...". It handles both JSON and SSE response bodies, carries the session id from initialize through later calls, serialises requests so nothing races ahead of that id, and turns the failures users actually hit (blank settings, scheme-less URL, unreachable host, rejected token) into messages that say what to do. The token reaches the bridge as an environment variable, never as an argument, so it stays out of the process list; the manifest marks it sensitive so Claude Desktop keeps it in the OS keychain. Verified by driving the packed bundle over stdio against a running instance: initialize, tools/list (22 tools), and tool calls returning real data. The install-and-click flow in Claude Desktop is untested -- there is no Claude Desktop in the build environment -- and the bundle is unsigned. Built artifacts are gitignored; rebuild with `cd mcpb && npx @anthropic-ai/mcpb pack .`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The shipped OpenSearch dashboard has panels for MITRE ATT&CK techniques, malware families, tools, threat actors, targeted sectors, kill-chain phases and sightings. Every one of them was empty: each aggregates tags.name.keyword on misp-attributes filtered by a prefix regex and reads nothing else, and the demo attributes carried no tags at all. The attribute fixture grows from 6 to 22, each carrying a TLP level, a kill-chain phase and the galaxy tags its panel looks for. The tag forms were taken from the visualisations themselves rather than guessed -- a malware panel that matches misp-galaxy:(malpedia|mitre-malware|ransomware|...) does not count a tag written any other way. The galaxy values are real cluster values checked against the loaded galaxies, so they resolve in the UI as well as counting on the dashboard; that check also caught threat-actor="Turla Group" in the events fixture, which is not a cluster -- the value is "Turla". Sightings are new: 107 across 45 days, each with a sensor, an observing organisation and a verdict, which fills the four sightings panels. They are written directly with a pinned id because create_sightings dispatches a notification task per sighting -- a hundred rows would mean a hundred notifications, and no way to re-seed without duplicating them. create_attribute always writes an empty tag list, so tags are applied after the attribute exists. Verified by running each panel's own aggregation against the index, then loading the dashboard: 66 panels render and nothing reports "No results found". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This reverts commit 477b461.
`python -m app.cli --help` crashed with
TypeError: Parameter.make_metavar() missing 1 required positional argument: 'ctx'
Click 8.2 added a required argument to make_metavar, and typer 0.15 calls it
without one. The installed click is 8.3.3, so every --help in the CLI was
broken -- including the subcommand help, which is the only place the flags on
seed-demo are documented, and which CLAUDE.md tells people to run.
typer 0.16 calls it correctly. Bumped the floor rather than pinning click back,
so the fix does not have to be undone later.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The published image still showed "Attributes 9" -- the number bf7ca54 changed to 2 -- so the spec asserted one thing and the documentation showed another. Captured on an instance carrying only the docs fixtures, so the Related Events panel lists just the Cobalt Strike fixture event and the tag list is the fixture's own. The demo seeder's events and the workflow tag its reactor script used to apply are both absent, which is what made earlier attempts unusable. Only the tab strip capture is included: the graph captures differ between runs by force-simulation jitter alone, and the spec's node-count assertion passes either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
righel
marked this pull request as ready for review
September 17, 2026 07:09
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
… pipelines)