Skip to content

Latest commit

 

History

History
424 lines (300 loc) · 110 KB

File metadata and controls

424 lines (300 loc) · 110 KB

macMCP

Standalone Swift MCP server exposing macOS-native tools via stdio, grouped by service under Sources/macMCP/Services/. No external dependencies.

Architecture

Single-threaded stdin/stdout MCP server. Newline-delimited JSON-RPC 2.0. Protocol version 2024-11-05.

main.swift           Stdin loop, JSON-RPC dispatch (initialize, tools/list, tools/call)
JSONRPCTypes.swift   Wire types, JSONValue enum, MCPTool, result helpers
ToolRegistry.swift   Tool registration map + JSON schema builder helpers
Services/            One file per service, each a caseless enum namespace

Entry point initialises NSApplication (.prohibited -- no dock icon) for macOS TCC permission support, then reads stdin line-by-line, dispatches to ToolRegistry, writes JSON to stdout. No async, no concurrency -- all async APIs bridged synchronously via CFRunLoopRunInMode.

Services

Service Tools Backend
Calendar 3 EventKit EKEventStore
Contacts 10 CNContactStore
Reminders 3 EventKit EKEventStore
Location 3 CoreLocation (RunLoop-pumped, 15s timeout)
Maps 3 CLGeocoder + NSWorkspace URL schemes
Capture 2 /usr/sbin/screencapture, /usr/bin/afrecord
Mail 14 JXA via /usr/bin/osascript -l JavaScript
Messages 5 SQLite3 on ~/Library/Messages/chat.db (read), AppleScript (send)
Shortcuts 2 /usr/bin/shortcuts CLI
Utilities 1 /usr/bin/afplay
Weather 3 api.open-meteo.com (free, no key)
Web 1 URLSession (http/https GET, 1 MB cap)
System 1 TCC status for every service these tools need

Key Patterns

  • Service = caseless enum with static register(_ registry: ToolRegistry) and private static handlers (JSONObject?) -> MCPCallResult.

  • Sync-over-async -- CFRunLoopRunInMode pumps the main RunLoop to deliver callbacks synchronously. All async APIs (EventKit, CoreLocation, CLGeocoder, URLSession, CNContactStore) use this pattern. Semaphores are not used because NSApplication.shared routes completions through the main RunLoop, which semaphores would deadlock.

  • Permissions re-requested on every tool call. macOS caches the grant, so this is idempotent.

  • No throws across service boundary -- all errors returned as MCPCallResult(isError: true).

  • Every tool states both annotation hints, and MCPAnnotations has no optionals so it cannot forget. readOnlyHint answers does this change anything; openWorldHint answers a second, orthogonal question -- is the call itself what reaches beyond this Mac -- and relay gates them separately (access: "read" on the first, an allow_external grant on the second). Both default, when absent, to the answer that costs the tool its availability or hands it a reach it should not have, so an omission is never inert; non-optional stored properties make the compiler refuse a registration that omits one, and ToolAnnotationTests pins the value tool by tool. The two axes are genuinely independent -- web_fetch and weather_* are read-only and open world, which is how a read-only mail profile came to hold an outbound HTTP channel. Two things deliberately do not count as open world, or the hint would be true of most of this server and mean nothing: replication of a local store to the user's own account (Mail's Maildir, EventKit, CNContactStore, chat.db -- that sync happens whether or not a tool is called, carries nothing the caller chose, and reaches only accounts this Mac is already logged into), and naming a remote thing without contacting it (maps_get_directions concatenates a maps.apple.com URL out of the caller's own strings and opens no socket). Twelve tools are open world: web_fetch, weather_* (3), location_* (3 -- a Mac has no GPS, so even location_get_current resolves position over the network), maps_search, maps_open, mail_send, messages_send, shortcuts_run. mail_send and mail_create_draft differ on this hint and on nothing else, which is the point of the feature: both mutate, so the access mode cannot separate them, and "may draft, may not send" is expressible only on the second axis. destructiveHint and idempotentHint are deliberately not modelled -- nothing consumes them, and an annotation nobody checks is a claim that rots unnoticed.

  • Messages reads require Full Disk Access (direct SQLite on chat.db); messages_send separately requires an Automation grant for Messages.app (com.apple.MobileSMS -- its bundle ID from its iOS-app origins, not a name mentioning Messages), reported by permissions_check as automation (Messages). Full Disk Access itself has no status API and is not reported there -- only Automation grants can be probed ahead of time.

    A caller-supplied limit/hours_ago/since is adversarial input, because a model can write any JSON number. Int32(limit) on the 64-bit Int JSONValue.intValue decodes traps -- and killed the whole process, every tool with it -- on anything past Int32.max; Int64((hours_ago * 3600 - epoch) * 1e9) traps the same way past about 2.6 million hours, and on a since far enough in the future (9999-12-31 parses fine as a date). extractDouble now rejects non-finite values before anything downstream converts them (Double("nan") and Double("inf") both parse successfully, which is exactly how a JSON string carries either), and every conversion to a bounded integer type clamps rather than converts -- boundedLimit, appleEpochNanos. limit is read as leniently as hours_ago always was (int, whole double, or numeric string), because JSONValue.intValue matching only .int meant "limit": "50" silently fell back to the default with no error, a narrower failure mode than the one coercedStringValue already exists to close for message_id in Mail.

    messages_send does not know whether Messages actually sent anything, because nothing in Messages' scripting dictionary says so. send declares no result, so osascript exiting 0 means only that the AppleScript didn't throw -- sending to a syntactically-plausible but unregistered address (nobody@nowhere.invalid) returned "message sent" and the message never left: chat.db's own is_sent/error columns told the true story (is_sent: 0, error: 22), just not to anything that read them. mail_send's whole design is verifying a send against the message Mail is actually holding; Messages offers no equivalent surface to verify against except the same database every read tool already uses. So sendMessage captures the moment before it runs the script and polls chat.db afterward (awaitDeliveryOutcome, newestOutgoingMessage) for the newest outgoing message in that recipient's chats created after that moment, and reports what it finds: sent (is_sent = 1), a hard error naming the code (error != 0), or -- if chat.db could not be opened at all, almost always missing Full Disk Access -- a distinct cannotVerify outcome, because "check messages_get_chat" is not useful advice to a caller who just learned that tool needs the same grant. Measured against this box's real Messages: a working send confirms in under 2 seconds; one Apple's servers reject can take 20+ seconds to carry an explicit error, which is why the wait is bounded (timeout_seconds, default 8) rather than open-ended -- a bound long enough to always catch a failure would make every successful send pay for it too, so the default favours the fast common case and reports the slow, rare one as unconfirmed instead of guessing. Both read tools that return message rows (messages_get_chat, messages_search) now surface a failed outgoing message's send_error for the same reason mail_get_email's fidelity fields exist: a message that never sent must not read identically to one that did.

    messages_send reaches osascript as a file, and its output through temp files, for the same reasons Mail's JXA calls do. A script passed as -e is bounded by ARG_MAX, so a long message text could fail with "Argument list too long" instead of sending -- the file has no such ceiling, and MailService.stripScriptPath (shared, not duplicated) undoes the path prefix osascript then puts on a thrown error line. Separately, and more seriously: AppleScript's own error text echoes back whatever it failed on (a bad buddy identifier, a value it "can't make ... into type"), so to or text past 64KB, echoed into a thrown error, used to fill a shared Pipe's buffer while nothing was draining it -- waitUntilExit() deadlocks forever in that shape, and with this server's single synchronous stdin loop, so does every other tool behind it. stdout and stderr now go to their own temp files (never read until after the process ends), and the wait itself is bounded (30s, then terminate(), then SIGKILL) so a stuck osascript -- most likely an unanswered Automation consent prompt -- cannot hang the call forever either.

    messages_list_chats orders by the newest message in a chat, not by chat.ROWID. ROWID order is chat creation, so a years-old conversation and a chat created moments ago by an unrelated failed send sort by which was created more recently, not by who actually said something last -- and it joins through chat_message_join (whose message_date column carries exactly this), which incidentally also drops a chat with zero messages, which "recent conversations" should not list to begin with.

    messages_search's scan cap applies with no query too. It used to bind SQL LIMIT to the caller's limit whenever there was no text query, on the premise that every row scanned is a row returned -- false, since a row with no text anywhere (an attachment, tapback or deleted message) is read and then dropped regardless of whether a query is filtering anything. A window with real messages past a run of such rows came back short of limit while scan_complete and its note blamed the window rather than the truth. The cap is now searchScanLimit unconditionally, with the early-exit on messages.count >= limit (checked before appending, not after -- limit: 0 used to let exactly one match through) doing the same job the old special case was trying to.

    Messages tools are deliberately outside the _meta/ResourceScope mechanism every other service here uses (see ContextSchema.swift's header): there is no resource axis short of per-chat, so a profile that needs to restrict Messages is confined by allowed_tools and the access mode alone, not by a messages_accounts/messages_chats-shaped field. This is a considered gap on the read side; messages_send additionally binds no recipient allowlist of any kind (unlike mail_send, which refuses a from no account owns), so a write profile holding messages_send can message anyone. file_dirs is the one exception -- a different axis (which host directories, not which chats) -- and governs messages_save_attachment the same way it governs mail_save_attachment; messages_send's image_path is confined the same way mail_send's attachments is, at the parameter rather than the tool.

    Images only, on both the read and the send side. messages_get_chat and messages_search join attachment/message_attachment_join and filter to mime_type LIKE 'image/%'; a video, PDF, vCard or sticker stays exactly as invisible as it always was rather than surfacing as a kind of attachment messages_save_attachment cannot save. An attachment is reported as attachment_id (an opaque handle, chat.db's own attachment.ROWID) plus filename (the attachment's own name, never the path it happens to sit at on this disk) and mime_type; the on-disk path is read only inside messages_save_attachment, by id, and is never itself a caller-visible value. A message with an image and no caption has empty text -- it is still returned when browsing (messages_search with no query), but a query cannot match it, since there is no text to match. messages_save_attachment validates the attachment is an image (refuses otherwise), reads the bytes off chat.db's own copy, and writes them to a caller-named destination; there is no directory-vs-file inference the way mail_save_attachment has, destination is always the exact file to write.

    Sending an image is a second send call in the same script, not the same call with text. Messages' send command takes either a block of text or a file, never both (confirmed against its own scripting dictionary), so a captioned photo is send "caption" to targetBuddy immediately followed by send (POSIX file "...") to targetBuddy. image_path is validated two ways before either line is written: HostFileScope.resolve (may this client read this path at all), then NSImage(contentsOfFile:) decoding it (is it actually an image, not a file merely named like one) -- send itself validates neither and will happily hand Messages a nonexistent path or a renamed text file.

    Messages.app is itself sandboxed, and a path outside its own directories is silently unreadable to it. Measured on this box: a real, valid image at /tmp/... or ~/Desktop/... makes send (POSIX file ...) exit 0 and chat.db later show is_sent: 0, error: 25 -- nothing surfaces to osascript at the time, so this is indistinguishable from a network failure until chat.db is checked. The fix macMCP takes (matching a workaround already shipped elsewhere for the identical problem) is to stage the image first: copy it into ~/Library/Messages/.send-staging/<uuid>-<name> -- a directory under Messages' own ~/Library/Messages/ tree, which its sandbox already permits -- and send that path instead, cleaning it up after.

    Staging fixes the read; it does not fix delivery, and on this box delivery never succeeded. Every real send attempted against a live remote number -- across four different source directories, two different images, before and after staging -- came back is_sent: 0 with error 25, 22 or 1 on different attempts, never is_sent: 1. The recipient confirmed nothing arrived. Text to the same number over the same period sent correctly every single time. So this is not "slower to confirm than text" (that was the working hypothesis before the recipient checked); on this VM, image delivery is not confirmed to work at all, and error changing across otherwise-identical attempts suggests an account- or network-level cause past anything client-side code can paper over -- possibly specific to running iMessage inside this VM's network path, untested on real hardware. messages_send's own description says this plainly (image_path is "unverified/likely broken") instead of the softer "expect it to be slow" wording an earlier pass shipped, which undersold what the recipient later disproved. Extraction is unaffected by any of this: messages_get_chat/messages_search/messages_save_attachment read chat.db and the Attachments folder directly, no send, no sandbox, no network.

    The delivery check has to confirm every row a send created, not the newest one by date. awaitDeliveryOutcome originally read back only the single newest outgoing message (by date DESC) to decide whether the send succeeded -- fine when there is one row, wrong once there can be two. Measured live: a text + image_path call creates two message rows, and once the text row's send is confirmed its date advances past the image row's, so ORDER BY date DESC LIMIT 1 returned the older, already-sent text row while the image sat at is_sent: 0 for over two minutes -- the call reported "message and image sent" for an image that never went out. newestOutgoingMessages now orders by m.ROWID DESC (assigned once, at insertion, and never revised) and reads back expectedCount rows -- 1 or 2, matching how many send lines the script issued -- requiring all of them is_sent for .sent and reporting .failed if any carries a nonzero error, even if another already succeeded.

  • Mail uses JXA because Mail.app has no public framework API. String escaping is manual.

  • Mail reads must use bulk column fetches. mbox.messages.subject() costs one Apple Event per column per mailbox regardless of message count (~0.07ms/message; 14k subjects in ~1s). Three things are forbidden in a read path, all measured against a 14,004-message mailbox:

    • messages[i].prop() in a loop — resolving one element is O(mailbox size), ~13ms on a small mailbox and ~150ms on a large one.
    • specifier .slice() — resolves element by element; 200 ids took 22s while all 14,004 took 1s.
    • whose() — an internal linear scan, ~10x slower than fetching the column and filtering in JS. whose({content: ...}) decodes every body and times out on ~251 messages.

    Columns from one mailbox must be checked for alignment before they are paired. Each is a separate Apple Event and the scan walks them by index, so a mailbox that changes between two of them pairs one message's id with another message's subject — and the id is the handle mail_move, mail_mark_read and mail_get_email act on. Mail orders the collection by date received, so an arrival is spliced into the middle: moving a message into Alice's INBOX between two fetches produced ids=9 subjects=10 with the new message at index 8 of 10, i.e. row 8 carrying id 412 under subject "Contract 2024-118". It is not theoretical — a scan run while a probe was arriving handed mail_get_source an id whose message measured 405 bytes against the probe's 300599. The scan therefore re-reads the id column after the others and requires it to come back identical (plus every column being the same length). Cost is one extra id column fetch per mailbox, the cheapest of them.

    Detection is a trigger, not a verdict — the rows are recoverable one message at a time. The check used to discard the whole mailbox, report it in unstable_mailboxes after three attempts, and (single-account scope) turn that into an error. Measured under sustained IMAP delivery into Alice's INBOX, 8 scans each: 5 of 8 refused with total_messages: 0, i.e. ~11,800 readable rows thrown away to avoid a mispairing that affects at most the rows after the splice — and Mail's collection is newest-first, so an arrival splices at index 0 and shifts every index, which is why no prefix salvage works. What the columns cannot be trusted about is the pairing; the ids arrived in a single Apple Event and are not in doubt. So on divergence each of the <= limit rows actually being returned is re-bound with messages.byId(id) and asked for its own subject, sender, date and read state, and for where it is (byId resolves globally, so a row is dropped rather than relabelled if the message has left the mailbox it is stamped with). The mailbox is named in changed_mailboxes with rows_reverified / rows_dropped. Same 8 scans, same delivery rate, after: 8 of 8 returned 20 verified rows with a correct total_messages, each row's id/subject pairing confirmed through the independent mail_get_email path. Cost is nothing on a quiet mailbox (0.90s vs 0.90s) and ~75ms per returned row when it fires (2.0s -> 3.5s at limit 20). A wall-clock budget (20s per account scan) caps the pathological case, past which a changed mailbox goes back to being reported as unread.

    A search's count is the one thing the salvage cannot recover, because the match is decided against the very columns shown not to line up: a row that really matches can be filtered out under its neighbour's subject and never reach the re-read. Every row returned is still verified, so nothing returned is wrong — but scan_complete is false for a filtered scan over a changed mailbox, which is exactly what that flag means everywhere else.

    A mailbox only claims a message it stands behind. The scan keeps a seen set so a message filed in two mailboxes at once — the normal shape of a Gmail account, where everything in INBOX is also in All Mail — comes back once and is counted once. The ids went into it as the rows were built, i.e. before the alignment check decided whether this mailbox's rows could be stood behind, so a mailbox that was then discarded had already claimed them: the counts were rolled back and the mailbox named in skipped_mailboxes, and the message was then dropped from the clean mailbox that also holds it. It appeared in no row and was counted in no total, under a total_messages that read as an answer. The claim is now committed only when the mailbox is kept, and a row the re-read lets go — the message is not here, is unreadable, or does not match after all — releases both its claim and its count, because a mailbox that ends up saying nothing about a message must not be what removes it from the answer. The rollback subtracts what is still counted rather than the whole batch, since the re-read may already have taken part of it off.

    A scan that read nothing is an error, not an empty result. total_messages: 0 beside isError: false is an affirmative claim that a mailbox holding thousands of messages is empty, and a caller who does not know to read the coverage fields cannot tell the two apart. When no mailbox in scope could be read the call refuses and says so; a partial read still returns its rows, with scan_complete (reported unconditionally, true or false) and a note saying the counts are a floor. Every reason is named, in one sentence. The refusal used to have a branch per cause, so a scan where one account timed out and one mailbox would not hold still reported only the mailbox and never mentioned the timeout; and no mailbox named "X" — a claim about every account in scope — was returned flatly with part of that scope unread, before the payload, so failed_accounts and note were never emitted either. skipped_mailboxes now carries why, the way failed_accounts always has: catch (err) { skipped.push(label) } read the label and threw the error away, and "gone" and "Mail was busy" are not the same answer.

    The body-search sweep is a second scan and reports its own coverage. mail_search's body pass runs scanAllAccounts again over the same scope and used to read only rows and total from it — skipped, failed and changed were discarded and scanFailure was never applied. A sweep that read nothing contributed no candidates and nothing to its own total, so rows.count >= total was satisfied by the failure and the answer came out body_scan_complete: true, bodies_read: 0, body_matches: 0: the completeness test passed because it was never reached. body_scan_complete now includes sweep.scanComplete, and body_search carries body_scan_skipped_mailboxes / body_scan_failed_accounts / body_scan_changed_mailboxes / body_scan_note separately from the metadata scan's, because the two are reads of the same scope at two moments and either can fall short alone.

    The sweep has to be WIDER than body_scan_limit. It used to run at exactly that limit, and every row already being returned plus every row that matched on its own subject or sender was then subtracted from that same set with nothing put back. Measured: body_scan_limit: 5 with a query matching subjects returned bodies_read: 0, body_matches: 0 — indistinguishable from five bodies read and none matching — and a body-only hit at position body_scan_limit + 1 was unreachable at every limit, because raising the limit widened the sweep that ate it by the same amount. The sweep now asks for body_scan_limit + min(metadata matches, max(4x the limit, 100)), which costs no extra Apple Events (a scan reads whole columns whatever the limit is; the limit only trims rows) and is capped so a changed mailbox's ~11ms-a-row salvage stays inside its 20s budget. Live: six subject matches in front of one body-only match, body_scan_limit: 5 — bodies_read 0 / body_matches 0 before, bodies_read 5 / body_matches 1 after. What is still unspent is reported as body_scan_shortfall, with a note saying whether the sweep was capped or that was every message in scope. matchBodies also groups its candidates in arrival order rather than iterating a Swift Dictionary, whose per-process hash seed made which mailboxes got their bodies read before the deadline vary between identical calls.

    Message bodies cost ~1.2s each individually and in bulk, so body search is a capped second pass (body_scan_limit), never a full scan. Scans run one osascript per account so a wedged account degrades to a failed_accounts entry instead of losing the request; per-mailbox processes would cost more in spawn overhead (~150ms each) than the scan itself.

  • A JXA specifier is not the object it resolves to, and a positional one is a bug. mail.accounts(), acct.mailboxes() and mbox.messages() all hand back specifiers indexed by position, and JXA re-evaluates a specifier on every property access rather than snapshotting. found = mbox.messages[k] therefore means "whatever is at position k right now": bind it, let one message leave the mailbox, and the same variable answers for a different message — no error, no warning. Measured on the fixture at ~3,000 messages under continuous delivery, mail_get_email was wrong in 26 of 60 calls, and in 18 of those it returned the right id, the right subject and the right rfc_message_id beside another message's body — a shape nothing in the response lets a caller detect. The exposure is the whole call, not a gap between two reads. Four places had it: findMessageJXA, the body-scan pass, mail_create_draft's find-back, and the mailbox collection itself.

    The fix is to bind by something that identifies the element. messages.byId(n) re-resolves by id, so it answers for the same message or raises, and it is cheaper than what it replaced: ~11ms on a 3,887-message mailbox against 54ms for the id column, and 28ms for messages[i]. Two things about it, both measured against Mail 16:

    • byId resolves globally. An id from Bob's INBOX resolves through Alice's, or through mail.inbox, or through a local On-My-Mac box. So a numeric lookup needs no mailbox enumeration at all — and the account/mailbox in the answer has to come off the message (msg.mailbox.name(), msg.mailbox.account.name(); the account read raises for a local mailbox, which is how On My Mac is told apart), not off whatever was searched. Any scoping the caller asked for is then checked against that, because the search no longer enforces it.
    • An RFC Message-ID cannot be resolved that way. That path still reads a column to translate the header value into a numeric id, then binds by that id and asks the bound message for its own messageId(). A message answering for itself is what makes it sound; no alignment guard is needed.
    • A re-bind is only as good as the scope check on it. Because byId is global, every place that binds by id has to read the account and the mailbox back off the message. findMessageJXA always did (fmLocate/fmInScope, account matched case-insensitively); the body-scan pass compared the mailbox name alone, and every account has an INBOX — so a message that moved Alice:INBOX -> Bob:INBOX between the metadata scan and the body pass passed the check and had its body returned under a row saying "account": "Alice". Both halves are now read through one shared mbWhere/mbSamePlace, which the scan's per-row re-verification uses too.

    A leaf name does not identify a mailbox; the path does. Mail flattens an account's mailbox tree and reports leaf names, so acct.mailboxes.name() can return ["Archive","Projects","Sub","Archive",…] — two different mailboxes called Archive (top-level and Projects/Archive) — and Mail enumerates children before parents with the special mailboxes last, so "the first match" is systematically the nested one. Every tool resolved, labelled and excluded a mailbox by that name, and each was a different way of guessing between two: mail_move to "Archive" filed into Projects/Archive and to "Trash" into R4-PROBE-Deep/Trash (a "delete this" left undeleted in a project folder) both reporting verified: true, because the read-back looked in the mailbox the pick had chosen; mailbox: "all" dropped any folder anywhere in the tree whose leaf name was Trash/Junk/Drafts, 35 messages reported against 38 on disk with scan_complete: true and skipped_mailboxes: [].

    The identity is the path — container leaf names, outermost first, joined with / — and it is Mail's own, not a label invented here. Measured against Mail 16.0 on the fixture (Bob, 33 mailboxes, 4 deep):

    • mailboxes.byName('Projects/Archive') resolves, so a path is a handle, not just a label. All 33 paths resolved, including names with quotes, apostrophes, spaces, ampersands, emoji and Hebrew, and including the nested ones whose leaf name does not (byName('Sub') for Archive/Sub is exists() false).
    • byName('Archive') resolves to the top-level one, so a bare name is a path with one component. That is the whole resolution rule: Trash means the account's Trash rather than Projects/Trash, and the special mailbox wins over a same-named user folder because it is at the root, not by a special case. A leaf name is then a fallback used only when exactly one mailbox carries it (so Sub still reaches Archive/Sub); two carriers is refused with both paths named, because filing into one of two is a coin toss the response cannot show.
    • / cannot occur inside a leaf name: creating a mailbox called a/b produces a mailbox b inside a mailbox a. So paths are unambiguous and siblings are unique.
    • It is cheaper than what it replaced. Containers come back as bulk columns — mailboxes.container.name(), then .container.container.name(), until a level is entirely null — one Apple Event per level for the whole collection. Measured: 7 Apple Events / 113ms against 25 / 415ms for the old bulk-name + collection() + per-name exists() probe. mbPathOf walks a single already-resolved mailbox (a message saying where it is) at one Apple Event per level plus one; the container of a top-level mailbox answers name() with null rather than raising, which is the stop condition.

    boundByName reads the name column, the container columns and (only when it must fall back to positional specifiers) the elements, then re-reads the name column and requires it identical — the message scan's guard, one level up, closing the ~400ms window in which a mailbox created or deleted mid-read stamped rows with another mailbox's name or dropped one from every list. The exists() probes that used to sit inside that window are gone: there is exactly one per collection, on the deepest unique path, saying whether this Mail resolves paths at all, and a Mail that does not degrades the whole collection to positional specifiers rather than losing it. A collection that will not hold still across three attempts yields entries with no element, which the caller's own filter then reports as unstable — so what is named is the mailboxes in scope, not every mailbox the account holds.

    Everything a caller sees is the path: mail_list_mailboxes, every row's mailbox, moved_from, and mail_move's destination — each of which is accepted back as mailbox/source_mailbox/target_mailbox. mail_move also reads the bound destination's own path back with mbPathOf and refuses before the move if it is not what the request resolved to: the old read-back looked inside the mailbox the pick had chosen, so it could only ever confirm where the message was put, never where the caller asked. all excludes on identity (top-level only) and names what it left out in excluded_mailboxes — deliberately not in skipped_mailboxes, because those are out of scope rather than unread and a scan_complete that is false for every all scan says nothing, the same reason mail_get_source's exact boolean was removed.

  • Mail composition binds by id, then verifies. The reference mail.OutgoingMessage() returns is not reliably the message outgoingMessages.push() added: with other compose messages present it can resolve to one of those instead. This shipped a real send to the recipient of an unrelated open window. Compose therefore re-finds the message by its read-only id in the live collection, sends invisibly (visible: false — a frontmost compose window is another thing Mail can act on in place of the script's reference), and reads the recipients, the subject and the sender back off the message immediately before send/save, aborting on any mismatch. The abort names the field that differed: it used to render the two recipient lists whatever the mismatch was, so a subject carrying a CR (Mail normalises it to a space, correctly tripping the guard) aborted with two identical recipient lists printed side by side — the scariest false positive this guard can raise. Note close({saving: 'no'}) does not remove an entry from outgoingMessages; they accumulate until Mail restarts, which is why binding cannot rely on position or on the collection being empty.

  • A from no account owns is refused, not substituted. Mail does not reject an unknown sender — it sends from the default account instead — so from: "nosuch@relaytest.local" returned {"status": "sent"} and went out as From: Alice Tester <alice@relaytest.local>, Return-Path alice@, filed in Alice's Sent. A caller asking to send as one identity sent as another, with nothing in the response saying so and the message already gone. The address is now checked against every account's emailAddresses() (which is every address an account can send as, aliases included) before mail.OutgoingMessage exists, so a rejected sender composes nothing at all — and the refusal lists the addresses that would have worked. Validating up front is the request, though, not the evidence: the pre-send guard also reads msg.sender() back and aborts on a mismatch, exactly as it does for a recipient. Both compose results now carry from and account, read off the message rather than off the request, which is the only thing that says which identity a call with neither argument used.

  • Compose owns the draft Mail writes behind its back. Mail autosaves whatever it is composing. A message typed by hand has that copy removed when its window closes; visible: false plus a scripted send() meant the close never happened, so every send left a permanent full copy — body, recipients, subject, X-Apple-Auto-Saved: 1 — in the sending account's Drafts: three copies on disk per send, Alice's Drafts going 13 → 14 → 15 across two sends, in a folder mailbox: "all" excludes so no tool here would have shown a caller it happened. Four things measured against Mail 16.0 on the fixture decide the shape of the fix, and none of them is guessable:

    • send() and save() each clear the autosaved copy that exists at that moment, and Mail writes a new one a few seconds later for the message that is still open — the leaked draft's own Date header was 7s after the sent copy's. Closing stops that, and closing immediately after the send is the whole of the fix for the success path: two sends, then three more, left Drafts unchanged. The gap is what matters — putting the Drafts check between the send and the close, about a second of Apple Events, was enough to let the copy through again, once for Alice and once for Bob.
    • close({saving: 'no'}) prevents a further autosave but does not delete one already written. So an abort, which already closed, still leaks.
    • Deleting the copy of a message Mail is still holding makes Mail write another — 3s later in one run, 12s in another, 15s in a third. An abort that deleted what it found reported "that copy has been moved to Trash" and took Drafts from 21 to 22: one in Trash and a fresh one in Drafts, where doing nothing leaves exactly one. mail.delete on the outgoing message does not help, nor does setting it visible and closing it. send() is the only thing that ends Mail's interest in the message, so a leftover is removed only after one; everywhere else it is reported. There is a second reason not to delete on an abort: the guard fires because what Mail is holding is not what was asked for, which is the worst possible moment to delete on the strength of having identified something.
    • A draft saved on purpose carries no X-Apple-Auto-Saved header and an autosaved one always does (checked across the fixture's 22 drafts) — but only when compose finishes quickly. On a slow one the autosave timer fires mid-compose and mail.save() saves over that copy: reproduced with a 300,000-character body and 40 attachments (~30s), one message reached Alice's Drafts carrying X-Apple-Auto-Saved: 1 and the Message-Id the call reports as its draft, and it came back in both draft and autosaved_draft.left_in_drafts. One copy on disk, nothing deleted: a false leak in the report. So the header is not the discriminator; identity is. The saved-draft lookup has already read the draft's numeric id and RFC Message-ID off the message, and either matching disowns the entry (the numeric id dies when the account re-uploads, the Message-ID survives that). The case is named rather than hidden: autosaved_draft.saved_over_autosave: true.

    autosaved_draft is reported unconditionally on both compose tools, because a leaked draft is invisible to every other tool here and "there was none" has to be distinguishable from "nobody looked". The identity is three-part — an id that was not in that account's Drafts before the compose message was created (the snapshot is taken first for that reason), this message's subject, and the header — and the subject matched is the one Mail holds, not the one that was asked for: matching only on the request found nothing for a CR-in-subject abort and reported a leak as a clean abort. What an abort says now is what is there — never that nothing was saved. An abort is the one path that hands the message to Mail neither by sending it nor by saving it, so Mail keeps the compose message (close does not remove it from outgoingMessages) and autosaves it whenever its timer next comes round: under a second on a Mail that has been running a while, and 30 seconds on one just relaunched, which is long enough for the check to look, find nothing, and say so truthfully about a copy that then appears. Sending and saving both do release the message on a quiet Mail — five sends and a 114-second watch on a saved draft produced no copy at all — but not under load: six sends run while Mail was being driven hard each left an autosaved copy whose own Date header was 7s after the sent copy's, appearing after the sweep had already looked and found nothing. Closing immediately narrows that window, it does not close it. So found: 0 says what was seen at that moment and that a copy can still appear, in the same voice the abort uses; the difference between the two paths is now only that a send removes what it finds.

  • One fetch per message, and the bytes are kept for the next call. mail_get_email ran two osascript processes with a findMessageJXA in each: one for Mail's own properties, one for the source those properties are checked against. The second is not optional — Mail answers body: "" and has_attachments: false for a message it has not finished downloading, without complaint — so the message was downloaded on every call regardless and the split bought a second process, a second bind, and two readings of messageSize taken at two moments. Both now come out of one script: the properties are written as one line of pure ASCII (MACMCP-META: + JSON with every scalar above U+007F escaped as \uXXXX) in front of the raw bytes, which is what lets it survive the UTF-8-in/Latin-1-out decoding the message behind it goes through, and it is read after the download wait so it describes the message the bytes describe. It fails closed exactly as the size marker does. The one path that still spawns twice is the failure path: if the bytes cannot be read at all, Mail is asked for the message on its own so the answer can come back with source_check saying nothing confirmed it.

    The source is then held for 60 seconds, one message, keyed on the id and account (not the mailbox: findMessageJXA ignores mailbox for a numeric id, so mail_get_email and a following mail_save_attachment have to reach the same entry — an RFC Message-ID is searched mailbox-first, so for one the mailbox stays in the key). Only a complete source is kept, because a fragment is what a caller retries for; every mutating tool calls invalidateSourceCache() first, because after one the id may name another message. Measured on a 7,082,933-byte message: mail_get_email → mail_save_attachment went 1.15s → 0.65s, the second call 0.30s → 0.02s, and the attachment is byte-identical either way (sha256 38d3b1d3…, the same bytes Python's email module extracts from the Maildir). A small mail_get_email went 0.71s → 0.45s. body is still Mail's content() and not the source's text/plain part, deliberately: Mail renders a plain-text body for an HTML-only message and the source has none to give.

  • A move's read-back does not depend on how big the destination is, twelve times over. Verifying a move used to fetch the destination's whole messageId() column and scan it — 555ms on Alice's 11,808-message INBOX — on every one of up to twelve attempts, ~9.4s of Apple Events inside a 120s script. Two things bound it. A cheap probe (destMbox.messages.byId(n) plus fmLocate, ~15-35ms, and fmLocate because byId resolves globally so exists() only says the message is somewhere) is tried first. And the column is fetched a fixed five times, front-loaded, so the later attempts spend wall clock rather than Apple Events. The cheap probe is a fast path, not a replacement: replacing the column scan for same-account moves assumes the numeric id survives one, and it does not — measured, moving 133106 from Alice's INBOX to Alice's Archive produced 133107 there, because an IMAP re-file is a new UID even inside one account. Making it the only verification turned a 0.55s move into 3.69s: twelve failed probes and their delays before the column scan that was always going to be the answer.

  • mail_mark_read reports what Mail says afterwards. It set readStatus and returned the sentence "marked read" — a claim about Mail assembled entirely out of what the caller had asked for, with nothing read back, naming neither the message it happened to nor the mailbox it was found in. For a call whose message_id may be an RFC Message-ID resolved across every account those are exactly what the caller does not know. found is already bound by id, so it now reads its own read state, numeric id, Message-ID and location back (one Apple Event) and returns them; a flag Mail did not take is an error, not a success sentence.

  • A mutating script is never re-run. runJXAData retries by running the whole script again, and mail_move and mail_mark_read were on the default of 2 — so a -1728 raised after found.mailbox = destMbox had executed re-ran the move. Same-account that is only wasteful; across accounts the move is a re-upload, the numeric id does not survive it, so the retry's findMessageJXA returns null and the caller is told "message not found with id: N" for a move that succeeded. Both now pass mutatingRetries (0), which is what mail_send and mail_create_draft have always passed.

  • On My Mac is an account name every mail tool accepts. Mail's app-level mailboxes belong to no account, and the scan has always labelled their rows On My Mac:<mailbox> — but mail_list_mailboxes enumerated mail.accounts() only, so a caller could be handed a row from a mailbox the one enumeration tool said did not exist. Naming it was only half: resolveTargets passed the string through to a scope lookup that walks mail.accounts(), so account: "On My Mac" threw account not found. It now maps to the local pass (nil) everywhere, findMessageJXA exempts it from the account-existence check, and the listing carries it as an entry whether or not there are any local boxes. Relay's resource scoping cannot scope what the enumeration does not name, and naming something that then cannot be asked for is worse than not naming it. Note the listing shows nested folders by leaf name, so one account really can show two mailboxes called Archive.

    One account that will not hold still costs that account and nothing else. mail_list_mailboxes built each account's listing through boundByName and threw when it returned null, out of the enumeration loop — so a single busy account replaced every other account's mailboxes with the mailbox list kept changing while it was being read. This is the discovery tool: every mailbox, source_mailbox and target_mailbox a caller passes anywhere else is a string they got from here, which makes it the costliest place to keep an all-or-nothing failure, and everywhere else in this file the choice is the opposite one (boundByNameOrReport, skipped_mailboxes). Such an account is now carried with an empty list and an unread sentence saying why — absent, not null, on an account that was read, so its presence is the whole test. Naming one account still refuses: there is nothing to degrade to, and an empty list for it would read as "this account holds no mailboxes".

  • Mail destinations resolve inside one account. Every account owns an Archive, Drafts, Sent, Trash and Junk, so resolving a mailbox by name across all accounts returns whichever account Mail lists first — which has nothing to do with the message. mail_move therefore resolves target_mailbox inside foundAccount (where findMessageJXA located the message) and refuses rather than borrowing another account's mailbox of the same name; crossing an account boundary requires an explicit target_account. It then reads the message back out of the destination by RFC Message-ID, because moved on its own says nothing about where it went. Note the reference is dead the instant msg.mailbox is assigned — every property read on it afterwards raises "Invalid index" — so identifiers must be captured before the move, and the numeric id does not survive an IMAP re-file.

  • A cross-account move is a re-upload of Mail's copy, not a server-side move, and target_account says so. Measured against the fixture's Maildir: a 254-byte-value probe moved between accounts arrives at the destination byte-for-byte identical to what mail_get_source returns for it — LF line endings, the NUL gone — which is the proof that Mail uploads what it holds rather than asking the servers to copy. Headers (including Return-Path), content and the RFC Message-Id all survive; the numeric id does not. A 2 MB message Mail held only as 271.partial.emlx (headers only) was uploaded in full, so a partial download is not a truncation risk. An earlier report of dropped Delivered-To-class headers did not reproduce: the 11-byte delta it described is what CRLF→LF accounts for on a message with 11 line breaks, and the destination copy matched the source header for header once line endings were normalised.

  • A script that throws is reported as the sentence it threw. osascript wraps a thrown value in its own text -- execution error: Error: Error: account "Alice" has no mailbox named "BobOnly" (-2700), doubled Error: and an OSStatus included -- and that reached callers verbatim from every path that refuses by throwing: mail_move's missing destination mailbox, and account not found from mail_move, mail_mark_read, mail_get_email and mail_get_source. scriptErrorMessage unwraps it in runJXAData, so scripts stay free to throw (the natural thing from inside an IIFE) and every one of them, including ones written later, comes out in the same voice as {error: ...} results. Only -2700 is unwrapped -- that is osascript's code for "the script threw", so the text is ours; -1712, -1728 and syntax errors keep their raw form because the number is the evidence. Which code it is is read at the position osascript writes it -- the trailing (-NNNN) -- never searched for in the text. Everything before that position is the message, and a message contains whatever the caller passed in: mail_move on a mailbox named Q (-1712) box reported "Mail timed out evaluating the request (-1712)" while the script had thrown no mailbox named "Q (-1712) box", and -1728 in a name bought two silent retries.

  • A mail timeout is checked against TCC before Mail is blamed. A consent-blocked osascript is indistinguishable from a wedged Apple Event from the outside, and the old message answered both with "narrow the scope" — advice mail_list_accounts (which takes no scope at all) cannot act on. runJXAData now takes the automation grant with AEDeterminePermissionToAutomateTarget(..., askUserIfNeeded: false) before running the script, and scopable says whether the calling tool has anything to narrow. The order matters: that check answers in ~10ms normally but blocks while a consent prompt is on screen (measured 12s and 73s, still blocked 20s after the script had been killed), so it is bounded by a 2s deadline and a blocked check is itself reported as a pending decision rather than as ignorance. permissions_check reports automation (Mail) for the same reason — it is the one TCC service macMCP hangs on rather than merely being refused by.

  • A deadline belongs to the call, not to the spawn. Per-spawn deadlines compose: a scan runs one osascript per account, mail_search runs two of those passes and then a body pass, and each was handed the full 120s independently. Read off the code that made mail_get_emails 406s at two accounts and 1398s at ten, mail_get_email 421s, mail_move 372s, mail_create_draft 308s, and mail_search with search_body about 1408s — 23.5 minutes, during which main.swift's single synchronous readLine() loop serves nobody else. MailCall is one wall-clock budget per tools/call; every runJXAData beneath a handler gets min(remaining, its own ceiling) via spawnAllowance, and running out degrades rather than raises — the accounts already read come back, the rest land in failed_accounts with the reason, scan_complete goes false and note says the counts are a floor. The defaults (MailService.Budget) are each the measured steady-state worst case times a contention allowance of 8, capped at 300s; the allowance is not padding, since the same script measured 2.16s alone and 17.58s with another client driving Mail. timeout_seconds is on every mail tool because which tool is the slow one is a property of the machine, not of the schema. Live: mailbox: "all" at timeout_seconds: 5 returns its rows in 3.9s with On My Mac named as unread; a full-scope mail_search with search_body under sustained Apple Event load returned in 20.04s at timeout_seconds: 20 and in 256s on the 300s default, both with honest coverage.

  • A refused Apple Events grant is diagnosed by a spawn and then latched. The pre-run TCC probe is a prediction; what settles it is osascript exiting with -1743, which is not -1712/-1728/-2700, so scriptErrorMessage declined and the raw line reached the caller. Measured against a real deny row (staged on an unrelated target, never on Mail): execution error: Error: Error: An error occurred. (-1743), exit 1 in 0.75s — so the caller was told "An error occurred" while the sentence written for exactly this case sat in a jxaTimeoutMessage branch only a timeout could reach, and a denial does not time out. -1743 is now recognised and answered with that sentence in 0.16s. The probe itself is read once per call, not once per spawn (it is a property of the process, and its 2s bound was being paid seven times over by a full-scope search), and a status in doubt — pendingConsent, checkBlocked, denied — cuts the first spawn to a 20s window, after which the reason is latched on the MailCall and every later script is refused without spawning. Not short-circuited before the first spawn on purpose: a stale probe must never be able to make Mail unreachable, and the first spawn costs ~150ms to be certain. targetNotRunning is excluded from the window — sending the event is what launches Mail.

  • The script reaches osascript as a file, not as -e. An -e argument is under ARG_MAX (1,048,576 here), and escapeJSString renders every non-ASCII UTF-16 unit as six ASCII bytes, so the ceiling was ~1 MB of ASCII but only ~175 KB of Hebrew. A mail_create_draft with 300,000 em dashes came back as "failed to run osascript: The operation couldn't be completed. Argument list too long" — an error naming an implementation detail rather than the body, for a limit no schema documented. As a 1.8 MB file the same script runs in 40ms; the draft now lands with all 300,000 em dashes intact in the Maildir. The one thing it changes is stderr: osascript prefixes its error line with the path it read, and scriptErrorMessage requires the text to begin with execution error: , so without stripScriptPath every thrown sentence would have stopped being unwrapped. Stripping is exact (the path is a fresh UUID under $TMPDIR), never pattern-based.

  • stderr is bytes with a fallback, and the wait is not a flat tick. stderr used to be read as UTF-8 only, so one invalid byte collapsed the whole of it to "" — which costs more than the prose: osaStatus("") is nil, so the -1712, -1728, -2700 and -1743 branches are all skipped and the caller gets "osascript exited with status 1" with the OSStatus the comments call "the evidence" gone. decodeStderr falls back to Latin-1, which cannot fail, as MIME.decodeString already did for message bytes. Separately, the wait for the child polled in flat 0.05s ticks, adding up to 50ms of pure sleep to every spawn — one per account per pass. It now starts at 1ms and backs off: measured 138ms per trivial spawn before, under 100ms after.

  • A TCC probe that overruns its deadline does not stop. boundedStatus gives up waiting; the probe is still inside AEDeterminePermissionToAutomateTarget, which returns only when the prompt is answered (12s, 73s, still blocked 20s after the script was killed). It used to sit on a DispatchQueue.global thread — a pool bounded at 64 that WebService, WeatherService, LocationService and the EventKit services all need for completions delivered under CFRunLoopRunInMode, so enough abandoned probes hang tools with nothing to do with Mail (200 of them starved it in a test). It now gets a dedicated Thread, hands its result over under a lock rather than through a bare var (the deadline path is a real data race, not a benign one), and maxConcurrentProbes caps it at 4 — past a handful there is nothing to learn, because a probe already blocked is the evidence .checkBlocked reports.

  • A revoked Apple Events grant cannot be restored by putting the database back. After a consent prompt was left to time out, tccd had written a deny row (auth_value=0, auth_reason=9) for the client, and it re-asserted that row over a restored copy of TCC.db twice. Copying the file back, even with tccd stopped, does not converge — the daemon's own view wins. What worked was raising a fresh prompt and approving it (vmallow watches for it; the prompt is owned by UserNotificationCenter, not by Mail or by the requesting app). So: never test consent by letting a prompt expire, and if a grant does go missing, re-prompt rather than reach for the database. Note also that ad-hoc signing pins a grant to the cdhash, so rebuilding Relay revokes macMCP's grants and they have to be re-granted through Relay > Settings > MCP Servers > macMCP > Reset Permissions. Verify functionally (mail_list_accounts returning accounts), not by reading a status field.

  • Mail escapes every non-ASCII character as \uXXXX when generating JXA. The reason it was written was the -e argument, decoded using the process locale, which the MCP server's host need not set — raw UTF-8 em dashes and Hebrew came out mangled. The script is a file now and a script file is decoded as UTF-8 whatever the locale says (verified under env -i and LC_ALL=C), so that hazard is gone; the escaping stays because U+2028/U+2029 terminate a JS string literal, quotes and backslashes need escaping regardless, and printable ASCII means nothing about how osascript reads a file can change what the script says.

  • Mail's own attachment APIs are unusable. save on a mail attachment fails with -10004 for every destination including ~/Downloads (Mail's sandbox, not Full Disk Access), and the MIME type property raises "AppleEvent handler failed" on any message that has an attachment. source works, so mail_save_attachment fetches raw RFC 822 and decodes it in MIME.swift — and mail_get_email now takes its mime_type from the same place, so the two agree. The filename guess (UTType) survives only as a fallback, and mime_type_source says which one a caller is looking at: guessing from the extension reported text/csv for a part the message declares as image/png; name="data.csv". A failed fetch leaves the guess in place rather than costing the caller the message.

    The fetch used to be skipped when Mail listed no attachments, which is exactly what a message Mail has not finished downloading looks like: content() returns '' and mailAttachments() returns [], without complaint. Severing the fixture's IMAPS proxy mid-fetch produced body: "", has_attachments: false, attachments: [] and no error for a 400 KB message carrying one attachment — beside a correct message_size: 400595 in the same response. So mail_get_email now checks every message against its own source and withholds negatives it cannot stand behind: an empty body and an empty attachment list are omitted (listed in omitted, with fidelity saying why) rather than returned as "" and false. Positive evidence — a partial body, an attachment Mail has already listed — is kept.

    Mail's attachment list is not authoritative even after the message arrives. The severed message's list stayed empty permanently once the download finished (mailAttachments() = 0 against a 400574-byte source declaring one), while mail_save_attachment extracted the attachment byte-exactly. The list is therefore reconciled with the message source, and anything the source declares that Mail does not list is added with listed_by_mail: false. Inline parts are excluded — Mail deliberately does not list a body image, and turning has_attachments true for every HTML message with a logo would be a new wrong answer in place of the old one. Note Mail's own list does not honour that rule (it listed a Content-Disposition: inline logo, and a CID-only PNG as "Mail Attachment.png"), so inline-ness is read off the message and never off whether Mail listed it.

  • There is one attachment list, and both tools work from it. mail_get_email used to report Mail's mailAttachments() rows with source-declared extras appended, while mail_save_attachment indexed MIME.attachments(of:) straight — different membership, different order, different names, with attachment_name documented as "as reported by mail_get_email". On a probe carrying an HTML body, a CID-only inline PNG and a report.txt: mail_get_email said [report.txt, Mail Attachment.png], index: 0 wrote the inline body image and reported success, and attachment_name: "Mail Attachment.png" — the name just handed to the caller — was rejected outright, naming two names it had never shown. MailService.attachmentList is now the single source: derived from the message, in document order, inline parts split off. index is an index into it, attachment_name is one of its names, and an unqualified save writes exactly what mail_get_email listed. An inline part is still reachable, by part_path and only by part_path, so a caller who wants the logo can have it without every HTML message growing an attachment.

  • The identity of an attachment is its MIME part path, not its filename. mail attachment.id is the part's position in the message — 2, 3, 1.2, the numbering IMAP BODYSTRUCTURE uses. Measured on Mail 16: three attachments of a flat multipart/mixed came back 2, 3, 4; an inline image inside a multipart/related that is part 1 of a multipart/mixed came back 1.2 in a message four levels deep. Reconciling on the filename instead emitted one part as two attachments whenever Mail rendered the name differently from the header, and the Mail-derived copy lost the declared type for one guessed off the extension — text/csv for a part headed image/png, which is verbatim the bug mime_type_source exists to prevent. Four independent triggers, all confirmed live and all fixed: an escaped quote in a quoted-string filename (MIME.parse stripped the quotes without undoing \"; MIME.unquote now does, and splitOutsideQuotes always handled escapes, so only the unquote was missing); a raw non-ASCII filename, where Mail returns its own Latin-1 mojibake — verified through plain JXA, so macMCP is relaying Mail faithfully and no decoding here will ever make the two strings equal; no filename parameter at all, where Mail invents "Mail Attachment" and the source has nothing to invent from; and a "/" in the filename, which Mail sanitises and the source keeps. A position is not a rendering of anything, so none of the four moves it. Name and size survive as later passes for a Mail that reports no id, and the size pass takes a match only when exactly one unclaimed part has that size. What the result reports is the message's name (that is the handle) with Mail's under mail_name when they differ (that is a label), and a row Mail lists that matches no part goes to attachments_mail_lists_only rather than being turned into an attachment with no bytes behind it.

  • A half-written save says what it wrote. mail_save_attachment returned a bare error sentence on the first failed write and discarded the saved array, so files already on disk were invisible to the caller who had just been told the call failed. The failure is still a failure — isError stays true — but it now carries saved and a count.

  • source() arrives one encoding layer removed. Mail builds the string for found.source() by decoding the message's raw bytes as ISO-8859-1, and osascript writes that string to stdout as UTF-8, so every byte above 0x7F comes back UTF-8 double-encoded. decodeSourceBytes undoes it (UTF-8 in, Latin-1 out) and also drops the single newline osascript appends after any result. Both halves are guarded: invalid UTF-8, or a scalar above U+00FF (what a Mail that decoded the source correctly would emit), returns the bytes untouched rather than mangling them again. Base64 attachments hid this for a long time because base64 is pure ASCII; the bug only shows on Content-Transfer-Encoding: 8bit.

  • A fetched source is still not byte-identical to the message, and says so. Measured against the fixture's Maildir with a message carrying every byte value except CR/LF: 253 of 254 round-trip, but a NUL comes back as 0x80 and every CRLF comes back as LF (890 bytes on disk, 869 returned, 21 CRs gone, one 0x00 arriving as 0x80). Both happen inside Mail — plain JXA emits NUL and CR fine, Swift's UTF-8/Latin-1 round trip preserves them, and Mail's own .emlx copy already holds 0 CRs and 0 NULs — so nothing here can undo them. sourceFidelity reports them instead, as a fidelity object on mail_get_source (both the inline and save_to paths) and on mail_save_attachment. The NUL case is ambiguous, not merely lossy: a returned 0x80 is either a real 0x80 or a lost NUL, so the count of candidates is reported rather than a claim about which — but only a 0x80 that is not a UTF-8 continuation byte counts, because Mail's replacement for a NUL always lands standalone and an em dash (E2 80 94) is not one. Counting every 0x80 reported three lost NULs for a body whose only sin was typography. source_encoding is not the place for this — it describes how the inline string was encoded for return, exists only on that path, and a source can be valid utf-8 and still be missing a NUL.

    There is deliberately no summary boolean. exact used to be one (complete + CRLF + no 0x80); Mail strips every CR, so it was false for every real message — including one whose fetched bytes matched the copy on disk exactly — and true only for data the pipeline cannot produce. What the object carries is facts that can go either way (complete, line_endings, ambiguous_nul_bytes, bytes_measured, message_size) plus a note, and the counts say they were measured over the whole source rather than over whatever slice max_bytes returned.

  • source() returns what Mail has downloaded so far. For a message still arriving that is the headers and a fragment: 838 bytes of a 300 KB message in one measurement, with truncated (which describes the max_bytes slice) saying false. Every consumer of the raw source inherits it — a truncated attachment on disk, a mime_type guessed from a filename, a body_check against a body that has not arrived. The fetch therefore reads messageSize (a bulk column, one Apple Event, giving the wire size) and waits, up to 10s, for source.count + LF count == messageSize — exact rather than heuristic, because every LF in a returned source stands for one CRLF on the wire. What is left is reported: fidelity.complete, and mail_save_attachment refuses rather than cutting a file out of a fragment. The size travels back on a MACMCP-SIZE:<n> line ahead of the source, stripped only when it matches exactly, because the source itself is raw bytes on stdout and a second osascript spawn would cost more than the fetch.

    messageSize is quoted in one of two units and Mail does not say which. A message the server holds is quoted in wire units — 375 bytes with 19 line breaks reported as 394. A local draft is quoted in the units Mail stores it in: bytes_measured: 1362 against message_size: 1362, matching the Maildir's S=1362 and not its W=1395. Counting every LF as a CRLF is what makes the first case come out right and is exactly what hands the second slack: a 1362-byte draft passes at 1362 + 33 >= 1362, and so would a fragment of it 33 bytes short. complete_basis says which reading complete rests on — bytes when the bytes reach the size on their own (assuming nothing), wire when they only reach it once each LF is counted, plus short, unchecked, none — and the note quantifies the slack in the wire case. Requiring an exact match on one of the two readings would close the hole and is deliberately not done: it turns any imprecision in messageSize into a permanent false incomplete, which costs a caller mail_save_attachment entirely.

    A CRLF that survived is not counted twice. The wire size is the bytes plus one CR for each bare LF, because a bare LF is what a CRLF came back as; a CRLF still in the bytes already weighs two and gets nothing added. Adding one for it too inflated the wire size by a byte per line — slack in the one direction this guard must never have, since wireSize >= messageSize is what promotes short to wire, and complete is what mail_save_attachment cuts a file on the strength of. Mail strips every CR today so nothing measured it, but line_endings models crlf and mixed as reachable and the arithmetic has to be right when they are.

    The closing MIME delimiter closes the hole instead, for free. The wire slack is one byte per line break, so it scales: 919,823 bytes on a 70.8 MB message, i.e. a fragment nearly a megabyte short passes. RFC 2046 requires a multipart body to end --<boundary>--, which is a structural end-of-message marker no fragment carries whatever its bytes add up to, and the bytes are already in hand. So on the wire reading only — never on bytes, where nothing is being assumed — a multipart is additionally required to end there: complete_basis is wire+terminated when it does and unterminated, not complete, when it does not. Measured against a truncated multipart delivered into Bob's Maildir: before, complete: true, complete_basis: "wire", bytes_measured: 69618, message_size: 70533 (915 bytes of slack, exactly the CRLF count) and mail_save_attachment wrote a 51,186-byte file for a 102,400-byte attachment; after, it is refused, while the intact copy reports wire+terminated and still saves all 102,400 bytes. The one message this can be wrong about is a sender that omits the closing line from a message that is all here, which RFC 2046 does not allow; the note says so, and mail_get_source still returns what there is.

    Two edges of that. Zero bytes is not a message — severing the fixture's IMAPS proxy mid-fetch produces exactly that, source() returning '' for a message Mail sizes at 400595 — so an empty source is never complete, is waited for even when messageSize could not be read, and mail_get_source errors rather than returning "source": "" with truncated: false. And messageSize raising is not evidence of anything: complete is then "nothing contradicts it" rather than a verified match, which fidelity.message_size: null and the note both say, since that is the one path left by which a fragment could pass for a message.

  • The MIME reader is bounded, and says where it stopped. Nesting is chosen by the sender, and MIME.parse used to recurse with no depth bound: a 929 KB message nested ~13,000 multipart/mixed levels deep exhausted the 8 MB main-thread stack and killed macmcp with signal 11 — no response, no error, and every tool gone with it, macmcp being one synchronous stdin loop. Reachable from mail_get_email and mail_save_attachment, both of which parse raw source; mail_get_source was unaffected because it does not parse. Measured against the release build: 13,000 levels survived, 40,000 exited 139. The parse is now iterative over an explicit work list rather than recursive-with-a-counter — one descent, one place the limits live, and stack usage that does not vary with depth — with attachments(of:) and firstPlainTextPart rewritten the same way, since each was a second walk over the same sender-chosen depth. Three ceilings, all reported: maxDepth 32 (multipart/signed over mixed over related over alternative over text/html is the deepest real shape at 5, and forward chains add none because nothing descends into message/rfc822), maxParts 10,000, and maxHeaderBytes 256 KB (a message with no blank line is all headers on purpose, which is what lets a sender make the header block the whole message). A part past the cap is kept unparsed rather than dropped or invented, and structure — parsed_complete, parts, depth, plus a note — rides on mail_get_email and mail_save_attachment beside fidelity, because the visible effect of a limit is a shorter attachment list, which is indistinguishable from a message with fewer attachments. mail_save_attachment refuses rather than reporting "has no attachments" for a message it could not read to the bottom. Note the cap is not a stack guard any more; it bounds work, since every level holds its own copy of the bytes below it.

    A boundary that never appears is not an empty message. A multipart/* whose declared boundary does not occur in its body — a sender that got it wrong, or bytes cut before the first delimiter — yields no pieces, and the parse then cleared the part's body (correct only when the children hold those bytes, and there were none) while leaving parsed_complete true. So the one flag mail_save_attachment reads to choose between "no attachment could be read out of this message" and the flat "this message has no attachments" said the second about a message none of whose content had been read. The bytes are kept, the part is marked unparsed, and the report says so. A multipart that really holds no parts is a different message and stays complete: its close-delimiter is there and was read, which is why the split now reports whether it saw a delimiter at all.

  • MACMCP-SIZE: is macMCP's own line and is always at offset 0. sourceScriptJXA writes it before a single byte of the message on every path, so splitSourceSizeMarker failing open — leaving the line in the bytes when the value on it did not parse — could never have protected a caller's own first line, which is never at offset 0. What it did instead was put MACMCP-SIZE:null into save_to files, bytes_total, sourceFidelity's counts, and MIME.parse as a bogus macmcp-size: header. It now fails closed: the line is stripped whatever is on it, -1 is accepted because the script writes it itself (the documented "messageSize() raised"), and any other value is an error naming what was seen — the wait that decides whether the whole message arrived ran against that same value, so nothing can say whether the bytes are the message or a fragment.

  • A multipart ends at its close-delimiter, wherever that sits. The terminator check that closes the wire slack (a fragment nearly a megabyte short of a 70 MB message passed as complete) first required --<boundary>-- to be the last thing in the source, stepping over trailing whitespace only. RFC 2046 §5.1.1 permits an epilogue after it, and such a message read as truncated: complete_basis: "unterminated", complete: false, and mail_save_attachment refusing outright — a legal message losing the one tool whose job is producing its attachment, under an error that told the caller RFC 2046 forbids what it in fact allows. The delimiter is now looked for at the start of a line anywhere in the last MiB, which is both the correct question (a close-delimiter at a line start is where the multipart ends) and a simpler one. The bound is deliberate: further back than that, a close-delimiter is not distinguishable from a message that never had one, and a whole-buffer walk to answer a tail question on a 70 MB message is work nobody asked for. Verified both ways against the fixture — a message with a legal epilogue saves its attachment byte-exactly, and a genuinely truncated multipart is still refused with nothing written.

  • Mail rewrites any body set through its scripting interface, and there is no way around it. Whatever is given as content (or html content) arrives inside <blockquote type="cite"> under Mail's Apple-Mail-URLShareWrapperClass scaffolding: a sent message's text/plain alternative gets > on every line, and a saved draft's text/plain part comes out empty. Ruled out on Darwin 27 / Mail 16.0 (3864.500.181), each verified against the fixture's Maildir: setting the body at creation, setting it after resolveOutgoingJXA, visible: true, html content, injecting closing tags to escape the blockquote (WebKit rebalances them), SendFormat = Plain with a Mail restart, and textbook AppleScript make new outgoing message with properties {content:…} — which reproduces it exactly, so this is not something macMCP's compose path chose. A message typed by hand in Mail comes out as a clean single-part text/plain, so it is specific to the scripting path. mailto: compose windows never appear in outgoing messages, so that route cannot be driven; content.paragraphs.push(…) kills osascript (SIGKILL, no output). What is left is not to lie about it: mail_create_draft re-reads the saved draft and reports body_check, because rendered_chars is measured off msg.content() before Mail generates the alternatives and will happily report a plausible number for a message whose plain part is empty. rendered_chars also counts whitespace-stripped characters — "just one line here" reports 15, not 18 — which is right for the question it answers ("did Mail render anything visible") and wrong for the one a caller is likely to ask it ("is this my body's length"), so both compose schemas now say which it is.

    body_check is behind verify_body and defaults to null, not false. It is a guard documented to fire on 100% of calls — the sentence above says the plain part comes back empty for every scripted draft — and it costs a full download of the saved draft to report that constant, which on a large draft is the whole draft. So the default says plain_text_matches_body: null, measured: false with the known behaviour in the detail. Null rather than a hardcoded false because a hardcoded claim becomes a confidently wrong answer the day a Mail release stops rewriting the body, which is precisely what this check exists to catch; verify_body: true takes the measurement.

  • HTML bodies use html content, which Mail's dictionary marks hidden and "does nothing at all (deprecated)" but which in fact still renders (verified on Darwin 27 — produces multipart/alternative with a Mail-generated plain-text part). It wins over content when both are set, so only one is ever sent. Compose checks the rendered body and errors rather than silently shipping an empty message if a future Mail makes good on the deprecation.

  • A message id that arrives as a number is still a message id. The schema said string and every handler read .stringValue, so a client emitting "message_id": 63926 unquoted was answered with message_id is required — a message about a missing argument, for an argument that was right there. Models write an id that looks like a number as a number, and several clients pass it through that way (LM Studio is the one this was found on). Both halves had to move: stringOrIntProp so a client validating the schema will send it at all, and coercedStringValue so the handler reads either rendering. It also unwraps a value the client quoted twice ("\"63926\"" decodes with the quote characters in the string) — safe because no id here begins and ends with a quote, an RFC Message-ID being delimited by angle brackets — and renders a whole double as 63926 rather than 63926.0, which would match no message at all.

  • A mail call is confined by the _meta it arrives with, and _meta present is what makes it confined. Relay injects _meta.project_id on every mediated call (ADR-007), so the presence of _meta at all is the signal that a chokepoint mediated this one — and MailScope.isScoped is that, not "did it carry one of my three restrict keys". The obvious reading fails open: relay failing to inject a field produces a call indistinguishable from an unmediated one, leaving relay's own call-time check as the only defence. So a mediated call whose profile sets no mail scope loses every mail_* tool until an operator sets one, and no _meta at all — an operator on a bare stdio pipe, which is same-user access equivalent to opening Mail.app — behaves exactly as macMCP always has. _meta present but not an object is a malformed mediated call and is refused (-32602, in main.swift, before dispatch): req.params?["_meta"]?.objectValue is nil for null, an array, a string and a number alike, so "_meta": null took the unmediated branch and was answered with every account, every mailbox and mail_send. Only a genuinely absent key means nobody mediated. Relay always sends an object, which is exactly why this had to be closed here rather than left to relay — a check that holds only because of what the other side happens to send is one check, not two, and that independence is the whole of ADR-011 decision 4.

    Which fields govern a tool is read off macmcpContextSchema's own applies_to (restrictFieldsGoverning), never re-typed per handler. Every handler used to name the fields it checked as accountKeys/mailboxKeys arguments, and the ones with nothing to reconcile passed [] and checked nothing: mail_list_mailboxes — the discovery tool — listed every mailbox on the machine to a mediated call carrying no mail_mailboxes, and mail_send, mail_create_draft and mail_list_accounts never consulted that field at all. Belt and braces, allowedList maps a field with no value to the empty list rather than to nil: nil means "no restriction" to every generated script, so one forgotten call site could widen a confinement into its opposite. An operator field (source: "operator") with no value refuses the whole tool; a project_path field governs the parameter instead — a distinction only macMCP can make, since relay's checkScopePresence reads applies_to as a tool-level denial and has no notion of an argument. file_dirs names a tool in applies_to only when the tool cannot function without it: mail_save_attachment's destination is required, so it is named and relay denies it outright with no value, which is ADR-011 finding 1's intended outcome. mail_get_source's save_to, and mail_send/mail_create_draft's attachments, are all optional — each tool works with the field unset — so none of the three is named, and each is confined at its own parameter (writeDestination, writeDestination, readableAttachment) instead. macMCP enforcing a parameter relay does not advertise is the fail-CLOSED direction of a mismatch between the two declarations; ADR-011 decision 9 forbids only the opposite one, where relay advertises a confinement that is not real.

    The reconciliation rule is stated once and applied identically at every seam (MailScope.accountTargets / mailboxTargets / writeDestination / readableAttachment): an absent or default-valued argument resolves to the scope (so a scoped mail_get_emails with no mailbox reads what the client may reach, not an INBOX that may not be one of them); a tool-level wildcard (mailbox: "all") resolves to the scope and does not error; an explicit argument outside it is an error, never a silent narrowing, because silent narrowing lets an agent build a false model of what it can reach and burn calls discovering the truth. Under a scope, all excludes nothing — a profile that names Trash was given it deliberately, and subtracting it would overrule the enumeration an operator typed. Values are mailbox paths, matched case-insensitively as every caller-supplied name in this file is; a scope resolving names by a different rule from the one the tools resolve arguments by is how an operator allows a mailbox the client cannot reach.

    Where the check has to happen is decided by how the thing is reached. The scan intersects the allowed paths with the mailbox list inside the generated script, before a single message column is fetched — post-filtering rows still reads every subject of every mailbox the caller may not touch. findMessageJXA is necessarily post-hoc, because byId resolves globally, and both halves are compared: the mailbox name alone admits another account's INBOX, which is verbatim a bug that already shipped here. mail_move checks the source off the message and the resolved destination (account and full path) at the last moment before found.mailbox = destMbox, since a leaf name and a defaulted target_account sit between the string the caller wrote and the mailbox that would be written to. senderJXA asks a second question beside ownership — Mail owning alice@ says the message can go out as it, not that this client may be Alice — and with neither from nor account it resolves the identity to the scope instead of letting Mail use its default account, which would have sent as someone else and reported {"status": "sent"}. The enumerators are scoped at both levels: mail_list_mailboxes is where every mailbox argument a caller ever passes comes from.

    A refusal raised inside generated JavaScript carries MailScopeRefusal.sentinel, which MailService.mailError strips back off and turns into _meta.scope_violation: true on the result — {"content": [...], "isError": true, "_meta": {"scope_violation": true}}. Relay surfaces it as an audit field with outcome staying tool_error (ADR-011 decision 7): the call completed and the MCP answered no. Absent rather than false when nothing set it. Two things that are deliberately not violations: a scope naming a mailbox Mail does not hold is an operator's typo, reported as an ordinary error (and reported at all — all resolving to nothing would otherwise return total_messages: 0 with scan_complete: true, an affirmative claim that the mailboxes the client may read are empty); and a message id nothing carries stays "not found", because reporting that as a refusal would tell a caller a message exists when it does not. A message that does exist and is out of reach is a refusal, worded so it never says where it is.

    The held source is keyed on the confinement (MailScope.cacheFingerprint in sourceCacheKey). A cache hit hands back bytes without running a script, and the script is where the scope is checked — two differently-scoped callers reaching one macmcp process would otherwise let the second be served a message it may not read, with nothing anywhere having decided that. Each value in that key is length-prefixed, not joined with a separator: ["a,b"] and ["a","b"] spelled the same fingerprint, and no character can be reserved because a mailbox path may contain any of them.

    file_dirs is ADR-011's write_dirs, and it is the client's filesystem foothold in both directions. It governed destination and save_to while mail_send and mail_create_draft took attachments: [absolute POSIX paths] and read whatever was named straight off the host into an outbound message, scoped by nothing — ADR-011 finding 1 on the read side and worse, since finding 1 was an arbitrary host write and this is an arbitrary host read wired to a channel that leaves the machine (mail_send {"attachments": ["/tmp/zsec-secret.txt"]} through the live write profile arrived base64'd in Alice's .Sent Maildir). One axis rather than two, because what an operator decides is which directories a client may touch. A client with no file_dirs may still send — it just may not attach a file off this host. The path check resolves symlinks component by component, not only when the whole path exists, because a destination normally does not exist yet (the tool creates it) and a symlink inside an allowed directory would otherwise let a write escape through a purely lexical comparison — the gap in fsMCP's validatePath, which this otherwise follows exactly. A bound that is not an absolute path is refused rather than resolved: the walk seeds /, so ".", "" and ".." all answered /, which is a prefix of every absolute path — one such entry made the whole check a no-op that still reported a confinement (file_dirs: ["."] let mail_get_source write to /tmp/zoutside/…). That is an operator's mistake, so it is an ordinary error and never scope_violation.

    Every comparison of an account name or a mailbox path goes through one fold: NFC, lowercased (MailScope.fold in Swift, scopeFold in every generated script). Swift's == compares by canonical equivalence and JavaScript's === by UTF-16 code units, so an accented or Hebrew name spelled NFC on one side and NFD on the other passed the Swift front door and was rejected by every JS seam behind it — a correct scope silently reaching nothing, with no violation logged because none had occurred.

    A leaf name cannot answer a scope, so a failed path walk refuses. fmLocate degraded mbPathOf returning null to the mailbox's leaf name and scope-checked that: an ancestor renamed or deleted mid-walk let an out-of-scope Projects/Archive satisfy a scope of ["Archive"]. The leaf is still the label reported when there is no mailbox scope to check against. It is answered as a transient error the caller can retry, not as a violation — nothing was probed, the check simply could not be made.

    A mailbox argument means the same mailbox to a move as to a read. mail_move's resolution took an exact path anywhere before a leaf name in scope, so under mail_mailboxes: ["Projects/Archive"] a mailbox: "Archive" read the nested one while target_mailbox: "Archive" resolved the top-level Bob:Archive and refused naming a mailbox the caller had never asked about — after mailboxTargets had already accepted the name. The scope filters the mailboxes first, at both steps, in both tools; an exact out-of-scope path resolves last, and only so the refusal names what the caller wrote instead of a "no such mailbox" that is untrue.

  • The scope mechanism is ResourceScope, and mail is one service using it. What was inside MailScope was never about mail: the three-state Access, isScoped meaning _meta was present at all, Decision with its misconfigured outcome, fold, the component-by-component path walk, the length-prefixed cache fingerprint, and the presence check driven by applies_to. It lives in ResourceScope.swift; MailScope is a typealias for it and MailScope.swift holds only which _meta keys a mail scope reads and how the reconciliation rule applies to a mail argument (575 lines to 296). A wrapper struct was rejected: there is one contextSchema, so a call carries one scope, and forwarding parse/isScoped/Access/Decision/fold/cacheFingerprint/.none by hand would be more mail-shaped code than the extraction removed, each forward a second place the rule could drift.

    A service declares a ScopeField — name, noun, operator-facing description, source, applies_to, depends_on, and an enumerator — and gets the machinery. enumerable is derived from carrying an enumerator rather than declared beside one, so "declares enumerable but does not implement it" is not a reachable state; the default: branch that used to answer exactly that is gone. Nine fields are declared: mail's three, plus calendar_accounts/calendars, contact_accounts/contact_groups, reminder_accounts/reminder_lists, each pair depends_on its account field. Five of the six govern their own <service>_* glob; contact_groups does not, for the reason below. All nine are enforced; the six were for one branch declared ahead of their enforcement — presence goes live the moment relay sees a field, the value half does not — which is ADR-011 decision 9's forbidden direction, and it is the reason a field and its enforcement land in one change. ContextSchema.swift's header is the single place that says what is enforced, so there is one such claim to keep true rather than a note per service drifting out of date.

    A calendar, a reminder list and a contact group are identified by Account/Name, never by a bare name (ScopePath). EKCalendar.title is not unique and CalendarService used to match with $0.title == name, which returns both when two sources hold a Work — the same "a leaf name does not identify a thing" bug the mailbox path work fixed, one framework over. A contact group is Account/Name rather than CNGroup.identifier deliberately: the UUID is more precise but unreadable, so a profile scoped by it cannot be reviewed and an audit line naming three UUIDs answers "was this call confined?" with a question — and it is re-issued when a container re-syncs, which fails silently. Unlike a mailbox path, the value is matched whole and never split, because a / can occur inside a calendar title or a group name; two rows generating one string are reported by ScopePath.ambiguousValues on the same footing as two resources genuinely sharing a name.

    A path is not by itself an identity, and one value admitting two resources is the same bug one level down. ScopedRows.allowed walked the rows and kept every row whose folded path matched a value, so a value naming two resources admitted both — and the operator saw one picker entry, because ScopePath.entries dedupes by path. Reproduced: two reminder lists both called ZSECDUP under Default, reminder_lists: ["Default/ZSECDUP"], and reminders_list returned the contents of both with every row labelled list_path: "Default/ZSECDUP"; two Work calendars in one account (an import, a shared-calendar subscription) do the same. It now walks the values and counts their carriers: more than one carrier admits none, and the whole call is .misconfigured rather than the bad value being dropped from an otherwise usable scope, because a profile that reads as granting three calendars and silently grants two is a confinement an operator cannot review. .misconfigured and not a violation — nobody probed, two resources were given one name. ContactScope.select has answered this way since it was written, and two services disagreeing about what one permission value means is worse than either answer. The advice on the read side changed with it: ambiguityMessage told a caller to "narrow the scope to the one this client should reach", and there is no such scope, so it now says to rename one of them in the app that owns it.

    Calendars and reminder lists enforce through one seam, ScopedRows (Services/EventKitScope.swift), which is to EventKit what MailScope.swift is to Mail. It is a pure function from [ScopePath.Row] plus two Access values to a set of row indices, and everything framework-shaped is one line per handler (store.calendars(for:), mapped positionally). Indices rather than paths, because two calendars can generate one Source/Title string and a seam returning strings would have to guess which object one meant — the bug the representation replaced. The two fields AND together, as ADR-011's worked example says: calendar_accounts: [iCloud] with calendars: [iCloud/Work, Exchange/Work] reaches one calendar. All six tools (calendars_list, calendars_list_events, calendars_create_event, reminders_list, reminders_create, reminders_complete) run the applies_to-driven presence check before the TCC check: the question is this call's authority, which does not depend on whether the Mac would have answered, and putting it second would make a client's refusal vary with a grant it has nothing to do with — and would put the one check that is macMCP's own behind a framework read no hermetic test can make.

    The enumerators are scoped. calendars_list and reminders_list report only what is in scope, because listing every calendar on the machine to a confined client is a disclosure and is how that client learns what to try next. calendars_list gained a path and reminder rows a list_path, since the value a caller is expected to pass back has to be the one the listing hands out; the bare title / list stay because they always have.

    Three answers that used to be one. An explicit name outside the scope is a refusal carrying scope_violation and never a silent narrowing or an empty result (decision 11: a "not found" is indistinguishable from a real miss); a name on no calendar at all stays a plain not-found; and a scope naming a calendar this Mac does not hold is an operator's typo, an ordinary error, explicitly not a violation, so a mistyped value cannot fill the security signal. Two real values whose cross-product is empty is a third, distinct sentence — saying "not on this Mac" about either would be false. An empty intersection is never returned as an empty list: [] with isError: false is an affirmative claim that this Mac holds no calendar the client may reach, the shape the mail scan work removed as total_messages: 0.

    Resolution is the mailbox two-step: path first, then leaf name, and two carriers is refused with both named. Matching runs inside the scope first, so calendar_name: "Work" with calendars: [iCloud/Work] on a Mac that also holds Exchange/Work is not ambiguous — only one is reachable — and all rows are consulted only to tell the out-of-scope refusal from the plain miss. Two rows that generate the same path get a different sentence, because no argument this tool takes can tell them apart and printing one string twice invites a retry that lands in the same place.

    A write with no calendar_name resolves to the scope or refuses; it never falls through to EventKit's default. defaultCalendarForNewEvents / defaultCalendarForNewReminders() answer whatever the scope says, so the obvious code silently files an event on a calendar the profile never granted — the same shape as Mail sending from the default account for a from no account owns, which is why that one is refused rather than substituted. The default is used when it is inside the scope (an absent argument resolves to the scope, and a default inside it satisfies that); a scope of exactly one calendar resolves too; anything else asks, naming the choices. reminders_complete matches a title across the lists in scope and then reads the list back off the reminder before writing, the way mail re-checks where a globally-resolved byId message actually lives — and only a miss re-reads every list, solely to tell "it is elsewhere" from "there is none", with the refusal never naming where it is.

    What tests can reach. EventKit cannot be driven hermetically — there is no stub store and the real one is the user's own data — so the split is deliberate: EventKitScopeTests drives ScopedRows from literal rows and covers every decision in both directions, and CalendarRemindersScopeWiringTests proves the check is wired in at each of the six tools by calling them through the registry, which is possible only because the refusal returns before store.calendars(for:). What no test here covers is the EventKit half: that rows(of:) stays positionally aligned with store.calendars(for:), that an index really selects the calendar the predicate then reads, and every admitting path through a handler. Those need a live store. ScopedRows.fetch is where an empty calendar or reminder-list selection is kept away from EventKit, whose predicates read nil as every calendar; its tests prove the helper, not that each handler calls it, which also needs a live store.

    A message is always in a mailbox; a card need not be in a group — so the two contacts fields do not share an applies_to. contact_accounts and contact_groups were first declared with applies_to: ["contacts_*"] both, copying mail's two-axis shape, and the enforcement faithfully implemented it: a card was reachable only as a member of a group in scope. That is not a tighter confinement, it is a missing one — "every card in this account, group or not" becomes inexpressible, so a profile could be granted only the cards somebody had remembered to file, and account-level scoping (the thing actually asked for) could not be said at all. The axes now bound what they name: contact_accounts governs all ten tools and bounds cards; contact_groups governs exactly the four that cannot function without naming a group (contacts_list_groups, contacts_create_group, contacts_add_to_group, contacts_remove_from_group) and bounds groups, still as a cross-product against the accounts. The four are named explicitly rather than matched by contacts_*group*, which would select the same four today but by the spelling of a tool name. contacts_create is the file_dirs shape one level down: it is not governed by contact_groups, because a create works without a group, and its group argument is what refuses when the grant is absent — the account is what places the card, and a new account argument names it when the scope holds more than one.

    file_dirs is not a mail field, and three tools outside mail had no bound at all. It arrived with the mail work and grep -rn file_dirs Sources/ outside ContextSchema.swift hit MailScope.swift and MailService.swift and nothing else — while capture_screenshot and capture_audio took an unbounded absolute path and wrote to it and utilities_play_sound took an unbounded absolute path and read from it. That is ADR-011 finding 1, the arbitrary host write it called escalation rather than exfiltration, one service over; the live attempt passed every relay and macMCP layer and died on the tool's backend. All three are now in applies_to, by the rule mail_get_source established (name a tool only if it cannot function at all without the field; a tool with merely a parameter needing it keeps working and the parameter refuses): utilities_play_sound requires path, and a capture's path is optional in the schema and unavoidable in fact — every call writes a file, a call with no path writes to ~/Desktop, so there is no form of it a client with no directory could be served. HostFileScope is the enforcement, sharing ResourceScope.bound/realPath with mail and carrying its own sentences and its own Outcome (not Decision, whose .refuse always means a violation, and "you may write in two directories and named neither" is not one). An absent path resolves to the scope in ScopedRows.defaultTarget's three steps: the tool's own default when the scope contains it, the single allowed directory when there is one, otherwise an error naming them — refusing outright would mean no scoped client could screenshot without first knowing its own file_dirs, which it has no way to ask for. The two sentences for an unusable file_dirs entry moved to ResourceScope and are shared, so one bad entry does not read two ways.

    Writing a message into a mailbox is reaching it. With mail_mailboxes: ["Archive", "INBOX"], three tools gave three answers about one mailbox: mail_move to Drafts refused, mail_create_draft wrote a complete message into that same Drafts, and mail_get_email on the id it handed back refused it as out of scope. Confirmed live — with the guard removed, that profile put two full drafts into Alice's Drafts on the fixture. The field is described to the operator as "mailbox paths within those accounts this client may reach", and decision 11's exemption does not cover this: it excuses bookkeeping about the message a tool just wrote (the autosave sweep, which can surface nothing but the client's own message), and the creation is the write. So mail_create_draft refuses a profile whose mail_mailboxes does not name Drafts — exactly the cost of not being able to mail_move into it, which such a profile already pays. mail_send is deliberately not bound by it: what it produces is a message on the wire, and the Sent copy is bookkeeping about a delivery that already happened, is not a destination the caller chose, and yields no handle; requiring Sent would take sending away from every write profile to bound a copy that changes nothing about who received the mail. The destination is matched as a path with no leaf-name fallback — nobody wrote it, and Projects/Drafts is a project folder — which is where mail_move also ends up, one step later, when fmAllowed checks the resolved path.

    Two smaller ones on the EventKit path. CalendarService.sourceName is the one place either service turns an EKSource.title into the account half of a path: source?.title ?? unknownSourceName fell back for an absent source and passed "" straight through, so an unnamed CalDAV account produced a container of "" and a path of "/Work" — ContactsService.containerName has handled empty since it was written. And reminders_complete matched titles with .lowercased() while the scope membership of the reminder it picked is decided by ResourceScope.fold. Measured rather than assumed: in Swift these agree — String == compares by canonical equivalence, so across every scalar in U+0000...U+2FFFF against its own decomposition zero verdicts differ, and the NFC/NFD divergence fold exists to end is a property of JavaScript's ===. The fix changes no behaviour; it removes the last comparison on a scoped path carrying its own case rule, which is how they come apart if one is ever moved somewhere == is not canonical.

    Both ends of a membership change are still checked, against different fields. contacts_add_to_group needs the card (its account) and the group (both fields). A client that could add any card in the address book to a group it holds could read it on the next call — widening its own scope by writing. Only the field the card end answers to moved.

    Cards are fetched per container and a handle is checked before it is read. cards(in:) runs one predicateForContactsInContainer fetch per in-scope account, so the store is never enumerated; CNContactFetchRequest takes one predicate and Contacts has no compound predicates, so a name or a phone number is matched in Swift over what came back (nameMatches, phoneMatches) — the choice that reads less. resolve asks CNContainer.predicateForContainerOfContact first, which returns containers and no contact data, and fetches the card only once one of them is in scope: the ordering is the confinement. Every scoped read sets unifyResults = false, because a unified contact merges linked cards across containers and CNContact has no per-value provenance to subtract an out-of-scope account's fields with.

    contacts_create_group stays refused, re-derived rather than inherited. The old reason (a new group is a dead end, since a card was reachable only through one) died with the old shape. Two reasons survive it. A contact_groups value is an Account/Group path matched whole — a group name may contain a /, unlike a mailbox leaf — so it cannot be split into an account and a name to create one from, which kills the one coherent case (a scope naming a group that does not exist yet). And a value matching no group is a stale or mistyped grant that decision 11 says to surface, not an instruction to make it real: creating iCloud/Familly on the strength of a typo deletes the misconfiguration signal that would have shown it. For any name the scope does not hold, the original dilemma stands unchanged.

    A contact_accounts value can be ambiguous, and not academically. containerName falls back to the container's type when CNContainer.name is empty — which macOS routinely leaves it — so two unnamed CardDAV accounts are both CardDAV. Such a value selects neither, for the reason a two-carrier group path does.

    An empty values filter on context/enumerate means ALL, never none. That is the picker's normal initial state — the dependent field is opened before its dependency is chosen — so reading it as "match nothing" shows an operator zero calendars at the moment they are trying to pick one, indistinguishable from a Mac that holds none. This is the opposite of decision 4's rule for a scope value, where empty refuses, and the two are not in tension: an authorisation must fail closed, a query must fail informative. A store that cannot be read is -32000 and never [], because a failure and "there are none" must not look the same.

  • messages_search filters in Swift, not in SQL, and says how far it got. A message's text lives in m.text or in the attributedBody typedstream archive, and everything a recent Messages build sends is in the second with the column empty — so a LIKE in the query would find some messages and silently miss the majority, and a SQL LIMIT would cut the candidates before any of them had been decoded. The query therefore selects newest-first over the time window and the matching happens here, bounded by searchScanLimit (5,000 rows) with messages_scanned and scan_complete reported on every answer. Two other things are not what the obvious implementation does: a contact is resolved by normalising both sides (+1 (555) 123-4567 and 5551234567 are one handle; h.id = ? matched exactly one spelling and returned nothing for the others), and naming a contact filters by chat membership rather than by each message's handle_id, because anything sent from this Mac carries handle_id 0 and would drop the caller's own half of the conversation. A handle nobody has is said so in a note rather than returned as an empty list, which would read as "you have no messages with them".

Build

swift build              # debug
./build.sh               # release, codesigned, installs to ~/.local/bin, registers with Relay

Requires Swift 5.9+, macOS 13+. System frameworks only: EventKit, Contacts, CoreLocation, Foundation, SQLite3, AppKit. The binary embeds an Info.plist via -sectcreate for macOS permission prompts (Location Services).

Tests

Real mail sends: only to the operator's own address, one to recipient, no cc/bcc. Read the recipient back from the message before sending where the tooling allows, and delete test drafts after.

swift test

Tests/macMCPTests tests the executable target directly (@testable import macmcp), so anything under test has to be at least internal — several Mail helpers are deliberately not private for that reason, and say so at their declaration.

The suite must stay hermetic: no test may talk to Mail.app, the network, or the user's own data. Two patterns make that possible for a service that is mostly generated JavaScript:

  • JXA.run executes a script through osascript -l JavaScript with mail bound to MailStubJS's fake object graph instead of Application('Mail'). Nothing calls Application(...), so no Apple Event is sent and no TCC prompt can appear — osascript is just a JavaScript engine. This is how the generated scripts themselves (which is where the real logic lives) get tested.
  • Pure seams. Byte-level behaviour that used to be inline in a handler — source decoding, truncation, attachment typing, the timeout message — is factored into small static functions that take data and return data, so the regression can be pinned without a mailbox.

The one exception is MailSourceOnDiskTests, which is in the suite and has to be: byte-identity claims cannot be substantiated by a test that synthesises its own input with the transform under test, which is how two real deviations went unnoticed for a release. It skips with instructions unless pointed at the fixture, and needs Mail.app and an Automation grant:

cd ~/source/barelyworkingcode/testMail && ./testmail.sh start
MACMCP_MAIL_FIXTURE=$HOME/source/barelyworkingcode/testMail \
    swift test --filter MailSourceOnDiskTests

It delivers a message carrying every byte value except CR/LF straight into the Maildir, fetches it back through mail_get_source, compares the bytes with the file on disk, and moves the probe to Trash afterwards. ~5s.

Other end-to-end checks against the fixture live outside this repo (~/source/barelyworkingcode/testMail); the ground truth for anything mail-shaped is the Maildir on disk, never a tool's own success return.

Adding a Service

  1. Create Sources/macMCP/Services/FooService.swift
  2. Define enum FooService with static func register(_ registry: ToolRegistry)
  3. Register tools using registry.register(MCPTool(...)) { params in ... }
  4. Use schema(), stringProp(), boolProp(), etc. from ToolRegistry for input schemas
  5. Return results via textResult(), errorResult(), or jsonResult()
  6. Call FooService.register(registry) in main.swift