Skip to content

[GOWS] - Messages carrying senderKeyDistributionMessage never decrypt — silent group message loss since @lid migration #2242

Description

@GalaxyRuler

Edit: re-posting the body — angle-bracket placeholders in the code blocks were stripped as HTML on first submit, which emptied out the SQL and log samples. Content is otherwise unchanged.

Describe the bug

Since WhatsApp migrated one of our group chats to @lid addressing, a consistent share of inbound group messages never decrypt. The message is stored, but Message (and RawMessage) contain only the encryption envelope and no content:

{
  "Message": {
    "senderKeyDistributionMessage": { "groupID": "GROUP_ID@g.us", "axolotlSenderKeyDistributionMessage": "..." },
    "messageContextInfo": { "deviceListMetadata": {} }
  },
  "IsReal": false,
  "Status": 4,
  "RetryCount": 0,
  "UnavailableRequestID": "",
  "Info": { "Type": "media", "MediaType": "document", "AddressingMode": "lid" }
}

The 1:1 part decrypts fine (that is where the SKDM comes from); the group skmsg payload does not. No error surfaces at the default log level, no webhook carries usable content, and GET /api/{session}/chats/{chatId}/messages/{messageId} later returns 404 Message not found — so the message cannot be recovered afterwards.

The predicate is exact

Analysing every group message since the addressing change, the split is total — this is not a correlation, it is an identity:

decrypted failed
Message contains senderKeyDistributionMessage 0 22
it does not 48 0

Every message that carries an SKDM fails. Every message that does not, succeeds.

Supporting observations, same dataset:

  • It starts exactly at the @lid cutover. 294 messages over the preceding three months, all with AddressingMode absent (legacy PN), zero failures. From the day the group flipped to AddressingMode: "lid", failures begin immediately and never stop. Overall since then: 46 ok / 20 failed (30%).
  • It is not sender-specific. 10 senders have both successes and failures under the same LID (one is 13 ok / 3 failed). Every sender that posted has a sender key stored for the group.
  • It tracks sender idleness. Median gap since that sender's previous message is 36 h before a failure versus 0 h before a success — consistent with WhatsApp re-distributing the sender key when someone has been quiet, and the content of that same message being encrypted under the newly distributed key.
  • The sender's next message recovers. After a failure, the same sender's following message decrypts in 12 of 14 cases — the key is usable from then on, just not for the message that delivered it.
  • No retry receipt is requested. All failures have RetryCount: 0, empty UnavailableRequestID, and an empty whatsmeow_retry_buffer. The retry path does work in general — it fired 3 times across 26,815 stored messages and recovered the message every time — it simply never engages for this case, so nothing self-heals.

Why this is painful

The lost message is, by construction, the first thing a participant posts after being idle. For any workflow where people drop a document into a group and then go quiet, that is precisely the message that matters. In our case roughly one in three shared documents is destroyed before the application ever sees it, silently.

Reproducing / verifying

Against the GOWS session store (gows.db), the failure set is:

SELECT count(*) FROM gows_messages
WHERE jid = 'GROUP_ID@g.us'
  AND json_extract(data,'$.Message.senderKeyDistributionMessage') IS NOT NULL
  AND json_extract(data,'$.Message.conversation') IS NULL
  AND json_extract(data,'$.Message.extendedTextMessage') IS NULL
  AND json_extract(data,'$.Message.documentMessage') IS NULL
  AND json_extract(data,'$.Message.imageMessage') IS NULL;

Anyone seeing silent group message loss on GOWS can run that to check whether they have the same thing.

Two smaller things

  1. The default log level hides it. With WAHA_LOG_LEVEL: error nothing at all is logged for these. Raising it to info surfaces GOWS engine lines. Given that the outcome is silent data loss, a warning at error level for an undecryptable message would have saved us a fortnight of not knowing.
  2. Upgrading did not help. We went gows-2026.7.1 to gows-2026.8.1 specifically for the LID fixes in that release. Behaviour is unchanged (2 ok / 2 failed since, including one more lost document). The session survived the upgrade and re-authenticated normally, so this is not a pairing problem.

Possibly related but, I believe, distinct: #1625 is about messages that do decrypt but are addressed inconsistently across @c.us / @lid. These never decrypt at all.

Happy to supply anything further — I have the full message store and can run whatever query is useful.

Version

{
  "version": "2026.8.1",
  "engine": "GOWS",
  "tier": "CORE"
}

Docker Logs

Nothing is emitted for a failing message at WAHA_LOG_LEVEL: error. At info the session behaves normally around the failure:

[Manager] Session started 'SESSION_NAME'
[Session/SESSION_NAME/Client] Successfully authenticated
[Session/SESSION_NAME/Client] NCT salt is empty - forcing regular_high app-state sync

The only trace of the failure is the stored message itself and, later, a 404 Message not found when the application tries to fetch it:

GET /api/SESSION_NAME/chats/GROUP_ID%40g.us/messages/false_GROUP_ID%40g.us_MESSAGE_ID_SENDER_LID%40lid?downloadMedia=true
   -> 404 NotFoundException "Message not found"

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions