You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We need a Lambda, user-bot, that sends each new IAM user a temporary console password and sign-in instructions by Slack DM, so that new members can log in to the incubator AWS account without a lead handing a password over by hand. Today the aws-users module already creates a console login profile for every user, but the password Terraform generates sits only in Terraform state and is never given to anyone.
Action Items
How it works. When Terraform creates a user, the module's aws_iam_user_login_profile produces a CreateLoginProfile call. An EventBridge rule catches that event and invokes the Lambda. The Lambda reads the user's tags and acts only if the user carries bothmanaged-by = terraform-devops-securityand a slack_id tag. It then sets a new temporary password with PasswordResetRequired: true and DMs it to that Slack member ID.
Why CreateLoginProfile and not CreateUser: Terraform calls CreateUser a moment before CreateLoginProfile, so a Lambda triggered by CreateUser could try to update a login profile that does not exist yet. By the time CreateLoginProfile fires, the user and all its tags exist, because the provider sends tags, including default_tags, with CreateUser.
Message sending goes through an interface with two implementations, both built here:
a Slack sender that calls Slack's Web API, fully unit-tested;
a stub sender that only logs that a DM would have been sent.
The deployed Lambda uses the stub. No Slack app or bot token exists yet, so switching the deployed Lambda to the Slack sender is out of scope (see the end of this section).
Deliver this as two PRs, in this order. The deploy workflow updates a function that the first PR's Terraform creates, so if both land in one merge, the deploy runs before the function exists and fails.
Add a validation block that accepts null or a value matching ^[UW][A-Z0-9]{8,}$.
The value is a Slack member ID, not a handle. Handles are not unique, people can change them, and Slack's API needs the ID to send a DM. The validation makes a pasted handle fail at terraform plan instead of silently going nowhere.
terraform/modules/aws-users/main.tf: in aws_iam_user.user, merge slack_id into the tags when it is set: tags = merge(var.user_tags, var.slack_id == null ? {} : { slack_id = var.slack_id }).
The tags = line was line 14 when this was written; line numbers may drift, so find it inside resource "aws_iam_user" "user".
Regenerate terraform/modules/aws-users/README.md with terraform-docs, using that directory's .terraform.docs.yml, so the new input is documented.
New file terraform/user-bot.tf. Put every resource in region = "us-east-1": IAM is a global service, and its CloudTrail events reach EventBridge only in us-east-1. The provider (~> 6.64) supports a per-resource region argument, so no provider alias is needed. Declare:
It depends on the existing management-events CloudTrail trail (terraform/cloudtrail.tf), which is multi-region and includes global service events. Do not change that trail.
aws_cloudwatch_event_target pointing at the Lambda, plus an aws_lambda_permission letting only this rule's ARN invoke the function.
aws_cloudwatch_log_group named /aws/lambda/user-bot, with a retention period set.
An execution role and its policy. Allow only:
iam:ListUserTags on users
iam:UpdateLoginProfile on users, with the condition StringEquals iam:ResourceTag/managed-by = terraform-devops-security
logs:CreateLogStream and logs:PutLogEvents on this function's log group
The tag condition matters. Without it, this role could reset the console password of any user in the account, including admins and the users tagged managed-by = exempt. It enforces in IAM what the Lambda also checks in code.
aws_lambda_functionuser-bot, on a Node.js runtime matching lambda/user-bot/.nvmrc from PR 2.
Create it from a placeholder handler packaged with an archive_file data source using inline source { content = ... }. Terraform never builds or ships the real code; the deploy workflow does.
Add lifecycle { ignore_changes = [filename, source_code_hash] }, so Terraform does not revert the deployed code and terraform plan shows no diff after a deploy.
Runtime, handler, memory, timeout, environment and role all stay managed by Terraform.
terraform/aws-gha-oidc-providers.tf: add the deploy role devops-security-user-bot-deploy.
Trust: the existing GitHub OIDC provider (module.iam_oidc_gha_incubator.provider_arn), with aud = sts.amazonaws.com and sub = repo:hackforla/devops-security:ref:refs/heads/main only.
Permissions: only lambda:UpdateFunctionCode and lambda:GetFunction, on this one function's ARN.
Unlike devops-security-tf-plan/-apply, this role can be declared in Terraform, because nothing assumes it until after Terraform has run. So it carries the normal managed-by tag, not exempt. Leave the comment at the top of that file about those two roles as it is.
.github/ISSUE_TEMPLATE/request-aws-iam-resources.yml: add a required input field "Slack member ID". Its description should explain how to find the ID (Slack profile → ⋮ → Copy member ID) and that it is not the @handle.
Do not modify .github/workflows/terraform-plan.yaml or terraform-apply.yaml. They trigger on **/*.tf, which already covers everything in this PR, and nothing under lambda/ contains a .tf file.
PR 2: The Lambda and its workflows
Scaffold lambda/user-bot/: TypeScript, package.json with a committed package-lock.json, tsconfig.json, .nvmrc matching the Lambda runtime in PR 1, and a build step that bundles the handler into a single file for zipping. Use AWS SDK v3 (@aws-sdk/client-iam).
The handler:
Read detail.requestParameters.userName from the event.
Call ListUserTags.
Exit without changing anything unless managed-by is exactly terraform-devops-securityandslack_id is present. Log which check failed and for which user name.
Generate a 20-character password with Node's crypto (not Math.random). The account has no custom password policy, so AWS's default applies; include upper, lower, digit and symbol characters.
Call UpdateLoginProfile with PasswordResetRequired: true.
Send the DM through a MessageSender interface. The message contains:
the sign-in URL https://hfla-incubator.signin.aws.amazon.com/console
the IAM user name
the temporary password
a note that they must change the password and set up MFA at first sign-in
Never log the password, on any path including errors.
Write SlackMessageSender, the real implementation:
It calls chat.postMessage with channel set to the member ID, which posts to the bot's DM with that user (bot scope chat:write).
Use the built-in fetch, so there is no Slack SDK dependency. The bot token is passed in through the constructor; where it comes from at runtime is part of the out-of-scope wiring.
Check the ok field of the response, not the HTTP status. Slack returns HTTP 200 with "ok": false and an error string for failures such as user_not_found, channel_not_found or invalid_auth. A sender that only checks the status reports success when nothing was delivered.
Never include the token or the message body in a thrown error or a log line.
Write StubMessageSender, which logs the user name and Slack ID but not the message body. It has no Slack code and no token. The deployed handler is wired to this one until the out-of-scope follow-up switches it.
Unit tests with Vitest, mocking AWS with aws-sdk-client-mock and using a mock MessageSender. Cover at least:
the happy path
managed-by missing
managed-by set to another value, e.g. exempt
slack_id missing
UpdateLoginProfile failing, in which case nothing is sent
the send failing
the generated password meeting the default policy
the password never appearing in any log output, asserted by capturing the logger
Unit tests for SlackMessageSender, mocking fetch (no real Slack calls in tests). Cover at least:
it sends the right URL, auth header, channel and text
"ok": true resolves
HTTP 200 with "ok": false rejects
a non-200 status rejects
a network error rejects
neither the token nor the password appears in a thrown error
New .github/workflows/user-bot-test.yml:
Trigger: pull_request with paths: ['lambda/user-bot/**', '.github/workflows/user-bot-test.yml'].
permissions: contents: read and no AWS credentials.
Steps: actions/checkout@v5, actions/setup-node with node-version-file: lambda/user-bot/.nvmrc, npm ci, typecheck, npm test, npm run build.
New .github/workflows/user-bot-deploy.yml:
Triggers: push to main with paths: ['lambda/user-bot/**', '.github/workflows/user-bot-deploy.yml'], plus workflow_dispatch for recovery.
permissions: id-token: write, contents: read, and a concurrency group so two merges cannot deploy over each other.
Steps:
npm ci, test, build, then zip the bundle. Tests run again on exactly what ships.
aws-actions/configure-aws-credentials@v6 assuming devops-security-user-bot-deploy in us-east-1.
aws lambda update-function-code --zip-file, then aws lambda wait function-updated.
Use checkout@v5 and configure-aws-credentials@v6, the org's settled targets, even though the existing workflows still use v4.
lambda/user-bot/README.md covering:
what the bot does and when it fires
the two tag conditions
how to run the tests locally
how deploys happen (merge to main, or workflow_dispatch to redeploy)
that Terraform owns everything about the function except its code
After the PRs merge
After PR 1 merges and Terraform Apply succeeds, confirm in us-east-1:
the rule, function, log group and both roles exist
terraform plan on a fresh PR shows no diff for them
After PR 2 merges, confirm user-bot-deploy ran green and that the function's CodeSha256 changed from the placeholder: aws lambda get-function --function-name user-bot --region us-east-1.
After both merge, test end to end:
Open a PR adding a throwaway user, e.g. user-bot-test, in terraform/aws-users.tf with a slack_id, and merge it.
In /aws/lambda/user-bot, confirm the run updated the login profile and called the stub sender, and that the password appears nowhere in the log.
Remove the throwaway user in a follow-up PR.
After both merge, confirm a run for a user withoutslack_id is logged as skipped and does not change that user's login profile.
Out of scope
Switching the deployed Lambda to SlackMessageSender: creating the Slack app, its bot token, a Secrets Manager secret the Lambda reads at runtime, and the execution-role permission to read it. Needs its own ticket; the sender code itself is built here.
Backfilling slack_id on existing users. The bot only fires for new login profiles.
The unused Terraform-generated password that remains in Terraform state.
Resources/Instructions
Module to change: terraform/modules/aws-users/ (main.tf, variables.tf, README.md)
Existing CloudTrail trail the rule depends on: terraform/cloudtrail.tf (management-events)
OIDC roles and the comment explaining why the two Terraform CI roles are not in Terraform: terraform/aws-gha-oidc-providers.tf
Overview
We need a Lambda,
user-bot, that sends each new IAM user a temporary console password and sign-in instructions by Slack DM, so that new members can log in to the incubator AWS account without a lead handing a password over by hand. Today theaws-usersmodule already creates a console login profile for every user, but the password Terraform generates sits only in Terraform state and is never given to anyone.Action Items
How it works. When Terraform creates a user, the module's
aws_iam_user_login_profileproduces aCreateLoginProfilecall. An EventBridge rule catches that event and invokes the Lambda. The Lambda reads the user's tags and acts only if the user carries bothmanaged-by = terraform-devops-securityand aslack_idtag. It then sets a new temporary password withPasswordResetRequired: trueand DMs it to that Slack member ID.Why
CreateLoginProfileand notCreateUser: Terraform callsCreateUsera moment beforeCreateLoginProfile, so a Lambda triggered byCreateUsercould try to update a login profile that does not exist yet. By the timeCreateLoginProfilefires, the user and all its tags exist, because the provider sends tags, includingdefault_tags, withCreateUser.Message sending goes through an interface with two implementations, both built here:
The deployed Lambda uses the stub. No Slack app or bot token exists yet, so switching the deployed Lambda to the Slack sender is out of scope (see the end of this section).
Deliver this as two PRs, in this order. The deploy workflow updates a function that the first PR's Terraform creates, so if both land in one merge, the deploy runs before the function exists and fails.
PR 1: Terraform and the request form
terraform/modules/aws-users/variables.tf: add an optionalslack_idvariable (type = string,default = null).validationblock that acceptsnullor a value matching^[UW][A-Z0-9]{8,}$.terraform planinstead of silently going nowhere.terraform/modules/aws-users/main.tf: inaws_iam_user.user, mergeslack_idinto the tags when it is set:tags = merge(var.user_tags, var.slack_id == null ? {} : { slack_id = var.slack_id }).slack_idis the one already used by hand on earlier accounts (see Create VRMS aws account for Incubator access devops#77 and Create AWS IAM user account for Antonina devops#80), and one live user still carries it.tags =line was line 14 when this was written; line numbers may drift, so find it insideresource "aws_iam_user" "user".terraform/modules/aws-users/README.mdwith terraform-docs, using that directory's.terraform.docs.yml, so the new input is documented.terraform/user-bot.tf. Put every resource inregion = "us-east-1": IAM is a global service, and its CloudTrail events reach EventBridge only in us-east-1. The provider (~> 6.64) supports a per-resourceregionargument, so no provider alias is needed. Declare:aws_cloudwatch_event_rulematchingsource = ["aws.iam"],detail.eventSource = ["iam.amazonaws.com"],detail.eventName = ["CreateLoginProfile"].management-eventsCloudTrail trail (terraform/cloudtrail.tf), which is multi-region and includes global service events. Do not change that trail.aws_cloudwatch_event_targetpointing at the Lambda, plus anaws_lambda_permissionletting only this rule's ARN invoke the function.aws_cloudwatch_log_groupnamed/aws/lambda/user-bot, with a retention period set.An execution role and its policy. Allow only:
iam:ListUserTagson usersiam:UpdateLoginProfileon users, with the conditionStringEquals iam:ResourceTag/managed-by = terraform-devops-securitylogs:CreateLogStreamandlogs:PutLogEventson this function's log groupThe tag condition matters. Without it, this role could reset the console password of any user in the account, including admins and the users tagged
managed-by = exempt. It enforces in IAM what the Lambda also checks in code.aws_lambda_functionuser-bot, on a Node.js runtime matchinglambda/user-bot/.nvmrcfrom PR 2.archive_filedata source using inlinesource { content = ... }. Terraform never builds or ships the real code; the deploy workflow does.lifecycle { ignore_changes = [filename, source_code_hash] }, so Terraform does not revert the deployed code andterraform planshows no diff after a deploy.terraform/aws-gha-oidc-providers.tf: add the deploy roledevops-security-user-bot-deploy.module.iam_oidc_gha_incubator.provider_arn), withaud = sts.amazonaws.comandsub = repo:hackforla/devops-security:ref:refs/heads/mainonly.lambda:UpdateFunctionCodeandlambda:GetFunction, on this one function's ARN.devops-security-tf-plan/-apply, this role can be declared in Terraform, because nothing assumes it until after Terraform has run. So it carries the normalmanaged-bytag, notexempt. Leave the comment at the top of that file about those two roles as it is..github/ISSUE_TEMPLATE/request-aws-iam-resources.yml: add a requiredinputfield "Slack member ID". Its description should explain how to find the ID (Slack profile → ⋮ → Copy member ID) and that it is not the @handle..github/workflows/terraform-plan.yamlorterraform-apply.yaml. They trigger on**/*.tf, which already covers everything in this PR, and nothing underlambda/contains a.tffile.PR 2: The Lambda and its workflows
lambda/user-bot/: TypeScript,package.jsonwith a committedpackage-lock.json,tsconfig.json,.nvmrcmatching the Lambda runtime in PR 1, and a build step that bundles the handler into a single file for zipping. Use AWS SDK v3 (@aws-sdk/client-iam).detail.requestParameters.userNamefrom the event.ListUserTags.managed-byis exactlyterraform-devops-securityandslack_idis present. Log which check failed and for which user name.crypto(notMath.random). The account has no custom password policy, so AWS's default applies; include upper, lower, digit and symbol characters.UpdateLoginProfilewithPasswordResetRequired: true.MessageSenderinterface. The message contains:https://hfla-incubator.signin.aws.amazon.com/consoleSlackMessageSender, the real implementation:chat.postMessagewithchannelset to the member ID, which posts to the bot's DM with that user (bot scopechat:write).fetch, so there is no Slack SDK dependency. The bot token is passed in through the constructor; where it comes from at runtime is part of the out-of-scope wiring.okfield of the response, not the HTTP status. Slack returns HTTP 200 with"ok": falseand anerrorstring for failures such asuser_not_found,channel_not_foundorinvalid_auth. A sender that only checks the status reports success when nothing was delivered.StubMessageSender, which logs the user name and Slack ID but not the message body. It has no Slack code and no token. The deployed handler is wired to this one until the out-of-scope follow-up switches it.aws-sdk-client-mockand using a mockMessageSender. Cover at least:managed-bymissingmanaged-byset to another value, e.g.exemptslack_idmissingUpdateLoginProfilefailing, in which case nothing is sentSlackMessageSender, mockingfetch(no real Slack calls in tests). Cover at least:channeland text"ok": trueresolves"ok": falserejects.github/workflows/user-bot-test.yml:pull_requestwithpaths: ['lambda/user-bot/**', '.github/workflows/user-bot-test.yml'].permissions: contents: readand no AWS credentials.actions/checkout@v5,actions/setup-nodewithnode-version-file: lambda/user-bot/.nvmrc,npm ci, typecheck,npm test,npm run build..github/workflows/user-bot-deploy.yml:pushtomainwithpaths: ['lambda/user-bot/**', '.github/workflows/user-bot-deploy.yml'], plusworkflow_dispatchfor recovery.permissions: id-token: write, contents: read, and aconcurrencygroup so two merges cannot deploy over each other.npm ci, test, build, then zip the bundle. Tests run again on exactly what ships.aws-actions/configure-aws-credentials@v6assumingdevops-security-user-bot-deployinus-east-1.aws lambda update-function-code --zip-file, thenaws lambda wait function-updated.checkout@v5andconfigure-aws-credentials@v6, the org's settled targets, even though the existing workflows still use v4.lambda/user-bot/README.mdcovering:main, orworkflow_dispatchto redeploy)After the PRs merge
Terraform Applysucceeds, confirm in us-east-1:terraform planon a fresh PR shows no diff for themuser-bot-deployran green and that the function'sCodeSha256changed from the placeholder:aws lambda get-function --function-name user-bot --region us-east-1.user-bot-test, interraform/aws-users.tfwith aslack_id, and merge it./aws/lambda/user-bot, confirm the run updated the login profile and called the stub sender, and that the password appears nowhere in the log.slack_idis logged as skipped and does not change that user's login profile.Out of scope
SlackMessageSender: creating the Slack app, its bot token, a Secrets Manager secret the Lambda reads at runtime, and the execution-role permission to read it. Needs its own ticket; the sender code itself is built here.slack_idon existing users. The bot only fires for new login profiles.Resources/Instructions
terraform/modules/aws-users/(main.tf,variables.tf,README.md)terraform/cloudtrail.tf(management-events)terraform/aws-gha-oidc-providers.tf.github/ISSUE_TEMPLATE/request-aws-iam-resources.ymlslack_idtag: Create VRMS aws account for Incubator access devops#77, Create AWS IAM user account for Antonina devops#80UpdateLoginProfileand controlling access withiam:ResourceTagregion(v6)chat.postMessage(DMing by member ID, theok/errorresponse shape)aws-sdk-client-mock