Vertex Macro | Financial Cloud Cloud · AWS Exam
AWS DevOps Certified Professional
| Article |
|---|
| 01 AWS Solutions Architect Professional SAP-C02 lecture notes |
| 02 AWS DevOps Certified Professional DOP-C02 lecture notes |
Lecture content for government technology advisors, public-welfare and education leaders, data and security owners, mental-health service networks, and large-institution decision makers.
Amazon Aurora Global Database
Storage-level, cross-Region replication provides very low RPO and fast managed failover for low RTO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Storage-level, cross-Region replication provides very low RPO and fast managed failover for low RTO.
● Scene fit: Storage-level, cross-Region replication provides very low RPO and fast managed failover for low RTO.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Asynchronous replication improves RPO, but promotion and catch-up increase RTO compared to purpose-built global replication.
● Automates recoveries but RPO depends on snapshot frequency and RTO includes full restore time.
● Multi-AZ standbys are confined to a single Region and cannot be placed cross-Region.
Workflow: Amazon Aurora Global Database.
AWS CodeDeploy blue/green with 60-minute blue termination
CodeDeploy blue/green for EC2/Auto Scaling swaps ALB target groups when ready and supports terminating the original fleet after a specified wait time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy blue/green for EC2/Auto Scaling swaps ALB target groups when ready and supports terminating the original fleet after a specified wait time.
● Scene fit: CodeDeploy blue/green for EC2/Auto Scaling swaps ALB target groups when ready and supports terminating the original fleet after a specified wait time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudFormation and Route 53 can create resources and update DNS but do not natively manage blue/green traffic shifting and timed old-fleet termination.
● Instance Refresh performs in-place rolling updates and does not create a separate green fleet or schedule post-cutover termination of the old fleet.
● Elastic Beanstalk can swap environments but does not provide an automatic 60-minute retirement of the previous environment.
Workflow: AWS CodeDeploy blue/green with 60-minute blue termination.
Update the deployment to create /health-ready.php only after initialization completes and point
Creating a deterministic readiness endpoint at the end of deployment ensures targets become healthy only when the app is actually ready.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Creating a deterministic readiness endpoint at the end of deployment ensures targets become healthy only when the app is actually ready.
● Scene fit: Creating a deterministic readiness endpoint at the end of deployment ensures targets become healthy only when the app.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Altering the status code and extending the timer does not guarantee the app is ready and can still mark unstable targets as healthy.
● Adding more time and checks delays rollouts but does not reliably correlate health with true application readiness.
● Route 53 health checks operate at the DNS level and do not control ALB target registration or per-target readiness.
Workflow: Update the deployment to create /health-ready.php only after initialization completes and point the ALB health check at that path.
an instance profile and retrieve credentials from AWS Secrets Manager
Instance profiles supply temporary credentials and Secrets Manager securely stores and rotates database passwords.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Instance profiles supply temporary credentials and Secrets Manager securely stores and rotates database passwords.
● Scene fit: Instance profiles supply temporary credentials and Secrets Manager securely stores and rotates database passwords.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● S3 is not a secrets store, lacks built-in secret rotation, and increases exposure risk compared to dedicated services.
● Embedding secrets in images complicates rotation and proliferates secrets across images.
● Static keys on instances are high risk and Parameter Store does not provide native rotation for RDS passwords.
Workflow: Use an instance profile and retrieve credentials from AWS Secrets Manager.
Amazon RDS with automated backups and deletion protection + AWS Elastic Beanstalk
Provides a managed relational database with automated backups and safeguards against accidental deletion. Managed platform for Node.js that integrates Application Load Balancer and Auto Scaling to reduce operational effort.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides a managed relational database with automated backups and safeguards against accidental deletion.
● Scene fit: Provides a managed relational database with automated backups and safeguards against accidental deletion. Managed platform for Node.js that integrates Application Load Balancer and Auto Scaling to reduce operational effort.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Kubernetes adds operational overhead for cluster, ingress, and scaling compared to a fully managed PaaS.
● DynamoDB is NoSQL and does not satisfy the relational database requirement.
● Still requires server management, patching, and capacity tuning, which increases operational burden.
Workflow: Amazon RDS with automated backups and deletion protection → AWS Elastic Beanstalk with ALB and Auto Scaling.
ALB and Auto Scaling in a second Region to place compute near
Running compute in another Region reduces latency for nearby users and increases regional resiliency. Latency-based routing directs clients to the lowest-latency healthy endpoint and supports failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Running compute in another Region reduces latency for nearby users and increases regional resiliency.
● Scene fit: Running compute in another Region reduces latency for nearby users and increases regional resiliency. Latency-based routing directs clients.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFront primarily helps cache static content and does not solve dynamic write latency or multi-Region availability.
● DynamoDB does not support ad-hoc cross-Region replication; Global Tables is the supported approach.
● With a single Region origin this improves edge-to-region path but does not provide local writes or multi-Region resiliency.
Workflow: Deploy ALB and Auto Scaling in a second Region to place compute near users → Use Route 53 latency-based routing with health checks to send users to the nearest regional ALB → Enable DynamoDB Global Tables with a replica in a second Region.
Ingest both app logs and CloudTrail into CloudWatch Logs; query with CloudWatch
Sending both datasets to CloudWatch Logs allows CloudWatch Logs Insights to query them together.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Sending both datasets to CloudWatch Logs allows CloudWatch Logs Insights to query them together.
● Scene fit: Sending both datasets to CloudWatch Logs allows CloudWatch Logs Insights to query them together.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch Logs Insights cannot query S3 data.
● CloudTrail Lake is for CloudTrail event data and does not ingest arbitrary application logs.
● The CloudWatch agent does not ship logs directly to S3.
Workflow: Ingest both app logs and CloudTrail into CloudWatch Logs → query with CloudWatch Logs Insights.
Block non-TLS via aws:SecureTransport, set default SSE-S3, and use S3 Cross-Region Replication
This enforces HTTPS, uses SSE-S3 which scales without KMS throttling, and replicates to a second Region.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This enforces HTTPS, uses SSE-S3 which scales without KMS throttling, and replicates to a second Region.
● Scene fit: This enforces HTTPS, uses SSE-S3 which scales without KMS throttling, and replicates to a second Region.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Enforces HTTPS and sets up CRR, but SSE-KMS introduces AWS KMS request-per-second limits that can bottleneck very high PUT rates.
● Multi-Region Access Points do not replicate data; without replication configured this does not satisfy cross-Region durability, and SSE-KMS can still throttle.
● RTC does not remove KMS TPS limits and this does not enforce TLS at the bucket level.
Workflow: Block non-TLS via aws:SecureTransport → set default SSE-S3, and use S3 Cross-Region Replication.
Ingest the stream with Amazon Kinesis Data Streams, use Amazon Managed Service
Kinesis Data Streams with Apache Flink provides scalable real-time processing, Firehose to S3 is low cost, and Glue plus Athena enables serverless ad hoc analytics.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Kinesis Data Streams with Apache Flink provides scalable real-time processing, Firehose to S3 is low cost, and Glue.
● Scene fit: Kinesis Data Streams with Apache Flink provides scalable real-time processing, Firehose to S3 is low cost, and Glue.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● EventBridge is not optimized for sustained high-throughput clickstream ingestion compared to purpose-built streaming services.
● Running EMR on large EC2 instances adds significant cost and inserting SQS in the pipeline is unnecessary for real-time.
● SQS is not designed for high-throughput streaming clickstreams and instance store is ephemeral and unsuitable for analytics storage.
Workflow: Ingest the stream with Amazon Kinesis Data Streams → use Amazon Managed Service for Apache Flink Studio for real-time sessionization, have AWS Lambda forward aggregates to Amazon Kinesis Data Firehose, land data.
EBS TagSpecifications with CostCenterId=9421 to the launch template
Launch template TagSpecifications tag EBS volumes at creation automatically.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Launch template TagSpecifications tag EBS volumes at creation automatically.
● Scene fit: Launch template TagSpecifications tag EBS volumes at creation automatically.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● ASG tag propagation applies to EC2 instances, not attached EBS volumes.
● Tags are applied after creation and add unnecessary operational overhead.
● Enforces but does not auto-tag and can break service-created volumes.
Workflow: Add EBS TagSpecifications with CostCenterId=9421 to the launch template.
Lower MinSuccessfulInstancesPercent to prevent full rollback when a few instances fail +
Reducing the success threshold lets the stack continue if only a minority of instances fail to launch or signal. Turning off signals helps identify whether signal.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Reducing the success threshold lets the stack continue if only a minority of instances fail to launch or signal.
● Scene fit: Reducing the success threshold lets the stack continue if only a minority of instances fail to launch or.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Suspending these essential processes stops instance replacement and registration, stalling the rollout.
● Replacing the entire group increases risk and hides the root cause rather than diagnosing rolling-update failures.
● Switching mechanisms does not address CloudFormation-specific stalls and does not directly reduce unnecessary rollbacks.
Workflow: Lower MinSuccessfulInstancesPercent to prevent full rollback when a few instances fail → Disable WaitOnResourceSignals for the rolling update → Suspend HealthCheck, ReplaceUnhealthy, AZRebalance, AlarmNotification, and ScheduledActions during the rollout.
Amazon DynamoDB and enable Global Tables spanning the five chosen Regions
DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform low-latency local reads and writes that are replicated worldwide.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform low-latency local reads and writes that are replicated worldwide.
● Scene fit: DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform low-latency local reads and writes that.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● ElastiCache replication groups are scoped to a single Region and even with Global Datastore do not support multi-active writes across Regions, so they cannot provide low-latency writes in every Region.
● Cross-Region read replicas are read-only and all writes must go to the primary Region, which increases write latency for users outside that Region.
● Aurora Global Database supports one primary write Region with cross-Region read replicas, so it does not meet the requirement for multi-Region active writes.
Workflow: Use Amazon DynamoDB and enable Global Tables spanning the five chosen Regions.
a new version of the existing launch template with the required instance
This updates the definition used for all future launches so new instances scale out with the new instance type.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This updates the definition used for all future launches so new instances scale out with the new instance type.
● Scene fit: This updates the definition used for all future launches so new instances scale out with the new instance type.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Instance refresh does not change instance type unless the launch template or configuration it references is updated first.
● A group configured for a launch template cannot simultaneously switch to a launch configuration and launch configurations are legacy.
● Overrides are for mixed instance types and will not force a single new type for all instances without reconfiguring the group.
Workflow: Create a new version of the existing launch template with the required instance type and update the Auto Scaling group to reference that version.
a rolling update on the green Auto Scaling group to roll out
This prepares the green environment first and then performs an immediate ALB switch for an all-at-once cutover without DNS propagation delays.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This prepares the green environment first and then performs an immediate ALB switch for an all-at-once cutover without.
● Scene fit: This prepares the green environment first and then performs an immediate ALB switch for an all-at-once cutover without.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Shifts focus to the blue group and proposes a DNS change that cannot directly target an ALB target group.
● Would route live traffic to green before the new version is fully deployed and validated, risking outages.
● Changing the Route 53 alias is unnecessary with a single ALB and cannot select a specific target group, and.
Workflow: Run a rolling update on the green Auto Scaling group to roll out the new build, → use the AWS CLI to move the ALB listener to the green target group.
Amazon Inspector to automatically evaluate applications for exposure, vulnerabilities, and deviations from
Inspector provides the required automated vulnerability assessment, the ALB plus Multi-AZ Auto Scaling design delivers high availability, Aurora is resilient, and an alias record is the.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Inspector provides the required automated vulnerability assessment, the ALB plus Multi-AZ Auto Scaling design delivers high availability, Aurora.
● Scene fit: Inspector provides the required automated vulnerability assessment, the ALB plus Multi-AZ Auto Scaling design delivers high availability, Aurora.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Macie focuses on sensitive data discovery and a non-alias A record cannot target an ALB, so this does not.
● GuardDuty is threat detection rather than application vulnerability assessment, and a CNAME cannot be used at the zone apex.
Workflow: Use Amazon Inspector to automatically evaluate applications for exposure, vulnerabilities, and deviations from AWS best practices → run an EC2 Auto Scaling group spread across three Availability Zones behind an Application Load.
Publish a new Lambda version and expose a new API Gateway stage
Both stages invoke one Lambda via an alias while the legacy stage adds a static field using a mapping template, preserving a single backend and long-term.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Both stages invoke one Lambda via an alias while the legacy stage adds a static field using a.
● Scene fit: Both stages invoke one Lambda via an alias while the legacy stage adds a static field using a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Creates two Lambda functions and introduces proxy logic, increasing maintenance and violating the requirement to manage only one function.
● Caching and stage variables do not reliably mutate request bodies, and removing the legacy path risks breaking older clients.
Workflow: Publish a new Lambda version and expose a new API Gateway stage named prod-v2 that, along with prod-v1, invokes the same Lambda alias → on prod-v1 add a request mapping template that.
statements to the destination bucket policy that allow the replication IAM role
For cross-account S3 Replication, the target bucket must trust the source account’s replication role to write and manage object ownership or ACLs. S3 uses a role.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: For cross-account S3 Replication, the target bucket must trust the source account’s replication role to write and manage.
● Scene fit: For cross-account S3 Replication, the target bucket must trust the source account’s replication role to write and manage.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Creates a custom copy pipeline rather than enabling native S3 Replication, so it does not satisfy the requirement for.
● S3 assumes a role in the source account for replication, not in the destination account, so creating the role.
Workflow: Add statements to the destination bucket policy that allow the replication IAM role in the source account to write objects and set object ownership → Create an IAM role in the source.
AWS Service Catalog
Publishes vetted CloudFormation products with constraints and required tags for preventative control at provisioning time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Publishes vetted CloudFormation products with constraints and required tags for preventative control at provisioning time.
● Scene fit: Publishes vetted CloudFormation products with constraints and required tags for preventative control at provisioning time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● SCPs and tag policies can restrict or validate tagging but do not provide a curated catalog of vetted architectures.
● Detective compliance after resources are created, not a preventative provisioning control.
● Template policy-as-code useful in pipelines but does not provide a self-service catalog or block ad hoc provisioning.
Workflow: AWS Service Catalog.
awslogs in the task definition; grant CloudWatch Logs permissions to the EC2
This is the simplest supported approach; on EC2 launch type, awslogs uses the container instance IAM role to write to CloudWatch Logs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is the simplest supported approach; on EC2 launch type, awslogs uses the container instance IAM role to write to CloudWatch Logs.
● Scene fit: This is the simplest supported approach; on EC2 launch type, awslogs uses the container instance IAM role to write to CloudWatch Logs.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but introduces extra configuration and components, so it is not the most straightforward option.
● On EC2, the awslogs driver does not use the task role; it uses the instance profile credentials.
● Can work but is operationally heavier than using the native awslogs log driver.
Workflow: Use awslogs in the task definition → grant CloudWatch Logs permissions to the EC2 instance profile.
Single repo with develop->main via PRs, CodeBuild on commit, CodeDeploy blue/green with
Blue/green with traffic shifting provides near-zero downtime and very fast rollback by flipping traffic between environments.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Blue/green with traffic shifting provides near-zero downtime and very fast rollback by flipping traffic between environments.
● Scene fit: Blue/green with traffic shifting provides near-zero downtime and very fast rollback by flipping traffic between environments.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● In-place rolling can cause partial interruption and rollback requires redeploying the old version across the fleet, which is slower.
● Multiple repos add coordination overhead without improving rollback speed or availability versus a single-repo approach.
● Rolling updates can reduce capacity and make rollback slower compared to an immediate blue/green traffic switch.
Workflow: Single repo with develop->main via PRs, CodeBuild on commit, CodeDeploy blue/green with traffic shifting.
All at once deployment policy for new versions
All at once updates all instances simultaneously for the fastest deployment, with brief downtime acceptable in staging.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: All at once updates all instances simultaneously for the fastest deployment, with brief downtime acceptable in staging.
● Scene fit: All at once updates all instances simultaneously for the fastest deployment, with brief downtime acceptable in staging.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Rolling updates proceed in batches, which slows the overall deployment and temporarily reduces capacity.
● Blue/green requires running a duplicate environment, increasing cost and adding provisioning time before the swap.
● Immutable spins up new instances for each release, which is safer but slower and incurs extra cost.
Workflow: All at once deployment policy for new versions.
Health endpoint returns 200 only if DB reachable; set ALB health check
Expose an app health URL that tests DB connectivity and returns non-200 when unreachable so the ALB marks the target unhealthy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Expose an app health URL that tests DB connectivity and returns non-200 when unreachable so the ALB marks the target unhealthy.
● Scene fit: Expose an app health URL that tests DB connectivity and returns non-200 when unreachable so the ALB marks the target unhealthy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Operates at DNS level and cannot remove individual ALB targets.
● Checks instance/system health, not application or database reachability.
● ALB health checks evaluate HTTP status codes and timeouts, not response payloads.
Workflow: Health endpoint returns 200 only if DB reachable → set ALB health check to that path.
CloudTrail with an EventBridge rule for DeleteTable to SNS
CloudTrail records the API call and EventBridge matches AWS API Call via CloudTrail events to send immediate notifications to SNS at low cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail records the API call and EventBridge matches AWS API Call via CloudTrail events to send immediate notifications to SNS at low cost.
● Scene fit: CloudTrail records the API call and EventBridge matches AWS API Call via CloudTrail events to send immediate notifications to SNS at low cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but requires delivering CloudTrail to CloudWatch Logs and using metric filters and alarms, adding cost and potential delay.
● Config rules evaluate periodically and are not intended for near real-time API call detection.
● Streams capture item-level changes and do not emit admin API events like DeleteTable.
Workflow: CloudTrail with an EventBridge rule for DeleteTable to SNS.
an Amazon EventBridge scheduled rule to invoke AWS Step Functions, which launches
This orchestrates a single ephemeral instance just for scanning the AMI, minimizing cost and impact while ensuring daily CVE coverage.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This orchestrates a single ephemeral instance just for scanning the AMI, minimizing cost and impact while ensuring daily.
● Scene fit: This orchestrates a single ephemeral instance just for scanning the AMI, minimizing cost and impact while ensuring daily.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Inspector cannot assess an AMI directly by ID and requires an EC2 instance as the target for evaluation.
● Scanning every instance repeatedly is redundant for a single golden AMI, drives up cost, and may cause unnecessary impact.
● Amazon Inspector does not accept an AMI ID as a direct target and cannot scan an image without an.
Workflow: Configure an Amazon EventBridge scheduled rule to invoke AWS Step Functions, which launches a short lived EC2 instance from the hardened AMI, tags it VulnScan: Yes, runs an Amazon Inspector assessment template.
EC2 Auto Scaling SNS notifications for EC2_INSTANCE_LAUNCH_ERROR
Built-in Auto Scaling notifications publish launch failure events directly to SNS for immediate alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Built-in Auto Scaling notifications publish launch failure events directly to SNS for immediate alerts.
● Scene fit: Built-in Auto Scaling notifications publish launch failure events directly to SNS for immediate alerts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Status check alarms fire after an instance is running, not when the launch attempt fails.
● Only captures API errors from RunInstances and can miss Auto Scaling launch failures that do not surface as API errors.
● May lag and indicates capacity shortfalls, not specifically failed launch attempts.
Workflow: Enable EC2 Auto Scaling SNS notifications for EC2_INSTANCE_LAUNCH_ERROR.
an EventBridge rule for CodeDeploy state-change events that invokes Lambda to post
EventBridge reacts to exact deployment state changes, Lambda sends Slack messages, and CodeDeploy handles native auto-rollback on failures.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge reacts to exact deployment state changes, Lambda sends Slack messages, and CodeDeploy handles native auto-rollback on failures.
● Scene fit: EventBridge reacts to exact deployment state changes, Lambda sends Slack messages, and CodeDeploy handles native auto-rollback on failures.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Notifies Slack but relies on custom rollback logic instead of CodeDeploy’s native automatic rollback.
● Metric alarms are indirect and can lag; they do not directly leverage state-change events and still require proper rollback configuration.
● Pipeline notifications are not specific to CodeDeploy state changes and require manual rollback scripting.
Workflow: Create an EventBridge rule for CodeDeploy state-change events that invokes Lambda to post to Slack, and enable CodeDeploy automatic rollback on failure.
Amazon ElastiCache for Redis (shared session store)
Provides an in-memory, shared session store with microsecond latency and decouples sessions from EC2 lifecycle.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides an in-memory, shared session store with microsecond latency and decouples sessions from EC2 lifecycle.
● Scene fit: Provides an in-memory, shared session store with microsecond latency and decouples sessions from EC2 lifecycle.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● In-memory Redis with durability across AZs; typically slightly higher latency than ElastiCache for Redis, not optimal for lowest-latency session reads.
● Binds users to instances but sessions are lost when instances are replaced or scaled in, so persistence still fails.
● Durable and scalable but per-request session lookups have higher latency than in-memory cache solutions.
Workflow: Amazon ElastiCache for Redis (shared session store).
a bucket policy that limits reads to the company's AWS accounts and
A scoped bucket policy plus removing the AuthenticatedUsers canned ACL ensures only approved principals can access objects and eliminates broad access.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A scoped bucket policy plus removing the AuthenticatedUsers canned ACL ensures only approved principals can access objects and eliminates broad access.
● Scene fit: A scoped bucket policy plus removing the AuthenticatedUsers canned ACL ensures only approved principals can access objects and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Encryption at rest does not alter object permissions, so anyone with current read privileges would still be able to access the data.
● Making the objects public would grant worldwide anonymous access and significantly reduce security.
● Adding CloudFront does not negate the S3 ACL, so the objects would remain readable by any AWS account unless the ACL and policy are corrected.
Workflow: Create a bucket policy that limits reads to the company's AWS accounts and remove the authenticated-read ACL from the upload step.
Connect Lambda validation to the AfterAllowTestTraffic lifecycle hook in AppSpec.yaml so tests
AfterAllowTestTraffic is designed to validate the replacement task set using test traffic, enabling safe rollback before any production cutover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AfterAllowTestTraffic is designed to validate the replacement task set using test traffic, enabling safe rollback before any production cutover.
● Scene fit: AfterAllowTestTraffic is designed to validate the replacement task set using test traffic, enabling safe rollback before any production cutover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● BeforeAllowTraffic runs after test validation should already be complete, so it is not the correct place to execute initial verification tests.
● AfterAllowTraffic occurs after production traffic is routed to the new task set, which is too late for pre-production validation.
● AfterInstall happens before the test listener sends traffic to the new task set, so validation using test traffic cannot occur here.
Workflow: Connect Lambda validation to the AfterAllowTestTraffic lifecycle hook in AppSpec.yaml so tests run against the test listener and can trigger automatic rollback.
an organization-wide AWS Config rule to evaluate EBS encryption by default and
This centrally deploys and enforces the compliance check across the organization with minimal overhead and protects the control from tampering.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This centrally deploys and enforces the compliance check across the organization with minimal overhead and protects the control from tampering.
● Scene fit: This centrally deploys and enforces the compliance check across the organization with minimal overhead and protects the control from tampering.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Prevents some new noncompliant launches but does not evaluate existing volumes or provide continuous compliance reporting across accounts.
● Is operationally heavy and requires custom code, scheduling, and per-account deployment and maintenance.
● Can work but is less efficient than organization-level AWS Config rules and still requires additional guardrails to prevent Config from being disabled.
Workflow: Create an organization-wide AWS Config rule to evaluate EBS encryption by default and attach an SCP that prevents disabling or deleting AWS Config in any account.
AWS Elastic Beanstalk with ALB/Auto Scaling, external Amazon RDS MySQL Multi-AZ, logs
Provides managed rolling or immutable deployments with quick rollback, decouples a shared production database, and uses CloudWatch Logs and Logs Insights for centralized, near-real-time search.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides managed rolling or immutable deployments with quick rollback, decouples a shared production database, and uses CloudWatch Logs and Logs Insights for centralized, near-real-time search.
● Scene fit: Provides managed rolling or immutable deployments with quick rollback, decouples a shared production database, and uses CloudWatch Logs and Logs Insights for centralized, near-real-time search.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Viable but higher operational complexity and cost compared to simpler managed app deployment and CloudWatch-based logging.
● Lacks Multi-AZ for the database and requires more custom deployment and rollback orchestration.
● Couples database lifecycle to the application environment, which is risky for shared databases and rollbacks.
Workflow: AWS Elastic Beanstalk with ALB/Auto Scaling, external Amazon RDS MySQL Multi-AZ, logs to CloudWatch Logs with 90-day retention.
AWS CodePipeline with CodeBuild and CodeDeploy using CodeDeployDefault.LambdaLinear10PercentEvery2Minutes
Combines managed CI/CD with automated tests and controlled linear traffic shifting plus built-in rollback via CloudWatch alarms.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Combines managed CI/CD with automated tests and controlled linear traffic shifting plus built-in rollback via CloudWatch alarms.
● Scene fit: Combines managed CI/CD with automated tests and controlled linear traffic shifting plus built-in rollback via CloudWatch alarms.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● SAM supports canary/linear deployments via CodeDeploy, but a CLI-driven flow is not a fully managed pipeline with automated tests and triggers.
● Custom orchestration adds undifferentiated work and lacks native rollback and deployment safety features.
● All-at-once moves 100% of traffic immediately and does not meet gradual shift requirements.
Workflow: AWS CodePipeline with CodeBuild and CodeDeploy using CodeDeployDefault.LambdaLinear10PercentEvery2Minutes.
EventBridge rule for aws.health AWS_RISK_CREDENTIALS_EXPOSED to Step Functions
Match aws.health events for AWS_RISK_CREDENTIALS_EXPOSED in EventBridge and orchestrate deletion, investigation, and notification with Step Functions and Lambda.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Match aws.health events for AWS_RISK_CREDENTIALS_EXPOSED in EventBridge and orchestrate deletion, investigation, and notification with Step Functions and Lambda.
● Scene fit: Match aws.health events for AWS_RISK_CREDENTIALS_EXPOSED in EventBridge and orchestrate deletion, investigation, and notification with Step Functions and Lambda.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● These services do not detect public IAM key exposure events nor emit the AWS Health event used for this workflow.
● AWS Config evaluates resource configurations and does not surface AWS Health events such as exposed credentials.
● Detective supports investigation but does not generate or route AWS Health exposed-credentials events for remediation.
Workflow: EventBridge rule for aws.health AWS_RISK_CREDENTIALS_EXPOSED to Step Functions.
BeforeAllowTraffic lifecycle hook
Pre-traffic hook for Lambda that can wait for DB migrations or seed data to complete before shifting traffic.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Pre-traffic hook for Lambda that can wait for DB migrations or seed data to complete before shifting traffic.
● Scene fit: Pre-traffic hook for Lambda that can wait for DB migrations or seed data to complete before shifting traffic.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs after traffic has already shifted, so it cannot block errors caused by incomplete DB changes.
● Canary alone does not validate DB readiness and won’t block the shift if changes aren’t complete.
● Not a supported lifecycle event for Lambda in CodeDeploy.
Workflow: BeforeAllowTraffic lifecycle hook.
a read replica with CloudFormation using SourceDBInstanceIdentifier, wait for it to catch
Upgrading a read replica and promoting it allows a short cutover window while the source continues serving traffic.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Upgrading a read replica and promoting it allows a short cutover window while the source continues serving traffic.
● Scene fit: Upgrading a read replica and promoting it allows a short cutover window while the source continues serving traffic.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Changing EngineVersion on the live instance through CloudFormation can trigger replacement or a disruptive in-place upgrade, which increases downtime.
● DBEngineVersion is not a valid AWS::RDS::DBInstance property, so this change would fail validation.
● While DMS can reduce downtime, it adds unnecessary complexity for a same-engine major upgrade that can be handled with RDS read replicas.
Workflow: Create a read replica with CloudFormation using SourceDBInstanceIdentifier, wait for it to catch up → update the replica's EngineVersion to 8.0 → promote it, → point applications to the promoted instance.
Target tracking at 75% CPU plus scheduled actions setting min 6 at
Target tracking maintains the utilization setpoint, while scheduled actions adjust baseline capacity for known peaks and troughs to improve efficiency and availability.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Target tracking maintains the utilization setpoint, while scheduled actions adjust baseline capacity for known peaks and troughs to improve efficiency and availability.
● Scene fit: Target tracking maintains the utilization setpoint, while scheduled actions adjust baseline capacity for known peaks and troughs to improve efficiency and availability.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● May reduce cost but risks interruptions and does not maintain a utilization target or handle predictable baselines.
● Forecasting desired capacity helps with timing but does not hold average CPU near a target or tune baseline minimums directly.
● Terminating instances out-of-band is brittle and the Auto Scaling group will replace them; this does not properly manage baseline capacity.
Workflow: Target tracking at 75% CPU → scheduled actions setting min 6 at peak and 3 off-peak.
an Amazon EventBridge rule that invokes an AWS Lambda function every 30
A scheduled EventBridge rule invoking Lambda can diff EC2 against SSM managed inventory and send alerts for uncovered instances. Systems Manager Inventory natively gathers software package.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A scheduled EventBridge rule invoking Lambda can diff EC2 against SSM managed inventory and send alerts for uncovered.
● Scene fit: A scheduled EventBridge rule invoking Lambda can diff EC2 against SSM managed inventory and send alerts for uncovered.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Run Command can only target instances that are already managed, so it cannot find unmanaged instances.
● Inspector focuses on vulnerability findings and coverage rather than producing a general-purpose software inventory or detecting instances not managed.
● Is higher effort and OS-specific compared to using built-in Inventory and does not inherently detect unmanaged instances.
Workflow: Create an Amazon EventBridge rule that invokes an AWS Lambda function every 30 minutes to compare EC2 instances with Systems Manager managed instances and notify on gaps → Install the SSM Agent.
AWS Step Functions with per-task retries
Provides serverless stateful orchestration, built-in retries and catchers, and the ability to re-run only failed steps.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides serverless stateful orchestration, built-in retries and catchers, and the ability to re-run only failed steps.
● Scene fit: Provides serverless stateful orchestration, built-in retries and catchers, and the ability to re-run only failed steps.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Older workflow service that requires deciders and workers, leading to more operational complexity than a serverless state machine.
● Decouples stages but needs custom state management and orchestration logic, increasing operational burden.
● Manages DAGs but adds environment and worker management overhead compared to a serverless orchestrator.
Workflow: AWS Step Functions with per-task retries.
CloudWatch agent procstat with a CloudWatch alarm that triggers Systems Manager Run
Procstat emits per-process metrics and an alarm can invoke SSM Run Command to restart only the failed process without affecting Auto Scaling capacity.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Procstat emits per-process metrics and an alarm can invoke SSM Run Command to restart only the failed process without affecting Auto Scaling capacity.
● Scene fit: Procstat emits per-process metrics and an alarm can invoke SSM Run Command to restart only the failed process without affecting Auto Scaling capacity.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Terminates and replaces instances, slowing recovery and reducing capacity unnecessarily.
● Auto Recovery responds to system status check failures, not individual process crashes, and reboots the whole instance.
● Cycling instances via Standby reduces capacity and reboots entire instances rather than just the failed process.
Workflow: Use CloudWatch agent procstat with a CloudWatch alarm that triggers Systems Manager Run Command to restart the worker.
AWS Config managed rule for EBS encryption with EventBridge filtering to SNS
The managed rule continuously evaluates EBS volume encryption, and EventBridge filters that rule’s NON_COMPLIANT events to SNS for targeted alerts and an auditable history in AWS Config.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The managed rule continuously evaluates EBS volume encryption, and EventBridge filters that rule’s NON_COMPLIANT events to SNS for targeted alerts and an auditable history in AWS Config.
● Scene fit: The managed rule continuously evaluates EBS volume encryption, and EventBridge filters that rule’s NON_COMPLIANT events to SNS for.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Security Hub aggregates findings from sources like AWS Config and adds setup and cost, making it less direct for a single control.
● Default encryption prevents new unencrypted volumes but does not detect or alert on existing unencrypted resources and provides no targeted notifications.
● AWS Config’s SNS delivery channel emits broad notifications and cannot isolate a single rule, leading to noisy, non-targeted alerts.
Workflow: AWS Config managed rule for EBS encryption with EventBridge filtering to SNS.
Define a patch baseline with AWS Systems Manager Patch Manager and use
Systems Manager Patch Manager automates patching via baselines while AWS Config managed rules such as approved-amis-by-id check AMI compliance, and alarms can notify on detected drift.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Systems Manager Patch Manager automates patching via baselines while AWS Config managed rules such as approved-amis-by-id check AMI.
● Scene fit: Systems Manager Patch Manager automates patching via baselines while AWS Config managed rules such as approved-amis-by-id check AMI.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● GuardDuty focuses on threat detection and does not assess patch levels or validate approved AMI usage, so it will.
● Denying launches with IAM blocks developers from using unapproved AMIs, which conflicts with the requirement to allow those launches.
Workflow: Define a patch baseline with AWS Systems Manager Patch Manager and use an AWS Config managed rule to evaluate instances against an approved AMI list, with CloudWatch alarms for any noncompliant resources.
PUT a success or failure to the CloudFormation ResponseURL from the Lambda
Custom resources must send a result to the pre-signed ResponseURL so CloudFormation can transition the stack state.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Custom resources must send a result to the pre-signed ResponseURL so CloudFormation can transition the stack state.
● Scene fit: Custom resources must send a result to the pre-signed ResponseURL so CloudFormation can transition the stack state.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Permission does not signal custom resource completion and is not required for finishing the stack operation.
● A normal Lambda exit does not notify CloudFormation; it still waits for the custom resource response.
● Wait conditions and cfn-signal apply to WaitCondition or EC2 CreationPolicy, not to custom resources.
Workflow: PUT a success or failure to the CloudFormation ResponseURL from the Lambda.
an Amazon EventBridge rule for EC2 Auto Scaling launch and terminate events
EventBridge can match Auto Scaling lifecycle events and invoke Systems Manager Run Command to execute an update on the managed batch instance in near real time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can match Auto Scaling lifecycle events and invoke Systems Manager Run Command to execute an update on.
● Scene fit: EventBridge can match Auto Scaling lifecycle events and invoke Systems Manager Run Command to execute an update on.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on constant polling and custom code, which is inefficient and not event driven.
● AWS Config focuses on configuration compliance and can introduce latency, making it unsuitable for near real-time Auto Scaling lifecycle reactions.
● Auto Scaling lifecycle events are delivered via EventBridge, not CloudWatch Logs, and having Lambda SSH into EC2 is brittle compared to using SSM.
Workflow: Create an Amazon EventBridge rule for EC2 Auto Scaling launch and terminate events that targets AWS Systems Manager Run Command to update the batch instance configuration.
a canary release on the API Gateway stage serving v2, deploy the
API Gateway stage canaries natively split a percentage of requests to a canary deployment and expose detailed metrics in CloudWatch.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: API Gateway stage canaries natively split a percentage of requests to a canary deployment and expose detailed metrics.
● Scene fit: API Gateway stage canaries natively split a percentage of requests to a canary deployment and expose detailed metrics.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Lambda aliases do not control API Gateway routing for a new path, and OpenSearch is not the native metrics.
● Route 53 weighting operates at DNS and cannot target a single API path, and CloudTrail is not intended for.
● Alias canaries affect Lambda versions but do not provide API Gateway stage-level traffic shifting for a new route.
Workflow: Enable a canary release on the API Gateway stage serving v2 → deploy the updated API to that stage, direct a small percentage of traffic to the canary deployment, and monitor CloudWatch metrics.
Tag Editor in each account and Region to find and bulk-tag existing
Tag Editor enables discovery and bulk application of tags, helping fix current untagged resources. An SCP using aws:RequestTag or aws:TagKeys conditions can block Create* actions that omit mandatory tags.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Tag Editor enables discovery and bulk application of tags, helping fix current untagged resources.
● Scene fit: Tag Editor enables discovery and bulk application of tags, helping fix current untagged resources. An SCP using aws:RequestTag.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Cost allocation tags only affect cost reporting once tags exist; they do not apply tags or block resource creation.
● Cost Categories group spend but neither tag resources nor enforce tag requirements.
● Tag Policies standardize and report on tags but do not prevent noncompliant resource creation or auto-apply tags.
Workflow: Use Tag Editor in each account and Region to find and bulk-tag existing resources → Apply an SCP at the Organizations root denying creates without required tags.
Amazon Route 53 latency-based routing with health checks to direct users to
Latency-based routing sends clients to the lowest-latency, healthy endpoint across Regions. Creating a regional ALB and Auto Scaling capacity in Europe brings compute closer to European.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency-based routing sends clients to the lowest-latency, healthy endpoint across Regions.
● Scene fit: Latency-based routing sends clients to the lowest-latency, healthy endpoint across Regions. Creating a regional ALB and Auto Scaling.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFront can help cache static content, but it will not solve write latency to a distant Region or provide.
● DynamoDB does not offer generic cross-Region replication; Global Tables is the supported approach.
● ALBs and target groups are Regional resources and cannot register instances from another Region.
Workflow: Configure Amazon Route 53 latency-based routing with health checks to direct users to the closest ALB → Deploy a new Application Load Balancer and an Auto Scaling group in eu-west-3 and configure.
AWS Service Catalog with constraints
Enables versioned products and launch constraints to enforce tags and control where products are available.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Enables versioned products and launch constraints to enforce tags and control where products are available.
● Scene fit: Enables versioned products and launch constraints to enforce tags and control where products are available.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides detective governance via rules but cannot prevent noncompliant launches or offer versioned templates.
● Can restrict regions and use tag conditions but lacks versioned product management and self-service catalog.
● Coordinates stack deployment across accounts and regions but does not enforce mandatory tags or template versioning.
Workflow: AWS Service Catalog with constraints.
BeforeAllowTraffic hook that invokes a validator Lambda for the live-v3 stage
Runs before the shift and allows gating until the API responds.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Runs before the shift and allows gating until the API responds.
● Scene fit: Runs before the shift and allows gating until the API responds.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Alarms can trigger rollback but do not block the pre-traffic phase for Lambda deployments.
● Tests earlier but does not integrate with CodeDeploy's pre-traffic gate.
● Runs after shifting traffic, so it cannot prevent an unhealthy cutover.
Workflow: BeforeAllowTraffic hook that invokes a validator Lambda for the live-v3 stage.
In account A, create a customer managed AWS KMS key that allows
The pipeline action in account B must be able to read and decrypt artifacts stored in account A, which requires both KMS permissions and an S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The pipeline action in account B must be able to read and decrypt artifacts stored in account A.
● Scene fit: The pipeline action in account B must be able to read and decrypt artifacts stored in account A.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● The CloudFormation execution role must exist in the target account where resources are created, not in the source account.
● Points trust in the wrong direction and does not provide a role in account B for the pipeline to.
● SCPs set permission guardrails and cannot grant the needed permissions or trust relationships for cross-account actions.
Workflow: In account A → create a customer managed AWS KMS key that allows use by the CodePipeline service role in account A and by principals in account B, and create an S3.
AWS App2Container to discover the Java workload, containerize it for Amazon ECS
App2Container inventories Java apps, builds container images and task definitions, and can bootstrap CodeBuild and CodeDeploy pipelines.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: App2Container inventories Java apps, builds container images and task definitions, and can bootstrap CodeBuild and CodeDeploy pipelines.
● Scene fit: App2Container inventories Java apps, builds container images and task definitions, and can bootstrap CodeBuild and CodeDeploy pipelines.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Proton organizes and deploys standardized infrastructure templates but does not refactor or containerize existing applications.
● Application Migration Service lifts and shifts VMs without converting applications to containers or generating container-focused CI/CD.
● Copilot streamlines deploying new container applications but does not analyze existing VMs or migrate legacy workloads.
Workflow: Use AWS App2Container to discover the Java workload, containerize it for Amazon ECS, and have A2C scaffold a CI/CD pipeline with CodeBuild and CodeDeploy.
Place the instance into Standby immediately after it becomes InService
Standby removes the instance from traffic and scaling activities while keeping it in the group, allowing unlimited time to troubleshoot.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Standby removes the instance from traffic and scaling activities while keeping it in the group, allowing unlimited time to troubleshoot.
● Scene fit: Standby removes the instance from traffic and scaling activities while keeping it in the group, allowing unlimited time to troubleshoot.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Prevents new instances from starting but does not isolate or preserve a specific instance that is failing health checks.
● Termination protection blocks manual termination but the Auto Scaling group can still terminate the instance on failed health checks.
● A termination hook is time limited and triggers only during termination, which does not provide an open-ended, in-service debugging window.
Workflow: Place the instance into Standby immediately after it becomes InService.
Amazon GuardDuty
GuardDuty is a managed threat detection service that analyzes CloudTrail events, VPC flow logs, and DNS logs to identify compromised instances and malicious activity like crypto-mining.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty is a managed threat detection service that analyzes CloudTrail events, VPC flow logs, and DNS logs to identify compromised instances and malicious activity like crypto-mining.
● Scene fit: GuardDuty is a managed threat detection service that analyzes CloudTrail events, VPC flow logs, and DNS logs to identify compromised instances and malicious activity like crypto-mining.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Macie focuses on discovering and protecting sensitive data in Amazon S3 and does not detect compromised EC2 instances or account-level threats.
● Inspector performs automated vulnerability and exposure assessments for workloads but does not provide continuous threat detection from CloudTrail, VPC flow logs, and DNS data.
● VPC Flow Logs capture network traffic metadata only and require external analysis, offering no built-in threat intelligence or anomaly detection.
Workflow: Amazon GuardDuty.
AWS WAF logging to Kinesis Data Firehose with S3 destination
Configuring AWS WAF to log to Kinesis Data Firehose delivers detailed per-request JSON records that can be stored durably in S3 for analysis and retention.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Configuring AWS WAF to log to Kinesis Data Firehose delivers detailed per-request JSON records that can be stored durably in S3 for analysis and retention.
● Scene fit: Configuring AWS WAF to log to Kinesis Data Firehose delivers detailed per-request JSON records that can be stored durably in S3 for analysis and retention.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch metrics show counters and rates, not per-request log records with matched rule details.
● AWS WAF writes directly to Kinesis Data Firehose; no Kinesis Data Streams stream is required.
● ALB access logs are separate and do not include AWS WAF rule evaluation details or matched rules.
Workflow: Enable AWS WAF logging to Kinesis Data Firehose with S3 destination.
GuardDuty organization with delegated admin; route findings via EventBridge to S3 through
GuardDuty supports org-wide enablement with a delegated admin and publishes findings to EventBridge, which can deliver to S3 via Firehose.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty supports org-wide enablement with a delegated admin and publishes findings to EventBridge, which can deliver to S3 via Firehose.
● Scene fit: GuardDuty supports org-wide enablement with a delegated admin and publishes findings to EventBridge, which can deliver to S3 via Firehose.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Inspector focuses on vulnerabilities/exposure and not managed threat detections from log analytics for EC2 attacks.
● Security Hub aggregates and normalizes findings but does not perform EC2 threat detection by itself.
● Lacks centralized org administration and Kinesis Data Streams does not write to S3 without a custom consumer.
Workflow: Enable GuardDuty organization with delegated admin → route findings via EventBridge to S3 through Kinesis Data Firehose.
a single buildspec.yml that reads the CODEBUILD_SOURCE_VERSION environment variable at runtime and
This uses a built-in variable that indicates the source version or branch, enabling simple and scalable artifact naming.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses a built-in variable that indicates the source version or branch, enabling simple and scalable artifact naming.
● Scene fit: This uses a built-in variable that indicates the source version or branch, enabling simple and scalable artifact naming.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Introduces unnecessary components and post-processing instead of using the context available during the build.
● Creates excessive project sprawl and manual management per branch, which is not the simplest solution.
● Adds significant operational overhead with many pipelines and is unnecessary for basic branch-based naming.
Workflow: Create a single buildspec.yml that reads the CODEBUILD_SOURCE_VERSION environment variable at runtime and reference it in the artifacts name.
up Amazon CloudFront for S3 and deploy DynamoDB Accelerator
CloudFront caches S3 objects at edge locations and DAX provides in-memory caching for DynamoDB reads, reducing duplicate read latency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFront caches S3 objects at edge locations and DAX provides in-memory caching for DynamoDB reads, reducing duplicate read latency.
● Scene fit: CloudFront caches S3 objects at edge locations and DAX provides in-memory caching for DynamoDB reads, reducing duplicate read latency.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Transfer Acceleration speeds uploads to S3 and increasing RCUs does not cache repeated reads or provide global edge caching.
● Lambda@Edge is not a large-scale cache and global tables replicate data but do not reduce repetitive read load.
● Redis does not natively cache DynamoDB API reads and MediaStore targets media workflows rather than general S3 static assets.
Workflow: Set up Amazon CloudFront for S3 and deploy DynamoDB Accelerator.
modular CloudFormation templates per logical component and share required values via Outputs
This approach cleanly decouples stacks and uses cross-stack references for reliable value sharing and reuse.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This approach cleanly decouples stacks and uses cross-stack references for reliable value sharing and reuse.
● Scene fit: This approach cleanly decouples stacks and uses cross-stack references for reliable value sharing and reuse.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Mixes application migration tooling with infrastructure-as-code and keeps a brittle monolith, which is not ideal for frequent changes.
● Nested stacks are valid but keep tight lifecycle coupling and lack the explicit exported outputs needed for clean inter-stack reuse.
● Storing outputs in DynamoDB is not a CloudFormation best practice for wiring stacks and bypasses native output and import mechanisms.
Workflow: Create modular CloudFormation templates per logical component and share required values via Outputs with Export and Fn::ImportValue, with templates versioned in GitHub.
Amazon FSx for NetApp ONTAP with SnapMirror cross-Region replication
Provides multi-protocol SMB and NFS on the same filesystem with efficient, incremental SnapMirror replication ideal for pilot-light DR.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides multi-protocol SMB and NFS on the same filesystem with efficient, incremental SnapMirror replication ideal for pilot-light DR.
● Scene fit: Provides multi-protocol SMB and NFS on the same filesystem with efficient, incremental SnapMirror replication ideal for pilot-light DR.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Lustre does not support SMB and AWS Backup copies are periodic, not near-continuous.
● Supports SMB only; no NFS and DFSR is not storage-level, near-continuous cross-Region replication.
● Splits protocols across two backends and uses batch sync, not a single multi-protocol store or near-continuous replication.
Workflow: Amazon FSx for NetApp ONTAP with SnapMirror cross-Region replication.
Single CodePipeline with cross-Region actions and per-Region artifact buckets
CodePipeline supports cross-Region actions and creates a separate artifact store in each Region, satisfying data residency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodePipeline supports cross-Region actions and creates a separate artifact store in each Region, satisfying data residency.
● Scene fit: CodePipeline supports cross-Region actions and creates a separate artifact store in each Region, satisfying data residency.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but adds unnecessary operational overhead; a single pipeline can orchestrate cross-Region actions more efficiently.
● Useful for multi-Region infrastructure provisioning, but not for CI/CD orchestration or artifact residency in CodePipeline.
● Centralizing artifacts in one bucket violates the requirement to keep artifacts within their respective Regions.
Workflow: Single CodePipeline with cross-Region actions and per-Region artifact buckets.
AWS Step Functions orchestrating AWS Lambda tasks, triggered by Amazon EventBridge with
This provides serverless orchestration with retries and conditional branches for multi-Region fallback, native execution history for auditing, scheduled runs, and a clean way to send final.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides serverless orchestration with retries and conditional branches for multi-Region fallback, native execution history for auditing, scheduled.
● Scene fit: This provides serverless orchestration with retries and conditional branches for multi-Region fallback, native execution history for auditing, scheduled.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Concentrates all logic in one Lambda with a 15-minute timeout and relies on AWS Config, which does not track.
● Approach introduces a single point of failure and operational overhead, and it lacks the managed retries and stateful orchestration.
● While capable, MWAA is heavier to operate and not as efficient for a simple daily backup workflow compared to.
Workflow: AWS Step Functions orchestrating AWS Lambda tasks, triggered by Amazon EventBridge with notifications via Amazon SNS.
Non-empty S3 bucket; add a Lambda-backed custom resource to empty objects and
CloudFormation cannot delete a non-empty bucket, so a custom resource must remove all objects, versions, and delete markers during the Delete event.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation cannot delete a non-empty bucket, so a custom resource must remove all objects, versions, and delete markers during the Delete event.
● Scene fit: CloudFormation cannot delete a non-empty bucket, so a custom resource must remove all objects, versions, and delete markers during the Delete event.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DeletionPolicy controls retain or delete of the resource but does not purge bucket contents.
● Stack policies restrict updates, not stack deletions, and won't resolve a bucket deletion failure.
● Wait conditions don't address the constraint that a bucket must be empty to delete.
Workflow: Non-empty S3 bucket → add a Lambda-backed custom resource to empty objects and versions on stack Delete.
Amazon Inspector for EC2 vulnerability and exposure scans, install the CloudWatch Agent
Amazon Inspector continuously detects CVEs and unintended network reachability on EC2, while CloudWatch Logs aggregates instance login logs and CloudTrail provides API activity auditing in one.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Amazon Inspector continuously detects CVEs and unintended network reachability on EC2, while CloudWatch Logs aggregates instance login logs.
● Scene fit: Amazon Inspector continuously detects CVEs and unintended network reachability on EC2, while CloudWatch Logs aggregates instance login logs.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● GuardDuty provides threat detection and anomaly alerts but does not perform host vulnerability scanning or collect operating system login.
● SSM Agent and Automation support patching workflows, but they are not a vulnerability management scanner and do not evaluate.
● ECR scanning evaluates container images and not EC2 host operating systems or their login activity.
Workflow: Deploy Amazon Inspector for EC2 vulnerability and exposure scans → install the CloudWatch Agent to forward login logs to CloudWatch Logs, and send CloudTrail events to CloudWatch Logs for centralized auditing.
Lambda precheck in CodePipeline using AWS Health API to block runs during
Implements a health-aware gate that fails fast or pauses when the Region has active incidents, preventing mid-run stalls and wasted cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Implements a health-aware gate that fails fast or pauses when the Region has active incidents, preventing mid-run stalls and wasted cost.
● Scene fit: Implements a health-aware gate that fails fast or pauses when the Region has active incidents, preventing mid-run stalls and wasted cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Retries later but does not block starting during an active regional incident, so runs can still fail mid-deployment.
● Reduces cutover risk but does not avoid starting during a regional health event and adds cost.
● Repeated retries during an ongoing Region event waste time and are unlikely to succeed.
Workflow: Lambda precheck in CodePipeline using AWS Health API to block runs during active regional events.
a CloudWatch Logs subscription filter that sends matching log events to an
This connects CloudWatch Logs to Lambda for tagging and uses an EventBridge schedule to automatically terminate tagged instances within the required time window.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This connects CloudWatch Logs to Lambda for tagging and uses an EventBridge schedule to automatically terminate tagged instances.
● Scene fit: This connects CloudWatch Logs to Lambda for tagging and uses an EventBridge schedule to automatically terminate tagged instances.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail records AWS API calls, not OS-level SSH or RDP logins, so this would not detect manual logins.
● CloudWatch Logs subscriptions cannot target Step Functions and a daily cadence risks missing a 12-hour termination requirement.
● AWS Config evaluates configuration state and does not ingest instance OS logs or runtime login events, so it cannot.
Workflow: Create a CloudWatch Logs subscription filter that sends matching log events to an AWS Lambda function, tag the instance that produced the login entry, and use an Amazon EventBridge scheduled rule to.
an AWS Config rule that flags any security group with port 22
An AWS Config rule can evaluate security groups for 0.0.0.0/0 on port 22 and publish compliance notifications to Amazon SNS. Using EventBridge with Lambda to poll.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An AWS Config rule can evaluate security groups for 0.0.0.0/0 on port 22 and publish compliance notifications to.
● Scene fit: An AWS Config rule can evaluate security groups for 0.0.0.0/0 on port 22 and publish compliance notifications to.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS Config remediation is designed to use Systems Manager Automation runbooks, not direct Lambda invocations, so this is not.
● Security Hub does not ingest Trusted Advisor findings and does not natively auto-remediate them, so it does not fulfill.
Workflow: Create an AWS Config rule that flags any security group with port 22 open to 0.0.0.0/0 and sends a notification to an SNS topic when noncompliant → Schedule an Amazon EventBridge rule.
AWS Health via EventBridge invoking Lambda to Slack
AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda, which posts to Slack.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda, which posts to Slack.
● Scene fit: AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda, which posts to Slack.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Instance status checks show health issues but do not deliver AWS Health scheduled maintenance or retirement events.
● Config and Trusted Advisor assess configuration and best practices, not AWS-scheduled maintenance notifications.
● EC2 state-change events do not include AWS Health scheduled maintenance or retirements.
Workflow: AWS Health via EventBridge invoking Lambda to Slack.
a CloudTrail trail that delivers management events to a CloudWatch Logs log
CloudTrail records S3 control plane API calls and, when streamed to CloudWatch Logs, can be filtered by a metric and alarmed for immediate notification.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail records S3 control plane API calls and, when streamed to CloudWatch Logs, can be filtered by a.
● Scene fit: CloudTrail records S3 control plane API calls and, when streamed to CloudWatch Logs, can be filtered by a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 server access logs record request access details, not control plane policy changes, and are not suitable for detecting policy updates.
● S3 Event Notifications do not emit events for policy changes such as PutBucketPolicy or DeleteBucketPolicy.
● EventBridge does not directly filter CloudWatch Logs group contents, so this setup would not detect policy changes from logs.
Workflow: Create a CloudTrail trail that delivers management events to a CloudWatch Logs log group → add a metric filter for PutBucketPolicy and DeleteBucketPolicy, and configure a CloudWatch alarm to notify on matches.
Stateless EC2 Auto Scaling with SQL Server on Amazon RDS; use Kinesis
Kinesis Data Firehose offers managed, near-real-time delivery to S3 and can load Redshift via an S3 staging bucket, aligning with the stateless and low-latency requirements.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Kinesis Data Firehose offers managed, near-real-time delivery to S3 and can load Redshift via an S3 staging bucket.
● Scene fit: Kinesis Data Firehose offers managed, near-real-time delivery to S3 and can load Redshift via an S3 staging bucket.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Glue crawlers only create metadata and do not load streaming data into Redshift; MSK would still require additional ETL and staging, adding complexity and latency.
● EventBridge does not deliver directly to S3 or Redshift and is not designed for high-throughput clickstream ingestion.
● Kinesis Data Streams does not natively sink to S3 or Redshift; it requires consumers such as Firehose, Lambda, or custom applications.
Workflow: Stateless EC2 Auto Scaling with SQL Server on Amazon RDS → use Kinesis Data Firehose to S3 and another Firehose to Redshift.
Initial rotation replaced the password while the app kept a cached value
Enabling rotation changes the database password and creates a new AWSCURRENT value.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Enabling rotation changes the database password and creates a new AWSCURRENT value.
● Scene fit: Enabling rotation changes the database password and creates a new AWSCURRENT value.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A KMS denial would normally prevent secret retrieval before rotation too.
● A VPC endpoint is optional if another network path exists.
● GetSecretValue returns AWSCURRENT by default when VersionStage is omitted.
Workflow: Initial rotation replaced the password while the app kept a cached value.
with AWS Elastic Beanstalk using a load-balanced, auto scaled Node.js environment and
Elastic Beanstalk supports zero-downtime deployment strategies and simple rollback, and hosting RDS externally allows sharing the database and avoids deletion if the environment is terminated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Elastic Beanstalk supports zero-downtime deployment strategies and simple rollback, and hosting RDS externally allows sharing the database and.
● Scene fit: Elastic Beanstalk supports zero-downtime deployment strategies and simple rollback, and hosting RDS externally allows sharing the database and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While ECS can scale and handle rolling updates, easy rollbacks typically require additional tooling such as CodeDeploy or custom.
● Attaching RDS to the Beanstalk environment couples its lifecycle to the app and risks deletion on environment teardown, which.
● EBS snapshots are a backup mechanism, not an application deployment rollback strategy, and they do not ensure zero-downtime releases.
Workflow: Deploy with AWS Elastic Beanstalk using a load-balanced, auto scaled Node.js environment and create the Amazon RDS MySQL instance outside the Beanstalk environment.
CloudWatch agent to CloudWatch Logs, Firehose to S3, query with Athena
Unified agent supports hybrid collection, S3 provides low-cost storage, and Athena offers serverless queries with minimal operations.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Unified agent supports hybrid collection, S3 provides low-cost storage, and Athena offers serverless queries with minimal operations.
● Scene fit: Unified agent supports hybrid collection, S3 provides low-cost storage, and Athena offers serverless queries with minimal operations.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Directly querying in CloudWatch Logs is feasible but long-term retention and query costs are higher for audit-style use cases.
● Works for search and dashboards but requires managing capacity and is typically higher cost than S3 plus Athena for audits.
● Focused on security sources using OCSF and not intended for general OS or application logs collection from hybrid servers.
Workflow: CloudWatch agent to CloudWatch Logs, Firehose to S3 → query with Athena.
AWS CloudFormation StackSets with a delegated admin to deploy identical stacks to
StackSets natively support automated, consistent multi-Region rollouts from a central admin while keeping each Region's stacks isolated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: StackSets natively support automated, consistent multi-Region rollouts from a central admin while keeping each Region's stacks isolated.
● Scene fit: StackSets natively support automated, consistent multi-Region rollouts from a central admin while keeping each Region's stacks isolated.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Can orchestrate steps across Regions but is not the most direct mechanism for uniform multi-Region infrastructure provisioning.
● Change sets preview updates for a single stack and do not orchestrate deployments to multiple Regions.
● Possible but more complex and lacks native fleet-style multi-Region stack management like StackSets.
Workflow: AWS CloudFormation StackSets with a delegated admin to deploy identical stacks to targeted Regions.
Amazon EventBridge rules for S3 object PUT to target ECS RunTask and
This is the most direct, event-driven setup using EventBridge with CloudTrail data events to start tasks and a Lambda to stop them.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is the most direct, event-driven setup using EventBridge with CloudTrail data events to start tasks and a Lambda to stop them.
● Scene fit: This is the most direct, event-driven setup using EventBridge with CloudTrail data events to start tasks and a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds unnecessary indirection and alarms are not the most direct way to respond to individual S3 object events.
● Capacity providers manage compute capacity, not event-driven scaling of desired task count, making this approach overly complex and misaligned.
● While powerful for queued batch workloads, this is not the simplest for S3 event triggers and introduces extra orchestration components.
Workflow: Create Amazon EventBridge rules for S3 object PUT to target ECS RunTask and for S3 object DELETE to invoke a Lambda that calls StopTask on all running tasks.
AWS SAM with CodeDeploy traffic shifting, pre-traffic and post-traffic hooks, and CloudWatch
SAM's DeploymentPreference integrates with CodeDeploy to canary or linearly shift Lambda alias traffic, run validation hooks, and trigger automatic rollback on CloudWatch alarms for fast detection.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SAM's DeploymentPreference integrates with CodeDeploy to canary or linearly shift Lambda alias traffic, run validation hooks, and trigger.
● Scene fit: SAM's DeploymentPreference integrates with CodeDeploy to canary or linearly shift Lambda alias traffic, run validation hooks, and trigger.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AppConfig can toggle runtime behavior, but it does not publish Lambda versions, shift alias traffic, or automate deployment rollback based on errors.
● Change sets preview infrastructure changes but do not provide Lambda canary traffic shifting or lifecycle tests, so detection and rollback remain slow and manual.
● CloudFormation change sets have no pre-traffic or post-traffic test hooks for Lambda, making this capability unsupported in CloudFormation alone.
Workflow: AWS SAM with CodeDeploy traffic shifting, pre-traffic and post-traffic hooks, and CloudWatch alarm rollback.
S3 buckets in two separate AWS Regions at least 700 miles apart
This is correct because it uses different Regions to meet the distance requirement, enforces in-transit encryption via a bucket policy, uses SSE-S3 for encryption at rest.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is correct because it uses different Regions to meet the distance requirement, enforces in-transit encryption via a.
● Scene fit: This is correct because it uses different Regions to meet the distance requirement, enforces in-transit encryption via a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Is incorrect because S3 is a regional service, AZs are not 700 miles apart, and Transfer Acceleration does not.
● Is incorrect because an IAM role cannot enforce TLS-only access; a bucket policy must be used to require HTTPS.
Workflow: Create S3 buckets in two separate AWS Regions at least 700 miles apart, enforce HTTPS-only access with a bucket policy, require SSE-S3 for all objects, and enable S3 cross-Region replication.
AWS Serverless Application Model to deploy API Gateway and Lambda with CodeDeploy
SAM integrates Lambda aliases with CodeDeploy to shift a small percentage of traffic with automated rollback and minimal configuration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SAM integrates Lambda aliases with CodeDeploy to shift a small percentage of traffic with automated rollback and minimal configuration.
● Scene fit: SAM integrates Lambda aliases with CodeDeploy to shift a small percentage of traffic with automated rollback and minimal configuration.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Blue/green swaps entire environments rather than providing a built-in, percentage-based canary for Lambda traffic shifting to a small user segment.
● Route 53 failover is health-check based for primary/secondary scenarios and is not designed for progressive canary traffic shifting.
● AppConfig manages configuration exposure rather than code deployment and introduces extra components compared to a built-in SAM canary.
Workflow: Use AWS Serverless Application Model to deploy API Gateway and Lambda with CodeDeploy canary by setting DeploymentPreference to Canary5Percent5Minutes.
a nightly EventBridge schedule to trigger a Lambda that calls Storage Gateway
Scheduling RefreshCache updates the file gateway's cached inventory so objects added directly to S3 overnight appear on the share by morning.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Scheduling RefreshCache updates the file gateway's cached inventory so objects added directly to S3 overnight appear on the share by morning.
● Scene fit: Scheduling RefreshCache updates the file gateway's cached inventory so objects added directly to S3 overnight appear on the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Changing the protocol used to land files in S3 does not cause the file gateway to update its cached directory listing or metadata.
● RefreshCache is only supported by file gateway shares and volume gateway does not expose file semantics for this operation.
● Storage Gateway cannot subscribe to SQS or S3 events to invalidate its cache, so this integration does not exist.
Workflow: Configure a nightly EventBridge schedule to trigger a Lambda that calls Storage Gateway RefreshCache for the share.
Application Load Balancer with HTTP listener and path-based routing to Lambda target
ALB supports HTTP listeners, path-based rules, and Lambda target groups.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB supports HTTP listeners, path-based rules, and Lambda target groups.
● Scene fit: ALB supports HTTP listeners, path-based rules, and Lambda target groups.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● NLB is Layer 4 and does not provide HTTP path rules or Lambda targets.
● API Gateway can route paths to Lambda, but its public endpoints use HTTPS rather than a plain HTTP listener.
● A Function URL maps to one function and uses HTTPS.
Workflow: Application Load Balancer with HTTP listener and path-based routing to Lambda target groups.
Elastic Beanstalk immutable updates
Launches a separate temporary Auto Scaling group, shifts traffic after health checks, preserves the CNAME, and discards the new group on failure.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Launches a separate temporary Auto Scaling group, shifts traffic after health checks, preserves the CNAME, and discards the new group on failure.
● Scene fit: Launches a separate temporary Auto Scaling group, shifts traffic after health checks, preserves the CNAME, and discards the new group on failure.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Requires swapping environment CNAMEs, which involves DNS changes.
● Updates existing instances in batches and can reduce capacity during deployment.
● Adds capacity but still performs in-place updates and does not fully isolate the new version.
Workflow: Elastic Beanstalk immutable updates.
CloudWatch cross-account observability via AWS Organizations
CloudWatch cross-account observability with Organizations and OAM provides one monitoring account for metrics, logs, and traces.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudWatch cross-account observability with Organizations and OAM provides one monitoring account for metrics, logs, and traces.
● Scene fit: CloudWatch cross-account observability with Organizations and OAM provides one monitoring account for metrics, logs, and traces.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is workable, but every account must be linked manually.
● OpenSearch can centralize logs, but it is not the native unified view for all CloudWatch metrics, logs, and X-Ray traces.
● Metric Streams cover metrics, not the complete logs-and-traces workflow.
Workflow: CloudWatch cross-account observability via AWS Organizations.
AWS VM Import/Export to import the on-prem VMware image as an EC2
VM Import/Export enables round-tripping VMware images to EC2 and exporting previously imported instances back to a vSphere-compatible format for parity testing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: VM Import/Export enables round-tripping VMware images to EC2 and exporting previously imported instances back to a vSphere-compatible format.
● Scene fit: VM Import/Export enables round-tripping VMware images to EC2 and exporting previously imported instances back to a vSphere-compatible format.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Outposts would work but introduces unnecessary cost and operational overhead for simple pre-migration image validation in an existing VMware.
● Different Linux distributions have distinct kernels, repos, and libraries, so results will not accurately reflect Amazon Linux 2 behavior.
Workflow: Use AWS VM Import/Export to import the on-prem VMware image as an EC2 AMI → validate on EC2, → export the imported instance as a VMware-compatible OVA to Amazon S3 and load.
EC2 Multi-AZ ASG + ALB, Aurora multi-writer cluster, Amazon Inspector
Provides Multi-AZ app scaling, write-scalable and highly available database, and continuous vulnerability assessments with Inspector.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides Multi-AZ app scaling, write-scalable and highly available database, and continuous vulnerability assessments with Inspector.
● Scene fit: Provides Multi-AZ app scaling, write-scalable and highly available database, and continuous vulnerability assessments with Inspector.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Macie discovers sensitive data in S3 and does not perform vulnerability scanning; Aurora Global Database does not provide multi-writer scaling.
● Security Hub aggregates findings and is not a scanner; RDS PostgreSQL has a single writer and does not scale writes.
● GuardDuty is threat detection, not vulnerability scanning; a single writer does not scale writes.
Workflow: EC2 Multi-AZ ASG + ALB, Aurora multi-writer cluster, Amazon Inspector.
CloudWatch Agent streaming app logs to CloudWatch Logs + ASG termination lifecycle
The agent pushes application logs off-instance continuously to a durable, centralized log store with retention. A termination lifecycle hook can trigger automation to collect logs before instance shutdown and upload them to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The agent pushes application logs off-instance continuously to a durable, centralized log store with retention.
● Scene fit: The agent pushes application logs off-instance continuously to a durable, centralized log store with retention. A termination lifecycle.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Flow logs capture network traffic metadata, not application or system logs.
● ALB access logs record request metadata at the load balancer, not instance-level application logs.
● Manual retrieval is not reliable at scale and may miss instances that terminate quickly.
Workflow: CloudWatch Agent streaming app logs to CloudWatch Logs → ASG termination lifecycle hook with EventBridge + Lambda + SSM Run Command to S3.
ALB target group health checks configured incorrectly
Wrong health check path, port, or success codes prevent the new targets from becoming healthy, causing AllowTraffic to fail without CodeDeploy log errors.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Wrong health check path, port, or success codes prevent the new targets from becoming healthy, causing AllowTraffic to fail without CodeDeploy log errors.
● Scene fit: Wrong health check path, port, or success codes prevent the new targets from becoming healthy, causing AllowTraffic to fail without CodeDeploy log errors.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● If instances are removed mid-deployment, CodeDeploy cannot complete the transition and the AllowTraffic stage can fail.
● Hook script errors from a prior revision can break deployments but they surface in logs rather than only at AllowTraffic.
● WAF filters client requests and does not block the ALB-to-target health probes, so it would not cause a silent AllowTraffic failure.
Workflow: ALB target group health checks configured incorrectly.
Instance profile IAM role permissions/trust + Bucket policy principals/conditions
If the instance role lacks s3:GetObject or has a broken trust/policy, S3 returns 403 AccessDenied. A restrictive bucket policy (principal limits, conditions like aws:SourceVpce or s3:prefix) can deny the role and cause AccessDenied.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: If the instance role lacks s3:GetObject or has a broken trust/policy, S3 returns 403 AccessDenied.
● Scene fit: If the instance role lacks s3:GetObject or has a broken trust/policy, S3 returns 403 AccessDenied. A restrictive bucket.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Block Public Access targets public ACLs and bucket policies and does not block authorized role access to private objects.
● Object Lock enforces retention/holds and does not block reads for principals with permission.
● Only relevant if an S3 VPC endpoint is used; otherwise it does not affect authorization and is not the typical root cause.
Workflow: Instance profile IAM role permissions/trust → Bucket policy principals/conditions.
Elastic Beanstalk with an external Amazon RDS for MySQL Multi-AZ and immutable
Immutable updates launch a new Auto Scaling group at full capacity, allow quick rollback if health checks fail, and avoid partial rolling update issues while limiting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Immutable updates launch a new Auto Scaling group at full capacity, allow quick rollback if health checks fail, and avoid partial rolling update issues while limiting extra cost to the deployment window.
● Scene fit: Immutable updates launch a new Auto Scaling group at full capacity, allow quick rollback if health checks fail.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Enables zero-downtime and rollback, but keeping the old environment after cutover is not the most cost-effective choice.
● Does not satisfy the requirement to host the application on Elastic Beanstalk and would add re-architecture overhead.
● Can still lead to partially applied updates and couples the database lifecycle to the environment, which is not recommended for production.
Workflow: Elastic Beanstalk with an external Amazon RDS for MySQL Multi-AZ and immutable deployments.
API Gateway canary release sending 10% to a parallel ALB/EC2 backend
API Gateway canary releases natively split traffic by weight and allow instant rollback by adjusting the canary weight.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: API Gateway canary releases natively split traffic by weight and allow instant rollback by adjusting the canary weight.
● Scene fit: API Gateway canary releases natively split traffic by weight and allow instant rollback by adjusting the canary weight.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Blue/green works but adds deployment tooling and configuration and does not leverage API Gateway's native traffic shifting.
● DNS weighting introduces TTL delays and requires duplicate domains/stacks, making rollback and control less precise.
● API Gateway cannot directly weight across ALB target groups; traffic splitting is done via API Gateway canary at the stage level.
Workflow: API Gateway canary release sending 10% to a parallel ALB/EC2 backend.
No outbound access to CodeDeploy endpoints + Missing IAM instance profile on
Without egress to CodeDeploy public endpoints or VPC endpoints, the agent cannot communicate and events are marked Skipped. Without an instance profile granting access to fetch the revision and report status, the agent cannot run lifecycle steps.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Without egress to CodeDeploy public endpoints or VPC endpoints, the agent cannot communicate and events are marked Skipped.
● Scene fit: Without egress to CodeDeploy public endpoints or VPC endpoints, the agent cannot communicate and events are marked Skipped.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● The deployment runs under the instance profile and CodeDeploy service role, not the initiating user.
● CodeDeploy supports both targeting methods and this does not cause lifecycle events to be skipped.
● Missing optional hooks may show Skipped for those hooks, but the revision can still be installed and not all events are skipped.
Workflow: No outbound access to CodeDeploy endpoints → Missing IAM instance profile on instances.
Kinesis Data Firehose with Lambda transform to S3 from Network Firewall
Network Firewall can deliver directly to Firehose, which supports inline Lambda transformations and near real-time S3 delivery.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Network Firewall can deliver directly to Firehose, which supports inline Lambda transformations and near real-time S3 delivery.
● Scene fit: Network Firewall can deliver directly to Firehose, which supports inline Lambda transformations and near real-time S3 delivery.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Network Firewall does not natively publish to Kinesis Data Streams and this adds custom components.
● Runs after objects are written to S3, not inline before delivery.
● Requires routing logs to CloudWatch and building a custom pipeline, which increases operational overhead.
Workflow: Kinesis Data Firehose with Lambda transform to S3 from Network Firewall.
Model the workflow in AWS Step Functions and run each stage as
Step Functions provides managed orchestration, per-task retries, and state management to reprocess only failed steps with minimal ops overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Step Functions provides managed orchestration, per-task retries, and state management to reprocess only failed steps with minimal ops overhead.
● Scene fit: Step Functions provides managed orchestration, per-task retries, and state management to reprocess only failed steps with minimal ops overhead.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Design decouples stages but requires custom orchestration and state tracking, increasing operational complexity.
● Airflow can orchestrate tasks but managing MWAA environments and EC2 workers adds operational burden compared to a serverless approach.
● A monolithic Lambda makes it hard to retry only the failed step and forces full re-execution on errors.
Workflow: Model the workflow in AWS Step Functions and run each stage as a separate task that invokes AWS Lambda with retries.
an IAM role for the EBS CSI driver using IRSA and attach
The EBS CSI controller needs an IRSA-bound IAM role with EC2 permissions to create and manage gp3 volumes, which resolves UnauthorizedOperation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The EBS CSI controller needs an IRSA-bound IAM role with EC2 permissions to create and manage gp3 volumes, which resolves UnauthorizedOperation.
● Scene fit: The EBS CSI controller needs an IRSA-bound IAM role with EC2 permissions to create and manage gp3 volumes.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Only modifies Kubernetes RBAC and does not grant the EC2 permissions required by the EBS CSI driver.
● Avoids provisioning but does not fix the underlying IAM authorization issue and reduces flexibility.
● Pods do not automatically use the node IAM role and the EBS CSI controller is designed to assume an IRSA role, so this is not the correct or recommended fix.
Workflow: Configure an IAM role for the EBS CSI driver using IRSA and attach it to the add-on so it can call the required EC2 APIs.
an IAM instance profile granting Systems Manager access + Use interface VPC
Instances need an IAM role (for example, AmazonSSMManagedInstanceCore) to authorize SSM operations without access keys. These PrivateLink endpoints keep Session Manager traffic inside the AWS network.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Instances need an IAM role (for example, AmazonSSMManagedInstanceCore) to authorize SSM operations without access keys.
● Scene fit: Instances need an IAM role (for example, AmazonSSMManagedInstanceCore) to authorize SSM operations without access keys. These PrivateLink endpoints keep Session Manager traffic inside the AWS network.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Gateway endpoints support only S3 and DynamoDB, not Systems Manager.
● Session Manager does not require inbound SSH and opening port 22 increases risk.
● Long-lived access keys on instances are insecure and unnecessary when using instance profiles.
Workflow: Attach an IAM instance profile granting Systems Manager access → Use interface VPC endpoints for SSM, SSMMessages, and EC2Messages.
CloudFront field-level encryption, require HTTPS to origin, and use long max-age
Field-level encryption protects specified fields at the edge, HTTPS secures transport, and long max-age increases cache residency for higher hit rate.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Field-level encryption protects specified fields at the edge, HTTPS secures transport, and long max-age increases cache residency for higher hit rate.
● Scene fit: Field-level encryption protects specified fields at the edge, HTTPS secures transport, and long max-age increases cache residency for higher hit rate.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● WAF can filter requests but does not encrypt sensitive fields; longer TTL helps caching but not PCI field protection.
● Signed URLs control access rather than encrypting form fields; longer max-age aids caching but not PCI encryption.
● OAI restricts S3 origin access and varying on headers fragments the cache; neither provides field-level encryption.
Workflow: Enable CloudFront field-level encryption, require HTTPS to origin, and use long max-age.
Publish event payloads to an Amazon SNS topic, subscribe an AWS Lambda
SNS to Lambda provides native event-driven processing and DynamoDB is a serverless key-value store ideal for per-event writes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SNS to Lambda provides native event-driven processing and DynamoDB is a serverless key-value store ideal for per-event writes.
● Scene fit: SNS to Lambda provides native event-driven processing and DynamoDB is a serverless key-value store ideal for per-event writes.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Athena is a query service, not a general-purpose compute runtime for per-event processing, and EventBridge is unnecessary for simple S3-to-compute workflows.
● ElastiCache is an in-memory cache running on managed instances and is not a serverless durable key-value database.
● RDS is a relational database that is neither serverless nor a key-value store, which does not meet the stated requirement.
Workflow: Publish event payloads to an Amazon SNS topic → subscribe an AWS Lambda function to process each message, and persist the output to an Amazon DynamoDB table named EventStore.
The S3 bucket is not empty and CloudFormation cannot delete it; use
CloudFormation cannot delete a non-empty S3 bucket, so a custom resource that deletes all objects, versions, and delete markers during the Delete event is the correct.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation cannot delete a non-empty S3 bucket, so a custom resource that deletes all objects, versions, and delete.
● Scene fit: CloudFormation cannot delete a non-empty S3 bucket, so a custom resource that deletes all objects, versions, and delete.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Stack policies restrict updates to protected resources and do not prevent stack deletions, so this would not explain the failure.
● There is no Delete: Force property in CloudFormation for S3 buckets, so this setting cannot resolve the issue.
● WaitCondition is not applied to Lambda and would not address the non-empty bucket constraint that actually blocks deletion.
Workflow: The S3 bucket is not empty and CloudFormation cannot delete it → use a Lambda-backed custom resource to purge the bucket on stack deletion.
lifecycle hooks to each Auto Scaling group and configure an Amazon EventBridge
Lifecycle hooks emit events that EventBridge can route to Lambda with the instance details and token, enabling reliable updates on both launch and terminate.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lifecycle hooks emit events that EventBridge can route to Lambda with the instance details and token, enabling reliable.
● Scene fit: Lifecycle hooks emit events that EventBridge can route to Lambda with the instance details and token, enabling reliable.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Alarms on aggregate metrics are not tied to individual lifecycle events and do not provide per-instance context or lifecycle tokens.
● Instance scripts are brittle and may not run on termination, making them unreliable for authoritative inventory updates.
● Lambda cannot be configured as a direct lifecycle hook notification target; you must use EventBridge, SNS, or SQS instead.
Workflow: Add lifecycle hooks to each Auto Scaling group and configure an Amazon EventBridge rule to invoke an AWS Lambda function on launch and terminate actions to update InfraInventoryV2.
Aurora Global Database with managed cross-Region failover and Route 53 health checks
Aurora Global Database is built for low-lag cross-Region replication and managed failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora Global Database is built for low-lag cross-Region replication and managed failover.
● Scene fit: Aurora Global Database is built for low-lag cross-Region replication and managed failover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Can work, but promotion and DNS changes need custom orchestration.
● Multi-AZ works only inside one Region.
● Snapshot copy and restore are technically possible but slow.
Workflow: Aurora Global Database with managed cross-Region failover and Route 53 health checks.
AWS Trusted Advisor with a Business or Enterprise Support plan integrated with
This leverages the Trusted Advisor Low Utilization EC2 check with EventBridge for event-driven remediation filtered by tags.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This leverages the Trusted Advisor Low Utilization EC2 check with EventBridge for event-driven remediation filtered by tags.
● Scene fit: This leverages the Trusted Advisor Low Utilization EC2 check with EventBridge for event-driven remediation filtered by tags.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Option builds dashboards and event-driven actions based on tags and metrics for EC2 utilization.
● Relies on Compute Optimizer recommendations and attempts to automate shutdowns using EventBridge and Lambda.
● Proposes a custom data collection pipeline into DynamoDB and QuickSight with Lambda-based remediation.
Workflow: Use AWS Trusted Advisor with a Business or Enterprise Support plan integrated with Amazon EventBridge, and invoke AWS Lambda to filter on tags and automatically terminate consistently low-utilization EC2 instances.
EventBridge to run a Lambda that calls RefreshCache on the file share
A scheduled EventBridge rule can invoke Lambda to call RefreshCache after uploads finish and before 9 AM.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A scheduled EventBridge rule can invoke Lambda to call RefreshCache after uploads finish and before 9 AM.
● Scene fit: A scheduled EventBridge rule can invoke Lambda to call RefreshCache after uploads finish and before 9 AM.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Changing the upload service does not refresh the File Gateway cache.
● File Gateway does not provide a built-in automatic refresh setting for direct S3 writes.
● S3 cannot directly make Storage Gateway refresh.
Workflow: Use EventBridge to run a Lambda that calls RefreshCache on the file share before 9 AM.
Subscribe to AWS Health events and trigger an EventBridge rule that invokes
This orchestrates on-demand instance replacement to avoid standby costs and provides cross-AZ Aurora failover for high availability.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This orchestrates on-demand instance replacement to avoid standby costs and provides cross-AZ Aurora failover for high availability.
● Scene fit: This orchestrates on-demand instance replacement to avoid standby costs and provides cross-AZ Aurora failover for high availability.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Keeps a warm standby and leaves the database as a single point of failure while incurring continuous compute cost.
● EC2 automatic recovery and Config remediation do not address the EFA and instance store constraints and a single Aurora.
● Licensing prevents using Auto Scaling and this approach does not ensure compatibility with EFA and instance store requirements.
Workflow: Subscribe to AWS Health events and trigger an EventBridge rule that invokes a Lambda function to create a replacement EC2 instance in another Availability Zone upon failure, and configure an Aurora cluster with one cross-AZ Aurora Replica that can be promoted.
an Amazon EventBridge rule that matches AWS Health EC2 events and publish
EventBridge natively receives AWS Health events, allowing an immediate rule-to-SNS pattern that rapidly delivers EC2 maintenance and retirement notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge natively receives AWS Health events, allowing an immediate rule-to-SNS pattern that rapidly delivers EC2 maintenance and retirement notifications.
● Scene fit: EventBridge natively receives AWS Health events, allowing an immediate rule-to-SNS pattern that rapidly delivers EC2 maintenance and retirement.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Sends broad account communications, including billing and security notices, and is neither targeted to AWS Health events nor the safest approach.
● AWS Health events do not appear in CloudWatch Logs by default, so this would require extra ingestion steps and is not the fastest path.
● Although it can relay Health notifications, it requires deploying and configuring the solution, which is slower than a simple EventBridge-to-SNS rule.
Workflow: Create an Amazon EventBridge rule that matches AWS Health EC2 events and publish to an SNS topic subscribed by the analytics team.
The S3 website bucket still contains objects and versions; update the Lambda
CloudFormation can only delete empty S3 buckets, so the custom resource should handle the Delete request by recursively removing all objects, versions, and delete markers.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation can only delete empty S3 buckets, so the custom resource should handle the Delete request by recursively.
● Scene fit: CloudFormation can only delete empty S3 buckets, so the custom resource should handle the Delete request by recursively.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Presumes Object Lock is blocking deletion, which is not indicated and is not the common reason for CloudFormation S3 deletions to fail.
● Static website hosting settings do not prevent CloudFormation from deleting an S3 bucket if it is empty.
● ForceDelete is not a valid CloudFormation DeletionPolicy value and does not exist for S3 buckets.
Workflow: The S3 website bucket still contains objects and versions → update the Lambda custom resource to empty the bucket on the Delete event so CloudFormation can remove it.
on AWS Elastic Beanstalk for the Go platform, upload the application bundle
Elastic Beanstalk supports Go and offers managed blue/green via environment cloning and CNAME swap with minimal operational overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Elastic Beanstalk supports Go and offers managed blue/green via environment cloning and CNAME swap with minimal operational overhead.
● Scene fit: Elastic Beanstalk supports Go and offers managed blue/green via environment cloning and CNAME swap with minimal operational overhead.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Approach replaces instances in place and does not natively keep two parallel environments for controlled traffic shifting.
● App Runner simplifies deployments but does not provide environment-level blue/green with parallel stacks and traffic swap semantics.
● CodeArtifact is a package repository and not intended for hosting application bundles or orchestrating blue/green environments.
Workflow: Deploy on AWS Elastic Beanstalk for the Go platform, upload the application bundle to Amazon S3, and use Elastic Beanstalk environment swap for blue/green deployments.
AWS Secrets Manager with KMS via ECS container secrets, rotation enabled
Secrets Manager is a dedicated secrets service with KMS encryption, native automatic rotation, ECS integration for environment variables, and supports secret sizes up to 64 KB.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Secrets Manager is a dedicated secrets service with KMS encryption, native automatic rotation, ECS integration for environment variables, and supports secret sizes up to 64 KB.
● Scene fit: Secrets Manager is a dedicated secrets service with KMS encryption, native automatic rotation, ECS integration for environment variables.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Parameter Store lacks native automatic rotation and Advanced parameters are limited to about 8 KB, making it unsuitable for 20 KB secrets even with custom rotation.
● AppConfig is for application configuration and does not serve as a dedicated secrets store with native rotation or ECS secret injection.
● S3 is not a secrets store and provides no native automatic rotation or ECS secret injection mechanisms.
Workflow: AWS Secrets Manager with KMS via ECS container secrets, rotation enabled.
Jenkins as a multi-master installation across multiple AZs and use the AWS
This provides HA for Jenkins masters and elastic, pay-per-use build execution with CodeBuild.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides HA for Jenkins masters and elastic, pay-per-use build execution with CodeBuild.
● Scene fit: This provides HA for Jenkins masters and elastic, pay-per-use build execution with CodeBuild.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Keeps builds on EC2 agents but a single-AZ control plane is not fault tolerant.
● Although HA is achieved, maintaining an EC2 agent fleet adds cost and operational overhead compared to managed builds.
● Offloading builds helps elasticity, but placing masters in a single AZ undermines availability.
Workflow: Run Jenkins as a multi-master installation across multiple AZs and use the AWS CodeBuild plugin for Jenkins so builds execute in CodeBuild.
the Cognito Post Authentication trigger to invoke Lambda that emails via Amazon
The post-authentication trigger runs immediately after successful sign-in and can call Lambda to send email via SES with minimal overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The post-authentication trigger runs immediately after successful sign-in and can call Lambda to send email via SES with minimal overhead.
● Scene fit: The post-authentication trigger runs immediately after successful sign-in and can call Lambda to send email via SES with minimal overhead.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Cognito User Pool end-user sign-ins are not emitted as CloudTrail events, so this is unreliable and indirect.
● Adds client logic and extra services instead of using Cognito's native trigger, increasing complexity.
● Custom Message is for verification/MFA/forgot password flows, not for post-authentication sign-ins.
Workflow: Use the Cognito Post Authentication trigger to invoke Lambda that emails via Amazon SES.
Rebuild the custom AMI with EC2 Image Builder to include the current
This uses SSM Session Manager over VPC endpoints without internet, grants the right instance permissions, and provides auditable logs and notifications via S3 and SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses SSM Session Manager over VPC endpoints without internet, grants the right instance permissions, and provides auditable.
● Scene fit: This uses SSM Session Manager over VPC endpoints without internet, grants the right instance permissions, and provides auditable.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces internet egress and a bastion pattern that conflicts with the isolation requirement and adds unnecessary components.
● SCPs cannot grant permissions and AWS Config cannot attach SCPs, so this would not enable the required access path.
● EC2 Instance Connect relies on SSH network paths and does not natively provide centralized session logging and access controls.
Workflow: Rebuild the custom AMI with EC2 Image Builder to include the current SSM Agent → attach the AmazonSSMManagedInstanceCore instance profile to the Auto Scaling groups → use Systems Manager Session Manager for.
the CodePipeline stage to run actions for each Lambda function in parallel
Using the same runOrder for multiple actions in the same stage runs them concurrently, which most directly reduces total pipeline duration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using the same runOrder for multiple actions in the same stage runs them concurrently, which most directly reduces total pipeline duration.
● Scene fit: Using the same runOrder for multiple actions in the same stage runs them concurrently, which most directly reduces total pipeline duration.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Upsizing compute may shorten individual build times but does not remove the sequential bottleneck across actions.
● A build graph enforces dependency sequencing and does not parallelize independent CodePipeline actions across multiple functions.
● Placing builds in a VPC and using dedicated hosts does not inherently improve pipeline speed and can add network overhead.
Workflow: Configure the CodePipeline stage to run actions for each Lambda function in parallel by assigning the same runOrder.
a service control policy at the organization root that denies resource creation
An SCP can block create actions that do not include mandatory tags, enforcing compliance and preventing future tagless resources. Tag Editor allows bulk discovery and application.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An SCP can block create actions that do not include mandatory tags, enforcing compliance and preventing future tagless.
● Scene fit: An SCP can block create actions that do not include mandatory tags, enforcing compliance and preventing future tagless.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Turning on cost allocation tags helps reporting once tags exist but neither fixes untagged resources nor prevents tagless creation.
● Cost Categories can group costs but they do not apply or enforce resource tags, so they cannot remediate untagged.
Workflow: Attach a service control policy at the organization root that denies resource creation when required tags are missing → Use AWS Resource Groups Tag Editor in each account and Region to locate.
CloudFormation; CodeDeploy blue/green (Lambda); Elastic Beanstalk Immutable
Blue/green for Lambda provides traffic shifting and instant rollback, and Beanstalk Immutable creates a parallel fleet for zero downtime and safe cutover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Blue/green for Lambda provides traffic shifting and instant rollback, and Beanstalk Immutable creates a parallel fleet for zero downtime and safe cutover.
● Scene fit: Blue/green for Lambda provides traffic shifting and instant rollback, and Beanstalk Immutable creates a parallel fleet for zero downtime and safe cutover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● In-place plus rolling with an extra batch keeps capacity but lacks environment isolation and atomic rollback, risking downtime and failed-fast gating.
● All at once replaces instances simultaneously, causing a brief outage and not meeting the zero-downtime requirement.
● Canary helps Lambda, but Beanstalk Rolling updates in-place and can reduce capacity or increase blast radius during failures.
Workflow: CloudFormation → CodeDeploy blue/green (Lambda) → Elastic Beanstalk Immutable.
SSM Agent and Systems Manager Maintenance Windows + AWS-RunPatchBaseline SSM document
Installing the SSM Agent and scheduling patch tasks via Maintenance Windows provides automated, controlled, and auditable patching with minimal effort. Run via Patch Manager to apply approved baselines to Windows and Linux across the fleet with centralized auditing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Installing the SSM Agent and scheduling patch tasks via Maintenance Windows provides automated, controlled, and auditable patching with minimal effort.
● Scene fit: Installing the SSM Agent and scheduling patch tasks via Maintenance Windows provides automated, controlled, and auditable patching with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Is fragmented across hosts and distros, increasing operational overhead and lacking centralized orchestration.
● Inspector identifies vulnerabilities but does not apply patches; Patch Manager is needed for remediation.
● Targets Windows only and is not suitable for a mixed Windows/Linux environment.
Workflow: SSM Agent and Systems Manager Maintenance Windows → AWS-RunPatchBaseline SSM document.
The target EC2 instances are missing an IAM instance profile that grants
Without an instance profile that allows actions such as retrieving the revision and reporting status, the CodeDeploy agent cannot run hooks and events are marked Skipped.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Without an instance profile that allows actions such as retrieving the revision and reporting status, the CodeDeploy agent.
● Scene fit: Without an instance profile that allows actions such as retrieving the revision and reporting status, the CodeDeploy agent.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CodeDeploy uses a service role and the instance profile at runtime, so the initiating user’s permissions do not cause.
● CodeDeploy supports both tag-based and Auto Scaling group targeting, so using tags alone is not a cause of skipped.
Workflow: The target EC2 instances are missing an IAM instance profile that grants the CodeDeploy agent required access → The EC2 instances cannot reach CodeDeploy public endpoints because they have no egress path.
SSM Parameter Store with a CloudFormation SSM parameter type and schedule UpdateStack
CloudFormation resolves SSM parameter types at update, so a scheduled update rolls all stacks to the latest AMI without template edits each cycle.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation resolves SSM parameter types at update, so a scheduled update rolls all stacks to the latest AMI without template edits each cycle.
● Scene fit: CloudFormation resolves SSM parameter types at update, so a scheduled update rolls all stacks to the latest AMI without template edits each cycle.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Needs each stack’s specific parameter name, which breaks when templates differ.
● Dynamic references resolve only at create or update time, not automatically.
● Rewriting templates is brittle and unnecessary when CloudFormation can resolve SSM parameters directly.
Workflow: Use SSM Parameter Store with a CloudFormation SSM parameter type and schedule UpdateStack every 15 days.
Build a new table with an LSI that reuses the partition key
LSIs share the partition key, allow an alternate sort key, and support strongly consistent reads, but must be defined at table creation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: LSIs share the partition key, allow an alternate sort key, and support strongly consistent reads, but must be defined at table creation.
● Scene fit: LSIs share the partition key, allow an alternate sort key, and support strongly consistent reads, but must be defined at table creation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● GSIs do not support strongly consistent reads.
● You cannot add an LSI to an existing table; it must be created with the table.
● DAX provides caching with eventually consistent reads, not strong consistency.
Workflow: Build a new table with an LSI that reuses the partition key and adds a new sort key → migrate data.
Pre-provision the stack in a different region with CloudFormation, create a cross-region
A cross-region RDS Read Replica plus S3 CRR provides low RPO, and pre-provisioning with promotion of the replica minimizes RTO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A cross-region RDS Read Replica plus S3 CRR provides low RPO, and pre-provisioning with promotion of the replica.
● Scene fit: A cross-region RDS Read Replica plus S3 CRR provides low RPO, and pre-provisioning with promotion of the replica.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Multi-AZ standby instances cannot be placed in a different region, so this design does not deliver cross-region DR.
● Cross-region snapshot copies and Glacier retrievals increase RPO and RTO, making them unsuitable for the lowest data loss and.
● ALB is regional and cannot route to another region, and Multi-AZ alone does not protect against regional failure.
Workflow: Pre-provision the stack in a different region with CloudFormation → create a cross-region RDS Read Replica → enable S3 cross-region replication to a destination bucket, and promote the replica during failover while pre-scaling the Auto Scaling group.
AWS Storage Gateway file gateway on-premises, have MAM read and write media
File Gateway provides SMB or NFS backed by S3, enabling Rekognition on S3 objects with minimal changes and low operational overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: File Gateway provides SMB or NFS backed by S3, enabling Rekognition on S3 objects with minimal changes and low operational overhead.
● Scene fit: File Gateway provides SMB or NFS backed by S3, enabling Rekognition on S3 objects with minimal changes and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces a streaming pipeline that increases complexity and does not naturally fit cold file-based archives on tape.
● Requires building and operating custom infrastructure and software, which increases ongoing management effort.
● Virtual tapes are archived to S3 Glacier storage classes and are not directly accessible for real-time Rekognition processing.
Workflow: Deploy AWS Storage Gateway file gateway on-premises, have MAM read and write media through the gateway so files land in Amazon S3, and use AWS Lambda to invoke Amazon Rekognition to index faces from S3 and update the MAM catalog.
Build the workflow with AWS Step Functions, use Amazon API Gateway with
Direct API Gateway to Step Functions supports long-running orchestrations and keeps the design simple while scaling and authenticating with Cognito.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Direct API Gateway to Step Functions supports long-running orchestrations and keeps the design simple while scaling and authenticating with Cognito.
● Scene fit: Direct API Gateway to Step Functions supports long-running orchestrations and keeps the design simple while scaling and authenticating.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Fails because Lambda cannot run for hours and will time out after 15 minutes even if called via API Gateway.
● Adding Lambda as a hop increases complexity and potential throttling points without any benefit over the native API Gateway to Step Functions integration.
● Is unsuitable for public consumer APIs and still violates Lambda’s 15-minute maximum execution time for multi-hour workflows.
Workflow: Build the workflow with AWS Step Functions → use Amazon API Gateway with a direct service integration to Step Functions, and secure the API with Amazon Cognito.
CloudWatch alarm on EC2 Max CPU with CodeDeploy automatic rollback enabled
Attach a CloudWatch alarm on EC2 CPUUtilization (Maximum) to the CodeDeploy deployment group and enable automatic rollback so a breach during shifting triggers rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Attach a CloudWatch alarm on EC2 CPUUtilization (Maximum) to the CodeDeploy deployment group and enable automatic rollback so a breach during shifting triggers rollback.
● Scene fit: Attach a CloudWatch alarm on EC2 CPUUtilization (Maximum) to the CodeDeploy deployment group and enable automatic rollback so a breach during shifting triggers rollback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● ALB error alarms can trigger rollbacks, but this does not address rollbacks based on high EC2 CPU as required.
● Scaling policies and health checks do not trigger CodeDeploy rollbacks based on CPU thresholds.
● Lifecycle hooks are not intended for live traffic rollback conditions; use CodeDeploy alarms for automatic rollback during shifting.
Workflow: CloudWatch alarm on EC2 Max CPU with CodeDeploy automatic rollback enabled.
Lambda alias with 20% weighted traffic + API Gateway stage canary at
Lambda aliases support routing configuration to shift a percentage of traffic between two versions and allow instant rollback. API Gateway canary settings let you send a percentage of traffic to a canary and quickly promote or revert.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda aliases support routing configuration to shift a percentage of traffic between two versions and allow instant rollback.
● Scene fit: Lambda aliases support routing configuration to shift a percentage of traffic between two versions and allow instant rollback.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Failover routing switches between primary and secondary based on health; it cannot do percentage splits.
● NLB does not provide weighted routing across target groups for precise canary splits via API Gateway.
● AppConfig manages feature flags but does not route API traffic at the gateway or enforce request-level percentage splits.
Workflow: Lambda alias with 20% weighted traffic → API Gateway stage canary at 20%.
a CloudWatch Logs subscription to Kinesis Data Firehose delivering to S3 +
CloudWatch Logs can stream to Kinesis Data Firehose, which reliably delivers to S3. Lifecycle transitions minimize cost after 90 days and expiration enforces 10-year retention.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudWatch Logs can stream to Kinesis Data Firehose, which reliably delivers to S3.
● Scene fit: CloudWatch Logs can stream to Kinesis Data Firehose, which reliably delivers to S3. Lifecycle transitions minimize cost after 90 days and expiration enforces 10-year retention.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DataSync is not a supported destination for CloudWatch Logs subscriptions.
● Exports are batch jobs requiring orchestration and are not continuous streaming.
● Retention only controls log deletion in CloudWatch Logs and does not archive to S3.
Workflow: Set a CloudWatch Logs subscription to Kinesis Data Firehose delivering to S3 → Add an S3 lifecycle rule: transition to S3 Glacier Deep Archive at 90 days, expire at 3,650 days.
EventBridge rule for CloudTrail iam:CreateUser events + Lambda to delete login profile
Match AWS API Call via CloudTrail with eventSource iam.amazonaws.com and eventName CreateUser to detect new IAM users quickly. Use DeleteLoginProfile and UpdateAccessKey to neutralize console and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Match AWS API Call via CloudTrail with eventSource iam.amazonaws.com and eventName CreateUser to detect new IAM users quickly.
● Scene fit: Match AWS API Call via CloudTrail with eventSource iam.amazonaws.com and eventName CreateUser to detect new IAM users quickly.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Config is for configuration compliance and may not provide immediate, event-driven remediation upon user creation.
● GetLoginProfile covers console password profiles only and misses users created without a console password.
● Prevents user creation rather than disabling credentials post-creation and provides no alerting.
Workflow: EventBridge rule for CloudTrail iam:CreateUser events → Lambda to delete login profile and deactivate access keys for new users → SNS topic target from EventBridge subscribed by security.
Read with one Lambda and fan out events via Amazon SNS
A single consumer avoids shard contention and SNS provides scalable fan-out to many processors.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A single consumer avoids shard contention and SNS provides scalable fan-out to many processors.
● Scene fit: A single consumer avoids shard contention and SNS provides scalable fan-out to many processors.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Increases concurrency per shard but does not remove shard reader limits and can worsen contention across multiple consumers.
● DAX speeds table reads and does not change how DynamoDB Streams are consumed.
● DynamoDB Streams is not a native EventBridge rule source and multiple rules would still compete for the same shards.
Workflow: Read with one Lambda and fan out events via Amazon SNS.
AWS Config managed rules for S3 with automatic remediation using Systems Manager
AWS Config continuously evaluates S3 buckets against managed rules and can invoke Systems Manager Automation to remediate noncompliant buckets automatically.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config continuously evaluates S3 buckets against managed rules and can invoke Systems Manager Automation to remediate noncompliant buckets automatically.
● Scene fit: AWS Config continuously evaluates S3 buckets against managed rules and can invoke Systems Manager Automation to remediate noncompliant buckets automatically.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Can react to API calls but is complex to maintain, event-focused rather than configuration-compliance focused, and does not provide native compliance evaluations across all resources.
● SCPs and Block Public Access can prevent certain actions but cannot automatically enable encryption, logging, and versioning or remediate existing buckets.
● Trusted Advisor does not provide the granular per-bucket compliance checks and native auto-remediation needed for this requirement.
Workflow: Configure AWS Config managed rules for S3 with automatic remediation using Systems Manager Automation runbooks.
AutomationAssumeRole to an IAM role for Systems Manager Automation and grant iam:PassRole
SSM Automation must assume a role to execute the remediation, and the initiator needs iam:PassRole permissions. This managed SSM Automation runbook turns on server access logging for noncompliant buckets.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SSM Automation must assume a role to execute the remediation, and the initiator needs iam:PassRole permissions.
● Scene fit: SSM Automation must assume a role to execute the remediation, and the initiator needs iam:PassRole permissions. This managed SSM Automation runbook turns on server access logging for noncompliant buckets.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Security Hub aggregates and prioritizes findings but does not change resource configurations.
● Possible but unnecessary because a managed remediation runbook already exists and is preferred.
● SCPs can restrict or allow APIs but cannot automatically enable logging on existing buckets.
Workflow: Set AutomationAssumeRole to an IAM role for Systems Manager Automation and grant iam:PassRole → Configure AWS Config auto-remediation for s3-bucket-logging-enabled with AWS-ConfigureS3BucketLogging.
an AWS Site-to-Site VPN between the data center and the VPC +
Site-to-Site VPN uses IPsec to encrypt traffic between the on-premises network and the VPC, satisfying the encryption requirement. Fargate removes the need to manage servers or.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Site-to-Site VPN uses IPsec to encrypt traffic between the on-premises network and the VPC, satisfying the encryption requirement.
● Scene fit: Site-to-Site VPN uses IPsec to encrypt traffic between the on-premises network and the VPC, satisfying the encryption requirement.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Increases operational burden because you must patch, scale, and manage the EC2 hosts yourself.
● An NLB does not provide an HTTPS listener and adding a load balancer does not fulfill the cross-network encryption requirement.
● Direct Connect does not encrypt traffic by default, so it does not meet the requirement for encrypted connectivity on its own.
Workflow: Configure an AWS Site-to-Site VPN between the data center and the VPC → Run the workload on Amazon ECS with the Fargate launch type in a dedicated VPC.
an EventBridge rule for AWS Glue job run events with a Lambda
EventBridge routes Glue job events to Lambda where code can check attempt and state details and publish to SNS only for last-attempt failures.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge routes Glue job events to Lambda where code can check attempt and state details and publish to.
● Scene fit: EventBridge routes Glue job events to Lambda where code can check attempt and state details and publish to.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Sends events directly to SNS based on a pattern, but EventBridge event patterns cannot evaluate attempt counts or conditional.
● The Personal Health Dashboard surfaces account and service health events, not per-job Glue retry outcomes, so it will not.
● Step Functions could handle this but requires redesigning orchestration and is unnecessary when EventBridge plus Lambda can filter Glue.
Workflow: Configure an EventBridge rule for AWS Glue job run events with a Lambda target that inspects the event details for a failed final attempt → publishes to SNS.
EventBridge (AWS Health) triggering Step Functions to orchestrate IAM, CloudTrail, and SNS
This uses the AWS Health event through EventBridge and Step Functions for reliable, auditable orchestration of alerts, activity collection, and key disablement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses the AWS Health event through EventBridge and Step Functions for reliable, auditable orchestration of alerts, activity collection, and key disablement.
● Scene fit: This uses the AWS Health event through EventBridge and Step Functions for reliable, auditable orchestration of alerts, activity collection, and key disablement.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A single Lambda can perform actions but lacks robust orchestration, retries, and a visual execution history for audit compared to Step Functions.
● AWS_RISK_CREDENTIALS_EXPOSED originates from AWS Health, not CloudTrail, so the rule would never match.
● GuardDuty does not emit the AWS Health leaked-credentials event and this path does not provide the required end-to-end orchestration or activity report.
Workflow: EventBridge (AWS Health) triggering Step Functions to orchestrate IAM, CloudTrail, and SNS.
Drift detection does not evaluate custom resources + The rule calls DetectStackDrift
CloudFormation drift detection excludes custom resources, which are not checked for drift. The managed rule relies on DetectStackDrift and defaults to NON_COMPLIANT on API errors or throttling, explaining an IN_SYNC console state mismatch.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation drift detection excludes custom resources, which are not checked for drift.
● Scene fit: CloudFormation drift detection excludes custom resources, which are not checked for drift. The managed rule relies on DetectStackDrift and defaults to NON_COMPLIANT on API errors or throttling, explaining an IN_SYNC console state mismatch.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Change sets preview proposed changes and do not detect out-of-band drift.
● Permission issues generally produce explicit access errors rather than defaulting the managed rule to NON_COMPLIANT.
● Custom resources are unsupported for drift detection regardless of how many properties are specified.
Workflow: Drift detection does not evaluate custom resources → The rule calls DetectStackDrift and treats API throttling or failures as NON_COMPLIANT.
Publish 90-second aggregated counts as CloudWatch custom metrics; use a CloudWatch alarm
Sending aggregated counts as a custom metric enables simple, low-cost alarms and instant SNS notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Sending aggregated counts as a custom metric enables simple, low-cost alarms and instant SNS notifications.
● Scene fit: Sending aggregated counts as a custom metric enables simple, low-cost alarms and instant SNS notifications.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch alarms evaluate metrics, not EventBridge events, so this cannot directly drive alarms without metrics.
● Introduces extra components and cost compared to native CloudWatch custom metrics and alarms.
● Feasible but adds log ingestion, filter management, and additional cost versus direct custom metrics.
Workflow: Publish 90-second aggregated counts as CloudWatch custom metrics → use a CloudWatch alarm with SNS.
Provision a new encrypted EBS volume, attach it to the instance, migrate
EBS encryption must be applied when a volume is created, so creating a new encrypted volume and moving the data achieves encryption at rest. Copying and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EBS encryption must be applied when a volume is created, so creating a new encrypted volume and moving.
● Scene fit: EBS encryption must be applied when a volume is created, so creating a new encrypted volume and moving.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Default EBS encryption affects only new volumes and snapshots, not existing resources.
● You cannot directly encrypt an existing snapshot or retroactively change an attached volume to encrypted.
● TLS protects data in transit and does not provide encryption at rest on the volume.
Workflow: Provision a new encrypted EBS volume → attach it to the instance, migrate the data, and retire the old unencrypted volume → Copy an unencrypted snapshot of the volume → encrypt the copied snapshot, and create a new volume from it.
Remove public access with a bucket policy and grant the CodeBuild service
This enforces least privilege and uses the CodeBuild service role for temporary credentials, eliminating public access and avoiding static keys.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This enforces least privilege and uses the CodeBuild service role for temporary credentials, eliminating public access and avoiding static keys.
● Scene fit: This enforces least privilege and uses the CodeBuild service role for temporary credentials, eliminating public access and avoiding static keys.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Amazon S3 does not support basic authentication, and this method deviates from AWS-recommended IAM and bucket policy controls.
● Embedding static credentials in builds increases exposure risk and is less secure than using a role with temporary credentials.
● Over-privileges the role and fails to close the public access gap flagged by the audit.
Workflow: Remove public access with a bucket policy and grant the CodeBuild service role least-privilege S3 permissions, → use the AWS CLI in the build to retrieve the object.
Stand up a parallel stack with an Application Load Balancer and Auto
Weighted alias records in Route 53 let you start with a small percentage to the new ALB and increase gradually, enabling a true canary for EC2/ASG.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Weighted alias records in Route 53 let you start with a small percentage to the new ALB and.
● Scene fit: Weighted alias records in Route 53 let you start with a small percentage to the new ALB and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● API Gateway private integration supports Network Load Balancers, not Application Load Balancers, and is oriented to APIs rather than.
● Latency-based routing selects endpoints by measured latency and does not provide explicit percentage controls required for a canary.
Workflow: Stand up a parallel stack with an Application Load Balancer and Auto Scaling group running the new version, → use Route 53 weighted alias records to split and progressively shift traffic between.
Amazon Inspector for EC2 plus CloudWatch Agent shipping login logs to CloudWatch
Inspector continuously scans EC2 for CVEs and exposure; CloudWatch Agent centralizes OS login logs in CloudWatch Logs where retention can be set to 30 days, and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Inspector continuously scans EC2 for CVEs and exposure; CloudWatch Agent centralizes OS login logs in CloudWatch Logs where retention can be set to 30 days, and CloudTrail complements with account activity auditing.
● Scene fit: Inspector continuously scans EC2 for CVEs and exposure; CloudWatch Agent centralizes OS login logs in CloudWatch Logs where.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Security Hub aggregates and prioritizes findings from other services but does not perform host vulnerability scanning or collect OS login logs.
● GuardDuty detects threats and anomalies but does not perform EC2 host vulnerability scans or capture OS login events.
● ECR scans container images, not EC2 operating systems or their login activity.
Workflow: Amazon Inspector for EC2 → CloudWatch Agent shipping login logs to CloudWatch Logs, and CloudTrail to the same logs.
a Lambda alias with weighted traffic to versions
Lambda alias routing supports weighted traffic between versions, enabling canary or linear rollouts behind a single API Gateway stage.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda alias routing supports weighted traffic between versions, enabling canary or linear rollouts behind a single API Gateway stage.
● Scene fit: Lambda alias routing supports weighted traffic between versions, enabling canary or linear rollouts behind a single API Gateway stage.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● API Gateway canaries work at the stage level and typically require a separate stage, not version-level traffic shifting within one stage.
● Splits DNS across endpoints and requires multiple API deployments; it does not split traffic between Lambda versions behind one stage.
● AppConfig targets configuration and feature flags, not code version traffic shifting for Lambda invocations.
Workflow: Configure a Lambda alias with weighted traffic to versions.
.ebextensions container_commands with leader_only: true
Container commands support leader_only and run a single time on the leader before deploying the app.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Container commands support leader_only and run a single time on the leader before deploying the app.
● Scene fit: Container commands support leader_only and run a single time on the leader before deploying the app.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs commands across instances and lacks Elastic Beanstalk leader election, risking parallel execution.
● The commands block doesn't support leader_only, so it would run on all instances.
● Predeploy hooks run on each instance unless you implement your own locking; not a built-in single-run mechanism.
Workflow: .ebextensions container_commands with leader_only: true.
Implement a pilot light in a secondary Region with a cross-Region RDS
This meets a minutes-level RPO via ongoing replication and minimizes steady-state cost by keeping most compute off until failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This meets a minutes-level RPO via ongoing replication and minimizes steady-state cost by keeping most compute off until failover.
● Scene fit: This meets a minutes-level RPO via ongoing replication and minimizes steady-state cost by keeping most compute off until.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Lowers recovery time but keeps an always-on environment that is costlier than necessary for hours-level RTO targets.
● Achieves excellent RTO/RPO but is the most expensive and unnecessary given the cost constraint and hours-level RTO.
● Often cannot meet a 20-minute RPO and may exceed a four-hour RTO due to snapshot copy and restore times.
Workflow: Implement a pilot light in a secondary Region with a cross-Region RDS PostgreSQL read replica, a minimal application stack ready to start → Route 53 health checks for failover, and promote the replica to primary during disaster.
ACM cert on ALB; ACM cert in us-east-1 on CloudFront custom domain
CloudFront requires the custom domain cert in us-east-1, enforce viewer HTTPS via Viewer Protocol Policy, and origin HTTPS via Origin Protocol Policy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFront requires the custom domain cert in us-east-1, enforce viewer HTTPS via Viewer Protocol Policy, and origin HTTPS via Origin Protocol Policy.
● Scene fit: CloudFront requires the custom domain cert in us-east-1, enforce viewer HTTPS via Viewer Protocol Policy, and origin HTTPS via Origin Protocol Policy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Self-signed ALB certs are untrusted by CloudFront and break origin HTTPS validation.
● Match Viewer permits HTTP from viewers, so HTTPS is not enforced end-to-end.
● CloudFront custom domain certificates must be in us-east-1, not the ALB's region.
Workflow: ACM cert on ALB → ACM cert in us-east-1 on CloudFront custom domain → Viewer HTTPS Only → origin HTTPS.
an EC2 Auto Scaling lifecycle hook that moves instances entering Terminating into
A Terminating:Wait lifecycle hook pauses termination so you can connect to the instance and debug before it is finally terminated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A Terminating:Wait lifecycle hook pauses termination so you can connect to the instance and debug before it is finally terminated.
● Scene fit: A Terminating:Wait lifecycle hook pauses termination so you can connect to the instance and debug before it is finally terminated.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Pausing AZRebalance only halts rebalancing across Availability Zones and does not prevent health-check-based termination events.
● Instance scale-in protection blocks scale-in but does not stop termination driven by failed health checks or instance replacement logic.
● A snapshot or AMI does not capture in-memory state and may be incomplete while the instance is terminating, which limits live debugging.
Workflow: Add an EC2 Auto Scaling lifecycle hook that moves instances entering Terminating into Terminating:Wait to allow troubleshooting access.
CloudWatch Logs subscription to Lambda tags instance; EventBridge every 30 minutes triggers
End-to-end automation that reacts to log events and periodically terminates tagged instances without human intervention.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: End-to-end automation that reacts to log events and periodically terminates tagged instances without human intervention.
● Scene fit: End-to-end automation that reacts to log events and periodically terminates tagged instances without human intervention.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EC2 alarm actions require per-instance alarms tied to instance metrics and are not practical or reliable with Logs-derived metrics for dynamic instances.
● Involves manual steps after notification, so it is not fully automated.
● CloudWatch Logs subscriptions cannot directly deliver to Step Functions, so this wiring is invalid and introduces unnecessary delay.
Workflow: CloudWatch Logs subscription to Lambda tags instance → EventBridge every 30 minutes triggers Lambda to terminate tagged instances.
an AWS CloudTrail trail and add an Amazon EventBridge rule for AWS
This uses CloudTrail to record the API call and EventBridge to deliver near real-time notifications to SNS at low cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses CloudTrail to record the API call and EventBridge to deliver near real-time notifications to SNS at low cost.
● Scene fit: This uses CloudTrail to record the API call and EventBridge to deliver near real-time notifications to SNS at.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail event selectors only control what is recorded and do not natively trigger functions, so this adds complexity and latency if you must parse logs after delivery.
● DynamoDB Streams capture item-level changes, not administrative actions like DeleteTable, so they will not reliably notify on table deletions.
● AWS Config evaluations are not intended for near real-time API auditing and introduce extra Lambda and evaluation costs.
Workflow: Create an AWS CloudTrail trail and add an Amazon EventBridge rule for AWS API Call via CloudTrail that matches DeleteTable and targets Amazon SNS.
Terminating:Wait lifecycle hook with EventBridge invoking SSM Automation to flush logs, then
A Terminating:Wait hook pauses scale-in, triggers automation to flush the CloudWatch agent, and completes the hook so logs are captured before termination.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A Terminating:Wait hook pauses scale-in, triggers automation to flush the CloudWatch agent, and completes the hook so logs are captured before termination.
● Scene fit: A Terminating:Wait hook pauses scale-in, triggers automation to flush the CloudWatch agent, and completes the hook so logs are captured before termination.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Pending:Wait applies during instance launch, not during scale-in or termination, so it will not delay termination to collect logs.
● Continuous streaming helps, but abrupt termination can prevent a final flush and still lose logs.
● Subscription filters move logs already in CloudWatch; they do not capture logs lost before the agent uploads them.
Workflow: Terminating:Wait lifecycle hook with EventBridge invoking SSM Automation to flush logs, → complete.
Invoke ContinueUpdateRollback on the stack in AWS CloudFormation + Manually remediate resources
ContinueUpdateRollback instructs CloudFormation to proceed with the rollback after blocking issues are addressed. Fixing out-of-band changes or missing dependencies enables the rollback to complete when you continue it.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ContinueUpdateRollback instructs CloudFormation to proceed with the rollback after blocking issues are addressed.
● Scene fit: ContinueUpdateRollback instructs CloudFormation to proceed with the rollback after blocking issues are addressed. Fixing out-of-band changes or missing dependencies enables the rollback to complete when you continue it.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Reusing the previous template does not advance a rollback that has failed and will not resolve the UPDATE_ROLLBACK_FAILED state.
● StackSets manage stacks across accounts and Regions and are not a mechanism to repair a single stack rollback failure.
● Drift detection reports configuration differences but does not clear an UPDATE_ROLLBACK_FAILED condition.
Workflow: Invoke ContinueUpdateRollback on the stack in AWS CloudFormation → Manually remediate resources so they match the stack's last good state.
Store secrets in AWS Secrets Manager encrypted with KMS, grant access via
Secrets Manager is a dedicated secrets service that supports automatic rotation, ECS integration via container secrets, KMS encryption, and larger secret sizes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Secrets Manager is a dedicated secrets service that supports automatic rotation, ECS integration via container secrets, KMS encryption.
● Scene fit: Secrets Manager is a dedicated secrets service that supports automatic rotation, ECS integration via container secrets, KMS encryption.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 with SSE-KMS and lifecycle rules does not provide native secret rotation or ECS secret injection and is not a dedicated secrets store.
● Parameter Store lacks built-in automatic secret rotation and has size limits that do not meet larger secret values.
● KMS manages encryption keys rather than storing and rotating application secrets for retrieval by ECS tasks.
Workflow: Store secrets in AWS Secrets Manager encrypted with KMS → grant access via the ECS task execution role, and reference secret ARNs in container definitions for environment variables with automatic rotation enabled.
DAX for DynamoDB and front S3 with CloudFront
DAX provides API-compatible, in-memory caching for DynamoDB reads and CloudFront caches S3 at the edge to reduce latency and cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: DAX provides API-compatible, in-memory caching for DynamoDB reads and CloudFront caches S3 at the edge to reduce latency and cost.
● Scene fit: DAX provides API-compatible, in-memory caching for DynamoDB reads and CloudFront caches S3 at the edge to reduce latency and cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Redis can cache reads but is not API-compatible with DynamoDB and adds operational overhead; CloudFront helps for S3.
● API Gateway caching requires the API to be on API Gateway and Transfer Acceleration speeds uploads, not edge caching for reads.
● CloudFront over ALB does not cache DynamoDB results and on-demand capacity only adjusts throughput, not read caching.
Workflow: Enable DAX for DynamoDB and front S3 with CloudFront.
CloudTrail log file integrity validation
Uses cryptographic digests and signatures to verify authenticity and preserve the ordered sequence of log files.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Uses cryptographic digests and signatures to verify authenticity and preserve the ordered sequence of log files.
● Scene fit: Uses cryptographic digests and signatures to verify authenticity and preserve the ordered sequence of log files.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides WORM immutability but does not prove CloudTrail log provenance or event sequencing.
● Encrypts and versions objects but does not provide cryptographic chain-of-custody or ordering.
● Enforces retention policies but does not validate who produced the logs or their sequence.
Workflow: CloudTrail log file integrity validation.
AutoScalingRollingUpdate
This UpdatePolicy performs rolling updates for Auto Scaling groups and supports MinInstancesInService and MaxBatchSize to preserve capacity during instance replacement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This UpdatePolicy performs rolling updates for Auto Scaling groups and supports MinInstancesInService and MaxBatchSize to preserve capacity during instance replacement.
● Scene fit: This UpdatePolicy performs rolling updates for Auto Scaling groups and supports MinInstancesInService and MaxBatchSize to preserve capacity during instance replacement.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● UpdatePolicy can replace the entire Auto Scaling group or all instances and is not suited for maintaining in-service capacity with a controlled batch rollout.
● Is a deployment service and not a CloudFormation UpdatePolicy used to orchestrate rolling updates within an Auto Scaling group resource in a template.
● Is not a valid CloudFormation UpdatePolicy, so it cannot be used to manage Auto Scaling group updates in a template.
Workflow: AutoScalingRollingUpdate.
Increase the stream's data retention to 72 hours + Run the KCL
Extending retention provides more time for backlogged consumers to catch up before records expire. Scaling KCL workers based on lag directly increases processing capacity only when.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Extending retention provides more time for backlogged consumers to catch up before records expire.
● Scene fit: Extending retention provides more time for backlogged consumers to catch up before records expire. Scaling KCL workers based.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adding shards increases producer throughput and potential parallelism but does not resolve a single under-provisioned consumer host and adds.
● Moving to Lambda requires architectural changes and does not inherently fix the root cause of consumer lag for this.
● Switching to SQS is a major redesign that loses shard semantics and is not a minimal-change reliability fix for.
Workflow: Increase the stream's data retention to 72 hours → Run the KCL app in an Auto Scaling group of EC2 instances and scale on the MillisBehindLatest CloudWatch metric.
AWS Storage Gateway File Gateway with S3 shares and Lambda-triggered Rekognition
Provides NFS/SMB backed by S3 for minimal workflow change and can invoke Rekognition via Lambda on new S3 objects.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides NFS/SMB backed by S3 for minimal workflow change and can invoke Rekognition via Lambda on new S3 objects.
● Scene fit: Provides NFS/SMB backed by S3 for minimal workflow change and can invoke Rekognition via Lambda on new S3 objects.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Designed for live or near-real-time streams, not bulk file archives, and adds unnecessary pipeline complexity.
● Rekognition cannot analyze objects directly in Glacier classes; restores and rehydration add overhead.
● Requires building and operating custom compute and software, increasing ongoing management effort.
Workflow: AWS Storage Gateway File Gateway with S3 shares and Lambda-triggered Rekognition.
an AWS SAM template for the serverless app and deploy with AWS
This performs a small canary shift with built-in alarms and automated rollback to minimize impact from defects.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This performs a small canary shift with built-in alarms and automated rollback to minimize impact from defects.
● Scene fit: This performs a small canary shift with built-in alarms and automated rollback to minimize impact from defects.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Enables weighted traffic shifting but lacks integrated health checks and automated rollback found in dedicated deployment strategies.
● AllAtOnce immediately shifts 100% of traffic, maximizing blast radius during failures.
● AppConfig manages configuration and feature flags but does not perform Lambda version traffic shifting or code rollbacks.
Workflow: Use an AWS SAM template for the serverless app and deploy with AWS CodeDeploy using the Canary5Percent15Minutes deployment type.
AWS Step Functions to start an EC2 instance from each baseline AMI
This orchestrates instance launches, installs the required agent, scopes targets via tags, and uses EventBridge for reliable scheduling.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This orchestrates instance launches, installs the required agent, scopes targets via tags, and uses EventBridge for reliable scheduling.
● Scene fit: This orchestrates instance launches, installs the required agent, scopes targets via tags, and uses EventBridge for reliable scheduling.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Uses CloudWatch Alarms for scheduling and omits ensuring the Inspector agent is present, making it unsuitable.
● Lacks tag-based scoping and misuses the event bus instead of a scheduled EventBridge rule.
● Patch Manager evaluates patch compliance of managed instances rather than performing Amazon Inspector CVE assessments of AMI-derived instances.
Workflow: Use AWS Step Functions to start an EC2 instance from each baseline AMI for Linux and Windows → install the Amazon Inspector agent → apply a tracking tag → invoke an Inspector.
post_build commands to push the image to ECR
post_build runs only if prior phases succeed, so the push occurs only on successful builds.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: post_build runs only if prior phases succeed, so the push occurs only on successful builds.
● Scene fit: post_build runs only if prior phases succeed, so the push occurs only on successful builds.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● The finally sequence runs regardless of success or failure, which can push even when the build fails.
● Pipeline orchestration does not change buildspec phase behavior; the success-only push should be controlled in buildspec.yml.
● Pre_build runs before the build completes, risking a push even if the build later fails.
Workflow: Use post_build commands to push the image to ECR.
Single ALB with two target groups; deploy to idle group and flip
Removes DNS from the cutover path by switching traffic at the ALB layer, simplifying and reducing cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Removes DNS from the cutover path by switching traffic at the ALB layer, simplifying and reducing cost.
● Scene fit: Removes DNS from the cutover path by switching traffic at the ALB layer, simplifying and reducing cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Still relies on DNS, so clients that pin or ignore updates may continue hitting the old ALB.
● Avoids DNS propagation issues but adds cost and complexity compared to a single ALB approach.
● Shorter TTL does not help clients that cache or pin DNS responses beyond TTL.
Workflow: Single ALB with two target groups → deploy to idle group and flip the listener to the new group.
CloudTrail log file integrity validation on the trails so CloudTrail delivers signed
CloudTrail's log file integrity validation is the native, least-effort way to create signed digest files that allow tamper-evident verification of log files.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail's log file integrity validation is the native, least-effort way to create signed digest files that allow tamper-evident.
● Scene fit: CloudTrail's log file integrity validation is the native, least-effort way to create signed digest files that allow tamper-evident.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 does not provide CloudTrail log integrity validation because the feature is configured in CloudTrail, and changing IAM permissions does not enable validation.
● State Manager cannot directly toggle CloudTrail log file integrity validation because that setting is enabled within CloudTrail itself or via its API.
● AWS Audit Manager helps manage compliance evidence but does not provide cryptographic integrity validation of CloudTrail log files.
Workflow: Enable CloudTrail log file integrity validation on the trails so CloudTrail delivers signed digest files you can use to validate the delivered log files.
AWS CodeBuild with KMS-encrypted artifacts
Fully managed builds with AWS KMS artifact encryption, reducing infrastructure management.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Fully managed builds with AWS KMS artifact encryption, reducing infrastructure management.
● Scene fit: Fully managed builds with AWS KMS artifact encryption, reducing infrastructure management.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Encrypts objects at rest but keeps self-managed Jenkins and associated operational overhead.
● Protects instances and block storage but does not ensure artifact encryption and still requires instance management.
● Moves Jenkins to containers but remains self-managed; SSE-S3 encrypts but operational burden persists.
Workflow: Use AWS CodeBuild with KMS-encrypted artifacts.
a role in the destination account that the source pipeline role can
Cross-account actions require an IAM role in the target account trusted by the source account's CodePipeline service role. CodePipeline artifact buckets must be versioned, and cross-account.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-account actions require an IAM role in the target account trusted by the source account's CodePipeline service role.
● Scene fit: Cross-account actions require an IAM role in the target account trusted by the source account's CodePipeline service role.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 buckets cannot be shared via AWS RAM; access is managed via IAM roles and bucket policies.
● Multi-Region keys are unnecessary unless spanning Regions; cross-account alone does not require them.
● The assumable role must reside in the target account; placing it in the source account reverses the trust direction.
Workflow: Configure a role in the destination account that the source pipeline role can assume via sts:AssumeRole → Enable versioning on the artifact bucket and use a customer managed KMS key for artifacts.
AWS Config approved-amis-by-id with SSM Automation remediation and SNS + EventBridge schedule
Continuously evaluates EC2 AMI compliance and can auto-remediate by stopping or terminating and notifying without blocking launches. A detective, scheduled check that compares AMIs and remediates keeps pipelines unblocked and minimizes disruption.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously evaluates EC2 AMI compliance and can auto-remediate by stopping or terminating and notifying without blocking launches.
● Scene fit: Continuously evaluates EC2 AMI compliance and can auto-remediate by stopping or terminating and notifying without blocking launches. A.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A preventative permission policy would block unapproved AMIs and disrupt staging and CI workflows.
● An org-wide preventative control that would stop experimental AMIs and interfere with developer workflows.
● Trusted Advisor does not provide AMI allowlist checks or automatic remediation, so it cannot enforce this requirement.
Workflow: AWS Config approved-amis-by-id with SSM Automation remediation and SNS → EventBridge schedule 15 min invoking Lambda to kill noncompliant EC2 and notify.
an AWS Config rule with a configuration change trigger scoped to the
A change-triggered AWS Config rule evaluates in near real time and can invoke SSM Automation via Lambda to automatically remediate drift.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A change-triggered AWS Config rule evaluates in near real time and can invoke SSM Automation via Lambda to.
● Scene fit: A change-triggered AWS Config rule evaluates in near real time and can invoke SSM Automation via Lambda to.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A scheduled scan introduces up to a 20-minute delay and requires custom parsing of policies, which is slower than.
● A periodic rule runs on a fixed interval and will not detect changes in near real time.
● While possible, this approach requires complex event patterns and lacks resource state evaluation, making it less reliable and not.
Workflow: Configure an AWS Config rule with a configuration change trigger scoped to the S3 bucket and federated role, and invoke an AWS Systems Manager Automation runbook via Lambda to restore approved settings.
CodeBuild to execute the tests, insert a Manual approval action in CodePipeline
This uses CodePipeline's built-in Manual approval action with optional SNS alerts, which is the simplest and lowest-cost way to enforce a human gate before production.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses CodePipeline's built-in Manual approval action with optional SNS alerts, which is the simplest and lowest-cost way.
● Scene fit: This uses CodePipeline's built-in Manual approval action with optional SNS alerts, which is the simplest and lowest-cost way.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Using Step Functions for tests adds extra complexity and cost compared to native CodeBuild and is unnecessary for this.
● Custom actions and job workers introduce development and operational overhead when a native Manual approval action already solves the.
Workflow: Use CodeBuild to execute the tests, insert a Manual approval action in CodePipeline immediately before the production CodeDeploy stage with SNS notifications to approvers, → proceed to the production deploy after approval.
Adopt the DynamoDB Streams Kinesis Adapter with KCL and offload heavy aggregations
The Kinesis Adapter with KCL provides coordinated, scalable shard consumption with checkpointing, and Flink efficiently handles stateful aggregations.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The Kinesis Adapter with KCL provides coordinated, scalable shard consumption with checkpointing, and Flink efficiently handles stateful aggregations.
● Scene fit: The Kinesis Adapter with KCL provides coordinated, scalable shard consumption with checkpointing, and Flink efficiently handles stateful aggregations.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Enhanced fan-out is not supported by DynamoDB Streams and adding consumers does not bypass per-shard read limits.
● Tuning Lambda and table capacity does not overcome shared shard throughput limits or multi-consumer contention.
● Adds cost and complexity and AWS Glue streaming does not natively consume DynamoDB Streams.
Workflow: Adopt the DynamoDB Streams Kinesis Adapter with KCL and offload heavy aggregations to Amazon Managed Service for Apache Flink.
ALB target group health checks are misconfigured
Incorrect path, port, or success codes keep new targets unhealthy, so AllowTraffic cannot complete even though CodeDeploy logs show no direct errors.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Incorrect path, port, or success codes keep new targets unhealthy, so AllowTraffic cannot complete even though CodeDeploy logs show no direct errors.
● Scene fit: Incorrect path, port, or success codes keep new targets unhealthy, so AllowTraffic cannot complete even though CodeDeploy logs show no direct errors.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● WAF evaluates client requests and does not interfere with ALB-to-target health probes.
● Scale-in can disrupt deployments, but it typically surfaces as instance or lifecycle event failures, not solely as a silent AllowTraffic failure.
● Missing permissions cause explicit IAM errors in the deployment logs rather than a quiet failure at AllowTraffic.
Workflow: ALB target group health checks are misconfigured.
an Amazon EventBridge rule that listens for Trusted Advisor check status updates
EventBridge can filter Trusted Advisor status-change events and forward them to SNS for immediate notifications. A scheduled Lambda can query Trusted Advisor and publish actionable findings.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can filter Trusted Advisor status-change events and forward them to SNS for immediate notifications.
● Scene fit: EventBridge can filter Trusted Advisor status-change events and forward them to SNS for immediate notifications. A scheduled Lambda.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Built-in Trusted Advisor notifications are sent weekly, not frequently enough for rapid cost control.
● EventBridge does not offer a one-click auto-notification for Trusted Advisor and requires explicit rules and targets.
● EventBridge does not natively target SES, so email alerts should be routed through SNS or Lambda instead.
Workflow: Create an Amazon EventBridge rule that listens for Trusted Advisor check status updates and route matches to an Amazon SNS topic for email alerts → Schedule an AWS Lambda function daily to.
Enforce HTTPS to the origins and enable CloudFront field-level encryption, and have
Field-level encryption protects specific fields at CloudFront while HTTPS secures transport, and long max-age keeps objects in edge caches longer to raise hit ratio.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Field-level encryption protects specific fields at CloudFront while HTTPS secures transport, and long max-age keeps objects in edge.
● Scene fit: Field-level encryption protects specific fields at CloudFront while HTTPS secures transport, and long max-age keeps objects in edge.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Signed URLs restrict who can access content but they do not encrypt sensitive form fields or satisfy field-level protection.
● An origin access identity secures S3 origins and varying on request headers usually fragments cache entries, and neither approach.
● KMS provides encryption at rest and Origin Shield optimizes origin fetches, but these do not encrypt fields at the.
Workflow: Enforce HTTPS to the origins and enable CloudFront field-level encryption, and have the origin send Cache-Control max-age with the longest safe value.
the site on stateless EC2 in an Auto Scaling group and move
Kinesis Data Firehose provides managed, near real time delivery to S3 and can load Redshift via an S3 staging bucket, aligning with the scale and analytics.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Kinesis Data Firehose provides managed, near real time delivery to S3 and can load Redshift via an S3.
● Scene fit: Kinesis Data Firehose provides managed, near real time delivery to S3 and can load Redshift via an S3.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Kinesis Data Streams does not natively deliver to S3 or Redshift and requires a consumer or Firehose to sink.
● Athena is a query service over S3 and is not used to ingest or load streaming data into S3.
Workflow: Run the site on stateless EC2 in an Auto Scaling group and move SQL Server to Amazon RDS → use Amazon Kinesis Data Firehose to deliver click events to Amazon S3 for.
Define AppSpec hooks for the ECS deployment and use the AfterAllowTestTraffic event
This uses the CodeDeploy lifecycle hook designed for validating the green environment before traffic and leverages hook failure to roll back automatically.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses the CodeDeploy lifecycle hook designed for validating the green environment before traffic and leverages hook failure.
● Scene fit: This uses the CodeDeploy lifecycle hook designed for validating the green environment before traffic and leverages hook failure.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Places testing in a separate CodePipeline stage using CodeBuild and aborts the deployment by calling the CLI.
● Runs tests only after production traffic is already shifted and then attempts to cancel the deployment via the CLI.
● Runs tests outside of the CodeDeploy lifecycle and relies on a separate CLI call to cancel the deployment.
Workflow: Define AppSpec hooks for the ECS deployment and use the AfterAllowTestTraffic event to invoke an AWS Lambda function that runs the tests → if any test fails, return an error from the function to trigger rollback.
an Amazon CloudFront distribution for the S3 static content and introduce DynamoDB
CloudFront globally caches S3 assets and DAX provides API-compatible in-memory caching to absorb repetitive DynamoDB reads with microsecond latency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFront globally caches S3 assets and DAX provides API-compatible in-memory caching to absorb repetitive DynamoDB reads with microsecond latency.
● Scene fit: CloudFront globally caches S3 assets and DAX provides API-compatible in-memory caching to absorb repetitive DynamoDB reads with microsecond latency.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Lambda@Edge is not intended to cache static files at scale and global tables do not remove duplicate read pressure at the application tier.
● MediaStore targets media workflows and scaling Redis does not directly solve duplicated DynamoDB reads or provide a global CDN for S3 objects.
● MediaPackage is built for video packaging and Memcached adds complexity without the DynamoDB-aware caching that DAX provides.
Workflow: Create an Amazon CloudFront distribution for the S3 static content and introduce DynamoDB Accelerator to offload repeated reads.
API Gateway stage canary, set 10% traffic, monitor in CloudWatch
API Gateway stage canaries natively split a defined percentage of requests and expose separate canary metrics in CloudWatch.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: API Gateway stage canaries natively split a defined percentage of requests and expose separate canary metrics in CloudWatch.
● Scene fit: API Gateway stage canaries natively split a defined percentage of requests and expose separate canary metrics in CloudWatch.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Lambda alias weighting shifts Lambda versions, not API Gateway stage deployments.
● Route 53 works at DNS or domain level, not for one API path, and client caching makes exact percentages less reliable.
● Usage plans apply quotas and throttling.
Workflow: Enable API Gateway stage canary → set 10% traffic → monitor in CloudWatch.
a redrive policy to send failed messages to a DLQ
A redrive policy moves messages that exceed maxReceiveCount to a dead-letter queue for analysis.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A redrive policy moves messages that exceed maxReceiveCount to a dead-letter queue for analysis.
● Scene fit: A redrive policy moves messages that exceed maxReceiveCount to a dead-letter queue for analysis.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● FIFO ordering does not help capture or inspect failed messages.
● Longer retention keeps messages longer but does not isolate failures for inspection.
● Long polling reduces empty receives and costs but does not capture failed messages.
Workflow: Enable a redrive policy to send failed messages to a DLQ.
AWS Config custom rule org-wide with an aggregator + SSM Automation builds
An organization-wide custom rule checks AMI IDs and the aggregator gives centralized compliance visibility. Centrally sharing the latest AMI and unsharing the previous one blocks new launches from the old AMI.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An organization-wide custom rule checks AMI IDs and the aggregator gives centralized compliance visibility.
● Scene fit: An organization-wide custom rule checks AMI IDs and the aggregator gives centralized compliance visibility. Centrally sharing the latest.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● SCPs can enforce launches only from approved AMIs but do not provide a centralized compliance view.
● Copying proliferates stale AMIs and does not prevent new launches from old copies.
● Decentralized AMI creation increases drift and complicates retiring old AMIs and compliance reporting.
Workflow: AWS Config custom rule org-wide with an aggregator → SSM Automation builds AMI centrally → share new AMI with org and unshare old.
Increase stream retention to 96 hours + Run KCL workers in an
Extending retention gives more time for backlogged consumers to process before records expire with minimal operational change. Scaling out KCL workers based on lag increases consumer capacity with minimal code changes since KCL manages shard leases.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Extending retention gives more time for backlogged consumers to process before records expire with minimal operational change.
● Scene fit: Extending retention gives more time for backlogged consumers to process before records expire with minimal operational change. Scaling.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Enhanced Fan-Out increases per-consumer throughput but does not fix a single-host bottleneck and requires consumer changes.
● More shards boost producer throughput and potential parallelism but do not resolve an under-provisioned single consumer.
● Requires a larger architectural change and is not the minimal path to address current lag.
Workflow: Increase stream retention to 96 hours → Run KCL workers in an Auto Scaling group and scale on MillisBehindLatest.
Initiate ContinueUpdateRollback for the stack + Fix resource issues to match the
ContinueUpdateRollback resumes the failed rollback after you correct blocking issues so CloudFormation can return the stack to a stable state. Manually correcting out-of-band changes or dependency issues enables the rollback to succeed when continued.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ContinueUpdateRollback resumes the failed rollback after you correct blocking issues so CloudFormation can return the stack to a stable state.
● Scene fit: ContinueUpdateRollback resumes the failed rollback after you correct blocking issues so CloudFormation can return the stack to a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Drift detection only reports differences between actual resources and the template; it does not change stack state or resolve rollback failures.
● Change sets preview proposed updates; they do not resolve a stack stuck in UPDATE_ROLLBACK_FAILED.
● Updates are blocked while the stack is in UPDATE_ROLLBACK_FAILED, and reusing the old template does not advance the rollback.
Workflow: Initiate ContinueUpdateRollback for the stack → Fix resource issues to match the last successful stack state.
Amazon MemoryDB
MemoryDB is a Redis-compatible durable primary database with Multi-AZ replication and strong read-after-write consistency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: MemoryDB is a Redis-compatible durable primary database with Multi-AZ replication and strong read-after-write consistency.
● Scene fit: MemoryDB is a Redis-compatible durable primary database with Multi-AZ replication and strong read-after-write consistency.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DAX is a cache for DynamoDB, not a standalone Redis-compatible durable database.
● Memcached has no durable persistence or native replication model for a primary database.
● ElastiCache for Redis is designed mainly as a cache, even though it supports replication and persistence options.
Workflow: Amazon MemoryDB.
Register the on-premises servers with AWS Systems Manager by using Systems Manager
Hybrid Activations enrolls on-premises machines as managed instances so Patch Manager can target and patch them alongside EC2. An instance profile grants the SSM Agent permissions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Hybrid Activations enrolls on-premises machines as managed instances so Patch Manager can target and patch them alongside EC2.
● Scene fit: Hybrid Activations enrolls on-premises machines as managed instances so Patch Manager can target and patch them alongside EC2.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Long-lived access keys on servers are insecure and not the supported method for onboarding on-premises instances to AWS Systems.
● EventBridge can trigger workflows, but patching schedules are best enforced with Systems Manager Maintenance Windows and Patch Manager.
Workflow: Register the on-premises servers with AWS Systems Manager by using Systems Manager Hybrid Activations → Attach an IAM instance profile to the EC2 instances that allows AWS Systems Manager to manage them.
AWS Config with approved-amis-by-id
This managed rule evaluates EC2 instances against a sanctioned AMI ID list and can push compliance changes to SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This managed rule evaluates EC2 instances against a sanctioned AMI ID list and can push compliance changes to SNS.
● Scene fit: This managed rule evaluates EC2 instances against a sanctioned AMI ID list and can push compliance changes to SNS.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Inspector focuses on vulnerability and exposure assessments, not validating instance AMI IDs against an allowlist.
● These are preventive controls that would block launching nonapproved AMIs, contrary to the non-blocking requirement.
● Security Hub aggregates findings but does not natively evaluate AMI allowlists; it relies on integrated sources like Config.
Workflow: AWS Config with approved-amis-by-id.
AWS CloudFormation StackSets from a central administrator account to roll out identical
StackSets natively support consistent, automated multi-Region deployments from an admin account while keeping each Region's stacks isolated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: StackSets natively support consistent, automated multi-Region deployments from an admin account while keeping each Region's stacks isolated.
● Scene fit: StackSets natively support consistent, automated multi-Region deployments from an admin account while keeping each Region's stacks isolated.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Change sets preview updates to a single stack and are not designed to orchestrate deployments across multiple Regions.
● While CodePipeline can coordinate actions across Regions, it is not the most direct mechanism to consistently provision identical regional infrastructure compared to StackSets.
● The AWS CLI supports only a single --region flag and does not provide a --regions parameter for parallel multi-Region stack creation.
Workflow: Use AWS CloudFormation StackSets from a central administrator account to roll out identical stacks to selected Regions.
EventBridge rule for AWS Health retirement events running an SSM Automation runbook
AWS Health precisely emits retirement events and SSM Automation reliably orchestrates the stop/start workflow with guardrails and logging.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health precisely emits retirement events and SSM Automation reliably orchestrates the stop/start workflow with guardrails and logging.
● Scene fit: AWS Health precisely emits retirement events and SSM Automation reliably orchestrates the stop/start workflow with guardrails and logging.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Auto Recovery addresses impaired system status checks, not scheduled retirements, and does not orchestrate a stop then start sequence.
● Can work but requires custom code and lacks SSM Automation's built-in controls, runbooks, and operational safety.
● ASG reacts to health issues and does not proactively handle scheduled retirements or guarantee a stop/start on the same instance.
Workflow: EventBridge rule for AWS Health retirement events running an SSM Automation runbook to stop → start instances.
CloudTrail S3 data events for the bucket, store in S3, and query
CloudTrail data events capture object-level API calls; storing in S3 and querying with Athena provides low-cost, on-demand search.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail data events capture object-level API calls; storing in S3 and querying with Athena provides low-cost, on-demand search.
● Scene fit: CloudTrail data events capture object-level API calls; storing in S3 and querying with Athena provides low-cost, on-demand search.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EventBridge does not directly receive S3 object-level API events unless CloudTrail emits them as data events.
● S3 access logs lack full API audit details and running OpenSearch increases ongoing cost.
● Management events do not include S3 data events, so object-level actions would be missed.
Workflow: Turn on CloudTrail S3 data events for the bucket → store in S3, and query with Athena.
interface VPC endpoints for AWS Systems Manager in the VPC hosting the
Interface VPC endpoints keep Session Manager traffic on the AWS private network path. The SSM Agent requires an instance role with permissions to communicate with Systems Manager.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Interface VPC endpoints keep Session Manager traffic on the AWS private network path.
● Scene fit: Interface VPC endpoints keep Session Manager traffic on the AWS private network path. The SSM Agent requires an.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Session Manager does not use SSH, so opening port 22 is unnecessary and reduces security.
● A NAT gateway provides outbound internet access and does not enable private inbound Session Manager connectivity.
● Client VPN is optional and does not ensure Session Manager uses a private VPC path without the necessary endpoints.
Workflow: Create interface VPC endpoints for AWS Systems Manager in the VPC hosting the instances → Attach an IAM instance profile to the EC2 instances that grants Systems Manager permissions.
Update the existing template to set DeletionPolicy Retain on every resource, delete
Retain ensures resources persist through stack deletion, allowing you to import them into a newly named stack and then revert template settings.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Retain ensures resources persist through stack deletion, allowing you to import them into a newly named stack and.
● Scene fit: Retain ensures resources persist through stack deletion, allowing you to import them into a newly named stack and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Snapshot will back up and then delete the EBS volume, which prevents importing the original volume and violates the.
● Hooks validate or block operations but do not retain resources or handle importing existing resources into a new stack.
● DependsOn only orders resource creation within a stack and does not facilitate renaming or importing existing resources.
Workflow: Update the existing template to set DeletionPolicy Retain on every resource, delete the stack so the resources are preserved → create a new stack with the new name, import the retained S3.
NLB access logs to S3 with SSE-S3, allow delivery.logs.amazonaws.com to write, and
This is the AWS-recommended configuration: NLB access logs to S3, encrypted at rest with SSE-S3, a bucket policy permitting the logging principal to write, and least-privilege.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is the AWS-recommended configuration: NLB access logs to S3, encrypted at rest with SSE-S3, a bucket policy permitting the logging principal to write, and least-privilege read access for the team.
● Scene fit: This is the AWS-recommended configuration: NLB access logs to S3, encrypted at rest with SSE-S3, a bucket policy.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● NLB access logs do not deliver directly to CloudWatch Logs; they are delivered only to Amazon S3.
● NLB access logs are delivered expecting SSE-S3; using SSE-KMS requires complex key policies and is not the standard supported path.
● VPC Flow Logs capture ENI-level traffic, not NLB access logs or per-connection load balancer details.
Workflow: Enable NLB access logs to S3 with SSE-S3 → allow delivery.logs.amazonaws.com to write, and grant team read via IAM.
In the source account, create an encrypted copy of the AMI using
You must encrypt with a customer managed key and allow the target account to create grants on that key before sharing the AMI. The Auto Scaling.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: You must encrypt with a customer managed key and allow the target account to create grants on that.
● Scene fit: You must encrypt with a customer managed key and allow the target account to create grants on that.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AMIs encrypted with the AWS managed KMS key cannot be shared across accounts, so this fails cross-account launches.
● Launch permissions do not grant rights to use the KMS key that encrypts the EBS snapshots, so instances will.
Workflow: In the source account → create an encrypted copy of the AMI using the customer managed KMS key → update that key policy to let the target account create grants, and share.
By default, CodeDeploy removes files installed by the latest deployment and a
CodeDeploy cleans up files from the last deployment and during rollback reapplies the earlier revision, so selecting Retain the content keeps files that were not part.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy cleans up files from the last deployment and during rollback reapplies the earlier revision, so selecting Retain.
● Scene fit: CodeDeploy cleans up files from the last deployment and during rollback reapplies the earlier revision, so selecting Retain.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Incorrectly assumes that CodeDeploy retains arbitrary files and that Overwrite preserves unexpected files, which it does not.
● Fail the deployment simply forces a failure and does not ensure manual files are preserved during rollback.
● Mischaracterizes the cleanup step and Overwrite does not guarantee that non-revision files survive rollback.
Workflow: By default, CodeDeploy removes files installed by the latest deployment and a rollback redeploys the previous revision cleanly → use Retain the content so files that were manually added are preserved.
CloudWatch agent to send OS auth logs to CloudWatch Logs; metric filter
The CloudWatch agent streams auth.log or /var/log/secure, a metric filter detects login patterns, and an alarm notifies SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The CloudWatch agent streams auth.log or /var/log/secure, a metric filter detects login patterns, and an alarm notifies SNS.
● Scene fit: The CloudWatch agent streams auth.log or /var/log/secure, a metric filter detects login patterns, and an alarm notifies SNS.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● GuardDuty reports suspicious behavior, not every successful operating-system login.
● CloudTrail records AWS control-plane API calls, not in-instance operating-system authentication.
● Scheduled upload and custom parsing can work but introduce delay, scripts, storage events, and maintenance.
Workflow: Use CloudWatch agent to send OS auth logs to CloudWatch Logs → metric filter + SNS alarm.
Amazon Route 53 ARC zonal shift for the ALB
ARC zonal shift temporarily evacuates one AZ for the ALB without changing infrastructure, steering requests only to healthy AZs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ARC zonal shift temporarily evacuates one AZ for the ALB without changing infrastructure, steering requests only to healthy AZs.
● Scene fit: ARC zonal shift temporarily evacuates one AZ for the ALB without changing infrastructure, steering requests only to healthy AZs.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Replaces unhealthy instances but continues to route traffic through the impaired AZ, so it does not isolate the zone.
● Can work but requires modifying the load balancer configuration or stack, which is not the minimal operational change asked for.
● Requires additional ALB endpoints and does not address isolating a single AZ behind one ALB, adding unnecessary complexity.
Workflow: Amazon Route 53 ARC zonal shift for the ALB.
Immutable as the deployment policy in the Elastic Beanstalk environment for upcoming
Immutable updates launch a parallel Auto Scaling group with the new version, keep the old fleet serving traffic, and enable fast rollback by terminating the new.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Immutable updates launch a parallel Auto Scaling group with the new version, keep the old fleet serving traffic.
● Scene fit: Immutable updates launch a parallel Auto Scaling group with the new version, keep the old fleet serving traffic.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Blue/green cutovers expect the database to be decoupled from an environment-managed RDS; keeping it tightly coupled risks data loss and complicates switching.
● Rolling with additional batch still updates in-place and can fail mid-deploy, so rollback remains slow and operationally complex.
● All at once replaces all instances simultaneously, which increases downtime risk and does not provide a safer rollback path.
Workflow: Configure Immutable as the deployment policy in the Elastic Beanstalk environment for upcoming releases.
Systems Manager Inventory on managed instances and sync to S3 + EventBridge
SSM Inventory collects package metadata from managed instances and Resource Data Sync exports it to S3 for reporting. A scheduled rule can compare EC2 instances with SSM Managed Instances and notify on uncovered hosts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SSM Inventory collects package metadata from managed instances and Resource Data Sync exports it to S3 for reporting.
● Scene fit: SSM Inventory collects package metadata from managed instances and Resource Data Sync exports it to S3 for reporting.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Inspector focuses on vulnerability findings and coverage, not full OS package inventory or SSM management gaps.
● Config does not gather OS package listings and by itself cannot detect SSM management coverage gaps.
● Run Command targets only already managed instances and cannot discover unmanaged ones.
Workflow: Enable Systems Manager Inventory on managed instances and sync to S3 → EventBridge schedule every 45 minutes invoking Lambda to diff EC2 vs SSM managed and alert.
the Lambda execution role in Account Y with permission to use the
The Lambda execution role must include the required VPC and elasticfilesystem permissions to mount and access EFS via an access point. You need network reachability via.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The Lambda execution role must include the required VPC and elasticfilesystem permissions to mount and access EFS via.
● Scene fit: The Lambda execution role must include the required VPC and elasticfilesystem permissions to mount and access EFS via.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● EFS does not support PrivateLink, so you cannot mount EFS over PrivateLink across accounts.
● EFS is not a shareable resource type in AWS RAM, so it cannot be shared directly to Lambda this.
● SCPs set permission guardrails but do not grant access, so resource and IAM policies are still required.
Workflow: Configure the Lambda execution role in Account Y with permission to use the VPC and mount the EFS access point → Establish VPC peering between the VPCs in Account X and Account.
AWS Systems Manager Parameter Store with CloudFormation parameters to resolve the latest
This uses CloudFormation support for SSM Parameter Store to dynamically resolve AMI IDs at create or update time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses CloudFormation support for SSM Parameter Store to dynamically resolve AMI IDs at create or update time.
● Scene fit: This uses CloudFormation support for SSM Parameter Store to dynamically resolve AMI IDs at create or update time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Proposes relying on AWS Service Catalog to obtain AMI IDs during CloudFormation deployments.
● Suggests using State Manager as a parameter source even though it is designed for enforcing instance state and compliance.
● Introduces a custom automation to modify templates rather than using CloudFormation's native parameter resolution.
Workflow: Use AWS Systems Manager Parameter Store with CloudFormation parameters to resolve the latest AMI IDs and run stack updates when rolling out new images.
IAM ABAC with matching tags
Attribute-based access control evaluates principal and resource tags so newly tagged resources are covered without policy edits.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Attribute-based access control evaluates principal and resource tags so newly tagged resources are covered without policy edits.
● Scene fit: Attribute-based access control evaluates principal and resource tags so newly tagged resources are covered without policy edits.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Permissions boundaries set a maximum permission scope and do not grant access or auto-include new resources.
● Wildcard resources broaden access and violate least privilege rather than providing targeted automatic coverage.
● SCPs define account guardrails and cannot grant resource access or remove the need for identity policy updates.
Workflow: IAM ABAC with matching tags.
Allow inbound traffic to the application tier security group on TCP 6379
Ensuring the correct security group rule on port 6379 allows the application layer to establish connections to the Redis service endpoint. Fargate tasks require task ENIs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Ensuring the correct security group rule on port 6379 allows the application layer to establish connections to the.
● Scene fit: Ensuring the correct security group rule on port 6379 allows the application layer to establish connections to the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While important, the error points to network reachability for registry auth rather than missing IAM permissions.
● Adding a Secrets Manager endpoint does not address the ECR registry authentication pull failure indicated by the error.
● Fargate tasks can run entirely in private subnets without public IPs when proper VPC endpoints or NAT are configured.
Workflow: Allow inbound traffic to the application tier security group on TCP 6379 from the Redis subnet CIDR → Enable awsvpc task networking with dedicated ENIs and add an ecr.api VPC endpoint.
EventBridge CodePipeline state-change events invoking Systems Manager Automation to start/stop EC2 and
Event-driven automation tied to pipeline lifecycle reliably starts and stops both EC2 and RDS without redesign.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Event-driven automation tied to pipeline lifecycle reliably starts and stops both EC2 and RDS without redesign.
● Scene fit: Event-driven automation tied to pipeline lifecycle reliably starts and stops both EC2 and RDS without redesign.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adds custom scripting in build steps, can miss shutdown on failures, and increases operational burden.
● Time-based scheduling risks running when no tests occur and missing ad hoc pipeline runs.
● Unnecessary architectural changes and added complexity that do not directly solve the event-driven requirement.
Workflow: EventBridge CodePipeline state-change events invoking Systems Manager Automation to start/stop EC2 and RDS.
All-at-once deployment
Deploys to all instances simultaneously for the fastest rollout with brief downtime.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Deploys to all instances simultaneously for the fastest rollout with brief downtime.
● Scene fit: Deploys to all instances simultaneously for the fastest rollout with brief downtime.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs a duplicate environment, increasing cost and time before swapping traffic.
● Deploys in batches to preserve capacity, which is slower than all-at-once.
● Creates a new fleet for safety and rollback but is slower and costlier.
Workflow: All-at-once deployment.
an external Lambda extension to emit X-Ray segments and subsegments, enable active
This provides full distributed tracing with per-request segments and subsegments, anomaly detection via X-Ray Insights, and near-real-time notifications using EventBridge and CloudWatch.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides full distributed tracing with per-request segments and subsegments, anomaly detection via X-Ray Insights, and near-real-time notifications.
● Scene fit: This provides full distributed tracing with per-request segments and subsegments, anomaly detection via X-Ray Insights, and near-real-time notifications.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Consolidating functions is unnecessary and reduces modularity and reuse when X-Ray groups can correlate traces across functions.
● An internal extension runs in-process and is less suitable for reliably buffering and exporting telemetry compared to an external.
● CloudWatch Logs Insights analyzes logs rather than trace data, so it cannot surface trace-level anomalies and causal relationships like.
Workflow: Use an external Lambda extension to emit X-Ray segments and subsegments → enable active tracing → grant X-Ray permissions → define X-Ray groups for each workflow → enable X-Ray Insights, and route.
CloudFront Origin Access Control and lock the S3 bucket policy to the
OAC signs origin requests and, with a bucket policy allowing only that distribution, blocks direct S3 access.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: OAC signs origin requests and, with a bucket policy allowing only that distribution, blocks direct S3 access.
● Scene fit: OAC signs origin requests and, with a bucket policy allowing only that distribution, blocks direct S3 access.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Signed URLs authorize CloudFront requests but do not, by themselves, prevent direct S3 URL access without a restrictive bucket policy.
● Blocks public ACLs and policies but does not ensure CloudFront is the only reader or stop authenticated direct S3 access.
● Referer-based controls are not secure or reliable for enforcing CloudFront-only access and can be spoofed.
Workflow: Configure CloudFront Origin Access Control and lock the S3 bucket policy to the distribution.
Activate DynamoDB Accelerator and put CloudFront in front of the S3 origin
DAX provides in-memory, API-compatible caching for DynamoDB reads and CloudFront caches S3 content at edge locations to reduce latency and egress cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: DAX provides in-memory, API-compatible caching for DynamoDB reads and CloudFront caches S3 content at edge locations to reduce latency and egress cost.
● Scene fit: DAX provides in-memory, API-compatible caching for DynamoDB reads and CloudFront caches S3 content at edge locations to reduce.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While Redis can be placed in front of DynamoDB, it is not natively integrated and typically costs more operational effort than DAX for this read-heavy pattern.
● The API is behind an ALB rather than API Gateway, and S3 Transfer Acceleration speeds uploads rather than caching reads for cost and latency gains.
● Memcached does not integrate to cache S3 objects directly, so it will not accelerate or reduce cost for static S3 content delivery.
Workflow: Activate DynamoDB Accelerator and put CloudFront in front of the S3 origin.
In the Audit account, set up an Amazon Kinesis Data Streams stream
Only Kinesis Data Streams is supported for cross-account CloudWatch Logs subscriptions, and Lambda can transform and index data into OpenSearch centrally.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Only Kinesis Data Streams is supported for cross-account CloudWatch Logs subscriptions, and Lambda can transform and index data.
● Scene fit: Only Kinesis Data Streams is supported for cross-account CloudWatch Logs subscriptions, and Lambda can transform and index data.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch Logs cross-account subscriptions cannot target Firehose, so this does not provide a supported cross-account delivery path.
● CloudWatch Logs subscription filters do not support SQS as a destination, making this integration unsupported.
● CloudWatch Logs cross-account subscriptions cannot deliver directly to Lambda, so this approach is not supported.
Workflow: In the Audit account → set up an Amazon Kinesis Data Streams stream with an AWS Lambda consumer that indexes records into an Amazon OpenSearch Service domain, and configure CloudWatch Logs subscription.
AWS Config with the approved-amis-by-id managed rule listing permitted AMI IDs, and
AWS Config provides detective compliance checks for AMI allowlists without blocking launches and can push compliance state changes to SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config provides detective compliance checks for AMI allowlists without blocking launches and can push compliance state changes to SNS.
● Scene fit: AWS Config provides detective compliance checks for AMI allowlists without blocking launches and can push compliance state changes.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Inspector focuses on vulnerability and exposure assessments rather than validating instance AMI IDs against an approved list.
● SCPs and IAM policies are preventive controls that would block the release team from launching nonapproved AMIs, which violates the requirement.
● Trusted Advisor does not offer a check to validate that EC2 instances were launched from a specific approved AMI list.
Workflow: Configure AWS Config with the approved-amis-by-id managed rule listing permitted AMI IDs, and send AWS Config compliance change notifications to an Amazon SNS topic subscribed by both teams.
AWS Systems Manager Run Command for instance changes instead of SSH or
Run Command executes commands on managed instances securely without distributing SSH keys. Relying on the CodeBuild service role replaces static credentials with short-lived role credentials. Parameter.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Run Command executes commands on managed instances securely without distributing SSH keys.
● Scene fit: Run Command executes commands on managed instances securely without distributing SSH keys. Relying on the CodeBuild service role.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Storing credentials in S3, even encrypted, expands exposure and is not the recommended secret management pattern for builds.
● Secrets Manager does not use the SecureString type; that is specific to Parameter Store.
● Still relies on SSH and key handling rather than using SSM native controls.
Workflow: Use AWS Systems Manager Run Command for instance changes instead of SSH or scp with an S3-stored key → Give the CodeBuild IAM role least-privilege and remove AWS keys from buildspec →.
Lambda target posting to Slack webhook + EventBridge rule for CodePipeline Pipeline
Use a Lambda function as the EventBridge target to call the Slack incoming webhook. Match pipeline-level failure events using the Pipeline Execution State Change pattern.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use a Lambda function as the EventBridge target to call the Slack incoming webhook.
● Scene fit: Use a Lambda function as the EventBridge target to call the Slack incoming webhook. Match pipeline-level failure events using the Pipeline Execution State Change pattern.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EventBridge has no native Slack target, so it cannot deliver directly to Slack.
● Slack does not poll SQS, so this will not deliver messages to Slack.
● Captures individual action failures, not overall pipeline execution failures.
Workflow: Lambda target posting to Slack webhook → EventBridge rule for CodePipeline Pipeline Execution FAILED.
EventBridge rule for AWS Health EC2 events to SNS
EventBridge natively receives AWS Health events, enabling a fast rule-to-SNS delivery pattern.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge natively receives AWS Health events, enabling a fast rule-to-SNS delivery pattern.
● Scene fit: EventBridge natively receives AWS Health events, enabling a fast rule-to-SNS delivery pattern.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but requires deploying and configuring a solution, which is slower than creating a simple rule.
● Status checks do not notify about scheduled maintenance or retirement from AWS Health.
● AWS Health events are not in CloudWatch Logs by default, adding unnecessary setup.
Workflow: EventBridge rule for AWS Health EC2 events to SNS.
EventBridge to run a Lambda that calls the file share RefreshCache API
Invoking RefreshCache updates the file gateway directory so new S3 objects appear on the SMB/NFS share.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Invoking RefreshCache updates the file gateway directory so new S3 objects appear on the SMB/NFS share.
● Scene fit: Invoking RefreshCache updates the file gateway directory so new S3 objects appear on the SMB/NFS share.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DataSync moves data but does not refresh a file gateway's directory cache or listings.
● Refreshes metadata when a file is opened but does not update directory listings for newly created objects.
● Replication copies objects between buckets and does not trigger a file gateway cache refresh.
Workflow: Use EventBridge to run a Lambda that calls the file share RefreshCache API.
Include buildspec.yml with Maven build, test, and package commands + Create an
CodeBuild uses buildspec.yml to run Maven phases and define artifacts. CodeBuild integrates with GitHub to pull source and produce artifacts. A GitHub webhook can automatically start a CodeBuild on push events.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeBuild uses buildspec.yml to run Maven phases and define artifacts.
● Scene fit: CodeBuild uses buildspec.yml to run Maven phases and define artifacts. CodeBuild integrates with GitHub to pull source and produce artifacts. A GitHub webhook can automatically start a CodeBuild on push events.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CodePipeline without a build stage cannot compile or test the code.
● Managing builds on EC2 adds overhead and does not provide native CI triggers.
● CodeDeploy handles deployments, not compiling or testing source code.
Workflow: Include buildspec.yml with Maven build, test, and package commands → Create an AWS CodeBuild project using GitHub as the source → Enable GitHub webhook trigger for pushes.
Trigger a Lambda from S3 object updates to parse the CSV and
Event-driven updates occur only when the list changes, keeping cost and operations minimal while enforcing IP allow listing at the network layer.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Event-driven updates occur only when the list changes, keeping cost and operations minimal while enforcing IP allow listing at the network layer.
● Scene fit: Event-driven updates occur only when the list changes, keeping cost and operations minimal while enforcing IP allow listing at the network layer.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Functionally possible but adds recurring WAF costs and extra components compared to security groups.
● Polling is costlier and more complex than reacting only when the file changes.
● Direct Connect introduces significant cost and is unnecessary for simple public ingress allow listing.
Workflow: Trigger a Lambda from S3 object updates to parse the CSV and sync ALB security group ingress to the listed proxy IPs only.
access logging on the Application Load Balancer
ALB access logs record each request with request_processing_time, target_processing_time, and response_processing_time fields for precise latency analysis.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB access logs record each request with request_processing_time, target_processing_time, and response_processing_time fields for precise latency analysis.
● Scene fit: ALB access logs record each request with request_processing_time, target_processing_time, and response_processing_time fields for precise latency analysis.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● The CloudWatch agent collects host-level metrics and logs from instances and does not expose per-request ALB latency details.
● CloudTrail captures control-plane API calls and not data-plane request timings flowing through an Application Load Balancer.
● CloudWatch provides aggregated ALB metrics and does not include per-request processing breakdowns unless you publish custom metrics.
Workflow: Enable access logging on the Application Load Balancer.
CloudFormation StackSets with service-managed permissions targeting OUs
Integrates with Organizations to auto-deploy to new accounts in target OUs and remove stack instances when accounts leave, ideal for deploying CloudTrail and IAM roles.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Integrates with Organizations to auto-deploy to new accounts in target OUs and remove stack instances when accounts leave, ideal for deploying CloudTrail and IAM roles.
● Scene fit: Integrates with Organizations to auto-deploy to new accounts in target OUs and remove stack instances when accounts leave, ideal for deploying CloudTrail and IAM roles.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Custom automation requires complex cross-account permissions and does not natively clean up when accounts leave.
● Baselines accounts but is heavier to adopt and does not inherently roll out custom IAM roles and cleanup for departing accounts without additional tooling.
● SCPs enforce permissions boundaries but cannot create CloudTrail trails or IAM roles.
Workflow: CloudFormation StackSets with service-managed permissions targeting OUs.
ECS with Fargate launch type + Site-to-Site VPN over IPsec
Fargate removes server and cluster management for containers, minimizing operational overhead. Site-to-Site VPN provides encrypted IPsec tunnels between on-premises networks and a VPC.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Fargate removes server and cluster management for containers, minimizing operational overhead.
● Scene fit: Fargate removes server and cluster management for containers, minimizing operational overhead. Site-to-Site VPN provides encrypted IPsec tunnels between on-premises networks and a VPC.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● PrivateLink exposes services privately within AWS but is not a general-purpose network connection from on-premises to a VPC.
● Direct Connect does not encrypt traffic by default, so it does not meet the encryption requirement on its own.
● Running EKS on self-managed EC2 nodes increases operational burden for node lifecycle and patching.
Workflow: ECS with Fargate launch type → Site-to-Site VPN over IPsec.
Define a custom action type in CodePipeline and run an associated on-premises
A CodePipeline custom action with an on-prem job worker polling for jobs is the supported pattern for long-running on-premises tasks and returning status.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A CodePipeline custom action with an on-prem job worker polling for jobs is the supported pattern for long-running.
● Scene fit: A CodePipeline custom action with an on-prem job worker polling for jobs is the supported pattern for long-running.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Lambda is limited to 15 minutes per invocation and is not suitable for a 90-minute on-premises scan.
● Exposing the service and using a public S3 bucket is insecure, and CodePipeline does not poll S3 buckets for results.
● Step Functions does not directly run on-prem tools and would still require a suitable compute runtime, leaving the core integration challenge unresolved.
Workflow: Define a custom action type in CodePipeline and run an associated on-premises job worker that polls CodePipeline, executes the scanner after the source stage, and reports the status back.
Link the codebase to GitHub via AWS CodeStar Connections, use AWS CodeBuild
Blue/green on Elastic Beanstalk with an external Multi-AZ RDS and CNAME swap enables zero-downtime releases and rapid rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Blue/green on Elastic Beanstalk with an external Multi-AZ RDS and CNAME swap enables zero-downtime releases and rapid rollback.
● Scene fit: Blue/green on Elastic Beanstalk with an external Multi-AZ RDS and CNAME swap enables zero-downtime releases and rapid rollback.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● In-place updates on Elastic Beanstalk can briefly remove capacity and risk downtime during deployments.
● Using ECR plus CodeDeploy here is mismatched for test automation and still relies on in-place style updates that can.
● Provisioning separate RDS instances per environment complicates data consistency for blue/green and is not the recommended decoupled database pattern.
Workflow: Link the codebase to GitHub via AWS CodeStar Connections → use AWS CodeBuild for automated unit and functional tests, stand up two Elastic Beanstalk environments wired to a shared external Amazon RDS.
AWS CodeDeploy blue/green for EC2/On-Premises with ALB and a 75 minute termination
Launches a parallel fleet, supports traffic shifting after validation, and can auto-terminate the original instances after a configured wait time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Launches a parallel fleet, supports traffic shifting after validation, and can auto-terminate the original instances after a configured wait time.
● Scene fit: Launches a parallel fleet, supports traffic shifting after validation, and can auto-terminate the original instances after a configured wait time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A custom approach that lacks native orchestration, health tracking, and timed termination controls.
● Updates the existing fleet in place and does not create a separate green environment or timed retirement.
● Performs in-place changes without a full blue/green environment or automatic termination of the old fleet.
Workflow: AWS CodeDeploy blue/green for EC2/On-Premises with ALB and a 75 minute termination wait.
Immutable deployments with external RDS Multi-AZ
Creates a new Auto Scaling group at full capacity, shifts traffic only after health checks pass, and enables quick rollback with temporary extra cost during deployment.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Creates a new Auto Scaling group at full capacity, shifts traffic only after health checks pass, and enables quick rollback with temporary extra cost during deployment.
● Scene fit: Creates a new Auto Scaling group at full capacity, shifts traffic only after health checks pass, and enables quick rollback with temporary extra cost during deployment.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Zero downtime and rollback are possible, but running a parallel environment and keeping it as standby increases cost.
● Can still leave partially updated instances and couples the database lifecycle to the environment, which is discouraged for production.
● Replaces instances in place, risking downtime and no safe rollback path if issues occur.
Workflow: Immutable deployments with external RDS Multi-AZ.
the DynamoDB Streams Kinesis Adapter with the Kinesis Client Library to scale
The Kinesis Adapter enables highly scalable, coordinated consumption of DynamoDB Streams with minimal ops and integrates well with Flink for real-time analytics.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The Kinesis Adapter enables highly scalable, coordinated consumption of DynamoDB Streams with minimal ops and integrates well with.
● Scene fit: The Kinesis Adapter enables highly scalable, coordinated consumption of DynamoDB Streams with minimal ops and integrates well with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Index and capacity changes shift to synchronous queries and do not solve stream-consumer throttling while increasing cost and operational burden.
● Running ECS and Glue adds complexity and cost, and Glue is not a native consumer for DynamoDB Streams.
● Tuning Lambda and table capacity does not remove shared shard limits that lead to throttling when multiple consumers read the same stream.
Workflow: Use the DynamoDB Streams Kinesis Adapter with the Kinesis Client Library to scale out consumption across shards and offload complex aggregations to Amazon Managed Service for Apache Flink Studio.
Amazon MemoryDB for Redis
This is a Redis-compatible, durable in-memory database that supports multi-AZ persistence, strong consistency, and large-scale clusters.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is a Redis-compatible, durable in-memory database that supports multi-AZ persistence, strong consistency, and large-scale clusters.
● Scene fit: This is a Redis-compatible, durable in-memory database that supports multi-AZ persistence, strong consistency, and large-scale clusters.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is a high-performance cache designed for speed and read scaling, but it is not intended as a durable, multi-AZ primary database.
● Is a visualization and dashboarding service for metrics and logs rather than a database or data store.
● Is an in-memory cache without persistence or replication features, making it unsuitable for durable storage needs.
Workflow: Amazon MemoryDB for Redis.
Rollback reapplies the last successful revision and cleans files not in it
Rollback restores the prior revision and removes files not tracked by it; RETAIN preserves existing files during deployment.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Rollback restores the prior revision and removes files not tracked by it; RETAIN preserves existing files during deployment.
● Scene fit: Rollback restores the prior revision and removes files not tracked by it; RETAIN preserves existing files during deployment.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Blue/green replaces instances and does not preserve manual state on the old fleet.
● CodeDeploy does not keep arbitrary files by default, and OVERWRITE does not protect non-revision files.
● Auto rollback controls when rollback occurs, not how files are retained or deleted.
Workflow: Rollback reapplies the last successful revision and cleans files not in it → set fileExistsBehavior to RETAIN.
Amazon GuardDuty
Managed threat detection that analyzes CloudTrail, VPC flow logs, and DNS logs to identify suspicious activity.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Managed threat detection that analyzes CloudTrail, VPC flow logs, and DNS logs to identify suspicious activity.
● Scene fit: Managed threat detection that analyzes CloudTrail, VPC flow logs, and DNS logs to identify suspicious activity.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Performs vulnerability and exposure assessments; not continuous threat detection from CloudTrail/VPC/DNS.
● Focuses on sensitive data discovery and protection for S3, not threat detection.
● Supports investigation of security findings; does not perform the initial continuous detection.
Workflow: Amazon GuardDuty.
S3 server access logging and analyze logs with Amazon Athena
S3 access logs record per-request object details and Athena queries them serverlessly and cheaply for quick ranking.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 access logs record per-request object details and Athena queries them serverlessly and cheaply for quick ranking.
● Scene fit: S3 access logs record per-request object details and Athena queries them serverlessly and cheaply for quick ranking.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Offers aggregated metrics at bucket or prefix level and updates periodically, not ideal for immediate per-object ranking.
● Provides object-level API logging but is costly at scale, making it less economical than access logs for high-volume analysis.
● Requires provisioning and paying for a Redshift cluster, which is excessive for ad hoc per-object access analysis.
Workflow: Turn on S3 server access logging and analyze logs with Amazon Athena.
a BeforeAllowTraffic hook in the AppSpec file that checks and waits for
BeforeAllowTraffic is the Lambda pre-traffic hook that can block the shift until schema and data changes are fully propagated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: BeforeAllowTraffic is the Lambda pre-traffic hook that can block the shift until schema and data changes are fully propagated.
● Scene fit: BeforeAllowTraffic is the Lambda pre-traffic hook that can block the shift until schema and data changes are fully propagated.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● For Lambda, AfterAllowTestTraffic is not a supported lifecycle event and would not gate production traffic before the shift.
● AfterAllowTraffic runs only after traffic has already been shifted, so it cannot prevent the temporary error window.
● BeforeInstall is not a lifecycle hook for Lambda deployments in CodeDeploy and will not execute in this deployment type.
Workflow: Add a BeforeAllowTraffic hook in the AppSpec file that checks and waits for required database updates before shifting traffic to the new Lambda version.
CloudTrail and create an Amazon EventBridge rule for AWS API Call via
EventBridge can match CloudTrail management events such as DeleteTable and directly route them to SNS with minimal cost and near real-time delivery.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can match CloudTrail management events such as DeleteTable and directly route them to SNS with minimal cost.
● Scene fit: EventBridge can match CloudTrail management events such as DeleteTable and directly route them to SNS with minimal cost.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Ingesting CloudTrail into CloudWatch Logs and filtering with a metric adds ongoing log ingestion and Lambda costs, making it.
● DynamoDB Streams capture item-level mutations, not table-level management calls like DeleteTable, so this will never detect table deletions.
● AWS Config focuses on configuration state and would require a custom rule with potential evaluation lag and extra costs.
Workflow: Enable CloudTrail and create an Amazon EventBridge rule for AWS API Call via CloudTrail that matches dynamodb DeleteTable and targets an SNS topic.
an AWS Config managed rule for EBS volume encryption in all accounts
This uses native configuration compliance with cross-account and cross-Region aggregation, minimizing custom code and ongoing operations while providing real-time detection and centralized reporting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses native configuration compliance with cross-account and cross-Region aggregation, minimizing custom code and ongoing operations while providing.
● Scene fit: This uses native configuration compliance with cross-account and cross-Region aggregation, minimizing custom code and ongoing operations while providing.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on custom log parsing and event correlation, which increases operational overhead and is less reliable for comprehensive compliance.
● Requires SSM agents, custom compliance items, and additional setup that is not optimized for EBS encryption checks at organization.
● While workable, maintaining StackSets for all accounts and Regions adds management overhead compared to using an organization-wide Config aggregator.
Workflow: Create an AWS Config managed rule for EBS volume encryption in all accounts and aggregate results with an AWS Config organization aggregator, exporting findings to Amazon S3 and sending notifications via Amazon.
Trusted Advisor low utilization with EventBridge to Lambda, which starts SSM Automation
Trusted Advisor emits check item change events to EventBridge, enabling a lightweight trigger to an SSM Automation runbook with an approval gate.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Trusted Advisor emits check item change events to EventBridge, enabling a lightweight trigger to an SSM Automation runbook with an approval gate.
● Scene fit: Trusted Advisor emits check item change events to EventBridge, enabling a lightweight trigger to an SSM Automation runbook with an approval gate.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Builds custom polling and state management, which is higher effort than using managed events.
● Trusted Advisor does not publish per-check notifications to SNS for direct automation.
● Compute Optimizer provides recommendations but does not emit actionable events for automated termination workflows.
Workflow: Trusted Advisor low utilization with EventBridge to Lambda, which starts SSM Automation with manual approval to terminate.
Keep database credentials in AWS Secrets Manager encrypted by AWS KMS, allow
Secrets Manager is designed for secrets lifecycle with native rotation and integrates directly with ECS for injecting secrets as environment variables.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Secrets Manager is designed for secrets lifecycle with native rotation and integrates directly with ECS for injecting secrets.
● Scene fit: Secrets Manager is designed for secrets lifecycle with native rotation and integrates directly with ECS for injecting secrets.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Secures retrieval, but Parameter Store is not a dedicated secrets service and lacks built-in rotation, so you must build.
● AppConfig targets application configuration rather than secrets and does not provide native secret rotation for database credentials.
● Docker Secrets is tied to Docker Swarm and is not the supported or integrated mechanism for secret delivery in.
Workflow: Keep database credentials in AWS Secrets Manager encrypted by AWS KMS → allow the ECS task execution role to read them, and inject them into the container through the task definition secrets mapping.
an Amazon EventBridge rule with source aws.health and event code AWS_RISK_CREDENTIALS_EXPOSED to
AWS Health emits AWS_RISK_CREDENTIALS_EXPOSED events that can be matched by EventBridge using source aws.health, and Step Functions can orchestrate key deletion, investigation, and notification.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health emits AWS_RISK_CREDENTIALS_EXPOSED events that can be matched by EventBridge using source aws.health, and Step Functions can.
● Scene fit: AWS Health emits AWS_RISK_CREDENTIALS_EXPOSED events that can be matched by EventBridge using source aws.health, and Step Functions can.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS Config does not evaluate or surface AWS Health events such as AWS_RISK_CREDENTIALS_EXPOSED, so it cannot detect leaked access.
● GuardDuty and Macie do not monitor public code hosting for exposed AWS access keys, so they will not emit.
Workflow: Create an Amazon EventBridge rule with source aws.health and event code AWS_RISK_CREDENTIALS_EXPOSED to start an AWS Step Functions state machine that runs three Lambda functions to remove the exposed access key, gather.
two CloudFormation StackSets, one to enable CloudTrail and another to enable AWS
This applies CloudTrail and AWS Config consistently across every account and Region, centralizes compliance via an aggregator, and uses EventBridge to automate alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This applies CloudTrail and AWS Config consistently across every account and Region, centralizes compliance via an aggregator, and.
● Scene fit: This applies CloudTrail and AWS Config consistently across every account and Region, centralizes compliance via an aggregator, and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces a new governance framework that may not be required and does not directly ensure immediate multi-Region CloudTrail enforcement.
● Only enabling AWS Config in a single account misses per-account resource evaluation and does not guarantee compliance checks across.
● Amazon SQS provides queues rather than topics, so this design misuses the service and is not the right eventing.
Workflow: Create two CloudFormation StackSets, one to enable CloudTrail and another to enable AWS Config with a rule that checks for CloudTrail, target all accounts and all Regions → set up an AWS.
Implement a readiness endpoint in the web service that verifies connectivity to
A custom readiness endpoint allows the ALB to mark web tasks healthy only when end-to-end dependencies are reachable.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A custom readiness endpoint allows the ALB to mark web tasks healthy only when end-to-end dependencies are reachable.
● Scene fit: A custom readiness endpoint allows the ALB to mark web tasks healthy only when end-to-end dependencies are reachable.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Route 53 checks and Application Auto Scaling do not validate cross-tier connectivity and cannot terminate or replace RDS instances.
● RDS instances cannot be registered as ALB targets and TCP checks do not validate application-level dependencies.
● ALBs cannot consume CloudWatch alarm states to set target health and these metrics do not confirm dependency reachability.
Workflow: Implement a readiness endpoint in the web service that verifies connectivity to the backend ECS service and the RDS database, and set the target group health check to that path.
AWS CodeDeploy with Auto Scaling, OneAtATime, CloudWatch CPU alarm at 95%, auto
CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to trigger automatic rollback on metric breaches.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to trigger automatic rollback on metric breaches.
● Scene fit: CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to trigger automatic rollback on metric breaches.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Elastic Beanstalk can do rolling updates and use health checks, but it does not natively roll back based on a CPU CloudWatch alarm for a single instance.
● Instance Refresh can progress gradually but does not support automatic rollback driven directly by a CloudWatch CPU alarm.
● Approach is complex and lacks native, metric-based automatic rollback for per-instance deployments.
Workflow: AWS CodeDeploy with Auto Scaling, OneAtATime, CloudWatch CPU alarm at 95%, auto rollback.
Commit .ebextensions/alb.config with option_settings for aws:elbv2:listener:default redirect
Placing redirect rules in .ebextensions option_settings lets Beanstalk apply the ALB listener configuration during deployments driven by the pipeline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Placing redirect rules in .ebextensions option_settings lets Beanstalk apply the ALB listener configuration during deployments driven by the pipeline.
● Scene fit: Placing redirect rules in .ebextensions option_settings lets Beanstalk apply the ALB listener configuration during deployments driven by the pipeline.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Elastic Beanstalk manages the ALB; out-of-band CloudFormation changes are brittle and can be overwritten by Beanstalk updates.
● Requires direct environment permissions and bypasses the pipeline, which does not meet the constraint.
● Imperative instance scripts create drift and are not the recommended or reliable way to manage Beanstalk ALB listener rules.
Workflow: Commit .ebextensions/alb.config with option_settings for aws:elbv2:listener:default redirect.
VPC endpoints: ECR API, ECR DKR, and S3 + Fix the task
ECR auth and image layers require ECR API and DKR interface endpoints plus an S3 gateway endpoint for private pulls. The execution role must allow required ECR (and secrets/KMS if used) actions to obtain tokens and pull images.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ECR auth and image layers require ECR API and DKR interface endpoints plus an S3 gateway endpoint for private pulls.
● Scene fit: ECR auth and image layers require ECR API and DKR interface endpoints plus an S3 gateway endpoint for.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Public IPs restore internet egress but violate the private/no-NAT constraint and are unnecessary if VPC endpoints are used.
● Secrets access alone is insufficient; registry auth and layer downloads still fail without ECR and S3 endpoints.
● Fargate already uses awsvpc by default; this does not address registry/auth connectivity.
Workflow: Create VPC endpoints: ECR API, ECR DKR, and S3 → Fix the task execution IAM role permissions for ECR and Secrets Manager.
Define AWS WAF rate-based and IP set rules to block abusive patterns
Layering AWS WAF rules with tight NACLs helps block malicious requests early and reduces unnecessary exposure at the network edge. Shield Advanced with CloudFront provides enhanced.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Layering AWS WAF rules with tight NACLs helps block malicious requests early and reduces unnecessary exposure at the.
● Scene fit: Layering AWS WAF rules with tight NACLs helps block malicious requests early and reduces unnecessary exposure at the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Focuses on instance management and capacity scaling, which absorbs traffic rather than shrinking the DDoS attack surface or providing.
● These actions improve governance and resilience but do not directly mitigate or reduce the surface area for DDoS attacks.
Workflow: Define AWS WAF rate-based and IP set rules to block abusive patterns and known bad sources, and restrict VPC network ACLs to only required ports and CIDRs → Enable AWS Shield Advanced.
a redrive policy on the standard queue to route failed messages to
A redrive policy sends repeatedly failing messages to a dead-letter queue where you can inspect payloads and debug.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A redrive policy sends repeatedly failing messages to a dead-letter queue where you can inspect payloads and debug.
● Scene fit: A redrive policy sends repeatedly failing messages to a dead-letter queue where you can inspect payloads and debug.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Long polling reduces empty receives and API costs but does not isolate failed messages for inspection.
● FIFO ordering does not address processing failures and switching queue types does not help analyze bad payloads.
● A delivery delay only postpones initial message visibility and does not help capture failed messages for analysis.
Workflow: Configure a redrive policy on the standard queue to route failed messages to a dead-letter queue.
EventBridge rule on CloudTrail UpdateStage; Lambda uses GetSdk to S3
UpdateStage captures stage changes including rollbacks; Lambda can call GetSdk and publish artifacts to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: UpdateStage captures stage changes including rollbacks; Lambda can call GetSdk and publish artifacts to S3.
● Scene fit: UpdateStage captures stage changes including rollbacks; Lambda can call GetSdk and publish artifacts to S3.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs only during pipeline executions and can miss manual or out-of-band stage rollbacks.
● CreateDeployment is not emitted for rollbacks that reuse a prior deployment, so SDKs would not update.
● Repository pushes do not reliably reflect stage rollbacks or direct stage updates in API Gateway.
Workflow: EventBridge rule on CloudTrail UpdateStage → Lambda uses GetSdk to S3.
AWS Step Functions with an EventBridge schedule to launch from each AMI
Coordinates ephemeral instances per AMI, ensures Inspector readiness, scopes by tags, and uses EventBridge for a recurring schedule.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Coordinates ephemeral instances per AMI, ensures Inspector readiness, scopes by tags, and uses EventBridge for a recurring schedule.
● Scene fit: Coordinates ephemeral instances per AMI, ensures Inspector readiness, scopes by tags, and uses EventBridge for a recurring schedule.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Uses CloudWatch Alarms instead of an EventBridge schedule and omits ensuring Inspector is enabled or agents are present.
● Image Builder can run image tests but does not perform Amazon Inspector CVE assessments.
● Lacks tag-based scoping and does not use a scheduled EventBridge rule for reliable cadence.
Workflow: AWS Step Functions with an EventBridge schedule to launch from each AMI, ensure Inspector is enabled via SSM, tag, and run a tag-scoped assessment.
an Amazon EventBridge schedule that invokes an AWS Lambda function to call
Scheduling RefreshCache updates the file gateway’s cached directory listing so newly added S3 objects are visible through the SMB or NFS share.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Scheduling RefreshCache updates the file gateway’s cached directory listing so newly added S3 objects are visible through the SMB or NFS share.
● Scene fit: Scheduling RefreshCache updates the file gateway’s cached directory listing so newly added S3 objects are visible through the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 replication only moves objects between S3 buckets and does not update or notify a file gateway to refresh its directory cache.
● DataSync performs bulk transfers but does not refresh a file gateway’s listing and adds unnecessary complexity for simple cache visibility.
● Volume Gateway provides block storage, not file shares, and S3 event notifications cannot directly instruct Storage Gateway to refresh file share listings.
Workflow: Create an Amazon EventBridge schedule that invokes an AWS Lambda function to call RefreshCache for the file share.
AWS CloudFormation StackSets with service-managed permissions from the management account to deploy
StackSets integrated with Organizations can auto-provision and optionally delete resources per account with minimal ongoing management.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: StackSets integrated with Organizations can auto-provision and optionally delete resources per account with minimal ongoing management.
● Scene fit: StackSets integrated with Organizations can auto-provision and optionally delete resources per account with minimal ongoing management.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Can baseline new accounts but is heavier to adopt in existing organizations and does not natively handle custom IAM.
● Introduces custom glue code with complex cross-account permissions and lacks native automatic cleanup when accounts are removed.
● An organization trail centralizes logging but sharing an IAM role across all accounts is not automatic and still requires.
Workflow: Use AWS CloudFormation StackSets with service-managed permissions from the management account to deploy a CloudTrail trail and required IAM roles to target OUs, enabling automatic deployments to new accounts and automatic stack instance removal when accounts are closed or leave.
Blue green, snapshot and deletion protection, new EB uses existing RDS, remove
Decouples the DB from EB lifecycle, protects data, and clears the security group reference before terminating.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Decouples the DB from EB lifecycle, protects data, and clears the security group reference before terminating.
● Scene fit: Decouples the DB from EB lifecycle, protects data, and clears the security group reference before terminating.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Skipping security group cleanup can leave a dependency and risk unintended deletion or blocked teardown.
● Still couples the DB to EB and adds unnecessary migration steps instead of directly decoupling.
● Elastic Beanstalk does not provide native canary and this ignores the SG dependency between EB and RDS.
Workflow: Blue green, snapshot and deletion protection, new EB uses existing RDS → remove old env SG from DB SG, → terminate.
CloudWatch Logs destination with Kinesis Data Firehose delivering to S3
Use a cross-account CloudWatch Logs destination subscribed to a Kinesis Data Firehose in the central account that writes to S3, minimizing provisioning and scaling needs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use a cross-account CloudWatch Logs destination subscribed to a Kinesis Data Firehose in the central account that writes to S3, minimizing provisioning and scaling needs.
● Scene fit: Use a cross-account CloudWatch Logs destination subscribed to a Kinesis Data Firehose in the central account that writes to S3, minimizing provisioning and scaling needs.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adds Kinesis Data Streams shard provisioning and scaling, increasing operational overhead.
● Batch exports are per log group, require scheduling and maintenance, and are not near real-time or easily scalable across many accounts.
● Requires managing OpenSearch capacity and indices, not minimal for long-term archival.
Workflow: CloudWatch Logs destination with Kinesis Data Firehose delivering to S3.
an Amazon EventBridge scheduled rule that invokes an AWS Lambda function daily
A scheduled EventBridge rule with Lambda can track first-seen dates via tags and reliably enforce deletion after 21 days of detachment.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A scheduled EventBridge rule with Lambda can track first-seen dates via tags and reliably enforce deletion after 21.
● Scene fit: A scheduled EventBridge rule with Lambda can track first-seen dates via tags and reliably enforce deletion after 21.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Trusted Advisor can highlight underutilized or unattached volumes, but it does not provide a 21-day detachment-age signal or native.
● Config flags configuration state but does not orchestrate per-resource 21-day timers, making a separate delayed rule per volume inefficient.
● DLM manages snapshot and AMI lifecycles rather than deleting orphaned EBS volumes based on detachment age.
Workflow: Create an Amazon EventBridge scheduled rule that invokes an AWS Lambda function daily to tag newly discovered detached volumes with the current date and delete any detached volumes whose tag shows they.
EventBridge schedule triggers Step Functions to launch a short-lived EC2 from the
Launch an ephemeral instance from the AMI on a schedule, allow Inspector to assess it for CVEs, and terminate to minimize cost and impact.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Launch an ephemeral instance from the AMI on a schedule, allow Inspector to assess it for CVEs, and terminate to minimize cost and impact.
● Scene fit: Launch an ephemeral instance from the AMI on a schedule, allow Inspector to assess it for CVEs, and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Patch Manager evaluates instance patch compliance, not the AMI artifact, and repeatedly scanning the whole fleet is costly and unnecessary.
● Amazon Inspector cannot scan an AMI by ID; it assesses running EC2 instances or container images.
● Rebuilding on a short cadence is heavy and does not directly validate the existing golden AMI for newly published CVEs.
Workflow: EventBridge schedule triggers Step Functions to launch a short-lived EC2 from the AMI, let Amazon Inspector scan it, → terminate.
Amazon GuardDuty organization-wide with a delegated admin; route findings via EventBridge to
GuardDuty supports AWS Organizations with a delegated administrator for centralized detection and EventBridge can deliver findings to S3 via Firehose.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty supports AWS Organizations with a delegated administrator for centralized detection and EventBridge can deliver findings to S3 via Firehose.
● Scene fit: GuardDuty supports AWS Organizations with a delegated administrator for centralized detection and EventBridge can deliver findings to S3 via Firehose.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Security Hub aggregates findings from other services but does not natively detect threats like SSH brute force or malware.
● Running GuardDuty only in the admin account will not generate findings for member accounts and adds unnecessary pipeline complexity.
● Inspector focuses on vulnerability and software inventory assessments rather than detecting network or account threats like SSH brute force.
Workflow: Enable Amazon GuardDuty organization-wide with a delegated admin → route findings via EventBridge to Kinesis Data Firehose to S3.
Lambda validation functions to AppSpec lifecycle hooks like BeforeAllowTraffic to run against
Pre-traffic AppSpec hooks allow test traffic validation and can fail the deployment to trigger rollback if checks fail. Deployment-group alarms can fail deployments, send SNS notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Pre-traffic AppSpec hooks allow test traffic validation and can fail the deployment to trigger rollback if checks fail.
● Scene fit: Pre-traffic AppSpec hooks allow test traffic validation and can fail the deployment to trigger rollback if checks fail.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch Alarms do not accept arbitrary pass or fail payloads directly from Lambda, so this cannot reliably drive rollback.
● Validating only after the traffic shift risks user impact and does not meet the requirement to test before live.
Workflow: Attach Lambda validation functions to AppSpec lifecycle hooks like BeforeAllowTraffic to run against test traffic and roll back on failures → Associate a CloudWatch alarm with the CodeDeploy deployment group and notify.
Package approved architectures as AWS CloudFormation templates and publish them as products
Service Catalog provides preventative governance by letting users provision only vetted products with enforced parameters and tags.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Service Catalog provides preventative governance by letting users provision only vetted products with enforced parameters and tags.
● Scene fit: Service Catalog provides preventative governance by letting users provision only vetted products with enforced parameters and tags.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Offers detective controls after creation and does not prevent noncompliant resources from being launched.
● IAM policies do not support approval workflows and SNS cannot enforce a pre-create approval gate.
● Direct CloudFormation access cannot reliably restrict which templates or resources are created, weakening compliance enforcement.
Workflow: Package approved architectures as AWS CloudFormation templates and publish them as products in AWS Service Catalog with required tags, → allow beginners to launch only Service Catalog products and deny write access to other services.
a single GitHub or GitLab repository with a develop branch for merges
A single-repo branching strategy plus CodeBuild on commit and CodeDeploy blue/green provides near-zero downtime and enables fast, safe rollbacks by shifting traffic between environments.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A single-repo branching strategy plus CodeBuild on commit and CodeDeploy blue/green provides near-zero downtime and enables fast, safe.
● Scene fit: A single-repo branching strategy plus CodeBuild on commit and CodeDeploy blue/green provides near-zero downtime and enables fast, safe.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Multiple repositories per developer add unnecessary coordination and complexity without improving rollback or uptime compared to a single shared.
● Amazon ECR is a container image registry and is not a suitable system for hosting application source code.
Workflow: Use a single GitHub or GitLab repository with a develop branch for merges → trigger AWS CodeBuild on each commit to that branch, require pull requests into main, and deploy to production.
Associate the function with a CloudFront distribution using Lambda@Edge
Lambda@Edge runs your code at CloudFront edge locations on viewer and origin events, enabling User-Agent based selection with globally low latency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda@Edge runs your code at CloudFront edge locations on viewer and origin events, enabling User-Agent based selection with globally low latency.
● Scene fit: Lambda@Edge runs your code at CloudFront edge locations on viewer and origin events, enabling User-Agent based selection with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Edge-optimized APIs route through CloudFront to a regional endpoint and do not execute your code at the edge, so latency and per-request customization at the CDN layer are limited.
● CloudFront Functions are designed for lightweight header or URL rewrites and cannot run existing Lambda code or perform heavier logic often needed for dynamic image selection.
● CloudFront requires an HTTP origin such as S3, an Application Load Balancer, or a custom server, and Lambda is not a valid origin type.
Workflow: Associate the function with a CloudFront distribution using Lambda@Edge.
Route 53 geoproximity routing with bias to ALBs across four Regions
Routes by proximity and supports bias to shift more or less traffic to chosen Regions for uneven demand.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Routes by proximity and supports bias to shift more or less traffic to chosen Regions for uneven demand.
● Scene fit: Routes by proximity and supports bias to shift more or less traffic to chosen Regions for uneven demand.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides lowest-latency routing but lacks controls to skew traffic toward or away from Regions.
● Routes by user location but offers no bias control for traffic shaping across Regions.
● Returns multiple records for simple load and health but does not provide regional bias or per-user latency optimization.
Workflow: Route 53 geoproximity routing with bias to ALBs across four Regions.
Define a BeforeAllowTraffic hook in the AppSpec that calls a validation Lambda
BeforeAllowTraffic runs before traffic shifting and is designed to block the shift until a validation Lambda signals readiness.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: BeforeAllowTraffic runs before traffic shifting and is designed to block the shift until a validation Lambda signals readiness.
● Scene fit: BeforeAllowTraffic runs before traffic shifting and is designed to block the shift until a validation Lambda signals readiness.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● AfterAllowTraffic runs only after traffic has already been shifted, so it cannot gate pre-traffic readiness.
● CloudWatch alarms can trigger rollback during or after the shift but do not provide a pre-traffic readiness gate for Lambda deployments.
● ValidateService is not a Lambda compute platform hook and does not act as a pre-traffic gate for Lambda deployments.
Workflow: Define a BeforeAllowTraffic hook in the AppSpec that calls a validation Lambda which returns Succeeded only when the live-v2 API Gateway stage responds to requests.
Amazon CloudFront origin failover with a primary and secondary origin for the
CloudFront origin failover can automatically switch to a secondary origin on specific 5xx responses such as 504, improving availability at low cost. Lambda@Edge executes close to.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFront origin failover can automatically switch to a secondary origin on specific 5xx responses such as 504, improving.
● Scene fit: CloudFront origin failover can automatically switch to a secondary origin on specific 5xx responses such as 504, improving.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While this can reduce latency, it is expensive and complex to operate and synchronize across Regions.
● Longer TTLs help with static content but do not accelerate dynamic login flows or prevent 504 errors.
● Global Accelerator may reduce network path latency but adds cost and does not address origin 504s or heavy authentication processing.
Workflow: Configure Amazon CloudFront origin failover with a primary and secondary origin for the login endpoint → Move authentication logic to Lambda@Edge so it runs in edge locations nearest to viewers.
EventBridge schedule targeting EC2 CreateSnapshot
Directly schedules the EC2 CreateSnapshot API for specific volume IDs at a fixed time without extra code or services.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Directly schedules the EC2 CreateSnapshot API for specific volume IDs at a fixed time without extra code or services.
● Scene fit: Directly schedules the EC2 CreateSnapshot API for specific volume IDs at a fixed time without extra code or services.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Automates EBS snapshot schedules and retention via policies, usually tag-based, which adds more setup than a single scheduled API call.
● Provides centralized backups with plans and vaults but is heavier to configure than a single scheduled rule.
● Adds an extra Systems Manager step and dependency on the SSM agent, which is unnecessary for a direct snapshot call.
Workflow: EventBridge schedule targeting EC2 CreateSnapshot.
a reader and switch to cluster writer and reader endpoints
Adding a reader enables failover and using the cluster writer and reader endpoints allows writes to fail over and reads to continue with minimal disruption.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Adding a reader enables failover and using the cluster writer and reader endpoints allows writes to fail over and reads to continue with minimal disruption.
● Scene fit: Adding a reader enables failover and using the cluster writer and reader endpoints allows writes to fail over.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● RDS Proxy pools connections but cannot prevent disconnects when a single writer instance is rebooted during maintenance.
● Serverless v2 helps with scaling but does not eliminate engine restarts during maintenance or provide read/write split by itself.
● Custom endpoints are for grouping instances and are unnecessary for basic read/write splitting, which is handled by the default cluster and reader endpoints.
Workflow: Add a reader and switch to cluster writer and reader endpoints.
Write the current AMI ID to SSM Parameter Store and reference it
Storing the AMI ID in Parameter Store and resolving it with a dynamic reference decouples teams and ensures stacks pick up the latest approved image at deploy time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Storing the AMI ID in Parameter Store and resolving it with a dynamic reference decouples teams and ensures stacks pick up the latest approved image at deploy time.
● Scene fit: Storing the AMI ID in Parameter Store and resolving it with a dynamic reference decouples teams and ensures.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFormation cannot natively resolve AMIs by tag without a custom resource or external logic, making this brittle.
● Using S3 artifacts and cross-stack exports adds operational overhead and lacks a first-class, secure parameter retrieval mechanism.
● Service Catalog and Launch Templates do not auto-resolve the latest AMI for CloudFormation without updates, and this increases coupling.
Workflow: Write the current AMI ID to SSM Parameter Store and reference it via CloudFormation SSM dynamic reference.
a post-deploy CodePipeline stage that triggers a Lambda to export the SDK
Automates export and cache refresh exactly when the deployment completes, ensuring clients get the newest SDK.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Automates export and cache refresh exactly when the deployment completes, ensuring clients get the newest SDK.
● Scene fit: Automates export and cache refresh exactly when the deployment completes, ensuring clients get the newest SDK.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Ties to deployments but does not invalidate CloudFront, so clients may receive cached SDK and not the latest immediately.
● Short TTLs reduce staleness but do not guarantee immediate freshness after deployments and increase origin load.
● S3 has no cache invalidation API; CloudFront invalidation must be used to refresh content.
Workflow: Add a post-deploy CodePipeline stage that triggers a Lambda to export the SDK from API Gateway, upload to S3, and create a CloudFront invalidation.
Install the AWS Systems Manager agent on all EC2 and on-premises servers
Patch Manager provides centralized hybrid patching with baselines, groups, and maintenance windows to automate consistent OS and application updates at scale.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Patch Manager provides centralized hybrid patching with baselines, groups, and maintenance windows to automate consistent OS and application updates at scale.
● Scene fit: Patch Manager provides centralized hybrid patching with baselines, groups, and maintenance windows to automate consistent OS and application updates at scale.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Approach is manual and relies on stored credentials, lacking managed patch baselines, scheduling, and built-in compliance reporting.
● Inspector surfaces vulnerabilities and prioritizes findings but does not perform operating system patch orchestration or maintenance-window scheduling.
● Custom scripting increases operational burden, is error-prone, and lacks centralized governance and standardized compliance controls.
Workflow: Install the AWS Systems Manager agent on all EC2 and on-premises servers → use Patch Manager with patch baselines and patch groups, and schedule updates during maintenance windows.
Amazon Macie
This discovers and classifies sensitive data in S3 and raises findings for potential public exposure or unusual access to PII.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This discovers and classifies sensitive data in S3 and raises findings for potential public exposure or unusual access to PII.
● Scene fit: This discovers and classifies sensitive data in S3 and raises findings for potential public exposure or unusual access to PII.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Detects malicious or anomalous behavior and includes S3 protection, but it does not classify or evaluate PII exposure against compliance needs in S3.
● Aggregates and prioritizes findings from other security services and cannot directly discover PII or assess S3 data exposure by itself.
● Focuses on vulnerability assessments for compute and container resources rather than S3 data classification or PII monitoring.
Workflow: Amazon Macie.
Console copy used multipart; checksum is aggregated from part checksums, not whole
Large console copies often use multipart, producing a checksum-of-parts that differs from a single-part whole-object checksum.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Large console copies often use multipart, producing a checksum-of-parts that differs from a single-part whole-object checksum.
● Scene fit: Large console copies often use multipart, producing a checksum-of-parts that differs from a single-part whole-object checksum.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Server-side encryption does not change checksum semantics in this way; checksums are not computed over ciphertext for SSE-S3 or SSE-KMS.
● S3 checksum fields are independent of ETag; checksums are not derived from ETag values.
● There is no universal threshold that invalidates a valid single-part checksum; the mismatch isn't due to an incorrect original checksum.
Workflow: Console copy used multipart → checksum is aggregated from part checksums, not whole object.
an Amazon Cognito identity pool for authenticated and guest users, use the
Cognito identity pools broker social or OIDC identities into STS temporary credentials mapped to IAM roles, enabling secure direct access to S3 and DynamoDB. AssumeRoleWithWebIdentity exchanges.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cognito identity pools broker social or OIDC identities into STS temporary credentials mapped to IAM roles, enabling secure.
● Scene fit: Cognito identity pools broker social or OIDC identities into STS temporary credentials mapped to IAM roles, enabling secure.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● SAML is for SAML 2.0 enterprise identity providers and is not the correct mechanism for social or OIDC logins.
● Embedding long-term access keys in a mobile app is insecure and violates best practices that require temporary credentials.
Workflow: Create an Amazon Cognito identity pool for authenticated and guest users → use the React Native AWS SDK to retrieve temporary AWS credentials, and assign IAM roles scoped to S3 and DynamoDB.
EventBridge rules: S3 PUT -> ECS RunTask; S3 DELETE -> Lambda StopTask
EventBridge can start an ad hoc Fargate task with RunTask for each S3 upload, avoiding a continuously running service.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can start an ad hoc Fargate task with RunTask for each S3 upload, avoiding a continuously running service.
● Scene fit: EventBridge can start an ad hoc Fargate task with RunTask for each S3 upload, avoiding a continuously running service.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch alarms are threshold based and CloudTrail is an indirect source for per-object processing.
● SQS is workable for buffering, but an ECS service remains a more persistent operating model.
● Capacity providers control how ECS obtains capacity, not when individual Fargate tasks start or stop.
Workflow: EventBridge rules: S3 PUT -> ECS RunTask → S3 DELETE -> Lambda StopTask.
a CloudWatch Logs metric filter to emit a custom metric, build a
Metric filters convert matching log events into metrics that can be alarmed and routed to SNS, which is the most direct and scalable approach for log-derived.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric filters convert matching log events into metrics that can be alarmed and routed to SNS, which is.
● Scene fit: Metric filters convert matching log events into metrics that can be alarmed and routed to SNS, which is.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● X-Ray focuses on tracing and service maps rather than counting log-line error occurrences from application logs, so it does.
● Embedding error counting and SetAlarmState in the app is brittle across multiple instances and misuses CloudWatch alarms, making it.
● Metric filters produce metrics and cannot directly target an alarm without first creating a corresponding CloudWatch metric.
Workflow: Create a CloudWatch Logs metric filter to emit a custom metric, build a CloudWatch alarm on that metric with a 10-minute evaluation, and send notifications through SNS email.
CloudFront with a response headers policy
Attach a response headers policy to the distribution to inject security headers on viewer responses without custom code.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Attach a response headers policy to the distribution to inject security headers on viewer responses without custom code.
● Scene fit: Attach a response headers policy to the distribution to inject security headers on viewer responses without custom code.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● These send headers from CloudFront to the origin on requests and do not set headers on viewer responses.
● Bucket policies control access and cannot modify HTTP response headers.
● Possible but adds code and operational overhead compared to using a native response headers policy.
Workflow: CloudFront with a response headers policy.
AWS Transit Gateway for transitive connectivity among VPCs and manage network access
Transit Gateway natively supports scalable transitive routing and Firewall Manager provides centralized policy enforcement for network controls across accounts and VPCs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Transit Gateway natively supports scalable transitive routing and Firewall Manager provides centralized policy enforcement for network controls across accounts and VPCs.
● Scene fit: Transit Gateway natively supports scalable transitive routing and Firewall Manager provides centralized policy enforcement for network controls across.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● VPC peering is not transitive and AWS WAF focuses on Layer 7 web traffic, not centralized network access controls across VPCs.
● PrivateLink enables private service access rather than full-mesh routing, and Security Hub aggregates findings but does not enforce network policies.
● AWS Site-to-Site VPN is designed for on-premises to AWS connectivity and building many VPC-to-VPC tunnels is operationally complex and inefficient.
Workflow: Use AWS Transit Gateway for transitive connectivity among VPCs and manage network access policies centrally with AWS Firewall Manager.
CloudFormation custom resource (Lambda) to fetch latest AMI
Resolves the AMI ID at stack create/update without editing templates.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Resolves the AMI ID at stack create/update without editing templates.
● Scene fit: Resolves the AMI ID at stack create/update without editing templates.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works only for vendor-published AMIs, not custom app images.
● Adds unnecessary scheduling, template churn, and cost.
● Cannot alter the AMI already used to launch the instance.
Workflow: CloudFormation custom resource (Lambda) to fetch latest AMI.
AWS Config with a custom rule and Lambda
Continuously records EC2 and Dedicated Host relationships and evaluates compliance via a custom rule, with built-in compliance summaries.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously records EC2 and Dedicated Host relationships and evaluates compliance via a custom rule, with built-in compliance summaries.
● Scene fit: Continuously records EC2 and Dedicated Host relationships and evaluates compliance via a custom rule, with built-in compliance summaries.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Tracks and enforces licenses but does not evaluate EC2 Dedicated Host placement for compliance.
● Possible but high maintenance and event-driven, not ideal for ongoing configuration compliance reporting.
● Focuses on patch and configuration baselines, not Dedicated Host placement validation.
Workflow: AWS Config with a custom rule and Lambda.
a Lambda that copies the AMI to each Region and stores each
Copies the AMI per Region and records the Region-specific AMI ID for programmatic lookup. Provides a consistent, repeatable workflow to patch, harden, and bake AMIs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Copies the AMI per Region and records the Region-specific AMI ID for programmatic lookup.
● Scene fit: Copies the AMI per Region and records the Region-specific AMI ID for programmatic lookup. Provides a consistent, repeatable.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Backups and snapshots are handled, but it does not create or register AMIs for launches.
● Configures instance state but does not build or copy AMIs.
● Copying a string does not distribute the AMI; AMI IDs are Region-scoped and require AMI copy.
Workflow: Create a Lambda that copies the AMI to each Region and stores each AMI ID in Parameter Store under a common key → Author an AWS Systems Manager Automation runbook to build and harden the AMI.
The additional Availability Zone is not enabled on the Application Load Balancer
ALB must have each Availability Zone explicitly enabled so it creates a load balancer node in that zone and can route traffic there.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB must have each Availability Zone explicitly enabled so it creates a load balancer node in that zone.
● Scene fit: ALB must have each Availability Zone explicitly enabled so it creates a load balancer node in that zone.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Disabling cross-zone load balancing affects distribution across zones but does not prevent routing to an Availability Zone that is enabled on the ALB.
● The Auto Scaling group is launching instances there, which implies subnets exist; the issue is more likely the ALB not enabling that zone.
● Failed health checks would mark targets unhealthy, but the symptom of zero traffic after expanding to a new zone most commonly indicates the ALB zone was never enabled.
Workflow: The additional Availability Zone is not enabled on the Application Load Balancer.
AWS Config required-tags for AWS::EC2::Volume with SSM Automation remediation to add BackupInterval=15d
Continuously evaluates EBS volumes and automatically adds the missing tag to noncompliant resources.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously evaluates EBS volumes and automatically adds the missing tag to noncompliant resources.
● Scene fit: Continuously evaluates EBS volumes and automatically adds the missing tag to noncompliant resources.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Blocks untagged creation but does not auto-tag existing resources or fix tag drift.
● Tag policies help standardize and report but do not enforce at the API or auto-remediate.
● Only handles creation-time events and will not fix existing or later-untagged volumes.
Workflow: Configure AWS Config required-tags for AWS::EC2::Volume with SSM Automation remediation to add BackupInterval=15d.
the application on AWS Elastic Beanstalk with load balancing and Auto Scaling
Elastic Beanstalk offers a managed Node.js platform with integrated Application Load Balancer and Auto Scaling, reducing operational burden. RDS is a managed relational database service, and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Elastic Beanstalk offers a managed Node.js platform with integrated Application Load Balancer and Auto Scaling, reducing operational burden.
● Scene fit: Elastic Beanstalk offers a managed Node.js platform with integrated Application Load Balancer and Auto Scaling, reducing operational burden.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach does not provide a managed web application platform or proper load balancing for a stateful web service and is unnecessary for Lambda.
● Still requires managing servers, patching, and capacity details, which conflicts with the goal of minimizing operations.
● DynamoDB is a NoSQL service and does not satisfy the requirement for a managed relational database.
Workflow: Deploy the application on AWS Elastic Beanstalk with load balancing and Auto Scaling enabled → Provision a standalone Amazon RDS instance in a VPC with automated backups and deletion protection enabled.
a custom AWS Config rule that evaluates S3 bucket policies for public
This uses AWS Config to continuously evaluate bucket ACLs and policies and notify via SNS when a bucket is noncompliant. EventBridge integrates with Trusted Advisor check.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses AWS Config to continuously evaluate bucket ACLs and policies and notify via SNS when a bucket.
● Scene fit: This uses AWS Config to continuously evaluate bucket ACLs and policies and notify via SNS when a bucket.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Inspector does not assess S3 bucket policies, so this would not identify or remediate public buckets.
● Account-level block public access would disable intended public buckets, which violates the requirement to keep some buckets public.
● Polling Trusted Advisor with Lambda is unnecessary and slow compared to EventBridge events, and digest emails are not real-time.
Workflow: Configure a custom AWS Config rule that evaluates S3 bucket policies for public access and publishes noncompliance notifications to an Amazon SNS topic → Use Amazon EventBridge to capture Trusted Advisor S3.
Inspect VPC Flow Logs for REJECTs; confirm SG allows outbound HTTPS
Flow Logs with REJECT entries directly reveal blocked traffic and validating outbound SG rules for TCP 443 targets the likely egress issue after moving to HTTPS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Flow Logs with REJECT entries directly reveal blocked traffic and validating outbound SG rules for TCP 443 targets the likely egress issue after moving to HTTPS.
● Scene fit: Flow Logs with REJECT entries directly reveal blocked traffic and validating outbound SG rules for TCP 443 targets.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Application logs may not reveal network egress denials, and default NACLs typically allow all; this is less targeted for pinpointing network blocks.
● The instances initiate the connection, so ingress rules on instances are not the issue; ACCEPT entries do not highlight denials.
● Analyzes paths between VPC resources and does not test reachability to external internet endpoints or diagnose SG egress denials.
Workflow: Inspect VPC Flow Logs for REJECTs → confirm SG allows outbound HTTPS.
CodeGuru Reviewer Secrets Detector on PRs and store DB credentials in AWS
CodeGuru Reviewer detects hardcoded secrets in PRs and can block merges via status checks, while Secrets Manager provides managed credential rotation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeGuru Reviewer detects hardcoded secrets in PRs and can block merges via status checks, while Secrets Manager provides managed credential rotation.
● Scene fit: CodeGuru Reviewer detects hardcoded secrets in PRs and can block merges via status checks, while Secrets Manager provides managed credential rotation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Macie analyzes S3 data, not GitHub PRs, and environment variables do not provide PR blocking or automatic rotation.
● CodeGuru Profiler is for performance profiling and Parameter Store lacks native automatic rotation and PR gating.
● Build-time checks do not enforce PR-time secret blocking and Parameter Store does not handle automatic rotation.
Workflow: Use CodeGuru Reviewer Secrets Detector on PRs and store DB credentials in AWS Secrets Manager with rotation, updating code to fetch the secret.
Install the Amazon CloudWatch agent with the procstat plugin on each instance
Procstat provides near-real-time process metrics, and a CloudWatch alarm invoking Systems Manager Run Command restarts only the failed process quickly without affecting Auto Scaling capacity.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Procstat provides near-real-time process metrics, and a CloudWatch alarm invoking Systems Manager Run Command restarts only the failed.
● Scene fit: Procstat provides near-real-time process metrics, and a CloudWatch alarm invoking Systems Manager Run Command restarts only the failed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Cycling instances via Standby and reboot is slow, impacts capacity during the cycle, and restarts entire instances rather than.
● Marking instances Unhealthy triggers termination and replacement, which is slower and temporarily reduces capacity compared to restarting only the.
Workflow: Install the Amazon CloudWatch agent with the procstat plugin on each instance to publish per-process metrics and use a CloudWatch alarm to invoke AWS Systems Manager Run Command to restart the worker.
an Amazon EventBridge rule that filters for restricted-ssh evaluations with complianceType set
This targets only the relevant NON_COMPLIANT events from the restricted-ssh rule and uses an input transformer to include the required identifiers before sending to SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This targets only the relevant NON_COMPLIANT events from the restricted-ssh rule and uses an input transformer to include.
● Scene fit: This targets only the relevant NON_COMPLIANT events from the restricted-ssh rule and uses an input transformer to include.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Is inefficient and relies on an SNS topic-level filter policy which is not supported since message filtering is configured.
● Is overly broad and again incorrectly assumes filtering at the SNS topic level rather than on subscriptions.
● ERROR indicates rule evaluation problems rather than a security group being noncompliant, so it does not satisfy the alerting.
Workflow: Create an Amazon EventBridge rule that filters for restricted-ssh evaluations with complianceType set to NON_COMPLIANT → add an input transformer to include the security group name and ID, and publish to the.
Build a CodePipeline with stages for prod, QA, and sandbox, and configure
CloudFormation action parameter overrides and Mappings allow one reusable template while each stage passes its own environment parameters.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation action parameter overrides and Mappings allow one reusable template while each stage passes its own environment parameters.
● Scene fit: CloudFormation action parameter overrides and Mappings allow one reusable template while each stage passes its own environment parameters.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Tightly couples the template to pipeline state and adds unnecessary operational complexity compared to native parameterization.
● Modifying UserData after instances are created is brittle and works against CloudFormation's declarative model.
● Still relies on manual coordination and does not leverage CodePipeline to supply environment values automatically.
Workflow: Build a CodePipeline with stages for prod, QA, and sandbox, and configure CloudFormation actions with parameter overrides → use CloudFormation Mappings and EC2 UserData to set environment-specific values.
Place the JSON files in an S3 bucket with versioning enabled and
S3 versioning provides audit and rollback, while EventBridge and Lambda enable automated, scalable, and restart-free refreshes using managed services.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 versioning provides audit and rollback, while EventBridge and Lambda enable automated, scalable, and restart-free refreshes using managed.
● Scene fit: S3 versioning provides audit and rollback, while EventBridge and Lambda enable automated, scalable, and restart-free refreshes using managed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach requires image rebuilds and task restarts, which prevents dynamic updates and increases operational toil.
● Although it supports versioning, Parameter Store has per-parameter size limits and becomes costly and harder to manage at large.
● Secrets Manager is optimized for sensitive credentials, not large and growing general configuration sets, making it a poor fit.
Workflow: Place the JSON files in an S3 bucket with versioning enabled and use Amazon EventBridge to invoke an AWS Lambda function on a schedule that refreshes the configuration inside running ECS tasks as needed.
CloudWatch cross-account observability with AWS Organizations to designate the central ops account
This natively unifies metrics, logs, and traces across accounts and automatically includes new organization accounts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This natively unifies metrics, logs, and traces across accounts and automatically includes new organization accounts.
● Scene fit: This natively unifies metrics, logs, and traces across accounts and automatically includes new organization accounts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides cross-account visibility but requires manual connections and will not automatically enroll newly created organization accounts.
● Streams only metrics and does not include logs or X-Ray traces, nor does it auto-onboard new accounts.
● EventBridge cannot deliver directly to S3 and this design does not aggregate X-Ray traces or auto-onboard new accounts.
Workflow: Use CloudWatch cross-account observability with AWS Organizations to designate the central ops account as the monitoring account and link all organization accounts.
Secrets Manager performs an initial rotation immediately after rotation is enabled, invalidating
When rotation is turned on, Secrets Manager triggers an initial rotation which can invalidate previously used credentials if the application was not updated to fetch the.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: When rotation is turned on, Secrets Manager triggers an initial rotation which can invalidate previously used credentials if.
● Scene fit: When rotation is turned on, Secrets Manager triggers an initial rotation which can invalidate previously used credentials if.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● By default, Secrets Manager returns the AWSCURRENT version when no version is specified, so omitting the version typically does.
● Fetching the secret value requires secretsmanager:GetSecretValue, and lacking DescribeSecret alone would not block retrieving the password if GetSecretValue is.
● If the initial release worked and networking did not change, a missing VPC endpoint is unlikely, and Secrets Manager.
Workflow: Secrets Manager performs an initial rotation immediately after rotation is enabled, invalidating the old password while the application still uses stale credentials.
AWS Application Discovery Service by deploying the agentless Discovery Connector (OVA) to
Application Discovery Service provides agentless collection for vCenter-managed VMs and augments EC2 with the agent, and Migration Hub offers built-in dashboards with minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Application Discovery Service provides agentless collection for vCenter-managed VMs and augments EC2 with the agent, and Migration Hub offers built-in dashboards with minimal setup.
● Scene fit: Application Discovery Service provides agentless collection for vCenter-managed VMs and augments EC2 with the agent, and Migration Hub.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Can work but requires deploying and managing an agent on every server, which is significant effort at this scale.
● AWS Config tracks AWS resource configurations and does not collect host-level OS, MAC, or IP details for on-premises VMware VMs.
● Custom ingestion requires substantial engineering, deployment, and maintenance compared to managed discovery tools.
Workflow: Use AWS Application Discovery Service by deploying the agentless Discovery Connector (OVA) to VMware vCenter and installing the Discovery Agent on the EC2 instances, → visualize in AWS Migration Hub.
lambda.amazonaws.com as a trusted principal in the IAM role used by the
The Lambda execution role must trust lambda.amazonaws.com so that the service can assume the role and run the function to forward logs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The Lambda execution role must trust lambda.amazonaws.com so that the service can assume the role and run the function to forward logs.
● Scene fit: The Lambda execution role must trust lambda.amazonaws.com so that the service can assume the role and run the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Using logs.amazonaws.com as the trusted principal is incorrect because the Lambda execution role must trust lambda.amazonaws.com, not CloudWatch Logs.
● An export task is batch-oriented and does not fix a broken real-time subscription filter to OpenSearch.
● You cannot attach IAM policies to an OpenSearch domain and this policy belongs on the Lambda function to enable VPC networking.
Workflow: Add lambda.amazonaws.com as a trusted principal in the IAM role used by the Lambda function created from the CloudWatch Logs subscription.
health checks in the ValidateService hook in appspec.yml and enable automatic rollback
ValidateService runs after the service is started and receiving traffic; failing here triggers CodeDeploy automatic rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ValidateService runs after the service is started and receiving traffic; failing here triggers CodeDeploy automatic rollback.
● Scene fit: ValidateService runs after the service is started and receiving traffic; failing here triggers CodeDeploy automatic rollback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● ALB health status alone does not instruct CodeDeploy to roll back unless tied to a failing lifecycle hook or alarm.
● Alarms can roll back deployments but this does not answer which lifecycle hook should run runtime health checks.
● Duplicates native CodeDeploy rollback and adds unnecessary complexity.
Workflow: Run health checks in the ValidateService hook in appspec.yml and enable automatic rollback on failures.
Readdress the overlapping CIDRs, deploy a centralized AWS Transit Gateway shared with
Renumbering removes the overlap and AWS Transit Gateway provides a scalable hub-and-spoke for multi-account VPC connectivity with private, routable paths.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Renumbering removes the overlap and AWS Transit Gateway provides a scalable hub-and-spoke for multi-account VPC connectivity with private.
● Scene fit: Renumbering removes the overlap and AWS Transit Gateway provides a scalable hub-and-spoke for multi-account VPC connectivity with private.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● PrivateLink keeps traffic private but becomes operationally heavy at scale due to per-service, per-VPC endpoints and does not provide.
● A peering mesh does not scale (N×(N−1)/2 links) and cannot connect VPCs with overlapping CIDRs.
● A service mesh manages service-to-service traffic policies but still requires underlying network connectivity and does not solve multi-VPC routing.
Workflow: Readdress the overlapping CIDRs → deploy a centralized AWS Transit Gateway shared with AWS RAM → attach each VPC and add routes to their CIDRs through the attachments, and front services with internal NLBs for private HTTPS.
CloudWatch Logs metric filter to a custom metric, CloudWatch alarm (15 min)
Metric filters convert matching log lines into a metric that a CloudWatch alarm can evaluate over 15 minutes and notify via SNS email.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric filters convert matching log lines into a metric that a CloudWatch alarm can evaluate over 15 minutes and notify via SNS email.
● Scene fit: Metric filters convert matching log lines into a metric that a CloudWatch alarm can evaluate over 15 minutes.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A custom Lambda counter is possible but adds operational overhead and state management for windowing, which is unnecessary for this use case.
● X-Ray is for tracing and latency analysis, not for counting log error occurrences and alerting from CloudWatch Logs.
● Embedding alarm logic in the application is brittle across instances and misuses CloudWatch alarms; it is not efficient or scalable.
Workflow: CloudWatch Logs metric filter to a custom metric, CloudWatch alarm (15 min), and SNS email.
a CloudWatch alarm for Lambda errors to the CodeDeploy deployment + LambdaCanary10Percent10Minutes
CodeDeploy can monitor CloudWatch alarms and automatically stop and roll back when thresholds are breached. Routes 10% of traffic to the new version for 10 minutes, then shifts the remainder.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy can monitor CloudWatch alarms and automatically stop and roll back when thresholds are breached.
● Scene fit: CodeDeploy can monitor CloudWatch alarms and automatically stop and roll back when thresholds are breached. Routes 10% of traffic to the new version for 10 minutes, then shifts the remainder.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● All traffic shifts immediately with no canary window.
● API Gateway canary does not integrate with CodeDeploy for Lambda rollbacks and only affects API Gateway traffic.
● Shifts traffic in repeated increments, not a single two-step canary.
Workflow: Attach a CloudWatch alarm for Lambda errors to the CodeDeploy deployment → LambdaCanary10Percent10Minutes.
Target account: IAM role trusted by the pipeline role from source; allow
Create a cross-account IAM role in the target that trusts the source account's pipeline role so CodePipeline can assume it. CloudFormation in the target account needs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Create a cross-account IAM role in the target that trusts the source account's pipeline role so CodePipeline can.
● Scene fit: Create a cross-account IAM role in the target that trusts the source account's pipeline role so CodePipeline can.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS RAM does not support sharing S3 buckets or KMS keys across accounts for CodePipeline artifacts.
● The CloudFormation execution role must exist in the target account where resources are created.
● SCPs set guardrails and cannot grant permissions or establish trust for cross-account actions.
Workflow: Target account: IAM role trusted by the pipeline role from source → allow AssumeRole → Target account: CloudFormation execution role for the stack → pipeline uses that role → Source account: KMS.
Make sure the Lambda function posts a success or failure result to
Custom resource providers must PUT a response to the pre-signed ResponseURL so CloudFormation can transition the stack state.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Custom resource providers must PUT a response to the pre-signed ResponseURL so CloudFormation can transition the stack state.
● Scene fit: Custom resource providers must PUT a response to the pre-signed ResponseURL so CloudFormation can transition the stack state.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A clean Lambda exit does not signal CloudFormation; custom resources require an explicit response to the provided URL.
● UpdateStack permission is unrelated to completing custom resource signaling and is not required here.
● While this permission can be needed to create the connector, it does not resolve the pending stack state without a custom resource response.
Workflow: Make sure the Lambda function posts a success or failure result to the pre-signed CloudFormation response URL.
The console used a multipart copy for the large object and generated
For console operations on objects over 16 MB, S3 often performs a multipart copy and computes a checksum-of-checksums rather than a single-part checksum.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: For console operations on objects over 16 MB, S3 often performs a multipart copy and computes a checksum-of-checksums rather than a single-part checksum.
● Scene fit: For console operations on objects over 16 MB, S3 often performs a multipart copy and computes a checksum-of-checksums.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Is wrong because the console-triggered multipart behavior threshold in this context is not 64 MB and a single-part upload provides a whole-object checksum.
● S3 does not change the checksum algorithm just because metadata was edited.
● Server-side encryption does not change the semantics of object checksums in this way and does not account for this discrepancy.
Workflow: The console used a multipart copy for the large object and generated a new checksum that aggregates the part checksums instead of representing a full-object checksum.
AWS Compute Optimizer EC2 recommendations
Uses machine learning on historical metrics to recommend optimal EC2 instance types and sizes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Uses machine learning on historical metrics to recommend optimal EC2 instance types and sizes.
● Scene fit: Uses machine learning on historical metrics to recommend optimal EC2 instance types and sizes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Enables monitoring and automation via thresholds but does not deliver ML-driven rightsizing across instance types.
● Offers cost-focused rightsizing suggestions but is less comprehensive than Compute Optimizer.
● Provides general checks including idle resources but not detailed ML-based rightsizing.
Workflow: AWS Compute Optimizer EC2 recommendations.
ELB health checks in the Auto Scaling group for the app's port
ASG ELB health checks ensure scaling decisions follow the load balancer's real application probe on the correct port and path. Target group health checks must probe.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ASG ELB health checks ensure scaling decisions follow the load balancer's real application probe on the correct port and path.
● Scene fit: ASG ELB health checks ensure scaling decisions follow the load balancer's real application probe on the correct port.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Switching to IP targets does not fix reachability caused by misaligned health checks or listener-to-target configuration.
● ALB supports only HTTP or HTTPS listeners; TCP is for NLB.
● If health checks are already succeeding, security group allowances are sufficient; the issue lies with health check configuration or mapping, not SGs.
Workflow: Use ELB health checks in the Auto Scaling group for the app's port and path → Set the target group health check to the app's custom port and health endpoint.
an AWS WAF IP set allowlist that includes the development team’s public
An IP set allowlist in WAF lets only approved source addresses from the blocked region reach the application while other traffic is filtered. A WAF geo.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An IP set allowlist in WAF lets only approved source addresses from the blocked region reach the application.
● Scene fit: An IP set allowlist in WAF lets only approved source addresses from the blocked region reach the application.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Improves DDoS resilience but does not implement country-based blocking or an allowlist for the team.
● ALB listener rules cannot evaluate country information and geo matching is enforced by AWS WAF or CloudFront, not the.
● NACLs are stateless and operate on IP addresses and ports only, so they cannot block by country and would.
Workflow: Use an AWS WAF IP set allowlist that includes the development team’s public IP addresses → Use an AWS WAF geo match rule to block requests from the specified countries.
During rebalancing, EC2 Auto Scaling may briefly go above the group maximum
Rebalancing launches replacements before terminating others, which can temporarily exceed MaxSize by up to 10% or one instance to maintain availability. When scaling in, Auto Scaling.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Rebalancing launches replacements before terminating others, which can temporarily exceed MaxSize by up to 10% or one instance.
● Scene fit: Rebalancing launches replacements before terminating others, which can temporarily exceed MaxSize by up to 10% or one instance.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Auto Scaling prioritizes balancing across Availability Zones over instance age, so age-based termination alone does not cause an AZ.
● The default termination policy evaluates AZ rebalancing before considering billing hour, so it would not preferentially remove an instance.
Workflow: During rebalancing, EC2 Auto Scaling may briefly go above the group maximum by up to 10 percent or one extra instance to avoid impact → With a mixed instances policy → scale-in.
an Amazon EventBridge rule for AWS Health EC2 retirement scheduled events that
AWS Health emits the scheduled retirement event and EventBridge can target it precisely, allowing SSM Automation to orchestrate the required stop/start to move the instance to.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health emits the scheduled retirement event and EventBridge can target it precisely, allowing SSM Automation to orchestrate.
● Scene fit: AWS Health emits the scheduled retirement event and EventBridge can target it precisely, allowing SSM Automation to orchestrate.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Auto Scaling replacement reacts to health failures and does not proactively stop and start instances in response to scheduled.
● EC2 state-change events do not indicate scheduled retirements and will fire for many unrelated stops, creating noisy and mis-scoped.
● Auto Recovery addresses impaired instances rather than scheduled retirements, and CloudWatch alarm actions cannot be time-scheduled for off-hours execution.
Workflow: Configure an Amazon EventBridge rule for AWS Health EC2 retirement scheduled events that invokes an AWS Systems Manager Automation runbook to stop → start the impacted instances.
Target tracking policy for Spot Fleet to maintain about 60% average CPU
Application Auto Scaling target tracking adjusts Spot Fleet capacity to hold CPU near the target. Scheduled actions set provisioned concurrency at specific times for predictable patterns.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Application Auto Scaling target tracking adjusts Spot Fleet capacity to hold CPU near the target.
● Scene fit: Application Auto Scaling target tracking adjusts Spot Fleet capacity to hold CPU near the target. Scheduled actions set.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Predictive scaling applies to EC2 Auto Scaling groups, not Spot Fleet, and does not directly control fleet CPU utilization.
● With target tracking, Application Auto Scaling creates and manages the CloudWatch alarms automatically.
● Lambda does not support step scaling with Application Auto Scaling; use provisioned concurrency with scheduled or target tracking instead.
Workflow: Target tracking policy for Spot Fleet to maintain about 60% average CPU → Scheduled scaling for Lambda provisioned concurrency on an alias during the midweek peak.
the CloudWatch agent to forward /var/log/secure or auth.log to CloudWatch Logs, create
Metric filters convert matching log lines into metrics that can be alarmed on, enabling near real-time SNS notifications for login events.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric filters convert matching log lines into metrics that can be alarmed on, enabling near real-time SNS notifications.
● Scene fit: Metric filters convert matching log lines into metrics that can be alarmed on, enabling near real-time SNS notifications.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail focuses on AWS API activity and delivery is not guaranteed within 60 seconds, so it is not suitable.
● Subscription filters forward matching events to a destination but do not create metrics or alarms for alerting by themselves.
● GuardDuty generates threat findings and anomalies rather than alerting on every successful host login in near real time.
Workflow: Deploy the CloudWatch agent to forward /var/log/secure or auth.log to CloudWatch Logs → create a metric filter for successful login patterns, and trigger an SNS alarm to alert the security team.
AWS Config cloudtrail-enabled or cloudtrail-security-trail-enabled with EventBridge-triggered remediation
Config continuously evaluates and records compliance and can trigger remediation via EventBridge to Lambda or SSM.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Config continuously evaluates and records compliance and can trigger remediation via EventBridge to Lambda or SSM.
● Scene fit: Config continuously evaluates and records compliance and can trigger remediation via EventBridge to Lambda or SSM.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Periodic polling can miss changes and does not provide a continuous compliance history.
● SCPs block actions but do not ensure CloudTrail is enabled, provide compliance history, or perform remediation.
● Security Hub detects and aggregates findings but does not automatically remediate or maintain per-rule configuration history.
Workflow: AWS Config cloudtrail-enabled or cloudtrail-security-trail-enabled with EventBridge-triggered remediation.
Amazon Macie
Performs automated sensitive data discovery in S3 and generates findings for potential exposure such as public or cross-account access.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Performs automated sensitive data discovery in S3 and generates findings for potential exposure such as public or cross-account access.
● Scene fit: Performs automated sensitive data discovery in S3 and generates findings for potential exposure such as public or cross-account access.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Aggregates findings from other services and frameworks; does not inspect S3 object contents or classify PII.
● Analyzes resource policies for public and cross-account access but does not classify or scan data for PII.
● Threat detection using CloudTrail, VPC Flow Logs, and DNS logs; not designed for S3 PII classification.
Workflow: Amazon Macie.
an AWS WAF web ACL to the CloudFront distribution + Deploy CloudFront
Provides edge inspection and mitigation for common web exploits before traffic reaches the origin. CloudFront reduces global latency, integrates with ACM for HTTPS, and benefits from AWS Shield Standard.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides edge inspection and mitigation for common web exploits before traffic reaches the origin.
● Scene fit: Provides edge inspection and mitigation for common web exploits before traffic reaches the origin. CloudFront reduces global latency.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Improves global performance and L3/L4 DDoS resilience but does not provide edge WAF inspection or caching.
● Secures at the regional ALB but lacks edge inspection and does not improve global latency.
● CloudFront cannot use an Auto Scaling group directly as an origin; use an ALB or instance endpoint.
Workflow: Attach an AWS WAF web ACL to the CloudFront distribution → Deploy CloudFront in front of the ALB and use ACM for the custom domain TLS.
an Organizations SCP that explicitly denies iam:CreateUser for all principals unless aws:username
An SCP with an explicit Deny on CreateUser and a condition that allows only listed usernames is a preventive, organization-wide control.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An SCP with an explicit Deny on CreateUser and a condition that allows only listed usernames is a preventive, organization-wide control.
● Scene fit: An SCP with an explicit Deny on CreateUser and a condition that allows only listed usernames is a preventive, organization-wide control.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is reactive and attempts cleanup after the fact rather than preventing the CreateUser call at the organization boundary.
● CreateLoginProfile controls passwords for existing users and does not block the creation of IAM users.
● Permissions boundaries must be attached to each principal, cannot constrain the root user, and are not an effective org-wide preventive control.
Workflow: Attach an Organizations SCP that explicitly denies iam:CreateUser for all principals unless aws:username matches an approved exception list by using a StringNotLike condition.
Kinesis Data Streams + Managed Service for Apache Flink + Firehose to
This combines purpose-built streaming with serverless delivery to S3 and a low-cost metastore and query layer for ad hoc analytics.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This combines purpose-built streaming with serverless delivery to S3 and a low-cost metastore and query layer for ad hoc analytics.
● Scene fit: This combines purpose-built streaming with serverless delivery to S3 and a low-cost metastore and query layer for ad hoc analytics.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EventBridge is not ideal for sustained high-throughput streaming and introduces unnecessary cost and limits for clickstreams.
● EMR on EC2 increases operational and compute cost, and SQS is not designed for real-time streaming ingestion.
● Viable but typically higher operational overhead and cost than Kinesis for this use case.
Workflow: Kinesis Data Streams + Managed Service for Apache Flink + Firehose to S3 + Glue + Athena.
Insufficient IAM permissions for the update + Resources changed outside CloudFormation
Missing permissions can block resource changes during update or rollback, leading to UPDATE_ROLLBACK_FAILED. Out-of-band changes create drift that can prevent CloudFormation from reverting resources during rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Missing permissions can block resource changes during update or rollback, leading to UPDATE_ROLLBACK_FAILED.
● Scene fit: Missing permissions can block resource changes during update or rollback, leading to UPDATE_ROLLBACK_FAILED. Out-of-band changes create drift that can prevent CloudFormation from reverting resources during rollback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Rollback triggers are optional and their absence does not cause rollback failures.
● CloudFormation uses regional public APIs; an interface endpoint is not required and its unavailability would not cause this state.
● Change sets are optional planning tools and not using one does not itself cause rollback failures.
Workflow: Insufficient IAM permissions for the update → Resources changed outside CloudFormation.
identical runOrder for actions in the stage
Actions sharing the same runOrder in a stage run in parallel, shortening overall pipeline time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Actions sharing the same runOrder in a stage run in parallel, shortening overall pipeline time.
● Scene fit: Actions sharing the same runOrder in a stage run in parallel, shortening overall pipeline time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Build batches parallelize CodeBuild jobs but do not parallelize distinct CodePipeline actions or deployments for separate functions.
● External orchestration adds complexity when CodePipeline supports native parallel actions.
● Caching can speed builds but does not make sequential CodePipeline actions run concurrently.
Workflow: Set identical runOrder for actions in the stage.
Schedule a Lambda to refresh AWS Trusted Advisor service quota checks with
Trusted Advisor provides built-in service quota checks and integrates with EventBridge, enabling simple notifications with minimal code.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Trusted Advisor provides built-in service quota checks and integrates with EventBridge, enabling simple notifications with minimal code.
● Scene fit: Trusted Advisor provides built-in service quota checks and integrates with EventBridge, enabling simple notifications with minimal code.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Demands significant custom coding to model every service's limits and usage, which is not the least effort.
● A custom Config rule would still require bespoke logic to discover and evaluate quotas, increasing complexity.
● AWS Health does not provide service quota monitoring, and mixing Health refresh with Trusted Advisor events is ineffective.
Workflow: Schedule a Lambda to refresh AWS Trusted Advisor service quota checks with EventBridge and add another EventBridge rule that matches Trusted Advisor limit events and publishes to an SNS topic subscribed by the operations manager.
BeforeAllowTraffic hook in the Lambda AppSpec
Runs pre-traffic checks and blocks the alias shift until readiness is confirmed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Runs pre-traffic checks and blocks the alias shift until readiness is confirmed.
● Scene fit: Runs pre-traffic checks and blocks the alias shift until readiness is confirmed.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Reduces cold starts but does not gate traffic or ensure dependencies are ready.
● Limits blast radius but still shifts some traffic before checks complete.
● Executes after traffic has already shifted, so it cannot prevent initial errors.
Workflow: Use BeforeAllowTraffic hook in the Lambda AppSpec.
The S3 event notification configuration was removed from the bucket + The
If the bucket notification configuration is deleted, S3 stops sending events to the Lambda function. Without the correct allow statement in the function resource policy, S3 cannot invoke the Lambda function.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: If the bucket notification configuration is deleted, S3 stops sending events to the Lambda function.
● Scene fit: If the bucket notification configuration is deleted, S3 stops sending events to the Lambda function. Without the correct.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Turning on default encryption does not interfere with S3 event notifications or Lambda triggers.
● Object Lock affects object retention and deletion behavior but does not block event notifications.
● S3 invokes Lambda using the function’s resource-based policy, so the execution role is not required for invocation to occur.
Workflow: The S3 event notification configuration was removed from the bucket → The Lambda function’s resource-based policy that permits s3.amazonaws.com to invoke it was deleted.
Two S3 buckets in separate Regions at least 900 miles apart; bucket
This enforces HTTPS with a bucket policy, encrypts at rest with SSE-S3, and uses CRR across Regions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This enforces HTTPS with a bucket policy, encrypts at rest with SSE-S3, and uses CRR across Regions.
● Scene fit: This enforces HTTPS with a bucket policy, encrypts at rest with SSE-S3, and uses CRR across Regions.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● No replication is configured, so objects are not copied between Regions.
● Multi-Region Access Points do not copy data and still require S3 replication rules.
● IAM roles cannot enforce TLS; HTTPS-only must be enforced via S3 bucket policy using aws:SecureTransport.
Workflow: Two S3 buckets in separate Regions at least 900 miles apart → bucket policy forces HTTPS → require SSE-S3 → enable S3 cross-Region replication.
a Lambda alias with weighted routing to split traffic between function versions
Lambda alias routing configuration supports weighted traffic between two versions, enabling a controlled canary and promotion using the same API endpoint.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda alias routing configuration supports weighted traffic between two versions, enabling a controlled canary and promotion using the same API endpoint.
● Scene fit: Lambda alias routing configuration supports weighted traffic between two versions, enabling a controlled canary and promotion using the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Splits DNS across endpoints but requires separate API deployments and does not control traffic between Lambda versions behind a single alias.
● Lambda does not support rolling updates of $LATEST and CodeDeploy manages version traffic shifting via aliases rather than updating $LATEST in place.
● API Gateway canaries operate at the stage level and typically require duplicate stages, which adds overhead and does not directly shift traffic between Lambda versions.
Workflow: Configure a Lambda alias with weighted routing to split traffic between function versions.
Elastic Beanstalk worker with SQS and S3 object keys; cron.yaml in worker
Decouples heavy work from the web tier, respects SQS size limits by passing S3 pointers, and uses cron.yaml where supported.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Decouples heavy work from the web tier, respects SQS size limits by passing S3 pointers, and uses cron.yaml where supported.
● Scene fit: Decouples heavy work from the web tier, respects SQS size limits by passing S3 pointers, and uses cron.yaml where supported.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Large, CPU-bound parsing can exceed Lambda time and memory limits and is less controllable for long-running workloads.
● Feasible but adds migration complexity and is unnecessary when Beanstalk worker tiers already provide the required pattern.
● Cron.yaml is a worker-tier feature; running it on the web tier is unsupported and unreliable.
Workflow: Elastic Beanstalk worker with SQS and S3 object keys → cron.yaml in worker for emails.
Decoupled stacks per domain with Export/ImportValue and separate versioning
Separate domain stacks with cross-stack exports/imports enable independent deployments and safe dependency management following best practices.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Separate domain stacks with cross-stack exports/imports enable independent deployments and safe dependency management following best practices.
● Scene fit: Separate domain stacks with cross-stack exports/imports enable independent deployments and safe dependency management following best practices.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Nested stacks couple lifecycles; parent updates increase coordination and blast radius.
● StackSets target multi-account/regional rollouts but still centralize cadence and do not simplify fine-grained inter-stack dependencies.
● A monolith couples all resources, increases blast radius, and blocks independent releases.
Workflow: Decoupled stacks per domain with Export/ImportValue and separate versioning.
EventBridge rule and Lambda to refresh Trusted Advisor via Support API every
Scheduling a refresh of Trusted Advisor checks and sending findings to SNS provides timely alerts. A managed or custom AWS Config rule can detect open SSH.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Scheduling a refresh of Trusted Advisor checks and sending findings to SNS provides timely alerts.
● Scene fit: Scheduling a refresh of Trusted Advisor checks and sending findings to SNS provides timely alerts. A managed or.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS Config native remediation uses AWS Systems Manager Automation runbooks, not direct Lambda invocations.
● GuardDuty detects threats, not configuration drift in security groups, and does not auto-remediate SG rules.
● Security Hub does not ingest Trusted Advisor findings and does not natively trigger TA-based remediation.
Workflow: EventBridge rule and Lambda to refresh Trusted Advisor via Support API every 10 minutes and publish to SNS → AWS Config rule flagging SGs with port 22 open to 0.0.0.0/0 and notifying.
Publish a new version and set the alias weight to 20% for
Weighted aliases natively shift a percentage of traffic between versions and allow instant rollback without client changes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Weighted aliases natively shift a percentage of traffic between versions and allow instant rollback without client changes.
● Scene fit: Weighted aliases natively shift a percentage of traffic between versions and allow instant rollback without client changes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adds API Gateway and requires client/API changes; unnecessary for direct Lambda alias traffic shifting.
● Works but introduces extra setup and management; not the lowest operational overhead for this simple weighted shift.
● Route 53 weights DNS traffic and cannot weight direct Lambda invocations.
Workflow: Publish a new version and set the alias weight to 20% for the new version.
S3 static hosting with CloudFront and an AWS WAF web ACL
S3 offers serverless static hosting, CloudFront accelerates globally, and AWS WAF mitigates common web exploits.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 offers serverless static hosting, CloudFront accelerates globally, and AWS WAF mitigates common web exploits.
● Scene fit: S3 offers serverless static hosting, CloudFront accelerates globally, and AWS WAF mitigates common web exploits.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Introduces unnecessary compute for static assets and GuardDuty detects threats but does not block web attacks.
● Provides CDN and static hosting but lacks AWS WAF for common web exploit mitigation; Shield Standard focuses on DDoS only.
● Not serverless, Redis is not needed for static site delivery, and Shield Advanced targets DDoS rather than general web exploits.
Workflow: S3 static hosting with CloudFront and an AWS WAF web ACL.
Amazon EFS replication to a cross-Region destination
Native EFS replication asynchronously maintains a replica in another Region without requiring VPC connectivity, supporting low RPO and fast failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Native EFS replication asynchronously maintains a replica in another Region without requiring VPC connectivity, supporting low RPO and fast failover.
● Scene fit: Native EFS replication asynchronously maintains a replica in another Region without requiring VPC connectivity, supporting low RPO and fast failover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DataSync requires network connectivity to both locations and typically runs scheduled or continuous tasks over an established network path.
● Backups provide point-in-time copies that must be restored, leading to higher RPO and RTO and not a hot standby.
● Object replication and Lambda-based copies are not file-system aware and add complexity and latency, undermining RPO/RTO and POSIX fidelity.
Workflow: Enable Amazon EFS replication to a cross-Region destination.
Route 53 latency-based routing to regional API Gateway with health checks; in-Region
Latency-based routing sends clients to the lowest-latency endpoint and health checks enable automatic failover; DynamoDB global tables support active-active data.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency-based routing sends clients to the lowest-latency endpoint and health checks enable automatic failover; DynamoDB global tables support active-active data.
● Scene fit: Latency-based routing sends clients to the lowest-latency endpoint and health checks enable automatic failover; DynamoDB global tables support active-active data.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Geolocation routes by user location, not measured latency, so it may not choose the fastest Region.
● CloudFront can cache and improve TLS/edge reach, but a single origin does not provide multi-Region latency steering.
● Failover is active-passive for resilience, not for distributing requests to the lowest-latency Region.
Workflow: Route 53 latency-based routing to regional API Gateway with health checks → in-Region Lambda → DynamoDB global tables.
interface VPC endpoints for Systems Manager, SSMMessages, and EC2Messages in the VPC
Interface VPC endpoints keep Session Manager traffic on the AWS private network, enabling private connectivity without traversing the internet. Granting permissions via an IAM instance profile.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Interface VPC endpoints keep Session Manager traffic on the AWS private network, enabling private connectivity without traversing the internet.
● Scene fit: Interface VPC endpoints keep Session Manager traffic on the AWS private network, enabling private connectivity without traversing the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Placing long-lived access keys on instances is insecure and not the recommended way to grant Systems Manager permissions.
● Session Manager does not require opening SSH inbound ports, so allowing TCP 22 is unnecessary and weakens security.
● A VPN is not required for private Session Manager connectivity because PrivateLink VPC endpoints handle this within AWS.
Workflow: Create interface VPC endpoints for Systems Manager, SSMMessages, and EC2Messages in the VPC → Attach an IAM policy with required Systems Manager permissions to the instances' IAM role or instance profile.
a parallel environment behind the ALB with the new build and configure
API Gateway canary releases allow routing a small percentage of traffic to a new backend and make rollback as simple as adjusting or disabling the canary.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: API Gateway canary releases allow routing a small percentage of traffic to a new backend and make rollback.
● Scene fit: API Gateway canary releases allow routing a small percentage of traffic to a new backend and make rollback.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While CodeDeploy blue/green can reduce disruption and enable rollback, it introduces additional tooling, agents, and configuration that are not.
● Flipping the DNS alias causes an all-at-once cutover that impacts all users and does not provide a gradual rollout.
● API Gateway cannot weight traffic across ALB target groups; weighted traffic splitting at the API tier is done with.
Workflow: Create a parallel environment behind the ALB with the new build and configure API Gateway canary release to send a small portion of requests to it.
subscription filters on all log groups to stream to Amazon Data Firehose
This uses native CloudWatch Logs subscriptions to Amazon Data Firehose for near real-time delivery to S3 and automates attachment for new log groups with EventBridge and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses native CloudWatch Logs subscriptions to Amazon Data Firehose for near real-time delivery to S3 and automates.
● Scene fit: This uses native CloudWatch Logs subscriptions to Amazon Data Firehose for near real-time delivery to S3 and automates.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch Logs exports are batch, per log group, time-bounded jobs and do not automatically cover newly created groups or.
● DataSync does not integrate with CloudWatch Logs as a source and is intended for file and object storage endpoints.
Workflow: Configure subscription filters on all log groups to stream to Amazon Data Firehose and set the delivery stream to write to the S3 bucket in the audit account, with an EventBridge rule.
Lambda role permissions and EFS access point mount + VPC peering plus
Lambda needs VPC and elasticfilesystem permissions and to mount via an access point. Provide network reachability and a file system policy allowing the other account to mount and write.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda needs VPC and elasticfilesystem permissions and to mount via an access point.
● Scene fit: Lambda needs VPC and elasticfilesystem permissions and to mount via an access point. Provide network reachability and a file system policy allowing the other account to mount and write.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EFS does not support PrivateLink.
● Network alone is insufficient; EFS resource policy is still required.
● SCPs are guardrails and do not grant access to resources.
Workflow: Lambda role permissions and EFS access point mount → VPC peering → EFS file system policy.
Application Load Balancer behind an Auto Scaling group + Amazon CloudFront with
Provides Layer 7 host and path-based routing to EC2 targets and scales cost-effectively with demand. Caches content close to users and runs edge logic to tailor behavior or origin selection by path or device, reducing latency and origin load.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides Layer 7 host and path-based routing to EC2 targets and scales cost-effectively with demand.
● Scene fit: Provides Layer 7 host and path-based routing to EC2 targets and scales cost-effectively with demand. Caches content close.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Improves availability and routing performance using Anycast, but it does not provide path-based routing or caching and adds cost.
● Operates at the DNS layer for domain, geolocation, and latency-based decisions, not per-request URL path routing.
● Works at Layer 4 and cannot inspect HTTP paths for content-based routing.
Workflow: Application Load Balancer behind an Auto Scaling group → Amazon CloudFront with Lambda@Edge.
Build an AWS CodePipeline that triggers from GitHub webhooks on a specific
This follows CI/CD best practices by using EventBridge for pipeline events and SNS for email notifications while gating S3 staging with a manual approval.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This follows CI/CD best practices by using EventBridge for pipeline events and SNS for email notifications while gating.
● Scene fit: This follows CI/CD best practices by using EventBridge for pipeline events and SNS for email notifications while gating.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Uses EventBridge correctly but relies on SES for event notifications, which is less suitable than SNS for direct, scalable.
● CloudTrail captures API audit logs and is not designed for real-time pipeline stage change detection.
● CloudWatch Logs would require custom parsing and is not an event source for CodePipeline state transitions.
Workflow: Build an AWS CodePipeline that triggers from GitHub webhooks on a specific branch with stages for security scan, unit and functional testing, and a manual approval step before pushing artifacts to Amazon.
Amazon ElastiCache for Redis (Multi-AZ) for sessions
A managed, in-memory, multi-AZ Redis provides low-latency, shared session storage with automatic failover across instance changes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A managed, in-memory, multi-AZ Redis provides low-latency, shared session storage with automatic failover across instance changes.
● Scene fit: A managed, in-memory, multi-AZ Redis provides low-latency, shared session storage with automatic failover across instance changes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Stickiness binds clients to specific instances and does not survive instance replacement, harming load distribution and scalability.
● DynamoDB can persist sessions but is higher-latency than in-memory caches and is less efficient for frequent session reads/writes.
● Relational databases add latency and overhead for ephemeral session data and are not optimized for session caching patterns.
Workflow: Use Amazon ElastiCache for Redis (Multi-AZ) for sessions.
Provision Aurora Global Database for catalog; keep other tables regional in Aurora
Aurora Global Database gives low-latency global reads and fast cross-Region replication while preserving the relational model and minimizing code changes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora Global Database gives low-latency global reads and fast cross-Region replication while preserving the relational model and minimizing code changes.
● Scene fit: Aurora Global Database gives low-latency global reads and fast cross-Region replication while preserving the relational model and minimizing code changes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Mixing NoSQL and SQL adds API and data model changes, increasing refactoring effort.
● Replacing the relational model with NoSQL requires significant redesign and code changes.
● Replicates all tables globally, not just catalog, which conflicts with keeping accounts and view_history regional.
Workflow: Provision Aurora Global Database for catalog → keep other tables regional in Aurora.
Provision Amazon Aurora MySQL with cross-Region read replicas for the catalog and
Aurora MySQL is MySQL-compatible for minimal code changes, supports cross-Region read replicas for low-latency catalog reads, and separate per-Region clusters keep order data resident for compliance.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora MySQL is MySQL-compatible for minimal code changes, supports cross-Region read replicas for low-latency catalog reads, and separate.
● Scene fit: Aurora MySQL is MySQL-compatible for minimal code changes, supports cross-Region read replicas for low-latency catalog reads, and separate.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Would require moving from a relational model to a NoSQL key-value store, which typically demands significant schema and code.
● While feasible and relational, this can introduce higher replica lag and more operational overhead than Aurora for global read.
● Global tables would replicate order data across Regions, violating the requirement to keep orders within each Region and also.
Workflow: Provision Amazon Aurora MySQL with cross-Region read replicas for the catalog and independent Aurora clusters in each Region for order data.
an AWS CodePipeline with a Git repository as the source to orchestrate
CodePipeline provides end-to-end orchestration, stage-level visibility, and integrates with automated failure handling and rollbacks. The BeforeInstall hook in CodeDeploy is the correct place to clear caches.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodePipeline provides end-to-end orchestration, stage-level visibility, and integrates with automated failure handling and rollbacks.
● Scene fit: CodePipeline provides end-to-end orchestration, stage-level visibility, and integrates with automated failure handling and rollbacks. The BeforeInstall hook in.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Systems Manager is not the primary AWS service for application deployments and lacks CodeDeploy lifecycle hooks and native rollback.
● EventBridge plus Lambda can build artifacts but does not provide full pipeline orchestration, approvals, or deployment tracking.
● EC2 user data runs only at first boot and cannot reliably handle repeated deployment steps or managed rollbacks.
Workflow: Create an AWS CodePipeline with a Git repository as the source to orchestrate build and deployment stages → Add a cache-purge script to the AppSpec BeforeInstall lifecycle hook for the deployment →.
Install the unified CloudWatch agent on all on-premises servers and EC2 instances
This centralizes hybrid log collection with minimal ops, uses low-cost S3 storage, and enables serverless analytics with Athena.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This centralizes hybrid log collection with minimal ops, uses low-cost S3 storage, and enables serverless analytics with Athena.
● Scene fit: This centralizes hybrid log collection with minimal ops, uses low-cost S3 storage, and enables serverless analytics with Athena.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Misses on-premises servers, relies on manual exports that do not scale, and creates operational overhead with EMR.
● Self-hosting storage and ELK increases cost and effort compared to managed, serverless alternatives.
● While workable, provisioning and running OpenSearch increases cost and management compared to S3 plus Athena for audit-style queries.
Workflow: Install the unified CloudWatch agent on all on-premises servers and EC2 instances to send logs to CloudWatch Logs → subscribe the log groups to Kinesis Data Firehose delivering to a central Amazon.
CloudFront origin failover for the login origin + Lambda@Edge authentication at the
Configures a secondary origin to automatically serve on 5xx (including 504), boosting availability at low cost. Executes auth logic near users to cut round trips and reduce load on the origin.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Configures a secondary origin to automatically serve on 5xx (including 504), boosting availability at low cost.
● Scene fit: Configures a secondary origin to automatically serve on 5xx (including 504), boosting availability at low cost. Executes auth logic near users to cut round trips and reduce load on the origin.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Improves network path but adds cost and does not resolve origin 5xx on authentication.
● Multi-Region deployments reduce latency but are complex and costly for authentication state.
● Optimizes cache hits for static content but does not accelerate dynamic login or prevent 504s.
Workflow: CloudFront origin failover for the login origin → Lambda@Edge authentication at the edge.
a new DynamoDB table that includes a Local Secondary Index on DocTitle
Local Secondary Indexes share the partition key and provide an alternate sort key with support for strongly consistent reads, but they must be created with the.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Local Secondary Indexes share the partition key and provide an alternate sort key with support for strongly consistent reads, but they must be created with the table so migration is required.
● Scene fit: Local Secondary Indexes share the partition key and provide an alternate sort key with support for strongly consistent.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● You cannot add a Local Secondary Index to an existing table because LSIs must be defined at table creation time.
● Global Secondary Indexes do not support strongly consistent reads, which the requirement mandates.
● Streams-based replication is asynchronous and cannot guarantee the immediate strong consistency required for the latest updates.
Workflow: Create a new DynamoDB table that includes a Local Secondary Index on DocTitle with a new sort key, → migrate existing data.
VM Import/Export: import VMware VM to AMI, validate on EC2, export as
Supports importing VMware VMs as AMIs and exporting previously imported instances as OVA for vSphere testing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Supports importing VMware VMs as AMIs and exporting previously imported instances as OVA for vSphere testing.
● Scene fit: Supports importing VMware VMs as AMIs and exporting previously imported instances as OVA for vSphere testing.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Builds and tests AMIs on AWS but does not export images back to VMware for on-prem parity.
● AWS does not provide a bare-metal ISO for Amazon Linux 2 and this does not mirror EC2 behavior.
● Brings AWS hardware on-prem but is costly and unnecessary for simple image validation and does not provide VM round-trip.
Workflow: VM Import/Export: import VMware VM to AMI → validate on EC2, export as OVA to vSphere.
CloudWatch agent -> CloudWatch Logs; cross-account subscription to Kinesis Data Firehose ->
This streams all logs into a single low-cost S3 data lake with serverless enrichment and ad hoc SQL across accounts and on-premises.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This streams all logs into a single low-cost S3 data lake with serverless enrichment and ad hoc SQL across accounts and on-premises.
● Scene fit: This streams all logs into a single low-cost S3 data lake with serverless enrichment and ad hoc SQL across accounts and on-premises.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Kinesis Data Analytics requires streaming inputs and cannot directly query data at rest in S3.
● Silos data by account and prevents a single, centralized analytics location.
● OpenSearch is not the lowest-cost long-term store and does not provide S3-based ad hoc SQL querying with Athena.
Workflow: CloudWatch agent -> CloudWatch Logs → cross-account subscription to Kinesis Data Firehose -> central S3 (logging account) → Lambda tagging → query with Athena.
AppSpec AfterAllowTestTraffic with Lambda
Intended for running validation after test traffic is enabled, allowing rollback before production cutover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Intended for running validation after test traffic is enabled, allowing rollback before production cutover.
● Scene fit: Intended for running validation after test traffic is enabled, allowing rollback before production cutover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs before any test traffic is routed, so it cannot validate using the test listener.
● Event shifts test traffic but occurs before validation should run.
● Occurs after test validation should already have completed and just before production traffic is shifted.
Workflow: AppSpec AfterAllowTestTraffic with Lambda.
Auto Scaling warm pools with lifecycle hooks
Warm pools keep pre-initialized instances for rapid scale-out, and lifecycle hooks gate InService until the app is ready.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Warm pools keep pre-initialized instances for rapid scale-out, and lifecycle hooks gate InService until the app is ready.
● Scene fit: Warm pools keep pre-initialized instances for rapid scale-out, and lifecycle hooks gate InService until the app is ready.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Pre-scaling far in advance overprovisions capacity and increases cost without solving readiness gating.
● Scheduled actions can add capacity at a known time but do not ensure instances serve traffic only after app bootstrap completes.
● Maintaining extra always-on capacity is costly and increases operational overhead.
Workflow: Auto Scaling warm pools with lifecycle hooks.
Publish a custom CloudWatch metric using statistic sets aggregated every 120 seconds
Statistic sets roll up many 10-second samples into one PutMetricData call per 120 seconds, enabling standard-resolution alarms at lower cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Statistic sets roll up many 10-second samples into one PutMetricData call per 120 seconds, enabling standard-resolution alarms at lower cost.
● Scene fit: Statistic sets roll up many 10-second samples into one PutMetricData call per 120 seconds, enabling standard-resolution alarms at lower cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Runs canaries externally and adds cost; it does not use the instance's probe output directly for low-cost aggregation.
● High-resolution metrics and frequent publishes increase cost and are unnecessary when 2-minute granularity is acceptable.
● Possible but adds log ingestion and filter costs and complexity versus directly publishing aggregated custom metrics.
Workflow: Publish a custom CloudWatch metric using statistic sets aggregated every 120 seconds.
AWS Elastic Beanstalk (load balanced, auto scaled) with Amazon RDS for MySQL
Elastic Beanstalk supports rolling and blue/green deployments with easy rollback, and an external RDS can be shared and is decoupled from app lifecycle.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Elastic Beanstalk supports rolling and blue/green deployments with easy rollback, and an external RDS can be shared and is decoupled from app lifecycle.
● Scene fit: Elastic Beanstalk supports rolling and blue/green deployments with easy rollback, and an external RDS can be shared and is decoupled from app lifecycle.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Supports rolling updates, but streamlined rollback typically needs additional deployment tooling and coordination.
● Delivers zero-downtime and rollback, but adds more operational complexity than a managed PaaS.
● Couples DB lifecycle to the app environment and risks deletion during environment teardown; sharing is harder.
Workflow: AWS Elastic Beanstalk (load balanced, auto scaled) with Amazon RDS for MySQL created outside the environment.
Rehost the Oracle RAC database on EBS-backed Amazon EC2, install the SSM
Running RAC on EC2 is supported, Patch Manager automates OS patching, and DLM provides scheduled EBS snapshots with minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Running RAC on EC2 is supported, Patch Manager automates OS patching, and DLM provides scheduled EBS snapshots with.
● Scene fit: Running RAC on EC2 is supported, Patch Manager automates OS patching, and DLM provides scheduled EBS snapshots with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Aurora does not support Oracle RAC, so this does not meet the migration requirement even though it automates backups and patching.
● Using Lambda for snapshots is heavier than DLM and CodeDeploy/CodePipeline are CI/CD tools rather than patch automation services.
● RDS for Oracle does not support Oracle RAC, so this option cannot satisfy the RAC migration requirement.
Workflow: Rehost the Oracle RAC database on EBS-backed Amazon EC2 → install the SSM agent → use AWS Systems Manager Patch Manager for OS patches, and configure Amazon Data Lifecycle Manager to schedule EBS snapshots.
UpdatePolicy: AutoScalingReplacingUpdate with WillReplace true
A replacing update with WillReplace true creates a new Auto Scaling group alongside the old one, preserving capacity and allowing immediate rollback if needed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A replacing update with WillReplace true creates a new Auto Scaling group alongside the old one, preserving capacity and allowing immediate rollback if needed.
● Scene fit: A replacing update with WillReplace true creates a new Auto Scaling group alongside the old one, preserving capacity and allowing immediate rollback if needed.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Rolling updates replace instances in batches, which can reduce available capacity and do not provide instant rollback to the previous group.
● CodeDeploy blue/green is a deployment strategy, not a CloudFormation UpdatePolicy for Auto Scaling groups.
● WillReplace false uses rolling or in-place behavior and does not create a parallel group to preserve full capacity or enable instant rollback.
Workflow: UpdatePolicy: AutoScalingReplacingUpdate with WillReplace true.
Elastic Beanstalk with environment swap
Elastic Beanstalk supports blue/green by cloning environments and performing a CNAME swap for managed traffic redirect.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Elastic Beanstalk supports blue/green by cloning environments and performing a CNAME swap for managed traffic redirect.
● Scene fit: Elastic Beanstalk supports blue/green by cloning environments and performing a CNAME swap for managed traffic redirect.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Instance refresh performs rolling replacements and does not maintain two parallel environments for traffic shifting.
● ECS rolling updates replace tasks in place and do not natively offer parallel environment traffic swap without additional tooling.
● App Runner simplifies deployment but does not provide environment-level blue/green with parallel stacks and controlled traffic swap.
Workflow: Elastic Beanstalk with environment swap.
unit and functional tests in CodeBuild and block promotion on failures +
CodeBuild can execute automated test suites and fail the pipeline gate if tests do not pass. A staged deployment with manual approval enables hands-on verification and prevents faulty releases from reaching production.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeBuild can execute automated test suites and fail the pipeline gate if tests do not pass.
● Scene fit: CodeBuild can execute automated test suites and fail the pipeline gate if tests do not pass. A staged deployment with manual approval enables hands-on verification and prevents faulty releases from reaching production.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● GuardDuty is for threat detection and does not validate application functionality.
● Inspector focuses on vulnerabilities and exposure, not functional correctness.
● AWS Config assesses resource configuration compliance, not application functional behavior.
Workflow: Run unit and functional tests in CodeBuild and block promotion on failures → Deploy to a stage with CodeDeploy → add a manual approval gate, → deploy to production.
CloudWatch metric streams to Kinesis Data Firehose delivering to S3, query with
Metric streams natively deliver metrics to S3 via Firehose without custom code, enabling durable multi-year storage, Athena queries, and QuickSight dashboards.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric streams natively deliver metrics to S3 via Firehose without custom code, enabling durable multi-year storage, Athena queries, and QuickSight dashboards.
● Scene fit: Metric streams natively deliver metrics to S3 via Firehose without custom code, enabling durable multi-year storage, Athena queries, and QuickSight dashboards.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Requires custom code, scheduling, and maintenance, which is higher effort than a native streaming solution.
● CloudWatch metrics retain at most about 15 months, which does not meet a 10-year requirement.
● CloudWatch dashboards cannot read data from S3, so dashboards cannot be built from exported files.
Workflow: Use CloudWatch metric streams to Kinesis Data Firehose delivering to S3 → query with Athena, and visualize in QuickSight.
Amazon Managed Service for Prometheus plus Amazon Managed Grafana
Amazon Managed Service for Prometheus ingests Prometheus metrics via remote write from EKS, ECS, and on-prem Kubernetes and Amazon Managed Grafana provides managed dashboards, queries, and alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Amazon Managed Service for Prometheus ingests Prometheus metrics via remote write from EKS, ECS, and on-prem Kubernetes and Amazon Managed Grafana provides managed dashboards, queries, and alerts.
● Scene fit: Amazon Managed Service for Prometheus ingests Prometheus metrics via remote write from EKS, ECS, and on-prem Kubernetes and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch agent and Athena/QuickSight are not Prometheus-native and do not provide a cohesive Prometheus scraping and remote write backend across clusters.
● SSM Agent does not scrape Prometheus metrics and Amazon Managed Service for Prometheus is not a visualization tool.
● Good for EKS/ECS metrics in CloudWatch, but it is not a Prometheus remote write backend and does not unify native Prometheus metrics across environments.
Workflow: Amazon Managed Service for Prometheus → Amazon Managed Grafana.
an ACM-issued certificate to an HTTPS listener on the ALB to offload
Placing the certificate on the ALB terminates TLS at the load balancer, protecting client connections while offloading encryption from the instances. AWS WAF can be attached.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Placing the certificate on the ALB terminates TLS at the load balancer, protecting client connections while offloading encryption.
● Scene fit: Placing the certificate on the ALB terminates TLS at the load balancer, protecting client connections while offloading encryption.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Terminating TLS on each instance consumes CPU on the hosts and conflicts with the requirement to avoid instance performance impact.
● EBS encryption safeguards data at rest on volumes and does not encrypt network traffic in transit.
● Shield Advanced primarily mitigates DDoS attacks and is not designed to filter common application-layer exploits like SQL injection or XSS.
Workflow: Attach an ACM-issued certificate to an HTTPS listener on the ALB to offload TLS → Create an AWS WAF web ACL with managed rules and associate it to the ALB.
EventBridge rule on CloudTrail sts:AssumeRole for the role, invoke Lambda to SNS
Filter AWS API Call via CloudTrail events for sts:AssumeRole with the role ARN and trigger Lambda to publish to SNS for immediate alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Filter AWS API Call via CloudTrail events for sts:AssumeRole with the role ARN and trigger Lambda to publish to SNS for immediate alerts.
● Scene fit: Filter AWS API Call via CloudTrail events for sts:AssumeRole with the role ARN and trigger Lambda to publish to SNS for immediate alerts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudTrail Insights detects anomalies, not every specific sts:AssumeRole event for a given role.
● Console sign-in events do not capture STS role assumptions, so many role uses would be missed.
● CloudTrail Lake relies on scheduled queries and is not real-time for per-event alerting.
Workflow: EventBridge rule on CloudTrail sts:AssumeRole for the role → invoke Lambda to SNS.
an Amazon EventBridge rule that matches AWS_RISK_CREDENTIALS_EXPOSED from AWS Health and start
AWS Health emits the exposure event and Step Functions provides auditable, reliable orchestration across IAM, CloudTrail, and notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health emits the exposure event and Step Functions provides auditable, reliable orchestration across IAM, CloudTrail, and notifications.
● Scene fit: AWS Health emits the exposure event and Step Functions provides auditable, reliable orchestration across IAM, CloudTrail, and notifications.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Points the rule at CloudTrail, but the exposure alert is generated by AWS Health, so the event pattern would.
● A single Lambda can perform actions, but it lacks native long-running orchestration, retries, and detailed stateful execution history for.
● Security Hub does not originate the AWS_RISK_CREDENTIALS_EXPOSED event and an SSM runbook alone does not address the specific AWS.
Workflow: Create an Amazon EventBridge rule that matches AWS_RISK_CREDENTIALS_EXPOSED from AWS Health and start an AWS Step Functions state machine that coordinates IAM, CloudTrail, and Amazon SNS.
Publish a new version and update the existing prod alias to route
Weighted aliases enable canary releases and quick rollback without modifying the EC2 application.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Weighted aliases enable canary releases and quick rollback without modifying the EC2 application.
● Scene fit: Weighted aliases enable canary releases and quick rollback without modifying the EC2 application.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Introduces a new component and requires client changes, which is unnecessary for direct Lambda canary releases.
● Route 53 cannot weight native Lambda invocations and this approach adds avoidable operational complexity.
● Layers are dependency bundles and aliases route to function versions, not layers.
Workflow: Publish a new version and update the existing prod alias to route 15% of traffic to the new version and 85% to the current version.
Publish vetted CloudFormation products in AWS Service Catalog
Service Catalog lets you offer versioned products backed by approved templates and apply constraints to enforce parameters, tags, and regional distribution.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Service Catalog lets you offer versioned products backed by approved templates and apply constraints to enforce parameters, tags, and regional distribution.
● Scene fit: Service Catalog lets you offer versioned products backed by approved templates and apply constraints to enforce parameters, tags, and regional distribution.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● StackSets coordinates deployments across accounts and regions but does not inherently enforce mandatory tags or limit where developers can launch stacks.
● Trusted Advisor provides best-practice checks but cannot block or govern CloudFormation usage, tag policies, or region constraints.
● Drift detection is a post-deployment check that cannot prevent out-of-policy resources from being created in the first place.
Workflow: Publish vetted CloudFormation products in AWS Service Catalog.
Amazon DynamoDB Global Tables
DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform local low-latency reads and writes with data replicated worldwide.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform local low-latency reads and writes with data replicated worldwide.
● Scene fit: DynamoDB Global Tables provide multi-Region, multi-active replication so each Region can perform local low-latency reads and writes with data replicated worldwide.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DAX is an in-memory cache for DynamoDB that improves read latency within a Region but does not provide multi-Region, multi-active writes or replication.
● Aurora Global Database uses a single write Region with cross-Region read replicas, so it is not multi-active for writes.
● RDS cross-Region read replicas are read-only; all writes go to a single primary Region, increasing write latency elsewhere.
Workflow: Amazon DynamoDB Global Tables.
an EventBridge rule for restricted-ssh with complianceType NON_COMPLIANT, use an input transformer
EventBridge can match the exact Config rule and compliance state, then format fields already present in the event for SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can match the exact Config rule and compliance state, then format fields already present in the event for SNS.
● Scene fit: EventBridge can match the exact Config rule and compliance state, then format fields already present in the event for SNS.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● The event pattern is too broad and SNS filtering is configured on subscriptions, not at the topic level.
● Is technically workable and may be needed if external lookup data is required, but the question only needs event-field formatting.
● Sends compliant and noncompliant events downstream and relies on later filtering.
Workflow: Create an EventBridge rule for restricted-ssh with complianceType NON_COMPLIANT → use an input transformer to add the group name and ID, and publish to SNS.
a CloudWatch Logs metric filter that matches CRITICAL, emit a custom metric
A CloudWatch Logs metric filter turns matching log lines into a custom metric, which a CloudWatch alarm can monitor and notify the team via SNS with.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A CloudWatch Logs metric filter turns matching log lines into a custom metric, which a CloudWatch alarm can.
● Scene fit: A CloudWatch Logs metric filter turns matching log lines into a custom metric, which a CloudWatch alarm can.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Transit Gateway flow logs capture IP flow information for the transit gateway and not the firewall appliance's log entries.
● EventBridge does not directly inspect raw CloudWatch Logs content without an intermediary such as a subscription filter or metric.
Workflow: Create a CloudWatch Logs metric filter that matches CRITICAL → emit a custom metric, and configure a CloudWatch alarm to publish to an Amazon SNS topic subscribed by the security team.
AWS Systems Manager Automation runbooks triggered by AWS CodePipeline via Amazon EventBridge
This uses managed services to automate AMI creation and stores identifiers centrally with minimal cost and overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses managed services to automate AMI creation and stores identifiers centrally with minimal cost and overhead.
● Scene fit: This uses managed services to automate AMI creation and stores identifiers centrally with minimal cost and overhead.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Increases cost and management overhead by operating Jenkins on EC2 and maintaining DynamoDB for simple lookups.
● Adds unnecessary complexity with OVF handling and EC2 build hosts when native AMI automation is available.
● Approach has many moving parts and stores metadata in S3 instead of a parameter store designed for configuration values.
Workflow: Use AWS Systems Manager Automation runbooks triggered by AWS CodePipeline via Amazon EventBridge to bake AMIs and publish the AMI IDs to Systems Manager Parameter Store.
Pilot light in another Region with cross-Region RDS read replica, minimal app
Asynchronous RDS replication supports minutes-level RPO and keeping most compute off minimizes cost while meeting hours-level RTO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Asynchronous RDS replication supports minutes-level RPO and keeping most compute off minimizes cost while meeting hours-level RTO.
● Scene fit: Asynchronous RDS replication supports minutes-level RPO and keeping most compute off minimizes cost while meeting hours-level RTO.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Snapshot copy and full environment rebuild often miss a 15-minute RPO and can exceed a 3-hour RTO.
● Meets RTO but costs more than necessary for the stated budget constraints.
● Protects from AZ failure but not a Regional disaster and does not meet cross-Region DR requirements.
Workflow: Pilot light in another Region with cross-Region RDS read replica, minimal app stack → Route 53 failover, and promote replica on disaster.
an Amazon EventBridge scheduled rule that directly targets EBS Create Snapshot for
EventBridge can invoke the built-in EBS Create Snapshot target on a schedule, providing the most direct and minimal-configuration solution.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can invoke the built-in EBS Create Snapshot target on a schedule, providing the most direct and minimal-configuration solution.
● Scene fit: EventBridge can invoke the built-in EBS Create Snapshot target on a schedule, providing the most direct and minimal-configuration solution.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Approach works but adds an extra Systems Manager step instead of using the native EventBridge EBS target.
● AWS Backup can do this but requires additional constructs like backup plans and vaults, which is heavier than a single EventBridge rule.
● Adds code, packaging, and IAM permissions management, making it more complex than a direct EventBridge EBS target.
Workflow: Create an Amazon EventBridge scheduled rule that directly targets EBS Create Snapshot for the specified volume IDs at 1:00 AM UTC.
Inspector, EC2 Auto Scaling across four AZs behind ALB, Aurora, Route 53
Inspector provides continuous vulnerability and exposure scanning; Multi-AZ ALB plus an alias apex record delivers HA and correct root-domain mapping.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Inspector provides continuous vulnerability and exposure scanning; Multi-AZ ALB plus an alias apex record delivers HA and correct root-domain mapping.
● Scene fit: Inspector provides continuous vulnerability and exposure scanning; Multi-AZ ALB plus an alias apex record delivers HA and correct root-domain mapping.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● GuardDuty is threat detection, not vulnerability scanning, and a CNAME cannot be used at the zone apex for an ALB.
● Security Hub aggregates and correlates findings but does not perform vulnerability scanning.
● Macie focuses on sensitive data discovery and a non-alias A record cannot target an ALB at the apex.
Workflow: Inspector, EC2 Auto Scaling across four AZs behind ALB, Aurora → Route 53 alias apex to ALB.
the CodePipeline stage to run each function's actions concurrently by assigning the
Using the same runOrder for multiple actions in a stage makes them run in parallel, reducing total pipeline wall-clock time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using the same runOrder for multiple actions in a stage makes them run in parallel, reducing total pipeline wall-clock time.
● Scene fit: Using the same runOrder for multiple actions in a stage makes them run in parallel, reducing total pipeline wall-clock time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Docker layer caching can accelerate container image rebuilds but does not make sequential CodePipeline actions execute in parallel.
● Larger build instances may speed up individual builds but the pipeline still completes serially across functions.
● Adds extra orchestration complexity outside CodePipeline even though native parallel actions already solve the problem.
Workflow: Configure the CodePipeline stage to run each function's actions concurrently by assigning the same runOrder value.
Organizations SCP denying iam:CreateUser except approved principals via aws:PrincipalArn condition
An SCP with an explicit Deny on iam:CreateUser that allows only listed caller ARNs enforces a preventive, organization-wide control.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An SCP with an explicit Deny on iam:CreateUser that allows only listed caller ARNs enforces a preventive, organization-wide control.
● Scene fit: An SCP with an explicit Deny on iam:CreateUser that allows only listed caller ARNs enforces a preventive, organization-wide control.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Local IAM policies are mutable by account admins, do not constrain the root user, and are not guaranteed as an org-wide preventive control.
● Is reactive and cannot block the CreateUser API call from succeeding.
● CreateLoginProfile manages console passwords and does not prevent creating IAM users.
Workflow: Organizations SCP denying iam:CreateUser except approved principals via aws:PrincipalArn condition.
Immutable updates in Elastic Beanstalk
Launches a parallel Auto Scaling group with the new version, keeps existing instances serving, and enables quick rollback by terminating the new group if unhealthy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Launches a parallel Auto Scaling group with the new version, keeps existing instances serving, and enables quick rollback by terminating the new group if unhealthy.
● Scene fit: Launches a parallel Auto Scaling group with the new version, keeps existing instances serving, and enables quick rollback by terminating the new group if unhealthy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Still an in-place update that can fail mid-rollout and does not provide immediate rollback to a healthy parallel fleet.
● Blue/green typically requires decoupling the database; keeping RDS attached to the environment complicates or blocks the cutover.
● Not an Elastic Beanstalk deployment policy and in-place updates lack a parallel fleet for rapid rollback.
Workflow: Immutable updates in Elastic Beanstalk.
a terminating lifecycle hook to place instances in Terminating:Wait, create an EventBridge
Using Terminating:Wait with EventBridge and SSM Automation ensures a controlled delay to flush logs to CloudWatch Logs before the instance is finally terminated.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using Terminating:Wait with EventBridge and SSM Automation ensures a controlled delay to flush logs to CloudWatch Logs before.
● Scene fit: Using Terminating:Wait with EventBridge and SSM Automation ensures a controlled delay to flush logs to CloudWatch Logs before.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Can stream logs but does not guarantee final log flush before immediate termination triggered by health checks, so logs.
● Pending:Wait applies to launch events, not scale-in or termination, so this does not solve the termination timing issue.
● Terminate Successful fires after the instance is deleted, so it is too late to retrieve logs from the instance.
Workflow: Add a terminating lifecycle hook to place instances in Terminating:Wait → create an EventBridge rule for the EC2 Instance-terminate lifecycle action → attach a Systems Manager Automation document to trigger the CloudWatch.
Secrets Manager for RDS; EC2 instance profile reads secret and calls DynamoDB
Store and optionally rotate the RDS credential in Secrets Manager and grant the EC2 role permissions to read the secret and access DynamoDB via IAM.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Store and optionally rotate the RDS credential in Secrets Manager and grant the EC2 role permissions to read the secret and access DynamoDB via IAM.
● Scene fit: Store and optionally rotate the RDS credential in Secrets Manager and grant the EC2 role permissions to read the secret and access DynamoDB via IAM.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DynamoDB uses IAM, not database-like credentials, so storing nonexistent DynamoDB credentials is incorrect.
● Parameter Store lacks native rotation for RDS and DynamoDB should be accessed via IAM without storing credentials.
● Hardcoding secrets in user data is insecure and user data can be retrieved; use a managed secrets service instead.
Workflow: Use Secrets Manager for RDS → EC2 instance profile reads secret and calls DynamoDB.
DynamoDB global tables with regional replicas
Provides native multi-Region, multi-writer replication with automatic propagation and conflict resolution for low-latency local reads and writes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides native multi-Region, multi-writer replication with automatic propagation and conflict resolution for low-latency local reads and writes.
● Scene fit: Provides native multi-Region, multi-writer replication with automatic propagation and conflict resolution for low-latency local reads and writes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Aurora Global Database is primarily single-writer across Regions; cross-Region writes flow to a primary Region and do not enable true multi-Region multi-writer.
● Is a cache with asynchronous cross-Region replication and a single primary Region, not a durable multi-writer data store.
● DIY replication via streams and Lambda is complex, fragile at scale, and lacks built-in conflict resolution.
Workflow: DynamoDB global tables with regional replicas.
Cross-account CloudWatch Logs destination feeding Kinesis Data Firehose to Amazon S3
Cross-account CloudWatch Logs subscriptions can deliver to Firehose, which writes to S3 for centralized, secure, low-cost archival.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-account CloudWatch Logs subscriptions can deliver to Firehose, which writes to S3 for centralized, secure, low-cost archival.
● Scene fit: Cross-account CloudWatch Logs subscriptions can deliver to Firehose, which writes to S3 for centralized, secure, low-cost archival.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Export tasks are periodic and not near real-time, and they do not provide continuous cross-account streaming.
● Redshift is for analytics and is costly for archival, not intended for inexpensive long-term log storage.
● Lambda-to-EFS adds complexity and EFS is not optimized for low-cost archival compared to S3.
Workflow: Cross-account CloudWatch Logs destination feeding Kinesis Data Firehose to Amazon S3.
Amazon FSx for NetApp ONTAP in the primary Region and the DR
FSx for NetApp ONTAP supports both SMB and NFS on the same platform and uses SnapMirror for efficient, incremental cross-Region replication ideal for pilot-light DR.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: FSx for NetApp ONTAP supports both SMB and NFS on the same platform and uses SnapMirror for efficient.
● Scene fit: FSx for NetApp ONTAP supports both SMB and NFS on the same platform and uses SnapMirror for efficient.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Splits SMB and NFS across two different file systems and relies on batch copies, which does not provide a.
● FSx for Lustre does not support SMB and AWS Backup provides periodic copies rather than efficient near-continuous cross-Region replication.
● Amazon S3 is object storage without native SMB or NFS mounts, and attempting to mount it introduces compatibility and.
Workflow: Deploy Amazon FSx for NetApp ONTAP in the primary Region and the DR Region and configure SnapMirror for cross-Region replication.
a Lambda-backed CloudFormation custom resource to resolve the latest AMI ID and
A Lambda-backed custom resource runs only during stack create or update to dynamically fetch the latest AMI and is cost-efficient.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A Lambda-backed custom resource runs only during stack create or update to dynamically fetch the latest AMI and is cost-efficient.
● Scene fit: A Lambda-backed custom resource runs only during stack create or update to dynamically fetch the latest AMI and is cost-efficient.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Introduces frequent scheduled execution and template churn when nothing changes, adding unnecessary cost and complexity.
● Cfn-init runs on the instance after launch and cannot alter the AMI that was already used to create the instance.
● Keeping an instance running for periodic checks is unnecessary and increases cost and operational overhead.
Workflow: Use a Lambda-backed CloudFormation custom resource to resolve the latest AMI ID and pass it into the launch template.
an AWS WAF web ACL with managed rules to the ALB +
AWS WAF inspects HTTP(S) requests on the ALB and blocks patterns like SQLi and XSS. Terminating TLS at the ALB with an ACM cert secures clients and offloads encryption from instances.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS WAF inspects HTTP(S) requests on the ALB and blocks patterns like SQLi and XSS.
● Scene fit: AWS WAF inspects HTTP(S) requests on the ALB and blocks patterns like SQLi and XSS. Terminating TLS at.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Terminating TLS on instances adds CPU overhead and violates the requirement to avoid host processing.
● Network Firewall operates at the VPC layer and is not the primary control for ALB L7 web exploit filtering.
● Shield Advanced focuses on DDoS mitigation, not application-layer exploit filtering.
Workflow: Attach an AWS WAF web ACL with managed rules to the ALB → Use an ACM certificate on an ALB HTTPS listener for TLS offload.
an EventBridge rule for Auto Scaling instance launch and terminate to invoke
EventBridge can invoke Lambda on Auto Scaling events; Lambda can modify S3 content and log details to CloudWatch Logs for search.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can invoke Lambda on Auto Scaling events; Lambda can modify S3 content and log details to CloudWatch Logs for search.
● Scene fit: EventBridge can invoke Lambda on Auto Scaling events; Lambda can modify S3 content and log details to CloudWatch.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Records events but does not update the S3 page in real time and relies on manual or scheduled exports.
● S3 is not a supported direct target for EventBridge, so you cannot update the page this way.
● Is not event driven and introduces delay; S3 alone is not natively searchable without extra tooling.
Workflow: Use an EventBridge rule for Auto Scaling instance launch and terminate to invoke Lambda → Lambda updates the S3 page and writes events to CloudWatch Logs.
External Lambda extension emitting X-Ray segments/subsegments, active tracing, X-Ray groups and Insights
External extensions with the Telemetry API can emit X-Ray segments/subsegments; X-Ray Insights detects anomalies and EventBridge/CloudWatch provide fast notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: External extensions with the Telemetry API can emit X-Ray segments/subsegments; X-Ray Insights detects anomalies and EventBridge/CloudWatch provide fast notifications.
● Scene fit: External extensions with the Telemetry API can emit X-Ray segments/subsegments; X-Ray Insights detects anomalies and EventBridge/CloudWatch provide fast notifications.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Internal extensions are less reliable for exporting telemetry and CloudWatch Logs Insights does not provide trace-level anomaly detection.
● ADOT and ServiceLens do not replace X-Ray Insights for trace anomaly detection and Contributor Insights analyzes log patterns, not trace anomalies.
● Collapsing functions removes cross-function visibility and is unnecessary when X-Ray can correlate distributed traces across multiple functions.
Workflow: External Lambda extension emitting X-Ray segments/subsegments, active tracing, X-Ray groups and Insights, alerts via EventBridge and CloudWatch.
Assume a deployment role in the cluster account via sts:AssumeRole, grant EKS
This enables the CodeBuild role to assume a cross-account role with EKS permissions and authorizes that role in Kubernetes via aws-auth (or access entries).
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This enables the CodeBuild role to assume a cross-account role with EKS permissions and authorizes that role in.
● Scene fit: This enables the CodeBuild role to assume a cross-account role with EKS permissions and authorizes that role in.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● IRSA is for pods running in the cluster and does not apply to CodeBuild; cross-account access and Kubernetes authorization are still required.
● The trust relationship must be on the target role in the cluster account that is being assumed, not on the DevOps role.
● CodeBuild does not assume roles via web identity, and EKS access still requires authorization through aws-auth or access entries.
Workflow: Assume a deployment role in the cluster account via sts:AssumeRole → grant EKS permissions, and map the role in aws-auth.
an Amazon EventBridge rule for CodePipeline state changes that invokes an AWS
This ties pipeline events to an automated runbook that programmatically manages both compute and database resources without altering the architecture.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This ties pipeline events to an automated runbook that programmatically manages both compute and database resources without altering the architecture.
● Scene fit: This ties pipeline events to an automated runbook that programmatically manages both compute and database resources without altering.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Changes pricing models and relies on manual scripting, and a Reserved Instance is inefficient for infrequent usage.
● Introduces an architectural change and still lacks a direct, event-driven trigger from CodePipeline.
● Is time-based rather than event-driven and may run when no pipeline tests are executing or miss ad hoc runs.
Workflow: Configure an Amazon EventBridge rule for CodePipeline state changes that invokes an AWS Systems Manager Automation runbook to start and stop the EC2 instances and the RDS DB instance at the beginning and end of the tests.
Bucket policy denies or omits required access + IAM role for the
A restrictive or incorrect bucket policy, including explicit Deny or unmet conditions, will return 403 Access Denied. Missing permissions, a permissions boundary, or an SCP Deny can cause S3 to return 403 Access Denied.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A restrictive or incorrect bucket policy, including explicit Deny or unmet conditions, will return 403 Access Denied.
● Scene fit: A restrictive or incorrect bucket policy, including explicit Deny or unmet conditions, will return 403 Access Denied. Missing.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Network egress blocks usually cause timeouts or connection errors rather than an authorization 403 from S3.
● S3 default encryption is transparent for authorized callers and does not by itself block access.
● Block Public Access restricts public access and does not block access for an authorized IAM role.
Workflow: Bucket policy denies or omits required access → IAM role for the instance lacks or is prevented from s3:GetObject.
an IAM role to instances granting SSM access + Add SSM interface
The SSM Agent needs an instance profile that allows it to call Systems Manager APIs. Create interface endpoints for ssm, ssmmessages, and ec2messages to keep traffic on AWS PrivateLink.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The SSM Agent needs an instance profile that allows it to call Systems Manager APIs.
● Scene fit: The SSM Agent needs an instance profile that allows it to call Systems Manager APIs. Create interface endpoints for ssm, ssmmessages, and ec2messages to keep traffic on AWS PrivateLink.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Client VPN is unrelated to keeping Session Manager traffic private and does not replace the need for VPC endpoints.
● Systems Manager uses interface VPC endpoints, not gateway endpoints.
● A NAT gateway enables internet egress, which violates the requirement.
Workflow: Attach an IAM role to instances granting SSM access → Add SSM interface VPC endpoints in the VPC.
Forward auth logs with CloudWatch agent, add a CloudWatch Logs metric filter
Metric filters turn matching log lines into metrics, enabling CloudWatch alarms to deliver near real-time SNS notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric filters turn matching log lines into metrics, enabling CloudWatch alarms to deliver near real-time SNS notifications.
● Scene fit: Metric filters turn matching log lines into metrics, enabling CloudWatch alarms to deliver near real-time SNS notifications.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Can work in near real time but adds custom code and operational overhead compared to native metric filters and alarms.
● CloudTrail records AWS API activity, not OS-level host logins, and pipeline latency can exceed the target window.
● ConsoleLogin indicates AWS account sign-in events, not Linux or Windows logins on EC2 hosts.
Workflow: Forward auth logs with CloudWatch agent → add a CloudWatch Logs metric filter for successful logins, and trigger a CloudWatch alarm to SNS.
AWS CloudTrail organization trail with EventBridge to SNS + AWS Config organization
An organization-level trail records Organizations API calls and EventBridge can route matching events to SNS for near-real-time alerts. Aggregates config data across accounts and can evaluate rules to alert on detected changes across the organization.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An organization-level trail records Organizations API calls and EventBridge can route matching events to SNS for near-real-time alerts.
● Scene fit: An organization-level trail records Organizations API calls and EventBridge can route matching events to SNS for near-real-time alerts.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● These aggregate and investigate security findings, not AWS Organizations membership or account change events.
● Control Tower governs landing zones but is not the authoritative or comprehensive source for Organizations membership change alerts.
● Useful for threat detection and log analytics but not targeted at precise AWS Organizations change events.
Workflow: AWS CloudTrail organization trail with EventBridge to SNS → AWS Config organization aggregator with rules to SNS or EventBridge.
In the ECS service, keep Platform version set to LATEST and choose
For services using LATEST, forcing a new deployment relaunches tasks onto the most current supported Fargate platform version.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: For services using LATEST, forcing a new deployment relaunches tasks onto the most current supported Fargate platform version.
● Scene fit: For services using LATEST, forcing a new deployment relaunches tasks onto the most current supported Fargate platform version.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Fargate platform version is not part of a task definition and cannot be specified via a task definition ARN.
● ECS provides no automatic upgrade switch for Fargate platform versions.
● While possible, creating a new service and blue/green shifting is unnecessary complexity compared to forcing a new deployment.
Workflow: In the ECS service → keep Platform version set to LATEST and choose Force new deployment to restart tasks on the newer runtime.
CAPABILITY_IAM or CAPABILITY_NAMED_IAM to the CloudFormation action in CodePipeline
CloudFormation requires explicit acknowledgement to create or modify IAM resources, otherwise it raises InsufficientCapabilitiesException.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation requires explicit acknowledgement to create or modify IAM resources, otherwise it raises InsufficientCapabilitiesException.
● Scene fit: CloudFormation requires explicit acknowledgement to create or modify IAM resources, otherwise it raises InsufficientCapabilitiesException.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A circular dependency causes a dependency cycle error, not InsufficientCapabilitiesException.
● PassRole fixes AccessDenied when passing roles but does not satisfy the required IAM capability acknowledgement.
● Broader permissions do not bypass the need to acknowledge IAM capabilities in CloudFormation.
Workflow: Add CAPABILITY_IAM or CAPABILITY_NAMED_IAM to the CloudFormation action in CodePipeline.
the same runOrder for actions in a single stage to run them
Actions within the same stage that share the same runOrder execute concurrently, reducing overall duration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Actions within the same stage that share the same runOrder execute concurrently, reducing overall duration.
● Scene fit: Actions within the same stage that share the same runOrder execute concurrently, reducing overall duration.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Coordinates dependent builds but does not make separate CodePipeline actions run concurrently.
● May speed individual builds but leaves the sequential pipeline as the bottleneck.
● Adds complexity and an external orchestrator without improving CodePipeline action parallelism.
Workflow: Set the same runOrder for actions in a single stage to run them in parallel.
Provision interface VPC endpoints for Systems Manager in the VPC to keep
Interface VPC endpoints enable private connectivity to Systems Manager APIs and data channels without using the internet. Instances need an IAM instance profile granting Systems Manager.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Interface VPC endpoints enable private connectivity to Systems Manager APIs and data channels without using the internet.
● Scene fit: Interface VPC endpoints enable private connectivity to Systems Manager APIs and data channels without using the internet. Instances.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A bastion host still relies on SSH and key pairs and does not enforce private Session Manager access.
● An EC2 API endpoint does not provide the private data channels required by Session Manager.
● Session Manager does not require opening SSH port 22 because it operates without inbound ports.
Workflow: Provision interface VPC endpoints for Systems Manager in the VPC to keep Session Manager traffic private → Associate an IAM instance profile with each instance that includes permissions such as AmazonSSMManagedInstanceCore.
SSM Automation golden AMI + user data reads env tag + Parameter
SSM Automation supports golden AMI workflows to pre-bake software, a bootstrap can read the environment tag to load settings, and Parameter Store SecureString with KMS secures secrets.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SSM Automation supports golden AMI workflows to pre-bake software, a bootstrap can read the environment tag to load settings, and Parameter Store SecureString with KMS secures secrets.
● Scene fit: SSM Automation supports golden AMI workflows to pre-bake software, a bootstrap can read the environment tag to load.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Session Manager does not create or bake AMIs, and adding Lambda from user data increases complexity without improving startup time.
● Patch Manager targets OS patching rather than application baking and AppConfig is for configuration flags, not secrets storage.
● Installing on first boot via cfn-init increases launch time and does not leverage a pre-baked AMI to reduce startup duration.
Workflow: SSM Automation golden AMI + user data reads env tag + Parameter Store SecureString.
CloudFormation nested stacks with one-instance Auto Scaling groups that reattach designated ENIs
This pattern gives self-healing, stable private IPs via ENI reuse, and minimal custom code using IaC.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This pattern gives self-healing, stable private IPs via ENI reuse, and minimal custom code using IaC.
● Scene fit: This pattern gives self-healing, stable private IPs via ENI reuse, and minimal custom code using IaC.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is ad-hoc and operationally heavy, lacking built-in self-healing and repeatability.
● Beanstalk abstracts instances and does not provide deterministic ENI reuse or fixed hostnames per node.
● Possible but adds custom logic and maintenance overhead versus a simpler native CloudFormation pattern.
Workflow: CloudFormation nested stacks with one-instance Auto Scaling groups that reattach designated ENIs and set hostnames at boot.
AWS CloudTrail data events for the bucket, store the logs in Amazon
CloudTrail data events capture S3 object-level API calls and storing them in S3 with Athena provides low-cost, on-demand search.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail data events capture S3 object-level API calls and storing them in S3 with Athena provides low-cost, on-demand search.
● Scene fit: CloudTrail data events capture S3 object-level API calls and storing them in S3 with Athena provides low-cost, on-demand search.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Running an OpenSearch domain adds ongoing cost and S3 access logs are not the most precise source for API-level auditing.
● EventBridge does not receive S3 object-level API events unless CloudTrail is configured to emit them.
● S3 cannot natively push object-level API logs directly to CloudWatch Logs.
Workflow: Enable AWS CloudTrail data events for the bucket → store the logs in Amazon S3, and query them ad hoc with Amazon Athena.
a single ALB with two target groups mapped to the two ASGs
Switching target groups on a single ALB removes DNS from the cutover path, reduces components and cost, and enables fast blue/green transitions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Switching target groups on a single ALB removes DNS from the cutover path, reduces components and cost, and.
● Scene fit: Switching target groups on a single ALB removes DNS from the cutover path, reduces components and cost, and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Can mask DNS issues with static anycast IPs, but it adds cost and complexity compared to simplifying the existing.
● Decreasing TTL does not fix clients that ignore or pin DNS responses and will still result in calls to.
● Instance-level proxies add operational overhead and complexity while creating fragile traffic paths that are hard to manage.
Workflow: Use a single ALB with two target groups mapped to the two ASGs → deploy to the idle ASG, → flip the ALB listener rule to the new target group while keeping.
EventBridge rule triggers Lambda to promote the replica and update Systems Manager
EventBridge detects RDS or Aurora failure events, Lambda automates promotion and updates a central endpoint value that clients read on reconnect.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge detects RDS or Aurora failure events, Lambda automates promotion and updates a central endpoint value that clients read on reconnect.
● Scene fit: EventBridge detects RDS or Aurora failure events, Lambda automates promotion and updates a central endpoint value that clients read on reconnect.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Global Database can switch write Region but does not auto-detect failure or update client endpoints; you still need orchestration.
● DNS failover cannot promote the replica and Aurora endpoints change on promotion, so DNS alone is insufficient.
● RDS Proxy is Regional and does not handle cross-Region promotion or endpoint switching.
Workflow: EventBridge rule triggers Lambda to promote the replica and update Systems Manager Parameter Store.
Scale-out occurred during deployment; new instances launched with the last successful revision
When the Auto Scaling group adds instances during a deployment, they are initialized with the most recent successful revision, leading to a mixed fleet even though the deployment completes successfully.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: When the Auto Scaling group adds instances during a deployment, they are initialized with the most recent successful revision, leading to a mixed fleet even though the deployment completes successfully.
● Scene fit: When the Auto Scaling group adds instances during a deployment, they are initialized with the most recent successful.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A stale launch template affects how instances boot, but CodeDeploy would still update targeted instances or fail the deployment.
● Missing IAM permissions would cause the deployment to fail, not succeed with mixed versions.
● Minimum healthy hosts controls batch availability, but a successful deployment still updates all targeted instances.
Workflow: Scale-out occurred during deployment → new instances launched with the last successful revision.
Apex latency alias to na.example.com and eu.example.com; each subdomain uses failover with
Latency decides Region; per-Region failover swaps to the other ALB on outage.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency decides Region; per-Region failover swaps to the other ALB on outage.
● Scene fit: Latency decides Region; per-Region failover swaps to the other ALB on outage.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Returns multiple healthy records; not latency-aware and no Regional failover coordination.
● Reverses the needed layering; does not give per-Region failover beneath latency.
● Removes unhealthy targets but lacks explicit cross-Region failover branches and can miss some Regional failure modes.
Workflow: Apex latency alias to na.example.com and eu.example.com → each subdomain uses failover with in-Region ALB primary and the other Region as secondary.
AWS CodeDeploy with a blue green deployment group for the Auto Scaling
CodeDeploy blue green supports shifting traffic to the new fleet and automatically terminating the original fleet after a configurable wait time such as 90 minutes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy blue green supports shifting traffic to the new fleet and automatically terminating the original fleet after a.
● Scene fit: CodeDeploy blue green supports shifting traffic to the new fleet and automatically terminating the original fleet after a.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFormation does not support time based retention or deletion policies for resources like ALB, and this does not orchestrate.
● Elastic Beanstalk application version lifecycle works on versions with minimum ages in days and does not schedule environment termination.
● Relies on custom orchestration and timers rather than using a native timed termination feature, making it more complex and.
Workflow: Use AWS CodeDeploy with a blue green deployment group for the Auto Scaling group and set BlueInstanceTerminationOption action to TERMINATE with terminationWaitTimeInMinutes set to 90.
AWS CloudFormation to manage Lambda versions and API Gateway stage canary at
CloudFormation can manage the API Gateway stage CanarySetting with a 0.15 traffic percentage and the related Lambda deployment configuration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation can manage the API Gateway stage CanarySetting with a 0.15 traffic percentage and the related Lambda deployment configuration.
● Scene fit: CloudFormation can manage the API Gateway stage CanarySetting with a 0.15 traffic percentage and the related Lambda deployment configuration.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CodeDeploy can shift Lambda alias traffic, but API Gateway canary configuration remains separate.
● Route 53 weighted routing works at DNS endpoint level and is affected by caching and TTL.
● Usage plans control API key quotas and throttling.
Workflow: Use AWS CloudFormation to manage Lambda versions and API Gateway stage canary at 15%, → promote.
Pre-provision a second region with CloudFormation, use cross-region RDS read replica, enable
Cross-region read replicas and S3 CRR minimize RPO, and having infrastructure pre-provisioned with warm capacity enables fast promotion and low RTO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-region read replicas and S3 CRR minimize RPO, and having infrastructure pre-provisioned with warm capacity enables fast promotion and low RTO.
● Scene fit: Cross-region read replicas and S3 CRR minimize RPO, and having infrastructure pre-provisioned with warm capacity enables fast promotion and low RTO.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Infrequent snapshot copies and Glacier retrievals increase RPO and RTO due to manual restore and retrieval delays.
● Failing over compute without cross-region database replication leaves no up-to-date database in the secondary region.
● RDS Multi-AZ is confined to a single region and cannot fail over across regions.
Workflow: Pre-provision a second region with CloudFormation → use cross-region RDS read replica → enable S3 CRR, warm ASG → promote on failover.
v1 and v2 stages with one Lambda alias; v1 mapping injects "color":"none"
Both stages invoke a single Lambda via an alias, and the legacy stage uses a mapping template to default the field for backward compatibility.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Both stages invoke a single Lambda via an alias, and the legacy stage uses a mapping template to default the field for backward compatibility.
● Scene fit: Both stages invoke a single Lambda via an alias, and the legacy stage uses a mapping template to default the field for backward compatibility.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Introduces two Lambda functions and proxy logic, violating the single-backend goal and adding operational overhead.
● Lambda authorizers cannot modify the request body; they only supply auth context.
● Models and validators validate shape and types but do not mutate or inject fields into the payload.
Workflow: Use v1 and v2 stages with one Lambda alias → v1 mapping injects "color":"none".
Maintain separate templates per logical domain and deploy each stack as it
This follows CloudFormation best practices by decoupling teams, enabling independent deployments, and using exports/imports for safe cross-stack references.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This follows CloudFormation best practices by decoupling teams, enabling independent deployments, and using exports/imports for safe cross-stack references.
● Scene fit: This follows CloudFormation best practices by decoupling teams, enabling independent deployments, and using exports/imports for safe cross-stack references.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Tightly couples releases across teams because any change forces a parent stack update, increasing blast radius and coordination overhead.
● A monolithic template couples all components, inflates failure blast radius, and blocks teams that are not ready from those that are.
● AWS Proton standardizes service and environment templates but does not replace the need for decoupled CloudFormation stacks with explicit cross-stack references.
Workflow: Maintain separate templates per logical domain and deploy each stack as it is ready, sharing values via cross-stack exports and imports and managing each template independently in version control.
AWS Config S3 managed rules with Systems Manager Automation remediation
Continuously evaluates S3 buckets against managed rules and auto-remediates drift via SSM Automation runbooks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously evaluates S3 buckets against managed rules and auto-remediates drift via SSM Automation runbooks.
● Scene fit: Continuously evaluates S3 buckets against managed rules and auto-remediates drift via SSM Automation runbooks.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Prevents certain actions and blocks public access but cannot enable encryption, logging, or versioning or fix existing buckets.
● Advisory checks lack granular, continuous per-bucket compliance and do not provide native auto-remediation.
● Custom event-driven logic can react to API calls but is complex, not configuration-compliance centric, and misses continuous evaluation.
Workflow: AWS Config S3 managed rules with Systems Manager Automation remediation.
a single AWS Lambda alias that references the current and new versions
Lambda aliases support routing configuration to weight traffic between up to two published versions and allow fast rollback. API Gateway canary settings natively support weighted traffic.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda aliases support routing configuration to weight traffic between up to two published versions and allow fast rollback.
● Scene fit: Lambda aliases support routing configuration to weight traffic between up to two published versions and allow fast rollback.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Network Load Balancers do not support weighted routing across target groups, so you cannot split 15% of traffic this.
● Failover routing is health-based primary and secondary behavior, not percentage-based traffic splitting.
● There is no built-in Canary routing option in Application Load Balancer.
Workflow: Use a single AWS Lambda alias that references the current and new versions → shift 15% of traffic to the new version and later route 100% when it proves stable → Enable.
AWS CloudTrail to send events to Amazon CloudWatch Logs, deploy the CloudWatch
Placing both CloudTrail events and instance logs in CloudWatch Logs enables CloudWatch Logs Insights to query across multiple log groups in one place.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Placing both CloudTrail events and instance logs in CloudWatch Logs enables CloudWatch Logs Insights to query across multiple.
● Scene fit: Placing both CloudTrail events and instance logs in CloudWatch Logs enables CloudWatch Logs Insights to query across multiple.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● The CloudWatch Agent cannot send logs directly to Amazon S3, so this pipeline is not supported end to end.
● CloudTrail does not natively deliver to Kinesis Data Streams and the CloudWatch Agent does not publish to Kinesis Data.
● Athena can query S3 objects but cannot directly query CloudWatch Logs, which prevents a single query interface across both.
Workflow: Configure AWS CloudTrail to send events to Amazon CloudWatch Logs → deploy the CloudWatch Agent on EC2 to push application logs to CloudWatch Logs, and run CloudWatch Logs Insights queries across both.
the application tier in a different AWS Region without RDS, create a
A cross-Region read replica with a cloned app tier and Route 53 failover provides a warm standby that meets a 10-minute RPO and 90-minute RTO with.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A cross-Region read replica with a cloned app tier and Route 53 failover provides a warm standby that.
● Scene fit: A cross-Region read replica with a cloned app tier and Route 53 failover provides a warm standby that.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● RDS Multi-AZ is confined to a single Region, so placing a standby in another Region is not supported and.
● An additional Availability Zone does not protect against regional outages and therefore does not meet the geographically isolated requirement.
Workflow: Deploy the application tier in a different AWS Region without RDS → create a cross-Region RDS MySQL read replica, point the DR app to the local replica, and use Route 53 failover.
WAF web ACL on API Gateway stage with managed SQLi rules; track
Directly protects API Gateway with AWS Managed Rules for SQLi and records configuration changes with AWS Config.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Directly protects API Gateway with AWS Managed Rules for SQLi and records configuration changes with AWS Config.
● Scene fit: Directly protects API Gateway with AWS Managed Rules for SQLi and records configuration changes with AWS Config.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Shield targets DDoS, and CloudTrail is not ideal for searchable configuration history compared to AWS Config.
● Adds cost and complexity; NACLs do not apply to API Gateway which is not in your VPC.
● Throttling does not stop SQL injection, and CloudWatch Logs do not provide configuration change history.
Workflow: WAF web ACL on API Gateway stage with managed SQLi rules → track via AWS Config.
AWS CodeDeploy with a blue/green deployment group for the Auto Scaling group
CodeDeploy blue/green creates a separate fleet with equal capacity, switches ALB target groups when ready, and can schedule termination of the original fleet after a specified.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy blue/green creates a separate fleet with equal capacity, switches ALB target groups when ready, and can schedule.
● Scene fit: CodeDeploy blue/green creates a separate fleet with equal capacity, switches ALB target groups when ready, and can schedule.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Elastic Beanstalk application version lifecycle rules manage stored versions, not timed environment termination, and do not natively enforce a.
● CloudFormation retention policies apply to stack deletions and do not orchestrate blue/green capacity duplication, traffic shifting, or timed instance.
Workflow: Use AWS CodeDeploy with a blue/green deployment group for the Auto Scaling group and set BlueInstanceTerminationOption to terminate the blue instances 90 minutes after the traffic shift.
AWS Elastic Disaster Recovery (DRS)
Continuously replicates to a staging area in another Region and enables quick failover and failback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously replicates to a staging area in another Region and enables quick failover and failback.
● Scene fit: Continuously replicates to a staging area in another Region and enables quick failover and failback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides point-in-time backups, not continuous replication or automated failback.
● Controls and coordinates failover routing but does not replicate EC2 data.
● Geared for migrations; less streamlined for ongoing DR and failback than DRS.
Workflow: AWS Elastic Disaster Recovery (DRS).
Rolling deployment with 30% batches
Updates instances in-place in batches within the same environment, preserving DNS and avoiding new resources while keeping capacity available.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Updates instances in-place in batches within the same environment, preserving DNS and avoiding new resources while keeping capacity available.
● Scene fit: Updates instances in-place in batches within the same environment, preserving DNS and avoiding new resources while keeping capacity available.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Creates a separate environment and then swaps traffic, which adds new resources.
● Updates all instances simultaneously, causing downtime.
● Temporarily launches extra instances for capacity, creating additional resources.
Workflow: Rolling deployment with 30% batches.
Amazon Macie and route findings and CloudTrail S3 data events via EventBridge
Macie automatically discovers PII in S3 and EventBridge can also handle CloudTrail S3 data events for access alerts with minimal engineering.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Macie automatically discovers PII in S3 and EventBridge can also handle CloudTrail S3 data events for access alerts with minimal engineering.
● Scene fit: Macie automatically discovers PII in S3 and EventBridge can also handle CloudTrail S3 data events for access alerts with minimal engineering.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● GuardDuty detects suspicious or anomalous S3 activity but does not scan object contents for PII.
● Security Hub aggregates findings but does not discover PII in S3; it needs another service like Macie for content classification.
● Custom pipelines add engineering overhead and Comprehend is text-focused, not a managed S3-wide PII discovery solution.
Workflow: Enable Amazon Macie and route findings and CloudTrail S3 data events via EventBridge.
Amazon Macie for the target buckets and route findings to EventBridge for
Macie automatically discovers and classifies PII in S3 and integrates with EventBridge for alerts with minimal engineering.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Macie automatically discovers and classifies PII in S3 and integrates with EventBridge for alerts with minimal engineering.
● Scene fit: Macie automatically discovers and classifies PII in S3 and integrates with EventBridge for alerts with minimal engineering.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Bucket policies cannot inspect object contents to identify PII, so they cannot detect or alert on newly stored sensitive data.
● GuardDuty focuses on threat and anomaly detection rather than discovering or classifying PII in object contents.
● A custom Lambda and SageMaker pipeline adds significant development and maintenance overhead compared to a managed service.
Workflow: Enable Amazon Macie for the target buckets and route findings to EventBridge for notifications.
a BeforeAllowTraffic hook in the Lambda AppSpec to run checks and wait
BeforeAllowTraffic lets you gate the shift so the new version only receives requests after prerequisite tasks, such as schema changes or warmup, are complete.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: BeforeAllowTraffic lets you gate the shift so the new version only receives requests after prerequisite tasks, such as schema changes or warmup, are complete.
● Scene fit: BeforeAllowTraffic lets you gate the shift so the new version only receives requests after prerequisite tasks, such as.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Provisioned Concurrency addresses cold starts but does not ensure database readiness or post-deployment data changes are complete.
● AfterAllowTraffic runs after traffic is already shifted, which is too late to prevent initial errors during the cutover.
● ValidateService confirms deployment status and does not provide a pre-traffic gating mechanism for Lambda deployments.
Workflow: Add a BeforeAllowTraffic hook in the Lambda AppSpec to run checks and wait for required database updates before shifting traffic.
EventBridge rules for CodePipeline state and approval events, route to SNS, then
EventBridge natively emits CodePipeline events, SNS provides fan-out and retries, and Lambda transforms payloads for the webhook.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge natively emits CodePipeline events, SNS provides fan-out and retries, and Lambda transforms payloads for the webhook.
● Scene fit: EventBridge natively emits CodePipeline events, SNS provides fan-out and retries, and Lambda transforms payloads for the webhook.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudTrail captures API calls and can be delayed, missing some non-API state transitions and not meeting near real-time needs.
● AWS Config evaluates resource configuration compliance and does not emit CodePipeline execution or approval events.
● CodeStar Notifications targets SNS but direct webhook subscriptions require SNS confirmation and specific formatting that most chat webhooks lack, and no transformation is provided.
Workflow: Use EventBridge rules for CodePipeline state and approval events → route to SNS, → Lambda posts to the webhook.
Amazon GuardDuty across all organization accounts with the security account as the
GuardDuty provides multi-account threat detection and can centrally route findings via EventBridge to Firehose for delivery to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty provides multi-account threat detection and can centrally route findings via EventBridge to Firehose for delivery to S3.
● Scene fit: GuardDuty provides multi-account threat detection and can centrally route findings via EventBridge to Firehose for delivery to S3.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Macie focuses on sensitive data discovery in S3 and does not detect threats like SSH brute force or malware.
● Running GuardDuty only in the administrator account does not create findings for member accounts, and this pipeline adds unnecessary.
● Security Hub aggregates and prioritizes findings but does not natively detect threats like SSH brute force without a detector.
Workflow: Enable Amazon GuardDuty across all organization accounts with the security account as the delegated administrator, and route GuardDuty findings from Amazon EventBridge to Amazon Kinesis Data Firehose that writes to the S3 bucket.
Lambda function policy no longer allows s3.amazonaws.com to invoke + Bucket notification
S3 requires a resource-based permission on the function that allows the s3.amazonaws.com principal with the bucket as source; without it, invocations stop. If the bucket's event.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 requires a resource-based permission on the function that allows the s3.amazonaws.com principal with the bucket as source; without it, invocations stop.
● Scene fit: S3 requires a resource-based permission on the function that allows the s3.amazonaws.com principal with the bucket as source.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Lifecycle transitions occur after object creation and do not prevent the initial PutObject event or its notification.
● SSE-S3 or SSE-KMS does not disable event notifications for object-created events.
● Invocation from S3 depends on the function's resource-based policy, not the execution role; missing execution role permissions would cause runtime errors, not stop invocations.
Workflow: Lambda function policy no longer allows s3.amazonaws.com to invoke → Bucket notification to Lambda was removed.
AWS CodeDeploy with Amazon EC2 Auto Scaling, configure a CloudWatch alarm for
CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to automatically roll back when a metric such as CPU breaches the threshold.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to automatically roll back when a metric such as.
● Scene fit: CodeDeploy supports one-at-a-time deployments and integrates with CloudWatch alarms to automatically roll back when a metric such as.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Elastic Beanstalk rolling updates rely on environment health and do not support CloudWatch alarm driven rollbacks based on CPU.
● Is overly complex and lacks native, metric-based automatic rollback tied to a CloudWatch CPU alarm during the deployment.
● AppConfig manages configuration flags and does not perform EC2 application code deployments or coordinate instance-by-instance rollouts with ASG and.
Workflow: Use AWS CodeDeploy with Amazon EC2 Auto Scaling → configure a CloudWatch alarm for CPU utilization, choose the CodeDeployDefault.OneAtATime deployment configuration, and enable automatic rollback when the alarm is triggered.
the EngineVersion in AWS::RDS::DBInstance to the target MySQL major version, provision a
Creating a read replica for cutover and then updating EngineVersion via CloudFormation minimizes downtime because the in-place Multi-AZ upgrade otherwise takes the deployment offline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Creating a read replica for cutover and then updating EngineVersion via CloudFormation minimizes downtime because the in-place Multi-AZ.
● Scene fit: Creating a read replica for cutover and then updating EngineVersion via CloudFormation minimizes downtime because the in-place Multi-AZ.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Only automates minor upgrades and does not perform a major engine version upgrade, so it will not achieve the.
● Is a valid migration tool, but it is not a CloudFormation-driven in-place upgrade strategy for the existing RDS instance.
● DBEngineVersion is not a valid CloudFormation property for AWS::RDS::DBInstance, and upgrading before creating a replica increases downtime risk.
Workflow: Set the EngineVersion in AWS::RDS::DBInstance to the target MySQL major version, provision a like-for-like Read Replica in a separate stack first, → perform an Update Stack to apply the change with a.
Amazon Inspector with CloudWatch agent to CloudWatch Logs
Inspector continuously detects software vulnerabilities and network exposure, while the CloudWatch agent sends in-guest authentication logs to centralized CloudWatch Logs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Inspector continuously detects software vulnerabilities and network exposure, while the CloudWatch agent sends in-guest authentication logs to centralized CloudWatch Logs.
● Scene fit: Inspector continuously detects software vulnerabilities and network exposure, while the CloudWatch agent sends in-guest authentication logs to centralized CloudWatch Logs.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Patch Manager measures and applies patches rather than performing full continuous vulnerability detection.
● Security Hub aggregates findings and CloudTrail records AWS API activity.
● GuardDuty detects suspicious activity and Detective supports investigation.
Workflow: Amazon Inspector with CloudWatch agent to CloudWatch Logs.
CodeDeploy blue/green with ALB and ASG; min healthy hosts 60%; cleanup in
CodeDeploy blue/green launches a replacement fleet behind the ALB, supports lifecycle hooks for pre-traffic cleanup, enforces 60% minimum healthy hosts, and can terminate the original instances.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy blue/green launches a replacement fleet behind the ALB, supports lifecycle hooks for pre-traffic cleanup, enforces 60% minimum.
● Scene fit: CodeDeploy blue/green launches a replacement fleet behind the ALB, supports lifecycle hooks for pre-traffic cleanup, enforces 60% minimum.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Beanstalk can swap environments but lacks CodeDeploy lifecycle hooks for pre-traffic cleanup and fine-grained old-fleet termination controls.
● Instance Refresh does not provide CodeDeploy lifecycle hooks for pre-traffic cleanup and does not manage blue/green termination behavior as required.
● In-place updates do not pre-provision a new fleet and the AllowTraffic event is not scriptable for cleanup.
Workflow: CodeDeploy blue/green with ALB and ASG → min healthy hosts 60% → cleanup in BeforeAllowTraffic → terminate original.
a single CodePipeline in the primary Region with GitHub as the source
Cross-region actions let one pipeline orchestrate builds and deployments across Regions while CodePipeline manages per-Region artifact buckets that satisfy data residency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-region actions let one pipeline orchestrate builds and deployments across Regions while CodePipeline manages per-Region artifact buckets that.
● Scene fit: Cross-region actions let one pipeline orchestrate builds and deployments across Regions while CodePipeline manages per-Region artifact buckets that.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Running and maintaining multiple regional pipelines increases operational complexity and is unnecessary when a single pipeline can orchestrate multi-Region actions.
● Service focuses on game server hosting and scaling rather than CI/CD orchestration or artifact residency.
● Centralizing artifacts into one bucket can break the requirement to keep artifacts and data within their respective Regions.
Workflow: Create a single CodePipeline in the primary Region with GitHub as the source → enable cross-region actions so CodeBuild and CodeDeploy run in secondary Regions, and allow CodePipeline to create a default artifact bucket in every Region.
a CloudWatch Logs subscription to Kinesis Data Firehose that delivers to S3
A subscription filter streams container logs from CloudWatch Logs to Firehose, which buffers and writes to S3 with near-real-time latency. The awslogs driver centralizes container logs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A subscription filter streams container logs from CloudWatch Logs to Firehose, which buffers and writes to S3 with.
● Scene fit: A subscription filter streams container logs from CloudWatch Logs to Firehose, which buffers and writes to S3 with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● ALB access logs cannot be delivered to CloudWatch Logs; they only write to S3.
● CloudTrail does not capture ALB access logs or ECS container stdout/stderr, so it cannot fulfill this requirement.
● CloudWatch Logs export tasks are batch-oriented and not suitable for near-real-time streaming to S3.
Workflow: Create a CloudWatch Logs subscription to Kinesis Data Firehose that delivers to S3 → Use the awslogs log driver in ECS task definitions to send stdout/stderr to CloudWatch Logs and ensure IAM.
Block public access and use a least-privileged CodeBuild role to get the
Remove public access and grant the CodeBuild service role minimal S3 permissions so the build uses temporary role credentials.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Remove public access and grant the CodeBuild service role minimal S3 permissions so the build uses temporary role credentials.
● Scene fit: Remove public access and grant the CodeBuild service role minimal S3 permissions so the build uses temporary role credentials.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Network controls do not remove public access and do not enforce IAM authentication.
● Static credentials increase exposure risk and violate least-privilege and credential rotation best practices.
● Over-privileged permissions and continued public access fail the security requirement.
Workflow: Block public access and use a least-privileged CodeBuild role to get the object.
AWS Elastic Beanstalk with load balancing and Auto Scaling, provision an Amazon
This provides managed deployments and rollbacks, keeps a production-grade shared database independent, and centralizes searchable logs with low overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides managed deployments and rollbacks, keeps a production-grade shared database independent, and centralizes searchable logs with low overhead.
● Scene fit: This provides managed deployments and rollbacks, keeps a production-grade shared database independent, and centralizes searchable logs with low overhead.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adds significant operational overhead to manage and cycle an OpenSearch domain purely for log retention.
● Placing the database inside the Beanstalk environment ties its lifecycle to the app and is risky for shared, production databases.
● Lacks Multi-AZ for the database and requires more custom deployment orchestration and rollback management.
Workflow: Use AWS Elastic Beanstalk with load balancing and Auto Scaling, provision an Amazon RDS MySQL Multi-AZ instance decoupled from the Beanstalk environment, and stream application logs to Amazon CloudWatch Logs with 120-day retention.
AWS Config recording for EC2 instances and Dedicated Hosts with a custom
AWS Config records relationships between instances and Dedicated Hosts and custom rules provide low-overhead, continuous compliance evaluation and reporting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config records relationships between instances and Dedicated Hosts and custom rules provide low-overhead, continuous compliance evaluation and reporting.
● Scene fit: AWS Config records relationships between instances and Dedicated Hosts and custom rules provide low-overhead, continuous compliance evaluation and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on Systems Manager compliance features intended for patch and configuration states and does not natively validate Dedicated Host placement.
● Service helps track and enforce license usage but does not evaluate or report EC2 instance placement on Dedicated Hosts.
● Is possible but creates a complex, high-maintenance log processing pipeline compared to a native configuration compliance service.
Workflow: Turn on AWS Config recording for EC2 instances and Dedicated Hosts with a custom AWS Config rule that invokes Lambda to evaluate host placement and mark noncompliant instances, → use AWS Config compliance reports.
a CloudWatch alarm on ALB TargetResponseTime to the CodeDeploy deployment group to
ALB publishes TargetResponseTime directly to CloudWatch, and CodeDeploy can monitor an attached alarm.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB publishes TargetResponseTime directly to CloudWatch, and CodeDeploy can monitor an attached alarm.
● Scene fit: ALB publishes TargetResponseTime directly to CloudWatch, and CodeDeploy can monitor an attached alarm.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Access logs arrive in batches and require parsing code.
● Target health checks detect availability, not a response-latency threshold.
● X-Ray can analyze latency but needs instrumentation and sampling.
Workflow: Attach a CloudWatch alarm on ALB TargetResponseTime to the CodeDeploy deployment group to auto-stop on breach.
Apply DeletionPolicy: Snapshot to AWS::EC2::Volume + Use DeletionPolicy: Retain on AWS::RDS::DBInstance
This instructs CloudFormation to create an EBS snapshot when the volume is deleted with the stack. Retain ensures CloudFormation will not delete the DB instance during stack deletion or replacement, preserving data.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This instructs CloudFormation to create an EBS snapshot when the volume is deleted with the stack.
● Scene fit: This instructs CloudFormation to create an EBS snapshot when the volume is deleted with the stack. Retain ensures.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A stack policy can block updates but does not create EBS snapshots or specifically preserve data during replacements.
● Termination protection only blocks stack deletion and does not affect replacement behavior or snapshots.
● AWS Backup is separate from CloudFormation and does not satisfy the requirement to use CFN settings to preserve data on deletion or replacement.
Workflow: Apply DeletionPolicy: Snapshot to AWS::EC2::Volume → Use DeletionPolicy: Retain on AWS::RDS::DBInstance.
Elastic Beanstalk with a prebuilt custom AMI
Prebakes dependencies for minimal startup time while Elastic Beanstalk delivers multi AZ, ALB, and health management.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Prebakes dependencies for minimal startup time while Elastic Beanstalk delivers multi AZ, ALB, and health management.
● Scene fit: Prebakes dependencies for minimal startup time while Elastic Beanstalk delivers multi AZ, ALB, and health management.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Still performs heavy bootstrap at launch, keeping instance warm-up slow.
● Spot capacity is interruptible and does not address slow bootstrap.
● Requires re-architecting to containers and does not use EC2 AMIs to solve slow instance provisioning.
Workflow: Elastic Beanstalk with a prebuilt custom AMI.
Move the instance to Standby state
Standby removes the instance from traffic and scaling while keeping it running for open-ended debugging.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Standby removes the instance from traffic and scaling while keeping it running for open-ended debugging.
● Scene fit: Standby removes the instance from traffic and scaling while keeping it running for open-ended debugging.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Pauses new launches but does not isolate a specific instance or stop health-check replacement.
● Applies only during termination and is time-limited, not for ongoing in-service troubleshooting.
● Prevents scale-in events but does not stop health-check-based replacement or remove from traffic.
Workflow: Move the instance to Standby state.
SNS topic, Lambda per message, DynamoDB table foo
SNS triggers Lambda for each published message and DynamoDB is a fully managed serverless key-value database.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SNS triggers Lambda for each published message and DynamoDB is a fully managed serverless key-value database.
● Scene fit: SNS triggers Lambda for each published message and DynamoDB is a fully managed serverless key-value database.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Athena runs queries on data at rest and is not intended for per-event compute; this adds unnecessary indirection for immediate processing.
● Timestream is a time series database, not a general-purpose key-value store as required.
● ElastiCache is an in-memory cache, not a serverless durable key-value database for persistence.
Workflow: SNS topic, Lambda per message, DynamoDB table foo.
a DynamoDB table with partition key instanceId and sort key eventTime +
Partitioning by instanceId with a time-based sort key enables efficient point lookups and range queries by date for each instance. S3 event notifications natively invoke Lambda.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Partitioning by instanceId with a time-based sort key enables efficient point lookups and range queries by date for.
● Scene fit: Partitioning by instanceId with a time-based sort key enables efficient point lookups and range queries by date for.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds unnecessary hops and is better suited for continuous log streaming rather than a one-time drain at termination.
● Using time as the partition key fragments data by timestamp and makes instance-centric queries and date ranges inefficient.
● Requires CloudTrail data events and adds complexity, while native S3 event notifications are simpler and purpose-built.
Workflow: Create a DynamoDB table with partition key instanceId and sort key eventTime → Configure an S3 PUT event notification to trigger a Lambda function that writes object metadata into DynamoDB → Use.
Amazon FSx for Lustre linked to S3 with assumable role and security
FSx for Lustre integrates with S3 so objects appear as files, and access across accounts is controlled via IAM roles and VPC security groups.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: FSx for Lustre integrates with S3 so objects appear as files, and access across accounts is controlled via IAM roles and VPC security groups.
● Scene fit: FSx for Lustre integrates with S3 so objects appear as files, and access across accounts is controlled via.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● EFS does not present S3 objects as files; lifecycle transitions only affect EFS storage classes.
● ONTAP does not expose an S3 bucket as an NFS file namespace and NFS clients do not authenticate with IAM access keys.
● Mountpoint offers file APIs to S3 but is not a shared POSIX file system and lacks full file-system semantics for HPC.
Workflow: Amazon FSx for Lustre linked to S3 with assumable role and security groups.
a post_build phase with a commands section that pushes the Docker image
Commands in post_build run only after earlier phases succeed, ensuring the push happens only on successful builds.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Commands in post_build run only after earlier phases succeed, ensuring the push happens only on successful builds.
● Scene fit: Commands in post_build run only after earlier phases succeed, ensuring the push happens only on successful builds.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● The finally sequence runs regardless of success or failure, so it could push even when the build fails.
● The install phase is for preparing the environment and would push before the build completes, which does not meet the requirement.
● CodePipeline can orchestrate stages, but it does not by itself ensure the push occurs from the build job only on success without configuring the buildspec appropriately.
Workflow: Add a post_build phase with a commands section that pushes the Docker image.
Amazon EventBridge with an AWS API call via CloudTrail event pattern that
EventBridge can filter CloudTrail events for sts:AssumeRole targeting the specific role and immediately invoke Lambda to publish notifications to SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can filter CloudTrail events for sts:AssumeRole targeting the specific role and immediately invoke Lambda to publish notifications.
● Scene fit: EventBridge can filter CloudTrail events for sts:AssumeRole targeting the specific role and immediately invoke Lambda to publish notifications.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail Insights surfaces anomalous activity patterns and is not designed to deterministically alert on every specific sts:AssumeRole event for.
● GuardDuty detects threats and anomalies and cannot be targeted to alert on the exact assumption of a specific IAM.
● Console sign-in events do not capture STS role assumption events, so this would miss users who sign in normally.
Workflow: Use Amazon EventBridge with an AWS API call via CloudTrail event pattern that filters sts:AssumeRole on the AdminElevate role and invoke AWS Lambda to send a message to an Amazon SNS topic.
EventBridge rule for Glue job run events invoking Lambda to check final-attempt
Lambda can inspect the event details (state, attempt, run ID) and publish to SNS only when the last retry fails.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda can inspect the event details (state, attempt, run ID) and publish to SNS only when the last retry fails.
● Scene fit: Lambda can inspect the event details (state, attempt, run ID) and publish to SNS only when the last retry fails.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● EventBridge patterns cannot evaluate attempt counts, so it cannot isolate only the final retry failure.
● Possible but unnecessarily complex when filtering can be done via EventBridge and Lambda.
● Metrics do not differentiate final retry from earlier failures, so alarms will fire on any failure.
Workflow: EventBridge rule for Glue job run events invoking Lambda to check final-attempt failure and publish to SNS.
separate patch groups with distinct tags, map to an approved baseline, and
Partitioning with patch groups and offset maintenance windows staggers patching and reboots while enforcing the baseline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Partitioning with patch groups and offset maintenance windows staggers patching and reboots while enforcing the baseline.
● Scene fit: Partitioning with patch groups and offset maintenance windows staggers patching and reboots while enforcing the baseline.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A single patch group and window risks many instances rebooting at once, reducing availability.
● Event orchestration with Run Command lacks maintenance-window rate controls and proper patch targeting.
● Running with 100% concurrency causes simultaneous patching and reboots.
Workflow: Create separate patch groups with distinct tags → map to an approved baseline, and schedule two non-overlapping maintenance windows that run AWS-RunPatchBaseline.
Store the RDS database credentials in AWS Secrets Manager and attach an
Secrets Manager securely stores and rotates the RDS password, while the instance role provides IAM-based access to both the secret and DynamoDB.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Secrets Manager securely stores and rotates the RDS password, while the instance role provides IAM-based access to both.
● Scene fit: Secrets Manager securely stores and rotates the RDS password, while the instance role provides IAM-based access to both.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● DynamoDB does not use database-style credentials and should be accessed via IAM, so storing nonexistent credentials adds no value.
● Long-lived IAM user keys are unnecessary and risky for EC2 workloads, which should use short-lived credentials from an instance role.
● While SecureString is encrypted, Secrets Manager is preferred for RDS due to native rotation and DynamoDB should be accessed with IAM rather than stored secrets.
Workflow: Store the RDS database credentials in AWS Secrets Manager and attach an EC2 instance profile that can read that secret and call DynamoDB APIs.
Launch EC2 instances on Dedicated Hosts and tag the application, then use
Dedicated Hosts provide host-level control and visibility of physical sockets and AWS Config custom rules can continuously flag noncompliant resources for a compliance view.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Dedicated Hosts provide host-level control and visibility of physical sockets and AWS Config custom rules can continuously flag.
● Scene fit: Dedicated Hosts provide host-level control and visibility of physical sockets and AWS Config custom rules can continuously flag.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Suggests using AWS Config correctly but Reserved Instances are only a billing discount and do not satisfy socket-based licensing.
● Dedicated Instances provide single-tenant isolation but do not expose or control physical socket inventory required for per-socket licensing.
● Service Catalog governs provisioning catalogs and constraints but is not the right tool for ongoing configuration compliance monitoring.
Workflow: Launch EC2 instances on Dedicated Hosts and tag the application, → use an AWS Config custom rule with a Lambda function to validate that instances run with the required launch mode.
Application Discovery Service with vCenter Discovery Connector and EC2 agent; view in
Agentless collection for vCenter VMs and agent for EC2 with built-in Migration Hub dashboards and minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Agentless collection for vCenter VMs and agent for EC2 with built-in Migration Hub dashboards and minimal setup.
● Scene fit: Agentless collection for vCenter VMs and agent for EC2 with built-in Migration Hub dashboards and minimal setup.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● AWS Config tracks AWS resources and does not inventory on-prem VMware VMs or host-level MAC/IP details.
● Requires deploying and managing the SSM Agent on every machine and building the reporting pipeline, which is more effort.
● Focuses on vulnerability findings for supported AWS resources and does not inventory on-prem VMs or MAC/IP attributes.
Workflow: Application Discovery Service with vCenter Discovery Connector and EC2 agent → view in Migration Hub.
AWS Config required-tags with an aggregator feeding QuickSight
AWS Config managed rules evaluate tag compliance across accounts and Regions; an aggregator centralizes results for reporting and dashboards.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config managed rules evaluate tag compliance across accounts and Regions; an aggregator centralizes results for reporting and dashboards.
● Scene fit: AWS Config managed rules evaluate tag compliance across accounts and Regions; an aggregator centralizes results for reporting and dashboards.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Resource Explorer helps find resources but does not perform rule-based, continuous tag compliance evaluations.
● Tag policies guide standardization and show coverage but do not enforce or continuously evaluate resource-level compliance suitable for a dashboard.
● CloudTrail records API activity, not current resource tag compliance state.
Workflow: AWS Config required-tags with an aggregator feeding QuickSight.
an Auto Scaling group with target-tracking CPU and scheduled min 3 after
Combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable off-hours, improving cost and resilience.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable off-hours, improving cost and resilience.
● Scene fit: Combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable off-hours, improving cost and resilience.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Lowers compute costs via commitment but adds no elasticity, health checks, or scale-in/out behavior.
● Time-based control only; no health-based replacement or automatic response to sudden spikes.
● Cheaper capacity but subject to interruptions and does not by itself set scaling policy for unpredictable bursts.
Workflow: Create an Auto Scaling group with target-tracking CPU and scheduled min 3 after hours and 6 before workday.
Readiness endpoint in the web service with ALB health check path set
Expose an endpoint that tests backend and RDS reachability; ALB health checks use this path to mark targets healthy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Expose an endpoint that tests backend and RDS reachability; ALB health checks use this path to mark targets healthy.
● Scene fit: Expose an endpoint that tests backend and RDS reachability; ALB health checks use this path to mark targets healthy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Does not drive ALB target health and does not validate cross-tier connectivity.
● ECS health checks affect task lifecycle, but ALB still uses its own health check and would not validate dependencies by default.
● ALB cannot consume CloudWatch alarm state to set target health and these metrics do not prove dependency connectivity.
Workflow: Readiness endpoint in the web service with ALB health check path set to it.
AWS Shield Advanced, front origins with Amazon CloudFront, and enforce least-privilege security
Shield Advanced adds managed DDoS protections and CloudFront reduces origin exposure, while least-privilege SGs constrain attack paths. WAF filters abusive patterns and known sources at the.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Shield Advanced adds managed DDoS protections and CloudFront reduces origin exposure, while least-privilege SGs constrain attack paths.
● Scene fit: Shield Advanced adds managed DDoS protections and CloudFront reduces origin exposure, while least-privilege SGs constrain attack paths. WAF.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● These are governance and resilience measures, not targeted DDoS protections or edge-layer mitigations.
● DNS failover can reroute but does not stop volumetric or application-layer DDoS and does not shield the origin.
● Scaling capacity attempts to absorb traffic rather than reducing attack surface or providing dedicated DDoS protections.
Workflow: Enable AWS Shield Advanced → front origins with Amazon CloudFront, and enforce least-privilege security groups → Configure AWS WAF rate-based rules and IP sets at the edge, and tighten VPC network ACLs to required ports and CIDRs.
a BeforeAllowTraffic hook in the Lambda AppSpec that waits for the Step
BeforeAllowTraffic executes before alias shifting, letting you block traffic until the Step Functions execution completes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: BeforeAllowTraffic executes before alias shifting, letting you block traffic until the Step Functions execution completes.
● Scene fit: BeforeAllowTraffic executes before alias shifting, letting you block traffic until the Step Functions execution completes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A canary still shifts some traffic to the new version before the workflow completes, violating the requirement.
● Triggering alias changes outside CodeDeploy bypasses deployment lifecycle control and is not integrated with deployment status.
● AfterAllowTraffic runs only after the alias has already shifted, so the new version would receive production traffic too early.
Workflow: Use a BeforeAllowTraffic hook in the Lambda AppSpec that waits for the Step Functions run to finish.
Restart the ECS agent on the container instances
Restarting the ECS agent clears stale state and forces fresh image pulls, fixing sporadic reuse of old images.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Restarting the ECS agent clears stale state and forces fresh image pulls, fixing sporadic reuse of old images.
● Scene fit: Restarting the ECS agent clears stale state and forces fresh image pulls, fixing sporadic reuse of old images.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Changing registries does not address cached images or agent state on ECS container instances.
● Digests prevent tag drift but do not resolve intermittent agent caching or pull failures.
● Deployment orchestration does not fix the agent reusing a locally cached image on EC2 hosts.
Workflow: Restart the ECS agent on the container instances.
Amazon GuardDuty and send findings via EventBridge to SNS
GuardDuty detects EC2 compromise behaviors and integrates with EventBridge and SNS for email alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty detects EC2 compromise behaviors and integrates with EventBridge and SNS for email alerts.
● Scene fit: GuardDuty detects EC2 compromise behaviors and integrates with EventBridge and SNS for email alerts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Surfaces API anomalies but misses network and runtime indicators of EC2 compromise.
● Inspector focuses on vulnerability scanning, not real-time compromise detection.
● Detective assists with investigation and visualization rather than primary detection and alerting.
Workflow: Turn on Amazon GuardDuty and send findings via EventBridge to SNS.
CloudTrail to CloudWatch Logs; metric filter on PutBucketPolicy/DeleteBucketPolicy; CloudWatch alarm
CloudTrail captures S3 control plane API calls; metric filters on these events can drive immediate CloudWatch alarms.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail captures S3 control plane API calls; metric filters on these events can drive immediate CloudWatch alarms.
● Scene fit: CloudTrail captures S3 control plane API calls; metric filters on these events can drive immediate CloudWatch alarms.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● S3 Event Notifications do not emit policy change events like PutBucketPolicy or DeleteBucketPolicy.
● Server access logs track object-level requests and do not include bucket policy changes.
● Checks for public read exposure, not all bucket policy updates or changes.
Workflow: CloudTrail to CloudWatch Logs → metric filter on PutBucketPolicy/DeleteBucketPolicy → CloudWatch alarm.
Keep Platform version LATEST and choose Force new deployment
Forcing a new deployment replaces tasks so they start on the current latest Fargate platform.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Forcing a new deployment replaces tasks so they start on the current latest Fargate platform.
● Scene fit: Forcing a new deployment replaces tasks so they start on the current latest Fargate platform.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but is unnecessary complexity for simply moving to the latest Fargate runtime.
● Platform version is not defined in a task definition and cannot be set there.
● Tasks will not switch platforms until a deployment replaces them.
Workflow: Keep Platform version LATEST and choose Force new deployment.
Introduce a single stream-consumer Lambda that republishes each record to an Amazon
A single consumer per shard prevents reader contention and SNS provides scalable fan-out so multiple Lambdas can process the same events without additional shard readers.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A single consumer per shard prevents reader contention and SNS provides scalable fan-out so multiple Lambdas can process.
● Scene fit: A single consumer per shard prevents reader contention and SNS provides scalable fan-out so multiple Lambdas can process.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● DAX accelerates table reads and does not change how DynamoDB Streams are consumed, so it will not address throttling on the stream.
● RCUs affect table read throughput, not the DynamoDB Streams reader limits, so this will not resolve stream consumer throttling.
● Parallelization increases concurrency for a single consumer but does not bypass the per-shard reader limits and can worsen contention across multiple consumers.
Workflow: Introduce a single stream-consumer Lambda that republishes each record to an Amazon SNS topic and have existing and future Lambdas subscribe to that topic.
Amazon FSx for Lustre linked to the S3 bucket, enable a cross-account
FSx for Lustre integrates with S3 via a data repository so objects appear as files, and cross-account access is appropriately handled with assumable IAM roles plus.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: FSx for Lustre integrates with S3 via a data repository so objects appear as files, and cross-account access.
● Scene fit: FSx for Lustre integrates with S3 via a data repository so objects appear as files, and cross-account access.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● FSx for NetApp ONTAP does not natively expose S3 objects as files and FSx services do not support the.
● FSx for Windows File Server targets SMB/NTFS workloads and does not provide file-based views of S3 objects.
● FSx for OpenZFS does not surface S3 objects as a native file namespace, so it fails the S3 file-based.
Workflow: Deploy Amazon FSx for Lustre linked to the S3 bucket → enable a cross-account assumable IAM role with an identity-based policy, and use security groups to manage file system access.
Amazon Elastic File System with EC2 Spot Instances
This combines a managed, shared POSIX file system with low-cost interruptible compute and supports checkpoint-based restarts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This combines a managed, shared POSIX file system with low-cost interruptible compute and supports checkpoint-based restarts.
● Scene fit: This combines a managed, shared POSIX file system with low-cost interruptible compute and supports checkpoint-based restarts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides a shared file system, but using On-Demand Instances does not minimize compute cost for interruptible workloads.
● EBS volumes are attached to individual instances and are not a concurrent shared file system, and Reserved Instances are not optimized for interruption-tolerant cost savings.
● S3 is object storage and not a POSIX mountable shared file system, and On-Demand Instances increase cost for workloads tolerant of interruptions.
Workflow: Amazon Elastic File System with EC2 Spot Instances.
AWS Organizations with cross-account IAM roles and tag-based TerminateInstances conditions
Use OUs and per-team roles in the shared account, applying ABAC with resource and principal tags to restrict TerminateInstances to owned instances.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use OUs and per-team roles in the shared account, applying ABAC with resource and principal tags to restrict TerminateInstances to owned instances.
● Scene fit: Use OUs and per-team roles in the shared account, applying ABAC with resource and principal tags to restrict TerminateInstances to owned instances.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adds global safeguards and approvals but does not enforce owner-based authorization across accounts.
● SCPs provide guardrails and cannot grant fine-grained owner-specific permissions; they still require IAM policy enforcement.
● These deliver visibility and governance but do not enforce per-resource termination authorization.
Workflow: AWS Organizations with cross-account IAM roles and tag-based TerminateInstances conditions.
a NAT gateway in a public subnet with an Elastic IP and
A NAT gateway with an Elastic IP in a public subnet enables instances in private subnets to reach the Internet while staying unreachable from it.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A NAT gateway with an Elastic IP in a public subnet enables instances in private subnets to reach.
● Scene fit: A NAT gateway with an Elastic IP in a public subnet enables instances in private subnets to reach.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A Customer Gateway is used for site-to-site VPNs and an ALB cannot route to it, so this does not.
● An Internet Gateway has no IP to whitelist and private subnets cannot route to an IGW without public IPs.
● VPC endpoints connect privately to supported AWS services, not arbitrary Internet hosts, and scripting EIPs does not fix private.
Workflow: Create a NAT gateway in a public subnet with an Elastic IP and update private subnet route tables to send Internet-bound traffic to the NAT gateway.
Two CodePipelines from one GitHub repo using env branches, STAGING auto on
It uses two pipelines, keeps one source repository, automatically deploys staging, and places a manual gate before production.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: It uses two pipelines, keeps one source repository, automatically deploys staging, and places a manual gate before production.
● Scene fit: It uses two pipelines, keeps one source repository, automatically deploys staging, and places a manual gate before production.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Uses one pipeline and automatically deploys production.
● Separate pipelines are workable, but manual approval in staging blocks automatic deployment.
● Production has no manual approval, so it fails the key safety control.
Workflow: Two CodePipelines from one GitHub repo using env branches, STAGING auto on push, PRODUCTION has Manual approval → deploy via CloudFormation.
AddToLoadBalancer
Prevents newly launched instances from being attached to the ALB target group, keeping them from receiving client traffic.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Prevents newly launched instances from being attached to the ALB target group, keeping them from receiving client traffic.
● Scene fit: Prevents newly launched instances from being attached to the ALB target group, keeping them from receiving client traffic.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Rebalances capacity across Availability Zones but does not control ALB target registration.
● Stops replacement of unhealthy instances but does not stop ALB registration for new instances.
● Blocks new instance launches entirely rather than preventing traffic to them.
Workflow: AddToLoadBalancer.
DynamoDB: partition key instanceId, sort key eventTime + S3 PutObject event to
instanceId as the partition key and eventTime as the sort key supports direct per-instance lookup and date-range queries. The S3 upload event can immediately invoke Lambda.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: instanceId as the partition key and eventTime as the sort key supports direct per-instance lookup and date-range queries.
● Scene fit: instanceId as the partition key and eventTime as the sort key supports direct per-instance lookup and date-range queries.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 Inventory is periodic and not immediate.
● The chain can be engineered, but CloudWatch Logs and Firehose add unnecessary hops for a one-time log drain.
● Using time as the partition key does not match the main query by instance ID and can create uneven or hot time partitions.
Workflow: DynamoDB: partition key instanceId, sort key eventTime → S3 PutObject event to Lambda writing metadata to DynamoDB → ASG termination lifecycle hook invoking Lambda to run SSM Run Command and copy logs to S3.
Amazon Cognito user pools with SAML IdP plus Cognito identity pools for
User pools deliver the branded sign-up/sign-in and SAML federation, and identity pools exchange tokens for short-lived AWS credentials.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: User pools deliver the branded sign-up/sign-in and SAML federation, and identity pools exchange tokens for short-lived AWS credentials.
● Scene fit: User pools deliver the branded sign-up/sign-in and SAML federation, and identity pools exchange tokens for short-lived AWS credentials.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● ALB authentication can gate access but does not provide a full sign-up/sign-in UI or exchange tokens for AWS credentials for the client.
● Lambda authorizers do not perform SAML web sign-in flows or provide a user-facing registration experience.
● AssumeRoleWithWebIdentity is for OIDC; SAML requires AssumeRoleWithSAML and does not provide a built-in sign-up/sign-in UI.
Workflow: Amazon Cognito user pools with SAML IdP → Cognito identity pools for temporary AWS credentials.
CodePipeline custom action with an on-prem job worker polling for jobs and
A custom action with a job worker is the supported pattern for long-running external tasks; it securely polls, executes on-prem, and reports status back to CodePipeline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A custom action with a job worker is the supported pattern for long-running external tasks; it securely polls, executes on-prem, and reports status back to CodePipeline.
● Scene fit: A custom action with a job worker is the supported pattern for long-running external tasks; it securely polls.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Step Functions orchestrates workflows but does not run on-prem tools; you still need a durable worker and callback integration.
● Lambda has a 15-minute max duration, which cannot accommodate a ~120-minute scan.
● CodePipeline has no native Run Command action; it would require additional orchestration and polling, making it less suitable for reliable job status reporting.
Workflow: CodePipeline custom action with an on-prem job worker polling for jobs and posting results.
an Amazon CloudWatch alarm on the EC2 instances' maximum CPU metric and
CodeDeploy can monitor CloudWatch alarms and automatically roll back a deployment when the EC2 CPU alarm breaches during traffic shifting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy can monitor CloudWatch alarms and automatically roll back a deployment when the EC2 CPU alarm breaches during traffic shifting.
● Scene fit: CodeDeploy can monitor CloudWatch alarms and automatically roll back a deployment when the EC2 CPU alarm breaches during.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Auto Scaling policies can add capacity but they do not trigger CodeDeploy rollbacks based on CPU thresholds.
● CodeDeploy rollback integration works with CloudWatch alarms and the ALB does not expose a CPU metric.
● ValidateService runs before traffic is shifted behind the ALB, so it cannot assess CPU under live load and is not the intended rollback mechanism.
Workflow: Create an Amazon CloudWatch alarm on the EC2 instances' maximum CPU metric and enable automatic rollback in the CodeDeploy deployment group using that alarm.
Amazon Route 53 latency-based routing with health checks to regional API Gateway
Latency-based routing directs clients to the lowest-latency API endpoint while health checks enable failover, and DynamoDB global tables provide multi-Region, active-active data access.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency-based routing directs clients to the lowest-latency API endpoint while health checks enable failover, and DynamoDB global tables provide multi-Region, active-active data access.
● Scene fit: Latency-based routing directs clients to the lowest-latency API endpoint while health checks enable failover, and DynamoDB global tables.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Geolocation routing sends users based on their geographic origin rather than the lowest latency path, so it does not optimize for minimal latency across Regions.
● Global Accelerator and ALB add complexity and do not directly provide the best multi-Region, per-request latency steering for API Gateway.
● Failover routing prioritizes active-passive resiliency rather than distributing traffic for the lowest latency across multiple Regions.
Workflow: Use Amazon Route 53 latency-based routing with health checks to regional API Gateway endpoints, with in-Region Lambda and a DynamoDB global table.
Analyze the logs using AWS Athena + Enable Access Logs on the
Athena is serverless and pay-per-query, making it ideal for occasional analysis of logs stored in S3. ALB access logging writes detailed request records to S3 for.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Athena is serverless and pay-per-query, making it ideal for occasional analysis of logs stored in S3.
● Scene fit: Athena is serverless and pay-per-query, making it ideal for occasional analysis of logs stored in S3. ALB access.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Target groups do not support access logs, which are available only on the load balancer itself.
● The CloudWatch Logs agent is intended for continuous streaming to CloudWatch Logs and is not suited for a one-time.
● EMR typically incurs higher costs and overhead for sporadic workloads compared to serverless query services.
Workflow: Analyze the logs using AWS Athena → Enable Access Logs on the Application Load Balancer → Create a termination lifecycle hook and trigger Lambda via EventBridge to run SSM Run Command to copy application logs to S3.
UpdatePolicy AutoScalingRollingUpdate with MinInstancesInService
This CloudFormation UpdatePolicy performs a rolling update of the Auto Scaling group while ensuring a minimum number of instances remain in service.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This CloudFormation UpdatePolicy performs a rolling update of the Auto Scaling group while ensuring a minimum number of instances remain in service.
● Scene fit: This CloudFormation UpdatePolicy performs a rolling update of the Auto Scaling group while ensuring a minimum number of instances remain in service.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● UpdateReplacePolicy governs retain/snapshot behavior on replacement and does not support rolling update controls like MinSuccessfulInstancesPercent or MaxBatchSize for Auto Scaling groups.
● Change Sets only preview changes and do not control rolling behavior or minimum in-service capacity during updates.
● Triggers full replacement of the Auto Scaling group rather than a controlled rolling update, risking service interruption.
Workflow: Configure UpdatePolicy AutoScalingRollingUpdate with MinInstancesInService.
CodeDeploy blue/green for ASG with BlueInstanceTerminationOption=TERMINATE and terminationWaitTimeInMinutes=120
CodeDeploy blue/green for EC2/Auto Scaling supports traffic shifting and timed termination of the original instances.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy blue/green for EC2/Auto Scaling supports traffic shifting and timed termination of the original instances.
● Scene fit: CodeDeploy blue/green for EC2/Auto Scaling supports traffic shifting and timed termination of the original instances.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Elastic Beanstalk can swap URLs, but application version lifecycle does not schedule environment termination in minutes.
● Is custom orchestration and lacks a native timed termination setting for the original fleet.
● ECS blue/green applies to ECS services and tasks, not EC2 Auto Scaling groups.
Workflow: CodeDeploy blue/green for ASG with BlueInstanceTerminationOption=TERMINATE and terminationWaitTimeInMinutes=120.
Amazon Inspector for continuous vulnerability scans on the EC2 instances and install
Amazon Inspector automates vulnerability detection and the CloudWatch agent centralizes OS logs such as auth and secure logs for auditing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Amazon Inspector automates vulnerability detection and the CloudWatch agent centralizes OS logs such as auth and secure logs.
● Scene fit: Amazon Inspector automates vulnerability detection and the CloudWatch agent centralizes OS logs such as auth and secure logs.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Security Hub aggregates findings and CloudTrail records API actions, but CloudTrail does not capture in-guest OS login activity and.
● GuardDuty detects threats not software vulnerabilities, and X-Ray traces applications rather than collecting OS login logs.
● Patch Manager evaluates and applies patches rather than performing full vulnerability scans, and KCL is for consuming Kinesis Data.
Workflow: Use Amazon Inspector for continuous vulnerability scans on the EC2 instances and install the Amazon CloudWatch agent to stream system logs, including login events, to Amazon CloudWatch Logs.
a CloudFormation child template with an Auto Scaling group set to MinSize=1
This provides self-healing, preserves static private IPs by reusing ENIs, and scales the pattern with minimal custom code via nested CloudFormation stacks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides self-healing, preserves static private IPs by reusing ENIs, and scales the pattern with minimal custom code.
● Scene fit: This provides self-healing, preserves static private IPs by reusing ENIs, and scales the pattern with minimal custom code.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Beanstalk abstracts instance details and does not natively manage static ENI attachment and deterministic per-instance hostnames with minimal effort.
● Relies on manual or ad-hoc execution and custom logic for recovery, which is not the least-effort nor fully automated.
● Adds complexity and does not inherently guarantee persistent ENI attachment or static hostnames during replacements compared to a simpler.
Workflow: Create a CloudFormation child template with an Auto Scaling group set to MinSize=1 and MaxSize=1 that attaches a designated ENI and sets the hostname at boot, → use nested stacks from a.
CloudFormation StackSets with self-managed permissions using cross-account roles
Supports deploying to external accounts by creating an admin role in the source and execution roles in each target account without needing AWS Organizations.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Supports deploying to external accounts by creating an admin role in the source and execution roles in each target account without needing AWS Organizations.
● Scene fit: Supports deploying to external accounts by creating an admin role in the source and execution roles in each target account without needing AWS Organizations.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Requires accounts to be enrolled in a Control Tower landing zone under a shared organization, which is not allowed here.
● Possible but adds significant setup and ongoing operations compared to StackSets and is not the most direct fit.
● Requires target accounts to be in the same AWS Organization with trusted access enabled, which is not permitted.
Workflow: CloudFormation StackSets with self-managed permissions using cross-account roles.
Modify the application health endpoint to return 200 only when it can
ALB target health is determined by HTTP status codes, so surfacing database reachability through the endpoint ensures instances without DB connectivity are marked unhealthy and removed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB target health is determined by HTTP status codes, so surfacing database reachability through the endpoint ensures instances.
● Scene fit: ALB target health is determined by HTTP status codes, so surfacing database reachability through the endpoint ensures instances.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● ALB target group health checks evaluate status codes and timeouts, not response bodies, so regex matching on payload content is not supported.
● Route 53 health checks operate at the DNS level and cannot choose individual instances behind an ALB target group.
● A canary improves monitoring but does not influence ALB routing decisions, so users would still be sent to unhealthy targets.
Workflow: Modify the application health endpoint to return 200 only when it can reach the database and configure the ALB health check to call that path.
The IAM instance profile role attached to the EC2 instance is misconfigured
If the instance role lacks s3:GetObject on the object or has incorrect trust or permission policies, S3 returns AccessDenied. A restrictive resource policy that limits principals.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: If the instance role lacks s3:GetObject on the object or has incorrect trust or permission policies, S3 returns.
● Scene fit: If the instance role lacks s3:GetObject on the object or has incorrect trust or permission policies, S3 returns.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Block Public Access prevents public ACLs and bucket policies but does not stop an authorized IAM role from accessing.
● Object Lock enforces retention and legal holds but does not prevent reads by principals with permission, so it is.
● Server-side encryption at rest does not change authorization outcomes for reads when permissions are correct, so it does not.
Workflow: The IAM instance profile role attached to the EC2 instance is misconfigured → The S3 bucket policy denies the request from the instance's identity.
lifecycle hooks to the Auto Scaling group and enable a warm pool
Warm pools keep pre-initialized instances ready and lifecycle hooks coordinate readiness so traffic is only sent when instances are healthy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Warm pools keep pre-initialized instances ready and lifecycle hooks coordinate readiness so traffic is only sent when instances are healthy.
● Scene fit: Warm pools keep pre-initialized instances ready and lifecycle hooks coordinate readiness so traffic is only sent when instances are healthy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Standby is intended for troubleshooting and maintenance, not rapid scale-out during traffic spikes.
● Manual scaling increases operational burden and risks overprovisioning and higher costs.
● Pre-scaling by overprovisioning raises costs and wastes capacity outside the event window.
Workflow: Add lifecycle hooks to the Auto Scaling group and enable a warm pool for faster scale-out.
The ALB hasn't enabled that Availability Zone
ALB requires each Availability Zone to be explicitly enabled with a subnet; otherwise, no load balancer nodes exist there and no requests are routed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB requires each Availability Zone to be explicitly enabled with a subnet; otherwise, no load balancer nodes exist there and no requests are routed.
● Scene fit: ALB requires each Availability Zone to be explicitly enabled with a subnet; otherwise, no load balancer nodes exist.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● If instances are in the wrong target group, they would not receive traffic; however, this usually is not isolated to a single new zone if the ASG uses one target group.
● Disabling cross-zone affects how traffic is balanced across enabled zones but does not stop routing to an enabled zone.
● Unhealthy targets would not receive requests, but this would be visible as failed health checks rather than an AZ enablement issue.
Workflow: The ALB hasn't enabled that Availability Zone.
a cross-Region Aurora read replica in eu-central-1, configure RDS event notifications to
Event-driven notifications from RDS combined with automated promotion and DNS failover minimize RTO and reduce downtime.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Event-driven notifications from RDS combined with automated promotion and DNS failover minimize RTO and reduce downtime.
● Scene fit: Event-driven notifications from RDS combined with automated promotion and DNS failover minimize RTO and reduce downtime.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Depends on human intervention and a temporary S3 page, which prolongs detection and recovery time and increases downtime.
● Scheduled polling introduces detection lag up to the interval, which delays failover and increases downtime.
● Restoring from snapshots can take hours and yields a higher RTO and RPO than replication, so it does not minimize downtime.
Workflow: Create a cross-Region Aurora read replica in eu-central-1 → configure RDS event notifications to an SNS topic → subscribe a Lambda function that promotes the replica on failure, and update Route 53 to direct traffic to the secondary Region.
a least-privilege CloudFormation execution role, attach to the stack, and allow iam:PassRole
Delegating to a minimally scoped execution role and granting iam:PassRole lets CloudFormation assume the role to create resources while honoring least privilege.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Delegating to a minimally scoped execution role and granting iam:PassRole lets CloudFormation assume the role to create resources while honoring least privilege.
● Scene fit: Delegating to a minimally scoped execution role and granting iam:PassRole lets CloudFormation assume the role to create resources while honoring least privilege.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Without iam:PassRole, CloudFormation cannot assume the execution role and act on the user's behalf.
● Capabilities acknowledge IAM changes but do not grant permissions or allow passing roles.
● Stack policies control updates and protections on stack resources; they do not grant caller permissions.
Workflow: Create a least-privilege CloudFormation execution role → attach to the stack, and allow iam:PassRole.
Amazon EC2 Auto Scaling to send notifications to an Amazon SNS topic
This directly alerts the team whenever the Auto Scaling group records a failed instance launch event.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This directly alerts the team whenever the Auto Scaling group records a failed instance launch event.
● Scene fit: This directly alerts the team whenever the Auto Scaling group records a failed instance launch event.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Detects issues after an instance is running and does not notify on failed launch attempts.
● Focuses on EC2 API errors and may miss Auto Scaling launch failures that do not surface as RunInstances errors.
● Status checks indicate problems with running instances, not failures during the launch process.
Workflow: Configure Amazon EC2 Auto Scaling to send notifications to an Amazon SNS topic for the EC2_INSTANCE_LAUNCH_ERROR event.
AWS Compute Optimizer to analyze historical utilization and follow its rightsizing recommendations
Compute Optimizer analyzes historical metrics and applies machine learning to suggest optimal instance types and sizes for better utilization and cost efficiency.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Compute Optimizer analyzes historical metrics and applies machine learning to suggest optimal instance types and sizes for better utilization and cost efficiency.
● Scene fit: Compute Optimizer analyzes historical metrics and applies machine learning to suggest optimal instance types and sizes for better.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds monitoring and automation but does not provide ML-based rightsizing recommendations and relies on brittle threshold rules.
● Control Tower and Image Builder address governance and AMI pipelines rather than workload rightsizing or utilization optimization.
● Trusted Advisor offers basic checks and may require higher support tiers, but it is less comprehensive for detailed rightsizing than Compute Optimizer.
Workflow: Use AWS Compute Optimizer to analyze historical utilization and follow its rightsizing recommendations to adjust EC2 instance families and sizes.
Send a SUCCESS or FAILED payload from the provider to the CloudFormation
CloudFormation waits for the provider to PUT a JSON response to the pre-signed ResponseURL, and without it the operation stays in progress.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation waits for the provider to PUT a JSON response to the pre-signed ResponseURL, and without it the operation stays in progress.
● Scene fit: CloudFormation waits for the provider to PUT a JSON response to the pre-signed ResponseURL, and without it the operation stays in progress.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudFormation ignores the Lambda return value for custom resources and requires a response sent to its ResponseURL instead.
● Choosing AWS::CloudFormation::CustomResource or a Custom:: type is valid but does not by itself complete the stack without the provider response.
● EventBridge events do not satisfy CloudFormation custom resource completion which explicitly requires a response to the ResponseURL.
Workflow: Send a SUCCESS or FAILED payload from the provider to the CloudFormation ResponseURL pre-signed S3 URL.
an environment tag to every instance, create the hardened AMI with AWS
Systems Manager Automation supports golden AMI workflows to cut boot time, while Parameter Store SecureString with KMS encryption securely provides environment secrets to the same AMI.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Systems Manager Automation supports golden AMI workflows to cut boot time, while Parameter Store SecureString with KMS encryption.
● Scene fit: Systems Manager Automation supports golden AMI workflows to cut boot time, while Parameter Store SecureString with KMS encryption.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Session Manager provides interactive, managed shell access rather than AMI creation capabilities, and adding Lambda from user data adds.
● Patch Manager targets OS patching not application baking, and AppConfig is designed for dynamic configuration rather than storing secrets.
Workflow: Add an environment tag to every instance → create the hardened AMI with AWS Systems Manager Automation runbooks → use a user data bootstrap script to read the tag and load settings.
Provision a read replica, upgrade the replica to the target major version
Creating and upgrading a read replica, then promoting it and switching traffic minimizes downtime and can be orchestrated with CloudFormation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Creating and upgrading a read replica, then promoting it and switching traffic minimizes downtime and can be orchestrated with CloudFormation.
● Scene fit: Creating and upgrading a read replica, then promoting it and switching traffic minimizes downtime and can be orchestrated with CloudFormation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● AutoMinorVersionUpgrade applies only minor version updates and cannot perform a major version upgrade.
● DMS can minimize downtime but adds migration complexity and is not an in-place CloudFormation-driven upgrade of the existing instance.
● An in-place major version upgrade causes downtime even with Multi-AZ, which is not minimal.
Workflow: Provision a read replica, upgrade the replica to the target major version → promote it, → cut over.
Elastic Beanstalk to use immutable environment updates
This spins up a separate temporary Auto Scaling group behind the same load balancer, preserves the CNAME, and discards the new group on failure.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This spins up a separate temporary Auto Scaling group behind the same load balancer, preserves the CNAME, and discards the new group on failure.
● Scene fit: This spins up a separate temporary Auto Scaling group behind the same load balancer, preserves the CNAME, and discards the new group on failure.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Maintains capacity by adding a batch but still updates the in-place fleet, so failures can affect current instances and rollback is slower.
● Requires a DNS CNAME swap between environments, which violates the requirement to avoid any DNS changes.
● Updates instances in place in batches, reducing capacity and potentially impacting availability during deployment.
Workflow: Configure Elastic Beanstalk to use immutable environment updates.
the CloudWatch agent to push application logs to CloudWatch Logs and set
Both data sources land in CloudWatch Logs, which allows CloudWatch Logs Insights to query them together.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Both data sources land in CloudWatch Logs, which allows CloudWatch Logs Insights to query them together.
● Scene fit: Both data sources land in CloudWatch Logs, which allows CloudWatch Logs Insights to query them together.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch Logs Insights can only query data in CloudWatch Logs, so splitting destinations with S3 prevents unified queries.
● The CloudWatch agent does not natively ship logs directly to S3, even though Athena can query S3 data.
● SQS is not a persistent log store or query service, and neither the agent nor CloudTrail supports SQS as a logging destination.
Workflow: Configure the CloudWatch agent to push application logs to CloudWatch Logs and set CloudTrail to deliver API events to a CloudWatch Logs log group, → analyze both with CloudWatch Logs Insights.
S3 server access logging and query the log files in place with
S3 server access logs contain detailed request records and Athena lets you analyze them serverlessly with minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 server access logs contain detailed request records and Athena lets you analyze them serverlessly with minimal setup.
● Scene fit: S3 server access logs contain detailed request records and Athena lets you analyze them serverlessly with minimal setup.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but introduces more setup and maintenance effort for ingestion and a data warehouse.
● While possible, this pipeline is more complex and can be costly for high-volume data events.
● S3 event notifications do not emit GET/read events, so you cannot capture access requests this way.
Workflow: Turn on S3 server access logging and query the log files in place with Amazon Athena using an external table to derive per-object access metrics.
S3 server access logging and use Amazon Athena with an external table
S3 access logs contain per-request object details and Athena can query them serverlessly with SQL, delivering fast and low-cost insights into the most accessed objects.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 access logs contain per-request object details and Athena can query them serverlessly with SQL, delivering fast and.
● Scene fit: S3 access logs contain per-request object details and Athena can query them serverlessly with SQL, delivering fast and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 Storage Lens provides aggregated metrics but does not offer per-object request details needed to rank specific videos, making.
● Redshift Spectrum requires a Redshift cluster which increases cost and setup time, making it excessive for simple, ad hoc.
● Building a Lambda plus Firehose plus OpenSearch pipeline adds operational complexity and ongoing costs that are unnecessary for batch-style.
Workflow: Enable S3 server access logging and use Amazon Athena with an external table to query the log files and identify top GETs and downloads.
Register a delegated administrator for CloudFormation StackSets and use service-managed permissions +
With Organizations trusted access, service-managed StackSets and delegated admin enable multi-account, multi-Region deployments with low overhead. All features is required to enable trusted access so CloudFormation StackSets can manage roles across member accounts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: With Organizations trusted access, service-managed StackSets and delegated admin enable multi-account, multi-Region deployments with low overhead.
● Scene fit: With Organizations trusted access, service-managed StackSets and delegated admin enable multi-account, multi-Region deployments with low overhead. All features.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Control Tower is not a general-purpose mechanism for arbitrary CloudFormation rollouts across many accounts and Regions.
● Consolidated billing mode does not enable trusted access required by service-managed StackSets.
● Self-managed permissions require creating and maintaining roles in every account, increasing operational overhead.
Workflow: Register a delegated administrator for CloudFormation StackSets and use service-managed permissions → Turn on all features in AWS Organizations.
an AWS CodePipeline with a CodeCommit source trigger, add AWS CodeBuild for
This provides an end-to-end managed CI/CD pipeline with automated testing and controlled linear traffic shifting for Lambda plus rollback support.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides an end-to-end managed CI/CD pipeline with automated testing and controlled linear traffic shifting for Lambda plus rollback support.
● Scene fit: This provides an end-to-end managed CI/CD pipeline with automated testing and controlled linear traffic shifting for Lambda plus.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on custom orchestration and manual traffic management rather than a managed deployment strategy with built-in rollback.
● Creates undifferentiated heavy lifting and lacks the standardized deployment controls and automated rollback provided by managed CI/CD services.
● Violates the requirement for incremental traffic shifting because all traffic moves to the new version immediately.
Workflow: Create an AWS CodePipeline with a CodeCommit source trigger → add AWS CodeBuild for automated tests, and use AWS CodeDeploy to update the Lambda functions with the predefined CodeDeployDefault.LambdaLinear10PercentEvery1Minute configuration.
a target tracking scaling policy for the Spot Fleet that keeps average
Target tracking automatically adjusts capacity to hold the fleet's average CPU near the specified target and can be managed via IaC. Scheduled actions are ideal for.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Target tracking automatically adjusts capacity to hold the fleet's average CPU near the specified target and can be.
● Scene fit: Target tracking automatically adjusts capacity to hold the fleet's average CPU near the specified target and can be.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Time-based approach cannot maintain a CPU utilization target as demand varies, so it does not meet the dynamic 55.
● Application Auto Scaling does not support step scaling for Lambda, so this will not work for that service.
Workflow: Create a target tracking scaling policy for the Spot Fleet that keeps average CPU close to 55 percent using the CLI, SDKs, or CloudFormation → Use scheduled scaling in Application Auto Scaling.
Cache purge in CodeDeploy AppSpec BeforeInstall + Build in AWS CodeBuild and
Use the BeforeInstall lifecycle hook to clear caches on each instance before installing the new version. CodeBuild produces artifacts and CodeDeploy rolls out to EC2 with.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use the BeforeInstall lifecycle hook to clear caches on each instance before installing the new version.
● Scene fit: Use the BeforeInstall lifecycle hook to clear caches on each instance before installing the new version. CodeBuild produces.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Can automate builds but lacks full pipeline orchestration, stage tracking, and native rollback.
● General workflow service without native CI/CD pipeline features, UI, or integrated deployment rollback like CodePipeline/CodeDeploy.
● Manages app environments but does not provide detailed multi-stage pipeline visibility or fine-grained cache-clearing hooks per deployment.
Workflow: Cache purge in CodeDeploy AppSpec BeforeInstall → Build in AWS CodeBuild and deploy with AWS CodeDeploy → AWS CodePipeline with Git source for orchestration.
NoEcho on sensitive CloudFormation parameters + Use Secrets Manager with CloudFormation dynamic
NoEcho masks parameter values with asterisks so they do not appear in events, console views, or Describe APIs. Dynamic references resolve at deploy time without persisting plaintext in the template, stack events, or metadata.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: NoEcho masks parameter values with asterisks so they do not appear in events, console views, or Describe APIs.
● Scene fit: NoEcho masks parameter values with asterisks so they do not appear in events, console views, or Describe APIs.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Encrypting the template at rest does not prevent plaintext values from showing in stack events or API outputs.
● Custom resource return data and logs can surface secret values in events and CloudWatch Logs.
● SSM parameter types can resolve to plaintext values that may appear in events unless combined with NoEcho or dynamic references.
Workflow: Set NoEcho on sensitive CloudFormation parameters → Use Secrets Manager with CloudFormation dynamic references.
EC2 Dedicated Hosts with tags and an AWS Config custom rule (Lambda)
Dedicated Hosts expose host/socket inventory and placement control; a custom AWS Config rule can continuously detect and surface noncompliant launches.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Dedicated Hosts expose host/socket inventory and placement control; a custom AWS Config rule can continuously detect and surface noncompliant launches.
● Scene fit: Dedicated Hosts expose host/socket inventory and placement control; a custom AWS Config rule can continuously detect and surface noncompliant launches.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● SCPs restrict APIs and Tag Policies validate tag formats, but they cannot enforce Dedicated Hosts tenancy or perform continuous configuration compliance checks.
● Dedicated Instances provide single-tenant isolation but lack host/socket visibility and control needed for per-socket licensing.
● Service Catalog constrains provisioning but does not provide ongoing drift detection and compliance monitoring.
Workflow: EC2 Dedicated Hosts with tags and an AWS Config custom rule (Lambda).
CloudFormation dynamic reference to SSM Parameter Store for AMI ID
CloudFormation can resolve SSM Parameter Store values at deploy time, including AWS::SSM::Parameter::Value<AWS::EC2::Image::Id>.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation can resolve SSM Parameter Store values at deploy time, including AWS::SSM::Parameter::Value<AWS::EC2::Image::Id>.
● Scene fit: CloudFormation can resolve SSM Parameter Store values at deploy time, including AWS::SSM::Parameter::Value<AWS::EC2::Image::Id>.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is a custom approach that rewrites templates and adds unnecessary complexity and maintenance.
● CloudFormation does not directly pull the latest AMI from Image Builder without publishing to SSM or using a custom resource.
● Service Catalog does not make CloudFormation templates auto-resolve latest AMI IDs by itself.
Workflow: CloudFormation dynamic reference to SSM Parameter Store for AMI ID.
Elastic Load Balancing health checks on the Auto Scaling group so instances
When ELB health checks are enabled for the Auto Scaling group, instances flagged unhealthy by the ALB are marked unhealthy and replaced automatically. CloudWatch does not.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: When ELB health checks are enabled for the Auto Scaling group, instances flagged unhealthy by the ALB are.
● Scene fit: When ELB health checks are enabled for the Auto Scaling group, instances flagged unhealthy by the ALB are.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Scaling on CPU does not repair failing instances or provide memory usage alerts, so it does not address the.
● EC2 auto recovery based on status checks addresses underlying host or OS issues rather than application memory leaks and.
Workflow: Enable Elastic Load Balancing health checks on the Auto Scaling group so instances failing the ALB target group check are terminated and replaced → Install and configure the CloudWatch agent on the.
Mirror the dependencies to Amazon S3, attach an IAM role to the
Hosting dependencies in S3 and using a gateway VPC endpoint keeps traffic on the AWS network with no public Internet path and uses least-privilege IAM access.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Hosting dependencies in S3 and using a gateway VPC endpoint keeps traffic on the AWS network with no.
● Scene fit: Hosting dependencies in S3 and using a gateway VPC endpoint keeps traffic on the AWS network with no.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● A NAT gateway still enables outbound Internet access, which violates the requirement to prevent any Internet connectivity.
● An egress-only internet gateway is IPv6-only and still allows outbound Internet traffic, which does not satisfy the no-Internet requirement.
● Temporarily whitelisting outbound addresses still requires Internet egress and introduces operational risk because AWS Config only monitors and does not enforce remediation.
Workflow: Mirror the dependencies to Amazon S3 → attach an IAM role to the instances, and access the bucket via an S3 gateway VPC endpoint.
an Application Load Balancer with a CodeDeploy blue/green deployment, associate the Auto
This launches a replacement fleet first, load-balances across AZs with auto healing, shifts at least half of traffic, cleans up before registration, and terminates the old.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This launches a replacement fleet first, load-balances across AZs with auto healing, shifts at least half of traffic.
● Scene fit: This launches a replacement fleet first, load-balances across AZs with auto healing, shifts at least half of traffic.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Performs an in-place update, uses a risky all-at-once configuration, and attempts to run scripts in a reserved BlockTraffic event.
● Instance Refresh cannot run CodeDeploy lifecycle hooks for pre-traffic cleanup and does not provide blue/green termination controls for the.
Workflow: Use an Application Load Balancer with a CodeDeploy blue/green deployment, associate the Auto Scaling group and target group to the deployment group → create a custom configuration with 50% minimum healthy hosts.
a Lambda gate in CodePipeline that calls the AWS Health API for
A health-aware precheck fails fast during active regional incidents to prevent mid-run stalls and unnecessary cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A health-aware precheck fails fast during active regional incidents to prevent mid-run stalls and unnecessary cost.
● Scene fit: A health-aware precheck fails fast during active regional incidents to prevent mid-run stalls and unnecessary cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Automatic retries can repeatedly fail during an ongoing regional health event and waste time and resources rather than preventing the issue.
● Relying on manual scheduling introduces delays and does not cover unplanned incidents or changing health events.
● Blue green reduces cutover risk but adds cost and does not prevent the pipeline from failing during a regional health event.
Workflow: Add a Lambda gate in CodePipeline that calls the AWS Health API for the target Region and stops the run early when an active event could affect the deployment.
In the pipeline, add an AWS CodeDeploy action to roll out the
A gated STAGE deployment with a manual approval action enables hands-on validation and blocks promotion of faulty releases. CodeBuild can run automated unit and functional tests.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A gated STAGE deployment with a manual approval action enables hands-on validation and blocks promotion of faulty releases.
● Scene fit: A gated STAGE deployment with a manual approval action enables hands-on validation and blocks promotion of faulty releases.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● GuardDuty is for threat detection and has no Runtime Behavior Analysis package nor functional test capability, so it will.
● Macie focuses on sensitive data discovery and classification, not application functionality or deployment verification.
● Amazon Inspector evaluates vulnerabilities and exposure rather than functional correctness, and Runtime Behavior Analysis is not used for app.
Workflow: In the pipeline → add an AWS CodeDeploy action to roll out the latest build to the STAGE environment, insert a manual approval step for QA validation, → add a CodeDeploy action.
AWS SAM with CodeDeploy canary shifts, hooks, and CloudWatch alarm rollback
SAM DeploymentPreference integrates CodeDeploy to shift Lambda alias traffic, run pre/post-traffic hooks, and rollback automatically on alarms.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SAM DeploymentPreference integrates CodeDeploy to shift Lambda alias traffic, run pre/post-traffic hooks, and rollback automatically on alarms.
● Scene fit: SAM DeploymentPreference integrates CodeDeploy to shift Lambda alias traffic, run pre/post-traffic hooks, and rollback automatically on alarms.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Change sets preview updates and stacks can roll back, but there is no Lambda alias traffic shifting or lifecycle validation hooks.
● AppConfig toggles configuration but does not publish Lambda versions, shift alias traffic, or auto-rollback based on Lambda errors.
● API Gateway canary targets stage configuration, not Lambda version deployments or rollback of Lambda code.
Workflow: AWS SAM with CodeDeploy canary shifts, hooks, and CloudWatch alarm rollback.
Auto Scaling lifecycle hooks with EventBridge to trigger Lambda and update table
Lifecycle hooks emit per-instance events and tokens, and EventBridge can reliably route to Lambda for launch and terminate updates.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lifecycle hooks emit per-instance events and tokens, and EventBridge can reliably route to Lambda for launch and terminate updates.
● Scene fit: Lifecycle hooks emit per-instance events and tokens, and EventBridge can reliably route to Lambda for launch and terminate updates.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Scripts may not execute on termination or during failures, making updates unreliable.
● DesiredCapacity is an aggregate metric and does not provide per-instance lifecycle context.
● CloudTrail delivery can lag and lacks lifecycle tokens, risking late or missed updates.
Workflow: Auto Scaling lifecycle hooks with EventBridge to trigger Lambda and update table foo.
AWS SAM with CodeDeploy Canary10Percent10Minutes
SAM integrates Lambda aliases with CodeDeploy, CloudWatch alarms, canary shifting, and automatic rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SAM integrates Lambda aliases with CodeDeploy, CloudWatch alarms, canary shifting, and automatic rollback.
● Scene fit: SAM integrates Lambda aliases with CodeDeploy, CloudWatch alarms, canary shifting, and automatic rollback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Weighted aliases are technically workable, but CloudFormation alone does not provide the same built-in health monitoring and automated rollback workflow.
● Supports incremental deployment, but moving 50% at each step exposes too much traffic.
● AppConfig controls configuration or feature exposure, not Lambda code-version traffic shifting.
Workflow: AWS SAM with CodeDeploy Canary10Percent10Minutes.
NLB access logs to an S3 bucket, turn on SSE-S3, allow delivery.logs.amazonaws.com
This follows AWS guidance by using S3 as the destination, granting the required service principal write permission, encrypting at rest, and limiting read access to the.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This follows AWS guidance by using S3 as the destination, granting the required service principal write permission, encrypting.
● Scene fit: This follows AWS guidance by using S3 as the destination, granting the required service principal write permission, encrypting.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Allows service writes but does not grant the team read access and may also require KMS key policy updates.
● NLB access logs do not deliver to CloudWatch Logs and are supported only for delivery to S3.
● Secures delivery and storage but fails to provide the platform team with permissions to read the logs.
Workflow: Enable NLB access logs to an S3 bucket → turn on SSE-S3 → allow delivery.logs.amazonaws.com to write via bucket policy, and grant the platform team read access with IAM.
Kinesis Data Firehose with a Lambda transform, ship logs with Kinesis Agent
Firehose can invoke Lambda to normalize records in-flight and write directly to S3, with per-GB pricing and minimal operations.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Firehose can invoke Lambda to normalize records in-flight and write directly to S3, with per-GB pricing and minimal operations.
● Scene fit: Firehose can invoke Lambda to normalize records in-flight and write directly to S3, with per-GB pricing and minimal operations.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Glue streaming jobs run continuously and incur DPU-hour costs plus Data Streams costs, typically higher than a managed delivery stream for simple in-flight transforms.
● Spectrum works at query time and typically requires a Redshift cluster and external tables, adding cost and not performing ingestion-time normalization.
● Object Lambda transforms on retrieval, not during ingestion, and does not persist normalized data back into S3.
Workflow: Use Kinesis Data Firehose with a Lambda transform, ship logs with Kinesis Agent, and deliver to S3.
Jenkins masters across multiple AZs; run builds in AWS CodeBuild via the
Multi-AZ controllers plus CodeBuild provides HA and elastic, pay-per-build execution with minimal ops.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Multi-AZ controllers plus CodeBuild provides HA and elastic, pay-per-build execution with minimal ops.
● Scene fit: Multi-AZ controllers plus CodeBuild provides HA and elastic, pay-per-build execution with minimal ops.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Can be HA, but you still operate and scale agent nodes, increasing cost and effort compared to managed builds.
● Single-AZ control plane is a single point of failure and not highly available.
● HA is achieved, but managing and scaling EC2 agents adds overhead and can be less cost-effective.
Workflow: Jenkins masters across multiple AZs → run builds in AWS CodeBuild via the Jenkins plugin.
Update the Auto Scaling launch template to include TagSpecifications for EBS volumes
TagSpecifications in the launch template ensure EBS volumes are tagged at creation by the service with no extra automation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: TagSpecifications in the launch template ensure EBS volumes are tagged at creation by the service with no extra automation.
● Scene fit: TagSpecifications in the launch template ensure EBS volumes are tagged at creation by the service with no extra automation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Works but adds operational complexity and tags after creation rather than at launch.
● Tag propagation from the Auto Scaling group only applies to EC2 instances and not to attached EBS volumes.
● AWS Config evaluates compliance but cannot block volume creation events.
Workflow: Update the Auto Scaling launch template to include TagSpecifications for EBS volumes with the required cost center tags.
an Auto Scaling group with the hardened AMI, apply target tracking on
This combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable periods, improving both cost efficiency and reliability.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable periods, improving both cost efficiency and.
● Scene fit: This combines dynamic scaling for unpredictable surges with scheduled right-sizing for predictable periods, improving both cost efficiency and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Provides time-based starts and stops but lacks elasticity and health-based recovery, so it does not improve reliability during unpredictable.
● Reduces compute cost commitments but does not add automatic scaling or health-driven replacement to handle variable demand.
● Scheduled capacity is inflexible for bursty global demand and still relies on scripts; moreover, Scheduled Instances have been retired.
Workflow: Create an Auto Scaling group with the hardened AMI → apply target tracking on average CPU, and add scheduled actions to set the minimum to 4 after hours and to 8 before the workday.
After upload, read the object ETag and compare it to a locally
For single-part uploads without special encryption, the ETag equals the MD5 of the object and can be compared to a locally computed digest. Setting Content-MD5 makes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: For single-part uploads without special encryption, the ETag equals the MD5 of the object and can be compared.
● Scene fit: For single-part uploads without special encryption, the ETag equals the MD5 of the object and can be compared.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 does not validate user metadata against the payload, so a metadata hash does not enforce integrity.
● Clients do not supply an MD5 trailing checksum to S3; trailing checksums are SDK-managed for other checksum algorithms, not.
● A version ID is an identifier for object versions and has no relationship to a content checksum.
Workflow: After upload, read the object ETag and compare it to a locally computed MD5 of the payload → Supply the MD5 value in the Content-MD5 header on PutObject and act on the.
EventBridge target: SSM Automation runbook to disjoin AD and create AMI +
EventBridge can trigger a Systems Manager Automation runbook to perform the pre-termination tasks. A lifecycle hook pauses termination in Terminating:Wait and emits an event that EventBridge can route to automation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can trigger a Systems Manager Automation runbook to perform the pre-termination tasks.
● Scene fit: EventBridge can trigger a Systems Manager Automation runbook to perform the pre-termination tasks. A lifecycle hook pauses termination.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● SNS termination notifications do not pause termination, so cleanup can race or be skipped without a lifecycle wait.
● Image Builder creates golden AMIs via pipelines and is not designed to run per-instance cleanup at scale-in.
● Maintenance Windows are scheduled and not tied to Auto Scaling scale-in events or lifecycle pauses.
Workflow: EventBridge target: SSM Automation runbook to disjoin AD and create AMI → Auto Scaling lifecycle hook (Terminating:Wait) with EventBridge rule.
the rule target to an AWS Lambda function that calls a Slack
Lambda can securely transform and forward EventBridge events to Slack, which EventBridge does not natively support. This pattern captures failures at the pipeline execution level, which is exactly what is needed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda can securely transform and forward EventBridge events to Slack, which EventBridge does not natively support.
● Scene fit: Lambda can securely transform and forward EventBridge events to Slack, which EventBridge does not natively support. This pattern.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Matches individual action failures rather than end-to-end pipeline failures, so it does not align with the requirement.
● EventBridge does not provide a native Slack target, so this configuration is not supported.
● An email subscription does not deliver messages to Slack, which is the stated destination.
Workflow: Set the rule target to an AWS Lambda function that calls a Slack incoming webhook for the #platform-eng channel → Configure an Amazon EventBridge rule that matches CodePipeline Pipeline Execution State Change events with state FAILED.
Manage Lambda versions/aliases with CloudFormation and use API Gateway stage canary
Lambda aliases support weighted traffic between versions and API Gateway stage canary can shift a small percentage before promotion; both are natively managed via CloudFormation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Lambda aliases support weighted traffic between versions and API Gateway stage canary can shift a small percentage before promotion; both are natively managed via CloudFormation.
● Scene fit: Lambda aliases support weighted traffic between versions and API Gateway stage canary can shift a small percentage before.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AppConfig controls configuration and feature toggles but does not route HTTP traffic or perform percentage-based traffic shifting at API Gateway or Lambda.
● Route 53 failover is binary active-passive, not percentage-based canary; CDK is an IaC tool and does not change that routing behavior.
● All-at-once and simple routing move 100% of users immediately, providing no limited exposure phase.
Workflow: Manage Lambda versions/aliases with CloudFormation and use API Gateway stage canary.
Place the EC2 fleet in a Multi-AZ Auto Scaling group behind an
This design provides fault-tolerant scaling for the app, high-performance multi-master Aurora for the database, and continuous vulnerability assessment with Amazon Inspector.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This design provides fault-tolerant scaling for the app, high-performance multi-master Aurora for the database, and continuous vulnerability assessment.
● Scene fit: This design provides fault-tolerant scaling for the app, high-performance multi-master Aurora for the database, and continuous vulnerability assessment.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Improves app and database resilience, but Macie focuses on sensitive data discovery in Amazon S3 and does not provide.
● RDS for PostgreSQL does not offer a true multi-master configuration, so this does not meet the stated high availability.
● GuardDuty is a threat detection service that analyzes telemetry rather than performing host and package vulnerability scanning.
Workflow: Place the EC2 fleet in a Multi-AZ Auto Scaling group behind an Application Load Balancer → use Amazon Aurora with multi-master writers for higher throughput and high availability, and enable Amazon Inspector for ongoing vulnerability scans.
Aurora MySQL Global Database for catalog; Region-local Aurora for PII and orders
This keeps the relational model, provides a single global catalog via Aurora Global Database, and stores regulated data in Region-local Aurora clusters.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This keeps the relational model, provides a single global catalog via Aurora Global Database, and stores regulated data in Region-local Aurora clusters.
● Scene fit: This keeps the relational model, provides a single global catalog via Aurora Global Database, and stores regulated data in Region-local Aurora clusters.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Redshift targets analytics, not OLTP, and mixing it with DynamoDB requires major refactoring.
● Moving from Aurora MySQL to DynamoDB changes data models and queries significantly.
● S3 is not a transactional relational store and does not meet OLTP requirements for the catalog.
Workflow: Aurora MySQL Global Database for catalog → Region-local Aurora for PII and orders.
the Auto Scaling group to use ELB health checks that validate the
Using ELB health checks ensures the ASG replaces instances based on the real application health probe on the correct port and path. Target group health checks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using ELB health checks ensures the ASG replaces instances based on the real application health probe on the correct port and path.
● Scene fit: Using ELB health checks ensures the ASG replaces instances based on the real application health probe on the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Switching to IP targets does not address application reachability when health checks are misconfigured.
● An ALB supports only HTTP or HTTPS listeners, not TCP.
● Migrating to NLB is unnecessary here and would remove HTTP-level path checks that help validate the application endpoint.
Workflow: Configure the Auto Scaling group to use ELB health checks that validate the service on the custom port and path → Update the target group health check to probe the application's custom port and a specific health endpoint.
an Amazon EventBridge rule that matches Auto Scaling instance launch and terminate
This wires scale events to Lambda to update the S3 page in real time and stores searchable logs in CloudWatch Logs.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This wires scale events to Lambda to update the S3 page in real time and stores searchable logs.
● Scene fit: This wires scale events to Lambda to update the S3 page in real time and stores searchable logs.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Captures events but relies on manual or scheduled exports and does not update the S3 webpage automatically.
● Amazon S3 is not a supported direct target for EventBridge rules, so the page cannot be updated this way.
● CloudTrail records API activity but does not update an S3 webpage and is not ideal for near real-time operational.
Workflow: Use an Amazon EventBridge rule that matches Auto Scaling instance launch and terminate events to invoke an AWS Lambda function → the function updates the S3 status page and writes event records to Amazon CloudWatch Logs.
an Amazon Aurora Global Database for the catalog table and keep accounts
Aurora Global Database provides low-latency global reads and fast cross-Region replication for catalog while preserving SQL and minimizing code changes for regional tables.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora Global Database provides low-latency global reads and fast cross-Region replication for catalog while preserving SQL and minimizing code changes for regional tables.
● Scene fit: Aurora Global Database provides low-latency global reads and fast cross-Region replication for catalog while preserving SQL and minimizing code changes for regional tables.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Mixes SQL and NoSQL and would force query and driver changes for the catalog portion, increasing refactoring effort.
● Migrating all tables to DynamoDB replaces the relational model and SQL with NoSQL, requiring major application rewrites.
● Combining Aurora and DynamoDB introduces two data models and APIs, which increases complexity and code changes.
Workflow: Use an Amazon Aurora Global Database for the catalog table and keep accounts and view_history in Amazon Aurora.
environment-specific patch baselines in AWS Systems Manager Patch Manager, tag instances by
This leverages Patch Manager with Patch Groups and per-environment baselines to target nonproduction first and enforce differing policies with minimal configuration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This leverages Patch Manager with Patch Groups and per-environment baselines to target nonproduction first and enforce differing policies with minimal configuration.
● Scene fit: This leverages Patch Manager with Patch Groups and per-environment baselines to target nonproduction first and enforce differing policies.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on custom scripting and manual baseline logic, increasing operational effort compared to using Patch Manager.
● Tagging only by OS cannot isolate Production from nonproduction, risking unintended patching of Production.
● Maintenance Windows provides scheduling but still requires custom patch selection and targeting logic without Patch Manager baselines.
Workflow: Create environment-specific patch baselines in AWS Systems Manager Patch Manager, tag instances by environment, business unit, and OS, organize them into Patch Groups, and apply the corresponding baseline to each group.
Golden AMI with app preinstalled; runtime config via user data
Prebaking the application eliminates most bootstrap time while user data injects environment-specific settings at boot.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Prebaking the application eliminates most bootstrap time while user data injects environment-specific settings at boot.
● Scene fit: Prebaking the application eliminates most bootstrap time while user data injects environment-specific settings at boot.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Still performs the lengthy install on each instance and does not remove the 27-minute setup time.
● Baking volatile cache data makes images stale and inconsistent across the fleet.
● Orchestrating Lambda to preload instance-local caches is brittle and still couples to frequently changing data.
Workflow: Golden AMI with app preinstalled → runtime config via user data.
Amazon Aurora MySQL with cross-Region read replicas for catalog; separate Aurora clusters
Aurora MySQL is MySQL-compatible for minimal changes, supports cross-Region replicas for low-latency reads, and per-Region clusters keep orders resident.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora MySQL is MySQL-compatible for minimal changes, supports cross-Region replicas for low-latency reads, and per-Region clusters keep orders resident.
● Scene fit: Aurora MySQL is MySQL-compatible for minimal changes, supports cross-Region replicas for low-latency reads, and per-Region clusters keep orders resident.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Switching from relational MySQL to NoSQL would require significant schema and code changes, not minimizing changes.
● Is relational and compatible but typically has higher replica lag and more operational overhead than Aurora for global read scaling.
● Global Database would replicate orders to multiple Regions, violating the need to keep order data in the Region of origin.
Workflow: Amazon Aurora MySQL with cross-Region read replicas for catalog → separate Aurora clusters per Region for orders.
a CloudWatch Logs subscription filter that invokes an AWS Lambda function to
This provides end-to-end automation by reacting to logs in real time and periodically terminating tagged instances without manual steps or extra servers.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This provides end-to-end automation by reacting to logs in real time and periodically terminating tagged instances without manual.
● Scene fit: This provides end-to-end automation by reacting to logs in real time and periodically terminating tagged instances without manual.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on human intervention after an SNS notification, which is not fully automated.
● CloudWatch Alarms cannot directly publish to SQS and maintaining a worker fleet is unnecessary for this use case.
● CloudWatch Logs subscriptions cannot directly deliver events to Step Functions, making this design invalid.
Workflow: Create a CloudWatch Logs subscription filter that invokes an AWS Lambda function to tag the source instance with MARK_FOR_TERMINATION when an SSH login is detected, and use an Amazon EventBridge scheduled rule.
Put the AMI ID in AWS Systems Manager Parameter Store and use
CloudFormation can directly resolve SSM Parameter Store values at stack update, enabling a simple scheduled UpdateStack to roll environments to the latest AMI without template edits.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation can directly resolve SSM Parameter Store values at stack update, enabling a simple scheduled UpdateStack to roll.
● Scene fit: CloudFormation can directly resolve SSM Parameter Store values at stack update, enabling a simple scheduled UpdateStack to roll.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach requires parsing and modifying each template on every cycle, which is brittle and does not scale given the.
● CloudFormation cannot natively resolve parameters from AWS AppConfig, so stacks would not automatically consume the value.
● Passing raw string parameters requires knowing each template’s parameter names, which is unmanageable when templates are inconsistent.
Workflow: Put the AMI ID in AWS Systems Manager Parameter Store and use an SSM parameter type in CloudFormation so the value is resolved at update time → trigger a scheduled EventBridge rule.
AWS DevOps Engineering Decision Patterns
Lecture materials covering 500 technical scenarios. Each lesson evaluates technical viability, requirement fit, scene fit, engineering common sense, rejected alternatives, and the practical workflow.
Install and configure the unified CloudWatch agent on each instance to stream
The unified CloudWatch agent natively collects both logs and additional metrics from Linux and Windows and integrates directly with CloudWatch Logs Insights for minimal operational overhead.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: The unified CloudWatch agent natively collects both logs and additional metrics from Linux and Windows and integrates directly.
● Scene fit: The unified CloudWatch agent natively collects both logs and additional metrics from Linux and Windows and integrates directly.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● SSM Agent is used for Systems Manager operations and is not the recommended or native agent for log shipping to CloudWatch Logs.
● Can work but introduces extra moving parts and setup compared to CloudWatch Agent and CloudWatch Logs Insights, so it is not the minimal-effort path.
● Amazon Inspector focuses on vulnerability and exposure assessments rather than collecting and analyzing system and application logs.
Workflow: Install and configure the unified CloudWatch agent on each instance to stream logs to CloudWatch Logs, → query with CloudWatch Logs Insights.
Build an image for the AWS X-Ray daemon, push it to ECR
Running the X-Ray daemon as a sidecar container and exposing UDP 2000 in the task definition is the correct pattern for collecting traces from ECS tasks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Running the X-Ray daemon as a sidecar container and exposing UDP 2000 in the task definition is the.
● Scene fit: Running the X-Ray daemon as a sidecar container and exposing UDP 2000 in the task definition is the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach places an Elastic Beanstalk-specific configuration file in the image and adjusts ports, but it is not the recommended.
● CloudTrail records AWS API activity and is not used for application-level distributed tracing or measuring request path latency.
Workflow: Build an image for the AWS X-Ray daemon, push it to ECR → run it as a sidecar in the ECS task, and open UDP 2000 via task definition port mappings.
AWS CloudTrail organization trail + AWS Config with organization aggregator
CloudTrail provides an organization-wide audit log of management events across all regions for who did what and when. Config records configuration changes and evaluates them against rules, with an aggregator for centralized multi-account compliance.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail provides an organization-wide audit log of management events across all regions for who did what and when.
● Scene fit: CloudTrail provides an organization-wide audit log of management events across all regions for who did what and when.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Audit Manager maps controls to frameworks and collects evidence but does not perform continuous configuration evaluation.
● EventBridge and Lambda can automate remediation, but this is not the primary service for continuous configuration assessment and auditing.
● Security Hub aggregates security findings from multiple services but does not track or evaluate detailed resource configuration drift.
Workflow: AWS CloudTrail organization trail → AWS Config with organization aggregator.
a nested CloudFormation template that defines an Auto Scaling group with min
A one-instance ASG per broker preserves state by reattaching the correct EBS volume via tags while providing health-based replacement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A one-instance ASG per broker preserves state by reattaching the correct EBS volume via tags while providing health-based.
● Scene fit: A one-instance ASG per broker preserves state by reattaching the correct EBS volume via tags while providing health-based.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFormation does not automatically recreate instances that are terminated outside of a stack update, so attachments and recovery would.
● A single ASG can rebalance into AZs without a matching preprovisioned volume, and EBS volumes cannot move across AZs.
● Drift detection is informational and does not guarantee automated recreation or stateful reattachment of EBS volumes.
Workflow: Use a nested CloudFormation template that defines an Auto Scaling group with min and max 1 and one EBS volume tagged to match that group, include a bootstrap script to attach the.
Install SSM Agent on on‑prem Windows servers, register with activation code/ID so
On‑prem servers must be registered via hybrid activation and then patched with Patch Manager. A single SSM service role trusted for STS AssumeRole is required to.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: On‑prem servers must be registered via hybrid activation and then patched with Patch Manager.
● Scene fit: On‑prem servers must be registered via hybrid activation and then patched with Patch Manager. A single SSM service.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Quick Setup does not eliminate the need for hybrid activation and a service role for on‑prem servers.
● Hybrid instances appear with mi‑ prefixes, and Patch Manager is the primary service for patching.
● Systems Manager management of on‑prem servers is not agentless; SSM Agent and hybrid activation are required.
Workflow: Install SSM Agent on on‑prem Windows servers → register with activation code/ID so they appear as mi‑ instances, → patch via Patch Manager → Create one IAM role for Systems Manager (assumable.
Amazon Cognito identity pool with role-based S3 and DynamoDB access + STS
Identity pools federate logins and return STS temporary credentials mapped to least-privilege IAM roles. Exchanges OIDC tokens for temporary role credentials granting scoped access to S3 and DynamoDB.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Identity pools federate logins and return STS temporary credentials mapped to least-privilege IAM roles.
● Scene fit: Identity pools federate logins and return STS temporary credentials mapped to least-privilege IAM roles. Exchanges OIDC tokens for.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Directory Service is for enterprise directories and does not broker social or OIDC identities to AWS credentials for a public mobile app.
● User pools handle authentication but their tokens cannot call AWS services directly without identity pools or STS.
● SAML targets enterprise IdPs and is not the right fit for typical social or OIDC mobile logins.
Workflow: Amazon Cognito identity pool with role-based S3 and DynamoDB access → STS AssumeRoleWithWebIdentity with OIDC providers.
Send Content-MD5 in PutObject and proceed only on success + Compare the
Supplying Content-MD5 makes S3 validate the MD5 server-side and reject mismatches. For single-part, unencrypted uploads, the ETag equals the object MD5, enabling verification.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Supplying Content-MD5 makes S3 validate the MD5 server-side and reject mismatches.
● Scene fit: Supplying Content-MD5 makes S3 validate the MD5 server-side and reject mismatches. For single-part, unencrypted uploads, the ETag equals the object MD5, enabling verification.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● User metadata is not validated against the payload, so a 200 response does not confirm integrity.
● Version IDs are identifiers, not content checksums.
● Multipart ETags are not the MD5 of the full object and cannot be used for MD5 validation.
Workflow: Send Content-MD5 in PutObject and proceed only on success → Compare the object ETag to a locally computed MD5 after upload.
a BeforeAllowTraffic lifecycle hook in the Lambda AppSpec that waits for the
BeforeAllowTraffic runs prior to alias shift, allowing you to block traffic until the Step Functions execution and Fargate task have fully completed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: BeforeAllowTraffic runs prior to alias shift, allowing you to block traffic until the Step Functions execution and Fargate task have fully completed.
● Scene fit: BeforeAllowTraffic runs prior to alias shift, allowing you to block traffic until the Step Functions execution and Fargate.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AfterAllowTraffic runs only after the alias has moved, so the new version would already be serving requests, which violates the requirement.
● A canary still routes a portion of traffic to the new version while the workflow runs, risking failures before the reorganization completes.
● There is no supported Step Functions-to-CodeDeploy signal to advance a Lambda deployment; CodeDeploy controls the lifecycle via hooks.
Workflow: Add a BeforeAllowTraffic lifecycle hook in the Lambda AppSpec that waits for the Step Functions execution to complete.
Resources were changed outside of CloudFormation and the template was not updated
Out-of-band changes create drift that can block CloudFormation from reverting resources to their previous configuration during rollback. Missing permissions can prevent CloudFormation from modifying or restoring.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Out-of-band changes create drift that can block CloudFormation from reverting resources to their previous configuration during rollback.
● Scene fit: Out-of-band changes create drift that can block CloudFormation from reverting resources to their previous configuration during rollback. Missing.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFormation uses regional service APIs and does not require an interface VPC endpoint to update VPC resources, so an endpoint outage is unlikely to cause this rollback state.
● Change sets are optional planning tools and their absence does not itself cause rollback failures.
● AWS Config is not a prerequisite for CloudFormation updates or rollbacks and its absence does not produce UPDATE_ROLLBACK_FAILED.
Workflow: Resources were changed outside of CloudFormation and the template was not updated → The IAM user or role that ran the update lacked some required permissions.
the app with AWS Elastic Beanstalk using a multi Availability Zone environment
A prebuilt custom AMI drastically shortens instance startup while Elastic Beanstalk provides resilient multi AZ scaling and managed load balancing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A prebuilt custom AMI drastically shortens instance startup while Elastic Beanstalk provides resilient multi AZ scaling and managed.
● Scene fit: A prebuilt custom AMI drastically shortens instance startup while Elastic Beanstalk provides resilient multi AZ scaling and managed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Reduces server management but requires containerizing the app and re-architecting, which may not immediately improve provisioning time.
● Installing dependencies at boot still leads to long warm-up times and does not solve the slow instance provisioning problem.
● Spot capacity can be interrupted and a fixed target capacity lacks elasticity, which undermines resilience and availability.
Workflow: Deploy the app with AWS Elastic Beanstalk using a multi Availability Zone environment behind a load balancer, and launch instances from an AWS Systems Manager Automation built custom AMI that includes all.
up a CloudWatch metric stream to a Kinesis Data Firehose delivery stream
Metric streams to Firehose deliver metrics to S3 without custom code, enabling durable multi-year storage queried by Athena and visualized in QuickSight with minimal effort.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric streams to Firehose deliver metrics to S3 without custom code, enabling durable multi-year storage queried by Athena.
● Scene fit: Metric streams to Firehose deliver metrics to S3 without custom code, enabling durable multi-year storage queried by Athena.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach works but requires custom code, scheduling, permissions, and ongoing maintenance, so it is not the lowest-effort option.
● CloudWatch metrics do not offer an Extended Retention setting, so this proposal cannot meet multi-year retention or compliance needs.
● CloudWatch dashboards cannot consume data stored in S3, so this design cannot produce dashboards from exported metrics.
Workflow: Set up a CloudWatch metric stream to a Kinesis Data Firehose delivery stream that stores all metrics in Amazon S3, analyze ranges with Athena, and create QuickSight dashboards.
Failover routing with alias targets and Evaluate Target Health + Route 53
Failover policy with alias records and Evaluate Target Health returns the secondary when the primary is unhealthy. Non-alias records require health checks and network allowlisting so Route 53 can assess health.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Failover policy with alias records and Evaluate Target Health returns the secondary when the primary is unhealthy.
● Scene fit: Failover policy with alias records and Evaluate Target Health returns the secondary when the primary is unhealthy. Non-alias records require health checks and network allowlisting so Route 53 can assess health.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Improves availability/performance but does not configure Route 53 DNS failover.
● Custom automation is unnecessary because Route 53 handles failover natively.
● Returns multiple healthy answers but does not enforce primary-to-secondary failover.
Workflow: Failover routing with alias targets and Evaluate Target Health → Route 53 health checks for non-alias records → allow health checker IPs.
the AWS-RunPatchBaseline SSM document to enforce approved patch baselines across the fleet
AWS-RunPatchBaseline works with Patch Manager to apply baselines to both Windows and Linux, providing centralized, auditable patching. SSM Agent plus Maintenance Windows enables controlled, auditable, and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS-RunPatchBaseline works with Patch Manager to apply baselines to both Windows and Linux, providing centralized, auditable patching.
● Scene fit: AWS-RunPatchBaseline works with Patch Manager to apply baselines to both Windows and Linux, providing centralized, auditable patching. SSM.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudFormation does not offer automatic OS patching capabilities, so this would not fulfill the requirement despite AWS Config providing.
● While feasible, building and operating host-level workflows across a mixed fleet is high effort and not aligned with the.
● AWS-ApplyPatchBaseline targets Windows only and therefore does not meet the mixed Windows and Linux requirement.
Workflow: Use the AWS-RunPatchBaseline SSM document to enforce approved patch baselines across the fleet → Install the AWS Systems Manager Agent on every instance → validate patches in staging, → schedule patching through.
Copy the AMI in source, encrypt with the source CMK, allow target
You must encrypt with a customer managed key and permit the target account (via key policy) to create grants before sharing. Auto Scaling needs a grant.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: You must encrypt with a customer managed key and permit the target account (via key policy) to create.
● Scene fit: You must encrypt with a customer managed key and permit the target account (via key policy) to create.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS RAM does not share AMIs or re-encrypt them; AMI sharing uses EC2 launch permissions and snapshot/KMS access.
● AMIs encrypted with AWS managed KMS keys cannot be shared across accounts for launches.
● EBS default encryption does not provide cross-account permissions to the source CMK required for the encrypted AMI.
Workflow: Copy the AMI in source → encrypt with the source CMK → allow target to create grants on the key, → share the AMI → In target → create a KMS grant.
AWS Systems Manager Patch Manager with SSM Agent and maintenance windows
Provides centralized hybrid patching with baselines, patch groups, maintenance windows, and compliance reporting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides centralized hybrid patching with baselines, patch groups, maintenance windows, and compliance reporting.
● Scene fit: Provides centralized hybrid patching with baselines, patch groups, maintenance windows, and compliance reporting.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Finds and prioritizes vulnerabilities but does not orchestrate patch deployment or maintenance windows.
● Enables ad-hoc command execution but lacks managed patch baselines, patch groups, scheduling, and compliance views.
● Tracks configuration and conformance but does not perform patch orchestration.
Workflow: AWS Systems Manager Patch Manager with SSM Agent and maintenance windows.
Install the SSM Agent on the on-prem Windows servers using the activation
This is correct because hybrid nodes must be registered via activation to show as mi- managed instances and Patch Manager is the right tool for OS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is correct because hybrid nodes must be registered via activation to show as mi- managed instances and.
● Scene fit: This is correct because hybrid nodes must be registered via activation to show as mi- managed instances and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Is incorrect because hybrid activation requires a single Systems Manager service role trusted for STS AssumeRole, not multiple roles.
● Is incorrect because hybrid managed instances display with an mi- prefix and State Manager is not the primary patching.
Workflow: Install the SSM Agent on the on-prem Windows servers using the activation code and ID → register them so they appear with an mi- prefix in the console, and apply updates with.
Install Amazon Kinesis Agent on the servers to send logs to an
Kinesis Data Firehose can invoke Lambda for in-flight transformations while Kinesis Agent ships logs, providing a low-cost, serverless path to normalized data in S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Kinesis Data Firehose can invoke Lambda for in-flight transformations while Kinesis Agent ships logs, providing a low-cost, serverless.
● Scene fit: Kinesis Data Firehose can invoke Lambda for in-flight transformations while Kinesis Agent ships logs, providing a low-cost, serverless.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● QuickSight is for visualization and ad hoc analysis, not for low-cost log ingestion or transformation before landing in S3.
● Redshift Spectrum requires a Redshift cluster and external tables, adding cost and complexity for a task that should be handled pre-ingest.
● OpenSearch introduces ongoing cluster costs and is intended for search and analytics, not economical pre-S3 normalization.
Workflow: Install Amazon Kinesis Agent on the servers to send logs to an Amazon Kinesis Data Firehose delivery stream and enable a Lambda transform to normalize before writing to S3.
Register on-prem servers with Systems Manager using Hybrid Activations + Attach an
Hybrid Activations enroll on-prem hosts as managed instances so Patch Manager can target them. An instance profile gives the SSM Agent permissions to register and run.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Hybrid Activations enroll on-prem hosts as managed instances so Patch Manager can target them.
● Scene fit: Hybrid Activations enroll on-prem hosts as managed instances so Patch Manager can target them. An instance profile gives.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● EventBridge can trigger tasks, but maintenance windows in Systems Manager are the native way to constrain patching timeframes.
● Inspector identifies vulnerabilities but does not directly perform patching; remediation uses SSM Patch Manager.
● Long-lived keys on servers are insecure and not the supported onboarding method for SSM.
Workflow: Register on-prem servers with Systems Manager using Hybrid Activations → Attach an IAM instance profile to EC2 for Systems Manager access → Configure a Systems Manager Maintenance Window for off-hours.
an EC2 Auto Scaling lifecycle hook to pause in Terminating:Wait for debugging
A termination lifecycle hook moves instances to Terminating:Wait, giving time to connect, inspect, and send heartbeats before continuing termination.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A termination lifecycle hook moves instances to Terminating:Wait, giving time to connect, inspect, and send heartbeats before continuing termination.
● Scene fit: A termination lifecycle hook moves instances to Terminating:Wait, giving time to connect, inspect, and send heartbeats before continuing termination.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Only halts Availability Zone rebalancing and does not prevent health-check-based terminations.
● Delays initial health checks after launch but does not pause termination once an instance is marked unhealthy.
● Instance protection blocks scale-in but does not stop replacements due to failed health checks.
Workflow: Add an EC2 Auto Scaling lifecycle hook to pause in Terminating:Wait for debugging.
PUT SUCCESS or FAILED JSON to the CloudFormation ResponseURL
CloudFormation requires the provider to send a response to the pre-signed ResponseURL to finish the operation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation requires the provider to send a response to the pre-signed ResponseURL to finish the operation.
● Scene fit: CloudFormation requires the provider to send a response to the pre-signed ResponseURL to finish the operation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Declares the resource type but does not signal completion to CloudFormation.
● CloudFormation ignores the Lambda return value and waits for the ResponseURL callback.
● WaitCondition is a different pattern and does not complete a custom resource lifecycle.
Workflow: PUT SUCCESS or FAILED JSON to the CloudFormation ResponseURL.
Route 53 weighted aliases to send a small and growing percentage to
Weighted records let you start with a small weight and increase it to shift traffic gradually without downtime.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Weighted records let you start with a small weight and increase it to shift traffic gradually without downtime.
● Scene fit: Weighted records let you start with a small weight and increase it to shift traffic gradually without downtime.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Latency-based routing chooses the lowest-latency endpoint and cannot do fixed percentage splits.
● In-place updates replace instances in batches and do not provide user-level percentage routing.
● Canary configuration applies only to Lambda, not EC2 Auto Scaling behind ALB.
Workflow: Use Route 53 weighted aliases to send a small and growing percentage to a second ALB.
a single script that reads the DEPLOYMENT_GROUP_NAME environment variable in CodeDeploy and
DEPLOYMENT_GROUP_NAME is a built-in variable and running the script in BeforeInstall lets you configure logging early without extra tooling or per-group scripts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: DEPLOYMENT_GROUP_NAME is a built-in variable and running the script in BeforeInstall lets you configure logging early without extra.
● Scene fit: DEPLOYMENT_GROUP_NAME is a built-in variable and running the script in BeforeInstall lets you configure logging early without extra.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Using the group ID and running at ApplicationStart is later in the lifecycle and adds complexity compared to simpler built-in options earlier in the process.
● Managing tags and running CLI calls during deployment adds credentials and operational overhead and occurs later than necessary.
● CodeDeploy hook scripts do not support arbitrary custom environment variables and relying on ValidateService is not the optimal timing.
Workflow: Use a single script that reads the DEPLOYMENT_GROUP_NAME environment variable in CodeDeploy and call it in the BeforeInstall hook to set NGINX logging per group.
Renumber VPCs to non-overlapping, use AWS Transit Gateway, expose services via internal
Readdressing removes overlap and Transit Gateway provides scalable hub-and-spoke routing; internal NLBs keep HTTPS private.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Readdressing removes overlap and Transit Gateway provides scalable hub-and-spoke routing; internal NLBs keep HTTPS private.
● Scene fit: Readdressing removes overlap and Transit Gateway provides scalable hub-and-spoke routing; internal NLBs keep HTTPS private.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Peering meshes do not scale operationally and cannot connect VPCs with overlapping CIDRs.
● Cloud WAN still requires non-overlapping routing domains for end-to-end reachability and does not by itself fix overlapping CIDRs.
● PrivateLink is private and works with overlapping CIDRs but requires per-service, per-VPC endpoints and lacks transitive routing, hindering scale.
Workflow: Renumber VPCs to non-overlapping → use AWS Transit Gateway, expose services via internal NLBs.
AWS Config cloudtrail-enabled (45-minute periodic) with EventBridge triggering Lambda to call StartLogging
Periodic compliance checks detect disabled logging independent of CloudTrail delivery and Lambda promptly re-enables logging.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Periodic compliance checks detect disabled logging independent of CloudTrail delivery and Lambda promptly re-enables logging.
● Scene fit: Periodic compliance checks detect disabled logging independent of CloudTrail delivery and Lambda promptly re-enables logging.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Polls and targets DeleteTrail rather than StopLogging, causing delay and not directly restoring logging.
● Prevents future disables but does not turn logging back on if it is already stopped.
● The rule evaluates periodically, not on configuration changes, and it does not auto-remediate by default.
Workflow: AWS Config cloudtrail-enabled (45-minute periodic) with EventBridge triggering Lambda to call StartLogging.
a CloudFormation stack that creates a CodePipeline to build hardened AMIs and
Parameter Store with dynamic references provides a scalable, decoupled, and secure way for stacks to consume the latest approved AMI at deployment time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Parameter Store with dynamic references provides a scalable, decoupled, and secure way for stacks to consume the latest.
● Scene fit: Parameter Store with dynamic references provides a scalable, decoupled, and secure way for stacks to consume the latest.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Relies on S3 artifacts and cross-stack plumbing, which increases operational overhead and is less direct than parameter lookups.
● CloudFormation cannot natively resolve AMIs by tag without custom logic, making this brittle for automated, account-wide consumption.
● Tightly couples responsibilities and slows delivery because application updates would depend on Security-managed stacks.
Workflow: Use a CloudFormation stack that creates a CodePipeline to build hardened AMIs and write the current AMI ID to AWS Systems Manager Parameter Store, and have the Platform team resolve that parameter.
Elastic Beanstalk blue/green with external RDS Multi-AZ; CodeStar Connections + CodeBuild
Blue/green on Elastic Beanstalk with CNAME swap and a shared external Multi-AZ RDS provides zero-downtime cutover and rapid rollback.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Blue/green on Elastic Beanstalk with CNAME swap and a shared external Multi-AZ RDS provides zero-downtime cutover and rapid rollback.
● Scene fit: Blue/green on Elastic Beanstalk with CNAME swap and a shared external Multi-AZ RDS provides zero-downtime cutover and rapid rollback.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● In-place updates can momentarily reduce capacity and risk downtime during deployments even with Multi-AZ RDS.
● Rolling updates may still cause brief service impact and this path requires more re-architecture than necessary for a fast migration.
● Separate databases per environment complicate data consistency and cutover for stateful apps.
Workflow: Elastic Beanstalk blue/green with external RDS Multi-AZ → CodeStar Connections + CodeBuild.
a launch lifecycle hook to the Auto Scaling group and complete it
A launch lifecycle hook pauses the instance in Pending:Wait so you can finish bootstrap and then allow registration when ready.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A launch lifecycle hook pauses the instance in Pending:Wait so you can finish bootstrap and then allow registration when ready.
● Scene fit: A launch lifecycle hook pauses the instance in Pending:Wait so you can finish bootstrap and then allow registration when ready.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Slow start meters traffic after registration and does not prevent targets from being registered prematurely.
● Changing health check thresholds alters how quickly a target is considered healthy but does not gate initial registration.
● The grace period only delays Auto Scaling health evaluations and terminations and does not control load balancer registration.
Workflow: Add a launch lifecycle hook to the Auto Scaling group and complete it only after the bootstrap process signals success.
UpdatePolicy: AutoScalingRollingUpdate
Enables rolling updates for Auto Scaling groups with MinInstancesInService and MaxBatchSize.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Enables rolling updates for Auto Scaling groups with MinInstancesInService and MaxBatchSize.
● Scene fit: Enables rolling updates for Auto Scaling groups with MinInstancesInService and MaxBatchSize.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A resource attribute that controls retain/delete behavior on replacement, not rolling updates.
● Can replace all instances or the whole group rather than doing controlled rolling batches.
● A deployment service, not a CloudFormation UpdatePolicy for ASG rolling updates.
Workflow: UpdatePolicy: AutoScalingRollingUpdate.
Bake AMI with latest SSM Agent using EC2 Image Builder; attach AmazonSSMManagedInstanceCore
Provides SSM Session Manager access without internet, grants required instance permissions, and delivers auditable S3 logs with SNS notifications.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides SSM Session Manager access without internet, grants required instance permissions, and delivers auditable S3 logs with SNS notifications.
● Scene fit: Provides SSM Session Manager access without internet, grants required instance permissions, and delivers auditable S3 logs with SNS.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Requires SSH network paths and does not provide native centralized Session Manager access controls or logging in isolated subnets.
● SCPs do not grant permissions and AWS Config cannot attach SCPs, leaving the required SSM access path unresolved.
● Detects API calls but does not establish SSM connectivity, ensure the agent/permissions, or capture session logs, so it is incomplete.
Workflow: Bake AMI with latest SSM Agent using EC2 Image Builder → attach AmazonSSMManagedInstanceCore → use Session Manager via VPC endpoints → log to S3 → SNS from S3 events.
a reader DB instance to the Aurora cluster and switch the app
Adding a reader and using the cluster and reader endpoints enables automatic failover and minimizes disruption for read traffic during maintenance.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Adding a reader and using the cluster and reader endpoints enables automatic failover and minimizes disruption for read traffic during maintenance.
● Scene fit: Adding a reader and using the cluster and reader endpoints enables automatic failover and minimizes disruption for read.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● RDS Proxy helps manage connections but it does not remove downtime from a single writer reboot during maintenance.
● Aurora does not provide a toggle to enable Multi-AZ with a special instance endpoint, so this approach is not applicable.
● Custom endpoints are intended for grouping by attributes rather than simple read/write splitting and add unnecessary complexity for this need.
Workflow: Add a reader DB instance to the Aurora cluster and switch the app to use the cluster writer endpoint for writes and the reader endpoint for reads.
Migrate the MySQL database to Amazon RDS for MySQL configured for Multi-AZ
RDS Multi-AZ provides synchronous standby and automatic failover across Availability Zones for high availability and fault tolerance. Fargate with an ECS service and ALB offers multi-AZ.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: RDS Multi-AZ provides synchronous standby and automatic failover across Availability Zones for high availability and fault tolerance.
● Scene fit: RDS Multi-AZ provides synchronous standby and automatic failover across Availability Zones for high availability and fault tolerance. Fargate.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● While highly available, this increases management overhead and is typically less cost-efficient for container workloads compared to serverless containers.
● DynamoDB is a NoSQL service and is not a drop-in replacement for a relational MySQL database without a full.
● Read replicas are asynchronous and do not provide automatic failover for HA, making them unsuitable as the primary HA.
Workflow: Migrate the MySQL database to Amazon RDS for MySQL configured for Multi-AZ failover → Run the containers on AWS Fargate with an ECS service integrated with an Application Load Balancer.
Replace the origin TLS certificate with a public CA-signed cert and install
CloudFront requires the origin certificate to be valid, unexpired, correctly named, and chained to a trusted public CA; renewing and installing the complete intermediate chain resolves.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFront requires the origin certificate to be valid, unexpired, correctly named, and chained to a trusted public CA.
● Scene fit: CloudFront requires the origin certificate to be valid, unexpired, correctly named, and chained to a trusted public CA.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adjusting the minimum TLS version does not fix an expired or incomplete origin certificate chain and won’t resolve a 502 caused by certificate validation failures.
● An expired viewer-facing certificate causes client TLS errors, not a 502 with X-Cache: Error from CloudFront which indicates an origin fetch/validation issue.
● ACM public certificates can’t be exported for installation on non-AWS servers, so this is not a viable remediation for an external origin.
Workflow: Replace the origin TLS certificate with a public CA-signed cert and install the full chain.
ALB access logs
Access logs include request_processing_time, target_processing_time, and response_processing_time for each request.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Access logs include request_processing_time, target_processing_time, and response_processing_time for each request.
● Scene fit: Access logs include request_processing_time, target_processing_time, and response_processing_time for each request.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Aggregated metrics only, not per-request timings.
● Typically requires code instrumentation and does not provide ALB per-request timing fields by itself.
● Collects host-level metrics and logs, not ALB per-request latency details.
Workflow: Turn on ALB access logs.
a CloudWatch alarm on the ALB TargetResponseTime metric and attach it to
CodeDeploy can stop a deployment when an associated CloudWatch alarm transitions to ALARM on ALB latency, providing the fastest built-in detection.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy can stop a deployment when an associated CloudWatch alarm transitions to ALARM on ALB latency, providing the fastest built-in detection.
● Scene fit: CodeDeploy can stop a deployment when an associated CloudWatch alarm transitions to ALARM on ALB latency, providing the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● ALBs already emit 1-minute metrics and do not support Detailed Monitoring, so this provides no faster signal nor special integration for auto-stop.
● Access logs are delivered with latency and require parsing, which delays detection compared to native metrics.
● X-Ray requires instrumentation and analysis and does not directly provide the quickest automatic stop for a deployment.
Workflow: Create a CloudWatch alarm on the ALB TargetResponseTime metric and attach it to the CodeDeploy deployment group to automatically halt the rollout on breach.
CodeDeploy AfterAllowTestTraffic hook invoking Lambda; fail the hook to auto-rollback
AfterAllowTestTraffic runs after the test listener routes test traffic to the green task set but before production traffic shifts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AfterAllowTestTraffic runs after the test listener routes test traffic to the green task set but before production traffic shifts.
● Scene fit: AfterAllowTestTraffic runs after the test listener routes test traffic to the green task set but before production traffic shifts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CodeBuild can run tests, but it is outside the CodeDeploy lifecycle.
● AfterAllowTraffic runs after production traffic has moved.
● CloudWatch alarms can roll back a deployment, but they do not run the required smoke tests against the test listener.
Workflow: CodeDeploy AfterAllowTestTraffic hook invoking Lambda → fail the hook to auto-rollback.
the unified CloudWatch agent to send logs and metrics to CloudWatch Logs
CloudWatch agent natively collects logs and additional metrics on Linux and Windows and enables direct analysis with Logs Insights with minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudWatch agent natively collects logs and additional metrics on Linux and Windows and enables direct analysis with Logs Insights with minimal setup.
● Scene fit: CloudWatch agent natively collects logs and additional metrics on Linux and Windows and enables direct analysis with Logs Insights with minimal setup.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Functional but adds extra services and configuration, not the minimal-effort path for EC2 logs and metrics.
● Possible but requires managing additional components and domains, increasing operational overhead compared to CloudWatch Logs Insights.
● SSM Agent is not the native log collection agent; it is for Systems Manager operations, not log shipping.
Workflow: Use the unified CloudWatch agent to send logs and metrics to CloudWatch Logs, → query with CloudWatch Logs Insights.
the CloudWatch agent on EC2 and on-premises hosts to stream to CloudWatch
This funnels all logs into a single low-cost S3 data lake with serverless processing and ad hoc SQL at scale across accounts and on-premises.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This funnels all logs into a single low-cost S3 data lake with serverless processing and ad hoc SQL.
● Scene fit: This funnels all logs into a single low-cost S3 data lake with serverless processing and ad hoc SQL.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Fragments data by account and complicates cross-account analytics and governance rather than centralizing it.
● Kinesis Data Analytics is for streaming inputs (for example Kinesis Data Streams or Firehose) and cannot directly query data.
● Retaining logs on-premises prevents a single AWS-centric aggregation point and is usually more complex and costly for long-term, multi-account.
Workflow: Use the CloudWatch agent on EC2 and on-premises hosts to stream to CloudWatch Logs → subscribe the log groups to Kinesis Data Firehose that delivers to a centralized S3 bucket in a.
Allowlist dev IPs with an AWS WAF IP set + AWS WAF
Create an IP set and an allow rule placed before block rules so developer IPs bypass geo blocking. Use a geo match statement to block requests from selected countries at the WAF layer.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Create an IP set and an allow rule placed before block rules so developer IPs bypass geo blocking.
● Scene fit: Create an IP set and an allow rule placed before block rules so developer IPs bypass geo blocking. Use a geo match statement to block requests from selected countries at the WAF layer.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Enhances DDoS protection but does not implement country blocking or allowlists.
● ALB listener rules lack geo conditions; use AWS WAF for geo-based filtering.
● Influences DNS responses, not a reliable security control for blocking requests by country.
Workflow: Allowlist dev IPs with an AWS WAF IP set → AWS WAF geo match rule to block those countries.
Restart the Amazon ECS agent on the EC2 container instances
Restarting the ECS agent can resolve cases where the agent fails to pull the updated tag and reuses a cached image, causing sporadic launches with old images.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Restarting the ECS agent can resolve cases where the agent fails to pull the updated tag and reuses a cached image, causing sporadic launches with old images.
● Scene fit: Restarting the ECS agent can resolve cases where the agent fails to pull the updated tag and reuses.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Moving registries does not address intermittent reuse of an old local image on the ECS hosts if the agent is unhealthy.
● While digests ensure immutability, the issue described is intermittent and points to an agent state problem rather than tag mutability.
● Using a mutable latest tag does not guarantee a pull if the agent is unhealthy or reuses a locally cached image.
Workflow: Restart the Amazon ECS agent on the EC2 container instances.
CloudTrail to CloudWatch Logs; EC2 logs to CloudWatch Logs via CloudWatch Agent
Both datasets reside in CloudWatch Logs, enabling cross-log-group queries with Logs Insights.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Both datasets reside in CloudWatch Logs, enabling cross-log-group queries with Logs Insights.
● Scene fit: Both datasets reside in CloudWatch Logs, enabling cross-log-group queries with Logs Insights.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● CloudWatch Agent does not send logs directly to S3; it publishes to CloudWatch Logs.
● CloudTrail Lake stores CloudTrail-style events and is not for arbitrary application log ingestion.
● Athena cannot directly query CloudWatch Logs; you would need to export logs to S3 first.
Workflow: CloudTrail to CloudWatch Logs → EC2 logs to CloudWatch Logs via CloudWatch Agent → query with CloudWatch Logs Insights.
Appoint a Firewall Manager administrator account in AWS Organizations and create AWS
AWS Firewall Manager centrally enforces organization-wide WAF policies for matching resources across accounts and Regions, using AWS Config for resource discovery.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Firewall Manager centrally enforces organization-wide WAF policies for matching resources across accounts and Regions, using AWS Config.
● Scene fit: AWS Firewall Manager centrally enforces organization-wide WAF policies for matching resources across accounts and Regions, using AWS Config.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS Config can detect and trigger remediation actions, but it is not the dedicated service for centralized, policy-based WAF.
● CloudFormation StackSets can roll out templates but does not continuously discover and auto-attach WAF to newly created resources without.
● GuardDuty is a threat detection service and does not provide proactive, policy-based attachment of WAF web ACLs across accounts.
Workflow: Appoint a Firewall Manager administrator account in AWS Organizations and create AWS Firewall Manager policies that automatically associate a WAF web ACL to all internet-facing ALBs and API Gateway APIs.
Define and publish the serverless application version with AWS CloudFormation and deploy
This preset performs a canary that shifts 20% of traffic for 10 minutes before routing the remainder, matching the small-subset-and-bake requirement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This preset performs a canary that shifts 20% of traffic for 10 minutes before routing the remainder, matching the small-subset-and-bake requirement.
● Scene fit: This preset performs a canary that shifts 20% of traffic for 10 minutes before routing the remainder, matching the small-subset-and-bake requirement.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● All-at-once shifts 100% of traffic immediately, which does not provide a limited canary exposure or bake time.
● A manual approval gate does not route a small subset of traffic for verification; once approved it still pushes to all users.
● AppConfig manages configuration and feature flags but does not perform Lambda alias traffic shifting in a CodeDeploy-driven deployment.
Workflow: Define and publish the serverless application version with AWS CloudFormation and deploy the Lambda updates with AWS CodeDeploy using CodeDeployDefault.LambdaCanary20Percent10Minutes.
an Amazon EventBridge rule for EC2 Instance State-change Notification events and target
EventBridge emits EC2 state-change events that can be sent to SNS for immediate notifications across all instance state transitions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge emits EC2 state-change events that can be sent to SNS for immediate notifications across all instance state transitions.
● Scene fit: EventBridge emits EC2 state-change events that can be sent to SNS for immediate notifications across all instance state transitions.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Action targets hardware-related failures by attempting instance recovery and does not notify for every state transition.
● AWS Health focuses on account and service-impacting events rather than routine EC2 start, stop, or terminate transitions.
● Addresses instance-level health issues by rebooting and does not provide alerts for all possible state changes.
Workflow: Create an Amazon EventBridge rule for EC2 Instance State-change Notification events and target an Amazon SNS topic.
Build an Amazon Aurora MySQL Global Database for the catalog with cross-Region
Staying on Aurora minimizes code changes while providing a single global catalog via Aurora Global Database and ensuring regulated data stays per Region in local Aurora.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Staying on Aurora minimizes code changes while providing a single global catalog via Aurora Global Database and ensuring.
● Scene fit: Staying on Aurora minimizes code changes while providing a single global catalog via Aurora Global Database and ensuring.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Switching both catalog and regulated data to DynamoDB would require significant application refactoring from relational SQL to NoSQL, so.
● Redshift is a data warehouse optimized for analytics rather than OLTP workloads, and pairing it with DynamoDB adds heavy.
● Placing PII on a global DynamoDB table violates data residency requirements and mixes engines unnecessarily, increasing refactor effort.
Workflow: Build an Amazon Aurora MySQL Global Database for the catalog with cross-Region readers and run Region-local Aurora MySQL clusters for PII and orders.
a CloudWatch alarm on Lambda metrics and associate it with the CodeDeploy
CodeDeploy integrates with CloudWatch alarms to automatically stop and roll back a deployment when error thresholds are breached. This canary strategy sends 10% of traffic to the new version for 5 minutes before shifting the rest, matching the requirement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy integrates with CloudWatch alarms to automatically stop and roll back a deployment when error thresholds are breached.
● Scene fit: CodeDeploy integrates with CloudWatch alarms to automatically stop and roll back a deployment when error thresholds are breached.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Shifts 100% of traffic immediately and does not provide a canary window.
● EventBridge rules can observe events but are not used by CodeDeploy to trigger automated rollbacks.
● A linear deployment moves traffic in repeated increments rather than a two-step canary shift.
Workflow: Create a CloudWatch alarm on Lambda metrics and associate it with the CodeDeploy deployment → Choose a deployment configuration of LambdaCanary10Percent5Minutes.
an AWS Config custom rule with AWS CloudFormation StackSets to check instance
An organization-wide AWS Config custom rule with an aggregator provides centralized compliance visibility for AMI usage. Centralizing and sharing the AMI lets you unshare the old.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An organization-wide AWS Config custom rule with an aggregator provides centralized compliance visibility for AMI usage.
● Scene fit: An organization-wide AWS Config custom rule with an aggregator provides centralized compliance visibility for AMI usage. Centralizing and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Building AMIs in each account increases drift and makes retiring older AMIs difficult across the organization.
● Copying AMIs proliferates stale versions and does not prevent launches from old copies that already exist in other accounts.
● Service Catalog helps with provisioning patterns but does not stop users from launching older AMIs directly or provide fleet-wide.
Workflow: Deploy an AWS Config custom rule with AWS CloudFormation StackSets to check instance AMI IDs against an approved list and aggregate results in an AWS Config aggregator in the management account →.
In mixed instances, scale-in chooses purchase option before AZ, causing transient AZ
Scale-in first selects Spot or On-Demand to remove, then applies AZ balancing within that option, which can momentarily unbalance AZs. Auto Scaling launches replacements before terminating during rebalancing, which can briefly go over MaxSize.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Scale-in first selects Spot or On-Demand to remove, then applies AZ balancing within that option, which can momentarily unbalance AZs.
● Scene fit: Scale-in first selects Spot or On-Demand to remove, then applies AZ balancing within that option, which can momentarily.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● False; rebalancing features can launch before terminating and temporarily exceed MaxSize.
● AZ rebalancing is evaluated before billing hour, so billing alignment does not explain choosing the smaller AZ.
● MaxSize is enforced for scaling activities; cooldown timing does not authorize exceeding MaxSize.
Workflow: In mixed instances → scale-in chooses purchase option before AZ, causing transient AZ skew → Rebalancing can temporarily exceed MaxSize by up to 10% or one instance to maintain availability.
Mirror dependencies to S3 and fetch via an S3 gateway VPC endpoint
S3 with a gateway VPC endpoint keeps traffic on the AWS network and uses IAM for least-privilege access without public Internet.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 with a gateway VPC endpoint keeps traffic on the AWS network and uses IAM for least-privilege access without public Internet.
● Scene fit: S3 with a gateway VPC endpoint keeps traffic on the AWS network and uses IAM for least-privilege access without public Internet.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is IPv6-only and still permits outbound Internet traffic, violating the no-Internet requirement.
● A NAT gateway enables Internet egress, which breaks the isolation requirement.
● There is no VPC endpoint for public OS repositories; those require Internet access unless mirrored privately.
Workflow: Mirror dependencies to S3 and fetch via an S3 gateway VPC endpoint using an instance role.
Systems Manager Patch Manager plus AWS Config approved-amis-by-id with alerts
Patch Manager automates OS updates and the approved-amis-by-id rule continuously evaluates AMI compliance, enabling alerts without blocking launches.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Patch Manager automates OS updates and the approved-amis-by-id rule continuously evaluates AMI compliance, enabling alerts without blocking launches.
● Scene fit: Patch Manager automates OS updates and the approved-amis-by-id rule continuously evaluates AMI compliance, enabling alerts without blocking launches.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Security Hub aggregates findings but does not automate OS patching or evaluate instances against a custom approved AMI list.
● GuardDuty detects threats, not patch posture or AMI allowlists.
● A deny policy prevents launching from unapproved AMIs, which conflicts with the requirement to allow launches and only alert.
Workflow: Systems Manager Patch Manager → AWS Config approved-amis-by-id with alerts.
AWS Personal Health Dashboard with Amazon EventBridge to match AWS Health events
AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda for Slack notifications with minimal setup.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda for Slack notifications with minimal setup.
● Scene fit: AWS Health publishes account-specific scheduled change events that EventBridge can route to Lambda for Slack notifications with minimal setup.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Status check alarms detect instance health issues but do not capture account-specific AWS-scheduled maintenance events across services.
● Config and Trusted Advisor focus on configuration and best practice checks, not delivery of AWS-scheduled maintenance notifications.
● CloudTrail logs API activity, not AWS-initiated maintenance, and Support API polling is not a push-based or comprehensive solution.
Workflow: Use AWS Personal Health Dashboard with Amazon EventBridge to match AWS Health events and trigger a Lambda function that relays them to Slack.
an Amazon Kinesis Data Firehose delivery stream with a Lambda transformation, set
Network Firewall can send logs directly to Kinesis Data Firehose, which supports inline Lambda transformations and reliable near real-time delivery to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Network Firewall can send logs directly to Kinesis Data Firehose, which supports inline Lambda transformations and reliable near.
● Scene fit: Network Firewall can send logs directly to Kinesis Data Firehose, which supports inline Lambda transformations and reliable near.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Performs batch processing after logs are already stored in S3, which fails the requirement to transform data before it lands in the bucket.
● Still processes logs after they have been written to S3 and introduces recursion and operational complexity.
● Network Firewall does not natively publish to Kinesis Data Streams, so this would require unsupported routing or extra components.
Workflow: Create an Amazon Kinesis Data Firehose delivery stream with a Lambda transformation → set the destination to the current S3 bucket, and update Network Firewall to publish logs to the stream.
an .ebextensions/db-migration.config with a container_commands block for the migration and set leader_only
This is correct because container_commands support leader_only and run once on the leader instance before the application is deployed.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This is correct because container_commands support leader_only and run once on the leader instance before the application is deployed.
● Scene fit: This is correct because container_commands support leader_only and run once on the leader instance before the application is deployed.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is incorrect because the commands block executes on every instance and does not support leader_only.
● Is incorrect because it does not integrate with Elastic Beanstalk leader election and can still trigger parallel runs.
● Is incorrect because lock_mode is not a valid attribute for Elastic Beanstalk container commands.
Workflow: Add an .ebextensions/db-migration.config with a container_commands block for the migration and set leader_only: true.
GuardDuty delegated administrator with member accounts; EventBridge filters Impact:IAMUser/AnomalousBehavior to invoke a
GuardDuty centrally detects findings across accounts and publishes to EventBridge, where a rule can match Impact:IAMUser/AnomalousBehavior and trigger Lambda to call the mapping API and send.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty centrally detects findings across accounts and publishes to EventBridge, where a rule can match Impact:IAMUser/AnomalousBehavior and trigger.
● Scene fit: GuardDuty centrally detects findings across accounts and publishes to EventBridge, where a rule can match Impact:IAMUser/AnomalousBehavior and trigger.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail Insights detects unusual API activity, but it does not emit the GuardDuty finding type targeted here and is.
● SNS cannot pattern-match GuardDuty findings for routing; EventBridge must be used to filter and invoke Lambda.
● Detective analyzes data and ingests GuardDuty findings but does not generate the IAM anomalous behavior findings needed to trigger.
Workflow: Use GuardDuty delegated administrator with member accounts → EventBridge filters Impact:IAMUser/AnomalousBehavior to invoke a Lambda for lookup and notifications.
a new S3 object key for each build + Change to a
Changing S3Key ensures CloudFormation detects a change and updates the function code. Altering S3Bucket also forces an update and refreshes the code, though it is heavy-handed. Specifying S3ObjectVersion makes CloudFormation pick the exact artifact version and redeploy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Changing S3Key ensures CloudFormation detects a change and updates the function code.
● Scene fit: Changing S3Key ensures CloudFormation detects a change and updates the function code. Altering S3Bucket also forces an update.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● SAM does not change how CloudFormation detects Lambda code changes for S3Bucket/S3Key.
● Waiting does not change the CloudFormation properties so no code update occurs.
● Without a changed S3Key or S3ObjectVersion, CloudFormation may not update the function code.
Workflow: Use a new S3 object key for each build → Change to a new S3 bucket name each deployment → Enable S3 versioning and reference S3ObjectVersion.
an egress-only internet gateway and route ::/0 to it
Egress-only internet gateway enables outbound-only IPv6 from private subnets via a ::/0 route.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Egress-only internet gateway enables outbound-only IPv6 from private subnets via a ::/0 route.
● Scene fit: Egress-only internet gateway enables outbound-only IPv6 from private subnets via a ::/0 route.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● NAT Gateway is for IPv4 egress and NAT64 (v6-to-v4), not IPv6-to-IPv6 routing.
● NAT instances only handle IPv4, not IPv6.
● Permits inbound IPv6 and effectively makes subnets public, not just egress.
Workflow: Add an egress-only internet gateway and route ::/0 to it.
provisioned concurrency and scale it with Application Auto Scaling
Provisioned concurrency keeps execution environments initialized and can be scaled to match predictable demand.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provisioned concurrency keeps execution environments initialized and can be scaled to match predictable demand.
● Scene fit: Provisioned concurrency keeps execution environments initialized and can be scaled to match predictable demand.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● More memory can reduce init and runtime but does not prevent cold starts under bursty traffic.
● Reserved concurrency limits and isolates capacity but does not keep environments warm to avoid cold starts.
● SnapStart reduces cold starts only for supported Java runtimes and is not a general solution for all functions.
Workflow: Use provisioned concurrency and scale it with Application Auto Scaling.
CloudTrail management events and add an EventBridge rule filtering dynamodb DeleteTable to
EventBridge can match CloudTrail management events for DeleteTable and route directly to SNS with near real-time delivery and minimal cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can match CloudTrail management events for DeleteTable and route directly to SNS with near real-time delivery and minimal cost.
● Scene fit: EventBridge can match CloudTrail management events for DeleteTable and route directly to SNS with near real-time delivery and minimal cost.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Config evaluates resource state with potential delay and added costs, not ideal for immediate API-call alerts.
● DynamoDB does not emit native service events for management API calls like DeleteTable; these are surfaced via CloudTrail.
● Ingesting logs adds ongoing costs and complexity versus direct EventBridge routing from CloudTrail.
Workflow: Enable CloudTrail management events and add an EventBridge rule filtering dynamodb DeleteTable to SNS.
Launch EC2 instances with an IAM role (instance profile) and retrieve the
Using an instance profile provides temporary credentials and Secrets Manager securely stores and rotates database passwords.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using an instance profile provides temporary credentials and Secrets Manager securely stores and rotates database passwords.
● Scene fit: Using an instance profile provides temporary credentials and Secrets Manager securely stores and rotates database passwords.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Long-lived access keys on instances increase risk of credential leakage even if Parameter Store is used.
● Embedding secrets in AMIs leads to distribution and rotation challenges and is not a best practice.
● User data is not intended for storing secrets and base64 is not encryption, making this insecure.
Workflow: Launch EC2 instances with an IAM role (instance profile) and retrieve the database credentials from AWS Secrets Manager.
an Amazon CloudWatch alarm on the StatusCheckFailed_System metric to trigger the EC2
A CloudWatch alarm on the system status check can invoke the EC2 recover action, which migrates the instance to new hardware to remediate host power or.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A CloudWatch alarm on the system status check can invoke the EC2 recover action, which migrates the instance.
● Scene fit: A CloudWatch alarm on the system status check can invoke the EC2 recover action, which migrates the instance.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● An ASG with size 1 can replace an unhealthy instance, but it launches a new instance rather than recovering the same one and may not preserve configuration or data attachments automatically.
● The instance status check and reboot action address OS-level issues but do not fix host-level power or networking problems.
● Frequent snapshots protect data but do not automatically recover a running instance or minimize downtime during host failures.
Workflow: Configure an Amazon CloudWatch alarm on the StatusCheckFailed_System metric to trigger the EC2 recover action.
interface VPC endpoints for Systems Manager + Attach IAM instance profile with
Interface endpoints for ssm, ec2messages, and ssmmessages keep Session Manager control and data channels private. Instances need this role so SSM Agent can register and establish sessions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Interface endpoints for ssm, ec2messages, and ssmmessages keep Session Manager control and data channels private.
● Scene fit: Interface endpoints for ssm, ec2messages, and ssmmessages keep Session Manager control and data channels private. Instances need this role so SSM Agent can register and establish sessions.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A bastion relies on SSH and keys and does not enforce private Session Manager access.
● Uses public SSM endpoints over the internet, violating the private-only requirement.
● The EC2 API endpoint does not provide the private data channels needed by Session Manager.
Workflow: Create interface VPC endpoints for Systems Manager → Attach IAM instance profile with AmazonSSMManagedInstanceCore.
a post-deploy stage in CodePipeline that invokes an AWS Lambda function to
Tying automation to the deployment event ensures the SDK is published and the cache is actively refreshed right after each release.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Tying automation to the deployment event ensures the SDK is published and the cache is actively refreshed right after each release.
● Scene fit: Tying automation to the deployment event ensures the SDK is published and the cache is actively refreshed right.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Polling on a timer is inefficient and may issue unnecessary exports and invalidations when no deployment occurred.
● S3 does not provide an API to invalidate CloudFront caches, so the distribution would not be refreshed.
● Short TTLs can still serve stale content and increase origin load, failing the requirement for immediate freshness after each deployment.
Workflow: Add a post-deploy stage in CodePipeline that invokes an AWS Lambda function to export the SDK from API Gateway, upload it to the S3 origin, and create a CloudFront invalidation for the SDK path.
an S3 replication rule on the source bucket for all objects +
Replication is configured on the source bucket and should target the desired prefix or all objects. The destination bucket must trust the source account's replication role.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Replication is configured on the source bucket and should target the desired prefix or all objects.
● Scene fit: Replication is configured on the source bucket and should target the desired prefix or all objects. The destination.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 assumes a role in the source account; a destination account role is not required.
● Replication is not configured on the destination; it is defined on the source bucket only.
● The role's attached policy typically grants source access; a source bucket policy change is not required for standard replication.
Workflow: Create an S3 replication rule on the source bucket for all objects → Configure the destination bucket policy to allow the source replication role to write and set ownership → Create a.
Lambda@Edge
Executes code at CloudFront edge locations on viewer/origin events for low-latency User-Agent based routing.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Executes code at CloudFront edge locations on viewer/origin events for low-latency User-Agent based routing.
● Scene fit: Executes code at CloudFront edge locations on viewer/origin events for low-latency User-Agent based routing.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Improves network path to regional endpoints but does not execute custom code or perform User-Agent based logic at the edge.
● Good for lightweight header/URL mods but not suited for heavier Lambda logic or dynamic image selection.
● Routes via CloudFront to a regional API; your code still runs in-region, not at the edge.
Workflow: Lambda@Edge.
DeletionPolicy: Snapshot on AWS::EC2::Volume resources to capture EBS snapshots on stack deletion
Setting DeletionPolicy to Snapshot on EBS volumes instructs CloudFormation to take a snapshot of the volume when the stack or resource is deleted. Using DeletionPolicy Retain.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Setting DeletionPolicy to Snapshot on EBS volumes instructs CloudFormation to take a snapshot of the volume when the.
● Scene fit: Setting DeletionPolicy to Snapshot on EBS volumes instructs CloudFormation to take a snapshot of the volume when the.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Termination or deletion protection does not create EBS snapshots and does not reliably prevent replacement-driven data loss in updates.
● CloudFormation does not support making individual resource properties read-only, so this would not prevent changes or data loss.
● A stack policy would block legitimate updates and still does not address the need to snapshot EBS volumes on.
Workflow: Configure DeletionPolicy: Snapshot on AWS::EC2::Volume resources to capture EBS snapshots on stack deletion → Set DeletionPolicy: Retain on the AWS::RDS::DBInstance so replacements and deletions do not remove the database.
one stage with test actions sharing the same runOrder
Actions in the same stage with the same runOrder execute in parallel, reducing total elapsed time.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Actions in the same stage with the same runOrder execute in parallel, reducing total elapsed time.
● Scene fit: Actions in the same stage with the same runOrder execute in parallel, reducing total elapsed time.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Service concurrency limits do not change CodePipeline sequencing when actions are configured to run one after another.
● Scaling CPU or memory does not help when tests are network bound.
● Stages execute sequentially, and runOrder only controls action order within a single stage.
Workflow: Use one stage with test actions sharing the same runOrder.
AWS CloudFormation to manage Lambda function versions and configure API Gateway stage
CloudFormation supports Lambda versions and API Gateway stage canary settings, enabling controlled percentage traffic shifting and easy promotion of the new release.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudFormation supports Lambda versions and API Gateway stage canary settings, enabling controlled percentage traffic shifting and easy promotion.
● Scene fit: CloudFormation supports Lambda versions and API Gateway stage canary settings, enabling controlled percentage traffic shifting and easy promotion.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Combines SAM with DNS-based traffic shifting, which is not suitable for fine-grained canaries on a single API Gateway endpoint.
● Failover routing moves 100% of traffic only after a health check fails and does not support progressive, percentage-based shifts.
● CodeDeploy can manage Lambda alias traffic shifting but does not coordinate API Gateway stage or resource updates as part.
Workflow: Use AWS CloudFormation to manage Lambda function versions and configure API Gateway stage canary settings → update the stack to roll out the new code and promote after validation.
Amazon EFS paired with EC2 Spot Instances
EFS offers a managed, shared POSIX file system for many instances and Spot minimizes cost for interruption-tolerant workloads.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EFS offers a managed, shared POSIX file system for many instances and Spot minimizes cost for interruption-tolerant workloads.
● Scene fit: EFS offers a managed, shared POSIX file system for many instances and Spot minimizes cost for interruption-tolerant workloads.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Provides a shared file system but On-Demand increases cost and Lustre is optimized for HPC rather than general checkpointed batch jobs.
● Orchestrates jobs but does not supply a shared POSIX file system and uses higher-cost compute.
● Multi-Attach is not a clustered shared file system and risks data corruption when used concurrently across instances.
Workflow: Amazon EFS paired with EC2 Spot Instances.
an Amazon EventBridge rule for CodeDeploy deployment state changes that invokes a
EventBridge directly captures CodeDeploy state-change events and, together with Lambda for Slack notifications and native rollback on failure, cleanly satisfies the requirements.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge directly captures CodeDeploy state-change events and, together with Lambda for Slack notifications and native rollback on failure.
● Scene fit: EventBridge directly captures CodeDeploy state-change events and, together with Lambda for Slack notifications and native rollback on failure.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach relies on audit logs and manual API calls rather than reacting to real-time deployment state events and native automatic rollback.
● While alarms can notify, this uses metrics instead of direct state-change events and adds indirection that is less precise than event-based detection.
● Adds unnecessary components and custom rollback logic instead of using CodeDeploy’s built-in automatic rollback based on deployment failure events.
Workflow: Set an Amazon EventBridge rule for CodeDeploy deployment state changes that invokes a Lambda function to notify Slack, and enable CodeDeploy Roll back when a deployment fails.
AWS Config to capture configuration changes, deliver snapshots and history to Amazon
AWS Config continuously records and evaluates resource configurations against rules and provides compliance status, and exporting to S3 enables QuickSight dashboards for organization-wide visibility.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config continuously records and evaluates resource configurations against rules and provides compliance status, and exporting to S3.
● Scene fit: AWS Config continuously records and evaluates resource configurations against rules and provides compliance status, and exporting to S3.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Inspector focuses on EC2 and ECR vulnerability and network exposure findings, not comprehensive resource configuration compliance or a.
● Trusted Advisor provides account-level best practice checks and cost/security recommendations, but it does not evaluate resource configuration rules or.
● Systems Manager Compliance targets patch and SSM configuration baselines for managed instances and does not provide broad AWS resource.
Workflow: Enable AWS Config to capture configuration changes → deliver snapshots and history to Amazon S3, and build near real-time compliance visuals in Amazon QuickSight.
a Failover routing policy and create alias records that point to the
Failover routing with alias targets and Evaluate Target Health enables Route 53 to return the secondary target when the primary is unhealthy. Health checks are required.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Failover routing with alias targets and Evaluate Target Health enables Route 53 to return the secondary target when.
● Scene fit: Failover routing with alias targets and Evaluate Target Health enables Route 53 to return the secondary target when.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds custom automation but is unnecessary for Route 53 DNS failover, which is handled natively by routing policies and.
● Weighted routing distributes load and is not designed to automatically fail over when the primary becomes unhealthy.
Workflow: Use a Failover routing policy and create alias records that point to the application resources, enabling Evaluate Target Health → Configure Route 53 health checks for non-alias endpoints on both the active.
AWS Config organization rule for EBS default encryption plus an SCP to
Workable with organization-level AWS Config and a central aggregator.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Workable with organization-level AWS Config and a central aggregator.
● Scene fit: Workable with organization-level AWS Config and a central aggregator.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Technically blocks some new unencrypted launches, but it does not evaluate existing resources or provide one central compliance view.
● Technically workable, but StackSets add per-account deployment and maintenance.
● Security Hub can centralize findings, but it does not enforce EBS default encryption and depends on underlying controls such as AWS Config.
Workflow: Use AWS Config organization rule for EBS default encryption → an SCP to block disabling Config.
Bucket still has objects/versions; make the custom resource empty the bucket on
S3 requires a bucket to be empty before deletion.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 requires a bucket to be empty before deletion.
● Scene fit: S3 requires a bucket to be empty before deletion.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Object Lock can technically block deletion, but it is a special case and retained objects may not be removable until retention expires.
● Permissions can cause failure, but granting delete permissions alone does not empty the bucket.
● Static website configuration does not prevent deletion of an empty bucket.
Workflow: Bucket still has objects/versions → make the custom resource empty the bucket on Delete.
an Auto Scaling lifecycle hook to place instances in Terminating:Wait and create
A lifecycle hook in Terminating:Wait pauses termination and EventBridge can react to that event so you can run cleanup before completing the lifecycle action. Systems Manager.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A lifecycle hook in Terminating:Wait pauses termination and EventBridge can react to that event so you can run.
● Scene fit: A lifecycle hook in Terminating:Wait pauses termination and EventBridge can react to that event so you can run.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Maintenance Windows are for scheduled operations and are not event-driven or able to pause Auto Scaling termination workflows.
● Terminating:Pending is not a valid Auto Scaling lifecycle state, so this hook cannot be used.
● EC2 Image Builder is not designed to run per-instance, event-driven cleanup before termination in an Auto Scaling scale-in event.
Workflow: Configure an Auto Scaling lifecycle hook to place instances in Terminating:Wait and create an Amazon EventBridge rule to monitor the Terminating:Wait event → Create an AWS Systems Manager Automation runbook and configure.
AWS Config managed and custom rules with both change-triggered and periodic evaluations
AWS Config continuously records resource states and evaluates them against rules, providing the authoritative service for configuration compliance and change auditing. CloudTrail provides the audit trail.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config continuously records resource states and evaluates them against rules, providing the authoritative service for configuration compliance.
● Scene fit: AWS Config continuously records resource states and evaluates them against rules, providing the authoritative service for configuration compliance.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Emphasizes automated remediation via EventBridge and Lambda, which is not required when the goal is continuous assessment and auditing.
● Collecting SDK logs is console-agnostic and misses full governance coverage, making it unreliable for organization-wide configuration assessment and compliance.
Workflow: Use AWS Config managed and custom rules with both change-triggered and periodic evaluations, and centralize compliance results with an aggregator across accounts → Enable an organization trail in AWS CloudTrail for all.
Store credentials in AWS Secrets Manager and resolve them in the template
Dynamic references fetch the secret at deploy time so the plaintext value is not embedded in the template or stack metadata. NoEcho masks parameter values with.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Dynamic references fetch the secret at deploy time so the plaintext value is not embedded in the template or stack metadata.
● Scene fit: Dynamic references fetch the secret at deploy time so the plaintext value is not embedded in the template.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Encrypting the template at rest does not stop secret values from appearing in stack events or Describe* API outputs.
● Tags cannot be used to resolve secure strings in CloudFormation and would not provide secure parameter retrieval.
● CloudTrail improves auditing but does not prevent secret values from being exposed during stack operations.
Workflow: Store credentials in AWS Secrets Manager and resolve them in the template using CloudFormation dynamic references → Configure NoEcho on sensitive CloudFormation parameters to hide their values in stack outputs and events.
EventBridge scheduled rule invoking Lambda that tags first-seen and deletes after 30
A scheduled EventBridge rule runs Lambda to track first-seen-detached via a tag and delete volumes whose detached age exceeds 30 days.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A scheduled EventBridge rule runs Lambda to track first-seen-detached via a tag and delete volumes whose detached age exceeds 30 days.
● Scene fit: A scheduled EventBridge rule runs Lambda to track first-seen-detached via a tag and delete volumes whose detached age exceeds 30 days.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Trusted Advisor can list unattached volumes but does not track 30-day detachment age or trigger automation directly.
● Config flags unattached volumes but does not manage per-volume 30-day timers; building a reliable delay/remediation flow is brittle.
● There is no native CloudWatch metric for EBS detachment age, so you cannot alarm on 30 days detached.
Workflow: EventBridge scheduled rule invoking Lambda that tags first-seen and deletes after 30 days.
a bucket policy that denies any request when aws:SecureTransport equals false, use
Denying when aws:SecureTransport is false enforces HTTPS, SSE-S3 scales without KMS quotas, and CRR meets the cross-continent DR requirement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Denying when aws:SecureTransport is false enforces HTTPS, SSE-S3 scales without KMS quotas, and CRR meets the cross-continent DR requirement.
● Scene fit: Denying when aws:SecureTransport is false enforces HTTPS, SSE-S3 scales without KMS quotas, and CRR meets the cross-continent DR.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Would block HTTPS traffic by denying when aws:SecureTransport is true, which violates the encryption-in-transit requirement.
● While HTTPS is enforced, SSE-KMS can introduce KMS request throttling at very high put rates like 15000 objects per second.
● CloudFront does not replicate S3 data across Regions and SSE-KMS can still bottleneck at very high ingest, so this does not satisfy DR or performance needs.
Workflow: Add a bucket policy that denies any request when aws:SecureTransport equals false → use SSE-S3 for default encryption, and configure S3 Cross-Region Replication.
CloudWatch Logs subscription filters to Kinesis Data Firehose (cross-account) delivering to S3
Use subscription filters to stream to Firehose for near real-time S3 delivery; an EventBridge rule on CreateLogGroup invokes Lambda to call PutSubscriptionFilter for new groups.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use subscription filters to stream to Firehose for near real-time S3 delivery; an EventBridge rule on CreateLogGroup invokes Lambda to call PutSubscriptionFilter for new groups.
● Scene fit: Use subscription filters to stream to Firehose for near real-time S3 delivery; an EventBridge rule on CreateLogGroup invokes.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail records API activity to S3 and is unrelated to streaming CloudWatch Logs log group data.
● Exports are batch, time-bounded per log group and do not automatically cover new groups or provide continuous delivery.
● While possible, this adds scaling and operational overhead compared to Firehose and is not the minimal-effort, organization-wide pattern.
Workflow: CloudWatch Logs subscription filters to Kinesis Data Firehose (cross-account) delivering to S3, auto-attached via EventBridge + Lambda on CreateLogGroup.
an egress-only internet gateway and add a ::/0 route in the private
An egress-only internet gateway enables outbound-only IPv6 traffic from private subnets when the route table sends ::/0 to it.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: An egress-only internet gateway enables outbound-only IPv6 traffic from private subnets when the route table sends ::/0 to it.
● Scene fit: An egress-only internet gateway enables outbound-only IPv6 traffic from private subnets when the route table sends ::/0 to it.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A NAT instance and a 0.0.0.0/0 route only address IPv4 egress and do not provide IPv6 connectivity.
● A 0.0.0.0/0 route is IPv4-only and does not enable IPv6 egress from private subnets.
● NAT Gateway does not provide IPv6-to-IPv6 egress and cannot be targeted by an IPv6 ::/0 route.
Workflow: Create an egress-only internet gateway and add a ::/0 route in the private subnet route table to it.
CAPABILITY_IAM or CAPABILITY_NAMED_IAM on the CloudFormation deploy action in CodePipeline
This acknowledges that the stack will create or modify IAM resources, which resolves the InsufficientCapabilitiesException in CloudFormation when run from CodePipeline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This acknowledges that the stack will create or modify IAM resources, which resolves the InsufficientCapabilitiesException in CloudFormation when run from CodePipeline.
● Scene fit: This acknowledges that the stack will create or modify IAM resources, which resolves the InsufficientCapabilitiesException in CloudFormation when run from CodePipeline.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● S3 service quotas are unrelated to the CloudFormation InsufficientCapabilitiesException error, which is about acknowledging capabilities for certain resource types, not storage limits.
● Granting broader permissions to the pipeline role does not resolve the requirement to explicitly acknowledge IAM-related capabilities in CloudFormation operations.
● A circular dependency would produce a different CloudFormation error, not InsufficientCapabilitiesException, which is specific to missing capability acknowledgments.
Workflow: Enable CAPABILITY_IAM or CAPABILITY_NAMED_IAM on the CloudFormation deploy action in CodePipeline.
Block S3 public access and grant least-privilege to the CodeBuild service role
Use S3 Block Public Access and a bucket/IAM policy that allows only the CodeBuild role to read the required objects.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use S3 Block Public Access and a bucket/IAM policy that allows only the CodeBuild role to read the required objects.
● Scene fit: Use S3 Block Public Access and a bucket/IAM policy that allows only the CodeBuild role to read the required objects.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Presigned URLs are temporary but still bearer tokens and add orchestration complexity compared to role-based access.
● Long-lived static credentials in env vars are high risk and against best practices; prefer short-lived role credentials.
● Overly broad permissions violate least-privilege and increase blast radius.
Workflow: Block S3 public access and grant least-privilege to the CodeBuild service role.
AWS Elastic Disaster Recovery to replicate EC2 data to a cross-Region staging
AWS Elastic Disaster Recovery continuously replicates to a low-cost staging area in another Region and supports rapid failover and failback with minimal operational effort.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Elastic Disaster Recovery continuously replicates to a low-cost staging area in another Region and supports rapid failover and failback with minimal operational effort.
● Scene fit: AWS Elastic Disaster Recovery continuously replicates to a low-cost staging area in another Region and supports rapid failover.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Resilience Hub evaluates and tracks resiliency but does not replicate EC2 data or provide native failover/failback automation.
● Backups provide point-in-time snapshots rather than continuous replication, and Fault Injection Simulator is for chaos experiments, not automated recovery.
● Application Migration Service is intended for migrations and is not the recommended service for streamlined, ongoing DR with fast failback compared to AWS Elastic Disaster Recovery.
Workflow: Use AWS Elastic Disaster Recovery to replicate EC2 data to a cross-Region staging area subnet using its agent.
Associate the GitHub repositories with Amazon CodeGuru Reviewer with Secrets Detector for
CodeGuru Reviewer Secrets Detector surfaces hardcoded secrets during PRs and can be used to enforce prevention, while Secrets Manager securely stores and automatically rotates credentials.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeGuru Reviewer Secrets Detector surfaces hardcoded secrets during PRs and can be used to enforce prevention, while Secrets.
● Scene fit: CodeGuru Reviewer Secrets Detector surfaces hardcoded secrets during PRs and can be used to enforce prevention, while Secrets.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Macie analyzes data in Amazon S3 rather than source code in GitHub, and environment variables do not provide automatic.
● CodeGuru Profiler focuses on performance profiling rather than secret detection, and Parameter Store lacks native automatic rotation and PR.
Workflow: Associate the GitHub repositories with Amazon CodeGuru Reviewer with Secrets Detector for pull request scanning, migrate the database credentials to AWS Secrets Manager with rotation enabled, and update SAM and Python code.
Store the credentials in AWS Secrets Manager, grant the Lambda role GetSecretValue
AWS Secrets Manager is designed for storing secrets with IAM-controlled access and supports automatic rotation for many databases including Amazon RDS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Secrets Manager is designed for storing secrets with IAM-controlled access and supports automatic rotation for many databases including Amazon RDS.
● Scene fit: AWS Secrets Manager is designed for storing secrets with IAM-controlled access and supports automatic rotation for many databases.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Secures retrieval but lacks native credential rotation for databases, so it does not meet the rotation requirement as cleanly as Secrets Manager.
● AWS KMS manages encryption keys rather than application secrets, so it is not appropriate for storing and rotating database passwords.
● AWS CloudHSM provides hardware-backed key management and cryptographic operations, not secret storage or rotation for database credentials.
Workflow: Store the credentials in AWS Secrets Manager → grant the Lambda role GetSecretValue on the secret ARN, and enable automatic rotation.
CodePipeline CloudFormation action with per-stage parameter overrides; use Parameters and Mappings
Per-stage parameter overrides pass environment values at deploy time, and CloudFormation Parameters/Mappings keep one reusable template.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Per-stage parameter overrides pass environment values at deploy time, and CloudFormation Parameters/Mappings keep one reusable template.
● Scene fit: Per-stage parameter overrides pass environment values at deploy time, and CloudFormation Parameters/Mappings keep one reusable template.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Dynamic references alone do not convey environment context from the pipeline; you still need stage-specific parameterization.
● Couples the template to pipeline internals and adds unnecessary complexity.
● Post-launch mutation is brittle and breaks CloudFormation’s declarative model.
Workflow: CodePipeline CloudFormation action with per-stage parameter overrides → use Parameters and Mappings.
Amazon API Gateway direct integration to AWS Step Functions secured by Amazon
Direct API Gateway to Step Functions keeps the design simple, supports long-running Standard workflows, scales, and uses Cognito for auth.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Direct API Gateway to Step Functions keeps the design simple, supports long-running Standard workflows, scales, and uses Cognito for auth.
● Scene fit: Direct API Gateway to Step Functions keeps the design simple, supports long-running Standard workflows, scales, and uses Cognito for auth.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Lambda has a hard 15-minute execution limit, so it cannot handle multi-hour workflows even when fronted by API Gateway.
● Can run long tasks but is not the simplest approach for a public API and introduces container/ALB management and scaling complexity.
● Adding Lambda as a hop adds latency, cost, and potential throttling without benefit over the native direct integration.
Workflow: Amazon API Gateway direct integration to AWS Step Functions secured by Amazon Cognito.
Request or import an ACM certificate for the ALB, associate an ACM
This configuration provides trusted certificates at both CloudFront and the ALB and enforces HTTPS on viewer and origin connections.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This configuration provides trusted certificates at both CloudFront and the ALB and enforces HTTPS on viewer and origin.
● Scene fit: This configuration provides trusted certificates at both CloudFront and the ALB and enforces HTTPS on viewer and origin.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Self-signed certificates on the ALB are not trusted by CloudFront for origin HTTPS and do not meet end-to-end encryption.
● Match Viewer allows HTTP from viewers and can forward HTTP to the origin, so HTTPS is not enforced end-to-end.
● CloudFront cannot use certificates stored in S3, the default certificate does not support custom CNAMEs, and HTTP to the.
Workflow: Request or import an ACM certificate for the ALB, associate an ACM certificate in us-east-1 with the CloudFront distribution's custom domain → set Viewer Protocol Policy to HTTPS Only, and configure the origin to use HTTPS.
AWS Config with Conformance Packs and Aggregators
Continuously records and evaluates resource configurations, aggregates multi-account compliance, and uses conformance packs for standardized policy reporting.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Continuously records and evaluates resource configurations, aggregates multi-account compliance, and uses conformance packs for standardized policy reporting.
● Scene fit: Continuously records and evaluates resource configurations, aggregates multi-account compliance, and uses conformance packs for standardized policy reporting.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Aggregates security findings and standards but relies on other services (like AWS Config) for configuration evaluations and is not the primary engine for near real-time resource config compliance.
● Focuses on vulnerability and exposure assessments for EC2 and ECR, not broad resource configuration compliance across AWS.
● Targets patch and SSM configuration baselines for managed instances rather than organization-wide AWS resource configuration compliance.
Workflow: AWS Config with Conformance Packs and Aggregators.
S3 with versioning + EventBridge schedule + Lambda to refresh tasks
S3 provides durable versioned storage for JSON with audit and rollback, and EventBridge with Lambda can trigger in-place config reloads for running tasks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 provides durable versioned storage for JSON with audit and rollback, and EventBridge with Lambda can trigger in-place config reloads for running tasks.
● Scene fit: S3 provides durable versioned storage for JSON with audit and rollback, and EventBridge with Lambda can trigger in-place config reloads for running tasks.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Secrets Manager is intended for credentials and rotates secrets; it is costly and unsuitable for large, growing general configuration sets.
● EFS allows shared files without restarts but lacks built-in versioning and audit history for rollbacks.
● Advanced tier increases cost and per-parameter size limits become cumbersome for many files; API throttling and hierarchy management add overhead at scale.
Workflow: S3 with versioning + EventBridge schedule + Lambda to refresh tasks.
AWS Database Migration Service (DMS)
Supports ongoing, highly available CDC from RDS Oracle and PostgreSQL into Amazon Redshift.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Supports ongoing, highly available CDC from RDS Oracle and PostgreSQL into Amazon Redshift.
● Scene fit: Supports ongoing, highly available CDC from RDS Oracle and PostgreSQL into Amazon Redshift.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Primarily batch ETL; not for continuous CDC replication into Amazon Redshift.
● Ingests streaming events but lacks native CDC from RDS to Redshift.
● Zero-ETL currently targets Aurora sources, not RDS Oracle or standard RDS PostgreSQL.
Workflow: AWS Database Migration Service (DMS).
Upload each release with a new object key in the same S3
Changing the S3Key forces CloudFormation to detect a new artifact and update the Lambda code. Providing S3ObjectVersion ensures CloudFormation recognizes the exact artifact version and updates.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Changing the S3Key forces CloudFormation to detect a new artifact and update the Lambda code.
● Scene fit: Changing the S3Key forces CloudFormation to detect a new artifact and update the Lambda code. Providing S3ObjectVersion ensures.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Requires reworking the template and deployment flow and does not address CloudFormation's change detection for Lambda code directly.
● Waiting does not cause CloudFormation to redeploy because it still sees the same bucket and key and S3 is.
● While CodeDeploy can deploy Lambda, switching tools is not a quick fix for CloudFormation not detecting a changed S3.
Workflow: Upload each release with a new object key in the same S3 bucket → Turn on S3 versioning for the artifacts bucket and reference the S3ObjectVersion in CloudFormation → Push the package.
an Amazon Aurora global database spanning two Regions and promote the secondary
Aurora Global Database uses low-latency storage-level replication and controlled promotion, which comfortably meets the stated RPO and RTO. Failover routing with health checks automatically shifts traffic.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Aurora Global Database uses low-latency storage-level replication and controlled promotion, which comfortably meets the stated RPO and RTO.
● Scene fit: Aurora Global Database uses low-latency storage-level replication and controlled promotion, which comfortably meets the stated RPO and RTO.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Approach relies on DMS and adds lag and a complex cutover that can jeopardize the 12-minute RTO.
● Latency-based routing optimizes for performance rather than disaster recovery and does not enforce a clear primary-to-secondary failover posture.
● Aurora multi-master is not supported across multiple Regions, so it cannot provide cross-Region active-active.
Workflow: Configure an Amazon Aurora global database spanning two Regions and promote the secondary to read/write if the primary Region fails → Deploy the web tier in two Regions and use Amazon Route.
Adopt attribute-based access control using tags for principals and resources with tag-matching
ABAC policies evaluate matching tags so newly tagged resources are covered automatically without updating the policy itself.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ABAC policies evaluate matching tags so newly tagged resources are covered automatically without updating the policy itself.
● Scene fit: ABAC policies evaluate matching tags so newly tagged resources are covered automatically without updating the policy itself.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Requires adding a new policy on every resource, which does not remove the ongoing maintenance burden.
● SCPs set maximum permission guardrails and cannot grant access to resources or eliminate policy updates.
● Permissions boundaries only define the maximum allowable permissions and do not grant access to new resources.
Workflow: Adopt attribute-based access control using tags for principals and resources with tag-matching policies.
provisioned concurrency for the function and configure Application Auto Scaling with a
Provisioned concurrency keeps execution environments initialized and ready, removing cold-start delays and allowing you to scale capacity to match predictable peaks.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provisioned concurrency keeps execution environments initialized and ready, removing cold-start delays and allowing you to scale capacity to match predictable peaks.
● Scene fit: Provisioned concurrency keeps execution environments initialized and ready, removing cold-start delays and allowing you to scale capacity to.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Reserved concurrency limits how many concurrent invocations a function can reach and isolates capacity, but it does not pre-initialize execution environments to avoid cold starts.
● More memory can reduce init and runtime duration but cannot reliably eliminate cold starts during bursty traffic patterns.
● Increasing ephemeral storage provides larger temporary disk space but does not affect function initialization latency or cold starts.
Workflow: Enable provisioned concurrency for the function and configure Application Auto Scaling with a minimum of 2 and a maximum of 120 provisioned instances.
AWS Trusted Advisor Low Utilization EC2 check via EventBridge with a Lambda
Trusted Advisor surfaces low-utilization EC2 across accounts (Business/Enterprise Support), EventBridge captures check item changes, and Lambda can filter on tags and take action.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Trusted Advisor surfaces low-utilization EC2 across accounts (Business/Enterprise Support), EventBridge captures check item changes, and Lambda can filter on tags and take action.
● Scene fit: Trusted Advisor surfaces low-utilization EC2 across accounts (Business/Enterprise Support), EventBridge captures check item changes, and Lambda can filter.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Uses CloudWatch utilization metrics and EventBridge to trigger a Lambda remediation, but lacks a native cross-account, cost-optimization signal specifically for low-utilization EC2.
● Monitors spend anomalies, not resource utilization levels, so it does not reliably target persistently low-utilization EC2 instances.
● Compute Optimizer provides recommendations, but it is not an event source for automatic shutdown and is not designed for direct auto-remediation of low utilization.
Workflow: AWS Trusted Advisor Low Utilization EC2 check via EventBridge with a Lambda remediation filtered by tags.
the web tier in two Regions with Amazon Route 53 failover to
Two Regional web tiers plus Route 53 failover provide a working cross-Region traffic path. Aurora Global Database continuously replicates to the second Region and supports fast promotion.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Two Regional web tiers plus Route 53 failover provide a working cross-Region traffic path.
● Scene fit: Two Regional web tiers plus Route 53 failover provide a working cross-Region traffic path. Aurora Global Database continuously replicates to the second Region and supports fast promotion.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DMS can replicate data, but cutover requires more coordination and can have lag.
● Multi-AZ protects against an Availability Zone failure inside one Region.
● Latency routing selects the fastest endpoint, not a deterministic primary-to-secondary disaster recovery path.
Workflow: Run the web tier in two Regions with Amazon Route 53 failover to the healthy ALB → Use Amazon Aurora Global Database across two Regions and promote the secondary during failure.
Cross-Region RDS MySQL read replica with pre-staged app and Route 53 failover
A cross-Region read replica plus a warm app tier and Route 53 failover meets a 15-minute RPO and 120-minute RTO with minimal changes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A cross-Region read replica plus a warm app tier and Route 53 failover meets a 15-minute RPO and 120-minute RTO with minimal changes.
● Scene fit: A cross-Region read replica plus a warm app tier and Route 53 failover meets a 15-minute RPO and 120-minute RTO with minimal changes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Delivers robust cross-Region DR but requires an engine migration and configuration changes.
● Multi-AZ is confined to a single Region and cannot span Regions; latency routing is not DR.
● Feasible but adds complexity and manual steps; RTO may be longer and changes are greater than using native read replicas.
Workflow: Cross-Region RDS MySQL read replica with pre-staged app and Route 53 failover.
Install CloudWatch agent to publish memory metrics and alarm to SNS +
CloudWatch does not collect memory by default; the agent publishes memory metrics and an alarm can notify via SNS. With ELB health checks enabled, the Auto.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudWatch does not collect memory by default; the agent publishes memory metrics and an alarm can notify via SNS.
● Scene fit: CloudWatch does not collect memory by default; the agent publishes memory metrics and an alarm can notify via.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Auto recovery addresses underlying host or hypervisor issues and does not replace instances based on ALB target group health.
● CPU scaling adds or removes tasks but does not replace unhealthy EC2 instances or provide memory alerts.
● Scaling tasks on memory utilization changes task count but does not replace failing EC2 instances or generate OS memory alerts.
Workflow: Install CloudWatch agent to publish memory metrics and alarm to SNS → Set ASG health check type to ELB using the ALB target group.
an Amazon EventBridge rule that matches CodePipeline Pipeline Execution State Change and
EventBridge is the supported event bus for CodePipeline state and approval events, SNS provides fan-out and retries, and Lambda can transform and post to the webhook.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge is the supported event bus for CodePipeline state and approval events, SNS provides fan-out and retries, and.
● Scene fit: EventBridge is the supported event bus for CodePipeline state and approval events, SNS provides fan-out and retries, and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● AWS Config evaluates resource configuration compliance and does not natively detect CodePipeline execution state change events.
● CloudTrail records API calls rather than all service-emitted execution state changes, and relying on it misses non-API state transitions.
● CodePipeline events are not emitted as CloudWatch Logs, and Logs subscription filters cannot target SNS.
Workflow: Create an Amazon EventBridge rule that matches CodePipeline Pipeline Execution State Change and Manual Approval Needed events → route them to an Amazon SNS topic, and have an AWS Lambda subscription that.
Stop using any ACL and allow reads only via a tightly scoped
Best practice is to avoid ACLs and enforce least-privilege access with bucket policies targeting specific principals.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Best practice is to avoid ACLs and enforce least-privilege access with bucket policies targeting specific principals.
● Scene fit: Best practice is to avoid ACLs and enforce least-privilege access with bucket policies targeting specific principals.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Encryption at rest does not change object permissions or grantees.
● Block Public Access covers anonymous/public access, not the AuthenticatedUsers group.
● Endpoint policies affect traffic via that endpoint only and do not override object ACLs or internet access.
Workflow: Stop using any ACL and allow reads only via a tightly scoped S3 bucket policy to selected accounts.
Send CloudWatch Logs to Lambda via a subscription; tag the instance on
A CloudWatch Logs subscription filter can invoke Lambda on matching login events; tagging and a scheduled EventBridge rule enable delayed termination within the required window.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A CloudWatch Logs subscription filter can invoke Lambda on matching login events; tagging and a scheduled EventBridge rule.
● Scene fit: A CloudWatch Logs subscription filter can invoke Lambda on matching login events; tagging and a scheduled EventBridge rule.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail records AWS API activity, not OS-level SSH/RDP logins found in system logs, so it will not detect the required event.
● ConsoleLogin pertains to AWS Management Console sign-ins and does not indicate OS-level logins on EC2 instances.
● CloudWatch Logs subscriptions do not target Step Functions directly and a daily schedule can exceed the 10-hour requirement.
Workflow: Send CloudWatch Logs to Lambda via a subscription → tag the instance on login → use an EventBridge schedule to invoke Lambda to terminate tagged instances within 10 hours.
API Gateway APIs in each target Region and use Amazon Route 53
Latency-based routing directs clients to the nearest Regional API while DynamoDB global tables enable multi-Region, low-latency data access with replication.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency-based routing directs clients to the nearest Regional API while DynamoDB global tables enable multi-Region, low-latency data access.
● Scene fit: Latency-based routing directs clients to the nearest Regional API while DynamoDB global tables enable multi-Region, low-latency data access.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Reduces some network latency to the API edge but keeps compute in one Region and does not provide a.
● Failover routing is active–passive and will not deliver low latency to all users because only the primary Region serves.
● Health checks alone do not route to the lowest-latency endpoint and Region-local tables fragment data, which breaks a single.
Workflow: Create API Gateway APIs in each target Region and use Amazon Route 53 latency-based routing with health checks → integrate each API with a same-Region Lambda function and access a DynamoDB global table.
Build a Step Functions workflow that invokes two Lambda functions to take
Automating snapshot, cross-Region copy, and restore via Step Functions and Lambda reduces manual effort for RTO and uses frequent scheduled snapshots to tighten RPO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Automating snapshot, cross-Region copy, and restore via Step Functions and Lambda reduces manual effort for RTO and uses.
● Scene fit: Automating snapshot, cross-Region copy, and restore via Step Functions and Lambda reduces manual effort for RTO and uses.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● RDS does not support scheduling snapshot creation via instance lifecycle events, and using CPU at 0% as a failover.
● Trusted Advisor does not emit AWS-initiated RDS event notifications and ECS adds unnecessary operational overhead compared to a serverless.
● RDS Multi-AZ failover is confined to a single Region and cannot place a standby in another Region.
Workflow: Build a Step Functions workflow that invokes two Lambda functions to take RDS snapshots → copy them to ap-northeast-1, and automate restore in the backup region → schedule snapshots every 30 minutes.
AWS CloudTrail to deliver logs to Amazon S3 and use CloudTrail log
CloudTrail log file integrity validation uses cryptographic digests and signatures to prove no logs were altered or removed and preserves event sequence.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudTrail log file integrity validation uses cryptographic digests and signatures to prove no logs were altered or removed and preserves event sequence.
● Scene fit: CloudTrail log file integrity validation uses cryptographic digests and signatures to prove no logs were altered or removed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 Object Lock provides immutability but does not cryptographically prove CloudTrail log provenance or event ordering, and AWS Config does not capture all API calls.
● Glacier Vault Lock enforces WORM retention but does not validate that CloudTrail produced the files or that their sequence is intact.
● CloudWatch Logs can alert on patterns but does not provide cryptographic integrity verification of log files or ordering.
Workflow: Configure AWS CloudTrail to deliver logs to Amazon S3 and use CloudTrail log file integrity validation to verify authenticity and ordering.
CODEBUILD_SOURCE_VERSION in buildspec.yml to name artifacts
This uses a built-in variable that resolves to the source branch for branch builds, enabling simple and scalable artifact naming.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses a built-in variable that resolves to the source branch for branch builds, enabling simple and scalable artifact naming.
● Scene fit: This uses a built-in variable that resolves to the source branch for branch builds, enabling simple and scalable artifact naming.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Yields a commit SHA, not the branch name, so it does not meet the branch-based naming requirement.
● Creates significant management overhead and does not scale well as branches change.
● Adds unnecessary components and complexity compared to using environment variables during the build.
Workflow: Use CODEBUILD_SOURCE_VERSION in buildspec.yml to name artifacts.
AWS CodeDeploy with CodeDeployDefault.LambdaCanary10Percent5Minutes for Lambda alias traffic shifting
Uses a predefined canary config to shift a small portion of traffic for a timed bake, then completes if healthy.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Uses a predefined canary config to shift a small portion of traffic for a timed bake, then completes if healthy.
● Scene fit: Uses a predefined canary config to shift a small portion of traffic for a timed bake, then completes if healthy.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A manual gate does not provide automated canary traffic shifting or timed bake before full rollout.
● Shifts 100% of traffic immediately, offering no limited exposure or bake time.
● API Gateway canaries target API stage configs, not Lambda version alias shifting in a deployment pipeline.
Workflow: AWS CodeDeploy with CodeDeployDefault.LambdaCanary10Percent5Minutes for Lambda alias traffic shifting.
the cloudtrail-enabled AWS Config managed rule with a 30-minute periodic evaluation and
This solution detects noncompliance on a schedule independent of CloudTrail event delivery and automatically re-enables logging via StartLogging, reducing downtime.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This solution detects noncompliance on a schedule independent of CloudTrail event delivery and automatically re-enables logging via StartLogging, reducing downtime.
● Scene fit: This solution detects noncompliance on a schedule independent of CloudTrail event delivery and automatically re-enables logging via StartLogging, reducing downtime.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Only alerts operators and still requires manual action to restore logging, increasing potential downtime.
● Focusing on DeleteTrail/CreateTrail is incorrect for restoring logging, and polling adds delay compared to reacting to compliance state.
● The cloudtrail-enabled rule supports only periodic evaluations and does not auto-remediate by default.
Workflow: Deploy the cloudtrail-enabled AWS Config managed rule with a 30-minute periodic evaluation and use an EventBridge rule for AWS Config compliance changes to invoke a Lambda function that calls StartLogging on the affected trail.
a CloudWatch Logs destination in the central account and subscribe a Kinesis
This uses a cross-account CloudWatch Logs subscription to a Firehose stream that delivers to S3, providing a secure, centralized, and near-serverless storage target.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses a cross-account CloudWatch Logs subscription to a Firehose stream that delivers to S3, providing a secure, centralized, and near-serverless storage target.
● Scene fit: This uses a cross-account CloudWatch Logs subscription to a Firehose stream that delivers to S3, providing a secure.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds an extra Kinesis Data Streams layer that requires shard provisioning and scaling, which conflicts with the minimal-provisioning requirement.
● OpenSearch Service requires provisioning and managing capacity, which does not meet the minimal-provisioning storage requirement.
● Amazon Redshift requires provisioning and operating a data warehouse cluster, which is not aligned with low-to-no provisioning for storage.
Workflow: Configure a CloudWatch Logs destination in the central account and subscribe a Kinesis Data Firehose delivery stream that writes directly to an Amazon S3 bucket.
Amazon Managed Service for Prometheus for collection of metrics and Amazon Managed
Amazon Managed Service for Prometheus centrally ingests Prometheus metrics from EKS, ECS, and on-prem clusters, and Amazon Managed Grafana provides managed dashboards, queries, and alerts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Amazon Managed Service for Prometheus centrally ingests Prometheus metrics from EKS, ECS, and on-prem clusters, and Amazon Managed Grafana provides managed dashboards, queries, and alerts.
● Scene fit: Amazon Managed Service for Prometheus centrally ingests Prometheus metrics from EKS, ECS, and on-prem clusters, and Amazon Managed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Stack is not Prometheus-native and does not provide cohesive scraping and aggregation across EKS, ECS, and on-prem Kubernetes for time-series metrics.
● The Systems Manager Agent does not scrape Prometheus metrics from containers, and Amazon Managed Service for Prometheus is not a visualization tool.
● AWS AppConfig manages application configurations rather than metrics, and OpenSearch is primarily for logs and search, not Prometheus time-series metrics.
Workflow: Amazon Managed Service for Prometheus for collection of metrics and Amazon Managed Grafana for visualization and analytics.
up a cross-account CloudWatch Logs destination in the log-archive account and attach
CloudWatch Logs cross-account subscriptions can deliver to Kinesis Data Firehose, and S3 provides secure, centralized, and cost-effective archival with lifecycle policies.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CloudWatch Logs cross-account subscriptions can deliver to Kinesis Data Firehose, and S3 provides secure, centralized, and cost-effective archival.
● Scene fit: CloudWatch Logs cross-account subscriptions can deliver to Kinesis Data Firehose, and S3 provides secure, centralized, and cost-effective archival.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon EFS is not a supported target for Kinesis Data Firehose and is not optimized for low-cost archival storage.
● Using Lambda to write to EFS increases complexity and cost compared to delivering directly to S3 for archival.
● Amazon Redshift is designed for analytics, not inexpensive long-term log archiving, and would be costly for this use case.
Workflow: Set up a cross-account CloudWatch Logs destination in the log-archive account and attach subscription filters from each member account → connect the destination to an Amazon Kinesis Data Firehose delivery stream that.
Amazon CloudFront with Lambda@Edge + Application Load Balancer with EC2 Auto Scaling
Caches content at the edge and allows custom logic at edge locations to tailor behavior by path or device, reducing latency and origin load. Offers Layer 7 host and path-based routing to EC2 targets and scales cost-effectively with demand.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Caches content at the edge and allows custom logic at edge locations to tailor behavior by path or device, reducing latency and origin load.
● Scene fit: Caches content at the edge and allows custom logic at edge locations to tailor behavior by path or.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Improves availability and routing performance but lacks caching and per-request path logic, and adds cost.
● Performs DNS-based decisions like geolocation or latency but cannot evaluate HTTP paths per request.
● Provides API front door features but is not needed for EC2-hosted web apps requiring edge caching and adds cost without path-aware caching.
Workflow: Amazon CloudFront with Lambda@Edge → Application Load Balancer with EC2 Auto Scaling.
The S3 bucket policy does not grant the required permission + There
A restrictive or incorrect bucket policy can deny s3:GetObject requests and result in a 403 Access Denied. If the role lacks s3:GetObject permission, has an invalid trust relationship, or an explicit deny applies, S3 returns 403 Access Denied.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A restrictive or incorrect bucket policy can deny s3:GetObject requests and result in a 403 Access Denied.
● Scene fit: A restrictive or incorrect bucket policy can deny s3:GetObject requests and result in a 403 Access Denied. If.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● S3 default encryption does not by itself block access because S3 transparently decrypts objects for authorized principals.
● Blocking egress would typically cause connectivity errors or timeouts rather than an S3 403 Access Denied authorization failure.
● Versioning changes how object versions are stored and retrieved, not whether the caller is authorized to access them.
Workflow: The S3 bucket policy does not grant the required permission → There is a misconfiguration in the instance profile IAM role.
Build a read replica in CloudFormation, upgrade it, promote, and switch clients
Upgrading a replica while the primary serves traffic, then promoting it, enables a brief cutover with minimal downtime.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Upgrading a replica while the primary serves traffic, then promoting it, enables a brief cutover with minimal downtime.
● Scene fit: Upgrading a replica while the primary serves traffic, then promoting it, enables a brief cutover with minimal downtime.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Updating EngineVersion directly on the primary can cause a disruptive outage during the major upgrade.
● Snapshot restore and endpoint changes require a longer cutover and manual steps, increasing downtime risk.
● DMS can work but adds unnecessary complexity for a same-engine major upgrade where native RDS methods suffice.
Workflow: Build a read replica in CloudFormation, upgrade it → promote, and switch clients.
a CloudWatch alarm on StatusCheckFailed_System to invoke EC2 recovery
Auto Recovery migrates the instance to new hardware on system status check failure and retains EBS volumes and instance attributes.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Auto Recovery migrates the instance to new hardware on system status check failure and retains EBS volumes and instance attributes.
● Scene fit: Auto Recovery migrates the instance to new hardware on system status check failure and retains EBS volumes and instance attributes.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Replaces the instance with a new one instead of migrating the same instance, and does not automatically preserve existing EBS attachments or attributes.
● A reboot addresses OS-level issues and does not fix host-level power or networking failures.
● The recover action is only supported on the system status check, not the instance status check.
Workflow: Configure a CloudWatch alarm on StatusCheckFailed_System to invoke EC2 recovery.
a CloudWatch Logs subscription filter to send login events to an AWS
This creates an automated pipeline from log detection to tagging and scheduled termination that satisfies the 45-minute requirement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This creates an automated pipeline from log detection to tagging and scheduled termination that satisfies the 45-minute requirement.
● Scene fit: This creates an automated pipeline from log detection to tagging and scheduled termination that satisfies the 45-minute requirement.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces unnecessary orchestration for a simple event-to-action workflow and adds complexity without benefit.
● Depends on human intervention and does not guarantee termination within the enforcement window.
● CloudWatch alarms cannot directly invoke Lambda and would require SNS or another indirection, making this approach unsuitable as stated.
Workflow: Use a CloudWatch Logs subscription filter to send login events to an AWS Lambda function that applies a quarantine tag to the source EC2 instance, and configure an Amazon EventBridge schedule to run every 45 minutes to invoke another Lambda that terminates all instances with that tag.
Amazon GuardDuty for the entire organization with a delegated administrator, capture GuardDuty
GuardDuty supports organization-wide aggregation with a delegated administrator and emits findings to EventBridge, which can route to Firehose for delivery to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: GuardDuty supports organization-wide aggregation with a delegated administrator and emits findings to EventBridge, which can route to Firehose.
● Scene fit: GuardDuty supports organization-wide aggregation with a delegated administrator and emits findings to EventBridge, which can route to Firehose.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Inspector focuses on vulnerability and network exposure assessments, not continuous threat detection from logs for EC2 attacks.
● Macie primarily discovers and classifies sensitive data in S3 and is not designed to detect EC2 attack patterns.
● Kinesis Data Streams does not natively deliver to S3 without a consumer and this setup misses the centralized organization administrator model.
Workflow: Enable Amazon GuardDuty for the entire organization with a delegated administrator → capture GuardDuty findings via an EventBridge rule in the admin account, and deliver them to an S3 bucket using Kinesis Data Firehose.
a CloudWatch Logs subscription filter that uses Amazon Kinesis Data Firehose to
Kinesis Data Firehose is a supported destination for CloudWatch Logs subscriptions and continuously streams logs into S3 for automation. Transitioning to S3 Glacier Deep Archive after.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Kinesis Data Firehose is a supported destination for CloudWatch Logs subscriptions and continuously streams logs into S3 for.
● Scene fit: Kinesis Data Firehose is a supported destination for CloudWatch Logs subscriptions and continuously streams logs into S3 for.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch Logs subscription filters do not support AWS DataSync as a delivery target, so this pairing is not valid.
● Export tasks are batch operations that require orchestration and are not as automated or continuous as a streaming subscription.
Workflow: Create a CloudWatch Logs subscription filter that uses Amazon Kinesis Data Firehose to deliver all logs to an S3 bucket → Configure an S3 lifecycle rule that transitions log objects to S3.
AWS Config change-triggered rule for the S3 bucket and federated role invoking
Change-triggered AWS Config rules evaluate in near real time and can invoke SSM Automation for automatic remediation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Change-triggered AWS Config rules evaluate in near real time and can invoke SSM Automation for automatic remediation.
● Scene fit: Change-triggered AWS Config rules evaluate in near real time and can invoke SSM Automation for automatic remediation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● A periodic schedule introduces delay and relies on custom scanning, so it is not near real time.
● Drift detection is not event-driven, has limited coverage, and is not immediate for access changes.
● API-event patterns can be brittle and lack resource state evaluation, making drift detection less reliable.
Workflow: AWS Config change-triggered rule for the S3 bucket and federated role invoking SSM Automation via Lambda to roll back.
a Lambda function that calls API Gateway GetSdk to fetch the client
UpdateStage fires on stage updates including rollbacks, allowing Lambda to consistently pull SDKs via GetSdk and publish them to S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: UpdateStage fires on stage updates including rollbacks, allowing Lambda to consistently pull SDKs via GetSdk and publish them to S3.
● Scene fit: UpdateStage fires on stage updates including rollbacks, allowing Lambda to consistently pull SDKs via GetSdk and publish them.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Only runs within pipeline executions and can miss rollbacks or stage changes that occur outside the pipeline.
● CreateDeployment is not emitted during stage rollbacks, so SDK updates would be skipped in those cases.
● CodePipeline does not natively support actions that run only on failure and this would not handle manual or automated rollbacks.
Workflow: Use a Lambda function that calls API Gateway GetSdk to fetch the client SDKs and upload them to S3, and invoke it from an EventBridge rule that matches API Gateway UpdateStage events.
DeletionPolicy Retain on all resources, delete the stack, create a new stack
This preserves resources during deletion and uses resource import to bring them under the new stack name.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This preserves resources during deletion and uses resource import to bring them under the new stack name.
● Scene fit: This preserves resources during deletion and uses resource import to bring them under the new stack name.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Hooks validate or block operations but cannot retain or import existing resources into a new stack.
● Stack names are immutable and change sets do not support renaming.
● Snapshot causes deletion of the original volume, violating the no-deletion requirement and preventing import of that resource.
Workflow: Set DeletionPolicy Retain on all resources, delete the stack → create a new stack, import the existing resources, → remove Retain.
EventBridge rule for EC2 Instance State-change Notification to SNS
EventBridge natively emits EC2 state-change events and can route them to SNS for real-time alerts across all transitions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge natively emits EC2 state-change events and can route them to SNS for real-time alerts across all transitions.
● Scene fit: EventBridge natively emits EC2 state-change events and can route them to SNS for real-time alerts across all transitions.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● AWS Health focuses on account or service-level issues and does not emit per-instance state-change events.
● Status checks detect health issues and recovery actions but do not notify on every start, stop, or terminate state change.
● CloudTrail is for audit logs and S3 delivery and does not provide timely, comprehensive notifications for all instance state transitions.
Workflow: EventBridge rule for EC2 Instance State-change Notification to SNS.
Build two AWS CodePipeline pipelines for STAGING and PRODUCTION, use a single
This design supports automatic STAGING releases, introduces an approval gate for PRODUCTION, and keeps source management simple with one repo and branches.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This design supports automatic STAGING releases, introduces an approval gate for PRODUCTION, and keeps source management simple with.
● Scene fit: This design supports automatic STAGING releases, introduces an approval gate for PRODUCTION, and keeps source management simple with.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Adds unnecessary complexity with multiple repos and violates the requirement for automatic STAGING deployments by forcing approval there as well.
● Contradicts the requirement for a manual approval in PRODUCTION and unnecessarily splits the source across multiple repositories.
● Fails to provide separate environment pipelines and omits the required manual approval for PRODUCTION.
Workflow: Build two AWS CodePipeline pipelines for STAGING and PRODUCTION → use a single GitHub repository with separate environment branches → trigger on branch commits → deploy via AWS CloudFormation, and require a manual approval only in the PRODUCTION pipeline.
EventBridge Health events + Lambda to recreate EC2 on failure; Aurora with
Event-driven instance replacement avoids warm-standby costs and an Aurora Replica enables automatic cross-AZ failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Event-driven instance replacement avoids warm-standby costs and an Aurora Replica enables automatic cross-AZ failover.
● Scene fit: Event-driven instance replacement avoids warm-standby costs and an Aurora Replica enables automatic cross-AZ failover.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● DRS focuses on EBS-backed replication and adds ongoing replication cost, and a single-instance Aurora lacks automatic cross-AZ failover.
● Licensing prohibits Auto Scaling and this does not guarantee EFA and instance store handling.
● Maintains a paid standby and leaves the database as a single point of failure.
Workflow: EventBridge Health events + Lambda to recreate EC2 on failure → Aurora with one cross-AZ Replica.
green with a rolling update, then change the ALB listener default action
This prepares green and performs an instant cutover at the load balancer without DNS propagation.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This prepares green and performs an instant cutover at the load balancer without DNS propagation.
● Scene fit: This prepares green and performs an instant cutover at the load balancer without DNS propagation.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Routes traffic to green before it is ready, risking downtime.
● Route 53 cannot alias to a target group and DNS TTL delays prevent an immediate cutover.
● Weighted DNS still relies on TTL and, with a single ALB, cannot directly target separate target groups.
Workflow: Deploy green with a rolling update, → change the ALB listener default action to the green target group.
During the rolling update, suspend HealthCheck, ReplaceUnhealthy, AZRebalance, AlarmNotification, and ScheduledActions +
Suspending these processes prevents unexpected Auto Scaling actions from conflicting with CloudFormation's rolling update steps. Adjusting the threshold prevents CloudFormation from rolling back the whole stack.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Suspending these processes prevents unexpected Auto Scaling actions from conflicting with CloudFormation's rolling update steps.
● Scene fit: Suspending these processes prevents unexpected Auto Scaling actions from conflicting with CloudFormation's rolling update steps. Adjusting the threshold.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Replacing the entire group is not a troubleshooting-focused step and can hide the root cause of rolling update failures.
● These processes are required for ELB-integrated rolling updates and suspending them would stop progress and break the rollout.
● Enabling success signals can prolong or block updates due to signal timeouts, which is counterproductive when diagnosing rollout stalls.
Workflow: During the rolling update → suspend HealthCheck, ReplaceUnhealthy, AZRebalance, AlarmNotification, and ScheduledActions → Set MinSuccessfulInstancesPercent in the AutoScalingRollingUpdate policy to avoid full stack rollback when a small portion of instances fail →.
AWS Config with required-tag rules, publish evaluations via Amazon SNS to an
AWS Config provides resource inventory and tag-compliance evaluations, which can be streamed and stored for reporting and dashboards in QuickSight.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config provides resource inventory and tag-compliance evaluations, which can be streamed and stored for reporting and dashboards in QuickSight.
● Scene fit: AWS Config provides resource inventory and tag-compliance evaluations, which can be streamed and stored for reporting and dashboards.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudTrail captures API activity logs rather than evaluating current resource tag compliance, so it is not suitable for a compliance dashboard.
● Service Catalog focuses on curated product portfolios and does not offer account-wide tag compliance evaluation across all resources.
● Resource Explorer helps discover resources but does not provide rule-based compliance checks or continuous evaluations needed for a compliance dashboard.
Workflow: Configure AWS Config with required-tag rules → publish evaluations via Amazon SNS to an AWS Lambda function that writes results to Amazon S3, and visualize in Amazon QuickSight.
a subdomain named na.shopwidget.net with failover routing, setting the us-west-2 ALB as
Stacking latency at the apex with Region-specific failover under each subdomain delivers lowest-latency routing with automatic cross-Region failover.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Stacking latency at the apex with Region-specific failover under each subdomain delivers lowest-latency routing with automatic cross-Region failover.
● Scene fit: Stacking latency at the apex with Region-specific failover under each subdomain delivers lowest-latency routing with automatic cross-Region failover.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Multivalue routing returns multiple healthy records and does not guarantee lowest-latency selection or coordinated Regional failover behavior.
● Inverts the required layering by putting latency at the subdomain level and failover at the apex, which does not.
● Weighted and geolocation routing do not ensure lowest latency or automatic failover between Regions.
Workflow: Create a subdomain named na.shopwidget.net with failover routing, setting the us-west-2 ALB as primary and the eu-central-1 ALB as secondary → create eu.shopwidget.net with failover routing, setting the eu-central-1 ALB as primary.
AWS Secrets Manager with rotation enabled
Purpose-built secret storage with IAM access controls and managed rotation workflows for RDS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Purpose-built secret storage with IAM access controls and managed rotation workflows for RDS.
● Scene fit: Purpose-built secret storage with IAM access controls and managed rotation workflows for RDS.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Secure storage and retrieval, but no native automatic rotation for RDS; would require custom rotation logic.
● Eliminates passwords by using short-lived tokens, but does not meet a policy that requires rotating a database password.
● Manages encryption keys, not application secrets; cannot store or rotate database credentials.
Workflow: AWS Secrets Manager with rotation enabled.
EC2 with SSM Patch Manager and Amazon Data Lifecycle Manager for EBS
SSM Patch Manager automates OS patching and DLM schedules EBS snapshots with retention, minimizing effort.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SSM Patch Manager automates OS patching and DLM schedules EBS snapshots with retention, minimizing effort.
● Scene fit: SSM Patch Manager automates OS patching and DLM schedules EBS snapshots with retention, minimizing effort.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● RDS for Oracle does not support Oracle RAC, so it cannot host a RAC deployment.
● Is heavier and misuses CI/CD tools for patching compared to native patch automation and DLM.
● Aurora is not Oracle and does not support Oracle RAC, so it is not a viable target for this workload.
Workflow: EC2 with SSM Patch Manager and Amazon Data Lifecycle Manager for EBS snapshots.
Host a static website in Amazon S3, distribute via Amazon CloudFront, and
S3 provides serverless static hosting, CloudFront accelerates global delivery, and AWS WAF helps protect against common web exploits.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: S3 provides serverless static hosting, CloudFront accelerates global delivery, and AWS WAF helps protect against common web exploits.
● Scene fit: S3 provides serverless static hosting, CloudFront accelerates global delivery, and AWS WAF helps protect against common web exploits.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces unnecessary compute for a static site and GuardDuty detects threats but does not block web exploits at the edge.
● EC2 is not serverless, Redis does not accelerate static asset delivery to users, and Shield focuses on DDoS rather than general web exploit mitigation.
● Fargate is serverless but unnecessary for static content, Redis is not used to serve static website assets, and this lacks the global CDN acceleration required.
Workflow: Host a static website in Amazon S3, distribute via Amazon CloudFront, and apply an AWS WAF web ACL.
EventBridge rule for EC2 Auto Scaling events targeting Systems Manager Run Command
Uses native Auto Scaling launch/terminate events to trigger an immediate Run Command on the batch instance to refresh the file.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Uses native Auto Scaling launch/terminate events to trigger an immediate Run Command on the batch instance to refresh the file.
● Scene fit: Uses native Auto Scaling launch/terminate events to trigger an immediate Run Command on the batch instance to refresh the file.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Relies on frequent API polling and custom code, which is inefficient and can miss rapid changes.
● AWS Config is for configuration compliance and is not optimized for near real-time scaling reactions.
● Lifecycle hooks and SNS notifications do not update the batch instance by themselves and add operational overhead.
Workflow: EventBridge rule for EC2 Auto Scaling events targeting Systems Manager Run Command.
Migrate the Jenkins build workload to AWS CodeBuild and enable encryption for
CodeBuild is a fully managed service that supports KMS-backed artifact encryption and removes the need to manage EC2 hosts.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeBuild is a fully managed service that supports KMS-backed artifact encryption and removes the need to manage EC2 hosts.
● Scene fit: CodeBuild is a fully managed service that supports KMS-backed artifact encryption and removes the need to manage EC2 hosts.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Bucket encryption protects objects in S3, but this keeps Jenkins to manage and does not guarantee artifacts are encrypted before upload.
● Patching and EBS encryption secure instances and disks but do not directly ensure artifact encryption and still leave significant operational overhead.
● ACM issues TLS certificates for in-transit security and does not provide at-rest encryption for build artifacts.
Workflow: Migrate the Jenkins build workload to AWS CodeBuild and enable encryption for build artifacts.
AWS Systems Manager Run Command to apply the one-time configuration on the
Run Command lets you execute ad hoc configuration on managed instances securely without distributing or storing SSH keys. Using the CodeBuild IAM role eliminates the need.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Run Command lets you execute ad hoc configuration on managed instances securely without distributing or storing SSH keys.
● Scene fit: Run Command lets you execute ad hoc configuration on managed instances securely without distributing or storing SSH keys.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Keeping credentials in S3, even encrypted, increases exposure and is not the recommended secret management approach for builds.
● Secrets Manager does not use the SecureString type, so this wording is incorrect even though Secrets Manager itself can.
Workflow: Use AWS Systems Manager Run Command to apply the one-time configuration on the instance instead of ssh and scp with a key from Amazon S3 → Grant the CodeBuild service role least-privilege.
Patch Manager with per-environment baselines and Patch Groups; tag by env and
Use Patch Manager baselines and Patch Groups mapped by tags to apply different rules and sequence environments.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Use Patch Manager baselines and Patch Groups mapped by tags to apply different rules and sequence environments.
● Scene fit: Use Patch Manager baselines and Patch Groups mapped by tags to apply different rules and sequence environments.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Requires custom scripts and manual patch selection, increasing effort and risk.
● Golden AMI replacement is heavier operationally and does not directly enforce per-environment patch rules.
● Provides scheduling but not approval rules and per-environment baselines without Patch Manager.
Workflow: Patch Manager with per-environment baselines and Patch Groups → tag by env and OS.
Regional API Gateways in each Region with Route 53 latency-based routing, local
Latency-based routing sends clients to the lowest-latency Regional API, and global tables provide multi-Region data replication.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Latency-based routing sends clients to the lowest-latency Regional API, and global tables provide multi-Region data replication.
● Scene fit: Latency-based routing sends clients to the lowest-latency Regional API, and global tables provide multi-Region data replication.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Edge-optimized reduces edge-to-Region latency but keeps compute in one Region, so it is not active-active multi-Region.
● Failover is active-passive, so only one Region serves traffic in steady state.
● Global Accelerator does not support API Gateway as an endpoint and this design will not work.
Workflow: Regional API Gateways in each Region with Route 53 latency-based routing, local Lambdas, DynamoDB global tables.
access logging on the Application Load Balancer and configure the destination to
ALB access logging writes detailed request logs directly to S3, which satisfies the requirement to collect load balancer logs. Using the awslogs driver centralizes container STDOUT/STDERR.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ALB access logging writes detailed request logs directly to S3, which satisfies the requirement to collect load balancer.
● Scene fit: ALB access logging writes detailed request logs directly to S3, which satisfies the requirement to collect load balancer.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Macie focuses on sensitive data discovery and classification rather than low-latency log ingestion or analysis.
● CloudWatch Detailed Monitoring provides higher-resolution metrics, not request-level access logs or delivery to S3.
● CloudWatch Logs export tasks are batch operations and not intended for near-real-time streaming to S3, making this unsuitable.
Workflow: Enable access logging on the Application Load Balancer and configure the destination to the specified S3 bucket → Use the awslogs log driver in ECS task definitions → install the CloudWatch Logs.
Parameter Store SecureString with KMS; use instance roles and CodeDeploy on-prem registration
SecureString uses KMS for encryption at rest, TLS provides encryption in transit, and instance roles (EC2 profile and on-prem via CodeDeploy registration) allow runtime retrieval without.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SecureString uses KMS for encryption at rest, TLS provides encryption in transit, and instance roles (EC2 profile and.
● Scene fit: SecureString uses KMS for encryption at rest, TLS provides encryption in transit, and instance roles (EC2 profile and.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Embedding secrets in deployment artifacts violates the requirement to avoid storing credentials in code or packages and risks exposure.
● CodeDeploy service roles do not push secrets into instances; applications should retrieve secrets at runtime using an attached role.
● You cannot attach a standalone policy to a server; you must use roles (instance profile for EC2 and on-prem registration) for least-privilege access.
Workflow: Parameter Store SecureString with KMS → use instance roles and CodeDeploy on-prem registration → apps fetch at runtime.
Designate a Firewall Manager admin and apply org-level WAF policies to auto-attach
Firewall Manager centrally enforces organization-wide WAF policies and automatically applies them to matching resources across accounts and Regions.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Firewall Manager centrally enforces organization-wide WAF policies and automatically applies them to matching resources across accounts and Regions.
● Scene fit: Firewall Manager centrally enforces organization-wide WAF policies and automatically applies them to matching resources across accounts and Regions.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Config can detect and remediate per rule, but it is not the purpose-built, continuous, org-wide policy engine for WAF across accounts and Regions.
● StackSets can roll out templates but does not continuously discover and auto-attach WAF for newly created resources or enforce ongoing policy.
● SCPs restrict permissions but cannot attach or enforce resource configurations like WAF associations.
Workflow: Designate a Firewall Manager admin and apply org-level WAF policies to auto-attach web ACLs to all public ALBs and API Gateway in every Region.
an Auto Scaling lifecycle hook on instance termination and use an Amazon
A termination lifecycle hook with EventBridge and Lambda invoking SSM Run Command lets you reliably retrieve instance logs before shutdown and store them durably in S3.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A termination lifecycle hook with EventBridge and Lambda invoking SSM Run Command lets you reliably retrieve instance logs.
● Scene fit: A termination lifecycle hook with EventBridge and Lambda invoking SSM Run Command lets you reliably retrieve instance logs.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● VPC Flow Logs record network traffic metadata rather than application or system logs, so they will not capture the.
● Health check settings do not address collecting or preserving instance logs and there is no evidence they are misconfigured.
Workflow: Configure an Auto Scaling lifecycle hook on instance termination and use an Amazon EventBridge rule to trigger an AWS Lambda function that runs AWS Systems Manager Run Command to pull application logs.
the application to persist session state in an Amazon ElastiCache for Redis
ElastiCache (Redis) is an in-memory, shared session store that delivers microsecond-level latency and survives instance replacement.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: ElastiCache (Redis) is an in-memory, shared session store that delivers microsecond-level latency and survives instance replacement.
● Scene fit: ElastiCache (Redis) is an in-memory, shared session store that delivers microsecond-level latency and survives instance replacement.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Stickiness ties a user to a single instance and local storage is lost on scale-in or replacement, so sessions still break.
● DynamoDB provides a shared, durable store but has higher latency than an in-memory cache for per-request session lookups.
● S3 is durable but request latency is too high for frequent session reads and writes during web traffic.
Workflow: Configure the application to persist session state in an Amazon ElastiCache for Redis cluster.
CloudWatch Logs metric filter with CloudWatch alarm to SNS
A metric filter matches "CRITICAL" log lines, a CloudWatch alarm triggers on the metric, and SNS sends the email.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: A metric filter matches "CRITICAL" log lines, a CloudWatch alarm triggers on the metric, and SNS sends the email.
● Scene fit: A metric filter matches "CRITICAL" log lines, a CloudWatch alarm triggers on the metric, and SNS sends the email.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Firewall Manager enforces security policies but does not generate alerts from arbitrary CloudWatch Logs severities.
● Logs Insights does not directly send results to SNS in real time and is not intended for per-line alerting.
● Synthetics tests endpoints and APIs, not the content of CloudWatch Logs.
Workflow: CloudWatch Logs metric filter with CloudWatch alarm to SNS.
AWS SAM with CodeDeploy canary (DeploymentPreference Canary10Percent10Minutes)
SAM integrates Lambda aliases with CodeDeploy for automated canary and rollback with minimal configuration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: SAM integrates Lambda aliases with CodeDeploy for automated canary and rollback with minimal configuration.
● Scene fit: SAM integrates Lambda aliases with CodeDeploy for automated canary and rollback with minimal configuration.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Handles configuration rollout, not Lambda version traffic shifting or deployment rollback.
● Swaps entire environments and lacks built-in percentage-based Lambda traffic shifting.
● Shifts API stage traffic but does not manage Lambda version deployment or automatic rollback.
Workflow: AWS SAM with CodeDeploy canary (DeploymentPreference Canary10Percent10Minutes).
An Auto Scaling scale-out event occurred during the deployment, so the new
When scale-out happens mid-deployment, new instances are initialized with the most recent successful revision, leading to a mixed fleet even if the deployment completes successfully.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: When scale-out happens mid-deployment, new instances are initialized with the most recent successful revision, leading to a mixed fleet even if the deployment completes successfully.
● Scene fit: When scale-out happens mid-deployment, new instances are initialized with the most recent successful revision, leading to a mixed.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch alarms do not cause CodeDeploy to report success while leaving some instances on an older version.
● If IAM access to the revision were blocked, CodeDeploy would mark the deployment as failed rather than successful.
● A stale template affects boot settings, but CodeDeploy would still update targeted instances, so this does not explain a successful deployment with mixed versions.
Workflow: An Auto Scaling scale-out event occurred during the deployment, so the new instances launched with the last successfully deployed revision.
AWS CodeDeploy Blue/Green for EC2/On-Premises with an ALB, set Original instances to
CodeDeploy Blue/Green provisions a replacement fleet behind the load balancer, supports traffic shifting after validation, and can automatically terminate the original instances after a configured wait.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy Blue/Green provisions a replacement fleet behind the load balancer, supports traffic shifting after validation, and can automatically.
● Scene fit: CodeDeploy Blue/Green provisions a replacement fleet behind the load balancer, supports traffic shifting after validation, and can automatically.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Rolling updates replace instances in batches within the same environment and do not provision a completely separate fleet for.
● An in-place rolling update modifies the existing fleet instead of creating a full green environment and lacks a built-in.
● Custom approach can work but lacks the native orchestration, status tracking, and rollback safety that CodeDeploy provides for Blue/Green.
Workflow: Use AWS CodeDeploy Blue/Green for EC2/On-Premises with an ALB → set Original instances to terminate after a 90 minute wait period.
an EventBridge rule for Trusted Advisor updates and send to SNS +
EventBridge can match Trusted Advisor events and route to SNS for immediate notifications. A periodic Lambda can query TA findings and push alerts to SNS subscribers.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: EventBridge can match Trusted Advisor events and route to SNS for immediate notifications.
● Scene fit: EventBridge can match Trusted Advisor events and route to SNS for immediate notifications. A periodic Lambda can query.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Weekly digests are delayed and not suitable for prompt alerts.
● EventBridge does not natively target SES; use SNS or Lambda for email delivery.
● AWS Health does not emit Trusted Advisor check change events.
Workflow: Create an EventBridge rule for Trusted Advisor updates and send to SNS → Schedule a Lambda every 30 minutes to call the Trusted Advisor API and publish to SNS → Run a Lambda hourly to write TA results to CloudWatch Logs → add a metric filter and alarm to notify.
AWS Config organization aggregator with managed encrypted-volumes rule; centralize results and SNS
Uses AWS Config managed compliance with org-wide aggregation for minimal operations and near real-time detection.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Uses AWS Config managed compliance with org-wide aggregation for minimal operations and near real-time detection.
● Scene fit: Uses AWS Config managed compliance with org-wide aggregation for minimal operations and near real-time detection.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Is custom log parsing with event correlation and is high effort, not a native compliance solution.
● Works but managing StackSets across many accounts and Regions adds operational overhead versus org-level aggregation.
● Security Hub relies on AWS Config for these checks and does not itself provide the managed rule or org-wide Config aggregation.
Workflow: AWS Config organization aggregator with managed encrypted-volumes rule → centralize results and SNS.
AWS Config with the cloudtrail-enabled or cloudtrail-security-trail-enabled managed rule and an EventBridge
AWS Config continuously evaluates compliance, provides an audit trail of evaluations, and with EventBridge-triggered Lambda can automatically remediate when CloudTrail is disabled.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: AWS Config continuously evaluates compliance, provides an audit trail of evaluations, and with EventBridge-triggered Lambda can automatically remediate.
● Scene fit: AWS Config continuously evaluates compliance, provides an audit trail of evaluations, and with EventBridge-triggered Lambda can automatically remediate.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Group-based denies can be bypassed by roles or the root user, and scheduled polling does not provide continuous compliance.
● A periodic Lambda check lacks configuration compliance history and can miss or delay detection compared to event-driven compliance evaluation.
● An SCP can block certain actions but does not ensure CloudTrail is enabled initially, does not create a compliance.
Workflow: Use AWS Config with the cloudtrail-enabled or cloudtrail-security-trail-enabled managed rule and an EventBridge rule that triggers a Lambda remediation on noncompliant evaluations to turn CloudTrail back on and record the change.
a custom CloudWatch metric and publish statistic sets that roll up the
Publishing statistic sets aggregates many samples into a single PutMetricData call per minute, providing 1 minute granularity at lower cost.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Publishing statistic sets aggregates many samples into a single PutMetricData call per minute, providing 1 minute granularity at lower cost.
● Scene fit: Publishing statistic sets aggregates many samples into a single PutMetricData call per minute, providing 1 minute granularity at.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Default metrics do not accept custom application data and dimensions do not create new metrics, so this would not capture the script output.
● Synthetics can run canaries but does not leverage the EC2 script output and generally costs more than aggregating custom metrics when 1 minute granularity is sufficient.
● High resolution and frequent publishing increase cost and are unnecessary when only 1 minute granularity is required.
Workflow: Create a custom CloudWatch metric and publish statistic sets that roll up the 5 second results, sending one update every 60 seconds.
Amazon EventBridge to trigger an AWS Lambda function every hour to enumerate
This detective control identifies and remediates drift on a schedule while keeping launches unblocked, minimizing impact on CI/CD. AWS Config continuously evaluates AMI compliance and, paired.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This detective control identifies and remediates drift on a schedule while keeping launches unblocked, minimizing impact on CI/CD.
● Scene fit: This detective control identifies and remediates drift on a schedule while keeping launches unblocked, minimizing impact on CI/CD.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Amazon Inspector focuses on vulnerability and exposure assessments and does not reliably enforce or detect AMI allowlists for compliance.
● An SCP is a preventative control that would block launches with unapproved AMIs and likely disrupt or slow the.
Workflow: Use Amazon EventBridge to trigger an AWS Lambda function every hour to enumerate running EC2 instances in the staging and production VPCs, compare their AMI IDs against the approved list → publish.
CodeDeploy pre-traffic hook to run Lambda validation and fail the deployment on
CodeDeploy lifecycle event BeforeAllowTraffic can invoke a validation Lambda and failing it aborts and rolls back. Deployment-group alarms can stop the deployment, send SNS notifications, and.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: CodeDeploy lifecycle event BeforeAllowTraffic can invoke a validation Lambda and failing it aborts and rolls back.
● Scene fit: CodeDeploy lifecycle event BeforeAllowTraffic can invoke a validation Lambda and failing it aborts and rolls back. Deployment-group alarms.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● CloudWatch Alarms evaluate metrics, not arbitrary payloads sent directly from Lambda.
● Post-shift validation does not meet the requirement to test before live traffic.
● AppConfig validators apply to configuration deployments, not Lambda code deployments via CodeDeploy.
Workflow: Use CodeDeploy pre-traffic hook to run Lambda validation and fail the deployment on errors → Attach a CloudWatch alarm to the deployment group so an alarm breach notifies SNS and triggers rollback → Run validation during deployment and drive a deployment-group CloudWatch alarm from the checks to stop and roll back.
EC2 replication workers with Auto Scaling in us-east-2 that continuously push deltas
Using S3 as a cross-region transport avoids VPC peering, and continuous EC2-based replication with a scaling metric maintains a hot EFS for low RPO and RTO.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Using S3 as a cross-region transport avoids VPC peering, and continuous EC2-based replication with a scaling metric maintains.
● Scene fit: Using S3 as a cross-region transport avoids VPC peering, and continuous EC2-based replication with a scaling metric maintains.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Introduces a scheduled 30-minute interval that prevents a hot copy and increases RPO.
● You cannot mount an EFS file system across AWS Regions or from an unpeered VPC, so direct cross-region mounting.
● DataSync requires network access to both locations and typically runs on schedules, so it cannot meet the no-connectivity constraint.
Workflow: Run EC2 replication workers with Auto Scaling in us-east-2 that continuously push deltas to an S3 bucket in ap-northeast-1 using a custom backlog metric for scaling, and operate a second replication fleet.
EventBridge schedule to Step Functions with Lambda tasks and SNS notifications
Provides stateful orchestration with retries and Choice branches for Region fallback, execution history, scheduling, and email via SNS.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Provides stateful orchestration with retries and Choice branches for Region fallback, execution history, scheduling, and email via SNS.
● Scene fit: Provides stateful orchestration with retries and Choice branches for Region fallback, execution history, scheduling, and email via SNS.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Powerful DAG orchestration but heavier to operate and unnecessary for a simple nightly backup workflow.
● Supports scheduled EBS backups and cross-Region copies with notifications, but does not offer conditional multi-Region fallback logic per run.
● Single Lambda risks timeouts and complex error handling, and AWS Config is not an execution audit trail.
Workflow: EventBridge schedule to Step Functions with Lambda tasks and SNS notifications.
Modular stacks with Outputs Export and Fn::ImportValue
Cross-stack references via exported Outputs and Fn::ImportValue decouple lifecycles while reliably sharing values.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-stack references via exported Outputs and Fn::ImportValue decouple lifecycles while reliably sharing values.
● Scene fit: Cross-stack references via exported Outputs and Fn::ImportValue decouple lifecycles while reliably sharing values.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Nested stacks share a lifecycle and rely on parameters, which tightly couples components and limits independent updates.
● StackSets target multi-account or multi-Region rollout, not inter-stack value sharing, and still couples stacks via parameters.
● Externalizes state and bypasses native cross-stack wiring, adding complexity and weaker change tracking.
Workflow: Modular stacks with Outputs Export and Fn::ImportValue.
Export Outputs from the database stack and import them into the application
Cross-stack exports and Fn::ImportValue allow the application stack to consume database resources while keeping stacks and change sets independent.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Cross-stack exports and Fn::ImportValue allow the application stack to consume database resources while keeping stacks and change sets independent.
● Scene fit: Cross-stack exports and Fn::ImportValue allow the application stack to consume database resources while keeping stacks and change sets independent.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● StackSets are intended for managing stacks across multiple accounts and Regions, not for cross-stack references within a single account and Region.
● Nested stacks tie resources into one parent stack where change sets cascade across the hierarchy, reducing independence between teams.
● Parameter Store does not create CloudFormation-managed dependencies or change set visibility between stacks, risking drift and coupling outside CloudFormation.
Workflow: Export Outputs from the database stack and import them into the application stack with Fn::ImportValue.
Systems Manager Automation running EC2Rescue via AWSSupport-ExecuteEC2Rescue, plus EventBridge aws.trustedadvisor
Executes EC2Rescue offline to fix Windows networking/RDP issues and surfaces Trusted Advisor events through EventBridge.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Executes EC2Rescue offline to fix Windows networking/RDP issues and surfaces Trusted Advisor events through EventBridge.
● Scene fit: Executes EC2Rescue offline to fix Windows networking/RDP issues and surfaces Trusted Advisor events through EventBridge.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Handles hardware or hypervisor issues, not OS-level misconfigurations like RDP or firewall.
● Relies on instance connectivity and lacks targeted offline OS repair for Windows networking problems.
● Replaces instances rather than repairing them and does not address Trusted Advisor event integration.
Workflow: Systems Manager Automation running EC2Rescue via AWSSupport-ExecuteEC2Rescue, → EventBridge aws.trustedadvisor.
one SSM service role and a single hybrid activation (limit 400), then
Hybrid activations at scale with one service role and an activation limit allow many servers to register via code/ID and show as mi- managed instances.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Hybrid activations at scale with one service role and an activation limit allow many servers to register via code/ID and show as mi- managed instances.
● Scene fit: Hybrid activations at scale with one service role and an activation limit allow many servers to register via.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Quick Setup cannot replace hybrid activations for on-prem, and EC2 instance profiles are only for EC2.
● Per-host roles and activations are unnecessary and on-prem managed instances use the mi- prefix, not i-.
● IAM users and long-term keys are not used for hybrid enrollment; activations are required.
Workflow: Use one SSM service role and a single hybrid activation (limit 400), → install SSM Agent and register each with the activation code/ID → they appear as mi-.
Define a CloudWatch Logs metric filter for 'CRITICAL', emit a custom metric
Metric filters turn matching log lines into metrics that alarms can watch and notify via SNS with low operational effort.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Metric filters turn matching log lines into metrics that alarms can watch and notify via SNS with low operational effort.
● Scene fit: Metric filters turn matching log lines into metrics that alarms can watch and notify via SNS with low operational effort.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Possible but adds code, scaling, and maintenance; not the simplest path for basic string-match alerting.
● Transit Gateway flow logs are IP flow records and won't include third-party firewall severity details like CRITICAL.
● EventBridge does not natively inspect CloudWatch Logs content; it needs a subscription or metric intermediary.
Workflow: Define a CloudWatch Logs metric filter for 'CRITICAL' → emit a custom metric, and use a CloudWatch alarm to notify an SNS topic.
an ASG launch lifecycle hook and complete it only after readiness
Pauses instances in Pending:Wait during launch so registration occurs only after signaling readiness via CompleteLifecycleAction.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Pauses instances in Pending:Wait during launch so registration occurs only after signaling readiness via CompleteLifecycleAction.
● Scene fit: Pauses instances in Pending:Wait during launch so registration occurs only after signaling readiness via CompleteLifecycleAction.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Adjusts when a target is considered healthy but does not prevent initial registration with the target group.
● Meters traffic after registration; it does not block initial registration.
● Controls connection draining during scale-in, not registration on scale-out.
Workflow: Add an ASG launch lifecycle hook and complete it only after readiness.
the Trusted Advisor low utilization check and create an Amazon EventBridge rule
This uses Trusted Advisor events with EventBridge and SSM Automation to add a human approval gate while minimizing custom code and orchestration.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This uses Trusted Advisor events with EventBridge and SSM Automation to add a human approval gate while minimizing.
● Scene fit: This uses Trusted Advisor events with EventBridge and SSM Automation to add a human approval gate while minimizing.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● Requires custom polling, state handling, and stream processing, which is high effort compared to managed signals from Trusted Advisor.
● Trusted Advisor does not emit per-check notifications to SNS, so this integration is not the supported path for reacting.
● An aggregated alarm cannot pinpoint specific idle instances and per-instance alarms are costly and complex to manage.
Workflow: Enable the Trusted Advisor low utilization check and create an Amazon EventBridge rule for that check's status changes → target a Lambda function that invokes an SSM Automation runbook with a manual.
Send aggregated per-minute counts of charges and refunds from the instances to
Publishing a single custom metric to CloudWatch with an alarm and SNS provides a simple, automated, and low-cost solution that meets the requirements.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: Publishing a single custom metric to CloudWatch with an alarm and SNS provides a simple, automated, and low-cost.
● Scene fit: Publishing a single custom metric to CloudWatch with an alarm and SNS provides a simple, automated, and low-cost.
● Engineering sense: Prefer the managed path with the fewest custom failure points.
Alternatives to reject
● EventBridge does not ingest or store CloudWatch custom metrics, so this cannot directly produce a metric for alarms.
● Can visualize data but is operationally heavy and more expensive than using native CloudWatch custom metrics and alarms.
● While feasible, adopting AMP and Grafana adds complexity and cost compared to a single CloudWatch custom metric and alarm.
Workflow: Send aggregated per-minute counts of charges and refunds from the instances to Amazon CloudWatch as a custom metric → create a CloudWatch alarm with Amazon SNS notifications, and use the AWS console for graphs.
.ebextensions/alb.config with an option_settings section that defines Rules for aws:elbv2:listener:default and rely
This declarative configuration is applied by Beanstalk during deployments, ensuring the ALB redirect is consistently managed via the pipeline.
Decision method:
● Technical viability: Native AWS services support the required workflow and integrations.
● Requirement fit: This declarative configuration is applied by Beanstalk during deployments, ensuring the ALB redirect is consistently managed via the pipeline.
● Scene fit: This declarative configuration is applied by Beanstalk during deployments, ensuring the ALB redirect is consistently managed via the pipeline.
● Engineering sense: Prefer the managed path that satisfies the stated cost, timing, security, and operational constraints with few failure points.
Alternatives to reject
● Relies on imperative instance scripts and out-of-band API calls, which is brittle and not the recommended way to manage Beanstalk ALB listener rules.
● Approach requires direct environment permissions and bypasses the pipeline, which the scenario explicitly restricts.
● CodePipeline does not natively orchestrate deployments using the EB CLI, so this integration is unsupported.
Workflow: Add .ebextensions/alb.config with an option_settings section that defines Rules for aws:elbv2:listener:default and rely on CodePipeline to deploy the change.