Vertex Macro | Financial Cloud Cloud · AWS Exam
AWS Solutions Architect Professional
| Article |
|---|
| 01 AWS Solutions Architect Professional SAP-C02 lecture notes |
| 02 AWS DevOps Certified Professional DOP-C02 lecture notes |
Lecture content for government technology advisors, public-welfare and education leaders, data and security owners, mental-health service networks, and large-institution decision makers.
One-Time EMR Capacity Mix
A one-time 300 TB analytics job should protect cluster control and HDFS while using Spot for restartable tasks.
Recommended architecture
● Use On-Demand Instances for EMR master and core nodes.
● Use Spot Instances for task nodes.
● Persist final results to durable storage.
● Terminate the cluster after completion.
Design reasoning
● Technical workability: Master and core nodes maintain cluster state and data; task nodes can be replaced after interruption.
● Requirement fit: The design lowers cost without risking the whole job through control-node interruption.
● Scene fit: The cluster runs for only eight hours and is not recurring.
● Engineering common sense: Do not purchase long-term commitments for a single execution.
Alternatives to reject
● Reserved master capacity is not cost-effective for one use.
● Spot master or core nodes can cause cluster failure and HDFS loss.
● Reserved master and core nodes require unnecessary commitment.
● All On-Demand capacity misses safe task-node savings.
Workflow: Launch stable master and core → add Spot task nodes → process 300 TB → save output → terminate.
Re-platforming Containerized Applications
A low-change migration preserves the application’s existing boundaries while replacing self-managed infrastructure with managed services.
Recommended architecture
● Store tested OpenJDK container images in Amazon ECR.
● Run the containers on Amazon ECS.
● Migrate MySQL to Amazon RDS with AWS Database Migration Service.
Design reasoning
● Technical workability: ECS runs Docker workloads, RDS supports MySQL, and DMS moves relational data with limited interruption.
● Requirement fit: OpenJDK removes commercial Java licensing, while managed compute and database services reduce operations.
● Scene fit: The application already uses containers and MySQL, so re-platforming preserves both patterns.
● Engineering common sense: Modernize one layer at a time rather than changing compute and the data model together.
Alternatives to reject
● EC2 rehosting works but retains server and database administration.
● Replacing MySQL with DynamoDB changes relational behavior and application code.
● Converting containers to Lambda is a refactor, not a minimal-change migration.
Workflow: Build OpenJDK image → test → push to ECR → deploy to ECS → replicate MySQL with DMS → validate → cut over.
Rolling AMI Updates with CloudFormation
An Auto Scaling group can replace instances gradually while maintaining a minimum healthy service capacity.
Recommended configuration
● Update the launch template or launch configuration to the new AMI.
● Define AutoScalingRollingUpdate in the CloudFormation update policy.
● Configure batch size, pause time, and minimum instances in service.
● Monitor health before continuing each batch.
Design reasoning
● Technical workability: CloudFormation coordinates controlled instance replacement in the Auto Scaling group.
● Requirement fit: The fleet receives the new AMI without replacing every instance simultaneously.
● Scene fit: The application already uses Auto Scaling and requires low downtime.
● Engineering common sense: Health checks must validate application readiness, not only instance launch.
Alternatives to reject
● A change set previews changes but does not control fleet replacement.
● A full blue/green stack can work but adds duplicated infrastructure and DNS switching.
● DeletionPolicy does not define rolling replacement behavior.
● An in-place AMI change cannot modify already running instances.
Workflow: Update template → replace one batch → wait for health → continue batches → finish or roll back.
SCP Allow Lists Must Permit Every Action
Under an SCP allow-list strategy, an action must be allowed at every applicable organization level before IAM permissions can use it.
Key behavior
● SCPs define the maximum permissions for member-account principals.
● IAM policies grant permissions only inside that maximum.
● S3 bucket creation must appear in applicable allow-list SCPs from root through OU and account.
Design reasoning
● Technical workability: Effective permission is the intersection of SCP boundaries and IAM grants.
● Requirement fit: Adding the missing S3 allow resolves the denial without changing unrelated permissions.
● Scene fit: The IAM identity already has S3 permission but an SCP limits it.
● Engineering common sense: Troubleshoot authorization from outer guardrails inward.
Alternatives to reject
● The IAM allow cannot override a missing SCP allow.
● SCPs support allow-list and deny-list strategies.
● IAM and SCP documents need not be identical.
● Adding more IAM policies cannot expand beyond the SCP boundary.
Workflow: Request → root SCP → OU SCP → account SCP → IAM policy → resource policy and conditions.
Resilient Global Web Platform
A public web platform needs global static delivery, web protection, resilient compute, and a managed relational database.
Recommended architecture
● Store static content in S3 and deliver it with CloudFront.
● Attach AWS WAF for common application attacks.
● Run web servers in a Multi-AZ Auto Scaling group.
● Use Aurora MySQL for managed database availability and scale.
Design reasoning
● Technical workability: CloudFront caches content, WAF filters HTTP requests, Auto Scaling replaces failed servers, and Aurora provides managed relational storage.
● Requirement fit: The design improves performance, security, availability, and operations together.
● Scene fit: Static content, web compute, and relational data need different services.
● Engineering common sense: Use managed layers rather than self-managing every component.
Alternatives to reject
● Self-managed MySQL on EC2 retains database operations and omits edge protection.
● Global Accelerator is not a static-content cache.
● S3 Transfer Acceleration speeds uploads, not website delivery.
● Standard RDS can work, but the selected Aurora design better fits managed scaling and availability.
Workflow: User → CloudFront and WAF → Auto Scaling web tier → Aurora MySQL.
Cross-Account Administration with IAM Roles
Consolidated billing does not grant administrative access to member accounts. Access must be explicitly delegated.
Recommended architecture
● Keep administrator IAM identities in the management account.
● Create administrative IAM roles in the Development and Test accounts.
● Configure each role’s trust policy to allow approved management-account principals to assume it.
● Attach the required administrative permissions in each target account.
Design reasoning
● Technical workability: AWS STS issues temporary credentials when a trusted administrator assumes the target-account role.
● Requirement fit: Administrators use one home identity and gain controlled access without duplicate IAM users.
● Scene fit: The accounts are linked for billing but remain separate security boundaries.
● Engineering common sense: Define permissions where the resources live and trust only the central identities that need them.
Alternatives to reject
● Consolidated billing changes payment aggregation, not IAM authorization.
● A role created only in the management account cannot automatically administer target-account resources.
● Duplicating long-lived IAM users in every account increases credential and offboarding work.
● Resource policies alone do not provide general account administration.
Workflow: Sign in centrally → request AssumeRole → target trust policy evaluates → receive temporary credentials → administer target account.
Improving Page Load with Two Cache Layers
CloudFront and ElastiCache solve different performance problems and can be used together.
Recommended architecture
● Cache reusable website content in CloudFront.
● Store sessions and frequently accessed query results in ElastiCache.
● Keep EC2 Auto Scaling and RDS as the authoritative compute and data layers.
Design reasoning
● Technical workability: CloudFront reduces network distance and origin requests; ElastiCache reduces backend database reads.
● Requirement fit: Page response improves without deploying a second Region.
● Scene fit: The application serves cacheable content and dynamic sessions or queries.
● Engineering common sense: Measure cache hit ratio and avoid caching personalized responses incorrectly.
Alternatives to reject
● State Manager configures instances and does not replace Auto Scaling.
● Lowering the scale trigger adds compute but does not eliminate repeated work.
● A second Region introduces cost and data-management complexity.
● Database scaling alone does not improve global static delivery.
Workflow: Viewer → CloudFront → application on miss → ElastiCache → RDS on cache miss.
Month-Partitioned Blog Storage and Private Delivery
Time-based blog content can use one partitioned bucket, automated lifecycle rules, and a private CloudFront origin.
Recommended architecture
● Store entries under month-based S3 prefixes.
● Apply lifecycle policies by prefix or tags.
● Use one CloudFront distribution for scalable delivery.
● Restrict the bucket to a CloudFront origin identity or equivalent private-origin control.
Design reasoning
● Technical workability: S3 prefixes organize objects, lifecycle rules transition them, and CloudFront caches delivery.
● Requirement fit: The design reduces storage and distribution management.
● Scene fit: Blog entries have time-based retention and access patterns.
● Engineering common sense: Logical partitioning does not require separate buckets or distributions.
Alternatives to reject
● Two distributions duplicate configuration without improving lifecycle management.
● A minimum TTL of zero prevents useful caching.
● Forwarding unnecessary query strings fragments the cache.
● Duplicating every entry across two buckets doubles storage and administration without a recovery requirement.
Workflow: Publish under monthly prefix → CloudFront retrieves privately → edge caches → lifecycle transitions older prefixes.
Serving Multiple TLS Domains
Different client capabilities and unrelated domains require certificate delivery methods that match each endpoint.
Recommended architecture
● Use CloudFront dedicated-IP SSL when legacy clients do not support SNI.
● Use an Application Load Balancer HTTPS listener with multiple certificates and SNI for modern clients.
Design reasoning
● Technical workability: Dedicated IPs support non-SNI viewers; ALB selects the correct certificate from the requested hostname.
● Requirement fit: Separate certificates can protect unrelated domains on shared managed endpoints.
● Scene fit: Some users have legacy TLS clients, while application domains remain distinct.
● Engineering common sense: Pay for dedicated IP delivery only where client compatibility requires it.
Alternatives to reject
● One SAN certificate could cover known names but does not directly preserve separate certificate administration.
● Classic Load Balancer does not support multiple SNI certificates on one listener.
● A wildcard certificate covers subdomains of one parent domain, not unrelated domains.
● Installing certificates on every backend increases rotation and exposure work.
Workflow: Viewer connects → endpoint determines SNI capability → dedicated-IP CloudFront or ALB listener presents matching certificate.
Durable Audit Logs for Regional and Global Services
A reliable security audit trail should cover every Region and global services while isolating logs from normal workload access.
Recommended architecture
● Create one multi-Region AWS CloudTrail trail.
● Include global service events such as IAM activity.
● Deliver logs to a dedicated Amazon S3 bucket.
● Protect the bucket with restrictive policies, versioning, encryption, and MFA Delete where appropriate.
Design reasoning
● Technical workability: CloudTrail records management activity for EC2, RDS, IAM, and other services.
● Requirement fit: S3 provides durable storage, while access restrictions and deletion controls protect confidentiality and integrity.
● Scene fit: The account uses resources across several Regions and requires central compliance evidence.
● Engineering common sense: Keep audit logs in a dedicated location with fewer administrators than production resources.
Alternatives to reject
● A trail that omits global events misses IAM activity.
● Reusing a general-purpose bucket and relying mainly on ACLs weakens isolation.
● Separate trails for console, SDK, and CLI are unnecessary because CloudTrail records all three channels.
● SNS delivery notifications do not replace log protection or validation.
Workflow: Account action → CloudTrail event → protected S3 storage → controlled auditor access → retention monitoring.
Month-End MySQL Read Scaling
A predictable read spike is best handled by managed database operations and temporary read replicas.
Recommended architecture
● Migrate self-managed MySQL from EC2 to Amazon RDS for MySQL.
● Add read replicas before month-end.
● Route reporting queries to replicas.
● Remove unnecessary replicas after the peak.
Design reasoning
● Technical workability: RDS manages backups and maintenance, while replicas offload reads from the writer.
● Requirement fit: Performance improves during a predictable window with lower operational risk.
● Scene fit: The bottleneck is read-heavy month-end processing.
● Engineering common sense: Scale the query path rather than repeatedly changing storage hardware.
Alternatives to reject
● Snapshotting and switching EBS types is slow and risky.
● Lambda cannot transparently resize a running EC2 database.
● Vertical scaling remains bounded.
● ALB pre-warming is unrelated to database reads.
● gp2 is not automatically faster for every heavy-I/O workload.
Workflow: Migrate to RDS → add replicas before peak → route reads → monitor lag → remove replicas afterward.
Testing CloudFormation Changes in CI/CD
Infrastructure changes should be tested and previewed before CloudFormation executes them.
Recommended architecture
● Trigger CodePipeline from GitHub changes.
● Use CodeBuild for tests and template validation.
● Create a CloudFormation change set.
● Review or automatically execute the approved change set using CloudFormation actions.
Design reasoning
● Technical workability: CodePipeline coordinates source, build, and CloudFormation deployment stages.
● Requirement fit: Changes are repeatable and their resource impact is visible before execution.
● Scene fit: The source contains CloudFormation templates.
● Engineering common sense: The service that owns the resource change should execute it.
Alternatives to reject
● Lambda and CodeArtifact are unnecessary for normal template deployment.
● CodeArtifact is a package repository, not the natural template execution service.
● CodeDeploy deploys application revisions and cannot execute CloudFormation change sets.
● Direct update without tests weakens safety.
Workflow: GitHub commit → CodePipeline → CodeBuild tests → change set → approval → CloudFormation execution.
Cost-Efficient OpsWorks Layers
Related application components can share one OpsWorks stack while using separate layers for different roles.
Recommended architecture
● Keep the existing customer-support web application in one layer.
● Add a second layer for the video-chat component.
● Use one custom Chef recipe for shared installation or integration tasks where practical.
● Assign instances and lifecycle events to the appropriate layer.
Design reasoning
● Technical workability: OpsWorks layers organize instances, recipes, packages, and lifecycle configuration within a stack.
● Requirement fit: The new component remains independently configurable without duplicating an entire stack.
● Scene fit: Web support and video chat belong to one application environment but have different server roles.
● Engineering common sense: Separate operational roles, not every feature, into different infrastructure estates.
Alternatives to reject
● A second complete stack duplicates configuration and management cost.
● Mixing both server roles in one layer reduces control over scaling and recipes.
● Separate recipes for every identical dependency add unnecessary maintenance.
● Replacing OpsWorks during a feature release expands scope.
Workflow: Update stack → create video layer → attach recipe → launch instances → validate integration → scale each layer independently.
Routing Static Paths to an S3 CloudFront Origin
A CloudFront distribution can use separate origins for dynamic application requests and static assets.
Recommended architecture
● Add the S3 bucket as a second CloudFront origin.
● Create a cache behavior for the static path pattern.
● Route that behavior to S3.
● Keep the ALB as the default origin for dynamic paths.
Design reasoning
● Technical workability: CloudFront cache behaviors select origins from URL path patterns.
● Requirement fit: Static assets stop returning application-origin 404 errors and gain edge caching.
● Scene fit: One domain serves both dynamic and static content.
● Engineering common sense: Perform origin selection at the CDN layer that knows both origins.
Alternatives to reject
● Global Accelerator does not cache content or route by HTTP path.
● ALB cannot use an S3 bucket as a target.
● Header conditions do not make S3 an ALB target.
● Duplicating static files on application instances defeats the S3 origin.
Workflow: Viewer request → CloudFront path match → S3 origin for static assets or ALB for dynamic content.
Warm Standby Disaster Recovery
Warm standby keeps a functional but scaled-down copy of the production environment ready for rapid expansion.
Recommended architecture
● Maintain a reduced application fleet in the recovery Region.
● Keep a synchronized standby database and current application artifacts.
● Continuously monitor replication and recovery health.
● During disaster, scale the standby environment and redirect traffic.
Design reasoning
● Technical workability: A running environment avoids building every component during the incident.
● Requirement fit: Warm standby balances recovery speed and ongoing cost.
● Scene fit: The business needs faster recovery than backup-and-restore but cannot justify full active-active capacity.
● Engineering common sense: Recovery time includes database promotion, capacity scaling, dependency validation, and DNS movement.
Alternatives to reject
● Full duplicate production capacity provides faster recovery but costs more.
● Pilot light keeps only core components running and may take longer to scale and validate.
● Backup and restore has the lowest steady cost but the longest recovery.
● Multi-AZ alone protects against an Availability Zone failure, not a regional disaster.
Workflow: Replicate continuously → detect disaster → promote database → scale application → validate dependencies → redirect traffic → monitor.
Global Anycast Entry Point for Regional ALBs
A multi-Region application can use AWS Global Accelerator as one public entry point for healthy regional Application Load Balancers.
Recommended architecture
● Create a Global Accelerator.
● Define endpoint groups for the required Regions.
● Register the regional ALBs as endpoints.
● Create a Route 53 public alias record at the apex domain pointing to the accelerator.
Design reasoning
● Technical workability: Global Accelerator uses static anycast IP addresses and routes users over the AWS network to healthy endpoints.
● Requirement fit: The design supports an apex domain, improves global routing, and minimizes DNS and failover operations.
● Scene fit: The application already has public regional ALBs and multi-Region data services.
● Engineering common sense: Use one managed global traffic layer instead of manually maintaining regional IP mappings.
Alternatives to reject
● Transit Gateway is for private network routing and cannot be a public application endpoint.
● Route 53 Resolver inbound endpoints serve private DNS queries, not public apex routing.
● A public CNAME at the zone apex is not the appropriate model; Route 53 alias records solve the apex requirement.
Workflow: Client → Global Accelerator anycast address → health-based regional endpoint group → ALB → application.
Vendor Access with External ID
A vendor should access customer resources through a least-privilege IAM role and an External ID condition.
Recommended architecture
● Create a role in the customer account.
● Grant only required actions and resources.
● Trust the vendor’s AWS account or role.
● Require a customer-specific External ID.
● Let the vendor call STS for temporary credentials.
Design reasoning
● Technical workability: The trust policy validates both vendor identity and External ID before role assumption.
● Requirement fit: Access is temporary, revocable, and protected against confused-deputy cross-customer mistakes.
● Scene fit: One vendor application serves several customers.
● Engineering common sense: Never share personal or root access keys with a service provider.
Alternatives to reject
● Long-term IAM user keys increase exposure and rotation work.
● Amazon Connect is unrelated to third-party API authorization.
● Personal access keys expose the customer’s own identity and permissions.
● Trusting the vendor without an External ID weakens customer separation.
Workflow: Vendor requests role with External ID → IAM trust evaluates → STS issues temporary session → vendor performs scoped work.
Why SCPs Do Not Restrict Service-Linked Roles
Service Control Policies restrict the permissions available to principals in member accounts, but AWS service-linked roles are exempt from SCP restrictions.
Key behavior
● A service-linked role is predefined by an AWS service.
● Its trust relationship allows that service to perform required actions.
● SCP denies do not limit actions performed through service-linked roles.
Design reasoning
● Technical workability: ECS can continue service-managed operations through its service-linked role even when account principals face an SCP deny.
● Requirement fit: This behavior explains why the expected restriction did not affect the observed actions.
● Scene fit: The ECS cluster uses a service-linked role created for AWS service integration.
● Engineering common sense: First identify which principal made the request before diagnosing policy evaluation.
Incorrect explanations to reject
● Resources do not operate outside the organization’s jurisdiction.
● A parent Allow cannot override an explicit Deny elsewhere in the applicable SCP path.
● The default FullAWSAccess SCP does not cancel an explicit Deny.
● Editing the default SCP is not the cause of service-linked-role behavior.
Investigation workflow: Review CloudTrail principal → identify service-linked role → map applicable SCPs → separate service actions from user or role actions.
Resolving Current AMIs in CloudFormation
CloudFormation can resolve AWS-managed public Systems Manager parameters that contain current AMI IDs.
Recommended architecture
● Reference the appropriate public SSM AMI parameter in the template.
● Resolve it during stack creation or update.
● Run update-stack intentionally when the fleet should adopt the newer image.
● Use a controlled rolling update policy.
Design reasoning
● Technical workability: CloudFormation parameter types can resolve current AMI identifiers from Parameter Store.
● Requirement fit: Templates avoid hard-coded regional AMI IDs.
● Scene fit: The organization wants current AWS images through deliberate stack updates.
● Engineering common sense: Latest should not mean untested; validate images before production rollout.
Alternatives to reject
● Service Catalog distributes products but does not itself resolve every latest AMI.
● AWS Config evaluates compliance and is not an AMI lookup service.
● State Manager maintains instance configuration and is not the CloudFormation AMI parameter source.
● Existing instances do not change until the stack is updated.
Workflow: Template resolves SSM parameter → launch template changes → rolling stack update → health validation.
Three-AZ Application and Database Availability
A resilient web platform distributes compute across three Availability Zones and uses a database design with writer availability.
Recommended architecture
● Run EC2 instances in a three-AZ Auto Scaling group.
● Place an Application Load Balancer before the fleet.
● Use an Aurora architecture that supports the required writer availability.
● Create a Route 53 alias to the ALB.
Design reasoning
● Technical workability: Auto Scaling replaces failed compute, ALB routes only to healthy targets, and Route 53 aliasing tracks the load balancer endpoint.
● Requirement fit: The design remains available through instance or AZ failures.
● Scene fit: The application requires managed relational writer resilience.
● Engineering common sense: Avoid unnecessary AZs or additional databases that do not improve the stated recovery target.
Alternatives to reject
● A generic A record should not hard-code ALB addresses.
● One database writer without suitable failover remains a risk.
● Four AZs add complexity without a stated need.
● Extra replicas do not automatically create writer availability.
Workflow: Route 53 alias → ALB → healthy Auto Scaling targets → available Aurora writer.
Event-Driven Processing for IoT Files
Files arriving every five minutes should be processed when they arrive rather than waiting for a nightly polling job.
Recommended architecture
● Convert the existing Python processing logic into an AWS Lambda function.
● Configure S3 object-created event notifications to invoke Lambda.
● Process each uploaded file and write its values to Amazon RDS.
● Delete or lifecycle the source object only after successful processing.
Design reasoning
● Technical workability: S3 can invoke Lambda directly for new objects, and Lambda can connect to the database with the correct networking and credentials.
● Requirement fit: Data becomes available shortly after upload with little infrastructure management.
● Scene fit: Daily processing takes only about ten minutes, indicating each individual file is small enough for event-based execution.
● Engineering common sense: Prefer push events over one-minute polling for discrete object arrivals.
Alternatives to reject
● Larger EC2 fleets and frequent cron execution waste capacity and require coordination.
● CloudTrail data events plus EventBridge can detect uploads but add cost and delay when native S3 notifications suffice.
● Multiple scheduled rules create duplicate-invocation risk without improving timeliness.
Workflow: Device upload → S3 event → Lambda processing → RDS write → success handling and object cleanup.
One Direct Connect for Multiple Regions
A Direct Connect Gateway extends one private connectivity design to VPCs in multiple AWS Regions.
Recommended architecture
● Create a Direct Connect Gateway.
● Associate regional virtual private gateways.
● Connect private virtual interfaces to the gateway.
● Advertise approved prefixes through BGP.
Design reasoning
● Technical workability: Direct Connect Gateway links private VIF connectivity to VGWs across supported Regions.
● Requirement fit: The data center reaches both regional VPCs through centralized, high-bandwidth private connectivity.
● Scene fit: One on-premises site needs access to east and west VPCs.
● Engineering common sense: Use one hub before purchasing duplicate circuits.
Alternatives to reject
● VPC peering does not connect the on-premises data center.
● One VPN to the east Region does not automatically reach the west VPC.
● VPN lacks the predictable performance of Direct Connect.
● Separate Direct Connect connections can work but cost more and add management.
Workflow: Data center → Direct Connect → private VIF → Direct Connect Gateway → regional VGW → VPC.
Buffering Donation Bursts
A donation campaign needs durable burst absorption, elastic workers, and predictable scalable writes.
Recommended architecture
● Put donation requests in Amazon SQS.
● Scale EC2 workers according to queue depth.
● Store processed donations in DynamoDB with provisioned throughput sized for the event.
● Use idempotency and dead-letter handling.
Design reasoning
● Technical workability: SQS protects requests, Auto Scaling adds consumers, and DynamoDB handles high write throughput.
● Requirement fit: Sudden traffic does not overload one synchronous database path.
● Scene fit: The campaign creates a temporary write burst.
● Engineering common sense: Accept requests durably before performing slower processing.
Alternatives to reject
● DynamoDB without a queue does not protect downstream processing from spikes.
● CloudFront does not buffer donation writes.
● One RDS instance remains vertically bounded.
● A dedicated extra-large Oracle host is costly and operationally heavy.
Workflow: Donor submits → SQS → workers scale → validate and write DynamoDB → acknowledge and notify.
Managed Intelligent Contact Center
Automated call handling needs durable work integration, speech understanding, intent recognition, and a managed agent platform.
Recommended architecture
● Use Amazon Connect for inbound calls, routing, and agent workflows.
● Use Amazon Lex for speech recognition and intent detection.
● Use Amazon SQS where backend work must be buffered asynchronously.
● Invoke controlled application integrations for business actions.
Design reasoning
● Technical workability: Connect provides contact-center infrastructure, Lex interprets caller requests, and SQS decouples backend tasks.
● Requirement fit: Common requests can be automated and escalated to agents when necessary.
● Scene fit: The workload is customer calls, not text-only analytics.
● Engineering common sense: Separate conversation state from durable business work.
Alternatives to reject
● Alexa for Business is not a customer contact-center platform.
● Kinesis and Comprehend do not provide call speech recognition and routing.
● Rekognition analyzes images and video.
● Polly produces speech but does not detect caller intent.
Workflow: Caller → Connect flow → Lex intent → business integration or SQS → response or agent transfer.
Secure Storage for Millions of User Preferences
Small per-user preferences are a direct key-value workload and should not be split across multiple storage systems.
Recommended architecture
● Store one item per user in Amazon DynamoDB.
● Authenticate social users through web identity federation.
● Use STS temporary credentials.
● Apply DynamoDB fine-grained access control so each user can access only the permitted item.
Design reasoning
● Technical workability: DynamoDB handles millions of small records with managed availability and horizontal scaling.
● Requirement fit: The design is highly available, cost-effective, scalable, and avoids long-lived mobile credentials.
● Scene fit: Each preference record is only about 4 KB and is naturally addressed by user ID.
● Engineering common sense: Keep a small atomic record in one database unless a second storage service solves a real limitation.
Alternatives to reject
● RDS read replicas can scale reads, but database accounts are not a practical identity model for millions of social users.
● A public application server plus RDS adds server and connection management.
● Storing each tiny preference object in S3 and its pointer in DynamoDB doubles requests and complexity.
Workflow: Social login → identity token → STS role session → item-level authorization → read or update preference item.
Understanding S3 Version IDs After Enabling Versioning
Objects created before S3 Versioning is enabled behave differently from versions created afterward.
Key behavior
● An object that existed before versioning has a null version ID.
● Updating that object after versioning creates a new version with a generated version ID.
● The original null version remains available unless deleted.
● Objects never updated after versioning still have only the original null version.
Design reasoning
● Technical workability: S3 maintains earlier object versions under the same key.
● Requirement fit: Version history explains which configuration files have one or two retrievable versions.
● Scene fit: Some files were updated after versioning, while others were not.
● Engineering common sense: Version IDs are opaque identifiers, not chronological sequence numbers.
Incorrect assumptions to reject
● Enabling versioning does not retroactively assign generated IDs to existing objects.
● Version IDs are not predictable or sequential.
● Updating an old object does not replace its original null version; it adds a new version.
● An unchanged pre-versioning object does not gain a second version automatically.
Workflow: Pre-versioning object → enable versioning → update key → preserve null version plus new generated version.
Efficient Migration of NGINX and MySQL
A small interactive web application can move to AWS by creating a resilient web tier and using a managed database migration path.
Recommended architecture
● Launch NGINX EC2 instances in two Availability Zones.
● Copy application files through an S3 staging location.
● Migrate MySQL data with AWS Database Migration Service.
● Place the web servers behind an Elastic Load Balancer.
● Create a Route 53 alias record for the load balancer.
Design reasoning
● Technical workability: EC2 preserves the existing NGINX application, DMS moves database records, and the load balancer distributes traffic.
● Requirement fit: The design improves availability without a major application rewrite.
● Scene fit: The source consists of one NGINX server and one MySQL database.
● Engineering common sense: Rehost the application tier first and manage traffic through a stable load-balanced endpoint.
Alternatives to reject
● A dynamic application cannot simply run as an S3 static website.
● Application Discovery Service inventories workloads but does not migrate the web server.
● A private hosted zone does not serve public users.
● A Multi-AZ database is not located in only one Availability Zone.
Workflow: Build two web nodes → copy files → migrate database → validate → place behind load balancer → switch Route 53.
Device-Specific Static Content at the Edge
Static content can be delivered globally while edge logic selects the appropriate version for each device type.
Recommended architecture
● Move static assets to Amazon S3.
● Configure the bucket as a CloudFront origin.
● Use Lambda@Edge to inspect the viewer’s User-Agent header.
● Rewrite or route the request to the correct device-specific object.
● Cache the selected response at the edge.
Design reasoning
● Technical workability: Lambda@Edge runs during CloudFront request processing and can alter the requested URI based on headers.
● Requirement fit: Content is selected near users and cached, reducing EC2 load and response time.
● Scene fit: The application serves different static assets to different device classes.
● Engineering common sense: Move static routing decisions to the content-delivery layer rather than scaling general-purpose servers.
Alternatives to reject
● Network Load Balancer operates at Layer 4 and cannot inspect User-Agent.
● Route 53 routes DNS queries and cannot evaluate HTTP headers.
● A simple CloudFront cache behavior cannot classify arbitrary device headers without edge logic.
● Keeping all static content on EC2 preserves the original load problem.
Workflow: Request → edge reads header → URI rewrite → S3 object → cached device-specific response.
Private Fargate Egress for ECR Image Pulls
A Fargate task in a private subnet needs an outbound path to retrieve its container image unless all required private endpoints are configured.
Recommended architecture
● Keep automatic public IP assignment disabled.
● Place a NAT Gateway in a public subnet.
● Give the public subnet a route to an internet gateway.
● Route private-subnet outbound traffic to the NAT Gateway.
● Allow required HTTPS traffic through network controls.
Design reasoning
● Technical workability: The task uses its private address and obtains outbound connectivity through NAT to reach ECR dependencies.
● Requirement fit: Containers launch without exposing tasks directly to the internet.
● Scene fit: The reported connection timeout occurs while pulling from the ECR registry endpoint.
● Engineering common sense: Public egress infrastructure belongs in a public subnet; workloads remain private.
Alternatives to reject
● ECR uses interface endpoints, not the gateway endpoint described in the alternative.
● Fargate requires awsvpc networking and cannot switch to bridge mode.
● A NAT Gateway in a private subnet cannot reach an internet gateway.
● Enabling task public IPs conflicts with the private design.
Workflow: Task starts → private route → NAT Gateway → ECR image pull → container launch.
Managed Encrypted Storage for Large Files
Large files needing fully managed storage, KMS encryption, and HTTPS-only access fit Amazon S3.
Recommended architecture
● Store files in S3.
● Use SSE-KMS with the approved customer-managed key.
● Add a bucket policy denying requests that do not use secure transport.
● Apply least-privilege IAM and key policies.
Design reasoning
● Technical workability: S3 supports large objects, managed durability, KMS encryption, and policy-based HTTPS enforcement.
● Requirement fit: The design removes server and volume administration.
● Scene fit: The data consists of files rather than database items.
● Engineering common sense: Encryption needs access control and transport policy to form a complete protection chain.
Alternatives to reject
● EC2 and EBS retain server and volume operations.
● File Gateway is unnecessary after migration when ongoing hybrid access is not required.
● SSE-S3 lacks customer-managed KMS control.
● DynamoDB is unsuitable for large files and would require application-level splitting.
Workflow: Authorized HTTPS PUT → S3 uses KMS key → encrypted object stored → policy denies insecure GET or PUT.
Streaming Genomic Analytics
Continuous genomic data needs streaming ingestion, scalable processing, and a warehouse for analytical queries.
Recommended architecture
● Ingest records with Amazon Kinesis Data Streams.
● Use stream consumers for near-real-time processing.
● Use Amazon EMR for large-scale transformation where required.
● Load curated analytical results into Amazon Redshift.
Design reasoning
● Technical workability: Kinesis handles continuous records, EMR processes large datasets, and Redshift supports analytical SQL.
● Requirement fit: The pipeline supports timely analysis and durable warehouse reporting.
● Scene fit: Genomic measurements arrive continuously rather than as occasional API requests.
● Engineering common sense: Separate ingestion from heavy analytics so processing delays do not block producers.
Alternatives to reject
● API Gateway and SQS add an indirect message workflow for continuous streaming.
● Kinesis does not analyze existing S3 objects simply by adding SQS.
● Firehose delivers data but is not the complete processing chain described by an active Kinesis client.
● QuickSight visualizes results and does not replace warehouse processing.
Workflow: Data producer → Kinesis → consumers or EMR → Redshift → analytical queries.
End-to-End HTTPS with CloudFront and ALB
CloudFront and an ALB origin need trusted certificates and policies that enforce HTTPS on both connections.
Recommended architecture
● Import or request a trusted certificate in ACM for the ALB.
● Configure CloudFront Viewer Protocol Policy as HTTPS Only.
● Use a supported ACM or IAM certificate for the CloudFront custom domain.
● Set the origin protocol to HTTPS.
Design reasoning
● Technical workability: CloudFront validates a trusted origin certificate and presents a trusted viewer certificate.
● Requirement fit: HTTP is not accepted anywhere in the delivery chain.
● Scene fit: The retail application uses CloudFront before an ALB.
● Engineering common sense: Certificate names must match both viewer and origin hostnames.
Alternatives to reject
● Self-signed origin certificates are not trusted by CloudFront.
● Viewer certificates cannot be imported from S3.
● Match Viewer permits HTTP when the viewer uses HTTP.
● TLS termination only at CloudFront leaves the origin connection unprotected.
Workflow: HTTPS viewer → CloudFront certificate → HTTPS origin request → ALB certificate → application.
Replacing Fragile NAT Instances
When half of outbound requests fail in a two-AZ application, one failed NAT path is a strong architectural signal.
Recommended architecture
● Replace EC2 NAT instances with managed NAT Gateways.
● Deploy one NAT Gateway in each active Availability Zone.
● Route each private subnet to the NAT Gateway in the same AZ.
● Monitor NAT metrics and third-party API connectivity.
Design reasoning
● Technical workability: NAT Gateway provides managed scaling and availability within its AZ for private-subnet internet egress.
● Requirement fit: The map API becomes reachable without maintaining NAT instance health or capacity.
● Scene fit: Approximately 50 percent success across two AZ paths indicates one broken egress route.
● Engineering common sense: Avoid cross-AZ dependency and remove self-managed network appliances where a managed equivalent exists.
Alternatives to reject
● Enlarging a NAT instance may address throughput but not a failed instance or route.
● Blaming the provider does not explain consistent success through the other application path.
● A network ACL error is possible, but the NAT-instance pattern more directly explains the observed split behavior.
Workflow: Create NAT Gateway per AZ → update private routes → test outbound calls → retire NAT instances.
Immediate Credential Scanning in CodeCommit
Exposed IAM access keys should be detected when code is pushed, then contained immediately.
Recommended architecture
● Configure CodeCommit push events to invoke AWS Lambda.
● Scan new commits for IAM access-key patterns.
● Notify the developer and security team when credentials are found.
● Disable the exposed IAM key and require controlled replacement.
Design reasoning
● Technical workability: CodeCommit events provide timely invocation, and Lambda can inspect submitted content and call IAM APIs.
● Requirement fit: Event-driven scanning responds faster than daily inspection and limits the exposure window.
● Scene fit: The security issue originates in source-code submissions.
● Engineering common sense: Contain a leaked credential before generating or distributing a replacement.
Alternatives to reject
● Amazon Macie discovers sensitive data in S3, not CodeCommit repositories.
● A daily EC2 and Systems Manager scan is delayed and operationally heavier.
● Rotating a key without immediately disabling the exposed key leaves an active risk.
● KMS stores encryption keys; it is not a repository for replacement IAM credentials.
Workflow: Developer push → event → Lambda scan → finding → disable key → notify owner → remove secret from history → issue replacement securely.
Blue/Green Delivery for OpsWorks
Separate OpsWorks stacks provide isolated blue and green environments for validation and rapid rollback.
Recommended architecture
● Clone the current OpsWorks stack.
● Deploy the new Chef and application revisions to the green stack.
● Validate it independently.
● Shift traffic using the existing routing layer.
● Preserve blue for rollback.
Design reasoning
● Technical workability: Independent stacks prevent changes from modifying serving instances.
● Requirement fit: The design supports controlled cutover and quick recovery.
● Scene fit: The application already uses OpsWorks and Chef.
● Engineering common sense: Deployment tooling should match the platform being operated.
Alternatives to reject
● Elastic Beanstalk rolling deployment changes the platform and modifies serving capacity.
● CodePipeline orchestrates stages but does not itself supply the proposed EC2 rolling strategy.
● CodeBuild builds and tests artifacts; it is not the traffic-shifting deployment service.
● In-place Chef updates weaken rollback.
Workflow: Clone blue → deploy green → test → shift traffic → monitor → return to blue if required.
Low-Disruption Media Cataloging with File Gateway
A large on-premises media archive can adopt cloud object storage and managed facial recognition while preserving its existing file-oriented workflow.
Recommended architecture
● Deploy AWS Storage Gateway File Gateway on premises.
● Let the Media Asset Management system write files through the familiar file interface.
● Store gateway-backed objects in Amazon S3.
● Use AWS Lambda to invoke Amazon Rekognition for media analysis.
● Return extracted metadata to the existing catalog system.
Design reasoning
● Technical workability: File Gateway exposes file protocols backed by S3, and Rekognition processes supported S3 media objects.
● Requirement fit: The workflow minimizes disruption and ongoing infrastructure management.
● Scene fit: Existing tools expect files, while the long-term direction is migration to AWS.
● Engineering common sense: Integrate at the existing interface first, then automate cloud-side processing.
Alternatives to reject
● Kinesis Video Streams targets live video, not a historical tape archive.
● Rekognition cannot process Glacier virtual tapes directly in real time.
● Snowball can move bulk data, but self-managed EC2 facial-recognition software adds maintenance.
Workflow: MAM exports file → File Gateway stores in S3 → event invokes processing → Rekognition returns labels or faces → metadata updates MAM.
Edge Offload for an Urgent Traffic Surge
When full migration is not possible, relieve the dominant traffic path without changing the application’s transactional core.
Recommended architecture
● Keep the dynamic website and payment workflow on premises.
● Place Amazon CloudFront in front of high-resolution images and other cacheable assets.
● Attach AWS WAF rules that block common SQL injection and cross-site scripting patterns.
Design reasoning
● Technical workability: CloudFront can use the existing site as an origin and cache static responses near users.
● Requirement fit: Edge caching reduces origin bandwidth and load quickly; WAF adds the requested web-layer protection.
● Scene fit: Large static assets are the immediate scaling pressure, while payments remain dynamic.
● Engineering common sense: Solve the bottleneck first instead of beginning a rushed full migration.
Alternatives to reject
● S3 static website hosting cannot replace dynamic payment processing.
● Building EC2 images, Auto Scaling, hybrid routing, and an ALB is too broad for the deadline.
● A full server migration introduces timing and cutover risk when edge services address the urgent need.
Workflow: Classify cacheable paths → configure origin and cache behavior → attach WAF → test payments → monitor origin offload.
Serving Large Forecast Datasets
Large forecast outputs shared by a server fleet can use EFS, elastic query servers, and short CloudFront caching.
Recommended architecture
● Store the generated 20 GB forecast on EFS.
● Run query servers in an Auto Scaling group behind an ELB.
● Cache responses in CloudFront for 15 minutes.
Design reasoning
● Technical workability: EFS gives all servers shared file access, Auto Scaling handles concurrency, and CloudFront absorbs repeated queries.
● Requirement fit: The design supports 1,500 to 15,000 concurrent users and frequent dataset replacement.
● Scene fit: One large forecast is replaced every 15 minutes.
● Engineering common sense: Avoid indexing data that is completely regenerated before the index earns its cost.
Alternatives to reject
● OpenSearch indexing one billion points every 15 minutes is expensive.
● Changing edge logic does not fix an inefficient index workflow.
● Creating one billion S3 objects per update creates extreme object and request overhead.
● One fixed query server cannot handle the range safely.
Workflow: Forecast written to EFS → elastic servers query it → CloudFront caches responses for 15 minutes.
Low-Cost Photo Processing and Archival
Photo processing is asynchronous, parallel, and tolerant of worker interruption, making it suitable for queued Spot capacity.
Recommended architecture
● Place photo-processing jobs in Amazon SQS.
● Scale EC2 Spot workers according to queue depth.
● Store active inputs in the appropriate online object class.
● Archive completed photos to S3 Glacier when immediate access is no longer required.
Design reasoning
● Technical workability: SQS preserves jobs and allows another worker to retry after interruption.
● Requirement fit: Spot reduces compute cost, while Glacier reduces long-term archive cost.
● Scene fit: Workers compete for independent photo jobs.
● Engineering common sense: Keep durable inputs and outputs outside disposable workers.
Alternatives to reject
● Moving unprocessed photos to Standard-IA creates retrieval charges before immediate processing.
● SNS does not provide a durable competing-consumer backlog or queue-depth scaling signal.
● Manually terminating workers is weaker than Auto Scaling.
● Standard-IA is less economical than Glacier for long-term finished archives that are rarely retrieved.
Workflow: Upload → SQS job → Spot worker → result stored → archive transition → worker scales down.
Cross-Account Private Hosted Zone Association
A Route 53 private hosted zone resolves records only for VPCs associated with that zone.
Recommended procedure
● In the hosted-zone owner account, authorize association with the application VPC in the second account.
● In the VPC owner account, associate the VPC with the private hosted zone.
● Delete the temporary association authorization after completion.
● Verify VPC DNS support and query the database CNAME.
Design reasoning
● Technical workability: Route 53 supports cross-account private hosted-zone associations through authorization followed by association.
● Requirement fit: EC2 instances resolve the centralized private database name without duplicating DNS zones.
● Scene fit: The record is correct, but the application VPC is not yet linked to the zone.
● Engineering common sense: Fix name-resolution scope rather than hard-coding changing database addresses.
Alternatives to reject
● VPC peering does not automatically associate a private hosted zone with another VPC.
● Private hosted zones cannot be associated with each other for record replication.
● Editing /etc/resolv.conf with an RDS IP is brittle because addresses can change during failover.
Workflow: Authorize → associate → remove authorization → resolve CNAME → connect to RDS.
Serverless CI/CD with Gradual Lambda Deployment
A serverless delivery pipeline should model the application, build and test artifacts, orchestrate releases, and shift production traffic safely.
Recommended architecture
● Define Lambda, API Gateway, DynamoDB, and related permissions with AWS SAM.
● Use CodeBuild to install dependencies, run tests, and package artifacts.
● Use CodePipeline to coordinate source, build, approval, and deployment stages.
● Use CodeDeploy deployment preferences for gradual Lambda traffic shifting and rollback.
Design reasoning
● Technical workability: SAM transforms into CloudFormation, while CodeDeploy can use Lambda aliases for canary or linear deployments.
● Requirement fit: The pipeline automates build and release while limiting production impact.
● Scene fit: The new architecture is entirely serverless.
● Engineering common sense: Keep infrastructure and application deployment versioned together and automate rollback on failed alarms.
Alternatives to reject
● Serverless Application Repository distributes reusable applications but is not a complete CI/CD pipeline.
● OpsWorks is not the natural configuration service for Lambda, API Gateway, and DynamoDB.
● Systems Manager Automation is unnecessary for building code or shifting Lambda alias traffic.
Workflow: Commit → build and test → package → deploy stack → shift alias traffic → monitor → complete or roll back.
Repeatable Highly Available Issue Tracking
A production issue tracker needs repeatable infrastructure, elastic application capacity, durable content delivery, and managed database failover.
Recommended architecture
● Define the environment in AWS CloudFormation.
● Run application instances in an Auto Scaling group across Availability Zones.
● Place an Elastic Load Balancer before the instances.
● Use CloudFront for cacheable content.
● Run the OLTP database on RDS Multi-AZ.
Design reasoning
● Technical workability: CloudFormation reproduces the stack, Auto Scaling replaces unhealthy compute, and RDS Multi-AZ provides managed standby failover.
● Requirement fit: The service scales and remains available with low manual recovery effort.
● Scene fit: Issue tracking is transactional and needs a relational OLTP database.
● Engineering common sense: Keep application instances stateless and store durable state in managed services.
Alternatives to reject
● A single EC2 instance or database creates a failure point.
● Read replicas alone do not provide writer failover.
● Hosting durable uploads on instance disks complicates replacement.
● Manual infrastructure deployment causes configuration drift and slow recovery.
Workflow: Route request → CloudFront or load balancer → Auto Scaling instance → RDS primary; standby promoted on failure.
Country-Level Content Restriction with CloudFront
When report files must be unavailable in specific countries, enforce geography at the content-delivery layer.
Recommended architecture
● Distribute the report files through Amazon CloudFront.
● Enable CloudFront geographic restrictions.
● Configure a deny list for the prohibited countries.
● Keep origin access restricted so users cannot bypass CloudFront.
Design reasoning
● Technical workability: CloudFront evaluates the viewer’s country and blocks or allows delivery according to the configured geographic policy.
● Requirement fit: The same distribution provides low-latency global delivery and country-level restriction.
● Scene fit: The protected resources are report files delivered to worldwide users.
● Engineering common sense: Apply the control where the content is delivered and prevent direct-origin access.
Alternatives to reject
● Route 53 geolocation selects endpoints but is not a direct deny control for individual report files.
● Geoproximity routing shifts traffic based on location and bias; it does not block countries.
● Network ACL country blocking requires maintaining large IP ranges and applies only at subnet boundaries.
● Elastic Beanstalk does not itself provide geographic content restriction.
Workflow: Viewer request → CloudFront determines country → allow or deny → retrieve permitted content from protected origin → cache near allowed users.
CloudFormation EC2 Roles with Instance Profiles
CloudFormation attaches an IAM role to EC2 through an instance profile.
Recommended architecture
● Define an IAM role trusted by EC2.
● Attach DynamoDB permissions to the role.
● Create AWS::IAM::InstanceProfile containing the role.
● Reference the instance profile from the EC2 launch configuration.
Design reasoning
● Technical workability: EC2 receives temporary role credentials through the attached instance profile.
● Requirement fit: The application accesses DynamoDB without long-term keys.
● Scene fit: Infrastructure is deployed with CloudFormation.
● Engineering common sense: Never pass secret access keys as stack parameters or user data.
Alternatives to reject
● User access keys can leak through deployment workflows.
● CloudFormation should not create and distribute IAM user secrets to instances.
● AWS::IAM::InstanceRoleName is not the valid EC2 attachment mechanism described.
● A role without an instance profile is not attached to EC2.
Workflow: Stack creates role → instance profile contains role → EC2 launches → SDK retrieves temporary credentials.
Encrypted Redshift Snapshot Copy Across Regions
A KMS-encrypted Redshift cluster needs destination-key authorization before native cross-Region snapshot copy can protect it.
Recommended architecture
● Create a snapshot copy grant in the destination Region for the destination KMS key.
● Enable Redshift cross-Region snapshot copy.
● Set retention appropriate to recovery objectives.
● Test restoring a snapshot in the recovery Region.
Design reasoning
● Technical workability: The copy grant lets Redshift use the destination key to encrypt copied snapshots.
● Requirement fit: Recovery snapshots survive loss of the source Region.
● Scene fit: The warehouse is encrypted with KMS.
● Engineering common sense: Backup existence is insufficient until restore permissions and timing are tested.
Alternatives to reject
● Custom Lambda copy logic duplicates the native feature and may omit key authorization.
● Enabling copy without the KMS grant fails for the encrypted workflow.
● S3 Cross-Region Replication is not the Redshift snapshot-copy mechanism.
● CloudFormation cannot restore Redshift from an arbitrary replicated snapshot file in S3.
Workflow: Redshift snapshot → destination copy grant → encrypted cross-Region copy → retention → recovery restore test.
Protecting an E-Commerce Trust Chain
Secure public access requires protection of both DNS resolution and the HTTPS path.
Recommended architecture
● Host and register the domain with Route 53.
● Enable DNSSEC signing for the public hosted zone and publish the required delegation records.
● Request a trusted public certificate from AWS Certificate Manager.
● Terminate HTTPS at the Application Load Balancer.
● Configure CloudFront with SNI and redirect HTTP viewers to HTTPS.
Design reasoning
● Technical workability: DNSSEC validates signed DNS responses; ACM and ALB establish trusted TLS; CloudFront enforces secure viewer connections.
● Requirement fit: The design reduces DNS spoofing, HTTPS spoofing, and downgrade exposure.
● Scene fit: It uses the existing Fargate, ALB, CloudFront, and Route 53 architecture.
● Engineering common sense: Use native managed controls instead of operating custom DNS servers or certificate processes.
Alternatives to reject
● Self-hosting BIND adds patching, availability, and key-management work.
● External DNSSEC can work but introduces another provider without necessity.
● Imported certificates can work but add renewal and deployment operations when ACM can issue and renew them.
Workflow: Sign DNS → validate delegation → issue certificate → configure ALB HTTPS → enforce CloudFront HTTPS.
Decoupled Processing with Durable Regional Storage
A single EC2 server with attached EBS cannot scale independently for uploads and processing. Separate object storage, work queuing, and worker capacity.
Recommended architecture
● Store applicant documents and photos in Amazon S3.
● Enable S3 Cross-Region Replication for regional data protection.
● Place verification tasks on Amazon SQS.
● Scale EC2 worker instances according to queue depth.
● Reproduce infrastructure in another Region with CloudFormation.
Design reasoning
● Technical workability: S3 provides durable scalable object storage, SQS buffers work, and Auto Scaling adds parallel processors.
● Requirement fit: Upload growth no longer depends on one EBS volume, and traffic surges do not overwhelm one server.
● Scene fit: Verification tasks can be processed asynchronously after files are accepted.
● Engineering common sense: Decouple ingestion from processing and keep durable data outside worker instances.
Alternatives to reject
● SNS is push notification, not a durable worker backlog with queue-depth scaling.
● Larger Provisioned IOPS EBS improves one volume but keeps the instance-attached storage limitation.
● EBS does not provide shared cross-Region object storage.
● Scaling from the number of SNS notifications lacks durable retry and backlog semantics.
Workflow: Upload to S3 → enqueue task → workers scale → process files → store result → replicate objects and recover stack if needed.
Multi-Account Deployment with StackSets
CloudFormation StackSets centrally deploy consistent infrastructure across accounts and Regions.
Recommended architecture
● Integrate StackSets with AWS Organizations.
● Define the required CloudFormation template once.
● Target organizational units, accounts, and Regions.
● Use controlled permissions, failure tolerances, and rollout order.
Design reasoning
● Technical workability: StackSets creates and updates stack instances in multiple accounts and Regions.
● Requirement fit: Central teams maintain consistent resources with low manual work.
● Scene fit: The deployment scope spans the organization.
● Engineering common sense: Standardize centrally while allowing account-specific parameters only where necessary.
Alternatives to reject
● Nested stacks provide template modularity but do not orchestrate organization-wide deployment.
● Regional parameters and IAM policies still require separate stack operations.
● Control Tower governs landing zones but is not the direct general StackSet deployment mechanism.
● Manual stacks are prone to drift.
Workflow: Update template → choose targets → deploy stack instances → monitor failures → remediate drift.
Forwarding Route 53 DNS Queries to Active Directory
VPC workloads can use Amazon-provided DNS while selected Active Directory names are forwarded to domain controllers.
Recommended architecture
● Create a Route 53 Resolver outbound endpoint.
● Create a forwarding rule for the AD domain, such as private.aws.com.
● Configure both domain-controller IP addresses as targets.
● Associate the rule with the required VPCs.
Design reasoning
● Technical workability: The outbound endpoint sends matching queries from Route 53 Resolver to AD DNS servers.
● Requirement fit: EC2 instances resolve AD names centrally without per-host DNS changes.
● Scene fit: Only one private domain requires forwarding.
● Engineering common sense: Forward the narrowest namespace and provide multiple DNS targets.
Alternatives to reject
● Per-client split-DNS settings create high administration and drift.
● An inbound Resolver endpoint accepts queries into the VPC; it does not forward VPC queries outward.
● A conditional forwarder on another DNS server is not the Route 53 Resolver rule required here.
● Public hosted zones should not expose internal AD records.
Workflow: EC2 query → AmazonProvidedDNS → forwarding rule match → outbound endpoint → AD DNS response.
Extending Active Directory Authentication to AWS
Windows workloads in AWS can use existing corporate credentials through a managed directory trust.
Recommended architecture
● Deploy AWS Managed Microsoft AD with AWS Directory Service.
● Establish the required trust relationship with the on-premises Microsoft Active Directory.
● Configure DNS and network connectivity between the directories.
● Join AWS-hosted Windows resources to the managed domain and enable the required SSO experience.
Design reasoning
● Technical workability: A directory trust allows identities from the corporate forest or domain to authenticate to trusted AWS-hosted resources.
● Requirement fit: Employees retain existing usernames and passwords while directory infrastructure in AWS is managed.
● Scene fit: The requirement is Windows domain extension and single sign-on, not consumer identity.
● Engineering common sense: Use a managed compatible directory rather than synchronizing passwords into application databases.
Alternatives to reject
● Amazon Cognito targets application users and does not extend a Windows domain.
● IAM roles authorize AWS API actions but do not provide domain authentication.
● IAM Identity Center can federate workforce access but does not replace the Windows directory needed for domain joins.
● Custom LDAP servers add availability and patching work.
Workflow: User authenticates to corporate AD → trust is evaluated → AWS Managed Microsoft AD authorizes access → Windows resource session.
Location-Based Mobile Alerts at Scale
Millions of users receiving time-sensitive local offers need durable buffering, scalable processing, fast lookup, and mobile push delivery.
Recommended architecture
● Buffer work in Amazon SQS.
● Scale EC2 workers according to queue depth.
● Store offers and location mappings in DynamoDB.
● Send alerts through Amazon SNS Mobile Push.
Design reasoning
● Technical workability: SQS absorbs bursts, workers process asynchronously, DynamoDB provides scalable lookups, and SNS delivers platform push notifications.
● Requirement fit: The pipeline can deliver alerts within the required short interval without losing burst traffic.
● Scene fit: More than two million mobile users may trigger irregular notification demand.
● Engineering common sense: Separate offer selection from notification delivery and preserve pending work in a queue.
Alternatives to reject
● AWS Device Farm tests applications; it does not deliver production push notifications.
● AWS AppSync provides managed GraphQL and subscriptions but is not the direct bulk mobile-push service.
● Direct synchronous processing lacks buffering during sudden demand.
● Pinpoint can support campaigns, but the SQS, worker, DynamoDB, and SNS chain more directly fits the stated transactional workflow.
Workflow: Location event → SQS → worker selects DynamoDB offer → SNS push → mobile device.
Reducing Serverless Cost by Removing Database Wait
Long Lambda duration caused by network waiting should be corrected at the latency source before compute settings are tuned.
Recommended architecture
● Migrate the on-premises MySQL database to Amazon RDS for MySQL.
● Use Multi-AZ for database availability.
● Enable API Gateway caching for safe, repeatable responses.
● Right-size Lambda memory and timeout after measuring the new execution profile.
● Use DynamoDB Auto Scaling for unpredictable growth.
Design reasoning
● Technical workability: Moving MySQL close to Lambda removes repeated hybrid latency; API caching reduces invocations.
● Requirement fit: Shorter executions and fewer calls directly lower cost.
● Scene fit: The application already scales serverlessly, so replacing Lambda with servers would reverse a working design.
● Engineering common sense: Remove 4.5-minute waiting before optimizing secondary resources.
Alternatives to reject
● Direct Connect may improve latency but is expensive for a single workload and preserves the remote dependency.
● Converting Lambda to EC2 adds capacity management.
● CloudFront caching is less direct than API Gateway stage caching here.
● DAX or ElastiCache adds cost without a stated low-latency cache requirement.
Workflow: Migrate DB → validate → cache API → measure Lambda → tune → monitor cost.
Central Auditor Account with Read-Only Roles
External audit access should use temporary cross-account roles from a dedicated auditor account.
Recommended architecture
● Create a dedicated auditor AWS account.
● Create read-only roles in each target account.
● Trust the auditor account’s approved principal.
● Require STS role assumption and record activity with CloudTrail.
Design reasoning
● Technical workability: STS issues temporary credentials scoped by each target role.
● Requirement fit: The auditor receives least-privilege access with traceable sessions.
● Scene fit: Multiple accounts must be reviewed consistently.
● Engineering common sense: Keep auditor identity, permissions, and session history separate from operational users.
Alternatives to reject
● Long-term IAM user keys increase rotation and exposure risk.
● Creating directory identities in every account adds unnecessary administration.
● Sharing existing employee passwords destroys accountability.
● One broad administrator role exceeds audit needs.
Workflow: Auditor signs in centrally → assumes target read-only role → reviews resources and logs → session expires.
Centralized Egress Inspection
More than 50 accounts should share one managed outbound inspection path instead of duplicating proxy or firewall fleets.
Recommended architecture
● Build a centralized egress VPC.
● Attach workload VPCs through Transit Gateway.
● Route outbound traffic through AWS Network Firewall.
● Use NAT Gateway for internet egress.
● Manage rules and routes centrally.
Design reasoning
● Technical workability: Transit Gateway provides hub routing, Network Firewall performs managed inspection, and NAT Gateway translates outbound traffic.
● Requirement fit: The design scales policy administration across many accounts with lower operations.
● Scene fit: All accounts require consistent outbound filtering.
● Engineering common sense: Preserve symmetric routing through the inspection path.
Alternatives to reject
● Proxy fleets in every account create patching and scaling work.
● A firewall endpoint in every account duplicates cost and policies.
● Central EC2 proxies can work but require capacity and software maintenance.
● Independent NAT paths can bypass centralized inspection.
Workflow: Spoke VPC → Transit Gateway → Network Firewall → NAT Gateway → internet → symmetric return path.
Scaling a Large and Growing MySQL Workload
A continuously growing 16 TiB database and 24-by-7 application need managed storage growth, database availability, and an elastic application tier.
Recommended architecture
● Run the application in an EC2 Auto Scaling group across Availability Zones behind an Application Load Balancer.
● Cover predictable application capacity with Reserved pricing and use elastic capacity for growth.
● Migrate MySQL to Amazon Aurora.
● Add Aurora replicas in another Availability Zone for availability and read scaling.
Design reasoning
● Technical workability: Aurora provides distributed managed storage, automatic growth, replicas, and managed failover.
● Requirement fit: The design supports continued dataset expansion with lower operational effort than self-managed MySQL.
● Scene fit: The Ruby on Rails application remains server based and can scale horizontally.
● Engineering common sense: Choose a database with growth headroom rather than repeatedly extending logical volumes.
Alternatives to reject
● Lambda@Edge is not a general replacement for a Rails processing tier.
● Self-managed source-replica MySQL requires failover, backup, and volume engineering.
● Manual RDS storage increases add operational work and less growth headroom.
● Reserved Instances reduce cost but do not provide database high availability by themselves.
Workflow: Scale application tier → migrate data → validate Aurora → add replicas → cut over → monitor storage and query load.
IAM-Authorized API Gateway Requests
AWS principals can call API Gateway securely through native IAM authorization and Signature Version 4.
Recommended architecture
● Set the API authorization type to AWS_IAM.
● Grant approved users or roles execute-api:Invoke on required stages and methods.
● Sign client requests with SigV4.
● Enable X-Ray when request tracing is required.
Design reasoning
● Technical workability: API Gateway validates the signed request against IAM permissions.
● Requirement fit: Existing IAM identities receive controlled API access without another credential database.
● Scene fit: The callers are AWS users or roles.
● Engineering common sense: Never transmit secret access keys to a custom authorizer.
Alternatives to reject
● CORS controls browser-origin behavior, not authentication.
● Client certificates authenticate API Gateway to selected backend integrations, not callers to the public API.
● Manual key validation duplicates AWS authentication and risks exposing secrets.
● API keys identify usage plans and are not a strong authorization mechanism.
Workflow: Client signs request → API Gateway verifies SigV4 → IAM evaluates execute-api:Invoke → backend runs.
Reducing Daily DynamoDB Costs
Predictable daily throughput benefits from reserved capacity, while historical table data can move to S3 and be deleted.
Recommended architecture
● Purchase reserved capacity for predictable provisioned DynamoDB usage.
● Use one daily table where that existing pattern is required.
● Export completed tables to S3.
● Delete old DynamoDB tables after validation.
Design reasoning
● Technical workability: Reserved capacity discounts predictable throughput, and S3 retains historical data at lower cost.
● Requirement fit: The application keeps current low-latency data while avoiding long-term DynamoDB storage charges.
● Scene fit: Biometric data is partitioned by day and older data is not actively queried.
● Engineering common sense: Verify exports before deleting the source table.
Alternatives to reject
● Redshift is not the low-latency operational database.
● S3 One Zone-IA reduces resilience unnecessarily for reports.
● ElastiCache is not durable primary storage.
● RDS adds fixed relational capacity and management without a relational requirement.
Workflow: Use daily table → export to S3 → verify objects → delete table → retain current provisioned capacity.
Restricting EC2 Egress by URL
Security groups and network ACLs filter addresses and ports, not full package-repository URLs.
Recommended architecture
● Deploy a forward proxy with an allow list for approved update URLs.
● Route EC2 outbound web requests through the proxy.
● Remove direct default internet routes from workload subnets where required.
● Log proxy requests and block every destination not explicitly approved.
Design reasoning
● Technical workability: An application-layer proxy can inspect hostnames and URLs before forwarding HTTPS or HTTP traffic according to policy.
● Requirement fit: Instances can retrieve approved software updates while other outbound destinations remain unavailable.
● Scene fit: The policy is destination-URL based, not merely IP or port based.
● Engineering common sense: Enforce controls at the protocol layer containing the information being evaluated.
Alternatives to reject
● Security groups are allow-only and cannot filter by URL.
● Network ACLs require IP ranges and cannot inspect HTTP paths.
● NAT Gateway provides egress translation but no destination URL policy.
● DNS controls alone do not prevent direct IP access or inspect URL paths.
Workflow: EC2 request → proxy → URL policy → approved repository → logged response; denied destinations stop at proxy.
Modernizing a Tape-Based Media Catalog
A file-oriented media system can move archived content into S3 and add managed facial analysis without replacing its existing interface.
Recommended architecture
● Deploy Storage Gateway File Gateway on premises.
● Let the media asset system copy tape-extracted files through the file share.
● Store the files in Amazon S3.
● Invoke Amazon Rekognition through Lambda.
● Write generated metadata back to the existing catalog.
Design reasoning
● Technical workability: File Gateway preserves file access, while Rekognition analyzes supported S3 media.
● Requirement fit: The design minimizes disruption and ongoing infrastructure management.
● Scene fit: The source is a historical tape archive, not a live stream.
● Engineering common sense: Introduce cloud services behind the interface the existing MAM system already understands.
Alternatives to reject
● Kinesis Video Streams is optimized for live video ingestion.
● Custom SageMaker computer-vision training adds effort when Rekognition already provides the capability.
● SFTP and self-managed EC2 vision software create patching, scaling, and recovery work.
● Replacing the MAM workflow expands scope.
Workflow: Extract tape file → File Gateway → S3 → Lambda → Rekognition → metadata returned to catalog.
Buffering Online Exam Submissions
A simultaneous submission spike needs durable queueing while static exam assets are handled separately.
Recommended architecture
● Place exam submissions in SQS.
● Process them with controlled consumers.
● Store images and diagrams in S3.
● Deliver static assets through CloudFront.
Design reasoning
● Technical workability: SQS absorbs the write burst, while S3 and CloudFront scale read-only exam content.
● Requirement fit: Submission data is preserved without overprovisioning a fixed write target.
● Scene fit: Thousands of candidates submit at the same time.
● Engineering common sense: Static content delivery and write processing need different scaling mechanisms.
Alternatives to reject
● A fixed 10,000 WCU may be too high or too low.
● CloudFront Functions do not host images.
● A 100 percent utilization target leaves no headroom.
● A low maximum capacity can still throttle the burst.
● Moving to RDS is a major redesign, and read replicas do not help writes.
Workflow: Candidate loads assets from CloudFront → submits to queue → consumers persist results safely.
Offloading RDS Batch Reads and Sending Completion Alerts
A batch report should read from replicas and notify subscribers directly when processing finishes.
Recommended architecture
● Add RDS read replicas.
● Direct batch analytics queries to replicas.
● Keep OLTP writes on the Multi-AZ primary.
● Publish completion to Amazon SNS for the on-premises dashboard subscriber.
Design reasoning
● Technical workability: Read replicas reduce primary read load, and SNS pushes notifications to subscribers.
● Requirement fit: The CRM remains responsive during batch work, and the dashboard receives timely completion notice.
● Scene fit: The workload is read-heavy analytics against an operational database.
● Engineering common sense: High availability and read scaling are different database functions.
Alternatives to reject
● Redshift should not replace the CRM OLTP database for this narrow need.
● Redshift Spectrum analyzes S3 data, not the stated RDS batch spike.
● SQS requires consumers to poll and is less direct for broadcast notification.
● Running analytics on the writer preserves the bottleneck.
Workflow: Batch starts → queries replica → report finishes → SNS publishes → dashboard receives notification.
Canary Lambda Deployment with X-Ray
A Lambda release can expose a small percentage of traffic first and trace requests through downstream services.
Recommended architecture
● Use CodeDeploy canary traffic shifting.
● Route 10 percent to the new Lambda version, then shift the remaining 90 percent after five minutes if healthy.
● Enable AWS X-Ray active tracing on the function and supported downstream calls.
● Attach alarms that trigger rollback.
Design reasoning
● Technical workability: Lambda aliases divide traffic between versions; CodeDeploy controls the shift; X-Ray records request segments.
● Requirement fit: The release limits blast radius and provides end-to-end diagnostics.
● Scene fit: The desired rollout is explicitly two-step canary exposure.
● Engineering common sense: Deployment safety requires both controlled traffic and observable success criteria.
Alternatives to reject
● EC2-style rolling deployment does not implement Lambda alias percentages.
● All-at-once deployment removes gradual validation.
● Linear shifting uses repeated equal increments rather than the requested canary pattern.
● AWS Config records configuration and does not trace requests.
Workflow: Publish version → alias sends 10 percent → monitor alarms and traces → shift 90 percent or roll back.
Regional Disaster Recovery with Route 53 Failover
A regional recovery design needs independent application stacks and health-based traffic control.
Recommended architecture
● Deploy an ALB and Auto Scaling group in the primary Region.
● Maintain a recovery ALB and Auto Scaling group in the secondary Region.
● Replicate or restore required data according to RPO.
● Configure Route 53 failover records and health checks.
● Test failover and failback regularly.
Design reasoning
● Technical workability: Each Region can serve traffic independently, and Route 53 changes DNS answers when the primary endpoint is unhealthy.
● Requirement fit: The application remains available after a complete primary-Region outage.
● Scene fit: Production runs in one Region and disaster recovery is required in another.
● Engineering common sense: Application failover succeeds only if data, secrets, certificates, and dependencies also exist in recovery.
Alternatives to reject
● Multi-AZ protects only inside one Region.
● One global load balancer with no secondary compute cannot recover the workload.
● Manual DNS changes increase RTO.
● Backups without a deployable application environment delay restoration.
Workflow: Health check fails → Route 53 returns secondary endpoint → recovery fleet scales → data promoted or restored → users reconnect.
Normalizing Query Strings Before CloudFront Cache Lookup
Equivalent query strings should map to one cache key when parameter order or letter case does not change the response.
Recommended architecture
● Use Lambda@Edge on a viewer request.
● Normalize approved parameter names, values, order, and case.
● Forward the canonical query string to CloudFront cache lookup.
● Preserve parameters that legitimately change content.
Design reasoning
● Technical workability: Viewer-request edge code runs before cache-key evaluation and can rewrite the request.
● Requirement fit: More semantically identical requests become cache hits.
● Scene fit: Clients send equivalent query strings in inconsistent forms.
● Engineering common sense: Canonicalization rules must match application semantics.
Alternatives to reject
● Ignoring all query parameters can serve incorrect content.
● CloudFront has no simple universal case-insensitive normalization switch.
● Origin-side normalization occurs after a cache miss and therefore cannot improve the initial cache lookup effectively.
● Forwarding every unnormalized variation fragments the cache.
Workflow: Viewer request → Lambda@Edge canonicalizes → CloudFront computes cache key → hit or origin fetch.
Enforcing HTTPS on CloudFront
CloudFront can provide trusted HTTPS on its default domain and enforce secure viewer connections.
Recommended configuration
● Use the default CloudFront certificate for the distribution’s cloudfront.net hostname.
● Set Viewer Protocol Policy to Redirect HTTP to HTTPS or HTTPS Only.
● Use ACM in us-east-1 when a custom domain is required.
Design reasoning
● Technical workability: Browsers trust the default CloudFront certificate for the default hostname, and the viewer policy controls accepted protocols.
● Requirement fit: Users reach content through HTTPS, supporting confidentiality and secure indexing.
● Scene fit: The immediate requirement can use the default distribution domain.
● Engineering common sense: Certificate hostname coverage must match the URL users visit.
Alternatives to reject
● Self-signed certificates are not publicly trusted.
● Storing a certificate in S3 does not establish browser trust.
● An ELB certificate is not automatically a CloudFront viewer certificate.
● A supposed generic ELB default certificate cannot secure the application’s custom hostname.
Workflow: Viewer uses HTTP or HTTPS → policy redirects or accepts → TLS terminates at CloudFront → origin request follows configured protocol.
Real-Time Alerts for Public S3 Objects
Compliance can receive immediate notification when an object is uploaded with public-read access.
Recommended architecture
● Enable CloudTrail S3 data events for PutObject.
● Create an EventBridge rule matching public-read ACL requests.
● Publish matching events to an SNS topic.
Design reasoning
● Technical workability: CloudTrail records object-level API parameters, EventBridge filters events, and SNS pushes alerts.
● Requirement fit: Detection is event driven rather than delayed by polling.
● Scene fit: The risk occurs during object upload.
● Engineering common sense: Alert on the authoritative API event closest to the policy violation.
Alternatives to reject
● Hourly polling delays response.
● Systems Manager Automation is unnecessary for direct notification.
● Lex is conversational AI.
● GuardDuty and Trusted Advisor do not identify this ACL event through the proposed workflow.
● CDK deploys infrastructure and is not the runtime detector.
Workflow: Public-read PUT → CloudTrail data event → EventBridge match → SNS compliance alert.
Governed Self-Service SageMaker Environments
Data scientists can receive approved machine-learning environments through a controlled self-service catalog.
Recommended architecture
● Define the SageMaker environment in CloudFormation.
● Encrypt required resources with a customer-controlled KMS key.
● Publish the approved template as an AWS Service Catalog product.
● Expose only mapped parameters that users are allowed to choose.
● Apply portfolio access controls and constraints.
Design reasoning
● Technical workability: Service Catalog provisions CloudFormation products under administrator-defined permissions and constraints.
● Requirement fit: Users launch standardized environments without broad infrastructure permissions.
● Scene fit: Multiple teams need repeatable SageMaker workspaces with governance and encryption.
● Engineering common sense: Separate product design from product consumption.
Alternatives to reject
● Giving every data scientist administrator permissions weakens governance.
● Manual ticket-based provisioning creates avoidable delay and inconsistency.
● Raw CloudFormation access may allow unapproved parameter and resource changes.
● Custom deployment scripts add maintenance where Service Catalog already provides portfolios and constraints.
Workflow: Administrator authors template → publishes catalog product → grants portfolio access → user selects approved parameters → governed stack deploys.
Cost-Effective Direct Connect Backup
A single Direct Connect circuit is a failure point, but the lowest-cost backup does not require purchasing another dedicated circuit.
Recommended architecture
● Establish AWS Site-to-Site VPN connections from the data center to each VPC.
● Terminate each VPN on the appropriate virtual private gateway.
● Use BGP to advertise and select backup routes.
● Prefer Direct Connect during normal operation and fail to VPN when needed.
Design reasoning
● Technical workability: Dynamic routing can withdraw the Direct Connect path and select the available VPN route.
● Requirement fit: The solution adds hybrid connectivity redundancy at lower cost than another Direct Connect connection.
● Scene fit: Ten VPCs already depend on one managed circuit.
● Engineering common sense: Match backup cost and performance to emergency use rather than duplicating full primary capacity automatically.
Alternatives to reject
● MPLS to every VPC is costly and slow to provision.
● A second Direct Connect provides stronger dedicated redundancy but costs more.
● VPN over a newly purchased second Direct Connect still pays for another circuit.
Workflow: Build VPNs → configure BGP preference → test route withdrawal → verify every VPC → monitor tunnel status.
Eliminating NAT Instance Download Timeouts
Long patch downloads can fail when a NAT instance times out and sends a TCP FIN.
Recommended architecture
● Replace the NAT instance with a NAT Gateway.
● Place the NAT Gateway in a public subnet with an Elastic IP.
● Update private-subnet default routes.
● Monitor gateway connections, errors, and throughput.
Design reasoning
● Technical workability: NAT Gateway provides managed scaling and connection handling suited to outbound package downloads.
● Requirement fit: Patch downloads become more reliable without maintaining NAT servers.
● Scene fit: Internet connectivity already exists, but the NAT instance interrupts long flows.
● Engineering common sense: Fix the failing egress component rather than unrelated placement or VPN settings.
Alternatives to reject
● An internet gateway is already implied by working NAT egress.
● Placement groups affect east-west EC2 latency, not internet patch downloads.
● A virtual private gateway serves VPN or Direct Connect, not public repositories.
● Increasing NAT instance size preserves timeout and administration risks.
Workflow: Private instance → route to NAT Gateway → internet gateway → repository → long download returns through NAT.
Regional ACM Certificates for Multi-Region ALBs
Application Load Balancers and the ACM certificates attached to them are regional resources.
Recommended architecture
● Request or import the required certificate in every application Region.
● Validate each fully qualified domain name.
● Attach the regional certificate to the Application Load Balancer in the same Region.
● Automate certificate and listener deployment with infrastructure as code.
Design reasoning
● Technical workability: An ALB can use an ACM certificate only when the certificate exists in the ALB’s Region.
● Requirement fit: Every regional HTTPS endpoint continues to present a trusted certificate.
● Scene fit: The application is expanding from one Region to several independent regional stacks.
● Engineering common sense: Replicate regional dependencies together rather than assuming one regional resource is global.
Alternatives to reject
● A certificate created in one Region cannot be attached to ALBs in other Regions.
● AWS KMS manages cryptographic keys but does not issue public website certificates.
● Reusing one regional ALB certificate does not provide regional independence.
Workflow: Define FQDNs → request ACM certificate per Region → complete DNS validation → attach to regional ALB listener → test TLS → route users to regional endpoints.
How a Zonal Reserved Instance Discount Applies
A zonal Reserved Instance discount applies to running instances that match its attributes, including Availability Zone.
Example outcome
● A reservation for r4.16xlarge in us-west-2a matches the DEV instance in that AZ.
● A similar UAT instance in another AZ does not match the zonal reservation.
● Under consolidated billing, eligible discounts can be shared when attributes match.
Design reasoning
● Technical workability: Billing evaluates instance family, size, platform, tenancy, Region, and zonal scope.
● Requirement fit: The organization can identify which account receives the discount.
● Scene fit: Accounts run similar instances in different AZs.
● Engineering common sense: Reservations are billing constructs; verify exact matching dimensions before purchase.
Alternatives to reject
● The discount is not limited to the purchaser when sharing is enabled.
● A different Availability Zone does not match a zonal RI.
● Both accounts cannot consume one zonal benefit when only one workload matches.
● An unused matching reservation does not attach manually to a named instance.
Workflow: Instance runs → consolidated billing searches matching RI → matching DEV usage receives discount.
Faster Global Uploads to S3
Users far from an S3 Region can upload through nearby AWS edge locations using S3 Transfer Acceleration.
Recommended architecture
● Enable Transfer Acceleration on the destination bucket.
● Update upload clients to use the s3-accelerate endpoint.
● Test performance from major user locations.
● Retain multipart upload for large objects.
Design reasoning
● Technical workability: The accelerated endpoint receives data at an AWS edge location and carries it over the AWS global network to the bucket Region.
● Requirement fit: European artists gain faster transfers without moving or duplicating the bucket.
● Scene fit: The latency is caused by long-distance internet paths to the current S3 Region.
● Engineering common sense: Validate acceleration benefit because performance and cost depend on source location and network conditions.
Alternatives to reject
● CloudFront is designed primarily for content delivery and is not the direct general upload accelerator here.
● Creating EC2 upload proxies adds scaling, failure, and data-handling work.
● S3 Cross-Region Replication copies objects after upload and does not accelerate the initial transfer.
● Changing storage class does not improve upload network latency.
Workflow: Client selects accelerated endpoint → nearest edge receives upload → AWS backbone transports object → S3 stores it.
Tools for Migration Discovery and Readiness
Migration planning should combine portfolio tracking, dependency discovery, and business-case analysis.
Recommended tools
● Use AWS Migration Hub to track applications and migration progress.
● Use AWS Application Discovery Service to collect server configuration, utilization, and dependency data.
● Use Cloud Adoption Readiness Tool to assess organizational readiness and identify capability gaps.
Design reasoning
● Technical workability: Discovery Service gathers estate data, Migration Hub centralizes visibility, and CART structures readiness assessment.
● Requirement fit: Together they support planning before workload movement begins.
● Scene fit: The organization needs portfolio-level understanding, not only one server transfer.
● Engineering common sense: Discover dependencies before sequencing migration waves.
Alternatives to reject
● Application Migration Service performs server migration but does not replace organizational readiness assessment.
● Database Migration Service focuses on database data movement.
● CloudFormation deploys infrastructure but does not discover the on-premises estate.
● Cost Explorer analyzes AWS spending after or during adoption and is not the primary migration discovery tool.
Workflow: Assess readiness → discover servers and dependencies → group applications → build waves → track execution in Migration Hub.
Oracle RAC Backups on EC2
Oracle RAC must remain on EC2 because Amazon RDS for Oracle does not provide RAC.
Recommended architecture
● Run the supported Oracle RAC cluster on EC2.
● Store database volumes on the required EBS architecture.
● Use Amazon Data Lifecycle Manager for scheduled EBS snapshots.
● Coordinate crash-consistent or application-consistent backup procedures.
Design reasoning
● Technical workability: EC2 preserves RAC control, while DLM automates repeatable snapshot policies.
● Requirement fit: The design reduces backup administration without changing the database architecture.
● Scene fit: RAC compatibility is mandatory.
● Engineering common sense: Snapshot orchestration must account for all cluster volumes and database consistency.
Alternatives to reject
● RDS Oracle RAC is not a supported service configuration.
● RDS Multi-AZ does not turn RDS into RAC.
● Manual shell snapshots are brittle and difficult to audit.
● Migrating to another engine expands scope and may break application compatibility.
Workflow: Quiesce or coordinate database → DLM snapshot policy → verify all volumes → test restore → retain by policy.
Layered Caching for Game Performance
Slow game asset loading and repeated dynamic lookups require two different caching layers.
Recommended architecture
● Deliver static game assets through Amazon CloudFront.
● Cache frequently accessed dynamic data in Amazon ElastiCache.
● Keep durable source data in its existing authoritative store.
Design reasoning
● Technical workability: CloudFront caches objects near global players; ElastiCache provides low-latency in-memory access inside the application architecture.
● Requirement fit: Both static download time and repeated backend reads improve.
● Scene fit: Game files and dynamic application data have different access patterns.
● Engineering common sense: Cache content at the layer closest to its consumers.
Alternatives to reject
● DynamoDB is a durable database, not a direct in-memory cache substitute.
● Using ElastiCache for static global file delivery lacks edge distribution.
● Using CloudFront as the primary cache for private rapidly changing backend records reverses the correct service roles.
● Selecting an unsupported or unsuitable ElastiCache engine makes an otherwise plausible design unworkable.
Workflow: Static request → CloudFront edge → origin on miss; dynamic request → application → ElastiCache → database on miss.
Stateless Scaling for a Hospital Application
A highly available hospital portal should scale application capacity and database reads without keeping user state on individual servers.
Recommended architecture
● Run stateless web and application tiers in Auto Scaling groups across Availability Zones.
● Use Elastic Load Balancing for request distribution.
● Store shared session state in Amazon ElastiCache.
● Use Amazon RDS with read replicas for read-heavy database access.
● Monitor capacity and health with CloudWatch.
Design reasoning
● Technical workability: Any application instance can serve a request because session state is externalized.
● Requirement fit: Auto Scaling adds capacity, while read replicas reduce writer load.
● Scene fit: A hospital information hub needs continuous availability and predictable response under growth.
● Engineering common sense: Replaceable compute must not own unique user state.
Alternatives to reject
● Stateful application servers make load balancing and replacement unreliable.
● RDS Multi-AZ improves database availability but does not scale read throughput.
● Increasing one instance size creates a larger failure domain.
● Local caches on each server can diverge and disappear during replacement.
Workflow: Request → load balancer → stateless instance → shared cache → writer or read replica.
Secure Third-Party Cross-Account Access
Third-party access should use temporary role credentials and an External ID to reduce confused-deputy risk.
Recommended architecture
● Create an IAM role in the resource-owning account.
● Grant only the required resource actions.
● Trust the provider’s AWS account in the role trust policy.
● Require a unique External ID supplied by the customer.
● Let the provider call STS AssumeRole.
Design reasoning
● Technical workability: STS issues temporary credentials only when the trusted principal and External ID conditions match.
● Requirement fit: The provider receives limited, revocable access without customer-created long-lived users.
● Scene fit: The external organization manages resources for multiple customers.
● Engineering common sense: Put permissions in the customer account and require a customer-specific trust signal.
Alternatives to reject
● IAM users create long-lived credentials and difficult offboarding.
● Trusting only the provider account can expose the role to confused-deputy scenarios.
● Resource policies do not provide general multi-service administration.
● Sharing access keys through Secrets Manager does not convert them into temporary sessions.
Workflow: Provider requests role with External ID → trust policy evaluates → STS returns temporary credentials → scoped actions occur.
Modernizing WebSphere, DB2, and IBM MQ
A managed replatform can preserve key application interfaces while reducing database and message-broker operations.
Recommended architecture
● Use SCT and DMS to migrate DB2 to Aurora.
● Run WebSphere on EC2 Auto Scaling behind an ELB.
● Replace self-managed IBM MQ infrastructure with Amazon MQ where compatibility permits.
Design reasoning
● Technical workability: SCT converts schema, DMS moves data, EC2 preserves WebSphere, and Amazon MQ supports managed broker compatibility.
● Requirement fit: Availability improves and operations decline without rewriting the whole application.
● Scene fit: Different tiers require different migration strategies.
● Engineering common sense: Modernize managed equivalents while preserving protocols that would be costly to rewrite.
Alternatives to reject
● Self-managing IBM MQ on EC2 preserves licensing and operations.
● Rehosting every IBM product gains little cloud benefit.
● Replacing MQ with SQS may require application changes.
● Keeping DB2 self-managed misses managed database savings.
Workflow: Convert schema → migrate data → deploy WebSphere fleet → configure Amazon MQ → test → cut over.
Burst-Ready Television Voting
A live voting platform needs global delivery, public authentication, durable burst buffering, and scalable vote storage.
Recommended architecture
● Use CloudFront before an ALB and EC2 Auto Scaling group.
● Authenticate viewers with Amazon Cognito.
● Place submitted votes in Amazon SQS.
● Process queued votes into DynamoDB.
Design reasoning
● Technical workability: CloudFront and Auto Scaling absorb viewing traffic, SQS buffers the voting spike, and DynamoDB scales writes.
● Requirement fit: Votes are preserved even when processing temporarily lags.
● Scene fit: Public viewers create a short, extreme write burst.
● Engineering common sense: Decouple vote acceptance from counting.
Alternatives to reject
● S3 static hosting cannot run the dynamic voting service.
● Generic IAM role assumption is not public-user authentication.
● SAML targets workforce federation, not viewers.
● RDS remains tightly coupled to sudden write volume.
● IAM does not directly authenticate millions of public users.
Workflow: Viewer → CloudFront → Cognito → ALB and application → SQS → workers → DynamoDB results.
Layered Web DDoS Mitigation
DDoS resilience combines exposure reduction, edge capacity, application filtering, and managed response.
Recommended architecture
● Restrict unnecessary ports with network ACLs and security groups.
● Place CloudFront before public web origins.
● Use AWS WAF to block malicious web patterns.
● Enable Shield Advanced when enhanced DDoS protection is required.
Design reasoning
● Technical workability: Network controls reduce attack surface, CloudFront absorbs edge traffic, WAF filters Layer 7 requests, and Shield adds infrastructure protection.
● Requirement fit: The controls cover complementary attack paths.
● Scene fit: The workload is an internet-facing web application.
● Engineering common sense: Scaling helps availability but does not replace filtering.
Alternatives to reject
● MFA, Config, Trusted Advisor, and Fraud Detector do not directly stop DDoS traffic.
● Larger instances can still be overwhelmed and increase attack cost.
● Session Manager is an administration tool.
● S3 versioning and operating-system patching improve durability and hygiene, not active traffic mitigation.
Workflow: Internet → Shield and CloudFront → WAF → allowed ports → load balancer and application.
Developer Access with PowerUserAccess
Application developers often need broad service creation permissions without the ability to administer IAM or the organization.
Recommended control
● Assign the AWS managed PowerUserAccess policy when its scope matches the development role.
● Add narrower denies or permission boundaries for company-specific guardrails.
● Keep IAM and Organizations administration with platform or security teams.
Design reasoning
● Technical workability: PowerUserAccess permits work with most AWS services while withholding broad identity administration.
● Requirement fit: Developers can build and configure applications without full account control.
● Scene fit: The team needs development autonomy, not operational or security administration.
● Engineering common sense: Managed policies are starting points; review their current permissions before production use.
Alternatives to reject
● AdministratorAccess is unrestricted and overprivileged.
● Wrapping AdministratorAccess in a role does not reduce its permissions.
● SystemAdministrator is oriented to operational administration and is less aligned with application development.
● Granting ad hoc IAM privileges can allow escalation.
Workflow: Developer assumes role → creates application resources → IAM administration remains denied → CloudTrail records actions.
Buffering DynamoDB Write Bursts
When DynamoDB throttles during bursts, durable queuing protects requests until consumers can write at an accepted rate.
Recommended architecture
● Put write requests into Amazon SQS.
● Process messages with controlled consumers.
● Use batch writes, retries, idempotency, and dead-letter handling.
● Scale DynamoDB capacity when sustained demand justifies it.
Design reasoning
● Technical workability: SQS retains messages during throttling or scaling delay.
● Requirement fit: Unexpected bursts do not cause data loss.
● Scene fit: The write spike is variable rather than a permanent data-model change.
● Engineering common sense: Capacity and buffering solve different problems and can complement each other.
Alternatives to reject
● More WCU alone does not durably protect requests during sudden bursts.
● Global tables provide regional replication, not local write buffering.
● Additional tables fragment the data model and retry logic.
● Immediate client retries can worsen throttling.
Workflow: Producer → SQS → consumer batch → DynamoDB → retry throttled items → DLQ persistent failures.
Scaling Daily Analytics Reports with S3 and CloudFront
Frequently read reports should be generated from the database, then served as durable static artifacts instead of repeatedly querying the database.
Recommended architecture
● Maintain source data in Multi-AZ Amazon RDS for MySQL.
● Use read replicas for report-generation workloads where appropriate.
● Generate reports in a scheduled batch process.
● Store completed reports in Amazon S3.
● Cache and distribute them with CloudFront using a daily TTL.
Design reasoning
● Technical workability: S3 durably stores report files, while CloudFront absorbs millions of global read requests.
● Requirement fit: Daily refresh aligns cache expiration with report production and reduces database cost.
● Scene fit: Statistical reports are read far more often than they change.
● Engineering common sense: Compute once, store once, and distribute many times.
Alternatives to reject
● Rebuilding and deleting DynamoDB tables daily adds unnecessary operations.
● ElastiCache is not durable report storage.
● Serving tens of millions of report queries directly from RDS is less scalable and more expensive.
● Read replicas protect the writer from report generation but should not become the public report-delivery tier.
Workflow: Update source data → generate report → write to S3 → invalidate or expire cache → serve through CloudFront.
Replatforming to Managed Application and Database Services
Replatforming changes the hosting platform while preserving the application’s core architecture and behavior.
Recommended migration
● Move the relational database to Amazon RDS.
● Deploy the application through AWS Elastic Beanstalk.
● Preserve application logic and database semantics.
● Replace server administration with managed deployment, scaling, backup, and failover capabilities.
Design reasoning
● Technical workability: Elastic Beanstalk runs supported application stacks, while RDS provides a managed relational engine.
● Requirement fit: The approach reduces operations and cost without a full redesign.
● Scene fit: The application can use AWS-managed equivalents with limited code change.
● Engineering common sense: Adopt managed services where they directly replace existing layers.
Alternatives to reject
● Rehosting copies servers to EC2 and retains most operating-system and database administration.
● Refactoring changes application architecture and code beyond the stated need.
● Repurchasing replaces the application with a different product and may not preserve required behavior.
● Calling the approach rehosting ignores the operational shift to managed platform services.
Workflow: Assess compatibility → migrate database to RDS → deploy application to Elastic Beanstalk → test → cut over → retire source systems.
Capturing Full Packets with Traffic Mirroring
Security analysis that requires packet payloads needs full packet copies, not connection metadata.
Recommended architecture
● Enable VPC Traffic Mirroring on selected EC2 elastic network interfaces.
● Send mirrored traffic to an inspection appliance or supported monitoring target.
● Filter sessions to capture only required protocols and sources.
Design reasoning
● Technical workability: Traffic Mirroring copies packet headers and payloads from ENIs.
● Requirement fit: Analysts can inspect complete network conversations.
● Scene fit: The requirement covers instance network traffic beyond HTTP requests.
● Engineering common sense: Limit capture scope because full packets consume bandwidth and may contain sensitive data.
Alternatives to reject
● VPC Flow Logs contain addresses, ports, actions, and metadata, not payloads.
● ALB access logs contain request metadata rather than complete packets.
● AppFlow is not a packet-analysis destination.
● WAF logs describe inspected web requests but do not capture all instance traffic.
Workflow: Instance ENI → mirror session and filter → mirror target → packet inspection → secured evidence retention.
Isolated Compliance Release with Gradual Traffic Shift
A new compliance boundary should be deployed as a parallel environment, not forced into the existing production stack.
Recommended architecture
● Build a separate OpsWorks stack for the compliant application version.
● Validate the new environment independently.
● Use Route 53 weighted routing to send a small percentage of production traffic to it.
● Increase traffic gradually and retain the original stack for rollback.
Design reasoning
● Technical workability: Parallel stacks can run different versions while DNS weights control exposure.
● Requirement fit: Gradual cutover protects millions of users and supports fast rollback.
● Scene fit: The existing platform already uses OpsWorks and EC2, so the release stays within the same operating model.
● Engineering common sense: Isolate regulatory changes and reduce blast radius during validation.
Alternatives to reject
● In-place upgrades expose all users and weaken rollback.
● Sending all traffic to the new stack immediately removes staged validation.
● Introducing Lambda into an OpsWorks release changes the platform without solving the stated isolation need.
Workflow: Build → compliance test → functional test → low traffic weight → observe → increase weight → retire old stack after acceptance.
Managed CI/CD Tests, Alerts, and Feature Switches
A managed delivery pipeline should run repeatable tests, alert failures, and deploy infrastructure-defined feature switches.
Recommended architecture
● Use CodePipeline to orchestrate stages.
● Use CodeBuild for tests and security scans.
● Use EventBridge and SNS for failure notifications.
● Model feature switches and infrastructure with AWS CDK.
Design reasoning
● Technical workability: CDK synthesizes deployment templates, CodeBuild executes arbitrary test commands, and CodePipeline coordinates promotion.
● Requirement fit: The workflow is automated and reduces server administration.
● Scene fit: Builds include custom tests and infrastructure changes.
● Engineering common sense: Keep feature configuration versioned with deployment code.
Alternatives to reject
● Lambda is not the best general full-build environment.
● Amplify plugins do not provide universal infrastructure feature toggles.
● Jenkins can work but adds server management.
● SES is unnecessary when SNS handles operational alerts.
● CodeArtifact stores packages but does not run all tests and scans.
Workflow: Commit → CodePipeline → CodeBuild tests → CDK deploy → EventBridge and SNS report failures.
Replacing NAT Instances with NAT Gateway
Private subnets need reliable outbound internet access without maintaining a custom NAT server fleet.
Recommended architecture
● Create a NAT Gateway in a public subnet.
● Associate an Elastic IP address.
● Ensure the public subnet routes to an internet gateway.
● Point private-subnet default routes to the NAT Gateway.
● Use one NAT Gateway per Availability Zone when AZ independence is required.
Design reasoning
● Technical workability: NAT Gateway performs managed source translation for outbound connections and return traffic.
● Requirement fit: It provides higher availability and bandwidth with less administration than a NAT instance.
● Scene fit: The existing NAT instance is unreliable and constrains throughput.
● Engineering common sense: Keep private workloads without public IPs and place internet-facing egress in public subnets.
Alternatives to reject
● Increasing NAT instance size preserves patching and failover responsibilities.
● A NAT Gateway in a private subnet cannot reach an internet gateway.
● Internet Gateway routes do not make private-addressed instances directly internet reachable.
● VPC peering does not provide internet egress.
Workflow: Private instance → default route → NAT Gateway → internet gateway → destination → return through NAT.
Read-Only Auditor Access to AWS Activity
External auditors need verifiable account activity and controlled read access, not administrative permissions.
Recommended architecture
● Enable AWS CloudTrail for the required Regions and services.
● Store trail logs in a protected Amazon S3 bucket.
● Create a dedicated IAM identity with read-only access to the required logs and resources.
● Provide credentials through the organization’s approved secure process.
Design reasoning
● Technical workability: CloudTrail records console, CLI, SDK, and API activity; IAM controls what the auditor can inspect.
● Requirement fit: The auditor receives evidence without modification rights.
● Scene fit: The task is account auditing, not operational administration.
● Engineering common sense: Separate audit identities from employee accounts and grant only the evidence needed.
Alternatives to reject
● AWS does not independently provide an outside auditor access to customer accounts.
● Emailing logs creates duplication, weak access control, and poor searchability.
● SNS notifications report events but do not provide complete historical audit evidence.
● Broad roles or administrator permissions exceed the auditor’s need.
Workflow: Account action → CloudTrail log → protected S3 storage → read-only auditor access → review and evidence collection.
Parallel Photo Metadata Orchestration
A photo list can be distributed across parallel serverless branches and joined before final metadata is written.
Recommended architecture
● Use Step Functions distributed processing for the photo list.
● Invoke specialized Lambda functions in parallel.
● Track each branch and handle retries.
● Combine results after all required branches complete.
Design reasoning
● Technical workability: Step Functions coordinates parallel Lambda execution and preserves workflow state.
● Requirement fit: Processing scales while producing one coordinated completion result.
● Scene fit: Each photo requires several independent metadata operations.
● Engineering common sense: Use an orchestrator when downstream work must join before completion.
Alternatives to reject
● Independent SQS-triggered functions do not automatically create one coordinated aggregate result.
● A custom list function plus queue leaves orchestration incomplete.
● Lambda functions cannot be attached to an AWS Batch compute environment as proposed.
● One serial Lambda wastes parallelism and risks timeout.
Workflow: Photo list → distributed map → parallel Lambda branches → retry failures → join results → store combined metadata.
Rotating S3 Credentials with an EC2 Role
EC2 applications should obtain short-lived S3 credentials from an attached IAM role.
Recommended architecture
● Create an IAM role trusted by EC2.
● Grant only required bucket listing and object-upload actions.
● Attach the role through an instance profile.
● Let the AWS SDK retrieve credentials from Instance Metadata Service.
Design reasoning
● Technical workability: The platform supplies and rotates temporary credentials automatically.
● Requirement fit: The application lists and uploads objects without storing secrets.
● Scene fit: The workload runs on EC2 and calls S3 APIs.
● Engineering common sense: Grant both the exact bucket-level and object-level actions required.
Alternatives to reject
● Static access keys on EC2 create long-term exposure.
● List-only permission cannot upload objects.
● User data is not a secure rotating credential provider.
● IAM users cannot be attached to EC2 instances, and metadata does not expose user credentials.
Workflow: Instance starts with role → SDK obtains temporary session → signs S3 request → credentials rotate before expiration.
Private Employee Access with Client VPN
Individual employees can reach a private application through an authenticated AWS Client VPN connection.
Recommended architecture
● Keep application servers in private subnets.
● Create an AWS Client VPN endpoint using certificate or approved identity authentication.
● Authorize employee networks and application routes.
● Apply security groups to limit reachable services.
Design reasoning
● Technical workability: Client VPN provides encrypted remote access from individual devices into the VPC.
● Requirement fit: The application is not exposed publicly and authorized employees can connect from the internet.
● Scene fit: Users are individual remote employees, not an entire branch network.
● Engineering common sense: Private subnet placement and user authentication solve different layers of access control.
Alternatives to reject
● Direct Connect is costly for temporary individual access.
● Site-to-Site VPN connects networks, not roaming users.
● Public subnets expose servers unnecessarily.
● A public load balancer alone does not restrict access to employees.
Workflow: Employee authenticates → encrypted Client VPN tunnel → authorized route → application security group → private server.
Immediate Response to Unauthorized IAM Users
CloudTrail events can trigger an automated approval and remediation workflow when a new IAM user is created.
Recommended architecture
● Match CloudTrail CreateUser events with EventBridge.
● Invoke Step Functions for approval and remediation.
● Remove or restrict permissions when approval is absent.
● Notify Security through SNS.
Design reasoning
● Technical workability: CloudTrail records the API, EventBridge reacts in near real time, and Step Functions coordinates actions.
● Requirement fit: Unauthorized users are contained quickly and Security is informed.
● Scene fit: The control focuses on one account-changing API event.
● Engineering common sense: Preserve event details before modifying the identity.
Alternatives to reject
● Audit Manager gathers evidence and is not immediate remediation.
● CloudTrail records events but does not independently filter and notify.
● Fargate adds unnecessary container startup and operations.
● Notification without permission removal leaves the risk active.
Workflow: CreateUser → CloudTrail → EventBridge → Step Functions → restrict user → SNS alert.
Non-Exportable TLS Keys and Durable Security Logs
Sensitive private keys should remain inside dedicated cryptographic hardware, while logs should be stored independently from compute instances.
Recommended architecture
● Use TCP load balancing so TLS operations reach the application tier without terminating at the load balancer.
● Use AWS CloudHSM for private-key operations.
● Deploy HSM capacity across two Availability Zones.
● Deliver encrypted application logs to a private Amazon S3 bucket.
Design reasoning
● Technical workability: CloudHSM performs cryptographic operations without exporting private-key material; S3 provides durable encrypted storage.
● Requirement fit: Multi-AZ HSM deployment supports availability, and IAM plus encryption controls log access.
● Scene fit: Traffic is predictable, so key custody and durability matter more than aggressive elasticity.
● Engineering common sense: Never store sensitive logs on ephemeral instance storage or copy private keys onto web servers.
Alternatives to reject
● TLS offload can work, but instance-store logs are not durable.
● A single HSM location creates an availability gap.
● Retrieving a key from S3 makes the key movable and exposes it to server processes.
Workflow: TCP pass-through → HSM-backed TLS → application → encrypted S3 logging → authorized audit access.
Immediate Sale Offload with CloudFront
A short preparation window favors edge caching over a full application and database migration.
Recommended architecture
● Place CloudFront before the on-premises website.
● Cache static and safely reusable content.
● Tune TTLs for the promotion period.
● Preserve dynamic gaming-store transactions at the existing origin.
Design reasoning
● Technical workability: CloudFront can use the on-premises site as an origin and absorb repeated global requests.
● Requirement fit: The solution reduces origin traffic quickly with minimal application change.
● Scene fit: The sale is imminent, and dynamic transactions must continue working.
● Engineering common sense: Address the immediate bottleneck rather than begin a risky full migration.
Alternatives to reject
● S3 static failover cannot preserve dynamic transactions.
● A hybrid ALB design requires private connectivity and duplicated servers.
● Full VM and Oracle migration is too large for the timeline.
● Database conversion during a sale introduces avoidable risk.
Workflow: User → CloudFront → edge hit for cacheable content → on-premises origin only on miss or dynamic request.
Private Remote Access to ERP
Roaming employees can reach private ERP servers through an authenticated SSL client VPN.
Recommended architecture
● Keep ERP servers in private subnets.
● Deploy an AWS Client VPN endpoint.
● Install client software on authorized devices.
● Configure identity authentication, routes, and restrictive security groups.
Design reasoning
● Technical workability: Client VPN creates encrypted user-to-VPC tunnels.
● Requirement fit: Managers and analysts connect remotely without exposing ERP servers publicly.
● Scene fit: Users work from changing home or travel networks.
● Engineering common sense: Network-to-network solutions are not ideal for individual roaming clients.
Alternatives to reject
● Site-to-Site VPN connects fixed networks, not individual employees.
● Direct Connect serves corporate sites and not roaming users directly.
● Public application servers increase exposure.
● HTTPS on a public ELB encrypts traffic but does not alone limit access to authorized staff.
Workflow: User authenticates → client VPN tunnel → authorized VPC route → ERP security group → private server.
Blocking Hostile Source Networks with NACLs
A network ACL can explicitly deny traffic from known attacking CIDR ranges at the subnet boundary.
Recommended control
● Add a numbered deny rule for the hostile source CIDR.
● Place the rule before broader allow entries.
● Apply it to the affected public subnets.
● Preserve required return-path rules because NACLs are stateless.
Design reasoning
● Technical workability: NACLs support explicit allow and deny rules.
● Requirement fit: Traffic from the identified source is blocked before reaching instances.
● Scene fit: The portal must remain public for other users.
● Engineering common sense: Treat this as immediate containment, not complete DDoS protection.
Alternatives to reject
● Moving the whole portal to a private subnet removes public availability without another ingress design.
● Route tables do not filter by source IP.
● Security groups are allow-only and cannot create an explicit deny.
● Host firewalls require per-instance management.
Workflow: Packet enters subnet → NACL evaluates lowest matching rule → hostile CIDR denied → other users continue.
Buffering Eventually Consistent On-Premises Writes
An eventually consistent BASE database can receive cloud-originated writes asynchronously through a durable queue.
Recommended architecture
● Place write requests in Amazon SQS.
● Run a consumer with connectivity to the on-premises database.
● Process messages idempotently and delete them only after successful persistence.
● Configure retries and a dead-letter queue.
Design reasoning
● Technical workability: SQS durably buffers messages when the database or network is slow or unavailable.
● Requirement fit: Producers remain responsive and the database converges asynchronously.
● Scene fit: Strict immediate consistency is not required.
● Engineering common sense: Preserve ordering only where business keys require it; prioritize durability and idempotency.
Alternatives to reject
● EventBridge does not replicate arbitrary database writes by itself.
● Elastic Transcoder is unrelated to databases.
● S3DistCp copies S3 and HDFS data, not application transactions.
● DynamoDB plus an EMR batch job adds unnecessary storage and delay.
Workflow: Application submits write → SQS retains → consumer writes on premises → acknowledges success → failed messages retry or enter DLQ.
Importing External HSM Key Material into KMS
Customer-generated HSM key material can be imported into a KMS key whose origin is external.
Recommended architecture
● Create a KMS key with EXTERNAL origin.
● Import key material generated by the approved on-premises HSM.
● Configure an S3 bucket policy that requires SSE-KMS with that key.
● Monitor expiration and reimport requirements for the key material.
Design reasoning
● Technical workability: KMS external-origin keys accept imported key material and can encrypt S3 objects.
● Requirement fit: The organization retains control over the original key generation process.
● Scene fit: Existing HSM-generated material must be used rather than replaced.
● Engineering common sense: Losing imported key material can make ciphertext unrecoverable, so preserve it securely.
Alternatives to reject
● Direct Connect does not import encryption keys.
● An AWS_KMS origin key cannot have its managed key material overwritten.
● A CloudHSM custom key store creates a new AWS HSM-backed system rather than importing the existing key as requested.
● SSE-S3 does not use the customer’s imported key.
Workflow: Generate material in HSM → create external-origin KMS key → import securely → enforce key in S3 policy.
Serverless Order and Shipment Workflow
A logistics workflow needs durable order state, explicit process orchestration, and event-driven shipment updates.
Recommended architecture
● Store orders and status in Amazon DynamoDB.
● Model processing steps with AWS Step Functions.
● Invoke Lambda functions for validation, state changes, and external service calls.
● Trigger a Lambda update when shipment scans or delivery events arrive.
Design reasoning
● Technical workability: Step Functions manages state, retries, branching, and service integration; DynamoDB provides scalable order records.
● Requirement fit: The workflow is serverless, event driven, and has low operational overhead.
● Scene fit: Orders move through defined stages and shipment events update their status asynchronously.
● Engineering common sense: Keep durable business state in a database and workflow progress in an orchestrator, not in compute memory.
Alternatives to reject
● AWS Batch is designed for batch compute, not interactive order-state orchestration.
● EFS is file storage and does not model workflow state.
● SQS can decouple steps but does not by itself express branching, retries, and end-to-end state.
● Notification services do not replace transactional order storage.
Workflow: Create order → write DynamoDB item → start state machine → execute steps → receive shipment event → Lambda updates status.
Patch Enforcement and Approved AMI Monitoring
Security patching and AMI compliance are separate controls: one changes instance state, while the other detects configuration drift.
Recommended architecture
● Define approved Windows patches with a Systems Manager Patch Manager baseline.
● Schedule or run patch operations across managed instances.
● Use an AWS Config managed rule to evaluate whether running EC2 instances use approved AMIs.
● Send notifications when noncompliant instances are detected.
Design reasoning
● Technical workability: Patch Manager installs approved updates; AWS Config evaluates resource configuration without blocking launches.
● Requirement fit: Instances receive current security fixes, and developers remain free to launch while violations are reported.
● Scene fit: The organization wants monitoring rather than preventive denial for AMI selection.
● Engineering common sense: Use preventive controls only when blocking is acceptable; otherwise detect and notify quickly.
Alternatives to reject
● GuardDuty detects suspicious behavior, not patch or approved-AMI compliance.
● IAM denial would impede developers, violating the operating requirement.
● Shield Advanced protects against DDoS attacks and does not patch operating systems or assess AMIs.
Workflow: Approve patches → deploy patches → evaluate AMI IDs → mark compliance → notify → remediate through a controlled replacement process.
Continuous VM Replication with AWS MGN
Server migration with minimal downtime requires continuous block-level replication rather than repeated full image imports.
Recommended architecture
● Install the AWS Application Migration Service replication agent on each source server.
● Continuously replicate root and attached data volumes.
● Launch test instances from the replicated state.
● Perform a final cutover after validation.
Design reasoning
● Technical workability: AWS MGN replicates complete server disks and launches equivalent EC2 instances.
● Requirement fit: Ongoing synchronization keeps the cutover state current and reduces downtime.
● Scene fit: Terabyte-scale root and data volumes belong in one coordinated migration workflow.
● Engineering common sense: Test the same replication stream that will be used for production cutover.
Alternatives to reject
● VM Import/Export is useful for image import, but repeated full imports are inefficient for continuous synchronization.
● Discovery Service and Migration Hub support planning and tracking, not block replication.
● Importing data-volume snapshots separately adds coordination, attachment, and consistency risk when MGN can replicate them.
Workflow: Install agent → replicate → test launch → remediate → final sync → cut over → verify application and data.
Cost-Effective Storage for Infrequent Oracle Data
Historical database data that is large, infrequently accessed, and throughput oriented does not need premium SSD performance.
Recommended architecture
● Use AWS Database Migration Service to move the Oracle workload or data.
● Place infrequently accessed historical data on EBS sc1 Cold HDD volumes when the database remains on EC2.
● Validate that the workload does not require boot-volume support or high random IOPS.
Design reasoning
● Technical workability: DMS migrates database records, while sc1 provides low-cost HDD storage for large, sequential workloads.
● Requirement fit: The design minimizes storage expense for cold historical data.
● Scene fit: Access is infrequent and throughput matters more than low-latency random I/O.
● Engineering common sense: Choose storage from the actual access pattern, not from maximum possible performance.
Alternatives to reject
● Server Migration Service moves servers rather than providing the database migration workflow.
● gp2 SSD costs more and targets general-purpose random I/O.
● st1 Throughput Optimized HDD suits frequently accessed throughput workloads and costs more than sc1.
● Provisioned IOPS SSD is unnecessary for cold historical records.
Workflow: Assess I/O → migrate with DMS → attach and format sc1 → validate throughput → monitor access pattern.
Shifting S3 Retrieval Charges with Requester Pays
When a partner frequently downloads objects from another company’s S3 buckets, Requester Pays can assign request and transfer charges to the requester.
Recommended architecture
● Enable S3 Requester Pays on the shared buckets.
● Grant the partner the required object permissions.
● Require requests to acknowledge requester billing.
● Monitor access and cost allocation.
Design reasoning
● Technical workability: Authorized requesters include the requester-pays parameter and are billed for eligible requests and data transfer.
● Requirement fit: The startup keeps one authoritative copy while the media company pays for its consumption.
● Scene fit: Frequent partner retrievals, rather than storage growth, are driving cost.
● Engineering common sense: Change the billing model before duplicating the entire dataset.
Alternatives to reject
● Synchronizing into another bucket duplicates storage and requires ongoing consistency management.
● AWS Organizations and SCPs govern accounts but do not shift S3 retrieval charges.
● Cross-account permissions alone do not make the requester pay; billing behavior must be explicitly enabled.
Workflow: Enable Requester Pays → update partner request process → test billing acknowledgement → observe usage → retain normal owner access for internal workflows.
Federated Employee Access to Personal S3 Prefixes
Directory users can receive temporary credentials limited to a matching personal S3 prefix.
Recommended architecture
● Use a federation broker to authenticate corporate directory users.
● Map identity attributes to an IAM role session.
● Use STS for temporary credentials.
● Apply an IAM policy with variables that map each user to an S3 prefix.
Design reasoning
● Technical workability: Policy variables scope object paths from federated identity context.
● Requirement fit: Employees use existing credentials without duplicate IAM users.
● Scene fit: Each person needs a separate folder in one bucket.
● Engineering common sense: Keep authentication centralized and authorize data paths dynamically.
Alternatives to reject
● Amazon Connect is a contact-center service.
● Mirroring every directory identity into IAM adds lifecycle overhead.
● An IAM user per employee creates duplicate passwords or access keys.
● Static shared credentials cannot isolate personal prefixes.
Workflow: Employee signs in → broker validates directory → STS issues session → policy variable selects prefix → S3 request.
CloudFront HTTPS and Cache Efficiency
A custom CloudFront domain needs a trusted certificate and cache instructions that keep reusable objects at edge locations.
Recommended architecture
● Request an ACM certificate in us-east-1 for the CloudFront hostname.
● Attach it to the distribution.
● Set a long practical Cache-Control: max-age for safe cacheable objects.
Design reasoning
● Technical workability: CloudFront uses ACM certificates from us-east-1; origin cache headers influence edge retention.
● Requirement fit: Managed TLS secures the custom domain, while longer TTLs improve cache hits and reduce origin traffic.
● Scene fit: The content can be cached and does not require per-request processing.
● Engineering common sense: Cache only content whose freshness and privacy permit reuse.
Alternatives to reject
● Storing a certificate in S3 does not attach it to CloudFront and creates private-key handling.
● OpenSearch is not a CDN cache.
● Lambda@Edge adds cost and latency without solving the basic certificate and TTL requirements.
● A very short TTL forces unnecessary origin requests.
Workflow: Viewer HTTPS → CloudFront validates certificate → edge checks freshness → serve cached object or request origin.
Fast Disaster Recovery for JBoss and Oracle
Existing backup artifacts can be restored into AWS faster than redesigning the application during a disaster.
Recommended architecture
● Launch EC2 capacity for the JBoss application and Oracle database.
● Restore Oracle with RMAN backups stored in Amazon S3.
● Restore the Storage Gateway volume snapshot as an Amazon EBS volume.
● Attach the EBS volume to the application server.
● Recreate networking and configuration from prepared templates.
Design reasoning
● Technical workability: RMAN restores Oracle data, and Storage Gateway snapshots can create EBS volumes for native EC2 block access.
● Requirement fit: The design uses existing backups to achieve a shorter recovery time.
● Scene fit: The source architecture already uses JBoss, Oracle, and Storage Gateway.
● Engineering common sense: Pre-document restore order and dependencies before an incident.
Alternatives to reject
● Keeping the application dependent on the on-premises gateway defeats regional disaster recovery.
● S3 object storage cannot directly replace the required attached block volume.
● Converting the workload to EFS during recovery adds unsupported transformation work.
● Refactoring Oracle and JBoss during an outage increases recovery time.
Workflow: Declare disaster → deploy EC2 → restore Oracle → create EBS from snapshot → attach → validate → switch traffic.
Private Cross-VPC Connectivity and Rejected-Traffic Monitoring
Same-Region VPCs can communicate privately through peering, while Flow Logs record accepted and rejected traffic.
Recommended architecture
● Peer each departmental VPC with the central application VPC.
● Add the required routes and security rules.
● Enable VPC Flow Logs to capture rejected records and source addresses.
● Deliver logs to CloudWatch Logs.
● Use a Logs subscription to forward records to the security account.
Design reasoning
● Technical workability: Peering routes traffic over AWS networking without the public Internet; Flow Logs record network-interface traffic decisions.
● Requirement fit: The design provides private communication and visibility into denied requests.
● Scene fit: The VPCs are in one Region and require access to one central service.
● Engineering common sense: Use direct AWS networking before introducing encryption appliances or external circuits that solve no stated problem.
Alternatives to reject
● Separate IPsec tunnels add gateway and routing operations.
● A third-party transit VPC is excessive and AWS Config does not record packet rejects.
● Direct Connect connects external networks to AWS; it is not a VPC-to-VPC service.
Workflow: Department VPC → peering route → central service; rejected flow → Flow Logs → CloudWatch → security subscription.
Scalable Media Website and Cost Review
A dynamic media site should separate durable media delivery from elastic application processing and use native cost-analysis services.
Recommended architecture
● Store media in S3 and deliver it through CloudFront.
● Run dynamic web servers in Auto Scaling behind an ELB.
● Use Cost Explorer and Trusted Advisor to identify savings.
Design reasoning
● Technical workability: S3 and CloudFront scale media delivery, while Auto Scaling handles variable server-side traffic.
● Requirement fit: The design improves performance, resilience, and cost visibility.
● Scene fit: The application contains both static media and server-side logic.
● Engineering common sense: Do not route heavy media through application servers.
Alternatives to reject
● CloudFront cannot use EFS directly as an origin.
● Storage-optimized web instances are unnecessary when media moves to S3.
● Consolidated Billing aggregates charges but is not a cost-analysis tool.
● A fully static S3 site cannot execute server-side processing.
● Cross-Region replication adds unrequested cost.
Workflow: Static media → CloudFront and S3; dynamic request → ELB and Auto Scaling; costs → Explorer and Advisor.
Layered DDoS Resilience Without Shield Advanced
When premium DDoS response services are outside the budget, web resilience should combine edge absorption, filtering, load distribution, monitoring, and elasticity.
Recommended architecture
● Use Amazon CloudFront for static and dynamic web delivery where appropriate.
● Place an Application Load Balancer in front of multiple application instances.
● Apply AWS WAF rules to CloudFront or the ALB for application-layer attack patterns.
● Create CloudWatch alarms for CPU and network pressure.
● Scale the EC2 fleet automatically when demand rises.
Design reasoning
● Technical workability: Edge capacity absorbs traffic, WAF filters requests, ALB spreads load, and Auto Scaling adds application capacity.
● Requirement fit: The controls improve availability without Shield Advanced cost.
● Scene fit: The protected workload is an EC2-hosted web application.
● Engineering common sense: No single control handles every DDoS layer; use complementary defenses.
Alternatives to reject
● Reserved Instances are a billing model and provide no extra performance.
● Additional ENIs or enhanced networking do not identify malicious requests.
● S3 is not POSIX storage, and patching does not mitigate an active traffic flood.
Workflow: Edge → WAF → ALB → targets → scaling and alarms.
Blocking SQL Injection at API Gateway
Injection attacks should be filtered at the managed API entry point, while firewall configuration changes should be recorded for audit.
Recommended architecture
● Associate an AWS WAF web ACL with Amazon API Gateway.
● Enable managed or custom rules that detect SQL injection patterns.
● Use AWS Config to record web ACL, rule, and configuration changes.
● Review WAF metrics and blocked-request logs.
Design reasoning
● Technical workability: AWS WAF integrates with API Gateway and inspects HTTP requests before backend invocation.
● Requirement fit: The solution blocks the demonstrated attack pattern and provides historical configuration tracking at low cost.
● Scene fit: The application already uses API Gateway and Lambda, so no new load-balancing tier is required.
● Engineering common sense: Block attack classes rather than one observed source IP.
Alternatives to reject
● API Gateway cannot be placed behind an ALB in the proposed manner.
● WAF does not attach directly to individual Lambda functions.
● Firewall Manager centralizes policy administration but is not the requested configuration-history recorder.
● VPC network ACLs do not protect a managed API Gateway endpoint from SQL injection content.
Workflow: Request → WAF inspection → API Gateway → Lambda → data tier; Config records policy changes.
Central Governance for Subsidiary Accounts
One AWS Organization can centralize subsidiary billing and apply service restrictions through organizational units.
Recommended architecture
● Place all subsidiaries in one AWS Organization.
● Group accounts into OUs by governance need.
● Attach Service Control Policies to define maximum services and actions.
● Use consolidated billing for centralized cost management.
Design reasoning
● Technical workability: SCPs constrain member-account permissions, while consolidated billing aggregates charges.
● Requirement fit: The parent company manages service boundaries and costs centrally.
● Scene fit: Subsidiaries remain separate AWS accounts.
● Engineering common sense: SCPs set guardrails; account IAM still grants actual permissions.
Alternatives to reject
● Consolidated billing alone does not impose service restrictions.
● One account cannot belong to multiple organizations.
● A separate organization per subsidiary prevents one parent governance hierarchy.
● Service quotas limit quantities, not which APIs or services are allowed.
Workflow: Invite accounts → arrange OUs → attach SCPs → configure account IAM → review consolidated costs.
Promoting OS Patches Across Environments
Critical patches should move through development and test before production, with separate baselines preserving each environment’s policy.
Recommended architecture
● Tag every instance by environment and operating system.
● Create separate Systems Manager Patch Manager baselines for development, test, and production.
● Map instances to Patch Groups using those tags.
● Patch development and test immediately.
● Patch production only after validation.
Design reasoning
● Technical workability: Patch Groups associate tagged instances with approved patch baselines.
● Requirement fit: Environment-specific control prevents unverified patches from reaching production.
● Scene fit: Each environment already has different baseline requirements.
● Engineering common sense: The same artifact should be promoted through increasingly critical environments after evidence of stability.
Alternatives to reject
● OS-only tags cannot distinguish production from non-production fleets.
● Maintenance Windows control timing but do not define which patches each environment approves.
● Custom shell scripts and Run Command duplicate managed patching functions and require ongoing maintenance.
Workflow: Tag → define baselines → patch development → validate → patch test → validate → approve production baseline → patch production → review compliance.
Real-Time Comments with AppSync Subscriptions
Clients can receive new comments in real time through GraphQL subscriptions over WebSockets.
Recommended architecture
● Publish comment mutations through AWS AppSync.
● Define GraphQL subscriptions linked to those mutations.
● Let connected clients receive pushed updates.
● Store comments in the existing durable data store.
Design reasoning
● Technical workability: AppSync maintains WebSocket connections and pushes subscription events.
● Requirement fit: New comments appear without repeated polling.
● Scene fit: The application needs real-time client updates.
● Engineering common sense: Push only events users are authorized to receive.
Alternatives to reject
● Polling every three seconds increases API and Lambda load.
● More Lambda concurrency does not push data to connected clients.
● CloudFront caching can serve stale comments and is not a bidirectional real-time channel.
● Database writes alone do not notify clients.
Workflow: User posts mutation → AppSync authorizes and stores → subscription event → connected clients update.
Real-Time Clickstream Processing
Web clickstream events require continuous ingestion and analysis rather than periodic batch collection.
Recommended architecture
● Publish click events to Amazon Kinesis Data Streams.
● Partition records using a key that distributes traffic evenly.
● Process events with stream consumers in near real time.
● Persist raw or aggregated results to durable analytical storage.
Design reasoning
● Technical workability: Kinesis accepts ordered records per shard and supports multiple low-latency consumers.
● Requirement fit: The platform can evaluate user behavior while sessions are active.
● Scene fit: Website clicks form a continuous event stream with bursty volume.
● Engineering common sense: Keep ingestion durable and decoupled from downstream analytical speed.
Alternatives to reject
● Nightly S3 batch jobs do not meet real-time requirements.
● SQS is useful for work queues but does not provide the same ordered streaming and multiple-consumer model.
● Direct writes to a warehouse couple website ingestion to analytical availability.
● CloudTrail records AWS account activity, not application clickstream events.
Workflow: Browser event → Kinesis partition → stream processor → real-time metric or action → durable archive and later analytics.
Multi-Region Transactions with DynamoDB Global Tables
A global marketplace needs local regional reads and writes with managed replication between Regions.
Recommended architecture
● Create a DynamoDB global table with replicas in all required Regions.
● Direct each regional API to its local replica.
● Let DynamoDB replicate item changes across all replica tables.
● Design transaction keys and conflict behavior for multi-Region writes.
Design reasoning
● Technical workability: Global tables provide managed multi-active replication and regional endpoints.
● Requirement fit: The design lowers latency for worldwide users and avoids custom replication infrastructure.
● Scene fit: Millions of marketplace users and multi-Region APIs require horizontally scalable storage.
● Engineering common sense: Prefer the database’s native replication feature over building a replay system.
Alternatives to reject
● The described Aurora Multi-Master architecture is not a valid multi-Region active-active design.
● Lambda-based replay duplicates built-in replication and introduces ordering, retry, and conflict-management work.
● Control Tower governs AWS accounts and does not replicate application records.
● Amazon Connect is a contact-center service and is unrelated to database replication.
Workflow: API writes local replica → stream-based global replication → remote replicas update → applications read locally → monitor replication and conflict metrics.
Managed Petabyte-Scale Batch Processing
Large parallel image workloads need managed job scheduling, elastic low-cost compute, durable object storage, and temporary local processing space.
Recommended architecture
● Package the processing executable for AWS Batch jobs.
● Use a managed compute environment with EC2 Spot capacity.
● Store raw and processed images in separate Amazon S3 locations.
● Use job queues to schedule thousands of parallel tasks.
● Use EBS volumes only as temporary local workspace.
Design reasoning
● Technical workability: AWS Batch provisions compute and schedules jobs, while S3 scales to petabytes with high durability.
● Requirement fit: Spot reduces compute cost and the managed scheduler minimizes operational overhead.
● Scene fit: Jobs are independent, parallel, and complete over about one week.
● Engineering common sense: Keep durable data outside disposable workers.
Alternatives to reject
● A custom SQS and Auto Scaling worker platform requires more scheduling and fleet logic.
● EFS is not the most economical shared source for a 10 PB batch dataset.
● EKS can run jobs but adds Kubernetes administration.
● EMR is appropriate for supported big-data frameworks, not automatically for an existing executable that would need rewriting for Spark.
Workflow: Upload input → submit jobs → Batch schedules Spot workers → stage to EBS → process → write results to S3.
Administrative Access to Invited Organization Accounts
Joining AWS Organizations centralizes governance and billing, but administrative access still requires a trusted role in each member account.
Recommended architecture
● Invite existing accounts from the organization management account.
● Ensure each member account has an OrganizationAccountAccessRole or equivalent administrative role.
● Trust approved management-account principals.
● Use STS role assumption for administration.
Design reasoning
● Technical workability: Organizations handles membership; IAM role trust grants actual cross-account permissions.
● Requirement fit: Central administrators can manage invited accounts without duplicate long-lived users.
● Scene fit: The accounts existed before joining the organization.
● Engineering common sense: Account membership and account authorization are separate control planes.
Alternatives to reject
● Membership alone does not grant administration.
● Invitations originate from the management account.
● Control Tower enrollment adds governance but should not be assumed to grant unrestricted access automatically.
● Sharing root credentials is unsafe and unnecessary.
Workflow: Management account sends invitation → member accepts → administrator assumes trusted role → temporary credentials manage account.
Rapid Global Offload for an On-Premises Site
CloudFront can use an existing on-premises website as a custom origin and reduce global traffic quickly.
Recommended architecture
● Configure the on-premises site as a CloudFront custom origin.
● Cache static and safely reusable dynamic HTTP responses.
● Tune cache behaviors and TTLs.
● Keep noncacheable application traffic at the origin.
Design reasoning
● Technical workability: CloudFront retrieves content from internet-accessible custom origins and serves cached responses globally.
● Requirement fit: The solution improves scale under a short deadline without full migration.
● Scene fit: The current application must remain on premises.
● Engineering common sense: Offload the dominant repeated traffic first.
Alternatives to reject
● Reversing static and dynamic service roles creates an unworkable design.
● Transit Gateway does not provide public CDN delivery.
● App Runner cannot manage on-premises servers.
● Replicating the full infrastructure and shifting half the traffic is slower and costlier.
Workflow: Viewer → CloudFront → edge hit or custom-origin request → on-premises application.
Cost-Effective Regional DR Artifacts
Recovery can avoid hot standby cost by continuously replicating the database and copying rebuild artifacts into the backup Region.
Recommended architecture
● Maintain a cross-Region Aurora replica.
● Copy web and application AMIs or snapshots to the recovery Region.
● Store templates and configuration needed to rebuild stateless tiers.
● Promote the replica and launch compute during disaster.
Design reasoning
● Technical workability: Aurora replication keeps data current, while regional images enable fast server recreation.
● Requirement fit: The design reduces steady compute cost while retaining practical recovery.
● Scene fit: Stateless tiers can be rebuilt; the database needs continuous replication.
● Engineering common sense: Recovery artifacts must already exist in the target Region.
Alternatives to reject
● AWS Backup does not simply place RDS and EBS backups into a customer S3 bucket as described.
● Constant five-minute Aurora snapshots are not the native regional replication method.
● Snapshot-only restoration increases RTO.
● Hot standby is faster but costs more than required.
Workflow: Replicate Aurora → copy images → disaster declared → promote replica → launch tiers → validate → switch traffic.
Infrastructure as Code with Global CDN
Repeatable infrastructure deployment and global content acceleration are separate needs served by CloudFormation and CloudFront.
Recommended architecture
● Define networks, compute, security, and managed services in AWS CloudFormation.
● Store templates in version control and deploy through controlled pipelines.
● Use Amazon CloudFront as the global CDN for cacheable content.
Design reasoning
● Technical workability: CloudFormation creates and updates AWS resources consistently; CloudFront caches content at edge locations.
● Requirement fit: The environment becomes repeatable while users receive lower-latency content.
● Scene fit: The solution needs full infrastructure automation, not only application deployment.
● Engineering common sense: Monitoring, deployment, and content delivery should not be confused.
Alternatives to reject
● CloudWatch observes systems but is neither infrastructure as code nor a CDN.
● Elastic Beanstalk manages an application platform but does not directly model the complete environment as comprehensively.
● Manual console deployment creates drift.
● A CDN does not replace infrastructure provisioning.
Workflow: Commit template → validate → deploy CloudFormation stack → publish content → CloudFront caches and delivers globally.
Organization-Wide CloudTrail Logging
A centralized organization trail records activity from member accounts without creating separate trails for each access method.
Recommended architecture
● Create a CloudTrail organization trail from the management account.
● Enable coverage for all organization accounts and required Regions.
● Deliver logs to a dedicated S3 bucket.
● Encrypt audit logs and enable strong deletion protection, including MFA Delete where appropriate.
Design reasoning
● Technical workability: One trail records console, SDK, CLI, and API activity across participating accounts.
● Requirement fit: Central storage improves audit coverage, integrity, and administration.
● Scene fit: The company manages multiple accounts through AWS Organizations.
● Engineering common sense: Keep audit evidence separate from workload administrators.
Alternatives to reject
● A single-account trail does not cover the organization.
● SNS delivery notifications are not the audit records.
● A regional trail may omit activity in other Regions.
● Separate trails for console, SDK, and CLI duplicate configuration unnecessarily.
Workflow: Member-account action → organization trail → encrypted S3 log → controlled auditor access → retention monitoring.
Separating Web Health from Database Health
A web-tier health check should test the web process without causing database overload to trigger an instance replacement storm.
Recommended architecture
● Change the ALB target health check to a simple static HTTP page.
● Add Amazon ElastiCache for frequently repeated database queries.
● Monitor the database-dependent application endpoint separately with a Route 53 health check.
● Alert administrators when the end-to-end check fails.
Design reasoning
● Technical workability: Static health checks verify the web server, while ElastiCache reduces repeated RDS reads.
● Requirement fit: Auto Scaling stops replacing healthy web instances, and the database can handle more traffic.
● Scene fit: RDS CPU saturation, not web-server failure, caused the timeouts.
● Engineering common sense: Health checks used for replacement should isolate the component being replaced.
Alternatives to reject
● Rebooting an overloaded database does not fix sustained demand.
● The proposed RDS MySQL single reader endpoint is not valid as described.
● A TCP check proves only that a port accepts connections and is weaker than a simple HTTP page.
Workflow: ALB checks static page → Route 53 checks full transaction → cache absorbs reads → alarms report database-path failure.
RDS Multi-AZ Failover and DNS
Applications should connect to an RDS endpoint rather than a database instance IP address.
Failover behavior
● RDS maintains a synchronous standby in another Availability Zone.
● When the primary fails, RDS promotes the standby.
● The database endpoint’s CNAME resolves to the new primary address.
● Applications reconnect after DNS and connection recovery.
Design reasoning
● Technical workability: RDS controls promotion and DNS remapping for the managed deployment.
● Requirement fit: Applications resume without manually changing connection strings.
● Scene fit: The design uses Multi-AZ for database availability.
● Engineering common sense: Connection pools must discard broken connections and resolve the endpoint again.
Incorrect assumptions to reject
● The application should not pin the old database IP.
● Multi-AZ standby instances are not normal read-scaling endpoints.
● Route 53 records do not need a manual operator change for standard RDS failover.
● Restoring a snapshot is not the normal Multi-AZ failover process.
Workflow: Primary failure → RDS detects → standby promoted → endpoint DNS changes → application retries and reconnects.
Separating Certificate Control from Application Operations
Application teams can manage EC2 workloads while cybersecurity retains exclusive control of production TLS certificates.
Recommended architecture
● Store and manage the certificate in AWS Certificate Manager.
● Restrict ACM certificate actions to the cybersecurity team through IAM policies.
● Terminate SSL or TLS on the Elastic Load Balancer.
● Keep certificate material off application EC2 instances.
● Let DevOps manage only the application and compute permissions it requires.
Design reasoning
● Technical workability: The load balancer integrates with ACM without distributing private keys to servers.
● Requirement fit: Duties remain separated while the portal provides HTTPS.
● Scene fit: DevOps owns compute operations, and cybersecurity owns X.509 certificate lifecycle.
● Engineering common sense: Enforce separation at the service and permission boundary, not through informal procedures.
Alternatives to reject
● Storing certificates in S3 creates retrieval and key-exposure paths.
● Installing certificates on EC2 gives application administrators access to private material.
● SCPs are broad organization guardrails and are not the normal tool for fine-grained team certificate access.
● AWS Config can detect changes but does not grant or deny certificate use.
Workflow: Cybersecurity manages ACM certificate → load balancer terminates TLS → DevOps manages backend targets without certificate access.
Rotating Database Passwords with Secrets Manager
Database passwords should be stored in a dedicated secret service, retrieved at runtime, and referenced without appearing in templates.
Recommended architecture
● Store the RDS master password in AWS Secrets Manager.
● Encrypt it with the approved KMS key.
● Configure managed rotation where supported.
● Let the application retrieve the secret at runtime through an IAM role.
● Use a CloudFormation dynamic reference for MasterUserPassword.
Design reasoning
● Technical workability: Dynamic references resolve the secret during provisioning without embedding plaintext in the template.
● Requirement fit: The password is centrally protected, access controlled, and rotatable.
● Scene fit: Both infrastructure deployment and application connections require the same credential lifecycle.
● Engineering common sense: Never copy a database password into user data, source repositories, or template parameters returned in outputs.
Alternatives to reject
● Plain CloudFormation parameters can expose values through handling and deployment workflows.
● A normal Ref does not provide secret rotation.
● Hard-coded application configuration becomes stale after rotation.
● Parameter storage without the required rotation workflow is less direct for this requirement.
Workflow: Create secret → deploy RDS with dynamic reference → application retrieves secret → rotate → application refreshes connection.
Stable Outbound IP for Third-Party Allowlisting
Application servers can share one predictable outbound public address through a NAT Gateway.
Recommended architecture
● Create a NAT Gateway in a public subnet.
● Associate an Elastic IP.
● Route private application subnets through the NAT Gateway.
● Give the Elastic IP to the payment provider for allowlisting.
Design reasoning
● Technical workability: Outbound connections are translated to the NAT Gateway’s Elastic IP.
● Requirement fit: The provider sees a stable source address while servers remain private.
● Scene fit: The integration initiates connections from application servers to a third party.
● Engineering common sense: Use managed egress rather than assigning public addresses to every server.
Alternatives to reject
● A NAT instance can provide a fixed address but requires patching and failover management.
● ELB handles inbound traffic and does not centralize server egress.
● An Internet Gateway does not provide one customer-controlled shared source IP.
● A customer gateway is a VPN endpoint, not the correct internet egress service.
Workflow: Private server → NAT Gateway → Elastic IP translation → payment endpoint → allowlist accepts connection.
Social Login with Temporary AWS Credentials
A mobile application can exchange an OIDC social identity token for temporary AWS permissions.
Recommended architecture
● Authenticate users with the supported social identity provider.
● Call STS AssumeRoleWithWebIdentity.
● Map the identity to an IAM role scoped for S3 and DynamoDB.
● Use the temporary credentials in the mobile SDK.
Design reasoning
● Technical workability: STS validates the web identity token and returns short-lived role credentials.
● Requirement fit: The app accesses AWS resources without embedded long-term keys.
● Scene fit: Social providers use OIDC-style web identity federation.
● Engineering common sense: Treat mobile devices as untrusted locations for permanent secrets.
Alternatives to reject
● Long-term access and secret keys remain insecure even when Cognito is mentioned.
● AssumeRoleWithSAML targets enterprise SAML federation.
● Generic AssumeRole with an IAM user does not directly validate social tokens.
● Distributing one shared key prevents safe user isolation.
Workflow: Social login → OIDC token → STS role exchange → temporary credentials → scoped S3 or DynamoDB request.
Protecting Payment Fields and Improving CloudFront Caching
Transport encryption protects connections, while field-level encryption protects selected sensitive values as they pass through intermediary systems.
Recommended architecture
● Require HTTPS from viewers to CloudFront.
● Require HTTPS from CloudFront to the origin.
● Configure CloudFront field-level encryption for credit-card fields using the approved public key.
● Set appropriate Cache-Control directives for content that is safe to cache.
Design reasoning
● Technical workability: CloudFront encrypts selected form fields before forwarding them, and only the private-key holder can decrypt them.
● Requirement fit: Card data remains protected through the delivery path, while longer safe TTLs improve cache hit ratio.
● Scene fit: The application uses CloudFront and accepts sensitive payment information.
● Engineering common sense: Never cache personalized payment responses merely to improve performance.
Alternatives to reject
● Signed URLs control access but do not encrypt selected payment fields.
● Custom TLS certificates secure connections but do not provide field-level protection.
● Origin Access Identity secures an S3 origin, not payment attributes.
● Forwarding unnecessary User-Agent or Host variations fragments the cache and lowers hit ratio.
Workflow: HTTPS request → CloudFront encrypts sensitive fields → origin processes protected values → cache only approved content.
Cost-Effective Recovery with Backups and Transaction Logs
Recovery objectives should determine backup frequency, transaction-log frequency, and storage tier.
Recommended architecture
● Create periodic database backups and store them in Amazon S3.
● Export transaction logs to a separate S3 location every five minutes.
● Restore the latest backup and replay logs during recovery.
● Use Amazon Macie to discover and classify sensitive data stored in S3.
Design reasoning
● Technical workability: A base backup plus frequent logs reconstructs the database close to the failure point.
● Requirement fit: Five-minute logs fit a 15-minute RPO, while S3 retrieval supports a recovery target under three hours more appropriately than archival storage.
● Scene fit: The application uses RDS MySQL and contains deliveries, transactions, and potentially sensitive data.
● Engineering common sense: RPO controls data capture frequency; RTO controls how quickly backups can be retrieved and restored.
Alternatives to reject
● Storage Gateway is unnecessary for an AWS-hosted database backup path.
● Multi-AZ is high availability, not a separate disaster-recovery backup.
● AWS Shield does not discover PII.
● Glacier retrieval can make a short recovery target difficult.
Workflow: Backup → capture logs → store securely → detect sensitive objects → restore base → replay logs → validate service.
Blue/Green Releases with Elastic Beanstalk
A managed blue/green release keeps the current environment available while a new version is validated independently.
Recommended architecture
● Clone or create a second Elastic Beanstalk environment.
● Deploy and test the new application version there.
● Swap environment CNAMEs when validation succeeds.
● Keep the old environment briefly for rollback.
Design reasoning
● Technical workability: Both environments run concurrently, and the CNAME swap redirects users without rebuilding the URL.
● Requirement fit: Deployment has no planned downtime and rollback is fast.
● Scene fit: The application already uses Elastic Beanstalk.
● Engineering common sense: Validate the exact production artifact before traffic moves.
Alternatives to reject
● Forced Auto Scaling replacement lacks a clean staged environment and instant URL swap.
● Lightsail and weighted DNS introduce unrelated infrastructure.
● Replacing all instances at once removes controlled validation.
● In-place updates increase rollback risk.
Workflow: Build green → deploy version → test → swap CNAME → monitor → swap back if needed → retire blue.
Long-Running Batch Processing with SQS
Long data-processing jobs need durable queuing, elastic workers, and durable output storage.
Recommended architecture
● Place work items in Amazon SQS.
● Scale EC2 workers based on queue length.
● Store processed files in Amazon S3.
● Configure visibility timeout, retries, and a dead-letter queue.
Design reasoning
● Technical workability: SQS decouples producers, Auto Scaling follows backlog, and S3 stores outputs outside workers.
● Requirement fit: The design handles variable long jobs cost-effectively.
● Scene fit: Processing may exceed Lambda duration or resource limits.
● Engineering common sense: Never rely on manual server shutdown for elasticity.
Alternatives to reject
● Amazon MQ and EFS can work but add broker and filesystem operations.
● Mixing an MQ queue with a Lambda consumer reading SQS is inconsistent.
● Lambda may exceed runtime limits.
● EFS is less operationally simple than S3 for completed object outputs.
Workflow: Submit job → SQS → worker fleet scales → process → write S3 result → delete message.
Applying Tags at Resource Creation
Governance tags should be applied during provisioning rather than discovered and repaired afterward.
Recommended architecture
● Publish approved products through AWS Service Catalog so provisioned resources inherit portfolio, product, and user metadata.
● Define required Tags properties in CloudFormation resources.
● Protect provisioning templates and tag modification permissions.
Design reasoning
● Technical workability: Service Catalog and CloudFormation apply supported tags as resources are created.
● Requirement fit: Cost allocation and ownership metadata exist immediately.
● Scene fit: Resources are provisioned through governed service and infrastructure templates.
● Engineering common sense: Prevent missing metadata where possible, then use detective controls for exceptions.
Alternatives to reject
● Systems Manager Automation adds tags after creation.
● AWS-generated cost-allocation tags do not cover every organization-specific requirement.
● AWS Config detects missing tags but does not add them at creation time.
● Manual tagging is inconsistent and difficult to audit.
Workflow: User selects catalog product or pipeline deploys stack → tags applied → resource created → compliance verified.
Protecting iSCSI Sessions with CHAP
Storage Gateway iSCSI sessions should authenticate initiators before block storage is exposed.
Recommended control
● Configure Challenge-Handshake Authentication Protocol for the iSCSI target.
● Store CHAP credentials securely on the gateway and approved initiators.
● Restrict network access to the required iSCSI ports and source systems.
● Rotate credentials through controlled maintenance.
Design reasoning
● Technical workability: CHAP authenticates an iSCSI initiator through a challenge-response exchange without sending the secret directly.
● Requirement fit: The control reduces unauthorized attachment and replay risk for gateway block volumes.
● Scene fit: The interface is iSCSI, so protection should use its supported authentication mechanism.
● Engineering common sense: Combine protocol authentication with network restrictions; neither replaces the other.
Alternatives to reject
● SMB and NFS authentication settings do not secure an iSCSI volume.
● HTTPS protects management or API traffic, not the iSCSI session itself.
● S3 bucket policies do not authenticate on-premises iSCSI initiators.
● Encryption alone does not prove which host is attaching the target.
Workflow: Initiator connects → gateway issues challenge → initiator responds using shared secret → gateway validates → iSCSI session opens.
Automated EC2 Rescue with Systems Manager
Impaired Windows or Linux EC2 instances can be diagnosed through a purpose-built Systems Manager Automation runbook.
Recommended procedure
● Run AWSSupport-ExecuteEC2Rescue through Systems Manager Automation.
● Supply the impaired instance and required permissions.
● Let the workflow collect diagnostics and apply supported remediation.
● Review outputs before returning the instance to service.
Design reasoning
● Technical workability: EC2Rescue automates common access and boot troubleshooting steps.
● Requirement fit: Recovery is faster and more consistent than manual repair.
● Scene fit: The issue is immediate instance impairment.
● Engineering common sense: Use a support runbook before building custom repair automation.
Alternatives to reject
● AWS Config and State Manager do not provide this rescue workflow.
● Maintenance Windows schedule tasks and are not required for immediate repair.
● OpsWorks Chef Automate adds unrelated configuration infrastructure.
● Session Manager may provide access but does not itself diagnose and remediate the problem.
Workflow: Start automation → create helper resources if needed → diagnose → remediate → review report → validate instance.
Least-Privilege Cross-Account S3 and KMS
Cross-account access to KMS-encrypted objects requires authorization from the key, the role, and the bucket.
Recommended controls
● Add the specific management-account reviewer role to the KMS key policy for decrypt.
● Give the reviewer role identity permissions for S3 read and KMS decrypt.
● Allow the reviewer role or Management account in the S3 bucket policy.
Design reasoning
● Technical workability: S3 and KMS evaluate separate permission chains.
● Requirement fit: Only the designated reviewer can read and decrypt the objects.
● Scene fit: Objects are owned by one account and reviewed from another.
● Engineering common sense: Use the narrowest principal in every policy.
Alternatives to reject
● Trusting the Marketing account grants the wrong side.
● Granting the whole Management account decrypt permission is broader than required.
● Full S3 permissions exceed read-only need.
● IAM decrypt permission cannot override a KMS key policy that does not trust the role.
Workflow: Reviewer assumes role → S3 bucket authorizes read → KMS key authorizes decrypt → object returned.
Secure Hybrid Deployment with CodeDeploy
A deployment platform spanning EC2 and on-premises servers needs different identity mechanisms but one release workflow.
Recommended architecture
● Store database credentials as KMS-encrypted SecureString parameters in Systems Manager Parameter Store.
● Install the required deployment and management agents.
● Give EC2 instances an instance profile with parameter and decrypt permissions.
● Register on-premises instances with the appropriate service identity.
● Use AWS CodeDeploy for both target environments.
Design reasoning
● Technical workability: CodeDeploy supports EC2 and registered on-premises instances; Parameter Store provides encrypted runtime configuration.
● Requirement fit: Credentials remain encrypted at rest and in transit, and releases are automated consistently.
● Scene fit: The application runs on both Reserved EC2 capacity and existing on-premises servers.
● Engineering common sense: Do not pretend an on-premises server can receive an EC2 instance profile.
Alternatives to reject
● Elastic Beanstalk does not deploy to arbitrary on-premises servers.
● Mixing Secrets Manager storage with Parameter Store permissions is inconsistent.
● Attaching an EC2 instance-profile policy to an on-premises machine is not the correct identity model.
Workflow: Store parameter → assign target identities → deploy revision → retrieve secret at runtime → validate application.
Managed Document OCR and Entity Extraction
Scanned forms can be processed with managed OCR, text analysis, and serverless workflow orchestration.
Recommended architecture
● Use Step Functions to coordinate processing stages.
● Invoke Lambda functions for integration and transformation.
● Use Amazon Textract for document OCR and structure extraction.
● Use Amazon Comprehend for entity or text analysis.
● Store structured results in RDS and durable artifacts in S3 as required.
Design reasoning
● Technical workability: Textract extracts document content, Comprehend analyzes text, and Step Functions handles retries and sequencing.
● Requirement fit: The design automates processing with low operational overhead.
● Scene fit: The inputs are documents, not speech or general photographs.
● Engineering common sense: Use purpose-built managed services before training or operating custom OCR systems.
Alternatives to reject
● A custom SageMaker model adds training and maintenance.
● Rekognition and Transcribe target images and speech, not structured document OCR.
● Self-managed OCR on EKS adds cluster and software operations.
● One large Lambda function weakens observability and retry control.
Workflow: Form upload → state machine → Textract → Comprehend → validation → RDS or S3 output.
PII Discovery, Storage Lifecycle, and Database Availability
A photo-sharing service needs separate controls for sensitive object discovery, storage-cost optimization, and relational database downtime.
Recommended architecture
● Use Amazon Macie to discover and classify PII in S3.
● Apply S3 lifecycle rules that transition older photos to S3 Standard-IA when access drops.
● Configure the RDS database for Multi-AZ deployment.
Design reasoning
● Technical workability: Macie analyzes S3 objects for sensitive data, lifecycle policies automate storage transition, and RDS Multi-AZ provides managed standby failover.
● Requirement fit: The platform reduces storage cost, improves database availability, and gains visibility into PII exposure.
● Scene fit: User photos reside in S3 while application records remain relational.
● Engineering common sense: Use different services for data classification, object lifecycle, and database high availability.
Alternatives to reject
● Amazon Inspector assesses compute workloads and does not classify PII in S3.
● Moving every new photo immediately to Standard-IA can create retrieval and minimum-duration charges.
● Glacier is unsuitable when users still expect routine photo access.
● Redshift is an analytics warehouse, not a high-availability replacement for the transactional database.
● EBS changes do not solve managed RDS downtime.
Workflow: Upload photo → Macie evaluates sensitive content → lifecycle transitions aging objects → RDS fails over automatically when needed.
Stopping S3 Hotlinking with CloudFront
Private S3 origin access and expiring CloudFront signed URLs prevent direct object access and limit link reuse.
Recommended architecture
● Remove public access from the S3 bucket.
● Allow only the CloudFront origin identity or origin access control.
● Require signed URLs with short, appropriate expiration.
● Distribute links only to authorized users.
Design reasoning
● Technical workability: Users cannot bypass CloudFront, and expired signatures stop later reuse.
● Requirement fit: Unauthorized hotlinking and direct S3 downloads are reduced.
● Scene fit: The content is delivered through CloudFront.
● Engineering common sense: Viewer authorization and origin protection must both be present.
Alternatives to reject
● S3 has no security groups.
● Blocking website IPs is brittle and affects legitimate users.
● EBS on web servers reduces durability and creates a bottleneck.
● Static IP deny lists cannot reliably identify every hotlinking site.
Workflow: Authorized app issues signed URL → viewer requests CloudFront → signature checked → private S3 origin read → URL expires.
Immutable Patching, Frequent Deployment, and Shared Data
A scalable EC2 fleet should launch from a patched image, deploy application revisions automatically, and mount large shared data instead of downloading it during boot.
Recommended architecture
● Use Systems Manager automation to patch and bake a new AMI.
● Update the Auto Scaling group and replace old instances.
● Deploy application revisions with AWS CodeDeploy.
● Store the 500 GB static dataset on Amazon EFS and mount it at startup.
Design reasoning
● Technical workability: A golden AMI creates a consistent OS state, CodeDeploy supports frequent releases, and EFS is immediately accessible to multiple instances.
● Requirement fit: Instances scale quickly, deployments can occur several times daily, and patches meet the required deadline.
● Scene fit: The dataset is shared and static, while compute instances are replaceable.
● Engineering common sense: Do not download 500 GB during every scale-out event.
Alternatives to reject
● A nightly deployment job cannot support several releases per day.
● In-place patching can leave a mixed fleet.
● Waiting for a new vendor AMI does not guarantee patch installation within 48 hours.
● Large S3 downloads in user data delay boot.
Workflow: Patch image → test → update launch template → roll fleet → mount EFS → deploy code with CodeDeploy.
Scalable Mobile Photo and Caption Platform
A consumer mobile application expecting millions of views should separate authentication, small structured records, object storage, and global delivery.
Recommended architecture
● Use Amazon Cognito for user authentication and identity management.
● Store captions and user metadata in Amazon DynamoDB.
● Store uploaded photos and static assets in Amazon S3.
● Deliver static content through Amazon CloudFront.
Design reasoning
● Technical workability: Cognito supports mobile sign-in, DynamoDB scales key-value records, and S3 with CloudFront handles large object traffic.
● Requirement fit: All components scale with low operational overhead and no fixed server fleet.
● Scene fit: Photos are objects, while short captions are small structured items.
● Engineering common sense: Match each data type to the service designed for it.
Alternatives to reject
● RDS is workable but adds capacity planning and database operations for simple caption data.
● Direct mobile access to RDS is unsuitable.
● Social sign-in does not require an on-premises Active Directory and custom LDAP code.
● SAML is not the direct pattern for common consumer social identity providers.
Workflow: User signs in → receives scoped identity → uploads photo to S3 → writes caption to DynamoDB → viewers receive cached assets through CloudFront.
Temporary Per-User S3 Access for Mobile Apps
A mobile application should authenticate users and issue short-lived, scoped AWS role credentials for direct S3 access.
Recommended architecture
● Authenticate users through the application’s identity system.
● Map each identity to an IAM role or scoped session.
● Use AWS STS to issue temporary credentials.
● Restrict S3 prefixes and actions by user context.
Design reasoning
● Technical workability: Mobile SDKs can use temporary STS credentials to sign S3 requests.
● Requirement fit: Each user accesses only approved objects without sharing one permanent key.
● Scene fit: The distributed client needs direct object access at scale.
● Engineering common sense: Credentials on untrusted devices must be short-lived and narrowly scoped.
Alternatives to reject
● Embedding an IAM user key exposes the same long-term secret to every installation.
● STS does not issue permanent credentials, and cached sessions expire.
● Static mobile keys are difficult to rotate and cannot safely isolate users.
● Public bucket access removes user-level authorization.
Workflow: User login → identity validation → STS role session → scoped S3 request → automatic credential expiry.
Enabling an Availability Zone on an ALB
An Application Load Balancer routes traffic only to targets in Availability Zones enabled for that load balancer.
Recommended procedure
● Review which subnets and Availability Zones are associated with the ALB.
● Add a suitable subnet from the missing Availability Zone.
● Verify that target instances are registered and healthy.
● Confirm route tables, security groups, and network ACLs permit the health-check and application paths.
Design reasoning
● Technical workability: Associating the AZ gives the ALB a node and network path for that zone.
● Requirement fit: Healthy Auto Scaling instances in all intended AZs can receive traffic.
● Scene fit: Instances launch successfully but one AZ receives no load-balanced requests.
● Engineering common sense: Auto Scaling placement and load-balancer zone enablement are separate configurations.
Alternatives to reject
● Manually adding more instances does not help when the ALB cannot route to the AZ.
● Auto Scaling does not automatically associate new ALB subnets.
● Cross-zone load balancing does not make a disabled AZ available to the load balancer.
● The issue is not resolved by changing the AWS Region.
Workflow: ASG launches target → target registers → ALB checks enabled AZ → health check succeeds → traffic begins.
Unified System Telemetry for Windows and Linux
Operating-system memory, disk, and log data require an in-guest collector because standard EC2 metrics do not expose every requested measurement.
Recommended architecture
● Install and configure the unified Amazon CloudWatch agent across Windows and Linux instances.
● Send system logs to CloudWatch Logs.
● Publish in-guest metrics such as memory and disk utilization.
● Analyze aggregated records with CloudWatch Logs Insights.
● Export required results to Amazon S3.
Design reasoning
● Technical workability: One supported agent handles both operating systems and centralizes telemetry.
● Requirement fit: The solution covers more than 200 instances with minimal custom development.
● Scene fit: Monthly performance analysis needs fleet-wide searchable data rather than individual host inspection.
● Engineering common sense: Standardize collection configuration and query centrally.
Alternatives to reject
● SSM Agent manages instances but is not the primary collector for the full log and metric set.
● Traffic Mirroring captures network packets, not memory or disk usage.
● A custom daemon and CDK deployment add unnecessary maintenance.
● Amazon Inspector assesses vulnerabilities, while dashboards do not replace log queries.
Workflow: Define agent configuration → deploy → verify ingestion → query with Logs Insights → retain or export results.
Low-Cost Searchable Glacier Archives
Large datasets can be compressed into fewer Glacier archives while searchable metadata remains in DynamoDB.
Recommended architecture
● Compress each dataset into one archive.
● Store the archive in S3 Glacier storage.
● Record filenames, metadata, and archive identifiers in DynamoDB.
● Search DynamoDB, then request restore using the archive identifier.
Design reasoning
● Technical workability: DynamoDB provides online metadata queries while Glacier provides low-cost archival storage.
● Requirement fit: Archive-count overhead falls and users can locate datasets before retrieval.
● Scene fit: Content is rarely accessed but must remain discoverable.
● Engineering common sense: Archive storage is not a search index.
Alternatives to reject
● Glacier vaults cannot be queried directly by filename inside archives.
● Metadata stored only as S3 objects still needs a searchable index.
● Two-bucket designs add complexity without improving archival discovery.
● Keeping all datasets in online S3 storage costs more than the required Glacier approach.
Workflow: Compress dataset → archive → index metadata in DynamoDB → search → retrieve archive on demand.
Continuous VMware Migration with AWS MGN
Physical or VMware servers can move to EC2 through continuous block-level replication and controlled cutover.
Recommended architecture
● Install the AWS Replication Agent on each source server.
● Use AWS Application Migration Service staging resources.
● Continuously replicate disk changes.
● Launch test instances, validate, and initiate cutover.
Design reasoning
● Technical workability: The agent sends block-level changes and MGN converts the replicated server into EC2 launch resources.
● Requirement fit: Testing and final synchronization reduce downtime.
● Scene fit: Existing virtual machines must migrate without application redesign.
● Engineering common sense: Test the same replication stream used for production cutover.
Alternatives to reject
● CloudFormation rebuilds infrastructure but does not replicate VM disks.
● Direct Connect and Service Catalog do not perform VM conversion.
● SAM deploys serverless applications.
● ECS hosts containers, not imported virtual machines.
Workflow: Install agent → replicate to staging → test launch → remediate → final sync → cut over → decommission source.
Troubleshooting SAML Role Federation
A SAML federation flow depends on trust configuration, a valid assertion, correct role mapping, and a properly formed STS request.
Validation steps
● Confirm the IAM role trust policy names the correct SAML provider and permits sts:AssumeRoleWithSAML.
● Check that the identity provider maps users or groups to the intended role.
● Verify the SAML assertion contains the required role and provider information.
● Confirm the AssumeRoleWithSAML call includes the role ARN, provider ARN, and assertion.
Design reasoning
● Technical workability: STS issues temporary credentials only when the assertion and IAM trust relationship agree.
● Requirement fit: These checks isolate failures in the complete federation chain.
● Scene fit: Authentication succeeds at the corporate identity provider but AWS role access fails.
● Engineering common sense: Troubleshoot identity claims and trust before investigating unrelated network services.
Alternatives to reject
● IAM user policies do not repair a SAML role trust failure.
● VPC DNS settings are unrelated to the STS federation decision.
● Placing users in an arbitrary AWS group does not change claims emitted by the external IdP.
Workflow: User → IdP authentication → SAML assertion → STS validation → role session.
Lowest-Cost Peak Capacity with Spot Fleet
Predictable baseline capacity can remain covered by reservations while diversified Spot handles temporary fault-tolerant peaks.
Recommended architecture
● Keep Reserved pricing for steady application demand.
● Launch a diversified Spot Fleet across Availability Zones for peak traffic.
● Use Auto Scaling and interruption-aware replacement.
Design reasoning
● Technical workability: Multiple Spot pools reduce dependence on one capacity source.
● Requirement fit: The design provides the lowest-cost peak capacity.
● Scene fit: The workload can replace interrupted peak instances.
● Engineering common sense: Never treat Reserved Instances as launchable temporary capacity.
Alternatives to reject
● Reserved Instances are billing commitments, not an Auto Scaling fleet type.
● Spot plus On-Demand improves reliability but costs more than the lowest-cost requirement.
● Reserved plus On-Demand preserves higher peak cost.
● One Spot pool increases interruption and capacity risk.
Workflow: Baseline runs → peak alarm → diversified Spot launches → traffic distributes → interrupted capacity replaced → fleet scales down.
Staged Windows Patching for High Uptime
Large Windows fleets should be patched in controlled waves so maintenance does not remove all service capacity at once.
Recommended architecture
● Divide instances into two Patch Groups using tags.
● Associate an approved Patch Baseline with each group.
● Create non-overlapping Systems Manager Maintenance Windows.
● Run AWS-RunPatchBaseline against each group at different start times.
Design reasoning
● Technical workability: Patch Manager identifies, installs, and reports approved patches; Maintenance Windows control execution timing.
● Requirement fit: Separate windows prevent fleet-wide simultaneous reboots.
● Scene fit: Hundreds of production instances need repeatable automation, not individual maintenance.
● Engineering common sense: Keep one serving group available while the other is patched and verified.
Alternatives to reject
● One Patch Group and one window can reboot the entire fleet together.
● CloudWatch scheduling plus custom State Manager commands is possible but indirect and operationally heavier.
● A design without separated windows does not control the disruption boundary.
Workflow: Tag → group → baseline → schedule group A → validate → schedule group B → review compliance.
Expanding an Existing VPC CIDR
A VPC can gain address space by associating secondary IPv4 CIDR blocks.
Recommended procedure
● Select a nonoverlapping secondary CIDR allowed by VPC rules.
● Associate it with the existing VPC.
● Create new subnets from the secondary range.
● Update routes, security rules, network appliances, and IP address management records.
Design reasoning
● Technical workability: AWS supports multiple associated CIDR blocks on one VPC.
● Requirement fit: The company adds addresses without rebuilding the network.
● Scene fit: Existing resources and connectivity should remain in the same VPC.
● Engineering common sense: Check overlap with peered, transit, VPN, and on-premises networks first.
Alternatives to reject
● Deleting subnets does not change the primary VPC CIDR.
● A second VPC plus peering creates a separate network and additional routing operations.
● The primary CIDR does not need to be replaced.
● Overlapping secondary ranges can break connectivity.
Workflow: Plan address space → associate secondary CIDR → create subnets → update controls → migrate or launch workloads.
Auditor Access with CloudTrail
Auditors need authoritative API history plus a dedicated read-only identity for inspecting logs and resources.
Recommended architecture
● Enable AWS CloudTrail across required Regions and global services.
● Deliver logs to a protected S3 bucket.
● Create a dedicated read-only auditor identity.
● Grant access only to required logs and resource metadata.
Design reasoning
● Technical workability: CloudTrail records account activity, while IAM controls evidence access.
● Requirement fit: The auditor can inspect events without changing production resources.
● Scene fit: The request is an external compliance review.
● Engineering common sense: Notification is not evidence; preserve complete logs and controlled search access.
Alternatives to reject
● A read-only role without CloudTrail omits the required activity history.
● SNS delivery notifications are not audit records.
● Denying the auditor log access prevents evidence review.
● AWS does not independently create customer-account permissions for third parties.
Workflow: API action → CloudTrail event → protected S3 log → auditor authenticates → read-only review.
Combining Online and Offline VM Migration
A constrained migration window may require continuous replication for critical systems and offline transfer for very large datasets.
Recommended architecture
● Use AWS Application Migration Service for critical VMs requiring synchronization and low-downtime cutover.
● Use AWS Snowball for bulk data that cannot move over the available network in time.
● Use VM Import/Export for supported image-import cases that do not require continuous replication.
● Coordinate testing and final cutover through one migration plan.
Design reasoning
● Technical workability: MGN replicates server disks, Snowball transports bulk bytes offline, and VM Import/Export creates AWS images from supported formats.
● Requirement fit: Each workload uses the migration path matching its size and downtime tolerance.
● Scene fit: The estate contains both critical virtual machines and large data volumes.
● Engineering common sense: Do not force every workload through one transfer method.
Alternatives to reject
● Provisioning a new Direct Connect may be slow and costly for a one-time migration.
● Refactoring every application exceeds the migration timeline.
● SFTP over limited bandwidth is unsuitable for very large datasets.
● Repeated full image imports do not provide efficient ongoing synchronization.
Workflow: Classify workloads → replicate critical VMs → ship bulk data → import remaining images → test → cut over.
Business-Unit Control of Resource Termination
A multi-account structure should make each business unit responsible for its own resources without granting destructive access to others.
Recommended architecture
● Organize accounts into business-unit OUs with AWS Organizations.
● Create cross-account IAM roles for approved unit administrators.
● Grant resource-level termination permissions only for resources and accounts owned by that unit.
● Use trust policies to restrict who can assume each role.
Design reasoning
● Technical workability: IAM roles and resource permissions authorize operations inside target accounts.
● Requirement fit: Each unit manages its own resources while account boundaries limit blast radius.
● Scene fit: Ownership aligns naturally with separate accounts or OUs.
● Engineering common sense: Use SCPs as maximum guardrails and IAM roles for actual operational access.
Alternatives to reject
● SCPs do not grant resource access and cannot replace target-account IAM permissions.
● A broad organization-wide administrator role violates business-unit isolation.
● Service-linked roles allow AWS services to operate; they are not human administration roles.
● Resource tags alone do not authorize termination unless IAM conditions explicitly use them.
Workflow: Administrator signs in → assumes unit role → IAM evaluates resource scope → permitted termination occurs → CloudTrail records action.
Serverless REST Session Service
State-based REST APIs can use managed request handling, serverless logic, and scalable session storage.
Recommended architecture
● Use API Gateway for REST resources, stages, API keys, and usage controls.
● Run state-transition logic in Lambda.
● Store sessions in DynamoDB with Auto Scaling.
Design reasoning
● Technical workability: API Gateway invokes Lambda, and DynamoDB provides low-operation scalable state.
● Requirement fit: The service supports REST APIs, test stages, and variable demand without server management.
● Scene fit: The existing interface is REST rather than GraphQL.
● Engineering common sense: Choose the API service matching the client contract.
Alternatives to reject
● EC2, NLB, and Aurora add server and database operations.
● AppSync exposes GraphQL rather than the required REST interface.
● ALB can invoke Lambda but lacks API Gateway’s direct API-key, usage-plan, and stage workflow.
● A fixed database capacity conflicts with variable session demand.
Workflow: REST request → API Gateway authorization and stage → Lambda logic → DynamoDB session update.
Monitoring Changes to AWS Organizations
Organization membership and policy changes need both authoritative API history and continuous compliance evaluation.
Recommended architecture
● Create a CloudTrail trail that records AWS Organizations console and API activity.
● Use EventBridge rules to match sensitive events such as account invitations or policy changes.
● Publish alerts through Amazon SNS.
● Use AWS Config for organization-related compliance monitoring where applicable.
Design reasoning
● Technical workability: CloudTrail records who performed an Organizations action; EventBridge reacts to matching events; Config evaluates configuration compliance.
● Requirement fit: Administrators receive timely notification and retain evidence for investigation.
● Scene fit: An external account was added without approval, making identity, time, and API action essential.
● Engineering common sense: A dashboard is not an audit source; collect events first, then visualize or alert.
Alternatives to reject
● Systems Manager does not provide the authoritative Organizations API history.
● Inspector assesses workload vulnerabilities, not account membership changes.
● Control Tower adds governance but does not replace CloudTrail evidence.
● A CloudWatch dashboard cannot detect changes unless a proper event source and rule already exist.
Workflow: Organizations action → CloudTrail record → EventBridge match → SNS notification → Config evaluation → investigation and remediation.
Safer Automated Deployments
A low-downtime delivery process should test code, preview infrastructure effects, and preserve a fast rollback path.
Recommended architecture
● Run automated nonproduction tests with CodeBuild.
● Create CloudFormation change sets before infrastructure updates.
● Use CodeDeploy blue/green deployment for the application.
● Shift traffic only after validation and roll back on failed alarms.
Design reasoning
● Technical workability: CodeBuild executes tests, change sets show planned resource changes, and blue/green keeps the old environment available.
● Requirement fit: The workflow reduces downtime and deployment risk.
● Scene fit: Both application and infrastructure changes occur through CI/CD.
● Engineering common sense: Syntax validation is necessary but cannot prove runtime behavior.
Alternatives to reject
● Template validation alone does not test the application or preview all impacts.
● Manual instance testing is slow and inconsistent.
● Helper scripts and manual QA do not provide automated traffic switching and rollback.
● In-place production updates increase blast radius.
Workflow: Commit → CodeBuild tests → change-set review → deploy green → validate → shift traffic → retain or remove blue.
Global Delivery of a Large Game Package
A static 5 GB game package should use durable object storage and a global CDN rather than file-transfer servers.
Recommended architecture
● Store the package in Amazon S3.
● Create a CloudFront distribution with S3 as the origin.
● Use Route 53 for the public domain.
● Configure caching and origin protection.
Design reasoning
● Technical workability: S3 stores the object durably, and CloudFront scales downloads through edge locations.
● Requirement fit: Global users receive lower-latency delivery with minimal server administration.
● Scene fit: The package is static and identical for every downloader.
● Engineering common sense: Do not operate FTP servers for globally cacheable public files.
Alternatives to reject
● Requester Pays expects authenticated AWS requesters and does not fit anonymous website downloads.
● EC2 FTP, EFS, and NLB require server scaling and higher operations.
● EBS is not shared automatically across an Auto Scaling fleet.
● ALB does not support FTP traffic.
Workflow: User resolves domain → CloudFront edge → cached package or S3 origin fetch → download.
Secure DynamoDB Access from EC2
An EC2 application should obtain temporary DynamoDB permissions through an instance profile rather than embedded access keys.
Recommended architecture
● Create an IAM role trusted by the EC2 service.
● Attach a least-privilege policy for the required DynamoDB table actions.
● Place the role in an instance profile.
● Launch or update application instances with that profile.
● Let the AWS SDK retrieve temporary credentials automatically.
Design reasoning
● Technical workability: EC2 exposes rotating role credentials through Instance Metadata Service to authorized local software.
● Requirement fit: The application accesses DynamoDB without storing static keys in code or configuration.
● Scene fit: A three-tier application performs AWS API calls from its compute tier.
● Engineering common sense: Bind permissions to workload identity and scope them to exact tables and actions.
Alternatives to reject
● Hard-coded API keys can leak and require manual rotation.
● An IAM user cannot be attached to an EC2 instance like a role.
● A DynamoDB resource policy is not the normal replacement for workload role credentials in this design.
● Granting administrator access violates least privilege.
Workflow: EC2 boots with profile → SDK obtains temporary credentials → signs DynamoDB request → IAM evaluates role policy → credentials rotate automatically.
Patching Linux Instances in OpsWorks
OpsWorks Linux instances can receive current packages through native dependency updates or controlled replacement.
Recommended approaches
● Run the OpsWorks Update Dependencies command for current package updates.
● Replace instances so setup and lifecycle recipes install current packages on fresh capacity.
● Patch in waves and validate application health.
Design reasoning
● Technical workability: OpsWorks integrates package update commands and instance lifecycle configuration.
● Requirement fit: Linux servers receive updates with lower risk to serving capacity.
● Scene fit: The application already uses OpsWorks stacks.
● Engineering common sense: Replacement provides a cleaner state, while in-place updates can address urgent packages.
Alternatives to reject
● CloudFormation is not the native dependency-update mechanism for existing OpsWorks instances.
● The command is not the equivalent Windows patch workflow.
● Deleting the entire stack is unnecessarily disruptive.
● AWS WAF filters requests and cannot install OS updates.
Workflow: Select wave → update dependencies or replace instance → validate → continue → review versions.
Modernizing Desktop Applications and MySQL
A managed modernization can stream desktop applications, move MySQL to Aurora, and host supporting web services across Availability Zones.
Recommended architecture
● Use Amazon AppStream 2.0 for centrally managed desktop application streaming.
● Migrate MySQL to Amazon Aurora.
● Run web or application services in a Multi-AZ Auto Scaling group behind an ALB.
Design reasoning
● Technical workability: AppStream streams applications without full desktop management; Aurora is MySQL compatible; ALB and Auto Scaling provide resilient hosting.
● Requirement fit: The design reduces infrastructure operations while improving scale and availability.
● Scene fit: Users need applications rather than complete persistent desktops.
● Engineering common sense: Select the narrowest managed end-user service that meets the access requirement.
Alternatives to reject
● WorkSpaces supplies full desktops and may add cost when only applications are required.
● Self-managed MySQL retains backup and failover work.
● CloudFront does not accelerate interactive desktop protocols.
● Redshift is not an OLTP MySQL replacement.
● ElastiCache cannot deliver desktop applications, and DynamoDB conversion requires major redesign.
Workflow: User launches stream → AppStream session → application reaches ALB services → Aurora stores relational data.
Searchable Archive for Scanned Documents
A scanned archive needs durable object storage, extracted searchable metadata, and a scalable web access layer.
Recommended architecture
● Store scanned files in Amazon S3.
● Index extracted text and metadata in Amazon CloudSearch.
● Host the web application in AWS Elastic Beanstalk.
● Keep search results linked to the authoritative S3 objects.
Design reasoning
● Technical workability: S3 provides durable storage, CloudSearch supports managed indexing and queries, and Elastic Beanstalk manages application deployment and scaling.
● Requirement fit: Users can search and retrieve files without operating custom search clusters.
● Scene fit: The source data consists of scanned files rather than relational transactions.
● Engineering common sense: Store originals separately from derived search indexes so either can be rebuilt safely.
Alternatives to reject
● A file server alone stores documents but does not provide scalable full-text search.
● Self-managed search software on EC2 adds patching, scaling, and recovery work.
● Database keyword fields without proper indexing scale poorly for a large archive.
● Object storage by itself does not make image text searchable.
Workflow: Upload scan → extract text through the existing OCR process → index metadata → query CloudSearch → retrieve original from S3.
Enforcing Cost-Allocation Tags
Accurate cost reports need existing resources corrected and future untagged creation prevented.
Recommended architecture
● Use Tag Editor to add required tags to existing RDS and DynamoDB resources.
● Activate those keys as cost-allocation tags.
● Use an SCP with tag conditions to deny future resource creation when required tags are missing.
Design reasoning
● Technical workability: Tag Editor performs bulk updates, billing exposes activated tags, and SCP conditions create preventive guardrails.
● Requirement fit: Historical resources become reportable and new drift is reduced.
● Scene fit: Cost center and project metadata are mandatory across accounts.
● Engineering common sense: Remediate the past and prevent recurrence.
Alternatives to reject
● Activating billing tags does not create missing resource tags.
● Lambda remediation adds custom code and acts after creation.
● Tagging only current resources does not enforce future behavior.
● AWS Config can detect missing tags but does not prevent creation.
Workflow: Inventory → bulk tag → activate billing keys → apply SCP → test approved and denied provisioning.
Restoring Gateway Block Storage to EBS
A Storage Gateway volume snapshot can become native AWS block storage for an EC2 recovery environment.
Recommended procedure
● Select the Storage Gateway snapshot.
● Restore it as an Amazon EBS volume.
● Create the volume in the required Availability Zone.
● Attach it to the EC2 application server.
● Mount and validate the filesystem.
Design reasoning
● Technical workability: Gateway snapshots integrate with EBS snapshot and volume workflows.
● Requirement fit: The application receives attachable block storage without depending on the on-premises gateway.
● Scene fit: The original workload expects an iSCSI-style block volume.
● Engineering common sense: Preserve the storage interface during disaster recovery before considering later modernization.
Alternatives to reject
● Direct attachment to the on-premises gateway preserves a hybrid dependency.
● S3 is object storage and cannot directly replace a mounted block device.
● Storage Gateway does not convert the volume directly into EFS.
● Copying files manually increases recovery time and may miss metadata.
Workflow: Gateway snapshot → EBS volume → EC2 attachment → filesystem check → application startup.
Controlling Reserved Instance Discount Sharing
Reserved Instance sharing is a billing allocation feature controlled from the AWS Organizations management account.
Recommended procedure
● Keep the business unit inside the organization.
● Disable RI discount sharing for the relevant member account or account scope through consolidated billing preferences.
● Verify how discounts are applied before the planned workload starts.
Design reasoning
● Technical workability: The management account controls whether eligible RI discounts are shared across linked accounts.
● Requirement fit: The purchasing business unit retains the intended billing benefit.
● Scene fit: No change to workloads, IAM, or account ownership is required.
● Engineering common sense: Solve a narrow cost-allocation issue through billing settings rather than restructuring governance.
Alternatives to reject
● RI discounts are not unavoidably shared; the preference can be changed.
● Member accounts do not have a special “private RI” setting.
● Removing the account from the organization would disrupt consolidated billing, policies, and central management.
Operational note: Reserved Instances affect billing benefits, not instance placement or technical capacity. Capacity planning and scaling must still be designed separately.
Workflow: Review ownership → change sharing preference → model effective coverage → monitor billing allocation.
Restricting Origins to CloudFront
CloudFront should be the only public path to both dynamic ALB content and static S3 objects.
Recommended architecture
● Configure CloudFront to add a secret custom header to ALB origin requests.
● Attach AWS WAF to the ALB and reject requests missing the header.
● Create an Origin Access Identity for the S3 origin.
● Update the S3 bucket policy to allow reads only from the OAI.
Design reasoning
● Technical workability: The custom-header check distinguishes CloudFront origin requests, while OAI signs requests to private S3 content.
● Requirement fit: Direct ALB and S3 access is blocked without changing the public CloudFront endpoint.
● Scene fit: One distribution serves dynamic and static paths from different origins.
● Engineering common sense: Protect each origin using controls supported by that origin type.
Alternatives to reject
● S3 ACLs alone provide weaker and less centralized control than a bucket policy for the OAI.
● Network ACLs cannot reliably allow changing CloudFront origin addresses as an application identity.
● WAF attached only to CloudFront filters viewers but does not prove the ALB request came through CloudFront.
● Route 53 does not prevent direct origin access.
Workflow: Viewer → CloudFront → custom-header ALB path or OAI-signed S3 path.
Trusted HTTPS for CloudFront WordPress
A custom WordPress domain should redirect viewers to HTTPS and use trusted ACM certificates for both delivery layers.
Recommended architecture
● Configure HTTP-to-HTTPS redirection in CloudFront.
● Use an ACM certificate for the CloudFront custom domain in the required Region.
● Use a trusted ACM certificate on the origin endpoint in its Region.
● Enforce HTTPS from CloudFront to the origin.
Design reasoning
● Technical workability: ACM integrates with CloudFront and AWS load-balancing origins.
● Requirement fit: Viewers and origin traffic receive trusted encryption with managed renewal.
● Scene fit: The site uses a custom WordPress domain.
● Engineering common sense: Certificate names must match the viewer and origin hostnames.
Alternatives to reject
● Self-signed certificates are not trusted by public browsers or CloudFront origins.
● Third-party certificates can work but add purchase and renewal operations.
● The default CloudFront certificate covers only cloudfront.net, not the custom domain.
Workflow: HTTP viewer → redirect → CloudFront TLS → HTTPS origin request → WordPress response.
PrivateLink Traffic Controls Behind an NLB
PrivateLink clients connect to a Network Load Balancer, so backend controls must account for the NLB-to-target network path.
Recommended controls
● Configure network ACLs to allow required traffic in both directions between the NLB subnets and logging-service subnets.
● Configure the EC2 target security group to allow inbound application traffic from the NLB subnet IP ranges.
● Keep unnecessary ports denied.
Design reasoning
● Technical workability: The NLB accepts endpoint-service connections and forwards them to registered targets.
● Requirement fit: Only the expected load-balancer path reaches the logging instances.
● Scene fit: Consumers use PrivateLink rather than connecting directly to backend EC2 addresses.
● Engineering common sense: Trace the complete packet path before selecting source addresses for firewall rules.
Alternatives to reject
● Client VPC CIDRs are not necessarily the direct source observed by the targets in this architecture.
● A Network Load Balancer in the described model does not use a security group as the primary filter.
● Allowing all private address ranges is broader than required.
● Configuring only one NACL direction fails because network ACLs are stateless.
Workflow: Interface endpoint → endpoint service → NLB subnet → target subnet NACL → EC2 security group → logging service.
Cost-Efficient Availability During Peak Load
A multi-AZ application should maintain enough headroom to survive one AZ loss while matching pricing models to load stability.
Recommended architecture
● Run Auto Scaling across all Availability Zones.
● Maintain recovery headroom for peak conditions.
● Cover steady usage with Reserved pricing.
● Use reliable On-Demand capacity for critical peaks and optional Spot for tolerant work.
● Scale down after demand falls.
Design reasoning
● Technical workability: Auto Scaling replaces capacity in healthy AZs when one AZ fails.
● Requirement fit: The application survives peak traffic without permanently overprovisioning every zone.
● Scene fit: High utilization during a failure leaves little spare capacity.
● Engineering common sense: Availability calculations must use the post-failure capacity, not normal capacity.
Alternatives to reject
● Fixed capacity overprovisions during normal operation.
● Too few instances per AZ cannot absorb peak traffic after failure.
● All-Spot capacity can disappear during interruption.
● A single-AZ Auto Scaling group is not highly available.
Workflow: Baseline covered → peak scaling starts → AZ fails → healthy AZs scale → load redistributes → fleet scales down later.
Long-Term Cost Optimization for a Global Game API
A predictable multi-year application can combine committed compute pricing with edge delivery for static assets.
Recommended architecture
● Run the GraphQL API on appropriately sized EC2 instances.
● Purchase Reserved pricing for the stable baseline over the expected three-year period.
● Place the API fleet behind a suitable load balancer.
● Store static assets in durable origin storage and distribute them through CloudFront.
Design reasoning
● Technical workability: EC2 supports the existing API runtime, and CloudFront serves cached assets from global edge locations.
● Requirement fit: Reserved pricing lowers long-term baseline cost, while edge caching improves worldwide responsiveness.
● Scene fit: The service has an expected life beyond three years and stable core demand.
● Engineering common sense: Commit only the predictable baseline and keep burst capacity flexible.
Alternatives to reject
● All On-Demand capacity ignores the long, predictable service lifetime.
● Serving static files directly from API servers wastes compute and bandwidth.
● Spot-only API capacity risks customer-visible interruption.
● Replatforming to unrelated services is unnecessary if the current API architecture meets functional needs.
Workflow: Client → CloudFront for assets → load balancer for GraphQL → reserved baseline EC2 → scale additional capacity as needed.
Managed Protection from Web Exploits and DDoS
Public applications need application-layer filtering and infrastructure-layer DDoS protection.
Recommended architecture
● Attach AWS WAF to supported web endpoints.
● Use managed or custom rules for SQL injection and common exploits.
● Enable AWS Shield Advanced for enhanced DDoS protection and monitoring.
Design reasoning
● Technical workability: WAF inspects HTTP requests, while Shield Advanced protects supported AWS edge and load-balancing resources.
● Requirement fit: The combination addresses web exploits and distributed attacks.
● Scene fit: The workload is an internet-facing AWS application.
● Engineering common sense: Reactive IP blocking cannot replace controls designed for distributed and changing sources.
Alternatives to reject
● Direct Connect and an on-premises hardware WAF add cost and do not protect the AWS public edge directly.
● One NACL deny blocks only known addresses and cannot inspect SQL payloads.
● A standby Region improves recovery but does not stop current malicious traffic.
● Scaling alone can increase attack cost without filtering requests.
Workflow: Internet → Shield protection → WAF inspection → allowed request → application; alerts initiate response.
Custom Federation from LDAP to IAM Roles
A custom mobile authentication solution can preserve LDAP as the credential source while issuing temporary AWS role credentials.
Supported patterns
● Build a custom OpenID Connect provider and use Cognito Identity Pools to map authenticated identities to IAM roles.
● Or build a SAML-compatible solution that authenticates against LDAP and sends assertions to an IAM SAML identity provider.
Design reasoning
● Technical workability: OIDC or SAML establishes federation; Cognito or IAM exchanges trusted identity information for temporary role sessions.
● Requirement fit: The mobile application uses custom authentication and IAM roles without storing AWS access keys.
● Scene fit: The company must retain LDAP and satisfy strict security controls.
● Engineering common sense: Reuse standard federation protocols rather than inventing a token database and authorization engine.
Alternatives to reject
● IAM Identity Center is better suited to workforce access portals than this direct mobile role-federation flow.
● A custom API Gateway and DynamoDB token service duplicates identity, session, and credential capabilities.
● OIDC combined only with Identity Center does not directly provide the requested mobile IAM-role workflow.
Workflow: User credentials → LDAP validation → OIDC token or SAML assertion → role mapping → temporary AWS credentials.
Private S3 Origins for CloudFront
CloudFront should read private S3 objects through an origin identity while direct public S3 access remains disabled.
Recommended architecture
● Create a CloudFront Origin Access Identity.
● Configure the S3 origin to use it.
● Grant the OAI read permission in the bucket policy.
● Remove public ACLs and other direct-access permissions.
Design reasoning
● Technical workability: CloudFront signs origin requests using the OAI, and S3 authorizes that identity.
● Requirement fit: Users receive objects through CloudFront but cannot bypass the distribution.
● Scene fit: S3 is a private CloudFront origin.
● Engineering common sense: Viewer authorization and origin authorization are separate controls.
Alternatives to reject
● Field-level encryption protects selected request fields, not S3 origin access.
● A TLS certificate does not authorize object retrieval.
● Signed URLs can restrict viewers, but origin protection still requires removing direct S3 permissions.
● Public-read access defeats the private-origin design.
Workflow: Viewer → CloudFront → OAI-signed request → S3 bucket policy → object response.
Publishing an S3 Static Website
An S3 website endpoint must be able to read the website objects. DNS routing alone does not grant object access.
Recommended configuration
● Enable static website hosting on the correctly named S3 bucket.
● Configure the index document.
● Point the Route 53 alias record to the regional S3 website endpoint.
● Allow public read access to the website objects through the required bucket policy and public-access settings.
Design reasoning
● Technical workability: The website endpoint serves the configured index object when bucket authorization permits anonymous reads.
● Requirement fit: Visitors can retrieve public website content through the custom domain.
● Scene fit: The site is intentionally public and is hosted directly from the S3 website endpoint.
● Engineering common sense: Check the full request chain: DNS resolution, endpoint selection, bucket policy, public-access block, and object existence.
Alternatives to reject
● A custom error document is optional and does not control index availability.
● Visitors do not need to append index.html when an index document is configured.
● Route 53 propagation is normally fast and waiting does not repair an authorization failure.
Workflow: Resolve domain → reach website endpoint → evaluate bucket access → return index object and referenced assets.
Monitoring Approved AMIs with AWS Config
AWS Config can evaluate running instances against an approved AMI list and notify teams about noncompliance.
Recommended architecture
● Configure the approved-amis-by-id managed rule.
● Provide the authorized AMI IDs.
● Send noncompliance notifications through SNS.
● Remediate through the approved replacement process.
Design reasoning
● Technical workability: Config evaluates the AMI associated with running EC2 instances.
● Requirement fit: The organization receives ongoing compliance visibility without blocking development launches.
● Scene fit: The concern is image approval, not software vulnerability scanning.
● Engineering common sense: Detecting a bad image and scanning an image for CVEs are different controls.
Alternatives to reject
● Inspector scans vulnerabilities, not membership in an approved AMI list.
● Preventive SCP and IAM restrictions can block development.
● CloudWatch does not natively evaluate approved AMIs.
● Trusted Advisor has no direct approved-AMI compliance check.
Workflow: Instance configuration recorded → Config rule evaluates AMI → noncompliant result → SNS notification → replacement.
Direct Connect with VPN Failover
A cost-conscious hybrid design can use one dedicated primary circuit and an internet VPN for backup.
Recommended architecture
● Use Direct Connect as the preferred path.
● Create a managed Site-to-Site VPN as redundant connectivity.
● Configure route preferences so Direct Connect is primary.
● Test withdrawal and failover behavior.
Design reasoning
● Technical workability: Dynamic routing selects the VPN when the Direct Connect route disappears.
● Requirement fit: The primary path is predictable, while backup cost remains lower than a second dedicated circuit.
● Scene fit: The company accepts lower backup performance during a failure.
● Engineering common sense: A backup must not depend on the same failed physical circuit.
Alternatives to reject
● One Direct Connect leaves a physical single point of failure.
● VPN CloudHub joins VPN sites and does not provide a dedicated primary path.
● A VPN carried on the same Direct Connect still fails with that circuit.
● Static routing can slow or complicate automatic failover.
Workflow: Normal traffic uses Direct Connect → circuit route withdrawn → VPN route selected → service continues.
Retaining Contest Data After Stack Deletion
CloudFormation deletion policies should preserve stored objects and create a database backup while removing ongoing database compute cost.
Recommended configuration
● Apply DeletionPolicy: Retain to the S3 bucket.
● Apply DeletionPolicy: Snapshot to the RDS resource.
● Validate policy changes before deleting the stack.
Design reasoning
● Technical workability: CloudFormation leaves the bucket and objects intact and creates a final RDS snapshot before deleting the database.
● Requirement fit: Contest images and recoverable database data survive stack termination.
● Scene fit: The application is no longer active, so the live database need not continue running.
● Engineering common sense: Preserve data through native lifecycle controls before building duplicate copy workflows.
Alternatives to reject
● S3 does not support Snapshot as a deletion policy.
● Retaining RDS keeps the running instance and its charges.
● Deleting both resources destroys required data.
● Cross-Region replication can create another copy but adds storage and operations not requested.
Workflow: Update template → inspect change → delete stack → verify retained bucket → verify RDS snapshot → manage retained assets separately.
Scaling License-Bound EC2 Applications
Applications licensed to network-interface identities need a controlled pool of reusable ENIs and centrally assigned license files.
Recommended architecture
● Maintain a pool of ENIs that preserve license-bound identities.
● Store license files securely in S3.
● Use bootstrap logic to assign one unused ENI and license during scale-out.
● Use Lambda to maintain current database addresses in Parameter Store.
● Retrieve current addresses during instance bootstrap.
Design reasoning
● Technical workability: ENIs preserve network identity across instance replacement, while Parameter Store separates changing configuration from images.
● Requirement fit: The fleet scales without duplicating licenses or baking stale database addresses.
● Scene fit: Licensing is tied to ENI identity.
● Engineering common sense: Coordinate allocation atomically so two instances cannot claim one license.
Alternatives to reject
● One AMI with one license cannot scale safely.
● Resolving DNS once at boot still leaves stale local addresses.
● Baking every license into the AMI risks duplicate use.
● Static database IPs fail after address changes.
Workflow: Scale-out → claim ENI and license → read Parameter Store → configure app → release assets on termination.
Dedicated Hybrid Connectivity with Direct Connect
A private, dedicated connection from a VPC to internal services uses Direct Connect and BGP-capable customer routing.
Recommended architecture
● Provision AWS Direct Connect.
● Configure the required private virtual interface and VPC attachment architecture.
● Use an on-premises router supporting BGP.
● Configure MD5 authentication for the BGP session where required.
Design reasoning
● Technical workability: Direct Connect provides dedicated transport, and BGP exchanges routes dynamically.
● Requirement fit: The path avoids dependence on public internet bandwidth.
● Scene fit: Internal services require stable private hybrid connectivity.
● Engineering common sense: Add redundant locations or VPN backup when the connection is business critical.
Alternatives to reject
● Internet Gateway plus VPN is encrypted but not dedicated bandwidth.
● A transit VPC still relies on VPN transport.
● Elastic IP is public addressing, not private connectivity.
● Static routes alone do not satisfy the required Direct Connect routing session.
Workflow: On-premises router → BGP over Direct Connect → private VIF → AWS gateway → VPC routes.
Reducing Database Read Latency
Repeated read traffic should be absorbed by cache and database replicas before considering sharding or engine replacement.
Recommended architecture
● Deploy ElastiCache in the required Availability Zones.
● Cache frequently requested records and query results.
● Add RDS read replicas and route read-only traffic to them.
Design reasoning
● Technical workability: ElastiCache provides in-memory responses, and replicas distribute database reads.
● Requirement fit: Announcement traffic receives lower latency with limited application change.
● Scene fit: The launch creates a read-heavy spike.
● Engineering common sense: Remove repeated work before increasing database size.
Alternatives to reject
● Sharding requires major code and operations changes.
● Keyspaces migration changes the data model unnecessarily.
● Larger instances and more IOPS scale vertically.
● CloudFront helps static content but not every dynamic database query.
● One cache node without AZ resilience can become a failure point.
Workflow: Application checks cache → hit returns → miss reads replica → updates cache → writes remain on primary.
Heterogeneous Database Migration: Oracle to PostgreSQL
A heterogeneous database migration has two distinct jobs: convert database objects, then move and synchronize the data.
Recommended architecture
● Use AWS Schema Conversion Tool to assess and convert Oracle schema and code.
● Resolve objects that require manual conversion.
● Create the target Amazon RDS for PostgreSQL database.
● Use AWS Database Migration Service for full load and ongoing change replication.
Design reasoning
● Technical workability: SCT converts structurally different database objects; DMS transports records between source and target engines.
● Requirement fit: The workflow supports migration with limited downtime and purpose-built tooling.
● Scene fit: Oracle and PostgreSQL are different engines, so schema conversion must precede data cutover.
● Engineering common sense: Do not expect a data movement service to redesign stored procedures or incompatible schemas.
Alternatives to reject
● SAM and Lambda are application-development tools, not database conversion services.
● Server Migration Service moves servers rather than transforming database engines.
● DMS alone does not complete heterogeneous schema and code conversion.
● Data Pipeline, CodeCommit, and Batch could support custom scripts but add unnecessary engineering.
Workflow: Assess → convert schema → remediate code → create target → full load → replicate changes → validate → cut over.
AWS CDK for Python and TypeScript Teams
AWS CDK lets teams define infrastructure in familiar languages while standardizing deployment on CloudFormation.
Recommended architecture
● Build CDK applications in Python and TypeScript.
● Synthesize CloudFormation templates.
● Run validation and synthesis in CodeBuild.
● Orchestrate promotion through CodePipeline.
Design reasoning
● Technical workability: CDK supports both languages and produces native CloudFormation deployments.
● Requirement fit: Teams retain their coding skills while sharing one infrastructure workflow.
● Scene fit: Existing deployment logic is split between Python and TypeScript.
● Engineering common sense: Standardize the deployment engine without forcing every team into one programming language.
Alternatives to reject
● Rewriting scripts manually into templates adds unnecessary conversion.
● EC2 user data is not an infrastructure provisioning engine.
● OpsWorks and Chef introduce a different configuration model.
● A third-party tool adds dependencies when CDK already fits both languages.
Workflow: Write CDK → test → synthesize → inspect template → deploy through pipeline.
Bulk Redshift Migration over Limited Bandwidth
A 60 TB initial transfer cannot meet a 30-day deadline over a 50 Mbps allocation, so bulk data and ongoing changes should use different paths.
Recommended architecture
● Use AWS SCT to assess the Oracle warehouse and prepare Redshift-compatible extraction and schema.
● Export the bulk dataset through an AWS Snowball Edge import job.
● Import the device into Amazon S3.
● Load the staged data into Amazon Redshift.
● Use AWS DMS to replicate ongoing daily changes until cutover.
Design reasoning
● Technical workability: Snowball moves the initial dataset offline; DMS carries the much smaller change stream over the network.
● Requirement fit: The approach meets the deadline without consuming business bandwidth or purchasing a permanent high-capacity link.
● Scene fit: Daily changes are minor, while the initial warehouse is extremely large.
● Engineering common sense: Separate bulk seeding from incremental synchronization.
Alternatives to reject
● Staging 60 TB through RDS Oracle adds cost and does not solve initial transfer limits.
● A Snowball workflow without reliable ongoing replication risks missing daily updates.
● Provisioning Direct Connect and self-managed Oracle RAC is expensive, slow, and outside the migration need.
Workflow: Convert → bulk export → ship → S3 load → DMS changes → reconcile → cut over.
The Two Essential EC2 IAM Role Policies
An EC2 role needs a permissions policy and a trust policy.
Required policies
● Attach a permissions policy granting the required S3 object actions.
● Configure the role trust policy so the EC2 service principal can assume the role.
● Attach the role through an instance profile.
Design reasoning
● Technical workability: The trust policy establishes who can assume the role, and the permissions policy defines what the session can do.
● Requirement fit: The Node.js application receives temporary S3 access.
● Scene fit: EC2 is the workload identity.
● Engineering common sense: Trust and permissions answer different authorization questions.
Alternatives to reject
● A bucket policy may also be needed, but it does not replace the role’s own permission and trust chain.
● The application is not an IAM service principal.
● EC2 assumes its instance role; it does not first assume a separate S3 role.
● Static IAM user keys are unnecessary.
Workflow: EC2 assumes role → metadata provides temporary credentials → role permissions authorize S3 request.
Cross-Account Resource Policies and Continuous Audit
Cross-account sharing is simplest when permissions are attached directly to resources that support resource-based policies.
Recommended architecture
● Add approved external account principals to S3, KMS, and OpenSearch resource policies.
● Limit actions, resources, and conditions to the required access path.
● Use AWS Config to record configuration changes and evaluate compliance.
Design reasoning
● Technical workability: Resource policies can trust principals from another AWS account.
● Requirement fit: External users retain their normal identity permissions and do not need to replace them with a role session for this access model.
● Scene fit: The named services support resource-level policy controls.
● Engineering common sense: Define the sharing boundary on the shared resource and continuously verify it.
Alternatives to reject
● An identity policy in the external account alone cannot authorize access to another account’s resources.
● SCPs define organization guardrails; they do not grant access to individual shared resources.
● Service-linked roles are for AWS service integrations, not this user-sharing pattern.
● Systems Manager is not the configuration-compliance recorder.
Workflow: Identify principals → write least-privilege resource policies → test access → record with Config → alert on noncompliance.
Managed Windows Desktops and Applications
A managed virtual desktop service can provide Windows access while reducing server maintenance and application-distribution work.
Recommended architecture
● Use Amazon WorkSpaces for managed Windows desktops.
● Use Amazon WorkSpaces Application Manager for controlled application delivery where applicable.
● Enable automatic Windows updates and defined maintenance windows.
● Apply directory, network, and access controls centrally.
Design reasoning
● Technical workability: WorkSpaces manages desktop infrastructure, while WAM packages and assigns applications.
● Requirement fit: Users receive managed Windows environments with lower administration.
● Scene fit: The need is secure desktop access, not a web development environment or bastion host.
● Engineering common sense: Use desktop services for desktops and keep administrative jump access as a separate security design.
Alternatives to reject
● Lightsail is not the managed enterprise desktop solution and lacks the described OS upgrade workflow.
● AppSync is a GraphQL service.
● Cloud9 is a development environment, not a hardened Windows desktop platform.
● AppStream streams applications but is not a bastion host.
Workflow: User authenticates → WorkSpaces session launches → assigned applications delivered → updates applied during maintenance.
Connecting SNS and SQS in CloudFormation
Infrastructure as code should create the notification topic, queue, and subscription using resource references rather than hard-coded identifiers.
Recommended template behavior
● Define an AWS::SNS::Topic resource.
● Define an AWS::SQS::Queue resource.
● Define an SNS subscription with protocol sqs.
● Use Fn::GetAtt to retrieve the queue ARN for the subscription endpoint.
● Allow the SNS topic to send messages through the SQS queue policy.
Design reasoning
● Technical workability: SNS publishes notifications and SQS durably buffers them for consumers.
● Requirement fit: CloudFormation creates the relationship repeatably across environments.
● Scene fit: The queue endpoint is generated during stack deployment.
● Engineering common sense: Reference deployed resource attributes instead of copying ARNs manually.
Alternatives to reject
● A queue URL is not the SNS subscription endpoint when an SQS ARN is required.
● Creating only the topic and queue does not subscribe them.
● Omitting the queue policy can prevent SNS delivery.
● Hard-coded ARNs reduce portability and may target the wrong account or Region.
Workflow: Deploy topic → deploy queue → resolve queue ARN → create subscription → authorize topic → test message delivery.
Accelerating Snowball Transfers of Small Files
Millions of small files can underuse Snowball bandwidth because per-file encryption and metadata work becomes the bottleneck.
Recommended approach
● Run multiple parallel copy sessions to the Snowball Edge device.
● Distribute source directories across workers.
● Keep each worker’s file list independent to avoid duplicate writes.
● Monitor aggregate throughput and device utilization.
● Validate transferred file counts and checksums.
Design reasoning
● Technical workability: Parallel clients overlap per-file processing and increase aggregate throughput.
● Requirement fit: The migration finishes faster without changing the dataset or purchasing another transfer method.
● Scene fit: The issue is many small files, not insufficient raw device capacity.
● Engineering common sense: Increase concurrency only until source disks, network, CPU, or device limits become saturated.
Alternatives to reject
● Compressing everything into one archive changes handling and may require extra temporary space and processing.
● A faster WAN does not improve a local copy bottleneck to the device.
● Copying files sequentially preserves the per-file overhead problem.
● Reformatting the device is unnecessary and risks restarting work.
Workflow: Partition file list → start parallel copy workers → monitor throughput → retry failures → reconcile file count → return device.
HPC Networking and Parallel Storage
Tightly coupled simulations need low-latency node communication, close physical placement, and a high-throughput shared filesystem.
Recommended architecture
● Use Elastic Fabric Adapter on supported EC2 instances.
● Place compute nodes in a cluster placement group within one Availability Zone.
● Use Amazon FSx for Lustre for parallel access to thousands of large simulation files.
Design reasoning
● Technical workability: EFA supports HPC communication, cluster placement minimizes network latency, and FSx for Lustre provides parallel filesystem throughput.
● Requirement fit: The combined design accelerates inter-node communication and shared data access.
● Scene fit: The workload is tightly coupled rather than independent batch processing.
● Engineering common sense: Optimize the full workflow; fast compute networking cannot compensate for serial storage.
Alternatives to reject
● RAID 0 EBS is local to an instance and is not a shared parallel filesystem.
● Spreading nodes across AZs increases latency.
● Multiple ENIs do not aggregate bandwidth like EFA.
● Per-instance disks complicate shared simulation data and recovery.
Workflow: Scheduler places nodes together → EFA carries inter-node traffic → FSx for Lustre supplies shared files → results persist.
Automatic Scale-In at Low CPU
Auto Scaling can remove excess EC2 capacity when a CloudWatch metric remains below a safe threshold.
Recommended architecture
● Create a CloudWatch alarm for CPU at or below 15 percent.
● Attach the alarm to an Auto Scaling scale-in policy.
● Configure cooldowns or instance warmup behavior.
● Preserve minimum capacity and availability constraints.
Design reasoning
● Technical workability: The alarm invokes the scaling policy directly when its evaluation criteria are met.
● Requirement fit: Unneeded instances terminate automatically and reduce cost.
● Scene fit: Capacity should follow measured utilization, not a fixed schedule.
● Engineering common sense: Avoid reacting to one short low-CPU sample.
Alternatives to reject
● Manual email-driven removal is slow.
● Lambda notification code is unnecessary for a native scaling action.
● Scheduled scaling follows time rather than actual demand.
● Terminating instances outside the Auto Scaling group can disrupt desired capacity management.
Workflow: CPU remains low → alarm enters ALARM → scale-in policy reduces desired capacity → Auto Scaling terminates safely.
Temporary S3 Credentials for Image Processing
An EC2 image-processing worker should use an IAM role to access S3 without handling long-lived credentials.
Recommended architecture
● Create an EC2 IAM role with exact read and write permissions for the required buckets and prefixes.
● Attach the role through an instance profile.
● Let the application SDK retrieve temporary credentials from Instance Metadata Service.
● Upload watermarked images to the authorized destination.
Design reasoning
● Technical workability: Instance-profile credentials rotate automatically and are used by standard AWS SDK credential providers.
● Requirement fit: The worker can download and upload photos without embedded access keys.
● Scene fit: Processing occurs on an EC2 instance that interacts with S3.
● Engineering common sense: Scope permissions separately for source reads and destination writes.
Alternatives to reject
● Hard-coded keys can leak into images, logs, AMIs, or source repositories.
● IAM users cannot be attached directly to EC2 instances.
● Public bucket permissions expose every object unnecessarily.
● Passing credentials from the mobile client to the worker breaks trust boundaries.
Workflow: EC2 receives job → SDK obtains role credentials → downloads source → watermarks image → uploads result → credentials rotate.
Routing Around Overlapping Peered VPC CIDRs
VPC peering is non-transitive, uses static routes, and applies longest-prefix matching when routes overlap.
Recommended routing
● In VPC A, add a host route such as 10.0.0.77/32 to the VPC B peering connection.
● Add the broader 10.0.0.0/16 route to the VPC C peering connection.
● Configure reciprocal routes in VPC B and VPC C.
● Ensure security groups and network ACLs permit the traffic.
Design reasoning
● Technical workability: The /32 route is more specific than /16, so traffic for the required VPC B host follows the correct peering connection.
● Requirement fit: The design reaches both destinations despite overlapping address ranges.
● Scene fit: Only a specific database address is required in one overlapping VPC.
● Engineering common sense: Prefer route specificity for a limited exception, but avoid overlapping CIDRs in new designs.
Alternatives to reject
● Peering does not exchange routes dynamically.
● Network ACLs cannot repair an incorrect route decision.
● A broad overlapping route to both peers is ambiguous and cannot provide full connectivity to both CIDR spaces.
Workflow: Destination lookup → longest-prefix match → selected peering connection → reciprocal route → security evaluation.
Third-Party Audit Through an Assumable Role
An external auditor should receive temporary cross-account access limited to the exact audit actions required.
Recommended architecture
● Create a cross-account IAM role in the audited account.
● Trust the auditor’s controlled AWS principal.
● Attach read-only, resource-scoped permissions.
● Require STS role assumption and record activity in CloudTrail.
Design reasoning
● Technical workability: STS returns temporary credentials without creating permanent users in the audited account.
● Requirement fit: Access is revocable, attributable, and least privilege.
● Scene fit: A third party needs temporary review access.
● Engineering common sense: Audit does not justify unrestricted administration.
Alternatives to reject
● Full access violates least privilege.
● Long-term IAM user keys are harder to rotate and revoke safely.
● Even limited IAM user keys remain inferior to temporary role sessions.
● Sharing existing employee credentials destroys accountability.
Workflow: Auditor authenticates in home account → assumes audit role → reviews evidence → CloudTrail records session → credentials expire.
SAML Federation for Hybrid Workforce Access
Employees can use corporate identities for AWS access through standards-based SAML federation.
Recommended architecture
● Configure the enterprise identity provider for SAML 2.0.
● Create the corresponding SAML provider and IAM roles in AWS.
● Map users or groups to approved roles.
● Exchange the SAML assertion through AWS federation and STS endpoints for temporary sessions.
Design reasoning
● Technical workability: AWS validates the signed assertion and role mapping before issuing temporary credentials.
● Requirement fit: Users retain corporate authentication and avoid separate long-lived IAM passwords.
● Scene fit: The company operates a hybrid environment with an existing enterprise identity source.
● Engineering common sense: Centralize authentication at the IdP and keep AWS authorization in IAM roles.
Alternatives to reject
● Creating IAM users duplicates password lifecycle and offboarding.
● Web identity federation targets consumer OIDC providers rather than this workforce SAML scenario.
● LDAP alone is not directly accepted by STS without a broker or federation layer.
● Network connectivity such as VPN does not provide identity federation.
Workflow: User authenticates to IdP → receives SAML assertion → selects mapped role → AWS validates assertion → temporary console or API session.
SCP Inheritance and Temporary Account Onboarding
An explicit deny inherited from the organization root cannot be overridden by an allow policy on a child organizational unit.
Recommended architecture
● Move production restrictions from the organization root to the Production OU.
● Create a temporary Onboarding OU with permissions needed to configure AWS Config.
● Place the new account in Onboarding while required controls are installed.
● Move the account to Production after validation.
Design reasoning
● Technical workability: SCP evaluation respects inherited denies; relocating the deny changes where it applies.
● Requirement fit: The onboarding account can be configured without weakening existing production accounts.
● Scene fit: The exception is temporary and limited to one newly acquired account.
● Engineering common sense: Place controls at the narrowest level that matches their intended scope.
Alternatives to reject
● Removing the root restriction for everyone creates an organization-wide exposure.
● An allow SCP on Onboarding cannot override a root-level explicit deny.
● A root allow list can work but increases long-term maintenance when new services are required.
● Service Catalog does not override SCP evaluation.
Workflow: Create OUs → relocate production SCP → onboard account → configure Config rules → verify compliance → move account to Production.
Active-Active DNS with Target Health
Public users can be routed to the lowest-latency healthy regional endpoint through Route 53.
Recommended architecture
● Create latency-based alias records for regional Elastic Load Balancers.
● Enable Evaluate Target Health.
● Keep both Regions active and capable of serving requests.
● Monitor application and data-layer health.
Design reasoning
● Technical workability: Route 53 chooses the lower-latency endpoint and removes unhealthy alias targets from answers.
● Requirement fit: The design supports active-active regional traffic and automatic endpoint avoidance.
● Scene fit: Public applications run behind load balancers in multiple Regions.
● Engineering common sense: DNS health routing works only when each Region has complete application dependencies.
Alternatives to reject
● A private hosted zone cannot route public users.
● DNSSEC protects DNS integrity but does not perform health-based failover.
● Disabling target-health evaluation weakens automatic recovery.
● A transit VPC connects networks and does not route public clients.
Workflow: DNS query → latency and health evaluation → regional ELB answer → client connects → unhealthy target removed.
Reducing RDS Failover Time
Fast database recovery requires resilient database targets and application connection management.
Recommended architecture
● Use RDS Proxy to pool connections and route applications to the healthy target.
● Migrate to Aurora MySQL when faster replica failover is required.
● Deploy Aurora replicas across Availability Zones.
● Test client retry behavior.
Design reasoning
● Technical workability: RDS Proxy reduces connection churn, while Aurora replicas provide promotion targets.
● Requirement fit: Applications recover more quickly after writer failure.
● Scene fit: The target failover is below standard DNS and reconnection behavior.
● Engineering common sense: Database promotion and client reconnection both contribute to outage time.
Alternatives to reject
● Optimized Writes, Optimized Reads, and caching improve performance, not failover.
● Standard read replicas do not provide automatic writer failover.
● Redis does not create a writable relational database after failure.
● Larger instances do not shorten failover.
Workflow: Writer fails → Aurora promotes replica → proxy routes pooled connections → applications retry → service resumes.
Operating Oracle RAC on EC2
When Oracle RAC must remain self-managed, automate infrastructure backups and operating-system patching around the cluster.
Recommended architecture
● Deploy the supported Oracle cluster on Amazon EC2.
● Use Amazon Data Lifecycle Manager for scheduled EBS snapshots.
● Use Systems Manager Patch Manager for approved operating-system patches.
● Coordinate patch waves with cluster failover and maintenance procedures.
Design reasoning
● Technical workability: DLM automates EBS snapshot policies, while Patch Manager installs and reports OS updates.
● Requirement fit: The design preserves RAC while reducing repetitive backup and patch operations.
● Scene fit: Managed RDS does not provide the required RAC architecture.
● Engineering common sense: Cloud automation does not remove the need to understand database quorum and node order.
Alternatives to reject
● RDS Multi-AZ provides managed availability but is not Oracle RAC.
● AWS Backup alone does not install OS patches.
● Manual snapshots and patch scripts increase missed schedules and inconsistent nodes.
● Replacing RAC with read replicas changes the availability and write architecture.
Workflow: Snapshot according to DLM policy → verify recovery point → drain one node → patch and test → continue through cluster nodes.
Preventing Production EC2 Termination
Developer access should permit normal work while explicitly preventing termination of production instances.
Recommended controls
● Remove production termination permission from developer roles.
● Add an explicit IAM deny for ec2:TerminateInstances when the target has the production tag.
● Protect tag-management permissions so developers cannot remove the control tag.
Design reasoning
● Technical workability: An explicit deny overrides allows, and resource-tag conditions scope the restriction.
● Requirement fit: Developers retain safe actions on nonproduction resources.
● Scene fit: The main risk is destructive API access to production instances.
● Engineering common sense: Least privilege is stronger than adding approval friction to an unnecessarily broad permission.
Alternatives to reject
● PowerUserAccess still permits many delete operations.
● Security groups affect network traffic, not EC2 API authorization.
● MFA conditions may reduce accidents but still permit termination.
● EC2 termination protection alone is not a complete IAM boundary if users can modify the setting.
Workflow: Developer calls termination → IAM evaluates role and production tag → explicit deny blocks request → CloudTrail records attempt.
Task-Level Security for ECS Microservices
Container security should isolate network access and AWS permissions at the task level rather than sharing host-level controls.
Recommended architecture
● Use awsvpc network mode in the ECS task definition.
● Assign security groups directly to ECS tasks.
● Use IAM task roles for access to AWS services.
● Grant each microservice only the actions and resources it requires.
Design reasoning
● Technical workability: Each task receives an elastic network interface, private IP address, and security-group controls.
● Requirement fit: Task roles and task security groups implement least privilege for both network and API access.
● Scene fit: Strict security policy requires standard network monitoring and controls at container level.
● Engineering common sense: Do not give every container the broad permissions of its EC2 host.
Alternatives to reject
● Bridge mode applies security groups at the host, not individual tasks.
● EC2 instance roles can expose broader permissions to multiple services on the host.
● Passing IAM credentials into containers creates long-lived secret risk.
● Moving to App Runner does not justify credentials in environment variables and changes the platform unnecessarily.
Workflow: Define task role → select awsvpc → attach task security group → deploy → inspect flow and application logs → refine permissions.
Cost-Efficient One-Time EMR Cluster
A one-time analytics cluster should protect control and data nodes while using interruptible capacity only for restartable work.
Recommended architecture
● Run EMR master and core nodes on On-Demand Instances.
● Run task nodes on Spot Instances.
● Store durable input and output outside disposable task nodes.
Design reasoning
● Technical workability: Master and core nodes maintain cluster control and HDFS data; task nodes perform repeatable computation.
● Requirement fit: Spot reduces cost without exposing critical cluster state to interruption.
● Scene fit: The cluster runs once, so long-term commitments are inefficient.
● Engineering common sense: Use cheap interruptible capacity only where lost work can be retried.
Alternatives to reject
● Reserved Instances require a commitment that does not fit one execution.
● Reserving only the master still wastes commitment cost.
● Spot master or core nodes can terminate the cluster or lose HDFS data.
● On-Demand task nodes miss the best cost-saving opportunity.
Workflow: Launch stable master and core → add Spot task fleet → process data → persist outputs → terminate cluster.
Two-Region Relational Application Resilience
A read-heavy application can use one highly available writer Region and secondary-Region readers while keeping the application tier active in both Regions.
Recommended architecture
● Deploy Auto Scaling application tiers in both Regions.
● Use an Amazon Aurora Global Database.
● Keep writes in the primary Region and use in-Region endpoints for reads.
● Route users by geography and configure health-checked failover to the healthy Region.
Design reasoning
● Technical workability: Aurora Global Database replicates relational data across Regions with one primary writer and regional readers.
● Requirement fit: Regional application fleets provide resilience, while the database preserves relational semantics.
● Scene fit: North American and Asian users benefit from regional reads and controlled failover.
● Engineering common sense: Avoid two independent writable databases unless conflict resolution is explicitly designed.
Alternatives to reject
● Independent writable RDS MySQL databases do not become safe active-active masters through simple replication.
● Multi-AZ protects against an AZ failure, not a complete Region failure.
● Snapshots are recovery artifacts, not an active regional-read solution.
● Multivalue routing may return multiple healthy endpoints rather than implement the intended regional failover policy.
Workflow: Write primary → replicate globally → read locally → monitor health → fail traffic to surviving Region.
Migrating Physical Servers with AWS MGN
Physical servers can move to EC2 through continuous replication, testing, and controlled cutover.
Recommended architecture
● Install the AWS Replication Agent on every source server.
● Configure AWS Application Migration Service staging resources.
● Replicate block-level changes continuously.
● Launch tests, remediate, and initiate cutover.
Design reasoning
● Technical workability: MGN turns replicated server disks into bootable EC2 launch resources.
● Requirement fit: Testing and final synchronization reduce migration downtime.
● Scene fit: Complete physical server workloads must move, not only files.
● Engineering common sense: Validate network, boot, and application dependencies before final cutover.
Alternatives to reject
● Application Discovery Service inventories servers but does not migrate them.
● DataSync moves files and does not create bootable server images.
● Outposts does not provide this replication-to-AMI workflow.
● Manual rebuilds create configuration drift.
Workflow: Install agent → replicate → launch test → validate → final sync → cut over → decommission source.
LDAP Federation with STS Credentials
LDAP authenticates corporate users, while IAM roles and STS authorize temporary access to AWS resources.
Recommended patterns
● Let the application validate credentials against LDAP, map the user to an IAM role, and call STS.
● Or use a custom identity broker that performs LDAP authentication and requests scoped federated credentials.
Design reasoning
● Technical workability: STS issues temporary role credentials after a trusted application or broker validates identity.
● Requirement fit: Users keep corporate credentials without permanent IAM users.
● Scene fit: The application already depends on LDAP.
● Engineering common sense: Separate authentication from AWS authorization and keep sessions short-lived.
Alternatives to reject
● STS does not authenticate LDAP passwords directly.
● IAM cannot natively validate corporate LDAP usernames and passwords.
● Long-lived access keys duplicate identity lifecycle and increase credential exposure.
● Network connectivity alone does not provide federation.
Workflow: User login → LDAP validation → role mapping → STS request → temporary credentials → signed AWS request.
Reducing Hybrid Database Load with Redis
A dynamic portal can reduce on-premises database pressure by caching sessions and frequently repeated query results.
Recommended architecture
● Deploy Amazon ElastiCache for Redis.
● Cache session state and suitable query results.
● Configure replication for cache availability and read scaling.
● Set expiration based on acceptable staleness.
Design reasoning
● Technical workability: Redis provides low-latency in-memory reads and data structures suitable for sessions.
● Requirement fit: Fewer requests cross the hybrid link or reach the database.
● Scene fit: The portal and comments are dynamic, so full static hosting is not viable.
● Engineering common sense: Cache only data the application can safely reconstruct after eviction.
Alternatives to reject
● S3 static hosting cannot replace dynamic portal and comment processing.
● Migrating to Aurora may work but is a larger database project than caching.
● SCT is unnecessary for a compatible database move.
● OpenSearch is a search engine, not a transactional store for comments and sessions.
Workflow: Application reads cache → hit returns immediately → miss queries database → result cached with TTL.
Low-Latency UDP Game Networking
UDP game traffic needs a Layer 4 load balancer and subnet controls that can explicitly reject unwanted protocols.
Recommended architecture
● Use a Network Load Balancer with UDP listeners.
● Assign static Elastic IP addresses where required.
● Create Route 53 records for the NLB.
● Use NACL deny rules to restrict unwanted non-UDP traffic.
Design reasoning
● Technical workability: NLB supports UDP and high-throughput low-latency transport; NACLs support explicit denies.
● Requirement fit: The service obtains stable addressing and protocol-appropriate scaling.
● Scene fit: The game protocol is UDP, not HTTP.
● Engineering common sense: Choose security and load-balancing controls that operate at the protocol layer used.
Alternatives to reject
● CloudFront focuses on HTTP and HTTPS delivery.
● AWS WAF inspects HTTP payloads, not UDP.
● ALB does not support UDP listeners.
● Security groups are allow-only and cannot express explicit denies.
Workflow: Player resolves Route 53 → NLB Elastic IP → UDP listener → healthy game target; NACL rejects unwanted traffic.
Detecting and Removing Unapproved AMIs
An agile pipeline can allow launches while automatically identifying and remediating instances built from unauthorized images.
Recommended patterns
● Use AWS Config to detect unapproved AMI IDs, then invoke Lambda to alert and terminate noncompliant instances.
● Or run a scheduled Lambda that inspects instance AMI IDs, notifies Security, and terminates unauthorized instances.
Design reasoning
● Technical workability: EC2 exposes source AMI IDs, and Lambda can evaluate and terminate instances.
● Requirement fit: CI/CD is not blocked, but unauthorized capacity is short-lived.
● Scene fit: The policy favors detective and corrective control over preventive denial.
● Engineering common sense: Preserve evidence before termination and protect the approved-image list.
Alternatives to reject
● Manual approval delays delivery.
● Amazon Inspector scans vulnerabilities but does not decide whether an AMI is approved.
● Preventive IAM restrictions conflict with the requirement not to stall launches.
● Notification without remediation leaves noncompliant workloads running.
Workflow: Instance launches → AMI evaluated → compliant instance remains; noncompliant instance logged, alerted, and terminated.
Private MySQL Replication to On-Premises
An on-premises read copy can follow RDS MySQL through native external replication over a private VPN.
Recommended architecture
● Establish private VPN connectivity.
● Seed the on-premises MySQL target with mysqldump.
● Configure RDS as the external replication source.
● Start native MySQL replication and monitor lag.
Design reasoning
● Technical workability: RDS supports replication to an external MySQL-compatible target under supported configurations.
● Requirement fit: The on-premises database remains current for reads.
● Scene fit: Continuous replication is required, not nightly batch export.
● Engineering common sense: Seed first, then apply the change stream from a known position.
Alternatives to reject
● An EC2 relay adds an unnecessary replication hop.
● Nightly Data Pipeline exports are not a current read replica.
● Replication over the open internet increases exposure.
● Repeated full exports waste bandwidth and extend lag.
Workflow: VPN → initial dump → capture replication coordinates → start external replica → monitor and repair lag.
Keeping TLS Keys in CloudHSM
A sensitive TLS private key should remain inside hardware security modules while application logs remain durable and access controlled.
Recommended architecture
● Use TCP load balancing so TLS reaches the application and HSM integration without load-balancer termination.
● Use AWS CloudHSM for private-key operations.
● Deploy HSM capacity for availability.
● Store encrypted logs in a private S3 bucket.
● Limit log decryption through IAM and key policies.
Design reasoning
● Technical workability: CloudHSM performs cryptographic operations without exporting private-key material, and S3 provides durable encrypted logging.
● Requirement fit: Administrators cannot casually copy the TLS key, while authorized users can access logs.
● Scene fit: The hybrid application has strict key-custody and audit requirements.
● Engineering common sense: Key storage, TLS execution, and log storage are separate security concerns.
Alternatives to reject
● Uploading the key to web servers makes it movable.
● Instance-store logs are not durable.
● A single HSM creates an availability risk.
● Standard load-balancer TLS offload does not meet a requirement for customer-controlled HSM custody.
Workflow: Client TLS → TCP load balancer → HSM-backed operation → application → encrypted S3 logs → authorized audit access.
Automated RDS Password Rotation
Secrets Manager can generate, store, and rotate an RDS password through native CloudFormation resources.
Recommended architecture
● Create the database secret in Secrets Manager.
● Let CloudFormation connect the secret to RDS.
● Define a RotationSchedule that invokes the rotation Lambda every 90 days.
● Give the application runtime permission to retrieve the secret.
Design reasoning
● Technical workability: Secrets Manager coordinates password updates and secret value changes.
● Requirement fit: Rotation is automated and the password is not embedded in templates.
● Scene fit: The secret is an RDS database credential.
● Engineering common sense: Applications must refresh connections after rotation.
Alternatives to reject
● Parameter Store has no equivalent native database RotationSchedule resource.
● KMS key rotation rotates encryption-key material, not the stored database password.
● EventBridge plus custom code duplicates native rotation and adds risk.
● Hard-coded credentials become stale after rotation.
Workflow: Generate secret → deploy RDS → rotation schedule triggers Lambda → password changes → application retrieves current value.
Immediate IP Blocking and Managed DDoS Protection
Block a known hostile source at the subnet boundary and use dedicated protection for broader DDoS attacks.
Recommended architecture
● Add an explicit deny for the offending CIDR to the relevant network ACL.
● Use AWS Shield Advanced for enhanced DDoS protection on supported resources.
Design reasoning
● Technical workability: Network ACLs are stateless and support explicit deny rules; Shield Advanced provides managed DDoS detection and response capabilities.
● Requirement fit: The source is blocked immediately, and the environment gains broader protection against common infrastructure attacks.
● Scene fit: Port scanning targets EC2 resources inside VPC subnets.
● Engineering common sense: Use a network control that can deny traffic before it reaches hosts.
Alternatives to reject
● Security groups are allow-only and cannot express an explicit source deny.
● Route 53 is not an IP firewall.
● Macie discovers sensitive S3 data, not DDoS traffic.
● Host firewalls require per-instance changes and may be reached only after traffic consumes network resources.
● GuardDuty and Patch Manager improve detection and hygiene but do not provide the requested immediate block and DDoS protection.
Workflow: Identify CIDR → deny in NACL → validate impact → enable managed protection → monitor findings.
Meeting a Weekend Migration Deadline with S3 Sync
A one-week head start allows most data to move before the final weekend, leaving only changed files for cutover.
Recommended approach
● Run an initial S3 sync one week before migration.
● Continue normal source operations.
● Run a final incremental sync on Friday.
● Validate counts and checksums.
● Point the EC2 application to the completed S3 dataset.
Design reasoning
● Technical workability: S3 sync transfers only missing or changed objects on the final pass.
● Requirement fit: The 1 TB migration gains schedule margin and can finish by Sunday.
● Scene fit: The dataset can be copied while the source remains online.
● Engineering common sense: Seed early and reserve the outage window for deltas and validation.
Alternatives to reject
● Copying the full dataset during the weekend leaves little margin.
● Gateway snapshot workflows add unnecessary steps for this file movement.
● Starting Saturday risks missing the deadline.
● Snowball ordering, shipping, import, and restore cannot reliably complete in one weekend.
Workflow: Initial sync → source changes → final delta sync → validate → application cutover.
How SSE-S3 Protects Objects
Amazon S3 server-side encryption with S3-managed keys encrypts every object using an individual data key.
Encryption model
● S3 generates a unique data key for each object.
● The object is encrypted with that data key.
● The data key is encrypted under an S3-managed root or master key.
● S3 rotates the master-key material regularly and manages the full key lifecycle.
● Decryption occurs transparently for authorized requests.
Design reasoning
● Technical workability: Envelope encryption limits the scope of each object data key while centralizing protection of those keys.
● Requirement fit: Data is encrypted at rest without customers storing or rotating encryption keys.
● Scene fit: The requirement is managed S3 encryption, not customer-controlled key policy or audit separation.
● Engineering common sense: Encryption does not replace IAM, bucket policies, versioning, or transport security.
Incorrect assumptions to reject
● One plaintext key is not reused directly for every object.
● Customers do not download the S3 master key.
● SSE-S3 is different from SSE-KMS, where customers can control KMS permissions and audit key use.
● Server-side encryption does not automatically make a public bucket private.
Workflow: Authorized PUT → S3 creates data key → encrypts object → wraps key → stores ciphertext → decrypts on authorized GET.
Scalable Hybrid Block Storage with Cached Volumes
Applications requiring an iSCSI block interface can use a small local cache while storing the dataset durably in AWS.
Recommended architecture
● Deploy AWS Storage Gateway Volume Gateway in cached-volume mode.
● Present iSCSI block volumes to on-premises servers.
● Keep frequently accessed blocks in local cache storage.
● Store the primary volume data in Amazon S3 through the managed gateway service.
● Use snapshots for recovery and AWS-side restoration.
Design reasoning
● Technical workability: Cached Volumes preserve block access while reducing the amount of local storage required.
● Requirement fit: The design scales a very large dataset and maintains frequent data locally.
● Scene fit: The existing application expects block storage rather than an object API.
● Engineering common sense: Preserve the interface required by the application while moving durable capacity to cloud storage.
Alternatives to reject
● Stored Volumes require the complete primary dataset to remain on premises.
● Direct S3 access is object based and does not present an iSCSI block device.
● S3 Glacier is archival object storage and cannot be mounted as an active block volume.
● File Gateway presents file protocols, not the required block interface.
Workflow: Application I/O → iSCSI volume → local cache → gateway-managed S3 backing → snapshot when required.
Protecting RDS from Flash-Sale Writes
A sudden sale can overwhelm a relational database unless submissions are durably buffered and drained at a controlled rate.
Recommended architecture
● Place sale submissions in Amazon SQS.
● Invoke Lambda consumers from the queue.
● Limit concurrency to a rate RDS can sustain.
● Use retries, idempotency, and a dead-letter queue.
Design reasoning
● Technical workability: SQS absorbs bursts, and Lambda processes messages asynchronously.
● Requirement fit: Customer submissions are preserved while database load remains controlled.
● Scene fit: The event is temporary and does not justify a database redesign.
● Engineering common sense: Durable queues protect writes; caches do not.
Alternatives to reject
● Vertical scaling requires planned changes and still couples traffic directly to RDS.
● Memcached can evict or lose entries and is not a durable write queue.
● Migrating PostgreSQL to DynamoDB is a major scope change.
● Synchronous retries can amplify overload.
Workflow: Customer submits → message enters SQS → Lambda consumes within concurrency limit → transaction writes to RDS → acknowledge success.
Virtual Tape Archives with Tape Gateway
Existing tape backup software can move to cloud-backed virtual tapes without changing its operating model.
Recommended architecture
● Deploy Storage Gateway Tape Gateway.
● Present the virtual tape library to existing backup software.
● Keep active virtual tapes in S3-backed storage.
● Eject tapes to the Virtual Tape Shelf for Glacier-class archival.
Design reasoning
● Technical workability: Tape Gateway emulates a tape library and integrates with common backup applications.
● Requirement fit: The company preserves current workflows while reducing physical tape operations.
● Scene fit: The source process already uses tape backup semantics.
● Engineering common sense: Choose the gateway type matching the interface expected by the backup software.
Alternatives to reject
● Virtual Tape Shelf is the archive location, not a generic S3 point-in-time backup.
● Stored Volume Gateway provides iSCSI block volumes, not tape emulation.
● File Gateway exposes file protocols and does not emulate a tape library.
● Rewriting the backup application is unnecessary.
Workflow: Backup writes virtual tape → active tape stored → eject → archived in Virtual Tape Shelf → retrieve when required.
Safe Enterprise Application Deployment
A complex application needs full infrastructure definition, controlled traffic movement, and immutable capacity replacement.
Recommended architecture
● Define the complete stack with CloudFormation.
● Use CodeDeploy blue/green for controlled production traffic shifting.
● Use Elastic Beanstalk immutable deployment for managed application capacity where applicable.
● Monitor health and retain a rollback target.
Design reasoning
● Technical workability: CloudFormation orchestrates resources, while blue/green and immutable deployment avoid modifying the serving fleet in place.
● Requirement fit: Releases minimize downtime and rollback risk.
● Scene fit: The stack includes DynamoDB, Lambda, OpenSearch, and Beanstalk resources.
● Engineering common sense: Infrastructure and application versions should be promoted together through tested stages.
Alternatives to reject
● SAM is optimized for serverless applications and is not the best complete model for this mixed stack.
● In-place deployment increases outage and rollback risk.
● Lightsail lacks the required enterprise orchestration.
● Manual cutovers create inconsistent recovery.
Workflow: Commit → build and test → update stack → deploy replacement capacity → shift traffic → monitor or roll back.
Moving Kafka Events into AWS
A hybrid Kafka pipeline needs reliable network connectivity and controlled ingestion into AWS streaming services.
Recommended architecture
● Establish AWS Direct Connect for predictable hybrid transport.
● Run EC2 consumers that read the on-premises Kafka topics.
● Publish consumed records to Amazon Kinesis for AWS processing.
● Use API Gateway WebSocket APIs with Lambda only for client-facing real-time interactions that require them.
Design reasoning
● Technical workability: Kafka consumers preserve the source protocol, Direct Connect provides capacity, and Kinesis supports scalable AWS consumers.
● Requirement fit: Existing Kafka remains usable while events become available to AWS workloads.
● Scene fit: The design bridges on-premises streaming and cloud-based real-time applications.
● Engineering common sense: Separate backend stream ingestion from browser or mobile WebSocket delivery.
Alternatives to reject
● API Gateway is not a direct Kafka broker replacement.
● Lambda cannot continuously poll arbitrary Kafka endpoints without supported connectivity and event-source design.
● S3 batch transfer loses real-time behavior.
● VPN alone may not meet the required sustained bandwidth and predictability.
Workflow: Kafka topic → EC2 consumer over Direct Connect → Kinesis stream → processors → optional Lambda and WebSocket updates.
Private S3 Access from Approved EC2 Instances
Secure S3 access requires network-path restriction and workload authorization.
Recommended architecture
● Create an S3 gateway VPC endpoint.
● Restrict the endpoint policy to the required bucket.
● Attach a least-privilege IAM role to approved EC2 instances.
● Require the endpoint and approved principal in the S3 bucket policy.
Design reasoning
● Technical workability: The endpoint keeps S3 traffic private, IAM grants positive authorization, and the bucket policy enforces the expected path.
● Requirement fit: Only the web portal instances can use the bucket through the approved VPC route.
● Scene fit: The application runs on EC2 in private networking.
● Engineering common sense: Network location is not a substitute for workload identity.
Alternatives to reject
● NACLs cannot identify an individual EC2 workload for S3 authorization.
● A bucket policy cannot securely treat a private subnet as the principal.
● SourceIp is unreliable when endpoint routing changes the observed path.
● Per-instance routes do not grant IAM permissions.
Workflow: EC2 role signs request → gateway endpoint policy → bucket policy path and principal checks → S3 action.
Separating TLS Certificate Control from Developers
Security can control private keys while developers manage application servers by terminating TLS at the load balancer.
Recommended architecture
● Store the certificate in AWS Certificate Manager or the supported AWS certificate store.
● Restrict certificate permissions to the Security team.
● Attach the certificate to the ALB HTTPS listener.
● Keep private-key material off EC2 instances.
Design reasoning
● Technical workability: The ALB performs TLS termination without distributing the private key to application hosts.
● Requirement fit: Security retains certificate lifecycle control while developers operate EC2.
● Scene fit: Teams require strict separation of duties.
● Engineering common sense: Host-level file permissions are weak when developers have broad server administration.
Alternatives to reject
● Storing the certificate on EC2 exposes it to privileged host users.
● Retrieving the private key from S3 breaks separation.
● Copying a key out of CloudHSM defeats HSM isolation.
● Application owners should not manage production certificate renewal.
Workflow: Security manages certificate → ALB terminates TLS → forwarded request reaches EC2 without private key exposure.
Redshift Regional Disaster Recovery
Redshift disaster recovery uses snapshots copied to another Region rather than continuous cluster replication.
Recommended architecture
● Enable automatic Redshift snapshots.
● Configure cross-Region snapshot copy to the recovery Region.
● Set retention to satisfy the 24-hour RPO.
● Test restoring a cluster within the one-hour RTO.
Design reasoning
● Technical workability: Native snapshot copy places recoverable warehouse data outside the source Region.
● Requirement fit: Automatic snapshots meet the data-loss target and regional copies protect against Region failure.
● Scene fit: The workload is Redshift, not a transactional database with a cross-Region replica.
● Engineering common sense: Recovery timing must include cluster restore and application reconnection.
Alternatives to reject
● Manual snapshot copying begins too late during a disaster.
● S3 Cross-Region Replication is not the Redshift cluster failover mechanism.
● Redshift has no continuous cross-Region cluster replication design here.
● Automatic snapshots remain in the source Region unless copy is enabled.
Workflow: Automatic snapshot → cross-Region copy → disaster → restore cluster → validate → reconnect analytics.
Tamper-Resistant Multi-Region API Audit Logging
Compliance logging requires complete API history across Regions and global services, plus protected long-term storage.
Recommended architecture
● Create an AWS CloudTrail multi-Region trail.
● Include global service events such as IAM activity.
● Deliver logs to a dedicated Amazon S3 bucket.
● Encrypt log files with AWS KMS.
● Restrict bucket and key access with least-privilege policies.
● Enable versioning and MFA Delete where operationally appropriate.
Design reasoning
● Technical workability: CloudTrail records console, SDK, CLI, and API activity across supported services.
● Requirement fit: Multi-Region coverage and global events provide the required audit history; S3 and KMS protect durability and confidentiality.
● Scene fit: EC2, S3, CloudFront, and IAM activity spans regional and global service boundaries.
● Engineering common sense: Centralize audit evidence separately from workload accounts and prevent easy deletion or alteration.
Alternatives to reject
● Excluding global service events omits required IAM activity.
● CloudWatch does not offer a “trail” resource for authoritative account API history.
● A single-Region configuration cannot cover activity in all application Regions.
Workflow: API call → CloudTrail event → encrypted S3 object → controlled audit access → retention and integrity monitoring.
Scaling Web Sessions and Aurora Reads
A stateful web application needs Layer 7 load balancing, elastic compute, and scalable database readers.
Recommended architecture
● Run web instances in an Auto Scaling group behind an ALB.
● Enable sticky sessions only when the application requires them.
● Add Aurora replicas and use Aurora Auto Scaling for read capacity.
● Keep writes directed to the writer endpoint.
Design reasoning
● Technical workability: ALB supports HTTP routing and stickiness; Aurora Auto Scaling adds replicas based on load.
● Requirement fit: Web demand and database reads scale independently.
● Scene fit: Sessions are stateful and read traffic varies.
● Engineering common sense: Aurora Auto Scaling scales replicas, not the writer.
Alternatives to reject
● NLB lacks the stated Layer 7 routing and sticky-session behavior.
● Aurora Auto Scaling does not scale the master database.
● Scaling only EC2 leaves read pressure on the database.
● Sticky sessions can reduce balancing efficiency, so external session storage is preferable when feasible.
Workflow: Client → ALB sticky routing → Auto Scaling instance → reader endpoint for reads → writer for updates.
Regional Recovery with Aurora Global Database
Low RPO and RTO across Regions require a continuously replicated database and health-based traffic movement.
Recommended architecture
● Use Aurora Global Database with a primary Region and secondary cluster.
● Deploy application capacity in both Regions.
● Configure Route 53 health checks and failover routing.
● Test managed regional promotion and application reconnection.
Design reasoning
● Technical workability: Aurora replicates storage changes across Regions with low latency and supports secondary promotion.
● Requirement fit: The design reduces data loss and recovery time compared with snapshot restoration.
● Scene fit: The application requires a relational database in two Regions.
● Engineering common sense: Database recovery and application traffic failover must be tested together.
Alternatives to reject
● RDS Multi-AZ remains inside one Region.
● A standard cross-Region read replica can work, but Aurora Global Database is more direct for fast regional recovery.
● Manual EC2 snapshot restoration adds operations and delays.
● Backups without ready application capacity extend RTO.
Workflow: Primary serves writes → global replication → health failure → promote secondary → Route 53 sends traffic to recovery Region.
Secure Highly Available Web Application
A production web application needs elastic ingress, private administration, database failover, edge acceleration, and web filtering.
Recommended architecture
● Use an ALB with an EC2 Auto Scaling group.
● Administer instances through Systems Manager Session Manager.
● Use RDS Multi-AZ.
● Add CloudFront and AWS WAF.
Design reasoning
● Technical workability: ALB and Auto Scaling distribute traffic, Session Manager removes inbound SSH, RDS provides standby failover, and CloudFront plus WAF improves delivery and protection.
● Requirement fit: The design addresses availability and security with managed services.
● Scene fit: The workload is an internet-facing EC2 application.
● Engineering common sense: Secure administration should not introduce another public server.
Alternatives to reject
● Direct SSH preserves key and port exposure.
● Single-AZ RDS remains a database failure point.
● A bastion can work but adds operations compared with Session Manager.
● Shield Standard alone does not provide the complete web-filtering architecture.
Workflow: User → CloudFront and WAF → ALB → EC2; administrators → Session Manager; data → Multi-AZ RDS.
Regional Recovery with Frequent Logs
A two-hour RTO and ten-minute RPO require recovery data already present in the disaster-recovery Region.
Recommended architecture
● Create hourly application or database backups in S3.
● Export transaction logs every five minutes.
● Enable S3 Cross-Region Replication to the recovery Region.
● Automate restore and log replay.
Design reasoning
● Technical workability: The base backup plus frequent logs reconstructs state near the failure time.
● Requirement fit: Five-minute log capture fits the RPO, and regional copies support the RTO.
● Scene fit: The requirement covers complete Region loss.
● Engineering common sense: Backup frequency and backup location solve different risks.
Alternatives to reject
● Regional EBS backups do not provide recovery data in another Region.
● Glacier restoration can exceed the RTO.
● Multi-AZ protects against an AZ failure, not regional disaster.
● Backups without tested restore automation may still miss the RTO.
Workflow: Hourly backup → five-minute logs → cross-Region replication → restore base → replay logs → validate.
Global Semi-Structured Data with DynamoDB
A multi-Region application needing local low-latency reads and writes can use DynamoDB global tables with managed regional compute.
Recommended architecture
● Deploy Fargate application services in each Region behind regional ALBs.
● Store semi-structured records in DynamoDB global tables.
● Use Global Accelerator to route users to healthy nearby ALBs.
Design reasoning
● Technical workability: Global tables provide multi-active replication; Fargate removes server management; Global Accelerator provides health-aware anycast routing.
● Requirement fit: Users receive regional writes and reads with high availability.
● Scene fit: The data is semi-structured and globally active.
● Engineering common sense: Align the database’s replication model with the application’s active-active behavior.
Alternatives to reject
● Aurora is relational and less natural for this data model.
● EC2 adds server operations.
● DocumentDB is regional in the described design.
● S3 replication is asynchronous object copying, not low-latency record access.
Workflow: User → Global Accelerator → regional ALB → Fargate → local DynamoDB replica → global replication.
Smoothing Smart-Meter Writes to DynamoDB
Write-throughput errors require more database capacity or a buffer that smooths bursts before writes reach DynamoDB.
Recommended architecture
● Increase or auto scale DynamoDB write capacity.
● Ingest meter events through Amazon Kinesis.
● Let Lambda consumers batch and write records to DynamoDB.
● Monitor throttling, iterator age, and consumed capacity.
Design reasoning
● Technical workability: Additional capacity removes the direct bottleneck; Kinesis durably buffers burst traffic for controlled consumption.
● Requirement fit: Meter events continue arriving without immediate write loss.
● Scene fit: Devices produce a high-volume stream with bursty throughput.
● Engineering common sense: Scale the constrained resource and decouple producers from downstream write speed.
Alternatives to reject
● More Lambda memory does not fix DynamoDB throttling when processing already succeeds.
● FIFO ordering and deduplication are unnecessary unless the business requires them and can constrain throughput.
● Reducing device reporting changes product behavior rather than fixing ingestion architecture.
● Retrying without buffering can amplify throttling.
Workflow: Meter event → Kinesis shard → Lambda batch → DynamoDB write → scale capacity from observed demand.
Durable Newspaper Search and OCR Modernization
A digital newspaper archive needs durable image storage, global delivery, scalable search, and a managed replacement for expiring OCR software.
Recommended architecture
● Store scanned PNG files in Amazon S3.
● Deliver images through Amazon CloudFront.
● Run the web application in a Multi-AZ Elastic Beanstalk environment.
● Index searchable content with Amazon CloudSearch.
● Extract text from scanned newspapers with Amazon Textract.
Design reasoning
● Technical workability: Textract performs document OCR, CloudSearch supports text queries, and S3 plus CloudFront provides durable global content delivery.
● Requirement fit: Managed services reduce administration while supporting availability and growth.
● Scene fit: The archive contains scanned documents that must become searchable.
● Engineering common sense: Separate durable source images, extracted text, search indexing, and web delivery.
Alternatives to reject
● Rekognition is not the primary OCR service for document text extraction.
● Glacier conflicts with immediate retrieval and does not solve OCR.
● Self-managed EC2, EBS, NGINX, and search software add operations and weaker shared-storage scaling.
Workflow: Ingest scan → store in S3 → extract text with Textract → index results → search through application → deliver images through CloudFront.
Burst HPC Storage with FSx for Lustre
A monthly 200 TB modeling job needs economical durable storage and a temporary parallel filesystem for intense processing.
Recommended architecture
● Keep source data in S3 Intelligent-Tiering.
● Create FSx for Lustre for the monthly compute window.
● Lazy-load only required S3 objects.
● Run parallel compute against the filesystem.
● Export results and delete the temporary filesystem afterward.
Design reasoning
● Technical workability: FSx for Lustre integrates with S3 and provides high-throughput parallel file access.
● Requirement fit: Intelligent-Tiering reduces idle storage cost, and temporary FSx avoids month-long filesystem charges.
● Scene fit: The job runs for a short monthly period but requires extreme parallel I/O.
● Engineering common sense: Separate durable storage from temporary performance storage.
Alternatives to reject
● EBS Multi-Attach has AZ and instance constraints and is not a 200 TB shared fleet filesystem.
● Glacier retrieval charges and restore behavior are poor for monthly bulk compute.
● EFS can work for general shared files but is less suited to this parallel modeling burst.
Workflow: S3 source → create FSx → lazy-load → compute → export results → delete FSx.
Scoped Mobile S3 Access with STS
A mobile application should receive temporary credentials after authenticating the user.
Recommended architecture
● Authenticate the user in the application identity system.
● Map the identity to an IAM role.
● Use STS to issue temporary, scoped credentials.
● Restrict access to the user’s authorized S3 objects.
Design reasoning
● Technical workability: Mobile SDKs can sign S3 requests with renewable STS sessions.
● Requirement fit: Users access their data without permanent embedded AWS keys.
● Scene fit: Direct mobile object access needs per-user isolation.
● Engineering common sense: Treat distributed clients as untrusted secret-storage environments.
Alternatives to reject
● Long-term IAM user keys can be extracted and reused.
● STS credentials are temporary and must not be treated as permanent embedded secrets.
● Storing user information in DynamoDB does not by itself map identities to safe AWS roles.
● One shared credential cannot isolate users.
Workflow: User login → role mapping → STS session → scoped S3 request → credential refresh or expiry.
Central Workforce Access with IAM Identity Center
A large multi-account environment needs centralized permission sets and federation with the existing corporate directory.
Recommended architecture
● Use AWS Organizations for the account hierarchy.
● Enable IAM Identity Center.
● Establish the required trust or integration with corporate Active Directory.
● Assign users and groups to permission sets across accounts.
Design reasoning
● Technical workability: Identity Center creates centrally managed account assignments and federated sessions.
● Requirement fit: Employees use existing credentials with lower administration across hundreds of accounts.
● Scene fit: The design is workforce SSO, not customer identity.
● Engineering common sense: Centralize identity and permission assignment rather than maintaining role mappings account by account.
Alternatives to reject
● Custom AD FS regex and role mapping require high manual effort.
● AD Connector can proxy authentication but does not replace centralized multi-account permission sets.
● A branded custom portal still requires integration and role-management code.
● Separate IAM users duplicate lifecycle management.
Workflow: Employee authenticates → directory trust → Identity Center assignment → temporary account session.
Improving Global Login Latency at the Edge
Global authentication performance can be improved by executing suitable request logic near users and providing origin failover for server errors.
Recommended architecture
● Use Lambda@Edge for authentication-related processing that is safe to execute at CloudFront edge locations.
● Configure a CloudFront origin group with primary and secondary origins.
● Fail over when the primary returns selected errors, including relevant gateway failures.
Design reasoning
● Technical workability: Lambda@Edge runs during CloudFront request or response events; origin failover redirects eligible requests to a backup origin.
● Requirement fit: The combination reduces distance-related login delay and limits the effect of HTTP 504 failures with lower cost than full global duplication.
● Scene fit: The application is already serverless and serves users worldwide.
● Engineering common sense: Move only suitable edge logic, and retain origins for operations requiring authoritative backend state.
Alternatives to reject
● Multiple VPCs and a transit VPC do not directly accelerate authentication.
● Full multi-Region application deployment can work but costs more.
● Increasing cache TTL helps static objects, not slow dynamic login processing, and authentication responses generally should not be broadly cached.
Workflow: Viewer request → edge authentication step → primary origin → automatic failover on configured errors.
Protecting Lambda Capacity with Reserved Concurrency
A high-volume replication function should have an explicit concurrency boundary so it cannot consume all regional Lambda capacity.
Recommended architecture
● Configure reserved concurrency for the replication function.
● Size the reservation from downstream capacity and required throughput.
● Monitor the Lambda Throttles, ConcurrentExecutions, duration, and error metrics.
● Create a CloudWatch alarm for sustained throttling.
Design reasoning
● Technical workability: Reserved concurrency guarantees capacity for the function and also caps its maximum concurrent executions.
● Requirement fit: Other Lambda functions retain capacity while replication load remains controlled.
● Scene fit: One function can burst heavily and affect unrelated workloads in the same Region.
● Engineering common sense: Protect both shared platform capacity and slower downstream systems.
Alternatives to reject
● Increasing the function timeout does not bound concurrency.
● SQS can buffer events but does not by itself enforce the function’s regional concurrency share.
● Retry backoff reduces repeated failures but does not guarantee isolation from capacity exhaustion.
● Provisioned concurrency improves startup readiness but is not the direct maximum-concurrency boundary requested.
Workflow: Event arrives → concurrency allocation checked → invoke within limit → throttle excess → alarm on sustained pressure → tune capacity.
Public HTTPS Termination with ACM and ALB
A public website behind an Application Load Balancer can use a trusted AWS Certificate Manager certificate for simple, low-cost HTTPS.
Recommended architecture
● Request a public ACM certificate for the university domain.
● Complete DNS validation.
● Attach the certificate to the ALB HTTPS listener.
● Redirect HTTP requests to HTTPS.
● Forward decrypted traffic to the application targets as required by the design.
Design reasoning
● Technical workability: ALB integrates directly with ACM and performs TLS termination.
● Requirement fit: Public ACM certificates are trusted, managed, and do not add certificate charges.
● Scene fit: The learning system already uses an ALB in front of EC2 instances.
● Engineering common sense: Terminate TLS at the managed load balancer unless end-to-end application encryption is explicitly required.
Alternatives to reject
● ACM Private CA certificates are not publicly trusted by default and introduce private-CA cost.
● Installing certificates on every EC2 instance adds renewal and deployment work.
● Self-signed certificates trigger browser trust failures.
● Purchasing and manually rotating third-party certificates is workable but less operationally efficient.
Workflow: Request certificate → validate domain → configure HTTPS listener → redirect HTTP → test trust and renewal.
Serverless Call Recording Processing
Call recordings should move from expensive primary storage into durable object storage, automated transcription, and lifecycle-based archive.
Recommended architecture
● Store recordings in Amazon S3.
● Trigger Lambda when a new recording arrives.
● Submit audio to Amazon Transcribe.
● Store transcripts and metadata in S3.
● Host a lightweight search or retrieval portal from S3 where suitable.
● Transition old recordings to S3 Glacier through lifecycle rules.
Design reasoning
● Technical workability: S3 events drive serverless processing, Transcribe converts speech to text, and lifecycle policies automate archival.
● Requirement fit: The design lowers operations and storage cost while making recordings searchable.
● Scene fit: Processing occurs per uploaded call and does not require always-on servers.
● Engineering common sense: Keep originals, derived transcripts, and archive lifecycle distinct.
Alternatives to reject
● EC2 transcription workers add capacity and patching work.
● Manual archive jobs are slower and less reliable than lifecycle policies.
● Keeping all recordings in premium active storage wastes cost.
● A relational database is not the correct primary store for large audio objects.
Workflow: Upload recording → Lambda trigger → Transcribe job → transcript to S3 → portal access → lifecycle archive.
Streaming IoT Data to an Analytics Warehouse
IoT telemetry needs managed streaming delivery, durable retention, archival lifecycle, and transformed warehouse analytics.
Recommended architecture
● Ingest records with Kinesis Data Firehose.
● Deliver raw data to S3.
● Use lifecycle policies to archive older data to Glacier classes.
● Process data with EMR and load curated results into Redshift.
Design reasoning
● Technical workability: Firehose buffers and delivers streams, S3 stores them durably, and EMR plus Redshift supports analytics.
● Requirement fit: The pipeline handles continuous events and long-term cost control.
● Scene fit: Devices emit an ongoing stream.
● Engineering common sense: Preserve raw data so transformations can be replayed.
Alternatives to reject
● Direct S3 ingestion lacks the managed streaming buffer.
● Athena does not accept streaming events or store them in DynamoDB.
● DynamoDB and Data Pipeline add write cost and scheduled orchestration.
● Glacier is an archive destination, not the ingestion layer.
Workflow: Device stream → Firehose → S3 raw zone → lifecycle archive → EMR transform → Redshift analytics.
Highly Available Three-Tier Web Architecture
A high-traffic web application needs elastic application capacity, an HTTP-aware load balancer, and a managed relational database with high availability.
Recommended architecture
● Run EC2 instances in an Auto Scaling group across multiple Availability Zones.
● Place an Application Load Balancer in front of the web tier.
● Use Amazon Aurora MySQL across Availability Zones.
● Use Route 53 alias records for the application endpoint.
● Retain the database if the infrastructure stack is deleted.
Design reasoning
● Technical workability: Auto Scaling and ALB distribute large HTTP workloads; Aurora provides managed availability and scalable storage.
● Requirement fit: The architecture handles peak users without paying for a second Region.
● Scene fit: The existing .NET application and MySQL model can migrate without a major redesign.
● Engineering common sense: Use the simplest multi-AZ design that meets the stated availability target.
Alternatives to reject
● An all-Spot application tier can lose too much capacity during interruptions.
● Network Load Balancer is less suitable than ALB for normal HTTP application routing.
● A two-Region application and cross-Region database replica add unnecessary cost and operational complexity.
Workflow: Route 53 → ALB → Auto Scaling web tier → Aurora writer and replicas.
Managed REST Backend for Mobile Content
A mobile REST service should use managed authentication, serverless APIs, scalable metadata storage, and direct object upload.
Recommended architecture
● Use Amazon API Gateway with AWS Lambda.
● Authenticate users with Amazon Cognito.
● Store application metadata in DynamoDB.
● Store files in Amazon S3.
● Issue pre-signed S3 URLs for authorized uploads and downloads.
Design reasoning
● Technical workability: API Gateway invokes Lambda, Cognito supplies identity, DynamoDB scales records, and pre-signed URLs provide time-limited object access.
● Requirement fit: The design scales without managing servers or proxying large objects through Lambda.
● Scene fit: REST operations manage metadata while users exchange files.
● Engineering common sense: Keep large payload transfer on S3 and keep API functions focused on authorization and business logic.
Alternatives to reject
● EC2 web fleets add patching and scaling work.
● Storing large files in DynamoDB is expensive and unnecessary.
● Embedding permanent AWS keys in the mobile app is insecure.
● Routing all file bytes through Lambda adds latency, cost, and runtime constraints.
Workflow: User authenticates → API authorizes request → Lambda creates pre-signed URL → client transfers object directly to S3.
Cost-Optimized Analytics and Continuous Reporting
Interruptible analytics and continuously available reporting have different compute requirements.
Recommended architecture
● Run large analytics jobs on an EC2 Spot-based Auto Scaling group.
● Design jobs with checkpoints and retry support.
● Run the continuous reporting service on Amazon ECS Fargate.
● Scale each component independently.
Design reasoning
● Technical workability: Spot provides discounted compute for restartable processing, while Fargate runs long-lived containers without server management.
● Requirement fit: Approximately 10,000 analytics compute-hours benefit from Spot savings, while reporting remains continuously available.
● Scene fit: Analytics is batch-oriented, but reporting serves ongoing user requests.
● Engineering common sense: Do not use interruptible capacity for a component that must always respond.
Alternatives to reject
● All On-Demand capacity is workable but unnecessarily expensive for fault-tolerant analytics.
● Reserved capacity is inefficient when analytics demand varies and the commitment is uncertain.
● Using Spot for the reporting service can cause visible interruptions.
● App Runner may host web services, but it does not improve the batch-compute economics of the analytics tier.
Workflow: Submit analytics work → scale Spot workers → persist results → reporting containers read results → Fargate scales service demand.
Reacting to DynamoDB Changes with Streams
When application behavior must follow every new DynamoDB item, use the table’s native change stream.
Recommended architecture
● Enable DynamoDB Streams with the required image view.
● Configure AWS Lambda as an event-source consumer.
● Process new-entry records idempotently.
● Monitor iterator age, errors, retries, and dead-letter handling.
Design reasoning
● Technical workability: DynamoDB Streams records item changes and Lambda polls the stream through a managed event source mapping.
● Requirement fit: Every table insert can trigger downstream processing with low operational effort.
● Scene fit: The event of interest is a database mutation, not an ECS application log.
● Engineering common sense: Capture change at the authoritative data source.
Alternatives to reject
● SNS requires the application to publish separately and can miss writes from other producers.
● Systems Manager Automation is for infrastructure operations, not change-data capture.
● Migrating to DocumentDB changes the database without need.
● Scheduled table scans introduce delay and duplicate reads.
Workflow: DynamoDB write → stream record → Lambda batch → idempotent action → checkpoint and retry management.
Centralized Egress for Thousands of AWS Accounts
Large multi-account environments need a scalable routing hub and centralized policy enforcement for outbound traffic.
Recommended architecture
● Attach spoke VPCs to AWS Transit Gateway.
● Route approved outbound traffic to a centralized egress or inspection VPC.
● Use managed or firewall VPN attachments according to the inspection design.
● Apply organization-controlled routes and firewall policies.
● Return permitted traffic to the appropriate spoke.
Design reasoning
● Technical workability: Transit Gateway provides hub-and-spoke routing without thousands of pairwise peering relationships.
● Requirement fit: Security teams manage egress controls centrally while application accounts retain separate VPCs.
● Scene fit: The design must scale across thousands of accounts.
● Engineering common sense: Centralize policy and inspection, not every workload’s application architecture.
Alternatives to reject
● VPC peering creates unmanageable connection and route growth.
● Shared VPCs do not fit every independent account and routing boundary.
● A self-managed transit VPC with EC2 appliances adds patching and scaling work.
● Independent NAT and firewall stacks in every account duplicate cost and policy administration.
Workflow: Spoke route → Transit Gateway → centralized inspection → approved internet egress → return path through hub.
Multiple Network Identities on One EC2 Prototype
A single prototype instance can expose components through separate IP identities without spanning Availability Zones.
Recommended architecture
● Attach multiple Elastic Network Interfaces to the EC2 instance.
● Assign distinct private IP addresses to the ENIs.
● Associate Elastic IP addresses where public static identities are required.
● Apply suitable security groups to each interface.
Design reasoning
● Technical workability: Multiple ENIs provide separate network identities while the instance remains in one Availability Zone.
● Requirement fit: Each component can bind to its own address on one prototype host.
● Scene fit: The design is a prototype, so separate instances may be unnecessary.
● Engineering common sense: Separate IP identity does not equal fault isolation; production may still require independent hosts.
Alternatives to reject
● Security groups filter traffic but do not create additional IP identities.
● NAT provides address translation, not multiple inbound component identities.
● One ENI cannot attach to subnets in two Availability Zones.
● An EC2 instance cannot span Availability Zones.
Workflow: Create ENIs → assign addresses and security groups → attach to instance → bind each component → test routing.
DynamoDB Keys for Sensor Time Series
Time-series queries per sensor need a partition key for the sensor and an ordered sort key for time.
Recommended design
● Create weekly DynamoDB tables to bound active data size.
● Use sensor ID as the partition key.
● Use timestamp as the sort key.
● Query one sensor with time-range conditions.
Design reasoning
● Technical workability: Items for one sensor are colocated logically and ordered by timestamp.
● Requirement fit: The application efficiently retrieves recent measurements for a sensor.
● Scene fit: Data is naturally grouped by device and time.
● Engineering common sense: Ensure sensor traffic is sufficiently distributed to avoid hot partitions.
Alternatives to reject
● Concatenating sensor and timestamp into one partition key makes range queries difficult.
● Weekly tables alone do not repair a poor key design.
● DynamoDB’s partition key is the hash key; reversing the key roles is incorrect.
● One global partition key creates a hot partition.
Workflow: Select weekly table → query sensor partition → apply timestamp range → return ordered measurements.
Reliable SQS Retries Before Dead-Letter Isolation
A dead-letter queue should isolate repeatedly failing work, not messages that experienced one transient processing error.
Recommended configuration
● Keep the visibility timeout longer than normal processing time.
● Increase the redrive policy maxReceiveCount from 1 to a reasonable retry value, such as 10.
● Continue alarming on dead-letter queue depth.
● Investigate messages only after normal retries are exhausted.
Design reasoning
● Technical workability: When processing fails or the message is not deleted, SQS makes it visible again until the receive threshold is reached.
● Requirement fit: Multiple attempts improve completion without changing the worker fleet.
● Scene fit: Videos normally finish in 20 to 40 minutes, below the one-hour visibility timeout.
● Engineering common sense: Distinguish transient worker failure from permanently invalid input.
Alternatives to reject
● Maintaining more idle EC2 capacity does not address failed message handling.
● Extending visibility to two hours delays retry even though normal processing already fits within one hour.
● Delivery delay postpones first availability and does not help after a consumer failure.
Workflow: Receive → hide during processing → delete on success → retry on failure → move to DLQ after threshold → alert developers.
Central Account Guardrails with Service Control Policies
Multi-account governance requires centrally enforced permission boundaries rather than duplicated account-level IAM policies.
Recommended architecture
● Organize departmental accounts with AWS Organizations and organizational units.
● Attach Service Control Policies to the appropriate root, OU, or account.
● Keep workload IAM policies inside each account for actual permission grants.
Design reasoning
● Technical workability: SCPs define the maximum permissions available to principals in member accounts.
● Requirement fit: Central administrators can allow or deny service categories with low maintenance.
● Scene fit: Departmental accounts remain independently administered while inheriting company guardrails.
● Engineering common sense: Separate organization-wide boundaries from local role permissions.
Alternatives to reject
● IAM policies attached inside each account do not centrally constrain every principal or the member-account root user.
● Identity federation solves authentication and workforce access, not account-wide service restrictions.
● Cross-account roles and per-resource policies create high complexity and incomplete governance coverage.
Important behavior: An SCP does not grant access. A request succeeds only when identity or resource policies allow it and no applicable SCP blocks it.
Workflow: Group accounts → design guardrails → test in a limited OU → attach SCPs → monitor denied actions.
Organization-Wide Sharing with AWS RAM
AWS Resource Access Manager integrates with AWS Organizations through trusted access, avoiding custom cross-account automation.
Recommended architecture
● Enable resource sharing with AWS Organizations using enable-sharing-with-aws-organization.
● Allow AWS RAM to create and use its service-linked role.
● Reproduce the established resource-share configuration for organizational accounts or OUs.
Design reasoning
● Technical workability: Trusted access lets RAM operate across organization accounts using the AWS-managed integration.
● Requirement fit: The native mechanism has lower ongoing administration than account-by-account role orchestration.
● Scene fit: The task is organizational resource sharing, not operating-system automation.
● Engineering common sense: Prefer a managed service integration when it already implements the required trust relationship.
Alternatives to reject
● Service-linked role trust policies are service controlled and are not a normal customization point.
● Generic cross-account access does not activate the RAM and Organizations integration.
● SSM agents, worker virtual machines, and automation documents do not contribute to RAM sharing.
Workflow: Enable trusted access → verify service-linked role → define resource share → select principals → validate access → audit changes.
URL-Based Egress with a Web Proxy
When instances accept inbound traffic but may download updates only from specified websites, outbound inspection must occur independently of inbound access.
Recommended architecture
● Place a managed or self-managed web proxy on the outbound path.
● Configure explicit URL or domain allow rules.
● Make private instances use the proxy for web access.
● Remove alternate direct egress paths.
● Retain inbound load-balancer and security-group rules separately.
Design reasoning
● Technical workability: A proxy evaluates application-layer destinations and forwards only compliant requests.
● Requirement fit: Inbound service availability remains unchanged while outbound access is tightly controlled.
● Scene fit: Package updates use known URLs whose underlying IP addresses can change.
● Engineering common sense: Do not maintain fragile IP allow lists for services identified by URL.
Alternatives to reject
● NAT Gateway does not filter websites.
● Security groups and NACLs cannot inspect requested URLs.
● Removing all internet access prevents required updates.
● S3 or service VPC endpoints help only when the approved repository is an endpoint-supported AWS service.
Workflow: Application receives inbound request → EC2 initiates update request → proxy checks URL → approved request exits → all other egress denied.
Organization-Wide Required Tags with SCPs
An organization SCP can deny resource creation when required tag keys or request tags are missing.
Recommended architecture
● Define required tag keys such as cost center and owner.
● Use SCP conditions with aws:TagKeys and request-tag keys.
● Attach the SCP to the appropriate OUs.
● Test services whose create APIs support tagging.
Design reasoning
● Technical workability: The explicit deny blocks noncompliant creation across member accounts.
● Requirement fit: Untagged resources cannot consume quotas or create unallocated cost.
● Scene fit: Enforcement must be centralized.
● Engineering common sense: Account for services that tag after creation or use different APIs.
Alternatives to reject
● AWS Config detects missing tags after creation.
● Systems Manager is not the organization-wide preventive policy service.
● Per-account IAM conditions require repeated administration.
● Tag policies standardize values but do not by themselves enforce every create action.
Workflow: Create request includes tags → SCP evaluates keys and values → compliant request continues; missing tags denied.
Inter-Region Private Connectivity with Direct Connect Gateway
A central office can reach VPCs in multiple Regions through one managed Direct Connect architecture.
Recommended architecture
● Use AWS Direct Connect Gateway.
● Attach a virtual private gateway to each regional VPC.
● Connect private virtual interfaces to the Direct Connect Gateway.
● Advertise approved routes with BGP.
Design reasoning
● Technical workability: Direct Connect Gateway associates private VIF connectivity with virtual private gateways across supported Regions.
● Requirement fit: Traffic uses dedicated connectivity with predictable performance and centralized management.
● Scene fit: The office needs private access to several regional VPCs, not only VPC-to-VPC communication.
● Engineering common sense: Use a hub service instead of building and maintaining a full mesh of point-to-point links.
Alternatives to reject
● Inter-Region VPC peering does not connect the on-premises office and creates many pairwise relationships.
● A public VIF is not the correct path for private VPC prefixes.
● A link aggregation group improves capacity or resilience at one Direct Connect location but does not replace the private-VIF design.
● Transit Gateway with internet-based VPN does not meet the dedicated-path requirement.
Workflow: Office router → Direct Connect → private VIF → Direct Connect Gateway → regional VGW → VPC.
Security-Group Referencing for EC2 to Aurora
Database access can be restricted to application instances by referencing security groups rather than broad CIDR ranges.
Recommended rules
● Allow outbound TCP 3306 from the EC2 security group to the Aurora security group.
● Allow inbound TCP 3306 on the Aurora security group from the EC2 security group.
● Remove broader database access rules.
Design reasoning
● Technical workability: Security groups are stateful and can reference other security groups.
● Requirement fit: Only approved application instances can initiate MySQL connections.
● Scene fit: EC2 is the client and Aurora is the server.
● Engineering common sense: Authorize workload identity groups instead of changing instance IP addresses.
Alternatives to reject
● The application does not need inbound database traffic on its own security group.
● Aurora does not initiate the application database session.
● NACL CIDR rules are broader and stateless.
● An inbound NACL alone omits return-path considerations and lacks security-group precision.
Workflow: EC2 initiates 3306 → outbound SG check → Aurora inbound SG check → stateful return traffic allowed.
Low-Latency Placement for an HPC Controller
A controller communicating frequently with clustered compute nodes should share the existing cluster placement group.
Recommended procedure
● Stop the controller instance if required.
● Move it into the compute fleet’s cluster placement group.
● Restart and measure latency and throughput.
Design reasoning
● Technical workability: Cluster placement groups keep supported instances physically close for low-latency networking.
● Requirement fit: Controller-to-node communication improves without rebuilding the whole cluster.
● Scene fit: Compute nodes already use the correct placement group.
● Engineering common sense: Change the outlier component rather than disrupting every healthy node.
Alternatives to reject
● An Elastic IP does not improve internal network performance.
● Spread placement intentionally separates instances and increases distance.
● ENA is not attached as a separate adapter in the proposed way.
● Rebuilding every compute instance creates unnecessary downtime.
Workflow: Stop controller → modify placement → start → validate network performance → resume workload.
Cost-Efficient Peak Capacity Across Availability Zones
A multi-AZ web tier should separate predictable base capacity from fault-tolerant peak capacity.
Recommended architecture
● Keep Reserved Instance coverage for steady-state usage.
● Use diversified Spot capacity for temporary peaks.
● Place Auto Scaling capacity across Availability Zones.
● Let scaling policies replace capacity lost during an AZ or Spot interruption.
Design reasoning
● Technical workability: Stateless web servers can be replaced without preserving local session state.
● Requirement fit: Spot lowers peak cost, while Auto Scaling supports rapid capacity recovery.
● Scene fit: The application already spans three AZs and reaches very high utilization during peaks.
● Engineering common sense: Do not depend on one Spot pool; diversify instance types and capacity pools.
Alternatives to reject
● Fixed Reserved and On-Demand capacity works but may overprovision and does not guarantee automated replacement.
● Reserved Instances are a pricing commitment, not a special Auto Scaling capacity type.
● Mixing Spot and On-Demand without Auto Scaling does not provide fast recovery.
Workflow: Measure baseline → cover baseline with commitments → diversify peak pools → scale across AZs → monitor interruptions and utilization.
DNS Distribution or Managed Load Balancing
A fixed public server fleet can use DNS-level multivalue responses, but an Application Load Balancer is the stronger managed design.
Recommended options
● Use Route 53 multivalue routing with health checks to return multiple healthy server IPs.
● Prefer an ALB with a Route 53 alias when managed request distribution is available.
Design reasoning
● Technical workability: Multivalue responses distribute client selections; ALB actively balances requests across healthy targets.
● Requirement fit: Both improve availability and distribution, while ALB reduces server-address management.
● Scene fit: The application has several web servers serving one domain.
● Engineering common sense: DNS is not a full load balancer because clients cache answers and choose addresses independently.
Alternatives to reject
● NAT provides outbound translation, not inbound load balancing.
● A non-alias record is less suitable for an AWS load balancer and cannot support every apex-domain case.
● CloudFront cannot use arbitrary private EC2 IP addresses as public origins.
● One static A record leaves a server failure point.
Workflow: DNS returns healthy ALB alias or multivalue addresses → client connects → health controls remove failed targets.
Managed Video Portal Processing
A video portal can reduce operations by separating the web tier, asynchronous analysis, and durable media storage.
Recommended architecture
● Run the dynamic web application on ECS Fargate.
● Store videos and static content in S3.
● Place analysis jobs in SQS.
● Use EC2 Spot workers for long processing.
● Use Amazon Rekognition instead of custom vision software.
Design reasoning
● Technical workability: Fargate removes web-server management, Spot reduces worker cost, and Rekognition provides managed analysis.
● Requirement fit: The design lowers cost and operational overhead.
● Scene fit: Upload handling is dynamic, while analysis is asynchronous and potentially long.
● Engineering common sense: Keep media outside compute instances.
Alternatives to reject
● S3 static hosting cannot run the upload application.
● Lambda may not fit long video jobs.
● EFS and EC2 web servers retain more infrastructure management.
● Elastic Beanstalk still operates EC2 fleets for both tiers.
Workflow: User uploads → Fargate → S3 and SQS → Spot worker → Rekognition → results stored.
Using SSE-C Through the S3 REST API
S3 server-side encryption with customer-provided keys requires the client to send the encryption key on every relevant request.
Required request headers
● x-amz-server-side-encryption-customer-algorithm
● x-amz-server-side-encryption-customer-key
● x-amz-server-side-encryption-customer-key-MD5
● Include the required SSE-C information when creating and using pre-signed requests.
Design reasoning
● Technical workability: S3 uses the supplied key to encrypt or decrypt the object but does not store the key.
● Requirement fit: The customer retains key custody while S3 performs object encryption.
● Scene fit: Uploads and downloads occur through REST APIs rather than only the console.
● Engineering common sense: Losing the customer key makes the object unrecoverable.
Alternatives to reject
● SSE-C is not limited to the AWS console.
● WebSocket Secure protects WebSocket transport and is unrelated to S3 encryption.
● The MD5 header alone is insufficient because S3 also needs the algorithm and key.
● Sending the key through an unprotected connection is unsafe; use HTTPS.
Workflow: Client builds HTTPS request → supplies all SSE-C headers → S3 validates MD5 → encrypts or decrypts object.
Preserving Failed Auto Scaling Instances for Diagnosis
Automated healing can erase the exact evidence needed to diagnose a failed deployment.
Recommended procedure
● Temporarily suspend the Auto Scaling group’s Terminate process.
● Allow an unhealthy instance to remain available.
● Connect through AWS Systems Manager Session Manager.
● Inspect process state, listening ports, dependencies, configuration, and local logs.
● Correct the launch template, AMI, or application package, then resume termination.
Design reasoning
● Technical workability: Suspending the process stops Auto Scaling from removing unhealthy instances while Session Manager provides private administrative access.
● Requirement fit: This is the quickest route to the original failed environment.
● Scene fit: The instances already have SSM Agent and do not require SSH exposure.
● Engineering common sense: Pause remediation briefly, collect evidence, fix the immutable source, then restore normal healing.
Alternatives to reject
● A separately created test instance may not reproduce the exact failure and adds setup time.
● More verbose logging does not preserve an instance long enough for inspection.
● EC2 termination protection does not stop Auto Scaling from terminating an instance it manages.
Workflow: Suspend → inspect → reproduce → correct → launch replacement → validate health → resume.
Queue-Based Media Processing
Variable, long-running media jobs need durable buffering and scalable workers rather than synchronous processing.
Recommended architecture
● Place jobs in Amazon SQS.
● Scale EC2 workers according to queue depth.
● Store source and processed media in Amazon S3.
● Set visibility timeout beyond expected processing time.
● Use retries and a dead-letter queue.
Design reasoning
● Technical workability: SQS decouples upload from processing, Auto Scaling follows backlog, and S3 provides durable shared storage.
● Requirement fit: The platform handles variable work at low cost without losing jobs.
● Scene fit: Media processing can exceed short serverless execution windows.
● Engineering common sense: Workers are disposable; messages and media must survive worker failure.
Alternatives to reject
● Lambda may be unsuitable for long-running jobs.
● EBS is not shared durable output storage for a scaling fleet.
● Amazon MQ adds broker compatibility and management not required here.
● Mixing MQ with an SQS-triggered design is internally inconsistent.
Workflow: Upload → SQS message → worker claims → process from S3 → write result → delete message.
Time-Limited Access to Private S3 Content
Private content can be shared temporarily through signatures while direct anonymous origin access remains blocked.
Recommended patterns
● Use S3 pre-signed URLs for direct, time-limited access to specific objects.
● Or use CloudFront signed URLs for edge-delivered private content.
● For CloudFront, protect the S3 origin with an Origin Access Identity and a restrictive bucket policy.
● Remove public S3 permissions.
Design reasoning
● Technical workability: Signatures encode an expiration and authorized resource; S3 or CloudFront validates them before delivery.
● Requirement fit: Only the approved client receives temporary access without permanent AWS credentials.
● Scene fit: Content is stored in S3 and may be distributed through CloudFront.
● Engineering common sense: Choose the signer at the delivery layer users actually access.
Alternatives to reject
● Public-read ACLs cannot restrict access to one client.
● Sharing IAM access keys creates long-lived credential risk.
● An OAI by itself protects the origin but does not authorize individual viewers.
● Network ACLs do not provide object-level time limits.
Workflow: Application authorizes client → generates S3 or CloudFront signature → client requests before expiry → service validates → object delivered.
Scaling Image Analysis by Queue Length
Shared object storage, durable work distribution, and backlog-based scaling reduce image-processing time safely.
Recommended architecture
● Store input and output files in Amazon S3.
● Put one processing task per image in Amazon SQS.
● Scale EC2 workers using SQS queue-depth metrics.
● Make workers idempotent and configure a dead-letter queue.
Design reasoning
● Technical workability: S3 is accessible to every worker, SQS distributes tasks, and Auto Scaling follows actual backlog.
● Requirement fit: More workers launch during peaks and terminate when work falls.
● Scene fit: Images are independent parallel tasks.
● Engineering common sense: Scale from pending work, not from notifications that may already have been consumed.
Alternatives to reject
● EBS is attached to instances and AZs, not shared across a dynamic fleet.
● SNS is a notification service, not a durable competing-consumer queue.
● SNS notification count is not the current processing backlog.
● Scaling from CPU alone may react too late or ignore queued demand.
Workflow: Image to S3 → task to SQS → queue grows → workers scale → results to S3 → queue drains.
One Client VPN for Peered VPCs
Employees can connect through one centrally managed Client VPN endpoint and reach applications in peered VPCs.
Recommended architecture
● Deploy the Client VPN endpoint in the main VPC.
● Install VPN clients on employee devices.
● Add Client VPN authorization and routes for remote VPC CIDRs.
● Configure reciprocal VPC peering routes and security rules.
Design reasoning
● Technical workability: Client VPN terminates remote-user sessions, while VPC peering carries traffic to connected application VPCs.
● Requirement fit: One endpoint reduces cost and administration across accounts.
● Scene fit: Internal applications already reside in peered VPCs.
● Engineering common sense: Verify that CIDRs do not overlap and peering is non-transitive.
Alternatives to reject
● One Client VPN per account duplicates management.
● Client software belongs on employee devices, not the data center.
● Site-to-Site VPN does not replace the Client VPN endpoint for roaming users.
● Missing return routes break the connection.
Workflow: Employee → Client VPN → main VPC route → peering connection → application VPC → return path.
DDoS Protection Across Network and Application Layers
Protection against Layer 3, Layer 4, and Layer 7 attacks requires complementary managed controls.
Recommended architecture
● Enable AWS Shield Advanced for supported public resources and enhanced DDoS response.
● Apply AWS WAF web ACLs to public HTTP endpoints.
● Configure alerts and response contacts for detected attacks.
Design reasoning
● Technical workability: Shield Advanced addresses infrastructure-layer attacks, while WAF filters malicious HTTP request patterns.
● Requirement fit: The combination covers network, transport, and application layers and supports attack notifications.
● Scene fit: A public cryptocurrency platform is a likely target for volumetric and web-layer attacks.
● Engineering common sense: No single web filter stops every infrastructure flood, and network protection cannot understand all HTTP payloads.
Alternatives to reject
● CloudFront improves resilience but does not alone provide the requested complete protection and response features.
● AWS Network Firewall controls VPC traffic but is not the primary managed DDoS service for public endpoints.
● Amazon Fraud Detector evaluates fraudulent business activity, not network attacks.
● Scaling only the origin can increase cost without filtering malicious traffic.
Workflow: Internet traffic → Shield protection → edge or load balancer → WAF inspection → application; alerts trigger response.
Trading Platform Recovery with Point-in-Time Data
A strict recovery plan needs frequent recoverability, protected backups, and copies outside the failed Region.
Recommended architecture
● Use AWS Backup with point-in-time recovery for supported databases.
● Store application backups and transaction logs in Amazon S3.
● Capture logs often enough to satisfy the ten-minute RPO.
● Replicate recovery objects to another Region with S3 Cross-Region Replication.
● Test restoration within the two-hour RTO.
Design reasoning
● Technical workability: PITR restores database state near a selected time, while replicated backups survive regional failure.
● Requirement fit: Frequent logs limit data loss and pre-positioned copies reduce recovery delay.
● Scene fit: A regulated trading system needs both operational recovery and durable evidence.
● Engineering common sense: RPO determines capture frequency; RTO determines restoration automation and data placement.
Alternatives to reject
● Multi-AZ protects against an AZ failure but is not regional disaster recovery.
● Daily snapshots cannot meet a ten-minute RPO.
● Glacier retrieval may delay a short RTO.
● Backups retained only in the production Region can disappear with a regional outage.
Workflow: Continuous transactions → PITR and frequent logs → cross-Region copy → restore → replay → validate → redirect traffic.
High-Volume Data Retention with DynamoDB TTL
A high-ingest workload with fixed 120-day retention needs scalable writes, low-latency reads, and automatic item expiration.
Recommended architecture
● Store records in Amazon DynamoDB using a well-distributed partition key.
● Add an expiration timestamp attribute to every item.
● Enable DynamoDB Time to Live on that attribute.
● Select on-demand or provisioned capacity according to traffic predictability.
Design reasoning
● Technical workability: DynamoDB scales horizontally and TTL removes expired items automatically without application delete jobs.
● Requirement fit: The design supports durable ingestion, responsive access, and low-operation retention management.
● Scene fit: Records have a uniform lifecycle and do not require relational joins.
● Engineering common sense: Encode lifecycle at item creation rather than scanning the table later for old data.
Alternatives to reject
● Relational databases add connection and scaling constraints for a simple high-volume key-value stream.
● Scheduled deletion jobs consume capacity and require checkpoints, retries, and maintenance.
● Keeping expired data indefinitely increases storage cost and violates retention intent.
● Archival object storage is less suitable when the active requirement includes low-latency record access.
Workflow: Ingest item → assign partition key and expiry time → serve low-latency queries → TTL removes data after 120 days.
Location-Based Mobile Offers
A short delivery window requires durable location buffering, fast offer lookup, scalable processing, and managed mobile push.
Recommended architecture
● Buffer incoming locations in SQS.
● Scale API or worker instances according to backlog.
● Store offers in DynamoDB.
● Send selected offers through SNS Mobile Push.
Design reasoning
● Technical workability: SQS absorbs bursts, DynamoDB provides low-latency lookup, and SNS reaches mobile push services.
● Requirement fit: Nearby offers can be selected and delivered within the short window.
● Scene fit: Millions of mobile locations may arrive irregularly.
● Engineering common sense: Separate location intake from push delivery.
Alternatives to reject
● Kinesis and SES do not form the direct mobile-push workflow.
● Direct Connect to mobile carriers is not how device GPS or platform push works.
● EC2 cannot directly replace managed mobile push channels.
● Synchronous processing without a queue risks dropped events.
Workflow: Device location → SQS → worker queries DynamoDB → SNS Mobile Push → phone notification.
Cost-Optimized Internal Container and Document Platform
A low-cost internal platform should use interruptible compute where safe, committed database pricing for steady demand, and lifecycle-managed object storage.
Recommended architecture
● Run ECS container instances on EC2 Spot capacity.
● Enable ECS Spot Instance draining.
● Use Reserved pricing for the steady Amazon RDS database tier.
● Store documents in encrypted Amazon S3.
● Transition documents to S3 Glacier after three months and expire them only after the required five-year retention period.
Design reasoning
● Technical workability: ECS drains tasks before Spot interruption, while S3 lifecycle policies automate storage transitions.
● Requirement fit: The design minimizes compute and storage cost while preserving documents for five years.
● Scene fit: Documents are frequently accessed only during their first three months.
● Engineering common sense: Align storage class with access age and avoid custom archival cron jobs.
Alternatives to reject
● On-Demand ECS and RDS cost more for predictable long-running use.
● EFS plus custom copy scripts adds file-system and archival operations.
● EKS adds Kubernetes management without a stated need.
● RDS cannot run on Spot Instances.
● Host-local container volumes are not durable document storage.
Workflow: Store active document in S3 → serve as authorized → lifecycle to Glacier → retain five years → expire by policy.
Fast Multi-Region Recovery for a NoSQL Application
A regional disaster-recovery design should reproduce infrastructure, replicate every data type, and automate traffic failover and failback.
Recommended architecture
● Use CloudFormation StackSets to deploy matching Auto Scaling web and application tiers in both Regions.
● Enable S3 Cross-Region Replication for static content.
● Use DynamoDB global tables for multi-Region NoSQL data replication.
● Configure Route 53 failover routing with health checks.
Design reasoning
● Technical workability: StackSets standardize regional infrastructure, S3 CRR copies objects, and global tables replicate DynamoDB records.
● Requirement fit: A ready secondary site allows fast recovery, while Route 53 automates traffic movement and return.
● Scene fit: The database requirement is explicitly NoSQL, making DynamoDB global tables the direct fit.
● Engineering common sense: Recovery speed depends on pre-provisioned infrastructure and continuously replicated data.
Alternatives to reject
● Aurora is relational and does not match the stated NoSQL tier.
● Service Catalog does not directly reproduce a complete regional stack for this workflow.
● Manual DNS changes slow recovery and failback.
● Periodic backups to S3 leave a larger data gap and require restoration before use.
Workflow: Deploy both Regions → replicate S3 and DynamoDB → health-check primary → fail to secondary → fail back after recovery.
Discovering Servers for Migration TCO
Accurate migration sizing and total-cost estimates require measured server configuration, utilization, and dependency data.
Recommended architecture
● Deploy AWS Application Discovery Service agents or agentless collectors.
● Gather CPU, memory, network, process, and connection information.
● Group dependent servers into applications.
● Use collected utilization to size AWS targets and estimate TCO.
Design reasoning
● Technical workability: Discovery Service collects detailed estate data needed for migration planning.
● Requirement fit: Recommendations are based on observed use rather than installed hardware alone.
● Scene fit: The organization is still assessing its data center before migration.
● Engineering common sense: Right-size from representative utilization periods and known business peaks.
Alternatives to reject
● Migration Hub tracks progress but does not independently collect all detailed utilization data.
● AWS SAM builds serverless applications.
● AWS MGN replicates servers and is not the primary discovery and TCO service.
● Manual spreadsheets quickly become stale and may miss dependencies.
Workflow: Install collector → observe workload cycle → review dependencies → group applications → right-size targets → estimate migration business case.
Patching EC2 and Recording Compliance
Urgent operating-system fixes require a deployment service and a separate compliance recorder.
Recommended architecture
● Define approved patches in Systems Manager Patch Manager.
● Target managed EC2 instances through patch groups and maintenance windows.
● Run patch operations promptly.
● Use AWS Config to record and evaluate compliance state.
Design reasoning
● Technical workability: Patch Manager installs updates; Config records resource and compliance history.
● Requirement fit: Running instances receive fixes and auditors can review fleet status.
● Scene fit: The vulnerability affects existing EC2 instances now.
● Engineering common sense: A new image helps future launches but does not repair the current fleet by itself.
Alternatives to reject
● Waiting for weekly AMI replacement may be too slow.
● OpenSearch is unnecessary for patch deployment.
● State Manager is useful for desired state, but Patch Manager is purpose built.
● Control Tower governs accounts and does not patch operating systems.
Workflow: Approve patch → target fleet → install and reboot → report status → Config records compliance → remediate failures.
Central Private DNS Across AWS Accounts
A shared-services account can own private DNS while application VPCs in other accounts resolve the same internal names.
Recommended architecture
● Create the private hosted zone in the shared-services account.
● Authorize cross-account associations for approved application VPCs.
● Associate each VPC with the central private hosted zone.
● Maintain records centrally and enable VPC DNS support.
Design reasoning
● Technical workability: Route 53 private hosted zones support association with VPCs owned by other AWS accounts.
● Requirement fit: DNS administration remains centralized without duplicating zones or running custom resolvers.
● Scene fit: Multiple accounts need consistent internal service discovery.
● Engineering common sense: Use one source of truth for common names and delegate only association rights.
Alternatives to reject
● Recreating the same private zone in every account risks inconsistent records.
● VPC peering alone does not share hosted-zone associations.
● Public hosted zones expose internal naming and do not provide private-only resolution.
● Self-managed DNS servers add patching, availability, forwarding, and scaling work.
Workflow: Create central zone → authorize VPC → associate from workload account → publish record → test resolution → audit associations.
Near-Real-Time Event Search and Dashboards
Semi-structured JSON events need scalable ingestion, transformation, indexing, and dashboard visualization.
Recommended architecture
● Use Kinesis Data Firehose to buffer and deliver records.
● Transform events with Lambda when required.
● Index and visualize data in Amazon OpenSearch Service.
Design reasoning
● Technical workability: Firehose manages delivery, Lambda reshapes records, and OpenSearch provides searchable indexes and dashboards.
● Requirement fit: The pipeline supports dynamic schemas and near-real-time operational views.
● Scene fit: The workload is event search, not relational transactions or graph traversal.
● Engineering common sense: The ingestion service should buffer temporary destination slowdown.
Alternatives to reject
● Aurora PostgreSQL remains a relational write bottleneck for this event pattern.
● QuickSight is less direct for operational near-real-time event dashboards.
● Neptune is a graph database.
● Kinesis Data Streams and Lambda alone omit durable searchable storage and visualization.
Workflow: Event → Firehose → Lambda transform → OpenSearch index → dashboard query.
Removing VPN Site and Tunnel Failure Points
Site-to-Site VPN resilience requires independent on-premises locations and redundant tunnels.
Recommended architecture
● Add a customer gateway at the second data center.
● Create a Site-to-Site VPN connection with both AWS-managed tunnels.
● Use dynamic routing where supported.
● Test loss of one tunnel and loss of one data center.
Design reasoning
● Technical workability: A second customer gateway creates a separate on-premises path, while dual tunnels protect against tunnel endpoint failure.
● Requirement fit: The connection survives both site and tunnel problems.
● Scene fit: The current single data center is the major failure domain.
● Engineering common sense: Redundancy must not share the component being protected.
Alternatives to reject
● A VPC does not attach separate virtual private gateways by Availability Zone.
● A second VGW is not the normal model for one VPC’s VPN resilience.
● NAT Gateway provides outbound internet translation and is not an on-premises VPN endpoint.
● Two tunnels terminating at one failed customer site do not provide site redundancy.
Workflow: Normal BGP path → tunnel failure uses peer tunnel → site failure uses second customer gateway → routes converge.
Recurring AMI Vulnerability Assessment
An approved-AMI pipeline should automate recurring scans, approval state, and replacement decisions.
Recommended architecture
● Use Amazon Inspector assessment templates to scan target EC2 instances for CVEs.
● Store approved AMI identifiers in Systems Manager Parameter Store.
● Use EventBridge to start a recurring workflow.
● Use Lambda for approval logic and Systems Manager Automation for image or fleet actions.
Design reasoning
● Technical workability: Inspector performs vulnerability assessment, while the other services coordinate scheduling and remediation.
● Requirement fit: The process regularly identifies vulnerable images and maintains an authoritative approved list.
● Scene fit: The organization needs automated CVE scanning rather than only launch auditing.
● Engineering common sense: Separate detection, approval decision, and remediation execution.
Alternatives to reject
● CloudTrail records launches but does not scan packages for CVEs.
● AWS Config records compliance but is not a detailed vulnerability scanner.
● SSM Agent alone does not perform the assessment.
● Manual review does not meet recurring automation requirements.
Workflow: EventBridge schedule → Inspector scan → Lambda evaluates findings → update approved parameter → SSM Automation remediates.
Encrypted VPN Transport over Direct Connect
Direct Connect provides private, predictable transport but does not encrypt traffic by default. Encryption can be layered through an AWS Site-to-Site VPN carried over the dedicated connection.
Recommended architecture
● Create a public virtual interface on the existing Direct Connect connection.
● Reach the public AWS VPN endpoints through that interface.
● Establish a BGP-based Site-to-Site VPN to the VPC.
● Route corporate employee traffic through the encrypted tunnel.
Design reasoning
● Technical workability: The public VIF provides access to AWS public service endpoints, including VPN termination addresses.
● Requirement fit: IPsec encrypts traffic while the underlying path retains Direct Connect performance characteristics.
● Scene fit: Employees already reach private EC2 applications from the corporate network.
● Engineering common sense: Add encryption to the existing reliable path rather than moving traffic back to the public Internet.
Alternatives to reject
● A VPN routed over the Internet loses the requested Direct Connect consistency.
● A private VIF reaches VPC private addresses but not the public VPN endpoints required for this pattern.
● Laptop-by-laptop VPN connections do not match the existing corporate network routing design.
Workflow: Corporate route → customer router → IPsec tunnel → public VIF over Direct Connect → AWS VPN endpoint → VPC.
LDAP Authentication with Temporary S3 Credentials
Existing LDAP identities can access S3 securely through a federation broker that maps users to IAM roles.
Recommended patterns
● Let the application authenticate the user against LDAP, map the identity to an IAM role, and call STS.
● Or use a dedicated identity broker to perform LDAP validation and role mapping.
● Return temporary AWS credentials for authorized S3 operations.
Design reasoning
● Technical workability: STS provides short-lived role credentials after the trusted application or broker validates the user.
● Requirement fit: Employees use existing credentials without creating permanent IAM users or embedding AWS keys.
● Scene fit: LDAP remains the enterprise identity source while S3 hosts protected content.
● Engineering common sense: Separate authentication from AWS authorization and keep sessions temporary.
Alternatives to reject
● Creating one IAM user per LDAP employee duplicates identity lifecycle management.
● Direct Connect Gateway and transit networking do not perform authentication.
● An S3 bucket policy cannot validate an LDAP password.
● Long-lived access keys in the application increase exposure and rotation work.
Workflow: User login → LDAP validation → role mapping → STS temporary credentials → signed S3 request → expiration.
Why SCPs Do Not Grant Permissions
A Service Control Policy defines the maximum permissions available in a member account but grants nothing by itself.
Required authorization chain
● Keep the applicable SCP allowance.
● Attach an identity policy to the IAM user or role granting required EC2 and S3 actions.
● Evaluate resource policies and other permission boundaries.
Design reasoning
● Technical workability: Effective access requires both an SCP boundary that permits the action and an IAM grant.
● Requirement fit: Adding identity permissions resolves access without weakening organization controls.
● Scene fit: The account inherits an SCP through its OU.
● Engineering common sense: Guardrails and grants have separate purposes.
Alternatives to reject
● SCPs are valid organization guardrails and should not be replaced merely because access is missing.
● Accounts automatically inherit SCPs attached to parent OUs.
● Member-account root users are also limited by SCPs.
● Root should not be used for normal resource creation.
Workflow: Request → SCP maximum checked → IAM grant checked → resource policy and conditions checked → allow or deny.
Private DNS Across Peered VPCs
Internal applications need both private network reachability and private name resolution.
Recommended architecture
● Create a Route 53 private hosted zone for the internal domain.
● Associate the required VPCs with the hosted zone.
● Create an A record for the database server’s private IP address.
● Enable enableDnsSupport and enableDnsHostnames where required.
● Maintain peering routes and security rules for database traffic.
Design reasoning
● Technical workability: Associated VPCs can resolve private hosted-zone records through the Amazon-provided resolver.
● Requirement fit: The name remains unavailable through public DNS.
● Scene fit: Existing VPC peering carries the traffic after resolution.
● Engineering common sense: Internal services should use private addresses and stable DNS names, not public endpoints.
Alternatives to reject
● A public hosted zone exposes the name publicly.
● A CNAME record cannot map directly to an IP address.
● An Elastic IP creates an unnecessary public addressing path.
● Disabling DNS support prevents the intended resolution.
Workflow: Associate zones → enable DNS → create private record → verify routes → verify security groups → test resolution and connection from each VPC.
One ECS Control Plane for On-Premises and Fargate
Hybrid container operations are simplest when on-premises servers participate in the same ECS operating model used by production.
Recommended architecture
● Use Amazon ECS Anywhere to register existing on-premises servers as external ECS instances.
● Manage development tasks through the ECS control plane.
● Build task definitions and container images that can also run on AWS Fargate.
● Promote validated workloads from on-premises development to regional Fargate production.
Design reasoning
● Technical workability: ECS Anywhere extends ECS scheduling and APIs to customer-managed servers.
● Requirement fit: Teams receive consistent tooling and an easy workload migration path while reusing existing capital equipment.
● Scene fit: Production already uses ECS Fargate, so ECS is the common orchestrator.
● Engineering common sense: Avoid introducing Kubernetes or dedicated hardware merely to host development containers.
Alternatives to reject
● EKS Anywhere changes the orchestration platform and does not create one ECS cluster.
● AWS Outposts requires additional hardware investment and conflicts with the cost-saving goal.
● ECS on Outposts does not use the existing servers in the direct manner requested.
Workflow: Register external instances → deploy development task → test image and task definition → publish revision → run on Fargate.
Burst-Ready E-Commerce Architecture
A major sale needs elastic web capacity, global content delivery, scalable identity, and a checkout path that absorbs sudden purchasing bursts.
Recommended architecture
● Place an Elastic Load Balancer before an EC2 Auto Scaling group.
● Cache static assets through Amazon CloudFront.
● Use Amazon Cognito for customer and social sign-in.
● Buffer checkout requests in Amazon SQS.
● Process queued purchases into Amazon DynamoDB.
Design reasoning
● Technical workability: Auto Scaling adds web capacity, SQS preserves requests during spikes, and DynamoDB scales transaction records horizontally.
● Requirement fit: The design supports millions of visitors without synchronously overloading checkout workers.
● Scene fit: Browsing traffic and purchase processing have different burst patterns.
● Engineering common sense: Decouple acceptance from processing so a temporary backend slowdown does not lose orders.
Alternatives to reject
● A fixed EC2 fleet cannot absorb an unpredictable sales surge.
● RDS without buffering can become a write bottleneck.
● Static S3 hosting alone cannot execute dynamic checkout logic.
● IAM users are not a scalable identity model for public customers.
Workflow: Customer → CloudFront or load balancer → Cognito session → checkout message in SQS → workers → DynamoDB order state.
Searchable 50 TB Document Platform
A large document archive needs durable object storage, a search index, dynamic application hosting, and repeatable infrastructure.
Recommended architecture
● Define the environment in AWS CloudFormation.
● Store the 50 TB document collection in Amazon S3.
● Index searchable metadata and text in Amazon CloudSearch.
● Host the dynamic website on Amazon EC2.
Design reasoning
● Technical workability: S3 stores large object collections, CloudSearch handles queries, and EC2 runs dynamic application logic.
● Requirement fit: The design scales while avoiding a large relational database for document binaries.
● Scene fit: Users search documents rather than process a real-time stream.
● Engineering common sense: Keep original objects separate from rebuildable search indexes.
Alternatives to reject
● S3 static website hosting cannot replace the dynamic application.
● S3 does not provide native full-text search.
● RDS is costly and unnecessary for 50 TB of document objects.
● Kinesis is a streaming service, not durable document storage or a search index.
Workflow: Upload document → store in S3 → extract and index metadata → query CloudSearch → retrieve object.
Scaling Read-Heavy Web Applications
Read-heavy applications improve through database replicas, caching, static offload, and right-sized compute before sharding or vertical scaling.
Recommended architecture
● Add RDS read replicas.
● Use Multi-AZ ElastiCache for repeated query results.
● Move static media behind CloudFront.
● Use Compute Optimizer to right-size EC2 capacity.
Design reasoning
● Technical workability: Replicas offload reads, cache removes repeated queries, and CloudFront reduces origin traffic.
● Requirement fit: Performance improves with lower cost and less application change.
● Scene fit: The bottleneck is high read volume.
● Engineering common sense: Remove repeated work before introducing data partitioning.
Alternatives to reject
● Sharding adds major application and operational complexity.
● Provisioned IOPS helps storage-bound workloads but may not solve repeated queries.
● A larger database instance provides only vertical scaling.
● Scaling application servers alone leaves RDS pressure unchanged.
Workflow: Viewer static request → CloudFront; dynamic request → application → ElastiCache → read replica on miss → writer for updates.
Serverless Delivery for a React Application
A static single-page application does not need an always-on web-server fleet.
Recommended architecture
● Host the React build in an Amazon S3 website bucket.
● Authenticate users through a supported web identity provider.
● Exchange identity tokens for temporary AWS credentials with STS.
● Grant narrowly scoped access to the required S3 objects and DynamoDB items.
Design reasoning
● Technical workability: Browsers can load static React assets from S3 and use temporary credentials for authorized AWS API calls.
● Requirement fit: The design scales without servers and minimizes fixed cost.
● Scene fit: Usage has no major spikes, captions are small, and DynamoDB already matches the data pattern.
● Engineering common sense: Use temporary credentials rather than embedded keys or a custom token service.
Alternatives to reject
● NGINX, EC2, a load balancer, and Auto Scaling add cost for static content.
● A single Token Vending Machine becomes a bottleneck and failure point.
● A custom token service plus an EC2 web fleet duplicates managed capabilities.
Workflow: Sign in → receive identity token → assume role with web identity → receive temporary credentials → access approved resources.
Allowing S3 Access Only Through CloudFront
A private S3 origin should trust CloudFront rather than anonymous viewers or the distribution’s public DNS name.
Recommended architecture
● Create a CloudFront Origin Access Identity.
● Configure the distribution to use the OAI for the S3 origin.
● Update the S3 bucket policy to grant object-read permission only to the OAI.
● Remove public-read ACLs and public bucket permissions.
Design reasoning
● Technical workability: CloudFront signs origin requests using the OAI identity, and S3 evaluates that identity in the bucket policy.
● Requirement fit: Users receive content through CloudFront but direct S3 URLs fail.
● Scene fit: The bucket is an origin, not a public website endpoint.
● Engineering common sense: Authorize the service identity that makes the origin request.
Alternatives to reject
● A CloudFront distribution ID is not the S3 principal used for OAI authorization.
● Creating a general IAM user for CloudFront introduces manual credentials unnecessarily.
● Viewer signed URLs control access to CloudFront but do not automatically restrict direct S3 access.
● Public bucket access defeats origin protection.
Workflow: Viewer → CloudFront authorization → OAI-signed origin request → S3 bucket policy → object response.
Bulk Database Migration with Snowball and VPN Catch-Up
A 25 TB initial database transfer over 50 Mbps needs offline bulk movement, followed by network replication of changes.
Recommended architecture
● Export the initial database dataset to Snowball.
● Ship and import the bulk data into AWS.
● Use VPN-based replication for changes generated during shipping.
● Cut over to Aurora after synchronization.
Design reasoning
● Technical workability: Snowball handles the large seed, while the smaller change stream fits the VPN.
● Requirement fit: The design reduces network usage and application downtime.
● Scene fit: The source continues growing during migration.
● Engineering common sense: Separate bulk seeding from ongoing deltas.
Alternatives to reject
● Sending the entire 25 TB over 50 Mbps takes too long.
● VM Import/Export moves machine images, not an efficient live database workflow.
● Stopping the application for the Snowball shipping cycle creates excessive downtime.
● Bulk import without catch-up loses changes.
Workflow: Export seed → Snowball shipment → load target → replicate VPN deltas → validate → cut over.
Shared High-Throughput Data with EFS
A fleet using the same dataset should mount one shared filesystem instead of keeping copies on each instance.
Recommended architecture
● Store the dataset on Amazon EFS.
● Mount it from all application instances.
● Use Provisioned Throughput when required throughput exceeds what filesystem size provides.
● Monitor throughput utilization and client performance.
Design reasoning
● Technical workability: EFS provides shared elastic file access, and Provisioned Throughput decouples performance from stored capacity.
● Requirement fit: Data duplication disappears and throughput becomes predictable.
● Scene fit: Multiple instances need concurrent access to the same files.
● Engineering common sense: Select performance mode separately from throughput mode.
Alternatives to reject
● Instance store is ephemeral and not shared.
● Max I/O improves aggregate scale but adds per-operation latency and may not fit this workload.
● Oversizing gp2 volumes wastes storage merely to gain performance.
● Separate EBS copies create synchronization and replacement problems.
Workflow: Instances mount EFS → all read shared dataset → provisioned throughput serves load → metrics guide tuning.
Durable and Scalable News Publishing Platform
A news site should distribute media globally, scale database reads, and keep static assets outside web-server disks.
Recommended architecture
● Store images, video, and static media in Amazon S3.
● Deliver content through Amazon CloudFront.
● Use Amazon RDS Multi-AZ for database high availability.
● Add RDS read replicas for read-heavy comments and article queries.
● Keep application servers stateless and scalable.
Design reasoning
● Technical workability: S3 and CloudFront scale content delivery, while RDS Multi-AZ and replicas address availability and read load separately.
● Requirement fit: The design is durable, globally responsive, and supports traffic growth.
● Scene fit: News media is static content, while comments and metadata remain relational.
● Engineering common sense: Multi-AZ protects the writer; read replicas scale reads. They solve different problems.
Alternatives to reject
● EBS RAID ties media durability and scaling to individual servers.
● EFS can share files but is not the best global content origin compared with S3.
● Lambda does not automatically repair a poorly designed stateful web tier.
● Oracle RAC adds cost and operational complexity outside the stated need.
Workflow: Publish media to S3 → cache through CloudFront → write metadata to primary DB → serve reads from replicas.
High Availability for Oracle with RDS Multi-AZ
A managed Oracle database requiring automatic failover should use an Amazon RDS Multi-AZ deployment.
Recommended architecture
● Run Oracle on Amazon RDS with Multi-AZ enabled.
● Connect applications through the RDS endpoint.
● Let RDS maintain a synchronous standby in another Availability Zone.
● Test application reconnect behavior during a controlled failover.
Design reasoning
● Technical workability: RDS detects infrastructure failure, promotes the standby, and updates the endpoint’s DNS mapping.
● Requirement fit: Database availability improves without customer-managed cluster and replication operations.
● Scene fit: The website needs continuity, not only analytical read scaling.
● Engineering common sense: Use managed failover where the database engine and edition support it.
Alternatives to reject
● RMAN backups support recovery but do not provide automatic failover.
● Oracle read replicas, where available, are not the same as a synchronous Multi-AZ standby.
● Self-managed Oracle RAC on EC2 adds licensing, cluster, storage, and patching complexity.
● A single RDS instance remains an availability risk.
Workflow: Application uses endpoint → primary failure → RDS promotes standby → DNS updates → application reconnects.
Reliable S3 Pre-Signed URLs
A pre-signed URL is valid only when created by an authorized AWS identity and used before expiration.
Recommended controls
● Give the portal a valid IAM role with permission for the requested S3 operation.
● Generate the URL with an expiration long enough for the user workflow.
● Keep bucket, key, method, and headers consistent with the signed request.
Design reasoning
● Technical workability: The signer’s credentials and permissions authorize the temporary request.
● Requirement fit: Users can upload or download without receiving permanent AWS credentials.
● Scene fit: Failures are intermittent, suggesting expiration or signing-identity problems.
● Engineering common sense: Time limits should reduce risk without expiring during normal transfers.
Alternatives to reject
● S3 versioning does not inherently invalidate an existing URL for the addressed operation.
● Developer console permissions are separate from the portal’s runtime identity.
● An ACL alone cannot compensate for missing signer credentials.
● A very short expiry can fail before a large transfer starts or completes.
Workflow: Portal role signs request → user receives URL → S3 validates signature and time → operation succeeds or expires.
Long-Running Human Workflows with SWF
Human tasks and batch activities need durable workflow state, retries, history, and coordination over long periods.
Recommended architecture
● Model the workflow in Amazon Simple Workflow Service.
● Use activity workers for automated tasks.
● Represent Mechanical Turk HIT creation and results as workflow activities.
● Store durable business outputs separately.
Design reasoning
● Technical workability: SWF preserves workflow history and coordinates long-running activities and retries.
● Requirement fit: Human responses and batch progress remain traceable.
● Scene fit: Mechanical Turk tasks can take unpredictable time.
● Engineering common sense: Do not keep long-running workflow state in compute memory.
Alternatives to reject
● RDS polling and Lambda workers require custom orchestration.
● AWS Config is a compliance service.
● Amazon MQ transports messages but does not manage workflow history and human task state.
● A simple queue does not express the complete process.
Workflow: Start batch → create HIT activities → wait durably → collect results → retry failures → complete batch.
Real-Time IoT Analytics Pipeline
Pet-collar telemetry needs a streaming path for current events and separate durable stores for raw or aggregated analytics.
Recommended architecture
● Ingest device events through Amazon Kinesis.
● Process or aggregate records in near real time.
● Store durable outputs and aggregates in Amazon S3.
● Load curated analytical data into Amazon Redshift for deeper queries.
Design reasoning
● Technical workability: Kinesis accepts continuous event streams, S3 provides durable object storage, and Redshift supports analytical SQL.
● Requirement fit: The pipeline supports timely monitoring and historical analysis at scale.
● Scene fit: IoT devices emit ongoing telemetry rather than occasional transactional records.
● Engineering common sense: Separate streaming ingestion from warehouse queries so analytical workloads do not block event intake.
Alternatives to reject
● Sending every device directly to a relational warehouse couples ingestion availability to warehouse capacity.
● Batch-only transfer delays operational insight.
● Using S3 alone stores records but does not provide real-time stream handling.
● Treating a queue as the full analytics platform still requires consumers, durable analytical storage, and query services.
Workflow: Device event → Kinesis stream → real-time processing → S3 archive and aggregates → Redshift loading → analytics.
Turning Application Logs into Immediate Alerts
Operational alerting requires centralized logs, a measurable error signal, a threshold, and a notification action.
Recommended architecture
● Install the CloudWatch agent on the on-premises servers.
● Send application logs to CloudWatch Logs.
● Create a metric filter for relevant error patterns.
● Create a CloudWatch alarm on the generated metric.
● Notify the operations team through the configured alarm action.
Design reasoning
● Technical workability: Metric filters convert matching log entries into numerical CloudWatch metrics that alarms can evaluate.
● Requirement fit: The design aggregates logs, automates analysis, and alerts immediately after the threshold is breached.
● Scene fit: The company needs actionable insight from application errors, not a general business-intelligence report.
● Engineering common sense: Alert on a defined service symptom and preserve the underlying logs for diagnosis.
Alternatives to reject
● The Kinesis agent and QuickSight add streaming and visualization components without simplifying alerting.
● CloudWatch Events is not a log storage destination.
● Athena is query based and does not continuously watch a metric filter.
● Managed Service for Prometheus is not installed as a general log agent, and Lambda transformation adds unnecessary custom code.
Workflow: Log event → CloudWatch Logs → metric filter → alarm threshold → operations notification → investigation.
Three-AZ Web and Database Availability
A 24-by-7 web application needs elastic compute across three Availability Zones and managed database failover.
Recommended architecture
● Run EC2 instances in a three-AZ Auto Scaling group.
● Place an Application Load Balancer before the fleet.
● Use RDS Multi-AZ for the relational database.
Design reasoning
● Technical workability: ALB routes to healthy targets, Auto Scaling replaces compute, and RDS promotes a synchronous standby.
● Requirement fit: Instance and AZ failures do not create a single service failure point.
● Scene fit: The application requires transactional database availability, not only read scaling.
● Engineering common sense: Every serving tier must remove its own failure point.
Alternatives to reject
● Read replicas scale reads but do not replace Multi-AZ writer failover.
● A design without a load balancer cannot distribute user traffic safely.
● One database instance remains a failure point even when the web tier spans AZs.
● More web servers do not repair database availability.
Workflow: User → ALB → healthy EC2 target → RDS writer; standby promoted on database failure.
Preserving Data During CloudFormation Stack Deletion
Infrastructure deletion should remove ongoing compute cost while preserving data through resource-appropriate deletion policies.
Recommended configuration
● Set the RDS resource to DeletionPolicy: Snapshot.
● Set the S3 bucket to DeletionPolicy: Retain.
● Delete the stack only after validating the policy changes.
Design reasoning
● Technical workability: CloudFormation creates a final RDS snapshot before deleting the database and leaves the S3 bucket outside stack deletion.
● Requirement fit: Database compute charges stop, while the database backup, images, and contestant data remain available.
● Scene fit: The application is finished and requires preservation rather than continued operation.
● Engineering common sense: Use native lifecycle behavior before creating duplicate storage workflows.
Alternatives to reject
● Snapshot is not supported as a deletion policy for S3 buckets.
● Retain on RDS keeps the live database and its ongoing cost instead of creating only a backup.
● S3 replication can preserve a second copy, but it adds another bucket, replication configuration, transfer, and storage cost without being required.
Workflow: Update template → review change → delete stack → verify RDS snapshot → verify retained bucket → manage retained resources separately.
Network Controls and Configuration History
Connectivity enforcement and resource-change history are different controls and should use different AWS services.
Recommended architecture
● Use security groups for stateful instance-level allow rules.
● Use network ACLs for stateless subnet-level allow and deny rules.
● Enable AWS Config to record resource configurations and historical changes.
● Add Config rules when compliance evaluation is required.
Design reasoning
● Technical workability: Security groups and NACLs evaluate traffic, while Config records how supported resources change over time.
● Requirement fit: The design provides both controlled EC2 communication and auditable security configuration history.
● Scene fit: The investigation needs current connectivity policy and previous configuration states.
● Engineering common sense: A traffic filter does not automatically provide historical governance evidence.
Alternatives to reject
● CloudTrail records API calls but is not the complete normalized configuration timeline provided by Config.
● VPC Flow Logs show accepted and rejected traffic, not full rule history.
● Systems Manager manages instances but does not replace VPC network controls.
● Route tables select paths and do not provide port-level policy.
Workflow: Packet → route → NACL → security group → instance; configuration changes → AWS Config history and compliance.
Automated Hybrid Patch Management
A hybrid estate can use one Systems Manager patch workflow for EC2 and registered on-premises servers.
Recommended architecture
● Install and register the SSM Agent on eligible servers.
● Define approved patch baselines in Systems Manager Patch Manager.
● Group servers with tags or patch-group attributes.
● Use Maintenance Windows to schedule AWS-RunPatchBaseline.
● Review patch compliance centrally.
Design reasoning
● Technical workability: Systems Manager manages both EC2 instances and hybrid activated servers through the agent.
● Requirement fit: Baselines and schedules remain synchronized and automated with low operational effort.
● Scene fit: The organization already maintains patch policy across on-premises and cloud servers.
● Engineering common sense: Standardize policy and reporting before increasing automation frequency.
Alternatives to reject
● Separate cron scripts create divergent patch logic and reporting.
● Rebuilding every server image does not patch long-lived on-premises machines directly.
● AWS Config records configuration but does not install operating-system patches.
● Manual console patching does not scale or prove compliance consistently.
Workflow: Register server → assign patch group → evaluate baseline → run in maintenance window → reboot as configured → report compliance.
Managed Secret Injection for ECS Fargate
Container credentials should be retrieved at runtime from a dedicated secret store rather than embedded in images or task-definition files.
Recommended architecture
● Store database credentials in AWS Secrets Manager.
● Encrypt the secret using AWS KMS.
● Configure managed credential rotation.
● Grant the ECS task execution role only the required secret and KMS permissions.
● Reference the secret ARN in the container definition for environment-variable injection.
Design reasoning
● Technical workability: ECS can resolve the secret during task startup and inject its value into the container environment.
● Requirement fit: Secrets Manager provides dedicated lifecycle management and rotation with minimal custom work.
● Scene fit: The workload already runs on Fargate, so no platform migration is needed.
● Engineering common sense: Keep plaintext out of source files, S3 task-definition artifacts, and container images.
Alternatives to reject
● Parameter Store SecureString is workable, but Secrets Manager better matches explicit managed rotation requirements.
● Migrating to EKS adds Kubernetes administration without improving this secret workflow.
● Encrypting credentials inside a task-definition file creates manual exposure and distribution risks.
Workflow: Create secret → configure rotation → grant role → reference ARN → deploy task → test rotation.
Managed Conversational Contact Center
A scalable call center needs managed telephony, speech understanding, conversational intent handling, and secure integration with business systems.
Recommended architecture
● Use Amazon Connect for inbound calls and contact flows.
● Use Amazon Lex for automatic speech recognition and natural-language intent handling.
● Invoke AWS Lambda functions to query or update business applications.
● Return relevant results through the contact flow.
Design reasoning
● Technical workability: Connect handles contact-center routing, Lex understands spoken requests, and Lambda integrates backend actions.
● Requirement fit: Callers can complete common tasks without an agent, while the service scales without call-center infrastructure management.
● Scene fit: Password changes and balance checks are intent-based self-service operations.
● Engineering common sense: Keep conversational interpretation separate from controlled business transactions.
Alternatives to reject
● MediaConnect transports professional video streams and is not a contact-center platform.
● Polly produces speech but does not perform speech recognition or intent detection.
● Ground Station supports satellite communications.
● Comprehend analyzes text but does not provide the interactive speech bot workflow supplied by Lex.
Workflow: Incoming call → Connect flow → Lex recognizes intent → Lambda validates and performs action → Connect returns response or routes to an agent.
Discovering PII and Reviewing S3 Access
Sensitive-data discovery and object-access investigation require different services.
Recommended architecture
● Use Amazon Macie to discover and classify PII in S3.
● Enable CloudTrail data events for relevant buckets.
● Review recent GetObject activity and the identities that performed it.
Design reasoning
● Technical workability: Macie analyzes object content and metadata; CloudTrail data events record object-level API activity.
● Requirement fit: Teams learn both where sensitive data exists and who accessed it.
● Scene fit: The protected information is stored in S3.
● Engineering common sense: Classification does not prove access, and access logs do not automatically identify PII.
Alternatives to reject
● GuardDuty detects threats but does not classify PII inside S3 objects.
● Inspector scans compute workloads and cannot install an agent on a bucket.
● Athena queries known data but does not automatically discover all PII.
● CloudWatch does not independently record S3 object GET calls.
Workflow: Macie scans bucket → findings identify sensitive objects → CloudTrail data events reveal recent access → investigate principals.
Preventing Accidental Public S3 Access
S3 Block Public Access is the direct preventive control for public ACLs and policies.
Recommended architecture
● Enable Block Public Access at the organization, account, bucket, or access-point scope required.
● Keep bucket policies least privilege.
● Use AWS Config or Security Hub for additional detection and reporting.
Design reasoning
● Technical workability: S3 rejects configurations and requests that would create blocked public access paths.
● Requirement fit: Users cannot accidentally expose protected buckets.
● Scene fit: The risk includes ACLs, bucket policies, and access points.
● Engineering common sense: Use the service-native preventive control before building custom evaluation logic.
Alternatives to reject
● A private canned ACL does not cover public bucket policies or access points.
● AWS Config detects noncompliance after the change.
● SCPs can deny APIs but cannot evaluate every resulting public configuration as completely.
● Manual reviews are slow and inconsistent.
Workflow: User attempts public setting → S3 Block Public Access evaluates → change or request rejected → monitoring records event.
Managed Document Versions with WorkDocs
A collaborative document service needs managed users, versions, encryption, APIs, and restoration of older content.
Recommended architecture
● Store documents in Amazon WorkDocs.
● Use its managed version history and access controls.
● Integrate application logic through WorkDocs APIs.
● Restore a selected older version as the current document when required.
Design reasoning
● Technical workability: WorkDocs provides document-oriented storage, users, versions, and managed encryption.
● Requirement fit: The platform avoids custom version and key-management code.
● Scene fit: Users collaborate on documents rather than generic objects or shared filesystem blocks.
● Engineering common sense: Use a document-management service when document lifecycle is the core requirement.
Alternatives to reject
● S3 can version objects, but distributing client-side master keys is unsafe and operationally heavy.
● S3 access logs do not provide rollback without versioning.
● EFS is a filesystem, not a managed document-version service.
● IAM cannot assign a different KMS key to each locked EFS file as proposed.
Workflow: User edits → WorkDocs creates version → application lists history → selected version restored as current.