Vertex Macro | Financial Cloud Cloud · Challenge
Weekend Productivity Challenge: Fab SPC Drift Synchronization Portal
Live application: Fab SPC Drift Synchronization Portal
Public source: GitHub — src/main.tsx
An AI-powered productivity application for semiconductor metrology engineers, developed with Kiro and implemented on AWS using Amazon S3, Amazon CloudFront, Amazon Route 53, AWS Certificate Manager, AWS WAF, Amazon MSK, AWS Lambda, AWS Glue, and AWS Lake Formation.
The Fab SPC Drift Synchronization Portal is an AWS-based system that helps semiconductor engineers identify equipment drift, review supporting evidence, and prioritize the safest next action. Manufacturing systems—including CD-SEM, SPC, FDC, APC, lot-route, and yield-watch systems—send operational events to Amazon MSK, which provides the real-time streaming backbone.
AWS Lambda processes these events by: Validating and normalizing incoming data Checking timestamps and data quality Calculating equipment-health and drift features Running AI-based risk inference Applying deterministic engineering rules Generating prioritized recommendations
The processed data, AI results, and audit evidence are stored in an Amazon S3 data lake. AWS Glue catalogs the datasets and manages schemas, while AWS Lake Formation controls access to sensitive manufacturing information. The engineer accesses the React portal through Amazon Route 53 and Amazon CloudFront. The React application is stored in a private Amazon S3 bucket, while AWS WAF protects the public entry point from unwanted traffic. The portal displays risk scores, reason codes, evidence, and recommended actions such as Release, Watch, Run SPC, Golden Wafer, Route Limit, APC Guard, or Hold Review. The system is advisory and human-in-the-loop: it helps engineers make faster, evidence-based decisions, but it cannot directly control equipment, hold lots, change routes, or modify APC settings. [codpayment...epoint.com]
Simple Data Flow
Manufacturing Systems
↓
Amazon MSK
↓
AWS Lambda Processing and AI
↓
Amazon S3 Data Lake
↓
AWS Glue + AWS Lake Formation
↓
Lambda API
↓
React Portal on S3 and CloudFront
↓
Engineer Reviews Recommendation
Vision & What the App Does
Statistical process control (SPC) is essential in semiconductor manufacturing, but a scheduled SPC result is only a snapshot. A CD-SEM can pass its last control check and then begin drifting while production lots continue to move. Gun-vacuum degradation, emission-current movement, deflector instability, stage vibration, contamination, matching changes, or simple measurement aging can emerge between scheduled checks. During this interval, an engineer may need to open several systems, compare timestamps manually, calculate fleet deviations, trace product routes, review equipment symptoms, and decide whether to release, watch, remeasure, reroute, or hold.
That investigation is both a manufacturing risk and a productivity problem.
The Fab SPC Drift Synchronization Portal is a personal AI-powered productivity workspace that compresses that investigation into one prioritized and explainable review experience. It synchronizes production-aware SPC indicators with fault detection and classification (FDC) telemetry, metrology results, tool-matching evidence, fleet behavior, and yield-watch context. Instead of requiring an engineer to reconstruct the complete situation manually, the portal provides a ranked risk board and makes the evidence behind each recommendation immediately visible.
The portal answers four practical questions:
● Which tool requires attention first?
● What changed during the SPC blind window?
● Which evidence supports the risk assessment?
● What is the safest next engineering action?
The current React experience contains six operational areas:
● Overview explains the control loop and displays fleet-level productivity indicators.
● Live Risk Board supports searching, sorting, ranking, and expanding individual CD-SEM records.
● FDC Health-Link maps equipment symptoms to metrology impact and potential yield exposure.
● Dynamic Matching compares each tool with fleet behavior and stable reference tools.
● Yield Triage assembles the evidence required to distinguish process movement from metrology error.
● Runbook converts the solution into a repeatable engineering workflow from proof of concept to production operation.
For each tool, the portal evaluates blind-window age, tool-matching gap, Mandel slope, fleet deviation, residual noise, trend behavior, layer criticality, and correlated FDC signals. It then produces a controlled recommendation:
● RELEASE
● WATCH
● RUN SPC
● GOLDEN WAFER
● ROUTE LIMIT
● APC GUARD
● HOLD REVIEW
The application is deliberately advisory. AI can detect patterns, rank work, summarize evidence, and recommend a next action, but it cannot command manufacturing equipment or bypass an engineer's approval. This human-in-the-loop boundary is part of the product design, not an afterthought.
From a productivity perspective, the portal works like a specialized task prioritizer, investigation assistant, and shift-handoff tool. It reduces context switching, makes each recommendation reproducible, and helps an engineer spend time on the highest-risk tool rather than manually searching for the next problem.
How I Built It
Kiro Development Workflow
I used Kiro as the AI-assisted development environment for requirements analysis, architecture design, implementation planning, code generation, refactoring, testing, and production hardening.
Rather than beginning with isolated UI code, I described the engineering outcome to Kiro:
Reduce the time required to identify CD-SEM drift between scheduled SPC checks, while preserving human authority over every manufacturing action.
I converted that outcome into a Kiro specification with three coordinated artifacts:
.kiro/specs/fab-spc-drift-portal/
├── requirements.md
├── design.md
└── tasks.md
Requirements created with Kiro
The requirements were written as testable behaviors:
● When normalized tool data arrives, the application must refresh the affected tool's risk state.
● When a user expands a tool, the application must display the value, engineering limit, event time, data-quality state, and reason code.
● When required data is stale, missing, incompatible, or outside the approved model domain, the application must return REVIEW_REQUIRED rather than infer a safe condition.
● Every recommendation must be reproducible from a versioned rule set, model artifact, feature set, and source-event range.
● No application component may initiate an equipment command, lot hold, route change, or APC change without an independently authenticated human approval process.
● The user interface must remain useful when AI explanation is unavailable by displaying deterministic evidence and a safe fallback message.
Design decisions captured with Kiro
Kiro helped separate the solution into five boundaries:
● React delivery and browser security;
● streaming ingestion with Amazon MSK;
● validation, normalization, feature calculation, and inference with AWS Lambda;
● immutable evidence storage in Amazon S3; and
● governed discovery and access with AWS Glue and AWS Lake Formation.
This separation prevents the front end from becoming a system of record. The browser displays results, but it does not possess Kafka credentials, access raw manufacturing data, calculate authoritative disposition, or invoke equipment interfaces.
Kiro steering and automated checks
Project steering files established persistent engineering rules:
.kiro/steering/
├── architecture.md
├── aws-security.md
├── react-typescript.md
├── streaming-contracts.md
├── ai-safety.md
└── testing-standards.md
The steering guidance required:
● TypeScript strictness and explicit domain types;
● accessible labels and keyboard-operable controls;
● private Amazon S3 origins rather than public website buckets;
● least-privilege AWS Identity and Access Management policies;
● schema versioning for every Amazon MSK event;
● idempotent AWS Lambda processing;
● no secrets or AWS credentials in React code;
● deterministic fallbacks for AI failure;
● model and rule version recording; and
● explicit human review for manufacturing-impacting recommendations.
Kiro hooks automated formatting, linting, unit tests, schema compatibility checks, dependency scanning, infrastructure validation, and a safety test that rejects any code path capable of issuing an autonomous equipment command.
React and TypeScript Front End
The user interface is implemented as a React single-page application. The supplied main.tsx defines typed objects for tools, actions, FDC mappings, runbook steps, tabs, sort keys, and risk tones. This is important because the interface represents controlled engineering states rather than arbitrary strings.
The implementation is componentized around operational responsibilities:
● Header and Hero establish current system context;
● KPIStrip summarizes the engineer's workload;
● LiveRiskBoard filters and ranks tools;
● Sparkline, LargeLineChart, and FleetBars visualize evidence;
● FDCHealth explains equipment-to-metrology relationships;
● DynamicMatching shows fleet comparisons;
● YieldTriage structures the investigation process;
● Runbook explains controlled rollout; and
● Sop embeds the expected user procedure.
The demonstration UI updates local sample values to show interactive behavior. In the production design, those fixture updates are replaced by authenticated API responses generated from Lambda-managed state. The presentation contract remains stable: the UI receives a sanitized tool summary, timestamp, risk score, confidence, reason codes, limits, recommendation, and evidence references.
The most important implementation decision was to separate three concepts:
● Observed facts — validated source values and timestamps;
● Analytical decisions — calculated features, engineering rules, and AI risk scores; and
● Presentation — React components that explain the result to a human.
A presentation bug cannot therefore change the authoritative recommendation, and an AI-generated explanation cannot overwrite a measured value.
AWS Services Used / Architecture Overview
The production architecture intentionally uses the following service scope:
● Kiro development
● Amazon S3, Amazon CloudFront, Amazon Route 53, AWS Certificate Manager, and AWS WAF for React delivery
● Amazon MSK as the streaming backbone
● AWS Lambda for data massage, feature engineering, AI inference, APIs, and event processing
● Amazon S3 data lake, AWS Glue, and AWS Lake Formation for governed evidence
Architecture at a glance
ENGINEER
|
v
Amazon Route 53
|
v
Amazon CloudFront ---- AWS WAF
| |
| +-- managed rules, rate control, request filtering
v
Private Amazon S3 bucket
React / TypeScript build artifacts
FAB EVENT PRODUCERS
CD-SEM | FDC | SPC | lot route | APC | yield-watch
|
v
Amazon MSK topics
|
v
AWS Lambda normalization and validation
|
+--> Amazon S3 raw evidence zone
|
v
AWS Lambda feature engineering and AI inference
|
+--> Amazon MSK recommendation topic
+--> Amazon S3 curated, feature, model-output, and audit zones
+--> React query API implemented with AWS Lambda
Amazon S3 data lake
|
v
AWS Glue Data Catalog and schema management
|
v
AWS Lake Formation governed access
It is a deployment-oriented view showing the web-delivery boundary, streaming boundary, processing boundary, and governance boundary.
Amazon S3, CloudFront, Route 53, ACM, and WAF: React Delivery
The production React build is generated by the CI/CD process and deployed to a private Amazon S3 bucket. The bucket is not configured as a publicly accessible S3 website. Amazon S3 Block Public Access remains enabled, and Amazon CloudFront uses Origin Access Control to retrieve objects from the bucket.
The deployment package contains content-hashed assets:
dist/
├── index.html
├── assets/index.<content-hash>.js
├── assets/index.<content-hash>.css
└── approved static assets
Content hashing allows JavaScript and CSS to use a long immutable cache lifetime. index.html receives a short cache lifetime because it points to the current asset versions. Deployment uploads immutable assets first and index.html last, preventing the HTML document from referencing files that are not yet available.
Amazon CloudFront configuration
Amazon CloudFront provides the global HTTPS entry point and performs:
● edge caching of React assets;
● compression;
● TLS termination;
● origin protection;
● security response headers;
● controlled invalidation during exceptional rollback; and
● a consistent application domain.
The response-headers policy includes:
Strict-Transport-Security
Content-Security-Policy
X-Content-Type-Options
Referrer-Policy
Permissions-Policy
frame-ancestors restriction
The content security policy limits scripts, styles, images, connections, and framing to explicitly approved sources. Source maps are not exposed in the general production distribution; they are retained separately for controlled debugging.
Route 53 and AWS Certificate Manager
Amazon Route 53 hosts the DNS record for the application and aliases the application hostname to the CloudFront distribution. AWS Certificate Manager supplies and renews the TLS certificate used by CloudFront. DNS validation keeps certificate renewal automated and auditable.
AWS WAF
AWS WAF is associated with CloudFront. Its web ACL includes:
● AWS managed common protections;
● known-bad-input protections;
● IP reputation protections;
● rate-based rules;
● size constraints for unexpected requests; and
● count-mode validation before a new blocking rule is promoted.
The React application is static, but WAF remains valuable because the same CloudFront entry point can route controlled API paths to Lambda-backed endpoints. WAF metrics are monitored before thresholds are tightened so legitimate engineers are not blocked by an untested rule.
S3 production controls
The application bucket uses:
● S3 Block Public Access;
● object versioning;
● default encryption;
● bucket-owner-enforced object ownership;
● access logging to a separate protected log bucket;
● lifecycle rules for superseded build artifacts; and
● a bucket policy that grants reads only through the designated CloudFront distribution.
A release manifest records the source commit, Kiro specification revision, build checksum, asset list, and deployment time. Rollback changes index.html to reference the previous immutable asset set.
Amazon MSK: The Streaming Backbone
Amazon Managed Streaming for Apache Kafka is the durable event backbone. It decouples high-frequency manufacturing producers from normalization, AI processing, evidence storage, and web presentation.
The production topic structure is domain based:
fab.fdc.telemetry.v1
fab.metrology.measurement.v1
fab.spc.result.v1
fab.lot.route.v1
fab.apc.feedback.v1
fab.yield.watch.v1
fab.tool.recommendation.v1
fab.processing.deadletter.v1
Event contract
Every event contains a common envelope:
{
"event_id": "source-generated-or-content-derived-id",
"event_type": "fdc.telemetry",
"schema_version": "1.2.0",
"event_time": "2026-07-12T10:15:32.481Z",
"ingest_time": "2026-07-12T10:15:33.014Z",
"fab_id": "tokenized-fab-reference",
"tool_id": "tokenized-tool-reference",
"lot_id": "tokenized-lot-reference",
"recipe_id": "approved-recipe-reference",
"sequence": 1845531,
"quality": {
"complete": true,
"clock_status": "synchronized"
},
"payload": {}
}
event_time represents when the physical or manufacturing event occurred. ingest_time represents when the platform received it. This distinction is essential: risk windows are based on event time, while platform latency is measured from ingest time.
Partitioning and ordering
Partition keys are selected by business ordering requirement:
● tool telemetry is partitioned by tool_id;
● lot lifecycle events are partitioned by lot_id;
● recommendations are partitioned by tool_id; and
● cross-fleet analytics uses curated feature events rather than assuming global Kafka order.
Ordering is guaranteed only within a partition. Any algorithm requiring several topics uses event-time windows, watermarks, and late-event handling instead of assuming that arrival order is correct.
Schema governance
AWS Glue Schema Registry manages compatible Avro or JSON schemas. The compatibility policy permits additive changes with defaults and rejects incompatible field reinterpretation. A breaking contract creates a new topic version. Lambda logs the writer schema, reader schema, and compatibility result so replay remains explainable.
Reliability and capacity
Partition count is calculated from peak event rate, average record size, target consumer parallelism, retention, and expected growth. The design monitors:
● producer error rate;
● broker storage usage;
● bytes in and out;
● under-replicated partitions;
● consumer lag;
● oldest unprocessed event age; and
● dead-letter volume.
Producers use TLS, authenticated access, acknowledgements appropriate to durability, retries with jitter, compression, bounded local buffering, and an explicit failure path. Multi-Availability Zone replication protects the event stream from a single-zone failure. Recovery testing includes restarting consumers, replaying offsets, and proving that recommendations are not duplicated.
AWS Lambda: Data Massage, AI, and Event Processing
AWS Lambda performs the serverless processing work between Amazon MSK and the governed S3 data lake. The design separates responsibilities into small functions rather than creating one untestable handler.
msk-normalizer
feature-builder
risk-inference
recommendation-publisher
evidence-writer
portal-query-api
feedback-processor
MSK event processing
An event source mapping polls the approved MSK topics and invokes the normalizer with bounded batches. Function settings are tuned for the event size and latency objective rather than left at arbitrary defaults. Reserved concurrency protects downstream storage, and the failure policy prevents a poison record from blocking a partition indefinitely.
For each Kafka record, the normalizer performs:
● payload decoding;
● writer-schema identification;
● schema validation;
● event-envelope validation;
● unit standardization;
● timestamp and clock-quality checks;
● tokenized identity enrichment;
● duplicate detection;
● data-quality classification;
● raw evidence persistence; and
● normalized event publication.
Malformed records are written to the dead-letter topic with the original topic, partition, offset, schema version, error code, and encrypted payload reference. They are never silently discarded.
Idempotency
Kafka consumers are designed for at-least-once delivery. Therefore, every side effect is idempotent. The processing key is derived from the source identity, partition, offset, schema version, and canonical event ID. S3 object keys and recommendation IDs incorporate this identity. A retry produces the same result instead of a second case.
The calculation output stores the exact contributing topic/partition/offset ranges. This allows a recommendation to be reproduced from source evidence after a replay or audit.
Feature engineering
The feature Lambda calculates point-in-time-correct values such as:
● elapsed time since validated SPC;
● volume of production exposure during the blind window;
● rolling mean and standard deviation;
● exponentially weighted movement;
● trend slope and step-change indicators;
● residual three-sigma estimate;
● site-to-site range;
● tool-matching gap;
● Mandel slope distance from one;
● robust fleet deviation;
● distance from approved master anchors;
● FDC baseline distance;
● missingness and lateness indicators; and
● layer-criticality and route-exposure features.
Features are calculated only from information available at the decision timestamp. This prevents future-data leakage during both production inference and offline model evaluation.
AI inference in Lambda
To keep the implementation inside the defined service scope, the approved AI model is packaged as a versioned Lambda layer or container-image dependency. The trained artifact and metadata are stored in a governed S3 model prefix. At cold start, the function loads the approved model version; warm invocations reuse it.
The AI output is structured rather than free-form:
{
"risk_probability": 0.91,
"severity": "HIGH",
"confidence": 0.84,
"model_version": "spc-risk-2026-07-12.3",
"reason_codes": [
"FLEET_DEVIATION_HIGH",
"MANDEL_SLOPE_OUTSIDE_RANGE",
"DEFLECTOR_SIGNAL_UNSTABLE"
],
"domain_status": "IN_DOMAIN"
}
A deterministic policy evaluator combines model output with approved engineering rules. AI may raise an early warning, but it cannot weaken a hard rule or issue an autonomous manufacturing action.
Safe decision policy
IF required data is missing, stale, or incompatible
RECOMMEND REVIEW_REQUIRED
ELSE IF a hard engineering limit is breached
AND an independent signal corroborates the breach
RECOMMEND HOLD_REVIEW
ELSE IF the model predicts a near-term breach
AND confidence and domain checks pass
RECOMMEND RUN_SPC or GOLDEN_WAFER
ELSE IF fleet or FDC warning criteria are met
RECOMMEND WATCH
ELSE
RECOMMEND RELEASE
Observability
Every Lambda invocation emits structured logs and operational metrics:
RecordsProcessed
ValidationFailureCount
DuplicateCount
LateEventCount
FeatureFreshnessSeconds
InferenceDurationMs
OutOfDomainCount
RecommendationCountByType
DeadLetterCount
EndToEndDecisionLatencyMs
Logs include correlation ID, model version, rule-set version, source offsets, and error classification, but exclude unnecessary sensitive manufacturing values. Alarms focus on user impact: growing MSK lag, stale recommendations, validation spikes, inference failures, and dead-letter growth.
S3 Data Lake, AWS Glue, and Lake Formation: Governed Evidence
A production AI application requires more than a model. It requires reproducible evidence, controlled training data, discoverable definitions, lineage, and access governance.
The Amazon S3 data lake is organized into zones:
s3://fab-spc-data/raw/
s3://fab-spc-data/validated/
s3://fab-spc-data/curated/
s3://fab-spc-data/features/
s3://fab-spc-data/models/
s3://fab-spc-data/predictions/
s3://fab-spc-data/evidence/
s3://fab-spc-data/feedback/
s3://fab-spc-data/quarantine/
Raw zone
The raw zone preserves source events with minimal transformation. Objects are immutable, encrypted, versioned, and partitioned by source domain and event date. The original Kafka metadata is retained for replay and audit.
Validated and curated zones
The validated zone contains schema-conforming events with quality flags. The curated zone contains standardized engineering entities such as tool state, metrology observations, matching results, route exposure, and approved outcomes. Columnar files and practical partitioning reduce scan cost and improve analytic performance.
Feature and model zones
The feature zone stores point-in-time-correct training and inference features. Each dataset records:
● feature definition version;
● source event range;
● generation code version;
● cutoff timestamp;
● quality results; and
● approved use.
The model zone stores the model artifact, preprocessing definition, feature order, evaluation report, threshold configuration, model card, approval state, checksum, and rollback predecessor. Lambda is allowed to load only an explicitly approved model prefix.
AWS Glue
AWS Glue provides the technical catalog and schema layer. Glue crawlers are used selectively; production tables with strict contracts are managed from explicit definitions so an unexpected file cannot silently redefine a critical column.
The Data Catalog records:
● database and table definitions;
● file formats and partitions;
● schema versions;
● owners and descriptions;
● data classification;
● quality status; and
● source-to-curated lineage references.
AWS Glue jobs support larger offline transformations when processing exceeds the appropriate Lambda execution profile. Glue-generated training datasets use the same feature definitions and event-time rules as the Lambda inference path, reducing training-serving skew.
AWS Lake Formation
AWS Lake Formation governs access to cataloged S3 data. Permissions are granted by role and data purpose rather than by broad bucket access. LF-tags classify domains such as fab, tool family, product sensitivity, evidence type, and approved use.
Example access boundaries include:
● front-end delivery roles cannot access the data lake;
● Lambda normalization can write raw and validated data but cannot approve models;
● the inference Lambda can read only approved feature definitions and model artifacts;
● engineering analysts can query authorized curated data;
● model-development roles can read approved training data but not unrestricted raw identifiers; and
● auditors can read evidence and lineage without modifying operational datasets.
Column- and row-level controls prevent unnecessary exposure. Cross-account sharing uses governed catalog permissions rather than copying uncontrolled datasets.
Retention and evidence integrity
Lifecycle rules move historical evidence to lower-cost S3 storage classes according to business and regulatory requirements. Quarantine retention is deliberately limited. Model, recommendation, and approval evidence remains available for the required audit period.
Checksums, object versioning, protected access logs, and separation of duties make the evidence chain tamper-evident. A recommendation can be traced from the UI to the Lambda output, model version, feature record, curated dataset, and original MSK offsets.
Complete AI Development Lifecycle
AI development covers much more than training an algorithm. For this application, the complete lifecycle is implemented around the AWS service scope.
1. Problem definition
The model predicts whether a tool is likely to produce unacceptable metrology behavior before the next scheduled SPC opportunity. The target is not “predict every anomaly.” The target is to provide enough lead time for a useful human review while controlling unnecessary investigations.
2. Label definition
Labels are derived from approved engineering outcomes:
● confirmed SPC out-of-control events;
● golden-wafer results;
● confirmed equipment findings;
● matching or linearity failures;
● approved route or hold decisions;
● false-alarm dispositions; and
● yield backtrace conclusions.
An engineer's initial recommendation is not automatically treated as truth. Final approved disposition and investigation outcome are stored separately to prevent self-reinforcing labels.
3. Data preparation
Amazon MSK captures operational events, Lambda validates and aligns them, S3 stores immutable history, AWS Glue creates reproducible datasets, and Lake Formation controls who may use each dataset. Training examples use event-time cutoffs so they cannot include information that became available after the prediction point.
4. Dataset splitting
Random row splitting would leak repeated tool behavior across train and test data. The production process therefore applies chronological splits and groups by tool or tool family where appropriate. A final untouched time period measures realistic forward performance.
5. Feature development
Features are documented with purpose, unit, source, expected range, missing-value behavior, owner, and leakage risk. The same versioned transformations are used for offline datasets and Lambda inference.
6. Model selection
The first production candidate favors a compact, interpretable model that can execute efficiently in Lambda. Candidate models are compared against deterministic baselines. A more complex model is accepted only when it improves operational metrics without unacceptable latency, instability, or loss of explainability.
7. Evaluation
Because true drift events are uncommon, accuracy is not the primary metric. Evaluation includes:
● precision and recall;
● precision-recall area;
● false-negative cost;
● recall at available engineer review capacity;
● probability calibration;
● average warning lead time;
● performance by tool family, layer, recipe, and event-quality state;
● out-of-domain detection; and
● comparison with the existing deterministic process.
Threshold selection is a manufacturing decision shared with engineering owners. It balances missed-drift risk against review workload.
8. Explainability
The model returns stable reason codes and contributing feature values. The UI presents these beside the approved limits and source time. The explanation is constrained to observed evidence; it does not invent a root cause.
9. Bias and coverage assessment
The evaluation checks whether the model performs consistently across tool families, layers, recipes, maintenance states, and data-quality conditions. In this industrial context, unfairness can appear as systematically poorer detection for a less-common tool or process family. Unsupported groups are marked out of domain and routed to human review.
10. Model approval and versioning
Each approved model package includes:
model artifact
preprocessor artifact
feature specification
training-data snapshot reference
evaluation report
slice metrics
threshold configuration
model card
known limitations
approval record
rollback model
artifact checksums
Only an approved S3 model prefix can be loaded by the inference Lambda. Model promotion is separated from model development through Lake Formation permissions and deployment controls.
11. Deployment
A candidate begins in shadow mode. It receives production features but cannot influence the displayed recommendation. Predictions are compared with the current rule set and later outcomes. After approval, a small controlled portion of requests uses the candidate while the previous model remains immediately available for rollback.
12. Monitoring
Production monitoring covers four categories:
● service health: Lambda errors, duration, throttles, cold starts, MSK lag;
● data health: missing fields, schema changes, late events, range violations;
● model health: feature drift, score distribution, out-of-domain rate, calibration;
● business health: warning lead time, accepted recommendations, false alarms, prevented exposure, and engineer review time.
When labels arrive later, the feedback processor joins predictions with approved outcomes and calculates delayed quality metrics.
13. Retraining
Retraining is triggered by an approved schedule, sufficient new labels, feature drift, performance degradation, or a meaningful equipment/process change. Retraining never automatically promotes a model. The full evaluation and approval gate runs again.
14. Responsible AI and safety
The application applies the following controls:
● AI is advisory only.
● Every recommendation displays uncertainty and source freshness.
● Missing or stale evidence produces review, not release.
● Hard engineering limits cannot be overridden by the model.
● Sensitive identifiers are minimized and tokenized.
● Training and inference access is governed through Lake Formation and IAM.
● Model artifacts, datasets, rules, and outputs are versioned.
● Engineers can accept, reject, or correct recommendations.
● AI failure falls back to deterministic rules and evidence.
● Autonomous equipment commands are outside the application's IAM permissions and network path.
Production Testing and Release Strategy
Kiro-generated tasks include tests for the UI, streams, Lambda processing, data governance, and AI behavior.
React tests
● component rendering;
● search and sort behavior;
● keyboard navigation;
● accessible names and contrast;
● stale-data and AI-unavailable states;
● threshold-boundary display; and
● mobile layout behavior.
MSK and Lambda tests
● compatible and incompatible schemas;
● duplicate delivery;
● reordered events;
● late arrivals;
● partial batch failure;
● poison messages;
● replay from earlier offsets;
● large batches;
● downstream timeout; and
● proof that retry does not duplicate a recommendation.
Data lake and governance tests
● S3 public-access denial;
● encryption and versioning;
● partition and schema validation;
● Glue catalog consistency;
● Lake Formation positive and negative authorization tests;
● lineage completeness; and
● retention-policy validation.
AI tests
● future-data leakage checks;
● feature parity between training and inference;
● model serialization and cold-start loading;
● boundary and missing-value behavior;
● probability calibration;
● slice evaluation;
● out-of-domain handling;
● reason-code stability;
● safe fallback; and
● model rollback.
The release pipeline promotes immutable artifacts through development, staging, and production. A production release records the Git commit, Kiro spec revision, React manifest, infrastructure revision, schema versions, Lambda versions, rule-set version, and approved model version.
What I Learned
The first lesson was that an industrial productivity application should reduce the number of decisions an engineer must reconstruct, not simply add another dashboard. The useful output is a prioritized case containing fresh evidence, explicit limits, uncertainty, and the smallest safe next action.
The second lesson was that streaming correctness is operational correctness. A correct formula can still produce the wrong decision if it uses arrival time instead of event time, assumes global ordering, duplicates a side effect after retry, or evaluates stale evidence. Amazon MSK and idempotent Lambda processing made replay, traceability, and failure recovery part of the design.
The third lesson was that AI development starts with data contracts and ends with monitored human outcomes. The model itself is only one artifact. S3 evidence zones, Glue metadata, Lake Formation controls, point-in-time feature construction, approval records, reason codes, feedback, drift monitoring, and rollback are equally important.
The fourth lesson came from Kiro. AI-assisted coding is most effective when requirements, architecture, security rules, and tests constrain generation. Kiro specs maintained traceability from the productivity problem to implementation tasks. Steering preserved AWS and safety conventions. Hooks placed repetitive quality checks directly inside the engineering workflow.
Finally, I learned that a production-ready AI system must be designed for uncertainty. The portal never treats missing data as healthy data, never treats a probability as an equipment command, and never hides the distinction between observed evidence, deterministic rules, and AI prediction.
The result is an AWS-focused productivity tool that helps an engineer understand what needs attention, why it matters, which evidence supports the decision, and what to review next.
Link to App or Repo
● Working deployed application: https://vertexmacro.com/cloud_club/demo/dashboard/aws_factory_automation_portal.html
● Public GitHub source: https://github.com/dchan-dev/aws-fab-spc-drift-synchronization-portal-ts/blob/main/src/main.tsx
Application SOP — Daily Fab-Duty Operation and Business Value
The portal delivers value only when its risk indicators lead to a consistent engineering routine. This standard operating procedure explains how metrology engineers, equipment engineers, process engineers, process-integration engineers, and yield engineers use the application during a normal shift. It also makes the productivity benefit visible to a challenge reviewer: the application does not merely display charts; it converts fragmented evidence into a prioritized, repeatable work process.
Safety boundary: The portal is an advisory decision-support application. Limits shown in the public demonstration are learning baselines. Production limits must be approved for the applicable node, product, layer, customer, and module. Final release, route-limit, APC, maintenance, and hold decisions remain with authorized fab personnel.
Intended audience and value by role
Metrology engineer or CD-SEM owner
The metrology owner uses the portal to detect drift between scheduled SPC checks, evaluate tool matching, compare a suspect tool with the fleet, and determine whether physical confirmation is required. The productivity value is a single evidence view for TMG, Mandel slope, residual noise, blind-window age, fleet deviation, and recent movement instead of a manual investigation across unrelated systems.
Equipment engineer
The equipment engineer works from the hardware-health perspective: Is the CD-SEM healthy enough to continue measuring production lots, which FDC signal is driving the loss of measurement credibility, and what containment is required? The FDC Health-Link panel directs the investigation toward gun vacuum, emission current, deflector DAC, or stage vibration and connects that path to an observable metrology effect.
Process or process-integration engineer
PE and PIE use the portal before changing lithography or etch settings. If the measurement itself is questionable, a process correction can move the process in the wrong direction. The portal provides the route history, metrology credibility indicators, APC-guard status, and engineering recommendation required to determine whether process action should wait for verification.
Yield engineer
The yield engineer uses the portal to backtrace a yield-watch lot to the measuring CD-SEM, the FDC state, the last trusted SPC point, the fleet comparison, and the APC exposure window. This accelerates the distinction among a real process excursion, a metrology false alarm, and process movement caused by suspect feedback.
Why the SOP matters
A scheduled SPC result may remain green while the physical tool changes during the following hours. That interval is the SPC blind window. The portal closes the operational gap by joining the last trusted qualification point with current production exposure, FDC movement, fleet behavior, matching quality, and yield context.
The expected productivity chain is:
Scattered tool, lot, SPC, FDC, matching, and yield evidence
↓
One ranked Live Risk Board
↓
One expandable evidence view with reason codes
↓
One controlled recommendation and named owner
↓
Faster verification, containment, passdown, and closure
The intended manufacturing defense is:
Detect hidden CD-SEM drift
↓
Protect measurement credibility
↓
Prevent suspect data from entering APC decisions
↓
Reduce unnecessary process changes and lot exposure
↓
Improve investigation speed and shift-to-shift continuity
Portal concepts an operator must understand
● Blind window: elapsed time between the last trusted SPC or golden-wafer result and the current production measurement.
● TMG: the tool-matching gap. The demonstration baseline is no more than 10% of target CD; the approved production percentage may be tighter.
● Mandel slope: multi-CD linearity indicator. The demonstration learning range is 0.98–1.02.
● Residual 3σ: random uncertainty after the fitted trend is removed. A rising residual can be more dangerous than a simple mean offset because the measurement becomes less repeatable.
● Fleet σ: distance from the governed fleet baseline. More than 2σ is an early warning; more than 3σ is a hold-review candidate in the demonstration logic.
● Site-to-site delta: spatial measurement range used to identify stage, vibration, scan-linearity, or center-edge artifacts.
● FDC health link: the relationship between a hardware signal and an observed metrology effect.
● Action label: an advisory next step that requires human review.
All displayed metrics must include a current timestamp and quality state in production. A stale green value is not evidence of health.
Start-of-shift SOP — first ten minutes
● Sign in through the approved application URL and verify that the displayed data-freshness timestamp is current.
● Review the KPI strip: tools watched, average SPC blind window, fleet out-of-control count, FDC health links, dynamic-limit state, and yield-watch workload.
● Open LIVE RISK BOARD.
● Sort by RISK and expand every high-risk tool, using 75 as the demonstration review threshold.
● Read the complete evidence, not only the color: layer, symptom, TMG, slope, fleet σ, residual, blind window, trajectory, reason codes, model version, and recommendation.
● Sort by BLIND WINDOW to identify tools whose last physical confidence point is old.
● Sort by FLEET σ to identify a tool separating from its peers even if its most recent SPC result remains green.
● Sort by RESIDUAL to identify developing repeatability, vacuum, vibration, charging, or beam-stability risk.
● Open FDC HEALTH-LINK for each high-risk tool and identify the most plausible hardware investigation path.
● Record the high-risk tool, affected layer, owner, current containment, exposed-lot window, and required exit criterion in the shift passdown.
Yellow trends are actionable information. The SOP does not require engineers to wait for a red alarm or a traditional three-sigma failure before reviewing a developing risk.
Before releasing a critical-layer lot
Before Gate, Fin, tight Contact/Via, risk-ramp, or yield-watch lots are measured or released, the responsible engineer verifies:
● the tool is not in HOLD REVIEW;
● the last trusted SPC or golden-wafer result is fresh enough for the layer;
● TMG remains within the approved layer-specific budget;
● Mandel slope remains within its approved linearity range;
● residual 3σ remains within the approved repeatability limit;
● site-to-site delta remains within the approved spatial limit;
● fleet deviation is not a hold candidate;
● the FDC fingerprint is stable or has an accepted explanation;
● the prediction is in domain and its required inputs are complete; and
● no unresolved APC guard applies to the measurement window.
If any condition is uncertain, the recommendation must become more conservative. Available containment includes running SPC, running a golden wafer, routing to a master tool, restricting the suspect tool to approved non-critical layers, guarding APC feedback, or calling a hold review.
How to interpret each portal action
RELEASE
Current evidence does not show a material tool-health threat to measurement credibility. Continue approved production use and trend monitoring. Release does not cancel normal qualification requirements and does not justify ignoring new FDC movement.
WATCH
The tool has an early drift signature, but a confirmed failure has not been established. Review the previous 8–24 hours of FDC behavior, compare the latest golden-wafer result, inspect maintenance and event logs, increase monitoring frequency, and prepare physical verification if the trend continues.
RUN SPC
The tool may not show a severe abnormality, but its physical confidence point is too old for the current exposure. Run the approved SPC or golden-wafer check before more critical lots. If constrained operation is necessary, only an authorized owner can approve an explicitly limited use.
GOLDEN WAFER
Virtual evidence is suspicious enough to require physical standard-wafer confirmation. Pause critical-layer measurement, use the approved recipe, and compare mean, three-sigma, TMG, slope, residual, site-to-site delta, and measurement profile with the governed baseline. A failed result escalates to hold review.
ROUTE LIMIT
The tool may remain usable for specifically approved, less-sensitive layers but is not suitable for critical work. Notify the dispatcher and module owner, restrict the affected layers, define an owner, and document measurable exit criteria before returning to full release.
APC GUARD
Measurement bias or noise may contaminate feedback. Identify lots measured during the suspect window, notify the APC owner, and decide whether affected measurements must be paused, excluded, reviewed, or repeated on a trusted tool. The portal cannot modify APC by itself.
HOLD REVIEW
The tool is a strong candidate for removal from critical measurement pending cross-functional review. Freeze the affected exposure window, check FDC, run physical confirmation, inspect image or waveform quality, compare against the master and fleet, determine affected lots, and select release, route limit, maintenance, calibration, or requalification through authorized procedures.
FDC-first diagnostic SOP
When the portal changes to WATCH, GOLDEN WAFER, ROUTE LIMIT, APC GUARD, or HOLD REVIEW, use the FDC Health-Link panel to select the first diagnostic branch.
Gun-vacuum branch
Typical portal evidence includes rising residual 3σ, worsening matching precision, random false alarms, apparent blur, or reduced peak-to-base ratio. Review chamber-vacuum trends, pump events, pressure spikes, recent venting or recovery, contamination indicators, and profile stability. Physical confirmation and stabilization take priority over re-baselining.
Emission-current branch
Typical evidence includes a mean-offset step, consecutive values on one side of the center line, an emission ramp, or concern that biased CD entered APC. Review emission current, extraction voltage, probe-current stability, gun age, approved recovery history, and changes after intervention. Guard the affected feedback window until credibility is restored.
Deflector-DAC branch
Typical evidence includes Mandel slope outside its approved range, dense/isolated bias, multi-CD mismatch, or site-to-site movement. Review deflector stability, scan-linearity calibration, image-shift correction, measurement-box placement, edge-algorithm configuration, and multi-CD standard-wafer behavior. Critical layers remain restricted until the linearity exit criterion passes.
Stage-vibration branch
Typical evidence includes widening residual, increasing site-to-site range, random site instability, or a false center-edge signature. Review the stage-vibration sensor, facility events, settling time, positioning logs, interferometer behavior, and repeated measurement at the same site. Repeatability and spatial stability must recover before full release.
Metric-driven response SOP
High blind window
Review all tool and facility events since the last trusted qualification point. Pay special attention to venting, beam restart, aperture work, preventive maintenance, vibration, temperature, and facility alarms. Run physical verification before critical work when the approved freshness boundary is exceeded.
TMG above the approved limit
Compare the suspect tool with its approved master, determine whether the change is a simple offset or includes wider residual error, verify recipe and correction-table versions, and run a multi-site or multi-CD standard wafer. Restrict cross-tool dispatch until matching returns within the approved rule.
Mandel slope outside the approved range
Perform a multi-CD linearity check, examine deflector and scan calibration evidence, compare dense and isolated features, and review edge-detection configuration. A tool that agrees at one CD but fails across CD sizes must not be treated as matched for critical layers.
Residual 3σ above the approved limit
Run a repeatability check on a standard wafer, inspect measurement profiles, review vacuum and emission stability, check facility and stage vibration, and verify focus and stigmator behavior. Guard APC if noisy measurements may have entered feedback.
Fleet deviation above the approved threshold
Compare product-lot means by tool, determine whether peers measuring the same product and layer remain stable, verify the suspect tool against the master, and backtrace lots measured during the deviation window. Fleet comparison is specifically intended to expose hidden inline drift before traditional SPC failure.
Site-to-site delta above the approved limit
Review site-level data rather than only the wafer mean. Check stage positioning, vibration, image-shift correction, interferometer evidence, and a multi-site standard-wafer result. Do not make a center-edge process disposition until a metrology artifact has been excluded.
HOLD REVIEW containment and recovery procedure
● Confirm the signal. Verify freshness, quality, risk, TMG, slope, fleet σ, residual, blind window, and reason codes.
● Freeze additional exposure. Stop unapproved critical-layer measurement and identify lots measured since the last trusted state.
● Select the FDC path. Review vacuum, emission/extraction, deflector, vibration, and applicable facility or thermal evidence.
● Run physical confirmation. Use an approved golden wafer, standard wafer, multi-CD wafer for slope concerns, or multi-site wafer for spatial concerns.
● Review image and waveform evidence. Inspect peak-to-base ratio, edge slope, baseline movement, X/Y asymmetry, blur, and tailing where those measurements are available.
● Compare with trusted peers. Use the master tool, governed fleet baseline, and comparable product distribution.
● Decide disposition. Select release, watch, route limit, APC guard, calibration, preventive maintenance, or continued hold through authorized review.
● Document and hand over. Record the hypothesis, evidence, affected-lot window, owner, due time, action, and exit criteria.
Exit criteria for return to full release
A tool does not return from route limit or hold review merely because an alarm clears. The owner verifies that:
● the approved standard or golden wafer passes;
● TMG is within the approved rule;
● Mandel slope is within the approved range;
● residual 3σ and site-to-site delta are within limits;
● fleet deviation has returned below the approved threshold;
● the abnormal FDC signal has returned to baseline or an accepted stable state;
● image or waveform quality has recovered where applicable;
● suspect APC feedback has been dispositioned;
● exposed lots have been reviewed by the responsible functions; and
● shift records contain the recovery evidence and exit criteria.
Yield-loss review evidence pack
When a yield engineer, process engineer, or process-integration engineer opens a case, the portal should produce a governed evidence pack rather than an informal collection of screenshots:
case and request identifier
product, layer, recipe, lot and wafer references
tool route and ADI/AEI tool pairing
last trusted SPC or golden-wafer timestamp
blind-window duration and exposed-lot window
TMG and matching history
Mandel slope and multi-CD evidence
residual 3σ and site-to-site delta
fleet deviation and peer distribution
FDC changes relative to baseline
maintenance and equipment-event history
APC feedback status
model and rule-set versions
reason codes and confidence
current containment, owner and due time
review disposition and exit criteria
The evidence pack supports one of five concise engineering conclusions:
● the tool is credible and a process cause is more likely;
● the tool is abnormal and a metrology false alarm is possible;
● the tool is abnormal and APC feedback contamination is possible;
● the evidence is mixed and master-tool remeasurement is required; or
● the evidence is mixed and standard-wafer verification is required.
End-of-shift passdown SOP
A passdown must describe measurement credibility, not merely “tool OK” or “tool NG.” The shift record contains:
Date and shift:
Engineer:
Tool-health summary:
High-risk CD-SEM tools:
Affected layers and lots:
FDC abnormal signals:
TMG / Mandel slope / residual 3σ / site delta / fleet σ:
Last trusted qualification time:
Current containment:
Golden-wafer or standard-wafer result:
APC guard status:
Maintenance or calibration action:
Owner and next action:
Due time:
Exit criteria for release:
This structured handoff is one of the application's most important productivity benefits. It prevents the next shift from repeating the same search and preserves an auditable connection between alert, evidence, action, and outcome.
Operational example — CDSEM-05
Assume the portal shows:
Tool: CDSEM-05
Layer: Contact/Via
Action: GOLDEN WAFER
Symptom: Vacuum degradation with rising residual
Blind window: 6.9 hours
TMG: 0.19 nm
Mandel slope: 1.006
Fleet deviation: 2.4 sigma
Residual 3σ: 0.21 nm
The slope remains inside the demonstration range, so multi-CD linearity is not the strongest signal. Residual is above the learning baseline and the tool is separating from its peers, while the vacuum trend provides a plausible hardware path. Because Contact/Via can be critical, the operator does not release blindly.
The engineer restricts critical use, runs the approved golden wafer, reviews the previous 24 hours of vacuum behavior, checks measurement-profile stability, performs a repeatability test, and compares the result with the master. If residual remains high, the owner opens the approved vacuum-recovery, contamination, or maintenance procedure. PE/PIE are notified about lots measured after the residual limit was crossed, and the APC owner reviews whether their feedback must be excluded.
This example shows the application's value clearly: one ranked case leads directly to the correct evidence, responsible roles, containment, and measurable exit criteria.
Actions the application must prevent
● Do not release a critical lot solely because the last SPC result was green when FDC changed afterward.
● Do not treat a mean-offset correction as sufficient when residual uncertainty is widening.
● Do not re-baseline before investigating the physical cause.
● Do not ignore an emerging fleet outlier solely because conventional SPC has not failed.
● Do not allow suspect CD data to influence APC without owner review.
● Do not treat one standard-wafer site as permanently stable; repeated exposure can age the reference.
● Do not close a hold review without documented exit criteria.
● Do not allow the AI score alone to make a hold or release decision.
SOP success metrics
The production team measures whether the portal improves work rather than simply generating alerts:
● drift cases detected before formal SPC failure;
● suspect APC feedback events prevented or reviewed;
● lots protected through route limit or APC guard;
● reduction in repeated CD-SEM false alarms;
● reduction in time to assemble a yield-review evidence pack;
● percentage of high-risk cases with complete evidence and named owners;
● alert-to-standard-wafer confirmation time;
● hold-review-to-disposition time;
● shift-passdown completeness;
● recommendation acceptance, rejection, and correction rates; and
● recurrence after tool recovery.
These outcomes are written to the governed S3 feedback zone. AWS Lambda aligns them with the original recommendation, and AWS Glue creates the quality dataset used for operational reporting and future model evaluation. Lake Formation ensures that only approved roles can use that feedback for model development.
One-page daily operating summary
● Start shift and verify data freshness.
● Open LIVE RISK BOARD and sort by risk.
● Expand high-risk tools and inspect TMG, slope, fleet σ, residual, blind window, trajectory, and reason codes.
● Sort by blind window, fleet deviation, and residual to find hidden exposure.
● Open FDC HEALTH-LINK and select the hardware investigation path.
● Before critical-lot release, verify that no hold, matching, linearity, repeatability, spatial, FDC, or APC concern remains unresolved.
● If evidence is suspicious, run physical confirmation or route to an approved master.
● If feedback is at risk, notify the APC owner and guard the affected window.
● If yield-watch lots are exposed, generate the governed evidence pack.
● End shift by recording credibility, affected lots, containment, owner, due time, and exit criteria.
This SOP turns the portal from a visual demonstration into a production-oriented productivity system: it shortens investigation time, standardizes daily decisions, improves cross-functional communication, preserves the evidence chain, and keeps every manufacturing action under human authority.