← Financial Cloud Cloud Cloud Club · Challenge

Vertex Macro | Financial Cloud Cloud · Challenge

Weekend Productivity Challenge: Fab SPC Drift Synchronization Portal

Series: Challenge

Challenge: 01

Article
Kiro workshop
01 Build with Kiro: Prompt-First Product Design for a Tagalog Learning App
Kiro workshop
02 Build with Kiro: Educational-First Dev Tips for a Tagalog Learning App
Kiro workshop
03 Build with Kiro: Deep-Dive Development Flow for a Tagalog Learning App
Kiro workshop
04 Build with Kiro: Localize a Tagalog Learning App into Chinese Variants Workshop
Kiro workshop
05 Build with Kiro: Grammar and Pronunciation Enrichment Pipeline for Tagalog Cards Workshop
Kiro workshop
06 Build with Kiro: Unique and Reviewable Extra Examples in a Tagalog Learning App Workshop
Kiro workshop
07 Build with Kiro: Factory Engineering Health Hooks Workshop
Kiro workshop
08 Build with Kiro: Etch Process Window Risk Test Automation Workshop
Kiro workshop
09 Build with Kiro: Photolithography Drift Risk Development Workshop
Kiro workshop
10 Engineering Team Get Started — Daily Fab-Duty Use of fab spc drift sync portal
Kiro workshop
11 Engineering Team Addendum — Daily Fab-Duty Use of fab spc drift sync portal
Kiro workshop
12 Kiro: Field Engineering Workshop for Spec-Driven Factory Software
Kiro workshop
13 Kiro: Hands-On Lab — Build a Typed Factory Risk Portal from Scratch
Kiro workshop
14 Kiro: Prompt, Code, and Type Standards Playbook for Engineering Developers
Kiro workshop
15 Kiro: Why a Strong React Prompt Prevents Type Declaration False-Starts
Kiro workshop
17 Build with Kiro: Create a Factory Automation Portal React UI
Kiro workshop
18 Build with Kiro: Create the Automation Analytics Engine Behind a Factory Automation Portal
Kiro workshop
19 Build with Kiro: Add an AI Factory Automation Assistant to a Factory Automation Portal
Kiro workshop
21 Kiro: 2-Hour Professional Developer Workshop Guide
Kiro workshop
22 Kiro: Build the Fab SPC Drift Synchronization Portal from Scratch
Kiro workshop
23 Kiro: Prompt Library and Deep Code Explanation Appendix
Kiro workshop
30 Build with Kiro: Create a Factory Automation Portal UI
Kiro workshop
31 Build with Kiro: Create the Automation Analytics Engine Behind a Factory Automation Portal
Kiro workshop
32 Build with Kiro: Add an AI Factory Automation Assistant to a Factory Automation Portal
Kiro workshop
33 Build with Kiro: Rebuild the CME Direct-Style Quant P&L Leaderboard UI
Kiro workshop
34 Build with Kiro: Recreate the Quant Analytics Engine Behind the P&L Board
Kiro workshop
35 Build with Kiro: AWS AI-Powered Trading Desk Assistant for the Quant Board
Kiro workshop
36 One-Page Trading Portal SOP
Kiro workshop
AgentCore
A1 Build with AgentCore & Strands: Gateway MCP Tool Fabric Developer Workshop
AgentCore
A2 Build with AgentCore & Strands: Governed Multi-Agent Risk System Developer Workshop
AgentCore
A3 Build with AgentCore & Strands: Runtime Sovereign Risk Agent Developer Workshop
AgentCore
Exam practice
E1 Build a Multilingual AWS Exam Practice Launch System with Vibe Coding
Exam practice
E2 Build an AWS Exam Practice Room with Vibe Coding Dev Tips
Exam practice
E3 Build the Practice Engine Behind a Static AWS Exam Room
Exam practice
Amazon Q
Q1 Amazon Q: CloudShell-First Developer Workshop for ACM Certificate Auto Renewal
Amazon Q
Tagalog Practice Room
T1 Build a Tagalog Learning App for AWS Manila Community Day with Prompt-First Product Design
Tagalog Practice Room
T2 Build Tagalog Learning App for AWS Manila Community Day with Educational-First Dev Tips
Tagalog Practice Room
T3 Deep Dive Development Flow for a Tagalog Learning App for AWS Manila Community Day
Tagalog Practice Room
T4 Build Localize a Tagalog Learning App into Chinese Variants for AWS Manila Community Day
Tagalog Practice Room
T5 Build a Grammar and Pronunciation Enrichment Pipeline for Tagalog Cards for AWS Manila Community Day
Tagalog Practice Room
T6 Make Extra Examples Unique and Reviewable in a Tagalog Learning App for AWS Manila Community Day
Tagalog Practice Room
Roadmap
R1 Enterprise Data Analytics Roadmap: 100 Deep Scenario Questions
Roadmap
R2 Front-End Development Roadmap: Real-World Enterprise Scenarios
Roadmap
Hong Kong Community Day
C1 A Hong Kong Weekend with AWS Community Day: From Cloud Sessions to Harbour Lights
Hong Kong Community Day
C2 The Speaker’s Luxury Weekend: Present an AWS Story, Then Let Hong Kong Take the Stage
Hong Kong Community Day
C3 Seventy-Two Hours in Hong Kong: The Grand Tour for an AWS Community Day Speaker
Hong Kong Community Day
Manila Community Day
C4 AWS Community Day Manila: A Joyful Weekend of Cloud, Culture, and True Friendship
Manila Community Day
C5 AWS Community Day Manila: Where Cloud Builders Find the Happiest Spirit of the Philippines
Manila Community Day
C6 AWS Community Day Manila: Build, Break, Repeat, and Belong in a City of Joy
Manila Community Day
C7 First-Time Visitor Tips for Manila, Philippines
Manila Community Day
Philippines × Hong Kong
C8 Philippines Hong Kong Capital Market Upgrade
Philippines × Hong Kong
Backtest
B1 Build Institutional Amazon Long-Only Backtesting Agents With Bedrock AgentCore And Strands Agents
Long-only AMZN agents with AgentCore, Strands, and a governed Backtrader ledger.
B2 Build Regime-Aware Amazon Position Management With Backtrader, AgentCore, And Strands Agents
Treat market regime as a position control, not a chart comment.
B3 Build Benchmark-Relative Amazon Timing Systems Using Nasdaq, S&P 500, Dow, AgentCore, And Strands
Time AMZN against Nasdaq, S&P 500, and Dow context.
B4 Build A Governed Amazon Trade-History Factory With Bedrock AgentCore, Strands Agents, And Backtrader
Turn backtests into an auditable trade-history factory.
B5 Build An Agentic Amazon Backtest Operating Model With Bedrock AgentCore And Strands Agents [Part 1]
Build the operating model before debating the result.
B6 Build A Custom Cerebro Code Talk For Amazon Timing And Position Management [Part 2]
Explain the Cerebro engine before explaining the chart.
B7 Build Trader Review Records For Amazon Strategy Results And Lessons Learned [Part 3]
Turn strategy ranks into trader review records.
B8 Build A Governed FSI Amazon Position Management Playbook With AgentCore And Strands [Part 4]
An FSI playbook for governed Amazon position management.
B9 Build a Sovereign Risk Trading Agent with Amazon Bedrock AgentCore for Yield Spreads, FX Hedging, and Debt Repricing
Sovereign-risk agent for yield spreads, FX hedges, and debt repricing.
B11 Build Modern Volatility Trading & Lawful Thailand Recovery Planning Agents: A Memory-Driven Strands Multi-Agent Risk Protection System
Memory-driven Strands agents for volatility and Thailand recovery.
B12 Build Short Straddle Trading-Risk Governance with Amazon Bedrock AgentCore Memory
Short-straddle risk governance with AgentCore Memory.
B13 Building Production-Ready Credit & Yield Staking AI Agents on Amazon EKS
Production credit and yield-staking agents on Amazon EKS.
Challenge
01 Weekend Productivity Challenge: Fab SPC Drift Synchronization Portal
Fab SPC drift review and recommendation portal.
02 Weekend Productivity Challenge: Quant P&L Commander — An AI-Powered Trading Productivity Portal on AWS
Quant P&L leaderboard and trading productivity portal.
03 Weekend Annoying Task Challenge: Trading Desk Execute Summary On Cloud, On Chain, On Air
DeskPulse daily execution communication.
04 Weekend Agent Challenge: The 6 AM Trading Risk Review
An unattended, evidence-backed morning credit and trading risk brief.
05 Weekend Creative Challenge: Leadership Card Game
A browser-based creative facilitation deck.
06 Full Stack Challenge: Community Day Board App
A browser-based event communication room.
Leadership Card Game
01 Leadership Card Game: Last Skill Cloud Did Not Automate
A field essay for Builders on language, courage, and the Leadership Card Game
02 Anatomy of a Leadership Round: How the Leadership Card Game Actually Plays
A facilitator’s field guide for Builders who want drills that fit inside real meetings
03 Leadership Card Game: When the Opportunity Stops Belonging to the Organizer
A field essay for Builders on power transfer, multilingual practice nights, and career arcs that complete Entrance, Resource, and Narrative
04 Weekend Creative Challenge: Leadership Card Game
Master high-stakes workplace conversations before they happen.
05 From a Weekend Challenge Project to $1,386 Crowdfunding: The Leadership Practice That Changes How You Show Up at Work
A weekend build became a live 600-card leadership practice room and reached $1,386 in crowdfunding.
06 From a Weekend Challenge Project to $1,386 Crowdfunding: A Day 1 Path Into the Tech Industry
How did a weekend challenge become a multilingual AWS-powered product with 600 cards and $1,386 in crowdfunding?
07 From a Weekend Challenge Project to $1,386 Crowdfunding: Build a Professional Brand by Transferring Opportunity
A weekend challenge reached $1,386 in crowdfunding by turning leadership ideas into a working multilingual product.
08 Leadership Card Game — Crowdfunding Campaign
Speak leadership before the room decides your career.
09 PR/FAQ 01 — Leadership Card Game launches for community builders
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Community managers, volunteer organizers, early-career…
10 PR/FAQ 02 — Enterprise facilitators adopt Leadership Card Game for live leadership drills
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Learning & development leads, people managers, agile…
10 PR/FAQ 03 — Multilingual Leadership Card Game opens global practice rooms for builder ownership
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Global AWS builders, bilingual communities, cross-border…
AWS Builder Center
01 AWS Builder Center, its community spirit, and AWS Builder Jacket
There are destinations you reach by plane, destinations you enter through a door, and destinations that begin with a sign-in screen and quickly feel like a…
02 Inside AWS Builder Center, where a global technical platform becomes a place to learn, contribute, and belong
A great journey does not always begin at an airport.
03 AWS Community Builder huge success
When builders share openly, the entire community moves forward.
04 AWS Builder Center huge success
A vibrant global district built for curiosity, public learning, and the AWS Builder Jacket.
05 A weekend inside AWS Builder Center, from community inspiration to unmistakable AWS Builder Jacket
Friday evening begins with a familiar builder feeling: there is an idea waiting somewhere between a problem and a possibility.

Live application: Fab SPC Drift Synchronization Portal
Public source: GitHub — src/main.tsx

An AI-powered productivity application for semiconductor metrology engineers, developed with Kiro and implemented on AWS using Amazon S3, Amazon CloudFront, Amazon Route 53, AWS Certificate Manager, AWS WAF, Amazon MSK, AWS Lambda, AWS Glue, and AWS Lake Formation.


The Fab SPC Drift Synchronization Portal is an AWS-based system that helps semiconductor engineers identify equipment drift, review supporting evidence, and prioritize the safest next action. Manufacturing systems—including CD-SEM, SPC, FDC, APC, lot-route, and yield-watch systems—send operational events to Amazon MSK, which provides the real-time streaming backbone.

AWS Lambda processes these events by: Validating and normalizing incoming data Checking timestamps and data quality Calculating equipment-health and drift features Running AI-based risk inference Applying deterministic engineering rules Generating prioritized recommendations

The processed data, AI results, and audit evidence are stored in an Amazon S3 data lake. AWS Glue catalogs the datasets and manages schemas, while AWS Lake Formation controls access to sensitive manufacturing information. The engineer accesses the React portal through Amazon Route 53 and Amazon CloudFront. The React application is stored in a private Amazon S3 bucket, while AWS WAF protects the public entry point from unwanted traffic. The portal displays risk scores, reason codes, evidence, and recommended actions such as Release, Watch, Run SPC, Golden Wafer, Route Limit, APC Guard, or Hold Review. The system is advisory and human-in-the-loop: it helps engineers make faster, evidence-based decisions, but it cannot directly control equipment, hold lots, change routes, or modify APC settings. [codpayment...epoint.com]

Simple Data Flow

Manufacturing Systems
        ↓
Amazon MSK
        ↓
AWS Lambda Processing and AI
        ↓
Amazon S3 Data Lake
        ↓
AWS Glue + AWS Lake Formation
        ↓
Lambda API
        ↓
React Portal on S3 and CloudFront
        ↓
Engineer Reviews Recommendation

Vision & What the App Does

Statistical process control (SPC) is essential in semiconductor manufacturing, but a scheduled SPC result is only a snapshot. A CD-SEM can pass its last control check and then begin drifting while production lots continue to move. Gun-vacuum degradation, emission-current movement, deflector instability, stage vibration, contamination, matching changes, or simple measurement aging can emerge between scheduled checks. During this interval, an engineer may need to open several systems, compare timestamps manually, calculate fleet deviations, trace product routes, review equipment symptoms, and decide whether to release, watch, remeasure, reroute, or hold.

That investigation is both a manufacturing risk and a productivity problem.

The Fab SPC Drift Synchronization Portal is a personal AI-powered productivity workspace that compresses that investigation into one prioritized and explainable review experience. It synchronizes production-aware SPC indicators with fault detection and classification (FDC) telemetry, metrology results, tool-matching evidence, fleet behavior, and yield-watch context. Instead of requiring an engineer to reconstruct the complete situation manually, the portal provides a ranked risk board and makes the evidence behind each recommendation immediately visible.

The portal answers four practical questions:

● Which tool requires attention first?

● What changed during the SPC blind window?

● Which evidence supports the risk assessment?

● What is the safest next engineering action?

The current React experience contains six operational areas:

● Overview explains the control loop and displays fleet-level productivity indicators.

● Live Risk Board supports searching, sorting, ranking, and expanding individual CD-SEM records.

● FDC Health-Link maps equipment symptoms to metrology impact and potential yield exposure.

● Dynamic Matching compares each tool with fleet behavior and stable reference tools.

● Yield Triage assembles the evidence required to distinguish process movement from metrology error.

● Runbook converts the solution into a repeatable engineering workflow from proof of concept to production operation.

For each tool, the portal evaluates blind-window age, tool-matching gap, Mandel slope, fleet deviation, residual noise, trend behavior, layer criticality, and correlated FDC signals. It then produces a controlled recommendation:

● RELEASE

● WATCH

● RUN SPC

● GOLDEN WAFER

● ROUTE LIMIT

● APC GUARD

● HOLD REVIEW

The application is deliberately advisory. AI can detect patterns, rank work, summarize evidence, and recommend a next action, but it cannot command manufacturing equipment or bypass an engineer's approval. This human-in-the-loop boundary is part of the product design, not an afterthought.

From a productivity perspective, the portal works like a specialized task prioritizer, investigation assistant, and shift-handoff tool. It reduces context switching, makes each recommendation reproducible, and helps an engineer spend time on the highest-risk tool rather than manually searching for the next problem.


How I Built It

Kiro Development Workflow

I used Kiro as the AI-assisted development environment for requirements analysis, architecture design, implementation planning, code generation, refactoring, testing, and production hardening.

Rather than beginning with isolated UI code, I described the engineering outcome to Kiro:

Reduce the time required to identify CD-SEM drift between scheduled SPC checks, while preserving human authority over every manufacturing action.

I converted that outcome into a Kiro specification with three coordinated artifacts:

.kiro/specs/fab-spc-drift-portal/
├── requirements.md
├── design.md
└── tasks.md

Requirements created with Kiro

The requirements were written as testable behaviors:

● When normalized tool data arrives, the application must refresh the affected tool's risk state.

● When a user expands a tool, the application must display the value, engineering limit, event time, data-quality state, and reason code.

● When required data is stale, missing, incompatible, or outside the approved model domain, the application must return REVIEW_REQUIRED rather than infer a safe condition.

● Every recommendation must be reproducible from a versioned rule set, model artifact, feature set, and source-event range.

● No application component may initiate an equipment command, lot hold, route change, or APC change without an independently authenticated human approval process.

● The user interface must remain useful when AI explanation is unavailable by displaying deterministic evidence and a safe fallback message.

Design decisions captured with Kiro

Kiro helped separate the solution into five boundaries:

● React delivery and browser security;

● streaming ingestion with Amazon MSK;

● validation, normalization, feature calculation, and inference with AWS Lambda;

● immutable evidence storage in Amazon S3; and

● governed discovery and access with AWS Glue and AWS Lake Formation.

This separation prevents the front end from becoming a system of record. The browser displays results, but it does not possess Kafka credentials, access raw manufacturing data, calculate authoritative disposition, or invoke equipment interfaces.

Kiro steering and automated checks

Project steering files established persistent engineering rules:

.kiro/steering/
├── architecture.md
├── aws-security.md
├── react-typescript.md
├── streaming-contracts.md
├── ai-safety.md
└── testing-standards.md

The steering guidance required:

● TypeScript strictness and explicit domain types;

● accessible labels and keyboard-operable controls;

● private Amazon S3 origins rather than public website buckets;

● least-privilege AWS Identity and Access Management policies;

● schema versioning for every Amazon MSK event;

● idempotent AWS Lambda processing;

● no secrets or AWS credentials in React code;

● deterministic fallbacks for AI failure;

● model and rule version recording; and

● explicit human review for manufacturing-impacting recommendations.

Kiro hooks automated formatting, linting, unit tests, schema compatibility checks, dependency scanning, infrastructure validation, and a safety test that rejects any code path capable of issuing an autonomous equipment command.

React and TypeScript Front End

The user interface is implemented as a React single-page application. The supplied main.tsx defines typed objects for tools, actions, FDC mappings, runbook steps, tabs, sort keys, and risk tones. This is important because the interface represents controlled engineering states rather than arbitrary strings.

The implementation is componentized around operational responsibilities:

● Header and Hero establish current system context;

● KPIStrip summarizes the engineer's workload;

● LiveRiskBoard filters and ranks tools;

● Sparkline, LargeLineChart, and FleetBars visualize evidence;

● FDCHealth explains equipment-to-metrology relationships;

● DynamicMatching shows fleet comparisons;

● YieldTriage structures the investigation process;

● Runbook explains controlled rollout; and

● Sop embeds the expected user procedure.

The demonstration UI updates local sample values to show interactive behavior. In the production design, those fixture updates are replaced by authenticated API responses generated from Lambda-managed state. The presentation contract remains stable: the UI receives a sanitized tool summary, timestamp, risk score, confidence, reason codes, limits, recommendation, and evidence references.

The most important implementation decision was to separate three concepts:

● Observed facts — validated source values and timestamps;

● Analytical decisions — calculated features, engineering rules, and AI risk scores; and

● Presentation — React components that explain the result to a human.

A presentation bug cannot therefore change the authoritative recommendation, and an AI-generated explanation cannot overwrite a measured value.


AWS Services Used / Architecture Overview

The production architecture intentionally uses the following service scope:

● Kiro development

● Amazon S3, Amazon CloudFront, Amazon Route 53, AWS Certificate Manager, and AWS WAF for React delivery

● Amazon MSK as the streaming backbone

● AWS Lambda for data massage, feature engineering, AI inference, APIs, and event processing

● Amazon S3 data lake, AWS Glue, and AWS Lake Formation for governed evidence

Architecture at a glance

ENGINEER
   |
   v
Amazon Route 53
   |
   v
Amazon CloudFront ---- AWS WAF
   |                         |
   |                         +-- managed rules, rate control, request filtering
   v
Private Amazon S3 bucket
React / TypeScript build artifacts

FAB EVENT PRODUCERS
CD-SEM | FDC | SPC | lot route | APC | yield-watch
   |
   v
Amazon MSK topics
   |
   v
AWS Lambda normalization and validation
   |
   +--> Amazon S3 raw evidence zone
   |
   v
AWS Lambda feature engineering and AI inference
   |
   +--> Amazon MSK recommendation topic
   +--> Amazon S3 curated, feature, model-output, and audit zones
   +--> React query API implemented with AWS Lambda

Amazon S3 data lake
   |
   v
AWS Glue Data Catalog and schema management
   |
   v
AWS Lake Formation governed access

It is a deployment-oriented view showing the web-delivery boundary, streaming boundary, processing boundary, and governance boundary.


Amazon S3, CloudFront, Route 53, ACM, and WAF: React Delivery

The production React build is generated by the CI/CD process and deployed to a private Amazon S3 bucket. The bucket is not configured as a publicly accessible S3 website. Amazon S3 Block Public Access remains enabled, and Amazon CloudFront uses Origin Access Control to retrieve objects from the bucket.

The deployment package contains content-hashed assets:

dist/
├── index.html
├── assets/index.<content-hash>.js
├── assets/index.<content-hash>.css
└── approved static assets

Content hashing allows JavaScript and CSS to use a long immutable cache lifetime. index.html receives a short cache lifetime because it points to the current asset versions. Deployment uploads immutable assets first and index.html last, preventing the HTML document from referencing files that are not yet available.

Amazon CloudFront configuration

Amazon CloudFront provides the global HTTPS entry point and performs:

● edge caching of React assets;

● compression;

● TLS termination;

● origin protection;

● security response headers;

● controlled invalidation during exceptional rollback; and

● a consistent application domain.

The response-headers policy includes:

Strict-Transport-Security
Content-Security-Policy
X-Content-Type-Options
Referrer-Policy
Permissions-Policy
frame-ancestors restriction

The content security policy limits scripts, styles, images, connections, and framing to explicitly approved sources. Source maps are not exposed in the general production distribution; they are retained separately for controlled debugging.

Route 53 and AWS Certificate Manager

Amazon Route 53 hosts the DNS record for the application and aliases the application hostname to the CloudFront distribution. AWS Certificate Manager supplies and renews the TLS certificate used by CloudFront. DNS validation keeps certificate renewal automated and auditable.

AWS WAF

AWS WAF is associated with CloudFront. Its web ACL includes:

● AWS managed common protections;

● known-bad-input protections;

● IP reputation protections;

● rate-based rules;

● size constraints for unexpected requests; and

● count-mode validation before a new blocking rule is promoted.

The React application is static, but WAF remains valuable because the same CloudFront entry point can route controlled API paths to Lambda-backed endpoints. WAF metrics are monitored before thresholds are tightened so legitimate engineers are not blocked by an untested rule.

S3 production controls

The application bucket uses:

● S3 Block Public Access;

● object versioning;

● default encryption;

● bucket-owner-enforced object ownership;

● access logging to a separate protected log bucket;

● lifecycle rules for superseded build artifacts; and

● a bucket policy that grants reads only through the designated CloudFront distribution.

A release manifest records the source commit, Kiro specification revision, build checksum, asset list, and deployment time. Rollback changes index.html to reference the previous immutable asset set.


Amazon MSK: The Streaming Backbone

Amazon Managed Streaming for Apache Kafka is the durable event backbone. It decouples high-frequency manufacturing producers from normalization, AI processing, evidence storage, and web presentation.

The production topic structure is domain based:

fab.fdc.telemetry.v1
fab.metrology.measurement.v1
fab.spc.result.v1
fab.lot.route.v1
fab.apc.feedback.v1
fab.yield.watch.v1
fab.tool.recommendation.v1
fab.processing.deadletter.v1

Event contract

Every event contains a common envelope:

{
  "event_id": "source-generated-or-content-derived-id",
  "event_type": "fdc.telemetry",
  "schema_version": "1.2.0",
  "event_time": "2026-07-12T10:15:32.481Z",
  "ingest_time": "2026-07-12T10:15:33.014Z",
  "fab_id": "tokenized-fab-reference",
  "tool_id": "tokenized-tool-reference",
  "lot_id": "tokenized-lot-reference",
  "recipe_id": "approved-recipe-reference",
  "sequence": 1845531,
  "quality": {
    "complete": true,
    "clock_status": "synchronized"
  },
  "payload": {}
}

event_time represents when the physical or manufacturing event occurred. ingest_time represents when the platform received it. This distinction is essential: risk windows are based on event time, while platform latency is measured from ingest time.

Partitioning and ordering

Partition keys are selected by business ordering requirement:

● tool telemetry is partitioned by tool_id;

● lot lifecycle events are partitioned by lot_id;

● recommendations are partitioned by tool_id; and

● cross-fleet analytics uses curated feature events rather than assuming global Kafka order.

Ordering is guaranteed only within a partition. Any algorithm requiring several topics uses event-time windows, watermarks, and late-event handling instead of assuming that arrival order is correct.

Schema governance

AWS Glue Schema Registry manages compatible Avro or JSON schemas. The compatibility policy permits additive changes with defaults and rejects incompatible field reinterpretation. A breaking contract creates a new topic version. Lambda logs the writer schema, reader schema, and compatibility result so replay remains explainable.

Reliability and capacity

Partition count is calculated from peak event rate, average record size, target consumer parallelism, retention, and expected growth. The design monitors:

● producer error rate;

● broker storage usage;

● bytes in and out;

● under-replicated partitions;

● consumer lag;

● oldest unprocessed event age; and

● dead-letter volume.

Producers use TLS, authenticated access, acknowledgements appropriate to durability, retries with jitter, compression, bounded local buffering, and an explicit failure path. Multi-Availability Zone replication protects the event stream from a single-zone failure. Recovery testing includes restarting consumers, replaying offsets, and proving that recommendations are not duplicated.


AWS Lambda: Data Massage, AI, and Event Processing

AWS Lambda performs the serverless processing work between Amazon MSK and the governed S3 data lake. The design separates responsibilities into small functions rather than creating one untestable handler.

msk-normalizer
feature-builder
risk-inference
recommendation-publisher
evidence-writer
portal-query-api
feedback-processor

MSK event processing

An event source mapping polls the approved MSK topics and invokes the normalizer with bounded batches. Function settings are tuned for the event size and latency objective rather than left at arbitrary defaults. Reserved concurrency protects downstream storage, and the failure policy prevents a poison record from blocking a partition indefinitely.

For each Kafka record, the normalizer performs:

● payload decoding;

● writer-schema identification;

● schema validation;

● event-envelope validation;

● unit standardization;

● timestamp and clock-quality checks;

● tokenized identity enrichment;

● duplicate detection;

● data-quality classification;

● raw evidence persistence; and

● normalized event publication.

Malformed records are written to the dead-letter topic with the original topic, partition, offset, schema version, error code, and encrypted payload reference. They are never silently discarded.

Idempotency

Kafka consumers are designed for at-least-once delivery. Therefore, every side effect is idempotent. The processing key is derived from the source identity, partition, offset, schema version, and canonical event ID. S3 object keys and recommendation IDs incorporate this identity. A retry produces the same result instead of a second case.

The calculation output stores the exact contributing topic/partition/offset ranges. This allows a recommendation to be reproduced from source evidence after a replay or audit.

Feature engineering

The feature Lambda calculates point-in-time-correct values such as:

● elapsed time since validated SPC;

● volume of production exposure during the blind window;

● rolling mean and standard deviation;

● exponentially weighted movement;

● trend slope and step-change indicators;

● residual three-sigma estimate;

● site-to-site range;

● tool-matching gap;

● Mandel slope distance from one;

● robust fleet deviation;

● distance from approved master anchors;

● FDC baseline distance;

● missingness and lateness indicators; and

● layer-criticality and route-exposure features.

Features are calculated only from information available at the decision timestamp. This prevents future-data leakage during both production inference and offline model evaluation.

AI inference in Lambda

To keep the implementation inside the defined service scope, the approved AI model is packaged as a versioned Lambda layer or container-image dependency. The trained artifact and metadata are stored in a governed S3 model prefix. At cold start, the function loads the approved model version; warm invocations reuse it.

The AI output is structured rather than free-form:

{
  "risk_probability": 0.91,
  "severity": "HIGH",
  "confidence": 0.84,
  "model_version": "spc-risk-2026-07-12.3",
  "reason_codes": [
    "FLEET_DEVIATION_HIGH",
    "MANDEL_SLOPE_OUTSIDE_RANGE",
    "DEFLECTOR_SIGNAL_UNSTABLE"
  ],
  "domain_status": "IN_DOMAIN"
}

A deterministic policy evaluator combines model output with approved engineering rules. AI may raise an early warning, but it cannot weaken a hard rule or issue an autonomous manufacturing action.

Safe decision policy

IF required data is missing, stale, or incompatible
    RECOMMEND REVIEW_REQUIRED
ELSE IF a hard engineering limit is breached
     AND an independent signal corroborates the breach
    RECOMMEND HOLD_REVIEW
ELSE IF the model predicts a near-term breach
     AND confidence and domain checks pass
    RECOMMEND RUN_SPC or GOLDEN_WAFER
ELSE IF fleet or FDC warning criteria are met
    RECOMMEND WATCH
ELSE
    RECOMMEND RELEASE

Observability

Every Lambda invocation emits structured logs and operational metrics:

RecordsProcessed
ValidationFailureCount
DuplicateCount
LateEventCount
FeatureFreshnessSeconds
InferenceDurationMs
OutOfDomainCount
RecommendationCountByType
DeadLetterCount
EndToEndDecisionLatencyMs

Logs include correlation ID, model version, rule-set version, source offsets, and error classification, but exclude unnecessary sensitive manufacturing values. Alarms focus on user impact: growing MSK lag, stale recommendations, validation spikes, inference failures, and dead-letter growth.


S3 Data Lake, AWS Glue, and Lake Formation: Governed Evidence

A production AI application requires more than a model. It requires reproducible evidence, controlled training data, discoverable definitions, lineage, and access governance.

The Amazon S3 data lake is organized into zones:

s3://fab-spc-data/raw/
s3://fab-spc-data/validated/
s3://fab-spc-data/curated/
s3://fab-spc-data/features/
s3://fab-spc-data/models/
s3://fab-spc-data/predictions/
s3://fab-spc-data/evidence/
s3://fab-spc-data/feedback/
s3://fab-spc-data/quarantine/

Raw zone

The raw zone preserves source events with minimal transformation. Objects are immutable, encrypted, versioned, and partitioned by source domain and event date. The original Kafka metadata is retained for replay and audit.

Validated and curated zones

The validated zone contains schema-conforming events with quality flags. The curated zone contains standardized engineering entities such as tool state, metrology observations, matching results, route exposure, and approved outcomes. Columnar files and practical partitioning reduce scan cost and improve analytic performance.

Feature and model zones

The feature zone stores point-in-time-correct training and inference features. Each dataset records:

● feature definition version;

● source event range;

● generation code version;

● cutoff timestamp;

● quality results; and

● approved use.

The model zone stores the model artifact, preprocessing definition, feature order, evaluation report, threshold configuration, model card, approval state, checksum, and rollback predecessor. Lambda is allowed to load only an explicitly approved model prefix.

AWS Glue

AWS Glue provides the technical catalog and schema layer. Glue crawlers are used selectively; production tables with strict contracts are managed from explicit definitions so an unexpected file cannot silently redefine a critical column.

The Data Catalog records:

● database and table definitions;

● file formats and partitions;

● schema versions;

● owners and descriptions;

● data classification;

● quality status; and

● source-to-curated lineage references.

AWS Glue jobs support larger offline transformations when processing exceeds the appropriate Lambda execution profile. Glue-generated training datasets use the same feature definitions and event-time rules as the Lambda inference path, reducing training-serving skew.

AWS Lake Formation

AWS Lake Formation governs access to cataloged S3 data. Permissions are granted by role and data purpose rather than by broad bucket access. LF-tags classify domains such as fab, tool family, product sensitivity, evidence type, and approved use.

Example access boundaries include:

● front-end delivery roles cannot access the data lake;

● Lambda normalization can write raw and validated data but cannot approve models;

● the inference Lambda can read only approved feature definitions and model artifacts;

● engineering analysts can query authorized curated data;

● model-development roles can read approved training data but not unrestricted raw identifiers; and

● auditors can read evidence and lineage without modifying operational datasets.

Column- and row-level controls prevent unnecessary exposure. Cross-account sharing uses governed catalog permissions rather than copying uncontrolled datasets.

Retention and evidence integrity

Lifecycle rules move historical evidence to lower-cost S3 storage classes according to business and regulatory requirements. Quarantine retention is deliberately limited. Model, recommendation, and approval evidence remains available for the required audit period.

Checksums, object versioning, protected access logs, and separation of duties make the evidence chain tamper-evident. A recommendation can be traced from the UI to the Lambda output, model version, feature record, curated dataset, and original MSK offsets.


Complete AI Development Lifecycle

AI development covers much more than training an algorithm. For this application, the complete lifecycle is implemented around the AWS service scope.

1. Problem definition

The model predicts whether a tool is likely to produce unacceptable metrology behavior before the next scheduled SPC opportunity. The target is not “predict every anomaly.” The target is to provide enough lead time for a useful human review while controlling unnecessary investigations.

2. Label definition

Labels are derived from approved engineering outcomes:

● confirmed SPC out-of-control events;

● golden-wafer results;

● confirmed equipment findings;

● matching or linearity failures;

● approved route or hold decisions;

● false-alarm dispositions; and

● yield backtrace conclusions.

An engineer's initial recommendation is not automatically treated as truth. Final approved disposition and investigation outcome are stored separately to prevent self-reinforcing labels.

3. Data preparation

Amazon MSK captures operational events, Lambda validates and aligns them, S3 stores immutable history, AWS Glue creates reproducible datasets, and Lake Formation controls who may use each dataset. Training examples use event-time cutoffs so they cannot include information that became available after the prediction point.

4. Dataset splitting

Random row splitting would leak repeated tool behavior across train and test data. The production process therefore applies chronological splits and groups by tool or tool family where appropriate. A final untouched time period measures realistic forward performance.

5. Feature development

Features are documented with purpose, unit, source, expected range, missing-value behavior, owner, and leakage risk. The same versioned transformations are used for offline datasets and Lambda inference.

6. Model selection

The first production candidate favors a compact, interpretable model that can execute efficiently in Lambda. Candidate models are compared against deterministic baselines. A more complex model is accepted only when it improves operational metrics without unacceptable latency, instability, or loss of explainability.

7. Evaluation

Because true drift events are uncommon, accuracy is not the primary metric. Evaluation includes:

● precision and recall;

● precision-recall area;

● false-negative cost;

● recall at available engineer review capacity;

● probability calibration;

● average warning lead time;

● performance by tool family, layer, recipe, and event-quality state;

● out-of-domain detection; and

● comparison with the existing deterministic process.

Threshold selection is a manufacturing decision shared with engineering owners. It balances missed-drift risk against review workload.

8. Explainability

The model returns stable reason codes and contributing feature values. The UI presents these beside the approved limits and source time. The explanation is constrained to observed evidence; it does not invent a root cause.

9. Bias and coverage assessment

The evaluation checks whether the model performs consistently across tool families, layers, recipes, maintenance states, and data-quality conditions. In this industrial context, unfairness can appear as systematically poorer detection for a less-common tool or process family. Unsupported groups are marked out of domain and routed to human review.

10. Model approval and versioning

Each approved model package includes:

model artifact
preprocessor artifact
feature specification
training-data snapshot reference
evaluation report
slice metrics
threshold configuration
model card
known limitations
approval record
rollback model
artifact checksums

Only an approved S3 model prefix can be loaded by the inference Lambda. Model promotion is separated from model development through Lake Formation permissions and deployment controls.

11. Deployment

A candidate begins in shadow mode. It receives production features but cannot influence the displayed recommendation. Predictions are compared with the current rule set and later outcomes. After approval, a small controlled portion of requests uses the candidate while the previous model remains immediately available for rollback.

12. Monitoring

Production monitoring covers four categories:

● service health: Lambda errors, duration, throttles, cold starts, MSK lag;

● data health: missing fields, schema changes, late events, range violations;

● model health: feature drift, score distribution, out-of-domain rate, calibration;

● business health: warning lead time, accepted recommendations, false alarms, prevented exposure, and engineer review time.

When labels arrive later, the feedback processor joins predictions with approved outcomes and calculates delayed quality metrics.

13. Retraining

Retraining is triggered by an approved schedule, sufficient new labels, feature drift, performance degradation, or a meaningful equipment/process change. Retraining never automatically promotes a model. The full evaluation and approval gate runs again.

14. Responsible AI and safety

The application applies the following controls:

● AI is advisory only.

● Every recommendation displays uncertainty and source freshness.

● Missing or stale evidence produces review, not release.

● Hard engineering limits cannot be overridden by the model.

● Sensitive identifiers are minimized and tokenized.

● Training and inference access is governed through Lake Formation and IAM.

● Model artifacts, datasets, rules, and outputs are versioned.

● Engineers can accept, reject, or correct recommendations.

● AI failure falls back to deterministic rules and evidence.

● Autonomous equipment commands are outside the application's IAM permissions and network path.


Production Testing and Release Strategy

Kiro-generated tasks include tests for the UI, streams, Lambda processing, data governance, and AI behavior.

React tests

● component rendering;

● search and sort behavior;

● keyboard navigation;

● accessible names and contrast;

● stale-data and AI-unavailable states;

● threshold-boundary display; and

● mobile layout behavior.

MSK and Lambda tests

● compatible and incompatible schemas;

● duplicate delivery;

● reordered events;

● late arrivals;

● partial batch failure;

● poison messages;

● replay from earlier offsets;

● large batches;

● downstream timeout; and

● proof that retry does not duplicate a recommendation.

Data lake and governance tests

● S3 public-access denial;

● encryption and versioning;

● partition and schema validation;

● Glue catalog consistency;

● Lake Formation positive and negative authorization tests;

● lineage completeness; and

● retention-policy validation.

AI tests

● future-data leakage checks;

● feature parity between training and inference;

● model serialization and cold-start loading;

● boundary and missing-value behavior;

● probability calibration;

● slice evaluation;

● out-of-domain handling;

● reason-code stability;

● safe fallback; and

● model rollback.

The release pipeline promotes immutable artifacts through development, staging, and production. A production release records the Git commit, Kiro spec revision, React manifest, infrastructure revision, schema versions, Lambda versions, rule-set version, and approved model version.


What I Learned

The first lesson was that an industrial productivity application should reduce the number of decisions an engineer must reconstruct, not simply add another dashboard. The useful output is a prioritized case containing fresh evidence, explicit limits, uncertainty, and the smallest safe next action.

The second lesson was that streaming correctness is operational correctness. A correct formula can still produce the wrong decision if it uses arrival time instead of event time, assumes global ordering, duplicates a side effect after retry, or evaluates stale evidence. Amazon MSK and idempotent Lambda processing made replay, traceability, and failure recovery part of the design.

The third lesson was that AI development starts with data contracts and ends with monitored human outcomes. The model itself is only one artifact. S3 evidence zones, Glue metadata, Lake Formation controls, point-in-time feature construction, approval records, reason codes, feedback, drift monitoring, and rollback are equally important.

The fourth lesson came from Kiro. AI-assisted coding is most effective when requirements, architecture, security rules, and tests constrain generation. Kiro specs maintained traceability from the productivity problem to implementation tasks. Steering preserved AWS and safety conventions. Hooks placed repetitive quality checks directly inside the engineering workflow.

Finally, I learned that a production-ready AI system must be designed for uncertainty. The portal never treats missing data as healthy data, never treats a probability as an equipment command, and never hides the distinction between observed evidence, deterministic rules, and AI prediction.

The result is an AWS-focused productivity tool that helps an engineer understand what needs attention, why it matters, which evidence supports the decision, and what to review next.


Link to App or Repo

● Working deployed application: https://vertexmacro.com/cloud_club/demo/dashboard/aws_factory_automation_portal.html

● Public GitHub source: https://github.com/dchan-dev/aws-fab-spc-drift-synchronization-portal-ts/blob/main/src/main.tsx


Application SOP — Daily Fab-Duty Operation and Business Value

The portal delivers value only when its risk indicators lead to a consistent engineering routine. This standard operating procedure explains how metrology engineers, equipment engineers, process engineers, process-integration engineers, and yield engineers use the application during a normal shift. It also makes the productivity benefit visible to a challenge reviewer: the application does not merely display charts; it converts fragmented evidence into a prioritized, repeatable work process.

Safety boundary: The portal is an advisory decision-support application. Limits shown in the public demonstration are learning baselines. Production limits must be approved for the applicable node, product, layer, customer, and module. Final release, route-limit, APC, maintenance, and hold decisions remain with authorized fab personnel.

Intended audience and value by role

Metrology engineer or CD-SEM owner

The metrology owner uses the portal to detect drift between scheduled SPC checks, evaluate tool matching, compare a suspect tool with the fleet, and determine whether physical confirmation is required. The productivity value is a single evidence view for TMG, Mandel slope, residual noise, blind-window age, fleet deviation, and recent movement instead of a manual investigation across unrelated systems.

Equipment engineer

The equipment engineer works from the hardware-health perspective: Is the CD-SEM healthy enough to continue measuring production lots, which FDC signal is driving the loss of measurement credibility, and what containment is required? The FDC Health-Link panel directs the investigation toward gun vacuum, emission current, deflector DAC, or stage vibration and connects that path to an observable metrology effect.

Process or process-integration engineer

PE and PIE use the portal before changing lithography or etch settings. If the measurement itself is questionable, a process correction can move the process in the wrong direction. The portal provides the route history, metrology credibility indicators, APC-guard status, and engineering recommendation required to determine whether process action should wait for verification.

Yield engineer

The yield engineer uses the portal to backtrace a yield-watch lot to the measuring CD-SEM, the FDC state, the last trusted SPC point, the fleet comparison, and the APC exposure window. This accelerates the distinction among a real process excursion, a metrology false alarm, and process movement caused by suspect feedback.

Why the SOP matters

A scheduled SPC result may remain green while the physical tool changes during the following hours. That interval is the SPC blind window. The portal closes the operational gap by joining the last trusted qualification point with current production exposure, FDC movement, fleet behavior, matching quality, and yield context.

The expected productivity chain is:

Scattered tool, lot, SPC, FDC, matching, and yield evidence
                            ↓
One ranked Live Risk Board
                            ↓
One expandable evidence view with reason codes
                            ↓
One controlled recommendation and named owner
                            ↓
Faster verification, containment, passdown, and closure

The intended manufacturing defense is:

Detect hidden CD-SEM drift
          ↓
Protect measurement credibility
          ↓
Prevent suspect data from entering APC decisions
          ↓
Reduce unnecessary process changes and lot exposure
          ↓
Improve investigation speed and shift-to-shift continuity

Portal concepts an operator must understand

● Blind window: elapsed time between the last trusted SPC or golden-wafer result and the current production measurement.

● TMG: the tool-matching gap. The demonstration baseline is no more than 10% of target CD; the approved production percentage may be tighter.

● Mandel slope: multi-CD linearity indicator. The demonstration learning range is 0.98–1.02.

● Residual 3σ: random uncertainty after the fitted trend is removed. A rising residual can be more dangerous than a simple mean offset because the measurement becomes less repeatable.

● Fleet σ: distance from the governed fleet baseline. More than 2σ is an early warning; more than 3σ is a hold-review candidate in the demonstration logic.

● Site-to-site delta: spatial measurement range used to identify stage, vibration, scan-linearity, or center-edge artifacts.

● FDC health link: the relationship between a hardware signal and an observed metrology effect.

● Action label: an advisory next step that requires human review.

All displayed metrics must include a current timestamp and quality state in production. A stale green value is not evidence of health.

Start-of-shift SOP — first ten minutes

● Sign in through the approved application URL and verify that the displayed data-freshness timestamp is current.

● Review the KPI strip: tools watched, average SPC blind window, fleet out-of-control count, FDC health links, dynamic-limit state, and yield-watch workload.

● Open LIVE RISK BOARD.

● Sort by RISK and expand every high-risk tool, using 75 as the demonstration review threshold.

● Read the complete evidence, not only the color: layer, symptom, TMG, slope, fleet σ, residual, blind window, trajectory, reason codes, model version, and recommendation.

● Sort by BLIND WINDOW to identify tools whose last physical confidence point is old.

● Sort by FLEET σ to identify a tool separating from its peers even if its most recent SPC result remains green.

● Sort by RESIDUAL to identify developing repeatability, vacuum, vibration, charging, or beam-stability risk.

● Open FDC HEALTH-LINK for each high-risk tool and identify the most plausible hardware investigation path.

● Record the high-risk tool, affected layer, owner, current containment, exposed-lot window, and required exit criterion in the shift passdown.

Yellow trends are actionable information. The SOP does not require engineers to wait for a red alarm or a traditional three-sigma failure before reviewing a developing risk.

Before releasing a critical-layer lot

Before Gate, Fin, tight Contact/Via, risk-ramp, or yield-watch lots are measured or released, the responsible engineer verifies:

● the tool is not in HOLD REVIEW;

● the last trusted SPC or golden-wafer result is fresh enough for the layer;

● TMG remains within the approved layer-specific budget;

● Mandel slope remains within its approved linearity range;

● residual 3σ remains within the approved repeatability limit;

● site-to-site delta remains within the approved spatial limit;

● fleet deviation is not a hold candidate;

● the FDC fingerprint is stable or has an accepted explanation;

● the prediction is in domain and its required inputs are complete; and

● no unresolved APC guard applies to the measurement window.

If any condition is uncertain, the recommendation must become more conservative. Available containment includes running SPC, running a golden wafer, routing to a master tool, restricting the suspect tool to approved non-critical layers, guarding APC feedback, or calling a hold review.

How to interpret each portal action

RELEASE

Current evidence does not show a material tool-health threat to measurement credibility. Continue approved production use and trend monitoring. Release does not cancel normal qualification requirements and does not justify ignoring new FDC movement.

WATCH

The tool has an early drift signature, but a confirmed failure has not been established. Review the previous 8–24 hours of FDC behavior, compare the latest golden-wafer result, inspect maintenance and event logs, increase monitoring frequency, and prepare physical verification if the trend continues.

RUN SPC

The tool may not show a severe abnormality, but its physical confidence point is too old for the current exposure. Run the approved SPC or golden-wafer check before more critical lots. If constrained operation is necessary, only an authorized owner can approve an explicitly limited use.

GOLDEN WAFER

Virtual evidence is suspicious enough to require physical standard-wafer confirmation. Pause critical-layer measurement, use the approved recipe, and compare mean, three-sigma, TMG, slope, residual, site-to-site delta, and measurement profile with the governed baseline. A failed result escalates to hold review.

ROUTE LIMIT

The tool may remain usable for specifically approved, less-sensitive layers but is not suitable for critical work. Notify the dispatcher and module owner, restrict the affected layers, define an owner, and document measurable exit criteria before returning to full release.

APC GUARD

Measurement bias or noise may contaminate feedback. Identify lots measured during the suspect window, notify the APC owner, and decide whether affected measurements must be paused, excluded, reviewed, or repeated on a trusted tool. The portal cannot modify APC by itself.

HOLD REVIEW

The tool is a strong candidate for removal from critical measurement pending cross-functional review. Freeze the affected exposure window, check FDC, run physical confirmation, inspect image or waveform quality, compare against the master and fleet, determine affected lots, and select release, route limit, maintenance, calibration, or requalification through authorized procedures.

FDC-first diagnostic SOP

When the portal changes to WATCH, GOLDEN WAFER, ROUTE LIMIT, APC GUARD, or HOLD REVIEW, use the FDC Health-Link panel to select the first diagnostic branch.

Gun-vacuum branch

Typical portal evidence includes rising residual 3σ, worsening matching precision, random false alarms, apparent blur, or reduced peak-to-base ratio. Review chamber-vacuum trends, pump events, pressure spikes, recent venting or recovery, contamination indicators, and profile stability. Physical confirmation and stabilization take priority over re-baselining.

Emission-current branch

Typical evidence includes a mean-offset step, consecutive values on one side of the center line, an emission ramp, or concern that biased CD entered APC. Review emission current, extraction voltage, probe-current stability, gun age, approved recovery history, and changes after intervention. Guard the affected feedback window until credibility is restored.

Deflector-DAC branch

Typical evidence includes Mandel slope outside its approved range, dense/isolated bias, multi-CD mismatch, or site-to-site movement. Review deflector stability, scan-linearity calibration, image-shift correction, measurement-box placement, edge-algorithm configuration, and multi-CD standard-wafer behavior. Critical layers remain restricted until the linearity exit criterion passes.

Stage-vibration branch

Typical evidence includes widening residual, increasing site-to-site range, random site instability, or a false center-edge signature. Review the stage-vibration sensor, facility events, settling time, positioning logs, interferometer behavior, and repeated measurement at the same site. Repeatability and spatial stability must recover before full release.

Metric-driven response SOP

High blind window

Review all tool and facility events since the last trusted qualification point. Pay special attention to venting, beam restart, aperture work, preventive maintenance, vibration, temperature, and facility alarms. Run physical verification before critical work when the approved freshness boundary is exceeded.

TMG above the approved limit

Compare the suspect tool with its approved master, determine whether the change is a simple offset or includes wider residual error, verify recipe and correction-table versions, and run a multi-site or multi-CD standard wafer. Restrict cross-tool dispatch until matching returns within the approved rule.

Mandel slope outside the approved range

Perform a multi-CD linearity check, examine deflector and scan calibration evidence, compare dense and isolated features, and review edge-detection configuration. A tool that agrees at one CD but fails across CD sizes must not be treated as matched for critical layers.

Residual 3σ above the approved limit

Run a repeatability check on a standard wafer, inspect measurement profiles, review vacuum and emission stability, check facility and stage vibration, and verify focus and stigmator behavior. Guard APC if noisy measurements may have entered feedback.

Fleet deviation above the approved threshold

Compare product-lot means by tool, determine whether peers measuring the same product and layer remain stable, verify the suspect tool against the master, and backtrace lots measured during the deviation window. Fleet comparison is specifically intended to expose hidden inline drift before traditional SPC failure.

Site-to-site delta above the approved limit

Review site-level data rather than only the wafer mean. Check stage positioning, vibration, image-shift correction, interferometer evidence, and a multi-site standard-wafer result. Do not make a center-edge process disposition until a metrology artifact has been excluded.

HOLD REVIEW containment and recovery procedure

● Confirm the signal. Verify freshness, quality, risk, TMG, slope, fleet σ, residual, blind window, and reason codes.

● Freeze additional exposure. Stop unapproved critical-layer measurement and identify lots measured since the last trusted state.

● Select the FDC path. Review vacuum, emission/extraction, deflector, vibration, and applicable facility or thermal evidence.

● Run physical confirmation. Use an approved golden wafer, standard wafer, multi-CD wafer for slope concerns, or multi-site wafer for spatial concerns.

● Review image and waveform evidence. Inspect peak-to-base ratio, edge slope, baseline movement, X/Y asymmetry, blur, and tailing where those measurements are available.

● Compare with trusted peers. Use the master tool, governed fleet baseline, and comparable product distribution.

● Decide disposition. Select release, watch, route limit, APC guard, calibration, preventive maintenance, or continued hold through authorized review.

● Document and hand over. Record the hypothesis, evidence, affected-lot window, owner, due time, action, and exit criteria.

Exit criteria for return to full release

A tool does not return from route limit or hold review merely because an alarm clears. The owner verifies that:

● the approved standard or golden wafer passes;

● TMG is within the approved rule;

● Mandel slope is within the approved range;

● residual 3σ and site-to-site delta are within limits;

● fleet deviation has returned below the approved threshold;

● the abnormal FDC signal has returned to baseline or an accepted stable state;

● image or waveform quality has recovered where applicable;

● suspect APC feedback has been dispositioned;

● exposed lots have been reviewed by the responsible functions; and

● shift records contain the recovery evidence and exit criteria.

Yield-loss review evidence pack

When a yield engineer, process engineer, or process-integration engineer opens a case, the portal should produce a governed evidence pack rather than an informal collection of screenshots:

case and request identifier
product, layer, recipe, lot and wafer references
tool route and ADI/AEI tool pairing
last trusted SPC or golden-wafer timestamp
blind-window duration and exposed-lot window
TMG and matching history
Mandel slope and multi-CD evidence
residual 3σ and site-to-site delta
fleet deviation and peer distribution
FDC changes relative to baseline
maintenance and equipment-event history
APC feedback status
model and rule-set versions
reason codes and confidence
current containment, owner and due time
review disposition and exit criteria

The evidence pack supports one of five concise engineering conclusions:

● the tool is credible and a process cause is more likely;

● the tool is abnormal and a metrology false alarm is possible;

● the tool is abnormal and APC feedback contamination is possible;

● the evidence is mixed and master-tool remeasurement is required; or

● the evidence is mixed and standard-wafer verification is required.

End-of-shift passdown SOP

A passdown must describe measurement credibility, not merely “tool OK” or “tool NG.” The shift record contains:

Date and shift:
Engineer:
Tool-health summary:
High-risk CD-SEM tools:
Affected layers and lots:
FDC abnormal signals:
TMG / Mandel slope / residual 3σ / site delta / fleet σ:
Last trusted qualification time:
Current containment:
Golden-wafer or standard-wafer result:
APC guard status:
Maintenance or calibration action:
Owner and next action:
Due time:
Exit criteria for release:

This structured handoff is one of the application's most important productivity benefits. It prevents the next shift from repeating the same search and preserves an auditable connection between alert, evidence, action, and outcome.

Operational example — CDSEM-05

Assume the portal shows:

Tool: CDSEM-05
Layer: Contact/Via
Action: GOLDEN WAFER
Symptom: Vacuum degradation with rising residual
Blind window: 6.9 hours
TMG: 0.19 nm
Mandel slope: 1.006
Fleet deviation: 2.4 sigma
Residual 3σ: 0.21 nm

The slope remains inside the demonstration range, so multi-CD linearity is not the strongest signal. Residual is above the learning baseline and the tool is separating from its peers, while the vacuum trend provides a plausible hardware path. Because Contact/Via can be critical, the operator does not release blindly.

The engineer restricts critical use, runs the approved golden wafer, reviews the previous 24 hours of vacuum behavior, checks measurement-profile stability, performs a repeatability test, and compares the result with the master. If residual remains high, the owner opens the approved vacuum-recovery, contamination, or maintenance procedure. PE/PIE are notified about lots measured after the residual limit was crossed, and the APC owner reviews whether their feedback must be excluded.

This example shows the application's value clearly: one ranked case leads directly to the correct evidence, responsible roles, containment, and measurable exit criteria.

Actions the application must prevent

● Do not release a critical lot solely because the last SPC result was green when FDC changed afterward.

● Do not treat a mean-offset correction as sufficient when residual uncertainty is widening.

● Do not re-baseline before investigating the physical cause.

● Do not ignore an emerging fleet outlier solely because conventional SPC has not failed.

● Do not allow suspect CD data to influence APC without owner review.

● Do not treat one standard-wafer site as permanently stable; repeated exposure can age the reference.

● Do not close a hold review without documented exit criteria.

● Do not allow the AI score alone to make a hold or release decision.

SOP success metrics

The production team measures whether the portal improves work rather than simply generating alerts:

● drift cases detected before formal SPC failure;

● suspect APC feedback events prevented or reviewed;

● lots protected through route limit or APC guard;

● reduction in repeated CD-SEM false alarms;

● reduction in time to assemble a yield-review evidence pack;

● percentage of high-risk cases with complete evidence and named owners;

● alert-to-standard-wafer confirmation time;

● hold-review-to-disposition time;

● shift-passdown completeness;

● recommendation acceptance, rejection, and correction rates; and

● recurrence after tool recovery.

These outcomes are written to the governed S3 feedback zone. AWS Lambda aligns them with the original recommendation, and AWS Glue creates the quality dataset used for operational reporting and future model evaluation. Lake Formation ensures that only approved roles can use that feedback for model development.

One-page daily operating summary

● Start shift and verify data freshness.

● Open LIVE RISK BOARD and sort by risk.

● Expand high-risk tools and inspect TMG, slope, fleet σ, residual, blind window, trajectory, and reason codes.

● Sort by blind window, fleet deviation, and residual to find hidden exposure.

● Open FDC HEALTH-LINK and select the hardware investigation path.

● Before critical-lot release, verify that no hold, matching, linearity, repeatability, spatial, FDC, or APC concern remains unresolved.

● If evidence is suspicious, run physical confirmation or route to an approved master.

● If feedback is at risk, notify the APC owner and guard the affected window.

● If yield-watch lots are exposed, generate the governed evidence pack.

● End shift by recording credibility, affected lots, containment, owner, due time, and exit criteria.

This SOP turns the portal from a visual demonstration into a production-oriented productivity system: it shortens investigation time, standardizes daily decisions, improves cross-functional communication, preserves the evidence chain, and keeps every manufacturing action under human authority.