← Financial Cloud Cloud Cloud Club · Builder Articles

Vertex Macro | Financial Cloud Cloud · Builder Articles

Build with Kiro: Grammar and Pronunciation Enrichment Pipeline for Tagalog Cards Workshop

Series: Kiro workshop

Article: 05

Article
Kiro workshop
01 Build with Kiro: Prompt-First Product Design for a Tagalog Learning App
Kiro workshop
02 Build with Kiro: Educational-First Dev Tips for a Tagalog Learning App
Kiro workshop
03 Build with Kiro: Deep-Dive Development Flow for a Tagalog Learning App
Kiro workshop
04 Build with Kiro: Localize a Tagalog Learning App into Chinese Variants Workshop
Kiro workshop
05 Build with Kiro: Grammar and Pronunciation Enrichment Pipeline for Tagalog Cards Workshop
Kiro workshop
06 Build with Kiro: Unique and Reviewable Extra Examples in a Tagalog Learning App Workshop
Kiro workshop
07 Build with Kiro: Factory Engineering Health Hooks Workshop
Kiro workshop
08 Build with Kiro: Etch Process Window Risk Test Automation Workshop
Kiro workshop
09 Build with Kiro: Photolithography Drift Risk Development Workshop
Kiro workshop
10 Engineering Team Get Started — Daily Fab-Duty Use of fab spc drift sync portal
Kiro workshop
11 Engineering Team Addendum — Daily Fab-Duty Use of fab spc drift sync portal
Kiro workshop
12 Kiro: Field Engineering Workshop for Spec-Driven Factory Software
Kiro workshop
13 Kiro: Hands-On Lab — Build a Typed Factory Risk Portal from Scratch
Kiro workshop
14 Kiro: Prompt, Code, and Type Standards Playbook for Engineering Developers
Kiro workshop
15 Kiro: Why a Strong React Prompt Prevents Type Declaration False-Starts
Kiro workshop
17 Build with Kiro: Create a Factory Automation Portal React UI
Kiro workshop
18 Build with Kiro: Create the Automation Analytics Engine Behind a Factory Automation Portal
Kiro workshop
19 Build with Kiro: Add an AI Factory Automation Assistant to a Factory Automation Portal
Kiro workshop
21 Kiro: 2-Hour Professional Developer Workshop Guide
Kiro workshop
22 Kiro: Build the Fab SPC Drift Synchronization Portal from Scratch
Kiro workshop
23 Kiro: Prompt Library and Deep Code Explanation Appendix
Kiro workshop
30 Build with Kiro: Create a Factory Automation Portal UI
Kiro workshop
31 Build with Kiro: Create the Automation Analytics Engine Behind a Factory Automation Portal
Kiro workshop
32 Build with Kiro: Add an AI Factory Automation Assistant to a Factory Automation Portal
Kiro workshop
33 Build with Kiro: Rebuild the CME Direct-Style Quant P&L Leaderboard UI
Kiro workshop
34 Build with Kiro: Recreate the Quant Analytics Engine Behind the P&L Board
Kiro workshop
35 Build with Kiro: AWS AI-Powered Trading Desk Assistant for the Quant Board
Kiro workshop
36 One-Page Trading Portal SOP
Kiro workshop
AgentCore
A1 Build with AgentCore & Strands: Gateway MCP Tool Fabric Developer Workshop
AgentCore
A2 Build with AgentCore & Strands: Governed Multi-Agent Risk System Developer Workshop
AgentCore
A3 Build with AgentCore & Strands: Runtime Sovereign Risk Agent Developer Workshop
AgentCore
Exam practice
E1 Build a Multilingual AWS Exam Practice Launch System with Vibe Coding
Exam practice
E2 Build an AWS Exam Practice Room with Vibe Coding Dev Tips
Exam practice
E3 Build the Practice Engine Behind a Static AWS Exam Room
Exam practice
Amazon Q
Q1 Amazon Q: CloudShell-First Developer Workshop for ACM Certificate Auto Renewal
Amazon Q
Tagalog Practice Room
T1 Build a Tagalog Learning App for AWS Manila Community Day with Prompt-First Product Design
Tagalog Practice Room
T2 Build Tagalog Learning App for AWS Manila Community Day with Educational-First Dev Tips
Tagalog Practice Room
T3 Deep Dive Development Flow for a Tagalog Learning App for AWS Manila Community Day
Tagalog Practice Room
T4 Build Localize a Tagalog Learning App into Chinese Variants for AWS Manila Community Day
Tagalog Practice Room
T5 Build a Grammar and Pronunciation Enrichment Pipeline for Tagalog Cards for AWS Manila Community Day
Tagalog Practice Room
T6 Make Extra Examples Unique and Reviewable in a Tagalog Learning App for AWS Manila Community Day
Tagalog Practice Room
Roadmap
R1 Enterprise Data Analytics Roadmap: 100 Deep Scenario Questions
Roadmap
R2 Front-End Development Roadmap: Real-World Enterprise Scenarios
Roadmap
Hong Kong Community Day
C1 A Hong Kong Weekend with AWS Community Day: From Cloud Sessions to Harbour Lights
Hong Kong Community Day
C2 The Speaker’s Luxury Weekend: Present an AWS Story, Then Let Hong Kong Take the Stage
Hong Kong Community Day
C3 Seventy-Two Hours in Hong Kong: The Grand Tour for an AWS Community Day Speaker
Hong Kong Community Day
Manila Community Day
C4 AWS Community Day Manila: A Joyful Weekend of Cloud, Culture, and True Friendship
Manila Community Day
C5 AWS Community Day Manila: Where Cloud Builders Find the Happiest Spirit of the Philippines
Manila Community Day
C6 AWS Community Day Manila: Build, Break, Repeat, and Belong in a City of Joy
Manila Community Day
C7 First-Time Visitor Tips for Manila, Philippines
Manila Community Day
Philippines × Hong Kong
C8 Philippines Hong Kong Capital Market Upgrade
Philippines × Hong Kong
Backtest
B1 Build Institutional Amazon Long-Only Backtesting Agents With Bedrock AgentCore And Strands Agents
Long-only AMZN agents with AgentCore, Strands, and a governed Backtrader ledger.
B2 Build Regime-Aware Amazon Position Management With Backtrader, AgentCore, And Strands Agents
Treat market regime as a position control, not a chart comment.
B3 Build Benchmark-Relative Amazon Timing Systems Using Nasdaq, S&P 500, Dow, AgentCore, And Strands
Time AMZN against Nasdaq, S&P 500, and Dow context.
B4 Build A Governed Amazon Trade-History Factory With Bedrock AgentCore, Strands Agents, And Backtrader
Turn backtests into an auditable trade-history factory.
B5 Build An Agentic Amazon Backtest Operating Model With Bedrock AgentCore And Strands Agents [Part 1]
Build the operating model before debating the result.
B6 Build A Custom Cerebro Code Talk For Amazon Timing And Position Management [Part 2]
Explain the Cerebro engine before explaining the chart.
B7 Build Trader Review Records For Amazon Strategy Results And Lessons Learned [Part 3]
Turn strategy ranks into trader review records.
B8 Build A Governed FSI Amazon Position Management Playbook With AgentCore And Strands [Part 4]
An FSI playbook for governed Amazon position management.
B9 Build a Sovereign Risk Trading Agent with Amazon Bedrock AgentCore for Yield Spreads, FX Hedging, and Debt Repricing
Sovereign-risk agent for yield spreads, FX hedges, and debt repricing.
B11 Build Modern Volatility Trading & Lawful Thailand Recovery Planning Agents: A Memory-Driven Strands Multi-Agent Risk Protection System
Memory-driven Strands agents for volatility and Thailand recovery.
B12 Build Short Straddle Trading-Risk Governance with Amazon Bedrock AgentCore Memory
Short-straddle risk governance with AgentCore Memory.
B13 Building Production-Ready Credit & Yield Staking AI Agents on Amazon EKS
Production credit and yield-staking agents on Amazon EKS.
Challenge
01 Weekend Productivity Challenge: Fab SPC Drift Synchronization Portal
Fab SPC drift review and recommendation portal.
02 Weekend Productivity Challenge: Quant P&L Commander — An AI-Powered Trading Productivity Portal on AWS
Quant P&L leaderboard and trading productivity portal.
03 Weekend Annoying Task Challenge: Trading Desk Execute Summary On Cloud, On Chain, On Air
DeskPulse daily execution communication.
04 Weekend Agent Challenge: The 6 AM Trading Risk Review
An unattended, evidence-backed morning credit and trading risk brief.
05 Weekend Creative Challenge: Leadership Card Game
A browser-based creative facilitation deck.
06 Full Stack Challenge: Community Day Board App
A browser-based event communication room.
Leadership Card Game
01 Leadership Card Game: Last Skill Cloud Did Not Automate
A field essay for Builders on language, courage, and the Leadership Card Game
02 Anatomy of a Leadership Round: How the Leadership Card Game Actually Plays
A facilitator’s field guide for Builders who want drills that fit inside real meetings
03 Leadership Card Game: When the Opportunity Stops Belonging to the Organizer
A field essay for Builders on power transfer, multilingual practice nights, and career arcs that complete Entrance, Resource, and Narrative
04 Weekend Creative Challenge: Leadership Card Game
Master high-stakes workplace conversations before they happen.
05 From a Weekend Challenge Project to $1,386 Crowdfunding: The Leadership Practice That Changes How You Show Up at Work
A weekend build became a live 600-card leadership practice room and reached $1,386 in crowdfunding.
06 From a Weekend Challenge Project to $1,386 Crowdfunding: A Day 1 Path Into the Tech Industry
How did a weekend challenge become a multilingual AWS-powered product with 600 cards and $1,386 in crowdfunding?
07 From a Weekend Challenge Project to $1,386 Crowdfunding: Build a Professional Brand by Transferring Opportunity
A weekend challenge reached $1,386 in crowdfunding by turning leadership ideas into a working multilingual product.
08 Leadership Card Game — Crowdfunding Campaign
Speak leadership before the room decides your career.
09 PR/FAQ 01 — Leadership Card Game launches for community builders
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Community managers, volunteer organizers, early-career…
10 PR/FAQ 02 — Enterprise facilitators adopt Leadership Card Game for live leadership drills
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Learning & development leads, people managers, agile…
10 PR/FAQ 03 — Multilingual Leadership Card Game opens global practice rooms for builder ownership
Working Backwards document · External press release + FAQ Product: Leadership Card Game Audience: Global AWS builders, bilingual communities, cross-border…
AWS Builder Center
01 AWS Builder Center, its community spirit, and AWS Builder Jacket
There are destinations you reach by plane, destinations you enter through a door, and destinations that begin with a sign-in screen and quickly feel like a…
02 Inside AWS Builder Center, where a global technical platform becomes a place to learn, contribute, and belong
A great journey does not always begin at an airport.
03 AWS Community Builder huge success
When builders share openly, the entire community moves forward.
04 AWS Builder Center huge success
A vibrant global district built for curiosity, public learning, and the AWS Builder Jacket.
05 A weekend inside AWS Builder Center, from community inspiration to unmistakable AWS Builder Jacket
Friday evening begins with a familiar builder feeling: there is an idea waiting somewhere between a problem and a possibility.

Audience: professional developers building content-enrichment pipelines and educational tooling Duration: 2 hours Primary AWS AI service: Kiro Project output: a Kiro-guided deterministic enrichment pipeline that adds grammar breakdowns and pronunciation guides to Tagalog cards.

Educational engineering workshop only. This is a software architecture exercise and not process-release advice.

Workshop Summary

This workshop guides developers through enriching Tagalog cards with grammar notes and pronunciation support. Participants use Kiro to define deterministic boundaries, build glossaries, generate helper text, patch structured content or HTML, and validate coverage. The exercises emphasize reviewable language assistance: automation can prepare explanations, but human oversight keeps learner-facing grammar, sound guidance, and examples clear, accurate, consistent, and useful overall.

Workshop objective

Developers build a pipeline that extracts Tagalog sentences, adds grammar explanations, generates pronunciation guides, patches HTML or structured card data, and validates coverage.

2-hour agenda

Time Module Developer outcome
0–10 min Kiro setup Steering and specs created
10–25 min Enrichment contract Define helper output shape
25–45 min Glossary Build local word definitions
45–65 min Pronunciation Add curated map and fallback
65–90 min HTML patching Enrich extra-example blocks
90–110 min Validation Check missing spans and sections
110–120 min Hooks and review Automate checks and create handoff

Step 1 — Create Kiro steering for enrichment boundaries

Developer action

● Generate steering docs in Kiro.

● Add enrichment-specific rules.

● Ask Kiro what should be deterministic and what needs review.

● Commit steering before writing scripts.

Kiro prompt sample

Create steering docs for a Tagalog grammar and pronunciation enrichment pipeline. Use deterministic Python helpers, topic glossaries, pronunciation maps, fallback rules, BeautifulSoup patching, validation checks, and native-speaker review reminders.

System design decision

● Make Step 1 — Create Kiro steering for enrichment boundaries explicit before coding: Professional developers should not rely on hidden assumptions when using AI-assisted engineering. The workshop first writes the rule into steering or specs so Kiro has durable project context. This makes generated code more consistent, gives reviewers something concrete to inspect, and prevents repeated explanation in every chat. The decision also helps new developers understand why a file exists, what problem it solves, and which behavior is allowed or disallowed.

● Keep the implementation deterministic and reviewable: Kiro can help generate code, tests, and documentation, but the workshop output should be reproducible. Deterministic scripts, explicit configuration, stable schemas, and validation reports make the result easier to debug. When every transformation has a visible input and output, developers can review diffs, rerun checks, and explain the system to another engineer. This is especially important for language-learning content where correctness and cultural context require human review.

● Attach validation to the workflow, not only the final demo: The workshop treats validation as part of system design. Each step has a check, a report, or a hook so defects appear close to the change that caused them. This approach lets Kiro act as a coding assistant and quality reviewer while developers stay in control. The result is a practical professional workflow: plan with specs, guide with steering, implement in small tasks, validate output, and document handoff.

Code sample — .kiro/steering/enrichment.md

# Enrichment Steering

- Keep grammar and pronunciation enrichment deterministic.
- Use glossary dictionaries for known words.
- Use pronunciation maps for common words.
- Use transparent fallback rules for unknown words.
- Do not rewrite the Tagalog sentence during enrichment.
- Mark helper output as draft until native-speaker reviewed.
- Print validation counts after every batch.

Code explanation

● Business logic: The steering file defines enrichment scope and quality expectations.

● Code logic: Kiro uses the file as persistent context when generating scripts, tests, hooks, and docs.

● Expected result: Kiro-generated enrichment code should preserve Tagalog sentences and add reviewable helper sections.


Step 2 — Build a local grammar glossary

Developer action

● Create glossary.py.

● Add common words and event loanwords.

● Add transparent fallback behavior.

● Ask Kiro to generate glossary tests.

Kiro prompt sample

Create a beginner-friendly Tagalog glossary module. Include po, opo, saan, ang, paki, puwede, salamat, tubig, bayad, workshop, badge, registration, volunteer. Return short explanations and fallback text for unknown words.

System design decision

● Make Step 2 — Build a local grammar glossary explicit before coding: Professional developers should not rely on hidden assumptions when using AI-assisted engineering. The workshop first writes the rule into steering or specs so Kiro has durable project context. This makes generated code more consistent, gives reviewers something concrete to inspect, and prevents repeated explanation in every chat. The decision also helps new developers understand why a file exists, what problem it solves, and which behavior is allowed or disallowed.

● Keep the implementation deterministic and reviewable: Kiro can help generate code, tests, and documentation, but the workshop output should be reproducible. Deterministic scripts, explicit configuration, stable schemas, and validation reports make the result easier to debug. When every transformation has a visible input and output, developers can review diffs, rerun checks, and explain the system to another engineer. This is especially important for language-learning content where correctness and cultural context require human review.

● Attach validation to the workflow, not only the final demo: The workshop treats validation as part of system design. Each step has a check, a report, or a hook so defects appear close to the change that caused them. This approach lets Kiro act as a coding assistant and quality reviewer while developers stay in control. The result is a practical professional workflow: plan with specs, guide with steering, implement in small tasks, validate output, and document handoff.

Code sample — glossary.py

import re

DEFINITIONS = {
    "po": "politeness marker used to show respect",
    "opo": "polite form of yes",
    "saan": "where; asks for a place or direction",
    "ang": "focus marker before the main noun or idea",
    "paki": "please; softens a request",
    "puwede": "may or can",
    "salamat": "thank you",
    "tubig": "water",
    "bayad": "payment",
    "workshop": "English loanword used locally for a workshop session",
    "registration": "English loanword used for event check-in"
}

def token_key(word):
    return re.sub(r"[^a-zA-ZñÑáéíóúÁÉÍÓÚ]", "", word).lower()

def explain_word(word):
    key = token_key(word)
    return DEFINITIONS.get(key, f"needs review; useful local word: {word}")

def grammar_breakdown(sentence):
    seen = set()
    output = []
    for word in sentence.split():
        key = token_key(word)
        if key and key not in seen:
            seen.add(key)
            output.append((word.strip(".,?!"), explain_word(word)))
    return output[:6]

Code explanation

● Business logic: The glossary creates consistent beginner grammar notes for known words and transparent fallback notes for unknown words.

● Code logic: It normalizes tokens, checks a dictionary, avoids duplicates, and returns up to six word explanations.

● Expected result: Calling grammar_breakdown('Saan po ang registration?') returns explanations for Saan, po, ang, and registration.


Step 3 — Generate pronunciation with curated and fallback rules

Developer action

● Create pronunciation.py.

● Add curated pronunciations for high-frequency words.

● Add vowel fallback for unknown words.

● Ask Kiro to test known and unknown tokens.

Kiro prompt sample

Create a pronunciation helper for Tagalog cards. Use curated pronunciations for common words and transparent fallback rules for unknown words. Return a full guide and word-level chunks.

System design decision

● Make Step 3 — Generate pronunciation with curated and fallback rules explicit before coding: Professional developers should not rely on hidden assumptions when using AI-assisted engineering. The workshop first writes the rule into steering or specs so Kiro has durable project context. This makes generated code more consistent, gives reviewers something concrete to inspect, and prevents repeated explanation in every chat. The decision also helps new developers understand why a file exists, what problem it solves, and which behavior is allowed or disallowed.

● Keep the implementation deterministic and reviewable: Kiro can help generate code, tests, and documentation, but the workshop output should be reproducible. Deterministic scripts, explicit configuration, stable schemas, and validation reports make the result easier to debug. When every transformation has a visible input and output, developers can review diffs, rerun checks, and explain the system to another engineer. This is especially important for language-learning content where correctness and cultural context require human review.

● Attach validation to the workflow, not only the final demo: The workshop treats validation as part of system design. Each step has a check, a report, or a hook so defects appear close to the change that caused them. This approach lets Kiro act as a coding assistant and quality reviewer while developers stay in control. The result is a practical professional workflow: plan with specs, guide with steering, implement in small tasks, validate output, and document handoff.

Code sample — pronunciation.py

import re

PRON = {"salamat": "sah-lah-maht", "po": "poh", "saan": "sah-ahn", "kayo": "kah-yoh", "tubig": "too-beeg", "bayad": "bah-yahd", "pumasok": "poo-mah-sohk"}
VOWELS = {"a": "ah", "e": "eh", "i": "ee", "o": "oh", "u": "oo"}

def clean(word):
    return re.sub(r"[^a-zA-ZñÑ]", "", word).lower()

def fallback(word):
    return "".join(VOWELS.get(ch, ch) for ch in clean(word)) or word

def pronounce_word(word):
    return PRON.get(clean(word), fallback(word))

def pronunciation_guide(sentence):
    chunks = [(w.strip(".,?!"), pronounce_word(w)) for w in sentence.split() if clean(w)]
    return {"full": " ".join(sound for _, sound in chunks), "chunks": chunks}

Code explanation

● Business logic: The helper gives every Tagalog sentence a pronunciation guide for practice.

● Code logic: Known words use curated values; unknown words use vowel fallback; output includes full and chunked forms.

● Expected result: Calling pronunciation_guide('Paki-check po kung pumasok ang bayad.') returns a readable guide and word chunks.


Step 4 — Add glossary coverage reporting

Developer action

● Scan all Tagalog sentences.

● Count known and unknown tokens.

● Export frequent unknown words.

● Ask Kiro to suggest glossary additions for reviewer approval.

Kiro prompt sample

Create a glossary coverage report. Scan Tagalog sentences, count words not found in the glossary, rank unknown tokens by frequency, and export a JSON report for reviewer approval. Do not automatically add definitions without review.

System design decision

● Make Step 4 — Add glossary coverage reporting explicit before coding: Professional developers should not rely on hidden assumptions when using AI-assisted engineering. The workshop first writes the rule into steering or specs so Kiro has durable project context. This makes generated code more consistent, gives reviewers something concrete to inspect, and prevents repeated explanation in every chat. The decision also helps new developers understand why a file exists, what problem it solves, and which behavior is allowed or disallowed.

● Keep the implementation deterministic and reviewable: Kiro can help generate code, tests, and documentation, but the workshop output should be reproducible. Deterministic scripts, explicit configuration, stable schemas, and validation reports make the result easier to debug. When every transformation has a visible input and output, developers can review diffs, rerun checks, and explain the system to another engineer. This is especially important for language-learning content where correctness and cultural context require human review.

● Attach validation to the workflow, not only the final demo: The workshop treats validation as part of system design. Each step has a check, a report, or a hook so defects appear close to the change that caused them. This approach lets Kiro act as a coding assistant and quality reviewer while developers stay in control. The result is a practical professional workflow: plan with specs, guide with steering, implement in small tasks, validate output, and document handoff.

Code sample — glossary_coverage.py

import json
from collections import Counter
from glossary import DEFINITIONS, token_key

def coverage_report(sentences, output="glossary-coverage.json"):
    unknown = Counter()
    total = 0
    for sentence in sentences:
        for word in sentence.split():
            key = token_key(word)
            if not key:
                continue
            total += 1
            if key not in DEFINITIONS:
                unknown[key] += 1
    payload = {"totalTokens": total, "knownDefinitionCount": len(DEFINITIONS), "unknownTokenCount": sum(unknown.values()), "topUnknown": unknown.most_common(30)}
    with open(output, "w", encoding="utf-8") as file:
        json.dump(payload, file, indent=2, ensure_ascii=False)
    return payload

Code explanation

● Business logic: The report tells developers and reviewers which glossary gaps are most important.

● Code logic: It tokenizes sentences, compares normalized tokens with glossary keys, counts unknown words, and writes JSON output.

● Expected result: Running the report produces a ranked list of unknown tokens that reviewers can approve for future definitions.


Additional Hands-on Developer Labs

These labs are unique to Workshop 5 — Grammar and Pronunciation Enrichment Pipeline. They extend the deterministic enrichment workflow with glossary governance, pronunciation confidence, patch provenance, reviewer exports, and regression checks. The focus is enrichment quality, not general workspace setup.

Hands-on Lab A — Add glossary decision states for reviewer governance

Developer action

● Ask Kiro to extend the glossary from a simple definition map into reviewable records.

● Add decision states for approved, needs-review, and blocked.

● Update explain_word so blocked words do not produce learner-facing explanations.

● Generate tests for approved, unknown, and blocked glossary entries.

Kiro prompt sample

Refactor the glossary into reviewable glossary records.
Each record should include definition, partOfSpeech, decisionState, reviewerNote, and lastReviewedAt.
Approved entries may appear in learner output.
Needs-review entries should be marked draft.
Blocked entries should be excluded from learner-facing grammar notes.
Create tests for all decision states.

System design decision

● Glossary governance protects learner output: Grammar explanations are teaching content, so a word definition needs review status, not only text.

● Blocked entries must fail closed: If a reviewer blocks an explanation, the enrichment pipeline should omit it or label it safely instead of publishing it.

● Structured records support future review tools: A record shape can be exported to CSV, reviewed by native speakers, and imported back into the pipeline.

Code sample — glossary_records.py

from dataclasses import dataclass, asdict
from typing import Literal

DecisionState = Literal["approved", "needs-review", "blocked"]

@dataclass(frozen=True)
class GlossaryRecord:
    term: str
    definition: str
    partOfSpeech: str
    decisionState: DecisionState = "needs-review"
    reviewerNote: str = "Needs native-speaker review."
    lastReviewedAt: str | None = None

GLOSSARY: dict[str, GlossaryRecord] = {
    "po": GlossaryRecord(
        term="po",
        definition="politeness marker used to show respect",
        partOfSpeech="particle",
        decisionState="approved",
        reviewerNote="Common beginner-safe explanation.",
        lastReviewedAt="2026-06-01"
    ),
    "bayad": GlossaryRecord(
        term="bayad",
        definition="payment",
        partOfSpeech="noun",
        decisionState="needs-review"
    )
}

def learner_definition(term: str) -> dict:
    record = GLOSSARY.get(term.lower())
    if record is None:
        return {"term": term, "definition": "needs review", "decisionState": "needs-review"}
    if record.decisionState == "blocked":
        return {"term": record.term, "definition": "blocked from learner output", "decisionState": "blocked"}
    return asdict(record)

Code explanation

● Business logic: The record model separates approved learner content from draft or blocked explanations.

● Code logic: GlossaryRecord stores review metadata, and learner_definition returns safe output based on decision state.

● Expected result: Enrichment can continue while clearly marking which glossary notes are approved, draft, or blocked.

Hands-on Lab B — Add pronunciation confidence and reviewer flags

Developer action

● Ask Kiro to add confidence metadata to pronunciation output.

● Mark curated pronunciations as high confidence and fallback pronunciations as low confidence.

● Export low-confidence tokens for review.

● Add a validation threshold for the maximum allowed low-confidence token ratio.

Kiro prompt sample

Add pronunciation confidence to the pronunciation helper.
Curated map entries should return confidence high.
Fallback entries should return confidence low and reviewRequired true.
Create a report of low-confidence tokens by frequency.
Fail validation if more than 25 percent of token pronunciations are low confidence.

System design decision

● Fallback output should be transparent: A fallback pronunciation is useful for coverage but should not appear equally trusted as curated guidance.

● Confidence supports prioritization: Reviewers can focus on high-frequency low-confidence tokens first.

● Thresholds make quality measurable: The pipeline can fail when too much output depends on fallback rules.

Code sample — pronunciation_confidence.py

import re
from collections import Counter

PRON = {"salamat": "sah-lah-maht", "po": "poh", "saan": "sah-ahn"}
VOWELS = {"a": "ah", "e": "eh", "i": "ee", "o": "oh", "u": "oo"}


def clean(word: str) -> str:
    return re.sub(r"[^a-zA-ZñÑ]", "", word).lower()


def fallback(word: str) -> str:
    return "".join(VOWELS.get(ch, ch) for ch in clean(word)) or word


def pronounce_token(word: str) -> dict:
    key = clean(word)
    if key in PRON:
        return {"word": word, "sound": PRON[key], "confidence": "high", "reviewRequired": False}
    return {"word": word, "sound": fallback(word), "confidence": "low", "reviewRequired": True}


def low_confidence_report(sentences: list[str]) -> dict:
    low = Counter()
    total = 0
    for sentence in sentences:
        for word in sentence.split():
            result = pronounce_token(word)
            if clean(word):
                total += 1
            if result["confidence"] == "low":
                low[clean(word)] += 1
    return {"totalTokens": total, "lowConfidenceTokens": sum(low.values()), "topLowConfidence": low.most_common(20)}

Code explanation

● Business logic: Pronunciation output now tells reviewers which guidance is curated and which needs review.

● Code logic: The helper returns structured token metadata and a ranked report of low-confidence words.

● Expected result: The enrichment pipeline can produce full pronunciation coverage while prioritizing human review.

Hands-on Lab C — Add enrichment provenance to every patched card

Developer action

● Ask Kiro to stamp each enriched card with pipeline version, source file, card ID, and enrichment timestamp.

● Render provenance as a hidden comment or structured JSON sidecar.

● Add validation that every enriched card has provenance.

● Generate a diff report that lists newly enriched cards.

Kiro prompt sample

Add enrichment provenance to HTML patching.
For every card patched with grammar or pronunciation, record cardId, sourceFile, enrichmentVersion, enrichedAt, grammarTermCount, and pronunciationTokenCount.
Write provenance to enrichment-provenance.json and validate that every patched card appears in it.

System design decision

● Patching needs traceability: When HTML is modified after generation, developers need to know which script changed which card.

● Sidecar files avoid UI clutter: Reviewers and developers can inspect provenance without adding noise to learner-facing pages.

● Versioning supports regression review: A changed enrichment version can trigger targeted review of affected cards.

Code sample — enrichment_provenance.py

from datetime import datetime, timezone
import json
from pathlib import Path

ENRICHMENT_VERSION = "grammar-pron-v1"


def provenance_record(card_id: str, source_file: str, grammar_terms: int, pronunciation_tokens: int) -> dict:
    return {
        "cardId": card_id,
        "sourceFile": source_file,
        "enrichmentVersion": ENRICHMENT_VERSION,
        "enrichedAt": datetime.now(timezone.utc).isoformat(),
        "grammarTermCount": grammar_terms,
        "pronunciationTokenCount": pronunciation_tokens
    }


def write_provenance(records: list[dict], path: str = "enrichment-provenance.json") -> None:
    ordered = sorted(records, key=lambda item: (item["sourceFile"], item["cardId"]))
    Path(path).write_text(json.dumps(ordered, indent=2, ensure_ascii=False), encoding="utf-8")

Code explanation

● Business logic: Provenance makes enrichment patches auditable after the workshop.

● Code logic: Records are sorted for stable diffs and written as UTF-8 JSON.

● Expected result: Reviewers can trace each grammar/pronunciation section back to a source file and pipeline version.

Hands-on Lab D — Build a grammar-note length and readability validator

Developer action

● Ask Kiro to create readability rules for beginner grammar notes.

● Limit each definition to a short phrase or sentence.

● Flag notes that contain advanced terminology unless approved.

● Export failures with card ID, term, and reason.

Kiro prompt sample

Create a validator for beginner-friendly grammar notes.
Each definition should be at most 18 words.
Flag advanced terms such as enclitic, absolutive, ergative, morphosyntax, and aspect unless allowAdvanced is true.
Return JSON failures with cardId, term, definition, and reason.

System design decision

● Beginner readability is testable: Grammar notes can be mechanically checked for length and advanced terminology.

● Validators complement human review: The script catches obvious complexity before reviewers spend time on nuance.

● Allowlist keeps flexibility: Advanced notes can still exist when deliberately approved.

Code sample — validate_grammar_notes.py

ADVANCED_TERMS = {"enclitic", "absolutive", "ergative", "morphosyntax", "aspect"}


def word_count(text: str) -> int:
    return len([part for part in text.split() if part.strip()])


def validate_note(card_id: str, term: str, definition: str, allow_advanced: bool = False) -> list[dict]:
    failures = []
    if word_count(definition) > 18:
        failures.append({"cardId": card_id, "term": term, "reason": "definition too long"})
    lowered = definition.lower()
    if not allow_advanced:
        blocked = sorted(word for word in ADVANCED_TERMS if word in lowered)
        if blocked:
            failures.append({"cardId": card_id, "term": term, "reason": f"advanced terminology: {', '.join(blocked)}"})
    return failures

Code explanation

● Business logic: The validator keeps helper text suitable for beginners.

● Code logic: It checks word count and blocked advanced terms, returning structured failures.

● Expected result: Long or overly technical grammar notes are flagged before publication.

Hands-on Lab E — Create enrichment snapshot tests for HTML patching

Developer action

● Ask Kiro to create a small fixture HTML file with one card.

● Run the patcher against the fixture.

● Assert that grammar, pronunciation, review note, and provenance markers exist.

● Avoid full-page snapshots unless the patcher is intentionally layout-sensitive.

Kiro prompt sample

Create regression tests for the HTML enrichment patcher.
Use a tiny fixture with one card and one Natural Tagalog span.
After patching, assert that grammar breakdown, pronunciation guide, draft review note, and card provenance marker are present.
Do not compare the entire HTML document.

System design decision

● Patchers are easy to break: A small selector change can silently stop enrichment from appearing.

● Targeted assertions reduce brittleness: Tests should protect required sections without failing on harmless formatting changes.

● Fixture-driven tests document assumptions: Future developers can see the expected HTML shape.

Code sample — tests/test_enrichment_patch.py

from bs4 import BeautifulSoup


def add_enrichment(html_text: str) -> str:
    soup = BeautifulSoup(html_text, "html.parser")
    card = soup.select_one(".sentence-card")
    grammar = soup.new_tag("section", **{"class": "grammar-breakdown"})
    grammar.string = "Grammar breakdown: draft"
    pronunciation = soup.new_tag("section", **{"class": "pronunciation-guide"})
    pronunciation.string = "Pronunciation guide: draft"
    card.append(grammar)
    card.append(pronunciation)
    card.append(soup.new_string("\n<!-- enrichmentVersion: grammar-pron-v1 -->\n"))
    return str(soup)


def test_patcher_adds_required_enrichment_sections():
    html = "<article class='sentence-card'><span lang='tl'>Saan po?</span></article>"
    patched = add_enrichment(html)
    assert "grammar-breakdown" in patched
    assert "pronunciation-guide" in patched
    assert "enrichmentVersion: grammar-pron-v1" in patched

Code explanation

● Business logic: The test protects required enrichment sections in generated cards.

● Code logic: A minimal fixture is patched and checked with targeted assertions.

● Expected result: Selector or patcher regressions are caught before a full batch run.

Hands-on Lab F — Export a reviewer packet for enrichment decisions

Developer action

● Ask Kiro to combine glossary gaps, low-confidence pronunciations, and long grammar notes into one reviewer packet.

● Write both JSON and CSV outputs.

● Include suggestedAction values such as approve-definition, fix-pronunciation, and shorten-note.

● Add a summary count by action type.

Kiro prompt sample

Create an enrichment reviewer packet.
Combine glossary unknowns, low-confidence pronunciation tokens, and grammar-note readability failures.
Export enrichment-review-packet.json and enrichment-review-packet.csv.
Each row should include issueType, tokenOrTerm, exampleSentence, frequency, suggestedAction, and reviewerDecision.

System design decision

● Review packets reduce reviewer friction: A single artifact is easier to review than separate console outputs.

● Suggested actions make triage faster: Reviewers can approve, correct, block, or request rewrite without interpreting raw validation logs.

● Empty reviewerDecision preserves human authority: The pipeline prepares work but does not claim final approval.

Code sample — review_packet.py

import csv
import json
from collections import Counter
from pathlib import Path

COLUMNS = ["issueType", "tokenOrTerm", "exampleSentence", "frequency", "suggestedAction", "reviewerDecision"]


def write_review_packet(rows: list[dict], json_path="enrichment-review-packet.json", csv_path="enrichment-review-packet.csv") -> dict:
    summary = Counter(row["suggestedAction"] for row in rows)
    payload = {"summary": dict(summary), "items": rows}
    Path(json_path).write_text(json.dumps(payload, indent=2, ensure_ascii=False), encoding="utf-8")
    with open(csv_path, "w", newline="", encoding="utf-8") as file:
        writer = csv.DictWriter(file, fieldnames=COLUMNS)
        writer.writeheader()
        for row in rows:
            writer.writerow({column: row.get(column, "") for column in COLUMNS})
    return {"json": json_path, "csv": csv_path, "summary": dict(summary)}

Code explanation

● Business logic: The reviewer packet converts enrichment issues into reviewable work items.

● Code logic: It writes stable JSON and CSV outputs and summarizes suggested actions.

● Expected result: Native-speaker and curriculum reviewers can work from one consolidated packet.


Reference architecture notes

● Kiro capabilities emphasized in this workshop: enrichment steering, deterministic glossary records, pronunciation confidence reporting, HTML patch provenance, validation generation, fixture-based regression tests, and reviewer-packet documentation.

● Product scope: grammar and pronunciation helper text for Tagalog learning cards. Helper output remains draft until reviewed by native Tagalog speakers.

● Runtime scope: local Python enrichment pipeline first. Optional automation can later run the same validators in CI before packaging static learning artifacts.