← Financial Cloud Cloud Cloud Club · AWS Re:cap

Vertex Macro | Financial Cloud Cloud · AWS Re:cap

AWS Re:cap 05: A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems

Speaker: Government data

Session: 05

Session
Summit Dev Lounge2026 Re:cap
01 Encode Architecture as Steering for AI Agents
Summit Dev Lounge2026 Re:cap
02 Agent Harness Is the Real Engineering Moat
Summit Dev Lounge2026 Re:cap
03 Ask Observability Data in Plain Language
Summit Dev Lounge2026 Re:cap
04 Serverless AR Game with Bedrock AgentCore
Summit Dev Lounge2026 Re:cap
05 Multi-Agent Quant Backtesting on AgentCore
Summit Dev Lounge2026 Re:cap
06 Blog to Slides in Three Minutes with Kiro
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 AWS Compliance with Terraform
AWS Community Day Hong Kong 2025 Re:cap
03 Beginner to Builder An AWeSome Cloud Journey
AWS Community Day Hong Kong 2025 Re:cap
04 Team-First Serverless Engineering with Laravel & Bref
AWS Community Day Hong Kong 2025 Re:cap
05 Event Opening - AWS Community Day Hong Kong 2025
AWS Community Day Hong Kong 2025 Re:cap
06 Agent-to-Agent: Building Interoperable AI on AWS
AWS Community Day Hong Kong 2025 Re:cap
07 Utilize another telemetry data for faster improvement with AI agent
AWS Community Day Hong Kong 2025 Re:cap
08 Graduating from Vibe Coding: Spec-Driven Development with Kiro
AWS Community Day Hong Kong 2025 Re:cap
09 Automated Testing using MCP & AI Agents
AWS Community Day Hong Kong 2025 Re:cap
10 Modernizing Telecom Security ML Powered Approach
AWS Community Day Hong Kong 2025 Re:cap
11 Rethinking GenAI Agent: RAG & MCP
AWS Community Day Hong Kong 2025 Re:cap
12 Disaster and Emergency Response with TAK and AWS
AWS Community Day Hong Kong 2025 Re:cap
13 Rethinking Serverless Application Workflows from a Testing Perspective
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 AI-Powered Global Pure-Alpha Macro Trades on AWS: Revolutionizing Risk-Adjusted Asset Returns
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 Modern Trade Lifecycle: Trading to Settlement
FSI Recap
02 Goldman Sachs: Fast Track your applications onto Cloud - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group: Building an Effective Log Management Solution on AWS
FSI Recap
04 FSI Meetup 2025 Q4 - Brex Database Disaster Recovery
FSI Recap
05 FSI Meetup 2025 Q4 - A Graviton Migration Success Story
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel Modern Data Platform
FSI Recap
07 FSI Meetup 2025 Q4 - Financial Transaction Data Reconciler PayPal
FSI Recap
08 FSI Meetup 2025 Q4 - Scaling Resilience
FSI Recap
09 Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances
FSI Recap
10 Advanced Agentic AI Design Patterns
FSI Recap
11 Build New Modern Apps on AWS
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent Recap (IND3312)
AWS re:Invent 2025
02 Building the Future Trading Platform Leveraging AI and AWS
AWS re:Invent 2025
03 Trading Innovation: Jefferies' AI Assistant on Amazon Bedrock (IND3315)
AWS re:Invent 2025
04 How FSI Revolutionized HFT Analytics with Agentic AI (GBL302)
AWS re:Invent 2025
05 Improving Distributed Systems with Amazon Time Sync Featuring Nasdaq
AWS re:Invent 2025
06 Amazon Aurora HA and DR Design Patterns for Global Resilience (DAT442)
AWS re:Invent 2025
07 Building Agentic AI: Amazon Nova Act and Strands Agents in Practice (DEV327)
AWS re:Invent 2025
08 Deep Dive into Amazon Aurora and Its Innovations (DAT441)
AWS re:Invent 2025
09 Deep dive on Amazon S3 (STG407)
AWS re:Invent 2025
10 Nasdaq: Build Resilient Infrastructure for Global Financial Services (HMC327)
AWS re:Invent 2025
11 What's New with AWS Lambda (CNS376)
AWS re:Invent 2025
12 Spec-Driven Development with Kiro (DEV314)
AWS re:Invent 2025
13 Amazon's finops: Cloud cost lessons from a global e-commerce giant (AMZ308)
AWS re:Invent 2025
14 Tick to trade latency trading platforms on aws
AWS re:Invent 2025
Government data
01 The AI Era: The Boundary Between Development and Design Is Disappearing
Government data
02 On-Device Multimodal AI and Smart-City Practice
Government data
03 Large-Model Capability Evaluation and a Method for Landing AI Projects
Government data
04 Controlled End-to-End Automation of Government Development with Cloud Agents
Government data
05 A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
Government data
06 AI-Driven Macro Quantitative Research and Smart Governance
Government data
07 Authorized Operation of Public Data and Smart-Government Practice
Government data
08 Putting Data Assetization into Practice: Rights, Compliance, Engineering Governance, and Digital-Government Cases
Government data
09 AI for Mental-Health Public Welfare: Governance, Architecture, and Practice of a Trustworthy Platform
Government data
Amarathon 2025 Recap
01 A Developer’s Roadmap to Architecting for Agents
Donnie Prakoso
02 Amazon Bedrock Data Automation
Hafiz Syed Ashir Hassan
03 Multi-Agent on AgentCore
Tan Xin
04 Building Agentic AI Nova Act and Strands Agents in Practice
Haowen Huang
04 Accelerating Migration Projects with Kiro using Spec-Driven Development
Sanchit Dilip Jain
06 From Matching to Understanding: Personalized AI Search Practice Driven by AgentCore Memory
Liu Cao
07 Observe to Optimize – LLM Observability to AIOps Turning real-time insights into intelligent automation
Jimmy Soh
08 Deploying TEAM and Building the Best Engineering Team
Yuji Oshima
09 Five Hard Lessons from Five Years of So-Called Serverless Databases
Renato Losio
14 What if AI does my job How Q Developer CLI and Kiro have changed my daily routine
Miguel Angel Muñoz
16 Velocity with Vigilance: Security Essentials for Amazon Bedrock Agent Development
Brian Tarbox
26 Run OSS LLMs on a Single H100 Smarter, Cheaper, Faster
Adit Modi
28 A Modern Unified Metadata Architecture: New Approaches to Breaking Down Data Silos
Shaofeng Shi
29 Serverless MediaOps: Automating Video Workflows with AI on Amazon Web Services
Luis Valdivia
30 Architecting for Efficiency and Reliability with Performance Testing at Scale
Luis Guirigay
31 Connecting the World Through Open Source: Practical Journey of Technology, Community and Global Developer Relations
Richard Lin
33 Building Streaming Iceberg Tables for Real-Time Logistics Analytics
Fahad Shah
34 Accelerating Large-Scale Robot Strategy Training: An Automated Closed-Loop Architecture Based on Kiro, Trainium, and EKS
Junjie Tang
35 From Vibe to Viable with spec driven development
Ricardo Sueiras
36 Making Cloud Cost Analysis Smarter: Building FinOps Intelligent Agents with Strands and AgentCore
Xiaofei Li
37 Transform Conversational Agentic AIOps for K8s Using CNCF Kagent, K8sGPT, and Nova Sonic
Shaoyi Li

A deep briefing for government technology advisers, digital-government officers, data-governance leads, cybersecurity officers, enterprise architects, and large-programme decision-makers.


The agent turning point in the government software ecosystem

Public value

Judge success by processing time, one-stop completion rate, and frontline burden. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Engineering essence

A multi-agent system is a distributed system with probabilistic behaviour. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Accountability red line

Statutory decisions must not be left to an agent vote. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Starting method

First decompose the process, then mark rule, reasoning, and human nodes. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


The correct division of labour between agents and microservices

Problem and judgement

Microservices carry transactions, ledgers, identity, and reproducible rules. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Agents handle semantic understanding, evidence collection, and draft recommendations. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

The core system remains the authoritative data state. The model must not overwrite it directly. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Every model call must prove more value than rules or search. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Fitness criteria for multi-agent systems

Problem and judgement

Multiple professions, data domains, and permission sets, plus a high exception rate, are positive signals. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Use traditional automation for fixed processes, a single data source, and low reasoning value. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Score by task variability, cross-domain dependency, exception rate, and reasoning gain. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Before a PoC, establish baselines for time, error, rejection, labour hours, and appeals. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Five roles and isolation of duties

Problem and judgement

The coordinator splits tasks and maintains state, but does not judge policy legality. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Expert agents provide policy, data, calculation, or text output. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Gatekeeper and audit agents check permission, purpose, and completeness of evidence. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

The human accountable owner holds approval, veto, and reversal rights for high-risk decisions. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Centralised, hierarchical, and decentralised collaboration

Problem and judgement

Centralised orchestration is best for formal government processes and consistent audit. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Hierarchical mode suits large cross-domain tasks, but must prevent summary distortion. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Decentralised negotiation is suitable only for low-risk exploration. It must not replace statutory accountability. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Every mode must limit rounds, cost, timeout, and conflict escalation. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Task contracts and structured schemas

Problem and judgement

Inputs include purpose, data scope, tools, deadline, cost, and risk. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Outputs include conclusion, citations, confidence, open items, and next steps. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Amounts, dates, eligibility, and identity must pass type and range validation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Contracts must have versions, compatibility policy, tests, and change approval. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Shared state and agent memory

Problem and judgement

Process state, business state, and semantic memory must be governed separately. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The core business system is the authoritative source. Vector memory is only an aid. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Long-running processes use event sourcing, idempotency keys, and Saga compensation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Each class of state must have create, use, archive, and delete cycles. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Security boundaries for tool calls

Problem and judgement

Split permissions across search, read, write, payment, notification, and delete. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Before execution, verify identity, delegation, purpose, data classification, and impact scope. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

After execution, verify the result, obtain a receipt, and write tamper-evident evidence. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Red-team test prompt injection, malicious files, out-of-scope parameters, and tool confusion. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Agent identity and zero trust

Problem and judgement

People, services, and agents use distinct and traceable identities. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Short-lived credentials bind task, purpose, data scope, and validity period. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Authorisation includes role, sensitivity, region, time, and risk conditions. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

When an agent is retired, revoke keys, tools, memory, and schedules in the same step. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


The full lifecycle of a digital worker

Problem and judgement

The job description lists who is served, what may be done, what must not be done, and the escalation path. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The permission card separates readable, writable, recommendable, and approvable. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Performance is judged on correctness, citations, refusal, appeals, and rework together. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

On decommission, hand over open cases, revoke rights, archive evidence, and update the catalogue. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


The real design of human review

Problem and judgement

Place humans before conflicts, irreversible tools, and low-confidence outputs. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The review interface shows raw data, policy, disagreements, and expected impact. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

The reviewer may veto, amend, request further evidence, or transfer. Reasons must be recorded. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

High risk is fully reviewed. Medium and low risk use thresholds, sampling, and consistency monitoring. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Turn agent conflict into a governance signal

Problem and judgement

Conflict often comes from version, date, data completeness, and jurisdiction differences. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The conflict pack lists the points of dispute, each party's basis, missing data, and impact. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

First check data and versions, then apply policy priority, and finally adjudicate by a human. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Regular analysis of conflict can reveal ambiguous policy and inconsistent cross-department standards. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Reliability, degradation, and compensation

Problem and judgement

Set timeout, backoff, and maximum retries for every agent and tool. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

When the model or an external cloud fails, switch to rules, search, a human, or deferral. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Unhandleable events go to a dead-letter queue with context and a handling deadline. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Chaos drills verify RTO, RPO, alerts, takeover, and recovery. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Multi-agent cost engineering

Problem and judgement

Split cost into inference, retrieval, API, compute, storage, labour, and operations. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Each task sets limits on rounds, context, tool calls, and total cost. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Use rules or a small model for classification and extraction. Use a high-capacity model only for complex reasoning. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Contracts include unit cost, peak throughput, and three-year total cost of ownership. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


End-to-end observability and audit

Problem and judgement

A case trace identifier links request, model, knowledge, tools, humans, and result. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Layer technical, agent, business, and governance metrics. Do not look only at latency. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Logs mask sensitive data and separate operations access from case access. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Review failures, cost, knowledge expiry, reversals, and vendor availability every day. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Evaluation is more than answer correctness

Problem and judgement

The benchmark set covers normal, boundary, missing-data, conflict, and adversarial inputs. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Evaluate task success, field correctness, citations, hallucination, and refusal. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Also test leakage, privilege escalation, fairness, latency, availability, and disaster recovery. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Re-run regression after any change to the model, knowledge, prompt, or tools. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Stage gates from PoC to production

Problem and judgement

In exploration, use synthetic, anonymised, or public data and set stop conditions. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

A controlled pilot limits business, population, and time window. Every result is checked by a human. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Before production, complete classified cybersecurity protection grading, commercial cryptography application security assessment, data classification, load testing, and disaster recovery. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Expand with canary release and rollback. Repeat risk assessment whenever an agent or tool is added. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Success case: Hong Kong digital government shared foundation

Problem and judgement

By the end of 2025, more than one hundred digital government and smart city measures will have been advanced. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

A next-generation government cloud, big data, shared blockchain, and a common chat service support departments. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

iAM Smart has more than four million users and covers more than 1,300 services and forms. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

The success lesson is to unify identity, data exchange, and security first, then extend intelligence. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Success case: consent-based data exchange gateway

Problem and judgement

With the citizen's consent, deliver authoritative-source data to a designated electronic service. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

About two million data exchanges a month reduce repeat submission and manual verification. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

A data agent may request only the fields required by policy. It must not extend the consented purpose. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Each exchange retains evidence of source, time, purpose, recipient, and withdrawal handling. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Success case: full digitalisation and enterprise identity

Problem and judgement

Electronic payment, electronic submission, and electronic approval documents have been fully digitalised. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

CorpID, the enterprise digital identity, is expected to launch by the end of 2026 and expand services in stages. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Core capabilities include enterprise verification, digital signing, pre-fill, and a document wallet. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

An agent may prepare materials, but formal signing and high-risk submission remain with the authorised person. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Success case: multi-vendor AI+ public services

Problem and judgement

The capability catalogue covers customer service, meetings, documents, writing, processes, creativity, and analysis. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Compare multiple vendors on the same field, and choose by Chinese capability, data, latency, cost, and exit. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Technical matching must include a real process, a baseline, and data constraints — not a demonstration alone. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Different models may connect through standard tools, but identity, logs, and evaluation must be unified. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Greater Bay Area cross-border data and rules

Problem and judgement

A memorandum of cooperation on promoting cross-border data flow in the Greater Bay Area was signed in 2023. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Standard-contract facilitation measures have been extended to all industries since November 2024. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Before an agent transmits, identify data type, origin, destination, purpose, and duration. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Pilot first with low-sensitivity services that have a clear purpose and a clear authoritative source. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Digital-first Northern Metropolis

Problem and judgement

The next five years plan more than 70,000 housing units and one million square metres of economic floor space. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The university town, the Hong Kong Park of the Loop, and San Tin Technopole connect research and industrial conversion. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Spatial, building, transport, energy, and environmental data form the city-operations foundation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Identity, data standards, communications, edge, disaster recovery, and cybersecurity should enter early planning. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Enterprise benefits without application: case panorama

Problem and judgement

Policy, data, compliance, calculation, notification, and audit agents divide professional labour. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The coordinator maintains the process but cannot approve subsidies or change enterprise data. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Agent conflicts are adjudicated by the business accountable owner. Payment runs through the core finance system. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Measure success by processing time, hit rate, error rate, and enterprise burden. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Enterprise-benefit case: policy knowledge engineering

Problem and judgement

Split provisions into subject, time, region, exclusion, evidence, formula, and discretion. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Each condition links back to the original text, issuing authority, version, and effective date. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Give clear thresholds to rules. Use a model only for semantic extraction. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Revision, suspension, and expiry trigger a process. Old versions are retained to reproduce historical cases. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Enterprise-benefit case: data verification

Problem and judgement

Enterprise registration, tax, employees, licences, and subsidies each have an authoritative source. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The agent requests the minimum fields required by the condition. Boolean or range answers can reduce disclosure. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Missing, delayed, or contradictory data is marked as undeterminable. Do not guess eligibility or refusal. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Verification results record source, time, transformation rule, and purpose of use. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Enterprise-benefit case: isolation of calculation and payment

Problem and judgement

Subsidy amounts are calculated by a versioned rule service. The model explains; it does not post to the ledger. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Candidate results are checked by an independent rule or a second calculation instance. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

The agent only generates a payment draft. The core finance system validates the budget and the approval chain. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Reconcile approvals, instructions, and bank results daily. Recovery and correction follow a compensation process. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Enterprise-benefit case: notification and appeal

Problem and judgement

The notice explains conditions, data, calculation, missing documents, and the ultimately accountable authority. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Website, SMS, email, hotline, and counter use a consistent case status. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

The enterprise may view the basis, supply further evidence, request human review, and lodge an appeal. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Track comprehension, number of supplementary submissions, cycle time, reversals, and satisfaction. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Agent boundaries in smart mobility

Problem and judgement

Agents integrate real-time traffic, parking, arrival, road-sensor, and incident data. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Signals, charging, and enforcement are controlled by deterministic systems, with human takeover retained. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

During major events, departments share incident state but keep their own statutory powers. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Metrics include journey reliability, response time, false alerts, availability, and complaints. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


High-risk layering in smart healthcare

Problem and judgement

Appointments and documents may be automated first. Diagnosis, medication, and treatment require clinical review. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Health data is accessed by purpose, consent, and the minimum-necessary principle. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Bed, examination, medication, and discharge agents may recommend. The workflow checks the rules. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Beyond accuracy, monitor missed detections, false alarms, downtime, leakage, and group differences. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Smart environment and enforceable governance

Problem and judgement

Air, water, noise, waste, energy, and ecological data need quality tags. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Agents may explain anomalies and schedule inspections, but retain sensor-calibration evidence. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Sensing and imagery only trigger verification. They must not form a penalty directly. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Assess by pollution improvement, resource savings, inspection efficiency, and public transparency. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Urban resilience and emergency coordination

Problem and judgement

Agents integrate forecasts, facility capacity, population, and historical events into scenarios. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Formal warnings are still issued by the authorised authority. Summaries must link back to the original message. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Assume cloud, model, and network failure. Retain offline, dedicated-network, and human takeover. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

After an incident, compare forecast, decision, execution, and outcome, and correct cross-department bottlenecks. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Open data and agent innovation

Problem and judgement

Open data has reached more than 5,700 datasets and more than 2,500 providers. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Downloads rose from about 5 billion in 2019 to more than 80 billion in 2025. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Machine-readable data needs field definitions, update frequency, licence, and stable identifiers. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Assess re-identification, misreading, and fraud risk when multiple datasets are combined. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Multi-cloud, hybrid cloud, and brand selection

Problem and judgement

Tencent Cloud, AWS, Google, ByteDance, and local clouds should be compared against governance objectives. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

First set residency, compliance, latency, capability, talent, cost, and exit requirements. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Keep sensitive transactions in a controlled domain. Elastic inference may use a compliant public cloud. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Accept valuable differentiation, but control lock-in with export, substitution, and migration drills. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Vendor assessment and procurement acceptance

Problem and judgement

Use de-identified real benchmarks to test Chinese, Cantonese, policy, long documents, and tool calls. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Check SLO, RTO, RPO, capacity, certifications, subcontractors, and data policy. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Pay in stages against data preparation, pilot, security, performance, disaster recovery, and operating results. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Contracts retain audit, version notice, cost breakdown, export, and termination assistance. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Organisational structure aligned with technical architecture

Problem and judgement

The central team provides identity, exchange, models, tools, evaluation, security, and audit. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Departments retain policy interpretation, process accountability, data quality, and outcome metrics. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

Projects include business, product, data, cybersecurity, legal, procurement, and operations roles. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Training covers process, contracts, risk, evaluation, incidents, and vendor management. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Red team and chaos engineering in practice

Problem and judgement

Attack with malicious attachments, web injection, out-of-scope queries, and poisoned summaries. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

Inject model delay, message duplication, stale knowledge, and permission-service interruption. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

People drill incorrect approval, stolen accounts, review backlog, and vendor loss of contact. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Issues need an accountable owner, a deadline, a retest, and risk acceptance — not a report alone. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


Five-layer reference architecture

Problem and judgement

The data layer manages master data, catalogue, lineage, quality, consent, and exchange. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

The model layer provides routing, security, and evaluation. The agent layer defines roles and tools. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

The workflow controls sequence, compensation, and human nodes. The runtime layer provides elasticity and observability. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Endpoints are designed for citizens, enterprises, and case officers, and support accessibility, weak networks, and humans. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


A 90-day and one-year implementation path

Problem and judgement

In the first 30 days, select the scenario, decompose the process, inventory data, define roles, and build the benchmark set. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Architecture and method

In the next 30 days, build orchestration, schema, identity, citations, tracing, and offline evaluation. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Governance and control

In the last 30 days, run a controlled pilot, measure effect, and rehearse interruption and takeover. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Implementation and acceptance

Within one year, institutionalise governance, a shared foundation, multi-vendor benchmarks, disaster recovery, and exit. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.


The leader's final decision framework

Value first, then technology

Confirm the bottleneck, baseline, beneficiaries, and alternatives. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.

Accountability first, then autonomy

Every decision, data set, and tool has a named accountable owner. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.

Foundation first, then expansion

Identity, contracts, evaluation, observability, and audit come first. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.

Least necessary intelligence

The system must be auditable, degradable, replaceable, and able to exit. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.