Vertex Macro | Financial Cloud Cloud · AWS Re:cap
AWS Re:cap 05: A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
A deep briefing for government technology advisers, digital-government officers, data-governance leads, cybersecurity officers, enterprise architects, and large-programme decision-makers.
The agent turning point in the government software ecosystem
Public value
Judge success by processing time, one-stop completion rate, and frontline burden. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Engineering essence
A multi-agent system is a distributed system with probabilistic behaviour. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Accountability red line
Statutory decisions must not be left to an agent vote. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Starting method
First decompose the process, then mark rule, reasoning, and human nodes. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
The correct division of labour between agents and microservices
Problem and judgement
Microservices carry transactions, ledgers, identity, and reproducible rules. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Agents handle semantic understanding, evidence collection, and draft recommendations. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
The core system remains the authoritative data state. The model must not overwrite it directly. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Every model call must prove more value than rules or search. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Fitness criteria for multi-agent systems
Problem and judgement
Multiple professions, data domains, and permission sets, plus a high exception rate, are positive signals. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Use traditional automation for fixed processes, a single data source, and low reasoning value. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Score by task variability, cross-domain dependency, exception rate, and reasoning gain. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Before a PoC, establish baselines for time, error, rejection, labour hours, and appeals. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Five roles and isolation of duties
Problem and judgement
The coordinator splits tasks and maintains state, but does not judge policy legality. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Expert agents provide policy, data, calculation, or text output. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Gatekeeper and audit agents check permission, purpose, and completeness of evidence. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
The human accountable owner holds approval, veto, and reversal rights for high-risk decisions. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Centralised, hierarchical, and decentralised collaboration
Problem and judgement
Centralised orchestration is best for formal government processes and consistent audit. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Hierarchical mode suits large cross-domain tasks, but must prevent summary distortion. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Decentralised negotiation is suitable only for low-risk exploration. It must not replace statutory accountability. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Every mode must limit rounds, cost, timeout, and conflict escalation. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Task contracts and structured schemas
Problem and judgement
Inputs include purpose, data scope, tools, deadline, cost, and risk. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Outputs include conclusion, citations, confidence, open items, and next steps. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Amounts, dates, eligibility, and identity must pass type and range validation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Contracts must have versions, compatibility policy, tests, and change approval. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Shared state and agent memory
Problem and judgement
Process state, business state, and semantic memory must be governed separately. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The core business system is the authoritative source. Vector memory is only an aid. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Long-running processes use event sourcing, idempotency keys, and Saga compensation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Each class of state must have create, use, archive, and delete cycles. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Security boundaries for tool calls
Problem and judgement
Split permissions across search, read, write, payment, notification, and delete. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Before execution, verify identity, delegation, purpose, data classification, and impact scope. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
After execution, verify the result, obtain a receipt, and write tamper-evident evidence. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Red-team test prompt injection, malicious files, out-of-scope parameters, and tool confusion. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Agent identity and zero trust
Problem and judgement
People, services, and agents use distinct and traceable identities. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Short-lived credentials bind task, purpose, data scope, and validity period. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Authorisation includes role, sensitivity, region, time, and risk conditions. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
When an agent is retired, revoke keys, tools, memory, and schedules in the same step. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
The full lifecycle of a digital worker
Problem and judgement
The job description lists who is served, what may be done, what must not be done, and the escalation path. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The permission card separates readable, writable, recommendable, and approvable. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Performance is judged on correctness, citations, refusal, appeals, and rework together. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
On decommission, hand over open cases, revoke rights, archive evidence, and update the catalogue. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
The real design of human review
Problem and judgement
Place humans before conflicts, irreversible tools, and low-confidence outputs. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The review interface shows raw data, policy, disagreements, and expected impact. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
The reviewer may veto, amend, request further evidence, or transfer. Reasons must be recorded. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
High risk is fully reviewed. Medium and low risk use thresholds, sampling, and consistency monitoring. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Turn agent conflict into a governance signal
Problem and judgement
Conflict often comes from version, date, data completeness, and jurisdiction differences. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The conflict pack lists the points of dispute, each party's basis, missing data, and impact. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
First check data and versions, then apply policy priority, and finally adjudicate by a human. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Regular analysis of conflict can reveal ambiguous policy and inconsistent cross-department standards. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Reliability, degradation, and compensation
Problem and judgement
Set timeout, backoff, and maximum retries for every agent and tool. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
When the model or an external cloud fails, switch to rules, search, a human, or deferral. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Unhandleable events go to a dead-letter queue with context and a handling deadline. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Chaos drills verify RTO, RPO, alerts, takeover, and recovery. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Multi-agent cost engineering
Problem and judgement
Split cost into inference, retrieval, API, compute, storage, labour, and operations. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Each task sets limits on rounds, context, tool calls, and total cost. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Use rules or a small model for classification and extraction. Use a high-capacity model only for complex reasoning. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Contracts include unit cost, peak throughput, and three-year total cost of ownership. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
End-to-end observability and audit
Problem and judgement
A case trace identifier links request, model, knowledge, tools, humans, and result. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Layer technical, agent, business, and governance metrics. Do not look only at latency. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Logs mask sensitive data and separate operations access from case access. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Review failures, cost, knowledge expiry, reversals, and vendor availability every day. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Evaluation is more than answer correctness
Problem and judgement
The benchmark set covers normal, boundary, missing-data, conflict, and adversarial inputs. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Evaluate task success, field correctness, citations, hallucination, and refusal. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Also test leakage, privilege escalation, fairness, latency, availability, and disaster recovery. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Re-run regression after any change to the model, knowledge, prompt, or tools. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Stage gates from PoC to production
Problem and judgement
In exploration, use synthetic, anonymised, or public data and set stop conditions. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
A controlled pilot limits business, population, and time window. Every result is checked by a human. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Before production, complete classified cybersecurity protection grading, commercial cryptography application security assessment, data classification, load testing, and disaster recovery. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Expand with canary release and rollback. Repeat risk assessment whenever an agent or tool is added. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Success case: Hong Kong digital government shared foundation
Problem and judgement
By the end of 2025, more than one hundred digital government and smart city measures will have been advanced. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
A next-generation government cloud, big data, shared blockchain, and a common chat service support departments. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
iAM Smart has more than four million users and covers more than 1,300 services and forms. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
The success lesson is to unify identity, data exchange, and security first, then extend intelligence. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Success case: consent-based data exchange gateway
Problem and judgement
With the citizen's consent, deliver authoritative-source data to a designated electronic service. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
About two million data exchanges a month reduce repeat submission and manual verification. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
A data agent may request only the fields required by policy. It must not extend the consented purpose. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Each exchange retains evidence of source, time, purpose, recipient, and withdrawal handling. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Success case: full digitalisation and enterprise identity
Problem and judgement
Electronic payment, electronic submission, and electronic approval documents have been fully digitalised. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
CorpID, the enterprise digital identity, is expected to launch by the end of 2026 and expand services in stages. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Core capabilities include enterprise verification, digital signing, pre-fill, and a document wallet. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
An agent may prepare materials, but formal signing and high-risk submission remain with the authorised person. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Success case: multi-vendor AI+ public services
Problem and judgement
The capability catalogue covers customer service, meetings, documents, writing, processes, creativity, and analysis. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Compare multiple vendors on the same field, and choose by Chinese capability, data, latency, cost, and exit. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Technical matching must include a real process, a baseline, and data constraints — not a demonstration alone. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Different models may connect through standard tools, but identity, logs, and evaluation must be unified. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Greater Bay Area cross-border data and rules
Problem and judgement
A memorandum of cooperation on promoting cross-border data flow in the Greater Bay Area was signed in 2023. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Standard-contract facilitation measures have been extended to all industries since November 2024. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Before an agent transmits, identify data type, origin, destination, purpose, and duration. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Pilot first with low-sensitivity services that have a clear purpose and a clear authoritative source. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Digital-first Northern Metropolis
Problem and judgement
The next five years plan more than 70,000 housing units and one million square metres of economic floor space. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The university town, the Hong Kong Park of the Loop, and San Tin Technopole connect research and industrial conversion. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Spatial, building, transport, energy, and environmental data form the city-operations foundation. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Identity, data standards, communications, edge, disaster recovery, and cybersecurity should enter early planning. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Enterprise benefits without application: case panorama
Problem and judgement
Policy, data, compliance, calculation, notification, and audit agents divide professional labour. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The coordinator maintains the process but cannot approve subsidies or change enterprise data. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Agent conflicts are adjudicated by the business accountable owner. Payment runs through the core finance system. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Measure success by processing time, hit rate, error rate, and enterprise burden. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Enterprise-benefit case: policy knowledge engineering
Problem and judgement
Split provisions into subject, time, region, exclusion, evidence, formula, and discretion. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Each condition links back to the original text, issuing authority, version, and effective date. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Give clear thresholds to rules. Use a model only for semantic extraction. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Revision, suspension, and expiry trigger a process. Old versions are retained to reproduce historical cases. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Enterprise-benefit case: data verification
Problem and judgement
Enterprise registration, tax, employees, licences, and subsidies each have an authoritative source. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The agent requests the minimum fields required by the condition. Boolean or range answers can reduce disclosure. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Missing, delayed, or contradictory data is marked as undeterminable. Do not guess eligibility or refusal. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Verification results record source, time, transformation rule, and purpose of use. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Enterprise-benefit case: isolation of calculation and payment
Problem and judgement
Subsidy amounts are calculated by a versioned rule service. The model explains; it does not post to the ledger. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Candidate results are checked by an independent rule or a second calculation instance. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
The agent only generates a payment draft. The core finance system validates the budget and the approval chain. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Reconcile approvals, instructions, and bank results daily. Recovery and correction follow a compensation process. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Enterprise-benefit case: notification and appeal
Problem and judgement
The notice explains conditions, data, calculation, missing documents, and the ultimately accountable authority. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Website, SMS, email, hotline, and counter use a consistent case status. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
The enterprise may view the basis, supply further evidence, request human review, and lodge an appeal. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Track comprehension, number of supplementary submissions, cycle time, reversals, and satisfaction. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Agent boundaries in smart mobility
Problem and judgement
Agents integrate real-time traffic, parking, arrival, road-sensor, and incident data. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Signals, charging, and enforcement are controlled by deterministic systems, with human takeover retained. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
During major events, departments share incident state but keep their own statutory powers. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Metrics include journey reliability, response time, false alerts, availability, and complaints. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
High-risk layering in smart healthcare
Problem and judgement
Appointments and documents may be automated first. Diagnosis, medication, and treatment require clinical review. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Health data is accessed by purpose, consent, and the minimum-necessary principle. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Bed, examination, medication, and discharge agents may recommend. The workflow checks the rules. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Beyond accuracy, monitor missed detections, false alarms, downtime, leakage, and group differences. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Smart environment and enforceable governance
Problem and judgement
Air, water, noise, waste, energy, and ecological data need quality tags. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Agents may explain anomalies and schedule inspections, but retain sensor-calibration evidence. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Sensing and imagery only trigger verification. They must not form a penalty directly. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Assess by pollution improvement, resource savings, inspection efficiency, and public transparency. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Urban resilience and emergency coordination
Problem and judgement
Agents integrate forecasts, facility capacity, population, and historical events into scenarios. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Formal warnings are still issued by the authorised authority. Summaries must link back to the original message. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Assume cloud, model, and network failure. Retain offline, dedicated-network, and human takeover. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
After an incident, compare forecast, decision, execution, and outcome, and correct cross-department bottlenecks. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Open data and agent innovation
Problem and judgement
Open data has reached more than 5,700 datasets and more than 2,500 providers. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Downloads rose from about 5 billion in 2019 to more than 80 billion in 2025. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Machine-readable data needs field definitions, update frequency, licence, and stable identifiers. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Assess re-identification, misreading, and fraud risk when multiple datasets are combined. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Multi-cloud, hybrid cloud, and brand selection
Problem and judgement
Tencent Cloud, AWS, Google, ByteDance, and local clouds should be compared against governance objectives. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
First set residency, compliance, latency, capability, talent, cost, and exit requirements. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Keep sensitive transactions in a controlled domain. Elastic inference may use a compliant public cloud. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Accept valuable differentiation, but control lock-in with export, substitution, and migration drills. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Vendor assessment and procurement acceptance
Problem and judgement
Use de-identified real benchmarks to test Chinese, Cantonese, policy, long documents, and tool calls. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Check SLO, RTO, RPO, capacity, certifications, subcontractors, and data policy. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Pay in stages against data preparation, pilot, security, performance, disaster recovery, and operating results. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Contracts retain audit, version notice, cost breakdown, export, and termination assistance. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Organisational structure aligned with technical architecture
Problem and judgement
The central team provides identity, exchange, models, tools, evaluation, security, and audit. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Departments retain policy interpretation, process accountability, data quality, and outcome metrics. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
Projects include business, product, data, cybersecurity, legal, procurement, and operations roles. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Training covers process, contracts, risk, evaluation, incidents, and vendor management. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Red team and chaos engineering in practice
Problem and judgement
Attack with malicious attachments, web injection, out-of-scope queries, and poisoned summaries. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
Inject model delay, message duplication, stale knowledge, and permission-service interruption. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
People drill incorrect approval, stolen accounts, review backlog, and vendor loss of contact. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Issues need an accountable owner, a deadline, a retest, and risk acceptance — not a report alone. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
Five-layer reference architecture
Problem and judgement
The data layer manages master data, catalogue, lineage, quality, consent, and exchange. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
The model layer provides routing, security, and evaluation. The agent layer defines roles and tools. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
The workflow controls sequence, compensation, and human nodes. The runtime layer provides elasticity and observability. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Endpoints are designed for citizens, enterprises, and case officers, and support accessibility, weak networks, and humans. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
A 90-day and one-year implementation path
Problem and judgement
In the first 30 days, select the scenario, decompose the process, inventory data, define roles, and build the benchmark set. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Architecture and method
In the next 30 days, build orchestration, schema, identity, citations, tracing, and offline evaluation. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Governance and control
In the last 30 days, run a controlled pilot, measure effect, and rehearse interruption and takeover. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Implementation and acceptance
Within one year, institutionalise governance, a shared foundation, multi-vendor benchmarks, disaster recovery, and exit. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.
The leader's final decision framework
Value first, then technology
Confirm the bottleneck, baseline, beneficiaries, and alternatives. The core of this design is to place technical capability back inside public accountability and business outcomes. Implementation must name an accountable owner, inputs and outputs, exception paths, and verifiable evidence. Maturity must not be judged by demonstration effect alone.
Accountability first, then autonomy
Every decision, data set, and tool has a named accountable owner. Implementation must distinguish deterministic process from probabilistic reasoning. The former controls transactions and permissions. The latter handles semantics and incomplete data. The two are connected by a structured contract, so they can be tested, replayed, and replaced.
Foundation first, then expansion
Identity, contracts, evaluation, observability, and audit come first. Risk control must not remain a principle. It must become identity, authorisation, versioning, logs, human veto, and an exit mechanism. Every high-impact action must leave a complete, auditable evidence chain before and after execution.
Least necessary intelligence
The system must be auditable, degradable, replaceable, and able to exit. Evaluation must check public value, model quality, engineering reliability, governance effect, and unit cost at the same time. If the baseline has not improved, or risk exceeds the threshold, degrade, correct, or stop — do not expand because investment has already been made.