← Financial Cloud Cloud Cloud Club · AWS Re:cap

极点宏观|Financial Cloud Cloud · AWS Re:cap

AWS Re:cap 01: The AI Era: The Boundary Between Development and Design Is Disappearing

讲者: 政务数据

场次: 01

场次
Summit Dev Lounge2026 Re:cap
01 把架构写成 Steering,引导 AI Agent
Summit Dev Lounge2026 Re:cap
02 Agent Harness 才是真正的工程护城河
Summit Dev Lounge2026 Re:cap
03 用白话询问可观测性数据
Summit Dev Lounge2026 Re:cap
04 以 Bedrock AgentCore 打造 Serverless AR 游戏
Summit Dev Lounge2026 Re:cap
05 AgentCore 上的多 Agent 量化回测
Summit Dev Lounge2026 Re:cap
06 三分钟用 Kiro 把博客变成幻灯片
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 使用 Terraform 实现 AWS 合规
AWS Community Day Hong Kong 2025 Re:cap
03 从初学者到构建者:一段精彩的 AWS 云之旅
AWS Community Day Hong Kong 2025 Re:cap
04 以团队为先:使用 Laravel 与 Bref 开展无服务器工程
AWS Community Day Hong Kong 2025 Re:cap
05 活动开幕式
AWS Community Day Hong Kong 2025 Re:cap
06 智能体到智能体:在 AWS 上构建可互操作的 AI
AWS Community Day Hong Kong 2025 Re:cap
07 利用另一类遥测数据,借助 AI 智能体更快改进
AWS Community Day Hong Kong 2025 Re:cap
08 告别氛围编程:使用 Kiro 进行规格驱动开发
AWS Community Day Hong Kong 2025 Re:cap
09 使用 MCP 与 AI 智能体进行自动化测试
AWS Community Day Hong Kong 2025 Re:cap
10 使用机器学习方法实现电信安全现代化
AWS Community Day Hong Kong 2025 Re:cap
11 重新思考生成式 AI 智能体:RAG 与 MCP
AWS Community Day Hong Kong 2025 Re:cap
12 使用 TAK 和 AWS 开展灾难与应急响应
AWS Community Day Hong Kong 2025 Re:cap
13 从测试视角重新思考无服务器应用程序工作流
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 基于 AWS 的 AI 驱动全球纯 Alpha 宏观交易:重塑风险调整后资产收益
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 现代交易生命周期:从交易到结算
FSI Recap
02 Goldman Sachs:通过 Fast Track 加速应用程序上云 - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group:在 AWS 上构建高效的日志管理解决方案
FSI Recap
04 FSI Meetup 2025 年第四季度 - Brex 数据库灾难恢复
FSI Recap
05 FSI Meetup 2025 Q4 - Graviton 迁移成功案例
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel 现代数据平台
FSI Recap
07 FSI Meetup 2025 Q4 - PayPal 金融交易数据对账系统
FSI Recap
08 FSI Meetup 2025 Q4 - 规模化提升韧性
FSI Recap
09 最大限度提高 AI 推理成本效益:战略性采用 AWS GPU 实例
FSI Recap
10 高级智能体 AI 设计模式
FSI Recap
11 在 AWS 上构建全新的现代化应用
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent 回顾 (IND3312)
AWS re:Invent 2025
02 利用 AI 和 AWS 构建未来交易平台
AWS re:Invent 2025
03 交易创新:Jefferies 基于 Amazon Bedrock 构建的 AI 助手 (IND3315)
AWS re:Invent 2025
04 FSI 如何通过 Agentic AI (GBL302) 彻底改变 HFT 分析
AWS re:Invent 2025
05 使用 Amazon Time Sync 改进分布式系统(采用 Nasdaq)
AWS re:Invent 2025
06 Amazon Aurora HA 和 DR 全球弹性设计模式 (DAT442)
AWS re:Invent 2025
07 构建智能体式 AI:Amazon Nova Act 与 Strands Agents 实践 (DEV327)
AWS re:Invent 2025
08 深入探讨 Amazon Aurora 及其创新 (DAT441)
AWS re:Invent 2025
09 深入探讨 Amazon S3(STG407)
AWS re:Invent 2025
10 Nasdaq:为全球金融服务构建弹性基础设施 (HMC327)
AWS re:Invent 2025
11 AWS Lambda 新功能 (CNS376)
AWS re:Invent 2025
12 使用 Kiro 进行规范驱动开发 (DEV314)
AWS re:Invent 2025
13 Amazon 的 FinOps:全球电商巨头的云成本管理经验 (AMZ308)
AWS re:Invent 2025
14 AWS 上交易平台的 Tick-to-Trade 延迟
AWS re:Invent 2025
政务数据
01 The AI Era: The Boundary Between Development and Design Is Disappearing
政务数据
02 端侧多模态 AI 与智慧城市实践
政务数据
03 大模型能力评测与 AI 项目落地方法论
政务数据
04 基于云代理的政府开发全链路受控自动化
政务数据
05 从多智能体看 Agent 时代软件新生态
政务数据
06 AI 驱动的宏观量化研究与智慧治理
政务数据
07 公共数据授权运营与智慧政务实践
政务数据
08 数据资产化落地实践:确权合规、工程治理与数字政府案例
政务数据
09 AI技术赋能心理健康公益:可信平台的治理、架构与实践
政务数据
Amarathon 2025 回顾
01 开发者的智能体架构设计路线图
Donnie Prakoso
02 Amazon Bedrock 数据自动化
Hafiz Syed Ashir Hassan
03 AgentCore 上的多智能体
Tan Xin
04 实践中构建智能体式 AI:Nova Act 与 Strands Agents
Haowen Huang
04 使用规格驱动开发,通过 Kiro 加速迁移项目
Sanchit Dilip Jain
06 从「匹配」到「理解」:由 AgentCore Memory 驱动的个性化 AI 搜索实践
Liu Cao
07 从观察到优化:从 LLM 可观测性迈向 AIOps,将实时洞察转化为智慧自动化
Jimmy Soh
08 部署 TEAM 并打造最佳工程团队
Yuji Oshima
09 五年来所谓无服务器数据库带来的五个惨痛教训
Renato Losio
14 如果 AI 替我工作会怎样:Q Developer CLI 与 Kiro 如何改变我的日常工作
Miguel Angel Muñoz
16 兼顾速度与警觉:Amazon Bedrock Agent 开发的安全要点
Brian Tarbox
26 在单张 H100 上运行 OSS LLM:更智能、更便宜、更快速
Adit Modi Adit Modi
28 现代统一元数据架构:打破数据孤岛的新方法
Shaofeng Shi
29 无服务器 MediaOps:使用 Amazon Web Services 上的 AI 自动化视频工作流
Luis Valdivia
30 通过大规模性能测试构建兼具效率与可靠性的架构
Luis Guirigay
31 通过开源连接世界:技术、社区与全球开发者关系的实践历程
Richard Lin
33 构建流式 Iceberg 表以进行实时物流分析
Fahad Shah
34 加速大规模机器人策略训练:基于 Kiro、Trainium 和 EKS 的自动化闭环架构
Junjie Tang
35 通过规格驱动开发,从 Vibe 走向可行方案
Ricardo Sueiras
36 让云成本分析更智能:使用 Strands 和 AgentCore 构建 FinOps 智能体
Xiaofei Li
37 使用 CNCF Kagent、K8sGPT 和 Nova Sonic 转型 K8s 对话式智能体 AIOps
Shaoyi Li

Speech manuscript for government technology advisers, digital government leads, data and security officers, digital-government architects, and large SOE technology decision-makers.


A New Product-Engineering Boundary in the AI Era: From Role Handoffs to Shared Accountability

The change is not “who replaces whom”

Generative AI connects policy interpretation, user research, service blueprints, interface prototypes, code, testing, documentation, and operations analysis into one computable chain. Work that once moved sequentially from role to role can now advance in parallel around the same structured task. The real meaning of a fading boundary is that the work object, the evidence, and the feedback are shared—not that designers, engineers, or domain specialists lose their value.

Distinct constraints in government settings

Public services cannot pursue demonstration speed alone. Model error, data leakage, service interruption, vendor lock-in, and unexplainable decisions can all become governance risks. Therefore every AI-assisted delivery must also satisfy legality and compliance, accessibility, auditability, graceful degradation, maintainability, and a right of appeal. Efficiency is only one outcome. Public accountability is the system boundary.

A judgement from four decades of engineering practice

The most common failure in large informatisation programmes is not that the technology cannot be built. It is that requirements are mistranslated, accountability is not assigned to named people, and acceptance looks only at a feature list. AI will amplify the right direction and will also harden the wrong direction faster. Leaders should first establish a shared language, accountability boundaries, and quality gates, and only then scale generation capacity.

What the audience can take away immediately

Write every task as eight fields: public purpose, service users, legal basis, input data, permitted actions, human approval, exception fallback, and acceptance evidence. This minimum task contract can be used at once by business, design, development, testing, cybersecurity, and procurement, so a cross-functional team works from the same facts from day one.


Why Traditional Delivery Distorts Between Requirements, Design, and Development

Three translations create three losses

Business writes policy as requirements; design turns requirements into flows and pages; development then turns pages into APIs and rules. Each translation can omit exception conditions, scope of application, and the original legal basis. When disputes arise late in a project, teams often have only meeting minutes to search and cannot quickly answer which policy a rule came from, who interpreted it, and when it took effect.

Prototypes can look consistent while the semantics differ

The same “Submit” button may mean “application completed” to business, “draft saved” to engineering, and “no legally effective service” to counsel. If the team synchronises only visual mock-ups and not the state machine, data contract, and legal effect, the more polished the interface, the harder hidden inconsistency is to find.

The new approach AI makes possible

Use a structured knowledge base and traceable generation to map policy clauses to service steps, data fields, API constraints, test cases, and citizen guidance. Any change to an artefact can return to the same source, instead of leaving each role to maintain its own interpretation. AI accelerates mapping and conflict discovery; specialists adjudicate meaning.

A hands-on inspection method

Select one high-frequency service and randomly sample ten page fields. For each field answer: what is the source clause, who maintains it, how errors are handled, whether it enters a downstream system, and whether the user must give explicit consent. If any question cannot be answered within ten minutes, the team lacks a unified semantic layer and should govern first rather than add more features.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Shared Work Objects: The First Principle for Removing Boundaries

From document handoff to model-based collaboration

Shared work objects may be service blueprints, policy rule libraries, domain models, design tokens, API contracts, acceptance scenarios, and runtime evidence. They must be versionable, comparable, and citable. Discussion no longer stops at “what I understood,” but locates a specific version, a specific rule, and a specific state.

One set of facts, multiple views

Leaders see public value and risk; business sees process and policy; design sees journeys and accessibility; development sees APIs and state; testing sees boundary conditions; operations sees SLOs and alerts. Roles see different views, but the underlying facts come from the same repository, so copies do not gradually diverge.

Capabilities the platform must provide

Version control, change-impact analysis, approval records, automated generation, diff review, citation back-links, and permission isolation are all required. After a policy change, the platform should flag affected pages, APIs, tests, guidance, and model evaluation sets, so change moves from manual notification to computable propagation.

Implementation experience

Do not start by building a vast unified model. First define twenty to thirty core objects and states around one cross-department service, and make design mock-ups, APIs, and tests all reference them. Expand only after two iterations have validated the approach. Semantic consistency in a small scope is more valuable than noun unification at a large scale.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Mission-Based Product Squads: How Organisational Boundaries Are Redrawn

Form teams around outcomes

Traditional projects queue by specialist department. A mission-based squad forms around an outcome such as “shorten the time to start a business” or “raise first-time completion.” Core members include the business owner, service design, front-end and back-end engineering, data and AI, testing, cybersecurity, legal, and operations representatives. Members are accountable for the same outcome, not only for their specialist artefacts.

A squad is not a boundary-free zone

Cross-functional work still needs clear decision rights. Business owns policy interpretation; the product owner ranks value; the architect holds the technical boundary; cybersecurity holds a veto on high-risk items; operations decides supportability. AI may propose options and generate drafts, but it must not replace statutory approval, risk acceptance, or production-release authority.

Platform teams and governance teams

Product squads pursue service speed; the platform team provides reusable identity, model gateway, logging, evaluation, and release pipelines; the governance team sets red lines, samples evidence, and handles exceptions. The three create speed, reuse, and checks and balances. Do not pile every duty onto a central platform, and do not let every squad rebuild the same wheels.

Team start-up checklist

Before the first iteration, write together the mission boundary, success metrics, key risks, data scope, human-review points, go-live authority, and fallback plan. Each item has one final owner, plus roles that must be consulted and roles that must be informed. This avoids the classic trap in which “everyone is involved, so no one is accountable.”

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


The Human–AI Accountability Boundary: What AI May Propose, Execute, and Must Not Do

A three-tier authorisation model

Low-risk tasks may be executed automatically by AI, such as format conversion, duplicate-field checks, and test-data generation. Medium-risk tasks allow AI to propose, subject to human confirmation, such as rewriting citizen guidance, recommending a process, and merging code. High-risk tasks allow retrieval and prompting only, never an automatic decision—for example eligibility determination, sanctions, fund disbursement, and cross-domain movement of sensitive data.

Authorisation depends on context

The same capability has different risk under different data, subjects, and impact. Summarising published policy is low risk; summarising unpublished case files may involve privacy. Auto-replying to general enquiries may be acceptable; auto-replying with an appeal decision is not. The accountability boundary must express the task, the data, the impact, reversibility, and legal consequence together.

Evidence that must be retained

Record inputs, knowledge sources, model and prompt versions, tool calls, outputs, confidence or risk labels, human edits, the approver, and the final action. Evidence is not collected for blame. It exists for retrospectives, appeals, model improvement, and vendor switching. The more automation proceeds without evidence, the harder the organisation is to control.

A practical judgement test

If an error can be discovered within minutes and rolled back automatically, the automation tier may rise. If an error would affect individual rights, is hard to detect, or cannot be remedied, lower the authorisation tier. Do not decide automation from model accuracy alone. Include reversibility, number of people affected, and the probability of human discovery.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Policy as Code: Make Rules Versionable, Testable, and Auditable

Do not bury policy permanently in prompts

Prompts are suitable for expressing how a task is done, not for serving as the only rule store. Eligibility thresholds, deadlines, document requirements, and exceptions should be maintained as structured rules, linked to original clauses, effective dates, applicable territories, and the interpreting authority. Models may invoke rules; they must not quietly rewrite them.

Three layers of rule testing

The first layer is input–output tests for a single rule. The second is conflict and priority tests for combinations of rules. The third is end-to-end tests of real service journeys. Every policy version change re-runs automatically and lists result changes. Legal, business, and development can then see actual impact before go-live.

Separate explanation from adjudication

The system may explain to the public “why this document is required,” but formal adjudication must rest on deterministic rules and an authorised process. For discretionary factors that cannot be structured, the model may only organise evidence, flag gaps, and generate a review summary. The final judgement is made by an authorised officer and recorded with reasons.

A migration method

First pick high-frequency, low-controversy, clearly specified services. Extract conditions scattered across code and documents, and assign rule identifiers and test samples. Do not try to cover every policy at once. Prioritise rules that change often, are re-implemented across systems, or attract many complaints. The return is most visible there.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


From Policy Text to Service Blueprint: AI-Assisted Requirements Discovery

Identify service users and events first

Policy is usually written by administrative duty. The public acts by life events or business events. AI can extract actors, triggering events, required conditions, processing bodies, and time limits from documents, then business specialists verify. The shift is from “what the department provides” to “what the user must complete.”

A blueprint must include front stage and back stage

The front stage records user actions, touchpoints, waiting, and emotion. The back stage records departmental processing, data exchange, decision rules, exception paths, and evidence. Page flows alone conceal the real bottlenecks. Many so-called experience problems in fact come from repeated back-office checks, permission boundaries, and cross-department waiting.

AI is suited to discovering conflict, not adjudicating it

A model can flag inconsistent descriptions of deadlines, documents, or names across files, and generate a confirmation list. Legal effect, scope of application, and priority must be decided by legal counsel and the competent authority. Making uncertainty explicit is safer than letting a model choose a plausible-looking answer.

Workshop outputs

Have business, counter staff, legal, data, and technical staff review one blueprint together. Mark every wait point, repeated-submission point, manual-transcription point, and high-risk judgement point. Each pain point must map to a testable hypothesis, such as “first-time completion rises after duplicate proofs are removed,” not a vague “improve the experience.”

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


A Prototype Is Not a Product: The Distance from Clickable to Production-Ready

A prototype validates understanding

Clickable pages are used to validate information architecture, task sequence, language readability, and key interactions early. They do not prove concurrency, data consistency, disaster recovery, security, accessibility, or operating cost. If leaders treat a successful demo as project completion, they postpone a large volume of risk to the most expensive stage.

The correct use of generative prototypes

AI can quickly generate multiple process options, content versions, and exception states to help the team compare. Each option must bind a hypothesis and a test audience—for example whether older users can complete independently, whether a corporate agent can switch identity, and whether work can continue after a network outage. The faster the prototype, the faster the validation cadence must also be.

The gate before development

At minimum complete usability tests of the core journey, confirmation of policy rules, definition of data fields, API responsibilities, an initial privacy-impact judgement, an accessibility plan, and exception handling. Also confirm which elements come from the design system and which are one-off exploration, so temporary prototype code is not carried into production.

Lessons learned

The most dangerous artefact is not a low-fidelity prototype, but a high-fidelity prototype that looks like a finished product. It creates a false sense of completion. Reviews should explicitly mark content as “validated,” “to be validated,” or “out of scope,” and require that any go-live commitment rest on engineering and governance evidence.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Design System 2.0: From Component Library to Governance Vehicle

Components are only the surface

A mature design system should include design tokens, interaction patterns, content standards, accessibility requirements, front-end components, usage constraints, version policy, and a migration guide. For government services it must also include public-service patterns such as identity switching, authorised agency, electronic signing, document upload, progress enquiry, and appeal.

AI generation must be constrained by the system

If a model may invent colours, components, and copy at will, speed is exchanged for chaos beyond mere sameness. A better approach is for AI to compose from approved components, tokens, and language patterns, then automatically detect unauthorised styles, contrast issues, keyboard operation, and responsive problems after generation.

Version governance

Component upgrades must state compatibility, deprecation windows, impact scope, and rollback method. Business pages must not remain locked on old versions indefinitely, nor may the platform team force every system to upgrade at once. Use a clear support window and migration tools so departments with different maintenance capacity can update on a plan.

Measuring the value of the system

Do not count components alone. Observe reduced duplicate development, fewer accessibility defects, lower cross-service learning cost, shorter change-propagation time, and brand consistency. If there are many components but projects still routinely bypass the system, it is not solving real tasks. Return to usage scenarios rather than keep expanding the catalogue.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Accessibility Is a Design Input, Not an Acceptance Attachment

Public services must be inclusive by default

Users may use a screen reader, keyboard, magnification, a slow network, or an old device, and they may be under stress, fatigue, or limited language comprehension. Accessibility is not only a need of particular groups. It improves understandability, operability, and error tolerance for everyone.

AI can help but cannot replace human verification

A model can check alternative text, heading hierarchy, field labels, contrast, and focus order, and can generate copy at different reading levels. Complex forms, identity verification, and error recovery still require testing with real users. Automated checks find rule-based issues; they cannot prove that a task can actually be completed.

Design decisions must leave a trail

Record why an interaction was chosen, which assistive technologies were tested, what barriers were found, how they were fixed, and which issues remain open. Procurement acceptance should also require an accessibility statement and a defect-remediation commitment, so teams do not discover after go-live that a vendor component cannot be remediated.

A field method that works

Have the team complete one application using only the keyboard, then magnify the browser to 200 percent, disable images, and simulate network delay. Each time work cannot continue, record the component, the flow, and the owner. Low-cost drills of this kind often build shared understanding faster than reading a standard.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Design the API Contract and the Interface Together: Front End and Back End No Longer Queue

Contract before implementation

Define core objects, states, error codes, permissions, and idempotency requirements at the prototype stage. The front end can validate flows against a mock service; the back end can develop to the same contract; testing can write cases early. Design, development, and testing then work in parallel around an executable agreement and reduce surprises at final integration.

Error states are part of the experience

API timeouts, duplicate submission, expired data, insufficient permission, and downstream unavailability must all have clear user feedback and a recovery path. Do not treat error codes as purely technical detail. Every error must state whether retry is allowed, whether completed fields are retained, when to escalate to a human, and how to enquire about processing status.

The boundary of AI-generated code

A model may generate a client, a server skeleton, and tests from a contract, but the output must pass static scanning, dependency checks, code review, and contract tests. If the contract itself is vague, AI will only generate mutually incompatible implementations faster. Stabilise semantics first, then accelerate coding.

A field check

Work backwards from one critical journey through every API. Confirm caller, owning system, data minimisation, timeout, retry, audit, and fallback for each. Pay particular attention to maintenance windows and version policy on cross-department APIs. If an internal API has no service commitment, it will still become an outage the public can see.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Engineering Discipline for Generated Code: The Faster the Speed, the Earlier the Guardrails

Treat AI as a pair engineer

It is suited to generating boilerplate, explaining legacy code, adding tests, and proposing refactoring options. It does not carry final accountability. Developers must provide clear context, architectural constraints, and acceptance conditions, and must understand every change. Code that cannot be explained should not be merged, even if it temporarily passes tests.

Shift security left to generation time

Before code enters the repository, run secret scanning, dependency-vulnerability checks, licence checks, static analysis, and security rules. For sensitive modules such as identity, payment, encryption, file upload, and command execution, limit the scope of automatic generation and require dual review by senior engineers.

Prevent technical debt from accelerating

AI tends to complete a local task. It may duplicate logic, bypass domain boundaries, or introduce new dependencies. Code review should examine overall consistency: whether platform capabilities are reused, whether observability is broken, whether running cost rises, and whether a migration burden is left behind. Fast completion does not equal a low lifecycle cost.

A team convention

All generated code must be labelled with its task source and accompanied by tests. Unauthorised sensitive code or data must not be entered. Any new dependency must state necessity and an exit path. Critical logic must have a human-readable design note. Make speed sustainable through institutional rules, not personal caution.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Testing Becomes a Shared Language Across the Whole Delivery, Not an End-Stage Safeguard

Drive collaboration with acceptance scenarios

At the requirements stage, write “given these conditions, when this action occurs, expect this result.” Business confirms meaning; design confirms interaction; development confirms implementation; testing confirms coverage. Scenarios connect policy rules, APIs, and user journeys at once, and are among the most effective shared languages for removing role boundaries.

AI widens the breadth of testing

A model can generate boundary cases, exception combinations, and regression sets from rules and historical defects, and can turn natural-language policy into candidate tests. Humans must review whether coverage is reasonable, especially uncommon but high-impact cases such as minority groups, extreme data, cross-year policy, and multiple agent identities.

Production quality cannot be judged by pass rate alone

All tests may pass and the system may still fail because of environment configuration, capacity, downstream dependencies, or a shift in data distribution. Before go-live, complete a performance baseline, fault injection, permission verification, disaster-recovery drills, and observability checks. After go-live, keep validating against real metrics and turn incidents into new regression cases.

How to run a defect retrospective

Do not only ask who wrote the wrong code. Trace which requirement was not expressed, which rule lacked a test, which gate failed to intercept, and which alert arrived too late. Repair the root cause in the working system, not only the current defect, or the same class of problem will be reproduced faster after AI accelerates delivery.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


RAG and Knowledge Engineering: Answers Need Sources, Currency, and Boundaries

Retrieval augmentation is not finished when files are uploaded

A high-quality knowledge base needs source grading, document chunking, metadata, validity periods, permissions, and an update owner. Policy, citizen guidance, internal operating manuals, and historical Q&A have different authority levels and must not be mixed for the model to choose among. Every passage must be able to return to its original source.

Answers must express uncertainty

When retrieval is insufficient, sources conflict, or the question exceeds permission, the system should refuse, ask for more information, or escalate to a human. Do not sacrifice correctness to raise answer rate. For public services, saying “uncertain” clearly is usually more responsible than giving a fluent but wrong answer.

Evaluation must stay close to the task

Beyond citation hit rate and factual correctness, assess whether a critical exception was omitted, whether sensitive data leaked, whether an ultra vires recommendation was given, whether expired policy was used, and whether refusal was correct. Build a real-question set, a hard-question set, and an adversarial set, and keep regressing after knowledge or model updates.

Operations experience

Assign a content owner and an update deadline to every knowledge domain. Monitor low-confidence answers, missed questions, and user corrections, and place them in a content-maintenance queue. A knowledge base is a long-term operating capability, not a one-off delivery attachment.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Model Gateway: A Unified Control Plane Across Brands and Models

Why a gateway is needed

Different tasks have different requirements for language, reasoning, vision, latency, cost, and deployment location. A model gateway unifies identity, routing, quotas, de-identification, logging, caching, evaluation, and failover, so applications need not bind directly to a single vendor API and can adjust when policy or price changes.

A brand-neutral view of selection

ByteDance, Google, AWS, Tencent Cloud, and other enterprise platforms each have ecosystem strengths. Government and large SOEs must not let brand volume replace scenario validation. Combine choices on data residency, compliance evidence, model effectiveness, integration difficulty, observability, supply continuity, and exit cost.

Routing must be explainable

Simple tasks may use a lower-cost model; complex reasoning uses a stronger model; sensitive tasks enter a dedicated network or on-premises capability. Every routing rule must record reason, version, and evaluation basis. “Intelligent routing” must not become an unauditable black box, or both cost and risk become hard to control.

Practical points

First unify the application calling convention and log format, then connect multiple models. Establish a replaceable request-and-response protocol and keep vendor-specific fields out of business code. Drill a switch once a quarter and verify that the standby model, knowledge base, and human process are genuinely usable.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Data Classification and Minimisation: What May Enter a Model

Start from the task, not from data convenience

First ask which data are the minimum needed to complete the task, then decide processing location and model. Public data may use public capability; internal data need a controlled environment; personally sensitive and critical business data may be processed only on a dedicated network, in a trusted execution environment, or after de-identification. Do not widen the data scope because a model “might be useful.”

Protect both input and output

On the input side, detect sensitive fields, malicious prompts, and out-of-scope data. On the output side, check privacy leakage, restricted content, and improper inference. Permission must run through retrieval, tool calls, and the final answer, not be verified only once at the entrance. Data the model cannot see must also be blocked from plugins and logs.

Retention and secondary use

State how long prompts, outputs, and runtime logs are kept, who may access them, and whether they may be used for training or evaluation. Vendor defaults may not meet institutional requirements. Both contract and technical configuration must forbid unauthorised secondary use and provide deletion, export, and audit mechanisms.

A field judgement method

For each AI use case, draw the data flow: source, transformation, storage, model, tools, recipients, and logs. Label data class, legal basis, and owner at every point. Any unexplained copy or cache should be treated as an item to remediate, not left until a security assessment.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Agent-Based AI: Tool Calls Carry Higher Risk Than Text Generation

From answerer to actor

An agent may query systems, create tickets, send notices, or even modify records. The closer the capability is to a real action, the more risk depends on authorisation, tool boundaries, and transaction integrity rather than language quality. Designing “what it may say” separately from “what it may do” is the first step in agent governance.

Least privilege and stepwise authorisation

Each agent receives only the minimum tools and data scope needed for the task. High-impact actions use preview, human confirmation, dual approval, or delayed execution. Set quantity and amount caps on batch operations, and require idempotency, timeout, and rollback so one error does not spread.

Defend against prompt injection

External web pages, email, and files may contain instructions that induce an agent to exceed its authority. The system must isolate data from instructions, restrict callable tools, validate parameters, and down-rank cross-domain content. Do not rely on the model itself to decide which text is trustworthy.

A recommended drill

Deliberately provide an attachment that contains a malicious instruction and test whether the agent leaks data, changes recipients, or performs extra actions. Also simulate tool timeouts, abnormal returns, and repeated callbacks. Only after failure paths have been verified does an agent meet the minimum condition for production.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Observability: Connect the Model, the Application, and Business Outcomes

Three layers of observation

The model layer watches latency, errors, tokens, refusals, and security blocks. The application layer watches retrieval, tool calls, cache, queues, and dependencies. The business layer watches task success, first-time completion, human handover, and appeals. Watching only model calls misses process problems that actually affect the public.

Trace one complete journey

From the user request, give every step a unified trace identifier that connects identity verification, knowledge retrieval, rule judgement, API calls, human review, and the final notice. When a dispute arises, the team can reconstruct the data seen at the time, the versions used, and the decision made, instead of seeing only a final answer.

Alerts must map to action

High latency, missing citations, abnormal refusals, cost spikes, and sensitive-content blocks should each have a different owner and playbook. An alert without a clear action is only noise. For critical public services, also monitor whether the human queue is overloaded after the model degrades.

Experience in brief

First define the ten most critical business and risk signals, then expand technical metrics. Dashboard data must be able to drive a decision: pause a model, switch a knowledge version, raise human review, or roll back a feature. The value of observability is shorter discovery and recovery, not more numbers on a screen.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


DevSecOps and the Software Supply Chain: Generated Dependencies Must Also Be Traceable

The supply-chain scope has widened

AI-generated code may introduce open-source packages, container images, models, datasets, and prompt templates. Each is part of the supply chain. Institutions need a software bill of materials, a model inventory, and a record of data provenance so they know what is actually running in production.

Gates that the pipeline must always pass

On commit: tests, static scanning, and dependency and licence checks. On build: generate an SBOM and sign it. Before deploy: verify artefact provenance and configuration. At runtime: monitor abnormal behaviour. Critical environments accept only pipeline-signed artefacts. Individuals must not bypass the process and upload directly.

Vulnerability response is not a one-time upgrade

When a new vulnerability appears, the platform should quickly locate affected systems, judge exploitability, and set mitigation and upgrade plans. If component versions and usage locations are unclear, response time is consumed by manual hunting. Transparent dependencies are the foundation of supply-chain resilience.

Procurement requirements

State in the contract the vulnerability-notification window, SBOM delivery, responsibility for third-party components, patch-support term, declarations of model and data provenance, and exit assistance. Without contractual support, technical governance often cannot require timely vendor cooperation in a high-risk event.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Cloud-Native Platform Capability: Not a Pile of Services, but Accountability Boundaries

Basic layers of platform capability

Container orchestration, service mesh, API gateway, GitOps, secrets management, model gateway, RAG, evaluation, design system, and immutable logs can form a shared foundation. Each capability needs a product owner, a support scope, an upgrade path, and service objectives.

The balance of sharing and autonomy

Identity, audit, secrets, network policy, and release evidence suit central governance. Business process, domain model, and service experience should be owned by the product squad. The platform provides secure defaults and self-service, but does not replace business decisions. Over-centralisation creates queues; over-autonomy fragments risk.

Capacity and cost must be assumed in advance

Define peak throughput, concurrency, data growth, inference budget, and scaling method for every component. The cost of an AI application may grow quickly with use. If estimates stay at PoC scale, formal rollout can easily lose control. Cost limits should be architectural constraints, not a finance review after the fact.

An exit plan

Ensure configuration can be exported, data uses open formats, infrastructure can be rebuilt from code, and business logic does not depend on proprietary APIs. Exit is not a prediction that a vendor will fail. It preserves bargaining power and continuity of public service. Portability that has never been drilled is only a paper promise.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Multi-Cloud, Hybrid Cloud, and On-Premises: Choose by Risk, Not by Slogan

Workload portraits

Choose deployment location by data sensitivity, latency, elasticity, ecosystem, network connectivity, and regulatory requirement. Public Q&A and internal high-sensitivity approvals should not share the same architecture. The value of a hybrid environment is placing different risks in the right place, not owning more brands at once.

Avoid surface multi-cloud

If an application is deployed on two clouds but depends on the same identity, the same network egress, or the same model vendor, a single point remains. True resilience analyses shared dependencies, operating capacity, and switch time. Duplicating resources is not duplicating capability.

A unified governance layer

Unify identity policy, log format, data classification, key rotation, artefact signing, and cost tags across environments. Without a unified control plane, multi-cloud multiplies the complexity of audit and incident handling. The platform team should reduce unnecessary difference while retaining essential local characteristics.

Selection experience

First validate portability, network, performance, cost, and operations on two or three representative workloads, then set a global policy. Do not first announce “full multi-cloud” or “everything on-premises.” Architecture should serve a public purpose and stand the test of three-year TCO and exit drills.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Brand-Neutral Technology Selection: Decide on Evidence, Not Position

Evaluate candidates on the same scenario

Run candidate platforms on the same data, the same tasks, the same security constraints, and the same peak conditions. Compare task success, citation accuracy, latency, cost, observability, deployment constraints, and staff learning cost. Promotional metrics from different brands cannot be compared sideways.

Look at the ecosystem and at the boundary

Large cloud providers typically have strengths in models, data, developer tools, global networks, or local services. Selection must weigh the delivery speed a mature ecosystem brings, and also identify proprietary APIs, data migration, talent dependence, and contract limits. Advantage and lock-in often appear together.

Procurement scores must be verifiable

Each scored item should correspond to evidence: a field test, a certification, a service commitment, a customer case, an incident report, or a contract clause. Do not award points directly for abstract phrases such as “industry-leading,” “sovereign and controllable,” or “international.” An unverifiable promise should be treated as a risk.

A combination strategy

An organisation may use several vendors under unified governance: procure general capability centrally, introduce specialist capability by scenario, and keep an alternative path for critical services. Brand diversity is not an equal budget split. It is a way to prevent any single choice from capturing the long-term roadmap.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


From PoC to Production: Six Stage Gates

The value gate

Confirm that the problem is real, the service users are clear, and a baseline can be measured, and prove that AI is more suitable than rules, process simplification, or ordinary search. A demonstration without clear public value should not enter a production budget.

The data and risk gate

Determine data sources, legal basis, classification, retention, and cross-border boundaries; complete an initial privacy and security assessment; define human accountability and prohibited actions. If any critical data source is unclear, pause expansion.

The engineering and operations gate

Complete architecture, testing, capacity, observability, disaster recovery, support, and cost validation. At the same time prepare training, operating manuals, incident response, and an appeal process. Go-live is the start of operations, not the end of the project.

Gradual volume increase

Start with internal users, then a small set of real users, then expand by risk and metrics. Set stop conditions at each stage—for example error rate, sensitive leakage, complaints, human-queue load, or unit cost exceeding a threshold. Stage gates let management continue, adjust, or exit on evidence, rather than be driven by sunk cost.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Practice: “Starting a Business as One Event”—Turn Policy into a Deliverable Chain

Start from the event, not from a departmental entrance

A business cares about completing incorporation, not how many departments process it internally. The team first lists applicant goals, agency relationships, required data, cross-department verification, and the final credential, then maps departmental duties onto the journey, instead of copying the organisation chart onto the internet.

One chain generates multiple artefacts

From an approved service blueprint, generate a clickable prototype, domain objects, API contracts, policy rules, test cases, and citizen guidance. Every artefact cites the same rule identifier. After a rule changes, affected interfaces, APIs, and guidance can be seen immediately.

Human review at high-risk nodes

Abnormal legal-person eligibility, name disputes, unclear authorisation, and sensitive-industry licences must not be decided automatically by a model. AI may organise materials, find gaps, and suggest a next step. Counter staff confirm against the rules and record edits and reasons. This reduces repetitive labour while retaining statutory accountability.

How the workshop is measured

Compare preparation time, defect count, rework cycles, and traceability time between the traditional and AI-assisted processes. More important is whether participants discover policy conflicts and exception paths earlier. Success is not generating pages faster. It is reaching a shared judgement of the facts across departments faster.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Success Case: Hong Kong Smart Government—From Shared Platforms to Scaled Services

A reusable digital foundation creates scale

Hong Kong’s smart-city programme supports departmental services with shared platforms such as digital identity, government cloud, big-data analytics, shared blockchain, and chatbots. By the end of 2025, more than one hundred digital-government and smart-city measures had been implemented, showing that building common capability centrally can reduce duplicate departmental investment and accelerate real-world delivery.

Digital identity drives service integration

iAM Smart covers more than four million registered users and more than 1,300 services and e-forms, and continues to add document, payment, stronger authentication, and enterprise-identity linkage. Success is not only user volume. Identity becomes a shared control point across services, giving pre-fill, signing, and personalised service a common foundation.

Data exchange improves the service experience

CDEG handles about two million data exchanges a month, sending already-verified data to the required service with user consent. This practice connects the identity, authorisation, provenance, and audit behind “fill one fewer form.” It is the key move from page-level convenience to institutionalised data reuse.

Lessons that can be borrowed, not copied

The lesson is to build shared capability and clear cross-department governance first, then expand applications—while still fitting local law, organisational maturity, and data boundaries. Copying a platform name is meaningless. Copying a reusable, auditable, extensible governance mechanism has value.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Success Case in Depth: AI for Government Work and a Multi-Vendor Catalogue

From point trials to a solution catalogue

Hong Kong’s AI-for-government work organises ready solutions from multiple vendors into a tools and solutions catalogue covering common areas such as intelligent customer service, meeting minutes, document processing, writing, process automation, creative work, and data analysis. Cataloguing reduces the cost of each department researching from zero, and creates conditions for side-by-side comparison and rapid trial.

Capability matching is more flexible than a single procurement

Through forums, seminars, and matching events, departments first state a pain point, then find a suitable capability. The mechanism accepts that different tasks need different technologies and does not cover every scenario with one brand. The central team’s role is to lower discovery and governance cost, not to decide every department’s business process.

Scale requires shared guardrails

Solutions in the catalogue still need data classification, permission, contract, evaluation, audit, and exit checks. Writing assistance and automatic execution should have different authorisation levels. Only when trial evidence is deposited as reusable evaluation and implementation guidance does the catalogue remain more than a product list.

Implications for Mainland government programmes

Build a cross-agency AI capability marketplace: unified access, a unified security baseline, unified procurement terms, and unified usage logs, while still allowing the best model to be chosen by task. First build experience on low-risk office scenarios, then enter rights-affecting and enforcement processes. Brand diversity can be kept while public risk is controlled.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Success Case in Depth: Open Data, Departmental Data Officers, and Consented Exchange

Openness and sharing are two kinds of governance

Open data faces society and emphasises public releasability, machine readability, and sustainable update. Inter-departmental sharing involves duty, permission, and purpose of use. Hong Kong both expands open-data resources and appoints departmental data officers and catalogue mechanisms, showing that data governance cannot rest on a technical portal alone.

A data catalogue makes accountability visible

Each department knows which data it owns, what the quality is, who maintains them, and whether they can be shared. A catalogue is not a static asset inventory. It is an entrance to service redesign. When a team designs a cross-department service, it can first look for reusable data, then decide whether to ask the public again.

An authorised gateway embodies minimisation

A service obtains already-verified data only for a stated purpose and with user consent. The exchange records source and recipient. Compared with pouring a full dataset into a single database, on-demand exchange makes purpose and accountability easier to control and retains evidence for appeal and audit.

An actionable recommendation

First establish a high-value data catalogue and named owners. Do not wait for whole-of-government governance to finish. Choose three cross-department services and validate consent, identity, field standards, error correction, and withdrawal. Driving data governance from real service improvement builds organisational consensus more readily than building a large platform in isolation.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Success Case in Depth: Local Large Models, Compute, and Application Ecosystem Alignment

Advance research, compute, and application together

Hong Kong’s AI ecosystem includes an AI and robotics research platform, a local generative-AI R&D centre, a supercomputing centre, funding schemes, and application products. The combination avoids investing only in models or only in compute, and connects research results, infrastructure, departmental demand, and industrial translation.

The value boundary of a local model

Local language, policy context, and data residency may require a dedicated model, but localisation does not automatically mean greater accuracy or safety. Still compare with other models on real tasks, and build update, evaluation, red-team, and operating capability. A model name cannot replace engineering evidence.

Government as an early user

Departments can provide real demand, a controlled trial environment, and scaled scenarios, while strict governance pushes products to maturity. If a funding scheme supports only training and not data engineering, evaluation, integration, and operations, results will be hard to turn into a stable service.

A lesson for regional collaboration

Organise regional universities, research, cloud services, industrial parks, and government scenarios into a mission alliance, delivering verifiable results around health, transport, and urban governance. Intellectual property, data use, and result-procurement rules should be clear in advance, so research success is not blocked from production.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


The Key to Smart-City Success: From a Technology List Back to Citizen Perception

Success is not a count of projects

Smart-city measures cover mobility, living, environment, talent, government, and the economy, but what the public actually perceives is whether waiting falls, whether information is accurate, and whether services are easier to obtain. Every technology must connect to an observable life outcome, so construction is not completed and then unused.

Shared infrastructure lowers the service threshold

Digital identity, real-time transport, electronic payment, government cloud, and open APIs mean new services need not rebuild foundational capability. Platform value shows in the speed and consistency of later innovation, not in the platform’s own feature list. The more departments reuse, the more a clear version, service objective, and support mechanism are needed.

Resilience and convenience are equally important

A smart city depends heavily on networks, data, and automation. Extreme weather, cyber attack, or vendor failure will amplify impact. City-scale systems must prepare alternative channels, offline continuity, cross-department warning, and coordinated recovery. Convenience must not rest on a single fragile chain.

Review questions

Every smart project should answer: who benefits, who may be excluded, who is accountable on failure, whether a non-digital channel exists, whether data can be corrected, and who maintains it in three years. Being able to keep answering these questions is the move from a technology demonstration to a public capability.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Public Safety and Cybersecurity: Development and Security Must Be Designed Together

Security is part of city capability

The more digital services concentrate, the more identity, cloud platforms, data exchange, and AI gateways become critical infrastructure. Complete threat modelling at design time. Identify how an attacker would use accounts, prompts, the supply chain, APIs, and internal permissions. Do not run one scan just before go-live.

Cross-department incident response

An attack may spread from one vendor or department. Unify incident classification, contacts, evidence format, and notification windows so technology, business, legal, and communications can act together. Major services must also connect to offline counters and emergency command.

Red teams must cover AI-specific risk

Test prompt injection, data leakage, ultra vires tool calls, refusal bypass, and knowledge poisoning, while retaining traditional identity, network, application, and supply-chain tests. AI security does not replace cybersecurity. It adds a new attack surface.

Recovery first

The playbook should state when to isolate a model, switch to read-only, disable tools, restore backups, and notify the public. Regular tabletop and technical drills verify that people know how to act under pressure. Security capability is finally shown in recovery speed and damage control.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Talent Transition: From Single Specialists to T-Shaped Full-Stack Creators

Deep specialism remains the root

A T-shaped person is not someone who knows a little of everything. They have sufficient depth in one profession and can also understand the language, constraints, and evidence of neighbouring fields. Designers must understand data and state; engineers must understand users and policy; business staff must understand system capability and risk.

Layered AI literacy

Everyone needs to recognise hallucination, privacy, and sources. Practitioners need task decomposition, evaluation, and review. Technical staff need to understand RAG, agents, security, and observability. Leaders need to set boundaries, a portfolio, and accountability mechanisms. Training cannot teach prompts alone.

Learn on a real task

Choose one low-risk process and have a cross-functional squad complete research, prototype, contract, tests, and a go-live plan in two weeks. Coaches review evidence at key nodes. Real delivery exposes organisational obstacles and turns skills into a shared way of working.

Career development

Build dual tracks: professional depth and cross-domain leadership are both recognised. Evaluation looks not only at individual output, but also at reusable assets, quality improvement, knowledge sharing, and risk reduction. The scarcest people in the AI era are those who can organise multiple professions into a reliable result.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


How Leaders Should Direct AI-Native Product Engineering

From approving features to approving risk

Leaders need not decide every technical detail, but they must be clear which outcomes are worth pursuing, which risks are unacceptable, and which matters must remain under human control. The focus of approval should move from page counts and feature lists to public value, accountability, evidence, and long-term operations.

Give the team a stable boundary

Make data red lines, authorisation tiers, platform standards, procurement rules, and go-live gates explicit, so the team can experiment inside the boundary. If every trial requires the policy to be re-explained, innovation stalls. If there is no boundary, risk fragments beyond governance.

Manage investment as a portfolio of problems

Classify programmes by task value and risk: low-risk efficiency, high-value service, foundation platform, and exploratory research. Different classes use different budgets, cycles, and success criteria. Do not load every AI wish onto one large programme, and do not leave scattered PoCs without an owner for long.

Four questions for the leadership standup

Which user outcome improved this period; which new risks appeared; which capabilities can be reused; if the model were withdrawn tomorrow, how would the service continue. Persistently asking these four points pulls discussion from technology heat back to governance and delivery.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Cost Governance and FinOps: Turn Every Inference into Manageable Unit Economics

Cost must attach to a task

Record model, token, retrieval, tool, storage, and human-review cost, and allocate them to a specific service and matter. Only when the true cost of completing one transaction is known can the organisation judge whether automation saves resources or merely moves cost from people to a cloud bill.

Optimise at the architecture layer

Use an appropriately sized model, prompt compression, caching, batching, retrieval filtering, and result reuse. Complex tasks can be routed in layers so simple steps do not call the most expensive model. Cost control cannot rely on a month-end cap. It should be validated in every design and release.

Prevent success from causing overspend

A PoC has few users. After formal rollout, call volume may grow quickly. Set budget thresholds, anomaly alerts, and per-user or per-matter quotas, and simulate peaks. When a cost ceiling is reached, degrade to a cheaper model or a static service rather than stop abruptly.

Assess TCO

Three-year cost includes platform, integration, network, security, data engineering, people, operations, migration, and exit. A low unit-price service that is highly proprietary, poorly logged, or needs heavy human remediation may be more expensive overall. Compare brands and architectures across the full lifecycle.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Outcome Metrics: From Model Accuracy Back to Public Value

Public-value metrics

Watch processing time, first-time completion, repeated submission, frontline burden, user help-seeking, and satisfaction. If an AI programme raises model scores but the service experience does not change, optimisation has not penetrated the business process. Metrics should have a pre-launch baseline and a clear target.

Model and engineering metrics

Task success, citation hit, hallucination, refusal, sensitive leakage, latency, availability, change failure, and time to repair reflect system quality. They are intermediate signals that explain business outcomes. They cannot define success alone.

Governance and fairness metrics

Record high-risk human review, traceability, policy-violation blocks, appeal closure, and error differences across groups. An overall average may hide disproportionate impact on a minority. Split by reasonable dimensions and interpret with specialists.

Against metric games

The more metrics, the easier it is to lose focus. Each stage should choose a small set that can trigger action, with an owner, a data source, and a threshold. If a metric does not affect any decision for several months, delete or redefine it. The purpose of measurement is not to prove project success. It is to help the programme become better.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Acceptance System: Function, Model, Engineering, and Governance in Parallel

Functional acceptance

Validate core and exception journeys, permissions, state, notices, and cross-system consistency. Do not only tick buttons against a requirements list. Complete end-to-end processing with real tasks and representative data.

Model acceptance

Test correctness, citation, refusal, safety, and stability on a frozen baseline set, a hard set, and an adversarial set, and record model and prompt versions. After a vendor updates a model, re-evaluate. Old conclusions must not be reused.

Engineering acceptance

Cover performance, capacity, availability, disaster recovery, backup and restore, observability, supply chain, and maintainability. Require an operations manual, alert rules, and evidence of failure drills. That a system can run once does not mean it can be operated for the long term.

Governance acceptance

Check data basis, human accountability, logs, appeals, contract, cost, and exit. If any of the four lines fails, do not release because the demonstration looked good. Acceptance conclusions must be traceable to evidence, and residual risk must have a named acceptor.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Common Failure Modes: What Errors AI Accelerates

Wrong requirements are implemented at scale

The team generates pages, code, and copy before validating the public problem, so rework grows larger. The remedy is to complete the task contract and user validation before generation, and to set stop conditions on hypotheses.

Prompts become a hidden rule store

Critical policy exists only in personal prompts or chat logs and cannot be versioned, tested, or audited. Structure the rules and link them to their basis. Prompts should only govern how they are invoked.

Automation exceeds authority

To demonstrate effect, an agent is allowed to modify records or send decisions without least privilege, confirmation, or rollback. Raise authorisation step by step from read-only, to suggestion, to preview, and decide whether to expand only after real failure drills.

The platform goes first and no one reuses it

A central team builds a complex platform without co-shaping it with real product squads, so projects bypass it. The right approach is to co-create a minimum platform capability around three to five high-value scenarios and attract reuse through quality and efficiency.

Looking only at the model, not at operations

After go-live no one maintains knowledge, evaluation, or cost, and effect declines quickly. Every AI capability must have a product owner, a content owner, a technical owner, and a risk owner, and must enter the regular budget.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


A 90-Day Path to Implementation: Build Reusable Capability from One Service

Month one: topic selection and baseline

Choose a high-frequency, relatively clear-ruled, controllable-risk, and measurable service. Establish the current journey, processing time, rejected applications, human burden, and enquiry baseline. Complete data classification, the task contract, and the human–AI accountability boundary. Do not rush to choose a model.

Month two: shared artefacts and prototype

Build a small policy-rule library, service blueprint, domain objects, API contracts, and evaluation set. Use multiple models to generate prototypes and tests, reviewed together by business, legal, design, engineering, and cybersecurity. Prepare logs, cost, and fallback in parallel.

Month three: controlled pilot

Let internal staff use it first, then open it to a small set of real users. Review errors, human handovers, sensitive blocks, and cost daily. Adjust knowledge, process, and prompts weekly. Expand only after quality thresholds are met. Do not force go-live by calendar date.

Deposit reusable assets

What the pilot delivers is not only an application. It also includes task templates, evaluation sets, rule patterns, components, procurement clauses, operating manuals, and a retrospective. The next programme should reuse these assets so organisational capability accumulates with each delivery.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Five Decision Checklists: Ready to Use on Return to the Organisation

Public value and risk checklist

State service users, current pain points, a quantifiable baseline, why AI is necessary, potential harm, affected groups, and stop conditions. Do not initiate any use case that cannot state its public value.

Layered-architecture checklist

List data, model, knowledge, agents, cloud, endpoints, identity, logs, and dependencies layer by layer, marking owner, residency, capacity, and alternatives. An architecture diagram must support decisions, not only reporting.

Stage-gate checklist

From value, data, prototype, engineering, and pilot to production, define evidence, approver, and return conditions for each stage. Let continued investment become a fact-based decision.

Vendor-assessment checklist

Compare effectiveness, compliance, ecosystem, integration, SLO, cost, proprietary dependence, data export, and exit support. Attach verification evidence to every score.

Accountability and exit checklist

List owners for AI recommendation, automatic execution, human approval, audit, appeal, incident, decommission, switch, and data deletion. A system that knows from day one how to end safely has long-term sustainability.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Ten Decision Questions Before a Closed-Door Q&A

Does it solve a real public bottleneck

Is the current problem policy complexity, process duplication, and data that cannot be shared—or merely the absence of a chat interface? If the root cause is not information generation, AI may not be the priority.

Can basic service be maintained during an outage

Have the standby model, static guidance, human counter, and offline process been drilled? During degradation, how are duplicate submissions and incorrect commitments prevented?

Who is accountable for the output

At which node does human review hold a veto; who signs risk acceptance; how does a user appeal; are the vendor and agency boundaries written into the contract?

Can lock-in be reduced

Are APIs open; can configuration be exported; can infrastructure be rebuilt; can knowledge and logs be migrated; has an alternative model been tested?

Is acceptance complete

Does it cover correctness, safety, fairness, latency, cost, availability, accessibility, and appeal together? If only function is accepted, risk moves into operations.

Does the benefit reach the public and the front line

The end state should show faster processing, higher first-time completion, fewer repeated submissions, and lower counter burden. If only generation volume and call volume grow, revisit the objective.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.


Closing: A Trustworthy, Restrained, and Sustainable Public AI Capability

Boundaries may fade; accountability must not

AI lets design, development, testing, and operations share the same work object, and also lets error propagate faster. The organisation must keep accountability through clear authorisation, specialist review, and traceable evidence. Fusion does not cancel specialism. It brings specialisms into collaboration earlier.

Platformisation avoids repeating risk

Government should build reusable identity, model gateway, data exchange, evaluation, design-system, and audit capability, so units innovate inside shared guardrails. The platform’s purpose is to lower the entry threshold and the cost of governance, not to monopolise every decision.

Validate direction against success cases

Smart-government digital identity, shared platforms, consented data exchange, a multi-vendor AI catalogue, and a local ecosystem show that scale comes from the alignment of infrastructure, governance, and real scenarios. What can be copied is the mechanism, not a brand or product name.

A final action

Start from one high-value service. Establish shared work objects, a human–AI accountability boundary, stage gates, and operating metrics. Form reusable assets in ninety days, then expand step by step. The goal of public AI is not the fastest generation. It is to create, under long-term constraints, public value that is explainable, maintainable, and appealable.

Takeaway

This page should be reused in solution reviews, procurement justifications, and go-live retrospectives. Teams must convert the argument into verifiable accountabilities, evidence, metrics, and exit actions, rather than stopping at conceptual consensus.