← Financial Cloud Cloud Cloud Club · AWS Re:cap

Vertex Macro | Financial Cloud Cloud · AWS Re:cap

AWS Re:cap 07: Authorized Operation of Public Data and Smart-Government Practice

Speaker: Government data

Session: 07

Session
Summit Dev Lounge2026 Re:cap
01 Encode Architecture as Steering for AI Agents
Summit Dev Lounge2026 Re:cap
02 Agent Harness Is the Real Engineering Moat
Summit Dev Lounge2026 Re:cap
03 Ask Observability Data in Plain Language
Summit Dev Lounge2026 Re:cap
04 Serverless AR Game with Bedrock AgentCore
Summit Dev Lounge2026 Re:cap
05 Multi-Agent Quant Backtesting on AgentCore
Summit Dev Lounge2026 Re:cap
06 Blog to Slides in Three Minutes with Kiro
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 AWS Compliance with Terraform
AWS Community Day Hong Kong 2025 Re:cap
03 Beginner to Builder An AWeSome Cloud Journey
AWS Community Day Hong Kong 2025 Re:cap
04 Team-First Serverless Engineering with Laravel & Bref
AWS Community Day Hong Kong 2025 Re:cap
05 Event Opening - AWS Community Day Hong Kong 2025
AWS Community Day Hong Kong 2025 Re:cap
06 Agent-to-Agent: Building Interoperable AI on AWS
AWS Community Day Hong Kong 2025 Re:cap
07 Utilize another telemetry data for faster improvement with AI agent
AWS Community Day Hong Kong 2025 Re:cap
08 Graduating from Vibe Coding: Spec-Driven Development with Kiro
AWS Community Day Hong Kong 2025 Re:cap
09 Automated Testing using MCP & AI Agents
AWS Community Day Hong Kong 2025 Re:cap
10 Modernizing Telecom Security ML Powered Approach
AWS Community Day Hong Kong 2025 Re:cap
11 Rethinking GenAI Agent: RAG & MCP
AWS Community Day Hong Kong 2025 Re:cap
12 Disaster and Emergency Response with TAK and AWS
AWS Community Day Hong Kong 2025 Re:cap
13 Rethinking Serverless Application Workflows from a Testing Perspective
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 AI-Powered Global Pure-Alpha Macro Trades on AWS: Revolutionizing Risk-Adjusted Asset Returns
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 Modern Trade Lifecycle: Trading to Settlement
FSI Recap
02 Goldman Sachs: Fast Track your applications onto Cloud - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group: Building an Effective Log Management Solution on AWS
FSI Recap
04 FSI Meetup 2025 Q4 - Brex Database Disaster Recovery
FSI Recap
05 FSI Meetup 2025 Q4 - A Graviton Migration Success Story
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel Modern Data Platform
FSI Recap
07 FSI Meetup 2025 Q4 - Financial Transaction Data Reconciler PayPal
FSI Recap
08 FSI Meetup 2025 Q4 - Scaling Resilience
FSI Recap
09 Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances
FSI Recap
10 Advanced Agentic AI Design Patterns
FSI Recap
11 Build New Modern Apps on AWS
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent Recap (IND3312)
AWS re:Invent 2025
02 Building the Future Trading Platform Leveraging AI and AWS
AWS re:Invent 2025
03 Trading Innovation: Jefferies' AI Assistant on Amazon Bedrock (IND3315)
AWS re:Invent 2025
04 How FSI Revolutionized HFT Analytics with Agentic AI (GBL302)
AWS re:Invent 2025
05 Improving Distributed Systems with Amazon Time Sync Featuring Nasdaq
AWS re:Invent 2025
06 Amazon Aurora HA and DR Design Patterns for Global Resilience (DAT442)
AWS re:Invent 2025
07 Building Agentic AI: Amazon Nova Act and Strands Agents in Practice (DEV327)
AWS re:Invent 2025
08 Deep Dive into Amazon Aurora and Its Innovations (DAT441)
AWS re:Invent 2025
09 Deep dive on Amazon S3 (STG407)
AWS re:Invent 2025
10 Nasdaq: Build Resilient Infrastructure for Global Financial Services (HMC327)
AWS re:Invent 2025
11 What's New with AWS Lambda (CNS376)
AWS re:Invent 2025
12 Spec-Driven Development with Kiro (DEV314)
AWS re:Invent 2025
13 Amazon's finops: Cloud cost lessons from a global e-commerce giant (AMZ308)
AWS re:Invent 2025
14 Tick to trade latency trading platforms on aws
AWS re:Invent 2025
Government data
01 The AI Era: The Boundary Between Development and Design Is Disappearing
Government data
02 On-Device Multimodal AI and Smart-City Practice
Government data
03 Large-Model Capability Evaluation and a Method for Landing AI Projects
Government data
04 Controlled End-to-End Automation of Government Development with Cloud Agents
Government data
05 A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
Government data
06 AI-Driven Macro Quantitative Research and Smart Governance
Government data
07 Authorized Operation of Public Data and Smart-Government Practice
Government data
08 Putting Data Assetization into Practice: Rights, Compliance, Engineering Governance, and Digital-Government Cases
Government data
09 AI for Mental-Health Public Welfare: Governance, Architecture, and Practice of a Trustworthy Platform
Government data
Amarathon 2025 Recap
01 A Developer’s Roadmap to Architecting for Agents
Donnie Prakoso
02 Amazon Bedrock Data Automation
Hafiz Syed Ashir Hassan
03 Multi-Agent on AgentCore
Tan Xin
04 Building Agentic AI Nova Act and Strands Agents in Practice
Haowen Huang
04 Accelerating Migration Projects with Kiro using Spec-Driven Development
Sanchit Dilip Jain
06 From Matching to Understanding: Personalized AI Search Practice Driven by AgentCore Memory
Liu Cao
07 Observe to Optimize – LLM Observability to AIOps Turning real-time insights into intelligent automation
Jimmy Soh
08 Deploying TEAM and Building the Best Engineering Team
Yuji Oshima
09 Five Hard Lessons from Five Years of So-Called Serverless Databases
Renato Losio
14 What if AI does my job How Q Developer CLI and Kiro have changed my daily routine
Miguel Angel Muñoz
16 Velocity with Vigilance: Security Essentials for Amazon Bedrock Agent Development
Brian Tarbox
26 Run OSS LLMs on a Single H100 Smarter, Cheaper, Faster
Adit Modi
28 A Modern Unified Metadata Architecture: New Approaches to Breaking Down Data Silos
Shaofeng Shi
29 Serverless MediaOps: Automating Video Workflows with AI on Amazon Web Services
Luis Valdivia
30 Architecting for Efficiency and Reliability with Performance Testing at Scale
Luis Guirigay
31 Connecting the World Through Open Source: Practical Journey of Technology, Community and Global Developer Relations
Richard Lin
33 Building Streaming Iceberg Tables for Real-Time Logistics Analytics
Fahad Shah
34 Accelerating Large-Scale Robot Strategy Training: An Automated Closed-Loop Architecture Based on Kiro, Trainium, and EKS
Junjie Tang
35 From Vibe to Viable with spec driven development
Ricardo Sueiras
36 Making Cloud Cost Analysis Smarter: Building FinOps Intelligent Agents with Strands and AgentCore
Xiaofei Li
37 Transform Conversational Agentic AIOps for K8s Using CNCF Kagent, K8sGPT, and Nova Sonic
Shaoyi Li

For government technology decision-makers, data directors, information-security directors, enterprise architects, and technology leaders in large state-owned enterprises


Authorized operation of public data: from a resource to a trusted service

Core judgement

The value of public data is not in moving raw data out of government. It is in combining statutory duties, data quality, use control, and service experience into a sustainable capability. Authorization is not a transfer of ownership, and it is not a simple sale of data. It is processing data into eligibility verification, risk alerts, aggregated analysis, decision support, or a trusted interface.

Decision focus

A government programme must answer public value, the authorizing party, controlled use, outcome accountability, and fail-safe degradation at the same time. Talking only about platforms produces an expensive data warehouse. Talking only about compliance may produce no usable product. Every technology choice must land on an accountability boundary, service level, cost cap, data residency, audit evidence, and an exit plan.

Practical starting point

First choose a high-frequency, low-controversy, measurable scenario. Draw the data flow, legal basis, accountable owners, purpose, retention period, metering, and exit. Then decide the platform and the brand.


Distinguish open data, sharing, and authorized operation first

Open data

Open data is addressed to unspecified social parties and usually provides low-sensitivity datasets or interfaces that can be reused publicly. The focus is open conditions, update frequency, machine readability, quality statements, and fair access. Data that should be open must not be packaged as an exclusive resource.

Government sharing

Government sharing mainly serves agencies in the discharge of their duties, and is exchanged on the basis of mandate, legal basis, and the least-necessary principle. Sharing is not unconditional retrieval. Purpose, fields, period, roles, and audit trail still need confirmation. Cross-department joint handling should prefer exchanging verification results, reducing repeat public submissions.

Authorized operation

Authorized operation is addressed to approved operators and users, and provides products through a defined scenario, use constraints, processing rules, metering method, and exit mechanism. When judging, ask whether users are specific, whether purpose is traceable, whether authorization can be revoked, and whether continuous supervision is required. The three mechanisms can coexist. Institutions and accountability must not be mixed.


Move from a data-volume mindset to scenario value

The wrong starting point

Measuring success by how much data is aggregated or how many subject databases are built often leads to data entering the lake first and requirements being written later. The competent authority bears concentration risk, business units see no benefit, the operator has no stable product, and what remains is storage and operating cost.

The right unit of demand

A scenario should describe who, at what point in time, on the basis of which controlled data, performs which reviewable action. Enterprise-support services, for example, are not the building of an enterprise big-data store. They are the conversion of policy conditions into verifiable rules that, under appropriate authorization, return eligible, missing-document, or needs-human-review.

Screening method

Ask first whether a high-frequency pain point exists, whether the minimum data can be obtained lawfully, whether the output can enter a business process, and whether benefit can be measured within twelve months. Then select lighthouse scenarios by public value, controllable risk, data availability, cross-department collaboration, and operating continuity.


Five-party governance: write accountability into institutions and systems

Role split

The competent party sets rules and approves boundaries. The provider ensures source, quality, and updates. The operator is accountable for processing, products, service, and day-to-day security. The user complies with purpose, retention, and onward-provision limits. The supervisor independently inspects authorization, operations, metering, returns, and incidents.

Where rights and duties land

Every role must have a clear decision right, execution duty, duty to be informed, veto, and incident accountability. The final decision on high-risk use cannot be fully outsourced to a platform or a model. A data provider also cannot, on the ground that delivery is complete, be excused from error correction and version notification.

Implementation tools

In institutions, form a responsibility matrix, approval forms, a register of data owners, and a joint emergency mechanism. In systems, configure role permissions, dual review, purpose tags, expiry, tickets, and non-repudiable logs. Institutions say who is accountable. Systems prove whether accountability is executed.


Data-rights boundary: authorization is not transfer

Define item by item

The authorization instrument should separately describe rights of access, processing, combination, derived indicators, external provision of results, re-authorization, retention, deletion, and model training. A vague statement that data may be used hides entirely different risks in a single sentence.

Retain control

Public bodies must retain approval of purpose change, emergency stop, audit sampling, error correction, version replacement, and expiry recovery. Labels, scores, or feature stores formed by the operator must also state intellectual property, portability, verifiability, and disposition after contract termination.

Prevent expansion

Any new purpose, new customer group, new region, new model, or new data combination should trigger a change assessment. Results that may affect individual rights, enterprise access, or the allocation of public resources must provide objection, correction, and human review. Reliable authorization can be understood, limited, withdrawn, and proven.


Data-product layering: do not treat a raw table as a product

Four-layer model

A data resource is raw or foundational data managed under law. A data product is a repeatably deliverable outcome after cleansing, standardization, de-identification, or aggregation. A data service is a capability continuously provided through an interface, a query, or trusted computation. An application outcome is the business effect formed when a user embeds the product in a process.

Security priority

At equal value, prefer statistical indicators over detail, verification results over fields, computation in a trusted environment over download, and short-lived credentials over long-lived copies. A smaller exposure surface usually buys more stable use.

Product specification

Every product should state scenario, legal basis, data scope, update frequency, quality, inputs and outputs, prohibited uses, retention period, service level, metering, cost, complaints, and unsubscribe. Only an outcome that can be clearly described, stably delivered, and continuously measured is an operable product.


Full lifecycle: from inventory to exit

Before admission

Complete a data inventory, classification and grading, legal-basis review, quality profiling, scenario assessment, and due diligence on the operating party. For personal data, trade secrets, or important data, first assess whether a result service, privacy-preserving computation, or a dedicated network can replace direct delivery.

During operation

Execute in sequence processing, product registration, user review, contract binding, purpose control, delivery, metering, quality monitoring, return accounting, and exception handling. Version upgrades must publish an impact statement, so that downstream parties do not make wrong decisions because of field, definition, or model change.

After exit

On expiry, breach, scenario cancellation, or vendor change, stop interfaces, revoke credentials, export evidence retained under law, delete copies and derivatives, verify backup disposition, settle fees, and notify all parties. Exit is not one sentence in a contract annex. It is an engineering capability that must be drilled and accepted.


Scenario design canvas: answer eight questions in one pass

Business and population

Who is the service for, which specific decision or process is to be improved, and how will public value be perceived after success. Do not write only "improve governance capability." Write which document is no longer submitted, which step is removed, and which class of error is avoided.

Data and control

Which minimum data or verification results are needed, on which statutory duty, consent, or contract, which outputs require human review, and which cases must refuse or transfer to a human. The data flow should mark source, processing points, models, interfaces, storage locations, and deletion nodes.

Operations and evidence

State how service, cost, and returns are metered, and how inputs, outputs, rule versions, model versions, knowledge sources, tool calls, human decisions, and exception events are retained. A workshop can have business, legal, security, and technology cross-challenge assumptions and quickly expose hidden premises.


Data contract: the technical contract for cross-department work

Beyond a dictionary

A data contract contains semantics, format, quality, updates, permissions, and change commitments at the same time. Every field needs a definition, source, allowed values, null rule, time basis, and sensitivity grade. The dataset as a whole needs update cadence, delay cap, availability, an accountable owner, and a traceability method.

Managing change

When a provider changes definitions, codes, or frequency, it first publishes a new version, a compatibility window, and a rollback plan. Users validate in a test environment. Unannounced replacement of production data is prohibited, because one field change can affect eligibility, risk, or funding decisions.

Acceptance advice

Test missing fields, duplicate subjects, cross-day updates, historical corrections, abnormal codes, late arrival, and permission revocation. Contract attainment is not only a successful interface. It is also semantic consistency, discoverable errors, notifiable impact, and locatable accountability.


Trusted data space: usable without free flow

Capability combination

A trusted data space is not a new centralized database. It is a combination of identity, authorization, policy, connectors, compute environment, logs, and evidence. Data can stay in its original domain. The user submits a lawful query or algorithm. The system returns a controlled result and records who, when, on what purpose, executed what.

Technology choice

Federated query, clean rooms, privacy-preserving computation, trusted execution environments, dynamic masking, watermarking, and differential privacy may be used, but every technique must map to a clear threat. Buying a privacy product is not completing governance. Output inference, repeated query, and bypass export still need to be prevented.

Minimum closed loop

First serve a single scenario with one identity source, one purpose policy, one result interface, and one complete audit chain. Expand multi-party collaboration only after revocation, rate limiting, blocking, and deletion have been verified. The standard of trust is that every use can be constrained by policy and replayed after the fact.


Minimum delivery: from data leaving the domain to a result service

Security gradient

Level one provides aggregated statistics, suited to planning and trends. Level two provides yes-or-no, grading, or missing-document prompts, suited to eligibility and risk verification. Level three provides detail query in a controlled environment only when there is no substitute. Each step up requires more approval, monitoring, and evidence.

Interface design

A result interface restricts fields, purpose, frequency, batch size, and return precision, and uses short-lived tokens and scenario credentials. High-risk queries set dual approval, anomaly thresholds, and human sampling. Interfaces that can be probed repeatedly add query budgets, result fuzzing, and behaviour analysis.

Value judgement

What a user truly needs is usually not ownership of the full dataset, but faster completion of a reliable decision. A result service can reduce storage, leakage, version mismatch, and deletion-proof cost, and it lets government retain correction and revocation capability. Product design should start from the task, not from how many fields can be provided.


Identity, consent, and purpose must not be treated as one thing

Identity answers who

Natural-person, legal-person, agency, and system identities must be managed in layers. A successful user login does not mean authority to represent an enterprise, and it does not mean a background job may continue to retrieve data. Identity verification, organizational authorization, post role, and machine credentials need to form a traceable chain.

Consent answers whether willing

Where the person's own authorization is involved, consent should be specific, informed, withdrawable, and time-bounded. The interface shows provider, recipient, purpose, scope, and retention method. After withdrawal, new exchanges stop, and disposition of lawfully deletable data and derivatives is triggered.

Purpose answers what may be done

Even when identity is genuine and consent has been given, the system must still check whether the purpose sits inside the legal basis and the contract. Purpose control needs policy tags, a scenario number, an interface allow-list, a retention period, and behaviour monitoring. The three together constitute verifiable minimum authorization.


Success case: digital identity as a service entry

Scale foundation

Hong Kong iAM Smart has formed a digital entry with more than four million registered users, supporting more than 1,300 services and e-forms, and has obtained ISO certifications related to information security and privacy management. Value is not only single sign-on. It is the gradual combination of identity, forms, documents, signing, payment, and personalized service.

Replicable design

Government digital identity should be platformized, not an account built by each department. The shared identity layer is accountable for authentication strength, stepped authentication, electronic signing, and authorization credentials. Business systems handle only their own duties. Personal codes, digital documents, a digital document wallet, and mini-programmes make services reusable.

Enterprise-identity implication

CorpID is planned around enterprise verification, digital signature, form pre-fill, and a digital document wallet, with a sandbox concept-proof first. Enterprise identity must also handle the legal person, authorized representatives, post changes, and multi-person signing. This shows that an identity platform should expand in steps through ecosystem access and testable interfaces.


Success case: a consented data-exchange gateway

Real outcomes

Hong Kong Consented Data Exchange Gateway (CDEG) exchanges verified data between government departments or authorized institutions on the premise of user consent, supporting a public-centred online application experience, at a scale of about two million data exchanges a month. This is a case that combines consent, a trusted source, and a service process.

Architecture focus

The gateway should not become a central store that permanently hoards all data. It is better suited to identity mapping, consent credentials, routing, format conversion, transport protection, status reporting, and audit. The provider maintains the authoritative source. The applying department is accountable for use purpose and the business decision.

Implementation experience

First choose a service that can reduce repeat submission, has a clear source, and has limited fields. The interface shows what is exchanged, to whom, why it is used, and how long it is retained. The back end handles refusal, withdrawal, source unavailability, data inconsistency, and human supplementary documents. Metrics include fewer fields filled, fewer proofs submitted, faster processing, and resolvable disputes.


Success case: open data from supply to use

Scale evolution

Hong Kong open data has grown to more than 5,700 datasets, about 110 interfaces, and more than 2,500 data providers. Downloads rose from about 5 billion in 2019 to more than 80 billion in 2025, showing that continuous update, machine readability, and broad supply can form real use.

Beyond download volume

High download volume may come from popular interfaces or machine polling. Also watch active applications, successful requests, data freshness, error reports, developer retention, and public-service improvement. Low-use data should first be checked for discoverability, format, authorization, updates, and documentation.

Implication for authorized operation

Open data remains fair and non-exclusive. Authorized operation is entered only when purpose limits, case-by-case verification, or continuous service assurance are required. The two can share a catalog, metadata, interface management, and quality mechanisms. The latter adds user review, purpose control, metering, contracts, and exit, avoiding the commercialization of data that should be free.


Success case: a combined approach to one hundred digital-government initiatives

From single points to combinations

After reviewing e-government across departments, Hong Kong advanced more than 100 digital-government and smart-city initiatives by the end of 2025, using big data, artificial intelligence, blockchain, and geospatial analysis, and building a government cloud, a big-data analytics platform, a shared blockchain, intelligent customer-service robots, and a unified service entry.

Key to success

A shared platform reduces duplicated construction, but it must be paired with demand inventory, shared-service accountability, access standards, and outcome tracking. Without process redesign, a department may only move paper online. The public still repeats filling, waits for human transcription, and reconciles across systems.

Combined method

Every initiative marks which journey it improves, which capabilities it reuses, which old steps it retires, and which evidence it produces. The platform team provides standard components and service levels. The business department is accountable for process and outcome. The governance team provides the boundary. The goal is that the next service is faster, cheaper, and more consistent.


Cross-border data flow: align rules first, then connect systems

Three layers of difference

Cross-domain cooperation has legal-system, technical-standard, and business-process differences at the same time. Opening the network does not mean data may circulate. First determine data type, statutory purpose, recipient qualification, storage location, onward transfer, data-subject rights, incident notification, and regulatory collaboration.

Regional practice

The Guangdong–Hong Kong–Macao Greater Bay Area (GBA) signed a cooperation arrangement to facilitate cross-border data flow in 2023, then launched a personal-information cross-border standard-contract pilot, and from November 2024 expanded it across industries. Institutional tools, contract templates, and staged expansion can reduce uncertainty.

Engineering implementation

Build a cross-domain data-flow map and a processing-activity register. For every flow, configure legal basis, minimum fields, encryption, key ownership, access region, retention period, and termination method. Start with low-risk scenarios, retain local degradation and a human alternative, and jointly drill regulatory queries, data-subject requests, and security incidents.


Enterprise-policy matching: output only the necessary eligibility result

Process breakdown

Turn policy text into traceable conditions, including region, industry, size, credit, project, time, and document requirements. Every rule retains the policy clause, version, effective date, and interpretation owner. After an enterprise applies, the minimum data is retrieved under appropriate authorization to complete a first check.

Result design

Outputs are eligible, possibly eligible, missing data, in conflict, or needs human judgement, with the basis and next step listed item by item. Conditions that cannot be determined automatically must be clearly marked. A model must not replace administrative discretion. High-impact refusal results provide human review and complaint.

Operating lessons

A policy update triggers a rule version and regression tests. Inconsistent enterprise data shows the source and a correction path. Models are suited to text extraction, Q&A, and document assistance, not to making a final eligibility decision alone. Success is judged by less data filled, policy reach, first-time completion, human return, and time to correct a mismatch.


Transport-operations product: aggregated indicators serving planning

Product boundary

Transport data includes road speed, incidents, public-transport arrivals, parking, passenger flow, and facility status. For planning, logistics, and the public, prefer segment-level, time-band, and area-level indicators. Avoid unnecessary vehicle or personal trajectories. Different purposes use different precision, latency, and retention strategies.

City practice

Hong Kong smart mobility covers real-time adaptive traffic lights, parking vacancies, traffic-data analysis, free-flow charging, electronic enforcement, automated parking, and electronic driving licences. These capabilities should jointly serve road safety, journey reliability, green travel, and accessibility—not isolated device displays.

Hands-on method

Choose one congested corridor. Define sources, time synchronization, outliers, missing-value compensation, and publication delay. Establish before-and-after event comparison, confidence, and human dispatch. Acceptance looks not only at average speed, but also at peak tails, incidents, sensor offline, and reliability in severe weather.


Enterprise risk-verification API: keep alerts and decisions separate

Minimum return

A verification service may return whether the subject exists, whether a licence is valid, whether a specific restriction is hit, the data as-of date, and source status. Unless the scenario has a clear legal basis, do not return a full case file, related persons, or unnecessary detail. Grading must carry a rule reason and confidence, so that a black-box label does not cause improper refusal.

Purpose isolation

Government procurement, finance, park investment attraction, and supply chain have different risk purposes and cannot share an undifferentiated score. Every credential binds scenario, fields, frequency, and institution. Batch queries, unusual hours, and high failure rates trigger alerts, preventing nominal verification from becoming data collection.

Accountability arrangement

The provider is accountable for the authoritative record and updates. The operator is accountable for the interface, rules, and logs. The user is accountable for the final decision and complaints. When data is delayed or in dispute, return indeterminate rather than force a determination. Retain positive and negative test samples, rule versions, and human conclusions so that a dispute can be reconstructed.


Healthcare data: high value with high-intensity governance

Scenario layering

Healthcare data can support clinical care, public health, research, drug and device evaluation, and health management, but legal basis, subject expectation, and risk differ. Clinical emergency care values high availability and immediacy. Research values de-identification, ethics review, and output control. The same authorization cannot handle both.

Hong Kong experience

Smart-health directions include eHEALTH+, public-hospital digitalization, public-health protection, and healthcare innovation. Existing practice covers electronic health records, smart hospitals, telemedicine, and a medical big-data research platform. The shared lesson is trusted identity, authoritative records, and sharing limits first, then analysis and AI.

Security implementation

Patients, clinicians, researchers, and systems have layered identities, authorized by treatment relationship, consent, ethics approval, and duty. Imaging, genetics, and free text are strictly isolated. A model must not directly replace diagnosis and treatment, and must show source, version, and limits. Download, query, export, and research output all require review and must be traceable.


Northern Metropolis: spatial governance and data governance in parallel

Planning opportunity

The Northern Metropolis is positioned as an important space for innovation and technology, post-secondary education, healthcare innovation, and GBA collaboration. The next five years plan to deliver more than 70,000 residential units and one million square metres of economic floor area. If data governance lags infrastructure, parks will easily duplicate platforms, interfaces will be incompatible, and accountability will be hard to trace.

Digital foundation

At the planning stage for land, transport, energy, buildings, environment, and enterprise services, establish address, spatial, facility, legal-person, and project master data in parallel. Share identity, a geospatial foundation, IoT access, an event bus, and a data catalog, while isolating healthcare, research, enterprise, and public-administration data by domain.

Operating model

A park company may be accountable for public facilities and platform operations. Government retains rules, supervision, and public-interest control. Investment attraction, construction, safety, energy, and transport products each define a contract and service level. Acceptance looks at cross-institution collaboration, energy and spatial efficiency, incident response, enterprise administration, and resident experience.


Finance and high-value-added supply chains: data supporting trusted transactions

Serviceable scenarios

Trade finance, insurance, inspection, logistics, maritime, and aviation need a trusted association among enterprise, cargo, documents, transport, and payment. Data products can provide document-authenticity verification, status attestation, risk-event alerts, and compliance checks—not centralized exposure of complete commercial data.

Architecture principles

Use legal-person digital identity and verifiable credentials to confirm who issued, holds, and uses. Use event interfaces to synchronize key status. Use data contracts to unify port, airport, bank, insurance, and regulatory semantics. Trade secrets are isolated by transaction and role. Cross-border flow is bound to contract and purpose.

Operating judgement

Value should appear as less manual document checking, shorter financing, lower fraud, and fewer repeat submissions—not pricing by field count. SMEs should not be permanently excluded for insufficient data. Supplementary proof and human review must be retained. The platform supports access by multiple service providers, avoiding a single cloud or data vendor monopolizing the entry.


Smart environment: credible monitoring matters more than a handsome display

Data chain

Air, water, noise, waste, energy, and ecological monitoring depend on sensors, calibration, communications, algorithms, and human sampling. Every value should trace to device, location, time, calibration status, correction, and a quality flag. Real-time data without a quality flag can cause wrong enforcement or public misunderstanding.

Practice direction

Hong Kong smart environment covers vessel emissions, environmental assessment, smart recycling, illegal dumping, green transport, and nature conservation. Five-year directions also include zero-carbon energy, electric vehicles, new-energy transport, hydrogen, and sustainable aviation fuel, providing cross-department scenarios for public data products.

Operating advice

At publication, provide measurement method, coverage, delay, missingness, and revision records. Enforcement use adds evidence-grade retention and a device-maintenance chain. Environment products can serve planning, enterprise disclosure, and community action, but pricing must not obstruct a basic right to environmental information. Revenue should be reinvested in quality, sensor maintenance, and public education.


Artificial intelligence entering government: establish risk grading first

Low risk first

Meeting summaries, document classification, knowledge retrieval, form extraction, and draft assistance are suited to go first, but still need to handle sensitive data, citations, and human confirmation. Medium risk includes customer-service advice, document pre-review, and process recommendation. High-risk scenarios involving eligibility, enforcement, healthcare, funds, and rights must be stricter.

Platform experience

Hong Kong AI+ Civil Services uses a multi-vendor catalog covering digital-human customer service, meeting summaries, document processing, writing, process automation, creative work, and data analysis, and helps departments select through forums, seminars, and matching events. This reduces single-brand dependence, but a unified governance threshold is still required.

Go-live conditions

Every use case needs a task boundary, allowed data, forbidden inputs, accuracy and refusal standards, human review, monitoring, a cost cap, and an exit. Before go-live, test prompt injection, over-privileged tool calls, sensitive-data leakage, wrong citations, and vendor unavailability. Average accuracy must not conceal high-risk errors.


Generative AI evidence chain: every answer can be replayed

Mandatory records

At minimum retain user and role, input, output, system-prompt version, knowledge sources, retrieved passages, model and parameters, tool calls, content filtering, human edits, the finally adopted result, latency, and cost. The records themselves need graded protection. Audit must not create a new pool of sensitive data.

Citation and refusal

Answers should prefer citing authoritative, valid, and applicable policy documents, and present version and date. When data is insufficient, sources conflict, the matter is outside the duty, or a high-risk judgement is involved, the system must refuse, state the limit, and transfer to a human. In government scenarios, a fluent but unfounded answer is more dangerous than a clear refusal.

Replay method

Build a fixed test set and incident samples, and reconstruct the then processing with the same versions. External-model contracts must guarantee version notification, log access, retention period, and incident cooperation. Every model or knowledge-base update runs regression tests and records differences, ensuring that improvement does not break existing safety boundaries.


Cloud-native foundation: assemble by accountability, do not stack by brand

Capability map

A data foundation needs object storage, a lakehouse, master data, metadata, lineage, quality, and data contracts. An exchange layer needs interface management, an event bus, trusted connectors, and batch and stream processing. A security layer needs identity, keys, policy, masking, and logs. An operations layer needs a catalog, subscription, metering, tickets, and settlement.

Multi-brand view

AWS, Google Cloud, Tencent Cloud, and other international and local clouds can be compared. Equivalent capability can also be deployed on a proprietary cloud, a government cloud, or a localized environment. Selection does not pursue identical service names. It confirms interfaces, data formats, observability, compliance support, geographic coverage, cost, and exit conditions.

Architecture discipline

Every component states an accountable owner, SLO, RTO, RPO, capacity assumptions, data residency, encryption boundary, and cost cap. Core data uses portable formats. Infrastructure is managed as code. Interfaces use open standards. Important processes prepare a degradation path for single-zone or cloud-service failure.


Multi-cloud and hybrid cloud: what is truly managed is difference

Not all workloads need multi-cloud

Running the same system across multiple clouds increases identity, network, data-consistency, monitoring, and troubleshooting cost. Adopt it only when regulation, geography, resilience, bargaining, or a special capability truly requires it. In other scenarios, a primary cloud plus a migratable design is usually more practical than surface dual-active.

Unify and retain

Unify container orchestration, identity federation, log format, tracing identifiers, infrastructure as code, secrets management, and data formats. Retain each cloud's differences in managed databases, AI, networking, and security. If abstraction is taken to the lowest common capability, cloud-native value is lost.

Operating playbook

Decide workload location by sensitivity, latency, dependency, cost, and egress traffic. Each quarter test backup restore, credential rotation, Region switch, and interface failure. Contracts require data and configuration export, log availability, defect handling, and termination assistance. The outcome of multi-cloud is controllable choice, not the number of consoles.


API and event architecture: turn exchange into a governable service

Synchronous and asynchronous

Use an API when immediate verification and a clear response are needed. Use events for status change, batch notification, and cross-system decoupling. Not every exchange should become a daily file, and not every process should depend on a synchronous interface. Before design, confirm timeliness, retry, order, consistency, and compensation.

Governance elements

Every interface and event has an owner, version, purpose, consumers, sensitivity, rate, service level, and retirement date. An API gateway is accountable for identity, authorization, rate limiting, and audit. An event platform is accountable for topic permissions, message retention, replay, and dead-letter handling. Business errors and technical errors are separated.

Live tests

Test duplicate requests, out-of-order events, timeouts, partial success, source rollback, consumer offline, and permission revocation. High-value operations use idempotency keys and business serial numbers, ensuring replay without duplicate handling. Before retiring an old version, identify all consumers and provide a migration window.


Zero-trust security: land on every use of data

Assume nothing is trusted

Whether a request comes from the intranet, the cloud, or a partner institution, verify identity, device, workload, risk, and purpose. Authorization uses least privilege, short-lived credentials, and dynamic policy. Sensitive operations add stepped verification or dual approval. Network segmentation is only one line of defence. It cannot replace data-level control.

Prevent abuse

Public-data risk often comes from a legitimate account over-querying, bulk export, purpose drift, or internal sharing. Combine behaviour baselines, field-level permissions, dynamic masking, watermarking, query budgets, and anomaly alerts. Administrator and service accounts are monitored independently, so that privilege does not become a blind spot.

Evidence and response

Centrally retain identity, policy decisions, data access, administrative operations, and export records. Connect one business journey with a unified tracing identifier. On an incident, immediately revoke tokens, block interfaces, preserve evidence, notify accountable owners, and start an alternative service. Regularly drill leakage, vendor compromise, and key failure.


Data quality: let business consequence set the threshold

Quality dimensions

Completeness, accuracy, consistency, timeliness, uniqueness, and validity are common dimensions, but they must not be averaged. If a field error can cause a mistaken subsidy, a healthcare risk, or a credit misjudgement, its standard should be higher than a general analytics field. Quality rules must connect to a concrete business consequence.

Closed-loop handling

After automated profiling finds an anomaly, a ticket points to the data owner and records cause, impact, temporary disposition, root cause, and permanent fix. For products already delivered, the provider notifies users and, where necessary, resends results or withdraws a version. A red-green light on a platform with no one acting is not enough.

Hands-on checklist

For critical fields, establish allowed values, association rules, time rules, and source consistency, and validate with gold samples, boundary samples, and historical incidents. After go-live, track returns, corrections, and complaints caused by quality issues. Prefer improving high-consequence, high-frequency, and cross-department shared data.


Observability: watch service, data, model, and cost together

Four signals

The service layer looks at availability, latency, errors, and capacity. The data layer looks at freshness, missingness, distribution drift, and lineage. The model layer looks at task success, citation hit, refusal, hallucination, and safety intercept. The cost layer looks at each query, each case, storage, network, and idle resources.

Journey linkage

A healthy single component does not mean a healthy public service. Use a unified business serial number through identity, consent, data retrieval, model, human review, payment, and notification, so that why one case failed can be located. Operations looks at technical dashboards. Management should look at public-service outcomes and material risk.

Alert discipline

Every alert has an accountable owner, a time limit, a grade, a silence rule, and a runbook, so that a flood of no-action alerts does not drown an incident. Set business thresholds for data delay, model drift, and cost anomalies. After an incident, complete a timeline, root cause, improvement, verification, and knowledge capture.


Metering and pricing: do not turn public value into a field unit price

Metering unit

Metering may be by valid call, verification case, compute duration, subscription period, or service grade, but exclude failure, retry, test, and the vendor's own error. Batch services state a metering boundary so that users can reconcile and supervisors can sample.

Pricing structure

Price may comprise base operating cost, value-added processing, service assurance, and a reasonable return. Basic public services, public-interest research, and SME innovation may use a free allowance, cost subsidy, or tiered prices. Scarce administrative data must not form an unreasonable monopoly, and unnecessary supply must not be added for revenue.

Decision principle

First compute the full cost of data governance, security, audit, customer service, incidents, exit, and long-term maintenance, then set the price. Evaluating an operator cannot look only at revenue. Also look at public-service improvement, inclusion, data quality, compliance, and ecosystem innovation. Revenue is only one part of sustainable operation, not the only goal.


Return allocation and public value: build an explainable mechanism

Allocation basis

The data provider contributes the authoritative source and continuous updates. The operator invests in processing, platform, service, and risk management. A partner may provide algorithms or a channel. Allocation should be designed on verifiable input, risk borne, service performance, and public mandate. It should not become a permanent ratio from administrative negotiation.

Public reinvestment

Part of the return should be clearly reinvested in data quality, standards governance, security capability, frontline digitalization, and inclusive services. If revenue is used entirely for platform expansion, data-providing units see no improvement and willingness to cooperate falls. Reinvestment projects should have a budget, an accountable owner, and outcome measurement.

Avoid perverse incentives

Transaction volume must not drive a department to expand collection or lower review. High social-value, low commercial-return scenarios can be supported through fiscal purchase of service or a performance contract. Regulatory reports disclose revenue, cost, service volume, beneficiary groups, risk events, complaints, and data-subject rights together.


From PoC to production: six stage gates

First four

Gate 1 validates scenario value and accountable owners. Gate 2 confirms legal basis, minimum data, and authorization. Gate 3 completes threat modelling, architecture, and vendor assessment. Gate 4 does a concept proof with synthetic, anonymized, or already-public data. A PoC should not connect directly to the full volume of real data.

Last two

Gate 5 is a limited real-traffic pilot, validating process, human review, service level, cost, complaints, and incident handling. Gate 6 is formal go-live, requiring operations, on-call, disaster recovery, contract, budget, and supervision all to be in place. Every gate has a clear veto condition.

Promotion judgement

Promote only when value is measurable, risk is controlled, accountability lands on a person, evidence is traceable, cost is bearable, and failure can degrade. A successful demonstration is not production-ready. At the end of a pilot, form a stop, redesign, or expand decision. Avoid an indefinite trial.


Acceptance framework: seven classes of metric decide pass together

Outcomes and models

Public value looks at processing time, first-time completion, repeat submission, errors, and frontline burden. Model quality looks at task success, citation hit, hallucination, refusal, and sensitive-data leakage. High-risk scenarios also test by group, boundary condition, and worst case.

Engineering and governance

Engineering quality looks at p95 latency, availability, change-failure rate, mean time to repair, RTO, RPO, and disaster-recovery drills. Governance quality looks at human review, traceability, violation intercept, complaint closure, and authorization revocation. Security acceptance covers identity, over-privilege, export, and supply chain.

Cost and experience

Cost looks at inference per case, unit throughput, idle resources, and three-year total cost of ownership. Experience looks at understandability, accessibility, less fill and less submit, and transparent progress. Samples are designed jointly by business, technology, security, and supervision. Every metric has a baseline, a target, a measurement method, and an accountable owner.


Procurement evaluation: do not be led by a product demonstration

Procure capability and outcomes

The requirements document first describes business outcomes, data boundaries, service levels, security, evidence, cost, and exit, then lets vendors propose a solution. If a string of brand services is specified directly, procurement becomes a technology list and loses the ability to compare architecture fit and total cost.

Evaluation dimensions

Compare function, open interfaces, data portability, compliance evidence, geographic support, capacity, observability, cost transparency, model governance, supply-chain security, and exit assistance. International clouds, local clouds, localized platforms, and self-build use the same logic, with weights adjusted by environment.

Contract levers

Require version notification, a vulnerability-fix time limit, no unauthorized training on the data, a sub-processor list, incident notification, log access, service termination, and a deletion certificate. Selection looks not only at live effect, but also at whether the system can degrade on failure, whether three-year cost is controllable, and whether the team can take over.


Exit and portability: design departure before signature

What to take away

Data, metadata, lineage, quality rules, model configuration, prompts, knowledge bases, interface definitions, infrastructure code, monitoring, logs, and tickets have different portability requirements. Being able to export CSV does not mean a system can migrate.

Prove exit is possible

The contract agrees standard formats, export frequency, a fee cap, assistance hours, a parallel period, and a deletion certificate. At least once a year, do a small-scale export and restore. Verify file completeness, key replaceability, identifiable dependencies, and readable history. Do not wait until the termination date for the first test.

Service continuity

During switchover, retain a read-only service, a human window, or a batch backup, giving priority to livelihood and high-risk business. Responsibilities, handover, incident attribution, and confidentiality between old and new vendors are written down. Successful exit is uninterrupted service, no lost data, revoked permissions, and retainable evidence.


Common failures: a platform, no product, missing accountability

Platform first

Building an all-domain platform from the start, without a first wave of scenarios, product managers, and business owners, easily becomes long-term construction and short-term display. The correction is to choose three to five high-value scenarios, support them with a minimum shared capability, then abstract a platform from real reuse.

Blurred rights and duties

The competent authority believes the operator is accountable. The operator believes the provider is accountable. The user treats a model result as authority. After an incident, no one can state authorization, rules, and the final decision. Place clear accountability, a human veto, and replayable evidence at process nodes.

Revenue or technology only

If transaction volume replaces public value, unnecessary supply is induced. If algorithm precision replaces process effect, complaints and availability are ignored. Success should appear as the public submitting fewer documents, departments doing less repeat verification, risk being supervisable, and cost being sustainable. Features that cannot support these outcomes should be re-prioritized.


Organization and talent: build a cross-profession product team

Core roles

Every data product needs a business product owner, a data owner, an architect, a security and privacy specialist, legal, operations, finance, and user research. AI scenarios add a model owner and a human-review representative. Roles may be concurrent, but decision rights and time commitment must be clear.

Way of working

Deliver verifiable outcomes in two-to-four-week iterations, rather than only reporting progress in large meetings. Each iteration updates requirements, data contracts, threat models, tests, cost, and runbooks together. Business participates in samples and acceptance. Engineering understands the statutory process. Governance people enter design early.

Capability building

Training does not only teach tools. It also covers data grading, purpose limits, interfaces, cloud cost, model limits, incident response, and vendor management. Build a practice community and capture templates, incident cases, and architecture decisions. Maturity is whether the team can independently judge, operate, and exit—not the number of certificates.


Twelve-month implementation: from a minimum institutional set to a replicable capability

First three months

Form a cross-department governance group. Complete a data inventory and three lighthouse scenarios. Set a role matrix, classification and grading, an authorization template, data contracts, and stage gates. Inventory existing identity, cloud, API, log, and catalog capability. Prefer reuse over new build.

Middle six months

Complete a PoC with synthetic or public data, then a limited real-traffic pilot. Establish a result service, purpose control, metering, and an evidence chain. Implement human review, complaints, fault degradation, and exit drills. Report value, risk, cost, and issues to the competent layer every month.

Last three months

Stop low-value scenarios on measured evidence, correct viable scenarios, and go live formally. Distil shared identity, consent, interfaces, monitoring, and contract clauses into standard capabilities, and expand the second wave of products. Year-end outcomes include a running service, auditable evidence, replicable templates, and next-year investment priorities.


Leader decision checklist: fifteen questions before approval

Value and accountability

Does it solve a real public-service bottleneck? Who benefits and who is disadvantaged? Which department is accountable for the final outcome? At which node does a human have a veto? If the programme stops, what public capability is lost?

Data and technology

Are legal basis, consent, and purpose clear? Can a result service replace raw data? Which data must not enter a model or an external cloud? How does the system degrade if a vendor is unavailable? Can interfaces, formats, infrastructure code, and data export reduce lock-in?

Acceptance and continuity

Does acceptance cover correctness, security, fairness, latency, cost, availability, complaints, and disaster recovery? Who bears the three-year total cost of ownership? Are returns reinvested in governance and quality? How is authorization revoked and data deleted on expiry? Can supervision obtain complete evidence? Only when every question has an accountable owner and a written answer does the programme have the basic conditions to enter production.


Closing: convert into a stable, trusted, restrained public capability

Three persistences

Persist in being driven by scenarios and data products, not by data volume and platform scale. Persist in controllable purpose, inspectable process, reviewable results, and traceable violations. Persist in assessing public value, risk, and sustainable operation together, and do not let revenue or technology heat dominate alone.

Shared answer from the cases

Hong Kong digital identity, consented data exchange, open data, one hundred digital-government initiatives, cross-border rule pilots, and smart-city practice show that long-term success comes from a shared foundation, clear governance, stepwise expansion, and perceptible service—not from a one-off mega-build.

Action commitment

Start from a high-frequency, low-controversy, measurable scenario. First draw the data flow, legal basis, accountability, purpose, retention, metering, and exit, then choose technology. A mature system is not the one with the most features. It is the one that keeps running under institutional constraints, protects the public when it fails, and remains under organizational control when it changes.