Vertex Macro | Financial Cloud Cloud · AWS Re:cap
AWS Re:cap 02: On-Device Multimodal AI and Smart-City Practice
On-Device Multimodal AI: From a Three-Day Prototype to City-Scale Public Capability
Core proposition
A mobile app that can recognise a plate of food and estimate nutrition looks like a small product. In practice it touches image understanding, knowledge retrieval, on-device inference, privacy protection, public health, and long-term operations at the same time. What government and large enterprises should actually learn is not how many features can be finished in three days, but how to make scope, risk, data, and accountability boundaries explicit on day one.
Decision lens
This briefing uses public value as the main thread, connecting Hong Kong smart-city, digital government, data governance, the Northern Metropolis, health and medical innovation, and regional collaboration. Technology choices are not centred on a single brand. They are assembled according to regulation, data residency, latency, cost, supply-chain resilience, talent capability, and exit conditions.
Takeaways
The audience will leave with a reusable method: break demonstration AI into acceptably scoped work packages; establish a division of labour between on-device and cloud; control health risk with refusal, uncertainty, and human review; then use evidence chains, stage gates, and service metrics to upgrade a prototype into an auditable, operable, and scalable formal capability.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Why Calorie Recognition Speaks to Smart Government
The amplification effect of a small scene
A plate photo contains uncertainty from lighting, occlusion, mixed dishes, portion size, regional diets, and personal preference. That is very similar to the blurred identity documents, field inspections, disaster photos, and agricultural pest images that government image recognition must handle. A small health app is therefore an ideal engineering sandbox: it can expose, at low cost, the boundaries that AI systems find hardest.
How public services differ
Internet products can fix problems with rapid updates. Government services must simultaneously carry availability, explainability, fairness, complaint handling, and vendor continuity. If a model misclassifies, the cloud drops offline, or data leaks, the impact is not only a single user experience; it can become a public-trust and governance risk.
Conversion method
Treat the calorie app as a scaled-down digital-government system. Practise, item by item, identity, authorisation, data minimisation, knowledge sources, model versions, human veto, and degraded service. Once those fundamentals are in place, the same architecture can be safely reused for school meals, primary health education, elderly services, and offline outreach in remote areas.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
From Policy to Product: How a Five-Year Plan Becomes an Engineering Backlog
Strategic translation
Hong Kong’s first five-year economic and social development plan places innovation and technology, people’s livelihood, the Northern Metropolis, regional cooperation, green transition, and security governance in the same development frame. Engineering teams must translate those macro directions into concrete requirements: reducing repeated citizen submissions, improving availability on weak networks, shortening processing time, and giving cross-department collaboration clear authorisation and traceability.
Requirement layers
The first layer is public outcomes, such as health-education coverage, service reachability, and frontline burden. The second layer is business capability, such as image pre-screening, trusted data retrieval, and human referral. The third layer is technical components, such as on-device models, APIs, vector indexes, identity platforms, and logs. This avoids buying a platform first and then looking for a use.
Landing discipline
Every policy objective must map to an accountable owner, a service population, measurable indicators, a data basis, a budget ceiling, and an exit condition. If a feature cannot explain how it improves a public result, or can only prove value through model accuracy, it should not enter formal procurement and large-scale deployment.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
First Principle of a Three-Day Sprint: Freeze the Problem, Not the Learning
Scope boundary
Three days validate only one complete journey: photograph the plate, quality check, food candidates, portion cues, nutrition retrieval, risk labelling, and result presentation. Explicitly exclude disease diagnosis, personal treatment advice, full cuisine coverage, and clinical-grade precision, so the team is not dragged down by an impossible commitment.
Assumption list
Before the sprint starts, write down assumptions that can be overturned: whether a single photo is enough; whether the on-device model can run within the target phone’s memory; whether users understand interval estimates; which functions still work on a weak network. Each assumption needs a test method and a stop condition.
Learning output
Success in three days is not the largest feature count. It is being able to answer whether it is worth continuing, where the largest risks sit, and what data and experts the next round needs. At the end, deliver a scope statement, test records, failure samples, model and knowledge versions, a cost estimate, a risk register, and a next-stage decision recommendation.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Day One: State Requirements, Data, and the Security Floor in One Pass
Morning work
Use half a day to complete the user journey, key roles, and error-consequence analysis. Nutrition-education users, campus administrators, health professionals, and system administrators see different information and have different permissions. Any high-risk prompt must predefine who can view it, who can change it, and who is responsible for answering complaints.
Data preparation
Use only public, synthetic, anonymised, or formally authorised images and nutrition data. Build a minimal data dictionary that records dish names, regional aliases, estimation units, source dates, applicable scope, and limits. Do not collect unrelated faces, locations, device identifiers, or complete health records at the prototype stage.
Security gates
Complete threat modelling and abuse scenarios, including malicious images, prompt injection, knowledge-base poisoning, model-file substitution, and log leakage. If sensitive-information masking, transport encryption, version lock, and basic audit cannot be achieved, the prototype may be demonstrated only in an isolated environment and must not touch real citizen data.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Day Two: Make On-Device Inference Run, and Make Failure Visible
Inference pipeline
First check image sharpness, brightness, composition, and sensitive content, then hand the image to an on-device multimodal model to produce a limited number of food candidates. The output should include candidate names, visible evidence, a confidence interval, and angles that need a retake—not a single answer that looks certain.
Device budget
Build a device matrix across different phone memory, processors, operating systems, and battery conditions. Measure, item by item, model load time, first inference, sustained inference, peak memory, power draw, and surface temperature. If a high-end phone can run the model but devices commonly used at the front line cannot, the public-service value must be reassessed.
Failure visualisation
Classify and keep cases of occlusion, glare, mixed dishes, sauces, distorted utensil scale, and local dish-name differences. The team reviews the distribution of error types and newly appearing boundaries every day, rather than selecting only successful photos for display. Repeatable failure classification has more engineering value than a single polished result.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Day Three: Productisation, Evaluation, and Deliverable Evidence
Experience convergence
On the last day, converge features into understandable operations: capture guidance, processing state, result intervals, data sources, risk prompts, retake, and human help. The interface must not disguise model confidence as medical credibility, nor use precise decimals to create false certainty.
Minimal evaluation
Build a test set that includes normal, difficult, out-of-scope, and malicious inputs. Beyond food-candidate hit rate, also test whether refusal is correct, whether nutrition citations correspond, whether sensitive information leaks, whether offline mode is complete, whether latency is acceptable, and whether the system can recover safely after an error.
Delivery pack
The demonstration version should be delivered together with a model card, data card, version list, architecture decision records, test results, known limits, cost assumptions, and follow-on work. What decision-makers see is not only a screen, but complete evidence for judging risk, budget, and sustainability.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
On-Device, Cloud, and Hybrid Inference: Choose by Accountability Boundary
When on-device fits
On-device inference has a clear advantage when offline operation, low latency, data minimisation, or immediate field response is required. It can first perform image-quality checks, sensitive-information masking, preliminary classification, and simple rules, reducing the export of raw data. Device fragmentation, model updates, and compute limits, however, raise testing and support cost.
When cloud fits
When large-model capability, centralised knowledge updates, complex retrieval, or cross-department service sharing is needed, the cloud makes unified governance and scale easier. Data residency, network dependence, per-inference cost, vendor failure, and cross-border data flows must still be handled explicitly. The cloud should not become the default destination for all data.
Hybrid principle
The most practical design is usually on-device filter and summarise first, uploading the minimum data only after consent and when necessary; the cloud handles tasks that need stronger capability; when disconnected, the system falls back to offline knowledge and conservative prompts. Every segment needs clear SLO, timeout, retry, degrade, and rollback rules.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Model Quantisation Is Not a Compression Job; It Is Service Design
Engineering trade-offs
Quantisation can reduce model size, memory, and inference time, but it may also affect fine-grained recognition and language-generation stability. Do not compare file size alone. Check, against real tasks, whether dish candidates, refusal, regional aliases, and mixed-dish descriptions have degraded.
Cohort strategy
Build high, medium, and low configurations by device capability, choosing different model sizes, image resolutions, context lengths, and parallelism. Low-end devices can use a two-stage flow: a lightweight model for screening first, then a stronger capability only when needed, so not every device bears the same resource cost.
Release experience
A quantised version must have its own identifier, its own evaluation, and a rollback path. In a canary release, start with internal test devices, then expand to a small share of the real environment; watch crashes, power draw, temperature rise, latency, and task success. Watching accuracy alone often misses the performance problems users feel first.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Image-Quality Gate: First Decide Whether It Can Be Seen, Then Decide What Is Seen
Pre-checks
A good recognition pipeline should not send every photo straight into the model. First decide whether it is blurry, too dark, reflective, incompletely cropped, too far away, or contains multiple plates, then give specific capture guidance. This step reduces later errors and also saves on-device battery and cloud cost.
Sensitive masking
A plate photo may also capture a face, staff badge, medical record, address, or screen content. The system should detect and mask unnecessary information on-device first; the user previews and then decides whether to continue. Masking results must also enter testing, so over-masking does not destroy food judgement.
Actionable feedback
Do not only display “image not acceptable.” Point to the reason and the next step, for example move closer to the plate, add light, remove packaging, retake from above, or place a reference object of known size. Specific feedback turns model limits into a process design that users can follow.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Food-Candidate Recognition: From a Single Answer to an Evidence Set
Candidate output
For similar dishes, local names, and mixed plates, the model should output a small ranked set of candidates and explain visible ingredients, cooking method, and uncertain parts. Users may correct candidates, but corrections must not enter the formal knowledge base directly; they must pass a quality process.
Classification system
Build a layered vocabulary of dishes, ingredients, cooking methods, portion units, and food cultures. Splitting “char siu rice” into staple, protein, sauce, and side dishes helps nutrition estimation and cross-region mapping, and lets existing components be reused when new dishes appear.
Bias control
The evaluation set must cover Cantonese cuisine, ethnic-minority diets, vegetarian meals, school meals, soft meals for older people, and different utensils. If data come only from influencer photos or standard plating, the model will underestimate occlusion and diversity in real settings and produce systematic errors for particular communities.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Portion Estimation: The Largest Source of Calorie Error
Why it is hard
Correct food recognition does not equal correct calories. A single two-dimensional photo lacks depth, density, and container-size information. The same bowl of rice can look very different by camera angle; sauce, oil, and hidden ingredients are even harder to judge from the surface.
Reducing error
Ask for two photos from above and from the side, use standard utensils or a reference card, ask about bowl or plate size, and let the user choose small, medium, large, or a gram interval. The system should store the estimation method, not only the final number.
Result expression
Output should prefer intervals, main assumptions, and sensitive factors, for example “if two tablespoons of sauce are included, the upper estimate rises.” For mixed dishes that cannot be reasonably estimated, refuse a precise calorie figure and instead provide food composition and general health-education information.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Separate Knowledge Base from Model: Make Every Answer Traceable
Why separate
Models are good at understanding images and language, but they should not invent nutrition numbers from memory. Nutrition master data, allergens, unit conversion, and policy prompts should live in a versioned knowledge base. The model only proposes queries, organises results, and explains limits.
Source governance
Every record must document the issuing body, update date, applicable region, food state, portion unit, and licence conditions. When sources conflict, keep the difference; do not average multiple values into an answer that looks authoritative. Expired data or data without a source must not be used for high-risk prompts.
Update mechanism
Knowledge updates and model updates follow different cadences and approval processes. Small data corrections can be released quickly, but any change that affects allergens, health warnings, or legal wording needs dual review, testing, and rollback. This decoupling lowers the risk and cost of every update.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Uncertainty, Refusal, and Human Review
Safety design
The maturity of government AI is not answering more; it is knowing when not to answer. When image quality is low, candidate margins are small, knowledge sources are missing, or the user asks about disease treatment, the system should stop giving a precise conclusion, state the limits clearly, and guide a retake or professional help.
Human veto
Human review is not decoration. The process must define which outputs require review, what evidence reviewers need to see, how quickly they must reply, and whether a published result can be withdrawn or corrected. Staff at high-risk nodes must have an explicit veto and must not be forced to pass quickly by performance metrics.
Evaluation method
Refusal must also be measured. Wrong refusals reduce usability; failing to refuse when required increases risk. The test set needs normal, ambiguous, out-of-scope, adversarial, and health-critical situations, with separate rates for appropriate answers, appropriate refusals, and erroneous confidence.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Health-Safety Boundary: An Education Tool Must Not Impersonate a Medical System
Use limits
The application is positioned as health education, research demonstration, and general dietary awareness. It does not diagnose disease, give treatment advice, adjust medication, or make personal medical decisions. Interface, publicity, data retention, and staff training must be consistent. A disclaimer cannot claim low risk while the actual flow encourages medical dependence.
High-risk handling
If a user mentions severe allergy, hypoglycaemia, swallowing difficulty, pregnancy, kidney disease, or similar conditions, the system must not judge safety from a photo. It should provide a clear, non-diagnostic risk reminder and establish a referral path to qualified professionals or emergency services.
Accountability evidence
Every output retains model version, knowledge source, rule version, confidence, refusal reason, and user correction. If a complaint or suspected harm occurs, the organisation can reconstruct what the system saw at the time, what it relied on, and who made the final decision.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Offline First: Weak-Network Areas Still Need Basic Service
Basic pack
Offline mode at least retains capture guidance, image-quality checks, sensitive masking, common food candidates, basic nutrition education, and emergency risk prompts. Functions that need live data or professional review should be clearly marked as temporarily unavailable. Old data must not be presented as the latest result.
Sync strategy
After reconnection, upload only necessary summaries, error codes, and consented samples. Use queues, retries, de-duplication, and conflict handling so the same event is not submitted twice. A sync failure must not block the user from viewing completed offline results.
Drill requirements
Testing is not only cutting Wi-Fi. Also simulate high latency, frequent disconnects, low battery, insufficient storage, and clock skew. For remote health outreach, disaster sites, or large events, offline capability is service resilience, not an add-on.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Battery, Thermal Management, and Sustainable Use
Measurement scenarios
Measure power draw and temperature rise under continuous capture, background download, model update, and long inference. A smooth one-off demo does not mean everyday usability, because overheating causes thermal throttling and low-power mode may also limit background work and camera capability.
Control measures
Use lower-resolution pre-checks, dynamic batching, on-demand model loading, caching of common knowledge, and updates during idle time. High-energy tasks should tell the user why they are needed and allow deferral until charging or a better network.
Public-procurement view
Acceptance should measure energy per task, peak temperature, average latency, and crash rate on a specified device matrix. If a vendor submits only laboratory results from flagship devices, that does not prove the service can cover equipment commonly used at the front line.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Evidence Chain: Turn Every AI Action into an Auditable Event
Event content
A complete event includes a de-identified input summary, image-quality result, model and prompt versions, knowledge-retrieval sources, tool calls, rule hits, human review, final output, latency, and cost. Not all raw content needs long-term retention. The point is being able to reconstruct the decision.
Retention strategy
Set different retention periods by data class, and separate security events, model quality, business transactions, and debug logs. Highly sensitive original images can be deleted after on-device processing, retaining only a hash, a feature summary, or approved anonymous samples.
Audit value
The evidence chain supports complaint investigation, version rollback, bias analysis, vendor acceptance, and cost reconciliation. An AI system without evidence, even with high average accuracy, cannot carry accountability in a government setting.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
From PoC to Formal Go-Live: Six Stage Gates
Gate design
Pass, in order, six gates: problem value, data legality, technical feasibility, security and privacy, operational readiness, and public accountability. Each gate has required documents, quantitative thresholds, approval roles, and return conditions, so a prototype does not jump to full launch because of senior attention.
Stop conditions
If there is no lawful data source, human review cannot be established, there is no safe degrade when offline, the vendor will not provide necessary version information, or three-year cost exceeds an affordable range, the project should pause or narrow its use. Stopping is not failure; it is governance maturity.
Expansion method
First run a trial in low-risk assistive scenes and a controlled population, observe error consequences and operating burden, then expand to more regions and devices. Every expansion reassesses data, capacity, fairness, and support capability. Do not assume prior conclusions still hold.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Service Metrics: From Model Scores to Public Value
Public outcomes
Core indicators include service processing time, first-time completion rate, reduction in repeated data submission, frontline workload, weak-network reachability, and citizen satisfaction. These indicators answer whether technology actually improves the service, rather than only adding a new interface.
Model and engineering
At the model layer, track task success, citation hit, erroneous confidence, appropriate refusal, and sensitive-data leakage. At the engineering layer, track p95 latency, availability, change-failure rate, mean time to repair, device crashes, and disaster-recovery results.
Governance and cost
At the governance layer, track high-risk human review, traceability rate, policy intercepts, and complaint closure. At the cost layer, track cost per inference, unit throughput, idle resources, and three-year TCO. Every metric needs an owner and a triggered action, so dashboards are not display-only.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Success Case: More Than One Hundred Digital-Government and Smart-City Measures
Outcome background
Hong Kong completed a whole-of-government review of electronic services and, by the end of 2025, advanced more than one hundred digital-government and smart-city measures, using big data, artificial intelligence, blockchain, and geospatial analysis to improve public services. The success of this case is not a single technology. It is the combination of shared platforms, cross-department coordination, and a goal of citizen convenience.
Reusable base
Government cloud, a big-data analytics platform, digital identity, shared blockchain, chatbot services, and a unified service portal mean departments do not have to build from zero every time. Shared capability reduces duplicate investment and also sets a consistent baseline for security, identity, logs, and service availability.
Implication for AI prototypes
If a calorie app is to enter a public-health scene, it should connect to existing identity, consent, cloud, and data-exchange capability rather than building another island. The success case shows that first building a governable shared base, then accommodating multi-vendor solutions, is more sustainable than chasing a single super-platform.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Success Case: Digital Identity Upgraded from a Login Tool to a Service Portal
Adoption scale
Hong Kong’s one-stop iAM Smart digital-identity platform has accumulated more than four million registered users, supports more than 1,300 services and electronic forms, and has obtained international-standard certification for information security and privacy information management. The key to scale is that identity, signing, form-filling, and document capability can be reused by many services.
Design implication
Public AI should not manage passwords, identity-document copies, and complete personal data itself. Obtaining the minimum necessary attributes through a trusted identity platform, and using step-up authentication for high-risk operations, reduces repeated collection and impersonation. Anonymous health-education functions should not force login.
Next linkage
As a corporate digital-identity platform is expected to launch by the end of 2026, government-to-business and business-to-business services can further use enterprise verification, digital signing, pre-fill, and a document wallet. If an AI agent acts on behalf of an organisation, authorisation scope and signing evidence will matter more than natural-language ability.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Success Case: CDEG Improves Once-Only Service
How it works
CDEG lets departments or authorised bodies exchange verified data with citizen consent, handling about two million data exchanges a month. It turns “do not submit again” from a slogan into a controlled process, and retains data source, purpose, and authorisation relationship.
Lesson for health applications
If campus or community health services need age group, service eligibility, or existing appointment status, they should obtain the necessary fields through consented exchange rather than requiring a full certificate upload. Nutrition photos and health data should still be handled separately, so convenience does not expand data linkage.
Governance focus
Consent must be specific, understandable, withdrawable, and time-bounded. A data recipient cannot extend a one-time authorisation to model training or commercial use. Every exchange needs a purpose, minimum fields, a retention period, and exception reporting if citizen trust is to be maintained.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Success Case: Open Data Moves from Supply to Use
Scale progress
As of December 2025, public data downloads had grown from about five billion in 2019 to more than 80 billion. The platform provides more than 5,700 datasets, about 110 APIs, and participation by more than 2,500 data providers. This shows that stable supply, machine readability, and continuous updates can form a usage ecosystem.
Quality matters more than quantity
AI applications need a data dictionary, update frequency, licence, lineage, quality rules, and a contact person. Uploading a file is not the same as being usable. For nutrition, transport, environment, and similar data, version and timestamp are especially important, because using expired data can create real risk.
Implementation advice
Appoint a data-product owner and track API availability, field changes, error reports, and downstream impact. When opening externally, provide samples, limits, change notices, and historical versions so multi-brand cloud, academic, and enterprise teams can reuse safely.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Success Case: AI+ Public Service Starts from a Tool Catalogue
A pragmatic entry
Hong Kong uses a catalogue of AI tools and solutions covering seven common work types: digital-human customer service, meeting summaries, document processing, writing, process automation, creative promotion, and data analysis. Forums, seminars, and matching events help departments understand available options. This is more efficient than asking every department to research every model itself.
Multi-vendor governance
A catalogue should not list only features and price. It should also mark data destination, deployment mode, model source, logging capability, portability, support level, and prohibited scenes. Keep at least one alternative for the same purpose and compare on a common test set, so brand recognition does not replace evidence.
Landing sequence
Start with low-risk internal work such as drafts, summaries, and classification, requiring staff confirmation before external issue; then gradually handle cross-department processes and citizen interaction. Every tool needs an exit path, data export, and prompt-version management.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Northern Metropolis: Place AI, Education, and Health Innovation inside Spatial Planning
Development positioning
The Northern Metropolis treats innovation and technology, post-secondary education, and health and medical innovation as important functions, and emphasises planning first, infrastructure-led development, industry drive, and a people-centred approach. The coming five-year plan proposes more than 70,000 housing units and one million square metres of economic floor space, creating capacity for industry and community to grow together.
Technology opportunity
A large new district can include data pipelines, digital identity, the Internet of Things, edge computing, green buildings, and public-health services in the foundation design rather than splicing them afterwards. Connecting a university town, research facilities, industrial parks, and communities also helps close the loop of real-scene testing and talent development.
Governance reminder
A living lab cannot become a justification for unrestricted data collection. Every pilot needs a clear scope, resident communication, an exit arrangement, and independent evaluation. New-district technology should support open interfaces and multi-vendor operations, so city infrastructure is not locked into a single solution for the long term.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Hetao and Cross-Boundary Innovation: Interoperable Rules Are Harder Than Interoperable Networks
Collaboration value
The Hetao Shenzhen–Hong Kong Science and Technology Innovation Co-operation Zone uses one zone, two parks to drive research, testing, translation, and industrialisation, and together with San Tin Technopole forms the northern innovation-and-technology engine. Health technology can combine Hong Kong research, rule of law, and international connectivity with Shenzhen engineering, manufacturing, and market capability.
Data boundary
Cross-boundary cooperation first draws data classes and flows, distinguishing public data, general business data, personal data, important data, and research samples. For each class, make legal basis, storage location, access roles, encryption, approval, and deletion explicit. A cooperation agreement cannot replace concrete controls.
Standard-contract experience
The GBA standard contract for cross-boundary flow of personal information began as a pilot in 2023 and, from November 2024, expanded to all GBA sectors. Engineering teams still need to land contract requirements in API fields, logs, permissions, and incident reporting. A legal document itself does not automatically become a secure system.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Smart Health: From Hospital Digitisation to Community Prevention
Policy direction
The five-year plan lists primary care, chronic-disease prevention and early diagnosis and treatment, smart healthcare, health-information infrastructure, and Chinese–Western medicine collaboration as priorities. Smart health therefore cannot concentrate only in large hospitals. It must also support continuous service for communities, older people, carers, and weak-network areas.
Application layers
The low-risk layer can provide health education, appointments, reminders, and general dietary information. The medium-risk layer can help professionals organise data and find anomalies. High-risk diagnosis and treatment must remain with qualified personnel. Different layers use different data, models, review, and acceptance standards.
Where the calorie case sits
On-device plate recognition is best placed at the health-education and behaviour-recording layer, helping users understand food composition and portion size, not making disease judgements. Any connection to electronic health records must separately complete clinical, security, privacy, and professional-liability assessment.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Smart Villages and Remote Service: High Public Value from Small Models
Scene demand
Smart-village pilots include public Wi-Fi, telemedicine, electronic payments, illegal-dumping and flooding detection, and early wildfire discovery assisted by robots and artificial intelligence. These scenes share unstable networks, limited maintenance resources, and the importance of field response time.
Architecture choice
Place initial detection on edge devices, retaining local rules and offline operation. The central cloud handles model management, cross-area analysis, and expert collaboration. Devices must support remote inventory, update, rollback, and disablement, and must preserve event sequences when communications break.
Operating experience
What remote deployments most often neglect is power, waterproofing, dust protection, spare parts, field training, and alert fatigue. Procurement scoring should include five-year maintainability and replacement cycles, not only model accuracy and a one-off quote.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Multi-Cloud and Hybrid Cloud: Do Not Bind the Brand; Bind the Standard
Selection philosophy
Different brand cloud services have different strengths in model ecosystems, data analysis, edge management, security, and regional coverage. Government does not need to split workloads evenly. It should choose the most suitable location by data residency, service level, cost, capability maturity, and existing talent.
Portable design
Use containers, standard APIs, infrastructure as code, open data formats, and externalised configuration, separating identity, logs, model interfaces, and business rules. Portability does not mean zero-cost move at any time. It means key dependencies can be replaced within reasonable time and budget.
Avoid fake multi-cloud
If two clouds appear together only on a slide, but data, monitoring, talent, and drills all concentrate on a single vendor, it is still a single point of dependence. Real multi-cloud needs explicit failover, data consistency, a common security baseline, and regular exit drills.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Enterprise Capability Stack: Every Layer Has Clear Accountability
End-to-end layers
The device layer handles camera, masking, offline inference, and device state. The API layer handles identity, rate limiting, and protocol. The data layer handles master data, vector indexes, versions, and quality. The model layer handles registration, evaluation, release, and rollback. The operations layer handles monitoring, incidents, and cost.
Shared platforms
Kubernetes fits workloads that need consistent deployment and long-lived service. Serverless fits event-driven and bursty traffic. Object storage fits versioned assets. Content delivery fits model and static-knowledge distribution. Choice should be decided by accountability and load characteristics, not by technology fashion.
Minimum documentation
Every component needs an owner, SLO, capacity assumptions, RTO, RPO, cost ceiling, data classification, external dependencies, update method, and exit plan. A component without these materials should not enter the formal architecture.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
MLOps: Device Cohorts, Canary Release, and Drift Management
Release unit
On-device models cannot be managed by application version alone. They must form a traceable release unit by device type, model format, quantisation method, knowledge version, and rule version. Any one change can alter final behaviour.
Canary strategy
Release first to internal and low-risk populations, set health indicators and automatic stop conditions, then expand gradually. If crash rate, latency, erroneous confidence, or power draw exceeds the threshold, the system should stop expansion and roll back, without waiting for large numbers of user complaints.
Drift monitoring
Monitor input drift from seasonal dishes, new packaging, camera hardware, and usage habits, and also monitor output change after knowledge updates. Drift is not retraining as soon as a number changes. First decide whether it affects public outcomes and particular groups.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Security Design: Protect Models, Devices, and the Supply Chain
On-device protection
Use secure storage, credential binding, program-integrity checks, and remote attestation where needed, reducing tampering of model files, rules, or API keys. When a device is lost, credentials can be revoked and sensitive caches cleared, without depending on the user to act.
Service protection
APIs implement least privilege, rate limiting, input validation, malicious-file scanning, and anomaly detection. Model and knowledge updates must be signed, verified, and released in batches, so supply-chain contamination does not affect every device at once.
Incident readiness
Build playbooks for model substitution, data leakage, prompt injection, vendor interruption, and bad updates. Drills must include technical repair, business degrade, management notification, citizen communication, and evidence preservation.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Privacy Engineering: Data Minimisation Is an Architecture Capability
Collect the least
If recognition only needs part of the plate, guide cropping in the camera interface. If statistics only need age group, do not collect date of birth. If error analysis only needs a feature summary, do not keep the original image. Every data item removed also lowers leakage, compliance, and operating cost.
Purpose limitation
Health education, service analysis, model improvement, and research are different purposes and cannot be covered by one vague consent. Users should be able to use the basic service without participating in model training, and there must be an executable deletion process after consent is withdrawn.
Verifiable controls
Privacy requirements must become tests: masking accuracy, logs without original images, deletion on schedule, permission changes taking effect immediately, and complete export content. Policy documents without technical verification cannot prove that data minimisation has actually landed.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Procurement and Acceptance: Buy Capability, Not a Demo
Tender requirements
Specifications should describe business outcomes, risk boundaries, interfaces, data rights, observability, and exit requirements, avoiding lock-in to a particular model name or proprietary service. Vendors may propose different technology combinations, but they must pass a common test set and field scenarios.
Acceptance combination
Accept accuracy, refusal, security, fairness, latency, power draw, availability, cost, disaster recovery, logs, and complaint process together. A single average score hides high-risk failures. Set non-negotiable hard thresholds.
Contract protection
Make explicit rights in data and derivatives, model-update notice, subcontractors, vulnerability patching, service termination, data export, deletion proof, and handover period. If exit cost is opaque, a low bid can become an expensive long-term dependency.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Organisation Model: Government, Industry, Academia, Research, and Investment in Concert
Role split
Government defines the public problem, rules, and adoption scenes. Enterprises handle engineering and ongoing operations. Universities and research institutions provide methods, evaluation, and talent. Investors support scalable translation of results. No single party should alone decide the success standard of a high-risk system.
Shared language
Use use cases, data contracts, service metrics, a risk register, and architecture decision records as cross-boundary communication tools. Research accuracy, commercial revenue, and public value are different goals and need an explicit ranking at the start of the project.
Knowledge transfer
Contracts require documents, training, joint duty, and delivery of code and configuration, so the public-service team has basic judgement and takeover capability. Outsourcing can supplement capability, but it cannot outsource final accountability.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Talent and Digital Literacy: Teach Users to Question AI
Capability layers
Leaders need to understand risk and the investment portfolio. Product owners need to define public outcomes. Engineers need data, models, security, and operations. Frontline staff need to recognise uncertainty, correct errors, and start a human process. Training cannot teach only prompts.
Practical training
Use real but anonymised failure cases for tabletop drills, including a model that is confident but wrong, conflicting data sources, vendor interruption, a bad update, and a citizen complaint. Participants decide stop, rollback, notification, and reply, not only how to operate the interface.
Ongoing mechanism
Build communities of practice, a tool catalogue, shared test sets, technology-matching events, and quarterly case reviews. This is consistent with smart-city experience of civil-service technology training and cross-department coordination, turning individual expert knowledge into organisational capability.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Green AI: Bring Compute Cost into Public Accountability
Link to city goals
Hong Kong’s direction is carbon neutrality before 2050, and it is advancing building energy efficiency, low-carbon transport, circular economy, sustainable aviation fuel, and hydrogen. AI projects should also measure the resource cost of compute, storage, network, and device refresh, rather than assuming digitisation is inherently green.
Engineering choices
Prefer small models suited to the task, on-device pre-screening, caching, batch processing, model quantisation, and automatic shutdown of idle resources. High-energy training needs a clear improvement goal and a stop condition, so disproportionate compute is not spent on a tiny score gain.
Procurement metrics
Require vendors to report resource use, hardware life, energy region, device replacement, and e-waste arrangements. Green metrics need not replace service quality, but they should be assessed together with cost, latency, and public value.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
City Resilience: Treat Failure as Inevitable, Not Exceptional
Threat range
Extreme weather, network interruption, power problems, vendor failure, cyber attack, and a bad release can all make digital services fail. The five-year plan emphasises city safety, cross-department warning, emergency plans, and rapid post-disaster recovery. AI systems must enter the same resilience system.
Degrade levels
Set four modes: full service, restricted service, offline basic service, and human substitute. Each level makes available functions, data freshness, accountable owner, and citizen prompts explicit, so decisions are not improvised at failure time.
Drills and after-action
Regularly drill regional-cloud failure, identity-platform unavailability, a bad model update, and a surge of requests. After-action review asks not only recovery time, but also whether critical livelihood services were retained, whether wrong outputs were produced, and whether cross-department notification was clear.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.
Closing: Advance AI with Restraint, Evidence, and Public Accountability
Redefine success
Finishing an on-device calorie app in three days can prove the team can integrate quickly. It cannot prove medical effectiveness, regulatory fitness, or large-scale reliability. True success is knowing clearly what can be used, what cannot, how citizens are protected when errors occur, and whether continued investment is justified.
City-scale implication
Hong Kong’s successful experience in digital identity, CDEG, open data, shared government platforms, more than one hundred smart-city measures, and AI-ecosystem building shows that long-term capability comes from a shared base, cross-department governance, and continuous operations—not from a single model.
Action principles
Start from low-risk assistive scenes. Build trust with on-device data minimisation, separation of model and knowledge, uncertainty, human veto, evidence chains, multi-vendor standards, and exit drills. Let AI become a stable, trustworthy, restrained, and sustainable public capability, not a short-lived demonstration.
Field prompt: Turn this page into an architecture-review question. Require the team to answer with evidence, and log unanswered items into the next worklist.