← Financial Cloud Cloud Cloud Club · AWS Re:cap

Vertex Macro | Financial Cloud Cloud · AWS Re:cap

AWS Re:cap 03: Large-Model Capability Evaluation and a Method for Landing AI Projects

Speaker: Government data

Session: 03

Session
Summit Dev Lounge2026 Re:cap
01 Encode Architecture as Steering for AI Agents
Summit Dev Lounge2026 Re:cap
02 Agent Harness Is the Real Engineering Moat
Summit Dev Lounge2026 Re:cap
03 Ask Observability Data in Plain Language
Summit Dev Lounge2026 Re:cap
04 Serverless AR Game with Bedrock AgentCore
Summit Dev Lounge2026 Re:cap
05 Multi-Agent Quant Backtesting on AgentCore
Summit Dev Lounge2026 Re:cap
06 Blog to Slides in Three Minutes with Kiro
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 AWS Compliance with Terraform
AWS Community Day Hong Kong 2025 Re:cap
03 Beginner to Builder An AWeSome Cloud Journey
AWS Community Day Hong Kong 2025 Re:cap
04 Team-First Serverless Engineering with Laravel & Bref
AWS Community Day Hong Kong 2025 Re:cap
05 Event Opening - AWS Community Day Hong Kong 2025
AWS Community Day Hong Kong 2025 Re:cap
06 Agent-to-Agent: Building Interoperable AI on AWS
AWS Community Day Hong Kong 2025 Re:cap
07 Utilize another telemetry data for faster improvement with AI agent
AWS Community Day Hong Kong 2025 Re:cap
08 Graduating from Vibe Coding: Spec-Driven Development with Kiro
AWS Community Day Hong Kong 2025 Re:cap
09 Automated Testing using MCP & AI Agents
AWS Community Day Hong Kong 2025 Re:cap
10 Modernizing Telecom Security ML Powered Approach
AWS Community Day Hong Kong 2025 Re:cap
11 Rethinking GenAI Agent: RAG & MCP
AWS Community Day Hong Kong 2025 Re:cap
12 Disaster and Emergency Response with TAK and AWS
AWS Community Day Hong Kong 2025 Re:cap
13 Rethinking Serverless Application Workflows from a Testing Perspective
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 AI-Powered Global Pure-Alpha Macro Trades on AWS: Revolutionizing Risk-Adjusted Asset Returns
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 Modern Trade Lifecycle: Trading to Settlement
FSI Recap
02 Goldman Sachs: Fast Track your applications onto Cloud - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group: Building an Effective Log Management Solution on AWS
FSI Recap
04 FSI Meetup 2025 Q4 - Brex Database Disaster Recovery
FSI Recap
05 FSI Meetup 2025 Q4 - A Graviton Migration Success Story
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel Modern Data Platform
FSI Recap
07 FSI Meetup 2025 Q4 - Financial Transaction Data Reconciler PayPal
FSI Recap
08 FSI Meetup 2025 Q4 - Scaling Resilience
FSI Recap
09 Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances
FSI Recap
10 Advanced Agentic AI Design Patterns
FSI Recap
11 Build New Modern Apps on AWS
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent Recap (IND3312)
AWS re:Invent 2025
02 Building the Future Trading Platform Leveraging AI and AWS
AWS re:Invent 2025
03 Trading Innovation: Jefferies' AI Assistant on Amazon Bedrock (IND3315)
AWS re:Invent 2025
04 How FSI Revolutionized HFT Analytics with Agentic AI (GBL302)
AWS re:Invent 2025
05 Improving Distributed Systems with Amazon Time Sync Featuring Nasdaq
AWS re:Invent 2025
06 Amazon Aurora HA and DR Design Patterns for Global Resilience (DAT442)
AWS re:Invent 2025
07 Building Agentic AI: Amazon Nova Act and Strands Agents in Practice (DEV327)
AWS re:Invent 2025
08 Deep Dive into Amazon Aurora and Its Innovations (DAT441)
AWS re:Invent 2025
09 Deep dive on Amazon S3 (STG407)
AWS re:Invent 2025
10 Nasdaq: Build Resilient Infrastructure for Global Financial Services (HMC327)
AWS re:Invent 2025
11 What's New with AWS Lambda (CNS376)
AWS re:Invent 2025
12 Spec-Driven Development with Kiro (DEV314)
AWS re:Invent 2025
13 Amazon's finops: Cloud cost lessons from a global e-commerce giant (AMZ308)
AWS re:Invent 2025
14 Tick to trade latency trading platforms on aws
AWS re:Invent 2025
Government data
01 The AI Era: The Boundary Between Development and Design Is Disappearing
Government data
02 On-Device Multimodal AI and Smart-City Practice
Government data
03 Large-Model Capability Evaluation and a Method for Landing AI Projects
Government data
04 Controlled End-to-End Automation of Government Development with Cloud Agents
Government data
05 A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
Government data
06 AI-Driven Macro Quantitative Research and Smart Governance
Government data
07 Authorized Operation of Public Data and Smart-Government Practice
Government data
08 Putting Data Assetization into Practice: Rights, Compliance, Engineering Governance, and Digital-Government Cases
Government data
09 AI for Mental-Health Public Welfare: Governance, Architecture, and Practice of a Trustworthy Platform
Government data
Amarathon 2025 Recap
01 A Developer’s Roadmap to Architecting for Agents
Donnie Prakoso
02 Amazon Bedrock Data Automation
Hafiz Syed Ashir Hassan
03 Multi-Agent on AgentCore
Tan Xin
04 Building Agentic AI Nova Act and Strands Agents in Practice
Haowen Huang
04 Accelerating Migration Projects with Kiro using Spec-Driven Development
Sanchit Dilip Jain
06 From Matching to Understanding: Personalized AI Search Practice Driven by AgentCore Memory
Liu Cao
07 Observe to Optimize – LLM Observability to AIOps Turning real-time insights into intelligent automation
Jimmy Soh
08 Deploying TEAM and Building the Best Engineering Team
Yuji Oshima
09 Five Hard Lessons from Five Years of So-Called Serverless Databases
Renato Losio
14 What if AI does my job How Q Developer CLI and Kiro have changed my daily routine
Miguel Angel Muñoz
16 Velocity with Vigilance: Security Essentials for Amazon Bedrock Agent Development
Brian Tarbox
26 Run OSS LLMs on a Single H100 Smarter, Cheaper, Faster
Adit Modi
28 A Modern Unified Metadata Architecture: New Approaches to Breaking Down Data Silos
Shaofeng Shi
29 Serverless MediaOps: Automating Video Workflows with AI on Amazon Web Services
Luis Valdivia
30 Architecting for Efficiency and Reliability with Performance Testing at Scale
Luis Guirigay
31 Connecting the World Through Open Source: Practical Journey of Technology, Community and Global Developer Relations
Richard Lin
33 Building Streaming Iceberg Tables for Real-Time Logistics Analytics
Fahad Shah
34 Accelerating Large-Scale Robot Strategy Training: An Automated Closed-Loop Architecture Based on Kiro, Trainium, and EKS
Junjie Tang
35 From Vibe to Viable with spec driven development
Ricardo Sueiras
36 Making Cloud Cost Analysis Smarter: Building FinOps Intelligent Agents with Strands and AgentCore
Xiaofei Li
37 Transform Conversational Agentic AIOps for K8s Using CNCF Kagent, K8sGPT, and Nova Sonic
Shaoyi Li

The Real Starting Point for Landing Large Models: First Define Acceptable Error

Decision premise

A government project cannot start from a model name. First write down the service population, the public problem, the consequences of error, and the human substitute. A wrong statutory answer in a policy Q&A, a missed sensitive item in document review, and a lost restriction in an internal summary have entirely different consequences, so they cannot share one abstract accuracy threshold.

Three boundaries

First define the answerable range, the must-refuse range, and the must-human-review range. For each error type, mark impact, reversibility, time to discovery, and remediation cost, then decide the control strength that the model, data, and process must reach.

Field practice

Require business, legal, data, security, and frontline staff to review together the ten cases most likely to go wrong. If the team cannot form a consensus on error consequences, do not procure a model yet. The definition of success should land on processing time, first-time completion rate, error rate, frontline burden, and public satisfaction.

Evaluate Capability before Choosing a Brand

A brand cannot replace evidence

Open weights, hosted APIs, proprietary cloud, and on-premises models each have strengths, but leaderboards, parameter counts, or market noise cannot directly prove fitness for government work. Models must be compared on local corpora, real documents, permission rules, and acceptable latency. Brand is only one piece of supply-risk information.

Seven faces of the scorecard

Build a weighted scorecard across task quality, Chinese capability, tool use, security, latency and throughput, deployment efficiency, and licence and supply chain. Every point needs test samples, repeat counts, and failure evidence, avoiding scores from subjective trial use.

Selection discipline

Keep at least one alternative, and record the reason for the choice, the conditions for abandonment, and the trigger for re-evaluation. If the model is upgraded, the price changes, the licence is updated, or data-residency requirements change, the original conclusion lapses automatically and the baseline must be re-run.

Split Model Capability into Verifiable Tasks

Do not test chat impression

“Can reason,” “supports multimodal,” and “can use tools” are not acceptance descriptions. Split them into atomic tasks such as policy-validity judgement, conflict detection among provisions, table-field extraction, image–text checking, function-parameter generation, citation location, and refusal of overreach.

Task contract

For each task, write the input format, allowed data, expected output, tolerated error, latency ceiling, cost ceiling, and human-takeover point. Measure composite flows in segments; otherwise, when the final answer is wrong, it is impossible to tell whether retrieval, the model, a tool, or a business rule caused it.

Reusable outcomes

Atomic tasks can form a common test set across vendors, supporting procurement acceptance, version regression, and canary release. Mature teams keep failure samples, not only average scores, because rare but high-consequence errors most deserve continuous monitoring.

Six-Stage Landing Method: From Problem to Operations

Stages one to two

First complete problem definition and data preparation, confirming public value, roles, data legality, quality, permissions, and retention period. Without a reliable data source or a human process, even a strong model cannot form a sustainable service.

Stages three to four

Establish a human-acceptable baseline, then run offline evaluation. Compare first against search, rules, or the existing process, and only then introduce the model. Evaluation covers task success, citation, refusal, sensitive data, latency, throughput, and cost at the same time.

Stages five to six

Enter operations only after canary release, setting automatic stop, rollback, version tracking, drift monitoring, incident handling, and periodic re-evaluation. Each stage has an approver, required evidence, and an exit condition. A PoC must not skip governance gates because a demonstration succeeded.

Golden Dataset: Government Must Own Its Own Examination Paper

Sample sources

The dataset should come from public, authorised, anonymised, or synthetic real-work situations, including common, difficult, boundary, expired, conflicting, sensitive, and malicious cases. Collecting only standard answers overestimates system capability.

Annotation governance

Every sample stores the question, basis, expected behaviour, acceptable variants, refusal requirements, risk level, and reviewer. When policy changes, update the answer and retain historical versions, so today’s rules are not used to mis-judge yesterday’s system.

Asset sovereignty

The test set, prompt assets, knowledge base, and audit data should be held by the institution, not locked on a single vendor platform. Vendors may help build them, but must not block export with a proprietary format. The golden dataset is a core asset for long-term procurement bargaining and model switching.

Reasoning Capability: Test Process Constraints, Do Not Worship Thinking Text

Test focus

Government reasoning must verify rule applicability, exception conditions, evidence completeness, and conclusion consistency, not require the model to output lengthy thinking. Structured basis, cited passages, applicable dates, and open items can be required so reviewers can check quickly.

High-risk samples

Add situations where provisions conflict, attachments are missing, a policy has lapsed, local rules differ, and the question’s premise is wrong. A strong system should point out the gap, ask for supplementation, or refuse to conclude, rather than fluently filling in the unknown.

Engineering experience

Splitting a complex problem into retrieval, rule judgement, calculation, and text generation reduces unexplained error. When precise calculation is needed, call a controlled tool; do not let a language model do mental arithmetic. When a policy conclusion is needed, force citation of a valid source.

Multimodal Evaluation: A Document Is Not One Picture

Difficulties in government documents

Scans, overlapping seals, tables, handwritten annotations, headers and footers, watermarks, and multi-column layout break reading order. Multimodal capability must separately test OCR, layout understanding, cross-page association, seal position, and image–text consistency.

Segmented pipeline

First perform document-quality detection, page classification, and text-and-coordinate extraction, then let the model understand semantics. Original image, OCR text, layout blocks, and final conclusion must stay linked so a human can return to the original page to verify.

Acceptance method

Do not only ask whether content was read. Test field completeness, table relationships, page-number citation, refusal on low quality, and masking of sensitive content. If the model is over-confident about a blurred seal, the system should require a rescan or human confirmation.

Tool Use: The Model May Only Propose a Call; the System Owns Authorisation

Separation of authority

Generating a function name and parameters does not mean the model has the right to execute. Real authorisation is decided by identity, role, data scope, transaction limit, and approval process. Model output must pass a whitelist and a policy engine.

Safety guardrails

Set different risk levels for query, write, delete, payment, notification, and similar tools. High-risk operations use preview, second confirmation, dual review, or draft-only generation. Tool returns must also prevent prompt injection and sensitive-data reflux.

Audit evidence

Retain model version, prompt version, tool, parameter summary, authorisation result, execution reply, human confirmation, and final state. An incident investigation must be able to answer who authorised, what the model recommended, what the system executed, and whether data were changed.

Long Context: Being Able to Fit It In Does Not Mean Being Able to Find It

Common misconception

A claim of very long context only means a large number of tokens is accepted. It does not guarantee that the model will find the key provision in the middle, correctly handle multi-version documents, or keep citations consistent. The longer the context, the higher cost, latency, and interference may also be.

Stress tests

Place key evidence at the beginning, middle, and end, mix in similar clauses, old versions, and irrelevant attachments, and test recall, citation, and refusal. Then compare three strategies: stuffing the full text, layered summaries, and retrieval augmentation.

Practical principle

Cut context by document structure, permissions, and task, provide only necessary passages, and retain source indicators. Long context fits cross-passage integration. It should not replace data governance, version control, and reliable retrieval.

Chinese Government Capability: Not a Translation Score

Language layers

Evaluation should cover Simplified and Traditional Chinese, fixed official phrasing, legal wording, abbreviations, place names, written Cantonese, numbers and dates, measure words, and mixed Chinese–English layout. General Chinese fluency is not the same as accurate handling of government materials.

Situational difficulties

The same word may have different definitions in different departments. Policy names may have colloquial forms. Table fields often omit the subject. The model must combine a terminology base, department data, and document versions, and must not guess on its own.

Improvement method

First establish terminology, entity, and format rules, then use retrieval to supplement local knowledge. Consider fine-tuning only for stable, high-volume repeated errors. For dialect input, retain the original, standardise the result, and allow human correction, so meaning is not lost in conversion.

RAG First: Give the Answer a Basis

When to use it

Policy Q&A, institutional lookup, document review, and internal knowledge support usually start with retrieval augmentation, because knowledge updates, citation is required, and permissions are involved. Fine-tuning all knowledge into the model is slow to update and hard to trace.

Core components

Complete RAG includes data cleaning, chunking, permission tags, hybrid keyword-and-vector retrieval, re-ranking, context assembly, citation, and no-answer handling. An error in any link can produce a wrong answer that looks sourced.

Field sequence

First use human-acceptable keyword search as the baseline, then add hybrid retrieval and re-ranking; change only one variable each time. If the answer is wrong, first check whether the evidence was found, then check whether the model used the evidence correctly.

Hybrid Retrieval: Keywords and Vectors Each Guard a Gate

Keyword strength

Statute numbers, organisation names, proper terms, exact dates, and fixed phrases fit keyword retrieval and give stable, explainable hits. Vectors alone may mix documents that are semantically close but legally different in effect.

Vector strength

Users often describe a problem in everyday language and may not know the formal policy name. Vector retrieval can find semantically close content and cover synonyms, different question forms, and cross-language expression.

Fusion practice

After two-way recall, re-rank by permission, validity period, document level, and relevance. When needed, keep multiple conflicting sources for later rule judgement. Evaluation should separately count recall failure and re-ranking failure, not only look at the final answer.

Chunking and Indexing: Document Structure Decides Retrieval Quality

Chunking principle

Do not chunk only by a fixed character count. Retain chapter, section, article, paragraph, table, and attachment relationships, and use title, issuing body, effective date, region, classification, and permission as metadata.

Cross-passage handling

Provisions often cite earlier definitions or attachments. A single-passage hit may lack necessary context. Parent–child blocks, neighbouring passages, and citation relationships can be built so that, after a precise passage is recalled, the minimum necessary background is added.

Quality checks

Sample-check sentence breaks, table splits, page numbers, character encoding, and duplicate content. Index updates use versioning and atomic switchover, so the service does not read old and new policy at the same time.

Re-ranking and Citation: Credibility in the Last Mile

Purpose of re-ranking

First recall aims not to miss. Re-ranking puts truly usable, valid, and permission-allowed evidence in front. Ranking features are not only semantic similarity; they also include legal effect, date, region, document type, and question intent.

Citation standard

An answer must at least provide document name, valid version, clause location, and a verifiable fragment. The citation must actually support the conclusion, not merely look related under the same topic.

Failure handling

If sources conflict, there is no valid document, or only low-permission content exists, the system should state the gap and hand off to a human. Citation hit rate and citation-support rate must be measured separately: the former finds a source; the latter judges whether the source is enough to support the answer.

Prompt Engineering: Write Business Rules as a Testable Asset

Structural design

The system prompt should include role, task, allowed data, prohibitions, output format, refusal conditions, and human referral. Project prompts are separated from general brand tone, so a style change does not accidentally change safety behaviour.

Version governance

Treat prompts as code: bring them into version control, review, test, release, and rollback. Every change records reason, affected tasks, evaluation results, and approver. Do not make ad-hoc production-interface edits.

Avoid over-prompting

An overly long prompt raises cost and interferes with knowledge content. Stable permissions, calculation, and policy judgement should be executed by programs and a rules engine. The prompt only carries the behavioural constraints the model needs to understand.

The Choice Order of Rules, Retrieval, Fine-Tuning, and Agents

Determinism first

Fixed calculation, format checks, permissions, validity period, and hard policy first use programs or rules. When knowledge updates frequently and citation is required, use retrieval. These two steps often already solve most government needs.

When to fine-tune

Evaluate fine-tuning only when the model repeatedly shows style, structure, or specific classification errors on a stable task, and when there is enough high-quality labelled data. Fine-tuning cannot repair a wrong knowledge source or a chaotic process.

When to use agents

Use an agent only when the task needs multi-step planning and tool collaboration, and limit steps, tools, budget, and write operations. Agents increase autonomy and also increase unpredictability and audit difficulty. They should not be the first architecture option.

Fine-Tuning Decision: Trade Stable Error for Controllable Improvement

Data threshold

Fine-tuning data must represent real input, be correctly labelled, cover boundaries, and have usage rights. Taking historical human answers straight into training may freeze expired policy, personal preference, and private data together.

Controlled experiment

Keep a test set that did not participate in training, and compare the base model, prompts, RAG, and the fine-tuned version. Besides target-task improvement, also check general-capability regression, safety bypass, Chinese format, and refusal behaviour.

Operating accountability

The fine-tuned weights, base version, data version, and parameters must be registered. A base-model upgrade cannot assume the old adaptation still works; retrain or run a full regression. Security patching and capacity accountability also move to the using party.

Agent Systems: Lock Autonomy inside an Observable Process

Process modelling

First draw the task as states, tools, decision points, stop conditions, and human nodes. The model may choose the next step within a limited range, but must not enlarge the goal, add data sources, or bypass approval on its own.

Budget and stop

Set maximum steps, tokens, time, external calls, and retry ceilings. Stop immediately on insufficient permission, conflicting evidence, tool failure, or cost overrun, and return completed steps and items pending human action.

Observability

Trace each step’s input summary, model decision, tool result, state change, and cost. Recording only the final answer cannot explain why an agent detoured, repeated a call, or produced a wrong operation.

Policy Q&A and Document-Review Assistants: Combined Field Practice

Step one: baseline

First build a human-acceptable search baseline, using public or authorised policy, and define validity period, department, region, and permission. Select real questions and check the time humans need to find an answer and common missed items.

Step two: enhancement

Add hybrid retrieval, re-ranking, passage citation, and structured output. Test expired policy, conflicting rules, sensitive matters, table attachments, and no-answer cases. The model must not fill missing documents with common sense.

Step three: red team

Test with prompt injection, overreach queries, data exfiltration, forged sources, and malicious attachments. Every input, output, model, prompt, source, tool, human review, exception, and cost forms an evidence chain.

Open Weights Are Not Zero Risk

Autonomy gained

Open weights can support local deployment, deep evaluation, customisation, and replacement of the inference framework, reducing some API lock-in and helping keep sensitive data in a controlled environment.

Responsibility assumed

The using party also assumes licence interpretation, vulnerability patching, model-source verification, capacity planning, inference security, content governance, version compatibility, and shutdown handling. Without a mature operations team, self-hosting may be riskier than a hosted service.

Supply-chain verification

Retain source, hash, licence version, dependency list, and build process; scan containers and packages; restrict model-download channels. Run a full evaluation before any new weights go live. Do not treat nearby names in the same series as equivalent.

Licences and Legal Terms: A Hard Gate on the Technical Scorecard

Items to check

Confirm commercial use, derivative models, redistribution, use restrictions, branding requirements, data use, and termination terms. Licences for the model, datasets, libraries, and service APIs must be checked separately.

Change management

Licence terms may change by version. Procurement and architecture decisions must pin the version used and establish periodic review. If terms affect government use or data rights, a migration plan is required.

Practical delivery

Build a readable licence inventory and risk classification, jointly signed by legal, procurement, and technology. Do not treat “open source,” “downloadable,” or “free trial” as proof of lawful long-term operation.

Deployment-Mode Comparison: Self-Hosted, API, Proprietary Cloud, and Hybrid

When self-hosting fits

Self-hosting has value when data are highly sensitive, load is stable, deep control is needed, and the organisation has GPU, platform, security, and model-operations teams. Up-front investment, idle capacity, patching, and talent cost cannot be ignored.

When hosting fits

When demand changes quickly, value is still being validated, or strong model capability is needed quickly, a hosted API can lower start-up cost. Data use, region, logs, service level, price change, and exit must still be checked.

Hybrid design

Sensitive pre-processing, identity, and rules stay on a private network or on premises; general inference uses multiple clouds on demand. Peak and disaster recovery can use an alternate path. Hybrid is not stacking platforms. It is allocating work by accountability boundary.

Model Gateway: A Unified Control Plane for Multi-Brand Capability

Unified interface

The gateway encapsulates different vendors’ requests, authentication, rate limiting, retry, routing, content safety, and logs. Applications do not bind directly to a particular SDK. Model differences are managed through capability descriptions and adapters.

Policy routing

Choose the model by data class, task, cost, latency, availability, and region. Highly sensitive content can only take a controlled deployment; low-risk summaries can use a lower-cost option. On failure, route to an evaluated alternative model.

Avoid the lowest common denominator

Unification does not mean flattening distinctive strengths. Keep controlled extensions for vendor-specific capability, while ensuring the core flow remains replaceable. Routing policy, prompts, and evaluation data are held by the institution.

GPUs and Inference Capacity: Reverse-Engineer Architecture from the Peak

Load profile

First measure request length, output length, concurrency, time of day, task priority, and latency targets. Average traffic cannot represent licensing peaks, sudden events, or batch document review.

Efficiency measures

Use dynamic batching, KV cache, quantisation, model parallelism, request queues, and small-model offload. Simple classification need not use the largest model. For long-document tasks, limit input and process in segments.

Capacity governance

Set a capacity ceiling, queue time limit, degraded model, batch window, and priority. Regularly run stress and failure tests, assessing GPU utilisation, cost per successful task, and timeout rate together.

Enterprise Data Platform: Knowledge Needs Catalogue, Lineage, and Quality

Data as product

Policies, guides, cases, forms, and frequently asked questions all need an owner, source, validity period, permissions, quality rules, and an update commitment. A lakehouse or other platform is only a carrier. Governance accountability cannot be handed to storage technology.

Lineage tracking

It must be possible to go from an answer back to the index block, the cleaned result, the original document, and the published source, and also to find affected indexes and answers when an original document is withdrawn. This is the basis of correction, audit, and impact analysis.

Data quality

Monitor missing fields, duplicates, expiry, encoding, parse failures, and permission mismatches. Quality anomalies can block an index release. They must not wait to be discovered after the model answers wrongly.

LLMOps: A Model Upgrade Is a Production Change

Registration requirements

Record model source, version, quantisation, inference framework, prompts, tools, knowledge snapshot, evaluation, and deployment environment. Writing only a market name is not enough to reproduce behaviour.

Release process

After offline regression passes, move to shadow traffic, internal canary, a small share of users, then gradual expansion. Each stage sets stop conditions for erroneous confidence, refusal, latency, cost, and safety.

Rollback capability

Retain the previous stable version of the model, prompts, index, and rules, and test one-click or controlled rollback. Data-structure changes must be backward compatible; otherwise, rolling back the model cannot restore the service.

Observability: From Tokens to Task Success

Four layers of signal

Infrastructure watches GPU, CPU, memory, and network. Service watches latency, errors, and availability. Model watches citation, refusal, and safety. Business watches task completion, human takeover, and public outcomes.

Distributed tracing

Use a consistent trace identifier to connect entry, retrieval, re-ranking, model, tools, and human process. Logs use summaries and de-identification, so observability does not copy sensitive data.

Action thresholds

Every metric must connect to an alert, an owner, and a playbook. A dashboard itself does not create reliability. Without rules to stop a release, scale out, switch, or intervene with a human, data is only watched.

Security Baseline: Identity, Keys, Network, DLP, and Red Team

Least privilege

Use workload identity rather than long-lived keys. Authorise model service, data, vector indexes, and tools separately. KMS manages keys, a private network restricts flows, and management operations use step-up authentication.

Data protection

DLP detects personal and sensitive information in input, logs, and output; necessary data are masked or replaced first. Content safety must distinguish illegal, sensitive, overreach, and inappropriate advice. A single blacklist cannot handle every risk.

Continuous red team

Test prompt injection, data exfiltration, tool abuse, overreach retrieval, model supply chain, and denial of service. After a fix, add the case to the regression set so the next upgrade does not repeat it.

Evidence Chain and Auditable AI

What to record

Input summary, identity and permissions, model and prompt versions, retrieval sources, tool calls, rule hits, human review, final output, exceptions, latency, and cost form a complete event.

Minimised retention

Not all original text needs long-term retention. By data class, retain hashes, source indicators, structured decisions, and necessary fragments. Original sensitive content uses a shorter period or does not land on disk.

Audit uses

The evidence chain supports complaints, version comparison, vendor acceptance, cost reconciliation, security threats, and policy correction. If the system’s behaviour at the time cannot be reconstructed, the model cannot be used in a high-accountability process.

Success Case: Hong Kong Digital Government Shared Base

Scale outcomes

Hong Kong completed a whole-of-government review of electronic services and, by the end of 2025, advanced more than one hundred digital-government and smart-city measures, using big data, artificial intelligence, blockchain, and geospatial analysis to improve services.

Platform method

Next-generation government cloud, a big-data analytics platform, digital identity, shared blockchain, chatbot services, and a unified portal let departments reuse identity, security, data, and operating capability rather than rebuilding them on every project.

Methodological implication

Landing large models should also first build shared capability such as a model gateway, data governance, evaluation, logs, and tool authorisation, then accommodate Google, AWS, Tencent Cloud, ByteDance, and other technology ecosystems. Success comes from standardised governance, not from a single brand doing everything.

Success Case: Digital Identity and Corporate Identity

Citizen service portal

The one-stop iAM Smart digital-identity platform already has more than four million registered users, supports more than 1,300 services and electronic forms, and has obtained international-standard certification for information security and privacy information management.

Requirements of the agent era

When an AI assistant acts for a person or an enterprise, it must know the agent, the principal, the authorisation scope, the validity period, and signing responsibility. The model cannot infer authorisation from conversation alone.

Corporate-service extension

A corporate digital-identity platform is expected to launch by the end of 2026. Its core includes enterprise verification, digital signing, pre-fill, and a document wallet. This builds a trusted identity base for controlled government-to-business and business-to-business AI agents.

Success Case: CDEG and Once-Only Service

Verified scale

CDEG lets departments or authorised bodies exchange verified data with citizen consent, handling about two million exchanges a month and reducing repeated citizen submissions.

Value for AI

The model obtains only the fields needed to complete the task and does not require the user to upload a full certificate. Verified data still need source and time labels. The model cannot extend a one-time consent to training or other uses.

Engineering landing

Write consent, purpose, fields, time limit, withdrawal, and recipient into the data contract and API controls. This success case shows that data sharing is realised by institution, identity, interface, and audit together.

Success Case: Open Data Forms an Innovation Ecosystem

Usage growth

Public data downloads grew from about five billion in 2019 to more than 80 billion by December 2025, with more than 5,700 datasets, about 110 APIs, and more than 2,500 data providers.

Large-model readiness

High usage does not mean natural fitness for AI. Every data product still needs a data dictionary, licence, update frequency, lineage, quality, change notice, and historical versions before it can safely enter retrieval and tool flows.

Ecosystem strategy

Use stable APIs and open formats so different clouds, research institutions, start-ups, and large enterprises can reuse. Government holds the standard and the service outcome; the market provides multi-brand innovation under common rules.

Success Case: AI+ Public-Service Tool Catalogue

Seven entry types

The tool catalogue covers digital-human customer service, meeting summaries, document processing, writing content, process automation, creative promotion, and data-analysis forecasting, and helps departments adopt through forums, seminars, and technology matching.

Catalogue elements

Besides features and price, mark deployment mode, data flow, model source, logs, support, portability, prohibited scenes, and exit path. Compare like-for-like uses on a common test set so procurement does not look only at a demo.

Risk sequence

First use low-risk internal drafts, summaries, and classification, requiring human confirmation before external issue. After evaluation and operating evidence accumulate, then enter cross-department and citizen services.

Northern Metropolis and Hetao: A Real Test Ground for Large Models

Industrial space

The Northern Metropolis focuses on innovation and technology, post-secondary education, and health and medical innovation, and forms a complete environment with a university town, industrial parks, research, and community. Hetao’s one zone, two parks drive Shenzhen–Hong Kong research and results translation.

Verifiable scenes

Within a controlled scope, test research knowledge assistants, clinical-trial material organisation, park services, cross-boundary professional information, and talent training, while assessing identity, data flow, weak networks, and multilingual needs.

Governance first

A living lab must have resident and user communication, data minimisation, exit, independent evaluation, and open interfaces. Cross-boundary projects handle rules, standards, and accountability first, then connect data and models.

Cross-Boundary Data Flows: Contract Requirements Must Land in the API

Institutional progress

The GBA standard contract for cross-boundary flow of personal information began as a pilot in 2023 and, from November 2024, expanded to all GBA sectors, providing an institutional tool for regional collaboration.

Technical translation

Engineering must turn purpose, data class, recipient, retention, onward transfer, security measures, and incident handling into field whitelists, permissions, encryption, logs, deletion, and alerts.

Model-specific risk

Confirm whether cross-boundary data enter prompts, caches, logs, vector indexes, or training. The processor location and subcontractors of the service provider must also be transparent. Checking only the main-database location is not enough.

Multi-Cloud and Domestic Compatibility: Do Not Bind the Brand; Bind the Standard

Combination principle

Different technology ecosystems have different strengths in models, data, edge, security, global regions, and local service. Nothing is split evenly. Choose by data residency, task quality, latency, cost, supply resilience, and talent.

Portability baseline

Use standard APIs, containers, open data formats, infrastructure as code, and externalised prompts and policy. Identity, logs, evaluation, and knowledge assets are controlled by the institution.

Verification method

Each year, drill transfer or alternate routing of a representative service, measuring time, cost, functional difference, and data integrity. Multi-cloud without drills is usually only multi-vendor at the contract layer.

Three-Year TCO: The Full Cost beyond the Price List

Direct cost

Includes tokens or GPUs, storage, network, software licences, platform, and support. Self-hosting also has facilities, idle capacity, hardware refresh, and spare parts. APIs have price change, cross-region, and peak cost.

Hidden cost

Data cleaning, labelling, evaluation, security, human review, training, customer service, audit, incidents, migration, and exit are often higher than the first prototype. Calculate per successful task, not per call.

Decision comparison

Compare at the same time doing nothing, improving the existing process, buying a mature product, self-building, and a hybrid option. For low-volume or unstable-demand projects, expensive self-hosting may not bring strategic autonomy. For stable, highly sensitive load, long-term self-hosting may be reasonable.

Procurement Acceptance: Turn a Demo into Accountable Capability

How to write the specification

Describe public outcomes, tasks, error boundaries, data rights, interfaces, observability, service levels, and exit. Do not write a particular brand or model name as the only answer.

Common examination paper

Every option runs the same golden dataset, red-team set, stress scenes, and failure drills. Set hard thresholds for accuracy, security, fairness, citation, refusal, latency, cost, availability, and complaint process.

Contract safeguards

Specify model-update notice, subcontractors, vulnerability patching, data export, deletion proof, handover period, and service termination. A low price accompanied by high exit cost cannot be treated as a real saving.

Closing: Turn Model Heat into Long-Term Public Capability

Final judgement

A new generation of models can raise reasoning, multimodal, tool, and deployment efficiency, but public-service success is still decided jointly by problem, data, process, governance, talent, and operations. A leaderboard can only provide clues. It cannot replace local evidence.

Successful experience

Hong Kong’s success cases in digital government, digital identity, CDEG, open data, an AI tool catalogue, and smart-city shared platforms show that scale depends on a common base, cross-department governance, and continuous service.

Action principles

First define acceptable error and build your own examination paper. Retrieval and rules first, then fine-tuning and agents. Every change is regressed, canaried, and reversibly rolled back. Maintain autonomy with open interfaces, multi-brand evaluation, and exit drills.