Vertex Macro | Financial Cloud Cloud · AWS Re:cap
AWS Re:cap 03: Large-Model Capability Evaluation and a Method for Landing AI Projects
The Real Starting Point for Landing Large Models: First Define Acceptable Error
Decision premise
A government project cannot start from a model name. First write down the service population, the public problem, the consequences of error, and the human substitute. A wrong statutory answer in a policy Q&A, a missed sensitive item in document review, and a lost restriction in an internal summary have entirely different consequences, so they cannot share one abstract accuracy threshold.
Three boundaries
First define the answerable range, the must-refuse range, and the must-human-review range. For each error type, mark impact, reversibility, time to discovery, and remediation cost, then decide the control strength that the model, data, and process must reach.
Field practice
Require business, legal, data, security, and frontline staff to review together the ten cases most likely to go wrong. If the team cannot form a consensus on error consequences, do not procure a model yet. The definition of success should land on processing time, first-time completion rate, error rate, frontline burden, and public satisfaction.
Evaluate Capability before Choosing a Brand
A brand cannot replace evidence
Open weights, hosted APIs, proprietary cloud, and on-premises models each have strengths, but leaderboards, parameter counts, or market noise cannot directly prove fitness for government work. Models must be compared on local corpora, real documents, permission rules, and acceptable latency. Brand is only one piece of supply-risk information.
Seven faces of the scorecard
Build a weighted scorecard across task quality, Chinese capability, tool use, security, latency and throughput, deployment efficiency, and licence and supply chain. Every point needs test samples, repeat counts, and failure evidence, avoiding scores from subjective trial use.
Selection discipline
Keep at least one alternative, and record the reason for the choice, the conditions for abandonment, and the trigger for re-evaluation. If the model is upgraded, the price changes, the licence is updated, or data-residency requirements change, the original conclusion lapses automatically and the baseline must be re-run.
Split Model Capability into Verifiable Tasks
Do not test chat impression
“Can reason,” “supports multimodal,” and “can use tools” are not acceptance descriptions. Split them into atomic tasks such as policy-validity judgement, conflict detection among provisions, table-field extraction, image–text checking, function-parameter generation, citation location, and refusal of overreach.
Task contract
For each task, write the input format, allowed data, expected output, tolerated error, latency ceiling, cost ceiling, and human-takeover point. Measure composite flows in segments; otherwise, when the final answer is wrong, it is impossible to tell whether retrieval, the model, a tool, or a business rule caused it.
Reusable outcomes
Atomic tasks can form a common test set across vendors, supporting procurement acceptance, version regression, and canary release. Mature teams keep failure samples, not only average scores, because rare but high-consequence errors most deserve continuous monitoring.
Six-Stage Landing Method: From Problem to Operations
Stages one to two
First complete problem definition and data preparation, confirming public value, roles, data legality, quality, permissions, and retention period. Without a reliable data source or a human process, even a strong model cannot form a sustainable service.
Stages three to four
Establish a human-acceptable baseline, then run offline evaluation. Compare first against search, rules, or the existing process, and only then introduce the model. Evaluation covers task success, citation, refusal, sensitive data, latency, throughput, and cost at the same time.
Stages five to six
Enter operations only after canary release, setting automatic stop, rollback, version tracking, drift monitoring, incident handling, and periodic re-evaluation. Each stage has an approver, required evidence, and an exit condition. A PoC must not skip governance gates because a demonstration succeeded.
Golden Dataset: Government Must Own Its Own Examination Paper
Sample sources
The dataset should come from public, authorised, anonymised, or synthetic real-work situations, including common, difficult, boundary, expired, conflicting, sensitive, and malicious cases. Collecting only standard answers overestimates system capability.
Annotation governance
Every sample stores the question, basis, expected behaviour, acceptable variants, refusal requirements, risk level, and reviewer. When policy changes, update the answer and retain historical versions, so today’s rules are not used to mis-judge yesterday’s system.
Asset sovereignty
The test set, prompt assets, knowledge base, and audit data should be held by the institution, not locked on a single vendor platform. Vendors may help build them, but must not block export with a proprietary format. The golden dataset is a core asset for long-term procurement bargaining and model switching.
Reasoning Capability: Test Process Constraints, Do Not Worship Thinking Text
Test focus
Government reasoning must verify rule applicability, exception conditions, evidence completeness, and conclusion consistency, not require the model to output lengthy thinking. Structured basis, cited passages, applicable dates, and open items can be required so reviewers can check quickly.
High-risk samples
Add situations where provisions conflict, attachments are missing, a policy has lapsed, local rules differ, and the question’s premise is wrong. A strong system should point out the gap, ask for supplementation, or refuse to conclude, rather than fluently filling in the unknown.
Engineering experience
Splitting a complex problem into retrieval, rule judgement, calculation, and text generation reduces unexplained error. When precise calculation is needed, call a controlled tool; do not let a language model do mental arithmetic. When a policy conclusion is needed, force citation of a valid source.
Multimodal Evaluation: A Document Is Not One Picture
Difficulties in government documents
Scans, overlapping seals, tables, handwritten annotations, headers and footers, watermarks, and multi-column layout break reading order. Multimodal capability must separately test OCR, layout understanding, cross-page association, seal position, and image–text consistency.
Segmented pipeline
First perform document-quality detection, page classification, and text-and-coordinate extraction, then let the model understand semantics. Original image, OCR text, layout blocks, and final conclusion must stay linked so a human can return to the original page to verify.
Acceptance method
Do not only ask whether content was read. Test field completeness, table relationships, page-number citation, refusal on low quality, and masking of sensitive content. If the model is over-confident about a blurred seal, the system should require a rescan or human confirmation.
Tool Use: The Model May Only Propose a Call; the System Owns Authorisation
Separation of authority
Generating a function name and parameters does not mean the model has the right to execute. Real authorisation is decided by identity, role, data scope, transaction limit, and approval process. Model output must pass a whitelist and a policy engine.
Safety guardrails
Set different risk levels for query, write, delete, payment, notification, and similar tools. High-risk operations use preview, second confirmation, dual review, or draft-only generation. Tool returns must also prevent prompt injection and sensitive-data reflux.
Audit evidence
Retain model version, prompt version, tool, parameter summary, authorisation result, execution reply, human confirmation, and final state. An incident investigation must be able to answer who authorised, what the model recommended, what the system executed, and whether data were changed.
Long Context: Being Able to Fit It In Does Not Mean Being Able to Find It
Common misconception
A claim of very long context only means a large number of tokens is accepted. It does not guarantee that the model will find the key provision in the middle, correctly handle multi-version documents, or keep citations consistent. The longer the context, the higher cost, latency, and interference may also be.
Stress tests
Place key evidence at the beginning, middle, and end, mix in similar clauses, old versions, and irrelevant attachments, and test recall, citation, and refusal. Then compare three strategies: stuffing the full text, layered summaries, and retrieval augmentation.
Practical principle
Cut context by document structure, permissions, and task, provide only necessary passages, and retain source indicators. Long context fits cross-passage integration. It should not replace data governance, version control, and reliable retrieval.
Chinese Government Capability: Not a Translation Score
Language layers
Evaluation should cover Simplified and Traditional Chinese, fixed official phrasing, legal wording, abbreviations, place names, written Cantonese, numbers and dates, measure words, and mixed Chinese–English layout. General Chinese fluency is not the same as accurate handling of government materials.
Situational difficulties
The same word may have different definitions in different departments. Policy names may have colloquial forms. Table fields often omit the subject. The model must combine a terminology base, department data, and document versions, and must not guess on its own.
Improvement method
First establish terminology, entity, and format rules, then use retrieval to supplement local knowledge. Consider fine-tuning only for stable, high-volume repeated errors. For dialect input, retain the original, standardise the result, and allow human correction, so meaning is not lost in conversion.
RAG First: Give the Answer a Basis
When to use it
Policy Q&A, institutional lookup, document review, and internal knowledge support usually start with retrieval augmentation, because knowledge updates, citation is required, and permissions are involved. Fine-tuning all knowledge into the model is slow to update and hard to trace.
Core components
Complete RAG includes data cleaning, chunking, permission tags, hybrid keyword-and-vector retrieval, re-ranking, context assembly, citation, and no-answer handling. An error in any link can produce a wrong answer that looks sourced.
Field sequence
First use human-acceptable keyword search as the baseline, then add hybrid retrieval and re-ranking; change only one variable each time. If the answer is wrong, first check whether the evidence was found, then check whether the model used the evidence correctly.
Hybrid Retrieval: Keywords and Vectors Each Guard a Gate
Keyword strength
Statute numbers, organisation names, proper terms, exact dates, and fixed phrases fit keyword retrieval and give stable, explainable hits. Vectors alone may mix documents that are semantically close but legally different in effect.
Vector strength
Users often describe a problem in everyday language and may not know the formal policy name. Vector retrieval can find semantically close content and cover synonyms, different question forms, and cross-language expression.
Fusion practice
After two-way recall, re-rank by permission, validity period, document level, and relevance. When needed, keep multiple conflicting sources for later rule judgement. Evaluation should separately count recall failure and re-ranking failure, not only look at the final answer.
Chunking and Indexing: Document Structure Decides Retrieval Quality
Chunking principle
Do not chunk only by a fixed character count. Retain chapter, section, article, paragraph, table, and attachment relationships, and use title, issuing body, effective date, region, classification, and permission as metadata.
Cross-passage handling
Provisions often cite earlier definitions or attachments. A single-passage hit may lack necessary context. Parent–child blocks, neighbouring passages, and citation relationships can be built so that, after a precise passage is recalled, the minimum necessary background is added.
Quality checks
Sample-check sentence breaks, table splits, page numbers, character encoding, and duplicate content. Index updates use versioning and atomic switchover, so the service does not read old and new policy at the same time.
Re-ranking and Citation: Credibility in the Last Mile
Purpose of re-ranking
First recall aims not to miss. Re-ranking puts truly usable, valid, and permission-allowed evidence in front. Ranking features are not only semantic similarity; they also include legal effect, date, region, document type, and question intent.
Citation standard
An answer must at least provide document name, valid version, clause location, and a verifiable fragment. The citation must actually support the conclusion, not merely look related under the same topic.
Failure handling
If sources conflict, there is no valid document, or only low-permission content exists, the system should state the gap and hand off to a human. Citation hit rate and citation-support rate must be measured separately: the former finds a source; the latter judges whether the source is enough to support the answer.
Prompt Engineering: Write Business Rules as a Testable Asset
Structural design
The system prompt should include role, task, allowed data, prohibitions, output format, refusal conditions, and human referral. Project prompts are separated from general brand tone, so a style change does not accidentally change safety behaviour.
Version governance
Treat prompts as code: bring them into version control, review, test, release, and rollback. Every change records reason, affected tasks, evaluation results, and approver. Do not make ad-hoc production-interface edits.
Avoid over-prompting
An overly long prompt raises cost and interferes with knowledge content. Stable permissions, calculation, and policy judgement should be executed by programs and a rules engine. The prompt only carries the behavioural constraints the model needs to understand.
The Choice Order of Rules, Retrieval, Fine-Tuning, and Agents
Determinism first
Fixed calculation, format checks, permissions, validity period, and hard policy first use programs or rules. When knowledge updates frequently and citation is required, use retrieval. These two steps often already solve most government needs.
When to fine-tune
Evaluate fine-tuning only when the model repeatedly shows style, structure, or specific classification errors on a stable task, and when there is enough high-quality labelled data. Fine-tuning cannot repair a wrong knowledge source or a chaotic process.
When to use agents
Use an agent only when the task needs multi-step planning and tool collaboration, and limit steps, tools, budget, and write operations. Agents increase autonomy and also increase unpredictability and audit difficulty. They should not be the first architecture option.
Fine-Tuning Decision: Trade Stable Error for Controllable Improvement
Data threshold
Fine-tuning data must represent real input, be correctly labelled, cover boundaries, and have usage rights. Taking historical human answers straight into training may freeze expired policy, personal preference, and private data together.
Controlled experiment
Keep a test set that did not participate in training, and compare the base model, prompts, RAG, and the fine-tuned version. Besides target-task improvement, also check general-capability regression, safety bypass, Chinese format, and refusal behaviour.
Operating accountability
The fine-tuned weights, base version, data version, and parameters must be registered. A base-model upgrade cannot assume the old adaptation still works; retrain or run a full regression. Security patching and capacity accountability also move to the using party.
Agent Systems: Lock Autonomy inside an Observable Process
Process modelling
First draw the task as states, tools, decision points, stop conditions, and human nodes. The model may choose the next step within a limited range, but must not enlarge the goal, add data sources, or bypass approval on its own.
Budget and stop
Set maximum steps, tokens, time, external calls, and retry ceilings. Stop immediately on insufficient permission, conflicting evidence, tool failure, or cost overrun, and return completed steps and items pending human action.
Observability
Trace each step’s input summary, model decision, tool result, state change, and cost. Recording only the final answer cannot explain why an agent detoured, repeated a call, or produced a wrong operation.
Policy Q&A and Document-Review Assistants: Combined Field Practice
Step one: baseline
First build a human-acceptable search baseline, using public or authorised policy, and define validity period, department, region, and permission. Select real questions and check the time humans need to find an answer and common missed items.
Step two: enhancement
Add hybrid retrieval, re-ranking, passage citation, and structured output. Test expired policy, conflicting rules, sensitive matters, table attachments, and no-answer cases. The model must not fill missing documents with common sense.
Step three: red team
Test with prompt injection, overreach queries, data exfiltration, forged sources, and malicious attachments. Every input, output, model, prompt, source, tool, human review, exception, and cost forms an evidence chain.
Open Weights Are Not Zero Risk
Autonomy gained
Open weights can support local deployment, deep evaluation, customisation, and replacement of the inference framework, reducing some API lock-in and helping keep sensitive data in a controlled environment.
Responsibility assumed
The using party also assumes licence interpretation, vulnerability patching, model-source verification, capacity planning, inference security, content governance, version compatibility, and shutdown handling. Without a mature operations team, self-hosting may be riskier than a hosted service.
Supply-chain verification
Retain source, hash, licence version, dependency list, and build process; scan containers and packages; restrict model-download channels. Run a full evaluation before any new weights go live. Do not treat nearby names in the same series as equivalent.
Licences and Legal Terms: A Hard Gate on the Technical Scorecard
Items to check
Confirm commercial use, derivative models, redistribution, use restrictions, branding requirements, data use, and termination terms. Licences for the model, datasets, libraries, and service APIs must be checked separately.
Change management
Licence terms may change by version. Procurement and architecture decisions must pin the version used and establish periodic review. If terms affect government use or data rights, a migration plan is required.
Practical delivery
Build a readable licence inventory and risk classification, jointly signed by legal, procurement, and technology. Do not treat “open source,” “downloadable,” or “free trial” as proof of lawful long-term operation.
Deployment-Mode Comparison: Self-Hosted, API, Proprietary Cloud, and Hybrid
When self-hosting fits
Self-hosting has value when data are highly sensitive, load is stable, deep control is needed, and the organisation has GPU, platform, security, and model-operations teams. Up-front investment, idle capacity, patching, and talent cost cannot be ignored.
When hosting fits
When demand changes quickly, value is still being validated, or strong model capability is needed quickly, a hosted API can lower start-up cost. Data use, region, logs, service level, price change, and exit must still be checked.
Hybrid design
Sensitive pre-processing, identity, and rules stay on a private network or on premises; general inference uses multiple clouds on demand. Peak and disaster recovery can use an alternate path. Hybrid is not stacking platforms. It is allocating work by accountability boundary.
Model Gateway: A Unified Control Plane for Multi-Brand Capability
Unified interface
The gateway encapsulates different vendors’ requests, authentication, rate limiting, retry, routing, content safety, and logs. Applications do not bind directly to a particular SDK. Model differences are managed through capability descriptions and adapters.
Policy routing
Choose the model by data class, task, cost, latency, availability, and region. Highly sensitive content can only take a controlled deployment; low-risk summaries can use a lower-cost option. On failure, route to an evaluated alternative model.
Avoid the lowest common denominator
Unification does not mean flattening distinctive strengths. Keep controlled extensions for vendor-specific capability, while ensuring the core flow remains replaceable. Routing policy, prompts, and evaluation data are held by the institution.
GPUs and Inference Capacity: Reverse-Engineer Architecture from the Peak
Load profile
First measure request length, output length, concurrency, time of day, task priority, and latency targets. Average traffic cannot represent licensing peaks, sudden events, or batch document review.
Efficiency measures
Use dynamic batching, KV cache, quantisation, model parallelism, request queues, and small-model offload. Simple classification need not use the largest model. For long-document tasks, limit input and process in segments.
Capacity governance
Set a capacity ceiling, queue time limit, degraded model, batch window, and priority. Regularly run stress and failure tests, assessing GPU utilisation, cost per successful task, and timeout rate together.
Enterprise Data Platform: Knowledge Needs Catalogue, Lineage, and Quality
Data as product
Policies, guides, cases, forms, and frequently asked questions all need an owner, source, validity period, permissions, quality rules, and an update commitment. A lakehouse or other platform is only a carrier. Governance accountability cannot be handed to storage technology.
Lineage tracking
It must be possible to go from an answer back to the index block, the cleaned result, the original document, and the published source, and also to find affected indexes and answers when an original document is withdrawn. This is the basis of correction, audit, and impact analysis.
Data quality
Monitor missing fields, duplicates, expiry, encoding, parse failures, and permission mismatches. Quality anomalies can block an index release. They must not wait to be discovered after the model answers wrongly.
LLMOps: A Model Upgrade Is a Production Change
Registration requirements
Record model source, version, quantisation, inference framework, prompts, tools, knowledge snapshot, evaluation, and deployment environment. Writing only a market name is not enough to reproduce behaviour.
Release process
After offline regression passes, move to shadow traffic, internal canary, a small share of users, then gradual expansion. Each stage sets stop conditions for erroneous confidence, refusal, latency, cost, and safety.
Rollback capability
Retain the previous stable version of the model, prompts, index, and rules, and test one-click or controlled rollback. Data-structure changes must be backward compatible; otherwise, rolling back the model cannot restore the service.
Observability: From Tokens to Task Success
Four layers of signal
Infrastructure watches GPU, CPU, memory, and network. Service watches latency, errors, and availability. Model watches citation, refusal, and safety. Business watches task completion, human takeover, and public outcomes.
Distributed tracing
Use a consistent trace identifier to connect entry, retrieval, re-ranking, model, tools, and human process. Logs use summaries and de-identification, so observability does not copy sensitive data.
Action thresholds
Every metric must connect to an alert, an owner, and a playbook. A dashboard itself does not create reliability. Without rules to stop a release, scale out, switch, or intervene with a human, data is only watched.
Security Baseline: Identity, Keys, Network, DLP, and Red Team
Least privilege
Use workload identity rather than long-lived keys. Authorise model service, data, vector indexes, and tools separately. KMS manages keys, a private network restricts flows, and management operations use step-up authentication.
Data protection
DLP detects personal and sensitive information in input, logs, and output; necessary data are masked or replaced first. Content safety must distinguish illegal, sensitive, overreach, and inappropriate advice. A single blacklist cannot handle every risk.
Continuous red team
Test prompt injection, data exfiltration, tool abuse, overreach retrieval, model supply chain, and denial of service. After a fix, add the case to the regression set so the next upgrade does not repeat it.
Evidence Chain and Auditable AI
What to record
Input summary, identity and permissions, model and prompt versions, retrieval sources, tool calls, rule hits, human review, final output, exceptions, latency, and cost form a complete event.
Minimised retention
Not all original text needs long-term retention. By data class, retain hashes, source indicators, structured decisions, and necessary fragments. Original sensitive content uses a shorter period or does not land on disk.
Audit uses
The evidence chain supports complaints, version comparison, vendor acceptance, cost reconciliation, security threats, and policy correction. If the system’s behaviour at the time cannot be reconstructed, the model cannot be used in a high-accountability process.
Success Case: Hong Kong Digital Government Shared Base
Scale outcomes
Hong Kong completed a whole-of-government review of electronic services and, by the end of 2025, advanced more than one hundred digital-government and smart-city measures, using big data, artificial intelligence, blockchain, and geospatial analysis to improve services.
Platform method
Next-generation government cloud, a big-data analytics platform, digital identity, shared blockchain, chatbot services, and a unified portal let departments reuse identity, security, data, and operating capability rather than rebuilding them on every project.
Methodological implication
Landing large models should also first build shared capability such as a model gateway, data governance, evaluation, logs, and tool authorisation, then accommodate Google, AWS, Tencent Cloud, ByteDance, and other technology ecosystems. Success comes from standardised governance, not from a single brand doing everything.
Success Case: Digital Identity and Corporate Identity
Citizen service portal
The one-stop iAM Smart digital-identity platform already has more than four million registered users, supports more than 1,300 services and electronic forms, and has obtained international-standard certification for information security and privacy information management.
Requirements of the agent era
When an AI assistant acts for a person or an enterprise, it must know the agent, the principal, the authorisation scope, the validity period, and signing responsibility. The model cannot infer authorisation from conversation alone.
Corporate-service extension
A corporate digital-identity platform is expected to launch by the end of 2026. Its core includes enterprise verification, digital signing, pre-fill, and a document wallet. This builds a trusted identity base for controlled government-to-business and business-to-business AI agents.
Success Case: CDEG and Once-Only Service
Verified scale
CDEG lets departments or authorised bodies exchange verified data with citizen consent, handling about two million exchanges a month and reducing repeated citizen submissions.
Value for AI
The model obtains only the fields needed to complete the task and does not require the user to upload a full certificate. Verified data still need source and time labels. The model cannot extend a one-time consent to training or other uses.
Engineering landing
Write consent, purpose, fields, time limit, withdrawal, and recipient into the data contract and API controls. This success case shows that data sharing is realised by institution, identity, interface, and audit together.
Success Case: Open Data Forms an Innovation Ecosystem
Usage growth
Public data downloads grew from about five billion in 2019 to more than 80 billion by December 2025, with more than 5,700 datasets, about 110 APIs, and more than 2,500 data providers.
Large-model readiness
High usage does not mean natural fitness for AI. Every data product still needs a data dictionary, licence, update frequency, lineage, quality, change notice, and historical versions before it can safely enter retrieval and tool flows.
Ecosystem strategy
Use stable APIs and open formats so different clouds, research institutions, start-ups, and large enterprises can reuse. Government holds the standard and the service outcome; the market provides multi-brand innovation under common rules.
Success Case: AI+ Public-Service Tool Catalogue
Seven entry types
The tool catalogue covers digital-human customer service, meeting summaries, document processing, writing content, process automation, creative promotion, and data-analysis forecasting, and helps departments adopt through forums, seminars, and technology matching.
Catalogue elements
Besides features and price, mark deployment mode, data flow, model source, logs, support, portability, prohibited scenes, and exit path. Compare like-for-like uses on a common test set so procurement does not look only at a demo.
Risk sequence
First use low-risk internal drafts, summaries, and classification, requiring human confirmation before external issue. After evaluation and operating evidence accumulate, then enter cross-department and citizen services.
Northern Metropolis and Hetao: A Real Test Ground for Large Models
Industrial space
The Northern Metropolis focuses on innovation and technology, post-secondary education, and health and medical innovation, and forms a complete environment with a university town, industrial parks, research, and community. Hetao’s one zone, two parks drive Shenzhen–Hong Kong research and results translation.
Verifiable scenes
Within a controlled scope, test research knowledge assistants, clinical-trial material organisation, park services, cross-boundary professional information, and talent training, while assessing identity, data flow, weak networks, and multilingual needs.
Governance first
A living lab must have resident and user communication, data minimisation, exit, independent evaluation, and open interfaces. Cross-boundary projects handle rules, standards, and accountability first, then connect data and models.
Cross-Boundary Data Flows: Contract Requirements Must Land in the API
Institutional progress
The GBA standard contract for cross-boundary flow of personal information began as a pilot in 2023 and, from November 2024, expanded to all GBA sectors, providing an institutional tool for regional collaboration.
Technical translation
Engineering must turn purpose, data class, recipient, retention, onward transfer, security measures, and incident handling into field whitelists, permissions, encryption, logs, deletion, and alerts.
Model-specific risk
Confirm whether cross-boundary data enter prompts, caches, logs, vector indexes, or training. The processor location and subcontractors of the service provider must also be transparent. Checking only the main-database location is not enough.
Multi-Cloud and Domestic Compatibility: Do Not Bind the Brand; Bind the Standard
Combination principle
Different technology ecosystems have different strengths in models, data, edge, security, global regions, and local service. Nothing is split evenly. Choose by data residency, task quality, latency, cost, supply resilience, and talent.
Portability baseline
Use standard APIs, containers, open data formats, infrastructure as code, and externalised prompts and policy. Identity, logs, evaluation, and knowledge assets are controlled by the institution.
Verification method
Each year, drill transfer or alternate routing of a representative service, measuring time, cost, functional difference, and data integrity. Multi-cloud without drills is usually only multi-vendor at the contract layer.
Three-Year TCO: The Full Cost beyond the Price List
Direct cost
Includes tokens or GPUs, storage, network, software licences, platform, and support. Self-hosting also has facilities, idle capacity, hardware refresh, and spare parts. APIs have price change, cross-region, and peak cost.
Hidden cost
Data cleaning, labelling, evaluation, security, human review, training, customer service, audit, incidents, migration, and exit are often higher than the first prototype. Calculate per successful task, not per call.
Decision comparison
Compare at the same time doing nothing, improving the existing process, buying a mature product, self-building, and a hybrid option. For low-volume or unstable-demand projects, expensive self-hosting may not bring strategic autonomy. For stable, highly sensitive load, long-term self-hosting may be reasonable.
Procurement Acceptance: Turn a Demo into Accountable Capability
How to write the specification
Describe public outcomes, tasks, error boundaries, data rights, interfaces, observability, service levels, and exit. Do not write a particular brand or model name as the only answer.
Common examination paper
Every option runs the same golden dataset, red-team set, stress scenes, and failure drills. Set hard thresholds for accuracy, security, fairness, citation, refusal, latency, cost, availability, and complaint process.
Contract safeguards
Specify model-update notice, subcontractors, vulnerability patching, data export, deletion proof, handover period, and service termination. A low price accompanied by high exit cost cannot be treated as a real saving.
Closing: Turn Model Heat into Long-Term Public Capability
Final judgement
A new generation of models can raise reasoning, multimodal, tool, and deployment efficiency, but public-service success is still decided jointly by problem, data, process, governance, talent, and operations. A leaderboard can only provide clues. It cannot replace local evidence.
Successful experience
Hong Kong’s success cases in digital government, digital identity, CDEG, open data, an AI tool catalogue, and smart-city shared platforms show that scale depends on a common base, cross-department governance, and continuous service.
Action principles
First define acceptable error and build your own examination paper. Retrieval and rules first, then fine-tuning and agents. Every change is regressed, canaried, and reversibly rolled back. Maintain autonomy with open interfaces, multi-brand evaluation, and exit drills.