← Financial Cloud Cloud Cloud Club · AWS Re:cap

Vertex Macro | Financial Cloud Cloud · AWS Re:cap

AWS Re:cap 09: Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances

Speaker: Troy Dai, Rex Law

Session: 09

Session
Summit Dev Lounge2026 Re:cap
01 Encode Architecture as Steering for AI Agents
Summit Dev Lounge2026 Re:cap
02 Agent Harness Is the Real Engineering Moat
Summit Dev Lounge2026 Re:cap
03 Ask Observability Data in Plain Language
Summit Dev Lounge2026 Re:cap
04 Serverless AR Game with Bedrock AgentCore
Summit Dev Lounge2026 Re:cap
05 Multi-Agent Quant Backtesting on AgentCore
Summit Dev Lounge2026 Re:cap
06 Blog to Slides in Three Minutes with Kiro
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 AWS Compliance with Terraform
AWS Community Day Hong Kong 2025 Re:cap
03 Beginner to Builder An AWeSome Cloud Journey
AWS Community Day Hong Kong 2025 Re:cap
04 Team-First Serverless Engineering with Laravel & Bref
AWS Community Day Hong Kong 2025 Re:cap
05 Event Opening - AWS Community Day Hong Kong 2025
AWS Community Day Hong Kong 2025 Re:cap
06 Agent-to-Agent: Building Interoperable AI on AWS
AWS Community Day Hong Kong 2025 Re:cap
07 Utilize another telemetry data for faster improvement with AI agent
AWS Community Day Hong Kong 2025 Re:cap
08 Graduating from Vibe Coding: Spec-Driven Development with Kiro
AWS Community Day Hong Kong 2025 Re:cap
09 Automated Testing using MCP & AI Agents
AWS Community Day Hong Kong 2025 Re:cap
10 Modernizing Telecom Security ML Powered Approach
AWS Community Day Hong Kong 2025 Re:cap
11 Rethinking GenAI Agent: RAG & MCP
AWS Community Day Hong Kong 2025 Re:cap
12 Disaster and Emergency Response with TAK and AWS
AWS Community Day Hong Kong 2025 Re:cap
13 Rethinking Serverless Application Workflows from a Testing Perspective
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 AI-Powered Global Pure-Alpha Macro Trades on AWS: Revolutionizing Risk-Adjusted Asset Returns
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 Modern Trade Lifecycle: Trading to Settlement
FSI Recap
02 Goldman Sachs: Fast Track your applications onto Cloud - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group: Building an Effective Log Management Solution on AWS
FSI Recap
04 FSI Meetup 2025 Q4 - Brex Database Disaster Recovery
FSI Recap
05 FSI Meetup 2025 Q4 - A Graviton Migration Success Story
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel Modern Data Platform
FSI Recap
07 FSI Meetup 2025 Q4 - Financial Transaction Data Reconciler PayPal
FSI Recap
08 FSI Meetup 2025 Q4 - Scaling Resilience
FSI Recap
09 Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances
FSI Recap
10 Advanced Agentic AI Design Patterns
FSI Recap
11 Build New Modern Apps on AWS
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent Recap (IND3312)
AWS re:Invent 2025
02 Building the Future Trading Platform Leveraging AI and AWS
AWS re:Invent 2025
03 Trading Innovation: Jefferies' AI Assistant on Amazon Bedrock (IND3315)
AWS re:Invent 2025
04 How FSI Revolutionized HFT Analytics with Agentic AI (GBL302)
AWS re:Invent 2025
05 Improving Distributed Systems with Amazon Time Sync Featuring Nasdaq
AWS re:Invent 2025
06 Amazon Aurora HA and DR Design Patterns for Global Resilience (DAT442)
AWS re:Invent 2025
07 Building Agentic AI: Amazon Nova Act and Strands Agents in Practice (DEV327)
AWS re:Invent 2025
08 Deep Dive into Amazon Aurora and Its Innovations (DAT441)
AWS re:Invent 2025
09 Deep dive on Amazon S3 (STG407)
AWS re:Invent 2025
10 Nasdaq: Build Resilient Infrastructure for Global Financial Services (HMC327)
AWS re:Invent 2025
11 What's New with AWS Lambda (CNS376)
AWS re:Invent 2025
12 Spec-Driven Development with Kiro (DEV314)
AWS re:Invent 2025
13 Amazon's finops: Cloud cost lessons from a global e-commerce giant (AMZ308)
AWS re:Invent 2025
14 Tick to trade latency trading platforms on aws
AWS re:Invent 2025
Government data
01 The AI Era: The Boundary Between Development and Design Is Disappearing
Government data
02 On-Device Multimodal AI and Smart-City Practice
Government data
03 Large-Model Capability Evaluation and a Method for Landing AI Projects
Government data
04 Controlled End-to-End Automation of Government Development with Cloud Agents
Government data
05 A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
Government data
06 AI-Driven Macro Quantitative Research and Smart Governance
Government data
07 Authorized Operation of Public Data and Smart-Government Practice
Government data
08 Putting Data Assetization into Practice: Rights, Compliance, Engineering Governance, and Digital-Government Cases
Government data
09 AI for Mental-Health Public Welfare: Governance, Architecture, and Practice of a Trustworthy Platform
Government data
Amarathon 2025 Recap
01 A Developer’s Roadmap to Architecting for Agents
Donnie Prakoso
02 Amazon Bedrock Data Automation
Hafiz Syed Ashir Hassan
03 Multi-Agent on AgentCore
Tan Xin
04 Building Agentic AI Nova Act and Strands Agents in Practice
Haowen Huang
04 Accelerating Migration Projects with Kiro using Spec-Driven Development
Sanchit Dilip Jain
06 From Matching to Understanding: Personalized AI Search Practice Driven by AgentCore Memory
Liu Cao
07 Observe to Optimize – LLM Observability to AIOps Turning real-time insights into intelligent automation
Jimmy Soh
08 Deploying TEAM and Building the Best Engineering Team
Yuji Oshima
09 Five Hard Lessons from Five Years of So-Called Serverless Databases
Renato Losio
14 What if AI does my job How Q Developer CLI and Kiro have changed my daily routine
Miguel Angel Muñoz
16 Velocity with Vigilance: Security Essentials for Amazon Bedrock Agent Development
Brian Tarbox
26 Run OSS LLMs on a Single H100 Smarter, Cheaper, Faster
Adit Modi
28 A Modern Unified Metadata Architecture: New Approaches to Breaking Down Data Silos
Shaofeng Shi
29 Serverless MediaOps: Automating Video Workflows with AI on Amazon Web Services
Luis Valdivia
30 Architecting for Efficiency and Reliability with Performance Testing at Scale
Luis Guirigay
31 Connecting the World Through Open Source: Practical Journey of Technology, Community and Global Developer Relations
Richard Lin
33 Building Streaming Iceberg Tables for Real-Time Logistics Analytics
Fahad Shah
34 Accelerating Large-Scale Robot Strategy Training: An Automated Closed-Loop Architecture Based on Kiro, Trainium, and EKS
Junjie Tang
35 From Vibe to Viable with spec driven development
Ricardo Sueiras
36 Making Cloud Cost Analysis Smarter: Building FinOps Intelligent Agents with Strands and AgentCore
Xiaofei Li
37 Transform Conversational Agentic AIOps for K8s Using CNCF Kagent, K8sGPT, and Nova Sonic
Shaoyi Li

@ AWSome Day Hong Kong 2025

This session explored how to select AWS GPU instances and purchasing options for AI and machine learning workloads. The central message was that cost efficiency depends on matching the model, throughput, latency, and workload predictability to the right instance family and pricing model.


GPU Instance Families

Amazon EC2 G Instances:

● Optimized for graphics-intensive applications and machine learning inference.

● G6 instances use NVIDIA L4 GPUs.

● G6e instances use NVIDIA L40S GPUs.

● Well suited to cost-effective, medium-throughput inference workloads.

Amazon EC2 P Instances:

● Designed for high-performance machine learning training and high-throughput inference.

● P4 instances use NVIDIA A100 GPUs.

● P5 instances use NVIDIA H100 GPUs.

● P6 instances use NVIDIA B200 GPUs based on the Blackwell architecture.

● Well suited to large-scale, latency-sensitive inference and model-training workloads.


Choosing Between G and P Instances

G Instances are a strong fit for:

● Cost-sensitive workloads.

● Small and medium-sized AI models.

● Small and medium language models with fewer than 30 billion parameters.

● Distilled large language models.

● Traditional machine learning models such as XGBoost and random forests.

● Chatbots, personalization engines, recommendation systems, and image recognition.

P Instances are a strong fit for:

● Large language models with more than 30 billion parameters.

● High-throughput inference.

● Latency-sensitive applications.

● Vision-language and multimodal AI systems.

● Training large AI models.


Amazon EC2 Purchase Options

On-Demand Instances:

● Provide compute capacity billed by the second without a long-term commitment.

● Offer maximum flexibility but have the highest hourly rate.

● Work well for testing, prototyping, new service rollouts, spiky demand, and unpredictable workloads.

Savings Plans:

● Require a one-year or three-year usage commitment.

● Can reduce costs by up to 72% compared with On-Demand pricing.

● Suit stable and predictable production inference.

● Compute Savings Plans provide flexibility across instance families and Regions.

● EC2 Instance Savings Plans provide deeper discounts for a specific instance family and Region.

Spot Instances:

● Use spare Amazon EC2 capacity at discounts of up to 90% compared with On-Demand pricing.

● May be interrupted with a two-minute notification.

● Suit fault-tolerant, flexible, and stateless workloads.

● Common use cases include batch processing, non-critical inference, and scaling during off-peak hours.

EC2 Capacity Blocks for ML:

● Reserve accelerated computing capacity for a future start date and a defined time window.

● Can be reserved up to eight weeks in advance.

● Reservations can run from 1 to 182 days.

● A block can contain between 1 and 64 instances.

● Uses fixed, upfront pricing.

● Supports instance families including P4d, P5, P5en, and P6.

● Suits scheduled GPU training or inference jobs that require predictable access to scarce capacity.


Strategic Purchasing Recommendations

On-Demand:

● Development and experimentation.

● New service launches.

● Workloads with unpredictable or highly variable demand.

Savings Plans:

● Production inference with stable, predictable utilization.

● Long-running workloads where a commitment can be supported by usage data.

Spot Instances:

● Non-critical inference.

● Batch jobs and off-peak processing.

● Interruptible workloads with checkpointing and retry mechanisms.

Capacity Blocks:

● Scheduled large-scale training.

● Time-bound inference campaigns.

● Workloads that must secure GPU capacity for a known period.


Recent GPU Pricing Changes

● AWS reduced prices for P4, P5, and P5en instances by approximately 25% to 45%.

● The reductions apply to On-Demand pricing, EC2 Instance Savings Plans, and Compute Savings Plans.

● One-year EC2 Instance Savings Plans became available for P5 and P5en.

● These one-year plans can provide savings of up to 40% compared with On-Demand pricing.

● P6 instances with NVIDIA B200 GPUs are included in Savings Plans.

Regional Availability:

● AWS reduced EC2 Capacity Blocks for ML pricing for P5, P5e, and P5en instances across multiple Regions outside the United States.

● More consistent regional pricing improves cost predictability.

● Standardized pricing simplifies multi-Region planning for machine learning workloads.

● Global customers can reserve GPU capacity with fewer location-based pricing differences.


Implications of the Pricing Changes

Lower Total Cost of Ownership:

● High-performance inference and training become more affordable.

● AI experimentation and deployment budgets can support more workloads.

Improved Global Access:

● Expanded regional availability supports global deployment strategies.

● More consistent prices simplify worldwide capacity planning.

Greater Pricing Flexibility:

● Organizations can combine On-Demand, Savings Plans, Spot Instances, and Capacity Blocks.

● A one-year commitment reduces financial risk compared with a three-year commitment.


Model and Workload Optimization

Model Optimization:

● Deploy smaller or optimized models on G6 or G6e to reduce cost while maintaining responsiveness.

● Use INT4 or INT8 quantization so larger models can run on smaller GPU instances.

● Consider distilled models that offer comparable quality with lower compute requirements.

● Apply model compression where appropriate.

Batch Processing:

● Combine multiple inference requests to improve GPU utilization.

● Use dynamic batching to balance throughput and latency.

● Run non-urgent processing during off-peak hours with Spot Instances.

Real-Time Inference:

● Use G6 or G6e with On-Demand Instances or Savings Plans.

● Design for low-latency service-level objectives.

● Scale automatically in response to demand.

● Use regional deployments where lower network latency is required.

Batch Inference:

● Consider P5 Spot Instances or Capacity Blocks.

● Process large volumes during off-peak periods.

● Use checkpointing to recover from Spot interruptions.

● Build queue-based architectures.

● Optimize primarily for throughput rather than request latency.


Decision Framework

1. Assess the model size and throughput requirements.

2. Determine whether workload demand is predictable.

3. Evaluate latency sensitivity.

4. Define budget and reliability constraints.

5. Select the appropriate instance family, such as G for cost efficiency or P for performance.

6. Choose the purchasing option that matches workload predictability and interruption tolerance.

7. Monitor utilization and continue optimizing.


Implementation Best Practices

Technical Optimization:

● Implement model compression and quantization.

● Use dynamic batch sizing.

● Enable GPU sharing when the workload supports it.

● Monitor GPU utilization to detect idle or underused capacity.

● Optimize container images to reduce startup time and storage overhead.

Financial Optimization:

● Review workload patterns regularly.

● Use AWS Cost Explorer to analyze spending.

● Configure budgets and cost alerts.

● Combine purchasing options instead of using one model for every workload.

● Reassess Savings Plans and other commitments annually.


Key Takeaways

● Recent AWS GPU price reductions create meaningful opportunities to lower inference and training costs.

● Choose G Instances for cost-efficient inference and P Instances for high-performance training or inference.

● Match the purchasing option to workload stability, flexibility, and interruption tolerance.

● Optimize models and batching before moving to a more expensive GPU instance.

● A mixed purchasing strategy can balance cost, performance, capacity assurance, and reliability.

● GPU infrastructure should be reviewed continuously as model requirements and demand patterns evolve.

Next Steps:

● Audit current GPU utilization and costs.

● Identify workloads that can benefit from updated pricing.

● Test quantization, distillation, and other optimization techniques.

● Evaluate regional deployment opportunities.

● Consider Capacity Blocks for upcoming large training or inference jobs.