← Financial Cloud Cloud Cloud Club · Amarathon 2025 Recap

Vertex Macro|Financial Cloud Cloud · AWS Amarathon 2025

Vertex Macro|Amarathon 2025 Recap 33:Building Streaming Iceberg Tables for Real-Time Logistics Analytics

Speaker: Fahad Shah

Session: 33

Session
Summit Dev Lounge2026 Re:cap
01 Encode Architecture as Steering for AI Agents
Summit Dev Lounge2026 Re:cap
02 Agent Harness Is the Real Engineering Moat
Summit Dev Lounge2026 Re:cap
03 Ask Observability Data in Plain Language
Summit Dev Lounge2026 Re:cap
04 Serverless AR Game with Bedrock AgentCore
Summit Dev Lounge2026 Re:cap
05 Multi-Agent Quant Backtesting on AgentCore
Summit Dev Lounge2026 Re:cap
06 Blog to Slides in Three Minutes with Kiro
Summit Dev Lounge2026 Re:cap
AWS Community Day Hong Kong 2025 Re:cap
02 AWS Compliance with Terraform
AWS Community Day Hong Kong 2025 Re:cap
03 Beginner to Builder An AWeSome Cloud Journey
AWS Community Day Hong Kong 2025 Re:cap
04 Team-First Serverless Engineering with Laravel & Bref
AWS Community Day Hong Kong 2025 Re:cap
05 Event Opening - AWS Community Day Hong Kong 2025
AWS Community Day Hong Kong 2025 Re:cap
06 Agent-to-Agent: Building Interoperable AI on AWS
AWS Community Day Hong Kong 2025 Re:cap
07 Utilize another telemetry data for faster improvement with AI agent
AWS Community Day Hong Kong 2025 Re:cap
08 Graduating from Vibe Coding: Spec-Driven Development with Kiro
AWS Community Day Hong Kong 2025 Re:cap
09 Automated Testing using MCP & AI Agents
AWS Community Day Hong Kong 2025 Re:cap
10 Modernizing Telecom Security ML Powered Approach
AWS Community Day Hong Kong 2025 Re:cap
11 Rethinking GenAI Agent: RAG & MCP
AWS Community Day Hong Kong 2025 Re:cap
12 Disaster and Emergency Response with TAK and AWS
AWS Community Day Hong Kong 2025 Re:cap
13 Rethinking Serverless Application Workflows from a Testing Perspective
AWS Community Day Hong Kong 2025 Re:cap
14 Practical AWS FinOps for Cloud Success
AWS Community Day Hong Kong 2025 Re:cap
15 AI-Powered Global Pure-Alpha Macro Trades on AWS: Revolutionizing Risk-Adjusted Asset Returns
AWS Community Day Hong Kong 2025 Re:cap
FSI Recap
01 Modern Trade Lifecycle: Trading to Settlement
FSI Recap
02 Goldman Sachs: Fast Track your applications onto Cloud - AWS Re:cap Q1/2023
FSI Recap
03 Zurich Insurance Group: Building an Effective Log Management Solution on AWS
FSI Recap
04 FSI Meetup 2025 Q4 - Brex Database Disaster Recovery
FSI Recap
05 FSI Meetup 2025 Q4 - A Graviton Migration Success Story
FSI Recap
06 FSI Meetup 2025 Q4 - Stifel Modern Data Platform
FSI Recap
07 FSI Meetup 2025 Q4 - Financial Transaction Data Reconciler PayPal
FSI Recap
08 FSI Meetup 2025 Q4 - Scaling Resilience
FSI Recap
09 Maximizing AI Inference Cost Efficiency: Strategic Adoption of AWS GPU Instances
FSI Recap
10 Advanced Agentic AI Design Patterns
FSI Recap
11 Build New Modern Apps on AWS
FSI Recap
AWS re:Invent 2025
01 Coinbase re:Invent Recap (IND3312)
AWS re:Invent 2025
02 Building the Future Trading Platform Leveraging AI and AWS
AWS re:Invent 2025
03 Trading Innovation: Jefferies' AI Assistant on Amazon Bedrock (IND3315)
AWS re:Invent 2025
04 How FSI Revolutionized HFT Analytics with Agentic AI (GBL302)
AWS re:Invent 2025
05 Improving Distributed Systems with Amazon Time Sync Featuring Nasdaq
AWS re:Invent 2025
06 Amazon Aurora HA and DR Design Patterns for Global Resilience (DAT442)
AWS re:Invent 2025
07 Building Agentic AI: Amazon Nova Act and Strands Agents in Practice (DEV327)
AWS re:Invent 2025
08 Deep Dive into Amazon Aurora and Its Innovations (DAT441)
AWS re:Invent 2025
09 Deep dive on Amazon S3 (STG407)
AWS re:Invent 2025
10 Nasdaq: Build Resilient Infrastructure for Global Financial Services (HMC327)
AWS re:Invent 2025
11 What's New with AWS Lambda (CNS376)
AWS re:Invent 2025
12 Spec-Driven Development with Kiro (DEV314)
AWS re:Invent 2025
13 Amazon's finops: Cloud cost lessons from a global e-commerce giant (AMZ308)
AWS re:Invent 2025
14 Tick to trade latency trading platforms on aws
AWS re:Invent 2025
Government data
01 The AI Era: The Boundary Between Development and Design Is Disappearing
Government data
02 On-Device Multimodal AI and Smart-City Practice
Government data
03 Large-Model Capability Evaluation and a Method for Landing AI Projects
Government data
04 Controlled End-to-End Automation of Government Development with Cloud Agents
Government data
05 A New Software Ecosystem for the Agent Era, Seen Through Multi-Agent Systems
Government data
06 AI-Driven Macro Quantitative Research and Smart Governance
Government data
07 Authorized Operation of Public Data and Smart-Government Practice
Government data
08 Putting Data Assetization into Practice: Rights, Compliance, Engineering Governance, and Digital-Government Cases
Government data
09 AI for Mental-Health Public Welfare: Governance, Architecture, and Practice of a Trustworthy Platform
Government data
Amarathon 2025 Recap
01 A Developer’s Roadmap to Architecting for Agents
Donnie Prakoso
02 Amazon Bedrock Data Automation
Hafiz Syed Ashir Hassan
03 Multi-Agent on AgentCore
Tan Xin
04 Building Agentic AI Nova Act and Strands Agents in Practice
Haowen Huang
04 Accelerating Migration Projects with Kiro using Spec-Driven Development
Sanchit Dilip Jain
06 From Matching to Understanding: Personalized AI Search Practice Driven by AgentCore Memory
Liu Cao
07 Observe to Optimize – LLM Observability to AIOps Turning real-time insights into intelligent automation
Jimmy Soh
08 Deploying TEAM and Building the Best Engineering Team
Yuji Oshima
09 Five Hard Lessons from Five Years of So-Called Serverless Databases
Renato Losio
14 What if AI does my job How Q Developer CLI and Kiro have changed my daily routine
Miguel Angel Muñoz
16 Velocity with Vigilance: Security Essentials for Amazon Bedrock Agent Development
Brian Tarbox
26 Run OSS LLMs on a Single H100 Smarter, Cheaper, Faster
Adit Modi
28 A Modern Unified Metadata Architecture: New Approaches to Breaking Down Data Silos
Shaofeng Shi
29 Serverless MediaOps: Automating Video Workflows with AI on Amazon Web Services
Luis Valdivia
30 Architecting for Efficiency and Reliability with Performance Testing at Scale
Luis Guirigay
31 Connecting the World Through Open Source: Practical Journey of Technology, Community and Global Developer Relations
Richard Lin
33 Building Streaming Iceberg Tables for Real-Time Logistics Analytics
Fahad Shah
34 Accelerating Large-Scale Robot Strategy Training: An Automated Closed-Loop Architecture Based on Kiro, Trainium, and EKS
Junjie Tang
35 From Vibe to Viable with spec driven development
Ricardo Sueiras
36 Making Cloud Cost Analysis Smarter: Building FinOps Intelligent Agents with Strands and AgentCore
Xiaofei Li
37 Transform Conversational Agentic AIOps for K8s Using CNCF Kagent, K8sGPT, and Nova Sonic
Shaoyi Li

Modern Logistics Challenges

● Managing multiple streams for trucks, drivers, routes, fuel, maintenance, shipments, and warehouses.

● Need for real-time operational views and long-term analytics.

Data Storage Requirements

● Fresh, joined views for immediate operations.

● Use of Apache Iceberg for long-term analytics.

Technology Stack

● RisingWave: Data platform for streaming capabilities.

● Lakekeeper: Open REST catalog for data management.

● Kafka: Event backbone for streaming data.

● Object Storage (e.g., MinIO): Storage solution for data.

Objective

● Demonstrate how to build streaming Iceberg tables using the specified open stack.

● Provide a simple and effective solution for modern logistics data management.

The Logistics Analytics Problem

● Today's logistics platforms generate:

● Trucks: fleet inventory and locations

● Drivers: rosters and assignments

● Shipments: origin, destination, and weight

● Warehouses: capacity and sites

● Routes: ETAs and distances

● Fuel & Maintenance: cost and reliability signals

● The challenge:

● Operational teams need fresh, joined views across all of these streams.

● Data teams need the same data in Iceberg for BI, AI, and historical analysis.

What We’ll Build (Streaming Iceberg Pattern)

● Kafka feeds seven logistics topics into RisingWave.

● A multi-way streaming join is expressed in SQL and materialized continuously inside RisingWave.

● The result is persisted from RisingWave as a native Apache Iceberg table in S3-compatible object store like MinIO.

● Engines like Spark, Trino, and DuckDB query the same Iceberg tables via an open REST catalog.

Why Streaming Iceberg Tables with RisingWave?

[ 1 ] Batch-first workflows:

● Periodic jobs, stale joins, and heavy pipelines.

● Separate ETL tools to write into Iceberg.

[ 2 ] RisingWave + streaming Iceberg tables:

● Continuously updated joins and aggregates in RisingWave MVs.

● Iceberg snapshots that are always “almost current.”

● One RisingWave pipeline that serves both real-time dashboards and offline analytics.

● Goal: Make Iceberg feel like a database by letting RisingWave own the streaming pipeline and Iceberg writes.


High-Level Architecture

● Our end-to-end stack:

● Kafka — event backbone for 7 logistics topics.

● RisingWave (streaming database) — ingest, join, and aggregate in SQL; manage materialized views.

● RisingWave Iceberg Table Engine + Lakekeeper — open REST catalog over Iceberg tables.

● MinIO — S3-compatible object storage.

● Pattern: Kafka → RisingWave → Iceberg in MinIO → Query from any engine via REST catalog.

Logistics streams in RisingWave & multi-way streaming joins

● The Seven Logistics Streams in RisingWave

● Our running example uses seven Kafka topics that become sources in RisingWave:

● trucks — fleet inventory, capacity, current location.

● driver — driver details and assigned_truck_id.

● shipments — origin, destination, weight, truck binding.

● warehouses — warehouse location and capacity.

● route — route_id, truck_id, driver_id, ETD/ETA, distance_km.

● fuel — refueling events (time, liters, station).

● maint — maintenance history and costs.

● RisingWave treats each one as a streaming table, ready to be joined with simple PostgreSQL-style SQL.

Pattern 1: Multi-Way Streaming Join in RisingWave

● In RisingWave, we express the core logistics logic as one multi-way streaming join.

● LEFT JOIN drivers → trucks to keep unmatched drivers visible.

● JOIN shipments to attach workload and destinations.

● JOIN warehouses to bring in capacity and location.

● JOIN route for ETD/ETA and distance.

● JOIN fuel and maint for cost and reliability signals.

● This becomes logistics_joined_mv — a continuously updated, denormalized logistics record per truck/driver/route inside RisingWave.


Fleet KPIs, native Iceberg tables & cross-engine reads

Pattern 2: Fleet KPIs View in RisingWave

● On top of the joined MV, we define another RisingWave MV for fleet KPIs:

● Capacity utilization (%) per truck.

● Total fuel cost and maintenance cost per truck.

● Combined total operational cost.

● Current route context (ID, ETD, ETA, distance_km).

● Associated driver details. overview in RisingWave becomes a live fleet performance table — for Grafana and operational dashboards.

Pattern 3: Streaming to Native Iceberg from RisingWave

● Instead of a custom writer service:

● [ 1 ] We define logistics_joined_iceberg as a native Iceberg table managed by RisingWave.

● [ 2 ] The schema mirrors logistics_joined_mv.

● [ 3 ] A small config in RisingWave controls how often streaming changes are committed as Iceberg snapshots.

Pattern 4: Cross-Engine Reads via REST Catalog

● With the Iceberg table created by RisingWave and registered in a Lakekeeper REST catalog:

● [ 1 ] Spark attaches lakekeeper as a catalog

● [ 2 ] Trino / DuckDB / Dremio can use their Iceberg connectors to read the same table.

[ 3 ] All engines see the same Iceberg data that RisingWave continuously updates.

● No copies, no proprietary table formats — just plain Iceberg, written by RisingWave.


From local laptop to production cluster: deployment options

● Deployment Options: From Laptop to Cluster

[ 1 ] Local (for learning and prototyping):

● Run RisingWave, Kafka, MinIO, and Lakekeeper with Docker.

● Perfect for experimenting with streaming joins and Iceberg tables on your laptop.

[ 2 ] Production (for real workloads):

● Deploy RisingWave and the rest of the stack via Kubernetes + Helm.

● Use storage classes, resource limits, and persistence suitable for your environment.

● Same SQL and patterns in RisingWave — just more durable, scalable, and automated.

Simplifying the Traditional Iceberg Stack

● Traditional Iceberg deployments often require:

● A separate stream processing engine.

● Standalone Iceberg writer jobs.

● External compaction and maintenance workflows.

● Extra glue to keep catalogs, writers, and storage aligned.

● With RisingWave:

● [ 1 ] The streaming database handles ingestion, joins, materialized views, and Iceberg writes.

● [ 2 ] The REST catalog + MinIO keep everything fully open and interoperable.

● Fewer moving parts, less operational overhead.

Reference architecture with RisingWave

● Think of the system in three layers, centered on RisingWave:

[ 1 ] Streams → RisingWave Tables.

● Kafka topics become streaming tables in RisingWave.

[ 2 ] Tables → RisingWave Materialized Views.

● Streaming joins and aggregates become live MVs (logistics_joined_mv, truck_fleet_overview).

[ 3 ] Views → Streaming Iceberg Tables.

● RisingWave turns an MV into a streaming Iceberg table with a small config and an INSERT....SELECT.

● Once you see RisingWave as the “streaming SQL + Iceberg engine”, you can reuse this model in many domains.

Reusable Patterns Beyond Logistics

● The RisingWave + Iceberg pattern applies to:

● E-commerce: orders, inventory, pricing, customer events.

● FinTech: transactions, balances, risk signals.

● Industrial IoT: machines, sensors, alerts, maintenance.

● Telecom: sessions, usage, QoS metrics.

● Anywhere you have multiple real-time streams plus a need for open, long-term storage, you can use RisingWave MVs and Iceberg tables the same way.

Key Takeaways (RisingWave + Iceberg)

● A reference architecture combining Kafka, RisingWave, REST catalog, MinIO, and Iceberg.

● Practical patterns: multi-way streaming joins, KPI views, and native Iceberg writes from RisingWave.

● Get real-time logistics analytics without custom writers, ad-hoc compaction jobs, or tight vendor lock-in.