# 高性能存储加速生成式 AI 全生命周期

> **合规声明（生成式 AI 服务）**：前述特定亚马逊云科技生成式人工智能相关的服务目前在亚马逊云科技海外区域可用。亚马逊云科技中国区域相关云服务由西云数据和光环新网运营，具体信息以中国区域官网为准。
> 
> **完整资料：如果想了解更多内容可以访问:https://events.amazoncloud.cn/china-summit-on-demand/
> 
> **免责声明**：本摘要仅为便于您了解峰会内容而提供，我们不对其完整性及准确性作出任何保证，且不构成任何建议或结论。相关内容仅供参考，实际内容以峰会回放视频为准。

## 一、基础信息

- **会议类型**：专题演讲
- **Persona**：数据团队
- **时间信息**：6月23日 | 15:00 - 15:30
- **标题**：高性能存储加速生成式 AI 全生命周期
- **PDF 资料**：有
- **视频回放**：有

## 演讲人信息

### 会议信息
- **地点**：上海世博中心（2026 亚马逊云科技中国峰会）
- **日期**：2026年6月23日（Day 1）
- **时间**：15:00 - 15:30
- **会议类型**：专题演讲
- **面向 Persona**：数据团队

### 亚马逊云科技演讲人
- **Christian Smith**：亚马逊云科技全球存储业务拓展总监

## 二、会议资料 PDF 转录内容

--- Page 1 ---
2026年6月23日-24日 上海·世博中心

--- Page 2 ---
高性能存储加速生成式 AI
Accelerating Generative AI with High-Performance Storage
Christian Smith
亚马逊云科技 WW Specialist Storage

--- Page 3 ---
数据是生成式 AI 的关键差异化优势
Your data is the differentiator for Generative AI
存储策略对生成式 AI 至关重要
• 生成式 AI 从非结构化数据中释放新的业务价值
• 您的数据才是关键差异化优势 ——而非模型
• 无论是构建、定制还是训练模型，数据可访问性决定了结果
• 高质量数据的存储策略至关重要
• S3 是数据存储的基石：全球百万客户的数据湖构建在 S3 上

--- Page 4 ---
生成式 AI 生命周期：存储的关键作用
Generative AI Lifecycle: Where Storage Matters
01 02 03
训练 微调 / 优化 智能体
利用海量数据集训练基础模型 利用行业数据调整预训练模型以提 使智能体能够从领域特定数据中提
（万亿级 token,PB 级 升模型性能 取洞察，从而简化复杂任务
数据，数千个 GPU 实例） （LoRA、RLHF、全量微调）

--- Page 5 ---
面向生成式 AI 的存储
Storage for Generative AI
S3 FSx S3+
Amazon S3 / S3 Amazon FSx S3 Vectors / S3 Files /
Express One Zone for Lustre Mountpoint
S3 Vectors:
适用于大规模数据湖和模型训练的可扩 完全托管的并行文件系统 适用于 RAG 工作负载的向量存储
展对象存储 POSIX 语义兼容
S3 Files:
S3 Express One Zone 提供 亚毫秒级延迟 通过 NFS/POSIX 访问 S3 数据个位数毫秒级延迟 高达 1 TB/s 的聚合吞吐
大规模人工智能场景（>1000 GPU） 适用于大规模分布式训练 Mountpoint:
为 ML 框架提供文件访问
无限容量 1 TB/s 吞吐量 向量搜索
10 倍速度提升（Express） 亚毫秒级延迟 NFS 访问
>1000 GPU 规模 POSIX 兼容 框架原生

--- Page 6 ---
分布式模型训练和调优的存储挑战
Distributed Training: Storage Challenges

--- Page 7 ---
为什么存储性能对模型训练调优至关重要
Why Storage Matters
模型调优在您的领域特定数据上重新训练基础模型，存储性能决定了您
$32/小时
昂贵的 GPU 是保持忙碌还是闲置
GPU 闲置小时的成本
数据吞吐必须匹配 GPU 计算能力， 如果存储无法足够快地向 GPU 提
供数据，您就在为闲置的 GPU 资源买单
#1
存储性能瓶颈是模型训练中的主要挑战
错误的存储选择 = 闲置的 GPU = 浪费的支出

--- Page 8 ---
Amazon S3 最佳适用场景
分布式模型训练/调优， 1-16 个实例的集群规模
能力
大规模数据场景
• 流式吞吐量
• 几乎无限的容量
• 按用量付费
• 与 SageMaker 和 PyTorch 原生集成
• 适用于：LoRA/QLoRA、模型调优
• 与 S3 Files 和 S3 Mountpoints 集成
挑战
小文件工作负载（亿级小文件）
频繁的检查点保存和恢复

--- Page 9 ---
最佳适用场景
FSx for Lustre
分布式模型训练/调优，16-2000 个实例的集群
能力
大规模训练场景
• 跨数千个客户端的并行 I/O
• 亚毫秒级延迟，保持 GPU 满负荷运行
• 扩展至 1+ TB/s 聚合吞吐量
• POSIX 兼容(PyTorch、DeepSpeed、Megatron)
• 与 S3 集成， 支持从 S3 延迟加载，无需手动管理数据复制和同步
• LZ4 压缩 + 原生 S3 检查点
• EFA+ NVIDIA GPU Direct Storage 支持
• 适用于：多节点大规模分布式训练
挑战
数十亿小文件（元数据服务器）场景，需要针对性的部署和规划

--- Page 10 ---
S3 Express 最佳适用场景
超大规模场景—数万个实例
One Zone
适合于云原生且偏好对象存储接口的数据科学团队
能力
超大规模训练 • 低延迟， 比标准 S3 低 10 倍以上
• 大规模并发下的一致高吞吐
• 数据并行访问，线性扩展
• 为机器学习优化的存储桶
• 适用于：超大规模集群，1000+ 或更多的 GPU 集群
挑战
需要使用 S3 API
与标准 S3 存储桶没有直接的数据集成，需要手动管理数据迁移

--- Page 11 ---
通用建议：从 FSx for Lustre 开始
Customers typically start with FSx for Lustre
• 简单易用：工具和数据加载器与文件系统原生集成
• 高性能：吞吐量随文件系统大小线性扩展
• 存储选择：持久化 vs. 临时： 按照工作负载生命周期来选择持久化存储或者临时存储
• 建议：从吞吐量需求开始规划，而非容量
• 成本优化：使用 Intelligent-Tiering 优化成本而不牺牲性能

--- Page 12 ---
Adobe 仅用九个月就发布了
Firefly 生成式 AI 模型
利用 Amazon FSx for Lustre 加速大规模模型训练
— Adobe Creative Cloud & AI

--- Page 13 ---
决策框架 Decision Framework
为训练选择合适的存储
就地使用您的数据
< 16 个实例
S3、EFS、EBS、FSxZ、FSxN 等 —— 无需移动数据
16-500:FSx for Lustre（除非团队或应用程序偏好对象存储）
16 - 1,000 个实例
500-1,000:FSx for Lustre 或 S3 Express OneZone
需要频繁检查点或 POSIX？→FSx for Lustre | 否则→团队对文件 vs 对象的偏好
S3 Express OneZone
> 1,000个实例
在此规模下，经济性通常超越其他考量因素

--- Page 14 ---
面向 Agentic 工作负载的存储服务
Storage for Agentic Workloads

--- Page 15 ---
AI 通过 Agent 变得更加自动化
AI gets more autonomous with agents
经典机器学习 生成式 AI 助手 生成式 AI 智能体 Agentic AI 系统
更多的
更少的
人工监督
人工监督
分析任务 内容生成 更复杂的功能 自动化决策
预定义规则 有限的决策功能 和任务 高度自适应性
有限的适应性 MCP 支持

--- Page 16 ---
Agents 需要多种类型的数据
短期 长期
S3 Express S3 Files S3 Vectors
记忆
React
框架
LLM 反思
（推理）
APIs
Tools/
MCP
Agents 思维链
Data
上下文
业务数据基础
Bedrock Knowledge Base
S3 Tables Amazon S3
S3 Vectors

--- Page 17 ---
数据如何支持 Agentic AI 应用
Natural Language Query
We just hired Amy
Please start onboarding
Orchestration Agent
任务规划 Agent
StrDucattuarbeads edsata
构建定制化的入职计划
Vector search
入职伙伴 Agent Unstructured
Data lakes Vectovria s tores
推荐需要联系的人 data
RAG
异常处理 Agent
ReaSl-ttriemame dinagta s estrrveicaems ing
立即适应变化

--- Page 18 ---
Agent 记忆组件
短期 长期
当前任务或当前会话（线程） 当前用户或持久化的应用级别
Agent state Messages Semantic Profiles Episodic Prompts
Action plan, Conversation Semantic User profiles Historical System Prompt,
data shared history, context relevant and similar data interactions and instructions
between loops, sequence of for the outcomes (RL)
scratchpad Messages requested task
通常使用 通常使用工具在向量 根据应用代码、
依赖 Agentic 框架 数据存储库和企业数据库中 文件或键值存储
的后端数据存储 进行检索 进行版本控制
构建者缺乏控制 数据可访问性和质量 提示词调参

--- Page 19 ---
不只是记住， 还要智能！
Agentic memory is key
情境智能 用户偏好 知识留存
通过情境理解和模式识别提 通过记住跨多个会话的偏好 通过持久记忆和持续学习
升响应的准确性和相关性 和历史对话实现个性化交互 能力实现复杂的问题解决

--- Page 20 ---
两种场景：有记忆 vs. 无记忆

--- Page 21 ---
S3 Files 像传统文件系统一样工作
Works like a traditional file system
文件和目录反映 S3 桶中的内容
• 从任何桶创建文件系统，无需数据迁移
• 智能地将数据存储在文件系统中，实现毫秒级延迟访问
• 文件系统更改立即对文件应用程序可见
• 更改在桶和文件系统之间自动同步
S3 Files

--- Page 22 ---
Amazon S3 Vectors
核心能力 典型适用场景
• 原生支持向量的云对象存储 • 数据湖上的语义搜索
• 向量上传、存储和查询成本降低高达 90% • 批量检索 pipelines
• 零基础设施管理
• 与向量数据库形成互补架构（冷热分层）
• 支持数十亿规模，并具备元数据过滤能力
• 对成本敏感的大规模向量存储
• 与 Amazon S3 相同的持久性和可用性
延迟（P95） 向量规模 QPS
>100 ms 数十亿 <100s/index

--- Page 23 ---
存储 — AI Agent 的外部大脑
Amazon Storage - The External Brain for AI Agents
可扩展的知识库
Amazon S3 Vectors
Long term memory and RAG 持久化记忆与上下文
Amazon S3 Files & FSx 工作记忆
AI Agent
Short Term Memory
多 Agent 协作
Amazon
Amazon
S3
S3
data lakes
Real-time data Tstarbealemsing
鲁棒性与弹性
Data Foundation

--- Page 24 ---
S3 Files — 大规模 Agent 场景的持久化存储
For agents operating at scale
• 每个 Agent 都需要相互隔
离的数据访问
• 通过 Access Points 实现数
据访问隔离的 Agent 工作
S3 Files 区，存储空间按需自动扩展
• 每个文件系统可扩展至
10,000 个 Access Points
Agents run time S3 Access Points

--- Page 25 ---
推理与数据管道
Inference and Data Pipelines
S3 Files S3 Express One Zone S3 Vectors
大规模存储和查询向量嵌入
为 ML 提供 S3 的文件系统访问个位数毫秒级延迟
用于 RAG 工作负载
框架和推理管道 用于模型服务和实时推理
相比专用向量数据库
S3 Files 提供低延迟性能的原生 NFS 大规模并发下的一致高吞吐量
成本降低多达 90%
挂载，无需数据复制或重新格式化 适用于需要快速读取的工作负载
同时具备完整的 S3 持久性
适用于模型加载和 适用于热模型缓存和实时服务
适用于语义搜索和检索增强生成
批量推理工作负载

--- Page 26 ---
RAG in Action
User Input
Text Generation User Prompt Large Response
augmentation Language
Workflow
Model
文本生成
Embeddings
Context
model
Embedding
0.89 -0.02 -0.53 0.95 0.17 -0.38
Data Ingestion Semantic
search
Workflow
数据摄入
S3 Vectors Embeddings Document chunks Data source
model

--- Page 27 ---
Agent 在 EKS 上的部署
Amazon Cognito Amazon S3 Files
(SessionManager)
用户认证 Session State
Artifacts & Scratch
旧金山的天气
怎么样？
Amazon EKS
聊天 MCP 天气
界面 天气 Server API
Alice Agent
Amazon S3 Amazon S3 Vectors
(Data Lakehouse)
RAG
LLM
User Content Long Term
Memory

--- Page 28 ---
开始使用 - Getting Started
1 2 3
识别您的工作负载 构建您的数据基础 加速投产
训练/微调？ 以 S3 作为您的数据湖起步 使用决策框架
将 GPU 数量匹配到存储层级 进行训练存储选择
使用 FSx for Lustre 用于训练
Agentic AI？使用 Agent 记忆分类法
将数据映射到记忆类型 添加 S3 Files + S3 Vectors 将存储原语映射
（短期 vs. 长期） 用于 Agent 记忆和 RAG 到 Agent 需求
联系您的亚马逊云科技存储专家 | aws.amazon.com/fsx | aws.amazon.com/s3

--- Page 29 ---
关键要点 - Key Takeaways
训练与微调 Agentic AI
01 存储匹配集群规模 01 您的数据是差异化优势
<16 GPU: S3 | 16-1000: FSx for Lustre | >1000: S3 Express Agent 的能力取决于它们能访问和推理的知识
02 消除 I/O 瓶颈 02 Agent 需要记忆
短期（S3 Files）用于工作状态+ 长期（S3 Vectors）用于
闲置 GPU 每小时 $32 ——存储吞吐必须匹配计算
RAG 和回忆
03 从 FSx for Lustre 开始 03 专用存储原语
S3 Files、S3 Vectors、S3 Tables ——每个解决不同的
POSIX、亚毫秒延迟、1 TB/s ——覆盖大多数训练工作负载
Agent 访问模式
存储是生成式 AI 的基础——从训练到生产 Agent

--- Page 30 ---
Thank you

## 三、会议回放视频转录内容

Hello there
This should be experience
Or Nihao I should say
This should be an experience
This should be the first time I get to present in front of a group of people
Where the presentation is entirely in Mandarin
And I get to look a little script here
Make sure that I'm saying the right things as we go
So first and foremost
Thank you
Like I said, my name is Christian Smith
I run the worldwide specialist team
For storage and data protection
Look, today we're going to talk about something today
That often gets overlooked
Which is storage in the conversations around GenAI
You know, everybody talks about models and GPUs and frameworks
And I'm actually going to show you how your storage strategy
Is actually going to be the make or break factor
In your GenAI performance and cost efficiency
So let's start with some of the fundamental truths
Look, when you're doing GenAI
The real differentiator is your data
It's not always the model
So everyone has pretty much the same access
To a set of models in their country or Region
And what makes the application unique
Is the quality, the relevance
And the accessibility of your data
And as a matter of fact, when I've been here
I've been talking to a lot of customers
And this has been the number one thing that comes up
Is like, how is my data quality
In order to feed it into my models
So that's always a great starting place
But whether you're building a model from scratch
Or customizing one with fine tuning
Or powering an AI agent
It's really going to be tied to
How well your storage layer delivers your compute
And this isn't hypothetical
The world's largest data lake today
Which store all of this data
Run on S3
And the question isn't whether your data
Is in the cloud, it's whether your storage
Architecture is optimized for the
Specific demands of GenAI workloads
So let's frame this up a little bit
There are three distinct phases
Or workloads
This one has a very distinct storage requirement
So first there's training
This is building a foundational model from scratch
We're talking trillions of tokens, petabytes of data
Thousands of GPUs running in parallel
Storage needs to deliver massive
Sustainable throughput
If you're building your own LLM
There's a good chance you're probably working with
Just due to the sheer quantity of LLM
Sure quantity of GPUs that you're using
Second, there's fine tuning
Where you're taking a pre-trained model
And you're applying your domain-specific data
Techniques like LoRA
and RLHF
Or full fine tuning
And the scale is smaller
But the GPU utilization is still present
Third is agentic
And with agentic, it's about
AI agents reasoning over your domain
Data using tools
Creating complex tasks autonomously
Agentists need fundamentally different
Storage patterns, persistent memory
Vector search, and file-based system access
Let me walk you through what storage
Looks like in each of these phases
Here's our portfolio we're going to talk about today
There's our object storage portfolio
There's Amazon S3
There's S3 Onezone Express
This is really about modern architectures
It's about storing data at object scale
And petabytes of data
Then we have FSx for Lustre
Fully managed parallel file system
Sub millisecond latency
Sub millisecond throughput
Aggregate throughput scales into the
terabytes of data per second
And then our newest addition
So we start to talk about platforms
We talk about data platform
We talk about S3 vectors
S3 files and mount points
Which is really about storing
Vector embeddings
Accessing your S3 data natively
As a file system, so think agents
Are providing mount points
To provide a file system interface
For high throughput, high bandwidth environments
Each of these will solve a distinct pattern
That we're going to talk about today
So we're going to dig into training first
So this is probably the highest
Throughput use case that we have
Which is high throughput, high bandwidth
And really this is important
Because this is where economics
Really matter
So fine tuning
Foundation models
On your domain specific data
Really storage becomes the
Bottle neck in almost all the cases
That we have seen today
And it's because of
You want to keep your GPUs
As busy as possible
Because if they're not busy
That's wasted dollars
So for example
This is from a while ago
Like a p4, p5 instance
Is about $32 per hour
And so every time that GPU
Is not fully utilized
That's money wasted
And storage again
Is the bottle neck here
But when you choose
Starve GPUs
Let's dive in a little bit here
So let's put the guide together
So the first one is
Let's Amazon S3
So architecting for scale
So when you're training
Smaller models and fine tuning
In the kind of 1 to 16 instance
S3 is your straightforward choice
So you get streaming throughput
Virtually unlimited capacity
And you only pay for what you use
S3 integrates natively
With SageMaker and PyTorch data
S3 和 贴文件
这项目的功能是
用 ML 和 自行管理的
ML 和 工程的 功能
这项目的功能是
用 LoRA、QLoRA
S3 和 单码的 训练
这些功能不需要
需要 训练的 技术
S3 会很难
如果 训练 技术
或 训练
是 功能需要
百分百的 技术
如果您有这些用途的情况，请请我们一起工作
我们有一个很好的团队的市场考验
可以建立最好的图案
Lustre，我们再次提到 Lustre
它会在16-2000个时期的时间中
提供的能量是 parallel IO
有千万的客人，在一秒钟的时间中
保持 GPUs 永远地浮现
调整一秒至一秒的空间
它是 natively with PyTorch
deep speed, Megatron, no code changes
一个重要的因素是
我们有了一些能量的功能
你可以连接到 S3 的预算
然后将资料放在资料中
这意味着你不需要
预算很多的资料
你只需要预算大量的资料
在你的热工程中
You get things like LZ4 compression to reduce storage
And it has things like EFA
So EFA is elastic fabric adapter
That's our high speed low latency network
You get NVIDIA GPU direct
This equates to getting 12 to 1300 GB per second per instance
The big operational planning here
If you get into the billions of files
Work with us
Because we're going to be in this
Working with you to make sure that data gets distributed
Across metadata servers
For the best throughput and performance
Now when you move beyond 1000 GPUs
Think P5s or Trainiums
This is where we talk about mega scale
So this is where S3 Onezone becomes your preferred choice
It delivers 10x lower latency than standard S3
It's consistently high throughput at high concurrency
It has request level parallelism
It scales linearly
The directory bucket structure is optimized for ML access patterns
This is for customers that are strongly cloud native
And they prefer an object interface
You can put mount points
You can put a file interface in front of S3
Express Onezone
But if you're at this scale
You're going to be building your own custom data loaders
And you're going to want to be working with us
To get the highest throughput possible
In practice
Almost all of our customers have started out with Lustre
This has been the preferred platform
Preferred platform
For all of our customers doing training and fine tuning
It's simple to get started
It works out of the box
It natively integrates with S3
You get linear scaling
You have a choice between persistent and scratch deployments
To match durability and workload lifecycle
My recommendations start sizing with throughput requirements
So this is an important point
So calculate how many GPUs you're going to need
Or you're going to get to
Size it for throughput first
Capacity second
Because this will make your life very easy
When you continue to add more and more data
And don't forget to use intelligent tiering
FSX intelligent tiering
For your underlying S3 data optimized cost
Without sacrificing performance
Adobe great example
They launched the Firefly generative AI models
In just nine months using FSx for Lustre
To power their high performance model training
When you're competing in the GenAI market
Time to model is everything
And storage throughput directly determines that timeline
We built a decision framework here
So just recapping everything
But I said at the beginning
If you have under 16
Just use S3
Because that's most likely where your data is at
Realistically, if your data is in EFS
Or EBS or FSXZ or N
Just work from there
Like don't spend the time moving your data around
Or shuffling your data around
Just to make it available to GPU frameworks
If you're in the 16 to 500
This is most likely going to be Lustre
If you start to get between 500 to 1000
You start to get in the place
Where you can make a decision of going between Lustre
Or S3 One Zone Express
And it really depends on
Are you going to be writing your own data loaders or not
And then of course you get bigger than that
You're going to want to go with S3 One Zone Express
It has the best economics at that scale
Write this down
Take a picture
It'll save you a lot of headaches later
Okay, let's shift gears completely
And talk about agentic AI
This is one of the fastest growing areas
In our organization
And it requires a fundamentally different
Storage architecture
I'm going to walk you through a little bit
Of a progression here
About where we've come from
And where we're going to
And why these storage classes matter
As we get to agentic
I don't know if we know everything yet
But these are the patterns that we're seeing emerge
So first
Look, in our life
This has been the journey we've been on
I can't read that
But I think it says something along the lines of
We started out with kind of classic machine learning
We focused on analytical tasks
Pre-defined rules
We then got to generative AI assistants
Think chatbots
All those websites you went to
And you asked it a question
And it didn't answer you
And then it directed you to customer service
That we thought were really cool
At the time, but they were really useless
Well, now we're getting to fully
Agentic systems that are emerging
Autonomous workflows
Where multiple agents are collaborating
And really with minimal human oversight
That's where we're heading to
As we move into the spectrum
The demands of your data foundation increase exponentially
You know, agents need far more data
And in far more ways
Than this like simple chatbot ever did
So this is the big picture
That kind of plays out what we knew today
You know, this is going to be start with your data foundation
Agent need to understand
And consume structured data
Like financial reports
Customer records
There's unstructured data like images
Videos, documents
This is really your data that sits in your lake house
Right?
This is your foundation of all your data
For RAG, agents need vector search
And knowledge bases
To access relevant content
S3 vectors and Bedrock
Knowledge base serve this role
For agentic memory
And this is critical
Agent need short term memory
For storing intermediate content
Like task state
Or enabling multiple agents
To exchange information
S3 files and S3 express
Are really going to be well suited here
Long term memory
Allows agents to recall information
Across sessions, users
And be more personalized over time
So stored as semantic embeddings
In S3 vectors
And finally here hosting the LLM models
That give agents their reasoning capacity
This requires fast model loading over here
And context retrieval
All of this is built on
Things like Bedrock, AgentCore, SageMaker
But a lot of customers are doing this themselves
In EKS with their own self managed models
So they're building these layers out
And using these primitives on their own
Let me give you a real concrete example here
So this is an example of an
Agentic orchestration
Task planning
So this is let's say we just hired Amy
Amy is going to be onboarding soon
A task planning agent right here
Is going to build a tailored plan
Via API calls using Amy's profile
To assign task provision tools
Schedule activities
Pulling from structured databases
An onboarding buddy curates a personal learning path
Using vector search across handbooks and wikis
And recommends coworkers based on roles and interests
And at the bottom you got an exception handling agent
Monitoring for disruptions
Like a delayed laptop or a changing start date
And adapts in real time feeding this state back up
Into the other layers
Each agent needs different data
Different access patterns and different storage services
To function effectively
So I'm leading to memory
And this is the path that we're on
Because this is the new area
Agent in order to get autonomous
In order to get agentic
Are going to need more and more memory
And memory of different types
And the way we look at it is two types of memory
There's short term memory
Which is your current task or session
This is going to store your agent state
It's going to store in-flight tasks
It's going to store Messages
Checkpoints for crash recovery
It's tightly integrated with frameworks
Like LangChain, CrewAI, Strands
And typically uses high performance shared storage
Like S3 Express, S3 files, or instance memory
Really, you're thinking minutes of retention here
But most people don't delete this
Becomes a good source to refine later
And create new contextual state
Long term memory
Is really about persistence across users and sessions
This includes semantic content
User profile, Historic interactions
So every time I'm my agent
And I have the Christian bot that I use
My Christian bot has Prompts
That tells
That I've gone through all my documents
It has all my writing style
It has all my Slack Messages
That I've ever written
It knows who I am
So my agent has this long term memory
That every time it invokes
It knows who I am
And how I write and how I respond
And it's pulling it from our long term memory
So that it has this Prompt
That says
Write like Christian
Or respond
How would Christian respond?
Typically condensed from short term memory
So we're going from short term memory
Down to long term memory
It's really embedded
It's stored in vector databases
At large scale and retrieval
This is where S3 vectors
Which is purpose built for this
The last category is Prompts
It's a form of long term memory
Stored in like key value stores
Like DynamoDB or ElastiCache
You can use it to reduce repeated LLM calls
Like in caching Prompts
To reduce agent operational costs
So agentic memory is the key
Look proper agentic memory
Unlocks three transformative capabilities
I keep stepping in front
If you want to take pictures
I'll move out of the way
Contextual intelligence here
Where the agents not just only understand
What you're asking
But why you're asking it
It's like pattern recognition
Across interactions
There's user preferences
Which is about personalized interactions
The agent remembers how you work
How you communicate across sessions
And then there's knowledge retention
Which is continuous learning
The agent gets smarter
With every interaction
So it's basically fine tuning itself
All as it goes
These aren't really nice to have
They're essential for enterprise adoption
Without memory, agents can't maintain
Conversational continuity
Can't adapt from feedback
And can't develop user-specific understanding
Here's an example
With memory and without memory
So without memory, you're asking questions
And you're always going to start out
The same place every time you start
It's going to always start out
With something like who are you
And it's going to go
Continue to ask those questions
From there
It's going to have no sense
Of your previous interaction
It's always
Think of it like a squirrel
Or like a gerbil
Where it's a goldfish
Which has a memory of two seconds
Like it's always excited to see you
It's always going to say the same thing
It's always going to respond the same way
Now you start to get it into memory
And this is where you start
To get that repeated interaction over time
So as you get into memory
You can interact with things
That remembers your contextual
History
It remembers how it responded to you
It remembers how you responded to it
So that when you come in the next time
It's going to have contextual awareness
Of who you are, what you do
And what types of questions that you're asking
So I've given you the history of memory
I think we've led up to that
Which have had nothing to do with storage yet
So let's dive into this first
So when we talk about this type of agent memory
We talk about really kind of two things here
We're going to first talk about
If we're using S3 as our data foundation
And it has all of our
Organizational line of business
Enterprise type data
S3 files is a great way
To connect that data to your agents
So your agents can directly interact
With that data through a file system interface
And almost all agents are built
Expecting a file system interface
It's the way they operate
Now from an agent perspective
Agent are going to be creating code
They're going to be executing code
They're going to be participating
They're going to be creating contextual memory
Via text files, MD files
They're going to be storing long term memory
In MD files or storing short term memory
In MD files
And so this is a great place
That you can then also interact with that data
But store that short term memory state
Directly to a file system
And then over time
That data gets synced directly into your S3 bucket
So that you have this long term
Agent state over time
When you deploy like this
There's no custom code
There's no manual sync jobs
It just works
So an example of this is
We were talking to a bioinformatics company
They have genomics data
They have agents start to parse
Through this genomics data
And each agent has a task
Like you're looking for this type of data
When you find this type of data
You're going to invoke this other agent
Who's going to go search for this type of cancer
Or tumor in that data
As it's going along
It's building up code
To go find that data
It's feeding that data into LLMs
It's persisting memory and state
So that the next time you interact with it
Or tell it to do something different
It knows where it started
It hands off that data
To another agent
Which is going to compare it against a known database
Of existing tumors and cancers
Looking for that genomic defect
And these things will run over and over and over again
Data is in S3
Agent accessing through a file system interface
Very easy to connect these two things together
And persist state
No manual custom code
No jobs, it just works
Now, I don't know
Most people here know what S3 vectors are
So S3 vectors
So we've had S3
As an object store for years
What you're seeing is
As a data platform
It's moving up into the value chain
And here's what's happening here
We notice that customers are running
A lot of things like OpenSearch
Or their vector databases
A lot of these vector databases
Get very costly to run
In the billions of rows
They're very expensive because they treat all data as hot
And they have to run on kind of like
The most expensive hardware that's out there
What we said was
You know, the semantics of puts and gets are sort of similar
We could add a k-top to it
We could add a
Similarity index to it
And we can create a vector database
In S3 through a vector bucket
The result is
You can store billions of vectors
Very inexpensively
It's like 90% lower cost
The tradeoff is you're going to have a higher query time
And so if you're dealing with things like agents
Does that matter?
If it's 100 milliseconds versus 1 millisecond
Does that matter if you're saving 90% on cost?
Most of the time, no
Really where that
High performance rag database
It needs to come in
Is when you're doing something interactive
Like with a website
Where there's a time to element
For customer experience
So again, very
Advantages for people that need multi-billion dollar
Multi-billion row vectors
It does work with ElastiCache
OpenSearch so you can tier vectors into it
But it's really about
How do I take that memory
Long term memory state
And store it in something cheap and inexpensive
That provides that performance
Necessary to meet the requirements
And so really
This is the layering of what it looks like
So how do we create storage
Is the brain for agents
It starts with long term memory and rag
And using S3 vectors
You go into short term memory
And it's S3 files
You can of course use FSX
Or EFS in here too
It's also for working memory too
So if you're just doing things
Like working with data
That you have stored in your bucket
It's robustness and reliability here
With things like
The 11 nines of S3s
Durability
And you can do multi-agent collaboration
We really recommend building
Your data foundation on S3
Because of the resilience, cost effectiveness
And integration with both
亚马逊云科技 analytic services
And the broad ecosystem
Of third party services out there
I'm coming to a close here soon
So we're going to cover
Two more use cases
One, we've seen a lot of people
Deploy open claw at scale
And they're thinking about
Doing it as a service
So they're going to provide it
As a service to their end users
We do think that open claw
Will become kind of the foundation
That a lot of ISVs
Will build their base off of
And turn it into a product
When you start to do that
And you're turning your
Open claw environment
Into a SaaS experience
There's something that you want to think about
Which is
Okay, I need a file interface
I need the ability to store that data
Eventually into S3
For long-term durability
And cost effectiveness
But I also need
Tenant isolation here
So I can't have
Customer X
Interfering with customer Y
Their data can't cross or commingle
So we give you a very cost-effective way
To create tenant isolation
You can create up to 10,000 access points
For file system
Each with its own isolated security model
You can almost think about it
Like giving each agent its own sandbox
You can also do this
Inside your organization too
So if you're building
A tenanted agent model
For your organization
This is a great way to do it
Last thing
Storage for inference
And data pipelines
So storage for inference
Specifically
Self-manage on EKS
It's really about three different services
By the way, if you go on
Our container on the Amazon
I can't say 亚马逊云科技
亚马逊云科技
And you go to the EKS
Landing page
There is a link
To go to workshops
Our top workshop right now
Is running inference
On EKS
And running fine-tuning on EKS
These are free
You can go take it anytime you want
You'll get like a hands-on lab through it
So walking through this
So inference
Especially on self-manage EKS
S3 files or EFS
For model loading and batch inference
The native NFS provides low latency
No copying of data
No reformatting needed
S3 express one zone
Is really about hot model caches
So if I'm deploying thousands of EKS nodes
And I change the library
That load time to reactivate
All those containers becomes really critical
So think of these as low capacity
But really high throughput
So I may only have a couple gigs
But I'm going to burst at times
Up to 200 to 300 gigabytes per second
As I flip my model
This is a great service to do that with here
And then S3 vectors for your rag pipeline
Really grounding your foundation models
And providing like the 90% lower cost
Over specialized vector databases
This is rag in action
So here's how it works
On the bottom here
You're going to be taking your data source
Turning it into document chunks
You're going to be running an embedding model in it
You're going to be storing those embeddings
In S3 vectors
At query time
User input goes through
It's going to query the same embedding model
And the vector database looking for semantic similarity
It retrieves the most relevant passages back up
And passes them to the context LLM
Vectors and semantic search
The foundation for rag S3 vectors
Makes this really simple and scalable and cost effective
In cases where kind of that access time
Can be built into what you're doing
Which tends to be a lot of time in Prompt documentation
And here is the state for agent deployment and EKS
When you put it all together
When I talk about EKS workshops
And doing self-manage
This is how we think about putting it all together
So you have here
An agent-maintaining state
In S3 accessing long term memory
In S3 vectors
They're storing user content
In the data lake
And connects to external servers
Like MCP servers
And it can use authentication through Cognito
So even in a simple like agent example here
You can see how storage services kind of map
Into each distinct phase of this
Short term memory, long term memory rag
Foundational data
And for frameworks if you're using things like strands
We have S3 session manager
Can provide native session state out of the box
How to get started
Identify your workload
For training, match GPU count to storage
Map your data to the data types
Memory types
Build your data foundation
Start with S3
Accelerate to production
Use the decision framework for what we showed earlier
For storage use agent memory taxonomy
To map storage, primitives to agent
And please engage your specialist if this gets hard
I think you're giving me a time
And I don't know what that says
But I think it's a minute
Key takeaways
Again fine tuning
Agentic AI
Great picture to take here
I won't rehash it
I think we've covered it enough
But your data is the differentiator here
And so agents are only going to be as good
As the data you provided
That's it
Thank you
I had a great time