# 从数周到数小时：借助 Amazon SageMaker AI 加速生成式 AI 的部署上线

> **合规声明（生成式 AI 服务）**：前述特定亚马逊云科技生成式人工智能相关的服务目前在亚马逊云科技海外区域可用。亚马逊云科技中国区域相关云服务由西云数据和光环新网运营，具体信息以中国区域官网为准。
> 
> **完整资料：如果想了解更多内容可以访问:https://events.amazoncloud.cn/china-summit-on-demand/
> 
> **免责声明**：本摘要仅为便于您了解峰会内容而提供，我们不对其完整性及准确性作出任何保证，且不构成任何建议或结论。相关内容仅供参考，实际内容以峰会回放视频为准。

## 一、基础信息

- **会议类型**：专题演讲
- **Persona**：模型实践者
- **时间信息**：6月23日 | 15:30 - 16:00
- **标题**：从数周到数小时：借助 Amazon SageMaker AI 加速生成式 AI 的部署上线
- **PDF 资料**：无
- **视频回放**：有

## 演讲人信息

### 会议信息
- **地点**：上海世博中心（2026 亚马逊云科技中国峰会）
- **日期**：2026年6月23日（Day 1）
- **时间**：15:30 - 16:00
- **会议类型**：专题演讲
- **面向 Persona**：模型实践者

### 亚马逊云科技演讲人
- **Susmita Marupaka**（来自旧金山）：Amazon SageMaker AI Inference 团队

## 二、会议资料 PDF 转录内容

（本场次无 PDF 资料）

## 三、会议回放视频转录内容

Hello everyone, welcome to 亚马逊云科技 Shanghai Summit.
I hope you are having a great day so far.
Spoiler alert, there won't be any cool demos like our previous presenter,
but I hope I can entertain you enough about Amazon SageMaker.
My name is Susmita Marupaka and I came all the way from San Francisco to meet you.
I'm excited to be here to share stories about Amazon SageMaker AI inference
and how it can help you deploy faster in your Agentic AI applications.
In the next 30 minutes or so, here is what we are going to discuss.
We are going to start by talking about the market trends of Agentic AI
followed by an introduction to SageMaker AI
and SageMaker Managed Inference.
I will give you a sneak peek of Agentic AI recommendations
which helps you save your benchmarking times from weeks into hours
followed by question and answers.
If you have any follow-ups on my presentation,
I'm happy to spend some time in the hallway discussing with you.
Let's get started.
What are the market trends that we are seeing with respect to Agentic AI?
We are living in a moment where the code, models and applications are all created by AI agents
than humans.
The proliferation of the AI artifacts and data tool calling,
querying databases, executing code,
all done by AI agents has increased in the order of magnitude over the past 18 months.
According to Gartner, by 2027, more than 33% of the workloads and applications will be Agentic AI-driven.
It goes without saying that that is the theme for 亚马逊云科技 Shanghai Summit today
and you are in the right place and the right room.
Now, as the data created by agents are increasing,
what are the challenges we are seeing?
As you see in this picture, it starts with zero-shot Prompting.
Very simple, easy compute.
Next comes RAG, retrieval augmented generation.
It introduces retrieval latencies,
long context, which increases the compute inferencing time by 2 to 5x.
A step further, let's talk about DeepSeek R1, which is a reasoning model,
which includes thinking and reasoning.
For every reasoning and thinking that the model does,
it consumes hundreds of tokens before generating responses.
That increases the compute, inference compute from 10 to 50x.
Welcome to Agentic AI, in which the multi-turn conversations
and the long context windows by Agentic AI workflows
exacerbate the consumption.
Every user Prompt has 5 to 10 model invocations
and in turn causing about hundreds of tokens
which proliferates the demand rising for compute inferencing.
Now, let's see what are the real challenges involved
in deploying your Agentic AI workflows from POCs to productions.
As we spoke to our customers in China and Asia Pacific,
here are the four reasons our customers are telling us.
Number one, while the model, which is the brain of AI,
lives and breathes on reasoning and intelligence.
Agentic AI goes one step forward.
It maintains context. It is stateful.
It needs runtime memory, which increases the challenges
fourfold. For example, when it comes to performance,
you might have deployed in your POC environment
at P99 latency for sub-hundred millisecond latencies.
But as it improves into production, it might be
up to ten thousand concurrent users.
Next up, when it comes to scalability, your workloads have
three x spike as compared to steady state.
And also the cost and complexity increases as well.
In the POC, you might have done for one p5.48xlarge instance, it is 10x. And the multimodality is
raising the complexity of deployment challenges.
According to Gartner, more than 40% of the POCs are bound to fail
and not to reach production by 2027.
Well, while this is the ground reality, let's understand
how SageMaker AI can help you alleviate your problems.
Let's start by understanding the SageMaker inferencing
and model development stack. As you see in this picture, it is
based on four core pillars. No-code ML, model customization
which is serverless. Recently, we have introduced
multi-turn reinforcement learning, in which, less than
ten lines of code, you will be able to incorporate new
skills for your model customization using SageMaker AI.
When it comes to training, our customers have deployed
SageMaker HyperPod over 500 GPU nodes and
Annapurna ASIC accelerators and chips. The SageMaker
HyperPod helps customers cut over 40% of the training
time with the built-in resiliency and task
governance capabilities. Today's topic of discussion
will be around deployment and inference. When it comes
to inference, that is when the rubber hits the road. Your
applications are deployed based on the training or
customization or the retraining capabilities which
you have done. When it comes to inferencing, there are
various workloads and capabilities which the
customers have come to expect from Amazon SageMaker.
For that, let's dive deep a little bit into how
Amazon SageMaker can help you balance your
performance, throughput latency and cost. All of this is
supported by SageMaker Studio, which provides access
to over thousands of models. We have, we provide choice
and flexibility to our customers. You can bring
your own agents or use the 亚马逊云科技 Hero Agent for
your agent work flows using SageMaker Studio. SageMaker
Jumpstart provides capability for day-zero support of
many models, including the open source NVIDIA and
Nemotron. The IDE, Studio notebooks, code
editors, and all of the studio will provide capability
for you to deploy, for you to discover the models
in SageMaker Studio. Let's have a key takeaway in
this slide. Whether you're training, experimenting, customizing
your models or deploying your models at scale using
SageMaker AI, it is built on governance and security
controls to ensure that your Agentic AI work flows
has the right context and right access to your data
backed up by 亚马逊云科技 Nitro Security Systems. We
understood Agentic AI challenges, why POC to
production fails in Agentic AI work flows, and how
SageMaker AI platform can help you alleviate those
challenges. Let's double-click on SageMaker AI
inference. There are two paths to production. If you
have existing EKS clusters or standardized on EKS
orchestrators in your environment, SageMaker
HyperPod inference is for you. Wherein, it provides persistent
dedicated clusters. One-click deployment from training
to your serving needs and HyperPod operator
which has out-of-box observability and
management simplicity for your inferencing
workloads. If you have challenges in terms of
standardizing on EKS and you want to fully manage
inference experience, SageMaker managed inference
is your path to production. It provides
dedicated inference endpoints, wherein you can
have customization and control over your
infrastructure. Now, the million-dollar question
whether I need to choose bedrock for inferencing
or SageMaker. The answer is it depends on
your operating model. If you want an easier access
to the foundational models with out-of-box
capabilities, faster time to deployment, using
API serverless access, Amazon Bedrock
inferencing is right for you. But as you deploy
your models, you scale your models. As you just
discuss from POC to production, the scale increases. The token
cost exacerbates. And you want to maintain
your price performance SLAs. Or you have custom
models, custom architectures, custom frameworks
that you want to deploy in your environment. SageMaker
inference is the right choice for you. Depending
on whether you want serverless or
dedicated endpoints and infrastructure
you can choose between your
operative choice, SageMaker inference offers two pathways
for you to choose for your deployment needs. Now, let's
take a look at SageMaker HyperPod architecture. As you
see in this picture, there are three VPCs. The one in
green on the right is the SageMaker HyperPod
data plane. It provides node-level access
to your environment. On the left hand side, you have
the purple customer control, your control, your
managed VPC, that is the user VPC. Above
that is the EKS control plane, which has
the control plane required for running your
EKS environment. Each of this multiple VPCs
are interconnected by an ENIs, which are provided
by both control plane and data plane. You don't have
to install multiple operators for your storage, observability
or management needs. SageMaker HyperPod EKS Kubernetes operator will take care of
that and abstract all of the underlying complexities
involved in managing your Kubernetes environment. And also
you can create your private subnets in two
series, in your user plane, and the
EKS control plane, and the SageMaker HyperPod
creates the ENIs, which helps you route
traffic efficiently among the control data and management plane, by
providing you node-level access for your inferencing needs. You
understood the EKS HyperPod architecture at
a bird's-eye view level. Now let's dive deep into how actually
inference operator works. Well, it is based on a single one-click deployment.
And it has auto-scaling and resource optimization
and built-in observability where it will
obtain information from multiple sources like NVIDIA
and 亚马逊云科技 observability metrics for you to have
a single pane of glass observability experience for your
inferencing needs to monitor your cluster health performance
and capacity. Alright, we discussed about
SageMaker HyperPod. Today, let's dive deep into SageMaker
Managed Inference. Let's take a look at the stack first. How do you
deploy your workloads or your endpoints in SageMaker
Managed Inference?Well, it is a three-simple process.
Step 1.Deploy your model. Step 1 is made based on model artifacts.
Provide your model objects, model artifacts
and inference code by pointing to S3
or Hugging Face or JumpStart. It is paired with the
ECR which has model registry and the runtime
container optimization code for you. Step 2.Deploy
SageMaker Endpoint by choosing your
inference stack, instance types, and then
SageMaker Managed Inference deploys the endpoints for you.
It does model deployment based on the source. It takes
care of the health checks to ensure it's the fully
managed experience and last but not the least, it
also ensures the observability dashboards are
given to ensure you have enough control over your infrastructure.
Now, after the deployment, you access the endpoint
by a simple API call. The key benefits of
SageMaker Managed Inference are six-fold, as you see
on the right-hand side. But to give you an overview, it
supports over 1000 plus instances. It auto-scales
and it has built-in observability dashboards using
CloudWatch, which I'll discuss in a bit so that you
don't have to worry about managing all of them by yourself.
This is fully managed inference optimization for you
using SageMaker Managed Inference. And there are
customers who deploy multiple models against our
endpoint. Our customers serve millions of user
requests behind dedicated endpoints. And those
dedicated endpoints are handful under five. It also
autoscales the model copies based on your
performance requirements. In the middle, you have
vLLM and SGLang and few other runtimes that
you support. If you have chosen vLLM
as excuse me, as your runtime container optimization, there
are two choices. You can bring your own container
or 亚马逊云科技 builds the deep learning container for you
based on three artifacts. Whether you want
latency optimized for which the tensor parallelism is
four or eight. Or you want a balanced workflow
in which you have expert parallelism along with
the tensor parallelism. Or you want a
throughput optimized inferencing work loads
for that you have pipeline optimization
along with expert optimization to ensure it
meets your customization need. Long answer
short depending on your workloads SageMaker
does the undifferentiated heavy lifting to ensure
you have right optimizations in place for your
customization architectures. All of it is backed
by the AI infrastructure, whether it is 亚马逊云科技
Annapurna Trainium/Inferentia
or NVIDIA GPUs or CPUs for your workloads. It
also supports streaming, bidirectional
streaming and nonstreaming responses. According
to Gartner, the multimodality of the
models has increased threefold as compared
to text only models. When it comes to
bidirectional streaming of the multimodality
request SageMaker AI has a
non-blocking process in which it is backed
by HTTPS or web socket wherein the input
and output streaming happens asynchronously
so that you can have a seamless experience
for your voice, video or any other
multimodal application. All right
we discussed about SageMaker
and how it can help you alleviate your
challenges on your path to production of
your Agentic AI workloads.I also want to dive
deep on Agentic AI recommendations, the
new features we have launched recently to
enhance your experience. When we
speak to our customers, our customers are
telling that the real challenge in
deploying an Agentic AI model is harder
not just to deploy but to have it done
right. What does that mean?There are five
different frameworks or
artifacts in which our customers
optimise whenever they deploy for
inferencing. What are they?Model
architectures, hardware types, frameworks, your
container runtime and your workload criteria. For each
of those combinations there are
hundred different permutations. Overall it is an
optimisation which our customers have to do
for more than 1600 combinations. Not only
that, some of the combinations are
conflicting. For example, if you have
RAG workflows and you done or
optimised your lengthier documents, it
doesn't work for smaller documents. And also
DeepSeek R1 which is an MoE and
reasoning model, it requires
different optimisations as I
discussed earlier. So your
reasoning model which is
output biased or output heavy. And what
is the impact?When customers
optimise, one it is manual, two they
try to optimise on whatever
instance it is available. So what
happens?They are seen customers who
have run inferencing by
deploying 1.5 billion parameter
model on a p5.48xlarge. What does
it mean?It means you are driving a
formula1 racecar to pick up
groceries. It is going to impact
in long time if you do not
optimise or right-size your
instances. Well SageMaker
Agentic AI recommendations ease
your answers. It is two fold. Number
one, it helps you benchmark
against your workload
preferences. SageMaker
AI collaborated with NVIDIA
AI Perf to benchmark your
workloads based on your S3
model artifact, your
workload preferences and your
business outcomes. Once the
benchmarking is completed, you
can review those results and
apply them using an S3
model artifact into your
dedicated endpoints. We also
go one step further. We provide
Agentic AI recommendations based
on the input parameters
which you specified. Our customers
were able to save weeks of
manual benchmarking
efforts into hours by using
Agentic AI recommendations
and benchmarking. Let's also
talk about other
challenges which our customers
have faced. Number one, what is
the most
scarcest resource when it
comes to GPUs?Capacity
availability. SageMaker
uses the same node pool as
EC2 compute. So if you
face insufficient capacity error
challenges, the endpoints
continue to try manual retries
until the instance is available. And
second, autoscaling does not
honour the priority. SageMaker
capacity aware inference helps
you provide and rank your
prioritised instances based on
your priority. Priority
one, for compute optimized. Priority
two, latency optimized. And
three, for your throughput
optimised. And it is
honoured during scale up and
scale down. So once you
define your steady state, the
inferencing workloads always
gravitate towards your best
case scenario in spite
of the spike in your
workloads. Last but not
the least, the observability
is at the granularity of
instance type so that
if you have a mixed fleet
instances in your workload, you
can have answers to
whether it is P5 instances
who are not providing P99
latency of sub
100 ms or
G instances. What's the
total utilization of my
fleet across P and
G instances. Let's
also take a step back and
understand about the SageMaker
AI OpenAI compatibility. Why
did we launch this
feature?Well, over
100% of open source <!-- 存疑：讲者原话，应为 "hundreds of" 或类似口误，保留原文 -->
frameworks, including
the Strands agent
and LangChain, they
use OpenAI
compatible APIs. By providing
support for the OpenAI
compatible, APIs, SageMaker
AI inferencing has
become
one click deployment with
just two lines of code. You
just change the base URL
and you provide
your endpoint and then the
tokens will take care of and
the key pair will take care
of the endpoint. Next
up. As I
called out earlier, observability
gives you granular
metrics at
instance level so
that you have full fleet visibility
of your multi-fleet deployments
wherein you can monitor
your health, you can monitor
your performance, and you can monitor
your capacity challenges. Step
2.It also provides
out of the box observability
capabilities, wherein
you don't have to separately
deploy the observability
dashboards or
operators to monitor
your cluster. Last
but not the least, it provides
inference component
visibility so
that you have knowledge
on the health and
performance monitoring of your
inference component visibility. Why
is it important?As we
have seen earlier, there are
various Agentic AI
workflows which failed from
POCs to production by
providing you the control
and configuration
of your
managed inference stack
SageMaker AI
helps you to
identify errors, identify
the bottlenecks, and then
help you solve those problems
so that you can efficiently translate
the price performance gains
over to your business outcomes
and extend your POCs
over to production workloads. And
it also gives you a
single pane of glass
for performance and health
troubleshooting. You
don't have to worry about
spinning up separate operators, separate
UI, separate dashboards
for monitoring your
observability metrics
of your SageMaker AI
for inference visibility. And
observability is
half of the problem
solved. SageMaker
managed inference lets
you have your dedicated
end points for
full control over your
container, whether you want
to optimize on DeepSeek
or Qwen or
Kimi K2.5 using
your custom containers, custom
model architectures, SageMaker
managed inference is the place for you.
We covered
the market trends for Agentic AI and SageMaker
managed inference, along
with understanding how
Agentic AI recommendations and
benchmarking helps you
save your
precious time and
increases time to value
for your Agentic AI applications to go to
production. With
all the
the session, I'm going to close
the session by giving you two key
take aways for you to deploy
your Agentic AI applications
in SageMaker AI.Whether
you have standardised on EKS
or you want to have a full
control over your deployment
architecture using
managed endpoints or
dedicated endpoints for
your workloads, SageMaker
has two paths to production for
your deployment and
the SageMaker AI platform provides
end to end
observability, granularity
and governance. You
can start with training. Serverless fine tuning, retraining, model
customisation and then extend
it all the way towards
deployment and inferencing. Thank
you so much for being a wonderful audience. If you
have any questions you can
find me in the hallway. 谢谢。