# 以 Agentic AI 重塑全球规模媒体运营

> **合规声明（生成式 AI 服务）**：前述特定亚马逊云科技生成式人工智能相关的服务目前在亚马逊云科技海外区域可用。亚马逊云科技中国区域相关云服务由西云数据和光环新网运营，具体信息以中国区域官网为准。
> 
> **完整资料：如果想了解更多内容可以访问:https://events.amazoncloud.cn/china-summit-on-demand/
> 
> **免责声明**：本摘要仅为便于您了解峰会内容而提供，我们不对其完整性及准确性作出任何保证，且不构成任何建议或结论。相关内容仅供参考，实际内容以峰会回放视频为准。

## 一、基础信息

- **会议类型**：专题演讲
- **Persona**：出海实践者
- **时间信息**：6月23日 | 15:00 - 15:30
- **标题**：以 Agentic AI 重塑全球规模媒体运营
- **PDF 资料**：有
- **视频回放**：有

## 演讲人信息

### 会议信息
- **地点**：上海世博中心（2026 亚马逊云科技中国峰会）
- **日期**：2026年6月23日（Day 1）
- **时间**：15:00 - 15:30
- **会议类型**：专题演讲
- **面向 Persona**：出海实践者

### 亚马逊云科技演讲人
- **Alastair Cousins**：Tech Leader - Media & Entertainment, 亚马逊云科技

## 二、会议资料 PDF 转录内容

--- Page 1 ---
Agentic AI
以 重塑全球规模媒体运营
Modernising global scale media operations with Agentic AI
Alastair Cousins
He/him
Tech Leader - Media & Entertainment
亚马逊云科技

--- Page 2 ---
Agenda • Transforming Media Supply Chains
with Agents
• Agentic Intelligent Operations for
Media & Entertainment

--- Page 3 ---
Agentic AI
使用 重塑媒体供应链
Transforming Media Supply Chains with Agentic AI
--- Page 4 ---
Sinclair: Vibe
编码供应链
Sinclair's Vibe-Coded Supply Chain
| Challenge Solution Outcomes
挑战 方案｜ 成果｜
• Kiro •
• 185
由 进行规格驱动的开发，由客户内部团队领导 需求驱动的工作流程
多个广播电台
• Amazon Step Functions •
•
将现有工作流程迁移到 编 使用内部开发的更易用的工作流程取代以前的商业工具
内容来自不同的制作人
Amazon Elemental MediaLive
排以及 和 • 100
•
MediaConvert 每年可节省 约 万美元
每天制作成千上万张全新的宣传片
• 185+ Broadcast stations • Spec-driven development with Kiro, led by • Demand-driven workflow. Only content needed for
internal teams broadcast is created.
• Content sourced from diverse range
of producers • Migration of existing workflows to Amazon • Entirely replaced previous proprietary tool with an
Step Functions orchestration and Amazon in-house developed workflow that is easy to modify
• Thousands of promos produced daily
Elemental MediaLive and MediaConvert
that are never broadcast • Estimated $1M/year in savings

--- Page 5 ---
现代化挑战：元数据
Modernization Challenge: Metadata
•
传统上，媒体档案是手动标记的
•
手动标记需要在采集时使用固定的分类法
•
缺乏灵活性
•
限制对内容的发现
• Media archives are traditionally manually tagged
• Manual tagging requires a fixed taxonomy at ingest
• Inflexible, not future proof
• Limits discovery of content

--- Page 6 ---
AI
将 应用于媒体分析
Applying AI to media analysis
特定于任务的模型 视觉语言模型 嵌入模型
Task-Specific models Vision-Language Embedding models
models (VLM)

--- Page 7 ---
详细任务分析
Task specific analysis
情绪分析
剧情简介和视频目录 Sentiment analysis
Synopsis & video catalog
（）
（）
审核与合规
Moderation & compliance
分类学注释
Taxonomy annotation ( )
（）
场景、镜头、主题和章节
Scene, shots, topic, and chapters
强迫叙事 （）
Forced narrative
（）
脚本
(Transcript)
字幕翻译
Translated caption
（）
品牌和赞助商商标
Brands & sponsor logos
（） 项目结构
Program structure (intro, recap, credit)
（）
名人和人物识别
Celebrity & People identification
（）
上下文元数据
Contextual metadata
（）
作品精选 广告片段
(Artwork selection (thumbnail)) Ad-breaks
（）

--- Page 8 ---
Media2Cloud
亚马逊云科技上 指南
Guidance for Media2Cloud

--- Page 9 ---
Vision-language models
视觉语言模型
Rich understanding Truly multimodal Zero-shot reasoning
丰富的理解力 真正的多模态 零样本推理
不只是 检测到人员，还有上下文描述
- →
Not just "person detected" but
视频 文字提示 自然语言回复 无需微调，只需询问即可
contextual descriptions Video + text Prompts → natural No fine-tuning needed—just ask
language responses
会议室里有两个人在讨论季度销售额
"Two people in conference room
discussing quarterly sales"

--- Page 10 ---
Vector embeddings
向量嵌入
数据的数字表示
A numerical representation of data
@@
必须将 非结构化数据（文本、图像、音频、视频）矢量
GenAI
化为矢量化才能在 应用程序中使用
Unstructured data (text, image, audio, video) must be vectorized
into vectors to be used in GenAI applications
嵌入模型
= Embedding model
相似向量 相似的含义和上下文
Similar vectors = similar meaning and context
通过比较向量距离，根据向量邻近度提供相似度搜索结果
Deliver similarity search results based on vector proximity by [0.743, 0.720, -0.325, 0.195, 0.835, -0.945]
comparing vector distances
n-dimensional vector

--- Page 11 ---
What is a vector embedding?
什么是向量嵌入
Query
查询
存在的本质和生命的意义 0.027 -0.011 … -0.023
Nature of existence and meaning of life
The Elegant Universe（优雅的宇宙） 0.025 -0.009 … -0.025
存在与时间 0.024 -0.012 … -0.021
Being and Time
嵌入模型
Embedding
Model
…
-0.011 0.021 0.013
哈利波特与魔法石
Harry Potter and the Philosopher's Stone
…
-0.009 0.019 0.015
魔术师
The Magicians
…
人类对意义的追求 -0.048 0.079 0.076
Man's search for meaning
Text
LLM
文本作为向量嵌入
文本
大语言模型

--- Page 12 ---
What is a vector embedding?
什么是向量嵌入
Query
查询
存在的本质和生命的意义 0.027 -0.011 … -0.023 Magic
Nature of existence and meaning of life
魔法 Universe
宇宙
Stars
…
0.025 -0.009 -0.025
星星
Harry
哈利波特
Potter Being
…
0.024 -0.012 -0.021
存在
Television
嵌入模型
Embedding 电视
…
Model -0.011 0.021 0.013
Time
哈利波特与魔法石 Movies
Harry Potter and the Philosopher's Stone 时间
电影
Film
-0.009 0.019 … 0.015 电影
Existence
存在
…
-0.048 0.079 0.076
Vector space
文本作为向量嵌入
LLM
大语言模型

--- Page 13 ---
Labels
标签
SMPTE 1,000
使用精确帧的 时间码，快速、一致地检测 多个物体、场景、活动
Fast, consistent detection of 1,000+ objects, scenes, activities, and
和人脸。
faces with frame-accurate SMPTE timecodes.
三者结合
Combine all three
非常适合结构索引，作为可靠搜索的基础。
Perfect for structural indexing and reliable search foundations.
Vision model descriptions
视觉模型描述
Natural language
通过多模态大语言模型对背景、情感和细微差别的自然语言理解。
三种方法结合：提供结构，视觉模型 understanding of context, sentiment, and nuance through multimodal LLMs.
增加了灵活性，嵌入支持语义搜索。
每种方法都捕捉了意义的不同维度。
支持灵活的分类和编辑评估。
Enables flexible classification and editorial assessment.
The optimal approach combines
all three: provides structure, vision
models add flexibility, and Vector embeddings
向量嵌入
embeddings enable semantic
Semantic similarity
search. Each captures different “查找类似内容” 查询的视觉、音频和文本模式的语义相似性。
across visual, audio, and text modalities for "find content like this" queries.
dimensions of meaning.
AI
为 驱动的推荐和发现提供支持。
Powers AI-driven recommendations and discovery.

--- Page 14 ---
我们需要一个媒体管理平台
We need a media management platform
特定任务：自动标记
Task-specific: auto-tagging
嵌入式：主搜索
Embeddings: primary search Rekognition；
可以检测物体、场景和面孔 结
→ openSearch/DynamoDB
Nova MME/Twelve Labs → k-nn vector
构化元数据
Rekognition detects objects, scenes, faces.
向量相似度（以毫秒为单位）
Structured metadata →
similarity in milliseconds
OpenSearch/DynamoDB
VLM
统一架构
Unified schema：丰富的元数据
VLM: rich metadata
Nova 2 /Twelve Labs /Anthropic；
组合矢量、标签、描述； 全面的查询结果
Combine vectors, tags, descriptions; 摘要 精选、点
播、优质内容
Comprehensive query results Nova 2 / Twelve Labs / Anthropic forsummaries;
Selective, on-demand, premium content

--- Page 15 ---
在亚马逊云科技上整合三者
Bringing the three together on 亚马逊云科技
Task-specific VLM Embeddings
特定任务 视觉模型 嵌入模型
Amazon Rekognition Amazon Bedrock Amazon Nova
Multimodal Embeddings
Amazon Transcribe Amazon Nova 2 Lite Twelve Labs Marengo 3.0
Amazon OpenSearch Service
Amazon S3 Vectors

--- Page 16 ---
Media Lake
亚马逊云科技 指南
Guidance for a Media Lake

--- Page 17 ---
— “
彭博媒体 新闻速度下的人工智能”
Bloomberg Media - “AI at the Speed of News”
NAB 2026 — NAB 2026 Project of the Year - Innovation Category
年度最佳项目 创新类别 ｜
Guidance for a Media Lake on Amazon •
自动制作多平台故事
Workflow Orchestration | Metadata Management | Semantic Understanding
• 13PB AI
来自直播和 存档的 驱动的片段汇编
1. Ingest 2. Analyze 3. Search 4. Understand 5. Create
•
用于故事差距检测的知识图谱
• +
垂直化 多格式再利用
Amazon Bedrock AgentCore — Orchestration + Parallel Agent Execution
• Automated multiplatform story production
Selection Assembly
Analysis Review Publish
Agent • AI-driven clip assembly from live and 13PB
Find Format cut, archive
Analyze clips using Content guardrails Push to channels
relevant Media platform
metadata + search & accuracy check + branding
(tone, sentiment) transcode
• Knowledge graph for story gap detection
Storage & Metadata Layer • Verticalization + multi format repurposing
testing Amazon S3
Live ingest | File import

--- Page 18 ---
Agentic
智能运营
Agentic Intelligent Operations
--- Page 19 ---
iHeartMedia
的人工智能运营革命
iHeartMedia's AIOps Revolution
| Challenge Solution Outcomes
挑战 方案｜ 成果｜
• 3000 • Amazon Bedrock AgentCore • 60%
多个广播电台和音频个性网站 使用 和 事件响应时间缩短
Strands SDK
• 1.5 实现了人工智能操作 • 30%
每月播客下载量超过 亿 待命负担减少了
•
• 集成了现有的可观测性工具 •
与数百台设备集成 知识保存和最佳做法应用的一致性
•
• 根据操作文档构建知识库
始终处于活动环境中，不容忍停机
• 3000+ radio stations and audio • Implemented AI Operations using • 60% reduction in incident response time
personality websites Amazon Bedrock AgentCore and
• Reduced on-call burden by 30%
Strands SDK
• > 150 million podcast downloads per
• Knowledge preservation and consistency
month • Integrated existing observability tools
of application of best practices
• Integrations with hundreds of devices • Built knowledgebase from
operational documentation
• Always live environment with no
tolerance for downtime

--- Page 20 ---
Agentic
智能运营
Agentic Intelligent Operations
Demo
演示 ｜

--- Page 22 ---
下一步行动
Next Steps
Guidance for Guidance for a DevOps Agent
Media2Cloud Media Lake

--- Page 23 ---
Thank you

## 三、会议回放视频转录内容

Thank you. Hello. It's working.
Nihao, this session is modernizing global scale media operations with Agentic AI.
My name is Alastair. My role at 亚马逊云科技 is I am the tech leader for media and entertainment.
And in that role what I do is I work with 亚马逊云科技 media and entertainment teams globally to understand customer trends.
I'm emerging solutions in the industry and I help to share that knowledge around the world with customers and our internal teams.
I'm really honoured to be here. The China team invited me to come. This is my second time in China presenting at a native 亚马逊云科技 event.
It's been a great trip so far. I am going to share with you some stories of global trends around applying Agentic AI to media and entertainment workloads in the next 30 minutes.
To start with, I'm going to talk about media supply chain.
So the media supply chain is everything that comes from acquiring content into a media business, processing and then finally producing that into a form that it can be delivered out to viewers and end users.
And this is an area where we are seeing a lot of transformation happen with media and entertainment customers through the application of generative AI and Agentic AI.
A customer example that I am really proud of, that this customer Sinclair broadcast group recently spoke at the NAB conference in Las Vegas with 亚马逊云科技 talking about this transformation that they went on over the last six months or so.
And during this journey Sinclair came to us to asking us to help them evaluate if they could move their media supply chain to the cloud.
And they brought some interesting challenges to us. They shared that they have over 185 broadcast stations. These are all in America.
Sinclair's business model is they are a local broadcaster in an affiliate model.
So 185 stations, 185 brands, all of them local.
What this leads to is that Sinclair's centralized media operations team received content from a huge variety of sources.
Thousands of vendors from very large production houses all the way down to individuals working in their garage or their back office.
And Sinclair over the years have developed and built a powerful pipeline for processing all of that media.
And they came to us and asked if we could look at moving that to the cloud.
And so the solution that we came up with and that we worked with Sinclair on was we introduced them to the idea of spec-driven development using a tool like Kiro.
And what we did with them is we worked with their engineers, many of whom were not developers, to learn about how they could translate their business requirements into specifications.
And then by bringing those specifications into their development environment, Kiro was then able to work with its agents to develop workflows that run in the 亚马逊云科技 cloud.
So what that resulted in is deployments using services like Amazon Step Functions and Elemental MediaConvert and MediaLive to perform all of the processing required on the media that comes into Sinclair on a daily basis.
So we worked with those teams to develop this and the net result was actually quite powerful.
One of the key things is we were able to move the workload to the cloud.
So that was powerful because it allowed Sinclair to retire some legacy equipment, some proprietary systems that were reaching end of life.
So they got a win by being able to move to a newer platform, a new more modern platform.
But interestingly as the engineers gained more skills with spec-driven development, they realized that they didn't just have to have the opportunity to lift and shift the existing workload,
but they were able to actually adapt the workflow to things that they really wanted to do but couldn't do with the old system.
And one of the key things that they implemented was they moved from a system where every single piece of media that came in was treated the same to a demand-driven workflow.
And in this case what this resulted in was as new content arrived at Sinclair, they would generate a promo, a short promo video for every piece of content
for every one of those 185 stations. Regardless of whether that content went to air on that station or not,
Sinclair were able to modify their pipeline to actually look at the schedules and understand if the content was needed
and only render out promotions that were needed for the stations that were going to air that content.
The result of that was estimated by Sinclair at over a million dollars a year in savings from compute and resources required to do that work.
So that's a massive dividend and being able to now move faster with agility thanks to the power of spec-driven development and Agentic AI.
So that's a real world story. We're having this conversation with many, many media companies around the world.
And I thought it would be good to go back to some of the theory and some of the common challenges we're seeing as we work with those customers.
Now for media organizations, a lot of the value in the business is the content, is the media assets that are being produced by the organization.
And most media companies then have an archive of this material.
They have a large archive, often petabyte scale of this material that they're keeping on tape or in a service like Glacier in the cloud.
When we work with customers to bring those archives to the cloud and to modernize, one of the things we consistently run into is challenges around metadata.
Because a lot of this content was created many years ago, 10, 20, 30 years ago.
And it was indexed using traditional techniques like you see visualized here.
A human watched the content and had to make notes on what was in the content, typically against a fixed taxonomy.
Title, date, location, etc.
That taxonomy was set at the time of ingest, not at the time of discovery.
And so that has downstream impacts.
When agents or automated systems want to locate media, they can only discover it based on the metadata that has been recorded against it.
And so for older content, it can make discovering relevant content difficult due to these, due to the limited metadata attached.
So, I've been at Amazon for 13 years.
I think I've been working with media customers for almost that entire time talking about how you apply AI to your media archive to discover more characteristics, more metadata.
And so, since the emergence of generative AI, this whole landscape has changed.
So the last two or three years, we've been seeing dramatic improvements around what we can do with AI.
And I'm going to talk about the three key techniques that we take customers through.
Task-specific models, Vision Language Models, and embedding models.
So, starting with task-specific analysis, this is what I would probably frame as traditional AI.
So these are domain-specific models that do functions like face detection, transcription, brand and logo detection, sentiment analysis.
So, typically we're talking about convolutional models, services like Amazon Rekognition, that are very powerful at doing one thing at a time when you pass media through them.
We've been using these for a decade, and we've been able to extract much richer datasets about our media by applying them.
However, what we've seen with customers is when we apply these task-specific models to their archive,
there's a lot of questions that come up about which models should we use, how often should we process our media,
at what indexing rate, like every frame, the entire video, key sections, how regularly do we sample our material to gather metadata.
And so, members of my team have put together a, what we call a guidance package, guidance for Media2Cloud, which is published on GitHub.
It's open source and available, so you can take it and extend it, which shows best practices.
On how to implement task-specific AI for extracting metadata from your media.
But now moving beyond task-specific AI, since, as I mentioned, since the emergence of generative AI, some new approaches have come along,
because those task-based models are still delivering us a fixed taxonomy, they're still delivering us, ok, these faces appeared, this is the transcript.
They're very much on rails, within some boundaries. Vision language models have been a new development.
Vision language models are models that have been trained on multiple modalities, text, image, video, audio.
So Vision Language Models are able to take from media in one modality and then describe it in another.
So the common use case here is taking an image or a video file and prompting the model and asking it to describe what's happening.
Very simple. And a model like Amazon Nova 2 is able to give a very effective rich text description of what's going on in an image or a video file that is unstructured.
It's not tied to a fixed taxonomy. And this becomes very, very powerful for discovery use cases downstream.
Because moving outside of a fixed taxonomy, we're now generating new keywords that we didn't anticipate.
Based on the actual media material, rather than a fixed taxonomy that we decided ten years ago was going to be the guardrails we were going to use here.
So applying Vision Language Models is proving to be very useful for discovery.
But this last technique, vector embeddings, is probably the most exciting development.
Embedding models are also multimodal models. They're models that have been trained on a variety of different data sets.
But embedding models allow us to extract from the model the vector representation of the input.
And this is a very powerful concept. When we're working with these generative AI models, one of the first steps that happens is the input, say an image of a dog needs to be translated into something that matters.
And this makes sense to the model. And the way these models work under the hood is they use vectors, mathematical representations of that input data set, these pixels, in a multi-dimensional space.
Thousands of dimensions. Now the interesting thing about embeddings is similar vectors to similar sets of embeddings that are close to each other, translate to similar meanings or context.
And this is really interesting from a search perspective. Because this is not being weighted by a human, by a librarian, by anyone in the business.
The relationship emerges from the training of the model originally. So it's outside of our control.
But those relationships because these vectors are thousands of dimensions deep, also can be very unexpected and subtle.
So, you know, a very simple example might be an image of a car and an image of an engine. Now you and I would probably relate a car to an engine quite naturally.
But unless that's in a fixed taxonomy and you search for car, you're not going to get engine in your results.
With embeddings, the car and the engine will be near each other in some of the dimensions of the vectors, which means that we can do a search within the vector space to identify things that are in close proximity and get related results that we didn't anticipate when designing our taxonomy up front.
To show a simple representation, hopefully simple. Sorry, there's a bit of math.
Taking a bunch of unrelated, similarly unrelated terms, like for example Harry Potter and the Philosopher's Stone, you know, a children's novel.
And the magicians, the concept of magic. These two are related, but the words don't necessarily have any tie to each other.
These all generate different embeddings. But if I was to visualize these, and I can't draw a thousand dimensional model on a screen, but let's just reduce this to two dimensions.
What you start to see is related terms sit close to each other in the space. So Harry Potter and magic are close to each other.
Television, movies and film are close to each other as words. And because embedding models are often multimodal, like Nova multimodal embeddings, we can translate from language to image to video naturally with these models. This becomes really, really powerful.
So customers then come to me and my team and ask which approach should we use. And the answer is really all three.I would love to say just use embeddings or just use a Vision Language Model.
But the reality is this search requires lots of nuance and getting the best results for our end users really requires that we bring all three approaches together.
And balance and weight between them to get the right result for the end user. That then leads our customers to the next challenge, which is ultimately we're going to need a tool.
We're going to need a platform to bring these, all these different techniques together.I now have fixed metadata.I have free text coming out of my Vision Language Model.
And I have embeddings, which are essentially numbers, don't seem to have any meaning to a human. Very three very different types of material that I now need to balance in order to get results.
So when we do this in 亚马逊云科技, the way we think about it is, on the top we have a variety of services that you can be used to actually obtain and extract this information.
We can use things like Rekognition and Transcribe as task specific services. We can use Bedrock and Nova 2 Lite for Vision Language Model.
We can use Nova multimodal embeddings or our partner Twelve Labs, the Marengo model to extract multimodal embeddings from our media content.
And then we need a storage service and a search service to put all this together. And OpenSearch is the most common place we're seeing customers deploy solutions today.
And then search lets us search across keywords, doing language searches, and also it supports vector embeddings.
Also our friends in the Amazon S3 team recently launched S3 Vectors. And S3 Vectors is a scalable vector store that allows search going up to petabyte scale, which is a really interesting problem that we're starting to see emerge with some of these large media archives.
Because when we start to extract embeddings from our content, you don't just take one embedding of a one hour video clip. You take an embedding for the audio, you take an embedding for the video, you take an embedding for the transcript.
But then you don't do that for the one hour of material, you do it every six seconds or every 10 seconds.
Which means that for a piece of media in your archive, you may have a couple of hundred kilobytes of fixed metadata, but you may be, you're going to be in the megabytes, possibly even the gigabytes of vectors very quickly, having all these samples to be able to detect key moments within your content.
So S3 Vectors is an emerging option here for the vector search at really large scale.
Now, my team has also put together a reference architecture for how to put all this together in the cloud. And this is something we call guidance for a media lake.
This is another open source project that we're working on internally, and we have a number of customers and partners also contributing to on GitHub.
And so guidance for a media lake shows how to extract vectors and embeddings from your media, as well as applying different AI services to extract fixed metadata.
And then it presents all of this into a simple application that allows users to discover their media and search for it with different techniques.
We released this earlier this year and this has been resonating a lot with customers around the world.
And I'm going to talk about an example now of a customer who's taken this and has extended it further in their own environment.
So Bloomberg Media. So Bloomberg is a financial news agency as well as data provider to the finance sector.
And Bloomberg have been working with us with guidance for a media lake to build a media management platform for their 13 petabyte archive of material spanning decades.
Bloomberg have been extending the guidance for a media lake to do exactly what I just talked about, to link those fixed taxonomies and fixed metadata into the embeddings that have been generated.
And then they're starting to build an Agentic platform using Bedrock AgentCore on top of this data to start facilitating the automated assembly and publication of content to different platforms.
Now Bloomberg have an extremely high bar for quality.
Being in the finance sector, financial news data is critical and reliability is critical to their brand and to their reputation.
So Bloomberg are working with us to actually take all of the data in the media lake and start to develop things called knowledge graphs.
And this is stuff we're working on in the open with the media lake GitHub project as well, which is to start extending all of the metadata we're now collecting in our ingest workflows in media lake into knowledge graphs that can then be used by agents to analyze the media, select content, compose it together into stories to do gap analysis and to work with journalists to then publish it.
To different channels in different formats. So vertical video, horizontal video, different voicings for premium data driven channels versus social media, for example as well.
So we won an award at NAB for project of the year around this collaboration we've been doing with Bloomberg and this is a pretty exciting story that we're partnering with them on and continuing to invest in.
Now I have about 10 minutes left and I wanted to talk to you about another workload that I think is this is close to my heart and I think is also probably one of the most impactful workloads for media organizations worldwide.
As I mentioned before, I've been at 亚马逊云科技 for about 13 years and I have seen a couple of waves of change come through to developers and operational teams.
In particular what I'd call out is the emergence of DevOps. 15 years ago most organizations had a development team and an operations team and they, the two teams often didn't talk to each other.
Code would be thrown over the wall and deployments were painful and difficult. DevOps changed the culture of organizations to have developers actually build and run systems at scale using automation and safeguards from cloud services to do that at scale.
I think the next layer of this is applying Agentic AI to how we deploy, scale and operate systems in the cloud because something that has occurred downstream from the adoption of DevOps in many organizations is engineers are on call and they're having to respond to incidents that are increasingly complicated.
Many of our media customers are now running events at tens of millions of concurrent user scale and what that means is there is always something that is breaking.
There's always a long tail event and it's not scalable or effective to tie up engineering resources on every single incident to troubleshoot those.
Agentic operations is applying agents to assist our engineers with troubleshooting and managing services at scale.
Now a customer who has implemented this with some success is iHeartMedia.
So iHeartMedia is a global radio platform.
They operate over 3000 radio stations around the world as well as running a podcast platform that delivers about 150 million downloads per month.
So pretty significant scale.
iHeartRadio's overall strategy is they want to be on every device.
So they integrate with over 150 devices and platforms worldwide.
So smart speakers, smart TVs, gaming platforms, all sorts of embedded devices have iHeartMedia in their marketplaces or embedded into their app stores.
When you do the math on this, when you've got thousands of stations, when you've got hundreds of device platforms
and you've got a global audience, that means there is no safe time to do maintenance or experience outages.
There is always going to be an audience who are impacted by service incidents.
And so iHeartMedia were experiencing increasing toil with their engineering and operations teams having to respond to these incidents around the clock.
Their on-call teams were being tied up with work keeping the platform running.
And so they started to experiment last year with AI operations using Bedrock AgentCore and the Strands SDK.
And what they did was they integrated all their existing observability tools.
So they were using a variety of proprietary systems for and partner systems for logging as well as internal tools.
And they also integrated their own knowledge bases, their own runbooks for how to respond to incidents in their wikis and other documents.
And they built a multi-agent system that was able to respond to incidents and events,
follow the runbooks to troubleshoot, extract logs, look at metrics,
and then make recommendations to on-call operators as to the next step to remediate incidents.
The dividend they've seen from this is quite remarkable.
They've seen a 60% reduction in incident response time,
which for a 24-7 business is pretty powerful.
That's a big impact.
But the thing I really like, because I came from a background doing this kind of work,
is they're also seeing a 30% reduction in on-call burden,
which basically means a lot less pages for engineers.
And when those engineers are being paged,
they're just there to approve a remediation for an incident.
So they're able to jump in, fix the thing,
and go back to their lives when they're working out of hours much more effectively.
That's a great win-win.
If you're responding to incidents more quickly and your engineers are happier
because they're not being paged as much,
that's great for the organization.
And that's going to deliver customer outcomes through
things like increased feature velocity
and delivering other things that are more important to the organization.
So my team have put together a demo of what this can look like.
It's kind of hard sometimes to talk about log analysis.
It's very text-heavy. It's not very exciting.
So we've put together a visual representation
of what this looks like for a media organization.
And I'm going to walk you through a quick tour
of what a multi-agent system
for media operations can look like.
And I just want to suggest that
I think in the next 12 months,
most media organizations will be starting
to do something like this
because this is going to be powerful
for all the engineers.
And what we've seen in the customers
who have adopted it is
as this workload is very close
to the people who are building or implementing it,
it's a great way to adopt
or to learn how to adopt Agentic AI.
It's an early workload
and the win goes directly back
to the builders who put it together.
The engineers who build the solution
get their time back in terms of reduced on-call workload.
So this is a really nice way
for an engineering team to get comfortable
with Agentic AI
and to deliver a business outcome,
which is pretty exciting.
So in our system,
we have a centralized coordinator agent
we've named El Capitan.
And then we have
individual specialist agents
for a number of services
within our architecture.
So we have an agent for MediaConnect
which is handling our media contribution,
an agent for MediaLive,
which is handling our live transcoding
and an agent for observing
the end user experience.
Looking at CMCD logs
and CDN logs to see how users
are actually experiencing the streaming video.
Now when I run the video
what's actually happening here now
is El Capitan is configured
with a number of tools.
So it's been given skills
on how to analyze requests
and how to actually start
troubleshooting problems
and has access to knowledge bases.
It is also connected via MCP
to each one of these
specialized agents
and each one of those agents
has its own tools configured to be able to
the state and configuration of the resources
and configuration
as well as to access tools
like Athena for example
to perform queries on logs
at scale and extract
insights from those logs.
We have also then integrated
an alarm flow. So we've integrated
both CloudWatch alarms
as well as an anomaly detection flow
taking the client-side logs
through Managed Flink
and performing anomaly detection
on the client-side logs
to detect things like buffering events.
These two alarm sources
ultimately go into
a decision tree
which then decides whether to trigger El Capitan
to investigate an incident.
So this is something we can develop
with spec-driven development as well.
Now in our demo environment
we have a sample channel here
a streaming channel which we've rigged
with logs coming from each of these sources
from the source all the way through to delivery.
And then we start to look at
an example of what the operator view could look like.
So from here we have the channel once again
but up here we can now interrogate
the agents as the channel is playing.
So we can chat
with El Capitan.
So an operator can ask El Capitan
what capabilities do you have
or run a full health check across the pipeline
because maybe there's an incident that needs investigating.
But El Capitan is also being triggered
by those alarms.
And so here we have logs of past
incidents that El Capitan has taken a look at.
So our on call operator might get paged
and get asked to investigate a specific ticket.
The view here is you pull it up
and then what we will see
here is the
incident report
that El Capitan has automatically generated
by working between all the specialist agents
to pull data from the
services and the logs
and then synthesizing that.
So it's made a key finding
that this is actually a true problem.
It's a real incident.
There's a split state condition
where the status of the flow is showing
active but there's actually no
video flowing through.
And then it gives us a more detailed
diagnosis. It's identified the
average bit rate dropped
and it's seen that the packet count
in our metrics has dropped at the same time.
It's seen downstream errors from this
and it provides to the operator
based on our run books and our knowledge base
recommended steps to remediate the problem.
So the experience I'm sharing with you is
imagine you're an on-call operator
you get paged into an incident
the ideal scenario is
rather than having to run all those
checks yourself
you log in and you immediately are
given this report. So you're immediately
given an investigation of what's been
going on in the system
and now you can potentially take action. You
have recommendations, you have the data
summarized and you can go deeper into it.
To assist those engineers as well
one thing our team has done and they've
learned this from our customers
is they've also included the Prompt
that was sent through the system to
allow that to be reproduced or to be
tweaked and tuned. If the report gives us
an inconclusive response I could copy
and paste this, modify it as
appropriate and send it back to El Capitan
to get a more detailed result. So this aids
the on-call operator with dealing
with the incident and trying to get
to the root cause. And so
that concludes
what I wanted to share with you today.
I already shared these
links but here they are again for
Media to Cloud and Media Lake. For the
intelligent operations piece, I believe
there's a number of talks going on
today and tomorrow talking about DevOps
Agent which is our new managed
service that focuses on this intelligent
operations space. We're seeing this
is a big trend across all industries, not just
media and entertainment and I really encourage
you to take a look at DevOps Agent and
see if it could work for your business. It's
going on there. If you're from the
media industry and you're interested in
these kind of customer stories, my
colleague Steve Sanwell is also
talking tomorrow down on
the Expo floor at 11am. He'll
be talking about some European
customers and some similar stories
around broadcast transformations
at scale that have happened in
Europe. I've used some American
examples. We also have a number
of great demos that the local
China team has put together
amongst various media workflows and
applying Agentic AI to those. So
really encourage you to spend a bit of
time with the team, learn a bit about
what's possible in Amazon Web
Services for your media workloads. So
with that, thank you for the time.
Really appreciate you all listening
and this has been fun. So
hope you enjoy the show.