A guide to the anatomy of effective commerce agents
高效电商智能体全景架构指南:从单模型闭环到毫秒级延迟与生产级评测
Over the past year, we've worked with teams across the commerce industry — retailers, marketplaces, travel, entertainment, and telecom providers — to build commerce agents using Claude.
These agents are in production, and enterprise customers have seen larger carts and more efficient seller operations when using them. They also share a simple architecture: Claude in an agent loop equipped with a set of skills, tools, and a strong eval suite.
This post is for the engineers and engineering leaders building these (or other consumer facing) agents. Part 1 covers the architecture, which you decide once. Part 2 covers latency and cost. Part 3 covers production: memory, safety, evals, and scaling the work across an organization.
在过去的一年中,我们与电商行业的众多顶尖团队展开了深度合作(涵盖零售商、综合交易市场、在线差旅平台、文娱演出票务以及电信运营商),基于 Claude 构建实战级电商智能体。
这些智能体目前均已投产上线,企业客户在使用后不仅见证了客单价与购物车容量的显著提升,还实现了更高效的商家端运营协同。它们共享着一套简洁统一的底层架构:处于智能体闭环中的 Claude,配备一组专项 Skills 技能、工具链以及完备严格的评测套件。
本文专为构建此类电商(或其他面向 C 端消费者)智能体的工程师与技术管理者撰写。第一部分探讨全景架构设计,这是只需一次定型长期受益的决策;第二部分聚焦延迟与成本优化;第三部分覆盖生产级落地运营:跨会话记忆、运行时安全防线、自动化评测以及跨大型组织的高效研发协同。
In this guide
- Part 1: The architecture
- What is a commerce agent?
- Skills, not subagents
- System prompt or skill: decide by frequency
- Engineering agent tooling
- The UI components are tools
- Part 2: Making it fast and affordable
- Minimizing task completion latency
- Perceived latency
- Prompt caching
- Choosing the model and its configuration
- Part 3: Running it in production
- Memory that survives the session
- Safety: enforcement lives in the harness
- Evals: shipping a non-deterministic system
- Shipping with a large organization
本指南核心目录 · In this guide
- 第一部分:全景架构设计(Part 1: The architecture)
- 什么是电商智能体?(What is a commerce agent?)
- 采用 Skills 技能,而非子智能体(Skills, not subagents)
- 写入系统提示词还是沉淀为 Skill:依据频次裁决(System prompt or skill: decide by frequency)
- 智能体工具链工程(Engineering agent tooling)
- 前端交互组件本身就是工具(The UI components are tools)
- 第二部分:极速与低成本优化(Part 2: Making it fast and affordable)
- 最小化任务完成物理耗时(Minimizing task completion latency)
- 感知延迟优化(Perceived latency)
- 提示词前缀缓存(Prompt caching)
- 模型选型与配置权衡(Choosing the model and its configuration)
- 第三部分:生产环境落地与运营(Part 3: Running it in production)
- 跨会话持久记忆(Memory that survives the session)
- 安全防线:在运行时底座中强制执行(Safety: enforcement lives in the harness)
- 评测体系:发布非确定性系统(Evals: shipping a non-deterministic system)
- 大型企业研发团队协同(Shipping with a large organization)
The architecture
全景架构设计
What is a commerce agent?
We define a commerce agent as an agent that simplifies buying and selling across an online catalog.
Some agents face consumers: they search, compare, substitute, and assemble the order. That could be a retail cart, a travel itinerary, a mobile plan change, or seats held for a show. Some agents face the business: they answer questions about sales, run promotions and campaigns, and manage inventory and pricing.
什么是电商智能体?
我们将电商智能体(Commerce Agent)定义为:能够围绕线上商品目录简化购买与销售全流程的智能体系统。
部分智能体面向消费者(C 端):负责商品搜索、比价推荐、智能替换与购物车组装。这可以涵盖日常零售加购、定制旅行行程、手机套餐变更,或是剧场选座锁单。另一部分智能体则面向商家与运营业务(B 端):负责解答销售数据疑问、配置促销与营销活动、以及统筹库存调度与动态调价。


- Standard Agent Loop
- 标准智能体闭环:单一大模型在循环中自主推理目标、探索上下文、调用工具、读取 Skills 并观察反馈。
- Skills vs Subagents
- 按需加载长尾能力:将各业务域规范沉淀为 Skills 指令按需注入,避免子智能体切换造成的上下文丢失与通信延迟。
- Harness & Memory
- 运行时底座与跨会话记忆:外围底座强制执行安全门禁与限额校验,记忆模块实现跨会话个性化偏好留存。
The core architecture is a model in a standard agent loop: reasoning about a goal, exploring context, taking actions through tools, learning procedures through skills, asking clarifying questions, and observing the results until the goal is accomplished.
There is no intent router in front of it that segments the conversation and no set of domain specific agents behind it.
其核心架构是处于标准智能体闭环(standard agent loop)中的单一基础大模型:围绕既定目标展开推理、探索上下文、通过 Tools 执行操作、借助 Skills 学习标准作业流程、提出澄清性提问,并持续观察环境反馈直至目标彻底达成。
在它的前方,既没有负责粗暴切分对话的意图路由器(Intent Router),后方也没有一整套割裂的垂直领域子智能体集群。
Engineering context
Skills, not subagents
A commerce agent has to cover a wide range of capabilities across many categories and intents, which makes it tempting to create one subagent per domain.
In practice this proves suboptimal, because a commerce conversation is one tightly coupled session across multiple intents and turns, and requires considerable shared context.
In a subagent architecture, the orchestrator holds the cart or staged changes, the user's preferences, and the conversation history.
Every handoff to a subagent is a state-lossy operation, which often impacts the quality of the subagent’s response and, consequently, the overall response. On top of that, each handoff can cost several times the tokens and adds seconds of latency.
The domains also rarely separate cleanly. A returns flow might need the order history, the current cart, and the product catalog, meaning a subagent-per-domain approach either duplicates that access everywhere or hands off mid-task.
As models get smarter, they also handle longer context, more skills, and more tools, so the limits behind today's placement rules loosen with each model generation.
Instead, agent skills give you similar per-domain modularity and context control without the handoff tax, because the skill instructions load into the main agent that already holds the entire history.
In our comparisons across several enterprise deployments, a single agent with skills consistently has outperformed both the one-prompt-for-everything design and the subagent design on quality, and often at a lower cost and latency per task.
上下文工程化设计(Engineering context)
采用 Skills 技能,而非子智能体(Skills, not subagents)
电商智能体必须横跨众多品类与复杂意图提供广泛能力,这很容易诱导开发者陷入“每个垂直业务域单独搭建一个子智能体”的设计陷阱。
然而在实际生产中,这种架构被证明是次优的:因为一次电商交互是一场跨多意图、多轮次高度紧密耦合的会话,需要大量共享的全局上下文。
在子智能体架构中,顶层编排器维护着购物车暂存变更、用户偏好与完整对话历史。
每一次向子智能体的上下文交接(Handoff)都是一次有损的状态压缩操作,往往会损害子智能体的回答质量,进而拖累整体体验。更严重的是,每次交接都会消耗数倍的 Token 并额外增加数秒的延迟。
此外,电商业务域之间极少能做到绝对干净的解耦。例如一个退换货流程可能同时需要查询历史订单、当前购物车和全量商品目录;采用多子智能体方案,要么在每个子智能体中重复冗余这些权限,要么在中途频繁进行上下文交接。
随着基础模型变得越来越强大,模型处理超长上下文、更多 Skills 和丰富工具链的能力持续跃升,限制过去放置规则的瓶颈在每一代新模型中都在不断被打破。
相比之下,Agent Skills(智能体技能)能在免除交接损耗的前提下,赋予你同样的模块化与上下文掌控力:因为 Skill 指令是直接按需加载到已经掌握全部会话历史的主智能体中的。
在我们在多家大型企业部署的横向对比中,搭载 Skills 的单一主智能体架构在回答质量上全面超越了“单一大杂烩提示词”和“多子智能体分发”设计,并且通常具备更低的单任务成本与延迟。
Where subagents do earn their place is when the orchestrator can call them as a tool for a narrow or self-contained task that would benefit from its own dedicated context window.
A common production example is a deep-research subagent, where the subagent searches and reads documents, writes and runs code, traverses data models, and hits dead ends. All the work happens inside one or more subagents, and only a compact answer comes back to the orchestrator.
The other exception is a domain that already has its own purpose-built agent. If your pharmacy or financial-services experience runs a dedicated agent with its own compliance surface, the right move can be a hand-off, where that agent takes over the task and works with the user directly through its own loop until the task is done.
The distinction is ownership of the conversation. A hand-off makes the domain agent the user's counterpart, while delegation keeps the orchestrator, bouncing the domain agent in and out within a single turn and degrading on every exchange.
子智能体真正能发挥价值的场景,是编排器将其作为一项独立工具调用,用于处理高度独立、在专属上下文窗口中封闭推演的窄任务。
一个典型的生产范例是深度调研子智能体(Deep-research Subagent):子智能体在后台海量检索阅读文档、编写并执行代码、遍历复杂数据模型并排查死路;所有探索性工作均在子智能体内部完成,最终仅向主编排器回传一份精炼的总结报告。
另一项例外是那些本身已具备专属合规门槛的独立领域系统。例如在线药房处方审核或金融信贷服务已拥有专属合规智能体,此时最佳方案是进行彻底的主权移交(Hand-off):由该领域智能体全面接管对话,直接与用户交互直至完成任务。
这里的关键区别在于对话所有权归属。彻底移交(Hand-off)让领域智能体成为用户的直接对话者;而委托调用(Delegation)则让主编排器在单轮内反复进出子智能体,在每一次交互中产生损耗。
System prompt or skill: decide by frequency
The main factor when deciding whether to put a set of instructions within a system prompt or skill is how often the agent will need it. Loading a skill costs a model turn, so anything the agent needs on most turns generally goes in the system prompt.
This does, however, depend on how your traffic is distributed, and what agent behavior your evals show. A good starting point is that anything relevant to a third or more of your traffic, whether anticipated before launch or observed in production, goes in the system prompt, and the rest goes in skills.
If a skill is predictable from a signal you already have, such as the page the user arrived from, we recommend injecting it from the harness before the first model call and skipping the extra turn to load the skill.
Critical instructions, such as safety and legal rules, brand constraints, and key user facts such as allergies, always go in the system prompt.
For commerce agents, this means product search lives in the prompt, since nearly every session touches it, and skills carry the long tail of features.
In our reference implementation, the shopping agent's prompt holds grounding, cart and checkout semantics, and presentation rules, and the following skills cover the rest: search-discovery, purchase-research, planning-goals, customer-care, and memory-personalization.
The merchant agent splits the same way, with performance-insights, catalog-listings, inventory-operations, pricing-promotions, and marketing-campaigns as its skills, one per operational domain.
写入系统提示词还是沉淀为 Skill:依据调用频次裁决
决定将一组指令写入系统提示词(System Prompt)还是沉淀为 Skill 的核心评判标准在于:智能体在日常运行中需要它的频次有多高。加载一项 Skill 会消耗一次模型交互轮次,因此智能体在绝大多数轮次都需要的内容通常直接置于系统提示词中。
当然,这也取决于你的实际流量分布以及评测中体现的智能体行为特征。一个实用的黄金分界法则是:凡是与 1/3(33%)及以上线上流量相关的通用规则,无论是上线前预估还是生产实测所得,一律放入系统提示词;其余长尾业务逻辑全部沉淀为 Skills。
如果你可以通过既有确定性信号(例如用户来源的落地页路径)提前预测所需 Skill,建议在发起第一次模型调用前由运行时底座(Harness)直接预注入,从而省去智能体主动加载 Skill 的额外轮次开销。
涉及安全合规底线、法律合规准则、品牌调性约束以及用户关键硬性事实(例如食物过敏史)的核心指令,必须始终常驻于系统提示词中。
对于电商智能体而言,这意味着商品搜索常驻提示词(因为几乎每个会话都会触发),而长尾特性则由 Skills 承载。
在我们的 官方参考实现 中,导购智能体的系统提示词固化了事实对齐、购物车与结账语义以及 UI 呈现规则;以下 5 大 Skills 则覆盖长尾场景:search-discovery(搜索发现)、purchase-research(深度导购对比)、planning-goals(目标规划)、customer-care(客户关怀)以及 memory-personalization(个性化记忆)。
商家运营智能体采用相同的拆分模式:以 performance-insights(经营分析)、catalog-listings(商品上架)、inventory-operations(库存操作)、pricing-promotions(定价促销)和 marketing-campaigns(营销活动)作为各个业务域的专用 Skills。
Engineering agent tooling
Our post on writing effective tools for agents covers tool design in general. Two points have mattered most in commerce:
Build agent tools on top of your core systems and logic.
A commerce company already has search and ranking, a cart, a preferences and profile store, an inventory system, promotion and campaign engines, sales analytics, and more, each encoding logic tuned over years and seeing signals the model never will.
The agent's tools should call those systems, not reimplement them, and the tool boundary is where their logic ends and the model's judgment takes over.
For example, when the agent calls search_products, the results should arrive already ranked; its job is to decide which results serve the user's goal, how many to show, and how to present them.
智能体工具链工程(Engineering agent tooling)
我们在《为智能体编写高效工具》一文中探讨了通用的工具设计原则。在电商场景中,以下两点至关重要:
基于企业既有核心系统与成熟业务逻辑构建工具。
一家成熟的电商公司早已拥有商品搜索与排序引擎、购物车系统、用户画像库、库存中心、营销与促销引擎以及销售分析大盘;这些系统凝聚了多年沉淀调优的算法逻辑,并能感知到大模型自身永远无法看到的实时特征信号。
智能体的 Tools 应当直接调用这些既有系统,而不是在模型侧拙劣地重新实现它们;工具边界正是传统确定性算法逻辑的终点,也是大模型常识推理与柔性判断力接管的起点。
例如,当智能体调用 search_products 时,返回的商品列表应当已经由后端算法完成了粗排与精排;智能体的核心任务是判断哪些候选最契合用户当下的真实意图、展示多少条以及如何组织呈现。
Tool results are context.
Return the fields the model reasons with and drop the rest. Image URLs on every search row are the usual offender.
As needed, reshape the raw response inside the tool, including appending a next step when it isn't obvious from the data.
This is especially relevant for error scenarios, where the model benefits from instructions instead of error codes. For example, add an error instruction "Include a product ID when querying availability," instead of a generic 403.
工具返回的每一项数据都是宝贵的上下文。
只返回模型推理真正需要的字段,果断剔除其余冗余数据。每行搜索结果中携带的图片 URL 通常就是最典型的冗余累赘。
根据实际需要,在工具内部对原始接口响应进行重塑加工,包括在数据含义不直观时代为追加下一步建议或引导指令。
这在异常与错误场景中尤为关键:大模型极度依赖清晰明确的恢复与纠错指令,而非晦涩难懂的冰冷错误码。例如,可以直接返回一段提示:“未找到符合条件的商品,建议引导用户放宽价格筛选或更换同类品类关键词”。
The UI components are tools
Most commerce agent responses are UI components rather than prose, whether a product carousel, an itinerary, a seat map, or a chart. That means the agent has to emit a schema rather than text.
Teams sometimes start by prompting the model to emit custom tags and parsing them on the client-side. This stops working as the surface grows, because:
- The model isn’t as well trained on your markup as it is on tool calls so reliability drops as nested components get added. Well-formed data is not guaranteed just through prompting.
- The tag definitions live in the system prompt, so every new component bloats context and every edit risks regressions elsewhere in the prompt.
- Past conversations end up stored in a format only your parser can read, so loading history means either parsing raw messages on the client or keeping a second copy in a format that isn't native to the model API.
The pattern that has held up is to make each UI component a tool. The model calls present_products, present_itinerary, or present_plan_comparison with typed arguments; your server validates and enriches the call and emits an event; and your client renders it.
As the components are tool calls, they're already in the messages array in native format, so you don’t need to re-parse when you reload an old conversation. An example presentation-tool contract is illustrated below and in the reference repo.
前端交互组件本身就是工具(The UI components are tools)
电商场景的绝大多数回答都是交互式 UI 组件而非冗长散文:无论是商品轮播卡片、行程卡片、选座地图还是数据大盘图表。这意味着工具不仅负责后端数据查询,更直接主导了前端界面的生成。
许多团队在起步时往往尝试让模型输出自定义 XML 标签并在客户端进行解析。但随着业务界面复杂度提升,这种模式很快会失效:
- 大模型在自定义标记语言上的预训练充分度远不及原生工具调用,因此一旦引入嵌套复杂组件,输出可靠性便会急剧崩塌。结构良好的 XML 标签非常容易被模型写坏。
- 自定义标签规范必须写在系统提示词中,每新增一个 UI 组件都会使全局上下文不断膨胀。
- 历史对话最终以只有你当前私有解析器才能看懂的专有格式存储在数据库中,未来若想将历史无缝迁移至新模型或新底座将面临沉重的二次改造成本。
经受住考验的最佳设计模式是将每一个 UI 交互组件本身定义为一项工具。模型直接调用 present_products、present_itinerary 或 present_chart,并传入强类型的格式化入参。
由于组件本身就是工具调用,它们以原生格式天然保存在消息历史数组中,因此在重新加载旧对话时完全无需进行任何二次脆弱解析。


- Presentation Tools
- 前端交互组件工具化:将商品卡片、尺码选择器、轮播图等 UI 组件定义为可调用的工具参数,由模型精准组装参数并在客户端渲染。
- Screen State Record
- 屏幕状态感知:工具调用的输入参数即成为对话上下文中“用户当前所见画面”的精确历史记录,模型无需盲猜界面状态。
The tradeoff is streaming granularity. Each top-level argument of a tool call buffers on the server for validation, so the sub-components of a presentation tool arrive in steps even with streaming on. This impacts perceived latency.
To get a token-level stream, set eager_input_streaming: true on the tool definition, which skips the buffering and with it the server-side schema guarantee.
In our evals, schema violations are very rare on Claude Sonnet-class models and up, but wrap the call in a retry for the cases where one slips through.
Presentation tools also give the agent a record of what's on screen. When a customer says "the first hotel" or "the third one down on the left," the layout is in the messages array, in the arguments of the last presentation call.
For that to work, the arguments have to reflect the rendered layout, so structure them the way the UI is structured, as ordered rows and carousels rather than a flat list the client rearranges.
这种设计的工程权衡在于流式传输粒度。工具调用的每个顶层参数默认会在服务端进行完整缓冲以供校验,因此展现型工具的内部子组件不会逐字实时流出,而是以组件为单位整块呈现。
若想获得 Token 级的极致流式体验,可以在工具定义中开启 eager_input_streaming: true 参数,跳过参数缓冲校验,直接以流式数据包形式推送到客户端边收边解。
在我们的基准评测中,Claude Sonnet 及以上级别的模型极少发生 Schema 违规;但在生产环境中仍建议为工具调用配置一层兜底重试机制,以捕获极低概率的漏网格式异常。
展现型工具还为智能体天然保留了屏幕感知记录。当顾客指示“首选第一家酒店”或“左侧第三个推荐”时,对应的界面布局已完整留存在消息数组中最近一次展示工具调用的参数里,模型无需盲猜当前屏幕画面。
要实现这一点,传入工具的参数必须如实映射渲染出的界面布局结构:将其组织为按序号排列的行与卡片流,而不是交给客户端在模型不知情的情况下私自重排。
Making it fast and affordable
极速与低成本优化
Latency matters in commerce, and consumer surfaces are the least forgiving. However, on agentic surfaces, what we have consistently seen move metrics like retention, engagement, and cart size is the quality of the outcome.
Whether the answer was relevant and the task actually completed was more critical to those metrics as compared to marginal latency gains.
So attack latency on two fronts. Minimize end-to-end latency through good engineering, and pair that with dropping perceived latency (since time spent watching an agent work reads as progress).
Every user has a latency budget, and the techniques below keep the agent inside it without spending intelligence to get there.
在电商业务中,延迟是决定生死的命脉,而 C 端消费者界面的容忍度最为严苛。然而在智能体交互界面上,我们一致观察到真正驱动留存率、参与度与客单容量的核心杠杆,依然是任务最终交付结果的质量。
回答是否高度相关、核心任务能否真正闭环完成,相比边缘微小的延迟增益对业务关键指标的影响要深远得多。
因此必须从两条战线发起攻坚:一方面通过扎实的工程优化最小化端到端物理耗时;另一方面深度压缩感知延迟(因为用户目睹智能体有节奏地推进工作本身就会被理解为实质性进展)。
每位用户都有一个心理延迟预算,以下技术能够确保智能体始终运行在此预算红线之内,且无需牺牲模型推理智能。
Minimizing task completion latency
Task completion latency is the sum, over model turns, of time to last token plus tool processing. That gives you three levers to work towards: fewer turns, faster tools, and faster tokens. These levers sometimes compete, so the thing to minimize is the sum rather than any one of them.
最小化任务完成物理耗时(Minimizing task completion latency)
任务完成物理耗时是所有模型交互轮次中“首字至尾字生成耗时 + 工具处理耗时”的累加和。这为你提供了三大优化抓手:减少交互轮次、加速工具执行、以及提升 Token 生成吞吐。这些杠杆有时此消彼长,因此最终需要优化的核心是三者的物理总和。
Fewer turns
Query complexity adds turns, and is generally out of your control. Model intelligence and relevant context help the agent get to task completion in fewer turns. Some of our key learnings in this area include:
- Load likely context up front. If the user opened the assistant from a product page, or a merchant opened it from a campaign dashboard, put that page's data in the session context. The conversation is likely about it, and answering from context costs no extra turns.
- Increase model intelligence. Smarter models can decrease overall turns in the completion of a task as the agent can more efficiently plan and issue its tool calls. That often outweighs their slower tokens. If your queries skew complex, or production shows more than about five turns per task, the faster model is frequently the smarter one. Which one that is depends on your traffic, so choose by sweep, as described under "Choosing the model" below.
- Have the model call independent tools in parallel. Commerce use cases often require many operations in parallel: be it searching for multiple products, querying many policy docs, or fetching records from many sources of sales data. Parallel tool ensures multiple independent queries don’t burn additional turns. Prompt the model to call many tools within a turn and return the results in one user message as an array of tool results (see the parallel tool use docs).
减少交互轮次(Fewer turns)
用户查询本身的复杂性会增加轮次,这通常不可控。更聪慧的基础模型与精准切题的上下文有助于智能体以更少的轮次完成任务。我们在这一领域的关键实践包括:
- 前置装载大概率上下文:若用户是从具体商品页唤起助手,或商家从营销控制台唤起,直接将该页面的核心数据注入会话上下文。对话极大概率围绕此展开,直接从上下文解答无需消耗额外的工具查询轮次。
- 提升基础模型智能水平:高阶模型通常能显著减少完成任务所需的总交互轮次,因为智能体能以更高远的大局观规划工具调用链条。这往往能轻松对冲掉高阶模型本身稍慢的字生成耗时。如果你的业务查询偏复杂,或生产数据显示平均每个任务需要 5 轮以上交互,那么更聪慧的模型往往反而能以更少轮次成为跑得最快的那一个。
- 让模型在单轮内并发调用互不依赖的工具:电商用例往往需要并发处理大量操作:跨多个独立数据源检索商品、查验运费与调取商家活动。单轮工具并发确保了多个独立查询不会白白烧掉多余轮次。
Faster tools
- Optimize the tool's own backend. Sometimes a tool genuinely fans out – a merchant agent with a "get today's snapshot" query reads sales, inventory, and campaign status in three independent calls. But we often see the tool boundary become the place where missing backend logic gets stitched together: an availability check that calls the catalog for the SKU, the inventory service per store, and the fulfillment service for cutoffs, then applies substitution rules and pickup eligibility in the tool's own code before answering. That tool is now overloaded with domain knowledge, hard to keep correct as the rules change, and is carrying logic that should sit in an upstream system. When you find yourself writing that logic in a tool, the fix is one backend endpoint that answers the question, and calling that with an agent tool.
- Dispatch tools eagerly. Tool arguments stream out of the model like any other tokens, so the harness can execute each tool’s call as its arguments complete and process it while the model is still streaming other, parallel tools or content blocks. We've seen this take multi-second gaps down to a few hundred milliseconds, and the Claude Agent SDK does it by default. You should prompt the model to emit its slowest call first for maximum latency gains.
加速工具执行响应(Faster tools)
- 深度优化工具自身的后端性能:当工具确实需要并发请求多个下游接口时(例如商家查询“今日经营快照”需同时读取销售、库存与营销状态),在工具内部实现异步并发聚合。但切忌将本应属于上游系统的复杂拼接逻辑硬塞进工具代码中。
- 基于流式输出激进预派发(Eager Dispatch):工具入参像普通文本 Token 一样从模型中流式吐出。底座可以在模型仍在流式吐出其他参数或并行工具调用的同时,一旦核心必需参数就绪便立即在后台提前触发网络请求。这能将多秒的空等直接压至数百毫秒。


- Eager Dispatch
- 流式推演与并发派发:在模型流式生成工具入参时,底座一旦解析出核心必需参数(如
product_id),立即提前向后端发起网络请求,无需等待模型输出完整 JSON 闭合标签。 - Latency Savings
- 端到端延迟节省:将模型生成耗时与后端接口查询耗时深度重叠,减少数百毫秒的空等时间。
Perceived latency
Perceived latency is the time a user feels until the screen does something. It’s especially critical in consumer-facing use cases where any transaction friction impacts checkout rates and revenue. Two techniques shorten it without touching the model:
- Stream components as they form. A rendered commerce response is typically 500–700 output tokens, which without streaming is five or more seconds of a spinner. Send each parameter of a presentation tool to the client as it streams and render the page progressively.
- Show the work. While the agent is gathering context, render a short progress line for each step in plain language (for example, "finding hotels near the water"). You can build it from the tool's existing arguments (such as the query for a product search), or add an additional user_facing_message parameter tool that prompts the model to write the line.
感知延迟优化(Perceived latency)
感知延迟是指用户从发出指令到屏幕产生实质性视觉反馈之间的心理等待时长。这在 C 端交易场景中极为致命,任何交互摩擦都会直接损耗结账率与营收。无需改动模型即可落地的两大提速技巧:
- 组件成型即刻流式上屏:电商回答通常包含 500–700 个输出 Token,如果不做流式处理,用户将面临长达 5 秒以上的空白转圈等待。在展现型工具的参数流式吐出时,将每个字段逐步推送给前端实现渐进式局部渲染。
- 亮出工作进度(Show the work):当智能体正在调用耗时较长的后台检索时,用通俗直白的自然语言实时渲染简短的进度提示条(例如“正在为您查找靠近水域的优选酒店…”)。


- Streaming UI Components
- 组件级流式渲染:首个组件数据就绪即刻在前端上屏渲染,无需等待全部商品批次加载完毕,大幅缩短首屏可见耗时(Time to First Visible Component)。
- Optimistic State
- 乐观状态展示:在模型生成决策前先行展示占位框架或加载骨架,降低用户等待焦虑。
Prompt caching
Prompt caching is your largest cost reduction candidate and commerce traffic is well-suited for it. Cached input token reads cost a tenth of fresh ones, and while cache-writes carry a premium of roughly 1.25x, a cached prefix pays for itself on its second use. In customer facing applications where volume is large, you have a unique opportunity to hit very high cache levels using the cheapest, default 5 minute cache expiration.
The best commerce deployments we've seen run at 90–99% cache hit rates, and that is the range to design for from the start. Our experience has shown cached token reads are also around 1.5 to 2x faster at ~100k tokens, with relatively linear scaling the more tokens there are.
Caching is prefix-based. A request reads from cache up to the first byte that differs from a previous request, so what matters is not just what is in the context but the order it is in. Think of a request as three segments, ordered by how often they change:
- Global: most of the system prompt and tool definitions, identical across every session. This is your warmest cache and, at scale, will likely not expire. Keep it byte-identical across turns and sessions and put a cache breakpoint at its end.
- Session: per-user context and conversation history, which differ across sessions but stay stable within one. This segment comes after the global one.
- Volatile: anything that changes within a session, such as the current time or the current page. Put it at the very end of the request, either as a tagged block in the newest user turn or, on models that support mid-conversation system messages, as a system-role message appended to the messages array. The most common mistake we see is a timestamp or the current page at the top of the system prompt, which silently breaks the cache on every request.
提示词前缀缓存(Prompt caching)
提示词缓存是智能体降低运行成本的最核心武器,而电商业务流量天然具备极高的缓存亲和性。缓存命中的输入 Token 读取成本仅为新输入 Token 的 1/10;在请求量庞大的 C 端应用中,即使使用默认最便宜的 5 分钟缓存过期策略,也能轻松达到极高的命中水平。
在我们见过的顶尖电商落地实践中,输入 Token 的缓存命中率可高达 90%–99%,这也是从第一天起就应当锁定的设计基准。实测表明,在约 100k Token 上下文规模下,缓存读取的速度还能带来约 1.5 到 2 倍的物理提速。
提示词缓存是严格基于前缀匹配的。一次请求会从头读取缓存直至遇到第一个发生变化的字节;因此上下文里放了什么固然重要,排列的先后顺序更为关键。应当将请求结构划分为三个分层:
- 全局前缀 (Global):系统提示词的大部分核心指令与全部工具定义,跨所有会话 100% 保持一致。这是最热的缓存层,在规模化流量下几乎永不过期。在末尾设置一个缓存断点。
- 会话前缀 (Session):单个用户的专属上下文与历史对话,跨会话不同但单会话内相对稳定。此层紧随全局前缀之后。
- 易变层 (Volatile):单轮内即时变化的动态信息(如精确时间戳或当前页面)。务必将其放置在请求的最末尾,切忌将动态时间戳置于系统提示词顶端,否则每一次请求都会悄无声息地彻底击穿缓存。


- Global Prefix (全局前缀)
- 系统提示词与全部工具定义,跨所有用户和会话 100% 相同,保持最长缓存命中。
- Session Prefix (会话前缀)
- 用户跨会话记忆、用户画像与当前激活的 Skills,在单个用户会话内保持稳定。
- Turn Prefix (轮次前缀)
- 实时对话历史,每轮滚动前移断点,确保缓存命中率稳定在 90%–99%。
There are two implementation details to remember here. First, skills should be loaded as tool results rather than appended to the system prompt. The skill body then lands in the conversation prefix and is cached along with it.
Second, roll your breakpoints forward in each turn: a request allows a limited number of breakpoints, so move the newest one to the end of each user turn. Each round then reads the accumulated history, including long tool results such as search responses, from cache.
在工程落地中有两个细节需要牢记:首先,动态加载的 Skills 应当作为工具调用返回结果注入,而绝不能直接拼接到系统提示词末尾,否则会破坏全局提示词缓存的一致性。
其次,在每轮交互中动态前移缓存断点(Rolling Breakpoints Forward):一次请求支持设置有限的缓存断点,因此在全局指令处常驻断点的同时,将最新断点前移至每轮用户最新输入末尾。


- Cache Breakpoint
- 缓存断点机制:请求最多支持设置 4 个缓存断点。在多轮对话中动态更新最新轮次的断点标记,使历史对话完整复用缓存。
- Cost Reduction
- 极致降本:缓存命中部分的输入 Token 享受大幅折扣,大幅降低高并发电商对话的运营成本。
Choosing the model and its configuration
Model size and the effort setting are the same tradeoff – intelligence against latency and cost – and you should choose both by measurement:
- Pick your metric and your floor. Pick the quality metrics your business runs on (task completion, answer relevance, grounded accuracy), the eval score you won't go below, and your p50 and p99 latency and cost budgets.
- Sweep. Run your entire eval suite across every model and effort level you'd consider. We recommend starting at Opus for merchant agents, whose tasks are analysis-heavy, and Sonnet for consumer agents, where latency weighs more. If you have production traffic, weigh the results by your real query mix. Then let the numbers decide. Sometimes Opus 5's lift on cart-driving tasks justifies the cost difference over Sonnet, and sometimes it doesn't.
- Read the results carefully. Two things regularly surprise teams. The first is that a prompt is tuned to a model, so a sweep run with one prompt may underperform other models that it wasn't written for. A smaller model usually needs instructions the current model infers on its own, and a larger one will follow instructions to the letter that the smaller one was ignoring. A few rounds of iteration on each candidate's failing cases is a cheap step before ruling any of them out. The second is that a more intelligent configuration sometimes wins on latency (most commonly on p90 and p99) despite slower tokens, because it plans its tool calls better and needs fewer rounds on the most complex requests.
Measure cost per completed task rather than per model call, since a cheaper model that needs more turns, or fails more often, is not cheaper. When the result is close, and the cost fits your per-task economics and latency, choose intelligence. Quality is what drives adoption and retention, and allows for room to build for the next 6 months as models become better.
模型选型与配置权衡(Choosing the model and its configuration)
模型规格与思考预算(Effort)设置本质上是同一组权衡:在推理智能深度、延迟与单次调用成本之间寻求最优平衡点:
- 确立核心业务指标与底线红线:明确业务最看重的质量指标(任务最终闭环率、回答相关性、商品推荐精准度、结账转化率),并设定绝不妥协的交付底线。
- 针对评测套件全量基准测试各候选配置:在完整的 Eval 用例集上横向比对不同模型档位与推理参数组合。
- 上线能够稳定突破质量底线且具备最低延迟与成本的最优配置。
务必衡量单任务闭环综合成本(Cost per completed task)而非单次 API 调用成本:一个单价便宜但需要更多交互轮次甚至频繁出错的小模型,其实际综合成本反而高得多。当多项配置结果接近时,优先选择更高智能等级的模型:卓越质量才是驱动用户长期采纳与复购的核心引擎。
Running it in production
生产环境落地与运营
Lastly, we talk about what gets an agent through production: memory, safety, evals, and scaling the work across an organization.
最后,我们将深入探讨支撑智能体平稳跨越生产门槛的核心支柱:长期记忆、安全防护底座、评测体系以及跨大型工程团队的高效协同机制。
Memory that survives the session
The relationship and interactions you have with your customers matter. Memory is what lets an agent pick up where the last conversation left off instead of starting from nothing. A shopper who mentioned a nut allergy in March shouldn't have to repeat it in June, and a merchant who checks the same three campaigns every Monday shouldn't have to name them each time. Long-term memory, the facts that should survive across sessions, is a system you build and it has three parts: how facts are stored, how they are written, and how they are read.
跨会话持久记忆(Memory that survives the session)
与客户建立长久信任关系至关重要。记忆机制让智能体能够在用户重返店铺时自然承接上一次对话,而非从零开始。3 月份提及坚果过敏的买家不应在 6 月被迫再次声明;每周一查验相同三组营销活动的商家也不必每次重新输入活动名称。跨会话留存的长期记忆是一个完整的工程系统,由三大模块组成:事实如何存储、如何写入以及如何读取。
Storing memories
Memory belongs in your systems, not in the model.
A flat markdown profile works when profiles are small and the agent is the only reader. Most production commerce agents outgrow it, and the practical replacement is the database you already operate. A fact is a small typed record: a key (such as shoe_size, default_store, preferred_report_cadence), a short value, a category, and the session it came from. Some keys you decide up front and every user gets; the rest the extractor discovers. A database stays queryable as the store grows, lets you build deterministic behavior on specific attributes, and joins to the user data you already have.
For merchant-facing agents, key memory by person rather than by account. Merchant logins are often shared between operators, so each operator needs their own profile, and reads have to respect that operator's permissions: a store manager's agent should not recall a fact a district manager stated.
In the commerce domain, agent memory holds personal data. The facts worth remembering are often the most regulated ones, and the rules between jurisdictions differ. Treat memory as a data-handling design problem and not just a storage one. In practice that means four things:
- Decide which types of memories you are willing to hold. Enforce that at the write path, with a validator that every save goes through, rather than in the prompt alone.
- Give users a way to see, correct, and delete what is stored. Wire deletion into your account-deletion and data-request flows.
- Set a retention period. A preference from a few years ago is likely to be outdated, so a retention period helps keep memory facts fresh.
- Memory should be a per-deployment switch. This allows regions that can't take on these obligations to run without it.
记忆的存储与治理(Storing memories)
记忆属于你的业务数据库,而非大模型自身。
当画像数据较小且智能体为唯一读取方时,扁平的 Markdown 画像文本简单有效;而绝大多数生产级电商智能体很快会超出这一承载上限,成熟替代方案是使用现有的关系型或文档数据库。将每条事实沉淀为类型化的小记录:键名(如 shoe_size、default_store)、简明键值、分类标签以及来源会话 ID。随着数据量膨胀,数据库依然支持高效检索,并能无缝连接既有企业用户表。
对于面向商家后台的 B 端智能体,请按真人(Person)而非商户大账号(Account)维度沉淀记忆。商户账号通常由多位运营者共享,每个员工必须拥有独立的个人偏好画像,且数据读取必须严格遵守该员工的权限隔离边界。
在电商领域,智能体记忆承载着敏感的个人数据,且不同司法管辖区的监管法规截然不同。请将记忆视为一项严格的数据合规架构课题:
- 明确划定允许留存的记忆分类白名单:在写入路径中通过校验器强制执行白名单过滤,而非仅仅依赖提示词约束。
- 确保记忆对用户透明可见、支持修正与随时删除:在设置界面向用户提供记忆看板,并将删除动作与账户注销流程深度集成。
- 设定记忆数据生命周期与过期机制(Retention Period):数年前的偏好极大概率已过时,设置有效期有助于保持记忆常新。
- 配置区域级记忆开关:允许无法履行特定隐私法定义务的区域一键禁用记忆功能。
Writing memory
Write memory asynchronously. At the end of each turn, or every few turns in a long session, an agent in a separate thread or process reads the conversation and creates, updates, or deletes facts in the store, keeping its own working context as the session goes on.
It adds nothing to the conversation's latency, and achieved 13% higher fact recall on our internal commerce memory eval suite.
The obvious alternative, a tool the agent calls to save a fact, is the wrong one for a latency-sensitive commerce agent. Every save is a tool call inside a user-facing turn, and unless the whole store is in context, a save needs a read first to update or dedupe, which is a round of its own.
It also puts one more decision in front of the agent on every turn, and in our evals that competition for attention showed up as missed memories.
Separating the extractor also lets you prompt it precisely. It reads only the user's and the assistant's text, never tool results, so a product description or a review can't become a fact about the user. Its prompt says what counts as a fact — a stated size, a dietary constraint, a fulfillment preference, a merchant’s usual materialized views — and what doesn't, such as anything from a listing or a one-off detail.
记忆的异步写入(Writing memory)
必须采用异步方式沉淀记忆。在每轮对话结束或每隔数轮时,由独立线程或后台进程自动审阅最新对话流,在记忆库中执行新增、更新或擦除操作,全程维护自己的工作上下文。
这种异步模式对主对话延迟带来零负担,并且在我们的电商记忆基准测试中,其事实召回准确度高出整整 13%。
许多团队最初尝试让主智能体通过 Tool 主动保存记忆,但这在延迟敏感的电商场景中完全是错误路线:每一次保存都是一次额外的用户端等待耗时,且去重更新往往需要先读后写,白白产生额外往返。
这还在每轮交互中给主模型凭空增加了额外的注意力竞争,导致模型在专注解答主问题时频繁遗漏记忆操作。
解耦独立的提取器还能让你进行精准纯粹的提示词调优:它只阅读用户与助手的对话文本(绝不阅读第三方工具结果,避免商品介绍被误当成用户偏好事实),严格依据事实规范完成提取归类。


- Extractor Agent
- 独立提取智能体:在后台异步执行,不占用主对话回路,从最近对话中提炼用户核心事实与偏好。
- Structured Memory Store
- 结构化记忆存储:按类别严格过滤敏感信息(如健康、政治倾向),仅保存合法可用的偏好数据(如尺码、常购品牌)。
Reading memory
Read memory in three layers.
记忆的三层读取机制(Reading memory)
采用三层递进架构读取用户记忆:
Since memory is per-user context, all of it goes in the session segment, below the global cache breakpoint.
由于记忆属于单用户维度的专属上下文,所有注入的记忆数据统一放置在会话前缀层(Session Segment),位于全局缓存指令下方。
Safety: enforcement lives in the harness
The prompt is where safe behavior starts, but in commerce it can't be where safety is enforced. The failures are financial and often irreversible, and a prompt rule is one injection or one bad sample away from being skipped. Every rule below is enforced in code, on both the consumer and the merchant agent, and defined once so every runtime shares it.
安全防线:在运行时底座中强制执行(Safety: enforcement lives in the harness)
提示词是引导合规行为的起点,但在电商业务中绝不能作为强制执行安全的终点。电商领域的故障直接涉及真金白银且往往不可逆转;任何写在 Prompt 里的规则,一次偶发的提示词注入或不良采样就可能导致其被彻底绕过。以下所有安全红线均在底座代码中强制执行,一次定义,全运行时共享:
The model stages; a person or a policy applies
No model tool call moves money or changes the business. Order placement, payments, refunds, price changes, and campaign launches all end in an action the harness controls instead of the model.
On the consumer side this is structural: the checkout tool renders the cart with a button to place the order, and the backend interface the agent calls has no charge method at all.
On the merchant side, every write tool produces a staged change with a server-generated ID, and apply_change succeeds only for IDs that have been approved through a real surface: a button in the operator's portal, a confirmation in the CLI, or the platform's own tool-approval prompt when the agent runs on Managed Agents.
The guardrails are re-checked at apply time against current limits, not the limits in force when the change was staged. Whatever the surface, the shape is the same: the model's most dangerous action is to propose, and the approval routes through the maker-checker flow your business already uses for that kind of change.
模型负责暂存推演,人或策略负责核准生效(The model stages; a person or a policy applies)
绝对严禁大模型通过单个工具调用直接划转资金或修改核心业务状态。下单结算、付款、退款、价格调整与营销上线,最终均交由底座而非模型控制的动作执行闭环。
在 C 端,这种结构是天然的:结账工具只负责组装渲染购物车核对页与“提交订单”按钮,智能体调用的后端接口根本不存在直接扣款扣款的接口方法。
在 B 端,所有写入工具均生成带有服务端随机唯一 ID 的暂存提案,且 apply_change 接口仅在提案经过权威界面由真人或策略审批核准后才允许调用成功。
在最终生效阶段,底座必须针对系统当下的实时限额与库存重新校验,而非依赖暂存推演时的旧状态。无论何种界面形态,核心原则始终如一:模型最具风险的操作仅限于“提议”,由企业成熟的双人复核流进行裁决。
Writes and renders accept only server-issued IDs
The harness keeps a per-session record of every ID the server has handed the model, and that record is the only key any write or render will accept.
The cart accepts only product IDs the server returned to this session, and the merchant tools accept only listing and campaign IDs the agent has actually read. An ID that arrived any other way — hallucinated, pasted by a user, planted in a review — is refused before the backend sees it.
The same rule covers the UI. Presentation tools take IDs, and the server fills in the product, order, or change records itself, so a card only renders records the server itself filled in.
It covers delegates too: the merchant analysis subagent reads data but never adds to the set of IDs the agent may write to.
For fees, disclosures, and other regulated content, the model chooses which product to disclose and the server supplies every word from approved copy. The same fee fields are on the merchant agent's protected list, so neither side of the counter can change or paraphrase them, and evals check the rendered strings byte for byte.
写入与渲染仅接收服务端签发的合法 ID
运行时底座在会话中严格维护一张“已由服务端下发给模型的合法 ID 白名单”,且该凭证是任何写入或渲染接口唯一认可的钥匙。
购物车接口只接收由服务端下发给当前会话的商品 ID,商家工具只接收智能体通过正规检索读到的活动与商品 ID。任何通过幻觉生成、用户在文本中粘贴或在评价中投毒注入的伪造 ID,在触达后端核心服务前就会被直接拒绝。
同样的铁律覆盖 UI 渲染:展现型工具仅接收对象 ID,商品标题、主图与实时价格由服务端权威数据填充,大模型绝无可能通过在参数中伪造字符串来擅自篡改展示价格。
这同样适用于委托调用:商家数据分析子智能体仅具备只读权限,永远无权向主智能体注入可写入的新 ID 集合。
对于法定披露条款、服务费明细等强监管文案:模型仅负责选择挂载哪个组件,文案内容 100% 由服务端权威合规模板直出。
Capped transactions must hold to repeated requests
Most commerce surfaces cap how many of an item one user can buy — for ticket allocations, promotional pricing, or fraud control — and an agent will retry, rephrase, and parallelize in ways a human clicking a button never did.
The cap is therefore enforced on the line as it would be after the write, so a second "add two more" can't stack past it, and cart writes for one session are serialized so parallel tool calls in a single turn can't combine to exceed it.
Merchant changes are checked the same way against caps on price movement, discount depth, restock size, and campaign budget, plus a list of protected fields no change may touch. The rule generalizes: enforce every limit on the resulting state rather than the request, and serialize writes per session.
交易限额在面对重复并发请求时必须坚不可摧
绝大多数电商场景都对单用户购买数量设定了严格上限(用于票务配额、促销限购或反黄牛欺诈),而智能体可能会面对用户连续发出多次“再加两件”的指令挑战。
因此,限额校验必须在写入数据库的单条事务内部原子执行:计算写入后的全量购物车与已成单数量总和,坚决杜绝多次连续请求穿透限额边界。
B 端商家的调价幅度、折扣深度、补货上限与营销活动预算同样在事务层执行硬门禁校验,并配置受保护核心品类白名单。
Third-party content is sanitized
In commerce most of the context is written by people who aren't you — sellers, reviewers, competitors — so every backend read is untrusted input and goes through one sanitizer.
Every tool result authored by a third party, such as listings, reviews, policies, seller messages, and stored memory, is sanitized and wrapped in a fence with a fixed label before the model sees it.
The sanitizer strips control and bidirectional characters, removes anything that imitates the fence markers, defuses text that imitates a conversation turn or a tool call, and caps the size, which is designed to stop a hostile listing from impersonating the system or filling the context.
The prompt carries the other half of the contract: fenced text is material to report on, never to act on.
第三方不可信内容必须严格清洗过滤
在电商系统中,绝大部分上下文均由外部第三方编写(第三方卖家、评价买家、竞争对手),因此每一次从后端读取不可信内容都伴随着潜在输入风险。
每一个由第三方生产的工具返回结果(如商品列表、评价、条款、买卖家聊天记录),在喂给大模型前必须经过底座的清洗管道过滤,并使用专属 XML 标签进行严格沙箱隔离。
清洗器会自动剔除控制字符与双向文本欺骗字符、剥离所有仿造隔离标签的代码、消除伪造系统对话的提示词注入内容。
系统提示词则约定另一半契约:被 XML 标签隔离的内容纯属待分析的事实材料,绝对不可作为系统指令执行。
Evals: shipping a non-deterministic system
Anything from a small prompt change to a new tool can change agent behavior in ways that are hard to predict, and the change you're shipping is often not the one that regresses. Evals are how you find that out before you deploy. Our earlier blog post on evals for agents covers the general practice. This section covers specifics for commerce agents.
Evaluate snapshots, not conversations
The model’s API is stateless, so what the agent outputs is a function of the system prompt, the tools, and the messages array. This means any state a commerce conversation can reach can be constructed directly. So creating an eval case means constructing the test state, appending the test user message, and letting the agent run from there.
Then grade the outcome: the final state and the rendered response, including the arguments of the last write. In most cases, we recommend against grading the path the agent took to get there as such test cases are brittle and restricting.
Simulated-user evals, in which a second model plays the user and a judge grades the whole conversation, are a poor tool for measurement. Two non-deterministic systems interacting need larger samples, cost more per trial, are harder to judge, and produce failures that are hard to attribute. They are useful for finding coverage gaps and for a general vibe check on the agent, so use them to discover cases, then write each case as a snapshot.
评测体系:发布非确定性系统(Evals: shipping a non-deterministic system)
从微调一行提示词到引入一个新工具,都可能以难以预测的方式改变智能体在全局场景中的行为。没有一套完备的自动化评测套件在用户前提前拦截倒退,你绝不能轻易发布任何变更。
评测单轮快照,而非冗长的多轮模拟会话(Evaluate snapshots, not conversations)
大模型 API 本身是无状态的:智能体在某一时刻的输出完全由系统提示词、可用工具以及当前传入的消息数组决定。这意味着任何多轮会话状态,都可以被精准压缩提炼为单次模型调用的快照。
随后对产出结果进行自动化打分:校验模型触发的最终状态、呈现格式以及工具调用的参数有效性。在绝大多数日常工程迭代中,无需在单轮评测中跑完耗时的完整工具链执行。
由另一模型扮演用户、裁判模型打分的多轮模拟评测固然有助于宏观体验查漏补缺,但两个非确定性系统相互碰撞需要极大的样本量、更高的开销且排查归因极其困难。单轮快照评测才是支撑高频日常工程迭代的主力核心。


- State Injection
- 状态注入评测:直接向智能体注入特定购物车状态、库存异常或歧义请求,测试单轮决策输出。
- Stateless Verification
- 无状态确定性核验:评估最终生成的参数、呈现格式与业务规则合规性,实现秒级高并发自动化回归测试。
Evaluate for behaviors in tough conditions
Most teams fail to properly test the injected state. A case should encode the preconditions of a failure, not just the task. If a behavior only emerges after a busy first turn with several tool calls, or after a contradiction earlier in the session, a case that starts from a clean state passes on every config and provides no meaningful data.
We've observed most suites to be heavy on such clean-state cases, so make sure a share of yours starts from long, messy, or contradictory histories.
针对严苛边缘场景与负向行为进行深度评测
许多团队往往忽视了对上下文注入状态的严苛测试。高质量的评测集必须在上下文中主动注入复杂边缘状况:如果某项异常行为仅在首轮调用了多次工具或会话前期存在自相矛盾时才会暴露,那么干净理想的用例对排查回归毫无帮助。
我们观察到绝大多数团队的用例库都充斥着大量过于干净的理想用例,请务必确保你的评测库中拥有充足的复杂、冗长且充满矛盾历史的极端用例。
Cover the different types of commerce agent evals
Effective evaluation requires testing both desired and undesired behaviors.
For every positive case, write its negative counterpart: a "should serve" for every "should refuse," a "should just do it" for every "should ask." Missing negatives are the most common gap we find in a suite.
Evaluate for the following:
- Core requests that make up the bulk of your traffic, since a failure here affects most sessions. These include simple lookups, multi-constraint requests, product and plan questions, and multi-intent messages. For the questions, check that every price, availability, and attribute traces back to returned data, and that the agent says when data is missing rather than inventing it.
- Context-dependent requests, such as references to what is on screen, constraints carried over from earlier turns, and writes against an existing cart. Evaluating memory falls into this bucket as well. Check that memories were extracted, retrieved, and changed the answer.
- Safety and brand cases, where a failure costs money or trust. These include attempted injection, attempts to read another user's data, and regulated language, which is checked byte for byte. Split injection into two cases: user-authored injection, where the directive comes from the user's own message, and data-plane injection, where it is planted in product names, reviews, or web snippets that arrive via tool results.
- Interface evaluations, to ensure the right component is rendered, item caps are respected, and there are no internal identifiers in user-facing text. Test for timeouts and empty results too.
- Requests that belong to multiple capabilities at once. An operator asks "if I mark this down 15%, do I have enough stock to cover the demand?" That is a pricing question and an inventory question together. The right answer stages the markdown with a stock projection attached; the wrong answers do one and skip the other. Evals written per capability won't catch this, because each grades only its own half. Write cases for the requests that need two neighboring capabilities together, and grade both halves of the answer.
覆盖电商智能体评测的各大核心维度
高效评测必须同时覆盖正向期望行为与负向拦截行为。
对于每一个正向推荐用例,都应编写对应的负向用例:有“应当满足”就必须有“应当拒绝”;有“应当直接执行”就必须有“应当主动向用户澄清”。缺乏负向测试是用例库最常见的漏洞。
评测必须覆盖以下维度:
- 核心主流请求:占线上流量绝大多数的基础用例,包括简单查询、多约束选购、产品对比与复合意图消息。严格校验所有价格、库存与属性均能溯源至真实数据,禁止凭空捏造。
- 上下文强相关用例:测试对屏幕所见元素的指代理解、前期轮次偏好的继承、以及基于既有购物车的更新。记忆机制的提取与召回效果也归属于此类。
- 安全与品牌底线用例:提示词注入抵御、越权数据嗅探拦截、以及严格合规法定义务用例。
- 交互界面合规性评测:确保展现型组件正确挂载、单人限额严格合规、且在用户端文本中绝无内部技术标识泄露。
- 跨多业务域复合请求:例如商家询问“如果我降价 15%,当前库存能否承接激增需求?”这同时涵盖了定价与库存双重业务域,评测必须同时核验两项复合决策。
Write evals with SMEs and use real incidents
Partner with the subject-matter experts who see the failures firsthand, such as team members in Product, Legal, Merchant Ops, Customer Care, and Category Management, to design test cases. Real failures make the best evals, and 50-100 eval cases per user flow is a good starting point.
Make sure to have a variety of cases, as outlined above. Production transcripts are a great stream for sourcing new cases, especially the tricky ones. Coding agents are good at generating additional cases and adversarial variants. The reference repository includes a Claude Code plugin with an eval-authoring skill built with our recommended approach.
与业务专家紧密协同,从真实故障中提炼评测
务必与一线直面故障的领域专家紧密合作:产品经理、法务合规、商家运营、一线客服与品类主管。生产真实客诉是打造评测用例最无价的源泉,每个核心业务流沉淀 50–100 个快照用例是一个极佳的基准点。
确保评测用例丰富多样。线上全量日志脱敏是挖掘前沿极端案例的绝佳渠道;借助编程智能体也能高产出大量对抗性测试变体。
Shipping with a large organization
In a commerce enterprise the agent is built by many engineering teams. Search, checkout, pricing, marketing tech, customer care, and the catalog platform each own systems the agent depends on, each ships on its own cadence, and each will want to add or change a tool, a skill, or a prompt rule.
Unlike a service, an agent has no strict module boundary protecting the others: a change made by the pricing team shares a context window with checkout.
The tempting fix is to break the system into many subagents, one per business unit. As discussed in Part 1, we recommend against it for quality reasons. Instead, we outline the process for de-risking multi-team collaboration:
- Ownership follows the systems. Every skill and tool has a single owner team. For example, pricing owns the promotion tools and the pricing skill, care owns the order and returns tools and the customer-care skill. The shared prompt has a single platform-level owner for the common parts and domain owner for the domain-specific section.
- A change ships with its cases and CI runs a set chosen for it. A team contributing a skill also contributes its cases, including the negative cases and the boundary cases against neighboring skills. Running the full suite on every pull request is too slow and too expensive to survive, so build a CI set from it instead. That set will consist of a core set of cases with the highest-traffic requests and every safety case. On top of that, run the cases for whatever the change touched. For a skill, that means its own cases and its neighbors' boundary cases. For a tool, it is every case that calls it. For the shared prompt, it is the full eval suite since everything reads the system prompt. We recommend gating the pass rate over a few trials, and on cache hit rate and cost per turn. It is also a good practice to run the full suite nightly and before every release. Cross-team regressions are caught in these runs.
- The agent should also be inside the release calendar. It's one deployment unit, so a bad change reaches every user at once. Roll prompt and skill changes to a canary cohort first, keep a switch that turns off one skill without a deploy, and freeze the agent ahead of peak periods the same way you freeze other systems.
For the human side of this arrangement, see Building effective human-agent teams.
大型企业组织内的多团队协同研发机制
在大型电商企业中,智能体通常由众多工程团队协同构建:搜索、结算、定价、营销技术、客服与商品目录中心分别维护着智能体依赖的底层系统,各团队交付节奏各异,并随时会提出新增工具或变更规则的诉求。
与传统微服务架构不同,大模型在同一个上下文中推理:某一个业务域编写不当的指令,极易污染并干扰模型在其他领域的判断决策。
很多团队试图通过“一个业务线拆一个子智能体”来强行隔离,但正如第一部分所述,这会引入严重的交接损耗与延迟惩罚。为此我们总结了化解多团队协同风险的标准作业范式:
- 代码所有权与底层系统严格对齐:每个 Skill 与 Tool 均有唯一的负责团队(例如定价团队掌管调价工具与 pricing skill,客服团队掌管退换货工具与 care skill)。共享系统提示词由平台架构团队负责公共基座,业务团队负责专属分区。
- 业务团队变更必须同步交付评测用例:新增 Skill 必须同步提交正向、负向与跨业务域边界用例;在 CI 管道中强制对改动相关业务域及全局核心用例运行回归断言。
- 智能体纳入统一发布节奏日历:由于模型上下文共享,一次不当改动可能波及全量用户。对提示词与 Skill 变更必须先推灰度集群(Canary Cohort),并在底座保留无需重新部署即可一键降级熔断单一 Skill 的动态开关。
关于团队组织架构与人机协同分工的深度实践,可进一步参阅官方指南:《构建高效的人机智能体协作团队》。
Looking ahead
Most of what this post describes is not about the model. The tools call systems you already run, the skills encode procedures you already follow, the evals are your product requirements doc written as tests, and the harness enforces policy you would enforce for any client. Models will keep improving, and when a better one ships, the architecture we describe adopts it as a config change with an eval sweep. Everything else keeps working.
It is also important to think about your roadmap for product surfaces. The architecture will outlast the chat panel. The same agent can work over voice, and it can proactively act on a fare drop before the user asks. For a team that already has the evals and the tools, those are presentation-layer projects. Further out, some of the traffic to your storefront will come from agents that shop on behalf of users. The same provenance, staging, and approval rules that keep your own agent in bounds are what will let you open your tools to those agents safely.
Commerce has always rewarded making the buying process as smooth as possible. Agents make that a lot easier. Check out the complete reference implementation, with both the consumer and the merchant agent and runnable examples for retail, travel, telecom, and entertainment.
前瞻与未来演进(Looking ahead)
本文探讨的核心内容,绝大部分并非关于基础模型本身。工具连接的是企业既有的成熟系统,Skills 固化的是团队已有的标准作业流程,评测集是用测试用例写就的产品需求文档,而底座则强制执行适用于任何生产客户端的安全合规策略。基础模型将持续飞速跃迁,而一旦更强大的新模型发布,本文阐述的这套架构只需极少改动就能无缝升级享受红利。真正的长期壁垒,在于工程脚手架的精细打磨。
前瞻性地规划产品交互形态同样至关重要。本文的架构生命周期将远超当前的对话侧边栏:相同的智能体可以通过语音流畅交互,或在机票降价时主动抢先在用户提问前完成锁单。对于已经沉淀完备工具与评测集的团队,这些都仅仅是前端呈现层的扩展项目。放眼未来,店铺的相当一部分流量将直接来自于代表用户自主采购的第三方智能体;正是本文所强调的来源追溯、暂存确认与审批流程,才能让你未来安全地将商业工具开放给全网外部智能体。
商业的演进历来奖赏那些将交易摩擦降至极致的创新。借助电商智能体,昔日繁琐割裂的搜索、筛选、对比、核验与结账流程,正彻底消融于一场自然直觉的对话之中。欢迎探索完整官方参考实现代码库,其中涵盖了面向消费者与商家的全套智能体,并附带针对零售、差旅、电信与文娱行业的开箱即可运行的实战范例。
Acknowledgements
Written by Matthew Koen and Ali Shazal. Special thanks to Michael Segner, Rodrigo Olivares, Amandeep Khurana, Aiza Usman, John Lopus and others for their contributions.
致谢与作者(Acknowledgements)
本文由 Anthropic 工程师 Matthew Koen 与 Ali Shazal 执笔撰写。特别鸣谢 Michael Segner、Rodrigo Olivares、Amandeep Khurana、Aiza Usman、John Lopus 及其他团队成员对本文作出的重要贡献与审阅支持。