ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

Claude Code 意外开源后,我用 OpenTelemetry 给企业级 Agent 补上行为分析

Claude Code 意外开源后,我用 OpenTelemetry 给企业级 Agent 补上行为分析 1. 从 Claude Code 源码泄露说起企业级 Agent 为什么需要行为分析Claude Code 的源码因为一个 npm source map 打包失误被完整暴露出来51.2 万行 TypeScript 摊在公网上。很多人第一反应是去看它的记忆系统、多 Agent 编排、权限控制但我翻完services/analytics/和utils/telemetry/这两个目录之后真正被触动的是另一件事一个生产级 Agent 的可观测性代码复杂度和它的核心功能代码是一个量级的。这件事对企业级 Agent 开发者来说是个提醒。你写了一个 TypeScript Agent接了模型、挂了工具、跑通了业务流程然后呢上线之后你会发现真正难回答的问题不是它能不能跑而是这一轮对话为什么花了 47 秒是模型慢、工具慢还是卡在等用户审批这个月 Token 账单涨了 3 倍是哪个 skill、哪个 prompt 版本、哪个 session 在烧钱多 Agent 协同的时候A 说交接完了B 说没收到到底断在哪一环某个 Agent 半夜连续调了 50 次高危工具为什么没有任何告警这些问题的共同点是它们都不是接口耗时能回答的。你需要的是对 Agent 行为本身的建模——一次 interaction 怎么展开、LLM request 和 tool call 怎么嵌套、blocked_on_user 占了多少时间、parent session 是谁。这就是 Agent Behavior AnalysisABA要解决的事。这篇教程面向的是正在用 TypeScript 写 Agent、准备把它推进生产环境的开发者。我会用 OpenTelemetry 从零搭一套 Agent 行为采集链路先讲清楚 Span 该怎么设计再给出可复制的 Collector 配置和埋点代码最后跑一次真实请求验证数据落库。全程不需要你改 Agent 的核心逻辑只需要在关键节点插桩。技术栈是 TypeScript OpenTelemetry SDK OTel Collector后端用任意兼容 OTLP 的服务即可。如果你手上有 TaoToken 的 API Key可以直接用它来跑通模型调用这一段省去自己搭模型网关的功夫。2. 前置准备TaoToken 接入与 OpenTelemetry 依赖安装在开始埋点之前先把两件事准备好一个能调通的模型入口和一套 OpenTelemetry 的依赖。2.1 为什么用 TaoToken 作为模型入口做 Agent 行为分析的时候模型调用这一段本身也是要被观测的对象。你需要一个稳定的、支持标准 OpenAI 兼容接口的入口这样 OpenTelemetry 的 LLM request span 才有东西可采。TaoToken 提供的就是这样一个入口Base URL 是https://taotoken.net/api兼容 OpenAI 的/v1/chat/completions协议TypeScript 里用openai这个 npm 包就能直接调。先去控制台创建一个 API Key控制台地址https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewrite创建完之后把 Key 存到环境变量里不要硬编码进代码export TAOTOKEN_API_KEYsk-xxxxxxxxxxxxxxxx export TAOTOKEN_BASE_URLhttps://taotoken.net/api如果你还没决定用哪个模型可以先去模型对话页面试一下不同模型在工具调用上的表现Agent 场景对 function calling 的稳定性要求比较高模型对话https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewrite2.2 安装 OpenTelemetry 依赖TypeScript 项目里需要装这几个包。注意版本opentelemetry/sdk-node在 0.5x 之后 API 有调整下面用的是当前稳定组合npm install opentelemetry/api \ opentelemetry/sdk-node \ opentelemetry/exporter-trace-otlp-http \ opentelemetry/exporter-metrics-otlp-http \ opentelemetry/resources \ opentelemetry/semantic-conventions \ opentelemetry/sdk-metrics \ opentelemetry/context-async-hooksopentelemetry/context-async-hooks这个包很关键。Claude Code 的sessionTracing.ts用AsyncLocalStorage保存 interactionContext 和 toolContext让后续的 llm_request、tool.execution 自动挂到正确的父节点下。OpenTelemetry 的AsyncLocalStorageContextManager做的就是同一件事——它让 context 在异步调用链里自动传递你不需要手动把 parent span 一层层往下传。2.3 目录结构约定为了让后面的配置能直接复制先约定一下项目结构agent-observability/ ├── src/ │ ├── telemetry/ │ │ ├── setup.ts # SDK 初始化 │ │ ├── attributes.ts # 统一关联键 │ │ └── spans.ts # 语义级 Span 封装 │ ├── agent/ │ │ └── runtime.ts # Agent 主逻辑 │ └── index.ts ├── otel-collector-config.yaml └── package.json这个结构对应 Claude Code 的三层分离思路setup.ts管标准化 telemetry 的接入attributes.ts管统一关联键spans.ts管会话级 tracing 的语义建模。三层各管各的事不要混在一起。3. 可复制配置Collector YAML 与 TypeScript 埋点片段这一节是全文的核心所有配置都可以直接复制到你的项目里改。3.1 OTel Collector 配置先写 Collector 的配置。它的作用是接收 Agent 发过来的 OTLP 数据做一次批处理和属性清洗再转发到后端。文件放在项目根目录命名otel-collector-config.yamlreceivers: otlp: protocols: http: endpoint: 0.0.0.0:4318 grpc: endpoint: 0.0.0.0:4317 processors: batch: timeout: 5s send_batch_size: 512 send_batch_max_size: 1024 # 高基数治理把 session.id 这类字段从指标里剔除只保留在 trace 上 attributes/strip_high_cardinality: actions: - key: session.id action: delete - key: agent.id action: delete - key: parent.session.id action: delete # 隐私前置把可能含敏感内容的字段在采集阶段就删掉 attributes/redact: actions: - key: tool.input action: delete - key: llm.prompt action: delete resource: attributes: - key: deployment.environment value: production action: upsert exporters: otlphttp: endpoint: ${OTEL_BACKEND_ENDPOINT} headers: Authorization: ${OTEL_BACKEND_TOKEN} debug: verbosity: detailed service: pipelines: traces: receivers: [otlp] processors: [attributes/redact, resource, batch] exporters: [otlphttp, debug] metrics: receivers: [otlp] processors: [attributes/strip_high_cardinality, resource, batch] exporters: [otlphttp]这里有两个设计点值得说明。第一attributes/strip_high_cardinality只挂在 metrics 管道上不挂 traces。原因和 Claude Code 用OTEL_METRICS_INCLUDE_*环境变量控制高基数字段是同一个道理session.id、agent.id这些字段在 trace 上是必须的你要靠它串因果链但一旦进了指标系统每个 session 都会生成一个独立的时间序列基数直接爆炸。所以 trace 保留、metric 剔除。第二attributes/redact放在 traces 管道的最前面。Claude Code 在sink.ts里用stripProtoFields()在 Datadog fanout 之前统一脱敏一个过滤点管住所有 sink。这里同理tool.input和llm.prompt在进 Collector 的第一时间就被删掉后面的 exporter 永远看不到。启动 Collectordocker run -d --name otel-collector \ -p 4317:4317 -p 4318:4318 \ -v $(pwd)/otel-collector-config.yaml:/etc/otelcol/config.yaml \ -e OTEL_BACKEND_ENDPOINThttps://your-backend.example.com \ -e OTEL_BACKEND_TOKENyour-token \ otel/opentelemetry-collector-contrib:latest \ --config/etc/otelcol/config.yaml3.2 统一关联键attributes.tsAgent 系统里最难的从来不是生成一条事件而是让不同事件能被关联起来。Claude Code 在metadata.ts里补了agentId、parentSessionId、agentType、teamName取值优先级是 AsyncLocalStorage → 环境变量 → bootstrap state。我们照这个思路写// src/telemetry/attributes.ts import { AsyncLocalStorage } from node:async_hooks; export interface AgentContext { sessionId: string; agentId: string; parentSessionId?: string; agentType: standalone | subagent | teammate; teamName?: string; } export const agentContextStorage new AsyncLocalStorageAgentContext(); export function getAgentAttributes(): Recordstring, string { const ctx agentContextStorage.getStore(); if (!ctx) { return { session.id: process.env.AGENT_SESSION_ID ?? unknown, agent.id: process.env.AGENT_ID ?? unknown, agent.type: standalone, }; } const attrs: Recordstring, string { session.id: ctx.sessionId, agent.id: ctx.agentId, agent.type: ctx.agentType, }; if (ctx.parentSessionId) attrs[parent.session.id] ctx.parentSessionId; if (ctx.teamName) attrs[team.name] ctx.teamName; return attrs; }agentType这个字段是必须的。同一个系统里Agent 可能是同进程 subagent也可能是 swarm teammate也可能是 standalone。可观测性层不能假定只有一种身份来源必须自己做统一归并。归并之后后端拿到的是一个最小可关联图当前 session 是谁、当前 agent 是谁、上游 parent session 是谁、属于哪个 team。3.3 语义级 Span 封装spans.ts这是整个埋点里最有价值的部分。Claude Code 把业务语义直接编码成了 span 类型claude_code.interaction、claude_code.llm_request、claude_code.tool、claude_code.tool.blocked_on_user、claude_code.tool.execution。我们照着建一套自己的// src/telemetry/spans.ts import { trace, SpanKind, SpanStatusCode, Span } from opentelemetry/api; import { getAgentAttributes } from ./attributes; const tracer trace.getTracer(agent-runtime, 1.0.0); // 根 span一次用户交互回合 export function startInteraction(userInput: string): Span { return tracer.startSpan(agent.interaction, { kind: SpanKind.SERVER, attributes: { ...getAgentAttributes(), interaction.input.length: userInput.length, }, }); } // 一次模型调用 export function startLLMRequest(model: string, promptTokens: number): Span { return tracer.startSpan(agent.llm_request, { kind: SpanKind.CLIENT, attributes: { ...getAgentAttributes(), llm.model: model, llm.prompt.tokens: promptTokens, }, }); } // 一次工具调用 export function startToolCall(toolName: string): Span { return tracer.startSpan(agent.tool, { kind: SpanKind.INTERNAL, attributes: { ...getAgentAttributes(), tool.name: toolName, }, }); } // 等待用户批准——这个 span 是 Agent 行为分析的关键 export function startBlockedOnUser(toolName: string): Span { return tracer.startSpan(agent.tool.blocked_on_user, { kind: SpanKind.INTERNAL, attributes: { ...getAgentAttributes(), tool.name: toolName, }, }); } // 工具实际执行 export function startToolExecution(toolName: string): Span { return tracer.startSpan(agent.tool.execution, { kind: SpanKind.INTERNAL, attributes: { ...getAgentAttributes(), tool.name: toolName, }, }); } export function endSpan(span: Span, error?: Error): void { if (error) { span.setStatus({ code: SpanStatusCode.ERROR, message: error.message }); span.recordException(error); } else { span.setStatus({ code: SpanStatusCode.OK }); } span.end(); }agent.tool.blocked_on_user这个 span 单独拎出来是整套设计里最容易被忽略但最有用的一个。很多 Agent 平台做 tracing 只追踪模型延迟和工具延迟但真实耗时不全是机器耗时——权限审批和人工中断本身就是主链路的一部分。把这类时间独立出来之后你才能回答慢是因为模型慢还是审批慢某类工具成功率低是执行问题还是审批被拒绝太多3.4 SDK 初始化setup.ts// src/telemetry/setup.ts import { NodeSDK } from opentelemetry/sdk-node; import { OTLPTraceExporter } from opentelemetry/exporter-trace-otlp-http; import { OTLPMetricExporter } from opentelemetry/exporter-metrics-otlp-http; import { PeriodicExportingMetricReader } from opentelemetry/sdk-metrics; import { AsyncLocalStorageContextManager } from opentelemetry/context-async-hooks; import { context } from opentelemetry/api; import { resourceFromAttributes } from opentelemetry/resources; const contextManager new AsyncLocalStorageContextManager(); context.setGlobalContextManager(contextManager); const sdk new NodeSDK({ resource: resourceFromAttributes({ service.name: agent-runtime, service.version: 1.0.0, }), traceExporter: new OTLPTraceExporter({ url: http://localhost:4318/v1/traces, }), metricReader: new PeriodicExportingMetricReader({ exporter: new OTLPMetricExporter({ url: http://localhost:4318/v1/metrics, }), exportIntervalMillis: 10000, }), }); sdk.start(); process.on(SIGTERM, () { sdk.shutdown().then(() process.exit(0)); });注意AsyncLocalStorageContextManager必须在 SDK 启动之前设置成全局 context manager否则后面startInteraction里创建的 span 不会自动成为子 span 的父节点。4. 验证请求跑一次真实 Agent 调用看数据落库配置写完了现在跑一次真实的 Agent 调用验证 trace 能不能正确串起来。4.1 Agent 主逻辑埋点// src/agent/runtime.ts import OpenAI from openai; import { agentContextStorage } from ../telemetry/attributes; import { startInteraction, startLLMRequest, startToolCall, startBlockedOnUser, startToolExecution, endSpan, } from ../telemetry/spans; const client new OpenAI({ apiKey: process.env.TAOTOKEN_API_KEY, baseURL: process.env.TAOTOKEN_BASE_URL, }); export async function runAgentTurn(userInput: string, sessionId: string) { const ctx { sessionId, agentId: agent-${Date.now()}, agentType: standalone as const, }; return agentContextStorage.run(ctx, async () { const interactionSpan startInteraction(userInput); try { // 1. 模型调用 const llmSpan startLLMRequest(gpt-4o-mini, userInput.length); const completion await client.chat.completions.create({ model: gpt-4o-mini, messages: [{ role: user, content: userInput }], }); llmSpan.setAttribute(llm.completion.tokens, completion.usage?.completion_tokens ?? 0); endSpan(llmSpan); const toolCall completion.choices[0]?.message?.tool_calls?.[0]; if (!toolCall) { endSpan(interactionSpan); return completion.choices[0]?.message?.content; } // 2. 工具调用 const toolSpan startToolCall(toolCall.function.name); // 3. 等待用户批准模拟 const blockedSpan startBlockedOnUser(toolCall.function.name); await new Promise((r) setTimeout(r, 800)); endSpan(blockedSpan); // 4. 工具执行 const execSpan startToolExecution(toolCall.function.name); await new Promise((r) setTimeout(r, 300)); endSpan(execSpan); endSpan(toolSpan); endSpan(interactionSpan); return tool executed; } catch (err) { endSpan(interactionSpan, err as Error); throw err; } }); }4.2 跑起来看结果npx tsx src/index.tsCollector 的debugexporter 会把收到的 span 打到日志里。你会看到类似这样的结构Span #0 Trace ID : 4f2a1b8c9d3e5f6a7b8c9d0e1f2a3b4c Parent ID : Name : agent.interaction Attributes : session.id: sess-20260331-001 agent.id: agent-1743400000000 agent.type: standalone Span #1 Parent ID : 4f2a1b8c9d3e5f6a7b8c9d0e1f2a3b4c Name : agent.llm_request Attributes : llm.model: gpt-4o-mini llm.prompt.tokens: 42 llm.completion.tokens: 87 Span #2 Parent ID : 4f2a1b8c9d3e5f6a7b8c9d0e1f2a3b4c Name : agent.tool Attributes : tool.name: query_order Span #3 Parent ID : (tool span id) Name : agent.tool.blocked_on_user Duration : 800ms Span #4 Parent ID : (tool span id) Name : agent.tool.execution Duration : 300ms关键看两点Parent ID 的嵌套关系对不对以及blocked_on_user 的耗时有没有被单独记录。如果 Parent ID 全是空的说明AsyncLocalStorageContextManager没生效检查setup.ts里context.setGlobalContextManager是不是在 SDK 启动前调用的。4.3 用指标看趋势Trace 看因果链Metric 看趋势。在setup.ts里加一个计数器统计每类工具的调用次数import { metrics } from opentelemetry/api; const meter metrics.getMeter(agent-runtime); const toolCallCounter meter.createCounter(agent.tool.calls, { description: Number of tool calls by name, }); // 在 startToolCall 里调用 toolCallCounter.add(1, { tool.name: toolName });注意这里只带了tool.name这一个低基数标签没有带session.id。因为 Collector 的attributes/strip_high_cardinality会把session.id从指标里删掉但更稳妥的做法是埋点的时候就不加——高基数字段不是默认全塞进指标系统的。5. 常见报错排查401、local proxy failed 与 OAuth 问题埋点跑起来之后最容易卡住的是下面这几类报错。我按实际遇到的频率排一下。5.1 401 UnauthorizedError: 401 Unauthorized {error:{message:Invalid API key provided,type:invalid_request_error}}这个报错来自模型调用那一段不是 OpenTelemetry 的问题。检查三件事第一TAOTOKEN_API_KEY环境变量有没有正确导出。在 Node 里process.env.TAOTOKEN_API_KEY如果是undefinedopenai包会直接抛 401。第二Base URL 有没有写错。TaoToken 的 API 地址是https://taotoken.net/api注意不要多加/v1openai包会自动拼/v1/chat/completions。如果你写成了https://taotoken.net/api/v1最终请求会变成/api/v1/v1/chat/completions直接 404 或 401。第三Key 有没有被复制时带上空格。从控制台复制的时候很容易多带一个换行符用echo $TAOTOKEN_API_KEY | xxd | tail -1看一下末尾有没有0a。如果 Key 本身没问题但还是 401去 API Keys 页面确认一下这个 Key 的状态和额度API Keyshttps://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewrite5.2 local proxy failed / connection refusedError: connect ECONNREFUSED 127.0.0.1:4318这是 Collector 没起来或者端口没对上。OTLP HTTP 默认端口是 4318gRPC 是 4317。检查docker ps | grep otel-collector curl -v http://localhost:4318/v1/traces如果容器在跑但端口不通大概率是otel-collector-config.yaml里receivers.otlp.protocols.http.endpoint写成了localhost:4318而不是0.0.0.0:4318。容器内的localhost指向容器自己宿主机访问不到。还有一种情况是setup.ts里 exporter 的 URL 写成了http://localhost:4318但 Agent 跑在另一个容器里。这时候要把localhost换成 Collector 的服务名或宿主机 IP。5.3 reading choices of undefinedTypeError: Cannot read properties of undefined (reading choices)这个报错通常出现在completion.choices[0]这一行。原因是模型返回体结构和预期不符。常见触发场景一是模型不支持 function calling但你传了tools参数返回体里没有choices字段而是error字段。先确认你用的模型支持工具调用。二是流式请求stream: true下返回的是 async iterator不是完整的 completion 对象。如果你开了 stream要改成for await (const chunk of stream)的方式消费。三是网络中断导致返回体被截断JSON 解析失败。这种情况在endSpan之前加一层 try-catch把原始响应打出来看。5.4 OAuth token expiredError: OAuth token has expired如果你用的是 Claude Code 或者类似的 CLI 工具做 Agent 入口可能会遇到 OAuth token 过期。这类工具通常把 token 存在本地配置目录里过期后需要重新走一次授权流程。排查思路是找到 token 的存储位置确认过期时间。以 Claude Code 为例配置一般在~/.claude/目录下。重新授权之后把新的 token 同步到你的环境变量里。如果你是通过 TaoToken 的 Coding Plan 来跑长期编码任务建议直接用 API Key 而不是 OAuth避免 token 过期打断 Agent 的长任务Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewrite5.5 Span 没有父子关系这个不算报错但数据看起来会很难受——所有 span 的 Parent ID 都是空的平铺在一起。原因几乎都是AsyncLocalStorageContextManager没生效。检查顺序context.setGlobalContextManager(contextManager)必须在new NodeSDK()之前调用。agentContextStorage.run(ctx, async () {...})必须包住整个 Agent 执行逻辑不能只包一部分。如果你用了Promise.all并发跑多个 LLM request注意 Claude Code 在endLLMRequestSpan()里明确提醒过多请求并发时必须传回原始 span 实例否则响应会被错误归因。OpenTelemetry 里对应的做法是用context.with(trace.setSpan(ctx, span), () {...})显式绑定。5.6 接入文档在哪上面这些排查如果还解决不了接入文档里有更完整的参数说明和示例接入文档https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewrite6. 把行为分析变成 Agent 的基础设施Claude Code 源码泄露这件事最有价值的不是那些被反复讨论的功能而是它展示了一个事实一个生产级 Agent 的可观测性代码需要和核心逻辑同步设计而不是事后补一个日志系统。三层分离、语义级 span、统一关联键、隐私治理前置、高基数管理——这五件事没有一件是加个 console.log能替代的。它们对应的是五个具体的工程决策产品事件和工程 trace 要不要分管道、span 的根节点选 interaction 还是 request、agentId 和 parentSessionId 从哪来、敏感字段在采集阶段还是查询阶段脱敏、session.id 要不要进指标系统。这套东西搭起来之后你回答问题的能力会完全不一样。以前你只能看到这个接口 P99 是 2 秒现在你能看到这一轮 interaction 里模型花了 1.2 秒工具执行花了 0.3 秒等用户审批花了 4.5 秒。以前你只能看到这个月 Token 涨了现在你能按 session、按 tool、按 prompt 版本拆开看是谁在烧。最后留一个实操建议先把agent.tool.blocked_on_user这个 span 加上。它是最容易被忽略、但信息量最大的一个。很多 Agent 的慢根本不是模型慢是卡在等人。把这个时间单独拎出来你会对 Agent 的真实瓶颈有完全不同的认知。
返回列表