Skip to content

The Case for a Cloud Native Agent Harness:为什么 AI Agent 不该再“裸跑”在 Kubernetes 上? ​

“Coding agents 真正变得实用,不是因为它们能聊天,而是因为它们终于‘落地’了——拥有了可调度的执行上下文、可版本化的技能仓库、可编排的子代理拓扑,以及可审计的工具调用链。云原生 Agent Harness 不是锦上添花,而是生产级 AI 工作流的基础设施刚需。”


背景动机:当 LLM 从“对话引擎”进化为“分布式协作者” ​

过去两年,我们见证了 LLM 应用范式的剧烈迁移:从 Chat UI → CLI 工具链 → 自动化 Pipeline → 最终走向 Autonomous Agent Orchestrator。但一个尖锐的现实被长期低估:绝大多数 Agent 实现仍运行在单机 Python 进程中,依赖 subprocess 启动 kubectl、git 或 curl,把 Pod 当作黑盒 API 调用目标,而非原生协作单元。

这导致三类典型故障模式在 SRE 团队中高频复现:

  • 可观测性黑洞:Agent 执行 kubectl apply -f manifest.yaml 后,无法关联其调用与下游 Pod 的 container_status、OOMKilled 或 ImagePullBackOff;Prometheus 中无 agent_task_duration_seconds 指标,只有 http_request_duration_seconds(且埋点在 LLM 推理层);
  • 权限爆炸:为支持 kubectl/helm/terraform 等工具,ServiceAccount 被授予 cluster-admin,违背最小权限原则;
  • 状态漂移失控:Agent 在 /tmp 写入中间产物,重启即丢失;多个 Agent 实例并发修改同一 Git repo,引发 merge conflict;子任务失败后无重试语义或断点续跑能力。

CNCF 这篇博文直指核心:Agent 不应是“运行在云上的程序”,而应是“云原生的一等公民”。它需要自己的 PodSpec、VolumeClaimTemplate、PodDisruptionBudget,甚至 HorizontalPodAutoscaler —— 就像当年 DaemonSet 之于日志采集器、StatefulSet 之于数据库一样。


核心技术:Cloud Native Agent Harness 的四大支柱 ​

Cloud Native Agent Harness(下文简称 CN-AH)并非新项目,而是一套设计范式 + 参考实现(如 agent-harness),其本质是 将 Agent 生命周期抽象为 Kubernetes 原生资源对象,并通过 Operator 实现闭环管控。它包含四个不可分割的技术支柱:

1. 可声明式定义的 Agent Workload(AgentJob CRD) ​

AgentJob 是 CN-AH 的核心资源,类似 Job,但专为长时、多阶段、带外部交互的 Agent 任务设计:

yaml
apiVersion: agent.cncf.io/v1alpha1
kind: AgentJob
metadata:
  name: pr-reviewer-2026-0928
  namespace: ai-platform
spec:
  # 驱动 Agent 的 LLM 模型服务(必须是集群内可访问的 vLLM/KServe endpoint)
  llmEndpoint: "http://vllm-gemma2-27b.ai-platform.svc.cluster.local:8000/v1/chat/completions"
  
  # Agent 的技能包(Skills Bundle)—— 一个 OCI 镜像,含 tools/、skills/、config.yaml
  skillsBundle: "ghcr.io/ai-platform/skills-pr-reviewer:v1.3.0@sha256:abc123..."
  
  # 声明式挂载:Git repo、Kubernetes config、Secrets(自动注入 token)
  volumes:
    - name: git-repo
      gitRepo:
        repository: "https://github.com/org/repo.git"
        revision: "main"
    - name: kube-config
      projected:
        sources:
        - serviceAccountToken:
            path: token
        - configMap:
            name: k8s-cluster-config
            items:
            - key: kubeconfig
              path: kubeconfig

  # 强制执行约束:GPU 资源、容忍污点、指定节点池
  podTemplate:
    spec:
      nodeSelector:
        cloud.google.com/gke-accelerator: nvidia-l4
      tolerations:
      - key: "ai-workload"
        operator: "Exists"
        effect: "NoSchedule"
      containers:
      - name: agent-runner
        resources:
          limits:
            nvidia.com/gpu: 1

✅ 关键洞察:skillsBundle 是 CN-AH 的灵魂。它不是 Docker 镜像(不包含 runtime),而是 OCI Artifact with application/vnd.cncf.agent.skills.v1+tar mediaType,由 skaffold build --platform=agent-skills 构建,内容包括:

  • tools/kubectl.py(封装 kubernetes.client 的安全 wrapper)
  • skills/code_review.py(带 context-aware retry 和 diff parsing)
  • config.yaml(定义 tool schema、rate limit、fallback LLM)

2. 技能驱动的工具注册中心(Skills Registry) ​

CN-AH 要求所有工具(tool)必须通过 Skills Registry 注册,禁止硬编码 subprocess.run("curl ...")。Registry 提供统一的 tool discovery、schema validation 和调用审计:

python
# tools/kubectl.py
from agent_harness.tool import Tool, ToolInput, ToolOutput

class KubectlApply(Tool):
    name = "kubectl_apply"
    description = "Apply Kubernetes manifests from local file or stdin"
    
    def execute(self, input: ToolInput) -> ToolOutput:
        # 自动注入 RBAC-contextualized kubeconfig
        # 自动记录 audit log 到 OpenTelemetry Collector
        # 自动检测 ImagePullBackOff 并触发 remediation hook
        ...

Operator 会扫描 skillsBundle 中所有 tools/*.py,生成 ToolSchema CR,并在 AgentJob 创建时校验 tool_calls 是否在白名单内。

3. 子代理(Subagent)拓扑即代码(Topology-as-Code) ​

复杂 Agent 任务需分治:code_analyzer → security_scanner → ci_runner → pr_commenter。CN-AH 通过 SubagentTemplate 定义拓扑:

yaml
spec:
  subagents:
  - name: code-analyzer
    template:
      spec:
        skillsBundle: "ghcr.io/ai-platform/skills-code-analyzer:v2.1.0"
        resources:
          limits: {cpu: "2", memory: "4Gi"}
  - name: security-scanner
    template:
      spec:
        skillsBundle: "ghcr.io/ai-platform/skills-trivy-scan:v0.4.2"
        dependsOn: ["code-analyzer"]  # DAG 依赖

Operator 动态创建 Job 或 StatefulSet(对有状态子 agent),并通过 agent-harness/bus(基于 NATS JetStream 的轻量消息总线)传递结构化 payload。

4. 可观测性原生集成(OpenTelemetry First) ​

CN-AH 的 agent-runner sidecar 默认注入 OpenTelemetry Collector,自动捕获:

  • agent.task.started/finished(含 agent_id, skill_name, tool_call_count)
  • tool.call.duration(按 tool name、status、error_type 分维度)
  • llm.request.token_usage(从 vLLM / KServe 响应头提取)

无需修改业务逻辑,即可在 Grafana 中构建 Agent SLO Dashboard:

  • rate(agent_task_failed_total{job="agent-runner"}[1h]) < 0.05
  • histogram_quantile(0.95, sum(rate(tool_call_duration_seconds_bucket[1h])) by (le, tool)) < 10

运维建议:SRE 如何渐进式落地 CN-AH ​

阶段行动项风险提示
L1:工具沙箱化将现有 subprocess 工具重构为 Tool 类,注册到 Skills Registry;禁用 shell=True;所有 kubectl 调用强制走 kubernetes.client wrapper切勿跳过 ToolInput schema 验证,否则可能绕过 RBAC
L2:CRD 托管化使用 AgentJob 替代 CronJob 触发 Agent;将 skillsBundle 镜像推送到受信 registry(如 Harbor with cosign verification)注意 AgentJob 的 ttlSecondsAfterFinished 必须设为 >0,避免历史 Job 泄露敏感日志
L3:生产加固启用 PodSecurityPolicy(或 PodSecurityAdmission)限制 CAP_SYS_ADMIN;为 agent-runner 设置 readOnlyRootFilesystem: true;skillsBundle 镜像启用 cosign verify webhookSubagent 间通信若跨 namespace,需显式配置 NATS 认证,避免消息泄露
L4:SLO 驱动运维将 agent_task_duration_seconds 95% 分位数纳入 SLO;对 tool_call_failed_total 设置告警(如 5m 内 >10 次);定期扫描 skillsBundle CVE(Trivy + Syft)切忌将 LLM 推理服务(vLLM)与 AgentRunner 部署在同一 Pod —— GPU 显存争抢会导致 OOMKill 链式故障

🔑 经验法则:Agent 的 P99 延迟 = max(P99 LLM latency, P99 tool_call latency, P99 subagent orchestration latency)。监控必须覆盖全链路,而非只盯 /v1/chat/completions。


延伸阅读 ​

  • 📘 规范文档:CNCF Agent Harness Spec v1.0(重点关注 SkillsBundle OCI Layout 和 ToolSchema CRD 定义)
  • ⚙️ 参考实现:agent-harness/operator(基于 Kubebuilder v4,支持 K8s 1.26+)
  • 🛠️ 工具链:agentctl CLI(agentctl bundle build / agentctl job create / agentctl trace job pr-reviewer-2026-0928)
  • 🧪 实验环境:kind + k3s 快速验证脚本见 cncf/agent-harness/examples/kind-setup.sh
  • 📊 性能基准:CNCF SIG-AI 发布的 AgentHarness vs Bare-Python Benchmark 显示:在 100 并发 pr-reviewer 任务下,CN-AH 的平均端到端延迟降低 42%,失败率下降至 0.3%(裸跑为 8.7%),且 kubectl 调用成功率从 91% 提升至 99.99%(得益于自动重试与 RBAC-aware 错误分类)。

💡 最后的技术判断:CN-AH 不是替代 LangChain/LlamaIndex 的框架,而是为其提供生产就绪的执行底座。未来 12 个月,我们预计:

  • 主流 MLOps 平台(KServe, BentoML, MLflow)将内置 AgentJob controller;
  • GitOps 工具(Argo CD, Flux)将支持 AgentJob 作为 first-class sync target;
  • 真正的“AI-native infrastructure”不是更聪明的 LLM,而是让 Agent 像 Pod 一样可靠、可观测、可编排。

—— KnoAI 技术站 · 2026.10.01