监控与告警

wesgine 提供多层次的可观测性支持:Prometheus 指标导出、健康检查端点、崩溃日志收集、运行追踪和证据链审计。


健康检查

Liveness 探针

GET /health

返回 Hypervisor 的生命周期状态,无需认证:

{
  "status": "ready",
  "version": "1.0.0",
  "started_at": "2026-09-14T10:00:00Z",
  "lifecycle": "ready"
}

lifecycle 取值:

状态说明
startingHypervisor 正在初始化
ready可以接受请求
draining正在排空,准备关闭

Readiness 探针

对于 Kubernetes / systemd 部署,使用 /health 端点:

INV-SERVE-REACHABLE-01

可达性不随数据量增长。wesgine serve 到达"可被探活"所需的时间是有界常数:

router.Start(绑端口)→ sdnotify.Ready() → ActivatePersisted(后台 goroutine)

ActivatePersisted 在监听之后才开始,不阻塞健康检查。首请求到达 Cool Cell 时同步 WarmUp,Cold Cell 返回 503 + Retry-After: 5。


Prometheus 指标

端点

GET /metrics

公开或需要 admin token,取决于配置。

核心指标

Hypervisor 级

指标类型说明
wesgine_cells_totalGaugeCell 总数
wesgine_cells_activeGauge活跃 Cell 数
wesgine_cells_by_temperatureGaugeVec各温度 Cell 数(hot/warm/cool/cold)
wesgine_provider_latency_secondsHistogramProvider 请求延迟
wesgine_provider_errors_totalCounterProvider 错误总数

Cell 级

指标类型说明
wesgine_cell_runs_totalCounterVecRun 总数(按 cell_id, status)
wesgine_cell_runs_activeGaugeVec活跃 Run 数(按 cell_id)
wesgine_cell_tokens_totalCounterVecToken 消耗(按 cell_id, model, direction)
wesgine_cell_tool_calls_totalCounterVec工具调用数(按 cell_id, tool_name)
wesgine_cell_memory_entriesGaugeVec记忆条目数(按 cell_id, layer)

数据韧性

指标类型说明
wesgine_db_snapshot_duration_secondsHistogram快照耗时
wesgine_db_snapshot_errors_totalCounter快照失败数
wesgine_db_degraded_layersGaugeVec降级层数

Prometheus 配置示例

# prometheus.yml
scrape_configs:
  - job_name: 'wesgine'
    scrape_interval: 15s
    static_configs:
      - targets: ['wesgine:9091']
    bearer_token: '<admin-token>'
    metrics_path: /metrics

Hypervisor 快照

全局观测接口,提供 Hypervisor 级别的运行时快照:

GET /admin/observe/snapshot

返回:

{
  "version": "1.0.0",
  "uptime_seconds": 3600,
  "cells_total": 5,
  "cells_active": 3,
  "cells_by_temperature": {
    "hot": 1,
    "warm": 2,
    "cool": 1,
    "cold": 1
  },
  "total_runs": 150,
  "active_runs": 2,
  "legacy_spec_rewrites": 0,
  "degraded_secrets": {}
}

Per-Cell 统计

GET /admin/observe/per-cell-stats?cell=dept-legal&cell=dept-hr

Provider 延迟统计

GET /admin/observe/provider-latency

返回各 Provider 的 P50/P95/P99 延迟。


崩溃日志

fd 2 落盘机制

wesgine 使用 dup2 把 fd 2 指向 {DataDir}/logs/wesgine-stderr.log。这是因为 Go 的 runtime.throw 绕过 os.Stderr 直写 fd 2——捕获崩溃是这段代码存在的唯一理由(INV-CRASH-STDERR-01)。

重要:journalctl 里看不到 panic 栈。排障第一步永远是查看引擎自己的日志文件。

崩溃日志查询

GET /admin/observe/crash-log
GET /cells/{id}/observe/crash-log

日志分级

输出目标级别说明
stderr log 文件Debug+完整日志
journal(systemd)Info+业务日志镜像

journal 对单元有速率限制且丢弃超额,Debug 灌进去会以"引擎突然安静了"的形状复现日志丢失。

每次启动的第一行

引擎每次启动会打印构建指纹:

wesgine: boot: version=1.0.0 commit=abc1234 built=2026-09-14T10:00:00Z

版本降级检测:对比启动日志中的 version,确保不会用旧二进制读新格式数据。


Cell 运行时统计

GET /cells/{id}/observe/stats

返回 Cell 的运行时统计快照:

{
  "cell_id": "dept-legal",
  "temperature": "warm",
  "active_runs": 1,
  "total_runs": 50,
  "total_tokens": 500000,
  "memory_entries": 1200,
  "knowledge_files": 30,
  "skills_installed": 5,
  "uptime_seconds": 7200,
  "last_activity": "2026-09-14T14:30:00Z"
}

Run 追踪

列出 Run

GET /cells/{id}/observe/runs?agent=legal-agent&limit=10

Run 详情

GET /cells/{id}/observe/runs/{runID}

包含:开始/结束时间、终止原因、Token 用量、Settlement 状态。

Run 工具调用追踪

GET /cells/{id}/observe/runs/{runID}/tools

返回每次工具调用的名称、参数摘要、结果摘要、耗时和错误信息。


证据链审计

wesgine 使用 hash chain 记录每次 Run 的证据,确保不可篡改:

GET /cells/{id}/observe/evidence

返回 JSONL 流。证据链按 Cell 全量不过滤(cell:admin scope),因为切片后缺环无法校验。

审计导出

GET /cells/{id}/audit/export?format=jsonl&from=2026-09-01&to=2026-09-14

支持 jsonl 和 csv 格式。

审计摘要

GET /cells/{id}/audit/summary

告警配置

Prometheus AlertManager 规则示例

groups:
  - name: wesgine
    rules:
      # Cell 长时间 Hot
      - alert: CellOverheated
        expr: wesgine_cells_by_temperature{temperature="hot"} > 0
        for: 30m
        labels:
          severity: warning
        annotations:
          summary: "Cell {{ $labels.cell_id }} has been hot for 30 minutes"

      # Provider 高延迟
      - alert: ProviderHighLatency
        expr: histogram_quantile(0.95, wesgine_provider_latency_seconds) > 10
        for: 5m
        labels:
          severity: critical

      # 降级层出现
      - alert: DegradedLayer
        expr: wesgine_db_degraded_layers > 0
        labels:
          severity: critical
        annotations:
          summary: "Cell {{ $labels.cell_id }} has degraded database layers"

      # 快照失败
      - alert: SnapshotFailed
        expr: increase(wesgine_db_snapshot_errors_total[1h]) > 0
        labels:
          severity: warning

      # Token 日用量接近上限
      - alert: TokenBudgetNearLimit
        expr: wesgine_cell_tokens_total / wesgine_cell_token_budget > 0.9
        labels:
          severity: warning

Grafana Dashboard 建议面板

面板数据源说明
Cell 温度分布wesgine_cells_by_temperature饼图
Run 吞吐量rate(wesgine_cell_runs_total[5m])时序图
Provider 延迟wesgine_provider_latency_seconds热力图
Token 消耗趋势wesgine_cell_tokens_total累积图
活跃 Run 数wesgine_cell_runs_active仪表盘
降级层告警wesgine_db_degraded_layers状态面板

最佳实践

  1. /health 只做 liveness 探针——不要拿它判断业务正确性
  2. 排障先看 stderr log——不是 journal,不是 HTTP 接口
  3. 监控 degraded_layers——降级是合法状态,但需要关注
  4. 设置 Token 日用量告警——防止失控 Run 造成账单异常
  5. 证据链定期导出备份——hash chain 是不可篡改的审计依据
  6. 版本升级后检查启动日志——确认构建指纹与预期一致