监控与告警
wesgine 提供多层次的可观测性支持:Prometheus 指标导出、健康检查端点、崩溃日志收集、运行追踪和证据链审计。
健康检查
Liveness 探针
GET /health
返回 Hypervisor 的生命周期状态,无需认证:
{
"status": "ready",
"version": "1.0.0",
"started_at": "2026-09-14T10:00:00Z",
"lifecycle": "ready"
}
lifecycle 取值:
| 状态 | 说明 |
|---|---|
starting | Hypervisor 正在初始化 |
ready | 可以接受请求 |
draining | 正在排空,准备关闭 |
Readiness 探针
对于 Kubernetes / systemd 部署,使用 /health 端点:
lifecycle=ready→ 返回 200,可接收流量lifecycle=starting→ 返回 503,等待就绪lifecycle=draining→ 返回 503,正在关闭
INV-SERVE-REACHABLE-01
可达性不随数据量增长。wesgine serve 到达"可被探活"所需的时间是有界常数:
router.Start(绑端口)→ sdnotify.Ready() → ActivatePersisted(后台 goroutine)
ActivatePersisted 在监听之后才开始,不阻塞健康检查。首请求到达 Cool Cell 时同步 WarmUp,Cold Cell 返回 503 + Retry-After: 5。
Prometheus 指标
端点
GET /metrics
公开或需要 admin token,取决于配置。
核心指标
Hypervisor 级
| 指标 | 类型 | 说明 |
|---|---|---|
wesgine_cells_total | Gauge | Cell 总数 |
wesgine_cells_active | Gauge | 活跃 Cell 数 |
wesgine_cells_by_temperature | GaugeVec | 各温度 Cell 数(hot/warm/cool/cold) |
wesgine_provider_latency_seconds | Histogram | Provider 请求延迟 |
wesgine_provider_errors_total | Counter | Provider 错误总数 |
Cell 级
| 指标 | 类型 | 说明 |
|---|---|---|
wesgine_cell_runs_total | CounterVec | Run 总数(按 cell_id, status) |
wesgine_cell_runs_active | GaugeVec | 活跃 Run 数(按 cell_id) |
wesgine_cell_tokens_total | CounterVec | Token 消耗(按 cell_id, model, direction) |
wesgine_cell_tool_calls_total | CounterVec | 工具调用数(按 cell_id, tool_name) |
wesgine_cell_memory_entries | GaugeVec | 记忆条目数(按 cell_id, layer) |
数据韧性
| 指标 | 类型 | 说明 |
|---|---|---|
wesgine_db_snapshot_duration_seconds | Histogram | 快照耗时 |
wesgine_db_snapshot_errors_total | Counter | 快照失败数 |
wesgine_db_degraded_layers | GaugeVec | 降级层数 |
Prometheus 配置示例
# prometheus.yml
scrape_configs:
- job_name: 'wesgine'
scrape_interval: 15s
static_configs:
- targets: ['wesgine:9091']
bearer_token: '<admin-token>'
metrics_path: /metrics
Hypervisor 快照
全局观测接口,提供 Hypervisor 级别的运行时快照:
GET /admin/observe/snapshot
返回:
{
"version": "1.0.0",
"uptime_seconds": 3600,
"cells_total": 5,
"cells_active": 3,
"cells_by_temperature": {
"hot": 1,
"warm": 2,
"cool": 1,
"cold": 1
},
"total_runs": 150,
"active_runs": 2,
"legacy_spec_rewrites": 0,
"degraded_secrets": {}
}
Per-Cell 统计
GET /admin/observe/per-cell-stats?cell=dept-legal&cell=dept-hr
Provider 延迟统计
GET /admin/observe/provider-latency
返回各 Provider 的 P50/P95/P99 延迟。
崩溃日志
fd 2 落盘机制
wesgine 使用 dup2 把 fd 2 指向 {DataDir}/logs/wesgine-stderr.log。这是因为 Go 的 runtime.throw 绕过 os.Stderr 直写 fd 2——捕获崩溃是这段代码存在的唯一理由(INV-CRASH-STDERR-01)。
重要:journalctl 里看不到 panic 栈。排障第一步永远是查看引擎自己的日志文件。
崩溃日志查询
GET /admin/observe/crash-log
GET /cells/{id}/observe/crash-log
日志分级
| 输出目标 | 级别 | 说明 |
|---|---|---|
| stderr log 文件 | Debug+ | 完整日志 |
| journal(systemd) | Info+ | 业务日志镜像 |
journal 对单元有速率限制且丢弃超额,Debug 灌进去会以"引擎突然安静了"的形状复现日志丢失。
每次启动的第一行
引擎每次启动会打印构建指纹:
wesgine: boot: version=1.0.0 commit=abc1234 built=2026-09-14T10:00:00Z
版本降级检测:对比启动日志中的 version,确保不会用旧二进制读新格式数据。
Cell 运行时统计
GET /cells/{id}/observe/stats
返回 Cell 的运行时统计快照:
{
"cell_id": "dept-legal",
"temperature": "warm",
"active_runs": 1,
"total_runs": 50,
"total_tokens": 500000,
"memory_entries": 1200,
"knowledge_files": 30,
"skills_installed": 5,
"uptime_seconds": 7200,
"last_activity": "2026-09-14T14:30:00Z"
}
Run 追踪
列出 Run
GET /cells/{id}/observe/runs?agent=legal-agent&limit=10
Run 详情
GET /cells/{id}/observe/runs/{runID}
包含:开始/结束时间、终止原因、Token 用量、Settlement 状态。
Run 工具调用追踪
GET /cells/{id}/observe/runs/{runID}/tools
返回每次工具调用的名称、参数摘要、结果摘要、耗时和错误信息。
证据链审计
wesgine 使用 hash chain 记录每次 Run 的证据,确保不可篡改:
GET /cells/{id}/observe/evidence
返回 JSONL 流。证据链按 Cell 全量不过滤(cell:admin scope),因为切片后缺环无法校验。
审计导出
GET /cells/{id}/audit/export?format=jsonl&from=2026-09-01&to=2026-09-14
支持 jsonl 和 csv 格式。
审计摘要
GET /cells/{id}/audit/summary
告警配置
Prometheus AlertManager 规则示例
groups:
- name: wesgine
rules:
# Cell 长时间 Hot
- alert: CellOverheated
expr: wesgine_cells_by_temperature{temperature="hot"} > 0
for: 30m
labels:
severity: warning
annotations:
summary: "Cell {{ $labels.cell_id }} has been hot for 30 minutes"
# Provider 高延迟
- alert: ProviderHighLatency
expr: histogram_quantile(0.95, wesgine_provider_latency_seconds) > 10
for: 5m
labels:
severity: critical
# 降级层出现
- alert: DegradedLayer
expr: wesgine_db_degraded_layers > 0
labels:
severity: critical
annotations:
summary: "Cell {{ $labels.cell_id }} has degraded database layers"
# 快照失败
- alert: SnapshotFailed
expr: increase(wesgine_db_snapshot_errors_total[1h]) > 0
labels:
severity: warning
# Token 日用量接近上限
- alert: TokenBudgetNearLimit
expr: wesgine_cell_tokens_total / wesgine_cell_token_budget > 0.9
labels:
severity: warning
Grafana Dashboard 建议面板
| 面板 | 数据源 | 说明 |
|---|---|---|
| Cell 温度分布 | wesgine_cells_by_temperature | 饼图 |
| Run 吞吐量 | rate(wesgine_cell_runs_total[5m]) | 时序图 |
| Provider 延迟 | wesgine_provider_latency_seconds | 热力图 |
| Token 消耗趋势 | wesgine_cell_tokens_total | 累积图 |
| 活跃 Run 数 | wesgine_cell_runs_active | 仪表盘 |
| 降级层告警 | wesgine_db_degraded_layers | 状态面板 |
最佳实践
/health只做 liveness 探针——不要拿它判断业务正确性- 排障先看 stderr log——不是 journal,不是 HTTP 接口
- 监控
degraded_layers——降级是合法状态,但需要关注 - 设置 Token 日用量告警——防止失控 Run 造成账单异常
- 证据链定期导出备份——hash chain 是不可篡改的审计依据
- 版本升级后检查启动日志——确认构建指纹与预期一致