故障排查

本文档汇总 wesgine 引擎的常见问题、诊断方法和解决方案。


排查原则

排障第一步永远是看引擎自己的日志,不是 journal。

引擎把 fd 2 dup2 到 {DataDir}/logs/wesgine-stderr.log,因为 Go 的 runtime.throw 直写 fd 2。崩溃现场只在那个文件里——journalctl 里永远看不到 panic 栈。

# 第一步:看崩溃日志
cat /var/lib/wesgine/logs/wesgine-stderr.log

# 第二步:看业务日志
journalctl -u wesgine -f --output=cat

# 第三步:看引擎版本和启动阶段
head -20 /var/lib/wesgine/logs/wesgine-stderr.log
# 输出包含:version=... commit=... 和各阶段里程碑

启动失败

症状:引擎进程启动后立即退出

检查清单:

  1. 数据目录权限
ls -la /var/lib/wesgine/
# 确保 wesgine 用户有读写权限
# 目录应为 700,文件应为 600
  1. 端口占用
ss -tlnp | grep 9091
# 如果已被占用,修改启动参数 --addr
  1. License 文件(如需要)
ls -la /opt/wesgine/license.yaml
# 确认文件存在且格式正确
  1. 数据目录锁
# 检查是否有残留的锁文件
ls /var/lib/wesgine/cells/*/. cell.lock

# 如果引擎已停止但锁文件存在,手动删除
rm /var/lib/wesgine/cells/*/.cell.lock

症状:systemd 报 timeout

# 查看启动超时
systemctl status wesgine
# 如果显示 "start operation timed out"

# 解决:增加 TimeoutStartSec
# /etc/systemd/system/wesgine.service
TimeoutStartSec=180

INV-SERVE-REACHABLE-01:热身(ActivatePersisted)在就绪之后后台运行。如果启动超时,检查是否有其他阻塞操作在就绪路径上。

症状:重复启动失败(StartLimitBurst 耗尽)

# 检查是否进入 failed 状态
systemctl status wesgine
# 如果显示 "start request repeated too quickly"

# 重置失败计数
sudo systemctl reset-failed wesgine

# 然后重启
sudo systemctl start wesgine

Cell 创建失败

症状:ErrCellLocked

error: cell "my-cell" is locked by another process

原因:另一个进程正在使用该 Cell(INV-RESIL-04 进程排他锁)。

解决:

# 确认没有其他 wesgine 进程
pgrep -a wesgine

# 如果确认无其他进程,删除锁文件
rm /var/lib/wesgine/cells/my-cell/.cell.lock

症状:ErrProviderStrategyUnset

error: provider strategy must be set when both private and shared providers exist

原因:Cell 同时有私有和共享 Provider,但没有指定 ProviderStrategy。

解决:

spec := wesgine.CellSpec{
	ProviderStrategy: "own_first",  // 或 "shared_first"
}

症状:ErrSpecUndecodable

error: spec bytes are not decodable

原因:引擎版本与 hypervisor.db 中的 CellSpec 格式不匹配(通常是版本降级导致)。

诊断:

# 检查引擎版本
/opt/wesgine/bin/wesgine version

# 检查 hypervisor.db 中的 spec
sqlite3 /var/lib/wesgine/hypervisor.db \
  "SELECT id, length(spec_json) FROM wes_cells;"

解决:使用与数据匹配的引擎版本,或从备份恢复。


Provider 连接问题

症状:ErrVisionNotSupported

error: model "gpt-3.5-turbo" does not support vision

原因:发送了包含图片的消息,但指定的模型未声明 SupportsVision。

解决:

Models: []config.ModelConfig{
	{
		Name:           "gpt-4o",
		SupportsVision: true,  // ★ 显式声明
	},
}

症状:Provider 认证失败

# 测试 Provider 连通性
curl -X POST http://localhost:9091/cells/my-cell/providers/test \
  -H "Authorization: Bearer <token>" \
  -d '{"provider": "openai"}'

# 带 vision 模式测试
curl -X POST http://localhost:9091/cells/my-cell/providers/test \
  -H "Authorization: Bearer <token>" \
  -d '{"provider": "openai", "mode": "vision"}'

症状:wes: Provider 认证失败(live-token)

error: provider "wes-openai" requires ProviderLiveKeyFn but none was bound

原因:wes: / org: 类型 Provider 需要动态 JWT,但 ProviderLiveKeyFn 未绑定(INV-CELL-08)。

解决:在 GetOrCreate 之前绑定:

liveTokens := &providerid.LiveTokenSource{}
spec.ProviderLiveKeyFn = liveTokens.Key
// auth 初始化后
liveTokens.Bind(auth.IdentityAccessToken, auth.ValidAccessToken)

症状:Provider fallback 不生效

检查 ProviderStrategy 配置:

# 查看 Cell Provider 列表
curl http://localhost:9091/cells/my-cell/providers \
  -H "Authorization: Bearer <token>"

性能问题

症状:Run 响应缓慢

诊断步骤:

  1. 检查 Provider 延迟
curl http://localhost:9091/admin/observe/provider-latency \
  -H "Authorization: Bearer <admin-token>"
  1. 检查 Cell 温度
curl http://localhost:9091/cells/my-cell/state \
  -H "Authorization: Bearer <token>"
# 如果 temperature 是 "cool" 或 "cold",首次请求会有额外延迟
  1. 检查并发 Run 数
curl http://localhost:9091/cells/my-cell/observe/stats \
  -H "Authorization: Bearer <token>"
# 查看 active_runs 是否接近 MaxConcurrentRuns

症状:内存使用过高

# 查看进程内存
ps aux | grep wesgine

# 检查 Cell 数量和状态
curl http://localhost:9091/admin/observe/snapshot \
  -H "Authorization: Bearer <admin-token>"

可能原因:

解决:

症状:SQLite 锁竞争

error: database is locked

原因:同一 Cell 被多个进程同时访问(违反 INV-RESIL-04)。

诊断:

# 检查是否有多个进程访问同一数据目录
fuser /var/lib/wesgine/cells/my-cell/sessions.db

解决:确保同一 Cell 同一时刻只有一个写入者。


日志分析

日志位置

日志位置内容
崩溃日志{DataDir}/logs/wesgine-stderr.logGo panic 栈(fd 2 直写)
业务日志journal(journalctl -u wesgine)slog 输出(Info+)
业务日志(文件){DataDir}/logs/wesgine-stderr.logslog 输出(Debug+,同文件)

关键日志模式

# 启动阶段
grep "boot:" /var/lib/wesgine/logs/wesgine-stderr.log
# 输出示例:
# wesgine: boot: hypervisor started
# wesgine: boot: cell "main" activated
# wesgine: boot: engine version sdk_version=v1.0.0

# Cell 生命周期
journalctl -u wesgine | grep "cell"

# Provider 错误
journalctl -u wesgine | grep "provider"

# Run 错误
journalctl -u wesgine | grep "run.*error"

# 降级事件
journalctl -u wesgine | grep "degraded"

审计日志

# 导出审计日志
curl "http://localhost:9091/cells/my-cell/audit/export?format=jsonl&from=2026-09-01" \
  -H "Authorization: Bearer <admin-token>" \
  -o audit.jsonl

# 审计摘要
curl http://localhost:9091/cells/my-cell/audit/summary \
  -H "Authorization: Bearer <admin-token>"

常用诊断命令

引擎状态

# 健康检查
curl -s http://localhost:9091/health | jq .

# Hypervisor 快照
curl -s http://localhost:9091/admin/observe/snapshot \
  -H "Authorization: Bearer <admin-token>" | jq .

# 引擎版本
/opt/wesgine/bin/wesgine version

Cell 状态

# 列出所有 Cell
curl -s http://localhost:9091/admin/cells \
  -H "Authorization: Bearer <admin-token>" | jq '.[] | {id, temperature}'

# Cell 详情
curl -s http://localhost:9091/admin/cells/my-cell \
  -H "Authorization: Bearer <admin-token>" | jq .

# Cell 运行时统计
curl -s http://localhost:9091/cells/my-cell/observe/stats \
  -H "Authorization: Bearer <token>" | jq .

Run 诊断

# 列出最近的 Run
curl -s "http://localhost:9091/cells/my-cell/observe/runs?limit=10" \
  -H "Authorization: Bearer <token>" | jq '.[] | {id, status, reason}'

# Run 详情
curl -s "http://localhost:9091/cells/my-cell/observe/runs/{runID}" \
  -H "Authorization: Bearer <token>" | jq .

# Run 工具调用追踪
curl -s "http://localhost:9091/cells/my-cell/observe/runs/{runID}/tools" \
  -H "Authorization: Bearer <token>" | jq .

记忆诊断

# 记忆统计
curl -s "http://localhost:9091/cells/my-cell/memory/stats" \
  -H "Authorization: Bearer <admin-token>" | jq .

# 记忆每层条数
curl -s "http://localhost:9091/cells/my-cell/memory/counts" \
  -H "Authorization: Bearer <token>" | jq .

# 验证记忆可召回性
curl -X POST "http://localhost:9091/cells/my-cell/memory/validate" \
  -H "Authorization: Bearer <admin-token>" \
  -d '{"entry_id": "mem-123"}'

数据库诊断

# SQLite 完整性检查
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA integrity_check;"
sqlite3 /var/lib/wesgine/cells/my-cell/state.db "PRAGMA integrity_check;"
sqlite3 /var/lib/wesgine/cells/my-cell/meta.db "PRAGMA integrity_check;"

# WAL 模式确认
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA journal_mode;"
# 应该返回 "wal"

# mmap 确认禁用
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA mmap_size;"
# 应该返回 0(INV-RESIL-08)

# 数据库大小
du -sh /var/lib/wesgine/cells/my-cell/*.db*

进程诊断

# 引擎进程状态
ps aux | grep wesgine

# 文件描述符使用
ls /proc/$(pgrep wesgine)/fd | wc -l

# goroutine 数量(如果引擎暴露 pprof)
curl http://localhost:9091/debug/pprof/goroutine?debug=1

# 内存使用
curl http://localhost:9091/debug/pprof/heap?debug=1

崩溃恢复

数据韧性

wesgine 使用三层 SQLite 数据库(INV-RESIL-01~10):

层文件损坏影响
metameta.dbCell 配置不可读
sessionssessions.db历史对话不可访问
statestate.db观测数据不可访问

关键不变量:

损坏恢复步骤

# 1. 检查损坏范围
for db in meta.db sessions.db state.db; do
  echo "=== $db ==="
  sqlite3 "/var/lib/wesgine/cells/my-cell/$db" "PRAGMA integrity_check;"
done

# 2. 损坏的数据库会被自动归档到 corrupt/ 目录
ls /var/lib/wesgine/cells/my-cell/corrupt/

# 3. 引擎会自动创建新的空数据库并降级运行
# 检查降级状态
curl http://localhost:9091/cells/my-cell/state | jq '.degraded'

获取帮助

如果以上方法无法解决问题:

  1. 收集以下信息:

    • 引擎版本:wesgine version
    • 崩溃日志:wesgine-stderr.log
    • 业务日志:journalctl -u wesgine --since "1 hour ago"
    • 健康检查结果:curl localhost:9091/health
    • Hypervisor 快照:curl localhost:9091/admin/observe/snapshot
  2. 检查已知问题列表

  3. 联系技术支持时附上上述信息