故障排查
本文档汇总 wesgine 引擎的常见问题、诊断方法和解决方案。
排查原则
排障第一步永远是看引擎自己的日志,不是 journal。
引擎把 fd 2 dup2 到 {DataDir}/logs/wesgine-stderr.log,因为 Go 的 runtime.throw 直写 fd 2。崩溃现场只在那个文件里——journalctl 里永远看不到 panic 栈。
# 第一步:看崩溃日志
cat /var/lib/wesgine/logs/wesgine-stderr.log
# 第二步:看业务日志
journalctl -u wesgine -f --output=cat
# 第三步:看引擎版本和启动阶段
head -20 /var/lib/wesgine/logs/wesgine-stderr.log
# 输出包含:version=... commit=... 和各阶段里程碑
启动失败
症状:引擎进程启动后立即退出
检查清单:
- 数据目录权限
ls -la /var/lib/wesgine/
# 确保 wesgine 用户有读写权限
# 目录应为 700,文件应为 600
- 端口占用
ss -tlnp | grep 9091
# 如果已被占用,修改启动参数 --addr
- License 文件(如需要)
ls -la /opt/wesgine/license.yaml
# 确认文件存在且格式正确
- 数据目录锁
# 检查是否有残留的锁文件
ls /var/lib/wesgine/cells/*/. cell.lock
# 如果引擎已停止但锁文件存在,手动删除
rm /var/lib/wesgine/cells/*/.cell.lock
症状:systemd 报 timeout
# 查看启动超时
systemctl status wesgine
# 如果显示 "start operation timed out"
# 解决:增加 TimeoutStartSec
# /etc/systemd/system/wesgine.service
TimeoutStartSec=180
INV-SERVE-REACHABLE-01:热身(
ActivatePersisted)在就绪之后后台运行。如果启动超时,检查是否有其他阻塞操作在就绪路径上。
症状:重复启动失败(StartLimitBurst 耗尽)
# 检查是否进入 failed 状态
systemctl status wesgine
# 如果显示 "start request repeated too quickly"
# 重置失败计数
sudo systemctl reset-failed wesgine
# 然后重启
sudo systemctl start wesgine
Cell 创建失败
症状:ErrCellLocked
error: cell "my-cell" is locked by another process
原因:另一个进程正在使用该 Cell(INV-RESIL-04 进程排他锁)。
解决:
# 确认没有其他 wesgine 进程
pgrep -a wesgine
# 如果确认无其他进程,删除锁文件
rm /var/lib/wesgine/cells/my-cell/.cell.lock
症状:ErrProviderStrategyUnset
error: provider strategy must be set when both private and shared providers exist
原因:Cell 同时有私有和共享 Provider,但没有指定 ProviderStrategy。
解决:
spec := wesgine.CellSpec{
ProviderStrategy: "own_first", // 或 "shared_first"
}
症状:ErrSpecUndecodable
error: spec bytes are not decodable
原因:引擎版本与 hypervisor.db 中的 CellSpec 格式不匹配(通常是版本降级导致)。
诊断:
# 检查引擎版本
/opt/wesgine/bin/wesgine version
# 检查 hypervisor.db 中的 spec
sqlite3 /var/lib/wesgine/hypervisor.db \
"SELECT id, length(spec_json) FROM wes_cells;"
解决:使用与数据匹配的引擎版本,或从备份恢复。
Provider 连接问题
症状:ErrVisionNotSupported
error: model "gpt-3.5-turbo" does not support vision
原因:发送了包含图片的消息,但指定的模型未声明 SupportsVision。
解决:
Models: []config.ModelConfig{
{
Name: "gpt-4o",
SupportsVision: true, // ★ 显式声明
},
}
症状:Provider 认证失败
# 测试 Provider 连通性
curl -X POST http://localhost:9091/cells/my-cell/providers/test \
-H "Authorization: Bearer <token>" \
-d '{"provider": "openai"}'
# 带 vision 模式测试
curl -X POST http://localhost:9091/cells/my-cell/providers/test \
-H "Authorization: Bearer <token>" \
-d '{"provider": "openai", "mode": "vision"}'
症状:wes: Provider 认证失败(live-token)
error: provider "wes-openai" requires ProviderLiveKeyFn but none was bound
原因:wes: / org: 类型 Provider 需要动态 JWT,但 ProviderLiveKeyFn 未绑定(INV-CELL-08)。
解决:在 GetOrCreate 之前绑定:
liveTokens := &providerid.LiveTokenSource{}
spec.ProviderLiveKeyFn = liveTokens.Key
// auth 初始化后
liveTokens.Bind(auth.IdentityAccessToken, auth.ValidAccessToken)
症状:Provider fallback 不生效
检查 ProviderStrategy 配置:
# 查看 Cell Provider 列表
curl http://localhost:9091/cells/my-cell/providers \
-H "Authorization: Bearer <token>"
性能问题
症状:Run 响应缓慢
诊断步骤:
- 检查 Provider 延迟
curl http://localhost:9091/admin/observe/provider-latency \
-H "Authorization: Bearer <admin-token>"
- 检查 Cell 温度
curl http://localhost:9091/cells/my-cell/state \
-H "Authorization: Bearer <token>"
# 如果 temperature 是 "cool" 或 "cold",首次请求会有额外延迟
- 检查并发 Run 数
curl http://localhost:9091/cells/my-cell/observe/stats \
-H "Authorization: Bearer <token>"
# 查看 active_runs 是否接近 MaxConcurrentRuns
症状:内存使用过高
# 查看进程内存
ps aux | grep wesgine
# 检查 Cell 数量和状态
curl http://localhost:9091/admin/observe/snapshot \
-H "Authorization: Bearer <admin-token>"
可能原因:
- 过多 Hot/Warm Cell 同时存在
- 大量 SSE 连接未关闭
- 知识库索引过大
解决:
- 调整温度调度策略(idle timeout)
- 检查客户端 SSE 连接管理
- 清理不需要的知识库文件
症状:SQLite 锁竞争
error: database is locked
原因:同一 Cell 被多个进程同时访问(违反 INV-RESIL-04)。
诊断:
# 检查是否有多个进程访问同一数据目录
fuser /var/lib/wesgine/cells/my-cell/sessions.db
解决:确保同一 Cell 同一时刻只有一个写入者。
日志分析
日志位置
| 日志 | 位置 | 内容 |
|---|---|---|
| 崩溃日志 | {DataDir}/logs/wesgine-stderr.log | Go panic 栈(fd 2 直写) |
| 业务日志 | journal(journalctl -u wesgine) | slog 输出(Info+) |
| 业务日志(文件) | {DataDir}/logs/wesgine-stderr.log | slog 输出(Debug+,同文件) |
关键日志模式
# 启动阶段
grep "boot:" /var/lib/wesgine/logs/wesgine-stderr.log
# 输出示例:
# wesgine: boot: hypervisor started
# wesgine: boot: cell "main" activated
# wesgine: boot: engine version sdk_version=v1.0.0
# Cell 生命周期
journalctl -u wesgine | grep "cell"
# Provider 错误
journalctl -u wesgine | grep "provider"
# Run 错误
journalctl -u wesgine | grep "run.*error"
# 降级事件
journalctl -u wesgine | grep "degraded"
审计日志
# 导出审计日志
curl "http://localhost:9091/cells/my-cell/audit/export?format=jsonl&from=2026-09-01" \
-H "Authorization: Bearer <admin-token>" \
-o audit.jsonl
# 审计摘要
curl http://localhost:9091/cells/my-cell/audit/summary \
-H "Authorization: Bearer <admin-token>"
常用诊断命令
引擎状态
# 健康检查
curl -s http://localhost:9091/health | jq .
# Hypervisor 快照
curl -s http://localhost:9091/admin/observe/snapshot \
-H "Authorization: Bearer <admin-token>" | jq .
# 引擎版本
/opt/wesgine/bin/wesgine version
Cell 状态
# 列出所有 Cell
curl -s http://localhost:9091/admin/cells \
-H "Authorization: Bearer <admin-token>" | jq '.[] | {id, temperature}'
# Cell 详情
curl -s http://localhost:9091/admin/cells/my-cell \
-H "Authorization: Bearer <admin-token>" | jq .
# Cell 运行时统计
curl -s http://localhost:9091/cells/my-cell/observe/stats \
-H "Authorization: Bearer <token>" | jq .
Run 诊断
# 列出最近的 Run
curl -s "http://localhost:9091/cells/my-cell/observe/runs?limit=10" \
-H "Authorization: Bearer <token>" | jq '.[] | {id, status, reason}'
# Run 详情
curl -s "http://localhost:9091/cells/my-cell/observe/runs/{runID}" \
-H "Authorization: Bearer <token>" | jq .
# Run 工具调用追踪
curl -s "http://localhost:9091/cells/my-cell/observe/runs/{runID}/tools" \
-H "Authorization: Bearer <token>" | jq .
记忆诊断
# 记忆统计
curl -s "http://localhost:9091/cells/my-cell/memory/stats" \
-H "Authorization: Bearer <admin-token>" | jq .
# 记忆每层条数
curl -s "http://localhost:9091/cells/my-cell/memory/counts" \
-H "Authorization: Bearer <token>" | jq .
# 验证记忆可召回性
curl -X POST "http://localhost:9091/cells/my-cell/memory/validate" \
-H "Authorization: Bearer <admin-token>" \
-d '{"entry_id": "mem-123"}'
数据库诊断
# SQLite 完整性检查
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA integrity_check;"
sqlite3 /var/lib/wesgine/cells/my-cell/state.db "PRAGMA integrity_check;"
sqlite3 /var/lib/wesgine/cells/my-cell/meta.db "PRAGMA integrity_check;"
# WAL 模式确认
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA journal_mode;"
# 应该返回 "wal"
# mmap 确认禁用
sqlite3 /var/lib/wesgine/cells/my-cell/sessions.db "PRAGMA mmap_size;"
# 应该返回 0(INV-RESIL-08)
# 数据库大小
du -sh /var/lib/wesgine/cells/my-cell/*.db*
进程诊断
# 引擎进程状态
ps aux | grep wesgine
# 文件描述符使用
ls /proc/$(pgrep wesgine)/fd | wc -l
# goroutine 数量(如果引擎暴露 pprof)
curl http://localhost:9091/debug/pprof/goroutine?debug=1
# 内存使用
curl http://localhost:9091/debug/pprof/heap?debug=1
崩溃恢复
数据韧性
wesgine 使用三层 SQLite 数据库(INV-RESIL-01~10):
| 层 | 文件 | 损坏影响 |
|---|---|---|
| meta | meta.db | Cell 配置不可读 |
| sessions | sessions.db | 历史对话不可访问 |
| state | state.db | 观测数据不可访问 |
关键不变量:
- INV-RESIL-01:Cell Boot 永远成功。无论何种数据损坏,Boot 都返回 nil error
- INV-RESIL-02:
Runtime().Run()永远可用。降级只影响历史数据 - INV-RESIL-03:单层损坏不扩散
损坏恢复步骤
# 1. 检查损坏范围
for db in meta.db sessions.db state.db; do
echo "=== $db ==="
sqlite3 "/var/lib/wesgine/cells/my-cell/$db" "PRAGMA integrity_check;"
done
# 2. 损坏的数据库会被自动归档到 corrupt/ 目录
ls /var/lib/wesgine/cells/my-cell/corrupt/
# 3. 引擎会自动创建新的空数据库并降级运行
# 检查降级状态
curl http://localhost:9091/cells/my-cell/state | jq '.degraded'
获取帮助
如果以上方法无法解决问题:
-
收集以下信息:
- 引擎版本:
wesgine version - 崩溃日志:
wesgine-stderr.log - 业务日志:
journalctl -u wesgine --since "1 hour ago" - 健康检查结果:
curl localhost:9091/health - Hypervisor 快照:
curl localhost:9091/admin/observe/snapshot
- 引擎版本:
-
检查已知问题列表
-
联系技术支持时附上上述信息