部署指南

本文介绍 wesgine 在生产环境中的部署方式,包括单机部署、多实例部署、systemd 集成、健康检查、性能调优和日志管理。

概览

wesgine 是一个独立的 Go 二进制,支持两种使用方式:

方式说明适用场景
HTTP 二进制wesgine serve 独立进程Teleclaw、SaaS 部署
go mod 嵌入库方式嵌入宿主进程wesclaw/wescode/wescraft 桌面应用

本文主要针对 HTTP 二进制模式的服务器部署。


单机部署

最小启动

wesgine serve \
    --data-dir /var/lib/wesgine \
    --addr :9091 \
    --admin-token-bootstrap

启动参数:

参数说明默认值
--data-dir数据目录$XDG_DATA_HOME/wesgine
--addr监听地址:9091
--admin-token-bootstrap启动时输出 admin token否
--license-pathLicense 文件路径无

数据目录结构

/var/lib/wesgine/
├── hypervisor.db           Cell 注册表
├── secrets/
│   └── spec.key            AES-256-GCM 密钥(0600)
├── cells/
│   ├── dept-legal/
│   │   ├── meta.db         Cell 元数据
│   │   ├── state.db        运行时状态
│   │   ├── sessions.db     会话与记忆
│   │   ├── skills/         技能包
│   │   ├── knowledge/      知识库
│   │   └── workspace/      工作区文件
│   └── dept-audit/
│       └── ...
├── logs/
│   └── wesgine-stderr.log  崩溃日志(fd 2 dup2)
└── crashes/                崩溃归档

文件权限

# 数据目录
chmod 700 /var/lib/wesgine
chown wesgine:wesgine /var/lib/wesgine

# 密钥目录
chmod 700 /var/lib/wesgine/secrets
chmod 600 /var/lib/wesgine/secrets/spec.key

systemd 集成

Unit 文件

# /etc/systemd/system/wesgine.service
[Unit]
Description=wesgine AI Agent Engine
After=network-online.target
Wants=network-online.target

[Service]
Type=notify
User=wesgine
Group=wesgine

ExecStart=/usr/local/bin/wesgine serve \
    --data-dir /var/lib/wesgine \
    --addr :9091 \
    --admin-token-bootstrap

# 超时配置
TimeoutStartSec=120
TimeoutStopSec=30

# 重启策略
Restart=on-failure
RestartSec=5
StartLimitBurst=5
StartLimitIntervalSec=300

# 安全加固
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/var/lib/wesgine
PrivateTmp=true

# 资源限制
LimitNOFILE=65535
MemoryMax=4G

[Install]
WantedBy=multi-user.target

关键配置说明

Type=notify

引擎使用 sd_notify 通知 systemd 就绪状态。就绪条件是绑端口完成,不是所有 Cell 热完——这确保启动时间是有界常数(INV-SERVE-REACHABLE-01)。

TimeoutStartSec=120

承重参数。默认 90s 可能不够(首次启动需要解压 runtime)。引擎 sdnotify.Ready() 在绑端口之后、热 Cell 之前发出。

TimeoutStopSec=30

SIGTERM 后给引擎 30s 完成:排空活跃 Run → checkpoint → 释放 flock。超时后 SIGKILL。

StartLimitBurst=5 / StartLimitIntervalSec=300

5 分钟内重启超过 5 次则 unit 进入 failed 状态。Java watchdog 重启前需 reset-failed。

操作命令

# 启动/停止
systemctl start wesgine
systemctl stop wesgine
systemctl restart wesgine

# 状态
systemctl status wesgine

# 日志
journalctl -u wesgine -f        # 实时
journalctl -u wesgine --since today  # 今日

# 重置失败计数
systemctl reset-failed wesgine

多实例部署

场景

当单机无法满足性能需求,或需要高可用时:

                Load Balancer
               /      |      \
        wesgine-1  wesgine-2  wesgine-3
              \       |       /
              Shared Storage (NFS/Ceph)
                     |
                  PostgreSQL

关键约束

INV-RESIL-04(进程排他锁):同一 Cell 同一时刻只有一个写入者。多实例部署时,每个 Cell 只能被一个 wesgine 进程打开。

分片策略:

实例 1:Cell dept-legal, dept-audit
实例 2:Cell dept-finance, dept-hr
实例 3:Cell dept-sales, dept-ops

多实例 systemd

# /etc/systemd/system/wesgine@.service(模板)
[Service]
ExecStart=/usr/local/bin/wesgine serve \
    --data-dir /var/lib/wesgine/%i \
    --addr :909%i
systemctl start wesgine@1  # 端口 9091
systemctl start wesgine@2  # 端口 9092

健康检查配置

端点

端点用途认证
GET /healthLiveness 探针无
GET /metricsPrometheus 指标可选 admin token

/health 响应

{
    "status": "ready",
    "lifecycle": "ready",
    "version": "1.0.0",
    "started_at": "2026-09-14T00:00:00Z",
    "uptime_seconds": 86400
}

lifecycle 状态:

Nginx 健康检查

upstream wesgine {
    server 127.0.0.1:9091;
    # 被动健康检查
}

server {
    location /health {
        proxy_pass http://wesgine;
        proxy_connect_timeout 5s;
        proxy_read_timeout 5s;
    }
}

Kubernetes 探针

livenessProbe:
  httpGet:
    path: /health
    port: 9091
  initialDelaySeconds: 10
  periodSeconds: 15
  timeoutSeconds: 5
  failureThreshold: 3

readinessProbe:
  httpGet:
    path: /health
    port: 9091
  initialDelaySeconds: 5
  periodSeconds: 10

性能调优

SQLite 优化

wesgine 使用 SQLite 作为存储(三层 DB:meta.db / state.db / sessions.db),以下参数已内置:

PRAGMA值说明
journal_modeWAL并发读
synchronousNORMAL平衡性能与安全
mmap_size0禁用 mmap(INV-RESIL-08)
cache_size-6400064MB 缓存
busy_timeout50005s 等待锁

文件句柄

生产环境推荐 LimitNOFILE=65535。每个活跃 Cell 约占 10-20 个文件句柄(3 个 DB + WAL + SHM + skills + knowledge)。

内存

场景推荐内存
单 Cell512MB
10 Cells2GB
50 Cells4-8GB
100+ Cells8GB+

主要内存消耗:per-Cell SQLite 缓存、LLM 请求缓冲、SSE 事件环。

并发 Run 限制

spec := wesgine.CellSpec{
    ID: "dept-legal",
    Quotas: wesgine.CellQuotas{
        MaxConcurrentRuns: 5,   // 建议 3-10
        TokensPerMinute:  100000,
        TokensPerDay:     5000000,
    },
}

日志管理

日志架构

wesgine 进程
    │
    ├── fd 2 (stderr) ──dup2──→ {DataDir}/logs/wesgine-stderr.log
    │                            (Go panic 栈直写 fd 2,不经 os.Stderr)
    │
    └── bootLogger
        ├── 文件 sink ────────→ {DataDir}/logs/wesgine-stderr.log(Debug 级别)
        └── journal sink ─────→ systemd journal(Info 级别)

INV-CRASH-STDERR-01:崩溃现场只在 wesgine-stderr.log 里。journalctl 看不到 Go panic 栈。

排障第一步

# 1. 看引擎自己的日志(崩溃栈在这里)
cat /var/lib/wesgine/logs/wesgine-stderr.log

# 2. 看 systemd journal(业务日志)
journalctl -u wesgine --since "10 minutes ago"

# 3. 看引擎版本(每次启动首行)
head -5 /var/lib/wesgine/logs/wesgine-stderr.log
# → version=1.0.0 commit=abc1234 ...

日志轮转

# /etc/logrotate.d/wesgine
/var/lib/wesgine/logs/*.log {
    daily
    missingok
    rotate 14
    compress
    delaycompress
    notifempty
    copytruncate
}

使用 copytruncate:因为 wesgine 通过 dup2 持有 fd 2,不能用 create + signal 方式轮转。

Prometheus 指标

curl http://localhost:9091/metrics

关键指标:


Nginx 反向代理

upstream wesgine_backend {
    server 127.0.0.1:9091;
    keepalive 32;
}

server {
    listen 443 ssl;
    server_name ai.example.com;

    # SSE 长连接(必须 > 引擎超时 86400s)
    proxy_read_timeout 90000s;
    proxy_send_timeout 90000s;

    # SSE 必须关闭缓冲
    proxy_buffering off;
    proxy_cache off;

    location / {
        proxy_pass http://wesgine_backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }

    # SSE 端点
    location ~ ^/cells/[^/]+/run$ {
        proxy_pass http://wesgine_backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_buffering off;
        chunked_transfer_encoding on;
    }
}

升级流程

滚动升级

# 1. 备份数据
cp -r /var/lib/wesgine /var/lib/wesgine.bak

# 2. 替换二进制
cp wesgine-new /usr/local/bin/wesgine

# 3. 重启
systemctl restart wesgine

# 4. 验证
curl http://localhost:9091/health

版本降级注意

引擎只升不降。 新版本写的 CellSpec(如 SecretRef 从 string 改为 object)旧版本无法解码,会导致 Cell 丢失。降级前必须确认目标版本能解码当前 hypervisor.db。


检查清单

部署前

部署后