性能调优
wescode 帮助你发现性能瓶颈、分析性能数据、生成优化方案,覆盖从代码级到架构级的性能优化。
功能概览
| 能力 | 说明 |
|---|---|
| 性能分析 | CPU、内存、I/O 性能剖析 |
| 瓶颈定位 | 基于调用图和执行数据定位热点 |
| 优化建议 | 针对具体瓶颈给出可执行的优化方案 |
| 基准测试 | 生成和运行性能基准测试 |
| 监控集成 | 添加性能指标采集和告警 |
性能分析
代码级性能分析
将性能剖析数据(如 Go pprof、Python cProfile)交给 AI 分析:
用户:帮我分析这个 pprof 结果
AI:CPU Profile 分析:
Top 5 热点函数:
┌─────────────────────────────┬───────┬───────┐
│ 函数 │ 占比 │ 累计 │
├─────────────────────────────┼───────┼───────┤
│ json.Marshal │ 35.2% │ 35.2% │
│ runtime.mallocgc │ 18.7% │ 53.9% │
│ database/sql.(*DB).Query │ 15.3% │ 69.2% │
│ crypto/sha256.block │ 8.1% │ 77.3% │
│ net/http.(*Transport).dial.. │ 6.4% │ 83.7% │
└─────────────────────────────┴───────┴───────┘
分析:
1. json.Marshal 占 35% — 序列化开销过大
→ 考虑使用 json-iterator 或 sonic 替代标准库
→ 或者对热路径使用手动序列化
2. runtime.mallocgc 占 19% — 内存分配频繁
→ 使用 sync.Pool 复用对象
→ 减少不必要的 string ↔ []byte 转换
3. database/sql.Query 占 15% — 数据库查询耗时
→ 检查是否有 N+1 查询
→ 考虑添加缓存层
内存分析
用户:程序内存使用越来越高,帮我分析
AI:内存问题排查指南:
1. 获取 heap profile
curl http://localhost:6060/debug/pprof/heap > heap.pb.gz
go tool pprof heap.pb.gz
2. 常见内存泄漏模式:
① Goroutine 泄漏 — goroutine 数量持续增长
检查:runtime.NumGoroutine()
原因:channel 未关闭、context 未取消
② 缓存无上限 — map 只增不减
检查:heap profile 中的 inuse_objects
修复:使用 LRU 或设置 TTL
③ 全局变量累积 — 全局 slice/map 不断 append
检查:heap profile 中的 alloc_space
修复:定期清理或使用弱引用
3. 工具建议:
go tool pprof -http=:8080 heap.pb.gz
→ 打开浏览器可视化内存分配图
瓶颈定位
基于 CKG 的调用路径分析
wescode 利用 CKG 结合运行时数据定位性能瓶颈:
用户:API 响应时间从 50ms 增长到 500ms,帮我定位
AI:性能退化分析:
请求路径追踪:
HandleGetOrder (2ms)
├── ValidateToken (5ms) ✓
├── GetOrder (480ms) ✗ ← 瓶颈
│ ├── QueryOrder (3ms) ✓
│ ├── QueryOrderItems (5ms) ✓
│ ├── QueryUserProfile (450ms) ✗ ← 根因
│ │ └── 外部 API 调用超时
│ └── CalculateDiscount (2ms) ✓
└── SerializeResponse (8ms) ✓
根因定位:
QueryUserProfile 调用外部用户服务,响应时间从 10ms 退化到 450ms
建议:
1. 添加缓存:用户 profile 短时间内不会变化
→ Redis 缓存,TTL 5 分钟
2. 添加超时:限制外部调用最大耗时
→ ctx, cancel := context.WithTimeout(ctx, 2*time.Second)
3. 降级策略:外部服务不可用时返回基本信息
并发瓶颈
用户:高并发下性能下降严重
AI:并发瓶颈分析:
可能的瓶颈点:
1. 锁竞争
→ 检查 sync.Mutex 的持有时间
→ 使用 go tool trace 分析 goroutine 阻塞
2. 数据库连接池
→ 检查 db.SetMaxOpenConns() 配置
→ 监控连接池等待队列
3. 外部服务串行调用
→ 多个独立外部调用应该并行化
→ 使用 errgroup.Group 管理并发
快速诊断命令:
# 查看 goroutine 阻塞情况
curl http://localhost:6060/debug/pprof/block > block.pb.gz
go tool pprof block.pb.gz
# 查看 mutex 竞争
curl http://localhost:6060/debug/pprof/mutex > mutex.pb.gz
go tool pprof mutex.pb.gz
优化建议
代码级优化
用户:优化这段 JSON 处理代码
AI:发现以下优化机会:
原始代码问题:
func processItems(data []byte) []Item {
var items []Item
json.Unmarshal(data, &items) // ① 标准库 JSON
var result []Item
for _, item := range items {
if item.Active {
result = append(result, item) // ② 多次扩容
}
}
return result
}
优化后:
var itemPool = sync.Pool{ // ③ 对象池
New: func() interface{} { return &[]Item{} },
}
func processItems(data []byte) []Item {
items := itemPool.Get().(*[]Item)
defer func() {
*items = (*items)[:0]
itemPool.Put(items)
}()
sonic.Unmarshal(data, items) // ① sonic 替代标准库
result := make([]Item, 0, len(*items)) // ② 预分配容量
for i := range *items {
if (*items)[i].Active {
result = append(result, (*items)[i])
}
}
return result
}
预期改善:
- JSON 解析:~3x 提升(sonic vs encoding/json)
- 内存分配:减少 ~60%(预分配 + 对象池)
- GC 压力:显著降低
架构级优化
用户:系统整体 QPS 不够,怎么优化?
AI:架构级优化建议:
当前架构瓶颈分析:
┌────────┐ ┌────────┐ ┌──────────┐
│ Client │ → │ Server │ → │ Database │ ← 瓶颈
└────────┘ └────────┘ └──────────┘
优化方案(按收益/成本排序):
1. 添加缓存层(收益高,成本低)
┌────────┐ ┌────────┐ ┌───────┐ ┌──────┐
│ Client │ → │ Server │ → │ Redis │ → │ DB │
└────────┘ └────────┘ └───────┘ └──────┘
→ 热数据缓存命中率预期 80%+
2. 读写分离(收益中,成本中)
读请求 → 从库(可水平扩展)
写请求 → 主库
3. 数据库优化(收益中,成本低)
→ 慢查询优化(添加索引)
→ 连接池调优
→ 减少不必要的字段查询
4. 异步化(收益高,成本中)
→ 非实时操作改为消息队列异步处理
→ 如:发送通知、生成报表、日志记录
基准测试
生成基准测试
用户:为 JSON 处理函数生成基准测试
AI:生成基准测试(benchmark_test.go):
func BenchmarkProcessItems(b *testing.B) {
data := generateTestData(1000) // 1000 条测试数据
b.ResetTimer()
b.ReportAllocs()
for i := 0; i < b.N; i++ {
processItems(data)
}
}
func BenchmarkProcessItems_Parallel(b *testing.B) {
data := generateTestData(1000)
b.ResetTimer()
b.RunParallel(func(pb *testing.PB) {
for pb.Next() {
processItems(data)
}
})
}
运行命令:
go test -bench=BenchmarkProcessItems -benchmem -count=5
基准对比
用户:对比优化前后的性能
AI:基准测试对比:
优化前:
BenchmarkProcessItems-8 5000 312456 ns/op 45678 B/op 234 allocs/op
优化后:
BenchmarkProcessItems-8 15000 105234 ns/op 12345 B/op 56 allocs/op
改善幅度:
├── 吞吐量:+200%(5000 → 15000 ops)
├── 延迟:-66%(312μs → 105μs)
├── 内存:-73%(45KB → 12KB per op)
└── 分配次数:-76%(234 → 56 allocs)
监控集成
添加性能指标
用户:帮我添加 Prometheus 性能监控
AI:添加关键性能指标采集:
// internal/metrics/metrics.go
var (
RequestDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "HTTP 请求耗时分布",
Buckets: []float64{.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5},
},
[]string{"method", "path", "status"},
)
DatabaseQueryDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "db_query_duration_seconds",
Help: "数据库查询耗时分布",
Buckets: []float64{.001, .005, .01, .025, .05, .1, .25, .5, 1},
},
[]string{"operation", "table"},
)
CacheHitRate = prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "cache_requests_total",
Help: "缓存请求计数",
},
[]string{"result"}, // "hit" or "miss"
)
ActiveGoroutines = prometheus.NewGauge(
prometheus.GaugeOpts{
Name: "active_goroutines",
Help: "当前活跃 goroutine 数量",
},
)
)
建议告警规则:
- P99 延迟 > 500ms → 告警
- 错误率 > 1% → 告警
- 缓存命中率 < 70% → 告警
- Goroutine 数量 > 10000 → 告警
注意事项
- 性能优化应基于实际数据,不要过早优化
- 基准测试结果受环境影响,建议在稳定环境中运行
- 优化前务必有完整的测试覆盖,确保不引入回归
- 架构级优化需要评估运维复杂度的增加
- 监控指标不宜过多,关注关键业务指标