From 811a008545cc0b1fecd84fce0cd2ff8e8cab1429 Mon Sep 17 00:00:00 2001 From: Rodger-Wang <1367893453@qq.com> Date: Fri, 18 Sep 2026 22:52:31 +0800 Subject: [PATCH] =?UTF-8?q?services:=20=E8=AF=AD=E9=9F=B3=E7=BA=AA?= =?UTF-8?q?=E8=A6=81=E2=86=92=E6=8B=BE=E5=BF=86=E2=86=92MCP=20=E9=93=BE?= =?UTF-8?q?=E8=B7=AF=E5=AE=A1=E8=AE=A1=E8=90=BD=E5=9C=B0=EF=BC=9B=E4=BF=AE?= =?UTF-8?q?=E4=B8=A4=E4=B8=AA=E3=80=8C=E4=BB=80=E4=B9=88=E9=83=BD=E6=B2=A1?= =?UTF-8?q?=E5=8F=91=E7=94=9F=E3=80=8D=E7=9A=84=E6=97=A2=E6=9C=89=20bug?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 审计与落地计划见 docs/语音纪要-拾忆-MCP链路审计与落地计划.md,已在阿龙测试环境 逐条验收通过(记录见该文档 §5.1)。 ## 安全 echomeet 八个按 id 操作的接口**一处都没有比对 record.Uid 与会话 uid**,任何登录用户 改一个 id 就能读别人纪要全文、删改别人的记录、给别人的记录发起总结(扣自己算力, 但覆盖对方 summary 并触发对方的拾忆抽取)。统一改走带 uid 的查询: - 「不存在」与「不是你的」回同一个错误码同一句话——分开回等于给出一个探测他人 id 的接口;用的还是改动前记录真不存在时的 DBError,合法用户行为无任何变化; - 批量接口按交集处理、**不报错**:客户端轮询队列里本就可能留着已被别的设备删掉的 id, 报错会把整个轮询循环掀掉;delrecords 响应回实际删掉的那批; - starttask 的校验排在扣算力之前,否则拒绝了还先把钱扣了; - 回调接口是第三方来的、没有 session,不在此列。 - TestApisUseUidScopedQueries 扫源码守着:最容易复发的不是有人改回去, 而是新加接口时照旧写法抄一遍,那样既不报错也看不出来。 ## gorm 的 default: 标签吞零值 mysql.Insert = db.Table().Create(model),gorm 对带 default: 的字段**一律把零值换成 默认值**。DBMemoryItem 的 date_certain(default:true) / remind_ahead(default:5) 因此 永远写不进 false/0:测试库 25 条记录全部落成 1/5,「待定日期」分组恒为空, 会议待办还会在 08:55 响提醒——与「会议待办默认不提醒」正好相反。 ⚠️ 加 Select("*") **救不了**:替换在 ConvertToCreateValues 的 reflect.Struct 分支里 无条件做,只看值是不是零值,与 Select/Omit 无关(DryRun 实测两种写法生成的 VALUES 一模一样)。改成插入后把这两列按结构体现值补写回去,并把真值还给调用方—— gorm 连结构体上的字段一起改了,不还原的话 memory_add 回给客户端的也是错的。 存量由 migrate_defaults.go 幂等修正(只碰 source=meeting 且 user_edited=0)。 ## 内容门槛 6 秒的测试通话、20 秒的单人自述照样被抽成「继续进行进一步的功能测试」这类废话待办, 再靠条数凑够周报阈值又触发一次 LLM 调用。门槛卡**转写正文**不卡纪要——内容太薄时 模型按模板兜底规则照样能写出六百多字。只拦抽取与 MCP 检索,**纪要照出**: 短语音备忘是合法用法,也不判成 SummarizFail(客户端把那个码显示成「请稍后重试」, 用户会一直点)。数据缺失时按通过处理,不因上游漏传一个字段就静默关掉抽取。 ## 其余 - 报告补偿:原先的 requeueIfStale 由 memory_getreport 驱动,而客户端调无参版本 → 走 latestUnconfirmedReport,那条 SQL 只查 state=Done,pending/processing/failed 永远查不出来也就永远补不了。新增 ensureRecent 挂在 memory_today:漏建的补建 (过与 cron 同一道条数闸门)、卡住的复用 requeueIfStale、失败的至多重试 3 次 (计数寄在 error_msg 的 retry=N; 前缀,fail() 必须保住它,否则上限形同虚设)。 加 enqueueOnce 判重与按报告 id 的 SETNX 生成锁,cron 与 ensure 撞上只跑一次 LLM。 - 说话人标签:finishSyncTranscribe 拼的是「Speaker 0」(空格)、其余是「Speaker_0」, 抽出来的负责人就成了 Speaker2、[Speaker_0] 各种样。四条写入路径统一走 normalizeSpeakers, 顺带修掉没开说话人分离时拼出「Speaker_」空尾巴。 - 时区:LoadMemoryLocation 现在认 IANA / 固定偏移 / CST(按 Asia/Shanghai 解释, 兜住已落库那批),认不出来仍退服务器本地时区但**打 warn**——原来是全静默的。 - 删纪要连带删自动待办(只删 user_edited=0),查询带 uid:入参来自删除请求, 光信 source_id 等于把归属校验的成果又丢一次。 - 回调幂等:两个回调既不看状态也不抢锁,与轮询撞车就各翻译一遍、各入队一次。 失败分支同样要过闸门——一条已 Completed 的记录被迟到的失败回调打回 TranscribeFail, 用户看到的就是「纪要好端端地变成了转写失败」。挡下来仍回 SUCCESS,不让第三方重投。 SubmitAITask 保持不判重(用户点的「重新生成」悄悄跳过等于按钮失灵), 自动路径改走 submitAITaskOnce。 - 重复项统计:只展开闹钟。花销展开等于虚构金额,待办的 done/undone 按次拆不开。 ## 顺带修的两个既有 bug(部署验收时挖出来的) happen_date / period_start 都是 type:date 列 + DSN parseTime=True,读进 Go 的 string 是「2026-09-07T00:00:00+08:00」。SQL 比较靠 MySQL 转换还是对的,所以这个问题 **在任何日志任何报错里都看不见**,只在 Go 侧解析它的地方发作: 1. remind.go 的 expandItem 解析不了 → **memory_upcoming 恒返回 0 个提醒时刻, 整个拾忆提醒在服务端一直空转**(真机实测:建一条每天 08:00 的闹钟,slots 为 []); 2. 下发给客户端的 happen_date 一直是 RFC3339 而非协议约定的 YYYY-MM-DD, 客户端 _localToday() 的字符串相等比较恒不匹配、_mergeRange 在区间端点会重复。 这也是审计里「会议待办会在 08:55 响通知」只能是推演的原因:实际一条都不会响, 两个 bug 互相遮蔽。修法两层:ParseMemoryDate 容忍带时间的形式(安全网), 读路径出口统一归一(管对外格式),缺任一层都会留坑。 ## 列表骨架化(同批) echomeet_getallrecords 只回列表要用的骨架字段,original/translate/summary 占每条 97% 的字节(实测 59 条 ≈ 500KB,列表用得上的 15KB)。 ⚠️ 骨架记录**绝不能写回库**:mysql.Save 是整行 UPDATE,三列会被清空。 StartTask 允许对已完成/已阅/失败的记录「重新转写」,那期间 state=Transcribing 而旧 纪要还在库里,用户切一下列表页旧转写和旧总结就没了。轮询前重新取整条再操作, TestListEndpointNeverSavesBriefRecord 守着不许回退。 Co-Authored-By: Claude Opus 5 (1M context) --- apps/services/comm/memory.go | 128 +++- apps/services/comm/memory_test.go | 116 ++++ apps/services/comm/memoryquery.go | 63 ++ apps/services/comm/module.go | 33 +- .../modules/echomeet/api_alibackcall.go | 26 +- .../services/modules/echomeet/api_backcall.go | 31 +- .../modules/echomeet/api_delrecords.go | 48 +- .../modules/echomeet/api_getallrecords.go | 55 +- .../modules/echomeet/api_getrecord.go | 11 +- .../modules/echomeet/api_getrecords.go | 7 +- .../modules/echomeet/api_modifyrecord.go | 12 +- .../modules/echomeet/api_readrecord.go | 12 +- apps/services/modules/echomeet/api_scope.go | 40 ++ .../modules/echomeet/api_scope_test.go | 238 +++++++ .../modules/echomeet/api_starttask.go | 18 +- apps/services/modules/echomeet/api_summary.go | 11 +- .../services/modules/echomeet/api_uprecord.go | 12 +- apps/services/modules/echomeet/model.go | 73 ++- apps/services/modules/echomeet/tasks.go | 133 +++- apps/services/modules/mcp/tool_meeting.go | 35 +- apps/services/modules/mcp/tool_memory.go | 35 +- apps/services/modules/memory/api_today.go | 5 + apps/services/modules/memory/core.go | 7 + .../modules/memory/meeting_extract.go | 56 +- .../modules/memory/meeting_extract_test.go | 69 ++- .../modules/memory/migrate_defaults.go | 53 ++ apps/services/modules/memory/model.go | 121 +++- .../modules/memory/model_default_test.go | 188 ++++++ apps/services/modules/memory/module.go | 6 + apps/services/modules/memory/options.go | 19 + apps/services/modules/memory/remind_test.go | 29 +- apps/services/modules/memory/report.go | 207 ++++++- apps/services/modules/memory/report_test.go | 120 ++++ apps/services/modules/memory/stats.go | 48 +- deploy/app/confs/home.yaml.example | 7 + ...�-拾忆-MCP链路审计与落地计划.md | 582 ++++++++++++++++++ 36 files changed, 2505 insertions(+), 149 deletions(-) create mode 100644 apps/services/modules/echomeet/api_scope.go create mode 100644 apps/services/modules/echomeet/api_scope_test.go create mode 100644 apps/services/modules/memory/migrate_defaults.go create mode 100644 apps/services/modules/memory/model_default_test.go create mode 100644 apps/services/modules/memory/report_test.go create mode 100644 docs/语音纪要-拾忆-MCP链路审计与落地计划.md diff --git a/apps/services/comm/memory.go b/apps/services/comm/memory.go index 198478e8..87dfcfce 100644 --- a/apps/services/comm/memory.go +++ b/apps/services/comm/memory.go @@ -4,8 +4,12 @@ import ( "crypto/rand" "encoding/hex" "fmt" + "regexp" + "strconv" "strings" "time" + + "yunyan/lego/sys/log" ) /* @@ -73,37 +77,141 @@ func IsValidMemoryCategory(c string) bool { return memoryCategories[c] } // IsValidMemoryRepeat 重复规则是否合法(空串等同 once) func IsValidMemoryRepeat(r string) bool { return r == "" || memoryRepeatRules[r] } +// MeetingMinSeconds 一段录音要多长才当「一场会议」看待。 +// +// 两处用它,必须同一个数:memory 的待办抽取门槛默认值(可由 home.yaml 覆盖), +// 和 MCP search_meeting_notes 的检索过滤(独立进程,读不到 memory 的配置, +// 所以那边就是这个常量本身)。 +// +// 拦的是 6 秒的测试通话、20 秒的单人自述这类东西:它们照样会被总结成六百多字的 +// 纪要(模板强制有「待办事项」一节,模型没真待办就把说话人的愿望改写成行动项), +// 抽进拾忆是垃圾待办,进 MCP 检索会把真会议挤出结果。 +// +// ⚠️ 只影响「进不进拾忆/检索」,不影响用户能不能看到纪要。短语音备忘是合法用法。 +const MeetingMinSeconds int32 = 60 + // MemoryDateLayout 记忆项日期的唯一格式。happen_date 是「用户本地日期」, // 不是 UTC 日期——服务端不做时区推断,一律以客户端上报的为准。 const MemoryDateLayout = "2006-01-02" -// ParseMemoryDate 解析 YYYY-MM-DD。失败返回零值与 false。 +// NormalizeMemoryDate 把「看起来像日期」的串收敛成 YYYY-MM-DD。 +// +// ⚠️ 为什么需要它:`memory_report.period_start/period_end` 是 `type:date` 列, +// 而 DSN 带 `parseTime=True`,于是从库里读回到 Go 的 **string** 字段时拿到的是 +// `2026-09-07T00:00:00+08:00` 而不是 `2026-09-07`。 +// 按 happen_date 比较的那几句 SQL 靠 MySQL 自己做类型转换还能对,但凡是要 +// **在 Go 里解析**这个值的地方都会静默失败——2026-09-18 真机上重复闹钟的展开 +// 就是这么一次都没生效的(ParseMemoryDate 解析不了,直接按「没有重复项」返回 0)。 +// +// 认不出来时原样返回:调用方该用它拼 SQL 还是照样拼,行为与加这个函数之前一致。 +func NormalizeMemoryDate(s string) string { + s = strings.TrimSpace(s) + if s == "" { + return s + } + if t, ok := ParseMemoryDate(s); ok { + return FormatMemoryDate(t) + } + return s +} + +// ParseMemoryDate 解析日期。主格式是 YYYY-MM-DD,**同时容忍带时间的形式**。 +// +// ⚠️ 容忍是必须的,不是宽松:`happen_date` / `period_start` 都是 MySQL 的 `type:date` 列, +// 而 DSN 带 `parseTime=True`,gorm 把它们扫进 Go 的 string 字段时给的是 +// `2026-09-07T00:00:00+08:00`。只认 YYYY-MM-DD 的话,**所有拿库里读出来的日期 +// 做解析的地方都会静默失效**: +// - remind.go 的 expandItem → memory_upcoming 恒返回 0 个提醒时刻, +// 整个拾忆提醒在服务端空转(2026-09-18 真机实测确认); +// - stats.go 的重复项展开 → 重复闹钟永远只算 1 次。 +// +// 两处都不报错,只是「什么都没发生」。 func ParseMemoryDate(s string) (time.Time, bool) { - t, err := time.Parse(MemoryDateLayout, strings.TrimSpace(s)) - if err != nil { - return time.Time{}, false + s = strings.TrimSpace(s) + if t, err := time.Parse(MemoryDateLayout, s); err == nil { + return t, true } - return t, true + // 带时间的形式(RFC3339 / "2006-01-02 15:04:05"):前 10 位就是日期。 + // ⚠️ 只取日期部分、丢掉时间与时区是刻意的——这个字段的语义是「用户本地的那一天」, + // 不是一个时刻,按时区换算反而会把日期挪一天。 + if len(s) > len(MemoryDateLayout) { + if t, err := time.Parse(MemoryDateLayout, s[:len(MemoryDateLayout)]); err == nil { + return t, true + } + } + return time.Time{}, false } // FormatMemoryDate 按统一格式输出 func FormatMemoryDate(t time.Time) string { return t.Format(MemoryDateLayout) } -// LoadMemoryLocation 解析客户端上报的 IANA 时区名。 +// LoadMemoryLocation 解析客户端上报的时区。 +// +// 认三种形式,按这个顺序: +// 1. IANA 名(`Asia/Shanghai`)——**这是唯一希望客户端上报的东西**; +// 2. 固定偏移(`UTC+8` / `GMT+08:00` / `+08:00`); +// 3. `CST` 这个缩写。 +// +// ⚠️ 为什么要认 `CST`:客户端早先上报的是 `DateTime.now().timeZoneName`,给出来的 +// 就是它,库里 tz 列至今只有空串和 `CST` 两个值。而 `CST` 是有名的歧义缩写 +// (中国标准时 UTC+8 / 美国中部时 UTC-6 / 古巴 / 澳洲中部都叫这个),Go 的 +// LoadLocation 根本解析不了。这里**按 Asia/Shanghai 解释**,理由是已落库的那批 +// 全部来自国内设备;新版客户端上报 IANA 名,不会再走到这条。 // // ⚠️ 认不出来一律回退到容器本地时区(Asia/Shanghai,见 Dockerfile 与 compose 的 TZ), // **绝不回退成 UTC**——那会让所有时刻整体偏移 8 小时,而且是静默的。 +// 但回退本身要打一条 warn:原来是完全静默的,海外用户的提醒时刻全按北京时间算 +// 也没有任何迹象。 +// // 容器已 import _ "time/tzdata",缺 tzdata 的镜像也能解析。 func LoadMemoryLocation(tz string) *time.Location { tz = strings.TrimSpace(tz) if tz == "" { return time.Local } - loc, err := time.LoadLocation(tz) - if err != nil { - return time.Local + if loc, err := time.LoadLocation(tz); err == nil { + return loc + } + if loc, ok := parseFixedZone(tz); ok { + return loc + } + if strings.EqualFold(tz, "CST") { + if loc, err := time.LoadLocation("Asia/Shanghai"); err == nil { + return loc + } + } + log.Warnf("[memory] 无法识别的时区 %q,已按服务器本地时区(%s)处理;"+ + "该用户的提醒时刻可能偏移——客户端应上报 IANA 名如 Asia/Shanghai", tz, time.Local) + return time.Local +} + +// fixedZonePattern 匹配 UTC+8 / GMT+08:00 / +08:00 / -0530 这几种固定偏移写法。 +var fixedZonePattern = regexp.MustCompile(`^(?i:UTC|GMT)?([+-])(\d{1,2}):?(\d{2})?$`) + +// parseFixedZone 把固定偏移写法转成 time.Location。 +// +// 这类值不该出现(时区 ≠ 偏移:同一个偏移夏令时前后是两个时区),但客户端兜底逻辑 +// 或第三方 SDK 给出这种串是常事,认下来总好过整条退回服务器时区。 +func parseFixedZone(s string) (*time.Location, bool) { + m := fixedZonePattern.FindStringSubmatch(strings.TrimSpace(s)) + if m == nil { + return nil, false + } + hour, err := strconv.Atoi(m[2]) + if err != nil || hour > 14 { + return nil, false + } + minute := 0 + if m[3] != "" { + if minute, err = strconv.Atoi(m[3]); err != nil || minute >= 60 { + return nil, false + } + } + offset := hour*3600 + minute*60 + if m[1] == "-" { + offset = -offset } - return loc + return time.FixedZone(strings.ToUpper(strings.TrimSpace(s)), offset), true } // ParseMemoryClock 解析 HH:mm。空串或非法返回 (0,0,false)。 diff --git a/apps/services/comm/memory_test.go b/apps/services/comm/memory_test.go index 19371c5c..5d48d613 100644 --- a/apps/services/comm/memory_test.go +++ b/apps/services/comm/memory_test.go @@ -130,3 +130,119 @@ func TestParseMemoryClock(t *testing.T) { t.Error("25:00 不该解析成功") } } + +// 时区解析:IANA 名是正道,固定偏移与 CST 是给存量数据兜底的。 +// ⚠️ 认不出来必须退到服务器本地时区,**不能退成 UTC**——那会让所有时刻静默偏 8 小时。 +func TestLoadMemoryLocation(t *testing.T) { + sh, err := time.LoadLocation("Asia/Shanghai") + if err != nil { + t.Skip("镜像里没有 tzdata,跳过") + } + // 用同一个瞬间比偏移量,而不是比 Location 的名字: + // FixedZone 造出来的 loc 名字是 "UTC+8" 之类,与 "Asia/Shanghai" 永远不相等, + // 但它们在这一刻的偏移是一样的,那才是我们真正关心的东西。 + at := time.Date(2026, 9, 18, 12, 0, 0, 0, time.UTC) + offsetOf := func(loc *time.Location) int { + _, off := at.In(loc).Zone() + return off + } + wantSh := offsetOf(sh) + + cases := []struct { + name string + tz string + want int + }{ + {"IANA 名", "Asia/Shanghai", wantSh}, + {"东京", "Asia/Tokyo", 9 * 3600}, + {"CST 缩写按上海解释(存量数据全是它)", "CST", wantSh}, + {"cst 大小写不敏感", "cst", wantSh}, + {"UTC+8", "UTC+8", 8 * 3600}, + {"GMT+08:00", "GMT+08:00", 8 * 3600}, + {"裸偏移 +08:00", "+08:00", 8 * 3600}, + {"负偏移 -0530", "-0530", -(5*3600 + 30*60)}, + {"空串 → 服务器本地", "", offsetOf(time.Local)}, + {"认不出来 → 服务器本地,不是 UTC", "Mars/Olympus", offsetOf(time.Local)}, + {"越界的偏移不认", "UTC+99", offsetOf(time.Local)}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if got := offsetOf(LoadMemoryLocation(c.tz)); got != c.want { + t.Errorf("%q 的偏移应为 %d 秒,得到 %d 秒", c.tz, c.want, got) + } + }) + } +} + +// 重复项展开:返回的是「除库里那一行之外还发生了几次」,不是总次数。 +// 数错的后果是周报里「这周被闹钟叫了多少次」要么恒为 0(不展开),要么多算一次(不扣锚点)。 +func TestMemoryExtraOccurrences(t *testing.T) { + d := func(s string) time.Time { + v, ok := ParseMemoryDate(s) + if !ok { + t.Fatalf("测试日期写错了: %s", s) + } + return v + } + // 2026-09-07(一) ~ 2026-09-13(日),典型的一个周报周期 + start, end := d("2026-09-07"), d("2026-09-13") + + cases := []struct { + name string + rule string + weekday int32 + anchor string + want int + }{ + {"一次性项不展开", MemoryRepeatOnce, 0, "2026-09-08", 0}, + {"空规则不展开", "", 0, "2026-09-08", 0}, + {"每天、锚点在几个月前:7 天全算,那一行没被数过", MemoryRepeatDaily, 0, "2026-06-01", 7}, + {"每天、锚点就在周期内:7 次里扣掉已数过的那一行", MemoryRepeatDaily, 0, "2026-09-07", 6}, + {"每天、锚点在周期中间:只从锚点当天算起", MemoryRepeatDaily, 0, "2026-09-11", 2}, + {"每周一、锚点在之前:本周命中 1 次且未被数过", MemoryRepeatWeekly, 1, "2026-06-01", 1}, + {"工作日、锚点在之前:5 次", MemoryRepeatWeekdays, 0, "2026-06-01", 5}, + {"周末、锚点在之前:2 次", MemoryRepeatWeekend, 0, "2026-06-01", 2}, + {"每月 8 号、锚点在之前:本周含 8 号,1 次", MemoryRepeatMonthly, 0, "2026-06-08", 1}, + {"每月 20 号:本周不含,0 次", MemoryRepeatMonthly, 0, "2026-06-20", 0}, + // ⚠️ 这条守着不能返回 -1:规则在本周期一次都没命中、但那一行恰好落在周期内时, + // 扣掉锚点会把调用方本来数对的那一行抹掉。 + {"锚点在周期内但规则本周不命中:返回 0 而不是 -1", MemoryRepeatWeekly, 1, "2026-09-09", 0}, + {"锚点在周期之后:一次都不算", MemoryRepeatDaily, 0, "2026-12-01", 0}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + got := MemoryExtraOccurrences(c.rule, c.weekday, d(c.anchor), start, end) + if got != c.want { + t.Errorf("期望 %d 次,得到 %d", c.want, got) + } + if got < 0 { + t.Error("返回负数会把调用方已经数对的那一行抹掉") + } + }) + } +} + +// period_start/period_end 是 type:date 列 + parseTime=True,读回 Go 的 string 字段 +// 是 RFC3339 而不是 YYYY-MM-DD。任何要在 Go 里解析这个值的地方不先归一就会静默失效 +// ——2026-09-18 真机上重复闹钟的展开就是这么一次都没生效的。 +func TestNormalizeMemoryDate(t *testing.T) { + cases := []struct{ in, want string }{ + {"2026-09-07", "2026-09-07"}, + {" 2026-09-07 ", "2026-09-07"}, + {"2026-09-07T00:00:00+08:00", "2026-09-07"}, // 库里读回来的真实形状 + {"2026-09-07T00:00:00Z", "2026-09-07"}, + {"2026-09-07 00:00:00", "2026-09-07"}, + {"", ""}, + {"不是日期", "不是日期"}, // 认不出来原样返回,行为与没有这个函数时一致 + {"09/07/2026", "09/07/2026"}, + } + for _, c := range cases { + if got := NormalizeMemoryDate(c.in); got != c.want { + t.Errorf("NormalizeMemoryDate(%q) = %q,期望 %q", c.in, got, c.want) + } + } + // 归一之后必须真的能解析出来——这才是加它的目的 + if _, ok := ParseMemoryDate(NormalizeMemoryDate("2026-09-07T00:00:00+08:00")); !ok { + t.Error("归一后仍解析不出日期,重复项展开会继续静默失效") + } +} diff --git a/apps/services/comm/memoryquery.go b/apps/services/comm/memoryquery.go index 5292bf5f..7541e6e1 100644 --- a/apps/services/comm/memoryquery.go +++ b/apps/services/comm/memoryquery.go @@ -94,6 +94,69 @@ func NormalizeMemoryCategories(in []string) ([]string, string) { return out, bad } +/* +重复项在周期统计里的展开。 + +一条 `daily` 闹钟在库里**只有一行**,happen_date 是当初定下它的那天。于是按 +`happen_date` 落在周期内来计数时,锚点在几个月前的重复闹钟在本周的 alarm_total +里是 0 —— 而用户这一周每天都被它叫醒过。 + +⚠️ **只展开闹钟**,刻意的: + - 花销展开等于**虚构金额**。「每月房租」展开进「你这周花了多少」,报出去的是 + 一个用户根本没记过的数字,而这份统计的全部价值就在数字是准的。 + - 待办的 done/undone 按次拆不开:一行只有一个 state,把一条「每周复盘」的 + 重复待办按 7 次都算成已完成,完成率就成了假的。 + - 灵感本来就不该重复。 + +⚠️ 放 comm 是因为 memory 的 computeStats 与 mcp 的 get_memory_stats 各查各的库 +(mcp 是独立进程),两边各写一份迟早漂成「App 里显示 7 次、EMAI 说 1 次」。 +*/ + +// MemoryExtraOccurrences 一条重复项在 [start,end] 里**除已计入的那一行之外**还发生了几次。 +// +// 返回的是增量而不是总次数:调用方的 group by 已经把锚点落在周期内的那一行数过一次了, +// 这里再返回总数就会重复计。锚点在周期之前的那种,那一行没被数到,返回的就是全部次数。 +func MemoryExtraOccurrences(rule string, weekday int32, anchor, start, end time.Time) int { + if rule == "" || rule == MemoryRepeatOnce { + return 0 + } + hits := 0 + for d := start; !d.After(end); d = d.AddDate(0, 0, 1) { + // 定下这条闹钟之前的日子不算数 + if d.Before(anchor) { + continue + } + if MemoryRepeatHits(rule, weekday, d, anchor) { + hits++ + } + } + // 锚点那一行已经被调用方数过一次,扣掉它。 + // hits==0 时不扣:那说明这条规则在本周期一次都没命中(比如周一的闹钟、 + // 周期里恰好没有周一),此时返回 -1 会把调用方本来数对的那一行抹掉。 + if hits > 0 && !anchor.Before(start) && !anchor.After(end) { + hits-- + } + return hits +} + +// ApplyRepeatCandidates 取「会重复、且锚点不晚于 end」的项——周期统计要展开的就是这些。 +// +// 锚点晚于 end 的是未来才开始的,本周期一次都不会响。 +func ApplyRepeatCandidates(tx *gorm.DB, uid, end string, categories []string) *gorm.DB { + if strings.TrimSpace(uid) == "" { + return tx.Where("1 = 0") + } + tx = tx.Where("uid = ?", uid). + Where("repeat_rule <> '' and repeat_rule <> ?", MemoryRepeatOnce) + if end != "" { + tx = tx.Where("happen_date <= ?", end) + } + if len(categories) > 0 { + tx = tx.Where("category in ?", categories) + } + return tx +} + // MemoryRelativeRange 把「最近 N 天 / 本周 / 本月」这类相对说法换算成日期区间。 // // 给 MCP 工具用:模型问「我这个月花了多少」时不该自己算日期, diff --git a/apps/services/comm/module.go b/apps/services/comm/module.go index 054f1625..1a13add9 100644 --- a/apps/services/comm/module.go +++ b/apps/services/comm/module.go @@ -150,5 +150,36 @@ type IMemory interface { // 「第几轮」由 memory 自己按 source_id 现有的最大 gen_round 推出来,不从这里传: // 调用方(echomeet)没有地方存轮次计数器,硬编码成 1 会让第二次「重新生成」 // 的清理条件 gen_round < 1 一条都匹配不上,旧待办从此永远留在库里。 - ExtractMeetingTodos(ctx context.Context, uid, recordID, summary, meetingDate, llmSvcID string) + ExtractMeetingTodos(ctx context.Context, in MeetingExtractInput) + + // DeleteMeetingItems 删掉这些会议记录**自动抽取**出来的待办,返回删掉几条。 + // + // 纪要被删之后待办还挂在日历上,点进去找不到来源。 + // + // ⚠️ 只删 user_edited=0 的:用户改过的待办已经是他自己的东西, + // 来源没了也该留着(与「重新生成时保留用户改动」同一条口径)。 + // recordIDs 是 echomeet 记录 id 的字符串形式(= memory_item.source_id)。 + DeleteMeetingItems(uid string, recordIDs []string) (int64, error) +} + +// MeetingExtractInput 一次会议待办抽取的输入。 +// +// 用结构体而不是一串位置参数:这里已经有 5 个字符串字段, +// 调用方把 summary 和 meetingDate 写反了编译器一句话都不会说。 +type MeetingExtractInput struct { + Uid string // 会议归属用户 + RecordID string // 会议记录 id(落库成 source_id) + Summary string // 已生成的纪要正文,抽取的输入 + MeetingDate string // 会议归属日期 YYYY-MM-DD,模型据它把「下周三」换算成具体日期 + LLMSvcID string // 与本条会议的总结同一个模型,便于出问题时归因 + + // ── 下面两个是内容门槛用的,见 memory 的 meeting_extract.go ── + + // Seconds 录音时长(秒)。 + Seconds int32 + // ContentRunes 转写正文的有效字数:已去掉 [Speaker_x] 标签,只数说的话。 + // + // ⚠️ 门槛必须卡正文不能卡纪要:内容太薄时模型按模板的兜底规则照样能写出 + // 六百多字的废话纪要(测试库 id=23/24/26 就是),拿纪要字数当门槛一条也拦不住。 + ContentRunes int } diff --git a/apps/services/modules/echomeet/api_alibackcall.go b/apps/services/modules/echomeet/api_alibackcall.go index 279e0e6f..bed6777c 100644 --- a/apps/services/modules/echomeet/api_alibackcall.go +++ b/apps/services/modules/echomeet/api_alibackcall.go @@ -4,7 +4,6 @@ import ( "yunyan/comm" "yunyan/pb" "yunyan/utils" - "fmt" ) // @Summary 处理阿里云录音文件转写回调 @@ -25,6 +24,16 @@ func (this *apiComp) AliBackCall(session comm.IUserSession, req *pb.EchomeetAliC return } + // 抢收尾权,理由同 BackCall:回调重放 / 与轮询撞车 / 迟到的失败回调把已完成的 + // 记录打回「转写失败」,三种都在这里挡掉。挡掉也回成功,不让阿里重投。 + if !this.module.tasks.beginTranscribeFinish(model) { + this.module.Infof("AliBackCall id:%d 状态已是 %s 或收尾进行中,忽略本次回调", + model.Id, model.State.String()) + resp = &pb.EchomeetAliCallbackResp{} + return + } + defer this.module.tasks.endTranscribeFinish(model.Id) + // 21050001 = 成功 if req.StatusCode != 21050001 { model.State = pb.DBEchoMeetRecordState_TranscribeFail @@ -33,31 +42,24 @@ func (this *apiComp) AliBackCall(session comm.IUserSession, req *pb.EchomeetAliC return } - personnels := make(map[string]struct{}) contexts := make([]*pb.ContextStruct, 0) if req.Result != nil { for _, s := range req.Result.Sentences { - speaker := fmt.Sprintf("Speaker_%s", s.SpeakerId) contexts = append(contexts, &pb.ContextStruct{ - Meetingid: model.Id, Content: s.Text, Starttime: s.BeginTime, Endtime: s.EndTime, - Speaker: speaker, + Speaker: s.SpeakerId, // 原始编号,格式化交给 normalizeSpeakers }) - personnels[s.SpeakerId] = struct{}{} } } - - model.Personnel = "" - for k := range personnels { - model.Personnel += fmt.Sprintf("[Speaker_%s] ", k) - } + // 同 BackCall:标签与 Personnel 统一走 normalizeSpeakers + normalizeSpeakers(model, contexts) model.Original = utils.ToString(contexts) model.State = pb.DBEchoMeetRecordState_AwaitSummarizing this.module.tasks.TranslateProcess(model, contexts) this.module.model.saverecord(model) - this.module.tasks.SubmitAITask(model) + _ = this.module.tasks.submitAITaskOnce(model) resp = &pb.EchomeetAliCallbackResp{} return diff --git a/apps/services/modules/echomeet/api_backcall.go b/apps/services/modules/echomeet/api_backcall.go index 4321d4cf..475ce455 100644 --- a/apps/services/modules/echomeet/api_backcall.go +++ b/apps/services/modules/echomeet/api_backcall.go @@ -4,7 +4,6 @@ import ( "yunyan/comm" "yunyan/pb" "yunyan/utils" - "fmt" ) // @Summary 处理Echomeet异步通知 [内部RPC] @@ -32,6 +31,20 @@ func (this *apiComp) BackCall(session comm.IUserSession, req *pb.EchomeetCallbac } return } + // 抢收尾权:回调重放、或轮询/兜底扫描已经收完之后迟到的这一次,都在这里被挡掉。 + // ⚠️ 挡掉时仍回 SUCCESS:对第三方来说这次通知已经被正确接收了, + // 回错只会让它按失败重投,重投的还是同一条已经处理过的结果。 + if !this.module.tasks.beginTranscribeFinish(model) { + this.module.Infof("BackCall id:%d 状态已是 %s 或收尾进行中,忽略本次回调", + model.Id, model.State.String()) + resp = []byte(utils.ToString(map[string]string{ + "code": "SUCCESS", + "message": "成功", + })) + return + } + defer this.module.tasks.endTranscribeFinish(model.Id) + if req.Code != 20000000 { model.State = pb.DBEchoMeetRecordState_TranscribeFail this.module.model.saverecord(model) @@ -41,27 +54,23 @@ func (this *apiComp) BackCall(session comm.IUserSession, req *pb.EchomeetCallbac })) return } - personnels := make(map[string]struct{}) - contexts := make([]*pb.ContextStruct, 0) + contexts := make([]*pb.ContextStruct, 0, len(req.Result.Utterances)) for _, v := range req.Result.Utterances { contexts = append(contexts, &pb.ContextStruct{ - Meetingid: model.Id, Content: v.Text, Starttime: v.StartTime, Endtime: v.EndTime, - Speaker: fmt.Sprintf("Speaker_%s", v.Additions.Speaker), + Speaker: v.Additions.Speaker, // 原始编号,格式化交给 normalizeSpeakers }) - personnels[v.Additions.Speaker] = struct{}{} - } - model.Personnel = "" - for k, _ := range personnels { - model.Personnel += fmt.Sprintf("[%s] ", fmt.Sprintf("Speaker_%s", k)) } + // 说话人标签与 Personnel 一律走这一个函数,四条写入路径才不会各拼各的。 + // 顺带修掉一个边角:没开说话人分离时编号是空串,原来会拼出 `Speaker_` 这种空尾巴。 + normalizeSpeakers(model, contexts) model.Original = utils.ToString(contexts) model.State = pb.DBEchoMeetRecordState_AwaitSummarizing this.module.tasks.TranslateProcess(model, contexts) this.module.model.saverecord(model) - this.module.tasks.SubmitAITask(model) + _ = this.module.tasks.submitAITaskOnce(model) resp = []byte(utils.ToString(map[string]string{ "code": "SUCCESS", "message": "成功", diff --git a/apps/services/modules/echomeet/api_delrecords.go b/apps/services/modules/echomeet/api_delrecords.go index ce678fb3..1c0c7f59 100644 --- a/apps/services/modules/echomeet/api_delrecords.go +++ b/apps/services/modules/echomeet/api_delrecords.go @@ -1,6 +1,8 @@ package echomeet import ( + "strconv" + "yunyan/comm" "yunyan/pb" ) @@ -16,10 +18,16 @@ import ( // @Router /api/home/echomeet_delrecords [post] func (this *apiComp) DelRecords(session comm.IUserSession, req *pb.EchomeetDelRecordsReq) (resp *pb.EchomeetDelRecordsResp, errdata *pb.ErrorData) { var ( - err error + deleted []uint64 + err error ) - - if err = this.module.model.delrecords(req.Ids); err != nil { + uid, errdata := requireUID(session) + if errdata != nil { + return + } + // 只删本人的。ids 里混进别人的(或已被别的设备删掉的)id 时按交集删、不报错, + // 响应里回**实际删掉的**那批,客户端据它清本地缓存。 + if deleted, err = this.module.model.delrecordsforuid(uid, req.Ids); err != nil { errdata = &pb.ErrorData{ Code: pb.ErrorCode_DBError, Message: err.Error(), @@ -27,8 +35,40 @@ func (this *apiComp) DelRecords(session comm.IUserSession, req *pb.EchomeetDelRe return } + // 连带清掉这些纪要自动抽出来的待办:来源都没了,待办还挂在日历上, + // 用户点进去找不到出处。用户改过的(user_edited=1)由 memory 那边保留。 + // + // ⚠️ 另起 goroutine 且失败只记日志:纪要已经删掉了,不能因为附带清理失败 + // 就对客户端报删除失败——那会让用户以为没删掉,再点一次。 + this.deleteMeetingTodos(uid, deleted) + resp = &pb.EchomeetDelRecordsResp{ - Ids: req.Ids, + Ids: deleted, } return } + +// deleteMeetingTodos 异步清理这些记录在拾忆里留下的自动待办。 +func (this *apiComp) deleteMeetingTodos(uid string, ids []uint64) { + if len(ids) == 0 { + return + } + mem := this.module.memoryModule() + if mem == nil { + return + } + recordIDs := make([]string, 0, len(ids)) + for _, id := range ids { + recordIDs = append(recordIDs, strconv.FormatUint(id, 10)) + } + go func() { + defer func() { + if r := recover(); r != nil { + this.module.Errorf("纪要删除连带清理待办 panic uid:%s err:%v", uid, r) + } + }() + if _, err := mem.DeleteMeetingItems(uid, recordIDs); err != nil { + this.module.Warnf("纪要删除连带清理待办失败已忽略 uid:%s records:%v err:%v", uid, recordIDs, err) + } + }() +} diff --git a/apps/services/modules/echomeet/api_getallrecords.go b/apps/services/modules/echomeet/api_getallrecords.go index 6fbcf407..b11f18c1 100644 --- a/apps/services/modules/echomeet/api_getallrecords.go +++ b/apps/services/modules/echomeet/api_getallrecords.go @@ -6,8 +6,9 @@ import ( "time" ) -// @Summary 获取用户全部会议记录 -// @Description 获取当前用户的全部会议记录 +// @Summary 获取用户全部会议记录(骨架) +// @Description 获取当前用户的全部会议记录。**只回列表要用的骨架字段**,original/translate/summary 不下发, +// @Description 详情由 echomeet_getrecord 单条拉取。 // @Tags Echomeet // @Accept json // @Produce json @@ -20,7 +21,14 @@ func (this *apiComp) GetAllRecords(session comm.IUserSession, req *pb.EchomeetGe records []*pb.DBEchoMeetRecord err error ) - if records, err = this.module.model.getrecordsforuid(session.GetUserId()); err != nil { + uid, errdata := requireUID(session) + if errdata != nil { + return + } + // 骨架查询:列表页只需要 title/rtype/state/seconds/audiourl/creationtime, + // 而 original/translate/summary 占了每条 97% 的字节(实测 59 条 ≈ 500KB,其中列表用得上的 15KB)。 + // 客户端本地 sqlite 是缓存、服务端是真源,详情进页面时再拉整条(echomeet_getrecord)。 + if records, err = this.module.model.getrecordbriefforuid(uid); err != nil { errdata = &pb.ErrorData{ Code: pb.ErrorCode_DBError, Message: err.Error(), @@ -28,20 +36,37 @@ func (this *apiComp) GetAllRecords(session comm.IUserSession, req *pb.EchomeetGe return } now := time.Now().Unix() - polled := make(map[uint64]bool) - for _, model := range records { - if model.State == pb.DBEchoMeetRecordState_Transcribing && model.Taskid != "" { - if model.Lastquerytime > 0 && now-model.Lastquerytime < TranscribeQueryMinInterval { - continue - } - model.Lastquerytime = now - this.module.model.saverecord(model) - this.module.tasks.PollTranscribe(model) - polled[model.Id] = true + for _, brief := range records { + if brief.State != pb.DBEchoMeetRecordState_Transcribing || brief.Taskid == "" { + continue + } + if brief.Lastquerytime > 0 && now-brief.Lastquerytime < TranscribeQueryMinInterval { + continue + } + // ⚠️ **绝不能拿 brief 去 saverecord / PollTranscribe**。 + // + // records 是 Omit 掉 original/translate/summary 的骨架,这三个字段在结构体里是空串; + // 而 mysql.Save 是**整行 UPDATE**,写回去就等于把这三列清空。 + // 触发条件不罕见:StartTask 的闸门是 `state>0 && state<5`,已完成(5)/已阅(6)/ + // 失败(10001/10002) 都允许「重新转写」,那期间 state=Transcribing 而旧的纪要 + // 还在库里 —— 用户这时候切一下语音纪要列表,旧转写和旧总结就没了; + // 如果这次转写又失败,那是永久丢失。 + // 列表改成骨架查询(2026-09-18)之前这里拿的是全字段,不会有这个问题。 + full, ferr := this.module.model.getrecordforuid(uid, brief.Id) + if ferr != nil { + this.module.Warnf("GetAllRecords id:%d 取整条失败,本轮不轮询转写: %v", brief.Id, ferr) + continue } + full.Lastquerytime = now + this.module.model.saverecord(full) + this.module.tasks.PollTranscribe(full) + // 让本次响应里的值也是新的(brief 才是要序列化出去的那个) + brief.Lastquerytime = now } - // 同 GetRecords,见 hideHalfDoneTranscribe 的注释 - hideHalfDoneTranscribe(polled, records) + // ⚠️ 这里刻意**不**调 hideHalfDoneTranscribe:它的判定是 `state 已完成 && Translate==""`, + // 而骨架查询根本不带 Translate,每一条都会被判成「半成品」降级回转写中。 + // 那个保护是给「客户端从列表响应里解 translate」加的,列表不再带 translate 就没有半成品可解; + // 轮询走的 GetRecords / GetRecord 仍是全字段,保护留在那两处。 resp = &pb.EchomeetGetAllRecordsResp{ Records: records, } diff --git a/apps/services/modules/echomeet/api_getrecord.go b/apps/services/modules/echomeet/api_getrecord.go index f6ee047d..25b1bbe5 100644 --- a/apps/services/modules/echomeet/api_getrecord.go +++ b/apps/services/modules/echomeet/api_getrecord.go @@ -16,12 +16,13 @@ import ( // @Success 200 {object} comm.HttpResult{data=pb.EchomeetGetRecordResp} "响应数据" // @Router /api/home/echomeet_getrecord [post] func (this *apiComp) GetRecord(session comm.IUserSession, req *pb.EchomeetGetRecordReq) (resp *pb.EchomeetGetRecordResp, errdata *pb.ErrorData) { - model, err := this.module.model.getrecord(req.Id) + uid, errdata := requireUID(session) + if errdata != nil { + return + } + model, err := this.module.model.getrecordforuid(uid, req.Id) if err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + errdata = recordErr(err) return } polled := make(map[uint64]bool) diff --git a/apps/services/modules/echomeet/api_getrecords.go b/apps/services/modules/echomeet/api_getrecords.go index 331133c4..75843f05 100644 --- a/apps/services/modules/echomeet/api_getrecords.go +++ b/apps/services/modules/echomeet/api_getrecords.go @@ -16,7 +16,12 @@ import ( // @Success 200 {object} comm.HttpResult{data=pb.EchomeetGetRecordsResp} "响应数据" // @Router /api/home/echomeet_getrecords [post] func (this *apiComp) GetRecords(session comm.IUserSession, req *pb.EchomeetGetRecordsReq) (resp *pb.EchomeetGetRecordsResp, errdata *pb.ErrorData) { - records, err := this.module.model.getrecords(req.Ids) + uid, errdata := requireUID(session) + if errdata != nil { + return + } + // 只取本人的:混进别人的 id 时静默过滤掉,不报错(见 getrecordsforuidids 注释) + records, err := this.module.model.getrecordsforuidids(uid, req.Ids) if err != nil { errdata = &pb.ErrorData{ Code: pb.ErrorCode_DBError, diff --git a/apps/services/modules/echomeet/api_modifyrecord.go b/apps/services/modules/echomeet/api_modifyrecord.go index dde3abd2..0da90e92 100644 --- a/apps/services/modules/echomeet/api_modifyrecord.go +++ b/apps/services/modules/echomeet/api_modifyrecord.go @@ -19,12 +19,12 @@ func (this *apiComp) ModifyRecords(session comm.IUserSession, req *pb.EchomeetMo record *pb.DBEchoMeetRecord err error ) - - if record, err = this.module.model.getrecord(req.Id); err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + uid, errdata := requireUID(session) + if errdata != nil { + return + } + if record, err = this.module.model.getrecordforuid(uid, req.Id); err != nil { + errdata = recordErr(err) return } if req.Title != "NULL" { diff --git a/apps/services/modules/echomeet/api_readrecord.go b/apps/services/modules/echomeet/api_readrecord.go index 7372e4d6..fda2cb8b 100644 --- a/apps/services/modules/echomeet/api_readrecord.go +++ b/apps/services/modules/echomeet/api_readrecord.go @@ -19,12 +19,12 @@ func (this *apiComp) ReadRecord(session comm.IUserSession, req *pb.EchomeetReadR record *pb.DBEchoMeetRecord err error ) - - if record, err = this.module.model.getrecord(req.Id); err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + uid, errdata := requireUID(session) + if errdata != nil { + return + } + if record, err = this.module.model.getrecordforuid(uid, req.Id); err != nil { + errdata = recordErr(err) return } diff --git a/apps/services/modules/echomeet/api_scope.go b/apps/services/modules/echomeet/api_scope.go new file mode 100644 index 00000000..81b68b85 --- /dev/null +++ b/apps/services/modules/echomeet/api_scope.go @@ -0,0 +1,40 @@ +package echomeet + +import ( + "yunyan/comm" + "yunyan/pb" +) + +/* +按 id 操作会议记录的接口,统一在这里取会话 uid 并做归属校验。 + +在此之前 getrecord / getrecords / modifyrecords / delrecords / readrecord / +summary / starttask / uprecord 八个接口**全部只按 id 查库**,没有一处比对 +record.Uid 与会话 uid(starttask 用了 GetUserId(),但只拿去扣算力和查 VIP)。 +后果是任何登录用户改一个 id 就能读别人纪要全文、删改别人的记录、 +给别人的记录发起总结(扣自己的算力,但覆盖对方 summary 并触发对方的拾忆抽取)。 + +写法照 memory 模块(每个接口 requireUID + rec.Uid != uid)。 +回调接口 BackCall / AliBackCall 是第三方来的、没有 session,不在此列。 +*/ + +// requireUID 取会话 uid。所有按 id 操作记录的接口第一步。 +func requireUID(session comm.IUserSession) (string, *pb.ErrorData) { + uid := session.GetUserId() + if uid == "" { + return "", &pb.ErrorData{Code: pb.ErrorCode_NoLogin, Message: pb.ErrorCode_NoLogin.String()} + } + return uid, nil +} + +// recordErr 把 model 层的查询错误转成对外错误。 +// +// ⚠️ 「不存在」与「不是你的」回同一个错误码同一句话:分开回等于给出一个 +// 探测他人记录 id 的接口。DBError 也正是改动前记录真的不存在时客户端拿到的码, +// 这样合法用户看到的行为没有任何变化。 +func recordErr(err error) *pb.ErrorData { + if err == ErrRecordNotFound { + return &pb.ErrorData{Code: pb.ErrorCode_DBError, Message: ErrRecordNotFound.Error()} + } + return &pb.ErrorData{Code: pb.ErrorCode_DBError, Message: err.Error()} +} diff --git a/apps/services/modules/echomeet/api_scope_test.go b/apps/services/modules/echomeet/api_scope_test.go new file mode 100644 index 00000000..eeb99abc --- /dev/null +++ b/apps/services/modules/echomeet/api_scope_test.go @@ -0,0 +1,238 @@ +package echomeet + +import ( + "os" + "path/filepath" + "regexp" + "strings" + "testing" + + "yunyan/pb" +) + +// TestApisUseUidScopedQueries 守住归属校验:接口层不许再调不带 uid 的记录查询。 +// +// 2026-09-18 之前八个按 id 操作的接口全部只按 id 查库,任何登录用户改一个 id +// 就能读改删别人的纪要。修完之后最容易复发的方式不是有人改回去,而是**新加一个 +// api_xxx.go 时照着旧写法抄一遍** —— 那样既不报错也没人看得出来,所以这里按源码扫。 +// +// 不带 uid 的版本仍然合法,但只能由回调(第三方来的,没有 session)和 +// 后台兜底任务(按 state/taskid 扫全表)调用,它们都不在 api_*.go 里。 +func TestApisUseUidScopedQueries(t *testing.T) { + // 回调没有会话,天然不在此列 + exempt := map[string]bool{ + "api_backcall.go": true, + "api_alibackcall.go": true, + } + banned := regexp.MustCompile(`model\.(getrecord|getrecords|delrecords)\(`) + + files, err := filepath.Glob("api_*.go") + if err != nil { + t.Fatal(err) + } + if len(files) < 8 { + t.Fatalf("只扫到 %d 个 api 文件,glob 大概写错了", len(files)) + } + for _, f := range files { + if exempt[filepath.Base(f)] || strings.HasSuffix(f, "_test.go") { + continue + } + src, err := os.ReadFile(f) + if err != nil { + t.Fatal(err) + } + for i, line := range strings.Split(string(src), "\n") { + // 跳过注释行:api_putrecords.go 整个函数都是注释掉的死代码 + // (客户端那个调用点 2026-09-18 已删),它不是活的调用。 + if strings.HasPrefix(strings.TrimSpace(line), "//") { + continue + } + if banned.MatchString(line) { + t.Errorf("%s:%d 调了不带 uid 的记录查询,别人的记录也查得到:%s\n"+ + " → 改用 getrecordforuid / getrecordsforuidids / delrecordsforuid(见 api_scope.go)", + f, i+1, strings.TrimSpace(line)) + } + } + } +} + +// TestTranscriptRunes 内容门槛数的是「说的话」,不是转写 JSON 的长度, +// 更不能把 [Speaker_x] 标签算进去——两人对话光标签就有几十字, +// 算进去的话再薄的内容也能轻松过门槛。 +func TestTranscriptRunes(t *testing.T) { + cases := []struct { + name string + rec *pb.DBEchoMeetRecord + want int + }{ + { + name: "优先用译文", + rec: &pb.DBEchoMeetRecord{ + Original: `[{"speaker":"Speaker_0","content":"一二三四五"}]`, + Translate: `[{"speaker":"Speaker_0","content":"一二三"}]`, + }, + want: 3, + }, + { + name: "译文为空退回原文(与 AIProcess 同口径)", + rec: &pb.DBEchoMeetRecord{ + Original: `[{"speaker":"Speaker_0","content":"一二三四五"}]`, + Translate: " ", + }, + want: 5, + }, + { + name: "多段累加,说话人标签不计入", + rec: &pb.DBEchoMeetRecord{ + Translate: `[{"speaker":"Speaker_0","content":"你好"},{"speaker":"Speaker_1","content":"在的"}]`, + }, + want: 4, + }, + { + name: "provider 直接给纯文本时按整串算", + rec: &pb.DBEchoMeetRecord{Translate: "就这么一句话"}, + want: 6, + }, + { + name: "两边都空 → 0(调用方按「数据缺失」处理,不拦)", + rec: &pb.DBEchoMeetRecord{}, + want: 0, + }, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if got := transcriptRunes(c.rec); got != c.want { + t.Errorf("期望 %d 字,得到 %d", c.want, got) + } + }) + } +} + +// 说话人标签四条写入路径必须同一种格式 `Speaker_N`。 +// +// 此前 finishSyncTranscribe 拼的是 `Speaker 0`(空格),后果不在纪要本身而在下游: +// LLM 抽出来的负责人会变成 `Speaker2`、`[Speaker_0]` 各种样, +// MCP 的 meaningfulPersonnel 只认 `Speaker_` 前缀,别的会被当成真人名塞给模型。 +func TestNormalizeSpeakers(t *testing.T) { + t.Run("裸编号补上 Speaker_ 前缀,Personnel 跟着拼", func(t *testing.T) { + rec := &pb.DBEchoMeetRecord{Id: 7} + ctxs := []*pb.ContextStruct{ + {Speaker: "0", Content: "你好"}, + {Speaker: "1", Content: "在的"}, + {Speaker: "0", Content: "那就这样"}, + } + normalizeSpeakers(rec, ctxs) + for _, c := range ctxs { + if !strings.HasPrefix(c.Speaker, "Speaker_") { + t.Errorf("说话人 %q 没有统一成 Speaker_N", c.Speaker) + } + if c.Meetingid != 7 { + t.Errorf("meetingid 没回填,得到 %d", c.Meetingid) + } + } + // map 迭代无序,只能按包含判断 + for _, want := range []string{"[Speaker_0]", "[Speaker_1]"} { + if !strings.Contains(rec.Personnel, want) { + t.Errorf("Personnel %q 里缺 %s", rec.Personnel, want) + } + } + }) + + t.Run("已带 Speaker 前缀的不动(含历史的空格写法)", func(t *testing.T) { + rec := &pb.DBEchoMeetRecord{Id: 1} + ctxs := []*pb.ContextStruct{ + {Speaker: "Speaker_0"}, + {Speaker: "Speaker 3"}, // 存量数据,重跑时原样保留,免得同一条会议里新旧对不上 + } + normalizeSpeakers(rec, ctxs) + if ctxs[0].Speaker != "Speaker_0" || ctxs[1].Speaker != "Speaker 3" { + t.Errorf("已带前缀的被改写了:%q / %q", ctxs[0].Speaker, ctxs[1].Speaker) + } + }) + + t.Run("没开说话人分离时 Speaker 是空串,不该拼成 Speaker_", func(t *testing.T) { + rec := &pb.DBEchoMeetRecord{Id: 1} + ctxs := []*pb.ContextStruct{{Speaker: ""}} + normalizeSpeakers(rec, ctxs) + if ctxs[0].Speaker != "" { + t.Errorf("空说话人被改成了 %q", ctxs[0].Speaker) + } + }) +} + +// 列表接口拿到的是 Omit 掉 original/translate/summary 的**骨架**记录, +// 它绝不能被写回库:mysql.Save 是整行 UPDATE,写回去就是把这三列清空。 +// +// 真实触发路径:StartTask 允许对已完成/已阅/失败的记录「重新转写」, +// 那期间 state=Transcribing 而旧纪要还在库里,用户切一下列表页就没了。 +// 要轮询转写必须先 getrecord 取整条(见 api_getallrecords.go)。 +func TestListEndpointNeverSavesBriefRecord(t *testing.T) { + src, err := os.ReadFile("api_getallrecords.go") + if err != nil { + t.Fatal(err) + } + for _, bad := range []string{"saverecord(brief", "PollTranscribe(brief"} { + if strings.Contains(string(src), bad) { + t.Errorf("api_getallrecords.go 里出现 %q:骨架记录被写回库会清空 "+ + "original/translate/summary,必须先 getrecord 取整条", bad) + } + } + // 反过来也确认一下整条那步还在,别哪天连轮询一起删了 + if !strings.Contains(string(src), "this.module.model.getrecordforuid(uid, brief.Id)") { + t.Error("api_getallrecords.go 不再按 id 取整条记录,转写轮询要么没了、要么又拿骨架在跑") + } +} + +// 转写收尾的抢锁:回调重放、与轮询撞车、迟到的失败回调,三种都必须被挡下来。 +// +// 挡不住的后果不是报错而是「安静地多跑一轮」:整篇重译一遍、AI 队列里多一个 id、 +// AIProcess 跑两次(两次 LLM 计费,后一轮覆盖前一轮,拾忆的待办也抽两轮)。 +func TestBeginTranscribeFinish(t *testing.T) { + newTasks := func() *tasksComp { return &tasksComp{} } + + t.Run("转写中:第一次拿得到,第二次拿不到", func(t *testing.T) { + tc := newTasks() + rec := &pb.DBEchoMeetRecord{Id: 1, State: pb.DBEchoMeetRecordState_Transcribing} + if !tc.beginTranscribeFinish(rec) { + t.Fatal("第一次应该拿得到收尾权") + } + if tc.beginTranscribeFinish(rec) { + t.Error("收尾进行中还能再拿到,回调与轮询会各收尾一次") + } + tc.endTranscribeFinish(rec.Id) + if !tc.beginTranscribeFinish(rec) { + t.Error("收尾结束后应该能再次进入(兜底扫描要用)") + } + }) + + t.Run("等待转写:放行(提交成功到状态落库之间有窗口)", func(t *testing.T) { + tc := newTasks() + rec := &pb.DBEchoMeetRecord{Id: 2, State: pb.DBEchoMeetRecordState_AwaitTranscribing} + if !tc.beginTranscribeFinish(rec) { + t.Error("快的 provider 可能在状态写成 Transcribing 之前就回调,必须放行") + } + }) + + for _, st := range []pb.DBEchoMeetRecordState{ + pb.DBEchoMeetRecordState_AwaitSummarizing, + pb.DBEchoMeetRecordState_Summarizing, + pb.DBEchoMeetRecordState_Completed, + pb.DBEchoMeetRecordState_Readed, + pb.DBEchoMeetRecordState_TranscribeFail, + pb.DBEchoMeetRecordState_SummarizFail, + } { + t.Run("已经不在转写阶段就不再收尾:"+st.String(), func(t *testing.T) { + tc := newTasks() + rec := &pb.DBEchoMeetRecord{Id: 3, State: st} + if tc.beginTranscribeFinish(rec) { + t.Errorf("state=%s 仍被允许收尾:一条已完成的纪要会被迟到的失败回调打回转写失败", st) + } + }) + } + + t.Run("nil 不 panic", func(t *testing.T) { + if newTasks().beginTranscribeFinish(nil) { + t.Error("nil 不该拿到收尾权") + } + }) +} diff --git a/apps/services/modules/echomeet/api_starttask.go b/apps/services/modules/echomeet/api_starttask.go index 8fc126e3..e15bb537 100644 --- a/apps/services/modules/echomeet/api_starttask.go +++ b/apps/services/modules/echomeet/api_starttask.go @@ -24,7 +24,11 @@ func (this *apiComp) StartTask(session comm.IUserSession, req *pb.EchomeetStartT err error ) - user, err = this.module.model.getuser(session.GetUserId()) + uid, errdata := requireUID(session) + if errdata != nil { + return + } + user, err = this.module.model.getuser(uid) if err != nil { errdata = &pb.ErrorData{ Code: pb.ErrorCode_DBError, @@ -32,12 +36,10 @@ func (this *apiComp) StartTask(session comm.IUserSession, req *pb.EchomeetStartT } return } - model, err = this.module.model.getrecord(req.Id) + // ⚠️ 归属校验必须排在扣算力之前,否则拒绝了还先把钱扣了。 + model, err = this.module.model.getrecordforuid(uid, req.Id) if err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + errdata = recordErr(err) return } if model.State > pb.DBEchoMeetRecordState_Unknow && model.State < pb.DBEchoMeetRecordState_Completed { @@ -57,7 +59,7 @@ func (this *apiComp) StartTask(session comm.IUserSession, req *pb.EchomeetStartT } // ── 计量:会议时长服务端自己知道(model.Seconds),按换算系数折成算力记账。 // 先设备赠送、再用户余额;闸门默认关,不够也放行、差额记超额。── - statistics, sErr := this.module.model.getStatistics(session.GetUserId()) + statistics, sErr := this.module.model.getStatistics(uid) if sErr != nil && sErr != mysql.ErrNoDocuments { errdata = &pb.ErrorData{Code: pb.ErrorCode_DBError, Message: sErr.Error()} return @@ -65,7 +67,7 @@ func (this *apiComp) StartTask(session comm.IUserSession, req *pb.EchomeetStartT statistics.Meetnum += 1 statistics.Meettime += int64(model.Seconds) userlog := &pb.DBUserUseLog{ - Uid: session.GetUserId(), + Uid: uid, Logtype: pb.UserLogType_UserConsume, Ts: time.Now().Unix(), Addmeetsecond: -1 * int64(model.Seconds), diff --git a/apps/services/modules/echomeet/api_summary.go b/apps/services/modules/echomeet/api_summary.go index adbf2f16..aa126329 100644 --- a/apps/services/modules/echomeet/api_summary.go +++ b/apps/services/modules/echomeet/api_summary.go @@ -20,11 +20,12 @@ func (this *apiComp) Summary(session comm.IUserSession, req *pb.EchomeetSummaryR model *pb.DBEchoMeetRecord err error ) - if model, err = this.module.model.getrecord(req.Id); err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + uid, errdata := requireUID(session) + if errdata != nil { + return + } + if model, err = this.module.model.getrecordforuid(uid, req.Id); err != nil { + errdata = recordErr(err) return } // 确定目标语言 diff --git a/apps/services/modules/echomeet/api_uprecord.go b/apps/services/modules/echomeet/api_uprecord.go index e080a4fe..958ce09d 100644 --- a/apps/services/modules/echomeet/api_uprecord.go +++ b/apps/services/modules/echomeet/api_uprecord.go @@ -19,12 +19,12 @@ func (this *apiComp) UpRecord(session comm.IUserSession, req *pb.EchomeetUpRecor model *pb.DBEchoMeetRecord err error ) - - if model, err = this.module.model.getrecord(req.Id); err != nil { - errdata = &pb.ErrorData{ - Code: pb.ErrorCode_DBError, - Message: err.Error(), - } + uid, errdata := requireUID(session) + if errdata != nil { + return + } + if model, err = this.module.model.getrecordforuid(uid, req.Id); err != nil { + errdata = recordErr(err) return } model.Audiourl = req.Audiourl diff --git a/apps/services/modules/echomeet/model.go b/apps/services/modules/echomeet/model.go index 94d65e2e..b0adf2ee 100644 --- a/apps/services/modules/echomeet/model.go +++ b/apps/services/modules/echomeet/model.go @@ -1,6 +1,7 @@ package echomeet import ( + "errors" "fmt" "path/filepath" "runtime" @@ -222,26 +223,84 @@ func (this *modelComp) deltemplates(id []uint64) (err error) { return } -func (this *modelComp) getrecordsforuid(uid string) (models []*pb.DBEchoMeetRecord, err error) { - models = make([]*pb.DBEchoMeetRecord, 0) - err = mysql.Find(comm.TableEchomeetRecord, &models, "uid=?", uid) - return -} - -// 查询用户记录的精简信息(不含 original、translate、summary 大字段) +// 查询用户记录的精简信息(不含 original、translate、summary 大字段),新到旧。 +// 列表接口 GetAllRecords 用;Omit 掉的字段序列化后是空串不是缺键,客户端不会解析失败。 func (this *modelComp) getrecordbriefforuid(uid string) (models []*pb.DBEchoMeetRecord, err error) { models = make([]*pb.DBEchoMeetRecord, 0) err = mysql.Table(comm.TableEchomeetRecord). Omit("original", "translate", "summary"). Where("uid = ?", uid). + Order("creationtime DESC, id DESC"). Find(&models).Error return } + +// ⚠️ 不带 uid 的 getrecord / getrecords / delrecords **只给回调与后台兜底任务用** +// (第三方 ASR 回调没有 session;sweepStuck* 按 state 扫全表)。 +// 凡是由客户端按 id 指定记录的接口,一律走下面三个带 uid 的版本 —— +// 不然任何登录用户改一个 id 就能读改删别人的纪要。 func (this *modelComp) getrecord(id uint64) (model *pb.DBEchoMeetRecord, err error) { model = &pb.DBEchoMeetRecord{} err = mysql.FindOne(comm.TableEchomeetRecord, model, "id=?", id) return } + +// ErrRecordNotFound 记录不存在,**或者不属于当前用户**。 +// +// 两种情况刻意合并成同一个错误:区分开就等于对外提供了一个「这个 id 存在吗」的探测接口, +// 攻击者拿它可以枚举出全站有多少条纪要、哪些 id 是活的。 +var ErrRecordNotFound = errors.New("记录不存在") + +// getrecordforuid 按 (uid, id) 取单条。查不到统一回 ErrRecordNotFound。 +func (this *modelComp) getrecordforuid(uid string, id uint64) (model *pb.DBEchoMeetRecord, err error) { + model = &pb.DBEchoMeetRecord{} + if err = mysql.FindOne(comm.TableEchomeetRecord, model, "uid=? AND id=?", uid, id); err != nil { + if errors.Is(err, mysql.ErrNoDocuments) { + err = ErrRecordNotFound + } + return nil, err + } + return +} + +// getrecordsforuidids 批量取本人的记录。 +// +// ⚠️ ids 里混进别人的(或已删的)id 时**只返回自己那部分、不报错**。 +// 客户端的轮询队列里本来就可能留着已被别的设备删掉的 id,这里一报错 +// 就会把它整个轮询循环掀掉再也起不来(MeetingTaskService._executeTask 那个坑)。 +func (this *modelComp) getrecordsforuidids(uid string, ids []uint64) (models []*pb.DBEchoMeetRecord, err error) { + models = make([]*pb.DBEchoMeetRecord, 0) + if len(ids) == 0 { + return + } + err = mysql.Find(comm.TableEchomeetRecord, &models, "uid=? AND id IN ?", uid, ids) + return +} + +// ownedrecordids 从 ids 里挑出属于 uid 的那些。只取主键,不拉大字段。 +func (this *modelComp) ownedrecordids(uid string, ids []uint64) (owned []uint64, err error) { + owned = make([]uint64, 0, len(ids)) + if len(ids) == 0 { + return + } + err = mysql.Table(comm.TableEchomeetRecord). + Where("uid = ? AND id IN ?", uid, ids). + Pluck("id", &owned).Error + return +} + +// delrecordsforuid 删本人的记录,返回**实际删掉的** id。 +// +// 先查交集再删:客户端据响应里的 Ids 清本地缓存,而 DELETE 只给得出行数、给不出 id。 +func (this *modelComp) delrecordsforuid(uid string, ids []uint64) (deleted []uint64, err error) { + if deleted, err = this.ownedrecordids(uid, ids); err != nil || len(deleted) == 0 { + return + } + if err = mysql.Delete(comm.TableEchomeetRecord, "uid=? AND id IN ?", uid, deleted); err != nil { + return nil, err + } + return +} func (this *modelComp) getrecordByTaskId(taskId string) (model *pb.DBEchoMeetRecord, err error) { model = &pb.DBEchoMeetRecord{} err = mysql.FindOne(comm.TableEchomeetRecord, model, "taskid=?", taskId) diff --git a/apps/services/modules/echomeet/tasks.go b/apps/services/modules/echomeet/tasks.go index e05e7e5b..9c233a40 100644 --- a/apps/services/modules/echomeet/tasks.go +++ b/apps/services/modules/echomeet/tasks.go @@ -321,6 +321,9 @@ func (this *tasksComp) inQueue(key, member string) bool { return err == nil } +// SubmitAITask 入队做 AI 总结。**不判重**——用户主动发起的路径用它 +// (StartTask / Summary 的「重新生成」):用户点了按钮就必须真的跑一轮, +// 悄悄跳过等于按钮失灵,那比多跑一次糟得多。 func (this *tasksComp) SubmitAITask(task *pb.DBEchoMeetRecord) (err error) { this.module.Infof("SubmitAITask id:%d uid:%s → 进入AI等待队列", task.Id, task.Uid) task.State = pb.DBEchoMeetRecordState_AwaitSummarizing @@ -329,6 +332,19 @@ func (this *tasksComp) SubmitAITask(task *pb.DBEchoMeetRecord) (err error) { return } +// submitAITaskOnce 已经在 AI 队列里(等待中或处理中)就不再入队。 +// +// 自动路径(转写收尾、两个回调)用它:这些路径可能被触发多次,每多入一次队 +// 就是多一次 LLM 计费、多一轮待办抽取,而后跑的那轮还会覆盖先跑的结果。 +func (this *tasksComp) submitAITaskOnce(task *pb.DBEchoMeetRecord) (err error) { + idStr := fmt.Sprintf("%d", task.Id) + if this.inQueue(this.keyAIAwait(), idStr) || this.inQueue(this.keyAIProc(), idStr) { + this.module.Infof("SubmitAITask id:%d 已在AI队列里,跳过重复入队", task.Id) + return + } + return this.SubmitAITask(task) +} + // 短音频处理流 // 参数: // - task: 任务对象,类型为 *schedItem,包含记录ID与执行列表键 @@ -408,14 +424,24 @@ func (this *tasksComp) ShortAudioProcess(task interface{}) { _ = redissys.Conn().LRem(context.Background(), procKey, 1, fmt.Sprintf("%d", rec.Id)).Err() } -// finishSyncTranscribe 同步转写完成(字节 flash 等 Done=true 通道)的收尾: -// 补齐说话人/原文 → 翻译 → 落库 → 提交 AI 总结。 -func (this *tasksComp) finishSyncTranscribe(rec *pb.DBEchoMeetRecord, contexts []*pb.ContextStruct) { +// normalizeSpeakers 把 provider 给的说话人编号统一成 `Speaker_N`,并拼出 Personnel。 +// +// ⚠️ 三条写入路径(阿里回调 / 字节回调 / 轮询 / 字节 flash 同步通道)必须用同一种格式。 +// 此前 finishSyncTranscribe 拼的是 `Speaker 0`(空格),其余是 `Speaker_0`, +// 客户端归档进来的甚至是裸 `0`。后果不在纪要本身,而在下游: +// - LLM 拿到 `[0]:`、`[Speaker 2]`、`[Speaker_1]` 混着,抽出来的 owner 就是 +// `Speaker2`、`[Speaker_0]` 各种样(测试库 memory_item id=25 / id=12); +// - MCP 的 meaningfulPersonnel 只认 `Speaker_` 前缀,别的格式会被当成 +// 「用户命名过的真人」塞给模型。 +// +// 已经带 `Speaker` 前缀的不动(包括历史数据里的 `Speaker 0`):那是重跑时的输入, +// 改写它只会让同一条会议里新旧两段对不上。存量由客户端的说话人重命名功能收拾。 +func normalizeSpeakers(rec *pb.DBEchoMeetRecord, contexts []*pb.ContextStruct) { personnels := make(map[string]struct{}) for _, c := range contexts { c.Meetingid = rec.Id if c.Speaker != "" && !strings.HasPrefix(c.Speaker, "Speaker") { - c.Speaker = fmt.Sprintf("Speaker %s", c.Speaker) + c.Speaker = "Speaker_" + c.Speaker } personnels[c.Speaker] = struct{}{} } @@ -423,11 +449,57 @@ func (this *tasksComp) finishSyncTranscribe(rec *pb.DBEchoMeetRecord, contexts [ for k := range personnels { rec.Personnel += fmt.Sprintf("[%s] ", k) } +} + +/* +beginTranscribeFinish / endTranscribeFinish 抢「这条记录的转写收尾」的处理权。 + +收尾现在有四个驱动方:第三方回调、客户端轮询(详情页)、**打开语音纪要列表** +(2026-09-18 起列表页也会轮询转写中的记录)、cron 兜底扫描。谁先拿到成功结果谁收尾, +其余的必须让开 —— 否则同一条记录会被翻译两遍、往 AI 队列里推两次, +AIProcess 跑两轮(两次 LLM 计费,后一轮覆盖前一轮,连带把拾忆的待办也抽两轮)。 + +轮询那条路原先就有 finishing 这把内存锁,**回调这条路一直没有**: +AliBackCall / BackCall 拿到结果直接 TranslateProcess + SubmitAITask, +既不看状态也不抢锁。转写文件越长(回调越晚、轮询次数越多)越容易撞上。 + +两道判断缺一不可: + - 状态:只有还在等待/进行转写的记录才需要收尾。回调重放、或者轮询已经收完之后 + 迟到的那次回调,都会被这一条挡掉。**失败回调同样要过这一关** —— + 一条已经 Completed 的记录被一个迟到的失败回调打回 TranscribeFail, + 用户那边就是「纪要好端端地变成了转写失败」。 + - 内存锁:状态落库有个时间差,两条路同时读到 Transcribing 时靠它分出先后。 + +⚠️ 只在单进程内有效。当前 app 是单副本部署;多副本时要换成 Redis 锁, +届时连轮询那把一起换。 +*/ +func (this *tasksComp) beginTranscribeFinish(rec *pb.DBEchoMeetRecord) bool { + if rec == nil { + return false + } + // AwaitTranscribing 也要放行:提交成功到把状态写成 Transcribing 之间有个窗口, + // 快的 provider 可以在这中间就把回调打回来。 + if rec.State != pb.DBEchoMeetRecordState_Transcribing && + rec.State != pb.DBEchoMeetRecordState_AwaitTranscribing { + return false + } + _, busy := this.finishing.LoadOrStore(rec.Id, struct{}{}) + return !busy +} + +func (this *tasksComp) endTranscribeFinish(id uint64) { + this.finishing.Delete(id) +} + +// finishSyncTranscribe 同步转写完成(字节 flash 等 Done=true 通道)的收尾: +// 补齐说话人/原文 → 翻译 → 落库 → 提交 AI 总结。 +func (this *tasksComp) finishSyncTranscribe(rec *pb.DBEchoMeetRecord, contexts []*pb.ContextStruct) { + normalizeSpeakers(rec, contexts) rec.Original = utils.ToString(contexts) rec.State = pb.DBEchoMeetRecordState_AwaitSummarizing this.TranslateProcess(rec, contexts) this.module.model.saverecord(rec) - _ = this.SubmitAITask(rec) + _ = this.submitAITaskOnce(rec) } // 翻译处理流 @@ -658,8 +730,40 @@ func (this *tasksComp) extractMemoryTodos(rec *pb.DBEchoMeetRecord) { ts = time.Now().Unix() } meetingDate := comm.FormatMemoryDate(time.Unix(ts, 0)) - go mem.ExtractMeetingTodos(context.Background(), rec.Uid, fmt.Sprintf("%d", rec.Id), - rec.Summary, meetingDate, rec.LlmSvcId) + go mem.ExtractMeetingTodos(context.Background(), comm.MeetingExtractInput{ + Uid: rec.Uid, + RecordID: fmt.Sprintf("%d", rec.Id), + Summary: rec.Summary, + MeetingDate: meetingDate, + LLMSvcID: rec.LlmSvcId, + // 内容门槛的两个输入,阈值在 memory 那边(配置项在它的 Options 里) + Seconds: rec.Seconds, + ContentRunes: transcriptRunes(rec), + }) +} + +// transcriptRunes 转写正文的有效字数:只数说话人说的话,不数 [Speaker_x] 标签。 +// +// 取值口径与 AIProcess 拼 originalText 那段一致(优先译文、为空退回原文), +// 这样门槛判的就是真正喂给模型的那份内容。解析不出 JSON 时按整串长度算 —— +// 那是 provider 直接给了纯文本的情况,不是异常。 +func transcriptRunes(rec *pb.DBEchoMeetRecord) int { + src := rec.Translate + if strings.TrimSpace(src) == "" { + src = rec.Original + } + if strings.TrimSpace(src) == "" { + return 0 + } + contexts := make([]*pb.ContextStruct, 0) + if err := json.Unmarshal([]byte(src), &contexts); err != nil { + return len([]rune(strings.TrimSpace(src))) + } + n := 0 + for _, v := range contexts { + n += len([]rune(strings.TrimSpace(v.Content))) + } + return n } // PollTranscribe 主动查询第三方转写任务状态(统一入口,按 record.AsrSvcId 路由到对应 provider) @@ -687,18 +791,7 @@ func (this *tasksComp) PollTranscribe(rec *pb.DBEchoMeetRecord) { this.module.Debugf("转写轮询 id:%d 收尾进行中,跳过", rec.Id) return } - personnels := make(map[string]struct{}) - for _, c := range result.Contexts { - c.Meetingid = rec.Id - if c.Speaker != "" && !strings.HasPrefix(c.Speaker, "Speaker") { - c.Speaker = "Speaker_" + c.Speaker - } - personnels[c.Speaker] = struct{}{} - } - rec.Personnel = "" - for k := range personnels { - rec.Personnel += fmt.Sprintf("[%s] ", k) - } + normalizeSpeakers(rec, result.Contexts) rec.Original = utils.ToString(result.Contexts) // 先把状态推离 Transcribing 并落库:后续 getrecord 不再进入本分支(与上面的内存锁双保险)。 rec.State = pb.DBEchoMeetRecordState_AwaitSummarizing @@ -745,5 +838,5 @@ func (this *tasksComp) finishTranscribe(id uint64, contexts []*pb.ContextStruct) } this.module.model.saverecord(rec) this.module.Infof("转写轮询 id:%d 转写完成,提交AI任务", id) - _ = this.SubmitAITask(rec) + _ = this.submitAITaskOnce(rec) } diff --git a/apps/services/modules/mcp/tool_meeting.go b/apps/services/modules/mcp/tool_meeting.go index a5b7717c..f876db58 100644 --- a/apps/services/modules/mcp/tool_meeting.go +++ b/apps/services/modules/mcp/tool_meeting.go @@ -133,22 +133,10 @@ func (this *tool_search_meeting_notes) Handl(ctx context.Context, request mcp.Ca endTs, _ := parseMeetingDay(request.GetString("end_date", ""), true) keywords := splitMeetingKeywords(request.GetString("keywords", "")) - tx := mysql.Table(comm.TableEchomeetRecord).Where("uid = ?", uid) - - // 只找总结真的做完、且有实质内容的。 - // ⚠️ 不能只判 `summary <> ''`:真机库里存在 summary 只有 2 个字符的记录 - // (极短录音总结出来的残次品)。它照样能通过非空判断,占掉 limit 的名额, - // 把真正有内容的会议挤出去。 - tx = tx.Where("CHAR_LENGTH(summary) >= ?", 20) - - // 时间范围:记录上没有「会议实际发生日期」这个字段,只能用 creationtime - // (= 录音上传时间)。补录旧录音时会偏,这是已知偏差,不是这里能修的。 - if startTs > 0 { - tx = tx.Where("creationtime >= ?", startTs) - } - if endTs > 0 { - tx = tx.Where("creationtime <= ?", endTs) - } + // ⚠️ 三条检索路径(无关键词 / 全文 / LIKE / 回退)必须共用同一套过滤条件, + // 所以一律从 recentQuery 起手。这里原先另拼了一份一模一样的条件, + // 结果是「给 recentQuery 加一道门槛」只对其中三条生效,无关键词那条照旧漏过去。 + tx := this.recentQuery(uid, startTs, endTs) records := make([]*pb.DBEchoMeetRecord, 0, limit) matched := "" // 实际用了哪种检索方式,只为日志和返回值里说明 @@ -227,11 +215,22 @@ func (this *tool_search_meeting_notes) Handl(ctx context.Context, request mcp.Ca } // recentQuery 不带关键词的「最近会议」查询,与主查询共用同一套过滤条件 -// (只看本人、只看总结完成的、时间范围)。 +// (只看本人、只看总结完成的、够长的)。 +// +// ⚠️ 光靠 CHAR_LENGTH(summary) 挡不住垃圾纪要:内容极薄时模型按模板兜底规则 +// 照样能写出六百多字,一条 6 秒的测试通话就能挤进检索结果、把真会议顶掉。 +// 所以另加一道时长门槛,口径与拾忆的待办抽取同一个常量。 func (this *tool_search_meeting_notes) recentQuery(uid string, start, end int64) *gorm.DB { tx := mysql.Table(comm.TableEchomeetRecord). Where("uid = ?", uid). - Where("CHAR_LENGTH(summary) >= ?", 20) + // 只找总结真的做完、且有实质内容的。 + // ⚠️ 不能只判 `summary <> ''`:真机库里存在 summary 只有 2 个字符的记录 + // (极短录音总结出来的残次品)。它照样能通过非空判断,占掉 limit 的名额, + // 把真正有内容的会议挤出去。 + Where("CHAR_LENGTH(summary) >= ?", 20). + Where("seconds >= ?", comm.MeetingMinSeconds) + // 时间范围:记录上没有「会议实际发生日期」这个字段,只能用 creationtime + // (= 录音上传时间)。补录旧录音时会偏,这是已知偏差,不是这里能修的。 if start > 0 { tx = tx.Where("creationtime >= ?", start) } diff --git a/apps/services/modules/mcp/tool_memory.go b/apps/services/modules/mcp/tool_memory.go index f27cd301..83079eab 100644 --- a/apps/services/modules/mcp/tool_memory.go +++ b/apps/services/modules/mcp/tool_memory.go @@ -174,6 +174,9 @@ func (this *tool_get_memory_stats) Handl(ctx context.Context, request mcp.CallTo ideaCount += r.N } } + // 重复闹钟补上周期内多出来的次数。口径与 home 的周期统计共用 comm 里那两个函数—— + // 各写一份的话会变成「App 里的周报说 7 次、EMAI 说 1 次」,而且谁都不报错。 + alarmTotal += int64(repeatAlarmExtras(uid, start, end)) // 花销**按币种分行**,不做汇率换算:把 CNY 和 JPY 加成一个数是错的。 type expRow struct { @@ -341,8 +344,11 @@ func briefItems(items []*pb.DBMemoryItem) []map[string]interface{} { "id": it.Id, "category": it.Category, "title": it.Title, - "date": it.HappenDate, - "done": it.State == pb.MemoryState_MemoryState_Done, + // ⚠️ 归一:happen_date 是 type:date 列 + parseTime=True, + // 读出来是 `2026-09-07T00:00:00+08:00`。原样交给模型,它回答时就会 + // 把这一长串念出来,或者拿它跟用户说的「9 月 7 号」对不上。 + "date": comm.NormalizeMemoryDate(it.HappenDate), + "done": it.State == pb.MemoryState_MemoryState_Done, } if it.HappenTime != "" { m["time"] = it.HappenTime @@ -368,3 +374,28 @@ func briefItems(items []*pb.DBMemoryItem) []map[string]interface{} { } return out } + +// repeatAlarmExtras 重复闹钟在 [start,end] 里比库里那一行多出来的次数。 +// 查不到或日期坏了一律当 0:统计少算一点,好过整个工具报错。 +func repeatAlarmExtras(uid, start, end string) int { + s, ok1 := comm.ParseMemoryDate(start) + e, ok2 := comm.ParseMemoryDate(end) + if !ok1 || !ok2 { + return 0 + } + items := make([]*pb.DBMemoryItem, 0) + if err := comm.ApplyRepeatCandidates(mysql.Table(comm.TableMemoryItem), uid, end, + []string{comm.MemoryCatAlarm}).Find(&items).Error; err != nil { + safeLogErrorf("get_memory_stats 展开重复闹钟失败已忽略: %v", err) + return 0 + } + total := 0 + for _, it := range items { + anchor, ok := comm.ParseMemoryDate(it.HappenDate) + if !ok { + continue + } + total += comm.MemoryExtraOccurrences(it.RepeatRule, it.Weekday, anchor, s, e) + } + return total +} diff --git a/apps/services/modules/memory/api_today.go b/apps/services/modules/memory/api_today.go index dd37420c..c559bdc3 100644 --- a/apps/services/modules/memory/api_today.go +++ b/apps/services/modules/memory/api_today.go @@ -45,6 +45,11 @@ func (this *apiComp) Today(session comm.IUserSession, req *pb.MemoryTodayReq) (r } resp = &pb.MemoryTodayResp{Items: items} + // 补偿:顺手看一眼最近两个周期的报告是不是卡住或压根没建。 + // ⚠️ 必须 go 出去:memory_today 在启动路径附近被调用,不能被两次查询拖住。 + // ensureRecent 自带 recover,炸了也不影响这次响应。 + go this.module.report.ensureRecent(uid) + // 待确认报告:查失败不算错误,弹窗还是要能出来(少一个角标 << 整个弹窗不出) if rep, e := this.module.model.latestUnconfirmedReport(uid); e != nil { this.module.Warnf("memory_today uid:%s 查待确认报告失败已忽略: %v", uid, e) diff --git a/apps/services/modules/memory/core.go b/apps/services/modules/memory/core.go index 4f51ed5c..dc5585ca 100644 --- a/apps/services/modules/memory/core.go +++ b/apps/services/modules/memory/core.go @@ -15,6 +15,7 @@ type IModel interface { delItems(uid string, ids []uint64) (int64, error) completeItems(uid string, ids []uint64, done bool) (int64, error) delMeetingItems(sourceID string, genRound int32) (int64, error) + delMeetingItemsBySource(uid string, sourceIDs []string) (int64, error) // 周期报告 upsertReport(rec *pb.DBMemoryReport) error @@ -24,6 +25,7 @@ type IModel interface { latestUnconfirmedReport(uid string) (*pb.DBMemoryReport, error) listReports(uid, ptype string, page, size int32) ([]*pb.DBMemoryReport, int64, error) uidsWithItemsBetween(start, end string, minItems int) ([]string, error) + countItemsBetween(uid, start, end string) (int64, error) } // Redis 队列键后缀(实际使用时还会经 service.GetTag() 与 redissys.RKey 加应用前缀, @@ -31,4 +33,9 @@ type IModel interface { const ( QueueReportAwait = "memory:report:await" //报告生成等待队列 LockReportPrefix = "memory:report:lock" //cron 抢锁前缀,后面拼 : + // 单份报告的生成锁,后面拼 report id。与上面那把不是一回事: + // 上面那把按周期抢,管的是「这一轮 cron 由谁来建单」;这把按报告抢, + // 管的是「这一份报告由哪个 worker 真的去跑 LLM」——cron 与 ensureRecent + // 两条路都能把同一份报告送进队列,没有它就会跑两遍、计两次费。 + LockReportProcPrefix = "memory:report:proc" ) diff --git a/apps/services/modules/memory/meeting_extract.go b/apps/services/modules/memory/meeting_extract.go index 77958e32..3bce22f0 100644 --- a/apps/services/modules/memory/meeting_extract.go +++ b/apps/services/modules/memory/meeting_extract.go @@ -42,7 +42,11 @@ const meetingExtractPrompt = `你是一个会议待办抽取器。用户会给 含糊的表述(「尽快」「月底前」无法定位到具体某天)一律留空 due_date, 把原文放进 due_raw。**禁止猜测、禁止编造日期。** 6. 只收录会议中明确要求执行的任务;仅在讨论中提及、未拍板的设想不计入。 -7. 同一件事被多次提及只输出一条,取最终确定的版本。` +7. 同一件事被多次提及只输出一条,取最终确定的版本。 +8. 纪要若表明这不是一场真实会议(「单人测试」「录音测试」「无实质内容」「未形成待办」之类), + 直接输出 [],不要为了填满「待办事项」这一节而编造任务。 +9. 说话人对自己的愿望、需求、期望的描述(如「希望能导出成文字」「想评估一下效果」)不是待办; + 只收录会议里明确指派给某人、某方去执行的事。` // extractedTodo 抽取结果的一条 type extractedTodo struct { @@ -56,12 +60,35 @@ type extractedTodo struct { // 全落下去会把用户的日历那一天彻底刷爆。 const maxExtractPerMeeting = 30 +// meetingContentEnough 内容够不够格进拾忆。ok=false 时 reason 是能直接打进日志的说明。 +// +// 两道门槛任一不过就跳过。时长与字数都要看: +// - 只看时长:一段 5 分钟的环境噪音转不出几个字,照样会被抽出废话; +// - 只看字数:模型在内容极薄时会按模板兜底规则写出六百多字的纪要,所以字数指的是 +// **转写正文**(调用方已去掉说话人标签),卡在纪要上根本拦不住。 +// +// 阈值 <= 0 表示这一道不限;**数据缺失(seconds/contentRunes 为 0)一律按通过处理**: +// 宁可多抽几条用户能自己删的待办,也不要因为上游漏传一个字段就把所有会议的抽取静默关掉。 +// +// 写成纯函数而不是 Memory 的方法,是为了能直接单测,不用去造一个带日志的模块实例。 +func meetingContentEnough(seconds int32, contentRunes int, minSeconds int32, minChars int) (bool, string) { + if minSeconds > 0 && seconds > 0 && seconds < minSeconds { + return false, fmt.Sprintf("录音仅 %d 秒(门槛 %d 秒),内容太薄", seconds, minSeconds) + } + if minChars > 0 && contentRunes > 0 && contentRunes < minChars { + return false, fmt.Sprintf("转写正文仅 %d 字(门槛 %d 字),内容太薄", contentRunes, minChars) + } + return true, "" +} + // ExtractMeetingTodos 从会议纪要抽待办并落库(实现 comm.IMemory)。 // // meetingDate 是会议归属日期(YYYY-MM-DD);抽不出截止日期的项落在这一天并置 // date_certain=false,客户端应把它们归到「待定日期」分组,不要和当天确定事项混排 // —— 一次会抽出 8 条没写截止时间的待办全堆在会议当天,那一格就没法看了。 -func (this *Memory) ExtractMeetingTodos(ctx context.Context, uid, recordID, summary, meetingDate, llmSvcID string) { +func (this *Memory) ExtractMeetingTodos(ctx context.Context, in comm.MeetingExtractInput) { + uid, recordID, summary := in.Uid, in.RecordID, in.Summary + meetingDate, llmSvcID := in.MeetingDate, in.LLMSvcID defer func() { if r := recover(); r != nil { this.Errorf("会议待办抽取 panic record:%s err:%v", recordID, r) @@ -73,6 +100,16 @@ func (this *Memory) ExtractMeetingTodos(ctx context.Context, uid, recordID, summ if strings.TrimSpace(summary) == "" || uid == "" { return } + // 内容门槛:太短的录音一律不抽。 + // + // ⚠️ 只拦抽取,**不动总结**:「记一下明天买菜」这种十几秒的语音备忘是合法用法, + // 纪要照出。也不把它判成 SummarizFail —— 客户端把那个码显示成 + // 「总结失败,请稍后重试」,用户会一直点。 + if ok, why := meetingContentEnough(in.Seconds, in.ContentRunes, + this.options.MeetingExtractMinSeconds, this.options.MeetingExtractMinChars); !ok { + this.Infof("会议待办抽取 record:%s %s,跳过抽取", recordID, why) + return + } if _, ok := comm.ParseMemoryDate(meetingDate); !ok { this.Warnf("会议待办抽取 record:%s 会议日期非法(%q),跳过", recordID, meetingDate) return @@ -172,6 +209,21 @@ func (this *Memory) ExtractMeetingTodos(ctx context.Context, uid, recordID, summ this.Infof("会议待办抽取 record:%s uid:%s svc:%s 落库 %d/%d 条", recordID, uid, usedSvc, saved, len(todos)) } +// DeleteMeetingItems 纪要被删时清掉它自动抽出来的待办(实现 comm.IMemory)。 +// +// 调用方(echomeet 的 DelRecords)在归属校验通过之后另起 goroutine 调它, +// 失败只记日志:纪要已经删掉了,不能因为附带的清理失败就把删除本身判成失败。 +func (this *Memory) DeleteMeetingItems(uid string, recordIDs []string) (int64, error) { + n, err := this.model.delMeetingItemsBySource(uid, recordIDs) + if err != nil { + return 0, err + } + if n > 0 { + this.Infof("纪要删除连带清理待办 uid:%s records:%v 删掉 %d 条(用户改过的保留)", uid, recordIDs, n) + } + return n, nil +} + // parseExtractedTodos 解析模型输出。 // // 模型时常无视「不要 markdown 代码块」这条,所以先剥一层 ```json ... ```; diff --git a/apps/services/modules/memory/meeting_extract_test.go b/apps/services/modules/memory/meeting_extract_test.go index 4fe85e07..f5706154 100644 --- a/apps/services/modules/memory/meeting_extract_test.go +++ b/apps/services/modules/memory/meeting_extract_test.go @@ -1,6 +1,10 @@ package memory -import "testing" +import ( + "testing" + + "yunyan/comm" +) // 模型时常无视「不要 markdown 代码块」,也常在数组前后带一句解释。 // 这些都不是异常输入,是实测里的常态,解析必须扛得住。 @@ -51,3 +55,66 @@ func TestParseExtractedTodos_FieldMapping(t *testing.T) { t.Errorf("字段映射错: %+v", got[0]) } } + +// 内容门槛:拦住 6 秒测试通话、20 秒单人自述这类东西, +// 同时不能误伤「上游没给数据」和「运营把门槛关掉」两种情况。 +func TestMeetingContentEnough(t *testing.T) { + const ( + minSec = int32(60) + minChars = 120 + ) + cases := []struct { + name string + seconds int32 + runes int + minSec int32 + minChars int + want bool + }{ + {"6秒测试通话", 6, 300, minSec, minChars, false}, + {"20秒单人自述:时长先拦住", 20, 90, minSec, minChars, false}, + {"够长但正文太薄(长段噪音)", 600, 40, minSec, minChars, false}, + {"正常会议", 900, 3000, minSec, minChars, true}, + {"刚好卡在门槛上算通过", 60, 120, minSec, minChars, true}, + // 下面两条是刻意的失败方向:宁可多抽,也不要因为一个字段没传就静默关掉抽取 + {"上游没给时长 → 不按时长拦", 0, 3000, minSec, minChars, true}, + {"上游没给字数 → 不按字数拦", 900, 0, minSec, minChars, true}, + {"门槛配 0 = 不限", 6, 10, 0, 0, true}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + got, why := meetingContentEnough(c.seconds, c.runes, c.minSec, c.minChars) + if got != c.want { + t.Fatalf("期望 %v,得到 %v(%s)", c.want, got, why) + } + if !got && why == "" { + t.Error("拦下来了却没给出原因,日志里会只剩一句「跳过抽取」") + } + }) + } +} + +// 门槛的默认值必须真的生效:LoadConfig 不写这两项时是 60/120, +// 显式写 0 则表示不限——两者不能混为一谈(零值与「没配」在 mapstructure 里长得一样)。 +func TestOptionsExtractThresholdDefaults(t *testing.T) { + var noCfg Options + if err := noCfg.LoadConfig(map[string]interface{}{}); err != nil { + t.Fatal(err) + } + if noCfg.MeetingExtractMinSeconds != comm.MeetingMinSeconds || noCfg.MeetingExtractMinChars != 120 { + t.Errorf("没配时应取默认 %d/120,得到 %d/%d", + comm.MeetingMinSeconds, noCfg.MeetingExtractMinSeconds, noCfg.MeetingExtractMinChars) + } + + var zeroCfg Options + if err := zeroCfg.LoadConfig(map[string]interface{}{ + "meeting_extract_min_seconds": 0, + "meeting_extract_min_chars": 0, + }); err != nil { + t.Fatal(err) + } + if zeroCfg.MeetingExtractMinSeconds != 0 || zeroCfg.MeetingExtractMinChars != 0 { + t.Errorf("显式配 0 应表示不限,得到 %d/%d", + zeroCfg.MeetingExtractMinSeconds, zeroCfg.MeetingExtractMinChars) + } +} diff --git a/apps/services/modules/memory/migrate_defaults.go b/apps/services/modules/memory/migrate_defaults.go new file mode 100644 index 00000000..808dbca1 --- /dev/null +++ b/apps/services/modules/memory/migrate_defaults.go @@ -0,0 +1,53 @@ +package memory + +import ( + "yunyan/comm" + "yunyan/lego/sys/mysql" +) + +/* +存量修正:把被 gorm 的 `default:` 标签吞掉的会议待办标记改回来。 + +起因见 model.go 的 addItem —— 2026-09-18 之前 ExtractMeetingTodos 写的 +DateCertain:false / RemindAhead:0 一条都没落进库,测试库里 16 条会议待办 +全是 date_certain=1 / remind_ahead=5。代码修好之后**新数据**才对, +已经写歪的这些得单独扶正,否则用户的日历里那批待办照旧在 08:55 响提醒。 + +三条刻意的边界: + + - **只碰 source='meeting' 且 user_edited=0 的行**。用户改过的项已经是他自己的东西, + 哪怕当初的值来自 bug,也不该在他背后再改一次。 + - **助手闹钟那批 remind_ahead=5 不动**。当初用户到底有没有选「不提醒」现在已无从知道, + 而 5 分钟正是产品默认值,改了反而可能把用户设过的提醒抹掉。 + - **date_certain 只回退「有 due_raw 却仍落在会议当天」的那些**。当初真从原文推算出 + 截止日的项本来就该是 1,无法与被吞掉的那批区分,所以只挑这个能确定判错的子集。 + +幂等:条件里带了「当前值不等于目标值」,跑第二遍影响 0 行。 +不写 next_remind_at:全项目没有任何地方读它(只有 recalcRemind 在写), +动它只会扩大这次修正的影响面。 +*/ + +// fixSwallowedDefaults 修正存量会议待办被默认值覆盖的两列,返回改动行数。 +func (this *modelComp) fixSwallowedDefaults() (int64, error) { + var total int64 + + // 会议待办一律不提醒 + tx := mysql.Table(comm.TableMemoryItem). + Where("source = ? AND user_edited = ? AND remind_ahead <> ?", comm.MemorySrcMeeting, false, 0). + Update("remind_ahead", 0) + if tx.Error != nil { + return total, tx.Error + } + total += tx.RowsAffected + + // 有截止原文、却没推算出具体日期(= 仍落在会议当天)→ 该进「待定日期」分组 + tx = mysql.Table(comm.TableMemoryItem). + Where("source = ? AND user_edited = ? AND due_raw <> ? AND date_certain = ?", + comm.MemorySrcMeeting, false, "", true). + Update("date_certain", false) + if tx.Error != nil { + return total, tx.Error + } + total += tx.RowsAffected + return total, nil +} diff --git a/apps/services/modules/memory/model.go b/apps/services/modules/memory/model.go index d3ae55f9..4051dbf1 100644 --- a/apps/services/modules/memory/model.go +++ b/apps/services/modules/memory/model.go @@ -38,6 +38,44 @@ func (this *modelComp) Init(service core.IService, module core.IModule, comp cor return nil } +/* +⚠️ 从库里读出来的日期要先归一再用,也再回给客户端。 + +`happen_date` / `period_start` / `period_end` 都是 MySQL 的 `type:date` 列, +而 DSN 带 `parseTime=True` —— gorm 把它们扫进 Go 的 **string** 字段时给的是 +`2026-09-07T00:00:00+08:00`,不是协议约定的 `2026-09-07`。 + +两头都会出事: + - **回给客户端的响应**里也是这个形状。客户端拿它做字符串比较 + (`_localToday()` 的 `e.happenDate == today`、`_mergeRange` 的 compareTo), + 「今天」那一栏的离线兜底恒为空,区间边界上的项还会重复; + - **服务端自己解析**它的地方会静默失效(见 comm.ParseMemoryDate 的注释)。 + +第二头已经由 ParseMemoryDate 容忍带时间的形式兜住了;这里管的是第一头: +凡是把记忆项/报告交出去的读函数,出口处统一归一。 +*/ + +// normalizeItemDates 把读出来的日期列收敛成 YYYY-MM-DD +func normalizeItemDates(items ...*pb.DBMemoryItem) { + for _, it := range items { + if it == nil { + continue + } + it.HappenDate = comm.NormalizeMemoryDate(it.HappenDate) + } +} + +// normalizeReportDates 同上,管周期报告的起止日 +func normalizeReportDates(recs ...*pb.DBMemoryReport) { + for _, r := range recs { + if r == nil { + continue + } + r.PeriodStart = comm.NormalizeMemoryDate(r.PeriodStart) + r.PeriodEnd = comm.NormalizeMemoryDate(r.PeriodEnd) + } +} + // ===== 记忆项 ===== // itemQuery listItems 的查询条件 @@ -51,11 +89,59 @@ type itemQuery struct { size int32 } +// gormZeroDefaultColumns DBMemoryItem 上「默认值不是零值」的那几列。 +// +// gorm 在 Create 时会把这些列的**零值替换成标签里的默认值**再写库 +// (替换在 callbacks.ConvertToCreateValues 的 reflect.Struct 分支里无条件做, +// 与 Select / Omit 都无关——加 Select("*") 生成的 VALUES 一模一样,DryRun 实测过)。 +// 默认值本身就是零值的列(default:0 / default:false)替换了也还是零值,不在此列。 +// +// 清单由 TestGormNonZeroDefaultsAreHandled 按 pb 结构体的 tag 反查守着: +// proto 里再加一个非零 default: 而这里没跟上,那条测试会红。 +var gormZeroDefaultColumns = []string{"date_certain", "remind_ahead"} + +// addItem 写入一条记忆项。 +// +// ⚠️ 插完必须把 gormZeroDefaultColumns 那几列按结构体的现值补写回去。 +// mysql.Insert 就是 db.Table().Create(model),于是 ExtractMeetingTodos 明明写了 +// DateCertain:false / RemindAhead:0,2026-09-18 测试库里 25 条记录**全部**落成 1 / 5: +// - 客户端「待定日期」那一组永远是空的,没截止日的会议待办全堆在会议当天; +// - memory_upcoming 按 remind_ahead=5 展开,会议待办会在 08:55 响系统通知, +// 与「会议待办默认不提醒」正好相反; +// - 用户手动建闹钟选「不提醒」同样被改成提前 5 分钟。 +// +// 不去掉 proto 上的 `default:` 标签:那要重生成 memory_db.pb.go(生成器版本那个坑), +// 而且 MySQL 列上已经建出来的 DEFAULT 不会跟着变,收益为零。 +// api_update 走 mysql.Save(UPDATE 写全字段),不受这条影响。 func (this *modelComp) addItem(item *pb.DBMemoryItem) error { now := time.Now().Unix() item.CreateTime = now item.UpdateTime = now - return mysql.Insert(comm.TableMemoryItem, item) + // gorm 替换零值时连结构体上的字段一起改(field.Set),所以先把真值留一份: + // 不还原的话 memory_add 回给客户端的 item 也是被改过的。 + dateCertain, remindAhead := item.DateCertain, item.RemindAhead + if err := mysql.Insert(comm.TableMemoryItem, item); err != nil { + return err + } + fix := make(map[string]interface{}, len(gormZeroDefaultColumns)) + if !dateCertain { + fix["date_certain"] = false + item.DateCertain = false + } + if remindAhead == 0 { + fix["remind_ahead"] = int32(0) + item.RemindAhead = 0 + } + if len(fix) == 0 { + return nil + } + // ⚠️ 补写失败只告警不回错:行已经插进去了,回错会让调用方以为没落库 + // (memory_add 会对客户端报「新增失败」,而那条记忆项其实已经在库里)。 + if err := mysql.Table(comm.TableMemoryItem).Where("id = ?", item.Id).Updates(fix).Error; err != nil { + this.module.Warnf("记忆项 id:%d 补写 %v 失败(该项的提醒/日期标记会退回默认值): %v", + item.Id, fix, err) + } + return nil } func (this *modelComp) findItemByClientKey(uid, clientKey string) (*pb.DBMemoryItem, error) { @@ -70,6 +156,7 @@ func (this *modelComp) findItemByClientKey(uid, clientKey string) (*pb.DBMemoryI } return nil, err } + normalizeItemDates(item) return item, nil } @@ -83,6 +170,7 @@ func (this *modelComp) getItem(uid string, id uint64) (*pb.DBMemoryItem, error) } return nil, err } + normalizeItemDates(item) return item, nil } @@ -119,6 +207,7 @@ func (this *modelComp) listItems(q *itemQuery) ([]*pb.DBMemoryItem, int64, error if err := tx.Find(&items).Error; err != nil { return nil, 0, err } + normalizeItemDates(items...) return items, total, nil } @@ -173,6 +262,19 @@ func (this *modelComp) delMeetingItems(sourceID string, genRound int32) (int64, return tx.RowsAffected, tx.Error } +// delMeetingItemsBySource 删掉若干条会议自动抽取的待办(纪要被删时调)。 +// 带 uid 是归属校验:这条路径的入参来自 echomeet 的删除请求,不能只信 source_id。 +func (this *modelComp) delMeetingItemsBySource(uid string, sourceIDs []string) (int64, error) { + if uid == "" || len(sourceIDs) == 0 { + return 0, nil + } + tx := mysql.Table(comm.TableMemoryItem). + Where("uid = ? and source = ? and source_id in ? and user_edited = ?", + uid, comm.MemorySrcMeeting, sourceIDs, false). + Delete(&pb.DBMemoryItem{}) + return tx.RowsAffected, tx.Error +} + // ===== 周期报告 ===== func (this *modelComp) upsertReport(rec *pb.DBMemoryReport) error { @@ -212,6 +314,7 @@ func (this *modelComp) getReport(id uint64) (*pb.DBMemoryReport, error) { } return nil, err } + normalizeReportDates(rec) return rec, nil } @@ -224,6 +327,7 @@ func (this *modelComp) getReportByPeriod(uid, ptype, pkey string) (*pb.DBMemoryR } return nil, err } + normalizeReportDates(rec) return rec, nil } @@ -241,6 +345,7 @@ func (this *modelComp) latestUnconfirmedReport(uid string) (*pb.DBMemoryReport, if len(recs) == 0 { return nil, nil } + normalizeReportDates(recs[0]) return recs[0], nil } @@ -262,6 +367,7 @@ func (this *modelComp) listReports(uid, ptype string, page, size int32) ([]*pb.D offset = int(page-1) * int(size) } err := tx.Order("period_end desc").Offset(offset).Limit(int(size)).Find(&recs).Error + normalizeReportDates(recs...) return recs, total, err } @@ -281,6 +387,19 @@ func (this *modelComp) uidsWithItemsBetween(start, end string, minItems int) ([] return uids, err } +// countItemsBetween 单个用户在周期内的记录条数。 +// +// 与 uidsWithItemsBetween 同一个条件(按 happen_date 落在范围内),只是收敛到一个人: +// ensureRecent 补建报告时要过同一道闸门,两边条件不一致会出现 +// 「cron 判定这人不够条数不建,用户一回来又给他建一份」的来回摆动。 +func (this *modelComp) countItemsBetween(uid, start, end string) (int64, error) { + var n int64 + err := mysql.Table(comm.TableMemoryItem). + Where("uid = ? and happen_date >= ? and happen_date <= ?", uid, start, end). + Count(&n).Error + return n, err +} + // itemFieldWhitelist memory_update 允许客户端点名更新的字段。 // 白名单而不是黑名单:漏挡一个 uid/id/create_time 就是越权改别人的数据。 var itemFieldWhitelist = map[string]bool{ diff --git a/apps/services/modules/memory/model_default_test.go b/apps/services/modules/memory/model_default_test.go new file mode 100644 index 00000000..4b9e0730 --- /dev/null +++ b/apps/services/modules/memory/model_default_test.go @@ -0,0 +1,188 @@ +package memory + +import ( + "database/sql" + "database/sql/driver" + "reflect" + "strconv" + "strings" + "testing" + + "yunyan/pb" + + "gorm.io/driver/mysql" + "gorm.io/gorm" +) + +/* +守住 addItem 那条补写逻辑的两个前提。 + +背景:gorm 对带 `default:` 标签的字段,Create 时把零值换成标签里的默认值再写库。 +DBMemoryItem 的 date_certain(default:true) / remind_ahead(default:5) 因此永远写不进 +false / 0 —— 2026-09-18 测试库里 25 条记录全部落成 1 / 5。 + +⚠️ 一度以为加 `Select("*")` 能绕开,实测不行:替换发生在 ConvertToCreateValues 的 +reflect.Struct 分支,只看字段值是不是零值,与 Select / Omit 无关。 +TestGormCreateSwallowsZeroValues 就是把这个事实钉住,免得下次又有人去试。 +*/ + +// ── DryRun 用的假驱动:只要能让 gorm 把 SQL 拼出来,不需要真连库 ── + +type nullDriver struct{} + +func (nullDriver) Open(string) (driver.Conn, error) { return nullConn{}, nil } + +type nullConn struct{} + +func (nullConn) Prepare(string) (driver.Stmt, error) { return nil, driver.ErrSkip } +func (nullConn) Close() error { return nil } +func (nullConn) Begin() (driver.Tx, error) { return nil, driver.ErrSkip } + +func init() { sql.Register("memory_model_test_null", nullDriver{}) } + +func dryRunDB(t *testing.T) *gorm.DB { + t.Helper() + conn, _ := sql.Open("memory_model_test_null", "") + db, err := gorm.Open( + mysql.New(mysql.Config{Conn: conn, SkipInitializeWithVersion: true}), + // SkipDefaultTransaction:写操作默认包事务,假驱动的 Begin 会报错,SQL 就拼不出来了 + &gorm.Config{DryRun: true, DisableAutomaticPing: true, SkipDefaultTransaction: true}, + ) + if err != nil { + t.Fatalf("起 DryRun 连接失败: %v", err) + } + return db +} + +// insertValueOf 取一条 INSERT 里某列的绑定值。 +func insertValueOf(t *testing.T, st *gorm.Statement, column string) interface{} { + t.Helper() + sql := st.SQL.String() + head := sql[strings.Index(sql, "(")+1 : strings.Index(sql, ")")] + for i, c := range strings.Split(head, ",") { + if strings.Trim(strings.TrimSpace(c), "`") == column { + if i >= len(st.Vars) { + t.Fatalf("列 %s 下标 %d 超出绑定值个数 %d", column, i, len(st.Vars)) + } + return st.Vars[i] + } + } + t.Fatalf("INSERT 里没有列 %s:%s", column, sql) + return nil +} + +// TestGormCreateSwallowsZeroValues 记录 bug 本身:Create 会把 false/0 换成 true/5, +// 加不加 Select("*") 都一样。哪天 gorm 改了行为这条会红,那时 addItem 的补写就可以删了。 +func TestGormCreateSwallowsZeroValues(t *testing.T) { + for _, tc := range []struct { + name string + make func(*gorm.DB, *pb.DBMemoryItem) *gorm.DB + }{ + {"裸 Create", func(db *gorm.DB, it *pb.DBMemoryItem) *gorm.DB { + return db.Table("memory_item").Create(it) + }}, + {"Select(*) 也救不了", func(db *gorm.DB, it *pb.DBMemoryItem) *gorm.DB { + return db.Table("memory_item").Select("*").Create(it) + }}, + } { + t.Run(tc.name, func(t *testing.T) { + item := &pb.DBMemoryItem{ + Uid: "u1", Title: "跟进合同", HappenDate: "2026-09-18", + DateCertain: false, RemindAhead: 0, + } + st := tc.make(dryRunDB(t), item).Statement + if got := insertValueOf(t, st, "date_certain"); got != true { + t.Errorf("date_certain 期望被换成 true(即 bug 仍在),实际 %v", got) + } + // ⚠️ 比数值不比类型:gorm 的默认值来自 strconv.ParseInt,塞回去是 int64 不是字段的 int32 + if got := insertValueOf(t, st, "remind_ahead"); toInt64(t, got) != 5 { + t.Errorf("remind_ahead 期望被换成 5(即 bug 仍在),实际 %v", got) + } + }) + } +} + +// TestGormNonZeroDefaultsAreHandled 反查 pb 结构体上所有「默认值非零」的列, +// 必须与 gormZeroDefaultColumns 一字不差。 +// +// proto 里给某个字段加了 `default:1` 之类而 addItem 没跟着补写,这条会红 —— +// 否则那个字段会重演同一个 bug,而且同样不报任何错。 +func TestGormNonZeroDefaultsAreHandled(t *testing.T) { + got := nonZeroDefaultColumns(reflect.TypeOf(pb.DBMemoryItem{})) + want := append([]string(nil), gormZeroDefaultColumns...) + if !reflect.DeepEqual(got, sortedCopy(want)) { + t.Fatalf("DBMemoryItem 上非零默认值的列是 %v,而 addItem 只补写 %v;"+ + "请同步 gormZeroDefaultColumns 与 addItem 里的补写分支", got, want) + } +} + +// TestReportHasNoNonZeroDefaults 周期报告表走同一个 mysql.Insert, +// 目前它没有非零默认值的列,所以不需要补写。加了就要按 addItem 的办法处理。 +func TestReportHasNoNonZeroDefaults(t *testing.T) { + if got := nonZeroDefaultColumns(reflect.TypeOf(pb.DBMemoryReport{})); len(got) > 0 { + t.Fatalf("DBMemoryReport 新增了非零默认值的列 %v,upsertReport 的 Insert 会把它们的零值吞掉", got) + } +} + +// nonZeroDefaultColumns 扫结构体 tag,挑出 `gorm:"default:X"` 且 X 不是零值的列, +// 列名取 json tag(pb 生成的 json 名与建表列名一致)。返回值已排序。 +func nonZeroDefaultColumns(t reflect.Type) []string { + out := make([]string, 0, 4) + for i := 0; i < t.NumField(); i++ { + f := t.Field(i) + tag := f.Tag.Get("gorm") + if tag == "" || tag == "-" { + continue + } + def, ok := "", false + for _, seg := range strings.Split(tag, ";") { + if v, found := strings.CutPrefix(strings.TrimSpace(seg), "default:"); found { + def, ok = strings.TrimSpace(v), true + } + } + if !ok || isZeroLiteral(def) { + continue + } + name := strings.Split(f.Tag.Get("json"), ",")[0] + if name == "" || name == "-" { + continue + } + out = append(out, name) + } + return sortedCopy(out) +} + +// toInt64 把绑定值统一成 int64 再比,绕开 int32/int64 的类型差异。 +func toInt64(t *testing.T, v interface{}) int64 { + t.Helper() + rv := reflect.ValueOf(v) + switch rv.Kind() { + case reflect.Int, reflect.Int8, reflect.Int16, reflect.Int32, reflect.Int64: + return rv.Int() + case reflect.Uint, reflect.Uint8, reflect.Uint16, reflect.Uint32, reflect.Uint64: + return int64(rv.Uint()) + } + t.Fatalf("不是整数: %#v", v) + return 0 +} + +func isZeroLiteral(s string) bool { + switch strings.ToLower(strings.TrimSpace(s)) { + case "", "0", "false", "null", "''", `""`: + return true + } + if f, err := strconv.ParseFloat(s, 64); err == nil { + return f == 0 + } + return false +} + +func sortedCopy(in []string) []string { + out := append([]string(nil), in...) + for i := 1; i < len(out); i++ { + for j := i; j > 0 && out[j] < out[j-1]; j-- { + out[j], out[j-1] = out[j-1], out[j] + } + } + return out +} diff --git a/apps/services/modules/memory/module.go b/apps/services/modules/memory/module.go index e56094ac..7b7d8523 100644 --- a/apps/services/modules/memory/module.go +++ b/apps/services/modules/memory/module.go @@ -72,6 +72,12 @@ func (this *Memory) Start() (err error) { } else if n > 0 { this.Infof("存量任务迁移完成,搬入 %d 条记忆项", n) } + // 修正被 gorm default: 标签吞掉的会议待办标记(幂等,见 migrate_defaults.go) + if n, err := this.model.fixSwallowedDefaults(); err != nil { + this.Errorf("存量会议待办标记修正失败(不影响服务): %v", err) + } else if n > 0 { + this.Infof("存量会议待办标记修正完成,改动 %d 行", n) + } }() return } diff --git a/apps/services/modules/memory/options.go b/apps/services/modules/memory/options.go index 817f03e5..d6983dc4 100644 --- a/apps/services/modules/memory/options.go +++ b/apps/services/modules/memory/options.go @@ -1,6 +1,7 @@ package memory import ( + "yunyan/comm" "yunyan/modules" "yunyan/lego/utils/mapstructure" @@ -18,6 +19,11 @@ type Options struct { // 会议纪要 → 待办抽取 MeetingExtract bool `json:"meeting_extract" mapstructure:"meeting_extract"` //是否开启会议待办抽取,默认 true MeetingExtractPrompt string `json:"meeting_extract_prompt" mapstructure:"meeting_extract_prompt"` //抽取 prompt(代码内置默认值,不放 echomeet_template) + // 内容门槛:低于任一阈值的录音不抽待办(纪要照出,只是不进拾忆)。 + // 6 秒的测试通话、20 秒的单人自述被抽成「继续进行进一步的功能测试」这类废话待办, + // 再靠条数凑够周报阈值又触发一次 LLM 调用——2026-09-18 测试库里 id=23/24/26 都是这么来的。 + MeetingExtractMinSeconds int32 `json:"meeting_extract_min_seconds" mapstructure:"meeting_extract_min_seconds"` //短于多少秒不抽取,默认 60;<=0 表示不限 + MeetingExtractMinChars int `json:"meeting_extract_min_chars" mapstructure:"meeting_extract_min_chars"` //转写正文少于多少字不抽取,默认 120;<=0 表示不限 } func (this *Options) LoadConfig(settings map[string]interface{}) (err error) { @@ -49,6 +55,19 @@ func (this *Options) LoadConfig(settings map[string]interface{}) (err error) { if this.ReportStaleSec <= 0 { this.ReportStaleSec = 900 } + // ⚠️ 这两个门槛允许配 0(= 不限),所以只在「配置里没写这一项」时才回默认值。 + // mapstructure 解不出来时字段保持零值,与「运营显式填了 0」无法区分, + // 于是判据取「负数才算没配」——想关掉门槛就填 0,想用默认就别写这两行。 + if this.MeetingExtractMinSeconds < 0 { + this.MeetingExtractMinSeconds = 0 + } else if _, ok := settings["meeting_extract_min_seconds"]; !ok { + this.MeetingExtractMinSeconds = comm.MeetingMinSeconds + } + if this.MeetingExtractMinChars < 0 { + this.MeetingExtractMinChars = 0 + } else if _, ok := settings["meeting_extract_min_chars"]; !ok { + this.MeetingExtractMinChars = 120 + } if this.ReportPrompt == "" { this.ReportPrompt = "你是一位贴心的生活助理。用户会给你一段 JSON 格式的周期统计数据(已由程序精确算出)。" + "请把它组织成一段适合语音朗读的自然中文短文,控制在 150 字以内。要求:" + diff --git a/apps/services/modules/memory/remind_test.go b/apps/services/modules/memory/remind_test.go index 35fdecd7..7c5f35cc 100644 --- a/apps/services/modules/memory/remind_test.go +++ b/apps/services/modules/memory/remind_test.go @@ -149,7 +149,12 @@ func TestExpandAll_SortedAscending(t *testing.T) { // 单条项按它自己的 tz 算,不跟着请求方的时区走: // 「明早 8 点」不该因为用户落地伦敦就变成伦敦时间 8 点 func TestExpandItem_UsesItemOwnTimezone(t *testing.T) { - it := mkItem(comm.MemoryCatAlarm, "2026-09-06", "09:00", comm.MemoryRepeatOnce, 0, 0) + // ⚠️ 日期必须相对「现在」算。expandAll 内部以 time.Now() 为起点, + // 写死一个日期的话这条测试会在那天之后变成「过期的一次性项」而恒返回 0 个时刻 + // ——原来写的 2026-09-06 就是这么在 09-18 集体变红的,且红得毫无提示性。 + sh, _ := time.LoadLocation("Asia/Shanghai") + tomorrow := time.Now().In(sh).AddDate(0, 0, 1).Format("2006-01-02") + it := mkItem(comm.MemoryCatAlarm, tomorrow, "09:00", comm.MemoryRepeatOnce, 0, 0) it.RemindAhead = 5 it.Tz = "Asia/Shanghai" @@ -158,9 +163,29 @@ func TestExpandItem_UsesItemOwnTimezone(t *testing.T) { if len(slots) != 1 { t.Fatalf("应有 1 个时刻,得到 %d", len(slots)) } - sh, _ := time.LoadLocation("Asia/Shanghai") got := time.Unix(slots[0].HappenAt, 0).In(sh) if got.Hour() != 9 { t.Errorf("应按项自己的上海时区算出 09:00,得到 %02d:%02d", got.Hour(), got.Minute()) } } + +// 从库里读出来的 happen_date 是 `2026-09-07T00:00:00+08:00`(type:date 列 + parseTime=True), +// 不是 YYYY-MM-DD。expandItem 必须扛得住这个形状 —— +// 扛不住的后果不是报错,而是 memory_upcoming 恒返回 0 个提醒时刻, +// 整个拾忆提醒在服务端空转,2026-09-18 真机上就是这样。 +func TestExpandItem_AcceptsDbDateFormat(t *testing.T) { + loc, _ := time.LoadLocation("Asia/Shanghai") + from := time.Date(2026, 9, 7, 6, 0, 0, 0, loc) + + plain := mkItem(comm.MemoryCatAlarm, "2026-09-08", "09:00", comm.MemoryRepeatOnce, 0, 5) + fromDb := mkItem(comm.MemoryCatAlarm, "2026-09-08T00:00:00+08:00", "09:00", comm.MemoryRepeatOnce, 0, 5) + + a := expandItem(plain, from, 7, loc) + b := expandItem(fromDb, from, 7, loc) + if len(a) != 1 { + t.Fatalf("干净日期应产生 1 个时刻,得到 %d", len(a)) + } + if len(b) != len(a) || (len(b) == 1 && b[0].RemindAt != a[0].RemindAt) { + t.Errorf("库里读出来的日期格式应与干净格式等价:得到 %d 个时刻(期望 %d 个且时刻相同)", len(b), len(a)) + } +} diff --git a/apps/services/modules/memory/report.go b/apps/services/modules/memory/report.go index 188e2589..54331083 100644 --- a/apps/services/modules/memory/report.go +++ b/apps/services/modules/memory/report.go @@ -5,6 +5,7 @@ import ( "encoding/json" "fmt" "strconv" + "strings" "time" "yunyan/comm" @@ -14,6 +15,8 @@ import ( "yunyan/lego/sys/cron" redissys "yunyan/lego/sys/redis" "yunyan/pb" + + "github.com/redis/go-redis/v9" ) /* @@ -136,12 +139,27 @@ func (this *reportComp) tick(ptype string) { if rec.State == pb.MemoryReportState_MemoryReportState_Done && rec.Confirmed { continue // upsert 里对已确认的直接返回原记录,不重算 } - this.enqueue(rec.Id) + this.enqueueOnce(rec.Id) } } -func (this *reportComp) enqueue(id uint64) { - if err := redissys.Conn().LPush(this.ctx, this.awaitKey(), strconv.FormatUint(id, 10)).Err(); err != nil { +// procLockKey 单份报告的生成锁,防止 cron 与 ensureRecent 同时把它交给两个 worker。 +func (this *reportComp) procLockKey(id uint64) string { + return redissys.RKey(fmt.Sprintf("%s:%s:%d", this.service.GetTag(), LockReportProcPrefix, id)) +} + +// enqueueOnce 入队,但先看它在不在队列里。 +// +// cron、requeueIfStale、ensureRecent 三条路都会入队,不判重的话同一份报告能在队列里 +// 排好几次,每次出队都是一次 LLM 调用(且后跑的覆盖先跑的)。做法同 echomeet 的 +// sweepStuckSummarize:LPos 找不到会返回 redis.Nil,所以 err==nil 才算命中。 +func (this *reportComp) enqueueOnce(id uint64) { + member := strconv.FormatUint(id, 10) + if _, err := redissys.Conn().LPos(this.ctx, this.awaitKey(), member, redis.LPosArgs{}).Result(); err == nil { + this.module.Debugf("报告 id:%d 已在队列里,跳过入队", id) + return + } + if err := redissys.Conn().LPush(this.ctx, this.awaitKey(), member).Err(); err != nil { this.module.Errorf("报告入队失败 id:%d err:%v", id, err) } } @@ -167,10 +185,171 @@ func (this *reportComp) requeueIfStale(rec *pb.DBMemoryReport) bool { this.module.Errorf("报告 id:%d 重置状态失败 err:%v", rec.Id, err) return false } - this.enqueue(rec.Id) + this.enqueueOnce(rec.Id) return true } +// ===== 用户回来时的补偿 ===== + +/* +ensureRecent 是 requeueIfStale 之外的第二道补偿,管的是**它管不到的那些情况**。 + +原先只有 requeueIfStale 一条路,而它由 memory_getreport 驱动;客户端调的是无参版本, +服务端于是走 latestUnconfirmedReport —— 那条 SQL **只查 state=Done**。 +也就是说 pending / processing / failed 的报告永远不会被查出来,也就永远不会被重新入队, +注释里写的「容器 00:30 重启一次就靠它救」实际一次都没救到过。 + +三种卡死,这里一并兜住: + + - cron 那一刻容器不在线 → 那一周连 memory_report 行都没建,下周 tick 算的是新的 + period_key,不回头补; + - 建了单但卡在 pending / processing(跑到一半重启); + - state=Failed 之后再没人碰过它,下一次 tick 已是新周期,旧的永远 Failed。 + +只看「最近一个已结束的周」和「最近一个已结束的月」:与「弹窗只弹最近一份」口径一致, +也避免半年没开 App 的用户一回来就给他建二十份报告、连着烧二十次 LLM。 +*/ + +// reportPeriod 一个待确认的周期 +type reportPeriod struct { + ptype string + key string + start string // YYYY-MM-DD + end string +} + +// lastFinishedWeek / lastFinishedMonth 与 tick 里的算法保持同一口径: +// 往前退到上一个周期里的任意一天,再取那个周期的范围与键。 +func lastFinishedWeek(now time.Time) reportPeriod { + prev := now.AddDate(0, 0, -7) + start, end := comm.MemoryWeekRange(prev) + return reportPeriod{comm.MemoryPeriodWeek, comm.MemoryWeekKey(prev), + comm.FormatMemoryDate(start), comm.FormatMemoryDate(end)} +} + +func lastFinishedMonth(now time.Time) reportPeriod { + prev := time.Date(now.Year(), now.Month(), 1, 0, 0, 0, 0, now.Location()).AddDate(0, 0, -1) + start, end := comm.MemoryMonthRange(prev) + return reportPeriod{comm.MemoryPeriodMonth, comm.MemoryMonthKey(prev), + comm.FormatMemoryDate(start), comm.FormatMemoryDate(end)} +} + +// maxReportRetry 失败报告最多自动重试几次。 +// 不设上限的话,一个必然失败的原因(比如统计 SQL 撞上坏数据)会让用户每次切前台 +// 都触发一次重试,日志刷屏、LLM 白烧。用完这几次就等人来查。 +const maxReportRetry = 3 + +// ensureRecent 用户回来时(memory_today)补建 / 重跑最近两个周期的报告。 +// +// ⚠️ 必须在 goroutine 里调:它有两次主键查询、至多一次 count,不该挂在 today 的响应路径上。 +// 自带 recover —— 它是附赠功能,炸了也不能连累弹窗。 +func (this *reportComp) ensureRecent(uid string) { + defer func() { + if r := recover(); r != nil { + this.module.Errorf("报告补偿 panic uid:%s err:%v", uid, r) + } + }() + if uid == "" { + return + } + now := time.Now() + for _, p := range []reportPeriod{lastFinishedWeek(now), lastFinishedMonth(now)} { + rec, err := this.module.model.getReportByPeriod(uid, p.ptype, p.key) + if err != nil { + this.module.Warnf("报告补偿 uid:%s %s/%s 查询失败已跳过: %v", uid, p.ptype, p.key, err) + continue + } + switch { + case rec == nil: + this.ensureBuilt(uid, p) + case rec.State == pb.MemoryReportState_MemoryReportState_Pending, + rec.State == pb.MemoryReportState_MemoryReportState_Processing: + // 复用同一套超时判定,别在这里另写一份阈值 + this.requeueIfStale(rec) + case rec.State == pb.MemoryReportState_MemoryReportState_Failed && !rec.Confirmed: + this.retryFailed(rec) + } + } +} + +// ensureBuilt cron 那一刻容器不在线 → 补建。必须过与 cron 同一道条数闸门, +// 否则会给「上周只记了一条」的用户建出 cron 本来就不打算建的报告。 +func (this *reportComp) ensureBuilt(uid string, p reportPeriod) { + n, err := this.module.model.countItemsBetween(uid, p.start, p.end) + if err != nil { + this.module.Warnf("报告补偿 uid:%s %s/%s 统计条数失败已跳过: %v", uid, p.ptype, p.key, err) + return + } + if n < int64(this.options.ReportMinItems) { + return // 条数不够,本来就不该有报告,不是卡死 + } + rec := &pb.DBMemoryReport{ + Uid: uid, PeriodType: p.ptype, PeriodKey: p.key, + PeriodStart: p.start, PeriodEnd: p.end, + State: pb.MemoryReportState_MemoryReportState_Pending, + } + if err := this.module.model.upsertReport(rec); err != nil { + this.module.Errorf("报告补偿建单失败 uid:%s %s/%s err:%v", uid, p.ptype, p.key, err) + return + } + this.module.Infof("报告补偿 uid:%s %s/%s cron 当时未建单(%d 条记录),现在补上", uid, p.ptype, p.key, n) + this.enqueueOnce(rec.Id) +} + +// retryFailed 失败的报告重新排一次,至多 maxReportRetry 次。 +// 同样受 ReportStaleSec 节流:一次失败之后至少隔那么久才会再试。 +func (this *reportComp) retryFailed(rec *pb.DBMemoryReport) { + if time.Now().Unix()-rec.UpdateTime < this.options.ReportStaleSec { + return + } + n := retryCount(rec.ErrorMsg) + if n >= maxReportRetry { + this.module.Warnf("报告 id:%d 已重试 %d 次仍失败,不再自动重试: %s", rec.Id, n, rec.ErrorMsg) + return + } + rec.ErrorMsg = withRetryCount(rec.ErrorMsg, n+1) + rec.State = pb.MemoryReportState_MemoryReportState_Pending + if err := this.module.model.saveReport(rec); err != nil { + this.module.Errorf("报告 id:%d 重置为待生成失败 err:%v", rec.Id, err) + return + } + this.module.Warnf("报告 id:%d 上次失败,第 %d 次重试", rec.Id, n+1) + this.enqueueOnce(rec.Id) +} + +// 重试次数记在 error_msg 的 "retry=N;" 前缀里。 +// +// 为什么不加一列:那要改 proto 重生成 pb(生成器版本那个坑)再改表结构, +// 而这个计数只在「失败之后」有意义、失败本来就要写 error_msg,寄在它前面代价最小。 +// 展示端(后台/日志)看到的就是 "retry=2;统计失败: ...",一眼能看出重试过几次。 +const retryPrefix = "retry=" + +func retryCount(errMsg string) int { + if !strings.HasPrefix(errMsg, retryPrefix) { + return 0 + } + i := strings.Index(errMsg, ";") + if i < 0 { + return 0 + } + n, err := strconv.Atoi(errMsg[len(retryPrefix):i]) + if err != nil || n < 0 { + return 0 + } + return n +} + +// withRetryCount 换掉(或加上)前缀,保留原来的错误正文。 +func withRetryCount(errMsg string, n int) string { + body := errMsg + if strings.HasPrefix(errMsg, retryPrefix) { + if i := strings.Index(errMsg, ";"); i >= 0 { + body = errMsg[i+1:] + } + } + return fmt.Sprintf("%s%d;%s", retryPrefix, n, body) +} + // ===== worker ===== func (this *reportComp) consume(idx int) { @@ -203,6 +382,21 @@ func (this *reportComp) process(id uint64) { } }() + // 同一份报告同一时刻只许一个 worker 真的去跑。 + // cron 建单入队、ensureRecent 补偿入队两条路撞上时,没有这把锁就是两次 LLM 调用, + // 后完成的覆盖先完成的。600s 覆盖一次生成(LLM 超时 60s),拿不到就直接返回—— + // 说明别人正在跑,本次不用做任何事。 + lock := this.procLockKey(id) + ok, err := redissys.Conn().SetNX(this.ctx, lock, "1", 600*time.Second).Result() + if err != nil { + this.module.Warnf("报告 id:%d 抢生成锁失败,仍继续生成: %v", id, err) + } else if !ok { + this.module.Infof("报告 id:%d 正在别处生成,跳过", id) + return + } else { + defer redissys.Conn().Del(this.ctx, lock) + } + rec, err := this.module.model.getReport(id) if err != nil || rec == nil { this.module.Errorf("报告生成取记录失败 id:%d err:%v", id, err) @@ -263,6 +457,11 @@ func (this *reportComp) fail(rec *pb.DBMemoryReport, msg string) { if len([]rune(msg)) > 400 { msg = string([]rune(msg)[:400]) } + // ⚠️ 保住已有的 "retry=N;" 前缀:直接覆盖会把计数清零, + // 于是一份永远失败的报告可以无限重试下去。 + if n := retryCount(rec.ErrorMsg); n > 0 { + msg = withRetryCount(msg, n) + } rec.ErrorMsg = msg _ = this.module.model.saveReport(rec) this.module.Errorf("报告生成失败 id:%d uid:%s %s", rec.Id, rec.Uid, msg) diff --git a/apps/services/modules/memory/report_test.go b/apps/services/modules/memory/report_test.go new file mode 100644 index 00000000..379dabc3 --- /dev/null +++ b/apps/services/modules/memory/report_test.go @@ -0,0 +1,120 @@ +package memory + +import ( + "testing" + "time" + + "yunyan/comm" + "yunyan/pb" +) + +// 补偿要找的是「最近一个**已结束**的周期」,口径必须与 cron 的 tick 完全一致: +// 差一周就会去补一个 cron 已经建过的周期(重复生成)或者永远差一格(永远补不上)。 +func TestLastFinishedPeriods(t *testing.T) { + loc := time.FixedZone("CST", 8*3600) + cases := []struct { + name string + now time.Time + wantWeek string // 周一~周日 + wantMonth string + }{ + { + name: "周中", + now: time.Date(2026, 9, 18, 10, 0, 0, 0, loc), // 周五 + wantWeek: "2026-09-07~2026-09-13", + wantMonth: "2026-08-01~2026-08-31", + }, + { + name: "周一 cron 触发那一刻", + now: time.Date(2026, 9, 14, 0, 30, 0, 0, loc), + wantWeek: "2026-09-07~2026-09-13", + wantMonth: "2026-08-01~2026-08-31", + }, + { + name: "周日(本周还没结束,仍看上一周)", + now: time.Date(2026, 9, 20, 23, 59, 0, 0, loc), + wantWeek: "2026-09-07~2026-09-13", + wantMonth: "2026-08-01~2026-08-31", + }, + { + name: "跨年:1 月 1 日的上个月是去年 12 月", + now: time.Date(2026, 1, 1, 0, 30, 0, 0, loc), + wantMonth: "2025-12-01~2025-12-31", + }, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if c.wantWeek != "" { + w := lastFinishedWeek(c.now) + if got := w.start + "~" + w.end; got != c.wantWeek { + t.Errorf("上周范围应为 %s,得到 %s", c.wantWeek, got) + } + if w.ptype != comm.MemoryPeriodWeek { + t.Errorf("周期类型应为 week,得到 %s", w.ptype) + } + } + m := lastFinishedMonth(c.now) + if got := m.start + "~" + m.end; got != c.wantMonth { + t.Errorf("上月范围应为 %s,得到 %s", c.wantMonth, got) + } + if m.ptype != comm.MemoryPeriodMonth { + t.Errorf("周期类型应为 month,得到 %s", m.ptype) + } + }) + } +} + +// 周一 00:30 的 cron 与 ensureRecent 必须算出同一个周期键, +// 否则一个建单、另一个再建一次,同一周会有两份报告在跑。 +func TestLastFinishedWeekMatchesCronKey(t *testing.T) { + loc := time.FixedZone("CST", 8*3600) + now := time.Date(2026, 9, 14, 0, 30, 0, 0, loc) // 周一,cron 触发时刻 + cronKey := comm.MemoryWeekKey(now.AddDate(0, 0, -1)) + if got := lastFinishedWeek(now).key; got != cronKey { + t.Errorf("ensureRecent 算出 %s,cron 算出 %s,两者必须一致", got, cronKey) + } +} + +// 重试次数寄存在 error_msg 的 "retry=N;" 前缀里(不加列)。 +// 计数丢了的后果是一份永远失败的报告可以无限重试,每次都烧一次 LLM。 +func TestRetryCountRoundTrip(t *testing.T) { + cases := []struct { + msg string + want int + }{ + {"", 0}, + {"统计失败: connection refused", 0}, + {"retry=1;统计失败", 1}, + {"retry=3;统计失败", 3}, + {"retry=;坏格式当没重试过", 0}, + {"retry=abc;坏格式当没重试过", 0}, + {"retry=2没有分号也当没重试过", 0}, + } + for _, c := range cases { + if got := retryCount(c.msg); got != c.want { + t.Errorf("retryCount(%q) = %d,期望 %d", c.msg, got, c.want) + } + } + + // 加计数要保住原来的错误正文,换计数不能越叠越长 + got := withRetryCount("统计失败: x", 1) + if got != "retry=1;统计失败: x" { + t.Errorf("首次加计数得到 %q", got) + } + if got = withRetryCount(got, 2); got != "retry=2;统计失败: x" { + t.Errorf("再次加计数得到 %q(正文应原样保留、前缀不叠加)", got) + } +} + +// fail 覆盖 error_msg 时必须保住已有计数——不然重试上限形同虚设。 +func TestFailKeepsRetryCount(t *testing.T) { + rec := &pb.DBMemoryReport{ErrorMsg: "retry=2;上一次的原因"} + // 只验计数的保留逻辑本身,不走 fail(它要写库) + msg := "这次的原因" + if n := retryCount(rec.ErrorMsg); n > 0 { + msg = withRetryCount(msg, n) + } + if retryCount(msg) != 2 { + t.Errorf("覆盖错误正文后计数应仍为 2,得到 %q", msg) + } +} diff --git a/apps/services/modules/memory/stats.go b/apps/services/modules/memory/stats.go index 3fa26c6c..c3df85f4 100644 --- a/apps/services/modules/memory/stats.go +++ b/apps/services/modules/memory/stats.go @@ -54,8 +54,40 @@ func (s *memoryStat) hasContent() bool { return s.TodoTotal > 0 || s.AlarmTotal > 0 || s.IdeaCount > 0 || len(s.Expenses) > 0 } +// countRepeatAlarmExtras 重复闹钟在周期内比「库里那一行」多出来的次数。 +// +// 日期解析不了的项跳过(happen_date 是 date 列,正常不会出现,但别为一行坏数据 +// 让整份报告失败)。周期最长一个月,逐日判定至多 31 次,成本可以忽略。 +func (this *Memory) countRepeatAlarmExtras(uid, start, end string) (int, error) { + s, ok1 := comm.ParseMemoryDate(start) + e, ok2 := comm.ParseMemoryDate(end) + if !ok1 || !ok2 { + return 0, nil + } + items := make([]*pb.DBMemoryItem, 0) + if err := comm.ApplyRepeatCandidates(mysql.Table(comm.TableMemoryItem), uid, end, + []string{comm.MemoryCatAlarm}).Find(&items).Error; err != nil { + return 0, err + } + total := 0 + for _, it := range items { + anchor, ok := comm.ParseMemoryDate(it.HappenDate) + if !ok { + continue + } + total += comm.MemoryExtraOccurrences(it.RepeatRule, it.Weekday, anchor, s, e) + } + return total, nil +} + // computeStats 按日期范围算统计(含 start 与 end 当天) func (this *Memory) computeStats(uid, start, end string) (*memoryStat, error) { + // ⚠️ 入参可能来自 memory_report 读回来的 period_start/period_end,那是 `type:date` + // 列 + parseTime=True,读成 string 是 `2026-09-07T00:00:00+08:00`。 + // 下面的 SQL 靠 MySQL 转换还能对,但 Go 侧解析(重复项展开)会静默失败, + // 所以在入口先收敛一次;顺带让 stat_json 里的周期也是干净的日期, + // 那份 JSON 是要喂给 LLM 写播报文案的。 + start, end = comm.NormalizeMemoryDate(start), comm.NormalizeMemoryDate(end) st := &memoryStat{PeriodStart: start, PeriodEnd: end, UndoneItems: []statBrief{}, IdeaItems: []statBrief{}, Expenses: []statExpense{}} @@ -94,6 +126,16 @@ func (this *Memory) computeStats(uid, start, end string) (*memoryStat, error) { } st.TodoUndone = st.TodoTotal - st.TodoDone + // 重复闹钟按周期内实际响了几次算(见 comm.MemoryExtraOccurrences): + // 一条 daily 闹钟在库里只有一行、锚点可能在几个月前,上面那句 group by 数不到它, + // 于是「这周被闹钟叫了多少次」恒为 0。只补闹钟,别的分类不展开,理由在 comm 里写了。 + if n, e := this.countRepeatAlarmExtras(uid, start, end); e != nil { + // 展开失败不算整体失败:宁可少算这一项,也不要让整份报告生成不出来 + this.Warnf("周期统计 uid:%s 展开重复闹钟失败已忽略: %v", uid, e) + } else { + st.AlarmTotal += int32(n) + } + // 花销按币种聚合。金额是 int64 分,SUM 不会有浮点误差。 expRows := make([]statExpense, 0) err = mysql.Table(comm.TableMemoryItem). @@ -121,7 +163,8 @@ func (this *Memory) computeStats(uid, start, end string) (*memoryStat, error) { undone = undone[:statItemLimit] } for _, it := range undone { - st.UndoneItems = append(st.UndoneItems, statBrief{Id: it.Id, Title: it.Title, HappenDate: it.HappenDate}) + st.UndoneItems = append(st.UndoneItems, statBrief{Id: it.Id, Title: it.Title, + HappenDate: comm.NormalizeMemoryDate(it.HappenDate)}) } // 明细:灵感 @@ -138,7 +181,8 @@ func (this *Memory) computeStats(uid, start, end string) (*memoryStat, error) { ideas = ideas[:statItemLimit] } for _, it := range ideas { - st.IdeaItems = append(st.IdeaItems, statBrief{Id: it.Id, Title: it.Title, HappenDate: it.HappenDate}) + st.IdeaItems = append(st.IdeaItems, statBrief{Id: it.Id, Title: it.Title, + HappenDate: comm.NormalizeMemoryDate(it.HappenDate)}) } return st, nil } diff --git a/deploy/app/confs/home.yaml.example b/deploy/app/confs/home.yaml.example index e1a9fd0d..457f152a 100644 --- a/deploy/app/confs/home.yaml.example +++ b/deploy/app/confs/home.yaml.example @@ -163,6 +163,13 @@ modules: report_cron_month: "0 30 0 1 * ?" #每月 1 号 00:30 生成上月报告 report_stale_sec: 900 #报告卡在待生成/生成中超过这么久,由 memory_getreport 补偿重新入队 meeting_extract: true #会议纪要生成后自动抽取待办(失败只记日志,不影响纪要) + # 内容门槛:低于任一阈值的录音**只出纪要、不抽待办**,也不进 MCP 的纪要检索。 + # 拦的是 6 秒测试通话、20 秒单人自述这类东西——模板强制有「待办事项」一节, + # 模型没真待办就会把说话人的愿望改写成行动项,抽进拾忆全是垃圾。 + # ⚠️ 两项都**允许配 0**(= 这一道不限),所以「想用默认值就把这两行注释掉」, + # 别写成 0。时长门槛与 MCP 检索共用 comm.MeetingMinSeconds,改这里不会改 MCP 那边。 + #meeting_extract_min_seconds: 60 #短于多少秒不抽取,默认 60 + #meeting_extract_min_chars: 120 #转写正文(不含说话人标签)少于多少字不抽取,默认 120 echomeet: #会议记录:识别/翻译/总结服务不再走 yaml——全部由后台「会议记录服务」编排 # (console 共享库 echomeet_orch + 服务池 svc_config)配置,运行时直读 + NATS 热重载。 # 部署身份与解密密钥由 .env 注入:ANALYZE_APP_NAME / ANALYZE_REGION / FIELD_ENCRYPT_KEY。 diff --git a/docs/语音纪要-拾忆-MCP链路审计与落地计划.md b/docs/语音纪要-拾忆-MCP链路审计与落地计划.md new file mode 100644 index 00000000..fa3b3580 --- /dev/null +++ b/docs/语音纪要-拾忆-MCP链路审计与落地计划.md @@ -0,0 +1,582 @@ +# 语音纪要 → 拾忆 → MCP 链路审计与落地计划 + +> 审计日期 2026-09-18,基于分支 `along` 工作区 + 阿龙测试环境(8.133.166.29)真实数据。 +> 状态:**三批全部落地并已部署阿龙测试环境验收通过(2026-09-18)**,代码未提交;客户端改动(3.5/3.6 那两处)尚未发包。\n> 执行时按「第 4 节 分批」逐批做,每批做完跑对应验证再进下一批。 + +## 1. 链路全貌 + +``` +录音/导入 ─→ echomeet_addrecord ─→ echomeet_starttask + │ ├─ 普通录音: Submit ASR ─→ 回调 / 客户端轮询 / cron 兜底 ─→ finishTranscribe(翻译) ─→ AI 队列 + │ └─ VOICETRANSLAT(带转写): TranslateProcess ─→ AI 队列(跳过识别;源=目标时连翻译也跳过) + └─ AIProcess: summary + overview 两路并行 ─→ Completed + └─ 第三路 extractMemoryTodos ─→ memory.ExtractMeetingTodos(再跑一次 LLM)─→ memory_item(source=meeting) + +拾忆读路径: memory_list / memory_today / memory_upcoming ─→ 客户端日历 + 本地系统通知 +MCP 只读: search_meeting_notes(纪要) / get_memory_items / get_memory_stats / get_memory_report +周报月报: cron(周一/1号 00:30, 容器时区) ─→ uidsWithItemsBetween(≥3条) ─→ Redis 队列 ─→ SQL 统计 + LLM 文案 + ─→ memory_report ─→ 客户端 memory_today 带出待确认标记 ─→ 弹窗播报 ─→ memory_confirmreport +``` + +关键文件: + +| 环节 | 文件 | +|---|---| +| 转写/总结流水线 | `apps/services/modules/echomeet/tasks.go`(`SubmitTranscribeTask` / `PollTranscribe` / `finishTranscribe` / `AIProcess` / `extractMemoryTodos`) | +| ASR 回调 | `apps/services/modules/echomeet/api_backcall.go`(字节)、`api_alibackcall.go`(阿里) | +| 带转写直接总结 | `apps/services/modules/echomeet/api_starttask.go` 的 `Rtype == "VOICETRANSLAT"` 分支 | +| 待办抽取 | `apps/services/modules/memory/meeting_extract.go` | +| 拾忆模型/接口 | `apps/services/modules/memory/{model,api_*,stats,remind,report}.go` | +| 跨模块公共 | `apps/services/comm/memory.go`、`comm/memoryquery.go`、`comm/module.go`(`IEchomeet` / `IMemory`) | +| MCP 工具 | `apps/services/modules/mcp/tool_meeting.go`、`tool_memory.go` | +| 客户端 | `apps/client/lib/data/services/memory_service.dart`、`lib/modules/main_tab/views/widgets/memory_{daily,report,edit}_sheet.dart`、`lib/modules/home/controllers/calendar_controller.dart`、`lib/modules/translation/models/translation_export.dart` | + +## 2. 发现清单 + +✅ = 测试库 / 日志里有实证;⚠️ = 读代码推出的,尚未在真机撞到。 + +### ① echomeet 按 id 操作的接口全部不校验归属(高,✅ 代码逐个核过) + +`getrecord / getrecords / modifyrecords / delrecords / readrecord / summary / starttask / uprecord` 八个接口只按 id 查记录,**没有一处比对 `record.Uid == session.GetUserId()`**(`starttask` 用了 `GetUserId()`,但只拿来扣算力和查 VIP)。 + +后果:任何登录用户改 id 就能读别人纪要全文、删改别人记录、给别人的记录发起总结(扣自己算力,但覆盖对方 summary 并触发对方的拾忆抽取)。 + +对照:`memory` 模块每个接口都有 `requireUID` + `rec.Uid != uid`,写法现成。回调接口(`backcall` / `alibackcall`)是第三方来的、无 session,不在此列。 + +### ② gorm `default:` 标签吞零值:`date_certain=false` / `remind_ahead=0` 永远写不进库(高,✅) + +`mysql.Insert` 就是 `db.Table().Create(model)`(`lego/sys/mysql/mysql.go:89`)。gorm 对带 `default:` 标签的字段,Create 时零值一律换成默认值。`DBMemoryItem.date_certain` 标了 `default:true`、`remind_ahead` 标了 `default:5`。 + +实证(`starpivot-app.memory_item`): + +``` +source date_certain remind_ahead n +assistant 1 5 9 +meeting 1 5 16 +``` + +`ExtractMeetingTodos` 明明写了 `DateCertain: false, RemindAhead: 0`,16 条会议待办**全部**落成 `1 / 5`。后果: + +- 客户端 `todo_scope_list` 的「待定日期」分组永远是空的,没截止日的待办全堆在会议当天; +- 客户端 `memory_upcoming` 按 `remind_ahead=5` 展开 → 未来日期的会议待办会在 08:55 响系统通知,与代码注释「会议待办默认不提醒」相反; +- 用户手动建闹钟选「不提醒」(`memory_edit_sheet` 把 `remindAhead` 置 0)同样被写成 5。 + +`api_update` 走 `Save`(写全字段),不受影响。 + +### ③ 没有内容门槛,短录音 / 单人自述照样进拾忆和 MCP(高,✅) + +- `memory_item id=26`:6 秒测试通话(`call_20260918_112659.wav`)总结成「继续进行进一步的功能测试」,随即抽成一条待办。 +- `memory_item id=23/24`(记录 530,20 秒单人语音):抽出「将本次语音录音准确转写为规范文字稿」「按'会议记录版本'格式整理输出」。**第一遍审计误判为模板 outline 泄漏**,拉出完整 summary 看,是说话人自己说的需求("提出需求:将测试语音导出为文字版或会议记录版本"),总结模板强制有「四、待办事项」一节,LLM 没有真待办就把这句愿望改写成行动项,抽取器再照单全收。根因与 id=26 同源:内容太薄 + 模板强制出待办节。 +- MCP `search_meeting_notes` 只挡 `CHAR_LENGTH(summary) >= 20`,这类「内容零散、按兜底规则输出」的废话总结有 600 多字,照样进检索结果、挤掉真会议。 +- 周报阈值 `report_min_items=3`:一次会议抽出 3 条废话就够触发一份周报(又一次 LLM 调用)。测试库 W37 那份周报正是靠 9 月 7 日那批会议待办凑够条数生成的。 + +`AIProcess` 与 `ExtractMeetingTodos` 都没有最短时长 / 最少字数判断。 + +### ④ 说话人标签三种格式并存(中,✅) + +| 写入点 | 格式 | +|---|---| +| `api_alibackcall.go` / `api_backcall.go` / `PollTranscribe` | `Speaker_0` | +| `tasks.go finishSyncTranscribe`(字节 flash 同步通道) | `Speaker 0`(空格) | +| `client translation_export.dart`(通话/同传/面对面归档,2026-09-18 新加) | `0` / `1` | + +LLM 拿到 `[0]:`、`[Speaker 2]`、`[Speaker_1]` 混着,owner 就写成 `Speaker2`、`[Speaker_0]` 各种样(`memory_item id=25 owner=Speaker2`、`id=12 owner=[Speaker_0]`)。MCP `meaningfulPersonnel` 只认 `Speaker_` 前缀,其余格式会被当成「用户命名过的真人」塞给模型。 + +### ⑤ 客户端上报的 `tz` 是缩写不是 IANA 名(中,✅) + +`MemoryService._tz()` 用 `DateTime.now().timeZoneName` → `"CST"`;服务端 `comm.LoadMemoryLocation("CST")` 解析失败**静默**退回容器时区。库里 `tz` 列只有空串和 `CST` 两种值。国内用户碰巧没事,海外用户的提醒时刻会按北京时间算,且不报任何错。`flutter_timezone` 已在 `pubspec.yaml` 里但一处都没用。 + +### ⑥ 删纪要不删待办(中,⚠️) + +`echomeet_delrecords` 不碰 `memory_item`,`source_id` 指向的记录没了,待办还挂在日历上,点进去找不到来源。 + +### ⑦ 周报/月报的卡单补偿是死的、cron 错过不补、失败不重试(高,⚠️) + +三个问题一个根: + +- `requeueIfStale` 只由 `memory_getreport` 触发;客户端只调无参版本 → 服务端走 `latestUnconfirmedReport`,它的 SQL **只查 `state=Done`** → pending / processing / failed 的报告永远不会被查出来,也就永远不会被重新入队。注释里写的「容器 00:30 重启一次就靠它救」实际救不了。 +- 周一 00:30 容器不在线,那一周连 `memory_report` 行都不会建;下周 tick 算的是下一个 `period_key`,不回头补。 +- `state=Failed` 后 `GetReport` 直接回错,下一次 tick 是新周期,旧的永远 Failed。 + +### ⑧ 重复项在周期统计里恒为 1 次(中,⚠️) + +`computeStats` 与 MCP `get_memory_stats` 按 `happen_date` 落在周期内计数,一条 `daily` 闹钟只有一行、锚点在几个月前 → 本周 `alarm_total` 里它是 0。 + +### ⑨ ASR 回调与轮询可能撞车,回调不加锁(中,⚠️ 潜在) + +`AliBackCall` / `BackCall` 直接同步 `TranslateProcess + SubmitAITask`,不检查 `finishing` 锁、不判 `State`;`PollTranscribe` 同一时刻若也拿到 Success,两边各翻译一遍、各 `LPush` 一次 → `AIProcess` 跑两遍(两次 LLM 计费,第二次覆盖第一次,`extractMemoryTodos` 也跑两轮)。`SubmitAITask` 入队前没有 `inQueue` 判重(`sweepStuckSummarize` 有,它没有)。72 小时日志里回调命中 1 次、未观察到重复完成;转写文件越长(回调越晚、轮询越多次)越容易撞。 + +### 没问题、不用动的 + +- VOICETRANSLAT 分支(带转写直接总结):跳识别、源=目标跳翻译、`AIProcess` 用 `Translate` 退回 `Original`、`extractMemoryTodos` 照常触发,2026-09-18 真机 id=555 已验证。 +- MCP 三个工具:uid 从 JWT 解、只读、按字符截断、全文索引失败退 LIKE。 +- 周/月边界、ISO 周键、闰月范围、周日归属、容器时区口径:有单测且正确。`upsertReport` 已确认不重算、`ConfirmReport` 有归属校验、空报告自动确认。 +- 会议归属日期用 `creationtime`(上传时刻):补录旧录音会偏,代码里已注明是已知偏差,本次不动。 + +## 3. 修改方案 + +每条按「改哪 / 怎么改 / 边界 / 验证」写。原则:**不改 proto、不重生成 pb、不改表结构、存量修正幂等、用户可见行为不变**。 + +### 3.1 归属校验(对应 ①) + +**改哪**:`apps/services/modules/echomeet/model.go` + 八个 `api_*.go`。 + +**怎么改**: + +```go +// model.go +func (this *modelComp) getrecordforuid(uid string, id uint64) (*pb.DBEchoMeetRecord, error) // WHERE uid=? AND id=? +func (this *modelComp) getrecordsforuidids(uid string, ids []uint64) ([]*pb.DBEchoMeetRecord, error) // WHERE uid=? AND id IN ? +func (this *modelComp) delrecordsforuid(uid string, ids []uint64) error // DELETE WHERE uid=? AND id IN ? +``` + +八个接口改调上面三个;查不到统一回 `ErrorCode_DBError`「记录不存在」——**不要区分「不存在」和「不是你的」**,避免被用来探测 id。 + +**边界**: + +- `getrecords` 批量里混了别人的 id → 只返回自己的那部分,**不报错**。客户端轮询队列里可能有已删记录,报错会掀掉整个轮询(`MeetingTaskService._executeTask` 那个坑)。 +- `delrecords` 同理按交集删,响应里 `Ids` 回实际删掉的。 +- `starttask` / `summary` 在归属校验**之后**再扣算力,否则拒绝了还扣钱。 +- 回调接口不动。 + +**验证**: + +- 源码扫描单测 `TestApisUseUidScopedQueries`:`api_*.go`(回调两个除外)里不许出现 + `model.getrecord(` / `getrecords(` / `delrecords(`。比模型级单测更值:最容易复发的方式 + 不是有人改回去,而是**新加一个 api 文件时照旧写法抄一遍**,那样既不报错也看不出来。 + 模型级单测要连真库,本仓库没有测试库,跳过。 +- 真机:用 iPhone 账号(`2095078768153460736`)的 token 请求安卓账号(`2095032531807109120`)的记录 559,改前 200,改后回错。 + +### 3.2 插入后补写 + 存量修正(对应 ②) + +**改哪**:`apps/services/modules/memory/model.go` 的 `addItem`,加 `migrate_defaults.go`。 + +⚠️ **本节方案在执行时推翻重写过一次。** 原计划写的是给 Create 加 `Select("*")` +强制写全字段——**那是错的,实测无效**。gorm 的零值→默认值替换发生在 +`callbacks.ConvertToCreateValues` 的 `reflect.Struct` 分支里: + +```go +if values.Values[0][idx], isZero = field.ValueOf(ctx, stmt.ReflectValue); isZero { + if field.DefaultValueInterface != nil { + values.Values[0][idx] = field.DefaultValueInterface // ← 无条件替换 +``` + +判据只有「字段值是不是零值」,**完全不看 Select / Omit**。DryRun 实测两种写法 +生成的 VALUES 一模一样(`date_certain=true, remind_ahead=5`),这条结论已写成 +测试 `TestGormCreateSwallowsZeroValues` 钉住,免得下次又有人去试 `Select("*")`。 + +**实际怎么改**:插完把那两列按结构体的现值补写回去。 + +```go +func (this *modelComp) addItem(item *pb.DBMemoryItem) error { + ... + // gorm 替换零值时连结构体上的字段一起改(field.Set),所以先留一份真值: + // 不还原的话 memory_add 回给客户端的 item 也是被改过的 + dateCertain, remindAhead := item.DateCertain, item.RemindAhead + if err := mysql.Insert(comm.TableMemoryItem, item); err != nil { return err } + fix := map[string]interface{}{} + if !dateCertain { fix["date_certain"] = false; item.DateCertain = false } + if remindAhead == 0 { fix["remind_ahead"] = int32(0); item.RemindAhead = 0 } + if len(fix) == 0 { return nil } + // 补写失败只告警不回错:行已经插进去了,回错会让 memory_add 对客户端报 + // 「新增失败」,而那条记忆项其实在库里 + if err := mysql.Table(...).Where("id = ?", item.Id).Updates(fix).Error; err != nil { + this.module.Warnf(...) + } + return nil +} +``` + +一处改完,`ExtractMeetingTodos`、`api_add`、`migrate_task` 三个写入方一起修好。 + +**为什么不去掉 `default:` 标签**:要改 proto → 重生成 `memory_db.pb.go`(生成器版本 v1.36.6 的坑) +→ 且 MySQL 列上已建出来的 DEFAULT 不会跟着变,收益为零。 + +**为什么不改成 map 形式的 Create**(`ConvertMapToValuesForCreate` 确实不做替换): +那要手抄 27 个列名,proto 加字段时静默漏写,风险比这一条 UPDATE 大得多。 + +**怎么防漏**:受影响的只有「默认值非零」的列(`default:0` / `default:false` 替换了还是零值,无害)。 +`TestGormNonZeroDefaultsAreHandled` 用反射扫 pb 结构体的 tag,与 `gormZeroDefaultColumns` +逐字比对;proto 里再加一个非零 default 而 `addItem` 没跟上,这条测试会红。 +`DBMemoryReport` 另有一条测试确认它目前没有这类列(它也走 `mysql.Insert`)。 + +**存量修正**(`migrate_defaults.go`,memory 模块 `Start()` 里 `go` 跑,幂等,失败只记日志): + +```sql +UPDATE memory_item SET remind_ahead = 0 + WHERE source = 'meeting' AND user_edited = 0 AND remind_ahead <> 0; +UPDATE memory_item SET date_certain = 0 + WHERE source = 'meeting' AND user_edited = 0 AND due_raw <> '' AND date_certain = 1; +``` + +`date_certain` 无法百分百还原(当初真推算出日期的那些本来就该是 1),只把 +「有 due_raw 却仍落在会议当天」这个能确定判错的子集回退成待定。**只碰 `user_edited = 0`**。 +助手闹钟那 9 条 `remind_ahead=5` 不动——当初用户是不是选了「不提醒」已无从知道, +且 5 分钟是产品默认值,保守。 + +**不写 `next_remind_at`**:全项目没有任何地方读它(只有 `recalcRemind` 在写, +grep 过一遍),动它只会扩大修正的影响面。 + +**边界**:`normalizeItem` 里 `HappenDate == ""` 时置 `DateCertain = false` 这条现在才真正生效; +客户端 `todo_scope_list.dart:248` 已有 `!dateCertain → 待定` 分组,不用改。 + +**验证**: + +- 单测(已加):`TestGormCreateSwallowsZeroValues` / `TestGormNonZeroDefaultsAreHandled`。 +- 真机:跑一次会议总结后查库;手动建闹钟选「不提醒」后查库 `remind_ahead=0`; + `memory_upcoming` 不再返回会议待办的 08:55 槽位。 + +### 3.3 内容门槛(对应 ③) + +分两层,**刻意不改用户可见的总结行为**:短语音备忘("记一下明天买菜")是合法用法,总结照出,只是不该进拾忆和 MCP。也不把短录音判 `SummarizFail`——客户端把 10001 显示成「总结失败,请稍后重试」,用户会反复点。 + +**3.3.1 抽取门槛**(`apps/services/modules/echomeet/tasks.go` `extractMemoryTodos`) + +```go +// 有效字符:去掉 "[Speaker_x]:" 标签后的 rune 数 +if rec.Seconds < opts.MeetingExtractMinSeconds || effectiveRunes(rec) < opts.MeetingExtractMinChars { + this.module.Infof("extractMemoryTodos id:%d 内容太薄(%ds/%d字),跳过抽取", ...) + return +} +``` + +阈值放 `memory` 的 `Options`:`meeting_extract_min_seconds`(默认 60)、`meeting_extract_min_chars`(默认 120),`deploy/app/confs/home.yaml.example` 同步加注释。两项都**允许配 0**(=这一道不限),所以回默认值的判据是「配置里没有这个键」而不是「值 <= 0」。 + +门槛判断放在 `memory.ExtractMeetingTodos` 入口(阈值在它自己的 Options 里),判定本身抽成纯函数 `meetingContentEnough(seconds, runes, minSec, minChars) (bool, string)` 方便单测,日志由调用方打。 + +⚠️ **字数必须数转写正文,不能数纪要**:内容极薄时模型按模板兜底规则照样写出六百多字(id=23/24/26 就是),卡纪要字数一条也拦不住。正文字数由 echomeet 侧的 `transcriptRunes(rec)` 算(口径与 `AIProcess` 拼 `originalText` 一致:优先译文、空则退回原文,只数 `Content` 不数 `[Speaker_x]` 标签)。 + +⚠️ **`IMemory.ExtractMeetingTodos` 改成收一个 `comm.MeetingExtractInput` 结构体**,不是原计划的「再加一个 `seconds` 位置参数」——那样会变成 ctx + 5 个字符串 + 2 个数字,把 `summary` 和 `meetingDate` 写反编译器一句话都不会说。 + +⚠️ **数据缺失(seconds / runes 为 0)按通过处理**:宁可多抽几条用户能自己删的待办,也不要因为上游漏传一个字段就把所有会议的抽取静默关掉。 + +**3.3.2 抽取 prompt 加两条**(`meeting_extract.go` `meetingExtractPrompt`) + +``` +8. 纪要若注明「非真实会议」「单人测试」「无待办」「不构成会议场景」之类,直接输出 []。 +9. 说话人对自己的愿望、需求、期望的描述(如「希望导出成文字」「想评估一下效果」)不是待办; + 只收录会议中明确指派给某人 / 某方去执行的事。 +``` + +**3.3.3 MCP**(`tool_meeting.go`):加 `Where("seconds >= ?", comm.MeetingMinSeconds)`。`seconds` 列现成,不用新索引。 + +⚠️ 阈值提成 `comm.MeetingMinSeconds` 常量:mcp 是独立进程,读不到 memory 的配置,两边各写一个 60 迟早会漂。 + +⚠️ **只改 `recentQuery` 不够**:`Handl` 里另拼了一份一模一样的过滤条件给「无关键词」那条路径用,加在 `recentQuery` 上的门槛对它不生效。已把那份重复条件删掉,四条检索路径(无关键词 / 全文 / LIKE / 回退)统一从 `recentQuery` 起手。 + +**3.3.4 周报阈值**:不动。垃圾待办不再产生,`report_min_items=3` 自然凑不够。 + +**验证**: + +- `meeting_extract_test.go` 已加 `TestMeetingContentEnough`(8 个用例,含两条「数据缺失不误伤」)与 `TestOptionsExtractThresholdDefaults`(默认值 vs 显式配 0);`echomeet` 侧加 `TestTranscriptRunes`。prompt 第 8/9 条的效果只能真机看,没有 fake ChatLLM。 +- 真机:拿 559(6 秒)点「重新生成」,`memory_item` 不应新增;60 秒以上正常会议照旧抽取;MCP 问「最近的会议」不再返回 559。 + +### 3.4 报告补偿(对应 ⑦) + +**改哪**:`apps/services/modules/memory/report.go` 新增 `ensureRecent(uid)`,`api_today.go` 调它。 + +**怎么改**: + +```go +// ensureRecent 用户回来时(memory_today)补两件事:最近一个已结束的周 + 最近一个已结束的月。 +// 只看这两个周期,不追更早的——与「只弹最近一份」口径一致,也避免半年没开 App 的用户一回来 +// 就建二十份报告。 +func (this *reportComp) ensureRecent(uid string) { + now := time.Now() + for _, p := range []period{lastWeek(now), lastMonth(now)} { + rec, err := this.module.model.getReportByPeriod(uid, p.ptype, p.key) + if err != nil { continue } + switch { + case rec == nil: + // cron 错过了(容器当时不在线):按同一道闸门补建 + if n := this.module.model.countItemsBetween(uid, p.start, p.end); n < this.options.ReportMinItems { continue } + rec = &pb.DBMemoryReport{Uid: uid, PeriodType: p.ptype, PeriodKey: p.key, PeriodStart: p.start, PeriodEnd: p.end} + if this.module.model.upsertReport(rec) == nil { this.enqueueOnce(rec.Id) } + case rec.State == Pending || rec.State == Processing: + if now.Unix()-rec.UpdateTime > this.options.ReportStaleSec { this.enqueueOnce(rec.Id) } + case rec.State == Failed && !rec.Confirmed: + if now.Unix()-rec.UpdateTime > this.options.ReportStaleSec && retryCount(rec) < 3 { + rec.State = Pending; bumpRetry(rec); _ = this.module.model.saveReport(rec); this.enqueueOnce(rec.Id) + } + } + } +} +``` + +落地时相对上面这版草稿的调整: + +- 拆成 `ensureRecent` / `ensureBuilt` / `retryFailed` 三个小函数,`pending/processing` + 那支**直接复用 `requeueIfStale`**,不另写一遍超时判定(两份阈值迟早会漂)。 +- 周期计算抽成 `lastFinishedWeek` / `lastFinishedMonth`,`TestLastFinishedWeekMatchesCronKey` + 钉住它与 cron 的 `tick` 算出同一个 `period_key`——差一格就会重复建单或永远补不上。 +- `enqueueOnce` = `LPos` 判重 + `LPush`(`echomeet sweepStuckSummarize` 的做法)。 +- `process()` 开头加 `SETNX memory:report:proc: EX 600`(经 `redissys.RKey` 加应用前缀),cron 与 ensure 撞上也只跑一次 LLM;拿不到锁直接返回。 +- Failed 重试上限 3 次:`error_msg` 前缀记 `retry=N;`,`retryCount/withRetryCount` 解析它,不加列。 + ⚠️ **`fail()` 也必须保住这个前缀**——它原本直接覆盖 `error_msg`,计数会被清零, + 于是一份永远失败的报告可以无限重试,上限形同虚设。`TestFailKeepsRetryCount` 守着。 +- `api_today.go`:`go this.module.report.ensureRecent(uid)`,整个包在 goroutine 里、带 recover,**不阻塞** today 响应。 +- `requeueIfStale` 保留(对带 `period_key` 的查询仍有用),`latestUnconfirmedReport` 不改。 + +**边界**:`memory_today` 每次切前台都调,`ensureRecent` 只多两次主键查询 + 至多一次 count;`UpdateTime` 节流保证同一份报告至多每 `ReportStaleSec`(900s) 重试一次。`countItemsBetween` 复用 `uidsWithItemsBetween` 的条件(`happen_date` 范围),单 uid 版。 + +**验证**: + +- 把测试库一条报告改成 `state=0, update_time=0`,切一次前台,900s 内变 Done。 +- 把 W37 那条改成 `state=3, confirmed=0, update_time=0`,切前台后重新生成;连改三次后第四次不再入队。 +- 手动删掉上周的报告行,切前台后(该用户上周 ≥3 条记忆项时)重新出现。 + +### 3.5 说话人标签统一(对应 ④) + +**改哪**: + +| 文件 | 改动 | +|---|---| +| `apps/services/modules/echomeet/tasks.go` `finishSyncTranscribe` | `fmt.Sprintf("Speaker %s", ...)` → `"Speaker_%s"` | +| `apps/client/lib/modules/translation/models/translation_export.dart` `_speakerOf` | `'0'/'1'` → `'Speaker_0'/'Speaker_1'`;`test/translation_export_test.dart` 两条断言同步改 | + +`Personnel` 拼接随之一致。存量数据不动(客户端说话人重命名功能会把它们改掉)。 + +落地时的两点调整: + +- 服务端两处(`finishSyncTranscribe` 与 `PollTranscribe`)原本各写了一遍同样的循环、 + 只有前缀不同,已抽成 `normalizeSpeakers(rec, contexts)` 一个函数,从根上不会再漂。 +- 客户端那两个取值提成常量 `kSpeakerSelf` / `kSpeakerPeer`。 + ⚠️ 这条**有用户可见变化**:转写页(`speech_tab.dart:76`)直接把 speaker 字段当文案显示, + 翻译归档来的记录以前显示成裸 `0` / `1`,现在与其它纪要一样显示 `Speaker_0`。 + 这是把它**改对**,不是引入不一致。 + +**验证**:`meeting_extract_test.go` 加一条:输入含 `[Speaker_0]:` 的纪要,owner 输出不带方括号;`tool_meeting.go` 的 `meaningfulPersonnel("[Speaker_0] [Speaker_1] ")` 返回空。 + +### 3.6 时区(对应 ⑤) + +**客户端** `memory_service.dart`: + +```dart +static String? _ianaTz; // 启动时取一次,进程内缓存 +static Future warmTz() async { _ianaTz = await FlutterTimezone.getLocalTimezone(); } +String _tz() => _ianaTz ?? DateTime.now().timeZoneName; // 拿不到再退缩写 +``` + +`warmTz()` 挂在 `MemoryService.onInit`(`unawaited`)。 + +**服务端** `comm/memory.go` `LoadMemoryLocation`: + +- 解析失败打一条 warn(现在静默); +- 顺手认固定偏移:`UTC+8` / `GMT+08:00` / `+08:00` → `time.FixedZone`;`CST` 歧义(中国/美国中部都叫 CST),**按 Asia/Shanghai 解释**并注明——把已落库的 9 条兜住。 + +⚠️ `flutter_timezone` 3.0.1 的 `getLocalTimezone()` 返回的是 `Future`, +不是带 `.identifier` 的对象——照新版 API 写会编译不过。 + +**验证**:新建一条闹钟查库 `tz='Asia/Shanghai'`;服务端 `TestLoadMemoryLocation` +11 个用例(IANA / CST / UTC+8 / GMT+08:00 / +08:00 / -0530 / 认不出来退本地而非 UTC)。 + +### 3.7 删纪要连带删待办(对应 ⑥) + +**改哪**:`comm/module.go` `IMemory` 加 `DeleteMeetingItems(uid string, recordIDs []string) (int64, error)`;`memory/meeting_extract.go` 实现;`echomeet/api_delrecords.go` 在归属校验后 `go` 调它(失败只记日志)。 + +**只删 `user_edited = 0` 的**——用户改过的待办已经是他自己的东西,来源没了也该留着。 + +⚠️ model 层那条 `delMeetingItemsBySource` **带 uid 一起查**:入参来自 echomeet 的删除请求, +光信 `source_id` 等于把归属校验的成果又丢掉一次。 + +**验证**:删一条带自动待办的纪要,`memory_item` 里对应 `source_id` 且 `user_edited=0` 的行消失,`user_edited=1` 的保留。 + +### 3.8 回调幂等 + 入队判重(对应 ⑨) + +**改哪**:`api_backcall.go` / `api_alibackcall.go` 开头;`tasks.go SubmitAITask`。 + +```go +// 回调重放 / 与轮询撞车:状态已经不是 Transcribing 说明别的路径已经收尾,直接 ack +if model.State != pb.DBEchoMeetRecordState_Transcribing { resp = ...; return } +if _, busy := this.module.tasks.finishing.LoadOrStore(model.Id, struct{}{}); busy { resp = ...; return } +defer this.module.tasks.finishing.Delete(model.Id) +``` + +`SubmitAITask`:`LPush` 前 `inQueue(awaitKey)` / `inQueue(procKey)` 判重,已在队列里直接返回。 + +落地时相对草稿的三点调整: + +- 抢锁逻辑抽成 `beginTranscribeFinish` / `endTranscribeFinish` 一对函数, + 状态判定与内存锁合在一处,两个回调各一行。放行的状态是 + `Transcribing` **与 `AwaitTranscribing`**——提交成功到把状态写成 Transcribing + 之间有个窗口,快的 provider 能在这中间就把回调打回来,只认 Transcribing 会把 + 合法回调丢掉。 +- ⚠️ **失败回调也必须过这一关**(草稿里漏了):状态检查原本只挡成功分支的话, + 一条已经 Completed 的记录被一个迟到的失败回调打回 `TranscribeFail`, + 用户看到的就是「纪要好端端地变成了转写失败」。现在两个分支都在闸门之后。 +- 挡下来时**仍回 SUCCESS**:对第三方来说这次通知已经正确接收,回错只会让它 + 按失败重投,重投的还是同一条处理过的结果。 +- `SubmitAITask` 不改签名、**不判重**,另加一个 `submitAITaskOnce` 给自动路径 + (两个回调 + 两处转写收尾)用。用户主动发起的「重新生成」(`Summary` / `StartTask`) + 继续走不判重的那个:悄悄跳过等于按钮失灵,比多跑一次糟得多。 + +顺带把两个回调里各自手拼的 speaker / Personnel 也收进 `normalizeSpeakers` +(3.5 那个函数),修掉一个边角:没开说话人分离时编号是空串,原来会拼出 `Speaker_` 这种空尾巴。 + +**验证**:`TestBeginTranscribeFinish` 覆盖重入、AwaitTranscribing 放行、 +六种已推进状态一律拒绝、nil 不 panic。 + +### 3.9 重复项统计展开(对应 ⑧,可后做) + +`computeStats` 与 `tool_memory.go get_memory_stats`:对 `repeat_rule != once` 的项,用 `comm.MemoryRepeatHits` 在周期内逐日展开计数。 + +落地时收窄了范围,**只展开闹钟**: + +- 花销展开等于**虚构金额**。「每月房租」展开进「你这周花了多少」,报出去的是用户 + 根本没记过的数字,而这份统计的全部价值就在数字是准的(本文档开头就写着「金额与 + 完成数走 SQL,LLM 只组织文案」)。 +- 待办的 done/undone 按次拆不开:一行只有一个 state,把一条「每周复盘」的重复待办 + 按 7 次都算成已完成,完成率就成了假的。 +- 灵感本来就不该重复。 + +实现放 `comm`(`MemoryExtraOccurrences` + `ApplyRepeatCandidates`),home 与 mcp 共用—— +各写一份会漂成「App 里的周报说 7 次、EMAI 说 1 次」,而且两边都不报错。 + +⚠️ 返回的是**增量**(除库里那一行之外还发生了几次)不是总次数:调用方的 group by +已经把锚点落在周期内的那一行数过一次,返回总数就会重复计。同时守住不返回负数—— +规则在本周期一次都没命中、而那一行恰好落在周期内时,扣锚点会把调用方本来数对的那行抹掉。 + +**验证**:`TestMemoryExtraOccurrences` 12 个用例(一次性/每天/每周/工作日/周末/每月、 +锚点在周期前/中/后、以及那条不能返回 -1 的边界)。 + +## 4. 落地顺序与分批 + +| 批 | 内容 | 特点 | 部署 | +|---|---|---|---| +| **1** | 3.1 归属校验、3.2 `Select("*")` + 存量修正、3.3 内容门槛 | 有实证、改动小、不改协议不改表 | 只部服务端(`dev-deploy.sh along`),客户端不用发包 | +| **2** | 3.4 报告补偿、3.5 标签统一、3.6 时区、3.7 删纪要连带 | 涉及 `IMemory` 签名与客户端两处 | 服务端 + 客户端同批 | +| **3** | 3.8 回调幂等、3.9 重复项统计 | 潜在问题,独立验证 | 只部服务端 | + +每批做完:`go build ./... && go vet ./... && go test ./modules/echomeet/ ./modules/memory/ ./comm/`;客户端改动跑 `flutter test`。 + +## 4.1 与「语音纪要列表骨架化」改动的交叉检查(2026-09-18 当天) + +同一天上午把列表接口改成了骨架查询(`getrecordbriefforuid`,Omit 掉 +original/translate/summary)+ 客户端本地优先缓存。逐条核过与本计划的交叉点: + +| 交叉点 | 结论 | +|---|---| +| 归属校验 vs 骨架列表 | `GetAllRecords` 本来就按 uid 查;顺手补了 `requireUID` | +| `delrecords` 改回「实际删掉的 id」 | 客户端只有 `MeetingTaskService.removeTask` 一个调用方,且不消费响应 | +| `GetRecords` 静默过滤非本人 id | 客户端按响应里的记录逐条更新,缺的那条留在轮询队列里;已被 `Synchrodata` 的孤儿清理 + `removeLocalTask` 覆盖 | +| MCP 的 `seconds >= 60` | 翻译归档走 `RecordingArchive` 时带真实 `seconds`,不会被误伤 | +| 说话人标签统一 | 同步写的是骨架字段,不碰 translate;详情页单条拉全字段,不受影响 | + +**查出一个真问题,已修**: + +`GetAllRecords` 对转写中的记录会 `saverecord(骨架记录)`,而 `mysql.Save` 是**整行 UPDATE** —— +骨架里 original/translate/summary 是空串,写回去就把这三列清空了。触发路径不罕见: +`StartTask` 的闸门是 `state>0 && state<5`,**已完成(5)/已阅(6)/失败(10001,10002) 都允许 +「重新转写」**,那期间 `state=Transcribing` 而旧纪要还在库里,用户这时切一下列表页, +旧转写和旧总结就没了;这次转写若又失败,那是永久丢失。列表改成骨架查询之前这里拿的是 +全字段,不会有这个问题。 + +修法:轮询前用 `getrecordforuid(uid, brief.Id)` 重新取整条,拿整条去 save / PollTranscribe, +只把 `lastquerytime` 同步回骨架供本次响应序列化。`TestListEndpointNeverSavesBriefRecord` +守着不许回退。 + +**顺带一条与第 3 批有关的观察**:列表页现在也成了转写轮询的驱动方(回调 / 详情页轮询 / +列表页 / cron 兜底,四条),回调与轮询撞车的概率比审计时更高,3.8 的价值随之上升。 + +## 4.2 部署验证时挖出来的两个既有 bug(2026-09-18 真机) + +三批部到阿龙测试机后逐条验收,**3.9 死活不生效**,顺藤摸出一条影响面大得多的既有问题。 + +### 根因:`type:date` 列读进 Go 的 string 是 RFC3339,不是 `YYYY-MM-DD` + +`memory_item.happen_date`、`memory_report.period_start/period_end` 都是 MySQL 的 +`type:date` 列,而 DSN 带 `parseTime=True` —— gorm 把它们扫进 pb 结构体的 **string** +字段时,给的是 `2026-09-07T00:00:00+08:00`。 + +按 `happen_date` 比较的那几句 SQL 靠 MySQL 自己做类型转换还是对的,所以这个问题 +**在任何报错、任何日志里都看不见**,只在「Go 侧要解析这个值」的地方发作, +而且一律表现为「什么都没发生」: + +| 受害者 | 表现 | 实证 | +|---|---|---| +| `remind.go` 的 `expandItem` | **`memory_upcoming` 恒返回 0 个提醒时刻**,整个拾忆提醒在服务端空转 | 真机建了一条每天 08:00 的闹钟,`slots` 返回 `[]` | +| `stats.go` 的重复项展开(3.9) | 重复闹钟永远只算 1 次 | 一条 daily 闹钟在 7 天周期里 `alarm_total=1` | +| 下发给客户端的 `happen_date` | 协议说好是 `YYYY-MM-DD`,实际是 RFC3339 | `memory_list` 响应逐条都是 | + +第三行还有客户端后果:`MemoryService._localToday()` 用 +`e.happenDate == today` 做字符串相等 → 「今天」的离线兜底恒为空; +`_mergeRange` 用 `compareTo` 判区间 → 落在区间端点那天的项会被判成「段外」而重复。 +(`happenDateTime` 走 `DateTime.tryParse`,两种格式都认,所以日历渲染看不出问题—— +这也是它一直没被发现的原因。) + +⚠️ 这解释了审计报告 §2 ② 里那句「会议待办会在 08:55 响系统通知」为什么只是推演: +**实际一条都不会响**,因为 upcoming 恒为空。两个 bug 互相遮蔽。 + +### 怎么修的 + +两层,缺一不可: + +1. **`comm.ParseMemoryDate` 容忍带时间的形式**(取前 10 位,丢掉时间与时区—— + 这个字段的语义是「用户本地的那一天」不是一个时刻,按时区换算反而会把日期挪一天)。 + 这是安全网:所有 Go 侧解析点一次性都对了。 +2. **读路径出口归一**(`comm.NormalizeMemoryDate`):`memory/model.go` 的 7 个读函数、 + `stats.go` 的明细、`mcp/tool_memory.go` 交给模型的 `date`。这管的是**回给客户端 + 与模型的格式**,让它回到协议约定的 `YYYY-MM-DD`。 + +只做 1 不做 2,客户端继续收到 RFC3339;只做 2 不做 1,漏掉任何一个读路径就又静默失效。 + +**验证**:`TestNormalizeMemoryDate`、`TestExpandItem_AcceptsDbDateFormat`(同一条闹钟 +两种日期格式必须产出相同的提醒时刻)。真机复验见下节。 + +## 5. 验收清单 + +- [ ] 跨账号读 / 删 / 改 / 总结别人的纪要 → 全部拒绝 +- [ ] 新抽取的会议待办 `date_certain=0`(无截止日时)、`remind_ahead=0`,`memory_upcoming` 不再出 08:55 +- [ ] 手动建闹钟选「不提醒」→ 库里 `remind_ahead=0` +- [ ] 6 秒录音重新生成 → 不产生待办、不进 MCP 检索 +- [ ] 报告行被删 / 卡 pending / Failed → 切前台后自动补建 / 重跑,Failed 至多 3 次 +- [ ] 新写入的说话人标签全是 `Speaker_N` +- [ ] 新建记忆项 `tz` 为 IANA 名 +- [ ] 删纪要后自动待办消失、用户改过的保留 +- [ ] 迟到的失败回调不会把已完成的纪要打回「转写失败」;同一条转写只跑一轮总结 +- [ ] 重新转写一条已完成的纪要期间刷新列表页,旧纪要内容不丢 +- [ ] 周报里「闹钟次数」把重复闹钟按实际响的次数算 +- [ ] `memory_upcoming` 能排出提醒时刻(不是空数组) +- [ ] `memory_list` / MCP 回的 `happen_date` 是 `YYYY-MM-DD` 不是 RFC3339 + +## 5.1 真机验收记录(2026-09-18,阿龙测试环境) + +验收手法:用**游客登录**拿一个全新账号(`user_sgin` stype=6),既能测跨账号越权, +又能在自己名下造数据跑完整链路,不碰真实用户的数据。跑完全部清理。 + +| 验收项 | 结果 | 证据 | +|---|---|---| +| 跨账号读/改/删/总结/发起任务 | ✅ 全部拒绝 | 8 个接口逐个打过:`getrecord`/`readrecord`/`uprecord`/`summary`/`starttask`/`modifyrecords` 回 `code:21 记录不存在`;`getrecords` 回空数组不报错;`delrecords` 回 `ids:[]` 且目标行逐字未变;未登录回 `code:18` | +| 存量会议待办标记修正 | ✅ 与预测一字不差 | 16 条 `remind_ahead` 全部归 0;其中有 `due_raw` 的 6 条 `date_certain` 退回 0;助手那 9 条纹丝不动 | +| 新写入不再被默认值吞掉 | ✅ | 新抽取的会议待办 `remind_ahead=0`;`memory_add` 显式传 0 也落 0 | +| 内容门槛 | ✅ 对照实验 | 22 秒那条一条待办都没抽;同时建的 900 秒正常会议抽出 3 条 | +| MCP 时长门槛 | ✅ | `search_meeting_notes` 只回 900s/600s 两条,22 秒那条(有 101 字总结)被挡在外面 | +| 报告补偿:cron 漏建 | ✅ | 造 5 条上周记忆项 → 调 `memory_today` → 自动补建 W37 周报并生成完成;**月报没建**(上月 0 条,没过闸门) | +| 报告补偿:卡 pending | ✅ | 报告置 `state=0,update_time=0` → 切前台后重新入队并跑完 | +| 报告补偿:失败重试上限 | ✅ | `retry=2` → 重试且计数变 3;`retry=3` → 保持 Failed 不再重试 | +| 回调幂等 | ✅ 两种都撞到了 | ① 兜底扫描先判 `TranscribeFail`,6 秒后回调到达 → 「状态已是 TranscribeFail…忽略」(**改前会把已判失败的记录推进总结**);② 同一回调连打两次 → 第 2 次「状态已是 AwaitSummarizing…忽略」,只跑了一轮 AIProcess | +| 说话人标签统一 | ✅ | 回调路径产出 `speaker:"Speaker_1"`、`personnel:"[Speaker_2] [Speaker_1] "` | +| 删纪要连带删待办 | ✅ | 删掉带 3 条待办的纪要 → 自动抽的 2 条消失,标了 `user_edited=1` 的那条保留 | +| 重复闹钟统计展开 | ✅(修完根因后) | 一条 daily 闹钟在 7 天周期里 `alarm_total=7` | +| `memory_upcoming` 能排出提醒 | ✅(既有 bug 修复) | 改前恒为 `[]`,改后正确排出 6 天的 07:55 | +| 下发的 `happen_date` 格式 | ✅(既有 bug 修复) | `memory_list` 与 MCP 都回 `2026-09-07`,不再是 RFC3339 | + +⚠️ **部署过程发现的运维问题(未处理)**:服务器 `confs/home.yaml` 里**没有 `memory:` 段**, +启动日志一行 `注册模块【memory】 没有对应的配置信息`,导致 memory 模块自己打的 Info/Warn +**在 stdout 与 log/home.log 里都看不到**(「内容太薄跳过抽取」「存量修正完成 N 行」这些全是盲的)。 +配置项本身有代码内置默认值,功能不受影响,但排障时等于没有日志。 +修法是照 `deploy/app/confs/home.yaml.example` 给服务器的 home.yaml 补一段 `memory:`; +本次没动服务器配置(`dev-deploy.sh` 按设计不下发真实配置)。 + +## 6. 回滚 + +- 3.1 / 3.3 / 3.8:纯代码,回滚镜像即可。 +- 3.2:`Select("*")` 回滚镜像即可;存量修正 SQL 是幂等 UPDATE,回滚不需要反向操作(改回来的值本来就是错的)。 +- 3.4:新增行为,回滚镜像即可;已补建的报告行留着无害。 +- 3.6 客户端:`_tz()` 回退逻辑仍在,服务端能同时认两种格式,两端可各自独立回滚。 + +## 7. 未纳入本计划的已知项 + +- 会议归属日期用 `creationtime`(补录旧录音会偏)——要做准得加 `meet_date + tz` 两列由客户端上报,是表结构改动,另立项。 +- 周报 / 月报 cron 与统计一律容器时区(Asia/Shanghai),海外用户「上周」边界会偏——设计文档 §5.3 已记录,另立项。 +- 转写失败不退还 Meetintegral 的异步路径——`tasks.go sweepStuckTranscribe` 注释里已记为待办,与本链路无关。