Browse Source

client: 通话翻译字幕错位/总结残留/退不出去/切后台被冻结,及 B 路说话人门限

真机实测(Android,恒玄耳机)逐条定位后修,原生能改的都在原生层改。

字幕乱(AliyunBailianE2EHelper,Android/iOS 同改)
- 原文流与译文流各自按到达顺序自己发号,两条流的切句本来就不一致,off-by-one
  一旦发生就一路错到底。改成认服务端的 item_id:译文帧的 item_id 经 itemParent
  (response item → input item)映射回原文那条,取不到才退回本地发号。
- response.audio_transcript 的 text 是累计值、delta 才是增量,此前一律当增量
  append,于是出现「Hello,Hello,Hello,」。改成按字段名分支:有 delta 就追加,
  只有 text 就整段替换。
- connectionGeneration 守住 onOpen/onMessage/onFailure/onClosed(iOS 用 activeTask
  弱引用按任务身份判),旧 socket 的迟到回调不再打进新会话;stopContinuousConversation
  同步清 webSocket/isStarted。

翻译结束后「总结内容」还留在屏幕上
- AliyunAstCallback.onSessionFinished 会把整场累计文本当作一条 translated 再发
  一次,表现就是一大坨总结。不再发;原文缓存改为按 utteranceId 存,上限 32 条。

结束后无法再翻译、返回按钮像被遮挡
- 退出时那颗延时 Get.delete 会误杀刚重新进来的页面。按 tag 记活跃页面数
  (_PageAliveMarker + _aliveCount),还有人用就跳过销毁。

重进页面提示「未在通话中」
- 耳机对 0xD1/0xD2(进/出通话模式)的应答被当成通话状态上报,把 isCallOngoing
  改掉了。这两个 type 改为只记日志,不动通话状态。

切后台功能被关掉
- 前台服务类型加 phoneCall:microphone|phoneCall (132),ROM 的进程冻结不再掐断
  通话翻译。startForeground 带类型失败时回退到仅 microphone。stopService 改用
  context.stopService,后台也能调。
- AST 掉线自动重连(1s/2s/4s,最多 3 次),回到前台按 _astHealthy 补一次;
  重连后的新会话用 800ms 宽限期,否则会被原来的 2s 窗口吞掉重试。

B 路说话人门限(PeerSpeechGate,新增 + 11 个单测)
- 声音复刻配的是 frequency=once,只在会话开头刻一次音色。B 路会话一启动就建立,
  此时对端多半还没开口,先收到的是机主的串音(实测 B 路 ASR 配 language=en 却
  识别出机主刚说的中文),于是把机主音色刻成了「对端音色」,两个人听起来一样。
- 未开门期间音频进环形缓冲,开门瞬间整段补发——对端先开口时第一句的开头一个字
  都不能丢。fallbackFrames 到时强行开门:门限太严会让对端整通电话不被翻译,
  太松只是音色可能刻错,后者轻得多。
- 每次会话重建(含断线重连)都 reset,否则重连后照样刻到串音。

已知未尽:Android 14+ 会拒绝 phoneCall 这个前台服务类型并回退成仅 microphone,
那里的冻结保护不生效;rmsThreshold=600 尚未在真机上标定(上一轮跑到了兜底强开)。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
main
Rodger-Wang 3 weeks ago
parent
commit
8b50e52572
  1. 130
      apps/client/lib/data/utils/peer_speech_gate.dart
  2. 20
      apps/client/lib/devices/bes/bes_bluetooth_service.dart
  3. 286
      apps/client/lib/modules/translation/controllers/translation_controller.dart
  4. 57
      apps/client/lib/modules/translation/views/translation_view.dart
  5. 7
      apps/client/local_plugins/azure_speech/android/src/main/AndroidManifest.xml
  6. 245
      apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/AliyunBailianE2EHelper.kt
  7. 51
      apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/AstCallbacks.kt
  8. 53
      apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/tools/AudioRecordingForegroundService.kt
  9. 214
      apps/client/local_plugins/azure_speech/ios/azure_speech/Sources/azure_speech/AliyunBailianE2EHelper.swift
  10. 42
      apps/client/local_plugins/azure_speech/ios/azure_speech/Sources/azure_speech/AzureSpeechPlugin.swift
  11. 117
      apps/client/test/peer_speech_gate_test.dart

130
apps/client/lib/data/utils/peer_speech_gate.dart

@ -0,0 +1,130 @@
import 'dart:typed_data';
/// 对端(B 路)说话人门限:**在对端真的开口之前,不要把音频推给端到端翻译**。
///
/// ## 为什么需要它
/// 通话翻译的 B 路会话在功能一启动就建立,而阿里的实时声音复刻配的是
/// `voice_clone_options.frequency = "once"` —— **只在会话开头刻一次音色,之后整通
/// 电话不再更新**。此时对端通常还没开口,B 路先收到的往往是**机主声音的串音**
/// (2026-09-19 真机实锤:B 路 ASR 配的是 `language: en`,却识别出了机主刚说的
/// 中文「你好,亮亮,听得到吗?」)。结果是 B 路把机主的音色刻成了「对端音色」,
/// 用户听到的对端译文用的是自己的声音——两个人听起来一模一样。
///
/// 这个门限保证「喂给 B 路的第一段音频来自对端本人」,从而刻对人。
///
/// ## ⚠️ 必须缓存补发,不能丢弃
/// 门限没过就把音频扔掉的话,**对端先开口时他第一句的前 150ms 会被吃掉**
/// (「你好,听得到吗」变成「好,听得到吗」,甚至整句识别失败)。所以未开门期间
/// 音频进环形缓冲,开门瞬间把缓冲整段补推上去——既刻对人,又一个字不丢。
/// 谁先说话都不影响:对端先说就是立刻开门,机主先说就被门限挡住等对端。
///
/// ## ⚠️ 失败方向是刻意选的
/// [fallbackFrames] 到时仍未开门就**强行开门**。因为两种失败的代价不对等:
/// 门限太严 → 对端整通电话都不被翻译(功能性失效)
/// 门限太松 → 音色可能刻错(体验问题)
/// 后者轻得多,所以宁可放宽。强行开门时会打日志,真机上据此调 [rmsThreshold]。
///
/// ## ⚠️ 每次会话重建都要 [reset]
/// 断线重连会新建 session、**重新刻一次音色**,门限不复位就等于白修。
class PeerSpeechGate {
PeerSpeechGate({
this.rmsThreshold = 600,
this.openFrames = 8,
this.prerollFrames = 25,
this.fallbackFrames = 500,
});
/// 判定为「有人在说话」的 RMS 阈值(PCM16)。
/// 经验值:静音 0~50,串音漏音通常 <300,正常说话 1000~8000。
/// **需要真机调**——开门时会把实测 RMS 打进日志。
final int rmsThreshold;
/// 连续多少帧超过阈值才开门。20ms/帧,8 帧 = 160ms。
/// 用「连续」而不是「单帧」是为了滤掉串音:串音是断续的小尖峰,真声是连续的。
final int openFrames;
/// 未开门时最多缓存多少帧(20ms/帧,25 帧 = 500ms)。开门时整段补发。
final int prerollFrames;
/// 兜底:收到这么多帧仍未开门就强行开门(500 帧 = 10s)。见上方「失败方向」。
final int fallbackFrames;
bool _open = false;
int _run = 0;
int _seen = 0;
int _openRms = 0;
bool _openedByFallback = false;
final List<Uint8List> _preroll = <Uint8List>[];
/// 门是否已开(开了之后本次会话一直直通)。
bool get isOpen => _open;
/// 开门时那一帧的 RMS;兜底开门时为最后一帧的 RMS。用于真机调阈值。
int get openRms => _openRms;
/// 是否是被 [fallbackFrames] 兜底强开的(说明阈值可能偏高,或者全程只有串音)。
bool get openedByFallback => _openedByFallback;
/// 喂一帧,返回**应当推给翻译服务**的帧(可能是 0 帧、1 帧,或开门瞬间的一整串)。
List<Uint8List> accept(Uint8List frame) {
if (_open) return <Uint8List>[frame];
_seen++;
final rms = rmsOfPcm16(frame);
if (rms >= rmsThreshold) {
_run++;
} else {
_run = 0;
}
_preroll.add(frame);
while (_preroll.length > prerollFrames) {
_preroll.removeAt(0);
}
final hitThreshold = _run >= openFrames;
final hitFallback = _seen >= fallbackFrames;
if (!hitThreshold && !hitFallback) return const <Uint8List>[];
_open = true;
_openRms = rms;
_openedByFallback = !hitThreshold;
final out = List<Uint8List>.from(_preroll);
_preroll.clear();
return out;
}
/// 会话重建(含断线重连)时必须调用,否则新会话照样可能刻到串音。
void reset() {
_open = false;
_run = 0;
_seen = 0;
_openRms = 0;
_openedByFallback = false;
_preroll.clear();
}
}
/// PCM16 小端单声道的 RMS。字节数为奇数时忽略末尾半个样本。
int rmsOfPcm16(Uint8List frame) {
final n = frame.length ~/ 2;
if (n == 0) return 0;
final view = ByteData.sublistView(frame, 0, n * 2);
var sum = 0.0;
for (var i = 0; i < n; i++) {
final s = view.getInt16(i * 2, Endian.little);
sum += s * s;
}
return (sum / n).isFinite ? _sqrtInt(sum / n) : 0;
}
int _sqrtInt(double v) {
if (v <= 0) return 0;
var x = v;
var y = (x + 1) / 2;
while (y < x) {
x = y;
y = (x + v / x) / 2;
}
return x.round();
}

20
apps/client/lib/devices/bes/bes_bluetooth_service.dart

@ -294,10 +294,18 @@ class BesBluetoothService extends GetxService {
final type = event['type']; final type = event['type'];
if (type is String && type.isNotEmpty) lastCmdType.value = type; if (type is String && type.isNotEmpty) lastCmdType.value = type;
switch (type) { switch (type) {
// 0x02/0x03 是去电起止,0xD1/0xD2 是通话接通/挂断, // 0x02/0x03 是去电起止 —— **只有它们**代表真实通话状态,
// 两类都算「通话中」——通话录音和通话翻译都以此为准 // 通话录音和通话翻译的准入都以此为准。
//
// ⚠️ **0xD1/0xD2 不在此列**(2026-09-19 真机定性):它们是
// `callModeStart(0x51)` / `callModeStop(0x52)` 的应答(`| 0x80`),
// 也就是**我们自己那条「进/退通话模式」指令的 ACK**,跟电话接没接通无关。
// 原来把它们一起当通话状态,后果是:通话翻译一结束,收尾发的 call.stop
// 换回一个 0xD2 → isCallOngoing 被打成 false → DeviceHub.inCall 跟着 false
// → 主界面那道「不在通话中」的闸门(home_controller.openTranslationFeature)
// 把用户挡在门外,而电话其实一直在通着;要等耳机下一次推真实状态帧
// (0x02/0x8E)才恢复,表现就是「关掉再点提示未在通话中,再点一次才进得去」。
case 'startMakeCall': case 'startMakeCall':
case 'startCall':
// ⚠️ 用 info 而不是 debug:这是排查「刚连上耳机就被判成通话中、 // ⚠️ 用 info 而不是 debug:这是排查「刚连上耳机就被判成通话中、
// 同传点不动」的关键线索——原生侧 `case 0x02` 对任何第二字节为 0x02 // 同传点不动」的关键线索——原生侧 `case 0x02` 对任何第二字节为 0x02
// 的帧都会上报 startMakeCall,且不做任何长度/上下文校验, // 的帧都会上报 startMakeCall,且不做任何长度/上下文校验,
@ -306,10 +314,14 @@ class BesBluetoothService extends GetxService {
isCallOngoing.value = true; isCallOngoing.value = true;
break; break;
case 'stopMakeCall': case 'stopMakeCall':
case 'stopCall':
Logger.i(_tag, 'isCallOngoing → false (type=$type, raw=[$hex])'); Logger.i(_tag, 'isCallOngoing → false (type=$type, raw=[$hex])');
isCallOngoing.value = false; isCallOngoing.value = false;
break; break;
case 'startCall':
case 'stopCall':
// 我们自己发的进/退通话模式指令的 ACK,只记一笔,**不动通话状态**(见上)
Logger.i(_tag, '通话模式指令已应答 (type=$type, raw=[$hex]),不改通话状态');
break;
case 'unknown': case 'unknown':
// 原始帧回包都以 type=unknown 上报,按 CMD 二级分发 // 原始帧回包都以 type=unknown 上报,按 CMD 二级分发
_handleRawFrame(event['data']); _handleRawFrame(event['data']);

286
apps/client/lib/modules/translation/controllers/translation_controller.dart

@ -6,6 +6,8 @@ import 'package:flutter/material.dart';
import 'package:flutter/services.dart'; import 'package:flutter/services.dart';
import 'package:get/get.dart'; import 'package:get/get.dart';
import 'package:get_storage/get_storage.dart'; import 'package:get_storage/get_storage.dart';
import '../../../core/utils/audio_foreground_service.dart';
import '../../../data/utils/peer_speech_gate.dart';
import '../../../core/utils/recent_languages.dart'; import '../../../core/utils/recent_languages.dart';
import 'package:intl/intl.dart'; import 'package:intl/intl.dart';
import 'package:floating_ui_plugin/native_plugin.dart'; import 'package:floating_ui_plugin/native_plugin.dart';
@ -92,6 +94,20 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
/// 后清掉;真正停止过之后(stopRecognition/_terminateCallOnError)也要清掉, /// 后清掉;真正停止过之后(stopRecognition/_terminateCallOnError)也要清掉,
/// 否则下次点开始会以为不需要再调,实际原生那边已经停了。 /// 否则下次点开始会以为不需要再调,实际原生那边已经停了。
bool _astAutoStartPending = false; bool _astAutoStartPending = false;
// AST 掉线自动重连:切后台被 ROM 冻结、弱网抖动都会让 WebSocket 断开,
// 那不是「功能该结束了」,用户还在通话里。见 _handleAstDrop。
int _astReconnectAttempts = 0;
bool _astReconnecting = false;
/// 对端(B 路)说话人门限,保证阿里刻到的第一段音频来自对端本人。
/// 见 [PeerSpeechGate];**每次 AST 会话重建都要 reset**,否则重连后照样刻到串音。
final PeerSpeechGate _peerGate = PeerSpeechGate();
/// AST 链路当前是否可用。收到任何一条 AST 事件即为 true,掉线置 false。
/// [didChangeAppLifecycleState] 靠它决定回到前台要不要补一次重连。
bool _astHealthy = false;
Timer? _astReconnectTimer;
static const int _astMaxReconnects = 3;
final GetStorage _storage = GetStorage(); final GetStorage _storage = GetStorage();
final _musiceManager = Get.find<MusicManager>(); final _musiceManager = Get.find<MusicManager>();
final _qqMusicManager = Get.find<QqMusicService>(); final _qqMusicManager = Get.find<QqMusicService>();
@ -991,7 +1007,7 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
} }
/// 初始化通话模式的语音翻译服务 /// 初始化通话模式的语音翻译服务
Future<void> _initializeCallModeTranslationService() async { Future<void> _initializeCallModeTranslationService({bool isRecovery = false}) async {
try { try {
Logger.info('开始初始化通话模式语音翻译服务'); Logger.info('开始初始化通话模式语音翻译服务');
@ -1029,13 +1045,22 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
await _astService.initialize(supportedLanguages: callModeLanguages, provider: astProvider); await _astService.initialize(supportedLanguages: callModeLanguages, provider: astProvider);
// ⚠️ 新会话 = 重新刻一次音色,门限必须跟着重新关上,否则重连后照样刻到串音。
_peerGate.reset();
// 获取识别事件流(端到端服务已包含 ASR+翻译+TTS,不需要单独启动 ASR) // 获取识别事件流(端到端服务已包含 ASR+翻译+TTS,不需要单独启动 ASR)
await _subscribeAstEvents(); await _subscribeAstEvents();
// 初始化完成后给 2s 宽限期,过滤掉底层异步 dispose 旧 provider session 时产生的 // 初始化完成后给一段宽限期,过滤掉底层异步 dispose 旧 provider session 时产生的
// cancelled(1012) / Socket not connected(1011) 残留事件,避免被误判为新 session 错误 // cancelled(1012) / Socket not connected(1011) 残留事件,避免被误判为新 session 错误。
_astErrorGraceUntil = //
DateTime.now().add(const Duration(milliseconds: 2000)); // ⚠️ **重连时必须把这个窗口收窄**。2026-09-19 真机:进程被 ROM 冻结 →
// 重连起来的新连接在 1.9s 后又被冻死,而那条致命错误正好落在 2s 宽限期里被
// 当成「旧会话残留」吞掉 → 第 2 次重连压根没排上,翻译就此停在那里、
// 回到前台也不会自己好。恢复场景下旧 session 早在上一轮就收干净了,
// 800ms 足够盖住 dispose 的尾巴。
_astErrorGraceUntil = DateTime.now().add(
isRecovery ? const Duration(milliseconds: 800) : const Duration(milliseconds: 2000));
Logger.info('通话模式语音翻译服务初始化完成'); Logger.info('通话模式语音翻译服务初始化完成');
} catch (e) { } catch (e) {
@ -1088,7 +1113,23 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
setActiveSpeaker(0); setActiveSpeaker(0);
}); });
} }
return;
} }
if (state != AppLifecycleState.resumed) return;
// ⚠️ 回到前台是通话翻译的**最后一道兜底**。
// 国内 ROM(鸿蒙的 Pged-Freezer 尤其激进)会在切后台几秒后直接冻结进程,
// 前台服务也挡不住(2026-09-19 真机:服务已按 serviceType:128 登记、通知也挂着,
// 照样 `Freeze process`)。进程被冻住时连重连定时器都不会跑,所以必须在解冻回来
// 的这一刻主动补一次——否则用户看到的就是「切出去一趟,回来翻译已经死了,
// 还得手动结束再重开」。
if (currentMode.value != 'call' || !isRecognizing.value) return;
if (_astHealthy) return;
Logger.w('Translation', '[STS] 回到前台且 AST 已掉线,立即重连');
_astReconnectTimer?.cancel();
_astReconnectTimer = null;
_astReconnecting = false;
_astReconnectAttempts = 0; // 用户回来了,重连预算给满
unawaited(_handleAstDrop('回到前台后恢复'));
} }
/// 切换TTS功能开关 /// 切换TTS功能开关
@ -1572,7 +1613,16 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
/// _initializeCallModeTranslationService 中初始化)。 /// _initializeCallModeTranslationService 中初始化)。
/// 需要一台具备 [DeviceCapability.callAudioTap] 的设备;没有就直接报错。 /// 需要一台具备 [DeviceCapability.callAudioTap] 的设备;没有就直接报错。
Future<void> _configureCallMode() async { Future<void> _configureCallMode() async {
_callSession = DeviceHub.to.sessionWith(DeviceCapability.callAudioTap); // ⚠️ **前台服务必须在这里拉起,不能指望 startContinuousTranslation**。
// 那个方法里确实起了前台服务,但通话翻译常常**根本不走它**:AST 已由
// initialize() 自动拉起时,startRecognition 会走「跳过重复调用」的分支
// (见 _astAutoStartPending),于是通话翻译成了唯一一个没有前台服务的模式。
// 后果是 2026-09-19 真机实录的这一串:按 Home 4.5 秒后
// `Pged-Freezer: Freeze process` 冻结进程 → 两条 WebSocket 当场
// `Software caused connection abort` → Dart 当成 AST 故障 →
// _terminateCallOnError 把整个通话翻译关掉。用户看到的是「切个后台功能就没了」。
await AudioForegroundService.start();
_callSession = await _awaitCallAudioSession();
_callModeUsingBes = _callSession != null; _callModeUsingBes = _callSession != null;
_callAudioTornDown = false; _callAudioTornDown = false;
if (_callModeUsingBes) { if (_callModeUsingBes) {
@ -1583,6 +1633,40 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
} }
} }
/// 等一台能抓通话双路的设备上线,最多 [timeout]。
///
/// ⚠️ 原来是「当场问一次 DeviceHub,拿不到就弹错抛异常」,而**设备会话不是立刻就绪的**:
/// 真机实测 App 启动到 `设备上线: bes Echo-one` 有 ~2.7s(蓝牙连上 → 服务发现 →
/// 握手),退出通话翻译时给耳机发过 `call.stop` 之后也有一小段空窗。用户在这个窗口里
/// 点进来,看到的就是「蓝牙通话状态异常」,而**过一会儿再点就好了**——这正是
/// 2026-09-19 真机报的「退出再进提示未在通话中,等一下又好了」。
///
/// 等待上限取 6s:比实测的 2.7s 留一倍余量,又不至于让真的没连耳机的人干等太久
/// (那种情况下多等这几秒,换来的是弱网/慢启动时不再误报)。
Future<DeviceSession?> _awaitCallAudioSession({
Duration timeout = const Duration(seconds: 6),
}) async {
final immediate = DeviceHub.to.sessionWith(DeviceCapability.callAudioTap);
if (immediate != null) return immediate;
Logger.w('Translation', '通话翻译:设备会话尚未就绪,等待最多 ${timeout.inSeconds}s');
final completer = Completer<DeviceSession?>();
// sessions 是 RxList,设备上线/下线都会触发
final worker = ever<List<DeviceSession>>(DeviceHub.to.sessions, (_) {
final s = DeviceHub.to.sessionWith(DeviceCapability.callAudioTap);
if (s != null && !completer.isCompleted) completer.complete(s);
});
final timer = Timer(timeout, () {
if (!completer.isCompleted) completer.complete(null);
});
final result = await completer.future;
timer.cancel();
worker.dispose();
Logger.i('Translation',
result != null ? '通话翻译:设备会话已就绪' : '通话翻译:等待 ${timeout.inSeconds}s 仍无可用设备');
return result;
}
/// 配置通话模式(设备双路通路):进通话模式拿双路上行 + 开译文下行通道 + /// 配置通话模式(设备双路通路):进通话模式拿双路上行 + 开译文下行通道 +
/// 通知设备 ASR 已开启,再把本端/对端 PCM 接上 AST。 /// 通知设备 ASR 已开启,再把本端/对端 PCM 接上 AST。
/// ///
@ -1637,8 +1721,24 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
(pcm) => _astService.pushExternalAudio(leg: 'A', pcm: pcm), (pcm) => _astService.pushExternalAudio(leg: 'A', pcm: pcm),
onError: (e) => Logger.error('[CALL-BRIDGE] local stream error: $e'), onError: (e) => Logger.error('[CALL-BRIDGE] local stream error: $e'),
); );
// ⚠️ B 路**不能拿到帧就推**:阿里的声音复刻是 frequency=once,只在会话开头刻一次
// 音色,而那时对端往往还没开口,先进来的是机主声音的串音(2026-09-19 真机实锤:
// B 路配的是 language:en,却识别出机主刚说的中文)。刻错的结果是用户听到的对端
// 译文用的是自己的声音——两个人音色一模一样。门限保证第一段来自对端本人,
// 未开门期间的音频进环形缓冲、开门时整段补发,所以对端先开口也不会被吃掉开头。
_peerGate.reset();
_besCallSpkSub = _callPeerSrc!.pcm.listen( _besCallSpkSub = _callPeerSrc!.pcm.listen(
(pcm) => _astService.pushExternalAudio(leg: 'B', pcm: pcm), (pcm) {
final wasOpen = _peerGate.isOpen;
for (final f in _peerGate.accept(pcm)) {
_astService.pushExternalAudio(leg: 'B', pcm: f);
}
if (!wasOpen && _peerGate.isOpen) {
// 真机调 rmsThreshold 的唯一依据,必须留在正式包(release 级别是 warning)
Logger.w('Translation',
'[CALL-BRIDGE] B 路门限开启:${_peerGate.openedByFallback ? "兜底强开(阈值可能偏高/全程只有串音)" : "检测到对端语音"} rms=${_peerGate.openRms}');
}
},
onError: (e) => Logger.error('[CALL-BRIDGE] peer stream error: $e'), onError: (e) => Logger.error('[CALL-BRIDGE] peer stream error: $e'),
); );
// 诊断:计数 + isFinal 必打,用于定位 "TTS 到底有没有被转发到 native"。 // 诊断:计数 + isFinal 必打,用于定位 "TTS 到底有没有被转发到 native"。
@ -1809,9 +1909,35 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
} }
} }
/// 找这条 AST 事件属于哪条**还没闭合**的字幕(识别与翻译分四个事件到齐才闭合)。
///
/// 匹配键是 **(utteranceId, serviceId)** 两项:
/// - `utteranceId` 由原生按**句**生成(`<sessionId>#<序号>`)。它曾经等于会话 id ——
/// 一整通电话一条腿只有一个值,于是这里必然匹配到上一句那条还开着的记录:
/// 端到端有 ~2.8s 语义延迟,第 N+1 句的中间结果总是赶在第 N 句译文之前到,
/// 原文被后一句覆盖、译文却是前一句的,字幕越说越乱(2026-09-19 修);
/// - `serviceId` 是 A/B 两条腿(己方 / 对方)。两腿 id 本就不同,带上它只是
/// 多一道保险:万一哪天 id 撞了,宁可多起一条,也不能把两个人的话并成一句。
TranslationItem? _findOpenAstItem(ASTEvent event) {
if (event.utteranceId.isEmpty) return null;
for (int i = translationHistory.length - 1; i >= 0; i--) {
final item = translationHistory[i];
if (item.isIntermediate &&
item.utteranceId == event.utteranceId &&
item.serviceId == event.serviceId) {
return item;
}
}
return null;
}
/// 处理 AST(语音识别+翻译一体)事件 /// 处理 AST(语音识别+翻译一体)事件
/// 按 serviceId 区分双路(A=己方, B=对方),用 utteranceId 匹配同一句话的事件 /// 按 serviceId 区分双路(A=己方, B=对方),用 utteranceId 匹配同一句话的事件
void _handleAstEvent(ASTEvent event) { void _handleAstEvent(ASTEvent event) {
// 有事件进来就说明链路是活的(错误/取消除外,它们在下面各自置回 false)
if (event.type != ASTEventType.error && event.type != ASTEventType.canceled) {
_astHealthy = true;
}
// 仅对终态事件打 info 日志,避免中间事件高频写 I/O 阻塞 event loop(会导致 AudioSendSlow 1011) // 仅对终态事件打 info 日志,避免中间事件高频写 I/O 阻塞 event loop(会导致 AudioSendSlow 1011)
if (event.type == ASTEventType.finalResult || if (event.type == ASTEventType.finalResult ||
event.type == ASTEventType.translationResult || event.type == ASTEventType.translationResult ||
@ -1833,19 +1959,8 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
if (event.text.isEmpty) break; if (event.text.isEmpty) break;
currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString(); currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString();
// 按 utteranceId 查找已有项 // 按 (utteranceId, serviceId) 查找已有项,见 _findOpenAstItem
TranslationItem? target; TranslationItem? target = _findOpenAstItem(event);
int targetIndex = -1;
if (event.utteranceId.isNotEmpty) {
for (int i = translationHistory.length - 1; i >= 0; i--) {
if (translationHistory[i].isIntermediate &&
translationHistory[i].utteranceId == event.utteranceId) {
target = translationHistory[i];
targetIndex = i;
break;
}
}
}
if (target != null) { if (target != null) {
target.sourceText = event.text; target.sourceText = event.text;
@ -1872,19 +1987,13 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
case ASTEventType.finalResult: case ASTEventType.finalResult:
if (event.text.isEmpty) break; if (event.text.isEmpty) break;
// 有终态结果回来 = 链路是通的,把重连预算还回去(否则一通长电话里
// 攒够 3 次零星抖动就再也不重连了)
_astReconnectAttempts = 0;
Logger.i('Translation', 'AST 识别完成 [${event.serviceId}]: ${event.text}'); Logger.i('Translation', 'AST 识别完成 [${event.serviceId}]: ${event.text}');
currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString(); currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString();
TranslationItem? target; TranslationItem? target = _findOpenAstItem(event);
if (event.utteranceId.isNotEmpty) {
for (int i = translationHistory.length - 1; i >= 0; i--) {
if (translationHistory[i].isIntermediate &&
translationHistory[i].utteranceId == event.utteranceId) {
target = translationHistory[i];
break;
}
}
}
if (target == null) { if (target == null) {
target = TranslationItem( target = TranslationItem(
@ -1925,16 +2034,7 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
if (event.text.isEmpty) break; if (event.text.isEmpty) break;
currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString(); currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString();
TranslationItem? target; TranslationItem? target = _findOpenAstItem(event);
if (event.utteranceId.isNotEmpty) {
for (int i = translationHistory.length - 1; i >= 0; i--) {
if (translationHistory[i].isIntermediate &&
translationHistory[i].utteranceId == event.utteranceId) {
target = translationHistory[i];
break;
}
}
}
if (target != null) { if (target != null) {
target.translatedText = event.text; target.translatedText = event.text;
@ -1963,16 +2063,7 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
Logger.i('Translation', 'AST 翻译完成 [${event.serviceId}]: ${event.text}'); Logger.i('Translation', 'AST 翻译完成 [${event.serviceId}]: ${event.text}');
currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString(); currentSessionId ??= DateTime.now().millisecondsSinceEpoch.toString();
TranslationItem? target; TranslationItem? target = _findOpenAstItem(event);
if (event.utteranceId.isNotEmpty) {
for (int i = translationHistory.length - 1; i >= 0; i--) {
if (translationHistory[i].isIntermediate &&
translationHistory[i].utteranceId == event.utteranceId) {
target = translationHistory[i];
break;
}
}
}
if (target == null) { if (target == null) {
target = TranslationItem( target = TranslationItem(
@ -2022,8 +2113,8 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
Logger.info('[STS] 初始化宽限期内,忽略旧会话错误: ${event.error}'); Logger.info('[STS] 初始化宽限期内,忽略旧会话错误: ${event.error}');
break; break;
} }
// 非重新初始化期间的错误:终止通话服务并还原状态 // 非重新初始化期间的错误:先尝试重连,连不上才真的关掉
_terminateCallOnError('AST 错误: ${event.error}'); unawaited(_handleAstDrop('AST 错误: ${event.error}'));
break; break;
case ASTEventType.canceled: case ASTEventType.canceled:
@ -2037,7 +2128,7 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
Logger.info('[STS] 初始化宽限期内,忽略旧会话取消: ${event.error}'); Logger.info('[STS] 初始化宽限期内,忽略旧会话取消: ${event.error}');
break; break;
} }
_terminateCallOnError('AST 取消: ${event.error}'); unawaited(_handleAstDrop('AST 取消: ${event.error}'));
break; break;
default: default:
@ -2061,6 +2152,10 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
await stopRecording(); await stopRecording();
await _asrService.stopContinuousRecognition(); await _asrService.stopContinuousRecognition();
if (currentMode.value == "call") { if (currentMode.value == "call") {
_astReconnectTimer?.cancel();
_astReconnectTimer = null;
_astReconnecting = false;
_astReconnectAttempts = 0;
await _astService.stopContinuousTranslation(); await _astService.stopContinuousTranslation();
_astAutoStartPending = false; _astAutoStartPending = false;
} }
@ -2549,10 +2644,63 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
Logger.info('[STS] 已订阅 astStream, subscription=${_astEventSubscription.hashCode}'); Logger.info('[STS] 已订阅 astStream, subscription=${_astEventSubscription.hashCode}');
} }
/// AST 链路掉了:**先当成掉线重连,不要当成「功能结束」**。
///
/// 2026-09-19 真机:按 Home 之后 4.5 秒,鸿蒙 `Pged-Freezer` 冻结进程,两条
/// WebSocket 同时 `Software caused connection abort (1011)`。原来这里直接
/// [_terminateCallOnError] 把整个通话翻译收掉——用户还在通话中,回到前台
/// 发现翻译没了。前台服务(见 _configureCallMode)能挡住大部分冻结,但弱网、
/// 服务端主动断开仍会发生,所以这一层重连是必须的。
///
/// 只在「用户确实还在通话翻译里」时重连;连 [_astMaxReconnects] 次仍不行,
/// 才按原来的方式终止——那时确实不是抖动了。
Future<void> _handleAstDrop(String reason) async {
_astHealthy = false;
if (currentMode.value != 'call' || !isRecognizing.value) {
await _terminateCallOnError(reason);
return;
}
// A/B 两条腿会各报一次,同一轮只处理一次
if (_astReconnecting) {
Logger.i('Translation', '[STS] 已在重连中,忽略重复上报: $reason');
return;
}
if (_astReconnectAttempts >= _astMaxReconnects) {
Logger.e('Translation', '[STS] 重连 $_astMaxReconnects 次仍失败,终止通话翻译: $reason');
await _terminateCallOnError(reason);
return;
}
_astReconnecting = true;
_astReconnectAttempts++;
final delay = Duration(seconds: 1 << (_astReconnectAttempts - 1)); // 1s / 2s / 4s
Logger.w('Translation',
'[STS] AST 掉线($reason),第 $_astReconnectAttempts 次重连将在 ${delay.inSeconds}s 后开始');
_astReconnectTimer?.cancel();
_astReconnectTimer = Timer(delay, () async {
try {
// 这段时间里用户可能已经自己结束了,别把它又拉起来
if (currentMode.value != 'call' || !isRecognizing.value) {
Logger.i('Translation', '[STS] 重连前发现已不在通话翻译中,放弃重连');
return;
}
await _initializeCallModeTranslationService(isRecovery: true);
Logger.i('Translation', '[STS] 第 $_astReconnectAttempts 次重连已发起');
} catch (e) {
Logger.e('Translation', '[STS] 重连失败: $e');
} finally {
_astReconnecting = false;
}
});
}
/// 通话模式发生错误时终止服务并还原状态 /// 通话模式发生错误时终止服务并还原状态
Future<void> _terminateCallOnError(String reason) async { Future<void> _terminateCallOnError(String reason) async {
Logger.error('[STS] 通话错误,终止服务: $reason'); Logger.error('[STS] 通话错误,终止服务: $reason');
try { try {
_astReconnectTimer?.cancel();
_astReconnectTimer = null;
_astReconnecting = false;
_astReconnectAttempts = 0;
_astEventSubscription?.cancel(); _astEventSubscription?.cancel();
_astEventSubscription = null; _astEventSubscription = null;
try { try {
@ -2603,6 +2751,28 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
// 之前挂着的"后台继续翻译"提示就没意义了,撤掉。 // 之前挂着的"后台继续翻译"提示就没意义了,撤掉。
unawaited(BackgroundSessionNotifier.cancel(kCallTranslationNotificationId)); unawaited(BackgroundSessionNotifier.cancel(kCallTranslationNotificationId));
} }
// 前台服务跟着这一轮通话翻译一起收(见 _configureCallMode 的说明)
await AudioForegroundService.stop();
// ⚠️ 悬浮字幕窗必须跟着一起收。它原来只在「用户手动 toggle / 切到不支持的模式 /
// controller 销毁」时才关,于是点完「结束」小窗还浮在页面上:安卓那个是
// TYPE_APPLICATION_OVERLAY 系统级覆盖层,压在它身上的点击**不会穿透**
// (FLAG_NOT_TOUCH_MODAL 只放行窗外的),盖住哪儿哪儿就点不动——用户报的
// 「返回按钮像被遮挡、退不出去」就是它;而退不出去 → controller 不销毁 →
// 小窗更关不掉,闭环。iOS 的 PiP 同理(系统窗口浮在 App 之上)。
//
// 放在这里是因为本方法正好覆盖「支持小窗的那两个模式」的全部收尾路径
// (stopAll / _abortRecognition / _terminateCallOnError),且自带一次性短路。
// ⚠️ 「后台继续翻译」那条路径不经过这里(它压根不停翻译),小窗照常留着——
// 那正是要的:人切到通话界面了,字幕只能靠小窗看。
try {
if (isFloatingWindowEnabled.value) {
await _disableFloatingWindow();
}
} catch (e) {
Logger.error('停止翻译时关闭悬浮字幕窗失败(忽略): $e');
}
} }
/// 设备通话翻译收尾:断开桥接 + 通知 ASR 关闭 + 关译文下行 + 退通话模式。 /// 设备通话翻译收尾:断开桥接 + 通知 ASR 关闭 + 关译文下行 + 退通话模式。
@ -2817,6 +2987,9 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
/// [language] 语言名称 /// [language] 语言名称
Future<void> setSourceLanguage(String language) => _runLanguageSwitch(() async { Future<void> setSourceLanguage(String language) => _runLanguageSwitch(() async {
if (sourceLanguage.value == language) return; if (sourceLanguage.value == language) return;
// 选中的正是目标语言 → 两边对调,而不是让源=目标。
// (语言弹窗左右两侧给的是同一张完整表,撞上是正常操作,见 LanguageSelectionDialog)
if (targetLanguage.value == language) return _swapLanguagesInner();
if (!guardLanguageChange()) return; if (!guardLanguageChange()) return;
final wasRecognizing = isRecognizing.value; final wasRecognizing = isRecognizing.value;
await stopAll(); await stopAll();
@ -2836,6 +3009,8 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
/// [language] 语言名称 /// [language] 语言名称
Future<void> setTargetLanguage(String language) => _runLanguageSwitch(() async { Future<void> setTargetLanguage(String language) => _runLanguageSwitch(() async {
if (targetLanguage.value == language) return; if (targetLanguage.value == language) return;
// 选中的正是源语言 → 两边对调(同上)
if (sourceLanguage.value == language) return _swapLanguagesInner();
if (!guardLanguageChange()) return; if (!guardLanguageChange()) return;
final wasRecognizing = isRecognizing.value; final wasRecognizing = isRecognizing.value;
await stopAll(); await stopAll();
@ -2987,9 +3162,6 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
Logger.d(TAG, '触发加载更多翻译历史记录(reverse模式)'); Logger.d(TAG, '触发加载更多翻译历史记录(reverse模式)');
loadMoreHistory(); loadMoreHistory();
} }
// 选中的正是目标语言 → 两边对调,而不是让源=目标。
// (语言弹窗左右两侧给的是同一张完整表,撞上是正常操作,见 LanguageSelectionDialog)
if (targetLanguage.value == language) return _swapLanguagesInner();
} }
}); });
@ -3009,8 +3181,6 @@ class TranslationController extends GetxController with WidgetsBindingObserver {
/// [timestamp] 时间戳 /// [timestamp] 时间戳
void deleteTranslationItem(String sessionId, DateTime timestamp) => void deleteTranslationItem(String sessionId, DateTime timestamp) =>
_historyManager.deleteTranslationItem(sessionId, timestamp); _historyManager.deleteTranslationItem(sessionId, timestamp);
// 选中的正是源语言 → 两边对调(同上)
if (sourceLanguage.value == language) return _swapLanguagesInner();
/// 获取所有历史记录 /// 获取所有历史记录
/// 返回所有翻译项列表 /// 返回所有翻译项列表

57
apps/client/lib/modules/translation/views/translation_view.dart

@ -133,8 +133,24 @@ class TranslationView extends GetView<TranslationController> {
(current != null && current.isAtSameMomentAs(ts)) ? null : ts; (current != null && current.isAtSameMomentAs(ts)) ? null : ts;
} }
/// 页面还活着的实例数。[_PageAliveMarker] 在 initState/dispose 里维护。
///
/// ⚠️ 这个计数是 [_deleteControllerAfterPop] 的**唯一保险**,别去掉。
static int _aliveCount = 0;
void _deleteControllerAfterPop() { void _deleteControllerAfterPop() {
Future.delayed(const Duration(milliseconds: 400), () { Future.delayed(const Duration(milliseconds: 400), () {
// ⚠️ 400ms 之内用户完全可能已经重新进来了——从「通话翻译」退出去紧接着点
// 「多媒体翻译」就是这条路径:四种模式共用同一个 permanent controller,
// 新页面复用的正是这个实例,而上一个页面留下的这颗延时炸弹会把它删掉。
// 后果不是抛在控制台就算了:GetView.controller 每次访问都 Get.find,删掉之后
// AppBar 里的 Obx 当场抛 "TranslationController not found",Flutter 用错误
// 控件顶替整条 AppBar —— 红框盖住标题和**返回键**,页面再也退不出去。
// (同款事故见 call_recording 页,那边是按 tag 记活跃页面数。)
if (_aliveCount > 0) {
Logger.i('Translation', '页面已被重新打开(活跃 $_aliveCount 个),跳过 controller 销毁');
return;
}
if (Get.isRegistered<TranslationController>()) { if (Get.isRegistered<TranslationController>()) {
Get.delete<TranslationController>(force: true); Get.delete<TranslationController>(force: true);
} }
@ -240,7 +256,8 @@ class TranslationView extends GetView<TranslationController> {
// 面对面翻译原来是上下分屏 + 两个「按住说话」按钮, // 面对面翻译原来是上下分屏 + 两个「按住说话」按钮,
// 现在与同声翻译统一:同一套 Scaffold、同一个开始/停止按钮, // 现在与同声翻译统一:同一套 Scaffold、同一个开始/停止按钮,
// 左右语言胶囊分别代表左耳 / 右耳。 // 左右语言胶囊分别代表左耳 / 右耳。
return PopScope( return _PageAliveMarker(
child: PopScope(
canPop: false, canPop: false,
onPopInvokedWithResult: (didPop, result) async { onPopInvokedWithResult: (didPop, result) async {
if (didPop) return; if (didPop) return;
@ -264,7 +281,7 @@ class TranslationView extends GetView<TranslationController> {
), ),
), ),
), ),
); ));
}); });
} }
@ -441,16 +458,16 @@ class TranslationView extends GetView<TranslationController> {
sourceText: () => controller.sourceLanguage.value.split(' ')[0], sourceText: () => controller.sourceLanguage.value.split(' ')[0],
targetText: () => controller.targetLanguage.value.split(' ')[0], targetText: () => controller.targetLanguage.value.split(' ')[0],
// 翻译进行中不开弹窗(见 controller.guardLanguageChange) // 翻译进行中不开弹窗(见 controller.guardLanguageChange)
// 左右两侧都给完整语言表(不传 excludeLanguage):选到对侧那项时
// controller 会把两边对调,不需要靠「从列表里删掉」来避免同语言
onTapSource: () { onTapSource: () {
if (!controller.guardLanguageChange()) return; if (!controller.guardLanguageChange()) return;
LanguageSelectionDialog.show(true, isDarkMode, LanguageSelectionDialog.show(true, isDarkMode,
excludeLanguage: controller.targetLanguage.value,
translationController: controller); translationController: controller);
}, },
onTapTarget: () { onTapTarget: () {
if (!controller.guardLanguageChange()) return; if (!controller.guardLanguageChange()) return;
LanguageSelectionDialog.show(false, isDarkMode, LanguageSelectionDialog.show(false, isDarkMode,
excludeLanguage: controller.sourceLanguage.value,
translationController: controller); translationController: controller);
}, },
onSwap: controller.swapLanguages, onSwap: controller.swapLanguages,
@ -1098,10 +1115,8 @@ class TranslationView extends GetView<TranslationController> {
sourceText: () => controller.sourceLanguage.value.split(' ')[0], sourceText: () => controller.sourceLanguage.value.split(' ')[0],
targetText: () => controller.targetLanguage.value.split(' ')[0], targetText: () => controller.targetLanguage.value.split(' ')[0],
onTapSource: () => LanguageSelectionDialog.show(true, isDarkMode, onTapSource: () => LanguageSelectionDialog.show(true, isDarkMode,
excludeLanguage: controller.targetLanguage.value,
translationController: controller), translationController: controller),
onTapTarget: () => LanguageSelectionDialog.show(false, isDarkMode, onTapTarget: () => LanguageSelectionDialog.show(false, isDarkMode,
excludeLanguage: controller.sourceLanguage.value,
translationController: controller), translationController: controller),
onSwap: controller.swapLanguages, onSwap: controller.swapLanguages,
translateLanguageNames: true, // 是否翻译语言名称 translateLanguageNames: true, // 是否翻译语言名称
@ -1110,3 +1125,33 @@ class TranslationView extends GetView<TranslationController> {
); );
} }
} }
/// 只做一件事:把「这个翻译页面还活着」这件事记进 [TranslationView._aliveCount]。
///
/// 单独拎成一个 StatefulWidget,是为了不把 1100 行的 TranslationView 整体改成
/// StatefulWidget —— 它是 GetView,改起来要连带挪十几个方法。
class _PageAliveMarker extends StatefulWidget {
const _PageAliveMarker({required this.child});
final Widget child;
@override
State<_PageAliveMarker> createState() => _PageAliveMarkerState();
}
class _PageAliveMarkerState extends State<_PageAliveMarker> {
@override
void initState() {
super.initState();
TranslationView._aliveCount++;
}
@override
void dispose() {
TranslationView._aliveCount--;
super.dispose();
}
@override
Widget build(BuildContext context) => widget.child;
}

7
apps/client/local_plugins/azure_speech/android/src/main/AndroidManifest.xml

@ -6,6 +6,11 @@
<!-- 添加前台服务权限 --> <!-- 添加前台服务权限 -->
<uses-permission android:name="android.permission.FOREGROUND_SERVICE" /> <uses-permission android:name="android.permission.FOREGROUND_SERVICE" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_MICROPHONE" /> <uses-permission android:name="android.permission.FOREGROUND_SERVICE_MICROPHONE" />
<!-- 通话类前台服务:国内 ROM(鸿蒙 Pged-Freezer 尤其激进)对「麦克风」类型照样冻结,
通话类通常更宽容。⚠️ Android 14+ 用这个类型还要求 App 是默认拨号或持有
MANAGE_OWN_CALLS,我们都不是,所以**运行时必然可能被拒**——
AudioRecordingForegroundService 里带了回退到纯麦克风的分支,别把它删了。 -->
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_PHONE_CALL" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_MEDIA_PROJECTION" /> <uses-permission android:name="android.permission.FOREGROUND_SERVICE_MEDIA_PROJECTION" />
<uses-permission android:name="android.permission.POST_NOTIFICATIONS" /> <uses-permission android:name="android.permission.POST_NOTIFICATIONS" />
@ -15,7 +20,7 @@
android:name=".tools.AudioRecordingForegroundService" android:name=".tools.AudioRecordingForegroundService"
android:enabled="true" android:enabled="true"
android:exported="false" android:exported="false"
android:foregroundServiceType="microphone" /> android:foregroundServiceType="microphone|phoneCall" />
<!-- 注册屏幕/系统音频捕获前台服务(Android 14+ 必须 mediaProjection 类型) --> <!-- 注册屏幕/系统音频捕获前台服务(Android 14+ 必须 mediaProjection 类型) -->
<service <service
android:name=".tools.ScreenCaptureForegroundService" android:name=".tools.ScreenCaptureForegroundService"

245
apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/AliyunBailianE2EHelper.kt

@ -21,21 +21,31 @@ class AliyunBailianE2EHelper(
) { ) {
private val TAG = "AliyunBailianHelper" private val TAG = "AliyunBailianHelper"
/**
* ⚠️ 四个文本回调都带 [utteranceId]:**一句话一个 id**,识别与翻译共用同一个,
* 上层据此把「原文 + 译文」并进同一条字幕。
*
* 原来这四个回调只有 sessionId,桥接层就拿它当 utteranceId 上报 —— 而 sessionId
* 是**每条 WebSocket 一个**,即一整通电话里一条腿只有一个 id。上层那套「按
* utteranceId 找同一句」的匹配因此全部失效:端到端有 ~2.8s 语义延迟,第 N+1 句的
* 中间结果必然赶在第 N 句译文回来之前,落到同一条未闭合的记录上 —— 表现是原文是
* 后一句、译文是前一句,字幕越说越乱。
*/
interface Callback { interface Callback {
/** 会话开始 */ /** 会话开始 */
fun onSessionStarted(sessionId: String) fun onSessionStarted(sessionId: String)
/** 增量文本到达(大模型回复的文本/翻译结果) */ /** 增量文本到达(大模型回复的文本/翻译结果)。[text] 是**本句累计**的译文,不是增量片段 */
fun onPartialText(sessionId: String, text: String) fun onPartialText(sessionId: String, utteranceId: String, text: String)
/** 增量源文本到达(用户语音的实时识别结果) */ /** 增量源文本到达(用户语音的实时识别结果) */
fun onPartialSourceText(sessionId: String, text: String) fun onPartialSourceText(sessionId: String, utteranceId: String, text: String)
/** 源文本最终结果到达(用户一句话识别完成) */ /** 源文本最终结果到达(用户一句话识别完成) */
fun onFinalSourceText(sessionId: String, finalText: String) fun onFinalSourceText(sessionId: String, utteranceId: String, finalText: String)
/** 翻译/回复最终结果到达(大模型回复完成) */ /** 翻译/回复最终结果到达(大模型回复完成) */
fun onFinalTranslatedText(sessionId: String, finalText: String) fun onFinalTranslatedText(sessionId: String, utteranceId: String, finalText: String)
/** 增量音频片段到达(服务端返回的 TTS 音频) */ /** 增量音频片段到达(服务端返回的 TTS 音频) */
fun onPartialAudio(sessionId: String, data: ByteArray) fun onPartialAudio(sessionId: String, data: ByteArray)
@ -103,6 +113,41 @@ class AliyunBailianE2EHelper(
private val isStarted = AtomicBoolean(false) private val isStarted = AtomicBoolean(false)
private val scope = CoroutineScope(Dispatchers.IO + SupervisorJob()) private val scope = CoroutineScope(Dispatchers.IO + SupervisorJob())
// 连接代次:每次 startContinuousConversation 自增一次,WebSocketListener 把自己那一代
// 捕获进闭包,代次对不上就整条回调忽略。
// ⚠️ 少了它,「结束 → 再开始」会被旧连接打断:OkHttp 的 close() 是**异步**的,
// 旧 socket 的 onClosed 完全可能晚于新 socket 的 onOpen 到达,进来就把 isStarted 打回
// false —— 之后 pushAudioData 一路静默丢弃(只有每 100 帧一条 Log.w),表现就是
// 「点了开始,一切日志正常,就是一个字都翻不出来」。
private val connectionGeneration = java.util.concurrent.atomic.AtomicInteger(0)
// ---- 按句切分的 utterance id(见 Callback 的说明)----
// 识别与翻译是两条各自流式、会互相重叠的流,服务端的事件里又不带任何可跨流对齐的
// item id,只能自己配对。
//
// ⚠️ **谁先到是不定的,两种时序都真实出现过**(2026-09-19 真机):
// - 这个模型(qwen3.5-livetranslate)**根本不发源语言的增量识别**,原文只在整句
// 结束时给一次,而译文早在 0.4s 时就开始流式吐字。所以通常是**译文先铸 id**,
// 原文终态到达时必须去**认领**它,不能另铸一个 —— 那正是上一版「原文 #2 / 译文 #1
// 恒定差一格、每句被拆成两条字幕」的原因。
// - 反过来(源终态先到、译文后来)也要支持,于是有 pending 队列。
// [transIdClaimedBySrc] 保证同一个译文 id 只被一个源终态认领:没有它,下一句的原文
// 会再认领同一条,把上一句的原文覆盖掉。
// 响应条目 id -> 它翻译的那个输入条目 id,来自 conversation.item.created 的
// previous_item_id。**这是跨两条流唯一可靠的对齐依据**(2026-09-19 真机抓到):
// 原文流 conversation.item.input_audio_transcription.* 带的是**输入条目** id
// 译文流 response.audio_transcript.* 带的是**响应条目** id
// 而响应条目的 created 事件里 previous_item_id 就是它回应的那个输入条目。
// 有了它就不需要任何「按到达顺序猜」的配对——那条路走不通,因为两条流的**断句
// 本来就不一致**(ASR 把几句合成一段,模型只对其中一部分出译文)。
private val itemParent = LinkedHashMap<String, String>()
private var utteranceCounter = 0
private var currentSrcUtteranceId: String? = null
private var currentTransUtteranceId: String? = null
private var transIdClaimedBySrc = false
private val pendingTransUtteranceIds = ArrayDeque<String>()
// 服务端实际使用的增量/终态文本事件名(首次出现即锁定,见 handleJsonMessage)。 // 服务端实际使用的增量/终态文本事件名(首次出现即锁定,见 handleJsonMessage)。
// 两套命名(.text / .delta)只会用一套,锁定是为了防止服务端兼容层双发时文本翻倍。 // 两套命名(.text / .delta)只会用一套,锁定是为了防止服务端兼容层双发时文本翻倍。
private var partialTextEvent: String? = null private var partialTextEvent: String? = null
@ -113,6 +158,99 @@ class AliyunBailianE2EHelper(
private var pushDroppedCount = 0L private var pushDroppedCount = 0L
private var recvMsgCount = 0L private var recvMsgCount = 0L
private fun nextUtteranceId(): String {
utteranceCounter += 1
return "$sessionId#$utteranceCounter"
}
/** 新会话 / 重连时把按句状态清干净,避免上一轮的 id 串到下一轮。 */
private fun resetUtteranceState() {
itemParent.clear()
utteranceCounter = 0
currentSrcUtteranceId = null
currentTransUtteranceId = null
transIdClaimedBySrc = false
pendingTransUtteranceIds.clear()
}
/** 记下 conversation.item.created 里的条目关系(响应条目 -> 输入条目)。 */
private fun rememberItemParent(json: JSONObject) {
val item = json.optJSONObject("item") ?: return
val id = item.optString("id", "")
val prev = json.optString("previous_item_id", "")
.ifEmpty { item.optString("previous_item_id", "") }
if (id.isEmpty() || prev.isEmpty()) return
itemParent[id] = prev
while (itemParent.size > MAX_ITEM_MAP) {
itemParent.remove(itemParent.keys.first())
}
}
/** 原文流的 utteranceId:直接用事件里的输入条目 id;没有才退回计数器。 */
private fun srcUtteranceIdOf(json: JSONObject, isFinal: Boolean): String {
val itemId = json.optString("item_id", "")
if (itemId.isNotEmpty()) return itemId
return if (isFinal) takeSrcFinalUtteranceId() else takeSrcUtteranceId()
}
/** 译文流的 utteranceId:把响应条目 id 换成它翻的那个输入条目 id。 */
private fun transUtteranceIdOf(json: JSONObject): String {
val itemId = json.optString("item_id", "")
if (itemId.isNotEmpty()) return itemParent[itemId] ?: itemId
return takeTransUtteranceId()
}
// 下面三个是**兜底路径**:服务端不带 item_id 时才用(按到达顺序猜,必然不完美)。
/** 译文流要用的 id:自己的 → 队列里等着的 → 正在识别的那句 → 新铸一个。 */
private fun takeTransUtteranceId(): String {
currentTransUtteranceId?.let { return it }
val queued = pendingTransUtteranceIds.removeFirstOrNull()
val id: String
when {
// 源终态早就发过了,排在队里等译文
queued != null -> { id = queued; transIdClaimedBySrc = true }
// 源正在识别这一句,译文抢先流式吐字,跟着它走
currentSrcUtteranceId != null -> { id = currentSrcUtteranceId!!; transIdClaimedBySrc = true }
// 谁都还没铸:新铸一个,等源终态来认领(这是本模型的常态)
else -> { id = nextUtteranceId(); transIdClaimedBySrc = false }
}
currentTransUtteranceId = id
return id
}
/** 源流要用的 id:自己的 → 认领正在翻译的那条(仅一次) → 新铸一个。
*
* ⚠️ **增量和终态必须走同一个函数**。上一版只让终态来认领,增量那边还是
* `currentSrcUtteranceId ?: nextUtteranceId()` 自己铸——而真机上译文比源的增量
* 还早(译文 #1 先到,源增量随后铸了 #2),于是终态拿到的是增量留下的 #2,
* 认领逻辑根本没机会生效,原文/译文照旧恒定差一格。 */
private fun takeSrcUtteranceId(): String {
currentSrcUtteranceId?.let { return it }
val trans = currentTransUtteranceId
val id: String
if (trans != null && !transIdClaimedBySrc) {
id = trans
transIdClaimedBySrc = true
} else {
id = nextUtteranceId()
}
currentSrcUtteranceId = id
return id
}
/** 源终态:取 id(同上),然后归零;没被译文认领的排进队等译文。 */
private fun takeSrcFinalUtteranceId(): String {
val id = takeSrcUtteranceId()
currentSrcUtteranceId = null
if (id != currentTransUtteranceId) {
pendingTransUtteranceIds.addLast(id)
while (pendingTransUtteranceIds.size > MAX_PENDING_UTTERANCES) {
Log.w(TAG, "pending utterance 队列超长,丢弃 ${pendingTransUtteranceIds.removeFirst()}")
}
}
return id
}
/** /**
* 初始化助手,设置配置与回调。 * 初始化助手,设置配置与回调。
* *
@ -136,6 +274,7 @@ class AliyunBailianE2EHelper(
recvTextBuffer.setLength(0) recvTextBuffer.setLength(0)
fullTextBuffer.setLength(0) fullTextBuffer.setLength(0)
fullAudioBuffer.reset() fullAudioBuffer.reset()
resetUtteranceState()
conf = config.copy(sourceLanguage = config.sourceLanguage, targetLanguage = config.targetLanguage) conf = config.copy(sourceLanguage = config.sourceLanguage, targetLanguage = config.targetLanguage)
callback = cb callback = cb
@ -159,11 +298,13 @@ class AliyunBailianE2EHelper(
if (isStarted.get()) return true if (isStarted.get()) return true
sessionId = UUID.randomUUID().toString() sessionId = UUID.randomUUID().toString()
Log.d(TAG, "startContinuousConversation: sessionId=${sessionId}") val generation = connectionGeneration.incrementAndGet()
Log.d(TAG, "startContinuousConversation: sessionId=${sessionId}, gen=$generation")
recvTextBuffer.setLength(0) recvTextBuffer.setLength(0)
fullTextBuffer.setLength(0) fullTextBuffer.setLength(0)
fullAudioBuffer.reset() fullAudioBuffer.reset()
resetUtteranceState()
// 构造 URL,必须包含 model 参数 // 构造 URL,必须包含 model 参数
// wss://dashscope.aliyuncs.com/api-ws/v1/realtime?model=<模型名> // wss://dashscope.aliyuncs.com/api-ws/v1/realtime?model=<模型名>
@ -188,6 +329,11 @@ class AliyunBailianE2EHelper(
val listener = object : WebSocketListener() { val listener = object : WebSocketListener() {
override fun onOpen(ws: WebSocket, response: Response) { override fun onOpen(ws: WebSocket, response: Response) {
if (generation != connectionGeneration.get()) {
Log.w(TAG, "onOpen: 旧连接(gen=$generation)回调,已被新连接取代,直接关闭")
try { ws.close(1000, "superseded") } catch (_: Exception) {}
return
}
Log.d(TAG, "onOpen: code=${response.code}") Log.d(TAG, "onOpen: code=${response.code}")
webSocket = ws webSocket = ws
isStarted.set(true) isStarted.set(true)
@ -199,6 +345,7 @@ class AliyunBailianE2EHelper(
} }
override fun onMessage(ws: WebSocket, text: String) { override fun onMessage(ws: WebSocket, text: String) {
if (generation != connectionGeneration.get()) return
handleJsonMessage(text) handleJsonMessage(text)
} }
@ -211,6 +358,10 @@ class AliyunBailianE2EHelper(
override fun onFailure(ws: WebSocket, t: Throwable, response: Response?) { override fun onFailure(ws: WebSocket, t: Throwable, response: Response?) {
val msg = t.message ?: "unknown" val msg = t.message ?: "unknown"
val code = response?.code ?: -1 val code = response?.code ?: -1
if (generation != connectionGeneration.get()) {
Log.w(TAG, "onFailure: 旧连接(gen=$generation)的错误,忽略: $msg")
return
}
Log.e(TAG, "onFailure: ${msg} code=${code}") Log.e(TAG, "onFailure: ${msg} code=${code}")
// // 检查是否需要自动重连 (例如网络异常) // // 检查是否需要自动重连 (例如网络异常)
@ -226,7 +377,10 @@ class AliyunBailianE2EHelper(
} }
override fun onClosed(ws: WebSocket, code: Int, reason: String) { override fun onClosed(ws: WebSocket, code: Int, reason: String) {
Log.d(TAG, "onClosed: code=${code} reason=${reason}") Log.d(TAG, "onClosed: code=${code} reason=${reason} gen=$generation")
// ⚠️ 代次对不上说明这是上一条连接的收尾,**绝不能**碰 isStarted:
// 新连接可能已经 onOpen 了,打回 false 就等于把新会话的音频上行掐死。
if (generation != connectionGeneration.get()) return
isStarted.set(false) isStarted.set(false)
callback?.onSessionFinished(sessionId, fullTextBuffer.toString(), fullAudioBuffer.toByteArray()) callback?.onSessionFinished(sessionId, fullTextBuffer.toString(), fullAudioBuffer.toByteArray())
} }
@ -360,6 +514,10 @@ class AliyunBailianE2EHelper(
if (recvMsgCount <= 5 || recvMsgCount % 50 == 1L) { if (recvMsgCount <= 5 || recvMsgCount % 50 == 1L) {
Log.d(TAG, "Received event #$recvMsgCount: type=$type (src=${conf.sourceLanguage}->${conf.targetLanguage})") Log.d(TAG, "Received event #$recvMsgCount: type=$type (src=${conf.sourceLanguage}->${conf.targetLanguage})")
} }
// 记下条目关系(响应条目 -> 输入条目),两条流靠它对齐,见 itemParent。
if (type == "conversation.item.created" || type == "response.output_item.added") {
rememberItemParent(json)
}
when (type) { when (type) {
"error" -> { "error" -> {
@ -379,8 +537,9 @@ class AliyunBailianE2EHelper(
"conversation.item.input_audio_transcription.text" -> { "conversation.item.input_audio_transcription.text" -> {
val txt = json.optString("text", "") val txt = json.optString("text", "")
if (txt.isNotEmpty()) { if (txt.isNotEmpty()) {
Log.d(TAG, "onPartialSourceText: $txt") val uttId = srcUtteranceIdOf(json, isFinal = false)
callback?.onPartialSourceText(sessionId, txt) Log.d(TAG, "onPartialSourceText[$uttId]: $txt")
callback?.onPartialSourceText(sessionId, uttId, txt)
} }
} }
// 源语言识别结果 (Final) // 源语言识别结果 (Final)
@ -404,8 +563,9 @@ class AliyunBailianE2EHelper(
} }
if (finalTxt.isNotEmpty()) { if (finalTxt.isNotEmpty()) {
Log.d(TAG, "sessionId: ${sessionId}, onFinalSourceText: $finalTxt") val uttId = srcUtteranceIdOf(json, isFinal = true)
callback?.onFinalSourceText(sessionId, finalTxt) Log.d(TAG, "sessionId: ${sessionId}, onFinalSourceText[$uttId]: $finalTxt")
callback?.onFinalSourceText(sessionId, uttId, finalTxt)
} }
} }
@ -430,12 +590,25 @@ class AliyunBailianE2EHelper(
Log.i(TAG, "增量文本事件名锁定为: $type") Log.i(TAG, "增量文本事件名锁定为: $type")
} }
if (type == partialTextEvent) { if (type == partialTextEvent) {
val delta = json.optString("delta", "") // ⚠️ **字段名就是语义,别混**(2026-09-19 真机踩实):
.ifEmpty { json.optString("text", "") } // `delta` 字段 = 增量片段,要自己累计;
if (delta.isNotEmpty()) { // `text` 字段 = 服务端已经累计好的本句全文,要整体替换。
recvTextBuffer.append(delta) // 实测这个模型发的是 `response.audio_transcript.text`,把它当增量
Log.d(TAG, "onPartialText: $delta") // 再 append 一遍,字幕就变成「Hello,Hello,Hello,」越滚越长。
callback?.onPartialText(sessionId, delta) val deltaField = json.optString("delta", "")
val textField = json.optString("text", "")
if (deltaField.isNotEmpty()) {
recvTextBuffer.append(deltaField)
} else if (textField.isNotEmpty()) {
recvTextBuffer.setLength(0)
recvTextBuffer.append(textField)
}
if (recvTextBuffer.isNotEmpty()) {
val uttId = transUtteranceIdOf(json)
// 上层的中间结果是整条覆盖上去的,所以这里上报**本句累计**
// (recvTextBuffer),两种字段形状在这里已经归一了。
Log.d(TAG, "onPartialText[$uttId]: $recvTextBuffer")
callback?.onPartialText(sessionId, uttId, recvTextBuffer.toString())
} }
} }
} }
@ -449,8 +622,11 @@ class AliyunBailianE2EHelper(
val transcript = json.optString("transcript", "") val transcript = json.optString("transcript", "")
.ifEmpty { json.optString("text", "") } .ifEmpty { json.optString("text", "") }
val finalText = if (transcript.isNotEmpty()) transcript else recvTextBuffer.toString() val finalText = if (transcript.isNotEmpty()) transcript else recvTextBuffer.toString()
Log.d(TAG, "sessionId: ${sessionId}, onFinalTranslatedText: $finalText") val uttId = transUtteranceIdOf(json)
callback?.onFinalTranslatedText(sessionId, finalText) currentTransUtteranceId = null
transIdClaimedBySrc = false
Log.d(TAG, "sessionId: ${sessionId}, onFinalTranslatedText[$uttId]: $finalText")
callback?.onFinalTranslatedText(sessionId, uttId, finalText)
if (fullTextBuffer.isNotEmpty()) fullTextBuffer.append(" ") if (fullTextBuffer.isNotEmpty()) fullTextBuffer.append(" ")
fullTextBuffer.append(finalText) fullTextBuffer.append(finalText)
@ -549,10 +725,23 @@ class AliyunBailianE2EHelper(
} }
fun stopContinuousConversation(): Boolean { fun stopContinuousConversation(): Boolean {
// LiveTranslate 没有明确的 "stop task" 指令,通常直接 close 连接即可 // LiveTranslate 没有明确的 "stop task" 指令,直接关连接即可。
// 或者发送 commit 强制模型生成(如果处于等待状态) //
// 这里简单处理为关闭连接 // ⚠️ **必须就地把 webSocket/isStarted 清掉,不能等 onClosed**。
webSocket?.close(1000, "User stopped") // OkHttp 的 close() 只是发出 Close 帧,onClosed 要等服务端回帧才触发:
// - 服务端还没回就点了「开始」→ startContinuousConversation 开头那句
// `if (isStarted.get()) return true` 直接短路,一次握手都不发;
// - 回得晚一点 → 旧的 onClosed 落在新连接之后,把 isStarted 打回 false,
// pushAudioData 从此静默丢弃全部音频。
// 两种都表现为「结束之后再也翻不了,且不报任何错」。代次守卫(见
// connectionGeneration)负责挡住迟到的回调,这里负责让重启立刻可用。
val ws = webSocket
webSocket = null
isStarted.set(false)
resetUtteranceState()
try { ws?.close(1000, "User stopped") } catch (e: Exception) {
Log.w(TAG, "stopContinuousConversation: close 失败(忽略): ${e.message}")
}
return true return true
} }
@ -560,6 +749,16 @@ class AliyunBailianE2EHelper(
try { webSocket?.close(1000, "dispose") } catch (_: Exception) {} try { webSocket?.close(1000, "dispose") } catch (_: Exception) {}
webSocket = null webSocket = null
isStarted.set(false) isStarted.set(false)
connectionGeneration.incrementAndGet()
resetUtteranceState()
scope.cancel() scope.cancel()
} }
private companion object {
/** 等译文的句子最多排多少个,见 onFinalSourceText 的兜底说明。 */
const val MAX_PENDING_UTTERANCES = 8
/** 条目关系表上限:一通长电话不能让它无限涨。 */
const val MAX_ITEM_MAP = 64
}
} }

51
apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/AstCallbacks.kt

@ -779,41 +779,46 @@ class AliyunAstCallback(
) )
} }
override fun onPartialSourceText(sessionId: String, text: String) { override fun onPartialSourceText(sessionId: String, utteranceId: String, text: String) {
eventSender.send( eventSender.send(
mapOf( mapOf(
"type" to "recognizing", "type" to "recognizing",
"serviceId" to serviceId, "serviceId" to serviceId,
"direction" to direction, "direction" to direction,
"utteranceId" to sessionId, "utteranceId" to utteranceId,
"text" to text, "text" to text,
"language" to direction.split("->").firstOrNull().orEmpty() "language" to direction.split("->").firstOrNull().orEmpty()
) )
) )
} }
override fun onFinalSourceText(sessionId: String, finalText: String) { override fun onFinalSourceText(sessionId: String, utteranceId: String, finalText: String) {
val key = "$serviceId:$sessionId" // ⚠️ 按 utteranceId 缓存,不是 sessionId:一整通电话只有一个 sessionId,
sourceTextCache[key] = finalText // 拿它做 key 会让每一句都覆盖上一句的原文。
sourceTextCache["$serviceId:$utteranceId"] = finalText
// 缓存只是给译文回调补 originalText 用的,句子没等到译文就会永远留着 —— 加个上限。
if (sourceTextCache.size > MAX_SOURCE_CACHE) {
sourceTextCache.keys.firstOrNull()?.let { sourceTextCache.remove(it) }
}
eventSender.send( eventSender.send(
mapOf( mapOf(
"type" to "recognized", "type" to "recognized",
"serviceId" to serviceId, "serviceId" to serviceId,
"direction" to direction, "direction" to direction,
"utteranceId" to sessionId, "utteranceId" to utteranceId,
"text" to finalText, "text" to finalText,
"language" to direction.split("->").firstOrNull().orEmpty() "language" to direction.split("->").firstOrNull().orEmpty()
) )
) )
} }
override fun onPartialText(sessionId: String, text: String) { override fun onPartialText(sessionId: String, utteranceId: String, text: String) {
eventSender.send( eventSender.send(
mapOf( mapOf(
"type" to "translatedInterim", "type" to "translatedInterim",
"serviceId" to serviceId, "serviceId" to serviceId,
"direction" to direction, "direction" to direction,
"utteranceId" to sessionId, "utteranceId" to utteranceId,
"originalText" to "", "originalText" to "",
"translatedText" to text, "translatedText" to text,
"targetLanguage" to targetLanguage "targetLanguage" to targetLanguage
@ -830,17 +835,14 @@ class AliyunAstCallback(
finalText: String, finalText: String,
finalAudio: ByteArray finalAudio: ByteArray
) { ) {
eventSender.send( // ⚠️ **不发 "translated"**:finalText 是 helper 的 fullTextBuffer——整个会话每一句
mapOf( // `.done` 译文拼起来的全文,而每一句早已各自经 onFinalTranslatedText 上报过。
"type" to "translated", // 它又只有会话级 id、对不上任何条目,Dart 侧只能新建一条「原文为空、译文是全部
"serviceId" to serviceId, // 句子拼接」的记录:通话一结束,界面上就多出 A/B 各一大段看着像「总结」的乱文。
"direction" to direction, // 与 iOS AzureSpeechPlugin.swift 的 AliyunCallbackProxy.onSessionFinished 同改。
"utteranceId" to sessionId, // (2026-09-18 豆包那条已这么修过,阿里这条漏了 —— 而线上跑的正是阿里。)
"originalText" to (sourceTextCache.remove("$serviceId:$sessionId") ?: ""), sourceTextCache.keys.filter { it.startsWith("$serviceId:") }
"translatedText" to finalText, .forEach { sourceTextCache.remove(it) }
"targetLanguage" to targetLanguage
)
)
// if (finalAudio.isNotEmpty()) { // if (finalAudio.isNotEmpty()) {
// audioWriter?.write(finalAudio) // audioWriter?.write(finalAudio)
@ -864,19 +866,22 @@ class AliyunAstCallback(
audioWriter?.markEnd() audioWriter?.markEnd()
} }
override fun onFinalTranslatedText(sessionId: String, finalText: String) { override fun onFinalTranslatedText(sessionId: String, utteranceId: String, finalText: String) {
val key = "$serviceId:$sessionId" val original = sourceTextCache.remove("$serviceId:$utteranceId") ?: ""
val original = sourceTextCache.remove(key) ?: ""
eventSender.send( eventSender.send(
mapOf( mapOf(
"type" to "translated", "type" to "translated",
"serviceId" to serviceId, "serviceId" to serviceId,
"direction" to direction, "direction" to direction,
"utteranceId" to sessionId, "utteranceId" to utteranceId,
"originalText" to original, "originalText" to original,
"translatedText" to finalText, "translatedText" to finalText,
"targetLanguage" to targetLanguage "targetLanguage" to targetLanguage
) )
) )
} }
private companion object {
const val MAX_SOURCE_CACHE = 32
}
} }

53
apps/client/local_plugins/azure_speech/android/src/main/kotlin/com/yunqiinnovation/azure_speech/tools/AudioRecordingForegroundService.kt

@ -7,6 +7,7 @@ import android.app.PendingIntent
import android.app.Service import android.app.Service
import android.content.Context import android.content.Context
import android.content.Intent import android.content.Intent
import android.content.pm.ServiceInfo
import android.os.Build import android.os.Build
import android.os.IBinder import android.os.IBinder
import androidx.core.app.NotificationCompat import androidx.core.app.NotificationCompat
@ -48,9 +49,17 @@ class AudioRecordingForegroundService : Service() {
*/ */
fun stopService(context: Context) { fun stopService(context: Context) {
val intent = Intent(context, AudioRecordingForegroundService::class.java) val intent = Intent(context, AudioRecordingForegroundService::class.java)
intent.action = ACTION_STOP // ⚠️ 用 stopService 而不是「startService + ACTION_STOP」。
context.startService(intent) // 后者在 App 已经切到后台时会被系统直接拒掉:
Log.d(TAG, "发送停止音频录制前台服务指令") // W/ActivityManager: Background start not allowed: service Intent {…STOP_SERVICE…}
// → Dart 侧收到 PlatformException,服务其实没停,通知一直挂着。
// stopService 从后台调用是允许的,效果等价(service 走 onDestroy)。
val stopped = context.stopService(intent)
if (!stopped) {
// 没停成一般是它本来就没在跑;留一条日志,别静默。
Log.d(TAG, "stopService: 服务未在运行")
}
Log.d(TAG, "已停止音频录制前台服务")
} }
} }
@ -60,6 +69,42 @@ class AudioRecordingForegroundService : Service() {
createNotificationChannel() createNotificationChannel()
} }
/**
* 以「麦克风 + 通话」类型进入前台,被系统拒绝则回退到纯麦克风。
*
* ⚠️ 为什么要带通话类型:真机实测(2026-09-19,HarmonyOS / ALN-AL80),
* 纯 microphone(128) 类型的前台服务**照样**被 `Pged-Freezer` 冻结——
* 服务已登记、常驻通知也挂着,按 Home 约 5s 后进程就被冻住,
* WebSocket 当场断,通话翻译整条停摆。通话类型在多数 ROM 上待遇更好。
*
* ⚠️ 为什么必须能回退:Android 14+ 要求用 phoneCall 类型的 App 是默认拨号应用
* 或持有 MANAGE_OWN_CALLS,我们两样都不是。那时 startForeground 会抛
* SecurityException / ForegroundServiceTypeNotAllowedException,
* **不接住就是前台服务起不来**,比原来还糟。
*/
private fun startForegroundCompat(notification: android.app.Notification) {
if (Build.VERSION.SDK_INT < Build.VERSION_CODES.Q) {
startForeground(NOTIFICATION_ID, notification)
return
}
val micOnly = ServiceInfo.FOREGROUND_SERVICE_TYPE_MICROPHONE
val withCall = micOnly or ServiceInfo.FOREGROUND_SERVICE_TYPE_PHONE_CALL
try {
startForeground(NOTIFICATION_ID, notification, withCall)
Log.d(TAG, "前台服务已启动(类型=麦克风+通话)")
} catch (e: Exception) {
Log.w(TAG, "通话类型被拒,回退纯麦克风: ${e.message}")
try {
startForeground(NOTIFICATION_ID, notification, micOnly)
Log.d(TAG, "前台服务已启动(类型=麦克风)")
} catch (e2: Exception) {
// 连纯麦克风都起不来(多半是通知权限被关 + ROM 限制),
// 别让它把整个进程带崩——录音/翻译在前台照样能跑。
Log.e(TAG, "前台服务启动失败: ${e2.message}")
}
}
}
override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int { override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int {
Log.d(TAG, "AudioRecordingForegroundService onStartCommand action=${intent?.action}") Log.d(TAG, "AudioRecordingForegroundService onStartCommand action=${intent?.action}")
@ -70,7 +115,7 @@ class AudioRecordingForegroundService : Service() {
// - 渠道为 IMPORTANCE_LOW,通常不会弹出“抬头横幅(Heads-up)”,但会在下拉通知栏显示常驻通知; // - 渠道为 IMPORTANCE_LOW,通常不会弹出“抬头横幅(Heads-up)”,但会在下拉通知栏显示常驻通知;
// - Android 13+ 如果未授予通知权限 / 用户关闭通知,可能出现通知不可见或启动前台服务失败(不同 ROM 行为可能不同)。 // - Android 13+ 如果未授予通知权限 / 用户关闭通知,可能出现通知不可见或启动前台服务失败(不同 ROM 行为可能不同)。
// 必须立即调用 startForeground 以满足 Android 8.0+ 的要求 // 必须立即调用 startForeground 以满足 Android 8.0+ 的要求
startForeground(NOTIFICATION_ID, notification) startForegroundCompat(notification)
// 检查是否是停止指令 // 检查是否是停止指令
if (intent?.action == ACTION_STOP) { if (intent?.action == ACTION_STOP) {

214
apps/client/local_plugins/azure_speech/ios/azure_speech/Sources/azure_speech/AliyunBailianE2EHelper.swift

@ -9,12 +9,21 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
/** /**
* 回调接口 * 回调接口
*/ */
/// ⚠️ 四个文本回调都带 `utteranceId`:**一句话一个 id**,识别与翻译共用同一个,
/// 上层据此把「原文 + 译文」并进同一条字幕。
///
/// 原来这四个回调只有 sessionId,桥接层就拿它当 utteranceId 上报 —— 而 sessionId
/// 是**每条 WebSocket 一个**,即一整通电话里一条腿只有一个 id。上层那套「按
/// utteranceId 找同一句」的匹配因此全部失效:端到端有 ~2.8s 语义延迟,第 N+1 句的
/// 中间结果必然赶在第 N 句译文回来之前,落到同一条未闭合的记录上 —— 表现是原文是
/// 后一句、译文是前一句,字幕越说越乱。与 Android 侧同构。
protocol Callback: AnyObject { protocol Callback: AnyObject {
func onSessionStarted(sessionId: String) func onSessionStarted(sessionId: String)
func onPartialText(sessionId: String, text: String) /// `text` 是**本句累计**的译文,不是增量片段
func onPartialSourceText(sessionId: String, text: String) func onPartialText(sessionId: String, utteranceId: String, text: String)
func onFinalSourceText(sessionId: String, finalText: String) func onPartialSourceText(sessionId: String, utteranceId: String, text: String)
func onFinalTranslatedText(sessionId: String, finalText: String) func onFinalSourceText(sessionId: String, utteranceId: String, finalText: String)
func onFinalTranslatedText(sessionId: String, utteranceId: String, finalText: String)
func onPartialAudio(sessionId: String, data: Data) func onPartialAudio(sessionId: String, data: Data)
func onSessionFinished(sessionId: String, finalText: String, finalAudio: Data) func onSessionFinished(sessionId: String, finalText: String, finalAudio: Data)
func onSessionError(sessionId: String, code: Int, message: String) func onSessionError(sessionId: String, code: Int, message: String)
@ -50,6 +59,13 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
private var urlSession: URLSession? private var urlSession: URLSession?
private var webSocket: URLSessionWebSocketTask? private var webSocket: URLSessionWebSocketTask?
/// 最近一次真正建立的那个 task。
/// ⚠️ 代理回调原来只判 `session === urlSession`,而 URLSession 是整个 helper 共用的、
/// 换连接时并不会变 —— 于是**旧 task 的 didClose 会被当成当前连接处理**,进来就把
/// isStarted 打回 false。「结束 → 再开始」时旧连接的收尾完全可能晚于新连接的 didOpen,
/// 结果是新会话的音频上行被静默掐死(pushAudioData 全部丢弃且不报错)。
/// 这里按 task 身份判,[webSocket] 不行:stop 时它已被置空,那样连正常收尾都收不到。
private weak var activeTask: URLSessionWebSocketTask?
private var sessionId: String = "" private var sessionId: String = ""
private var fullTextBuffer = "" private var fullTextBuffer = ""
@ -63,6 +79,32 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
private var audioChunkBuffer = Data() private var audioChunkBuffer = Data()
private var isStarted = false private var isStarted = false
// ---- 按句切分的 utterance id(见 Callback 的说明)----
// 识别与翻译是两条各自流式、且会**互相重叠**的流(第 N+1 句开始识别时,第 N 句的
// 译文往往还没回来),服务端的译文事件里又不带任何 item id,所以只能按 FIFO 对齐:
// 一句识别完就把它的 id 排进队,译文按先进先出取用。
/// 响应条目 id -> 它翻译的那个输入条目 id,来自 conversation.item.created 的
/// previous_item_id。**这是跨两条流唯一可靠的对齐依据**(2026-09-19 安卓真机抓到):
/// 原文流 conversation.item.input_audio_transcription.* 带的是**输入条目** id
/// 译文流 response.audio_transcript.* 带的是**响应条目** id
/// 而响应条目的 created 事件里 previous_item_id 就是它回应的那个输入条目。
/// 有了它就不需要任何「按到达顺序猜」的配对——那条路走不通,因为两条流的**断句
/// 本来就不一致**(ASR 把几句合成一段,模型只对其中一部分出译文)。
private var itemParent: [String: String] = [:]
private var itemParentOrder: [String] = []
private let maxItemMap = 64
private var utteranceCounter = 0
private var currentSrcUtteranceId: String?
private var currentTransUtteranceId: String?
/// 同一个译文 id 只允许被一个源终态认领;没有它,下一句的原文会再认领同一条,
/// 把上一句的原文覆盖掉。
private var transIdClaimedBySrc = false
private var pendingTransUtteranceIds: [String] = []
/// 等译文的句子最多排多少个,见 onFinalSourceText 分支的兜底说明。
private let maxPendingUtterances = 8
private var openContinuation: CheckedContinuation<Bool, Never>? private var openContinuation: CheckedContinuation<Bool, Never>?
// isStarted 成立前的音频缓冲,didOpen 后再冲刷,避免首包被丢 // isStarted 成立前的音频缓冲,didOpen 后再冲刷,避免首包被丢
private var pendingAudioChunks: [Data] = [] private var pendingAudioChunks: [Data] = []
@ -130,6 +172,105 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
/** /**
* 启动会话并建立 WebSocket 连接 * 启动会话并建立 WebSocket 连接
*/ */
private func nextUtteranceId() -> String {
utteranceCounter += 1
return "\(sessionId)#\(utteranceCounter)"
}
/// 新会话 / 重连时把按句状态清干净,避免上一轮的 id 串到下一轮。
private func resetUtteranceState() {
itemParent.removeAll()
itemParentOrder.removeAll()
utteranceCounter = 0
currentSrcUtteranceId = nil
currentTransUtteranceId = nil
transIdClaimedBySrc = false
pendingTransUtteranceIds.removeAll()
}
/// 记下 conversation.item.created 里的条目关系(响应条目 -> 输入条目)。
private func rememberItemParent(_ obj: [String: Any]) {
guard let item = obj["item"] as? [String: Any] else { return }
let id = item["id"] as? String ?? ""
var prev = obj["previous_item_id"] as? String ?? ""
if prev.isEmpty { prev = item["previous_item_id"] as? String ?? "" }
guard !id.isEmpty, !prev.isEmpty else { return }
if itemParent[id] == nil { itemParentOrder.append(id) }
itemParent[id] = prev
while itemParentOrder.count > maxItemMap {
let oldest = itemParentOrder.removeFirst()
itemParent.removeValue(forKey: oldest)
}
}
/// 原文流的 utteranceId:直接用事件里的输入条目 id;没有才退回计数器。
private func srcUtteranceId(_ obj: [String: Any], isFinal: Bool) -> String {
if let itemId = obj["item_id"] as? String, !itemId.isEmpty { return itemId }
return isFinal ? takeSrcFinalUtteranceId() : takeSrcUtteranceId()
}
/// 译文流的 utteranceId:把响应条目 id 换成它翻的那个输入条目 id。
private func transUtteranceId(_ obj: [String: Any]) -> String {
if let itemId = obj["item_id"] as? String, !itemId.isEmpty {
return itemParent[itemId] ?? itemId
}
return takeTransUtteranceId()
}
// 下面三个是**兜底路径**:服务端不带 item_id 时才用(按到达顺序猜,必然不完美)。
/// 译文流要用的 id:自己的 → 队列里等着的 → 正在识别的那句 → 新铸一个。
private func takeTransUtteranceId() -> String {
if let cur = currentTransUtteranceId { return cur }
let id: String
if !pendingTransUtteranceIds.isEmpty {
// 源终态早就发过了,排在队里等译文
id = pendingTransUtteranceIds.removeFirst()
transIdClaimedBySrc = true
} else if let srcId = currentSrcUtteranceId {
// 源正在识别这一句,译文抢先流式吐字,跟着它走
id = srcId
transIdClaimedBySrc = true
} else {
// 谁都还没铸:新铸一个,等源终态来认领(这是本模型的常态)
id = nextUtteranceId()
transIdClaimedBySrc = false
}
currentTransUtteranceId = id
return id
}
/// 源流要用的 id:自己的 → 认领正在翻译的那条(仅一次) → 新铸一个。
///
/// ⚠️ **增量和终态必须走同一个函数**。只让终态认领是不够的:真机上译文比源的
/// 增量还早(译文 #1 先到,源增量随后铸了 #2),终态拿到的就是增量留下的 #2,
/// 认领逻辑根本没机会生效,原文/译文照旧恒定差一格。
private func takeSrcUtteranceId() -> String {
if let cur = currentSrcUtteranceId { return cur }
let id: String
if let trans = currentTransUtteranceId, !transIdClaimedBySrc {
id = trans
transIdClaimedBySrc = true
} else {
id = nextUtteranceId()
}
currentSrcUtteranceId = id
return id
}
/// 源终态:取 id(同上),然后归零;没被译文认领的排进队等译文。
private func takeSrcFinalUtteranceId() -> String {
let id = takeSrcUtteranceId()
currentSrcUtteranceId = nil
if id != currentTransUtteranceId {
pendingTransUtteranceIds.append(id)
while pendingTransUtteranceIds.count > maxPendingUtterances {
let dropped = pendingTransUtteranceIds.removeFirst()
os_log("pending utterance 队列超长,丢弃 %{public}@", log: log, type: .info, dropped)
}
}
return id
}
func startContinuousConversation() -> Bool { func startContinuousConversation() -> Bool {
guard !isStarted else { guard !isStarted else {
os_log("Already started, ignore start request", log: log, type: .info) os_log("Already started, ignore start request", log: log, type: .info)
@ -169,13 +310,16 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
finalTextEvent = nil finalTextEvent = nil
fullAudioBuffer.removeAll() fullAudioBuffer.removeAll()
audioChunkBuffer.removeAll() audioChunkBuffer.removeAll()
resetUtteranceState()
var request = URLRequest(url: url) var request = URLRequest(url: url)
request.timeoutInterval = 60 request.timeoutInterval = 60
request.addValue("Bearer \(conf.apiKey)", forHTTPHeaderField: "Authorization") request.addValue("Bearer \(conf.apiKey)", forHTTPHeaderField: "Authorization")
webSocket = session.webSocketTask(with: request) let task = session.webSocketTask(with: request)
webSocket?.resume() webSocket = task
activeTask = task
task.resume()
// isStarted 在 didOpenWithProtocol 中设置,确保 session.update 先于音频数据发送 // isStarted 在 didOpenWithProtocol 中设置,确保 session.update 先于音频数据发送
return true return true
} }
@ -294,6 +438,12 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
os_log("handleJsonMessage: type=%{public}@", log: log, type: .info, type) os_log("handleJsonMessage: type=%{public}@", log: log, type: .info, type)
// 两条流的对齐全靠它(见 itemParent 的说明),必须在分发之前记下来:
// 响应条目的 created 事件总是先于它自己的 audio_transcript 事件到达。
if type == "conversation.item.created" || type == "response.output_item.added" {
rememberItemParent(obj)
}
switch type { switch type {
case "error": case "error":
if let errorObj = obj["error"] as? [String: Any] { if let errorObj = obj["error"] as? [String: Any] {
@ -308,8 +458,9 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
os_log("Session updated", log: log, type: .info) os_log("Session updated", log: log, type: .info)
case "conversation.item.input_audio_transcription.text": case "conversation.item.input_audio_transcription.text":
if let txt = obj["text"] as? String, !txt.isEmpty { if let txt = obj["text"] as? String, !txt.isEmpty {
let uttId = srcUtteranceId(obj, isFinal: false)
// 高频事件,取消 info 日志 // 高频事件,取消 info 日志
callback?.onPartialSourceText(sessionId: sessionId, text: txt) callback?.onPartialSourceText(sessionId: sessionId, utteranceId: uttId, text: txt)
} }
case "conversation.item.input_audio_transcription.completed": case "conversation.item.input_audio_transcription.completed":
var finalTxt = "" var finalTxt = ""
@ -326,8 +477,9 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
finalTxt = obj["transcript"] as? String ?? "" finalTxt = obj["transcript"] as? String ?? ""
} }
if !finalTxt.isEmpty { if !finalTxt.isEmpty {
os_log("onFinalSourceText: %{public}@", log: log, type: .info, finalTxt) let uttId = srcUtteranceId(obj, isFinal: true)
callback?.onFinalSourceText(sessionId: sessionId, finalText: finalTxt) os_log("onFinalSourceText[%{public}@]: %{public}@", log: log, type: .info, uttId, finalTxt)
callback?.onFinalSourceText(sessionId: sessionId, utteranceId: uttId, finalText: finalTxt)
} }
// ⚠️ 增量文本的事件名有两套,之前只认错的那一套: // ⚠️ 增量文本的事件名有两套,之前只认错的那一套:
// 实时语音翻译(livetranslate):audio+text 模态是 // 实时语音翻译(livetranslate):audio+text 模态是
@ -349,13 +501,24 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
os_log("增量文本事件名锁定为: %{public}@", log: log, type: .info, type) os_log("增量文本事件名锁定为: %{public}@", log: log, type: .info, type)
} }
if type == partialTextEvent { if type == partialTextEvent {
// ⚠️ **字段名就是语义,别混**(2026-09-19 安卓真机踩实,iOS 同构):
// `delta` 字段 = 增量片段,要自己累计;
// `text` 字段 = 服务端已经累计好的本句全文,要整体替换。
// 实测这个模型发的是 `response.audio_transcript.text`,把它当增量
// 再 append 一遍,字幕就变成「Hello,Hello,Hello,」越滚越长。
let deltaField = obj["delta"] as? String ?? "" let deltaField = obj["delta"] as? String ?? ""
let textField = obj["text"] as? String ?? "" let textField = obj["text"] as? String ?? ""
let delta = deltaField.isEmpty ? textField : deltaField if !deltaField.isEmpty {
if !delta.isEmpty { recvTextBuffer.append(deltaField)
recvTextBuffer.append(delta) } else if !textField.isEmpty {
recvTextBuffer = textField
}
if !recvTextBuffer.isEmpty {
let uttId = transUtteranceId(obj)
// 上层的中间结果是整条覆盖上去的,所以这里上报**本句累计**
// (recvTextBuffer),两种字段形状在这里已经归一了。
// 高频事件,取消 info 日志 // 高频事件,取消 info 日志
callback?.onPartialText(sessionId: sessionId, text: delta) callback?.onPartialText(sessionId: sessionId, utteranceId: uttId, text: recvTextBuffer)
} }
} }
// 同上,`.done` 也按模态分两个名字;字段名两种都取。 // 同上,`.done` 也按模态分两个名字;字段名两种都取。
@ -367,8 +530,11 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
let textField = obj["text"] as? String ?? "" let textField = obj["text"] as? String ?? ""
let transcript = transcriptField.isEmpty ? textField : transcriptField let transcript = transcriptField.isEmpty ? textField : transcriptField
let finalText = transcript.isEmpty ? recvTextBuffer : transcript let finalText = transcript.isEmpty ? recvTextBuffer : transcript
os_log("onFinalTranslatedText: %{public}@", log: log, type: .info, finalText) let uttId = transUtteranceId(obj)
callback?.onFinalTranslatedText(sessionId: sessionId, finalText: finalText) currentTransUtteranceId = nil
transIdClaimedBySrc = false
os_log("onFinalTranslatedText[%{public}@]: %{public}@", log: log, type: .info, uttId, finalText)
callback?.onFinalTranslatedText(sessionId: sessionId, utteranceId: uttId, finalText: finalText)
if !fullTextBuffer.isEmpty { if !fullTextBuffer.isEmpty {
fullTextBuffer.append(" ") fullTextBuffer.append(" ")
@ -457,6 +623,7 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
} }
webSocket = nil webSocket = nil
isStarted = false isStarted = false
resetUtteranceState()
pendingLock.lock() pendingLock.lock()
pendingAudioChunks.removeAll() pendingAudioChunks.removeAll()
pendingLock.unlock() pendingLock.unlock()
@ -489,8 +656,8 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
* WebSocket 打开回调 * WebSocket 打开回调
*/ */
func urlSession(_ session: URLSession, webSocketTask: URLSessionWebSocketTask, didOpenWithProtocol protocol: String?) { func urlSession(_ session: URLSession, webSocketTask: URLSessionWebSocketTask, didOpenWithProtocol protocol: String?) {
// 忽略旧会话的回调 // 忽略旧会话 / 旧连接的回调(见 activeTask 的说明)
guard session === urlSession else { return } guard session === urlSession, webSocketTask === activeTask else { return }
os_log("WebSocket didOpen", log: log, type: .info) os_log("WebSocket didOpen", log: log, type: .info)
sendSessionUpdate() sendSessionUpdate()
isStarted = true // 在 session.update 发送后才允许推送音频,与 Android 行为对齐 isStarted = true // 在 session.update 发送后才允许推送音频,与 Android 行为对齐
@ -518,8 +685,11 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
* WebSocket 接收循环 * WebSocket 接收循环
*/ */
private func receiveLoop() { private func receiveLoop() {
webSocket?.receive { [weak self] result in guard let task = webSocket else { return }
task.receive { [weak self] result in
guard let self = self else { return } guard let self = self else { return }
// 这一帧属于旧连接就整条丢掉,别把上一轮的残留混进新会话
guard task === self.activeTask else { return }
switch result { switch result {
case .failure(let error): case .failure(let error):
// 忽略旧连接的错误 // 忽略旧连接的错误
@ -546,8 +716,8 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
* WebSocket 关闭回调 * WebSocket 关闭回调
*/ */
func urlSession(_ session: URLSession, webSocketTask: URLSessionWebSocketTask, didCloseWith closeCode: URLSessionWebSocketTask.CloseCode, reason: Data?) { func urlSession(_ session: URLSession, webSocketTask: URLSessionWebSocketTask, didCloseWith closeCode: URLSessionWebSocketTask.CloseCode, reason: Data?) {
// 忽略旧会话的回调 // 忽略旧会话 / 旧连接的回调(见 activeTask 的说明)
guard session === urlSession else { return } guard session === urlSession, webSocketTask === activeTask else { return }
let reasonStr = String(data: reason ?? Data(), encoding: .utf8) ?? "" let reasonStr = String(data: reason ?? Data(), encoding: .utf8) ?? ""
os_log("WebSocket didClose code=%{public}d reason=%{public}@", log: log, type: .info, closeCode.rawValue, reasonStr) os_log("WebSocket didClose code=%{public}d reason=%{public}@", log: log, type: .info, closeCode.rawValue, reasonStr)
isStarted = false isStarted = false
@ -563,8 +733,8 @@ class AliyunBailianE2EHelper: NSObject, URLSessionWebSocketDelegate {
* 任务完成回调(错误处理) * 任务完成回调(错误处理)
*/ */
func urlSession(_ session: URLSession, task: URLSessionTask, didCompleteWithError error: Error?) { func urlSession(_ session: URLSession, task: URLSessionTask, didCompleteWithError error: Error?) {
// 忽略旧会话的回调 // 忽略旧会话 / 旧连接的回调(见 activeTask 的说明)
guard session === urlSession else { return } guard session === urlSession, task === activeTask else { return }
if let e = error { if let e = error {
os_log("WebSocket task error: %{public}@", log: log, type: .error, e.localizedDescription) os_log("WebSocket task error: %{public}@", log: log, type: .error, e.localizedDescription)
callback?.onSessionError(sessionId: sessionId, code: 1012, message: e.localizedDescription) callback?.onSessionError(sessionId: sessionId, code: 1012, message: e.localizedDescription)

42
apps/client/local_plugins/azure_speech/ios/azure_speech/Sources/azure_speech/AzureSpeechPlugin.swift

@ -1969,7 +1969,7 @@ private class AliyunCallbackProxy: AliyunBailianE2EHelper.Callback {
]) ])
} }
func onPartialSourceText(sessionId: String, text: String) { func onPartialSourceText(sessionId: String, utteranceId: String, text: String) {
// 高频事件,取消 info 日志 // 高频事件,取消 info 日志
plugin?.sendAstEvent([ plugin?.sendAstEvent([
"type": "recognizing", "type": "recognizing",
@ -1977,25 +1977,31 @@ private class AliyunCallbackProxy: AliyunBailianE2EHelper.Callback {
"direction": direction, "direction": direction,
"text": text, "text": text,
"language": sourceLanguage, "language": sourceLanguage,
"utteranceId": sessionId "utteranceId": utteranceId
]) ])
} }
func onFinalSourceText(sessionId: String, finalText: String) { func onFinalSourceText(sessionId: String, utteranceId: String, finalText: String) {
os_log("[AliyunCallback-%{public}@] onFinalSourceText text=%{public}@", os_log("[AliyunCallback-%{public}@] onFinalSourceText[%{public}@] text=%{public}@",
log: ctLog, type: .info, serviceId, finalText) log: ctLog, type: .info, serviceId, utteranceId, finalText)
sourceTextCache["\(serviceId):\(sessionId)"] = finalText // ⚠️ 按 utteranceId 缓存,不是 sessionId:一整通电话只有一个 sessionId,
// 拿它做 key 会让每一句都覆盖上一句的原文。
sourceTextCache["\(serviceId):\(utteranceId)"] = finalText
// 缓存只是给译文回调补 originalText 用的,句子没等到译文就会永远留着 —— 加个上限。
if sourceTextCache.count > 32, let oldest = sourceTextCache.keys.first {
sourceTextCache.removeValue(forKey: oldest)
}
plugin?.sendAstEvent([ plugin?.sendAstEvent([
"type": "recognized", "type": "recognized",
"serviceId": serviceId, "serviceId": serviceId,
"direction": direction, "direction": direction,
"text": finalText, "text": finalText,
"language": sourceLanguage, "language": sourceLanguage,
"utteranceId": sessionId "utteranceId": utteranceId
]) ])
} }
func onPartialText(sessionId: String, text: String) { func onPartialText(sessionId: String, utteranceId: String, text: String) {
// 高频事件,取消 info 日志 // 高频事件,取消 info 日志
plugin?.sendAstEvent([ plugin?.sendAstEvent([
"type": "translatedInterim", "type": "translatedInterim",
@ -2004,7 +2010,7 @@ private class AliyunCallbackProxy: AliyunBailianE2EHelper.Callback {
"translatedText": text, "translatedText": text,
"originalText": "", "originalText": "",
"targetLanguage": targetLanguage, "targetLanguage": targetLanguage,
"utteranceId": sessionId "utteranceId": utteranceId
]) ])
} }
@ -2020,10 +2026,9 @@ private class AliyunCallbackProxy: AliyunBailianE2EHelper.Callback {
} }
func onSessionFinished(sessionId: String, finalText: String, finalAudio: Data) { func onSessionFinished(sessionId: String, finalText: String, finalAudio: Data) {
let key = "\(serviceId):\(sessionId)" sourceTextCache.removeAll()
let original = sourceTextCache.removeValue(forKey: key) ?? "" os_log("[AliyunCallback-%{public}@] onSessionFinished finalText=%{public}@ audioLen=%d",
os_log("[AliyunCallback-%{public}@] onSessionFinished finalText=%{public}@ original=%{public}@ audioLen=%d", log: ctLog, type: .info, serviceId, finalText, finalAudio.count)
log: ctLog, type: .info, serviceId, finalText, original, finalAudio.count)
// 段尾合约(与 Android markEnd 对齐):**不**重复写 finalAudio,只 emit 空 PCM + isFinal=true。 // 段尾合约(与 Android markEnd 对齐):**不**重复写 finalAudio,只 emit 空 PCM + isFinal=true。
plugin?.pushAstTtsFrame(leg: serviceId, pcm: Data(), isFinal: true) plugin?.pushAstTtsFrame(leg: serviceId, pcm: Data(), isFinal: true)
// ⚠️ 这里**不能**再发 "translated" 事件。finalText 是 helper 的 fullTextBuffer—— // ⚠️ 这里**不能**再发 "translated" 事件。finalText 是 helper 的 fullTextBuffer——
@ -2048,18 +2053,17 @@ private class AliyunCallbackProxy: AliyunBailianE2EHelper.Callback {
]) ])
} }
func onFinalTranslatedText(sessionId: String, finalText: String) { func onFinalTranslatedText(sessionId: String, utteranceId: String, finalText: String) {
let key = "\(serviceId):\(sessionId)" let original = sourceTextCache.removeValue(forKey: "\(serviceId):\(utteranceId)") ?? ""
let original = sourceTextCache.removeValue(forKey: key) ?? "" os_log("[AliyunCallback-%{public}@] onFinalTranslatedText[%{public}@] finalText=%{public}@ original=%{public}@",
os_log("[AliyunCallback-%{public}@] onFinalTranslatedText finalText=%{public}@ original=%{public}@", log: ctLog, type: .info, serviceId, utteranceId, finalText, original)
log: ctLog, type: .info, serviceId, finalText, original)
plugin?.sendAstEvent([ plugin?.sendAstEvent([
"type": "translated", "type": "translated",
"serviceId": serviceId, "serviceId": serviceId,
"direction": direction, "direction": direction,
"translatedText": finalText, "translatedText": finalText,
"targetLanguage": targetLanguage, "targetLanguage": targetLanguage,
"utteranceId": sessionId, "utteranceId": utteranceId,
"originalText": original "originalText": original
]) ])
} }

117
apps/client/test/peer_speech_gate_test.dart

@ -0,0 +1,117 @@
import 'dart:math';
import 'dart:typed_data';
import 'package:flutter_test/flutter_test.dart';
import 'package:eaimar/data/utils/peer_speech_gate.dart';
/// 造一帧 20ms/16k/单声道 PCM16(320 样本 = 640 字节),幅度固定。
Uint8List frame(int amplitude) {
const samples = 320;
final b = ByteData(samples * 2);
for (var i = 0; i < samples; i++) {
// 用正弦而不是常数:常数信号的 RMS 等于幅值,掩盖不了算错的情况
final v = (amplitude * sin(2 * pi * 8 * i / samples)).round();
b.setInt16(i * 2, v.clamp(-32768, 32767), Endian.little);
}
return b.buffer.asUint8List();
}
void main() {
group('rmsOfPcm16', () {
test('静音为 0', () => expect(rmsOfPcm16(frame(0)), 0));
test('正弦的 RMS 约为幅值的 0.707 倍', () {
// 3000 * 0.707 ≈ 2121
expect(rmsOfPcm16(frame(3000)), closeTo(2121, 40));
});
test('空帧与半个样本不崩', () {
expect(rmsOfPcm16(Uint8List(0)), 0);
expect(rmsOfPcm16(Uint8List.fromList([0x11])), 0);
});
});
group('PeerSpeechGate 开门条件', () {
test('静音不开门,一帧都不放行', () {
final g = PeerSpeechGate();
for (var i = 0; i < 50; i++) {
expect(g.accept(frame(0)), isEmpty);
}
expect(g.isOpen, isFalse);
});
test('断续的串音尖峰不开门——这是本门限存在的理由', () {
final g = PeerSpeechGate(openFrames: 8);
// 每 3 帧来一个响帧,永远凑不满 8 帧连续
for (var i = 0; i < 90; i++) {
g.accept(frame(i % 3 == 0 ? 5000 : 0));
}
expect(g.isOpen, isFalse, reason: '断续尖峰不该被当成有人在说话');
});
test('连续超阈值 openFrames 帧即开门', () {
final g = PeerSpeechGate(openFrames: 8);
for (var i = 0; i < 7; i++) {
expect(g.accept(frame(5000)), isEmpty, reason: '第 ${i + 1} 帧还不该开');
}
final out = g.accept(frame(5000));
expect(g.isOpen, isTrue);
expect(out, isNotEmpty);
expect(g.openedByFallback, isFalse);
});
});
group('PeerSpeechGate 缓存补发(对端先说话时不能吃掉开头)', () {
test('开门时把此前缓存的帧整段补发,一帧不丢', () {
final g = PeerSpeechGate(openFrames: 8, prerollFrames: 25);
var emitted = 0;
for (var i = 0; i < 8; i++) {
emitted += g.accept(frame(5000)).length;
}
expect(emitted, 8, reason: '开门前的 8 帧必须在开门瞬间补发出来');
});
test('缓存上限生效:更早的帧被丢弃,不会无限涨', () {
final g = PeerSpeechGate(openFrames: 2, prerollFrames: 5);
for (var i = 0; i < 30; i++) {
g.accept(frame(0)); // 静音,只进缓存不开门
}
final out = [g.accept(frame(5000)), g.accept(frame(5000))]
.expand((e) => e)
.toList();
expect(out.length, 5, reason: '只应补发最近 prerollFrames 帧');
});
test('开门之后直通,每帧进一出一', () {
final g = PeerSpeechGate(openFrames: 1);
g.accept(frame(5000));
expect(g.accept(frame(0)).length, 1, reason: '开门后静音帧也要照推');
expect(g.accept(frame(9000)).length, 1);
});
});
group('PeerSpeechGate 兜底与复位', () {
test('阈值一直不过时兜底强开——宁可音色刻错也不能整路不翻译', () {
final g = PeerSpeechGate(rmsThreshold: 30000, fallbackFrames: 20);
List<Uint8List> out = const [];
for (var i = 0; i < 20; i++) {
out = g.accept(frame(100));
}
expect(g.isOpen, isTrue);
expect(g.openedByFallback, isTrue, reason: '要能区分是正常开门还是兜底强开');
expect(out, isNotEmpty, reason: '兜底开门同样要把缓存补发出去');
});
test('reset 之后回到初始态(断线重连会重新刻音色,必须重新把门关上)', () {
final g = PeerSpeechGate(openFrames: 2);
g.accept(frame(5000));
g.accept(frame(5000));
expect(g.isOpen, isTrue);
g.reset();
expect(g.isOpen, isFalse);
expect(g.openedByFallback, isFalse);
expect(g.accept(frame(0)), isEmpty, reason: 'reset 后静音不该直通');
});
});
}
Loading…
Cancel
Save