强制对齐

把已知文本对齐到音频

POST /v1/audio/alignments

传入 file 与对应的 text,用于定位已知文本在音频中的时间戳,而非识别未知语音。示例中关闭了文本归一化以保留原始拼写;省略该配置时默认启用文本归一化。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world." -F language=en \
  -F 'text_normalization={"enabled":false}'

输出示例:

JSON
{
  "task": "align",
  "duration": 2.5,
  "text": "Hello world.",
  "segments": [
    {
      "id": 0,
      "start": 0.2,
      "end": 1.4,
      "text": "Hello world.",
      "words": [
        {
          "word": "Hello",
          "start": 0.2,
          "end": 0.6,
          "score": 0.98
        },
        {
          "word": "world.",
          "start": 0.7,
          "end": 1.4,
          "score": 0.96
        }
      ]
    }
  ],
  "usage": {
    "type": "duration",
    "seconds": 2.5
  }
}

可选参数 start 与 end 必须成对提供(单位为秒)。返回的时间戳仍相对于完整音频起始点,duration 为源文件总时长,usage.seconds 为实际处理的窗口时长。桌面网关额外支持传入 text_path,Headless 服务仅接受 text。高级配置见归一化、词典、保护词及轨道/声道选项。

参数 model 指定对齐模型,模型决定支持的语言范围。省略时默认使用 EdgeSpeak/Lattice-2,它覆盖语言目录里的全部语言、混合语言文本以及 und。EdgeSpeak/Lattice-1 仅覆盖中文、英文与德文,只在显式传 model=EdgeSpeak/Lattice-1 时使用;在它上面传其他 language 或 und 会被拒绝。这两个模型均可在 GET /v1/models 中查询,其 supported_endpoints 均包含 /v1/audio/alignments,default_for 标出当前缺省档。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@weather.wav -F text="今日はいい天気ですね。" \
  -F language=ja

输出示例:

JSON
{
  "task": "align",
  "duration": 2.1,
  "language": "ja",
  "text": "今日はいい天気ですね。",
  "segments": [
    {
      "id": 0,
      "start": 0.3,
      "end": 1.9,
      "text": "今日はいい天気ですね。",
      "words": [
        { "word": "今日は", "start": 0.3, "end": 0.7, "score": 0.97 },
        { "word": "いい", "start": 0.7, "end": 0.9, "score": 0.95 },
        { "word": "天気", "start": 0.9, "end": 1.4, "score": 0.98 },
        { "word": "ですね。", "start": 1.4, "end": 1.9, "score": 0.94 }
      ]
    }
  ],
  "usage": { "type": "duration", "seconds": 2.1 }
}

显式指定 model=EdgeSpeak/Lattice-1 再传 language=ja,接口会返回 HTTP 422 language_unsupported。遇到此错误时应更换模型,而非修改语言代码。

language 是传给 Tokenizer 的语言提示,而非硬性门槛。传入 zh、yue、ja 等代码表示显式指定语言;省略表示由运行时根据参考文本判定语言,判定不出时按不带语言对齐,不会报错;只看文本可能混淆书写系统相同的语言(如都用汉字书写的粤语和普通话),所以已知语言时请显式传入;传入 und 表示不判定语言,也不预设是哪门语言,适用于混合语言、音译文本以及没有对应语言代码的录音。und 也并非全部回退到 G2P:分词器会在所有语言的词典中逐词匹配(例如中英混排中的英文单词仍会命中英文词条,所有词典均未收录的词才会回退到通用音素集),文本归一化则按文本证据逐段选择读法,例如 我想要$100, I want $100. 会分别处理为 我想要一百美元, I want one hundred dollars.。语言目录之外的代码(如 nan、hak、wuu 等)在带 G2P 兜底的模型上同样支持,会保留文字与地区子标签(如 nan-Hant),但缺乏词典读法支持。词典里查不到的代码,效果弱于 und。

文本归一化的规则集只覆盖语言目录的一部分。对于目录外的代码,以及像粤语 yue 这样在表内但缺乏规则的语种,默认请求会自动跳过文本归一化,避免整次对齐失败;但若显式指定了 top_k、classes、pins、apply_dictionary 或 dictionary_ids,接口会返回 HTTP 400,不会静默丢弃你的配置。缺省的 EdgeSpeak/Lattice-2 支持上述全部输入;显式选择的 EdgeSpeak/Lattice-1 没有 G2P 兜底,会以 HTTP 400 bad_request 拒绝 und 与目录外代码,以 HTTP 422 language_unsupported 拒绝未覆盖的语言。

读取对齐进度

在请求中添加请求头 progress: 1。响应格式为 NDJSON,每行需独立解析为 JSON,非 SSE 格式。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -N -H "progress: 1" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world." \
  -F 'text_normalization={"enabled":false}'

输出示例:

{  "progress": 0.5}
{  "progress": 1.0}
{"task":"align","duration":2.5,"text":"Hello world.","segments":[{"id":0,"start":0.2,"end":1.4,"text":"Hello world.","words":[{"word":"Hello","start":0.2,"end":0.6,"score":0.98},{"word":"world.","start":0.7,"end":1.4,"score":0.96}]}],"usage":{"type":"duration","seconds":2.5}}

包含 "task": "align" 的最终文稿行才是有效结果。仅收到进度 1.0 并不代表对齐成功。数据流有三种结束方式:

  • "task": "align" 行:对齐成功;
  • {"error": {...}} 行:对齐失败,形状与 422 响应体相同,但数据流的 HTTP 状态仍是 200;
  • 数据流关闭时两者都没有:按失败处理。

使用 search_effort=auto 时,{"stage":"extended_search"} 这一行表示开始扩大搜索范围,之后进度从 0 重新计。

并发请求与等待

Headless 服务允许并发提交对齐请求,上传槽或执行槽占满时自动等待调度,不再仅因槽满而立即失败。槽位参数、默认值与作用范围见 CLI 并发参数。

请求与响应格式同上。建议根据音频时长与排队情况调大客户端超时,例如预留 30 分钟:

Shell
curl --fail-with-body --max-time 1800 "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world."

任务完成后返回包含 task: "align"、segments 与 usage 的对齐结果。设置客户端超时时,需将排队、模型加载与推理三段时间一并计入;在等待执行槽期间,进度模式不会输出进度行。

处理对齐失败

本地引擎没能把文本对齐到音频时,请求返回 HTTP 422 和 error.alignment 对象,不返回结果,也不计入 usage。请按 error.code 和 error.alignment 分支;message 是英文诊断文本,可能变化。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@song.mp3 -F text="$(cat lyrics.txt)"

输出示例(HTTP 422;数值仅作示意):

JSON
{
  "error": {
    "message": "Alignment could not be completed; the search stopped before the end of the text.",
    "type": "invalid_request_error",
    "code": "alignment_failed",
    "alignment": {
      "reason": "search_incomplete",
      "search_effort_used": "standard",
      "search_path": "whole_audio",
      "search_limit": "none",
      "retry_recommended": true,
      "estimated_extra_bytes": 3221225472
    }
  }
}
错误码含义
alignment_failed对齐跑了但没能完成,细节见 reason。
alignment_search_budget_exceeded扩大搜索范围所需的内存超出本机可用范围,因此没有运行。可以截短音频后再试,例如用 start / end。
字段说明
reasonno_path 或 no_words:音频里没有找到与文本对应的内容;collapsed:词被挤进了不合理的极短时长;search_incomplete:搜索在文本结尾之前停下,可能是伴奏或长静音所致。
search_effort_used本次运行的搜索档位:standard 或 extended。
search_pathwhole_audio 表示整段一次搜索;streaming 表示长音频分段对齐。
search_limitnone,或搜索碰到的限制:budget(内存预算)、graph_size(搜索空间过大,无法扩大搜索)、estimate_overflow(无法算出内存估算)、streaming(分段对齐没有扩大搜索的版本)。
retry_recommended始终出现。只有码是 alignment_failed、reason 是 search_incomplete、本次跑的是 standard、没有碰到 none 以外的限制,且扩大搜索的估算落在本机内存预算内时才为 true。
estimated_extra_bytes扩大搜索预计额外占用的内存。只覆盖搜索本身,是估算,不是上限。

除 retry_recommended 外,服务不知道的字段会省略。reason 不是 search_incomplete 时,请确认音频与文本对应同一段内容;扩大搜索范围帮不上忙。

对齐成功但没有对上任何词(例如文本里没有可对齐的内容)时,返回 HTTP 200 且不带 segments;segments 缺失时按空列表读取。

扩大搜索范围重试

multipart 字段 search_effort 决定搜索范围:

取值行为
standard默认值,传空值也按它处理。按标准范围搜索一次。
extended按更大的范围搜索一次。更慢,更耗内存;失败那次的 estimated_extra_bytes 给出大约多用多少。
auto先跑 standard;仅当它失败且 retry_recommended: true 时,再跑一次 extended。其他失败原样返回。

其他取值返回 400 bad_request,param 为 "search_effort"。error.alignment.retry_recommended 为 true 时,用 search_effort=extended 重发同一请求:

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@song.mp3 -F text="$(cat lyrics.txt)" \
  -F search_effort=extended

输出示例(响应节选):

JSON
{
  "task": "align",
  "duration": 214.6,
  "text": "...",
  "segments": [{ "id": 0, "start": 12.1, "end": 16.8, "text": "...", "words": [{ "word": "...", "start": 12.1, "end": 12.5, "score": 0.93 }] }],
  "search_effort_used": "extended",
  "usage": { "type": "duration", "seconds": 214.6 }
}

服务报告了实际档位时,成功响应带 search_effort_used。使用 search_effort=auto 时,只有第二遍确实跑了才出现 search_retried: true,否则省略。extended 仍然失败时,retry_recommended 为 false。

继续阅读:转录与说话人 · API 索引