强制对齐
把已知文本对齐到音频
POST /v1/audio/alignments
传入 file 与对应的 text,用于定位已知文本在音频中的时间戳,而非识别未知语音。示例中关闭了文本归一化以保留原始拼写;省略该配置时默认启用文本归一化。
请求示例:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@hello.wav -F text="Hello world." -F language=en \
-F 'text_normalization={"enabled":false}'输出示例:
{
"task": "align",
"duration": 2.5,
"text": "Hello world.",
"segments": [
{
"id": 0,
"start": 0.2,
"end": 1.4,
"text": "Hello world.",
"words": [
{
"word": "Hello",
"start": 0.2,
"end": 0.6,
"score": 0.98
},
{
"word": "world.",
"start": 0.7,
"end": 1.4,
"score": 0.96
}
]
}
],
"usage": {
"type": "duration",
"seconds": 2.5
}
}可选参数 start 与 end 必须成对提供(单位为秒)。返回的时间戳仍相对于完整音频起始点,duration 为源文件总时长,usage.seconds 为实际处理的窗口时长。桌面网关额外支持传入 text_path,Headless 服务仅接受 text。高级配置见归一化、词典、保护词及轨道/声道选项。
参数 model 指定对齐模型,模型决定支持的语言范围。省略时默认使用 EdgeSpeak/Lattice-2,它覆盖语言目录里的全部语言、混合语言文本以及 und。EdgeSpeak/Lattice-1 仅覆盖中文、英文与德文,只在显式传 model=EdgeSpeak/Lattice-1 时使用;在它上面传其他 language 或 und 会被拒绝。这两个模型均可在 GET /v1/models 中查询,其 supported_endpoints 均包含 /v1/audio/alignments,default_for 标出当前缺省档。
请求示例:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@weather.wav -F text="今日はいい天気ですね。" \
-F language=ja输出示例:
{
"task": "align",
"duration": 2.1,
"language": "ja",
"text": "今日はいい天気ですね。",
"segments": [
{
"id": 0,
"start": 0.3,
"end": 1.9,
"text": "今日はいい天気ですね。",
"words": [
{ "word": "今日は", "start": 0.3, "end": 0.7, "score": 0.97 },
{ "word": "いい", "start": 0.7, "end": 0.9, "score": 0.95 },
{ "word": "天気", "start": 0.9, "end": 1.4, "score": 0.98 },
{ "word": "ですね。", "start": 1.4, "end": 1.9, "score": 0.94 }
]
}
],
"usage": { "type": "duration", "seconds": 2.1 }
}显式指定 model=EdgeSpeak/Lattice-1 再传 language=ja,接口会返回 HTTP 422 language_unsupported。遇到此错误时应更换模型,而非修改语言代码。
language 是传给 Tokenizer 的语言提示,而非硬性门槛。传入 zh、yue、ja 等代码表示显式指定语言;省略表示由运行时根据参考文本判定语言,判定不出时按不带语言对齐,不会报错;只看文本可能混淆书写系统相同的语言(如都用汉字书写的粤语和普通话),所以已知语言时请显式传入;传入 und 表示不判定语言,也不预设是哪门语言,适用于混合语言、音译文本以及没有对应语言代码的录音。und 也并非全部回退到 G2P:分词器会在所有语言的词典中逐词匹配(例如中英混排中的英文单词仍会命中英文词条,所有词典均未收录的词才会回退到通用音素集),文本归一化则按文本证据逐段选择读法,例如 我想要$100, I want $100. 会分别处理为 我想要一百美元, I want one hundred dollars.。语言目录之外的代码(如 nan、hak、wuu 等)在带 G2P 兜底的模型上同样支持,会保留文字与地区子标签(如 nan-Hant),但缺乏词典读法支持。词典里查不到的代码,效果弱于 und。
文本归一化的规则集只覆盖语言目录的一部分。对于目录外的代码,以及像粤语 yue 这样在表内但缺乏规则的语种,默认请求会自动跳过文本归一化,避免整次对齐失败;但若显式指定了 top_k、classes、pins、apply_dictionary 或 dictionary_ids,接口会返回 HTTP 400,不会静默丢弃你的配置。缺省的 EdgeSpeak/Lattice-2 支持上述全部输入;显式选择的 EdgeSpeak/Lattice-1 没有 G2P 兜底,会以 HTTP 400 bad_request 拒绝 und 与目录外代码,以 HTTP 422 language_unsupported 拒绝未覆盖的语言。
读取对齐进度
在请求中添加请求头 progress: 1。响应格式为 NDJSON,每行需独立解析为 JSON,非 SSE 格式。
请求示例:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-N -H "progress: 1" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@hello.wav -F text="Hello world." \
-F 'text_normalization={"enabled":false}'输出示例:
{ "progress": 0.5}
{ "progress": 1.0}
{"task":"align","duration":2.5,"text":"Hello world.","segments":[{"id":0,"start":0.2,"end":1.4,"text":"Hello world.","words":[{"word":"Hello","start":0.2,"end":0.6,"score":0.98},{"word":"world.","start":0.7,"end":1.4,"score":0.96}]}],"usage":{"type":"duration","seconds":2.5}}包含 "task": "align" 的最终文稿行才是有效结果。仅收到进度 1.0 并不代表对齐成功。数据流有三种结束方式:
"task": "align"行:对齐成功;{"error": {...}}行:对齐失败,形状与 422 响应体相同,但数据流的 HTTP 状态仍是 200;- 数据流关闭时两者都没有:按失败处理。
使用 search_effort=auto 时,{"stage":"extended_search"} 这一行表示开始扩大搜索范围,之后进度从 0 重新计。
并发请求与等待
Headless 服务允许并发提交对齐请求,上传槽或执行槽占满时自动等待调度,不再仅因槽满而立即失败。槽位参数、默认值与作用范围见 CLI 并发参数。
请求与响应格式同上。建议根据音频时长与排队情况调大客户端超时,例如预留 30 分钟:
curl --fail-with-body --max-time 1800 "$EDGESPEAK_BASE_URL/audio/alignments" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@hello.wav -F text="Hello world."任务完成后返回包含 task: "align"、segments 与 usage 的对齐结果。设置客户端超时时,需将排队、模型加载与推理三段时间一并计入;在等待执行槽期间,进度模式不会输出进度行。
处理对齐失败
本地引擎没能把文本对齐到音频时,请求返回 HTTP 422 和 error.alignment 对象,不返回结果,也不计入 usage。请按 error.code 和 error.alignment 分支;message 是英文诊断文本,可能变化。
请求示例:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@song.mp3 -F text="$(cat lyrics.txt)"输出示例(HTTP 422;数值仅作示意):
{
"error": {
"message": "Alignment could not be completed; the search stopped before the end of the text.",
"type": "invalid_request_error",
"code": "alignment_failed",
"alignment": {
"reason": "search_incomplete",
"search_effort_used": "standard",
"search_path": "whole_audio",
"search_limit": "none",
"retry_recommended": true,
"estimated_extra_bytes": 3221225472
}
}
}| 错误码 | 含义 |
|---|---|
alignment_failed | 对齐跑了但没能完成,细节见 reason。 |
alignment_search_budget_exceeded | 扩大搜索范围所需的内存超出本机可用范围,因此没有运行。可以截短音频后再试,例如用 start / end。 |
| 字段 | 说明 |
|---|---|
reason | no_path 或 no_words:音频里没有找到与文本对应的内容;collapsed:词被挤进了不合理的极短时长;search_incomplete:搜索在文本结尾之前停下,可能是伴奏或长静音所致。 |
search_effort_used | 本次运行的搜索档位:standard 或 extended。 |
search_path | whole_audio 表示整段一次搜索;streaming 表示长音频分段对齐。 |
search_limit | none,或搜索碰到的限制:budget(内存预算)、graph_size(搜索空间过大,无法扩大搜索)、estimate_overflow(无法算出内存估算)、streaming(分段对齐没有扩大搜索的版本)。 |
retry_recommended | 始终出现。只有码是 alignment_failed、reason 是 search_incomplete、本次跑的是 standard、没有碰到 none 以外的限制,且扩大搜索的估算落在本机内存预算内时才为 true。 |
estimated_extra_bytes | 扩大搜索预计额外占用的内存。只覆盖搜索本身,是估算,不是上限。 |
除 retry_recommended 外,服务不知道的字段会省略。reason 不是 search_incomplete 时,请确认音频与文本对应同一段内容;扩大搜索范围帮不上忙。
对齐成功但没有对上任何词(例如文本里没有可对齐的内容)时,返回 HTTP 200 且不带 segments;segments 缺失时按空列表读取。
扩大搜索范围重试
multipart 字段 search_effort 决定搜索范围:
| 取值 | 行为 |
|---|---|
standard | 默认值,传空值也按它处理。按标准范围搜索一次。 |
extended | 按更大的范围搜索一次。更慢,更耗内存;失败那次的 estimated_extra_bytes 给出大约多用多少。 |
auto | 先跑 standard;仅当它失败且 retry_recommended: true 时,再跑一次 extended。其他失败原样返回。 |
其他取值返回 400 bad_request,param 为 "search_effort"。error.alignment.retry_recommended 为 true 时,用 search_effort=extended 重发同一请求:
请求示例:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@song.mp3 -F text="$(cat lyrics.txt)" \
-F search_effort=extended输出示例(响应节选):
{
"task": "align",
"duration": 214.6,
"text": "...",
"segments": [{ "id": 0, "start": 12.1, "end": 16.8, "text": "...", "words": [{ "word": "...", "start": 12.1, "end": 12.5, "score": 0.93 }] }],
"search_effort_used": "extended",
"usage": { "type": "duration", "seconds": 214.6 }
}服务报告了实际档位时,成功响应带 search_effort_used。使用 search_effort=auto 时,只有第二遍确实跑了才出现 search_retried: true,否则省略。extended 仍然失败时,retry_recommended 为 false。