转录与说话人

转录音频文件

POST /v1/audio/transcriptions

将 hello.wav 替换为自己的音频或视频文件。multipart 的 file 字段必填;省略 model 时使用服务配置的默认模型。仅需纯文本时使用 response_format=json,需要时间戳或说话人信息时请参考下一示例。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F response_format=json

输出示例:

JSON
{
  "text": "Hello world.",
  "start": 0.0,
  "duration": 2.5,
  "language": "en",
  "usage": {"type": "duration", "seconds": 2.5}
}

response_format 设为 text 返回纯文本,verbose_json 返回结构化文稿,diarized_json 会附带说话人区分。各字段定义见文稿字段定义,完整请求参数见可选参数。

参数 language 可省略,省略时从音频自动检测语言。它会按支持的转录语言校验,代码写法见统一的语言代码写法。传入不支持的语言代码会返回 HTTP 400 asr_language_unsupported,并在 param=language 中定位错误字段。响应中的 message 会回显你传入的取值,仅供调试诊断。客户端应按 error.code 处理业务分支,不要解析错误文案。

一次获取文字、时间戳与说话人

需要同时获取文字、时间戳并区分「谁说了什么」时,指定 response_format=diarized_json。返回的说话人标签代表当前录音内的声纹聚类,不对应真实个人身份。以下时间与标签仅为示意。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F response_format=diarized_json \
  -F "timestamp_granularities[]=word"

输出示例:

JSON
{
  "text": "Hello world.",
  "segments": [
    {
      "type": "transcript.text.segment",
      "id": "seg_0",
      "start": 0.2,
      "end": 1.4,
      "text": "Hello world.",
      "speaker": "SPEAKER_00",
      "words": [
        {
          "word": "Hello",
          "start": 0.2,
          "end": 0.6
        },
        {
          "word": "world.",
          "start": 0.7,
          "end": 1.4
        }
      ]
    }
  ],
  "usage": {
    "type": "duration",
    "seconds": 2.5
  }
}

CLI 的 transcribe --diarize 命令底层调用此接口,参数映射与默认配置差异见 API 复用。参数 candidates(取值 2–8)仅适用于非流式 verbose_json,且不能与逐词时间戳同时使用。

读取流式转录事件

该模式仅桌面网关支持(Headless 服务目前对此请求返回 HTTP 400)。调用时设置 stream=true 与 response_format=json,且不能指定时间戳粒度。使用 curl -N 可关闭客户端缓冲。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
  -F model=EdgeSpeak/Lattice-2 \
  -N -F file=@hello.wav -F response_format=json -F stream=true

输出示例:

event: transcript.text.delta
data: {"type":"transcript.text.delta","delta":"Hello world."}

event: transcript.text.done
data: {"type":"transcript.text.done","text":"Hello world.","usage":{"type":"duration","seconds":2.5}}

客户端将 delta 文本增量追加到界面显示,并以 done 事件中的完整 text 为最终结果。同时应处理 error 事件及连接异常中断的情况。

获取说话人活动时间轴

POST /v1/speaker/diarizations

仅需获取「谁在何时说话」的时间区间而不需要文字内容时,调用此接口。上传音频文件 file;可选参数 num_speakers(1–32)用于指导说话人聚类数量。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/diarizations" \
  -F file=@meeting.wav -F num_speakers=2

输出示例:

JSON
{
  "duration": 10.0,
  "speakers": [
    "SPEAKER_00",
    "SPEAKER_01"
  ],
  "segments": [
    {
      "start": 0.4,
      "end": 3.2,
      "speaker": "SPEAKER_00"
    },
    {
      "start": 4.0,
      "end": 8.1,
      "speaker": "SPEAKER_01"
    }
  ]
}

响应中不包含转录文字,num_speakers 仅作为聚类参考,返回的说话人标签不对应真实个人身份。

提取说话人声纹向量

POST /v1/speaker/embeddings

支持通过多个 file 字段批量上传音频样本。每段样本不超过 30 秒,且至少包含 4 秒有效的单一说话人语音。请保存完整响应供后续相似度计算使用;示例中使用 jq 仅提取并打印向量长度,避免输出全部 256 个数值。运行示例需先安装 jq。

请求示例:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/embeddings" \
  -F model=EdgeSpeak/Lattice-1 \
  -F file=@speaker-a.wav -F file=@speaker-b.wav \
  -o embeddings.json
jq '.data | map({index, dimensions: (.embedding | length), effective_speech_seconds})' embeddings.json

输出示例:

JSON
[
  {
    "index": 0,
    "dimensions": 256,
    "effective_speech_seconds": 5.2
  },
  {
    "index": 1,
    "dimensions": 256,
    "effective_speech_seconds": 6.1
  }
]

上方输出为 jq 处理后的结果,原始 HTTP 响应结构为 {object:"list",model,data:[{index,embedding,effective_speech_seconds}]}。可选参数 model 默认为 EdgeSpeak/Lattice-1。请保留保存的 embeddings.json,供下一步比对使用。

比较两个声纹向量

POST /v1/speaker/similarity

承接上一步,使用实际返回的向量构造请求体。参与比对的两个向量必须均为 256 维有限数值,且模长不能为零。

请求示例:

Shell
jq '{model, embedding_a: .data[0].embedding, embedding_b: .data[1].embedding}' \
  embeddings.json > similarity-request.json
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/similarity" \
  -H "Content-Type: application/json" -d @similarity-request.json

输出示例:

JSON
{
  "object": "speaker_similarity",
  "model": "EdgeSpeak/Lattice-1",
  "similarity": 0.82
}

输出中的相似度分数为示意值,实际余弦相似度取值在 -1 到 1 之间。接口不接受 threshold 或 match 参数,由应用决定如何使用分数。

继续阅读:强制对齐 · API 索引