Text processing

Segment sentences

POST /v1/text/segmentations

Requests support plain text or timed segments[]. Desktop also accepts the file parameter, a server-local Transcript JSON path; Headless services require inline content. The threshold defaults to 0.35. The shown segmentation boundaries only illustrate the model output format.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/segmentations" \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello world. Let us begin.","threshold":0.35}'

Example output:

JSON
{
  "task": "segment",
  "text": "Hello world. Let us begin.",
  "segments": [
    {
      "text": "Hello world."
    },
    {
      "text": "Let us begin."
    }
  ]
}

Plain text inputs do not include timestamps or speakers after segmentation; timed inputs can preserve these fields. All timing values and margins are in seconds; see length and margin settings.

Segment paragraphs with the same API

Pass paragraph_threshold to enable native model paragraph boundaries. Each output item is a paragraph, potentially containing several sentences; the top-level text joins paragraphs with \n\n (a blank line). If the input text is short, it may return only one paragraph.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/segmentations" \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello world. Let us begin. Now a different topic: tomorrow’s trip.","threshold":0.35,"paragraph_threshold":0.5}'

Example output:

JSON
{
  "task": "segment",
  "text": "Hello world. Let us begin.\n\nNow a different topic: tomorrow’s trip.",
  "segments": [
    {
      "text": "Hello world. Let us begin."
    },
    {
      "text": "Now a different topic: tomorrow’s trip."
    }
  ]
}

For the threshold ranges, the paragraph-mode input restrictions and the CLI equivalent, see sentence and paragraph segmentation. The illustrated boundary is not a guaranteed split for this text. Omit paragraph_threshold for sentence mode. Corresponding CLI command: edgespeak-cli segment --text "Hello world. Let us begin." --paragraph-threshold 0.5.

Normalize text: TN and ITN

POST /v1/text/normalizations

ParameterDescription
textRequired, non-empty, up to 256 KiB UTF-8.
languageRequired: a language code, or und to choose readings segment by segment from the text. See supported languages.
modelDefaults to built-in EdgeSpeak/Skylark.
modetn (default): written → spoken; itn: spoken → written.
top_kDefault 8; range 1–16.
classesOptional non-empty array, e.g. cardinal, date, money.

The examples below set top_k: 1 to request a single candidate.

This workshop announcement contains two expressions to normalize: the fee $25 and the time 09:30. Surrounding prose stays intact. The complete TN and ITN responses below were checked against the running local gateway.

TN: written → spoken

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/normalizations" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "EdgeSpeak/Skylark",
  "text": "Registration is open for the workshop. The fee is $25 and the session starts at 09:30.",
  "language": "en",
  "mode": "tn",
  "top_k": 1
}'

Example output (complete response):

JSON
{
  "task": "normalize",
  "model": "EdgeSpeak/Skylark",
  "language": "en",
  "mode": "tn",
  "text": "Registration is open for the workshop. The fee is $25 and the session starts at 09:30.",
  "alternatives": [
    {
      "rank": 0,
      "text": "Registration is open for the workshop. The fee is twenty five dollars and the session starts at nine thirty.",
      "spans": [
        {
          "start_byte": 50,
          "end_byte": 53,
          "class": "money",
          "source": "$25",
          "output": "twenty five dollars"
        },
        {
          "start_byte": 80,
          "end_byte": 85,
          "class": "time",
          "source": "09:30",
          "output": "nine thirty"
        }
      ]
    }
  ]
}

Top-level text preserves the whole input; alternatives[0].text is the converted paragraph. The two spans record details for the money and time replacements. Offsets use half-open UTF-8 byte ranges in the original input, not character counts. For example, [50, 53) selects $25.

ITN: spoken → written

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/normalizations" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "EdgeSpeak/Skylark",
  "text": "Registration is open for the workshop. The fee is twenty five dollars and the session starts at nine thirty.",
  "language": "en",
  "mode": "itn",
  "top_k": 1
}'

Example output (complete response):

JSON
{
  "task": "normalize",
  "model": "EdgeSpeak/Skylark",
  "language": "en",
  "mode": "itn",
  "text": "Registration is open for the workshop. The fee is twenty five dollars and the session starts at nine thirty.",
  "alternatives": [
    {
      "rank": 0,
      "text": "Registration is open for the workshop. The fee is $25 and the session starts at 09:30.",
      "spans": [
        {
          "start_byte": 50,
          "end_byte": 69,
          "class": "money",
          "source": "twenty five dollars",
          "output": "$25"
        },
        {
          "start_byte": 96,
          "end_byte": 107,
          "class": "time",
          "source": "nine thirty",
          "output": "09:30"
        }
      ]
    }
  ]
}

ITN normalizes spoken text back to written form; it is not guaranteed to restore the original spelling from the TN input. Candidate results vary by language, category, and rules. Supported languages and categories are in Skylark’s x_edgespeak.normalization response metadata; for limitations, see all constraints.

Continue: Transcription and speakers · API index