Models and service

Check service health

GET /health

GET /v1/health

Both endpoints require no API key. A running service can return HTTP 200 with an inactive license; check both fields. A successful response does not mean a model is loaded.

Request:

Shell
curl --fail-with-body "${EDGESPEAK_BASE_URL%/v1}/health"
curl --fail-with-body "$EDGESPEAK_BASE_URL/health"

Example output:

JSON
{
  "status": "ok",
  "license": "active"
}

license is active or inactive. Use edgespeak-cli status on the service host for activation details.

Discover models and capabilities

GET /v1/models

The model catalog lists models known to this service, including entries that are not downloaded or loaded. Filter supported_endpoints for the API you want, read default_for for defaults, and check execution_location before sending requests. The response below is a single-entry excerpt; your catalog contents will differ. Alignment models list /v1/audio/alignments under supported_endpoints; for which one to choose, see alignment requests.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models"

Example output:

JSON
{
  "object": "list",
  "data": [
    {
      "id": "Qwen/Qwen3.5-4B",
      "object": "model",
      "created": 0,
      "owned_by": "EdgeSpeak",
      "supported_endpoints": [
        "/v1/chat/completions",
        "/v1/responses",
        "/v1/messages",
        "/v1/tokenize"
      ],
      "features": [
        "reasoning",
        "vision",
        "tool_calling"
      ],
      "execution_location": "local",
      "default_for": []
    }
  ]
}

When calling subsequent endpoints, use an actual ID from the model catalog. local models execute on the service host; configured remote models receive the request. Built-in EdgeSpeak/Skylark is already available and has no download/load/unload lifecycle. Speech and language models can expose different lifecycle behavior; the following sequence is for a supported local language model.

Start a model download

POST /v1/models/download

Downloads consume network bandwidth and disk space on the service host. This endpoint starts a background job and returns 202 Accepted, not a ready model. All four POST lifecycle endpoints accept exactly {"model":"catalog ID"}; passing extra fields is rejected.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/download" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B"}'

Example output:

JSON
{
  "model": "Qwen/Qwen3.5-4B",
  "state": "downloading",
  "bytes_downloaded": 0,
  "total_bytes": 0,
  "error": null
}

total_bytes: 0 can mean that the total is not known yet. Before loading the model, poll the query endpoint for progress.

Check download progress

GET /v1/models/downloads

Poll once per second and match the task by model; returning an empty list does not mean the model is installed.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/downloads"

Example output:

JSON
{
  "object": "list",
  "data": [
    {
      "model": "Qwen/Qwen3.5-4B",
      "state": "completed",
      "bytes_downloaded": 2500000000,
      "total_bytes": 2500000000,
      "error": null
    }
  ]
}

States are downloading, completed, failed, and cancelled. Byte counts here are illustrative. If the download fails, inspect the error field; retrying requires a new POST request.

Cancel an active download

POST /v1/models/download/cancel

This endpoint applies only to active jobs; cancellation is also asynchronous.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/download/cancel" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B"}'

Example output:

JSON
{
  "model": "Qwen/Qwen3.5-4B",
  "state": "cancelling"
}

Poll /models/downloads until the job reaches a terminal state.

Load a local model

POST /v1/models/load

After download completion, load the model before language generation or Realtime. This allocates memory on the service host. The response shows the stable core fields; the service host may also return extended fields.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/load" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B"}'

Example output:

JSON
{
  "success": true,
  "model": "Qwen/Qwen3.5-4B",
  "status": "loaded"
}

Inspect running models

GET /v1/models/running

Catalog presence does not mean a model is ready. Inspect runtime state here. The example shows the state after a successful load; the status content for each entry is model-specific.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/running"

Example output:

JSON
{
  "object": "list",
  "data": [{"id":"Qwen/Qwen3.5-4B","status":{"value":"loaded"}}],
  "runtime": {
    "state": "running"
  }
}

If data is empty and the runtime is not loading, load the model before Realtime generation.

Request queue status

The same authenticated endpoint reports runtime.queues on the Headless service; the desktop gateway does not include this field. This is an excerpt extracted with jq:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/running" \
  -H "Authorization: Bearer $EDGESPEAK_API_KEY" | jq '.runtime.queues.native_broadcast'
JSON
{"active":1,"waiting":3,"execution_capacity":1,"oldest_wait_ms":1200}

During cold loading or model switching, runtime.state can be loading. The model list is temporarily empty, but queues are still returned; do not treat a temporarily empty list as completed unloading.

Each queue includes active, waiting, execution_capacity, and oldest_wait_ms. http_transcribe, http_alignment, http_segmentation, http_normalization, http_diarization, http_speaker_embedding, http_broadcast, and http_realtime are admission scheduling queues; native_transcribe / native_broadcast represent native model occupancy across REST, Realtime, MCP, and preloading. active includes admitted tasks still preparing or cleaning up, which is not equivalent to concurrent GPU inference. The same task can appear in both queue layers at the same time; do not sum queue values as total requests. Waiting before the global upload slot is not included in these counts.

Cancelling a waiting operation removes its entry; started blocking tasks can continue until completion or native cancellation cleanup. Realtime sessions hold occupancy until closed, so switching transcription models can wait for a long time. For configuration details, see concurrency and sleep options.

Unload a model

POST /v1/models/unload

Release a model from memory when it is no longer in use. This does not delete its downloaded files. Active sessions can prevent unloading; end the sessions and retry. The example response shows core fields.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models/unload" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B"}'

Example output:

JSON
{
  "success": true,
  "model": "Qwen/Qwen3.5-4B",
  "status": "unloaded"
}

Continue: Text processing · API index