Skip to content

LiteLLM compatibility contract ​

This document defines the Cloud-facing contract that Go Feather Route must match before it can be considered as a future LiteLLM replacement. Thingd Cloud continues to use LiteLLM while this contract is implemented and validated.

Thingd Cloud remains the authority for users, projects, operation labels, reservations, quotas, billing attribution, usage rollups, caching, and Thingd tool execution. Go Feather Route is only a provider gateway.

Required operations ​

OperationWire endpointRequired behaviorCurrent Go status
Chat completionPOST /v1/chat/completionsOpenAI-compatible request/response, model routing, usage passthroughImplemented; contract tests required
Structured completionPOST /v1/chat/completionsresponse_format.type=json_object, bounded max_tokensImplemented as request passthrough; fixture coverage required
Agent streamingPOST /v1/chat/completionsSSE forwarding, cancellation, [DONE], optional final usageImplemented; Cloud-shaped stream coverage required
NLQPOST /v1/chat/completionsStructured JSON response and normal error semanticsUses chat contract
ClassificationPOST /v1/chat/completionsStructured JSON response and bounded retriesUses chat contract
SummarizationPOST /v1/chat/completionsStructured JSON response and cancellationUses chat contract
EmbeddingsPOST /v1/embeddingsSingle/batch input, ordered vectors, dimensions, usageImplemented; workload integration required

Chat request contract ​

Cloud may send:

  • model
  • messages
  • max_tokens
  • response_format: {"type":"json_object"}
  • stream

The gateway must preserve the selected model, provider usage, request ID, and OpenAI-compatible status/error behavior. Cloud application logic performs any bounded Thingd capability execution from structured model output; native provider tool calling is not a prerequisite for this contract.

Streaming contract ​

The gateway must:

  • Forward SSE chunks promptly without buffering the complete response.
  • Preserve JSON chunk structure and choices[0].delta.content.
  • Preserve optional final usage metadata.
  • Preserve [DONE].
  • Cancel upstream work when the Cloud client disconnects.
  • Apply request, idle, and concurrency limits.
  • Surface provider failures before or during a stream.

Cloud accumulates the streamed content and records usage once after successful completion. The gateway must not create Cloud usage events.

Embedding contract ​

Cloud sends model and input, including batches of approved text values. A compatible response must provide one vector per input, preserve indexes, use a consistent dimension, and include provider usage when available. Cloud checks ordering, dimensions, finite numeric values, and batch limits before storing vectors.

For local Ollama, Go Feather exposes stable ollama-qwen3 and ollama-nomic-embed aliases and translates them to the configured upstream model names. It also translates Cloud's reasoning_effort: "none" to Ollama's think: false for Qwen3-style thinking models. This is standalone provider support, not Cloud replacement qualification.

Operational contract ​

The gateway must provide:

  • GET /health/liveliness
  • GET /ready
  • GET /status
  • GET /metrics
  • GET /v1/models

Authentication, provider credentials, timeouts, bounded retries, body limits, concurrency limits, request IDs, readiness, and resource diagnostics belong to the gateway boundary. Tenant identity, reservations, quotas, usage settlement, and reporting remain outside it.

Replacement gate ​

This contract is necessary but not sufficient for replacement. A future Cloud integration requires passing the project-memory, embedding-worker, accounting, cancellation, reliability, and benchmark gates described in the project plan. Until then, LiteLLM remains the only active Cloud gateway.