Inference

Reasoning

Use reasoning_effort and read reasoning traces when models support them.

Updated

Models that advertise reasoning can expose the model’s thinking trace as reasoning_content on the assistant message. LLMBase also includes a compatibility alias named reasoning when returned by the selected model.

Use reasoning_effort for portable effort control when the selected model lists that parameter in supported_parameters. As with OpenAI, reasoning_effort on its own turns thinking on; you do not need an extra reasoning object or extra_body flag. The reasoning: { "enabled": true } object is accepted as well and uses the model’s lowest published effort.

Effort control

{
  "model": "<model-id-from-/v1/models>",
  "messages": [
    { "role": "user", "content": "Compare these two migration plans." }
  ],
  "reasoning_effort": "high"
}

Effort values are model-dependent, so check metadata before requiring a value in production. When present, supported_reasoning_efforts lists the model-native values; use one of those values instead of carrying assumptions from another model family.

For compatibility with OpenAI-compatible clients, the standard OpenAI values minimal, low, medium, and high are accepted whenever the model publishes a low, medium, or high tier at least as strong. If the selected model has no native tier with that name, LLMBase uses the model’s weakest published tier that is at least as strong (for example high on a model that publishes ["high", "max"]). Native tiers are used unchanged, and max is only used when you request it. Values the model cannot serve, and unknown values, return 400.

To turn thinking off, send reasoning_effort: "none".

Thinking and max_tokens

Reasoning tokens count against max_tokens (or max_completion_tokens), as with OpenAI. When you set your own limit on a plain-text request, thinking stays on and your limit caps the total output (reasoning plus answer):

{
  "model": "<model-id-from-/v1/models>",
  "messages": [{ "role": "user", "content": "What is 17 * 23?" }],
  "reasoning_effort": "high",
  "max_tokens": 2000
}
  • If your limit is smaller than the requested effort usually needs, LLMBase may use a lower published effort so an answer still fits. Usage reports the reasoning tokens in usage.completion_tokens_details.reasoning_tokens.
  • If thinking uses the whole limit, the response is a normal 200 with finish_reason: "length" and empty content. Raise max_tokens or lower reasoning_effort if that happens.
  • If you omit max_tokens, LLMBase reserves room for both thinking and the answer.
  • Very small limits (below 16 tokens, for example a max_tokens: 1 health check) skip thinking, so health checks stay fast.
  • Requests with tools, JSON output (response_format that the model applies), tool results, or image input skip thinking when your limit is below the room the effort needs, because a cut-off tool call or JSON document cannot be returned as a valid answer. Omit max_tokens or raise it to keep thinking on for those requests. On a model without JSON support, a response_format: {"type": "json_object"} is ignored (see Structured outputs), so the request is treated as plain text.

Thinking flags

Some model families also expose thinking through chat template flags. The portable flags below are supported through extra_body.chat_template_kwargs when advertised by the selected model:

{
  "model": "<model-id-from-/v1/models>",
  "messages": [
    { "role": "user", "content": "Solve 17 * 23 and show the final answer." }
  ],
  "extra_body": {
    "chat_template_kwargs": {
      "enable_thinking": true,
      "thinking": true,
      "preserve_thinking": true
    }
  }
}

Response shape

Typical non-streaming responses include the final answer in content and the reasoning trace separately:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "17 * 23 = 391.",
        "reasoning_content": "Compute 17 * 20 = 340 and 17 * 3 = 51, then add them.",
        "reasoning": "Compute 17 * 20 = 340 and 17 * 3 = 51, then add them."
      }
    }
  ]
}

Reasoning support is model-dependent. If your request depends on reasoning, choose a model whose metadata includes supported_features: ["reasoning"] and supported_parameters containing reasoning_effort. If the model includes supported_reasoning_efforts, prefer one of those values.