Inference
Reasoning
Use reasoning_effort and read reasoning traces when models support them.
Updated
Models that advertise reasoning can expose the model’s thinking trace as
reasoning_content on the assistant message. LLMBase also includes a
compatibility alias named reasoning when returned by the selected model.
Use reasoning_effort for portable effort control when the selected model lists
that parameter in supported_parameters. As with OpenAI, reasoning_effort on
its own turns thinking on; you do not need an extra reasoning object or
extra_body flag. The reasoning: { "enabled": true } object
is accepted as well and uses the model’s lowest published effort.
Effort control
{
"model": "<model-id-from-/v1/models>",
"messages": [
{ "role": "user", "content": "Compare these two migration plans." }
],
"reasoning_effort": "high"
}
Effort values are model-dependent, so check metadata before requiring a value in
production. When present, supported_reasoning_efforts lists the model-native
values; use one of those values instead of carrying assumptions from another
model family.
For compatibility with OpenAI-compatible clients, the standard OpenAI values
minimal, low, medium, and high are accepted whenever the model publishes
a low, medium, or high tier at least as strong. If the selected model has no native tier with that
name, LLMBase uses the model’s weakest published tier that is at least as strong
(for example high on a model that publishes ["high", "max"]). Native tiers
are used unchanged, and max is only used when you request it. Values the model
cannot serve, and unknown values, return 400.
To turn thinking off, send reasoning_effort: "none".
Thinking and max_tokens
Reasoning tokens count against max_tokens (or max_completion_tokens), as
with OpenAI. When you set your own limit on a plain-text request, thinking
stays on and your limit caps the total output (reasoning plus answer):
{
"model": "<model-id-from-/v1/models>",
"messages": [{ "role": "user", "content": "What is 17 * 23?" }],
"reasoning_effort": "high",
"max_tokens": 2000
}
- If your limit is smaller than the requested effort usually needs, LLMBase may
use a lower published effort so an answer still fits. Usage reports the
reasoning tokens in
usage.completion_tokens_details.reasoning_tokens. - If thinking uses the whole limit, the response is a normal
200withfinish_reason: "length"and emptycontent. Raisemax_tokensor lowerreasoning_effortif that happens. - If you omit
max_tokens, LLMBase reserves room for both thinking and the answer. - Very small limits (below 16 tokens, for example a
max_tokens: 1health check) skip thinking, so health checks stay fast. - Requests with
tools, JSON output (response_formatthat the model applies), tool results, or image input skip thinking when your limit is below the room the effort needs, because a cut-off tool call or JSON document cannot be returned as a valid answer. Omitmax_tokensor raise it to keep thinking on for those requests. On a model without JSON support, aresponse_format: {"type": "json_object"}is ignored (see Structured outputs), so the request is treated as plain text.
Thinking flags
Some model families also expose thinking through chat template flags. The
portable flags below are supported through extra_body.chat_template_kwargs
when advertised by the selected model:
{
"model": "<model-id-from-/v1/models>",
"messages": [
{ "role": "user", "content": "Solve 17 * 23 and show the final answer." }
],
"extra_body": {
"chat_template_kwargs": {
"enable_thinking": true,
"thinking": true,
"preserve_thinking": true
}
}
}
Response shape
Typical non-streaming responses include the final answer in content and the
reasoning trace separately:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "17 * 23 = 391.",
"reasoning_content": "Compute 17 * 20 = 340 and 17 * 3 = 51, then add them.",
"reasoning": "Compute 17 * 20 = 340 and 17 * 3 = 51, then add them."
}
}
]
}
Reasoning support is model-dependent. If your request depends on reasoning,
choose a model whose metadata includes supported_features: ["reasoning"] and
supported_parameters containing reasoning_effort. If the model includes
supported_reasoning_efforts, prefer one of those values.