Skip to main content
Some AI models use thinking, also called reasoning, before answering or calling tools. HQ configures this automatically from the model’s provider catalog. Choose your model as usual; there is no separate thinking token budget to configure in Desktop.

Automatic reasoning

HQ starts with thinking off and adjusts that setting to what the model supports:
  • Required reasoning: HQ uses the lowest supported thinking level. A model that cannot disable reasoning stays enabled.
  • Optional reasoning: Thinking stays off by default.
These rules apply to conversations from every client connected to the same HQ. Background tasks, such as generating chat titles and extracting memories, follow the same rules for the model they use.

Response time and cost

Reasoning can increase response time and consume output tokens, even when the provider does not return visible thinking text. The lowest supported level reduces reasoning effort but does not guarantee a fixed response time or token count. If you need lower latency or cost, choose a model that supports disabling reasoning or a faster model. Models that require reasoning cannot run with it turned off.

Provider plugins

If you add models through a provider plugin, its catalog must describe which reasoning settings the provider accepts. See the example plugin for the model metadata.