What's new
Parameters & Settings

Latency

Definition

The time delay between sending a request to an AI model and receiving the first token of the response. Lower latency means faster-feeling interactions. Latency depends on model size, server load, and network distance.

In Plain English

How long you wait before the AI starts answering. When you press Send and there's a pause before text appears — that's latency. Smaller, faster models have lower latency. Bigger, smarter models often have higher latency.