Parameters & Settings
Latency
Definition
The time delay between sending a request to an AI model and receiving the first token of the response. Lower latency means faster-feeling interactions. Latency depends on model size, server load, and network distance.
In Plain English
How long you wait before the AI starts answering. When you press Send and there's a pause before text appears — that's latency. Smaller, faster models have lower latency. Bigger, smarter models often have higher latency.