◈ AI GLOSSARY ◈

Latency

The delay between sending a request to a model and getting a response back.

WHY IT MATTERS

It shapes how an app feels. Reasoning models are more accurate but slower, a real trade-off to weigh.

Frequently asked questions

Why does one AI chatbot feel instant while another makes me wait?

That waiting time is latency, the gap between hitting send and seeing a response. A model that thinks harder before answering, like a reasoning model, will usually feel slower, so the delay often reflects the kind of model doing the work.

Does low latency mean the AI gives better answers?

No, speed and quality are separate things. A fast reply can still be wrong, and the most accurate models are sometimes the slowest, so low latency just means it feels snappy, not that it is smarter.

Why do chat assistants seem to type out their answer word by word?

That effect is called streaming, and it is a trick to hide latency. Instead of making you wait for the whole answer, the tool shows each piece as it is produced, so it feels faster even when the total time is the same.

New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.

← Back to the glossary