◈ AI GLOSSARY ◈

Streaming

Delivering a model's answer token by token as it is generated, rather than waiting for the whole reply to finish.

WHY IT MATTERS

It is why chat assistants appear to type in real time, which feels faster and more responsive.

Frequently asked questions

Why does an AI chat look like it is typing in real time?

That is streaming, which delivers the answer token by token as it is generated instead of waiting for the whole reply to finish. It makes the assistant feel faster and more responsive even before the full answer is ready.

Does streaming make the AI actually answer faster?

It does not speed up the total work, but you start seeing words sooner, so the wait feels shorter. It mainly improves the experience and perceived latency rather than the true processing time.

Is streaming always the best choice?

For chat where people watch the reply appear, yes, it feels great. But when the output is being fed straight into other software that needs the complete, finished result, waiting for the whole response can be simpler to handle.

New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.

← Back to the glossary