What is Inference?
In one line
The moment an AI model is actually used to produce an answer, as opposed to being trained.
Speed and cost announcements, like “300 tokens per second”, usually refer to inference.
The moment an AI model is actually used to produce an answer, as opposed to being trained.
Speed and cost announcements, like “300 tokens per second”, usually refer to inference.