Inference is the use of an already-trained model: given an input, it produces its output. Every chatbot answer is an inference. Its cost (latency, energy, price per token) is a central criterion when sizing an AI solution for production.
In practice at Gensai
Gensai optimises inference costs by choosing the right model size, down to the fully local inference of the Oceans installation.