Glossary

Red Hat AI Inference Server

RHAIIS

Red Hat's model serving product, built on the open source vLLM engine. IBM lists it as one of two supported ways to drive a Spyre Accelerator on IBM Power.

Red Hat AI Inference Server accepts requests for a model, packs them onto the accelerator efficiently, manages the memory each conversation needs while it runs, and returns the answers. It is built on vLLM with optimisation work from Neural Magic, which Red Hat acquired. IBM's Spyre documentation names it, or Red Hat OpenShift AI, as a requirement for the accelerator on Power 11. How this layer is configured has a direct effect on how many concurrent users a single card supports, and therefore on how many cards a project needs.

Related Terms