Red Hat OpenShift AI
The other product IBM accepts for driving a Spyre card. It is the heavier option: a full container platform for teams running several models, several projects, and more than one server.
IBM gives you two approved ways to drive a Spyre card. This is the bigger one.
The difference is roughly the difference between a good oven and a commercial kitchen. Red Hat AI Inference Server serves models. Red Hat OpenShift AI serves models and also manages who owns which model, how new versions get promoted, how several teams share the same hardware without stepping on each other, and how the whole thing scales onto more machines.
How to tell which one you need
Count your models
One model serving one application points at Red Hat AI Inference Server. Several models, or several versions of the same model in flight, points at OpenShift AI.
Count your teams
If one person owns the whole thing, you do not need project separation. If three departments want their own space, you do.
Count your servers
A single Power system with a card or two does not need container orchestration. A growing estate does, and retrofitting it later is the expensive way round.
Most AS/400 shops starting their first AI project answer one, one, and one. That answer is not embarrassing and it does not need a platform. Start with the lighter option and move up when the counts change.
OpenShift is a substantial commitment in its own right, with its own skills requirement. Do not adopt it purely to satisfy a Spyre prerequisite when Red Hat AI Inference Server also satisfies it.
Related
Red Hat AI Inference Server
One of the two products IBM names as required before a Spyre card will do anything. It is the driver, the traffic controller, and the reason a 128 GB card can serve more than one person at a time.
IBM Spyre Accelerator for Power
A 75 watt PCIe card that adds more than 300 TOPS of AI inference to a Power 11 server. It is the only IBM AI chip an AS/400 shop can actually order as an add-on, and it comes with a short list of things it refuses to work without.
vLLM
The open source engine inside the product IBM requires for a Spyre card. Worth knowing by name, because it is the reason the same card can answer many people at once instead of queuing them.
IBM Granite Models
IBM's own family of open language models, and the cargo an accelerator card is built to carry. You can download them free, run them on your own hardware, and nobody meters the tokens.