Red Hat AI Inference 3.5
What's New
Get started
Plan
Inference Operations
Deploy the standalone Red Hat AI Inference container in OpenShift Container Platform
Deploy the standalone Red Hat AI Inference container in OpenShift Container Platform clusters that have supported AI accelerators installed
Deploy the standalone Red Hat AI Inference container in a disconnected environment
Deploy Red Hat AI Inference in a disconnected environment using OpenShift Container Platform and a disconnected mirror image registry
Inference serving language models in OCI-compliant model containers
Inferencing OCI-compliant models in Red Hat AI Inference
Speculative decoding
Speculative decoding with Red Hat AI Inference
Inference serving Mistral 3 models
Inference serving Mistral 3 models with Red Hat AI Inference
Inference serving geospatial foundation models
Inference serving geospatial foundation models with Red Hat AI Inference
Red Hat AI Model Optimization Toolkit
Compressing large language models with the LLM Compressor library
vLLM server arguments
Server arguments for running Red Hat AI Inference
Extending Red Hat AI Inference with tool calling capabilities
Configuring tool calling and chat templates for AI Inference
Distributed Inference Operations
Deploy Distributed Inference with llm-d on OpenShift Container Platform
Deploy and serve large language models at scale on OpenShift Container Platform
Deploy Distributed Inference with llm-d on AWS, Azure, or CoreWeave Kubernetes Service
Deploy Distributed Inference with llm-d on Azure or CoreWeave Kubernetes Service
Monitor and troubleshoot Distributed Inference with llm-d deployments
Monitor and troubleshoot Distributed Inference with llm-d deployments
Batch process inference requests with Distributed Inference with llm-d
Batch process inference requests with Distributed Inference with llm-d
Deploy large MoE models across multiple GPUs with WideEP
Deploy Distributed Inference with llm-d on Azure or CoreWeave Kubernetes Service
Deploy Models as a Service on non-OpenShift Kubernetes
Deploy and configure Models as a Service on non-OpenShift Kubernetes environments
Manage mixed workloads by using priority queuing
Manage mixed workloads by using priority queuing
Tool calling for Distributed Inference with llm-d deployments
Tool calling for Distributed Inference with llm-d deployments
Upgrade Distributed Inference with llm-d on managed Kubernetes
Upgrade the Distributed Inference with llm-d infrastructure stack on managed Kubernetes clusters
Additional Resources
Red Hat AI Inference Server 3.3
Switch to the Red Hat AI Inference Server 3.3 documentation
This content is not included.Red Hat AI Foundations
Explore no-cost courses to boost your AI knowledge and get hands-on experience with Red Hat AI products while earning a certificate
This content is not included.Red Hat AI learning hub
Explore the curated set of third‑party models validated for Red Hat AI products, ready for fast, reliable deployment