Skip to content

Instantly share code, notes, and snippets.

View aneeshkp's full-sized avatar

Aneesh Puttur aneeshkp

View GitHub Profile
@aneeshkp
aneeshkp / deploy-llmisvc-local-models.sh
Last active May 19, 2026 20:46
Deploy LLMInferenceService with local NVMe model cache on AKS
#!/bin/bash
# Deploy LLMInferenceService using pre-downloaded models from local NVMe storage
# Models are in HuggingFace cache format at /mnt/local-nvme-storage/models/hub/
#
# The operator normally generates "vllm serve /mnt/models" which breaks with HF cache
# (symlinks in snapshots/ point to ../../blobs/ which is outside the mount boundary).
#
# Workaround: override the container command to use the model NAME instead of a path.
# Mount the HF cache and set HF_HUB_CACHE so vLLM resolves from cache — no download.
# Must include SSL flags since the operator's default command template is bypassed.
@aneeshkp
aneeshkp / deploy-llmisvc-hf-download.sh
Created May 19, 2026 20:00
Deploy LLMInferenceService with HuggingFace download (storage-initializer) on AKS
#!/bin/bash
# Deploy LLMInferenceService — download model from HuggingFace via storage-initializer
# No local models needed. The storage-initializer init container downloads the model
# before vLLM starts.
# 1. Create namespace
kubectl create namespace llm-test 2>/dev/null || true
# 2. Copy pull secret
kubectl create secret generic rhai-pull-secret \
@aneeshkp
aneeshkp / xks-https-workaround.md
Last active May 29, 2026 16:41
HTTPS on xKS Gateway — Manual Workaround for RHOAI 3.4 TP (pre-PR #90)

HTTPS on xKS Gateway — Manual Workaround for RHOAI 3.4 TP (pre-PR #90)

HTTPS on xKS Gateway — Manual Workaround (RHOAI 3.4 TP)

The rhai-on-xks helm chart creates the inference gateway with HTTP only (port 80). These steps add HTTPS manually after the chart is installed.

Verified working on AKS (May 28–29, 2026).

Which case should I use?