Deploying AI Models on Kubernetes
A practical field guide to GPU node pools, model serving with vLLM and Triton, and the dark art of autoscaling inference workloads on GKE and EKS — without setting your cloud bill on fire. (Day 22)
A practical field guide to GPU node pools, model serving with vLLM and Triton, and the dark art of autoscaling inference workloads on GKE and EKS — without setting your cloud bill on fire. (Day 22)
Self-hosting LLMs Local Machines and on EC2/GKE with Ollama — when open-source beats API services (Day 15)
What vectors are, why semantic search matters, and how Pinecone/Weaviate/pgvector fit in (Day 8)
Why AI lies confidently and how to build guardrails as an infrastructure problem (Day 7)
How to think about context limits, pricing models, and request optimization (Day 6)
Comparing model APIs like you compare managed services (Day 5)
Example: Kubernetes, Terraform, Docker, AWS, MLOps...