vLLM Scale-to-Zero on GKE: KEDA Meets Datadog
Monitoring and scaling a vLLM stack on Google Kubernetes Engine with Datadog, using a custom proxy for synchronous scale-to-zero.
Senior Tech Lead — DataOps, DevOps & AI. Writing about Data, Cloud, and AI development.
Monitoring and scaling a vLLM stack on Google Kubernetes Engine with Datadog, using a custom proxy for synchronous scale-to-zero.
A high-level overview of the transition from Amazon DataZone to SageMaker Unified Studio, exploring the structural changes and migration impact.
Connecting a generative UI host to the real AWS MCP Server: CORS proxying, OAuth 2.1 sign-in, and live AWS data rendered as widgets.
A deep dive into reducing idle costs by scaling a vLLM GPU node to zero using KEDA and its HTTP add-on in GKE.
Deep dive into project creation with the Lakehouse Database blueprint across accounts and regions from a single SageMaker Unified Studio domain.