vLLM Scale-to-Zero on GKE: KEDA Meets Datadog
Monitoring and scaling a vLLM stack on Google Kubernetes Engine with Datadog, using a custom proxy for synchronous scale-to-zero.
Thoughts on Data Engineering, Cloud Architecture, and AI Development.
Monitoring and scaling a vLLM stack on Google Kubernetes Engine with Datadog, using a custom proxy for synchronous scale-to-zero.
A high-level overview of the transition from Amazon DataZone to SageMaker Unified Studio, exploring the structural changes and migration impact.
Connecting a generative UI host to the real AWS MCP Server: CORS proxying, OAuth 2.1 sign-in, and live AWS data rendered as widgets.
A deep dive into reducing idle costs by scaling a vLLM GPU node to zero using KEDA and its HTTP add-on in GKE.
Deep dive into project creation with the Lakehouse Database blueprint across accounts and regions from a single SageMaker Unified Studio domain.
Exploring how AI chats generate dynamic user interfaces using MCP Apps and the BYOK model.
Deep dive into the DataLakehouse Blueprint differences between DataZone and SageMaker Unified Studio (SMUS) and how to reproduce the v1 layout.
A practical guide to deploying and serving open-weight models on Google Kubernetes Engine using vLLM for agentic coding workflows.
A deep dive into the architecture and lifecycle of projects in SageMaker Unified Studio.
A deep dive into the Model Context Protocol (MCP) and building WebMCP tools for AI agents.
A deep dive into AWS-managed and custom blueprints in SageMaker Unified Studio.
How to securely sync your Obsidian vault with Git and integrate AI features via Copilot for a powerful second brain.
Welcome to my blog about Data Engineering, Cloud Architecture, and AI Development.