waybarrios/
vllm-mlx
waybarrios/vllm-mlxTools
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
1.6k
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: π AI Gateway (LiteLLM) π€ LLM Serving (vLLM, SGLang, Ollama) π Vector Databases, π Embedding Models (TEI) π Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes