High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: π AI Gateway (LiteLLM) π€ LLM Serving (vLLM, SGLang, Ollama) π Vector Databases, π Embedding Models (TEI) π Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
This repository contains hands-on projects, code examples, and deployment workflows. Explore multi-agent systems, LangChain, LangGraph, AutoGen, CrewAI, RAG, MCP, automation with n8n, and scalable agent deployment using Docker, AWS, and BentoML.
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.