Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, Slurm, 20+ clouds, on-prem).
-
Updated
Jul 21, 2026 - Python
Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, Slurm, 20+ clouds, on-prem).
Build, Manage and Deploy AI/ML Systems
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
FinOps and cloud cost optimization tool. Supports AWS, Azure, GCP, Alibaba Cloud and Kubernetes.
Automatic LLM router — 82% cost savings, 79.4% accuracy, 93.4% pass rate. Drop-in OpenAI proxy.
Open-source LLM router & AI cost optimizer. Routes simple prompts to cheap/local models, complex ones to premium — automatically. Drop-in OpenAI-compatible proxy for Claude Code, Codex, Cursor, OpenClaw. Saves 40-70% on AI API costs. Self-hosted, no middleman.
A tool for customers to evaluate their AWS service configurations based on AWS and community best practices and receive recommendations on potential improvements.
Programmatically delete AWS resources based on an allowlist and time to live (TTL) settings
System of Record for Kubernetes cost accounting: per-namespace CPU, memory and GPU usage, with the 30% non-allocatable overhead made visible. Connects to AI assistants via MCP (Claude, Gemini, Mistral, Cursor) for plain-language analysis. Formerly Kube-Opex-Analytics.
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.
An open source solution for cost management of cloud and hybrid cloud environments.
Scan 30+ AWS services. Find cost waste. Detect security gaps. Map your attack surface. One command.
A tool to convert AWS EC2 instances back and forth between On-Demand and Spot billing models.
Free, local, open-source LLM cost analyzer - see where your LLM bill leaks, on your machine.
Intelligent multi-LLM router with task-aware routing strategies, cost optimization, and production safety controls — drop-in OpenAI-compatible API
Generate Sheets (Google, Excel and CSV) with useful information about your AWS spendings.
This solution analyzes all of your Amazon WorkSpaces usage data and automatically converts the WorkSpace to the most cost-effective billing option (hourly or monthly), depending on your individual usage. Use this with a single account, or with AWS Organizations across multiple accounts, to help you monitor your WorkSpace usage and optimize costs.
⚡ Unit tests for AI. Test prompts, compare models, save money.
The first Task-Aware MCP server and automated VRAM calculator for LLM fine-tuning. Instantly snipe the cheapest, fastest GPUs across 10+ cloud providers.
AI-powered text compression library for RAG systems and API calls. Reduce token usage by up to 50-60% while preserving semantic meaning with advanced compression strategies.
Add a description, image, and links to the cost-optimization topic page so that developers can more easily learn about it.
To associate your repository with the cost-optimization topic, visit your repo's landing page and select "manage topics."