Skip to content

Repository files navigation

NexusAI-Client - Unified Multi-Provider AI Gateway

NexusAI-Client ⚑

An ultra-lightweight, strictly-typed, asynchronous Python 3.12+ gateway for multi-provider AI APIs.
Unify Cerebras, Cohere, DeepSeek, Google Gemini (Free & Pro), Groq, Mistral, Nvidia NIM, OpenRouter, and OrcaRouter behind a single, elegant interface with zero heavy SDK dependencies.

PyPI version Documentation Website Python 3.12+ uv httpx Typing Tests License MIT

🌐 Interactive Documentation Website: https://nexus-ai-client-doc.vercel.app/


πŸ’‘ Why NexusAI-Client?

Integrating multiple AI providers in modern Python applications usually requires installing 9 or 10 separate proprietary SDKs (google-genai, openai, groq, cohere, mistralai, etc.). This creates dozens of transitive dependencies, version conflicts, memory overhead, and fragmented codebases.

NexusAI-Client solves this at the core:

  • πŸ”„ Dynamic Model Management & Auto-Rotation: Automatic failover on HTTP 404/400 (deprecated models), HTTP 429 rate limits, and timeouts across models and providers.
  • πŸͺΆ Zero Heavyweight Dependencies: Powered purely by httpx and python-dotenv.
  • ⚑ Native Asynchronous & SSE Streaming: Stream responses token-by-token in real time via stream_text() and stream_chat().
  • πŸ”„ Zero-Cost-First Smart Fallback: Automatic progression from 100% Free Tiers (Gemini, Groq, Cerebras, Cohere, Nvidia, OpenRouter, OrcaRouter, Mistral) to Paid Backups (AIGateway.auto_fallback()).
  • πŸ› οΈ Universal Tool Calling / Function Calling: Seamless tool definitions, structured function arguments parsing, and multi-turn agent loops across Groq, Cerebras, Mistral, DeepSeek, Gemini REST, Cohere V2, and Nvidia NIM.
  • πŸš€ World-Record Hardware Accelerators: Native support for Groq LPUs and Cerebras CS-3 Wafer-Scale engines (2,000+ tokens/sec).
  • 🧠 Enterprise Reasoning & Search Models: Native Cohere Command R+, DeepSeek R1, and Qwen 3.8 models.
  • 🎯 Guaranteed JSON Outputs: Native json_mode=True across all supported providers.
  • πŸ’° Live Account & Budget Inspection: Inspect real-time balances (USD, NGC credits) and rate limits (RPM, TPM, RPD).
  • πŸ” 670+ Models Discovered Live: Automatic detection of free-tier models (:free, -free) and accurate per-million-token pricing.

🌟 Spotlight: Zero-Cost-First Smart Fallback Routing

Why pay for AI calls when you can leverage high-throughput free tiers first, with seamless automatic fallback to paid commercial models?

NexusAI-Client automatically prioritizes zero-cost models before touching your wallet:

  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                                               100% FREE ZERO-COST TIERS                                                β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 1. Gemini    β”‚ 2. Groq LPU  β”‚ 3. Cerebras CS-3 β”‚ 4. Nvidia    β”‚ 5. OpenRouterβ”‚ 6. OrcaRouterβ”‚ 7. Cohere    β”‚ 8. Mistralβ”‚
  β”‚ (1M Context) β”‚ (Ultra-Fast) β”‚ (2000+ tok/s)    β”‚ (1k Credits) β”‚ (Free Hub)   β”‚ (Qwen/DeepS) β”‚ (Command R+) β”‚ (Dev Free)β”‚
  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
         β”‚              β”‚                β”‚                β”‚              β”‚              β”‚              β”‚             β”‚
         β–Ό (If Rate-Limited / 429 Quota Exceeded / Network Outage / Timeout) ────────────────────────────────────────▼
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                                           ULTRA-LOW-COST PAID BACKUP TIERS                                             β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 9. DeepSeek ($0.27 / 1M tokens)                            β”‚ 10. Gemini Pro (Enterprise GCP)                           β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1-Line Zero-Cost Failover in Your Code:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Automatically discovers active keys in .env and routes: Free -> Free -> Paid
    async with AIGateway.auto_fallback() as client:
        response = await client.generate_text("Explain quantum computing in 2 sentences.")
        print(f"βœ… Served by [{response.provider}] with zero downtime:")
        print(response.text)

if __name__ == "__main__":
    asyncio.run(main())

🎯 Supported Providers Matrix

Provider Identifier (provider) Tier Protocol Default Model Live Budget & Quota Detection
Cerebras "cerebras" (or "cerebras_free") Free (CS-3) OpenAI Chat API gpt-oss-120b Quotas: 30 RPM | 60k TPM | 1M tok/day
Cohere "cohere" (or "cohere_free") Free Trial Cohere V2 REST command-r-plus-08-2024 Quotas: 20 RPM | 1,000 calls/month
DeepSeek "deepseek" Paid OpenAI Chat API deepseek-chat Real-time USD Balance (GET /user/balance)
Gemini Free "gemini_free" Free (AI Studio) Gemini REST gemini-3.5-flash-lite Auto-rotation 429 | 15 RPM | 500 RPD (Lite) / 20 RPD (Flash)
Gemini Pro "gemini_pro" Paid Gemini REST gemini-3.1-pro-preview Google Cloud Pay-as-you-go Billing
Groq "groq" (or "groq_free") Free (LPU) OpenAI Chat API openai/gpt-oss-120b Quotas: 30 RPM | 14,400 RPD | 30k TPM
Mistral AI "mistral" Free / Platform OpenAI Chat API mistral-small-latest Free Dev Models (codestral-latest, etc.)
Nvidia NIM "nvidia_free" Free (NGC) OpenAI Chat API meta/llama-3.1-8b-instruct 1,000 Free GPU Inference Credits (NGC)
OpenRouter "openrouter" Free & Paid OpenAI Chat API openrouter/free 19 Free models live + 390 Commercial models
OrcaRouter "orcarouter" (or "orcarouter_free") Free & Paid OpenAI Chat API qwen/qwen3.8-27b-free Zero-margin routing + Free tier models (-free)

πŸš€ Quickstart (1 Minute)

1. Installation

# With pip
pip install nexusai-client

# With uv (Recommended)
uv add nexusai-client

# With poetry
poetry add nexusai-client

# Local / Editable mode (development)
uv add --editable /path/to/NexusAI-Client

2. Configure API Keys (.env)

Create a .env file at the root of your project:

# Free Tiers (Priority 1)
GEMINI_FREE_API_KEY=your_google_ai_studio_key
GROQ_API_KEY=gsk_your_groq_key
CEREBRAS_API_KEY=csk-your_cerebras_key
COHERE_API_KEY=your_cohere_key
NVIDIA_API_KEY=nvapi-your_nvidia_nim_key
OPENROUTER_API_KEY=sk-or-v1-your_openrouter_key
ORCAROUTER_API_KEY=sk-orca-your_orcarouter_key
MISTRAL_API_KEY=your_mistral_api_key

# Paid Tiers (Backup Priority 2)
DEEPSEEK_API_KEY=sk-your_deepseek_key
GEMINI_PRO_API_KEY=your_gemini_pro_key

3. Basic Generation

import asyncio
from nexusai_client import AIGateway

async def main():
    async with AIGateway("cerebras") as client:
        response = await client.generate_text("Explain the theory of relativity in 2 sentences.")
        print(response.text)

if __name__ == "__main__":
    asyncio.run(main())

🍳 Cookbooks & Common Patterns

1. Real-Time Token Streaming (SSE)

import asyncio
from nexusai_client import AIGateway

async def main():
    async with AIGateway("groq") as client:
        async for chunk in client.stream_text("Write a short poem about lightning fast LPUs."):
            print(chunk, end="", flush=True)

if __name__ == "__main__":
    asyncio.run(main())

2. Custom Fallback Chain (Fine-Grained Strategy)

Define your own explicit priority list:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Priority: Free Gemini -> Free Groq -> Free Cerebras -> Free Cohere -> Paid DeepSeek
    custom_chain = ["gemini_free", "groq", "cerebras", "cohere", "nvidia_free", "openrouter", "deepseek"]
    async with AIGateway.with_fallback(custom_chain) as client:
        res = await client.generate_text("Summarize the key advantages of Python 3.14.")
        print(f"[{res.provider}] {res.text}")

if __name__ == "__main__":
    asyncio.run(main())

3. Multi-Turn Conversation (Chat)

import asyncio
from nexusai_client import AIGateway, ChatMessage

async def main():
    history = [
        ChatMessage(role="system", content="You are a senior algorithms instructor."),
        ChatMessage(role="user", content="How does QuickSort work?"),
    ]
    async with AIGateway("cohere") as client:
        response = await client.chat(history)
        print(response.text)

if __name__ == "__main__":
    asyncio.run(main())

4. Guaranteed Structured JSON Output

import asyncio, json
from nexusai_client import AIGateway

async def main():
    async with AIGateway("groq") as client:
        res = await client.generate_text(
            prompt="Extract profile data: Alice, 28 years old, Software Engineer.",
            json_mode=True,
        )
        data = json.loads(res.text)
        print("Parsed JSON:", data)

if __name__ == "__main__":
    asyncio.run(main())

5. Inspect Real-Time Account Balances & Quotas

import asyncio
from nexusai_client import AIGateway

async def main():
    async with AIGateway("deepseek") as client:
        account = await client.get_account_info()
        print(account.format_summary())
        # Output: "Solde restant: $4.99 | (Offert: $0.00)"

if __name__ == "__main__":
    asyncio.run(main())

6. Multimodal Vision Analysis (Images, Charts, PDFs)

Pass a local file path (Path or str), raw bytes, or web URL:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Automatically selects the best Vision model (Gemini 2.5 Flash, Llama 3.2 Vision, Qwen 3.8 Vision, Aya Vision, Pixtral)
    async with AIGateway.auto_fallback_vision() as client:
        res = await client.analyze_image(
            prompt="Extract the invoice total and line items formatted as JSON.",
            image="invoice.png", # or "https://example.com/chart.jpg" or raw bytes
            json_mode=True,
        )
        print(f"[{res.provider} / {res.model}]:")
        print(res.text)

if __name__ == "__main__":
    asyncio.run(main())

7. Universal Tool Calling (Autonomous AI Agents)

Equip AI models with callable tools across all providers (Groq, Cerebras, Mistral, DeepSeek, Gemini, Cohere, etc.):

import asyncio
from nexusai_client import AIGateway, ChatMessage, FunctionDefinition, ToolDefinition

# Define tool schema
weather_tool = ToolDefinition(
    function=FunctionDefinition(
        name="get_current_weather",
        description="Get current temperature and conditions for a given city.",
        parameters={
            "type": "object",
            "properties": {
                "location": {"type": "string", "description": "City name, e.g. Tokyo, Paris"},
                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
            },
            "required": ["location"],
        },
    )
)

async def main():
    async with AIGateway.auto_fallback() as client:
        messages = [ChatMessage(role="user", content="What is the weather in Tokyo?")]
        response = await client.chat(messages=messages, tools=[weather_tool])

        if response.has_tool_calls:
            for call in response.tool_calls:
                print(f"πŸ”§ Tool Requested: {call.name}")
                print(f"πŸ“¦ Arguments: {call.arguments}")
                # Execute your local Python function and return result back to agent loop!

if __name__ == "__main__":
    asyncio.run(main())

8. Intelligent Gemini Free Model Rotation (Auto 429 Quota Failover)

Google AI Studio Free Tier offers world-class models (gemini-3.5-flash-lite, gemini-3.7-flash, etc.) with massive context windows (up to 1M tokens) at zero cost. However, Google enforces strict per-model rate limits and daily quota pools:

  • Flash-Lite Models: ~500 Requests/day (RPD), 15 RPM, 250k TPM
  • Flash Models: ~20 Requests/day (RPD), 15 RPM, 1M TPM
  • Gemma Open Models: ~14,400 Requests/day (RPD), 30 RPM, 30k TPM

When a single model hits its quota limit (HTTP 429 RESOURCE_EXHAUSTED), traditional SDKs fail immediately. NexusAI-Client solves this natively with a 2-Tier Fallback Hierarchy:

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚                                   LEVEL 1: INTRA-PROVIDER SMART MODEL ROTATION                             β”‚
 β”‚                                                                                                             β”‚
 β”‚  1. gemini-3.5-flash-lite (500 RPD) ──► 2. gemini-3.1-flash-lite (500 RPD) ──► 3. gemini-flash-lite-latest   β”‚
 β”‚                                                                                                             β”‚
 β”‚       β–Ό (If 429 Quota Exceeded)                                                                             β”‚
 β”‚  4. gemini-3.7-flash (20 RPD)       ──► 5. gemini-3.6-flash (20 RPD)       ──► 6. gemini-3.5-flash (20 RPD)   β”‚
 β”‚                                                                                                             β”‚
 β”‚       β–Ό (If 429 Quota Exceeded)                                                                             β”‚
 β”‚  7. gemini-flash-latest             ──► 8. gemini-2.5-flash-lite (500 RPD) ──► 9. gemini-2.5-flash (20 RPD)   β”‚
 β”‚                                                                                                             β”‚
 β”‚       β–Ό (If 429 Quota Exceeded)                                                                             β”‚
 β”‚  10. gemma-4-31b-it (14.4k RPD)     ──► 11. gemma-4-26b-a4b-it (14.4k RPD)                                  β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                        β”‚ (Only if ALL 11 Gemini models exhausted)
                                                        β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚                               LEVEL 2: INTER-PROVIDER AUTO-FALLBACK FAILOVER                                β”‚
 β”‚   Groq LPU ──► Cerebras CS-3 ──► Nvidia NIM ──► OrcaRouter ──► Mistral ──► Cohere ──► OpenRouter ──► DeepSeekβ”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“‹ Complete 11-Model Cascade Chain

Step Model Identifier Quotas (Free Tier) Context Window Primary Use Case
1 gemini-3.5-flash-lite (Default) 500 RPD | 15 RPM | 250k TPM 1,048,576 tokens Ultra-fast default general queries & tool calling
2 gemini-3.1-flash-lite 500 RPD | 15 RPM | 250k TPM 1,048,576 tokens Fast secondary lightweight failover
3 gemini-flash-lite-latest 500 RPD | 15 RPM | 250k TPM 1,048,576 tokens Latest stable flash-lite pointer alias
4 gemini-3.7-flash 20 RPD | 15 RPM | 1,000,000 TPM 1,048,576 tokens High-reasoning & complex logic tasks
5 gemini-3.6-flash 20 RPD | 15 RPM | 1,000,000 TPM 1,048,576 tokens Advanced multimodal & code synthesis
6 gemini-3.5-flash 20 RPD | 15 RPM | 1,000,000 TPM 1,048,576 tokens General multimodal & vision failover
7 gemini-flash-latest 20 RPD | 15 RPM | 1,000,000 TPM 1,048,576 tokens Latest stable flash pointer alias
8 gemini-2.5-flash-lite 500 RPD | 15 RPM | 250k TPM 1,048,576 tokens Previous-generation fast fallback
9 gemini-2.5-flash 20 RPD | 15 RPM | 1,000,000 TPM 1,048,576 tokens Previous-generation robust fallback
10 gemma-4-31b-it 14,400 RPD | 30 RPM | 30k TPM 131,072 tokens High-volume open-weight model with massive RPD
11 gemma-4-26b-a4b-it 14,400 RPD | 30 RPM | 30k TPM 131,072 tokens Final emergency high-RPD free tier model

πŸ‘οΈ Multimodal Vision Rotation Sequence

For analyze_image(), the rotation automatically restricts itself to the 7 vision-capable Gemini models: gemini-3.5-flash-lite βž” gemini-3.1-flash-lite βž” gemini-3.7-flash βž” gemini-3.6-flash βž” gemini-3.5-flash βž” gemini-2.5-flash-lite βž” gemini-2.5-flash.

⏱️ Stateful Cooldown & Zero Latency Penalty

  • Automatic Cooldowns: When a model returns HTTP 429, it is immediately placed in cooldown (cooldown_seconds=60.0 by default). Deprecated/404 models are cooled down for 1 hour.
  • Zero Retrial Overhead: Subsequent requests during the same process lifecycle instantly skip cooled-down models, routing directly to the first available operational model with 0ms penalty.
  • Real-Time Visibility: Inspect live model statuses and cooldown timers with get_account_info().

πŸ’» Code Examples

1. Transparent Auto-Failover (Zero Configuration):

import asyncio
from nexusai_client import AIGateway

async def main():
    # If the primary model (gemini-3.5-flash-lite) hits a 429 limit,
    # it immediately retries with gemini-3.1-flash-lite, gemini-3.7-flash, etc.
    async with AIGateway("gemini_free") as client:
        res = await client.generate_text("Explain quantum entanglement simply.")
        print(f"βœ… Served by model [{res.model}] (Provider: {res.provider}):")
        print(res.text)

if __name__ == "__main__":
    asyncio.run(main())

2. Live Rotation & Cooldown Diagnostics:

import asyncio
from nexusai_client import AIGateway

async def main():
    async with AIGateway("gemini_free") as client:
        info = await client.get_account_info()
        print(f"Platform: {info.extra_details['platform']}")
        print(f"Current Active Model: {info.extra_details['current_active_model']}")
        print(f"Models in Cooldown: {info.extra_details['models_in_cooldown']}")
        print(f"Quota Limits: {info.rate_limit_info}")

if __name__ == "__main__":
    asyncio.run(main())

3. Customizing Fallback Sequence & Cooldowns:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Define custom priority models and a 120s cooldown period
    custom_models = ["gemini-3.7-flash", "gemini-3.5-flash-lite", "gemma-4-31b-it"]
    async with AIGateway(
        "gemini_free",
        fallback_models=custom_models,
        cooldown_seconds=120.0,
        auto_rotate_models=True,
    ) as client:
        res = await client.generate_text("Write an async Python pipeline.")
        print(f"[{res.model}]: {res.text}")

if __name__ == "__main__":
    asyncio.run(main())

4. Real-Time Token Streaming with Model Rotation:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Streaming seamlessly falls back to candidate models if 429 is encountered at stream start
    async with AIGateway("gemini_free") as client:
        async for chunk in client.stream_text("Explain distributed systems in 3 bullet points."):
            print(chunk, end="", flush=True)

if __name__ == "__main__":
    asyncio.run(main())

9. 100% Free Multi-Provider Fallback (auto_fallback_free)

Build zero-cost resilient workflows by cascading across all free providers present in your .env:

import asyncio
from nexusai_client import AIGateway

async def main():
    # Cascade: Gemini Free (11 models) -> Groq LPU -> Cerebras CS-3 -> Nvidia NIM -> OrcaRouter -> Mistral -> Cohere -> OpenRouter
    async with AIGateway.auto_fallback_free() as client:
        res = await client.generate_text("Write a concise summary of AI agents architecture.")
        print(f"βœ… Served by [{res.provider} / {res.model}]:\n{res.text}")

if __name__ == "__main__":
    asyncio.run(main())

πŸ‘‰ Read the Full Integration Guide (FastAPI, Background Workers, Chat Sessions)


πŸ› οΈ CLI Utilities Included

The package includes CLI diagnostic tools to audit your accesses and explore live models:

1. Test & Benchmark Access in Real-Time

uv run python verify_access.py

Validates .env keys, inspects real-time balances, tests inference, and measures network latency in milliseconds.

2. Live Catalog Explorer (670+ Models)

# List free-tier models only
uv run python list_all_models.py --free

# Search by keyword (e.g., llama, r1, command, sonnet)
uv run python list_all_models.py --search llama

# Export complete catalog with pricing to JSON
uv run python list_all_models.py --export models_catalog.json

3. Automated Unit Test Suite

uv run pytest -v

πŸ›‘οΈ Strongly-Typed Exceptions

All exceptions inherit from NexusAIError for clean error handling:

from nexusai_client import (
    AIGateway,
    NexusAIError,
    MissingAPIKeyError,    # Missing API key in environment
    AuthenticationError,   # Invalid key (HTTP 401/403)
    RateLimitError,        # Quota exceeded (HTTP 429)
    APITimeoutError,       # Network timeout
    APIConnectionError,    # Unreachable provider host
    ProviderNotFoundError, # Unknown provider requested
)

πŸ“„ License

This project is licensed under the MIT License. Free for personal and commercial use.

About

An ultra-lightweight, asynchronous Python 3.14 client to centralize and unify calls across AI providers (Cerebras, Cohere, DeepSeek, Gemini, Groq, Mistral, Nvidia NIM, OpenRouter, OrcaRouter). Powered by uv.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages