An ultra-lightweight, strictly-typed, asynchronous Python 3.12+ gateway for multi-provider AI APIs.
Unify Cerebras, Cohere, DeepSeek, Google Gemini (Free & Pro), Groq, Mistral, Nvidia NIM, OpenRouter, and OrcaRouter behind a single, elegant interface with zero heavy SDK dependencies.
π Interactive Documentation Website: https://nexus-ai-client-doc.vercel.app/
Integrating multiple AI providers in modern Python applications usually requires installing 9 or 10 separate proprietary SDKs (google-genai, openai, groq, cohere, mistralai, etc.). This creates dozens of transitive dependencies, version conflicts, memory overhead, and fragmented codebases.
NexusAI-Client solves this at the core:
- π Dynamic Model Management & Auto-Rotation: Automatic failover on HTTP 404/400 (deprecated models), HTTP 429 rate limits, and timeouts across models and providers.
- πͺΆ Zero Heavyweight Dependencies: Powered purely by
httpxandpython-dotenv. - β‘ Native Asynchronous & SSE Streaming: Stream responses token-by-token in real time via
stream_text()andstream_chat(). - π Zero-Cost-First Smart Fallback: Automatic progression from 100% Free Tiers (Gemini, Groq, Cerebras, Cohere, Nvidia, OpenRouter, OrcaRouter, Mistral) to Paid Backups (
AIGateway.auto_fallback()). - π οΈ Universal Tool Calling / Function Calling: Seamless tool definitions, structured function arguments parsing, and multi-turn agent loops across Groq, Cerebras, Mistral, DeepSeek, Gemini REST, Cohere V2, and Nvidia NIM.
- π World-Record Hardware Accelerators: Native support for Groq LPUs and Cerebras CS-3 Wafer-Scale engines (2,000+ tokens/sec).
- π§ Enterprise Reasoning & Search Models: Native Cohere Command R+, DeepSeek R1, and Qwen 3.8 models.
- π― Guaranteed JSON Outputs: Native
json_mode=Trueacross all supported providers. - π° Live Account & Budget Inspection: Inspect real-time balances (USD, NGC credits) and rate limits (RPM, TPM, RPD).
- π 670+ Models Discovered Live: Automatic detection of free-tier models (
:free,-free) and accurate per-million-token pricing.
Why pay for AI calls when you can leverage high-throughput free tiers first, with seamless automatic fallback to paid commercial models?
NexusAI-Client automatically prioritizes zero-cost models before touching your wallet:
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β 100% FREE ZERO-COST TIERS β
ββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββ¬βββββββββββββββ¬βββββββββββββββ¬ββββββββββββ€
β 1. Gemini β 2. Groq LPU β 3. Cerebras CS-3 β 4. Nvidia β 5. OpenRouterβ 6. OrcaRouterβ 7. Cohere β 8. Mistralβ
β (1M Context) β (Ultra-Fast) β (2000+ tok/s) β (1k Credits) β (Free Hub) β (Qwen/DeepS) β (Command R+) β (Dev Free)β
ββββββββ¬ββββββββ΄βββββββ¬ββββββββ΄βββββββββ¬ββββββββββ΄βββββββ¬ββββββββ΄βββββββ¬ββββββββ΄βββββββ¬ββββββββ΄βββββββ¬ββββββββ΄ββββββ¬ββββββ
β β β β β β β β
βΌ (If Rate-Limited / 429 Quota Exceeded / Network Outage / Timeout) βββββββββββββββββββββββββββββββββββββββββΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β ULTRA-LOW-COST PAID BACKUP TIERS β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 9. DeepSeek ($0.27 / 1M tokens) β 10. Gemini Pro (Enterprise GCP) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
import asyncio
from nexusai_client import AIGateway
async def main():
# Automatically discovers active keys in .env and routes: Free -> Free -> Paid
async with AIGateway.auto_fallback() as client:
response = await client.generate_text("Explain quantum computing in 2 sentences.")
print(f"β
Served by [{response.provider}] with zero downtime:")
print(response.text)
if __name__ == "__main__":
asyncio.run(main())| Provider | Identifier (provider) |
Tier | Protocol | Default Model | Live Budget & Quota Detection |
|---|---|---|---|---|---|
| Cerebras | "cerebras" (or "cerebras_free") |
Free (CS-3) | OpenAI Chat API | gpt-oss-120b |
Quotas: 30 RPM | 60k TPM | 1M tok/day |
| Cohere | "cohere" (or "cohere_free") |
Free Trial | Cohere V2 REST | command-r-plus-08-2024 |
Quotas: 20 RPM | 1,000 calls/month |
| DeepSeek | "deepseek" |
Paid | OpenAI Chat API | deepseek-chat |
Real-time USD Balance (GET /user/balance) |
| Gemini Free | "gemini_free" |
Free (AI Studio) | Gemini REST | gemini-3.5-flash-lite |
Auto-rotation 429 | 15 RPM | 500 RPD (Lite) / 20 RPD (Flash) |
| Gemini Pro | "gemini_pro" |
Paid | Gemini REST | gemini-3.1-pro-preview |
Google Cloud Pay-as-you-go Billing |
| Groq | "groq" (or "groq_free") |
Free (LPU) | OpenAI Chat API | openai/gpt-oss-120b |
Quotas: 30 RPM | 14,400 RPD | 30k TPM |
| Mistral AI | "mistral" |
Free / Platform | OpenAI Chat API | mistral-small-latest |
Free Dev Models (codestral-latest, etc.) |
| Nvidia NIM | "nvidia_free" |
Free (NGC) | OpenAI Chat API | meta/llama-3.1-8b-instruct |
1,000 Free GPU Inference Credits (NGC) |
| OpenRouter | "openrouter" |
Free & Paid | OpenAI Chat API | openrouter/free |
19 Free models live + 390 Commercial models |
| OrcaRouter | "orcarouter" (or "orcarouter_free") |
Free & Paid | OpenAI Chat API | qwen/qwen3.8-27b-free |
Zero-margin routing + Free tier models (-free) |
# With pip
pip install nexusai-client
# With uv (Recommended)
uv add nexusai-client
# With poetry
poetry add nexusai-client
# Local / Editable mode (development)
uv add --editable /path/to/NexusAI-ClientCreate a .env file at the root of your project:
# Free Tiers (Priority 1)
GEMINI_FREE_API_KEY=your_google_ai_studio_key
GROQ_API_KEY=gsk_your_groq_key
CEREBRAS_API_KEY=csk-your_cerebras_key
COHERE_API_KEY=your_cohere_key
NVIDIA_API_KEY=nvapi-your_nvidia_nim_key
OPENROUTER_API_KEY=sk-or-v1-your_openrouter_key
ORCAROUTER_API_KEY=sk-orca-your_orcarouter_key
MISTRAL_API_KEY=your_mistral_api_key
# Paid Tiers (Backup Priority 2)
DEEPSEEK_API_KEY=sk-your_deepseek_key
GEMINI_PRO_API_KEY=your_gemini_pro_keyimport asyncio
from nexusai_client import AIGateway
async def main():
async with AIGateway("cerebras") as client:
response = await client.generate_text("Explain the theory of relativity in 2 sentences.")
print(response.text)
if __name__ == "__main__":
asyncio.run(main())import asyncio
from nexusai_client import AIGateway
async def main():
async with AIGateway("groq") as client:
async for chunk in client.stream_text("Write a short poem about lightning fast LPUs."):
print(chunk, end="", flush=True)
if __name__ == "__main__":
asyncio.run(main())Define your own explicit priority list:
import asyncio
from nexusai_client import AIGateway
async def main():
# Priority: Free Gemini -> Free Groq -> Free Cerebras -> Free Cohere -> Paid DeepSeek
custom_chain = ["gemini_free", "groq", "cerebras", "cohere", "nvidia_free", "openrouter", "deepseek"]
async with AIGateway.with_fallback(custom_chain) as client:
res = await client.generate_text("Summarize the key advantages of Python 3.14.")
print(f"[{res.provider}] {res.text}")
if __name__ == "__main__":
asyncio.run(main())import asyncio
from nexusai_client import AIGateway, ChatMessage
async def main():
history = [
ChatMessage(role="system", content="You are a senior algorithms instructor."),
ChatMessage(role="user", content="How does QuickSort work?"),
]
async with AIGateway("cohere") as client:
response = await client.chat(history)
print(response.text)
if __name__ == "__main__":
asyncio.run(main())import asyncio, json
from nexusai_client import AIGateway
async def main():
async with AIGateway("groq") as client:
res = await client.generate_text(
prompt="Extract profile data: Alice, 28 years old, Software Engineer.",
json_mode=True,
)
data = json.loads(res.text)
print("Parsed JSON:", data)
if __name__ == "__main__":
asyncio.run(main())import asyncio
from nexusai_client import AIGateway
async def main():
async with AIGateway("deepseek") as client:
account = await client.get_account_info()
print(account.format_summary())
# Output: "Solde restant: $4.99 | (Offert: $0.00)"
if __name__ == "__main__":
asyncio.run(main())Pass a local file path (Path or str), raw bytes, or web URL:
import asyncio
from nexusai_client import AIGateway
async def main():
# Automatically selects the best Vision model (Gemini 2.5 Flash, Llama 3.2 Vision, Qwen 3.8 Vision, Aya Vision, Pixtral)
async with AIGateway.auto_fallback_vision() as client:
res = await client.analyze_image(
prompt="Extract the invoice total and line items formatted as JSON.",
image="invoice.png", # or "https://example.com/chart.jpg" or raw bytes
json_mode=True,
)
print(f"[{res.provider} / {res.model}]:")
print(res.text)
if __name__ == "__main__":
asyncio.run(main())Equip AI models with callable tools across all providers (Groq, Cerebras, Mistral, DeepSeek, Gemini, Cohere, etc.):
import asyncio
from nexusai_client import AIGateway, ChatMessage, FunctionDefinition, ToolDefinition
# Define tool schema
weather_tool = ToolDefinition(
function=FunctionDefinition(
name="get_current_weather",
description="Get current temperature and conditions for a given city.",
parameters={
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name, e.g. Tokyo, Paris"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
)
)
async def main():
async with AIGateway.auto_fallback() as client:
messages = [ChatMessage(role="user", content="What is the weather in Tokyo?")]
response = await client.chat(messages=messages, tools=[weather_tool])
if response.has_tool_calls:
for call in response.tool_calls:
print(f"π§ Tool Requested: {call.name}")
print(f"π¦ Arguments: {call.arguments}")
# Execute your local Python function and return result back to agent loop!
if __name__ == "__main__":
asyncio.run(main())Google AI Studio Free Tier offers world-class models (gemini-3.5-flash-lite, gemini-3.7-flash, etc.) with massive context windows (up to 1M tokens) at zero cost. However, Google enforces strict per-model rate limits and daily quota pools:
- Flash-Lite Models: ~500 Requests/day (RPD), 15 RPM, 250k TPM
- Flash Models: ~20 Requests/day (RPD), 15 RPM, 1M TPM
- Gemma Open Models: ~14,400 Requests/day (RPD), 30 RPM, 30k TPM
When a single model hits its quota limit (HTTP 429 RESOURCE_EXHAUSTED), traditional SDKs fail immediately. NexusAI-Client solves this natively with a 2-Tier Fallback Hierarchy:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LEVEL 1: INTRA-PROVIDER SMART MODEL ROTATION β
β β
β 1. gemini-3.5-flash-lite (500 RPD) βββΊ 2. gemini-3.1-flash-lite (500 RPD) βββΊ 3. gemini-flash-lite-latest β
β β
β βΌ (If 429 Quota Exceeded) β
β 4. gemini-3.7-flash (20 RPD) βββΊ 5. gemini-3.6-flash (20 RPD) βββΊ 6. gemini-3.5-flash (20 RPD) β
β β
β βΌ (If 429 Quota Exceeded) β
β 7. gemini-flash-latest βββΊ 8. gemini-2.5-flash-lite (500 RPD) βββΊ 9. gemini-2.5-flash (20 RPD) β
β β
β βΌ (If 429 Quota Exceeded) β
β 10. gemma-4-31b-it (14.4k RPD) βββΊ 11. gemma-4-26b-a4b-it (14.4k RPD) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β (Only if ALL 11 Gemini models exhausted)
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LEVEL 2: INTER-PROVIDER AUTO-FALLBACK FAILOVER β
β Groq LPU βββΊ Cerebras CS-3 βββΊ Nvidia NIM βββΊ OrcaRouter βββΊ Mistral βββΊ Cohere βββΊ OpenRouter βββΊ DeepSeekβ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
| Step | Model Identifier | Quotas (Free Tier) | Context Window | Primary Use Case |
|---|---|---|---|---|
| 1 | gemini-3.5-flash-lite (Default) |
500 RPD | 15 RPM | 250k TPM | 1,048,576 tokens | Ultra-fast default general queries & tool calling |
| 2 | gemini-3.1-flash-lite |
500 RPD | 15 RPM | 250k TPM | 1,048,576 tokens | Fast secondary lightweight failover |
| 3 | gemini-flash-lite-latest |
500 RPD | 15 RPM | 250k TPM | 1,048,576 tokens | Latest stable flash-lite pointer alias |
| 4 | gemini-3.7-flash |
20 RPD | 15 RPM | 1,000,000 TPM | 1,048,576 tokens | High-reasoning & complex logic tasks |
| 5 | gemini-3.6-flash |
20 RPD | 15 RPM | 1,000,000 TPM | 1,048,576 tokens | Advanced multimodal & code synthesis |
| 6 | gemini-3.5-flash |
20 RPD | 15 RPM | 1,000,000 TPM | 1,048,576 tokens | General multimodal & vision failover |
| 7 | gemini-flash-latest |
20 RPD | 15 RPM | 1,000,000 TPM | 1,048,576 tokens | Latest stable flash pointer alias |
| 8 | gemini-2.5-flash-lite |
500 RPD | 15 RPM | 250k TPM | 1,048,576 tokens | Previous-generation fast fallback |
| 9 | gemini-2.5-flash |
20 RPD | 15 RPM | 1,000,000 TPM | 1,048,576 tokens | Previous-generation robust fallback |
| 10 | gemma-4-31b-it |
14,400 RPD | 30 RPM | 30k TPM | 131,072 tokens | High-volume open-weight model with massive RPD |
| 11 | gemma-4-26b-a4b-it |
14,400 RPD | 30 RPM | 30k TPM | 131,072 tokens | Final emergency high-RPD free tier model |
For analyze_image(), the rotation automatically restricts itself to the 7 vision-capable Gemini models:
gemini-3.5-flash-lite β gemini-3.1-flash-lite β gemini-3.7-flash β gemini-3.6-flash β gemini-3.5-flash β gemini-2.5-flash-lite β gemini-2.5-flash.
- Automatic Cooldowns: When a model returns
HTTP 429, it is immediately placed in cooldown (cooldown_seconds=60.0by default). Deprecated/404 models are cooled down for 1 hour. - Zero Retrial Overhead: Subsequent requests during the same process lifecycle instantly skip cooled-down models, routing directly to the first available operational model with 0ms penalty.
- Real-Time Visibility: Inspect live model statuses and cooldown timers with
get_account_info().
1. Transparent Auto-Failover (Zero Configuration):
import asyncio
from nexusai_client import AIGateway
async def main():
# If the primary model (gemini-3.5-flash-lite) hits a 429 limit,
# it immediately retries with gemini-3.1-flash-lite, gemini-3.7-flash, etc.
async with AIGateway("gemini_free") as client:
res = await client.generate_text("Explain quantum entanglement simply.")
print(f"β
Served by model [{res.model}] (Provider: {res.provider}):")
print(res.text)
if __name__ == "__main__":
asyncio.run(main())2. Live Rotation & Cooldown Diagnostics:
import asyncio
from nexusai_client import AIGateway
async def main():
async with AIGateway("gemini_free") as client:
info = await client.get_account_info()
print(f"Platform: {info.extra_details['platform']}")
print(f"Current Active Model: {info.extra_details['current_active_model']}")
print(f"Models in Cooldown: {info.extra_details['models_in_cooldown']}")
print(f"Quota Limits: {info.rate_limit_info}")
if __name__ == "__main__":
asyncio.run(main())3. Customizing Fallback Sequence & Cooldowns:
import asyncio
from nexusai_client import AIGateway
async def main():
# Define custom priority models and a 120s cooldown period
custom_models = ["gemini-3.7-flash", "gemini-3.5-flash-lite", "gemma-4-31b-it"]
async with AIGateway(
"gemini_free",
fallback_models=custom_models,
cooldown_seconds=120.0,
auto_rotate_models=True,
) as client:
res = await client.generate_text("Write an async Python pipeline.")
print(f"[{res.model}]: {res.text}")
if __name__ == "__main__":
asyncio.run(main())4. Real-Time Token Streaming with Model Rotation:
import asyncio
from nexusai_client import AIGateway
async def main():
# Streaming seamlessly falls back to candidate models if 429 is encountered at stream start
async with AIGateway("gemini_free") as client:
async for chunk in client.stream_text("Explain distributed systems in 3 bullet points."):
print(chunk, end="", flush=True)
if __name__ == "__main__":
asyncio.run(main())Build zero-cost resilient workflows by cascading across all free providers present in your .env:
import asyncio
from nexusai_client import AIGateway
async def main():
# Cascade: Gemini Free (11 models) -> Groq LPU -> Cerebras CS-3 -> Nvidia NIM -> OrcaRouter -> Mistral -> Cohere -> OpenRouter
async with AIGateway.auto_fallback_free() as client:
res = await client.generate_text("Write a concise summary of AI agents architecture.")
print(f"β
Served by [{res.provider} / {res.model}]:\n{res.text}")
if __name__ == "__main__":
asyncio.run(main())π Read the Full Integration Guide (FastAPI, Background Workers, Chat Sessions)
The package includes CLI diagnostic tools to audit your accesses and explore live models:
uv run python verify_access.pyValidates .env keys, inspects real-time balances, tests inference, and measures network latency in milliseconds.
# List free-tier models only
uv run python list_all_models.py --free
# Search by keyword (e.g., llama, r1, command, sonnet)
uv run python list_all_models.py --search llama
# Export complete catalog with pricing to JSON
uv run python list_all_models.py --export models_catalog.jsonuv run pytest -vAll exceptions inherit from NexusAIError for clean error handling:
from nexusai_client import (
AIGateway,
NexusAIError,
MissingAPIKeyError, # Missing API key in environment
AuthenticationError, # Invalid key (HTTP 401/403)
RateLimitError, # Quota exceeded (HTTP 429)
APITimeoutError, # Network timeout
APIConnectionError, # Unreachable provider host
ProviderNotFoundError, # Unknown provider requested
)This project is licensed under the MIT License. Free for personal and commercial use.
