Free AI inference for real work β as a coding agent, or as a gateway for your own tools.
FreeRide routes requests across free-tier providers β OpenRouter, Groq, NVIDIA NIM, HuggingFace, Cerebras, Cloudflare Workers AI, and your own Ollama β with automatic failover when one rate-limits or errors. No vendor subscription, no cloud middleman: your machine talks to the providers directly with your own free keys. There are two ways in; both run on the same engine.
102M+ tokens served in 35 days. $0 spent. Routed through community free-tier keys via this gateway. Daily traffic: free-ride.xyz/models
A fast, native agent (our fork of vercel-labs/fx, Apache-2.0) that reads files, edits code, and runs commands β every model call served free through FreeRide. One command installs the agent and the gateway, supervises the daemon, and walks you through keys on first run:
curl -sSL https://api.free-ride.xyz/ridex.sh | sh
ridex # interactive agent
ridex ask "reply with the single word pong" # or one-shotAgent releases: github.com/Shaivpidadi/ridex. macOS + Linux (arm64/x86_64).
A local, OpenAI-compatible endpoint on localhost:11343 that any agent, SDK, or script can point at β Aider, Continue.dev, LangChain, raw openai clients, your own code:
curl -sSL https://api.free-ride.xyz/install.sh | sh
freeride init && freeride serveOPENAI_API_BASE=http://localhost:11343/v1
OPENAI_API_KEY=any-string-here # inbound auth is ignored; your real keys stay server-sideIt also natively serves the Anthropic (/v1/messages), OpenAI Responses (/v1/responses), Gemini (:generateContent), and fx gateway wire protocols, plus /v1/embeddings β so most tools work unmodified. Prefer the big-vendor CLIs' UX? freeride run claude|codex|gemini wraps them with no vendor login. macOS + Linux + Windows.
Already running ridex? You have the gateway too β one shared daemon on :11343 serves your agent sessions and your own tools with the same keys, failover chain, and cooldowns.
Any one is enough; more = better failover. Collected by freeride init (or ridex's first run) and stored only on your machine in ~/.freeride/.env:
| Provider | Free tier | Get a key |
|---|---|---|
| OpenRouter | rotating free models | openrouter.ai/keys |
| Groq | daily token cap | console.groq.com/keys |
| NVIDIA NIM | credits per account | build.nvidia.com |
| HuggingFace | $0.10/mo Free, $2/mo PRO | huggingface.co/settings/tokens |
| Cerebras | RPM / TPM caps | cloud.cerebras.ai |
| Cloudflare Workers AI | 10K neurons/day | dash.cloudflare.com |
| Ollama (local) | no quota | install from ollama.com |
Free tiers are flaky β that's the whole reason FreeRide exists. Whichever path you picked, the same engine keeps a single provider's bad minute from ever reaching you:
- The gateway is a supervised daemon (via the ridex installer): registered with launchd (macOS) or a systemd user unit (Linux), a crash restarts it in seconds.
ridex start|stop|restart|doctormanage it, andridex stopsticks until you say otherwise. (Standalonefreeride serveusers manage the process themselves, as before.) - Every request carries a fallback ladder. If the serving provider rate-limits, runs out of free inference, or retires the model mid-session, the gateway silently retries on the next provider's best tool-capable model β inside the same response. Failed candidates are remembered for a few minutes so consecutive turns don't re-pay the cost.
- The agent can diagnose its own plumbing. ridex ships with a
freerideskill: when requests fail it runs the (pre-approved, read-only) diagnostics βfreeride doctor,freeride keys, the health probe β reads the structured error taxonomy, and tells you the exact fix. - Tool calls are non-negotiable. The default
freeride/codingroute pins to models proven to emit correct tool calls; providers whose catalogs can't do tools are never handed agent traffic.
Inside a session, /model switches routing per request: freeride/coding (default), freeride/fast (Groq-first, low TTFT), freeride/quality (OpenRouter-first, widest catalog), freeride/free / auto (pure smart-routing), or any concrete model id from ridex models.
Prefer Claude Code, OpenAI Codex, or Gemini CLI's UX? freeride run points them at the gateway β no per-vendor key, no login:
freeride run claude # /model freeride/coding etc. inside the session
freeride run codex # Responses-API wire format translated natively
freeride run gemini # Google's {contents, tools, generationConfig} shape both waysGuides: docs/agents/claude-code.md Β· docs/agents/codex.md Β· docs/agents/gemini.md. For Aider, Continue.dev, and friends: freeride bind <agent> (docs/agents/binders.md).
Per-request the chain is (provider, key), sorted by recent health:
- Try the head pair.
RATE_LIMITorAUTHerror β mark the key as cooling, try the next key on the same provider.MODEL_NOT_FOUNDorQUOTA_EXHAUSTEDβ skip to the next provider.- 5xx / TIMEOUT β next pair.
- First successful response β stamp
X-FreeRide-Provider+X-FreeRide-Request-Idheaders and ship.
Agent traffic gets a second layer on top: the candidate ladder walks (provider, tool-capable model) pairs, so even a model that exists on only one cooling provider falls through to a working equivalent elsewhere β silently, under streaming keepalives. If every pair fails, you get a structured 503 with a per-provider breakdown so debugging is one log line, not five round-trips. An upstream dying mid-stream before any output switches candidates invisibly; after output the turn ends as an explicit error (agents retry it) rather than a silently truncated answer.
Smart routing for model: "auto": the resolver scores every free model in the catalog by health Γ popularity (from the public models leaderboard) and picks the best one. Run freeride audit-models once after install to cache health probes locally so the first real request isn't a cold start.
Deeper: docs/architecture/failover.md.
| Provider | Surface | Notes |
|---|---|---|
| OpenRouter | chat, streaming, tools, vision, structured outputs, embeddings | full surface β the most-used provider in our routing |
| NVIDIA NIM | chat + embeddings | curated free-model allowlist; NVIDIA_NIM_FREE_MODELS_OVERRIDE to expand |
| Groq | chat | Llama 3.x, Gemma 2, Mixtral, DeepSeek-R1-distill; daily token cap |
| Cloudflare Workers AI | chat | cheap-per-neuron models; needs CLOUDFLARE_ACCOUNT_ID |
| HuggingFace Inference | chat + embeddings | full HF router catalog; budget governs access |
| Cerebras | chat | fastest Llama / Qwen inference; no embeddings |
| Ollama (local) | chat | local-only; can mix with remote in the same failover chain |
Adding a new provider: implement freeride.core.provider.Provider in freeride/providers/<name>.py, register it in the conformance suite. See CONTRIBUTING.md.
Provide more than one key per provider with a numbered suffix:
OPENROUTER_API_KEY=sk-or-v1-aaa # primary
OPENROUTER_API_KEY_2=sk-or-v1-bbb
OPENROUTER_API_KEY_3=sk-or-v1-cccThe router tries them in health order. A 429 on one key cools it for the next 60s and rotates to the sibling key β no provider switch needed. On startup freeride keys shows which keys are available vs cooling.
ridex doctor # agent binary + daemon + key status in one report
freeride doctor # static checks: keys, ports, /etc/hosts, common gotchas
freeride audit-models # probe every free model on every key; cache the results
freeride bench # measure p50/p95/tok-s per providerTail live events:
tail -f ~/.freeride/events.jsonlEach line is a JSON event: routing decisions, provider attempts, ladder fallbacks, response statuses, mid-stream errors. Same schema the marketing site reads to render the live token counter and provider leaderboard.
A small beacon ships hourly with counts only: tokens served, request count, active providers, uptime hours, OS, version, and a per-install UUID. Never sent: prompts, completions, model IDs, API keys, hostname, IP.
freeride telemetry # audit what the next beacon would post
freeride telemetry off # opt outThe aggregate is what powers free-ride.xyz/models. Default on; explicit disclosure banner prints on first run.
ridex interactive coding agent (auto-starts the gateway daemon)
ridex ask <prompt> one noninteractive agent request
ridex models list available models
ridex start|stop|restart manage the gateway daemon (stop sticks)
ridex doctor agent + daemon + key health report
freeride init interactive setup wizard β prompts for keys, writes ~/.freeride/.env
freeride serve start the gateway on :11343 (the daemon runs this for you)
freeride run <cli> wrap a CLI (claude / codex / gemini) β points it at the gateway
freeride bind <agent> write the agent's config so it uses the gateway permanently
freeride doctor pre-flight checks: keys, ports, hosts file, common gotchas
freeride keys which provider keys are available vs cooling
freeride reload hot-reload provider keys on a running gateway
freeride audit-models probe every free model; cache health locally
freeride bench measure p50/p95/tok-s per provider
freeride list list available free models
freeride telemetry manage the hourly aggregate beacon
- The agent
- github.com/Shaivpidadi/ridex β the ridex agent (fork of vercel-labs/fx)
internal-docs/RIDEX_PLAN.mdβ architecture decisions + verification log
- Wrapped CLIs
docs/agents/claude-code.mdβ Claude Code setup,/modelmodes, troubleshootingdocs/agents/codex.mdβ OpenAI Codex setup, bwrap notes, model selectiondocs/agents/gemini.mdβ Google Gemini CLI setup, auth flow, model selectiondocs/agents/binders.mdβ Aider, Continue, OpenClaw β per-agentfreeride bindreferencedocs/agents/hermes.mdβ NousResearch Hermes agent integration
- Providers
docs/providers/SURVEY.mdβ per-provider fit (auth, free-tier semantics, error mapping)docs/providers/nvidia_nim.mdβ NVIDIA NIM specifics
- Architecture
docs/architecture/failover.mdβ failover chain, cooldown, health trackingdocs/architecture/translators.mdβ how the Anthropic / Google / OpenAI-Responses / fx translators work
- Other
CONTRIBUTING.mdβ adding a provider, a CLI wrapper, or a binderSECURITY.mdβ reporting vulnerabilities
MIT. The ridex agent is a fork of vercel-labs/fx (Apache-2.0); its license and notices ship with every release tarball.
