Skip to content

Standardize Fleet AI on direct free-model clients and retire the shared gateway #61

Description

@sarthakagrawal927

Why

The owner superseded the 2026-08-29 gateway decision on 2026-08-30: Fleet will not use a shared AI gateway. Projects should use free model paths directly, with the Vercel AI SDK wherever the runtime supports it.

The current ratified standard is therefore stale and actively points migrations in the wrong direction. The corrected local audit inspected all 56 catalog projects: 4 use the Vercel AI SDK, 19 use raw HTTP, 2 use provider SDKs, and 31 do not call a hosted model. The old detector also reports gateway references and direct-provider calls using gateway-era compliance language, which is no longer trustworthy for the new decision.

What

In scope:

  • Replace the gateway-first canonical standard with direct free-provider or local-inference clients.
  • Keep exact pinned versions for JavaScript and TypeScript: ai 6.0.168 and @ai-sdk/openai-compatible 2.0.41, plus @ai-sdk/react 3.0.86 where needed.
  • Prefer an official Vercel AI SDK provider adapter when one fits; otherwise use the OpenAI-compatible adapter against the selected provider or local endpoint.
  • Keep native Swift, Rust, and Python call paths explicit when the JavaScript SDK cannot apply.
  • Treat ai-gateway.sassmaker.com and gateway-only environment variables as retired references in active runtime code.
  • Reclassify direct provider calls: they are acceptable when they are the deliberate project-owned free-model path, not gateway bypasses.
  • Migrate each affected repository in small, independently verified changes.
  • Update the Free AI project dossier, canonical catalog, public technology metadata, and generated public projection if its public form, domain, deployment target, or prominent tools change.

Out of scope:

  • Paid provider adoption or new paid spend.
  • Sharing provider credentials between projects.
  • Production deploys, credential rotation, provider-resource deletion, DNS changes, or gateway decommissioning without separate explicit authorization.
  • Turso account closure, D1 query hardening, and the Fleet credential incident response; those have separate verification and authorization boundaries.

Design

flowchart LR
  A[JS or TS product] --> B[Vercel AI SDK pinned]
  B --> C[Direct free provider or local endpoint]
  D[Swift Rust Python product] --> E[Native project-owned client]
  E --> C
  F[Shared AI gateway] --> G[Retired reference and later decommission gate]
Loading

Each product owns its model selection, free-tier limits, runtime credentials, and fallback behavior. There is no Fleet-wide request proxy. The shared tooling owns only the credential-free detector, pinned client standard, and migration evidence.

Specs

Requirement: No shared gateway

Active Fleet runtime code SHALL NOT call ai-gateway.sassmaker.com or depend on gateway-only URL or token variables.

Scenario: Existing gateway caller

  • WHEN an active project currently calls the shared gateway
  • THEN it is migrated to a direct free-provider or local endpoint and the gateway configuration is removed from that project

Requirement: Pinned Vercel AI SDK where applicable

JavaScript and TypeScript hosted-model clients SHALL use the exact canonical Vercel AI SDK pins unless a dated project-specific exception explains why the runtime cannot.

Scenario: Hand-rolled JavaScript request

  • WHEN a JS or TS project constructs a hosted-model HTTP request directly
  • THEN it is moved to the pinned Vercel AI SDK and receives focused build and behavior checks

Requirement: Honest non-JavaScript boundary

Swift, Rust, Python, and local-inference surfaces SHALL use the smallest maintained native client appropriate to their runtime and SHALL NOT be forced through a JavaScript package.

Scenario: Native client

  • WHEN the Vercel AI SDK cannot run in the product runtime
  • THEN the audit records the native path and verifies direct free-provider or local-endpoint routing

Requirement: Project-owned credentials and cost limits

Every direct provider credential SHALL remain project-scoped and runtime-only. The migration SHALL NOT introduce paid spend or a shared credential.

Scenario: Provider requires a key

  • WHEN a selected free-tier provider requires authentication
  • THEN the project uses its own runtime secret binding and documents the free-tier budget without committing the value

Requirement: Detector matches the new policy

The Fleet audit SHALL classify SDK use, native exceptions, retired gateway references, direct provider or local endpoints, and credential literals without treating all direct-provider calls as bypasses.

Scenario: Retired gateway reference

  • WHEN active source still references the gateway host or gateway-only variables
  • THEN the audit fails with the repository and file location but never prints a credential value

Tasks

  • 1. Rewrite the canonical standard and tests for direct free-model clients; mark the gateway host and variables retired.
  • 2. Update detector verdicts and report language, then run its focused test suite and the 56-project audit.
  • 3. Re-audit the 25 hosted-model callers and split them into JS/TS SDK migrations, already-compliant direct clients, native clients, and no-longer-needed calls.
  • 4. Migrate the 19 hand-rolled projects in small repository-owned batches, starting with active JS/TS runtime call sites.
  • 5. Remove active gateway host and environment references from SDK adopters and native clients.
  • 6. Run the smallest relevant build, typecheck, and behavior test in every changed repository; record skipped hardware or provider checks honestly.
  • 7. Refresh the Free AI dossier, canonical catalog, public projection, and durable project status if its public form or tooling changes.
  • 8. Prepare a separate, explicitly authorized decommission checklist for the gateway deployment, domain, secrets, and provider resources.

No deploy, credential mutation, provider deletion, commit, push, or release is authorized by this issue update.

Activity

  1. sarthakagrawal927 commented on Aug 29, 2026

    @sarthakagrawal927
    MemberAuthor

    Audit shipped in #81 (merged). This issue stays open on one decision that is yours: ratifying the canonical model-calling path. Nothing is enforced until you do — tooling/config/ai-client-standard.json is "status": "unratified", and the validator exits 0.

    What the audit found across the 56 catalog projects: 3 compliant (karte, reader, swe-interview-prep), 1 drifted (rolepatch, carrying ranges instead of pins), 19 hand-rolled raw HTTP, 2 exceptions (free-ai is the gateway itself; posttrainllm runs in-browser), 31 not applicable. Ten projects call a provider API host directly rather than the gateway — including all three "compliant" ones, so adopting the SDK did not by itself move anyone onto the gateway.

    Recommendation on the table: Vercel AI SDK pinned at ai@6.0.168 / @ai-sdk/openai-compatible@2.0.41 against the gateway base URL, scoped to JS/TS, with raw fetch recorded as the standard for runtimes without JS. It is the only pattern with existing convergence, every adopter already ships on Workers, and scoping to JS/TS keeps pace (Swift) and the Python harnesses from becoming permanent exceptions. The rejected alternative was a published in-house wrapper — @sass-maker/ai-gateway is private: true and 404s on npm, and its one real benefit (hiding the gateway behind a seam) is not worth paying for while the gateway stays OpenAI-compatible.

    To ratify: set status and ratifiedAt in tooling/config/ai-client-standard.json. The validator refuses a ratified standard without a date.

    Two follow-ups noted but deliberately not done: converging the rolepatch pins, and deciding whether the committed report should name all 56 projects rather than only the 49 public ones (it currently runs --omit-private because saas-maker is public and the Site Health catalog is not; all 7 withheld are not-applicable, so no signal is lost).

  2. sarthakagrawal927 commented on Aug 29, 2026

    @sarthakagrawal927
    MemberAuthor

    Status update — ratified, now blocked on a fleet-wide migration

    The standard was ratified 2026-08-29 (#83). tooling/config/ai-client-standard.json now carries status: ratified with the decision recorded: every project calling a hosted model goes through the free-ai gateway (ai-gateway.sassmaker.com) via its OpenAI-compatible paths; in JS/TS the client is the Vercel AI SDK at the pinned canonical versions; runtimes without a JavaScript client still target the gateway and only the client library differs.

    Blocked on: migrating the non-compliant projects. Current audit (tooling/reports/ai-clients/latest.json):

    verdict count
    compliant 3 (karte, reader, swe-interview-prep)
    drifted 1 (rolepatch — ranges, not pins)
    hand-rolled 19
    exception 2 (free-ai is the gateway; posttrainllm runs in-browser)
    not-applicable 31

    The finding that reframes this issue: ten projects call provider API hosts directly rather than the gateway — including all three that count as compliant on declared packages alone. Adopting the SDK never actually moved anyone onto the gateway. Package choice was never the problem; gateway adoption is. That is what the migration has to fix.

    Drift is deliberately reported, not failed. Reddening tooling:check on 19 projects of known debt would train everyone to ignore it. Blocking findings stay narrow: a credential literal in tracked source, or a reference to a retired gateway host.

    Who can unblock: this is now a scheduling decision — it is ~29 repos and should be batched, not done in one sweep. Two follow-ups are already recorded in the config: converging rolepatch's pins, and whether the committed report should name all 56 projects rather than only the 49 public ones (it runs --omit-private because this repo is public and the Site Health catalog is not; all 7 withheld are not-applicable, so no signal is lost).

  3. sarthakagrawal927 commented on Aug 29, 2026

    @sarthakagrawal927
    MemberAuthor

    Correction: the "ten projects call provider hosts directly" figure is wrong

    An earlier comment on this issue (and the audit's framing generally) claimed that ten projects bypass the gateway by calling provider API hosts directly, "including all three that count as compliant". That is substantially incorrect, and it came from misreading the audit's own counter.

    evidence.providerHostFiles counts files containing a provider-host string, not files that make a request to one. Reading every hit:

    Project providerHostFiles What it actually is
    rolepatch 2 BYO-key settings placeholder text (settings-form.tsx:43, ai-settings.tsx:96)
    swe-interview-prep 3 a provider picker list of endpoint presets (SettingsModal.tsx:24-28)
    reader 1 placeholder text (NotesAIChat.tsx:489)
    psi-swarm 1 README/docs env-var examples
    pace 3 PaceLocalEndpointGuardTests.swift — a test asserting Pace blocks these hosts
    codevetter 9 mostly docs and a governance test fixture, plus two calls to Anthropic/OpenAI billing endpoints (observability.rs:191,292) — cost reporting, not model calls

    pace is the clearest illustration: the audit flagged as non-compliant a test whose entire purpose is proving Pace refuses non-local endpoints.

    What actually survives as a real gateway bypass

    • high-signal — scripts/auto-publish-drafts.ts:88 and .github/workflows/cron-publish.yml:36 default AI_BASE_URL to https://api.deepseek.com/v1. A genuine bypass.
    • issue-pages — src/lib/moderation.ts:28 calls https://api.openai.com/v1/moderations directly. Real, but note the gateway's declared OpenAI-compatible paths are /v1/chat/completions, /v1/responses, /v1/embeddings, /v1/messages, /v1/models — /v1/moderations is not among them, so this likely warrants a recorded exception rather than a migration, unless the gateway is extended.

    So the migration's "bypass" batch is roughly two projects, not ten. The 19 hand-rolled verdicts are unaffected by this correction — those are real, and they are about using raw fetch instead of the pinned SDK, which is a separate (and more mechanical) matter.

    Two follow-ups this exposes

    1. ai-client-audit.mjs has a precision problem. It cannot distinguish a call site from a placeholder literal, a documentation example, or a test asserting the opposite. Before anyone migrates on its output, the detector should count a provider host only when it is the target of a request, and should exclude docs, tests, and UI placeholder strings. Until then, treat providerHostFiles > 0 as "worth reading", not "non-compliant".
    2. Env-var name drift. rolepatch reads AI_BASE_URL, and high-signal does too, while the standard names AI_GATEWAY_BASE_URL as canonical.baseUrlEnv and does not list AI_BASE_URL among detection.gatewayEnvNames. That means gateway usage via AI_BASE_URL is invisible to the audit's gateway-env detection. Worth reconciling — either add it to the recognised names or rename it in the consuming repos.

    Progress

    rolepatch pins converged in Significant-Hobbies/rolepatch#66 (merged): ai ^6.0.97 → 6.0.168, @ai-sdk/openai-compatible ^2.0.41 → 2.0.41. The SDK bump was clean — no source changes, 442 tests green before and after. Its two "violations" were the placeholders above and were deliberately left alone: that field exists so a user can point at their own provider, and suggesting the fleet gateway there would be misleading.

  4. changed the title [-]Establish and audit the canonical model-calling package across every Fleet project[/-] [+]Standardize Fleet AI on direct free-model clients and retire the shared gateway[/+] on Aug 30, 2026
  5. sarthakagrawal927 commented on Sep 1, 2026

    @sarthakagrawal927
    MemberAuthor

    Execution update — 2026-09-01

    The direct free-model standard, detector language, project-owned credential boundary, and gateway retirement runbook are now recorded and pushed. The bounded Fleet scan reports no active source routing through the retired shared gateway.

    Free AI remains intentionally running: source cleanliness is not decommission proof. The remaining closure path is freeze, drain and observe traffic, revoke project credentials, delete compute only after rollback evidence, then apply the separately approved retention policy. No credentials, provider resources, DNS, data, or production bindings were changed in this run. This issue remains open for the repository-by-repository client migrations and separately authorized runtime decommission.

  6. sarthakagrawal927 commented on Sep 6, 2026

    @sarthakagrawal927
    MemberAuthor

    Blocker (2026-09-06): deferred: portfolio-cut — proceeds only if saas-maker's shared-gateway retirement work is kept in the 2026-09-06 portfolio cut.

  7. sarthakagrawal927 commented on Sep 6, 2026

    @sarthakagrawal927
    MemberAuthor

    Closing — the standard was ratified on 2026-08-30 and is fully implemented:

    • tooling/config/ai-client-standard.json — ratified standard with direct-free-model Vercel AI SDK as canonical, gateway retired
    • tooling/docs/ai-client-standard.md — full documentation with routing rules and exceptions
    • tooling/scripts/ai-client-audit.mjs — detector updated to use new terminology (retired hosts, gateway env names as blocking findings, direct provider calls as routing evidence)
    • 8 project-specific exceptions recorded (free-ai, posttrainllm, on-record, codevetter, mashup, pace, research-papers, issue-pages)
    • Audit passes with no blocking findings; remaining drift is reported as migration debt, not failures
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

deferred: portfolio-cutP3 L/XL bet that only proceeds if the product survives the 2026-09-06 portfolio cut.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions