Skip to content

providers: hardcoded 120s ResponseHeaderTimeout still too short for slow local Ollama (follow-up to #349) #1038

Description

@FabioLeitao

Context

PR #349 (thank you for that fix) raised the shared transport's ResponseHeaderTimeout from 60s
to 120s specifically to tolerate a slow cloud proxy (ollama *:cloud) that withholds its 200
response header until the upstream model emits a first token, without wrongly aborting an
otherwise-alive request.

The gap

A throttled local Ollama deployment (low num_gpu/num_thread on constrained hardware, not
a cloud proxy) can see real time-to-first-token of 1-5 minutes on a cold model load. That still
exceeds the 120s ceiling PR #349 established, on a request that is alive, not hung — the same
class of problem #349 fixed, just a slower real-world case than the cloud-proxy scenario it was
tuned for.

Proposal

Rather than raising the shared default again for every deployment (which would just shift the
same trade-off), make it configurable via an opt-in env var, ZERO_RESPONSE_HEADER_TIMEOUT
(same accepted forms as the existing ZERO_STREAM_IDLE_TIMEOUT: a Go duration string or bare
seconds). The default stays exactly 120s — nothing changes for anyone who doesn't set it.

Validated locally against a throttled local Ollama deployment:

  • ZERO_RESPONSE_HEADER_TIMEOUT=5s fails at ~6.4s (expected — proves the override takes effect)
  • ZERO_RESPONSE_HEADER_TIMEOUT=300s succeeds at ~223s (a request that would have hit the old
    120s ceiling)

I have a patch + regression test ready locally (mirrors the existing
TestResolveStreamIdleTimeout pattern — fmt-check/vet/package tests all green). Per
AGENTS.md, I understand a community PR needs an issue-approved label first — happy to open
the PR as soon as this is approved.

Activity

  1. FabioLeitao commented on Sep 9, 2026

    @FabioLeitao
    Author

    Discovered, reviewed, suggested and tested by human (HITL) in bare metal running LMDE (Linux Mint Debian Edition) - @FabioLeitao

    -----BEGIN SSH SIGNATURE-----
    U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgHSwaWVCd3rALjtmwINNtVKRX3t
    ucF7TfXNZFALyRo4EAAAAEZmlsZQAAAAAAAAAGc2hhNTEyAAAAUwAAAAtzc2gtZWQyNTUx
    OQAAAECebXFbE0OASoBEfQb9xSuNpvR3fio1n0EP5627jHzim7H7o04fVd0sf4UGlisH2Z
    7Qcq0evFfQstiDx+7SeyMF
    -----END SSH SIGNATURE-----
    
  2. Vasanthdev2004 commented on Sep 12, 2026

    @Vasanthdev2004
    Collaborator

    Approving. The 120s ceiling in providerio.go is one hardcoded number on the shared transport, and the repo already has the exact pattern you propose for the stream idle timeout: ResolveStreamIdleTimeout reads ZERO_STREAM_IDLE_TIMEOUT, accepts a Go duration or bare seconds, and treats 0, off, none and disabled as unlimited. Mirror that for ZERO_RESPONSE_HEADER_TIMEOUT with the default unchanged at 120s, document it next to the idle timeout, and cover the parsing with a table test the way the idle one is covered. Open the PR whenever you are ready; the patch you describe is the right size.

  3. added
    issue-approvedReviewed and approved by the core team; community PRs may implement this issue.
    on Sep 12, 2026
  4. Vasanthdev2004 commented on Sep 26, 2026

    @Vasanthdev2004
    Collaborator

    @FabioLeitao checking in, since it's been two weeks: are you still planning to open this? If you'd rather not, say so and it goes in the pool with the approval already on it.

  5. FabioLeitao commented on Sep 29, 2026

    @FabioLeitao
    Author

    @Vasanthdev2004 yes, still planning it, sorry for the delay. PR is up: #1106. It follows your outline: mirrors ZERO_STREAM_IDLE_TIMEOUT (Go duration or bare seconds), 0/off/none/disabled remove the limit, an unparseable value falls back to the 120s default, the default is unchanged, documented next to the idle-timeout constant, with a table test.

    One note for review: go test ./... is not fully green on my machine, but the same 8 packages fail on an unmodified upstream/main archive (listed in the PR description); ./internal/providers/... passes, including -race on providerio.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    issue-approvedReviewed and approved by the core team; community PRs may implement this issue.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions