Skip to content

Android device video never starts when the server runs without XDG_RUNTIME_DIR #12313

Description

@mchisolm0

What happened

"I've noticed for the past few days that device hub doesn't always work. When I try to connect a simulator it will just hang and never actually show. Sometimes it works. It specifically hangs on connecting video."

The user also saw [apple-utils] Failed to run \xcrun simctl list devices --json`` at the top of the T3 Code window on Linux and reasonably assumed that was the cause.

Diagnosis

The xcrun message is a red herring. expo-device-hub enumerates iOS simulators unconditionally on every platform, and on Linux the spawn xcrun ENOENT is collected into the errors[] array of /api/devices and never touches the Android path. Worth suppressing on non-darwin for noise reasons, but unrelated to this bug.

The real cause is bearer-token discovery for the emulator's gRPC endpoint.

The Android emulator secures its gRPC endpoint and publishes the token in a per-run file at $XDG_RUNTIME_DIR/avd/running/pid_<pid>.ini (grpc.port, grpc.token). In expo-device-hub@0.9.0, vendor/serve-emu/dist/emulator-grpc.js:20, discoveryDirs() builds its candidate list like this:

const dirs = [join(home, "Library", "Caches", "TemporaryItems", "avd", "running")];
if (process.env.XDG_RUNTIME_DIR) dirs.push(join(process.env.XDG_RUNTIME_DIR, "avd", "running"));
if (process.env.LOCALAPPDATA)    dirs.push(join(process.env.LOCALAPPDATA, "Temp", "avd", "running"));
dirs.push(join(home, ".android", "avd", "running"));

On Linux the only unconditional candidate is ~/.android/avd/running, which the emulator does not write to. So if XDG_RUNTIME_DIR is absent from the hub process environment, discovery finds nothing. It falls back to adb -s <serial> emu grpc <port>, retried 5 times with 40 probes each (this is the visible hang), and at probe === 39 returns { port: activePort, token: null }. useEndpoint() only warns when the token is null, and the authorization header is only set when a token exists, so every subsequent call fails grpc-status 16 UNAUTHENTICATED.

T3 Code's contribution is that it does not pass XDG_RUNTIME_DIR down. apps/server/src/device/LocalDeviceHost.ts:232 (v0.0.42; :245 on main, unchanged as of 2026-09-17):

const hubEnvironment = (): NodeJS.ProcessEnv => ({
  ...hostEnvironment,
  FORCE_COLOR: "0",
  NO_COLOR: "1",
});

hostEnvironment is inherited from the server process. deviceHostEnvironment() adds only ANDROID_HOME and PATH entries. So the hub inherits whatever the server had. A server started from a desktop session has XDG_RUNTIME_DIR; one started over SSH typically does not.

This explains "sometimes it works." On this machine two servers were running: one from the desktop app (has XDG_RUNTIME_DIR) and one t3 serve launched over SSH (does not). The device hub is shared and reaped through device/agent-device/hub.json, so whichever server spawns the hub first silently decides whether device video works for the whole machine.

Three secondary problems make this hard to see:

  1. DeviceHubProxy.proxyWebSocket in apps/server/src/device/DeviceHubProxy.ts ends with .pipe(Effect.catchCause(() => Effect.void)), so the upstream failure is swallowed. Every one of 199 logged proxy requests exits Success in server.trace.ndjson while the panel is visibly broken. Nothing surfaces in the UI or the trace.
  2. serve-emu treats a null token as a warning rather than an error, so the failure only appears later as a generic grpc-status 16 from getScreenshot.
  3. More generally, the device path drops causes on the way to the user. Separately from this bug, the same session hit an unreachable SSH device host, and the server had the exact reason in hand (ssh: connect to host <host> port 22: No route to host, later Operation timed out) while the user was shown only The device command failed (exit code 255). and then Could not communicate with device support. Try refreshing devices. The real stderr was sitting in server.trace.ndjson the whole time. Surfacing the underlying SshCommandError message would have made that self-service.

Steps to reproduce

  1. On Linux, start an Android emulator.
  2. Start the server from an environment with no XDG_RUNTIME_DIR (an SSH session does this by default): env -u XDG_RUNTIME_DIR npx t3 serve --port 3775.
  3. Make sure that server is the one that spawns the device hub (delete ~/.t3/userdata/device/agent-device/hub.json first if a hub is already running).
  4. Open the device panel and select the Android device.
  5. It hangs on "Connecting video" indefinitely. Confirm with curl 'http://127.0.0.1:<hubPort>/vendor/serve-emu/health?device=emulator-5554', which returns getScreenshot: grpc-status 16.

Starting the same server with XDG_RUNTIME_DIR set makes it work.

Version

0.0.42 (runtime 0.0.43-nightly.20260917.1866), expo-device-hub 0.9.0, agent-device 0.20.10

Environment

Linux x64, Nobara 7.2.3 (7.2.3-200.nobara.fc44.x86_64), Node v26.8.2, emulator 37.1.11.0, opencode as the agent CLI. Desktop app on macOS connecting to this Linux server, which also has a remote SSH device host configured.

Evidence

# hub health while broken
{"ok":false,"error":"getScreenshot: grpc-status 16 (Missing the 'authorization' header with security credentials.)"}
{"ok":false,"error":"serve-emu start for emulator-5554 is cooling down after a failure"}

# proven on the wire, token from the discovery ini, values redacted
POST /android.emulation.control.EmulatorController/getStatus
  no authorization header           -> grpc-status: 16 "Missing the 'authorization' header with security credentials."
  Bearer <token from pid_<pid>.ini> -> grpc-status: 0

# discovery file the hub never reads without XDG_RUNTIME_DIR
$XDG_RUNTIME_DIR/avd/running/pid_<pid>.ini
  grpc.port=8554
  grpc.token=<REDACTED>
  avd.id=Medium_Phone

# hub process environment
XDG_RUNTIME_DIR absent; SSH_CLIENT and SSH_CONNECTION present

# server.trace.ndjson: 199 proxy requests, every one Success, while the panel is broken
GET /api/device-hub/vendor/serve-emu/ws?device=emulator-5554&frame-meta=1&wsTicket=<REDACTED>&hostId=local
  exit._tag: "Success"   durationMs: 27..2519

# point 3: cause present in the trace, absent from the UI
DeviceService.fetchDevices  Failure
  DeviceOperationError: Device list failed: The device command failed (exit code 255).
ssh/command.runSshCommand   Failure
  SshCommandError: ssh: connect to host <host> port 22: No route to host

Related issues

#12229, same user-visible symptom ("Connecting video") but a different bug. There the device streams and is interactive and the failure is a keyframe/decoder restart deadlock. Here getScreenshot fails outright and serve-emu never starts. #10810 (Honor XDG Base Directory paths on Linux) is adjacent but about config paths, not process environment inheritance.

Fix applied or workaround

Symlinked the unconditional discovery directory to the real one, which works regardless of how the server was launched and needs no restart:

ln -s "$XDG_RUNTIME_DIR/avd/running" ~/.android/avd/running

Both devices went to status: "streaming" immediately.

Suggested real fixes, in the order I'd rank them:

  1. hubEnvironment() should pass XDG_RUNTIME_DIR through explicitly, and fall back to /run/user/$(id -u) on Linux when the server's own environment lacks it.
  2. serve-emu should treat a null token on Linux as a hard error rather than a warning, so the failure names itself.
  3. Surface the underlying cause in the device path rather than discarding it: log the swallowed cause in proxyWebSocket, and include the SshCommandError message instead of a bare exit code.
  4. Skip the xcrun probe on non-darwin.

Filed by

Claude Opus 5 via Claude Code, through t3 triage

Activity

  1. mchisolm0 commented on Sep 17, 2026

    @mchisolm0
    Author

    Posting as myself for context. Apparently Opus tried/failed labeling the issue with via-triage and got blocked because I don't have write permissions. Just surfacing that as possible friction.

    Also, this is the state I was describing the UI getting stuck in.

    Image
  2. juliusmarminge commented on Sep 17, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on current main (4749035bd) and on v0.0.42. This is a real Linux device-hub bug, not the xcrun banner and not a duplicate of #12229. Not already fixed. Pinned tools still match the report: expo-device-hub@0.9.0, agent-device@0.20.10.

    What happens

    The Android emulator publishes grpc.port / grpc.token under $XDG_RUNTIME_DIR/avd/running/pid_<pid>.ini. In expo-device-hub@0.9.0 vendor/serve-emu/dist/emulator-grpc.js, discoveryDirs() only adds that path when process.env.XDG_RUNTIME_DIR is set. The unconditional Linux candidate is ~/.android/avd/running, which the emulator does not write. With no ini, serve-emu falls back to adb emu grpc, retries, and at probe === 39 returns { token: null }. useEndpoint() only warns; later getScreenshot fails grpc-status 16 (missing authorization). That matches the health JSON and the wire capture in this report.

    T3 starts the hub with:

      const hubEnvironment = (): NodeJS.ProcessEnv => ({
        ...hostEnvironment,
        FORCE_COLOR: "0",
        NO_COLOR: "1",
      });

    hostEnvironment is process.env plus ANDROID_HOME / PATH. There is no Linux fallback. A desktop-launched server usually has XDG_RUNTIME_DIR; t3 serve over SSH usually does not. The desktop shell already knows the fallback (linuxRuntimeDirCandidates → /run/user/${uid} in DesktopShellEnvironment.ts); the hub child does not.

    The “sometimes it works” story is right in spirit, with one correction: reapStaleHub kills a leftover hub on the same userdata and spawnHub starts a new one. On a shared ~/.t3/userdata, the last server to call ensureHubReady wins, and that process’s env decides whether video works for everyone.

    The xcrun simctl line is noise. fetchDevices joins /api/devices errors[] into hostStatusDetail, and DevicePanel shows that banner while the host is otherwise ready. expo-device-hub enumerates iOS on every platform; on Linux that is spawn xcrun ENOENT and never touches Android.

    Why this is hard to see

    DeviceHubProxy.proxyWebSocket ends with Effect.catchCause(() => Effect.void), so the 199 proxy rows in server.trace.ndjson are Success while the canvas stays on “Connecting video…”. Separately, DeviceOperationError renders command_failed / request_failed as a bare exit code or “Could not communicate with device support,” even when SshCommandError already has ssh: connect to host … No route to host in the cause. That SSH host is a different failure in the same session, not this video bug.

    Not a duplicate

    Suggested smallest scope

    In hubEnvironment(), on Linux, if XDG_RUNTIME_DIR is missing, set it to /run/user/${uid} when that directory exists (HostProcessUserId is already in hostProcess.ts). Same pattern as the desktop shell. That is enough to find the emulator’s ini when the emulator was started from a graphical session.

    Do not block on expo-device-hub. Useful follow-ups, not required to unblock: log the swallowed proxy cause; put the underlying SshCommandError / stderr in the device error string; skip xcrun on non-darwin (upstream).

    Workaround

    The report’s symlink is correct and does not need a hub restart:

    ln -s "$XDG_RUNTIME_DIR/avd/running" ~/.android/avd/running

    ~/.android/avd/running is always on the discovery list. Starting the server with XDG_RUNTIME_DIR set also works, as long as that server is the one that last spawned the hub.

  3. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 17, 2026
  4. cestercian commented on Sep 17, 2026

    @cestercian
    Contributor

    Looking into starting Android device video when the server has no XDG_RUNTIME_DIR.

  5. juliusmarminge commented on Oct 5, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Fixed by #12402 (merged in f729e0e): when the server starts on Linux without XDG_RUNTIME_DIR, the device hub now falls back to /run/user/<uid> if that directory exists and belongs to the user. Values you set yourself are left alone. That's the smallest fix suggested in the triage above, so the hub can find the emulator's gRPC token again and Android video shouldn't hang on "Connecting video" anymore.

    The other follow-ups (logging the proxy error that gets swallowed, showing SSH error details, and skipping xcrun on non-macOS) were optional and aren't part of this fix. Feel free to open a separate issue for any of them. Closing this as completed. If video still hangs on a nightly that includes #12402, please reopen.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions