Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,7 @@ Local run (`launch --new`):
- **Skip checklist box (`F01.onboarding-skip-checklist`).** Run `control-openhands browser goto /conversations` and `control-openhands browser count 'testid=sidebar-onboarding-checklist'` (`1`). Run `control-openhands browser goto '/?previewOnboardingStep=3'`, then `control-openhands browser click 'text=Skip Getting Started checklist'`. `control-openhands browser eval "document.querySelector('[data-testid=onboarding-skip-getting-started-checklist]').checked"` is `true`. Run `control-openhands browser goto /conversations` and `control-openhands browser reload`; the checklist count is `0`. On `/settings/app`, `control-openhands browser eval "document.querySelector('[data-testid=show-getting-started-checklist-switch]').checked"` is `false`. Restore it: go back to `/?previewOnboardingStep=3`, where the box now reads `true`, click the same text again, and the checklist count on `/conversations` is `1`.
- **Phone (`F01.onboarding-phone`).** Run `control-openhands browser viewport phone`. Then for N in 1, 2 and 3 run `control-openhands browser goto '/?previewOnboardingStep=N'`, `control-openhands browser bbox 'testid=onboarding-modal'` and `control-openhands browser screenshot --feature F01.onboarding-phone --name step-N`. The modal is 351 px wide with `insideViewport` `true` and `pageHorizontalOverflow` `false`, and Back/Next or Back/Close are visible. Return with `control-openhands browser viewport desktop`.
- **Error page (`F01.route-error-boundary`).** Run `control-openhands browser errors --clear`, `control-openhands browser goto /this-route-does-not-exist` and `control-openhands browser snapshot 'testid=not-found-screen'`. It shows `heading "Page not found"`, the paragraph `This address does not match any page. Check the URL, or go back to the home page.` and `link "Home"` (`/url: /`). `control-openhands browser count 'aside[data-collapsed]'` is `1`: the sidebar stays. `control-openhands browser screenshot --feature F01.route-error-boundary --name not-found` shows the message and the Home button centered beside the sidebar. `control-openhands browser errors --app-only` reports `pageErrors` `0` and `appErrors` `0`. Run `control-openhands browser click 'testid=not-found-home-link' --expect-url '/$'`; `control-openhands browser count 'testid=home-screen'` is `1`.
- **Onboarding again with the saved endpoint (`F01.onboarding-repeat-endpoint`).** Two more first runs on this stack, both through the All view, which shows the Base URL field. Pass 1 saves a custom endpoint: `control-openhands browser reset`, `control-openhands browser goto /`, `control-openhands browser click 'testid=onboarding-agent-next'`, `control-openhands browser wait '[data-testid=onboarding-modal][data-current-step="1"]'`, `control-openhands browser click 'testid=onboarding-step-setup-llm >> testid=sdk-section-all-toggle'` and `control-openhands browser value 'testid=onboarding-step-setup-llm >> testid=base-url-input'` (empty). Run `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=llm-custom-model-input' openai/deepseek-chat`, `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=base-url-input' https://api.deepseek.com/v1`, `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=llm-api-key-input' --value-env DEEPSEEK_API_KEY`, `control-openhands browser click 'testid=onboarding-llm-next'` and `control-openhands browser wait '[data-testid=onboarding-modal][data-current-step="2"]'`. `control-openhands llm show` lists `deepseek-chat` (`openai/deepseek-chat`, `base_url` `https://api.deepseek.com/v1`) as `active_profile`. Say hello as above (`browser fill 'testid=onboarding-hello-input' 'Reply with only the word hello. Do not run any tools.'`, `browser press Enter --selector 'testid=onboarding-hello-input'`, `browser wait-url '/conversations/[0-9a-f-]+'`), then `control-openhands conversation wait <id> --timeout 180` is `finished` and the agent `MessageEvent` is `hello`. Pass 2 types the same endpoint again: repeat the steps up to the All view. `browser value 'testid=onboarding-step-setup-llm >> testid=base-url-input'` is now `https://api.deepseek.com/v1`, the saved value. Fill the model `openai/deepseek-v4-flash` (a DeepSeek alias, so the new profile gets its own name), the same Base URL and the key, then Next and wait for step 2 (`control-openhands browser screenshot 'testid=onboarding-step-setup-llm' --feature F01.onboarding-repeat-endpoint --name second-pass-form` before Next). Expected: `control-openhands llm show` lists `deepseek-v4-flash` with `base_url` `https://api.deepseek.com/v1`, and the say-hello conversation finishes with `hello`. Known failure (reproduced 2026-10-08 at `53c8b4d`): the profile is saved with `base_url` `null`. The say-hello conversation then ends in `error`, and `control-openhands conversation events <id> --last 10` shows `ConversationErrorEvent` `LLMAuthenticationError: litellm.AuthenticationError: ... OpenAIException - Incorrect API key provided`: the DeepSeek key went to OpenAI. The chat shows `Your LLM API key appears to be invalid or has expired.` (`browser screenshot --feature F01.onboarding-repeat-endpoint --name second-pass-error`). Issue #17884, fix in #17889. Restore: run `control-openhands llm preset deepseek`, then point the `default` agent profile back at it. Run `control-openhands browser goto /settings/agents`, `control-openhands browser click 'testid=agent-profile-row >> has-text=default >> testid=agent-profile-menu-trigger'` and `control-openhands browser click 'testid=agent-profile-actions-menu >> text=Edit'`. Then run `control-openhands browser choose 'testid=agent-profile-llm-selector' 'deepseek-flash (deepseek/deepseek-flash)'` and `control-openhands browser click 'testid=save-agent-profile-btn'`; the toast reads `Profile "default" updated`. Only then do `control-openhands api DELETE /api/profiles/deepseek-v4-flash --write` and `.../deepseek-chat --write` answer `200`. Delete the two say-hello conversations with Clean up's steps.
- **Onboarding again with the saved endpoint (`F01.onboarding-repeat-endpoint`).** Two more first runs on this stack, both through the All view, which shows the Base URL field. Pass 1 saves a custom endpoint: `control-openhands browser reset`, `control-openhands browser goto /`, `control-openhands browser click 'testid=onboarding-agent-next'`, `control-openhands browser wait '[data-testid=onboarding-modal][data-current-step="1"]'`, `control-openhands browser click 'testid=onboarding-step-setup-llm >> testid=sdk-section-all-toggle'` and `control-openhands browser value 'testid=onboarding-step-setup-llm >> testid=base-url-input'` (empty). Run `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=llm-custom-model-input' openai/deepseek-chat`, `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=base-url-input' https://api.deepseek.com/v1`, `control-openhands browser fill 'testid=onboarding-step-setup-llm >> testid=llm-api-key-input' --value-env DEEPSEEK_API_KEY`, `control-openhands browser click 'testid=onboarding-llm-next'` and `control-openhands browser wait '[data-testid=onboarding-modal][data-current-step="2"]'`. `control-openhands llm show` lists `deepseek-chat` (`openai/deepseek-chat`, `base_url` `https://api.deepseek.com/v1`) as `active_profile`. Say hello as above (`browser fill 'testid=onboarding-hello-input' 'Reply with only the word hello. Do not run any tools.'`, `browser press Enter --selector 'testid=onboarding-hello-input'`, `browser wait-url '/conversations/[0-9a-f-]+'`), then `control-openhands conversation wait <id> --timeout 180` is `finished` and the agent `MessageEvent` is `hello`. Pass 2 types the same endpoint again: repeat the steps up to the All view. `browser value 'testid=onboarding-step-setup-llm >> testid=base-url-input'` is now `https://api.deepseek.com/v1`, the saved value. Fill the model `openai/deepseek-v4-flash` (a DeepSeek alias, so the new profile gets its own name), the same Base URL and the key, then Next and wait for step 2 (`control-openhands browser screenshot 'testid=onboarding-step-setup-llm' --feature F01.onboarding-repeat-endpoint --name second-pass-form` before Next). Expected: `control-openhands llm show` lists `deepseek-v4-flash` with `base_url` `https://api.deepseek.com/v1`, and the say-hello conversation finishes with `hello`. Known failure (reproduced 2026-10-08 at `53c8b4d`): the profile is saved with `base_url` `null`. The say-hello conversation then ends in `error`, and `control-openhands conversation events <id> --last 10` shows `ConversationErrorEvent` `LLMAuthenticationError: litellm.AuthenticationError: ... OpenAIException - Incorrect API key provided`: the DeepSeek key went to OpenAI. The chat shows `Your LLM API key appears to be invalid or has expired.` (`browser screenshot --feature F01.onboarding-repeat-endpoint --name second-pass-error`). Issue #17884, fix in #17889. Restore: run `control-openhands llm preset deepseek`. It also points the `default` agent profile back at `deepseek-flash`: its output has `repointed` from `deepseek-v4-flash` to `deepseek-flash`. Then `control-openhands api DELETE /api/profiles/deepseek-v4-flash --write` and `.../deepseek-chat --write` answer `200`. Delete the two say-hello conversations with Clean up's steps.
- **Clean up.** Delete the hello conversation from its header menu: `control-openhands browser goto /conversations/<id>` (the hello conversation's `<id>` from Say hello; the Error page bullet left `/`), `control-openhands browser click 'testid=chat-pane-header >> testid=ellipsis-button'`, `control-openhands browser click 'testid=conversation-name-context-menu >> testid=delete-button'`, `control-openhands browser click 'role=button[name="Confirm Delete"]'`. `control-openhands conversation list` then shows `count` `0`.

Public run A (`launch --new --public`), backend step walked:
Expand Down Expand Up @@ -102,7 +102,7 @@ Not reachable locally:
- Slide indices renumber. After the public backend step succeeds, "Choose your agent" becomes index `0` and has no Back. `previewOnboardingStep` counts phases with the backend step included, so `0` and `1` look identical on a healthy backend.
- The consent modal (z-70) sits over the onboarding modal, and clicks at the screen center land on agent tiles once it closes. A forced click on `first-run-onboarding-screen` selected Codex. Check backdrop non-dismissal with `browser mouse-click 40 40`, a corner outside the modal.
- The say-hello default message ("Create a basic webpage…") makes the agent work for a while. Replace it with a tiny prompt to keep model cost down.
- On a local backend, Next on the LLM step also points the `default` agent profile at the profile it just created. Running `llm preset deepseek` later does not move that pointer back. Deleting the onboarding profile then answers `409` `LLM profile is referenced by 1 agent profile(s): default`. Re-point it first in Settings → Agents (the Restore steps of `F01.onboarding-repeat-endpoint`).
- On a local backend, Next on the LLM step also points the `default` agent profile at the profile it just created. Activating another LLM profile does not move that pointer, and while it stays, deleting the onboarding profile answers `409` `LLM profile is referenced by 1 agent profile(s): default`. `control-openhands llm preset deepseek` moves it to `deepseek-flash` and reports `repointed`; `llm set` does not. In the UI, edit `default` in Settings → Agents and pick the profile in `testid=agent-profile-llm-selector`.
- The LLM step saves only the fields that differ from the saved settings, and the profile it creates is built from that same diff. That is the cause of the `F01.onboarding-repeat-endpoint` failure (#17884, fix in #17889). It is also why the mock-LLM onboarding e2e failed after earlier specs had saved the mock endpoint. The first pass of that bullet passes on a run whose saved Base URL is still empty.
- The tab title prefixes a status emoji to the stored title. A generated title that already starts with an emoji therefore shows twice (`✅ ✅ Reply hello without tools | OpenHands`). Right after launch the title is `Conversation <id prefix>` until a reload picks up the generated one.
- Right after Next, the ACP step reads "Checking for an existing Claude Code login…" with Next disabled. Wait for `testid=onboarding-acp-auth-detected` before reading the banner. The "already signed in" banner depends on the host's CLI login. On a clean CI machine the fields become required, and Next is blocked until they are filled.
Expand Down
39 changes: 38 additions & 1 deletion .agents/skills/verify-openhands/scripts/control-openhands.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,10 @@ import {
} from "./lib/testids.mjs";
import { tmuxPathFor } from "./lib/tmux-path.mjs";
import { browserCallLimit } from "./lib/call-limit.mjs";
import {
DEFAULT_AGENT_PROFILE,
agentProfileRepoint,
} from "./lib/agent-profile-repoint.mjs";
import {
PAGE_LIMIT,
collectEvents,
Expand Down Expand Up @@ -1492,6 +1496,30 @@ async function activateProfile(run, name) {
);
}

// Onboarding pins the `default` agent profile to the LLM profile it created;
// `llm preset` moves it to the one it activates (see
// lib/agent-profile-repoint.mjs). `llm set` leaves it: its throwaway profiles
// would otherwise become `default`'s and refuse deletion.
async function repointDefaultAgentProfile(run, target) {
const path = `/api/agent-profiles/${DEFAULT_AGENT_PROFILE}`;
const detail = await http(run, "GET", path);
if (detail.status === 404) return undefined;
if (!detail.ok)
throw new CliError(
`Reading agent profile ${DEFAULT_AGENT_PROFILE} failed: ${detail.status} ${detail.text.slice(0, 300)}`,
);
const llm = await http(run, "GET", "/api/profiles");
const names = (llm.json?.profiles ?? []).map((p) => p.name);
const plan = agentProfileRepoint(detail.json?.profile, names, target);
if (!plan) return undefined;
const saved = await http(run, "POST", path, { body: plan.body });
if (!saved.ok)
throw new CliError(
`Pointing agent profile ${DEFAULT_AGENT_PROFILE} at ${target} failed: ${saved.status} ${saved.text.slice(0, 300)}`,
);
return { agentProfile: DEFAULT_AGENT_PROFILE, from: plan.from, to: target };
}

const PRESETS = {
deepseek: {
envVar: "DEEPSEEK_API_KEY",
Expand Down Expand Up @@ -1571,14 +1599,19 @@ async function cmdLlm({ positional, flags }) {
created.push(profile.name);
}
const active = preset.profiles.find((p) => p.activate);
if (active) await activateProfile(run, active.name);
let repointed;
if (active) {
await activateProfile(run, active.name);
repointed = await repointDefaultAgentProfile(run, active.name);
}
const settings = await http(run, "GET", "/api/settings");
out({
ok: true,
preset: presetName,
profiles: created,
active: active?.name,
activeModel: settings.json?.agent_settings?.llm?.model,
...(repointed ? { repointed } : {}),
});
return;
}
Expand Down Expand Up @@ -4072,6 +4105,10 @@ set/preset validate with a 1-token completion first (skip with --no-validate).
Keys are read from an environment variable or file, never from argv.
'preset deepseek' saves deepseek-flash (deepseek/deepseek-flash, activated) and
deepseek-pro (deepseek/deepseek-v4-pro). Prefer flash; it is cheaper.
'preset' also points the 'default' agent profile at deepseek-flash when it
references another LLM profile that exists, as onboarding leaves it; the output
then has 'repointed'. A reference to a missing profile (a fresh run's seed),
named agent profiles and 'set' leave agent profiles as they are.

Examples:
DEEPSEEK_API_KEY=... control-openhands llm preset deepseek
Expand Down
67 changes: 67 additions & 0 deletions .agents/skills/verify-openhands/scripts/control-openhands.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ import { routePattern } from "./lib/route-pattern.mjs";
import { resolveTestids } from "./lib/testids.mjs";
import { tmuxPathFor } from "./lib/tmux-path.mjs";
import { browserCallLimit } from "./lib/call-limit.mjs";
import { agentProfileRepoint } from "./lib/agent-profile-repoint.mjs";
import {
canvasSize,
encodeRecording,
Expand Down Expand Up @@ -1444,6 +1445,72 @@ test("fixture skill --commit records the project skill once and reports a re-run
assert.match(personal.json.error, /needs --repo/);
});

test("llm preset repoints default only off another live LLM profile", () => {
const seeded = {
id: "a1",
name: "default",
revision: 3,
schema_version: 1,
agent_kind: "openhands",
llm_profile_ref: "deepseek-chat",
enable_sub_agents: true,
mcp_server_refs: ["qa_mcp_a"],
};
const names = ["deepseek-chat", "deepseek-flash", "deepseek-pro"];
// Onboarding left default on its own profile: move it, keep every other field.
assert.deepEqual(agentProfileRepoint(seeded, names, "deepseek-flash"), {
from: "deepseek-chat",
body: {
schema_version: 1,
agent_kind: "openhands",
llm_profile_ref: "deepseek-flash",
enable_sub_agents: true,
mcp_server_refs: ["qa_mcp_a"],
},
});
// Already there, a fresh run's missing `default` ref, or no ref: unchanged.
assert.equal(agentProfileRepoint(seeded, names, "deepseek-chat"), null);
assert.equal(
agentProfileRepoint(
{ ...seeded, llm_profile_ref: "default" },
names,
"deepseek-flash",
),
null,
);
assert.equal(
agentProfileRepoint(
{ ...seeded, llm_profile_ref: null },
names,
"deepseek-flash",
),
null,
);
// An ACP default owns its own model; an older backend has no profile.
assert.equal(
agentProfileRepoint(
{ ...seeded, agent_kind: "acp" },
names,
"deepseek-flash",
),
null,
);
assert.equal(agentProfileRepoint(undefined, names, "deepseek-flash"), null);
});

test("llm help says when preset repoints the default agent profile", () => {
const help = spawnSync(process.execPath, [cli, "help", "llm"], {
encoding: "utf8",
});
assert.equal(help.status, 0);
assert.match(
help.stdout,
/'preset' also points the 'default' agent profile at deepseek-flash/,
);
assert.match(help.stdout, /and 'set' leave agent profiles as they are/);
assert.match(help.stdout, /'repointed'/);
});

test("a recording's timeline cuts the lead-in and pauses, and holds the end", () => {
// Changes at 0, 3, 3.2 and 10 s; stopped at 10.5 s.
const frames = [{ t: 0 }, { t: 3000 }, { t: 3200 }, { t: 10000 }];
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
// Whether `llm preset` must also move the seeded `default` agent profile onto
// the LLM profile it activates.
//
// On a local backend, onboarding points `default` at the LLM profile it
// creates (useApplyOnboardingAgentProfile). Activating another LLM profile
// leaves that reference behind. The Agents page then still shows the
// onboarding model on `default`, a launch that picks `default` explicitly
// uses it, and deleting the onboarding profile answers 409 ("referenced by
// 1 agent profile(s)").
//
// Only `default` moves, and only when it is an OpenHands profile whose
// reference names another LLM profile that exists. Named agent profiles are
// deliberate picks. `llm set` never moves it: recipes activate throwaway
// profiles with it and delete them afterwards. A reference to a missing
// profile is the fresh-run seed: the app already falls back to the active LLM
// profile for it, and F13.stale-llm-ref drives that state.

export const DEFAULT_AGENT_PROFILE = "default";

/**
* @param {object | undefined} agentProfile `profile` from
* `GET /api/agent-profiles/default`
* @param {string[]} llmProfileNames names from `GET /api/profiles`
* @param {string} target the LLM profile just activated
* @returns {{ from: string, body: object } | null} the save body for
* `POST /api/agent-profiles/default`, or null when nothing should move
*/
export function agentProfileRepoint(agentProfile, llmProfileNames, target) {
if (agentProfile?.agent_kind !== "openhands") return null;
const from = agentProfile.llm_profile_ref;
if (!from || from === target || !llmProfileNames.includes(from)) return null;
// The save is a whole-profile overwrite: keep every stored field and drop
// only the identity the server owns (as the Agents editor does).
const { id, name, revision, ...stored } = agentProfile;
return { from, body: { ...stored, llm_profile_ref: target } };
}
Loading