You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Strictly and forcefully read ./RULES.md BEFORE doing anything else — and stay fully aligned with it for the entire task.
Scope: RULES.md applies to real, complex work — especially complex back-end work (features, architecture, database/migrations, security, performance, infrastructure, difficult debugging, cross-system changes, and anything where testing/verification is part of the core workflow). This is mandatory there — no skipping.
Out of scope: do NOT apply RULES.md ceremony to super-simple front-end minor edits, small code edits, typo/copy fixes, trivial styling tweaks, or other tiny low-risk changes. For those, just inspect → edit → verify.
This file is the fastest correct path into this repository. It describes the code that exists, not the code we wish existed. TODO.md holds the work plan, RULES.md holds the process, DESIGN.md owns every native UI decision, ENHANCE-DESIGN.md owns the UI polish pass, HANDOFF.md holds the previous agent's handoff, Docs/ holds deeper protocol/architecture notes.
Snapshot basis: Package.swift + all of Sources/ — 128 Swift files, ~28,650 lines. All numbers below were counted from the files, not estimated.
0. What this project is
One native macOS browser, two first-class users. Humans get an exceptionally clean, fast browser. Terminal AI agents (Codex, Claude Code, OpenCode) get a programmable execution environment with the same engine underneath.
The agent owns its sessions. The browser is an execution environment, not the agent's supervisor.
No Chromium, WebKit, Electron, Tauri, Rust, React, webview, or Cloud control plane. Apple frameworks are used where they are the correct systems primitive: Metal, CoreText/CoreGraphics, ImageIO, Foundation/Network, Swift structured concurrency.
Current honest state: Engine 0 — a real end-to-end engine (network bytes → DOM → CSS → layout → display list → pixels) with a compact-but-incomplete JS interpreter, a 76-method local agent protocol, fleet/lifecycle management, and design capture. It is not Chrome/Safari-class and is not a hardened sandbox for hostile content.
1. Engineering laws (non-negotiable)
Build the engine as a serious systems project, not a demo.
Keep the architecture modern, reliable, direct, and simpler than the problem naturally wants to become.
Do not over-engineer. Every abstraction must earn its existence through correctness, performance, isolation, or reuse.
Use modularity aggressively. Keep modules cohesive, dependencies one-directional, and public surfaces small.
Keep source code self-explanatory. Do not add explanatory source comments. Toolchain-required directives are the only exception.
Treat memory, latency, frame time, and unnecessary work as correctness concerns.
Prefer structured concurrency and value semantics where they improve clarity. Do not create actors, tasks, allocations, processes, or abstractions merely because they are available.
Keep agent control native to the engine. Do not bolt automation on through screenshots when structured engine state exists.
Never hide incomplete behavior behind fake compatibility claims. Implement, test, measure, then expand.
Preserve the existing architecture unless a measured or correctness-driven reason justifies changing it.
Additional laws specific to this repository:
One engine, two faces. Human UI and agent automation must call the same runtime. Never build a second automation path that bypasses BrowserRuntime.
No AI inference in the browser command path. Commands are deterministic. The model that decides what to do lives outside the engine.
Structured state beats pixels. Screenshots are a fallback for genuinely visual tasks, never the primary control protocol.
No UI exists yet. There is currently zero SwiftUI/AppKit code in this package (verified: no import SwiftUI, AppKit, NSWindow, or CAMetalLayer anywhere under Sources/, Tests/, Benchmarks/). Any UI work starts from scratch and must obey DESIGN.md.
2. Architecture diagrams
2.1 Target dependency graph (from Package.swift)
Dependencies are strictly one-directional. Arrows mean "imports".
agent (Codex / Claude Code / OpenCode / any process)
│ newline-delimited JSON over a Unix-domain socket
▼
browserd --socket /tmp/native-browser-engine.sock
→ AgentSocketServer.run { request → AgentCommandDispatcher.handle(request) }
→ switch on AgentMethod (76 methods)
→ BrowserRuntime public API (80 public funcs) or CaptureCoordinator
→ AgentResponse { id, result | error { code, message } }
browserctl speaks the same protocol over the same socket (79 CLI subcommands) and also has socket-free local commands: inspect, render, eval, shell, capture, bench-info.
Raw evidence from a past agent: plan/plan.md, results/*.json (RSS/CPU watchdog samples)
evidence
Sources/NativeCapture/
Nested sub-package: own Package.swift, README.md, Examples/, Sources/AetherCapture/, Tests/
vendored
.build/
SwiftPM build output (git-ignored)
generated
Trap:Sources/NativeCapture/ contains a nested Package.swift. SwiftPM ignores nested manifests; the root manifest pulls in only Sources/NativeCapture/Sources/AetherCapture via path:. Do not add files expecting the root target to pick them up outside that path.
BrowserCaptureEngine + actor NativeCaptureSession: adapts the runtime to the AetherCapture protocols
4.22 BrowserEngine (2 files, 1,017 lines — public face)
File
Role
BrowserEngine.swift
NativeBrowserEngine facade over BrowserRuntime (+ typealias BrowserEngine). Start here for the supported surface
AgentCommandDispatcher.swift
AgentCommandDispatcher.handle: the 76-method switch plus every JSON projection (pageJSON, nodeJSON, snapshotJSON, metricsJSON, …) and DispatchError (736 lines)
runAgentCapture(...) example binding for a CLI/MCP registry
../Tests/AetherCaptureTests/CaptureTests.swift
6 tests with a fake encoder/engine — fixtures only, not a renderer
4.24 browserctl (1 file, 692 lines)
Sources/browserctl/main.swift — @main BrowserControl. Local mode (inspect, render, eval, shell, capture, bench-info) plus remote mode over --socket with 79 subcommands mirroring AgentMethod. Its usage string (bottom of the file) is what agents see on a usage error — keep it in sync when adding commands.
4.25 browserd (1 file, 30 lines)
Sources/browserd/main.swift — @main BrowserDaemon. Parses --socket (default /tmp/native-browser-engine.sock), builds one NativeBrowserEngine, one AgentCommandDispatcher, one AgentSocketServer, serves until terminated. This is the entire daemon.
4.26 Benchmarks (1 file, 46 lines)
Benchmarks/enginebench/main.swift — generates 10,000 .row divs, prints nodes, parse_ms, style_ms, layout_ms, boxes. This is the only regression benchmark; extend it instead of inventing a new harness.
Host measured for every number below: macOS 27.0 (26A5425a), arm64, 8 CPUs, 8 GiB RAM, Swift 6.4, Command Line Tools only (xcode-select -p → /Library/Developer/CommandLineTools; no /Applications/Xcode*.app).
swift build -c release # engine + browserctl + browserd + enginebench
swift test --no-parallel # 138 @Test cases across 14 test targets
.build/release/browserd --socket /tmp/native-browser-engine.sock &
.build/release/browserctl --socket /tmp/native-browser-engine.sock ping
.build/release/enginebench # parse/style/layout regression baseline
Hard-won environment facts — read before "debugging" the toolchain
swift test currently fails on this host for a toolchain reason, not a code reason. With Command Line Tools only, the Swift Testing macro plugin is missing: external macro implementation type 'TestingMacros.TestDeclarationMacro' could not be found for macro 'Test'; plugin for module 'TestingMacros' not found. Fix the host, not the tests: install full Xcode and sudo xcode-select --switch /Applications/Xcode.app/Contents/Developer, or run tests on an Xcode machine. Do not migrate tests to XCTest to work around a missing plugin without a human decision.
Never run builds and tests in parallel, and never let swift test run unbounded on an 8 GiB host. A prior agent's watchdog run (evidence in agents/codex/test-freeze/results/) measured peak tree RSS ≈ 410 MiB, ≈ 89% CPU, plus heavy system swap and compressor pressure. Use nice -n 10 swift test --no-parallel --jobs 1, prefer --skip-build when the build is current, and do not start a build while a test run is active.
A fresh full release build takes minutes and streams enormous compiler command lines. Redirect and grep instead of reading output: swift build -c release > /tmp/build.log 2>&1, then search the log for error:.
Documentation/ does not exist. Docs live in Docs/ and at the repository root.
The Git repository root is /Users/harshitduggal, not this folder.Aether/ is currently untracked (git ls-files returns nothing for engine paths; git log shows unrelated home-directory commits). Do not rely on git diff, git checkout, or git log -- Aether/... to recover engine files. Copy before overwriting.
Expected benign noise: ld: warning: search path '/Library/Developer/CommandLineTools/Developer/Library/Frameworks' not found. Ignore it. Real, worth-fixing warnings exist in the JavaScript module (~55 non-linker warnings: unused/never-mutated variables, JSError carrying non-SendableJSValue, write-only self/runtime captures).
Known code defects found during this snapshot
Location
Defect
State
Sources/JavaScript/JSTextCodec.swift:42
writeTypedElement(...) is throws but called without try → hard compile error in TextEncoder.encodeInto
fixed in this snapshot (added try)
Sources/EngineCore/Geometry.swift
Point had no .zero while Size, EdgeInsets, StyleLength all do; SoftwareRenderer.render(..., origin: Point = .zero) therefore failed to compile and SoftwareRenderer did not conform to OffscreenRendering
fixed in this snapshot (added public static let zero = Point())
Whole tree
Because of the two defects above, the tree did not compile as it stood. Treat pre-existing "build green / 27 tests passed" claims in Docs/VALIDATION.md as stale.
needs a clean build + full test run on an Xcode host
7. Non-negotiable contracts
BrowserRuntime is the only owner of browser state. UI and agents both go through it.
NodeID is (index, generation). Never store only the index; never resurrect a removed node.
Structured inspection is the control surface; pixels are the fallback.
Every agent method returns AgentResponse with either result or error { code, message }. Keep error codes stable.
EngineCore stays a leaf. Adding a dependency to it is an architecture change, not a convenience.
AgentProtocol must never import engine state types; it models transport only.
AetherCapture stays engine-agnostic; adapt it through EnginePort.swift, never by importing EngineRuntime into it.
No explanatory source comments (repo law). Names and types carry the meaning.
Keep public surfaces small: adding a capability is BrowserRuntime + one DTO + one dispatcher case + one AgentMethod + one CLI subcommand + one test, not a new type hierarchy.
8. Change-routing table
I want to…
Touch these
Add an agent capability
EngineRuntime/BrowserRuntime.swift → RuntimeTypes.swift (if a new DTO) → BrowserEngine/BrowserEngine.swift → BrowserEngine/AgentCommandDispatcher.swift → AgentProtocol/AgentMessages.swift (new AgentMethod) → browserctl/main.swift (subcommand + usage text) → Tests/AgentTests/
Add a CLI-only command
browserctl/main.swift, if it needs no protocol round-trip
Fix HTML parsing
HTML/HTMLTreeBuilder.swift + tokenizer; prove it in Tests/HTMLTests/ with a malformed fixture
Fix CSS or style
CSS/ → Style/StyleResolver.swift; prove it in Tests/CSSTests/
Fix layout
Layout/LayoutEngine.swift; prove it in Tests/LayoutTests/
Fix painting/rendering
Display/DisplayListBuilder.swift → Graphics/SoftwareRenderer.swift; prove it in Tests/RenderingTests/
Add JS semantics or builtins
the matching JavaScript/JSBuiltins*.swift or JSRuntime.swift; prove it in Tests/JavaScriptTests/
Improve performance
measure first with enginebench + FrameRecorder/EngineMetrics, then Layout/PipelineInvalidation.swift, Display/PaintChunks.swift, Graphics/RasterTiles.swift, Graphics/DamageCulling.swift
Change lifecycle/fleet policy
Scheduler/FleetScheduler.swift + BrowserRuntime.sweepFleet; prove it in Tests/AgentTests/AgentFleetTests.swift
Change persistence
Persistence/ProfileStore.swift → SQLiteStore.swift; prove it in Tests/PersistenceTests/
Change capture output
NativeCapture/Sources/AetherCapture/ (plus Types.swift for the manifest) and EngineRuntime/CaptureSessionAdapter.swift; prove it in the nested CaptureTests.swift
Build the human UI
Nothing exists yet. Follow DESIGN.md, then ENHANCE-DESIGN.md. Add a new app target; never put UI inside engine modules
Add a test
Swift Testing (import Testing, @Test, #expect) — that is the convention in all 138 existing tests
9. Doc map and staleness warnings
File
Authority
Staleness
AGENTS.md (this file)
operating contract + code snapshot
current
TODO.md
the work plan
current
RULES.md
process rigor
current
DESIGN.md
all native UI/visual decisions (tokens, palette, motion, bans)
current; UI not implemented
ENHANCE-DESIGN.md
UI polish pass (run only after the UI works)
current; UI not implemented
Docs/ARCHITECTURE.md
pipeline + module layering
accurate, but predates runtime/fleet/capture work
Docs/AGENT_PROTOCOL.md
protocol overview
stale: documents ~25 methods; 76 exist. AgentMessages.swift is authoritative
Docs/IMPLEMENTATION_STATUS.md
capability status table
mostly accurate
Docs/ROADMAP.md
Engine 0/1/2/3 phases
accurate as intent
Docs/VALIDATION.md
past validation run
stale: claims 27 tests passed, a Linux host, and a green release build; the current tree has 138 tests, runs on macOS, and had two compile errors fixed today
Docs/PERFORMANCE.md
optimization order + rules
accurate as intent
HANDOFF.md
previous agent's handoff, 11-item priority list
accurate; still the best ordered engineering backlog
README.md
public framing + CLI examples
accurate
SECURITY.md
security boundary
accurate: no sandbox, no process isolation
Sources/NativeCapture/README.md
capture module contract
accurate
When you change behavior, update the doc that owns it. When you learn a non-obvious repository fact (toolchain trap, surprising API behavior, module gotcha), write it down here or in Docs/ instead of leaving it in a chat log. That is the difference between this repo getting faster for the next agent and getting slower.
About
AI-native browser built for agents to do real professional work like humans. ( Currently under construction & Not ready for PR yet) )