Files
Nithilan Kumaranandmorluto 9eee709f0c fix(wasm): accept unpadded WABT section rows (#1963)
* fix(wasm): accept unpadded WABT section rows

WABT 1.0.42 right-aligns section names in a nine-column field. DataCount
fills that field, so its objdump row has no leading whitespace. Accept
that legal row. Header recognition, the required field shape, and section
range checks still reject malformed output.

Fixes #1956

* fix(ci): refresh MIPS expectations and Windows test paths

Assert admitted soft-float ABI metadata and its committed storage policy through the public target resolver, retain FP64 rejection, and point the curated Windows lane at the moved filesystem reconstruction tests.

Validation: reproduced all four prior failures; 97 focused tests and check:ci pass. Native Windows and real Ghidra execution were not performed locally.

---------

Co-authored-by: morluto <[email protected]>
2026-10-12 05:29:04 +08:00

90 KiB
Raw Permalink Blame History

Testing REA

Use this guide when choosing coverage, pruning tests, or running a verification lane. For commands, jump to developer commands; for engine prerequisites, use real-toolchain lanes.

Prefer real public CLI/MCP workflows, then production-boundary integration, then goldens captured from real producers. Keep module tests for distinct failures or semantics that stronger workflows cannot reliably reproduce. Paths, compiled imports, and suite names do not establish test depth.

Prune tests that mirror implementation details or duplicate a stronger claim. Before deleting a distinct regression, identify the replacement scenario and assertions. Preserve malformed formats, cancellation, cleanup, permissions, capacity, target identity, and combinations with distinct interactions. Trivial helper assertions need no replacement. Remove abandoned production scaffolding only after checking runtime and verifier consumers, including compiled imports.

Use representative scenarios rather than Cartesian matrices of equivalent cases. Share setup without hiding inputs or expected evidence. For successful Result values, throw the typed error before asserting the value; a preceding .ok assertion adds no coverage. Protocol fixtures must reject unmodeled commands and preserve actual producer reply shapes.

Stateless source-analysis matrices may share a file-owned stdio MCP connection and retain representative compiled CLI parity checks. Every distinct source scenario still runs through the public workflow; session-lifecycle scenarios need their own sessions and cleanup assertions.

Report unavailable authority as a named skip. Await owned cleanup before a verifier reports success. Native process inspectors are released at Vitest worker teardown; fork termination does not run Node exit hooks, and per-file teardown would retire inspectors still needed by other files.

Measure slow cases before pruning capacity coverage. Retain inputs beyond the formerly failing size/depth and distinguish cold runs, warm builds, coverage, and toolchain conditions when comparing timings.

Behavioral depths

Depth Location What it proves
Module src/**/*.test.ts One domain, service, or adapter through its narrow public surface
Composition tests/composition/** Provider-neutral session and registry wiring; filesystem access is limited to fixture materialization
Boundary tests/boundary/** Exactly one production filesystem, process, network, browser, CLI, or provider boundary
MCP boundary tests/boundary/mcp/** MCP transport and tool-contract boundaries with isolated sessions and explicit cleanup
Acceptance tests/acceptance/** A compiled CLI or MCP journey; injected providers still make it integration
Conformance tests/conformance/** Shared provider contracts, parameterized by declared capabilities and explicit opt-outs
Evaluation tests/evaluation/** Deterministic evaluator parsing, scoring, and report generation

Colocated domain tests may import only their owning domain and inward dependencies. They do not construct sessions or servers and do not start subprocesses or network listeners. Colocated application tests exercise one service through explicit ports and recording adapters, never a CLI or MCP entrypoint. Composition tests may assemble provider-neutral sessions and registries but do not cross production filesystem, process, socket, or browser boundaries. Boundary tests cross one production boundary. Only acceptance tests assemble the complete runtime or invoke the compiled product surface.

Focused immutable builders and recording ports shared by one test family live beside their production owner as src/**/*.fixture.ts. They are typechecked with the suite and excluded from package builds; broader runtime and provider fixtures remain under tests/fixtures/**.

tests/process-global/** is reserved for cases with a demonstrated dependency on process-global state. Those tests use isolated forks so environment and exit-status changes cannot leak between files. Serialize a case only when it demonstrably shares an external resource that cannot be isolated. Reusable, test-scoped fixtures live under tests/support/**; immutable source artifacts remain under tests/fixtures/**. The process-global Vitest configuration contract rejects new direct temporary-root creation outside the workspace seam and its narrowly documented boundary/package exceptions.

Real Hopper, Ghidra, IDA, browser, package, and managed-code claims belong to their explicit npm run verify:* lanes. The reconstruction-readiness lane also checks deterministic rerun, tamper, and stale-input handling; those checks do not execute extracted JavaScript modules. When application runtime behavior is needed, exercise the actual target through browser, Electron, or process capture. Real model trials are manual; Vitest covers deterministic evaluator logic.

verify:managed runs the portable PE byte-fixture conformance entrypoint under scripts/verify/managed/, with its byte builder under scripts/fixtures/managed/. It checks static classification, members, reconstruction, native-boundary relationships and application graphs without executing fixture PE files. Operator-local manifests and actual ILSpy oracles remain optional, separately reported checks; the real Ghidra NativeAOT lane has its own toolchain prerequisites. See the managed guide for those configurations. Generated completion-ledger checks use the same owning entrypoint and include its verifier/fixture files in their cache inputs.

End-to-end, integration and golden evidence

The optional verify:qwen-client, verify:pi-client, and verify:hermes-client lanes require an installed native client. Select its executable with REA_VERIFY_QWEN_COMMAND, REA_VERIFY_PI_COMMAND, or REA_VERIFY_HERMES_COMMAND. They configure disposable profiles through the public REA CLI, preserve caller settings and backups, and exercise real stdio MCP calls and skill loading. The call mode checks complete catalog discovery and actual JavaScript Evidence; pass -- chat to check plain chat with REA enabled. Qwen also reads the full result that its client offloads, and Pi exercises default codemode execution. Use REA_VERIFY_RUNTIME_ROOT to select a production-only installed REA package. Receipts retain client versions, result digests, host coverage and owned-process cleanup. These POSIX lanes use a deterministic loopback model, so they do not prove live model-provider or native Windows compatibility. Qwen, Pi and Hermes bind both HOME and USERPROFILE to the disposable account. Qwen and Pi discover the shared personal skill through their default search paths without an explicit skill setting; Hermes discovers it through the selected HERMES_HOME. Set REA_VERIFY_HERMES_STICKY_PROFILE=1 to exercise native Hermes selection of a named sticky profile with HERMES_HOME still pointing to its root.

verify:gemini-client requires an installed Gemini CLI (verified with @google/[email protected]); select it with REA_VERIFY_GEMINI_COMMAND. The optional POSIX lane uses native GEMINI_CLI_HOME discovery in an isolated Git project, checks setup plans, backups and idempotence, activates the installed personal skill, validates all forwarded input JSON Schemas against their declared dialect, forwards the complete REA catalog and checks full JavaScript Evidence for a Unicode path. Use -- chat for ordinary chat with REA enabled, or REA_VERIFY_RUNTIME_ROOT for a production-only installed REA package. The local fixture exercises the real Gemini API adapter, including its untrusted_context tool-result envelope. Token counts are synthetic; live Google API and native Windows compatibility remain unverified. The lane disables the client's memory-based relaunch to preserve the caller's Node heap budget.

verify:opencode-client requires an installed OpenCode (verified with [email protected]); select it with REA_VERIFY_OPENCODE_COMMAND. The optional POSIX lane configures isolated XDG roots and OPENCODE_CONFIG_DIR, checks setup plans, backups, preserved JSONC comments and idempotence, activates the installed skill, validates all forwarded input schemas and verifies complete JavaScript Evidence for a Unicode path. JSONC is the default fixture; set REA_VERIFY_OPENCODE_CONFIG_FORMAT=json for JSON. Use -- chat for ordinary chat or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The native core runs with external plugins disabled through OPENCODE_PURE; the local OpenAI-compatible fixture has a caller-declared one-million-token context and synthetic usage. Skills use an explicit isolated skills.paths directory, so default OS-home discovery remains unverified. Live model APIs and native Windows are also unverified.

verify:claude-client requires installed Claude Code (verified with @anthropic-ai/[email protected] on POSIX and native Windows); select it with REA_VERIFY_CLAUDE_COMMAND. On Windows, select claude.exe: npm installs put only a claude.cmd shim on PATH, and the real executable is under node_modules/@anthropic-ai/claude-code/bin. The optional lane uses native CLAUDE_CONFIG_DIR discovery, an independent Git workspace, user settings and default built-in tools. It preserves caller preferences and backups, checks idempotent setup, waits through the native WaitForMcpServers tool when discovery is pending, and validates the complete catalog and forwarded input schemas. Call mode loads the personal skill through native Skill, checks its full body and verifies named-schema JavaScript Evidence for a Unicode path. Use -- chat for ordinary chat with REA enabled, or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The local Anthropic Messages/SSE fixture preserves native resource-hint envelopes in its request artifacts. Usage is synthetic; live Anthropic API and bare-mode personal-skill activation remain unverified by this lane.

verify:copilot-client requires installed GitHub Copilot CLI (verified with @github/[email protected]); select it with REA_VERIFY_COPILOT_COMMAND. The optional POSIX lane uses native COPILOT_HOME discovery, an independent Git workspace, guarded setup plans, preserved registrations, backups and idempotence. Native skill add registers the isolated installed skill; call mode loads its full body, validates all forwarded input schemas and checks named-schema JavaScript Evidence for a Unicode path. When Copilot spills a large MCP result, the model requests native view with forceReadLargeFiles and validates the complete returned Evidence rather than its preview. Use -- chat for ordinary chat or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The native offline BYOK adapter uses a loopback OpenAI completions/SSE fixture and gpt-5.4 model metadata; model inference and token usage are synthetic. Set REA_VERIFY_COPILOT_MODEL to examine another client model ID and REA_VERIFY_COPILOT_WIRE_MODEL for its backend model name (defaults to the client ID). Every actual HTTP request's model is checked and recorded separately. Both HOME and USERPROFILE point to the disposable account; native skill add still selects the installed skill directory explicitly.

The same native client passed full-profile chat and analysis with the actual wire model gpt-4.1 under a distinct BYOK model ID:

REA_VERIFY_COPILOT_MODEL=rea-gpt-4.1 REA_VERIFY_COPILOT_WIRE_MODEL=gpt-4.1 npm run verify:copilot-client -- chat
REA_VERIFY_COPILOT_MODEL=rea-gpt-4.1 REA_VERIFY_COPILOT_WIRE_MODEL=gpt-4.1 npm run verify:copilot-client -- call

These runs retain full canonical input/output schemas, all tools, installed-skill loading and complete Evidence readback. They use the fixture's requested one-million-token BYOK prompt setting; effective capacity and live API acceptance remain unknown. With the built-in gpt-4.1 model ID, the verified client still blocks the full input-schema profile before HTTP with compaction_static_context_blocked, even when requesting a larger BYOK prompt capacity. The effective capacity is unknown. The default verifier preserves the complete catalog and full schemas. To verify the existing compact profile:

REA_VERIFY_COPILOT_MODEL=gpt-4.1 REA_VERIFY_COPILOT_SCHEMA_PROFILE=compact npm run verify:copilot-client -- chat
REA_VERIFY_COPILOT_MODEL=gpt-4.1 REA_VERIFY_COPILOT_SCHEMA_PROFILE=compact npm run verify:copilot-client -- call

Both native workflows passed with the complete tool inventory. The verifier selects the profile only in its disposable Copilot registration and records it in the receipt. Compact advertisements can reduce nested validation structure; the server still validates complete canonical inputs. This is a tested alternate workflow, not a fix for the full-profile client/model limit tracked in #1554. Live model APIs, native Windows and default OS-home skill discovery remain unverified.

verify:grok-client requires installed Grok Build (verified with the official Linux x64 1.0.50 binary); select it with REA_VERIFY_GROK_COMMAND. This optional POSIX lane isolates GROK_HOME, an independent Git workspace and additional skill roots, disables foreign configuration discovery, guards setup writes, preserves unrelated registrations/backups and checks idempotence. Call mode loads the full skill through native read_file, queries all REA names and input schemas through native search_tool, and calls analysis through use_tool. If discovery reports partial, the native agent diagnoses REA registration and retries discovery. No startup-timeout override or fixed readiness delay is used. A larger Unicode-path fixture exercises native result offloading: the full retained Evidence is validated against its named schema, then native terminal queries recover the selected export facts, subject and full artifact digest in the next model request. This verifies useful artifact recovery without claiming that every offloaded graph fact enters model context. Native line-number and truncation envelopes remain in the request artifacts. Use -- chat for ordinary chat, or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The loopback OpenAI completions/SSE custom model declares a one-million-token context and synthetic usage. Live xAI APIs, native Windows and default OS-home skill discovery remain unverified.

verify:commandcode-client requires installed Command Code (verified with npm 1.79.2); select it with REA_VERIFY_COMMANDCODE_COMMAND. Auth, MCP tokens and sessions use the actual OS home, so provision a disposable POSIX account rather than overriding HOME. Set REA_VERIFY_COMMANDCODE_ACCOUNT_HOME to that account's actual home and create .rea-client-verification there with exactly Disposable REA client verification account followed by a newline. The lane owns that account's .commandcode configuration and shared REA skill. Build the runtime as the checkout owner, then run node scripts/verify-commandcode-client.mjs as the disposable account; the npm shortcut also needs a writable checkout. The verifier does not create accounts or install clients. It guards setup targets, preserves an unrelated registration and backups, checks idempotence, and verifies native default shared-skill discovery and complete activation. Default deferred-schema delivery stays enabled: the native prompt advertises every REA tool, exact-name search_tools returns every input schema, and a native shell query recovers the saved catalog's count and digest when the client spills it. Actual analysis of a Unicode-path fixture delivers full named-schema Evidence to the next model request. Use -- chat for ordinary chat or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The loopback OpenAI completions/SSE model is a keyless BYOK endpoint declaring a million-token context with synthetic usage. A synthetic account-key value satisfies the client's print-mode gate; native local-only mode refuses hosted API calls. Updates, telemetry and cron are disabled. Hosted authentication, live models and Windows remain unverified.

verify:omp-client requires installed OMP (verified with the official Linux x64 18.8.7 binary); select it with REA_VERIFY_OMP_COMMAND. This optional POSIX lane isolates the default agent, global config and XDG roots, guards setup targets, preserves unrelated registrations and backups, and checks idempotence. It compares native rejection of an invalid profile with REA's refusal to plan fallback writes. OMP's default xd:// interface mounts the complete REA catalog as devices: call mode reads the complete installed skill and tool documentation, then dispatches analysis through native write. It checks full named-schema JavaScript Evidence with a Unicode path in the next model request. Device metadata and documentation are distinct from forwarding all JSON schemas as model functions. Native print-mode MCP readiness uses its defaults. Use -- chat for ordinary chat or REA_VERIFY_RUNTIME_ROOT for a production-only installed package. The loopback OpenAI completions/SSE model fixture declares a one-million-token context and synthetic usage. Live model APIs, native Windows, named-profile native execution and default OS-home skill discovery are unverified; skill discovery uses an explicit isolated custom directory.

verify:deepseek-client requires an installed DeepSeek Harness (dsh; verified with @deepseek-ai/[email protected]) and Git. Set REA_VERIFY_DEEPSEEK_COMMAND to select its executable. This optional POSIX lane installs the skill through rea setup --skill, writes the native Cordis MCP patch into an isolated DSH_HOME, and binds HOME and USERPROFILE to the disposable account. The native default ~/.agents/skills search path discovers the installed skill without a DSH_AGENTS_HOME override. A separate Git root prevents inherited project skills from masking a missing installation. The default call mode checks skill loading, complete catalog discovery, the actual forwarded input schemas and regexes, and full JavaScript Evidence for a Unicode path. Use -- chat for ordinary chat with REA enabled, or REA_VERIFY_RUNTIME_ROOT for an installed package. The fixture uses Harness's custom OpenAI adapter and the one-million-token context capacity declared by its default DeepSeek Flash model; it does not establish live DeepSeek API or native Windows coverage. Receipts include request artifacts, result digests and owned-process lineage.

Full E2E tests invoke the production command dispatcher and real providers, without fake launchers, runners or responses. verify:keyed-archive writes Foundation binary, XML, and integer archives using one run-owned Swift module cache. It checks CLI and separate stdio MCP parity, pagination, escaping-path rejection, an XML graph golden, and numeric meaning for safe controls, unsafe integers, boxed numbers, and integral-valued reals. verify:asset-catalog compiles source-owned colors with actool, invokes real assetutil, then checks CLI/MCP results, exact catalog digest, every raw metadata field, pagination and malformed input rejection. Neither artifact workflow requires Hopper or launches it. Both run in macOS CI. verify:macos-bundle needs only macOS with Command Line Tools. It compiles a source-owned app with clang: a versioned framework, XPC services, an app extension, a login item, a privileged helper, launchd plists, and a helper tool, signed ad hoc. It packs the app as a directory, a ditto ZIP, and an APFS DMG, then checks that inspect-artifact plus project-apple-application-graph report the same bundle anatomy for all three through the CLI, with stdio MCP parity. It also checks that the DMG is detached afterwards. The same app covers trace-dylib-resolution:

  • each resolution status and finding, with CLI/MCP parity;
  • for every traced image, dependencies, rpaths, and install names equal to otool -l;
  • for the main executable and an XPC service, a predicted load order equal to the images dyld actually loads under DYLD_PRINT_LIBRARIES.

The lane also compiles executable/library pairs with empty embedded directory, versioned-path, and suffix settings. It checks present and removed dependencies through CLI and MCP, and compares actual DYLD_PRINT_SEARCHING diagnostics for root-level candidates. These cases distinguish an empty search directory from an empty versioned scan or a suffix that only repeats the original path.

It runs in macOS CI.

Prepare the native inspector before cases that measure producer output or exit behavior; keep startup deadlines and cancellation in distinct cases. Run real process-capture verification separately from package or build checks. On macOS, new npm processes can become token-unreadable after changing their display title and prevent verified cleanup during a capture's ownership sweep.

MCP SDK transport tests with recording providers remain integration tests. They are useful for schema drift and failure projection but do not prove that Hopper, Ghidra or another substituted engine works. verify:package proves packaging/install behavior and fake-provider integration; use the corresponding real-provider lanes for engine claims. Packed-bridge checks verify shipped files and Python syntax. Real Apple dispatch and Interface Builder verifiers currently prove format integration through production readers.

verify:hopper exercises an installed Hopper through the production stdio MCP server and CLI. It checks source-owned call chains, CFG edges, references and complete large inventories, then probes unnamed bookmarks, annotation clearing, batch validation before mutation, malformed addresses and regexes, segment-end partial reads, and synthetic file-offset rejection. Advertised schemas are validated in their JSON Schema dialect and successful replies are checked against their advertised output schemas. Navigation checks cover interior-object cursor readback and mapped-memory boundaries. Annotation probes verify invalid native text and duplicate destinations/names before mutation, preserve unselected label owners, and exercise explicit batch label swaps. Function locals retain observed names and stack displacements. Graph probes check symbol/interior-address parity and a source-owned recursive cycle; literal tracing preserves complete queries and whitespace. Disposable binary copies prove that switching and closing actually removes the native document, and that CLI byte, function-dossier, literal-trace results and invalid-address diagnostics agree with MCP. No provider is mocked in this lane. Real search probes cover Unicode names, literal metacharacters, case and regex modes, annotation cache invalidation, complete native fragments of long literals checked against byte reads, escaped UTF-8/control text, byte-preserving Latin-1 decoding, and Hopper's UTF-16 symbol-name truncation boundary. It verifies native CallReference classifications across reference and dossier results, retains long string fragment metadata in dossiers, and exercises pathological regex deadline and cancellation followed by successful requests in the same native session. When the macOS Objective-C fixture is present, the lane also verifies native UTF-16 string objects and their inferred encodings against their actual bytes. Native terminal calls are checked across reference, instruction, assembly, block-range and procedure-length projections; block endpoints are normalized using actual native membership. Exact Objective-C names also exercise named CLI selectors for function, instruction, decompilation, reference and search operations. A literal --help trace query proves that selector data is preserved independently of global CLI flags. Unmapped annotation destinations and oversized later batch names fail before any earlier edit is applied.

verify:hopper:deadlines checks the native client's optional request deadlines on macOS with an owned source-built fixture. Zero and elapsed deadlines must leave native comments unchanged, including when a synchronous progress observer delays timer dispatch. It also observes a short analysis deadline and verifies subsequent wire recovery and clean shutdown. Caller timeout does not interrupt Hopper's synchronous native operation. Socket boundary tests deterministically cover active and queued expiry, late replies and timer cleanup.

The Linux demo lane remains a separate verify:hopper:linux command. Its lifecycle checks attempt competing CLI launches for the active target and a different target, requiring rejection before either can change the owning MCP session's documents or procedures. Lease boundary tests also preserve unresponsive or malformed live endpoints while allowing confirmed stale sockets to recover.

verify:hopper:fat is a separate macOS lane requiring installed Hopper and the existing Xcode clang/lipo toolchain. It compiles arm64/x86-64 thin executables and one- and two-slice FAT32 containers, verifies exact/interior address mappings against bytes in the original files, checks CLI/MCP parity, and checks owned runtime cleanup. Source byte changes, removal, permission denial (for non-root callers), and nonregular replacement must retain native partial mapping facts, reject unverified original-file coordinates, and recover after restoration. Single-slice FAT cases also relocate the slice without changing its loaded bytes. verify:hopper:fat64 additionally checks FAT64 preparation through Hopper's native Mach-O loader, source-container mappings, profile identity, malformed and ambiguous slice rejection, and temporary-image lifetime. Both lanes have been verified on Hopper 6.1.0-demo; this establishes REA's prepared FAT64 workflow, not native FAT64-loader support. Cross-architecture fixture compilation is not required by verify:hopper.

Golden tests use immutable captured text inputs with producer/source provenance under tests/fixtures/golden/. Expected results are reviewed for the semantic claim; capture commands do not automatically approve new expected outputs. Do not call handcrafted utility output or synthetic binary builders real-data goldens. Keep unsupported binary layouts and malformed boundaries as targeted regressions until a real fixture establishes equivalent coverage.

verify:browser exercises source-map failure isolation and expanded-output limits through the compiled CLI and stdio MCP with real Chrome. It also submits five 2 MiB source-map annotations, checks retained script identities and explicit map omissions, and closes its owned fixture target. Replacement and document-reset behavior has separate producer-boundary coverage.

The raw Chrome fixture launchers use --password-store=basic so Linux startup does not wait for an unlock prompt from a reachable but locked desktop keyring. Such a stall leaves the initial page unloaded with empty or placeholder CDP URLs; it does not establish permission to capture that page. See the keyring diagnosis.

verify:browser also captures a source-owned noise canvas as a real PNG above 8 MiB through the CLI and stdio MCP, with complete byte/digest parity and real PNG decoding. Its SDK client explicitly permits the larger inline JSON response; this lane does not establish large image-comparison request transport coverage.

The scenario checks read locale, timezone, device scale and viewport dimensions from actual page JavaScript through CLI and stdio MCP DOM captures. They also verify that attached-page emulation is released after success, initialization failure, action failure and cancellation while the external target stays open.

verify:browser:network is a focused real-browser lane for transaction identity, selected request/response bytes, binary and compressed responses, duplicate headers, credential and declared-secret redaction, redirects, streaming cutoff, CLI/MCP parity, and owned-profile cleanup. Set REA_BROWSER_EXECUTABLE to an installed Chrome-family browser. An optional script argument selects an already installed package's scripts/rea.mjs entry point for packaged-artifact checks. The complete verify:browser lane includes these same checks.

After building, verify:browser:dom checks empty and HTML-whitespace form destinations against a native Chrome DOM-property oracle through CLI and stdio MCP. It requires REA_BROWSER_EXECUTABLE and accepts an optional installed REA entrypoint. The full browser lane includes the same public-adapter assertions before other fixtures navigate the selected page.

verify:browser:scripts checks active script capture → exact-byte export → existing static JavaScript analysis through CLI and stdio MCP, including manifest readback, competing query variants, and resolved relative imports. It uses an installed browser and accepts an optional installed REA entrypoint. The complete verify:browser lane also exercises passive script export through both public adapters. See website script export.

verify:browser:modules compares CLI and stdio MCP traces against an independent real Chromium module-loading fixture: import-map scopes, package prefixes, null/backtracking rejection, query/fragment identities, repeated module instances, lazy and computed unknowns, exact source/map identities and no implicit refetch. Set REA_BROWSER_EXECUTABLE; the optional script argument selects an installed package entrypoint. The complete verify:browser lane and existing conditional Chrome CI include this verifier. See module relationships.

Real-toolchain verification lanes

Each real-toolchain command must require only the host tools needed to prove its stated claim. Use a host-native fixture for host/provider acceptance, and place optional cross-target formats or platform-specific runners in separate commands. Check prerequisites before starting expensive work and name the missing command, target, and lane in any failure message. A lane must not imply that a host or target is covered when it was skipped.

Provider admission accepts Ghidra 12.1.x and the JDK range declared by that installation (application.java.min through application.java.max). Current 12.1 releases require JDK 21 or newer and set no maximum. The lanes below still prove behavior on the verified Ghidra 12.1.4 and JDK 21 build.

Ghidra lane Supported runner/target Additional local tools
npm run verify:ghidra Linux x64/arm64 ELF or macOS x64/arm64 Mach-O Host C compiler, Ghidra 12.1.4, and full JDK 21
npm run verify:ghidra:seeds POSIX; raw x86 fixture; seed ordering, BSS labels, analyzer readback and unknown-name rejection Ghidra 12.1.4 and full JDK 21
npm run verify:ghidra:swift macOS x64/arm64 Mach-O Host Swift compiler, Ghidra 12.1.4, and full JDK 21
npm run verify:ghidra:switch Linux x64 ELF; GCC/Clang optimized and stripped switch fixtures GCC, Clang, GNU nm/objdump/strip, Ghidra 12.1.4, and full JDK 21
npm run verify:ghidra:aarch64-jump-table Any supported Ghidra host; AArch64 ELF; signed/unsigned byte/halfword tables; host ARM64 Mach-O Clang with AArch64 target support, Ghidra 12.1.4, and full JDK 21
npm run verify:ghidra:cross-format Any supported Ghidra host; also analyzes AArch64 ELF, x86-64 PE, and x86-64 Mach-O clang, LLD, and lld-link in addition to host-lane prerequisites
npm run verify:ghidra:windows Controlled Windows x64 with native x86-64 PE; -- --x86 selects native x86 PE Ghidra 12.1.4, full JDK 21, and the matching native artifact

The optional Swift lane compiles a source-owned fixture and checks modern and provider-demangled procedure names through real stdio MCP, with and without the dSYM companion. It checks unfiltered, structs, and literal-name filters, same-address alias evidence, deduplication, and explicit unresolved categories. It does not add a Swift prerequisite to the host C fixture lane.

The host lane also verifies namespaced C++ ABI symbols and annotation identity. That facet requires Ghidra's matching demangler_gnu_v2_41 native component in GPL/DemanglerGnu/os/<platform>/ or build/os/<platform>/, in addition to the native decompiler. The namespace verifier reports missing or non-executable components before testing the fixture. REA does not build or install them.

Windows native conformance runs with npm run verify:windows-native and does not require Ghidra or Java. It includes Job Object commit-limit controls (a limited child that reaches the limit, a control that does not, and invalid limit rejection). An optional independently compiled Windows fixture adds in-place reparse, breakaway, environment, and token observations. The separate scripts/verify-windows-native-inspection.mjs lane uses the .cache/windows-native-inspection-fixture artifact built on Linux with node scripts/build-windows-native.mjs --inspection-fixture. Its controlled first deletion failure checks retained NTFS ownership, later-query recovery, successful-admission caching, and environment-teardown cleanup. It neither establishes the cause of the intermittent #1701 host failure nor replaces verification with the shipped production artifact. npm run verify:ghidra:windows:package additionally packs and installs REA into an isolated prefix and runs ordinary-user CLI/MCP operations against their canonical schemas. Run the controlled fixture generator before that lane, or supply an installed package root and an explicit fixture as arguments. The host-native Ghidra lane also verifies native value tracing through the production CLI and a separate stdio MCP process. It compares complete dependency graphs, validates Evidence and upstream/workflow profiles, checks capability discovery, and closes the MCP session. No provider or transport is mocked. It also validates every advertised input/output JSON Schema and the exercised MCP outputs, probes address spelling and name/address ambiguity, and checks direct versus targetless calls, byte-read completeness, invalid input diagnostics, CLI/MCP parity, atomic annotation rollback, refreshed inventories, unchanged executable bytes, and discarded edits after reopen. A deliberately long temporary path exercises private Unix socket allocation and cleanup, including cancellation after a real headless process launches. Native annotation probes reject NUL and unpaired Unicode surrogates without partial edits or a broken bridge, preserve supported Unicode and control text, and check lossless malformed-text diagnostics. Memory-to-file mapping is checked against original artifact bytes. The fixture also stores a pointer one byte past a function entry; exact xrefs, raw procedure references, and CLI/MCP dossiers must retain that data edge. A valid legacy snapshot reconstructs the former omitted edge under its older profile; CLI and MCP must reject that binding with a mismatch reason and recovery advice. The rejected open must preserve the active live session. Exact external entries must resolve while retaining an empty body; unknown external addresses remain unresolved and external annotations are rejected. An adversarial regex over a full 12 KB literal must report stack exhaustion as a resource constraint, preserve live annotations, and allow complete literal searches afterward; CLI and MCP must agree on both results and recovery advice. Real snapshot lifecycle checks retain edited API results as Evidence while rejecting immutable snapshot saves and imports before and after a repeated open of the same target. They verify unchanged live annotations and run identity, absent rejected output files, an unchanged source snapshot, and successful snapshot import/save after closing and recreating the database. The pristine snapshot is written by an independent real CLI session. They also start a real annotation and snapshot close concurrently: the edit must succeed, the snapshot must be rejected without creating a file, and the edited session must remain usable until explicitly closed. Source-admission probes change a caller-owned fixture after open_binary but before the first Ghidra query. They require an actionable artifact_changed error preserving both digests and the selected path, unchanged provider availability, failed-copy cleanup, and successful recovery after reopening. After import, deleting that source must preserve the captured database identity. Instruction inspection and containing-function lookup also agree across hexadecimal case, leading zeros, and encoded default address-space spellings. The same source-acquisition workflow exercises missing and directory-replaced inputs, plus real read-permission denial on a non-root host. Root runs report that permission-denial check as unverified. Filesystem integration separately checks selected-platform routing and exclusive creation. Namespace annotation probes compile a separate host C fixture with C++ ABI symbols, avoiding a C++ runtime prerequisite. Real Ghidra demangling supplies duplicate leaf names in two top-level namespaces and a nested namespace. The workflow verifies leaf and qualified renames, repeated reuse of fully qualified readback, lookup by the returned name, literal namespace-like leaf names, rejection of empty qualified leaf names without changing comments, CLI behavior, and independent CLI/MCP database ownership. Large-result probes compile initialized host-native data sized from the pinned MCP SDK receive budget. Real byte reads, annotation edits and function dossiers exceed that budget while preserving the connection and active analysis run. Successful-result and oversized-error delivery constraints must identify their successfully retained Evidence records; export must recover every source byte, complete annotation and original error diagnostic, with CLI parity and an unchanged executable. Formatter integration separately checks unacknowledged Evidence recording. Long ordinary procedure names and encoded address-space selectors must produce normal validation errors without exhausting Java's regex stack or losing the private bridge connection. Each rejection is followed by a real provider lookup; address-like literal names still resolve exactly after annotation. Malformed annotation readback, memory completeness, and inventory data remain separate SDK/provider integration cases; success from a real provider cannot establish rejection of a contradictory provider response.

The Linux switch lane checks dense, sparse-with-holes, shared-body, nonzero, negative, and nonexact JSON integer labels plus a comparison-only control. Independent source labels, ELF file bytes, table slots, and bounds branches define expected case/default destinations. Production CLI and MCP must agree; debug labels must retain their signed values. For stripped negative fixtures, the independently checked 32-bit dispatch permits equivalent unsigned labels only alongside the reported low-confidence undefined4 parameter; original source signedness remains unknown and the ABI residual must remain visible. unsafe labels remain unresolved and every recovered destination is retained. The compiler oracle also injects malformed records to check that its assertions reject missing/default-confused labels, wrong destinations, and numeric guesses. It also invokes the actual bridge methods on detached Ghidra model objects to check ambiguous dispatches, signed literals, precision bounds, shared targets, and conflicting labels. This reflection fixture depends on the pinned Ghidra model and does not claim a compiler produced those synthetic graph shapes. Pass --entrypoint /path/to/installed/rea-agents/scripts/rea.mjs directly to scripts/verify-real-ghidra-switch.mjs to verify an installed package through the same compiler oracles and CLI/MCP checks.

The cross-format Ghidra lane also analyzes an optimized AArch64 ELF switch fixture. It checks the recovered case values against the source cases and requires unresolved table bounds or case mappings to remain visible as residual unknowns.

npm run verify:inspector requires the supported Node.js runtime and installed REA dependencies. CI runs it on Linux and macOS x64/arm64 and Windows x64. It starts owned loopback Node Inspector fixtures and verifies discovery and passive observation through the CLI and stdio MCP, including special filenames, unresolved discovery locations, and independently resolved loaded scripts. Double-quote filenames are tested on POSIX only because Windows does not support them.

Node Inspector verification does not establish Electron GUI or other engine behavior. Browser and Inspector share CDP transport/value and file-location helpers; their producer workflows remain distinct.

Android APK analysis

The real Android lane also runs on Windows x64 with its matching bundled native controls. It verifies bidirectional protocol input, CLI/MCP result parity, real MCP cancellation/disconnect and Java exit after abrupt CLI owner termination. The forced Windows exit does not exercise the POSIX SIGTERM handler and can leave temporary workspace files. verify:windows-native separately checks binary input, backpressure, EOF and pending-write job closure.

npm run verify:android requires an existing Java 17+ and an explicit REA_JADX_MCP_JAR for jadx-headless-mcp 0.7.1. Set REA_ANDROID_TEST_APK to the fixed public ApiDemos v6.0.18 fixture. Obtain both with the explicit npm run fixtures:android command; files are SHA-256 verified and kept under ignored _reference/. No Gradle build, Android SDK, emulator or application execution is required. The lane compares real CLI/MCP package, class search, class inventory, method decompilation and incoming references. See Android analysis for boundaries and resource budgets.

verify:jeb requires a caller-started JEB client serving MCP at REA_JEB_MCP_URL (default http://127.0.0.1:8425/mcp) with a project already open, and verifies real CLI client inspection, unit listing coverage, and method decompilation. The script records the engine's exact response to open_jeb_project; JEB 5.48.0 headless instances do not advertise that tool. REA does not install or launch JEB; see JEB analysis for the bring-your-own boundary. Verified against JEB 5.48.0 serving jeb-mcp-server 1.3.0 at four scales: a compiled Java class fixture, a locally built signed probe APK with a launcher activity (manifest, v1/v2/v3 certificates, dex bytecode, filtered and paginated listing), the published Signal 8.30.3 universal release APK (109 MB, R8-processed Kotlin/Compose, four signature schemes, native arm64 ELF units) with MainActivity method decompilation through both CLI and MCP, and the official Flutter Gallery 2.9.2 release APK (112 MB): per-ABI libapp.so Dart AOT snapshot images analyzed as native code (75k methods on arm64) and decompiled through both CLI and MCP with explicit unit selection, plus the thin Java wrapper and decoded manifest. The dedicated Dart snapshot processor and run_script remain GUI-only surfaces in JEB 5.48.0 headless. Authenticated IPA and macOS application inventory projection is documented in Apple application analysis.

npm run verify:harmony requires no SDK or engine. It projects the pinned harmony-next-pipeline 1.0.1 release HAP from an MIT-licensed source repository, fetched by npm run fixtures:harmony into ignored _reference/harmony-integration/, set as REA_HARMONY_TEST_HAP. The lane exercises inspect_artifact, projection determinism, Stage-model path hints, exact component sets, signing-entry limitations, and CLI/MCP parity. The release artifact has no independently established compiler/SDK or reproducible source build; the default lane does not verify an installed package. See HarmonyOS package inventory for those proof boundaries. HarmonyOS .har libraries are deliberately not suffix-classified because .har also names HTTP Archive JSON.

Synthetic producer regressions run independently:

npm run test:focused -- tests/boundary/process/jadxIntegration.test.ts tests/boundary/process/androidAnalysisMcp.test.ts

Apktool resource decoding

npm run verify:apktool -- --apk PATH exercises both Apktool operations against a real launcher (REA_APKTOOL_COMMAND or PATH) and one APK. No download or build step is involved. Record: apktool 2.7.0-dirty (Debian packaging) on Linux with OpenJDK 25, against a signed aapt2-built probe APK carrying two resource locales — launcher version, on-disk digest agreement with the reported target identity, apktool.yml metadata projection (1.2.3, SDK 24–34), manifest package agreement, and both the default and de string tables projected through the locale option. Parser goldens in src/apktool/ApktoolDecodeOutput.test.ts come from the same decode. See Apktool resource analysis for boundaries and budgets.

ADB device acquisition

npm run verify:adb exercises every ADB operation against a live device or emulator through the caller's adb binary (REA_ADB_PATH or PATH). No SDK installation or emulator management is involved; the device-mutating tools run only in the explicit lifecycle section (--install-apk drives a lane-owned install → resolve → start → observe → force-stop → uninstall probe). --pull acquires a real package and re-digests the pulled files on disk against the returned SHA-256 values; --serial selects a device when several are attached.

Record: adb 34.0.5-debian on Linux against an Android 14 (API 34) x86_64 emulator. The lane covered the complete observation surface — 337-process listing, 287 binder services including AIDL /-suffixed names, 92 features including hex GL versions, display size/density, window focus (legitimately null on headless devices), settings get global adb_enabled, bounded logcat, directory listings — plus a real two-APK split set (base.apk plus split_probe.apk, built and installed through install-multiple) pulled with byte-exact digests, a push/pull roundtrip verified by the device's own sha256sum, a 1.3 MB screen capture with PNG dimensions, dumpsys package projection, and the full lifecycle: unique resolution before launch, launcher-activity start through cmd package resolve-activity with am start -W, the started app visible in the process listing, force-stop, uninstall, and zero matches after removal. System-package pulls whose APKs keep non-base.apk names report the unknown role with the file-name basis, verified with com.android.settings (single 73.9 MB APK). Parser goldens for adb devices -l, getprop, pm list packages -f, and pm path come from the same device. See ADB device analysis for boundaries and budgets.

Optional NativeAOT Ghidra analysis

The source parser lane is independent of the optional annotation JAR. Its synthetic RTR 9.1 PE golden checks captured-byte digest admission, rehydration into an immutable parser overlay, MethodTable/slot relationships, and frozen literal extraction; provider tests exercise inspect_native_load_image and address-based inspect_native_data_type with the database result still unavailable. These tests do not establish Windows Ghidra host behavior.

This lane is separate from the default native lane. build:fixtures:nativeaot requires an existing .NET SDK 8.0.416, the platform NativeAOT compiler/linker and runtime pack 8.0.22. It builds benign sources into ignored _reference/ and records independent symbol/directory/SHA oracles; it never executes the target. Linux also builds stripped, ordinary-native and small malformed/ambiguous/layout negative inputs. The optional Windows fixture workflow builds a PE on a Windows runner; analyzing that PE on Linux does not verify a Windows Ghidra host.

Build the clean pinned upstream adapter with build:ghidra:nativeaot, then set REA_GHIDRA_NATIVEAOT_JAR. Run verify:ghidra:nativeaot -- symbols, -- stripped, -- ordinary, -- unsupported, -- malformed, -- ambiguous, and -- loader-failure, and -- default-native separately. The loader-failure mode source-builds a deliberately failing JDK 21 initializer and checks the actual loader cause and cleanup. The default-native mode verifies ordinary analysis with the optional extension disabled. The real MCP lane checks source identity, inline format discovery, metadata relationships/slots against independent compiler symbols, frozen strings, pseudocode and owned cleanup. Set REA_NATIVEAOT_PROOF_CLI=1 for one equivalent CLI type inspection; this costs an additional full import. Select an unpacked installed package with REA_NATIVEAOT_PROOF_PACKAGE_ROOT, a fixture directory with REA_NATIVEAOT_PROOF_FIXTURE_ROOT, and optional evidence capture directory with REA_NATIVEAOT_PROOF_CAPTURE_DIR (absolute paths).

Keep builds/imports sequential on small hosts; scope GHIDRA_HEADLESS_MAXMEM (e.g. 768M) to this command and use CPU affinity if needed. REA does not install or upgrade Java, Ghidra, .NET or native toolchains. See the supported layout and provenance.

IDA MCP adapter

npm run verify:ida -- --target /absolute/path/to/program --procedure main uses the existing REA_IDA_MCP_CONFIG registration. It installs no engine, Python package, or compiler. The target must already be open in the GUI for the legacy attached profile; the database-supervisor headless profile opens a digest-verified private copy. A caller-supplied fixture keeps prerequisites limited to the selected engine and host. tests/conformance/ida/inventory.c provides an optional small native fixture source with an exported rea_fixture_add function.

The lane invokes the production CLI dispatcher and connects the pinned MCP client SDK to the production REA server. It verifies function Evidence and CLI/MCP parity, inventory/search, pseudocode, instructions, xrefs, malformed input, original-input preservation, and lifecycle cleanup. For headless analysis it confirms that the owned database IDs disappear from upstream discovery and private workspaces are removed. For attached analysis it confirms the existing GUI target remains reachable with the same input identity. --package-root selects an installed/extracted REA artifact. --report writes private local observations with mode 0600; the console summary contains no target paths or upstream output.

Adapter and composition tests cover producer parsing, pagination, canonical entries, external callees, target switches, cancellation draining, snapshot replay exclusion, ownership failures, and incomplete cleanup. They do not establish real IDA operation. The initial real workflows cover legacy upstream 1.4.0 on a Windows GUI and the modern supervisor at upstream commit c133c3853faa111a9b00ee615c013b720d0c4acd with Windows x64 IDA 9.3, and the headless profile on macOS arm64 with IDA Professional 9.1 (build 9.1.250226) at the same upstream commit, using tests/conformance/ida/inventory.c built with clang -O0 -arch arm64 and the procedure _rea_fixture_add (Mach-O C symbols carry a leading underscore). Before the headless session waited out the transient busy health a worker can report right after opening, that lane failed intermittently on that host. Linux headless, modern attached GUI tools, other engine versions and architectures remain unverified; see the provider guide.

DOS Ghidra analysis

npm run verify:ghidra:dos requires the supported Ghidra and JDK installation on Linux x64 or macOS x64/arm64. It generates a source-owned MZ fixture without a DOS emulator or compiler, then checks real 16-bit decoding, segment relocation, near/far calls, decompilation, disjoint function body ranges, stable CLI/MCP observations, unchanged source bytes, and owned process/project cleanup. Raw p-code address-space selector tokens are reported separately from the stable observation comparison. Linux x64 and macOS arm64 are verified; macOS x64 remains unverified. This lane is separate from host-native and optional cross-format verification. See DOS analysis.

npm run verify:ghidra:com uses a generated headerless fixture with no compiler, DOS emulator or game data. It exercises explicit admission, BinaryLoader entry preparation, measured register context, whole-file byte readback, source offsets, unmapped PSP/partial reads, actual decompilation, CLI/MCP parity and owned cleanup. Both segmented-address lanes reject oversized default, explicit-space and encoded-space coordinates through real reads, function queries and annotation attempts. Rejected annotations must preserve the live function dossier; CLI and MCP must report the truncation constraint, while leading-zero coordinates still resolve correctly. It has the same Ghidra/JDK prerequisites as the MZ lane. Neither lane claims DOS runtime or PC-98 device execution.

npm run verify:ghidra:raw uses a generated ARM little-endian headerless image through the compiled CLI and stdio MCP server. With the same Ghidra/JDK prerequisites, it checks preparation and bridge startup, verified load mappings, instruction and byte readback, canonical-entry session reuse, replacement when the entry changes without provider_id, unchanged source bytes, and owned cleanup. Six-byte x86 and x86_64 controls also verify supported compiler specifications, independent mappings, unchanged bytes, and decompilation of a known return value. AArch64 remains covered by deterministic tests rather than this real-provider lane. See raw binary analysis.

Apple Interface Builder archives

npm run verify:interface-builder compiles the source-owned AppKit XIB into a real .nib with Xcode ibtool, wraps it in a temporary app bundle, and runs the compiled CLI and stdio MCP server with production providers. It compares their decoded results and checks the view hierarchy, outlet, action, evidence coverage, and truncation status. Storyboard compilation additionally requires an installed iOS platform.

Keep the provider-specific acceptance path independent from optional cross-compilers. Cross-format failures belong to the cross-format lane and must not make native host acceptance unavailable.

Native platform baseline in CI

Pull requests that change only root README*.md files, docs/, AGENTS.md, or CONTRIBUTING.md run formatting and generated-document validation without the source-test shards or native package lanes. Classification compares the PR head with its merge base, so later base-branch changes do not expand that scope. Static and coverage aggregate jobs remain present and fail if classification fails. Every code change retains the complete four-shard Linux suite; individual test files are not selected by imports or filenames.

scripts/ci/plan.mjs selects lanes using the ownership table in scripts/ci/scopes.mjs. Provider implementation changes run their real-provider verifiers and the Linux x64 installed-package/Inspector baseline. Test-only changes run the Linux suite without native package verification. Website changes run website checks. Provider workflow changes validate the workflow and run that provider's verifier. Manual-provider/Android workflow changes run workflow validation only. Shared process/contract/CLI boundaries, dependencies, packaging inputs, unknown paths, and changes to the planner select the full baseline. Package script-only edits select their owning verifier when its entrypoint is identifiable; unknown or shared script changes select everything. Android build-script edits retain the Linux suite and workflow validation; real Android verification stays manual. Renames account for both paths.

Main pushes and manual dispatch retain the full baseline, including all portable provider lanes and the native matrix below. A ci:full PR label selects the full baseline on the next PR run; manual dispatch can request it immediately. Each classification job records its selection and broad-fallback reasons in the run summary. Representative Git-diff cases live in tests/boundary/filesystem/ciChangeScope.test.ts.

CI required checks classification and every selected lane, rejecting failed, cancelled, or unexpectedly skipped work. Existing required check names are preserved; changing branch protection to require the new aggregate is a separate repository-settings operation. Newer commits cancel superseded CI for the same PR or main branch. Release publication has its own concurrency group and retains its package verification and public-registry canary. Licensed/self-hosted provider workflows and the Android smoke test retain their explicit triggers.

CI exercises the pinned Node.js runtime on native hosted runners:

Host Runner Baseline checks
Linux x64 ubuntu-latest Installed package and real Node Inspector CLI/MCP
Linux arm64 (aarch64) ubuntu-24.04-arm Installed package and real Node Inspector CLI/MCP
macOS 15 arm64 macos-15 Installed package and real Node Inspector CLI/MCP
macOS 15 x64 macos-15-intel Installed package and real Node Inspector CLI/MCP
Windows x64 windows-latest Curated capabilities, native controls, installed package and real Node Inspector CLI/MCP

Package lanes assert the actual Node platform/architecture before verification and record those values with the Node version. The four-host package matrix runs up to four jobs concurrently on independent hosted runners with explicit timeouts and per-runner Node heap/thread limits; Windows retains its curated capability lane. Package checks cover installation, CLI/MCP discovery, target-free analysis, configuration backups/recovery, Evidence and owned lifecycle; Inspector checks execute source-owned loopback targets and special filename cases.

POSIX package verification resolves the runner's configured npm cache before creating disposable client homes and passes that location to its child installs. This reuses the download cache restored by setup-node while keeping client configuration isolated. Installs retain their existing registry and integrity checks.

The Apple artifact CLI/MCP lane uses one macos-15 job to avoid adding demand for scarce macOS runner capacity. It runs the artifact-format and metadata verifiers, followed by one consolidated focused invocation containing all existing native and process-boundary regression paths.

The Linux source-test shards download a shared runtime and generated test metadata snapshot produced by the build job with npm run test:prepare. The snapshot includes dist/, generated skills, the MCP tool catalog, product catalog, and managed-conformance metadata. Each shard still installs the locked dependencies and runs its assigned coverage-enabled tests; it no longer regenerates the same metadata independently.

The build job also retrieves the merged timing report from the last successful default-branch push and includes that same report in each shard's snapshot. CI assigns discovered files by historical duration, with a deterministic fallback weight for new files, rather than dividing only by file count. Every discovered file still belongs to exactly one shard; worker limits, project ordering, coverage instrumentation, and aggregate thresholds remain unchanged. Missing or unusable history falls back to Vitest's default assignment. Only timing data is reused; persistent compile caches remain disabled.

These native baseline checks complement the Linux source-test shards and the separate Apple-artifact and real-provider lanes. Actual Hopper, Ghidra, IDA, browser and managed-tool claims require their corresponding verification lanes. Runner labels follow the GitHub hosted-runner reference.

Termux Android browser smoke test

The Real Termux browser verification workflow is manual-only while its emulator setup is being validated. After the workflow reaches the default branch, select a revision under Actions → Real Termux browser verification → Run workflow. It boots an Android 11 (API 30) x86_64 emulator, installs the checksum-verified Termux 0.118.3 APK, and lets the app initialize its bootstrap. The host stages the selected revision using git archive HEAD; no workstation build outputs or dependencies are copied into Android.

scripts/verify/termux/emulator.sh sends a RUN_COMMAND intent to Termux's app service. Root access is limited to staging files, sending the intent, and reading diagnostics; installation, compilation, Node, and Chromium run as Termux's app UID. The in-app script installs Termux's Android-linked Node distribution, checks its version against .nvmrc, pins npm from packageManager, runs npm ci, and builds with build:termux. Termux repository packages, including Chromium, are resolved at run time; their installed versions are logged.

The verifier leaves Playwright's cache environment overrides unset, requires a healthy public MCP doctor result, and captures URL plus DOM from a loopback HTTP fixture through capture-browser-scenario. It checks completed steps, the exact URL, a DOM marker, and reported browser cleanup. This lane covers one Android emulator and Termux combination, not physical ARM devices, Electron, or the complete desktop browser suite. A workflow definition alone does not establish passing Android coverage; inspect its execution receipt.

The termux-browser-diagnostics artifact retains the Termux command log, exit status, and Android logcat on success or failure. The emulator is disposable.

Developer commands

Use source feedback while editing, explicit boundary checks for the changed behavior, and complete CI evidence before merging.

Command Scope
npm run test:local Dirty source tests via the import graph, without build; explicit source paths run regardless of Git status
npm run test:focused -- PATH... Exact existing test files; compiled boundaries build first, and unmatched paths fail
npm run test:changed Source tests affected since the merge base with origin/main, including committed and dirty changes
npm run test:fast All domain, service, adapter, composition, conformance, and evaluation tests without build
npm run test:boundary Boundary, CLI boundary, MCP boundary, process boundary, and process-global projects
npm run test:mcp MCP boundary project
npm run test:acceptance Complete compiled CLI and MCP acceptance workflows
npm run test:watch Dirty source tests in watch mode, without build
npm run test:watch:all Changed tests from every project; builds at startup, so rebuild after production edits before relying on compiled tests
npm run check:changed Cached typecheck/lint and branch-related source feedback
npm run check:pr Opt-in complete local deterministic gate and generated-file checks
npm run docs:check Generated-document validation from current source and build outputs

For example:

npm run test:local -- src/config.test.ts
npm run test:focused -- tests/acceptance/applications/runtime.test.ts
npm run test:changed -- --base origin/main
npm run test:changed -- --dry-run

test:focused accepts exact repository-relative test file paths. Source-only paths do not build; boundary, acceptance, process-global, or unfamiliar tests/ paths build conservatively. Tests use Vitest concurrency and isolated workspaces; commands do not hold a broad test lock. Build and documentation writers retain checkout-local locks for their shared output files. Explicit selections do not use --changed or permit zero-test success. The dry-run option reports the chosen merge base, scope and build prerequisite without executing tests or building. A missing Git base reports how to fetch it or select another revision.

Source typechecking needs no compiled runtime or generated test metadata. npm run build:cached compiles the CLI/MCP runtime and its packaged skill. Turbo's strict task environment preserves caller-selected GOMAXPROCS, GOMEMLIMIT, and UV_THREADPOOL_SIZE for resource-constrained builds and checks. These execution controls do not change artifact content, so changing their values retains cache reuse. They control individual runtimes; monitor the complete process family separately when enforcing a total memory budget. npm run test:prepare also generates the MCP contract test catalog, product catalog, and portable managed evidence used by the complete suite. Focused tests prepare those extra outputs only when their selected files consume them. The MCP test catalog is JSON in .cache/mcp-tool-catalog.json; the tracked test loader supplies its types without making generation a source-check prerequisite. Run npm run mcp-catalog:generate to refresh it independently. Documentation generation and managed conformance run through their separate docs:generate and evidence:generate task graphs.

Changed selection can miss runtime registration, generated data, shell entrypoints, bridges, or other relationships absent from the import graph. An empty changed selection means no tests were selected, not verified correctness. Select relevant boundary files and real-provider lanes explicitly. Source projects also contain large capacity regressions; test:fast promises no compiled-runtime prerequisite, not a fixed time budget.

Routine iterations and rebases need focused regressions and relevant checks. Before handing off a PR, run npm run check and generated-document checks when applicable; CI owns the full suite and coverage. Use the full local gate for broad changes or diagnosing CI, rather than after every edit. Package/install changes additionally need package verification; provider changes need actual provider evidence.

Vitest projects use up to two workers, bounded by host parallelism. process-boundary runs later with serial files because process-tree sampling shares host resources. Acceptance and process-global cases use isolated forks without serial scheduling. Pure domain/service, composition, conformance, evaluation, CLI boundary, and MCP boundary projects share module graphs. Their tests avoid process-global state and keep resource cleanup test-scoped. Each MCP session still owns its resources, and CLI cases retain fresh command instances or subprocesses. CLI module mocks belong in tests/process-global/ so they keep per-file isolation. Product catalog verification stays with the isolated filesystem boundary tests because it loads both source and compiled module graphs. Adapter, other boundary, acceptance, and process-global projects retain per-file isolation. See vitest.config.ts for current settings.

The reused CLI boundary workers cap V8 old space at 384 MiB, based on their profiled allocation and cleanup behavior. This encourages collection between files; spawned REA processes retain their own heap settings.

Vitest worker limits do not bound subprocesses launched within a test. Bound those batches separately so repeated CLI validation cannot oversubscribe the runner. Preserve simultaneous launches where concurrency is itself under test.

Build/documentation writers hold checkout-local output locks. Tests do not hold a broad command lock; check:pr finishes tests before document validation.

Vitest's persistent compile cache remains disabled. The cliTest fixture shares a temporary Node compile cache among its fresh CLI subprocesses. The worker removes it after test-scoped processes finish. Explicit child environment settings take precedence; children requesting NODE_V8_COVERAGE do not receive the fixture cache. These subprocesses are outside aggregate Vitest coverage; subprocess coverage attachment remains disabled. Cold-start import checks launch independently without this fixture cache.

The Apple artifact verifier step shares a job-temporary bytecode cache among its fresh CLI/MCP children via REA_ARTIFACT_NODE_COMPILE_CACHE. The artifact helper maps this setting to child NODE_COMPILE_CACHE; npm, Turbo, and nested Vitest runs do not inherit that setting. Explicit NODE_COMPILE_CACHE takes precedence, and coverage collection or NODE_DISABLE_COMPILE_CACHE=1 disables the helper's cache. The cache is discarded with the CI job.

To evaluate persistent compile caching for repeated local runs, opt in for both cold and warm measurements with an isolated cache:

NODE_COMPILE_CACHE=.cache/node-compile npm test

Do not report the warm result as a cold-suite improvement, and do not enable the cache in coverage or benchmark CI without first showing that its instrumentation remains equivalent.

Coverage and timing

CI owns coverage; aggregate and domain/contract thresholds are maintained in vitest.config.ts. They are glob-specific and never updated automatically. Coverage does not replace named boundary or real-provider scenarios.

CI runs four native Vitest shards without retries. Each shard emits a blob report; the merge job produces aggregate coverage plus JUnit and JSON timing reports, uploads them together, and writes the slowest files to the workflow summary. Static checks, documentation, build, package verification, and real-system lanes remain separate jobs so one kind of evidence cannot stand in for another.

The build job saves portable Turbo outputs across runs. Eight Linux real-system lanes restore this cache without waiting for the build job. The shared action derives its key from test:prepare --dry=json task hashes, runner OS and architecture, and .npmrc; test-only changes can reuse compiled outputs. Every build and verification command still runs, so Turbo checks task inputs and rebuilds misses. There is no prefix fallback to accumulate older build outputs.

Native package jobs reuse the same workflow's Turbo cache for portable compiled JavaScript, skills, and catalogs. They retain the normal build and prepack commands, which restore matching outputs or rebuild on a cache miss. Native artifacts, dependency installs, package archives, and runtime verification remain host-local. Keep platform-dependent outputs out of this shared cache.

The PR acceptance target is a median npm run check:pr wall time below three minutes across three warm-build runs on the benchmark host. Keep Vitest caches cold unless separately identified. A PR that touches packaging or real-system behavior requires the applicable verify:* lanes; packaging and installation changes also require npm run verify:package. The full-gate benchmark measures that explicit lane, not the routine iteration requirement.

Apple native metadata and UI

npm run verify:apple-dispatch compiles the Objective-C fixture (classes, protocols, a property and an NSString category) and the Swift conformance/vtable fixture.

  • Each fixture is linked with legacy LC_DYLD_INFO binds and with chained fixups; on Apple silicon the ObjC fixture is also built as arm64e, which uses authenticated pointers.
  • The lane inspects the bytes and repeats after stripping local symbols.
  • It requires the bound NSObject superclass, the external category, and the matching pointer_fixups coverage. It requires macOS and the host Xcode toolchain; targets are not executed. npm run verify:native-ui launches exactly one source-owned fixture window and requires successful selected-window capture and selected actions. An OS permission denial fails the positive lane. npm run verify:native-ui:permissions allows a host-permission-boundary-only result and explicitly reports positive_e2e: false; it must not be reported as capture/action proof. Both commands reject a changed executable digest and clean up the fixture process and helper. These lanes require an interactive macOS desktop. See native investigation

npm run verify:native-calls needs macOS with Command Line Tools (clang, lldb, codesign, nm) and actual debugger access to the owned fixtures. The lane reports Developer Mode status without treating it as proof of access or denial, and does not change host settings. LLDB and target entitlements establish the tested permission boundary. It compiles tests/conformance/native/calls.m and runs observe-native-calls through the CLI and stdio MCP. It checks:

  • the receiver class, selector and argument registers of every entry, and that breakpoint addresses equal nm's symbol addresses;
  • overlapping breakpoint selections retain each selection's hits and share the aggregate event limit; callback capture reads the stopped frame before LLDB resumes, avoiding stale frame metadata from asynchronous stop events;
  • captured stdout and an environment override, plus a 2 MiB flood on each output stream with bounded retained prefixes and exact drained-byte counts;
  • the event-limit and duration outcomes, with the process confirmed gone;
  • that a hardened-runtime copy is refused with debugger-attach-denied, and that the same copy signed with get-task-allow is traced.

Firmware adapters

Linux-owned provider launches require a procps-compatible ps on REA's PATH so ownership can be inspected before launching a child. A missing or incompatible command fails with an actionable capability error before the launch.

npm run fixtures:firmware uses existing Python 3 and a host C compiler to make an ignored gzip/USTAR firmware fixture and independent offset/hash oracle. npm run verify:firmware requires caller-supplied Binwalk 3.1.0, Unblob 26.6.4 and util-linux prlimit on Linux, with 7z on PATH for the gzip/USTAR fixture's Unblob extractor. The provider also accepts other 3.1.x and 26.6.x builds and reports them as unverified; this lane proves the audited releases. It verifies CLI/MCP parity, selected ranges, unknown chunks, depth limits and extracted child digests. The optional REA_FIRMWARE_VERIFY_EXT4=1 lane requires existing mke2fs/debugfs; the separate REA_FIRMWARE_VERIFY_GHIDRA=1 lane checks a selected host ELF through real Ghidra. Neither optional toolchain is a base-lane prerequisite. See firmware analysis for limits and unverified formats.

JavaScript source recovery

Build once with npm run build:cached, then run npm run verify:javascript:recovery. This focused lane requires Linux x64, util-linux prlimit, REA_WAKARU_COMMAND pointing to the official Wakaru 1.14.0 Linux x64 binary, and REA_JAVASCRIPT_FIXTURE_TOOLS pointing to an isolated npm prefix containing esbuild 0.25.10 and webpack 5.101.3. No global installation is required. The lane compiles source-owned fixtures, exercises CLI and stdio MCP, verifies published bytes and provenance, feeds recovered modules into existing analysis, and compares a finite set of known fixture results. A webpack fixture with separate runtime, shared, entry and lazy chunks verifies multi-input recovery through both surfaces, per-file provenance and resolved cross-chunk imports in the downstream application graph. The provider also accepts other ^1.13.0 releases and reports them as unverified; this lane proves the audited release. It does not establish arbitrary recovered-application equivalence. CI installs these prerequisites only in .github/workflows/real-javascript-recovery.yml; the existing real-browser lane uses real Chrome for browser capture and website workflows.

Captured website source-map lane

npm run verify:browser:source-maps checks actual Chromium capture/export and source-map point tracing through CLI and stdio MCP. It requires absolute REA_BROWSER_EXECUTABLE and REA_WEB_SOURCE_MAP_COMPILER pointing to esbuild 0.25.10's lib/main.js in a caller-owned isolated installation. Preflight reports missing prerequisites for this lane. The compiler is used only to generate the source-owned fixture. No Hopper, Ghidra or application dependency installation is required.

REA_BROWSER_EXECUTABLE=/absolute/path/to/chromium \
REA_WEB_SOURCE_MAP_COMPILER=/absolute/path/to/fixture-tools/node_modules/esbuild/lib/main.js \
npm run verify:browser:source-maps

scripts/verify-browser-source-maps.mjs /absolute/path/to/installed/rea.mjs checks an installed package after building the verifier dependencies. The separate conditional real-web-source-map CI job supplies Chrome and an isolated pinned fixture compiler; static/unit checks do not acquire a browser. See the source location guide for the verified decoder profile.

JavaScript large-output lane

npm run verify:javascript:output exercises the CLI JSON result surface beyond the running Node engine's single-string limit. Shared input leaves keep the fixture's graph small; the verifier writes one temporary output file, checks its complete byte count and an independent digest, then removes it. It requires only Node and the built REA runtime, with space for the output plus a 1 GiB reserve. Run npm run verify:javascript:output -- jsonl for compact JSONL coverage. This opt-in lane is separate from routine tests and the canonical hash check npm run verify:javascript:digests. It verifies serialization rather than an arbitrary third-party application's parsing cost or MCP client capacity.

Website runtime attribution lane

npm run verify:browser:runtime uses caller-supplied REA_BROWSER_EXECUTABLE and an owned synthetic site/profile. It exercises public CLI and stdio MCP for precise execution and native listener source locations, including actual armed progress, Unicode/CRLF digests, repeated source URLs with distinct script IDs, zero branches and function-only unknowns on repeated coverage, request initiators and an externally owned page that remains open.

An optional entrypoint argument to scripts/verify-browser-runtime.mjs runs the same checks through an isolated installed package. The conditional real-web-runtime CI job runs only for relevant changes and needs no fixture compiler. Ordinary unit/static gates acquire no browser. See website runtime attribution for effects, resource bounds and coverage limits.

Offline binary layout

npm run verify:binary:layout requires Linux x64, GCC/binutils, absolute REA_PWNTOOLS_PYTHON with pwntools 4.15.0/pyelftools 0.33/Unicorn 2.1.2 and absolute REA_VERIFY_STRACE_COMMAND. It compiles ephemeral source-owned ELF fixtures and checks public CLI/MCP, lossless addresses/names, file ranges, mitigation inferences, malformed/unsupported input, original file hashes and released process ownership. Exec syscall tracing must identify only the declared Node/Python launchers; no target binary is executed. Core/debugger claims need separate verification lanes. Pass an installed package entrypoint as the script's first argument to verify packaging independently of the checkout. The valid SHN_XINDEX fixture has 65,281 full section rows; CLI is checked in the ordinary lane. Its large MCP transfer is opt-in with REA_VERIFY_LARGE_ELF_MCP=1 (or the workflow dispatch large_mcp input), an explicit 256 MiB SDK receive buffer and five-minute request timeout. Ordinary MCP fixtures retain the pinned SDK defaults.

Offline EVM interface

npm run verify:evm:interface requires Linux x64, an absolute REA_VERIFY_STRACE_COMMAND, caller-supplied util-linux prlimit and REA_VERIFY_SOLC_MODULE selecting the absolute module path for solc 0.8.30. It compiles source-owned plain/optimized/via-IR Cancun fixtures in private storage and checks actual CLI/MCP selector evidence, raw/hex identity, unknowns, malformed carriers and independent cleanup. An optional positional entrypoint verifies a fresh installed package. It acquires no engine, compiler or chain dependency and does not execute a contract on a chain.

Recorded crash evidence

npm run verify:recorded:crash is a separate Linux x64 lane. It requires GCC, GDB, absolute REA_PWNTOOLS_PYTHON with the offline ELF profile above, REA_PWNDBG_GDBINIT and REA_PWNDBG_VENV_PATH with unchanged pwndbg 2026.09.15, and REA_VERIFY_STRACE_COMMAND. Its disposable CI runner installs GDB, checks out the exact upstream commit and installs its frozen lockfile in isolated runner storage. No developer host configuration or core-pattern setting changes.

Fixture generation explicitly runs an owned source-built two-thread program under GDB to create a recording. Subsequent public CLI/MCP inspection verifies lossless high registers, signed signals, note source bytes, malformed/missing notes, unfamiliar owners, optional core-only mapping context and actionable missing/unsupported plugin errors. A historical-PID collision fixture references an owned live sentinel; inspection syscall traces reject process attach/memory access, provider lookups of that PID's /proc files and attempted Internet sockets. Traces admit the observed upstream startup helpers (iconv -l, the selected checkout's Git version lookup) and REA ownership inspection separately from target execution. This is fixture evidence, not a sandbox claim. Inputs remain unchanged and the sentinel must stay alive; owned cleanup and empty verifier descendants are required. Pass an installed package entrypoint as the script's first argument for package coverage.

Agent evaluation and conformance records

Evaluate native, JavaScript, managed and browser investigation tasks through a real local Codex CLI with:

npm run verify:agent

The lane runs scenarios sequentially in standalone Codex (--no-daemon) with a disposable HOME and CODEX_HOME under ~/.cache/rea-agent-evaluations/. Set REA_AGENT_EVAL_ROOT to choose another parent directory outside the OS temporary directory: current Codex refuses helper aliases under /tmp, preventing shell commands and skill reads even if MCP works. Each owned run directory is removed after verification unless fixture retention is requested. The lane installs the packaged skill and MCP registration through public rea setup, rather than copying a skill into a project that also inherits the caller's skills and configuration. Only authentication is copied, when present, and it is removed even with REA_AGENT_EVAL_KEEP_FIXTURES=true. No user configuration, rules, plugins or history is copied; shell snapshots and additional agents are disabled. Existing API-key environment authentication remains available. The native missing-provider scenario preapproves only open_binary and close_binary in the disposable client so it reaches REA's actual provider boundary instead of stopping at Codex approval. This does not validate a real native engine or authorize arbitrary runtime execution.

Select a smaller run with REA_AGENT_EVAL_SCENARIOS, a comma-separated list of scenario IDs. asar-zh exercises the packaged desktop rubric from a Chinese request. javascript-module-view exercises summary analysis, module paging and item inspection in a tree with unrelated modules; its routing and text checks remain heuristic, so review the retained transcript for selector choice and coverage claims. Set REA_CODEX_CLI to an installed Codex executable and REA_AGENT_EVAL_MODEL to an account-supported model. Use an external resource limit for the whole process tree when testing on a constrained host; a Node heap limit alone does not bound client and child-process memory.

missing-target leaves desktop fixtures visible without selecting one in the request. It requires zero REA calls and a structured request for an app name or artifact path. Its expectedFirstTool is null, and targetClarificationPassed grades that limited routing outcome independently of analysis prose heuristics. Analyzing a nearby example cannot pass this case.

The packaged desktop and parser-comparison scenarios now assess a closed set of known fixture claims. Each final answer must be a strict JSON manifest containing exactly the requested claim IDs, values, Evidence IDs, and Evidence authority and confidence. Golden values are not included in the prompt. The evaluator checks exact values against both the fixture oracle and authenticated successful REA results; an Evidence ID by itself is insufficient. Missing, duplicate, extra, contradictory, incorrectly sourced, or unsupported claims fail the scenario.

The desktop rubric covers the exposed bridge API and members, renderer and main IPC operations, and the resolved preload. It binds its Evidence to the exact packaged artifact path and SHA-256. The parser rubric checks the precise heading depth addition, discriminant, complete comparison counts, and the explicit limit on runtime semantics. Its comparison must link to the two delivered source analyses and use the requested module and export selectors. Boundary tests package the same source-owned desktop fixture and compare the same source-owned parser files through actual MCP tool results, then verify that fabricated answers fail.

Per-scenario factualCorrectness is passed, failed, or not_assessed within the configured_fixture_claims scope. A pass establishes only the configured claims, not unrestricted factual correctness. Native, managed, browser, navigation-context, and address-context scenarios have no factual rubric and remain not_assessed; the managed workflow requires both artifact and member inspection to answer its type and entry-point question.

The closed factual scenarios request complete producer Evidence on the initial analysis so their source selectors can authenticate all configured facts. Summary/page/item workflows are exercised separately by javascript-module-view; its heuristic gate does not establish full-producer factual correctness.

All scenarios retain routing, workflow, validation, repetition, process-exit, and token-use gates. Configured factual scenarios additionally require a factual pass. Their answer-text heuristics are diagnostic and do not affect the gate. Validation counts include complete text-only REA invalid_request envelopes and failed SDK calls whose arguments violate the named current REA contract. Client approval rejections and provider unavailability remain distinct failures; truncated JSON previews are not parsed into invented error codes. Repeated binary_session observations separated by a successful open_binary or close_binary are treated as lifecycle verification. An unchanged retry, including one after a failed lifecycle call, still counts as repetition. Scenarios without a rubric retain the legacy text-heuristic gate. answerTermCoverageMet checks case-insensitive substrings, epistemicCuePresent checks keywords, and finalCitesEvidence checks only an ID's presence. These metrics can still accept fabricated prose and must not be read as factual assessments. Set REA_AGENT_EVAL_TRANSCRIPT_DIR to retain complete tool results and final answers for review.

Report schema version 4 adds negative routing scenarios with nullable expectedFirstTool and a targetClarificationPassed outcome. Consumers must allow an expected absence of REA calls; this does not make an analysis scenario pass without its required tools. Version 3 changed the top-level factualCorrectness from a constant string to an assessment summary with status, scope, and assessed/passed/failed/ not-assessed scenario counts. Scenario records include the factual assessment and configured claim IDs. The evaluation scope is routing_workflow_and_configured_fixture_claims. Update report consumers for these changes. The narrower heuristic field names introduced in version 2 remain: answerHeuristicsMet, epistemicCuePresent, answerTermCoverageMet, and requiredAnswerTermGroups.

Regenerate the managed conformance manifest and Evidence completion ledger from live verification results, or check them for drift:

npm run evidence:generate
npm run evidence:check

The records preserve unsupported and unverified coverage as explicit unknowns. Run the matching real-tool prerequisites described in this guide.

Offline WASM artifacts

npm run verify:wasm:artifact -- --require-tools uses an absolute REA_WABT_BIN_DIRECTORY containing WABT 1.0.42, including wat2wasm for fixture generation. It installs nothing. Without configuration, the default invocation reports a named skip; --require-tools makes missing configuration fail.

The verifier creates and validates real modules with custom sections, multiple import/export kinds, escaped names, an empty module, and a bulk-memory DataCount section. It compares exact WAT, section output and selected/tool digests through the compiled CLI and real MCP SDK, validates advertised schemas, checks identical CLI/MCP Evidence IDs and session bundle readback, retains invalid/truncated module diagnostics, preserves all inputs and checks workspace cleanup. Pass an installed package's scripts/rea.mjs as the first argument to run the same lane against the package. npm run verify:wasm:package packs and installs REA into an isolated temporary prefix with lifecycle scripts disabled and runs the same required real-tool lane. It requires WABT configuration and makes no host tool installation.

Source-owned provider tests separately exercise cancellation, deadline propagation, malformed producer output and cleanup uncertainty; shared owned-command tests exercise actual subprocess timeout and cancellation.