Files
deepagentsjs/internal/eval-harness
github-actions[bot] a9e1ba1b89 chore: version packages (#605)
This PR was opened by the [Changesets
release](https://github.com/changesets/action) GitHub action. When
you're ready to do a release, you can merge this and the packages will
be published to npm automatically. If you're not ready to do a release
yet, that's fine, whenever you add more changesets to main, this PR will
be updated.


# Releases
## deepagents-acp@0.1.15

### Patch Changes

- Updated dependencies
[[`7c4a11e`](https://github.com/langchain-ai/deepagentsjs/commit/7c4a11eacc11c3720b70d802068300ac3b4d8651)]:
  - deepagents@1.10.5
## deepagents@1.10.5

### Patch Changes

- [#598](https://github.com/langchain-ai/deepagentsjs/pull/598)
[`7c4a11e`](https://github.com/langchain-ai/deepagentsjs/commit/7c4a11eacc11c3720b70d802068300ac3b4d8651)
Thanks [@christian-bromann](https://github.com/christian-bromann)! -
refactor(stream): use langchain `run.subagents` instead of bespoke
transformer

Remove deepagents' custom `createSubagentTransformer` and rely on the
native
  subagent stream that `createAgent` registers (langchain#37739). Keep
`DeepAgentRunStream` as a compile-time overlay that narrows
`run.subagents` to
declared subagent specs. Update streaming tests for `cause` and
per-subagent
  message coverage.
## @langchain/quickjs@0.5.1

### Patch Changes

- [#602](https://github.com/langchain-ai/deepagentsjs/pull/602)
[`204cb27`](https://github.com/langchain-ai/deepagentsjs/commit/204cb27414c82c34a0c681c7e5a5336e1834f058)
Thanks [@colifran](https://github.com/colifran)! - chore(quickjs):
refine dynamic subagent prompt to trigger on workflow keyword and to
improve iterative eval behavior

- [#604](https://github.com/langchain-ai/deepagentsjs/pull/604)
[`0971007`](https://github.com/langchain-ai/deepagentsjs/commit/0971007ab2491e73eb1a78c5426ab934470d5620)
Thanks [@colifran](https://github.com/colifran)! - fix(quickjs): unwrap
Command/ToolMessage envelopes from tool and subagent results
## @deepagents/evals@0.0.14

### Patch Changes

- Updated dependencies
[[`7c4a11e`](https://github.com/langchain-ai/deepagentsjs/commit/7c4a11eacc11c3720b70d802068300ac3b4d8651)]:
  - deepagents@1.10.5

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-18 09:38:18 -07:00
..
2026-04-01 17:10:16 -07:00
2026-06-18 09:38:18 -07:00
2026-06-18 09:38:18 -07:00

@deepagents/evals

Generic eval harness for deepagents. Provides the runner interface, runner registry, trajectory parsing, and custom vitest matchers with LangSmith feedback integration.

Package exports

Export Description
@deepagents/evals Core harness — EvalRunner, registry functions, parseTrajectory, getFinalText, custom matchers
@deepagents/evals/deepagent registerDeepAgentRunner() — wires createDeepAgent into the generic runner interface
@deepagents/evals/setup Side-effect import that registers all concrete runners (sonnet-4-5, gpt-4.1, etc.)

Core concepts

EvalRunner

The central interface. A runner has a name, a run() method, and an extend() method:

interface EvalRunner {
  name: string;
  run(params: RunAgentParams): Promise<AgentTrajectory>;
  extend(overrides: Record<string, unknown>): EvalRunner;
}
  • run() takes invocation params only — { query, initialFiles? }.
  • extend() returns a new runner with agent configuration overrides (e.g. systemPrompt, tools, subagents) baked in. This keeps "what the agent is" separate from "what to ask it".

RunAgentParams

interface RunAgentParams {
  query: string;
  initialFiles?: Record<string, string>;
}

Pure invocation inputs. query is the user message; initialFiles seeds the agent's virtual file system.

AgentTrajectory

interface AgentTrajectory {
  steps: AgentStep[];
  files: Record<string, string>;
}

The output of every run() call. steps is an ordered list of agent turns (each containing an AIMessage action and any ToolMessage observations). files is the final file-system snapshot.

Runner registry

flowchart LR
    setup["setup.ts<br/><i>vitest setupFile</i>"]
    deepagent["deepagent.ts<br/>registerDeepAgentRunner()"]
    registry["index.ts<br/>registerRunner()"]
    resolve["resolveRunner(name)"]
    default["getDefaultRunner()<br/>reads EVAL_RUNNER env"]
    test["test file<br/>runner.run() / runner.extend()"]

    setup -- "calls for each model" --> deepagent
    deepagent -- "wraps in DeepAgentEvalRunner" --> registry
    default -- "looks up by name" --> registry
    resolve -- "looks up by name" --> registry
    test --> default
    test --> resolve
  1. setup.ts is loaded as a vitest setupFile. It calls registerDeepAgentRunner() for each model, which internally calls registerRunner().
  2. getDefaultRunner() reads the EVAL_RUNNER env var, looks up the runner by name, and returns it (cached).
  3. resolveRunner(name) is the lower-level lookup if you need a specific runner by name.

registerDeepAgentRunner

Bridges createDeepAgent to the generic EvalRunner interface:

registerDeepAgentRunner("sonnet-4-5", (config) =>
  createDeepAgent({
    ...config,
    model: new ChatAnthropic({ model: "claude-sonnet-4-5-20250929" }),
  }),
);

The factory receives optional overrides (from extend()) and must return an invokable agent. A default agent (no overrides) is built eagerly at registration time and reused across run() calls for performance. When extend() is called, a fresh agent is constructed with the overrides.

The DeepAgentEvalRunner handles:

  • Converting initialFiles strings into the FileData format expected by the agent's state backend
  • Generating a unique thread_id per invocation
  • Calling ls.logOutputs() for LangSmith experiment tracking
  • Parsing the raw LangGraph result into an AgentTrajectory

Custom vitest matchers

Imported automatically when you import from @deepagents/evals. Each matcher also logs LangSmith feedback (scores) so results appear in the experiment dashboard.

Matcher Description
toHaveAgentSteps(n) Trajectory has exactly n steps
toHaveToolCallRequests(n) Total tool-call count equals n
toHaveToolCallInStep(step, match) Step (1-indexed) contains a tool call matching { name, argsContains?, argsEquals? }
toHaveFinalTextContaining(text, caseInsensitive?) Last step's text content includes text

Adding a new runner

To add support for a new model, add a registerDeepAgentRunner call in src/setup.ts:

registerDeepAgentRunner("my-model", (config) =>
  createDeepAgent({ ...config, model: new ChatMyProvider({ model: "my-model" }) }),
);

Then run evals with EVAL_RUNNER=my-model.

To add a completely different runner backend (not deepagents), implement the EvalRunner interface directly and call registerRunner(), or use the runner directly inside of the eval suite.

Development

pnpm build       # Build with tsdown
pnpm typecheck   # Type-check without emitting