Files
deepagentsjs/internal/eval-harness
github-actions[bot] 5af75fffbe chore: version packages (#726)
This PR was opened by the [Changesets
release](https://github.com/changesets/action) GitHub action. When
you're ready to do a release, you can merge this and the packages will
be published to npm automatically. If you're not ready to do a release
yet, that's fine, whenever you add more changesets to main, this PR will
be updated.


# Releases
## deepagents-acp@0.1.23

### Patch Changes

- Updated dependencies
[[`590c2a5`](https://github.com/langchain-ai/deepagentsjs/commit/590c2a5042473f096d5fac5ddbb4be96e2ace0f2)]:
  - deepagents@1.12.2
## deepagents@1.12.2

### Patch Changes

- [#723](https://github.com/langchain-ai/deepagentsjs/pull/723)
[`590c2a5`](https://github.com/langchain-ai/deepagentsjs/commit/590c2a5042473f096d5fac5ddbb4be96e2ace0f2)
Thanks
[@thushanth-bengre-langchain](https://github.com/thushanth-bengre-langchain)!
- fix(deepagents): prevent stack overflow in CompositeBackend grep/glob
on huge result sets, and add a grep match-count cap

`CompositeBackend` accumulated merged `ls`/`grep`/`glob` results with
`push(...entries)`, which passes every entry as a separate function
argument and overflows the call stack (RangeError: Maximum call stack
size exceeded) when a broad search over a large tree returns hundreds of
thousands of entries. Results are now accumulated with a plain loop, so
no result-set size can overflow the stack.

`grep` also gains an optional `maxCount` (backend) / `max_count` (tool)
cap, mirroring the Python SDK. When the cap is hit, results are flagged
`truncated: true` on `GrepResult`/`GlobResult` and the grep tool appends
a note telling the model to narrow the search. The cap defaults to 1000
via the `grepMaxCount` middleware option (set to `null` to disable).
`CompositeBackend` splits the budget across routed backends and
OR-propagates the `truncated` flag on merged results.
## @deepagents/evals@0.0.22

### Patch Changes

- Updated dependencies
[[`590c2a5`](https://github.com/langchain-ai/deepagentsjs/commit/590c2a5042473f096d5fac5ddbb4be96e2ace0f2)]:
  - deepagents@1.12.2

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-04 15:20:01 -04:00
..
2026-04-01 17:10:16 -07:00
2026-08-04 15:20:01 -04:00
2026-08-04 15:20:01 -04:00

@deepagents/evals

Generic eval harness for deepagents. Provides the runner interface, runner registry, trajectory parsing, and custom vitest matchers with LangSmith feedback integration.

Package exports

Export Description
@deepagents/evals Core harness — EvalRunner, registry functions, parseTrajectory, getFinalText, custom matchers
@deepagents/evals/deepagent registerDeepAgentRunner() — wires createDeepAgent into the generic runner interface
@deepagents/evals/setup Side-effect import that registers all concrete runners (sonnet-4-5, gpt-4.1, etc.)

Core concepts

EvalRunner

The central interface. A runner has a name, a run() method, and an extend() method:

interface EvalRunner {
  name: string;
  run(params: RunAgentParams): Promise<AgentTrajectory>;
  extend(overrides: Record<string, unknown>): EvalRunner;
}
  • run() takes invocation params only — { query, initialFiles? }.
  • extend() returns a new runner with agent configuration overrides (e.g. systemPrompt, tools, subagents) baked in. This keeps "what the agent is" separate from "what to ask it".

RunAgentParams

interface RunAgentParams {
  query: string;
  initialFiles?: Record<string, string>;
}

Pure invocation inputs. query is the user message; initialFiles seeds the agent's virtual file system.

AgentTrajectory

interface AgentTrajectory {
  steps: AgentStep[];
  files: Record<string, string>;
}

The output of every run() call. steps is an ordered list of agent turns (each containing an AIMessage action and any ToolMessage observations). files is the final file-system snapshot.

Runner registry

flowchart LR
    setup["setup.ts<br/><i>vitest setupFile</i>"]
    deepagent["deepagent.ts<br/>registerDeepAgentRunner()"]
    registry["index.ts<br/>registerRunner()"]
    resolve["resolveRunner(name)"]
    default["getDefaultRunner()<br/>reads EVAL_RUNNER env"]
    test["test file<br/>runner.run() / runner.extend()"]

    setup -- "calls for each model" --> deepagent
    deepagent -- "wraps in DeepAgentEvalRunner" --> registry
    default -- "looks up by name" --> registry
    resolve -- "looks up by name" --> registry
    test --> default
    test --> resolve
  1. setup.ts is loaded as a vitest setupFile. It calls registerDeepAgentRunner() for each model, which internally calls registerRunner().
  2. getDefaultRunner() reads the EVAL_RUNNER env var, looks up the runner by name, and returns it (cached).
  3. resolveRunner(name) is the lower-level lookup if you need a specific runner by name.

registerDeepAgentRunner

Bridges createDeepAgent to the generic EvalRunner interface:

registerDeepAgentRunner("sonnet-4-5", (config) =>
  createDeepAgent({
    ...config,
    model: new ChatAnthropic({ model: "claude-sonnet-4-5-20250929" }),
  }),
);

The factory receives optional overrides (from extend()) and must return an invokable agent. A default agent (no overrides) is built eagerly at registration time and reused across run() calls for performance. When extend() is called, a fresh agent is constructed with the overrides.

The DeepAgentEvalRunner handles:

  • Converting initialFiles strings into the FileData format expected by the agent's state backend
  • Generating a unique thread_id per invocation
  • Calling ls.logOutputs() for LangSmith experiment tracking
  • Parsing the raw LangGraph result into an AgentTrajectory

Custom vitest matchers

Imported automatically when you import from @deepagents/evals. Each matcher also logs LangSmith feedback (scores) so results appear in the experiment dashboard.

Matcher Description
toHaveAgentSteps(n) Trajectory has exactly n steps
toHaveToolCallRequests(n) Total tool-call count equals n
toHaveToolCallInStep(step, match) Step (1-indexed) contains a tool call matching { name, argsContains?, argsEquals? }
toHaveFinalTextContaining(text, caseInsensitive?) Last step's text content includes text

Adding a new runner

To add support for a new model, add a registerDeepAgentRunner call in src/setup.ts:

registerDeepAgentRunner("my-model", (config) =>
  createDeepAgent({ ...config, model: new ChatMyProvider({ model: "my-model" }) }),
);

Then run evals with EVAL_RUNNER=my-model.

To add a completely different runner backend (not deepagents), implement the EvalRunner interface directly and call registerRunner(), or use the runner directly inside of the eval suite.

Development

pnpm build       # Build with tsdown
pnpm typecheck   # Type-check without emitting