Files
deepagentsjs/internal/eval-harness
github-actions[bot] ff8cc5c4b4 chore: version packages (#607)
This PR was opened by the [Changesets
release](https://github.com/changesets/action) GitHub action. When
you're ready to do a release, you can merge this and the packages will
be published to npm automatically. If you're not ready to do a release
yet, that's fine, whenever you add more changesets to main, this PR will
be updated.


# Releases
## @langchain/quickjs@0.6.0

### Minor Changes

- [#606](https://github.com/langchain-ai/deepagentsjs/pull/606)
[`3c8f8b2`](https://github.com/langchain-ai/deepagentsjs/commit/3c8f8b2ea6f0204353c63bf49c3fdc6655bf1069)
Thanks [@colifran](https://github.com/colifran)! - chore(quickjs):
disallow task as a configurable ptc tool
## deepagents-acp@0.1.16

### Patch Changes

- Updated dependencies
[[`d7ecab2`](https://github.com/langchain-ai/deepagentsjs/commit/d7ecab2d9f9d41321a043eed6edc3366a1381a67),
[`1a2b2df`](https://github.com/langchain-ai/deepagentsjs/commit/1a2b2df5528f0f61870b054fff8291355f6a2a0b),
[`42f34b6`](https://github.com/langchain-ai/deepagentsjs/commit/42f34b65ededf4a1fbf3cd4bbff486ddfeb320e9),
[`0ae10d7`](https://github.com/langchain-ai/deepagentsjs/commit/0ae10d7e26c84203a5273939c9ad7a9c8c8661c6)]:
  - deepagents@1.10.6
## deepagents@1.10.6

### Patch Changes

- [#608](https://github.com/langchain-ai/deepagentsjs/pull/608)
[`d7ecab2`](https://github.com/langchain-ai/deepagentsjs/commit/d7ecab2d9f9d41321a043eed6edc3366a1381a67)
Thanks [@aolsenjazz](https://github.com/aolsenjazz)! - fix(deepagents):
forward subagent results as text

Fixed a 400 `invalid_request_error` that occurred when a subagent used
an Anthropic server-side tool (web search, web fetch, or code
execution): the subagent's `server_tool_use`/`*_tool_result` blocks were
forwarded to the parent agent as `tool_result` content, which the API
rejects. Subagent results are now passed back to the parent as their
text content (matching the Python implementation), which resolves the
error and also handles a trailing empty `end_turn` message.

- [#656](https://github.com/langchain-ai/deepagentsjs/pull/656)
[`1a2b2df`](https://github.com/langchain-ai/deepagentsjs/commit/1a2b2df5528f0f61870b054fff8291355f6a2a0b)
Thanks [@colifran](https://github.com/colifran)! - fix(deepagents):
default unknown file extensions to text/plain

- [#611](https://github.com/langchain-ai/deepagentsjs/pull/611)
[`42f34b6`](https://github.com/langchain-ai/deepagentsjs/commit/42f34b65ededf4a1fbf3cd4bbff486ddfeb320e9)
Thanks [@aolsenjazz](https://github.com/aolsenjazz)! - feat(deepagents):
add bedrockPromptCachingMiddleware to default stack

Add bedrockPromptCachingMiddleware to default middleware stack. This
automatically opts-in to Bedrock prompt caching for Nova and Anthropic
models

- [#613](https://github.com/langchain-ai/deepagentsjs/pull/613)
[`0ae10d7`](https://github.com/langchain-ai/deepagentsjs/commit/0ae10d7e26c84203a5273939c9ad7a9c8c8661c6)
Thanks [@christian-bromann](https://github.com/christian-bromann)! -
fix(deepagents): declare LangChain runtime packages as peer dependencies

Move `@langchain/core`, `@langchain/langgraph`,
`@langchain/langgraph-sdk`, and
`langchain` from `dependencies` to `peerDependencies`, and also declare
`@langchain/langgraph-checkpoint` as a peer (its
`BaseCheckpointSaver`/`BaseStore`
types are part of the public API), so they resolve to a single shared
instance in
  the consumer's tree. Previously they were bundled as regular
dependencies, which let a consumer end up with two copies of
`@langchain/core`
(e.g. `1.2.0` vs `1.2.1`). Because these packages ship classes with
private/
protected fields, the duplicate copies are treated as nominally distinct
types,
producing errors like passing a `ChatOpenAI` model to `createDeepAgent`
or a
compiled graph to the local protocol helpers. As peers, the app controls
the
version and bumping `@langchain/core` no longer requires a `deepagents`
release.
## @langchain/daytona@0.2.1

### Patch Changes

- [#614](https://github.com/langchain-ai/deepagentsjs/pull/614)
[`5b462f2`](https://github.com/langchain-ai/deepagentsjs/commit/5b462f2ab2400fba52906e8f1485a500ec5b6e17)
Thanks [@christian-bromann](https://github.com/christian-bromann)! -
Replace deprecated `@daytonaio/sdk` dependency with `@daytona/sdk`.
## @langchain/modal@0.1.5

### Patch Changes

- [#657](https://github.com/langchain-ai/deepagentsjs/pull/657)
[`5f93c11`](https://github.com/langchain-ai/deepagentsjs/commit/5f93c114759399002a3981a93e3df13981514b1d)
Thanks [@colifran](https://github.com/colifran)! - fix(modal): modal low
level file handle api `sandbox.open` has been removed
## @deepagents/evals@0.0.15

### Patch Changes

- Updated dependencies
[[`d7ecab2`](https://github.com/langchain-ai/deepagentsjs/commit/d7ecab2d9f9d41321a043eed6edc3366a1381a67),
[`1a2b2df`](https://github.com/langchain-ai/deepagentsjs/commit/1a2b2df5528f0f61870b054fff8291355f6a2a0b),
[`42f34b6`](https://github.com/langchain-ai/deepagentsjs/commit/42f34b65ededf4a1fbf3cd4bbff486ddfeb320e9),
[`0ae10d7`](https://github.com/langchain-ai/deepagentsjs/commit/0ae10d7e26c84203a5273939c9ad7a9c8c8661c6)]:
  - deepagents@1.10.6

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Colin Francis <colin.francis@langchain.dev>
2026-07-08 09:16:25 -07:00
..
2026-04-01 17:10:16 -07:00
2026-07-08 09:16:25 -07:00
2026-07-08 09:16:25 -07:00

@deepagents/evals

Generic eval harness for deepagents. Provides the runner interface, runner registry, trajectory parsing, and custom vitest matchers with LangSmith feedback integration.

Package exports

Export Description
@deepagents/evals Core harness — EvalRunner, registry functions, parseTrajectory, getFinalText, custom matchers
@deepagents/evals/deepagent registerDeepAgentRunner() — wires createDeepAgent into the generic runner interface
@deepagents/evals/setup Side-effect import that registers all concrete runners (sonnet-4-5, gpt-4.1, etc.)

Core concepts

EvalRunner

The central interface. A runner has a name, a run() method, and an extend() method:

interface EvalRunner {
  name: string;
  run(params: RunAgentParams): Promise<AgentTrajectory>;
  extend(overrides: Record<string, unknown>): EvalRunner;
}
  • run() takes invocation params only — { query, initialFiles? }.
  • extend() returns a new runner with agent configuration overrides (e.g. systemPrompt, tools, subagents) baked in. This keeps "what the agent is" separate from "what to ask it".

RunAgentParams

interface RunAgentParams {
  query: string;
  initialFiles?: Record<string, string>;
}

Pure invocation inputs. query is the user message; initialFiles seeds the agent's virtual file system.

AgentTrajectory

interface AgentTrajectory {
  steps: AgentStep[];
  files: Record<string, string>;
}

The output of every run() call. steps is an ordered list of agent turns (each containing an AIMessage action and any ToolMessage observations). files is the final file-system snapshot.

Runner registry

flowchart LR
    setup["setup.ts<br/><i>vitest setupFile</i>"]
    deepagent["deepagent.ts<br/>registerDeepAgentRunner()"]
    registry["index.ts<br/>registerRunner()"]
    resolve["resolveRunner(name)"]
    default["getDefaultRunner()<br/>reads EVAL_RUNNER env"]
    test["test file<br/>runner.run() / runner.extend()"]

    setup -- "calls for each model" --> deepagent
    deepagent -- "wraps in DeepAgentEvalRunner" --> registry
    default -- "looks up by name" --> registry
    resolve -- "looks up by name" --> registry
    test --> default
    test --> resolve
  1. setup.ts is loaded as a vitest setupFile. It calls registerDeepAgentRunner() for each model, which internally calls registerRunner().
  2. getDefaultRunner() reads the EVAL_RUNNER env var, looks up the runner by name, and returns it (cached).
  3. resolveRunner(name) is the lower-level lookup if you need a specific runner by name.

registerDeepAgentRunner

Bridges createDeepAgent to the generic EvalRunner interface:

registerDeepAgentRunner("sonnet-4-5", (config) =>
  createDeepAgent({
    ...config,
    model: new ChatAnthropic({ model: "claude-sonnet-4-5-20250929" }),
  }),
);

The factory receives optional overrides (from extend()) and must return an invokable agent. A default agent (no overrides) is built eagerly at registration time and reused across run() calls for performance. When extend() is called, a fresh agent is constructed with the overrides.

The DeepAgentEvalRunner handles:

  • Converting initialFiles strings into the FileData format expected by the agent's state backend
  • Generating a unique thread_id per invocation
  • Calling ls.logOutputs() for LangSmith experiment tracking
  • Parsing the raw LangGraph result into an AgentTrajectory

Custom vitest matchers

Imported automatically when you import from @deepagents/evals. Each matcher also logs LangSmith feedback (scores) so results appear in the experiment dashboard.

Matcher Description
toHaveAgentSteps(n) Trajectory has exactly n steps
toHaveToolCallRequests(n) Total tool-call count equals n
toHaveToolCallInStep(step, match) Step (1-indexed) contains a tool call matching { name, argsContains?, argsEquals? }
toHaveFinalTextContaining(text, caseInsensitive?) Last step's text content includes text

Adding a new runner

To add support for a new model, add a registerDeepAgentRunner call in src/setup.ts:

registerDeepAgentRunner("my-model", (config) =>
  createDeepAgent({ ...config, model: new ChatMyProvider({ model: "my-model" }) }),
);

Then run evals with EVAL_RUNNER=my-model.

To add a completely different runner backend (not deepagents), implement the EvalRunner interface directly and call registerRunner(), or use the runner directly inside of the eval suite.

Development

pnpm build       # Build with tsdown
pnpm typecheck   # Type-check without emitting