Files
deepagentsjs/internal/eval-harness
github-actions[bot] c231aed3ee chore: version packages (#525)
This PR was opened by the [Changesets
release](https://github.com/changesets/action) GitHub action. When
you're ready to do a release, you can merge this and the packages will
be published to npm automatically. If you're not ready to do a release
yet, that's fine, whenever you add more changesets to main, this PR will
be updated.


# Releases
## @langchain/quickjs@0.4.0

### Minor Changes

- [#531](https://github.com/langchain-ai/deepagentsjs/pull/531)
[`a76b7df`](https://github.com/langchain-ai/deepagentsjs/commit/a76b7df62310e7f2dd49bb1ea5f1b3ee6c8590b6)
Thanks [@colifran](https://github.com/colifran)! - chore(quickjs):
update `REPLMiddleware` to be named `CodeInterpreterMiddleware`

### Patch Changes

- [#524](https://github.com/langchain-ai/deepagentsjs/pull/524)
[`2cbd524`](https://github.com/langchain-ai/deepagentsjs/commit/2cbd5245a43fb1ba97fa532c1942a8903e090cfa)
Thanks [@colifran](https://github.com/colifran)! - fix(quickjs):
individual repl sessions use individual wasm module causing inefficient
memory usage

## deepagents-acp@0.1.11

### Patch Changes

- Updated dependencies
\[[`f164f99`](https://github.com/langchain-ai/deepagentsjs/commit/f164f992e06a157573612fb2640232f44d9daa18)]:
    -   deepagents@1.10.1

## deepagents@1.10.1

### Patch Changes

- [#479](https://github.com/langchain-ai/deepagentsjs/pull/479)
[`f164f99`](https://github.com/langchain-ai/deepagentsjs/commit/f164f992e06a157573612fb2640232f44d9daa18)
Thanks [@ramon-langchain](https://github.com/ramon-langchain)! -
feat(deepagents): add snapshot/start/stop lifecycle to LangSmithSandbox

## @deepagents/evals@0.0.10

### Patch Changes

- Updated dependencies
\[[`f164f99`](https://github.com/langchain-ai/deepagentsjs/commit/f164f992e06a157573612fb2640232f44d9daa18)]:
    -   deepagents@1.10.1

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Colin Francis <colin.francis@langchain.dev>
2026-05-11 14:32:28 -07:00
..
2026-04-01 17:10:16 -07:00
2026-05-11 14:32:28 -07:00
2026-05-11 14:32:28 -07:00

@deepagents/evals

Generic eval harness for deepagents. Provides the runner interface, runner registry, trajectory parsing, and custom vitest matchers with LangSmith feedback integration.

Package exports

Export Description
@deepagents/evals Core harness — EvalRunner, registry functions, parseTrajectory, getFinalText, custom matchers
@deepagents/evals/deepagent registerDeepAgentRunner() — wires createDeepAgent into the generic runner interface
@deepagents/evals/setup Side-effect import that registers all concrete runners (sonnet-4-5, gpt-4.1, etc.)

Core concepts

EvalRunner

The central interface. A runner has a name, a run() method, and an extend() method:

interface EvalRunner {
  name: string;
  run(params: RunAgentParams): Promise<AgentTrajectory>;
  extend(overrides: Record<string, unknown>): EvalRunner;
}
  • run() takes invocation params only — { query, initialFiles? }.
  • extend() returns a new runner with agent configuration overrides (e.g. systemPrompt, tools, subagents) baked in. This keeps "what the agent is" separate from "what to ask it".

RunAgentParams

interface RunAgentParams {
  query: string;
  initialFiles?: Record<string, string>;
}

Pure invocation inputs. query is the user message; initialFiles seeds the agent's virtual file system.

AgentTrajectory

interface AgentTrajectory {
  steps: AgentStep[];
  files: Record<string, string>;
}

The output of every run() call. steps is an ordered list of agent turns (each containing an AIMessage action and any ToolMessage observations). files is the final file-system snapshot.

Runner registry

flowchart LR
    setup["setup.ts<br/><i>vitest setupFile</i>"]
    deepagent["deepagent.ts<br/>registerDeepAgentRunner()"]
    registry["index.ts<br/>registerRunner()"]
    resolve["resolveRunner(name)"]
    default["getDefaultRunner()<br/>reads EVAL_RUNNER env"]
    test["test file<br/>runner.run() / runner.extend()"]

    setup -- "calls for each model" --> deepagent
    deepagent -- "wraps in DeepAgentEvalRunner" --> registry
    default -- "looks up by name" --> registry
    resolve -- "looks up by name" --> registry
    test --> default
    test --> resolve
  1. setup.ts is loaded as a vitest setupFile. It calls registerDeepAgentRunner() for each model, which internally calls registerRunner().
  2. getDefaultRunner() reads the EVAL_RUNNER env var, looks up the runner by name, and returns it (cached).
  3. resolveRunner(name) is the lower-level lookup if you need a specific runner by name.

registerDeepAgentRunner

Bridges createDeepAgent to the generic EvalRunner interface:

registerDeepAgentRunner("sonnet-4-5", (config) =>
  createDeepAgent({
    ...config,
    model: new ChatAnthropic({ model: "claude-sonnet-4-5-20250929" }),
  }),
);

The factory receives optional overrides (from extend()) and must return an invokable agent. A default agent (no overrides) is built eagerly at registration time and reused across run() calls for performance. When extend() is called, a fresh agent is constructed with the overrides.

The DeepAgentEvalRunner handles:

  • Converting initialFiles strings into the FileData format expected by the agent's state backend
  • Generating a unique thread_id per invocation
  • Calling ls.logOutputs() for LangSmith experiment tracking
  • Parsing the raw LangGraph result into an AgentTrajectory

Custom vitest matchers

Imported automatically when you import from @deepagents/evals. Each matcher also logs LangSmith feedback (scores) so results appear in the experiment dashboard.

Matcher Description
toHaveAgentSteps(n) Trajectory has exactly n steps
toHaveToolCallRequests(n) Total tool-call count equals n
toHaveToolCallInStep(step, match) Step (1-indexed) contains a tool call matching { name, argsContains?, argsEquals? }
toHaveFinalTextContaining(text, caseInsensitive?) Last step's text content includes text

Adding a new runner

To add support for a new model, add a registerDeepAgentRunner call in src/setup.ts:

registerDeepAgentRunner("my-model", (config) =>
  createDeepAgent({ ...config, model: new ChatMyProvider({ model: "my-model" }) }),
);

Then run evals with EVAL_RUNNER=my-model.

To add a completely different runner backend (not deepagents), implement the EvalRunner interface directly and call registerRunner(), or use the runner directly inside of the eval suite.

Development

pnpm build       # Build with tsdown
pnpm typecheck   # Type-check without emitting