Compare commits

...

265 Commits

Author SHA1 Message Date
Kit Langton 691e630ea5 refactor(tui): remove unavailable sharing commands 2026-07-31 21:45:11 +00:00
opencode-agent[bot] 56c6add5c3 refactor(core): remove unused helpers (#39943)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-31 16:17:14 -05:00
Kit Langton 2cfe8883ea feat(plugin): expose session rename and wait to plugins (#39932) 2026-07-31 21:08:33 +00:00
Kit Langton 102086c50f feat(tui): hot-reload local TUI plugins (#39776) 2026-07-31 16:35:05 -04:00
Kit Langton 06e26d89ad feat(protocol): accept title on session create (#39935) 2026-07-31 19:39:54 +00:00
Aiden Cline 84dd56ed34 fix(core): bound outbound image history (#39929) 2026-07-31 13:57:04 -05:00
opencode-agent[bot] f92e490bb6 fix(tui): preserve diff hunk boundaries (#39693)
Co-authored-by: James Long <17031+jlongster@users.noreply.github.com>
2026-07-31 14:47:36 -04:00
Kit Langton 0481dba88e fix(tui): improve narrow layouts (#39918) 2026-07-31 14:44:13 -04:00
Kit Langton dd0e41c633 fix(tui): render plugin slots and routes with component semantics (#39917) 2026-07-31 14:30:57 -04:00
opencode-agent[bot] b69e51d835 refactor(ai): simplify provider options (#39924)
Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
2026-07-31 13:27:32 -05:00
opencode-agent[bot] e60e8d9387 refactor(core): rename attachments config to media (#39927)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-31 13:26:53 -05:00
Kit Langton 3bfce3fd2d feat(tui): expand pasted text (#39920) 2026-07-31 18:17:35 +00:00
Aiden Cline b498a5c6c4 feat(ai): expand OpenRouter native support (#39907) 2026-07-31 12:57:18 -05:00
Kit Langton db9e942398 fix(core): resize large image attachments (#39919) 2026-07-31 13:24:05 -04:00
Kit Langton 0e26116f68 fix(tui): stabilize open picker 2026-07-31 16:38:48 +00:00
Aiden Cline 3468aa0140 feat(ai): configure chat max tokens field (#39909) 2026-07-31 11:14:08 -05:00
Kit Langton 7757f712a2 test(app): restore stable offset observer ordering (#39901) 2026-07-31 12:13:16 -04:00
Aiden Cline 7129ed6e62 fix(core): respect model input limits (#39797) 2026-07-31 10:41:36 -05:00
Kit Langton 6c37842520 fix(cli): deduplicate Node asset destinations (#39900) 2026-07-31 11:40:38 -04:00
Kit Langton dc3c996892 fix(tui): stabilize generated session titles (#39894) 2026-07-31 11:13:25 -04:00
Kit Langton 7fd12c560c fix(tui): fit project paths in open menu (#39887) 2026-07-31 11:07:12 -04:00
Kit Langton 7aaf4e7750 refactor(session): centralize fallback title policy (#39890) 2026-07-31 10:56:38 -04:00
Kit Langton 1311d909a7 fix(tui): preserve session location during handoff (#39886) 2026-07-31 10:54:36 -04:00
Kit Langton d07d9ae0da fix(tui): remove shells from their location (#39885) 2026-07-31 10:46:43 -04:00
Kit Langton 1d18459cdc fix(tui): align session picker scope (#39784) 2026-07-31 10:32:16 -04:00
Kit Langton db45026c6c fix(tui): default tabs to global scope (#39783) 2026-07-31 10:15:53 -04:00
Kit Langton 77df98db51 Revert "feat(tui): delete current session (#39750)"
This reverts commit 7814568ba0.
2026-07-31 10:01:35 -04:00
Aiden Cline ff2b184af1 feat(ai): support Gemini thinking levels (#39796) 2026-07-30 23:38:00 -05:00
Aiden Cline 671e164f8d refactor(core): contain Codex in OpenAI plugin (#39734) 2026-07-30 21:43:38 -05:00
Aiden Cline 5b4ebf3e9d fix(core): map xAI native options (#39787) 2026-07-30 21:42:53 -05:00
Aiden Cline 98229d466d feat(plugin): add session request hook (#39764) 2026-07-30 21:11:22 -05:00
Kit Langton 5c7e2b5042 fix(tui): clarify open menu project labels (#39780) 2026-07-30 21:34:29 -04:00
Kit Langton 0a6a5d3e80 fix(session): retry failed title generation (#39748) 2026-07-30 21:26:44 -04:00
Kit Langton eb95bd27fe feat(session): make generated titles optional (#39747) 2026-07-30 21:24:34 -04:00
Kit Langton 865f512a44 fix(tui): preserve current selection across list updates (#39774) 2026-07-31 00:48:51 +00:00
Kit Langton a460f02f67 feat(tui): inherit session directory when creating a new session (#39753) 2026-07-30 20:31:38 -04:00
Kit Langton 146fdb9de1 feat(tui): add open menu for sessions and projects (#39752) 2026-07-30 20:08:12 -04:00
Kit Langton b866900417 fix(tui): name deleted session in toast (#39768) 2026-07-30 19:59:27 -04:00
Kit Langton 8e3e94aa26 fix(tui): focus palette settings after layout (#39585) 2026-07-30 19:39:16 -04:00
Aiden Cline cc1289048e refactor(core): isolate AI SDK native mappings (#39761) 2026-07-30 18:38:29 -05:00
Kit Langton 7814568ba0 feat(tui): delete current session (#39750) 2026-07-30 17:06:41 -04:00
Kit Langton 16b247f756 fix(tui): smooth new session tab handoff (#39745) 2026-07-30 16:17:01 -04:00
Kit Langton 9abd9594de fix(simulation): mirror kitty keyboard protocol (#39741) 2026-07-30 20:03:21 +00:00
Kit Langton 22d2012f75 fix(tui): register storybook only for story runs (#39733) 2026-07-30 15:55:05 -04:00
Kit Langton cc883058df feat(tui): reopen closed session tabs (#39731) 2026-07-30 15:54:43 -04:00
Kit Langton 8ebc4c6f85 fix(tui): call session tabs just tabs (#39730) 2026-07-30 15:52:58 -04:00
Kit Langton b0ed9990b3 feat(tui): add option tab shortcuts (#39725) 2026-07-30 15:51:10 -04:00
Kit Langton e1c04dcce6 feat(tui): add temporary new session tab (#39735) 2026-07-30 15:05:34 -04:00
Kit Langton 188c642d8c refactor(tui): flatten project picker list (#39726) 2026-07-30 14:18:16 -04:00
Kit Langton 9d55b223bb fix(core): stop repository discovery at nearest marker (#39714) 2026-07-30 13:28:29 -04:00
Kit Langton ce7a7e4e23 fix(tui): truncate project picker paths (#39678) 2026-07-30 13:13:31 -04:00
Kit Langton cba5ba03e3 fix(tui): show shells from session location (#39691) 2026-07-30 13:12:01 -04:00
Kit Langton 49081a4e24 fix(tui): flash tabs on prompt admission (#39702) 2026-07-30 16:57:36 +00:00
opencode-agent[bot] d6371f2fcd fix(session): update activity on prompt (#39539)
Co-authored-by: James Long <17031+jlongster@users.noreply.github.com>
2026-07-30 16:28:10 +00:00
Aiden Cline 7625cbdf47 feat(core): route providers through native AI (#39615) 2026-07-30 11:26:18 -05:00
opencode-agent[bot] 0ffec67ca3 refactor: remove unused V2 code (#39699)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-30 11:25:03 -05:00
Kit Langton bd132f2614 refactor(core): skip redundant worktree lookup (#39692) 2026-07-30 16:15:07 +00:00
opencode-agent[bot] dcb0df4ac1 fix(core): stop MCP SSE error reconnect loops (#39671)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-30 10:53:33 -05:00
Kit Langton 02f3504055 fix(tui): hide redundant session directory labels (#39686) 2026-07-30 11:37:33 -04:00
James Long f23ee5e4a8 refactor(core): simplify formatter selection (#39575) 2026-07-30 10:59:01 -04:00
Kit Langton 77aa85c589 feat(tui): project picker with footer crossfade (#39566) 2026-07-30 10:39:22 -04:00
Dax Raad a618946b7e fix(tui): reload plugins with config 2026-07-30 10:31:09 -04:00
opencode-agent[bot] f9de608dea chore: refresh bun lockfile (#39617)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-30 04:36:21 +00:00
Aiden Cline 906dc8f5b2 fix(ai): apply catalog settings to provider models (#39613) 2026-07-29 22:42:47 -05:00
Aiden Cline 12971935ab feat(core): parse shell permission commands (#39567) 2026-07-29 22:42:18 -05:00
Kit Langton 78c139b8b5 feat(plugin): add ui.tabs API for session tab control (#39591) 2026-07-29 22:57:39 -04:00
Kit Langton c161978852 feat(tui): prefetch open session tabs after connect (#39589) 2026-07-29 22:23:13 -04:00
Kit Langton cb37a7166a feat(tui): make session tab switching fast for long transcripts (#39568) 2026-07-29 22:23:10 -04:00
Dax Raad 488445a679 fix(tui): correct project-aware session lists 2026-07-29 21:59:33 -04:00
opencode-agent[bot] 97786afdd8 refactor(core): share file diff construction (#39586)
Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
2026-07-30 00:28:33 +00:00
Aiden Cline b1a5e8a6ae test(tui): restore compaction event lifecycle (#39581) 2026-07-29 18:11:58 -05:00
Aiden Cline 0ece10af43 fix(core): add mutation permission previews (#39578) 2026-07-29 18:04:13 -05:00
Aiden Cline d1a02b149c fix(core): clarify subagent tool guidance (#39572) 2026-07-29 16:48:49 -05:00
James Long 94ee274aeb feat(core): add V2 formatter runtime (#39564) 2026-07-29 17:23:04 -04:00
Kit Langton 3c259fc552 feat(tui): replace scrap screen with component storybook (#39548) 2026-07-29 15:51:15 -04:00
Aiden Cline 210be4b749 fix(core): preserve shell output on timeout (#39559) 2026-07-29 14:31:26 -05:00
James Long 464649e67e feat(tui): batch event delivery (#39551) 2026-07-29 15:20:39 -04:00
Aiden Cline c8dca936b1 feat(plugin): add shell.create.before hook (#39547) 2026-07-29 14:11:52 -05:00
Kit Langton 4f871906bc fix(tui): support cd before session creation (#39555) 2026-07-29 18:45:42 +00:00
Aiden Cline f7ea2fc346 test(cli): update ACP fork expectation (#39554) 2026-07-29 13:21:34 -05:00
Aiden Cline 5cb633a48e feat(core): support pinned Code Mode tools (#39550) 2026-07-29 13:16:49 -05:00
Kit Langton b985d2eb8e feat(tui): polish session tab animations (#39542) 2026-07-29 13:50:49 -04:00
Dax Raad 014908a8d7 feat(tui): reload config file changes 2026-07-29 13:05:48 -04:00
Dax Raad f599f8f3d3 fix(tui): guard Bun runtime plugin support 2026-07-29 12:57:35 -04:00
Aiden Cline b2010220f9 fix(core): clarify Code Mode tool boundary (#39540) 2026-07-29 11:56:53 -05:00
James Long c2e975c4e6 refactor(plugin): expose resolved TUI theme (#39536) 2026-07-29 12:51:55 -04:00
Dax Raad 6aa250ee5d refactor(tui): flatten state storage path 2026-07-29 12:39:21 -04:00
Dax Raad 8f1e3ff75c docs: record v1 to v2 database migration decisions 2026-07-29 12:37:00 -04:00
Dax Raad 5438dfb751 fix(tui): remove invalid model toasts 2026-07-29 12:36:56 -04:00
Dax Raad 4bd16d6f47 feat(tui): default tabs to cwd scope 2026-07-29 12:35:32 -04:00
Dax Raad 9ee337469d feat(tui): add persistent storage context 2026-07-29 12:33:18 -04:00
Aiden Cline 813c41ff6c fix(core): simplify shell execution boundary (#39530) 2026-07-29 11:16:45 -05:00
Dax Raad 247f14f955 feat(tui): add replaceable prompt footer slot 2026-07-29 10:16:34 -04:00
Dax Raad bd906d468d refactor(core): make watcher subscription effectful 2026-07-29 09:50:24 -04:00
Dax Raad fc11ed3838 feat(session): define explicit fork boundaries 2026-07-29 09:50:24 -04:00
Dax Raad 2a85c861e0 fix(session): hide pending admission sequence 2026-07-29 09:50:24 -04:00
Shoubhit Dash 9554f9a16e feat(ai): type Vertex request options (#39499) 2026-07-29 18:23:00 +05:30
Shoubhit Dash d72b428061 feat(ai): type Vertex Chat request options (#39503) 2026-07-29 18:22:41 +05:30
Shoubhit Dash fea17b4a0e feat(ai): type Cloudflare request options (#39507) 2026-07-29 18:22:16 +05:30
Shoubhit Dash 9038e44a68 feat(ai): type Copilot request options (#39496) 2026-07-29 18:22:01 +05:30
Shoubhit Dash 3b8299e3f2 feat(ai): type Azure request options (#39498) 2026-07-29 18:21:43 +05:30
Shoubhit Dash f5cdf0f056 feat(ai): type Anthropic request options (#39502) 2026-07-29 18:21:19 +05:30
Shoubhit Dash 224feff7c4 feat(ai): type compatible Responses options (#39506) 2026-07-29 18:20:55 +05:30
Shoubhit Dash ce2c9e7e26 feat(ai): type Google request options (#39504) 2026-07-29 18:20:30 +05:30
Shoubhit Dash cb80f47112 feat(ai): type Vertex Messages request options (#39501) 2026-07-29 18:20:04 +05:30
Shoubhit Dash b4ac939537 feat(ai): type xAI request options (#39505) 2026-07-29 18:19:44 +05:30
Shoubhit Dash 8a96b80aec feat(ai): type OpenAI request options (#39495) 2026-07-29 18:19:26 +05:30
Shoubhit Dash 333a090975 feat(ai): type Vertex Responses request options (#39500) 2026-07-29 18:19:02 +05:30
Shoubhit Dash 5a78a17e49 feat(ai): type OpenRouter request options (#39508) 2026-07-29 18:18:43 +05:30
Shoubhit Dash 9d6af6afa4 feat(ai): type compatible request options (#39509) 2026-07-29 18:18:24 +05:30
Shoubhit Dash 8c3e06798c refactor(ai): limit provider option inference (#39510) 2026-07-29 18:18:00 +05:30
Shoubhit Dash 5882b64612 feat(ai): type Anthropic-compatible request options (#39497) 2026-07-29 18:07:24 +05:30
Shoubhit Dash 309c4fe6f0 feat(ai): infer model provider options (#39493) 2026-07-29 18:05:56 +05:30
Shoubhit Dash d9555f138b refactor(ai): internalize request compilation (#39132) 2026-07-29 16:51:11 +05:30
Kit Langton b47cfbee7c fix(tui): reduce tab pulse allocations (#39433) 2026-07-28 22:43:46 -04:00
Kit Langton 5504245f7b feat(tui): add session tab playground (#39432) 2026-07-28 22:36:53 -04:00
Kit Langton f64b50d71b feat(tui): add unread tab glow (#39428) 2026-07-28 22:34:33 -04:00
Kit Langton a7b2ea94e5 fix(tui): always show session tab (#39429) 2026-07-28 22:33:34 -04:00
Kit Langton 7b775c2582 fix(cli): embed native watcher binding 2026-07-28 22:32:08 -04:00
Dax 12a931a220 feat(tui): filter subagents by activity 2026-07-28 22:13:46 -04:00
Dax Raad 90100c1365 docs: clarify side-by-side V1 and V2 installs 2026-07-28 22:02:19 -04:00
Dax Raad 139c9febe4 refactor(tui): group tab settings 2026-07-28 21:23:53 -04:00
Dax Raad 06290907a9 feat(tui): restore plugin manager dialog 2026-07-28 21:19:17 -04:00
Kit Langton 1c8175a61a fix(tui): preserve tab context on home and close (#39421) 2026-07-29 01:03:17 +00:00
Dax Raad fe91698ed6 fix(tui): initialize external plugin runtime 2026-07-28 21:00:09 -04:00
Dax Raad 068c32df39 feat(tui): discover project plugins 2026-07-28 20:21:38 -04:00
Kit Langton a2885d1662 feat(tui): add session tab history (#39411) 2026-07-28 19:38:46 -04:00
Kit Langton 38a3dbb4c4 fix(tui): fade full-width tab titles (#39409) 2026-07-28 19:38:28 -04:00
Kit Langton 43383d4fba fix(tui): hide single session tab (#39408) 2026-07-28 18:25:46 -04:00
opencode-agent[bot] 40c4c3918a feat(core): enable fff in node runtimes (#38776)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-28 17:14:11 -05:00
Aiden Cline 754ea99d86 fix(core): preserve shell output tail (#39403) 2026-07-28 16:32:11 -05:00
Kit Langton 37a1b80d5a feat(tui): add adaptive session tabs (#39396) 2026-07-28 17:12:07 -04:00
Aiden Cline f95d04fea0 feat(core): improve shell tool guidance (#39401) 2026-07-28 15:53:19 -05:00
James Long 08b80da931 refactor(tui): split theme hooks (#39395) 2026-07-28 16:25:32 -04:00
Aiden Cline f6fb1a7cdd fix(ai): retry transient client statuses (#39391) 2026-07-28 14:25:33 -05:00
Dax Raad 5bcc0016a6 feat(tui): add plugin context hook 2026-07-28 14:57:59 -04:00
Aiden Cline 771174b5c3 fix(cli): align auto permission flags (#39384) 2026-07-28 13:31:57 -05:00
Dax Raad 44cd984589 feat(tui): refine plugin context slots 2026-07-28 14:19:31 -04:00
James Long c445d98188 feat(theme): extract TUI theme package (#39378) 2026-07-28 14:14:39 -04:00
Kit Langton 27e7b0558a fix(core): preserve plugin update order (#39372) 2026-07-28 13:00:24 -04:00
Kit Langton fb975eeb7c test(ai): add scoped test LLM (#39223) 2026-07-28 16:41:09 +00:00
Kit Langton e556aca833 refactor(core): simplify plugin reload loop (#39356) 2026-07-28 12:21:29 -04:00
opencode-agent[bot] 73bd8a264b fix(app): keep new tab button visible (#39366)
Co-authored-by: Brendan Allan <14191578+Brendonovich@users.noreply.github.com>
2026-07-28 16:18:26 +00:00
opencode-agent[bot] 077338fcc8 fix(app): hide delete for provided servers (#39363)
Co-authored-by: Brendan Allan <14191578+Brendonovich@users.noreply.github.com>
2026-07-28 16:00:47 +00:00
Dax Raad b671a77145 fix(tui): simplify form field labels 2026-07-28 11:56:11 -04:00
Dax Raad 30d09a7d7e fix(core): improve web search consent flow 2026-07-28 11:54:44 -04:00
Dax Raad ee02fb4fce Websearch tweaks 2026-07-28 11:10:33 -04:00
Dax Raad 010133f6df feat(tui): expand v2 plugin context 2026-07-28 11:04:16 -04:00
opencode-agent[bot] 3b0d8f0e6f fix(tui): clear rehydrated compaction state (#39336)
Co-authored-by: Simon Klee <hello@simonklee.dk>
2026-07-28 16:04:11 +02:00
Simon Klee 4d59b059ee feat(tui): add verbose turn token usage (#39281) 2026-07-28 13:57:19 +02:00
Luke Parker 1be6d94267 fix(desktop): bootstrap v2 background service (#39309) 2026-07-28 21:03:34 +10:00
opencode-agent[bot] 302e9b45ab chore: merge dev into v2 (#39290)
Co-authored-by: Luke Parker <10430890+Hona@users.noreply.github.com>
Co-authored-by: opencode-agent[bot] <219766164+opencode-agent[bot]@users.noreply.github.com>
Co-authored-by: Aarav Sareen <96787824+arvsrn@users.noreply.github.com>
Co-authored-by: Brendan Allan <14191578+Brendonovich@users.noreply.github.com>
Co-authored-by: opencode-agent[bot] <opencode-agent[bot]@users.noreply.github.com>
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
Co-authored-by: opencode <opencode@sst.dev>
Co-authored-by: Jay <53023+jayair@users.noreply.github.com>
Co-authored-by: David Hill <1879069+iamdavidhill@users.noreply.github.com>
Co-authored-by: BB84 <110078428+BB-84C@users.noreply.github.com>
Co-authored-by: Aiden Cline <aidenpcline@gmail.com>
Co-authored-by: Dax <mail@thdxr.com>
Co-authored-by: Brendan Allan <git@brendonovich.dev>
Co-authored-by: Jay V <air@live.ca>
Co-authored-by: Adam <2363879+adamdotdevin@users.noreply.github.com>
Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
Co-authored-by: Dustin Deus <deusdustin@gmail.com>
Co-authored-by: Frank <frank@anoma.ly>
Co-authored-by: usrnk1 <7547651+usrnk1@users.noreply.github.com>
Co-authored-by: Jack <jack@anoma.ly>
Co-authored-by: Sebastian <hasta84@gmail.com>
Co-authored-by: Jérôme Benoit <jerome.benoit@sap.com>
Co-authored-by: Test User <test@test.com>
Co-authored-by: Simon Klee <hello@simonklee.dk>
Co-authored-by: Rahul A Mistry <149420892+ProdigyRahul@users.noreply.github.com>
Co-authored-by: Qiping Li <liqiping1991@gmail.com>
Co-authored-by: liqiping <liqiping@msh.team>
Co-authored-by: OpeOginni <107570612+OpeOginni@users.noreply.github.com>
Co-authored-by: Matthias Reso <13337103+mreso@users.noreply.github.com>
Co-authored-by: tobwen <1864057+tobwen@users.noreply.github.com>
Co-authored-by: Daniel Polito <danielbpolito@gmail.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: Devin R Leopold <devin.leopold@gmail.com>
Co-authored-by: Zach Bruggeman <mail@bruggie.com>
Co-authored-by: Zach Bruggeman <zbruggeman@ramp.com>
Co-authored-by: Kit Langton <kit.langton@gmail.com>
Co-authored-by: David Siewert <david1gruppenplan@gmail.com>
Co-authored-by: Andrei Dziahel <develop7@develop7.info>
Co-authored-by: adityachaudhary99 <adityaachaudhary2003@gmail.com>
Co-authored-by: Vladimir Glafirov <vglafirov@gitlab.com>
Co-authored-by: Matt Carey <mcarey@cloudflare.com>
2026-07-28 18:36:55 +10:00
Aiden Cline 7211c9934a feat(core): make edit matching forgiving (#39258) 2026-07-28 00:06:43 -05:00
opencode-agent[bot] 3bda0ce123 fix(core): bound search tool execution (#39238)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-27 23:21:15 -05:00
Aiden Cline 62320947d9 fix(core): refresh system prompt references (#39245) 2026-07-27 21:52:31 -05:00
Aiden Cline 0cb9bb567e fix(core): align Meta system prompt (#39240) 2026-07-27 21:31:49 -05:00
Kit Langton b14adcaf83 docs: forbid type-position import references (#39234) 2026-07-27 22:30:26 -04:00
Kit Langton 5bd3da40a5 fix(core): keep config root watches alive and ignore vendored trees (#39239) 2026-07-27 22:30:21 -04:00
Aiden Cline abcbdad530 fix(core): refresh Meta system prompt (#39237) 2026-07-27 21:14:16 -05:00
Kit Langton debdea40ea feat(core): reload configured plugins from source edits (#39224) 2026-07-27 22:04:28 -04:00
Kit Langton 775f24f049 test(core): add native watcher command reload test (#39216) 2026-07-27 21:08:10 -04:00
Kit Langton 470e360942 test(core): align tool contract expectations (#39172) 2026-07-28 00:48:29 +00:00
Aiden Cline f15398efc3 feat(core): improve edit tool output (#39211) 2026-07-27 19:11:04 -05:00
Kit Langton 4333a44e65 refactor(core): manage watcher lifecycle with RcMap (#39203) 2026-07-27 20:01:17 -04:00
Kit Langton 8b4b0d67d7 feat(core): reload discovered plugins from source edits (#39174) 2026-07-27 18:20:13 -04:00
Aiden Cline 9200e353bf feat(core): improve edit tool guidance (#39198) 2026-07-27 17:03:00 -05:00
Aiden Cline 31124312f6 fix(core): simplify tool schemas (#39184) 2026-07-27 16:04:48 -05:00
James Long 4f622fa7cd fix(tui): reference inferred hues in migrated themes (#39183) 2026-07-27 17:02:35 -04:00
James Long 4e4cf9e25e refactor(tui): extract event stream connection (#38872) 2026-07-27 16:31:21 -04:00
Aiden Cline 6da2f3c38f feat(core): improve read model output (#39146) 2026-07-27 14:52:50 -05:00
Kit Langton 1b39d364bd fix(core): align command reload pipeline and repair plugin fixture (#39171) 2026-07-27 15:29:27 -04:00
Kit Langton 856c569458 feat(core): reload agents from config change feed (#39167) 2026-07-27 15:14:08 -04:00
Kit Langton 92807d0bb9 feat(core): reload commands from config change feed (#39160) 2026-07-27 15:02:40 -04:00
Kit Langton 713658c07b test(core): add config and watcher test services (#39157) 2026-07-27 14:32:38 -04:00
opencode-agent[bot] f5700808c5 fix(tui): preserve subagent list position (#39156)
Co-authored-by: Kit Langton <kit.langton@gmail.com>
2026-07-27 14:08:00 -04:00
Kit Langton 7eb51d0507 feat(core): diagnose prompt cache prefix changes (#39139) 2026-07-27 13:44:02 -04:00
Kit Langton 02c37c401a feat(core): expose config changes stream (#39131) 2026-07-27 13:38:08 -04:00
opencode-agent[bot] 33e3d1ebca fix(core): tolerate missing tool input schemas (#39130)
Co-authored-by: Dax Raad <d@ironbay.co>
2026-07-27 11:21:47 -04:00
Kit Langton 1f2c59a1b6 fix(core): commit state before finalize publishes (#38983) 2026-07-27 11:04:04 -04:00
Aiden Cline 65d2a4e00c feat(core): improve read tool parity (#39126) 2026-07-27 09:58:37 -05:00
Shoubhit Dash 7d4de3d9e4 fix(core): clarify web search provider prompt (#39123) 2026-07-27 19:52:36 +05:30
Aiden Cline 9977ef0160 refactor(core): tag read outputs (#39122) 2026-07-27 09:09:44 -05:00
Simon Klee 9a55d125f6 tui: render mini compaction boundaries (#39103) 2026-07-27 14:19:03 +02:00
Simon Klee 766aaf448d tui: add settings to command palette search (#39058) 2026-07-27 10:45:35 +02:00
Simon Klee f14d78afeb tui: skip abort on mini session close (#39067) 2026-07-27 10:42:37 +02:00
Dax Raad 0261f04b90 fix(core): handle oversized ripgrep matches 2026-07-27 03:17:08 -04:00
Aiden Cline 9b49e7bec9 test(core): implement catalog host model list (#39053) 2026-07-26 23:32:33 -05:00
Aiden Cline 93cb113cef fix(util): declare node tracing dependency (#39050) 2026-07-26 23:12:36 -05:00
Aiden Cline 4216d35e4b fix(server): declare schema dependency (#39043) 2026-07-26 22:46:46 -05:00
Dax Raad 863645c671 test(core): update grep error assertion 2026-07-26 20:17:31 -04:00
Dax Raad 5592f5225b fix(app): update remote sdk contracts 2026-07-26 20:16:58 -04:00
Dax Raad 8db7487c89 refactor(core): consolidate tool architecture 2026-07-26 20:16:58 -04:00
Aiden Cline 0fd73a2976 fix(core): align grep behavior and guidance (#38999) 2026-07-26 17:29:27 -05:00
Dax Raad 80865407e0 refactor(sdk): remove local legacy package 2026-07-26 02:47:43 -04:00
Dax Raad 28f4284bd7 fix(www): canonicalize production routes 2026-07-26 02:39:12 -04:00
Aiden Cline 7affee529b fix(core): harden grep search behavior (#38922) 2026-07-25 23:23:53 -05:00
Aiden Cline 79c7e9446e fix(core): clarify custom question answers (#38919) 2026-07-25 23:00:26 -05:00
Shoubhit Dash efb629a33a feat(core): add pluggable web search (#35558)
Co-authored-by: Dax Raad <d@ironbay.co>
2026-07-26 03:55:05 +00:00
Dax Raad 2ddc91a0e8 fix(tui): show shell working directory in prompt 2026-07-25 23:42:53 -04:00
Dax Raad c7871e14d4 fix(www): remove deployment environment gate 2026-07-25 21:21:57 -04:00
Dax Raad 9840f63b12 chore(www): simplify worker routes 2026-07-25 21:17:25 -04:00
Dax Raad 56a9c0150a fix(www): mark deploy script as module 2026-07-25 21:17:25 -04:00
Dax Raad 203a0613b8 feat(www): migrate docs to Blume 2026-07-25 21:17:25 -04:00
Aiden Cline 7d8f1bdab3 tweak(core): simplify skill tool description (#38900) 2026-07-25 17:09:20 -05:00
Aiden Cline 9eea5bc925 fix(core): tweak glob tool description/parameters (#38899) 2026-07-25 16:52:36 -05:00
Aiden Cline f753103e82 fix(core): reject file glob roots (#38890) 2026-07-25 14:43:18 -05:00
Aiden Cline 1e35d33ecb fix(codemode): search nested namespaces (#38887) 2026-07-25 14:17:49 -05:00
Aiden Cline c5bf4edb10 fix(ai): preserve response message phases (#38777) 2026-07-25 14:06:00 -05:00
Aiden Cline cce8bb0e1c fix(core): clarify empty Code Mode guidance (#38883) 2026-07-25 13:20:16 -05:00
Dax Raad 02c66c5fc1 docs(core): fix OpenCode skill links 2026-07-25 14:08:29 -04:00
Aiden Cline 33390cc457 fix(core): keep execute tool cache stable (#38783) 2026-07-25 13:01:05 -05:00
Kit Langton b2afb35527 refactor(core): settle steps lock-free by joining tool fibers first (#38743) 2026-07-25 13:43:17 -04:00
opencode-agent[bot] 828148909d fix(tui): preserve workspace while reconnecting (#38788) 2026-07-25 03:48:42 +00:00
opencode-agent[bot] 454145fe65 fix(tui): preserve workspace while reconnecting (#38788) 2026-07-25 02:22:47 +00:00
Aiden Cline 1291dc1f11 fix(core): clarify code mode tool boundary (#38785) 2026-07-24 20:49:55 -05:00
Aiden Cline c7d7f61146 fix(ai): preserve Anthropic usage metadata (#38751) 2026-07-24 16:23:51 -05:00
Aiden Cline 13b6845e7e fix(core): scope MCP execute guidance to Code Mode (#38753) 2026-07-24 16:02:57 -05:00
Aiden Cline d1d97014b4 refactor(core): move static Code Mode guidance (#38746) 2026-07-24 14:46:59 -05:00
Aiden Cline b31747124b fix(core): stream Code Mode tool progress (#38718) 2026-07-24 14:46:46 -05:00
Aiden Cline 49bec25ae5 fix(ai): align Anthropic stream handling (#38733) 2026-07-24 14:23:35 -05:00
Aiden Cline d66d0cb904 fix(core): clarify Code Mode tool availability (#38745) 2026-07-24 13:44:59 -05:00
Kit Langton 3193f3aa95 fix(tui): flag likely cache busts accurately (#38727) 2026-07-24 18:35:13 +00:00
Aiden Cline ee5460a152 fix(codemode): report interrupted tool calls (#38741) 2026-07-24 13:32:54 -05:00
Kit Langton b09a066fb5 fix(ai): report OpenAI cache writes (#38735) 2026-07-24 18:26:30 +00:00
Aiden Cline 423fad730c fix(core): authorize external glob paths (#38714) 2026-07-24 13:24:23 -05:00
Kit Langton 5ae2d6d3f6 refactor(core): settle declined tool calls durably with typed reasons (#38734) 2026-07-24 18:18:02 +00:00
Kit Langton 993f046dd9 fix(ai): layer prompt cache breakpoints (#38725) 2026-07-24 18:02:35 +00:00
opencode-agent[bot] 9b640cf97d test(core): remove flaky npm install test (#38729)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-24 12:33:28 -05:00
Kit Langton 2200d100d0 refactor(core): name the unsettled-tool sweep scope and untangle hosted settlement (#38724) 2026-07-24 13:32:14 -04:00
Kit Langton 68ef893818 fix(core): start all suspended sessions promptly (#38720) 2026-07-24 16:58:50 +00:00
Kit Langton 7ab3dd04ad refactor(core): unify tool fiber bookkeeping into one owned structure (#38719) 2026-07-24 16:55:43 +00:00
Aiden Cline 4f201f87a9 fix(ai): align Bedrock stream handling (#38712) 2026-07-24 11:15:23 -05:00
Kit Langton 5aa276c117 refactor(core): mint assistant message identity before the step runs (#38717) 2026-07-24 16:11:31 +00:00
Aiden Cline 4605308be2 feat(core): render CodeMode catalog deltas from structured snapshots (#38183) 2026-07-24 10:20:26 -05:00
Aiden Cline c64d813347 fix(core): report truncated glob results (#38631) 2026-07-24 10:17:37 -05:00
Kit Langton 6e4a972bb9 refactor(core): clean up callModel readability (#38706) 2026-07-24 11:17:20 -04:00
Kit Langton 835149e42b refactor(core): select small model without sorting (#38707) 2026-07-24 11:08:57 -04:00
Kit Langton 0f3c30118c refactor(core): simplify session runner loop and pending input scopes (#38602) 2026-07-24 14:27:23 +00:00
Shoubhit Dash 35d31d8ec1 refactor(ai): remove dead LLM exports (#38700) 2026-07-24 14:19:07 +00:00
Shoubhit Dash 0374d29232 refactor(ai): colocate provider options (#38698) 2026-07-24 19:08:36 +05:30
Shoubhit Dash d90da82be2 refactor(ai): normalize provider option parsing (#38695) 2026-07-24 19:02:02 +05:30
Shoubhit Dash c06186a9d9 fix(ai): forward Anthropic provider options (#38694) 2026-07-24 18:50:20 +05:30
Shoubhit Dash c5680a206e refactor(ai): make OpenAI Responses extend Open Responses (#38681) 2026-07-24 18:36:35 +05:30
Simon Klee edaee143d9 mini: reserve headroom before showing usage (#38659) 2026-07-24 11:23:53 +02:00
Simon Klee 4184149b90 mini: monochrome rendered markdown only (#38656) 2026-07-24 11:14:48 +02:00
Simon Klee 00f063b381 mini: pack statusline by content width (#38646) 2026-07-24 10:54:30 +02:00
Aiden Cline ea010ab3a4 feat(ai): round-trip Bedrock redacted reasoning (#38623) 2026-07-24 01:06:58 -05:00
Aiden Cline 65c5c7e3f6 feat(ai): round-trip Anthropic redacted thinking blocks (#38614) 2026-07-24 00:42:44 -05:00
opencode-agent[bot] ee69a91f26 docs: add ideal pseudocode skill (#38611)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-23 23:02:44 -05:00
Kit Langton 7456598cde fix(core): share one tool snapshot per request (#38596) 2026-07-24 02:54:22 +00:00
Kit Langton c228fc4886 fix(core): stabilize tool definition ordering (#38590) 2026-07-23 21:30:06 -04:00
Kit Langton bb3f4cc3c7 fix(codemode): stabilize catalog ordering (#38588) 2026-07-23 20:56:18 -04:00
Aiden Cline 8d9727be9f fix(core): improve patch errors (#38369) 2026-07-23 18:59:29 -05:00
Aiden Cline 6401eeaea0 feat(ai): preserve raw finish reasons (#38423) 2026-07-23 18:16:11 -05:00
Kit Langton e7ecee5df2 fix(core): isolate tool hook outcomes (#38571) 2026-07-23 17:40:42 -04:00
opencode-agent[bot] 02f2725154 chore: merge dev into v2 (#38563)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
Co-authored-by: opencode-agent[bot] <219766164+opencode-agent[bot]@users.noreply.github.com>
Co-authored-by: Dax Raad <d@ironbay.co>
Co-authored-by: Dax <mail@thdxr.com>
Co-authored-by: opencode-agent[bot] <opencode-agent[bot]@users.noreply.github.com>
Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
Co-authored-by: Nabs <nabil@instafork.com>
Co-authored-by: usrnk1 <7547651+usrnk1@users.noreply.github.com>
Co-authored-by: Brendan Allan <14191578+Brendonovich@users.noreply.github.com>
Co-authored-by: Aarav Sareen <96787824+arvsrn@users.noreply.github.com>
Co-authored-by: Brendan Allan <git@brendonovich.dev>
Co-authored-by: Victor Navarro <vn4varro@gmail.com>
Co-authored-by: Vladimir Glafirov <vglafirov@gitlab.com>
Co-authored-by: AidenGeunGeun <eastlandwyvern@gmail.com>
Co-authored-by: Mark <geraint0923@users.noreply.github.com>
Co-authored-by: Aiden Cline <aidenpcline@gmail.com>
Co-authored-by: Adam <2363879+adamdotdevin@users.noreply.github.com>
Co-authored-by: opencode <opencode@sst.dev>
Co-authored-by: Luke Parker <10430890+Hona@users.noreply.github.com>
Co-authored-by: David Hill <1879069+iamdavidhill@users.noreply.github.com>
Co-authored-by: Jay <air@live.ca>
Co-authored-by: Jay <53023+jayair@users.noreply.github.com>
Co-authored-by: BB84 <110078428+BB-84C@users.noreply.github.com>
Co-authored-by: Dustin Deus <deusdustin@gmail.com>
Co-authored-by: Frank <frank@anoma.ly>
Co-authored-by: Jack <jack@anoma.ly>
Co-authored-by: Sebastian <hasta84@gmail.com>
Co-authored-by: Jérôme Benoit <jerome.benoit@sap.com>
Co-authored-by: Test User <test@test.com>
Co-authored-by: Simon Klee <hello@simonklee.dk>
Co-authored-by: Rahul A Mistry <149420892+ProdigyRahul@users.noreply.github.com>
Co-authored-by: Qiping Li <liqiping1991@gmail.com>
Co-authored-by: liqiping <liqiping@msh.team>
Co-authored-by: OpeOginni <107570612+OpeOginni@users.noreply.github.com>
Co-authored-by: Matthias Reso <13337103+mreso@users.noreply.github.com>
Co-authored-by: tobwen <1864057+tobwen@users.noreply.github.com>
Co-authored-by: Daniel Polito <danielbpolito@gmail.com>
Co-authored-by: opencode <noreply@opencode.ai>
2026-07-23 16:20:37 -05:00
Kit Langton 79c1544072 refactor(tools): unify tool APIs and result handling (#38367) 2026-07-23 21:13:31 +00:00
Aiden Cline 8cac010bac fix(core): stop forcing toolChoice none on session.generate (#38557) 2026-07-23 14:12:23 -05:00
Aiden Cline 193f6be99c fix(ai): keep tools when Gemini tool choice is none (#38556) 2026-07-23 14:09:25 -05:00
opencode-agent[bot] 2a9f8e3a2c fix(tui): manage focus in devtools panels (#38555)
Co-authored-by: James Long <17031+jlongster@users.noreply.github.com>
2026-07-23 18:58:02 +00:00
Aiden Cline 2c814120c7 fix(ai): keep tools when Anthropic tool_choice is none (#38553) 2026-07-23 13:52:20 -05:00
opencode-agent[bot] 18fccac6ff fix(tui): preserve first message in new sessions (#38542)
Co-authored-by: Kit Langton <kit.langton@gmail.com>
Co-authored-by: James Long <17031+jlongster@users.noreply.github.com>
2026-07-23 13:59:51 -04:00
opencode-agent[bot] 360e7b412d feat(tui): expose debug settings (#38546)
Co-authored-by: James Long <longster@gmail.com>
2026-07-23 13:37:16 -04:00
opencode-agent[bot] ad596fb42b chore(core): upgrade fff to 0.10.1 (#38545)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-23 12:34:24 -05:00
opencode-agent[bot] 74e92f73e0 refactor(ai): remove unused response format (#38540)
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
2026-07-23 12:32:33 -05:00
1307 changed files with 91653 additions and 112893 deletions
-7
View File
@@ -1,7 +0,0 @@
---
"@opencode-ai/client": patch
"@opencode-ai/protocol": patch
"@opencode-ai/cli": patch
---
Expose background-service lifecycle status, preserve one process-held owner through startup and failure, reconnect TUIs without activating replacement, and stop exact service instances gracefully.
-5
View File
@@ -1,5 +0,0 @@
---
"@opencode-ai/cli": patch
---
Expose a TUI plugin slot at the top of the session view.
-7
View File
@@ -1,7 +0,0 @@
---
"@opencode-ai/client": patch
"@opencode-ai/plugin": patch
"@opencode-ai/protocol": patch
---
Expose transient, read-only session generation through the HTTP API, generated clients, and V2 plugin session context.
-5
View File
@@ -1,5 +0,0 @@
---
"@opencode-ai/cli": patch
---
Expose a TUI plugin slot above the session composer.
+37
View File
@@ -0,0 +1,37 @@
name: deploy-www
on:
push:
branches:
- dev
- v2
workflow_dispatch:
concurrency:
group: deploy-www-${{ github.ref_name }}
cancel-in-progress: false
permissions:
contents: read
jobs:
deploy:
if: github.repository == 'anomalyco/opencode' && (github.ref_name == 'dev' || github.ref_name == 'v2')
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@f43a0e5ff2bd294095638e18286ca9a3d1956744 # v3.6.0
- uses: ./.github/actions/setup-bun
- name: Build
working-directory: packages/www
run: bun run build
env:
BLUME_ENV: ${{ github.ref_name == 'v2' && 'production' || 'dev' }}
CLOUDFLARE_ENV: ${{ github.ref_name == 'v2' && 'production' || 'dev' }}
- name: Deploy
working-directory: packages/www
run: bun run deploy
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
+1 -1
View File
@@ -99,7 +99,7 @@ jobs:
- name: Check generated documentation - name: Check generated documentation
if: runner.os == 'Linux' if: runner.os == 'Linux'
working-directory: packages/docs working-directory: packages/www
run: bun run check:generated run: bun run check:generated
e2e: e2e:
+13
View File
@@ -0,0 +1,13 @@
import type { Context } from "../../../packages/plugin/src/tui/context"
export default {
id: "test.tui-discovery-smoke",
setup(_context: Context) {
// context.ui.toast.show({
// title: "TUI plugin discovery works",
// message: "Loaded .opencode/plugins/tui/discovery-smoke.ts",
// variant: "success",
// duration: 30_000,
// })
},
}
@@ -0,0 +1,68 @@
---
name: ideal-pseudocode
description: Function-by-function refactoring loop driven by ideal pseudocode. Use when the user says "ideal pseudocode", asks to make a function read like its pseudocode, or wants a dense module cleaned up one function at a time.
---
# Ideal Pseudocode
Clean up one function at a time by writing the pseudocode it _should_ read as, naming every delta between that and the real code, and closing only the gaps the user approves.
## Loop
One function per round. Never touch code before the user picks a direction.
1. **Pick the target** with the user — usually the next function up or down the call chain from the last round.
2. **Read the current code** fresh from disk. It may have unsaved or parallel edits; ask before overwriting anything unexpected.
3. **Distill.** Write the function's ideal pseudocode in a `ts`-fenced code block — TypeScript-flavored for syntax highlighting, but pseudocode: comments over mechanics, one line per idea, every arm of a loop visible as an arm. For a dense or unfamiliar function, first show the _current_ structure as pseudocode, then the ideal.
4. **Name the deltas.** A numbered list; each delta is one concrete gap: control flow smuggled through combinators, a flag-typed parameter, colliding names, side-effect plumbing drowning a loop body, twin functions differing only by a filter. For each, give an honest cost/benefit and a decisive recommendation — including "leave it" when the density has reasons (state the reasons).
5. **Verify before proposing.** Any API the proposal builds on (a combinator, schedule, matcher) gets checked against the installed version's source, not memory.
6. **Get buy-in, then apply.** Minimal edits. Quarantine irreducible ugliness inside a small helper whose name is the pseudocode line it implements. Collapse shallow sibling functions into one deep function only when the distinguishing argument is a domain noun, not a behavior switch.
7. **Validate.** Typecheck and the focused test suites from the package directory; rerun timing-sensitive tests repeatedly. Flag any semantic change the restructure introduced, even when tests pass.
8. **Offer the next target** and stop.
## Style
- Keep each round tight: pseudocode block, deltas, recommendation, one question.
- State mutation belongs in one visible place (usually the loop tail), never spread across error channels or callback closures.
- Prefer early return and `if` chains for control flow; matchers are for producing values.
- Compare names against the question each caller is actually asking; rename toward the domain noun.
## Example round
Every round should read like this (abridged from a real one, on an Effect step-runner):
> `runStep` it is. The ideal pseudocode:
>
> ```ts
> // One logical step: one settled model call, however many attempts it takes.
> // An attempt can end without settling in two ways:
> // transient provider failure -> retry the same call, same assistant message
> // compaction rewrote history -> rebuild the request and call again
> function runStep(session, promotable, step) {
> while (true) {
> const result = callModel(session, promotable, step)
>
> if (result.completed) return result
>
> if (result.retryable) {
> wait(retrySchedule) // exhausted -> publish Step.Failed, fail drain
> promotable = none // never re-promote on a second attempt
> continue
> }
>
> // compaction restarted the step: fresh request from rewritten history
> promotable = none
> step = result.step
> }
> }
> ```
>
> Comparing against the real thing, three deltas:
>
> 1. **The retry arm doesn't live in the loop.** In the pseudocode, retry is just another loop arm. In the real code it's smuggled through the error channel: `Effect.tapError` mutating loop variables via `Effect.sync` closures, then `Effect.retryOrElse` re-running a suspended call. Four mutable variables mutated from three different places — the state flow is invisible unless you already know the trick.
> 2. **`recoverOverflow: typeof compaction.compact | undefined`** — a function-typed parameter used as a one-shot armed/disarmed flag. The callee already closes over `compaction`; a boolean says what it is.
> 3. **`assistantMessageID`** exists solely to resume the same durable message across retries — fine, but only legible once the retry arm is visible in the loop.
>
> The unifying move for (1): a schedule driver, so a retryable failure becomes a caught value handled in the loop body exactly like the compaction arm. One honest trade: it's more manual than the blessed combinator — but the manual version is the loop the combinator was hiding, and the loop already exists for restarts. Two restart mechanisms, one control structure.
>
> Want me to apply it — unified loop, simplified schedule input, boolean `recoverOverflow`?
+1 -1
View File
@@ -1,4 +1,3 @@
- To regenerate the legacy JavaScript SDK, run `./packages/sdk/js/script/build.ts`.
- After changing the public Protocol or Server `HttpApi`, run `bun run generate` from `packages/client`. Do not edit `src/generated` or `src/generated-effect` directly. - After changing the public Protocol or Server `HttpApi`, run `bun run generate` from `packages/client`. Do not edit `src/generated` or `src/generated-effect` directly.
- Keep runtime dependencies directed from Schema to Core and Protocol, then from Core and Protocol to Server. Client runtime code may depend on Schema and Protocol but never Core or Server; `sdk-next` composes Client, Core, and Server. - Keep runtime dependencies directed from Schema to Core and Protocol, then from Core and Protocol to Server. Client runtime code may depend on Schema and Protocol but never Core or Server; `sdk-next` composes Client, Core, and Server.
- Do not modify `packages/opencode` unless the user explicitly asks for V1 work. `packages/opencode` is the V1 implementation and is present for reference only. New implementation changes should land in the V2 package set: `packages/core`, `packages/cli`, `packages/server`, `packages/protocol`, `packages/schema`, and related generated client surfaces when required. - Do not modify `packages/opencode` unless the user explicitly asks for V1 work. `packages/opencode` is the V1 implementation and is present for reference only. New implementation changes should land in the V2 package set: `packages/core`, `packages/cli`, `packages/server`, `packages/protocol`, `packages/schema`, and related generated client surfaces when required.
@@ -61,6 +60,7 @@ const { a, b } = obj
### Imports ### Imports
- Never alias imports. Do not use `import { foo as bar } from "..."` or renamed imports like `resolve as pathResolve`. - Never alias imports. Do not use `import { foo as bar } from "..."` or renamed imports like `resolve as pathResolve`.
- Never use type-position `import("...")` references such as `Schema.declare<import("@opencode-ai/plugin/effect/plugin").Plugin["effect"]>`. Only when two imports genuinely collide on a name and no other option exists, an aliased type import (`import type { Plugin as PluginDefinition } from "..."`) is permitted as a last resort — still strongly preferred not to.
- Never use star imports. Do not use `import * as Foo from "..."` or `import type * as Foo from "..."`. - Never use star imports. Do not use `import * as Foo from "..."` or `import type * as Foo from "..."`.
- If a namespace-style value is needed, import the module's own exported namespace by name, for example `import { Project } from "@opencode-ai/core/project"`, then reference `Project.ID`. - If a namespace-style value is needed, import the module's own exported namespace by name, for example `import { Project } from "@opencode-ai/core/project"`, then reference `Project.ID`.
- Prefer dynamic imports for heavy modules that are only needed in selected code paths, especially in startup-sensitive entrypoints. Destructure dynamic import bindings near the top of the narrowest scope that needs them so they read like normal imports. Avoid inline chains such as `await import("./module").then((mod) => mod.value())` or `(await import("./module")).value()`. Keep branch-specific imports inside the branch that needs them to preserve lazy loading. - Prefer dynamic imports for heavy modules that are only needed in selected code paths, especially in startup-sensitive entrypoints. Destructure dynamic import bindings near the top of the narrowest scope that needs them so they read like normal imports. Avoid inline chains such as `await import("./module").then((mod) => mod.value())` or `(await import("./module")).value()`. Keep branch-specific imports inside the branch that needs them to preserve lazy loading.
+1934 -1851
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -2,7 +2,7 @@
exact = true exact = true
# Only install newly resolved package versions published at least 3 days ago. # Only install newly resolved package versions published at least 3 days ago.
minimumReleaseAge = 259200 minimumReleaseAge = 259200
minimumReleaseAgeExcludes = ["@ai-sdk/amazon-bedrock", "@ai-sdk/anthropic", "@opentui/core", "@opentui/core-darwin-arm64", "@opentui/core-darwin-x64", "@opentui/core-linux-arm64", "@opentui/core-linux-arm64-musl", "@opentui/core-linux-x64", "@opentui/core-linux-x64-musl", "@opentui/core-win32-arm64", "@opentui/core-win32-x64", "@opentui/keymap", "@opentui/solid", "opentui-spinner", "gitlab-ai-provider", "opencode-gitlab-auth", "@ff-labs/fff-node", "@ff-labs/fff-bun", "@ff-labs/fff-bin-darwin-arm64", "@ff-labs/fff-bin-darwin-x64", "@ff-labs/fff-bin-linux-arm64-gnu", "@ff-labs/fff-bin-linux-arm64-musl", "@ff-labs/fff-bin-linux-x64-gnu", "@ff-labs/fff-bin-linux-x64-musl", "@ff-labs/fff-bin-win32-arm64", "@ff-labs/fff-bin-win32-x64", "@pierre/diffs", "@pierre/theming", "app-builder-lib", "dmg-builder", "electron-builder", "electron-publish"] minimumReleaseAgeExcludes = ["@ai-sdk/amazon-bedrock", "@ai-sdk/anthropic", "@opencode-ai/sdk", "@opentui/core", "@opentui/core-darwin-arm64", "@opentui/core-darwin-x64", "@opentui/core-linux-arm64", "@opentui/core-linux-arm64-musl", "@opentui/core-linux-x64", "@opentui/core-linux-x64-musl", "@opentui/core-win32-arm64", "@opentui/core-win32-x64", "@opentui/keymap", "@opentui/solid", "opentui-spinner", "gitlab-ai-provider", "opencode-gitlab-auth", "@ff-labs/fff-node", "@ff-labs/fff-bun", "@ff-labs/fff-bin-darwin-arm64", "@ff-labs/fff-bin-darwin-x64", "@ff-labs/fff-bin-linux-arm64-gnu", "@ff-labs/fff-bin-linux-arm64-musl", "@ff-labs/fff-bin-linux-x64-gnu", "@ff-labs/fff-bin-linux-x64-musl", "@ff-labs/fff-bin-win32-arm64", "@ff-labs/fff-bin-win32-x64", "@pierre/diffs", "@pierre/theming", "app-builder-lib", "dmg-builder", "electron-builder", "electron-publish"]
[test] [test]
root = "./do-not-run-tests-from-root" root = "./do-not-run-tests-from-root"
+118
View File
@@ -0,0 +1,118 @@
# V1 to V2 Database Migration
## Approach
- Use the `dev` branch database schema and migration registry as the V1 baseline.
- Remove migrations that exist only on the V2 branch.
- Generate one canonical migration from the `dev` schema to the final V2 schema.
- Add explicit data operations to that migration where generated DDL is insufficient.
- Test the migration against a populated database at the exact `dev` schema.
## Preserve
The canonical V1 data remains in its existing tables. In particular, preserve `session`, `message`, and `part` rows.
Preserve `workspace` rows and existing `session.workspace_id` values unchanged. The migration must not clear or rebuild
workspace relationships.
Keep the `todo` table and its data unchanged. V2 does not currently migrate todos into another representation, and the
generated migration must not drop the table.
## Truncate
Truncate these pre-launch V2 tables before applying schema changes:
- `event`
- `event_sequence`
- `session_message`
These rows are not canonical V1 data. Truncating `event` before adding the required `event.created` column means the
column needs neither a backfill nor a default. After truncation, rebuild `session_message` from canonical V1 `message`
and `part` rows rather than retaining its pre-launch V2 contents.
## Message Backfill
Backfill canonical V1 history from `message` and `part` into `session_message`. This is the main data transformation in
the migration. Preserving the V1 tables alone keeps the data safe but does not make existing history visible through the
V2 session APIs, which read `session_message`.
Reuse each V1 `message.id` as the corresponding `session_message.id`. Stable IDs keep the migration deterministic and
avoid rewriting other persisted state that may refer to a message.
Within each session, order V1 messages by `time_created` and then `id`, matching the existing V1 message index. Assign
contiguous `session_message.seq` values starting at `0`.
Map ordinary V1 messages one-to-one by role. Each ordinary V1 user message becomes one V2 `user` row, and each ordinary
V1 assistant message becomes one V2 `assistant` row. Fold the source message's ordered V1 parts into that row's V2
payload.
Handle semantic marker parts before applying the ordinary mapping. In particular, a V1 user message containing a
`compaction` part and its paired assistant summary represent one compaction operation, not two ordinary messages. Special
part mappings must be decided explicitly before implementing the backfill.
V1 synthetic content is represented by user text parts with `synthetic: true`, not by a separate message role. A V1 user
message whose visible text parts are all synthetic should become a V2 `synthetic` message. If a V1 user message mixes
ordinary and synthetic content, preserve the ordinary content in the V2 `user` row and emit the synthetic content as an
adjacent V2 `synthetic` row. Ignore text parts marked `ignored`, matching V1 model-history behavior.
Use the V1 compaction user message ID as the ID of the collapsed V2 compaction message. This matches V2's use of the
admitted compaction input ID and preserves references to the initiating message.
For a completed compaction, create one V2 `compaction` row with `status: "completed"`. Set `reason` from the V1
compaction part's `auto` flag, join the paired summary assistant's nonempty text parts with blank lines for `summary`, and
serialize the retained V1 tail beginning at `tail_start_id` for `recent`. Use an empty `recent` value when no tail was
retained, and use the compaction user message creation time. Do not emit the paired summary assistant as a separate V2
assistant row.
After rebuilding `session_message`, seed `event_sequence` with one row per migrated session. Set its watermark to that
session's maximum backfilled `session_message.seq`. This prevents new V2 events from reusing sequence numbers or sorting
before migrated history. The `event` table remains empty.
## Drop
Drop these pre-launch V2 tables without preserving or transforming their rows:
- `session_input`
- `session_context_epoch`
Do not transfer `session_input` rows into `session_pending`.
## Create Empty
Let the generated migration create these tables empty:
- `instruction_blob`
- `instruction_entry`
- `instruction_state`
- `session_pending`
- `kv`
V1 has no canonical data to backfill into these tables. V2 initializes their state as it runs.
## Fork Storage
V1 has no fork-boundary state to backfill. New V2 forks use a required message boundary and persist it in
`session.fork_boundary`. The durable fork event contains no parent sequence. Its resolved boundary is one of:
- `before`: copy messages before the identified message.
- `through`: copy messages through the identified message.
Forking an empty session is not supported. `session.fork_seq` and `session.fork_message_id` are not part of the final V2
schema.
New nullable session columns, including `fork_session_id`, `fork_boundary`, and `time_suspended`, require no explicit
backfill. Existing rows naturally receive `NULL` when the generated migration adds the columns.
## Verification
The canonical migration test should seed representative V1 sessions, messages, parts, todos, projects, accounts,
credentials, permissions, shares, and workspaces. After migration, it should verify:
- Preserved rows and encoded values remain unchanged.
- Todo rows remain available in the unchanged `todo` table.
- `event` is empty, and stale pre-launch rows are absent from the rebuilt projections.
- Backfilled `session_message` rows represent the canonical V1 `message` and `part` history.
- Each migrated session's `event_sequence` watermark matches its maximum backfilled message sequence.
- Dropped tables no longer exist.
- New tables exist and are empty.
- The final schema has no ungenerated changes.
+1 -1
View File
@@ -15,6 +15,6 @@
"@actions/github": "6.0.1", "@actions/github": "6.0.1",
"@octokit/graphql": "9.0.1", "@octokit/graphql": "9.0.1",
"@octokit/rest": "catalog:", "@octokit/rest": "catalog:",
"@opencode-ai/sdk": "workspace:*" "@opencode-ai/sdk": "1.18.5"
} }
} }
+4 -4
View File
@@ -1,8 +1,8 @@
{ {
"nodeModules": { "nodeModules": {
"x86_64-linux": "sha256-qt11SKmOjq0KU542QFbs+u7YyJicn4drCcwCdg325yk=", "x86_64-linux": "sha256-RFek0QoEEjsgbqmTE/SxQAmPtYyzs0IPR2ugFn5Okrs=",
"aarch64-linux": "sha256-z68doReXTrWS7HeiAjc0btIjAsvzeZZ7hXAlHr0c77Q=", "aarch64-linux": "sha256-BmAxapY1YrAFn7mVq3/6A9+6Au5UIvSqBboHMkyJH3I=",
"aarch64-darwin": "sha256-PILYH1Pi8XBvSkuZ+1sNnUTao5kba+m5Z8iJKx6YXPo=", "aarch64-darwin": "sha256-Sx3bGWQqLlgoa/RudJxanjSzhFRNklckT2ffnO2I5F4=",
"x86_64-darwin": "sha256-KpcJzP4m0SUavu/WaSffgzOxrHq8ljdy0GOzs9p16lo=" "x86_64-darwin": "sha256-CMOhiisHNowg06qadvgg4K+60zrynglwiT0qKYQ4NiA="
} }
} }
+10 -9
View File
@@ -33,13 +33,12 @@
"packages/*", "packages/*",
"packages/console/*", "packages/console/*",
"packages/stats/*", "packages/stats/*",
"packages/sdk/js",
"packages/slack" "packages/slack"
], ],
"catalog": { "catalog": {
"@effect/opentelemetry": "4.0.0-beta.98", "@effect/opentelemetry": "4.0.0-beta.101",
"@effect/platform-node": "4.0.0-beta.98", "@effect/platform-node": "4.0.0-beta.101",
"@effect/sql-sqlite-bun": "4.0.0-beta.98", "@effect/sql-sqlite-bun": "4.0.0-beta.101",
"@npmcli/arborist": "9.4.0", "@npmcli/arborist": "9.4.0",
"@types/bun": "1.3.13", "@types/bun": "1.3.13",
"@types/cross-spawn": "6.0.6", "@types/cross-spawn": "6.0.6",
@@ -51,6 +50,7 @@
"@opentui/solid": "0.4.5", "@opentui/solid": "0.4.5",
"@tanstack/solid-virtual": "3.13.32", "@tanstack/solid-virtual": "3.13.32",
"@shikijs/stream": "4.2.0", "@shikijs/stream": "4.2.0",
"@standard-schema/spec": "1.1.0",
"ulid": "3.0.1", "ulid": "3.0.1",
"@kobalte/core": "0.13.11", "@kobalte/core": "0.13.11",
"@corvu/drawer": "0.2.4", "@corvu/drawer": "0.2.4",
@@ -69,7 +69,7 @@
"dompurify": "3.3.1", "dompurify": "3.3.1",
"drizzle-kit": "1.0.0-rc.2", "drizzle-kit": "1.0.0-rc.2",
"drizzle-orm": "1.0.0-rc.2", "drizzle-orm": "1.0.0-rc.2",
"effect": "4.0.0-beta.98", "effect": "4.0.0-beta.101",
"ai": "6.0.168", "ai": "6.0.168",
"cross-spawn": "7.0.6", "cross-spawn": "7.0.6",
"hono": "4.10.7", "hono": "4.10.7",
@@ -124,7 +124,7 @@
"@aws-sdk/client-s3": "3.933.0", "@aws-sdk/client-s3": "3.933.0",
"@opencode-ai/plugin": "workspace:*", "@opencode-ai/plugin": "workspace:*",
"@opencode-ai/script": "workspace:*", "@opencode-ai/script": "workspace:*",
"@opencode-ai/sdk": "workspace:*", "@opencode-ai/sdk": "1.18.5",
"heap-snapshot-toolkit": "1.1.3", "heap-snapshot-toolkit": "1.1.3",
"typescript": "catalog:" "typescript": "catalog:"
}, },
@@ -152,21 +152,22 @@
"@opentui/keymap": "catalog:", "@opentui/keymap": "catalog:",
"@opentui/solid": "catalog:", "@opentui/solid": "catalog:",
"@types/bun": "catalog:", "@types/bun": "catalog:",
"@types/node": "catalog:" "@types/node": "catalog:",
"effect": "catalog:"
}, },
"patchedDependencies": { "patchedDependencies": {
"@ff-labs/fff-bun@0.9.3": "patches/@ff-labs%2Ffff-bun@0.9.3.patch",
"@npmcli/agent@4.0.2": "patches/@npmcli%2Fagent@4.0.2.patch", "@npmcli/agent@4.0.2": "patches/@npmcli%2Fagent@4.0.2.patch",
"@silvia-odwyer/photon-node@0.3.4": "patches/@silvia-odwyer%2Fphoton-node@0.3.4.patch", "@silvia-odwyer/photon-node@0.3.4": "patches/@silvia-odwyer%2Fphoton-node@0.3.4.patch",
"@standard-community/standard-openapi@0.2.9": "patches/@standard-community%2Fstandard-openapi@0.2.9.patch", "@standard-community/standard-openapi@0.2.9": "patches/@standard-community%2Fstandard-openapi@0.2.9.patch",
"solid-js@1.9.10": "patches/solid-js@1.9.10.patch", "solid-js@1.9.10": "patches/solid-js@1.9.10.patch",
"@ai-sdk/xai@3.0.102": "patches/@ai-sdk%2Fxai@3.0.102.patch", "@ai-sdk/xai@3.0.102": "patches/@ai-sdk%2Fxai@3.0.102.patch",
"@ai-sdk/mistral@3.0.51": "patches/@ai-sdk%2Fmistral@3.0.51.patch",
"gcp-metadata@8.1.2": "patches/gcp-metadata@8.1.2.patch", "gcp-metadata@8.1.2": "patches/gcp-metadata@8.1.2.patch",
"pacote@21.5.0": "patches/pacote@21.5.0.patch", "pacote@21.5.0": "patches/pacote@21.5.0.patch",
"@ai-sdk/google@3.0.73": "patches/@ai-sdk%2Fgoogle@3.0.73.patch", "@ai-sdk/google@3.0.73": "patches/@ai-sdk%2Fgoogle@3.0.73.patch",
"@pierre/trees@1.0.0-beta.4": "patches/@pierre%2Ftrees@1.0.0-beta.4.patch", "@pierre/trees@1.0.0-beta.4": "patches/@pierre%2Ftrees@1.0.0-beta.4.patch",
"@modelcontextprotocol/sdk@1.29.0": "patches/@modelcontextprotocol%2Fsdk@1.29.0.patch", "@modelcontextprotocol/sdk@1.29.0": "patches/@modelcontextprotocol%2Fsdk@1.29.0.patch",
"effect@4.0.0-beta.98": "patches/effect@4.0.0-beta.98.patch", "effect@4.0.0-beta.101": "patches/effect@4.0.0-beta.101.patch",
"@tanstack/virtual-core@3.17.3": "patches/@tanstack%2Fvirtual-core@3.17.3.patch" "@tanstack/virtual-core@3.17.3": "patches/@tanstack%2Fvirtual-core@3.17.3.patch"
} }
} }
+11 -8
View File
@@ -12,6 +12,8 @@
Per-type constructors live on the type, not as top-level re-exports. Use `Message.system(...)`, `Message.user(...)`, `Message.assistant(...)`, `Message.tool(...)`, `Model.make(...)`, `ToolDefinition.make(...)`, `ToolCallPart.make(...)`, `ToolResultPart.make(...)`, `ToolChoice.make(...)`, `ToolChoice.named(...)`, `SystemPart.make(...)`, and `GenerationOptions.make(...)` directly. The top-level `LLM` namespace is reserved for request-shaped call APIs: `LLM.request`, `LLM.generate`, `LLM.stream`, `LLM.updateRequest`, and `LLM.generateObject`. Two ways to construct the same thing is one too many. Per-type constructors live on the type, not as top-level re-exports. Use `Message.system(...)`, `Message.user(...)`, `Message.assistant(...)`, `Message.tool(...)`, `Model.make(...)`, `ToolDefinition.make(...)`, `ToolCallPart.make(...)`, `ToolResultPart.make(...)`, `ToolChoice.make(...)`, `ToolChoice.named(...)`, `SystemPart.make(...)`, and `GenerationOptions.make(...)` directly. The top-level `LLM` namespace is reserved for request-shaped call APIs: `LLM.request`, `LLM.generate`, `LLM.stream`, `LLM.updateRequest`, and `LLM.generateObject`. Two ways to construct the same thing is one too many.
- Keep provider-defined string enums forward-compatible. Expose known values for autocomplete while accepting future values with `Known | (string & {})`; use `Schema.String` at runtime unless rejecting unknown values is required for correctness.
## Tests ## Tests
- Use `testEffect(...)` from `test/lib/effect.ts` for tests requiring Effect layers. - Use `testEffect(...)` from `test/lib/effect.ts` for tests requiring Effect layers.
@@ -46,7 +48,7 @@ const response = yield * LLMClient.generate(request)
`LLM.request(...)` builds an `LLMRequest`. `LLMClient.generate(...)` reads the executable route carried by `request.model.route`, builds the provider-native body, asks the route's transport for a real `HttpClientRequest.HttpClientRequest`, sends it through `RequestExecutor.Service`, parses the provider stream into common `LLMEvent`s, and finally returns an `LLMResponse`. `LLM.request(...)` builds an `LLMRequest`. `LLMClient.generate(...)` reads the executable route carried by `request.model.route`, builds the provider-native body, asks the route's transport for a real `HttpClientRequest.HttpClientRequest`, sends it through `RequestExecutor.Service`, parses the provider stream into common `LLMEvent`s, and finally returns an `LLMResponse`.
Use `LLMClient.stream(request)` when callers want incremental `LLMEvent`s. Use `LLMClient.generate(request)` when callers want those same events collected into an `LLMResponse`. Use `LLMClient.prepare<Body>(request)` to compile a request through the route pipeline without sending it — the optional `Body` type argument narrows `.body` to the route's native shape (e.g. `prepare<OpenAIChatBody>(...)` returns a `PreparedRequestOf<OpenAIChatBody>`). The runtime body is identical; the generic is a type-level assertion. Use `LLMClient.stream(request)` when callers want incremental `LLMEvent`s. Use `LLMClient.generate(request)` when callers want those same events collected into an `LLMResponse`.
Filter or narrow `LLMEvent` streams with `LLMEvent.is.*` (camelCase guards, e.g. `events.filter(LLMEvent.is.toolCall)`). The kebab-case `LLMEvent.guards["tool-call"]` form also works but prefer `is.*` in new code. Filter or narrow `LLMEvent` streams with `LLMEvent.is.*` (camelCase guards, e.g. `events.filter(LLMEvent.is.toolCall)`). The kebab-case `LLMEvent.guards["tool-call"]` form also works but prefer `is.*` in new code.
@@ -54,7 +56,7 @@ Filter or narrow `LLMEvent` streams with `LLMEvent.is.*` (camelCase guards, e.g.
A route is the registered, runnable composition of four orthogonal pieces: A route is the registered, runnable composition of four orthogonal pieces:
- **`Protocol`** (`src/route/protocol.ts`) — semantic API contract. Owns request body construction (`body.from`), the body schema (`body.schema`), the streaming-event schema (`stream.event`), and the event-to-`LLMEvent` state machine (`stream.step`). `Route.make(...)` validates and JSON-encodes the body from `body.schema` and decodes frames with `stream.event`. Examples: `OpenAIChat.protocol`, `OpenAIResponses.protocol`, `AnthropicMessages.protocol`, `Gemini.protocol`, `BedrockConverse.protocol`. - **`Protocol`** (`src/route/protocol.ts`) — semantic API contract. Owns request body construction (`body.from`), the body schema (`body.schema`), the streaming-event schema (`stream.event`), and the event-to-`LLMEvent` state machine (`stream.step`). `Route.make(...)` validates and JSON-encodes the body from `body.schema` and decodes frames with `stream.event`. Examples: `OpenAIChat.protocol`, `OpenResponses.protocol`, `OpenAIResponses.protocol`, `AnthropicMessages.protocol`, `Gemini.protocol`, `BedrockConverse.protocol`.
- **`Endpoint`** (`src/route/endpoint.ts`) — URL construction. The host, path, and route query live on the endpoint. `Endpoint.path("/chat/completions", { baseURL })` is the common case; pass a function for paths that embed the model id or a body field (e.g. `Endpoint.path(({ body }) => `/model/${body.modelId}/converse-stream`)`). - **`Endpoint`** (`src/route/endpoint.ts`) — URL construction. The host, path, and route query live on the endpoint. `Endpoint.path("/chat/completions", { baseURL })` is the common case; pass a function for paths that embed the model id or a body field (e.g. `Endpoint.path(({ body }) => `/model/${body.modelId}/converse-stream`)`).
- **`Auth`** (`src/route/auth.ts`) — per-request transport authentication. Provider facades configure credentials onto the route before model selection, usually via `Auth.bearer(apiKey)` or `Auth.header(name, apiKey)`. Routes that need per-request signing (Bedrock SigV4, future Vertex IAM, Azure AAD) implement `Auth` as a function that signs the body and merges signed headers into the result. - **`Auth`** (`src/route/auth.ts`) — per-request transport authentication. Provider facades configure credentials onto the route before model selection, usually via `Auth.bearer(apiKey)` or `Auth.header(name, apiKey)`. Routes that need per-request signing (Bedrock SigV4, future Vertex IAM, Azure AAD) implement `Auth` as a function that signs the body and merges signed headers into the result.
- **`Framing`** (`src/route/framing.ts`) — bytes → frames. SSE (`Framing.sse`) is shared; Bedrock keeps its AWS event-stream framing as a typed `Framing<object>` value alongside its protocol. - **`Framing`** (`src/route/framing.ts`) — bytes → frames. SSE (`Framing.sse`) is shared; Bedrock keeps its AWS event-stream framing as a typed `Framing<object>` value alongside its protocol.
@@ -138,13 +140,13 @@ packages/ai/src/
ids.ts branded IDs, literal types, ProviderMetadata ids.ts branded IDs, literal types, ProviderMetadata
options.ts Generation/Provider/Http options, Limits, Model, cache policy options.ts Generation/Provider/Http options, Limits, Model, cache policy
messages.ts content parts, Message, ToolDefinition, LLMRequest messages.ts content parts, Message, ToolDefinition, LLMRequest
events.ts Usage, individual events, LLMEvent, PreparedRequest, LLMResponse events.ts Usage, individual events, LLMEvent, LLMResponse
errors.ts error reasons, LLMError, ToolFailure errors.ts error reasons, LLMError, ToolFailure
index.ts barrel index.ts barrel
llm.ts request constructors and convenience helpers llm.ts request constructors and convenience helpers
route/ route/
index.ts @opencode-ai/ai/route advanced barrel index.ts @opencode-ai/ai/route advanced barrel
client.ts Route.make + LLMClient.prepare/stream/generate client.ts Route.make + LLMClient.stream/generate
executor.ts RequestExecutor service + transport error mapping executor.ts RequestExecutor service + transport error mapping
protocol.ts Protocol type + Protocol.make protocol.ts Protocol type + Protocol.make
endpoint.ts Endpoint type + Endpoint.path endpoint.ts Endpoint type + Endpoint.path
@@ -158,13 +160,14 @@ packages/ai/src/
protocols/ protocols/
shared.ts ProviderShared toolkit used inside protocol impls shared.ts ProviderShared toolkit used inside protocol impls
openai-chat.ts protocol + route (compose OpenAIChat.protocol) openai-chat.ts protocol + route (compose OpenAIChat.protocol)
openai-responses.ts open-responses.ts provider-neutral Responses protocol baseline
openai-responses.ts OpenAI tools/events/transports composed over OpenResponses
anthropic-messages.ts anthropic-messages.ts
gemini.ts gemini.ts
bedrock-converse.ts bedrock-converse.ts
bedrock-event-stream.ts framing for AWS event-stream binary frames bedrock-event-stream.ts framing for AWS event-stream binary frames
openai-compatible-chat.ts route that reuses OpenAIChat.protocol, no canonical URL openai-compatible-chat.ts route that reuses OpenAIChat.protocol, no canonical URL
openai-compatible-responses.ts route that reuses OpenAIResponses.protocol, no canonical URL openai-compatible-responses.ts deployment adapter that reuses OpenResponses.protocol, no canonical URL
utils/ per-protocol helpers (auth, cache, media, tool-stream, ...) utils/ per-protocol helpers (auth, cache, media, tool-stream, ...)
providers/ providers/
openai-compatible.ts generic Chat helper + family model helpers openai-compatible.ts generic Chat helper + family model helpers
@@ -175,7 +178,7 @@ packages/ai/src/
tool-runtime.ts narrow one-call typed tool dispatcher tool-runtime.ts narrow one-call typed tool dispatcher
``` ```
The dependency arrow points down: `providers/*.ts` files import protocol routes and auth-option utilities; protocol modules import `endpoint`, `auth`, `framing`, and transport pieces. Protocols do not import provider facades. Lower-level modules know nothing about provider catalog metadata. The dependency arrow points down: `providers/*.ts` files import protocol routes and auth-option utilities; protocol modules import `endpoint`, `auth`, `framing`, and transport pieces. Protocols do not import provider facades. Lower-level modules know nothing about provider catalog metadata. `OpenAIResponses` composes the provider-neutral `OpenResponses` protocol; the baseline never imports the OpenAI extension.
### Shared protocol helpers ### Shared protocol helpers
@@ -240,7 +243,7 @@ const get_weather = tool({
const tools = { get_weather, get_time, ... } const tools = { get_weather, get_time, ... }
const events = yield* LLM.stream( const events = yield* LLM.stream(
LLM.updateRequest(request, { tools: Tool.toDefinitions(tools) }), LLMRequest.update(request, { tools: Tool.toDefinitions(tools) }),
).pipe(Stream.runCollect) ).pipe(Stream.runCollect)
const call = Array.from(events).find(LLMEvent.is.toolCall) const call = Array.from(events).find(LLMEvent.is.toolCall)
+3 -2
View File
@@ -315,7 +315,8 @@ const longer = {
} }
``` ```
There is no `LLM.updateRequest(...)` helper and no request Schema class. There is no `LLM.updateRequest(...)` helper. The current Schema-backed implementation
uses `LLMRequest.update(...)` when canonical request data must be derived.
### Conversation history ### Conversation history
@@ -436,7 +437,7 @@ const call = Array.from(events).find(LLMEvent.is.toolCall)
if (call && !call.providerExecuted) { if (call && !call.providerExecuted) {
const dispatched = yield * ToolRuntime.dispatch(tools, call) const dispatched = yield * ToolRuntime.dispatch(tools, call)
const followUp = LLM.updateRequest(request, { const followUp = LLMRequest.update(request, {
messages: [...request.messages, Message.assistant([call]), Message.tool({ ...call, result: dispatched.result })], messages: [...request.messages, Message.assistant([call]), Message.tool({ ...call, result: dispatched.result })],
}) })
// Caller must invoke the provider again and repeat the loop. // Caller must invoke the provider again and repeat the loop.
+7 -6
View File
@@ -196,7 +196,6 @@ The hosted result is represented as a provider-executed tool call and tool resul
- **`LLM.generate` / `LLM.stream`** — re-exported from `LLMClient` for one-import use. - **`LLM.generate` / `LLM.stream`** — re-exported from `LLMClient` for one-import use.
- **`Message.user(...)` / `Message.assistant(...)` / `Message.tool(...)`** — message constructors from the canonical schema model. - **`Message.user(...)` / `Message.assistant(...)` / `Message.tool(...)`** — message constructors from the canonical schema model.
- **`Model.make(...)` / `ToolCallPart.make(...)` / `ToolResultPart.make(...)` / `ToolDefinition.make(...)`** — model and tool-related constructors from the canonical schema model. - **`Model.make(...)` / `ToolCallPart.make(...)` / `ToolResultPart.make(...)` / `ToolDefinition.make(...)`** — model and tool-related constructors from the canonical schema model.
- **`LLMClient.prepare(request)`** — compile a request through protocol body construction, validation, and HTTP preparation without sending. Useful for inspection and testing.
- **`LLMEvent.is.*`** — typed guards (`is.textDelta`, `is.toolCall`, `is.finish`, …) for filtering streams. - **`LLMEvent.is.*`** — typed guards (`is.textDelta`, `is.toolCall`, `is.finish`, …) for filtering streams.
- **`Image.generate({...})`** — generate images through a provider-neutral image request and response model. - **`Image.generate({...})`** — generate images through a provider-neutral image request and response model.
- **`ImageClient`** — Effect service and layer for image execution, parallel to `LLMClient`. - **`ImageClient`** — Effect service and layer for image execution, parallel to `LLMClient`.
@@ -207,7 +206,9 @@ Prompt caching is **on by default**. Every `LLMRequest` resolves to `cache: "aut
### Auto placement ### Auto placement
`"auto"` places three breakpoints — last tool definition, last system part, latest user message. The last-user-message boundary is the load-bearing detail: in a tool-use loop, a single user turn expands into many assistant/tool round-trips, all sharing that prefix. Caching at that boundary lets every intra-turn API call hit. `"auto"` places up to four breakpoints — the last tool definition, the first system part, the last system part when distinct, and the final message boundary. These expose successively larger reusable prefixes for tools, the base agent, project instructions, and the active conversation. The rolling final-message boundary is the load-bearing detail in tool loops: it advances on every request so the previous cache entry stays within Anthropic's 20-block lookback.
Tools precede every system and conversation block in the provider prefix, so tool definitions must remain byte-stable and deterministically ordered for downstream breakpoints to remain reusable.
The math justifies the default: Anthropic's 5-minute cache write is 1.25× base, read is 0.1×, so a single reuse within 5 minutes already wins. One-shot completions below the per-model minimum-cacheable-token threshold silently no-op on the wire, so the worst case is harmless. The math justifies the default: Anthropic's 5-minute cache write is 1.25× base, read is 0.1×, so a single reuse within 5 minutes already wins. One-shot completions below the per-model minimum-cacheable-token threshold silently no-op on the wire, so the worst case is harmless.
@@ -235,7 +236,7 @@ cache: {
### Manual hints ### Manual hints
Inline `CacheHint` on any text / system / tool / tool-result part overrides automatic placement. The auto policy preserves manual hints; it only fills gaps. Inline `CacheHint` on any text / system / tool / tool-result part overrides automatic placement. The auto policy preserves manual hints, counts them against Anthropic and Bedrock's four-breakpoint limit, and only fills the remaining slots.
```ts ```ts
LLM.request({ LLM.request({
@@ -251,8 +252,8 @@ LLM.request({
| Protocol | `cache: "auto"` | | Protocol | `cache: "auto"` |
| ----------------------- | ------------------------------------------------------------------------- | | ----------------------- | ------------------------------------------------------------------------- |
| Anthropic Messages | emits up to 3 `cache_control` markers (4-breakpoint cap enforced) | | Anthropic Messages | emits up to 4 `cache_control` markers (4-breakpoint cap enforced) |
| Bedrock Converse | emits up to 3 `cachePoint` blocks (4-breakpoint cap enforced) | | Bedrock Converse | emits up to 4 `cachePoint` blocks (4-breakpoint cap enforced) |
| OpenAI Chat / Responses | no-op (implicit caching above 1024 tokens) | | OpenAI Chat / Responses | no-op (implicit caching above 1024 tokens) |
| Gemini | no-op (implicit caching on 2.5+; explicit `CachedContent` is out-of-band) | | Gemini | no-op (implicit caching on 2.5+; explicit `CachedContent` is out-of-band) |
@@ -300,7 +301,7 @@ OpenAI Chat and OpenAI Responses are separate semantic entrypoints:
- `@opencode-ai/ai/providers/google-vertex/responses` - `@opencode-ai/ai/providers/google-vertex/responses`
- `@opencode-ai/ai/providers/google-vertex/messages` - `@opencode-ai/ai/providers/google-vertex/messages`
Responses HTTP versus WebSocket is a scoped `transport` setting on the OpenAI Responses entrypoint, not another entrypoint. Azure follows the same Chat/Responses split at `providers/azure/chat` and `providers/azure/responses`. Generic OpenAI-compatible Chat remains at `providers/openai-compatible`; compatible Responses is separate at `providers/openai-compatible/responses`. Generic Anthropic Messages-compatible providers use `providers/anthropic-compatible`, which the named Anthropic provider composes. Google Gemini and Amazon Bedrock expose their single native API through their existing provider paths. Responses HTTP versus WebSocket is a scoped `transport` setting on the OpenAI Responses entrypoint, not another entrypoint. Azure follows the same Chat/Responses split at `providers/azure/chat` and `providers/azure/responses`. Generic OpenAI-compatible Chat remains at `providers/openai-compatible`; the Responses adapter at `providers/openai-compatible/responses` uses the provider-neutral Open Responses protocol. OpenAI Responses extends that baseline with OpenAI tools, event variants, metadata, defaults, and transports. Generic Anthropic Messages-compatible providers use `providers/anthropic-compatible`, which the named Anthropic provider composes. Google Gemini and Amazon Bedrock expose their single native API through their existing provider paths.
Vertex Gemini, Vertex Chat, Vertex Responses, and Vertex Messages are separate API entrypoints. All accept `project`, `location`, and an optional `accessToken`; when no explicit token or auth override is supplied they lazily use Google Application Default Credentials. Vertex Gemini instead selects express mode when `apiKey` or `GOOGLE_VERTEX_API_KEY` is present. Vertex Chat targets MaaS models through the OpenAI-compatible Chat Completions endpoint, while Vertex Responses targets Grok models and defaults `store` to `false` as required by Vertex. `providers/google-vertex` remains the default alias for `providers/google-vertex/gemini`. Vertex Gemini, Vertex Chat, Vertex Responses, and Vertex Messages are separate API entrypoints. All accept `project`, `location`, and an optional `accessToken`; when no explicit token or auth override is supplied they lazily use Google Application Default Credentials. Vertex Gemini instead selects express mode when `apiKey` or `GOOGLE_VERTEX_API_KEY` is present. Vertex Chat targets MaaS models through the OpenAI-compatible Chat Completions endpoint, while Vertex Responses targets Grok models and defaults `store` to `false` as required by Vertex. `providers/google-vertex` remains the default alias for `providers/google-vertex/gemini`.
+24 -24
View File
@@ -1,6 +1,6 @@
# LLM Provider Parity Status # LLM Provider Parity Status
Last reviewed: 2026-07-17 Last reviewed: 2026-07-24
This file tracks the gap between the native `@opencode-ai/ai` package and the AI SDK provider packages that opencode still depends on for many catalog/runtime paths. This file tracks the gap between the native `@opencode-ai/ai` package and the AI SDK provider packages that opencode still depends on for many catalog/runtime paths.
@@ -13,26 +13,26 @@ This file tracks the gap between the native `@opencode-ai/ai` package and the AI
## Current Implementation Snapshot ## Current Implementation Snapshot
| Native slice | Source | Current state | Main gaps | | Native slice | Source | Current state | Main gaps |
| ---------------------------------- | ---------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| OpenAI Chat | `src/protocols/openai-chat.ts`, `src/providers/openai.ts` | Usable. Streams text, reasoning deltas, tool calls, usage, images, and common generation controls. | No typed structured-output / `response_format` path. Limited typed OpenAI option surface compared with SDK escape hatches. | | OpenAI Chat | `src/protocols/openai-chat.ts`, `src/providers/openai.ts` | Usable. Streams text, reasoning deltas, tool calls, usage, images, and common generation controls. | No typed structured-output / `response_format` path. Limited typed OpenAI option surface compared with SDK escape hatches. |
| OpenAI Responses HTTP | `src/protocols/openai-responses.ts`, `src/providers/openai.ts` | Usable. Supports hosted-tool event surfacing, reasoning replay metadata, GPT-5 defaults, and cache usage. | No explicit `previous_response_id` path. Typed options cover only a subset of Responses fields. Structured output is still mostly synthetic-tool based. | | OpenAI Responses HTTP | `src/protocols/open-responses.ts`, `src/protocols/openai-responses.ts`, `src/providers/openai.ts` | Usable. Extends the Open Responses baseline with hosted-tool event surfacing, reasoning replay metadata, GPT-5 defaults, and cache usage. | No explicit `previous_response_id` path. Typed options cover only a subset of Responses fields. Structured output is still mostly synthetic-tool based. |
| OpenAI Responses WebSocket | `src/protocols/openai-responses.ts`, `src/route/transport/websocket.ts` | Present as `OpenAI.responsesWebSocket(...)`. | Runner/catalog support explicitly must not downgrade WebSocket routes; broader runtime selection is not complete. | | OpenAI Responses WebSocket | `src/protocols/openai-responses.ts`, `src/route/transport/websocket.ts` | Present as `OpenAI.responsesWebSocket(...)`. | Runner/catalog support explicitly must not downgrade WebSocket routes; broader runtime selection is not complete. |
| OpenAI-compatible Chat | `src/protocols/openai-compatible-chat.ts`, `src/providers/openai-compatible.ts` | Usable for generic Chat and several profiles: Baseten, Cerebras, DeepInfra, DeepSeek, Fireworks, Groq, TogetherAI. | Family quirks are mostly endpoint defaults, not full typed behavior. | | OpenAI-compatible Chat | `src/protocols/openai-compatible-chat.ts`, `src/providers/openai-compatible.ts` | Usable for generic Chat and several profiles: Baseten, Cerebras, DeepInfra, DeepSeek, Fireworks, Groq, TogetherAI. | Family quirks are mostly endpoint defaults, not full typed behavior. |
| OpenAI-compatible Responses | `src/protocols/openai-compatible-responses.ts`, `src/providers/openai-compatible-responses.ts` | Usable for deployments that implement the OpenAI Responses wire protocol. | No named family profiles or recorded deployment coverage yet. | | Open Responses-compatible | `src/protocols/open-responses.ts`, `src/protocols/openai-compatible-responses.ts`, `src/providers/openai-compatible-responses.ts` | Usable for deployments that implement the provider-neutral Open Responses protocol. The deployment adapter does not inherit OpenAI tools, events, metadata, or defaults. | No named family profiles or recorded deployment coverage yet. |
| Anthropic-compatible Messages | `src/protocols/anthropic-messages.ts`, `src/providers/anthropic-compatible.ts` | Usable for deployments that implement the Anthropic Messages wire protocol. Named Anthropic composes this base; MiniMax M3 has recorded text and tool-loop coverage. | No named compatible family profiles yet. | | Anthropic-compatible Messages | `src/protocols/anthropic-messages.ts`, `src/providers/anthropic-compatible.ts` | Usable for deployments that implement the Anthropic Messages wire protocol. Named Anthropic composes this base; MiniMax M3 has recorded text and tool-loop coverage. | No named compatible family profiles yet. |
| Anthropic Messages | `src/protocols/anthropic-messages.ts`, `src/providers/anthropic.ts` | Usable. Supports tools, thinking, cache control, images, server-hosted tool events, and usage. | Provider option surface is small. Beta/header handling, metadata, and newer Messages fields need a typed parity pass. | | Anthropic Messages | `src/protocols/anthropic-messages.ts`, `src/providers/anthropic.ts` | Usable. Supports tools, thinking, cache control, images, server-hosted tool events, and usage. | Provider option surface is small. Beta/header handling, metadata, and newer Messages fields need a typed parity pass. |
| Gemini Developer API | `src/protocols/gemini.ts`, `src/providers/google.ts` | Usable for Google API key flow. Supports text, images, tools, thinking signatures, and cache usage. | This is not Vertex. Typed provider options are narrow; many Gemini request fields currently require raw `http.body` overlays. | | Gemini Developer API | `src/protocols/gemini.ts`, `src/providers/google.ts` | Usable for Google API key flow. Supports text, images, tools, thinking signatures, and cache usage. | This is not Vertex. Typed provider options are narrow; many Gemini request fields currently require raw `http.body` overlays. |
| Vertex Gemini | `src/protocols/gemini.ts`, `src/providers/google-vertex.ts` | Usable through API-key express mode, explicit OAuth tokens, or ADC with project/location endpoint derivation, including tuned `endpoints/...` deployments. | Core runner/catalog mapping and recorded provider coverage are missing. | | Vertex Gemini | `src/protocols/gemini.ts`, `src/providers/google-vertex.ts` | Usable through API-key express mode, explicit OAuth tokens, or ADC with project/location endpoint derivation, including tuned `endpoints/...` deployments. | Core runner/catalog mapping and recorded provider coverage are missing. |
| Vertex Chat | `src/protocols/openai-chat.ts`, `src/providers/google-vertex-chat.ts` | Usable for MaaS models through OpenAI-compatible Chat Completions with explicit OAuth tokens or ADC and project/location endpoint derivation. | Core runner/catalog mapping and recorded provider coverage are missing; MaaS family-specific request parity needs review. | | Vertex Chat | `src/protocols/openai-chat.ts`, `src/providers/google-vertex-chat.ts` | Usable for MaaS models through OpenAI-compatible Chat Completions with explicit OAuth tokens or ADC and project/location endpoint derivation. | Core runner/catalog mapping and recorded provider coverage are missing; MaaS family-specific request parity needs review. |
| Vertex Responses | `src/protocols/openai-responses.ts`, `src/providers/google-vertex-responses.ts` | Usable for Grok models through OpenAI-compatible Responses with explicit OAuth tokens or ADC, project/location endpoint derivation, and storage disabled by default. | Core runner/catalog mapping and recorded provider coverage are missing; stateful continuation is not supported by Vertex. | | Vertex Responses | `src/protocols/open-responses.ts`, `src/providers/google-vertex-responses.ts` | Usable for Grok models through Open Responses with explicit OAuth tokens or ADC, project/location endpoint derivation, and an explicit `store: false` Vertex default. | Core runner/catalog mapping and recorded provider coverage are missing; stateful continuation is not supported by Vertex. |
| Vertex Messages | `src/protocols/anthropic-messages.ts`, `src/providers/google-vertex-messages.ts` | Usable through explicit OAuth tokens or ADC, including global, regional, and `eu`/`us` multi-region endpoints. | Core runner/catalog mapping and recorded provider coverage are missing; Vertex-specific hosted-tool parity needs review. | | Vertex Messages | `src/protocols/anthropic-messages.ts`, `src/providers/google-vertex-messages.ts` | Usable through explicit OAuth tokens or ADC, including global, regional, and `eu`/`us` multi-region endpoints. | Core runner/catalog mapping and recorded provider coverage are missing; Vertex-specific hosted-tool parity needs review. |
| Bedrock Converse | `src/protocols/bedrock-converse.ts`, `src/providers/amazon-bedrock.ts` | Partial but real. Supports AWS event-stream framing, SigV4 with supplied credentials, bearer auth, tools, reasoning signatures, media, cache points, and recorded tests. | Native facade does not mirror the AI SDK plugin's default AWS credential chain/profile behavior. Runner/catalog mapping is missing. Guardrails, inference profiles, region-specific model ID fixes, and model-specific request fields need a parity pass. | | Bedrock Converse | `src/protocols/bedrock-converse.ts`, `src/providers/amazon-bedrock.ts` | Partial but real. Supports AWS event-stream framing, SigV4 with supplied credentials, bearer auth, tools, reasoning signatures, media, cache points, and recorded tests. | Native facade does not mirror the AI SDK plugin's default AWS credential chain/profile behavior. Runner/catalog mapping is missing. Guardrails, inference profiles, region-specific model ID fixes, and model-specific request fields need a parity pass. |
| Azure OpenAI | `src/providers/azure.ts` using OpenAI Chat/Responses protocols | Partial. Supports resource/base URL setup, API key auth, API version query, Chat, and Responses selectors. | Core runner does not map `@ai-sdk/azure` to this native facade. AAD/token auth and Azure-specific endpoint variants need review. | | Azure OpenAI | `src/providers/azure.ts` using OpenAI Chat/Responses protocols | Partial. Supports resource/base URL setup, API key auth, API version query, Chat, and Responses selectors. | Core runner does not map `@ai-sdk/azure` to this native facade. AAD/token auth and Azure-specific endpoint variants need review. |
| Cloudflare AI Gateway / Workers AI | `src/providers/cloudflare.ts` | Present via OpenAI-compatible Chat routes. | Useful but not part of the critical AI SDK replacement set yet. Needs per-product recorded coverage before relying on it broadly. | | Cloudflare AI Gateway / Workers AI | `src/providers/cloudflare.ts` | Present via OpenAI-compatible Chat routes. | Useful but not part of the critical AI SDK replacement set yet. Needs per-product recorded coverage before relying on it broadly. |
| OpenRouter | `src/providers/openrouter.ts` | Present with OpenRouter-specific usage/reasoning/prompt-cache options over Chat. | Responses-style OpenRouter support is absent. | | OpenRouter | `src/providers/openrouter.ts` | Present with OpenRouter-specific usage/reasoning/prompt-cache options over Chat. | Responses-style OpenRouter support is absent. |
| xAI | `src/providers/xai.ts` | Present with Responses and Chat selectors. | Needs package-parity review against the AI SDK xAI provider. | | xAI | `src/providers/xai.ts` | Present with Responses and Chat selectors. | Needs package-parity review against the AI SDK xAI provider. |
| GitHub Copilot | `src/providers/github-copilot.ts` | Present as explicit-base-URL OpenAI Chat/Responses facade. | Runtime/catalog integration remains specialized and should stay separate from public OpenAI-compatible defaults. | | GitHub Copilot | `src/providers/github-copilot.ts` | Present as explicit-base-URL OpenAI Chat/Responses facade. | Runtime/catalog integration remains specialized and should stay separate from public OpenAI-compatible defaults. |
## V2 Runner Status ## V2 Runner Status
@@ -65,7 +65,7 @@ Other `aisdk:` packages, including Google Vertex, Azure, and Bedrock, currently
## Highest-Risk Gaps ## Highest-Risk Gaps
1. Runner support is narrower than the LLM package. The package has native provider facades for Google, Azure, and Bedrock, but the V2 Session runner only maps OpenAI, Anthropic, and explicit OpenAI-compatible Chat from `aisdk` catalog metadata. 1. Runner support is narrower than the LLM package. The package has native provider facades for Google, Azure, and Bedrock, but the V2 Session runner only maps OpenAI, Anthropic, and explicit OpenAI-compatible Chat from `aisdk` catalog metadata.
2. OpenAI-compatible Responses is available as a separate package entrypoint, but the V2 runner still maps `@ai-sdk/openai-compatible` to Chat only. Catalog selection must become API-aware before Responses deployments can use it. 2. The Open Responses adapter is available through a separate package entrypoint, but the V2 runner still maps `@ai-sdk/openai-compatible` to Chat only. Catalog selection must become API-aware before Responses deployments can use it.
3. Bedrock native auth is not AI SDK parity. The AI SDK plugin uses the default AWS provider chain, profile, container credentials, and Bedrock bearer token env behavior. Native Bedrock currently expects explicit credentials or bearer auth on the facade. 3. Bedrock native auth is not AI SDK parity. The AI SDK plugin uses the default AWS provider chain, profile, container credentials, and Bedrock bearer token env behavior. Native Bedrock currently expects explicit credentials or bearer auth on the facade.
4. Vertex Gemini, Vertex Chat, Vertex Responses, and Vertex Messages now have native package entrypoints, but the core runner does not map catalog metadata to them yet and recorded provider coverage is still missing. 4. Vertex Gemini, Vertex Chat, Vertex Responses, and Vertex Messages now have native package entrypoints, but the core runner does not map catalog metadata to them yet and recorded provider coverage is still missing.
5. Azure is only a provider facade, not a full runtime replacement. Native Azure exists, but the catalog runner does not select it, and token auth/resource variants need review. 5. Azure is only a provider facade, not a full runtime replacement. Native Azure exists, but the catalog runner does not select it, and token auth/resource variants need review.
@@ -83,13 +83,13 @@ These are implementation/API slices, not separate npm packages.
| OpenAI Chat | `@opencode-ai/ai/providers/openai/chat` | OpenAI `/chat/completions` semantics. | | OpenAI Chat | `@opencode-ai/ai/providers/openai/chat` | OpenAI `/chat/completions` semantics. |
| OpenAI Responses | `@opencode-ai/ai/providers/openai/responses` | OpenAI `/responses` semantics with HTTP/WebSocket selected through settings. | | OpenAI Responses | `@opencode-ai/ai/providers/openai/responses` | OpenAI `/responses` semantics with HTTP/WebSocket selected through settings. |
| OpenAI-compatible Chat | `@opencode-ai/ai/providers/openai-compatible` | Generic OpenAI-compatible `/chat/completions`. | | OpenAI-compatible Chat | `@opencode-ai/ai/providers/openai-compatible` | Generic OpenAI-compatible `/chat/completions`. |
| OpenAI-compatible Responses | `@opencode-ai/ai/providers/openai-compatible/responses` | Generic OpenAI-compatible `/responses`. | | Open Responses-compatible | `@opencode-ai/ai/providers/openai-compatible/responses` | Generic provider-neutral `/responses`. |
| Anthropic-compatible Messages | `@opencode-ai/ai/providers/anthropic-compatible` | Generic Anthropic-compatible `/messages`. | | Anthropic-compatible Messages | `@opencode-ai/ai/providers/anthropic-compatible` | Generic Anthropic-compatible `/messages`. |
| Anthropic Messages | `@opencode-ai/ai/providers/anthropic` | Anthropic Messages API. | | Anthropic Messages | `@opencode-ai/ai/providers/anthropic` | Anthropic Messages API. |
| Gemini Developer API | `@opencode-ai/ai/providers/google` | Google AI Studio Gemini API. | | Gemini Developer API | `@opencode-ai/ai/providers/google` | Google AI Studio Gemini API. |
| Vertex Gemini | `@opencode-ai/ai/providers/google-vertex/gemini` | Vertex Gemini API; `providers/google-vertex` is the default alias. | | Vertex Gemini | `@opencode-ai/ai/providers/google-vertex/gemini` | Vertex Gemini API; `providers/google-vertex` is the default alias. |
| Vertex Chat | `@opencode-ai/ai/providers/google-vertex/chat` | Vertex OpenAI-compatible Chat Completions for MaaS models. | | Vertex Chat | `@opencode-ai/ai/providers/google-vertex/chat` | Vertex OpenAI-compatible Chat Completions for MaaS models. |
| Vertex Responses | `@opencode-ai/ai/providers/google-vertex/responses` | Vertex OpenAI-compatible Responses for Grok models. | | Vertex Responses | `@opencode-ai/ai/providers/google-vertex/responses` | Vertex Open Responses for Grok models. |
| Vertex Messages | `@opencode-ai/ai/providers/google-vertex/messages` | Vertex-hosted Anthropic Messages API. | | Vertex Messages | `@opencode-ai/ai/providers/google-vertex/messages` | Vertex-hosted Anthropic Messages API. |
| Bedrock Converse | `@opencode-ai/ai/providers/amazon-bedrock` | AWS Bedrock Converse API. | | Bedrock Converse | `@opencode-ai/ai/providers/amazon-bedrock` | AWS Bedrock Converse API. |
| Bedrock Mantle | Missing | AWS Bedrock Mantle OpenAI-compatible APIs. | | Bedrock Mantle | Missing | AWS Bedrock Mantle OpenAI-compatible APIs. |
+1 -1
View File
@@ -568,7 +568,7 @@ App boundary = explicit durable-config -> typed-provider call
calling `.model(...)`. calling `.model(...)`.
- [x] Remove request-shaping defaults from `Model`; selected models now carry only - [x] Remove request-shaping defaults from `Model`; selected models now carry only
id, provider, and configured route while defaults live on routes or requests. id, provider, and configured route while defaults live on routes or requests.
- [x] Rework `LLMClient.prepare` / `stream` / `generate` to read - [x] Rework `LLMClient.stream` / `generate` to read
`request.model.route` directly instead of calling `registeredRoute(...)`. `request.model.route` directly instead of calling `registeredRoute(...)`.
- [x] Remove `Route.make(...)` global registration from the normal execution - [x] Remove `Route.make(...)` global registration from the normal execution
path; keep route ids only as diagnostics/provider API labels. path; keep route ids only as diagnostics/provider API labels.
+9 -36
View File
@@ -1,5 +1,5 @@
import { Config, Effect, Formatter, Layer, Schema, Stream } from "effect" import { Config, Effect, Formatter, Layer, Schema, Stream } from "effect"
import { LLM, LLMClient, Message, ProviderID, Tool, ToolRuntime } from "@opencode-ai/ai" import { LLM, LLMClient, LLMRequest, Message, ProviderID, Tool, ToolRuntime } from "@opencode-ai/ai"
import { Route, Auth, Endpoint, Framing, Protocol, RequestExecutor, WebSocketExecutor } from "@opencode-ai/ai/route" import { Route, Auth, Endpoint, Framing, Protocol, RequestExecutor, WebSocketExecutor } from "@opencode-ai/ai/route"
import { OpenAI } from "@opencode-ai/ai/providers" import { OpenAI } from "@opencode-ai/ai/providers"
@@ -50,18 +50,6 @@ const request = LLM.request({
}, },
}) })
// `http` is intentionally not needed for normal calls. This shows the shape for
// newly released provider fields before they deserve a typed provider option.
const rawOverlayExample = LLM.request({
model,
prompt: "Show the final HTTP overlay shape.",
http: {
body: { metadata: { example: "tutorial" } },
headers: { "x-opencode-tutorial": "1" },
query: { debug: "1" },
},
})
// 3. `generate` sends the request and collects the event stream into one // 3. `generate` sends the request and collects the event stream into one
// response object. `response.text` is the collected text output. // response object. `response.text` is the collected text output.
const generateOnce = Effect.gen(function* () { const generateOnce = Effect.gen(function* () {
@@ -78,7 +66,10 @@ const streamText = LLM.stream(request).pipe(
Stream.tap((event) => Stream.tap((event) =>
Effect.sync(() => { Effect.sync(() => {
if (event.type === "text-delta") process.stdout.write(`\ntext: ${event.text}`) if (event.type === "text-delta") process.stdout.write(`\ntext: ${event.text}`)
if (event.type === "finish") process.stdout.write(`\nfinish: ${event.reason}\n`) if (event.type === "finish")
process.stdout.write(
`\nfinish: ${event.reason.normalized}${event.reason.raw ? ` (${event.reason.raw})` : ""}\n`,
)
}), }),
), ),
Stream.runDrain, Stream.runDrain,
@@ -113,7 +104,7 @@ const streamWithTools = Effect.gen(function* () {
// A durable agent would persist these messages before starting another // A durable agent would persist these messages before starting another
// raw model turn. This tutorial keeps the boundary visible instead. // raw model turn. This tutorial keeps the boundary visible instead.
const followUp = LLM.updateRequest(request, { const followUp = LLMRequest.update(request, {
messages: [ messages: [
...request.messages, ...request.messages,
Message.assistant([event]), Message.assistant([event]),
@@ -194,7 +185,7 @@ const FakeProtocol = Protocol.make<FakeBody, string, string, void>({
event: Schema.String, event: Schema.String,
initial: () => undefined, initial: () => undefined,
step: (_, frame) => Effect.succeed([undefined, [{ type: "text-delta", id: "text-0", text: frame }]] as const), step: (_, frame) => Effect.succeed([undefined, [{ type: "text-delta", id: "text-0", text: frame }]] as const),
onHalt: () => [{ type: "finish", reason: "stop" }], onHalt: () => [{ type: "finish", reason: { normalized: "stop" } }],
}, },
}) })
@@ -219,33 +210,15 @@ const FakeEcho = {
}), }),
} }
// `LLMClient.prepare` is the lower-level inspection hook: it compiles through
// body conversion, validation, endpoint, auth, and HTTP construction without
// sending anything over the network.
const inspectFakeProvider = Effect.gen(function* () {
const prepared = yield* LLMClient.prepare(
LLM.request({
model: FakeEcho.configure().model("tiny-echo"),
prompt: "Show me the provider pipeline.",
}),
)
console.log("\n== fake provider prepare ==")
console.log("route:", prepared.route)
console.log("body:", Formatter.formatJson(prepared.body, { space: 2 }))
})
// Provide the LLM runtime and the HTTP request executor once. Keep one path // Provide the LLM runtime and the HTTP request executor once. Keep one path
// enabled at a time so the tutorial can demonstrate generate, prepare, stream, // enabled at a time so the tutorial can demonstrate generate, stream, or
// or tool-loop behavior without spending tokens on every example. // tool-loop behavior without spending tokens on every example.
const requestExecutorLayer = RequestExecutor.fetchLayer const requestExecutorLayer = RequestExecutor.fetchLayer
const llmDeps = Layer.mergeAll(requestExecutorLayer, WebSocketExecutor.layer) const llmDeps = Layer.mergeAll(requestExecutorLayer, WebSocketExecutor.layer)
const llmClientLayer = LLMClient.layer.pipe(Layer.provide(llmDeps)) const llmClientLayer = LLMClient.layer.pipe(Layer.provide(llmDeps))
const program = Effect.gen(function* () { const program = Effect.gen(function* () {
// yield* generateOnce // yield* generateOnce
// yield* inspectFakeProvider
// yield* LLMClient.prepare(rawOverlayExample).pipe(Effect.andThen((prepared) => Effect.sync(() => console.log(prepared.body))))
// yield* streamText // yield* streamText
// yield* generateStructuredObject // yield* generateStructuredObject
// yield* generateDynamicObject.pipe(Effect.andThen((response) => Effect.sync(() => console.log(response.object)))) // yield* generateDynamicObject.pipe(Effect.andThen((response) => Effect.sync(() => console.log(response.object))))
+1
View File
@@ -15,6 +15,7 @@
], ],
"exports": { "exports": {
".": "./src/index.ts", ".": "./src/index.ts",
"./testing": "./src/testing.ts",
"./*": "./src/*.ts" "./*": "./src/*.ts"
}, },
"devDependencies": { "devDependencies": {
+63 -26
View File
@@ -2,32 +2,31 @@
// the policy designates. Runs once at compile time, before the per-protocol // the policy designates. Runs once at compile time, before the per-protocol
// body builder, so the existing inline-hint lowering path handles the rest. // body builder, so the existing inline-hint lowering path handles the rest.
// //
// The default `"auto"` shape places one breakpoint at the last tool definition, // The default `"auto"` shape places breakpoints at the last tool definition,
// one at the last system part, and one at the latest user message. This // the first and last distinct system parts, and the conversation tail. This
// matches what production agent harnesses (LangChain's caching middleware, // exposes reusable tool, base-agent, project, and session prefixes while
// kern-ai's 10x cost-reduction playbook) converge on for tool-use loops: the // advancing the tail after each tool result keeps the previous cache entry
// latest user message stays put while a single turn explodes into many // within Anthropic's 20-block lookback during long agent turns.
// assistant/tool round-trips, so caching at that boundary lets every
// intra-turn API call hit the prefix.
// //
// Manual `cache: CacheHint` placements on individual parts are preserved // Manual `cache: CacheHint` placements on individual parts are preserved and
// this function only fills gaps the caller left empty. // count against the four-breakpoint budget; auto only fills remaining slots.
import { CacheHint, type CachePolicy, type CachePolicyObject } from "./schema/options" import { CacheHint, type CachePolicy, type CachePolicyObject } from "./schema/options"
import { LLMRequest, Message, ToolDefinition, type ContentPart } from "./schema/messages" import { LLMRequest, Message, ToolDefinition, type ContentPart } from "./schema/messages"
const AUTO: CachePolicyObject = { const AUTO: CachePolicyObject = {
tools: true, tools: true,
system: true, system: true,
messages: "latest-user-message", messages: { tail: 1 },
} }
const NONE: CachePolicyObject = {} const NONE: CachePolicyObject = {}
const BREAKPOINT_CAP = 4
// Resolution rules: // Resolution rules:
// - undefined → "auto" — caching is on by default. The math favors it: // - undefined → "auto" — caching is on by default. The math favors it:
// Anthropic 5m-cache write is 1.25x base, read is 0.1x, // Anthropic 5m-cache write is 1.25x base, read is 0.1x,
// so a single reuse within 5 minutes already wins. // so a single reuse within 5 minutes already wins.
// - "auto" → tools + system + latest user msg. // - "auto" → tools + first/last system + final message boundary.
// - "none" → no auto placement; manual `CacheHint`s still flow. // - "none" → no auto placement; manual `CacheHint`s still flow.
// - object form → exactly what the caller asked for. // - object form → exactly what the caller asked for.
const resolve = (policy: CachePolicy | undefined): CachePolicyObject => { const resolve = (policy: CachePolicy | undefined): CachePolicyObject => {
@@ -39,23 +38,37 @@ const resolve = (policy: CachePolicy | undefined): CachePolicyObject => {
// Protocols whose wire format ignores inline cache markers (OpenAI's implicit // Protocols whose wire format ignores inline cache markers (OpenAI's implicit
// prefix caching, Gemini's implicit + out-of-band CachedContent). Skip the // prefix caching, Gemini's implicit + out-of-band CachedContent). Skip the
// whole policy pass for these — emitting hints would be harmless but pointless. // whole policy pass for these — emitting hints would be harmless but pointless.
const RESPECTS_INLINE_HINTS = new Set(["anthropic-messages", "bedrock-converse"]) const RESPECTS_INLINE_HINTS = new Set(["anthropic-messages", "bedrock-converse", "openrouter"])
const makeHint = (ttlSeconds: number | undefined): CacheHint => const makeHint = (ttlSeconds: number | undefined): CacheHint =>
ttlSeconds !== undefined ? new CacheHint({ type: "ephemeral", ttlSeconds }) : new CacheHint({ type: "ephemeral" }) ttlSeconds !== undefined ? new CacheHint({ type: "ephemeral", ttlSeconds }) : new CacheHint({ type: "ephemeral" })
const markLastTool = (tools: ReadonlyArray<ToolDefinition>, hint: CacheHint): ReadonlyArray<ToolDefinition> => { interface Budget {
remaining: number
}
const markLastTool = (
tools: ReadonlyArray<ToolDefinition>,
hint: CacheHint,
budget: Budget,
): ReadonlyArray<ToolDefinition> => {
if (tools.length === 0) return tools if (tools.length === 0) return tools
const last = tools.length - 1 const last = tools.length - 1
if (tools[last]!.cache) return tools if (tools[last]!.cache || budget.remaining === 0) return tools
budget.remaining -= 1
return tools.map((tool, i) => (i === last ? new ToolDefinition({ ...tool, cache: hint }) : tool)) return tools.map((tool, i) => (i === last ? new ToolDefinition({ ...tool, cache: hint }) : tool))
} }
const markLastSystem = (system: LLMRequest["system"], hint: CacheHint): LLMRequest["system"] => { const markSystemBoundaries = (system: LLMRequest["system"], hint: CacheHint, budget: Budget): LLMRequest["system"] => {
if (system.length === 0) return system if (system.length === 0) return system
const last = system.length - 1 let changed = false
if (system[last]!.cache) return system const next = system.map((part, index) => {
return system.map((part, i) => (i === last ? { ...part, cache: hint } : part)) if ((index !== 0 && index !== system.length - 1) || part.cache || budget.remaining === 0) return part
budget.remaining -= 1
changed = true
return { ...part, cache: hint }
})
return changed ? next : system
} }
const lastIndexOfRole = (messages: ReadonlyArray<Message>, role: Message["role"]): number => const lastIndexOfRole = (messages: ReadonlyArray<Message>, role: Message["role"]): number =>
@@ -64,14 +77,20 @@ const lastIndexOfRole = (messages: ReadonlyArray<Message>, role: Message["role"]
// Mark the last text part of `messages[index]`. If no text part exists, mark // Mark the last text part of `messages[index]`. If no text part exists, mark
// the last content part regardless of type — that's the breakpoint position // the last content part regardless of type — that's the breakpoint position
// in tool-result-only messages too. // in tool-result-only messages too.
const markMessageAt = (messages: ReadonlyArray<Message>, index: number, hint: CacheHint): ReadonlyArray<Message> => { const markMessageAt = (
messages: ReadonlyArray<Message>,
index: number,
hint: CacheHint,
budget: Budget,
): ReadonlyArray<Message> => {
if (index < 0 || index >= messages.length) return messages if (index < 0 || index >= messages.length) return messages
const target = messages[index]! const target = messages[index]!
if (target.content.length === 0) return messages if (target.content.length === 0) return messages
const lastTextIndex = target.content.findLastIndex((part) => part.type === "text") const lastTextIndex = target.content.findLastIndex((part) => part.type === "text")
const markAt = lastTextIndex >= 0 ? lastTextIndex : target.content.length - 1 const markAt = lastTextIndex >= 0 ? lastTextIndex : target.content.length - 1
const existing = target.content[markAt]! const existing = target.content[markAt]!
if ("cache" in existing && existing.cache) return messages if (("cache" in existing && existing.cache) || budget.remaining === 0) return messages
budget.remaining -= 1
const nextContent = target.content.map((part, i) => (i === markAt ? ({ ...part, cache: hint } as ContentPart) : part)) const nextContent = target.content.map((part, i) => (i === markAt ? ({ ...part, cache: hint } as ContentPart) : part))
const next = new Message({ ...target, content: nextContent }) const next = new Message({ ...target, content: nextContent })
// Single pass over `messages`, substituting the one updated entry. Long // Single pass over `messages`, substituting the one updated entry. Long
@@ -86,25 +105,43 @@ const markMessages = (
messages: ReadonlyArray<Message>, messages: ReadonlyArray<Message>,
strategy: NonNullable<CachePolicyObject["messages"]>, strategy: NonNullable<CachePolicyObject["messages"]>,
hint: CacheHint, hint: CacheHint,
budget: Budget,
): ReadonlyArray<Message> => { ): ReadonlyArray<Message> => {
if (messages.length === 0) return messages if (messages.length === 0) return messages
if (strategy === "latest-user-message") return markMessageAt(messages, lastIndexOfRole(messages, "user"), hint) if (strategy === "latest-user-message")
if (strategy === "latest-assistant") return markMessageAt(messages, lastIndexOfRole(messages, "assistant"), hint) return markMessageAt(messages, lastIndexOfRole(messages, "user"), hint, budget)
if (strategy === "latest-assistant")
return markMessageAt(messages, lastIndexOfRole(messages, "assistant"), hint, budget)
const start = Math.max(0, messages.length - strategy.tail) const start = Math.max(0, messages.length - strategy.tail)
let next = messages let next = messages
for (let i = start; i < messages.length; i++) next = markMessageAt(next, i, hint) for (let i = start; i < messages.length; i++) next = markMessageAt(next, i, hint, budget)
return next return next
} }
const countHints = (request: LLMRequest) =>
request.tools.reduce((count, tool) => count + (tool.cache === undefined ? 0 : 1), 0) +
request.system.reduce((count, part) => count + (part.cache === undefined ? 0 : 1), 0) +
request.messages.reduce(
(count, message) =>
count +
message.content.reduce(
(contentCount, part) => contentCount + ("cache" in part && part.cache !== undefined ? 1 : 0),
0,
),
0,
)
export const applyCachePolicy = (request: LLMRequest): LLMRequest => { export const applyCachePolicy = (request: LLMRequest): LLMRequest => {
if (!RESPECTS_INLINE_HINTS.has(request.model.route.id)) return request if (!RESPECTS_INLINE_HINTS.has(request.model.route.id)) return request
if (request.model.route.id === "openrouter" && (request.cache === undefined || request.cache === "auto")) return request
const policy = resolve(request.cache) const policy = resolve(request.cache)
if (!policy.tools && !policy.system && !policy.messages) return request if (!policy.tools && !policy.system && !policy.messages) return request
const hint = makeHint(policy.ttlSeconds) const hint = makeHint(policy.ttlSeconds)
const tools = policy.tools ? markLastTool(request.tools, hint) : request.tools const budget = { remaining: Math.max(0, BREAKPOINT_CAP - countHints(request)) }
const system = policy.system ? markLastSystem(request.system, hint) : request.system const tools = policy.tools ? markLastTool(request.tools, hint, budget) : request.tools
const messages = policy.messages ? markMessages(request.messages, policy.messages, hint) : request.messages const system = policy.system ? markSystemBoundaries(request.system, hint, budget) : request.system
const messages = policy.messages ? markMessages(request.messages, policy.messages, hint, budget) : request.messages
if (tools === request.tools && system === request.system && messages === request.messages) return request if (tools === request.tools && system === request.system && messages === request.messages) return request
return LLMRequest.update(request, { tools, system, messages }) return LLMRequest.update(request, { tools, system, messages })
+18 -31
View File
@@ -9,36 +9,28 @@ import {
LLMRequest, LLMRequest,
LLMResponse, LLMResponse,
Message, Message,
type ModelInput as SchemaModelInput, Model,
SystemPart, SystemPart,
ToolChoice, ToolChoice,
ToolDefinition, ToolDefinition,
type ContentPart, type ContentPart,
ToolResultPart, type ModelProviderOptions,
} from "./schema" } from "./schema"
import { make as makeTool, toDefinitions, type ToolSchema } from "./tool" import { make as makeTool, toDefinitions, type ToolSchema } from "./tool"
export type ModelInput = SchemaModelInput
export type MessageInput = Message.Input
export type ToolChoiceInput = ToolChoice.Input
export type ToolChoiceMode = ToolChoice.Mode
export type ToolResultInput = Parameters<typeof ToolResultPart.make>[0]
/** Input accepted by `LLM.request`, normalized into the canonical `LLMRequest` class. */ /** Input accepted by `LLM.request`, normalized into the canonical `LLMRequest` class. */
export type RequestInput = Omit< export type RequestInput<SelectedModel extends Model = Model> = Omit<
ConstructorParameters<typeof LLMRequest>[0], ConstructorParameters<typeof LLMRequest>[0],
"system" | "messages" | "tools" | "toolChoice" | "generation" | "http" | "providerOptions" "model" | "system" | "messages" | "tools" | "toolChoice" | "generation" | "http" | "providerOptions"
> & { > & {
readonly model: SelectedModel
readonly system?: string | SystemPart | ReadonlyArray<SystemPart> readonly system?: string | SystemPart | ReadonlyArray<SystemPart>
readonly prompt?: string | ContentPart | ReadonlyArray<ContentPart> readonly prompt?: string | ContentPart | ReadonlyArray<ContentPart>
readonly messages?: ReadonlyArray<Message | MessageInput> readonly messages?: ReadonlyArray<Message | Message.Input>
readonly tools?: ReadonlyArray<ToolDefinition.Input> readonly tools?: ReadonlyArray<ToolDefinition.Input>
readonly toolChoice?: ToolChoiceInput readonly toolChoice?: ToolChoice.Input
readonly generation?: GenerationOptions.Input readonly generation?: GenerationOptions.Input
readonly providerOptions?: ConstructorParameters<typeof LLMRequest>[0]["providerOptions"] readonly providerOptions?: NoInfer<ModelProviderOptions<SelectedModel>>
readonly http?: HttpOptions.Input readonly http?: HttpOptions.Input
} }
@@ -46,11 +38,7 @@ export const generate = LLMClient.generate
export const stream = LLMClient.stream export const stream = LLMClient.stream
export const requestInput = (input: LLMRequest): RequestInput => ({ export const request = <const SelectedModel extends Model>(input: RequestInput<SelectedModel>) => {
...LLMRequest.input(input),
})
export const request = (input: RequestInput) => {
const { const {
system: requestSystem, system: requestSystem,
prompt, prompt,
@@ -74,14 +62,11 @@ export const request = (input: RequestInput) => {
}) })
} }
export const updateRequest = (input: LLMRequest, patch: Partial<RequestInput>) =>
request({ ...requestInput(input), ...patch })
const GENERATE_OBJECT_TOOL_NAME = "generate_object" const GENERATE_OBJECT_TOOL_NAME = "generate_object"
const GENERATE_OBJECT_TOOL_DESCRIPTION = "Return the structured result by calling this tool." const GENERATE_OBJECT_TOOL_DESCRIPTION = "Return the structured result by calling this tool."
type GenerateObjectBase = Omit<RequestInput, "tools" | "toolChoice" | "responseFormat"> type GenerateObjectBase<SelectedModel extends Model = Model> = Omit<RequestInput<SelectedModel>, "tools" | "toolChoice">
export class GenerateObjectResponse<T> { export class GenerateObjectResponse<T> {
constructor( constructor(
@@ -98,11 +83,13 @@ export class GenerateObjectResponse<T> {
} }
} }
export interface GenerateObjectOptions<S extends ToolSchema<any>> extends GenerateObjectBase { export interface GenerateObjectOptions<S extends ToolSchema<any>, SelectedModel extends Model = Model>
extends GenerateObjectBase<SelectedModel> {
readonly schema: S readonly schema: S
} }
export interface GenerateObjectDynamicOptions extends GenerateObjectBase { export interface GenerateObjectDynamicOptions<SelectedModel extends Model = Model>
extends GenerateObjectBase<SelectedModel> {
/** Raw JSON Schema object describing the expected output shape. */ /** Raw JSON Schema object describing the expected output shape. */
readonly jsonSchema: JsonSchema.JsonSchema readonly jsonSchema: JsonSchema.JsonSchema
} }
@@ -155,11 +142,11 @@ const runGenerateObject = Effect.fn("LLM.generateObject")(function* (
* 2. `jsonSchema: JsonSchema.JsonSchema` — `.object` is `unknown`. Use when * 2. `jsonSchema: JsonSchema.JsonSchema` — `.object` is `unknown`. Use when
* the schema is only available at runtime (MCP, plugin manifests). Caller validates. * the schema is only available at runtime (MCP, plugin manifests). Caller validates.
*/ */
export function generateObject<S extends ToolSchema<any>>( export function generateObject<const SelectedModel extends Model, S extends ToolSchema<any>>(
options: GenerateObjectOptions<S>, options: GenerateObjectOptions<S, SelectedModel>,
): Effect.Effect<GenerateObjectResponse<Schema.Schema.Type<S>>, LLMError> ): Effect.Effect<GenerateObjectResponse<Schema.Schema.Type<S>>, LLMError>
export function generateObject( export function generateObject<const SelectedModel extends Model>(
options: GenerateObjectDynamicOptions, options: GenerateObjectDynamicOptions<SelectedModel>,
): Effect.Effect<GenerateObjectResponse<unknown>, LLMError> ): Effect.Effect<GenerateObjectResponse<unknown>, LLMError>
export function generateObject(options: GenerateObjectOptions<ToolSchema<any>> | GenerateObjectDynamicOptions) { export function generateObject(options: GenerateObjectOptions<ToolSchema<any>> | GenerateObjectDynamicOptions) {
if ("schema" in options) { if ("schema" in options) {
+206 -69
View File
@@ -1,4 +1,5 @@
import { Effect, Schema } from "effect" import { Effect, Schema } from "effect"
import { Tool } from "@opencode-ai/schema/tool"
import { Route } from "../route/client" import { Route } from "../route/client"
import { Auth } from "../route/auth" import { Auth } from "../route/auth"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
@@ -7,16 +8,18 @@ import { Protocol } from "../route/protocol"
import { import {
LLMError, LLMError,
LLMEvent, LLMEvent,
mergeJsonRecords,
Usage, Usage,
type CacheHint, type CacheHint,
type FinishReasonDetails,
type FinishReason, type FinishReason,
type JsonSchema, type JsonSchema,
type LLMRequest, type LLMRequest,
type MediaPart, type MediaPart,
type ProviderOptions,
type ProviderMetadata, type ProviderMetadata,
type ToolCallPart, type ToolCallPart,
type ToolDefinition, type ToolDefinition,
type ToolContent,
type ToolResultPart, type ToolResultPart,
} from "../schema" } from "../schema"
import { JsonObject, optionalArray, optionalNull, ProviderShared } from "./shared" import { JsonObject, optionalArray, optionalNull, ProviderShared } from "./shared"
@@ -31,6 +34,29 @@ const MEDIA_MIMES = new Set<string>([...ProviderShared.IMAGE_MIMES, ...ProviderS
export const DEFAULT_BASE_URL = "https://api.anthropic.com/v1" export const DEFAULT_BASE_URL = "https://api.anthropic.com/v1"
export const PATH = "/messages" export const PATH = "/messages"
export type ThinkingInput =
| {
readonly type: "adaptive"
readonly display?: "summarized" | "omitted"
}
| {
readonly type: "disabled"
}
| ({ readonly type: "enabled" } & (
| { readonly budgetTokens: number; readonly budget_tokens?: number }
| { readonly budgetTokens?: number; readonly budget_tokens: number }
))
export interface OptionsInput {
readonly [key: string]: unknown
readonly thinking?: ThinkingInput
readonly effort?: string
}
export type ProviderOptionsInput = ProviderOptions & {
readonly anthropic?: OptionsInput
}
// ============================================================================= // =============================================================================
// Request Body Schema // Request Body Schema
// ============================================================================= // =============================================================================
@@ -75,6 +101,15 @@ const AnthropicThinkingBlock = Schema.Struct({
cache_control: Schema.optional(AnthropicCacheControl), cache_control: Schema.optional(AnthropicCacheControl),
}) })
// Safety-filtered thinking arrives as an opaque encrypted `data` payload with
// no visible text. It must round-trip verbatim so multi-turn thinking + tool
// use conversations keep their reasoning continuity.
const AnthropicRedactedThinkingBlock = Schema.Struct({
type: Schema.tag("redacted_thinking"),
data: Schema.String,
cache_control: Schema.optional(AnthropicCacheControl),
})
const AnthropicToolUseBlock = Schema.Struct({ const AnthropicToolUseBlock = Schema.Struct({
type: Schema.tag("tool_use"), type: Schema.tag("tool_use"),
id: Schema.String, id: Schema.String,
@@ -136,6 +171,7 @@ type AnthropicUserBlock = Schema.Schema.Type<typeof AnthropicUserBlock>
const AnthropicAssistantBlock = Schema.Union([ const AnthropicAssistantBlock = Schema.Union([
AnthropicTextBlock, AnthropicTextBlock,
AnthropicThinkingBlock, AnthropicThinkingBlock,
AnthropicRedactedThinkingBlock,
AnthropicToolUseBlock, AnthropicToolUseBlock,
AnthropicServerToolUseBlock, AnthropicServerToolUseBlock,
AnthropicServerToolResultBlock, AnthropicServerToolResultBlock,
@@ -159,7 +195,7 @@ const AnthropicTool = Schema.Struct({
type AnthropicTool = Schema.Schema.Type<typeof AnthropicTool> type AnthropicTool = Schema.Schema.Type<typeof AnthropicTool>
const AnthropicToolChoice = Schema.Union([ const AnthropicToolChoice = Schema.Union([
Schema.Struct({ type: Schema.Literals(["auto", "any"]) }), Schema.Struct({ type: Schema.Literals(["auto", "any", "none"]) }),
Schema.Struct({ type: Schema.tag("tool"), name: Schema.String }), Schema.Struct({ type: Schema.tag("tool"), name: Schema.String }),
]) ])
@@ -199,12 +235,27 @@ const AnthropicBodyFields = {
export const AnthropicMessagesBody = Schema.Struct(AnthropicBodyFields) export const AnthropicMessagesBody = Schema.Struct(AnthropicBodyFields)
export type AnthropicMessagesBody = Schema.Schema.Type<typeof AnthropicMessagesBody> export type AnthropicMessagesBody = Schema.Schema.Type<typeof AnthropicMessagesBody>
const AnthropicUsage = Schema.Struct({ const AnthropicUsage = Schema.StructWithRest(
input_tokens: Schema.optional(Schema.Number), Schema.Struct({
output_tokens: Schema.optional(Schema.Number), input_tokens: Schema.optional(Schema.Number),
cache_creation_input_tokens: optionalNull(Schema.Number), output_tokens: Schema.optional(Schema.Number),
cache_read_input_tokens: optionalNull(Schema.Number), cache_creation_input_tokens: optionalNull(Schema.Number),
}) cache_read_input_tokens: optionalNull(Schema.Number),
server_tool_use: optionalNull(
Schema.StructWithRest(
Schema.Struct({ web_search_requests: Schema.optional(Schema.Number) }),
[Schema.Record(Schema.String, Schema.Unknown)],
),
),
output_tokens_details: optionalNull(
Schema.StructWithRest(
Schema.Struct({ thinking_tokens: Schema.optional(Schema.Number) }),
[Schema.Record(Schema.String, Schema.Unknown)],
),
),
}),
[Schema.Record(Schema.String, Schema.Unknown)],
)
type AnthropicUsage = Schema.Schema.Type<typeof AnthropicUsage> type AnthropicUsage = Schema.Schema.Type<typeof AnthropicUsage>
const AnthropicStreamBlock = Schema.Struct({ const AnthropicStreamBlock = Schema.Struct({
@@ -214,6 +265,9 @@ const AnthropicStreamBlock = Schema.Struct({
text: Schema.optional(Schema.String), text: Schema.optional(Schema.String),
thinking: Schema.optional(Schema.String), thinking: Schema.optional(Schema.String),
signature: Schema.optional(Schema.String), signature: Schema.optional(Schema.String),
// redacted_thinking blocks arrive whole in content_block_start with the
// encrypted payload in `data`; there is no streaming delta sequence.
data: Schema.optional(Schema.String),
input: Schema.optional(Schema.Unknown), input: Schema.optional(Schema.Unknown),
// *_tool_result blocks arrive whole as content_block_start (no streaming // *_tool_result blocks arrive whole as content_block_start (no streaming
// delta) with the structured payload in `content` and the originating // delta) with the structured payload in `content` and the originating
@@ -251,7 +305,12 @@ type AnthropicEvent = Schema.Schema.Type<typeof AnthropicEvent>
interface ParserState { interface ParserState {
readonly tools: ToolStream.State<number> readonly tools: ToolStream.State<number>
readonly reasoningSignatures: Readonly<Record<number, string>>
readonly usage?: Usage readonly usage?: Usage
readonly pendingFinish?: {
readonly reason: FinishReasonDetails
readonly providerMetadata?: ProviderMetadata
}
readonly lifecycle: Lifecycle.State readonly lifecycle: Lifecycle.State
} }
@@ -287,6 +346,12 @@ const signatureFromMetadata = (metadata: ProviderMetadata | undefined): string |
return typeof anthropic.signature === "string" ? anthropic.signature : undefined return typeof anthropic.signature === "string" ? anthropic.signature : undefined
} }
const redactedDataFromMetadata = (metadata: ProviderMetadata | undefined): string | undefined => {
const anthropic = metadata?.anthropic
if (!ProviderShared.isRecord(anthropic)) return undefined
return typeof anthropic.redactedData === "string" ? anthropic.redactedData : undefined
}
const lowerTool = (breakpoints: Cache.Breakpoints, tool: ToolDefinition, inputSchema: JsonSchema): AnthropicTool => ({ const lowerTool = (breakpoints: Cache.Breakpoints, tool: ToolDefinition, inputSchema: JsonSchema): AnthropicTool => ({
name: tool.name, name: tool.name,
description: tool.description, description: tool.description,
@@ -297,7 +362,7 @@ const lowerTool = (breakpoints: Cache.Breakpoints, tool: ToolDefinition, inputSc
const lowerToolChoice = (toolChoice: NonNullable<LLMRequest["toolChoice"]>) => const lowerToolChoice = (toolChoice: NonNullable<LLMRequest["toolChoice"]>) =>
ProviderShared.matchToolChoice("Anthropic Messages", toolChoice, { ProviderShared.matchToolChoice("Anthropic Messages", toolChoice, {
auto: () => ({ type: "auto" as const }), auto: () => ({ type: "auto" as const }),
none: () => undefined, none: () => ({ type: "none" as const }),
required: () => ({ type: "any" as const }), required: () => ({ type: "any" as const }),
tool: (name) => ({ type: "tool" as const, name }), tool: (name) => ({ type: "tool" as const, name }),
}) })
@@ -330,7 +395,10 @@ const lowerServerToolResult = Effect.fn("AnthropicMessages.lowerServerToolResult
const wireType = serverToolResultType(part.name) const wireType = serverToolResultType(part.name)
if (!wireType) if (!wireType)
return yield* invalid(`Anthropic Messages does not know how to round-trip server tool result for ${part.name}`) return yield* invalid(`Anthropic Messages does not know how to round-trip server tool result for ${part.name}`)
return { type: wireType, tool_use_id: part.id, content: part.result.value } satisfies AnthropicServerToolResultBlock // Prefer the provider-owned replay payload; fall back to the result value for
// histories constructed directly from provider events.
const payload = part.providerMetadata?.anthropic?.["result"] ?? part.result.value
return { type: wireType, tool_use_id: part.id, content: payload } satisfies AnthropicServerToolResultBlock
}) })
const lowerMedia = Effect.fn("AnthropicMessages.lowerMedia")(function* (part: MediaPart) { const lowerMedia = Effect.fn("AnthropicMessages.lowerMedia")(function* (part: MediaPart) {
@@ -357,7 +425,7 @@ const lowerMedia = Effect.fn("AnthropicMessages.lowerMedia")(function* (part: Me
// Tool results may carry structured text, images, and documents. Keep media as provider-native // Tool results may carry structured text, images, and documents. Keep media as provider-native
// content instead of JSON-stringifying base64 into a prompt string. // content instead of JSON-stringifying base64 into a prompt string.
const lowerToolResultContentItem = Effect.fn("AnthropicMessages.lowerToolResultContentItem")(function* ( const lowerToolResultContentItem = Effect.fn("AnthropicMessages.lowerToolResultContentItem")(function* (
item: ToolContent, item: Tool.Content,
) { ) {
if (item.type === "text") return { type: "text" as const, text: item.text } satisfies AnthropicTextBlock if (item.type === "text") return { type: "text" as const, text: item.text } satisfies AnthropicTextBlock
return yield* lowerMedia({ type: "media", mediaType: item.mime, data: item.uri, filename: item.name }) return yield* lowerMedia({ type: "media", mediaType: item.mime, data: item.uri, filename: item.name })
@@ -368,7 +436,7 @@ const lowerToolResultContent = Effect.fn("AnthropicMessages.lowerToolResultConte
// with existing cassettes and provider expectations. // with existing cassettes and provider expectations.
if (part.result.type !== "content") return ProviderShared.toolResultText(part) if (part.result.type !== "content") return ProviderShared.toolResultText(part)
// Preserve the narrowed array element type when compiled through a consumer package. // Preserve the narrowed array element type when compiled through a consumer package.
const content: ReadonlyArray<ToolContent> = part.result.value const content: ReadonlyArray<Tool.Content> = part.result.value
return yield* Effect.forEach(content, lowerToolResultContentItem) return yield* Effect.forEach(content, lowerToolResultContentItem)
}) })
@@ -469,11 +537,16 @@ const lowerMessages = Effect.fn("AnthropicMessages.lowerMessages")(function* (
continue continue
} }
if (part.type === "reasoning") { if (part.type === "reasoning") {
content.push({ // Mirrors Vercel's @ai-sdk/anthropic: a signature marks visible
type: "thinking", // thinking; only signature-less parts carrying redactedData
thinking: part.text, // round-trip as opaque redacted_thinking blocks.
signature: part.encrypted ?? signatureFromMetadata(part.providerMetadata), const signature = part.encrypted ?? signatureFromMetadata(part.providerMetadata)
}) const redactedData = redactedDataFromMetadata(part.providerMetadata)
if (signature === undefined && redactedData !== undefined) {
content.push({ type: "redacted_thinking", data: redactedData })
continue
}
content.push({ type: "thinking", thinking: part.text, signature })
continue continue
} }
if (part.type === "tool-call") { if (part.type === "tool-call") {
@@ -510,39 +583,39 @@ const lowerMessages = Effect.fn("AnthropicMessages.lowerMessages")(function* (
return messages return messages
}) })
const anthropicOptions = (request: LLMRequest) => request.providerOptions?.anthropic const resolveOptions = Effect.fn("AnthropicMessages.resolveOptions")(function* (request: LLMRequest) {
const input = request.providerOptions?.anthropic
return {
thinking: yield* resolveThinking(input?.thinking),
effort: typeof input?.effort === "string" ? input.effort : undefined,
}
})
const lowerThinking = Effect.fn("AnthropicMessages.lowerThinking")(function* (request: LLMRequest) { const resolveThinking = Effect.fn("AnthropicMessages.resolveThinking")(function* (input: unknown) {
const thinking = anthropicOptions(request)?.thinking if (!ProviderShared.isRecord(input)) return undefined
if (!ProviderShared.isRecord(thinking)) return undefined if (input.type === "adaptive") {
if (thinking.type === "adaptive") {
const display = const display =
thinking.display === "summarized" input.display === "summarized"
? ("summarized" as const) ? ("summarized" as const)
: thinking.display === "omitted" : input.display === "omitted"
? ("omitted" as const) ? ("omitted" as const)
: undefined : undefined
return { type: "adaptive" as const, ...(display === undefined ? {} : { display }) } return { type: "adaptive" as const, ...(display === undefined ? {} : { display }) }
} }
if (thinking.type === "disabled") return { type: "disabled" as const } if (input.type === "disabled") return { type: "disabled" as const }
if (thinking.type !== "enabled") return undefined if (input.type !== "enabled") return undefined
const budget = const budget =
typeof thinking.budgetTokens === "number" typeof input.budgetTokens === "number"
? thinking.budgetTokens ? input.budgetTokens
: typeof thinking.budget_tokens === "number" : typeof input.budget_tokens === "number"
? thinking.budget_tokens ? input.budget_tokens
: undefined : undefined
if (budget === undefined) return yield* invalid("Anthropic thinking provider option requires budgetTokens") if (budget === undefined)
return yield* ProviderShared.invalidRequest("Anthropic thinking provider option requires budgetTokens")
return { type: "enabled" as const, budget_tokens: budget } return { type: "enabled" as const, budget_tokens: budget }
}) })
const outputConfig = (request: LLMRequest) => {
const effort = anthropicOptions(request)?.effort
return typeof effort === "string" ? { effort } : undefined
}
const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (request: LLMRequest) { const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (request: LLMRequest) {
const toolChoice = request.toolChoice ? yield* lowerToolChoice(request.toolChoice) : undefined
const generation = request.generation const generation = request.generation
const toolSchemaCompatibility = request.model.compatibility?.toolSchema const toolSchemaCompatibility = request.model.compatibility?.toolSchema
const outputLimit = request.model.defaults?.limits?.output ?? request.model.route.defaults.limits?.output ?? 4096 const outputLimit = request.model.defaults?.limits?.output ?? request.model.route.defaults.limits?.output ?? 4096
@@ -551,7 +624,7 @@ const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (reques
// over-mark we keep their tool hints and shed the message-tail ones first. // over-mark we keep their tool hints and shed the message-tail ones first.
const breakpoints = Cache.newBreakpoints(ANTHROPIC_BREAKPOINT_CAP) const breakpoints = Cache.newBreakpoints(ANTHROPIC_BREAKPOINT_CAP)
const tools = const tools =
request.tools.length === 0 || request.toolChoice?.type === "none" request.tools.length === 0
? undefined ? undefined
: request.tools.map((tool) => : request.tools.map((tool) =>
lowerTool( lowerTool(
@@ -560,6 +633,8 @@ const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (reques
ToolSchemaProjection.modelCompatibility(tool.inputSchema, toolSchemaCompatibility), ToolSchemaProjection.modelCompatibility(tool.inputSchema, toolSchemaCompatibility),
), ),
) )
// Anthropic rejects tool_choice when tools are absent; "none" is only meaningful with tools present.
const toolChoice = tools === undefined || !request.toolChoice ? undefined : yield* lowerToolChoice(request.toolChoice)
const system = const system =
request.system.length === 0 request.system.length === 0
? undefined ? undefined
@@ -574,6 +649,7 @@ const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (reques
`Anthropic Messages: dropped ${breakpoints.dropped} cache breakpoint(s); the API allows at most ${ANTHROPIC_BREAKPOINT_CAP} per request.`, `Anthropic Messages: dropped ${breakpoints.dropped} cache breakpoint(s); the API allows at most ${ANTHROPIC_BREAKPOINT_CAP} per request.`,
) )
} }
const options = yield* resolveOptions(request)
return { return {
model: request.model.id, model: request.model.id,
system, system,
@@ -586,8 +662,8 @@ const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (reques
top_p: generation?.topP, top_p: generation?.topP,
top_k: generation?.topK, top_k: generation?.topK,
stop_sequences: generation?.stop, stop_sequences: generation?.stop,
thinking: yield* lowerThinking(request), thinking: options.thinking,
output_config: outputConfig(request), output_config: options.effort === undefined ? undefined : { effort: options.effort },
} }
}) })
@@ -596,7 +672,7 @@ const fromRequest = Effect.fn("AnthropicMessages.fromRequest")(function* (reques
// ============================================================================= // =============================================================================
const mapFinishReason = (reason: string | null | undefined): FinishReason => { const mapFinishReason = (reason: string | null | undefined): FinishReason => {
if (reason === "end_turn" || reason === "stop_sequence" || reason === "pause_turn") return "stop" if (reason === "end_turn" || reason === "stop_sequence" || reason === "pause_turn") return "stop"
if (reason === "max_tokens") return "length" if (reason === "max_tokens" || reason === "model_context_window_exceeded") return "length"
if (reason === "tool_use") return "tool-calls" if (reason === "tool_use") return "tool-calls"
if (reason === "refusal") return "content-filter" if (reason === "refusal") return "content-filter"
return "unknown" return "unknown"
@@ -606,9 +682,8 @@ const mapFinishReason = (reason: string | null | undefined): FinishReason => {
// `input_tokens` is the *non-cached* count per the Messages API docs, with // `input_tokens` is the *non-cached* count per the Messages API docs, with
// cache reads and writes as separate fields. We sum them to derive the // cache reads and writes as separate fields. We sum them to derive the
// inclusive `inputTokens` the rest of the contract expects. Extended // inclusive `inputTokens` the rest of the contract expects. Extended
// thinking tokens are *not* broken out by Anthropic — they're billed as // thinking tokens are included in `output_tokens`; newer responses also
// part of `output_tokens`, so `reasoningTokens` stays `undefined` and // expose that subset through `output_tokens_details.thinking_tokens`.
// `outputTokens` carries the combined total.
const mapUsage = (usage: AnthropicUsage | undefined): Usage | undefined => { const mapUsage = (usage: AnthropicUsage | undefined): Usage | undefined => {
if (!usage) return undefined if (!usage) return undefined
const nonCached = usage.input_tokens const nonCached = usage.input_tokens
@@ -621,6 +696,7 @@ const mapUsage = (usage: AnthropicUsage | undefined): Usage | undefined => {
nonCachedInputTokens: nonCached, nonCachedInputTokens: nonCached,
cacheReadInputTokens: cacheRead, cacheReadInputTokens: cacheRead,
cacheWriteInputTokens: cacheWrite, cacheWriteInputTokens: cacheWrite,
reasoningTokens: usage.output_tokens_details?.thinking_tokens,
totalTokens: ProviderShared.totalTokens(inputTokens, usage.output_tokens, undefined), totalTokens: ProviderShared.totalTokens(inputTokens, usage.output_tokens, undefined),
providerMetadata: { anthropic: usage }, providerMetadata: { anthropic: usage },
}) })
@@ -639,18 +715,18 @@ const mergeUsage = (left: Usage | undefined, right: Usage | undefined) => {
const cacheWriteInputTokens = right.cacheWriteInputTokens ?? left.cacheWriteInputTokens const cacheWriteInputTokens = right.cacheWriteInputTokens ?? left.cacheWriteInputTokens
const inputTokens = ProviderShared.sumTokens(nonCachedInputTokens, cacheReadInputTokens, cacheWriteInputTokens) const inputTokens = ProviderShared.sumTokens(nonCachedInputTokens, cacheReadInputTokens, cacheWriteInputTokens)
const outputTokens = right.outputTokens ?? left.outputTokens const outputTokens = right.outputTokens ?? left.outputTokens
const reasoningTokens = right.reasoningTokens ?? left.reasoningTokens
return new Usage({ return new Usage({
inputTokens, inputTokens,
outputTokens, outputTokens,
nonCachedInputTokens, nonCachedInputTokens,
cacheReadInputTokens, cacheReadInputTokens,
cacheWriteInputTokens, cacheWriteInputTokens,
reasoningTokens,
totalTokens: ProviderShared.totalTokens(inputTokens, outputTokens, undefined), totalTokens: ProviderShared.totalTokens(inputTokens, outputTokens, undefined),
providerMetadata: { providerMetadata: {
anthropic: { anthropic:
...left.providerMetadata?.["anthropic"], mergeJsonRecords(left.providerMetadata?.["anthropic"], right.providerMetadata?.["anthropic"]) ?? {},
...right.providerMetadata?.["anthropic"],
},
}, },
}) })
} }
@@ -680,7 +756,9 @@ const serverToolResultEvent = (block: NonNullable<AnthropicEvent["content_block"
name: SERVER_TOOL_RESULT_NAMES[block.type], name: SERVER_TOOL_RESULT_NAMES[block.type],
result: isError ? { type: "error", value: block.content } : { type: "json", value: block.content }, result: isError ? { type: "error", value: block.content } : { type: "json", value: block.content },
providerExecuted: true, providerExecuted: true,
providerMetadata: anthropicMetadata({ blockType: block.type }), // The complete payload is irreducible provider replay state: subsequent
// stateless requests must round-trip the typed result block verbatim.
providerMetadata: anthropicMetadata({ blockType: block.type, result: block.content }),
}) })
} }
@@ -707,6 +785,10 @@ const onContentBlockStart = (state: ParserState, event: AnthropicEvent): StepRes
tools: ToolStream.start(state.tools, event.index, { tools: ToolStream.start(state.tools, event.index, {
id: block.id ?? String(event.index), id: block.id ?? String(event.index),
name: block.name ?? "", name: block.name ?? "",
input:
block.input !== undefined && (!ProviderShared.isRecord(block.input) || Object.keys(block.input).length > 0)
? ProviderShared.encodeJson(block.input)
: undefined,
providerExecuted: block.type === "server_tool_use", providerExecuted: block.type === "server_tool_use",
}), }),
}, },
@@ -721,20 +803,50 @@ const onContentBlockStart = (state: ParserState, event: AnthropicEvent): StepRes
] ]
} }
if (block.type === "text" && block.text) { if (block.type === "text" && block.text !== undefined) {
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const id = `text-${event.index ?? 0}`
const lifecycle = Lifecycle.textStart(state.lifecycle, events, id)
return [ return [
{ ...state, lifecycle: Lifecycle.textDelta(state.lifecycle, events, `text-${event.index ?? 0}`, block.text) }, { ...state, lifecycle: block.text ? Lifecycle.textDelta(lifecycle, events, id, block.text) : lifecycle },
events, events,
] ]
} }
if (block.type === "thinking" && block.thinking) { if (block.type === "thinking" && block.thinking !== undefined) {
const events: LLMEvent[] = []
const id = `reasoning-${event.index ?? 0}`
const providerMetadata = block.signature === undefined ? undefined : anthropicMetadata({ signature: block.signature })
const lifecycle = Lifecycle.reasoningStart(state.lifecycle, events, id, providerMetadata)
return [
{
...state,
lifecycle: block.thinking
? Lifecycle.reasoningDelta(lifecycle, events, id, block.thinking, providerMetadata)
: lifecycle,
reasoningSignatures:
event.index === undefined || block.signature === undefined
? state.reasoningSignatures
: { ...state.reasoningSignatures, [event.index]: block.signature },
},
events,
]
}
// Redacted thinking surfaces as an empty reasoning part carrying the opaque
// payload as `redactedData` metadata (same model as Vercel's
// @ai-sdk/anthropic). The existing content_block_stop closes the part.
if (block.type === "redacted_thinking" && block.data !== undefined) {
const events: LLMEvent[] = [] const events: LLMEvent[] = []
return [ return [
{ {
...state, ...state,
lifecycle: Lifecycle.reasoningDelta(state.lifecycle, events, `reasoning-${event.index ?? 0}`, block.thinking), lifecycle: Lifecycle.reasoningStart(
state.lifecycle,
events,
`reasoning-${event.index ?? 0}`,
anthropicMetadata({ redactedData: block.data }),
),
}, },
events, events,
] ]
@@ -772,18 +884,13 @@ const onContentBlockDelta = Effect.fn("AnthropicMessages.onContentBlockDelta")(f
} }
if (delta?.type === "signature_delta" && delta.signature) { if (delta?.type === "signature_delta" && delta.signature) {
const events: LLMEvent[] = [] const index = event.index ?? 0
return [ return [
{ {
...state, ...state,
lifecycle: Lifecycle.reasoningEnd( reasoningSignatures: { ...state.reasoningSignatures, [index]: delta.signature },
state.lifecycle,
events,
`reasoning-${event.index ?? 0}`,
anthropicMetadata({ signature: delta.signature }),
),
}, },
events, NO_EVENTS,
] satisfies StepResult ] satisfies StepResult
} }
@@ -814,28 +921,53 @@ const onContentBlockStop = Effect.fn("AnthropicMessages.onContentBlockStop")(fun
const result = yield* ToolStream.finish(ADAPTER, state.tools, event.index) const result = yield* ToolStream.finish(ADAPTER, state.tools, event.index)
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const resultEvents = result.events ?? [] const resultEvents = result.events ?? []
const signature = state.reasoningSignatures[event.index]
const lifecycle = resultEvents.length const lifecycle = resultEvents.length
? Lifecycle.stepStart(state.lifecycle, events) ? Lifecycle.stepStart(state.lifecycle, events)
: Lifecycle.reasoningEnd( : Lifecycle.reasoningEnd(
Lifecycle.textEnd(state.lifecycle, events, `text-${event.index}`), Lifecycle.textEnd(state.lifecycle, events, `text-${event.index}`),
events, events,
`reasoning-${event.index}`, `reasoning-${event.index}`,
signature === undefined ? undefined : anthropicMetadata({ signature }),
) )
events.push(...resultEvents) events.push(...resultEvents)
return [{ ...state, lifecycle, tools: result.tools }, events] satisfies StepResult const reasoningSignatures = { ...state.reasoningSignatures }
delete reasoningSignatures[event.index]
return [{ ...state, lifecycle, tools: result.tools, reasoningSignatures }, events] satisfies StepResult
}) })
const onMessageDelta = (state: ParserState, event: AnthropicEvent): StepResult => { const onMessageDelta = (state: ParserState, event: AnthropicEvent): StepResult => {
const usage = mergeUsage(state.usage, mapUsage(event.usage)) const usage = mergeUsage(state.usage, mapUsage(event.usage))
return [
{
...state,
usage,
pendingFinish: {
reason: {
normalized: mapFinishReason(event.delta?.stop_reason),
raw: event.delta?.stop_reason ?? undefined,
},
providerMetadata:
event.delta?.stop_sequence === null || event.delta?.stop_sequence === undefined
? undefined
: anthropicMetadata({ stopSequence: event.delta.stop_sequence }),
},
},
NO_EVENTS,
]
}
const onMessageStop = (state: ParserState): StepResult => {
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const lifecycle = Lifecycle.finish(state.lifecycle, events, { const lifecycle = Lifecycle.finish(state.lifecycle, events, {
reason: mapFinishReason(event.delta?.stop_reason), reason: state.pendingFinish?.reason ?? {
usage, normalized: "unknown",
providerMetadata: event.delta?.stop_sequence raw: undefined,
? anthropicMetadata({ stopSequence: event.delta.stop_sequence }) },
: undefined, usage: state.usage,
providerMetadata: state.pendingFinish?.providerMetadata,
}) })
return [{ ...state, lifecycle, usage }, events] return [{ ...state, lifecycle }, events]
} }
// Prefix `error.type` so overloads, rate limits, and quota errors are visible // Prefix `error.type` so overloads, rate limits, and quota errors are visible
@@ -860,6 +992,7 @@ const step = (state: ParserState, event: AnthropicEvent) => {
if (event.type === "content_block_delta") return onContentBlockDelta(state, event) if (event.type === "content_block_delta") return onContentBlockDelta(state, event)
if (event.type === "content_block_stop") return onContentBlockStop(state, event) if (event.type === "content_block_stop") return onContentBlockStop(state, event)
if (event.type === "message_delta") return Effect.succeed(onMessageDelta(state, event)) if (event.type === "message_delta") return Effect.succeed(onMessageDelta(state, event))
if (event.type === "message_stop") return Effect.succeed(onMessageStop(state))
if (event.type === "error") return onError(event) if (event.type === "error") return onError(event)
return Effect.succeed<StepResult>([state, NO_EVENTS]) return Effect.succeed<StepResult>([state, NO_EVENTS])
} }
@@ -880,7 +1013,11 @@ export const protocol = Protocol.make({
}, },
stream: { stream: {
event: Protocol.jsonEvent(AnthropicEvent), event: Protocol.jsonEvent(AnthropicEvent),
initial: () => ({ tools: ToolStream.empty<number>(), lifecycle: Lifecycle.initial() }), initial: () => ({
tools: ToolStream.empty<number>(),
reasoningSignatures: {},
lifecycle: Lifecycle.initial(),
}),
step, step,
}, },
}) })
+90 -28
View File
@@ -8,6 +8,7 @@ import {
Usage, Usage,
type CacheHint, type CacheHint,
type FinishReason, type FinishReason,
type FinishReasonDetails,
type JsonSchema, type JsonSchema,
type LLMRequest, type LLMRequest,
type ModelToolSchemaCompatibility, type ModelToolSchemaCompatibility,
@@ -65,14 +66,15 @@ const BedrockToolResultBlock = Schema.Struct({
type BedrockToolResultBlock = Schema.Schema.Type<typeof BedrockToolResultBlock> type BedrockToolResultBlock = Schema.Schema.Type<typeof BedrockToolResultBlock>
const BedrockReasoningBlock = Schema.Struct({ const BedrockReasoningBlock = Schema.Struct({
reasoningContent: Schema.Struct({ reasoningContent: Schema.Union([
reasoningText: Schema.optional( Schema.Struct({
Schema.Struct({ reasoningText: Schema.Struct({
text: Schema.String, text: Schema.String,
signature: Schema.optional(Schema.String), signature: Schema.optional(Schema.String),
}), }),
), }),
}), Schema.Struct({ redactedContent: Schema.String }),
]),
}) })
const BedrockUserBlock = Schema.Union([ const BedrockUserBlock = Schema.Union([
@@ -153,6 +155,12 @@ const BedrockUsageSchema = Schema.Struct({
}) })
type BedrockUsageSchema = Schema.Schema.Type<typeof BedrockUsageSchema> type BedrockUsageSchema = Schema.Schema.Type<typeof BedrockUsageSchema>
const BedrockStreamException = Schema.Struct({
message: Schema.optional(Schema.String),
originalMessage: Schema.optional(Schema.String),
originalStatusCode: Schema.optional(Schema.Number),
})
// Streaming event shape — the AWS event stream wraps each JSON payload by its // Streaming event shape — the AWS event stream wraps each JSON payload by its
// `:event-type` header (e.g. `messageStart`, `contentBlockDelta`). We // `:event-type` header (e.g. `messageStart`, `contentBlockDelta`). We
// reconstruct that wrapping in `decodeFrames` below so the event schema can // reconstruct that wrapping in `decodeFrames` below so the event schema can
@@ -180,6 +188,11 @@ const BedrockEvent = Schema.Struct({
Schema.Struct({ Schema.Struct({
text: Schema.optional(Schema.String), text: Schema.optional(Schema.String),
signature: Schema.optional(Schema.String), signature: Schema.optional(Schema.String),
// Blob fields in Bedrock's JSON event stream are base64 strings.
redactedContent: Schema.optional(Schema.String),
// Vercel's Bedrock provider exposes the same delta under
// Anthropic's shorter `data` spelling.
data: Schema.optional(Schema.String),
}), }),
), ),
}), }),
@@ -199,11 +212,11 @@ const BedrockEvent = Schema.Struct({
metrics: Schema.optional(Schema.Unknown), metrics: Schema.optional(Schema.Unknown),
}), }),
), ),
internalServerException: Schema.optional(Schema.Struct({ message: Schema.String })), internalServerException: Schema.optional(BedrockStreamException),
modelStreamErrorException: Schema.optional(Schema.Struct({ message: Schema.String })), modelStreamErrorException: Schema.optional(BedrockStreamException),
validationException: Schema.optional(Schema.Struct({ message: Schema.String })), validationException: Schema.optional(BedrockStreamException),
throttlingException: Schema.optional(Schema.Struct({ message: Schema.String })), throttlingException: Schema.optional(BedrockStreamException),
serviceUnavailableException: Schema.optional(Schema.Struct({ message: Schema.String })), serviceUnavailableException: Schema.optional(BedrockStreamException),
}) })
type BedrockEvent = Schema.Schema.Type<typeof BedrockEvent> type BedrockEvent = Schema.Schema.Type<typeof BedrockEvent>
@@ -259,6 +272,13 @@ const reasoningSignature = (part: ReasoningPart) => {
) )
} }
const reasoningRedactedData = (part: ReasoningPart) => {
const bedrock = part.providerMetadata?.bedrock
return ProviderShared.isRecord(bedrock) && typeof bedrock.redactedData === "string"
? bedrock.redactedData
: undefined
}
const lowerToolCall = (part: ToolCallPart): BedrockToolUseBlock => ({ const lowerToolCall = (part: ToolCallPart): BedrockToolUseBlock => ({
toolUse: { toolUse: {
toolUseId: part.id, toolUseId: part.id,
@@ -348,11 +368,13 @@ const lowerMessages = Effect.fn("BedrockConverse.lowerMessages")(function* (
continue continue
} }
if (part.type === "reasoning") { if (part.type === "reasoning") {
content.push({ const signature = reasoningSignature(part)
reasoningContent: { const redactedData = reasoningRedactedData(part)
reasoningText: { text: part.text, signature: reasoningSignature(part) }, if (signature === undefined && redactedData !== undefined) {
}, content.push({ reasoningContent: { redactedContent: redactedData } })
}) continue
}
content.push({ reasoningContent: { reasoningText: { text: part.text, signature } } })
continue continue
} }
if (part.type === "tool-call") { if (part.type === "tool-call") {
@@ -392,8 +414,13 @@ const fromRequest = Effect.fn("BedrockConverse.fromRequest")(function* (request:
// tools → system → messages order to favour the highest-impact prefixes. // tools → system → messages order to favour the highest-impact prefixes.
const breakpoints = BedrockCache.breakpoints() const breakpoints = BedrockCache.breakpoints()
const toolConfig = const toolConfig =
request.tools.length > 0 && request.toolChoice?.type !== "none" request.tools.length > 0
? { tools: lowerTools(request.model.compatibility?.toolSchema, breakpoints, request.tools), toolChoice } ? {
tools: lowerTools(request.model.compatibility?.toolSchema, breakpoints, request.tools),
// Converse has no native "none". Keep definitions stable for prompt
// caching and omit only the unsupported choice.
toolChoice,
}
: undefined : undefined
const system = request.system.length === 0 ? undefined : lowerSystem(breakpoints, request.system) const system = request.system.length === 0 ? undefined : lowerSystem(breakpoints, request.system)
const messages = yield* lowerMessages(request, breakpoints) const messages = yield* lowerMessages(request, breakpoints)
@@ -430,9 +457,10 @@ const fromRequest = Effect.fn("BedrockConverse.fromRequest")(function* (request:
// ============================================================================= // =============================================================================
const mapFinishReason = (reason: string): FinishReason => { const mapFinishReason = (reason: string): FinishReason => {
if (reason === "end_turn" || reason === "stop_sequence") return "stop" if (reason === "end_turn" || reason === "stop_sequence") return "stop"
if (reason === "max_tokens") return "length" if (reason === "max_tokens" || reason === "model_context_window_exceeded") return "length"
if (reason === "tool_use") return "tool-calls" if (reason === "tool_use") return "tool-calls"
if (reason === "content_filtered" || reason === "guardrail_intervened") return "content-filter" if (reason === "content_filtered" || reason === "guardrail_intervened") return "content-filter"
if (reason === "malformed_model_output" || reason === "malformed_tool_use") return "error"
return "unknown" return "unknown"
} }
@@ -461,7 +489,7 @@ interface ParserState {
// Bedrock splits the finish into `messageStop` (carries `stopReason`) and // Bedrock splits the finish into `messageStop` (carries `stopReason`) and
// `metadata` (carries usage). Hold the terminal event in state so `onHalt` // `metadata` (carries usage). Hold the terminal event in state so `onHalt`
// can emit exactly one finish after both chunks have had a chance to arrive. // can emit exactly one finish after both chunks have had a chance to arrive.
readonly pendingFinish: { readonly reason: FinishReason; readonly usage?: Usage } | undefined readonly pendingFinish: { readonly reason: FinishReasonDetails; readonly usage?: Usage } | undefined
readonly hasToolCalls: boolean readonly hasToolCalls: boolean
readonly lifecycle: Lifecycle.State readonly lifecycle: Lifecycle.State
readonly reasoningSignatures: Readonly<Record<number, string>> readonly reasoningSignatures: Readonly<Record<number, string>>
@@ -512,12 +540,26 @@ const step = (state: ParserState, event: BedrockEvent) =>
const index = event.contentBlockDelta.contentBlockIndex const index = event.contentBlockDelta.contentBlockIndex
const reasoning = event.contentBlockDelta.delta.reasoningContent const reasoning = event.contentBlockDelta.delta.reasoningContent
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const redactedData = reasoning.redactedContent ?? reasoning.data
const providerMetadata = reasoning.signature
? bedrockMetadata({ signature: reasoning.signature })
: redactedData !== undefined
? bedrockMetadata({ redactedData })
: undefined
const lifecycle =
reasoning.text !== undefined || providerMetadata !== undefined
? Lifecycle.reasoningDelta(
state.lifecycle,
events,
`reasoning-${index}`,
reasoning.text ?? "",
providerMetadata,
)
: state.lifecycle
return [ return [
{ {
...state, ...state,
lifecycle: reasoning.text lifecycle,
? Lifecycle.reasoningDelta(state.lifecycle, events, `reasoning-${index}`, reasoning.text)
: state.lifecycle,
reasoningSignatures: reasoning.signature reasoningSignatures: reasoning.signature
? { ...state.reasoningSignatures, [index]: reasoning.signature } ? { ...state.reasoningSignatures, [index]: reasoning.signature }
: state.reasoningSignatures, : state.reasoningSignatures,
@@ -578,15 +620,30 @@ const step = (state: ParserState, event: BedrockEvent) =>
return [ return [
{ {
...state, ...state,
pendingFinish: { reason: mapFinishReason(event.messageStop.stopReason), usage: state.pendingFinish?.usage }, pendingFinish: {
reason: {
normalized: mapFinishReason(event.messageStop.stopReason),
raw: event.messageStop.stopReason,
},
usage: state.pendingFinish?.usage,
},
}, },
[], [],
] as const ] as const
} }
if (event.metadata) { if (event.metadata) {
const usage = mapUsage(event.metadata.usage) const usage = mapUsage(event.metadata.usage) ?? state.pendingFinish?.usage
return [{ ...state, pendingFinish: { reason: state.pendingFinish?.reason ?? "stop", usage } }, []] as const return [
{
...state,
pendingFinish: {
reason: state.pendingFinish?.reason ?? { normalized: "stop" },
usage,
},
},
[],
] as const
} }
const exception = ( const exception = (
@@ -603,7 +660,7 @@ const step = (state: ParserState, event: BedrockEvent) =>
module: ADAPTER, module: ADAPTER,
method: "stream", method: "stream",
reason: classifyProviderFailure({ reason: classifyProviderFailure({
message: exception[1]?.message ?? "Bedrock Converse stream error", message: exception[1]?.message ?? exception[1]?.originalMessage ?? "Bedrock Converse stream error",
code: exception[0], code: exception[0],
}), }),
}) })
@@ -619,8 +676,13 @@ const onHalt = (state: ParserState): ReadonlyArray<LLMEvent> =>
? (() => { ? (() => {
const events: LLMEvent[] = [] const events: LLMEvent[] = []
Lifecycle.finish(state.lifecycle, events, { Lifecycle.finish(state.lifecycle, events, {
reason: reason: {
state.pendingFinish.reason === "stop" && state.hasToolCalls ? "tool-calls" : state.pendingFinish.reason, ...state.pendingFinish.reason,
normalized:
state.pendingFinish.reason.normalized === "stop" && state.hasToolCalls
? "tool-calls"
: state.pendingFinish.reason.normalized,
},
usage: state.pendingFinish.usage, usage: state.pendingFinish.usage,
}) })
return events return events
@@ -53,8 +53,22 @@ const consumeFrames = (route: string) => (state: FrameBufferState, chunk: Uint8A
}) })
cursor = { buffer: cursor.buffer, offset: cursor.offset + totalLength } cursor = { buffer: cursor.buffer, offset: cursor.offset + totalLength }
if (decoded.headers[":message-type"]?.value !== "event") continue const messageType = decoded.headers[":message-type"]?.value
const eventType = decoded.headers[":event-type"]?.value if (messageType === "error") {
const code = decoded.headers[":error-code"]?.value
const message = decoded.headers[":error-message"]?.value
return yield* ProviderShared.eventError(
route,
[code, message].filter((value): value is string => typeof value === "string").join(": ") ||
"Bedrock Converse event-stream error",
)
}
const eventType =
messageType === "event"
? decoded.headers[":event-type"]?.value
: messageType === "exception"
? decoded.headers[":exception-type"]?.value
: undefined
if (typeof eventType !== "string") continue if (typeof eventType !== "string") continue
const payload = utf8.decode(decoded.body) const payload = utf8.decode(decoded.body)
if (!payload) continue if (!payload) continue
+104 -19
View File
@@ -1,4 +1,5 @@
import { Effect, Schema } from "effect" import { Effect, Schema } from "effect"
import { Tool } from "@opencode-ai/schema/tool"
import { Route } from "../route/client" import { Route } from "../route/client"
import { Auth } from "../route/auth" import { Auth } from "../route/auth"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
@@ -11,11 +12,11 @@ import {
type JsonSchema, type JsonSchema,
type LLMRequest, type LLMRequest,
type MediaPart, type MediaPart,
type ProviderOptions,
type ProviderMetadata, type ProviderMetadata,
type TextPart, type TextPart,
type ToolCallPart, type ToolCallPart,
type ToolDefinition, type ToolDefinition,
type ToolContent,
} from "../schema" } from "../schema"
import { JsonObject, optionalArray, ProviderShared } from "./shared" import { JsonObject, optionalArray, ProviderShared } from "./shared"
import { GeminiToolSchema } from "./utils/gemini-tool-schema" import { GeminiToolSchema } from "./utils/gemini-tool-schema"
@@ -26,6 +27,39 @@ const ADAPTER = "gemini"
const MEDIA_MIMES = new Set<string>(ProviderShared.MEDIA_MIMES) const MEDIA_MIMES = new Set<string>(ProviderShared.MEDIA_MIMES)
export const DEFAULT_BASE_URL = "https://generativelanguage.googleapis.com/v1beta" export const DEFAULT_BASE_URL = "https://generativelanguage.googleapis.com/v1beta"
export interface OptionsInput {
readonly [key: string]: unknown
readonly cachedContent?: string
readonly safetySettings?: ReadonlyArray<{
readonly category:
| "HARM_CATEGORY_UNSPECIFIED"
| "HARM_CATEGORY_HATE_SPEECH"
| "HARM_CATEGORY_DANGEROUS_CONTENT"
| "HARM_CATEGORY_HARASSMENT"
| "HARM_CATEGORY_SEXUALLY_EXPLICIT"
| "HARM_CATEGORY_CIVIC_INTEGRITY"
| (string & {})
readonly threshold:
| "HARM_BLOCK_THRESHOLD_UNSPECIFIED"
| "BLOCK_LOW_AND_ABOVE"
| "BLOCK_MEDIUM_AND_ABOVE"
| "BLOCK_ONLY_HIGH"
| "BLOCK_NONE"
| "OFF"
| (string & {})
}>
readonly serviceTier?: "standard" | "flex" | "priority" | (string & {})
readonly thinkingConfig?: {
readonly thinkingBudget?: number
readonly includeThoughts?: boolean
readonly thinkingLevel?: "minimal" | "low" | "medium" | "high" | (string & {})
}
}
export type ProviderOptionsInput = ProviderOptions & {
readonly gemini?: OptionsInput
}
// ============================================================================= // =============================================================================
// Request Body Schema // Request Body Schema
// ============================================================================= // =============================================================================
@@ -98,6 +132,12 @@ const GeminiToolConfig = Schema.Struct({
const GeminiThinkingConfig = Schema.Struct({ const GeminiThinkingConfig = Schema.Struct({
thinkingBudget: Schema.optional(Schema.Number), thinkingBudget: Schema.optional(Schema.Number),
includeThoughts: Schema.optional(Schema.Boolean), includeThoughts: Schema.optional(Schema.Boolean),
thinkingLevel: Schema.optional(Schema.String),
})
const GeminiSafetySetting = Schema.Struct({
category: Schema.String,
threshold: Schema.String,
}) })
const GeminiGenerationConfig = Schema.Struct({ const GeminiGenerationConfig = Schema.Struct({
@@ -110,7 +150,10 @@ const GeminiGenerationConfig = Schema.Struct({
}) })
const GeminiBodyFields = { const GeminiBodyFields = {
cachedContent: Schema.optional(Schema.String),
contents: Schema.Array(GeminiContent), contents: Schema.Array(GeminiContent),
safetySettings: optionalArray(GeminiSafetySetting),
serviceTier: Schema.optional(Schema.String),
systemInstruction: Schema.optional(GeminiSystemInstruction), systemInstruction: Schema.optional(GeminiSystemInstruction),
tools: optionalArray(GeminiTool), tools: optionalArray(GeminiTool),
toolConfig: Schema.optional(GeminiToolConfig), toolConfig: Schema.optional(GeminiToolConfig),
@@ -203,7 +246,9 @@ const thoughtSignature = (providerMetadata: ProviderMetadata | undefined) => {
const functionCallId = (providerMetadata: ProviderMetadata | undefined) => { const functionCallId = (providerMetadata: ProviderMetadata | undefined) => {
const google = providerMetadata?.google const google = providerMetadata?.google
return ProviderShared.isRecord(google) && typeof google.functionCallId === "string" ? google.functionCallId : undefined return ProviderShared.isRecord(google) && typeof google.functionCallId === "string"
? google.functionCallId
: undefined
} }
const lowerToolCall = (part: ToolCallPart) => ({ const lowerToolCall = (part: ToolCallPart) => ({
@@ -274,7 +319,7 @@ const lowerMessages = Effect.fn("Gemini.lowerMessages")(function* (request: LLMR
}) })
continue continue
} }
const content: ReadonlyArray<ToolContent> = part.result.value const content: ReadonlyArray<Tool.Content> = part.result.value
const text = content.filter((item) => item.type === "text").map((item) => item.text) const text = content.filter((item) => item.type === "text").map((item) => item.text)
const media: GeminiInlineDataPart[] = [] const media: GeminiInlineDataPart[] = []
for (const item of content) { for (const item of content) {
@@ -300,21 +345,43 @@ const lowerMessages = Effect.fn("Gemini.lowerMessages")(function* (request: LLMR
return contents return contents
}) })
const geminiOptions = (request: LLMRequest) => request.providerOptions?.gemini const resolveOptions = (request: LLMRequest) => {
const input = request.providerOptions?.gemini
const thinkingConfig = (request: LLMRequest) => { const value = input?.thinkingConfig
const value = geminiOptions(request)?.thinkingConfig const thinkingConfig = {
if (!ProviderShared.isRecord(value)) return undefined thinkingBudget:
const result = { ProviderShared.isRecord(value) && typeof value.thinkingBudget === "number" ? value.thinkingBudget : undefined,
thinkingBudget: typeof value.thinkingBudget === "number" ? value.thinkingBudget : undefined, includeThoughts:
includeThoughts: typeof value.includeThoughts === "boolean" ? value.includeThoughts : undefined, ProviderShared.isRecord(value) && typeof value.includeThoughts === "boolean"
? value.includeThoughts
: ProviderShared.isRecord(value)
? true
: undefined,
thinkingLevel:
ProviderShared.isRecord(value) && typeof value.thinkingLevel === "string" ? value.thinkingLevel : undefined,
} }
return Object.values(result).some((item) => item !== undefined) ? result : undefined return {
cachedContent: typeof input?.cachedContent === "string" ? input.cachedContent : undefined,
safetySettings: mapSafetySettings(input?.safetySettings),
serviceTier: typeof input?.serviceTier === "string" ? input.serviceTier : undefined,
thinkingConfig: Object.values(thinkingConfig).some((item) => item !== undefined) ? thinkingConfig : undefined,
}
}
function mapSafetySettings(value: unknown) {
if (!Array.isArray(value)) return undefined
const settings = value.flatMap((item) =>
ProviderShared.isRecord(item) && typeof item.category === "string" && typeof item.threshold === "string"
? [{ category: item.category, threshold: item.threshold }]
: [],
)
return settings
} }
const fromRequest = Effect.fn("Gemini.fromRequest")(function* (request: LLMRequest) { const fromRequest = Effect.fn("Gemini.fromRequest")(function* (request: LLMRequest) {
const toolsEnabled = request.tools.length > 0 && request.toolChoice?.type !== "none" const hasTools = request.tools.length > 0
const generation = request.generation const generation = request.generation
const options = resolveOptions(request)
const toolSchemaCompatibility = request.model.compatibility?.toolSchema const toolSchemaCompatibility = request.model.compatibility?.toolSchema
const generationConfig = { const generationConfig = {
maxOutputTokens: generation?.maxTokens, maxOutputTokens: generation?.maxTokens,
@@ -322,14 +389,17 @@ const fromRequest = Effect.fn("Gemini.fromRequest")(function* (request: LLMReque
topP: generation?.topP, topP: generation?.topP,
topK: generation?.topK, topK: generation?.topK,
stopSequences: generation?.stop, stopSequences: generation?.stop,
thinkingConfig: thinkingConfig(request), thinkingConfig: options.thinkingConfig,
} }
return { return {
cachedContent: options.cachedContent,
contents: yield* lowerMessages(request), contents: yield* lowerMessages(request),
safetySettings: options.safetySettings,
serviceTier: options.serviceTier,
systemInstruction: systemInstruction:
request.system.length === 0 ? undefined : { parts: [{ text: ProviderShared.joinText(request.system) }] }, request.system.length === 0 ? undefined : { parts: [{ text: ProviderShared.joinText(request.system) }] },
tools: toolsEnabled tools: hasTools
? [ ? [
{ {
functionDeclarations: request.tools.map((tool) => functionDeclarations: request.tools.map((tool) =>
@@ -338,7 +408,7 @@ const fromRequest = Effect.fn("Gemini.fromRequest")(function* (request: LLMReque
}, },
] ]
: undefined, : undefined,
toolConfig: toolsEnabled && request.toolChoice ? yield* lowerToolConfig(request.toolChoice) : undefined, toolConfig: hasTools && request.toolChoice ? yield* lowerToolConfig(request.toolChoice) : undefined,
generationConfig: Object.values(generationConfig).some((value) => value !== undefined) generationConfig: Object.values(generationConfig).some((value) => value !== undefined)
? generationConfig ? generationConfig
: undefined, : undefined,
@@ -382,10 +452,22 @@ const mapFinishReason = (finishReason: string | undefined, hasToolCalls: boolean
finishReason === "SAFETY" || finishReason === "SAFETY" ||
finishReason === "BLOCKLIST" || finishReason === "BLOCKLIST" ||
finishReason === "PROHIBITED_CONTENT" || finishReason === "PROHIBITED_CONTENT" ||
finishReason === "SPII" finishReason === "SPII" ||
finishReason === "MODEL_ARMOR" ||
finishReason === "IMAGE_PROHIBITED_CONTENT" ||
finishReason === "IMAGE_RECITATION" ||
finishReason === "LANGUAGE"
) )
return "content-filter" return "content-filter"
if (finishReason === "MALFORMED_FUNCTION_CALL") return "error" if (
finishReason === "MALFORMED_FUNCTION_CALL" ||
finishReason === "UNEXPECTED_TOOL_CALL" ||
finishReason === "NO_IMAGE" ||
finishReason === "TOO_MANY_TOOL_CALLS" ||
finishReason === "MISSING_THOUGHT_SIGNATURE" ||
finishReason === "MALFORMED_RESPONSE"
)
return "error"
return "unknown" return "unknown"
} }
@@ -402,7 +484,10 @@ const finish = (state: ParserState): ReadonlyArray<LLMEvent> =>
) )
: state.lifecycle : state.lifecycle
Lifecycle.finish(lifecycle, events, { Lifecycle.finish(lifecycle, events, {
reason: mapFinishReason(state.finishReason, state.hasToolCalls), reason: {
normalized: mapFinishReason(state.finishReason, state.hasToolCalls),
raw: state.finishReason,
},
usage: state.usage, usage: state.usage,
}) })
return events return events
+1
View File
@@ -6,3 +6,4 @@ export * as OpenAIImages from "./openai-images"
export * as OpenAICompatibleChat from "./openai-compatible-chat" export * as OpenAICompatibleChat from "./openai-compatible-chat"
export * as OpenAICompatibleResponses from "./openai-compatible-responses" export * as OpenAICompatibleResponses from "./openai-compatible-responses"
export * as OpenAIResponses from "./openai-responses" export * as OpenAIResponses from "./openai-responses"
export * as OpenResponses from "./open-responses"
File diff suppressed because it is too large Load Diff
+174 -44
View File
@@ -1,13 +1,17 @@
import { Effect, Schema } from "effect" import { Effect, Schema } from "effect"
import { Tool } from "@opencode-ai/schema/tool"
import { Route } from "../route/client" import { Route } from "../route/client"
import { Auth } from "../route/auth" import { Auth } from "../route/auth"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
import { HttpTransport } from "../route/transport" import { HttpTransport } from "../route/transport"
import { Protocol } from "../route/protocol" import { Protocol } from "../route/protocol"
import { import {
LLMError,
LLMEvent, LLMEvent,
Usage, Usage,
type FinishReason, type FinishReason,
type FinishReasonDetails,
type CacheHint,
type JsonSchema, type JsonSchema,
type LLMRequest, type LLMRequest,
type MediaPart, type MediaPart,
@@ -15,8 +19,8 @@ import {
type TextPart, type TextPart,
type ToolCallPart, type ToolCallPart,
type ToolDefinition, type ToolDefinition,
type ToolContent,
} from "../schema" } from "../schema"
import { classifyProviderFailure } from "../provider-error"
import { isRecord, JsonObject, optionalArray, optionalNull, ProviderShared } from "./shared" import { isRecord, JsonObject, optionalArray, optionalNull, ProviderShared } from "./shared"
import { OpenAIOptions } from "./utils/openai-options" import { OpenAIOptions } from "./utils/openai-options"
import { Lifecycle } from "./utils/lifecycle" import { Lifecycle } from "./utils/lifecycle"
@@ -35,6 +39,11 @@ export const PATH = "/chat/completions"
// The body schema is the provider-native JSON body. `fromRequest` below builds // The body schema is the provider-native JSON body. `fromRequest` below builds
// this shape from the common `LLMRequest`, then `Route.make` validates and // this shape from the common `LLMRequest`, then `Route.make` validates and
// JSON-encodes it before transport. // JSON-encodes it before transport.
const OpenAIChatCacheControl = Schema.Struct({
type: Schema.Literal("ephemeral"),
ttl: Schema.optional(Schema.String),
})
const OpenAIChatFunction = Schema.Struct({ const OpenAIChatFunction = Schema.Struct({
name: Schema.String, name: Schema.String,
description: Schema.String, description: Schema.String,
@@ -44,6 +53,7 @@ const OpenAIChatFunction = Schema.Struct({
const OpenAIChatTool = Schema.Struct({ const OpenAIChatTool = Schema.Struct({
type: Schema.tag("function"), type: Schema.tag("function"),
function: OpenAIChatFunction, function: OpenAIChatFunction,
cache_control: Schema.optional(OpenAIChatCacheControl),
}) })
type OpenAIChatTool = Schema.Schema.Type<typeof OpenAIChatTool> type OpenAIChatTool = Schema.Schema.Type<typeof OpenAIChatTool>
@@ -58,7 +68,11 @@ const OpenAIChatAssistantToolCall = Schema.Struct({
type OpenAIChatAssistantToolCall = Schema.Schema.Type<typeof OpenAIChatAssistantToolCall> type OpenAIChatAssistantToolCall = Schema.Schema.Type<typeof OpenAIChatAssistantToolCall>
const OpenAIChatUserContent = Schema.Union([ const OpenAIChatUserContent = Schema.Union([
Schema.Struct({ type: Schema.Literal("text"), text: Schema.String }), Schema.Struct({
type: Schema.Literal("text"),
text: Schema.String,
cache_control: Schema.optional(OpenAIChatCacheControl),
}),
Schema.Struct({ Schema.Struct({
type: Schema.Literal("image_url"), type: Schema.Literal("image_url"),
image_url: Schema.Struct({ url: Schema.String }), image_url: Schema.Struct({ url: Schema.String }),
@@ -66,7 +80,10 @@ const OpenAIChatUserContent = Schema.Union([
]) ])
const OpenAIChatMessage = Schema.Union([ const OpenAIChatMessage = Schema.Union([
Schema.Struct({ role: Schema.Literal("system"), content: Schema.String }), Schema.Struct({
role: Schema.Literal("system"),
content: Schema.Union([Schema.String, Schema.Array(OpenAIChatUserContent)]),
}),
Schema.Struct({ Schema.Struct({
role: Schema.Literal("user"), role: Schema.Literal("user"),
content: Schema.Union([Schema.String, Schema.Array(OpenAIChatUserContent)]), content: Schema.Union([Schema.String, Schema.Array(OpenAIChatUserContent)]),
@@ -80,10 +97,16 @@ const OpenAIChatMessage = Schema.Union([
reasoning: Schema.optional(Schema.String), reasoning: Schema.optional(Schema.String),
reasoning_text: Schema.optional(Schema.String), reasoning_text: Schema.optional(Schema.String),
reasoning_details: Schema.optional(Schema.Unknown), reasoning_details: Schema.optional(Schema.Unknown),
cache_control: Schema.optional(OpenAIChatCacheControl),
}), }),
[Schema.Record(Schema.String, Schema.Unknown)], [Schema.Record(Schema.String, Schema.Unknown)],
), ),
Schema.Struct({ role: Schema.Literal("tool"), tool_call_id: Schema.String, content: Schema.String }), Schema.Struct({
role: Schema.Literal("tool"),
tool_call_id: Schema.String,
content: Schema.String,
cache_control: Schema.optional(OpenAIChatCacheControl),
}),
]).pipe(Schema.toTaggedUnion("role")) ]).pipe(Schema.toTaggedUnion("role"))
type OpenAIChatMessage = Schema.Schema.Type<typeof OpenAIChatMessage> type OpenAIChatMessage = Schema.Schema.Type<typeof OpenAIChatMessage>
@@ -104,6 +127,7 @@ export const bodyFields = {
stream_options: Schema.optional(Schema.Struct({ include_usage: Schema.Boolean })), stream_options: Schema.optional(Schema.Struct({ include_usage: Schema.Boolean })),
store: Schema.optional(Schema.Boolean), store: Schema.optional(Schema.Boolean),
reasoning_effort: Schema.optional(OpenAIOptions.OpenAIReasoningEffort), reasoning_effort: Schema.optional(OpenAIOptions.OpenAIReasoningEffort),
max_completion_tokens: Schema.optional(Schema.Number),
max_tokens: Schema.optional(Schema.Number), max_tokens: Schema.optional(Schema.Number),
temperature: Schema.optional(Schema.Number), temperature: Schema.optional(Schema.Number),
top_p: Schema.optional(Schema.Number), top_p: Schema.optional(Schema.Number),
@@ -128,6 +152,7 @@ const OpenAIChatUsage = Schema.Struct({
prompt_tokens_details: optionalNull( prompt_tokens_details: optionalNull(
Schema.Struct({ Schema.Struct({
cached_tokens: Schema.optional(Schema.Number), cached_tokens: Schema.optional(Schema.Number),
cache_write_tokens: Schema.optional(Schema.Number),
}), }),
), ),
completion_tokens_details: optionalNull( completion_tokens_details: optionalNull(
@@ -164,11 +189,18 @@ const OpenAIChatDelta = Schema.StructWithRest(
const OpenAIChatChoice = Schema.Struct({ const OpenAIChatChoice = Schema.Struct({
delta: optionalNull(OpenAIChatDelta), delta: optionalNull(OpenAIChatDelta),
finish_reason: optionalNull(Schema.String), finish_reason: optionalNull(Schema.String),
native_finish_reason: optionalNull(Schema.String),
})
const OpenAIChatError = Schema.Struct({
code: optionalNull(Schema.Union([Schema.String, Schema.Number])),
message: Schema.String,
}) })
export const OpenAIChatEvent = Schema.Struct({ export const OpenAIChatEvent = Schema.Struct({
choices: Schema.Array(OpenAIChatChoice), choices: optionalNull(Schema.Array(OpenAIChatChoice)),
usage: optionalNull(OpenAIChatUsage), usage: optionalNull(OpenAIChatUsage),
error: optionalNull(OpenAIChatError),
}) })
export type OpenAIChatEvent = Schema.Schema.Type<typeof OpenAIChatEvent> export type OpenAIChatEvent = Schema.Schema.Type<typeof OpenAIChatEvent>
type OpenAIChatRequestMessage = LLMRequest["messages"][number] type OpenAIChatRequestMessage = LLMRequest["messages"][number]
@@ -184,7 +216,7 @@ export interface ParserState {
readonly pendingTools: Partial<Record<number, PendingToolDelta>> readonly pendingTools: Partial<Record<number, PendingToolDelta>>
readonly toolCallEvents: ReadonlyArray<LLMEvent> readonly toolCallEvents: ReadonlyArray<LLMEvent>
readonly usage?: Usage readonly usage?: Usage
readonly finishReason?: FinishReason readonly finishReason?: FinishReasonDetails
readonly lifecycle: Lifecycle.State readonly lifecycle: Lifecycle.State
readonly reasoningField?: string readonly reasoningField?: string
readonly reasoningDetails: Array<unknown> readonly reasoningDetails: Array<unknown>
@@ -198,13 +230,20 @@ export interface ParserState {
// Lowering is the only place that knows how common LLM messages map onto the // Lowering is the only place that knows how common LLM messages map onto the
// OpenAI Chat wire format. Keep provider quirks here instead of leaking native // OpenAI Chat wire format. Keep provider quirks here instead of leaking native
// fields into `LLMRequest`. // fields into `LLMRequest`.
const lowerTool = (tool: ToolDefinition, inputSchema: JsonSchema): OpenAIChatTool => ({ interface LoweringOptions {
readonly cacheControl?: (
cache: CacheHint | undefined,
) => Schema.Schema.Type<typeof OpenAIChatCacheControl> | undefined
}
const lowerTool = (tool: ToolDefinition, inputSchema: JsonSchema, options: LoweringOptions): OpenAIChatTool => ({
type: "function", type: "function",
function: { function: {
name: tool.name, name: tool.name,
description: tool.description, description: tool.description,
parameters: ToolSchemaProjection.openAI(inputSchema), parameters: ToolSchemaProjection.openAI(inputSchema),
}, },
cache_control: options.cacheControl?.(tool.cache),
}) })
const lowerToolChoice = (toolChoice: NonNullable<LLMRequest["toolChoice"]>) => const lowerToolChoice = (toolChoice: NonNullable<LLMRequest["toolChoice"]>) =>
@@ -246,11 +285,14 @@ const reasoningDetails = (parts: ReadonlyArray<ReasoningPart>, native: unknown)
if (isRecord(native) && Array.isArray(native.reasoning_details)) return native.reasoning_details if (isRecord(native) && Array.isArray(native.reasoning_details)) return native.reasoning_details
} }
const lowerUserMessage = Effect.fn("OpenAIChat.lowerUserMessage")(function* (message: OpenAIChatRequestMessage) { const lowerUserMessage = Effect.fn("OpenAIChat.lowerUserMessage")(function* (
message: OpenAIChatRequestMessage,
options: LoweringOptions,
) {
const content: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = [] const content: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = []
for (const part of message.content) { for (const part of message.content) {
if (part.type === "text") { if (part.type === "text") {
content.push({ type: "text", text: part.text }) content.push({ type: "text", text: part.text, cache_control: options.cacheControl?.(part.cache) })
continue continue
} }
if (part.type === "media") { if (part.type === "media") {
@@ -259,14 +301,18 @@ const lowerUserMessage = Effect.fn("OpenAIChat.lowerUserMessage")(function* (mes
} }
return yield* ProviderShared.unsupportedContent("OpenAI Chat", "user", ["text", "media"]) return yield* ProviderShared.unsupportedContent("OpenAI Chat", "user", ["text", "media"])
} }
if (content.every((part) => part.type === "text")) if (content.every((part) => part.type === "text" && part.cache_control === undefined))
return { role: "user" as const, content: content.map((part) => part.text).join("") } return {
role: "user" as const,
content: content.map((part) => (part.type === "text" ? part.text : "")).join(""),
}
return { role: "user" as const, content } return { role: "user" as const, content }
}) })
const lowerAssistantMessage = Effect.fn("OpenAIChat.lowerAssistantMessage")(function* ( const lowerAssistantMessage = Effect.fn("OpenAIChat.lowerAssistantMessage")(function* (
message: OpenAIChatRequestMessage, message: OpenAIChatRequestMessage,
configuredField?: string, configuredField?: string,
options: LoweringOptions = {},
) { ) {
const content: TextPart[] = [] const content: TextPart[] = []
const reasoning: ReasoningPart[] = [] const reasoning: ReasoningPart[] = []
@@ -304,29 +350,44 @@ const lowerAssistantMessage = Effect.fn("OpenAIChat.lowerAssistantMessage")(func
if (reasoning.length === 0) return nativeReasoning if (reasoning.length === 0) return nativeReasoning
return text return text
})() })()
const cached = message.content.findLast((part) => "cache" in part && part.cache !== undefined)
const result = { const result = {
role: "assistant" as const, role: "assistant" as const,
content: content.length === 0 ? null : ProviderShared.joinText(content), content: content.length === 0 ? null : ProviderShared.joinText(content),
tool_calls: toolCalls.length === 0 ? undefined : toolCalls, tool_calls: toolCalls.length === 0 ? undefined : toolCalls,
reasoning_details: details, reasoning_details: details,
cache_control: options.cacheControl?.(cached && "cache" in cached ? cached.cache : undefined),
} }
if (field === undefined || reasoningText === undefined) return result if (field === undefined || reasoningText === undefined) return result
return { ...result, [field]: reasoningText } return { ...result, [field]: reasoningText }
}) })
const lowerToolMessages = Effect.fn("OpenAIChat.lowerToolMessages")(function* (message: OpenAIChatRequestMessage) { const lowerToolMessages = Effect.fn("OpenAIChat.lowerToolMessages")(function* (
message: OpenAIChatRequestMessage,
options: LoweringOptions,
) {
const messages: OpenAIChatMessage[] = [] const messages: OpenAIChatMessage[] = []
const images: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = [] const images: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = []
for (const part of message.content) { for (const part of message.content) {
if (!ProviderShared.supportsContent(part, ["tool-result"])) if (!ProviderShared.supportsContent(part, ["tool-result"]))
return yield* ProviderShared.unsupportedContent("OpenAI Chat", "tool", ["tool-result"]) return yield* ProviderShared.unsupportedContent("OpenAI Chat", "tool", ["tool-result"])
if (part.result.type !== "content") { if (part.result.type !== "content") {
messages.push({ role: "tool", tool_call_id: part.id, content: ProviderShared.toolResultText(part) }) messages.push({
role: "tool",
tool_call_id: part.id,
content: ProviderShared.toolResultText(part),
cache_control: options.cacheControl?.(part.cache),
})
continue continue
} }
const content: ReadonlyArray<ToolContent> = part.result.value const content: ReadonlyArray<Tool.Content> = part.result.value
const text = content.filter((item) => item.type === "text").map((item) => item.text) const text = content.filter((item) => item.type === "text").map((item) => item.text)
messages.push({ role: "tool", tool_call_id: part.id, content: text.join("\n") }) messages.push({
role: "tool",
tool_call_id: part.id,
content: text.join("\n"),
cache_control: options.cacheControl?.(part.cache),
})
const files = content.filter((item) => item.type === "file") const files = content.filter((item) => item.type === "file")
images.push( images.push(
...(yield* Effect.forEach(files, (item) => ...(yield* Effect.forEach(files, (item) =>
@@ -340,15 +401,29 @@ const lowerToolMessages = Effect.fn("OpenAIChat.lowerToolMessages")(function* (m
const lowerMessage = Effect.fn("OpenAIChat.lowerMessage")(function* ( const lowerMessage = Effect.fn("OpenAIChat.lowerMessage")(function* (
message: OpenAIChatRequestMessage, message: OpenAIChatRequestMessage,
reasoningField?: string, reasoningField?: string,
options: LoweringOptions = {},
) { ) {
if (message.role === "user") return [yield* lowerUserMessage(message)] if (message.role === "user") return [yield* lowerUserMessage(message, options)]
if (message.role === "assistant") return [yield* lowerAssistantMessage(message, reasoningField)] if (message.role === "assistant") return [yield* lowerAssistantMessage(message, reasoningField, options)]
return (yield* lowerToolMessages(message)).messages return (yield* lowerToolMessages(message, options)).messages
}) })
const lowerMessages = Effect.fn("OpenAIChat.lowerMessages")(function* (request: LLMRequest) { const lowerMessages = Effect.fn("OpenAIChat.lowerMessages")(function* (request: LLMRequest, options: LoweringOptions) {
const system: OpenAIChatMessage[] = const system: OpenAIChatMessage[] =
request.system.length === 0 ? [] : [{ role: "system", content: ProviderShared.joinText(request.system) }] request.system.length === 0
? []
: request.system.some((part) => part.cache !== undefined) && options.cacheControl !== undefined
? [
{
role: "system",
content: request.system.map((part) => ({
type: "text",
text: part.text,
cache_control: options.cacheControl?.(part.cache),
})),
},
]
: [{ role: "system", content: ProviderShared.joinText(request.system) }]
const messages = [...system] const messages = [...system]
const pendingImages: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = [] const pendingImages: Array<Schema.Schema.Type<typeof OpenAIChatUserContent>> = []
const flushImages = () => { const flushImages = () => {
@@ -359,43 +434,70 @@ const lowerMessages = Effect.fn("OpenAIChat.lowerMessages")(function* (request:
if (message.role === "system") { if (message.role === "system") {
const part = yield* ProviderShared.wrappedSystemUpdate("OpenAI Chat", message) const part = yield* ProviderShared.wrappedSystemUpdate("OpenAI Chat", message)
if (pendingImages.length > 0) { if (pendingImages.length > 0) {
messages.push({ role: "user", content: [...pendingImages.splice(0), { type: "text", text: part.text }] }) messages.push({
role: "user",
content: [
...pendingImages.splice(0),
{ type: "text", text: part.text, cache_control: options.cacheControl?.(part.cache) },
],
})
continue continue
} }
const previous = messages.at(-1) const previous = messages.at(-1)
if (previous?.role === "user" && typeof previous.content === "string") if (previous?.role === "user" && typeof previous.content === "string")
messages[messages.length - 1] = { role: "user", content: `${previous.content}\n${part.text}` } messages[messages.length - 1] = options.cacheControl?.(part.cache)
? {
role: "user",
content: [
{ type: "text", text: previous.content },
{ type: "text", text: part.text, cache_control: options.cacheControl(part.cache) },
],
}
: { role: "user", content: `${previous.content}\n${part.text}` }
else if (previous?.role === "user" && Array.isArray(previous.content)) else if (previous?.role === "user" && Array.isArray(previous.content))
messages[messages.length - 1] = { messages[messages.length - 1] = {
role: "user", role: "user",
content: [...previous.content, { type: "text", text: part.text }], content: [
...previous.content,
{ type: "text", text: part.text, cache_control: options.cacheControl?.(part.cache) },
],
} }
else messages.push({ role: "user", content: part.text }) else
messages.push(
options.cacheControl?.(part.cache)
? {
role: "user",
content: [{ type: "text", text: part.text, cache_control: options.cacheControl(part.cache) }],
}
: { role: "user", content: part.text },
)
continue continue
} }
if (message.role === "tool") { if (message.role === "tool") {
const lowered = yield* lowerToolMessages(message) const lowered = yield* lowerToolMessages(message, options)
messages.push(...lowered.messages) messages.push(...lowered.messages)
pendingImages.push(...lowered.images) pendingImages.push(...lowered.images)
continue continue
} }
flushImages() flushImages()
messages.push(...(yield* lowerMessage(message, request.model.compatibility?.reasoningField))) messages.push(...(yield* lowerMessage(message, request.model.compatibility?.reasoningField, options)))
} }
flushImages() flushImages()
return messages return messages
}) })
const lowerOptions = Effect.fn("OpenAIChat.lowerOptions")(function* (request: LLMRequest) { const lowerOptions = (request: LLMRequest) => {
const store = OpenAIOptions.store(request) const options = OpenAIOptions.resolve(request)
const reasoningEffort = OpenAIOptions.reasoningEffort(request)
return { return {
...(store !== undefined ? { store } : {}), ...(options.store !== undefined ? { store: options.store } : {}),
...(reasoningEffort ? { reasoning_effort: reasoningEffort } : {}), ...(options.reasoningEffort ? { reasoning_effort: options.reasoningEffort } : {}),
} }
}) }
const fromRequest = Effect.fn("OpenAIChat.fromRequest")(function* (request: LLMRequest) { export const fromRequest = Effect.fn("OpenAIChat.fromRequest")(function* (
request: LLMRequest,
options: LoweringOptions = {},
) {
// `fromRequest` returns the provider body only. Endpoint, auth, framing, // `fromRequest` returns the provider body only. Endpoint, auth, framing,
// validation, and HTTP execution are composed by `Route.make`. // validation, and HTTP execution are composed by `Route.make`.
const reasoningField = request.model.compatibility?.reasoningField const reasoningField = request.model.compatibility?.reasoningField
@@ -405,26 +507,33 @@ const fromRequest = Effect.fn("OpenAIChat.fromRequest")(function* (request: LLMR
) )
const generation = request.generation const generation = request.generation
const toolSchemaCompatibility = request.model.compatibility?.toolSchema const toolSchemaCompatibility = request.model.compatibility?.toolSchema
const maxTokensField = request.model.compatibility?.maxTokensField ?? "max_tokens"
return { return {
model: request.model.id, model: request.model.id,
messages: yield* lowerMessages(request), messages: yield* lowerMessages(request, options),
tools: tools:
request.tools.length === 0 request.tools.length === 0
? undefined ? undefined
: request.tools.map((tool) => : request.tools.map((tool) =>
lowerTool(tool, ToolSchemaProjection.modelCompatibility(tool.inputSchema, toolSchemaCompatibility)), lowerTool(
tool,
ToolSchemaProjection.modelCompatibility(tool.inputSchema, toolSchemaCompatibility),
options,
),
), ),
tool_choice: request.toolChoice ? yield* lowerToolChoice(request.toolChoice) : undefined, tool_choice: request.toolChoice ? yield* lowerToolChoice(request.toolChoice) : undefined,
stream: true as const, stream: true as const,
stream_options: { include_usage: true }, stream_options: { include_usage: true },
max_tokens: generation?.maxTokens, ...(maxTokensField === "max_completion_tokens"
? { max_completion_tokens: generation?.maxTokens }
: { max_tokens: generation?.maxTokens }),
temperature: generation?.temperature, temperature: generation?.temperature,
top_p: generation?.topP, top_p: generation?.topP,
frequency_penalty: generation?.frequencyPenalty, frequency_penalty: generation?.frequencyPenalty,
presence_penalty: generation?.presencePenalty, presence_penalty: generation?.presencePenalty,
seed: generation?.seed, seed: generation?.seed,
stop: generation?.stop, stop: generation?.stop,
...(yield* lowerOptions(request)), ...lowerOptions(request),
} }
}) })
@@ -439,24 +548,27 @@ const mapFinishReason = (reason: string | null | undefined): FinishReason => {
if (reason === "length") return "length" if (reason === "length") return "length"
if (reason === "content_filter") return "content-filter" if (reason === "content_filter") return "content-filter"
if (reason === "function_call" || reason === "tool_calls") return "tool-calls" if (reason === "function_call" || reason === "tool_calls") return "tool-calls"
if (reason === "error") return "error"
return "unknown" return "unknown"
} }
// OpenAI Chat reports `prompt_tokens` (inclusive total) with a // OpenAI Chat reports `prompt_tokens` (inclusive total) with a
// `cached_tokens` subset, and `completion_tokens` (inclusive total) with // cached-read and cache-write subsets, and `completion_tokens` (inclusive
// a `reasoning_tokens` subset. We pass the inclusive totals through and // total) with a `reasoning_tokens` subset. We pass the inclusive totals
// derive the non-cached breakdown so the `LLM.Usage` contract is // through and derive the non-cached breakdown so the `LLM.Usage` contract is
// satisfied on both sides. // satisfied on both sides.
const mapUsage = (usage: OpenAIChatEvent["usage"]): Usage | undefined => { const mapUsage = (usage: OpenAIChatEvent["usage"]): Usage | undefined => {
if (!usage) return undefined if (!usage) return undefined
const cached = usage.prompt_tokens_details?.cached_tokens const cached = usage.prompt_tokens_details?.cached_tokens
const cacheWrite = usage.prompt_tokens_details?.cache_write_tokens
const reasoning = usage.completion_tokens_details?.reasoning_tokens const reasoning = usage.completion_tokens_details?.reasoning_tokens
const nonCached = ProviderShared.subtractTokens(usage.prompt_tokens, cached) const nonCached = ProviderShared.subtractTokens(usage.prompt_tokens, ProviderShared.sumTokens(cached, cacheWrite))
return new Usage({ return new Usage({
inputTokens: usage.prompt_tokens, inputTokens: usage.prompt_tokens,
outputTokens: usage.completion_tokens, outputTokens: usage.completion_tokens,
nonCachedInputTokens: nonCached, nonCachedInputTokens: nonCached,
cacheReadInputTokens: cached, cacheReadInputTokens: cached,
cacheWriteInputTokens: cacheWrite,
reasoningTokens: reasoning, reasoningTokens: reasoning,
totalTokens: ProviderShared.totalTokens(usage.prompt_tokens, usage.completion_tokens, usage.total_tokens), totalTokens: ProviderShared.totalTokens(usage.prompt_tokens, usage.completion_tokens, usage.total_tokens),
providerMetadata: { openai: usage }, providerMetadata: { openai: usage },
@@ -532,10 +644,22 @@ const reasoningMetadata = (field: ParserState["reasoningField"], details?: Reado
const step = (state: ParserState, event: OpenAIChatEvent) => const step = (state: ParserState, event: OpenAIChatEvent) =>
Effect.gen(function* () { Effect.gen(function* () {
if (event.error)
return yield* new LLMError({
module: ADAPTER,
method: "stream",
reason: classifyProviderFailure({
message: event.error.message,
code: event.error.code === undefined || event.error.code === null ? undefined : String(event.error.code),
status: typeof event.error.code === "number" ? event.error.code : undefined,
}),
})
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const usage = mapUsage(event.usage) ?? state.usage const usage = mapUsage(event.usage) ?? state.usage
const choice = event.choices[0] const choice = event.choices?.[0]
const finishReason = choice?.finish_reason ? mapFinishReason(choice.finish_reason) : state.finishReason const finishReason = choice?.finish_reason
? { normalized: mapFinishReason(choice.finish_reason), raw: choice.native_finish_reason ?? choice.finish_reason }
: state.finishReason
const delta = choice?.delta const delta = choice?.delta
const toolDeltas = delta?.tool_calls ?? [] const toolDeltas = delta?.tool_calls ?? []
let tools = state.tools let tools = state.tools
@@ -627,7 +751,13 @@ const step = (state: ParserState, event: OpenAIChatEvent) =>
const finishEvents = (state: ParserState): ReadonlyArray<LLMEvent> => { const finishEvents = (state: ParserState): ReadonlyArray<LLMEvent> => {
const events: LLMEvent[] = [] const events: LLMEvent[] = []
const hasToolCalls = state.toolCallEvents.length > 0 const hasToolCalls = state.toolCallEvents.length > 0
const reason = state.finishReason === "stop" && hasToolCalls ? "tool-calls" : state.finishReason const reason = state.finishReason
? {
...state.finishReason,
normalized:
state.finishReason.normalized === "stop" && hasToolCalls ? "tool-calls" : state.finishReason.normalized,
}
: undefined
const metadata = reasoningMetadata( const metadata = reasoningMetadata(
state.reasoningField, state.reasoningField,
state.reasoningDetailsObserved ? state.reasoningDetails : undefined, state.reasoningDetailsObserved ? state.reasoningDetails : undefined,
@@ -1,23 +1,22 @@
import { Route, type RouteRoutedModelInput } from "../route/client" import { Route, type RouteRoutedModelInput } from "../route/client"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
import { OpenAIResponses } from "./openai-responses" import { OpenResponses } from "./open-responses"
const ADAPTER = "openai-compatible-responses" const ADAPTER = "openai-compatible-responses"
export type OpenAICompatibleResponsesModelInput = RouteRoutedModelInput export type OpenAICompatibleResponsesModelInput = RouteRoutedModelInput
/** /**
* Route for providers that expose an OpenAI Responses-compatible `/responses` * Deployment adapter for providers that expose an Open Responses-compatible
* endpoint. Provider helpers configure identity, endpoint, and auth before * `/responses` endpoint. Provider helpers configure identity, endpoint, and
* model selection while this route reuses the OpenAI Responses protocol. * auth while the semantic protocol remains provider-neutral.
*/ */
export const route = Route.make({ export const route = Route.make({
id: ADAPTER, id: ADAPTER,
providerMetadataKey: "openai", providerMetadataKey: "openresponses",
protocol: OpenAIResponses.protocol, protocol: OpenResponses.protocol,
endpoint: Endpoint.path(OpenAIResponses.PATH), endpoint: Endpoint.path(OpenResponses.PATH),
transport: OpenAIResponses.httpTransport, transport: OpenResponses.httpTransport,
defaults: { providerOptions: { openai: { store: false } } },
}) })
export * as OpenAICompatibleResponses from "./openai-compatible-responses" export * as OpenAICompatibleResponses from "./openai-compatible-responses"
File diff suppressed because it is too large Load Diff
+2 -2
View File
@@ -1,4 +1,5 @@
import { Buffer } from "node:buffer" import { Buffer } from "node:buffer"
import { Tool } from "@opencode-ai/schema/tool"
import { Effect, Schema, Stream } from "effect" import { Effect, Schema, Stream } from "effect"
import * as Sse from "effect/unstable/encoding/Sse" import * as Sse from "effect/unstable/encoding/Sse"
import { Headers, HttpClientRequest } from "effect/unstable/http" import { Headers, HttpClientRequest } from "effect/unstable/http"
@@ -9,7 +10,6 @@ import {
type ContentPart, type ContentPart,
type LLMRequest, type LLMRequest,
type MediaPart, type MediaPart,
type ToolFileContent,
type TextPart, type TextPart,
type ToolResultPart, type ToolResultPart,
} from "../schema" } from "../schema"
@@ -206,7 +206,7 @@ export const validateMedia = Effect.fn("ProviderShared.validateMedia")(function*
return { mime, base64, dataUrl: `data:${mime};base64,${base64}`, bytes } satisfies ValidatedMedia return { mime, base64, dataUrl: `data:${mime};base64,${base64}`, bytes } satisfies ValidatedMedia
}) })
export const validateToolFile = (route: string, part: ToolFileContent, supportedMimes: ReadonlySet<string>) => export const validateToolFile = (route: string, part: Tool.FileContent, supportedMimes: ReadonlySet<string>) =>
validateMedia(route, { type: "media", mediaType: part.mime, data: part.uri, filename: part.name }, supportedMimes) validateMedia(route, { type: "media", mediaType: part.mime, data: part.uri, filename: part.name }, supportedMimes)
export const trimBaseUrl = (value: string) => value.replace(/\/+$/, "") export const trimBaseUrl = (value: string) => value.replace(/\/+$/, "")
+11 -8
View File
@@ -1,4 +1,4 @@
import { LLMEvent, type FinishReason, type ProviderMetadata, type Usage } from "../../schema" import { LLMEvent, type FinishReasonDetails, type ProviderMetadata, type Usage } from "../../schema"
export interface State { export interface State {
readonly stepStarted: boolean readonly stepStarted: boolean
@@ -14,16 +14,19 @@ export const stepStart = (state: State, events: LLMEvent[]): State => {
return { ...state, stepStarted: true } return { ...state, stepStarted: true }
} }
export const textDelta = (state: State, events: LLMEvent[], id: string, text: string): State => { export const textStart = (state: State, events: LLMEvent[], id: string, providerMetadata?: ProviderMetadata): State => {
if (state.text.has(id)) return state
const stepped = stepStart(state, events) const stepped = stepStart(state, events)
if (stepped.text.has(id)) { events.push(LLMEvent.textStart({ id, providerMetadata }))
events.push(LLMEvent.textDelta({ id, text }))
return stepped
}
events.push(LLMEvent.textStart({ id }), LLMEvent.textDelta({ id, text }))
return { ...stepped, text: new Set([...stepped.text, id]) } return { ...stepped, text: new Set([...stepped.text, id]) }
} }
export const textDelta = (state: State, events: LLMEvent[], id: string, text: string): State => {
const started = textStart(state, events, id)
events.push(LLMEvent.textDelta({ id, text }))
return started
}
export const reasoningStart = ( export const reasoningStart = (
state: State, state: State,
events: LLMEvent[], events: LLMEvent[],
@@ -81,7 +84,7 @@ export const finish = (
state: State, state: State,
events: LLMEvent[], events: LLMEvent[],
input: { input: {
readonly reason: FinishReason readonly reason: FinishReasonDetails
readonly usage?: Usage readonly usage?: Usage
readonly providerMetadata?: ProviderMetadata readonly providerMetadata?: ProviderMetadata
}, },
@@ -0,0 +1,65 @@
import { Schema } from "effect"
import { TextVerbosity, type LLMRequest } from "../../schema"
export const ResponseIncludables = [
"file_search_call.results",
"web_search_call.results",
"web_search_call.action.sources",
"message.input_image.image_url",
"computer_call_output.output.image_url",
"code_interpreter_call.outputs",
"reasoning.encrypted_content",
"message.output_text.logprobs",
] as const
export type ResponseIncludable = (typeof ResponseIncludables)[number]
export const ServiceTiers = ["auto", "default", "flex", "priority"] as const
export type ServiceTier = (typeof ServiceTiers)[number]
const TEXT_VERBOSITY = new Set<string>(["low", "medium", "high"])
const INCLUDABLES = new Set<string>(ResponseIncludables)
const SERVICE_TIERS = new Set<string>(ServiceTiers)
const isTextVerbosity = (value: unknown): value is Schema.Schema.Type<typeof TextVerbosity> =>
typeof value === "string" && TEXT_VERBOSITY.has(value)
const isServiceTier = (value: unknown): value is ServiceTier => typeof value === "string" && SERVICE_TIERS.has(value)
export const ReasoningEffort = Schema.String
export const TextVerbositySchema = TextVerbosity
export const ResponseIncludableSchema = Schema.Literals(ResponseIncludables)
export const ServiceTierSchema = Schema.Literals(ServiceTiers)
export interface Resolved {
readonly instructions?: string
readonly store?: boolean
readonly promptCacheKey?: string
readonly reasoningEffort?: string
readonly reasoningSummary?: "auto" | "concise" | "detailed"
readonly include?: ReadonlyArray<ResponseIncludable>
readonly textVerbosity?: Schema.Schema.Type<typeof TextVerbosity>
readonly serviceTier?: ServiceTier
}
export const resolve = (request: LLMRequest): Resolved => {
const input = request.providerOptions?.[request.model.route.providerMetadataKey ?? "openresponses"]
const include = Array.isArray(input?.include)
? input.include.filter((entry): entry is ResponseIncludable => INCLUDABLES.has(entry))
: []
const reasoningSummary = input?.reasoningSummary
return {
instructions: typeof input?.instructions === "string" ? input.instructions : undefined,
store: typeof input?.store === "boolean" ? input.store : undefined,
promptCacheKey: typeof input?.promptCacheKey === "string" ? input.promptCacheKey : undefined,
reasoningEffort: typeof input?.reasoningEffort === "string" ? input.reasoningEffort : undefined,
reasoningSummary:
reasoningSummary === "auto" || reasoningSummary === "concise" || reasoningSummary === "detailed"
? reasoningSummary
: undefined,
include: include.length > 0 ? include : undefined,
textVerbosity: isTextVerbosity(input?.textVerbosity) ? input.textVerbosity : undefined,
serviceTier: isServiceTier(input?.serviceTier) ? input.serviceTier : undefined,
}
}
export * as OpenResponsesOptions from "./open-responses-options"
@@ -1,85 +1,23 @@
import { Schema } from "effect" import { ReasoningEfforts } from "../../schema"
import type { LLMRequest, TextVerbosity as TextVerbosityValue } from "../../schema" import { OpenResponsesOptions } from "./open-responses-options"
import { ReasoningEfforts, TextVerbosity } from "../../schema"
export const OpenAIReasoningEfforts = ReasoningEfforts export const OpenAIReasoningEfforts = ReasoningEfforts
export type OpenAIReasoningEffort = string export type OpenAIReasoningEffort = string
// Mirrors OpenAI's `ResponseIncludable` union from the official SDK. Keep this // Mirrors OpenAI's `ResponseIncludable` union from the official SDK. Keep this
// in lockstep with `openai-node/src/resources/responses/responses.ts`. // in lockstep with `openai-node/src/resources/responses/responses.ts`.
export const OpenAIResponseIncludables = [ export const OpenAIResponseIncludables = OpenResponsesOptions.ResponseIncludables
"file_search_call.results", export type OpenAIResponseIncludable = OpenResponsesOptions.ResponseIncludable
"web_search_call.results", export const OpenAIServiceTiers = OpenResponsesOptions.ServiceTiers
"web_search_call.action.sources", export type OpenAIServiceTier = OpenResponsesOptions.ServiceTier
"message.input_image.image_url",
"computer_call_output.output.image_url",
"code_interpreter_call.outputs",
"reasoning.encrypted_content",
"message.output_text.logprobs",
] as const
export type OpenAIResponseIncludable = (typeof OpenAIResponseIncludables)[number]
export const OpenAIServiceTiers = ["auto", "default", "flex", "priority"] as const
export type OpenAIServiceTier = (typeof OpenAIServiceTiers)[number]
const TEXT_VERBOSITY = new Set<string>(["low", "medium", "high"]) export const OpenAIReasoningEffort = OpenResponsesOptions.ReasoningEffort
const INCLUDABLES = new Set<string>(OpenAIResponseIncludables) export const OpenAITextVerbosity = OpenResponsesOptions.TextVerbositySchema
const SERVICE_TIERS = new Set<string>(OpenAIServiceTiers) export const OpenAIResponseIncludable = OpenResponsesOptions.ResponseIncludableSchema
export const OpenAIServiceTier = OpenResponsesOptions.ServiceTierSchema
export const OpenAIReasoningEffort = Schema.String
export const OpenAITextVerbosity = TextVerbosity
export const OpenAIResponseIncludable = Schema.Literals(OpenAIResponseIncludables)
export const OpenAIServiceTier = Schema.Literals(OpenAIServiceTiers)
export const isReasoningEffort = (effort: unknown): effort is OpenAIReasoningEffort => typeof effort === "string" export const isReasoningEffort = (effort: unknown): effort is OpenAIReasoningEffort => typeof effort === "string"
const isTextVerbosity = (value: unknown): value is TextVerbosityValue => export const resolve = OpenResponsesOptions.resolve
typeof value === "string" && TEXT_VERBOSITY.has(value)
const options = (request: LLMRequest) => request.providerOptions?.openai
export const store = (request: LLMRequest): boolean | undefined => {
const value = options(request)?.store
return typeof value === "boolean" ? value : undefined
}
export const reasoningEffort = (request: LLMRequest): string | undefined => {
const value = options(request)?.reasoningEffort
return typeof value === "string" ? value : undefined
}
export const reasoningSummary = (request: LLMRequest): "auto" | undefined =>
options(request)?.reasoningSummary === "auto" ? "auto" : undefined
// Resolve the OpenAI Responses `include` field. Filters out unknown
// includable values defensively so a typo in upstream config drops the
// invalid entry instead of poisoning the wire body. An empty array (either
// passed directly or produced by filtering) is treated as "no include" and
// returns undefined so the request body omits the field entirely.
export const include = (request: LLMRequest): ReadonlyArray<OpenAIResponseIncludable> | undefined => {
const value = options(request)?.include
if (!Array.isArray(value)) return undefined
const filtered = value.filter((entry): entry is OpenAIResponseIncludable => INCLUDABLES.has(entry))
return filtered.length > 0 ? filtered : undefined
}
export const promptCacheKey = (request: LLMRequest) => {
const value = options(request)?.promptCacheKey
return typeof value === "string" ? value : undefined
}
export const textVerbosity = (request: LLMRequest) => {
const value = options(request)?.textVerbosity
return isTextVerbosity(value) ? value : undefined
}
export const serviceTier = (request: LLMRequest) => {
const value = options(request)?.serviceTier
return typeof value === "string" && SERVICE_TIERS.has(value) ? (value as OpenAIServiceTier) : undefined
}
export const instructions = (request: LLMRequest) => {
const value = options(request)?.instructions
return typeof value === "string" ? value : undefined
}
export * as OpenAIOptions from "./openai-options" export * as OpenAIOptions from "./openai-options"
@@ -63,6 +63,8 @@ const openAI = (schema: JsonSchema): JsonSchema => {
return isRecord(normalized) ? normalized : { type: "object" } return isRecord(normalized) ? normalized : { type: "object" }
} }
const responses = openAI
const gemini = (schema: JsonSchema): JsonSchema => GeminiToolSchema.convert(schema) ?? {} const gemini = (schema: JsonSchema): JsonSchema => GeminiToolSchema.convert(schema) ?? {}
const modelCompatibility = ( const modelCompatibility = (
@@ -83,4 +85,5 @@ export const ToolSchemaProjection = {
modelCompatibility, modelCompatibility,
moonshot, moonshot,
openAI, openAI,
responses,
} as const } as const
+1 -2
View File
@@ -135,7 +135,7 @@ export function classifyProviderFailure(input: ProviderFailure): LLMError["reaso
rateLimit: input.rateLimit, rateLimit: input.rateLimit,
}) })
} }
if (input.status !== undefined && input.status >= 500) if (input.status === 408 || input.status === 409 || (input.status !== undefined && input.status >= 500))
return new ProviderInternalReason({ return new ProviderInternalReason({
...common, ...common,
status: input.status, status: input.status,
@@ -145,7 +145,6 @@ export function classifyProviderFailure(input: ProviderFailure): LLMError["reaso
if ( if (
input.status === 400 || input.status === 400 ||
input.status === 404 || input.status === 404 ||
input.status === 409 ||
input.status === 413 || input.status === 413 ||
input.status === 422 input.status === 422
) )
+8 -3
View File
@@ -1,16 +1,21 @@
import type { Model } from "./schema" import type { Model, ProviderOptions } from "./schema"
export interface Settings extends Readonly<Record<string, unknown>> { export interface Settings extends Readonly<Record<string, unknown>> {
readonly baseURL?: string
readonly headers?: Readonly<Record<string, string>> readonly headers?: Readonly<Record<string, string>>
readonly body?: Readonly<Record<string, unknown>> readonly body?: Readonly<Record<string, unknown>>
readonly limits?: { readonly limits?: {
readonly context: number readonly context: number
readonly input?: number
readonly output: number readonly output: number
} }
} }
export interface Definition<ProviderSettings extends Settings = Settings> { export interface Definition<
readonly model: (modelID: string, settings: ProviderSettings) => Model ProviderSettings extends Settings = Settings,
Options extends ProviderOptions = ProviderOptions,
> {
readonly model: (modelID: string, settings: ProviderSettings) => Model<Options>
} }
export * as ProviderPackage from "./provider-package" export * as ProviderPackage from "./provider-package"
@@ -5,12 +5,17 @@ import type { ProviderAuthOption } from "../route/auth-options"
import type { RouteDefaultsInput } from "../route/client" import type { RouteDefaultsInput } from "../route/client"
import { ProviderID, type ModelID } from "../schema" import { ProviderID, type ModelID } from "../schema"
export type AnthropicOptionsInput = AnthropicMessages.OptionsInput
export type AnthropicProviderOptionsInput = AnthropicMessages.ProviderOptionsInput
export type AnthropicThinkingInput = AnthropicMessages.ThinkingInput
export const id = ProviderID.make("anthropic-compatible") export const id = ProviderID.make("anthropic-compatible")
export type Config = RouteDefaultsInput & export type Config = RouteDefaultsInput &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly provider?: string readonly provider?: string
readonly baseURL: string readonly baseURL: string
readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
} }
export type Settings = ProviderPackage.Settings & export type Settings = ProviderPackage.Settings &
@@ -20,6 +25,7 @@ export type Settings = ProviderPackage.Settings &
) & { ) & {
readonly baseURL: string readonly baseURL: string
readonly provider?: string readonly provider?: string
readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
} }
export const routes = [AnthropicMessages.route] export const routes = [AnthropicMessages.route]
@@ -41,7 +47,7 @@ export const configure = (input: Config) => {
}) })
return { return {
id: ProviderID.make(provider), id: ProviderID.make(provider),
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<AnthropicMessages.ProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -51,7 +57,10 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, AnthropicMessages.ProviderOptionsInput>["model"] = (
modelID,
settings,
) => {
if (settings.apiKey !== undefined && settings.authToken !== undefined) if (settings.apiKey !== undefined && settings.authToken !== undefined)
throw new Error("Anthropic-compatible apiKey cannot be combined with authToken") throw new Error("Anthropic-compatible apiKey cannot be combined with authToken")
return configure({ return configure({
@@ -61,6 +70,7 @@ export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, se
http: settings.body === undefined ? undefined : { body: { ...settings.body } }, http: settings.body === undefined ? undefined : { body: { ...settings.body } },
limits: settings.limits, limits: settings.limits,
provider: settings.provider, provider: settings.provider,
providerOptions: settings.providerOptions,
}).model(modelID) }).model(modelID)
} }
+15 -2
View File
@@ -6,11 +6,19 @@ import { ProviderID, type ModelID } from "../schema"
import { AnthropicMessages } from "../protocols/anthropic-messages" import { AnthropicMessages } from "../protocols/anthropic-messages"
import { AnthropicCompatible } from "./anthropic-compatible" import { AnthropicCompatible } from "./anthropic-compatible"
export type AnthropicOptionsInput = AnthropicMessages.OptionsInput
export type AnthropicProviderOptionsInput = AnthropicMessages.ProviderOptionsInput
export type AnthropicThinkingInput = AnthropicMessages.ThinkingInput
export const id = ProviderID.make("anthropic") export const id = ProviderID.make("anthropic")
export const routes = [AnthropicMessages.route] export const routes = [AnthropicMessages.route]
export type Config = RouteDefaultsInput & ProviderAuthOption<"optional"> & { readonly baseURL?: string } export type Config = RouteDefaultsInput &
ProviderAuthOption<"optional"> & {
readonly baseURL?: string
readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
}
export type Settings = ProviderPackage.Settings & export type Settings = ProviderPackage.Settings &
( (
@@ -18,6 +26,7 @@ export type Settings = ProviderPackage.Settings &
| { readonly apiKey?: never; readonly authToken?: string } | { readonly apiKey?: never; readonly authToken?: string }
) & { ) & {
readonly baseURL?: string readonly baseURL?: string
readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
} }
const auth = (options: ProviderAuthOption<"optional">) => { const auth = (options: ProviderAuthOption<"optional">) => {
@@ -43,7 +52,10 @@ export const configure = (input: Config = {}) => {
} }
export const provider = configure() export const provider = configure()
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, AnthropicMessages.ProviderOptionsInput>["model"] = (
modelID,
settings,
) => {
if (settings.apiKey !== undefined && settings.authToken !== undefined) if (settings.apiKey !== undefined && settings.authToken !== undefined)
throw new Error("Anthropic apiKey cannot be combined with authToken") throw new Error("Anthropic apiKey cannot be combined with authToken")
return configure({ return configure({
@@ -52,5 +64,6 @@ export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, se
headers: settings.headers === undefined ? undefined : { ...settings.headers }, headers: settings.headers === undefined ? undefined : { ...settings.headers },
http: settings.body === undefined ? undefined : { body: { ...settings.body } }, http: settings.body === undefined ? undefined : { body: { ...settings.body } },
limits: settings.limits, limits: settings.limits,
providerOptions: settings.providerOptions,
}).model(modelID) }).model(modelID)
} }
+14 -6
View File
@@ -99,10 +99,14 @@ export const configure = (input: Config) => {
const modelDefaults = defaults(input) const modelDefaults = defaults(input)
const responses = (modelID: string | ModelID) => const responses = (modelID: string | ModelID) =>
configuredResponsesRoute.with(withOpenAIOptions(modelID, modelDefaults)).model({ id: modelID }) configuredResponsesRoute
.with(withOpenAIOptions(modelID, modelDefaults))
.model<OpenAIProviderOptionsInput>({ id: modelID })
const chat = (modelID: string | ModelID) => const chat = (modelID: string | ModelID) =>
configuredChatRoute.with(withOpenAIOptions(modelID, modelDefaults)).model({ id: modelID }) configuredChatRoute
.with(withOpenAIOptions(modelID, modelDefaults))
.model<OpenAIProviderOptionsInput>({ id: modelID })
return { return {
id, id,
@@ -133,8 +137,12 @@ const config = (settings: Settings): Config => {
throw new Error("Azure requires resourceName or baseURL") throw new Error("Azure requires resourceName or baseURL")
} }
export const responsesModel: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => export const responsesModel: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (
configure(config(settings)).responses(modelID) modelID,
export const chatModel: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => settings,
configure(config(settings)).chat(modelID) ) => configure(config(settings)).responses(modelID)
export const chatModel: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (
modelID,
settings,
) => configure(config(settings)).chat(modelID)
export const model = responsesModel export const model = responsesModel
+10 -4
View File
@@ -4,6 +4,7 @@ import { Auth } from "../route/auth"
import { AuthOptions, type AtLeastOne, type ProviderAuthOption } from "../route/auth-options" import { AuthOptions, type AtLeastOne, type ProviderAuthOption } from "../route/auth-options"
import type { RouteDefaultsInput } from "../route/client" import type { RouteDefaultsInput } from "../route/client"
import { ProviderID, type ModelID } from "../schema" import { ProviderID, type ModelID } from "../schema"
import type { OpenAIProviderOptionsInput } from "./openai-options"
export const aiGatewayID = ProviderID.make("cloudflare-ai-gateway") export const aiGatewayID = ProviderID.make("cloudflare-ai-gateway")
export const workersAIID = ProviderID.make("cloudflare-workers-ai") export const workersAIID = ProviderID.make("cloudflare-workers-ai")
@@ -20,10 +21,11 @@ type GatewayURL = AtLeastOne<{
} }
export type AIGatewayOptions = GatewayURL & export type AIGatewayOptions = GatewayURL &
RouteDefaultsInput & Omit<RouteDefaultsInput, "providerOptions"> &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
/** Cloudflare AI Gateway authentication token. Sent as `cf-aig-authorization`. */ /** Cloudflare AI Gateway authentication token. Sent as `cf-aig-authorization`. */
readonly gatewayApiKey?: CloudflareSecret readonly gatewayApiKey?: CloudflareSecret
readonly providerOptions?: OpenAIProviderOptionsInput
} }
type WorkersAIURL = AtLeastOne<{ type WorkersAIURL = AtLeastOne<{
@@ -31,7 +33,11 @@ type WorkersAIURL = AtLeastOne<{
readonly baseURL: string readonly baseURL: string
}> }>
export type WorkersAIOptions = WorkersAIURL & RouteDefaultsInput & ProviderAuthOption<"optional"> export type WorkersAIOptions = WorkersAIURL &
Omit<RouteDefaultsInput, "providerOptions"> &
ProviderAuthOption<"optional"> & {
readonly providerOptions?: OpenAIProviderOptionsInput
}
export const aiGatewayBaseURL = (input: GatewayURL) => { export const aiGatewayBaseURL = (input: GatewayURL) => {
if (input.baseURL) return input.baseURL if (input.baseURL) return input.baseURL
@@ -98,7 +104,7 @@ const configureAIGateway = (options: AIGatewayOptions) => {
}) })
return { return {
id: aiGatewayID, id: aiGatewayID,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenAIProviderOptionsInput>({ id: modelID }),
configure: configureAIGateway, configure: configureAIGateway,
} }
} }
@@ -111,7 +117,7 @@ const configureWorkersAI = (options: WorkersAIOptions) => {
}) })
return { return {
id: workersAIID, id: workersAIID,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenAIProviderOptionsInput>({ id: modelID }),
configure: configureWorkersAI, configure: configureWorkersAI,
} }
} }
+4 -2
View File
@@ -50,9 +50,11 @@ export const configure = (options: ModelOptions) => {
const responsesRoute = configuredResponsesRoute(options) const responsesRoute = configuredResponsesRoute(options)
const chatRoute = configuredChatRoute(options) const chatRoute = configuredChatRoute(options)
const responses = (modelID: string | ModelID) => const responses = (modelID: string | ModelID) =>
responsesRoute.with(withOpenAIOptions(modelID, defaults(options))).model({ id: modelID }) responsesRoute
.with(withOpenAIOptions(modelID, defaults(options)))
.model<OpenAIProviderOptionsInput>({ id: modelID })
const chat = (modelID: string | ModelID) => const chat = (modelID: string | ModelID) =>
chatRoute.with(withOpenAIOptions(modelID, defaults(options))).model({ id: modelID }) chatRoute.with(withOpenAIOptions(modelID, defaults(options))).model<OpenAIProviderOptionsInput>({ id: modelID })
return { return {
id, id,
model: (modelID: string | ModelID) => model: (modelID: string | ModelID) =>
@@ -1,8 +1,9 @@
import type { ProviderPackage } from "../provider-package" import type { ProviderPackage } from "../provider-package"
import { OpenAICompatibleChat } from "../protocols/openai-compatible-chat" import { OpenAICompatibleChat } from "../protocols/openai-compatible-chat"
import type { RouteDefaultsInput } from "../route/client" import type { RouteDefaultsInput } from "../route/client"
import { ProviderID, type ModelID, type ProviderOptions } from "../schema" import { ProviderID, type ModelID } from "../schema"
import { GoogleVertexShared } from "./google-vertex-shared" import { GoogleVertexShared } from "./google-vertex-shared"
import type { OpenAIProviderOptionsInput } from "./openai-options"
export const id = ProviderID.make("google-vertex") export const id = ProviderID.make("google-vertex")
@@ -11,6 +12,7 @@ export type Config = RouteDefaultsInput &
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: OpenAIProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
@@ -19,7 +21,7 @@ export interface Settings extends ProviderPackage.Settings {
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: ProviderOptions readonly providerOptions?: OpenAIProviderOptionsInput
} }
const route = OpenAICompatibleChat.route.with({ const route = OpenAICompatibleChat.route.with({
@@ -56,7 +58,7 @@ export const configure = (input: Config = {}) => {
const route = configuredRoute(input) const route = configuredRoute(input)
return { return {
id, id,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenAIProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -66,7 +68,7 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (modelID, settings) => {
if (settings.apiKey !== undefined) throw new Error("Google Vertex Chat does not support API keys") if (settings.apiKey !== undefined) throw new Error("Google Vertex Chat does not support API keys")
return configure({ return configure({
accessToken: settings.accessToken, accessToken: settings.accessToken,
@@ -6,9 +6,13 @@ import { Route, type RouteDefaultsInput } from "../route/client"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
import { Framing } from "../route/framing" import { Framing } from "../route/framing"
import { Protocol } from "../route/protocol" import { Protocol } from "../route/protocol"
import { ProviderID, type ModelID, type ProviderOptions } from "../schema" import { ProviderID, type ModelID } from "../schema"
import { GoogleVertexShared } from "./google-vertex-shared" import { GoogleVertexShared } from "./google-vertex-shared"
export type AnthropicOptionsInput = AnthropicMessages.OptionsInput
export type AnthropicProviderOptionsInput = AnthropicMessages.ProviderOptionsInput
export type AnthropicThinkingInput = AnthropicMessages.ThinkingInput
const VERSION = "vertex-2023-10-16" as const const VERSION = "vertex-2023-10-16" as const
// models.dev uses this provider id even though the API contract is Anthropic Messages. // models.dev uses this provider id even though the API contract is Anthropic Messages.
@@ -19,6 +23,7 @@ export type Config = RouteDefaultsInput &
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
@@ -27,7 +32,7 @@ export interface Settings extends ProviderPackage.Settings {
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: ProviderOptions readonly providerOptions?: AnthropicMessages.ProviderOptionsInput
} }
const route = Route.make({ const route = Route.make({
@@ -86,7 +91,7 @@ export const configure = (input: Config = {}) => {
const route = configuredRoute(input) const route = configuredRoute(input)
return { return {
id, id,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<AnthropicMessages.ProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -96,7 +101,10 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, AnthropicMessages.ProviderOptionsInput>["model"] = (
modelID,
settings,
) => {
if (settings.apiKey !== undefined) throw new Error("Google Vertex Messages does not support API keys") if (settings.apiKey !== undefined) throw new Error("Google Vertex Messages does not support API keys")
return configure({ return configure({
accessToken: settings.accessToken, accessToken: settings.accessToken,
@@ -1,8 +1,9 @@
import type { ProviderPackage } from "../provider-package" import type { ProviderPackage } from "../provider-package"
import { OpenAICompatibleResponses } from "../protocols/openai-compatible-responses" import { OpenAICompatibleResponses } from "../protocols/openai-compatible-responses"
import type { RouteDefaultsInput } from "../route/client" import type { RouteDefaultsInput } from "../route/client"
import { ProviderID, type ModelID, type ProviderOptions } from "../schema" import { ProviderID, type ModelID } from "../schema"
import { GoogleVertexShared } from "./google-vertex-shared" import { GoogleVertexShared } from "./google-vertex-shared"
import type { OpenResponsesProviderOptionsInput } from "./open-responses-options"
export const id = ProviderID.make("google-vertex") export const id = ProviderID.make("google-vertex")
@@ -11,6 +12,7 @@ export type Config = RouteDefaultsInput &
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: OpenResponsesProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
@@ -19,12 +21,13 @@ export interface Settings extends ProviderPackage.Settings {
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: ProviderOptions readonly providerOptions?: OpenResponsesProviderOptionsInput
} }
const route = OpenAICompatibleResponses.route.with({ const route = OpenAICompatibleResponses.route.with({
id: "google-vertex-responses", id: "google-vertex-responses",
provider: id, provider: id,
providerOptions: { openresponses: { store: false } },
}) })
export const routes = [route] export const routes = [route]
@@ -57,7 +60,7 @@ export const configure = (input: Config = {}) => {
const route = configuredRoute(input) const route = configuredRoute(input)
return { return {
id, id,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenResponsesProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -67,7 +70,10 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, OpenResponsesProviderOptionsInput>["model"] = (
modelID,
settings,
) => {
if (settings.apiKey !== undefined) throw new Error("Google Vertex Responses does not support API keys") if (settings.apiKey !== undefined) throw new Error("Google Vertex Responses does not support API keys")
return configure({ return configure({
accessToken: settings.accessToken, accessToken: settings.accessToken,
+12 -4
View File
@@ -4,9 +4,12 @@ import { Auth } from "../route/auth"
import { Route, type RouteDefaultsInput } from "../route/client" import { Route, type RouteDefaultsInput } from "../route/client"
import { Endpoint } from "../route/endpoint" import { Endpoint } from "../route/endpoint"
import { Framing } from "../route/framing" import { Framing } from "../route/framing"
import { ProviderID, type ModelID, type ProviderOptions } from "../schema" import { ProviderID, type ModelID } from "../schema"
import { GoogleVertexShared } from "./google-vertex-shared" import { GoogleVertexShared } from "./google-vertex-shared"
export type GeminiOptionsInput = Gemini.OptionsInput
export type GeminiProviderOptionsInput = Gemini.ProviderOptionsInput
export const id = ProviderID.make("google-vertex") export const id = ProviderID.make("google-vertex")
export type Config = RouteDefaultsInput & export type Config = RouteDefaultsInput &
@@ -14,6 +17,7 @@ export type Config = RouteDefaultsInput &
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: Gemini.ProviderOptionsInput
} }
export type Settings = ProviderPackage.Settings & export type Settings = ProviderPackage.Settings &
@@ -24,7 +28,7 @@ export type Settings = ProviderPackage.Settings &
readonly baseURL?: string readonly baseURL?: string
readonly location?: string readonly location?: string
readonly project?: string readonly project?: string
readonly providerOptions?: ProviderOptions readonly providerOptions?: Gemini.ProviderOptionsInput
} }
const route = Route.make({ const route = Route.make({
@@ -73,7 +77,8 @@ const configuredRoute = (input: Config, modelID: string | ModelID) => {
export const configure = (input: Config = {}) => { export const configure = (input: Config = {}) => {
return { return {
id, id,
model: (modelID: string | ModelID) => configuredRoute(input, modelID).model({ id: modelID }), model: (modelID: string | ModelID) =>
configuredRoute(input, modelID).model<Gemini.ProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -82,7 +87,10 @@ export const provider = {
id, id,
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, Gemini.ProviderOptionsInput>["model"] = (
modelID,
settings,
) => {
if (settings.apiKey !== undefined && settings.accessToken !== undefined) if (settings.apiKey !== undefined && settings.accessToken !== undefined)
throw new Error("Google Vertex apiKey cannot be combined with accessToken or auth") throw new Error("Google Vertex apiKey cannot be combined with accessToken or auth")
return configure({ return configure({
+7 -4
View File
@@ -2,11 +2,13 @@ import type { RouteDefaultsInput } from "../route/client"
import { Auth } from "../route/auth" import { Auth } from "../route/auth"
import type { ProviderAuthOption } from "../route/auth-options" import type { ProviderAuthOption } from "../route/auth-options"
import type { ProviderPackage } from "../provider-package" import type { ProviderPackage } from "../provider-package"
import { HttpOptions, ProviderID, mergeHttpOptions, type ModelID, type ProviderOptions } from "../schema" import { HttpOptions, ProviderID, mergeHttpOptions, type ModelID } from "../schema"
import { Gemini } from "../protocols/gemini" import { Gemini } from "../protocols/gemini"
import { GoogleImages } from "../protocols/google-images" import { GoogleImages } from "../protocols/google-images"
export type { GoogleImageOptions } from "../protocols/google-images" export type { GoogleImageOptions } from "../protocols/google-images"
export type GeminiOptionsInput = Gemini.OptionsInput
export type GeminiProviderOptionsInput = Gemini.ProviderOptionsInput
export const id = ProviderID.make("google") export const id = ProviderID.make("google")
@@ -15,12 +17,13 @@ export const routes = [Gemini.route]
export type Config = RouteDefaultsInput & export type Config = RouteDefaultsInput &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly baseURL?: string readonly baseURL?: string
readonly providerOptions?: Gemini.ProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
readonly apiKey?: string readonly apiKey?: string
readonly baseURL?: string readonly baseURL?: string
readonly providerOptions?: ProviderOptions readonly providerOptions?: Gemini.ProviderOptionsInput
} }
const auth = (options: ProviderAuthOption<"optional">) => { const auth = (options: ProviderAuthOption<"optional">) => {
@@ -47,14 +50,14 @@ export const configure = (input: Config = {}) => {
}) })
return { return {
id, id,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<Gemini.ProviderOptionsInput>({ id: modelID }),
image, image,
configure, configure,
} }
} }
export const provider = configure() export const provider = configure()
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => export const model: ProviderPackage.Definition<Settings, Gemini.ProviderOptionsInput>["model"] = (modelID, settings) =>
configure({ configure({
apiKey: settings.apiKey, apiKey: settings.apiKey,
baseURL: settings.baseURL, baseURL: settings.baseURL,
@@ -0,0 +1,20 @@
import type { ResponseIncludable, ServiceTier } from "../protocols/utils/open-responses-options"
import type { ProviderOptions, ReasoningEffort, TextVerbosity } from "../schema"
export interface OpenResponsesOptionsInput {
readonly [key: string]: unknown
readonly instructions?: string
readonly store?: boolean
readonly promptCacheKey?: string
readonly reasoningEffort?: ReasoningEffort
readonly reasoningSummary?: "auto" | "concise" | "detailed"
readonly include?: ReadonlyArray<ResponseIncludable>
readonly textVerbosity?: TextVerbosity
readonly serviceTier?: ServiceTier
}
export type OpenResponsesProviderOptionsInput = ProviderOptions & {
readonly openresponses?: OpenResponsesOptionsInput
}
export * as OpenResponsesProviderOptions from "./open-responses-options"
@@ -3,7 +3,9 @@ import { OpenAICompatibleResponses } from "../protocols/openai-compatible-respon
import { AuthOptions, type ProviderAuthOption } from "../route/auth-options" import { AuthOptions, type ProviderAuthOption } from "../route/auth-options"
import type { RouteDefaultsInput } from "../route/client" import type { RouteDefaultsInput } from "../route/client"
import { ProviderID, type ModelID } from "../schema" import { ProviderID, type ModelID } from "../schema"
import type { OpenAIProviderOptionsInput } from "./openai-options" import type { OpenResponsesProviderOptionsInput } from "./open-responses-options"
export type { OpenResponsesOptionsInput, OpenResponsesProviderOptionsInput } from "./open-responses-options"
export const id = ProviderID.make("openai-compatible") export const id = ProviderID.make("openai-compatible")
@@ -11,13 +13,14 @@ export type Config = RouteDefaultsInput &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly provider?: string readonly provider?: string
readonly baseURL: string readonly baseURL: string
readonly providerOptions?: OpenResponsesProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
readonly apiKey?: string readonly apiKey?: string
readonly baseURL: string readonly baseURL: string
readonly provider?: string readonly provider?: string
readonly providerOptions?: OpenAIProviderOptionsInput readonly providerOptions?: OpenResponsesProviderOptionsInput
} }
export const routes = [OpenAICompatibleResponses.route] export const routes = [OpenAICompatibleResponses.route]
@@ -33,7 +36,7 @@ export const configure = (input: Config) => {
}) })
return { return {
id: ProviderID.make(provider), id: ProviderID.make(provider),
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenResponsesProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
@@ -43,7 +46,10 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => export const model: ProviderPackage.Definition<Settings, OpenResponsesProviderOptionsInput>["model"] = (
modelID,
settings,
) =>
configure({ configure({
apiKey: settings.apiKey, apiKey: settings.apiKey,
baseURL: settings.baseURL, baseURL: settings.baseURL,
@@ -4,13 +4,15 @@ import type { RouteDefaultsInput } from "../route/client"
import { AuthOptions, type ProviderAuthOption } from "../route/auth-options" import { AuthOptions, type ProviderAuthOption } from "../route/auth-options"
import type { ProviderPackage } from "../provider-package" import type { ProviderPackage } from "../provider-package"
import { profiles, type OpenAICompatibleProfile } from "./openai-compatible-profile" import { profiles, type OpenAICompatibleProfile } from "./openai-compatible-profile"
import type { OpenAIProviderOptionsInput } from "./openai-options"
export const id = ProviderID.make("openai-compatible") export const id = ProviderID.make("openai-compatible")
type GenericModelOptions = RouteDefaultsInput & type GenericModelOptions = Omit<RouteDefaultsInput, "providerOptions"> &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly provider?: string readonly provider?: string
readonly baseURL: string readonly baseURL: string
readonly providerOptions?: OpenAIProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings { export interface Settings extends ProviderPackage.Settings {
@@ -19,9 +21,10 @@ export interface Settings extends ProviderPackage.Settings {
readonly provider?: string readonly provider?: string
} }
export type FamilyModelOptions = RouteDefaultsInput & export type FamilyModelOptions = Omit<RouteDefaultsInput, "providerOptions"> &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly baseURL?: string readonly baseURL?: string
readonly providerOptions?: OpenAIProviderOptionsInput
} }
export const routes = [OpenAICompatibleChat.route] export const routes = [OpenAICompatibleChat.route]
@@ -37,7 +40,8 @@ export const configure = (input: GenericModelOptions) => {
}) })
return { return {
id: ProviderID.make(provider), id: ProviderID.make(provider),
model: (modelID: string | ModelID) => route.model({ id: modelID, provider: ProviderID.make(provider) }), model: (modelID: string | ModelID) =>
route.model<OpenAIProviderOptionsInput>({ id: modelID, provider: ProviderID.make(provider) }),
configure, configure,
} }
} }
@@ -63,7 +67,7 @@ export const provider = {
configure, configure,
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => export const model: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (modelID, settings) =>
configure({ configure({
apiKey: settings.apiKey, apiKey: settings.apiKey,
baseURL: settings.baseURL, baseURL: settings.baseURL,
+3 -15
View File
@@ -1,22 +1,10 @@
import type { ProviderOptions, ReasoningEffort, TextVerbosity } from "../schema" import type { ProviderOptions } from "../schema"
import { mergeProviderOptions } from "../schema" import { mergeProviderOptions } from "../schema"
import type { OpenAIResponseIncludable, OpenAIServiceTier } from "../protocols/utils/openai-options" import type { OpenResponsesOptionsInput } from "./open-responses-options"
export type { OpenAIResponseIncludable, OpenAIServiceTier } from "../protocols/utils/openai-options" export type { OpenAIResponseIncludable, OpenAIServiceTier } from "../protocols/utils/openai-options"
export interface OpenAIOptionsInput { export type OpenAIOptionsInput = OpenResponsesOptionsInput
readonly [key: string]: unknown
readonly store?: boolean
readonly promptCacheKey?: string
readonly reasoningEffort?: ReasoningEffort
readonly reasoningSummary?: "auto"
// OpenAI Responses `include` wire field. Mirrors the official SDK's
// `ResponseIncludable[]` union exactly so AI SDK callers and direct
// native-SDK callers share one shape and no translation is required.
readonly include?: ReadonlyArray<OpenAIResponseIncludable>
readonly textVerbosity?: TextVerbosity
readonly serviceTier?: OpenAIServiceTier
}
export type OpenAIProviderOptionsInput = ProviderOptions & { export type OpenAIProviderOptionsInput = ProviderOptions & {
readonly openai?: OpenAIOptionsInput readonly openai?: OpenAIOptionsInput
+13 -6
View File
@@ -86,10 +86,15 @@ export const configure = (input: Config = {}) => {
const chatRoute = configuredRoute(OpenAIChat.route, input) const chatRoute = configuredRoute(OpenAIChat.route, input)
const modelDefaults = defaults(input) const modelDefaults = defaults(input)
const responses = (id: string | ModelID) => const responses = (id: string | ModelID) =>
responsesRoute.with(withOpenAIOptions(id, modelDefaults, { textVerbosity: true })).model({ id }) responsesRoute
.with(withOpenAIOptions(id, modelDefaults, { textVerbosity: true }))
.model<OpenAIProviderOptionsInput>({ id })
const responsesWebSocket = (id: string | ModelID) => const responsesWebSocket = (id: string | ModelID) =>
responsesWebSocketRoute.with(withOpenAIOptions(id, modelDefaults, { textVerbosity: true })).model({ id }) responsesWebSocketRoute
const chat = (id: string | ModelID) => chatRoute.with(withOpenAIOptions(id, modelDefaults)).model({ id }) .with(withOpenAIOptions(id, modelDefaults, { textVerbosity: true }))
.model<OpenAIProviderOptionsInput>({ id })
const chat = (id: string | ModelID) =>
chatRoute.with(withOpenAIOptions(id, modelDefaults)).model<OpenAIProviderOptionsInput>({ id })
const image = (modelID: string | ModelID) => const image = (modelID: string | ModelID) =>
OpenAIImages.model({ OpenAIImages.model({
id: modelID, id: modelID,
@@ -132,15 +137,17 @@ const config = (settings: Settings): Config => {
} }
} }
export const model: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => { export const model: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (modelID, settings) => {
const configured = configure(config(settings)) const configured = configure(config(settings))
if (settings.transport === undefined || settings.transport === "http") return configured.responses(modelID) if (settings.transport === undefined || settings.transport === "http") return configured.responses(modelID)
if (settings.transport === "websocket") return configured.responsesWebSocket(modelID) if (settings.transport === "websocket") return configured.responsesWebSocket(modelID)
throw new Error(`Unsupported OpenAI Responses transport: ${String(settings.transport)}`) throw new Error(`Unsupported OpenAI Responses transport: ${String(settings.transport)}`)
} }
export const chatModel: ProviderPackage.Definition<Settings>["model"] = (modelID, settings) => export const chatModel: ProviderPackage.Definition<Settings, OpenAIProviderOptionsInput>["model"] = (
configure(config(settings)).chat(modelID) modelID,
settings,
) => configure(config(settings)).chat(modelID)
export const responses = provider.responses export const responses = provider.responses
export const responsesWebSocket = provider.responsesWebSocket export const responsesWebSocket = provider.responsesWebSocket
export const chat = provider.chat export const chat = provider.chat
+104 -12
View File
@@ -4,20 +4,72 @@ import { Endpoint } from "../route/endpoint"
import { Framing } from "../route/framing" import { Framing } from "../route/framing"
import { Protocol } from "../route/protocol" import { Protocol } from "../route/protocol"
import { AuthOptions, type ProviderAuthOption } from "../route/auth-options" import { AuthOptions, type ProviderAuthOption } from "../route/auth-options"
import { ProviderID, type ModelID, type ProviderOptions } from "../schema" import { ProviderID, type CacheHint, type ModelID, type ProviderOptions } from "../schema"
import type { ProviderPackage } from "../provider-package"
import * as OpenAICompatibleProfiles from "./openai-compatible-profile" import * as OpenAICompatibleProfiles from "./openai-compatible-profile"
import * as OpenAIChat from "../protocols/openai-chat" import * as OpenAIChat from "../protocols/openai-chat"
import { newBreakpoints, ttlBucket } from "../protocols/utils/cache"
import { isRecord } from "../protocols/shared" import { isRecord } from "../protocols/shared"
export const profile = OpenAICompatibleProfiles.profiles.openrouter export const profile = OpenAICompatibleProfiles.profiles.openrouter
export const id = ProviderID.make(profile.provider) export const id = ProviderID.make(profile.provider)
const ADAPTER = "openrouter" const ADAPTER = "openrouter"
type OpenRouterString<Known extends string> = Known | (string & {})
export interface OpenRouterProviderRouting {
readonly [key: string]: unknown
readonly order?: ReadonlyArray<string>
readonly allow_fallbacks?: boolean
readonly require_parameters?: boolean
readonly data_collection?: OpenRouterString<"allow" | "deny">
readonly only?: ReadonlyArray<string>
readonly ignore?: ReadonlyArray<string>
readonly quantizations?: ReadonlyArray<string>
readonly sort?: OpenRouterString<"price" | "throughput" | "latency">
readonly max_price?: Readonly<{
prompt?: number | string
completion?: number | string
image?: number | string
audio?: number | string
request?: number | string
}>
readonly zdr?: boolean
}
export type OpenRouterPlugin =
| Readonly<{
id: "web"
max_results?: number
search_prompt?: string
engine?: OpenRouterString<"native" | "exa">
}>
| Readonly<{ id: "file-parser"; max_files?: number; pdf?: { engine?: string } }>
| Readonly<{ id: "moderation" }>
| Readonly<{ id: "response-healing" }>
| Readonly<{ id: "auto-router"; allowed_models?: ReadonlyArray<string> }>
| Readonly<{ id: string & {}; [key: string]: unknown }>
export interface OpenRouterOptions { export interface OpenRouterOptions {
readonly [key: string]: unknown readonly [key: string]: unknown
readonly usage?: boolean | Record<string, unknown> readonly debug?: Readonly<{ echo_upstream_body?: boolean }>
readonly reasoning?: Record<string, unknown> readonly models?: ReadonlyArray<string>
readonly plugins?: ReadonlyArray<OpenRouterPlugin>
readonly promptCacheKey?: string readonly promptCacheKey?: string
readonly provider?: OpenRouterProviderRouting
readonly reasoning?: Readonly<{
enabled?: boolean
exclude?: boolean
effort?: OpenRouterString<"none" | "minimal" | "low" | "medium" | "high" | "xhigh" | "max">
max_tokens?: number
}>
readonly usage?: boolean | Readonly<{ include: boolean }>
readonly user?: string
readonly web_search_options?: Readonly<{
max_results?: number
search_prompt?: string
engine?: OpenRouterString<"native" | "exa">
}>
} }
export type OpenRouterProviderOptionsInput = ProviderOptions & { export type OpenRouterProviderOptionsInput = ProviderOptions & {
@@ -30,6 +82,12 @@ export type ModelOptions = Omit<RouteDefaultsInput, "providerOptions"> &
readonly providerOptions?: OpenRouterProviderOptionsInput readonly providerOptions?: OpenRouterProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings {
readonly apiKey?: string
readonly baseURL?: string
readonly providerOptions?: OpenRouterProviderOptionsInput
}
const OpenRouterBody = Schema.StructWithRest(Schema.Struct(OpenAIChat.bodyFields), [ const OpenRouterBody = Schema.StructWithRest(Schema.Struct(OpenAIChat.bodyFields), [
Schema.Record(Schema.String, Schema.Any), Schema.Record(Schema.String, Schema.Any),
]) ])
@@ -40,7 +98,7 @@ export const protocol = Protocol.make({
body: { body: {
schema: OpenRouterBody, schema: OpenRouterBody,
from: (request) => from: (request) =>
OpenAIChat.protocol.body.from(request).pipe( OpenAIChat.fromRequest(request, { cacheControl: cacheControl() }).pipe(
Effect.map((body) => { Effect.map((body) => {
const sourceAssistants = request.messages.filter((message) => message.role === "assistant") const sourceAssistants = request.messages.filter((message) => message.role === "assistant")
let assistantIndex = 0 let assistantIndex = 0
@@ -71,16 +129,39 @@ export const protocol = Protocol.make({
stream: OpenAIChat.protocol.stream, stream: OpenAIChat.protocol.stream,
}) })
const cacheControl = () => {
const breakpoints = newBreakpoints(4)
return (cache: CacheHint | undefined) => {
if (cache === undefined || breakpoints.remaining === 0) return undefined
breakpoints.remaining -= 1
return {
type: "ephemeral" as const,
...(ttlBucket(cache.ttlSeconds) === "1h" ? { ttl: "1h" } : {}),
}
}
}
const bodyOptions = (input: unknown) => { const bodyOptions = (input: unknown) => {
const openrouter = isRecord(input) ? input : {} const openrouter = isRecord(input) ? input : {}
const { usage, models, provider, plugins, web_search_options, debug, user, reasoning, promptCacheKey, ...options } =
openrouter
return { return {
...(openrouter.usage === true ...options,
...(usage === undefined || usage === true
? { usage: { include: true } } ? { usage: { include: true } }
: isRecord(openrouter.usage) : usage === false
? { usage: openrouter.usage } ? { usage: { include: false } }
: {}), : isRecord(usage)
...(isRecord(openrouter.reasoning) ? { reasoning: openrouter.reasoning } : {}), ? { usage }
...(typeof openrouter.promptCacheKey === "string" ? { prompt_cache_key: openrouter.promptCacheKey } : {}), : {}),
...(Array.isArray(models) ? { models } : {}),
...(isRecord(provider) ? { provider } : {}),
...(Array.isArray(plugins) ? { plugins } : {}),
...(isRecord(web_search_options) ? { web_search_options } : {}),
...(isRecord(debug) ? { debug } : {}),
...(typeof user === "string" ? { user } : {}),
...(isRecord(reasoning) ? { reasoning } : {}),
...(typeof promptCacheKey === "string" ? { prompt_cache_key: promptCacheKey } : {}),
} }
} }
@@ -107,10 +188,21 @@ export const configure = (input: ModelOptions = {}) => {
const route = configuredRoute(input) const route = configuredRoute(input)
return { return {
id, id,
model: (modelID: string | ModelID) => route.model({ id: modelID }), model: (modelID: string | ModelID) => route.model<OpenRouterProviderOptionsInput>({ id: modelID }),
configure, configure,
} }
} }
export const provider = configure() export const provider = configure()
export const model = provider.model export const model: ProviderPackage.Definition<Settings, OpenRouterProviderOptionsInput>["model"] = (
modelID,
settings,
) =>
configure({
apiKey: settings.apiKey,
baseURL: settings.baseURL,
headers: settings.headers,
http: settings.body === undefined ? undefined : { body: { ...settings.body } },
limits: settings.limits,
providerOptions: settings.providerOptions,
}).model(modelID)
+51 -11
View File
@@ -1,29 +1,62 @@
import { AuthOptions, type ProviderAuthOption } from "../route/auth-options" import { AuthOptions, type ProviderAuthOption } from "../route/auth-options"
import type { RouteDefaultsInput } from "../route/client" import { Route, type RouteDefaultsInput } from "../route/client"
import { HttpOptions, ProviderID, type ModelID } from "../schema" import { Endpoint } from "../route/endpoint"
import { HttpOptions, ProviderID, type ModelID, type ProviderOptions } from "../schema"
import * as OpenAICompatibleProfiles from "./openai-compatible-profile" import * as OpenAICompatibleProfiles from "./openai-compatible-profile"
import * as OpenAICompatibleChat from "../protocols/openai-compatible-chat" import * as OpenAICompatibleChat from "../protocols/openai-compatible-chat"
import * as OpenAIChat from "../protocols/openai-chat"
import * as OpenAIResponses from "../protocols/openai-responses" import * as OpenAIResponses from "../protocols/openai-responses"
import { XAIImages } from "../protocols/xai-images" import { XAIImages } from "../protocols/xai-images"
import type { OpenAIOptionsInput } from "./openai-options"
import type { ProviderPackage } from "../provider-package"
export const id = ProviderID.make("xai") export const id = ProviderID.make("xai")
export type ModelOptions = RouteDefaultsInput & export type XAIProviderOptionsInput = ProviderOptions & {
readonly xai?: OpenAIOptionsInput
}
export type ModelOptions = Omit<RouteDefaultsInput, "providerOptions"> &
ProviderAuthOption<"optional"> & { ProviderAuthOption<"optional"> & {
readonly baseURL?: string readonly baseURL?: string
readonly providerOptions?: XAIProviderOptionsInput
} }
export interface Settings extends ProviderPackage.Settings {
readonly apiKey?: string
readonly baseURL?: string
readonly providerOptions?: XAIProviderOptionsInput
}
export type { XAIImageOptions } from "../protocols/xai-images" export type { XAIImageOptions } from "../protocols/xai-images"
export const routes = [OpenAIResponses.route, OpenAICompatibleChat.route] const responsesRoute = Route.make({
id: "openai-responses",
provider: id,
providerMetadataKey: "xai",
protocol: OpenAIResponses.protocol,
endpoint: Endpoint.path("/responses", { baseURL: OpenAICompatibleProfiles.profiles.xai.baseURL }),
transport: OpenAIResponses.httpTransport,
defaults: { providerOptions: { xai: { store: false } } },
})
const chatRoute = Route.make({
id: "openai-compatible-chat",
provider: id,
providerMetadataKey: "xai",
protocol: OpenAIChat.protocol,
endpoint: Endpoint.path("/chat/completions", { baseURL: OpenAICompatibleProfiles.profiles.xai.baseURL }),
transport: OpenAICompatibleChat.route.transport,
})
export const routes = [responsesRoute, chatRoute]
const auth = (options: ProviderAuthOption<"optional">) => AuthOptions.bearer(options, "XAI_API_KEY") const auth = (options: ProviderAuthOption<"optional">) => AuthOptions.bearer(options, "XAI_API_KEY")
const configuredResponsesRoute = (input: ModelOptions) => { const configuredResponsesRoute = (input: ModelOptions) => {
const { apiKey: _, auth: _auth, baseURL, ...rest } = input const { apiKey: _, auth: _auth, baseURL, ...rest } = input
return OpenAIResponses.route.with({ return responsesRoute.with({
...rest, ...rest,
provider: id,
endpoint: { baseURL: baseURL ?? OpenAICompatibleProfiles.profiles.xai.baseURL }, endpoint: { baseURL: baseURL ?? OpenAICompatibleProfiles.profiles.xai.baseURL },
auth: auth(input), auth: auth(input),
}) })
@@ -31,9 +64,8 @@ const configuredResponsesRoute = (input: ModelOptions) => {
const configuredChatRoute = (input: ModelOptions) => { const configuredChatRoute = (input: ModelOptions) => {
const { apiKey: _, auth: _auth, baseURL, ...rest } = input const { apiKey: _, auth: _auth, baseURL, ...rest } = input
return OpenAICompatibleChat.route.with({ return chatRoute.with({
...rest, ...rest,
provider: id,
endpoint: { baseURL: baseURL ?? OpenAICompatibleProfiles.profiles.xai.baseURL }, endpoint: { baseURL: baseURL ?? OpenAICompatibleProfiles.profiles.xai.baseURL },
auth: auth(input), auth: auth(input),
}) })
@@ -42,8 +74,8 @@ const configuredChatRoute = (input: ModelOptions) => {
export const configure = (input: ModelOptions = {}) => { export const configure = (input: ModelOptions = {}) => {
const responsesRoute = configuredResponsesRoute(input) const responsesRoute = configuredResponsesRoute(input)
const chatRoute = configuredChatRoute(input) const chatRoute = configuredChatRoute(input)
const responses = (modelID: string | ModelID) => responsesRoute.model({ id: modelID }) const responses = (modelID: string | ModelID) => responsesRoute.model<XAIProviderOptionsInput>({ id: modelID })
const chat = (modelID: string | ModelID) => chatRoute.model({ id: modelID }) const chat = (modelID: string | ModelID) => chatRoute.model<XAIProviderOptionsInput>({ id: modelID })
const image = (modelID: string | ModelID) => const image = (modelID: string | ModelID) =>
XAIImages.model({ XAIImages.model({
id: modelID, id: modelID,
@@ -63,7 +95,15 @@ export const configure = (input: ModelOptions = {}) => {
} }
export const provider = configure() export const provider = configure()
export const model = provider.model export const model: ProviderPackage.Definition<Settings, XAIProviderOptionsInput>["model"] = (modelID, settings) =>
configure({
apiKey: settings.apiKey,
baseURL: settings.baseURL,
headers: settings.headers,
http: settings.body === undefined ? undefined : { body: { ...settings.body } },
limits: settings.limits,
providerOptions: settings.providerOptions,
}).model(modelID)
export const responses = provider.responses export const responses = provider.responses
export const chat = provider.chat export const chat = provider.chat
export const image = provider.image export const image = provider.image
+37 -46
View File
@@ -5,12 +5,12 @@ import { Endpoint, type EndpointPatch } from "./endpoint"
import { RequestExecutor } from "./executor" import { RequestExecutor } from "./executor"
import { Framing } from "./framing" import { Framing } from "./framing"
import { HttpTransport } from "./transport" import { HttpTransport } from "./transport"
import type { Transport, TransportRuntime } from "./transport" import type { HttpRequestTransform, Transport, TransportRuntime } from "./transport"
import { WebSocketExecutor } from "./transport" import { WebSocketExecutor } from "./transport"
import type { Protocol } from "./protocol" import type { Protocol } from "./protocol"
import { applyCachePolicy } from "../cache-policy" import { applyCachePolicy } from "../cache-policy"
import * as ProviderShared from "../protocols/shared" import * as ProviderShared from "../protocols/shared"
import type { LLMError, PreparedRequestOf, ProtocolID, ProviderOptions } from "../schema" import type { LLMError, ProtocolID, ProviderOptions } from "../schema"
import { import {
GenerationOptions, GenerationOptions,
HttpOptions, HttpOptions,
@@ -20,7 +20,6 @@ import {
ModelLimits, ModelLimits,
LLMError as LLMErrorClass, LLMError as LLMErrorClass,
LLMEvent, LLMEvent,
PreparedRequest,
ProviderID, ProviderID,
mergeGenerationOptions, mergeGenerationOptions,
mergeHttpOptions, mergeHttpOptions,
@@ -46,8 +45,12 @@ export interface Route<Body, Prepared = unknown> {
readonly defaults: RouteDefaults readonly defaults: RouteDefaults
readonly body: RouteBody<Body> readonly body: RouteBody<Body>
readonly with: (patch: RoutePatch<Body, Prepared>) => Route<Body, Prepared> readonly with: (patch: RoutePatch<Body, Prepared>) => Route<Body, Prepared>
readonly model: (input: RouteMappedModelInput) => Model readonly model: <Options extends ProviderOptions = ProviderOptions>(input: RouteMappedModelInput) => Model<Options>
readonly prepareTransport: (body: Body, request: LLMRequest) => Effect.Effect<Prepared, LLMError> readonly prepareTransport: (
body: Body,
request: LLMRequest,
options?: StreamOptions,
) => Effect.Effect<Prepared, LLMError>
readonly streamPrepared: ( readonly streamPrepared: (
prepared: Prepared, prepared: Prepared,
request: LLMRequest, request: LLMRequest,
@@ -93,12 +96,12 @@ export interface RoutePatch<Body, Prepared> extends RouteDefaultsInput {
type RouteMappedModelInput = RouteModelInput | RouteRoutedModelInput type RouteMappedModelInput = RouteModelInput | RouteRoutedModelInput
const makeRouteModel = (route: AnyRoute, mapped: RouteMappedModelInput) => { const makeRouteModel = <Options extends ProviderOptions = ProviderOptions>(route: AnyRoute, mapped: RouteMappedModelInput) => {
const provider = route.provider ?? ("provider" in mapped ? mapped.provider : undefined) const provider = route.provider ?? ("provider" in mapped ? mapped.provider : undefined)
if (!provider) throw new Error(`Route.model(${route.id}) requires a provider`) if (!provider) throw new Error(`Route.model(${route.id}) requires a provider`)
if (!endpointBaseURL(route.endpoint)) if (!endpointBaseURL(route.endpoint))
throw new Error(`Route.model(${route.id}) requires an endpoint baseURL — configure it on the route first`) throw new Error(`Route.model(${route.id}) requires an endpoint baseURL — configure it on the route first`)
return Model.make({ return Model.make<Options>({
...mapped, ...mapped,
provider, provider,
route, route,
@@ -142,27 +145,20 @@ export const httpOptions = (input: HttpOptionsInput | undefined) => {
} }
export interface Interface { export interface Interface {
/**
* Compile a request through protocol body construction, validation, and HTTP
* preparation without sending it. Returns the prepared request including the
* provider-native body.
*
* Pass a `Body` type argument to statically expose the route's body
* shape (e.g. `prepare<OpenAIChatBody>(...)`) — the runtime body is
* identical, so this is a type-level assertion the caller makes about which
* route the request will resolve to.
*/
readonly prepare: <Body = unknown>(request: LLMRequest) => Effect.Effect<PreparedRequestOf<Body>, LLMError>
readonly stream: StreamMethod readonly stream: StreamMethod
readonly generate: GenerateMethod readonly generate: GenerateMethod
} }
export interface StreamOptions {
readonly transform?: HttpRequestTransform
}
export interface StreamMethod { export interface StreamMethod {
(request: LLMRequest): Stream.Stream<LLMEvent, LLMError> (request: LLMRequest, options?: StreamOptions): Stream.Stream<LLMEvent, LLMError>
} }
export interface GenerateMethod { export interface GenerateMethod {
(request: LLMRequest): Effect.Effect<LLMResponse, LLMError> (request: LLMRequest, options?: StreamOptions): Effect.Effect<LLMResponse, LLMError>
} }
export class Service extends Context.Service<Service, Interface>()("@opencode/LLMClient") {} export class Service extends Context.Service<Service, Interface>()("@opencode/LLMClient") {}
@@ -296,8 +292,9 @@ function makeFromTransport<Body, Prepared, Frame, Event, State>(
defaults: mergeRouteDefaults(route.defaults, defaults), defaults: mergeRouteDefaults(route.defaults, defaults),
}) })
}, },
model: (input) => makeRouteModel(route, input), model: <Options extends ProviderOptions = ProviderOptions>(input: RouteMappedModelInput) =>
prepareTransport: (body, request) => makeRouteModel<Options>(route, input),
prepareTransport: (body, request, options) =>
routeInput.transport.prepare({ routeInput.transport.prepare({
body, body,
request, request,
@@ -305,6 +302,7 @@ function makeFromTransport<Body, Prepared, Frame, Event, State>(
auth: routeInput.auth ?? Auth.none, auth: routeInput.auth ?? Auth.none,
encodeBody, encodeBody,
headers: routeInput.headers, headers: routeInput.headers,
transform: options?.transform,
}), }),
streamPrepared: (prepared: Prepared, request: LLMRequest, runtime: TransportRuntime) => { streamPrepared: (prepared: Prepared, request: LLMRequest, runtime: TransportRuntime) => {
const route = `${request.model.provider}/${request.model.route.id}` const route = `${request.model.provider}/${request.model.route.id}`
@@ -370,17 +368,14 @@ export function make<Body, Prepared, Frame, Event, State>(
}) })
} }
// `compile` is the important boundary: it turns a common `LLMRequest` into a const compile = Effect.fn("LLM.compile")(function* (request: LLMRequest, options?: StreamOptions) {
// validated provider body plus transport-private prepared data, but does not
// execute transport.
const compile = Effect.fn("LLM.compile")(function* (request: LLMRequest) {
const resolved = applyCachePolicy(resolveRequestOptions(request)) const resolved = applyCachePolicy(resolveRequestOptions(request))
const route = resolved.model.route const route = resolved.model.route
const body = yield* route.body const body = yield* route.body
.from(resolved) .from(resolved)
.pipe(Effect.flatMap(ProviderShared.validateWith(Schema.decodeUnknownEffect(route.body.schema)))) .pipe(Effect.flatMap(ProviderShared.validateWith(Schema.decodeUnknownEffect(route.body.schema))))
const prepared = yield* route.prepareTransport(body, resolved) const prepared = yield* route.prepareTransport(body, resolved, options)
return { return {
request: resolved, request: resolved,
@@ -390,30 +385,30 @@ const compile = Effect.fn("LLM.compile")(function* (request: LLMRequest) {
} }
}) })
const prepareWith = Effect.fn("LLMClient.prepare")(function* (request: LLMRequest) { /** @internal Test-only projection of the execution compiler; not exported from package barrels. */
export const compileRequest = Effect.fn("LLM.compileRequest")(function* (request: LLMRequest) {
const compiled = yield* compile(request) const compiled = yield* compile(request)
return {
return new PreparedRequest({
id: compiled.request.id ?? "request", id: compiled.request.id ?? "request",
route: compiled.route.id, route: compiled.route.id,
protocol: compiled.route.protocol, protocol: compiled.route.protocol,
model: compiled.request.model, model: compiled.request.model,
body: compiled.body, body: compiled.body,
metadata: { transport: compiled.route.transport.id }, metadata: { transport: compiled.route.transport.id },
}) }
}) })
const streamRequestWith = (runtime: TransportRuntime) => (request: LLMRequest) => const streamRequestWith = (runtime: TransportRuntime) => (request: LLMRequest, options?: StreamOptions) =>
Stream.unwrap( Stream.unwrap(
Effect.gen(function* () { Effect.gen(function* () {
const compiled = yield* compile(request) const compiled = yield* compile(request, options)
return compiled.route.streamPrepared(compiled.prepared, compiled.request, runtime) return compiled.route.streamPrepared(compiled.prepared, compiled.request, runtime)
}), }),
) )
const generateWith = (stream: Interface["stream"]) => const generateWith = (stream: Interface["stream"]) =>
Effect.fn("LLM.generate")(function* (request: LLMRequest) { Effect.fn("LLM.generate")(function* (request: LLMRequest, options?: StreamOptions) {
const state = yield* stream(request).pipe(Stream.runFold(LLMResponse.empty, LLMResponse.reduce)) const state = yield* stream(request, options).pipe(Stream.runFold(LLMResponse.empty, LLMResponse.reduce))
const response = LLMResponse.complete(state) const response = LLMResponse.complete(state)
if (response) return response if (response) return response
return yield* ProviderShared.eventError( return yield* ProviderShared.eventError(
@@ -422,27 +417,24 @@ const generateWith = (stream: Interface["stream"]) =>
) )
}) })
export const prepare = <Body = unknown>(request: LLMRequest) => export function stream(request: LLMRequest, options?: StreamOptions): Stream.Stream<LLMEvent, LLMError> {
prepareWith(request) as Effect.Effect<PreparedRequestOf<Body>, LLMError>
export function stream(request: LLMRequest): Stream.Stream<LLMEvent, LLMError> {
return Stream.unwrap( return Stream.unwrap(
Effect.gen(function* () { Effect.gen(function* () {
return (yield* Service).stream(request) return (yield* Service).stream(request, options)
}), }),
) as Stream.Stream<LLMEvent, LLMError> ) as Stream.Stream<LLMEvent, LLMError>
} }
export function generate(request: LLMRequest): Effect.Effect<LLMResponse, LLMError> { export function generate(request: LLMRequest, options?: StreamOptions): Effect.Effect<LLMResponse, LLMError> {
return Effect.gen(function* () { return Effect.gen(function* () {
return yield* (yield* Service).generate(request) return yield* (yield* Service).generate(request, options)
}) as Effect.Effect<LLMResponse, LLMError> }) as Effect.Effect<LLMResponse, LLMError>
} }
export const streamRequest = (request: LLMRequest) => export const streamRequest = (request: LLMRequest, options?: StreamOptions) =>
Stream.unwrap( Stream.unwrap(
Effect.gen(function* () { Effect.gen(function* () {
return (yield* Service).stream(request) return (yield* Service).stream(request, options)
}), }),
) )
@@ -453,7 +445,7 @@ export const layer: Layer.Layer<Service, never, RequestExecutor.Service> = Layer
http: yield* RequestExecutor.Service, http: yield* RequestExecutor.Service,
webSocket: Option.getOrUndefined(yield* Effect.serviceOption(WebSocketExecutor.Service)), webSocket: Option.getOrUndefined(yield* Effect.serviceOption(WebSocketExecutor.Service)),
}) })
return Service.of({ prepare: prepareWith as Interface["prepare"], stream, generate: generateWith(stream) }) return Service.of({ stream, generate: generateWith(stream) })
}), }),
) )
@@ -462,7 +454,6 @@ export const Route = { make } as const
export const LLMClient = { export const LLMClient = {
Service, Service,
layer, layer,
prepare,
stream, stream,
generate, generate,
} as const } as const
+2 -1
View File
@@ -8,6 +8,7 @@ export type {
AnyRoute, AnyRoute,
Interface as LLMClientShape, Interface as LLMClientShape,
Service as LLMClientService, Service as LLMClientService,
StreamOptions,
} from "./client" } from "./client"
export * from "./executor" export * from "./executor"
export { Auth } from "./auth" export { Auth } from "./auth"
@@ -22,4 +23,4 @@ export type { ApiKeyMode, AuthOverride, ProviderAuthOption } from "./auth-option
export type { Definition as EndpointFn, EndpointInput } from "./endpoint" export type { Definition as EndpointFn, EndpointInput } from "./endpoint"
export type { Definition as FramingDef } from "./framing" export type { Definition as FramingDef } from "./framing"
export type { Protocol as ProtocolDef } from "./protocol" export type { Protocol as ProtocolDef } from "./protocol"
export type { Transport as TransportDef, TransportRuntime } from "./transport" export type { HttpRequest, HttpRequestTransform, Transport as TransportDef, TransportRuntime } from "./transport"
+2 -1
View File
@@ -12,7 +12,8 @@ import type { LLMError, LLMEvent, LLMRequest, ProtocolID } from "../schema"
* Examples: * Examples:
* *
* - `OpenAIChat.protocol` — chat completions style * - `OpenAIChat.protocol` — chat completions style
* - `OpenAIResponses.protocol` — responses API * - `OpenResponses.protocol` — provider-neutral Responses API baseline
* - `OpenAIResponses.protocol` — OpenAI extensions to that baseline
* - `AnthropicMessages.protocol` — messages API with content blocks * - `AnthropicMessages.protocol` — messages API with content blocks
* - `Gemini.protocol` — generateContent * - `Gemini.protocol` — generateContent
* - `BedrockConverse.protocol` — Converse with binary event-stream framing * - `BedrockConverse.protocol` — Converse with binary event-stream framing
+12 -55
View File
@@ -28,57 +28,9 @@ const applyQuery = (url: string, query: Record<string, string> | undefined) => {
return next.toString() return next.toString()
} }
const PROTOCOL_BODY_OVERLAY_DENYLIST = new Set([
"anthropic_version",
"content",
"contents",
"frequencyPenalty",
"frequency_penalty",
"generationConfig",
"inferenceConfig",
"input",
"maxTokens",
"max_tokens",
"messages",
"model",
"presencePenalty",
"presence_penalty",
"responseFormat",
"response_format",
"seed",
"stop",
"stopSequences",
"stop_sequences",
"stream",
"streamOptions",
"stream_options",
"system",
"systemInstruction",
"system_instruction",
"temperature",
"thinking",
"toolChoice",
"toolConfig",
"tool_choice",
"tool_config",
"tools",
"topK",
"topP",
"top_k",
"top_p",
])
const forbiddenBodyOverlayKeys = (body: Record<string, unknown>) =>
Object.keys(body).filter((key) => PROTOCOL_BODY_OVERLAY_DENYLIST.has(key))
const bodyWithOverlay = <Body>(body: Body, request: LLMRequest, encodeBody: (body: Body) => string) => const bodyWithOverlay = <Body>(body: Body, request: LLMRequest, encodeBody: (body: Body) => string) =>
Effect.gen(function* () { Effect.gen(function* () {
if (request.http?.body === undefined) return { jsonBody: body, bodyText: encodeBody(body) } if (request.http?.body === undefined) return { jsonBody: body, bodyText: encodeBody(body) }
const forbiddenKeys = forbiddenBodyOverlayKeys(request.http.body)
if (forbiddenKeys.length > 0)
return yield* ProviderShared.invalidRequest(
`http.body cannot overlay protocol-owned field(s): ${forbiddenKeys.join(", ")}`,
)
if (ProviderShared.isRecord(body)) { if (ProviderShared.isRecord(body)) {
const overlaid = mergeJsonRecords(body, request.http.body) ?? {} const overlaid = mergeJsonRecords(body, request.http.body) ?? {}
return { jsonBody: overlaid, bodyText: ProviderShared.encodeJson(overlaid) } return { jsonBody: overlaid, bodyText: ProviderShared.encodeJson(overlaid) }
@@ -120,14 +72,19 @@ export const httpJson = <Body, Frame>(input: HttpJsonInput<Body, Frame>): HttpJs
id: "http-json", id: "http-json",
with: (patch) => httpJson({ ...input, ...patch }), with: (patch) => httpJson({ ...input, ...patch }),
prepare: (prepareInput) => prepare: (prepareInput) =>
jsonRequestParts({ Effect.gen(function* () {
...prepareInput, const parts = yield* jsonRequestParts({ ...prepareInput })
}).pipe( const request = { url: parts.url, method: "POST", headers: { ...parts.headers }, body: parts.bodyText }
Effect.map((parts) => ({ yield* (prepareInput.transform?.(request) ?? Effect.void)
request: ProviderShared.jsonPost({ url: parts.url, body: parts.bodyText, headers: parts.headers }), return {
request: ProviderShared.jsonPost({
url: request.url,
body: request.body ?? "",
headers: Headers.fromInput(request.headers),
}),
framing: input.framing, framing: input.framing,
})), }
), }),
frames: (prepared, request, runtime) => frames: (prepared, request, runtime) =>
Stream.unwrap( Stream.unwrap(
runtime.http runtime.http
+10
View File
@@ -10,6 +10,15 @@ export interface TransportRuntime {
readonly webSocket?: WebSocketExecutorInterface readonly webSocket?: WebSocketExecutorInterface
} }
export interface HttpRequest {
url: string
readonly method: string
headers: Record<string, string>
body: string | undefined
}
export type HttpRequestTransform = (request: HttpRequest) => Effect.Effect<void>
export interface Transport<Body, Prepared, Frame> { export interface Transport<Body, Prepared, Frame> {
readonly id: string readonly id: string
readonly prepare: (input: TransportPrepareInput<Body>) => Effect.Effect<Prepared, LLMError> readonly prepare: (input: TransportPrepareInput<Body>) => Effect.Effect<Prepared, LLMError>
@@ -27,6 +36,7 @@ export interface TransportPrepareInput<Body> {
readonly auth: Auth.Definition readonly auth: Auth.Definition
readonly encodeBody: (body: Body) => string readonly encodeBody: (body: Body) => string
readonly headers?: (input: { readonly request: LLMRequest }) => Record<string, string> readonly headers?: (input: { readonly request: LLMRequest }) => Record<string, string>
readonly transform?: HttpRequestTransform
} }
export * as HttpTransport from "./http" export * as HttpTransport from "./http"
+2 -5
View File
@@ -1,4 +1,5 @@
import { Schema } from "effect" import { Schema } from "effect"
import { Tool } from "@opencode-ai/schema/tool"
import { ModelID, ProviderID, ProviderMetadata, RouteID } from "./ids" import { ModelID, ProviderID, ProviderMetadata, RouteID } from "./ids"
export const ProviderFailureClassification = Schema.Literal("context-overflow") export const ProviderFailureClassification = Schema.Literal("context-overflow")
@@ -152,8 +153,4 @@ export class LLMError extends Schema.TaggedErrorClass<LLMError>()("LLM.Error", {
* Anything thrown or yielded by a handler that is not a `ToolFailure` is * Anything thrown or yielded by a handler that is not a `ToolFailure` is
* treated as a defect and fails the stream. * treated as a defect and fails the stream.
*/ */
export class ToolFailure extends Schema.TaggedErrorClass<ToolFailure>()("LLM.ToolFailure", { export class ToolFailure extends Tool.Error {}
message: Schema.String,
error: Schema.optional(Schema.Defect()),
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
}) {}
+15 -33
View File
@@ -1,6 +1,5 @@
import { Schema } from "effect" import { Schema } from "effect"
import { ContentBlockID, FinishReason, ProtocolID, ProviderMetadata, RouteID, ToolCallID } from "./ids" import { ContentBlockID, FinishReason, ProviderMetadata, ToolCallID } from "./ids"
import { ModelSchema } from "./options"
import { Message, ToolCallPart, ToolOutput, ToolResultPart, ToolResultValue, type ContentPart } from "./messages" import { Message, ToolCallPart, ToolOutput, ToolResultPart, ToolResultValue, type ContentPart } from "./messages"
import { ProviderFailureClassification } from "./errors" import { ProviderFailureClassification } from "./errors"
@@ -40,9 +39,9 @@ import { ProviderFailureClassification } from "./errors"
* - Anthropic and Bedrock report the input breakdown natively: Anthropic's * - Anthropic and Bedrock report the input breakdown natively: Anthropic's
* `input_tokens` and Bedrock's `inputTokens` are non-cached only. Their * `input_tokens` and Bedrock's `inputTokens` are non-cached only. Their
* mappers sum the breakdown to derive the inclusive `inputTokens`. * mappers sum the breakdown to derive the inclusive `inputTokens`.
* Anthropic does *not* break extended-thinking out of `output_tokens`, so * Anthropic's `outputTokens` includes extended thinking. Newer responses
* `reasoningTokens` is `undefined` and `outputTokens` carries the * expose that subset as `output_tokens_details.thinking_tokens`, which maps
* combined total — a documented limitation of the Anthropic API. * to `reasoningTokens`; older responses leave it undefined.
* *
* `providerMetadata` always carries the provider's raw usage payload — * `providerMetadata` always carries the provider's raw usage payload —
* keyed by provider name (`{ openai: ... }`, `{ anthropic: ... }`, etc.) * keyed by provider name (`{ openai: ... }`, `{ anthropic: ... }`, etc.)
@@ -191,10 +190,16 @@ export const ToolError = Schema.Struct({
}).annotate({ identifier: "LLM.Event.ToolError" }) }).annotate({ identifier: "LLM.Event.ToolError" })
export type ToolError = Schema.Schema.Type<typeof ToolError> export type ToolError = Schema.Schema.Type<typeof ToolError>
export const FinishReasonDetails = Schema.Struct({
normalized: FinishReason,
raw: Schema.optional(Schema.String),
}).annotate({ identifier: "LLM.FinishReasonDetails" })
export type FinishReasonDetails = Schema.Schema.Type<typeof FinishReasonDetails>
export const StepFinish = Schema.Struct({ export const StepFinish = Schema.Struct({
type: Schema.tag("step-finish"), type: Schema.tag("step-finish"),
index: Schema.Number, index: Schema.Number,
reason: FinishReason, reason: FinishReasonDetails,
usage: Schema.optional(Usage), usage: Schema.optional(Usage),
providerMetadata: Schema.optional(ProviderMetadata), providerMetadata: Schema.optional(ProviderMetadata),
}).annotate({ identifier: "LLM.Event.StepFinish" }) }).annotate({ identifier: "LLM.Event.StepFinish" })
@@ -202,7 +207,7 @@ export type StepFinish = Schema.Schema.Type<typeof StepFinish>
export const Finish = Schema.Struct({ export const Finish = Schema.Struct({
type: Schema.tag("finish"), type: Schema.tag("finish"),
reason: FinishReason, reason: FinishReasonDetails,
usage: Schema.optional(Usage), usage: Schema.optional(Usage),
providerMetadata: Schema.optional(ProviderMetadata), providerMetadata: Schema.optional(ProviderMetadata),
}).annotate({ identifier: "LLM.Event.Finish" }) }).annotate({ identifier: "LLM.Event.Finish" })
@@ -308,29 +313,6 @@ export const LLMEvent = Object.assign(llmEventTagged, {
}) })
export type LLMEvent = Schema.Schema.Type<typeof llmEventTagged> export type LLMEvent = Schema.Schema.Type<typeof llmEventTagged>
export class PreparedRequest extends Schema.Class<PreparedRequest>("LLM.PreparedRequest")({
id: Schema.String,
route: RouteID,
protocol: ProtocolID,
model: ModelSchema,
body: Schema.Unknown,
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
}) {}
/**
* A `PreparedRequest` whose `body` is typed as `Body`. Use with the generic
* on `LLMClient.prepare<Body>(...)` when the caller knows which route their
* request will resolve to and wants its native shape statically exposed
* (debug UIs, request previews, plan rendering).
*
* The runtime body is identical — the route still emits `body: unknown` — so
* this is a type-level assertion the caller makes about what they expect to
* find. The prepare runtime does not validate the assertion.
*/
export type PreparedRequestOf<Body> = Omit<PreparedRequest, "body"> & {
readonly body: Body
}
const responseText = (events: ReadonlyArray<LLMEvent>) => const responseText = (events: ReadonlyArray<LLMEvent>) =>
events events
.filter(LLMEvent.is.textDelta) .filter(LLMEvent.is.textDelta)
@@ -365,7 +347,7 @@ interface ResponseState {
readonly events: ReadonlyArray<LLMEvent> readonly events: ReadonlyArray<LLMEvent>
readonly message: Message readonly message: Message
readonly usage?: Usage readonly usage?: Usage
readonly finishReason?: FinishReason readonly finishReason?: FinishReasonDetails
readonly textParts: Readonly<Record<string, ContentAssembly>> readonly textParts: Readonly<Record<string, ContentAssembly>>
readonly reasoningParts: Readonly<Record<string, ContentAssembly>> readonly reasoningParts: Readonly<Record<string, ContentAssembly>>
readonly toolInputs: Readonly<Record<string, ToolInputAssembly>> readonly toolInputs: Readonly<Record<string, ToolInputAssembly>>
@@ -393,7 +375,7 @@ const appendEvent = (state: ResponseState, event: LLMEvent): ResponseState => {
return { return {
...state, ...state,
events, events,
finishReason: state.finishReason ?? "error", finishReason: state.finishReason ?? { normalized: "error" },
} }
} }
return { return {
@@ -580,7 +562,7 @@ export class LLMResponse extends Schema.Class<LLMResponse>("LLM.Response")({
message: Message, message: Message,
events: Schema.Array(LLMEvent), events: Schema.Array(LLMEvent),
usage: Schema.optional(Usage), usage: Schema.optional(Usage),
finishReason: FinishReason, finishReason: FinishReasonDetails,
}) { }) {
/** Concatenated assistant text assembled from streamed `text-delta` events. */ /** Concatenated assistant text assembled from streamed `text-delta` events. */
get text() { get text() {
+7 -16
View File
@@ -1,5 +1,5 @@
import { Schema } from "effect" import { Schema } from "effect"
import { ToolContent, ToolFileContent, ToolTextContent } from "@opencode-ai/schema/llm" import { Tool } from "@opencode-ai/schema/tool"
import { JsonSchema, MessageRole, ProviderMetadata } from "./ids" import { JsonSchema, MessageRole, ProviderMetadata } from "./ids"
import { CacheHint, CachePolicy, GenerationOptions, HttpOptions, ModelSchema, ProviderOptions } from "./options" import { CacheHint, CachePolicy, GenerationOptions, HttpOptions, ModelSchema, ProviderOptions } from "./options"
import { isRecord } from "../utils/record" import { isRecord } from "../utils/record"
@@ -40,8 +40,6 @@ export const MediaPart = Schema.Struct({
}).annotate({ identifier: "LLM.Content.Media" }) }).annotate({ identifier: "LLM.Content.Media" })
export type MediaPart = Schema.Schema.Type<typeof MediaPart> export type MediaPart = Schema.Schema.Type<typeof MediaPart>
export { ToolContent, ToolFileContent, ToolTextContent }
const isToolResultValue = (value: unknown): value is ToolResultValue => const isToolResultValue = (value: unknown): value is ToolResultValue =>
isRecord(value) && isRecord(value) &&
(value.type === "text" || value.type === "json" || value.type === "error" || value.type === "content") && (value.type === "text" || value.type === "json" || value.type === "error" || value.type === "content") &&
@@ -63,7 +61,7 @@ export const ToolResultValue = Object.assign(
}), }),
Schema.Struct({ Schema.Struct({
type: Schema.Literal("content"), type: Schema.Literal("content"),
value: Schema.Array(ToolContent), value: Schema.Array(Tool.Content),
}), }),
]).annotate({ identifier: "LLM.ToolResult" }), ]).annotate({ identifier: "LLM.ToolResult" }),
{ {
@@ -79,16 +77,16 @@ export type ToolResultValue = Schema.Schema.Type<typeof ToolResultValue>
export interface ToolOutput { export interface ToolOutput {
readonly structured: unknown readonly structured: unknown
readonly content: ReadonlyArray<ToolContent> readonly content: ReadonlyArray<Tool.Content>
} }
export const ToolOutput = Object.assign( export const ToolOutput = Object.assign(
Schema.Struct({ Schema.Struct({
structured: Schema.Unknown, structured: Schema.Unknown,
content: Schema.Array(ToolContent), content: Schema.Array(Tool.Content),
}).annotate({ identifier: "LLM.ToolOutput" }), }).annotate({ identifier: "LLM.ToolOutput" }),
{ {
make: (structured: unknown, content: ReadonlyArray<ToolContent> = []): ToolOutput => ({ structured, content }), make: (structured: unknown, content: ReadonlyArray<Tool.Content> = []): ToolOutput => ({ structured, content }),
fromResultValue: (result: ToolResultValue): ToolOutput | undefined => { fromResultValue: (result: ToolResultValue): ToolOutput | undefined => {
switch (result.type) { switch (result.type) {
case "json": case "json":
@@ -126,6 +124,7 @@ export const ToolCallPart = Object.assign(
name: Schema.String, name: Schema.String,
input: Schema.Unknown, input: Schema.Unknown,
providerExecuted: Schema.optional(Schema.Boolean), providerExecuted: Schema.optional(Schema.Boolean),
cache: Schema.optional(CacheHint),
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)), metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
providerMetadata: Schema.optional(ProviderMetadata), providerMetadata: Schema.optional(ProviderMetadata),
}).annotate({ identifier: "LLM.Content.ToolCall" }), }).annotate({ identifier: "LLM.Content.ToolCall" }),
@@ -170,6 +169,7 @@ export const ReasoningPart = Schema.Struct({
type: Schema.Literal("reasoning"), type: Schema.Literal("reasoning"),
text: Schema.String, text: Schema.String,
encrypted: Schema.optional(Schema.String), encrypted: Schema.optional(Schema.String),
cache: Schema.optional(CacheHint),
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)), metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
providerMetadata: Schema.optional(ProviderMetadata), providerMetadata: Schema.optional(ProviderMetadata),
}).annotate({ identifier: "LLM.Content.Reasoning" }) }).annotate({ identifier: "LLM.Content.Reasoning" })
@@ -261,13 +261,6 @@ export namespace ToolChoice {
} }
} }
export const ResponseFormat = Schema.Union([
Schema.Struct({ type: Schema.Literal("text") }),
Schema.Struct({ type: Schema.Literal("json"), schema: JsonSchema }),
Schema.Struct({ type: Schema.Literal("tool"), tool: ToolDefinition }),
]).pipe(Schema.toTaggedUnion("type"))
export type ResponseFormat = Schema.Schema.Type<typeof ResponseFormat>
export class LLMRequest extends Schema.Class<LLMRequest>("LLM.Request")({ export class LLMRequest extends Schema.Class<LLMRequest>("LLM.Request")({
id: Schema.optional(Schema.String), id: Schema.optional(Schema.String),
model: ModelSchema, model: ModelSchema,
@@ -278,7 +271,6 @@ export class LLMRequest extends Schema.Class<LLMRequest>("LLM.Request")({
generation: Schema.optional(GenerationOptions), generation: Schema.optional(GenerationOptions),
providerOptions: Schema.optional(ProviderOptions), providerOptions: Schema.optional(ProviderOptions),
http: Schema.optional(HttpOptions), http: Schema.optional(HttpOptions),
responseFormat: Schema.optional(ResponseFormat),
cache: Schema.optional(CachePolicy), cache: Schema.optional(CachePolicy),
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)), metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
}) {} }) {}
@@ -296,7 +288,6 @@ export namespace LLMRequest {
generation: request.generation, generation: request.generation,
providerOptions: request.providerOptions, providerOptions: request.providerOptions,
http: request.http, http: request.http,
responseFormat: request.responseFormat,
cache: request.cache, cache: request.cache,
metadata: request.metadata, metadata: request.metadata,
}) })
+19 -11
View File
@@ -123,6 +123,7 @@ export const mergeGenerationOptions = (...items: ReadonlyArray<GenerationOptions
export class ModelLimits extends Schema.Class<ModelLimits>("LLM.ModelLimits")({ export class ModelLimits extends Schema.Class<ModelLimits>("LLM.ModelLimits")({
context: Schema.optional(Schema.Number), context: Schema.optional(Schema.Number),
input: Schema.optional(Schema.Number),
output: Schema.optional(Schema.Number), output: Schema.optional(Schema.Number),
}) {} }) {}
@@ -166,9 +167,13 @@ export namespace ModelDefaults {
export const ModelToolSchemaCompatibility = Schema.Literals(["gemini", "moonshot"]) export const ModelToolSchemaCompatibility = Schema.Literals(["gemini", "moonshot"])
export type ModelToolSchemaCompatibility = Schema.Schema.Type<typeof ModelToolSchemaCompatibility> export type ModelToolSchemaCompatibility = Schema.Schema.Type<typeof ModelToolSchemaCompatibility>
export const ModelMaxTokensFieldCompatibility = Schema.Literals(["max_completion_tokens", "max_tokens"])
export type ModelMaxTokensFieldCompatibility = Schema.Schema.Type<typeof ModelMaxTokensFieldCompatibility>
export class ModelCompatibility extends Schema.Class<ModelCompatibility>("LLM.ModelCompatibility")({ export class ModelCompatibility extends Schema.Class<ModelCompatibility>("LLM.ModelCompatibility")({
toolSchema: Schema.optional(ModelToolSchemaCompatibility), toolSchema: Schema.optional(ModelToolSchemaCompatibility),
reasoningField: Schema.optional(Schema.String), reasoningField: Schema.optional(Schema.String),
maxTokensField: Schema.optional(ModelMaxTokensFieldCompatibility),
}) {} }) {}
export namespace ModelCompatibility { export namespace ModelCompatibility {
@@ -178,7 +183,8 @@ export namespace ModelCompatibility {
export const make = (input: Input) => (input instanceof ModelCompatibility ? input : new ModelCompatibility(input)) export const make = (input: Input) => (input instanceof ModelCompatibility ? input : new ModelCompatibility(input))
} }
export class Model { export class Model<Options extends ProviderOptions = ProviderOptions> {
declare protected readonly _ProviderOptions: Options
readonly id: ModelID readonly id: ModelID
readonly provider: ProviderID readonly provider: ProviderID
readonly route: AnyRoute readonly route: AnyRoute
@@ -193,8 +199,8 @@ export class Model {
this.compatibility = input.compatibility this.compatibility = input.compatibility
} }
static make(input: Model.Input) { static make<Options extends ProviderOptions = ProviderOptions>(input: Model.Input) {
return new Model({ return new Model<Options>({
id: ModelID.make(input.id), id: ModelID.make(input.id),
provider: ProviderID.make(input.provider), provider: ProviderID.make(input.provider),
route: input.route, route: input.route,
@@ -203,7 +209,7 @@ export class Model {
}) })
} }
static input(model: Model): Model.ConstructorInput { static input<Options extends ProviderOptions>(model: Model<Options>): Model.ConstructorInput {
return { return {
id: model.id, id: model.id,
provider: model.provider, provider: model.provider,
@@ -213,9 +219,9 @@ export class Model {
} }
} }
static update(model: Model, patch: Partial<Model.Input>) { static update<Options extends ProviderOptions>(model: Model<Options>, patch: Partial<Model.Input>) {
if (Object.keys(patch).length === 0) return model if (Object.keys(patch).length === 0) return model
return Model.make({ return Model.make<Options>({
...Model.input(model), ...Model.input(model),
...patch, ...patch,
}) })
@@ -241,6 +247,8 @@ export namespace Model {
export type ModelInput = Model.Input export type ModelInput = Model.Input
export type ModelProviderOptions<SelectedModel> = SelectedModel extends Model<infer Options> ? Options : never
export const ModelSchema = Schema.declare((value): value is Model => value instanceof Model, { expected: "LLM.Model" }) export const ModelSchema = Schema.declare((value): value is Model => value instanceof Model, { expected: "LLM.Model" })
export class CacheHint extends Schema.Class<CacheHint>("LLM.CacheHint")({ export class CacheHint extends Schema.Class<CacheHint>("LLM.CacheHint")({
@@ -251,11 +259,11 @@ export class CacheHint extends Schema.Class<CacheHint>("LLM.CacheHint")({
// Auto-placement policy for prompt caching. The protocol-neutral lowering step // Auto-placement policy for prompt caching. The protocol-neutral lowering step
// reads this and injects `CacheHint`s at the configured boundaries; the // reads this and injects `CacheHint`s at the configured boundaries; the
// per-protocol body builders then translate those hints into wire markers as // per-protocol body builders then translate those hints into wire markers as
// usual. `"auto"` is the recommended default for agent loops — it places one // usual. `"auto"` is the recommended default for agent loops — it places
// breakpoint at the last tool definition, one at the last system part, and one // breakpoints at the last tool definition, the first and last distinct system
// at the latest user message. The combination of provider invalidation // parts, and the conversation tail. The rolling message breakpoint keeps a
// hierarchy (tools → system → messages) and Anthropic/Bedrock's 20-block // prior cache entry within Anthropic/Bedrock's 20-block lookback during long
// lookback means three trailing breakpoints reliably cover the static prefix. // tool loops.
// //
// Pass `"none"` to opt out entirely (the legacy behavior). Pass the granular // Pass `"none"` to opt out entirely (the legacy behavior). Pass the granular
// object form to override individual choices. // object form to override individual choices.
+156
View File
@@ -0,0 +1,156 @@
export * as TestLLM from "./testing"
import { LLMClient, type Interface as LLMClientShape } from "./route/client"
import {
LLMEvent,
LLMResponse,
type FinishReasonDetails,
type LLMError,
type LLMRequest,
type UsageInput,
} from "./schema"
import { Context, Deferred, Effect, Latch, Layer, Queue, Scope, Stream } from "effect"
export type Response = readonly LLMEvent[] | Stream.Stream<LLMEvent, LLMError>
export type Gate = Readonly<{ started: Effect.Effect<void>; release: Effect.Effect<void> }>
export interface Interface {
readonly requests: LLMRequest[]
readonly push: (...responses: readonly Response[]) => Effect.Effect<void>
readonly always: (response: Response) => Effect.Effect<void>
readonly wait: (count: number) => Effect.Effect<void>
readonly gate: Effect.Effect<Gate, never, Scope.Scope>
readonly client: LLMClientShape
}
export interface LayerOptions {
readonly transformRequest?: (request: LLMRequest) => LLMRequest
/** Used after the one-shot response queue is exhausted. Omit to defect on unexpected requests. */
readonly fallback?: Response
}
export class Service extends Context.Service<Service, Interface>()("@opencode/ai/TestLLM") {}
export const complete = (
options: { readonly reason: FinishReasonDetails; readonly usage?: UsageInput },
...events: readonly LLMEvent[]
) => [
LLMEvent.stepStart({ index: 0 }),
...events,
LLMEvent.stepFinish({ index: 0, reason: options.reason, usage: options.usage }),
LLMEvent.finish({ reason: options.reason }),
]
export const stop = (...events: readonly LLMEvent[]) => complete({ reason: { normalized: "stop" } }, ...events)
export const toolCalls = (...events: readonly LLMEvent[]) =>
complete({ reason: { normalized: "tool-calls" } }, ...events)
const textEvents = (value: string, id: string) => [
LLMEvent.textStart({ id }),
LLMEvent.textDelta({ id, text: value }),
LLMEvent.textEnd({ id }),
]
export const text = (value: string, id: string) => stop(...textEvents(value, id))
export const textWithUsage = (value: string, id: string, inputTokens: number) =>
complete(
{ reason: { normalized: "stop" }, usage: { inputTokens, nonCachedInputTokens: inputTokens } },
...textEvents(value, id),
)
export const tool = (id: string, name: string, input: unknown) => toolCalls(LLMEvent.toolCall({ id, name, input }))
export const failAfter = (error: LLMError, ...events: readonly LLMEvent[]) =>
Stream.fromIterable(events).pipe(Stream.concat(Stream.fail(error)))
export const hangAfter = (...events: readonly LLMEvent[]) => Stream.concat(Stream.fromIterable(events), Stream.never)
const toStream = (response: Response) => (Stream.isStream(response) ? response : Stream.fromIterable(response))
export const layer = (options: LayerOptions = {}) =>
Layer.effect(
Service,
Effect.gen(function* () {
const requests: LLMRequest[] = []
const responses: Response[] = []
let started = Deferred.makeUnsafe<void>()
let fallback = options.fallback
let activeGate: { readonly started: Queue.Queue<void>; readonly release: Latch.Latch } | undefined
const wait = (count: number): Effect.Effect<void> =>
Effect.suspend(() =>
requests.length >= count ? Effect.void : Deferred.await(started).pipe(Effect.andThen(wait(count))),
)
const stream = ((request: LLMRequest) => {
requests.push(options.transformRequest?.(request) ?? request)
const waiting = started
started = Deferred.makeUnsafe()
Deferred.doneUnsafe(waiting, Effect.void)
const response = responses.shift() ?? fallback
if (!response) return Stream.die(new Error(`TestLLM has no response for request ${requests.length}`))
const streamed = toStream(response)
const gate = activeGate
if (!gate) return streamed
return Stream.unwrap(
Queue.offer(gate.started, undefined).pipe(Effect.andThen(gate.release.await), Effect.as(streamed)),
)
}) as LLMClientShape["stream"]
const client = LLMClient.Service.of({
stream,
generate: (request) =>
stream(request).pipe(
Stream.runFold(LLMResponse.empty, LLMResponse.reduce),
Effect.flatMap((state) => {
const response = LLMResponse.complete(state)
if (response) return Effect.succeed(response)
return Effect.die("TestLLM response ended without a terminal finish event")
}),
),
})
return Service.of({
requests,
push: (...input) =>
Effect.sync(() => {
responses.push(...input)
}),
always: (response) =>
Effect.sync(() => {
fallback = response
}),
wait,
gate: Effect.gen(function* () {
const gate = {
started: yield* Effect.acquireRelease(Queue.unbounded<void>(), Queue.shutdown),
release: yield* Latch.make(),
}
activeGate = gate
const release = Effect.sync(() => {
if (activeGate === gate) activeGate = undefined
}).pipe(Effect.andThen(gate.release.open), Effect.asVoid)
yield* Effect.addFinalizer(() => release)
return {
started: Queue.take(gate.started),
release,
}
}),
client,
})
}),
)
export const clientLayer = Layer.effect(
LLMClient.Service,
Effect.map(Service, (service) => service.client),
)
export const push = (...responses: readonly Response[]) => Service.use((service) => service.push(...responses))
export const always = (response: Response) => Service.use((service) => service.always(response))
export const wait = (count: number) => Service.use((service) => service.wait(count))
export const gate = Service.use((service) => service.gate)
+1 -1
View File
@@ -28,7 +28,7 @@ export const dispatch = (tools: Tools, call: ToolCallPart): Effect.Effect<Dispat
return decodeAndExecute(tool, call).pipe( return decodeAndExecute(tool, call).pipe(
Effect.map((value) => result(call, value)), Effect.map((value) => result(call, value)),
Effect.catchTag("LLM.ToolFailure", (failure) => Effect.catchTag("Tool.Error", (failure) =>
Effect.succeed(result(call, { type: "error", value: failure.message }, failure.error)), Effect.succeed(result(call, { type: "error", value: failure.message }, failure.error)),
), ),
) )
+6 -6
View File
@@ -1,7 +1,7 @@
import { Effect, JsonSchema, Schema } from "effect" import { Effect, JsonSchema, Schema } from "effect"
import { Tool } from "@opencode-ai/schema/tool"
import type { import type {
ToolCallPart, ToolCallPart,
ToolContent,
ToolDefinition as ToolDefinitionClass, ToolDefinition as ToolDefinitionClass,
ToolOutput as ToolOutputType, ToolOutput as ToolOutputType,
} from "./schema" } from "./schema"
@@ -31,7 +31,7 @@ export interface ToolModelOutputInput<Parameters, Output> {
export type ToolToModelOutput<Parameters extends ToolSchema<any>, Success extends ToolSchema<any>> = ( export type ToolToModelOutput<Parameters extends ToolSchema<any>, Success extends ToolSchema<any>> = (
input: ToolModelOutputInput<Schema.Schema.Type<Parameters>, Success["Encoded"]>, input: ToolModelOutputInput<Schema.Schema.Type<Parameters>, Success["Encoded"]>,
) => ReadonlyArray<ToolContent> ) => ReadonlyArray<Tool.Content>
/** /**
* A type-safe LLM tool. Each tool bundles its own description, parameter * A type-safe LLM tool. Each tool bundles its own description, parameter
@@ -95,7 +95,7 @@ type DynamicToolConfig = {
readonly jsonSchema: JsonSchema.JsonSchema readonly jsonSchema: JsonSchema.JsonSchema
readonly outputSchema?: JsonSchema.JsonSchema readonly outputSchema?: JsonSchema.JsonSchema
readonly execute?: (params: unknown, context?: ToolExecuteContext) => Effect.Effect<unknown, ToolFailure> readonly execute?: (params: unknown, context?: ToolExecuteContext) => Effect.Effect<unknown, ToolFailure>
readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<ToolContent> readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<Tool.Content>
readonly toStructuredOutput?: (output: unknown) => unknown readonly toStructuredOutput?: (output: unknown) => unknown
} }
@@ -151,7 +151,7 @@ export function make(config: {
readonly jsonSchema: JsonSchema.JsonSchema readonly jsonSchema: JsonSchema.JsonSchema
readonly outputSchema?: JsonSchema.JsonSchema readonly outputSchema?: JsonSchema.JsonSchema
readonly execute: (params: unknown, context?: ToolExecuteContext) => Effect.Effect<unknown, ToolFailure> readonly execute: (params: unknown, context?: ToolExecuteContext) => Effect.Effect<unknown, ToolFailure>
readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<ToolContent> readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<Tool.Content>
readonly toStructuredOutput?: (output: unknown) => unknown readonly toStructuredOutput?: (output: unknown) => unknown
}): AnyExecutableTool }): AnyExecutableTool
export function make(config: { export function make(config: {
@@ -159,7 +159,7 @@ export function make(config: {
readonly jsonSchema: JsonSchema.JsonSchema readonly jsonSchema: JsonSchema.JsonSchema
readonly outputSchema?: JsonSchema.JsonSchema readonly outputSchema?: JsonSchema.JsonSchema
readonly execute?: undefined readonly execute?: undefined
readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<ToolContent> readonly toModelOutput?: (input: ToolModelOutputInput<unknown, unknown>) => ReadonlyArray<Tool.Content>
readonly toStructuredOutput?: (output: unknown) => unknown readonly toStructuredOutput?: (output: unknown) => unknown
}): AnyTool }): AnyTool
export function make(config: TypedToolConfig | DynamicToolConfig): AnyTool { export function make(config: TypedToolConfig | DynamicToolConfig): AnyTool {
@@ -236,7 +236,7 @@ const toJsonSchema = (schema: Schema.Top): JsonSchema.JsonSchema => {
} }
const project = ( const project = (
toModelOutput: ((input: ToolModelOutputInput<any, any>) => ReadonlyArray<ToolContent>) | undefined, toModelOutput: ((input: ToolModelOutputInput<any, any>) => ReadonlyArray<Tool.Content>) | undefined,
toStructuredOutput: ((output: unknown) => unknown) | undefined, toStructuredOutput: ((output: unknown) => unknown) | undefined,
parameters: unknown, parameters: unknown,
callID: ToolCallPart["id"], callID: ToolCallPart["id"],
+7 -7
View File
@@ -1,7 +1,8 @@
import { describe, expect } from "bun:test" import { describe, expect } from "bun:test"
import { Effect, Schema, Stream } from "effect" import { Effect, Schema, Stream } from "effect"
import { LLM, LLMResponse } from "../src" import { LLM, LLMRequest, LLMResponse } from "../src"
import { Route, Endpoint, LLMClient, Protocol, type FramingDef } from "../src/route" import { Route, Endpoint, LLMClient, Protocol, type FramingDef } from "../src/route"
import { compileRequest } from "../src/route/client"
import { Model } from "../src/schema" import { Model } from "../src/schema"
import { testEffect } from "./lib/effect" import { testEffect } from "./lib/effect"
import { dynamicResponse } from "./lib/http" import { dynamicResponse } from "./lib/http"
@@ -40,7 +41,7 @@ const fakeFraming: FramingDef<FakeEvent> = {
const raiseEvent = (event: FakeEvent): import("../src/schema").LLMEvent => const raiseEvent = (event: FakeEvent): import("../src/schema").LLMEvent =>
event.type === "finish" event.type === "finish"
? { type: "finish", reason: event.reason } ? { type: "finish", reason: { normalized: event.reason } }
: { type: "text-delta", id: "text-0", text: event.text } : { type: "text-delta", id: "text-0", text: event.text }
const fakeProtocol = Protocol.make<FakeBody, FakeEvent, FakeEvent, void>({ const fakeProtocol = Protocol.make<FakeBody, FakeEvent, FakeEvent, void>({
@@ -139,9 +140,8 @@ describe("llm route", () => {
it.effect("selects routes by model route value", () => it.effect("selects routes by model route value", () =>
Effect.gen(function* () { Effect.gen(function* () {
const llm = yield* LLMClient.Service const prepared = yield* compileRequest(
const prepared = yield* llm.prepare( LLMRequest.update(request, { model: updateModel(request.model, { route: configuredGemini }) }),
LLM.updateRequest(request, { model: updateModel(request.model, { route: configuredGemini }) }),
) )
expect(prepared.route).toBe("gemini-fake") expect(prepared.route).toBe("gemini-fake")
@@ -173,8 +173,8 @@ describe("llm route", () => {
framing: fakeFraming, framing: fakeFraming,
}) })
const prepared = yield* (yield* LLMClient.Service).prepare( const prepared = yield* compileRequest(
LLM.updateRequest(request, { model: updateModel(request.model, { route: duplicate }) }), LLMRequest.update(request, { model: updateModel(request.model, { route: duplicate }) }),
) )
expect(prepared.body).toEqual({ body: "late-default" }) expect(prepared.body).toEqual({ body: "late-default" })
+26 -2
View File
@@ -137,15 +137,26 @@ Azure.configure({ apiKey: "azure-key", resourceName: "resource" }).chat("deploym
Azure.configure({ resourceName: "resource", apiKey: "azure-key", auth: Auth.header("api-key", "override") }) Azure.configure({ resourceName: "resource", apiKey: "azure-key", auth: Auth.header("api-key", "override") })
Anthropic.configure({ apiKey: "anthropic-key" }).model("claude-haiku") Anthropic.configure({ apiKey: "anthropic-key" }).model("claude-haiku")
Anthropic.configure({
apiKey: "anthropic-key",
providerOptions: {
anthropic: { thinking: { type: "enabled", budgetTokens: 1_024 }, effort: "high" },
},
}).model("claude-haiku")
// @ts-expect-error Anthropic model selectors only accept model ids. // @ts-expect-error Anthropic model selectors only accept model ids.
Anthropic.configure({ apiKey: "anthropic-key" }).model("claude-haiku", {}) Anthropic.configure({ apiKey: "anthropic-key" }).model("claude-haiku", {})
// @ts-expect-error Anthropic package settings accept only one auth source. // @ts-expect-error Anthropic package settings accept only one auth source.
Anthropic.model("claude-sonnet-4-6", { apiKey: "anthropic-key", authToken: "anthropic-token" }) Anthropic.model("claude-sonnet-4-6", { apiKey: "anthropic-key", authToken: "anthropic-token" })
// @ts-expect-error Enabled Anthropic thinking requires a token budget.
Anthropic.configure({ providerOptions: { anthropic: { thinking: { type: "enabled" } } } })
// @ts-expect-error Anthropic thinking budgets must be numbers.
Anthropic.configure({ providerOptions: { anthropic: { thinking: { type: "enabled", budgetTokens: "large" } } } })
AnthropicCompatible.configure({ AnthropicCompatible.configure({
apiKey: "messages-key", apiKey: "messages-key",
baseURL: "https://messages.example.com/v1", baseURL: "https://messages.example.com/v1",
provider: "example", provider: "example",
providerOptions: { anthropic: { thinking: { type: "disabled" } } },
}).model("compatible-model") }).model("compatible-model")
// @ts-expect-error Anthropic-compatible providers require a base URL. // @ts-expect-error Anthropic-compatible providers require a base URL.
AnthropicCompatible.configure({ apiKey: "messages-key" }) AnthropicCompatible.configure({ apiKey: "messages-key" })
@@ -159,10 +170,19 @@ AnthropicCompatible.model("compatible-model", {
}) })
Google.configure({ apiKey: "google-key" }).model("gemini-2.5-flash") Google.configure({ apiKey: "google-key" }).model("gemini-2.5-flash")
Google.configure({
apiKey: "google-key",
providerOptions: { gemini: { thinkingConfig: { thinkingBudget: 0, includeThoughts: false } } },
}).model("gemini-2.5-flash")
// @ts-expect-error Google model selectors only accept model ids. // @ts-expect-error Google model selectors only accept model ids.
Google.configure({ apiKey: "google-key" }).model("gemini-2.5-flash", {}) Google.configure({ apiKey: "google-key" }).model("gemini-2.5-flash", {})
// @ts-expect-error Gemini thinking budgets must be numbers.
Google.configure({ providerOptions: { gemini: { thinkingConfig: { thinkingBudget: "large" } } } })
GoogleVertex.configure({ apiKey: "vertex-key" }).model("gemini-3.5-flash") GoogleVertex.configure({
apiKey: "vertex-key",
providerOptions: { gemini: { thinkingConfig: { thinkingBudget: 1_024 } } },
}).model("gemini-3.5-flash")
GoogleVertex.configure({ accessToken: "vertex-token", project: "project" }).model("gemini-3.5-flash") GoogleVertex.configure({ accessToken: "vertex-token", project: "project" }).model("gemini-3.5-flash")
GoogleVertex.configure({ auth: Auth.bearer("vertex-token"), project: "project" }).model("gemini-3.5-flash") GoogleVertex.configure({ auth: Auth.bearer("vertex-token"), project: "project" }).model("gemini-3.5-flash")
// @ts-expect-error Vertex Gemini model selectors only accept model ids. // @ts-expect-error Vertex Gemini model selectors only accept model ids.
@@ -208,7 +228,11 @@ GoogleVertexResponses.configure({
project: "project", project: "project",
}) })
GoogleVertexMessages.configure({ accessToken: "vertex-token", project: "project" }).model("claude-sonnet-4-6") GoogleVertexMessages.configure({
accessToken: "vertex-token",
project: "project",
providerOptions: { anthropic: { thinking: { type: "adaptive", display: "omitted" }, effort: "low" } },
}).model("claude-sonnet-4-6")
// @ts-expect-error Vertex Messages package settings do not accept API keys. // @ts-expect-error Vertex Messages package settings do not accept API keys.
GoogleVertexMessages.model("claude-sonnet-4-6", { apiKey: "vertex-key", project: "project" }) GoogleVertexMessages.model("claude-sonnet-4-6", { apiKey: "vertex-key", project: "project" })
GoogleVertexMessages.configure({ auth: Auth.bearer("vertex-token"), project: "project" }).model("claude-sonnet-4-6") GoogleVertexMessages.configure({ auth: Auth.bearer("vertex-token"), project: "project" }).model("claude-sonnet-4-6")
+81 -20
View File
@@ -1,7 +1,8 @@
import { describe, expect, test } from "bun:test" import { describe, expect, test } from "bun:test"
import { Effect } from "effect" import { Effect } from "effect"
import { CacheHint, LLM, Message } from "../src" import { CacheHint, LLM, Message } from "../src"
import { Auth, LLMClient } from "../src/route" import { Auth } from "../src/route"
import { compileRequest } from "../src/route/client"
import { AmazonBedrock } from "../src/providers" import { AmazonBedrock } from "../src/providers"
import * as AnthropicMessages from "../src/protocols/anthropic-messages" import * as AnthropicMessages from "../src/protocols/anthropic-messages"
import * as Gemini from "../src/protocols/gemini" import * as Gemini from "../src/protocols/gemini"
@@ -31,7 +32,7 @@ const geminiModel = Gemini.route
describe("applyCachePolicy", () => { describe("applyCachePolicy", () => {
it.effect("undefined cache resolves to 'auto' (the recommended default)", () => it.effect("undefined cache resolves to 'auto' (the recommended default)", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: "You are concise.", system: "You are concise.",
@@ -39,8 +40,8 @@ describe("applyCachePolicy", () => {
}), }),
) )
// No explicit cache field → auto policy fires → last system part + latest // A single system block is both the first and last boundary, so the auto
// user message both get cache_control markers. // policy deduplicates it and still marks the conversation tail.
expect(prepared.body).toMatchObject({ expect(prepared.body).toMatchObject({
system: [{ type: "text", text: "You are concise.", cache_control: { type: "ephemeral" } }], system: [{ type: "text", text: "You are concise.", cache_control: { type: "ephemeral" } }],
messages: [{ role: "user", content: [{ type: "text", text: "hi", cache_control: { type: "ephemeral" } }] }], messages: [{ role: "user", content: [{ type: "text", text: "hi", cache_control: { type: "ephemeral" } }] }],
@@ -48,12 +49,15 @@ describe("applyCachePolicy", () => {
}), }),
) )
it.effect("'auto' marks the last tool, last system part, and latest user message on Anthropic", () => it.effect("'auto' marks the last tool, first and last system parts, and final message boundary on Anthropic", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: "Sys A", system: [
{ type: "text", text: "Base agent" },
{ type: "text", text: "Project instructions" },
],
tools: [{ name: "t1", description: "t1", inputSchema: { type: "object", properties: {} } }], tools: [{ name: "t1", description: "t1", inputSchema: { type: "object", properties: {} } }],
messages: [ messages: [
Message.user("first user"), Message.user("first user"),
@@ -66,7 +70,10 @@ describe("applyCachePolicy", () => {
expect(prepared.body).toMatchObject({ expect(prepared.body).toMatchObject({
tools: [{ name: "t1", cache_control: { type: "ephemeral" } }], tools: [{ name: "t1", cache_control: { type: "ephemeral" } }],
system: [{ type: "text", text: "Sys A", cache_control: { type: "ephemeral" } }], system: [
{ type: "text", text: "Base agent", cache_control: { type: "ephemeral" } },
{ type: "text", text: "Project instructions", cache_control: { type: "ephemeral" } },
],
messages: [ messages: [
{ role: "user", content: [{ type: "text", text: "first user" }] }, { role: "user", content: [{ type: "text", text: "first user" }] },
{ role: "assistant", content: [{ type: "text", text: "assistant reply" }] }, { role: "assistant", content: [{ type: "text", text: "assistant reply" }] },
@@ -81,7 +88,7 @@ describe("applyCachePolicy", () => {
it.effect("'auto' is a no-op on OpenAI (implicit caching protocol)", () => it.effect("'auto' is a no-op on OpenAI (implicit caching protocol)", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: openaiModel, model: openaiModel,
system: "Sys", system: "Sys",
@@ -100,7 +107,7 @@ describe("applyCachePolicy", () => {
it.effect("'auto' is a no-op on Gemini (out-of-band caching protocol)", () => it.effect("'auto' is a no-op on Gemini (out-of-band caching protocol)", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: geminiModel, model: geminiModel,
system: "Sys", system: "Sys",
@@ -117,10 +124,13 @@ describe("applyCachePolicy", () => {
it.effect("'auto' on Bedrock emits cachePoint markers in the right places", () => it.effect("'auto' on Bedrock emits cachePoint markers in the right places", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: bedrockModel, model: bedrockModel,
system: "Sys", system: [
{ type: "text", text: "Base agent" },
{ type: "text", text: "Project instructions" },
],
tools: [{ name: "t1", description: "t1", inputSchema: { type: "object", properties: {} } }], tools: [{ name: "t1", description: "t1", inputSchema: { type: "object", properties: {} } }],
messages: [Message.user("first user"), Message.assistant("reply"), Message.user("latest user")], messages: [Message.user("first user"), Message.assistant("reply"), Message.user("latest user")],
cache: "auto", cache: "auto",
@@ -131,7 +141,12 @@ describe("applyCachePolicy", () => {
toolConfig: { toolConfig: {
tools: [{ toolSpec: { name: "t1" } }, { cachePoint: { type: "default" } }], tools: [{ toolSpec: { name: "t1" } }, { cachePoint: { type: "default" } }],
}, },
system: [{ text: "Sys" }, { cachePoint: { type: "default" } }], system: [
{ text: "Base agent" },
{ cachePoint: { type: "default" } },
{ text: "Project instructions" },
{ cachePoint: { type: "default" } },
],
messages: [ messages: [
{ role: "user", content: [{ text: "first user" }] }, { role: "user", content: [{ text: "first user" }] },
{ role: "assistant", content: [{ text: "reply" }] }, { role: "assistant", content: [{ text: "reply" }] },
@@ -143,7 +158,7 @@ describe("applyCachePolicy", () => {
it.effect("'none' disables auto placement even when manual hints exist", () => it.effect("'none' disables auto placement even when manual hints exist", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: "Sys", system: "Sys",
@@ -162,7 +177,7 @@ describe("applyCachePolicy", () => {
it.effect("granular object form: tools-only marks just tools", () => it.effect("granular object form: tools-only marks just tools", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: "Sys", system: "Sys",
@@ -181,7 +196,7 @@ describe("applyCachePolicy", () => {
it.effect("auto policy preserves manual CacheHints on other parts", () => it.effect("auto policy preserves manual CacheHints on other parts", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: [ system: [
@@ -193,15 +208,61 @@ describe("applyCachePolicy", () => {
}), }),
) )
const body = prepared.body as { system: Array<{ text: string; cache_control?: unknown }> } const body = prepared.body as {
system: Array<{ text: string; cache_control?: unknown }>
messages: Array<{ content: Array<{ cache_control?: unknown }> }>
}
expect(body.system[0]?.cache_control).toEqual({ type: "ephemeral", ttl: "1h" }) expect(body.system[0]?.cache_control).toEqual({ type: "ephemeral", ttl: "1h" })
expect(body.system[1]?.cache_control).toEqual({ type: "ephemeral" }) expect(body.system[1]?.cache_control).toEqual({ type: "ephemeral" })
expect(body.messages[0]?.content[0]?.cache_control).toEqual({ type: "ephemeral" })
}),
)
it.effect("auto policy stays within the four-breakpoint cap when preserving manual hints", () =>
Effect.gen(function* () {
const request = LLM.request({
model: anthropicModel,
system: [
{ type: "text", text: "Base agent" },
{
type: "text",
text: "Manual context",
cache: new CacheHint({ type: "ephemeral", ttlSeconds: 3600 }),
},
{ type: "text", text: "Project instructions" },
],
tools: [{ name: "t1", description: "t1", inputSchema: { type: "object", properties: {} } }],
prompt: "hi",
cache: "auto",
})
const applied = applyCachePolicy(request)
expect(applied.tools[0]?.cache).toBeDefined()
expect(applied.system.map((part) => part.cache !== undefined)).toEqual([true, true, true])
const tail = applied.messages[0]!.content[0]!
expect("cache" in tail ? tail.cache : undefined).toBeUndefined()
expect(applyCachePolicy(applied)).toBe(applied)
const prepared = yield* compileRequest(request)
const body = prepared.body as {
tools: Array<{ cache_control?: unknown }>
system: Array<{ cache_control?: unknown }>
messages: Array<{ content: Array<{ cache_control?: unknown }> }>
}
const marked = [
...body.tools.map((tool) => tool.cache_control),
...body.system.map((part) => part.cache_control),
...body.messages.flatMap((message) => message.content.map((part) => part.cache_control)),
].filter((cache) => cache !== undefined)
expect(marked).toHaveLength(4)
expect(body.system[1]?.cache_control).toEqual({ type: "ephemeral", ttl: "1h" })
expect(body.messages[0]?.content[0]?.cache_control).toBeUndefined()
}), }),
) )
it.effect("ttlSeconds in the policy flows through to wire markers", () => it.effect("ttlSeconds in the policy flows through to wire markers", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
system: "Sys", system: "Sys",
@@ -218,7 +279,7 @@ describe("applyCachePolicy", () => {
it.effect("messages: { tail: 2 } marks the last 2 message boundaries", () => it.effect("messages: { tail: 2 } marks the last 2 message boundaries", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
messages: [Message.user("u1"), Message.assistant("a1"), Message.user("u2"), Message.assistant("a2")], messages: [Message.user("u1"), Message.assistant("a1"), Message.user("u2"), Message.assistant("a2")],
@@ -236,7 +297,7 @@ describe("applyCachePolicy", () => {
it.effect("'latest-assistant' marks the last assistant message", () => it.effect("'latest-assistant' marks the last assistant message", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: anthropicModel, model: anthropicModel,
messages: [Message.user("u1"), Message.assistant("a1"), Message.user("u2")], messages: [Message.user("u1"), Message.assistant("a1"), Message.user("u2")],
@@ -4,6 +4,7 @@ import { HttpClientRequest } from "effect/unstable/http"
import { LLM, mergeProviderOptions } from "../src" import { LLM, mergeProviderOptions } from "../src"
import { AnthropicMessages, OpenAIChat } from "../src/protocols" import { AnthropicMessages, OpenAIChat } from "../src/protocols"
import { Auth, LLMClient } from "../src/route" import { Auth, LLMClient } from "../src/route"
import { compileRequest } from "../src/route/client"
import { it } from "./lib/effect" import { it } from "./lib/effect"
import { dynamicResponse } from "./lib/http" import { dynamicResponse } from "./lib/http"
import { deltaChunk } from "./lib/openai-chunks" import { deltaChunk } from "./lib/openai-chunks"
@@ -44,7 +45,7 @@ describe("request option precedence", () => {
}) })
}) })
it.effect("prepares bodies with route defaults, model defaults, and call options in order", () => it.effect("compiles bodies with route defaults, model defaults, and call options in order", () =>
Effect.gen(function* () { Effect.gen(function* () {
const route = OpenAIChat.route.with({ const route = OpenAIChat.route.with({
endpoint: { baseURL: "https://api.openai.test/v1/" }, endpoint: { baseURL: "https://api.openai.test/v1/" },
@@ -59,7 +60,7 @@ describe("request option precedence", () => {
providerOptions: { openai: { reasoningEffort: "medium" } }, providerOptions: { openai: { reasoningEffort: "medium" } },
}, },
}) })
const prepared = yield* LLMClient.prepare<OpenAIChat.OpenAIChatBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
prompt: "Say hello.", prompt: "Say hello.",
@@ -136,24 +137,61 @@ describe("request option precedence", () => {
), ),
) )
it.effect("rejects raw body overlays for protocol-owned roots", () => it.effect("transforms the final HTTP request after serialization and authentication", () =>
Effect.gen(function* () { LLMClient.generate(
const model = OpenAIChat.route LLM.request({
.with({ endpoint: { baseURL: "https://api.openai.test/v1/" }, auth: Auth.bearer("test") }) model: OpenAIChat.route
.model({ id: "gpt-4o-mini" }) .with({ endpoint: { baseURL: "https://api.openai.test/v1/" }, auth: Auth.bearer("fresh-key") })
const error = yield* LLMClient.prepare( .model({ id: "gpt-4o-mini" }),
LLM.request({ prompt: "Say hello.",
model, }),
prompt: "Say hello.", {
http: { body: { model: "gpt-5", messages: [], tools: [] } }, transform: (request) =>
}), Effect.sync(() => {
).pipe(Effect.flip) expect(request.headers.authorization).toBe("Bearer fresh-key")
request.url = "https://proxy.test/v1/chat/completions"
request.headers["x-plugin"] = "transformed"
request.body = JSON.stringify({ transformed: true })
}),
},
).pipe(
Effect.provide(
dynamicResponse((input) =>
Effect.gen(function* () {
const web = yield* HttpClientRequest.toWeb(input.request).pipe(Effect.orDie)
expect(web.url).toBe("https://proxy.test/v1/chat/completions")
expect(web.headers.get("x-plugin")).toBe("transformed")
expect(decodeJson(input.text)).toEqual({ transformed: true })
return input.respond(sseEvents(deltaChunk({}, "stop")), {
headers: { "content-type": "text/event-stream" },
})
}),
),
),
),
)
expect(error.reason).toMatchObject({ it.effect("applies raw body overlays after protocol lowering", () =>
_tag: "InvalidRequest", LLMClient.generate(
message: "http.body cannot overlay protocol-owned field(s): model, messages, tools", LLM.request({
}) model: OpenAIChat.route
}), .with({ endpoint: { baseURL: "https://api.openai.test/v1/" }, auth: Auth.bearer("test") })
.model({ id: "gpt-4o-mini" }),
prompt: "Say hello.",
http: { body: { model: "gpt-5", messages: [], tools: [] } },
}),
).pipe(
Effect.provide(
dynamicResponse((input) =>
Effect.gen(function* () {
expect(decodeJson(input.text)).toMatchObject({ model: "gpt-5", messages: [], tools: [] })
return input.respond(sseEvents(deltaChunk({}, "stop")), {
headers: { "content-type": "text/event-stream" },
})
}),
),
),
),
) )
it.effect("uses model output limits after route limits and before call maxTokens", () => it.effect("uses model output limits after route limits and before call maxTokens", () =>
@@ -164,10 +202,8 @@ describe("request option precedence", () => {
limits: { output: 128 }, limits: { output: 128 },
}) })
const model = route.model({ id: "claude-sonnet-4-5", defaults: { limits: { output: 64 } } }) const model = route.model({ id: "claude-sonnet-4-5", defaults: { limits: { output: 64 } } })
const withoutMaxTokens = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const withoutMaxTokens = yield* compileRequest(LLM.request({ model, prompt: "Say hello.", cache: "none" }))
LLM.request({ model, prompt: "Say hello.", cache: "none" }), const withMaxTokens = yield* compileRequest(
)
const withMaxTokens = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>(
LLM.request({ model, prompt: "Say hello.", cache: "none", generation: { maxTokens: 32 } }), LLM.request({ model, prompt: "Say hello.", cache: "none", generation: { maxTokens: 32 } }),
) )
+11 -3
View File
@@ -11,8 +11,15 @@ import {
XAI, XAI,
} from "@opencode-ai/ai/providers" } from "@opencode-ai/ai/providers"
import * as GitHubCopilot from "@opencode-ai/ai/providers/github-copilot" import * as GitHubCopilot from "@opencode-ai/ai/providers/github-copilot"
import { OpenAIChat, OpenAICompatibleChat, OpenAICompatibleResponses, OpenAIResponses } from "@opencode-ai/ai/protocols" import {
OpenAIChat,
OpenAICompatibleChat,
OpenAICompatibleResponses,
OpenAIResponses,
OpenResponses,
} from "@opencode-ai/ai/protocols"
import * as AnthropicMessages from "@opencode-ai/ai/protocols/anthropic-messages" import * as AnthropicMessages from "@opencode-ai/ai/protocols/anthropic-messages"
import { TestLLM } from "@opencode-ai/ai/testing"
describe("public exports", () => { describe("public exports", () => {
test("root exposes app-facing runtime APIs", () => { test("root exposes app-facing runtime APIs", () => {
@@ -22,6 +29,7 @@ describe("public exports", () => {
expect(ImageInput.bytes).toBeFunction() expect(ImageInput.bytes).toBeFunction()
expect(Provider.make).toBeFunction() expect(Provider.make).toBeFunction()
expect(ProviderSubpath.make).toBe(Provider.make) expect(ProviderSubpath.make).toBe(Provider.make)
expect(TestLLM.layer).toBeFunction()
}) })
test("route barrel exposes route-authoring APIs", () => { test("route barrel exposes route-authoring APIs", () => {
@@ -45,9 +53,7 @@ describe("public exports", () => {
expect(CloudflareWorkersAI.configure).toBeFunction() expect(CloudflareWorkersAI.configure).toBeFunction()
expect(CloudflareWorkersAI.configure({ accountId: "fixture", apiKey: "fixture" }).model).toBeFunction() expect(CloudflareWorkersAI.configure({ accountId: "fixture", apiKey: "fixture" }).model).toBeFunction()
expect(OpenRouter.model).toBeFunction() expect(OpenRouter.model).toBeFunction()
expect(OpenRouter.provider.model).toBe(OpenRouter.model)
expect(XAI.model).toBeFunction() expect(XAI.model).toBeFunction()
expect(XAI.provider.model).toBe(XAI.model)
expect(XAI.provider.responses).toBe(XAI.responses) expect(XAI.provider.responses).toBe(XAI.responses)
expect(XAI.provider.chat).toBe(XAI.chat) expect(XAI.provider.chat).toBe(XAI.chat)
expect(XAI.configure({ apiKey: "fixture" }).responses("grok-4.3").route.id).toBe("openai-responses") expect(XAI.configure({ apiKey: "fixture" }).responses("grok-4.3").route.id).toBe("openai-responses")
@@ -74,7 +80,9 @@ describe("public exports", () => {
test("protocol barrels expose supported low-level routes", () => { test("protocol barrels expose supported low-level routes", () => {
expect(OpenAIChat.route.id).toBe("openai-chat") expect(OpenAIChat.route.id).toBe("openai-chat")
expect(OpenAICompatibleChat.route.id).toBe("openai-compatible-chat") expect(OpenAICompatibleChat.route.id).toBe("openai-compatible-chat")
expect(OpenResponses.protocol.id).toBe("open-responses")
expect(OpenAICompatibleResponses.route.id).toBe("openai-compatible-responses") expect(OpenAICompatibleResponses.route.id).toBe("openai-compatible-responses")
expect(OpenAICompatibleResponses.route.protocol).toBe("open-responses")
expect(OpenAIResponses.route.id).toBe("openai-responses") expect(OpenAIResponses.route.id).toBe("openai-responses")
expect(OpenAIResponses.webSocketRoute.id).toBe("openai-responses-websocket") expect(OpenAIResponses.webSocketRoute.id).toBe("openai-responses-websocket")
expect(AnthropicMessages.route.id).toBe("anthropic-messages") expect(AnthropicMessages.route.id).toBe("anthropic-messages")
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -83,7 +83,7 @@ const indexStep = (event: LLMEvent, index: number): LLMEvent => {
const stepState = (events: ReadonlyArray<LLMEvent>) => { const stepState = (events: ReadonlyArray<LLMEvent>) => {
const assistantContent: ContentPart[] = [] const assistantContent: ContentPart[] = []
const toolCalls: ToolCallPart[] = [] const toolCalls: ToolCallPart[] = []
let reason: Extract<LLMEvent, { type: "finish" }>["reason"] = "unknown" let reason: Extract<LLMEvent, { type: "finish" }>["reason"] = { normalized: "unknown" }
let usage: Usage | undefined let usage: Usage | undefined
let providerMetadata: ProviderMetadata | undefined let providerMetadata: ProviderMetadata | undefined
@@ -0,0 +1,47 @@
import { Schema } from "effect"
import { LLM, type Model, type ModelProviderOptions, type ProviderOptions } from "../src"
import { OpenAIChat } from "../src/protocols"
interface ExampleOptions {
readonly [key: string]: unknown
readonly mode?: "fast" | "thorough"
}
type ExampleProviderOptions = ProviderOptions & {
readonly example?: ExampleOptions
}
const model = OpenAIChat.route
.with({ endpoint: { baseURL: "https://example.com/v1" } })
.model<ExampleProviderOptions>({ id: "example" })
LLM.request({ model, prompt: "Hello", providerOptions: { example: { mode: "fast" } } })
LLM.request({ model, prompt: "Hello", providerOptions: { future: { option: true } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Known provider options preserve their value types.
providerOptions: { example: { mode: "slow" } },
})
LLM.generateObject({
model,
prompt: "Hello",
schema: Schema.Struct({ answer: Schema.String }),
providerOptions: { example: { mode: "thorough" } },
})
LLM.generateObject({
model,
prompt: "Hello",
jsonSchema: { type: "object" },
// @ts-expect-error Dynamic object generation uses the selected model's provider options.
providerOptions: { example: { mode: false } },
})
declare const generic: Model
LLM.request({ model: generic, prompt: "Hello", providerOptions: { arbitrary: { option: true } } })
const options: ModelProviderOptions<typeof model> = { example: { mode: "fast" } }
void options
+13 -4
View File
@@ -2,7 +2,16 @@ import { describe, expect, test } from "bun:test"
import { CacheHint, LLM, LLMResponse } from "../src" import { CacheHint, LLM, LLMResponse } from "../src"
import * as OpenAIChat from "../src/protocols/openai-chat" import * as OpenAIChat from "../src/protocols/openai-chat"
import * as OpenAIResponses from "../src/protocols/openai-responses" import * as OpenAIResponses from "../src/protocols/openai-responses"
import { LLMRequest, Message, Model, ToolCallPart, ToolChoice, ToolDefinition, ToolResultPart } from "../src/schema" import {
GenerationOptions,
LLMRequest,
Message,
Model,
ToolCallPart,
ToolChoice,
ToolDefinition,
ToolResultPart,
} from "../src/schema"
const chatRoute = OpenAIChat.route const chatRoute = OpenAIChat.route
const responsesRoute = OpenAIResponses.route const responsesRoute = OpenAIResponses.route
@@ -31,8 +40,8 @@ describe("llm constructors", () => {
model: Model.make({ id: "fake-model", provider: "fake", route: chatRoute }), model: Model.make({ id: "fake-model", provider: "fake", route: chatRoute }),
prompt: "Say hello.", prompt: "Say hello.",
}) })
const updated = LLM.updateRequest(base, { const updated = LLMRequest.update(base, {
generation: { maxTokens: 20 }, generation: GenerationOptions.make({ maxTokens: 20 }),
messages: [...base.messages, Message.assistant("Hi.")], messages: [...base.messages, Message.assistant("Hi.")],
}) })
@@ -191,7 +200,7 @@ describe("llm constructors", () => {
LLMResponse.text({ LLMResponse.text({
events: [ events: [
{ type: "text-delta", id: "text-0", text: "hi" }, { type: "text-delta", id: "text-0", text: "hi" },
{ type: "finish", reason: "stop" }, { type: "finish", reason: { normalized: "stop" } },
], ],
}), }),
).toBe("hi") ).toBe("hi")
+6
View File
@@ -58,6 +58,12 @@ describe("provider error classification", () => {
).toEqual(["ProviderInternal", "ProviderInternal"]) ).toEqual(["ProviderInternal", "ProviderInternal"])
}) })
test("classifies transient client statuses as provider internal", () => {
expect(
[408, 409].map((status) => classifyProviderFailure({ message: `HTTP ${status}`, status })._tag),
).toEqual(["ProviderInternal", "ProviderInternal"])
})
test("classifies nested provider codes when a top-level code is also present", () => { test("classifies nested provider codes when a top-level code is also present", () => {
expect( expect(
[ [
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { AnthropicCompatible } from "../../src/providers"
const model = AnthropicCompatible.configure({ baseURL: "https://example.com" }).model("claude")
LLM.request({ model, prompt: "Hello", providerOptions: { anthropic: { effort: "high" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Anthropic effort must be a string.
providerOptions: { anthropic: { effort: 1 } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { Anthropic } from "../../src/providers"
const model = Anthropic.provider.model("claude-sonnet-4-5")
LLM.request({ model, prompt: "Hello", providerOptions: { anthropic: { thinking: { type: "adaptive" } } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Anthropic thinking modes are a fixed union.
providerOptions: { anthropic: { thinking: { type: "automatic" } } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { Azure } from "../../src/providers"
const model = Azure.configure({ resourceName: "example" }).responses("deployment")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { store: false } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Azure OpenAI store must be boolean.
providerOptions: { openai: { store: "false" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { CloudflareWorkersAI } from "../../src/providers"
const model = CloudflareWorkersAI.configure({ accountId: "account", apiKey: "test" }).model("model")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { promptCacheKey: "cache" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Cloudflare's OpenAI-compatible prompt cache key must be a string.
providerOptions: { openai: { promptCacheKey: 1 } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { GitHubCopilot } from "../../src/providers"
const model = GitHubCopilot.configure({ baseURL: "https://example.com" }).model("gpt-5")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { reasoningSummary: "auto" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Copilot reasoning summaries use the OpenAI union.
providerOptions: { openai: { reasoningSummary: "full" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { GoogleVertexChat } from "../../src/providers"
const model = GoogleVertexChat.configure({ accessToken: "test", project: "project" }).model("gemini")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { serviceTier: "priority" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Vertex OpenAI-compatible service tiers use the OpenAI union.
providerOptions: { openai: { serviceTier: "premium" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { GoogleVertexMessages } from "../../src/providers"
const model = GoogleVertexMessages.configure({ accessToken: "test", project: "project" }).model("claude")
LLM.request({ model, prompt: "Hello", providerOptions: { anthropic: { effort: "medium" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Vertex Anthropic effort must be a string.
providerOptions: { anthropic: { effort: false } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { GoogleVertexResponses } from "../../src/providers"
const model = GoogleVertexResponses.configure({ accessToken: "test", project: "project" }).model("gemini")
LLM.request({ model, prompt: "Hello", providerOptions: { openresponses: { textVerbosity: "high" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Vertex Responses verbosity uses the Open Responses union.
providerOptions: { openresponses: { textVerbosity: "verbose" } },
})
@@ -0,0 +1,17 @@
import { LLM } from "../../src"
import { GoogleVertex } from "../../src/providers"
const model = GoogleVertex.provider.configure({ apiKey: "test" }).model("gemini-2.5-pro")
LLM.request({
model,
prompt: "Hello",
providerOptions: { gemini: { thinkingConfig: { includeThoughts: true } } },
})
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Vertex Gemini includeThoughts must be boolean.
providerOptions: { gemini: { thinkingConfig: { includeThoughts: "yes" } } },
})
@@ -0,0 +1,47 @@
import { LLM } from "../../src"
import { Google } from "../../src/providers"
const model = Google.provider.model("gemini-2.5-pro")
LLM.request({
model,
prompt: "Hello",
providerOptions: { gemini: { thinkingConfig: { thinkingBudget: 1024 } } },
})
LLM.request({
model,
prompt: "Hello",
providerOptions: {
gemini: {
// @ts-expect-error Gemini safety settings require a threshold for every category.
safetySettings: [{ category: "HARM_CATEGORY_HATE_SPEECH" }],
},
},
})
LLM.request({
model,
prompt: "Hello",
providerOptions: {
gemini: {
cachedContent: "cachedContents/example",
safetySettings: [{ category: "HARM_CATEGORY_HATE_SPEECH", threshold: "BLOCK_ONLY_HIGH" }],
serviceTier: "future-tier",
thinkingConfig: { thinkingLevel: "high", includeThoughts: true },
},
},
})
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Gemini thinking budgets must be numeric.
providerOptions: { gemini: { thinkingConfig: { thinkingBudget: "large" } } },
})
LLM.request({
model,
prompt: "Hello",
providerOptions: { gemini: { thinkingConfig: { thinkingLevel: "maximum" } } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { OpenAICompatibleResponses } from "../../src/providers"
const model = OpenAICompatibleResponses.configure({ baseURL: "https://example.com" }).model("model")
LLM.request({ model, prompt: "Hello", providerOptions: { openresponses: { reasoningSummary: "detailed" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error Open Responses reasoning summaries use a fixed union.
providerOptions: { openresponses: { reasoningSummary: "full" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { OpenAICompatible } from "../../src/providers"
const model = OpenAICompatible.deepseek.model("deepseek-chat")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { store: false } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error OpenAI-compatible store must be boolean.
providerOptions: { openai: { store: "false" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { OpenAI } from "../../src/providers"
const model = OpenAI.responses("gpt-5")
LLM.request({ model, prompt: "Hello", providerOptions: { openai: { reasoningEffort: "high" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error OpenAI reasoning effort must be a string.
providerOptions: { openai: { reasoningEffort: 1 } },
})
@@ -0,0 +1,35 @@
import { LLM } from "../../src"
import { OpenRouter } from "../../src/providers"
const model = OpenRouter.provider.model("anthropic/claude-sonnet-4.5")
LLM.request({ model, prompt: "Hello", providerOptions: { openrouter: { usage: true } } })
LLM.request({
model,
prompt: "Hello",
providerOptions: {
openrouter: {
models: ["google/gemini-3.1-pro"],
provider: {
order: ["anthropic"],
require_parameters: true,
data_collection: "future-policy",
sort: "future-sort",
max_price: { prompt: "0.50" },
},
reasoning: { effort: "future-effort", exclude: false },
plugins: [{ id: "future-plugin", enabled: true }],
web_search_options: { engine: "future-engine" },
debug: { echo_upstream_body: true },
user: "user_123",
},
},
})
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error OpenRouter usage must be boolean or an option record.
providerOptions: { openrouter: { usage: "yes" } },
})
@@ -0,0 +1,13 @@
import { LLM } from "../../src"
import { XAI } from "../../src/providers"
const model = XAI.provider.model("grok-4")
LLM.request({ model, prompt: "Hello", providerOptions: { xai: { reasoningEffort: "high" } } })
LLM.request({
model,
prompt: "Hello",
// @ts-expect-error xAI's OpenAI-compatible reasoning effort must be a string.
providerOptions: { xai: { reasoningEffort: true } },
})
+49 -4
View File
@@ -21,6 +21,8 @@ describe("provider package entrypoints", () => {
import("@opencode-ai/ai/providers/google-vertex/chat"), import("@opencode-ai/ai/providers/google-vertex/chat"),
import("@opencode-ai/ai/providers/google-vertex/responses"), import("@opencode-ai/ai/providers/google-vertex/responses"),
import("@opencode-ai/ai/providers/google-vertex/messages"), import("@opencode-ai/ai/providers/google-vertex/messages"),
import("@opencode-ai/ai/providers/openrouter"),
import("@opencode-ai/ai/providers/xai"),
]) ])
for (const module of modules) expect(module.model).toBeFunction() for (const module of modules) expect(module.model).toBeFunction()
@@ -29,6 +31,35 @@ describe("provider package entrypoints", () => {
expect(modules[12].model).toBe(modules[13].model) expect(modules[12].model).toBe(modules[13].model)
}) })
test("maps OpenRouter and xAI package settings onto executable models", async () => {
const OpenRouter = await import("@opencode-ai/ai/providers/openrouter")
const XAI = await import("@opencode-ai/ai/providers/xai")
const settings = {
apiKey: "fixture",
baseURL: "https://provider.example.test/v1",
headers: { "x-application": "opencode" },
body: { service_tier: "priority" },
limits: { context: 200_000, output: 64_000 },
}
const openrouter = OpenRouter.model("anthropic/claude-sonnet-4", {
...settings,
providerOptions: { openrouter: { usage: true } },
})
const xai = XAI.model("grok-4", {
...settings,
providerOptions: { xai: { reasoningEffort: "high" } },
})
for (const selected of [openrouter, xai]) {
expect(selected.route.endpoint.baseURL).toBe(settings.baseURL)
expect(selected.route.defaults.headers).toEqual(settings.headers)
expect(selected.route.defaults.http?.body).toEqual(settings.body)
expect(selected.route.defaults.limits).toEqual(settings.limits)
}
expect(openrouter.route.defaults.providerOptions).toEqual({ openrouter: { usage: true } })
expect(xai.route.defaults.providerOptions).toMatchObject({ xai: { reasoningEffort: "high", store: false } })
})
test("maps package settings onto the executable model", () => { test("maps package settings onto the executable model", () => {
const selected = model("gpt-5", { const selected = model("gpt-5", {
apiKey: "fixture", apiKey: "fixture",
@@ -59,7 +90,7 @@ describe("provider package entrypoints", () => {
headers: { "x-application": "opencode" }, headers: { "x-application": "opencode" },
body: { service_tier: "priority" }, body: { service_tier: "priority" },
limits: { context: 200_000, output: 64_000 }, limits: { context: 200_000, output: 64_000 },
providerOptions: { openai: { reasoningEffort: "low", store: true } }, providerOptions: { openresponses: { reasoningEffort: "low", store: true } },
}) })
expect(String(selected.provider)).toBe("example") expect(String(selected.provider)).toBe("example")
@@ -72,7 +103,7 @@ describe("provider package entrypoints", () => {
expect(selected.route.defaults.http?.body).toEqual({ service_tier: "priority" }) expect(selected.route.defaults.http?.body).toEqual({ service_tier: "priority" })
expect(selected.route.defaults.limits).toEqual({ context: 200_000, output: 64_000 }) expect(selected.route.defaults.limits).toEqual({ context: 200_000, output: 64_000 })
expect(selected.route.defaults.providerOptions).toEqual({ expect(selected.route.defaults.providerOptions).toEqual({
openai: { reasoningEffort: "low", store: true }, openresponses: { reasoningEffort: "low", store: true },
}) })
}) })
@@ -85,6 +116,7 @@ describe("provider package entrypoints", () => {
headers: { "x-application": "opencode" }, headers: { "x-application": "opencode" },
body: { metadata: { user_id: "user_1" } }, body: { metadata: { user_id: "user_1" } },
limits: { context: 200_000, output: 64_000 }, limits: { context: 200_000, output: 64_000 },
providerOptions: { anthropic: { effort: "low" } },
}) })
expect(String(selected.provider)).toBe("example") expect(String(selected.provider)).toBe("example")
@@ -96,6 +128,19 @@ describe("provider package entrypoints", () => {
expect(selected.route.defaults.headers).toEqual({ "x-application": "opencode" }) expect(selected.route.defaults.headers).toEqual({ "x-application": "opencode" })
expect(selected.route.defaults.http?.body).toEqual({ metadata: { user_id: "user_1" } }) expect(selected.route.defaults.http?.body).toEqual({ metadata: { user_id: "user_1" } })
expect(selected.route.defaults.limits).toEqual({ context: 200_000, output: 64_000 }) expect(selected.route.defaults.limits).toEqual({ context: 200_000, output: 64_000 })
expect(selected.route.defaults.providerOptions).toEqual({ anthropic: { effort: "low" } })
})
test("maps Anthropic provider options onto the executable model", async () => {
const Anthropic = await import("@opencode-ai/ai/providers/anthropic")
const selected = Anthropic.model("claude-sonnet-4-6", {
apiKey: "fixture",
providerOptions: { anthropic: { thinking: { type: "adaptive" } } },
})
expect(selected.route.defaults.providerOptions).toEqual({
anthropic: { thinking: { type: "adaptive" } },
})
}) })
test("requires an Anthropic-compatible base URL at runtime", async () => { test("requires an Anthropic-compatible base URL at runtime", async () => {
@@ -235,12 +280,12 @@ describe("provider package entrypoints", () => {
path: "/chat/completions", path: "/chat/completions",
}) })
expect(responses.route.id).toBe("google-vertex-responses") expect(responses.route.id).toBe("google-vertex-responses")
expect(responses.route.protocol).toBe("openai-responses") expect(responses.route.protocol).toBe("open-responses")
expect(responses.route.endpoint).toMatchObject({ expect(responses.route.endpoint).toMatchObject({
baseURL: "https://aiplatform.googleapis.com/v1/projects/vertex-project/locations/global/endpoints/openapi", baseURL: "https://aiplatform.googleapis.com/v1/projects/vertex-project/locations/global/endpoints/openapi",
path: "/responses", path: "/responses",
}) })
expect(responses.route.defaults.providerOptions).toEqual({ openai: { store: false } }) expect(responses.route.defaults.providerOptions).toEqual({ openresponses: { store: false } })
}) })
test("rejects conflicting Vertex auth settings at runtime", async () => { test("rejects conflicting Vertex auth settings at runtime", async () => {
@@ -1,6 +1,6 @@
import { describe, expect } from "bun:test" import { describe, expect } from "bun:test"
import { Effect } from "effect" import { Effect } from "effect"
import { CacheHint, LLM } from "../../src" import { CacheHint, LLM, LLMRequest, Message, ToolCallPart, ToolDefinition } from "../../src"
import { LLMClient } from "../../src/route" import { LLMClient } from "../../src/route"
import * as Anthropic from "../../src/providers/anthropic" import * as Anthropic from "../../src/providers/anthropic"
import { LARGE_CACHEABLE_SYSTEM } from "../recorded-scenarios" import { LARGE_CACHEABLE_SYSTEM } from "../recorded-scenarios"
@@ -24,6 +24,39 @@ const cacheRequest = LLM.request({
generation: { maxTokens: 16, temperature: 0 }, generation: { maxTokens: 16, temperature: 0 },
}) })
const lookup = ToolDefinition.make({
name: "lookup",
description: "Look up a fixture value.",
inputSchema: {
type: "object",
properties: { index: { type: "number" } },
required: ["index"],
additionalProperties: false,
},
})
const longToolTurn = [
Message.user("Run the fixture lookups."),
...Array.from({ length: 11 }, (_, index) => {
const id = `lookup_${index}`
return [
Message.assistant(ToolCallPart.make({ id, name: lookup.name, input: { index } })),
Message.tool({
id,
name: lookup.name,
result: `Fixture result ${index}. `.repeat(80),
}),
]
}).flat(),
]
const longToolTurnRequest = LLM.request({
id: "recorded_anthropic_cache_long_tool_turn",
model,
system: LARGE_CACHEABLE_SYSTEM,
messages: longToolTurn,
tools: [lookup],
generation: { maxTokens: 16, temperature: 0 },
})
const recorded = recordedTests({ const recorded = recordedTests({
prefix: "anthropic-messages-cache", prefix: "anthropic-messages-cache",
provider: "anthropic", provider: "anthropic",
@@ -50,4 +83,28 @@ describe("Anthropic Messages cache recorded", () => {
expect(second.usage?.cacheReadInputTokens ?? 0).toBeGreaterThan(0) expect(second.usage?.cacheReadInputTokens ?? 0).toBeGreaterThan(0)
}), }),
) )
recorded.effect.with("keeps a long tool turn inside the cache lookback", { tags: ["cache", "tool"] }, () =>
Effect.gen(function* () {
const first = yield* LLMClient.generate(longToolTurnRequest)
const firstRead = first.usage?.cacheReadInputTokens ?? 0
const firstWrite = first.usage?.cacheWriteInputTokens ?? 0
const firstCached = firstRead + firstWrite
// The prefix may already be warm when recording, so either a read or a
// write establishes that Anthropic recognized the cache boundary.
expect(firstCached).toBeGreaterThan(0)
const second = yield* LLMClient.generate(
LLMRequest.update(longToolTurnRequest, {
messages: [
...longToolTurn,
Message.assistant("The fixture lookups are complete."),
Message.user("Reply exactly: OK"),
],
}),
)
expect(second.usage?.cacheReadInputTokens ?? 0).toBeGreaterThanOrEqual(firstCached)
expect(second.usage?.cacheWriteInputTokens ?? 0).toBeLessThan(firstCached)
}),
)
}) })
@@ -1,8 +1,9 @@
import { describe, expect } from "bun:test" import { describe, expect } from "bun:test"
import { Effect } from "effect" import { Effect } from "effect"
import { HttpClientRequest } from "effect/unstable/http" import { HttpClientRequest } from "effect/unstable/http"
import { CacheHint, LLM, LLMError, Message, ToolCallPart, Usage } from "../../src" import { CacheHint, LLM, LLMError, LLMRequest, Message, ToolCallPart, ToolDefinition, Usage } from "../../src"
import { Auth, LLMClient } from "../../src/route" import { Auth, LLMClient } from "../../src/route"
import { compileRequest } from "../../src/route/client"
import * as AnthropicMessages from "../../src/protocols/anthropic-messages" import * as AnthropicMessages from "../../src/protocols/anthropic-messages"
import { continuationRequest, nativeAnthropicMessagesContinuation } from "../continuation-scenarios" import { continuationRequest, nativeAnthropicMessagesContinuation } from "../continuation-scenarios"
import { it } from "../lib/effect" import { it } from "../lib/effect"
@@ -44,7 +45,7 @@ const expectToolResult = (body: AnthropicMessages.AnthropicMessagesBody): Anthro
describe("Anthropic Messages route", () => { describe("Anthropic Messages route", () => {
it.effect("prepares Anthropic Messages target", () => it.effect("prepares Anthropic Messages target", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare(request) const prepared = yield* compileRequest(request)
expect(prepared.body).toEqual({ expect(prepared.body).toEqual({
model: "claude-sonnet-4-5", model: "claude-sonnet-4-5",
@@ -59,8 +60,8 @@ describe("Anthropic Messages route", () => {
it.effect("lowers adaptive thinking settings with effort", () => it.effect("lowers adaptive thinking settings with effort", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.updateRequest(request, { LLMRequest.update(request, {
providerOptions: { providerOptions: {
anthropic: { thinking: { type: "adaptive", display: "summarized" }, effort: "low" }, anthropic: { thinking: { type: "adaptive", display: "summarized" }, effort: "low" },
}, },
@@ -74,9 +75,45 @@ describe("Anthropic Messages route", () => {
}), }),
) )
it.effect("normalizes enabled and disabled thinking settings", () =>
Effect.gen(function* () {
const enabled = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { anthropic: { thinking: { type: "enabled", budgetTokens: 1_024 } } },
}),
)
const legacy = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { anthropic: { thinking: { type: "enabled", budget_tokens: 2_048 } } },
}),
)
const disabled = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { anthropic: { thinking: { type: "disabled" } } },
}),
)
expect(enabled.body.thinking).toEqual({ type: "enabled", budget_tokens: 1_024 })
expect(legacy.body.thinking).toEqual({ type: "enabled", budget_tokens: 2_048 })
expect(disabled.body.thinking).toEqual({ type: "disabled" })
}),
)
it.effect("rejects enabled thinking without a budget", () =>
Effect.gen(function* () {
const error = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { anthropic: { thinking: { type: "enabled" } } },
}),
).pipe(Effect.flip)
expect(error.message).toContain("Anthropic thinking provider option requires budgetTokens")
}),
)
it.effect("lowers chronological system updates natively for Claude Opus 4.8 with cache hints", () => it.effect("lowers chronological system updates natively for Claude Opus 4.8 with cache hints", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: opus48, model: opus48,
messages: [ messages: [
@@ -101,7 +138,7 @@ describe("Anthropic Messages route", () => {
it.effect("lowers chronological system updates to wrapped user text for unsupported Anthropic models", () => it.effect("lowers chronological system updates to wrapped user text for unsupported Anthropic models", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -128,7 +165,7 @@ describe("Anthropic Messages route", () => {
it.effect("rejects non-text chronological system update content before send", () => it.effect("rejects non-text chronological system update content before send", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
model: opus48, model: opus48,
messages: [ messages: [
@@ -145,7 +182,7 @@ describe("Anthropic Messages route", () => {
it.effect("falls back for unsupported native chronological system update placement", () => it.effect("falls back for unsupported native chronological system update placement", () =>
Effect.gen(function* () { Effect.gen(function* () {
expect( expect(
(yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( (yield* compileRequest(
LLM.request({ LLM.request({
model: opus48, model: opus48,
messages: [Message.assistant("Plain."), Message.system("After plain assistant.")], messages: [Message.assistant("Plain."), Message.system("After plain assistant.")],
@@ -160,12 +197,11 @@ describe("Anthropic Messages route", () => {
}, },
]) ])
expect( expect(
(yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( (yield* compileRequest(LLM.request({ model: opus48, messages: [Message.system("First.")], cache: "none" })))
LLM.request({ model: opus48, messages: [Message.system("First.")], cache: "none" }), .body.messages,
)).body.messages,
).toEqual([{ role: "user", content: [{ type: "text", text: "<system-update>\nFirst.\n</system-update>" }] }]) ).toEqual([{ role: "user", content: [{ type: "text", text: "<system-update>\nFirst.\n</system-update>" }] }])
expect( expect(
(yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( (yield* compileRequest(
LLM.request({ LLM.request({
model: opus48, model: opus48,
messages: [Message.user("Before."), Message.system("One."), Message.system("Two.")], messages: [Message.user("Before."), Message.system("One."), Message.system("Two.")],
@@ -187,7 +223,7 @@ describe("Anthropic Messages route", () => {
it.effect("rejects a system update between a local tool call and its result", () => it.effect("rejects a system update between a local tool call and its result", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
model: opus48, model: opus48,
messages: [ messages: [
@@ -206,7 +242,7 @@ describe("Anthropic Messages route", () => {
it.effect("prepares tool call and tool result messages", () => it.effect("prepares tool call and tool result messages", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_result", id: "req_tool_result",
model, model,
@@ -235,11 +271,39 @@ describe("Anthropic Messages route", () => {
}), }),
) )
it.effect("keeps tools and sends tool_choice none", () =>
Effect.gen(function* () {
const prepared = yield* compileRequest(
LLM.request({
id: "req_tool_choice_none",
model,
tools: [{ name: "lookup", description: "Look things up", inputSchema: { type: "object", properties: {} } }],
messages: [
Message.user("What is the weather?"),
Message.assistant([ToolCallPart.make({ id: "call_1", name: "lookup", input: { query: "weather" } })]),
Message.tool({ id: "call_1", name: "lookup", result: { forecast: "sunny" } }),
],
toolChoice: "none",
cache: "none",
}),
)
expect(prepared.body.tools).toEqual([
{
name: "lookup",
description: "Look things up",
input_schema: { type: "object", properties: {} },
},
])
expect(prepared.body.tool_choice).toEqual({ type: "none" })
}),
)
// Regression: read tool results must stay structured so base64 media data is // Regression: read tool results must stay structured so base64 media data is
// not JSON-stringified into `tool_result.content`. // not JSON-stringified into `tool_result.content`.
it.effect("lowers media tool-result content as structured blocks", () => it.effect("lowers media tool-result content as structured blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_result_image", id: "req_tool_result_image",
model, model,
@@ -271,7 +335,7 @@ describe("Anthropic Messages route", () => {
it.effect("lowers single-image tool-result content as a structured image block", () => it.effect("lowers single-image tool-result content as a structured image block", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_result_image_only", id: "req_tool_result_image_only",
model, model,
@@ -296,7 +360,7 @@ describe("Anthropic Messages route", () => {
it.effect("rejects unsupported media in tool-result content with a clear error", () => it.effect("rejects unsupported media in tool-result content with a clear error", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_result_unsupported_media", id: "req_tool_result_unsupported_media",
model, model,
@@ -320,7 +384,7 @@ describe("Anthropic Messages route", () => {
it.effect("prepares the composed native continuation request", () => it.effect("prepares the composed native continuation request", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<AnthropicMessages.AnthropicMessagesBody>( const prepared = yield* compileRequest(
continuationRequest({ continuationRequest({
id: "req_native_continuation_anthropic", id: "req_native_continuation_anthropic",
model, model,
@@ -364,7 +428,7 @@ describe("Anthropic Messages route", () => {
it.effect("lowers preserved Anthropic reasoning signature metadata", () => it.effect("lowers preserved Anthropic reasoning signature metadata", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -381,6 +445,34 @@ describe("Anthropic Messages route", () => {
}), }),
) )
it.effect("round-trips redacted thinking as redacted_thinking blocks", () =>
Effect.gen(function* () {
const prepared = yield* compileRequest(
LLM.request({
model,
messages: [
Message.assistant([
{ type: "reasoning", text: "", providerMetadata: { anthropic: { redactedData: "opaque_1" } } },
{ type: "reasoning", text: "visible", providerMetadata: { anthropic: { signature: "sig_1" } } },
]),
],
}),
)
expect(prepared.body).toMatchObject({
messages: [
{
role: "assistant",
content: [
{ type: "redacted_thinking", data: "opaque_1" },
{ type: "thinking", thinking: "visible", signature: "sig_1" },
],
},
],
})
}),
)
it.effect("parses text, reasoning, and usage stream fixtures", () => it.effect("parses text, reasoning, and usage stream fixtures", () =>
Effect.gen(function* () { Effect.gen(function* () {
const body = sseEvents( const body = sseEvents(
@@ -414,18 +506,349 @@ describe("Anthropic Messages route", () => {
expect(response.events.find((event) => event.type === "reasoning-end")).toMatchObject({ expect(response.events.find((event) => event.type === "reasoning-end")).toMatchObject({
providerMetadata: { anthropic: { signature: "sig_1" } }, providerMetadata: { anthropic: { signature: "sig_1" } },
}) })
expect(response.events.find((event) => event.type === "reasoning-delta" && event.text === "")).toBeUndefined()
expect(response.message.content).toEqual([ expect(response.message.content).toEqual([
{ type: "text", text: "Hello!" }, { type: "text", text: "Hello!" },
{ type: "reasoning", text: "thinking", providerMetadata: { anthropic: { signature: "sig_1" } } }, { type: "reasoning", text: "thinking", providerMetadata: { anthropic: { signature: "sig_1" } } },
]) ])
expect(response.events.at(-1)).toMatchObject({ expect(response.events.at(-1)).toMatchObject({
type: "finish", type: "finish",
reason: "stop", reason: { normalized: "stop", raw: "end_turn" },
providerMetadata: { anthropic: { stopSequence: "\n\nHuman:" } }, providerMetadata: { anthropic: { stopSequence: "\n\nHuman:" } },
}) })
}), }),
) )
it.effect("requires message_stop before completing a streamed message", () =>
Effect.gen(function* () {
const error = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "content_block_start", index: 0, content_block: { type: "text", text: "" } },
{ type: "content_block_delta", index: 0, delta: { type: "text_delta", text: "Hello" } },
{ type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
),
),
),
Effect.flip,
)
expect(error.reason).toMatchObject({
_tag: "InvalidProviderOutput",
message: "Provider stream ended without a terminal finish event",
})
}),
)
it.effect("maps thinking tokens and preserves unknown Anthropic usage fields", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{
type: "message_start",
message: {
usage: {
input_tokens: 5,
cache_read_input_tokens: 2,
service_tier: "standard",
cache_creation: { ephemeral_5m_input_tokens: 1 },
server_tool_use: { web_search_requests: 1, start_counter: 2 },
output_tokens_details: { thinking_tokens: 3, start_detail: "preserved" },
},
},
},
{
type: "message_delta",
delta: { stop_reason: "end_turn" },
usage: {
output_tokens: 8,
server_tool_use: { web_search_requests: 2, terminal_counter: 3 },
output_tokens_details: { terminal_detail: "preserved" },
future_terminal: { requests: 4 },
},
},
{ type: "message_stop" },
),
),
),
)
expect(response.usage).toMatchObject({
inputTokens: 7,
outputTokens: 8,
reasoningTokens: 3,
totalTokens: 15,
providerMetadata: {
anthropic: {
input_tokens: 5,
cache_read_input_tokens: 2,
service_tier: "standard",
cache_creation: { ephemeral_5m_input_tokens: 1 },
server_tool_use: { web_search_requests: 2, start_counter: 2, terminal_counter: 3 },
output_tokens: 8,
output_tokens_details: {
thinking_tokens: 3,
start_detail: "preserved",
terminal_detail: "preserved",
},
future_terminal: { requests: 4 },
},
},
})
}),
)
it.effect("round-trips omitted thinking carried only by a signature delta", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{
type: "content_block_start",
index: 0,
content_block: { type: "thinking", thinking: "", signature: "" },
},
{ type: "content_block_delta", index: 0, delta: { type: "signature_delta", signature: "sig_1" } },
{ type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
expect(response.message.content).toEqual([
{ type: "reasoning", text: "", providerMetadata: { anthropic: { signature: "sig_1" } } },
])
const prepared = yield* compileRequest(LLM.request({ model, messages: [response.message], cache: "none" }))
expect(prepared.body.messages).toEqual([
{ role: "assistant", content: [{ type: "thinking", thinking: "", signature: "sig_1" }] },
])
}),
)
it.effect("retains a thinking signature supplied in content_block_start", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{
type: "content_block_start",
index: 0,
content_block: { type: "thinking", thinking: "", signature: "sig_1" },
},
{ type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
expect(response.message.content).toEqual([
{ type: "reasoning", text: "", providerMetadata: { anthropic: { signature: "sig_1" } } },
])
expect(response.events.find((event) => event.type === "reasoning-end")).toMatchObject({
providerMetadata: { anthropic: { signature: "sig_1" } },
})
}),
)
it.effect("retains complete tool input from content_block_start", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{
type: "content_block_start",
index: 0,
content_block: { type: "tool_use", id: "call_1", name: "lookup", input: { query: "weather" } },
},
{ type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "tool_use" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
expect(response.toolCalls).toMatchObject([
{ id: "call_1", name: "lookup", input: { query: "weather" } },
])
}),
)
it.effect("retains empty text blocks", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "content_block_start", index: 0, content_block: { type: "text", text: "" } },
{ type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
expect(response.message.content).toEqual([{ type: "text", text: "" }])
}),
)
it.effect("parses redacted thinking into empty reasoning with redactedData metadata", () =>
Effect.gen(function* () {
const body = sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "content_block_start", index: 0, content_block: { type: "redacted_thinking", data: "opaque_1" } },
{ type: "content_block_stop", index: 0 },
{ type: "content_block_start", index: 1, content_block: { type: "text", text: "" } },
{ type: "content_block_delta", index: 1, delta: { type: "text_delta", text: "Hello" } },
{ type: "content_block_stop", index: 1 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 2 } },
{ type: "message_stop" },
)
const response = yield* LLMClient.generate(request).pipe(Effect.provide(fixedResponse(body)))
expect(response.events.find((event) => event.type === "reasoning-start")).toMatchObject({
providerMetadata: { anthropic: { redactedData: "opaque_1" } },
})
expect(response.message.content).toEqual([
{ type: "reasoning", text: "", providerMetadata: { anthropic: { redactedData: "opaque_1" } } },
{ type: "text", text: "Hello" },
])
}),
)
it.effect("round-trips streamed redacted thinking with tool use into a continuation request", () =>
Effect.gen(function* () {
// Anthropic types `redacted_thinking.data` as an opaque string. Its
// contents are provider-owned and must be replayed without inspection.
const redactedData = "cmVkYWN0ZWQtdGhpbmtpbmc="
const response = yield* LLMClient.generate(
LLMRequest.update(request, {
tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}),
).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{
type: "content_block_start",
index: 0,
content_block: { type: "redacted_thinking", data: redactedData },
},
{ type: "content_block_stop", index: 0 },
{
type: "content_block_start",
index: 1,
content_block: { type: "tool_use", id: "call_1", name: "lookup" },
},
{
type: "content_block_delta",
index: 1,
delta: { type: "input_json_delta", partial_json: '{"query":"weather"}' },
},
{ type: "content_block_stop", index: 1 },
{ type: "message_delta", delta: { stop_reason: "tool_use" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
const prepared = yield* compileRequest(
LLM.request({
model,
messages: [
Message.user("Say hello."),
response.message,
Message.tool({ id: "call_1", name: "lookup", result: "sunny", resultType: "text" }),
],
tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
cache: "none",
}),
)
expect(prepared.body.messages).toEqual([
{ role: "user", content: [{ type: "text", text: "Say hello." }] },
{
role: "assistant",
content: [
{ type: "redacted_thinking", data: redactedData },
{ type: "tool_use", id: "call_1", name: "lookup", input: { query: "weather" } },
],
},
{
role: "user",
content: [
{
type: "tool_result",
tool_use_id: "call_1",
content: "sunny",
is_error: undefined,
cache_control: undefined,
},
],
},
])
}),
)
it.effect("maps context-window truncation to length", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{
type: "message_delta",
delta: { stop_reason: "model_context_window_exceeded" },
usage: { output_tokens: 1 },
},
{ type: "message_stop" },
),
),
),
)
expect(response.finishReason).toEqual({ normalized: "length", raw: "model_context_window_exceeded" })
}),
)
it.effect("preserves pause_turn while normalizing it to stop", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(
sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "message_delta", delta: { stop_reason: "pause_turn" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
),
),
),
)
expect(response.finishReason).toEqual({ normalized: "stop", raw: "pause_turn" })
}),
)
it.effect("assembles streamed tool call input", () => it.effect("assembles streamed tool call input", () =>
Effect.gen(function* () { Effect.gen(function* () {
const body = sseEvents( const body = sseEvents(
@@ -435,10 +858,11 @@ describe("Anthropic Messages route", () => {
{ type: "content_block_delta", index: 0, delta: { type: "input_json_delta", partial_json: ':"weather"}' } }, { type: "content_block_delta", index: 0, delta: { type: "input_json_delta", partial_json: ':"weather"}' } },
{ type: "content_block_stop", index: 0 }, { type: "content_block_stop", index: 0 },
{ type: "message_delta", delta: { stop_reason: "tool_use" }, usage: { output_tokens: 1 } }, { type: "message_delta", delta: { stop_reason: "tool_use" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
) )
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
const usage = new Usage({ const usage = new Usage({
@@ -475,10 +899,16 @@ describe("Anthropic Messages route", () => {
providerExecuted: undefined, providerExecuted: undefined,
providerMetadata: undefined, providerMetadata: undefined,
}, },
{ type: "step-finish", index: 0, reason: "tool-calls", usage, providerMetadata: undefined }, {
type: "step-finish",
index: 0,
reason: { normalized: "tool-calls", raw: "tool_use" },
usage,
providerMetadata: undefined,
},
{ {
type: "finish", type: "finish",
reason: "tool-calls", reason: { normalized: "tool-calls", raw: "tool_use" },
providerMetadata: undefined, providerMetadata: undefined,
usage, usage,
}, },
@@ -614,10 +1044,13 @@ describe("Anthropic Messages route", () => {
{ type: "content_block_delta", index: 2, delta: { type: "text_delta", text: "Found it." } }, { type: "content_block_delta", index: 2, delta: { type: "text_delta", text: "Found it." } },
{ type: "content_block_stop", index: 2 }, { type: "content_block_stop", index: 2 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 8 } }, { type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 8 } },
{ type: "message_stop" },
) )
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "web_search", description: "Web search", inputSchema: { type: "object" } }], tools: [
ToolDefinition.make({ name: "web_search", description: "Web search", inputSchema: { type: "object" } }),
],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
@@ -636,10 +1069,20 @@ describe("Anthropic Messages route", () => {
name: "web_search", name: "web_search",
result: { type: "json", value: [{ type: "web_search_result", url: "https://example.com", title: "Example" }] }, result: { type: "json", value: [{ type: "web_search_result", url: "https://example.com", title: "Example" }] },
providerExecuted: true, providerExecuted: true,
providerMetadata: { anthropic: { blockType: "web_search_tool_result" } }, // The complete payload rides in provider metadata as irreducible replay
// state for later stateless requests.
providerMetadata: {
anthropic: {
blockType: "web_search_tool_result",
result: [{ type: "web_search_result", url: "https://example.com", title: "Example" }],
},
},
}) })
expect(response.text).toBe("Found it.") expect(response.text).toBe("Found it.")
expect(response.events.at(-1)).toMatchObject({ type: "finish", reason: "stop" }) expect(response.events.at(-1)).toMatchObject({
type: "finish",
reason: { normalized: "stop", raw: "end_turn" },
})
}), }),
) )
@@ -665,10 +1108,13 @@ describe("Anthropic Messages route", () => {
}, },
{ type: "content_block_stop", index: 1 }, { type: "content_block_stop", index: 1 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } }, { type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
{ type: "message_stop" },
) )
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "web_search", description: "Web search", inputSchema: { type: "object" } }], tools: [
ToolDefinition.make({ name: "web_search", description: "Web search", inputSchema: { type: "object" } }),
],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
@@ -685,7 +1131,7 @@ describe("Anthropic Messages route", () => {
it.effect("round-trips provider-executed assistant content into server tool blocks", () => it.effect("round-trips provider-executed assistant content into server tool blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_round_trip", id: "req_round_trip",
model, model,
@@ -736,7 +1182,7 @@ describe("Anthropic Messages route", () => {
it.effect("rejects round-trip for unknown server tool names", () => it.effect("rejects round-trip for unknown server tool names", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_unknown_server_tool", id: "req_unknown_server_tool",
model, model,
@@ -784,7 +1230,10 @@ describe("Anthropic Messages route", () => {
content: [ content: [
{ type: "text", text: "What is in this image?" }, { type: "text", text: "What is in this image?" },
{ type: "image", source: { type: "base64", media_type: "image/png", data: "AAECAw==" } }, { type: "image", source: { type: "base64", media_type: "image/png", data: "AAECAw==" } },
{ type: "document", source: { type: "base64", media_type: "application/pdf", data: "JVBERi0xLjQ=" } }, {
type: "document",
source: { type: "base64", media_type: "application/pdf", data: "JVBERi0xLjQ=" },
},
], ],
}, },
], ],
@@ -810,7 +1259,7 @@ describe("Anthropic Messages route", () => {
it.effect("maps ttlSeconds >= 3600 to cache_control ttl: '1h'", () => it.effect("maps ttlSeconds >= 3600 to cache_control ttl: '1h'", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
system: { type: "text", text: "system", cache: new CacheHint({ type: "ephemeral", ttlSeconds: 3600 }) }, system: { type: "text", text: "system", cache: new CacheHint({ type: "ephemeral", ttlSeconds: 3600 }) },
@@ -826,7 +1275,7 @@ describe("Anthropic Messages route", () => {
it.effect("emits cache_control on tool definitions and tool-result blocks", () => it.effect("emits cache_control on tool definitions and tool-result blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
tools: [ tools: [
@@ -867,7 +1316,7 @@ describe("Anthropic Messages route", () => {
it.effect("drops cache_control breakpoints past the 4-per-request cap", () => it.effect("drops cache_control breakpoints past the 4-per-request cap", () =>
Effect.gen(function* () { Effect.gen(function* () {
const hint = new CacheHint({ type: "ephemeral" }) const hint = new CacheHint({ type: "ephemeral" })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
system: [ system: [
@@ -893,7 +1342,7 @@ describe("Anthropic Messages route", () => {
it.effect("spends breakpoint budget on tools before system before messages", () => it.effect("spends breakpoint budget on tools before system before messages", () =>
Effect.gen(function* () { Effect.gen(function* () {
const hint = new CacheHint({ type: "ephemeral" }) const hint = new CacheHint({ type: "ephemeral" })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
tools: [ tools: [
@@ -2,8 +2,18 @@ import { EventStreamCodec } from "@smithy/eventstream-codec"
import { fromUtf8, toUtf8 } from "@smithy/util-utf8" import { fromUtf8, toUtf8 } from "@smithy/util-utf8"
import { describe, expect } from "bun:test" import { describe, expect } from "bun:test"
import { Effect } from "effect" import { Effect } from "effect"
import { CacheHint, LLM, Message, ToolCallPart, ToolChoice } from "../../src" import {
CacheHint,
GenerationOptions,
LLM,
LLMRequest,
Message,
ToolCallPart,
ToolChoice,
ToolDefinition,
} from "../../src"
import { LLMClient } from "../../src/route" import { LLMClient } from "../../src/route"
import { compileRequest } from "../../src/route/client"
import { AmazonBedrock } from "../../src/providers" import { AmazonBedrock } from "../../src/providers"
import * as BedrockConverse from "../../src/protocols/bedrock-converse" import * as BedrockConverse from "../../src/protocols/bedrock-converse"
import { it } from "../lib/effect" import { it } from "../lib/effect"
@@ -34,6 +44,26 @@ const eventFrame = (type: string, payload: object) =>
body: utf8Encoder.encode(JSON.stringify(payload)), body: utf8Encoder.encode(JSON.stringify(payload)),
}) })
const exceptionFrame = (type: string, payload: object) =>
codec.encode({
headers: {
":message-type": { type: "string", value: "exception" },
":exception-type": { type: "string", value: type },
":content-type": { type: "string", value: "application/json" },
},
body: utf8Encoder.encode(JSON.stringify(payload)),
})
const errorFrame = (code: string, message: string) =>
codec.encode({
headers: {
":message-type": { type: "string", value: "error" },
":error-code": { type: "string", value: code },
":error-message": { type: "string", value: message },
},
body: new Uint8Array(),
})
const concat = (frames: ReadonlyArray<Uint8Array>) => { const concat = (frames: ReadonlyArray<Uint8Array>) => {
const total = frames.reduce((sum, frame) => sum + frame.length, 0) const total = frames.reduce((sum, frame) => sum + frame.length, 0)
const out = new Uint8Array(total) const out = new Uint8Array(total)
@@ -72,7 +102,7 @@ const baseRequest = LLM.request({
describe("Bedrock Converse route", () => { describe("Bedrock Converse route", () => {
it.effect("prepares Converse target with system, inference config, and messages", () => it.effect("prepares Converse target with system, inference config, and messages", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare(baseRequest) const prepared = yield* compileRequest(baseRequest)
expect(prepared.body).toEqual({ expect(prepared.body).toEqual({
modelId: "anthropic.claude-3-5-sonnet-20240620-v1:0", modelId: "anthropic.claude-3-5-sonnet-20240620-v1:0",
@@ -85,8 +115,10 @@ describe("Bedrock Converse route", () => {
it.effect("passes topK through additionalModelRequestFields as top_k", () => it.effect("passes topK through additionalModelRequestFields as top_k", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>( const prepared = yield* compileRequest(
LLM.updateRequest(baseRequest, { generation: { maxTokens: 64, temperature: 0, topK: 40 } }), LLMRequest.update(baseRequest, {
generation: GenerationOptions.make({ maxTokens: 64, temperature: 0, topK: 40 }),
}),
) )
// Converse's inferenceConfig has no topK; Anthropic/Nova read it from // Converse's inferenceConfig has no topK; Anthropic/Nova read it from
@@ -98,14 +130,14 @@ describe("Bedrock Converse route", () => {
it.effect("omits additionalModelRequestFields when topK is unset", () => it.effect("omits additionalModelRequestFields when topK is unset", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>(baseRequest) const prepared = yield* compileRequest(baseRequest)
expect(prepared.body.additionalModelRequestFields).toBeUndefined() expect(prepared.body.additionalModelRequestFields).toBeUndefined()
}), }),
) )
it.effect("lowers chronological system updates to wrapped user text in order", () => it.effect("lowers chronological system updates to wrapped user text in order", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [Message.user("Before."), Message.system("Update."), Message.assistant("After.")], messages: [Message.user("Before."), Message.system("Update."), Message.assistant("After.")],
@@ -122,14 +154,14 @@ describe("Bedrock Converse route", () => {
it.effect("prepares tool config with toolSpec and toolChoice", () => it.effect("prepares tool config with toolSpec and toolChoice", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.updateRequest(baseRequest, { LLMRequest.update(baseRequest, {
tools: [ tools: [
{ ToolDefinition.make({
name: "lookup", name: "lookup",
description: "Lookup data", description: "Lookup data",
inputSchema: { type: "object", properties: { query: { type: "string" } }, required: ["query"] }, inputSchema: { type: "object", properties: { query: { type: "string" } }, required: ["query"] },
}, }),
], ],
toolChoice: ToolChoice.make({ type: "required" }), toolChoice: ToolChoice.make({ type: "required" }),
}), }),
@@ -154,9 +186,39 @@ describe("Bedrock Converse route", () => {
}), }),
) )
it.effect("keeps tools and omits the unsupported choice when tool choice is none", () =>
Effect.gen(function* () {
const prepared = yield* compileRequest(
LLMRequest.update(baseRequest, {
tools: [
ToolDefinition.make({
name: "lookup",
description: "Lookup data",
inputSchema: { type: "object", properties: { query: { type: "string" } } },
}),
],
toolChoice: ToolChoice.make({ type: "none" }),
}),
)
expect(prepared.body.toolConfig).toMatchObject({
tools: [
{
toolSpec: {
name: "lookup",
description: "Lookup data",
inputSchema: { json: { type: "object", properties: { query: { type: "string" } } } },
},
},
],
})
expect(prepared.body.toolConfig?.toolChoice).toBeUndefined()
}),
)
it.effect("lowers assistant tool-call + tool-result message history", () => it.effect("lowers assistant tool-call + tool-result message history", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_history", id: "req_history",
model, model,
@@ -195,7 +257,7 @@ describe("Bedrock Converse route", () => {
it.effect("lowers image content in tool-result messages", () => it.effect("lowers image content in tool-result messages", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_image", id: "req_tool_image",
model, model,
@@ -260,7 +322,10 @@ describe("Bedrock Converse route", () => {
// `metadata` (carries usage). We consolidate them into a single // `metadata` (carries usage). We consolidate them into a single
// terminal `finish` event with both. // terminal `finish` event with both.
expect(finishes).toHaveLength(1) expect(finishes).toHaveLength(1)
expect(finishes[0]).toMatchObject({ type: "finish", reason: "stop" }) expect(finishes[0]).toMatchObject({
type: "finish",
reason: { normalized: "stop", raw: "end_turn" },
})
expect(response.usage).toMatchObject({ expect(response.usage).toMatchObject({
inputTokens: 5, inputTokens: 5,
outputTokens: 2, outputTokens: 2,
@@ -269,6 +334,23 @@ describe("Bedrock Converse route", () => {
}), }),
) )
it.effect("maps truncation and malformed output stop reasons", () =>
Effect.gen(function* () {
const reasons = [
["model_context_window_exceeded", "length"],
["malformed_model_output", "error"],
["malformed_tool_use", "error"],
] as const
for (const [raw, normalized] of reasons) {
const response = yield* LLMClient.generate(baseRequest).pipe(
Effect.provide(fixedBytes(eventStreamBody(["messageStop", { stopReason: raw }]))),
)
expect(response.finishReason).toEqual({ normalized, raw })
}
}),
)
it.effect("adds cache reads and writes to Bedrock input usage", () => it.effect("adds cache reads and writes to Bedrock input usage", () =>
Effect.gen(function* () { Effect.gen(function* () {
const body = eventStreamBody( const body = eventStreamBody(
@@ -302,6 +384,19 @@ describe("Bedrock Converse route", () => {
}), }),
) )
it.effect("preserves usage across later metadata events without usage", () =>
Effect.gen(function* () {
const body = eventStreamBody(
["messageStop", { stopReason: "end_turn" }],
["metadata", { usage: { inputTokens: 5, outputTokens: 2, totalTokens: 7 } }],
["metadata", { metrics: { latencyMs: 100 } }],
)
const response = yield* LLMClient.generate(baseRequest).pipe(Effect.provide(fixedBytes(body)))
expect(response.usage).toMatchObject({ inputTokens: 5, outputTokens: 2, totalTokens: 7 })
}),
)
it.effect("assembles streamed tool call input", () => it.effect("assembles streamed tool call input", () =>
Effect.gen(function* () { Effect.gen(function* () {
const body = eventStreamBody( const body = eventStreamBody(
@@ -319,8 +414,8 @@ describe("Bedrock Converse route", () => {
["messageStop", { stopReason: "tool_use" }], ["messageStop", { stopReason: "tool_use" }],
) )
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(baseRequest, { LLMRequest.update(baseRequest, {
tools: [{ name: "lookup", description: "Lookup", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup", inputSchema: { type: "object" } })],
}), }),
).pipe(Effect.provide(fixedBytes(body))) ).pipe(Effect.provide(fixedBytes(body)))
@@ -332,7 +427,10 @@ describe("Bedrock Converse route", () => {
{ type: "tool-input-delta", id: "tool_1", name: "lookup", text: '{"query"' }, { type: "tool-input-delta", id: "tool_1", name: "lookup", text: '{"query"' },
{ type: "tool-input-delta", id: "tool_1", name: "lookup", text: ':"weather"}' }, { type: "tool-input-delta", id: "tool_1", name: "lookup", text: ':"weather"}' },
]) ])
expect(response.events.at(-1)).toMatchObject({ type: "finish", reason: "tool-calls" }) expect(response.events.at(-1)).toMatchObject({
type: "finish",
reason: { normalized: "tool-calls", raw: "tool_use" },
})
}), }),
) )
@@ -358,7 +456,7 @@ describe("Bedrock Converse route", () => {
name: "lookup", name: "lookup",
raw: '{"query":"partial', raw: '{"query":"partial',
}) })
expect(response.finishReason).toBe("tool-calls") expect(response.finishReason).toEqual({ normalized: "tool-calls", raw: "end_turn" })
}), }),
) )
@@ -394,7 +492,7 @@ describe("Bedrock Converse route", () => {
providerMetadata: { bedrock: { signature: "sig_1" } }, providerMetadata: { bedrock: { signature: "sig_1" } },
}) })
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -414,12 +512,168 @@ describe("Bedrock Converse route", () => {
}), }),
) )
it.effect("classifies throttlingException as a rate limit", () => it.effect("preserves reasoning signatures when contentBlockStop is missing", () =>
Effect.gen(function* () {
const response = yield* LLMClient.generate(baseRequest).pipe(
Effect.provide(
fixedBytes(
eventStreamBody(
["messageStart", { role: "assistant" }],
[
"contentBlockDelta",
{ contentBlockIndex: 0, delta: { reasoningContent: { text: "Let me think." } } },
],
[
"contentBlockDelta",
{ contentBlockIndex: 0, delta: { reasoningContent: { signature: "sig_1" } } },
],
["messageStop", { stopReason: "end_turn" }],
),
),
),
)
expect(response.events.find((event) => event.type === "reasoning-delta" && event.text === "")).toEqual({
type: "reasoning-delta",
id: "reasoning-0",
text: "",
providerMetadata: { bedrock: { signature: "sig_1" } },
})
expect(response.message.content).toEqual([
{
type: "reasoning",
text: "Let me think.",
providerMetadata: { bedrock: { signature: "sig_1" } },
},
])
const prepared = yield* compileRequest(LLM.request({ model, messages: [response.message], cache: "none" }))
expect(prepared.body.messages).toEqual([
{
role: "assistant",
content: [{ reasoningContent: { reasoningText: { text: "Let me think.", signature: "sig_1" } } }],
},
])
}),
)
it.effect("preserves signature-only reasoning blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const body = eventStreamBody( const body = eventStreamBody(
["messageStart", { role: "assistant" }], ["messageStart", { role: "assistant" }],
["throttlingException", { message: "Slow down" }], [
"contentBlockDelta",
{ contentBlockIndex: 0, delta: { reasoningContent: { signature: "sig_1" } } },
],
["contentBlockStop", { contentBlockIndex: 0 }],
["messageStop", { stopReason: "end_turn" }],
) )
const response = yield* LLMClient.generate(baseRequest).pipe(Effect.provide(fixedBytes(body)))
expect(response.message.content).toEqual([
{ type: "reasoning", text: "", providerMetadata: { bedrock: { signature: "sig_1" } } },
])
}),
)
it.effect("accepts Vercel-compatible redacted reasoning data deltas", () =>
Effect.gen(function* () {
const redactedData = "cmVkYWN0ZWQtdGhpbmtpbmc="
const body = eventStreamBody(
["messageStart", { role: "assistant" }],
["contentBlockDelta", { contentBlockIndex: 0, delta: { reasoningContent: { data: redactedData } } }],
["contentBlockStop", { contentBlockIndex: 0 }],
["messageStop", { stopReason: "end_turn" }],
)
const response = yield* LLMClient.generate(baseRequest).pipe(Effect.provide(fixedBytes(body)))
expect(response.events.find((event) => event.type === "reasoning-delta" && event.text === "")).toEqual({
type: "reasoning-delta",
id: "reasoning-0",
text: "",
providerMetadata: { bedrock: { redactedData } },
})
expect(response.message.content).toEqual([
{ type: "reasoning", text: "", providerMetadata: { bedrock: { redactedData } } },
])
}),
)
it.effect("round-trips streamed redacted reasoning with tool use into a continuation request", () =>
Effect.gen(function* () {
// Bedrock represents redactedContent blobs as base64 strings on its JSON
// wire. The provider owns the payload and requires byte-exact replay.
const redactedData = "cmVkYWN0ZWQtdGhpbmtpbmc="
const response = yield* LLMClient.generate(
LLMRequest.update(baseRequest, {
tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}),
).pipe(
Effect.provide(
fixedBytes(
eventStreamBody(
["messageStart", { role: "assistant" }],
[
"contentBlockDelta",
{ contentBlockIndex: 0, delta: { reasoningContent: { redactedContent: redactedData } } },
],
["contentBlockStop", { contentBlockIndex: 0 }],
[
"contentBlockStart",
{
contentBlockIndex: 1,
start: { toolUse: { toolUseId: "tool_1", name: "lookup" } },
},
],
["contentBlockDelta", { contentBlockIndex: 1, delta: { toolUse: { input: '{"query":"weather"}' } } }],
["contentBlockStop", { contentBlockIndex: 1 }],
["messageStop", { stopReason: "tool_use" }],
),
),
),
)
expect(response.events.find((event) => event.type === "reasoning-delta" && event.text === "")).toEqual({
type: "reasoning-delta",
id: "reasoning-0",
text: "",
providerMetadata: { bedrock: { redactedData } },
})
const prepared = yield* compileRequest(
LLM.request({
model,
messages: [
Message.user("Say hello."),
response.message,
Message.tool({ id: "tool_1", name: "lookup", result: "sunny", resultType: "text" }),
],
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }],
cache: "none",
}),
)
expect(prepared.body.messages).toEqual([
{ role: "user", content: [{ text: "Say hello." }] },
{
role: "assistant",
content: [
{ reasoningContent: { redactedContent: redactedData } },
{ toolUse: { toolUseId: "tool_1", name: "lookup", input: { query: "weather" } } },
],
},
{
role: "user",
content: [{ toolResult: { toolUseId: "tool_1", content: [{ text: "sunny" }], status: "success" } }],
},
])
}),
)
it.effect("classifies throttlingException as a rate limit", () =>
Effect.gen(function* () {
const body = concat([
eventFrame("messageStart", { role: "assistant" }),
exceptionFrame("throttlingException", { message: "Slow down" }),
])
const error = yield* LLMClient.generate(baseRequest).pipe(Effect.provide(fixedBytes(body)), Effect.flip) const error = yield* LLMClient.generate(baseRequest).pipe(Effect.provide(fixedBytes(body)), Effect.flip)
expect(error.reason).toMatchObject({ _tag: "RateLimit", message: "Slow down" }) expect(error.reason).toMatchObject({ _tag: "RateLimit", message: "Slow down" })
@@ -430,7 +684,7 @@ describe("Bedrock Converse route", () => {
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.generate(baseRequest).pipe( const error = yield* LLMClient.generate(baseRequest).pipe(
Effect.provide( Effect.provide(
fixedBytes(eventStreamBody(["validationException", { message: "Input is too long for requested model" }])), fixedBytes(exceptionFrame("validationException", { message: "Input is too long for requested model" })),
), ),
Effect.flip, Effect.flip,
) )
@@ -443,12 +697,44 @@ describe("Bedrock Converse route", () => {
}), }),
) )
it.effect("uses originalMessage from model stream exception frames", () =>
Effect.gen(function* () {
const error = yield* LLMClient.generate(baseRequest).pipe(
Effect.provide(
fixedBytes(
exceptionFrame("modelStreamErrorException", {
originalMessage: "Upstream model failed",
originalStatusCode: 500,
}),
),
),
Effect.flip,
)
expect(error.reason).toMatchObject({ _tag: "ProviderInternal", message: "Upstream model failed" })
}),
)
it.effect("fails unmodeled AWS event-stream errors", () =>
Effect.gen(function* () {
const error = yield* LLMClient.generate(baseRequest).pipe(
Effect.provide(fixedBytes(errorFrame("BadStream", "Stream failed"))),
Effect.flip,
)
expect(error.reason).toMatchObject({
_tag: "InvalidProviderOutput",
message: "BadStream: Stream failed",
})
}),
)
it.effect("rejects requests with no auth path", () => it.effect("rejects requests with no auth path", () =>
Effect.gen(function* () { Effect.gen(function* () {
const unsignedModel = AmazonBedrock.configure({ const unsignedModel = AmazonBedrock.configure({
baseURL: "https://bedrock-runtime.test", baseURL: "https://bedrock-runtime.test",
}).model("anthropic.claude-3-5-sonnet-20240620-v1:0") }).model("anthropic.claude-3-5-sonnet-20240620-v1:0")
const error = yield* LLMClient.generate(LLM.updateRequest(baseRequest, { model: unsignedModel })).pipe( const error = yield* LLMClient.generate(LLMRequest.update(baseRequest, { model: unsignedModel })).pipe(
Effect.provide(fixedBytes(eventStreamBody(["messageStop", { stopReason: "end_turn" }]))), Effect.provide(fixedBytes(eventStreamBody(["messageStop", { stopReason: "end_turn" }]))),
Effect.flip, Effect.flip,
) )
@@ -467,7 +753,7 @@ describe("Bedrock Converse route", () => {
secretAccessKey: "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY", secretAccessKey: "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
}, },
}).model("anthropic.claude-3-5-sonnet-20240620-v1:0") }).model("anthropic.claude-3-5-sonnet-20240620-v1:0")
const prepared = yield* LLMClient.prepare(LLM.updateRequest(baseRequest, { model: signed })) const prepared = yield* compileRequest(LLMRequest.update(baseRequest, { model: signed }))
expect(prepared.route).toBe("bedrock-converse") expect(prepared.route).toBe("bedrock-converse")
expect(prepared.model).toBe(signed) expect(prepared.model).toBe(signed)
@@ -477,7 +763,7 @@ describe("Bedrock Converse route", () => {
it.effect("emits cachePoint markers after system, user-text, and assistant-text with cache hints", () => it.effect("emits cachePoint markers after system, user-text, and assistant-text with cache hints", () =>
Effect.gen(function* () { Effect.gen(function* () {
const cache = new CacheHint({ type: "ephemeral" }) const cache = new CacheHint({ type: "ephemeral" })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_cache", id: "req_cache",
model, model,
@@ -509,7 +795,7 @@ describe("Bedrock Converse route", () => {
it.effect("does not emit cachePoint when no cache hint is set", () => it.effect("does not emit cachePoint when no cache hint is set", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare(baseRequest) const prepared = yield* compileRequest(baseRequest)
expect(prepared.body).toMatchObject({ expect(prepared.body).toMatchObject({
system: [{ text: "You are concise." }], system: [{ text: "You are concise." }],
messages: [{ role: "user", content: [{ text: "Say hello." }] }], messages: [{ role: "user", content: [{ text: "Say hello." }] }],
@@ -519,7 +805,7 @@ describe("Bedrock Converse route", () => {
it.effect("lowers image media into Bedrock image blocks", () => it.effect("lowers image media into Bedrock image blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_image", id: "req_image",
model, model,
@@ -556,7 +842,7 @@ describe("Bedrock Converse route", () => {
it.effect("base64-encodes Uint8Array image bytes", () => it.effect("base64-encodes Uint8Array image bytes", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_image_bytes", id: "req_image_bytes",
model, model,
@@ -578,7 +864,7 @@ describe("Bedrock Converse route", () => {
it.effect("lowers document media into Bedrock document blocks with format and name", () => it.effect("lowers document media into Bedrock document blocks with format and name", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_doc", id: "req_doc",
model, model,
@@ -610,7 +896,7 @@ describe("Bedrock Converse route", () => {
it.effect("requires names for document media", () => it.effect("requires names for document media", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [Message.user({ type: "media", mediaType: "application/pdf", data: "UERGREFUQQ==" })], messages: [Message.user({ type: "media", mediaType: "application/pdf", data: "UERGREFUQQ==" })],
@@ -623,7 +909,7 @@ describe("Bedrock Converse route", () => {
it.effect("passes named document-only messages through for provider validation", () => it.effect("passes named document-only messages through for provider validation", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
cache: "none", cache: "none",
@@ -649,9 +935,10 @@ describe("Bedrock Converse route", () => {
it.effect("lowers document media in tool results", () => it.effect("lowers document media in tool results", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<BedrockConverse.BedrockConverseBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
cache: "none",
messages: [ messages: [
Message.assistant([ToolCallPart.make({ id: "call_1", name: "read", input: { path: "report.pdf" } })]), Message.assistant([ToolCallPart.make({ id: "call_1", name: "read", input: { path: "report.pdf" } })]),
Message.tool({ Message.tool({
@@ -700,7 +987,7 @@ describe("Bedrock Converse route", () => {
it.effect("rejects unsupported image media types", () => it.effect("rejects unsupported image media types", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_bad_image", id: "req_bad_image",
model, model,
@@ -714,7 +1001,7 @@ describe("Bedrock Converse route", () => {
it.effect("rejects unsupported document media types", () => it.effect("rejects unsupported document media types", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_bad_doc", id: "req_bad_doc",
model, model,
@@ -729,7 +1016,7 @@ describe("Bedrock Converse route", () => {
it.effect("maps ttlSeconds >= 3600 to cachePoint ttl: '1h'", () => it.effect("maps ttlSeconds >= 3600 to cachePoint ttl: '1h'", () =>
Effect.gen(function* () { Effect.gen(function* () {
const cache = new CacheHint({ type: "ephemeral", ttlSeconds: 3600 }) const cache = new CacheHint({ type: "ephemeral", ttlSeconds: 3600 })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
system: [{ type: "text", text: "system", cache }], system: [{ type: "text", text: "system", cache }],
@@ -746,7 +1033,7 @@ describe("Bedrock Converse route", () => {
it.effect("appends cachePoint after marked tool definitions and tool-result blocks", () => it.effect("appends cachePoint after marked tool definitions and tool-result blocks", () =>
Effect.gen(function* () { Effect.gen(function* () {
const cache = new CacheHint({ type: "ephemeral" }) const cache = new CacheHint({ type: "ephemeral" })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
tools: [{ name: "lookup", description: "lookup", inputSchema: { type: "object", properties: {} }, cache }], tools: [{ name: "lookup", description: "lookup", inputSchema: { type: "object", properties: {} }, cache }],
@@ -778,7 +1065,7 @@ describe("Bedrock Converse route", () => {
it.effect("drops cachePoint markers past the 4-per-request cap", () => it.effect("drops cachePoint markers past the 4-per-request cap", () =>
Effect.gen(function* () { Effect.gen(function* () {
const cache = new CacheHint({ type: "ephemeral" }) const cache = new CacheHint({ type: "ephemeral" })
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
system: [ system: [
+5 -5
View File
@@ -3,7 +3,7 @@ import { ConfigProvider, Effect, Schema } from "effect"
import { HttpClientRequest } from "effect/unstable/http" import { HttpClientRequest } from "effect/unstable/http"
import { LLM, LLMEvent } from "../../src" import { LLM, LLMEvent } from "../../src"
import { CloudflareAIGateway, CloudflareWorkersAI } from "../../src/providers/cloudflare" import { CloudflareAIGateway, CloudflareWorkersAI } from "../../src/providers/cloudflare"
import { LLMClient } from "../../src/route" import { compileRequest } from "../../src/route/client"
import { it } from "../lib/effect" import { it } from "../lib/effect"
import { dynamicResponse } from "../lib/http" import { dynamicResponse } from "../lib/http"
import { sseEvents } from "../lib/sse" import { sseEvents } from "../lib/sse"
@@ -34,7 +34,7 @@ describe("Cloudflare", () => {
}) })
expect(model.route.endpoint.baseURL).toBe("https://gateway.ai.cloudflare.com/v1/test-account/test-gateway/compat") expect(model.route.endpoint.baseURL).toBe("https://gateway.ai.cloudflare.com/v1/test-account/test-gateway/compat")
const prepared = yield* LLMClient.prepare(LLM.request({ model, prompt: "Say hello." })) const prepared = yield* compileRequest(LLM.request({ model, prompt: "Say hello." }))
expect(prepared.route).toBe("cloudflare-ai-gateway") expect(prepared.route).toBe("cloudflare-ai-gateway")
expect(prepared.body).toMatchObject({ expect(prepared.body).toMatchObject({
@@ -129,7 +129,7 @@ describe("Cloudflare", () => {
openai: { reasoningField: "reasoning", reasoningDetails: merged }, openai: { reasoningField: "reasoning", reasoningDetails: merged },
}) })
const replay = yield* LLMClient.prepare(LLM.request({ model, messages: [response.message] })) const replay = yield* compileRequest(LLM.request({ model, messages: [response.message] }))
expect(replay.body.messages).toEqual([ expect(replay.body.messages).toEqual([
{ role: "assistant", content: "Hello", reasoning: "Thinking", reasoning_details: merged }, { role: "assistant", content: "Hello", reasoning: "Thinking", reasoning_details: merged },
]) ])
@@ -180,7 +180,7 @@ describe("Cloudflare", () => {
it.effect("allows a fully configured baseURL override", () => it.effect("allows a fully configured baseURL override", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model: CloudflareAIGateway.configure({ model: CloudflareAIGateway.configure({
baseURL: "https://gateway.proxy.test/v1/custom/compat", baseURL: "https://gateway.proxy.test/v1/custom/compat",
@@ -208,7 +208,7 @@ describe("Cloudflare", () => {
}) })
expect(model.route.endpoint.baseURL).toBe("https://api.cloudflare.com/client/v4/accounts/test-account/ai/v1") expect(model.route.endpoint.baseURL).toBe("https://api.cloudflare.com/client/v4/accounts/test-account/ai/v1")
const prepared = yield* LLMClient.prepare(LLM.request({ model, prompt: "Say hello." })) const prepared = yield* compileRequest(LLM.request({ model, prompt: "Say hello." }))
expect(prepared.route).toBe("cloudflare-workers-ai") expect(prepared.route).toBe("cloudflare-workers-ai")
expect(prepared.body).toMatchObject({ expect(prepared.body).toMatchObject({
+128 -29
View File
@@ -1,7 +1,8 @@
import { describe, expect } from "bun:test" import { describe, expect } from "bun:test"
import { Effect } from "effect" import { Effect } from "effect"
import { LLM, LLMError, Message, ToolCallPart, Usage } from "../../src" import { LLM, LLMError, LLMRequest, Message, ToolCallPart, ToolDefinition, Usage } from "../../src"
import { Auth, LLMClient } from "../../src/route" import { Auth, LLMClient } from "../../src/route"
import { compileRequest } from "../../src/route/client"
import * as Gemini from "../../src/protocols/gemini" import * as Gemini from "../../src/protocols/gemini"
import { ProviderShared } from "../../src/protocols/shared" import { ProviderShared } from "../../src/protocols/shared"
import { it } from "../lib/effect" import { it } from "../lib/effect"
@@ -26,7 +27,7 @@ const request = LLM.request({
describe("Gemini route", () => { describe("Gemini route", () => {
it.effect("prepares Gemini target", () => it.effect("prepares Gemini target", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare(request) const prepared = yield* compileRequest(request)
expect(prepared.body).toEqual({ expect(prepared.body).toEqual({
contents: [{ role: "user", parts: [{ text: "Say hello." }] }], contents: [{ role: "user", parts: [{ text: "Say hello." }] }],
@@ -36,9 +37,58 @@ describe("Gemini route", () => {
}), }),
) )
it.effect("normalizes Gemini thinking options", () =>
Effect.gen(function* () {
const prepared = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: {
gemini: {
cachedContent: "cachedContents/example",
safetySettings: [{ category: "HARM_CATEGORY_HATE_SPEECH", threshold: "BLOCK_ONLY_HIGH" }],
serviceTier: "priority",
thinkingConfig: { thinkingBudget: 0, includeThoughts: false, thinkingLevel: "high" },
},
},
}),
)
const filtered = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { gemini: { thinkingConfig: { thinkingBudget: "invalid", includeThoughts: false } } },
}),
)
const defaulted = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { gemini: { thinkingConfig: { thinkingLevel: "high" } } },
}),
)
const emptySafetySettings = yield* compileRequest(
LLMRequest.update(request, {
providerOptions: { gemini: { safetySettings: [] } },
}),
)
expect(prepared.body.generationConfig?.thinkingConfig).toEqual({
thinkingBudget: 0,
includeThoughts: false,
thinkingLevel: "high",
})
expect(prepared.body.cachedContent).toBe("cachedContents/example")
expect(prepared.body.safetySettings).toEqual([
{ category: "HARM_CATEGORY_HATE_SPEECH", threshold: "BLOCK_ONLY_HIGH" },
])
expect(prepared.body.serviceTier).toBe("priority")
expect(filtered.body.generationConfig?.thinkingConfig).toEqual({ includeThoughts: false })
expect(defaulted.body.generationConfig?.thinkingConfig).toEqual({
includeThoughts: true,
thinkingLevel: "high",
})
expect(emptySafetySettings.body.safetySettings).toEqual([])
}),
)
it.effect("lowers chronological system updates to wrapped user text in order", () => it.effect("lowers chronological system updates to wrapped user text in order", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<Gemini.GeminiBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [Message.user("Before."), Message.system("Update."), Message.assistant("After.")], messages: [Message.user("Before."), Message.system("Update."), Message.assistant("After.")],
@@ -54,7 +104,7 @@ describe("Gemini route", () => {
it.effect("prepares multimodal user input and tool history", () => it.effect("prepares multimodal user input and tool history", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_tool_result", id: "req_tool_result",
model, model,
@@ -122,7 +172,7 @@ describe("Gemini route", () => {
it.effect("continues media tool results as inline model input without base64 text", () => it.effect("continues media tool results as inline model input without base64 text", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<Gemini.GeminiBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -167,7 +217,7 @@ describe("Gemini route", () => {
it.effect("strips matching data URLs to raw base64 inlineData", () => it.effect("strips matching data URLs to raw base64 inlineData", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare<Gemini.GeminiBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -208,7 +258,7 @@ describe("Gemini route", () => {
] as const) ] as const)
it.effect(`rejects ${name}`, () => it.effect(`rejects ${name}`, () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ model, messages: [Message.user({ type: "media", ...media })] }), LLM.request({ model, messages: [Message.user({ type: "media", ...media })] }),
).pipe(Effect.flip) ).pipe(Effect.flip)
expect(error.message).toMatch(/does not support|does not match|valid base64/) expect(error.message).toMatch(/does not support|does not match|valid base64/)
@@ -217,7 +267,7 @@ describe("Gemini route", () => {
it.effect("rejects oversized image input", () => it.effect("rejects oversized image input", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -233,27 +283,29 @@ describe("Gemini route", () => {
}), }),
) )
it.effect("omits tools when tool choice is none", () => it.effect("keeps tools and sends function calling mode NONE", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_no_tools", id: "req_tool_choice_none",
model, model,
prompt: "Say hello.", prompt: "Say hello.",
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
toolChoice: { type: "none" }, toolChoice: { type: "none" },
}), }),
) )
expect(prepared.body).toEqual({ expect(prepared.body).toMatchObject({
contents: [{ role: "user", parts: [{ text: "Say hello." }] }], contents: [{ role: "user", parts: [{ text: "Say hello." }] }],
tools: [{ functionDeclarations: [{ name: "lookup", description: "Lookup data" }] }],
toolConfig: { functionCallingConfig: { mode: "NONE" } },
}) })
}), }),
) )
it.effect("sanitizes integer enums, dangling required, untyped arrays, and scalar object keys", () => it.effect("sanitizes integer enums, dangling required, untyped arrays, and scalar object keys", () =>
Effect.gen(function* () { Effect.gen(function* () {
const prepared = yield* LLMClient.prepare( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_schema_patch", id: "req_schema_patch",
model, model,
@@ -371,10 +423,16 @@ describe("Gemini route", () => {
{ type: "text-delta", id: "text-0", text: "Hello" }, { type: "text-delta", id: "text-0", text: "Hello" },
{ type: "text-delta", id: "text-0", text: "!" }, { type: "text-delta", id: "text-0", text: "!" },
{ type: "text-end", id: "text-0" }, { type: "text-end", id: "text-0" },
{ type: "step-finish", index: 0, reason: "stop", usage, providerMetadata: undefined }, {
type: "step-finish",
index: 0,
reason: { normalized: "stop", raw: "STOP" },
usage,
providerMetadata: undefined,
},
{ {
type: "finish", type: "finish",
reason: "stop", reason: { normalized: "stop", raw: "STOP" },
usage, usage,
}, },
]) ])
@@ -402,8 +460,8 @@ describe("Gemini route", () => {
], ],
}) })
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
const reasoning = response.events.find((event) => event.type === "reasoning-start") const reasoning = response.events.find((event) => event.type === "reasoning-start")
@@ -428,7 +486,7 @@ describe("Gemini route", () => {
response.events.findIndex((event) => event.type === "tool-call"), response.events.findIndex((event) => event.type === "tool-call"),
) )
const prepared = yield* LLMClient.prepare<Gemini.GeminiBody>( const prepared = yield* compileRequest(
LLM.request({ LLM.request({
model, model,
messages: [ messages: [
@@ -493,8 +551,8 @@ describe("Gemini route", () => {
usageMetadata: { promptTokenCount: 5, candidatesTokenCount: 1 }, usageMetadata: { promptTokenCount: 5, candidatesTokenCount: 1 },
}) })
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
const usage = new Usage({ const usage = new Usage({
@@ -527,10 +585,16 @@ describe("Gemini route", () => {
providerExecuted: undefined, providerExecuted: undefined,
providerMetadata: undefined, providerMetadata: undefined,
}, },
{ type: "step-finish", index: 0, reason: "tool-calls", usage, providerMetadata: undefined }, {
type: "step-finish",
index: 0,
reason: { normalized: "tool-calls", raw: "STOP" },
usage,
providerMetadata: undefined,
},
{ {
type: "finish", type: "finish",
reason: "tool-calls", reason: { normalized: "tool-calls", raw: "STOP" },
usage, usage,
}, },
]) ])
@@ -554,8 +618,8 @@ describe("Gemini route", () => {
], ],
}) })
const response = yield* LLMClient.generate( const response = yield* LLMClient.generate(
LLM.updateRequest(request, { LLMRequest.update(request, {
tools: [{ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } }], tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}), }),
).pipe(Effect.provide(fixedResponse(body))) ).pipe(Effect.provide(fixedResponse(body)))
@@ -569,7 +633,10 @@ describe("Gemini route", () => {
}, },
{ type: "tool-call", id: "tool_1", name: "lookup", input: { query: "news" } }, { type: "tool-call", id: "tool_1", name: "lookup", input: { query: "news" } },
]) ])
expect(response.events.at(-1)).toMatchObject({ type: "finish", reason: "tool-calls" }) expect(response.events.at(-1)).toMatchObject({
type: "finish",
reason: { normalized: "tool-calls", raw: "STOP" },
})
}), }),
) )
@@ -589,9 +656,41 @@ describe("Gemini route", () => {
) )
expect(length.events.map((event) => event.type)).toEqual(["step-start", "step-finish", "finish"]) expect(length.events.map((event) => event.type)).toEqual(["step-start", "step-finish", "finish"])
expect(length.events.at(-1)).toMatchObject({ type: "finish", reason: "length" }) expect(length.events.at(-1)).toMatchObject({
type: "finish",
reason: { normalized: "length", raw: "MAX_TOKENS" },
})
expect(filtered.events.map((event) => event.type)).toEqual(["step-start", "step-finish", "finish"]) expect(filtered.events.map((event) => event.type)).toEqual(["step-start", "step-finish", "finish"])
expect(filtered.events.at(-1)).toMatchObject({ type: "finish", reason: "content-filter" }) expect(filtered.events.at(-1)).toMatchObject({
type: "finish",
reason: { normalized: "content-filter", raw: "SAFETY" },
})
}),
)
it.effect("maps current blocking and invalid-output finish reasons", () =>
Effect.gen(function* () {
const reasons = [
["MODEL_ARMOR", "content-filter"],
["IMAGE_PROHIBITED_CONTENT", "content-filter"],
["IMAGE_RECITATION", "content-filter"],
["LANGUAGE", "content-filter"],
["UNEXPECTED_TOOL_CALL", "error"],
["NO_IMAGE", "error"],
["IMAGE_OTHER", "unknown"],
["TOO_MANY_TOOL_CALLS", "error"],
["MISSING_THOUGHT_SIGNATURE", "error"],
["MALFORMED_RESPONSE", "error"],
] as const
for (const [raw, normalized] of reasons) {
const response = yield* LLMClient.generate(request).pipe(
Effect.provide(
fixedResponse(sseEvents({ candidates: [{ content: { role: "model", parts: [] }, finishReason: raw }] })),
),
)
expect(response.finishReason).toEqual({ normalized, raw })
}
}), }),
) )
@@ -621,7 +720,7 @@ describe("Gemini route", () => {
it.effect("rejects unsupported assistant media content", () => it.effect("rejects unsupported assistant media content", () =>
Effect.gen(function* () { Effect.gen(function* () {
const error = yield* LLMClient.prepare( const error = yield* compileRequest(
LLM.request({ LLM.request({
id: "req_media", id: "req_media",
model, model,

Some files were not shown because too many files have changed in this diff Show More