Commit Graph

48 Commits

Author SHA1 Message Date
Adrian Lyjak b1a48e88bf copier-update downstream templates with WorkflowProgress failure cards (#263) v0.3.2 2026-04-17 14:49:41 -04:00
Adrian Lyjak 67a75b4f81 surface failed workflow handlers in data-extraction UI (LI-6906) (#260) 2026-04-17 13:06:24 -04:00
Adrian Lyjak 111144991c fix agent-data reads on extraction-review UIs deployed to local clusters (#259) v0.3.1 2026-04-16 23:55:40 -04:00
Adrian Lyjak 77a668b337 llama-cloud v2: classify-extract-sec + extract-reconcile-invoice (#257) 2026-04-16 19:37:26 -04:00
Adrian Lyjak 9e47fd558e extract-basic: align config with llama-cloud SDK *V2Parameters (#254) v0.3.0 2026-04-15 16:59:31 -04:00
Adrian Lyjak 730d00183a extract-basic: migrate to llama-cloud v2 SDK (#249)
* extract-basic: migrate workflow to llama-cloud v2 SDK

Switch extract-basic to llama-cloud>=2.3.0 and the v2 extract API:
configuration_id replaces extraction_agent_id, ExtractConfig mirrors
ExtractConfiguration (tier/extraction_target/etc), and process_file now
uses client.extract.create + wait_for_completion + get with
expand=["extract_metadata"] and ExtractedData.from_extract_job.
metadata_workflow resolves schemas via client.configurations.retrieve
gated on ExtractV2Parameters.

* extract-basic: update fake server and tests for llama-cloud v2

Rewrite FakeLlamaCloudServer.extract for v2 routes (POST/GET
/api/v2/extract, schema validate/generate, /api/v1/beta/configurations
CRUD) and update split to read categories from the nested
configuration payload. Rewrite test_extract and test_split to the v2
SDK shapes (client.extract.run with inline configuration and
configuration_id; split with configuration={'categories': [...]}).

* extract-basic: split configurations namespace out of extract fake

Move /api/v1/beta/configurations CRUD handlers and StoredConfiguration
state into their own FakeConfigurationsNamespace. FakeExtractNamespace
now takes a FakeConfigurationsNamespace and delegates configuration_id
lookups to it. Server wires both namespaces independently.
2026-04-15 11:12:11 -04:00
Adrian Lyjak e6850493be fix: update classify target_pages to string format matching v2 API (#246) v0.2.9 2026-04-01 23:30:28 -04:00
Adrian Lyjak 9f3557b6e0 Pin llama-cloud version <2 (#243) v0.2.8 2026-03-30 15:29:14 -04:00
Adrian Lyjak 3e1789812b upgrade file api from deprecated query (#242) v0.2.7 2026-03-18 19:20:32 -04:00
Adrian Lyjak e826f6cbf2 feat: copier update child templates from data-extraction v0.5.0 (#241) v0.2.6 2026-03-18 19:11:41 -04:00
Adrian Lyjak b4464b5ff6 feat: migrate templates from llama-cloud-services to llama-cloud (#240) v0.2.4 2026-03-18 16:28:30 -04:00
Adrian Lyjak c26b4aa631 Fix testing_utils import error for llama-cloud 1.3.0 and bump extract… (#218) v0.2.3 2026-02-04 20:31:08 -05:00
Adrian Lyjak 8d1bbb50ff Split agents.md up more (#215)
* improving ergonomics

* Split up instructions and adjust discriminator
2026-02-04 15:45:49 -05:00
Adrian Lyjak f512dcab5e version-bump (#217) 2026-02-04 12:56:50 -05:00
Adrian Lyjak 241c502393 Update llama-cloud-services to @llamaindex/llama-cloud (#216) 2026-02-04 12:10:40 -05:00
Adrian Lyjak 9a9650de8e Use client-side content hash for file deduplication in extract-basic (#214)
- Enable contentHash option in WorkflowTrigger to compute SHA-256 hash of file contents before upload
- Pass file_hash from UI to backend workflow for deduplication
- Remove file download from backend start_extraction step - no longer needed since hash comes from UI
- Fall back to external_file_id from file metadata when file_hash is not provided
- Remove unused imports (hashlib, tempfile, os, httpx, Path) and file_path from state

This improves efficiency by avoiding redundant file downloads just to compute the content hash.

https://claude.ai/code/session_01EX9SjB2rHLqBYLRvYVoSKy

Co-authored-by: Claude <noreply@anthropic.com>
2026-02-03 12:39:13 -05:00
Adrian Lyjak 65e1f6fc6b Update dependencies and migrate to new @llamaindex/ui API (#210)
* feat(extract-basic): update UI to @llamaindex/ui 4.x

Update the extract-basic template UI to use the latest @llamaindex/ui
version (4.1.2) with the following changes:

- Update dependencies to latest versions (@llamaindex/ui ~4.1.2,
  @llamaindex/workflows-client ^1.8.3, react-router-dom ^7.8.0, etc.)
- Migrate API client from createCloudAgentClient/cloudApiClient to
  configureCloudClient/getCloudClient/createAgentDataConfig
- Update TypedAgentData to AgentDataItem
- Update Button component to use label/startIcon props instead of children
- Update imports to use types from @llamaindex/ui instead of llama-cloud-services

https://claude.ai/code/session_01Y5nbgonCDBWGEtEzN3Ua6f

* fix(extract-basic): update @llamaindex/ui to ^4.1.3

Update @llamaindex/ui from ~4.1.2 to ^4.1.3 minimum version.
Version 4.1.3 includes project ID support via ApiProvider.

https://claude.ai/code/session_01Y5nbgonCDBWGEtEzN3Ua6f

* fix(extract-basic): configure cloud client with project ID

Add projectId option to configureCloudClient() to include Project-Id
header in all cloud API requests. This properly scopes requests to
the agent's project when authenticating with a user cookie.

https://claude.ai/code/session_01Y5nbgonCDBWGEtEzN3Ua6f

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-02-03 11:58:54 -05:00
Adrian Lyjak 1a665377c0 fix discriminator access (#213) 2026-02-03 09:51:52 -05:00
Adrian Lyjak 20fe453628 Add LlamaParse configuration support to extraction workflows (#211)
* Add parse configuration to coding agent templates and extract-basic

Add ParseConfig and ParseSettings models to support LlamaParse configuration
through config.json. This allows configuring parse tier (fast/agentic),
version, max_pages, and target_pages via ResourceConfig annotations.

- Add ParseConfig and ParseSettings Pydantic models to config.py
- Add parse section to extract-basic configs/config.json
- Update AGENTS.extraction.md with parse configuration documentation

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Remove target_pages from parse config, add lang setting

target_pages is job-specific (varies per file) and doesn't belong in
global configuration. Added lang setting for document language hints.

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Remove lang from parse config

Language detection is handled automatically by LlamaParse, so there's
no need to expose it as a required configuration option.

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Add lang as optional parse setting

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* chore: auto-fix Lint issues (#212)

Co-authored-by: adrianlyjak <2024018+adrianlyjak@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: llama-org-ci-bot[bot] <231146559+llama-org-ci-bot[bot]@users.noreply.github.com>
Co-authored-by: adrianlyjak <2024018+adrianlyjak@users.noreply.github.com>
2026-02-03 09:37:01 -05:00
Adrian Lyjak 39c533fef3 Add discriminator pattern support for multi-schema extraction workflows (#208) 2026-02-02 22:54:06 -05:00
Adrian Lyjak ef3e3ad632 Handle union types in JSON schema value generation (#202)
* fix: handle nullable union types in _generate_value()

JSON Schema represents nullable fields as type arrays (e.g.,
["number", "null"]). The code compared these arrays against strings,
causing all nullable fields to fall through to the text blob default,
producing invalid data that always failed Pydantic validation.

Extract the first concrete (non-null) type from the array so the
correct generator branch is used.

https://claude.ai/code/session_01NMzCmbJGa6WNS6ceb1aXT7

* test: add unit tests for _generate_value schema handling

Cover nullable union types (["number", "null"]), all basic types,
string formats (date-time, email, uri), composite schemas (object,
array, oneOf, anyOf), and edge cases (depth limit, empty enum,
unknown mapping). 35 tests total.

https://claude.ai/code/session_01NMzCmbJGa6WNS6ceb1aXT7

* test: rewrite deterministic tests as flat parametrized functions

Replace class-based tests with idiomatic pytest: parametrized cases
for basic/nullable types and numeric bounds, a kitchen-sink
integration test, and minimal standalone tests for edge cases.

https://claude.ai/code/session_01NMzCmbJGa6WNS6ceb1aXT7

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-01-27 20:49:20 -05:00
Adrian Lyjak 758c5cf294 Configure Vite HMR client port via environment variable (#201)
* feat: add LLAMA_DEPLOY_SERVER_PORT HMR clientPort to template vite configs

Set hmr.clientPort to LLAMA_DEPLOY_SERVER_PORT when defined, enabling
Vite HMR websocket connections to work through the server proxy.

https://claude.ai/code/session_01Hq5ZQxquwtyJVJ2qw48bbt

* feat: add LLAMA_DEPLOY_SERVER_PORT HMR support to template vite configs

Add hmr.port and hmr.clientPort to all template vite configs so that
Vite HMR websocket connections work through the server proxy. When
LLAMA_DEPLOY_SERVER_PORT is set, clientPort directs the browser's HMR
websocket to connect via the proxy server port.

Bump patch versions for all affected templates.

https://claude.ai/code/session_01Hq5ZQxquwtyJVJ2qw48bbt

---------

Co-authored-by: Claude <noreply@anthropic.com>
v0.2.2
2026-01-27 17:58:46 -05:00
Adrian Lyjak 428152d82f feat: extend mock server with files list, split index, and sheets APIs (#197) 2026-01-27 13:06:19 -05:00
Clelia (Astra) Bertelli 6f4e661fb4 chore: add split to mock server (#199) 2026-01-27 14:17:04 +01:00
Adrian Lyjak 9b970b7c83 Slim the steps down (#193) 2026-01-26 11:10:14 -05:00
Clelia (Astra) Bertelli c5713aa659 chore: add classify config (#195) 2026-01-26 11:07:04 -05:00
Adrian Lyjak f1dbe49802 Enforce ResourceConfig for extraction configs in workflows (#189)
* Refactor extract-basic to prevent load_config anti-pattern

- Remove bad example from create_union_schema() docstring that showed
  using a load_config() function
- Add comprehensive example showing the correct Resource-based pattern
  for multi-document-type workflows
- Update AGENTS.extraction.md with prominent warnings about the
  load_config anti-pattern and proper ResourceConfig usage
- Add "Common Mistakes" section documenting this critical anti-pattern

The load_config pattern bypasses ResourceConfig, causing extraction
configs to be invisible in the workflow representation and preventing
users from editing extraction schemas in the UI.

* Add DRY pattern for ResourceConfig type aliases

Show how to store ResourceConfig annotations as type aliases
(e.g., Extract10KConfig) that can be reused across multiple workflow
steps and Resource functions, reducing duplication.

* Simplify ResourceConfig documentation

Remove explicit anti-pattern examples that may inadvertently teach
the pattern. Keep warnings brief and general, focus on showing the
correct Resource function pattern.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-01-25 10:59:21 -05:00
Adrian Lyjak 2107116dc5 Make workflow representation export programmatic via ResourceConfig (#187) 2026-01-25 09:56:26 -05:00
Adrian Lyjak c2dede1a58 improve .env documentation comment (#182) v0.2.1 2026-01-23 12:42:52 -05:00
Clelia (Astra) Bertelli adda5d6b32 feat: template with new sdk (#178)
* feat: template with new sdk

* chore: more test updates

* chore: more tests

* chore: update system prompt

* fix: generate_value for nested pydantic models

* chore: template validation

* chore: add ExtractedData

* chore: last tweaks to new sdk template

* feat: add --template flag to bundle-coder.sh (#181)

* feat: add --template flag to bundle-coder.sh

Add a -T/--template flag to select which template to bundle instead of
creating a separate script. This consolidates bundle-coder.sh and
bundle-coder-new-sdk.sh into a single script.

- Default template remains 'extract-basic'
- Use --template extract-basic-new to bundle the new SDK template
- Template name suffix is included in the e2b template name

* Update bundle-coder.sh

---------

Co-authored-by: Claude <noreply@anthropic.com>

* chore: pr suggestions

* chore: package versioning

* chore: move extract-basic-new to extract-basic

* chore: prompt tweaks

* fix: typo

* fix: typo pt2

* chore: pr suggestions

---------

Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
v0.2.0
2026-01-23 16:12:00 +01:00
Adrian Lyjak 8513a08de0 Configure INFO level logging to stdout in extract-basic test conftest (#177)
Co-authored-by: Claude <noreply@anthropic.com>
2026-01-15 10:35:49 -05:00
Adrian Lyjak f43a898b90 Update workflow and resource docstrings for UI export (#176)
Improve docstrings throughout extract-basic workflow to be more
domain-expert friendly, using product terms instead of programming
jargon. Also document resource function naming behavior in AGENTS.md.

Co-authored-by: Claude <noreply@anthropic.com>
2026-01-15 09:42:10 -05:00
Adrian Lyjak 7d8c8a5e85 removing llamactl - it locks old versions, I guess causing issues if we update llamactl in the container? The injected llama-deploy-appserver isn't updated if its locked, leading to errors like where the app server version doesn't support the --skip-env-validation flag (#174) 2026-01-14 19:22:32 -05:00
Adrian Lyjak cfb4963e23 fix parse mock, and test the fake server (#171) 2026-01-14 16:59:51 -05:00
Adrian Lyjak 8bf8720c2f Validate extract-basic tests (#170) 2026-01-14 15:37:17 -05:00
Adrian Lyjak 33e6a39188 modify gitignore (#168) v0.1.8 2026-01-14 11:52:59 -05:00
Adrian Lyjak 0ecd2e1f06 Add copier-update script to help run mass copier updates (#166)
* add updater script

* bump versions

* fixer

* chore(classify-extract-sec): copier update

* chore(extract-basic): copier update

* chore(extract-reconcile-invoice): copier update

* fix format
2026-01-13 07:59:26 -05:00
Adrian Lyjak 8586b25e8a Make the llama-ui version range more strict (#164)
* fix: pin @llamaindex/ui to ~3.6.1 to avoid breaking changes in minors

Changed from ^x.y.z (caret) to ~3.6.1 (tilde) versioning for
@llamaindex/ui dependency across all template UI packages to prevent
breaking changes from minor version bumps.

* chore: bump template versions for @llamaindex/ui update

Incremented patch versions for all templates with updated UI dependency:
- basic-ui: 0.2.4 → 0.2.5
- classify-extract-sec: 0.2.6 → 0.2.7
- data-extraction: 0.3.7 → 0.3.8
- document-qa: 0.2.6 → 0.2.7
- extract-basic: 0.1.6 → 0.1.7
- extract-reconcile-invoice: 0.3.2 → 0.3.3
- showcase: 0.3.1 → 0.3.2

* fix: revert document-qa and basic-ui to @llamaindex/ui ^2.1.1

These templates still use v2 of the UI package, so keeping them on their
original caret range and reverting their version bumps.

---------

Co-authored-by: Claude <noreply@anthropic.com>
v0.1.7
2026-01-12 17:31:09 -05:00
Clelia (Astra) Bertelli d6978b3847 fix: testing utils (#154)
* fix: handle files.get when ID is uuidv4 in mock utilities

* chore: tweaks to test and prompt

* fix: fix problem with nested objects

* chore: use jsonref!

* chore: remove mention to uuid4

* chore: resort to preloading files to get their ID instead of uuid

* chore: warn explicitly about test being skipped
v0.1.6
2026-01-06 15:04:14 -05:00
Adrian Lyjak 2baff57886 Add test timeout functionality and configuration (#145)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-01-03 11:45:03 -05:00
Clelia (Astra) Bertelli 67d5f0d5c6 chore: reimplement visualization (#123)
* chore: reimplement visualization

* fix: track viz task

* chore: add suggestion

Co-authored-by: Logan <logan.markewich@live.com>

* ci: format

* chore: switch to global task

* do not push

* cleanup

* fix: testing and fixing broken viz components

* ci: format

* chore: vbump

---------

Co-authored-by: Logan <logan.markewich@live.com>
v0.1.5
2025-12-19 20:46:37 +01:00
Adrian Lyjak 66741c399d Cursor behavior issue (#122)
* Bump ty dependency to 0.0.2

Co-authored-by: adrian <adrian@runllama.ai>

* Update template versions and fix agent context typing

Co-authored-by: adrian <adrian@runllama.ai>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-12-16 16:38:12 -05:00
Clelia (Astra) Bertelli 49cdf6c40f feat: running tests in the background (#90)
* wip: offload ui tests running to a subagent

* chore: clarification comment

* chore: automated ui tests with mocks

* feat: background test runner

* chore: prepare experiment for bg test runner

* fix: use threading for bg test runner, update test template

* ci: ruff format and check

* fix: import uuid in test template + try fixing race conditions in file watcher

* ci: format

* feat: experiment

* fix: threading approach to log collection; chore: prepare experiment

* chore: migrate to watchfiles and asyncio tasks

* chore: clean experiment and prepare to run again

* feat: experiment result; fix: fix test results hook to read from the correct file

* fix: ui test while loop breaking

* chore: reset experiment ground

* fix: use sys.executable to run tests

* fix: bug hunting

* ci: format

* fix: try adding a debugging layer

* yay it worked 🎉

* chore: refactor to be async, make filewatcher a singleton shared across modules (hopefully)

* chore: prepare experiment

* chore: try make tests complete

* fix: merge conflicts mess

* chore: prepare for testing

* feat: instrument test runner + perform experiment

* ci: format

* chore: re-instrument _read_logs_stream

* chore: PR comments
v0.1.3
2025-12-05 22:56:53 +01:00
Clelia (Astra) Bertelli 5355f9b48c Adding benchmarking utils and scripts (#86) 2025-12-01 16:12:11 -05:00
Clelia (Astra) Bertelli 7519d9a12b chore: add testing utils + adapt agent structure (#84)
* chore: add testing utils + adapt agent structure

* Playwright tester for UI (#85)

* feat: add playwright testing command

* fix: small fixes

* feat: add multi-turn runner; chore: prepare for exp 2

* chore: correct exp 2 name

* chore: experiment 2 and more trial and error with the playwright tester

* ci: format

* docs: document multi-turn runner
v0.1.2
2025-11-27 16:47:32 +01:00
Clelia (Astra) Bertelli 9a6ce4da33 Modify extract-basic template (#82)
* chore: modify extract basic template

* ci: format

* chore: version bump

* docs: mention only stateless extraction

* chore: prepare for first experiment with extract-basic

* chore: experiment 1 + cleanup as per PR comments

* ci: format and lint

* Tweaks to code and docs to encourage better coding agent behavior

---------

Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
v0.1.1
2025-11-24 19:11:51 +01:00
Adrian Lyjak a44e95123f initialize extract-basic from data-extraction 2025-11-21 16:15:37 -05:00
Adrian Lyjak 8d235a02bf Initial commit 2025-11-21 16:11:56 -05:00