Adrian Lyjak 20fe453628 Add LlamaParse configuration support to extraction workflows (#211)
* Add parse configuration to coding agent templates and extract-basic

Add ParseConfig and ParseSettings models to support LlamaParse configuration
through config.json. This allows configuring parse tier (fast/agentic),
version, max_pages, and target_pages via ResourceConfig annotations.

- Add ParseConfig and ParseSettings Pydantic models to config.py
- Add parse section to extract-basic configs/config.json
- Update AGENTS.extraction.md with parse configuration documentation

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Remove target_pages from parse config, add lang setting

target_pages is job-specific (varies per file) and doesn't belong in
global configuration. Added lang setting for document language hints.

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Remove lang from parse config

Language detection is handled automatically by LlamaParse, so there's
no need to expose it as a required configuration option.

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* Add lang as optional parse setting

https://claude.ai/code/session_01HCw4cnMMFpKYgRA9m2ryC6

* chore: auto-fix Lint issues (#212)

Co-authored-by: adrianlyjak <2024018+adrianlyjak@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: llama-org-ci-bot[bot] <231146559+llama-org-ci-bot[bot]@users.noreply.github.com>
Co-authored-by: adrianlyjak <2024018+adrianlyjak@users.noreply.github.com>
2026-02-03 09:37:01 -05:00
2026-01-14 11:52:59 -05:00

Data Extraction and Ingestion

This is a starter for LlamaAgents. See the LlamaAgents (llamactl) getting started guide for context on local development and deployment.

To run the application, install uv and run uvx llamactl serve.

Simple customizations

For some basic customizations, you can modify src/extraction_review/config.py

  • EXTRACTION_AGENT_NAME: Logical name for your Extraction Agent. When USE_REMOTE_EXTRACTION_SCHEMA is False, this name is used to upsert the agent with your local schema; when True, it is used to fetch an existing agent.
  • EXTRACTED_DATA_COLLECTION: The Agent Data collection name used to store extractions (namespaced by agent name and environment).
  • ExtractionSchema: When using a local schema, edit this Pydantic model to match the fields you want extracted. Prefer optional types where possible to allow for partial extractions. Note that the extraction process requires all values! so you must explicitly set values to be optional if they are not required. (pydantic default factories will not work, as pydantic only uses default values for missing fields).

The UI fetches the JSON Schema and collection name from the backend metadata workflow at runtime, and dynamically generates an editing UI based on the schema. If you customize this application to have a different extraction schema from the presentation schema rendered in the UI, for example if you customize the extraction process to add additional fields or otherwise transforma it, then you must return the presentation schema from the metadata workflow.

Complex customizations

For more complex customizations, you can edit the rest of the application. For example, you could

  • Modify the existing file processing workflow to provide additional context for the extraction process
  • Take further action based on the extracted data.
  • Add additional workflows to submit data upon approval.

Linting and type checking

Python and javascript packages contain helpful scripts to lint, format, and type check the code.

To check and fix python code:

uv run hatch run lint
uv run hatch run typecheck
uv run hatch run test
# run all at once
uv run hatch run all-fix

To check and fix javascript code, within the ui directory:

pnpm run lint
pnpm run typecheck
pnpm run test
# run all at once
pnpm run all-fix
S
Description
Llama Index Workflow Template
Readme 247 KiB
Languages
Python 48.3%
TypeScript 43.7%
CSS 7.3%
HTML 0.4%
Jinja 0.2%
Other 0.1%