Compare commits

...

125 Commits

Author SHA1 Message Date
Neeraj Pradhan f310aea5da fix prod/staging env 2026-01-22 18:14:50 -08:00
Neeraj Pradhan a503329e84 fix url 2026-01-22 17:57:46 -08:00
Neeraj Pradhan 87f6ce8200 Run hourly extract tests to catch SDK schema drifts 2026-01-22 17:53:18 -08:00
Neeraj Pradhan 013b689812 Bump up minor version for python packages (#1085) 2026-01-21 12:13:13 -08:00
Adrian Lyjak 3040951cb8 Use error description in invalid extraction error (#1081)
* fix: display extraction job error in InvalidExtractionData exception

Refactored InvalidExtractionData to read the `error` field from
ExtractRun and prominently display it in the exception message.
The job-level error is now stored in the `extraction_error` attribute
and included in the invalid_item's metadata as `job_error`.

* Create three-yaks-beg.md

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-01-18 17:43:21 -05:00
github-actions[bot] 9239498945 chore: version packages (#1076)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-01-14 19:15:05 +01:00
Pierre-Loic Doulcet 19cbb25631 remove extension filter (#1075)
* remove extension filter

* changeset

* Update ninety-goats-look.md

Make it a patch version

* Update package.json

back out of version bump

* Update pyproject.toml

back out of version bump

* Update package.json

back out of version bump

* Update pyproject.toml

back out of version bump

---------

Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2026-01-14 19:13:39 +01:00
github-actions[bot] 812e2f7d72 chore: version packages (#1073)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-01-12 19:03:13 +01:00
Clelia (Astra) Bertelli d7864afe3f fix: bug fix retry logic in Classify and Extract (#1066)
* fix: bug fix retry logic in Classify and Extract

* chore: apply suggestion

* chore: add PARTIAL_SUCCESS to classify
2026-01-12 18:57:40 +01:00
github-actions[bot] ade8d027a5 chore: version packages (#1071)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-01-09 20:29:00 -05:00
Adrian Lyjak 997bcc8531 forgot ts changeset (#1070) 2026-01-09 20:23:29 -05:00
github-actions[bot] 8be554c234 chore: version packages (#1068)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-01-09 18:56:51 -05:00
Adrian Lyjak f777cab0c5 Add bounding box type support to TS too (#1069)
ts too
2026-01-09 18:55:16 -05:00
Adrian Lyjak b9b83c953d Parse bounding boxes from extract jobs results in agent data (#1067) 2026-01-09 18:47:57 -05:00
github-actions[bot] 3ec7024626 chore: version packages (#1058) 2025-12-10 11:53:30 -06:00
Logan d5b18a03fa Remove generate from build path to fix publishing (#1057) 2025-12-10 11:52:43 -06:00
Clelia (Astra) Bertelli 18dd04b6de docs: correct links in readme (#1056) 2025-12-10 17:08:58 +01:00
github-actions[bot] 685a5e6ccc chore: version packages (#1054) 2025-12-09 15:30:13 -06:00
Jim Geurts 576c3d9076 feat: support zod v4 & v3 (#1052) 2025-12-09 15:29:23 -06:00
Logan c8321d2bc5 improve parse ts polling (#1053) 2025-12-09 15:21:19 -06:00
Tuana Çelik 131bbed7aa batch parse sctript with asyncio (#1051)
* batch parse sctript with asyncio

* lint

---------

Co-authored-by: Logan Markewich <logan.markewich@live.com>
2025-12-08 18:50:11 +01:00
Javier Torres 41c8ac2348 docs: Split Example Notebook (#1044)
* split notebook

* Lint
2025-12-08 13:57:20 +01:00
github-actions[bot] 32c53cdf96 chore: version packages (#1046) 2025-12-04 20:43:29 -06:00
Logan 71db318fc2 add tier/version to api (#1045) 2025-12-04 20:42:17 -06:00
George He dac0f79e51 Fix sheets API client (#1032) 2025-12-03 16:39:47 -06:00
github-actions[bot] 32487763d5 chore: version packages (#1043) 2025-12-03 14:52:26 -06:00
Daniel Bustamante Ospina 06c3c556e6 Add new fields to SpreadsheetParsingConfig and update validation tests (#1042) 2025-12-03 14:50:23 -06:00
github-actions[bot] e5dcaa83df chore: version packages (#1041)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-12-03 11:03:36 -08:00
Neeraj Pradhan 1b7198dc62 Bump llama cloud services and parse versions (#1040) 2025-12-03 10:39:35 -08:00
github-actions[bot] 9cfe074206 chore: version packages (#1039) 2025-12-02 12:16:50 -06:00
Logan ae30990ada line level bbox (#1038) 2025-12-02 12:12:17 -06:00
github-actions[bot] 8f1c359abc chore: version packages (#1037) 2025-12-02 09:50:07 -06:00
Logan 0a110de9c7 Dummy release (#1036) 2025-12-02 09:45:52 -06:00
github-actions[bot] d705b16923 chore: version packages (#1035) 2025-12-02 09:43:20 -06:00
Logan ca781132c8 No more presigned URLs by default (#1034) 2025-12-02 09:41:49 -06:00
Roman Isecke 7a68b0fb68 docs: add batch parse directory example notebook (#1009)
* create notebook to parse a batch of documents

* remove local dev code

* tidy

* don't git track the sample pdfs

* update notebook to use client

* add logic to fetch parse results using job id from batch item

* generate example for fetching results via parse job id

* fix linting

* convert notebook to use httpx rather than client for now

* fix linting
2025-12-01 13:57:18 -05:00
George He 87dec5433d Add timeouts to E2E GHA (#1031)
* Add timeouts

* Session timeouts too
2025-11-27 14:57:59 -08:00
Pierre-Loic Doulcet 99f4eba8d0 Pierre/more parse parameters (#1027)
* up python sdk

* bupmVErsion

* Update py/llama_cloud_services/parse/base.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Update py/llama_cloud_services/parse/base.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-11-25 14:43:27 +01:00
github-actions[bot] 54561e2dd2 chore: version packages (#1025) 2025-11-24 16:41:22 -06:00
Logan Markewich bfaec79a8f changeset 2025-11-24 16:37:58 -06:00
Logan Markewich 3e0e522a6b update ts 2025-11-24 16:36:31 -06:00
Logan Markewich f70b6d87ec update py 2025-11-24 16:31:15 -06:00
Logan Markewich 693b5b83b1 improve llama-sheets example 2025-11-24 09:44:11 -06:00
Neeraj Pradhan ad38ef5cd7 Add notebook for tabular extraction (#1017) 2025-11-18 09:47:07 -08:00
Logan Markewich 4c4c6e6575 fix sheets test 2025-11-17 16:14:29 -06:00
github-actions[bot] 740b47d9dc chore: version packages (#1016)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-17 16:11:18 -06:00
Logan f3233deb2e propagate retrieval metadata to retrieved nodes (#1015) 2025-11-17 16:06:52 -06:00
github-actions[bot] fd45127678 chore: version packages (#1014)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-17 21:18:09 +01:00
Clelia (Astra) Bertelli 0506c88735 chore: rename classifyclient and keep it backward compatible (#1013)
* chore: rename classifyclient and keep it backward compatible

* chore: Replace ClassifyClient in notebooks

* chore: changesets
2025-11-17 21:16:23 +01:00
Logan 4bc9eb6c0d beta sheets API (#992) 2025-11-17 11:32:06 -06:00
Patricia 5a3dac655c Add support for custom metadata in file upload methods (#1012) 2025-11-17 11:18:11 -06:00
github-actions[bot] 519254efbe chore: version packages (#999)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-04 14:18:27 -05:00
Adrian Lyjak 6ab56b79f3 fix version breaking (#998) 2025-11-04 14:14:38 -05:00
Adrian Lyjak e020e3e2b1 Remove organization id from classify (#997) 2025-11-04 14:05:19 -05:00
Adrian Lyjak f293547910 destructured keyword params for classify (#996) 2025-11-04 14:04:41 -05:00
github-actions[bot] 662bc37462 chore: version packages (#995) 2025-11-03 20:15:50 -06:00
Neeraj Pradhan 9f1ef4ef1f Bump to version 0.6.78 (#994) 2025-11-03 20:11:18 -06:00
github-actions[bot] 1243573924 chore: version packages (#991) 2025-10-30 10:11:16 -06:00
Preston Carlson 407292b177 Fix: Return partial results on job failure (#990)
* Return partial result on failed job, especially job id

* Maintains NO_DATA_FOUND_IN_FILE throw behavior
2025-10-23 13:44:41 -07:00
Clelia (Astra) Bertelli a7df7c0912 docs: add llamaclassify demo (#989) 2025-10-23 17:38:57 +02:00
github-actions[bot] c758144bfe chore: version packages (#988)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-22 14:41:44 +02:00
Clelia (Astra) Bertelli fee516dd19 feat: add classify to ts sdk (#985)
* feat: add classify to ts sdk

* ci: changesets

* chore: camelCase for everyone; refactor: slimmer logic for fileContents/filePaths handling

* chore: implement claude suggestions
2025-10-22 14:39:20 +02:00
Neeraj Pradhan 032fbd5768 Add common SourceText class for classify/extract text inputs (#986) 2025-10-21 13:37:41 -07:00
Jerry Liu 970e864514 improve classify notebook (#983) 2025-10-20 10:07:35 -07:00
github-actions[bot] d0649ece6e chore: version packages (#982) 2025-10-16 16:58:29 -06:00
MartijnLeplae 5d4cabd843 Add ImageNode support in TypeScript (#969) 2025-10-16 16:56:28 -06:00
github-actions[bot] 9070a6ac16 chore: version packages (#981) 2025-10-15 12:01:34 -06:00
Bogdan Gheorghe 4f24f537f6 Add agressive table extraction argument (#980) 2025-10-15 11:57:34 -06:00
github-actions[bot] 8859a203e2 chore: version packages (#977) 2025-10-14 19:03:36 -06:00
dependabot[bot] b091364054 build(deps): bump astral-sh/setup-uv from 6 to 7 (#974) 2025-10-14 19:02:32 -06:00
dependabot[bot] 43b1a013ca build(deps): bump github/codeql-action from 3 to 4 (#973) 2025-10-14 19:02:20 -06:00
Logan f81532e7f2 safest types possible for parse (#976) 2025-10-14 19:02:07 -06:00
github-actions[bot] 986d3987d3 chore: version packages (#965)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-14 08:14:49 -06:00
Logan 1bf522311f fix default bbox values (#975) 2025-10-14 07:44:35 -06:00
Preston Carlson 24166dcfc8 Only escape single dollar sign in notebook md (#964)
* Limit escaping to lone dollar signs - preserve double dollar for latex equations

* Updated uv.lock via make lint

* Patch bump

* Unit test for _format_markdown_for_notebook

Test doesn't depend on getting real results/is just testing a string manipulation function, so inserting before other tests. Should move to its own file if we add additional formatting configurations
2025-10-07 08:06:03 -07:00
github-actions[bot] bfb7f3973f chore: version packages (#956)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-06 11:15:55 -04:00
dependabot[bot] 979f643c77 build(deps): bump actions/checkout from 4 to 5 (#961) 2025-10-06 09:12:38 -06:00
dependabot[bot] aefd89cf1b build(deps): bump actions/setup-python from 5 to 6 (#960) 2025-10-06 09:12:30 -06:00
dependabot[bot] 8ea2b2c64e build(deps): bump pnpm/action-setup from 3 to 4 (#959) 2025-10-06 09:12:20 -06:00
dependabot[bot] 4a9a2a21d8 build(deps): bump astral-sh/setup-uv from 3 to 6 (#958) 2025-10-06 09:12:08 -06:00
Logan e6a7939206 loosen packaging requirements (#962) 2025-10-06 09:11:57 -06:00
Adrian Lyjak 104a03e829 fix: re-enable js publishing (#963) 2025-10-06 11:10:46 -04:00
Terry Zhao 6e0f2f4ca0 citation can be null (#869)
* citation can be null

* Add changeset

---------

Co-authored-by: Terry Zhao <terryzhao@runllama.ai>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-10-04 16:26:11 -04:00
dependabot[bot] 0708d11f8a Bump actions/setup-node from 4 to 5 (#909)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 4 to 5.
- [Release notes](https://github.com/actions/setup-node/releases)
- [Commits](https://github.com/actions/setup-node/compare/v4...v5)

---
updated-dependencies:
- dependency-name: actions/setup-node
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-10-04 16:21:50 -04:00
github-actions[bot] be19185503 chore: version packages (#954)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-03 20:14:04 -04:00
Adrian Lyjak 7571b0d6c4 Missed some things again with tag fixes (#955)
guh
2025-10-03 20:12:53 -04:00
Adrian Lyjak ad6734bf80 fixup tagging more better (#953)
* fix: correct private field type in py/package.json to be recognized by pnpm

* use packages more directly, make public

* add bump

* fix crash
2025-10-03 19:53:57 -04:00
github-actions[bot] 9ec2a8322e chore: version packages (#952) 2025-10-03 15:11:14 -06:00
Logan 51011b9f30 fix changeset harder (#951) 2025-10-03 15:09:58 -06:00
Logan 09805f9e15 swap changesets (#949) 2025-10-03 15:06:00 -06:00
Adrian Lyjak 8ced6f6eab fix: explicitly tag. I thought the action did this (#948) 2025-10-03 16:59:41 -04:00
Preston Carlson 081ddeca34 Escaping dollar signs in md output when running in a jupyter notebook (#945) 2025-10-03 14:52:26 -06:00
Adrian Lyjak 2460908789 Disable npm release (#946) 2025-10-03 16:13:16 -04:00
Adrian Lyjak c226d6a54c Fix more bugs in publishing (#944) 2025-10-03 11:16:43 -04:00
Adrian Lyjak 5d4c682eb2 fix: theres just one publish token (#943) 2025-10-03 10:56:10 -04:00
github-actions[bot] f72d3535c8 chore: version packages (#941)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-03 10:25:11 -04:00
Adrian Lyjak 1ea09a366e Update llama-cloud dep (#940) 2025-10-03 09:56:56 -04:00
Adrian Lyjak d4bbeb6389 ignore nvmrc (#942)
ignore npmrc
2025-10-03 00:21:32 -04:00
Adrian Lyjak d028397603 version and release via changesets (#849) 2025-10-03 00:08:52 -04:00
Emanuel Ferreira 35ea8476db docs: parse -> classify -> extract (#931) 2025-09-24 18:52:15 -03:00
Logan 3e5f7c4f1e Update parse.md 2025-09-24 11:35:13 -06:00
Adrian Lyjak 9d9b816644 Handle reasoning field conflict (#929)
* Handle reasoning field conflict

* update version to 0.6.69
2025-09-22 11:29:11 -04:00
Adrian Lyjak 83555f76e6 Handle validation errors for agent data retrieval (#928)
* feat: Add untyped agent data retrieval and handling

Introduces methods to retrieve agent data as untyped dictionaries,
handling validation errors gracefully. This allows for more flexible
data access when strict typing is not required or when data may be
malformed.

Co-authored-by: adrian <adrian@runllama.ai>

* Expose raw api result

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-09-22 11:28:49 -04:00
Adrian Lyjak 5edf5f914a Support creating indexes in a specified project_id (#924)
* Support creating indexes in a specified project_id

* Bump
2025-09-18 11:07:07 -04:00
Adrian Lyjak 22e4975cb2 Refactor agent fields in llama_cloud_services (#921) 2025-09-17 15:14:40 -04:00
Peter Rowlands (변기호) bc2f04379b py: bump version to v.0.6.66 (#920) 2025-09-16 19:34:18 +09:00
Peter Rowlands (변기호) f9f951d5d8 parse: expose spreadsheet_force_formula_computation option (#919) 2025-09-16 19:28:03 +09:00
Emmanuel Ferdman 355129fea5 Fix colab broken links (#750)
Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>
2025-09-14 23:10:21 +02:00
Adrian Lyjak d9aed80ded fix: v prefix goes deeper. Fix more (#899) 2025-09-08 17:45:06 -04:00
Pierre-Loic Doulcet c07d2d70a8 update parse package (#911) 2025-09-08 09:46:32 -06:00
Neeraj Pradhan ed6937a5a9 Fix uv sync; remove poetry lock (#906) 2025-09-05 17:13:31 -07:00
Neeraj Pradhan 34c15932a3 Bump version to 0.6.64 (#904) 2025-09-05 17:05:21 -07:00
Neeraj Pradhan b18ea96d11 Remove report generation related code from llama_cloud_services (#905) 2025-09-05 16:41:28 -07:00
Clelia (Astra) Bertelli 196ab827f5 fix: make ts release beautiful again (#902) 2025-09-05 10:41:39 -06:00
Peter Rowlands (변기호) ba4cb4d5e9 parse: expose page.slideSpeakerNotes (#889) 2025-09-05 15:48:44 +09:00
Adrian Lyjak 58d883b825 fix: "v" prefix being added to js versions (#898) 2025-09-04 15:39:27 -04:00
Adrian Lyjak 5fc5ebfc6c client unification (#895)
read from the shared client
2025-09-04 14:12:28 -04:00
Adrian Lyjak fe3e20fd53 Update version script, and unify the linting so that prettier is more consistent (#897)
Add version script, and unify the linting so that prettier is more consistent
2025-09-04 14:09:27 -04:00
Jerry Liu e7e59459ab getting started LlamaCloudIndex notebook (#891) 2025-09-02 14:52:39 -06:00
Logan Markewich f4d7c84e19 remove stale param 2025-09-02 13:37:16 -06:00
Yannis Panagis 9050a346e4 Added "SourceText" to __init__.py (#892) 2025-09-02 13:28:24 -06:00
Sourabh Desai 9690ccf4ea Fix tag push command in CONTRIBUTING.md (#894)
seems to be missing one little `v`
2025-09-02 10:49:14 -07:00
Sourabh Desai 97745f0f1c version bump to 0.6.63 (#893) 2025-09-02 10:36:51 -07:00
Sourabh Desai 61a696b9db add file names in return values (#888) 2025-08-29 15:55:18 -07:00
Sourabh Desai 3e01adaf0e add alternative builder method (#887)
* add alternative builder method

* fix test
2025-08-29 15:55:04 -07:00
171 changed files with 43260 additions and 38387 deletions
+8
View File
@@ -0,0 +1,8 @@
# Changesets
Hello and welcome! This folder has been automatically generated by `@changesets/cli`, a build tool that works
with multi-package repos, or single-package repos to help you version and publish your code. You can
find the full documentation for it [in our repository](https://github.com/changesets/changesets)
We have a quick list of common questions to get you started engaging with this project in
[our documentation](https://github.com/changesets/changesets/blob/main/docs/common-questions.md)
+11
View File
@@ -0,0 +1,11 @@
{
"$schema": "https://unpkg.com/@changesets/config@3.1.1/schema.json",
"changelog": "@changesets/cli/changelog",
"commit": false,
"fixed": [],
"linked": [],
"access": "restricted",
"baseBranch": "main",
"updateInternalDependencies": "patch",
"ignore": []
}
+6
View File
@@ -0,0 +1,6 @@
---
"llama-cloud-services-py": minor
"llama_parse": minor
---
Incorporate updates from llama-cloud
+5
View File
@@ -0,0 +1,5 @@
---
"llama-cloud-services-py": patch
---
Use error description in ExtractedData invalid extraction error
+1 -1
View File
@@ -27,7 +27,7 @@ jobs:
- uses: actions/checkout@v5
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+1 -3
View File
@@ -19,11 +19,9 @@ jobs:
uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
+2 -2
View File
@@ -30,12 +30,12 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v3
uses: github/codeql-action/init@v4
with:
languages: python
dependency-caching: true
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v3
uses: github/codeql-action/analyze@v4
with:
category: "/language:python"
+159
View File
@@ -0,0 +1,159 @@
name: Hourly Extract E2E Tests
on:
schedule:
- cron: "0 * * * *" # Runs every hour at the top of the hour
workflow_dispatch:
inputs:
environment:
description: "Environment to run the tests in"
required: false
default: staging
type: choice
options:
- staging
- production
notify_slack:
description: "Notify Slack"
required: false
default: false
type: boolean
workflow_call:
env:
UV_VERSION: "0.7.20"
PYTHON_VERSION: "3.12"
SLACK_CHANNEL_ID: C078PHNTF44 # Extract channel ID
API_E2E_LOG_PATH: ${{ github.workspace }}/extract-e2e.log
jobs:
extract-e2e:
name: "Hourly Extract E2E Tests (${{ matrix.environment }})"
runs-on: ubuntu-latest
timeout-minutes: 30
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}-${{ matrix.environment }}
cancel-in-progress: true
strategy:
fail-fast: false
matrix:
environment: ${{ github.event_name == 'schedule' && fromJson('["staging", "production"]') || fromJson(format('["{0}"]', github.event.inputs.environment || 'staging')) }}
steps:
- name: Set runtime inputs
id: runtime
run: |
environment=${{ matrix.environment }}
notify_slack=${{ github.event.inputs.notify_slack || github.event_name == 'schedule' }}
echo "environment=${environment}" >> $GITHUB_OUTPUT
echo "notify_slack=${notify_slack}" >> $GITHUB_OUTPUT
if [ "${environment}" = "production" ]; then
echo "LLAMA_CLOUD_BASE_URL=https://api.cloud.llamaindex.ai" >> $GITHUB_ENV
api_key_secret="${{ secrets.LLAMA_CLOUD_API_KEY }}"
project_id_secret="${{ secrets.LLAMA_CLOUD_PROJECT_ID }}"
else
echo "LLAMA_CLOUD_BASE_URL=https://api.staging.llamaindex.ai" >> $GITHUB_ENV
api_key_secret="${{ secrets.LLAMA_CLOUD_API_KEY_STAGING }}"
project_id_secret="${{ secrets.LLAMA_CLOUD_PROJECT_ID_STAGING }}"
fi
if [ -n "$api_key_secret" ]; then
echo "LLAMA_CLOUD_API_KEY=$api_key_secret" >> $GITHUB_ENV
fi
if [ -n "$project_id_secret" ]; then
echo "LLAMA_CLOUD_PROJECT_ID=$project_id_secret" >> $GITHUB_ENV
fi
- uses: actions/checkout@v5
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ env.PYTHON_VERSION }} && uv python pin ${{ env.PYTHON_VERSION }}
- name: Run Extract E2E tests
id: extract-tests
continue-on-error: true
working-directory: py
run: |
set -o pipefail
rm -f "$API_E2E_LOG_PATH"
uv run pytest -v -n 8 --timeout=300 --session-timeout=1740 tests/extract/ 2>&1 | tee "$API_E2E_LOG_PATH"
- name: Extract pytest failure summary
id: failed-tests
if: steps.extract-tests.outcome == 'failure' || cancelled()
run: |
summary="$(python3 - <<'PY'
import os
from pathlib import Path
log_path = Path(os.environ["API_E2E_LOG_PATH"])
if not log_path.exists():
print("Test log not found.")
raise SystemExit(0)
lines = log_path.read_text(errors="ignore").splitlines()
def find_section(keyword):
for i, line in enumerate(lines):
if line.startswith("=") and keyword in line:
return i
return None
start = (
find_section("FAILURES")
or find_section("ERRORS")
or find_section("short test summary info")
)
if start is not None:
snippet = lines[start : start + 200]
else:
snippet = lines[-200:]
print("\n".join(snippet).strip())
PY
)"
if [ -z "$summary" ]; then
summary="Failed test summary not available. Review the full run logs."
fi
{
printf 'summary<<EOF\n%s\nEOF\n' "$summary"
} >> "$GITHUB_OUTPUT"
- name: Check test results
if: always()
run: |
if [ "${{ steps.extract-tests.outcome }}" == "failure" ]; then
echo "Extract E2E tests failed"
exit 1
fi
- name: Post to Extract Slack channel
id: slack
if: (failure() || cancelled()) && steps.runtime.outputs.notify_slack == 'true'
uses: slackapi/slack-github-action@v1.27.0
with:
channel-id: ${{ env.SLACK_CHANNEL_ID }}
slack-message: |
ALERT: *Hourly Extract E2E Tests Failure*
*Environment*: ${{ steps.runtime.outputs.environment }}
*Repository*: ${{ github.repository }}
*Branch*: ${{ github.ref_name }}
<${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}|View Details>
Failed tests:
```
${{ steps.failed-tests.outputs.summary }}
```
env:
SLACK_BOT_TOKEN: ${{ secrets.SLACK_BOT_TOKEN }}
@@ -1,4 +1,4 @@
name: Lint - Python
name: Lint
on:
push:
@@ -22,14 +22,25 @@ jobs:
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }}
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
run: pnpm install --no-frozen-lockfile
- name: Run linter
shell: bash
working-directory: py
run: uv run -- pre-commit run -a
# the js checks are run roundaboutly through lint-staged, and -a doesn't run it. Run them directly.
- run: pnpm -w --filter llama-cloud-services run lint
- run: pnpm -w --filter llama-cloud-services run format:check
-37
View File
@@ -1,37 +0,0 @@
name: Lint - TypeScript
on:
push:
branches:
- main
paths:
- "ts/**"
pull_request:
paths:
- "ts/**"
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
run: pnpm install --no-frozen-lockfile
- name: Run lint
working-directory: ts/llama_cloud_services/
run: pnpm run lint
- name: Run Prettier
working-directory: ts/llama_cloud_services/
run: pnpm run format
-66
View File
@@ -1,66 +0,0 @@
name: Publish Release - Python
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
UV_VERSION: "0.7.20"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Publish package
shell: bash
working-directory: py
run: uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Build and publish llama-parse
working-directory: py/llama_parse/
run: |
uv build
uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }} - LlamaCloud Services PY
artifacts: "py/**/dist/*"
generateReleaseNotes: true
draft: false
prerelease: false
-54
View File
@@ -1,54 +0,0 @@
name: Publish Release - TypeScript
on:
push:
tags:
- "llama-cloud-services@*"
jobs:
build-and-publish:
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
run: pnpm install --no-frozen-lockfile
- name: Run Build
working-directory: ts/llama_cloud_services/
run: pnpm build
- name: Build tarball
run: |
pnpm pack
working-directory: ts/llama_cloud_services
- name: Setup npm authentication
run: echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc
env:
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Release
working-directory: ts/llama_cloud_services
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
run: pnpm publish --access public --no-git-checks
- name: Create release
uses: ncipollo/release-action@v1
with:
artifacts: "ts/llama_cloud_services/llama-cloud-services*.tgz"
name: Release ${{ github.ref }} - LlamaCloud Services TS
bodyFile: "ts/llama_cloud_services/CHANGELOG.md"
token: ${{ secrets.GITHUB_TOKEN }}
+2 -1
View File
@@ -12,6 +12,7 @@ env:
jobs:
test_e2e:
runs-on: ubuntu-latest
timeout-minutes: 30
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
@@ -22,7 +23,7 @@ jobs:
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+5 -8
View File
@@ -1,4 +1,4 @@
name: Lint - TypeScript
name: Test - TypeScript
on:
push:
@@ -23,17 +23,14 @@ jobs:
steps:
- uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
run: pnpm install --no-frozen-lockfile
- name: Run Build
working-directory: ts/llama_cloud_services/
run: pnpm build
run: pnpm -r install --no-frozen-lockfile
- name: Build package
run: pnpm --filter llama-cloud-services build
- name: Run Tests
working-directory: ts/llama_cloud_services/
run: pnpm test
@@ -0,0 +1,61 @@
name: Version Bump and Release
on:
push:
branches:
- main
concurrency: ${{ github.workflow }}-${{ github.ref }}
jobs:
release:
name: Release
runs-on: ubuntu-latest
# Only run on main branch pushes
if: github.ref == 'refs/heads/main'
steps:
- name: Checkout Repo
uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v5
with:
node-version: "22"
cache: "pnpm"
- name: Setup Python
uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Install uv
uses: astral-sh/setup-uv@v7
- name: Install dependencies
run: pnpm install
- name: Add auth token to .npmrc file
run: |
cat << EOF >> ".npmrc"
//registry.npmjs.org/:_authToken=$NPM_TOKEN
EOF
env:
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create Release Pull Request or Publish packages
id: changesets
uses: changesets/action@v1
with:
commit: "chore: version packages"
title: "chore: version packages"
# Custom version script
version: pnpm -w run version
# Custom publish script
publish: pnpm -w run publish
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
UV_PUBLISH_TOKEN: ${{ secrets.PYPI_TOKEN }}
LLAMA_PARSE_PYPI_TOKEN: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
+1
View File
@@ -9,3 +9,4 @@ __pycache__/
node_modules/
.turbo/
dist/
.npmrc
+8 -6
View File
@@ -29,12 +29,12 @@ repos:
- id: black-jupyter
name: black-src
alias: black
exclude: ".*uv.lock"
exclude: ".*uv.lock|examples/extract/solar_panel_e2e_comparison.ipynb"
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.0.1
hooks:
- id: mypy
exclude: ^py/tests|^py/unit_tests
exclude: ^py/tests|^py/unit_tests|^examples
additional_dependencies:
[
"types-requests",
@@ -60,11 +60,13 @@ repos:
additional_dependencies: [black==23.10.1]
# Using PEP 8's line length in docs prevents excess left/right scrolling
args: [--line-length=79]
- repo: https://github.com/pre-commit/mirrors-prettier
rev: v3.0.3
- repo: local
hooks:
- id: prettier
exclude: ^(uv.lock|ts/llama_cloud_services/pnpm-lock.yaml|ts/e2e-tests)
- id: lint-staged
name: Run lint-staged for TS files
entry: pnpm -w exec lint-staged
language: system
pass_filenames: false
- repo: https://github.com/codespell-project/codespell
rev: v2.2.6
hooks:
+1 -1
View File
@@ -18,7 +18,7 @@ versions need to be kept consistent to sidecar it with `llama_cloud_services`. B
You can also do this with `./scripts/version-bump.py set 0.x.x` if you have `uv` installed.
Once the change is merged, push a tag `git tag -a v0.x.x -m 0.x.x` and `git push origin 0.x.x`.
Once the change is merged, push a tag `git tag -a v0.x.x -m 0.x.x` and `git push origin v0.x.x`.
This tagging step can be done with `./scripts/version-bump tag`.
-6
View File
@@ -9,7 +9,6 @@ This repository contains the code for hand-written SDKs and clients for interact
This includes:
- [LlamaParse](./parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](./report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](./index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
@@ -28,13 +27,11 @@ Then, you can use the services in your code:
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
)
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
@@ -44,7 +41,6 @@ index = LlamaCloudIndex(
See the quickstart guides for each service for more information:
- [LlamaParse](./parse.md)
- [LlamaReport (beta/invite-only)](./report.md)
- [LlamaExtract](./extract.md)
- [LlamaCloud Index](./index.md)
@@ -57,13 +53,11 @@ You can also create your API key in the EU region [here](https://cloud.eu.llamai
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
EU_BASE_URL,
)
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
+21
View File
@@ -0,0 +1,21 @@
node_modules
package-lock.json
yarn.lock
.DS_Store
.cache
.env
.vercel
.output
.nitro
/build/
/api/
/server/build
/public/build# Sentry Config File
.env.sentry-build-plugin
/test-results/
/playwright-report/
/blob-report/
/playwright/.cache/
.tanstack
.vscode
+4
View File
@@ -0,0 +1,4 @@
**/build
**/public
pnpm-lock.yaml
routeTree.gen.ts
+88
View File
@@ -0,0 +1,88 @@
# LlamaClassify Demo
A TypeScript demo application showcasing the power of **LlamaClassify** - an agentic documents classification service from [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to classify financial documents among three different types (Cash flow statement, Income Statement and Balance Sheet).
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [How It Works](#how-it-works)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 📄 **Documemt Classification**: Classify files based on well-defined rules you can customized and play around with.
- 🤖 **Reasoning-based Actionable Insights**: Get in-depth, reasoning based insights on the document classification, accompanied by confidence scores.
- 🎨 **Beautiful UI**: [DaisyUI](https://daisyui.com)-based interface powered by [TanStack](https://tanstack.com)
-**Fast Development**: Hot reload support with development mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 22 or higher)
- pnpm package manager
- LlamaCloud API key
## Installation
1. Clone the repository:
```bash
git clone https://github.com/run-llama/llama_cloud_services
cd lama_cloud_services/examples-ts/classify/
```
2. Install dependencies:
```bash
npm install
```
3. Set up your environment variables:
```bash
# Add your API key to your environment
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
```
## Usage
### Start the Demo
```bash
npm run dev
```
The application will be up and running on http://localhost:3000
## How It Works
1. **Document Input**: Enter the path to your document when prompted
2. **Parsing**: LlamaClassify, based on the rules you can find [here](./src/utils/classifier.ts), processes the document and classifies it
3. **Results**: The classification outcome, as well as the reasoning behind it and the confidence score, are displayed in the UI.
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 22+ and have all dependencies installed
2. **API Key Issues**: Verify your LlamaCloud API key is correctly set
3. **File Path Errors**: Use absolute paths or ensure relative paths are correct from the project root
## License
MIT License - see the [LICENSE](../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `npm run format` and `npm run lint`
5. Submit a pull request
+34
View File
@@ -0,0 +1,34 @@
{
"name": "tanstack-start-example-basic",
"private": true,
"sideEffects": false,
"type": "module",
"scripts": {
"dev": "vite dev",
"build": "vite build && tsc --noEmit",
"start": "node .output/server/index.mjs"
},
"dependencies": {
"@tanstack/react-router": "^1.133.22",
"@tanstack/react-router-devtools": "^1.133.22",
"@tanstack/react-start": "^1.133.22",
"llama-cloud-services": "file:../../ts/llama_cloud_services",
"react": "^19.0.0",
"react-dom": "^19.0.0",
"tailwind-merge": "^2.6.0",
"zod": "^3.24.2"
},
"devDependencies": {
"@tailwindcss/postcss": "^4.1.15",
"@types/node": "^22.5.4",
"@types/react": "^19.0.8",
"@types/react-dom": "^19.0.3",
"@vitejs/plugin-react": "^4.6.0",
"daisyui": "^5.3.7",
"postcss": "^8.5.1",
"tailwindcss": "^4.1.15",
"typescript": "^5.7.2",
"vite": "^7.1.7",
"vite-tsconfig-paths": "^5.1.4"
}
}
+5
View File
@@ -0,0 +1,5 @@
export default {
plugins: {
'@tailwindcss/postcss': {},
},
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 21 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 862 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.0 KiB

@@ -0,0 +1,19 @@
{
"name": "",
"short_name": "",
"icons": [
{
"src": "/android-chrome-192x192.png",
"sizes": "192x192",
"type": "image/png"
},
{
"src": "/android-chrome-512x512.png",
"sizes": "512x512",
"type": "image/png"
}
],
"theme_color": "#ffffff",
"background_color": "#ffffff",
"display": "standalone"
}
@@ -0,0 +1,53 @@
import {
ErrorComponent,
Link,
rootRouteId,
useMatch,
useRouter,
} from '@tanstack/react-router'
import type { ErrorComponentProps } from '@tanstack/react-router'
export function DefaultCatchBoundary({ error }: ErrorComponentProps) {
const router = useRouter()
const isRoot = useMatch({
strict: false,
select: (state) => state.id === rootRouteId,
})
console.error('DefaultCatchBoundary Error:', error)
return (
<div className="min-w-0 flex-1 p-4 flex flex-col items-center justify-center gap-6">
<ErrorComponent error={error} />
<div className="flex gap-2 items-center flex-wrap">
<button
onClick={() => {
router.invalidate()
}}
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
>
Try Again
</button>
{isRoot ? (
<Link
to="/"
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
>
Home
</Link>
) : (
<Link
to="/"
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
onClick={(e) => {
e.preventDefault()
window.history.back()
}}
>
Go Back
</Link>
)}
</div>
</div>
)
}
@@ -0,0 +1,25 @@
import { Link } from '@tanstack/react-router'
export function NotFound({ children }: { children?: any }) {
return (
<div className="space-y-2 p-2">
<div className="text-gray-600 dark:text-gray-400">
{children || <p>The page you are looking for does not exist.</p>}
</div>
<p className="flex items-center gap-2 flex-wrap">
<button
onClick={() => window.history.back()}
className="bg-emerald-500 text-white px-2 py-1 rounded-sm uppercase font-black text-sm"
>
Go back
</button>
<Link
to="/"
className="bg-cyan-600 text-white px-2 py-1 rounded-sm uppercase font-black text-sm"
>
Start Over
</Link>
</p>
</div>
)
}
+225
View File
@@ -0,0 +1,225 @@
/* eslint-disable */
// @ts-nocheck
// noinspection JSUnusedGlobalSymbols
// This file was automatically generated by TanStack Router.
// You should NOT make any changes in this file as it will be overwritten.
// Additionally, you should also exclude this file from your linter and/or formatter to prevent it from being checked or modified.
import { Route as rootRouteImport } from './routes/__root'
import { Route as UsersRouteImport } from './routes/users'
import { Route as IndexRouteImport } from './routes/index'
import { Route as UsersIndexRouteImport } from './routes/users.index'
import { Route as PostsIndexRouteImport } from './routes/posts.index'
import { Route as UsersUserIdRouteImport } from './routes/users.$userId'
import { Route as PostsPostIdRouteImport } from './routes/posts.$postId'
import { Route as ApiClassifyRouteImport } from './routes/api/classify'
import { Route as PostsPostIdDeepRouteImport } from './routes/posts_.$postId.deep'
const UsersRoute = UsersRouteImport.update({
id: '/users',
path: '/users',
getParentRoute: () => rootRouteImport,
} as any)
const IndexRoute = IndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => rootRouteImport,
} as any)
const UsersIndexRoute = UsersIndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => UsersRoute,
} as any)
const PostsIndexRoute = PostsIndexRouteImport.update({
id: '/posts/',
path: '/posts/',
getParentRoute: () => rootRouteImport,
} as any)
const UsersUserIdRoute = UsersUserIdRouteImport.update({
id: '/$userId',
path: '/$userId',
getParentRoute: () => UsersRoute,
} as any)
const PostsPostIdRoute = PostsPostIdRouteImport.update({
id: '/posts/$postId',
path: '/posts/$postId',
getParentRoute: () => rootRouteImport,
} as any)
const ApiClassifyRoute = ApiClassifyRouteImport.update({
id: '/api/classify',
path: '/api/classify',
getParentRoute: () => rootRouteImport,
} as any)
const PostsPostIdDeepRoute = PostsPostIdDeepRouteImport.update({
id: '/posts_/$postId/deep',
path: '/posts/$postId/deep',
getParentRoute: () => rootRouteImport,
} as any)
export interface FileRoutesByFullPath {
'/': typeof IndexRoute
'/users': typeof UsersRouteWithChildren
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts': typeof PostsIndexRoute
'/users/': typeof UsersIndexRoute
'/posts/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRoutesByTo {
'/': typeof IndexRoute
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts': typeof PostsIndexRoute
'/users': typeof UsersIndexRoute
'/posts/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRoutesById {
__root__: typeof rootRouteImport
'/': typeof IndexRoute
'/users': typeof UsersRouteWithChildren
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts/': typeof PostsIndexRoute
'/users/': typeof UsersIndexRoute
'/posts_/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRouteTypes {
fileRoutesByFullPath: FileRoutesByFullPath
fullPaths:
| '/'
| '/users'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts'
| '/users/'
| '/posts/$postId/deep'
fileRoutesByTo: FileRoutesByTo
to:
| '/'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts'
| '/users'
| '/posts/$postId/deep'
id:
| '__root__'
| '/'
| '/users'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts/'
| '/users/'
| '/posts_/$postId/deep'
fileRoutesById: FileRoutesById
}
export interface RootRouteChildren {
IndexRoute: typeof IndexRoute
UsersRoute: typeof UsersRouteWithChildren
ApiClassifyRoute: typeof ApiClassifyRoute
PostsPostIdRoute: typeof PostsPostIdRoute
PostsIndexRoute: typeof PostsIndexRoute
PostsPostIdDeepRoute: typeof PostsPostIdDeepRoute
}
declare module '@tanstack/react-router' {
interface FileRoutesByPath {
'/users': {
id: '/users'
path: '/users'
fullPath: '/users'
preLoaderRoute: typeof UsersRouteImport
parentRoute: typeof rootRouteImport
}
'/': {
id: '/'
path: '/'
fullPath: '/'
preLoaderRoute: typeof IndexRouteImport
parentRoute: typeof rootRouteImport
}
'/users/': {
id: '/users/'
path: '/'
fullPath: '/users/'
preLoaderRoute: typeof UsersIndexRouteImport
parentRoute: typeof UsersRoute
}
'/posts/': {
id: '/posts/'
path: '/posts'
fullPath: '/posts'
preLoaderRoute: typeof PostsIndexRouteImport
parentRoute: typeof rootRouteImport
}
'/users/$userId': {
id: '/users/$userId'
path: '/$userId'
fullPath: '/users/$userId'
preLoaderRoute: typeof UsersUserIdRouteImport
parentRoute: typeof UsersRoute
}
'/posts/$postId': {
id: '/posts/$postId'
path: '/posts/$postId'
fullPath: '/posts/$postId'
preLoaderRoute: typeof PostsPostIdRouteImport
parentRoute: typeof rootRouteImport
}
'/api/classify': {
id: '/api/classify'
path: '/api/classify'
fullPath: '/api/classify'
preLoaderRoute: typeof ApiClassifyRouteImport
parentRoute: typeof rootRouteImport
}
'/posts_/$postId/deep': {
id: '/posts_/$postId/deep'
path: '/posts/$postId/deep'
fullPath: '/posts/$postId/deep'
preLoaderRoute: typeof PostsPostIdDeepRouteImport
parentRoute: typeof rootRouteImport
}
}
}
interface UsersRouteChildren {
UsersUserIdRoute: typeof UsersUserIdRoute
UsersIndexRoute: typeof UsersIndexRoute
}
const UsersRouteChildren: UsersRouteChildren = {
UsersUserIdRoute: UsersUserIdRoute,
UsersIndexRoute: UsersIndexRoute,
}
const UsersRouteWithChildren = UsersRoute._addFileChildren(UsersRouteChildren)
const rootRouteChildren: RootRouteChildren = {
IndexRoute: IndexRoute,
UsersRoute: UsersRouteWithChildren,
ApiClassifyRoute: ApiClassifyRoute,
PostsPostIdRoute: PostsPostIdRoute,
PostsIndexRoute: PostsIndexRoute,
PostsPostIdDeepRoute: PostsPostIdDeepRoute,
}
export const routeTree = rootRouteImport
._addFileChildren(rootRouteChildren)
._addFileTypes<FileRouteTypes>()
import type { getRouter } from './router.tsx'
import type { createStart } from '@tanstack/react-start'
declare module '@tanstack/react-start' {
interface Register {
ssr: true
router: Awaited<ReturnType<typeof getRouter>>
}
}
+15
View File
@@ -0,0 +1,15 @@
import { createRouter } from '@tanstack/react-router'
import { routeTree } from './routeTree.gen'
import { DefaultCatchBoundary } from './components/DefaultCatchBoundary'
import { NotFound } from './components/NotFound'
export function getRouter() {
const router = createRouter({
routeTree,
defaultPreload: 'intent',
defaultErrorComponent: DefaultCatchBoundary,
defaultNotFoundComponent: () => <NotFound />,
scrollRestoration: true,
})
return router
}
+128
View File
@@ -0,0 +1,128 @@
/// <reference types="vite/client" />
import {
HeadContent,
Scripts,
createRootRoute,
} from '@tanstack/react-router'
import * as React from 'react'
import { DefaultCatchBoundary } from '~/components/DefaultCatchBoundary'
import { NotFound } from '~/components/NotFound'
import { seo } from '~/utils/seo'
export const Route = createRootRoute({
head: () => ({
meta: [
{
charSet: 'utf-8',
},
{
name: 'viewport',
content: 'width=device-width, initial-scale=1',
},
...seo({
title:
'Financial Documents Classification Agent',
description: `Classify financial documents as balance sheets, income statements and cash flow statemets. `,
}),
],
links: [
{ rel: 'stylesheet', href: "https://cdn.jsdelivr.net/npm/daisyui@5" },
{
rel: 'apple-touch-icon',
sizes: '180x180',
href: '/apple-touch-icon.png',
},
{
rel: 'icon',
type: 'image/png',
sizes: '32x32',
href: '/favicon-32x32.png',
},
{
rel: 'icon',
type: 'image/png',
sizes: '16x16',
href: '/favicon-16x16.png',
},
{ rel: 'manifest', href: '/site.webmanifest', color: '#fffff' },
{ rel: 'icon', href: '/favicon.ico' },
],
scripts: [
{
src: '/customScript.js',
type: 'text/javascript',
},
{
src: "https://cdn.jsdelivr.net/npm/@tailwindcss/browser@4",
type: "text/javascript",
}
],
}),
errorComponent: DefaultCatchBoundary,
notFoundComponent: () => <NotFound />,
shellComponent: RootDocument,
})
function RootDocument({ children }: { children: React.ReactNode }) {
return (
<html>
<head>
<HeadContent />
</head>
<body>
<div className="navbar bg-base-100 shadow-sm">
<div className="navbar-start">
<div className="dropdown">
<div tabIndex={0} role="button" className="btn btn-ghost btn-circle">
<svg
xmlns="http://www.w3.org/2000/svg"
className="h-5 w-5"
fill="none"
viewBox="0 0 24 24"
stroke="currentColor"
>
<path
strokeLinecap="round"
strokeLinejoin="round"
strokeWidth="2"
d="M4 6h16M4 12h16M4 18h7"
/>
</svg>
</div>
<ul
tabIndex={0}
className="menu menu-lg dropdown-content bg-base-100 rounded-box z-1 mt-3 w-80 p-2 shadow"
>
<li><a href="/">Home</a></li>
<li><a href="https://cloud.llamaindex.ai">Get Started with LlamaCloud</a></li>
<li><a href="https://developers.llamaindex.ai/python/cloud/llamaclassify/getting_started/">LlamaClassify Docs</a></li>
</ul>
</div>
</div>
<div className="navbar-center">
<a className="btn btn-ghost text-xl" href="/">Financial Documents Classification Agent</a>
</div>
<div className="navbar-end">
<a href="https://github.com/run-llama/llama_cloud_services/main/blob/examples-ts/classify">
<button className="btn btn-ghost btn-circle">
<div className="indicator">
<svg
xmlns="http://www.w3.org/2000/svg"
className="h-10 w-10"
fill="currentColor"
viewBox="0 0 640 512"
>
<path d="M237.9 461.4C237.9 463.4 235.6 465 232.7 465C229.4 465.3 227.1 463.7 227.1 461.4C227.1 459.4 229.4 457.8 232.3 457.8C235.3 457.5 237.9 459.1 237.9 461.4zM206.8 456.9C206.1 458.9 208.1 461.2 211.1 461.8C213.7 462.8 216.7 461.8 217.3 459.8C217.9 457.8 216 455.5 213 454.6C210.4 453.9 207.5 454.9 206.8 456.9zM251 455.2C248.1 455.9 246.1 457.8 246.4 460.1C246.7 462.1 249.3 463.4 252.3 462.7C255.2 462 257.2 460.1 256.9 458.1C256.6 456.2 253.9 454.9 251 455.2zM316.8 72C178.1 72 72 177.3 72 316C72 426.9 141.8 521.8 241.5 555.2C254.3 557.5 258.8 549.6 258.8 543.1C258.8 536.9 258.5 502.7 258.5 481.7C258.5 481.7 188.5 496.7 173.8 451.9C173.8 451.9 162.4 422.8 146 415.3C146 415.3 123.1 399.6 147.6 399.9C147.6 399.9 172.5 401.9 186.2 425.7C208.1 464.3 244.8 453.2 259.1 446.6C261.4 430.6 267.9 419.5 275.1 412.9C219.2 406.7 162.8 398.6 162.8 302.4C162.8 274.9 170.4 261.1 186.4 243.5C183.8 237 175.3 210.2 189 175.6C209.9 169.1 258 202.6 258 202.6C278 197 299.5 194.1 320.8 194.1C342.1 194.1 363.6 197 383.6 202.6C383.6 202.6 431.7 169 452.6 175.6C466.3 210.3 457.8 237 455.2 243.5C471.2 261.2 481 275 481 302.4C481 398.9 422.1 406.6 366.2 412.9C375.4 420.8 383.2 435.8 383.2 459.3C383.2 493 382.9 534.7 382.9 542.9C382.9 549.4 387.5 557.3 400.2 555C500.2 521.8 568 426.9 568 316C568 177.3 455.5 72 316.8 72zM169.2 416.9C167.9 417.9 168.2 420.2 169.9 422.1C171.5 423.7 173.8 424.4 175.1 423.1C176.4 422.1 176.1 419.8 174.4 417.9C172.8 416.3 170.5 415.6 169.2 416.9zM158.4 408.8C157.7 410.1 158.7 411.7 160.7 412.7C162.3 413.7 164.3 413.4 165 412C165.7 410.7 164.7 409.1 162.7 408.1C160.7 407.5 159.1 407.8 158.4 408.8zM190.8 444.4C189.2 445.7 189.8 448.7 192.1 450.6C194.4 452.9 197.3 453.2 198.6 451.6C199.9 450.3 199.3 447.3 197.3 445.4C195.1 443.1 192.1 442.8 190.8 444.4zM179.4 429.7C177.8 430.7 177.8 433.3 179.4 435.6C181 437.9 183.7 438.9 185 437.9C186.6 436.6 186.6 434 185 431.7C183.6 429.4 181 428.4 179.4 429.7z" />
</svg>
</div>
</button>
</a>
</div>
</div>
<hr />
{children}
<Scripts />
</body>
</html>
)
}
@@ -0,0 +1,45 @@
import { createFileRoute } from '@tanstack/react-router'
import { classifier, classificationRules, parsingConfig } from '~/utils/classifier'
export const Route = createFileRoute('/api/classify')({
component: RouteComponent,
server: {
handlers: {
POST: async ({ request }) => {
const body = await request.formData()
const fl = body.get("file") as File;
if (!fl) {
return new Response(JSON.stringify({"result": "you need to provide a file"}))
}
const buff = await fl.arrayBuffer()
const rawRes = await classifier.classify(
classificationRules,
parsingConfig,
{ fileContents: [new Uint8Array(buff)] },
)
const results = rawRes.items
let classification = ""
for (const result of results) {
if ("result" in result && result.result) {
classification += `
<div class="card bg-base-100 shadow-xl p-6 mb-4">
<div class="space-y-3">
<p><span class="font-semibold">📄 Document:</span> ${fl.name}</p>
<p><span class="font-semibold">🏷️ Type:</span> <span class="badge badge-primary">${result.result.type}</span></p>
<p><span class="font-semibold">📊 Confidence:</span> ${result.result.confidence*100}%</p>
<p><span class="font-semibold">💭 Reasoning:</span> ${result.result.reasoning}</p>
</div>
</div>
`
}
}
return new Response(JSON.stringify({"result": classification}))
},
},
},
})
function RouteComponent() {
return
}
+99
View File
@@ -0,0 +1,99 @@
import { createFileRoute } from '@tanstack/react-router'
import { useRef, useState } from 'react'
export const Route = createFileRoute('/')({
component: Home,
})
function Home() {
const [file, setFile] = useState<null | File>(null)
const fileInputRef = useRef<HTMLInputElement>(null)
const [reply, setReply] = useState<null | string>(null)
const [loading, setLoading] = useState<boolean>(false)
const handleFileChange = (event: React.ChangeEvent<HTMLInputElement>) => {
const selectedFile = event.target.files?.[0]
if (selectedFile) {
setFile(selectedFile)
}
}
const handleClearFile = () => {
if (file) {
setFile(null)
}
if (fileInputRef.current) {
fileInputRef.current.value = ''
}
if (reply) {
setReply(null)
}
}
const handleClassify = async () => {
if (!file) return
if (reply) {
setReply(null)
}
setLoading(true)
try {
const formData = new FormData()
formData.append('file', file)
const res = await fetch('/api/classify', {
method: 'POST',
body: formData,
})
const data = await res.json()
setReply(data.result)
} catch (error) {
console.error('Error:', error)
} finally {
setLoading(false)
}
}
return (
<div className="flex flex-col justify-center items-center gap-y-8">
<br />
<h1 className="text-xl font-bold text-gray-700">AI-Powered finacial document classification</h1>
<h2 className="text-lg font-semibold text-gray-500">Need help sorting out the financial documents jungle? Let our classification agent handle it!</h2>
<fieldset className="fieldset bg-base-100 border-base-300 rounded-box w-200 border p-4">
<legend className="fieldset-legend text-lg">Upload your financial document here</legend>
<label className="label flex justify-center">
<input type="file" className="file-input" onChange={handleFileChange} accept='application/pdf' ref={fileInputRef} />
</label>
</fieldset>
{file && (
<div className="flex flex-col justify-center items-center gap-y-8">
<p className="text-sm text-gray-600">Selected file: {file.name}</p>
<div className='grid grid-cols-2 gap-x-6'>
<button
type="button"
className='btn bg-gray-500 text-white shadow-lg hover:bg-gray-600 hover:shadow-xl rounded'
onClick={handleClassify}
>
Classify
</button>
<button
onClick={handleClearFile}
type="button"
className="px-4 py-2 bg-red-300 text-black rounded hover:bg-red-400 hover:shadow-xl shadow-lg"
>
Clear
</button>
</div>
</div>
)}
{loading && (
<span className="loading loading-spinner text-primary"></span>
)}
{reply && (
<div
className="max-w-2xl w-full"
dangerouslySetInnerHTML={{ __html: reply }}
/>
)}
</div>
)
}
@@ -0,0 +1,23 @@
import { LlamaClassify, ClassifierRule, ClassifyParsingConfiguration } from "llama-cloud-services"
export const classifier = new LlamaClassify(process.env.LLAMA_CLOUD_API_KEY);
export const classificationRules: ClassifierRule[] = [
{
description: "Shows a company's assets, liabilities, and shareholders' equity at a specific point in time, providing a snapshot of financial position.",
type: "balance_sheet"
},
{
description: "Reports cash inflows and outflows from operating, investing, and financing activities, highlighting liquidity and cash management.",
type: "cash_flow_statement"
},
{
description: "Summarizes revenues, expenses, and profits over a period, indicating financial performance and profitability.",
type: "income_statement"
},
];
export const parsingConfig: ClassifyParsingConfiguration = {
lang: "en",
max_pages: 20,
}
+33
View File
@@ -0,0 +1,33 @@
export const seo = ({
title,
description,
keywords,
image,
}: {
title: string
description?: string
image?: string
keywords?: string
}) => {
const tags = [
{ title },
{ name: 'description', content: description },
{ name: 'keywords', content: keywords },
{ name: 'twitter:title', content: title },
{ name: 'twitter:description', content: description },
{ name: 'twitter:creator', content: '@tannerlinsley' },
{ name: 'twitter:site', content: '@tannerlinsley' },
{ name: 'og:type', content: 'website' },
{ name: 'og:title', content: title },
{ name: 'og:description', content: description },
...(image
? [
{ name: 'twitter:image', content: image },
{ name: 'twitter:card', content: 'summary_large_image' },
{ name: 'og:image', content: image },
]
: []),
]
return tags
}
+22
View File
@@ -0,0 +1,22 @@
{
"include": ["**/*.ts", "**/*.tsx"],
"compilerOptions": {
"strict": true,
"esModuleInterop": true,
"jsx": "react-jsx",
"module": "ESNext",
"moduleResolution": "Bundler",
"lib": ["DOM", "DOM.Iterable", "ES2022"],
"isolatedModules": true,
"resolveJsonModule": true,
"skipLibCheck": true,
"target": "ES2022",
"allowJs": true,
"forceConsistentCasingInFileNames": true,
"baseUrl": ".",
"paths": {
"~/*": ["./src/*"]
},
"noEmit": true
}
}
+19
View File
@@ -0,0 +1,19 @@
import { tanstackStart } from '@tanstack/react-start/plugin/vite'
import { defineConfig } from 'vite'
import tsConfigPaths from 'vite-tsconfig-paths'
import viteReact from '@vitejs/plugin-react'
export default defineConfig({
server: {
port: 3000,
},
plugins: [
tsConfigPaths({
projects: ['./tsconfig.json'],
}),
tanstackStart({
srcDirectory: 'src',
}),
viteReact(),
],
})
+1 -1
View File
@@ -4,6 +4,6 @@ In this folder you will find several python notebooks that contain examples rega
- [LlamaParse](./parse/)
- [LlamaExtract](./extract/)
- [LlamaReport](./report/)
- [LlamaCloudIndex](./index/)
Follow the instructions in each notebook to get started!
+1
View File
@@ -0,0 +1 @@
sample_files/
@@ -0,0 +1,807 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "cell-0",
"metadata": {},
"source": [
"# Batch Parse with LlamaCloud Directories\n",
"\n",
"This notebook demonstrates how to use LlamaCloud's batch processing API to parse multiple files in a directory. The workflow includes:\n",
"\n",
"1. **Creating a Directory** - Set up a directory to organize your files\n",
"2. **Uploading Files** - Upload multiple files to the directory\n",
"3. **Starting a Batch Parse Job** - Kick off batch processing on all files\n",
"4. **Monitoring Progress** - Check the status and view results\n",
"\n",
"This is useful when you need to parse many documents at once, as the batch API handles the orchestration and provides progress tracking."
]
},
{
"cell_type": "markdown",
"id": "cell-1",
"metadata": {},
"source": [
"## Setup and Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-2",
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-cloud python-dotenv"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-3",
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"from dotenv import load_dotenv\n",
"import httpx\n",
"\n",
"# Load environment variables\n",
"load_dotenv()\n",
"\n",
"# Set your API key\n",
"LLAMA_CLOUD_API_KEY = os.environ.get(\"LLAMA_CLOUD_API_KEY\", \"llx-...\")\n",
"\n",
"# Optional: Set base URL (defaults to https://api.cloud.llamaindex.ai if not set)\n",
"LLAMA_CLOUD_BASE_URL = os.environ.get(\n",
" \"LLAMA_CLOUD_BASE_URL\", \"https://api.cloud.llamaindex.ai\"\n",
")\n",
"\n",
"# Optional: Set project_id if you have one, otherwise it will use your default project\n",
"PROJECT_ID = os.environ.get(\"LLAMA_CLOUD_PROJECT_ID\", None)\n",
"\n",
"print(\"✅ API key configured\")\n",
"print(f\" Base URL: {LLAMA_CLOUD_BASE_URL}\")"
]
},
{
"cell_type": "markdown",
"id": "cell-4",
"metadata": {},
"source": [
"## Setup HTTP Client\n",
"\n",
"Since the current version of the llama-cloud SDK has some issues with the beta endpoints, we'll use direct HTTP requests with httpx for reliability."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-5",
"metadata": {},
"outputs": [],
"source": [
"# Create HTTP client with authentication\n",
"headers = {\n",
" \"Authorization\": f\"Bearer {LLAMA_CLOUD_API_KEY}\",\n",
"}\n",
"\n",
"print(\"✅ HTTP client configured\")\n",
"print(f\" Using base URL: {LLAMA_CLOUD_BASE_URL}\")"
]
},
{
"cell_type": "markdown",
"id": "cell-6",
"metadata": {},
"source": [
"## Step 1: Create a Directory\n",
"\n",
"First, we'll create a directory to organize our files. Directories help you group related files together for batch processing."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-7",
"metadata": {},
"outputs": [],
"source": [
"from datetime import datetime\n",
"\n",
"# Create a directory with a timestamp in the name\n",
"timestamp = datetime.now().strftime(\"%Y%m%d-%H%M%S\")\n",
"directory_name = f\"batch-parse-demo-{timestamp}\"\n",
"\n",
"# Create directory using HTTP request\n",
"response = httpx.post(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/directories\",\n",
" headers=headers,\n",
" params={\"project_id\": PROJECT_ID},\n",
" json={\n",
" \"name\": directory_name,\n",
" \"description\": \"Demo directory for batch parse example\",\n",
" },\n",
" timeout=60.0,\n",
")\n",
"\n",
"if response.status_code in [200, 201]:\n",
" directory = response.json()\n",
" directory_id = directory[\"id\"]\n",
" project_id = directory[\"project_id\"]\n",
"\n",
" print(f\"✅ Created directory: {directory['name']}\")\n",
" print(f\" Directory ID: {directory_id}\")\n",
" print(f\" Project ID: {project_id}\")\n",
"else:\n",
" raise Exception(\n",
" f\"Failed to create directory: {response.status_code} - {response.text}\"\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "cell-8",
"metadata": {},
"source": [
"## Step 2: Upload Files to the Directory\n",
"\n",
"Now we'll upload some files to our directory. For this demo, we'll download some sample PDFs and upload them.\n",
"\n",
"You can replace these with your own files."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-9",
"metadata": {},
"outputs": [],
"source": [
"# Create a directory for sample files\n",
"import requests\n",
"\n",
"os.makedirs(\"sample_files\", exist_ok=True)\n",
"\n",
"# Sample documents to download\n",
"sample_docs = {\n",
" \"attention.pdf\": \"https://arxiv.org/pdf/1706.03762.pdf\",\n",
" \"bert.pdf\": \"https://arxiv.org/pdf/1810.04805.pdf\",\n",
"}\n",
"\n",
"# Download sample documents\n",
"for filename, url in sample_docs.items():\n",
" filepath = f\"sample_files/{filename}\"\n",
" if not os.path.exists(filepath):\n",
" print(f\"📥 Downloading {filename}...\")\n",
" response = requests.get(url)\n",
" if response.status_code == 200:\n",
" with open(filepath, \"wb\") as f:\n",
" f.write(response.content)\n",
" print(f\" ✅ Downloaded {filename}\")\n",
" else:\n",
" print(f\" ❌ Failed to download {filename}\")\n",
" else:\n",
" print(f\"📁 {filename} already exists\")\n",
"\n",
"print(\"\\n✅ Sample files ready!\")"
]
},
{
"cell_type": "markdown",
"id": "cell-10",
"metadata": {},
"source": [
"### Upload Files to Directory\n",
"\n",
"Now let's upload the files to our directory using the `upload_file_to_directory` endpoint."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-11",
"metadata": {},
"outputs": [],
"source": [
"uploaded_files = []\n",
"\n",
"# Workaround: Use direct HTTP requests instead of SDK due to SDK bug\n",
"import httpx\n",
"\n",
"for filename in os.listdir(\"sample_files\"):\n",
" if filename.endswith(\".pdf\"):\n",
" filepath = f\"sample_files/{filename}\"\n",
"\n",
" print(f\"📤 Uploading {filename}...\")\n",
"\n",
" # Upload file using direct HTTP request (SDK has a bug with file uploads)\n",
" with open(filepath, \"rb\") as f:\n",
" # Prepare the multipart form data correctly\n",
" files = {\"upload_file\": (filename, f, \"application/pdf\")}\n",
"\n",
" # Make the request directly\n",
" response = httpx.post(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/directories/{directory_id}/files/upload\",\n",
" params={\"project_id\": project_id},\n",
" files=files,\n",
" headers={\"Authorization\": f\"Bearer {LLAMA_CLOUD_API_KEY}\"},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code in [200, 201]:\n",
" directory_file = response.json()\n",
" uploaded_files.append(directory_file)\n",
" print(f\" ✅ Uploaded: {directory_file.get('display_name')}\")\n",
" print(f\" File ID: {directory_file.get('id')}\")\n",
" else:\n",
" print(f\" ❌ Upload failed: {response.status_code}\")\n",
" print(f\" Error: {response.text[:200]}\")\n",
"\n",
"print(f\"\\n✅ Uploaded {len(uploaded_files)} files to directory\")"
]
},
{
"cell_type": "markdown",
"id": "cell-12",
"metadata": {},
"source": [
"## Step 3: Create a Batch Parse Job\n",
"\n",
"Now that we have files in our directory, let's create a batch parse job to process them all at once.\n",
"\n",
"The batch processing API uses the same configuration as LlamaParse."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-13",
"metadata": {},
"outputs": [],
"source": [
"# Configure the parse job\n",
"# This configuration will apply to all files in the directory\n",
"job_config = {\n",
" \"job_name\": \"parse_raw_file_job\", # Must match the JobNames enum value\n",
" \"partitions\": {},\n",
" \"parameters\": {\n",
" \"type\": \"parse\",\n",
" \"lang\": \"en\",\n",
" \"fast_mode\": True,\n",
" },\n",
"}\n",
"\n",
"print(\"✅ Job configuration created\")\n",
"print(f\" Language: {job_config['parameters']['lang']}\")\n",
"print(f\" Fast mode: {job_config['parameters']['fast_mode']}\")"
]
},
{
"cell_type": "markdown",
"id": "cell-14",
"metadata": {},
"source": [
"### Submit the Batch Job\n",
"\n",
"Now let's submit the batch job to process all files in the directory."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-15",
"metadata": {},
"outputs": [],
"source": [
"print(f\"🚀 Submitting batch parse job for directory: {directory_id}\")\n",
"print(f\" Processing {len(uploaded_files)} files...\\n\")\n",
"\n",
"# Submit batch job using HTTP request\n",
"response = httpx.post(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/batch-processing\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" json={\n",
" \"directory_id\": directory_id,\n",
" \"job_config\": job_config,\n",
" \"page_size\": 100, # Number of files to fetch per batch\n",
" \"continue_as_new_threshold\": 10, # Workflow continuation threshold\n",
" },\n",
" timeout=60.0,\n",
")\n",
"\n",
"if response.status_code in [200, 201]:\n",
" batch_job = response.json()\n",
" batch_job_id = batch_job[\"id\"]\n",
"\n",
" print(\"✅ Batch job submitted successfully!\")\n",
" print(f\" Batch Job ID: {batch_job_id}\")\n",
" print(f\" Workflow ID: {batch_job.get('workflow_id')}\")\n",
" print(f\" Status: {batch_job.get('status')}\")\n",
" print(f\" Total Items: {batch_job.get('total_items')}\")\n",
"else:\n",
" raise Exception(\n",
" f\"Failed to create batch job: {response.status_code} - {response.text}\"\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "cell-16",
"metadata": {},
"source": [
"## Step 4: Monitor Job Progress\n",
"\n",
"Now let's monitor the batch job progress. We'll poll the status endpoint to see how the job is progressing."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-17",
"metadata": {},
"outputs": [],
"source": [
"import time\n",
"\n",
"\n",
"def print_job_status(status_data):\n",
" \"\"\"Helper function to print job status in a readable format.\"\"\"\n",
" job = status_data[\"job\"]\n",
" progress_pct = status_data[\"progress_percentage\"]\n",
"\n",
" print(f\"\\n{'='*60}\")\n",
" print(f\"Job Status: {job['status']}\")\n",
" print(f\"{'='*60}\")\n",
" print(f\"Total Items: {job['total_items']}\")\n",
" print(f\"Completed: {job['processed_items']}\")\n",
" print(f\"Failed: {job['failed_items']}\")\n",
" print(f\"Skipped: {job['skipped_items']}\")\n",
" print(f\"Progress: {progress_pct:.1f}%\")\n",
"\n",
" if job.get(\"completed_at\"):\n",
" print(f\"Completed At: {job['completed_at']}\")\n",
" elif job.get(\"started_at\"):\n",
" print(f\"Started At: {job['started_at']}\")\n",
"\n",
" print(f\"{'='*60}\")\n",
"\n",
"\n",
"# Poll for status updates\n",
"print(\"🔄 Monitoring batch job progress...\")\n",
"print(\n",
" \"Note: It may take a few seconds for the workflow to initialize and count files.\\n\"\n",
")\n",
"\n",
"max_polls = 60 # Maximum number of status checks (increased for longer jobs)\n",
"poll_interval = 10 # Seconds between checks\n",
"\n",
"for i in range(max_polls):\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/batch-processing/{batch_job_id}\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" status_data = response.json()\n",
" print_job_status(status_data)\n",
"\n",
" # Check if job is complete\n",
" job_status = status_data[\"job\"][\"status\"]\n",
" if job_status in [\"completed\", \"failed\", \"cancelled\"]:\n",
" print(f\"\\n✅ Job finished with status: {job_status}\")\n",
" break\n",
"\n",
" if i < max_polls - 1:\n",
" print(f\"\\n⏳ Waiting {poll_interval} seconds before next check...\")\n",
" time.sleep(poll_interval)\n",
" else:\n",
" print(f\"Error getting status: {response.status_code} - {response.text}\")\n",
" break\n",
"else:\n",
" print(f\"\\n⚠️ Reached maximum polling attempts. Job may still be running.\")"
]
},
{
"cell_type": "markdown",
"id": "cell-18",
"metadata": {},
"source": [
"## Step 5: View Job Items\n",
"\n",
"Let's look at the individual items in the batch job to see which files were processed successfully."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-19",
"metadata": {},
"outputs": [],
"source": [
"# Get all items in the batch job\n",
"response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/batch-processing/{batch_job_id}/items\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id, \"limit\": 100},\n",
" timeout=60.0,\n",
")\n",
"\n",
"if response.status_code == 200:\n",
" items_response = response.json()\n",
"\n",
" print(f\"\\n📋 Batch Job Items ({items_response['total_size']} total)\")\n",
" print(f\"{'='*80}\\n\")\n",
"\n",
" for item in items_response[\"items\"]:\n",
" status_emoji = (\n",
" \"✅\"\n",
" if item[\"status\"] == \"completed\"\n",
" else \"❌\"\n",
" if item[\"status\"] == \"failed\"\n",
" else \"⏳\"\n",
" )\n",
" print(f\"{status_emoji} {item['item_name']}\")\n",
" print(f\" Status: {item['status']}\")\n",
" print(f\" Item ID: {item['item_id']}\")\n",
"\n",
" if item.get(\"error_message\"):\n",
" print(f\" Error: {item['error_message']}\")\n",
"\n",
" print()\n",
"else:\n",
" print(f\"Error listing items: {response.status_code} - {response.text}\")"
]
},
{
"cell_type": "markdown",
"id": "cell-20",
"metadata": {},
"source": [
"## Step 6: Retrieve Processing Results\n",
"\n",
"For each completed file, we can retrieve the processing results to see where the parsed output is stored."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-21",
"metadata": {},
"outputs": [],
"source": [
"# Get processing results for a specific item\n",
"if items_response[\"items\"]:\n",
" first_item = items_response[\"items\"][0]\n",
"\n",
" print(f\"\\n🔍 Processing results for: {first_item['item_name']}\")\n",
" print(f\"{'='*80}\\n\")\n",
"\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/batch-processing/items/{first_item['item_id']}/processing-results\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" results = response.json()\n",
"\n",
" print(f\"Item: {results['item_name']}\")\n",
" print(f\"Total processing runs: {len(results['processing_results'])}\\n\")\n",
"\n",
" for i, result in enumerate(results[\"processing_results\"], 1):\n",
" print(f\"Run {i}:\")\n",
" print(f\" Job Type: {result['job_type']}\")\n",
" print(f\" Processed At: {result['processed_at']}\")\n",
" print(f\" Parameters Hash: {result['parameters_hash']}\")\n",
"\n",
" if result.get(\"output_s3_path\"):\n",
" print(f\" Output S3 Path: {result['output_s3_path']}\")\n",
"\n",
" if result.get(\"output_metadata\"):\n",
" print(f\" Output Metadata: {result['output_metadata']}\")\n",
"\n",
" print()\n",
" else:\n",
" print(f\"Error getting results: {response.status_code} - {response.text}\")"
]
},
{
"cell_type": "markdown",
"id": "cell-22",
"metadata": {},
"source": [
"## Optional: List All Batch Jobs\n",
"\n",
"You can also list all batch jobs in your project to see the history of batch processing operations."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-23",
"metadata": {},
"outputs": [],
"source": [
"# List all parse jobs in the project\n",
"response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/beta/batch-processing\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id, \"job_type\": \"parse\", \"limit\": 10},\n",
" timeout=60.0,\n",
")\n",
"\n",
"if response.status_code == 200:\n",
" jobs_response = response.json()\n",
"\n",
" print(f\"\\n📊 Recent Batch Parse Jobs ({jobs_response['total_size']} total)\")\n",
" print(f\"{'='*80}\\n\")\n",
"\n",
" for job in jobs_response[\"items\"]:\n",
" status_emoji = (\n",
" \"✅\"\n",
" if job[\"status\"] == \"completed\"\n",
" else \"❌\"\n",
" if job[\"status\"] == \"failed\"\n",
" else \"⏳\"\n",
" )\n",
" print(f\"{status_emoji} Job ID: {job['id']}\")\n",
" print(f\" Status: {job['status']}\")\n",
" print(f\" Directory: {job['directory_id']}\")\n",
" print(f\" Total Items: {job['total_items']}\")\n",
" print(f\" Completed: {job['processed_items']}\")\n",
" print(f\" Created: {job['created_at']}\")\n",
" print()\n",
"else:\n",
" print(f\"Error listing jobs: {response.status_code} - {response.text}\")"
]
},
{
"cell_type": "markdown",
"id": "uug7591rkq",
"metadata": {},
"source": [
"## Step 7: Retrieve Parsed Text Results\n",
"\n",
"Once the batch job is complete, each BatchJobItem will have a `job_id` field that maps to a parse job ID. We can use this ID with the standard parse client methods to fetch the actual parsed text results."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "vpp0vxtc0y",
"metadata": {},
"outputs": [],
"source": [
"# Get all completed items and their job IDs\n",
"completed_items = [\n",
" item for item in items_response[\"items\"] if item[\"status\"] == \"completed\"\n",
"]\n",
"\n",
"print(f\"📄 Found {len(completed_items)} completed items\\n\")\n",
"print(f\"{'='*80}\\n\")\n",
"\n",
"# Display the job_id for each completed item\n",
"for item in completed_items:\n",
" print(f\"📝 {item['item_name']}\")\n",
" print(f\" Item ID: {item['item_id']}\")\n",
" print(f\" Parse Job ID: {item['job_id']}\")\n",
" print()"
]
},
{
"cell_type": "markdown",
"id": "4gck6hwpnl6",
"metadata": {},
"source": [
"### Fetch Parsed Text for a Specific Document\n",
"\n",
"Now let's use the `job_id` to retrieve the actual parsed text content using the parse client methods."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "g191kvgxxvk",
"metadata": {},
"outputs": [],
"source": [
"# Get the parsed text for the first completed item\n",
"if completed_items:\n",
" first_completed = completed_items[0]\n",
"\n",
" print(f\"📖 Retrieving parsed text for: {first_completed['item_name']}\")\n",
" print(f\" Using Parse Job ID: {first_completed['job_id']}\\n\")\n",
" print(f\"{'='*80}\\n\")\n",
"\n",
" # Use the job_id to fetch the parse result\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/parsing/job/{first_completed['job_id']}/result/text\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" parse_result = response.text\n",
"\n",
" print(f\"✅ Retrieved parsed text ({len(parse_result)} characters)\\n\")\n",
"\n",
" # Display first 1000 characters as a preview\n",
" print(\"Preview (first 1000 characters):\")\n",
" print(\"-\" * 80)\n",
" print(parse_result[:1000])\n",
" print(\"-\" * 80)\n",
"\n",
" if len(parse_result) > 1000:\n",
" print(f\"\\n... and {len(parse_result) - 1000} more characters\")\n",
" else:\n",
" print(\n",
" f\"Error retrieving parse result: {response.status_code} - {response.text}\"\n",
" )\n",
"else:\n",
" print(\"⚠️ No completed items found to retrieve results from\")"
]
},
{
"cell_type": "markdown",
"id": "2olccb4l8fj",
"metadata": {},
"source": [
"### Retrieve Parsed Results in Other Formats\n",
"\n",
"You can also retrieve the parsed results in JSON or Markdown format using different client methods."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "lcqsfxiw0sr",
"metadata": {},
"outputs": [],
"source": [
"if completed_items:\n",
" first_completed = completed_items[0]\n",
"\n",
" print(\n",
" f\"📋 Retrieving parse results in different formats for: {first_completed['item_name']}\\n\"\n",
" )\n",
"\n",
" # Get as JSON (includes structured data with pages, images, etc.)\n",
" print(\"1️⃣ Retrieving as JSON...\")\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/parsing/job/{first_completed['job_id']}/result/json\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" json_result = response.json()\n",
" print(f\" ✅ JSON result with {len(json_result['pages'])} pages\")\n",
" print(f\" Keys: {list(json_result.keys())}\\n\")\n",
" else:\n",
" print(f\" Error: {response.status_code}\\n\")\n",
"\n",
" # Get as Markdown\n",
" print(\"2️⃣ Retrieving as Markdown...\")\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/parsing/job/{first_completed['job_id']}/result/markdown\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" markdown_result = response.text\n",
" print(f\" ✅ Markdown result ({len(markdown_result)} characters)\\n\")\n",
"\n",
" # Display markdown preview\n",
" print(\"Markdown Preview (first 500 characters):\")\n",
" print(\"-\" * 80)\n",
" print(markdown_result[:500])\n",
" print(\"-\" * 80)\n",
"\n",
" if len(markdown_result) > 500:\n",
" print(f\"\\n... and {len(markdown_result) - 500} more characters\")\n",
" else:\n",
" print(f\" Error: {response.status_code}\")\n",
"else:\n",
" print(\"⚠️ No completed items found to retrieve results from\")"
]
},
{
"cell_type": "markdown",
"id": "lr61wqkfq3",
"metadata": {},
"source": [
"### Batch Process All Parsed Results\n",
"\n",
"You can also loop through all completed items to retrieve and process all the parsed results."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "kltydf9xzkl",
"metadata": {},
"outputs": [],
"source": [
"# Process all completed items\n",
"print(f\"🔄 Processing all {len(completed_items)} completed items...\\n\")\n",
"print(f\"{'='*80}\\n\")\n",
"\n",
"all_results = {}\n",
"\n",
"for item in completed_items:\n",
" print(f\"📄 Processing: {item['item_name']}\")\n",
" print(f\" Parse Job ID: {item['job_id']}\")\n",
"\n",
" try:\n",
" # Retrieve the parsed text for this item\n",
" response = httpx.get(\n",
" f\"{LLAMA_CLOUD_BASE_URL}/api/v1/parsing/job/{item['job_id']}/result/text\",\n",
" headers=headers,\n",
" params={\"project_id\": project_id},\n",
" timeout=60.0,\n",
" )\n",
"\n",
" if response.status_code == 200:\n",
" parsed_text = response.text\n",
"\n",
" all_results[item[\"item_name\"]] = {\n",
" \"job_id\": item[\"job_id\"],\n",
" \"text\": parsed_text,\n",
" \"length\": len(parsed_text),\n",
" }\n",
"\n",
" print(f\" ✅ Retrieved {len(parsed_text)} characters\")\n",
" else:\n",
" all_results[item[\"item_name\"]] = {\n",
" \"job_id\": item[\"job_id\"],\n",
" \"error\": f\"HTTP {response.status_code}\",\n",
" }\n",
" print(f\" ❌ Error: HTTP {response.status_code}\")\n",
"\n",
" except Exception as e:\n",
" print(f\" ❌ Error: {str(e)}\")\n",
" all_results[item[\"item_name\"]] = {\"job_id\": item[\"job_id\"], \"error\": str(e)}\n",
"\n",
" print()\n",
"\n",
"print(f\"{'='*80}\")\n",
"print(f\"\\n✅ Processed {len(all_results)} items\")\n",
"print(f\"\\nSummary:\")\n",
"for name, result in all_results.items():\n",
" if \"error\" in result:\n",
" print(f\" ❌ {name}: Error - {result['error']}\")\n",
" else:\n",
" print(f\" ✅ {name}: {result['length']:,} characters\")"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
@@ -7,7 +7,7 @@
"source": [
"# Extraction and Analysis over a Fidelity Multi-Fund Annual Report\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/extract/asset_manager_fund_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/asset_manager_fund_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook we show you how to create an agentic document workflow over a complex document that contains annual reports for multiple funds - each fund reports financials in a standardized reporting structure, and it's all consolidated in the same document.\n",
"\n",
@@ -7,7 +7,7 @@
"source": [
"# Automotive Equity Research: A Multi-Step Agentic Workflow\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/extract/automotive_sector_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/automotive_sector_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates an endtoend agentic workflow using LlamaExtract and the LlamaIndex eventdriven workflow framework for automotive sector analysis.\n",
"\n",
Binary file not shown.

After

Width:  |  Height:  |  Size: 287 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 769 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 942 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.5 MiB

+508
View File
@@ -0,0 +1,508 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "a7oq3cfnync",
"metadata": {},
"source": [
"# Extracting Repeating Entities from Documents\n",
"\n",
"This notebook demonstrates how to use the `PER_TABLE_ROW` extraction target to extract structured data from documents containing repeating entities like tables, lists, or catalogs.\n",
"\n",
"## Why Use the Tabular Extraction Target?\n",
"\n",
"`PER_DOC` (refer to the table below for a quick overview of the different extraction targets) is the default extraction target in LlamaExtract, which looks at the entire document's context when doing an extraction. When extracting lists of entities, LLM-based extraction has a critical failure mode — it often **only extracts the first few tens of entries** from a long list. This happens because LLMs have limited attention spans for repetitive data. Document-level extraction doesn't guarantee exhaustive coverage, and long lists lead to incomplete extractions.\n",
"\n",
"**The Solution**: `PER_TABLE_ROW` solves this by processing each entity individually or in smaller batches, ensuring **exhaustive extraction** of all entries regardless of list length.\n",
"\n",
"### Entity-Level Extraction\n",
"\n",
"When using `extraction_target=ExtractTarget.PER_TABLE_ROW`, you define a schema for a **single entity** (e.g., one hospital, one product, one invoice line item), not the full document. LlamaExtract automatically:\n",
"- Detects the formatting patterns that distinguish individual entities (table rows, list items, section headers, etc.)\n",
"- Applies your schema to each identified entity\n",
"- Returns a `list[YourSchema]` with one object per entity\n",
"\n",
"This approach is ideal when each entity locally contains all the information needed for your schema.\n",
"\n",
"### Choosing the Right Extraction Target\n",
"\n",
"| Extraction Target | Best For | Returns |\n",
"|-------------------|----------|---------|\n",
"| `PER_DOC` | Single-entity documents, summaries, or short lists | One JSON object for entire document |\n",
"| `PER_PAGE` | Multi-page documents where each page is independent | One JSON object per page |\n",
"| `PER_TABLE_ROW` | **Long lists, tables, catalogs with repeating entities** | List of JSON objects (one per entity) |\n",
"\n",
"📖 For more details, see the [Extraction Target documentation](https://developers.llamaindex.ai/python/cloud/llamaextract/features/concepts/#extraction-target)."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "9427d1de",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"\n",
"\n",
"# Load environment variables (put LLAMA_CLOUD_API_KEY in your .env file)\n",
"load_dotenv(override=True)\n",
"\n",
"# Optionally, add your project id/organization id\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"id": "4426b360",
"metadata": {},
"source": [
"## Table of Hospitals by County and Insurance Plans\n",
"\n",
"We have a PDF document with a list of hospitals by county and different insurance plans offered by Blue Shield of California. \n",
"\n",
"\n",
"![First few entries from the PDF](./data/tables/bsc_page1.png)"
]
},
{
"cell_type": "markdown",
"id": "c86sjymhn1r",
"metadata": {},
"source": [
"We want to extract each hospital from this table along with a list of applicable insurance plans. \n",
"\n",
"### Example 1: Structured Table\n",
"\n",
"This is an ideal use case for `PER_TABLE_ROW` extraction:\n",
"- **Clear structure**: The document has explicit table formatting with rows and columns\n",
"- **Repeating entities**: Each row represents one hospital with consistent attributes\n",
"- **Local information**: All data for each hospital (county, name, plans) is contained within its row\n",
"\n",
"Notice that our `Hospital` schema describes a **single hospital**, not the full document. LlamaExtract will return a `list[Hospital]` with one entry per table row."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7c61a802",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"\n",
"\n",
"class Hospital(BaseModel):\n",
" \"\"\"List of hospitals by county available for different BSC plans\"\"\"\n",
"\n",
" county: str = Field(description=\"County name\")\n",
" hospital_name: str = Field(description=\"Name of the hospital\")\n",
" plan_names: list[str] = Field(\n",
" description=\"List of plans available at the hospital. One of: Trio HMO, SaveNet, Access+ HMO, BlueHPN PPO, Tandem PPO, PPO\"\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b8a69b7a",
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services.extract import ExtractConfig, ExtractMode, ExtractTarget\n",
"\n",
"\n",
"result = await llama_extract.aextract(\n",
" data_schema=Hospital,\n",
" files=\"./data/tables/BSC-Hospital-List-by-County.pdf\",\n",
" config=ExtractConfig(\n",
" extraction_mode=ExtractMode.PREMIUM,\n",
" extraction_target=ExtractTarget.PER_TABLE_ROW,\n",
" parse_model=\"anthropic-sonnet-4.5\",\n",
" ),\n",
")"
]
},
{
"cell_type": "markdown",
"id": "43722cda",
"metadata": {},
"source": [
"### Results"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "95b5aca6",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"380"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"len(result.data)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1e355770",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[{'county': 'Alameda',\n",
" 'hospital_name': 'Alameda Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Med Ctr Herrick Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Med Ctr Alta Bates Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Med Ctr Summit Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Medical Center',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'BHC Fremont Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Centre For Neuro Skills San Francisco',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Eden Medical Center',\n",
" 'plan_names': ['Trio HMO', 'Access+ HMO', 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Fairmont Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Highland Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']}]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"result.data[:10]"
]
},
{
"cell_type": "markdown",
"id": "e28f0de8",
"metadata": {},
"source": [
"![](./data/tables/bsc_results.png)"
]
},
{
"cell_type": "markdown",
"id": "di156pb7s6j",
"metadata": {},
"source": [
"**Success!** We extracted all **380 hospitals** from the multi-page PDF. Each entity was correctly parsed with its county, hospital name, and applicable insurance plans. With `PER_DOC`, we would likely have only gotten the first 20-30 entries."
]
},
{
"cell_type": "markdown",
"id": "gelvl6db268",
"metadata": {},
"source": [
"## Extracting from a Toy Catalog\n",
"\n",
"### Example 2: Semi-Structured List\n",
"\n",
"The `PER_TABLE_ROW` extraction target also works well for documents that aren't explicit tables but have similar properties:\n",
"- **Ordered listing**: The toys are listed sequentially with visual separation (section headers, spacing)\n",
"- **Repeating pattern**: Each toy entry has a consistent structure (code, name, specs, description)\n",
"- **Local information**: All attributes for each toy are grouped together in its entry\n",
"\n",
"Even though this isn't a traditional table format, each toy entity locally contains all the information needed for our schema. LlamaExtract detects the formatting patterns that distinguish each toy and extracts them as separate entities.\n",
"\n",
"![](./data/tables/toy_catalog_page.png)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "8cf0b2db",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"\n",
"\n",
"class ToyCatalog(BaseModel):\n",
" \"\"\"Product information from a toy catalog.\"\"\"\n",
"\n",
" section_name: str = Field(\n",
" description=\"The name of the toy section (e.g. Table Toys, Active Toys).\"\n",
" )\n",
" product_code: str = Field(\n",
" description=\"The unique product code for the toy (e.g., GA457).\"\n",
" )\n",
" toy_name: str = Field(description=\"The name of the toy.\")\n",
" age_range: str = Field(\n",
" description=\"The recommended age range for the toy (e.g., 6 +, 4 +).\",\n",
" )\n",
" player_range: str = Field(\n",
" description=\"The number of players the toy is designed for (e.g., 2, 2-4, 1-6).\",\n",
" )\n",
" material: str = Field(\n",
" description=\"The primary material(s) the toy is made of (e.g., wood, cardboard).\",\n",
" )\n",
" description: str = Field(\n",
" description=\"A brief description of the toy and its components and dimensions.\",\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "mysu1i2qo9e",
"metadata": {},
"source": [
"### Results\n",
"\n",
"Again, our schema represents a **single toy product**, not the entire catalog. The system will return a `list[ToyCatalog]` with one entry per toy."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "5b38b806",
"metadata": {},
"outputs": [],
"source": [
"result = await llama_extract.aextract(\n",
" data_schema=ToyCatalog,\n",
" files=\"./data/tables/Click-BS-Toys-Catalogue-2024.pdf\",\n",
" config=ExtractConfig(\n",
" extraction_mode=ExtractMode.PREMIUM,\n",
" extraction_target=ExtractTarget.PER_TABLE_ROW,\n",
" parse_model=\"anthropic-sonnet-4.5\",\n",
" ),\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "91aface0",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"153"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"len(result.data)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "51278736",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[{'section_name': 'Table Toys',\n",
" 'product_code': 'GA457',\n",
" 'toy_name': 'Dots and Boxes',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': 'base 17x17 cm\\n50 border pieces 4x1,2x0,3 cm\\n34 trees 2,6x1,4 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA456',\n",
" 'toy_name': '3 In a Row',\n",
" 'age_range': '8+',\n",
" 'player_range': '2',\n",
" 'material': 'wood, pine, cardboard',\n",
" 'description': 'base 24x22,5x2,5 cm\\n30 cards 5,5x5 cm\\n6 chips'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA467',\n",
" 'toy_name': 'Which Cow am i?',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood, beech',\n",
" 'description': '2 cow bases 56x4x4,5 cm\\n16 cards 4x5 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA460',\n",
" 'toy_name': 'Balance Bunnies',\n",
" 'age_range': '4+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': '1 base 35x12x25 cm\\n7 bunnies 7 foxes\\n1 dice 3 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA462',\n",
" 'toy_name': 'Color Combination Race',\n",
" 'age_range': '4+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood, cardboard',\n",
" 'description': 'base 6,5x6,5x15 cm, rings 5,5x5,5x0,5 mm\\ncardholder 6x6x2 cm, cards 5,5x5,5 cm\\ncolor cards Ø 15,5 cm - Ø 7 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA465',\n",
" 'toy_name': 'Plop It',\n",
" 'age_range': '6+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood, elastic, cardboard',\n",
" 'description': 'Catch the right balls and plop them in the net!\\n* 2 ploppers 8x5 cm\\n* 2 net holders Ø 5cm, length 55 cm\\n* 6 cards 1,5x2,5 cm, 30 balls Ø 2,5 cm\\n* 1 rope 120 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA466',\n",
" 'toy_name': 'Whack a Shape',\n",
" 'age_range': '4+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood',\n",
" 'description': '* base 38,5x15,5 cm\\n* 2 stands 36 half balls, 4 hammers\\n* 1 dice 2,5 cm\\n* 4 cards'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA458',\n",
" 'toy_name': 'Sling Puck | Table Hockey',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': '* double sides base 39x21x3 cm\\n* 10 chips Ø 2,5 cm\\n* 2 pushers 4x4x3 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA039',\n",
" 'toy_name': 'DIY Birdhouse',\n",
" 'age_range': '3+',\n",
" 'player_range': '1',\n",
" 'material': 'wood',\n",
" 'description': '* house 9x9x13 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA319',\n",
" 'toy_name': 'Triangle Domino',\n",
" 'age_range': '6+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood',\n",
" 'description': '* 35 triangles 10x10 x10 cm'}]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"result.data[:10]"
]
},
{
"cell_type": "markdown",
"id": "d1810c0a",
"metadata": {},
"source": [
"![](./data/tables/toy_catalog_results.png)"
]
},
{
"cell_type": "markdown",
"id": "ezur9gnhmsb",
"metadata": {},
"source": [
"**Success!** Despite the semi-structured format, we extracted all **152 toy products** from the catalog (there's an extra repeated extracted toy from the Appendix section). LlamaExtract automatically detected the visual patterns separating each toy entry and applied our schema to each one."
]
},
{
"cell_type": "markdown",
"id": "aeyr3io29u",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"The `PER_TABLE_ROW` extraction target is powerful for extracting repeating structured entities from documents. Key takeaways:\n",
"\n",
"1. **Schema design**: Define your schema for a single entity, not the full document. The system returns `list[YourSchema]`.\n",
"\n",
"2. **Works with various formats**: Not just traditional tables—any document with distinguishable repeating entities (bullets, numbering, headers, visual separation, etc.). The common requirement is that each entity should contain all the necessary data for your schema within its local context.\n",
"\n",
"3. **Automatic pattern detection**: LlamaExtract identifies the formatting patterns that distinguish entities and applies your schema to each one."
]
}
],
"metadata": {
"kernelspec": {
"display_name": ".venv",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -7,7 +7,7 @@
"source": [
"# Dynamic Section Retrieval with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/parse/advanced_rag/dynamic_section_retrieval.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/advanced_rag/dynamic_section_retrieval.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook showcases a concept called \"dynamic section retrieval\".\n",
"\n",
+183
View File
@@ -0,0 +1,183 @@
"""
Example: Batch Processing a Folder of PDFs with LlamaParse
This script demonstrates how to process multiple PDFs from a folder
using LlamaParse with controlled concurrency using asyncio and semaphores.
Usage:
python batch_parse_folder.py --input-dir ./pdfs --max-concurrent 5
"""
import asyncio
import argparse
from pathlib import Path
from typing import List, Dict, Any
from datetime import datetime
from dotenv import load_dotenv
import os
from llama_cloud_services import LlamaParse
# Load environment variables from .env file
load_dotenv()
async def parse_single_file(
parser: LlamaParse,
file_path: Path,
semaphore: asyncio.Semaphore,
) -> Dict[str, Any]:
"""
Parse a single PDF file with concurrency control.
Args:
parser: LlamaParse instance
file_path: Path to the PDF file
semaphore: Semaphore to control concurrent requests
Returns:
Dictionary with file info and parse result
"""
async with semaphore:
try:
print(f"Starting parse: {file_path.name}")
result = await parser.aparse(str(file_path))
print(f"✓ Completed: {file_path.name} ({len(result.pages)} pages)")
return {
"file": file_path.name,
"status": "success",
"result": result,
"pages": len(result.pages) if result.pages else 0,
}
except Exception as e:
print(f"✗ Error parsing {file_path.name}: {str(e)}")
return {
"file": file_path.name,
"status": "error",
"error": str(e),
}
async def parse_folder(
input_dir: Path,
max_concurrent: int = 5,
api_key: str = None,
) -> List[Dict[str, any]]:
"""
Parse all PDFs in a folder with controlled concurrency.
Args:
input_dir: Directory containing PDF files
max_concurrent: Maximum number of concurrent parse operations
api_key: LlamaCloud API key (loaded from .env file)
Returns:
List of parse results for each file
"""
# Find all PDF files
pdf_files = list(input_dir.glob("*.pdf"))
if not pdf_files:
print(f"No PDF files found in {input_dir}")
return []
print(f"Found {len(pdf_files)} PDF files to parse")
# Initialize parser
parser = LlamaParse(
api_key=api_key,
num_workers=1, # We control concurrency with semaphore
show_progress=False, # We'll show our own progress
)
# Create semaphore to limit concurrent requests
semaphore = asyncio.Semaphore(max_concurrent)
# Create tasks for all files
tasks = [parse_single_file(parser, pdf_file, semaphore) for pdf_file in pdf_files]
# Run all tasks concurrently (but limited by semaphore)
print(
f"Processing {len(tasks)} files with max {max_concurrent} concurrent operations..."
)
start_time = datetime.now()
results = await asyncio.gather(*tasks)
end_time = datetime.now()
duration = (end_time - start_time).total_seconds()
# Process results
successful = [
r for r in results if isinstance(r, dict) and r.get("status") == "success"
]
failed = [r for r in results if isinstance(r, dict) and r.get("status") == "error"]
# Print summary
print("PARSE SUMMARY \n")
print(f"Total files: {len(pdf_files)}")
print(f"Successful: {len(successful)}")
print(f"Failed: {len(failed)}")
print(f"Total time: {duration:.2f} seconds")
print(f"Average time per file: {duration / len(pdf_files):.2f} seconds")
if failed:
print("\nFailed files:")
for result in failed:
print(f" - {result['file']}: {result.get('error', 'Unknown error')}")
return results
def main():
"""Main entry point for the script."""
parser = argparse.ArgumentParser(
description="Batch process PDFs in a folder with LlamaParse"
)
parser.add_argument(
"--input-dir",
type=str,
required=True,
help="Directory containing PDF files to parse",
)
parser.add_argument(
"--max-concurrent",
type=int,
default=5,
help="Maximum number of concurrent parse operations (default: 5)",
)
args = parser.parse_args()
input_dir = Path(args.input_dir)
# Validate input directory
if not input_dir.exists():
print(f"Error: Input directory does not exist: {input_dir}")
return
if not input_dir.is_dir():
print(f"Error: Input path is not a directory: {input_dir}")
return
# Get API key from environment (loaded from .env file)
api_key = os.getenv("LLAMA_CLOUD_API_KEY")
if not api_key:
print("Error: LLAMA_CLOUD_API_KEY not found. Please set it in your .env file")
return
# Run async function
asyncio.run(
parse_folder(
input_dir=input_dir,
max_concurrent=args.max_concurrent,
api_key=api_key,
)
)
if __name__ == "__main__":
main()
+1 -1
View File
@@ -6,7 +6,7 @@
"source": [
"# Advanced RAG with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/parse/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook is a complete walkthrough for using LlamaParse with advanced indexing/retrieval techniques in LlamaIndex over the Apple 10K Filing. \n",
"\n",
+2 -2
View File
@@ -6,7 +6,7 @@
"source": [
"# RAG with Excel Spreadsheet using LlamaPrase\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you using LlamaParse with Excel Spreadsheet.\n",
"\n",
@@ -43,7 +43,7 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"api_key = \"llx-jwAQZL8T38onyL9hKBOXyRtnuCU0Fk3z7tmDhIT3L0GEfohJ\" # get from cloud.llamaindex.ai"
"api_key = \"llx-...\" # get from cloud.llamaindex.ai"
]
},
{
+1 -1
View File
@@ -7,7 +7,7 @@
"source": [
"# Download Charts\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_get_charts.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_get_charts.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates how to download charts from a document using the result object.\n",
"\n",
+4 -4
View File
@@ -6,7 +6,7 @@
"source": [
"# LlamaParse - Fast checking Insurance Contract for Coverage\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook we will look at how LlamaParse can be used to extract structured coverage information from an insurance policy.\n",
"\n",
@@ -36,7 +36,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"## Download an insurance policy fron IRDAI\n",
"## Download an insurance policy from IRDAI\n",
"\n",
"The Insurance Regulatory and Development Authority of India (IRDAI) maintains a great resource: https://policyholder.gov.in/web/guest/non-life-insurance-products where all insurance policies available in India are publicly available for download! Let's download a complex health insurance policy as an example."
]
@@ -228,11 +228,11 @@
" result_type=\"markdown\",\n",
" system_prompt_append=\"\"\"\n",
"This document is an insurance policy.\n",
"When a benefits/coverage/exlusion is describe in the document ammend to it add a text in the follwing benefits string format (where coverage could be an exclusion).\n",
"When a benefits/coverage/exlusion is describe in the document amend to it add a text in the following benefits string format (where coverage could be an exclusion).\n",
"\n",
"For {nameofrisk} and in this condition {whenDoesThecoverageApply} the coverage is {coverageDescription}. \n",
" \n",
"If the document contain a benefits TABLE that describe coverage amounts, do not ouput it as a table, but instead as a list of benefits string.\n",
"If the document contain a benefits TABLE that describe coverage amounts, do not output it as a table, but instead as a list of benefits string.\n",
" \n",
"\"\"\",\n",
").aparse(\"./policy.pdf\")\n",
+1 -1
View File
@@ -7,7 +7,7 @@
"source": [
"# LlamaParse `JobResult` Tour\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"The `JobResult` object is the main object returned by the LlamaParse API. It contains all the information about the job, including the parsed data, metadata, and any errors.\n",
"\n",
+1 -1
View File
@@ -9,7 +9,7 @@
"\n",
"LlamaParse supports users to specify a `language` parameter before uploading documents, giving users better OCR capabilities over non-English PDFs, parsing images into more accurate representations.\n",
"\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_cloud_services/blob/main/llama_parse/base.py.\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_cloud_services/blob/main/py/llama_cloud_services/parse/base.py.\n",
"\n",
"This notebook shows a demo of this in action. \n",
"\n",
-1
View File
@@ -75,7 +75,6 @@
" adaptive_long_table=True,\n",
" outlined_table_extraction=True,\n",
" output_tables_as_HTML=True,\n",
" api_key=\"llx-jwAQZL8T38onyL9hKBOXyRtnuCU0Fk3z7tmDhIT3L0GEfohJ\",\n",
")\n",
"\n",
"result = await parser.aparse(\"./dcf_template.xlsx\")\n",
+1 -1
View File
@@ -4,7 +4,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/excel/o1_excel_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/excel/o1_excel_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
@@ -740,7 +740,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"In this example, these pages aren't going to be that different when parsed, but we can verify which pages triggered auto-made by looking at the [JSON output](https://github.com/run-llama/llama_cloud_services/blob/main/examples/demo_json_tour.ipynb) of LlamaParse:"
"In this example, these pages aren't going to be that different when parsed, but we can verify which pages triggered auto-made by looking at the [JSON output](https://github.com/run-llama/llama_cloud_services/blob/main/examples/parse/demo_json_tour.ipynb) of LlamaParse:"
]
},
{
-762
View File
@@ -1,762 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Report Generation with LlamaReport\n",
"\n",
"In this notebook, we'll walk through the basic process of generating a report with LlamaReport, and highlight some of the key features of the library.\n",
"\n",
"TLDR:\n",
"1. Download source data to use as knowledge base for the report\n",
"2. Kick off report generation with a template\n",
"3. Get the plan and review/accept/reject suggestions\n",
"4. Get the final report\n",
"5. Review/accept/reject suggestions to edit the final report\n",
"6. Print the final report"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-cloud-services"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 1. Download Source Data\n",
"\n",
"Here, we download the `Attention is All You Need` paper as a PDF.\n",
"\n",
"LlamaReport currently supports up to 5 files as input, and essentially any file type that can be parsed by LlamaParse.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/1706.03762.pdf\" -O \"./attention.pdf\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 2. Kick off Report Generation\n",
"\n",
"Here, we kick off report generation with a template.\n",
"\n",
"The template can either be a string or a file path, but here we'll use a string.\n",
"\n",
"In our experiments, anything works as a template, but some general guidelines:\n",
"\n",
"- Use markdown formatting + instructions in each section to guide the report generation\n",
"- If using an existing file as a template, provide extra instructions to guide the report generation\n",
"\n",
"**NOTE:** Since we are in a notebook, we will use async functions and `await` throughout. Synchronous methods that work without `await` are available by just removing the `a` from the method name and removing the `await` keyword."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaReport\n",
"\n",
"llama_report = LlamaReport(\n",
" api_key=\"llx-...\",\n",
")\n",
"\n",
"report_client = await llama_report.acreate_report(\n",
" name=\"my_cool_report_on_attention\",\n",
" # can pass in file paths or bytes\n",
" input_files=[\"./attention.pdf\"],\n",
" template_text=\"\"\"\\\n",
"# [Some title]\\n\\n\n",
"## TLDR\\n\n",
"A quick summary of the paper.\\n\\n\n",
"## Details\\n\n",
"More details about the paper, possibly more than one section here.\\n\n",
"\"\"\",\n",
" # optional additional instructions for the report generation\n",
" # template_instructions=None,\n",
" # optional file path to an existing template instead of template_text\n",
" # template_file=None,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The returned `ReportClient` object is used to interact with the report generation process for this specific report."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Report(id=0a394b33-1a3e-463c-b5cb-7ff8ab827d0a, name=my_cool_report_on_attention)\n"
]
}
],
"source": [
"print(report_client)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 3. Get the plan\n",
"\n",
"The first phases of report generation involve ingesting the source data and generating a plan.\n",
"\n",
"The plan is a list of instructions for the report generation, and can be reviewed/accepted/rejected by the user.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"plan = await report_client.await_for_plan(\n",
" timeout=10000,\n",
" poll_interval=10,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# {title}\n",
"[ReportQuery(field='title', prompt='Generate a clear and concise title for this paper about the Transformer model and attention mechanisms', context='The paper discusses the Transformer architecture for sequence transduction using attention mechanisms, focusing on machine translation applications')]\n",
"==================\n",
"## TLDR\n",
"\n",
"{tldr_content}\n",
"[ReportQuery(field='tldr_content', prompt='Write a brief, clear summary of the key points about the Transformer model', context='Focus on the main innovations: attention mechanisms, efficiency improvements, and state-of-the-art results in machine translation')]\n",
"==================\n",
"## Details\n",
"\n",
"{details_content}\n",
"[ReportQuery(field='details_content', prompt='Provide detailed information about the Transformer model architecture and its applications', context='Include information about:\\n- The attention mechanism implementation\\n- Advantages over recurrent and convolutional models\\n- Performance in machine translation tasks\\n- Training efficiency improvements')]\n",
"==================\n"
]
}
],
"source": [
"for plan_block in plan.blocks:\n",
" print(plan_block.block.template)\n",
" print(plan_block.queries)\n",
" print(\"==================\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"With the plan, we can either use it to kick off generation of the final report, or we can edit the plan and adjust it as needed.\n",
"\n",
"While we could manually edit the objects here and use `await report_client.aupdate_plan(action=\"edit\", updated_plan=plan)`, we can also use `LlamaReport` to agentically edit the plan."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"suggestions = await report_client.asuggest_edits(\n",
" \"Can you split the details section into two sections?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Justification for change: \n",
"I'll help you break down the details section into two distinct parts - one focusing on the architecture and another on the practical applications and performance. This will make the content more organized and easier to follow. The original block at index 2 will be replaced with these two new sections.\n",
"\n",
"Proposed changes:\n",
"\n",
"## Architecture Details\n",
"\n",
"{architecture_content}\n",
"\n",
"[ReportQuery(field='architecture_content', prompt='Describe the technical details of the Transformer model architecture', context='Focus on:\\n- Core components of the Transformer architecture\\n- Self-attention mechanism implementation\\n- Multi-head attention details\\n- Position encoding approach\\n- Feed-forward network structure')]\n",
"==================\n",
"\n",
"## Performance and Applications\n",
"\n",
"{applications_content}\n",
"\n",
"[ReportQuery(field='applications_content', prompt='Explain the practical applications and performance advantages of the Transformer model', context='Cover:\\n- Comparison with RNN and CNN models\\n- Machine translation results and benchmarks\\n- Training efficiency improvements\\n- Real-world applications and use cases\\n- Scalability benefits')]\n",
"==================\n"
]
}
],
"source": [
"for suggestion in suggestions:\n",
" print(\"Justification for change:\", suggestion.justification)\n",
" print(\"Proposed changes:\")\n",
" for plan_block in suggestion.blocks:\n",
" print(plan_block.block.template)\n",
" print(plan_block.queries)\n",
" print(\"==================\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This looks pretty good! We can also use the client to automatically accept and apply, or reject, these suggestions.\n",
"\n",
"This will (locally) keep track of the history of changes, so that future suggestions can be based on the previous changes."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"for suggestion in suggestions:\n",
" await report_client.aaccept_edit(suggestion)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"What effect did that have on the tracked local history? Let's see!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[EditAction(block_idx=2, old_content='## Details\\n\\n{details_content}\\n\\nField: details_content, Prompt: Provide detailed information about the Transformer model architecture and its applications, Context: Include information about:\\n- The attention mechanism implementation\\n- Advantages over recurrent and convolutional models\\n- Performance in machine translation tasks\\n- Training efficiency improvements\\nDepends on: none', new_content='\\n## Architecture Details\\n\\n{architecture_content}\\n\\n\\nField: architecture_content, Prompt: Describe the technical details of the Transformer model architecture, Context: Focus on:\\n- Core components of the Transformer architecture\\n- Self-attention mechanism implementation\\n- Multi-head attention details\\n- Position encoding approach\\n- Feed-forward network structure\\nDepends on: none', action='approved', timestamp=datetime.datetime(2025, 2, 4, 20, 59, 55, 773558)),\n",
" EditAction(block_idx=3, old_content='[No old content]', new_content='\\n## Performance and Applications\\n\\n{applications_content}\\n\\n\\nField: applications_content, Prompt: Explain the practical applications and performance advantages of the Transformer model, Context: Cover:\\n- Comparison with RNN and CNN models\\n- Machine translation results and benchmarks\\n- Training efficiency improvements\\n- Real-world applications and use cases\\n- Scalability benefits\\nDepends on: previous', action='approved', timestamp=datetime.datetime(2025, 2, 4, 20, 59, 55, 773687))]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"report_client.edit_history"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[Message(role=<MessageRole.USER: 'user'>, content='Can you split the details section into two sections?', timestamp=datetime.datetime(2025, 2, 4, 20, 59, 47, 754848)),\n",
" Message(role=<MessageRole.ASSISTANT: 'assistant'>, content=\"\\nI'll help you break down the details section into two distinct parts - one focusing on the architecture and another on the practical applications and performance. This will make the content more organized and easier to follow. The original block at index 2 will be replaced with these two new sections.\\n\", timestamp=datetime.datetime(2025, 2, 4, 20, 59, 55, 482070))]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"report_client.chat_history"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"These two items are used to provide context for future suggestions! You can always clear this, or provide your own history."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# report_client.suggest_edits(\"....\", chat_history=[{\"role\": \"user\", \"content\": \"...\"}, ...])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 4. Get the final report\n",
"\n",
"Now that we have a plan, we can kick off generation of the final report."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# kicks off report generation\n",
"await report_client.aupdate_plan(action=\"approve\")\n",
"\n",
"# waits for report generation to complete\n",
"report = await report_client.await_completion(\n",
" timeout=10000,\n",
" poll_interval=10,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Attention Is All You Need: A Pure Attention-Based Architecture for Neural Machine Translation\n",
"\n",
"## TLDR\n",
"\n",
"The Transformer introduced a revolutionary architecture that relies entirely on attention mechanisms, eliminating the need for recurrence or convolution in sequence processing. Its key innovations include multi-head self-attention for parallel processing of input sequences, scaled dot-product attention for efficient computation, and positional encodings for sequence order awareness. The model achieved breakthrough results in machine translation (28.4 BLEU on English-to-German, 41.8 BLEU on English-to-French) while requiring significantly less training time than previous approaches, training in 3.5 days on 8 GPUs. This architecture demonstrated that attention mechanisms alone are sufficient for state-of-the-art sequence modeling, setting a new direction for natural language processing.\n",
"\n",
"\n",
"## Architecture Details\n",
"\n",
"The Transformer architecture represents a groundbreaking approach to sequence processing, built entirely on attention mechanisms without recurrence or convolution. Here are its key technical details:\n",
"\n",
"Core Components:\n",
"- Encoder-decoder architecture with stacked self-attention and point-wise feed-forward layers\n",
"- Each layer contains two main sub-layers: multi-head self-attention mechanism and position-wise feed-forward network\n",
"- Layer normalization and residual connections between sub-layers\n",
"- No recurrent or convolutional elements, enabling parallel processing\n",
"\n",
"Self-Attention Mechanism:\n",
"- Processes relationships between all positions in a sequence simultaneously\n",
"- Computes attention weights using queries, keys, and values derived from input representations\n",
"- Implements scaled dot-product attention to prevent gradient issues with large input dimensions\n",
"- Allows direct modeling of dependencies regardless of positional distance\n",
"- Uses masking in decoder to prevent leftward information flow and maintain auto-regressive property\n",
"\n",
"Multi-Head Attention:\n",
"- Employs multiple attention heads operating in parallel\n",
"- Each head processes information in different representation subspaces\n",
"- Three types of attention applications:\n",
" 1. Encoder self-attention (all positions attend to each other)\n",
" 2. Decoder self-attention (each position attends to previous positions)\n",
" 3. Encoder-decoder attention (decoder queries attend to encoder outputs)\n",
"- Counteracts reduced resolution from attention averaging through parallel processing\n",
"\n",
"Position-wise Feed-Forward Network:\n",
"- Applied identically to each position separately\n",
"- Consists of two linear transformations with ReLU activation\n",
"- Structure: FFN(x) = max(0, xW1 + b1)W2 + b2\n",
"- Input and output dimensionality: dmodel = 512\n",
"- Inner-layer dimensionality: dff = 2048\n",
"- Parameters vary between layers but remain constant across positions\n",
"\n",
"Position Encoding:\n",
"- Adds positional information to input embeddings\n",
"- Enables the model to consider sequential order without recurrence\n",
"- Implements sinusoidal position encodings to allow model to attend to relative positions\n",
"- Maintains constant number of operations between any two positions, unlike convolutional approaches\n",
"- Allows effective modeling of both local and long-range dependencies\n",
"\n",
"\n",
"\n",
"## Performance and Applications\n",
"\n",
"The Transformer model demonstrates significant performance advantages and practical applications across multiple domains:\n",
"\n",
"Performance Advantages over RNN/CNN Models:\n",
"- Eliminates sequential computation constraints present in RNNs, enabling superior parallelization\n",
"- Reduces operations needed for relating distant positions to a constant number, compared to linear/logarithmic scaling in CNNs\n",
"- Processes all input and output positions simultaneously through self-attention mechanisms\n",
"- Achieves state-of-the-art results while requiring significantly less computational resources\n",
"\n",
"Machine Translation Benchmarks:\n",
"- WMT 2014 English-to-German: 28.4 BLEU score, exceeding previous best results by over 2 BLEU points\n",
"- WMT 2014 English-to-French: 41.8 BLEU score (single-model state-of-the-art)\n",
"- Surpasses performance of existing model ensembles in translation tasks\n",
"\n",
"Training Efficiency:\n",
"- Requires only 3.5 days of training on eight GPUs for state-of-the-art performance\n",
"- Achieves superior results at \"a small fraction of the training costs\" compared to previous models\n",
"- Enables significantly faster training through parallel processing of input/output sequences\n",
"- Can reach production-quality performance in as little as twelve hours on modern GPU hardware\n",
"\n",
"Real-world Applications:\n",
"- Machine translation systems\n",
"- Natural language understanding tasks\n",
"- Reading comprehension\n",
"- Abstractive summarization\n",
"- Text entailment analysis\n",
"- Constituency parsing (achieving 92.7 F1 score in semi-supervised settings)\n",
"- Adaptable to both large and limited training data scenarios\n",
"\n",
"Scalability Benefits:\n",
"- Highly parallelizable architecture enables efficient scaling across multiple GPUs\n",
"- Constant computational complexity for relating any input/output positions\n",
"- Effective handling of long-range dependencies in sequences\n",
"- Maintains performance quality while scaling to larger datasets and model sizes\n",
"- Generalizes well across different tasks and domains without architectural changes\n",
"- Supports efficient inference and deployment in production environments\n",
"\n"
]
}
],
"source": [
"report_text = \"\\n\\n\".join([block.template for block in report.blocks])\n",
"print(report_text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 5. Edit the final report\n",
"\n",
"Now that we have a report, we can edit it.\n",
"\n",
"We can use the `asuggest_edits` method to get suggestions for edits, and then use the `aaccept_edit`/`areject_edit` methods to apply them.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Justification for change: \n",
"I'd suggest changing \"TLDR\" to \"Executive Summary\" which is more appropriate for a professional or academic report. This term is widely used in formal documents and better reflects the nature of this concise overview section while maintaining the same function of providing a quick summary of the key points.\n",
"\n",
"Proposed changes:\n",
"## Executive Summary\n",
"\n",
"The Transformer introduced a revolutionary architecture that relies entirely on attention mechanisms, eliminating the need for recurrence or convolution in sequence processing. Its key innovations include multi-head self-attention for parallel processing of input sequences, scaled dot-product attention for efficient computation, and positional encodings for sequence order awareness. The model achieved breakthrough results in machine translation (28.4 BLEU on English-to-German, 41.8 BLEU on English-to-French) while requiring significantly less training time than previous approaches, training in 3.5 days on 8 GPUs. This architecture demonstrated that attention mechanisms alone are sufficient for state-of-the-art sequence modeling, setting a new direction for natural language processing.\n",
"==================\n"
]
}
],
"source": [
"suggestions = await report_client.asuggest_edits(\n",
" \"Can you change the TLDR header to something more professional?\"\n",
")\n",
"for suggestion in suggestions:\n",
" print(\"Justification for change:\", suggestion.justification)\n",
" print(\"Proposed changes:\")\n",
" for block in suggestion.blocks:\n",
" print(block.template)\n",
" print(\"==================\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Changing to \"Executive Summary\" sounds reasonable, lets accept that!\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"for suggestion in suggestions:\n",
" await report_client.aaccept_edit(suggestion)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 7. Print the final report\n",
"\n",
"Now that we have a report, we can print it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Attention Is All You Need: A Pure Attention-Based Architecture for Neural Machine Translation\n",
"\n",
"## Executive Summary\n",
"\n",
"The Transformer introduced a revolutionary architecture that relies entirely on attention mechanisms, eliminating the need for recurrence or convolution in sequence processing. Its key innovations include multi-head self-attention for parallel processing of input sequences, scaled dot-product attention for efficient computation, and positional encodings for sequence order awareness. The model achieved breakthrough results in machine translation (28.4 BLEU on English-to-German, 41.8 BLEU on English-to-French) while requiring significantly less training time than previous approaches, training in 3.5 days on 8 GPUs. This architecture demonstrated that attention mechanisms alone are sufficient for state-of-the-art sequence modeling, setting a new direction for natural language processing.\n",
"\n",
"\n",
"## Architecture Details\n",
"\n",
"The Transformer architecture represents a groundbreaking approach to sequence processing, built entirely on attention mechanisms without recurrence or convolution. Here are its key technical details:\n",
"\n",
"Core Components:\n",
"- Encoder-decoder architecture with stacked self-attention and point-wise feed-forward layers\n",
"- Each layer contains two main sub-layers: multi-head self-attention mechanism and position-wise feed-forward network\n",
"- Layer normalization and residual connections between sub-layers\n",
"- No recurrent or convolutional elements, enabling parallel processing\n",
"\n",
"Self-Attention Mechanism:\n",
"- Processes relationships between all positions in a sequence simultaneously\n",
"- Computes attention weights using queries, keys, and values derived from input representations\n",
"- Implements scaled dot-product attention to prevent gradient issues with large input dimensions\n",
"- Allows direct modeling of dependencies regardless of positional distance\n",
"- Uses masking in decoder to prevent leftward information flow and maintain auto-regressive property\n",
"\n",
"Multi-Head Attention:\n",
"- Employs multiple attention heads operating in parallel\n",
"- Each head processes information in different representation subspaces\n",
"- Three types of attention applications:\n",
" 1. Encoder self-attention (all positions attend to each other)\n",
" 2. Decoder self-attention (each position attends to previous positions)\n",
" 3. Encoder-decoder attention (decoder queries attend to encoder outputs)\n",
"- Counteracts reduced resolution from attention averaging through parallel processing\n",
"\n",
"Position-wise Feed-Forward Network:\n",
"- Applied identically to each position separately\n",
"- Consists of two linear transformations with ReLU activation\n",
"- Structure: FFN(x) = max(0, xW1 + b1)W2 + b2\n",
"- Input and output dimensionality: dmodel = 512\n",
"- Inner-layer dimensionality: dff = 2048\n",
"- Parameters vary between layers but remain constant across positions\n",
"\n",
"Position Encoding:\n",
"- Adds positional information to input embeddings\n",
"- Enables the model to consider sequential order without recurrence\n",
"- Implements sinusoidal position encodings to allow model to attend to relative positions\n",
"- Maintains constant number of operations between any two positions, unlike convolutional approaches\n",
"- Allows effective modeling of both local and long-range dependencies\n",
"\n",
"\n",
"\n",
"## Performance and Applications\n",
"\n",
"The Transformer model demonstrates significant performance advantages and practical applications across multiple domains:\n",
"\n",
"Performance Advantages over RNN/CNN Models:\n",
"- Eliminates sequential computation constraints present in RNNs, enabling superior parallelization\n",
"- Reduces operations needed for relating distant positions to a constant number, compared to linear/logarithmic scaling in CNNs\n",
"- Processes all input and output positions simultaneously through self-attention mechanisms\n",
"- Achieves state-of-the-art results while requiring significantly less computational resources\n",
"\n",
"Machine Translation Benchmarks:\n",
"- WMT 2014 English-to-German: 28.4 BLEU score, exceeding previous best results by over 2 BLEU points\n",
"- WMT 2014 English-to-French: 41.8 BLEU score (single-model state-of-the-art)\n",
"- Surpasses performance of existing model ensembles in translation tasks\n",
"\n",
"Training Efficiency:\n",
"- Requires only 3.5 days of training on eight GPUs for state-of-the-art performance\n",
"- Achieves superior results at \"a small fraction of the training costs\" compared to previous models\n",
"- Enables significantly faster training through parallel processing of input/output sequences\n",
"- Can reach production-quality performance in as little as twelve hours on modern GPU hardware\n",
"\n",
"Real-world Applications:\n",
"- Machine translation systems\n",
"- Natural language understanding tasks\n",
"- Reading comprehension\n",
"- Abstractive summarization\n",
"- Text entailment analysis\n",
"- Constituency parsing (achieving 92.7 F1 score in semi-supervised settings)\n",
"- Adaptable to both large and limited training data scenarios\n",
"\n",
"Scalability Benefits:\n",
"- Highly parallelizable architecture enables efficient scaling across multiple GPUs\n",
"- Constant computational complexity for relating any input/output positions\n",
"- Effective handling of long-range dependencies in sequences\n",
"- Maintains performance quality while scaling to larger datasets and model sizes\n",
"- Generalizes well across different tasks and domains without architectural changes\n",
"- Supports efficient inference and deployment in production environments\n",
"\n"
]
}
],
"source": [
"report_response = await report_client.aget()\n",
"report_text = \"\\n\\n\".join([block.template for block in report_response.report.blocks])\n",
"print(report_text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can also see the sources for each block!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"0.99687636\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutiona\n",
"==================\n",
"0.99591404\n",
"# 2 Background\n",
"\n",
"The goal of reducing sequential computation also forms the foundation of the Extende\n",
"==================\n",
"0.9951325\n",
"# 1 Introduction\n",
"\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neu\n",
"==================\n",
"0.99442345\n",
"# 7 Conclusion\n",
"\n",
"In this work, we presented the Transformer, the first sequence transduction model ba\n",
"==================\n",
"0.9967649\n",
"# 3.2.3 Applications of Attention in our Model\n",
"\n",
"The Transformer uses multi-head attention in three d\n",
"==================\n",
"0.99533635\n",
"# 2 Background\n",
"\n",
"The goal of reducing sequential computation also forms the foundation of the Extende\n",
"==================\n",
"0.9935868\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutiona\n",
"==================\n",
"0.98780584\n",
"# Outputs\n",
"\n",
"(shifted right)\n",
"\n",
"Figure 1: The Transformer - model architecture.\n",
"\n",
"The Transformer follows\n",
"==================\n",
"0.9205043\n",
"# 3.3 Position-wise Feed-Forward Networks\n",
"\n",
"In addition to attention sub-layers, each of the layers i\n",
"==================\n",
"0.79581684\n",
"# 1 Introduction\n",
"\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neu\n",
"==================\n",
"0.9946774\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutiona\n",
"==================\n",
"0.97079873\n",
"# 7 Conclusion\n",
"\n",
"In this work, we presented the Transformer, the first sequence transduction model ba\n",
"==================\n",
"0.9535353\n",
"# 6.3 English Constituency Parsing\n",
"\n",
"To evaluate if the Transformer can generalize to other tasks we \n",
"==================\n",
"0.9514138\n",
"# 2 Background\n",
"\n",
"The goal of reducing sequential computation also forms the foundation of the Extende\n",
"==================\n",
"0.9790758\n",
"# 1 Introduction\n",
"\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neu\n",
"==================\n",
"0.92262185\n",
"# Outputs\n",
"\n",
"(shifted right)\n",
"\n",
"Figure 1: The Transformer - model architecture.\n",
"\n",
"The Transformer follows\n",
"==================\n"
]
}
],
"source": [
"for block in report_response.report.blocks:\n",
" # Each block has a list of sources, which are the nodes that were used to generate the block\n",
" for source in block.sources:\n",
" print(source.score)\n",
" print(source.node.text[:100])\n",
" print(\"==================\")"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama-parse-aNC435Vv-py3.10",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
@@ -0,0 +1,73 @@
This project uses LlamaSheets to extract data from spreadsheets for analysis.
## Current Project Structure
- `data/` - Contains extracted parquet files from LlamaSheets
- `{name}_region_{N}.parquet` - Table data files
- `{name}_metadata_{N}.parquet` - Cell metadata files
- `{name}_job_metadata.json` - Extraction job information
- `scripts/` - Analysis and helper scripts
- `reports/` - Your generated reports and outputs
## Working with LlamaSheets Data
### Understanding the Files
When a spreadsheet is extracted, you'll find:
1. **Table parquet files** (`region_*.parquet`): The actual table data
- Columns correspond to spreadsheet columns
- Data types are preserved (dates, numbers, strings, booleans)
2. **Metadata parquet files** (`metadata_*.parquet`): Rich cell-level metadata
- Formatting: `font_bold`, `font_italic`, `font_size`, `background_color_rgb`
- Position: `row_number`, `column_number`, `coordinate` (e.g., "A1")
- Type detection: `data_type`, `is_date_like`, `is_percentage`, `is_currency`
- Layout: `is_in_first_row`, `is_merged_cell`, `horizontal_alignment`
- Content: `cell_value`, `raw_cell_value`
3. **Job metadata JSON** (`job_metadata.json`): Overall extraction results
- `regions[]`: List of extracted regions with IDs, locations, and titles/descriptions
- `worksheet_metadata[]`: Generated titles and descriptions
- `status`: Success/failure status
### Key Principles
1. **Use metadata to understand structure**: Bold cells often indicate headers, colors indicate groupings
2. **Validate before analysis**: Check data types, look for missing values
3. **Preserve formatting context**: The metadata tells you what the spreadsheet author emphasized
4. **Save intermediate results**: Store cleaned data as new parquet files
### Common Patterns
**Loading data:**
```python
import pandas as pd
df = pd.read_parquet("data/region_1_Sheet1.parquet")
meta_df = pd.read_parquet("data/metadata_1_Sheet1.parquet")
```
**Finding headers:**
```python
headers = meta_df[meta_df["font_bold"] == True]["cell_value"].tolist()
```
**Finding date columns:**
```python
date_cols = meta_df[meta_df["is_date_like"] == True]["column_number"].unique()
```
## Tools Available
- **Python 3.11+**: For data analysis
- **pandas**: DataFrame manipulation
- **pyarrow**: Parquet file reading
- **matplotlib**: Visualization (optional)
## Guidelines
- Always read the job_metadata.json first to understand what was extracted
- Check both table data and metadata before making assumptions
- Write reusable functions for common operations
- Document any data quality issues discovered
@@ -0,0 +1,278 @@
"""
Generate sample spreadsheets for LlamaSheets + Claude workflows.
This script creates example Excel files that demonstrate different use cases:
1. Simple data table (for Workflow 1)
2. Regional sales data (for Workflow 2)
3. Complex budget with formatting (for Workflow 3)
4. Weekly sales report (for Workflow 4)
Usage:
python generate_sample_data.py
"""
import random
from datetime import datetime, timedelta
from pathlib import Path
import pandas as pd
from openpyxl import Workbook
from openpyxl.styles import Font, PatternFill, Alignment
def generate_workflow_1_data(output_dir: Path) -> None:
"""Generate simple financial report for Workflow 1."""
print("📊 Generating Workflow 1: financial_report_q1.xlsx")
# Create sample quarterly data
months = ["January", "February", "March"]
categories = ["Revenue", "Cost of Goods Sold", "Operating Expenses", "Net Income"]
data = []
for category in categories:
row: dict[str, str | int] = {"Category": category}
for month in months:
if category == "Revenue":
value = random.randint(80000, 120000)
elif category == "Cost of Goods Sold":
value = random.randint(30000, 50000)
elif category == "Operating Expenses":
value = random.randint(20000, 35000)
else: # Net Income
value = int(
int(row.get("January", 0))
+ int(row.get("February", 0))
+ int(row.get("March", 0))
)
value = random.randint(15000, 40000)
row[month] = value
data.append(row)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / "financial_report_q1.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Q1 Summary", index=False)
# Format it nicely
worksheet = writer.sheets["Q1 Summary"]
for cell in worksheet[1]: # Header row
cell.font = Font(bold=True)
cell.fill = PatternFill(
start_color="4F81BD", end_color="4F81BD", fill_type="solid"
)
cell.font = Font(color="FFFFFF", bold=True)
print(f" ✅ Created {output_file}")
def generate_workflow_2_data(output_dir: Path) -> None:
"""Generate regional sales data for Workflow 2."""
print("\n📊 Generating Workflow 2: Regional sales data")
regions = ["northeast", "southeast", "west"]
products = ["Widget A", "Widget B", "Widget C", "Gadget X", "Gadget Y"]
for region in regions:
data = []
start_date = datetime(2024, 1, 1)
# Generate 90 days of sales data
for day in range(90):
date = start_date + timedelta(days=day)
# Random number of sales per day (3-8)
for _ in range(random.randint(3, 8)):
product = random.choice(products)
units_sold = random.randint(1, 20)
price_per_unit = random.randint(50, 200)
revenue = units_sold * price_per_unit
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units_Sold": units_sold,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / f"sales_{region}.xlsx"
df.to_excel(output_file, sheet_name="Sales", index=False)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def generate_workflow_3_data(output_dir: Path) -> None:
"""Generate complex budget spreadsheet with formatting for Workflow 3."""
print("\n📊 Generating Workflow 3: company_budget_2024.xlsx")
wb = Workbook()
ws = wb.active
ws.title = "Budget"
# Define departments with colors
departments = {
"Engineering": "C6E0B4",
"Marketing": "FFD966",
"Sales": "F4B084",
"Operations": "B4C7E7",
}
# Define categories
categories = {
"Personnel": ["Salaries", "Benefits", "Training"],
"Infrastructure": ["Office Rent", "Equipment", "Software Licenses"],
"Operations": ["Travel", "Supplies", "Miscellaneous"],
}
# Styles
header_font = Font(bold=True, size=12)
category_font = Font(bold=True, size=11)
row = 1
# Title
ws.merge_cells(f"A{row}:E{row}")
ws[f"A{row}"] = "2024 Annual Budget"
ws[f"A{row}"].font = Font(bold=True, size=14)
ws[f"A{row}"].alignment = Alignment(horizontal="center")
row += 2
# Headers
ws[f"A{row}"] = "Category"
ws[f"B{row}"] = "Item"
for i, dept in enumerate(departments.keys()):
ws.cell(row, 3 + i, dept)
ws.cell(row, 3 + i).font = header_font
for cell in ws[row]:
cell.font = header_font
row += 1
# Data
for category, items in categories.items():
# Category header (bold)
ws[f"A{row}"] = category
ws[f"A{row}"].font = category_font
row += 1
# Items with department budgets
for item in items:
ws[f"A{row}"] = ""
ws[f"B{row}"] = item
# Add budget amounts for each department (with color)
for i, (dept, color) in enumerate(departments.items()):
amount = random.randint(5000, 50000)
cell = ws.cell(row, 3 + i, amount)
cell.fill = PatternFill(
start_color=color, end_color=color, fill_type="solid"
)
cell.number_format = "$#,##0"
row += 1
row += 1 # Blank row between categories
# Adjust column widths
ws.column_dimensions["A"].width = 20
ws.column_dimensions["B"].width = 25
for i in range(len(departments)):
ws.column_dimensions[chr(67 + i)].width = 15 # C, D, E, F
output_file = output_dir / "company_budget_2024.xlsx"
wb.save(output_file)
print(f" ✅ Created {output_file}")
print(" • Bold categories, colored departments, merged title cell")
def generate_workflow_4_data(output_dir: Path) -> None:
"""Generate weekly sales report for Workflow 4."""
print("\n📊 Generating Workflow 4: sales_weekly.xlsx")
products = [
"Product A",
"Product B",
"Product C",
"Product D",
"Product E",
"Product F",
"Product G",
"Product H",
]
# Generate one week of data
data = []
start_date = datetime(2024, 11, 4) # Monday
for day in range(7):
date = start_date + timedelta(days=day)
# Each product has 3-10 transactions per day
for product in products:
for _ in range(random.randint(3, 10)):
units = random.randint(1, 15)
price = random.randint(20, 150)
revenue = units * price
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units": units,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel with some formatting
output_file = output_dir / "sales_weekly.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Weekly Sales", index=False)
# Format header
worksheet = writer.sheets["Weekly Sales"]
for cell in worksheet[1]:
cell.font = Font(bold=True)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def main() -> None:
"""Generate all sample data files."""
print("=" * 60)
print("Generating Sample Data for LlamaSheets + Coding Agent Workflows")
print("=" * 60)
# Create output directory
output_dir = Path("input_data")
output_dir.mkdir(exist_ok=True)
# Generate data for each workflow
generate_workflow_1_data(output_dir)
generate_workflow_2_data(output_dir)
generate_workflow_3_data(output_dir)
generate_workflow_4_data(output_dir)
print("\n" + "=" * 60)
print("✅ All sample data generated!")
print("=" * 60)
print(f"\nFiles created in {output_dir.absolute()}:")
print("\nWorkflow 1 (Understanding a New Spreadsheet):")
print(" • financial_report_q1.xlsx")
print("\nWorkflow 2 (Generating Analysis Scripts):")
print(" • sales_northeast.xlsx")
print(" • sales_southeast.xlsx")
print(" • sales_west.xlsx")
print("\nWorkflow 3 (Using Cell Metadata):")
print(" • company_budget_2024.xlsx")
print("\nWorkflow 4 (Complete Automation):")
print(" • sales_weekly.xlsx")
print("\nYou can now use these files with the workflows in the documentation!")
if __name__ == "__main__":
main()
@@ -0,0 +1,5 @@
llama-cloud-services # LlamaSheets SDK
pandas>=2.0.0
pyarrow>=12.0.0
openpyxl>=3.0.0 # For Excel file support
matplotlib>=3.7.0 # For visualizations (optional)
@@ -0,0 +1,100 @@
"""Helper script to extract spreadsheets using LlamaSheets."""
import asyncio
import json
import os
import dotenv
from pathlib import Path
from llama_cloud_services.beta.sheets import LlamaSheets
from llama_cloud_services.beta.sheets.types import (
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
dotenv.load_dotenv()
async def extract_spreadsheet(
file_path: str, output_dir: str = "data", generate_metadata: bool = True
) -> dict:
"""Extract a spreadsheet using LlamaSheets."""
client = LlamaSheets(
base_url="https://api.cloud.llamaindex.ai",
api_key=os.getenv("LLAMA_CLOUD_API_KEY"),
)
print(f"Extracting {file_path}...")
# Extract regions
config = SpreadsheetParsingConfig(
sheet_names=None, # Extract all sheets
generate_additional_metadata=generate_metadata,
)
job_result = await client.aextract_regions(file_path, config=config)
print(f"Extracted {len(job_result.regions)} region(s)")
# Create output directory
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
# Get base name for files
base_name = Path(file_path).stem
# Save job metadata
job_metadata_path = output_path / f"{base_name}_job_metadata.json"
with open(job_metadata_path, "w") as f:
json.dump(job_result.model_dump(mode="json"), f, indent=2)
print(f"Saved job metadata to {job_metadata_path}")
# Download each region
for idx, region in enumerate(job_result.regions, 1):
sheet_name = region.sheet_name.replace(" ", "_")
# Download region data
region_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=region.region_type,
)
region_path = output_path / f"{base_name}_region_{idx}_{sheet_name}.parquet"
with open(region_path, "wb") as f:
f.write(region_bytes)
print(f" Table {idx}: {region_path}")
# Download metadata
metadata_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=SpreadsheetResultType.CELL_METADATA,
)
metadata_path = output_path / f"{base_name}_metadata_{idx}_{sheet_name}.parquet"
with open(metadata_path, "wb") as f:
f.write(metadata_bytes)
print(f" Metadata {idx}: {metadata_path}")
print(f"\nAll files saved to {output_path}/")
return job_result.model_dump(mode="json")
if __name__ == "__main__":
import sys
if len(sys.argv) < 2:
print("Usage: python scripts/extract.py <spreadsheet_file>")
sys.exit(1)
file_path = sys.argv[1]
if not Path(file_path).exists():
print(f"❌ File not found: {file_path}")
sys.exit(1)
result = asyncio.run(extract_spreadsheet(file_path))
print(f"\n✅ Extraction complete! Job ID: {result['id']}")
@@ -0,0 +1,278 @@
"""
Generate sample spreadsheets for LlamaSheets + LlamaIndex Agent workflows.
This script creates example Excel files that demonstrate different use cases:
1. Simple data table (for Workflow 1)
2. Regional sales data (for Workflow 2)
3. Complex budget with formatting (for Workflow 3)
4. Weekly sales report (for Workflow 4)
Usage:
python generate_sample_data.py
"""
import random
from datetime import datetime, timedelta
from pathlib import Path
import pandas as pd
from openpyxl import Workbook
from openpyxl.styles import Font, PatternFill, Alignment
def generate_workflow_1_data(output_dir: Path) -> None:
"""Generate simple financial report for Workflow 1."""
print("📊 Generating Workflow 1: financial_report_q1.xlsx")
# Create sample quarterly data
months = ["January", "February", "March"]
categories = ["Revenue", "Cost of Goods Sold", "Operating Expenses", "Net Income"]
data = []
for category in categories:
row: dict[str, str | int] = {"Category": category}
for month in months:
if category == "Revenue":
value = random.randint(80000, 120000)
elif category == "Cost of Goods Sold":
value = random.randint(30000, 50000)
elif category == "Operating Expenses":
value = random.randint(20000, 35000)
else: # Net Income
value = int(
int(row.get("January", 0))
+ int(row.get("February", 0))
+ int(row.get("March", 0))
)
value = random.randint(15000, 40000)
row[month] = value
data.append(row)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / "financial_report_q1.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Q1 Summary", index=False)
# Format it nicely
worksheet = writer.sheets["Q1 Summary"]
for cell in worksheet[1]: # Header row
cell.font = Font(bold=True)
cell.fill = PatternFill(
start_color="4F81BD", end_color="4F81BD", fill_type="solid"
)
cell.font = Font(color="FFFFFF", bold=True)
print(f" ✅ Created {output_file}")
def generate_workflow_2_data(output_dir: Path) -> None:
"""Generate regional sales data for Workflow 2."""
print("\n📊 Generating Workflow 2: Regional sales data")
regions = ["northeast", "southeast", "west"]
products = ["Widget A", "Widget B", "Widget C", "Gadget X", "Gadget Y"]
for region in regions:
data = []
start_date = datetime(2024, 1, 1)
# Generate 90 days of sales data
for day in range(90):
date = start_date + timedelta(days=day)
# Random number of sales per day (3-8)
for _ in range(random.randint(3, 8)):
product = random.choice(products)
units_sold = random.randint(1, 20)
price_per_unit = random.randint(50, 200)
revenue = units_sold * price_per_unit
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units_Sold": units_sold,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / f"sales_{region}.xlsx"
df.to_excel(output_file, sheet_name="Sales", index=False)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def generate_workflow_3_data(output_dir: Path) -> None:
"""Generate complex budget spreadsheet with formatting for Workflow 3."""
print("\n📊 Generating Workflow 3: company_budget_2024.xlsx")
wb = Workbook()
ws = wb.active
ws.title = "Budget"
# Define departments with colors
departments = {
"Engineering": "C6E0B4",
"Marketing": "FFD966",
"Sales": "F4B084",
"Operations": "B4C7E7",
}
# Define categories
categories = {
"Personnel": ["Salaries", "Benefits", "Training"],
"Infrastructure": ["Office Rent", "Equipment", "Software Licenses"],
"Operations": ["Travel", "Supplies", "Miscellaneous"],
}
# Styles
header_font = Font(bold=True, size=12)
category_font = Font(bold=True, size=11)
row = 1
# Title
ws.merge_cells(f"A{row}:E{row}")
ws[f"A{row}"] = "2024 Annual Budget"
ws[f"A{row}"].font = Font(bold=True, size=14)
ws[f"A{row}"].alignment = Alignment(horizontal="center")
row += 2
# Headers
ws[f"A{row}"] = "Category"
ws[f"B{row}"] = "Item"
for i, dept in enumerate(departments.keys()):
ws.cell(row, 3 + i, dept)
ws.cell(row, 3 + i).font = header_font
for cell in ws[row]:
cell.font = header_font
row += 1
# Data
for category, items in categories.items():
# Category header (bold)
ws[f"A{row}"] = category
ws[f"A{row}"].font = category_font
row += 1
# Items with department budgets
for item in items:
ws[f"A{row}"] = ""
ws[f"B{row}"] = item
# Add budget amounts for each department (with color)
for i, (dept, color) in enumerate(departments.items()):
amount = random.randint(5000, 50000)
cell = ws.cell(row, 3 + i, amount)
cell.fill = PatternFill(
start_color=color, end_color=color, fill_type="solid"
)
cell.number_format = "$#,##0"
row += 1
row += 1 # Blank row between categories
# Adjust column widths
ws.column_dimensions["A"].width = 20
ws.column_dimensions["B"].width = 25
for i in range(len(departments)):
ws.column_dimensions[chr(67 + i)].width = 15 # C, D, E, F
output_file = output_dir / "company_budget_2024.xlsx"
wb.save(output_file)
print(f" ✅ Created {output_file}")
print(" • Bold categories, colored departments, merged title cell")
def generate_workflow_4_data(output_dir: Path) -> None:
"""Generate weekly sales report for Workflow 4."""
print("\n📊 Generating Workflow 4: sales_weekly.xlsx")
products = [
"Product A",
"Product B",
"Product C",
"Product D",
"Product E",
"Product F",
"Product G",
"Product H",
]
# Generate one week of data
data = []
start_date = datetime(2024, 11, 4) # Monday
for day in range(7):
date = start_date + timedelta(days=day)
# Each product has 3-10 transactions per day
for product in products:
for _ in range(random.randint(3, 10)):
units = random.randint(1, 15)
price = random.randint(20, 150)
revenue = units * price
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units": units,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel with some formatting
output_file = output_dir / "sales_weekly.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Weekly Sales", index=False)
# Format header
worksheet = writer.sheets["Weekly Sales"]
for cell in worksheet[1]:
cell.font = Font(bold=True)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def main() -> None:
"""Generate all sample data files."""
print("=" * 60)
print("Generating Sample Data for LlamaSheets + Coding Agent Workflows")
print("=" * 60)
# Create output directory
output_dir = Path("input_data")
output_dir.mkdir(exist_ok=True)
# Generate data for each workflow
generate_workflow_1_data(output_dir)
generate_workflow_2_data(output_dir)
generate_workflow_3_data(output_dir)
generate_workflow_4_data(output_dir)
print("\n" + "=" * 60)
print("✅ All sample data generated!")
print("=" * 60)
print(f"\nFiles created in {output_dir.absolute()}:")
print("\nWorkflow 1 (Understanding a New Spreadsheet):")
print(" • financial_report_q1.xlsx")
print("\nWorkflow 2 (Generating Analysis Scripts):")
print(" • sales_northeast.xlsx")
print(" • sales_southeast.xlsx")
print(" • sales_west.xlsx")
print("\nWorkflow 3 (Using Cell Metadata):")
print(" • company_budget_2024.xlsx")
print("\nWorkflow 4 (Complete Automation):")
print(" • sales_weekly.xlsx")
print("\nYou can now use these files with the workflows in the documentation!")
if __name__ == "__main__":
main()
@@ -0,0 +1,292 @@
"""
LlamaSheets Agent with LlamaIndex
This example shows how to build an agent that can work with spreadsheet data
extracted by LlamaSheets using Python code execution.
The agent has minimal tools but maximum flexibility - it can execute arbitrary
pandas code against the extracted data, similar to a coding agent.
NOTE: Code execution should be handled safely in a sandboxed environment for security.
"""
import io
import json
import sys
from pathlib import Path
from typing import Any, Dict, Optional
import dotenv
import pandas as pd
from llama_index.core.agent import FunctionAgent, ToolCall, ToolCallResult, AgentStream
from llama_index.llms.openai import OpenAI
from workflows import Context
dotenv.load_dotenv()
# Global context for executed code
_code_context: Dict[str, Any] = {}
# Helper function for initial agent context
def list_extracted_data(data_dir: str = "data") -> str:
"""
List all regions and metadata files extracted by LlamaSheets.
This helps discover what data is available to work with.
Args:
data_dir: Directory containing extracted parquet files (default: "data")
Returns:
JSON string with information about available files
"""
data_path = Path(data_dir)
if not data_path.exists():
return json.dumps({"error": f"Data directory '{data_dir}' not found"})
# Find all parquet and metadata files
region_files = list(data_path.glob("*_region_*.parquet"))
job_metadata_files = list(data_path.glob("*_job_metadata.json"))
regions = []
for region_file in region_files:
# Quick peek at dimensions
df = pd.read_parquet(region_file)
# Find corresponding metadata file
base_name = region_file.stem.replace("_region_", "_metadata_")
metadata_path = region_file.parent / f"{base_name}.parquet"
regions.append(
{
"region_file": str(region_file),
"metadata_file": str(metadata_path) if metadata_path.exists() else None,
"shape": {"rows": len(df), "columns": len(df.columns)},
"columns": list(df.columns),
}
)
result = {
"data_directory": str(data_path.absolute()),
"num_regions": len(regions),
"regions": regions,
"job_metadata_files": [str(f) for f in job_metadata_files],
}
return json.dumps(result, indent=2)
# Agent tool for code execution against dataframes
def execute_code(code: str) -> str:
"""
Execute Python pandas code against LlamaSheets extracted data.
This tool allows flexible data analysis by executing arbitrary pandas code.
You can load parquet files, manipulate dataframes, and return results.
The code executes in a context where:
- pandas is available as 'pd'
- json is available for formatting output
Args:
code: Python code to execute. Any print() statements or stdout/stderr
will be captured and returned. Optionally set a 'result' variable
for structured output.
Returns:
String containing:
- Any stdout/stderr output from the code execution
- The 'result' variable if it was set (formatted appropriately)
- Error message if execution failed
Example usage:
code = '''
# Load and inspect data
df = pd.read_parquet("data/sales_region_1.parquet")
print(f"Loaded {len(df)} rows")
result = {
"shape": df.shape,
"columns": list(df.columns),
"sample": df.head(3).to_dict(orient="records")
}
'''
"""
global _code_context
# Capture stdout and stderr
stdout_capture = io.StringIO()
stderr_capture = io.StringIO()
old_stdout = sys.stdout
old_stderr = sys.stderr
try:
# Redirect stdout/stderr
sys.stdout = stdout_capture
sys.stderr = stderr_capture
# Create execution context with pandas, json, and previously loaded dfs
exec_context = {
"pd": pd,
"json": json,
"Path": Path,
**_code_context, # Include previously loaded dataframes
}
# Execute the code
exec(code, exec_context)
# Update global context with any new variables (excluding built-ins and modules)
for key, value in exec_context.items():
if not key.startswith("_") and key not in ["pd", "json", "Path"]:
_code_context[key] = value
# Restore stdout/stderr
sys.stdout = old_stdout
sys.stderr = old_stderr
# Collect output
stdout_output = stdout_capture.getvalue()
stderr_output = stderr_capture.getvalue()
output_parts = []
# Add stdout if any
if stdout_output:
output_parts.append(f"<stdout>{stdout_output}</stdout>")
# Add stderr if any
if stderr_output:
output_parts.append(f"<stderr>{stderr_output}</stderr>")
# Try to get a result (if code set a 'result' variable)
if "result" in exec_context:
result = exec_context["result"]
result_str = None
if isinstance(result, pd.DataFrame):
# Convert DataFrame to readable format
result_str = result.to_string()
elif isinstance(result, (dict, list)):
result_str = json.dumps(result, indent=2, default=str)
else:
result_str = str(result)
if result_str:
output_parts.append(f"<result_var>{result_str}</result_var>")
# Return combined output or success message
if output_parts:
return "\n\n".join(output_parts)
else:
return "Code executed successfully (no output or result)"
except Exception as e:
# Restore stdout/stderr in case of error
sys.stdout = old_stdout
sys.stderr = old_stderr
# Get any partial output
stdout_output = stdout_capture.getvalue()
stderr_output = stderr_capture.getvalue()
error_parts = []
if stdout_output:
error_parts.append(f"=== STDOUT (before error) ===\n{stdout_output}")
if stderr_output:
error_parts.append(f"=== STDERR (before error) ===\n{stderr_output}")
error_parts.append(f"=== ERROR ===\n{str(e)}")
error_parts.append(f"\n=== CODE ===\n{code}")
return "\n\n".join(error_parts)
def create_llamasheets_agent(
llm_model: str = "gpt-4.1", api_key: Optional[str] = None
) -> FunctionAgent:
# Initialize LLM
llm = OpenAI(model=llm_model, api_key=api_key)
# Create tools list
tools = [execute_code]
# System prompt to guide the agent
available_regions = list_extracted_data()
system_prompt = f"""You are an AI assistant that helps analyze spreadsheet data extracted by LlamaSheets.
LlamaSheets extracts messy spreadsheets into clean parquet files with two types of outputs:
1. Region files (*_region_*.parquet) - The actual data with columns and rows
2. Metadata files (*_metadata_*.parquet) - Rich cell-level metadata including:
- Formatting: font_bold, font_italic, font_size, background_color_rgb
- Position: row_number, column_number, coordinate
- Type detection: data_type, is_date_like, is_percentage, is_currency
- Layout: is_in_first_row, is_merged_cell, horizontal_alignment
You have access to tools that allow you to execute Python pandas code against these files.
Use these tools to load the parquet files, analyze the data, and return results.
Key tips:
- Bold cells in metadata often indicate headers
- Background colors often indicate groupings or departments
- Load both region and metadata files for complete analysis
- Write clear pandas code - you have full pandas functionality available
- Store results in variables for reuse across multiple code executions
Existing Processed Regions:
{available_regions}
"""
# Configure agent
return FunctionAgent(tools=tools, llm=llm, system_prompt=system_prompt)
async def main():
"""Example of using the LlamaSheets agent."""
# Create the agent
agent = create_llamasheets_agent()
ctx = Context(agent)
# Example queries the agent can handle:
queries = [
# Discovery
"What spreadsheet data is available?",
# Simple analysis
"Load the sales data and show me the first few rows with column info",
# Using metadata
"Find all bold cells in the metadata - these are likely headers",
]
# Example: Run a query
for query in queries:
print(f"\n=== Query: {query} ===")
handler = agent.run(query, ctx=ctx)
async for ev in handler.stream_events():
if isinstance(ev, ToolCall):
tool_kwargs_str = (
str(ev.tool_kwargs)[:500] + " ..."
if len(str(ev.tool_kwargs)) > 500
else str(ev.tool_kwargs)
)
print(f"\n[Tool Call] {ev.tool_name} with args:\n{tool_kwargs_str}\n\n")
elif isinstance(ev, ToolCallResult):
result_str = (
str(ev.tool_output)[:500] + " ..."
if len(str(ev.tool_output)) > 500
else str(ev.tool_output)
)
print(f"\n[Tool Result] {ev.tool_name}:\n{result_str}\n\n")
elif isinstance(ev, AgentStream):
print(ev.delta, end="", flush=True)
_ = await handler
print("\n=== End Query ===\n")
if __name__ == "__main__":
import asyncio
asyncio.run(main())
@@ -0,0 +1,7 @@
llama-cloud-services # LlamaSheets SDK
llama-index-core
llama-index-llms-openai
pandas>=2.0.0
pyarrow>=12.0.0
openpyxl>=3.0.0 # For Excel file support
matplotlib>=3.7.0 # For visualizations (optional)
@@ -0,0 +1,100 @@
"""Helper script to extract spreadsheets using LlamaSheets."""
import asyncio
import json
import os
import dotenv
from pathlib import Path
from llama_cloud_services.beta.sheets import LlamaSheets
from llama_cloud_services.beta.sheets.types import (
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
dotenv.load_dotenv()
async def extract_spreadsheet(
file_path: str, output_dir: str = "data", generate_metadata: bool = True
) -> dict:
"""Extract a spreadsheet using LlamaSheets."""
client = LlamaSheets(
base_url="https://api.cloud.llamaindex.ai",
api_key=os.getenv("LLAMA_CLOUD_API_KEY"),
)
print(f"Extracting {file_path}...")
# Extract regions
config = SpreadsheetParsingConfig(
sheet_names=None, # Extract all sheets
generate_additional_metadata=generate_metadata,
)
job_result = await client.aextract_regions(file_path, config=config)
print(f"Extracted {len(job_result.regions)} region(s)")
# Create output directory
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
# Get base name for files
base_name = Path(file_path).stem
# Save job metadata
job_metadata_path = output_path / f"{base_name}_job_metadata.json"
with open(job_metadata_path, "w") as f:
json.dump(job_result.model_dump(mode="json"), f, indent=2)
print(f"Saved job metadata to {job_metadata_path}")
# Download each region
for idx, region in enumerate(job_result.regions, 1):
sheet_name = region.sheet_name.replace(" ", "_")
# Download region data
region_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=region.region_type,
)
region_path = output_path / f"{base_name}_region_{idx}_{sheet_name}.parquet"
with open(region_path, "wb") as f:
f.write(region_bytes)
print(f" Table {idx}: {region_path}")
# Download metadata
metadata_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=SpreadsheetResultType.CELL_METADATA,
)
metadata_path = output_path / f"{base_name}_metadata_{idx}_{sheet_name}.parquet"
with open(metadata_path, "wb") as f:
f.write(metadata_bytes)
print(f" Metadata {idx}: {metadata_path}")
print(f"\nAll files saved to {output_path}/")
return job_result.model_dump(mode="json")
if __name__ == "__main__":
import sys
if len(sys.argv) < 2:
print("Usage: python scripts/extract.py <spreadsheet_file>")
sys.exit(1)
file_path = sys.argv[1]
if not Path(file_path).exists():
print(f"❌ File not found: {file_path}")
sys.exit(1)
result = asyncio.run(extract_spreadsheet(file_path))
print(f"\n✅ Extraction complete! Job ID: {result['id']}")
@@ -0,0 +1,540 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Document Splitting with LlamaCloud\n",
"\n",
"This notebook demonstrates how to use the LlamaCloud **Split** API to automatically segment a concatenated PDF into logical document sections based on content categories.\n",
"\n",
"## Use Case\n",
"\n",
"When dealing with large PDFs that contain multiple distinct documents or sections (e.g., a bundle of research papers, a collection of reports), you often need to split them into individual segments. The Split API uses AI to:\n",
"\n",
"1. Analyze each page's content\n",
"2. Classify pages into user-defined categories\n",
"3. Group consecutive pages of the same category into segments\n",
"\n",
"## Example Document\n",
"\n",
"We'll use a PDF containing three concatenated documents:\n",
"- **Alan Turing's essay** \"Intelligent Machinery, A Heretical Theory\" (an essay)\n",
"- **ImageNet paper** (a research paper)\n",
"- **\"Attention is All You Need\"** paper (a research paper)\n",
"\n",
"We'll split this into segments categorized as either `essay` or `research_paper`.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Requirement already satisfied: llama-cloud in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (0.1.44)\n",
"Requirement already satisfied: python-dotenv in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (1.2.1)\n",
"Requirement already satisfied: requests in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (2.32.5)\n",
"Requirement already satisfied: certifi>=2024.7.4 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from llama-cloud) (2025.11.12)\n",
"Requirement already satisfied: httpx>=0.20.0 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from llama-cloud) (0.28.1)\n",
"Requirement already satisfied: pydantic>=1.10 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from llama-cloud) (2.12.5)\n",
"Requirement already satisfied: charset_normalizer<4,>=2 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from requests) (3.4.4)\n",
"Requirement already satisfied: idna<4,>=2.5 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from requests) (3.11)\n",
"Requirement already satisfied: urllib3<3,>=1.21.1 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from requests) (2.5.0)\n",
"Requirement already satisfied: anyio in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from httpx>=0.20.0->llama-cloud) (4.11.0)\n",
"Requirement already satisfied: httpcore==1.* in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from httpx>=0.20.0->llama-cloud) (1.0.9)\n",
"Requirement already satisfied: h11>=0.16 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from httpcore==1.*->httpx>=0.20.0->llama-cloud) (0.16.0)\n",
"Requirement already satisfied: annotated-types>=0.6.0 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from pydantic>=1.10->llama-cloud) (0.7.0)\n",
"Requirement already satisfied: pydantic-core==2.41.5 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from pydantic>=1.10->llama-cloud) (2.41.5)\n",
"Requirement already satisfied: typing-extensions>=4.14.1 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from pydantic>=1.10->llama-cloud) (4.15.0)\n",
"Requirement already satisfied: typing-inspection>=0.4.2 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from pydantic>=1.10->llama-cloud) (0.4.2)\n",
"Requirement already satisfied: sniffio>=1.1 in /Users/javier/llama_cloud_services/.venv/lib/python3.11/site-packages (from anyio->httpx>=0.20.0->llama-cloud) (1.3.1)\n",
"\n",
"\u001b[1m[\u001b[0m\u001b[34;49mnotice\u001b[0m\u001b[1;39;49m]\u001b[0m\u001b[39;49m A new release of pip is available: \u001b[0m\u001b[31;49m25.0.1\u001b[0m\u001b[39;49m -> \u001b[0m\u001b[32;49m25.3\u001b[0m\n",
"\u001b[1m[\u001b[0m\u001b[34;49mnotice\u001b[0m\u001b[1;39;49m]\u001b[0m\u001b[39;49m To update, run: \u001b[0m\u001b[32;49mpip install --upgrade pip\u001b[0m\n",
"Note: you may need to restart the kernel to use updated packages.\n"
]
}
],
"source": [
"# Install required packages\n",
"%pip install llama-cloud python-dotenv requests"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"✅ API configured with base URL: https://api.cloud.llamaindex.ai\n",
"✅ Project ID: using default project\n"
]
}
],
"source": [
"import os\n",
"import time\n",
"import requests\n",
"from dotenv import load_dotenv\n",
"\n",
"# Load environment variables\n",
"load_dotenv()\n",
"\n",
"# Configuration\n",
"LLAMA_CLOUD_API_KEY = os.environ.get(\"LLAMA_CLOUD_API_KEY\", \"llx-...\")\n",
"BASE_URL = os.environ.get(\"LLAMA_CLOUD_BASE_URL\", \"https://api.cloud.llamaindex.ai\")\n",
"PROJECT_ID = os.environ.get(\"LLAMA_CLOUD_PROJECT_ID\", None)\n",
"\n",
"# Headers for API requests\n",
"headers = {\n",
" \"Authorization\": f\"Bearer {LLAMA_CLOUD_API_KEY}\",\n",
" \"Content-Type\": \"application/json\",\n",
"}\n",
"\n",
"print(f\"✅ API configured with base URL: {BASE_URL}\")\n",
"print(f\"✅ Project ID: {PROJECT_ID or 'using default project'}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 1: Upload the PDF File\n",
"\n",
"First, we'll upload our concatenated PDF to LlamaCloud using the Files API. This can be done using the `llama-cloud` SDK.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📤 Uploading ./data/turing+imagenet+attention.pdf...\n",
"✅ File uploaded successfully!\n",
" File name: turing+imagenet+attention.pdf\n"
]
}
],
"source": [
"from llama_cloud.client import LlamaCloud\n",
"\n",
"# Initialize the client\n",
"client = LlamaCloud(token=LLAMA_CLOUD_API_KEY, base_url=BASE_URL)\n",
"\n",
"# Path to the PDF file\n",
"pdf_path = \"./data/turing+imagenet+attention.pdf\"\n",
"\n",
"# Upload the file\n",
"print(f\"📤 Uploading {pdf_path}...\")\n",
"\n",
"with open(pdf_path, \"rb\") as f:\n",
" uploaded_file = client.files.upload_file(upload_file=f, project_id=PROJECT_ID)\n",
"\n",
"file_id = uploaded_file.id\n",
"print(f\"✅ File uploaded successfully!\")\n",
"print(f\" File name: {uploaded_file.name}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Create a Split Job\n",
"\n",
"Now we'll create a split job using the Split API. Since the Split API is in beta and not yet available in the SDK, we'll use raw HTTP requests.\n",
"\n",
"We define two categories:\n",
"- **essay**: For philosophical or reflective writing\n",
"- **research_paper**: For formal academic documents with methodology and citations\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🔄 Creating split job...\n",
"✅ Split job created!\n",
" Job ID: spl-zsssb632a742aikliu96pqkb56t5\n",
" Status: pending\n",
" Categories: ['essay', 'research_paper']\n"
]
}
],
"source": [
"# Define the split job request\n",
"split_request = {\n",
" \"document_input\": {\n",
" \"type\": \"file_id\", # only file_id is supported for now\n",
" \"value\": file_id,\n",
" },\n",
" \"categories\": [\n",
" {\n",
" \"name\": \"essay\",\n",
" \"description\": \"A philosophical or reflective piece of writing that presents personal viewpoints, arguments, or thoughts on a topic without strict formal structure\",\n",
" },\n",
" {\n",
" \"name\": \"research_paper\",\n",
" \"description\": \"A formal academic document presenting original research, methodology, experiments, results, and conclusions with citations and references\",\n",
" },\n",
" ],\n",
"}\n",
"\n",
"# Create the split job\n",
"print(\"🔄 Creating split job...\")\n",
"response = requests.post(\n",
" f\"{BASE_URL}/api/v1/beta/split/jobs\",\n",
" params={\"project_id\": PROJECT_ID},\n",
" headers=headers,\n",
" json=split_request,\n",
")\n",
"response.raise_for_status()\n",
"\n",
"split_job = response.json()\n",
"job_id = split_job[\"id\"]\n",
"\n",
"print(f\"✅ Split job created!\")\n",
"print(f\" Job ID: {job_id}\")\n",
"print(f\" Status: {split_job['status']}\")\n",
"print(f\" Categories: {[c['name'] for c in split_job['categories']]}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Poll for Job Completion\n",
"\n",
"The split job runs asynchronously. We'll poll the job status until it completes.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"⏳ Waiting for split job to complete...\n",
" Status: processing (elapsed: 0s)\n",
" Status: processing (elapsed: 5s)\n",
" Status: processing (elapsed: 11s)\n",
" Status: completed (elapsed: 16s)\n",
"\n",
"✅ Split job completed successfully!\n"
]
}
],
"source": [
"def poll_split_job(job_id: str, max_wait_seconds: int = 180, poll_interval: int = 5):\n",
" \"\"\"\n",
" Poll a split job until it reaches a terminal state.\n",
"\n",
" Args:\n",
" job_id: The split job ID\n",
" max_wait_seconds: Maximum time to wait for completion\n",
" poll_interval: Seconds between poll attempts\n",
"\n",
" Returns:\n",
" The completed job response\n",
" \"\"\"\n",
" start_time = time.time()\n",
"\n",
" while (time.time() - start_time) < max_wait_seconds:\n",
" response = requests.get(\n",
" f\"{BASE_URL}/api/v1/beta/split/jobs/{job_id}\",\n",
" params={\"project_id\": PROJECT_ID},\n",
" headers=headers,\n",
" )\n",
" response.raise_for_status()\n",
" job = response.json()\n",
"\n",
" status = job[\"status\"]\n",
" elapsed = int(time.time() - start_time)\n",
" print(f\" Status: {status} (elapsed: {elapsed}s)\")\n",
"\n",
" if status in [\"completed\", \"failed\"]:\n",
" return job\n",
"\n",
" time.sleep(poll_interval)\n",
"\n",
" raise TimeoutError(f\"Job did not complete within {max_wait_seconds} seconds\")\n",
"\n",
"\n",
"print(\"⏳ Waiting for split job to complete...\")\n",
"completed_job = poll_split_job(job_id)\n",
"\n",
"if completed_job[\"status\"] == \"completed\":\n",
" print(\"\\n✅ Split job completed successfully!\")\n",
"else:\n",
" print(\n",
" f\"\\n❌ Split job failed: {completed_job.get('error_message', 'Unknown error')}\"\n",
" )"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Analyze the Results\n",
"\n",
"Let's examine the split results to see how the document was segmented.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📊 Split Results Summary\n",
"==================================================\n",
"Total segments found: 3\n",
"\n",
"Segments by category:\n",
" • essay: 1 segment(s)\n",
" • research_paper: 2 segment(s)\n"
]
}
],
"source": [
"# Get the segments from the result\n",
"segments = completed_job.get(\"result\", {}).get(\"segments\", [])\n",
"\n",
"print(f\"📊 Split Results Summary\")\n",
"print(f\"=\" * 50)\n",
"print(f\"Total segments found: {len(segments)}\")\n",
"print()\n",
"\n",
"# Count by category\n",
"category_counts = {}\n",
"for segment in segments:\n",
" cat = segment[\"category\"]\n",
" category_counts[cat] = category_counts.get(cat, 0) + 1\n",
"\n",
"print(\"Segments by category:\")\n",
"for cat, count in category_counts.items():\n",
" print(f\" • {cat}: {count} segment(s)\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"📄 Segment Details\n",
"==================================================\n",
"\n",
"Segment 1:\n",
" Category: essay\n",
" Pages 1-4 (4 pages)\n",
" Confidence: high\n",
"\n",
"Segment 2:\n",
" Category: research_paper\n",
" Pages 5-13 (9 pages)\n",
" Confidence: high\n",
"\n",
"Segment 3:\n",
" Category: research_paper\n",
" Pages 14-24 (11 pages)\n",
" Confidence: high\n"
]
}
],
"source": [
"# Display detailed segment information\n",
"print(f\"\\n📄 Segment Details\")\n",
"print(f\"=\" * 50)\n",
"\n",
"for i, segment in enumerate(segments, 1):\n",
" category = segment[\"category\"]\n",
" pages = segment[\"pages\"]\n",
" confidence = segment[\"confidence_category\"]\n",
"\n",
" # Format page range\n",
" if len(pages) == 1:\n",
" page_range = f\"Page {pages[0]}\"\n",
" else:\n",
" page_range = f\"Pages {min(pages)}-{max(pages)}\"\n",
"\n",
" print(f\"\\nSegment {i}:\")\n",
" print(f\" Category: {category}\")\n",
" print(f\" {page_range} ({len(pages)} page{'s' if len(pages) > 1 else ''})\")\n",
" print(f\" Confidence: {confidence}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Expected Results\n",
"\n",
"Based on our test document, we expect:\n",
"- **1 essay segment**: Alan Turing's \"Intelligent Machinery, A Heretical Theory\"\n",
"- **2 research paper segments**: ImageNet paper and \"Attention is All You Need\" paper\n",
"\n",
"The pages should be grouped consecutively, with no overlap between segments.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"✅ Validation\n",
"==================================================\n",
"Total pages assigned: 24\n",
"Unique pages: 24\n",
"✅ No page overlap detected - each page belongs to exactly one segment\n"
]
}
],
"source": [
"# Verify no page overlap\n",
"all_pages = []\n",
"for segment in segments:\n",
" all_pages.extend(segment[\"pages\"])\n",
"\n",
"unique_pages = set(all_pages)\n",
"\n",
"print(f\"\\n✅ Validation\")\n",
"print(f\"=\" * 50)\n",
"print(f\"Total pages assigned: {len(all_pages)}\")\n",
"print(f\"Unique pages: {len(unique_pages)}\")\n",
"\n",
"if len(all_pages) == len(unique_pages):\n",
" print(f\"✅ No page overlap detected - each page belongs to exactly one segment\")\n",
"else:\n",
" print(\n",
" f\"⚠️ Page overlap detected - {len(all_pages) - len(unique_pages)} duplicate assignments\"\n",
" )"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using `allow_uncategorized` Strategy\n",
"\n",
"You can also use the `allow_uncategorized` splitting strategy. This is useful when you want to capture pages that don't match any defined category.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📝 With allow_uncategorized=True and only 'essay' category defined,\n",
" pages that don't match 'essay' will be grouped as 'uncategorized'.\n"
]
}
],
"source": [
"# Example with allow_uncategorized strategy\n",
"split_request_uncategorized = {\n",
" \"document_input\": {\"type\": \"file_id\", \"value\": file_id},\n",
" \"categories\": [\n",
" {\n",
" \"name\": \"essay\",\n",
" \"description\": \"A philosophical or reflective piece of writing that presents personal viewpoints, arguments, or thoughts on a topic\",\n",
" }\n",
" # Note: We only define 'essay' category\n",
" # Research papers will be classified as 'uncategorized'\n",
" ],\n",
" \"splitting_strategy\": {\"allow_uncategorized\": True},\n",
"}\n",
"\n",
"print(\"📝 With allow_uncategorized=True and only 'essay' category defined,\")\n",
"print(\" pages that don't match 'essay' will be grouped as 'uncategorized'.\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conclusion\n",
"\n",
"The LlamaCloud Split API provides a powerful way to automatically segment concatenated documents based on content categories. This is useful for:\n",
"\n",
"- **Document processing pipelines**: Automatically separate bundled documents before further processing\n",
"- **Content organization**: Categorize and organize mixed document collections\n",
"- **Information extraction**: Identify different document types within a single file\n",
"\n",
"### Key Features\n",
"\n",
"- **AI-powered classification**: Uses LLMs to understand page content and assign categories\n",
"- **Flexible categories**: Define any categories relevant to your use case\n",
"- **Confidence scoring**: Each segment includes a confidence level\n",
"- **Page-level granularity**: Results include exact page numbers for each segment\n",
"\n",
"### API Reference\n",
"\n",
"- **Create Split Job**: `POST /api/v1/beta/split/jobs`\n",
"- **Get Split Job**: `GET /api/v1/beta/split/jobs/{job_id}`\n",
"- **List Split Jobs**: `GET /api/v1/beta/split/jobs`\n"
]
}
],
"metadata": {
"kernelspec": {
"display_name": ".venv",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
+1 -1
View File
@@ -398,6 +398,6 @@ Another option (orthogonal to the above) is to break the document into smaller s
## Additional Resources
- [Extract Documentation](https://docs.cloud.llamaindex.ai/llamaextract/getting_started) - Details on Extract features, API and examples.
- [Example Notebook](docs/examples-py/extract/resume_screening.ipynb) - Detailed walkthrough of resume parsing
- [Example Notebook](examples/extract/resume_screening.ipynb) - Detailed walkthrough of resume parsing
- [Example Application with TypeScript](./examples-ts/extract/) - End-to-end examples using LlamaExtract TypeScript client.
- [Discord Community](https://discord.com/invite/eN6D2HQ4aX) - Get help and share feedback
+26
View File
@@ -0,0 +1,26 @@
{
"name": "llama-cloud-services-workspace",
"version": "0.0.1",
"description": "",
"private": true,
"keywords": [],
"author": "",
"scripts": {
"pre-commit-version": "pnpm changeset",
"version": "./scripts/changeset-version.py version",
"publish": "./scripts/changeset-version.py publish --tag"
},
"devDependencies": {
"prettier": "^3.6.2",
"lint-staged": "^15.4.2",
"@changesets/cli": "^2.29.5",
"changesets": "^1.0.2"
},
"lint-staged": {
"ts/llama_cloud_services/src/**/*.{ts,tsx,js,jsx}": [
"pnpm --filter llama-cloud-services exec eslint --fix",
"pnpm --filter llama-cloud-services exec prettier --write src/ tests/"
]
},
"packageManager": "pnpm@10.11.1+sha512.e519b9f7639869dc8d5c3c5dfef73b3f091094b0a006d7317353c72b124e80e1afd429732e28705ad6bfa1ee879c1fce46c128ccebd3192101f43dd67c667912"
}
+6 -6
View File
@@ -97,7 +97,7 @@ for page in result.pages:
print(page.structuredData)
```
See more details about the result object in the [example notebook](./docs/examples-py/parse/demo_json_tour.ipynb).
See more details about the result object in the [example notebook](./examples/parse/demo_json_tour.ipynb).
### Using with file object / bytes
@@ -147,16 +147,16 @@ documents = SimpleDirectoryReader(
).load_data()
```
Full documentation for `SimpleDirectoryReader` can be found on the [LlamaIndex Documentation](https://docs.llamaindex.ai/en/stable/module_guides/loading/simpledirectoryreader.html).
Full documentation for `SimpleDirectoryReader` can be found on the [LlamaIndex Documentation](https://developers.llamaindex.ai/python/framework/module_guides/loading/simpledirectoryreader/).
## Examples
Several end-to-end indexing examples can be found in the examples folder
- [Getting Started](docs/examples-py/parse/demo_basic.ipynb)
- [Advanced RAG Example](docs/examples-py/parse/demo_advanced.ipynb)
- [Raw API Usage](docs/examples-py/parse/demo_api.ipynb)
- [Result Object Tour](docs/examples-py/parse/demo_json_tour.ipynb)
- [Getting Started](examples/parse/demo_basic.ipynb)
- [Advanced RAG Example](examples/parse/demo_advanced.ipynb)
- [Raw API Usage](examples/parse/demo_api.ipynb)
- [Result Object Tour](examples/parse/demo_json_tour.ipynb)
## Documentation
+1049 -88
View File
File diff suppressed because it is too large Load Diff
+3 -1
View File
@@ -1,2 +1,4 @@
packages:
- "ts/**"
- "ts/*"
- "py"
- "py/*"
+128
View File
@@ -0,0 +1,128 @@
# llama-cloud-services-py
## 0.6.90
### Patch Changes
- 19cbb25: Remove extension filter
## 0.6.89
### Patch Changes
- b9b83c9: Parse bounding boxes from extract jobs results in agent data
## 0.6.88
### Patch Changes
- 71db318: Add tier and version
## 0.6.87
### Patch Changes
- 06c3c55: Update spreadsheet parsing config
## 0.6.86
### Patch Changes
- 1b7198d: Update extract to have confidence scores available in all modes
## 0.6.85
### Patch Changes
- ae30990: Add line-level bbox support
## 0.6.84
### Patch Changes
- 0a110de: Release to re-align versions
## 0.6.83
### Patch Changes
- ca78113: Do not use presigned URLs by default in files client
## 0.6.82
### Patch Changes
- bfaec79: Update for new page number params
## 0.6.81
### Patch Changes
- f3233de: Propagate retrieval metadata to retriever nodes
## 0.6.80
### Patch Changes
- 0506c88: Moved ClassifyClient to LlamaClassify (backward compatible)
## 0.6.79
### Patch Changes
- e020e3e: Remove unneeded organization_id param from beta classifier client
## 0.6.78
### Patch Changes
- 9f1ef4e: Fix extract
## 0.6.77
### Patch Changes
- 407292b: Now return partial results on job failure
## 0.6.76
### Patch Changes
- 4f24f53: Add aggressive_table_extraction flag in python sdk
## 0.6.75
### Patch Changes
- f81532e: Safest types possible for parse
## 0.6.74
### Patch Changes
- 1bf5223: Fix default bbox values
- 24166dc: Now only escape single dollar signs - preserve double for latex equations
## 0.6.73
### Patch Changes
- e6a7939: Loosen packaging dep requirement
## 0.6.72
### Patch Changes
- ad6734b: Fixup and test versioning
## 0.6.71
### Patch Changes
- 51011b9: Escape dollar signs in jupyter notebooks
## 0.6.70
### Patch Changes
- d028397: Update llama-cloud api version, and integrate with agent data deletion
+1 -1
View File
@@ -15,4 +15,4 @@ test: ## Run unit tests via pytest
.PHONY: e2e
e2e: ## Run all tests. Run with high parallelism using xdist since tests are bottlenecked bound by the slow backend parsing
uv run pytest -v -n 32 tests/
uv run pytest -v -n 32 --timeout=300 --session-timeout=1740 tests/
+1 -7
View File
@@ -9,7 +9,6 @@ This repository contains the code for hand-written SDKs and clients for interact
This includes:
- [LlamaParse](../parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](../report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract](../extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](../index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
@@ -28,14 +27,12 @@ Then, you can use the services in your code:
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
)
from llama_cloud_services import LlamaParse, LlamaReport, LlamaExtract
from llama_cloud_services import LlamaParse, LlamaExtract
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
@@ -45,7 +42,6 @@ index = LlamaCloudIndex(
See the quickstart guides for each service for more information:
- [LlamaParse](../parse.md)
- [LlamaReport (beta/invite-only)](../report.md)
- [LlamaExtract](../extract.md)
- [LlamaCloud Index](../index.md)
@@ -58,13 +54,11 @@ You can also create your API key in the EU region [here](https://cloud.eu.llamai
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
EU_BASE_URL,
)
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
+3 -3
View File
@@ -1,6 +1,6 @@
from llama_cloud_services.parse import LlamaParse
from llama_cloud_services.report import ReportClient, LlamaReport
from llama_cloud_services.extract import LlamaExtract, ExtractionAgent
from llama_cloud_services.utils import SourceText, FileInput
from llama_cloud_services.constants import EU_BASE_URL
from llama_cloud_services.index import (
LlamaCloudCompositeRetriever,
@@ -10,10 +10,10 @@ from llama_cloud_services.index import (
__all__ = [
"LlamaParse",
"ReportClient",
"LlamaReport",
"LlamaExtract",
"ExtractionAgent",
"SourceText",
"FileInput",
"EU_BASE_URL",
"LlamaCloudIndex",
"LlamaCloudRetriever",
@@ -11,6 +11,9 @@ from .schema import (
InvalidExtractionData,
ExtractedFieldMetadata,
ExtractedFieldMetaDataDict,
FieldCitation,
BoundingBox,
PageDimensions,
)
from .client import AsyncAgentDataClient
@@ -28,4 +31,7 @@ __all__ = [
"InvalidExtractionData",
"ExtractedFieldMetadata",
"ExtractedFieldMetaDataDict",
"FieldCitation",
"BoundingBox",
"PageDimensions",
]
@@ -1,6 +1,11 @@
import os
from typing import Any, Dict, Generic, List, Optional, Type
from llama_cloud import (
AgentData,
PaginatedResponseAgentData,
PaginatedResponseAggregateGroup,
)
from llama_cloud.client import AsyncLlamaCloud
from tenacity import (
WrappedFn,
@@ -86,7 +91,7 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
client=llama_client,
type=ExtractedPerson,
collection="extracted_people",
agent_url_id="person-extraction-agent"
deployment_name="person-extraction-agent"
)
# Create data
@@ -109,10 +114,12 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
self,
type: Type[AgentDataT],
collection: str = "default",
agent_url_id: Optional[str] = None,
deployment_name: Optional[str] = None,
client: Optional[AsyncLlamaCloud] = None,
token: Optional[str] = None,
base_url: Optional[str] = None,
# deprecated, use deployment_name instead
agent_url_id: Optional[str] = None,
):
"""
Initialize the AsyncAgentDataClient.
@@ -123,11 +130,11 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
collection: Named collection within the agent for organizing data.
Defaults to "default". Collections allow logical separation of
different data types or workflows within the same agent.
agent_url_id: Unique identifier for the agent. This normally appears in the
url of an agent within the llama cloud platform. If not provided,
will attempt to use the LLAMA_DEPLOY_DEPLOYMENT_NAME environment
variable. Data can only be added to an already existing agent in the
platform.
deployment_name: Unique identifier for the agent deployment. This normally
appears in the URL of an agent within the Llama Cloud platform. If not
provided, will attempt to use the LLAMA_DEPLOY_DEPLOYMENT_NAME
environment variable. Data can only be added to an already existing
agent in the platform.
client: AsyncLlamaCloud client instance for API communication. If not provided, will
construct one from the provided api token and base url
token: Llama Cloud API token. Reads from LLAMA_CLOUD_API_KEY if not provided
@@ -135,15 +142,14 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
defaults to https://api.cloud.llamaindex.ai
Raises:
ValueError: If agent_url_id is not provided and the
ValueError: If deployment_name is not provided and the
LLAMA_DEPLOY_DEPLOYMENT_NAME environment variable is not set
Note:
The client automatically applies retry logic to all API calls with
exponential backoff for timeout, connection, and HTTP status errors.
"""
self.agent_url_id = agent_url_id or get_default_agent_id()
self.deployment_name = deployment_name or agent_url_id or get_default_agent_id()
self.collection = collection
if not client:
@@ -156,15 +162,19 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
@agent_data_retry
async def get_item(self, item_id: str) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.get_agent_data(
raw_data = await self.untyped_get_item(item_id)
return TypedAgentData.from_raw(raw_data, self.type)
@agent_data_retry
async def untyped_get_item(self, item_id: str) -> AgentData:
return await self.client.beta.get_agent_data(
item_id=item_id,
)
return TypedAgentData.from_raw(raw_data, validator=self.type)
@agent_data_retry
async def create_item(self, data: AgentDataT) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.create_agent_data(
agent_slug=self.agent_url_id,
deployment_name=self.deployment_name,
collection=self.collection,
data=data.model_dump(),
)
@@ -184,6 +194,21 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
async def delete_item(self, item_id: str) -> None:
await self.client.beta.delete_agent_data(item_id=item_id)
@agent_data_retry
async def delete(
self, filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None
) -> int:
"""
Delete agent data by query, similar to search.
Returns the number of deleted items.
"""
response = await self.client.beta.delete_agent_data_by_query_api_v_1_beta_agent_data_delete_post(
deployment_name=self.deployment_name,
collection=self.collection,
filter=filter,
)
return response.deleted_count
@agent_data_retry
async def search(
self,
@@ -210,9 +235,7 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
offset: Number of items to skip from the beginning. Defaults to 0.
include_total: Whether to include the total count in the response. Defaults to False to improve performance. It's recommended to only request on the first page.
"""
raw = await self.client.beta.search_agent_data_api_v_1_beta_agent_data_search_post(
agent_slug=self.agent_url_id,
collection=self.collection,
raw = await self.untyped_search(
filter=filter,
order_by=order_by,
offset=offset,
@@ -227,6 +250,25 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
total=raw.total_size,
)
@agent_data_retry
async def untyped_search(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
include_total: bool = False,
) -> PaginatedResponseAgentData:
return await self.client.beta.search_agent_data_api_v_1_beta_agent_data_search_post(
deployment_name=self.deployment_name,
collection=self.collection,
filter=filter,
order_by=order_by,
offset=offset,
page_size=page_size,
include_total=include_total,
)
@agent_data_retry
async def aggregate(
self,
@@ -253,8 +295,38 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
offset: Number of groups to skip from the beginning. Defaults to 0.
page_size: Maximum number of groups to return per page.
"""
raw = await self.client.beta.aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
agent_slug=self.agent_url_id,
raw = await self.untyped_aggregate(
filter=filter,
group_by=group_by,
count=count,
first=first,
order_by=order_by,
offset=offset,
page_size=page_size,
)
return TypedAggregateGroupItems(
items=[
TypedAggregateGroup.from_raw(grp, validator=self.type)
for grp in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@agent_data_retry
async def untyped_aggregate(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
group_by: Optional[List[str]] = None,
count: Optional[bool] = None,
first: Optional[bool] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
) -> PaginatedResponseAggregateGroup:
return await self.client.beta.aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
deployment_name=self.deployment_name,
collection=self.collection,
page_size=page_size,
filter=filter,
@@ -264,11 +336,3 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
first=first,
offset=offset,
)
return TypedAggregateGroupItems(
items=[
TypedAggregateGroup.from_raw(item, validator=self.type)
for item in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@@ -10,7 +10,7 @@ CRUD operations, search capabilities, filtering, and aggregation functionality
for managing agent-generated data at scale.
Key Concepts:
- Agent Slug: Unique identifier for an agent instance
- Deployment Name: Unique identifier for an agent deployment
- Collection: Named grouping of data within an agent (defaults to "default"). Data within a collection should be of the same type.
- Agent Data: Individual structured data records with metadata and timestamps
@@ -26,7 +26,7 @@ Example Usage:
client=async_llama_cloud,
type=Person,
collection="people",
agent_url_id="my-extraction-agent-xyz"
deployment_name="my-extraction-agent-xyz"
)
# Create typed data
@@ -56,7 +56,6 @@ from typing import (
# Type variable for user-defined data models
AgentDataT = TypeVar("AgentDataT", bound=BaseModel)
# Type variable for extracted data (can be dict or Pydantic model)
ExtractedT = TypeVar("ExtractedT", bound=Union[BaseModel, dict])
@@ -78,7 +77,7 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
Attributes:
id: Unique identifier for this data record
agent_url_id: Identifier of the agent that created this data
deployment_name: Identifier of the agent deployment that created this data
collection: Named collection within the agent (used for organization)
data: The actual structured data payload (typed as AgentDataT)
created_at: Timestamp when the record was first created
@@ -94,8 +93,8 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
"""
id: Optional[str] = Field(description="Unique identifier for this data record")
agent_url_id: str = Field(
description="Identifier of the agent that created this data"
deployment_name: str = Field(
description="Identifier of the agent deployment that created this data"
)
collection: Optional[str] = Field(
description="Named collection within the agent for data organization"
@@ -116,15 +115,15 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
Args:
raw_data: Raw agent data from the API
validator: Pydantic model class to validate the data field
Returns:
TypedAgentData instance with validated data
"""
data: AgentDataT = validator.model_validate(raw_data.data)
return cls(
id=raw_data.id,
agent_url_id=raw_data.agent_slug,
deployment_name=raw_data.deployment_name,
collection=raw_data.collection,
data=data,
created_at=raw_data.created_at,
@@ -175,6 +174,22 @@ class TypedAgentDataItems(BaseModel, Generic[AgentDataT]):
)
class BoundingBox(BaseModel):
"""Bounding box coordinates for a citation location on a page."""
x: float = Field(description="X coordinate of the bounding box origin")
y: float = Field(description="Y coordinate of the bounding box origin")
w: float = Field(description="Width of the bounding box")
h: float = Field(description="Height of the bounding box")
class PageDimensions(BaseModel):
"""Dimensions of a page in the source document."""
width: float = Field(description="Width of the page")
height: float = Field(description="Height of the page")
class FieldCitation(BaseModel):
page: Optional[int] = Field(
None, description="The page number that the field occurred on"
@@ -183,6 +198,14 @@ class FieldCitation(BaseModel):
None,
description="The original text this field's value was derived from",
)
bounding_boxes: Optional[List[BoundingBox]] = Field(
None,
description="Bounding boxes indicating where the citation appears on the page",
)
page_dimensions: Optional[PageDimensions] = Field(
None,
description="Dimensions of the page containing the citation",
)
class ExtractedFieldMetadata(BaseModel):
@@ -202,6 +225,10 @@ class ExtractedFieldMetadata(BaseModel):
None,
description="The confidence score for the field based on the extracted text only",
)
parsing_confidence: Optional[float] = Field(
None,
description="The confidence score for the field based on the parsing/OCR quality",
)
citation: Optional[List[FieldCitation]] = Field(
None,
description="The citation for the field, including page number and matching text",
@@ -222,12 +249,16 @@ def parse_extracted_field_metadata(
return {
k: _parse_extracted_field_metadata_recursive(v)
for k, v in field_metadata.items()
if k not in _METADATA_FIELDS_SIBLING_TO_LEAF
and k not in _ADDITIONAL_ROOT_METADATA_FIELDS
if not _is_reasoning_field(k, v) and k not in _ADDITIONAL_ROOT_METADATA_FIELDS
}
_METADATA_FIELDS_SIBLING_TO_LEAF = {"reasoning"}
def _is_reasoning_field(field_name: str, field_value: Any) -> bool:
# There can either be a user specified reasoning field (from the schema), or a reasoning metadata field for the
# dict of values
return field_name == "reasoning" and isinstance(field_value, str)
_ADDITIONAL_ROOT_METADATA_FIELDS = {"error"}
@@ -257,14 +288,12 @@ def _parse_extracted_field_metadata_recursive(
except ValidationError:
pass
additional_fields = {
k: v
for k, v in field_value.items()
if k in _METADATA_FIELDS_SIBLING_TO_LEAF
k: v for k, v in field_value.items() if _is_reasoning_field(k, v)
}
return {
k: _parse_extracted_field_metadata_recursive(v, additional_fields)
for k, v in field_value.items()
if k not in _METADATA_FIELDS_SIBLING_TO_LEAF
if not _is_reasoning_field(k, v)
}
elif isinstance(field_value, list):
return [_parse_extracted_field_metadata_recursive(item) for item in field_value]
@@ -446,26 +475,49 @@ class ExtractedData(BaseModel, Generic[ExtractedT]):
},
)
except ValidationError as e:
# Capture the job-level error from the extraction run if available
job_error = result.error
invalid_item = ExtractedData[Dict[str, Any]].create(
data=result.data or {},
status="error",
field_metadata=field_metadata,
metadata={"extraction_error": str(e), **(metadata or {})},
metadata={
"extraction_error": str(e),
**({"job_error": job_error} if job_error else {}),
**(metadata or {}),
},
file_id=file_id,
file_name=file_name,
file_hash=file_hash,
)
raise InvalidExtractionData(invalid_item) from e
raise InvalidExtractionData(invalid_item, extraction_error=job_error) from e
class InvalidExtractionData(Exception):
"""
Exception raised when the extracted data does not conform to the schema.
Attributes:
invalid_item: The ExtractedData instance containing the invalid data and metadata
extraction_error: The error message from the extraction job, if available
"""
def __init__(self, invalid_item: ExtractedData[Dict[str, Any]]):
def __init__(
self,
invalid_item: ExtractedData[Dict[str, Any]],
extraction_error: Optional[str] = None,
):
self.invalid_item = invalid_item
super().__init__("Not able to parse the extracted data, parsed invalid format")
self.extraction_error = extraction_error
# Build an informative error message
if extraction_error:
message = f"Extraction error: {extraction_error}"
else:
message = "Not able to parse the extracted data, parsed invalid format"
super().__init__(message)
def calculate_overall_confidence(
@@ -0,0 +1,11 @@
from llama_cloud_services.beta.classifier.client import LlamaClassify, ClassifyClient
from llama_cloud_services.beta.classifier.types import ClassifyJobResultsWithFiles
from llama_cloud_services.utils import SourceText, FileInput
__all__ = [
"LlamaClassify",
"ClassifyClient",
"ClassifyJobResultsWithFiles",
"SourceText",
"FileInput",
]
+176 -40
View File
@@ -1,6 +1,7 @@
import asyncio
import time
from typing import Optional
import warnings
from typing import Optional, List, Union
from pydantic import BaseModel
from llama_cloud.client import AsyncLlamaCloud
from llama_cloud.types import (
@@ -9,14 +10,20 @@ from llama_cloud.types import (
ClassifyJobResults,
ClassifyParsingConfiguration,
StatusEnum,
ClassifyJobWithStatus,
File,
)
from llama_cloud.resources.classifier.client import OMIT
from llama_cloud_services.files.client import FileClient
from llama_cloud_services.constants import POLLING_TIMEOUT_SECONDS
from llama_cloud_services.utils import is_terminal_status, augment_async_errors
from llama_cloud_services.utils import (
is_terminal_status,
augment_async_errors,
FileInput,
)
from llama_index.core.async_utils import DEFAULT_NUM_WORKERS, run_jobs
from llama_cloud_services.beta.classifier.types import (
ClassifyJobResultsWithFiles,
)
class ClassificationOutput(BaseModel):
@@ -24,7 +31,7 @@ class ClassificationOutput(BaseModel):
classification: str
class ClassifyClient:
class LlamaClassify:
"""
Experimental - Client for interacting with the LlamaCloud Classifier API.
The Classification API is currently in beta and may change in the future without notice.
@@ -32,7 +39,6 @@ class ClassifyClient:
Args:
client: The LlamaCloud client to use.
project_id: The project ID to use.
organization_id: The organization ID to use.
polling_interval: The interval to poll for job completion in seconds.
polling_timeout: The timeout for the job to complete in seconds.
"""
@@ -41,17 +47,31 @@ class ClassifyClient:
self,
client: AsyncLlamaCloud,
project_id: Optional[str] = None,
organization_id: Optional[str] = None,
polling_interval: float = 1.0,
polling_timeout: float = POLLING_TIMEOUT_SECONDS,
):
self.client = client
self.project_id = project_id
self.organization_id = organization_id
self.polling_interval = polling_interval
self.file_client = FileClient(client, project_id, organization_id)
self.file_client = FileClient(client, project_id)
self.polling_timeout = polling_timeout
@classmethod
def from_api_key(
cls,
api_key: str,
project_id: Optional[str] = None,
base_url: Optional[str] = None,
) -> "ClassifyClient":
"""
Create a classify client from an API key.
"""
client = AsyncLlamaCloud(token=api_key, base_url=base_url)
return cls(
client,
project_id,
)
async def acreate_classify_job(
self,
rules: list[ClassifierRule],
@@ -76,7 +96,6 @@ class ClassifyClient:
file_ids=file_ids,
parsing_configuration=parsing_configuration or OMIT,
project_id=self.project_id,
organization_id=self.organization_id,
)
def create_classify_job(
@@ -127,7 +146,6 @@ class ClassifyClient:
results = await self.client.classifier.get_classification_job_results(
classify_job_with_status.id,
project_id=self.project_id,
organization_id=self.organization_id,
)
return results
@@ -146,16 +164,115 @@ class ClassifyClient:
)
)
async def aclassify(
self,
rules: list[ClassifierRule],
files: Union[FileInput, List[FileInput]],
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResultsWithFiles:
"""
Classify one or more files from various input types.
Args:
rules: The rules to use for classification.
files: The file(s) to classify. Can be a single file or list of files. Each can be:
- str/Path: File path
- SourceText: Text content or file with explicit filename
- File: Already uploaded file
- BufferedIOBase: File-like object
parsing_configuration: The parsing configuration to use for classification.
raise_on_error: Whether to raise an error if the classification job fails.
workers: Number of parallel workers for uploading files.
show_progress: Whether to show progress bars.
Returns:
The results of the classification job with file metadata.
"""
# Normalize to list
if not isinstance(files, list):
files = [files]
# Upload all files
coroutines = [
self.file_client.upload_content(file_input) for file_input in files
]
uploaded_files: List[File] = await run_jobs(
coroutines,
show_progress=show_progress,
workers=workers,
desc="Uploading files for classification",
)
# Classify
results = await self.aclassify_file_ids(
rules,
[file.id for file in uploaded_files],
parsing_configuration,
raise_on_error,
)
return ClassifyJobResultsWithFiles.from_classify_job_results(
results, uploaded_files
)
def classify(
self,
rules: list[ClassifierRule],
files: Union[FileInput, List[FileInput]],
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResultsWithFiles:
"""
Classify one or more files from various input types (synchronous version).
Args:
rules: The rules to use for classification.
files: The file(s) to classify. Can be a single file or list of files. Each can be:
- str/Path: File path
- SourceText: Text content or file with explicit filename
- File: Already uploaded file
- BufferedIOBase: File-like object
parsing_configuration: The parsing configuration to use for classification.
raise_on_error: Whether to raise an error if the classification job fails.
workers: Number of parallel workers for uploading files.
show_progress: Whether to show progress bars.
Returns:
The results of the classification job with file metadata.
"""
with augment_async_errors():
return asyncio.run(
self.aclassify(
rules,
files,
parsing_configuration,
raise_on_error,
workers,
show_progress,
)
)
async def aclassify_file_path(
self,
rules: list[ClassifierRule],
file_input_path: str,
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResults:
file = await self.file_client.upload_file(file_input_path)
return await self.aclassify_file_ids(
rules, [file.id], parsing_configuration, raise_on_error
) -> ClassifyJobResultsWithFiles:
"""
Deprecated: Use aclassify() instead.
"""
warnings.warn(
"aclassify_file_path is deprecated, use aclassify() instead",
DeprecationWarning,
stacklevel=2,
)
return await self.aclassify(
rules, file_input_path, parsing_configuration, raise_on_error
)
def classify_file_path(
@@ -164,13 +281,18 @@ class ClassifyClient:
file_input_path: str,
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResults:
with augment_async_errors():
return asyncio.run(
self.aclassify_file_path(
rules, file_input_path, parsing_configuration, raise_on_error
)
)
) -> ClassifyJobResultsWithFiles:
"""
Deprecated: Use classify() instead.
"""
warnings.warn(
"classify_file_path is deprecated, use classify() instead",
DeprecationWarning,
stacklevel=2,
)
return self.classify(
rules, file_input_path, parsing_configuration, raise_on_error
)
async def aclassify_file_paths(
self,
@@ -180,16 +302,22 @@ class ClassifyClient:
raise_on_error: bool = True,
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResults:
coroutines = [self.file_client.upload_file(path) for path in file_input_paths]
files: list[File] = await run_jobs(
coroutines,
show_progress=show_progress,
workers=workers,
desc="Uploading files for classification",
) -> ClassifyJobResultsWithFiles:
"""
Deprecated: Use aclassify() instead.
"""
warnings.warn(
"aclassify_file_paths is deprecated, use aclassify() instead",
DeprecationWarning,
stacklevel=2,
)
return await self.aclassify_file_ids(
rules, [file.id for file in files], parsing_configuration, raise_on_error
return await self.aclassify(
rules,
file_input_paths,
parsing_configuration,
raise_on_error,
workers,
show_progress,
)
def classify_file_paths(
@@ -198,15 +326,20 @@ class ClassifyClient:
file_input_paths: list[str],
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResults:
with augment_async_errors():
return asyncio.run(
self.aclassify_file_paths(
rules, file_input_paths, parsing_configuration, raise_on_error
)
)
) -> ClassifyJobResultsWithFiles:
"""
Deprecated: Use classify() instead.
"""
warnings.warn(
"classify_file_paths is deprecated, use classify() instead",
DeprecationWarning,
stacklevel=2,
)
return self.classify(
rules, file_input_paths, parsing_configuration, raise_on_error
)
async def wait_for_job_completion(self, job_id: str) -> ClassifyJobWithStatus:
async def wait_for_job_completion(self, job_id: str) -> ClassifyJob:
"""
Wait for a classify job to complete.
Meant to expose lower level access to classifier jobs for advanced use cases.
@@ -219,7 +352,7 @@ class ClassifyClient:
The classify job with status.
"""
job = await self.client.classifier.get_classify_job(
job_id, project_id=self.project_id, organization_id=self.organization_id
job_id, project_id=self.project_id
)
start_time = time.time()
while not is_terminal_status(job.status):
@@ -230,6 +363,9 @@ class ClassifyClient:
)
await asyncio.sleep(self.polling_interval)
job = await self.client.classifier.get_classify_job(
job_id, project_id=self.project_id, organization_id=self.organization_id
job_id, project_id=self.project_id
)
return job
ClassifyClient = LlamaClassify
@@ -0,0 +1,59 @@
from llama_cloud.types.classify_job_results import ClassifyJobResults
from llama_cloud.types.file_classification import FileClassification
from llama_cloud.types.file import File
class FileClassificationWithFile(FileClassification):
"""
File classification with file object.
"""
file: File
@classmethod
def from_file_classification(
cls, file_classification: FileClassification, file: File
) -> "FileClassificationWithFile":
if file_classification.file_id != file.id:
raise ValueError(
f"File classification ID {file_classification.id} does not match file ID {file.id}"
)
ctor_args = {
**file_classification.dict(),
"file": file,
}
return cls(**ctor_args)
class ClassifyJobResultsWithFiles(ClassifyJobResults):
"""
Classify job results with file objects.
"""
items: list[FileClassificationWithFile]
@classmethod
def from_classify_job_results(
cls, classify_job_results: ClassifyJobResults, files: list[File]
) -> "ClassifyJobResultsWithFiles":
if len(classify_job_results.items) != len(files):
raise ValueError(
f"Number of classify job results {len(classify_job_results.items)} does not match number of files {len(files)}"
)
# create mapping of file classification result to file object
file_id_to_file: dict[str, File] = {file.id: file for file in files}
file_classification_to_file: list[tuple[FileClassification, File]] = []
for item in classify_job_results.items:
if item.file_id not in file_id_to_file:
raise ValueError(
f"File classification result {item.id} has file ID {item.file_id} that does not match any provided file ID"
)
file_classification_to_file.append((item, file_id_to_file[item.file_id]))
# create a list of file classification with file objects
ctor_args = classify_job_results.dict()
ctor_args["items"] = [
FileClassificationWithFile.from_file_classification(item, file)
for item, file in file_classification_to_file
]
return cls(**ctor_args)
@@ -0,0 +1,43 @@
"""LlamaCloud Spreadsheet API SDK
This module provides a Python SDK for the LlamaCloud Spreadsheet API.
"""
from llama_cloud_services.beta.sheets.client import (
LlamaSheets,
SpreadsheetAPIError,
SpreadsheetJobError,
SpreadsheetTimeoutError,
)
from llama_cloud_services.beta.sheets.types import (
ExtractedRegionSummary,
FileUploadResponse,
JobStatus,
PresignedUrlResponse,
SpreadsheetJob,
SpreadsheetJobResult,
SpreadsheetParseResult,
SpreadsheetParsingConfig,
SpreadsheetResultType,
WorksheetMetadata,
)
__all__ = [
# Client
"LlamaSheets",
# Exceptions
"SpreadsheetAPIError",
"SpreadsheetJobError",
"SpreadsheetTimeoutError",
# Types
"ExtractedRegionSummary",
"FileUploadResponse",
"JobStatus",
"PresignedUrlResponse",
"SpreadsheetJob",
"SpreadsheetJobResult",
"SpreadsheetParseResult",
"SpreadsheetParsingConfig",
"SpreadsheetResultType",
"WorksheetMetadata",
]
@@ -0,0 +1,550 @@
import asyncio
import io
import os
import time
from typing import Any, Dict, TYPE_CHECKING
import httpx
from llama_cloud.client import AsyncLlamaCloud
from tenacity import (
AsyncRetrying,
retry_if_exception,
stop_after_attempt,
wait_exponential,
)
from llama_cloud_services.beta.sheets.types import (
FileUploadResponse,
JobStatus,
PresignedUrlResponse,
SpreadsheetJob,
SpreadsheetJobResult,
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
from llama_cloud_services.constants import BASE_URL
from llama_cloud_services.files.client import FileClient
from llama_cloud_services.utils import (
augment_async_errors,
FileInput,
)
if TYPE_CHECKING:
import pandas as pd
def _should_retry_exception(exception: BaseException) -> bool:
"""Determine if an exception should be retried."""
if isinstance(exception, httpx.HTTPStatusError):
return exception.response.status_code in (429, 500, 502, 503, 504)
return False
class SpreadsheetAPIError(Exception):
"""Base exception for spreadsheet API errors"""
pass
class SpreadsheetJobError(SpreadsheetAPIError):
"""Exception raised when a spreadsheet job fails"""
pass
class SpreadsheetTimeoutError(SpreadsheetAPIError):
"""Exception raised when a job times out"""
pass
class LlamaSheets:
"""Client for the LlamaCloud Spreadsheet API"""
def __init__(
self,
api_key: str | None = None,
base_url: str | None = None,
max_timeout: int = 300,
poll_interval: int = 5,
max_retries: int = 3,
project_id: str | None = None,
organization_id: str | None = None,
async_httpx_client: httpx.AsyncClient | None = None,
) -> None:
"""Initialize the LlamaSheets client.
Args:
api_key: API key for authentication. If not provided, will use LLAMA_CLOUD_API_KEY env var
base_url: Base URL for the API
max_timeout: Maximum time to wait for job completion in seconds
poll_interval: Interval between status checks in seconds
max_retries: Maximum number of retries for failed requests
project_id: Project ID for file operations. If not provided, will use LLAMA_CLOUD_PROJECT_ID env var
organization_id: Organization ID for file operations. If not provided, will use LLAMA_CLOUD_ORGANIZATION_ID env var
async_httpx_client: Optional custom async httpx client
"""
self.api_key = api_key or os.environ.get("LLAMA_CLOUD_API_KEY")
if not self.api_key:
raise ValueError(
"An API key must be provided either as an argument or via the LLAMA_CLOUD_API_KEY environment variable."
)
base_url = base_url or os.environ.get("LLAMA_CLOUD_BASE_URL", BASE_URL)
self.base_url = str(base_url).rstrip("/")
self.max_timeout = max_timeout
self.poll_interval = poll_interval
self.max_retries = max_retries
self.project_id = project_id or os.environ.get("LLAMA_CLOUD_PROJECT_ID")
self.organization_id = organization_id or os.environ.get(
"LLAMA_CLOUD_ORGANIZATION_ID"
)
self._async_client: httpx.AsyncClient | None = async_httpx_client
self._files_client = FileClient(
AsyncLlamaCloud(
token=self.api_key,
base_url=self.base_url,
httpx_client=async_httpx_client,
),
project_id=self.project_id,
organization_id=self.organization_id,
)
def _get_default_params(self) -> dict[str, str]:
"""Get default query parameters for API requests"""
params = {}
if self.project_id is not None:
params["project_id"] = self.project_id
if self.organization_id is not None:
params["organization_id"] = self.organization_id
return params
def _get_async_client(self) -> httpx.AsyncClient:
"""Get or create the async httpx client"""
if self._async_client is None:
self._async_client = httpx.AsyncClient(
timeout=httpx.Timeout(60.0),
follow_redirects=True,
)
return self._async_client
def _get_headers(self) -> dict[str, str]:
"""Get common headers for API requests"""
return {
"Authorization": f"Bearer {self.api_key}",
"Content-Type": "application/json",
}
# Sync methods
def upload_file(
self, file_obj: FileInput, file_name: str | None = None
) -> FileUploadResponse:
"""Upload a file to the Files API.
Args:
file_obj: File to upload (path, bytes, or file-like object)
file_name: Optional name for the uploaded filename
Returns:
FileUploadResponse with the uploaded file ID
"""
with augment_async_errors():
return asyncio.run(self.aupload_file(file_obj))
def create_job(
self,
file_id: str,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJob:
"""Create a new spreadsheet parsing job.
Args:
file_id: ID of the uploaded file
config: Parsing configuration
Returns:
SpreadsheetJob with job details
"""
with augment_async_errors():
return asyncio.run(self.acreate_job(file_id, config))
def get_job(
self, job_id: str, include_results_metadata: bool = True
) -> SpreadsheetJobResult:
"""Get the status of a spreadsheet parsing job.
Args:
job_id: ID of the job
include_results_metadata: Whether to include results metadata in the response
Returns:
SpreadsheetJobResult with job status and optionally results
"""
with augment_async_errors():
return asyncio.run(self.aget_job(job_id, include_results_metadata))
def wait_for_completion(self, job_id: str) -> SpreadsheetJobResult:
"""Wait for a job to complete by polling.
Args:
job_id: ID of the job to wait for
Returns:
SpreadsheetJobResult when job is complete
Raises:
SpreadsheetTimeoutError: If job doesn't complete within max_timeout
SpreadsheetJobError: If job fails
"""
with augment_async_errors():
return asyncio.run(self.await_for_completion(job_id))
def download_region_result(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> bytes:
"""Download a region result (either region data or cell metadata).
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
Raw bytes of the parquet file
"""
with augment_async_errors():
return asyncio.run(
self.adownload_region_result(job_id, region_id, result_type)
)
def download_region_as_dataframe(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> "pd.DataFrame":
"""Download a region result as a pandas DataFrame.
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
pandas DataFrame
"""
with augment_async_errors():
return asyncio.run(
self.adownload_region_as_dataframe(job_id, region_id, result_type)
)
def extract_regions(
self,
file_obj: FileInput,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJobResult:
"""High-level method to parse a spreadsheet file.
This method handles the entire workflow:
1. Upload the file
2. Create a parsing job
3. Wait for completion
4. Return results
Args:
file_obj: File to parse (path, bytes, or file-like object)
config: Parsing configuration
Returns:
SpreadsheetJobResult with parsing results
"""
with augment_async_errors():
return asyncio.run(self.aextract_regions(file_obj, config))
# Async methods
async def aupload_file(
self, file_obj: FileInput, file_name: str | None = None
) -> FileUploadResponse:
"""Upload a file to the Files API.
Args:
file_obj: File to upload (path, bytes, or file-like object)
file_name: Optional name for the uploaded filename
Returns:
FileUploadResponse with the uploaded file ID
"""
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
return await self._files_client.upload_content(
file_obj, external_file_id=file_name
)
except Exception as e:
raise SpreadsheetAPIError(f"Failed to upload file: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def acreate_job(
self,
file_id: str,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJob:
"""Create a new spreadsheet parsing job.
Args:
file_id: ID of the uploaded file
config: Parsing configuration
Returns:
SpreadsheetJob with job details
"""
if config is None:
config = SpreadsheetParsingConfig()
elif isinstance(config, dict):
config = SpreadsheetParsingConfig.model_validate(config)
if not isinstance(config, SpreadsheetParsingConfig):
raise ValueError(
"config must be a dict or SpreadsheetParsingConfig instance"
)
payload = {
"file_id": file_id,
"config": config.model_dump(mode="json", exclude_none=True),
}
params = self._get_default_params()
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
response = await client.post(
f"{self.base_url}/api/v1/beta/sheets/jobs",
headers=self._get_headers(),
params=params,
json=payload,
)
response.raise_for_status()
return SpreadsheetJob.model_validate(response.json())
except Exception as e:
raise SpreadsheetAPIError(f"Failed to create job: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def aget_job(
self, job_id: str, include_results_metadata: bool = True
) -> SpreadsheetJobResult:
"""Get the status of a spreadsheet parsing job.
Args:
job_id: ID of the job
include_results_metadata: Whether to include results in the response
Returns:
SpreadsheetJobResult with job status and optionally results
"""
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
params: Dict[str, Any] = {
"include_results": include_results_metadata,
**self._get_default_params(),
}
response = await client.get(
f"{self.base_url}/api/v1/beta/sheets/jobs/{job_id}",
headers=self._get_headers(),
params=params,
)
response.raise_for_status()
return SpreadsheetJobResult.model_validate(response.json())
except Exception as e:
raise SpreadsheetAPIError(f"Failed to get job status: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def await_for_completion(self, job_id: str) -> SpreadsheetJobResult:
"""Wait for a job to complete by polling.
Args:
job_id: ID of the job to wait for
Returns:
SpreadsheetJobResult when job is complete
Raises:
SpreadsheetTimeoutError: If job doesn't complete within max_timeout
SpreadsheetJobError: If job fails
"""
start_time = time.time()
while (time.time() - start_time) < self.max_timeout:
job_result = await self.aget_job(job_id, include_results_metadata=True)
if job_result.status in (
JobStatus.SUCCESS,
JobStatus.PARTIAL_SUCCESS,
JobStatus.ERROR,
JobStatus.FAILURE,
):
if job_result.status in (JobStatus.SUCCESS, JobStatus.PARTIAL_SUCCESS):
return job_result
else:
error_msg = f"Job failed with status: {job_result.status}"
if job_result.errors:
error_msg += f"\nErrors: {', '.join(job_result.errors)}"
raise SpreadsheetJobError(error_msg)
await asyncio.sleep(self.poll_interval)
raise SpreadsheetTimeoutError(
f"Job did not complete within {self.max_timeout} seconds"
)
async def adownload_region_result(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> bytes:
"""Download a region result (either region data or cell metadata).
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
Raw bytes of the parquet file
"""
# Get presigned URL
presigned_response = None
result_type_str = str(result_type)
params = self._get_default_params()
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
response = await client.get(
f"{self.base_url}/api/v1/beta/sheets/jobs/{job_id}/regions/{region_id}/result/{result_type_str}",
headers=self._get_headers(),
params=params,
)
response.raise_for_status()
presigned_response = PresignedUrlResponse.model_validate(
response.json()
)
except Exception as e:
raise SpreadsheetAPIError(f"Failed to get presigned URL: {e}") from e
# Download using presigned URL
if presigned_response is None:
raise SpreadsheetAPIError("Failed to obtain presigned URL.")
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
download_response = await client.get(presigned_response.url)
download_response.raise_for_status()
return download_response.content
except Exception as e:
raise SpreadsheetAPIError(f"Failed to download result: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def adownload_region_as_dataframe(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> "pd.DataFrame":
"""Download a region result as a pandas DataFrame.
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
pandas DataFrame
"""
import pandas as pd
parquet_bytes = await self.adownload_region_result(
job_id, region_id, result_type
)
return pd.read_parquet(io.BytesIO(parquet_bytes))
async def aextract_regions(
self,
file_obj: FileInput,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJobResult:
"""High-level method to parse a spreadsheet file.
This method handles the entire workflow:
1. Upload the file
2. Create a parsing job
3. Wait for completion
4. Return results
Args:
file_obj: File to parse (path, bytes, or file-like object)
config: Parsing configuration
Returns:
SpreadsheetJobResult with parsing results
"""
# Upload file
file_response = await self.aupload_file(file_obj)
# Create job
job = await self.acreate_job(file_response.id, config)
# Wait for completion
return await self.await_for_completion(job.id)
async def aclose(self) -> None:
"""Close all HTTP clients (async)"""
if self._async_client:
await self._async_client.aclose()
async def __aenter__(self) -> "LlamaSheets":
return self
async def __aexit__(self, _exc_type, _exc_val, _exc_tb) -> None: # type: ignore
await self.aclose()

Some files were not shown because too many files have changed in this diff Show More