Compare commits

...

96 Commits

Author SHA1 Message Date
Clelia (Astra) Bertelli 3fd5127252 fix: skip examples for codespell 2025-07-31 18:12:22 +02:00
Clelia (Astra) Bertelli 43e74c36eb chore: correct broken link in readmes 2025-07-31 17:45:19 +02:00
Clelia (Astra) Bertelli e021446901 feat: restructure docs/examples 2025-07-31 17:35:46 +02:00
Clelia (Astra) Bertelli ee4e565604 Example Notebooks (#829)
* fix: add symlink to avoid breaking links

* feat: copy examples
2025-07-31 16:54:12 +02:00
Clelia (Astra) Bertelli 6dbb089f4c delete examples (#830) 2025-07-31 16:53:54 +02:00
Logan Markewich c4b694db8d update symlink 2025-07-31 08:44:30 -06:00
Clelia (Astra) Bertelli 97f428ad06 fix: add symlink to avoid breaking links (#828) 2025-07-31 08:39:44 -06:00
Clelia (Astra) Bertelli ef92ee5408 feat: add ts examples (clean) (#822)
* feat: add ts examples (clean)

* chore: correct title
2025-07-31 11:25:29 +02:00
Logan d094668d03 Update extract.md 2025-07-30 14:58:25 -06:00
Logan 5bb5fc1625 Update parse.md 2025-07-30 14:58:09 -06:00
Logan 1d57e0071d Update parse.md 2025-07-30 14:57:31 -06:00
Logan 2a344c4f5c Update extract.md 2025-07-30 14:56:33 -06:00
Logan ce02559b8d Update README.md (#824) 2025-07-30 14:55:21 -06:00
Harshit Budhiraja e42746e372 docs(readme): update hyperlinks to correct targets (#820) 2025-07-30 14:53:43 -06:00
Clelia (Astra) Bertelli 3149dfd03a fix: no git checks on pnpm publish (#823) 2025-07-30 21:25:23 +02:00
Clelia (Astra) Bertelli e499fdbdab fix: add release to NPM (#819) 2025-07-30 20:55:41 +02:00
Clelia (Astra) Bertelli e57df39248 Merge index into main (#821)
* wip: monorepo changes

* fix ci for the time being

* fix ci for the time being pt2

* wip: first cloud refactoring for ts

* chore: restore original package

* fix: imports, package.json, tsconfig.json, client, reader

* feat: adjustments after local testing

* ci: github actions for typescript

* ci: typescript ci

* ci: nvmrc 🤦

* ci: remove cache 🤦

* ci: actions

* ci: actions (i lost count)

* ci: pnpm run format

* ci: pnpm run format

* chore: migrate llama-parse to uv

* add tests

* remove unneeded readme

* update workflows

* feat: modify py release workflow, adding uv version, bump version for llama-cloud-services to latest

* uv lock

* ci: python tests all tests

* fix: lock file pulling in wrong version of numpy

* feat: add index to llama-cloud-services (#817)

---------

Co-authored-by: Logan Markewich <logan.markewich@live.com>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-07-30 19:46:36 +02:00
Clelia (Astra) Bertelli 09b192b98b Adding TS llama-cloud-services and moving llama-parse to uv (#811)
* wip: monorepo changes

* fix ci for the time being

* fix ci for the time being pt2

* wip: first cloud refactoring for ts

* chore: restore original package

* fix: imports, package.json, tsconfig.json, client, reader

* feat: adjustments after local testing

* ci: github actions for typescript

* ci: typescript ci

* ci: nvmrc 🤦

* ci: remove cache 🤦

* ci: actions

* ci: actions (i lost count)

* ci: pnpm run format

* ci: pnpm run format

* chore: migrate llama-parse to uv

* add tests

* remove unneeded readme

* update workflows

* feat: modify py release workflow, adding uv version, bump version for llama-cloud-services to latest

* uv lock

* ci: python tests all tests

* fix: lock file pulling in wrong version of numpy

---------

Co-authored-by: Logan Markewich <logan.markewich@live.com>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-07-30 17:59:08 +02:00
Adrian Lyjak 13f01a0621 Adding support for page citations, and refactor the confidence into the field metadata (#815) 2025-07-30 10:55:29 -04:00
Javier Torres cf879a1a58 Bump llama-cloud version (#814) 2025-07-28 16:06:31 -05:00
Tuana Çelik fcdf2ab63e Fixes to multimodal report generation (#809) 2025-07-23 16:28:53 -06:00
Adrian Lyjak 083d8109c2 Make versioning a little easier, and fix llama_parse version (#808)
* Make versioning a little easier

* fix up ci
2025-07-21 18:49:07 -04:00
Adrian Lyjak 89cfc8b25f feat: default to _public agent data (#803)
* feat: default to _public agent data
* version bump
2025-07-21 15:58:03 -04:00
Peter Rowlands (변기호) c46e157f92 parse: expose preserve_very_small_text option (#806) 2025-07-21 14:19:15 +09:00
Peter Rowlands (변기호) 05d6026d37 bump version to v0.6.50 (#802) 2025-07-18 18:59:25 +09:00
Peter Rowlands (변기호) 8e98d5c146 parse: expose functionality to get raw job results (#801)
* add LlamaParse.get_result()

* add JobResult.get_text/get_markdown/get_json

* add tests
2025-07-18 18:50:29 +09:00
Adrian Lyjak 3f311c0669 Bump v0.6.49 (#797) 2025-07-16 19:42:09 -04:00
Adrian Lyjak b1a2f9d42b Add new method to fetch the full, non-paginated markdown (#796)
Add new method to fetch the full, non-paginated markdown for proper merge_tables_across_pages_in_markdown support
2025-07-16 19:29:57 -04:00
Neeraj Pradhan 142f55c94c Update to version 0.6.48 (#795)
* Update to version 0.6.48

* pin version

* poetry lock

* adjust warnings

* collect all agents for cleanup
2025-07-16 13:24:44 -07:00
Clelia (Astra) Bertelli 230a110e52 chore: vbump to 0.6.47 and example notebook (#794)
* chore: vbump to 0.6.47 and example notebook

* chore: update llama-parse pyproject.toml
2025-07-16 19:08:44 +02:00
Clelia (Astra) Bertelli 83e2b031cd feat: add table extraction for LlamaParse as CSV files (#793)
* feat: add table extraction for LlamaParse as CSV files

* chore: poetry lock

* chore: add tests

* fix: handle the case where no tables are present

* chore: implement suggestions
2025-07-16 17:08:09 +02:00
Adrian Lyjak 4844e26e5c Improve Agent Data interface, and add file related fields to extracted data for file tracking (#785)
Add file related fields for file tracking. Simplify API
2025-07-09 14:27:24 -04:00
Pierre-Loic Doulcet 70a049af3c merge_tables_across_pages_in_markdown parse parameter (#786)
* merge_tables_across_pages_in_markdown parse parameter

* base.py
2025-07-09 19:03:48 +02:00
Adrian Lyjak dc11776c86 Add nicer hand-written agent data interface (#782)
* Add nicer hand-written agent data interface

* bump to 0.6.44
2025-07-08 17:49:00 -04:00
Logan 2448a42b90 relax pydantic job object (#784) 2025-07-08 12:12:56 -06:00
Neeraj Pradhan c75a900174 Bump up version to 0.6.42 (#783) 2025-07-08 09:16:46 -07:00
Peter Rowlands (변기호) 2fb7adfe0e parse: loosen PageItem.rows type hint (v0.6.41) (#776)
* parse: loosen PageItem.rows type hint

* bump version to 0.6.41
2025-06-30 21:47:40 +09:00
Pierre-Loic Doulcet dc82270724 header footer control in llamaparse (#775) 2025-06-30 16:02:59 +08:00
Neeraj Pradhan d880a48dd0 Bump to version 0.6.39 (#772)
* Bump to version 0.6.39

* lock file update
2025-06-27 16:04:40 -07:00
Logan 7567e8b45e except one more error type (#771) 2025-06-27 10:17:57 -06:00
Neeraj Pradhan 0d59a90151 Relax tenacity version; bump up version to 0.6.37 (#769) 2025-06-25 15:32:20 -07:00
Neeraj Pradhan 98ad550b1a Manage extract agent lifecycle in pytest (#766) 2025-06-24 08:59:38 -07:00
Neeraj Pradhan b58f43ce9f Bump up version to 0.6.36 (#763) 2025-06-23 14:26:05 -07:00
Neeraj Pradhan acf6adcd91 Make job fetching more robust to connection errors (#764) 2025-06-23 13:17:28 -07:00
Neeraj Pradhan daf6576c3c Bump version to 0.6.35 (#762) 2025-06-20 09:33:21 -07:00
Logan 8caa4defa6 fix partition (#758) 2025-06-16 17:37:52 -06:00
Pierre-Loic Doulcet 26918b8de4 add high_res_ocr to the package (#757) 2025-06-16 16:28:23 +08:00
Pierre-Loic Doulcet 6fb5ebe2f9 6.32 warning on unused parameters (#755) 2025-06-12 22:35:48 -06:00
dependabot[bot] c0aa67995b Bump requests from 2.32.3 to 2.32.4 in /llama_parse (#754) 2025-06-10 18:14:44 -06:00
dependabot[bot] 9f841f8328 Bump tornado from 6.4.2 to 6.5.1 in /llama_parse (#753) 2025-06-10 18:14:35 -06:00
dependabot[bot] 99c75eece9 Bump h11 from 0.14.0 to 0.16.0 in /llama_parse (#752) 2025-06-10 18:14:27 -06:00
Logan 57d2586ee3 v0.6.31 (#751) 2025-06-10 17:58:36 -06:00
Jerry Liu 4280a43ec8 add multi-fund analysis notebook (#739) 2025-06-07 11:25:25 -07:00
Neeraj Pradhan 7f1082bbb2 Bump to version 0.6.30 (#748) 2025-06-05 14:34:20 -07:00
Simon Suo 57cfc45804 Directly pass None project_id (#743) 2025-06-05 14:16:54 -07:00
Soumil.Binhani 30e8913875 0.6.29: Standerdize the parsing input format for both .aget_json() and .aload_data() (#745) 2025-06-05 10:58:07 -06:00
Logan 0ce6d4d7a4 more optional types marked (#747) 2025-06-05 10:50:29 -06:00
Peter Rowlands (변기호) 584ba8d48e 0.6.28: fix job result format after partitioning changes (#741)
* parse: fix job result format

* bump to 0.6.28
2025-06-02 15:25:30 -07:00
Peter Rowlands (변기호) 925805ee11 parse: support partitioning files before parsing (#709)
* parse: add utils for handling target_pages

* parse: support partitioning docs into multiple parse jobs

* tests: add tests for partitioned parse

* drop unneeded get_job_result call

* add parse JobFailedException and expected error handling

* bump to 0.6.27
2025-06-02 12:27:58 -07:00
Logan 76fb73c971 v0.6.26 (#740) 2025-06-02 09:59:45 -06:00
Abhik Bhattacharjee 6d19ea9ac0 parse: fix the "model" parameter mismatch between playground and Python client (#737) 2025-06-02 09:35:30 -06:00
Pierre-Loic Doulcet 90431090e9 0.6.25 outlined_table_extraction (#736) 2025-05-30 11:37:21 +02:00
Neeraj Pradhan 6dff35b204 Add notebook for Form 4 extraction (#731)
* Add notebook for Form 4 extraction

* fix comments

* heavier caching; add mermaid diag

* add output directory

* save notebook
2025-05-29 18:31:56 -07:00
Logan e634c7978d v0.6.24 (#732) 2025-05-28 20:11:51 -06:00
Neeraj Pradhan 7a9e99bba2 Bump to version 0.6.23 (#729) 2025-05-20 09:43:06 -07:00
Adrian Lyjak efcdd4405b Pass through verify and timeout config to the extraction agent (#726) 2025-05-17 12:51:16 -07:00
Javier Torres bf3614690f Remove credits from parse metadata (#720) 2025-05-09 16:03:09 -05:00
Logan 7463e00da3 v0.6.22 (#718) 2025-05-08 11:44:41 -06:00
Tuana Çelik cbe9de0c57 Adding example for extracting with citations (#716)
* Adding example for extracting with citations

* removing TOC and installation output
2025-05-06 23:32:17 +02:00
Logan a023507d42 even more optional (#711) 2025-05-01 15:52:38 -06:00
Peter Rowlands (변기호) e48f544ddc parse: fix num_workers/parse job batching (#708) 2025-05-01 09:30:35 -06:00
Logan 4aa7ad5642 v0.6.20 (#707) 2025-04-29 08:53:55 -06:00
Sacha Bron c39cdbcd01 v0.6.19 (#706) 2025-04-29 12:28:21 +02:00
Pierre-Loic Doulcet 71eaa8bcc6 add auto_mode_configuration_jon for llamaParse (#704) 2025-04-29 12:23:03 +02:00
Pierre-Loic Doulcet 1e1cbdfc79 add support for presets (#703) 2025-04-29 11:54:54 +08:00
Logan cc8af4a43a make original height + width optional in the parse result (#702) 2025-04-27 18:31:35 -06:00
dependabot[bot] 43fbd48ab8 Bump actions/setup-python from 4 to 5 (#701)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 4 to 5.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v4...v5)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-04-27 13:08:44 -06:00
dependabot[bot] 5ec66e9452 Bump actions/checkout from 3 to 4 (#700)
Bumps [actions/checkout](https://github.com/actions/checkout) from 3 to 4.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v3...v4)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-04-27 13:08:31 -06:00
Scott Brenner 211521c82e Dependabot configuration to update actions in workflow (#698) 2025-04-27 12:52:11 -06:00
Scott Brenner 4ddaab1efb Refactor CodeQL workflow (#699)
* Refactor CodeQL workflow

* Update .github/workflows/codeql.yml
2025-04-27 12:51:56 -06:00
Neeraj Pradhan 53e5ce2e83 Bump to v0.6.16 (#697) 2025-04-25 14:39:52 -07:00
Neeraj Pradhan 9f4bd1cb64 Update to latest version of llama-cloud (#696)
update to latest version of llama-cloud
2025-04-25 14:14:49 -07:00
Logan 456863752b small enum nit for FailedPageMode (#693) 2025-04-23 21:34:26 -06:00
Pierre-Loic Doulcet c2dc34bbd6 Page error parameters (#691) 2025-04-23 20:47:57 -06:00
Logan fcabb04baf skip llama-report tests in cicd (#692)
* skip llama-report tests in cicd

* skip llama-report tests in cicd
2025-04-23 20:47:00 -06:00
Sacha Bron 8e7c32d3d6 Add markdown_table_multiline_header_separator support (#683)
* Add markdown_table_multiline_header_separator support

* Lint
2025-04-15 17:39:46 +02:00
Neeraj Pradhan 7e3013d914 Use unique filename to avoid db collisions later (#682)
* Use unique filename to avoid db collisions later

* add xfail marker to test_create_and_delete_report
2025-04-11 11:03:15 -07:00
Logan 4a664c33d2 parse readme nits (#681) 2025-04-10 19:25:06 -06:00
Logan 6d049ee2e4 v0.6.12 (#680) 2025-04-10 19:18:49 -06:00
Logan fa73e73664 new result object (#650) 2025-04-10 19:17:23 -06:00
Neeraj Pradhan bf67ee6056 Update docs for LlamaExtract (#679) 2025-04-10 12:16:32 -07:00
Neeraj Pradhan a1abef2ee9 Bump version to v0.6.11 (#678) 2025-04-10 11:23:06 -07:00
Neeraj Pradhan a753e01d3c Support text as input directly in the SDK (#676) 2025-04-09 21:40:56 -07:00
Logan 9b15065b24 v0.6.10 (#677) 2025-04-09 19:30:59 -06:00
Pierre-Loic Doulcet 6e4150537c Add compact_markdown_table parameter (#675) 2025-04-09 19:19:19 -06:00
Neeraj Pradhan 233d715a14 Better connection management on llamaextract client (#674) 2025-04-09 14:26:52 -07:00
191 changed files with 125137 additions and 12177 deletions
+11
View File
@@ -0,0 +1,11 @@
# Please see the documentation for all configuration options:
# https://docs.github.com/github/administering-a-repository/configuration-options-for-dependency-updates
# and
# https://docs.github.com/code-security/dependabot/dependabot-version-updates/configuration-options-for-the-dependabot.yml-file
version: 2
updates:
- package-ecosystem: "github-actions"
directory: "/"
schedule:
interval: "weekly"
-48
View File
@@ -1,48 +0,0 @@
name: Build Package
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install
- name: Ensure lock works
shell: bash
run: poetry lock
- name: Build
shell: bash
run: poetry build
- name: Test installing built package
shell: bash
run: python -m pip install .
- name: Test import
shell: bash
working-directory: ${{ vars.RUNNER_TEMP }}
run: python -c "import llama_cloud_services"
+50
View File
@@ -0,0 +1,50 @@
name: Build Package - Python
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Test import
working-directory: py
run: uv run -- python -c "import llama_cloud_services"
+28
View File
@@ -0,0 +1,28 @@
name: Build Package - TypeScript
on: [pull_request]
jobs:
pre_release:
name: Pre Release
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Build
working-directory: ts/llama_cloud_services/
run: pnpm run build
+8 -48
View File
@@ -1,14 +1,3 @@
# For most projects, this workflow file will not need changing; you simply need
# to commit it to your repository.
#
# You may wish to alter this file to override the set of languages analyzed,
# or to provide custom queries or build logic.
#
# ******** NOTE ********
# We have attempted to detect the languages in your repository. Please check
# the `language` matrix defined below to confirm you have the correct set of
# supported CodeQL languages.
#
name: "CodeQL"
on:
@@ -28,54 +17,25 @@ jobs:
# - https://gh.io/supported-runners-and-hardware-resources
# - https://gh.io/using-larger-runners
# Consider using larger runners for possible analysis time improvements.
runs-on: ${{ (matrix.language == 'swift' && 'macos-latest') || 'ubuntu-latest' }}
timeout-minutes: ${{ (matrix.language == 'swift' && 120) || 360 }}
runs-on: "ubuntu-latest"
timeout-minutes: 360
permissions:
actions: read
contents: read
security-events: write
strategy:
fail-fast: false
matrix:
language: ["python"]
# CodeQL supports [ 'cpp', 'csharp', 'go', 'java', 'javascript', 'python', 'ruby', 'swift' ]
# Use only 'java' to analyze code written in Java, Kotlin or both
# Use only 'javascript' to analyze code written in JavaScript, TypeScript or both
# Learn more about CodeQL language support at https://aka.ms/codeql-docs/language-support
steps:
- name: Checkout repository
uses: actions/checkout@v3
uses: actions/checkout@v4
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v2
uses: github/codeql-action/init@v3
with:
languages: ${{ matrix.language }}
# If you wish to specify custom queries, you can do so here or in a config file.
# By default, queries listed here will override any specified in a config file.
# Prefix the list here with "+" to use these queries and those in the config file.
# For more details on CodeQL's query packs, refer to: https://docs.github.com/en/code-security/code-scanning/automatically-scanning-your-code-for-vulnerabilities-and-errors/configuring-code-scanning#using-queries-in-ql-packs
# queries: security-extended,security-and-quality
# Autobuild attempts to build any compiled languages (C/C++, C#, Go, Java, or Swift).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@v2
# ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
# If the Autobuild fails above, remove it and uncomment the following three lines.
# modify them (or add more) to build your code if your project, please refer to the EXAMPLE below for guidance.
# - run: |
# echo "Run, Build Application using script"
# ./location_of_script_within_repo/buildscript.sh
languages: python
dependency-caching: true
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v2
uses: github/codeql-action/analyze@v3
with:
category: "/language:${{matrix.language}}"
category: "/language:python"
-37
View File
@@ -1,37 +0,0 @@
name: Linting
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install pre-commit
shell: bash
run: poetry run pip install pre-commit
- name: Run linter
shell: bash
run: poetry run make lint
+35
View File
@@ -0,0 +1,35 @@
name: Lint - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }}
- name: Run linter
shell: bash
working-directory: py
run: uv run -- pre-commit run -a
+36
View File
@@ -0,0 +1,36 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run lint
working-directory: ts/llama_cloud_services/
run: pnpm run lint
- name: Run Prettier
working-directory: ts/llama_cloud_services/
run: pnpm run format
-75
View File
@@ -1,75 +0,0 @@
name: Publish llama-parse to PyPI / GitHub
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
POETRY_VERSION: "1.6.1"
PYTHON_VERSION: "3.9"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ env.PYTHON_VERSION }}
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: pip install -e .
- name: Build and publish llama-cloud-services
uses: JRubics/poetry-publish@v2.1
with:
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
poetry_install_options: "--without dev"
- name: Build and publish llama-parse
uses: JRubics/poetry-publish@v2.1
with:
working_directory: "llama_parse"
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
poetry_install_options: "--without dev"
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }}
draft: false
prerelease: false
- name: Get Asset name
run: |
export PKG=$(ls dist/ | grep tar)
set -- $PKG
echo "name=$1" >> $GITHUB_ENV
- name: Upload Release Asset (sdist) to GitHub
id: upload-release-asset
uses: actions/upload-release-asset@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
upload_url: ${{ steps.create_release.outputs.upload_url }}
asset_path: dist/${{ env.name }}
asset_name: ${{ env.name }}
asset_content_type: application/zip
+66
View File
@@ -0,0 +1,66 @@
name: Publish Release - Python
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
UV_VERSION: "0.7.20"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Publish package
shell: bash
working-directory: py
run: uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Build and publish llama-parse
working-directory: py/llama_parse/
run: |
uv build
uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }} - LlamaCloud Services PY
artifacts: "py/**/dist/*"
generateReleaseNotes: true
draft: false
prerelease: false
+51
View File
@@ -0,0 +1,51 @@
name: Publish Release - TypeScript
on:
push:
tags:
- "llama-cloud-services@*"
jobs:
build-and-publish:
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services
run: pnpm install --no-frozen-lockfile
- name: Build tarball
run: |
pnpm pack
working-directory: ts/llama_cloud_services
- name: Setup npm authentication
run: echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc
env:
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Release
working-directory: ts/llama_cloud_services
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
run: pnpm publish --access public --no-git-checks
- name: Create release
uses: ncipollo/release-action@v1
with:
artifacts: "ts/llama_cloud_services/llama-cloud-services*.tgz"
name: Release ${{ github.ref }} - LlamaCloud Services TS
bodyFile: "ts/llama_cloud_services/CHANGELOG.md"
token: ${{ secrets.GITHUB_TOKEN }}
+39
View File
@@ -0,0 +1,39 @@
name: Test - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }} && uv python pin ${{ matrix.python-version }}
- name: Run Tests
working-directory: py
run: uv run -- pytest tests/**/test_*.py
- name: Remove virtual environment
working-directory: py
run: rm -rf .venv/
+34
View File
@@ -0,0 +1,34 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run tests
working-directory: ts/llama_cloud_services/
run: pnpm test --run
-40
View File
@@ -1,40 +0,0 @@
name: Unit Testing
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: 0
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install --with dev
- name: Run testing
env:
CI: true
shell: bash
run: poetry run pytest tests
+4
View File
@@ -5,3 +5,7 @@ __pycache__/
.idea
.env*
.ipynb_checkpoints*
*_cache/
node_modules/
.turbo/
dist/
+6 -6
View File
@@ -21,19 +21,19 @@ repos:
hooks:
- id: ruff
args: [--fix, --exit-non-zero-on-fix]
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
- repo: https://github.com/psf/black-pre-commit-mirror
rev: 23.10.1
hooks:
- id: black-jupyter
name: black-src
alias: black
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.0.1
hooks:
- id: mypy
exclude: ^tests/
exclude: ^py/tests/
additional_dependencies:
[
"types-requests",
@@ -63,13 +63,13 @@ repos:
rev: v3.0.3
hooks:
- id: prettier
exclude: poetry.lock
exclude: uv.lock
- repo: https://github.com/codespell-project/codespell
rev: v2.2.6
hooks:
- id: codespell
additional_dependencies: [tomli]
exclude: ^(poetry.lock|examples)
exclude: ^(uv.lock|examples|ts)
args:
[
"--ignore-words-list",
@@ -84,6 +84,6 @@ repos:
rev: v0.23.1
hooks:
- id: toml-sort-fix
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
exclude: .github/ISSUE_TEMPLATE
+33
View File
@@ -0,0 +1,33 @@
# Python
## Installation
This project uses uv. Create a virtual environment, and run `uv sync`
## Versioning (Maintainers only)
Before merging your changes, make sure to bump the versions.
Make a version bump to `pyproject.toml`. If the underlying dependency on the llamacloud platform OpenAPI
sdk needs bumping, make sure to bring that in as well. If updating dependencies, run `uv lock`.
The legacy `llama_parse` package re-exports some of `llama_cloud_services` in the old namespace. The
versions need to be kept consistent to sidecar it with `llama_cloud_services`. Bump it's version in `llama_parse/pyproject.toml`, and also bump it's dependency version of `llama-cloud-services` to match.
**Note**: Don't worry about updating the `llama_parse/poetry.lock` file when bumping versions. The GitHub action will automatically run `poetry lock` for the llama_parse package during the build process (though it doesn't commit the updated lockfile back to the repo).
You can also do this with `./scripts/version-bump.py set 0.x.x` if you have `uv` installed.
Once the change is merged, push a tag `git tag -a v0.x.x -m 0.x.x` and `git push origin 0.x.x`.
This tagging step can be done with `./scripts/version-bump tag`.
# Typescript
## Installation
...
## Versioning
...
+19 -3
View File
@@ -10,7 +10,8 @@ This includes:
- [LlamaParse](./parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](./report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract (beta/invite-only)](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaExtract](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](./index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
## Getting Started
@@ -25,18 +26,27 @@ Then, get your API key from [LlamaCloud](https://cloud.llamaindex.ai/).
Then, you can use the services in your code:
```python
from llama_cloud_services import LlamaParse, LlamaReport, LlamaExtract
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
)
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
)
```
See the quickstart guides for each service for more information:
- [LlamaParse](./parse.md)
- [LlamaReport (beta/invite-only)](./report.md)
- [LlamaExtract (beta/invite-only)](./extract.md)
- [LlamaExtract](./extract.md)
- [LlamaCloud Index](./index.md)
## Switch to EU SaaS 🇪🇺
@@ -55,6 +65,12 @@ from llama_cloud_services import (
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
project_name="default",
api_key="YOUR_API_KEY",
base_url=EU_BASE_URL,
)
```
## Documentation
+11
View File
@@ -0,0 +1,11 @@
# LlamaCloud Services Examples
In this folder you will find several python notebooks and two end-to-end typescript applications that contain examples regarding:
- [LlamaParse - Python](./parse/)
- [LlamaParse - TypeScript](./parse-ts/)
- [LlamaExtract - Python](./extract/)
- [LlamaReport - Python](./report/)
- [LlamaCloud Index - TypeScript](./index-ts/)
Follow the instructions of each notebook/application to get started!
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 MiB

@@ -0,0 +1 @@
sec_form_4_dump.json
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 202 KiB

@@ -0,0 +1,440 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Extract Data from Financial Reports - with Citations and Reasoning\n",
"\n",
"Given complex files like financial reports, contracts, invoices etc, Llama Extract allows you to make use of an LLM to extract the information relevant to you, in a structured format.\n",
"\n",
"In this example, we'll be using [LlamaExtract](https://docs.cloud.llamaindex.ai/llamaextract/getting_started?utm_campaign=extract&utm_medium=recipe) to extract structured data from an SEC filing (specifically, the filing by Nvidia for fiscal year 2025).\n",
"\n",
"On top of simple data extraction, we'll ask our extraction agent to provide citations and reasoning for each extracted field. This allows us to:\n",
"- Confirm the accuracy of the extracted field\n",
"- Understand the reasoning behind why the LLM extracted a given piece of information\n",
"- This last point allows us an opportunity to adjust the system prompt or field descriptions and improve on results where needed.\n",
"\n",
"\n",
"The example we go through below is also replicable within Llama Cloud as well, where you will also be able to pick between a number of pre-defined schemas, instead of building your own."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-cloud-services"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Connect to Llama Cloud\n",
"\n",
"To get started, make sure you provide your [Llama Cloud](https://cloud.llamaindex.ai?utm_campaign=extract&utm_medium=recipe) API key."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Enter your Llama Cloud API Key: ··········\n"
]
}
],
"source": [
"import os\n",
"from getpass import getpass\n",
"\n",
"if \"LLAMA_CLOUD_API_KEY\" not in os.environ:\n",
" os.environ[\"LLAMA_CLOUD_API_KEY\"] = getpass(\"Enter your Llama Cloud API Key: \")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Extract Data with Llama Extract Agent"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"No project_id provided, fetching default project.\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaExtract\n",
"\n",
"# Optionally, provide your project id, if not, it will use the 'Default' project\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Provide Your Custom Schema\n",
"\n",
"When using LlamaExtract via the API, you provide your own schema that describes what you want extracted from files and data provided to your agent. Here, we are essentially building an SEC filings extraction agent."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from enum import Enum\n",
"\n",
"\n",
"class FilingType(str, Enum):\n",
" ten_k = \"10 K\"\n",
" ten_q = \"10-Q\"\n",
" ten_ka = \"10-K/A\"\n",
" ten_qa = \"10-Q/A\"\n",
"\n",
"\n",
"class FinancialReport(BaseModel):\n",
" company_name: str = Field(description=\"The name of the company\")\n",
" description: str = Field(\n",
" description=\"Short description of the filing and what it contains\"\n",
" )\n",
" filing_type: FilingType = Field(description=\"Type of SEC filing\")\n",
" filing_date: str = Field(description=\"Date when filing was submitted to SEC\")\n",
" fiscal_year: int = Field(description=\"Fiscal year\")\n",
" unit: str = Field(\n",
" description=\"Unit of financial figures (thousands, millions, etc.)\"\n",
" )\n",
" revenue: int = Field(description=\"Total revenue for period\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Set Up Citations and Reasoning\n",
"\n",
"Optionally, we can set the `ExtractConfig` to extract citations for each field the agent extracts. These cications will cite the specific pages and sections of the file from which a given field was extractedd.\n",
"\n",
"By setting `use_reasoning` to True, we als ask the agent to do an additional reasoning step, explaining why a given field was extracted."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud.types import ExtractConfig, ExtractMode\n",
"\n",
"config = ExtractConfig(\n",
" use_reasoning=True, cite_sources=True, extraction_mode=ExtractMode.MULTIMODAL\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:127: ExperimentalWarning: `use_reasoning` is an experimental feature. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n",
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:133: ExperimentalWarning: `cite_sources` is an experimental feature. This may greatly increase the size of the response, and slow down the extraction. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n"
]
}
],
"source": [
"agent = llama_extract.create_agent(\n",
" name=\"filing-parser\", data_schema=FinancialReport, config=config\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Demo Time - Download a PDF and Extract Data with Citations"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"PDF downloaded successfully.\n"
]
}
],
"source": [
"import requests\n",
"\n",
"url = \"https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf\"\n",
"\n",
"response = requests.get(url)\n",
"\n",
"if response.status_code == 200:\n",
" with open(\"/content/nvda_10k.pdf\", \"wb\") as f:\n",
" f.write(response.content)\n",
" print(\"PDF downloaded successfully.\")\n",
"else:\n",
" print(f\"Failed to download. Status code: {response.status_code}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|██████████| 1/1 [00:00<00:00, 1.83it/s]\n",
"Creating extraction jobs: 100%|██████████| 1/1 [00:00<00:00, 4.38it/s]\n",
"Extracting files: 100%|██████████| 1/1 [02:03<00:00, 123.40s/it]\n"
]
}
],
"source": [
"filing_info = agent.extract(\"/content/nvda_10k.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'company_name': 'NVIDIA Corporation',\n",
" 'description': \"The filing provides a detailed overview of NVIDIA's business as a full-stack computing infrastructure company, discusses various technologies including digital avatars and autonomous vehicles, outlines numerous risk factors affecting operations such as supply chain issues and geopolitical tensions, and describes employee stock purchase plans and related compliance requirements.\",\n",
" 'filing_type': '10 K',\n",
" 'filing_date': 'February 26, 2025',\n",
" 'fiscal_year': 2025,\n",
" 'unit': 'millions',\n",
" 'revenue': 130497}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Inspect Citations and Reasoning"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'field_metadata': {'company_name': {'reasoning': 'VERBATIM EXTRACTION',\n",
" 'citation': [{'page': 1, 'matching_text': 'NVIDIA CORPORATION'},\n",
" {'page': 2, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 3,\n",
" 'matching_text': 'All references to \"NVIDIA,\" \"we,\" \"us,\" \"our,\" or the \"Company\" mean NVIDIA Corporation and its subsidiaries.'},\n",
" {'page': 35,\n",
" 'matching_text': 'Comparison of 5 Year Cumulative Total Return* Among NVIDIA Corporation'},\n",
" {'page': 49,\n",
" 'matching_text': 'To the Board of Directors and Shareholders of NVIDIA Corporation'},\n",
" {'page': 90, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 119,\n",
" 'matching_text': '*\"Company\"* means NVIDIA Corporation, a Delaware corporation.'},\n",
" {'page': 126,\n",
" 'matching_text': 'Annual Report on Form 10-K of NVIDIA Corporation'}]},\n",
" 'filing_type': {'reasoning': \"VERBATIM EXTRACTION from multiple sources confirming the filing type as '10 K'.\",\n",
" 'citation': [{'page': 1, 'matching_text': 'FORM 10-K'},\n",
" {'page': 2, 'matching_text': 'Item 16. | Form 10-K Summary'},\n",
" {'page': 3,\n",
" 'matching_text': 'This Annual Report on Form 10-K contains forward-looking statements...'},\n",
" {'page': 13, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 15, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 32,\n",
" 'matching_text': 'Annual Report on Form 10-K, which information is hereby incorporated by reference.'},\n",
" {'page': 36, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 43,\n",
" 'matching_text': 'Annual Report on Form 10-K for additional information'},\n",
" {'page': 45, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 46, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 62, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 83,\n",
" 'matching_text': 'Restated Certificate of Incorporation | 10-K'},\n",
" {'page': 84, 'matching_text': 'Item 16. Form 10-K Summary'},\n",
" {'page': 126, 'matching_text': 'which appears in this Form 10-K'},\n",
" {'page': 127, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 128, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 129, 'matching_text': \"The Company's Annual Report on Form 10-K\"},\n",
" {'page': 130,\n",
" 'matching_text': \"The Company's Annual Report on Form 10-K for the year ended January 26, 2025\"}]},\n",
" 'fiscal_year': {'reasoning': 'The fiscal year ended January 26, 2025, indicates the fiscal year is 2025. Additionally, multiple references throughout the text confirm the fiscal year 2025 in various contexts.',\n",
" 'citation': [{'page': 1,\n",
" 'matching_text': 'For the fiscal year ended January 26, 2025'},\n",
" {'page': 6,\n",
" 'matching_text': 'In fiscal year 2025, we launched the NVIDIA Blackwell architecture'},\n",
" {'page': 12, 'matching_text': 'fiscal year 2025'},\n",
" {'page': 17,\n",
" 'matching_text': 'our gross margins in the second quarter of fiscal year 2025 were negatively impacted'},\n",
" {'page': 20,\n",
" 'matching_text': 'we generated 53% of our revenue in fiscal year 2025 from sales outside the United States.'},\n",
" {'page': 23,\n",
" 'matching_text': 'For fiscal year 2025, an indirect customer which primarily purchases our products through system integrators...'},\n",
" {'page': 33,\n",
" 'matching_text': 'In fiscal year 2025, we repurchased 310 million shares of our common stock for $34.0 billion.'},\n",
" {'page': 37,\n",
" 'matching_text': 'Our Data Center revenue in China grew in fiscal year 2025.'},\n",
" {'page': 44,\n",
" 'matching_text': 'Cash provided by operating activities increased in fiscal year 2025 compared to fiscal year 2024'},\n",
" {'page': 57,\n",
" 'matching_text': 'Fiscal years 2025, 2024 and 2023 were all 52-week years.'},\n",
" {'page': 65,\n",
" 'matching_text': 'Beginning in the second quarter of fiscal year 2025'},\n",
" {'page': 69, 'matching_text': 'In the fourth quarter of fiscal year 2025'},\n",
" {'page': 78,\n",
" 'matching_text': 'Depreciation and amortization expense attributable to our Compute and Networking segment for fiscal years 2025'},\n",
" {'page': 129, 'matching_text': 'for the year ended January 26, 2025'}]},\n",
" 'description': {'reasoning': 'The extracted data combines multiple descriptions from the source text, ensuring no duplication while maintaining the order and context of the information. Each section of the filing is summarized to reflect the key points without losing the essence of the original text.',\n",
" 'citation': [{'page': 4,\n",
" 'matching_text': 'NVIDIA is now a full-stack computing infrastructure company with data-center-scale offerings that are reshaping industry.'},\n",
" {'page': 8,\n",
" 'matching_text': 'a suite of technologies that help developers bring digital avatars to life with generative Al...autonomous vehicles, or AV, and electric vehicles, or EV, is revolutionizing the transportation industry...Our worldwide sales and marketing strategy is key to achieving our objective of providing markets with our high-performance and efficient computing platforms and software.'},\n",
" {'page': 14, 'matching_text': 'Risk Factors Summary'},\n",
" {'page': 16,\n",
" 'matching_text': 'Risks Related to Demand, Supply, and Manufacturing\\n\\nLong manufacturing lead times and uncertain supply and component availability...'},\n",
" {'page': 18,\n",
" 'matching_text': 'cryptocurrency mining, on demand for our products. Volatility in the cryptocurrency market, including new compute technologies...'},\n",
" {'page': 21,\n",
" 'matching_text': 'supply-chain attacks or other business disruptions. We cannot guarantee that third parties and infrastructure in our supply chain...'},\n",
" {'page': 22,\n",
" 'matching_text': 'We are monitoring the impact of the geopolitical conflict in and around Israel on our operations... Climate change may have a long-term impact on our business.'},\n",
" {'page': 25,\n",
" 'matching_text': 'We are subject to complex laws, rules, regulations, and political and other actions, including restrictions on the export of our products, which may adversely impact our business.'},\n",
" {'page': 28,\n",
" 'matching_text': 'Our competitive position has been harmed by the existing export controls, and our competitive position and future results may be further harmed'},\n",
" {'page': 29,\n",
" 'matching_text': 'restrictions imposed by the Chinese government on the duration of gaming activities and access to games may adversely affect our Gaming revenue'},\n",
" {'page': 29,\n",
" 'matching_text': 'our business depends on our ability to receive consistent and reliable supply from our overseas partners, especially in Taiwan and South Korea'},\n",
" {'page': 29,\n",
" 'matching_text': 'Increased scrutiny from shareholders, regulators and others regarding our corporate sustainability practices could result in additional costs'},\n",
" {'page': 29,\n",
" 'matching_text': 'Concerns relating to the responsible use of new and evolving technologies, such as Al, in our products and services may result in reputational or financial harm'},\n",
" {'page': 31,\n",
" 'matching_text': 'Data protection laws around the world are quickly changing and may be interpreted and applied in an increasingly stringent fashion...'}]},\n",
" 'filing_date': {'reasoning': 'The filing date is consistently mentioned as February 26, 2025 across multiple entries, making it the most reliable date for the filing.',\n",
" 'citation': [{'page': 51, 'matching_text': 'February 26, 2025'},\n",
" {'page': 86, 'matching_text': 'on February 26, 2025.'},\n",
" {'page': 87, 'matching_text': 'February 26, 2025'},\n",
" {'page': 126, 'matching_text': 'our report dated February 26, 2025'},\n",
" {'page': 127, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 128, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 129, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 130, 'matching_text': 'Date: February 26, 2025'}]},\n",
" 'unit': {'reasoning': \"The unit of financial figures is explicitly mentioned multiple times in the text as 'millions', including in table headers and notes. This is confirmed by various citations from pages 38, 42, 43, 52, 53, 54, 56, 65, 71, 72, 73, 75, 77, 79, 80, and 82.\",\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': '($ in millions, except per share data)'},\n",
" {'page': 42, 'matching_text': '($ in millions)'},\n",
" {'page': 43, 'matching_text': '($ in millions)'},\n",
" {'page': 52, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 53,\n",
" 'matching_text': 'Consolidated Statements of Comprehensive Income (In millions)'},\n",
" {'page': 54,\n",
" 'matching_text': 'Consolidated Balance Sheets (In millions, except par value)'},\n",
" {'page': 55, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 56,\n",
" 'matching_text': 'Consolidated Statements of Cash Flows (In millions)'},\n",
" {'page': 65,\n",
" 'matching_text': 'Year Ended<br/>Jan 26, 2025<br/>(In millions, except per share data)'},\n",
" {'page': 71, 'matching_text': '(In millions) | (In millions)'},\n",
" {'page': 72, 'matching_text': '(In millions)'}]},\n",
" 'revenue': {'reasoning': 'The total revenue for fiscal year 2025 is extracted from multiple sources within the text, all confirming the same figure of $130,497 million. The revenue recognized for fiscal year 2025 is also noted as $4,607 million, which is a separate figure. However, the primary focus is on the total revenue figure, which is consistently cited.',\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': 'Revenue for fiscal year 2025 was $130.5 billion'},\n",
" {'page': 41,\n",
" 'matching_text': 'Total | $ 130,497 | $ | 60,922'},\n",
" {'page': 52, 'matching_text': 'Revenue | $ 130,497'},\n",
" {'page': 78,\n",
" 'matching_text': 'Revenue | $ 116,193 | $ 14,304 | $ - | $ 130,497'},\n",
" {'page': 79, 'matching_text': 'Total revenue | $ 130,497'},\n",
" {'page': 80, 'matching_text': 'Total revenue | $ 130,497'}]}},\n",
" 'usage': {'num_pages_extracted': 130,\n",
" 'num_document_tokens': 105932,\n",
" 'num_output_tokens': 31306}}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.extraction_metadata"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## What's Next?\n",
"\n",
"In this example, we built an Extraction Agent that is capable of citing it's sources from the document it's extracting data from, and reasoning about its reponse. To further customize and improve on the results, you can also try to customize the `system_prompt` in the `ExtractConfig`.\n",
"\n",
"#### Learn More\n",
"\n",
"- [LlamaExtract Documentation](https://docs.cloud.llamaindex.ai/llamaextract/getting_started)\n",
"- [Example Notebooks](https://github.com/run-llama/llama_cloud_services/tree/main/examples/extract)"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
File diff suppressed because it is too large Load Diff
+131
View File
@@ -0,0 +1,131 @@
# LlamaCloud Index Demo
A TypeScript demo application showcasing the power of **LlamaCloud Index** - a fully automated document ingestion and retrieval serviced offered within [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to ask questions, retrieve relevant contextual information and generate AI-powered responses using OpenAI's GPT models.
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [Development Mode](#development-mode)
- [Build the Project](#build-the-project)
- [Code Quality](#code-quality)
- [Quick Commands Reference](#quick-commands-reference)
- [How It Works](#how-it-works)
- [API Dependencies](#api-dependencies)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 🤖 **RAG**: Simple-yet-effective Retrieval Augmented Generation pipeline built on top of LlamaCloud Index and OpenAI
- 🎨 **Beautiful CLI**: Styled console interface with colors and ASCII art
-**Fast Development**: Hot reload support with watch mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 18 or higher)
- pnpm package manager
- OpenAI API key
- LlamaCloud API key
- An existing LlamaCloud Index pipeline
## Installation
1. Clone the repository:
```bash
git clone <repository-url>
cd llamaparse-demo
```
2. Install dependencies:
```bash
pnpm install
```
3. Set up your environment variables:
```bash
export OPENAI_API_KEY="your-openai-api-key"
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
export PIPELINE_NAME="your-pipeline-name"
```
4. Or write them into a `.env` file:
```env
OPENAI_API_KEY="your-openai-api-key"
LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
PIPELINE_NAME="your-pipeline-name"
```
## Usage
### Start the Demo
```bash
pnpm run start
```
The application will display a welcome screen and prompt you to start chatting!
### Development Mode
For development with hot reload:
```bash
pnpm run dev
```
### Build the Project
```bash
pnpm run build
```
### Code Quality
Format code:
```bash
pnpm run format
```
Lint code:
```bash
pnpm run lint
```
## How It Works
1. **Message Input**: Enter a message
2. **Retrieval**: Several nodes are retrieved from the LlamaCloud index you specified
3. **AI Response Generation**: The retrieved information is passed on to the AI model, along with its relevance score, and a reply to your original message is generated starting from that.
4. **Results**: View the AI-generated summary in your terminal
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 18+ and have all dependencies installed
2. **API Key Issues**: Verify your OpenAI and LlamaCloud API keys are correctly set
## License
MIT License - see the [LICENSE](../../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `pnpm format` and `pnpm lint`
5. Submit a pull request
+15
View File
@@ -0,0 +1,15 @@
import js from "@eslint/js";
import globals from "globals";
import tseslint from "typescript-eslint";
import { defineConfig } from "eslint/config";
export default defineConfig([
{
files: ["**/*.{js,mjs,cjs,ts,mts,cts}"],
plugins: { js },
extends: ["js/recommended"],
languageOptions: { globals: globals.browser },
},
{ files: ["**/*.js"], languageOptions: { sourceType: "script" } },
tseslint.configs.recommended,
]);
+48
View File
@@ -0,0 +1,48 @@
{
"name": "llama-chat",
"version": "0.1.0",
"description": "Demo for LlamaCloud Index in TypeScript",
"type": "module",
"main": "index.js",
"scripts": {
"test": "echo \"There are no tests\"",
"start": "pnpm exec tsx src/index.ts",
"lint": "eslint ./src/",
"format": "prettier --write ./src/",
"build": "tsc",
"dev": "pnpm exec tsx --watch src/index.ts"
},
"keywords": [
"ai",
"rag",
"retrieval",
"pipeline",
"llms",
"chatbot"
],
"author": "LlamaIndex",
"license": "MIT",
"packageManager": "pnpm@10.12.4",
"devDependencies": {
"@eslint/js": "^9.32.0",
"@types/figlet": "^1.7.0",
"@types/node": "^24.1.0",
"@typescript-eslint/eslint-plugin": "^8.38.0",
"@typescript-eslint/parser": "^8.38.0",
"eslint": "^9.32.0",
"globals": "^16.3.0",
"jiti": "^2.5.1",
"prettier": "^3.6.2",
"typescript": "^5.8.3",
"typescript-eslint": "^8.38.0"
},
"dependencies": {
"@ai-sdk/openai": "^1.3.23",
"ai": "^4.3.19",
"consola": "^3.4.2",
"dotenv": "^17.2.1",
"figlet": "^1.8.2",
"llama-cloud-services": "link:../../../ts/llama_cloud_services",
"picocolors": "^1.1.1"
}
}
+1770
View File
File diff suppressed because it is too large Load Diff
+48
View File
@@ -0,0 +1,48 @@
import { LlamaCloudIndex } from "llama-cloud-services";
import { logger } from "./logger";
import pc from "picocolors";
import {
consoleInput,
retrievalAugmentedGeneration,
renderLogo,
} from "./utils";
import dotenv from "dotenv";
dotenv.config();
export async function main(): Promise<number> {
const index = new LlamaCloudIndex({
name: process.env.PIPELINE_NAME as string,
projectName: "Default",
apiKey: process.env.LLAMA_CLOUD_API_KEY, // can provide API-key in the constructor or in the env
});
const retriever = index.asRetriever({
similarityTopK: 5,
});
await renderLogo();
logger.log(
`Welcome to ${pc.bold(
pc.magentaBright("✨LlamaChat✨"),
)}, our demo for ${pc.bold(pc.green("Index🦙"))}, a ${pc.bold(
pc.cyan("LlamaCloud☁️"),
)} (https://cloud.llamaindex.ai) product!.\nType a question below, and you will get an answer!👇\nIf you wish to exit, just type ${pc.bold(
pc.gray("quit"),
)}.\n`,
);
while (true) {
const userInput = await consoleInput();
if (userInput.toLowerCase() == "quit") {
break;
}
try {
const nodes = await retriever.retrieve(userInput);
const summary = await retrievalAugmentedGeneration(nodes, userInput);
logger.log(`${pc.bold(pc.magentaBright("LlamaChat✨:"))}\n${summary}`);
} catch (error) {
logger.error(`Error processing your request: ${error}`);
}
}
return 0;
}
main().catch(console.error);
+8
View File
@@ -0,0 +1,8 @@
import { createConsola } from "consola";
import type { ConsolaInstance } from "consola";
export const logger: ConsolaInstance = createConsola({
formatOptions: {
date: false,
},
});
+56
View File
@@ -0,0 +1,56 @@
import { generateText } from "ai";
import { openai } from "@ai-sdk/openai";
import { NodeWithScore, MetadataMode } from "llamaindex";
import * as readline from "readline/promises";
import figlet from "figlet";
import pc from "picocolors";
export async function renderLogo(): Promise<void> {
const logoText = figlet.textSync("LlamaChat", {
font: "ANSI Shadow",
horizontalLayout: "default",
verticalLayout: "default",
width: 100,
whitespaceBreak: true,
});
// Add some styling with picocolors
const styledLogo = pc.bold(pc.yellowBright(logoText));
// Add some padding/margin
console.log("\n");
console.log(styledLogo);
console.log(pc.gray("─".repeat(60)));
console.log("\n");
}
export async function consoleInput(): Promise<string> {
const rl = readline.createInterface({
input: process.stdin,
output: process.stdout,
});
const answer = await rl.question(pc.cyanBright("You✨:"));
rl.close();
return answer;
}
export async function retrievalAugmentedGeneration(
nodes: NodeWithScore[],
prompt: string,
): Promise<string> {
let mainText: string = "";
for (const node of nodes) {
mainText += `\t{information: '${node.node.getContent(
MetadataMode.ALL,
)}', relevanceScore: '${node.score ?? "no score"}'}\n`;
}
const { text } = await generateText({
model: openai("gpt-4.1"),
prompt: `[\n${mainText}\n]\n\nBased on the information you are given and on the relevance score of that (where -1 means no score available), answer to this user prompt: '${prompt}'`,
});
return text;
}
+22
View File
@@ -0,0 +1,22 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "ES2022",
"lib": ["ES2022"],
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"declaration": true,
"declarationMap": true,
"sourceMap": true,
"types": ["node"],
"moduleResolution": "bundler",
"allowSyntheticDefaultImports": true,
"resolveJsonModule": true
},
"include": ["src/**/*"],
"exclude": ["node_modules", "dist"]
}
+124
View File
@@ -0,0 +1,124 @@
# LlamaParse Demo
A TypeScript demo application showcasing the power of **LlamaParse** - an intelligent document parsing service from [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to parse various document formats and generate AI-powered summaries using OpenAI's GPT models.
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [Development Mode](#development-mode)
- [Build the Project](#build-the-project)
- [Code Quality](#code-quality)
- [Quick Commands Reference](#quick-commands-reference)
- [How It Works](#how-it-works)
- [API Dependencies](#api-dependencies)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 📄 **Document Parsing**: Parse PDFs, Word docs, and other formats using LlamaParse
- 🤖 **AI Summaries**: Generate intelligent summaries using OpenAI GPT-4
- 🎨 **Beautiful CLI**: Styled console interface with colors and ASCII art
-**Fast Development**: Hot reload support with watch mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 18 or higher)
- pnpm package manager
- OpenAI API key
- LlamaCloud API key
## Installation
1. Clone the repository:
```bash
git clone <repository-url>
cd llamaparse-demo
```
2. Install dependencies:
```bash
pnpm install
```
3. Set up your environment variables:
```bash
# Add your API keys to your environment
export OPENAI_API_KEY="your-openai-api-key"
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
```
## Usage
### Start the Demo
```bash
pnpm run start
```
The application will display a welcome screen and prompt you to enter the path to a document you'd like to process.
### Development Mode
For development with hot reload:
```bash
pnpm run dev
```
### Build the Project
```bash
pnpm run build
```
### Code Quality
Format code:
```bash
pnpm run format
```
Lint code:
```bash
pnpm run lint
```
## How It Works
1. **Document Input**: Enter the path to your document when prompted
2. **Parsing**: LlamaParse processes the document and extracts structured content
3. **AI Summary**: The extracted content is sent to OpenAI GPT-4 for summarization
4. **Results**: View the AI-generated summary in your terminal
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 18+ and have all dependencies installed
2. **API Key Issues**: Verify your OpenAI and LlamaCloud API keys are correctly set
3. **File Path Errors**: Use absolute paths or ensure relative paths are correct from the project root
## License
MIT License - see the [LICENSE](../../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `pnpm format` and `pnpm lint`
5. Submit a pull request
Binary file not shown.
+15
View File
@@ -0,0 +1,15 @@
import js from "@eslint/js";
import globals from "globals";
import tseslint from "typescript-eslint";
import { defineConfig } from "eslint/config";
export default defineConfig([
{
files: ["**/*.{js,mjs,cjs,ts,mts,cts}"],
plugins: { js },
extends: ["js/recommended"],
languageOptions: { globals: globals.browser },
},
{ files: ["**/*.js"], languageOptions: { sourceType: "script" } },
tseslint.configs.recommended,
]);
+47
View File
@@ -0,0 +1,47 @@
{
"name": "llamaparse-demo",
"version": "0.1.0",
"description": "Demo for LlamaParse in TypeScript",
"type": "module",
"main": "index.js",
"scripts": {
"test": "echo \"There are no tests\"",
"start": "pnpm exec tsx src/index.ts",
"lint": "eslint ./src/",
"format": "prettier --write ./src/",
"build": "tsc",
"dev": "pnpm exec tsx --watch src/index.ts"
},
"keywords": [
"ai",
"ocr",
"parsing",
"intelligent-document-processing",
"pdf",
"llms"
],
"author": "LlamaIndex",
"license": "MIT",
"packageManager": "pnpm@10.12.4",
"devDependencies": {
"@eslint/js": "^9.32.0",
"@types/figlet": "^1.7.0",
"@types/node": "^24.1.0",
"@typescript-eslint/eslint-plugin": "^8.38.0",
"@typescript-eslint/parser": "^8.38.0",
"eslint": "^9.32.0",
"globals": "^16.3.0",
"jiti": "^2.5.1",
"prettier": "^3.6.2",
"typescript": "^5.8.3",
"typescript-eslint": "^8.38.0"
},
"dependencies": {
"@ai-sdk/openai": "^1.3.23",
"ai": "^4.3.19",
"consola": "^3.4.2",
"figlet": "^1.8.2",
"llama-cloud-services": "link:../../../ts/llama_cloud_services",
"picocolors": "^1.1.1"
}
}
+1758
View File
File diff suppressed because it is too large Load Diff
+34
View File
@@ -0,0 +1,34 @@
import { LlamaParseReader } from "llama-cloud-services";
import { logger } from "./logger";
import pc from "picocolors";
import { consoleInput, generateSummary, renderLogo } from "./utils";
export async function main(): Promise<number> {
const reader = new LlamaParseReader({ resultType: "markdown" });
await renderLogo();
logger.log(
`Welcome to ${pc.bold(
pc.magentaBright("✨LlamaParse Demo✨"),
)}, our demo for ${pc.bold(pc.green("LlamaParse🦙"))}, a ${pc.bold(
pc.cyan("LlamaCloud☁️"),
)} (https://cloud.llamaindex.ai) product!.\nType the path to the document you would like to process below👇\nIf you wish to exit, just type ${pc.bold(
pc.gray("quit"),
)}.\n`,
);
while (true) {
const userInput = await consoleInput();
if (userInput.toLowerCase() == "quit") {
break;
}
try {
const documents = await reader.loadData(userInput);
const summary = await generateSummary(documents); // Added await here
logger.log(`${pc.bold(pc.cyan("AI-generated summary✨"))}:\n${summary}`);
} catch (error) {
logger.error(`Error processing file: ${error}`);
}
}
return 0;
}
main().catch(console.error);
+8
View File
@@ -0,0 +1,8 @@
import { createConsola } from "consola";
import type { ConsolaInstance } from "consola";
export const logger: ConsolaInstance = createConsola({
formatOptions: {
date: false,
},
});
+51
View File
@@ -0,0 +1,51 @@
import { generateText } from "ai";
import { openai } from "@ai-sdk/openai";
import { Document } from "llamaindex";
import * as readline from "readline/promises";
import figlet from "figlet";
import pc from "picocolors";
export async function renderLogo(): Promise<void> {
const logoText = figlet.textSync("LlamaParse Demo", {
font: "ANSI Shadow",
horizontalLayout: "default",
verticalLayout: "default",
width: 100,
whitespaceBreak: true,
});
// Add some styling with picocolors
const styledLogo = pc.bold(pc.magentaBright(logoText));
// Add some padding/margin
console.log("\n");
console.log(styledLogo);
console.log(pc.gray("─".repeat(60)));
console.log("\n");
}
export async function consoleInput(): Promise<string> {
const rl = readline.createInterface({
input: process.stdin,
output: process.stdout,
});
const answer = await rl.question("Path to your file: ");
rl.close();
return answer;
}
export async function generateSummary(documents: Document[]): Promise<string> {
let mainText: string = "";
for (const document of documents) {
mainText += `${document.text}\n\n---\n\n`;
}
const { text } = await generateText({
model: openai("gpt-4.1"),
prompt: `</chat>\n\t<text>${mainText}</text>\n\t<instructions>Could you please generate a summary of the given text?</instructions>\n</chat>`,
});
return text;
}
+22
View File
@@ -0,0 +1,22 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "ES2022",
"lib": ["ES2022"],
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"declaration": true,
"declarationMap": true,
"sourceMap": true,
"types": ["node"],
"moduleResolution": "bundler",
"allowSyntheticDefaultImports": true,
"resolveJsonModule": true
},
"include": ["src/**/*"],
"exclude": ["node_modules", "dist"]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
-295
View File
@@ -1,295 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Using llama-parse with AstraDB"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"In this notebook, we show a basic RAG-style example that uses `llama-parse` to parse a PDF document, store the corresponding document into a vector store (`AstraDB`) and finally, perform some basic queries against that store. The notebook is modeled after the quick start notebooks and hence is meant as a way of getting started with `llama-parse`, backed by a vector database."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Requirements"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# First, install the required dependencies\n",
"%pip install --quiet llama-index llama-parse llama-index-vector-stores-astra-db llama-index-llms-openai"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Configuration"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"import openai\n",
"\n",
"from getpass import getpass\n",
"\n",
"# Get all required API keys and parameters\n",
"llama_cloud_api_key = getpass(\"Enter your Llama Index Cloud API Key: \")\n",
"api_endpoint = input(\"Enter your Astra DB API Endpoint: \")\n",
"token = getpass(\"Enter your Astra DB Token: \")\n",
"namespace = (\n",
" input(\"Enter your Astra DB namespace (optional, must exist on Astra): \") or None\n",
")\n",
"openai_api_key = getpass(\"Enter your OpenAI API Key: \")\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = llama_cloud_api_key\n",
"openai.api_key = openai_api_key"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Using llama-parse to parse a PDF"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Download complete.\n"
]
}
],
"source": [
"# Grab a PDF from Arxiv for indexing\n",
"import requests\n",
"\n",
"# The URL of the file you want to download\n",
"url = \"https://arxiv.org/pdf/1706.03762.pdf\"\n",
"# The local path where you want to save the file\n",
"file_path = \"./attention.pdf\"\n",
"\n",
"# Perform the HTTP request\n",
"response = requests.get(url)\n",
"\n",
"# Check if the request was successful\n",
"if response.status_code == 200:\n",
" # Open the file in binary write mode and save the content\n",
" with open(file_path, \"wb\") as file:\n",
" file.write(response.content)\n",
" print(\"Download complete.\")\n",
"else:\n",
" print(\"Error downloading the file.\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id ce3909a7-54cf-438b-849a-fe9a903b0c71\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'rmer - model architecture.\\nThe Transformer follows this overall architecture using stacked self-attention and point-wise, fully\\nconnected layers for both the encoder and decoder, shown in the left and right halves of Figure 1,\\nrespectively.\\n3.1 Encoder and Decoder Stacks\\nEncoder: The encoder is composed of a stack of N = 6 identical layers. Each layer has two\\nsub-layers. The first is a multi-head self-attention mechanism, and the second is a simple, position-\\nwise fully connected feed-forward network. We employ a residual connection [11] around each of\\nthe two sub-layers, followed by layer normalization [1]. That is, the output of each sub-layer is\\nLayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer\\nitself. To facilitate these residual connections, all sub-layers in the model, as well as the embedding\\nlayers, produce outputs of dimension dmodel = 512.\\nDecoder: The decoder is also composed of a stack of N = 6 identical layers. In addition '"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Take a quick look at some of the parsed text from the document:\n",
"documents[0].get_content()[10000:11000]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Storing into Astra DB"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.vector_stores.astra_db import AstraDBVectorStore\n",
"\n",
"astra_db_store = AstraDBVectorStore(\n",
" token=token,\n",
" api_endpoint=api_endpoint,\n",
" namespace=namespace,\n",
" collection_name=\"astra_v_table_llamaparse\",\n",
" embedding_dimension=1536,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.node_parser import SimpleNodeParser\n",
"\n",
"node_parser = SimpleNodeParser()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"\n",
"storage_context = StorageContext.from_defaults(vector_store=astra_db_store)\n",
"\n",
"index = VectorStoreIndex(\n",
" nodes=nodes,\n",
" storage_context=storage_context,\n",
" embed_model=OpenAIEmbedding(api_key=openai_api_key),\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Simple RAG Example"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"query_engine = index.as_query_engine(similarity_top_k=15)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********New LlamaParse+ Basic Query Engine***********\n",
"Multi-Head Attention is also known as multi-headed self-attention.\n"
]
}
],
"source": [
"query = \"What is Multi-Head Attention also known as?\"\n",
"\n",
"response_1 = query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Basic Query Engine***********\")\n",
"print(response_1)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'We used beam search as described in the previous section, but no\\ncheckpoint averaging. We present these results in Table 3.\\nIn Table 3 rows (A), we vary the number of attention heads and the attention key and value dimensions,\\nkeeping the amount of computation constant, as described in Section 3.2.2. While single-head\\nattention is 0.9 BLEU worse than the best setting, quality also drops off with too many heads.\\nIn Table 3 rows (B), we observe that reducing the attention key size dk hurts model quality. This\\nsuggests that determining compatibility is not easy and that a more sophisticated compatibility\\nfunction than dot product may be beneficial. We further observe in rows (C) and (D) that, as expected,\\nbigger models are better, and dropout is very helpful in avoiding over-fitting. In row (E) we replace our\\nsinusoidal positional encoding with learned positional embeddings [9], and observe nearly identical\\nresults to the base model.\\n6.3 English Constituency Parsing\\nTo evaluate if the Transformer can generalize to other tasks we performed experiments on English\\nconstituency parsing. This task presents specific challenges: the output is subject to strong structural\\nconstraints and is significantly longer than the input. Furthermore, RNN sequence-to-sequence\\nmodels have not been able to attain state-of-the-art results in small-data regimes [37].\\nWe trained a 4-layer transformer with dmodel = 1024 on the Wall Street Journal (WSJ) portion of the\\nPenn Treebank [25], about 40K training sentences. We also trained it in a semi-supervised setting,\\nusing the larger high-confidence and BerkleyParser corpora from with approximately 17M sentences\\n[37]. We used a vocabulary of 16K tokens for the WSJ only setting and a vocabulary of 32K tokens\\nfor the semi-supervised setting.\\nWe performed only a small number of experiments to select the dropout, both attention and residual\\n(section 5.4), learning rates and beam size on the Section 22 development set, all other parameters\\nremained unchanged from the English-to-German base translation model. During inference, we\\n 9\\n---\\nTable 4: The Transformer generalizes well to English constituency parsing (Results are on Section 23\\nof WSJ)\\n Parser Training WSJ 23 F1\\n Vinyals & Kaiser el al. (2014) [37] WSJ only, discriminative 88.3\\n Petrov et al. (2006) [29] WSJ only, discriminative 90.4\\n Zhu et al. (2013) [40] WSJ only, discriminative 90.4\\n Dyer et al. (2016) [8] WSJ only, discriminative 91.7\\n Transformer (4 layers) WSJ only, discriminative 91.3\\n Zhu et al. (2013) [40] semi-supervised 91.3\\n Huang & Harper (2009) [14] semi-supervised 91.3\\n McClosky et al. (2006) [26] semi-supervised 92.1\\n Vinyals & Kaiser el al. (2014) [37] semi-supervised 92.1\\n Transformer (4 layers) semi-supervised 92.7\\n Luong et al. (2015) [23] multi-task 93.0\\n Dyer et al. (2016) [8] generative 93.3\\nincreased the maximum output length to input length + 300. We used a beam size of 21 and α = 0.3\\nfor both WSJ only and the semi-supervised setting.\\nOur results in Table 4 show that despite the lack of task-specific tuning our model performs sur-\\nprisingly well, yielding better results than all previously reported models with the exception of the\\nRecurrent Neural Network Grammar [8].\\nIn contrast to RNN sequence-to-sequence models [37], the Transformer outperforms the Berkeley-\\nParser [29] even when training only on the WSJ training set of 40K sentences.\\n7 Conclusion\\nIn this work, we presented the Transformer, the first sequence transduction model based entirely on\\nattention, replacing the recurrent layers most commonly used in encoder-decoder architectures with\\nmulti-headed self-attention.\\nFor translation tasks, the Transformer can be trained significantly faster than architectures based\\non recurrent or convolutional layers.'"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Take a look at one of the source nodes from the response\n",
"response_1.source_nodes[0].get_content()"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
+82 -69
View File
@@ -13,32 +13,14 @@
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-parse"
"%pip install llama-index llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-02-02 11:10:10-- https://arxiv.org/pdf/1706.03762.pdf\n",
"Resolving arxiv.org (arxiv.org)... 151.101.131.42, 151.101.3.42, 151.101.67.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.131.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 2215244 (2.1M) [application/pdf]\n",
"Saving to: ./attention.pdf\n",
"\n",
"./attention.pdf 100%[===================>] 2.11M --.-KB/s in 0.08s \n",
"\n",
"2024-02-02 11:10:10 (25.9 MB/s) - ./attention.pdf saved [2215244/2215244]\n",
"\n"
]
}
],
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/1706.03762.pdf\" -O \"./attention.pdf\""
]
@@ -49,11 +31,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
@@ -68,14 +45,14 @@
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id dd0b8e31-0c09-4497-b78a-cc1c92f1d6cf\n"
"Started parsing the file under job_id 79ae653c-4598-4bd0-ba6e-b3dab7eab57e\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"text\").load_data(\"./attention.pdf\")"
"result = await LlamaParse().aparse(\"./attention.pdf\")"
]
},
{
@@ -87,23 +64,62 @@
"name": "stdout",
"output_type": "stream",
"text": [
"ad\n",
"1 Introduction\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neural networks\n",
"in particular, have been firmly established as state of the art approaches in sequence modeling and\n",
"transduction problems such as language modeling and machine translation [35, 2, 5]. Numerous\n",
"efforts have since continued to push the boundaries of recurrent language models and encoder-decoder\n",
"architectures [38, 24, 15].\n",
"Recurrent models typically factor computation along the symbol positions of the input and output\n",
"sequences. Aligning the positions to steps in computation time, they generate a sequence of hidden\n",
"states ht, as a function of the previous hidden state ht1 and the input for position t. This inherently\n",
"sequential nature precludes parallelization within training examples, which becomes critical at longer\n",
"sequence lengths, as memory constraints limit batching across examples. Recent work has achieved\n",
"significant improvements in computational efficiency through factorization tricks [21] and conditional\n",
"computation [32], while also improving model performance in case of the latter. The fundamental\n",
"constraint of sequential computation, however, remains.\n",
"Attention mechanisms have become an integral part of compelling sequence modeling and transduc-\n",
"tion models in various tasks, allowing modeling of dependencies without regard to their distance in\n",
"the input or output sequences [2, 19]. In all but a few cases [27], however, such attention mechanisms\n",
"are used in conjunction with a recurrent network.\n",
"In this work we propose the Transformer, a model architecture eschewing recurrence and instead\n",
"relying entirely on an attention mechanism to draw global dependencies between input and output.\n",
"The Transformer allows for significantly more parallelization and can reach a new state of the art in\n",
"translation quality after being trained for as little as twelve hours on eight P100 GPUs.\n",
"2 Background\n",
"2 Background\n",
"The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU\n",
"[16], ByteNet [18] and ConvS2S [9], all of which use convolutional neural networks as basic building\n",
"block, computing hidden representations in parallel for all input and output positions. In these models,\n",
"the number of operations required to relate signals from two arbitrary input or output positions grows\n",
"in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes\n",
"it more difficult to learn dependencies between distant positions [12]. In the Transformer this is\n",
"reduced to a constant number of operations, albeit at the cost of reduced effective res\n"
"reduced to a constant number of operations, albeit at the cost of reduced effective resolution due\n",
"to averaging attention-weighted positions, an effect we counteract with Multi-Head Attention as\n",
"described in section 3.2.\n",
"Self-attention, sometimes called intra-attention is an attention mechanism relating different positions\n",
"of a single sequence in order to compute a representation of the sequence. Self-attention has been\n",
"used successfully in a variety of tasks including reading comprehension, abstractive summarization,\n",
"textual entailment and learning task-independent sentence representations [4, 27, 28, 22].\n",
"End-to-end memory networks are based on a recurrent attention mechanism instead of sequence-\n",
"aligned recurrence and have been shown to perform well on simple-language question answering and\n",
"language modeling tasks [34].\n",
"To the best of our knowledge, however, the Transformer is the first transduction model relying\n",
"entirely on self-attention to compute representations of its input and output without using sequence-\n",
"aligned RNNs or convolution. In the following sections, we will describe the Transformer, motivate\n",
"self-attention and discuss its advantages over models such as [17, 18] and [9].\n",
"3 Model Architecture\n",
"Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].\n",
"Here, the encoder maps an input sequence of symbol representations (x1, ..., xn) to a sequence\n",
"of continuous representations z = (z1, ..., zn). Given z, the decoder then generates an output\n",
"sequence (y1, ..., ym) of symbols one element at a time. At each step the model is auto-regressive\n",
"[10], consuming the previously generated symbols as additional input when generating the next.\n",
" 2\n"
]
}
],
"source": [
"print(documents[0].text[6000:7000])"
"documents = result.get_text_documents(split_by_page=True)\n",
"print(documents[1].text)"
]
},
{
@@ -115,48 +131,45 @@
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id d4531453-1bbb-48c4-8324-ae9fea9f2fa2\n"
"arXiv:1706.03762v7 [cs.CL] 2 Aug 2023\n",
"\n",
"Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.\n",
"\n",
"# Attention Is All You Need\n",
"\n",
"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit\n",
"\n",
"Google Brain Google Brain Google Research Google Research\n",
"\n",
"avaswani@google.com noam@google.com nikip@google.com usz@google.com\n",
"\n",
"Llion Jones Aidan N. Gomez † Łukasz Kaiser\n",
"\n",
"Google Research University of Toronto Google Brain\n",
"\n",
"llion@google.com aidan@cs.toronto.edu lukaszkaiser@google.com\n",
"\n",
"Illia Polosukhin ‡\n",
"\n",
"illia.polosukhin@gmail.com\n",
"\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.\n",
"\n",
"Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research.\n",
"\n",
"†Work performed while at Google Brain.\n",
"\n",
"‡Work performed while at Google Research.\n",
"\n",
"31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"markdown\").load_data(\"./attention.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"ction describes the training regime for our models.\n",
"\n",
"##### Training Data and Batching\n",
"\n",
"We trained on the standard WMT 2014 English-German dataset consisting of about 4.5 million\n",
"sentence pairs. Sentences were encoded using byte-pair encoding [3], which has a shared source-\n",
"target vocabulary of about 37000 tokens. For English-French, we used the significantly larger WMT\n",
"2014 English-French dataset consisting of 36M sentences and split tokens into a 32000 word-piece\n",
"vocabulary [38]. Sentence pairs were batched together by approximate sequence length. Each training\n",
"batch contained a set of sentence pairs containing approximately 25000 source tokens and 25000\n",
"target tokens.\n",
"\n",
"##### Hardware and Schedule\n",
"\n",
"We trained our models on one machine with 8 NVIDIA P100 GPUs. For our base models using\n",
"the hyperparameters described throughout the paper, each training step took about 0.4 seconds. We\n",
"trained the base models for a total of 100,000 steps or 12 hours. For our big models,(described on the\n",
"bo...\n"
]
}
],
"source": [
"print(documents[0].text[20000:21000] + \"...\")"
"documents = result.get_markdown_documents(split_by_page=True)\n",
"print(documents[0].text)"
]
}
],
File diff suppressed because one or more lines are too long
+35 -50
View File
@@ -31,11 +31,11 @@
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-index\n",
"!pip install llama-index-core\n",
"!pip install llama-index-llms-anthropic llama-index-multi-modal-llms-anthropic\n",
"!pip install llama-index-embeddings-huggingface\n",
"!pip install llama-cloud-services"
"%pip install llama-index\n",
"%pip install llama-index-core\n",
"%pip install llama-index-llms-anthropic\n",
"%pip install llama-index-embeddings-huggingface\n",
"%pip install llama-cloud-services"
]
},
{
@@ -45,11 +45,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -68,7 +63,7 @@
"source": [
"from llama_index.llms.anthropic import Anthropic\n",
"\n",
"llm = Anthropic(model=\"claude-3-opus-20240229\", temperature=0.0)"
"llm = Anthropic(model=\"claude-3-5-sonnet-20241022\")"
]
},
{
@@ -131,28 +126,8 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(verbose=True)\n",
"json_objs = parser.get_json_result(\"./uber_10q_march_2022.pdf\")\n",
"json_list = json_objs[0][\"pages\"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b26d21d1-05b5-4f49-b937-c13106a84015",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"text\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
"parser = LlamaParse(take_screenshot=True)\n",
"result = await parser.aparse(\"./uber_10q_march_2022.pdf\")"
]
},
{
@@ -162,7 +137,12 @@
"metadata": {},
"outputs": [],
"source": [
"text_nodes = get_text_nodes(json_list)"
"text_nodes = await result.aget_text_nodes(split_by_page=True)\n",
"image_nodes = await result.aget_image_nodes(\n",
" include_screenshot_images=True,\n",
" include_object_images=True,\n",
" image_download_dir=\"./uber_10q_images\",\n",
")"
]
},
{
@@ -172,7 +152,7 @@
"source": [
"## Extract/Index images from image dicts\n",
"\n",
"Here we use a multimodal model to extract and index images from image dictionaries."
"Here we use a multimodal model to caption images and create text nodes for indexing."
]
},
{
@@ -190,27 +170,32 @@
}
],
"source": [
"# call get_images on parser, convert to ImageDocuments\n",
"!mkdir llama2_images\n",
"!mkdir -p llama2_images\n",
"\n",
"from llama_index.core.schema import ImageDocument\n",
"from llama_index.multi_modal_llms.anthropic import AnthropicMultiModal\n",
"from llama_index.core.llms import ChatMessage, ImageBlock, TextBlock\n",
"from llama_index.core.schema import ImageNode, TextNode\n",
"from llama_index.llms.anthropic import Anthropic\n",
"\n",
"\n",
"def get_image_text_nodes(json_objs: List[dict]):\n",
"def get_image_text_nodes(image_nodes: list[ImageNode]):\n",
" \"\"\"Extract out text from images using a multimodal model.\"\"\"\n",
" anthropic_mm_llm = AnthropicMultiModal(max_tokens=300)\n",
" image_dicts = parser.get_images(json_objs, download_path=\"llama2_images\")\n",
" image_documents = []\n",
" llm = Anthropic(model=\"claude-3-5-haiku-20241022\", max_tokens=300)\n",
" img_text_nodes = []\n",
" for image_dict in image_dicts:\n",
" image_doc = ImageDocument(image_path=image_dict[\"path\"])\n",
" response = anthropic_mm_llm.complete(\n",
" prompt=\"Describe the images as alt text\",\n",
" image_documents=[image_doc],\n",
" for image_node in image_nodes:\n",
" image_path = image_node.image_path\n",
" message = ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" TextBlock(text=\"Describe the images as alt text\"),\n",
" ImageBlock(path=image_path),\n",
" ],\n",
" )\n",
" response = llm.chat([message])\n",
" text_node = TextNode(\n",
" text=str(response.message.content), metadata={\"path\": image_path}\n",
" )\n",
" text_node = TextNode(text=str(response), metadata={\"path\": image_dict[\"path\"]})\n",
" img_text_nodes.append(text_node)\n",
"\n",
" return img_text_nodes"
]
},
@@ -221,7 +206,7 @@
"metadata": {},
"outputs": [],
"source": [
"image_text_nodes = get_image_text_nodes(json_objs)"
"image_text_nodes = get_image_text_nodes(image_nodes)"
]
},
{
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+10 -12
View File
@@ -31,14 +31,9 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
@@ -79,8 +74,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"text\", language=\"fr\")\n",
"documents = parser.load_data(\"./treasury_report.pdf\")"
"parser = LlamaParse(language=\"fr\")\n",
"result = await parser.aparse(\"./treasury_report.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -252,8 +248,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"text\", language=\"ch_sim\")\n",
"documents = parser.load_data(\"./chinese_pdf.pdf\")"
"parser = LlamaParse(language=\"ch_sim\")\n",
"result = await parser.aparse(\"./chinese_pdf.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -406,8 +403,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"base_parser = LlamaParse(result_type=\"text\", language=\"en\")\n",
"base_documents = parser.load_data(\"./chinese_pdf2.pdf\")"
"base_parser = LlamaParse(language=\"en\")\n",
"result = await base_parser.aparse(\"./chinese_pdf2.pdf\")\n",
"base_documents = result.get_text_documents(split_by_page=False)"
]
},
{
+4 -8
View File
@@ -60,11 +60,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import requests\n",
"import pymongo\n",
"\n",
@@ -72,7 +67,7 @@
"from llama_cloud_services import LlamaParse\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"from llama_index.core.node_parser import SimpleNodeParser"
"from llama_index.core.node_parser import SentenceSplitter"
]
},
{
@@ -137,7 +132,8 @@
}
],
"source": [
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
"result = await LlamaParse().aparse(file_path)\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -203,7 +199,7 @@
"metadata": {},
"outputs": [],
"source": [
"node_parser = SimpleNodeParser()\n",
"node_parser = SentenceSplitter()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
@@ -1,544 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse - Parsing comic books with parsing intructions\n",
"Parsing intructions allow you to instruct our parsing model the same way you would instruct an LLM!\n",
"\n",
"They can be useful to help the parser get better results on complex document layouts, to extract data in a specific format, or to transform the document in other ways.\n",
"\n",
"Using Parsing Instruction you will get better results out of LlamaParse on complicated documents, and also be able to simplify your application code."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Installation\n",
"\n",
"Parsing instructions are part of the llamaParse API. They can be accessed by directly specifying the parsing_instruction parameter in the API or by using the LlamaParse python module (which we will use for this tutorial).\n",
"\n",
"To install llama-parse, just get it from PIP:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Collecting llama-parse\n",
" Downloading llama_parse-0.3.8-py3-none-any.whl (6.7 kB)\n",
"Collecting llama-index-core>=0.10.7 (from llama-parse)\n",
" Downloading llama_index_core-0.10.19-py3-none-any.whl (15.3 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m15.3/15.3 MB\u001b[0m \u001b[31m31.9 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: PyYAML>=6.0.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (6.0.1)\n",
"Requirement already satisfied: SQLAlchemy[asyncio]>=1.4.49 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.0.28)\n",
"Requirement already satisfied: aiohttp<4.0.0,>=3.8.6 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.9.3)\n",
"Collecting dataclasses-json (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dataclasses_json-0.6.4-py3-none-any.whl (28 kB)\n",
"Collecting deprecated>=1.2.9.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading Deprecated-1.2.14-py2.py3-none-any.whl (9.6 kB)\n",
"Collecting dirtyjson<2.0.0,>=1.0.8 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dirtyjson-1.0.8-py3-none-any.whl (25 kB)\n",
"Requirement already satisfied: fsspec>=2023.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2023.6.0)\n",
"Collecting httpx (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpx-0.27.0-py3-none-any.whl (75 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m75.6/75.6 kB\u001b[0m \u001b[31m6.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hCollecting llamaindex-py-client<0.2.0,>=0.1.13 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading llamaindex_py_client-0.1.13-py3-none-any.whl (107 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m108.0/108.0 kB\u001b[0m \u001b[31m10.0 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: nest-asyncio<2.0.0,>=1.5.8 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.6.0)\n",
"Requirement already satisfied: networkx>=3.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.2.1)\n",
"Requirement already satisfied: nltk<4.0.0,>=3.8.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.8.1)\n",
"Requirement already satisfied: numpy in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.25.2)\n",
"Collecting openai>=1.1.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading openai-1.13.3-py3-none-any.whl (227 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m227.4/227.4 kB\u001b[0m \u001b[31m16.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: pandas in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.5.3)\n",
"Requirement already satisfied: pillow>=9.0.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (9.4.0)\n",
"Requirement already satisfied: requests>=2.31.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.31.0)\n",
"Requirement already satisfied: tenacity<9.0.0,>=8.2.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (8.2.3)\n",
"Collecting tiktoken>=0.3.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading tiktoken-0.6.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.8 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m1.8/1.8 MB\u001b[0m \u001b[31m43.1 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: tqdm<5.0.0,>=4.66.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.66.2)\n",
"Requirement already satisfied: typing-extensions>=4.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.10.0)\n",
"Collecting typing-inspect>=0.8.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading typing_inspect-0.9.0-py3-none-any.whl (8.8 kB)\n",
"Requirement already satisfied: aiosignal>=1.1.2 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Requirement already satisfied: attrs>=17.3.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (23.2.0)\n",
"Requirement already satisfied: frozenlist>=1.1.1 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.4.1)\n",
"Requirement already satisfied: multidict<7.0,>=4.5 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (6.0.5)\n",
"Requirement already satisfied: yarl<2.0,>=1.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.9.4)\n",
"Requirement already satisfied: async-timeout<5.0,>=4.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (4.0.3)\n",
"Requirement already satisfied: wrapt<2,>=1.10 in /usr/local/lib/python3.10/dist-packages (from deprecated>=1.2.9.3->llama-index-core>=0.10.7->llama-parse) (1.14.1)\n",
"Requirement already satisfied: pydantic>=1.10 in /usr/local/lib/python3.10/dist-packages (from llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.6.3)\n",
"Requirement already satisfied: anyio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.7.1)\n",
"Requirement already satisfied: certifi in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (2024.2.2)\n",
"Collecting httpcore==1.* (from httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpcore-1.0.4-py3-none-any.whl (77 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m77.8/77.8 kB\u001b[0m \u001b[31m8.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: idna in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.6)\n",
"Requirement already satisfied: sniffio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Collecting h11<0.15,>=0.13 (from httpcore==1.*->httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading h11-0.14.0-py3-none-any.whl (58 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m58.3/58.3 kB\u001b[0m \u001b[31m5.7 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: click in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (8.1.7)\n",
"Requirement already satisfied: joblib in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (1.3.2)\n",
"Requirement already satisfied: regex>=2021.8.3 in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (2023.12.25)\n",
"Requirement already satisfied: distro<2,>=1.7.0 in /usr/lib/python3/dist-packages (from openai>=1.1.0->llama-index-core>=0.10.7->llama-parse) (1.7.0)\n",
"Requirement already satisfied: charset-normalizer<4,>=2 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (3.3.2)\n",
"Requirement already satisfied: urllib3<3,>=1.21.1 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (2.0.7)\n",
"Requirement already satisfied: greenlet!=0.4.17 in /usr/local/lib/python3.10/dist-packages (from SQLAlchemy[asyncio]>=1.4.49->llama-index-core>=0.10.7->llama-parse) (3.0.3)\n",
"Collecting mypy-extensions>=0.3.0 (from typing-inspect>=0.8.0->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading mypy_extensions-1.0.0-py3-none-any.whl (4.7 kB)\n",
"Collecting marshmallow<4.0.0,>=3.18.0 (from dataclasses-json->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading marshmallow-3.21.1-py3-none-any.whl (49 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m49.4/49.4 kB\u001b[0m \u001b[31m4.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: python-dateutil>=2.8.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2.8.2)\n",
"Requirement already satisfied: pytz>=2020.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2023.4)\n",
"Requirement already satisfied: exceptiongroup in /usr/local/lib/python3.10/dist-packages (from anyio->httpx->llama-index-core>=0.10.7->llama-parse) (1.2.0)\n",
"Requirement already satisfied: packaging>=17.0 in /usr/local/lib/python3.10/dist-packages (from marshmallow<4.0.0,>=3.18.0->dataclasses-json->llama-index-core>=0.10.7->llama-parse) (23.2)\n",
"Requirement already satisfied: annotated-types>=0.4.0 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (0.6.0)\n",
"Requirement already satisfied: pydantic-core==2.16.3 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.16.3)\n",
"Requirement already satisfied: six>=1.5 in /usr/local/lib/python3.10/dist-packages (from python-dateutil>=2.8.1->pandas->llama-index-core>=0.10.7->llama-parse) (1.16.0)\n",
"Installing collected packages: dirtyjson, mypy-extensions, marshmallow, h11, deprecated, typing-inspect, tiktoken, httpcore, httpx, dataclasses-json, openai, llamaindex-py-client, llama-index-core, llama-parse\n",
"Successfully installed dataclasses-json-0.6.4 deprecated-1.2.14 dirtyjson-1.0.8 h11-0.14.0 httpcore-1.0.4 httpx-0.27.0 llama-index-core-0.10.19 llama-parse-0.3.8 llamaindex-py-client-0.1.13 marshmallow-3.21.1 mypy-extensions-1.0.0 openai-1.13.3 tiktoken-0.6.0 typing-inspect-0.9.0\n"
]
}
],
"source": [
"%pip install llama-cloud-services"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## API key\n",
"\n",
"The use of LlamaParse requires an API key which you can get here: https://cloud.llamaindex.ai/parse"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Async (Notebook only)\n",
"llama-parse is async-first, so running the code in a notebook requires the use of nest_asyncio\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Import the package"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using llamaparse for getting better results (on Manga!)\n",
"\n",
"Sometimes the layout of a page is unusual and you will get sub-optimal reading order results with LlamaParse. For example, when parsing manga you expect the reading order to be right to left even if the content is in English!"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let's download an extract of a great manga \"The manga guide to calculus\", by Hiroyuki Kojima (https://www.amazon.com/Manga-Guide-Calculus-Hiroyuki-Kojima/dp/1593271948)\n",
"\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Resolving drive.usercontent.google.com (drive.usercontent.google.com)... 173.194.211.132, 2607:f8b0:400c:c10::84\n",
"Connecting to drive.usercontent.google.com (drive.usercontent.google.com)|173.194.211.132|:443... connected.\n",
"HTTP request sent, awaiting response... 303 See Other\n",
"Location: https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download [following]\n",
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Reusing existing connection to drive.usercontent.google.com:443.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 3041634 (2.9M) [application/octet-stream]\n",
"Saving to: ./manga.pdf\n",
"\n",
"./manga.pdf 100%[===================>] 2.90M --.-KB/s in 0.04s \n",
"\n",
"2024-03-13 13:57:20 (78.6 MB/s) - ./manga.pdf saved [3041634/3041634]\n",
"\n"
]
}
],
"source": [
"! wget \"https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\" -O ./manga.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Without parsing instructions\n",
"For the sake of comparison, let's first parse without any instructions."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 25bf4202-78d8-4705-88cf-c616ae7c82af\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"As you can see below, LlamaParse is not doing a great job here. It is interpreting the grid of comic panels as a table, and trying to fit the dialogue into a table. It's very hard to follow."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"---\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = \n"
]
}
],
"source": [
"print(vanilaParsing[0].text[100:1000])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Using parsing instructions\n",
"Let's try to parse the manga with custom instructions:\n",
"\n",
"\"The provided document is a manga comic book. Most pages do NOT have a title. It does not contain tables. Try to reconstruct the dialogue spoken in a cohesive way.\"\n",
"\n",
"To do so just pass the parsing instruction as a parameter to LlamaParse:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 88ab273e-b2a7-4f84-8e72-e9367cf6b114\n",
"."
]
}
],
"source": [
"parsingInstructionManga = \"\"\"The provided document is a manga comic book. Most pages do NOT have a title.\n",
"It does not contain tables.\n",
"Try to reconstruct the dialogue spoken in a cohesive way.\"\"\"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstructionManga\n",
").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let's see how it compare with page 3! We encourage you to play with the target page and explore other pages. As you will see, the parsing instruction allowed LlamaParse to make sense of the document!\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1M87rXTIZE8d5v7aHmVZVW6gW3eDGq6ks&authuser=0\" />\n",
"\n",
"\n",
"\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"\n",
"\n",
"------------------------------------------------------------\n",
"\n",
"\n",
"# The Asagake Times\n",
"\n",
"Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor?\n",
"\n",
"Do I have the wrong map?\n",
"\n",
"You're looking for the Sanda-cho branch office?\n",
"\n",
"It's next door.\n",
"\n",
"Everybody mistakes us for the office because we are larger.\n",
"\n",
"What Is a Function? 3\n"
]
}
],
"source": [
"target_page = 1\n",
"print(vanilaParsing[0].text.split(\"\\n---\\n\")[target_page])\n",
"print(\"\\n\\n------------------------------------------------------------\\n\\n\")\n",
"print(withInstructionParsing[0].text.split(\"\\n---\\n\")[target_page])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Math - doing more with parsing instuction!\n",
"\n",
"But this manga is about math and full of equations, why not ask the parser to output them in **LaTeX**?\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1tze3xcQ7axVA-vC_iZeAj_GvYcyNuYDa&authuser=0\" />"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 3a055e64-d91e-484e-b9b0-99a2e637c08d\n",
"."
]
}
],
"source": [
"parsingInstructionMangaLatex = \"\"\"The provided document is a manga comic book. Most pages do NOT have a title.\n",
"It does not contain tables.\n",
"Try to reconstruct the dialogue spoken in a cohesive way.\n",
"Output any math equation in LATEX markdown (between $$)\"\"\"\n",
"withLatex = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstructionMangaLatex\n",
").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"\n",
"[Without instruction]------------------------------------------------------------\n",
"\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = lim ε→0 (α(a + ε) + β - (αa + β)) = lim ε→0 α = α|\n",
"| |Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.|\n",
"\n",
"|3.|Lets find the derivative of f(x) = x^2, which appeared in the story. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 ((a + ε)^2 - a^2) / ε = lim (a^2 + 2aε + ε^2 - a^2) / ε = lim (2aε + ε^2) = lim (2a + ε) = 2a|\n",
"| |Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.|\n",
"\n",
"## Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n",
"\n",
"## Chapter 1 Lets Differentiate a Function!\n",
"\n",
"\n",
"[With instruction to output math in LATEX!]------------------------------------------------------------\n",
"\n",
"\n",
"# Derivative of Constant, Linear, or Quadratic Function\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"1. Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha - \\alpha}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} 0 = 0 \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"2. Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(\\alpha + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha(a + \\varepsilon) + \\beta - (\\alpha a + \\beta)}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} \\alpha = \\alpha \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.\n",
"\n",
"3. Lets find the derivative of f(x) = x2. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\left( (a + \\varepsilon)^2 - a^2 \\right) = \\lim_{{\\varepsilon \\to 0}} 2a\\varepsilon + \\varepsilon = \\lim_{{\\varepsilon \\to 0}} (2a + \\varepsilon) = 2a \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.\n",
"\n",
"### Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n"
]
}
],
"source": [
"target_page = 2\n",
"print(\n",
" \"\\n\\n[Without instruction]------------------------------------------------------------\\n\\n\"\n",
")\n",
"print(vanilaParsing[0].text.split(\"\\n---\\n\")[target_page])\n",
"print(\n",
" \"\\n\\n[With instruction to output math in LATEX!]------------------------------------------------------------\\n\\n\"\n",
")\n",
"print(withLatex[0].text.split(\"\\n---\\n\")[target_page])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"And here is the result as rendered by https://upmath.me/ .\n",
"\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1qGo5bMGYOiIC9MnprcgEByaYjU9YII2Q&authuser=0\" />\n",
"\n",
"\n",
"Over this short notebook we saw how to use parsing instructions to increase the quality and accuracy of parsing with LLamaParse!"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
+7 -52
View File
@@ -55,10 +55,6 @@
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -80,25 +76,7 @@
"execution_count": null,
"id": "IjtKDQRLrylI",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 18:54:24-- https://arxiv.org/pdf/2409.18486\n",
"Resolving arxiv.org (arxiv.org)... 151.101.67.42, 151.101.131.42, 151.101.3.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.67.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 13986265 (13M) [application/pdf]\n",
"Saving to: o1.pdf\n",
"\n",
"o1.pdf 100%[===================>] 13.34M 11.8MB/s in 1.1s \n",
"\n",
"2024-12-05 18:54:26 (11.8 MB/s) - o1.pdf saved [13986265/13986265]\n",
"\n"
]
}
],
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/2409.18486\" -O \"o1.pdf\""
]
@@ -118,25 +96,6 @@
"Using your own API key may incur additional costs from your model provider and could result in failed pages or documents if you do not have sufficient usage limits."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"md\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
"cell_type": "markdown",
"id": "1b5d6da6",
@@ -163,15 +122,13 @@
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model_name=\"anthropic-sonnet-3.5\",\n",
" target_pages=\"24\"\n",
" # invalidate_cache=True\n",
")\n",
"json_objs = parser.get_json_result(\"o1.pdf\")\n",
"json_list = json_objs[0][\"pages\"]\n",
"docs = get_text_nodes(json_list)"
"result = await parser.aparse(\"o1.pdf\")\n",
"nodes = result.get_text_nodes(split_by_page=False)"
]
},
{
@@ -202,15 +159,13 @@
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model=\"openai-gpt4o\",\n",
" target_pages=\"24\",\n",
" # invalidate_cache=True\n",
")\n",
"json_objs_gpt4o = parser_gpt4o.get_json_result(\"o1.pdf\")\n",
"json_list_gpt4o = json_objs_gpt4o[0][\"pages\"]\n",
"docs_gpt4o = get_text_nodes(json_list_gpt4o)"
"result = await parser_gpt4o.aparse(\"o1.pdf\")\n",
"nodes = result.get_markdown_nodes(split_by_page=False)"
]
},
{
@@ -268,7 +223,7 @@
],
"source": [
"# using Sonnet-3.5\n",
"print(docs[0].get_content(metadata_mode=\"all\"))"
"print(nodes[0].get_content(metadata_mode=\"all\"))"
]
},
{
@@ -327,7 +282,7 @@
],
"source": [
"# using GPT-4o\n",
"print(docs_gpt4o[0].get_content(metadata_mode=\"all\"))"
"print(nodes[0].get_content(metadata_mode=\"all\"))"
]
}
],
@@ -47,11 +47,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -71,25 +66,7 @@
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 11:40:59-- https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf\n",
"Resolving raw.githubusercontent.com (raw.githubusercontent.com)... 2606:50c0:8000::154, 2606:50c0:8002::154, 2606:50c0:8003::154, ...\n",
"Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|2606:50c0:8000::154|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 1880483 (1.8M) [application/octet-stream]\n",
"Saving to: ./uber_2021.pdf\n",
"\n",
"./uber_2021.pdf 100%[===================>] 1.79M --.-KB/s in 0.1s \n",
"\n",
"2024-12-05 11:40:59 (14.2 MB/s) - ./uber_2021.pdf saved [1880483/1880483]\n",
"\n"
]
}
],
"outputs": [],
"source": [
"!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf' -O './uber_2021.pdf'"
]
@@ -119,9 +96,10 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(target_pages=\"0,1,2\", result_type=\"markdown\")\n",
"parser = LlamaParse(target_pages=\"0,1,2\")\n",
"\n",
"documents = parser.load_data(\"./uber_2021.pdf\")"
"results = await parser.aparse(\"./uber_2021.pdf\")\n",
"documents = results.get_text_documents(split_by_page=True)"
]
},
{
-367
View File
@@ -1,367 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# RAG for Table Comparisons with LlamaParse + LlamaIndex\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_table_comparisons.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you how to do comparisons across both tabular and text data across multiple PDF documents.\n",
"\n",
"We load in multiple PDFs with embedded tables (2021 and 2020 10K filings for Apple) using LlamaParse, parse each into a hierarchy of tables/text objects, define a recursive retriever over each, and then compose both with a SubQuestionQueryEngine."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup\n",
"\n",
"Install core packages, download files, parse documents."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-index-core\n",
"%pip install llama-index-embeddings-openai\n",
"%pip install llama-index-question-gen-openai\n",
"%pip install llama-index-postprocessor-flag-embedding-reranker\n",
"%pip install git+https://github.com/FlagOpen/FlagEmbedding.git\n",
"%pip install llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2020/ar/_10-K-2020-(As-Filed).pdf\" -O apple_2020_10k.pdf\n",
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2021/q4/_10-K-2021-(As-Filed).pdf\" -O apple_2021_10k.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Some OpenAI and LlamaParse details"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-\"\n",
"\n",
"# Using OpenAI API for embeddings/llms\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex\n",
"from llama_index.core import Settings\n",
"\n",
"embed_model = OpenAIEmbedding(model=\"text-embedding-3-small\")\n",
"llm = OpenAI(model=\"gpt-3.5-turbo-0125\")\n",
"\n",
"Settings.llm = llm\n",
"Settings.embed_model = embed_model"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using brand new `LlamaParse` PDF reader for PDF Parsing\n",
"\n",
"we also compare two different retrieval/query engine strategies:\n",
"1. Using raw Markdown text as nodes for building index and apply simple query engine for generating the results;\n",
"2. Using `MarkdownElementNodeParser` for parsing the `LlamaParse` output Markdown results and building recursive retriever query engine for generation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"docs_2021 = LlamaParse(result_type=\"markdown\").load_data(\"./apple_2021_10k.pdf\")\n",
"docs_2020 = LlamaParse(result_type=\"markdown\").load_data(\"./apple_2020_10k.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Create Recursive Retriever over each Document\n",
"\n",
"We define a function to get a recursive retriever from each document. The steps are the following:\n",
"- Hierarchically parse the document using our `MarkdownElementNodeParser`, which will embed/summarize embedded tables.\n",
"- Load into a vector store. Under the hood we will automatically store links between nodes (e.g. table summary to table text).\n",
"- Get a query engine over the vector store, which performs retrieval/synthesis. Under the hood we will automatically perform recursive retrieval if there are links."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.node_parser import MarkdownElementNodeParser\n",
"\n",
"node_parser = MarkdownElementNodeParser(\n",
" llm=OpenAI(model=\"gpt-3.5-turbo-0125\"), num_workers=8\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import pickle\n",
"from llama_index.postprocessor.flag_embedding_reranker import (\n",
" FlagEmbeddingReranker,\n",
")\n",
"\n",
"reranker = FlagEmbeddingReranker(\n",
" top_n=5,\n",
" model=\"BAAI/bge-reranker-large\",\n",
")\n",
"\n",
"\n",
"def create_query_engine_over_doc(docs, nodes_save_path=None):\n",
" \"\"\"Big function to go from document path -> recursive retriever.\"\"\"\n",
" if nodes_save_path is not None and os.path.exists(nodes_save_path):\n",
" raw_nodes = pickle.load(open(nodes_save_path, \"rb\"))\n",
" else:\n",
" raw_nodes = node_parser.get_nodes_from_documents(docs)\n",
" if nodes_save_path is not None:\n",
" pickle.dump(raw_nodes, open(nodes_save_path, \"wb\"))\n",
"\n",
" base_nodes, objects = node_parser.get_nodes_and_objects(raw_nodes)\n",
"\n",
" ### Construct Retrievers\n",
" # construct top-level vector index + query engine\n",
" vector_index = VectorStoreIndex(nodes=base_nodes + objects)\n",
" query_engine = vector_index.as_query_engine(\n",
" similarity_top_k=15, node_postprocessors=[reranker]\n",
" )\n",
" return query_engine, base_nodes"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"query_engine_2021, nodes_2021 = create_query_engine_over_doc(\n",
" docs_2021, nodes_save_path=\"2021_nodes.pkl\"\n",
")\n",
"query_engine_2020, nodes_2020 = create_query_engine_over_doc(\n",
" docs_2020, nodes_save_path=\"2020_nodes.pkl\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.tools import QueryEngineTool, ToolMetadata\n",
"from llama_index.core.query_engine import SubQuestionQueryEngine\n",
"\n",
"\n",
"# setup base query engine as tool\n",
"query_engine_tools = [\n",
" QueryEngineTool(\n",
" query_engine=query_engine_2021,\n",
" metadata=ToolMetadata(\n",
" name=\"apple_2021_10k\",\n",
" description=(\"Provides information about Apple financials for year 2021\"),\n",
" ),\n",
" ),\n",
" QueryEngineTool(\n",
" query_engine=query_engine_2020,\n",
" metadata=ToolMetadata(\n",
" name=\"apple_2020_10k\",\n",
" description=(\"Provides information about Apple financials for year 2020\"),\n",
" ),\n",
" ),\n",
"]\n",
"\n",
"sub_query_engine = SubQuestionQueryEngine.from_defaults(\n",
" query_engine_tools=query_engine_tools,\n",
" llm=llm,\n",
" use_async=True,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Try out Some Comparisons"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 4 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What are the deferred assets in 2021?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2021_10k] Q: What are the deferred liabilities in 2021?\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203m[apple_2020_10k] Q: What are the deferred assets in 2020?\n",
"\u001b[0m\u001b[1;3;38;2;155;135;227m[apple_2020_10k] Q: What are the deferred liabilities in 2020?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2021_10k] A: $7,200\n",
"\u001b[0m\u001b[1;3;38;2;155;135;227m[apple_2020_10k] A: $10,138\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: $25,176 million\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203m[apple_2020_10k] A: $19,336\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the deferred assets and liabilities in 2021 with 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"In 2021, the deferred assets increased by $5,840 million compared to 2020, while the deferred liabilities decreased by $2,938 million in the same period.\n"
]
}
],
"source": [
"print(str(response))"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 2 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What is the total number of RSUs in Apple's 2021 financials?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] Q: What is the total number of RSUs in Apple's 2020 financials?\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: The total number of RSUs in Apple's 2021 financials is 240,427.\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] A: The total number of RSUs in Apple's 2020 financials is 310,778.\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the total number of RSUs in 2021 and 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 2 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What are the risk factors mentioned in the 2021 financial report of Apple?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] Q: What are the risk factors mentioned in the 2020 financial report of Apple?\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: The risk factors mentioned in the 2021 financial report of Apple include risks related to COVID-19, macroeconomic and industry risks, political events, trade and international disputes, natural disasters, public health issues, industrial accidents, credit risk, fluctuations in foreign currency exchange rates, changes in tax rates and legislation, volatility in the price of the company's stock, and exposure to legal proceedings and claims.\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] A: The risk factors mentioned in the 2020 financial report of Apple include the impact of the COVID-19 pandemic on the company's business operations, financial condition, and stock price; global and regional economic conditions affecting demand for products and services; competition in global markets with rapid technological changes; potential disruptions in the supply chain due to industrial accidents or public health issues; information technology system failures or network disruptions affecting business operations; risks associated with confidential information security and potential unauthorized access; fluctuations in quarterly net sales and operating results due to various factors; stock price volatility impacting investor confidence and employee retention; financial performance risks related to changes in foreign currency exchange rates affecting sales and earnings.\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the risk factors in 2021 vs. 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The risk factors mentioned in the 2021 financial report of Apple include risks related to COVID-19, macroeconomic and industry risks, political events, trade and international disputes, natural disasters, public health issues, industrial accidents, credit risk, fluctuations in foreign currency exchange rates, changes in tax rates and legislation, volatility in the price of the company's stock, and exposure to legal proceedings and claims. In contrast, the risk factors mentioned in the 2020 financial report of Apple focused more on the impact of the COVID-19 pandemic on the company's business operations, financial condition, and stock price; global and regional economic conditions affecting demand for products and services; competition in global markets with rapid technological changes; potential disruptions in the supply chain due to industrial accidents or public health issues; information technology system failures or network disruptions affecting business operations; risks associated with confidential information security and potential unauthorized access; fluctuations in quarterly net sales and operating results due to various factors; stock price volatility impacting investor confidence and employee retention; financial performance risks related to changes in foreign currency exchange rates affecting sales and earnings.\n"
]
}
],
"source": [
"print(str(response))"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
+516
View File
@@ -0,0 +1,516 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Table Extraction with LlamaParse\n",
"\n",
"This notebook will show you how to extract tables and save them as CSV files thanks to LlamaParse advanced parsing capabilities."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**1. Install needed dependencies**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"! pip install llama-cloud-services pandas"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**2. Set you LLAMA_CLOUD_API_KEY as env variable**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"LLAMA_CLOUD_API_KEY: ··········\n"
]
}
],
"source": [
"import os\n",
"from getpass import getpass\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = getpass(\"LLAMA_CLOUD_API_KEY: \")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**3. Initialiaze the parser**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"markdown\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**4. Get data**\n",
"\n",
"This is a PDF with _lots_ of tables!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2025-07-16 16:20:41-- https://assets.accessible-digital-documents.com/uploads/2017/01/sample-tables.pdf\n",
"Resolving assets.accessible-digital-documents.com (assets.accessible-digital-documents.com)... 3.166.135.2, 3.166.135.62, 3.166.135.51, ...\n",
"Connecting to assets.accessible-digital-documents.com (assets.accessible-digital-documents.com)|3.166.135.2|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 145494 (142K) [application/pdf]\n",
"Saving to: sample-tables.pdf\n",
"\n",
"sample-tables.pdf 100%[===================>] 142.08K --.-KB/s in 0.04s \n",
"\n",
"2025-07-16 16:20:41 (3.72 MB/s) - sample-tables.pdf saved [145494/145494]\n",
"\n"
]
}
],
"source": [
"! wget https://assets.accessible-digital-documents.com/uploads/2017/01/sample-tables.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**5. Parse document**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b53949f7-9017-4b6a-b30c-be6227271ed2\n"
]
}
],
"source": [
"json_result = parser.get_json_result(\"sample-tables.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**6. Get tables!**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"tables = parser.get_tables(json_result, \"tables/\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**7. Load tables**\n",
"\n",
"Let's show one example table!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"application/vnd.google.colaboratory.intrinsic+json": {
"summary": "{\n \"name\": \"display(df\",\n \"rows\": 8,\n \"fields\": [\n {\n \"column\": \"Rainfall\",\n \"properties\": {\n \"dtype\": \"string\",\n \"num_unique_values\": 5,\n \"samples\": [\n \"Average\",\n \"\",\n \"24 hour high\"\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Americas\",\n \"properties\": {\n \"dtype\": \"number\",\n \"std\": 908,\n \"min\": 9,\n \"max\": 2010,\n \"num_unique_values\": 8,\n \"samples\": [\n 104,\n 133,\n 2010\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Asia\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 201.0,\n 28.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Europe\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 193.0,\n 29.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Africa\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 144.0,\n 20.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n }\n ]\n}",
"type": "dataframe"
},
"text/html": [
"\n",
" <div id=\"df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb\" class=\"colab-df-container\">\n",
" <div>\n",
"<style scoped>\n",
" .dataframe tbody tr th:only-of-type {\n",
" vertical-align: middle;\n",
" }\n",
"\n",
" .dataframe tbody tr th {\n",
" vertical-align: top;\n",
" }\n",
"\n",
" .dataframe thead th {\n",
" text-align: right;\n",
" }\n",
"</style>\n",
"<table border=\"1\" class=\"dataframe\">\n",
" <thead>\n",
" <tr style=\"text-align: right;\">\n",
" <th></th>\n",
" <th>Rainfall</th>\n",
" <th>Americas</th>\n",
" <th>Asia</th>\n",
" <th>Europe</th>\n",
" <th>Africa</th>\n",
" </tr>\n",
" </thead>\n",
" <tbody>\n",
" <tr>\n",
" <th>0</th>\n",
" <td>(inches)</td>\n",
" <td>2010</td>\n",
" <td></td>\n",
" <td></td>\n",
" <td></td>\n",
" </tr>\n",
" <tr>\n",
" <th>1</th>\n",
" <td>Average</td>\n",
" <td>104</td>\n",
" <td>201.0</td>\n",
" <td>193.0</td>\n",
" <td>144.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>2</th>\n",
" <td>24 hour high</td>\n",
" <td>15</td>\n",
" <td>26.0</td>\n",
" <td>27.0</td>\n",
" <td>18.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>3</th>\n",
" <td>12 hour high</td>\n",
" <td>9</td>\n",
" <td>10.0</td>\n",
" <td>11.0</td>\n",
" <td>12.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>4</th>\n",
" <td></td>\n",
" <td>2009</td>\n",
" <td></td>\n",
" <td></td>\n",
" <td></td>\n",
" </tr>\n",
" <tr>\n",
" <th>5</th>\n",
" <td>Average</td>\n",
" <td>133</td>\n",
" <td>244.0</td>\n",
" <td>155.0</td>\n",
" <td>166.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>6</th>\n",
" <td>24 hour high</td>\n",
" <td>27</td>\n",
" <td>28.0</td>\n",
" <td>29.0</td>\n",
" <td>20.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>7</th>\n",
" <td>12 hour high</td>\n",
" <td>11</td>\n",
" <td>12.0</td>\n",
" <td>13.0</td>\n",
" <td>16.0</td>\n",
" </tr>\n",
" </tbody>\n",
"</table>\n",
"</div>\n",
" <div class=\"colab-df-buttons\">\n",
"\n",
" <div class=\"colab-df-container\">\n",
" <button class=\"colab-df-convert\" onclick=\"convertToInteractive('df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb')\"\n",
" title=\"Convert this dataframe to an interactive table.\"\n",
" style=\"display:none;\">\n",
"\n",
" <svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\" viewBox=\"0 -960 960 960\">\n",
" <path d=\"M120-120v-720h720v720H120Zm60-500h600v-160H180v160Zm220 220h160v-160H400v160Zm0 220h160v-160H400v160ZM180-400h160v-160H180v160Zm440 0h160v-160H620v160ZM180-180h160v-160H180v160Zm440 0h160v-160H620v160Z\"/>\n",
" </svg>\n",
" </button>\n",
"\n",
" <style>\n",
" .colab-df-container {\n",
" display:flex;\n",
" gap: 12px;\n",
" }\n",
"\n",
" .colab-df-convert {\n",
" background-color: #E8F0FE;\n",
" border: none;\n",
" border-radius: 50%;\n",
" cursor: pointer;\n",
" display: none;\n",
" fill: #1967D2;\n",
" height: 32px;\n",
" padding: 0 0 0 0;\n",
" width: 32px;\n",
" }\n",
"\n",
" .colab-df-convert:hover {\n",
" background-color: #E2EBFA;\n",
" box-shadow: 0px 1px 2px rgba(60, 64, 67, 0.3), 0px 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
" fill: #174EA6;\n",
" }\n",
"\n",
" .colab-df-buttons div {\n",
" margin-bottom: 4px;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-convert {\n",
" background-color: #3B4455;\n",
" fill: #D2E3FC;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-convert:hover {\n",
" background-color: #434B5C;\n",
" box-shadow: 0px 1px 3px 1px rgba(0, 0, 0, 0.15);\n",
" filter: drop-shadow(0px 1px 2px rgba(0, 0, 0, 0.3));\n",
" fill: #FFFFFF;\n",
" }\n",
" </style>\n",
"\n",
" <script>\n",
" const buttonEl =\n",
" document.querySelector('#df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb button.colab-df-convert');\n",
" buttonEl.style.display =\n",
" google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
"\n",
" async function convertToInteractive(key) {\n",
" const element = document.querySelector('#df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb');\n",
" const dataTable =\n",
" await google.colab.kernel.invokeFunction('convertToInteractive',\n",
" [key], {});\n",
" if (!dataTable) return;\n",
"\n",
" const docLinkHtml = 'Like what you see? Visit the ' +\n",
" '<a target=\"_blank\" href=https://colab.research.google.com/notebooks/data_table.ipynb>data table notebook</a>'\n",
" + ' to learn more about interactive tables.';\n",
" element.innerHTML = '';\n",
" dataTable['output_type'] = 'display_data';\n",
" await google.colab.output.renderOutput(dataTable, element);\n",
" const docLink = document.createElement('div');\n",
" docLink.innerHTML = docLinkHtml;\n",
" element.appendChild(docLink);\n",
" }\n",
" </script>\n",
" </div>\n",
"\n",
"\n",
" <div id=\"df-54b2aa43-838b-47d3-9209-2fb18153cf87\">\n",
" <button class=\"colab-df-quickchart\" onclick=\"quickchart('df-54b2aa43-838b-47d3-9209-2fb18153cf87')\"\n",
" title=\"Suggest charts\"\n",
" style=\"display:none;\">\n",
"\n",
"<svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\"viewBox=\"0 0 24 24\"\n",
" width=\"24px\">\n",
" <g>\n",
" <path d=\"M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zM9 17H7v-7h2v7zm4 0h-2V7h2v10zm4 0h-2v-4h2v4z\"/>\n",
" </g>\n",
"</svg>\n",
" </button>\n",
"\n",
"<style>\n",
" .colab-df-quickchart {\n",
" --bg-color: #E8F0FE;\n",
" --fill-color: #1967D2;\n",
" --hover-bg-color: #E2EBFA;\n",
" --hover-fill-color: #174EA6;\n",
" --disabled-fill-color: #AAA;\n",
" --disabled-bg-color: #DDD;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-quickchart {\n",
" --bg-color: #3B4455;\n",
" --fill-color: #D2E3FC;\n",
" --hover-bg-color: #434B5C;\n",
" --hover-fill-color: #FFFFFF;\n",
" --disabled-bg-color: #3B4455;\n",
" --disabled-fill-color: #666;\n",
" }\n",
"\n",
" .colab-df-quickchart {\n",
" background-color: var(--bg-color);\n",
" border: none;\n",
" border-radius: 50%;\n",
" cursor: pointer;\n",
" display: none;\n",
" fill: var(--fill-color);\n",
" height: 32px;\n",
" padding: 0;\n",
" width: 32px;\n",
" }\n",
"\n",
" .colab-df-quickchart:hover {\n",
" background-color: var(--hover-bg-color);\n",
" box-shadow: 0 1px 2px rgba(60, 64, 67, 0.3), 0 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
" fill: var(--button-hover-fill-color);\n",
" }\n",
"\n",
" .colab-df-quickchart-complete:disabled,\n",
" .colab-df-quickchart-complete:disabled:hover {\n",
" background-color: var(--disabled-bg-color);\n",
" fill: var(--disabled-fill-color);\n",
" box-shadow: none;\n",
" }\n",
"\n",
" .colab-df-spinner {\n",
" border: 2px solid var(--fill-color);\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" animation:\n",
" spin 1s steps(1) infinite;\n",
" }\n",
"\n",
" @keyframes spin {\n",
" 0% {\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" border-left-color: var(--fill-color);\n",
" }\n",
" 20% {\n",
" border-color: transparent;\n",
" border-left-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" }\n",
" 30% {\n",
" border-color: transparent;\n",
" border-left-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" border-right-color: var(--fill-color);\n",
" }\n",
" 40% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" }\n",
" 60% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" }\n",
" 80% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" border-bottom-color: var(--fill-color);\n",
" }\n",
" 90% {\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" }\n",
" }\n",
"</style>\n",
"\n",
" <script>\n",
" async function quickchart(key) {\n",
" const quickchartButtonEl =\n",
" document.querySelector('#' + key + ' button');\n",
" quickchartButtonEl.disabled = true; // To prevent multiple clicks.\n",
" quickchartButtonEl.classList.add('colab-df-spinner');\n",
" try {\n",
" const charts = await google.colab.kernel.invokeFunction(\n",
" 'suggestCharts', [key], {});\n",
" } catch (error) {\n",
" console.error('Error during call to suggestCharts:', error);\n",
" }\n",
" quickchartButtonEl.classList.remove('colab-df-spinner');\n",
" quickchartButtonEl.classList.add('colab-df-quickchart-complete');\n",
" }\n",
" (() => {\n",
" let quickchartButtonEl =\n",
" document.querySelector('#df-54b2aa43-838b-47d3-9209-2fb18153cf87 button');\n",
" quickchartButtonEl.style.display =\n",
" google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
" })();\n",
" </script>\n",
" </div>\n",
"\n",
" </div>\n",
" </div>\n"
],
"text/plain": [
" Rainfall Americas Asia Europe Africa\n",
"0 (inches) 2010 \n",
"1 Average 104 201.0 193.0 144.0\n",
"2 24 hour high 15 26.0 27.0 18.0\n",
"3 12 hour high 9 10.0 11.0 12.0\n",
"4 2009 \n",
"5 Average 133 244.0 155.0 166.0\n",
"6 24 hour high 27 28.0 29.0 20.0\n",
"7 12 hour high 11 12.0 13.0 16.0"
]
},
"metadata": {},
"output_type": "display_data"
}
],
"source": [
"import pandas as pd\n",
"from IPython.display import display\n",
"\n",
"df = pd.read_csv(\n",
" \"/content/tables/table_2025_16_07_16_30_01_569.csv\",\n",
")\n",
"display(df.fillna(\"\"))"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
File diff suppressed because one or more lines are too long
+36 -7
View File
@@ -1,6 +1,6 @@
# LlamaExtract
LlamaExtract provides a simple API for extracting structured data from unstructured documents like PDFs, text files and images (upcoming).
LlamaExtract provides a simple API for extracting structured data from unstructured documents like PDFs, text files and images.
## Quick Start
@@ -97,11 +97,12 @@ _LlamaExtract only supports a subset of the JSON Schema specification._ While li
be sufficient for a wide variety of use-cases.
- All fields are required by default. Nullable fields must be explicitly marked as such,
using `"anyOf"` with a `"null"` type. See `"start_date"` field above.
- Root node must be of type `"object"`.
using `anyOf` with a `null` type. See `"start_date"` field above.
- Root node must be of type `object`.
- Schema nesting must be limited to within 5 levels.
- The important fields are key names/titles, type and description. Fields for
formatting, default values, etc. are not supported.
formatting, default values, etc. are **not supported**. If you need these, you can add the
restrictions to your field description and/or use a post-processing step. e.g. default values can be supported by making a field optional and then setting `"null"` values from the extraction result to the default value.
- There are other restrictions on number of keys, size of the schema, etc. that you may
hit for complex extraction use cases. In such cases, it is worth thinking how to restructure
your extraction workflow to fit within these constraints, e.g. by extracting subset of fields
@@ -109,6 +110,23 @@ be sufficient for a wide variety of use-cases.
## Other Extraction APIs
### Extraction over bytes or text
You can use the `SourceText` class to extract from bytes or text directly without using a file. If passing the file bytes,
you will need to pass the filename to the `SourceText` class.
```python
with open("resume.pdf", "rb") as f:
file_bytes = f.read()
result = test_agent.extract(SourceText(file=file_bytes, filename="resume.pdf"))
```
```python
result = test_agent.extract(
SourceText(text_content="Candidate Name: Jane Doe")
)
```
### Batch Processing
Process multiple files asynchronously:
@@ -159,16 +177,18 @@ pip install llama-extract==0.1.0
## Tips & Best Practices
At the core of LlamaExtract is the schema, which defines the structure of the data you want to extract from your documents.
1. **Schema Design**:
- Try to limit schema nesting to 3-4 levels.
- Make fields optional when data might not always be present. Having required fields may force the model
to hallucinate when these fields are not present in the documents.
- When you want to extract a variable number of entities, use an `array` type. Note that you cannot use
- When you want to extract a variable number of entities, use an `array` type. However, note that you cannot use
an `array` type for the root node.
- Use descriptive field names and detailed descriptions. Use descriptions to pass formatting
instructions or few-shot examples.
- Start simple and iteratively build your schema to incorporate requirements.
- Above all, start simple and iteratively build your schema to incorporate requirements.
2. **Running Extractions**:
- Note that resetting `agent.schema` will not save the schema to the database,
@@ -177,7 +197,16 @@ pip install llama-extract==0.1.0
part of `job.error` or `extraction_run.error` fields for debugging.
- Consider async operations (`queue_extraction`) for large-scale extraction once you have finalized your schema.
### Hitting "The response was too long to be processed" Error
This implies that the extraction response is hitting output token limits of the LLM. In such cases, it is worth rethinking the design of your schema to enable a more efficient/scalable extraction. e.g.
- Instead of one field that extracts a complex object, you can use multiple fields to distribute the extraction logic.
- You can also use multiple schemas to extract different subsets of fields from the same document and merge them later.
Another option (orthogonal to the above) is to break the document into smaller sections and extract from each section individually, when possible. LlamaExtract will in most cases be able to handle both document and schema chunking automatically, but there are cases where you may need to do this manually.
## Additional Resources
- [Example Notebook](examples/resume_screening.ipynb) - Detailed walkthrough of resume parsing
- [Example Notebook](examples/extract/resume_screening.ipynb) - Detailed walkthrough of resume parsing
- [Discord Community](https://discord.com/invite/eN6D2HQ4aX) - Get help and share feedback
+86
View File
@@ -0,0 +1,86 @@
# LlamaCloud Index + Retriever
LlamaCloud is a new generation of managed parsing, ingestion, and retrieval services, designed to bring production-grade context-augmentation to your LLM and RAG applications.
Currently, LlamaCloud supports
- Managed Ingestion API, handling parsing and document management
- Managed Retrieval API, configuring optimal retrieval for your RAG system
## Access
We are opening up a private beta to a limited set of enterprise partners for the managed ingestion and retrieval API. If youre interested in centralizing your data pipelines and spending more time working on your actual RAG use cases, come [talk to us.](https://www.llamaindex.ai/contact)
If you have access to LlamaCloud, you can visit [LlamaCloud](https://cloud.llamaindex.ai) to sign in and get an API key.
## Setup
First, make sure you have the latest LlamaIndex version installed.
```
pip uninstall llama-index # run this if upgrading from v0.9.x or older
pip install -U llama-index --upgrade --no-cache-dir --force-reinstall
```
The `llama-index-indices-managed-llama-cloud` package is included with the above install, but you can also install directly
```
pip install -U llama-index-indices-managed-llama-cloud
```
## Usage
You can create an index on LlamaCloud using the following code. By default, new indexes use managed embeddings (OpenAI text-embedding-3-small, 1536 dimensions, 1 credit/page):
```python
import os
os.environ[
"LLAMA_CLOUD_API_KEY"
] = "llx-..." # can provide API-key in env or in the constructor later on
from llama_index.core import SimpleDirectoryReader
from llama_cloud_services import LlamaCloudIndex
# create a new index (uses managed embeddings by default)
index = LlamaCloudIndex.from_documents(
documents,
"my_first_index",
project_name="default",
api_key="llx-...",
verbose=True,
)
# connect to an existing index
index = LlamaCloudIndex("my_first_index", project_name="default")
```
You can also configure a retriever for managed retrieval:
```python
# from the existing index
index.as_retriever()
# from scratch
from llama_index.indices.managed.llama_cloud import LlamaCloudRetriever
retriever = LlamaCloudRetriever("my_first_index", project_name="default")
```
And of course, you can use other index shortcuts to get use out of your new managed index:
```python
query_engine = index.as_query_engine(llm=llm)
chat_engine = index.as_chat_engine(llm=llm)
```
## Retriever Settings
A full list of retriever settings/kwargs is below:
- `dense_similarity_top_k`: Optional[int] -- If greater than 0, retrieve `k` nodes using dense retrieval
- `sparse_similarity_top_k`: Optional[int] -- If greater than 0, retrieve `k` nodes using sparse retrieval
- `enable_reranking`: Optional[bool] -- Whether to enable reranking or not. Sacrifices some speed for accuracy
- `rerank_top_n`: Optional[int] -- The number of nodes to return after reranking initial retrieval results
- `alpha` Optional[float] -- The weighting between dense and sparse retrieval. 1 = Full dense retrieval, 0 = Full sparse retrieval.
-7
View File
@@ -1,7 +0,0 @@
from llama_cloud_services.extract.extract import (
LlamaExtract,
ExtractionAgent,
SourceText,
)
__all__ = ["LlamaExtract", "ExtractionAgent", "SourceText"]
-3
View File
@@ -1,3 +0,0 @@
from llama_cloud_services.parse.base import LlamaParse, ResultType
__all__ = ["LlamaParse", "ResultType"]
-211
View File
@@ -1,211 +0,0 @@
from enum import Enum
# Asyncio error messages
nest_asyncio_err = "cannot be called from a running event loop"
nest_asyncio_msg = "The event loop is already running. Add `import nest_asyncio; nest_asyncio.apply()` to your code to fix this issue."
class ResultType(str, Enum):
"""The result type for the parser."""
TXT = "text"
MD = "markdown"
JSON = "json"
STRUCTURED = "structured"
class ParsingMode(str, Enum):
"""The parsing mode for the parser."""
parse_page_without_llm = "parse_page_without_llm"
parse_page_with_llm = "parse_page_with_llm"
parse_page_with_lvm = "parse_page_with_lvm"
parse_page_with_agent = "parse_page_with_agent"
parse_document_with_llm = "parse_document_with_llm"
class Language(str, Enum):
BAZA = "abq"
ADYGHE = "ady"
AFRIKAANS = "af"
ANGIKA = "ang"
ARABIC = "ar"
ASSAMESE = "as"
AVAR = "ava"
AZERBAIJANI = "az"
BELARUSIAN = "be"
BULGARIAN = "bg"
BIHARI = "bh"
BHOJPURI = "bho"
BENGALI = "bn"
BOSNIAN = "bs"
SIMPLIFIED_CHINESE = "ch_sim"
TRADITIONAL_CHINESE = "ch_tra"
CHECHEN = "che"
CZECH = "cs"
WELSH = "cy"
DANISH = "da"
DARGWA = "dar"
GERMAN = "de"
ENGLISH = "en"
SPANISH = "es"
ESTONIAN = "et"
PERSIAN_FARSI = "fa"
FRENCH = "fr"
IRISH = "ga"
GOAN_KONKANI = "gom"
HINDI = "hi"
CROATIAN = "hr"
HUNGARIAN = "hu"
INDONESIAN = "id"
INGUSH = "inh"
ICELANDIC = "is"
ITALIAN = "it"
JAPANESE = "ja"
KABARDIAN = "kbd"
KANNADA = "kn"
KOREAN = "ko"
KURDISH = "ku"
LATIN = "la"
LAK = "lbe"
LEZGHIAN = "lez"
LITHUANIAN = "lt"
LATVIAN = "lv"
MAGAHI = "mah"
MAITHILI = "mai"
MAORI = "mi"
MONGOLIAN = "mn"
MARATHI = "mr"
MALAY = "ms"
MALTESE = "mt"
NEPALI = "ne"
NEWARI = "new"
DUTCH = "nl"
NORWEGIAN = "no"
OCCITAN = "oc"
PALI = "pi"
POLISH = "pl"
PORTUGUESE = "pt"
ROMANIAN = "ro"
RUSSIAN = "ru"
SERBIAN_CYRILLIC = "rs_cyrillic"
SERBIAN_LATIN = "rs_latin"
NAGPURI = "sck"
SLOVAK = "sk"
SLOVENIAN = "sl"
ALBANIAN = "sq"
SWEDISH = "sv"
SWAHILI = "sw"
TAMIL = "ta"
TABASSARAN = "tab"
TELUGU = "te"
THAI = "th"
TAJIK = "tjk"
TAGALOG = "tl"
TURKISH = "tr"
UYGHUR = "ug"
UKRAINIAN = "uk"
URDU = "ur"
UZBEK = "uz"
VIETNAMESE = "vi"
SUPPORTED_FILE_TYPES = [
".pdf",
# document and presentations
".602",
".abw",
".cgm",
".cwk",
".doc",
".docx",
".docm",
".dot",
".dotm",
".hwp",
".key",
".lwp",
".mw",
".mcw",
".pages",
".pbd",
".ppt",
".pptm",
".pptx",
".pot",
".potm",
".potx",
".rtf",
".sda",
".sdd",
".sdp",
".sdw",
".sgl",
".sti",
".sxi",
".sxw",
".stw",
".sxg",
".txt",
".uof",
".uop",
".uot",
".vor",
".wpd",
".wps",
".xml",
".zabw",
".epub",
# images
".jpg",
".jpeg",
".png",
".gif",
".bmp",
".svg",
".tiff",
".webp",
# web
".htm",
".html",
# spreadsheets
".xlsx",
".xls",
".xlsm",
".xlsb",
".xlw",
".csv",
".dif",
".sylk",
".slk",
".prn",
".numbers",
".et",
".ods",
".fods",
".uos1",
".uos2",
".dbf",
".wk1",
".wk2",
".wk3",
".wk4",
".wks",
".123",
".wq1",
".wq2",
".wb1",
".wb2",
".wb3",
".qpw",
".xlr",
".eth",
".tsv",
".mp3",
".mp4",
".mpeg",
".mpga",
".m4a",
".wav",
".webm",
]
-3
View File
@@ -1,3 +0,0 @@
from llama_cloud_services.parse import LlamaParse, ResultType
__all__ = ["LlamaParse", "ResultType"]
-24
View File
@@ -1,24 +0,0 @@
[build-system]
requires = ["poetry-core"]
build-backend = "poetry.core.masonry.api"
[tool.poetry]
name = "llama-parse"
version = "0.6.9"
description = "Parse files into RAG-Optimized formats."
authors = ["Logan Markewich <logan@llamaindex.ai>"]
license = "MIT"
readme = "README.md"
packages = [{include = "llama_parse"}]
[tool.poetry.dependencies]
python = ">=3.9,<4.0"
llama-cloud-services = ">=0.6.9"
[tool.poetry.group.dev.dependencies]
pytest = "^8.0.0"
pytest-asyncio = "*"
ipykernel = "^6.29.0"
[tool.poetry.scripts]
llama-parse = "llama_parse.cli.main:parse"
+46 -25
View File
@@ -25,6 +25,8 @@ Then, install the package:
`pip install llama-cloud-services`
## CLI Usage
Now you can parse your first PDF file using the command line interface. Use the command `llama-parse [file_paths]`. See the help text with `llama-parse --help`.
```bash
@@ -40,50 +42,72 @@ llama-parse my_file.pdf --result-type markdown --output-file output.md
llama-parse my_file.pdf --output-raw-json --output-file output.json
```
## Python Usage
You can also create simple scripts:
```python
import nest_asyncio
nest_asyncio.apply()
from llama_cloud_services import LlamaParse
parser = LlamaParse(
api_key="llx-...", # can also be set in your env as LLAMA_CLOUD_API_KEY
result_type="markdown", # "markdown" and "text" are available
num_workers=4, # if multiple files passed, split in `num_workers` API calls
verbose=True,
language="en", # Optionally you can define a language, default=en
)
# sync
documents = parser.load_data("./my_file.pdf")
result = parser.parse("./my_file.pdf")
# sync batch
documents = parser.load_data(["./my_file1.pdf", "./my_file2.pdf"])
results = parser.parse(["./my_file1.pdf", "./my_file2.pdf"])
# async
documents = await parser.aload_data("./my_file.pdf")
result = await parser.aparse("./my_file.pdf")
# async batch
documents = await parser.aload_data(["./my_file1.pdf", "./my_file2.pdf"])
results = await parser.aparse(["./my_file1.pdf", "./my_file2.pdf"])
```
## Using with file object
The result object is a fully typed `JobResult` object, and you can interact with it to parse and transform various parts of the result:
```python
# get the llama-index markdown documents
markdown_documents = result.get_markdown_documents(split_by_page=True)
# get the llama-index text documents
text_documents = result.get_text_documents(split_by_page=False)
# get the image documents
image_documents = result.get_image_documents(
include_screenshot_images=True,
include_object_images=False,
# Optional: download the images to a directory
# (default is to return the image bytes in ImageDocument objects)
image_download_dir="./images",
)
# access the raw job result
# Items will vary based on the parser configuration
for page in result.pages:
print(page.text)
print(page.md)
print(page.images)
print(page.layout)
print(page.structuredData)
```
See more details about the result object in the [example notebook](./examples/parse/demo_json_tour.ipynb).
### Using with file object / bytes
You can parse a file object directly:
```python
import nest_asyncio
nest_asyncio.apply()
from llama_cloud_services import LlamaParse
parser = LlamaParse(
api_key="llx-...", # can also be set in your env as LLAMA_CLOUD_API_KEY
result_type="markdown", # "markdown" and "text" are available
num_workers=4, # if multiple files passed, split in `num_workers` API calls
verbose=True,
language="en", # Optionally you can define a language, default=en
@@ -94,24 +118,20 @@ extra_info = {"file_name": file_name}
with open(f"./{file_name}", "rb") as f:
# must provide extra_info with file_name key with passing file object
documents = parser.load_data(f, extra_info=extra_info)
result = parser.parse(f, extra_info=extra_info)
# you can also pass file bytes directly
with open(f"./{file_name}", "rb") as f:
file_bytes = f.read()
# must provide extra_info with file_name key with passing file bytes
documents = parser.load_data(file_bytes, extra_info=extra_info)
result = parser.parse(file_bytes, extra_info=extra_info)
```
## Using with `SimpleDirectoryReader`
### Using with `SimpleDirectoryReader`
You can also integrate the parser as the default PDF loader in `SimpleDirectoryReader`:
```python
import nest_asyncio
nest_asyncio.apply()
from llama_cloud_services import LlamaParse
from llama_index.core import SimpleDirectoryReader
@@ -133,9 +153,10 @@ Full documentation for `SimpleDirectoryReader` can be found on the [LlamaIndex D
Several end-to-end indexing examples can be found in the examples folder
- [Getting Started](examples/parse/demo_basic.ipynb)
- [Advanced RAG Example](examples/parse/demo_advanced.ipynb)
- [Raw API Usage](examples/parse/demo_api.ipynb)
- [Getting Started](./examples/parse/demo_basic.ipynb)
- [Advanced RAG Example](./examples/parse/demo_advanced.ipynb)
- [Raw API Usage](./examples/parse/demo_api.ipynb)
- [Result Object Tour](./examples/parse/demo_json_tour.ipynb)
## Documentation
+1107
View File
File diff suppressed because it is too large Load Diff
Generated
-4411
View File
File diff suppressed because it is too large Load Diff
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2024 LlamaIndex
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
View File
+87
View File
@@ -0,0 +1,87 @@
[![PyPI - Downloads](https://img.shields.io/pypi/dm/llama-cloud-services)](https://pypi.org/project/llama-cloud-services/)
[![GitHub contributors](https://img.shields.io/github/contributors/run-llama/llama_cloud_services)](https://github.com/run-llama/llama_cloud_services/graphs/contributors)
[![Discord](https://img.shields.io/discord/1059199217496772688)](https://discord.gg/dGcwcsnxhU)
# Llama Cloud Services
This repository contains the code for hand-written SDKs and clients for interacting with LlamaCloud.
This includes:
- [LlamaParse](../parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](../report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract](../extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](../index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
## Getting Started
Install the package:
```bash
pip install llama-cloud-services
```
Then, get your API key from [LlamaCloud](https://cloud.llamaindex.ai/).
Then, you can use the services in your code:
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
)
from llama_cloud_services import LlamaParse, LlamaReport, LlamaExtract
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
)
```
See the quickstart guides for each service for more information:
- [LlamaParse](../parse.md)
- [LlamaReport (beta/invite-only)](../report.md)
- [LlamaExtract](../extract.md)
- [LlamaCloud Index](../index.md)
## Switch to EU SaaS 🇪🇺
If you are interested in using LlamaCloud services in the EU, you can adjust your base URL to `https://api.cloud.eu.llamaindex.ai`.
You can also create your API key in the EU region [here](https://cloud.eu.llamaindex.ai).
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
EU_BASE_URL,
)
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
project_name="default",
api_key="YOUR_API_KEY",
base_url=EU_BASE_URL,
)
```
## Documentation
You can see complete SDK and API documentation for each service on [our official docs](https://docs.cloud.llamaindex.ai/).
## Terms of Service
See the [Terms of Service Here](../TOS.pdf).
## Get in Touch (LlamaCloud)
You can get in touch with us by following our [contact link](https://www.llamaindex.ai/contact).
@@ -2,6 +2,11 @@ from llama_cloud_services.parse import LlamaParse
from llama_cloud_services.report import ReportClient, LlamaReport
from llama_cloud_services.extract import LlamaExtract, ExtractionAgent
from llama_cloud_services.constants import EU_BASE_URL
from llama_cloud_services.index import (
LlamaCloudCompositeRetriever,
LlamaCloudIndex,
LlamaCloudRetriever,
)
__all__ = [
"LlamaParse",
@@ -10,4 +15,7 @@ __all__ = [
"LlamaExtract",
"ExtractionAgent",
"EU_BASE_URL",
"LlamaCloudIndex",
"LlamaCloudRetriever",
"LlamaCloudCompositeRetriever",
]
@@ -0,0 +1,31 @@
from .schema import (
TypedAgentData,
ExtractedData,
TypedAgentDataItems,
StatusType,
ExtractedT,
AgentDataT,
ComparisonOperator,
parse_extracted_field_metadata,
calculate_overall_confidence,
InvalidExtractionData,
ExtractedFieldMetadata,
ExtractedFieldMetaDataDict,
)
from .client import AsyncAgentDataClient
__all__ = [
"TypedAgentData",
"AsyncAgentDataClient",
"ExtractedData",
"TypedAgentDataItems",
"StatusType",
"ExtractedT",
"AgentDataT",
"ComparisonOperator",
"parse_extracted_field_metadata",
"calculate_overall_confidence",
"InvalidExtractionData",
"ExtractedFieldMetadata",
"ExtractedFieldMetaDataDict",
]
@@ -0,0 +1,274 @@
import os
from typing import Any, Dict, Generic, List, Optional, Type
from llama_cloud.client import AsyncLlamaCloud
from tenacity import (
WrappedFn,
retry,
stop_after_attempt,
wait_exponential,
retry_if_exception_type,
)
import httpx
from .schema import (
AgentDataT,
ComparisonOperator,
TypedAgentData,
TypedAgentDataItems,
TypedAggregateGroup,
TypedAggregateGroupItems,
)
def agent_data_retry(func: WrappedFn) -> WrappedFn:
"""
Decorator that adds automatic retry logic to agent data API calls.
Applies exponential backoff retry strategy for common network-related exceptions:
- Up to 3 retry attempts
- Exponential wait time between 0.5s and 10s
- Retries on timeout, connection, and HTTP status errors
This ensures resilient API communication in distributed environments where
temporary network issues or service unavailability may occur.
"""
return retry(
stop=stop_after_attempt(3),
wait=wait_exponential(min=0.5, max=10),
retry=retry_if_exception_type(
(httpx.TimeoutException, httpx.ConnectError, httpx.HTTPStatusError)
),
)(func)
def get_default_agent_id() -> str:
"""
Retrieve the default agent ID from environment variables.
Returns:
The value of LLAMA_DEPLOY_DEPLOYMENT_NAME environment variable,
or None if not set
Note:
This provides a convenient way to configure agent ID globally
via environment variables instead of passing it explicitly
to each client instance.
"""
return os.getenv("LLAMA_DEPLOY_DEPLOYMENT_NAME") or "_public"
class AsyncAgentDataClient(Generic[AgentDataT]):
"""
Async client for managing agent-generated structured data with type safety.
This client provides a high-level interface for CRUD operations, searching, and
aggregation of structured data created by agents. It enforces type safety by
validating all data against a specified Pydantic model type.
The client is generic over AgentDataT, which must be a Pydantic BaseModel that
defines the structure of your agent's data output.
Example:
```python
from pydantic import BaseModel
from llama_cloud.client import AsyncLlamaCloud
from llama_cloud_services.beta.agent_data import AsyncAgentDataClient
class ExtractedPerson(BaseModel):
name: str
age: int
email: str
# Initialize client
llama_client = AsyncLlamaCloud(token="your-api-key")
agent_client = AsyncAgentDataClient(
client=llama_client,
type=ExtractedPerson,
collection="extracted_people",
agent_url_id="person-extraction-agent"
)
# Create data
person = ExtractedPerson(name="John Doe", age=30, email="john@example.com")
result = await agent_client.create_agent_data(person)
# Search data
results = await agent_client.search(
filter={"age": {"gt": 25}},
order_by="data.name",
page_size=20
)
```
Type Parameters:
AgentDataT: Pydantic BaseModel type that defines the structure of agent data
"""
def __init__(
self,
type: Type[AgentDataT],
collection: str = "default",
agent_url_id: Optional[str] = None,
client: Optional[AsyncLlamaCloud] = None,
token: Optional[str] = None,
base_url: Optional[str] = None,
):
"""
Initialize the AsyncAgentDataClient.
Args:
type: Pydantic BaseModel class that defines the data structure.
All agent data will be validated against this type.
collection: Named collection within the agent for organizing data.
Defaults to "default". Collections allow logical separation of
different data types or workflows within the same agent.
agent_url_id: Unique identifier for the agent. This normally appears in the
url of an agent within the llama cloud platform. If not provided,
will attempt to use the LLAMA_DEPLOY_DEPLOYMENT_NAME environment
variable. Data can only be added to an already existing agent in the
platform.
client: AsyncLlamaCloud client instance for API communication. If not provided, will
construct one from the provided api token and base url
token: Llama Cloud API token. Reads from LLAMA_CLOUD_API_KEY if not provided
base_url: Llama Cloud API token. Reads from LLAMA_CLOUD_BASE_URL if not provided, and
defaults to https://api.cloud.llamaindex.ai
Raises:
ValueError: If agent_url_id is not provided and the
LLAMA_DEPLOY_DEPLOYMENT_NAME environment variable is not set
Note:
The client automatically applies retry logic to all API calls with
exponential backoff for timeout, connection, and HTTP status errors.
"""
self.agent_url_id = agent_url_id or get_default_agent_id()
self.collection = collection
if not client:
client = AsyncLlamaCloud(
token=token or os.getenv("LLAMA_CLOUD_API_KEY"),
base_url=base_url or os.getenv("LLAMA_CLOUD_BASE_URL"),
)
self.client = client
self.type = type
@agent_data_retry
async def get_item(self, item_id: str) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.get_agent_data(
item_id=item_id,
)
return TypedAgentData.from_raw(raw_data, validator=self.type)
@agent_data_retry
async def create_item(self, data: AgentDataT) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.create_agent_data(
agent_slug=self.agent_url_id,
collection=self.collection,
data=data.model_dump(),
)
return TypedAgentData.from_raw(raw_data, validator=self.type)
@agent_data_retry
async def update_item(
self, item_id: str, data: AgentDataT
) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.update_agent_data(
item_id=item_id,
data=data.model_dump(),
)
return TypedAgentData.from_raw(raw_data, validator=self.type)
@agent_data_retry
async def delete_item(self, item_id: str) -> None:
await self.client.beta.delete_agent_data(item_id=item_id)
@agent_data_retry
async def search(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
include_total: bool = False,
) -> TypedAgentDataItems[AgentDataT]:
"""
Search agent data with filtering, sorting, and pagination.
Args:
filter: Filter conditions to apply to the search. Dict mapping field names to FilterOperation objects. Filters only by data fields
Examples:
- {"age": {"gt": 18}} - age greater than 18
- {"status": {"eq": "active"}} - status equals "active"
- {"tags": {"includes": ["python", "ml"]}} - tags include "python" or "ml"
- {"created_at": {"gte": "2024-01-01"}} - created after date
- {"score": {"lt": 100, "gte": 50}} - score between 50 and 100
order_by: Comma delimited list of fields to sort results by. Can order by standard agent fields like created_at, or by data fields. Data fields must be prefixed with "data.". If ordering desceding, use a " desc" suffix.
Examples:
- "data.name desc, created_at" - sort by name in descending order, and then by creation date
page_size: Maximum number of items to return per page. Defaults to 10.
offset: Number of items to skip from the beginning. Defaults to 0.
include_total: Whether to include the total count in the response. Defaults to False to improve performance. It's recommended to only request on the first page.
"""
raw = await self.client.beta.search_agent_data_api_v_1_beta_agent_data_search_post(
agent_slug=self.agent_url_id,
collection=self.collection,
filter=filter,
order_by=order_by,
offset=offset,
page_size=page_size,
include_total=include_total,
)
return TypedAgentDataItems(
items=[
TypedAgentData.from_raw(item, validator=self.type) for item in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@agent_data_retry
async def aggregate(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
group_by: Optional[List[str]] = None,
count: Optional[bool] = None,
first: Optional[bool] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
) -> TypedAggregateGroupItems[AgentDataT]:
"""
Aggregate agent data into groups according to the group_by fields.
Args:
filter: Filter conditions to apply to the search. Dict mapping field names to FilterOperation objects. Filters only by data fields
See search for more details on filtering.
group_by: List of fields to group by. Groups strictly by equality. Can only group by data fields.
Examples:
- ["name"] - group by name
- ["name", "age"] - group by name and age
count: Whether to include the count of items in each group.
first: Whether to include the first item in each group.
order_by: Comma delimited list of fields to sort results by. See search for more details on ordering.
offset: Number of groups to skip from the beginning. Defaults to 0.
page_size: Maximum number of groups to return per page.
"""
raw = await self.client.beta.aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
agent_slug=self.agent_url_id,
collection=self.collection,
page_size=page_size,
filter=filter,
order_by=order_by,
group_by=group_by,
count=count,
first=first,
offset=offset,
)
return TypedAggregateGroupItems(
items=[
TypedAggregateGroup.from_raw(item, validator=self.type)
for item in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@@ -0,0 +1,563 @@
"""
Agent Data API Schema Definitions
This module provides typed wrappers around the raw LlamaCloud agent data API,
enabling type-safe interactions with agent-generated structured data.
The agent data API serves as a persistent storage system for structured data
produced by LlamaCloud agents (particularly extraction agents). It provides
CRUD operations, search capabilities, filtering, and aggregation functionality
for managing agent-generated data at scale.
Key Concepts:
- Agent Slug: Unique identifier for an agent instance
- Collection: Named grouping of data within an agent (defaults to "default"). Data within a collection should be of the same type.
- Agent Data: Individual structured data records with metadata and timestamps
Example Usage:
```python
from pydantic import BaseModel
class Person(BaseModel):
name: str
age: int
client = AsyncAgentDataClient(
client=async_llama_cloud,
type=Person,
collection="people",
agent_url_id="my-extraction-agent-xyz"
)
# Create typed data
person = Person(name="John", age=30)
result = await client.create_agent_data(person)
print(result.data.name) # Type-safe access
```
"""
from datetime import datetime
import numbers
from llama_cloud import ExtractRun
from llama_cloud.types.agent_data import AgentData
from llama_cloud.types.aggregate_group import AggregateGroup
from pydantic import BaseModel, Field, ValidationError
from typing import (
Generic,
List,
Literal,
Optional,
Dict,
Type,
TypeVar,
Union,
Any,
)
# Type variable for user-defined data models
AgentDataT = TypeVar("AgentDataT", bound=BaseModel)
# Type variable for extracted data (can be dict or Pydantic model)
ExtractedT = TypeVar("ExtractedT", bound=Union[BaseModel, dict])
# Status types for extracted data workflow
StatusType = Union[Literal["error", "accepted", "rejected", "pending_review"], str]
ComparisonOperator = Dict[
str, Dict[Literal["gt", "gte", "lt", "lte", "eq", "includes"], Any]
]
class TypedAgentData(BaseModel, Generic[AgentDataT]):
"""
Type-safe wrapper for agent data records.
This class represents a single data record stored in the agent data API,
combining the structured data payload with metadata about when and where
it was created.
Attributes:
id: Unique identifier for this data record
agent_url_id: Identifier of the agent that created this data
collection: Named collection within the agent (used for organization)
data: The actual structured data payload (typed as AgentDataT)
created_at: Timestamp when the record was first created
updated_at: Timestamp when the record was last modified
Example:
```python
# Access typed data
person_data: TypedAgentData[Person] = await client.get_agent_data(id)
print(person_data.data.name) # Type-safe access to Person fields
print(person_data.created_at) # Access metadata
```
"""
id: Optional[str] = Field(description="Unique identifier for this data record")
agent_url_id: str = Field(
description="Identifier of the agent that created this data"
)
collection: Optional[str] = Field(
description="Named collection within the agent for data organization"
)
data: AgentDataT = Field(description="The structured data payload")
created_at: Optional[datetime] = Field(description="When this record was created")
updated_at: Optional[datetime] = Field(
description="When this record was last modified"
)
@classmethod
def from_raw(
cls, raw_data: AgentData, validator: Type[AgentDataT]
) -> "TypedAgentData[AgentDataT]":
"""
Convert raw API response to typed agent data.
Args:
raw_data: Raw agent data from the API
validator: Pydantic model class to validate the data field
Returns:
TypedAgentData instance with validated data
"""
data: AgentDataT = validator.model_validate(raw_data.data)
return cls(
id=raw_data.id,
agent_url_id=raw_data.agent_slug,
collection=raw_data.collection,
data=data,
created_at=raw_data.created_at,
updated_at=raw_data.updated_at,
)
class TypedAgentDataItems(BaseModel, Generic[AgentDataT]):
"""
Paginated collection of agent data records.
This class represents a page of search results from the agent data API,
providing both the data records and pagination metadata.
Attributes:
items: List of agent data records in this page
total: Total number of records matching the query (only present if requested)
has_more: Whether there are more records available beyond this page
Example:
```python
# Search with pagination
results = await client.search(
page_size=10,
include_total=True
)
for item in results.items:
print(item.data.name)
if results.has_more:
# Load next page
next_page = await client.search(
page_size=10,
offset=10
)
```
"""
items: List[TypedAgentData[AgentDataT]] = Field(
description="List of agent data records in this page"
)
total: Optional[int] = Field(
description="Total number of records matching the query (only present if requested)"
)
has_more: bool = Field(
description="Whether there are more records available beyond this page"
)
class ExtractedFieldMetadata(BaseModel):
"""
Metadata for an extracted data field, such as confidence, and citation information.
"""
confidence: Optional[float] = Field(
None,
description="The confidence score for the field, combined with parsing confidence if applicable",
)
extracted_confidence: Optional[float] = Field(
None,
description="The confidence score for the field based on the extracted text only",
)
page_number: Optional[int] = Field(
None, description="The page number that the field occurred on"
)
matching_text: Optional[str] = Field(
None,
description="The original text this field's value was derived from",
)
ExtractedFieldMetaDataDict = Dict[
str, Union[ExtractedFieldMetadata, Dict[str, Any], list[Any]]
]
def parse_extracted_field_metadata(
field_metadata: dict[str, Any],
) -> ExtractedFieldMetaDataDict:
"""
Parse the extracted field metadata into a dictionary of field names to field metadata.
"""
result: ExtractedFieldMetaDataDict = {}
for field_name, field_value in field_metadata.items():
if isinstance(field_value, ExtractedFieldMetadata):
# support running this multiple times
result[field_name] = field_value
elif isinstance(field_value, dict):
if "confidence" in field_value or "citations" in field_value:
try:
validated = ExtractedFieldMetadata.model_validate(field_value)
# grab the citation from the array. This is just an array for backwards compatibility.
if "citations" in field_value and len(field_value["citations"]) > 0:
first_citation = field_value["citations"][0]
if "page_number" in first_citation and isinstance(
first_citation["page_number"], numbers.Number
):
validated.page_number = int(first_citation["page_number"]) # type: ignore
if "matching_text" in first_citation and isinstance(
first_citation["matching_text"], str
):
validated.matching_text = first_citation["matching_text"]
result[field_name] = validated
continue
except ValidationError:
pass
result[field_name] = parse_extracted_field_metadata(field_value)
elif isinstance(field_value, list):
result[field_name] = [
parse_extracted_field_metadata(item) for item in field_value
]
else:
result[field_name] = field_value
return result
class ExtractedData(BaseModel, Generic[ExtractedT]):
"""
Wrapper for extracted data with workflow status tracking.
This class is designed for extraction workflows where data goes through
review and approval stages. It maintains both the original extracted data
and the current state after any modifications.
Attributes:
original_data: The data as originally extracted from the source
data: The current state of the data (may differ from original after edits)
status: Current workflow status (in_review, accepted, rejected, error)
confidence: Confidence scores for individual fields (if available)
file_id: The llamacloud file ID of the file that was used to extract the data
file_name: The name of the file that was used to extract the data
file_hash: A content hash of the file that was used to extract the data, for de-duplication
Status Workflow:
- "pending_review": Initial state, awaiting human review
- "accepted": Data approved and ready for use
- "rejected": Data rejected, needs re-extraction or manual fix
- "error": Processing error occurred
Example:
```python
# Create extracted data for review
extracted = ExtractedData.create(
data=person_data,
status="pending_review",
confidence={"name": 0.95, "age": 0.87}
)
# Later, after review
if extracted.status == "accepted":
# Use the data
process_person(extracted.data)
```
"""
original_data: ExtractedT = Field(
description="The original data that was extracted from the document"
)
data: ExtractedT = Field(
description="The latest state of the data. Will differ if data has been updated"
)
status: StatusType = Field(description="The status of the extracted data")
overall_confidence: Optional[float] = Field(
None,
description="The overall confidence score for the extracted data",
)
field_metadata: ExtractedFieldMetaDataDict = Field(
default_factory=dict,
description="Page links, and perhaps eventually bounding boxes, for individual fields in the extracted data. Structure is expected to have a ",
)
file_id: Optional[str] = Field(
None, description="The ID of the file that was used to extract the data"
)
file_name: Optional[str] = Field(
None, description="The name of the file that was used to extract the data"
)
file_hash: Optional[str] = Field(
None, description="The hash of the file that was used to extract the data"
)
metadata: Optional[Dict[str, Any]] = Field(
default_factory=dict,
description="Additional metadata about the extracted data, such as errors, tokens, etc.",
)
@classmethod
def create(
cls,
data: ExtractedT,
status: StatusType = "pending_review",
field_metadata: ExtractedFieldMetaDataDict = {},
file_id: Optional[str] = None,
file_name: Optional[str] = None,
file_hash: Optional[str] = None,
metadata: Optional[Dict[str, Any]] = None,
) -> "ExtractedData[ExtractedT]":
"""
Create a new ExtractedData instance with sensible defaults.
Args:
extracted_data: The extracted data payload
status: Initial workflow status
field_metadata: Optional confidence scores, citations, and other metadata for fields
file_id: The llamacloud file ID of the file that was used to extract the data
file_name: The name of the file that was used to extract the data
file_hash: A content hash of the file that was used to extract the data, for de-duplication
metadata: Arbitrary additional application-specific data about the extracted data
Returns:
New ExtractedData instance ready for storage
"""
normalized_field_metadata = parse_extracted_field_metadata(field_metadata)
return cls(
original_data=data,
data=data,
status=status,
field_metadata=normalized_field_metadata,
overall_confidence=calculate_overall_confidence(normalized_field_metadata),
file_id=file_id,
file_name=file_name,
file_hash=file_hash,
metadata=metadata or {},
)
@classmethod
def from_extraction_result(
cls,
result: ExtractRun,
schema: Type[ExtractedT],
file_hash: Optional[str] = None,
file_name: Optional[str] = None,
file_id: Optional[str] = None,
status: StatusType = "pending_review",
metadata: Optional[Dict[str, Any]] = None,
) -> "ExtractedData[ExtractedT]":
"""
Create an ExtractedData instance from an extraction result.
"""
file_id = file_id or result.file.id
file_name = file_name or result.file.name
try:
field_metadata = parse_extracted_field_metadata(
result.extraction_metadata.get("field_metadata", {})
)
except ValidationError:
field_metadata = {}
try:
data = schema.model_validate(result.data) # type: ignore
return cls.create(
data=data,
status=status,
field_metadata=field_metadata,
file_id=file_id,
file_name=file_name,
file_hash=file_hash,
metadata=metadata or {},
)
except ValidationError as e:
invalid_item = ExtractedData[Dict[str, Any]].create(
data=result.data or {},
status="error",
field_metadata=field_metadata,
metadata={"extraction_error": str(e), **(metadata or {})},
file_id=file_id,
file_name=file_name,
file_hash=file_hash,
)
raise InvalidExtractionData(invalid_item) from e
class InvalidExtractionData(Exception):
"""
Exception raised when the extracted data does not conform to the schema.
"""
def __init__(self, invalid_item: ExtractedData[Dict[str, Any]]):
self.invalid_item = invalid_item
super().__init__("Not able to parse the extracted data, parsed invalid format")
def calculate_overall_confidence(
metadata: ExtractedFieldMetaDataDict,
) -> Optional[float]:
"""
Calculate the overall confidence score for the extracted data.
"""
numerator, denominator = _calculate_overall_confidence_recursive(metadata)
if denominator == 0:
return None
return numerator / denominator
def _calculate_overall_confidence_recursive(
confidence: Union[ExtractedFieldMetadata, Dict[str, Any], list[Any]],
) -> tuple[float, int]:
"""
Calculate the overall confidence score for the extracted data.
"""
if isinstance(confidence, ExtractedFieldMetadata):
if confidence.confidence is not None:
return confidence.confidence, 1
else:
return 0, 0
if isinstance(confidence, dict):
numerator: float = 0
denominator: int = 0
for value in confidence.values():
num, den = _calculate_overall_confidence_recursive(value)
numerator += num
denominator += den
return numerator, denominator
elif isinstance(confidence, list):
numerator = 0
denominator = 0
for value in confidence:
num, den = _calculate_overall_confidence_recursive(value)
numerator += num
denominator += den
return numerator, denominator
else:
return 0, 0
class TypedAggregateGroup(BaseModel, Generic[AgentDataT]):
"""
Represents a group of agent data records aggregated by common field values.
This class is used for grouping and analyzing agent data based on shared
characteristics. It's particularly useful for generating summaries and
statistics across large datasets.
Attributes:
group_key: The field values that define this group
count: Number of records in this group (if count aggregation was requested)
first_item: Representative data record from this group (if requested)
Example:
```python
# Group by age range
groups = await client.aggregate_agent_data(
group_by=["age_range"],
count=True,
first=True
)
for group in groups.items:
print(f"Age range {group.group_key['age_range']}: {group.count} people")
if group.first_item:
print(f"Example: {group.first_item.name}")
```
"""
group_key: Dict[str, Any] = Field(
description="The field values that define this group"
)
count: Optional[int] = Field(
description="Number of records in this group (if count aggregation was requested)"
)
first_item: Optional[AgentDataT] = Field(
description="Representative data record from this group (if requested)"
)
@classmethod
def from_raw(
cls, raw_data: AggregateGroup, validator: Type[AgentDataT]
) -> "TypedAggregateGroup[AgentDataT]":
"""
Convert raw API response to typed aggregate group.
Args:
raw_data: Raw aggregate group from the API
validator: Pydantic model class to validate the first_item field
Returns:
TypedAggregateGroup instance with validated first_item
"""
first_item: Optional[AgentDataT] = raw_data.first_item
if first_item is not None:
first_item = validator.model_validate(first_item)
return cls(
group_key=raw_data.group_key,
count=raw_data.count,
first_item=first_item,
)
class TypedAggregateGroupItems(BaseModel, Generic[AgentDataT]):
"""
Paginated collection of aggregate groups.
This class represents a page of aggregation results from the agent data API,
providing both the grouped data and pagination metadata.
Attributes:
items: List of aggregate groups in this page
total: Total number of groups matching the query (only present if requested)
has_more: Whether there are more groups available beyond this page
Example:
```python
# Get first page of groups
results = await client.aggregate_agent_data(
group_by=["department"],
count=True,
page_size=20
)
for group in results.items:
dept = group.group_key["department"]
print(f"{dept}: {group.count} employees")
# Load more if needed
if results.has_more:
next_page = await client.aggregate_agent_data(
group_by=["department"],
count=True,
page_size=20,
offset=20
)
```
"""
items: List[TypedAggregateGroup[AgentDataT]] = Field(
description="List of aggregate groups in this page"
)
total: Optional[int] = Field(
description="Total number of groups matching the query (only present if requested)"
)
has_more: bool = Field(
description="Whether there are more groups available beyond this page"
)
@@ -0,0 +1,17 @@
from llama_cloud_services.extract.extract import (
LlamaExtract,
ExtractConfig,
ExtractionAgent,
SourceText,
ExtractTarget,
ExtractMode,
)
__all__ = [
"LlamaExtract",
"ExtractionAgent",
"SourceText",
"ExtractConfig",
"ExtractTarget",
"ExtractMode",
]
@@ -1,31 +1,39 @@
import asyncio
import os
import time
from io import BufferedIOBase, BufferedReader, BytesIO
from io import BufferedIOBase, BufferedReader, BytesIO, TextIOWrapper
from pathlib import Path
from typing import List, Optional, Type, Union, Coroutine, Any, TypeVar
import secrets
import warnings
import httpx
from pydantic import BaseModel
from tenacity import (
retry_if_exception,
stop_after_attempt,
wait_exponential_jitter,
AsyncRetrying,
)
from llama_cloud import (
ExtractAgent as CloudExtractAgent,
ExtractAgentCreate,
ExtractConfig,
ExtractJob,
ExtractJobCreate,
ExtractRun,
ExtractSchemaValidateRequest,
ExtractAgentUpdate,
File,
ExtractMode,
StatusEnum,
Project,
ExtractTarget,
LlamaExtractSettings,
PaginatedExtractRunsResponse,
)
from llama_cloud.client import AsyncLlamaCloud
from llama_cloud_services.extract.utils import JSONObjectType, augment_async_errors
from llama_cloud.core.api_error import ApiError
from llama_cloud_services.extract.utils import (
JSONObjectType,
augment_async_errors,
ExperimentalWarning,
)
from llama_index.core.schema import BaseComponent
from llama_index.core.async_utils import run_jobs
from llama_index.core.bridge.pydantic import Field, PrivateAttr
@@ -43,30 +51,101 @@ DEFAULT_EXTRACT_CONFIG = ExtractConfig(
)
def _is_retryable_error(exception: BaseException) -> bool:
"""Check if an exception is retryable."""
if isinstance(exception, ApiError):
return exception.status_code in (502, 503, 504, 425, 408)
elif isinstance(
exception, (httpx.HTTPStatusError, httpx.RequestError, httpx.TimeoutException)
):
return True
return False
class SourceText:
def __init__(
self,
file: Union[bytes, BufferedIOBase, str, Path],
*,
file: Union[bytes, BufferedIOBase, TextIOWrapper, str, Path, None] = None,
text_content: Optional[str] = None,
filename: Optional[str] = None,
):
self.file = file
self.filename = filename
self.text_content = text_content
self._validate()
def _validate(self) -> None:
"""Ensure filename is provided when needed."""
if isinstance(self.file, (bytes, BufferedIOBase)):
if not ((self.file is None) ^ (self.text_content is None)):
raise ValueError("Either file or text_content must be provided.")
if self.text_content is not None:
if not self.filename:
random_hex = secrets.token_hex(4)
self.filename = f"text_input_{random_hex}.txt"
return
if isinstance(self.file, (bytes, BufferedIOBase, TextIOWrapper)):
if not self.filename and hasattr(self.file, "name"):
self.filename = os.path.basename(str(self.file.name))
elif not hasattr(self.file, "name") and self.filename is None:
raise ValueError(
"filename must be provided when file is bytes or a file-like object without a name"
)
elif isinstance(self.file, (str, Path)):
if not self.filename:
self.filename = os.path.basename(str(self.file))
else:
raise ValueError(f"Unsupported file type: {type(self.file)}")
FileInput = Union[str, Path, BufferedIOBase, SourceText]
def run_in_thread(
coro: Coroutine[Any, Any, T],
thread_pool: ThreadPoolExecutor,
verify: bool,
httpx_timeout: float,
client_wrapper: Any,
) -> T:
"""Run coroutine in a thread with proper client management."""
async def wrapped_coro() -> T:
client = httpx.AsyncClient(
verify=verify,
timeout=httpx_timeout,
limits=httpx.Limits(max_keepalive_connections=100, max_connections=100),
)
original_client = client_wrapper.httpx_client
try:
client_wrapper.httpx_client = client
return await coro
finally:
client_wrapper.httpx_client = original_client
await client.aclose()
def run_coro() -> T:
try:
return asyncio.run(wrapped_coro())
except httpx.TimeoutException as e:
raise TimeoutError(f"Request timed out: {str(e)}") from e
except httpx.NetworkError as e:
raise ConnectionError(f"Network error: {str(e)}") from e
return thread_pool.submit(run_coro).result()
def _extraction_config_warning(config: ExtractConfig) -> None:
if config.cite_sources or config.confidence_scores:
warnings.warn(
"`cite_sources`/`confidence_scores` could greatly increase the "
"size of the response, and slow down the extraction. Results will be "
"available in the `extraction_metadata` field for the extraction run.",
ExperimentalWarning,
)
class ExtractionAgent:
"""Class representing a single extraction agent with methods for extraction operations."""
@@ -101,31 +180,6 @@ class ExtractionAgent:
max_workers=min(10, (os.cpu_count() or 1) + 4)
)
def _run_in_thread(self, coro: Coroutine[Any, Any, T]) -> T:
"""Run coroutine in a separate thread to avoid event loop issues"""
def run_coro() -> T:
async def wrapped_coro() -> T:
# Get the original client to preserve its configuration
original_client = self._client._client_wrapper.httpx_client
# Create a new client with the same configuration as the original
async with httpx.AsyncClient(
verify=self.verify,
timeout=self.httpx_timeout,
) as client:
# Temporarily replace the client
self._client._client_wrapper.httpx_client = client
try:
return await coro
finally:
# Restore the original client
self._client._client_wrapper.httpx_client = original_client
return asyncio.run(wrapped_coro())
return self._thread_pool.submit(run_coro).result()
@property
def id(self) -> str:
return self._agent.id
@@ -152,7 +206,7 @@ class ExtractionAgent:
)
validated_schema = self._run_in_thread(
self._client.llama_extract.validate_extraction_schema(
request=ExtractSchemaValidateRequest(data_schema=processed_schema)
data_schema=processed_schema
)
)
self._data_schema = validated_schema.data_schema
@@ -163,8 +217,19 @@ class ExtractionAgent:
@config.setter
def config(self, config: ExtractConfig) -> None:
_extraction_config_warning(config)
self._config = config
def _run_in_thread(self, coro: Coroutine[Any, Any, T]) -> T:
"""Run coroutine in a separate thread to avoid event loop issues"""
return run_in_thread(
coro,
self._thread_pool,
self.verify, # type: ignore
self.httpx_timeout, # type: ignore
self._client._client_wrapper,
)
async def upload_file(self, file_input: SourceText) -> File:
"""Upload a file for extraction.
@@ -175,14 +240,26 @@ class ExtractionAgent:
ValueError: If filename is not provided for bytes input or for file-like objects
without a name attribute.
"""
file_contents: Optional[Union[BufferedIOBase, BytesIO]] = None
try:
file_contents: Union[BufferedIOBase, BytesIO]
if isinstance(file_input.file, (str, Path)):
if file_input.text_content is not None:
# Handle direct text content
file_contents = BytesIO(file_input.text_content.encode("utf-8"))
elif isinstance(file_input.file, TextIOWrapper):
# Handle text-based IO objects
file_contents = BytesIO(file_input.file.read().encode("utf-8"))
elif isinstance(file_input.file, (str, Path)):
# Handle file paths
file_contents = open(file_input.file, "rb")
elif isinstance(file_input.file, bytes):
# Handle bytes
file_contents = BytesIO(file_input.file)
else:
elif isinstance(file_input.file, BufferedIOBase):
# Handle binary IO objects
file_contents = file_input.file
else:
raise ValueError(f"Unsupported file type: {type(file_input.file)}")
# Add name attribute to file object if needed
if not hasattr(file_contents, "name"):
file_contents.name = file_input.filename # type: ignore
@@ -191,7 +268,7 @@ class ExtractionAgent:
project_id=self._project_id, upload_file=file_contents
)
finally:
if isinstance(file_contents, BufferedReader):
if file_contents is not None and isinstance(file_contents, BufferedReader):
file_contents.close()
async def _upload_file(self, file_input: FileInput) -> File:
@@ -219,35 +296,60 @@ class ExtractionAgent:
return await self.upload_file(source_text)
async def _get_job_with_retry(self, job_id: str) -> ExtractJob:
"""Get job with retry logic for transient errors."""
async for attempt in AsyncRetrying(
retry=retry_if_exception(_is_retryable_error),
stop=stop_after_attempt(5),
wait=wait_exponential_jitter(initial=1, max=60, jitter=5),
reraise=True,
):
with attempt:
return await self._client.llama_extract.get_job(job_id=job_id)
async def _get_run_with_retry(self, job_id: str) -> ExtractRun:
"""Get extraction run with retry logic for transient errors."""
async for attempt in AsyncRetrying(
retry=retry_if_exception(_is_retryable_error),
stop=stop_after_attempt(3),
wait=wait_exponential_jitter(initial=1, max=20, jitter=3),
reraise=True,
):
with attempt:
return await self._client.llama_extract.get_run_by_job_id(job_id=job_id)
async def _wait_for_job_result(self, job_id: str) -> Optional[ExtractRun]:
"""Wait for and return the results of an extraction job."""
start = time.perf_counter()
tries = 0
while True:
await asyncio.sleep(self.check_interval)
tries += 1
job = await self._client.llama_extract.get_job(
job_id=job_id,
)
if job.status == StatusEnum.SUCCESS:
return await self._client.llama_extract.get_run_by_job_id(
job_id=job_id,
)
elif job.status == StatusEnum.PENDING:
end = time.perf_counter()
if end - start > self.max_timeout:
raise Exception(f"Timeout while extracting the file: {job_id}")
if self._verbose and tries % 10 == 0:
print(".", end="", flush=True)
continue
else:
warnings.warn(
f"Failure in job: {job_id}, status: {job.status}, error: {job.error}"
)
return await self._client.llama_extract.get_run_by_job_id(
job_id=job_id,
)
try:
job = await self._get_job_with_retry(job_id)
if job.status == StatusEnum.SUCCESS:
return await self._get_run_with_retry(job_id)
elif job.status == StatusEnum.PENDING:
end = time.perf_counter()
if end - start > self.max_timeout:
raise Exception(f"Timeout while extracting the file: {job_id}")
if self._verbose and tries % 10 == 0:
print(".", end="", flush=True)
continue
else:
warnings.warn(
f"Failure in job: {job_id}, status: {job.status}, error: {job.error}"
)
return await self._get_run_with_retry(job_id)
except Exception as e:
# If we get a non-retryable error or all retries are exhausted, re-raise
if self._verbose:
print(f"\nError in job polling for {job_id}: {e}")
raise e
def save(self) -> None:
"""Persist the extraction agent's schema and config to the database.
@@ -258,10 +360,8 @@ class ExtractionAgent:
self._agent = self._run_in_thread(
self._client.llama_extract.update_extraction_agent(
extraction_agent_id=self.id,
request=ExtractAgentUpdate(
data_schema=self.data_schema,
config=self.config,
),
data_schema=self.data_schema,
config=self.config,
)
)
@@ -475,6 +575,14 @@ class ExtractionAgent:
def __repr__(self) -> str:
return f"ExtractionAgent(id={self.id}, name={self.name})"
def __del__(self) -> None:
"""Cleanup resources properly."""
try:
if hasattr(self, "_thread_pool"):
self._thread_pool.shutdown(wait=True)
except Exception:
pass # Suppress exceptions during cleanup
class LlamaExtract(BaseComponent):
"""Factory class for creating and managing extraction agents."""
@@ -545,7 +653,7 @@ class LlamaExtract(BaseComponent):
httpx_timeout=httpx_timeout,
verbose=verbose,
)
self._httpx_client = httpx.AsyncClient(verify=verify, timeout=httpx_timeout)
self._httpx_client = httpx.AsyncClient(verify=verify, timeout=httpx_timeout) # type: ignore
self.verify = verify
self.httpx_timeout = httpx_timeout
@@ -557,51 +665,20 @@ class LlamaExtract(BaseComponent):
self._thread_pool = ThreadPoolExecutor(
max_workers=min(10, (os.cpu_count() or 1) + 4)
)
# Fetch default project id if not provided
if not project_id:
project_id = os.getenv("LLAMA_CLOUD_PROJECT_ID", None)
if not project_id:
print("No project_id provided, fetching default project.")
projects: List[Project] = self._run_in_thread(
self._async_client.projects.list_projects()
)
default_project = [p for p in projects if p.is_default]
if not default_project:
raise ValueError(
"No default project found. Please provide a project_id."
)
project_id = default_project[0].id
self._project_id = project_id
self._organization_id = organization_id
def _run_in_thread(self, coro: Coroutine[Any, Any, T]) -> T:
"""Run coroutine in a separate thread to avoid event loop issues"""
def run_coro() -> T:
# Create a new client for this thread
async def wrapped_coro() -> T:
assert (
self._httpx_client is not None
), "httpx_client should be initialized"
# Create a new client with the same configuration as the original
async with httpx.AsyncClient(
verify=self.verify,
timeout=self.httpx_timeout,
) as client:
# Temporarily replace the client
self._async_client._client_wrapper.httpx_client = client
try:
return await coro
finally:
# Restore the original client
self._async_client._client_wrapper.httpx_client = (
self._httpx_client
)
return asyncio.run(wrapped_coro())
return self._thread_pool.submit(run_coro).result()
return run_in_thread(
coro,
self._thread_pool,
self.verify, # type: ignore
self.httpx_timeout, # type: ignore
self._async_client._client_wrapper,
)
def create_agent(
self,
@@ -620,11 +697,7 @@ class LlamaExtract(BaseComponent):
ExtractionAgent: The created extraction agent
"""
if config is not None:
if config.extraction_mode == ExtractMode.ACCURATE:
warnings.warn(
"ACCURATE extraction mode is deprecated. Using BALANCED instead."
)
config.extraction_mode = ExtractMode.BALANCED
_extraction_config_warning(config)
else:
config = DEFAULT_EXTRACT_CONFIG
@@ -641,11 +714,9 @@ class LlamaExtract(BaseComponent):
self._async_client.llama_extract.create_extraction_agent(
project_id=self._project_id,
organization_id=self._organization_id,
request=ExtractAgentCreate(
name=name,
data_schema=data_schema,
config=config,
),
name=name,
data_schema=data_schema,
config=config,
)
)
@@ -659,6 +730,8 @@ class LlamaExtract(BaseComponent):
num_workers=self.num_workers,
show_progress=self.show_progress,
verbose=self.verbose,
verify=self.verify,
httpx_timeout=self.httpx_timeout,
)
def get_agent(
@@ -746,6 +819,14 @@ class LlamaExtract(BaseComponent):
)
)
def __del__(self) -> None:
"""Cleanup resources properly."""
try:
if hasattr(self, "_thread_pool"):
self._thread_pool.shutdown(wait=True)
except Exception:
pass # Suppress exceptions during cleanup
if __name__ == "__main__":
from dotenv import load_dotenv
@@ -32,3 +32,9 @@ def augment_async_errors() -> Generator[None, None, None]:
JSONType = Union[Dict[str, Any], List[Any], str, int, float, bool, None]
JSONObjectType = Dict[str, JSONType]
class ExperimentalWarning(Warning):
"""Warning for experimental features."""
pass
+11
View File
@@ -0,0 +1,11 @@
from .base import LlamaCloudIndex
from .retriever import LlamaCloudRetriever
from .composite_retriever import (
LlamaCloudCompositeRetriever,
)
__all__ = [
"LlamaCloudIndex",
"LlamaCloudRetriever",
"LlamaCloudCompositeRetriever",
]
+422
View File
@@ -0,0 +1,422 @@
import base64
import urllib.parse
import uuid
from httpx import Request, HTTPStatusError
from tenacity import (
retry,
wait_exponential_jitter,
stop_after_attempt,
retry_if_exception,
)
from typing import Any, Optional, Tuple, List, Union, Dict, Callable
from llama_index.core.async_utils import run_jobs
from llama_index.core.schema import NodeWithScore, ImageNode
from llama_cloud import (
AutoTransformConfig,
PageFigureNodeWithScore,
PageScreenshotNodeWithScore,
Pipeline,
PipelineCreateTransformConfig,
PipelineType,
Project,
Retriever,
)
from llama_cloud.core import remove_none_from_dict
from llama_cloud.client import LlamaCloud, AsyncLlamaCloud
from llama_cloud.core.api_error import ApiError
def is_retryable_http_error(exception: Exception) -> bool:
# Retry for ApiError with 5xx status codes
if isinstance(exception, ApiError):
return 500 <= exception.status_code < 600
# Also retry for HTTPError with 5xx status codes
elif isinstance(exception, HTTPStatusError):
return 500 <= exception.response.status_code < 600
return False
def retry_on_failure(func: Callable) -> Callable:
"""Decorator to apply tenacity retry with exponential backoff and jitter for 5xx errors."""
return retry(
wait=wait_exponential_jitter(exp_base=2, max=10),
stop=stop_after_attempt(5),
retry=retry_if_exception(is_retryable_http_error),
reraise=True,
)(func)
def default_transform_config() -> PipelineCreateTransformConfig:
return AutoTransformConfig()
def resolve_retriever(
client: LlamaCloud,
project: Project,
retriever_name: Optional[str] = None,
retriever_id: Optional[str] = None,
persisted: bool = True,
) -> Optional[Retriever]:
if not persisted:
return Retriever(
id=str(uuid.uuid4()),
project_id=project.id,
name=retriever_name,
pipelines=[],
)
if retriever_id:
return client.retrievers.get_retriever(
retriever_id=retriever_id, project_id=project.id
)
elif retriever_name:
retrievers = client.retrievers.list_retrievers(
project_id=project.id, name=retriever_name
)
return next(
(retriever for retriever in retrievers if retriever.name == retriever_name),
None,
)
else:
return None
def resolve_project(
client: LlamaCloud,
project_name: Optional[str],
project_id: Optional[str],
organization_id: Optional[str],
) -> Project:
if project_id is not None:
return client.projects.get_project(project_id=project_id)
else:
projects = client.projects.list_projects(
project_name=project_name, organization_id=organization_id
)
if len(projects) == 0:
raise ValueError(f"No project found with name {project_name}")
elif len(projects) > 1:
raise ValueError(
f"Multiple projects found with name {project_name}. Please specify organization_id."
)
return projects[0]
def resolve_pipeline(
client: LlamaCloud,
pipeline_id: Optional[str],
project: Optional[Project],
pipeline_name: Optional[str],
) -> Pipeline:
if pipeline_id is not None:
return client.pipelines.get_pipeline(pipeline_id=pipeline_id)
else:
pipelines = client.pipelines.search_pipelines(
project_id=project.id, # type: ignore [union-attr]
pipeline_name=pipeline_name,
pipeline_type=PipelineType.MANAGED.value, # type: ignore [union-attr]
)
if len(pipelines) == 0:
raise ValueError(
f"Unknown index name {pipeline_name}. Please confirm an index with this name exists."
)
elif len(pipelines) > 1:
raise ValueError(
f"Multiple pipelines found with name {pipeline_name} in project {project.name}" # type: ignore [union-attr]
)
return pipelines[0]
def resolve_project_and_pipeline(
client: LlamaCloud,
pipeline_name: Optional[str],
pipeline_id: Optional[str],
project_name: Optional[str],
project_id: Optional[str],
organization_id: Optional[str],
) -> Tuple[Project, Pipeline]:
# resolve pipeline by ID
if pipeline_id is not None:
pipeline = resolve_pipeline(
client, pipeline_id=pipeline_id, project=None, pipeline_name=None
)
project_id = pipeline.project_id
# resolve project
project = resolve_project(client, project_name, project_id, organization_id)
# resolve pipeline by name
if pipeline_id is None:
pipeline = resolve_pipeline(
client, pipeline_id=None, project=project, pipeline_name=pipeline_name
)
return project, pipeline
def _build_get_page_screenshot_request(
client: Union[LlamaCloud, AsyncLlamaCloud],
file_id: str,
page_index: int,
project_id: str,
) -> Request:
return client._client_wrapper.httpx_client.build_request(
"GET",
urllib.parse.urljoin(
f"{client._client_wrapper.get_base_url()}/",
f"api/v1/files/{file_id}/page_screenshots/{page_index}",
),
params=remove_none_from_dict({"project_id": project_id}),
headers=client._client_wrapper.get_headers(),
timeout=60,
)
def _build_get_page_figure_request(
client: Union[LlamaCloud, AsyncLlamaCloud],
file_id: str,
page_index: int,
figure_name: str,
project_id: str,
) -> Request:
return client._client_wrapper.httpx_client.build_request(
"GET",
urllib.parse.urljoin(
f"{client._client_wrapper.get_base_url()}/",
f"api/v1/files/{file_id}/page-figures/{page_index}/{figure_name}",
),
params=remove_none_from_dict({"project_id": project_id}),
headers=client._client_wrapper.get_headers(),
timeout=60,
)
@retry_on_failure
def get_page_screenshot(
client: LlamaCloud, file_id: str, page_index: int, project_id: str
) -> str:
"""Get the page screenshot."""
# TODO: this currently uses requests, should be replaced with the client
request = _build_get_page_screenshot_request(
client, file_id, page_index, project_id
)
_response = client._client_wrapper.httpx_client.send(request)
if 200 <= _response.status_code < 300:
return _response.content
else:
raise ApiError(status_code=_response.status_code, body=_response.text)
@retry_on_failure
def get_page_figure(
client: LlamaCloud, file_id: str, page_index: int, figure_name: str, project_id: str
) -> str:
request = _build_get_page_figure_request(
client, file_id, page_index, figure_name, project_id
)
_response = client._client_wrapper.httpx_client.send(request)
if 200 <= _response.status_code < 300:
return _response.content
else:
raise ApiError(status_code=_response.status_code, body=_response.text)
@retry_on_failure
async def aget_page_screenshot(
client: AsyncLlamaCloud, file_id: str, page_index: int, project_id: str
) -> str:
"""Get the page screenshot (async)."""
request = _build_get_page_screenshot_request(
client, file_id, page_index, project_id
)
_response = await client._client_wrapper.httpx_client.send(request)
if 200 <= _response.status_code < 300:
return _response.content
else:
raise ApiError(status_code=_response.status_code, body=_response.text)
@retry_on_failure
async def aget_page_figure(
client: AsyncLlamaCloud,
file_id: str,
page_index: int,
figure_name: str,
project_id: str,
) -> str:
request = _build_get_page_figure_request(
client, file_id, page_index, figure_name, project_id
)
_response = await client._client_wrapper.httpx_client.send(request)
if 200 <= _response.status_code < 300:
return _response.content
else:
raise ApiError(status_code=_response.status_code, body=_response.text)
def page_screenshot_nodes_to_node_with_score(
client: LlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
if not raw_image_nodes:
return []
image_nodes = []
for raw_image_node in raw_image_nodes:
image_bytes = get_page_screenshot(
client=client,
file_id=raw_image_node.node.file_id,
page_index=raw_image_node.node.page_index,
project_id=project_id,
)
image_base64 = base64.b64encode(image_bytes).decode("utf-8")
image_node_metadata: Dict[str, Any] = {
**(raw_image_node.node.metadata or {}),
"file_id": raw_image_node.node.file_id,
"page_index": raw_image_node.node.page_index,
}
image_node_with_score = NodeWithScore(
node=ImageNode(image=image_base64, metadata=image_node_metadata),
score=raw_image_node.score,
)
image_nodes.append(image_node_with_score)
return image_nodes
def image_nodes_to_node_with_score(
client: LlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
"""
Legacy method to alias page_screenshot_nodes_to_node_with_score.
"""
if not raw_image_nodes:
return []
return page_screenshot_nodes_to_node_with_score(
client=client, raw_image_nodes=raw_image_nodes, project_id=project_id
)
def page_figure_nodes_to_node_with_score(
client: LlamaCloud,
raw_figure_nodes: Optional[List[PageFigureNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
if not raw_figure_nodes:
return []
figure_nodes = []
for raw_figure_node in raw_figure_nodes:
figure_bytes = get_page_figure(
client=client,
file_id=raw_figure_node.node.file_id,
page_index=raw_figure_node.node.page_index,
figure_name=raw_figure_node.node.figure_name,
project_id=project_id,
)
figure_base64 = base64.b64encode(figure_bytes).decode("utf-8")
figure_node_metadata: Dict[str, Any] = {
**(raw_figure_node.node.metadata or {}),
"file_id": raw_figure_node.node.file_id,
"page_index": raw_figure_node.node.page_index,
"figure_name": raw_figure_node.node.figure_name,
}
figure_node_with_score = NodeWithScore(
node=ImageNode(image=figure_base64, metadata=figure_node_metadata),
score=raw_figure_node.score,
)
figure_nodes.append(figure_node_with_score)
return figure_nodes
async def apage_screenshot_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
if not raw_image_nodes:
return []
image_nodes = []
tasks = [
aget_page_screenshot(
client=client,
file_id=raw_image_node.node.file_id,
page_index=raw_image_node.node.page_index,
project_id=project_id,
)
for raw_image_node in raw_image_nodes
]
image_bytes_list = await run_jobs(tasks)
for image_bytes, raw_image_node in zip(image_bytes_list, raw_image_nodes):
image_base64 = base64.b64encode(image_bytes).decode("utf-8")
image_node_metadata: Dict[str, Any] = {
**(raw_image_node.node.metadata or {}),
"file_id": raw_image_node.node.file_id,
"page_index": raw_image_node.node.page_index,
}
image_node_with_score = NodeWithScore(
node=ImageNode(image=image_base64, metadata=image_node_metadata),
score=raw_image_node.score,
)
image_nodes.append(image_node_with_score)
return image_nodes
async def aimage_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
"""
Legacy method to alias apage_screenshot_nodes_to_node_with_score.
"""
if not raw_image_nodes:
return []
return await apage_screenshot_nodes_to_node_with_score(
client=client, raw_image_nodes=raw_image_nodes, project_id=project_id
)
async def apage_figure_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_figure_nodes: Optional[List[PageFigureNodeWithScore]],
project_id: str,
) -> List[NodeWithScore]:
if not raw_figure_nodes:
return []
figure_nodes = []
tasks = [
aget_page_figure(
client=client,
file_id=raw_figure_node.node.file_id,
page_index=raw_figure_node.node.page_index,
figure_name=raw_figure_node.node.figure_name,
project_id=project_id,
)
for raw_figure_node in raw_figure_nodes
]
figure_bytes_list = await run_jobs(tasks)
for figure_bytes, raw_figure_node in zip(figure_bytes_list, raw_figure_nodes):
figure_base64 = base64.b64encode(figure_bytes).decode("utf-8")
figure_node_metadata: Dict[str, Any] = {
**(raw_figure_node.node.metadata or {}),
"file_id": raw_figure_node.node.file_id,
"page_index": raw_figure_node.node.page_index,
"figure_name": raw_figure_node.node.figure_name,
}
figure_node_with_score = NodeWithScore(
node=ImageNode(image=figure_base64, metadata=figure_node_metadata),
score=raw_figure_node.score,
)
figure_nodes.append(figure_node_with_score)
return figure_nodes
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,291 @@
from typing import Any, List, Optional
import httpx
from llama_cloud import (
CompositeRetrievalMode,
CompositeRetrievedTextNodeWithScore,
RetrieverCreate,
Retriever,
RetrieverPipeline,
PresetRetrievalParams,
ReRankConfig,
)
from llama_cloud.resources.pipelines.client import OMIT
from llama_index.core.base.base_retriever import BaseRetriever
from llama_index.core.constants import DEFAULT_PROJECT_NAME
from llama_index.core.ingestion.api_utils import get_aclient, get_client
from llama_index.core.schema import NodeWithScore, QueryBundle, TextNode
from .base import LlamaCloudIndex
from .api_utils import (
resolve_project,
resolve_retriever,
page_screenshot_nodes_to_node_with_score,
)
class LlamaCloudCompositeRetriever(BaseRetriever):
def __init__(
self,
# retriever identifier
name: Optional[str] = None,
retriever_id: Optional[str] = None,
# project identifier
project_name: Optional[str] = DEFAULT_PROJECT_NAME,
project_id: Optional[str] = None,
organization_id: Optional[str] = None,
# creation options
create_if_not_exists: bool = False,
# connection params
api_key: Optional[str] = None,
base_url: Optional[str] = None,
app_url: Optional[str] = None,
timeout: int = 60,
httpx_client: Optional[httpx.Client] = None,
async_httpx_client: Optional[httpx.AsyncClient] = None,
# composite retrieval params
mode: Optional[CompositeRetrievalMode] = None,
rerank_top_n: Optional[int] = None,
rerank_config: Optional[ReRankConfig] = None,
persisted: Optional[bool] = True,
**kwargs: Any,
) -> None:
"""Initialize the Composite Retriever."""
# initialize clients
self._client = get_client(api_key, base_url, app_url, timeout, httpx_client)
self._aclient = get_aclient(
api_key, base_url, app_url, timeout, async_httpx_client
)
self.project = resolve_project(
self._client, project_name, project_id, organization_id
)
self.name = name
self.project_name = self.project.name
self._persisted = persisted
self.retriever = resolve_retriever(
self._client, self.project, name, retriever_id, persisted # type: ignore [arg-type]
)
if self.retriever is None and persisted:
if create_if_not_exists:
self.retriever = self._client.retrievers.upsert_retriever(
project_id=self.project.id,
request=RetrieverCreate(name=self.name, pipelines=[]),
)
else:
raise ValueError(
f"Retriever with name '{self.name}' does not exist in project."
)
# composite retrieval params
self._mode = mode if mode is not None else OMIT
self._rerank_top_n = rerank_top_n if rerank_top_n is not None else OMIT
self._rerank_config = rerank_config if rerank_config is not None else OMIT
super().__init__(
callback_manager=kwargs.get("callback_manager"),
verbose=kwargs.get("verbose", False),
)
@property
def retriever_pipelines(self) -> List[RetrieverPipeline]:
return self.retriever.pipelines or [] # type: ignore [union-attr]
def update_retriever_pipelines(
self, pipelines: List[RetrieverPipeline]
) -> Retriever:
if self._persisted:
self.retriever = self._client.retrievers.update_retriever(
self.retriever.id, pipelines=pipelines # type: ignore [union-attr]
)
else:
# Update in-memory retriever for non-persisted case using copy
self.retriever = self.retriever.copy(update={"pipelines": pipelines}) # type: ignore [union-attr]
return self.retriever
def add_index(
self,
index: LlamaCloudIndex,
name: Optional[str] = None,
description: Optional[str] = None,
preset_retrieval_parameters: Optional[PresetRetrievalParams] = None,
) -> Retriever:
name = name or index.name
preset_retrieval_parameters = (
preset_retrieval_parameters or index.pipeline.preset_retrieval_parameters
)
retriever_pipeline = RetrieverPipeline(
pipeline_id=index.id,
name=name,
description=description,
preset_retrieval_parameters=preset_retrieval_parameters,
)
current_retriever_pipelines_by_name = {
pipeline.name: pipeline for pipeline in (self.retriever_pipelines or [])
}
current_retriever_pipelines_by_name[
retriever_pipeline.name
] = retriever_pipeline
return self.update_retriever_pipelines(
list(current_retriever_pipelines_by_name.values())
)
def remove_index(self, name: str) -> bool:
current_retriever_pipeline_names = self.retriever.pipelines or [] # type: ignore [union-attr]
new_retriever_pipelines = [
pipeline
for pipeline in current_retriever_pipeline_names
if pipeline.name != name
]
if len(new_retriever_pipelines) == len(current_retriever_pipeline_names):
return False
self.update_retriever_pipelines(new_retriever_pipelines)
return True
async def aupdate_retriever_pipelines(
self, pipelines: List[RetrieverPipeline]
) -> Retriever:
if self._persisted:
self.retriever = await self._aclient.retrievers.update_retriever(
self.retriever.id, pipelines=pipelines # type: ignore [union-attr]
)
else:
# Update in-memory retriever for non-persisted case using copy
self.retriever = self.retriever.copy(update={"pipelines": pipelines}) # type: ignore [union-attr]
return self.retriever
async def async_add_index(
self,
index: LlamaCloudIndex,
name: Optional[str] = None,
description: Optional[str] = None,
preset_retrieval_parameters: Optional[PresetRetrievalParams] = None,
) -> Retriever:
name = name or index.name
preset_retrieval_parameters = (
preset_retrieval_parameters or index.pipeline.preset_retrieval_parameters
)
retriever_pipeline = RetrieverPipeline(
pipeline_id=index.id,
name=name,
description=description,
preset_retrieval_parameters=preset_retrieval_parameters,
)
current_retriever_pipelines_by_name = {
pipeline.name: pipeline for pipeline in (self.retriever_pipelines or [])
}
current_retriever_pipelines_by_name[
retriever_pipeline.name
] = retriever_pipeline
return await self.aupdate_retriever_pipelines(
list(current_retriever_pipelines_by_name.values())
)
async def aremove_index(self, name: str) -> bool:
current_retriever_pipeline_names = self.retriever.pipelines or [] # type: ignore [union-attr]
new_retriever_pipelines = [
pipeline
for pipeline in current_retriever_pipeline_names
if pipeline.name != name
]
if len(new_retriever_pipelines) == len(current_retriever_pipeline_names):
return False
await self.aupdate_retriever_pipelines(new_retriever_pipelines)
return True
def _result_nodes_to_node_with_score(
self, composite_retrieval_node: CompositeRetrievedTextNodeWithScore
) -> NodeWithScore:
return NodeWithScore(
node=TextNode(
id=composite_retrieval_node.node.id,
text=composite_retrieval_node.node.text,
metadata=composite_retrieval_node.node.metadata,
),
score=composite_retrieval_node.score,
)
def _retrieve(
self,
query_bundle: QueryBundle,
mode: Optional[CompositeRetrievalMode] = None,
rerank_top_n: Optional[int] = None,
rerank_config: Optional[ReRankConfig] = None,
) -> List[NodeWithScore]:
mode = mode if mode is not None else self._mode
rerank_top_n = rerank_top_n if rerank_top_n is not None else self._rerank_top_n
rerank_config = (
rerank_config if rerank_config is not None else self._rerank_config
)
if self._persisted:
result = self._client.retrievers.retrieve(
self.retriever.id, # type: ignore [union-attr]
mode=mode,
rerank_top_n=rerank_top_n,
rerank_config=rerank_config,
query=query_bundle.query_str,
)
else:
result = self._client.retrievers.direct_retrieve(
project_id=self.project.id,
mode=mode,
rerank_top_n=rerank_top_n,
rerank_config=rerank_config,
query=query_bundle.query_str,
pipelines=self.retriever.pipelines, # type: ignore [union-attr]
)
node_w_scores = [
self._result_nodes_to_node_with_score(node) for node in result.nodes # type: ignore [union-attr]
]
image_nodes_w_scores = page_screenshot_nodes_to_node_with_score(
self._client, result.image_nodes, self.retriever.project_id # type: ignore [union-attr]
)
return sorted(
node_w_scores + image_nodes_w_scores, key=lambda x: x.score, reverse=True
)
async def _aretrieve(
self,
query_bundle: QueryBundle,
mode: Optional[CompositeRetrievalMode] = None,
rerank_top_n: Optional[int] = None,
rerank_config: Optional[ReRankConfig] = None,
) -> List[NodeWithScore]:
mode = mode if mode is not None else self._mode
rerank_top_n = rerank_top_n if rerank_top_n is not None else self._rerank_top_n
rerank_config = (
rerank_config if rerank_config is not None else self._rerank_config
)
if self._persisted:
result = await self._aclient.retrievers.retrieve(
self.retriever.id, # type: ignore [union-attr]
mode=mode,
rerank_config=rerank_config,
rerank_top_n=rerank_top_n,
query=query_bundle.query_str,
)
else:
result = await self._aclient.retrievers.direct_retrieve(
project_id=self.project.id,
mode=mode,
rerank_top_n=rerank_top_n,
rerank_config=rerank_config,
query=query_bundle.query_str,
pipelines=self.retriever.pipelines, # type: ignore [union-attr]
)
node_w_scores = [
self._result_nodes_to_node_with_score(node) for node in result.nodes # type: ignore [union-attr]
]
image_nodes_w_scores = page_screenshot_nodes_to_node_with_score(
self._aclient, result.image_nodes, self.retriever.project_id # type: ignore [union-attr]
)
return sorted(
node_w_scores + image_nodes_w_scores, key=lambda x: x.score, reverse=True
)
+217
View File
@@ -0,0 +1,217 @@
from typing import Any, List, Optional
import httpx
from llama_cloud import (
TextNodeWithScore,
)
from llama_cloud.resources.pipelines.client import OMIT
from llama_index.core.base.base_retriever import BaseRetriever
from llama_index.core.bridge.pydantic import BaseModel
from llama_index.core.constants import DEFAULT_PROJECT_NAME
from llama_index.core.ingestion.api_utils import get_aclient, get_client
from llama_index.core.schema import NodeWithScore, QueryBundle, TextNode
from llama_index.core.vector_stores.types import MetadataFilters
from .api_utils import (
resolve_project_and_pipeline,
page_screenshot_nodes_to_node_with_score,
page_figure_nodes_to_node_with_score,
apage_screenshot_nodes_to_node_with_score,
apage_figure_nodes_to_node_with_score,
)
import logging
logger = logging.getLogger(__name__)
class LlamaCloudRetriever(BaseRetriever):
def __init__(
self,
# index identifier
name: Optional[str] = None,
index_id: Optional[str] = None, # alias for pipeline_id
id: Optional[str] = None, # alias for pipeline_id
pipeline_id: Optional[str] = None,
# project identifier
project_name: Optional[str] = DEFAULT_PROJECT_NAME,
project_id: Optional[str] = None,
organization_id: Optional[str] = None,
# connection params
api_key: Optional[str] = None,
base_url: Optional[str] = None,
app_url: Optional[str] = None,
timeout: int = 60,
httpx_client: Optional[httpx.Client] = None,
async_httpx_client: Optional[httpx.AsyncClient] = None,
# retrieval params
dense_similarity_top_k: Optional[int] = None,
sparse_similarity_top_k: Optional[int] = None,
enable_reranking: Optional[bool] = None,
rerank_top_n: Optional[int] = None,
alpha: Optional[float] = None,
filters: Optional[MetadataFilters] = None,
retrieval_mode: Optional[str] = None,
files_top_k: Optional[int] = None,
retrieve_image_nodes: Optional[bool] = None,
retrieve_page_screenshot_nodes: Optional[bool] = None,
retrieve_page_figure_nodes: Optional[bool] = None,
search_filters_inference_schema: Optional[BaseModel] = None,
**kwargs: Any,
) -> None:
"""Initialize the Platform Retriever."""
if sum([bool(id), bool(index_id), bool(pipeline_id), bool(name)]) != 1:
raise ValueError(
"Exactly one of `name`, `id`, `pipeline_id` or `index_id` must be provided to identify the index."
)
# initialize clients
self._httpx_client = httpx_client
self._async_httpx_client = async_httpx_client
self._client = get_client(api_key, base_url, app_url, timeout, httpx_client)
self._aclient = get_aclient(
api_key, base_url, app_url, timeout, async_httpx_client
)
pipeline_id = id or index_id or pipeline_id
self.project, self.pipeline = resolve_project_and_pipeline(
self._client, name, pipeline_id, project_name, project_id, organization_id
)
self.name = self.pipeline.name
self.project_name = self.project.name
# retrieval params
self._dense_similarity_top_k = (
dense_similarity_top_k if dense_similarity_top_k is not None else OMIT
)
self._sparse_similarity_top_k = (
sparse_similarity_top_k if sparse_similarity_top_k is not None else OMIT
)
self._enable_reranking = (
enable_reranking if enable_reranking is not None else OMIT
)
self._rerank_top_n = rerank_top_n if rerank_top_n is not None else OMIT
self._alpha = alpha if alpha is not None else OMIT
self._filters = filters if filters is not None else OMIT
self._retrieval_mode = retrieval_mode if retrieval_mode is not None else OMIT
self._files_top_k = files_top_k if files_top_k is not None else OMIT
if retrieve_image_nodes is not None:
logger.warning(
"The `retrieve_image_nodes` parameter is deprecated. "
"Use `retrieve_page_screenshot_nodes` and `retrieve_page_figure_nodes` instead."
)
if retrieve_image_nodes:
if (
retrieve_page_screenshot_nodes is False
or retrieve_page_figure_nodes is False
):
raise ValueError(
"If `retrieve_image_nodes` is set to True, "
"both `retrieve_page_screenshot_nodes` and `retrieve_page_figure_nodes` must also be set to True or omitted."
)
retrieve_page_screenshot_nodes = True
retrieve_page_figure_nodes = True
self._retrieve_page_screenshot_nodes = (
retrieve_page_screenshot_nodes
if retrieve_page_screenshot_nodes is not None
else OMIT
)
self._retrieve_page_figure_nodes = (
retrieve_page_figure_nodes
if retrieve_page_figure_nodes is not None
else OMIT
)
self._search_filters_inference_schema = search_filters_inference_schema
super().__init__(
callback_manager=kwargs.get("callback_manager"),
verbose=kwargs.get("verbose", False),
)
def _result_nodes_to_node_with_score(
self, result_nodes: List[TextNodeWithScore]
) -> List[NodeWithScore]:
nodes = []
for res in result_nodes:
text_node = TextNode.parse_obj(res.node.dict())
nodes.append(NodeWithScore(node=text_node, score=res.score))
return nodes
def _retrieve(self, query_bundle: QueryBundle) -> List[NodeWithScore]:
"""Retrieve from the platform."""
search_filters_inference_schema = OMIT
if self._search_filters_inference_schema is not None:
search_filters_inference_schema = (
self._search_filters_inference_schema.model_json_schema()
)
results = self._client.pipelines.run_search(
query=query_bundle.query_str,
pipeline_id=self.pipeline.id,
dense_similarity_top_k=self._dense_similarity_top_k,
sparse_similarity_top_k=self._sparse_similarity_top_k,
enable_reranking=self._enable_reranking,
rerank_top_n=self._rerank_top_n,
alpha=self._alpha,
search_filters=self._filters,
files_top_k=self._files_top_k,
retrieval_mode=self._retrieval_mode,
retrieve_page_screenshot_nodes=self._retrieve_page_screenshot_nodes,
retrieve_page_figure_nodes=self._retrieve_page_figure_nodes,
search_filters_inference_schema=search_filters_inference_schema,
)
result_nodes = self._result_nodes_to_node_with_score(results.retrieval_nodes)
if self._retrieve_page_screenshot_nodes:
result_nodes.extend(
page_screenshot_nodes_to_node_with_score(
self._client, results.image_nodes, self.project.id
)
)
if self._retrieve_page_figure_nodes:
result_nodes.extend(
page_figure_nodes_to_node_with_score(
self._client, results.page_figure_nodes, self.project.id
)
)
return result_nodes
async def _aretrieve(self, query_bundle: QueryBundle) -> List[NodeWithScore]:
"""Asynchronously retrieve from the platform."""
search_filters_inference_schema = OMIT
if self._search_filters_inference_schema is not None:
search_filters_inference_schema = (
self._search_filters_inference_schema.model_json_schema()
)
results = await self._aclient.pipelines.run_search(
query=query_bundle.query_str,
pipeline_id=self.pipeline.id,
dense_similarity_top_k=self._dense_similarity_top_k,
sparse_similarity_top_k=self._sparse_similarity_top_k,
enable_reranking=self._enable_reranking,
rerank_top_n=self._rerank_top_n,
alpha=self._alpha,
search_filters=self._filters,
files_top_k=self._files_top_k,
retrieval_mode=self._retrieval_mode,
retrieve_page_screenshot_nodes=self._retrieve_page_screenshot_nodes,
retrieve_page_figure_nodes=self._retrieve_page_figure_nodes,
search_filters_inference_schema=search_filters_inference_schema,
)
result_nodes = self._result_nodes_to_node_with_score(results.retrieval_nodes)
if self._retrieve_page_screenshot_nodes:
result_nodes.extend(
await apage_screenshot_nodes_to_node_with_score(
self._aclient, results.image_nodes, self.project.id
)
)
if self._retrieve_page_figure_nodes:
result_nodes.extend(
await apage_figure_nodes_to_node_with_score(
self._aclient, results.page_figure_nodes, self.project.id
)
)
return result_nodes
@@ -0,0 +1,8 @@
from llama_cloud_services.parse.base import (
LlamaParse,
ResultType,
ParsingMode,
FailedPageMode,
)
__all__ = ["LlamaParse", "ResultType", "ParsingMode", "FailedPageMode"]
@@ -2,28 +2,42 @@ import asyncio
import mimetypes
import os
import time
import warnings
from contextlib import asynccontextmanager
from copy import deepcopy
from enum import Enum
from io import BufferedIOBase
from pathlib import Path, PurePath, PurePosixPath
from typing import Any, AsyncGenerator, Dict, List, Optional, Union
from typing import Any, AsyncGenerator, Dict, List, Optional, Tuple, Union
from urllib.parse import urlparse
import httpx
from fsspec import AbstractFileSystem
from llama_index.core.async_utils import asyncio_run, run_jobs
from llama_index.core.bridge.pydantic import Field, PrivateAttr, field_validator
from llama_index.core.bridge.pydantic import (
Field,
PrivateAttr,
field_validator,
model_validator,
)
from llama_index.core.constants import DEFAULT_BASE_URL
from llama_index.core.readers.base import BasePydanticReader
from llama_index.core.readers.file.base import get_default_fs
from llama_index.core.schema import Document
from llama_cloud_services.utils import check_extra_params
from llama_cloud_services.parse.types import JobResult
from llama_cloud_services.parse.utils import (
SUPPORTED_FILE_TYPES,
ResultType,
ParsingMode,
FailedPageMode,
expand_target_pages,
nest_asyncio_err,
nest_asyncio_msg,
make_api_request,
partition_pages,
extract_tables_from_json_results,
)
# can put in a path to the file or the file bytes itself
@@ -53,6 +67,36 @@ def build_url(
return base_url
class JobFailedException(Exception):
"""Parse job failed exception."""
def __init__(
self,
job_id: str,
status: str,
error_code: Optional[str] = None,
error_message: Optional[str] = None,
):
exception_str = (
f"Job ID: {job_id} failed with status: {status}, "
f'Error code: {error_code or "No error code found"}, '
f'Error message: {error_message or "No error message found"}'
)
super().__init__(exception_str)
self.job_id = job_id
self.status = status
self.error_code = error_code
self.error_message = error_message
@classmethod
def from_result(cls, result_json: Dict[str, Any]) -> "JobFailedException":
job_id = result_json["id"]
status = result_json["status"]
error_code = result_json.get("error_code")
error_message = result_json.get("error_message")
return cls(job_id, status, error_code=error_code, error_message=error_message)
class BackoffPattern(str, Enum):
"""Backoff pattern for polling."""
@@ -111,7 +155,7 @@ class LlamaParse(BasePydanticReader):
num_workers: int = Field(
default=4,
gt=0,
lt=10,
lt=20,
description="The number of workers to use sending API requests for parsing.",
)
result_type: ResultType = Field(
@@ -141,6 +185,10 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, the parser will automatically select the best mode to extract text from documents based on the rules provide. Will use the 'accurate' default mode by default and will upgrade page that match the rule to Premium mode.",
)
auto_mode_configuration_json: Optional[str] = Field(
default=None,
description="A JSON string containing the configuration for the auto mode. If set, the parser will use the provided configuration for the auto mode.",
)
auto_mode_trigger_on_image_in_page: Optional[bool] = Field(
default=False,
description="If auto_mode is set to true, the parser will upgrade the page that contain an image to Premium mode.",
@@ -185,7 +233,10 @@ class LlamaParse(BasePydanticReader):
default=None,
description="The top margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page height.",
)
compact_markdown_table: Optional[bool] = Field(
default=False,
description="If set to true, the parser will output compact markdown table (without trailing spaces in cells).",
)
continuous_mode: Optional[bool] = Field(
default=False,
description="Parse documents continuously, leading to better results on documents where tables span across two pages.",
@@ -223,6 +274,10 @@ class LlamaParse(BasePydanticReader):
default=False,
description="Whether to guess the sheet names of the xlsx file.",
)
high_res_ocr: Optional[bool] = Field(
default=False,
description="If set to true, the parser will use high resolution OCR to extract text from images. This will increase the accuracy of the parsing job, but reduce the speed.",
)
html_make_all_elements_visible: Optional[bool] = Field(
default=False,
description="If set to true, when parsing HTML the parser will consider all elements display not element as display block.",
@@ -262,10 +317,18 @@ class LlamaParse(BasePydanticReader):
language: Optional[str] = Field(
default="en", description="The language of the text to parse."
)
markdown_table_multiline_header_separator: Optional[str] = Field(
default=None,
description="The separator to use to split the header of the markdown table into multiple lines. Default is: <br/>",
)
max_pages: Optional[int] = Field(
default=None,
description="The maximum number of pages to extract text from documents. If set to 0 or not set, all pages will be that should be extracted will be extracted (can work in combination with targetPages).",
)
merge_tables_across_pages_in_markdown: Optional[bool] = Field(
default=False,
description="If set to true, the parser will merge tables across pages in the markdown output. This is useful for documents with tables that span across multiple pages.",
)
output_pdf_of_document: Optional[bool] = Field(
default=False,
description="If set to true, the parser will also output a PDF of the document. (except for spreadsheets)",
@@ -282,6 +345,14 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, the parser will output tables as HTML in the markdown.",
)
outlined_table_extraction: Optional[bool] = Field(
default=False,
description="If set to true, the parser will use a dedicated approach to extract tables with outlined cells. This is useful for documents with spreadsheet-like tables where cells are outlined with borders. This could lead to false positives, so use with caution.",
)
page_error_tolerance: Optional[float] = Field(
default=None,
description="The error tolerance for the number of pages with error in a doc (percentage express as 0-1). If we fail to parse a greater percentage of pages than the tolerance value we fail the job.",
)
page_prefix: Optional[str] = Field(
default=None,
description="A templated prefix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
@@ -294,7 +365,7 @@ class LlamaParse(BasePydanticReader):
default=None,
description="A templated suffix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
)
parse_mode: Optional[str] = Field(
parse_mode: Optional[Union[ParsingMode, str]] = Field(
default=None,
description="The parsing mode to use, see ParsingMode enum for possible values ",
)
@@ -302,10 +373,30 @@ class LlamaParse(BasePydanticReader):
default=False,
description="Use our best parser mode if set to True.",
)
preset: Optional[str] = Field(
default=None,
description="The preset to use for the parser. If set, the parser will use the preset configuration. See LlamaParse documentation for available presets. Preset override most other parameters.",
)
preserve_layout_alignment_across_pages: Optional[bool] = Field(
default=False,
description="Preserve grid alignment across page in text mode.",
)
preserve_very_small_text: Optional[bool] = Field(
default=False,
description="If set, the parser will try to preserve very small text lines. This can be useful for documents containing vector graphics with very small text lines that may not be recognized by OCR or a vision model (such as in CAD drawings).",
)
replace_failed_page_mode: Optional[FailedPageMode] = Field(
default=None,
description="The mode to use to replace the failed page, see FailedPageMode enum for possible value. If set, the parser will replace the failed page with the specified mode. If not set, the default mode (raw_text) will be used.",
)
replace_failed_page_with_error_message_prefix: Optional[str] = Field(
default=None,
description="A prefix to add before error message in failed pages. If not set, no prefix will be used.",
)
replace_failed_page_with_error_message_suffix: Optional[str] = Field(
default=None,
description="A suffix to add after error message in failed pages. If not set, no suffix will be used.",
)
skip_diagonal_text: Optional[bool] = Field(
default=False,
description="If set to true, the parser will ignore diagonal text (when the text rotation in degrees modulo 90 is not 0).",
@@ -375,10 +466,42 @@ class LlamaParse(BasePydanticReader):
default=None,
description="The model name for the vendor multimodal API.",
)
model: Optional[str] = Field(
default=None,
description="The document model name to be used with `parse_with_agent`.",
)
webhook_url: Optional[str] = Field(
default=None,
description="A URL that needs to be called at the end of the parsing job.",
)
partition_pages: Optional[int] = Field(
default=None,
description="If set, documents will automatically be partitioned into segments containing the specified number of pages at most. Parsing will be split into separate jobs for each partition segment. Can be used in combination with targetPages and maxPages.",
)
hide_headers: Optional[bool] = Field(
default=False,
description="Whether to hide page header in output markdown.",
)
hide_footers: Optional[bool] = Field(
default=False,
description="Whether to hide page footers in output markdown.",
)
page_header_suffix: Optional[str] = Field(
default=None,
description="A suffix to add to the page header in the output markdown.",
)
page_header_prefix: Optional[str] = Field(
default=None,
description="A prefix to add to the page header in the output markdown.",
)
page_footer_suffix: Optional[str] = Field(
default=None,
description="A suffix to add to the page footer in the output markdown.",
)
page_footer_prefix: Optional[str] = Field(
default=None,
description="A prefix to add to the page footer in the output markdown.",
)
# Deprecated
bounding_box: Optional[str] = Field(
@@ -418,6 +541,21 @@ class LlamaParse(BasePydanticReader):
description="Whether to use the vendor multimodal API.",
)
@model_validator(mode="before")
@classmethod
def warn_extra_params(cls, data: Dict[str, Any]) -> Dict[str, Any]:
extra_params, suggestions = check_extra_params(cls, data)
if extra_params:
suggestions = [f"\n - {suggestion}" for suggestion in suggestions]
suggestions_str = "".join(suggestions)
warnings.warn(
"The following parameters are unused: "
+ ", ".join(extra_params)
+ f".\n{suggestions_str}",
)
return data
@field_validator("api_key", mode="before", check_fields=True)
@classmethod
def validate_api_key(cls, v: str) -> str:
@@ -505,6 +643,7 @@ class LlamaParse(BasePydanticReader):
file_input: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
partition_target_pages: Optional[str] = None,
) -> str:
files = None
file_handle = None
@@ -555,6 +694,9 @@ class LlamaParse(BasePydanticReader):
if self.auto_mode:
data["auto_mode"] = self.auto_mode
if self.auto_mode_configuration_json is not None:
data["auto_mode_configuration_json"] = self.auto_mode_configuration_json
if self.auto_mode_trigger_on_image_in_page:
data[
"auto_mode_trigger_on_image_in_page"
@@ -599,6 +741,9 @@ class LlamaParse(BasePydanticReader):
if self.bbox_top is not None:
data["bbox_top"] = self.bbox_top
if self.compact_markdown_table:
data["compact_markdown_table"] = self.compact_markdown_table
if self.complemental_formatting_instruction:
print(
"WARNING: complemental_formatting_instruction is deprecated and may be remove in a future release. Use system_prompt, system_prompt_append or user_prompt instead."
@@ -649,6 +794,9 @@ class LlamaParse(BasePydanticReader):
if self.html_make_all_elements_visible:
data["html_make_all_elements_visible"] = self.html_make_all_elements_visible
if self.high_res_ocr:
data["high_res_ocr"] = self.high_res_ocr
if self.html_remove_fixed_elements:
data["html_remove_fixed_elements"] = self.html_remove_fixed_elements
@@ -711,9 +859,38 @@ class LlamaParse(BasePydanticReader):
if self.output_tables_as_HTML:
data["output_tables_as_HTML"] = self.output_tables_as_HTML
if self.outlined_table_extraction:
data["outlined_table_extraction"] = self.outlined_table_extraction
if self.page_error_tolerance is not None:
data["page_error_tolerance"] = self.page_error_tolerance
if self.page_prefix is not None:
data["page_prefix"] = self.page_prefix
if self.merge_tables_across_pages_in_markdown:
data[
"merge_tables_across_pages_in_markdown"
] = self.merge_tables_across_pages_in_markdown
if self.hide_headers:
data["hide_headers"] = self.hide_headers
if self.hide_footers:
data["hide_footers"] = self.hide_footers
if self.page_header_suffix is not None:
data["page_header_suffix"] = self.page_header_suffix
if self.page_header_prefix is not None:
data["page_header_prefix"] = self.page_header_prefix
if self.page_footer_suffix is not None:
data["page_footer_suffix"] = self.page_footer_suffix
if self.page_footer_prefix is not None:
data["page_footer_prefix"] = self.page_footer_prefix
# only send page separator to server if it is not None
# as if a null, "" string is sent the server will then ignore the page separator instead of using the default
if self.page_separator is not None:
@@ -739,6 +916,25 @@ class LlamaParse(BasePydanticReader):
"preserve_layout_alignment_across_pages"
] = self.preserve_layout_alignment_across_pages
if self.preserve_very_small_text:
data["preserve_very_small_text"] = self.preserve_very_small_text
if self.preset is not None:
data["preset"] = self.preset
if self.replace_failed_page_mode is not None:
data["replace_failed_page_mode"] = self.replace_failed_page_mode.value
if self.replace_failed_page_with_error_message_prefix is not None:
data[
"replace_failed_page_with_error_message_prefix"
] = self.replace_failed_page_with_error_message_prefix
if self.replace_failed_page_with_error_message_suffix is not None:
data[
"replace_failed_page_with_error_message_suffix"
] = self.replace_failed_page_with_error_message_suffix
if self.skip_diagonal_text:
data["skip_diagonal_text"] = self.skip_diagonal_text
@@ -774,7 +970,9 @@ class LlamaParse(BasePydanticReader):
if self.take_screenshot:
data["take_screenshot"] = self.take_screenshot
if self.target_pages is not None:
if partition_target_pages is not None:
data["target_pages"] = partition_target_pages
elif self.target_pages is not None:
data["target_pages"] = self.target_pages
if self.user_prompt is not None:
data["user_prompt"] = self.user_prompt
@@ -787,9 +985,17 @@ class LlamaParse(BasePydanticReader):
if self.vendor_multimodal_model_name is not None:
data["vendor_multimodal_model_name"] = self.vendor_multimodal_model_name
if self.model is not None:
data["model"] = self.model
if self.webhook_url is not None:
data["webhook_url"] = self.webhook_url
if self.markdown_table_multiline_header_separator is not None:
data[
"markdown_table_multiline_header_separator"
] = self.markdown_table_multiline_header_separator
# Deprecated
if self.bounding_box is not None:
data["bounding_box"] = self.bounding_box
@@ -802,7 +1008,7 @@ class LlamaParse(BasePydanticReader):
try:
url = build_url(JOB_UPLOAD_ROUTE, self.organization_id, self.project_id)
resp = await self.aclient.post(url, files=files, data=data) # type: ignore
resp = await make_api_request(self.aclient, "POST", url, timeout=self.max_timeout, files=files, data=data) # type: ignore
resp.raise_for_status() # this raises if status is not 2xx
return resp.json()["id"]
except httpx.HTTPStatusError as err: # this catches it
@@ -863,15 +1069,7 @@ class LlamaParse(BasePydanticReader):
print(".", end="", flush=True)
current_interval = self._calculate_backoff(current_interval)
else:
error_code = result_json.get("error_code", "No error code found")
error_message = result_json.get(
"error_message", "No error message found"
)
exception_str = (
f"Job ID: {job_id} failed with status: {status}, "
f"Error code: {error_code}, Error message: {error_message}"
)
raise Exception(exception_str)
raise JobFailedException.from_result(result_json)
except (
httpx.ConnectError,
httpx.ReadError,
@@ -880,6 +1078,7 @@ class LlamaParse(BasePydanticReader):
httpx.ReadTimeout,
httpx.WriteTimeout,
httpx.HTTPStatusError,
httpx.RemoteProtocolError,
) as err:
error_count += 1
end = time.time()
@@ -894,26 +1093,151 @@ class LlamaParse(BasePydanticReader):
)
current_interval = self._calculate_backoff(current_interval)
async def _parse_one(
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
result_type: Optional[str] = None,
num_workers: Optional[int] = None,
) -> List[Tuple[str, Dict[str, Any]]]:
if self.partition_pages is None:
job_results = [
await self._parse_one_unpartitioned(
file_path,
extra_info=extra_info,
fs=fs,
result_type=result_type,
)
]
else:
job_results = await self._parse_one_partitioned(
file_path,
extra_info,
fs=fs,
result_type=result_type,
num_workers=num_workers,
)
return job_results
async def _parse_one_unpartitioned(
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
result_type: Optional[str] = None,
**create_kwargs: Any,
) -> Tuple[str, Dict[str, Any]]:
"""Create one parse job and wait for the result."""
job_id = await self._create_job(
file_path, extra_info=extra_info, fs=fs, **create_kwargs
)
if self.verbose:
print("Started parsing the file under job_id %s" % job_id)
result = await self._get_job_result(
job_id, result_type or self.result_type.value, verbose=self.verbose
)
return job_id, result
async def _parse_one_partitioned(
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
result_type: Optional[str] = None,
num_workers: Optional[int] = None,
) -> List[Tuple[str, Dict[str, Any]]]:
"""Partition a file and run separate parse jobs per partition segment."""
assert self.partition_pages is not None
num_workers = num_workers or self.num_workers
if num_workers < 1:
raise ValueError("Invalid number of workers")
if self.target_pages is not None:
jobs = [
self._parse_one_unpartitioned(
file_path,
extra_info=extra_info,
fs=fs,
result_type=result_type,
partition_target_pages=target_pages,
)
for target_pages in partition_pages(
expand_target_pages(self.target_pages),
self.partition_pages,
max_pages=self.max_pages,
)
]
return await run_jobs(
jobs,
workers=num_workers,
desc="Getting job results",
show_progress=self.show_progress,
)
total = 0
results: List[Tuple[str, Dict[str, Any]]] = []
while self.max_pages is None or total < self.max_pages:
if (
self.max_pages is not None
and total + self.partition_pages >= self.max_pages
):
size = self.max_pages - total
else:
size = self.partition_pages
if not size:
break
try:
# Fetch JSON result type first to get accurate pagination data
# and then fetch the user's desired result type if needed
job_id, json_result = await self._parse_one_unpartitioned(
file_path,
extra_info=extra_info,
fs=fs,
result_type=ResultType.JSON.value,
partition_target_pages=f"{total}-{total + size - 1}",
)
result_type = result_type or self.result_type.value
if result_type == ResultType.JSON.value:
job_result = json_result
else:
job_result = await self._get_job_result(
job_id, result_type, verbose=self.verbose
)
except JobFailedException as e:
if results and e.error_code == "NO_DATA_FOUND_IN_FILE":
# Expected when we try to read past the end of the file
return results
raise
results.append((job_id, job_result))
if len(json_result["pages"]) < size:
break
total += size
return results
async def _aload_data(
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
verbose: bool = False,
num_workers: Optional[int] = None,
) -> List[Document]:
"""Load data from the input path."""
try:
job_id = await self._create_job(file_path, extra_info=extra_info, fs=fs)
if verbose:
print("Started parsing the file under job_id %s" % job_id)
result = await self._get_job_result(
job_id, self.result_type.value, verbose=verbose
)
results = [
job_result
for _, job_result in await self._parse_one(
file_path, extra_info, fs=fs, num_workers=num_workers
)
]
# Flatten the resulting doc if it was partitioned
separator = self.page_separator or _DEFAULT_SEPARATOR
docs = [
Document(
text=result[self.result_type.value],
text=separator.join(
result[self.result_type.value] for result in results
),
metadata=extra_info or {},
)
]
@@ -936,7 +1260,11 @@ class LlamaParse(BasePydanticReader):
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> List[Document]:
"""Load data from the input path."""
"""Load data from the input path.
File(s) which were partitioned before parsing will be loaded as a single
re-assembled Document.
"""
if isinstance(file_path, (str, PurePosixPath, Path, bytes, BufferedIOBase)):
return await self._aload_data(
file_path, extra_info=extra_info, fs=fs, verbose=self.verbose
@@ -948,6 +1276,7 @@ class LlamaParse(BasePydanticReader):
extra_info=extra_info,
fs=fs,
verbose=self.verbose and not self.show_progress,
num_workers=1,
)
for f in file_path
]
@@ -986,21 +1315,161 @@ class LlamaParse(BasePydanticReader):
else:
raise e
async def _aparse_one(
self,
file_path: FileInput,
file_name: str,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
num_workers: Optional[int] = None,
) -> List[JobResult]:
job_results = await self._parse_one(
file_path,
extra_info,
fs=fs,
result_type=ResultType.JSON.value,
num_workers=num_workers,
)
return [
JobResult(
job_id=job_id,
file_name=file_name,
job_result=job_result,
api_key=self.api_key,
base_url=self.base_url,
client=self.aclient,
page_separator=self.page_separator or _DEFAULT_SEPARATOR,
)
for job_id, job_result in job_results
]
async def aparse(
self,
file_path: Union[List[FileInput], FileInput],
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> Union[List["JobResult"], "JobResult"]:
"""
Parse the file and return a JobResult object instead of Document objects.
This method is similar to aload_data but returns JobResult objects that provide
direct access to the various output formats (text, markdown, json, etc.)
Args:
file_path: Path to the file to parse. Can be a string, path, bytes, file-like object, or a list of these.
extra_info: Additional metadata to include in the result.
fs: Optional filesystem to use for reading files.
Returns:
JobResult object or list of JobResult objects if either multiple files were provided or file(s) were partitioned before parsing.
"""
if isinstance(file_path, (str, PurePosixPath, Path, bytes, BufferedIOBase)):
if isinstance(file_path, (bytes, BufferedIOBase)):
if not extra_info or "file_name" not in extra_info:
raise ValueError(
"file_name must be provided in extra_info when passing bytes"
)
file_name = extra_info["file_name"]
else:
file_name = str(file_path)
result = await self._aparse_one(
file_path, file_name, extra_info=extra_info, fs=fs
)
return result[0] if len(result) == 1 else result
elif isinstance(file_path, list):
file_names = []
for f in file_path:
if isinstance(f, (bytes, BufferedIOBase)):
if not extra_info or "file_name" not in extra_info:
raise ValueError(
"file_name must be provided in extra_info when passing bytes"
)
file_names.append(extra_info["file_name"])
else:
file_names.append(str(f))
job_results = []
try:
for result in await run_jobs(
[
self._aparse_one(
f,
file_names[i],
extra_info=extra_info,
fs=fs,
num_workers=1,
)
for i, f in enumerate(file_path)
],
workers=self.num_workers,
desc="Getting job results",
show_progress=self.show_progress,
):
job_results.extend(result)
return job_results
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
else:
raise ValueError(
"The input file_path must be a string or a list of strings."
)
def parse(
self,
file_path: Union[List[FileInput], FileInput],
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> Union[List["JobResult"], "JobResult"]:
"""
Parse the file and return a JobResult object instead of Document objects.
This method is similar to load_data but returns JobResult objects that provide
direct access to the various output formats (text, markdown, json, etc.)
Args:
file_path: Path to the file to parse. Can be a string, path, bytes, file-like object, or a list of these.
extra_info: Additional metadata to include in the result.
fs: Optional filesystem to use for reading files.
Returns:
JobResult object or list of JobResult objects if multiple files were provided
"""
try:
return asyncio_run(self.aparse(file_path, extra_info, fs=fs))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
async def _aget_json(
self, file_path: FileInput, extra_info: Optional[dict] = None
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
num_workers: Optional[int] = None,
) -> List[dict]:
"""Load data from the input path."""
try:
job_id = await self._create_job(file_path, extra_info=extra_info)
if self.verbose:
print("Started parsing the file under job_id %s" % job_id)
result = await self._get_job_result(job_id, "json")
result["job_id"] = job_id
job_results = await self._parse_one(
file_path,
extra_info=extra_info,
result_type=ResultType.JSON.value,
num_workers=num_workers,
)
if not isinstance(file_path, (bytes, BufferedIOBase)):
result["file_path"] = str(file_path)
return [result]
results = []
for job_id, job_result in job_results:
job_result["job_id"] = job_id
if not isinstance(file_path, (bytes, BufferedIOBase)):
job_result["file_path"] = str(file_path)
results.append(job_result)
return results
except Exception as e:
file_repr = file_path if isinstance(file_path, str) else "<bytes/buffer>"
print(f"Error while parsing the file '{file_repr}':", e)
@@ -1015,7 +1484,7 @@ class LlamaParse(BasePydanticReader):
extra_info: Optional[dict] = None,
) -> List[dict]:
"""Load data from the input path."""
if isinstance(file_path, (str, Path)):
if isinstance(file_path, (str, PurePosixPath, Path, bytes, BufferedIOBase)):
return await self._aget_json(file_path, extra_info=extra_info)
elif isinstance(file_path, list):
jobs = [self._aget_json(f, extra_info=extra_info) for f in file_path]
@@ -1036,7 +1505,7 @@ class LlamaParse(BasePydanticReader):
raise e
else:
raise ValueError(
"The input file_path must be a string or a list of strings."
"The input file_path must be a string, Path, bytes, BufferedIOBase, or a list of these types."
)
def get_json_result(
@@ -1053,6 +1522,14 @@ class LlamaParse(BasePydanticReader):
else:
raise e
def get_json(
self,
file_path: Union[List[FileInput], FileInput],
extra_info: Optional[dict] = None,
) -> List[dict]:
"""Load data from the input path."""
return self.get_json_result(file_path, extra_info)
async def aget_assets(
self, json_result: List[dict], download_path: str, asset_key: str
) -> List[dict]:
@@ -1091,7 +1568,9 @@ class LlamaParse(BasePydanticReader):
with open(asset_path, "wb") as f:
asset_url = f"{self.base_url}/api/parsing/job/{job_id}/result/image/{asset_name}"
resp = await client.get(asset_url)
resp = await make_api_request(
client, "GET", asset_url, timeout=self.max_timeout
)
resp.raise_for_status()
f.write(resp.content)
assets.append(asset)
@@ -1149,6 +1628,16 @@ class LlamaParse(BasePydanticReader):
else:
raise e
def get_tables(self, json_results: List[dict], download_path: str) -> List[str]:
if not os.path.exists(download_path):
os.makedirs(download_path)
return extract_tables_from_json_results(json_results, download_path)
async def aget_tables(
self, json_result: List[dict], download_path: str
) -> List[str]:
return await asyncio.to_thread(self.get_tables, json_result, download_path)
async def aget_xlsx(
self, json_result: List[dict], download_path: str
) -> List[dict]:
@@ -1176,7 +1665,9 @@ class LlamaParse(BasePydanticReader):
xlsx_url = (
f"{self.base_url}/api/parsing/job/{job_id}/result/raw/xlsx"
)
res = await client.get(xlsx_url)
res = await make_api_request(
client, "GET", xlsx_url, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
xlsx_list.append(xlsx)
@@ -1213,3 +1704,79 @@ class LlamaParse(BasePydanticReader):
sub_docs.append(sub_doc)
return sub_docs
async def aget_result(
self, job_id: Union[str, List[str]]
) -> Union[JobResult, List[JobResult]]:
"""
Return JobResult object for previously parsed job(s).
If the job is still pending, the result will not be returned until it is completed.
Args:
job_id: Job ID or list of multiple Job IDs to be retrieved.
Returns:
JobResult object or list of JobResult objects if multiple job IDs were provided.
"""
if isinstance(job_id, str):
result = await self._get_job_result(
job_id, ResultType.JSON.value, verbose=self.verbose
)
return JobResult(
job_id=job_id,
file_name="",
job_result=result,
api_key=self.api_key,
base_url=self.base_url,
client=self.aclient,
page_separator=self.page_separator or _DEFAULT_SEPARATOR,
)
elif isinstance(job_id, list):
results = []
jobs = [
self._get_job_result(id_, ResultType.JSON.value, verbose=self.verbose)
for id_ in job_id
]
results = await run_jobs(
jobs,
workers=self.num_workers,
desc="Getting job results",
show_progress=self.show_progress,
)
return [
JobResult(
job_id=job_id[i],
file_name="",
job_result=result,
api_key=self.api_key,
base_url=self.base_url,
client=self.aclient,
page_separator=self.page_separator or _DEFAULT_SEPARATOR,
)
for i, result in enumerate(results)
]
else:
raise ValueError("The input job_id must be a string or a list of strings.")
def get_result(
self, job_id: Union[str, List[str]]
) -> Union[JobResult, List[JobResult]]:
"""
Return JobResult object for previously parsed job(s).
If the job is still pending, the result will not be returned until it is completed.
Args:
job_id: Job ID or list of multiple Job IDs to be retrieved.
Returns:
JobResult object or list of JobResult objects if multiple job IDs were provided.
"""
try:
return asyncio_run(self.aget_result(job_id))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
+644
View File
@@ -0,0 +1,644 @@
import httpx
import os
import re
from pydantic import BaseModel, Field, SerializeAsAny
from typing import Dict, Any, List, Optional
from llama_cloud_services.parse.utils import make_api_request
from llama_index.core.async_utils import asyncio_run
from llama_index.core.schema import Document, ImageDocument, ImageNode, TextNode
PAGE_REGEX = r"page[-_](\d+)\.jpg$"
class JobMetadata(BaseModel):
"""Metadata about the job."""
job_pages: int = Field(default=0, description="The number of pages in the job.")
job_auto_mode_triggered_pages: Optional[int] = Field(
default=None,
description="The number of pages that triggered auto mode (thus increasing the cost).",
)
job_is_cache_hit: bool = Field(
default=False, description="Whether the job was a cache hit."
)
class BBox(BaseModel):
"""A bounding box."""
x: float = Field(description="The x-coordinate of the bounding box.")
y: float = Field(description="The y-coordinate of the bounding box.")
w: float = Field(description="The width of the bounding box.")
h: float = Field(description="The height of the bounding box.")
class PageItem(BaseModel):
"""An item in a page."""
type: str = Field(description="The type of the item.")
lvl: Optional[int] = Field(
default=None, description="The level of indentation of the item."
)
value: Optional[str] = Field(
default=None, description="The text content of the item."
)
md: Optional[str] = Field(
default=None, description="The markdown-formatted content of the item."
)
rows: Optional[List[List[Any]]] = Field(
default=None, description="The rows of the item."
)
bBox: Optional[BBox] = Field(
default=None, description="The bounding box of the item."
)
class ImageItem(BaseModel):
"""An image in a page."""
name: str = Field(description="The name of the image.")
height: Optional[float] = Field(
default=None, description="The height of the image."
)
width: Optional[float] = Field(default=None, description="The width of the image.")
x: Optional[float] = Field(
default=None, description="The x-coordinate of the image."
)
y: Optional[float] = Field(
default=None, description="The y-coordinate of the image."
)
original_width: Optional[int] = Field(
default=None, description="The original width of the image."
)
original_height: Optional[int] = Field(
default=None, description="The original height of the image."
)
type: Optional[str] = Field(default=None, description="The type of the image.")
class LayoutItem(BaseModel):
"""The layout of a page."""
image: str = Field(description="The name of the image containing the layout item")
confidence: float = Field(description="The confidence of the layout item.")
label: str = Field(description="The label of the layout item.")
bbox: Optional[BBox] = Field(
default=None, description="The bounding box of the layout item."
)
isLikelyNoise: bool = Field(description="Whether the layout item is likely noise.")
class ChartItem(BaseModel):
"""A chart in a page."""
name: str = Field(description="The name of the chart.")
x: Optional[float] = Field(
default=None, description="The x-coordinate of the chart."
)
y: Optional[float] = Field(
default=None, description="The y-coordinate of the chart."
)
width: Optional[float] = Field(default=None, description="The width of the chart.")
height: Optional[float] = Field(
default=None, description="The height of the chart."
)
class Page(BaseModel):
"""A page of the document."""
page: int = Field(description="The page number.")
text: Optional[str] = Field(default=None, description="The text of the page.")
md: Optional[str] = Field(default=None, description="The markdown of the page.")
images: List[ImageItem] = Field(
default_factory=list,
description="The names of the image IDs in the page, including both objects and page screenshots.",
)
charts: List[ChartItem] = Field(
default_factory=list, description="The charts in the page."
)
tables: List[str] = Field(
default_factory=list, description="The names of the table IDs in the page."
)
layout: List[LayoutItem] = Field(
default_factory=list, description="The layout of the page."
)
items: List[PageItem] = Field(
default_factory=list, description="The items in the page."
)
status: Optional[str] = Field(default=None, description="The status of the page.")
links: List[SerializeAsAny[Any]] = Field(
default_factory=list, description="The links in the page."
)
width: Optional[float] = Field(default=None, description="The width of the page.")
height: Optional[float] = Field(default=None, description="The height of the page.")
triggeredAutoMode: Optional[bool] = Field(
default=False,
description="Whether the page triggered auto mode (thus increasing the cost).",
)
parsingMode: str = Field(
default="", description="The parsing mode used for the page."
)
structuredData: Optional[Dict[str, Any]] = Field(
default=None, description="The structured data of the page."
)
noStructuredContent: bool = Field(
default=True, description="Whether the page has no structured data."
)
noTextContent: bool = Field(
default=False, description="Whether the page has no text content."
)
class JobResult(BaseModel):
"""The raw JSON result from the LlamaParse API."""
pages: List[Page] = Field(
default_factory=list, description="The pages of the document."
)
job_metadata: JobMetadata = Field(
default_factory=JobMetadata, description="The metadata of the job."
)
file_name: str = Field(
default="", description="The path to the file that was parsed."
)
job_id: str = Field(default="", description="The ID of the job.")
is_done: bool = Field(default=False, description="Whether the job is done.")
error: Optional[str] = Field(
default=None, description="The error message if the job failed."
)
def __init__(
self,
job_id: str,
file_name: str,
job_result: Dict[str, Any],
api_key: Optional[str] = None,
base_url: Optional[str] = None,
client: Optional[httpx.AsyncClient] = None,
page_separator: str = "\n\n",
):
"""
Initialize JobResult with job_id and job_result.
Args:
job_id: The job ID of the parsing task
job_result: The JSON response from the parsing job or a JobResult instance (optional)
api_key: The API key for the LlamaParse API
base_url: The base URL of the Llama Parsing API
page_separator: The separator that was used to define page splits in the result
"""
super().__init__(job_id=job_id, file_name=file_name, **job_result)
self._api_key = api_key or os.environ.get("LLAMA_CLOUD_API_KEY", "")
self._base_url = base_url or os.environ.get(
"LLAMA_CLOUD_BASE_URL", "https://api.llama-parse.ai"
)
self._client = client or httpx.AsyncClient()
self._client.base_url = self._base_url
self._client.headers["Authorization"] = f"Bearer {self._api_key}"
self._page_separator = page_separator
def get_text_documents(self, split_by_page: bool = False) -> List[Document]:
"""
Get the documents from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
if split_by_page:
return [
Document(
text=page.text,
metadata={"page_number": page.page, "file_name": self.file_name},
)
for page in self.pages
]
else:
text = self._page_separator.join(
[page.text if page.text is not None else "" for page in self.pages]
)
return [Document(text=text, metadata={"file_name": self.file_name})]
async def aget_text_documents(self, split_by_page: bool = False) -> List[Document]:
"""
Get the documents from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
# No async needed, but here for consistency
return self.get_text_documents(split_by_page)
def get_text_nodes(self, split_by_page: bool = False) -> List[TextNode]:
"""
Get the text nodes from the job.
"""
documents = self.get_text_documents(split_by_page)
return [TextNode(text=doc.text, metadata=doc.metadata) for doc in documents]
async def aget_text_nodes(self, split_by_page: bool = False) -> List[TextNode]:
"""
Get the text nodes from the job.
"""
documents = await self.aget_text_documents(split_by_page)
return [TextNode(text=doc.text, metadata=doc.metadata) for doc in documents]
def get_markdown_documents(self, split_by_page: bool = False) -> List[Document]:
"""
Get the markdown documents from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
if split_by_page:
return [
Document(
text=page.md,
metadata={"page_number": page.page, "file_name": self.file_name},
)
for page in self.pages
]
else:
return [
Document(
text=self._page_separator.join(
[page.md if page.md is not None else "" for page in self.pages]
),
metadata={"file_name": self.file_name},
)
]
async def aget_markdown_documents(
self, split_by_page: bool = False
) -> List[Document]:
"""
Get the markdown documents from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
# No async needed, but here for consistency
return self.get_markdown_documents(split_by_page)
def get_markdown_nodes(self, split_by_page: bool = False) -> List[TextNode]:
"""
Get the markdown nodes from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
documents = self.get_markdown_documents(split_by_page)
return [TextNode(text=doc.text, metadata=doc.metadata) for doc in documents]
async def aget_markdown_nodes(self, split_by_page: bool = False) -> List[TextNode]:
"""
Get the markdown nodes from the job.
Args:
split_by_page: Whether to split the pages into separate documents
"""
documents = await self.aget_markdown_documents(split_by_page)
return [TextNode(text=doc.text, metadata=doc.metadata) for doc in documents]
def get_markdown(self) -> str:
"""
Get the raw parsed markdown from the job, distinct from the markdown documents.
This does not include page separators, e.g. if merge_tables_across_pages_in_markdown is True
"""
return asyncio_run(self.aget_markdown())
async def aget_markdown(self) -> str:
"""
Get the raw parsed markdown from the job, distinct from the markdown documents.
This does not include page separators, e.g. if merge_tables_across_pages_in_markdown is True
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/raw/markdown"
response = await make_api_request(self._client, "GET", url)
return response.content.decode("utf-8")
def get_text(self) -> str:
"""
Get the raw parsed text from the job.
"""
return asyncio_run(self.aget_text())
async def aget_text(self) -> str:
"""
Get the raw parsed text from the job.
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/raw/text"
response = await make_api_request(self._client, "GET", url)
return response.content.decode("utf-8")
def get_json(self) -> Dict[str, Any]:
"""
Get the full parsed JSON result from the job.
Note:
This is not the same as JobResult.json(), which is a
JSON serialized version of the JobResult Page Documents.
"""
return asyncio_run(self.aget_json())
async def aget_json(self) -> Dict[str, Any]:
"""
Get the full parsed JSON result from the job.
Note:
This is not the same as JobResult.json(), which is a
JSON serialized version of the JobResult Page Documents.
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/json"
response = await make_api_request(self._client, "GET", url)
return response.json()
async def _get_image_document_with_bytes(
self, image: ImageItem, page: Page
) -> ImageDocument:
image_data = await self.aget_image_data(image.name)
return ImageDocument(
image=image_data,
metadata={
"page_number": page.page,
"file_name": self.file_name,
"width": image.original_width,
"height": image.original_height,
"x": image.x,
"y": image.y,
},
excluded_embed_metadata_keys=["width", "height", "x", "y"],
excluded_llm_metadata_keys=["width", "height", "x", "y"],
)
async def _get_image_document_with_path(
self, image: ImageItem, page: Page, image_download_dir: str
) -> ImageDocument:
image_path = await self.asave_image(image.name, image_download_dir)
return ImageDocument(
image_path=image_path,
metadata={
"page_number": page.page,
"file_name": self.file_name,
"width": image.original_width,
"height": image.original_height,
"x": image.x,
"y": image.y,
},
excluded_embed_metadata_keys=["width", "height", "x", "y"],
excluded_llm_metadata_keys=["width", "height", "x", "y"],
)
def get_image_documents(
self,
include_screenshot_images: bool = True,
include_object_images: bool = True,
image_download_dir: Optional[str] = None,
) -> List[ImageDocument]:
"""
Get the image documents from the job.
Args:
include_screenshot_images (bool):
Whether to include screenshot images. Default is True.
include_object_images (bool):
Whether to include object images. Default is True.
image_download_dir (Optional[str]):
The directory to save the images to. If not provided, the images will be loaded into memory.
Default is None.
"""
return asyncio_run(
self.aget_image_documents(
include_screenshot_images, include_object_images, image_download_dir
)
)
async def aget_image_documents(
self,
include_screenshot_images: bool = True,
include_object_images: bool = True,
image_download_dir: Optional[str] = None,
) -> List[ImageDocument]:
"""
Get the image documents from the job.
Args:
include_screenshot_images (bool):
Whether to include screenshot images. Default is True.
include_object_images (bool):
Whether to include object images. Default is True.
image_download_dir (Optional[str]):
The directory to save the images to. If not provided, the images will be loaded into memory.
Default is None.
"""
documents = []
for page in self.pages:
for image in page.images:
is_screenshot = re.search(PAGE_REGEX, image.name) is not None
# Skip images that don't match the inclusion criteria
if (is_screenshot and not include_screenshot_images) or (
not is_screenshot and not include_object_images
):
continue
# Get image document using appropriate method based on download_dir
get_document = (
self._get_image_document_with_path
if image_download_dir
else self._get_image_document_with_bytes
)
documents.append(
await get_document(image, page, image_download_dir) # type: ignore
if image_download_dir
else await get_document(image, page) # type: ignore
)
return documents
def get_image_nodes(
self,
include_screenshot_images: bool = True,
include_object_images: bool = True,
image_download_dir: Optional[str] = None,
) -> List[ImageNode]:
"""
Get the image nodes from the job.
Args:
include_screenshot_images (bool):
Whether to include screenshot images. Default is True.
include_object_images (bool):
Whether to include object images. Default is True.
image_download_dir (Optional[str]):
The directory to save the images to. If not provided, the images will be loaded into memory.
Default is None.
"""
documents = self.get_image_documents(
include_screenshot_images, include_object_images, image_download_dir
)
return [
ImageNode(
image=doc.image,
image_path=doc.image_path,
image_url=doc.image_url,
metadata=doc.metadata,
)
for doc in documents
]
async def aget_image_nodes(
self,
include_screenshot_images: bool = True,
include_object_images: bool = True,
image_download_dir: Optional[str] = None,
) -> List[ImageNode]:
"""
Get the image nodes from the job.
Args:
include_screenshot_images (bool):
Whether to include screenshot images. Default is True.
include_object_images (bool):
Whether to include object images. Default is True.
image_download_dir (Optional[str]):
The directory to save the images to. If not provided, the images will be loaded into memory.
Default is None.
"""
documents = await self.aget_image_documents(
include_screenshot_images, include_object_images, image_download_dir
)
return [
ImageNode(
image=doc.image,
image_path=doc.image_path,
image_url=doc.image_url,
metadata=doc.metadata,
)
for doc in documents
]
async def aget_image_data(self, image_name: str) -> bytes:
"""
Get image data by name using the job ID.
Args:
image_name: The name of the image to fetch
Returns:
The image data as bytes
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/image/{image_name}"
response = await make_api_request(self._client, "GET", url)
return response.content
def get_image_data(self, image_name: str) -> bytes:
"""
Get image data by name using the job ID (synchronous version).
Args:
image_name: The name of the image to fetch
Returns:
The image data as bytes
"""
return asyncio_run(self.aget_image_data(image_name))
async def aget_xlsx_data(self) -> bytes:
"""
Get the XLSX data for the job.
Returns:
The XLSX data as bytes
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/xlsx"
response = await make_api_request(self._client, "GET", url)
return response.content
def get_xlsx_data(self) -> bytes:
"""
Get the XLSX data for the job (synchronous version).
Returns:
The XLSX data as bytes
"""
return asyncio_run(self.aget_xlsx_data())
async def asave_image(self, image_name: str, output_dir: str) -> str:
"""
Save an image to a file.
Args:
image_name: The name of the image to fetch
output_dir: The directory to save the image to
Returns:
The path to the saved image
"""
image_data = await self.aget_image_data(image_name)
# Create output directory if it doesn't exist
os.makedirs(output_dir, exist_ok=True)
# Save image to file
output_path = os.path.join(output_dir, image_name)
with open(output_path, "wb") as f:
f.write(image_data)
return output_path
def save_image(self, image_name: str, output_dir: str) -> str:
"""
Save an image to a file (synchronous version).
Args:
image_name: The name of the image to fetch
output_dir: The directory to save the image to
Returns:
The path to the saved image
"""
return asyncio_run(self.asave_image(image_name, output_dir))
def get_image_names(self) -> List[str]:
"""
Get the names of all images in the job.
Returns:
A list of image names
"""
return [image.name for page in self.pages for image in page.images]
async def asave_all_images(self, output_dir: str) -> List[str]:
"""
Save all images to files.
Args:
output_dir: The directory to save the images to
Returns:
A list of paths to the saved images
"""
image_names = self.get_image_names()
saved_paths = []
for name in image_names:
path = await self.asave_image(name, output_dir)
saved_paths.append(path)
return saved_paths
def save_all_images(self, output_dir: str) -> List[str]:
"""
Save all images to files (synchronous version).
Args:
output_dir: The directory to save the images to
Returns:
A list of paths to the saved images
"""
return asyncio_run(self.asave_all_images(output_dir))
+383
View File
@@ -0,0 +1,383 @@
import httpx
import itertools
import logging
import os
from enum import Enum
from pathlib import Path
from datetime import datetime
from tenacity import (
retry,
stop_after_attempt,
wait_exponential,
retry_if_exception,
before_sleep_log,
)
from typing import Any, Iterable, Iterator, Optional, List, cast
logger = logging.getLogger(__name__)
# Asyncio error messages
nest_asyncio_err = "cannot be called from a running event loop"
nest_asyncio_msg = "The event loop is already running. Add `import nest_asyncio; nest_asyncio.apply()` to your code to fix this issue."
class ResultType(str, Enum):
"""The result type for the parser."""
TXT = "text"
MD = "markdown"
JSON = "json"
STRUCTURED = "structured"
class ParsingMode(str, Enum):
"""The parsing mode for the parser."""
parse_page_without_llm = "parse_page_without_llm"
parse_page_with_llm = "parse_page_with_llm"
parse_page_with_lvm = "parse_page_with_lvm"
parse_page_with_agent = "parse_page_with_agent"
parse_document_with_llm = "parse_document_with_llm"
parse_document_with_agent = "parse_document_with_agent"
class FailedPageMode(str, Enum):
"""
Enum for representing the different available page error handling modes
"""
raw_text = "raw_text"
blank_page = "blank_page"
error_message = "error_message"
class Language(str, Enum):
BAZA = "abq"
ADYGHE = "ady"
AFRIKAANS = "af"
ANGIKA = "ang"
ARABIC = "ar"
ASSAMESE = "as"
AVAR = "ava"
AZERBAIJANI = "az"
BELARUSIAN = "be"
BULGARIAN = "bg"
BIHARI = "bh"
BHOJPURI = "bho"
BENGALI = "bn"
BOSNIAN = "bs"
SIMPLIFIED_CHINESE = "ch_sim"
TRADITIONAL_CHINESE = "ch_tra"
CHECHEN = "che"
CZECH = "cs"
WELSH = "cy"
DANISH = "da"
DARGWA = "dar"
GERMAN = "de"
ENGLISH = "en"
SPANISH = "es"
ESTONIAN = "et"
PERSIAN_FARSI = "fa"
FRENCH = "fr"
IRISH = "ga"
GOAN_KONKANI = "gom"
HINDI = "hi"
CROATIAN = "hr"
HUNGARIAN = "hu"
INDONESIAN = "id"
INGUSH = "inh"
ICELANDIC = "is"
ITALIAN = "it"
JAPANESE = "ja"
KABARDIAN = "kbd"
KANNADA = "kn"
KOREAN = "ko"
KURDISH = "ku"
LATIN = "la"
LAK = "lbe"
LEZGHIAN = "lez"
LITHUANIAN = "lt"
LATVIAN = "lv"
MAGAHI = "mah"
MAITHILI = "mai"
MAORI = "mi"
MONGOLIAN = "mn"
MARATHI = "mr"
MALAY = "ms"
MALTESE = "mt"
NEPALI = "ne"
NEWARI = "new"
DUTCH = "nl"
NORWEGIAN = "no"
OCCITAN = "oc"
PALI = "pi"
POLISH = "pl"
PORTUGUESE = "pt"
ROMANIAN = "ro"
RUSSIAN = "ru"
SERBIAN_CYRILLIC = "rs_cyrillic"
SERBIAN_LATIN = "rs_latin"
NAGPURI = "sck"
SLOVAK = "sk"
SLOVENIAN = "sl"
ALBANIAN = "sq"
SWEDISH = "sv"
SWAHILI = "sw"
TAMIL = "ta"
TABASSARAN = "tab"
TELUGU = "te"
THAI = "th"
TAJIK = "tjk"
TAGALOG = "tl"
TURKISH = "tr"
UYGHUR = "ug"
UKRAINIAN = "uk"
URDU = "ur"
UZBEK = "uz"
VIETNAMESE = "vi"
SUPPORTED_FILE_TYPES = [
".pdf",
# document and presentations
".602",
".abw",
".cgm",
".cwk",
".doc",
".docx",
".docm",
".dot",
".dotm",
".hwp",
".key",
".lwp",
".mw",
".mcw",
".pages",
".pbd",
".ppt",
".pptm",
".pptx",
".pot",
".potm",
".potx",
".rtf",
".sda",
".sdd",
".sdp",
".sdw",
".sgl",
".sti",
".sxi",
".sxw",
".stw",
".sxg",
".txt",
".uof",
".uop",
".uot",
".vor",
".wpd",
".wps",
".xml",
".zabw",
".epub",
# images
".jpg",
".jpeg",
".png",
".gif",
".bmp",
".svg",
".tiff",
".webp",
# web
".htm",
".html",
# spreadsheets
".xlsx",
".xls",
".xlsm",
".xlsb",
".xlw",
".csv",
".dif",
".sylk",
".slk",
".prn",
".numbers",
".et",
".ods",
".fods",
".uos1",
".uos2",
".dbf",
".wk1",
".wk2",
".wk3",
".wk4",
".wks",
".123",
".wq1",
".wq2",
".wb1",
".wb2",
".wb3",
".qpw",
".xlr",
".eth",
".tsv",
".mp3",
".mp4",
".mpeg",
".mpga",
".m4a",
".wav",
".webm",
]
def should_retry(exception: Exception) -> bool:
"""Check if the exception should be retried.
Args:
exception: The exception to check.
"""
# Retry on connection errors (network issues)
if isinstance(
exception,
(
httpx.ConnectError,
httpx.ConnectTimeout,
httpx.ReadTimeout,
httpx.WriteTimeout,
httpx.RemoteProtocolError,
),
):
return True
# Retry on specific HTTP status codes
if isinstance(exception, httpx.HTTPStatusError):
status_code = exception.response.status_code
# Retry on rate limiting or temporary server errors
return status_code in (429, 500, 502, 503, 504)
return False
async def make_api_request(
client: httpx.AsyncClient,
method: str,
url: str,
timeout: float = 60.0,
max_retries: int = 5,
**httpx_kwargs: Any,
) -> httpx.Response:
"""Make an retrying API request to the LlamaParse API.
Args:
client: The httpx.AsyncClient to use for the request.
url: The URL to request.
headers: The headers to include in the request.
timeout: The timeout for the request.
max_retries: The maximum number of retries for the request.
"""
@retry(
stop=stop_after_attempt(max_retries),
wait=wait_exponential(multiplier=1, min=4, max=timeout),
retry=retry_if_exception(should_retry),
before_sleep=before_sleep_log(logger, logging.WARNING),
)
async def _make_request(url: str, **httpx_kwargs: Any) -> httpx.Response:
if method == "GET":
response = await client.get(url, **httpx_kwargs)
elif method == "POST":
response = await client.post(url, **httpx_kwargs)
else:
raise ValueError(f"Invalid method: {method}")
response.raise_for_status()
return response
return await _make_request(url, **httpx_kwargs)
def expand_target_pages(target_pages: str) -> Iterator[int]:
"""Yield all values in target_pages."""
for target in target_pages.strip().split(","):
if "-" in target:
try:
start, end = map(int, target.strip().split("-"))
if start > end:
raise ValueError
yield from range(start, end + 1)
except ValueError as e:
raise ValueError(f"Invalid page range: {target}") from e
else:
try:
yield int(target)
except ValueError as e:
raise ValueError(f"Invalid page number: {target}") from e
def partition_pages(
pages: Iterable[int], size: int, max_pages: Optional[int] = None
) -> Iterator[str]:
"""Yield partitioned target_pages segments."""
if size < 1:
raise ValueError(f"Invalid partition segment size: {size}")
if max_pages is not None and max_pages < 1:
raise ValueError("Max pages must be > 0")
it = iter(pages)
total = 0
while max_pages is None or total < max_pages:
segment = tuple(itertools.islice(it, size))
if segment:
targets = []
for _k, g in itertools.groupby(enumerate(segment), lambda x: x[0] - x[1]):
group = [item[1] for item in g]
if len(group) > 1:
start, end = group[0], group[-1]
group_size = end - start + 1
if max_pages is not None and total + group_size > max_pages:
end -= total + group_size - max_pages
group_size = end - start + 1
if group_size > 1:
targets.append(f"{start}-{end}")
else:
targets.append(str(start))
total += group_size
else:
targets.append(str(group[0]))
total += 1
yield ",".join(targets)
else:
return
def extract_tables_from_json_results(
json_results: List[dict], download_path: str
) -> List[str]:
tables = []
for json_result in json_results:
pages = json_result["pages"]
for page in pages:
page = cast(dict, page)
items = page.get("items", [])
if items:
for i, item in enumerate(items):
item = cast(dict, item)
if item.get("type", "") == "table" and item.get("csv", ""):
savepath = os.path.join(
download_path,
f"table_{datetime.now().strftime('%Y_%d_%m_%H_%M_%S_%f')[:-3]}.csv",
)
if Path(savepath).exists():
savepath = (
savepath.replace(".csv", "_")[0] + str(i) + ".csv"
)
with open(savepath, "w") as f:
f.write(item["csv"])
tables.append(savepath)
return tables
+29
View File
@@ -0,0 +1,29 @@
import difflib
from pydantic import BaseModel
from typing import Any, Dict, List, Tuple, Type
def check_extra_params(
model_cls: Type[BaseModel], data: Dict[str, Any]
) -> Tuple[List[str], List[str]]:
# check if one of the parameters is unused, and warn the user
model_attributes = set(model_cls.model_fields.keys())
extra_params = [param for param in data.keys() if param not in model_attributes]
suggestions: List[str] = []
if extra_params:
# for each unused parameter, check if it is similar to a valid parameter and suggest a typo correction, else suggest to check the documentation / update the package
for param in extra_params:
similar_params = difflib.get_close_matches(
param, model_attributes, n=1, cutoff=0.8
)
if similar_params:
suggestions.append(
f"'{param}' is not a valid parameter. Did you mean '{similar_params[0]}' instead of '{param}'?"
)
else:
suggestions.append(
f"'{param}' is not a valid parameter. Please check the documentation or update the package."
)
return extra_params, suggestions
@@ -146,9 +146,9 @@ Full documentation for `SimpleDirectoryReader` can be found on the [LlamaIndex D
Several end-to-end indexing examples can be found in the examples folder
- [Getting Started](/examples/parse/demo_basic.ipynb)
- [Advanced RAG Example](/examples/parse/demo_advanced.ipynb)
- [Raw API Usage](/examples/parse/demo_api.ipynb)
- [Getting Started](/docs/examples-py/parse/demo_basic.ipynb)
- [Advanced RAG Example](/docs/examples-py/parse/demo_advanced.ipynb)
- [Raw API Usage](/docs/examples-py/parse/demo_api.ipynb)
## Documentation
+8
View File
@@ -0,0 +1,8 @@
from llama_cloud_services.parse import (
LlamaParse,
ResultType,
ParsingMode,
FailedPageMode,
)
__all__ = ["LlamaParse", "ResultType", "ParsingMode", "FailedPageMode"]
@@ -1,6 +1,8 @@
from llama_cloud_services.parse.base import (
LlamaParse,
ResultType,
ParsingMode,
FailedPageMode,
FileInput,
_DEFAULT_SEPARATOR,
JOB_RESULT_URL,
@@ -12,6 +14,8 @@ __all__ = [
"LlamaParse",
"ResultType",
"FileInput",
"ParsingMode",
"FailedPageMode",
"_DEFAULT_SEPARATOR",
"JOB_RESULT_URL",
"JOB_STATUS_ROUTE",

Some files were not shown because too many files have changed in this diff Show More