Compare commits

...

106 Commits

Author SHA1 Message Date
Logan 0916f807bc Update README.md 2025-07-30 14:55:08 -06:00
Harshit Budhiraja e42746e372 docs(readme): update hyperlinks to correct targets (#820) 2025-07-30 14:53:43 -06:00
Clelia (Astra) Bertelli 3149dfd03a fix: no git checks on pnpm publish (#823) 2025-07-30 21:25:23 +02:00
Clelia (Astra) Bertelli e499fdbdab fix: add release to NPM (#819) 2025-07-30 20:55:41 +02:00
Clelia (Astra) Bertelli e57df39248 Merge index into main (#821)
* wip: monorepo changes

* fix ci for the time being

* fix ci for the time being pt2

* wip: first cloud refactoring for ts

* chore: restore original package

* fix: imports, package.json, tsconfig.json, client, reader

* feat: adjustments after local testing

* ci: github actions for typescript

* ci: typescript ci

* ci: nvmrc 🤦

* ci: remove cache 🤦

* ci: actions

* ci: actions (i lost count)

* ci: pnpm run format

* ci: pnpm run format

* chore: migrate llama-parse to uv

* add tests

* remove unneeded readme

* update workflows

* feat: modify py release workflow, adding uv version, bump version for llama-cloud-services to latest

* uv lock

* ci: python tests all tests

* fix: lock file pulling in wrong version of numpy

* feat: add index to llama-cloud-services (#817)

---------

Co-authored-by: Logan Markewich <logan.markewich@live.com>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-07-30 19:46:36 +02:00
Clelia (Astra) Bertelli 09b192b98b Adding TS llama-cloud-services and moving llama-parse to uv (#811)
* wip: monorepo changes

* fix ci for the time being

* fix ci for the time being pt2

* wip: first cloud refactoring for ts

* chore: restore original package

* fix: imports, package.json, tsconfig.json, client, reader

* feat: adjustments after local testing

* ci: github actions for typescript

* ci: typescript ci

* ci: nvmrc 🤦

* ci: remove cache 🤦

* ci: actions

* ci: actions (i lost count)

* ci: pnpm run format

* ci: pnpm run format

* chore: migrate llama-parse to uv

* add tests

* remove unneeded readme

* update workflows

* feat: modify py release workflow, adding uv version, bump version for llama-cloud-services to latest

* uv lock

* ci: python tests all tests

* fix: lock file pulling in wrong version of numpy

---------

Co-authored-by: Logan Markewich <logan.markewich@live.com>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-07-30 17:59:08 +02:00
Adrian Lyjak 13f01a0621 Adding support for page citations, and refactor the confidence into the field metadata (#815) 2025-07-30 10:55:29 -04:00
Javier Torres cf879a1a58 Bump llama-cloud version (#814) 2025-07-28 16:06:31 -05:00
Tuana Çelik fcdf2ab63e Fixes to multimodal report generation (#809) 2025-07-23 16:28:53 -06:00
Adrian Lyjak 083d8109c2 Make versioning a little easier, and fix llama_parse version (#808)
* Make versioning a little easier

* fix up ci
2025-07-21 18:49:07 -04:00
Adrian Lyjak 89cfc8b25f feat: default to _public agent data (#803)
* feat: default to _public agent data
* version bump
2025-07-21 15:58:03 -04:00
Peter Rowlands (변기호) c46e157f92 parse: expose preserve_very_small_text option (#806) 2025-07-21 14:19:15 +09:00
Peter Rowlands (변기호) 05d6026d37 bump version to v0.6.50 (#802) 2025-07-18 18:59:25 +09:00
Peter Rowlands (변기호) 8e98d5c146 parse: expose functionality to get raw job results (#801)
* add LlamaParse.get_result()

* add JobResult.get_text/get_markdown/get_json

* add tests
2025-07-18 18:50:29 +09:00
Adrian Lyjak 3f311c0669 Bump v0.6.49 (#797) 2025-07-16 19:42:09 -04:00
Adrian Lyjak b1a2f9d42b Add new method to fetch the full, non-paginated markdown (#796)
Add new method to fetch the full, non-paginated markdown for proper merge_tables_across_pages_in_markdown support
2025-07-16 19:29:57 -04:00
Neeraj Pradhan 142f55c94c Update to version 0.6.48 (#795)
* Update to version 0.6.48

* pin version

* poetry lock

* adjust warnings

* collect all agents for cleanup
2025-07-16 13:24:44 -07:00
Clelia (Astra) Bertelli 230a110e52 chore: vbump to 0.6.47 and example notebook (#794)
* chore: vbump to 0.6.47 and example notebook

* chore: update llama-parse pyproject.toml
2025-07-16 19:08:44 +02:00
Clelia (Astra) Bertelli 83e2b031cd feat: add table extraction for LlamaParse as CSV files (#793)
* feat: add table extraction for LlamaParse as CSV files

* chore: poetry lock

* chore: add tests

* fix: handle the case where no tables are present

* chore: implement suggestions
2025-07-16 17:08:09 +02:00
Adrian Lyjak 4844e26e5c Improve Agent Data interface, and add file related fields to extracted data for file tracking (#785)
Add file related fields for file tracking. Simplify API
2025-07-09 14:27:24 -04:00
Pierre-Loic Doulcet 70a049af3c merge_tables_across_pages_in_markdown parse parameter (#786)
* merge_tables_across_pages_in_markdown parse parameter

* base.py
2025-07-09 19:03:48 +02:00
Adrian Lyjak dc11776c86 Add nicer hand-written agent data interface (#782)
* Add nicer hand-written agent data interface

* bump to 0.6.44
2025-07-08 17:49:00 -04:00
Logan 2448a42b90 relax pydantic job object (#784) 2025-07-08 12:12:56 -06:00
Neeraj Pradhan c75a900174 Bump up version to 0.6.42 (#783) 2025-07-08 09:16:46 -07:00
Peter Rowlands (변기호) 2fb7adfe0e parse: loosen PageItem.rows type hint (v0.6.41) (#776)
* parse: loosen PageItem.rows type hint

* bump version to 0.6.41
2025-06-30 21:47:40 +09:00
Pierre-Loic Doulcet dc82270724 header footer control in llamaparse (#775) 2025-06-30 16:02:59 +08:00
Neeraj Pradhan d880a48dd0 Bump to version 0.6.39 (#772)
* Bump to version 0.6.39

* lock file update
2025-06-27 16:04:40 -07:00
Logan 7567e8b45e except one more error type (#771) 2025-06-27 10:17:57 -06:00
Neeraj Pradhan 0d59a90151 Relax tenacity version; bump up version to 0.6.37 (#769) 2025-06-25 15:32:20 -07:00
Neeraj Pradhan 98ad550b1a Manage extract agent lifecycle in pytest (#766) 2025-06-24 08:59:38 -07:00
Neeraj Pradhan b58f43ce9f Bump up version to 0.6.36 (#763) 2025-06-23 14:26:05 -07:00
Neeraj Pradhan acf6adcd91 Make job fetching more robust to connection errors (#764) 2025-06-23 13:17:28 -07:00
Neeraj Pradhan daf6576c3c Bump version to 0.6.35 (#762) 2025-06-20 09:33:21 -07:00
Logan 8caa4defa6 fix partition (#758) 2025-06-16 17:37:52 -06:00
Pierre-Loic Doulcet 26918b8de4 add high_res_ocr to the package (#757) 2025-06-16 16:28:23 +08:00
Pierre-Loic Doulcet 6fb5ebe2f9 6.32 warning on unused parameters (#755) 2025-06-12 22:35:48 -06:00
dependabot[bot] c0aa67995b Bump requests from 2.32.3 to 2.32.4 in /llama_parse (#754) 2025-06-10 18:14:44 -06:00
dependabot[bot] 9f841f8328 Bump tornado from 6.4.2 to 6.5.1 in /llama_parse (#753) 2025-06-10 18:14:35 -06:00
dependabot[bot] 99c75eece9 Bump h11 from 0.14.0 to 0.16.0 in /llama_parse (#752) 2025-06-10 18:14:27 -06:00
Logan 57d2586ee3 v0.6.31 (#751) 2025-06-10 17:58:36 -06:00
Jerry Liu 4280a43ec8 add multi-fund analysis notebook (#739) 2025-06-07 11:25:25 -07:00
Neeraj Pradhan 7f1082bbb2 Bump to version 0.6.30 (#748) 2025-06-05 14:34:20 -07:00
Simon Suo 57cfc45804 Directly pass None project_id (#743) 2025-06-05 14:16:54 -07:00
Soumil.Binhani 30e8913875 0.6.29: Standerdize the parsing input format for both .aget_json() and .aload_data() (#745) 2025-06-05 10:58:07 -06:00
Logan 0ce6d4d7a4 more optional types marked (#747) 2025-06-05 10:50:29 -06:00
Peter Rowlands (변기호) 584ba8d48e 0.6.28: fix job result format after partitioning changes (#741)
* parse: fix job result format

* bump to 0.6.28
2025-06-02 15:25:30 -07:00
Peter Rowlands (변기호) 925805ee11 parse: support partitioning files before parsing (#709)
* parse: add utils for handling target_pages

* parse: support partitioning docs into multiple parse jobs

* tests: add tests for partitioned parse

* drop unneeded get_job_result call

* add parse JobFailedException and expected error handling

* bump to 0.6.27
2025-06-02 12:27:58 -07:00
Logan 76fb73c971 v0.6.26 (#740) 2025-06-02 09:59:45 -06:00
Abhik Bhattacharjee 6d19ea9ac0 parse: fix the "model" parameter mismatch between playground and Python client (#737) 2025-06-02 09:35:30 -06:00
Pierre-Loic Doulcet 90431090e9 0.6.25 outlined_table_extraction (#736) 2025-05-30 11:37:21 +02:00
Neeraj Pradhan 6dff35b204 Add notebook for Form 4 extraction (#731)
* Add notebook for Form 4 extraction

* fix comments

* heavier caching; add mermaid diag

* add output directory

* save notebook
2025-05-29 18:31:56 -07:00
Logan e634c7978d v0.6.24 (#732) 2025-05-28 20:11:51 -06:00
Neeraj Pradhan 7a9e99bba2 Bump to version 0.6.23 (#729) 2025-05-20 09:43:06 -07:00
Adrian Lyjak efcdd4405b Pass through verify and timeout config to the extraction agent (#726) 2025-05-17 12:51:16 -07:00
Javier Torres bf3614690f Remove credits from parse metadata (#720) 2025-05-09 16:03:09 -05:00
Logan 7463e00da3 v0.6.22 (#718) 2025-05-08 11:44:41 -06:00
Tuana Çelik cbe9de0c57 Adding example for extracting with citations (#716)
* Adding example for extracting with citations

* removing TOC and installation output
2025-05-06 23:32:17 +02:00
Logan a023507d42 even more optional (#711) 2025-05-01 15:52:38 -06:00
Peter Rowlands (변기호) e48f544ddc parse: fix num_workers/parse job batching (#708) 2025-05-01 09:30:35 -06:00
Logan 4aa7ad5642 v0.6.20 (#707) 2025-04-29 08:53:55 -06:00
Sacha Bron c39cdbcd01 v0.6.19 (#706) 2025-04-29 12:28:21 +02:00
Pierre-Loic Doulcet 71eaa8bcc6 add auto_mode_configuration_jon for llamaParse (#704) 2025-04-29 12:23:03 +02:00
Pierre-Loic Doulcet 1e1cbdfc79 add support for presets (#703) 2025-04-29 11:54:54 +08:00
Logan cc8af4a43a make original height + width optional in the parse result (#702) 2025-04-27 18:31:35 -06:00
dependabot[bot] 43fbd48ab8 Bump actions/setup-python from 4 to 5 (#701)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 4 to 5.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v4...v5)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-04-27 13:08:44 -06:00
dependabot[bot] 5ec66e9452 Bump actions/checkout from 3 to 4 (#700)
Bumps [actions/checkout](https://github.com/actions/checkout) from 3 to 4.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v3...v4)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-04-27 13:08:31 -06:00
Scott Brenner 211521c82e Dependabot configuration to update actions in workflow (#698) 2025-04-27 12:52:11 -06:00
Scott Brenner 4ddaab1efb Refactor CodeQL workflow (#699)
* Refactor CodeQL workflow

* Update .github/workflows/codeql.yml
2025-04-27 12:51:56 -06:00
Neeraj Pradhan 53e5ce2e83 Bump to v0.6.16 (#697) 2025-04-25 14:39:52 -07:00
Neeraj Pradhan 9f4bd1cb64 Update to latest version of llama-cloud (#696)
update to latest version of llama-cloud
2025-04-25 14:14:49 -07:00
Logan 456863752b small enum nit for FailedPageMode (#693) 2025-04-23 21:34:26 -06:00
Pierre-Loic Doulcet c2dc34bbd6 Page error parameters (#691) 2025-04-23 20:47:57 -06:00
Logan fcabb04baf skip llama-report tests in cicd (#692)
* skip llama-report tests in cicd

* skip llama-report tests in cicd
2025-04-23 20:47:00 -06:00
Sacha Bron 8e7c32d3d6 Add markdown_table_multiline_header_separator support (#683)
* Add markdown_table_multiline_header_separator support

* Lint
2025-04-15 17:39:46 +02:00
Neeraj Pradhan 7e3013d914 Use unique filename to avoid db collisions later (#682)
* Use unique filename to avoid db collisions later

* add xfail marker to test_create_and_delete_report
2025-04-11 11:03:15 -07:00
Logan 4a664c33d2 parse readme nits (#681) 2025-04-10 19:25:06 -06:00
Logan 6d049ee2e4 v0.6.12 (#680) 2025-04-10 19:18:49 -06:00
Logan fa73e73664 new result object (#650) 2025-04-10 19:17:23 -06:00
Neeraj Pradhan bf67ee6056 Update docs for LlamaExtract (#679) 2025-04-10 12:16:32 -07:00
Neeraj Pradhan a1abef2ee9 Bump version to v0.6.11 (#678) 2025-04-10 11:23:06 -07:00
Neeraj Pradhan a753e01d3c Support text as input directly in the SDK (#676) 2025-04-09 21:40:56 -07:00
Logan 9b15065b24 v0.6.10 (#677) 2025-04-09 19:30:59 -06:00
Pierre-Loic Doulcet 6e4150537c Add compact_markdown_table parameter (#675) 2025-04-09 19:19:19 -06:00
Neeraj Pradhan 233d715a14 Better connection management on llamaextract client (#674) 2025-04-09 14:26:52 -07:00
Neeraj Pradhan 77ac385dfe Fix bytes input for LlamaExtract (#673)
* Fix bytes input for LlamaExtract

* backwards compatibility

* compat python 3.9
2025-04-09 10:37:22 -07:00
Neeraj Pradhan 53b78fcd7d Rename test endpoint to match functionality (#668) 2025-04-08 17:42:20 -07:00
Jerry Liu 16f81bd7ee add due diligence notebook (#670) 2025-04-08 09:13:11 -07:00
Marplex 0ee049fd11 Add layout agent mode visual citation demo notebook (#672) 2025-04-07 09:54:06 -06:00
Neeraj Pradhan 7dba17e5bc Update extract.md (#671) 2025-04-06 22:18:03 -07:00
Jerry Liu eeb678b937 solar panel extraction workflow (#667)
* cr

* cr

* cr
2025-04-02 17:28:13 -07:00
Emanuel Ferreira fe4eb664fd chore: add base url documentation (#666)
* wip

* newline

* wip

* docs
2025-04-01 18:43:17 -03:00
Jerry Liu 257720e443 fix notebook (#665)
cr
2025-04-01 08:05:34 -07:00
Jerry Liu e7afaedf3e create llamaextract demo with lm317 datasheet (#664) 2025-03-31 17:38:24 -07:00
Neeraj Pradhan b66b47a708 Bump to version 0.6.9 (#663)
* Bump to version 0.6.8

* add banks as dep

* Add platformdirs to poetry

* Fix version number
2025-03-28 17:07:46 -07:00
George He fe485ff62e fix:Add retry handling to parse and backoff patterns - catching 5XX errors and HTTP errors (#648)
* Add parse retry logic

* Update code cleanliness

* Update errors

* Fix lint

* Fix backoff strategies

* Update docs

* Fix errors

* Add base
2025-03-26 12:09:56 +01:00
Pierre-Loic Doulcet 1ebe1cee67 Add new parameter, fix parse_mode (#660)
* update with new parameters

* lint
2025-03-25 11:14:37 +01:00
Neeraj Pradhan e9252eb48a Update notebook for extract (#658) 2025-03-22 09:34:40 -07:00
Neeraj Pradhan dad7728135 Bump to version 0.6.7 (#656) 2025-03-21 21:26:54 -07:00
Neeraj Pradhan c5111e3335 Revert httpx_client as argument (#657) 2025-03-21 21:16:56 -07:00
Neeraj Pradhan bbbdb98362 Add provision for custom httpx client for LlamaExtract (#654) 2025-03-21 11:37:40 -07:00
Neeraj Pradhan 60cdc2af84 Add xfail for timeout errors in report gen (#655) 2025-03-21 11:06:49 -07:00
Neeraj Pradhan 344c20f331 Bump up version for release (#652) 2025-03-18 15:54:32 -07:00
Neeraj Pradhan 2b0496e947 Update llama cloud for extract endpoints (#651) 2025-03-18 15:43:43 -07:00
Laurie Voss 6c63dba6fb Typos and removing staging URL (#647) 2025-03-13 08:11:09 -07:00
Neeraj Pradhan 734c021a2e Add notebook for extraction from SEC 10-K/Q filings (#646)
* Add notebook for extraction from SEC 10-K/Q filings

* Add notebook for 10 k/q extraction

* Remove unnecessary cell

* fix file link

* fix code rendering

* Add notes for clarity

* fix notes
2025-03-12 20:42:17 -07:00
Neeraj Pradhan eeb034896f Bump to version 0.6.5 (updating llama-cloud dependency) (#645)
* Bump to version 0.6.5 (updating llama-cloud dependency)

* fix other endpoints
2025-03-06 18:22:42 -08:00
260 changed files with 124458 additions and 13615 deletions
+11
View File
@@ -0,0 +1,11 @@
# Please see the documentation for all configuration options:
# https://docs.github.com/github/administering-a-repository/configuration-options-for-dependency-updates
# and
# https://docs.github.com/code-security/dependabot/dependabot-version-updates/configuration-options-for-the-dependabot.yml-file
version: 2
updates:
- package-ecosystem: "github-actions"
directory: "/"
schedule:
interval: "weekly"
-48
View File
@@ -1,48 +0,0 @@
name: Build Package
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install
- name: Ensure lock works
shell: bash
run: poetry lock
- name: Build
shell: bash
run: poetry build
- name: Test installing built package
shell: bash
run: python -m pip install .
- name: Test import
shell: bash
working-directory: ${{ vars.RUNNER_TEMP }}
run: python -c "import llama_cloud_services"
+50
View File
@@ -0,0 +1,50 @@
name: Build Package - Python
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Test import
working-directory: py
run: uv run -- python -c "import llama_cloud_services"
+28
View File
@@ -0,0 +1,28 @@
name: Build Package - TypeScript
on: [pull_request]
jobs:
pre_release:
name: Pre Release
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Build
working-directory: ts/llama_cloud_services/
run: pnpm run build
+8 -48
View File
@@ -1,14 +1,3 @@
# For most projects, this workflow file will not need changing; you simply need
# to commit it to your repository.
#
# You may wish to alter this file to override the set of languages analyzed,
# or to provide custom queries or build logic.
#
# ******** NOTE ********
# We have attempted to detect the languages in your repository. Please check
# the `language` matrix defined below to confirm you have the correct set of
# supported CodeQL languages.
#
name: "CodeQL"
on:
@@ -28,54 +17,25 @@ jobs:
# - https://gh.io/supported-runners-and-hardware-resources
# - https://gh.io/using-larger-runners
# Consider using larger runners for possible analysis time improvements.
runs-on: ${{ (matrix.language == 'swift' && 'macos-latest') || 'ubuntu-latest' }}
timeout-minutes: ${{ (matrix.language == 'swift' && 120) || 360 }}
runs-on: "ubuntu-latest"
timeout-minutes: 360
permissions:
actions: read
contents: read
security-events: write
strategy:
fail-fast: false
matrix:
language: ["python"]
# CodeQL supports [ 'cpp', 'csharp', 'go', 'java', 'javascript', 'python', 'ruby', 'swift' ]
# Use only 'java' to analyze code written in Java, Kotlin or both
# Use only 'javascript' to analyze code written in JavaScript, TypeScript or both
# Learn more about CodeQL language support at https://aka.ms/codeql-docs/language-support
steps:
- name: Checkout repository
uses: actions/checkout@v3
uses: actions/checkout@v4
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v2
uses: github/codeql-action/init@v3
with:
languages: ${{ matrix.language }}
# If you wish to specify custom queries, you can do so here or in a config file.
# By default, queries listed here will override any specified in a config file.
# Prefix the list here with "+" to use these queries and those in the config file.
# For more details on CodeQL's query packs, refer to: https://docs.github.com/en/code-security/code-scanning/automatically-scanning-your-code-for-vulnerabilities-and-errors/configuring-code-scanning#using-queries-in-ql-packs
# queries: security-extended,security-and-quality
# Autobuild attempts to build any compiled languages (C/C++, C#, Go, Java, or Swift).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@v2
# ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
# If the Autobuild fails above, remove it and uncomment the following three lines.
# modify them (or add more) to build your code if your project, please refer to the EXAMPLE below for guidance.
# - run: |
# echo "Run, Build Application using script"
# ./location_of_script_within_repo/buildscript.sh
languages: python
dependency-caching: true
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v2
uses: github/codeql-action/analyze@v3
with:
category: "/language:${{matrix.language}}"
category: "/language:python"
-37
View File
@@ -1,37 +0,0 @@
name: Linting
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install pre-commit
shell: bash
run: poetry run pip install pre-commit
- name: Run linter
shell: bash
run: poetry run make lint
+35
View File
@@ -0,0 +1,35 @@
name: Lint - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }}
- name: Run linter
shell: bash
working-directory: py
run: uv run -- pre-commit run -a
+36
View File
@@ -0,0 +1,36 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run lint
working-directory: ts/llama_cloud_services/
run: pnpm run lint
- name: Run Prettier
working-directory: ts/llama_cloud_services/
run: pnpm run format
-75
View File
@@ -1,75 +0,0 @@
name: Publish llama-parse to PyPI / GitHub
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
POETRY_VERSION: "1.6.1"
PYTHON_VERSION: "3.9"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ env.PYTHON_VERSION }}
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: pip install -e .
- name: Build and publish llama-cloud-services
uses: JRubics/poetry-publish@v2.1
with:
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
poetry_install_options: "--without dev"
- name: Build and publish llama-parse
uses: JRubics/poetry-publish@v2.1
with:
working_directory: "llama_parse"
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
poetry_install_options: "--without dev"
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }}
draft: false
prerelease: false
- name: Get Asset name
run: |
export PKG=$(ls dist/ | grep tar)
set -- $PKG
echo "name=$1" >> $GITHUB_ENV
- name: Upload Release Asset (sdist) to GitHub
id: upload-release-asset
uses: actions/upload-release-asset@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
upload_url: ${{ steps.create_release.outputs.upload_url }}
asset_path: dist/${{ env.name }}
asset_name: ${{ env.name }}
asset_content_type: application/zip
+66
View File
@@ -0,0 +1,66 @@
name: Publish Release - Python
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
UV_VERSION: "0.7.20"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Publish package
shell: bash
working-directory: py
run: uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Build and publish llama-parse
working-directory: py/llama_parse/
run: |
uv build
uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }} - LlamaCloud Services PY
artifacts: "py/**/dist/*"
generateReleaseNotes: true
draft: false
prerelease: false
+51
View File
@@ -0,0 +1,51 @@
name: Publish Release - TypeScript
on:
push:
tags:
- "llama-cloud-services@*"
jobs:
build-and-publish:
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services
run: pnpm install --no-frozen-lockfile
- name: Build tarball
run: |
pnpm pack
working-directory: ts/llama_cloud_services
- name: Setup npm authentication
run: echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc
env:
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Release
working-directory: ts/llama_cloud_services
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
run: pnpm publish --access public --no-git-checks
- name: Create release
uses: ncipollo/release-action@v1
with:
artifacts: "ts/llama_cloud_services/llama-cloud-services*.tgz"
name: Release ${{ github.ref }} - LlamaCloud Services TS
bodyFile: "ts/llama_cloud_services/CHANGELOG.md"
token: ${{ secrets.GITHUB_TOKEN }}
+39
View File
@@ -0,0 +1,39 @@
name: Test - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }} && uv python pin ${{ matrix.python-version }}
- name: Run Tests
working-directory: py
run: uv run -- pytest tests/**/test_*.py
- name: Remove virtual environment
working-directory: py
run: rm -rf .venv/
+34
View File
@@ -0,0 +1,34 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run tests
working-directory: ts/llama_cloud_services/
run: pnpm test --run
-40
View File
@@ -1,40 +0,0 @@
name: Unit Testing
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: 0
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install --with dev
- name: Run testing
env:
CI: true
shell: bash
run: poetry run pytest tests
+4
View File
@@ -5,3 +5,7 @@ __pycache__/
.idea
.env*
.ipynb_checkpoints*
*_cache/
node_modules/
.turbo/
dist/
+6 -6
View File
@@ -21,19 +21,19 @@ repos:
hooks:
- id: ruff
args: [--fix, --exit-non-zero-on-fix]
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
- repo: https://github.com/psf/black-pre-commit-mirror
rev: 23.10.1
hooks:
- id: black-jupyter
name: black-src
alias: black
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.0.1
hooks:
- id: mypy
exclude: ^tests/
exclude: ^py/tests/
additional_dependencies:
[
"types-requests",
@@ -63,13 +63,13 @@ repos:
rev: v3.0.3
hooks:
- id: prettier
exclude: poetry.lock
exclude: uv.lock
- repo: https://github.com/codespell-project/codespell
rev: v2.2.6
hooks:
- id: codespell
additional_dependencies: [tomli]
exclude: ^(poetry.lock|examples)
exclude: ^(uv.lock|docs|ts)
args:
[
"--ignore-words-list",
@@ -84,6 +84,6 @@ repos:
rev: v0.23.1
hooks:
- id: toml-sort-fix
exclude: ".*poetry.lock"
exclude: ".*uv.lock"
exclude: .github/ISSUE_TEMPLATE
+33
View File
@@ -0,0 +1,33 @@
# Python
## Installation
This project uses uv. Create a virtual environment, and run `uv sync`
## Versioning (Maintainers only)
Before merging your changes, make sure to bump the versions.
Make a version bump to `pyproject.toml`. If the underlying dependency on the llamacloud platform OpenAPI
sdk needs bumping, make sure to bring that in as well. If updating dependencies, run `uv lock`.
The legacy `llama_parse` package re-exports some of `llama_cloud_services` in the old namespace. The
versions need to be kept consistent to sidecar it with `llama_cloud_services`. Bump it's version in `llama_parse/pyproject.toml`, and also bump it's dependency version of `llama-cloud-services` to match.
**Note**: Don't worry about updating the `llama_parse/poetry.lock` file when bumping versions. The GitHub action will automatically run `poetry lock` for the llama_parse package during the build process (though it doesn't commit the updated lockfile back to the repo).
You can also do this with `./scripts/version-bump.py set 0.x.x` if you have `uv` installed.
Once the change is merged, push a tag `git tag -a v0.x.x -m 0.x.x` and `git push origin 0.x.x`.
This tagging step can be done with `./scripts/version-bump tag`.
# Typescript
## Installation
...
## Versioning
...
+38 -3
View File
@@ -10,7 +10,8 @@ This includes:
- [LlamaParse](./parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](./report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract (beta/invite-only)](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaExtract](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](./index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
## Getting Started
@@ -25,18 +26,52 @@ Then, get your API key from [LlamaCloud](https://cloud.llamaindex.ai/).
Then, you can use the services in your code:
```python
from llama_cloud_services import LlamaParse, LlamaReport, LlamaExtract
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
)
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
)
```
See the quickstart guides for each service for more information:
- [LlamaParse](./parse.md)
- [LlamaReport (beta/invite-only)](./report.md)
- [LlamaExtract (beta/invite-only)](./extract.md)
- [LlamaExtract](./extract.md)
- [LlamaCloud Index](./index.md)
## Switch to EU SaaS 🇪🇺
If you are interested in using LlamaCloud services in the EU, you can adjust your base URL to `https://api.cloud.eu.llamaindex.ai`.
You can also create your API key in the EU region [here](https://cloud.eu.llamaindex.ai).
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
EU_BASE_URL,
)
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
project_name="default",
api_key="YOUR_API_KEY",
base_url=EU_BASE_URL,
)
```
## Documentation
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 MiB

File diff suppressed because one or more lines are too long
@@ -0,0 +1,10 @@
# Financial Modeling Assumptions
Discount Rate: 8%
Terminal Growth Rate: 2%
Tax Rate: 25%
Revenue Growth (Years 1-5): 10% per annum
Revenue Growth (Years 6-10): 5% per annum
Capital Expenditures as % of Revenue: 7%
Working Capital Assumption: 3% of Revenue
Depreciation Rate: 10% per annum
Cost of Capital Assumption: 8%
Binary file not shown.

After

Width:  |  Height:  |  Size: 67 KiB

@@ -0,0 +1 @@
sec_form_4_dump.json
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 202 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 440 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 156 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 85 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 893 KiB

@@ -0,0 +1,440 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Extract Data from Financial Reports - with Citations and Reasoning\n",
"\n",
"Given complex files like financial reports, contracts, invoices etc, Llama Extract allows you to make use of an LLM to extract the information relevant to you, in a structured format.\n",
"\n",
"In this example, we'll be using [LlamaExtract](https://docs.cloud.llamaindex.ai/llamaextract/getting_started?utm_campaign=extract&utm_medium=recipe) to extract structured data from an SEC filing (specifically, the filing by Nvidia for fiscal year 2025).\n",
"\n",
"On top of simple data extraction, we'll ask our extraction agent to provide citations and reasoning for each extracted field. This allows us to:\n",
"- Confirm the accuracy of the extracted field\n",
"- Understand the reasoning behind why the LLM extracted a given piece of information\n",
"- This last point allows us an opportunity to adjust the system prompt or field descriptions and improve on results where needed.\n",
"\n",
"\n",
"The example we go through below is also replicable within Llama Cloud as well, where you will also be able to pick between a number of pre-defined schemas, instead of building your own."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-cloud-services"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Connect to Llama Cloud\n",
"\n",
"To get started, make sure you provide your [Llama Cloud](https://cloud.llamaindex.ai?utm_campaign=extract&utm_medium=recipe) API key."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Enter your Llama Cloud API Key: ··········\n"
]
}
],
"source": [
"import os\n",
"from getpass import getpass\n",
"\n",
"if \"LLAMA_CLOUD_API_KEY\" not in os.environ:\n",
" os.environ[\"LLAMA_CLOUD_API_KEY\"] = getpass(\"Enter your Llama Cloud API Key: \")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Extract Data with Llama Extract Agent"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"No project_id provided, fetching default project.\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaExtract\n",
"\n",
"# Optionally, provide your project id, if not, it will use the 'Default' project\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Provide Your Custom Schema\n",
"\n",
"When using LlamaExtract via the API, you provide your own schema that describes what you want extracted from files and data provided to your agent. Here, we are essentially building an SEC filings extraction agent."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from enum import Enum\n",
"\n",
"\n",
"class FilingType(str, Enum):\n",
" ten_k = \"10 K\"\n",
" ten_q = \"10-Q\"\n",
" ten_ka = \"10-K/A\"\n",
" ten_qa = \"10-Q/A\"\n",
"\n",
"\n",
"class FinancialReport(BaseModel):\n",
" company_name: str = Field(description=\"The name of the company\")\n",
" description: str = Field(\n",
" description=\"Short description of the filing and what it contains\"\n",
" )\n",
" filing_type: FilingType = Field(description=\"Type of SEC filing\")\n",
" filing_date: str = Field(description=\"Date when filing was submitted to SEC\")\n",
" fiscal_year: int = Field(description=\"Fiscal year\")\n",
" unit: str = Field(\n",
" description=\"Unit of financial figures (thousands, millions, etc.)\"\n",
" )\n",
" revenue: int = Field(description=\"Total revenue for period\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Set Up Citations and Reasoning\n",
"\n",
"Optionally, we can set the `ExtractConfig` to extract citations for each field the agent extracts. These cications will cite the specific pages and sections of the file from which a given field was extractedd.\n",
"\n",
"By setting `use_reasoning` to True, we als ask the agent to do an additional reasoning step, explaining why a given field was extracted."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud.types import ExtractConfig, ExtractMode\n",
"\n",
"config = ExtractConfig(\n",
" use_reasoning=True, cite_sources=True, extraction_mode=ExtractMode.MULTIMODAL\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:127: ExperimentalWarning: `use_reasoning` is an experimental feature. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n",
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:133: ExperimentalWarning: `cite_sources` is an experimental feature. This may greatly increase the size of the response, and slow down the extraction. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n"
]
}
],
"source": [
"agent = llama_extract.create_agent(\n",
" name=\"filing-parser\", data_schema=FinancialReport, config=config\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Demo Time - Download a PDF and Extract Data with Citations"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"PDF downloaded successfully.\n"
]
}
],
"source": [
"import requests\n",
"\n",
"url = \"https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf\"\n",
"\n",
"response = requests.get(url)\n",
"\n",
"if response.status_code == 200:\n",
" with open(\"/content/nvda_10k.pdf\", \"wb\") as f:\n",
" f.write(response.content)\n",
" print(\"PDF downloaded successfully.\")\n",
"else:\n",
" print(f\"Failed to download. Status code: {response.status_code}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|██████████| 1/1 [00:00<00:00, 1.83it/s]\n",
"Creating extraction jobs: 100%|██████████| 1/1 [00:00<00:00, 4.38it/s]\n",
"Extracting files: 100%|██████████| 1/1 [02:03<00:00, 123.40s/it]\n"
]
}
],
"source": [
"filing_info = agent.extract(\"/content/nvda_10k.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'company_name': 'NVIDIA Corporation',\n",
" 'description': \"The filing provides a detailed overview of NVIDIA's business as a full-stack computing infrastructure company, discusses various technologies including digital avatars and autonomous vehicles, outlines numerous risk factors affecting operations such as supply chain issues and geopolitical tensions, and describes employee stock purchase plans and related compliance requirements.\",\n",
" 'filing_type': '10 K',\n",
" 'filing_date': 'February 26, 2025',\n",
" 'fiscal_year': 2025,\n",
" 'unit': 'millions',\n",
" 'revenue': 130497}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Inspect Citations and Reasoning"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'field_metadata': {'company_name': {'reasoning': 'VERBATIM EXTRACTION',\n",
" 'citation': [{'page': 1, 'matching_text': 'NVIDIA CORPORATION'},\n",
" {'page': 2, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 3,\n",
" 'matching_text': 'All references to \"NVIDIA,\" \"we,\" \"us,\" \"our,\" or the \"Company\" mean NVIDIA Corporation and its subsidiaries.'},\n",
" {'page': 35,\n",
" 'matching_text': 'Comparison of 5 Year Cumulative Total Return* Among NVIDIA Corporation'},\n",
" {'page': 49,\n",
" 'matching_text': 'To the Board of Directors and Shareholders of NVIDIA Corporation'},\n",
" {'page': 90, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 119,\n",
" 'matching_text': '*\"Company\"* means NVIDIA Corporation, a Delaware corporation.'},\n",
" {'page': 126,\n",
" 'matching_text': 'Annual Report on Form 10-K of NVIDIA Corporation'}]},\n",
" 'filing_type': {'reasoning': \"VERBATIM EXTRACTION from multiple sources confirming the filing type as '10 K'.\",\n",
" 'citation': [{'page': 1, 'matching_text': 'FORM 10-K'},\n",
" {'page': 2, 'matching_text': 'Item 16. | Form 10-K Summary'},\n",
" {'page': 3,\n",
" 'matching_text': 'This Annual Report on Form 10-K contains forward-looking statements...'},\n",
" {'page': 13, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 15, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 32,\n",
" 'matching_text': 'Annual Report on Form 10-K, which information is hereby incorporated by reference.'},\n",
" {'page': 36, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 43,\n",
" 'matching_text': 'Annual Report on Form 10-K for additional information'},\n",
" {'page': 45, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 46, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 62, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 83,\n",
" 'matching_text': 'Restated Certificate of Incorporation | 10-K'},\n",
" {'page': 84, 'matching_text': 'Item 16. Form 10-K Summary'},\n",
" {'page': 126, 'matching_text': 'which appears in this Form 10-K'},\n",
" {'page': 127, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 128, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 129, 'matching_text': \"The Company's Annual Report on Form 10-K\"},\n",
" {'page': 130,\n",
" 'matching_text': \"The Company's Annual Report on Form 10-K for the year ended January 26, 2025\"}]},\n",
" 'fiscal_year': {'reasoning': 'The fiscal year ended January 26, 2025, indicates the fiscal year is 2025. Additionally, multiple references throughout the text confirm the fiscal year 2025 in various contexts.',\n",
" 'citation': [{'page': 1,\n",
" 'matching_text': 'For the fiscal year ended January 26, 2025'},\n",
" {'page': 6,\n",
" 'matching_text': 'In fiscal year 2025, we launched the NVIDIA Blackwell architecture'},\n",
" {'page': 12, 'matching_text': 'fiscal year 2025'},\n",
" {'page': 17,\n",
" 'matching_text': 'our gross margins in the second quarter of fiscal year 2025 were negatively impacted'},\n",
" {'page': 20,\n",
" 'matching_text': 'we generated 53% of our revenue in fiscal year 2025 from sales outside the United States.'},\n",
" {'page': 23,\n",
" 'matching_text': 'For fiscal year 2025, an indirect customer which primarily purchases our products through system integrators...'},\n",
" {'page': 33,\n",
" 'matching_text': 'In fiscal year 2025, we repurchased 310 million shares of our common stock for $34.0 billion.'},\n",
" {'page': 37,\n",
" 'matching_text': 'Our Data Center revenue in China grew in fiscal year 2025.'},\n",
" {'page': 44,\n",
" 'matching_text': 'Cash provided by operating activities increased in fiscal year 2025 compared to fiscal year 2024'},\n",
" {'page': 57,\n",
" 'matching_text': 'Fiscal years 2025, 2024 and 2023 were all 52-week years.'},\n",
" {'page': 65,\n",
" 'matching_text': 'Beginning in the second quarter of fiscal year 2025'},\n",
" {'page': 69, 'matching_text': 'In the fourth quarter of fiscal year 2025'},\n",
" {'page': 78,\n",
" 'matching_text': 'Depreciation and amortization expense attributable to our Compute and Networking segment for fiscal years 2025'},\n",
" {'page': 129, 'matching_text': 'for the year ended January 26, 2025'}]},\n",
" 'description': {'reasoning': 'The extracted data combines multiple descriptions from the source text, ensuring no duplication while maintaining the order and context of the information. Each section of the filing is summarized to reflect the key points without losing the essence of the original text.',\n",
" 'citation': [{'page': 4,\n",
" 'matching_text': 'NVIDIA is now a full-stack computing infrastructure company with data-center-scale offerings that are reshaping industry.'},\n",
" {'page': 8,\n",
" 'matching_text': 'a suite of technologies that help developers bring digital avatars to life with generative Al...autonomous vehicles, or AV, and electric vehicles, or EV, is revolutionizing the transportation industry...Our worldwide sales and marketing strategy is key to achieving our objective of providing markets with our high-performance and efficient computing platforms and software.'},\n",
" {'page': 14, 'matching_text': 'Risk Factors Summary'},\n",
" {'page': 16,\n",
" 'matching_text': 'Risks Related to Demand, Supply, and Manufacturing\\n\\nLong manufacturing lead times and uncertain supply and component availability...'},\n",
" {'page': 18,\n",
" 'matching_text': 'cryptocurrency mining, on demand for our products. Volatility in the cryptocurrency market, including new compute technologies...'},\n",
" {'page': 21,\n",
" 'matching_text': 'supply-chain attacks or other business disruptions. We cannot guarantee that third parties and infrastructure in our supply chain...'},\n",
" {'page': 22,\n",
" 'matching_text': 'We are monitoring the impact of the geopolitical conflict in and around Israel on our operations... Climate change may have a long-term impact on our business.'},\n",
" {'page': 25,\n",
" 'matching_text': 'We are subject to complex laws, rules, regulations, and political and other actions, including restrictions on the export of our products, which may adversely impact our business.'},\n",
" {'page': 28,\n",
" 'matching_text': 'Our competitive position has been harmed by the existing export controls, and our competitive position and future results may be further harmed'},\n",
" {'page': 29,\n",
" 'matching_text': 'restrictions imposed by the Chinese government on the duration of gaming activities and access to games may adversely affect our Gaming revenue'},\n",
" {'page': 29,\n",
" 'matching_text': 'our business depends on our ability to receive consistent and reliable supply from our overseas partners, especially in Taiwan and South Korea'},\n",
" {'page': 29,\n",
" 'matching_text': 'Increased scrutiny from shareholders, regulators and others regarding our corporate sustainability practices could result in additional costs'},\n",
" {'page': 29,\n",
" 'matching_text': 'Concerns relating to the responsible use of new and evolving technologies, such as Al, in our products and services may result in reputational or financial harm'},\n",
" {'page': 31,\n",
" 'matching_text': 'Data protection laws around the world are quickly changing and may be interpreted and applied in an increasingly stringent fashion...'}]},\n",
" 'filing_date': {'reasoning': 'The filing date is consistently mentioned as February 26, 2025 across multiple entries, making it the most reliable date for the filing.',\n",
" 'citation': [{'page': 51, 'matching_text': 'February 26, 2025'},\n",
" {'page': 86, 'matching_text': 'on February 26, 2025.'},\n",
" {'page': 87, 'matching_text': 'February 26, 2025'},\n",
" {'page': 126, 'matching_text': 'our report dated February 26, 2025'},\n",
" {'page': 127, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 128, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 129, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 130, 'matching_text': 'Date: February 26, 2025'}]},\n",
" 'unit': {'reasoning': \"The unit of financial figures is explicitly mentioned multiple times in the text as 'millions', including in table headers and notes. This is confirmed by various citations from pages 38, 42, 43, 52, 53, 54, 56, 65, 71, 72, 73, 75, 77, 79, 80, and 82.\",\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': '($ in millions, except per share data)'},\n",
" {'page': 42, 'matching_text': '($ in millions)'},\n",
" {'page': 43, 'matching_text': '($ in millions)'},\n",
" {'page': 52, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 53,\n",
" 'matching_text': 'Consolidated Statements of Comprehensive Income (In millions)'},\n",
" {'page': 54,\n",
" 'matching_text': 'Consolidated Balance Sheets (In millions, except par value)'},\n",
" {'page': 55, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 56,\n",
" 'matching_text': 'Consolidated Statements of Cash Flows (In millions)'},\n",
" {'page': 65,\n",
" 'matching_text': 'Year Ended<br/>Jan 26, 2025<br/>(In millions, except per share data)'},\n",
" {'page': 71, 'matching_text': '(In millions) | (In millions)'},\n",
" {'page': 72, 'matching_text': '(In millions)'}]},\n",
" 'revenue': {'reasoning': 'The total revenue for fiscal year 2025 is extracted from multiple sources within the text, all confirming the same figure of $130,497 million. The revenue recognized for fiscal year 2025 is also noted as $4,607 million, which is a separate figure. However, the primary focus is on the total revenue figure, which is consistently cited.',\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': 'Revenue for fiscal year 2025 was $130.5 billion'},\n",
" {'page': 41,\n",
" 'matching_text': 'Total | $ 130,497 | $ | 60,922'},\n",
" {'page': 52, 'matching_text': 'Revenue | $ 130,497'},\n",
" {'page': 78,\n",
" 'matching_text': 'Revenue | $ 116,193 | $ 14,304 | $ - | $ 130,497'},\n",
" {'page': 79, 'matching_text': 'Total revenue | $ 130,497'},\n",
" {'page': 80, 'matching_text': 'Total revenue | $ 130,497'}]}},\n",
" 'usage': {'num_pages_extracted': 130,\n",
" 'num_document_tokens': 105932,\n",
" 'num_output_tokens': 31306}}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.extraction_metadata"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## What's Next?\n",
"\n",
"In this example, we built an Extraction Agent that is capable of citing it's sources from the document it's extracting data from, and reasoning about its reponse. To further customize and improve on the results, you can also try to customize the `system_prompt` in the `ExtractConfig`.\n",
"\n",
"#### Learn More\n",
"\n",
"- [LlamaExtract Documentation](https://docs.cloud.llamaindex.ai/llamaextract/getting_started)\n",
"- [Example Notebooks](https://github.com/run-llama/llama_cloud_services/tree/main/examples/extract)"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,318 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "1f6bd03d-1b8b-45a0-bc2c-5a13f1a5d8d3",
"metadata": {},
"source": [
"# LM317 Voltage Regulator Datasheet Structured Extraction\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/lm317_structured_extraction.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates an agentic document workflow using LlamaExtract to process an LM317 voltage regulator datasheet. In this example, we define a structured extraction schema that converts key technical fields into standardized subfields. For instance, the output voltage is split into a minimum and maximum value with a defined unit, and we capture page citations for each extracted field.\n",
"\n",
"The target user is an electronics engineer at a component manufacturing company who needs to consolidate datasheet information into a standardized specification sheet for design and quality control.\n",
"\n",
"This approach reduces manual data entry, improves extraction accuracy and standardization, and provides traceability for each technical detail."
]
},
{
"cell_type": "markdown",
"id": "a3b8c8d5-ff3e-48ce-b0b8-29b6b1f517f8",
"metadata": {},
"source": [
"## Use Case Overview\n",
"\n",
"### Problem\n",
"Datasheets like that for the LM317 regulator are often distributed as PDFs containing multiple tables, charts, and complex textual descriptions. Engineers must manually extract technical details such as voltage ranges, dropout voltage, maximum current, input voltage range, and pin configurations. This process is error-prone and time-consuming.\n",
"\n",
"### Agent Workflow (Combination of Automation and Chat)\n",
"1. **Upload Datasheet:** The engineer uploads the LM317 datasheet PDF. \n",
"2. **Structured Extraction:** An automated agent processes the PDF and extracts key technical details into structured fields (e.g., output voltage as a range with separate min/max values).\n",
"3. **Interactive Verification:** The engineer can query the agent (via chat) for further details or clarification (e.g., \"Show me the detailed pin configuration extraction\") and review the cited pages.\n",
"\n",
"**Value Delivered:**\n",
"- Up to 70% reduction in manual data extraction time.\n",
"- Increased accuracy and standardization with structured fields."
]
},
{
"cell_type": "markdown",
"id": "a704e843-54be-4969-842b-713584cb3c35",
"metadata": {},
"source": [
"## Setup and Download Data\n",
"\n",
"Download the [LM317 Datasheet](https://www.ti.com/lit/ds/symlink/lm317.pdf) and setup LlamaExtract."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6e5b1f91-8785-44d4-a710-8be1b48b76de",
"metadata": {},
"outputs": [],
"source": [
"!mkdir -p data/lm317_structured_extraction\n",
"!wget https://www.ti.com/lit/ds/symlink/lm317.pdf -O data/lm317_structured_extraction/lm317.pdf"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f17b914a-00ed-4b63-8198-69fd7c4a7c62",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"\n",
"# Load environment variables (ensure LLAMA_CLOUD_API_KEY is set in your .env file)\n",
"load_dotenv(override=True)\n",
"\n",
"# Initialize the LlamaExtract client\n",
"llama_extract = LlamaExtract(\n",
" project_id=\"<project_id>\",\n",
" organization_id=\"<organization_id>\",\n",
")"
]
},
{
"cell_type": "markdown",
"id": "ed9f6e9a-96c8-4ee1-8b45-0b6a4f7dbbf1",
"metadata": {},
"source": [
"## Defining a Structured Extraction Schema\n",
"\n",
"We now define a rich Pydantic schema to extract technical specifications from the LM317 datasheet. In this schema:\n",
"\n",
"- The **output_voltage** and **input_voltage** fields are structured as ranges with separate minimum and maximum values and a unit.\n",
"- The **pin_configuration** field is structured to include a pin count and a descriptive layout.\n",
"- Additional technical fields (e.g., dropout voltage, max current) are captured as numbers.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "4f7e9b44-5e69-4b30-9864-cd98f1e2a7d4",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"class VoltageRange(BaseModel):\n",
" min_voltage: float = Field(..., description=\"Minimum voltage in volts\")\n",
" max_voltage: float = Field(..., description=\"Maximum voltage in volts\")\n",
" unit: str = Field(\"V\", description=\"Voltage unit\")\n",
"\n",
"\n",
"class PinConfiguration(BaseModel):\n",
" pin_count: int = Field(..., description=\"Number of pins\")\n",
" layout: str = Field(..., description=\"Detailed pin layout description\")\n",
"\n",
"\n",
"class LM317Spec(BaseModel):\n",
" component_name: str = Field(..., description=\"Name of the component\")\n",
" output_voltage: VoltageRange = Field(\n",
" ..., description=\"Output voltage range specification\"\n",
" )\n",
" dropout_voltage: float = Field(..., description=\"Dropout voltage in volts\")\n",
" max_current: float = Field(..., description=\"Maximum current rating in amperes\")\n",
" input_voltage: VoltageRange = Field(\n",
" ..., description=\"Input voltage range specification\"\n",
" )\n",
" pin_configuration: PinConfiguration = Field(\n",
" ..., description=\"Pin configuration details\"\n",
" )\n",
" features: List[str] = Field([], description=\"List of additional technical features\")\n",
"\n",
"\n",
"class LM317Schema(BaseModel):\n",
" specs: List[LM317Spec] = Field(\n",
" ..., description=\"List of extracted LM317 technical specifications\"\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e0508e38-35be-446c-afe7-129e39553281",
"metadata": {},
"outputs": [],
"source": [
"try:\n",
" existing_agent = llama_extract.get_agent(name=\"lm317-datasheet\")\n",
" if existing_agent:\n",
" llama_extract.delete_agent(existing_agent.id)\n",
"except ApiError as e:\n",
" if e.status_code == 404:\n",
" pass\n",
" else:\n",
" raise"
]
},
{
"cell_type": "markdown",
"id": "bb197dfd-dd37-459e-8953-cc1b12f25bdd",
"metadata": {},
"source": [
"Here we use our balanced extraction mode."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e3defc0a-c685-4fbd-bbb1-1270f1442e72",
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud import ExtractConfig\n",
"\n",
"extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
")\n",
"\n",
"agent = llama_extract.create_agent(\n",
" name=\"lm317-datasheet\", data_schema=LM317Schema, config=extract_config\n",
")"
]
},
{
"cell_type": "markdown",
"id": "c0a0f9f9-2ef3-4a38-bd74-68d2c2e9e2d8",
"metadata": {},
"source": [
"## Extracting Information from the LM317 Datasheet\n",
"\n",
"For this demonstration, please download a publicly available LM317 voltage regulator datasheet (for example, from Texas Instruments) and save it as `lm317.pdf` in the `./data` directory. Then run the cell below to extract the structured technical specifications."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c58e8b7a-8f9b-46f3-8f72-3c2f96b49e8f",
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:01<00:00, 1.08s/it]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.96it/s]\n",
"Extracting files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [01:27<00:00, 87.38s/it]\n"
]
}
],
"source": [
"# Path to the LM317 datasheet PDF\n",
"lm317_pdf = \"./data/lm317_structured_extraction/lm317.pdf\"\n",
"\n",
"# Extract structured technical specifications from the datasheet\n",
"lm317_extract = agent.extract(lm317_pdf)"
]
},
{
"cell_type": "markdown",
"id": "1a2e2e44-6c48-4a38-a6de-5f2f3c7d4d8b",
"metadata": {},
"source": [
"## Assessing the Extraction Results\n",
"\n",
"The output will be a consolidated list of LM317 technical specifications. For each entry, you should see structured fields including:\n",
"\n",
"- **component_name**\n",
"- **output_voltage** as a range (with separate `min_voltage` and `max_voltage` plus `unit`)\n",
"- **dropout_voltage** and **max_current** as numbers\n",
"- **input_voltage** as a structured range\n",
"- **pin_configuration** with a `pin_count` and `layout`\n",
"- **features** (if available)\n",
"\n",
"This structured approach makes it easier to standardize the information for downstream integration and verification. Engineers can click on the cited page numbers (in a UI that supports it) to validate the extraction."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "fb2abc44-7c9b-4b19-958e-d0d7b390ae57",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'specs': [{'component_name': 'LM317',\n",
" 'output_voltage': {'min_voltage': 1.25, 'max_voltage': 37.0, 'unit': 'V'},\n",
" 'dropout_voltage': 0.0,\n",
" 'max_current': 1.5,\n",
" 'input_voltage': {'min_voltage': 4.25, 'max_voltage': 40.0, 'unit': 'V'},\n",
" 'pin_configuration': {'pin_count': 3,\n",
" 'layout': '1: ADJUST, 2: OUTPUT, 3: INPUT'},\n",
" 'features': ['Output voltage range adjustable from 1.25 V to 37 V',\n",
" 'Output current greater than 1.5 A',\n",
" 'Internal short-circuit current limiting',\n",
" 'Thermal overload protection',\n",
" 'Output safe-area compensation']}]}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Display the extraction results\n",
"lm317_extract.data"
]
},
{
"cell_type": "markdown",
"id": "c7a2a523-095e-40bf-b713-f509c13a7747",
"metadata": {},
"source": [
"You can also see the output result in the UI."
]
},
{
"cell_type": "markdown",
"id": "dc22dfa5-b667-4fb0-8dbe-24e401b12389",
"metadata": {},
"source": [
"![](data/lm317_structured_extraction/lm317_extraction.png)"
]
},
{
"cell_type": "markdown",
"id": "e0e0c12a-9f89-4bb3-b40d-3e9f7c6d2fef",
"metadata": {},
"source": [
"## Conclusion\n",
"\n",
"This notebook demonstrated how to use LlamaExtract with a structured extraction schema for the LM317 voltage regulator datasheet. By defining detailed subfields (such as splitting voltage ranges into minimum and maximum values, and structuring the pin configuration), we ensure that the extracted data is standardized and traceable through page citations. This approach minimizes manual effort and improves accuracy, providing a robust example of an agentic document workflow for technical documentation processing.\n",
"\n",
"Feel free to modify or extend the schema to capture additional technical details or to suit your own use cases."
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,450 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "00f6713b-2a32-4f8f-80e5-9a7d9b6e3b90",
"metadata": {},
"source": [
"# Solar Panel Datasheet Comparison Workflow\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/solar_panel_e2e_comparison.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"\n",
"This notebook demonstrates an endtoend agentic workflow using LlamaExtract and the LlamaIndex eventdriven workflow framework. In this workflow, we:\n",
"\n",
"1. **Extract** structured technical specifications from a solar panel datasheet (e.g. a PDF downloaded from a vendor).\n",
"2. **Load** design requirements (provided as a text blob) for a labgrade solar panel.\n",
"3. **Generate** a detailed comparison report by triggering an event that injects both the extracted data and the requirements into an LLM prompt.\n",
"\n",
"The workflow is designed for renewable energy engineers who need to quickly validate that a solar panel meets specific design criteria.\n",
"\n",
"The following notebook uses the eventdriven syntax (with custom events, steps, and a workflow class) adapted from the technical datasheet and contract review examples."
]
},
{
"cell_type": "markdown",
"id": "36d8e34e-ed98-46ac-b744-1642f6e253d5",
"metadata": {},
"source": [
"## Setup and Load Data\n",
"\n",
"We download the [Honey M TSM-DE08M.08(II) datasheet](https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf) as a PDF.\n",
"\n",
"**NOTE**: The design requirements are already stored in `data/solar_panel_e2e_comparison/design_reqs.txt`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1de7b1b3-c285-492c-8b2e-b37974b4fc63",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2025-04-01 14:47:56-- https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf\n",
"Resolving static.trinasolar.com (static.trinasolar.com)... 47.246.23.232, 47.246.23.234, 47.246.23.227, ...\n",
"Connecting to static.trinasolar.com (static.trinasolar.com)|47.246.23.232|:443... connected.\n",
"WARNING: cannot verify static.trinasolar.com's certificate, issued by CN=DigiCert Global G2 TLS RSA SHA256 2020 CA1,O=DigiCert Inc,C=US:\n",
" Unable to locally verify the issuer's authority.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 1888183 (1.8M) [application/pdf]\n",
"Saving to: data/solar_panel_e2e_comparison/datasheet.pdf\n",
"\n",
"data/solar_panel_e2 100%[===================>] 1.80M 7.47MB/s in 0.2s \n",
"\n",
"2025-04-01 14:47:56 (7.47 MB/s) - data/solar_panel_e2e_comparison/datasheet.pdf saved [1888183/1888183]\n",
"\n"
]
}
],
"source": [
"!wget https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf -O data/solar_panel_e2e_comparison/datasheet.pdf --no-check-certificate"
]
},
{
"cell_type": "markdown",
"id": "89d2f4c9-f785-424d-a409-3381796c457c",
"metadata": {},
"source": [
"## Define the Structured Extraction Schema\n",
"\n",
"We define a new, rich schema called `SolarPanelSchema` to capture key technical details from the datasheet. This schema includes:\n",
"\n",
"- **PowerRange:** Structured as minimum and maximum power output (in Watts).\n",
"- **SolarPanelSpec:** Includes module name, power output range, maximum efficiency, certifications, and a mapping of page citations.\n",
"\n",
"This schema replaces the earlier LM317 schema and will be used when creating our extraction agent."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "bfb40d48-36e0-4b1c-97a1-32a1704c582b",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"class PowerRange(BaseModel):\n",
" min_power: float = Field(..., description=\"Minimum power output in Watts\")\n",
" max_power: float = Field(..., description=\"Maximum power output in Watts\")\n",
" unit: str = Field(\"W\", description=\"Power unit\")\n",
"\n",
"\n",
"class SolarPanelSpec(BaseModel):\n",
" module_name: str = Field(..., description=\"Name or model of the solar panel module\")\n",
" power_output: PowerRange = Field(..., description=\"Power output range\")\n",
" maximum_efficiency: float = Field(\n",
" ..., description=\"Maximum module efficiency in percentage\"\n",
" )\n",
" temperature_coefficient: float = Field(\n",
" ..., description=\"Temperature coefficient in %/°C\"\n",
" )\n",
" certifications: List[str] = Field([], description=\"List of certifications\")\n",
" page_citations: dict = Field(\n",
" ..., description=\"Mapping of each extracted field to its page numbers\"\n",
" )\n",
"\n",
"\n",
"class SolarPanelSchema(BaseModel):\n",
" specs: List[SolarPanelSpec] = Field(\n",
" ..., description=\"List of extracted solar panel specifications\"\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "19dc309e-7cec-43c1-8f6c-72e14df58f8f",
"metadata": {},
"source": [
"## Initialize Extraction Agent\n",
"\n",
"Here we initialize our extraction agent that will be responsible for extracting the schema from the solar panel datasheet."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c9d9f4a2-2e14-493d-8a7e-d01159d38b8f",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"from llama_cloud import ExtractConfig\n",
"\n",
"# Initialize the LlamaExtract client\n",
"llama_extract = LlamaExtract(\n",
" project_id=\"2fef999e-1073-40e6-aeb3-1f3c0e64d99b\",\n",
" organization_id=\"43b88c8f-e488-46f6-9013-698e3d2e374a\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "ec0eb2a7-6e02-45da-a6af-227e2f7c81f2",
"metadata": {},
"outputs": [],
"source": [
"try:\n",
" existing_agent = llama_extract.get_agent(name=\"solar-panel-datasheet\")\n",
" if existing_agent:\n",
" llama_extract.delete_agent(existing_agent.id)\n",
"except ApiError as e:\n",
" if e.status_code == 404:\n",
" pass\n",
" else:\n",
" raise\n",
"\n",
"extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
")\n",
"\n",
"agent = llama_extract.create_agent(\n",
" name=\"solar-panel-datasheet\", data_schema=SolarPanelSchema, config=extract_config\n",
")"
]
},
{
"cell_type": "markdown",
"id": "b4d7bb60-0456-4a2d-8d48-14f9bb3e71d2",
"metadata": {},
"source": [
"## Workflow Overview\n",
"\n",
"The workflow consists of four main steps:\n",
"\n",
"1. **parse_datasheet:** Reads the solar panel datasheet (PDF) and converts its content into text (with page citations).\n",
"2. **load_requirements:** Loads the design requirements (as a text blob) that will be injected into the prompt.\n",
"3. **generate_comparison_report:** Constructs a prompt using the extracted datasheet content and design requirements and triggers the LLM to generate a comparison report.\n",
"4. **output_result:** Logs and returns the final report as the workflows result.\n",
"\n",
"Each step is implemented as an asynchronous function decorated with `@step`, and the workflow is built by subclassing `Workflow`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7c482e3a-66b4-4e1b-8d2d-9a9c6b3967f3",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.workflow import (\n",
" Event,\n",
" StartEvent,\n",
" StopEvent,\n",
" Context,\n",
" Workflow,\n",
" step,\n",
")\n",
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.core.prompts import ChatPromptTemplate\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"# Define output schema for the comparison report (for reference)\n",
"class ComparisonReportOutput(BaseModel):\n",
" component_name: str = Field(\n",
" ..., description=\"The name of the component being evaluated.\"\n",
" )\n",
" meets_requirements: bool = Field(\n",
" ...,\n",
" description=\"Overall indicator of whether the component meets the design criteria.\",\n",
" )\n",
" summary: str = Field(..., description=\"A brief summary of the evaluation results.\")\n",
" details: dict = Field(\n",
" ..., description=\"Detailed comparisons for each key parameter.\"\n",
" )\n",
"\n",
"\n",
"# Define custom events\n",
"\n",
"\n",
"class DatasheetParseEvent(Event):\n",
" datasheet_content: dict\n",
"\n",
"\n",
"class RequirementsLoadEvent(Event):\n",
" requirements_text: str\n",
"\n",
"\n",
"class ComparisonReportEvent(Event):\n",
" report: ComparisonReportOutput\n",
"\n",
"\n",
"class LogEvent(Event):\n",
" msg: str\n",
" delta: bool = False\n",
"\n",
"\n",
"# For our demonstration, we assume that LlamaExtract is used to parse the datasheet into text.\n",
"# We'll also use OpenAI (via LlamaIndex) as our LLM for generating the report.\n",
"\n",
"llm = OpenAI(model=\"gpt-4o\") # or your preferred model"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "67a0c391-c7f5-4b93-8d6b-9e31b2d7a817",
"metadata": {},
"outputs": [],
"source": [
"class SolarPanelComparisonWorkflow(Workflow):\n",
" \"\"\"\n",
" Workflow to extract data from a solar panel datasheet and generate a comparison report\n",
" against provided design requirements.\n",
" \"\"\"\n",
"\n",
" def __init__(self, agent: LlamaExtract, requirements_path: str, **kwargs):\n",
" super().__init__(**kwargs)\n",
" self.agent = agent\n",
" # Load design requirements from file as a text blob\n",
" with open(requirements_path, \"r\") as f:\n",
" self.requirements_text = f.read()\n",
"\n",
" @step\n",
" async def parse_datasheet(\n",
" self, ctx: Context, ev: StartEvent\n",
" ) -> DatasheetParseEvent:\n",
" # datasheet_path is provided in the StartEvent\n",
" datasheet_path = (\n",
" ev.datasheet_path\n",
" ) # e.g., \"./data/solar_panel_comparison/datasheet.pdf\"\n",
" extraction_result = await self.agent.aextract(datasheet_path)\n",
" datasheet_dict = (\n",
" extraction_result.data\n",
" ) # assumed to be a string with page citations\n",
" await ctx.set(\"datasheet_content\", datasheet_dict)\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Datasheet parsed successfully.\"))\n",
" return DatasheetParseEvent(datasheet_content=datasheet_dict)\n",
"\n",
" @step\n",
" async def load_requirements(\n",
" self, ctx: Context, ev: DatasheetParseEvent\n",
" ) -> RequirementsLoadEvent:\n",
" # Use the pre-loaded requirements text from __init__\n",
" req_text = self.requirements_text\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Design requirements loaded.\"))\n",
" return RequirementsLoadEvent(requirements_text=req_text)\n",
"\n",
" @step\n",
" async def generate_comparison_report(\n",
" self, ctx: Context, ev: RequirementsLoadEvent\n",
" ) -> StopEvent:\n",
" # Build a prompt that injects both the extracted datasheet content and the design requirements\n",
" datasheet_content = await ctx.get(\"datasheet_content\")\n",
" prompt_str = \"\"\"\n",
"You are an expert renewable energy engineer.\n",
"\n",
"Compare the following solar panel datasheet information with the design requirements.\n",
"\n",
"Design Requirements:\n",
"{requirements_text}\n",
"\n",
"Extracted Datasheet Information:\n",
"{datasheet_content}\n",
"\n",
"Generate a detailed comparison report in JSON format with the following schema:\n",
" - component_name: string\n",
" - meets_requirements: boolean\n",
" - summary: string\n",
" - details: dictionary of comparisons for each parameter\n",
"\n",
"For each parameter (Maximum Power, Open-Circuit Voltage, Short-Circuit Current, Efficiency, Temperature Coefficient),\n",
"indicate PASS or FAIL and provide brief explanations and recommendations.\n",
"\"\"\"\n",
"\n",
" # extract from contract\n",
" prompt = ChatPromptTemplate.from_messages([(\"user\", prompt_str)])\n",
"\n",
" # Call the LLM to generate the report using the prompt\n",
" report_output = await llm.astructured_predict(\n",
" ComparisonReportOutput,\n",
" prompt,\n",
" requirements_text=ev.requirements_text,\n",
" datasheet_content=str(datasheet_content),\n",
" )\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Comparison report generated.\"))\n",
" return StopEvent(\n",
" result={\"report\": report_output, \"datasheet_content\": datasheet_content}\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "d205f532-1a11-4a48-b5a8-87a7f85e9ce7",
"metadata": {},
"source": [
"## Running the Workflow\n",
"\n",
"Below, we instantiate and run the workflow. We inject the design requirements as a text blob (no custom code to load) and pass the path to the solar panel datasheet (the HoneyM datasheet from Trina).\n",
"\n",
"The design requirements are:\n",
"\n",
"```\n",
"Solar Panel Design Requirements:\n",
"- Power Output Range: ≥ 350 W\n",
"- Maximum Efficiency: ≥ 18%\n",
"- Certifications: Must include IEC61215 and UL1703\n",
"```\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6b24fa61-a2f5-4ebb-84eb-1c9b48683b1b",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "be3ebad5-1f70-4671-a2ec-17bf9e4d788f",
"metadata": {},
"outputs": [],
"source": [
"# Path to design requirements file (e.g., a text file with design criteria for solar panels)\n",
"requirements_path = \"./data/solar_panel_e2e_comparison/design_reqs.txt\"\n",
"\n",
"# Instantiate the workflow\n",
"workflow = SolarPanelComparisonWorkflow(\n",
" agent=agent, requirements_path=requirements_path, verbose=True, timeout=120\n",
")\n",
"\n",
"# Run the workflow; pass the datasheet path in the StartEvent\n",
"result = await workflow.run(\n",
" datasheet_path=\"./data/solar_panel_e2e_comparison/datasheet.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e1e61f1e-8701-4acc-8f99-cc89d8aae535",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"********Final Comparison Report:********\n",
"\n",
"{\n",
" \"component_name\": \"TSM-DE08M.08(II)\",\n",
" \"meets_requirements\": true,\n",
" \"summary\": \"The solar panel TSM-DE08M.08(II) meets all the design requirements, making it a suitable choice for the intended application.\",\n",
" \"details\": {\n",
" \"Maximum Power Output\": \"PASS - The panel's power output ranges from 360 W to 385 W, exceeding the minimum requirement of 350 W.\",\n",
" \"Open-Circuit Voltage\": \"PASS - The datasheet does not specify Voc, but the panel meets other critical requirements. Verification of Voc is recommended.\",\n",
" \"Short-Circuit Current\": \"PASS - The datasheet does not specify Isc, but the panel meets other critical requirements. Verification of Isc is recommended.\",\n",
" \"Efficiency\": \"PASS - The panel's efficiency is 21.0%, which is above the required 18%.\",\n",
" \"Temperature Coefficient\": \"PASS - The temperature coefficient is -0.34%/°C, which is better than the maximum allowable -0.5%/°C.\"\n",
" }\n",
"}\n"
]
}
],
"source": [
"print(\"\\n********Final Comparison Report:********\\n\")\n",
"print(result[\"report\"].model_dump_json(indent=4))\n",
"# print(\"\\n********Datasheet Content:********\\n\", result[\"datasheet_content\"])"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}

Before

Width:  |  Height:  |  Size: 6.9 MiB

After

Width:  |  Height:  |  Size: 6.9 MiB

+618
View File
@@ -0,0 +1,618 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Advanced RAG with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook is a complete walkthrough for using LlamaParse with advanced indexing/retrieval techniques in LlamaIndex over the Apple 10K Filing. \n",
"\n",
"This allows us to ask sophisticated questions that aren't possible with \"naive\" parsing/indexing techniques with existing models."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2021/q4/_10-K-2021-(As-Filed).pdf\" -O apple_2021_10k.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Some OpenAI and LlamaParse details"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\"\n",
"\n",
"# Using OpenAI API for embeddings/llms\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-proj-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import Settings\n",
"\n",
"embed_model = OpenAIEmbedding(model_name=\"text-embedding-3-small\")\n",
"llm = OpenAI(model=\"gpt-4o-mini\")\n",
"\n",
"Settings.llm = llm\n",
"Settings.embed_model = embed_model"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using brand new `LlamaParse` PDF reader for PDF Parsing\n",
"\n",
"We also compare three different retrieval/query engine strategies:\n",
"1. Baseline using default parsing from `SimpleDirectoryReader`\n",
"2. Using raw markdown text as nodes for building index and apply simple query engine for generating the results;\n",
"3. Using markdown + page screenshots to help retrieve the proper nodes."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id e403a457-1721-4093-82bf-4a316d2d637a\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"result = await LlamaParse(take_screenshot=True).aparse(\"./apple_2021_10k.pdf\")\n",
"\n",
"markdown_nodes = await result.aget_markdown_nodes(split_by_page=True)\n",
"screenshot_image_nodes = await result.aget_image_nodes(\n",
" include_screenshot_images=True,\n",
" include_object_images=False,\n",
" image_download_dir=\"./images\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import SimpleDirectoryReader\n",
"\n",
"baseline_documents = SimpleDirectoryReader(\n",
" input_files=[\"apple_2021_10k.pdf\"]\n",
").load_data()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup Baseline Index\n",
"\n",
"For comparison, we setup a naive RAG pipeline with default parsing and standard chunking, indexing, retrieval."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import VectorStoreIndex\n",
"\n",
"baseline_index = VectorStoreIndex.from_documents(baseline_documents)\n",
"baseline_query_engine = baseline_index.as_query_engine(similarity_top_k=3)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup our LlamaParse Indexes\n",
"\n",
"Using both the markdown and screenshot images, we can build two different indexes.\n",
"\n",
"1. An index over just the markdown documents\n",
"2. A custom index that uses the markdown + screenshot images to help with response quality."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import VectorStoreIndex\n",
"\n",
"markdown_index = VectorStoreIndex(nodes=markdown_nodes)\n",
"markdown_query_engine = markdown_index.as_query_engine(similarity_top_k=3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.indices import MultiModalVectorStoreIndex\n",
"from llama_index.embeddings.huggingface import HuggingFaceEmbedding\n",
"from llama_index.core import Settings\n",
"\n",
"# could also use other API-based multimodal models like voyageai or jinaai\n",
"# Note: this may take quite a while if running on CPU!\n",
"image_embed_model = HuggingFaceEmbedding(\n",
" model_name=\"llamaindex/vdr-2b-multi-v1\",\n",
" embed_batch_size=2,\n",
" trust_remote_code=True,\n",
" cache_folder=\"./hf_cache_2\",\n",
" device=\"cpu\", # set to \"cuda\" if you have a GPU or remove to auto-detect\n",
")\n",
"\n",
"multi_modal_index = MultiModalVectorStoreIndex(\n",
" nodes=[*markdown_nodes, *screenshot_image_nodes],\n",
" embed_model=Settings.embed_model,\n",
" image_embed_model=image_embed_model,\n",
" show_progress=True,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Below, we will create a custom query engine that does a few things\n",
"1. Retrieves both image nodes and text nodes\n",
"2. Combines them into two lists -- one where images and texts come from the same page, and one where we have texts alone\n",
"3. Use a Jinja-based `RichPromptTemplate` to format the retrieved content automatically into a list of multimodal chat messages\n",
"4. Send our messages to the LLM and return a result\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.async_utils import asyncio_run\n",
"from llama_index.core.llms import LLM\n",
"from llama_index.core.query_engine import CustomQueryEngine\n",
"from llama_index.core.prompts import RichPromptTemplate\n",
"from llama_index.core.response import Response\n",
"from llama_index.core.schema import NodeWithScore\n",
"from llama_index.core import Settings\n",
"\n",
"TEXT_IMAGE_PROMPT_TEMPLATE = RichPromptTemplate(\n",
" \"\"\"\n",
"<context>\n",
"Here is some retrieved content from a knowledge base:\n",
"{% for image_path, text in images_and_texts %}\n",
"<page>\n",
"<text>{{ text }}</text>\n",
"<image>{{ image_path | image }}</image>\n",
"</page>\n",
"{% endfor %}\n",
"{% for text in texts %}\n",
"<page>\n",
"<text>{{ text }}</text>\n",
"</page>\n",
"{% endfor %}\n",
"</context>\n",
"\n",
"Using the context, answer the following question:\n",
"<query>{{ query_str }}</query>\n",
"\"\"\"\n",
")\n",
"\n",
"\n",
"class SimpleMultiModalQueryEngine(CustomQueryEngine):\n",
" def __init__(\n",
" self,\n",
" index: MultiModalVectorStoreIndex,\n",
" image_top_k: int = 4,\n",
" text_top_k: int = 4,\n",
" llm: LLM | None = None,\n",
" **kwargs\n",
" ):\n",
" super().__init__(**kwargs)\n",
" self._retriever = index.as_retriever(\n",
" similarity_top_k=text_top_k, image_similarity_top_k=image_top_k\n",
" )\n",
" self._llm = llm or Settings.llm\n",
"\n",
" def _match_images_and_texts(\n",
" self, text_results: list[NodeWithScore], image_results: list[NodeWithScore]\n",
" ) -> tuple[list[NodeWithScore], list[NodeWithScore]]:\n",
" # combine results, prioritize images and texts\n",
" # if both an image and matching text was retrieved, that is a strong indicator\n",
" images_and_texts = []\n",
" text_keys = {\n",
" (x.metadata[\"page_number\"], x.metadata[\"file_name\"]): x\n",
" for x in text_results\n",
" }\n",
" for image_result in image_results:\n",
" key = (\n",
" image_result.metadata[\"page_number\"],\n",
" image_result.metadata[\"file_name\"],\n",
" )\n",
" # add matching text to results if available\n",
" if key in text_keys:\n",
" text_result = text_keys[key]\n",
" images_and_texts.append(\n",
" (image_result.node.image_path, text_result.node.text)\n",
" )\n",
"\n",
" # remove from list\n",
" text_keys.pop(key)\n",
"\n",
" # get the remaining texts as a fallback\n",
" texts = [result.node.text for result in text_keys.values()]\n",
"\n",
" return images_and_texts, texts\n",
"\n",
" def custom_query(self, query_str: str) -> Response:\n",
" # wrap the async method to avoid code duplication\n",
" # asyncio_run is a slightly safer asyncio.run() call\n",
" return asyncio_run(self.acustom_query(query_str))\n",
"\n",
" async def acustom_query(self, query_str: str) -> Response:\n",
" text_results = await self._retriever.atext_retrieve(query_str)\n",
" image_results = await self._retriever.atext_to_image_retrieve(query_str)\n",
"\n",
" images_and_texts, texts = self._match_images_and_texts(\n",
" text_results, image_results\n",
" )\n",
" messages = TEXT_IMAGE_PROMPT_TEMPLATE.format_messages(\n",
" images_and_texts=images_and_texts, texts=texts, query_str=str(query_str)\n",
" )\n",
"\n",
" response = await self._llm.achat(messages)\n",
"\n",
" return Response(\n",
" response.message.content, source_nodes=[*text_results, *image_results]\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"multimodal_query_engine = SimpleMultiModalQueryEngine(\n",
" index=multi_modal_index,\n",
" image_top_k=3,\n",
" text_top_k=3,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Try out the Query Engines and Compare!\n",
"\n",
"Now with our three query engines assembled, we can compare each approach with a rough \"vibes-based\" evaluation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $190,516 million.\n",
"\n",
"***********Markdown Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $191,830 million.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $191,830 million.\n"
]
}
],
"source": [
"query = \"What were the total fair value of marketable securities in 2020\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"As we can see, the multimodal and markdown query engines are able to retrieve the correct content, while the default query engine struggles to find the correct total value."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can also inspect the source nodes, and see the pages that were retrieved. Here is the correct page for the total fair value of marketable securities in 2020:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'images/page_41.jpg'"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"response_3.source_nodes[4].node.image_path"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Lets try a few more queries to see how the query engines perform."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The effective interest rates for the debt issuances in 2021 were as follows:\n",
"\n",
"- Floating-rate notes: 0.48% 0.63%\n",
"- Fixed-rate notes: 0.03% 4.78% for maturities from 2022 to 2060\n",
"- Fixed-rate notes issued in the second quarter: 0.75% 2.81% for maturities from 2026 to 2061\n",
"- Fixed-rate notes issued in the fourth quarter: 1.43% 2.86% for maturities from 2028 to 2061\n",
"\n",
"***********Markdown Query Engine***********\n",
"The effective interest rates for the debt issuances in 2021 were as follows:\n",
"\n",
"- Floating-rate notes: 0.48% 0.63%\n",
"- Fixed-rate notes: 0.03% 4.78% for the 0.000% 4.650% notes, 0.75% 2.81% for the 0.700% 2.800% notes, and 1.43% 2.86% for the 1.400% 2.850% notes.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The effective interest rates of all debt issuances in 2021 were as follows:\n",
"\n",
"1. **Floating-rate notes**: 0.48% 0.63%\n",
"2. **Fixed-rate 0.000% 4.650% notes**: 0.03% 4.78%\n",
"3. **Fixed-rate 0.700% 2.800% notes**: 0.75% 2.81%\n",
"4. **Fixed-rate 1.400% 2.850% notes**: 1.43% 2.86%\n"
]
}
],
"source": [
"query = \"What were the effective interest rates of all debt issuances in 2021\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The federal deferred tax amounts for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2019**: $(2,939)\n",
"- **2020**: $(3,619)\n",
"- **2021**: $(7,176)\n",
"\n",
"These figures represent the deferred tax expense for each respective year.\n",
"\n",
"***********Markdown Query Engine***********\n",
"As of September 25, 2021, the total deferred tax assets and liabilities for the years 2021 and 2020 are as follows:\n",
"\n",
"**Deferred Tax Assets:**\n",
"- 2021: $25,176 million\n",
"- 2020: $19,336 million\n",
"\n",
"**Deferred Tax Liabilities:**\n",
"- 2021: $7,200 million\n",
"- 2020: $10,138 million\n",
"\n",
"**Net Deferred Tax Assets:**\n",
"- 2021: $13,073 million\n",
"- 2020: $8,157 million\n",
"\n",
"The information for 2019 is not provided in the context.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The federal deferred tax assets and liabilities for the years 2019 to 2021 are as follows:\n",
"\n",
"### Deferred Tax Assets (in millions):\n",
"- **2021**: $25,176\n",
"- **2020**: $19,336\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"### Deferred Tax Liabilities (in millions):\n",
"- **2021**: $7,200\n",
"- **2020**: $10,138\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"### Net Deferred Tax Assets (in millions):\n",
"- **2021**: $13,073\n",
"- **2020**: $8,157\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"The significant components of deferred tax assets and liabilities reflect the effects of tax credits and temporary differences between financial statement carrying amounts and their respective tax bases.\n"
]
}
],
"source": [
"query = \"federal deferred tax in 2019-2021\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- 2021: $1,620\n",
"- 2020: $455\n",
"- 2019: $475\n",
"\n",
"This indicates an increase of $1,165 million from 2020 to 2021, a decrease of $20 million from 2018 to 2019, and an increase of $80 million from 2019 to 2020.\n",
"\n",
"***********Markdown Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2021**: $1,620\n",
"- **2020**: $455\n",
"- **2019**: $475\n",
"\n",
"The changes in current state taxes from year to year are:\n",
"\n",
"- From 2019 to 2020: Decrease of $20 million\n",
"- From 2020 to 2021: Increase of $1,165 million\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2021**: $1,620\n",
"- **2020**: $455\n",
"- **2019**: $475\n",
"\n",
"So, the changes are:\n",
"- From 2019 to 2020: Decrease of $20 million\n",
"- From 2020 to 2021: Increase of $1,165 million\n"
]
}
],
"source": [
"query = \"current state taxes per year in 2019-2021 (include +/-)\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama-parse-aNC435Vv-py3.10",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
+196
View File
@@ -0,0 +1,196 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse Usage"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/1706.03762.pdf\" -O \"./attention.pdf\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 79ae653c-4598-4bd0-ba6e-b3dab7eab57e\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"result = await LlamaParse().aparse(\"./attention.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"1 Introduction\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neural networks\n",
"in particular, have been firmly established as state of the art approaches in sequence modeling and\n",
"transduction problems such as language modeling and machine translation [35, 2, 5]. Numerous\n",
"efforts have since continued to push the boundaries of recurrent language models and encoder-decoder\n",
"architectures [38, 24, 15].\n",
"Recurrent models typically factor computation along the symbol positions of the input and output\n",
"sequences. Aligning the positions to steps in computation time, they generate a sequence of hidden\n",
"states ht, as a function of the previous hidden state ht1 and the input for position t. This inherently\n",
"sequential nature precludes parallelization within training examples, which becomes critical at longer\n",
"sequence lengths, as memory constraints limit batching across examples. Recent work has achieved\n",
"significant improvements in computational efficiency through factorization tricks [21] and conditional\n",
"computation [32], while also improving model performance in case of the latter. The fundamental\n",
"constraint of sequential computation, however, remains.\n",
"Attention mechanisms have become an integral part of compelling sequence modeling and transduc-\n",
"tion models in various tasks, allowing modeling of dependencies without regard to their distance in\n",
"the input or output sequences [2, 19]. In all but a few cases [27], however, such attention mechanisms\n",
"are used in conjunction with a recurrent network.\n",
"In this work we propose the Transformer, a model architecture eschewing recurrence and instead\n",
"relying entirely on an attention mechanism to draw global dependencies between input and output.\n",
"The Transformer allows for significantly more parallelization and can reach a new state of the art in\n",
"translation quality after being trained for as little as twelve hours on eight P100 GPUs.\n",
"2 Background\n",
"The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU\n",
"[16], ByteNet [18] and ConvS2S [9], all of which use convolutional neural networks as basic building\n",
"block, computing hidden representations in parallel for all input and output positions. In these models,\n",
"the number of operations required to relate signals from two arbitrary input or output positions grows\n",
"in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes\n",
"it more difficult to learn dependencies between distant positions [12]. In the Transformer this is\n",
"reduced to a constant number of operations, albeit at the cost of reduced effective resolution due\n",
"to averaging attention-weighted positions, an effect we counteract with Multi-Head Attention as\n",
"described in section 3.2.\n",
"Self-attention, sometimes called intra-attention is an attention mechanism relating different positions\n",
"of a single sequence in order to compute a representation of the sequence. Self-attention has been\n",
"used successfully in a variety of tasks including reading comprehension, abstractive summarization,\n",
"textual entailment and learning task-independent sentence representations [4, 27, 28, 22].\n",
"End-to-end memory networks are based on a recurrent attention mechanism instead of sequence-\n",
"aligned recurrence and have been shown to perform well on simple-language question answering and\n",
"language modeling tasks [34].\n",
"To the best of our knowledge, however, the Transformer is the first transduction model relying\n",
"entirely on self-attention to compute representations of its input and output without using sequence-\n",
"aligned RNNs or convolution. In the following sections, we will describe the Transformer, motivate\n",
"self-attention and discuss its advantages over models such as [17, 18] and [9].\n",
"3 Model Architecture\n",
"Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].\n",
"Here, the encoder maps an input sequence of symbol representations (x1, ..., xn) to a sequence\n",
"of continuous representations z = (z1, ..., zn). Given z, the decoder then generates an output\n",
"sequence (y1, ..., ym) of symbols one element at a time. At each step the model is auto-regressive\n",
"[10], consuming the previously generated symbols as additional input when generating the next.\n",
" 2\n"
]
}
],
"source": [
"documents = result.get_text_documents(split_by_page=True)\n",
"print(documents[1].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"arXiv:1706.03762v7 [cs.CL] 2 Aug 2023\n",
"\n",
"Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.\n",
"\n",
"# Attention Is All You Need\n",
"\n",
"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit\n",
"\n",
"Google Brain Google Brain Google Research Google Research\n",
"\n",
"avaswani@google.com noam@google.com nikip@google.com usz@google.com\n",
"\n",
"Llion Jones Aidan N. Gomez † Łukasz Kaiser\n",
"\n",
"Google Research University of Toronto Google Brain\n",
"\n",
"llion@google.com aidan@cs.toronto.edu lukaszkaiser@google.com\n",
"\n",
"Illia Polosukhin ‡\n",
"\n",
"illia.polosukhin@gmail.com\n",
"\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.\n",
"\n",
"Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research.\n",
"\n",
"†Work performed while at Google Brain.\n",
"\n",
"‡Work performed while at Google Research.\n",
"\n",
"31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.\n"
]
}
],
"source": [
"documents = result.get_markdown_documents(split_by_page=True)\n",
"print(documents[0].text)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
@@ -31,11 +31,11 @@
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-index\n",
"!pip install llama-index-core\n",
"!pip install llama-index-llms-anthropic llama-index-multi-modal-llms-anthropic\n",
"!pip install llama-index-embeddings-huggingface\n",
"!pip install llama-cloud-services"
"%pip install llama-index\n",
"%pip install llama-index-core\n",
"%pip install llama-index-llms-anthropic\n",
"%pip install llama-index-embeddings-huggingface\n",
"%pip install llama-cloud-services"
]
},
{
@@ -45,11 +45,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -68,7 +63,7 @@
"source": [
"from llama_index.llms.anthropic import Anthropic\n",
"\n",
"llm = Anthropic(model=\"claude-3-opus-20240229\", temperature=0.0)"
"llm = Anthropic(model=\"claude-3-5-sonnet-20241022\")"
]
},
{
@@ -131,28 +126,8 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(verbose=True)\n",
"json_objs = parser.get_json_result(\"./uber_10q_march_2022.pdf\")\n",
"json_list = json_objs[0][\"pages\"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b26d21d1-05b5-4f49-b937-c13106a84015",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"text\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
"parser = LlamaParse(take_screenshot=True)\n",
"result = await parser.aparse(\"./uber_10q_march_2022.pdf\")"
]
},
{
@@ -162,7 +137,12 @@
"metadata": {},
"outputs": [],
"source": [
"text_nodes = get_text_nodes(json_list)"
"text_nodes = await result.aget_text_nodes(split_by_page=True)\n",
"image_nodes = await result.aget_image_nodes(\n",
" include_screenshot_images=True,\n",
" include_object_images=True,\n",
" image_download_dir=\"./uber_10q_images\",\n",
")"
]
},
{
@@ -172,7 +152,7 @@
"source": [
"## Extract/Index images from image dicts\n",
"\n",
"Here we use a multimodal model to extract and index images from image dictionaries."
"Here we use a multimodal model to caption images and create text nodes for indexing."
]
},
{
@@ -190,27 +170,32 @@
}
],
"source": [
"# call get_images on parser, convert to ImageDocuments\n",
"!mkdir llama2_images\n",
"!mkdir -p llama2_images\n",
"\n",
"from llama_index.core.schema import ImageDocument\n",
"from llama_index.multi_modal_llms.anthropic import AnthropicMultiModal\n",
"from llama_index.core.llms import ChatMessage, ImageBlock, TextBlock\n",
"from llama_index.core.schema import ImageNode, TextNode\n",
"from llama_index.llms.anthropic import Anthropic\n",
"\n",
"\n",
"def get_image_text_nodes(json_objs: List[dict]):\n",
"def get_image_text_nodes(image_nodes: list[ImageNode]):\n",
" \"\"\"Extract out text from images using a multimodal model.\"\"\"\n",
" anthropic_mm_llm = AnthropicMultiModal(max_tokens=300)\n",
" image_dicts = parser.get_images(json_objs, download_path=\"llama2_images\")\n",
" image_documents = []\n",
" llm = Anthropic(model=\"claude-3-5-haiku-20241022\", max_tokens=300)\n",
" img_text_nodes = []\n",
" for image_dict in image_dicts:\n",
" image_doc = ImageDocument(image_path=image_dict[\"path\"])\n",
" response = anthropic_mm_llm.complete(\n",
" prompt=\"Describe the images as alt text\",\n",
" image_documents=[image_doc],\n",
" for image_node in image_nodes:\n",
" image_path = image_node.image_path\n",
" message = ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" TextBlock(text=\"Describe the images as alt text\"),\n",
" ImageBlock(path=image_path),\n",
" ],\n",
" )\n",
" response = llm.chat([message])\n",
" text_node = TextNode(\n",
" text=str(response.message.content), metadata={\"path\": image_path}\n",
" )\n",
" text_node = TextNode(text=str(response), metadata={\"path\": image_dict[\"path\"]})\n",
" img_text_nodes.append(text_node)\n",
"\n",
" return img_text_nodes"
]
},
@@ -221,7 +206,7 @@
"metadata": {},
"outputs": [],
"source": [
"image_text_nodes = get_image_text_nodes(json_objs)"
"image_text_nodes = get_image_text_nodes(image_nodes)"
]
},
{
+553
View File
@@ -0,0 +1,553 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "d27f1082-cd10-405e-9570-6f0e934bba8b",
"metadata": {},
"source": [
"# LlamaParse `JobResult` Tour\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"The `JobResult` object is the main object returned by the LlamaParse API. It contains all the information about the job, including the parsed data, metadata, and any errors.\n",
"\n",
"This notebook walks through each component of the `JobResult` object and shows you what it contains."
]
},
{
"cell_type": "markdown",
"id": "a004db48-8d3f-421c-915a-477692f71b90",
"metadata": {},
"source": [
"## Setup\n",
"\n",
"Let's bring in our imports and set up our API keys."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "bc6a7a4b-b568-4db5-bcba-62f5c517ff3a",
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "0879301c-ff91-4431-941a-6c0ef7cd8fe2",
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-..\""
]
},
{
"cell_type": "markdown",
"id": "b411d2ee-3e6b-45b0-b532-4a8e3abcdea0",
"metadata": {},
"source": [
"## Load Data\n",
"\n",
"Let's load a large and complex PDF, San Francisco's 2023 proposed budget."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c39d408f-e885-4940-85c7-b09ca3bc7cb7",
"metadata": {},
"outputs": [],
"source": [
"!wget 'https://www.dropbox.com/scl/fi/vip161t63s56vd94neqlt/2023-CSF_Proposed_Budget_Book_June_2023_Master_Web.pdf?rlkey=hemoce3w1jsuf6s2bz87g549i&dl=0' -O './san_francisco_budget_2023.pdf'"
]
},
{
"cell_type": "markdown",
"id": "c2f42af8-afb3-4b3b-82d3-6b332fb38aa4",
"metadata": {},
"source": [
"## Using LlamaParse for Basic PDF Parsing\n",
"\n",
"Let's parse our document!"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "9c9cd670-8229-4ad6-99a9-845bd82b7ec1",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id d12d419a-52fc-400c-9f88-f61b352d3fb2\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse()\n",
"result = await parser.aparse(\"./san_francisco_budget_2023.pdf\")"
]
},
{
"cell_type": "markdown",
"id": "11c22bab",
"metadata": {},
"source": [
"Every job will come back with some metadata about the job:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c588c578",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"JobMetadata(job_credits_usage=0, job_pages=0, job_auto_mode_triggered_pages=0, job_is_cache_hit=True)"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"result.job_metadata"
]
},
{
"cell_type": "markdown",
"id": "1e96b7c9",
"metadata": {},
"source": [
"Since this was a re-run, I can see that a cache hit occurred. Jobs are cached for 48 hours by default."
]
},
{
"cell_type": "markdown",
"id": "6543d2c6",
"metadata": {},
"source": [
"Beyond this, we can explore the parsed data per-page:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "af9f3717",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"362\n"
]
}
],
"source": [
"print(len(result.pages))"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f8845fac",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"dict_keys(['page', 'text', 'md', 'images', 'charts', 'tables', 'layout', 'items', 'status', 'links', 'width', 'height', 'triggeredAutoMode', 'parsingMode', 'structuredData', 'noStructuredContent', 'noTextContent'])\n"
]
}
],
"source": [
"print(result.pages[0].model_dump().keys())"
]
},
{
"cell_type": "markdown",
"id": "6261f5e3",
"metadata": {},
"source": [
"Inside the page object, you can see nearly every detail about the page.\n",
"\n",
"Most of these will depend on the settings you used when parsing. Since we used the default settings, we get the text and markdown for each page, as well as a list of all the elements on the page.\n",
"\n",
"* `page`: this is simply the page number, starting at 1.\n",
"* `text`: this is the text of the page, as extracted by the parser.\n",
"* `images`: this is an array of all the images on the page, including metadata and text OCRed out of the images, as well as a full-page screenshot of the entire page.\n",
"* `charts`: this is an array of all the charts on the page, including metadata and text OCRed out of the charts, as well as a full-page screenshot of the entire chart.\n",
"* `layout`: this is an array of all the layout elements on the page, if you are using layout mode.\n",
"* `items`: This is an array of all the parsed elements on the page, as used to render the markdown, but separated out into their own objects. This is useful if you want to do more processing on the data.\n",
"* `links`: this is an array of all the links on the page, if you are used `annotate_links=True`\n",
"* `status`: this is the status of the page, which is usually \"OK\" unless there was an error processing the page.\n",
"* `width` and `height`: these are the dimensions of the page in pixels.\n",
"* `parsingMode`: Contains the specific parsing mode that was used for the page.\n",
"* `triggeredAutoMode`: this indicates whether the page triggered auto mode; see [LlamaParse docs](https://docs.cloud.llamaindex.ai/llamaparse/getting_started) for more details.\n",
"* `structuredData`/`noStructuredContent`: these are set if you are using structured mode; see [LlamaParse docs](https://docs.cloud.llamaindex.ai/llamaparse/getting_started) for more details.\n",
"* `noTextContent`: this is true if the page was empty of text.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7a4cc901",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
" CITY & COUNTY OF SAN FRANCISCO, CALIFORNIA\n",
" PROPOSED BUDGET\n",
" FISCAL YEARS 2023-2024 & 2024-2025\n",
" LONDON N. BREED\n",
" MAYORS OFFICE OF PUBLIC POLICY AND FINANCE\n",
" Anna Duning, Director of Mayors Fisher Zhu, Fiscal and Policy Analyst\n",
" Office of Public Policy and Finance Anya Shutovska, Fiscal and Policy Analyst\n",
" Sally Ma, Deputy Budget Director\n",
"Radhika Mehlotra, Senior Fiscal and Policy Analyst Jack English, Fiscal and Policy Analyst\n",
" Damon Daniels, Fiscal and Policy Analyst Xang Hang, Junior Fiscal and Policy Analyst\n",
" Matthew Puckett, Fiscal and Policy Analyst Tabitha Romero-Bothi, Fiscal and Policy Assistant\n"
]
}
],
"source": [
"print(result.pages[0].text[:1000])"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "2d5a5bc2",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# CITY & COUNTY OF SAN FRANCISCO, CALIFORNIA\n",
"\n",
"# PROPOSED BUDGET\n",
"\n",
"# FISCAL YEARS 2023-2024 & 2024-2025\n",
"\n",
"# LONDON N. BREED\n",
"\n",
"# MAYORS OFFICE OF PUBLIC POLICY AND FINANCE\n",
"\n",
"Anna Duning, Director of Mayors Office of Public Policy and Finance\n",
"\n",
"Fisher Zhu, Fiscal and Policy Analyst\n",
"\n",
"Anya Shutovska, Fiscal and Policy Analyst\n",
"\n",
"Sally Ma, Deputy Budget Director\n",
"\n",
"Radhika Mehlotra, Senior Fiscal and Policy Analyst\n",
"\n",
"Jack English, Fiscal and Policy Analyst\n",
"\n",
"Damon Daniels, Fiscal and Policy Analyst\n",
"\n",
"Xang Hang, Junior Fiscal and Policy Analyst\n",
"\n",
"Matthew Puckett, Fiscal and Policy Analyst\n",
"\n",
"Tabitha Romero-Bothi, Fiscal and Policy Assistant\n"
]
}
],
"source": [
"print(result.pages[0].md[:1000])"
]
},
{
"cell_type": "markdown",
"id": "32de4c62",
"metadata": {},
"source": [
"## Images\n",
"\n",
"By default, images embedded in documents that can be extracted are part of the result object."
]
},
{
"cell_type": "markdown",
"id": "802d4a98",
"metadata": {},
"source": [
"We can also specify to take screenshots of every page:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "2ee78f2f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id e6332422-803b-404d-8d0d-ad510fa56c09\n",
"..."
]
}
],
"source": [
"parser = LlamaParse(take_screenshot=True)\n",
"result = await parser.aparse(\"./san_francisco_budget_2023.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "fab32886",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"[ImageItem(name='page_1.jpg', height=792.0, width=612.0, x=0.0, y=0.0, original_width=1236, original_height=1600, type='full_page_screenshot')]\n"
]
}
],
"source": [
"print(result.pages[0].images)"
]
},
{
"cell_type": "markdown",
"id": "9eba9e52",
"metadata": {},
"source": [
"We can download images (either their bytes or to a local file) using the `JobResult` object as well!"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "a7aa0a29",
"metadata": {},
"outputs": [],
"source": [
"# single image\n",
"image_data = await result.aget_image_data(result.pages[0].images[0].name)\n",
"\n",
"# save an image to a file\n",
"output_path = await result.asave_image(\n",
" result.pages[0].images[0].name, \"./json_tour_screenshots\"\n",
")\n",
"\n",
"# save all images\n",
"output_paths = await result.asave_all_images(\"./json_tour_screenshots\")"
]
},
{
"cell_type": "markdown",
"id": "eae4ece3",
"metadata": {},
"source": [
"## Items\n",
"\n",
"This is an array of all the parsed elements on the page, as used to render the markdown, but separated out into their own objects. This is useful if you want to do more processing on the data. Let's take a look:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c10b9d7d",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"type='heading' lvl=1 value='CITY & COUNTY OF SAN FRANCISCO, CALIFORNIA' md='# CITY & COUNTY OF SAN FRANCISCO, CALIFORNIA' rows=None bBox=BBox(x=176.0, y=52.0, w=277.0, h=12.0)\n"
]
}
],
"source": [
"import json\n",
"\n",
"print(result.pages[0].items[0])"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "dcb9f832",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"type='heading' lvl=1 value='PROPOSED BUDGET' md='# PROPOSED BUDGET' rows=None bBox=BBox(x=89.0, y=118.0, w=451.0, h=47.0)\n"
]
}
],
"source": [
"print(result.pages[0].items[1])"
]
},
{
"cell_type": "markdown",
"id": "a7f64443",
"metadata": {},
"source": [
"As you can see you get different element types: text, headings, and tables. Each comes with its own `md` key containing a Markdown representation of that element, allowing you to easily summarize with only headings, tables only, etc..\n",
"\n",
"The ability to extract tables from visual data is really powerful. Let's take a look at page 35, which has some bar charts that get automatically converted into tables:\n",
"\n",
"<img src=\"./json_tour_screenshots/page_35.png\" alt=\"Page 35\" width=\"300\"/>\n"
]
},
{
"cell_type": "markdown",
"id": "e4ccee76",
"metadata": {},
"source": [
"The bar chart has been converted into a table, and even though explicit values are not included, the bar chart has been read and approximate values for each bar on the chart have been included!"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7d6404a5",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"type='table' lvl=None value=None md=\"Source: U.S. Census Bureau, 2017-2021 American Community Survey 5-years Estimate.\\n|Race|Educational Level|Number of Residents| | | | |\\n|---|---|---|---|---|---|---|\\n|Age Group| | | | | | |\\n|Under 5 Years|5 to 19 Years|20 to 34 Years|35 to 59 Years|60 and Over| | |\\n|Graduate or professional degree|Bachelor's degree|Associate's degree|Some college, no degree|High school graduate (includes equivalency)|9th to 12th grade, no diploma|Less than 9th grade|\" rows=[[], ['Race', 'Educational Level', 'Number of Residents', '', '', '', ''], ['---', '---', '---', '---', '---', '---', '---'], ['Age Group', '', '', '', '', '', ''], ['Under 5 Years', '5 to 19 Years', '20 to 34 Years', '35 to 59 Years', '60 and Over', '', ''], ['Graduate or professional degree', \"Bachelor's degree\", \"Associate's degree\", 'Some college, no degree', 'High school graduate (includes equivalency)', '9th to 12th grade, no diploma', 'Less than 9th grade']] bBox=BBox(x=68.0, y=129.0, w=613.0, h=3067.0)\n"
]
}
],
"source": [
"print(result.pages[34].items[6])"
]
},
{
"cell_type": "markdown",
"id": "9570d3b8",
"metadata": {},
"source": [
"### `links`\n",
"\n",
"Our budget PDF doesn't have any links, so let's load a different PDF with links and see what we get.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "fb0da11a",
"metadata": {},
"outputs": [],
"source": [
"!wget 'https://www.dropbox.com/scl/fi/hay06lyxc49gkuh91oek6/basic-link-1.pdf?rlkey=uije7yb0lxqgqwk7p7hnqepdx&dl=0' -O './basic-link-1.pdf'"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e7e393e6",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 9b2df975-af3c-4868-99e2-520ce0b21f4d\n"
]
}
],
"source": [
"parser = LlamaParse(annotate_links=True)\n",
"result = await parser.aparse(\"./basic-link-1.pdf\")"
]
},
{
"cell_type": "markdown",
"id": "701ada4b",
"metadata": {},
"source": [
"This is a very simple document with some internal and external links:\n",
"\n",
"<img src=\"./json_tour_screenshots/links_page.png\" alt=\"Page 1\" width=\"300\"/>\n"
]
},
{
"cell_type": "markdown",
"id": "2e4de7de",
"metadata": {},
"source": [
"The parser finds the external links and their labels and includes them in the `links` section:"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "29bf7e3c",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"[{'url': 'https://www.antennahouse.com/', 'text': 'Antenna House, Inc.'}, {'url': 'https://www.antennahouse.com/', 'text': 'Linking to a website (https://www.antennahouse.com/)'}]\n"
]
}
],
"source": [
"print(result.pages[0].links)"
]
},
{
"cell_type": "markdown",
"id": "ac9088a2",
"metadata": {},
"source": [
"This concludes our tour! I hope this makes clear the power of JSON mode and the flexibility it gives you over what parts of your documents you can use."
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
@@ -31,14 +31,9 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
@@ -79,8 +74,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"text\", language=\"fr\")\n",
"documents = parser.load_data(\"./treasury_report.pdf\")"
"parser = LlamaParse(language=\"fr\")\n",
"result = await parser.aparse(\"./treasury_report.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -252,8 +248,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"text\", language=\"ch_sim\")\n",
"documents = parser.load_data(\"./chinese_pdf.pdf\")"
"parser = LlamaParse(language=\"ch_sim\")\n",
"result = await parser.aparse(\"./chinese_pdf.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -406,8 +403,9 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"base_parser = LlamaParse(result_type=\"text\", language=\"en\")\n",
"base_documents = parser.load_data(\"./chinese_pdf2.pdf\")"
"base_parser = LlamaParse(language=\"en\")\n",
"result = await base_parser.aparse(\"./chinese_pdf2.pdf\")\n",
"base_documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -60,11 +60,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import requests\n",
"import pymongo\n",
"\n",
@@ -72,7 +67,7 @@
"from llama_cloud_services import LlamaParse\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"from llama_index.core.node_parser import SimpleNodeParser"
"from llama_index.core.node_parser import SentenceSplitter"
]
},
{
@@ -137,7 +132,8 @@
}
],
"source": [
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
"result = await LlamaParse().aparse(file_path)\n",
"documents = result.get_text_documents(split_by_page=False)"
]
},
{
@@ -203,7 +199,7 @@
"metadata": {},
"outputs": [],
"source": [
"node_parser = SimpleNodeParser()\n",
"node_parser = SentenceSplitter()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
@@ -55,10 +55,6 @@
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -80,25 +76,7 @@
"execution_count": null,
"id": "IjtKDQRLrylI",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 18:54:24-- https://arxiv.org/pdf/2409.18486\n",
"Resolving arxiv.org (arxiv.org)... 151.101.67.42, 151.101.131.42, 151.101.3.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.67.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 13986265 (13M) [application/pdf]\n",
"Saving to: o1.pdf\n",
"\n",
"o1.pdf 100%[===================>] 13.34M 11.8MB/s in 1.1s \n",
"\n",
"2024-12-05 18:54:26 (11.8 MB/s) - o1.pdf saved [13986265/13986265]\n",
"\n"
]
}
],
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/2409.18486\" -O \"o1.pdf\""
]
@@ -118,25 +96,6 @@
"Using your own API key may incur additional costs from your model provider and could result in failed pages or documents if you do not have sufficient usage limits."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"md\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
"cell_type": "markdown",
"id": "1b5d6da6",
@@ -163,15 +122,13 @@
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model_name=\"anthropic-sonnet-3.5\",\n",
" target_pages=\"24\"\n",
" # invalidate_cache=True\n",
")\n",
"json_objs = parser.get_json_result(\"o1.pdf\")\n",
"json_list = json_objs[0][\"pages\"]\n",
"docs = get_text_nodes(json_list)"
"result = await parser.aparse(\"o1.pdf\")\n",
"nodes = result.get_text_nodes(split_by_page=False)"
]
},
{
@@ -202,15 +159,13 @@
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model=\"openai-gpt4o\",\n",
" target_pages=\"24\",\n",
" # invalidate_cache=True\n",
")\n",
"json_objs_gpt4o = parser_gpt4o.get_json_result(\"o1.pdf\")\n",
"json_list_gpt4o = json_objs_gpt4o[0][\"pages\"]\n",
"docs_gpt4o = get_text_nodes(json_list_gpt4o)"
"result = await parser_gpt4o.aparse(\"o1.pdf\")\n",
"nodes = result.get_markdown_nodes(split_by_page=False)"
]
},
{
@@ -268,7 +223,7 @@
],
"source": [
"# using Sonnet-3.5\n",
"print(docs[0].get_content(metadata_mode=\"all\"))"
"print(nodes[0].get_content(metadata_mode=\"all\"))"
]
},
{
@@ -327,7 +282,7 @@
],
"source": [
"# using GPT-4o\n",
"print(docs_gpt4o[0].get_content(metadata_mode=\"all\"))"
"print(nodes[0].get_content(metadata_mode=\"all\"))"
]
}
],
@@ -47,11 +47,6 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -71,25 +66,7 @@
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 11:40:59-- https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf\n",
"Resolving raw.githubusercontent.com (raw.githubusercontent.com)... 2606:50c0:8000::154, 2606:50c0:8002::154, 2606:50c0:8003::154, ...\n",
"Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|2606:50c0:8000::154|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 1880483 (1.8M) [application/octet-stream]\n",
"Saving to: ./uber_2021.pdf\n",
"\n",
"./uber_2021.pdf 100%[===================>] 1.79M --.-KB/s in 0.1s \n",
"\n",
"2024-12-05 11:40:59 (14.2 MB/s) - ./uber_2021.pdf saved [1880483/1880483]\n",
"\n"
]
}
],
"outputs": [],
"source": [
"!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf' -O './uber_2021.pdf'"
]
@@ -119,9 +96,10 @@
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(target_pages=\"0,1,2\", result_type=\"markdown\")\n",
"parser = LlamaParse(target_pages=\"0,1,2\")\n",
"\n",
"documents = parser.load_data(\"./uber_2021.pdf\")"
"results = await parser.aparse(\"./uber_2021.pdf\")\n",
"documents = results.get_text_documents(split_by_page=True)"
]
},
{
@@ -0,0 +1,516 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Table Extraction with LlamaParse\n",
"\n",
"This notebook will show you how to extract tables and save them as CSV files thanks to LlamaParse advanced parsing capabilities."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**1. Install needed dependencies**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"! pip install llama-cloud-services pandas"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**2. Set you LLAMA_CLOUD_API_KEY as env variable**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"LLAMA_CLOUD_API_KEY: ··········\n"
]
}
],
"source": [
"import os\n",
"from getpass import getpass\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = getpass(\"LLAMA_CLOUD_API_KEY: \")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**3. Initialiaze the parser**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"parser = LlamaParse(result_type=\"markdown\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**4. Get data**\n",
"\n",
"This is a PDF with _lots_ of tables!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2025-07-16 16:20:41-- https://assets.accessible-digital-documents.com/uploads/2017/01/sample-tables.pdf\n",
"Resolving assets.accessible-digital-documents.com (assets.accessible-digital-documents.com)... 3.166.135.2, 3.166.135.62, 3.166.135.51, ...\n",
"Connecting to assets.accessible-digital-documents.com (assets.accessible-digital-documents.com)|3.166.135.2|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 145494 (142K) [application/pdf]\n",
"Saving to: sample-tables.pdf\n",
"\n",
"sample-tables.pdf 100%[===================>] 142.08K --.-KB/s in 0.04s \n",
"\n",
"2025-07-16 16:20:41 (3.72 MB/s) - sample-tables.pdf saved [145494/145494]\n",
"\n"
]
}
],
"source": [
"! wget https://assets.accessible-digital-documents.com/uploads/2017/01/sample-tables.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**5. Parse document**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b53949f7-9017-4b6a-b30c-be6227271ed2\n"
]
}
],
"source": [
"json_result = parser.get_json_result(\"sample-tables.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**6. Get tables!**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"tables = parser.get_tables(json_result, \"tables/\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"**7. Load tables**\n",
"\n",
"Let's show one example table!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"application/vnd.google.colaboratory.intrinsic+json": {
"summary": "{\n \"name\": \"display(df\",\n \"rows\": 8,\n \"fields\": [\n {\n \"column\": \"Rainfall\",\n \"properties\": {\n \"dtype\": \"string\",\n \"num_unique_values\": 5,\n \"samples\": [\n \"Average\",\n \"\",\n \"24 hour high\"\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Americas\",\n \"properties\": {\n \"dtype\": \"number\",\n \"std\": 908,\n \"min\": 9,\n \"max\": 2010,\n \"num_unique_values\": 8,\n \"samples\": [\n 104,\n 133,\n 2010\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Asia\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 201.0,\n 28.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Europe\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 193.0,\n 29.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n },\n {\n \"column\": \"Africa\",\n \"properties\": {\n \"dtype\": \"object\",\n \"num_unique_values\": 7,\n \"samples\": [\n \"\",\n 144.0,\n 20.0\n ],\n \"semantic_type\": \"\",\n \"description\": \"\"\n }\n }\n ]\n}",
"type": "dataframe"
},
"text/html": [
"\n",
" <div id=\"df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb\" class=\"colab-df-container\">\n",
" <div>\n",
"<style scoped>\n",
" .dataframe tbody tr th:only-of-type {\n",
" vertical-align: middle;\n",
" }\n",
"\n",
" .dataframe tbody tr th {\n",
" vertical-align: top;\n",
" }\n",
"\n",
" .dataframe thead th {\n",
" text-align: right;\n",
" }\n",
"</style>\n",
"<table border=\"1\" class=\"dataframe\">\n",
" <thead>\n",
" <tr style=\"text-align: right;\">\n",
" <th></th>\n",
" <th>Rainfall</th>\n",
" <th>Americas</th>\n",
" <th>Asia</th>\n",
" <th>Europe</th>\n",
" <th>Africa</th>\n",
" </tr>\n",
" </thead>\n",
" <tbody>\n",
" <tr>\n",
" <th>0</th>\n",
" <td>(inches)</td>\n",
" <td>2010</td>\n",
" <td></td>\n",
" <td></td>\n",
" <td></td>\n",
" </tr>\n",
" <tr>\n",
" <th>1</th>\n",
" <td>Average</td>\n",
" <td>104</td>\n",
" <td>201.0</td>\n",
" <td>193.0</td>\n",
" <td>144.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>2</th>\n",
" <td>24 hour high</td>\n",
" <td>15</td>\n",
" <td>26.0</td>\n",
" <td>27.0</td>\n",
" <td>18.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>3</th>\n",
" <td>12 hour high</td>\n",
" <td>9</td>\n",
" <td>10.0</td>\n",
" <td>11.0</td>\n",
" <td>12.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>4</th>\n",
" <td></td>\n",
" <td>2009</td>\n",
" <td></td>\n",
" <td></td>\n",
" <td></td>\n",
" </tr>\n",
" <tr>\n",
" <th>5</th>\n",
" <td>Average</td>\n",
" <td>133</td>\n",
" <td>244.0</td>\n",
" <td>155.0</td>\n",
" <td>166.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>6</th>\n",
" <td>24 hour high</td>\n",
" <td>27</td>\n",
" <td>28.0</td>\n",
" <td>29.0</td>\n",
" <td>20.0</td>\n",
" </tr>\n",
" <tr>\n",
" <th>7</th>\n",
" <td>12 hour high</td>\n",
" <td>11</td>\n",
" <td>12.0</td>\n",
" <td>13.0</td>\n",
" <td>16.0</td>\n",
" </tr>\n",
" </tbody>\n",
"</table>\n",
"</div>\n",
" <div class=\"colab-df-buttons\">\n",
"\n",
" <div class=\"colab-df-container\">\n",
" <button class=\"colab-df-convert\" onclick=\"convertToInteractive('df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb')\"\n",
" title=\"Convert this dataframe to an interactive table.\"\n",
" style=\"display:none;\">\n",
"\n",
" <svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\" viewBox=\"0 -960 960 960\">\n",
" <path d=\"M120-120v-720h720v720H120Zm60-500h600v-160H180v160Zm220 220h160v-160H400v160Zm0 220h160v-160H400v160ZM180-400h160v-160H180v160Zm440 0h160v-160H620v160ZM180-180h160v-160H180v160Zm440 0h160v-160H620v160Z\"/>\n",
" </svg>\n",
" </button>\n",
"\n",
" <style>\n",
" .colab-df-container {\n",
" display:flex;\n",
" gap: 12px;\n",
" }\n",
"\n",
" .colab-df-convert {\n",
" background-color: #E8F0FE;\n",
" border: none;\n",
" border-radius: 50%;\n",
" cursor: pointer;\n",
" display: none;\n",
" fill: #1967D2;\n",
" height: 32px;\n",
" padding: 0 0 0 0;\n",
" width: 32px;\n",
" }\n",
"\n",
" .colab-df-convert:hover {\n",
" background-color: #E2EBFA;\n",
" box-shadow: 0px 1px 2px rgba(60, 64, 67, 0.3), 0px 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
" fill: #174EA6;\n",
" }\n",
"\n",
" .colab-df-buttons div {\n",
" margin-bottom: 4px;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-convert {\n",
" background-color: #3B4455;\n",
" fill: #D2E3FC;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-convert:hover {\n",
" background-color: #434B5C;\n",
" box-shadow: 0px 1px 3px 1px rgba(0, 0, 0, 0.15);\n",
" filter: drop-shadow(0px 1px 2px rgba(0, 0, 0, 0.3));\n",
" fill: #FFFFFF;\n",
" }\n",
" </style>\n",
"\n",
" <script>\n",
" const buttonEl =\n",
" document.querySelector('#df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb button.colab-df-convert');\n",
" buttonEl.style.display =\n",
" google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
"\n",
" async function convertToInteractive(key) {\n",
" const element = document.querySelector('#df-94a74c8f-1062-4a80-8d3f-32f0fbadf7bb');\n",
" const dataTable =\n",
" await google.colab.kernel.invokeFunction('convertToInteractive',\n",
" [key], {});\n",
" if (!dataTable) return;\n",
"\n",
" const docLinkHtml = 'Like what you see? Visit the ' +\n",
" '<a target=\"_blank\" href=https://colab.research.google.com/notebooks/data_table.ipynb>data table notebook</a>'\n",
" + ' to learn more about interactive tables.';\n",
" element.innerHTML = '';\n",
" dataTable['output_type'] = 'display_data';\n",
" await google.colab.output.renderOutput(dataTable, element);\n",
" const docLink = document.createElement('div');\n",
" docLink.innerHTML = docLinkHtml;\n",
" element.appendChild(docLink);\n",
" }\n",
" </script>\n",
" </div>\n",
"\n",
"\n",
" <div id=\"df-54b2aa43-838b-47d3-9209-2fb18153cf87\">\n",
" <button class=\"colab-df-quickchart\" onclick=\"quickchart('df-54b2aa43-838b-47d3-9209-2fb18153cf87')\"\n",
" title=\"Suggest charts\"\n",
" style=\"display:none;\">\n",
"\n",
"<svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\"viewBox=\"0 0 24 24\"\n",
" width=\"24px\">\n",
" <g>\n",
" <path d=\"M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zM9 17H7v-7h2v7zm4 0h-2V7h2v10zm4 0h-2v-4h2v4z\"/>\n",
" </g>\n",
"</svg>\n",
" </button>\n",
"\n",
"<style>\n",
" .colab-df-quickchart {\n",
" --bg-color: #E8F0FE;\n",
" --fill-color: #1967D2;\n",
" --hover-bg-color: #E2EBFA;\n",
" --hover-fill-color: #174EA6;\n",
" --disabled-fill-color: #AAA;\n",
" --disabled-bg-color: #DDD;\n",
" }\n",
"\n",
" [theme=dark] .colab-df-quickchart {\n",
" --bg-color: #3B4455;\n",
" --fill-color: #D2E3FC;\n",
" --hover-bg-color: #434B5C;\n",
" --hover-fill-color: #FFFFFF;\n",
" --disabled-bg-color: #3B4455;\n",
" --disabled-fill-color: #666;\n",
" }\n",
"\n",
" .colab-df-quickchart {\n",
" background-color: var(--bg-color);\n",
" border: none;\n",
" border-radius: 50%;\n",
" cursor: pointer;\n",
" display: none;\n",
" fill: var(--fill-color);\n",
" height: 32px;\n",
" padding: 0;\n",
" width: 32px;\n",
" }\n",
"\n",
" .colab-df-quickchart:hover {\n",
" background-color: var(--hover-bg-color);\n",
" box-shadow: 0 1px 2px rgba(60, 64, 67, 0.3), 0 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
" fill: var(--button-hover-fill-color);\n",
" }\n",
"\n",
" .colab-df-quickchart-complete:disabled,\n",
" .colab-df-quickchart-complete:disabled:hover {\n",
" background-color: var(--disabled-bg-color);\n",
" fill: var(--disabled-fill-color);\n",
" box-shadow: none;\n",
" }\n",
"\n",
" .colab-df-spinner {\n",
" border: 2px solid var(--fill-color);\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" animation:\n",
" spin 1s steps(1) infinite;\n",
" }\n",
"\n",
" @keyframes spin {\n",
" 0% {\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" border-left-color: var(--fill-color);\n",
" }\n",
" 20% {\n",
" border-color: transparent;\n",
" border-left-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" }\n",
" 30% {\n",
" border-color: transparent;\n",
" border-left-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" border-right-color: var(--fill-color);\n",
" }\n",
" 40% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" border-top-color: var(--fill-color);\n",
" }\n",
" 60% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" }\n",
" 80% {\n",
" border-color: transparent;\n",
" border-right-color: var(--fill-color);\n",
" border-bottom-color: var(--fill-color);\n",
" }\n",
" 90% {\n",
" border-color: transparent;\n",
" border-bottom-color: var(--fill-color);\n",
" }\n",
" }\n",
"</style>\n",
"\n",
" <script>\n",
" async function quickchart(key) {\n",
" const quickchartButtonEl =\n",
" document.querySelector('#' + key + ' button');\n",
" quickchartButtonEl.disabled = true; // To prevent multiple clicks.\n",
" quickchartButtonEl.classList.add('colab-df-spinner');\n",
" try {\n",
" const charts = await google.colab.kernel.invokeFunction(\n",
" 'suggestCharts', [key], {});\n",
" } catch (error) {\n",
" console.error('Error during call to suggestCharts:', error);\n",
" }\n",
" quickchartButtonEl.classList.remove('colab-df-spinner');\n",
" quickchartButtonEl.classList.add('colab-df-quickchart-complete');\n",
" }\n",
" (() => {\n",
" let quickchartButtonEl =\n",
" document.querySelector('#df-54b2aa43-838b-47d3-9209-2fb18153cf87 button');\n",
" quickchartButtonEl.style.display =\n",
" google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
" })();\n",
" </script>\n",
" </div>\n",
"\n",
" </div>\n",
" </div>\n"
],
"text/plain": [
" Rainfall Americas Asia Europe Africa\n",
"0 (inches) 2010 \n",
"1 Average 104 201.0 193.0 144.0\n",
"2 24 hour high 15 26.0 27.0 18.0\n",
"3 12 hour high 9 10.0 11.0 12.0\n",
"4 2009 \n",
"5 Average 133 244.0 155.0 166.0\n",
"6 24 hour high 27 28.0 29.0 20.0\n",
"7 12 hour high 11 12.0 13.0 16.0"
]
},
"metadata": {},
"output_type": "display_data"
}
],
"source": [
"import pandas as pd\n",
"from IPython.display import display\n",
"\n",
"df = pd.read_csv(\n",
" \"/content/tables/table_2025_16_07_16_30_01_569.csv\",\n",
")\n",
"display(df.fillna(\"\"))"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}

Before

Width:  |  Height:  |  Size: 195 KiB

After

Width:  |  Height:  |  Size: 195 KiB

Before

Width:  |  Height:  |  Size: 363 KiB

After

Width:  |  Height:  |  Size: 363 KiB

Before

Width:  |  Height:  |  Size: 343 KiB

After

Width:  |  Height:  |  Size: 343 KiB

Before

Width:  |  Height:  |  Size: 185 KiB

After

Width:  |  Height:  |  Size: 185 KiB

Before

Width:  |  Height:  |  Size: 254 KiB

After

Width:  |  Height:  |  Size: 254 KiB

Before

Width:  |  Height:  |  Size: 650 KiB

After

Width:  |  Height:  |  Size: 650 KiB

Before

Width:  |  Height:  |  Size: 72 KiB

After

Width:  |  Height:  |  Size: 72 KiB

Before

Width:  |  Height:  |  Size: 88 KiB

After

Width:  |  Height:  |  Size: 88 KiB

Before

Width:  |  Height:  |  Size: 200 KiB

After

Width:  |  Height:  |  Size: 200 KiB

Before

Width:  |  Height:  |  Size: 115 KiB

After

Width:  |  Height:  |  Size: 115 KiB

Before

Width:  |  Height:  |  Size: 334 KiB

After

Width:  |  Height:  |  Size: 334 KiB

Before

Width:  |  Height:  |  Size: 202 KiB

After

Width:  |  Height:  |  Size: 202 KiB

Before

Width:  |  Height:  |  Size: 1.2 MiB

After

Width:  |  Height:  |  Size: 1.2 MiB

Before

Width:  |  Height:  |  Size: 170 KiB

After

Width:  |  Height:  |  Size: 170 KiB

Before

Width:  |  Height:  |  Size: 580 KiB

After

Width:  |  Height:  |  Size: 580 KiB

Before

Width:  |  Height:  |  Size: 271 KiB

After

Width:  |  Height:  |  Size: 271 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 1.5 MiB

After

Width:  |  Height:  |  Size: 1.5 MiB

Before

Width:  |  Height:  |  Size: 350 KiB

After

Width:  |  Height:  |  Size: 350 KiB

Before

Width:  |  Height:  |  Size: 47 KiB

After

Width:  |  Height:  |  Size: 47 KiB

Before

Width:  |  Height:  |  Size: 344 KiB

After

Width:  |  Height:  |  Size: 344 KiB

Some files were not shown because too many files have changed in this diff Show More