Compare commits

..

6 Commits

Author SHA1 Message Date
Logan Markewich b706363f11 vbump 2024-09-10 11:45:07 -06:00
Logan Markewich f603f507b6 linting 2024-09-10 11:32:50 -06:00
Pierre-Loic Doulcet 8aa08c590c Update llama_parse/base.py
Co-authored-by: Logan <logan.markewich@live.com>
2024-09-10 10:30:53 -07:00
Pierre-Loic Doulcet 42a98f921a Update llama_parse/base.py
Co-authored-by: Logan <logan.markewich@live.com>
2024-09-10 10:30:48 -07:00
Pierre-Loic Doulcet 1c453397de trailing whitespaces 2024-09-10 10:30:28 -07:00
Pierre-Loic Doulcet 3207daa7bd do not attach a filepath when a stram of bytes is passed 2024-09-10 10:20:57 -07:00
276 changed files with 11458 additions and 145520 deletions
+10 -4
View File
@@ -7,6 +7,8 @@ assignees: ''
---
_Note: we're aware of some missing content in the output and layout issues on tables. Please refrain from opening new issues on this topic unless if you think it's different from what has already been reported._
**Describe the bug**
Write a concise description of what the bug is.
@@ -17,15 +19,19 @@ If possible, please provide the PDF file causing the issue.
If you have it, please provide the ID of the job you ran.
You can find it here: https://cloud.llamaindex.ai/parse in the "History" tab.
**Screenshots**
Feel free to also provide screenshots if relevant.
**Client:**
Please remove untested options:
- Python Library
- API
- Frontend (cloud.llamaindex.ai)
- Python Library
- Typescript Library
- Notebook
- API
**Options**
What options did you use? Multimodal, fast mode, parsing instructions, etc.
**Additional context**
Add any additional context about the problem here.
What options did you use? Premium mode, multimodal, fast mode, parsing instructions, etc.
Screenshots, code snippets, etc.
-11
View File
@@ -1,11 +0,0 @@
# Please see the documentation for all configuration options:
# https://docs.github.com/github/administering-a-repository/configuration-options-for-dependency-updates
# and
# https://docs.github.com/code-security/dependabot/dependabot-version-updates/configuration-options-for-the-dependabot.yml-file
version: 2
updates:
- package-ecosystem: "github-actions"
directory: "/"
schedule:
interval: "weekly"
+48
View File
@@ -0,0 +1,48 @@
name: Build Package
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install
- name: Ensure lock works
shell: bash
run: poetry lock
- name: Build
shell: bash
run: poetry build
- name: Test installing built package
shell: bash
run: python -m pip install .
- name: Test import
shell: bash
working-directory: ${{ vars.RUNNER_TEMP }}
run: python -c "import llama_parse"
-50
View File
@@ -1,50 +0,0 @@
name: Build Package - Python
# Build package on its own without additional pip install
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ${{ matrix.os }}
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Test import
working-directory: py
run: uv run -- python -c "import llama_cloud_services"
-28
View File
@@ -1,28 +0,0 @@
name: Build Package - TypeScript
on: [pull_request]
jobs:
pre_release:
name: Pre Release
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Build
working-directory: ts/llama_cloud_services/
run: pnpm run build
+48 -8
View File
@@ -1,3 +1,14 @@
# For most projects, this workflow file will not need changing; you simply need
# to commit it to your repository.
#
# You may wish to alter this file to override the set of languages analyzed,
# or to provide custom queries or build logic.
#
# ******** NOTE ********
# We have attempted to detect the languages in your repository. Please check
# the `language` matrix defined below to confirm you have the correct set of
# supported CodeQL languages.
#
name: "CodeQL"
on:
@@ -17,25 +28,54 @@ jobs:
# - https://gh.io/supported-runners-and-hardware-resources
# - https://gh.io/using-larger-runners
# Consider using larger runners for possible analysis time improvements.
runs-on: "ubuntu-latest"
timeout-minutes: 360
runs-on: ${{ (matrix.language == 'swift' && 'macos-latest') || 'ubuntu-latest' }}
timeout-minutes: ${{ (matrix.language == 'swift' && 120) || 360 }}
permissions:
actions: read
contents: read
security-events: write
strategy:
fail-fast: false
matrix:
language: ["python"]
# CodeQL supports [ 'cpp', 'csharp', 'go', 'java', 'javascript', 'python', 'ruby', 'swift' ]
# Use only 'java' to analyze code written in Java, Kotlin or both
# Use only 'javascript' to analyze code written in JavaScript, TypeScript or both
# Learn more about CodeQL language support at https://aka.ms/codeql-docs/language-support
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@v3
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v3
uses: github/codeql-action/init@v2
with:
languages: python
dependency-caching: true
languages: ${{ matrix.language }}
# If you wish to specify custom queries, you can do so here or in a config file.
# By default, queries listed here will override any specified in a config file.
# Prefix the list here with "+" to use these queries and those in the config file.
# For more details on CodeQL's query packs, refer to: https://docs.github.com/en/code-security/code-scanning/automatically-scanning-your-code-for-vulnerabilities-and-errors/configuring-code-scanning#using-queries-in-ql-packs
# queries: security-extended,security-and-quality
# Autobuild attempts to build any compiled languages (C/C++, C#, Go, Java, or Swift).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@v2
# ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
# If the Autobuild fails above, remove it and uncomment the following three lines.
# modify them (or add more) to build your code if your project, please refer to the EXAMPLE below for guidance.
# - run: |
# echo "Run, Build Application using script"
# ./location_of_script_within_repo/buildscript.sh
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v3
uses: github/codeql-action/analyze@v2
with:
category: "/language:python"
category: "/language:${{matrix.language}}"
+37
View File
@@ -0,0 +1,37 @@
name: Linting
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install pre-commit
shell: bash
run: poetry run pip install pre-commit
- name: Run linter
shell: bash
run: poetry run make lint
-35
View File
@@ -1,35 +0,0 @@
name: Lint - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
jobs:
build:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }}
- name: Run linter
shell: bash
working-directory: py
run: uv run -- pre-commit run -a
-36
View File
@@ -1,36 +0,0 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run lint
working-directory: ts/llama_cloud_services/
run: pnpm run lint
- name: Run Prettier
working-directory: ts/llama_cloud_services/
run: pnpm run format
+64
View File
@@ -0,0 +1,64 @@
name: Publish llama-parse to PyPI / GitHub
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
POETRY_VERSION: "1.6.1"
PYTHON_VERSION: "3.9"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_parse'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up python ${{ env.PYTHON_VERSION }}
uses: actions/setup-python@v4
with:
python-version: ${{ env.PYTHON_VERSION }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: pip install -e .
- name: Build and publish to pypi
uses: JRubics/poetry-publish@v1.17
with:
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
ignore_dev_requirements: "yes"
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }}
draft: false
prerelease: false
- name: Get Asset name
run: |
export PKG=$(ls dist/ | grep tar)
set -- $PKG
echo "name=$1" >> $GITHUB_ENV
- name: Upload Release Asset (sdist) to GitHub
id: upload-release-asset
uses: actions/upload-release-asset@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
upload_url: ${{ steps.create_release.outputs.upload_url }}
asset_path: dist/${{ env.name }}
asset_name: ${{ env.name }}
asset_content_type: application/zip
-66
View File
@@ -1,66 +0,0 @@
name: Publish Release - Python
on:
push:
tags:
- "v*"
workflow_dispatch:
env:
UV_VERSION: "0.7.20"
jobs:
build-n-publish:
name: Build and publish to PyPI
if: github.repository == 'run-llama/llama_cloud_services'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install
- name: Display Python version
run: python --version
- name: Build
working-directory: py
run: uv build
- name: Test installing built package
shell: bash
working-directory: py
run: |
uv venv
uv pip install dist/*.whl
- name: Publish package
shell: bash
working-directory: py
run: uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Build and publish llama-parse
working-directory: py/llama_parse/
run: |
uv build
uv publish --token ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
- name: Create GitHub Release
id: create_release
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # This token is provided by Actions, you do not need to create your own token
with:
tag_name: ${{ github.ref }}
release_name: ${{ github.ref }} - LlamaCloud Services PY
artifacts: "py/**/dist/*"
generateReleaseNotes: true
draft: false
prerelease: false
-51
View File
@@ -1,51 +0,0 @@
name: Publish Release - TypeScript
on:
push:
tags:
- "llama-cloud-services@*"
jobs:
build-and-publish:
runs-on: ubuntu-latest
steps:
- name: Checkout Repo
uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services
run: pnpm install --no-frozen-lockfile
- name: Build tarball
run: |
pnpm pack
working-directory: ts/llama_cloud_services
- name: Setup npm authentication
run: echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc
env:
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Release
working-directory: ts/llama_cloud_services
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
run: pnpm publish --access public --no-git-checks
- name: Create release
uses: ncipollo/release-action@v1
with:
artifacts: "ts/llama_cloud_services/llama-cloud-services*.tgz"
name: Release ${{ github.ref }} - LlamaCloud Services TS
bodyFile: "ts/llama_cloud_services/CHANGELOG.md"
token: ${{ secrets.GITHUB_TOKEN }}
-39
View File
@@ -1,39 +0,0 @@
name: Test - Python
on:
push:
branches:
- main
pull_request:
env:
UV_VERSION: "0.7.20"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
version: ${{ env.UV_VERSION }}
- name: Set up Python
run: uv python install ${{ matrix.python-version }} && uv python pin ${{ matrix.python-version }}
- name: Run Tests
working-directory: py
run: uv run -- pytest tests/**/test_*.py
- name: Remove virtual environment
working-directory: py
run: rm -rf .venv/
-34
View File
@@ -1,34 +0,0 @@
name: Lint - TypeScript
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
TURBO_TOKEN: ${{ secrets.TURBO_TOKEN }}
TURBO_TEAM: ${{ vars.TURBO_TEAM }}
TURBO_REMOTE_ONLY: true
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
version: 10
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
working-directory: ts/llama_cloud_services/
run: pnpm install --no-frozen-lockfile
- name: Run tests
working-directory: ts/llama_cloud_services/
run: pnpm test --run
+40
View File
@@ -0,0 +1,40 @@
name: Unit Testing
on:
push:
branches:
- main
pull_request:
env:
POETRY_VERSION: "1.6.1"
LLAMA_CLOUD_API_KEY: ${{ secrets.LLAMA_CLOUD_API_KEY }}
jobs:
test:
runs-on: ubuntu-latest
strategy:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.8", "3.10", "3.11"]
steps:
- uses: actions/checkout@v3
with:
fetch-depth: 0
- name: Set up python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install Poetry
uses: snok/install-poetry@v1
with:
version: ${{ env.POETRY_VERSION }}
- name: Install deps
shell: bash
run: poetry install --with dev
- name: Run testing
env:
CI: true
shell: bash
run: poetry run pytest tests
-6
View File
@@ -3,9 +3,3 @@ __pycache__/
*.pyc
.DS_Store
.idea
.env*
.ipynb_checkpoints*
*_cache/
node_modules/
.turbo/
dist/
+6 -7
View File
@@ -21,19 +21,18 @@ repos:
hooks:
- id: ruff
args: [--fix, --exit-non-zero-on-fix]
exclude: ".*uv.lock"
exclude: ".*poetry.lock"
- repo: https://github.com/psf/black-pre-commit-mirror
rev: 23.10.1
hooks:
- id: black-jupyter
name: black-src
alias: black
exclude: ".*uv.lock"
exclude: ".*poetry.lock"
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.0.1
hooks:
- id: mypy
exclude: ^py/tests/
additional_dependencies:
[
"types-requests",
@@ -47,7 +46,7 @@ repos:
[
--disallow-untyped-defs,
--ignore-missing-imports,
--python-version=3.10,
--python-version=3.8,
]
- repo: https://github.com/adamchainz/blacken-docs
rev: 1.16.0
@@ -63,13 +62,13 @@ repos:
rev: v3.0.3
hooks:
- id: prettier
exclude: uv.lock
exclude: poetry.lock
- repo: https://github.com/codespell-project/codespell
rev: v2.2.6
hooks:
- id: codespell
additional_dependencies: [tomli]
exclude: ^(uv.lock|examples|ts)
exclude: ^(poetry.lock|examples)
args:
[
"--ignore-words-list",
@@ -84,6 +83,6 @@ repos:
rev: v0.23.1
hooks:
- id: toml-sort-fix
exclude: ".*uv.lock"
exclude: ".*poetry.lock"
exclude: .github/ISSUE_TEMPLATE
-33
View File
@@ -1,33 +0,0 @@
# Python
## Installation
This project uses uv. Create a virtual environment, and run `uv sync`
## Versioning (Maintainers only)
Before merging your changes, make sure to bump the versions.
Make a version bump to `pyproject.toml`. If the underlying dependency on the llamacloud platform OpenAPI
sdk needs bumping, make sure to bring that in as well. If updating dependencies, run `uv lock`.
The legacy `llama_parse` package re-exports some of `llama_cloud_services` in the old namespace. The
versions need to be kept consistent to sidecar it with `llama_cloud_services`. Bump it's version in `llama_parse/pyproject.toml`, and also bump it's dependency version of `llama-cloud-services` to match.
**Note**: Don't worry about updating the `llama_parse/poetry.lock` file when bumping versions. The GitHub action will automatically run `poetry lock` for the llama_parse package during the build process (though it doesn't commit the updated lockfile back to the repo).
You can also do this with `./scripts/version-bump.py set 0.x.x` if you have `uv` installed.
Once the change is merged, push a tag `git tag -a v0.x.x -m 0.x.x` and `git push origin 0.x.x`.
This tagging step can be done with `./scripts/version-bump tag`.
# Typescript
## Installation
...
## Versioning
...
View File
+111 -52
View File
@@ -1,81 +1,138 @@
[![PyPI - Downloads](https://img.shields.io/pypi/dm/llama-cloud-services)](https://pypi.org/project/llama-cloud-services/)
[![GitHub contributors](https://img.shields.io/github/contributors/run-llama/llama_cloud_services)](https://github.com/run-llama/llama_cloud_services/graphs/contributors)
# LlamaParse
[![PyPI - Downloads](https://img.shields.io/pypi/dm/llama-parse)](https://pypi.org/project/llama-parse/)
[![GitHub contributors](https://img.shields.io/github/contributors/run-llama/llama_parse)](https://github.com/run-llama/llama_parse/graphs/contributors)
[![Discord](https://img.shields.io/discord/1059199217496772688)](https://discord.gg/dGcwcsnxhU)
# Llama Cloud Services
LlamaParse is a **GenAI-native document parser** that can parse complex document data for any downstream LLM use case (RAG, agents).
This repository contains the code for hand-written SDKs and clients for interacting with LlamaCloud.
It is really good at the following:
This includes:
-**Broad file type support**: Parsing a variety of unstructured file types (.pdf, .pptx, .docx, .xlsx, .html) with text, tables, visual elements, weird layouts, and more.
-**Table recognition**: Parsing embedded tables accurately into text and semi-structured representations.
-**Multimodal parsing and chunking**: Extracting visual elements (images/diagrams) into structured formats and return image chunks using the latest multimodal models.
-**Custom parsing**: Input custom prompt instructions to customize the output the way you want it.
- [LlamaParse](./parse.md) - A GenAI-native document parser that can parse complex document data for any downstream LLM use case (Agents, RAG, data processing, etc.).
- [LlamaReport (beta/invite-only)](./report.md) - A prebuilt agentic report builder that can be used to build reports from a variety of data sources.
- [LlamaExtract](./extract.md) - A prebuilt agentic data extractor that can be used to transform data into a structured JSON representation.
- [LlamaCloud Index](./index.md) - A widely customizable and fully automated document ingestion pipeline that also serves retrieval purposes.
LlamaParse directly integrates with [LlamaIndex](https://github.com/run-llama/llama_index).
The free plan is up to 1000 pages a day. Paid plan is free 7k pages per week + 0.3c per additional page by default. There is a sandbox available to test the API [**https://cloud.llamaindex.ai/parse ↗**](https://cloud.llamaindex.ai/parse).
Read below for some quickstart information, or see the [full documentation](https://docs.cloud.llamaindex.ai/).
If you're a company interested in enterprise RAG solutions, and/or high volume/on-prem usage of LlamaParse, come [talk to us](https://www.llamaindex.ai/contact).
## Getting Started
Install the package:
First, login and get an api-key from [**https://cloud.llamaindex.ai/api-key ↗**](https://cloud.llamaindex.ai/api-key).
```bash
pip install llama-cloud-services
Then, make sure you have the latest LlamaIndex version installed.
**NOTE:** If you are upgrading from v0.9.X, we recommend following our [migration guide](https://pretty-sodium-5e0.notion.site/v0-10-0-Migration-Guide-6ede431dcb8841b09ea171e7f133bd77), as well as uninstalling your previous version first.
```
pip uninstall llama-index # run this if upgrading from v0.9.x or older
pip install -U llama-index --upgrade --no-cache-dir --force-reinstall
```
Then, get your API key from [LlamaCloud](https://cloud.llamaindex.ai/).
Lastly, install the package:
Then, you can use the services in your code:
`pip install llama-parse`
Now you can run the following to parse your first PDF file:
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
LlamaCloudIndex,
import nest_asyncio
nest_asyncio.apply()
from llama_parse import LlamaParse
parser = LlamaParse(
api_key="llx-...", # can also be set in your env as LLAMA_CLOUD_API_KEY
result_type="markdown", # "markdown" and "text" are available
num_workers=4, # if multiple files passed, split in `num_workers` API calls
verbose=True,
language="en", # Optionally you can define a language, default=en
)
parser = LlamaParse(api_key="YOUR_API_KEY")
report = LlamaReport(api_key="YOUR_API_KEY")
extract = LlamaExtract(api_key="YOUR_API_KEY")
index = LlamaCloudIndex(
"my_first_index", project_name="default", api_key="YOUR_API_KEY"
)
# sync
documents = parser.load_data("./my_file.pdf")
# sync batch
documents = parser.load_data(["./my_file1.pdf", "./my_file2.pdf"])
# async
documents = await parser.aload_data("./my_file.pdf")
# async batch
documents = await parser.aload_data(["./my_file1.pdf", "./my_file2.pdf"])
```
See the quickstart guides for each service for more information:
## Using with file object
- [LlamaParse](./parse.md)
- [LlamaReport (beta/invite-only)](./report.md)
- [LlamaExtract](./extract.md)
- [LlamaCloud Index](./index.md)
## Switch to EU SaaS 🇪🇺
If you are interested in using LlamaCloud services in the EU, you can adjust your base URL to `https://api.cloud.eu.llamaindex.ai`.
You can also create your API key in the EU region [here](https://cloud.eu.llamaindex.ai).
You can parse a file object directly:
```python
from llama_cloud_services import (
LlamaParse,
LlamaReport,
LlamaExtract,
EU_BASE_URL,
import nest_asyncio
nest_asyncio.apply()
from llama_parse import LlamaParse
parser = LlamaParse(
api_key="llx-...", # can also be set in your env as LLAMA_CLOUD_API_KEY
result_type="markdown", # "markdown" and "text" are available
num_workers=4, # if multiple files passed, split in `num_workers` API calls
verbose=True,
language="en", # Optionally you can define a language, default=en
)
parser = LlamaParse(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
report = LlamaReport(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
extract = LlamaExtract(api_key="YOUR_API_KEY", base_url=EU_BASE_URL)
index = LlamaCloudIndex(
"my_first_index",
project_name="default",
api_key="YOUR_API_KEY",
base_url=EU_BASE_URL,
)
with open("./my_file1.pdf", "rb") as f:
documents = parser.load_data(f)
# you can also pass file bytes directly
with open("./my_file1.pdf", "rb") as f:
file_bytes = f.read()
documents = parser.load_data(file_bytes)
```
## Using with `SimpleDirectoryReader`
You can also integrate the parser as the default PDF loader in `SimpleDirectoryReader`:
```python
import nest_asyncio
nest_asyncio.apply()
from llama_parse import LlamaParse
from llama_index.core import SimpleDirectoryReader
parser = LlamaParse(
api_key="llx-...", # can also be set in your env as LLAMA_CLOUD_API_KEY
result_type="markdown", # "markdown" and "text" are available
verbose=True,
)
file_extractor = {".pdf": parser}
documents = SimpleDirectoryReader(
"./data", file_extractor=file_extractor
).load_data()
```
Full documentation for `SimpleDirectoryReader` can be found on the [LlamaIndex Documentation](https://docs.llamaindex.ai/en/stable/module_guides/loading/simpledirectoryreader.html).
## Examples
Several end-to-end indexing examples can be found in the examples folder
- [Getting Started](examples/demo_basic.ipynb)
- [Advanced RAG Example](examples/demo_advanced.ipynb)
- [Raw API Usage](examples/demo_api.ipynb)
## Documentation
You can see complete SDK and API documentation for each service on [our official docs](https://docs.cloud.llamaindex.ai/).
[https://docs.cloud.llamaindex.ai/](https://docs.cloud.llamaindex.ai/)
## Terms of Service
@@ -83,4 +140,6 @@ See the [Terms of Service Here](./TOS.pdf).
## Get in Touch (LlamaCloud)
You can get in touch with us by following our [contact link](https://www.llamaindex.ai/contact).
LlamaParse is part of LlamaCloud, our e2e enterprise RAG platform that provides out-of-the-box, production-ready connectors, indexing, and retrieval over your complex data sources. We offer SaaS and VPC options.
LlamaCloud is currently available via waitlist (join by [creating an account](https://cloud.llamaindex.ai/)). If you're interested in state-of-the-art quality and in centralizing your RAG efforts, come [get in touch with us](https://www.llamaindex.ai/contact).
-11
View File
@@ -1,11 +0,0 @@
# LlamaCloud Services Examples
In this folder you will find several python notebooks and two end-to-end typescript applications that contain examples regarding:
- [LlamaParse - Python](./parse/)
- [LlamaParse - TypeScript](./parse-ts/)
- [LlamaExtract - Python](./extract/)
- [LlamaReport - Python](./report/)
- [LlamaCloud Index - TypeScript](./index-ts/)
Follow the instructions of each notebook/application to get started!
@@ -22,7 +22,7 @@
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-cloud-services llama-index llama-index-postprocessor-sbert-rerank"
"!pip install llama-parse llama-index llama-index-postprocessor-sbert-rerank"
]
},
{
@@ -82,7 +82,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -7,7 +7,7 @@
"source": [
"# RAG over the Caltrain Weekend Schedule \n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/caltrain/caltrain_text_mode.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/caltrain/caltrain_text_mode.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This example shows off LlamaParse parsing capabilities to build a functioning query pipeline over the Caltrain weekend schedule, a big timetable containing all trains northbound and southbound and their stops in various cities.\n",
"\n",
@@ -81,7 +81,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"docs = LlamaParse(result_type=\"text\").load_data(\"./caltrain_schedule_weekend.pdf\")"
]
+759
View File
@@ -0,0 +1,759 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Advanced RAG with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook is a complete walkthrough for using LlamaParse with advanced indexing/retrieval techniques in LlamaIndex over the Apple 10K Filing. \n",
"\n",
"This allows us to ask sophisticated questions that aren't possible with \"naive\" parsing/indexing techniques with existing models.\n",
"\n",
"Note for this example, we are using the `llama_index >=0.10.4` version"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-index\n",
"!pip install llama-index-core==0.10.6.post1\n",
"!pip install llama-index-embeddings-openai\n",
"!pip install llama-index-postprocessor-flag-embedding-reranker\n",
"!pip install git+https://github.com/FlagOpen/FlagEmbedding.git\n",
"!pip install llama-parse"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2021/q4/_10-K-2021-(As-Filed).pdf\" -O apple_2021_10k.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Some OpenAI and LlamaParse details"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\"\n",
"\n",
"# Using OpenAI API for embeddings/llms\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex\n",
"from llama_index.core import Settings\n",
"\n",
"embed_model = OpenAIEmbedding(model=\"text-embedding-3-small\")\n",
"llm = OpenAI(model=\"gpt-3.5-turbo-0125\")\n",
"\n",
"Settings.llm = llm\n",
"Settings.embed_model = embed_model"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using brand new `LlamaParse` PDF reader for PDF Parsing\n",
"\n",
"we also compare two different retrieval/query engine strategies:\n",
"1. Using raw Markdown text as nodes for building index and apply simple query engine for generating the results;\n",
"2. Using `MarkdownElementNodeParser` for parsing the `LlamaParse` output Markdown results and building recursive retriever query engine for generation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id cac11eca-71db-4dab-b72b-c67d31e551f3\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"markdown\").load_data(\"./apple_2021_10k.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from copy import deepcopy\n",
"from llama_index.core.schema import TextNode\n",
"from llama_index.core import VectorStoreIndex\n",
"\n",
"\n",
"def get_page_nodes(docs, separator=\"\\n---\\n\"):\n",
" \"\"\"Split each document into page node, by separator.\"\"\"\n",
" nodes = []\n",
" for doc in docs:\n",
" doc_chunks = doc.text.split(separator)\n",
" for doc_chunk in doc_chunks:\n",
" node = TextNode(\n",
" text=doc_chunk,\n",
" metadata=deepcopy(doc.metadata),\n",
" )\n",
" nodes.append(node)\n",
"\n",
" return nodes"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"page_nodes = get_page_nodes(documents)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.node_parser import MarkdownElementNodeParser\n",
"\n",
"node_parser = MarkdownElementNodeParser(\n",
" llm=OpenAI(model=\"gpt-3.5-turbo-0125\"), num_workers=8\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"nodes = node_parser.get_nodes_from_documents(documents)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"base_nodes, objects = node_parser.get_nodes_and_objects(nodes)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"\"This table provides information about a company's state of incorporation or organization and its corresponding I.R.S. Employer Identification Number.,\\nwith the following table title:\\nCompany Incorporation Information,\\nwith the following columns:\\n- California: None\\n- 94-2404110: None\\n\""
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"objects[0].get_content()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# dump both indexed tables and page text into the vector index\n",
"recursive_index = VectorStoreIndex(nodes=base_nodes + objects + page_nodes)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Apple Inc.\n",
"\n",
"**CONSOLIDATED STATEMENTS OF OPERATIONS (In millions, except number of shares which are reflected in thousands and per share amounts)**\n",
"| |September 25, 2021|September 26, 2020|September 28, 2019|\n",
"|---|---|---|---|\n",
"|Net sales:|$297,392|$220,747|$213,883|\n",
"|Products| | | |\n",
"|Services|$68,425|$53,768|$46,291|\n",
"|Total net sales|$365,817|$274,515|$260,174|\n",
"|Cost of sales:| | | |\n",
"|Products|$192,266|$151,286|$144,996|\n",
"|Services|$20,715|$18,273|$16,786|\n",
"|Total cost of sales|$212,981|$169,559|$161,782|\n",
"|Gross margin|$152,836|$104,956|$98,392|\n",
"|Operating expenses:| | | |\n",
"|Research and development|$21,914|$18,752|$16,217|\n",
"|Selling, general and administrative|$21,973|$19,916|$18,245|\n",
"|Total operating expenses|$43,887|$38,668|$34,462|\n",
"|Operating income|$108,949|$66,288|$63,930|\n",
"|Other income/(expense), net|$258|$803|$1,807|\n",
"|Income before provision for income taxes|$109,207|$67,091|$65,737|\n",
"|Provision for income taxes|$14,527|$9,680|$10,481|\n",
"|Net income|$94,680|$57,411|$55,256|\n",
"|Earnings per share:| | | |\n",
"|Basic|$5.67|$3.31|$2.99|\n",
"|Diluted|$5.61|$3.28|$2.97|\n",
"|Shares used in computing earnings per share:| | | |\n",
"|Basic|16,701,272|17,352,119|18,471,336|\n",
"|Diluted|16,864,919|17,528,214|18,595,651|\n",
"\n",
"See accompanying Notes to Consolidated Financial Statements.\n"
]
}
],
"source": [
"print(page_nodes[31].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.postprocessor.flag_embedding_reranker import FlagEmbeddingReranker\n",
"\n",
"reranker = FlagEmbeddingReranker(\n",
" top_n=5,\n",
" model=\"BAAI/bge-reranker-large\",\n",
")\n",
"\n",
"recursive_query_engine = recursive_index.as_query_engine(\n",
" similarity_top_k=5, node_postprocessors=[reranker], verbose=True\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"233\n"
]
}
],
"source": [
"print(len(nodes))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup Baseline\n",
"\n",
"For comparison, we setup a naive RAG pipeline with default parsing and standard chunking, indexing, retrieval."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import SimpleDirectoryReader\n",
"\n",
"reader = SimpleDirectoryReader(input_files=[\"apple_2021_10k.pdf\"])\n",
"base_docs = reader.load_data()\n",
"raw_index = VectorStoreIndex.from_documents(base_docs)\n",
"raw_query_engine = raw_index.as_query_engine(\n",
" similarity_top_k=5, node_postprocessors=[reranker]\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using `new LlamaParse` as pdf data parsing methods and retrieve tables with two different methods\n",
"we compare base query engine vs recursive query engine with tables"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Table Query Task: Queries for Table Question Answering"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"The purchases of marketable securities in 2020 amounted to $163.4 billion.\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering 59368b87-e602-4bd1-88a7-7526fd6ab83f: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Purchases of marketable securities in 2020\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering dfd97f47-eb4d-4bab-8a22-9bbbc0096a4b: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Purchases of marketable securities in 2020\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"$114,938\n"
]
}
],
"source": [
"query = \"Purchases of marketable securities in 2020\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"This table provides information on hedged assets and liabilities for the years 2021 and 2020, including current and non-current marketable securities and term debt.,\n",
"with the following table title:\n",
"Hedged Assets and Liabilities Summary,\n",
"with the following columns:\n",
"- 2021: None\n",
"- 2020: None\n",
"\n",
"| |2021|2020|\n",
"|---|---|---|\n",
"|Hedged assets/(liabilities):| | |\n",
"|Current and non-current marketable securities|$15,954|$16,270|\n",
"|Current and non-current term debt|$(17,857)|$(21,033)|\n",
"\n"
]
}
],
"source": [
"print(response_2.source_nodes[2].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"0.03%, 0.75%, 1.43%\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering a5afa785-217f-4e72-87cf-15da11632ec0: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query effective interest rates of all debt issuances in 2021\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"0.48% 0.63%, 0.03% 4.78%, 0.75% 2.81%, 1.43% 2.86%\n"
]
}
],
"source": [
"query = \"effective interest rates of all debt issuances in 2021\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Term Debt\n",
"As of September 25, 2021 , the Company had outstanding floating- and fixed-rate notes with varying maturities for an aggregate \n",
"principal amount of $118.1 billion (collectively the “Notes”). The Notes are senior unsecured obligations and interest is payable in \n",
"arrears. The following table provides a summary of the Companys term debt as of September 25, 2021 and September 26, \n",
"2020 :\n",
"Maturities\n",
"(calendar year)2021 2020\n",
"Amount\n",
"(in millions)Effective\n",
"Interest RateAmount\n",
"(in millions)Effective\n",
"Interest Rate\n",
"2013 2020 debt issuances:\n",
"Floating-rate notes 2022 $ 1,750 0.48% 0.63% $ 2,250 0.60% 1.39%\n",
"Fixed-rate 0.000% 4.650% notes 2022 2060 95,813 0.03% 4.78% 103,828 0.03% 4.78%\n",
"Second quarter 2021 debt issuance:\n",
"Fixed-rate 0.700% 2.800% notes 2026 2061 14,000 0.75% 2.81% — — %\n",
"Fourth quarter 2021 debt issuance:\n",
"Fixed-rate 1.400% 2.850% notes 2028 2061 6,500 1.43% 2.86% — — %\n",
"Total term debt 118,063 106,078 \n",
"Unamortized premium/(discount) and issuance \n",
"costs, net (380) (314) \n",
"Hedge accounting fair value adjustments 1,036 1,676 \n",
"Less: Current portion of term debt (9,613) (8,773) \n",
"Total non-current portion of term debt $ 109,106 $ 98,667 \n",
"To manage interest rate risk on certain of its U.S. dollardenominated fixed- or floating-rate notes, the Company has entered into \n",
"interest rate swaps to effectively convert the fixed interest rates to floating interest rates or the floating interest rates to fixed \n",
"interest rates on a portion of these notes. Additionally, to manage foreign currency risk on certain of its foreign currency\n",
"denominated notes, the Company has entered into foreign currency swaps to effectively convert these notes to U.S. dollar\n",
"denominated notes.\n",
"The effective interest rates for the Notes include the interest on the Notes, amortization of the discount or premium and, if \n",
"applicable, adjustments related to hedging. The Company recogni zed $2.6 billion , $2.8 billion and $3.2 billion of interest expense \n",
"on its term debt for 2021 , 2020 and 2019 , respectively.\n",
"The future principal payments for the Companys Notes as of September 25, 2021 , are as follows (in millions):\n",
"2022 $ 9,583 \n",
"2023 11,391 \n",
"2024 10,202 \n",
"2025 10,914 \n",
"2026 11,408 \n",
"Thereafter 64,565 \n",
"Total term debt $ 118,063 \n",
"As of September 25, 2021 and September 26, 2020 , the fair value of the Companys Notes, based on Level 2 inputs, was $125.3 \n",
"billion and $117.1 billion , respectively.\n",
"Apple Inc. | 2021 Form 10-K | 45\n"
]
}
],
"source": [
"print(response_1.source_nodes[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"The U.S. Tax Cuts and Jobs Act of 2017 had an impact on income taxes in 2020, as evidenced by a decrease in the provision for income taxes compared to the prior year.\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering b9416f35-ebf1-45d6-9a29-b59e435ab42d: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Impacts of the U.S. Tax Cuts and Jobs Act of 2017 on income taxes in 2020\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 8d8d5733-ff30-4535-9376-7f761b5900ea: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Impacts of the U.S. Tax Cuts and Jobs Act of 2017 on income taxes in 2020\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 82f301e5-199a-4aa2-bbdf-ef97898c0326: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Impacts of the U.S. Tax Cuts and Jobs Act of 2017 on income taxes in 2020\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 86f666b4-254b-487f-9870-8ee09aef07a9: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query Impacts of the U.S. Tax Cuts and Jobs Act of 2017 on income taxes in 2020\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"The U.S. Tax Cuts and Jobs Act of 2017 had a negative impact on income taxes in 2020.\n"
]
}
],
"source": [
"query = \"Impacts of the U.S. Tax Cuts and Jobs Act of 2017 on income taxes in 2020\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Other Income/(Expense), Net\n",
"The following table shows the detail of OI&E for 2021 , 2020 and 2019 (in millions):\n",
"2021 2020 2019\n",
"Interest and dividend income $ 2,843 $ 3,763 $ 4,961 \n",
"Interest expense (2,645) (2,873) (3,576) \n",
"Other income/(expense), net 60 (87) 422 \n",
"Total other income/(expense), net $ 258 $ 803 $ 1,807 \n",
"Note 5 Income Taxe s\n",
"Provision for Income Taxes and Effective Tax Rat e\n",
"The provision for income taxes for 2021 , 2020 and 2019 , consisted of the following (in millions):\n",
"2021 2020 2019\n",
"Federal:\n",
"Current $ 8,257 $ 6,306 $ 6,384 \n",
"Deferred (7,176) (3,619) (2,939) \n",
"Total 1,081 2,687 3,445 \n",
"State:\n",
"Current 1,620 455 475 \n",
"Deferred (338) 21 (67) \n",
"Total 1,282 476 408 \n",
"Foreign:\n",
"Current 9,424 3,134 3,962 \n",
"Deferred 2,740 3,383 2,666 \n",
"Total 12,164 6,517 6,628 \n",
"Provision for income taxes $ 14,527 $ 9,680 $ 10,481 \n",
"The foreign provision for income taxes is based on foreign pretax earnings of $68.7 billion , $38.1 billion and $44.3 billion in 2021 , \n",
"2020 and 2019 , respectively.\n",
"A reconciliation of the provision for income taxes, with the amount computed by applying the statutory federal income tax rate \n",
"(21% in 2021 , 2020 and 2019 ) to income before provision for income taxes for 2021 , 2020 and 2019 , is as follows (dollars in \n",
"millions):\n",
"2021 2020 2019\n",
"Computed expected tax $ 22,933 $ 14,089 $ 13,805 \n",
"State taxes, net of federal effect 1,151 423 423 \n",
"Impacts of the U.S. Tax Cuts and Jobs Act of 2017 — (582) — \n",
"Earnings of foreign subsidiaries (4,715) (2,534) (2,625) \n",
"Foreign-derived intangible income deduction (1,372) (169) (149) \n",
"Research and development credit, net (1,033) (728) (548) \n",
"Excess tax benefits from equity awards (2,137) (930) (639) \n",
"Other (300) 111 214 \n",
"Provision for income taxes $ 14,527 $ 9,680 $ 10,481 \n",
"Effective tax rate 13.3% 14.4% 15.9% \n",
"Apple Inc. | 2021 Form 10-K | 41\n"
]
}
],
"source": [
"print(response_1.source_nodes[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"$3,619 million in 2019, $7,176 million in 2020, and $1,081 million in 2021\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering 12b1355a-f9e6-4b08-a19a-3ffc00dc5b9f: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query federal deferred tax in 2019-2021\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 82f301e5-199a-4aa2-bbdf-ef97898c0326: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query federal deferred tax in 2019-2021\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 8d8d5733-ff30-4535-9376-7f761b5900ea: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query federal deferred tax in 2019-2021\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"$2,939, $3,619, $7,176\n"
]
}
],
"source": [
"query = \"federal deferred tax in 2019-2021\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"State deferred income tax for 2019: $454 million\n",
"State deferred income tax for 2020: $21 million\n",
"State deferred income tax for 2021: -$338 million\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering 12b1355a-f9e6-4b08-a19a-3ffc00dc5b9f: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query give me the deferred state income tax in 2019-2021 (include +/-)\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 8d8d5733-ff30-4535-9376-7f761b5900ea: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query give me the deferred state income tax in 2019-2021 (include +/-)\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"Deferred state income tax for the years 2019-2021:\n",
"- 2019: ($67) million\n",
"- 2020: $21 million\n",
"- 2021: ($338) million\n"
]
}
],
"source": [
"query = \"give me the deferred state income tax in 2019-2021 (include +/-)\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Summary of income tax provisions for Federal, State, and Foreign entities over the years 2019, 2020, and 2021.,\n",
"with the following table title:\n",
"Income Tax Provisions by Entity and Year,\n",
"with the following columns:\n",
"- Entity: The type of entity (Federal, State, Foreign)\n",
"- 2019: Income tax provisions for the year 2019\n",
"- 2020: Income tax provisions for the year 2020\n",
"- 2021: Income tax provisions for the year 2021\n",
"\n",
"| |2021|2020|2019|\n",
"|---|---|---|---|\n",
"|Federal:| | | |\n",
"|Current|$8,257|$6,306|$6,384|\n",
"|Deferred|(7,176)|(3,619)|(2,939)|\n",
"|Total|1,081|2,687|3,445|\n",
"|State:| | | |\n",
"|Current|1,620|455|475|\n",
"|Deferred|(338)|21|(67)|\n",
"|Total|1,282|476|408|\n",
"|Foreign:| | | |\n",
"|Current|9,424|3,134|3,962|\n",
"|Deferred|2,740|3,383|2,666|\n",
"|Total|12,164|6,517|6,628|\n",
"|Provision for income taxes|$14,527|$9,680|$10,481|\n",
"\n"
]
}
],
"source": [
"print(response_2.source_nodes[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Basic Query Engine***********\n",
"$1,620 million in 2019, $455 million in 2020, $475 million in 2021\n",
"\u001b[1;3;38;2;11;159;203mRetrieval entering 82f301e5-199a-4aa2-bbdf-ef97898c0326: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query current state taxes per year in 2019-2021 (include +/-)\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering 8d8d5733-ff30-4535-9376-7f761b5900ea: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query current state taxes per year in 2019-2021 (include +/-)\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering b9416f35-ebf1-45d6-9a29-b59e435ab42d: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query current state taxes per year in 2019-2021 (include +/-)\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203mRetrieval entering a029e464-575f-4dd6-afad-7cc0bbc5dbf9: TextNode\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200mRetrieving from object TextNode with query current state taxes per year in 2019-2021 (include +/-)\n",
"\u001b[0m\n",
"***********New LlamaParse+ Recursive Retriever Query Engine***********\n",
"$475 in 2019, $455 in 2020, $1,620 in 2021.\n"
]
}
],
"source": [
"query = \"current state taxes per year in 2019-2021 (include +/-)\"\n",
"\n",
"response_1 = raw_query_engine.query(query)\n",
"print(\"\\n***********Basic Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = recursive_query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Recursive Retriever Query Engine***********\")\n",
"print(response_2)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+295
View File
@@ -0,0 +1,295 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Using llama-parse with AstraDB"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"In this notebook, we show a basic RAG-style example that uses `llama-parse` to parse a PDF document, store the corresponding document into a vector store (`AstraDB`) and finally, perform some basic queries against that store. The notebook is modeled after the quick start notebooks and hence is meant as a way of getting started with `llama-parse`, backed by a vector database."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Requirements"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# First, install the required dependencies\n",
"%pip install --quiet llama-index llama-parse llama-index-vector-stores-astra-db llama-index-llms-openai"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Configuration"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"import openai\n",
"\n",
"from getpass import getpass\n",
"\n",
"# Get all required API keys and parameters\n",
"llama_cloud_api_key = getpass(\"Enter your Llama Index Cloud API Key: \")\n",
"api_endpoint = input(\"Enter your Astra DB API Endpoint: \")\n",
"token = getpass(\"Enter your Astra DB Token: \")\n",
"namespace = (\n",
" input(\"Enter your Astra DB namespace (optional, must exist on Astra): \") or None\n",
")\n",
"openai_api_key = getpass(\"Enter your OpenAI API Key: \")\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = llama_cloud_api_key\n",
"openai.api_key = openai_api_key"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Using llama-parse to parse a PDF"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Download complete.\n"
]
}
],
"source": [
"# Grab a PDF from Arxiv for indexing\n",
"import requests\n",
"\n",
"# The URL of the file you want to download\n",
"url = \"https://arxiv.org/pdf/1706.03762.pdf\"\n",
"# The local path where you want to save the file\n",
"file_path = \"./attention.pdf\"\n",
"\n",
"# Perform the HTTP request\n",
"response = requests.get(url)\n",
"\n",
"# Check if the request was successful\n",
"if response.status_code == 200:\n",
" # Open the file in binary write mode and save the content\n",
" with open(file_path, \"wb\") as file:\n",
" file.write(response.content)\n",
" print(\"Download complete.\")\n",
"else:\n",
" print(\"Error downloading the file.\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id ce3909a7-54cf-438b-849a-fe9a903b0c71\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'rmer - model architecture.\\nThe Transformer follows this overall architecture using stacked self-attention and point-wise, fully\\nconnected layers for both the encoder and decoder, shown in the left and right halves of Figure 1,\\nrespectively.\\n3.1 Encoder and Decoder Stacks\\nEncoder: The encoder is composed of a stack of N = 6 identical layers. Each layer has two\\nsub-layers. The first is a multi-head self-attention mechanism, and the second is a simple, position-\\nwise fully connected feed-forward network. We employ a residual connection [11] around each of\\nthe two sub-layers, followed by layer normalization [1]. That is, the output of each sub-layer is\\nLayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer\\nitself. To facilitate these residual connections, all sub-layers in the model, as well as the embedding\\nlayers, produce outputs of dimension dmodel = 512.\\nDecoder: The decoder is also composed of a stack of N = 6 identical layers. In addition '"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Take a quick look at some of the parsed text from the document:\n",
"documents[0].get_content()[10000:11000]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Storing into Astra DB"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.vector_stores.astra_db import AstraDBVectorStore\n",
"\n",
"astra_db_store = AstraDBVectorStore(\n",
" token=token,\n",
" api_endpoint=api_endpoint,\n",
" namespace=namespace,\n",
" collection_name=\"astra_v_table_llamaparse\",\n",
" embedding_dimension=1536,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.node_parser import SimpleNodeParser\n",
"\n",
"node_parser = SimpleNodeParser()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"\n",
"storage_context = StorageContext.from_defaults(vector_store=astra_db_store)\n",
"\n",
"index = VectorStoreIndex(\n",
" nodes=nodes,\n",
" storage_context=storage_context,\n",
" embed_model=OpenAIEmbedding(api_key=openai_api_key),\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Simple RAG Example"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"query_engine = index.as_query_engine(similarity_top_k=15)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********New LlamaParse+ Basic Query Engine***********\n",
"Multi-Head Attention is also known as multi-headed self-attention.\n"
]
}
],
"source": [
"query = \"What is Multi-Head Attention also known as?\"\n",
"\n",
"response_1 = query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Basic Query Engine***********\")\n",
"print(response_1)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'We used beam search as described in the previous section, but no\\ncheckpoint averaging. We present these results in Table 3.\\nIn Table 3 rows (A), we vary the number of attention heads and the attention key and value dimensions,\\nkeeping the amount of computation constant, as described in Section 3.2.2. While single-head\\nattention is 0.9 BLEU worse than the best setting, quality also drops off with too many heads.\\nIn Table 3 rows (B), we observe that reducing the attention key size dk hurts model quality. This\\nsuggests that determining compatibility is not easy and that a more sophisticated compatibility\\nfunction than dot product may be beneficial. We further observe in rows (C) and (D) that, as expected,\\nbigger models are better, and dropout is very helpful in avoiding over-fitting. In row (E) we replace our\\nsinusoidal positional encoding with learned positional embeddings [9], and observe nearly identical\\nresults to the base model.\\n6.3 English Constituency Parsing\\nTo evaluate if the Transformer can generalize to other tasks we performed experiments on English\\nconstituency parsing. This task presents specific challenges: the output is subject to strong structural\\nconstraints and is significantly longer than the input. Furthermore, RNN sequence-to-sequence\\nmodels have not been able to attain state-of-the-art results in small-data regimes [37].\\nWe trained a 4-layer transformer with dmodel = 1024 on the Wall Street Journal (WSJ) portion of the\\nPenn Treebank [25], about 40K training sentences. We also trained it in a semi-supervised setting,\\nusing the larger high-confidence and BerkleyParser corpora from with approximately 17M sentences\\n[37]. We used a vocabulary of 16K tokens for the WSJ only setting and a vocabulary of 32K tokens\\nfor the semi-supervised setting.\\nWe performed only a small number of experiments to select the dropout, both attention and residual\\n(section 5.4), learning rates and beam size on the Section 22 development set, all other parameters\\nremained unchanged from the English-to-German base translation model. During inference, we\\n 9\\n---\\nTable 4: The Transformer generalizes well to English constituency parsing (Results are on Section 23\\nof WSJ)\\n Parser Training WSJ 23 F1\\n Vinyals & Kaiser el al. (2014) [37] WSJ only, discriminative 88.3\\n Petrov et al. (2006) [29] WSJ only, discriminative 90.4\\n Zhu et al. (2013) [40] WSJ only, discriminative 90.4\\n Dyer et al. (2016) [8] WSJ only, discriminative 91.7\\n Transformer (4 layers) WSJ only, discriminative 91.3\\n Zhu et al. (2013) [40] semi-supervised 91.3\\n Huang & Harper (2009) [14] semi-supervised 91.3\\n McClosky et al. (2006) [26] semi-supervised 92.1\\n Vinyals & Kaiser el al. (2014) [37] semi-supervised 92.1\\n Transformer (4 layers) semi-supervised 92.7\\n Luong et al. (2015) [23] multi-task 93.0\\n Dyer et al. (2016) [8] generative 93.3\\nincreased the maximum output length to input length + 300. We used a beam size of 21 and α = 0.3\\nfor both WSJ only and the semi-supervised setting.\\nOur results in Table 4 show that despite the lack of task-specific tuning our model performs sur-\\nprisingly well, yielding better results than all previously reported models with the exception of the\\nRecurrent Neural Network Grammar [8].\\nIn contrast to RNN sequence-to-sequence models [37], the Transformer outperforms the Berkeley-\\nParser [29] even when training only on the WSJ training set of 40K sentences.\\n7 Conclusion\\nIn this work, we presented the Transformer, the first sequence transduction model based entirely on\\nattention, replacing the recurrent layers most commonly used in encoder-decoder architectures with\\nmulti-headed self-attention.\\nFor translation tasks, the Transformer can be trained significantly faster than architectures based\\non recurrent or convolutional layers.'"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Take a look at one of the source nodes from the response\n",
"response_1.source_nodes[0].get_content()"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
+183
View File
@@ -0,0 +1,183 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse Usage"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-parse"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-02-02 11:10:10-- https://arxiv.org/pdf/1706.03762.pdf\n",
"Resolving arxiv.org (arxiv.org)... 151.101.131.42, 151.101.3.42, 151.101.67.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.131.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 2215244 (2.1M) [application/pdf]\n",
"Saving to: ./attention.pdf\n",
"\n",
"./attention.pdf 100%[===================>] 2.11M --.-KB/s in 0.08s \n",
"\n",
"2024-02-02 11:10:10 (25.9 MB/s) - ./attention.pdf saved [2215244/2215244]\n",
"\n"
]
}
],
"source": [
"!wget \"https://arxiv.org/pdf/1706.03762.pdf\" -O \"./attention.pdf\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id dd0b8e31-0c09-4497-b78a-cc1c92f1d6cf\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"text\").load_data(\"./attention.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"ad\n",
"relying entirely on an attention mechanism to draw global dependencies between input and output.\n",
"The Transformer allows for significantly more parallelization and can reach a new state of the art in\n",
"translation quality after being trained for as little as twelve hours on eight P100 GPUs.\n",
"2 Background\n",
"The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU\n",
"[16], ByteNet [18] and ConvS2S [9], all of which use convolutional neural networks as basic building\n",
"block, computing hidden representations in parallel for all input and output positions. In these models,\n",
"the number of operations required to relate signals from two arbitrary input or output positions grows\n",
"in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes\n",
"it more difficult to learn dependencies between distant positions [12]. In the Transformer this is\n",
"reduced to a constant number of operations, albeit at the cost of reduced effective res\n"
]
}
],
"source": [
"print(documents[0].text[6000:7000])"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id d4531453-1bbb-48c4-8324-ae9fea9f2fa2\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"markdown\").load_data(\"./attention.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"ction describes the training regime for our models.\n",
"\n",
"##### Training Data and Batching\n",
"\n",
"We trained on the standard WMT 2014 English-German dataset consisting of about 4.5 million\n",
"sentence pairs. Sentences were encoded using byte-pair encoding [3], which has a shared source-\n",
"target vocabulary of about 37000 tokens. For English-French, we used the significantly larger WMT\n",
"2014 English-French dataset consisting of 36M sentences and split tokens into a 32000 word-piece\n",
"vocabulary [38]. Sentence pairs were batched together by approximate sequence length. Each training\n",
"batch contained a set of sentence pairs containing approximately 25000 source tokens and 25000\n",
"target tokens.\n",
"\n",
"##### Hardware and Schedule\n",
"\n",
"We trained our models on one machine with 8 NVIDIA P100 GPUs. For our base models using\n",
"the hyperparameters described throughout the paper, each training step took about 0.4 seconds. We\n",
"trained the base models for a total of 100,000 steps or 12 hours. For our big models,(described on the\n",
"bo...\n"
]
}
],
"source": [
"print(documents[0].text[20000:21000] + \"...\")"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
File diff suppressed because one or more lines are too long
@@ -6,7 +6,7 @@
"source": [
"# RAG with Excel Spreadsheet using LlamaPrase\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you using LlamaParse with Excel Spreadsheet.\n",
"\n",
@@ -21,7 +21,7 @@
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-cloud-services"
"%pip install llama-parse"
]
},
{
@@ -41,7 +41,7 @@
"\n",
"nest_asyncio.apply()\n",
"\n",
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"api_key = \"llx-\" # get from cloud.llamaindex.ai"
]
@@ -6,7 +6,7 @@
"source": [
"# LlamaParse - Fast checking Insurance Contract for Coverage\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook we will look at how LlamaParse can be used to extract structured coverage information from an insurance policy."
]
@@ -116,7 +116,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"documents = LlamaParse(result_type=\"markdown\").load_data(\"./policy.pdf\")"
]
@@ -7,7 +7,7 @@
"source": [
"# LlamaParse JSON Mode + Multimodal RAG\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you how to use LlamaParse JSON mode with LlamaIndex to build a simple multimodal RAG pipeline.\n",
"\n",
@@ -31,11 +31,11 @@
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-index-core\n",
"%pip install llama-index-llms-anthropic\n",
"%pip install llama-index-embeddings-huggingface\n",
"%pip install llama-cloud-services"
"!pip install llama-index\n",
"!pip install llama-index-core\n",
"!pip install llama-index-llms-anthropic llama-index-multi-modal-llms-anthropic\n",
"!pip install llama-index-embeddings-huggingface\n",
"!pip install llama-parse"
]
},
{
@@ -45,6 +45,11 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
@@ -63,7 +68,7 @@
"source": [
"from llama_index.llms.anthropic import Anthropic\n",
"\n",
"llm = Anthropic(model=\"claude-3-5-sonnet-20241022\")"
"llm = Anthropic(model=\"claude-3-opus-20240229\", temperature=0.0)"
]
},
{
@@ -124,10 +129,30 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(take_screenshot=True)\n",
"result = await parser.aparse(\"./uber_10q_march_2022.pdf\")"
"parser = LlamaParse(verbose=True)\n",
"json_objs = parser.get_json_result(\"./uber_10q_march_2022.pdf\")\n",
"json_list = json_objs[0][\"pages\"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b26d21d1-05b5-4f49-b937-c13106a84015",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"text\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
@@ -137,12 +162,7 @@
"metadata": {},
"outputs": [],
"source": [
"text_nodes = await result.aget_text_nodes(split_by_page=True)\n",
"image_nodes = await result.aget_image_nodes(\n",
" include_screenshot_images=True,\n",
" include_object_images=True,\n",
" image_download_dir=\"./uber_10q_images\",\n",
")"
"text_nodes = get_text_nodes(json_list)"
]
},
{
@@ -152,7 +172,7 @@
"source": [
"## Extract/Index images from image dicts\n",
"\n",
"Here we use a multimodal model to caption images and create text nodes for indexing."
"Here we use a multimodal model to extract and index images from image dictionaries."
]
},
{
@@ -170,32 +190,27 @@
}
],
"source": [
"!mkdir -p llama2_images\n",
"# call get_images on parser, convert to ImageDocuments\n",
"!mkdir llama2_images\n",
"\n",
"from llama_index.core.llms import ChatMessage, ImageBlock, TextBlock\n",
"from llama_index.core.schema import ImageNode, TextNode\n",
"from llama_index.llms.anthropic import Anthropic\n",
"from llama_index.core.schema import ImageDocument\n",
"from llama_index.multi_modal_llms.anthropic import AnthropicMultiModal\n",
"\n",
"\n",
"def get_image_text_nodes(image_nodes: list[ImageNode]):\n",
"def get_image_text_nodes(json_objs: List[dict]):\n",
" \"\"\"Extract out text from images using a multimodal model.\"\"\"\n",
" llm = Anthropic(model=\"claude-3-5-haiku-20241022\", max_tokens=300)\n",
" anthropic_mm_llm = AnthropicMultiModal(max_tokens=300)\n",
" image_dicts = parser.get_images(json_objs, download_path=\"llama2_images\")\n",
" image_documents = []\n",
" img_text_nodes = []\n",
" for image_node in image_nodes:\n",
" image_path = image_node.image_path\n",
" message = ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" TextBlock(text=\"Describe the images as alt text\"),\n",
" ImageBlock(path=image_path),\n",
" ],\n",
" )\n",
" response = llm.chat([message])\n",
" text_node = TextNode(\n",
" text=str(response.message.content), metadata={\"path\": image_path}\n",
" for image_dict in image_dicts:\n",
" image_doc = ImageDocument(image_path=image_dict[\"path\"])\n",
" response = anthropic_mm_llm.complete(\n",
" prompt=\"Describe the images as alt text\",\n",
" image_documents=[image_doc],\n",
" )\n",
" text_node = TextNode(text=str(response), metadata={\"path\": image_dict[\"path\"]})\n",
" img_text_nodes.append(text_node)\n",
"\n",
" return img_text_nodes"
]
},
@@ -206,7 +221,7 @@
"metadata": {},
"outputs": [],
"source": [
"image_text_nodes = get_image_text_nodes(image_nodes)"
"image_text_nodes = get_image_text_nodes(json_objs)"
]
},
{
@@ -327,7 +342,7 @@
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"display_name": "llama-parse-aNC435Vv-py3.10",
"language": "python",
"name": "python3"
},
File diff suppressed because one or more lines are too long
@@ -9,7 +9,7 @@
"\n",
"LlamaParse supports users to specify a `language` parameter before uploading documents, giving users better OCR capabilities over non-English PDFs, parsing images into more accurate representations.\n",
"\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_cloud_services/blob/main/llama_parse/base.py.\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_parse/blob/main/llama_parse/base.py.\n",
"\n",
"This notebook shows a demo of this in action. "
]
@@ -31,9 +31,14 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
"# os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
@@ -72,11 +77,10 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(language=\"fr\")\n",
"result = await parser.aparse(\"./treasury_report.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
"parser = LlamaParse(result_type=\"text\", language=\"fr\")\n",
"documents = parser.load_data(\"./treasury_report.pdf\")"
]
},
{
@@ -246,11 +250,10 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(language=\"ch_sim\")\n",
"result = await parser.aparse(\"./chinese_pdf.pdf\")\n",
"documents = result.get_text_documents(split_by_page=False)"
"parser = LlamaParse(result_type=\"text\", language=\"ch_sim\")\n",
"documents = parser.load_data(\"./chinese_pdf.pdf\")"
]
},
{
@@ -401,11 +404,10 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"base_parser = LlamaParse(language=\"en\")\n",
"result = await base_parser.aparse(\"./chinese_pdf2.pdf\")\n",
"base_documents = result.get_text_documents(split_by_page=False)"
"base_parser = LlamaParse(result_type=\"text\", language=\"en\")\n",
"base_documents = parser.load_data(\"./chinese_pdf2.pdf\")"
]
},
{
@@ -7,7 +7,7 @@
"source": [
"# LlamaParse With MongoDB\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_mongodb.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_mongodb.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook, we provide a straightforward example of using LlamaParse with MongoDB Atlas VectorSearch.\n",
"\n",
@@ -60,14 +60,19 @@
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import requests\n",
"import pymongo\n",
"\n",
"from llama_index.vector_stores.mongodb import MongoDBAtlasVectorSearch\n",
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"from llama_index.core.node_parser import SentenceSplitter"
"from llama_index.core.node_parser import SimpleNodeParser"
]
},
{
@@ -132,8 +137,7 @@
}
],
"source": [
"result = await LlamaParse().aparse(file_path)\n",
"documents = result.get_text_documents(split_by_page=False)"
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
]
},
{
@@ -199,7 +203,7 @@
"metadata": {},
"outputs": [],
"source": [
"node_parser = SentenceSplitter()\n",
"node_parser = SimpleNodeParser()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
+544
View File
@@ -0,0 +1,544 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse - Parsing comic books with parsing intructions\n",
"Parsing intructions allow you to instruct our parsing model the same way you would instruct an LLM!\n",
"\n",
"They can be useful to help the parser get better results on complex document layouts, to extract data in a specific format, or to transform the document in other ways.\n",
"\n",
"Using Parsing Instruction you will get better results out of LlamaParse on complicated documents, and also be able to simplify your application code."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Installation\n",
"\n",
"Parsing instructions are part of the llamaParse API. They can be accessed by directly specifying the parsing_instruction parameter in the API or by using the LlamaParse python module (which we will use for this tutorial).\n",
"\n",
"To install llama-parse, just get it from PIP:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Collecting llama-parse\n",
" Downloading llama_parse-0.3.8-py3-none-any.whl (6.7 kB)\n",
"Collecting llama-index-core>=0.10.7 (from llama-parse)\n",
" Downloading llama_index_core-0.10.19-py3-none-any.whl (15.3 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m15.3/15.3 MB\u001b[0m \u001b[31m31.9 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: PyYAML>=6.0.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (6.0.1)\n",
"Requirement already satisfied: SQLAlchemy[asyncio]>=1.4.49 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.0.28)\n",
"Requirement already satisfied: aiohttp<4.0.0,>=3.8.6 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.9.3)\n",
"Collecting dataclasses-json (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dataclasses_json-0.6.4-py3-none-any.whl (28 kB)\n",
"Collecting deprecated>=1.2.9.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading Deprecated-1.2.14-py2.py3-none-any.whl (9.6 kB)\n",
"Collecting dirtyjson<2.0.0,>=1.0.8 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dirtyjson-1.0.8-py3-none-any.whl (25 kB)\n",
"Requirement already satisfied: fsspec>=2023.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2023.6.0)\n",
"Collecting httpx (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpx-0.27.0-py3-none-any.whl (75 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m75.6/75.6 kB\u001b[0m \u001b[31m6.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hCollecting llamaindex-py-client<0.2.0,>=0.1.13 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading llamaindex_py_client-0.1.13-py3-none-any.whl (107 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m108.0/108.0 kB\u001b[0m \u001b[31m10.0 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: nest-asyncio<2.0.0,>=1.5.8 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.6.0)\n",
"Requirement already satisfied: networkx>=3.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.2.1)\n",
"Requirement already satisfied: nltk<4.0.0,>=3.8.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.8.1)\n",
"Requirement already satisfied: numpy in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.25.2)\n",
"Collecting openai>=1.1.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading openai-1.13.3-py3-none-any.whl (227 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m227.4/227.4 kB\u001b[0m \u001b[31m16.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: pandas in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.5.3)\n",
"Requirement already satisfied: pillow>=9.0.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (9.4.0)\n",
"Requirement already satisfied: requests>=2.31.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.31.0)\n",
"Requirement already satisfied: tenacity<9.0.0,>=8.2.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (8.2.3)\n",
"Collecting tiktoken>=0.3.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading tiktoken-0.6.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.8 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m1.8/1.8 MB\u001b[0m \u001b[31m43.1 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: tqdm<5.0.0,>=4.66.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.66.2)\n",
"Requirement already satisfied: typing-extensions>=4.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.10.0)\n",
"Collecting typing-inspect>=0.8.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading typing_inspect-0.9.0-py3-none-any.whl (8.8 kB)\n",
"Requirement already satisfied: aiosignal>=1.1.2 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Requirement already satisfied: attrs>=17.3.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (23.2.0)\n",
"Requirement already satisfied: frozenlist>=1.1.1 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.4.1)\n",
"Requirement already satisfied: multidict<7.0,>=4.5 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (6.0.5)\n",
"Requirement already satisfied: yarl<2.0,>=1.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.9.4)\n",
"Requirement already satisfied: async-timeout<5.0,>=4.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (4.0.3)\n",
"Requirement already satisfied: wrapt<2,>=1.10 in /usr/local/lib/python3.10/dist-packages (from deprecated>=1.2.9.3->llama-index-core>=0.10.7->llama-parse) (1.14.1)\n",
"Requirement already satisfied: pydantic>=1.10 in /usr/local/lib/python3.10/dist-packages (from llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.6.3)\n",
"Requirement already satisfied: anyio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.7.1)\n",
"Requirement already satisfied: certifi in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (2024.2.2)\n",
"Collecting httpcore==1.* (from httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpcore-1.0.4-py3-none-any.whl (77 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m77.8/77.8 kB\u001b[0m \u001b[31m8.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: idna in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.6)\n",
"Requirement already satisfied: sniffio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Collecting h11<0.15,>=0.13 (from httpcore==1.*->httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading h11-0.14.0-py3-none-any.whl (58 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m58.3/58.3 kB\u001b[0m \u001b[31m5.7 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: click in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (8.1.7)\n",
"Requirement already satisfied: joblib in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (1.3.2)\n",
"Requirement already satisfied: regex>=2021.8.3 in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (2023.12.25)\n",
"Requirement already satisfied: distro<2,>=1.7.0 in /usr/lib/python3/dist-packages (from openai>=1.1.0->llama-index-core>=0.10.7->llama-parse) (1.7.0)\n",
"Requirement already satisfied: charset-normalizer<4,>=2 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (3.3.2)\n",
"Requirement already satisfied: urllib3<3,>=1.21.1 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (2.0.7)\n",
"Requirement already satisfied: greenlet!=0.4.17 in /usr/local/lib/python3.10/dist-packages (from SQLAlchemy[asyncio]>=1.4.49->llama-index-core>=0.10.7->llama-parse) (3.0.3)\n",
"Collecting mypy-extensions>=0.3.0 (from typing-inspect>=0.8.0->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading mypy_extensions-1.0.0-py3-none-any.whl (4.7 kB)\n",
"Collecting marshmallow<4.0.0,>=3.18.0 (from dataclasses-json->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading marshmallow-3.21.1-py3-none-any.whl (49 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m49.4/49.4 kB\u001b[0m \u001b[31m4.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: python-dateutil>=2.8.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2.8.2)\n",
"Requirement already satisfied: pytz>=2020.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2023.4)\n",
"Requirement already satisfied: exceptiongroup in /usr/local/lib/python3.10/dist-packages (from anyio->httpx->llama-index-core>=0.10.7->llama-parse) (1.2.0)\n",
"Requirement already satisfied: packaging>=17.0 in /usr/local/lib/python3.10/dist-packages (from marshmallow<4.0.0,>=3.18.0->dataclasses-json->llama-index-core>=0.10.7->llama-parse) (23.2)\n",
"Requirement already satisfied: annotated-types>=0.4.0 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (0.6.0)\n",
"Requirement already satisfied: pydantic-core==2.16.3 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.16.3)\n",
"Requirement already satisfied: six>=1.5 in /usr/local/lib/python3.10/dist-packages (from python-dateutil>=2.8.1->pandas->llama-index-core>=0.10.7->llama-parse) (1.16.0)\n",
"Installing collected packages: dirtyjson, mypy-extensions, marshmallow, h11, deprecated, typing-inspect, tiktoken, httpcore, httpx, dataclasses-json, openai, llamaindex-py-client, llama-index-core, llama-parse\n",
"Successfully installed dataclasses-json-0.6.4 deprecated-1.2.14 dirtyjson-1.0.8 h11-0.14.0 httpcore-1.0.4 httpx-0.27.0 llama-index-core-0.10.19 llama-parse-0.3.8 llamaindex-py-client-0.1.13 marshmallow-3.21.1 mypy-extensions-1.0.0 openai-1.13.3 tiktoken-0.6.0 typing-inspect-0.9.0\n"
]
}
],
"source": [
"%pip install llama-parse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## API key\n",
"\n",
"The use of LlamaParse requires an API key which you can get here: https://cloud.llamaindex.ai/parse"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Async (Notebook only)\n",
"llama-parse is async-first, so running the code in a notebook requires the use of nest_asyncio\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Import the package"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_parse import LlamaParse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using llamaparse for getting better results (on Manga!)\n",
"\n",
"Sometimes the layout of a page is unusual and you will get sub-optimal reading order results with LlamaParse. For example, when parsing manga you expect the reading order to be right to left even if the content is in English!"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let's download an extract of a great manga \"The manga guide to calculus\", by Hiroyuki Kojima (https://www.amazon.com/Manga-Guide-Calculus-Hiroyuki-Kojima/dp/1593271948)\n",
"\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Resolving drive.usercontent.google.com (drive.usercontent.google.com)... 173.194.211.132, 2607:f8b0:400c:c10::84\n",
"Connecting to drive.usercontent.google.com (drive.usercontent.google.com)|173.194.211.132|:443... connected.\n",
"HTTP request sent, awaiting response... 303 See Other\n",
"Location: https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download [following]\n",
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Reusing existing connection to drive.usercontent.google.com:443.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 3041634 (2.9M) [application/octet-stream]\n",
"Saving to: ./manga.pdf\n",
"\n",
"./manga.pdf 100%[===================>] 2.90M --.-KB/s in 0.04s \n",
"\n",
"2024-03-13 13:57:20 (78.6 MB/s) - ./manga.pdf saved [3041634/3041634]\n",
"\n"
]
}
],
"source": [
"! wget \"https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\" -O ./manga.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Without parsing instructions\n",
"For the sake of comparison, let's first parse without any instructions."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 25bf4202-78d8-4705-88cf-c616ae7c82af\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"As you can see below, LlamaParse is not doing a great job here. It is interpreting the grid of comic panels as a table, and trying to fit the dialogue into a table. It's very hard to follow."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"---\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = \n"
]
}
],
"source": [
"print(vanilaParsing[0].text[100:1000])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Using parsing instructions\n",
"Let's try to parse the manga with custom instructions:\n",
"\n",
"\"The provided document is a manga comic book. Most pages do NOT have a title. It does not contain tables. Try to reconstruct the dialogue spoken in a cohesive way.\"\n",
"\n",
"To do so just pass the parsing instruction as a parameter to LlamaParse:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 88ab273e-b2a7-4f84-8e72-e9367cf6b114\n",
"."
]
}
],
"source": [
"parsingInstructionManga = \"\"\"The provided document is a manga comic book. Most pages do NOT have a title.\n",
"It does not contain tables.\n",
"Try to reconstruct the dialogue spoken in a cohesive way.\"\"\"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstructionManga\n",
").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let's see how it compare with page 3! We encourage you to play with the target page and explore other pages. As you will see, the parsing instruction allowed LlamaParse to make sense of the document!\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1M87rXTIZE8d5v7aHmVZVW6gW3eDGq6ks&authuser=0\" />\n",
"\n",
"\n",
"\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"\n",
"\n",
"------------------------------------------------------------\n",
"\n",
"\n",
"# The Asagake Times\n",
"\n",
"Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor?\n",
"\n",
"Do I have the wrong map?\n",
"\n",
"You're looking for the Sanda-cho branch office?\n",
"\n",
"It's next door.\n",
"\n",
"Everybody mistakes us for the office because we are larger.\n",
"\n",
"What Is a Function? 3\n"
]
}
],
"source": [
"target_page = 1\n",
"print(vanilaParsing[0].text.split(\"\\n---\\n\")[target_page])\n",
"print(\"\\n\\n------------------------------------------------------------\\n\\n\")\n",
"print(withInstructionParsing[0].text.split(\"\\n---\\n\")[target_page])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Math - doing more with parsing instuction!\n",
"\n",
"But this manga is about math and full of equations, why not ask the parser to output them in **LaTeX**?\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1tze3xcQ7axVA-vC_iZeAj_GvYcyNuYDa&authuser=0\" />"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 3a055e64-d91e-484e-b9b0-99a2e637c08d\n",
"."
]
}
],
"source": [
"parsingInstructionMangaLatex = \"\"\"The provided document is a manga comic book. Most pages do NOT have a title.\n",
"It does not contain tables.\n",
"Try to reconstruct the dialogue spoken in a cohesive way.\n",
"Output any math equation in LATEX markdown (between $$)\"\"\"\n",
"withLatex = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstructionMangaLatex\n",
").load_data(\"./manga.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"\n",
"[Without instruction]------------------------------------------------------------\n",
"\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = lim ε→0 (α(a + ε) + β - (αa + β)) = lim ε→0 α = α|\n",
"| |Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.|\n",
"\n",
"|3.|Lets find the derivative of f(x) = x^2, which appeared in the story. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 ((a + ε)^2 - a^2) / ε = lim (a^2 + 2aε + ε^2 - a^2) / ε = lim (2aε + ε^2) = lim (2a + ε) = 2a|\n",
"| |Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.|\n",
"\n",
"## Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n",
"\n",
"## Chapter 1 Lets Differentiate a Function!\n",
"\n",
"\n",
"[With instruction to output math in LATEX!]------------------------------------------------------------\n",
"\n",
"\n",
"# Derivative of Constant, Linear, or Quadratic Function\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"1. Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha - \\alpha}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} 0 = 0 \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"2. Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(\\alpha + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha(a + \\varepsilon) + \\beta - (\\alpha a + \\beta)}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} \\alpha = \\alpha \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.\n",
"\n",
"3. Lets find the derivative of f(x) = x2. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\left( (a + \\varepsilon)^2 - a^2 \\right) = \\lim_{{\\varepsilon \\to 0}} 2a\\varepsilon + \\varepsilon = \\lim_{{\\varepsilon \\to 0}} (2a + \\varepsilon) = 2a \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.\n",
"\n",
"### Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n"
]
}
],
"source": [
"target_page = 2\n",
"print(\n",
" \"\\n\\n[Without instruction]------------------------------------------------------------\\n\\n\"\n",
")\n",
"print(vanilaParsing[0].text.split(\"\\n---\\n\")[target_page])\n",
"print(\n",
" \"\\n\\n[With instruction to output math in LATEX!]------------------------------------------------------------\\n\\n\"\n",
")\n",
"print(withLatex[0].text.split(\"\\n---\\n\")[target_page])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"And here is the result as rendered by https://upmath.me/ .\n",
"\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1qGo5bMGYOiIC9MnprcgEByaYjU9YII2Q&authuser=0\" />\n",
"\n",
"\n",
"Over this short notebook we saw how to use parsing instructions to increase the quality and accuracy of parsing with LLamaParse!"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
+367
View File
@@ -0,0 +1,367 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# RAG for Table Comparisons with LlamaParse + LlamaIndex\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_table_comparisons.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you how to do comparisons across both tabular and text data across multiple PDF documents.\n",
"\n",
"We load in multiple PDFs with embedded tables (2021 and 2020 10K filings for Apple) using LlamaParse, parse each into a hierarchy of tables/text objects, define a recursive retriever over each, and then compose both with a SubQuestionQueryEngine."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup\n",
"\n",
"Install core packages, download files, parse documents."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-index-core\n",
"%pip install llama-index-embeddings-openai\n",
"%pip install llama-index-question-gen-openai\n",
"%pip install llama-index-postprocessor-flag-embedding-reranker\n",
"%pip install git+https://github.com/FlagOpen/FlagEmbedding.git\n",
"%pip install llama-parse"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2020/ar/_10-K-2020-(As-Filed).pdf\" -O apple_2020_10k.pdf\n",
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2021/q4/_10-K-2021-(As-Filed).pdf\" -O apple_2021_10k.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Some OpenAI and LlamaParse details"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-\"\n",
"\n",
"# Using OpenAI API for embeddings/llms\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex\n",
"from llama_index.core import Settings\n",
"\n",
"embed_model = OpenAIEmbedding(model=\"text-embedding-3-small\")\n",
"llm = OpenAI(model=\"gpt-3.5-turbo-0125\")\n",
"\n",
"Settings.llm = llm\n",
"Settings.embed_model = embed_model"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using brand new `LlamaParse` PDF reader for PDF Parsing\n",
"\n",
"we also compare two different retrieval/query engine strategies:\n",
"1. Using raw Markdown text as nodes for building index and apply simple query engine for generating the results;\n",
"2. Using `MarkdownElementNodeParser` for parsing the `LlamaParse` output Markdown results and building recursive retriever query engine for generation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"docs_2021 = LlamaParse(result_type=\"markdown\").load_data(\"./apple_2021_10k.pdf\")\n",
"docs_2020 = LlamaParse(result_type=\"markdown\").load_data(\"./apple_2020_10k.pdf\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Create Recursive Retriever over each Document\n",
"\n",
"We define a function to get a recursive retriever from each document. The steps are the following:\n",
"- Hierarchically parse the document using our `MarkdownElementNodeParser`, which will embed/summarize embedded tables.\n",
"- Load into a vector store. Under the hood we will automatically store links between nodes (e.g. table summary to table text).\n",
"- Get a query engine over the vector store, which performs retrieval/synthesis. Under the hood we will automatically perform recursive retrieval if there are links."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.node_parser import MarkdownElementNodeParser\n",
"\n",
"node_parser = MarkdownElementNodeParser(\n",
" llm=OpenAI(model=\"gpt-3.5-turbo-0125\"), num_workers=8\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import pickle\n",
"from llama_index.postprocessor.flag_embedding_reranker import (\n",
" FlagEmbeddingReranker,\n",
")\n",
"\n",
"reranker = FlagEmbeddingReranker(\n",
" top_n=5,\n",
" model=\"BAAI/bge-reranker-large\",\n",
")\n",
"\n",
"\n",
"def create_query_engine_over_doc(docs, nodes_save_path=None):\n",
" \"\"\"Big function to go from document path -> recursive retriever.\"\"\"\n",
" if nodes_save_path is not None and os.path.exists(nodes_save_path):\n",
" raw_nodes = pickle.load(open(nodes_save_path, \"rb\"))\n",
" else:\n",
" raw_nodes = node_parser.get_nodes_from_documents(docs)\n",
" if nodes_save_path is not None:\n",
" pickle.dump(raw_nodes, open(nodes_save_path, \"wb\"))\n",
"\n",
" base_nodes, objects = node_parser.get_nodes_and_objects(raw_nodes)\n",
"\n",
" ### Construct Retrievers\n",
" # construct top-level vector index + query engine\n",
" vector_index = VectorStoreIndex(nodes=base_nodes + objects)\n",
" query_engine = vector_index.as_query_engine(\n",
" similarity_top_k=15, node_postprocessors=[reranker]\n",
" )\n",
" return query_engine, base_nodes"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"query_engine_2021, nodes_2021 = create_query_engine_over_doc(\n",
" docs_2021, nodes_save_path=\"2021_nodes.pkl\"\n",
")\n",
"query_engine_2020, nodes_2020 = create_query_engine_over_doc(\n",
" docs_2020, nodes_save_path=\"2020_nodes.pkl\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.tools import QueryEngineTool, ToolMetadata\n",
"from llama_index.core.query_engine import SubQuestionQueryEngine\n",
"\n",
"\n",
"# setup base query engine as tool\n",
"query_engine_tools = [\n",
" QueryEngineTool(\n",
" query_engine=query_engine_2021,\n",
" metadata=ToolMetadata(\n",
" name=\"apple_2021_10k\",\n",
" description=(\"Provides information about Apple financials for year 2021\"),\n",
" ),\n",
" ),\n",
" QueryEngineTool(\n",
" query_engine=query_engine_2020,\n",
" metadata=ToolMetadata(\n",
" name=\"apple_2020_10k\",\n",
" description=(\"Provides information about Apple financials for year 2020\"),\n",
" ),\n",
" ),\n",
"]\n",
"\n",
"sub_query_engine = SubQuestionQueryEngine.from_defaults(\n",
" query_engine_tools=query_engine_tools,\n",
" llm=llm,\n",
" use_async=True,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Try out Some Comparisons"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 4 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What are the deferred assets in 2021?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2021_10k] Q: What are the deferred liabilities in 2021?\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203m[apple_2020_10k] Q: What are the deferred assets in 2020?\n",
"\u001b[0m\u001b[1;3;38;2;155;135;227m[apple_2020_10k] Q: What are the deferred liabilities in 2020?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2021_10k] A: $7,200\n",
"\u001b[0m\u001b[1;3;38;2;155;135;227m[apple_2020_10k] A: $10,138\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: $25,176 million\n",
"\u001b[0m\u001b[1;3;38;2;11;159;203m[apple_2020_10k] A: $19,336\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the deferred assets and liabilities in 2021 with 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"In 2021, the deferred assets increased by $5,840 million compared to 2020, while the deferred liabilities decreased by $2,938 million in the same period.\n"
]
}
],
"source": [
"print(str(response))"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 2 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What is the total number of RSUs in Apple's 2021 financials?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] Q: What is the total number of RSUs in Apple's 2020 financials?\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: The total number of RSUs in Apple's 2021 financials is 240,427.\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] A: The total number of RSUs in Apple's 2020 financials is 310,778.\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the total number of RSUs in 2021 and 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Generated 2 sub questions.\n",
"\u001b[1;3;38;2;237;90;200m[apple_2021_10k] Q: What are the risk factors mentioned in the 2021 financial report of Apple?\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] Q: What are the risk factors mentioned in the 2020 financial report of Apple?\n",
"\u001b[0m\u001b[1;3;38;2;237;90;200m[apple_2021_10k] A: The risk factors mentioned in the 2021 financial report of Apple include risks related to COVID-19, macroeconomic and industry risks, political events, trade and international disputes, natural disasters, public health issues, industrial accidents, credit risk, fluctuations in foreign currency exchange rates, changes in tax rates and legislation, volatility in the price of the company's stock, and exposure to legal proceedings and claims.\n",
"\u001b[0m\u001b[1;3;38;2;90;149;237m[apple_2020_10k] A: The risk factors mentioned in the 2020 financial report of Apple include the impact of the COVID-19 pandemic on the company's business operations, financial condition, and stock price; global and regional economic conditions affecting demand for products and services; competition in global markets with rapid technological changes; potential disruptions in the supply chain due to industrial accidents or public health issues; information technology system failures or network disruptions affecting business operations; risks associated with confidential information security and potential unauthorized access; fluctuations in quarterly net sales and operating results due to various factors; stock price volatility impacting investor confidence and employee retention; financial performance risks related to changes in foreign currency exchange rates affecting sales and earnings.\n",
"\u001b[0m"
]
}
],
"source": [
"response = sub_query_engine.query(\n",
" \"Can you compare and contrast the risk factors in 2021 vs. 2020?\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The risk factors mentioned in the 2021 financial report of Apple include risks related to COVID-19, macroeconomic and industry risks, political events, trade and international disputes, natural disasters, public health issues, industrial accidents, credit risk, fluctuations in foreign currency exchange rates, changes in tax rates and legislation, volatility in the price of the company's stock, and exposure to legal proceedings and claims. In contrast, the risk factors mentioned in the 2020 financial report of Apple focused more on the impact of the COVID-19 pandemic on the company's business operations, financial condition, and stock price; global and regional economic conditions affecting demand for products and services; competition in global markets with rapid technological changes; potential disruptions in the supply chain due to industrial accidents or public health issues; information technology system failures or network disruptions affecting business operations; risks associated with confidential information security and potential unauthorized access; fluctuations in quarterly net sales and operating results due to various factors; stock price volatility impacting investor confidence and employee retention; financial performance risks related to changes in foreign currency exchange rates affecting sales and earnings.\n"
]
}
],
"source": [
"print(str(response))"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
@@ -7,7 +7,7 @@
"source": [
"# RAG with Excel Spreadsheet using LlamaPrase\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/excel/dcf_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook constructs a RAG pipeline over a simple DCF template [here](https://eqvista.com/app/uploads/2020/09/Eqvista_DCF-Excel-Template.xlsx).\n",
"\n"
@@ -31,7 +31,7 @@
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-cloud-services"
"%pip install llama-parse"
]
},
{
@@ -53,7 +53,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"# api_key = \"llx-\" # get from cloud.llamaindex.ai"
]
File diff suppressed because it is too large Load Diff
Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.3 MiB

File diff suppressed because one or more lines are too long
@@ -1,10 +0,0 @@
# Financial Modeling Assumptions
Discount Rate: 8%
Terminal Growth Rate: 2%
Tax Rate: 25%
Revenue Growth (Years 1-5): 10% per annum
Revenue Growth (Years 6-10): 5% per annum
Capital Expenditures as % of Revenue: 7%
Working Capital Assumption: 3% of Revenue
Depreciation Rate: 10% per annum
Cost of Capital Assumption: 8%
Binary file not shown.

Before

Width:  |  Height:  |  Size: 67 KiB

@@ -1 +0,0 @@
sec_form_4_dump.json
File diff suppressed because it is too large Load Diff
Binary file not shown.

Before

Width:  |  Height:  |  Size: 202 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 440 KiB

Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 156 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 85 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 893 KiB

@@ -1,440 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Extract Data from Financial Reports - with Citations and Reasoning\n",
"\n",
"Given complex files like financial reports, contracts, invoices etc, Llama Extract allows you to make use of an LLM to extract the information relevant to you, in a structured format.\n",
"\n",
"In this example, we'll be using [LlamaExtract](https://docs.cloud.llamaindex.ai/llamaextract/getting_started?utm_campaign=extract&utm_medium=recipe) to extract structured data from an SEC filing (specifically, the filing by Nvidia for fiscal year 2025).\n",
"\n",
"On top of simple data extraction, we'll ask our extraction agent to provide citations and reasoning for each extracted field. This allows us to:\n",
"- Confirm the accuracy of the extracted field\n",
"- Understand the reasoning behind why the LLM extracted a given piece of information\n",
"- This last point allows us an opportunity to adjust the system prompt or field descriptions and improve on results where needed.\n",
"\n",
"\n",
"The example we go through below is also replicable within Llama Cloud as well, where you will also be able to pick between a number of pre-defined schemas, instead of building your own."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-cloud-services"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Connect to Llama Cloud\n",
"\n",
"To get started, make sure you provide your [Llama Cloud](https://cloud.llamaindex.ai?utm_campaign=extract&utm_medium=recipe) API key."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Enter your Llama Cloud API Key: ··········\n"
]
}
],
"source": [
"import os\n",
"from getpass import getpass\n",
"\n",
"if \"LLAMA_CLOUD_API_KEY\" not in os.environ:\n",
" os.environ[\"LLAMA_CLOUD_API_KEY\"] = getpass(\"Enter your Llama Cloud API Key: \")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Extract Data with Llama Extract Agent"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"No project_id provided, fetching default project.\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaExtract\n",
"\n",
"# Optionally, provide your project id, if not, it will use the 'Default' project\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Provide Your Custom Schema\n",
"\n",
"When using LlamaExtract via the API, you provide your own schema that describes what you want extracted from files and data provided to your agent. Here, we are essentially building an SEC filings extraction agent."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from enum import Enum\n",
"\n",
"\n",
"class FilingType(str, Enum):\n",
" ten_k = \"10 K\"\n",
" ten_q = \"10-Q\"\n",
" ten_ka = \"10-K/A\"\n",
" ten_qa = \"10-Q/A\"\n",
"\n",
"\n",
"class FinancialReport(BaseModel):\n",
" company_name: str = Field(description=\"The name of the company\")\n",
" description: str = Field(\n",
" description=\"Short description of the filing and what it contains\"\n",
" )\n",
" filing_type: FilingType = Field(description=\"Type of SEC filing\")\n",
" filing_date: str = Field(description=\"Date when filing was submitted to SEC\")\n",
" fiscal_year: int = Field(description=\"Fiscal year\")\n",
" unit: str = Field(\n",
" description=\"Unit of financial figures (thousands, millions, etc.)\"\n",
" )\n",
" revenue: int = Field(description=\"Total revenue for period\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Set Up Citations and Reasoning\n",
"\n",
"Optionally, we can set the `ExtractConfig` to extract citations for each field the agent extracts. These cications will cite the specific pages and sections of the file from which a given field was extractedd.\n",
"\n",
"By setting `use_reasoning` to True, we als ask the agent to do an additional reasoning step, explaining why a given field was extracted."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud.types import ExtractConfig, ExtractMode\n",
"\n",
"config = ExtractConfig(\n",
" use_reasoning=True, cite_sources=True, extraction_mode=ExtractMode.MULTIMODAL\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:127: ExperimentalWarning: `use_reasoning` is an experimental feature. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n",
"/usr/local/lib/python3.11/dist-packages/llama_cloud_services/extract/extract.py:133: ExperimentalWarning: `cite_sources` is an experimental feature. This may greatly increase the size of the response, and slow down the extraction. Results will be available in the `extraction_metadata` field for the extraction run.\n",
" warnings.warn(\n"
]
}
],
"source": [
"agent = llama_extract.create_agent(\n",
" name=\"filing-parser\", data_schema=FinancialReport, config=config\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Demo Time - Download a PDF and Extract Data with Citations"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"PDF downloaded successfully.\n"
]
}
],
"source": [
"import requests\n",
"\n",
"url = \"https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf\"\n",
"\n",
"response = requests.get(url)\n",
"\n",
"if response.status_code == 200:\n",
" with open(\"/content/nvda_10k.pdf\", \"wb\") as f:\n",
" f.write(response.content)\n",
" print(\"PDF downloaded successfully.\")\n",
"else:\n",
" print(f\"Failed to download. Status code: {response.status_code}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|██████████| 1/1 [00:00<00:00, 1.83it/s]\n",
"Creating extraction jobs: 100%|██████████| 1/1 [00:00<00:00, 4.38it/s]\n",
"Extracting files: 100%|██████████| 1/1 [02:03<00:00, 123.40s/it]\n"
]
}
],
"source": [
"filing_info = agent.extract(\"/content/nvda_10k.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'company_name': 'NVIDIA Corporation',\n",
" 'description': \"The filing provides a detailed overview of NVIDIA's business as a full-stack computing infrastructure company, discusses various technologies including digital avatars and autonomous vehicles, outlines numerous risk factors affecting operations such as supply chain issues and geopolitical tensions, and describes employee stock purchase plans and related compliance requirements.\",\n",
" 'filing_type': '10 K',\n",
" 'filing_date': 'February 26, 2025',\n",
" 'fiscal_year': 2025,\n",
" 'unit': 'millions',\n",
" 'revenue': 130497}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Inspect Citations and Reasoning"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'field_metadata': {'company_name': {'reasoning': 'VERBATIM EXTRACTION',\n",
" 'citation': [{'page': 1, 'matching_text': 'NVIDIA CORPORATION'},\n",
" {'page': 2, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 3,\n",
" 'matching_text': 'All references to \"NVIDIA,\" \"we,\" \"us,\" \"our,\" or the \"Company\" mean NVIDIA Corporation and its subsidiaries.'},\n",
" {'page': 35,\n",
" 'matching_text': 'Comparison of 5 Year Cumulative Total Return* Among NVIDIA Corporation'},\n",
" {'page': 49,\n",
" 'matching_text': 'To the Board of Directors and Shareholders of NVIDIA Corporation'},\n",
" {'page': 90, 'matching_text': 'NVIDIA Corporation'},\n",
" {'page': 119,\n",
" 'matching_text': '*\"Company\"* means NVIDIA Corporation, a Delaware corporation.'},\n",
" {'page': 126,\n",
" 'matching_text': 'Annual Report on Form 10-K of NVIDIA Corporation'}]},\n",
" 'filing_type': {'reasoning': \"VERBATIM EXTRACTION from multiple sources confirming the filing type as '10 K'.\",\n",
" 'citation': [{'page': 1, 'matching_text': 'FORM 10-K'},\n",
" {'page': 2, 'matching_text': 'Item 16. | Form 10-K Summary'},\n",
" {'page': 3,\n",
" 'matching_text': 'This Annual Report on Form 10-K contains forward-looking statements...'},\n",
" {'page': 13, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 15, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 32,\n",
" 'matching_text': 'Annual Report on Form 10-K, which information is hereby incorporated by reference.'},\n",
" {'page': 36, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 43,\n",
" 'matching_text': 'Annual Report on Form 10-K for additional information'},\n",
" {'page': 45, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 46, 'matching_text': 'this Annual Report on Form 10-K'},\n",
" {'page': 62, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 83,\n",
" 'matching_text': 'Restated Certificate of Incorporation | 10-K'},\n",
" {'page': 84, 'matching_text': 'Item 16. Form 10-K Summary'},\n",
" {'page': 126, 'matching_text': 'which appears in this Form 10-K'},\n",
" {'page': 127, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 128, 'matching_text': 'Annual Report on Form 10-K'},\n",
" {'page': 129, 'matching_text': \"The Company's Annual Report on Form 10-K\"},\n",
" {'page': 130,\n",
" 'matching_text': \"The Company's Annual Report on Form 10-K for the year ended January 26, 2025\"}]},\n",
" 'fiscal_year': {'reasoning': 'The fiscal year ended January 26, 2025, indicates the fiscal year is 2025. Additionally, multiple references throughout the text confirm the fiscal year 2025 in various contexts.',\n",
" 'citation': [{'page': 1,\n",
" 'matching_text': 'For the fiscal year ended January 26, 2025'},\n",
" {'page': 6,\n",
" 'matching_text': 'In fiscal year 2025, we launched the NVIDIA Blackwell architecture'},\n",
" {'page': 12, 'matching_text': 'fiscal year 2025'},\n",
" {'page': 17,\n",
" 'matching_text': 'our gross margins in the second quarter of fiscal year 2025 were negatively impacted'},\n",
" {'page': 20,\n",
" 'matching_text': 'we generated 53% of our revenue in fiscal year 2025 from sales outside the United States.'},\n",
" {'page': 23,\n",
" 'matching_text': 'For fiscal year 2025, an indirect customer which primarily purchases our products through system integrators...'},\n",
" {'page': 33,\n",
" 'matching_text': 'In fiscal year 2025, we repurchased 310 million shares of our common stock for $34.0 billion.'},\n",
" {'page': 37,\n",
" 'matching_text': 'Our Data Center revenue in China grew in fiscal year 2025.'},\n",
" {'page': 44,\n",
" 'matching_text': 'Cash provided by operating activities increased in fiscal year 2025 compared to fiscal year 2024'},\n",
" {'page': 57,\n",
" 'matching_text': 'Fiscal years 2025, 2024 and 2023 were all 52-week years.'},\n",
" {'page': 65,\n",
" 'matching_text': 'Beginning in the second quarter of fiscal year 2025'},\n",
" {'page': 69, 'matching_text': 'In the fourth quarter of fiscal year 2025'},\n",
" {'page': 78,\n",
" 'matching_text': 'Depreciation and amortization expense attributable to our Compute and Networking segment for fiscal years 2025'},\n",
" {'page': 129, 'matching_text': 'for the year ended January 26, 2025'}]},\n",
" 'description': {'reasoning': 'The extracted data combines multiple descriptions from the source text, ensuring no duplication while maintaining the order and context of the information. Each section of the filing is summarized to reflect the key points without losing the essence of the original text.',\n",
" 'citation': [{'page': 4,\n",
" 'matching_text': 'NVIDIA is now a full-stack computing infrastructure company with data-center-scale offerings that are reshaping industry.'},\n",
" {'page': 8,\n",
" 'matching_text': 'a suite of technologies that help developers bring digital avatars to life with generative Al...autonomous vehicles, or AV, and electric vehicles, or EV, is revolutionizing the transportation industry...Our worldwide sales and marketing strategy is key to achieving our objective of providing markets with our high-performance and efficient computing platforms and software.'},\n",
" {'page': 14, 'matching_text': 'Risk Factors Summary'},\n",
" {'page': 16,\n",
" 'matching_text': 'Risks Related to Demand, Supply, and Manufacturing\\n\\nLong manufacturing lead times and uncertain supply and component availability...'},\n",
" {'page': 18,\n",
" 'matching_text': 'cryptocurrency mining, on demand for our products. Volatility in the cryptocurrency market, including new compute technologies...'},\n",
" {'page': 21,\n",
" 'matching_text': 'supply-chain attacks or other business disruptions. We cannot guarantee that third parties and infrastructure in our supply chain...'},\n",
" {'page': 22,\n",
" 'matching_text': 'We are monitoring the impact of the geopolitical conflict in and around Israel on our operations... Climate change may have a long-term impact on our business.'},\n",
" {'page': 25,\n",
" 'matching_text': 'We are subject to complex laws, rules, regulations, and political and other actions, including restrictions on the export of our products, which may adversely impact our business.'},\n",
" {'page': 28,\n",
" 'matching_text': 'Our competitive position has been harmed by the existing export controls, and our competitive position and future results may be further harmed'},\n",
" {'page': 29,\n",
" 'matching_text': 'restrictions imposed by the Chinese government on the duration of gaming activities and access to games may adversely affect our Gaming revenue'},\n",
" {'page': 29,\n",
" 'matching_text': 'our business depends on our ability to receive consistent and reliable supply from our overseas partners, especially in Taiwan and South Korea'},\n",
" {'page': 29,\n",
" 'matching_text': 'Increased scrutiny from shareholders, regulators and others regarding our corporate sustainability practices could result in additional costs'},\n",
" {'page': 29,\n",
" 'matching_text': 'Concerns relating to the responsible use of new and evolving technologies, such as Al, in our products and services may result in reputational or financial harm'},\n",
" {'page': 31,\n",
" 'matching_text': 'Data protection laws around the world are quickly changing and may be interpreted and applied in an increasingly stringent fashion...'}]},\n",
" 'filing_date': {'reasoning': 'The filing date is consistently mentioned as February 26, 2025 across multiple entries, making it the most reliable date for the filing.',\n",
" 'citation': [{'page': 51, 'matching_text': 'February 26, 2025'},\n",
" {'page': 86, 'matching_text': 'on February 26, 2025.'},\n",
" {'page': 87, 'matching_text': 'February 26, 2025'},\n",
" {'page': 126, 'matching_text': 'our report dated February 26, 2025'},\n",
" {'page': 127, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 128, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 129, 'matching_text': 'Date: February 26, 2025'},\n",
" {'page': 130, 'matching_text': 'Date: February 26, 2025'}]},\n",
" 'unit': {'reasoning': \"The unit of financial figures is explicitly mentioned multiple times in the text as 'millions', including in table headers and notes. This is confirmed by various citations from pages 38, 42, 43, 52, 53, 54, 56, 65, 71, 72, 73, 75, 77, 79, 80, and 82.\",\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': '($ in millions, except per share data)'},\n",
" {'page': 42, 'matching_text': '($ in millions)'},\n",
" {'page': 43, 'matching_text': '($ in millions)'},\n",
" {'page': 52, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 53,\n",
" 'matching_text': 'Consolidated Statements of Comprehensive Income (In millions)'},\n",
" {'page': 54,\n",
" 'matching_text': 'Consolidated Balance Sheets (In millions, except par value)'},\n",
" {'page': 55, 'matching_text': '(In millions, except per share data)'},\n",
" {'page': 56,\n",
" 'matching_text': 'Consolidated Statements of Cash Flows (In millions)'},\n",
" {'page': 65,\n",
" 'matching_text': 'Year Ended<br/>Jan 26, 2025<br/>(In millions, except per share data)'},\n",
" {'page': 71, 'matching_text': '(In millions) | (In millions)'},\n",
" {'page': 72, 'matching_text': '(In millions)'}]},\n",
" 'revenue': {'reasoning': 'The total revenue for fiscal year 2025 is extracted from multiple sources within the text, all confirming the same figure of $130,497 million. The revenue recognized for fiscal year 2025 is also noted as $4,607 million, which is a separate figure. However, the primary focus is on the total revenue figure, which is consistently cited.',\n",
" 'citation': [{'page': 38,\n",
" 'matching_text': 'Revenue for fiscal year 2025 was $130.5 billion'},\n",
" {'page': 41,\n",
" 'matching_text': 'Total | $ 130,497 | $ | 60,922'},\n",
" {'page': 52, 'matching_text': 'Revenue | $ 130,497'},\n",
" {'page': 78,\n",
" 'matching_text': 'Revenue | $ 116,193 | $ 14,304 | $ - | $ 130,497'},\n",
" {'page': 79, 'matching_text': 'Total revenue | $ 130,497'},\n",
" {'page': 80, 'matching_text': 'Total revenue | $ 130,497'}]}},\n",
" 'usage': {'num_pages_extracted': 130,\n",
" 'num_document_tokens': 105932,\n",
" 'num_output_tokens': 31306}}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"filing_info.extraction_metadata"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## What's Next?\n",
"\n",
"In this example, we built an Extraction Agent that is capable of citing it's sources from the document it's extracting data from, and reasoning about its reponse. To further customize and improve on the results, you can also try to customize the `system_prompt` in the `ExtractConfig`.\n",
"\n",
"#### Learn More\n",
"\n",
"- [LlamaExtract Documentation](https://docs.cloud.llamaindex.ai/llamaextract/getting_started)\n",
"- [Example Notebooks](https://github.com/run-llama/llama_cloud_services/tree/main/examples/extract)"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
},
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
File diff suppressed because it is too large Load Diff
@@ -1,318 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "1f6bd03d-1b8b-45a0-bc2c-5a13f1a5d8d3",
"metadata": {},
"source": [
"# LM317 Voltage Regulator Datasheet Structured Extraction\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/lm317_structured_extraction.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates an agentic document workflow using LlamaExtract to process an LM317 voltage regulator datasheet. In this example, we define a structured extraction schema that converts key technical fields into standardized subfields. For instance, the output voltage is split into a minimum and maximum value with a defined unit, and we capture page citations for each extracted field.\n",
"\n",
"The target user is an electronics engineer at a component manufacturing company who needs to consolidate datasheet information into a standardized specification sheet for design and quality control.\n",
"\n",
"This approach reduces manual data entry, improves extraction accuracy and standardization, and provides traceability for each technical detail."
]
},
{
"cell_type": "markdown",
"id": "a3b8c8d5-ff3e-48ce-b0b8-29b6b1f517f8",
"metadata": {},
"source": [
"## Use Case Overview\n",
"\n",
"### Problem\n",
"Datasheets like that for the LM317 regulator are often distributed as PDFs containing multiple tables, charts, and complex textual descriptions. Engineers must manually extract technical details such as voltage ranges, dropout voltage, maximum current, input voltage range, and pin configurations. This process is error-prone and time-consuming.\n",
"\n",
"### Agent Workflow (Combination of Automation and Chat)\n",
"1. **Upload Datasheet:** The engineer uploads the LM317 datasheet PDF. \n",
"2. **Structured Extraction:** An automated agent processes the PDF and extracts key technical details into structured fields (e.g., output voltage as a range with separate min/max values).\n",
"3. **Interactive Verification:** The engineer can query the agent (via chat) for further details or clarification (e.g., \"Show me the detailed pin configuration extraction\") and review the cited pages.\n",
"\n",
"**Value Delivered:**\n",
"- Up to 70% reduction in manual data extraction time.\n",
"- Increased accuracy and standardization with structured fields."
]
},
{
"cell_type": "markdown",
"id": "a704e843-54be-4969-842b-713584cb3c35",
"metadata": {},
"source": [
"## Setup and Download Data\n",
"\n",
"Download the [LM317 Datasheet](https://www.ti.com/lit/ds/symlink/lm317.pdf) and setup LlamaExtract."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6e5b1f91-8785-44d4-a710-8be1b48b76de",
"metadata": {},
"outputs": [],
"source": [
"!mkdir -p data/lm317_structured_extraction\n",
"!wget https://www.ti.com/lit/ds/symlink/lm317.pdf -O data/lm317_structured_extraction/lm317.pdf"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f17b914a-00ed-4b63-8198-69fd7c4a7c62",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"\n",
"# Load environment variables (ensure LLAMA_CLOUD_API_KEY is set in your .env file)\n",
"load_dotenv(override=True)\n",
"\n",
"# Initialize the LlamaExtract client\n",
"llama_extract = LlamaExtract(\n",
" project_id=\"<project_id>\",\n",
" organization_id=\"<organization_id>\",\n",
")"
]
},
{
"cell_type": "markdown",
"id": "ed9f6e9a-96c8-4ee1-8b45-0b6a4f7dbbf1",
"metadata": {},
"source": [
"## Defining a Structured Extraction Schema\n",
"\n",
"We now define a rich Pydantic schema to extract technical specifications from the LM317 datasheet. In this schema:\n",
"\n",
"- The **output_voltage** and **input_voltage** fields are structured as ranges with separate minimum and maximum values and a unit.\n",
"- The **pin_configuration** field is structured to include a pin count and a descriptive layout.\n",
"- Additional technical fields (e.g., dropout voltage, max current) are captured as numbers.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "4f7e9b44-5e69-4b30-9864-cd98f1e2a7d4",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"class VoltageRange(BaseModel):\n",
" min_voltage: float = Field(..., description=\"Minimum voltage in volts\")\n",
" max_voltage: float = Field(..., description=\"Maximum voltage in volts\")\n",
" unit: str = Field(\"V\", description=\"Voltage unit\")\n",
"\n",
"\n",
"class PinConfiguration(BaseModel):\n",
" pin_count: int = Field(..., description=\"Number of pins\")\n",
" layout: str = Field(..., description=\"Detailed pin layout description\")\n",
"\n",
"\n",
"class LM317Spec(BaseModel):\n",
" component_name: str = Field(..., description=\"Name of the component\")\n",
" output_voltage: VoltageRange = Field(\n",
" ..., description=\"Output voltage range specification\"\n",
" )\n",
" dropout_voltage: float = Field(..., description=\"Dropout voltage in volts\")\n",
" max_current: float = Field(..., description=\"Maximum current rating in amperes\")\n",
" input_voltage: VoltageRange = Field(\n",
" ..., description=\"Input voltage range specification\"\n",
" )\n",
" pin_configuration: PinConfiguration = Field(\n",
" ..., description=\"Pin configuration details\"\n",
" )\n",
" features: List[str] = Field([], description=\"List of additional technical features\")\n",
"\n",
"\n",
"class LM317Schema(BaseModel):\n",
" specs: List[LM317Spec] = Field(\n",
" ..., description=\"List of extracted LM317 technical specifications\"\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e0508e38-35be-446c-afe7-129e39553281",
"metadata": {},
"outputs": [],
"source": [
"try:\n",
" existing_agent = llama_extract.get_agent(name=\"lm317-datasheet\")\n",
" if existing_agent:\n",
" llama_extract.delete_agent(existing_agent.id)\n",
"except ApiError as e:\n",
" if e.status_code == 404:\n",
" pass\n",
" else:\n",
" raise"
]
},
{
"cell_type": "markdown",
"id": "bb197dfd-dd37-459e-8953-cc1b12f25bdd",
"metadata": {},
"source": [
"Here we use our balanced extraction mode."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e3defc0a-c685-4fbd-bbb1-1270f1442e72",
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud import ExtractConfig\n",
"\n",
"extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
")\n",
"\n",
"agent = llama_extract.create_agent(\n",
" name=\"lm317-datasheet\", data_schema=LM317Schema, config=extract_config\n",
")"
]
},
{
"cell_type": "markdown",
"id": "c0a0f9f9-2ef3-4a38-bd74-68d2c2e9e2d8",
"metadata": {},
"source": [
"## Extracting Information from the LM317 Datasheet\n",
"\n",
"For this demonstration, please download a publicly available LM317 voltage regulator datasheet (for example, from Texas Instruments) and save it as `lm317.pdf` in the `./data` directory. Then run the cell below to extract the structured technical specifications."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c58e8b7a-8f9b-46f3-8f72-3c2f96b49e8f",
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:01<00:00, 1.08s/it]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.96it/s]\n",
"Extracting files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [01:27<00:00, 87.38s/it]\n"
]
}
],
"source": [
"# Path to the LM317 datasheet PDF\n",
"lm317_pdf = \"./data/lm317_structured_extraction/lm317.pdf\"\n",
"\n",
"# Extract structured technical specifications from the datasheet\n",
"lm317_extract = agent.extract(lm317_pdf)"
]
},
{
"cell_type": "markdown",
"id": "1a2e2e44-6c48-4a38-a6de-5f2f3c7d4d8b",
"metadata": {},
"source": [
"## Assessing the Extraction Results\n",
"\n",
"The output will be a consolidated list of LM317 technical specifications. For each entry, you should see structured fields including:\n",
"\n",
"- **component_name**\n",
"- **output_voltage** as a range (with separate `min_voltage` and `max_voltage` plus `unit`)\n",
"- **dropout_voltage** and **max_current** as numbers\n",
"- **input_voltage** as a structured range\n",
"- **pin_configuration** with a `pin_count` and `layout`\n",
"- **features** (if available)\n",
"\n",
"This structured approach makes it easier to standardize the information for downstream integration and verification. Engineers can click on the cited page numbers (in a UI that supports it) to validate the extraction."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "fb2abc44-7c9b-4b19-958e-d0d7b390ae57",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'specs': [{'component_name': 'LM317',\n",
" 'output_voltage': {'min_voltage': 1.25, 'max_voltage': 37.0, 'unit': 'V'},\n",
" 'dropout_voltage': 0.0,\n",
" 'max_current': 1.5,\n",
" 'input_voltage': {'min_voltage': 4.25, 'max_voltage': 40.0, 'unit': 'V'},\n",
" 'pin_configuration': {'pin_count': 3,\n",
" 'layout': '1: ADJUST, 2: OUTPUT, 3: INPUT'},\n",
" 'features': ['Output voltage range adjustable from 1.25 V to 37 V',\n",
" 'Output current greater than 1.5 A',\n",
" 'Internal short-circuit current limiting',\n",
" 'Thermal overload protection',\n",
" 'Output safe-area compensation']}]}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# Display the extraction results\n",
"lm317_extract.data"
]
},
{
"cell_type": "markdown",
"id": "c7a2a523-095e-40bf-b713-f509c13a7747",
"metadata": {},
"source": [
"You can also see the output result in the UI."
]
},
{
"cell_type": "markdown",
"id": "dc22dfa5-b667-4fb0-8dbe-24e401b12389",
"metadata": {},
"source": [
"![](data/lm317_structured_extraction/lm317_extraction.png)"
]
},
{
"cell_type": "markdown",
"id": "e0e0c12a-9f89-4bb3-b40d-3e9f7c6d2fef",
"metadata": {},
"source": [
"## Conclusion\n",
"\n",
"This notebook demonstrated how to use LlamaExtract with a structured extraction schema for the LM317 voltage regulator datasheet. By defining detailed subfields (such as splitting voltage ranges into minimum and maximum values, and structuring the pin configuration), we ensure that the extracted data is standardized and traceable through page citations. This approach minimizes manual effort and improves accuracy, providing a robust example of an agentic document workflow for technical documentation processing.\n",
"\n",
"Feel free to modify or extend the schema to capture additional technical details or to suit your own use cases."
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
-834
View File
@@ -1,834 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Extracting data from Resumes\n",
"\n",
"Let us assume that we are running a hiring process for a company and we have received a list of resumes from candidates. We want to extract structured data from the resumes so that we can run a screening process and shortlist candidates. \n",
"\n",
"Take a look at one of the resumes in the `data/resumes` directory. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/html": [
"\n",
" <iframe\n",
" width=\"600\"\n",
" height=\"400\"\n",
" src=\"./data/resumes/ai_researcher.pdf\"\n",
" frameborder=\"0\"\n",
" allowfullscreen\n",
" \n",
" ></iframe>\n",
" "
],
"text/plain": [
"<IPython.lib.display.IFrame at 0x109a7dcd0>"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"from IPython.display import IFrame\n",
"\n",
"IFrame(src=\"./data/resumes/ai_researcher.pdf\", width=600, height=400)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"You will notice that all the resumes have different layouts but contain common information like name, email, experience, education, etc. \n",
"\n",
"With LlamaExtract, we will show you how to:\n",
"- *Define* a data schema to extract the information of interest. \n",
"- *Iterate* over the data schema to generalize the schema for multiple resumes.\n",
"- *Finalize* the schema and schedule extractions for multiple resumes.\n",
"\n",
"We will start by defining a `LlamaExtract` client which provides a Python interface to the LlamaExtract API. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"\n",
"\n",
"# Load environment variables (put LLAMA_CLOUD_API_KEY in your .env file)\n",
"load_dotenv(override=True)\n",
"\n",
"# Optionally, add your project id/organization id\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Defining the data schema\n",
"\n",
"Next, let us try to extract two fields from the resume: `name` and `email`. We can either use a Python dictionary structure to define the `data_schema` as a JSON or use a Pydantic model instead, for brevity and convenience. In either case, our output is guaranteed to validate against this schema."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"\n",
"\n",
"class Resume(BaseModel):\n",
" name: str = Field(description=\"The name of the candidate\")\n",
" email: str = Field(description=\"The email address of the candidate\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:02<00:00, 2.20s/it]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:02<00:00, 2.93s/it]\n",
"Extracting files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:02<00:00, 2.94s/it]\n",
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.13it/s]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.80it/s]\n",
"Extracting files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:15<00:00, 15.18s/it]\n",
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.16it/s]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 2.33it/s]\n",
"Extracting files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:32<00:00, 32.86s/it]\n"
]
}
],
"source": [
"from llama_cloud.core.api_error import ApiError\n",
"\n",
"try:\n",
" existing_agent = llama_extract.get_agent(name=\"resume-screening\")\n",
" if existing_agent:\n",
" llama_extract.delete_agent(existing_agent.id)\n",
"except ApiError as e:\n",
" if e.status_code == 404:\n",
" pass\n",
" else:\n",
" raise\n",
"\n",
"agent = llama_extract.create_agent(name=\"resume-screening\", data_schema=Resume)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[ExtractionAgent(id=1fef43b5-8230-43b4-9e80-c1cddf53889c, name=resume-screening),\n",
" ExtractionAgent(id=93f8508b-3570-46f0-ae62-6315b40043bd, name=receipt/noisebridge_receipt.pdf_56db3d92),\n",
" ExtractionAgent(id=08315f0e-7146-430b-99b8-9701cb3ace6a, name=receipt/noisebridge_receipt.pdf_5c4730a7),\n",
" ExtractionAgent(id=cfcd7756-015d-4dbd-b142-a3eefcb16cd3, name=resume/software_architect_resume.html_4a11cf15),\n",
" ExtractionAgent(id=17cb83d9-601e-4f5c-a7aa-286e3045bcb4, name=resume/software_architect_resume.html_0b7d84a8),\n",
" ExtractionAgent(id=adc8e88c-44d3-4613-a5aa-d666ef007494, name=slide/saas_slide.pdf_bcc627a5),\n",
" ExtractionAgent(id=189f14cd-6370-4476-a6ad-36eafbc62618, name=slide/saas_slide.pdf_065aa22b),\n",
" ExtractionAgent(id=b9938ca5-6225-43cb-89ea-b0065237792f, name=test2),\n",
" ExtractionAgent(id=574d37b8-59dc-41e9-bde0-5c506a8eb670, name=test)]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"llama_extract.list_agents()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Dr. Rachel Zhang', 'email': 'rachel.zhang@email.com'}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"resume = agent.extract(\"./data/resumes/ai_researcher.pdf\")\n",
"resume.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Iterating over the data schema\n",
"\n",
"Now that we have created a data schema, let us add more fields to the schema. We will add `experience` and `education` fields to the schema. \n",
"- We can create a new Pydantic model for each of these fields and represent `experience` and `education` as lists of these models. Doing this will allow us to extract multiple entities from the resume without having to pre-define how many experiences or education the candidate has. \n",
"- We have added a `description` parameter to provide more context for extraction. We can use `description` to provide example inputs/outputs for the extraction. \n",
"- Note that we have annotated the `start_date` and `end_date` fields with `Optional[str]` to indicate that these fields are optional. This is *important* because the schema will be used to extract data from multiple resumes and not all resumes will have the same format. A field must only be required if it is guaranteed to be present in all the resumes. \n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from typing import List, Optional\n",
"\n",
"\n",
"class Education(BaseModel):\n",
" institution: str = Field(description=\"The institution of the candidate\")\n",
" degree: str = Field(description=\"The degree of the candidate\")\n",
" start_date: Optional[str] = Field(\n",
" default=None, description=\"The start date of the candidate's education\"\n",
" )\n",
" end_date: Optional[str] = Field(\n",
" default=None, description=\"The end date of the candidate's education\"\n",
" )\n",
"\n",
"\n",
"class Experience(BaseModel):\n",
" company: str = Field(description=\"The name of the company\")\n",
" title: str = Field(description=\"The title of the candidate\")\n",
" description: Optional[str] = Field(\n",
" default=None, description=\"The description of the candidate's experience\"\n",
" )\n",
" start_date: Optional[str] = Field(\n",
" default=None, description=\"The start date of the candidate's experience\"\n",
" )\n",
" end_date: Optional[str] = Field(\n",
" default=None, description=\"The end date of the candidate's experience\"\n",
" )\n",
"\n",
"\n",
"class Resume(BaseModel):\n",
" name: str = Field(description=\"The name of the candidate\")\n",
" email: str = Field(description=\"The email address of the candidate\")\n",
" links: List[str] = Field(\n",
" description=\"The links to the candidate's social media profiles\"\n",
" )\n",
" experience: List[Experience] = Field(description=\"The candidate's experience\")\n",
" education: List[Education] = Field(description=\"The candidate's education\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Next, we will update the `data_schema` for the `resume-screening` agent to use the new `Resume` model. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Dr. Rachel Zhang',\n",
" 'email': 'rachel.zhang@email.com',\n",
" 'links': ['linkedin.com/in/rachelzhang',\n",
" 'github.com/rzhang-ai',\n",
" 'scholar.google.com/rachelzhang'],\n",
" 'experience': [{'company': 'DeepMind',\n",
" 'title': 'Senior Research Scientist',\n",
" 'description': '- Lead researcher on large-scale multi-task learning systems, developing novel architectures that improve cross-task generalization by 40%\\n- Pioneered new approach to zero-shot learning using contrastive training, published in NeurIPS 2023\\n- Built and led team of 6 researchers working on foundational ML models\\n- Developed novel regularization techniques for large language models, reducing catastrophic forgetting by 35%',\n",
" 'start_date': '2019',\n",
" 'end_date': 'Present'},\n",
" {'company': 'Google Research',\n",
" 'title': 'Research Scientist',\n",
" 'description': '- Developed probabilistic frameworks for robust ML, published in ICML 2018\\n- Created novel attention mechanisms for computer vision models, improving accuracy by 25%\\n- Led collaboration with Google Brain team on efficient training methods for transformer models\\n- Mentored 4 PhD interns and collaborated with academic institutions',\n",
" 'start_date': '2015',\n",
" 'end_date': '2019'},\n",
" {'company': 'Columbia University',\n",
" 'title': 'Research Assistant Professor',\n",
" 'description': '- Published seminal work on Bayesian optimization methods (cited 1000+ times)\\n- Taught graduate-level courses in Machine Learning and Statistical Learning Theory\\n- Supervised 5 PhD students and 3 MSc students\\n- Secured $500K in research grants for probabilistic ML research',\n",
" 'start_date': '2011',\n",
" 'end_date': '2015'}],\n",
" 'education': [{'institution': 'Columbia University',\n",
" 'degree': 'Ph.D. in Computer Science',\n",
" 'start_date': '2007',\n",
" 'end_date': '2011'},\n",
" {'institution': 'Stanford University',\n",
" 'degree': 'M.S. in Computer Science',\n",
" 'start_date': '2005',\n",
" 'end_date': '2007'}]}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"agent.data_schema = Resume\n",
"resume = agent.extract(\"./data/resumes/ai_researcher.pdf\")\n",
"resume.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This is a good start. Let us add a few more fields to the schema and re-run the extraction. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"class TechnicalSkills(BaseModel):\n",
" programming_languages: List[str] = Field(\n",
" description=\"The programming languages the candidate is proficient in.\"\n",
" )\n",
" frameworks: List[str] = Field(\n",
" description=\"The tools/frameworks the candidate is proficient in, e.g. React, Django, PyTorch, etc.\"\n",
" )\n",
" skills: List[str] = Field(\n",
" description=\"Other general skills the candidate is proficient in, e.g. Data Engineering, Machine Learning, etc.\"\n",
" )\n",
"\n",
"\n",
"class Resume(BaseModel):\n",
" name: str = Field(description=\"The name of the candidate\")\n",
" email: str = Field(description=\"The email address of the candidate\")\n",
" links: List[str] = Field(\n",
" description=\"The links to the candidate's social media profiles\"\n",
" )\n",
" experience: List[Experience] = Field(description=\"The candidate's experience\")\n",
" education: List[Education] = Field(description=\"The candidate's education\")\n",
" technical_skills: TechnicalSkills = Field(\n",
" description=\"The candidate's technical skills\"\n",
" )\n",
" key_accomplishments: str = Field(\n",
" description=\"Summarize the candidates highest achievements.\"\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Dr. Rachel Zhang, Ph.D.',\n",
" 'email': 'rachel.zhang@email.com',\n",
" 'links': ['linkedin.com/in/rachelzhang',\n",
" 'github.com/rzhang-ai',\n",
" 'scholar.google.com/rachelzhang'],\n",
" 'experience': [{'company': 'DeepMind',\n",
" 'title': 'Senior Research Scientist',\n",
" 'description': 'Lead researcher on large-scale multi-task learning systems, developing novel architectures that improve cross-task generalization by 40%\\nPioneered new approach to zero-shot learning using contrastive training, published in NeurIPS 2023\\nBuilt and led team of 6 researchers working on foundational ML models\\nDeveloped novel regularization techniques for large language models, reducing catastrophic forgetting by 35%',\n",
" 'start_date': '2019',\n",
" 'end_date': 'Present'},\n",
" {'company': 'Google Research',\n",
" 'title': 'Research Scientist',\n",
" 'description': 'Developed probabilistic frameworks for robust ML, published in ICML 2018\\nCreated novel attention mechanisms for computer vision models, improving accuracy by 25%\\nLed collaboration with Google Brain team on efficient training methods for transformer models\\nMentored 4 PhD interns and collaborated with academic institutions',\n",
" 'start_date': '2015',\n",
" 'end_date': '2019'},\n",
" {'company': 'Columbia University',\n",
" 'title': 'Research Assistant Professor',\n",
" 'description': 'Published seminal work on Bayesian optimization methods (cited 1000+ times)\\nTaught graduate-level courses in Machine Learning and Statistical Learning Theory\\nSupervised 5 PhD students and 3 MSc students\\nSecured $500K in research grants for probabilistic ML research',\n",
" 'start_date': '2011',\n",
" 'end_date': '2015'}],\n",
" 'education': [{'institution': 'Columbia University',\n",
" 'degree': 'Ph.D. in Computer Science',\n",
" 'start_date': '2007',\n",
" 'end_date': '2011'},\n",
" {'institution': 'Stanford University',\n",
" 'degree': 'M.S. in Computer Science',\n",
" 'start_date': '2005',\n",
" 'end_date': '2007'}],\n",
" 'technical_skills': {'programming_languages': ['Python',\n",
" 'C++',\n",
" 'Julia',\n",
" 'CUDA'],\n",
" 'frameworks': ['PyTorch', 'TensorFlow', 'JAX', 'Ray'],\n",
" 'skills': ['Deep Learning',\n",
" 'Reinforcement Learning',\n",
" 'Probabilistic Models',\n",
" 'Multi-Task Learning',\n",
" 'Zero-Shot Learning',\n",
" 'Neural Architecture Search']},\n",
" 'key_accomplishments': 'AI researcher with 12+ years of experience spanning classical machine learning, deep learning, and probabilistic modeling. Led groundbreaking research in reinforcement learning, generative models, and multi-task learning. Published 25+ papers in top-tier conferences (NeurIPS, ICML, ICLR). Strong track record of transitioning theoretical advances into practical applications in both academic and industrial settings.'}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"agent.data_schema = Resume\n",
"resume = agent.extract(\"./data/resumes/ai_researcher.pdf\")\n",
"resume.data"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Finalizing the schema\n",
"\n",
"This is great! We have extracted a lot of key information from the resume that is well-typed and can be used downstream for further processing. Until now, this data is ephemeral and will be lost if we close the session. Let us save the state of our extraction and use it to extract data from multiple resumes. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"agent.save()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'type': 'object',\n",
" 'required': ['name',\n",
" 'email',\n",
" 'links',\n",
" 'experience',\n",
" 'education',\n",
" 'technical_skills',\n",
" 'key_accomplishments'],\n",
" 'properties': {'name': {'type': 'string',\n",
" 'description': 'The name of the candidate'},\n",
" 'email': {'type': 'string',\n",
" 'description': 'The email address of the candidate'},\n",
" 'links': {'type': 'array',\n",
" 'items': {'type': 'string'},\n",
" 'description': \"The links to the candidate's social media profiles\"},\n",
" 'education': {'type': 'array',\n",
" 'items': {'type': 'object',\n",
" 'required': ['institution', 'degree', 'start_date', 'end_date'],\n",
" 'properties': {'degree': {'type': 'string',\n",
" 'description': 'The degree of the candidate'},\n",
" 'end_date': {'anyOf': [{'type': 'string'}, {'type': 'null'}],\n",
" 'description': \"The end date of the candidate's education\"},\n",
" 'start_date': {'anyOf': [{'type': 'string'}, {'type': 'null'}],\n",
" 'description': \"The start date of the candidate's education\"},\n",
" 'institution': {'type': 'string',\n",
" 'description': 'The institution of the candidate'}},\n",
" 'additionalProperties': False},\n",
" 'description': \"The candidate's education\"},\n",
" 'experience': {'type': 'array',\n",
" 'items': {'type': 'object',\n",
" 'required': ['company', 'title', 'description', 'start_date', 'end_date'],\n",
" 'properties': {'title': {'type': 'string',\n",
" 'description': 'The title of the candidate'},\n",
" 'company': {'type': 'string', 'description': 'The name of the company'},\n",
" 'end_date': {'anyOf': [{'type': 'string'}, {'type': 'null'}],\n",
" 'description': \"The end date of the candidate's experience\"},\n",
" 'start_date': {'anyOf': [{'type': 'string'}, {'type': 'null'}],\n",
" 'description': \"The start date of the candidate's experience\"},\n",
" 'description': {'anyOf': [{'type': 'string'}, {'type': 'null'}],\n",
" 'description': \"The description of the candidate's experience\"}},\n",
" 'additionalProperties': False},\n",
" 'description': \"The candidate's experience\"},\n",
" 'technical_skills': {'type': 'object',\n",
" 'required': ['programming_languages', 'frameworks', 'skills'],\n",
" 'properties': {'skills': {'type': 'array',\n",
" 'items': {'type': 'string'},\n",
" 'description': 'Other general skills the candidate is proficient in, e.g. Data Engineering, Machine Learning, etc.'},\n",
" 'frameworks': {'type': 'array',\n",
" 'items': {'type': 'string'},\n",
" 'description': 'The tools/frameworks the candidate is proficient in, e.g. React, Django, PyTorch, etc.'},\n",
" 'programming_languages': {'type': 'array',\n",
" 'items': {'type': 'string'},\n",
" 'description': 'The programming languages the candidate is proficient in.'}},\n",
" 'description': \"The candidate's technical skills\",\n",
" 'additionalProperties': False},\n",
" 'key_accomplishments': {'type': 'string',\n",
" 'description': 'Summarize the candidates highest achievements.'}},\n",
" 'additionalProperties': False}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"agent = llama_extract.get_agent(\"resume-screening\")\n",
"agent.data_schema # Latest schema should be returned"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"#### Queueing extractions"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"For multiple resumes, we can use the `queue_extraction` method to run extractions asynchronously. This is ideal for processing batch extraction jobs."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"Uploading files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:01<00:00, 2.13it/s]\n",
"Creating extraction jobs: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:00<00:00, 5.83it/s]\n"
]
}
],
"source": [
"import os\n",
"\n",
"# All resumes in the data/resumes directory\n",
"resumes = []\n",
"\n",
"with os.scandir(\"./data/resumes\") as entries:\n",
" for entry in entries:\n",
" if entry.is_file():\n",
" resumes.append(entry.path)\n",
"\n",
"jobs = await agent.queue_extraction(resumes)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"To get the latest status of the extractions for any `job_id`, we can use the `get_extraction_job` method. \n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[<StatusEnum.PENDING: 'PENDING'>,\n",
" <StatusEnum.PENDING: 'PENDING'>,\n",
" <StatusEnum.PENDING: 'PENDING'>]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"[agent.get_extraction_job(job_id=job.id).status for job in jobs]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We notice that all extraction runs are in a PENDING state. We can check back again to see if the extractions have completed. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[<StatusEnum.SUCCESS: 'SUCCESS'>,\n",
" <StatusEnum.SUCCESS: 'SUCCESS'>,\n",
" <StatusEnum.SUCCESS: 'SUCCESS'>]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"[agent.get_extraction_job(job_id=job.id).status for job in jobs]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"#### Retrieving results\n",
"\n",
"Let us now retrieve the results of the extractions. If the status of the extraction is `SUCCESS`, we can retrieve the data from the `data` field. In case there are errors (status = `ERROR`), we can retrieve the error message from the `error` field. \n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"results = []\n",
"for job in jobs:\n",
" extract_run = agent.get_extraction_run_for_job(job.id)\n",
" if extract_run.status == \"SUCCESS\":\n",
" results.append(extract_run.data)\n",
" else:\n",
" print(f\"Extraction status for job {job.id}: {extract_run.status}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Dr. Rachel Zhang, Ph.D.',\n",
" 'email': 'rachel.zhang@email.com',\n",
" 'links': ['linkedin.com/in/rachelzhang',\n",
" 'github.com/rzhang-ai',\n",
" 'scholar.google.com/rachelzhang'],\n",
" 'education': [{'degree': 'Ph.D. in Computer Science',\n",
" 'end_date': '2011',\n",
" 'start_date': '2007',\n",
" 'institution': 'Columbia University'},\n",
" {'degree': 'M.S. in Computer Science',\n",
" 'end_date': '2007',\n",
" 'start_date': '2005',\n",
" 'institution': 'Stanford University'}],\n",
" 'experience': [{'title': 'Senior Research Scientist',\n",
" 'company': 'DeepMind',\n",
" 'end_date': None,\n",
" 'start_date': '2019',\n",
" 'description': '- Lead researcher on large-scale multi-task learning systems, developing novel architectures that improve cross-task generalization by 40%\\n- Pioneered new approach to zero-shot learning using contrastive training, published in NeurIPS 2023\\n- Built and led team of 6 researchers working on foundational ML models\\n- Developed novel regularization techniques for large language models, reducing catastrophic forgetting by 35%'},\n",
" {'title': 'Research Scientist',\n",
" 'company': 'Google Research',\n",
" 'end_date': '2019',\n",
" 'start_date': '2015',\n",
" 'description': '- Developed probabilistic frameworks for robust ML, published in ICML 2018\\n- Created novel attention mechanisms for computer vision models, improving accuracy by 25%\\n- Led collaboration with Google Brain team on efficient training methods for transformer models\\n- Mentored 4 PhD interns and collaborated with academic institutions'},\n",
" {'title': 'Research Assistant Professor',\n",
" 'company': 'Columbia University',\n",
" 'end_date': '2015',\n",
" 'start_date': '2011',\n",
" 'description': '- Published seminal work on Bayesian optimization methods (cited 1000+ times)\\n- Taught graduate-level courses in Machine Learning and Statistical Learning Theory\\n- Supervised 5 PhD students and 3 MSc students\\n- Secured $500K in research grants for probabilistic ML research'}],\n",
" 'technical_skills': {'skills': ['Deep Learning',\n",
" 'Reinforcement Learning',\n",
" 'Probabilistic Models',\n",
" 'Multi-Task Learning',\n",
" 'Zero-Shot Learning',\n",
" 'Neural Architecture Search'],\n",
" 'frameworks': ['PyTorch', 'TensorFlow', 'JAX', 'Ray'],\n",
" 'programming_languages': ['Python', 'C++', 'Julia', 'CUDA']},\n",
" 'key_accomplishments': 'AI researcher with 12+ years of experience spanning classical machine learning, deep learning, and probabilistic modeling. Led groundbreaking research in reinforcement learning, generative models, and multi-task learning. Published 25+ papers in top-tier conferences (NeurIPS, ICML, ICLR). Strong track record of transitioning theoretical advances into practical applications in both academic and industrial settings.'}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"results[0]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Alex Park',\n",
" 'email': 'alex park@email.com',\n",
" 'links': ['linkedin.com/in/alexpark'],\n",
" 'education': [{'degree': 'M.S. Computer Science',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'institution': 'University of California, Berkeley'},\n",
" {'degree': 'B.S. Computer Science',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'institution': 'University of California, Berkeley'}],\n",
" 'experience': [{'title': 'Senior Machine Learning Engineer',\n",
" 'company': 'SearchTech AI',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'description': 'Led development of next-generation learning-to-rank system using BER\\nArchitected and deployed real-time personalization system processing 10\\nIncreasing CTR by 15%\\nImproving search relevance by 24% (NDCG@10)'},\n",
" {'title': '',\n",
" 'company': 'Commerce Corp',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'description': 'Developed semantic search system using transformer models and approximate nearest neighbors, reducing null search results by 35%'},\n",
" {'title': 'Machine Learning Engineer',\n",
" 'company': 'Tech Solutions Inc',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'description': 'Implemented query understanding pipeline'},\n",
" {'title': 'Software Engineer',\n",
" 'company': '',\n",
" 'end_date': None,\n",
" 'start_date': None,\n",
" 'description': 'Built data pipelines and Flasticsearch'}],\n",
" 'technical_skills': {'skills': ['Elasticsearch',\n",
" 'Solr',\n",
" 'Lucene',\n",
" 'Python',\n",
" 'SQL',\n",
" 'Java',\n",
" 'Scala',\n",
" 'Shell Scripting'],\n",
" 'frameworks': ['PyTorch',\n",
" 'TensorFlow',\n",
" 'Scikit-learn',\n",
" 'BERT',\n",
" 'Word2Vec',\n",
" 'FastAI',\n",
" 'BM25',\n",
" 'FAISS',\n",
" 'Docker',\n",
" 'Kubernetes'],\n",
" 'programming_languages': []},\n",
" 'key_accomplishments': 'Machine Learning Engineer with 5 years of experience building and deploying large-scale search and relevance systems: Specialized in developing personalized search algorithms, learning-to-rank models; and recommendation systems. Strong track record of improving search relevance metrics and user engagement through ML-driven solutions:'}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"results[1]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'name': 'Sarah Chen',\n",
" 'email': 'sarah.chen@email.com',\n",
" 'links': [],\n",
" 'education': [{'degree': 'Master of Science in Computer Science',\n",
" 'end_date': '2013',\n",
" 'start_date': None,\n",
" 'institution': 'Stanford University'},\n",
" {'degree': 'Bachelor of Science in Computer Engineering',\n",
" 'end_date': '2011',\n",
" 'start_date': None,\n",
" 'institution': 'University of California, Berkeley'}],\n",
" 'experience': [{'title': 'Senior Software Architect',\n",
" 'company': 'TechCorp Solutions',\n",
" 'end_date': None,\n",
" 'start_date': '2020',\n",
" 'description': '- Led architectural design and implementation of a cloud-native platform serving 2M+ users\\n- Established architectural guidelines and best practices adopted across 12 development teams\\n- Reduced system latency by 40% through implementation of event-driven architecture\\n- Mentored 15+ senior developers in cloud-native development practices'},\n",
" {'title': 'Lead Software Engineer',\n",
" 'company': 'DataFlow Systems',\n",
" 'end_date': '2020',\n",
" 'start_date': '2016',\n",
" 'description': '- Architected and led development of distributed data processing platform handling 5TB daily\\n- Designed microservices architecture reducing deployment time by 65%\\n- Led migration of legacy monolith to cloud-native architecture\\n- Managed team of 8 engineers across 3 international locations'},\n",
" {'title': 'Senior Software Engineer',\n",
" 'company': 'InnovateTech',\n",
" 'end_date': '2016',\n",
" 'start_date': '2013',\n",
" 'description': '- Developed high-performance trading platform processing 100K transactions per second\\n- Implemented real-time analytics engine reducing processing latency by 75%\\n- Led adoption of container orchestration reducing deployment costs by 35%'}],\n",
" 'technical_skills': {'skills': ['Architecture & Design',\n",
" 'Microservices',\n",
" 'Event-Driven Architecture',\n",
" 'Domain-Driven Design',\n",
" 'REST APIs',\n",
" 'Cloud Platforms'],\n",
" 'frameworks': ['AWS (Advanced)', 'Azure', 'Google Cloud Platform'],\n",
" 'programming_languages': ['Java', 'Python', 'Go', 'JavaScript/TypeScript']},\n",
" 'key_accomplishments': '- Co-inventor on three patents for distributed systems architecture\\n- Published paper on \"Scalable Microservices Architecture\" at IEEE Cloud Computing Conference 2022\\n- Keynote Speaker, CloudCon 2023: \"Future of Cloud-Native Architecture\"\\n- Regular presenter at local tech meetups and conferences'}"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"results[2]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Congratulations! You now have an agent that can extract structured data from resumes. \n",
"- You can now use this agent to extract data from more resumes and use the extracted data for further processing. \n",
"- To update the schema, you can simply update the `data_schema` attribute of the agent and re-run the extraction. \n",
"- You can also use the `save` method to save the state of the agent and persist changes to the schema for future use. \n",
"\n"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
File diff suppressed because it is too large Load Diff
@@ -1,450 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "00f6713b-2a32-4f8f-80e5-9a7d9b6e3b90",
"metadata": {},
"source": [
"# Solar Panel Datasheet Comparison Workflow\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/solar_panel_e2e_comparison.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"\n",
"This notebook demonstrates an endtoend agentic workflow using LlamaExtract and the LlamaIndex eventdriven workflow framework. In this workflow, we:\n",
"\n",
"1. **Extract** structured technical specifications from a solar panel datasheet (e.g. a PDF downloaded from a vendor).\n",
"2. **Load** design requirements (provided as a text blob) for a labgrade solar panel.\n",
"3. **Generate** a detailed comparison report by triggering an event that injects both the extracted data and the requirements into an LLM prompt.\n",
"\n",
"The workflow is designed for renewable energy engineers who need to quickly validate that a solar panel meets specific design criteria.\n",
"\n",
"The following notebook uses the eventdriven syntax (with custom events, steps, and a workflow class) adapted from the technical datasheet and contract review examples."
]
},
{
"cell_type": "markdown",
"id": "36d8e34e-ed98-46ac-b744-1642f6e253d5",
"metadata": {},
"source": [
"## Setup and Load Data\n",
"\n",
"We download the [Honey M TSM-DE08M.08(II) datasheet](https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf) as a PDF.\n",
"\n",
"**NOTE**: The design requirements are already stored in `data/solar_panel_e2e_comparison/design_reqs.txt`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1de7b1b3-c285-492c-8b2e-b37974b4fc63",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2025-04-01 14:47:56-- https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf\n",
"Resolving static.trinasolar.com (static.trinasolar.com)... 47.246.23.232, 47.246.23.234, 47.246.23.227, ...\n",
"Connecting to static.trinasolar.com (static.trinasolar.com)|47.246.23.232|:443... connected.\n",
"WARNING: cannot verify static.trinasolar.com's certificate, issued by CN=DigiCert Global G2 TLS RSA SHA256 2020 CA1,O=DigiCert Inc,C=US:\n",
" Unable to locally verify the issuer's authority.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 1888183 (1.8M) [application/pdf]\n",
"Saving to: data/solar_panel_e2e_comparison/datasheet.pdf\n",
"\n",
"data/solar_panel_e2 100%[===================>] 1.80M 7.47MB/s in 0.2s \n",
"\n",
"2025-04-01 14:47:56 (7.47 MB/s) - data/solar_panel_e2e_comparison/datasheet.pdf saved [1888183/1888183]\n",
"\n"
]
}
],
"source": [
"!wget https://static.trinasolar.com/sites/default/files/EU_Datasheet_HoneyM_DE08M.08%28II%29_2021_A.pdf -O data/solar_panel_e2e_comparison/datasheet.pdf --no-check-certificate"
]
},
{
"cell_type": "markdown",
"id": "89d2f4c9-f785-424d-a409-3381796c457c",
"metadata": {},
"source": [
"## Define the Structured Extraction Schema\n",
"\n",
"We define a new, rich schema called `SolarPanelSchema` to capture key technical details from the datasheet. This schema includes:\n",
"\n",
"- **PowerRange:** Structured as minimum and maximum power output (in Watts).\n",
"- **SolarPanelSpec:** Includes module name, power output range, maximum efficiency, certifications, and a mapping of page citations.\n",
"\n",
"This schema replaces the earlier LM317 schema and will be used when creating our extraction agent."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "bfb40d48-36e0-4b1c-97a1-32a1704c582b",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"class PowerRange(BaseModel):\n",
" min_power: float = Field(..., description=\"Minimum power output in Watts\")\n",
" max_power: float = Field(..., description=\"Maximum power output in Watts\")\n",
" unit: str = Field(\"W\", description=\"Power unit\")\n",
"\n",
"\n",
"class SolarPanelSpec(BaseModel):\n",
" module_name: str = Field(..., description=\"Name or model of the solar panel module\")\n",
" power_output: PowerRange = Field(..., description=\"Power output range\")\n",
" maximum_efficiency: float = Field(\n",
" ..., description=\"Maximum module efficiency in percentage\"\n",
" )\n",
" temperature_coefficient: float = Field(\n",
" ..., description=\"Temperature coefficient in %/°C\"\n",
" )\n",
" certifications: List[str] = Field([], description=\"List of certifications\")\n",
" page_citations: dict = Field(\n",
" ..., description=\"Mapping of each extracted field to its page numbers\"\n",
" )\n",
"\n",
"\n",
"class SolarPanelSchema(BaseModel):\n",
" specs: List[SolarPanelSpec] = Field(\n",
" ..., description=\"List of extracted solar panel specifications\"\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "19dc309e-7cec-43c1-8f6c-72e14df58f8f",
"metadata": {},
"source": [
"## Initialize Extraction Agent\n",
"\n",
"Here we initialize our extraction agent that will be responsible for extracting the schema from the solar panel datasheet."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "c9d9f4a2-2e14-493d-8a7e-d01159d38b8f",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"from llama_cloud import ExtractConfig\n",
"\n",
"# Initialize the LlamaExtract client\n",
"llama_extract = LlamaExtract(\n",
" project_id=\"2fef999e-1073-40e6-aeb3-1f3c0e64d99b\",\n",
" organization_id=\"43b88c8f-e488-46f6-9013-698e3d2e374a\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "ec0eb2a7-6e02-45da-a6af-227e2f7c81f2",
"metadata": {},
"outputs": [],
"source": [
"try:\n",
" existing_agent = llama_extract.get_agent(name=\"solar-panel-datasheet\")\n",
" if existing_agent:\n",
" llama_extract.delete_agent(existing_agent.id)\n",
"except ApiError as e:\n",
" if e.status_code == 404:\n",
" pass\n",
" else:\n",
" raise\n",
"\n",
"extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
")\n",
"\n",
"agent = llama_extract.create_agent(\n",
" name=\"solar-panel-datasheet\", data_schema=SolarPanelSchema, config=extract_config\n",
")"
]
},
{
"cell_type": "markdown",
"id": "b4d7bb60-0456-4a2d-8d48-14f9bb3e71d2",
"metadata": {},
"source": [
"## Workflow Overview\n",
"\n",
"The workflow consists of four main steps:\n",
"\n",
"1. **parse_datasheet:** Reads the solar panel datasheet (PDF) and converts its content into text (with page citations).\n",
"2. **load_requirements:** Loads the design requirements (as a text blob) that will be injected into the prompt.\n",
"3. **generate_comparison_report:** Constructs a prompt using the extracted datasheet content and design requirements and triggers the LLM to generate a comparison report.\n",
"4. **output_result:** Logs and returns the final report as the workflows result.\n",
"\n",
"Each step is implemented as an asynchronous function decorated with `@step`, and the workflow is built by subclassing `Workflow`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7c482e3a-66b4-4e1b-8d2d-9a9c6b3967f3",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.workflow import (\n",
" Event,\n",
" StartEvent,\n",
" StopEvent,\n",
" Context,\n",
" Workflow,\n",
" step,\n",
")\n",
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.core.prompts import ChatPromptTemplate\n",
"from llama_cloud_services import LlamaExtract\n",
"from llama_cloud.core.api_error import ApiError\n",
"from pydantic import BaseModel, Field\n",
"from typing import List\n",
"\n",
"\n",
"# Define output schema for the comparison report (for reference)\n",
"class ComparisonReportOutput(BaseModel):\n",
" component_name: str = Field(\n",
" ..., description=\"The name of the component being evaluated.\"\n",
" )\n",
" meets_requirements: bool = Field(\n",
" ...,\n",
" description=\"Overall indicator of whether the component meets the design criteria.\",\n",
" )\n",
" summary: str = Field(..., description=\"A brief summary of the evaluation results.\")\n",
" details: dict = Field(\n",
" ..., description=\"Detailed comparisons for each key parameter.\"\n",
" )\n",
"\n",
"\n",
"# Define custom events\n",
"\n",
"\n",
"class DatasheetParseEvent(Event):\n",
" datasheet_content: dict\n",
"\n",
"\n",
"class RequirementsLoadEvent(Event):\n",
" requirements_text: str\n",
"\n",
"\n",
"class ComparisonReportEvent(Event):\n",
" report: ComparisonReportOutput\n",
"\n",
"\n",
"class LogEvent(Event):\n",
" msg: str\n",
" delta: bool = False\n",
"\n",
"\n",
"# For our demonstration, we assume that LlamaExtract is used to parse the datasheet into text.\n",
"# We'll also use OpenAI (via LlamaIndex) as our LLM for generating the report.\n",
"\n",
"llm = OpenAI(model=\"gpt-4o\") # or your preferred model"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "67a0c391-c7f5-4b93-8d6b-9e31b2d7a817",
"metadata": {},
"outputs": [],
"source": [
"class SolarPanelComparisonWorkflow(Workflow):\n",
" \"\"\"\n",
" Workflow to extract data from a solar panel datasheet and generate a comparison report\n",
" against provided design requirements.\n",
" \"\"\"\n",
"\n",
" def __init__(self, agent: LlamaExtract, requirements_path: str, **kwargs):\n",
" super().__init__(**kwargs)\n",
" self.agent = agent\n",
" # Load design requirements from file as a text blob\n",
" with open(requirements_path, \"r\") as f:\n",
" self.requirements_text = f.read()\n",
"\n",
" @step\n",
" async def parse_datasheet(\n",
" self, ctx: Context, ev: StartEvent\n",
" ) -> DatasheetParseEvent:\n",
" # datasheet_path is provided in the StartEvent\n",
" datasheet_path = (\n",
" ev.datasheet_path\n",
" ) # e.g., \"./data/solar_panel_comparison/datasheet.pdf\"\n",
" extraction_result = await self.agent.aextract(datasheet_path)\n",
" datasheet_dict = (\n",
" extraction_result.data\n",
" ) # assumed to be a string with page citations\n",
" await ctx.set(\"datasheet_content\", datasheet_dict)\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Datasheet parsed successfully.\"))\n",
" return DatasheetParseEvent(datasheet_content=datasheet_dict)\n",
"\n",
" @step\n",
" async def load_requirements(\n",
" self, ctx: Context, ev: DatasheetParseEvent\n",
" ) -> RequirementsLoadEvent:\n",
" # Use the pre-loaded requirements text from __init__\n",
" req_text = self.requirements_text\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Design requirements loaded.\"))\n",
" return RequirementsLoadEvent(requirements_text=req_text)\n",
"\n",
" @step\n",
" async def generate_comparison_report(\n",
" self, ctx: Context, ev: RequirementsLoadEvent\n",
" ) -> StopEvent:\n",
" # Build a prompt that injects both the extracted datasheet content and the design requirements\n",
" datasheet_content = await ctx.get(\"datasheet_content\")\n",
" prompt_str = \"\"\"\n",
"You are an expert renewable energy engineer.\n",
"\n",
"Compare the following solar panel datasheet information with the design requirements.\n",
"\n",
"Design Requirements:\n",
"{requirements_text}\n",
"\n",
"Extracted Datasheet Information:\n",
"{datasheet_content}\n",
"\n",
"Generate a detailed comparison report in JSON format with the following schema:\n",
" - component_name: string\n",
" - meets_requirements: boolean\n",
" - summary: string\n",
" - details: dictionary of comparisons for each parameter\n",
"\n",
"For each parameter (Maximum Power, Open-Circuit Voltage, Short-Circuit Current, Efficiency, Temperature Coefficient),\n",
"indicate PASS or FAIL and provide brief explanations and recommendations.\n",
"\"\"\"\n",
"\n",
" # extract from contract\n",
" prompt = ChatPromptTemplate.from_messages([(\"user\", prompt_str)])\n",
"\n",
" # Call the LLM to generate the report using the prompt\n",
" report_output = await llm.astructured_predict(\n",
" ComparisonReportOutput,\n",
" prompt,\n",
" requirements_text=ev.requirements_text,\n",
" datasheet_content=str(datasheet_content),\n",
" )\n",
" ctx.write_event_to_stream(LogEvent(msg=\"Comparison report generated.\"))\n",
" return StopEvent(\n",
" result={\"report\": report_output, \"datasheet_content\": datasheet_content}\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "d205f532-1a11-4a48-b5a8-87a7f85e9ce7",
"metadata": {},
"source": [
"## Running the Workflow\n",
"\n",
"Below, we instantiate and run the workflow. We inject the design requirements as a text blob (no custom code to load) and pass the path to the solar panel datasheet (the HoneyM datasheet from Trina).\n",
"\n",
"The design requirements are:\n",
"\n",
"```\n",
"Solar Panel Design Requirements:\n",
"- Power Output Range: ≥ 350 W\n",
"- Maximum Efficiency: ≥ 18%\n",
"- Certifications: Must include IEC61215 and UL1703\n",
"```\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6b24fa61-a2f5-4ebb-84eb-1c9b48683b1b",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "be3ebad5-1f70-4671-a2ec-17bf9e4d788f",
"metadata": {},
"outputs": [],
"source": [
"# Path to design requirements file (e.g., a text file with design criteria for solar panels)\n",
"requirements_path = \"./data/solar_panel_e2e_comparison/design_reqs.txt\"\n",
"\n",
"# Instantiate the workflow\n",
"workflow = SolarPanelComparisonWorkflow(\n",
" agent=agent, requirements_path=requirements_path, verbose=True, timeout=120\n",
")\n",
"\n",
"# Run the workflow; pass the datasheet path in the StartEvent\n",
"result = await workflow.run(\n",
" datasheet_path=\"./data/solar_panel_e2e_comparison/datasheet.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "e1e61f1e-8701-4acc-8f99-cc89d8aae535",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"********Final Comparison Report:********\n",
"\n",
"{\n",
" \"component_name\": \"TSM-DE08M.08(II)\",\n",
" \"meets_requirements\": true,\n",
" \"summary\": \"The solar panel TSM-DE08M.08(II) meets all the design requirements, making it a suitable choice for the intended application.\",\n",
" \"details\": {\n",
" \"Maximum Power Output\": \"PASS - The panel's power output ranges from 360 W to 385 W, exceeding the minimum requirement of 350 W.\",\n",
" \"Open-Circuit Voltage\": \"PASS - The datasheet does not specify Voc, but the panel meets other critical requirements. Verification of Voc is recommended.\",\n",
" \"Short-Circuit Current\": \"PASS - The datasheet does not specify Isc, but the panel meets other critical requirements. Verification of Isc is recommended.\",\n",
" \"Efficiency\": \"PASS - The panel's efficiency is 21.0%, which is above the required 18%.\",\n",
" \"Temperature Coefficient\": \"PASS - The temperature coefficient is -0.34%/°C, which is better than the maximum allowable -0.5%/°C.\"\n",
" }\n",
"}\n"
]
}
],
"source": [
"print(\"\\n********Final Comparison Report:********\\n\")\n",
"print(result[\"report\"].model_dump_json(indent=4))\n",
"# print(\"\\n********Datasheet Content:********\\n\", result[\"datasheet_content\"])"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama_parse",
"language": "python",
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
-131
View File
@@ -1,131 +0,0 @@
# LlamaCloud Index Demo
A TypeScript demo application showcasing the power of **LlamaCloud Index** - a fully automated document ingestion and retrieval serviced offered within [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to ask questions, retrieve relevant contextual information and generate AI-powered responses using OpenAI's GPT models.
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [Development Mode](#development-mode)
- [Build the Project](#build-the-project)
- [Code Quality](#code-quality)
- [Quick Commands Reference](#quick-commands-reference)
- [How It Works](#how-it-works)
- [API Dependencies](#api-dependencies)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 🤖 **RAG**: Simple-yet-effective Retrieval Augmented Generation pipeline built on top of LlamaCloud Index and OpenAI
- 🎨 **Beautiful CLI**: Styled console interface with colors and ASCII art
-**Fast Development**: Hot reload support with watch mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 18 or higher)
- pnpm package manager
- OpenAI API key
- LlamaCloud API key
- An existing LlamaCloud Index pipeline
## Installation
1. Clone the repository:
```bash
git clone <repository-url>
cd llamaparse-demo
```
2. Install dependencies:
```bash
pnpm install
```
3. Set up your environment variables:
```bash
export OPENAI_API_KEY="your-openai-api-key"
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
export PIPELINE_NAME="your-pipeline-name"
```
4. Or write them into a `.env` file:
```env
OPENAI_API_KEY="your-openai-api-key"
LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
PIPELINE_NAME="your-pipeline-name"
```
## Usage
### Start the Demo
```bash
pnpm run start
```
The application will display a welcome screen and prompt you to start chatting!
### Development Mode
For development with hot reload:
```bash
pnpm run dev
```
### Build the Project
```bash
pnpm run build
```
### Code Quality
Format code:
```bash
pnpm run format
```
Lint code:
```bash
pnpm run lint
```
## How It Works
1. **Message Input**: Enter a message
2. **Retrieval**: Several nodes are retrieved from the LlamaCloud index you specified
3. **AI Response Generation**: The retrieved information is passed on to the AI model, along with its relevance score, and a reply to your original message is generated starting from that.
4. **Results**: View the AI-generated summary in your terminal
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 18+ and have all dependencies installed
2. **API Key Issues**: Verify your OpenAI and LlamaCloud API keys are correctly set
## License
MIT License - see the [LICENSE](../../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `pnpm format` and `pnpm lint`
5. Submit a pull request
-15
View File
@@ -1,15 +0,0 @@
import js from "@eslint/js";
import globals from "globals";
import tseslint from "typescript-eslint";
import { defineConfig } from "eslint/config";
export default defineConfig([
{
files: ["**/*.{js,mjs,cjs,ts,mts,cts}"],
plugins: { js },
extends: ["js/recommended"],
languageOptions: { globals: globals.browser },
},
{ files: ["**/*.js"], languageOptions: { sourceType: "script" } },
tseslint.configs.recommended,
]);
-48
View File
@@ -1,48 +0,0 @@
{
"name": "llama-chat",
"version": "0.1.0",
"description": "Demo for LlamaCloud Index in TypeScript",
"type": "module",
"main": "index.js",
"scripts": {
"test": "echo \"There are no tests\"",
"start": "pnpm exec tsx src/index.ts",
"lint": "eslint ./src/",
"format": "prettier --write ./src/",
"build": "tsc",
"dev": "pnpm exec tsx --watch src/index.ts"
},
"keywords": [
"ai",
"rag",
"retrieval",
"pipeline",
"llms",
"chatbot"
],
"author": "LlamaIndex",
"license": "MIT",
"packageManager": "pnpm@10.12.4",
"devDependencies": {
"@eslint/js": "^9.32.0",
"@types/figlet": "^1.7.0",
"@types/node": "^24.1.0",
"@typescript-eslint/eslint-plugin": "^8.38.0",
"@typescript-eslint/parser": "^8.38.0",
"eslint": "^9.32.0",
"globals": "^16.3.0",
"jiti": "^2.5.1",
"prettier": "^3.6.2",
"typescript": "^5.8.3",
"typescript-eslint": "^8.38.0"
},
"dependencies": {
"@ai-sdk/openai": "^1.3.23",
"ai": "^4.3.19",
"consola": "^3.4.2",
"dotenv": "^17.2.1",
"figlet": "^1.8.2",
"llama-cloud-services": "link:../../../ts/llama_cloud_services",
"picocolors": "^1.1.1"
}
}
-1770
View File
File diff suppressed because it is too large Load Diff
-48
View File
@@ -1,48 +0,0 @@
import { LlamaCloudIndex } from "llama-cloud-services";
import { logger } from "./logger";
import pc from "picocolors";
import {
consoleInput,
retrievalAugmentedGeneration,
renderLogo,
} from "./utils";
import dotenv from "dotenv";
dotenv.config();
export async function main(): Promise<number> {
const index = new LlamaCloudIndex({
name: process.env.PIPELINE_NAME as string,
projectName: "Default",
apiKey: process.env.LLAMA_CLOUD_API_KEY, // can provide API-key in the constructor or in the env
});
const retriever = index.asRetriever({
similarityTopK: 5,
});
await renderLogo();
logger.log(
`Welcome to ${pc.bold(
pc.magentaBright("✨LlamaChat✨"),
)}, our demo for ${pc.bold(pc.green("Index🦙"))}, a ${pc.bold(
pc.cyan("LlamaCloud☁️"),
)} (https://cloud.llamaindex.ai) product!.\nType a question below, and you will get an answer!👇\nIf you wish to exit, just type ${pc.bold(
pc.gray("quit"),
)}.\n`,
);
while (true) {
const userInput = await consoleInput();
if (userInput.toLowerCase() == "quit") {
break;
}
try {
const nodes = await retriever.retrieve(userInput);
const summary = await retrievalAugmentedGeneration(nodes, userInput);
logger.log(`${pc.bold(pc.magentaBright("LlamaChat✨:"))}\n${summary}`);
} catch (error) {
logger.error(`Error processing your request: ${error}`);
}
}
return 0;
}
main().catch(console.error);
-8
View File
@@ -1,8 +0,0 @@
import { createConsola } from "consola";
import type { ConsolaInstance } from "consola";
export const logger: ConsolaInstance = createConsola({
formatOptions: {
date: false,
},
});
-56
View File
@@ -1,56 +0,0 @@
import { generateText } from "ai";
import { openai } from "@ai-sdk/openai";
import { NodeWithScore, MetadataMode } from "llamaindex";
import * as readline from "readline/promises";
import figlet from "figlet";
import pc from "picocolors";
export async function renderLogo(): Promise<void> {
const logoText = figlet.textSync("LlamaChat", {
font: "ANSI Shadow",
horizontalLayout: "default",
verticalLayout: "default",
width: 100,
whitespaceBreak: true,
});
// Add some styling with picocolors
const styledLogo = pc.bold(pc.yellowBright(logoText));
// Add some padding/margin
console.log("\n");
console.log(styledLogo);
console.log(pc.gray("─".repeat(60)));
console.log("\n");
}
export async function consoleInput(): Promise<string> {
const rl = readline.createInterface({
input: process.stdin,
output: process.stdout,
});
const answer = await rl.question(pc.cyanBright("You✨:"));
rl.close();
return answer;
}
export async function retrievalAugmentedGeneration(
nodes: NodeWithScore[],
prompt: string,
): Promise<string> {
let mainText: string = "";
for (const node of nodes) {
mainText += `\t{information: '${node.node.getContent(
MetadataMode.ALL,
)}', relevanceScore: '${node.score ?? "no score"}'}\n`;
}
const { text } = await generateText({
model: openai("gpt-4.1"),
prompt: `[\n${mainText}\n]\n\nBased on the information you are given and on the relevance score of that (where -1 means no score available), answer to this user prompt: '${prompt}'`,
});
return text;
}
-22
View File
@@ -1,22 +0,0 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "ES2022",
"lib": ["ES2022"],
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"declaration": true,
"declarationMap": true,
"sourceMap": true,
"types": ["node"],
"moduleResolution": "bundler",
"allowSyntheticDefaultImports": true,
"resolveJsonModule": true
},
"include": ["src/**/*"],
"exclude": ["node_modules", "dist"]
}
@@ -7,7 +7,7 @@
"source": [
"# Knowledge Graph Agent with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/knowledge_graphs/kg_agent.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/knowledge_graphs/kg_agent.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"Here we build a knowledge graph agent over the SF 2023 Budget Proposal. We use LlamaIndex abstractions to construct a knowledge graph, and we store the property graph in neo4j. We then build an agent that can interact with the knowledge graph as a tool."
]
@@ -33,7 +33,7 @@
"!pip install llama-index-postprocessor-flag-embedding-reranker\n",
"!pip install git+https://github.com/FlagOpen/FlagEmbedding.git\n",
"!pip install llama-index-graph-stores-neo4j\n",
"!pip install llama-cloud-services"
"!pip install llama-parse"
]
},
{
@@ -125,7 +125,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"docs = LlamaParse(result_type=\"text\").load_data(\"./data/budget_2023.pdf\")"
]

Before

Width:  |  Height:  |  Size: 334 KiB

After

Width:  |  Height:  |  Size: 334 KiB

@@ -7,7 +7,7 @@
"source": [
"# Multimodal Parsing using Anthropic Claude (Sonnet 3.5)\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/claude_parse.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/claude_parse.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This cookbook shows you how to use LlamaParse to parse any document with the multimodal capabilities of Sonnet 3.5. \n",
"\n",
@@ -141,7 +141,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -205,7 +205,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -7,7 +7,7 @@
"source": [
"# Multimodal Parsing using GPT4o-mini\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/gpt4o_mini.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/gpt4o_mini.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This cookbook shows you how to use LlamaParse to parse any document with the multimodal capabilities of GPT4o-mini.\n",
"\n",
@@ -118,7 +118,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -181,7 +181,7 @@
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -6,7 +6,7 @@
"source": [
"# Building a Multimodal RAG Pipeline over an Auto Insurance Claim\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/insurance_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/insurance_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
@@ -99,7 +99,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -6,7 +6,7 @@
"source": [
"# Building a RAG Pipeline over Legal Documents\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/legal_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/legal_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This example shows how LlamaParse and LlamaIndex can be used to parse various types of legal documents, which may contain complex tabular data. The advantage of this is being able to quickly retrieve a specific answer to a legal question with comprehensive context — knowledge of precedents, statutes, and cases presented in the given documents. A user can quickly find the answer to or find out more details about a specific legal question without having to read through the often long documents by using LLMs.\n",
"\n",
@@ -102,7 +102,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",

Before

Width:  |  Height:  |  Size: 1.2 MiB

After

Width:  |  Height:  |  Size: 1.2 MiB

Before

Width:  |  Height:  |  Size: 170 KiB

After

Width:  |  Height:  |  Size: 170 KiB

@@ -7,7 +7,7 @@
"source": [
"# Building a Natively Multimodal RAG Pipeline (over a Slide Deck)\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/multimodal_rag_slide_deck.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/multimodal_rag_slide_deck.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this cookbook we show you how to build a multimodal RAG pipeline over a slide deck, with text, tables, images, diagrams, and complex layouts.\n",
"\n",
@@ -153,7 +153,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"\n",
"parser_text = LlamaParse(result_type=\"text\")\n",
@@ -165,18 +165,7 @@
"execution_count": null,
"id": "ef82a985-4088-4bb7-9a21-0318e1b9207d",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Parsing text...\n",
"Started parsing the file under job_id 62f157a9-9ef9-4e5b-95ac-67093fa25800\n",
"..........Parsing PDF file...\n",
"Started parsing the file under job_id 1ddd5654-062b-4e19-b488-d66efc9c509d\n"
]
}
],
"outputs": [],
"source": [
"print(f\"Parsing text...\")\n",
"docs_text = parser_text.load_data(\"data/conocophillips.pdf\")\n",
@@ -185,36 +174,42 @@
"md_json_list = md_json_objs[0][\"pages\"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7506b603-c01f-45de-b354-4a0728dde03c",
"metadata": {},
"outputs": [],
"source": [
"print(docs_text[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "5318fb7b-fe6a-4a8a-b82e-4ed7b4512c37",
"metadata": {},
"outputs": [],
"source": [
"print(md_json_list[10][\"md\"])"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7a46a73e-c6e2-4b0b-bd10-31b0d3e4b70f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Commitment to Disciplined Reinvestment Rate\n",
"\n",
"| Period | Description | Reinvestment Rate | WTI Average |\n",
"|--------------|--------------------------------------|-------------------|-------------|\n",
"| 2012-2016 | Industry Growth Focus | >100% | ~$75/BBL |\n",
"| 2017-2022 | ConocoPhillips Strategy Reset | <60% | ~$63/BBL |\n",
"| 2023E | | | at $80/BBL |\n",
"| 2024-2028 | Disciplined Reinvestment Rate | ~50% | at $60/BBL |\n",
"| 2029-2032 | | ~6% CFO CAGR | at $60/BBL |\n",
"\n",
"- **Historic Reinvestment Rate**: Gray bars\n",
"- **Reinvestment Rate at $60/BBL WTI**: Blue bars\n",
"- **Reinvestment Rate at $80/BBL WTI**: Dashed blue lines\n",
"\n",
"Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n"
"dict_keys(['page', 'text', 'md', 'images', 'items'])\n"
]
}
],
"source": [
"print(md_json_list[10][\"md\"])"
"print(md_json_list[1].keys())"
]
},
{
@@ -304,7 +299,7 @@
" image_files = _get_sorted_image_files(image_dir) if image_dir is not None else None\n",
" md_texts = [d[\"md\"] for d in json_dicts] if json_dicts is not None else None\n",
"\n",
" doc_chunks = [c for d in docs for c in d.text.split(\"---\")]\n",
" doc_chunks = docs[0].text.split(\"---\")\n",
" for idx, doc_chunk in enumerate(doc_chunks):\n",
" chunk_metadata = {\"page_num\": idx + 1}\n",
" if image_files is not None:\n",
@@ -344,23 +339,25 @@
"output_type": "stream",
"text": [
"page_num: 11\n",
"image_path: data_images/1ddd5654-062b-4e19-b488-d66efc9c509d-page_39.jpg\n",
"image_path: data_images/d9137e19-3974-4b5d-998f-dac0cf29dd9d-page-10.jpg\n",
"parsed_text_markdown: # Commitment to Disciplined Reinvestment Rate\n",
"\n",
"| Period | Description | Reinvestment Rate | WTI Average |\n",
"|--------------|--------------------------------------|-------------------|-------------|\n",
"| 2012-2016 | Industry Growth Focus | >100% | ~$75/BBL |\n",
"| 2017-2022 | ConocoPhillips Strategy Reset | <60% | ~$63/BBL |\n",
"| 2023E | | | at $80/BBL |\n",
"| 2024-2028 | Disciplined Reinvestment Rate | ~50% | at $60/BBL |\n",
"| 2029-2032 | | ~6% CFO CAGR | at $60/BBL |\n",
"| Year | Reinvestment Rate | WTI Average Price | Reinvestment Rate at $60/BBL WTI | Reinvestment Rate at $80/BBL WTI |\n",
"|------------|-------------------|-------------------|----------------------------------|----------------------------------|\n",
"| 2012-2016 | >100% | ~$75/BBL | | |\n",
"| 2017-2022 | <60% | ~$63/BBL | | |\n",
"| 2023E | | | | at $80/BBL WTI |\n",
"| 2024-2028 | | | at $60/BBL WTI | at $80/BBL WTI |\n",
"| 2029-2032 | | | at $60/BBL WTI | at $80/BBL WTI |\n",
"\n",
"- **Historic Reinvestment Rate**: Gray bars\n",
"- **Reinvestment Rate at $60/BBL WTI**: Blue bars\n",
"- **Reinvestment Rate at $80/BBL WTI**: Dashed blue lines\n",
"**Disciplined Reinvestment Rate is the Foundation for Superior Returns on and of Capital, while Driving Durable CFO Growth**\n",
"\n",
"Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n",
"parsed_text: Commitment to Disciplined Reinvestment Rate\n",
"- ~50% 10-Year Reinvestment Rate\n",
"- ~6% CFO CAGR 2024-2032 at $60/BBL WTI Mid-Cycle Planning Price\n",
"\n",
"**Note:** Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n",
"parsed_text: \n",
"Commitment to Disciplined Reinvestment Rate\n",
" Industry ConocoPhillips\n",
" Strategy Reset Disciplined Reinvestment Rate is the Foundation for Superior\n",
" Growth Focus Returns on and of Capital, while Driving Durable CFO Growth\n",
@@ -377,7 +374,7 @@
" 0%\n",
" 2012-2016 2017-2022 2023E 2024-2028 2029-2032\n",
" Historic Reinvestment Rate Reinvestment Rate at $60/BBL WTI Reinvestment Rate at $80/BBL WTI\n",
" Reinvestment rate and cash from operations (CFO) are non-GAAP measures: Definitions and reconciliations are included in the Appendix ConocoPhillips\n"
" Reinvestment rate andcashfrom operations (CFO) are non-GAAP measures: Definitions and reconciliations are included in the Appendix ConocoPhillips\n"
]
}
],
@@ -400,17 +397,7 @@
"execution_count": null,
"id": "6ea53c31-0e38-421c-8d9b-0e3adaa1677e",
"metadata": {},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"/Users/jerryliu/Programming/gpt_index/.venv/lib/python3.10/site-packages/tiktoken/core.py:50: RuntimeWarning: coroutine 'LlamaParse.aload_data' was never awaited\n",
" self._core_bpe = _tiktoken.CoreBPE(mergeable_ranks, special_tokens, pat_str)\n",
"RuntimeWarning: Enable tracemalloc to get the object allocation traceback\n"
]
}
],
"outputs": [],
"source": [
"import os\n",
"from llama_index.core import (\n",
@@ -601,7 +588,7 @@
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3 $40 GKA GWA\n",
" 3$40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
@@ -612,7 +599,7 @@
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Resource (BBOE) Eagle Ford Other MalaysiaChina Surmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
@@ -700,126 +687,70 @@
{
"cell_type": "code",
"execution_count": null,
"id": "d78e53cf-35cb-4ef8-b03e-1b47ba15ae64",
"id": "1cdce5d8-6bb3-4cd3-929d-1cec249d9052",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Added user message to memory: Tell me about the diverse geographies where Conoco Phillips has a production base\n",
"Added user message to memory: How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips production base geographies\"}\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"ConocoPhillips' production base geographies include:\n",
"The ConocoPhillips capex/EUR in the Delaware Basin is $10/BOE.\n",
"\n",
"1. **Lower 48** (Permian, Eagle Ford, Bakken, Other)\n",
"2. **Alaska** (GKA, GWA, GPA, WNS)\n",
"3. **EMENA** (Norway, Libya, Qatar)\n",
"4. **Asia Pacific** (APLNG, Malaysia, China)\n",
"5. **Canada** (Montney, Surmont)\n",
"\n",
"This information was derived from the image on page 14, which provides a detailed breakdown of the diverse production base and the regions involved. The parsed markdown and raw text also support this information, but the image provides the clearest and most comprehensive view. There are no discrepancies between the image and the parsed text in this case.\n",
"=== LLM Response ===\n",
"ConocoPhillips has a diverse production base spread across various geographies, including:\n",
"\n",
"1. **Lower 48**:\n",
" - Permian Basin\n",
" - Eagle Ford\n",
" - Bakken\n",
" - Other regions within the continental United States\n",
"\n",
"2. **Alaska**:\n",
" - Greater Kuparuk Area (GKA)\n",
" - Greater Prudhoe Area (GPA)\n",
" - Greater Willow Area (GWA)\n",
" - Western North Slope (WNS)\n",
"\n",
"3. **EMENA (Europe, Middle East, and North Africa)**:\n",
" - Norway\n",
" - Libya\n",
" - Qatar\n",
"\n",
"4. **Asia Pacific**:\n",
" - Australia Pacific LNG (APLNG)\n",
" - Malaysia\n",
" - China\n",
"\n",
"5. **Canada**:\n",
" - Montney\n",
" - Surmont\n",
"\n",
"These regions highlight the global reach and diverse geographical footprint of ConocoPhillips' production operations.\n",
"Added user message to memory: Tell me about the diverse geographies where Conoco Phillips has a production base\n",
"I obtained this information from the image provided. The image clearly shows a bar chart under the section \"Delaware Basin Well Capex/EUR ($/BOE)\" where ConocoPhillips is listed with a capex/EUR of $10/BOE. This information is consistent with the parsed markdown text, which also lists ConocoPhillips' capex/EUR as $10/BOE in the Delaware Basin. There are no discrepancies between the image and the parsed markdown text in this case.\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"diverse geographies where Conoco Phillips has a production base\"}\n",
"Calling function: vector_tool with args: {\"input\": \"competitors capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"ConocoPhillips has a diverse production base that includes the Lower 48 (Permian, Bakken, Eagle Ford), Alaska, Canada (Montney, Surmont), EMENA (Norway, Libya), Asia Pacific (Malaysia, China, APLNG), and Qatar.\n",
"The competitors' Capex/EUR in the Delaware Basin can be found in the image on the slide titled \"Delaware: Vast Inventory with Proven Track Record of Performance.\" The relevant information is presented in a bar chart under the section \"Delaware Basin Well Capex/EUR ($/BOE)\".\n",
"\n",
"Here are the details:\n",
"\n",
"- ConocoPhillips: $10/BOE\n",
"- Competitor 1: $15/BOE\n",
"- Competitor 2: $20/BOE\n",
"- Competitor 3: $25/BOE\n",
"- Competitor 4: $30/BOE\n",
"- Competitor 5: $35/BOE\n",
"- Competitor 6: $40/BOE\n",
"- Competitor 7: $45/BOE\n",
"\n",
"This information was obtained directly from the image, which provides a clear visual representation of the Capex/EUR values for ConocoPhillips and its competitors in the Delaware Basin. The parsed markdown text also confirms these values, ensuring consistency between the image and the text.\n",
"=== LLM Response ===\n",
"ConocoPhillips has a diverse production base spanning several key geographies:\n",
"The capital expenditure per estimated ultimate recovery (capex/EUR) for ConocoPhillips in the Delaware Basin is $10 per barrel of oil equivalent (BOE). When compared to its competitors, ConocoPhillips has a significantly lower capex/EUR. Here are the capex/EUR values for ConocoPhillips and its competitors:\n",
"\n",
"1. **Lower 48 (United States)**: This includes major production areas such as the Permian Basin, Bakken Formation, and Eagle Ford Shale.\n",
"2. **Alaska**: Significant operations in the North Slope region.\n",
"3. **Canada**: Operations in the Montney Formation and the Surmont oil sands project.\n",
"4. **EMENA (Europe, Middle East, and North Africa)**: Notable operations in Norway and Libya.\n",
"5. **Asia Pacific**: Includes operations in Malaysia, China, and the Australia Pacific LNG (APLNG) project.\n",
"6. **Qatar**: Involvement in the country's energy sector.\n",
"- **ConocoPhillips**: $10/BOE\n",
"- **Competitor 1**: $15/BOE\n",
"- **Competitor 2**: $20/BOE\n",
"- **Competitor 3**: $25/BOE\n",
"- **Competitor 4**: $30/BOE\n",
"- **Competitor 5**: $35/BOE\n",
"- **Competitor 6**: $40/BOE\n",
"- **Competitor 7**: $45/BOE\n",
"\n",
"These regions highlight the company's extensive and varied geographical footprint in the energy production industry.\n"
"This data indicates that ConocoPhillips has a more cost-efficient operation in the Delaware Basin compared to its competitors.\n",
"The capital expenditure per estimated ultimate recovery (capex/EUR) for ConocoPhillips in the Delaware Basin is $10 per barrel of oil equivalent (BOE). When compared to its competitors, ConocoPhillips has a significantly lower capex/EUR. Here are the capex/EUR values for ConocoPhillips and its competitors:\n",
"\n",
"- **ConocoPhillips**: $10/BOE\n",
"- **Competitor 1**: $15/BOE\n",
"- **Competitor 2**: $20/BOE\n",
"- **Competitor 3**: $25/BOE\n",
"- **Competitor 4**: $30/BOE\n",
"- **Competitor 5**: $35/BOE\n",
"- **Competitor 6**: $40/BOE\n",
"- **Competitor 7**: $45/BOE\n",
"\n",
"This data indicates that ConocoPhillips has a more cost-efficient operation in the Delaware Basin compared to its competitors.\n"
]
}
],
"source": [
"query = (\n",
" \"Tell me about the diverse geographies where Conoco Phillips has a production base\"\n",
"# response = agent.query(\"Tell me about the different regions and subregions where Conoco Phillips has a production base.\")\n",
"response = agent.query(\n",
" \"How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\"\n",
")\n",
"response = agent.query(query)\n",
"base_response = base_agent.query(query)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "355d2aa4-c26f-480e-b512-4446acbd9227",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"ConocoPhillips has a diverse production base spread across various geographies, including:\n",
"\n",
"1. **Lower 48**:\n",
" - Permian Basin\n",
" - Eagle Ford\n",
" - Bakken\n",
" - Other regions within the continental United States\n",
"\n",
"2. **Alaska**:\n",
" - Greater Kuparuk Area (GKA)\n",
" - Greater Prudhoe Area (GPA)\n",
" - Greater Willow Area (GWA)\n",
" - Western North Slope (WNS)\n",
"\n",
"3. **EMENA (Europe, Middle East, and North Africa)**:\n",
" - Norway\n",
" - Libya\n",
" - Qatar\n",
"\n",
"4. **Asia Pacific**:\n",
" - Australia Pacific LNG (APLNG)\n",
" - Malaysia\n",
" - China\n",
"\n",
"5. **Canada**:\n",
" - Montney\n",
" - Surmont\n",
"\n",
"These regions highlight the global reach and diverse geographical footprint of ConocoPhillips' production operations.\n"
]
}
],
"source": [
"print(str(response))"
]
},
@@ -833,82 +764,85 @@
"name": "stdout",
"output_type": "stream",
"text": [
"page_num: 14\n",
"image_path: data_images/1ddd5654-062b-4e19-b488-d66efc9c509d-page_12.jpg\n",
"parsed_text_markdown: # Our Differentiated Portfolio: Deep, Durable and Diverse\n",
"page_num: 38\n",
"image_path: data_images/d9137e19-3974-4b5d-998f-dac0cf29dd9d-page-37.jpg\n",
"parsed_text_markdown: # Delaware: Vast Inventory with Proven Track Record of Performance\n",
"\n",
"## ~20 BBOE of Resource\n",
"Under $40/BBL Cost of Supply\n",
"## Prolific Acreage Spanning Over ~659,000 Net Acres¹\n",
"\n",
"### ~ $32/BBL\n",
"Average Cost of Supply\n",
"![Map of Delaware Basin](image)\n",
"\n",
"### WTI Cost of Supply ($/BBL)\n",
"### Total 10-Year Operated Permian Inventory\n",
"\n",
"| Cost ($/BBL) | Resource (BBOE) |\n",
"|--------------|-----------------|\n",
"| $0 | 0 |\n",
"| $10 | |\n",
"| $20 | |\n",
"| $30 | |\n",
"| $40 | |\n",
"| $50 | |\n",
"- Delaware Basin: 65%\n",
"- Midland Basin: 35%\n",
"\n",
"- **Legend:**\n",
" - Lower 48\n",
" - Canada\n",
" - Alaska\n",
" - EMENA\n",
" - Asia Pacific\n",
"### High Single-Digit Production Growth\n",
"\n",
"*Costs assume a mid-cycle price environment of $60/BBL WTI.*\n",
"## 12-Month Cumulative Production³ (BOE/FT)\n",
"\n",
"## Diverse Production Base\n",
"10-Year Plan Cumulative Production (BBOE)\n",
"| Months | 2019 | 2020 | 2021 | 2022 |\n",
"|--------|------|------|------|------|\n",
"| 1 | 0 | 0 | 0 | 0 |\n",
"| 2 | 5 | 6 | 7 | 8 |\n",
"| 3 | 10 | 12 | 14 | 16 |\n",
"| 4 | 15 | 18 | 21 | 24 |\n",
"| 5 | 20 | 24 | 28 | 32 |\n",
"| 6 | 25 | 30 | 35 | 40 |\n",
"| 7 | 30 | 36 | 42 | 48 |\n",
"| 8 | 35 | 42 | 49 | 56 |\n",
"| 9 | 40 | 48 | 56 | 64 |\n",
"| 10 | 45 | 54 | 63 | 72 |\n",
"| 11 | 50 | 60 | 70 | 80 |\n",
"| 12 | 55 | 66 | 77 | 88 |\n",
"\n",
"| Region | Sub-region |\n",
"|--------------|-----------------|\n",
"| Lower 48 | Permian |\n",
"| | Eagle Ford |\n",
"| | Bakken |\n",
"| | Other |\n",
"| Alaska | GKA |\n",
"| | GWA |\n",
"| | GPA |\n",
"| | WNS |\n",
"| EMENA | Norway |\n",
"| | Libya |\n",
"| | Qatar |\n",
"| Asia Pacific | APLNG |\n",
"| | Malaysia |\n",
"| | China |\n",
"| Canada | Montney |\n",
"| | Surmont |\n",
"parsed_text: Our Differentiated Portfolio: Deep; Durable and Diverse\n",
" 20 BBOE of Resource Diverse Production Base\n",
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3 $40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
" 8 $20\n",
" E Qatar Libya\n",
" Asia Pacific Canada\n",
" $10 Permian\n",
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
"~30% Improved Performance from 2019 to 2022\n",
"\n",
"## Delaware Basin Well Capex/EUR⁴ ($/BOE)\n",
"\n",
"| Company | Capex/EUR |\n",
"|------------------|-----------|\n",
"| ConocoPhillips | 10 |\n",
"| Competitor 1 | 15 |\n",
"| Competitor 2 | 20 |\n",
"| Competitor 3 | 25 |\n",
"| Competitor 4 | 30 |\n",
"| Competitor 5 | 35 |\n",
"| Competitor 6 | 40 |\n",
"| Competitor 7 | 45 |\n",
"\n",
"---\n",
"\n",
"¹ Unconventional acres. \n",
"² Source: Enverus and ConocoPhillips (March 2023). \n",
"³ Source: Enverus (March 2023) based on wells online year. \n",
"⁴ Source: Enverus (March 2023). Average single well capex/EUR. Top eight public operators based on wells online in years 2021-2022, greater than 50% oil weight. COP based on COP well design. Competitors include: CVX, DVN, EOG, MTDR, OXY, PR and XOM.\n",
"parsed_text: \n",
"Delaware: Vast Inventory with Proven Track Record of Performance\n",
" New Prolific Acreage Spanning Over 12-Month Cumulative Production? (BOE/FT)\n",
" Mexico 659,000 Net Acres' 40\n",
" Texas 3828\n",
" 30 2019\n",
" 20 30%\n",
" 10 Improved Performancefrom 2019 to 2022\n",
" Total\n",
" Permian Inventory\n",
" 10-Year Operated\n",
" 2 10 11 12\n",
" Months\n",
" Delaware Basin Well Capex/EUR4 (S/BOE)\n",
" 65% 25\n",
" Delaware Basin 20\n",
" Midland Basin 15\n",
" Low HighCost of Supplyz 10 ConocoPhillips\n",
" High Single-Digit Production Growth\n",
" \"Unconventional acres. 2Source: Enverus and ConocoPhillips (March 2023). 3SourceEnverus (March 2023) based on wells online year: \"Source; Enverus (March 2023). Average single well capex/EUR Top eight public operators based on\n",
"wells online in years 2021-2022, greater than 50% oil weight; COP based on COP well design: Competitors include; CVX DVN, EOG; MTDR, OXY, PR and XOM: ConocoPhillips\n"
]
}
],
"source": [
"print(response.source_nodes[7].get_content(metadata_mode=\"all\"))"
"print(response.source_nodes[0].get_content(metadata_mode=\"all\"))"
]
},
{
@@ -921,20 +855,26 @@
"name": "stdout",
"output_type": "stream",
"text": [
"ConocoPhillips has a diverse production base spanning several key geographies:\n",
"\n",
"1. **Lower 48 (United States)**: This includes major production areas such as the Permian Basin, Bakken Formation, and Eagle Ford Shale.\n",
"2. **Alaska**: Significant operations in the North Slope region.\n",
"3. **Canada**: Operations in the Montney Formation and the Surmont oil sands project.\n",
"4. **EMENA (Europe, Middle East, and North Africa)**: Notable operations in Norway and Libya.\n",
"5. **Asia Pacific**: Includes operations in Malaysia, China, and the Australia Pacific LNG (APLNG) project.\n",
"6. **Qatar**: Involvement in the country's energy sector.\n",
"\n",
"These regions highlight the company's extensive and varied geographical footprint in the energy production industry.\n"
"Added user message to memory: How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20/BOE.\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"competitors capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"The average single well capex/EUR for competitors in the Delaware Basin is between $10 and $25 per BOE.\n",
"=== LLM Response ===\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20 per BOE. In comparison, the average capex/EUR for competitors in the Delaware Basin ranges between $10 and $25 per BOE. This places ConocoPhillips' capex/EUR towards the higher end of the competitive range.\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20 per BOE. In comparison, the average capex/EUR for competitors in the Delaware Basin ranges between $10 and $25 per BOE. This places ConocoPhillips' capex/EUR towards the higher end of the competitive range.\n"
]
}
],
"source": [
"# base_response = base_agent.query(\"Tell me about the different regions and subregions where Conoco Phillips has a production base.\")\n",
"base_response = base_agent.query(\n",
" \"How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\"\n",
")\n",
"print(str(base_response))"
]
},
@@ -948,31 +888,30 @@
"name": "stdout",
"output_type": "stream",
"text": [
"Our Differentiated Portfolio: Deep; Durable and Diverse\n",
" 20 BBOE of Resource Diverse Production Base\n",
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3 $40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
" 8 $20\n",
" E Qatar Libya\n",
" Asia Pacific Canada\n",
" $10 Permian\n",
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
"Deep, Durable and Diverse Portfolio with Significant Growth Runway\n",
" 1,2002022 Lower 48 Unconventional Production' (MBOED S50 ~S32/BBL\n",
" 000 ConocoPhillips Cost of SupplyAverage\n",
" 00 S40\n",
" 500 3\n",
" 400 1 S30\n",
" 200\n",
" 5\n",
" 15,000ConocoPhillipsNet Remaining Well Inventory? 1 S20\n",
" 12,000 S10\n",
" 000\n",
" 0o0 SO\n",
" 3,000 10\n",
" Resource (BBOE)\n",
" Delaware Basin Midland Basin Eagle Ford Bakken Other\n",
" Largest Lower 48 Unconventional Producer; Growing into the Next Decade\n",
" onshore operated inventory that achieves 15% IRR at $SO/BBL WTI, Competitors include CVX, DVN, EOG, FANG, MRO, OXY, PXD,and XOM:\n",
" Source: Wood Mackenzie Lower 48 Unconventional Plays 2022 ProductionCompetitors include CVX, DVN; EOG, FANG, MRO, OXY, PXD and XOM; greaterthan50% liquids weight: ?Source: Wood Mackenzie (March 2023), Lower 48\n",
" ConocoPhillips\n"
]
}
],
"source": [
"print(base_response.source_nodes[1].get_content(metadata_mode=\"all\"))"
"print(base_response.source_nodes[0].get_content(metadata_mode=\"llm\"))"
]
}
],

Before

Width:  |  Height:  |  Size: 271 KiB

After

Width:  |  Height:  |  Size: 271 KiB

@@ -7,7 +7,7 @@
"source": [
"# Multimodal Report Generation (from a Slide Deck)\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/multimodal_report_generation.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/multimodal_report_generation.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this cookbook we show you how to build a multimodal report generator. The pipeline parses a slide deck and stores both text and image chunks. It generates a detailed response that contains interleaving text and images.\n",
"\n",
@@ -143,7 +143,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
File diff suppressed because one or more lines are too long
@@ -6,7 +6,7 @@
"source": [
"# Building a RAG Pipeline over IKEA Product Instruction Manuals\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/multimodal/product_manual_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/multimodal/product_manual_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
@@ -104,7 +104,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse\n",
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
@@ -25,7 +25,7 @@
"\n",
"nest_asyncio.apply()\n",
"\n",
"from llama_cloud_services import LlamaParse"
"from llama_parse import LlamaParse"
]
},
{
@@ -27,7 +27,7 @@
"outputs": [],
"source": [
"%pip install llama-index\n",
"%pip install llama-cloud-services\n",
"%pip install llama-parse\n",
"%pip install torch transformers python-pptx Pillow"
]
},
@@ -85,7 +85,7 @@
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services import LlamaParse"
"from llama_parse import LlamaParse"
]
},
{
-124
View File
@@ -1,124 +0,0 @@
# LlamaParse Demo
A TypeScript demo application showcasing the power of **LlamaParse** - an intelligent document parsing service from [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to parse various document formats and generate AI-powered summaries using OpenAI's GPT models.
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [Development Mode](#development-mode)
- [Build the Project](#build-the-project)
- [Code Quality](#code-quality)
- [Quick Commands Reference](#quick-commands-reference)
- [How It Works](#how-it-works)
- [API Dependencies](#api-dependencies)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 📄 **Document Parsing**: Parse PDFs, Word docs, and other formats using LlamaParse
- 🤖 **AI Summaries**: Generate intelligent summaries using OpenAI GPT-4
- 🎨 **Beautiful CLI**: Styled console interface with colors and ASCII art
-**Fast Development**: Hot reload support with watch mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 18 or higher)
- pnpm package manager
- OpenAI API key
- LlamaCloud API key
## Installation
1. Clone the repository:
```bash
git clone <repository-url>
cd llamaparse-demo
```
2. Install dependencies:
```bash
pnpm install
```
3. Set up your environment variables:
```bash
# Add your API keys to your environment
export OPENAI_API_KEY="your-openai-api-key"
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
```
## Usage
### Start the Demo
```bash
pnpm run start
```
The application will display a welcome screen and prompt you to enter the path to a document you'd like to process.
### Development Mode
For development with hot reload:
```bash
pnpm run dev
```
### Build the Project
```bash
pnpm run build
```
### Code Quality
Format code:
```bash
pnpm run format
```
Lint code:
```bash
pnpm run lint
```
## How It Works
1. **Document Input**: Enter the path to your document when prompted
2. **Parsing**: LlamaParse processes the document and extracts structured content
3. **AI Summary**: The extracted content is sent to OpenAI GPT-4 for summarization
4. **Results**: View the AI-generated summary in your terminal
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 18+ and have all dependencies installed
2. **API Key Issues**: Verify your OpenAI and LlamaCloud API keys are correctly set
3. **File Path Errors**: Use absolute paths or ensure relative paths are correct from the project root
## License
MIT License - see the [LICENSE](../../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `pnpm format` and `pnpm lint`
5. Submit a pull request
Binary file not shown.
-15
View File
@@ -1,15 +0,0 @@
import js from "@eslint/js";
import globals from "globals";
import tseslint from "typescript-eslint";
import { defineConfig } from "eslint/config";
export default defineConfig([
{
files: ["**/*.{js,mjs,cjs,ts,mts,cts}"],
plugins: { js },
extends: ["js/recommended"],
languageOptions: { globals: globals.browser },
},
{ files: ["**/*.js"], languageOptions: { sourceType: "script" } },
tseslint.configs.recommended,
]);
-47
View File
@@ -1,47 +0,0 @@
{
"name": "llamaparse-demo",
"version": "0.1.0",
"description": "Demo for LlamaParse in TypeScript",
"type": "module",
"main": "index.js",
"scripts": {
"test": "echo \"There are no tests\"",
"start": "pnpm exec tsx src/index.ts",
"lint": "eslint ./src/",
"format": "prettier --write ./src/",
"build": "tsc",
"dev": "pnpm exec tsx --watch src/index.ts"
},
"keywords": [
"ai",
"ocr",
"parsing",
"intelligent-document-processing",
"pdf",
"llms"
],
"author": "LlamaIndex",
"license": "MIT",
"packageManager": "pnpm@10.12.4",
"devDependencies": {
"@eslint/js": "^9.32.0",
"@types/figlet": "^1.7.0",
"@types/node": "^24.1.0",
"@typescript-eslint/eslint-plugin": "^8.38.0",
"@typescript-eslint/parser": "^8.38.0",
"eslint": "^9.32.0",
"globals": "^16.3.0",
"jiti": "^2.5.1",
"prettier": "^3.6.2",
"typescript": "^5.8.3",
"typescript-eslint": "^8.38.0"
},
"dependencies": {
"@ai-sdk/openai": "^1.3.23",
"ai": "^4.3.19",
"consola": "^3.4.2",
"figlet": "^1.8.2",
"llama-cloud-services": "link:../../../ts/llama_cloud_services",
"picocolors": "^1.1.1"
}
}
-1758
View File
File diff suppressed because it is too large Load Diff
-34
View File
@@ -1,34 +0,0 @@
import { LlamaParseReader } from "llama-cloud-services";
import { logger } from "./logger";
import pc from "picocolors";
import { consoleInput, generateSummary, renderLogo } from "./utils";
export async function main(): Promise<number> {
const reader = new LlamaParseReader({ resultType: "markdown" });
await renderLogo();
logger.log(
`Welcome to ${pc.bold(
pc.magentaBright("✨LlamaParse Demo✨"),
)}, our demo for ${pc.bold(pc.green("LlamaParse🦙"))}, a ${pc.bold(
pc.cyan("LlamaCloud☁️"),
)} (https://cloud.llamaindex.ai) product!.\nType the path to the document you would like to process below👇\nIf you wish to exit, just type ${pc.bold(
pc.gray("quit"),
)}.\n`,
);
while (true) {
const userInput = await consoleInput();
if (userInput.toLowerCase() == "quit") {
break;
}
try {
const documents = await reader.loadData(userInput);
const summary = await generateSummary(documents); // Added await here
logger.log(`${pc.bold(pc.cyan("AI-generated summary✨"))}:\n${summary}`);
} catch (error) {
logger.error(`Error processing file: ${error}`);
}
}
return 0;
}
main().catch(console.error);
-8
View File
@@ -1,8 +0,0 @@
import { createConsola } from "consola";
import type { ConsolaInstance } from "consola";
export const logger: ConsolaInstance = createConsola({
formatOptions: {
date: false,
},
});
-51
View File
@@ -1,51 +0,0 @@
import { generateText } from "ai";
import { openai } from "@ai-sdk/openai";
import { Document } from "llamaindex";
import * as readline from "readline/promises";
import figlet from "figlet";
import pc from "picocolors";
export async function renderLogo(): Promise<void> {
const logoText = figlet.textSync("LlamaParse Demo", {
font: "ANSI Shadow",
horizontalLayout: "default",
verticalLayout: "default",
width: 100,
whitespaceBreak: true,
});
// Add some styling with picocolors
const styledLogo = pc.bold(pc.magentaBright(logoText));
// Add some padding/margin
console.log("\n");
console.log(styledLogo);
console.log(pc.gray("─".repeat(60)));
console.log("\n");
}
export async function consoleInput(): Promise<string> {
const rl = readline.createInterface({
input: process.stdin,
output: process.stdout,
});
const answer = await rl.question("Path to your file: ");
rl.close();
return answer;
}
export async function generateSummary(documents: Document[]): Promise<string> {
let mainText: string = "";
for (const document of documents) {
mainText += `${document.text}\n\n---\n\n`;
}
const { text } = await generateText({
model: openai("gpt-4.1"),
prompt: `</chat>\n\t<text>${mainText}</text>\n\t<instructions>Could you please generate a summary of the given text?</instructions>\n</chat>`,
});
return text;
}
-22
View File
@@ -1,22 +0,0 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "ES2022",
"lib": ["ES2022"],
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"declaration": true,
"declarationMap": true,
"sourceMap": true,
"types": ["node"],
"moduleResolution": "bundler",
"allowSyntheticDefaultImports": true,
"resolveJsonModule": true
},
"include": ["src/**/*"],
"exclude": ["node_modules", "dist"]
}
File diff suppressed because it is too large Load Diff
Binary file not shown.

Before

Width:  |  Height:  |  Size: 6.9 MiB

Binary file not shown.
-618
View File
@@ -1,618 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Advanced RAG with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook is a complete walkthrough for using LlamaParse with advanced indexing/retrieval techniques in LlamaIndex over the Apple 10K Filing. \n",
"\n",
"This allows us to ask sophisticated questions that aren't possible with \"naive\" parsing/indexing techniques with existing models."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://s2.q4cdn.com/470004039/files/doc_financials/2021/q4/_10-K-2021-(As-Filed).pdf\" -O apple_2021_10k.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Some OpenAI and LlamaParse details"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\"\n",
"\n",
"# Using OpenAI API for embeddings/llms\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-proj-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.openai import OpenAI\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import Settings\n",
"\n",
"embed_model = OpenAIEmbedding(model_name=\"text-embedding-3-small\")\n",
"llm = OpenAI(model=\"gpt-4o-mini\")\n",
"\n",
"Settings.llm = llm\n",
"Settings.embed_model = embed_model"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Using brand new `LlamaParse` PDF reader for PDF Parsing\n",
"\n",
"We also compare three different retrieval/query engine strategies:\n",
"1. Baseline using default parsing from `SimpleDirectoryReader`\n",
"2. Using raw markdown text as nodes for building index and apply simple query engine for generating the results;\n",
"3. Using markdown + page screenshots to help retrieve the proper nodes."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id e403a457-1721-4093-82bf-4a316d2d637a\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"result = await LlamaParse(take_screenshot=True).aparse(\"./apple_2021_10k.pdf\")\n",
"\n",
"markdown_nodes = await result.aget_markdown_nodes(split_by_page=True)\n",
"screenshot_image_nodes = await result.aget_image_nodes(\n",
" include_screenshot_images=True,\n",
" include_object_images=False,\n",
" image_download_dir=\"./images\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import SimpleDirectoryReader\n",
"\n",
"baseline_documents = SimpleDirectoryReader(\n",
" input_files=[\"apple_2021_10k.pdf\"]\n",
").load_data()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup Baseline Index\n",
"\n",
"For comparison, we setup a naive RAG pipeline with default parsing and standard chunking, indexing, retrieval."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import VectorStoreIndex\n",
"\n",
"baseline_index = VectorStoreIndex.from_documents(baseline_documents)\n",
"baseline_query_engine = baseline_index.as_query_engine(similarity_top_k=3)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup our LlamaParse Indexes\n",
"\n",
"Using both the markdown and screenshot images, we can build two different indexes.\n",
"\n",
"1. An index over just the markdown documents\n",
"2. A custom index that uses the markdown + screenshot images to help with response quality."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import VectorStoreIndex\n",
"\n",
"markdown_index = VectorStoreIndex(nodes=markdown_nodes)\n",
"markdown_query_engine = markdown_index.as_query_engine(similarity_top_k=3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.indices import MultiModalVectorStoreIndex\n",
"from llama_index.embeddings.huggingface import HuggingFaceEmbedding\n",
"from llama_index.core import Settings\n",
"\n",
"# could also use other API-based multimodal models like voyageai or jinaai\n",
"# Note: this may take quite a while if running on CPU!\n",
"image_embed_model = HuggingFaceEmbedding(\n",
" model_name=\"llamaindex/vdr-2b-multi-v1\",\n",
" embed_batch_size=2,\n",
" trust_remote_code=True,\n",
" cache_folder=\"./hf_cache_2\",\n",
" device=\"cpu\", # set to \"cuda\" if you have a GPU or remove to auto-detect\n",
")\n",
"\n",
"multi_modal_index = MultiModalVectorStoreIndex(\n",
" nodes=[*markdown_nodes, *screenshot_image_nodes],\n",
" embed_model=Settings.embed_model,\n",
" image_embed_model=image_embed_model,\n",
" show_progress=True,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Below, we will create a custom query engine that does a few things\n",
"1. Retrieves both image nodes and text nodes\n",
"2. Combines them into two lists -- one where images and texts come from the same page, and one where we have texts alone\n",
"3. Use a Jinja-based `RichPromptTemplate` to format the retrieved content automatically into a list of multimodal chat messages\n",
"4. Send our messages to the LLM and return a result\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.async_utils import asyncio_run\n",
"from llama_index.core.llms import LLM\n",
"from llama_index.core.query_engine import CustomQueryEngine\n",
"from llama_index.core.prompts import RichPromptTemplate\n",
"from llama_index.core.response import Response\n",
"from llama_index.core.schema import NodeWithScore\n",
"from llama_index.core import Settings\n",
"\n",
"TEXT_IMAGE_PROMPT_TEMPLATE = RichPromptTemplate(\n",
" \"\"\"\n",
"<context>\n",
"Here is some retrieved content from a knowledge base:\n",
"{% for image_path, text in images_and_texts %}\n",
"<page>\n",
"<text>{{ text }}</text>\n",
"<image>{{ image_path | image }}</image>\n",
"</page>\n",
"{% endfor %}\n",
"{% for text in texts %}\n",
"<page>\n",
"<text>{{ text }}</text>\n",
"</page>\n",
"{% endfor %}\n",
"</context>\n",
"\n",
"Using the context, answer the following question:\n",
"<query>{{ query_str }}</query>\n",
"\"\"\"\n",
")\n",
"\n",
"\n",
"class SimpleMultiModalQueryEngine(CustomQueryEngine):\n",
" def __init__(\n",
" self,\n",
" index: MultiModalVectorStoreIndex,\n",
" image_top_k: int = 4,\n",
" text_top_k: int = 4,\n",
" llm: LLM | None = None,\n",
" **kwargs\n",
" ):\n",
" super().__init__(**kwargs)\n",
" self._retriever = index.as_retriever(\n",
" similarity_top_k=text_top_k, image_similarity_top_k=image_top_k\n",
" )\n",
" self._llm = llm or Settings.llm\n",
"\n",
" def _match_images_and_texts(\n",
" self, text_results: list[NodeWithScore], image_results: list[NodeWithScore]\n",
" ) -> tuple[list[NodeWithScore], list[NodeWithScore]]:\n",
" # combine results, prioritize images and texts\n",
" # if both an image and matching text was retrieved, that is a strong indicator\n",
" images_and_texts = []\n",
" text_keys = {\n",
" (x.metadata[\"page_number\"], x.metadata[\"file_name\"]): x\n",
" for x in text_results\n",
" }\n",
" for image_result in image_results:\n",
" key = (\n",
" image_result.metadata[\"page_number\"],\n",
" image_result.metadata[\"file_name\"],\n",
" )\n",
" # add matching text to results if available\n",
" if key in text_keys:\n",
" text_result = text_keys[key]\n",
" images_and_texts.append(\n",
" (image_result.node.image_path, text_result.node.text)\n",
" )\n",
"\n",
" # remove from list\n",
" text_keys.pop(key)\n",
"\n",
" # get the remaining texts as a fallback\n",
" texts = [result.node.text for result in text_keys.values()]\n",
"\n",
" return images_and_texts, texts\n",
"\n",
" def custom_query(self, query_str: str) -> Response:\n",
" # wrap the async method to avoid code duplication\n",
" # asyncio_run is a slightly safer asyncio.run() call\n",
" return asyncio_run(self.acustom_query(query_str))\n",
"\n",
" async def acustom_query(self, query_str: str) -> Response:\n",
" text_results = await self._retriever.atext_retrieve(query_str)\n",
" image_results = await self._retriever.atext_to_image_retrieve(query_str)\n",
"\n",
" images_and_texts, texts = self._match_images_and_texts(\n",
" text_results, image_results\n",
" )\n",
" messages = TEXT_IMAGE_PROMPT_TEMPLATE.format_messages(\n",
" images_and_texts=images_and_texts, texts=texts, query_str=str(query_str)\n",
" )\n",
"\n",
" response = await self._llm.achat(messages)\n",
"\n",
" return Response(\n",
" response.message.content, source_nodes=[*text_results, *image_results]\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"multimodal_query_engine = SimpleMultiModalQueryEngine(\n",
" index=multi_modal_index,\n",
" image_top_k=3,\n",
" text_top_k=3,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Try out the Query Engines and Compare!\n",
"\n",
"Now with our three query engines assembled, we can compare each approach with a rough \"vibes-based\" evaluation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $190,516 million.\n",
"\n",
"***********Markdown Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $191,830 million.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The total fair value of marketable securities in 2020 was $191,830 million.\n"
]
}
],
"source": [
"query = \"What were the total fair value of marketable securities in 2020\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"As we can see, the multimodal and markdown query engines are able to retrieve the correct content, while the default query engine struggles to find the correct total value."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can also inspect the source nodes, and see the pages that were retrieved. Here is the correct page for the total fair value of marketable securities in 2020:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'images/page_41.jpg'"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"response_3.source_nodes[4].node.image_path"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Lets try a few more queries to see how the query engines perform."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The effective interest rates for the debt issuances in 2021 were as follows:\n",
"\n",
"- Floating-rate notes: 0.48% 0.63%\n",
"- Fixed-rate notes: 0.03% 4.78% for maturities from 2022 to 2060\n",
"- Fixed-rate notes issued in the second quarter: 0.75% 2.81% for maturities from 2026 to 2061\n",
"- Fixed-rate notes issued in the fourth quarter: 1.43% 2.86% for maturities from 2028 to 2061\n",
"\n",
"***********Markdown Query Engine***********\n",
"The effective interest rates for the debt issuances in 2021 were as follows:\n",
"\n",
"- Floating-rate notes: 0.48% 0.63%\n",
"- Fixed-rate notes: 0.03% 4.78% for the 0.000% 4.650% notes, 0.75% 2.81% for the 0.700% 2.800% notes, and 1.43% 2.86% for the 1.400% 2.850% notes.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The effective interest rates of all debt issuances in 2021 were as follows:\n",
"\n",
"1. **Floating-rate notes**: 0.48% 0.63%\n",
"2. **Fixed-rate 0.000% 4.650% notes**: 0.03% 4.78%\n",
"3. **Fixed-rate 0.700% 2.800% notes**: 0.75% 2.81%\n",
"4. **Fixed-rate 1.400% 2.850% notes**: 1.43% 2.86%\n"
]
}
],
"source": [
"query = \"What were the effective interest rates of all debt issuances in 2021\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The federal deferred tax amounts for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2019**: $(2,939)\n",
"- **2020**: $(3,619)\n",
"- **2021**: $(7,176)\n",
"\n",
"These figures represent the deferred tax expense for each respective year.\n",
"\n",
"***********Markdown Query Engine***********\n",
"As of September 25, 2021, the total deferred tax assets and liabilities for the years 2021 and 2020 are as follows:\n",
"\n",
"**Deferred Tax Assets:**\n",
"- 2021: $25,176 million\n",
"- 2020: $19,336 million\n",
"\n",
"**Deferred Tax Liabilities:**\n",
"- 2021: $7,200 million\n",
"- 2020: $10,138 million\n",
"\n",
"**Net Deferred Tax Assets:**\n",
"- 2021: $13,073 million\n",
"- 2020: $8,157 million\n",
"\n",
"The information for 2019 is not provided in the context.\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The federal deferred tax assets and liabilities for the years 2019 to 2021 are as follows:\n",
"\n",
"### Deferred Tax Assets (in millions):\n",
"- **2021**: $25,176\n",
"- **2020**: $19,336\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"### Deferred Tax Liabilities (in millions):\n",
"- **2021**: $7,200\n",
"- **2020**: $10,138\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"### Net Deferred Tax Assets (in millions):\n",
"- **2021**: $13,073\n",
"- **2020**: $8,157\n",
"- **2019**: Not specified in the provided content.\n",
"\n",
"The significant components of deferred tax assets and liabilities reflect the effects of tax credits and temporary differences between financial statement carrying amounts and their respective tax bases.\n"
]
}
],
"source": [
"query = \"federal deferred tax in 2019-2021\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********Baseline Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- 2021: $1,620\n",
"- 2020: $455\n",
"- 2019: $475\n",
"\n",
"This indicates an increase of $1,165 million from 2020 to 2021, a decrease of $20 million from 2018 to 2019, and an increase of $80 million from 2019 to 2020.\n",
"\n",
"***********Markdown Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2021**: $1,620\n",
"- **2020**: $455\n",
"- **2019**: $475\n",
"\n",
"The changes in current state taxes from year to year are:\n",
"\n",
"- From 2019 to 2020: Decrease of $20 million\n",
"- From 2020 to 2021: Increase of $1,165 million\n",
"\n",
"***********MultiModal Query Engine***********\n",
"The current state taxes for the years 2019 to 2021 are as follows (in millions):\n",
"\n",
"- **2021**: $1,620\n",
"- **2020**: $455\n",
"- **2019**: $475\n",
"\n",
"So, the changes are:\n",
"- From 2019 to 2020: Decrease of $20 million\n",
"- From 2020 to 2021: Increase of $1,165 million\n"
]
}
],
"source": [
"query = \"current state taxes per year in 2019-2021 (include +/-)\"\n",
"\n",
"response_1 = await baseline_query_engine.aquery(query)\n",
"print(\"\\n***********Baseline Query Engine***********\")\n",
"print(response_1)\n",
"\n",
"response_2 = await markdown_query_engine.aquery(query)\n",
"print(\"\\n***********Markdown Query Engine***********\")\n",
"print(response_2)\n",
"\n",
"response_3 = await multimodal_query_engine.aquery(query)\n",
"print(\"\\n***********MultiModal Query Engine***********\")\n",
"print(response_3)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama-parse-aNC435Vv-py3.10",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
-196
View File
@@ -1,196 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse Usage"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index llama-cloud-services"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget \"https://arxiv.org/pdf/1706.03762.pdf\" -O \"./attention.pdf\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 79ae653c-4598-4bd0-ba6e-b3dab7eab57e\n"
]
}
],
"source": [
"from llama_cloud_services import LlamaParse\n",
"\n",
"result = await LlamaParse().aparse(\"./attention.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"1 Introduction\n",
"Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neural networks\n",
"in particular, have been firmly established as state of the art approaches in sequence modeling and\n",
"transduction problems such as language modeling and machine translation [35, 2, 5]. Numerous\n",
"efforts have since continued to push the boundaries of recurrent language models and encoder-decoder\n",
"architectures [38, 24, 15].\n",
"Recurrent models typically factor computation along the symbol positions of the input and output\n",
"sequences. Aligning the positions to steps in computation time, they generate a sequence of hidden\n",
"states ht, as a function of the previous hidden state ht1 and the input for position t. This inherently\n",
"sequential nature precludes parallelization within training examples, which becomes critical at longer\n",
"sequence lengths, as memory constraints limit batching across examples. Recent work has achieved\n",
"significant improvements in computational efficiency through factorization tricks [21] and conditional\n",
"computation [32], while also improving model performance in case of the latter. The fundamental\n",
"constraint of sequential computation, however, remains.\n",
"Attention mechanisms have become an integral part of compelling sequence modeling and transduc-\n",
"tion models in various tasks, allowing modeling of dependencies without regard to their distance in\n",
"the input or output sequences [2, 19]. In all but a few cases [27], however, such attention mechanisms\n",
"are used in conjunction with a recurrent network.\n",
"In this work we propose the Transformer, a model architecture eschewing recurrence and instead\n",
"relying entirely on an attention mechanism to draw global dependencies between input and output.\n",
"The Transformer allows for significantly more parallelization and can reach a new state of the art in\n",
"translation quality after being trained for as little as twelve hours on eight P100 GPUs.\n",
"2 Background\n",
"The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU\n",
"[16], ByteNet [18] and ConvS2S [9], all of which use convolutional neural networks as basic building\n",
"block, computing hidden representations in parallel for all input and output positions. In these models,\n",
"the number of operations required to relate signals from two arbitrary input or output positions grows\n",
"in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes\n",
"it more difficult to learn dependencies between distant positions [12]. In the Transformer this is\n",
"reduced to a constant number of operations, albeit at the cost of reduced effective resolution due\n",
"to averaging attention-weighted positions, an effect we counteract with Multi-Head Attention as\n",
"described in section 3.2.\n",
"Self-attention, sometimes called intra-attention is an attention mechanism relating different positions\n",
"of a single sequence in order to compute a representation of the sequence. Self-attention has been\n",
"used successfully in a variety of tasks including reading comprehension, abstractive summarization,\n",
"textual entailment and learning task-independent sentence representations [4, 27, 28, 22].\n",
"End-to-end memory networks are based on a recurrent attention mechanism instead of sequence-\n",
"aligned recurrence and have been shown to perform well on simple-language question answering and\n",
"language modeling tasks [34].\n",
"To the best of our knowledge, however, the Transformer is the first transduction model relying\n",
"entirely on self-attention to compute representations of its input and output without using sequence-\n",
"aligned RNNs or convolution. In the following sections, we will describe the Transformer, motivate\n",
"self-attention and discuss its advantages over models such as [17, 18] and [9].\n",
"3 Model Architecture\n",
"Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].\n",
"Here, the encoder maps an input sequence of symbol representations (x1, ..., xn) to a sequence\n",
"of continuous representations z = (z1, ..., zn). Given z, the decoder then generates an output\n",
"sequence (y1, ..., ym) of symbols one element at a time. At each step the model is auto-regressive\n",
"[10], consuming the previously generated symbols as additional input when generating the next.\n",
" 2\n"
]
}
],
"source": [
"documents = result.get_text_documents(split_by_page=True)\n",
"print(documents[1].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"arXiv:1706.03762v7 [cs.CL] 2 Aug 2023\n",
"\n",
"Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.\n",
"\n",
"# Attention Is All You Need\n",
"\n",
"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit\n",
"\n",
"Google Brain Google Brain Google Research Google Research\n",
"\n",
"avaswani@google.com noam@google.com nikip@google.com usz@google.com\n",
"\n",
"Llion Jones Aidan N. Gomez † Łukasz Kaiser\n",
"\n",
"Google Research University of Toronto Google Brain\n",
"\n",
"llion@google.com aidan@cs.toronto.edu lukaszkaiser@google.com\n",
"\n",
"Illia Polosukhin ‡\n",
"\n",
"illia.polosukhin@gmail.com\n",
"\n",
"# Abstract\n",
"\n",
"The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.\n",
"\n",
"Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research.\n",
"\n",
"†Work performed while at Google Brain.\n",
"\n",
"‡Work performed while at Google Research.\n",
"\n",
"31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.\n"
]
}
],
"source": [
"documents = result.get_markdown_documents(split_by_page=True)\n",
"print(documents[0].text)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}

Some files were not shown because too many files have changed in this diff Show More