Compare commits

...

42 Commits

Author SHA1 Message Date
Pierre-Loic Doulcet 032c8173f4 Extract layout, audio files 2024-12-18 16:20:04 +01:00
Bharath Lakshman Kumar 6d62fb89c3 Fix docstring for aget_xlsx method (#551)
Updated docstring to describe xlsx download instead of image download
2024-12-13 20:35:18 +05:30
Ravi Theja 7d4df3b6e5 Add cookbook for parsing instructions (#550) 2024-12-13 06:44:17 -08:00
Ravi Theja bc28db5b92 Update cache parameter (#548) 2024-12-11 16:15:59 +01:00
Jerry Liu f78186c0f7 update auto-mode (#545) 2024-12-09 16:13:09 -06:00
Laurie Voss e3292f5566 Expanding auto mode notebook with strings and regex triggers (#544) 2024-12-09 12:03:09 -08:00
Jerry Liu 58f980f411 auto-mode notebook (#540)
Co-authored-by: Laurie Voss <github@seldo.com>
2024-12-09 08:59:21 -08:00
Ravi Theja 4740d0611d Add get charts function (#542)
* Add get charts function

* code refactoring

* solve linting

* Add cookbook
2024-12-09 21:28:48 +05:30
Laurie Voss 3651a10e80 JSON mode tour notebook (#531) 2024-12-06 14:21:15 -08:00
Pierre-Loic Doulcet 483b51c51c Add support for html_remove_navigation_elements. (#532) 2024-12-06 12:05:46 +01:00
Ravi Theja cdbddef86d Add demo videos notebooks (#529) 2024-12-05 08:38:34 -08:00
Pierre-Loic Doulcet 3690109abf Add more parameters (#525)
* add after revert

* 3.8 so numpy work

* change defaults

* change requested

* change requested
2024-12-04 15:39:00 +01:00
Pierre-Loic Doulcet 2e322b4fc8 Revert "Add more paramerters"
This reverts commit 735e5f3ddc.
2024-12-04 10:20:07 +01:00
Pierre-Loic Doulcet 735e5f3ddc Add more paramerters 2024-12-04 10:17:08 +01:00
Logan e4cb4c75e5 add test for downloading images (#506) 2024-11-21 13:08:29 -06:00
Jerry Liu 1693deff72 dynamic section retrieval nb (#484) 2024-11-13 13:29:30 +01:00
Jerry Liu 3270f1228d multimodal report generation image (#461)
* cr

* cr
2024-11-13 13:28:07 +01:00
Pierre-Loic Doulcet eeabf48d29 add input url and http_proxy (#475) 2024-11-12 12:56:58 -06:00
Pierre-Loic Doulcet 89348aa8e5 add xlsx support (#472) 2024-11-01 10:09:17 -06:00
Thiago Salvatore 3ab2ce27b5 Add PurePosixPath to list of allowed file-paths (#464) 2024-10-25 10:45:47 -06:00
Sacha Bron 265261862f Add continuous_mode (#460) 2024-10-22 19:45:46 +02:00
Sacha Bron 66cf052b8c Update issue templates (#457)
* Update issue templates

* Update issue templates
2024-10-21 19:51:46 +02:00
Jerry Liu 2ca2d81e58 fix RFP example (#455) 2024-10-21 09:13:24 -07:00
Sacha Bron 951ba4dfd8 Release is_formatting_instruction parameter (#446)
* Release is_formatting_instruction parameter

* Add annotate links
2024-10-17 12:29:05 +02:00
Adam Reichert 386d210e8b CLI Testing Tool for Parsing Results to Standard Output (#363) 2024-10-16 12:40:00 -06:00
Sacha Bron 9321602845 Add missing parameters (#441) 2024-10-15 10:57:32 -06:00
Jerry Liu 26c06353f0 Add RFP Response generation workflow (#438) 2024-10-14 08:45:04 -07:00
Jerry Liu 62cf12d6eb add multimodal RAG pipeline with contextual retrieval (#429) 2024-10-06 15:25:57 -07:00
Logan 253ee61463 improve error handling for jobs (#426) 2024-10-02 18:57:46 -06:00
Sourabh Desai 2ccd2a9397 Update README.md to convey need to specify extra_info["file_name"] (#417) 2024-09-29 17:07:12 -07:00
Jerry Liu c139e8e3e6 fix excel notebook (#416) 2024-09-24 17:11:33 -07:00
Ravi Theja 6e6e96c422 Update excel rag with o1 notebook (#415) 2024-09-24 07:42:00 -07:00
Jerry Liu b677e5226d nit: move o1 excel notebook (#414) 2024-09-23 10:51:51 -07:00
Ravi Theja df723584b6 Compare Excel RAG with o1 models (#409) 2024-09-23 10:42:47 -07:00
Sacha Bron efe06ffff0 Bump to v0.5.6 2024-09-19 14:33:30 +02:00
Pierre-Loic Doulcet 6ba052d58f add premium mode support (#406) 2024-09-18 11:48:45 +02:00
Sacha Bron 8cf52058b5 Remove JSON from valid result types (#400) 2024-09-18 11:48:22 +02:00
Jerry Liu 1bae09126c fix multimodal RAG over slide deck (#402) 2024-09-17 13:08:10 +08:00
Pierre-Loic Doulcet bbbae9de9d do not attach a filepath when a stram of bytes is passed (#394) 2024-09-10 11:53:53 -06:00
Thiago Salvatore 7cb6d06316 Enable support for custom filesystem (#117) 2024-09-10 10:39:46 -06:00
Jerry Liu bca5492829 update README (#386) 2024-09-09 17:31:58 -06:00
Sourabh Desai f6a4d8681f bump to 0.5.3 (#388) 2024-09-09 11:06:58 -06:00
54 changed files with 10838 additions and 1611 deletions
+4 -10
View File
@@ -7,8 +7,6 @@ assignees: ''
---
_Note: we're aware of some missing content in the output and layout issues on tables. Please refrain from opening new issues on this topic unless if you think it's different from what has already been reported._
**Describe the bug**
Write a concise description of what the bug is.
@@ -19,19 +17,15 @@ If possible, please provide the PDF file causing the issue.
If you have it, please provide the ID of the job you ran.
You can find it here: https://cloud.llamaindex.ai/parse in the "History" tab.
**Screenshots**
Feel free to also provide screenshots if relevant.
**Client:**
Please remove untested options:
- Frontend (cloud.llamaindex.ai)
- Python Library
- API
- Frontend (cloud.llamaindex.ai)
- Typescript Library
- Notebook
- API
**Options**
What options did you use? Multimodal, fast mode, parsing instructions, etc.
**Additional context**
Add any additional context about the problem here.
What options did you use? Premium mode, multimodal, fast mode, parsing instructions, etc.
Screenshots, code snippets, etc.
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.8", "3.10", "3.11"]
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v3
with:
+1
View File
@@ -2,3 +2,4 @@
__pycache__/
*.pyc
.DS_Store
.idea
+46 -9
View File
@@ -1,15 +1,26 @@
# LlamaParse
LlamaParse is an API created by LlamaIndex to efficiently parse and represent files for efficient retrieval and context augmentation using LlamaIndex frameworks.
[![PyPI - Downloads](https://img.shields.io/pypi/dm/llama-parse)](https://pypi.org/project/llama-parse/)
[![GitHub contributors](https://img.shields.io/github/contributors/run-llama/llama_parse)](https://github.com/run-llama/llama_parse/graphs/contributors)
[![Discord](https://img.shields.io/discord/1059199217496772688)](https://discord.gg/dGcwcsnxhU)
LlamaParse is a **GenAI-native document parser** that can parse complex document data for any downstream LLM use case (RAG, agents).
It is really good at the following:
-**Broad file type support**: Parsing a variety of unstructured file types (.pdf, .pptx, .docx, .xlsx, .html) with text, tables, visual elements, weird layouts, and more.
-**Table recognition**: Parsing embedded tables accurately into text and semi-structured representations.
-**Multimodal parsing and chunking**: Extracting visual elements (images/diagrams) into structured formats and return image chunks using the latest multimodal models.
-**Custom parsing**: Input custom prompt instructions to customize the output the way you want it.
LlamaParse directly integrates with [LlamaIndex](https://github.com/run-llama/llama_index).
Free plan is up to 1000 pages a day. Paid plan is free 7k pages per week + 0.3c per additional page.
There is a sandbox available to test the API [**https://cloud.llamaindex.ai/parse ↗**](https://cloud.llamaindex.ai/parse).
The free plan is up to 1000 pages a day. Paid plan is free 7k pages per week + 0.3c per additional page by default. There is a sandbox available to test the API [**https://cloud.llamaindex.ai/parse ↗**](https://cloud.llamaindex.ai/parse).
Read below for some quickstart information, or see the [full documentation](https://docs.cloud.llamaindex.ai/).
If you're a company interested in enterprise RAG solutions, and/or high volume/on-prem usage of LlamaParse, come [talk to us](https://www.llamaindex.ai/contact).
## Getting Started
First, login and get an api-key from [**https://cloud.llamaindex.ai/api-key ↗**](https://cloud.llamaindex.ai/api-key).
@@ -27,7 +38,22 @@ Lastly, install the package:
`pip install llama-parse`
Now you can run the following to parse your first PDF file:
Now you can parse your first PDF file using the command line interface. Use the command `llama-parse [file_paths]`. See the help text with `llama-parse --help`.
```bash
export LLAMA_CLOUD_API_KEY='llx-...'
# output as text
llama-parse my_file.pdf --result-type text --output-file output.txt
# output as markdown
llama-parse my_file.pdf --result-type markdown --output-file output.md
# output as raw json
llama-parse my_file.pdf --output-raw-json --output-file output.json
```
You can also create simple scripts:
```python
import nest_asyncio
@@ -76,13 +102,18 @@ parser = LlamaParse(
language="en", # Optionally you can define a language, default=en
)
with open("./my_file1.pdf", "rb") as f:
documents = parser.load_data(f)
file_name = "my_file1.pdf"
extra_info = {"file_name": file_name}
with open(f"./{file_name}", "rb") as f:
# must provide extra_info with file_name key with passing file object
documents = parser.load_data(f, extra_info=extra_info)
# you can also pass file bytes directly
with open("./my_file1.pdf", "rb") as f:
with open(f"./{file_name}", "rb") as f:
file_bytes = f.read()
documents = parser.load_data(file_bytes)
# must provide extra_info with file_name key with passing file bytes
documents = parser.load_data(file_bytes, extra_info=extra_info)
```
## Using with `SimpleDirectoryReader`
@@ -126,3 +157,9 @@ Several end-to-end indexing examples can be found in the examples folder
## Terms of Service
See the [Terms of Service Here](./TOS.pdf).
## Get in Touch (LlamaCloud)
LlamaParse is part of LlamaCloud, our e2e enterprise RAG platform that provides out-of-the-box, production-ready connectors, indexing, and retrieval over your complex data sources. We offer SaaS and VPC options.
LlamaCloud is currently available via waitlist (join by [creating an account](https://cloud.llamaindex.ai/)). If you're interested in state-of-the-art quality and in centralizing your RAG efforts, come [get in touch with us](https://www.llamaindex.ai/contact).
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 6.9 MiB

Binary file not shown.
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -342,7 +342,7 @@
],
"metadata": {
"kernelspec": {
"display_name": "llama-parse-aNC435Vv-py3.10",
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
File diff suppressed because it is too large Load Diff
+357
View File
@@ -0,0 +1,357 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "97c79c38-38a3-40f3-ba2e-250649347d63",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_starter_multimodal.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
"cell_type": "markdown",
"id": "4e081457",
"metadata": {},
"source": [
"# Multimodal Parsing using LlamaParse\n",
"\n",
"This cookbook shows you how to use LlamaParse to parse any document with the multimodal capabilities of Multi-Modal LLMs from Anthropic/ OpenAI.\n",
"\n",
"LlamaParse allows you to plug in external, multimodal model vendors for parsing - we handle the error correction, validation, and scalability/reliability for you.\n"
]
},
{
"cell_type": "markdown",
"id": "qOdqBxCS51Ow",
"metadata": {},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "H_Vqcylb50vm",
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"id": "15e60ecf-519c-41fc-911b-765adaf8bad4",
"metadata": {},
"source": [
"### Setup\n",
"\n",
"Here we setup `LLAMA_CLOUD_API_KEY` for using `LlamaParse`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "91a9e532-1454-40e0-bbf0-fd442c350121",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<YOUR LLAMACLOUD API KEY>\""
]
},
{
"cell_type": "markdown",
"id": "LGwBNPNotZRQ",
"metadata": {},
"source": [
"## Download Data\n",
"\n",
"For this demonstration, we will use OpenAI's recent paper `Evaluation of OpenAI o1: Opportunities and Challenges of AGI`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "IjtKDQRLrylI",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 18:54:24-- https://arxiv.org/pdf/2409.18486\n",
"Resolving arxiv.org (arxiv.org)... 151.101.67.42, 151.101.131.42, 151.101.3.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.67.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 13986265 (13M) [application/pdf]\n",
"Saving to: o1.pdf\n",
"\n",
"o1.pdf 100%[===================>] 13.34M 11.8MB/s in 1.1s \n",
"\n",
"2024-12-05 18:54:26 (11.8 MB/s) - o1.pdf saved [13986265/13986265]\n",
"\n"
]
}
],
"source": [
"!wget \"https://arxiv.org/pdf/2409.18486\" -O \"o1.pdf\""
]
},
{
"cell_type": "markdown",
"id": "4e29a9d7-5bd9-4fb8-8ec1-4c128a748662",
"metadata": {},
"source": [
"## Initialize LlamaParse\n",
"\n",
"Initialize LlamaParse in multimodal mode, and specify the vendor.\n",
"\n",
"**NOTE**: optionally you can specify the Anthropic/ OpenAI API key. If you choose to do so LlamaParse will only charge you 1 credit (0.3c) per page. \n",
"\n",
"\n",
"Using your own API key may incur additional costs from your model provider and could result in failed pages or documents if you do not have sufficient usage limits."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"md\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
"cell_type": "markdown",
"id": "1b5d6da6",
"metadata": {},
"source": [
"### With anthropic-sonnet-3.5"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f2e9d9cf-8189-4fcb-b34f-cde6cc0b59c8",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id dd9d5e0f-160e-486a-89a2-6005e5a1c2ac\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model_name=\"anthropic-sonnet-3.5\",\n",
" target_pages=\"24\"\n",
" # invalidate_cache=True\n",
")\n",
"json_objs = parser.get_json_result(\"o1.pdf\")\n",
"json_list = json_objs[0][\"pages\"]\n",
"docs = get_text_nodes(json_list)"
]
},
{
"cell_type": "markdown",
"id": "4f3c51b0-7878-48d7-9bc3-02b516500128",
"metadata": {},
"source": [
"### With GPT-4o\n",
"\n",
"For comparison, we will also parse the document using GPT-4o."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6fc3f258-50ae-4988-b904-c105463a498f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 6a4dea44-4f90-406b-b290-9e98620b1232\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model=\"openai-gpt4o\",\n",
" target_pages=\"24\",\n",
" # invalidate_cache=True\n",
")\n",
"json_objs_gpt4o = parser_gpt4o.get_json_result(\"o1.pdf\")\n",
"json_list_gpt4o = json_objs_gpt4o[0][\"pages\"]\n",
"docs_gpt4o = get_text_nodes(json_list_gpt4o)"
]
},
{
"cell_type": "markdown",
"id": "44c20f7a-2901-4dd0-b635-a4b33c5664c1",
"metadata": {},
"source": [
"### View Results\n",
"\n",
"Let's visualize the results along with the original document page."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "778698aa-da7e-4081-b3b5-0372f228536f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|-----------------|----------------------------------|\n",
"| Attribute | Value | Basic Personal Information: Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia. |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White | Biomarker Measurements: The subject's genetic profile includes an ApoE4 status of 0.0... |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... | Cognitive and Neurofunctional Assessments: The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 | Volumetric Data: Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 54422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 2782.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0.... |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-preview's performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the model's logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-preview's solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-preview's solutions. By comparing o1-preview's solutions with the dataset's solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-preview's overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using Sonnet-3.5\n",
"print(docs[0].get_content(metadata_mode=\"all\"))"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1511a30f-3efc-4142-9668-7dc056a24d0c",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|----------------|--------------------------------|\n",
"| **Attribute** | **Value** |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"**Basic Personal Information:** Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia.\n",
"\n",
"**Biomarker Measurements:** The subject's genetic profile includes an ApoE4 status of 0.0...\n",
"\n",
"**Cognitive and Neurofunctional Assessments:** The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall...\n",
"\n",
"**Volumetric Data:** Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross-Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 84422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 27820.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0...\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"----\n",
"\n",
"Skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-previews performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the models logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-previews solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-previews solutions. By comparing o1-previews solutions with the datasets solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-previews overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using GPT-4o\n",
"print(docs_gpt4o[0].get_content(metadata_mode=\"all\"))"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
@@ -0,0 +1,170 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_starter_parse_selected_pages.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Parse Selected Pages \n",
"\n",
"In this notebook we will demonstrate how to parse selected pages in a document using LlamaParse."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Installation\n",
"\n",
"Here we install `llama-parse` used for parsing the document"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Set API Key"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# llama-parse is async-first, running the async code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<YOUR LLAMACLOUD API KEY>\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Download Data\n",
"\n",
"Here we download Uber 2021 10K SEC filings data for the demonstration."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 11:40:59-- https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf\n",
"Resolving raw.githubusercontent.com (raw.githubusercontent.com)... 2606:50c0:8000::154, 2606:50c0:8002::154, 2606:50c0:8003::154, ...\n",
"Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|2606:50c0:8000::154|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 1880483 (1.8M) [application/octet-stream]\n",
"Saving to: ./uber_2021.pdf\n",
"\n",
"./uber_2021.pdf 100%[===================>] 1.79M --.-KB/s in 0.1s \n",
"\n",
"2024-12-05 11:40:59 (14.2 MB/s) - ./uber_2021.pdf saved [1880483/1880483]\n",
"\n"
]
}
],
"source": [
"!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf' -O './uber_2021.pdf'"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Parse the PDF file in selected pages\n",
"\n",
"Here we will parse the PDF file in selected pages and get the text in `markdown` format."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id ad1087c1-b085-4dc7-9aa8-d13cdd440f2b\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(target_pages=\"0,1,2\", result_type=\"markdown\")\n",
"\n",
"documents = parser.load_data(\"./uber_2021.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[Document(id_='d0b34f4a-27ef-48e2-a92a-386e5e265f4c', embedding=None, metadata={}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, metadata_template='{key}: {value}', metadata_separator='\\n', text='# UNITED STATES SECURITIES AND EXCHANGE COMMISSION\\n\\n# Washington, D.C. 20549\\n\\n# FORM 10-K\\n\\n(Mark One)\\n\\n☒ ANNUAL REPORT PURSUANT TO SECTION 13 OR 15(d) OF THE SECURITIES EXCHANGE ACT OF 1934\\n\\nFor the fiscal year ended December 31, 2021\\n\\nOR\\n\\n☐ TRANSITION REPORT PURSUANT TO SECTION 13 OR 15(d) OF THE SECURITIES EXCHANGE ACT OF 1934\\n\\nFor the transition period from _____ to _____\\n\\nCommission File Number: 001-38902\\n\\n# UBER TECHNOLOGIES, INC.\\n\\n(Exact name of registrant as specified in its charter)\\n\\nDelaware\\n\\n45-2647441\\n\\n(State or other jurisdiction of incorporation or organization) (I.R.S. Employer Identification No.)\\n\\n1515 3rd Street\\n\\nSan Francisco, California 94158\\n\\n(Address of principal executive offices, including zip code)\\n\\n(415) 612-8582\\n\\n(Registrants telephone number, including area code)\\n\\n# Securities registered pursuant to Section 12(b) of the Act:\\n\\n|Title of each class|Trading Symbol(s)|Name of each exchange on which registered|\\n|---|---|---|\\n|Common Stock, par value $0.00001 per share|UBER|New York Stock Exchange|\\n\\nSecurities registered pursuant to Section 12(g) of the Act: None\\n\\nIndicate by check mark whether the registrant is a well-known seasoned issuer, as defined in Rule 405 of the Securities Act. Yes ☒ No ☐\\n\\nIndicate by check mark whether the registrant is not required to file reports pursuant to Section 13 or Section 15(d) of the Act. Yes ☐ No ☒\\n\\nIndicate by check mark whether the registrant (1) has filed all reports required to be filed by Section 13 or 15(d) of the Securities Exchange Act of 1934 during the preceding 12 months (or for such shorter period that the registrant was required to file such reports), and (2) has been subject to such filing requirements for the past 90 days. Yes ☒ No ☐\\n\\nIndicate by check mark whether the registrant has submitted electronically every Interactive Data File required to be submitted pursuant to Rule 405 of Regulation S-T (§232.405 of this chapter) during the preceding 12 months (or for such shorter period that the registrant was required to submit such files). Yes ☒ No ☐\\n\\nIndicate by check mark whether the registrant is a large accelerated filer, an accelerated filer, a non-accelerated filer, a smaller reporting company, or an emerging growth company. See the definitions of “large accelerated filer,” “accelerated filer,” “smaller reporting company,” and “emerging growth company” in Rule 12b-2 of the Exchange Act.', mimetype='text/plain', start_char_idx=None, end_char_idx=None, metadata_seperator='\\n', text_template='{metadata_str}\\n\\n{content}'),\n",
" Document(id_='253b1141-a260-466e-b164-b39df67ef799', embedding=None, metadata={}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, metadata_template='{key}: {value}', metadata_separator='\\n', text=\"# Large accelerated filer\\n\\n☒\\n\\n# Accelerated filer\\n\\n☐\\n\\n# Non-accelerated filer\\n\\n☐\\n\\n# Smaller reporting company\\n\\n☐\\n\\n# Emerging growth company\\n\\n☐\\n\\nIf an emerging growth company, indicate by check mark if the registrant has elected not to use the extended transition period for complying with any new or revised financial accounting standards provided pursuant to Section 13(a) of the Exchange Act.\\n\\n☐\\n\\nIndicate by check mark whether the registrant has filed a report on and attestation to its managements assessment of the effectiveness of its internal control over financial reporting under Section 404(b) of the Sarbanes-Oxley Act (15 U.S.C. 7262(b)) by the registered public accounting firm that prepared or issued\\n\\n☒\\n\\nIndicate by check mark whether the registrant is a shell company (as defined in Rule 12b-2 of the Exchange Act). Yes\\n\\n☐\\n\\nNo\\n\\n☒\\n\\nThe aggregate market value of the voting and non-voting common equity held by non-affiliates of the registrant as of June 30, 2021, the last business day of the registrant's most recently completed second fiscal quarter, was approximately $90.5 billion based upon the closing price reported for such date on the New York Stock Exchange.\\n\\nThe number of shares of the registrant's common stock outstanding as of February 22, 2022 was 1,954,464,088.\\n\\n# DOCUMENTS INCORPORATED BY REFERENCE\\n\\nPortions of the registrants Definitive Proxy Statement relating to the Annual Meeting of Stockholders are incorporated by reference into Part III of this Annual Report on Form 10-K where indicated. Such Definitive Proxy Statement will be filed with the Securities and Exchange Commission within 120 days after the end of the registrants fiscal year ended December 31, 2021.\", mimetype='text/plain', start_char_idx=None, end_char_idx=None, metadata_seperator='\\n', text_template='{metadata_str}\\n\\n{content}'),\n",
" Document(id_='ad988239-3ab5-498d-85ba-a29241db24d4', embedding=None, metadata={}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, metadata_template='{key}: {value}', metadata_separator='\\n', text='# UBER TECHNOLOGIES, INC.\\n\\n# TABLE OF CONTENTS\\n\\n|Special Note Regarding Forward-Looking Statements|2|\\n|---|---|\\n|PART I|PART I|\\n|Item 1. Business|4|\\n|Item 1A. Risk Factors|11|\\n|Item 1B. Unresolved Staff Comments|46|\\n|Item 2. Properties|46|\\n|Item 3. Legal Proceedings|46|\\n|Item 4. Mine Safety Disclosures|47|\\n|PART II|PART II|\\n|Item 5. Market for Registrants Common Equity, Related Stockholder Matters and Issuer Purchases of Equity Securities|47|\\n|Item 6. [Reserved]|48|\\n|Item 7. Managements Discussion and Analysis of Financial Condition and Results of Operations|48|\\n|Item 7A. Quantitative and Qualitative Disclosures About Market Risk|69|\\n|Item 8. Financial Statements and Supplementary Data|70|\\n|Item 9. Changes in and Disagreements with Accountants on Accounting and Financial Disclosure|146|\\n|Item 9A. Controls and Procedures|147|\\n|Item 9B. Other Information|147|\\n|Item 9C. Disclosure Regarding Foreign Jurisdictions that Prevent Inspections|147|\\n|PART III|PART III|\\n|Item 10. Directors, Executive Officers and Corporate Governance|147|\\n|Item 11. Executive Compensation|147|\\n|Item 12. Security Ownership of Certain Beneficial Owners and Management and Related Stockholder Matters|148|\\n|Item 13. Certain Relationships and Related Transactions, and Director Independence|148|\\n|Item 14. Principal Accounting Fees and Services|148|\\n|PART IV|PART IV|\\n|Item 15. Exhibits, Financial Statement Schedules|148|\\n|Item 16. Form 10-K Summary|148|\\n|Exhibit Index|149|\\n|Signatures|152|', mimetype='text/plain', start_char_idx=None, end_char_idx=None, metadata_seperator='\\n', text_template='{metadata_str}\\n\\n{content}')]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"documents"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
+1515
View File
File diff suppressed because one or more lines are too long
Binary file not shown.

After

Width:  |  Height:  |  Size: 195 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 363 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 343 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 185 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 254 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 650 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 173 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 200 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 115 KiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 580 KiB

@@ -165,7 +165,18 @@
"execution_count": null,
"id": "ef82a985-4088-4bb7-9a21-0318e1b9207d",
"metadata": {},
"outputs": [],
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Parsing text...\n",
"Started parsing the file under job_id 62f157a9-9ef9-4e5b-95ac-67093fa25800\n",
"..........Parsing PDF file...\n",
"Started parsing the file under job_id 1ddd5654-062b-4e19-b488-d66efc9c509d\n"
]
}
],
"source": [
"print(f\"Parsing text...\")\n",
"docs_text = parser_text.load_data(\"data/conocophillips.pdf\")\n",
@@ -174,42 +185,36 @@
"md_json_list = md_json_objs[0][\"pages\"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7506b603-c01f-45de-b354-4a0728dde03c",
"metadata": {},
"outputs": [],
"source": [
"print(docs_text[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "5318fb7b-fe6a-4a8a-b82e-4ed7b4512c37",
"metadata": {},
"outputs": [],
"source": [
"print(md_json_list[10][\"md\"])"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7a46a73e-c6e2-4b0b-bd10-31b0d3e4b70f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"dict_keys(['page', 'text', 'md', 'images', 'items'])\n"
"# Commitment to Disciplined Reinvestment Rate\n",
"\n",
"| Period | Description | Reinvestment Rate | WTI Average |\n",
"|--------------|--------------------------------------|-------------------|-------------|\n",
"| 2012-2016 | Industry Growth Focus | >100% | ~$75/BBL |\n",
"| 2017-2022 | ConocoPhillips Strategy Reset | <60% | ~$63/BBL |\n",
"| 2023E | | | at $80/BBL |\n",
"| 2024-2028 | Disciplined Reinvestment Rate | ~50% | at $60/BBL |\n",
"| 2029-2032 | | ~6% CFO CAGR | at $60/BBL |\n",
"\n",
"- **Historic Reinvestment Rate**: Gray bars\n",
"- **Reinvestment Rate at $60/BBL WTI**: Blue bars\n",
"- **Reinvestment Rate at $80/BBL WTI**: Dashed blue lines\n",
"\n",
"Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n"
]
}
],
"source": [
"print(md_json_list[1].keys())"
"print(md_json_list[10][\"md\"])"
]
},
{
@@ -299,7 +304,7 @@
" image_files = _get_sorted_image_files(image_dir) if image_dir is not None else None\n",
" md_texts = [d[\"md\"] for d in json_dicts] if json_dicts is not None else None\n",
"\n",
" doc_chunks = docs[0].text.split(\"---\")\n",
" doc_chunks = [c for d in docs for c in d.text.split(\"---\")]\n",
" for idx, doc_chunk in enumerate(doc_chunks):\n",
" chunk_metadata = {\"page_num\": idx + 1}\n",
" if image_files is not None:\n",
@@ -339,25 +344,23 @@
"output_type": "stream",
"text": [
"page_num: 11\n",
"image_path: data_images/d9137e19-3974-4b5d-998f-dac0cf29dd9d-page-10.jpg\n",
"image_path: data_images/1ddd5654-062b-4e19-b488-d66efc9c509d-page_39.jpg\n",
"parsed_text_markdown: # Commitment to Disciplined Reinvestment Rate\n",
"\n",
"| Year | Reinvestment Rate | WTI Average Price | Reinvestment Rate at $60/BBL WTI | Reinvestment Rate at $80/BBL WTI |\n",
"|------------|-------------------|-------------------|----------------------------------|----------------------------------|\n",
"| 2012-2016 | >100% | ~$75/BBL | | |\n",
"| 2017-2022 | <60% | ~$63/BBL | | |\n",
"| 2023E | | | | at $80/BBL WTI |\n",
"| 2024-2028 | | | at $60/BBL WTI | at $80/BBL WTI |\n",
"| 2029-2032 | | | at $60/BBL WTI | at $80/BBL WTI |\n",
"| Period | Description | Reinvestment Rate | WTI Average |\n",
"|--------------|--------------------------------------|-------------------|-------------|\n",
"| 2012-2016 | Industry Growth Focus | >100% | ~$75/BBL |\n",
"| 2017-2022 | ConocoPhillips Strategy Reset | <60% | ~$63/BBL |\n",
"| 2023E | | | at $80/BBL |\n",
"| 2024-2028 | Disciplined Reinvestment Rate | ~50% | at $60/BBL |\n",
"| 2029-2032 | | ~6% CFO CAGR | at $60/BBL |\n",
"\n",
"**Disciplined Reinvestment Rate is the Foundation for Superior Returns on and of Capital, while Driving Durable CFO Growth**\n",
"- **Historic Reinvestment Rate**: Gray bars\n",
"- **Reinvestment Rate at $60/BBL WTI**: Blue bars\n",
"- **Reinvestment Rate at $80/BBL WTI**: Dashed blue lines\n",
"\n",
"- ~50% 10-Year Reinvestment Rate\n",
"- ~6% CFO CAGR 2024-2032 at $60/BBL WTI Mid-Cycle Planning Price\n",
"\n",
"**Note:** Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n",
"parsed_text: \n",
"Commitment to Disciplined Reinvestment Rate\n",
"Reinvestment rate and cash from operations (CFO) are non-GAAP measures. Definitions and reconciliations are included in the Appendix.\n",
"parsed_text: Commitment to Disciplined Reinvestment Rate\n",
" Industry ConocoPhillips\n",
" Strategy Reset Disciplined Reinvestment Rate is the Foundation for Superior\n",
" Growth Focus Returns on and of Capital, while Driving Durable CFO Growth\n",
@@ -374,7 +377,7 @@
" 0%\n",
" 2012-2016 2017-2022 2023E 2024-2028 2029-2032\n",
" Historic Reinvestment Rate Reinvestment Rate at $60/BBL WTI Reinvestment Rate at $80/BBL WTI\n",
" Reinvestment rate andcashfrom operations (CFO) are non-GAAP measures: Definitions and reconciliations are included in the Appendix ConocoPhillips\n"
" Reinvestment rate and cash from operations (CFO) are non-GAAP measures: Definitions and reconciliations are included in the Appendix ConocoPhillips\n"
]
}
],
@@ -397,7 +400,17 @@
"execution_count": null,
"id": "6ea53c31-0e38-421c-8d9b-0e3adaa1677e",
"metadata": {},
"outputs": [],
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
"/Users/jerryliu/Programming/gpt_index/.venv/lib/python3.10/site-packages/tiktoken/core.py:50: RuntimeWarning: coroutine 'LlamaParse.aload_data' was never awaited\n",
" self._core_bpe = _tiktoken.CoreBPE(mergeable_ranks, special_tokens, pat_str)\n",
"RuntimeWarning: Enable tracemalloc to get the object allocation traceback\n"
]
}
],
"source": [
"import os\n",
"from llama_index.core import (\n",
@@ -588,7 +601,7 @@
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3$40 GKA GWA\n",
" 3 $40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
@@ -599,7 +612,7 @@
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other MalaysiaChina Surmont\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
@@ -687,70 +700,126 @@
{
"cell_type": "code",
"execution_count": null,
"id": "1cdce5d8-6bb3-4cd3-929d-1cec249d9052",
"id": "d78e53cf-35cb-4ef8-b03e-1b47ba15ae64",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Added user message to memory: How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\n",
"Added user message to memory: Tell me about the diverse geographies where Conoco Phillips has a production base\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips capex/EUR in the Delaware Basin\"}\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips production base geographies\"}\n",
"=== Function Output ===\n",
"The ConocoPhillips capex/EUR in the Delaware Basin is $10/BOE.\n",
"ConocoPhillips' production base geographies include:\n",
"\n",
"I obtained this information from the image provided. The image clearly shows a bar chart under the section \"Delaware Basin Well Capex/EUR ($/BOE)\" where ConocoPhillips is listed with a capex/EUR of $10/BOE. This information is consistent with the parsed markdown text, which also lists ConocoPhillips' capex/EUR as $10/BOE in the Delaware Basin. There are no discrepancies between the image and the parsed markdown text in this case.\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"competitors capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"The competitors' Capex/EUR in the Delaware Basin can be found in the image on the slide titled \"Delaware: Vast Inventory with Proven Track Record of Performance.\" The relevant information is presented in a bar chart under the section \"Delaware Basin Well Capex/EUR ($/BOE)\".\n",
"1. **Lower 48** (Permian, Eagle Ford, Bakken, Other)\n",
"2. **Alaska** (GKA, GWA, GPA, WNS)\n",
"3. **EMENA** (Norway, Libya, Qatar)\n",
"4. **Asia Pacific** (APLNG, Malaysia, China)\n",
"5. **Canada** (Montney, Surmont)\n",
"\n",
"Here are the details:\n",
"\n",
"- ConocoPhillips: $10/BOE\n",
"- Competitor 1: $15/BOE\n",
"- Competitor 2: $20/BOE\n",
"- Competitor 3: $25/BOE\n",
"- Competitor 4: $30/BOE\n",
"- Competitor 5: $35/BOE\n",
"- Competitor 6: $40/BOE\n",
"- Competitor 7: $45/BOE\n",
"\n",
"This information was obtained directly from the image, which provides a clear visual representation of the Capex/EUR values for ConocoPhillips and its competitors in the Delaware Basin. The parsed markdown text also confirms these values, ensuring consistency between the image and the text.\n",
"This information was derived from the image on page 14, which provides a detailed breakdown of the diverse production base and the regions involved. The parsed markdown and raw text also support this information, but the image provides the clearest and most comprehensive view. There are no discrepancies between the image and the parsed text in this case.\n",
"=== LLM Response ===\n",
"The capital expenditure per estimated ultimate recovery (capex/EUR) for ConocoPhillips in the Delaware Basin is $10 per barrel of oil equivalent (BOE). When compared to its competitors, ConocoPhillips has a significantly lower capex/EUR. Here are the capex/EUR values for ConocoPhillips and its competitors:\n",
"ConocoPhillips has a diverse production base spread across various geographies, including:\n",
"\n",
"- **ConocoPhillips**: $10/BOE\n",
"- **Competitor 1**: $15/BOE\n",
"- **Competitor 2**: $20/BOE\n",
"- **Competitor 3**: $25/BOE\n",
"- **Competitor 4**: $30/BOE\n",
"- **Competitor 5**: $35/BOE\n",
"- **Competitor 6**: $40/BOE\n",
"- **Competitor 7**: $45/BOE\n",
"1. **Lower 48**:\n",
" - Permian Basin\n",
" - Eagle Ford\n",
" - Bakken\n",
" - Other regions within the continental United States\n",
"\n",
"This data indicates that ConocoPhillips has a more cost-efficient operation in the Delaware Basin compared to its competitors.\n",
"The capital expenditure per estimated ultimate recovery (capex/EUR) for ConocoPhillips in the Delaware Basin is $10 per barrel of oil equivalent (BOE). When compared to its competitors, ConocoPhillips has a significantly lower capex/EUR. Here are the capex/EUR values for ConocoPhillips and its competitors:\n",
"2. **Alaska**:\n",
" - Greater Kuparuk Area (GKA)\n",
" - Greater Prudhoe Area (GPA)\n",
" - Greater Willow Area (GWA)\n",
" - Western North Slope (WNS)\n",
"\n",
"- **ConocoPhillips**: $10/BOE\n",
"- **Competitor 1**: $15/BOE\n",
"- **Competitor 2**: $20/BOE\n",
"- **Competitor 3**: $25/BOE\n",
"- **Competitor 4**: $30/BOE\n",
"- **Competitor 5**: $35/BOE\n",
"- **Competitor 6**: $40/BOE\n",
"- **Competitor 7**: $45/BOE\n",
"3. **EMENA (Europe, Middle East, and North Africa)**:\n",
" - Norway\n",
" - Libya\n",
" - Qatar\n",
"\n",
"This data indicates that ConocoPhillips has a more cost-efficient operation in the Delaware Basin compared to its competitors.\n"
"4. **Asia Pacific**:\n",
" - Australia Pacific LNG (APLNG)\n",
" - Malaysia\n",
" - China\n",
"\n",
"5. **Canada**:\n",
" - Montney\n",
" - Surmont\n",
"\n",
"These regions highlight the global reach and diverse geographical footprint of ConocoPhillips' production operations.\n",
"Added user message to memory: Tell me about the diverse geographies where Conoco Phillips has a production base\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"diverse geographies where Conoco Phillips has a production base\"}\n",
"=== Function Output ===\n",
"ConocoPhillips has a diverse production base that includes the Lower 48 (Permian, Bakken, Eagle Ford), Alaska, Canada (Montney, Surmont), EMENA (Norway, Libya), Asia Pacific (Malaysia, China, APLNG), and Qatar.\n",
"=== LLM Response ===\n",
"ConocoPhillips has a diverse production base spanning several key geographies:\n",
"\n",
"1. **Lower 48 (United States)**: This includes major production areas such as the Permian Basin, Bakken Formation, and Eagle Ford Shale.\n",
"2. **Alaska**: Significant operations in the North Slope region.\n",
"3. **Canada**: Operations in the Montney Formation and the Surmont oil sands project.\n",
"4. **EMENA (Europe, Middle East, and North Africa)**: Notable operations in Norway and Libya.\n",
"5. **Asia Pacific**: Includes operations in Malaysia, China, and the Australia Pacific LNG (APLNG) project.\n",
"6. **Qatar**: Involvement in the country's energy sector.\n",
"\n",
"These regions highlight the company's extensive and varied geographical footprint in the energy production industry.\n"
]
}
],
"source": [
"# response = agent.query(\"Tell me about the different regions and subregions where Conoco Phillips has a production base.\")\n",
"response = agent.query(\n",
" \"How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\"\n",
"query = (\n",
" \"Tell me about the diverse geographies where Conoco Phillips has a production base\"\n",
")\n",
"response = agent.query(query)\n",
"base_response = base_agent.query(query)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "355d2aa4-c26f-480e-b512-4446acbd9227",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"ConocoPhillips has a diverse production base spread across various geographies, including:\n",
"\n",
"1. **Lower 48**:\n",
" - Permian Basin\n",
" - Eagle Ford\n",
" - Bakken\n",
" - Other regions within the continental United States\n",
"\n",
"2. **Alaska**:\n",
" - Greater Kuparuk Area (GKA)\n",
" - Greater Prudhoe Area (GPA)\n",
" - Greater Willow Area (GWA)\n",
" - Western North Slope (WNS)\n",
"\n",
"3. **EMENA (Europe, Middle East, and North Africa)**:\n",
" - Norway\n",
" - Libya\n",
" - Qatar\n",
"\n",
"4. **Asia Pacific**:\n",
" - Australia Pacific LNG (APLNG)\n",
" - Malaysia\n",
" - China\n",
"\n",
"5. **Canada**:\n",
" - Montney\n",
" - Surmont\n",
"\n",
"These regions highlight the global reach and diverse geographical footprint of ConocoPhillips' production operations.\n"
]
}
],
"source": [
"print(str(response))"
]
},
@@ -764,85 +833,82 @@
"name": "stdout",
"output_type": "stream",
"text": [
"page_num: 38\n",
"image_path: data_images/d9137e19-3974-4b5d-998f-dac0cf29dd9d-page-37.jpg\n",
"parsed_text_markdown: # Delaware: Vast Inventory with Proven Track Record of Performance\n",
"page_num: 14\n",
"image_path: data_images/1ddd5654-062b-4e19-b488-d66efc9c509d-page_12.jpg\n",
"parsed_text_markdown: # Our Differentiated Portfolio: Deep, Durable and Diverse\n",
"\n",
"## Prolific Acreage Spanning Over ~659,000 Net Acres¹\n",
"## ~20 BBOE of Resource\n",
"Under $40/BBL Cost of Supply\n",
"\n",
"![Map of Delaware Basin](image)\n",
"### ~ $32/BBL\n",
"Average Cost of Supply\n",
"\n",
"### Total 10-Year Operated Permian Inventory\n",
"### WTI Cost of Supply ($/BBL)\n",
"\n",
"- Delaware Basin: 65%\n",
"- Midland Basin: 35%\n",
"| Cost ($/BBL) | Resource (BBOE) |\n",
"|--------------|-----------------|\n",
"| $0 | 0 |\n",
"| $10 | |\n",
"| $20 | |\n",
"| $30 | |\n",
"| $40 | |\n",
"| $50 | |\n",
"\n",
"### High Single-Digit Production Growth\n",
"- **Legend:**\n",
" - Lower 48\n",
" - Canada\n",
" - Alaska\n",
" - EMENA\n",
" - Asia Pacific\n",
"\n",
"## 12-Month Cumulative Production³ (BOE/FT)\n",
"*Costs assume a mid-cycle price environment of $60/BBL WTI.*\n",
"\n",
"| Months | 2019 | 2020 | 2021 | 2022 |\n",
"|--------|------|------|------|------|\n",
"| 1 | 0 | 0 | 0 | 0 |\n",
"| 2 | 5 | 6 | 7 | 8 |\n",
"| 3 | 10 | 12 | 14 | 16 |\n",
"| 4 | 15 | 18 | 21 | 24 |\n",
"| 5 | 20 | 24 | 28 | 32 |\n",
"| 6 | 25 | 30 | 35 | 40 |\n",
"| 7 | 30 | 36 | 42 | 48 |\n",
"| 8 | 35 | 42 | 49 | 56 |\n",
"| 9 | 40 | 48 | 56 | 64 |\n",
"| 10 | 45 | 54 | 63 | 72 |\n",
"| 11 | 50 | 60 | 70 | 80 |\n",
"| 12 | 55 | 66 | 77 | 88 |\n",
"## Diverse Production Base\n",
"10-Year Plan Cumulative Production (BBOE)\n",
"\n",
"~30% Improved Performance from 2019 to 2022\n",
"\n",
"## Delaware Basin Well Capex/EUR⁴ ($/BOE)\n",
"\n",
"| Company | Capex/EUR |\n",
"|------------------|-----------|\n",
"| ConocoPhillips | 10 |\n",
"| Competitor 1 | 15 |\n",
"| Competitor 2 | 20 |\n",
"| Competitor 3 | 25 |\n",
"| Competitor 4 | 30 |\n",
"| Competitor 5 | 35 |\n",
"| Competitor 6 | 40 |\n",
"| Competitor 7 | 45 |\n",
"\n",
"---\n",
"\n",
"¹ Unconventional acres. \n",
"² Source: Enverus and ConocoPhillips (March 2023). \n",
"³ Source: Enverus (March 2023) based on wells online year. \n",
"⁴ Source: Enverus (March 2023). Average single well capex/EUR. Top eight public operators based on wells online in years 2021-2022, greater than 50% oil weight. COP based on COP well design. Competitors include: CVX, DVN, EOG, MTDR, OXY, PR and XOM.\n",
"parsed_text: \n",
"Delaware: Vast Inventory with Proven Track Record of Performance\n",
" New Prolific Acreage Spanning Over 12-Month Cumulative Production? (BOE/FT)\n",
" Mexico 659,000 Net Acres' 40\n",
" Texas 3828\n",
" 30 2019\n",
" 20 30%\n",
" 10 Improved Performancefrom 2019 to 2022\n",
" Total\n",
" Permian Inventory\n",
" 10-Year Operated\n",
" 2 10 11 12\n",
" Months\n",
" Delaware Basin Well Capex/EUR4 (S/BOE)\n",
" 65% 25\n",
" Delaware Basin 20\n",
" Midland Basin 15\n",
" Low HighCost of Supplyz 10 ConocoPhillips\n",
" High Single-Digit Production Growth\n",
" \"Unconventional acres. 2Source: Enverus and ConocoPhillips (March 2023). 3SourceEnverus (March 2023) based on wells online year: \"Source; Enverus (March 2023). Average single well capex/EUR Top eight public operators based on\n",
"wells online in years 2021-2022, greater than 50% oil weight; COP based on COP well design: Competitors include; CVX DVN, EOG; MTDR, OXY, PR and XOM: ConocoPhillips\n"
"| Region | Sub-region |\n",
"|--------------|-----------------|\n",
"| Lower 48 | Permian |\n",
"| | Eagle Ford |\n",
"| | Bakken |\n",
"| | Other |\n",
"| Alaska | GKA |\n",
"| | GWA |\n",
"| | GPA |\n",
"| | WNS |\n",
"| EMENA | Norway |\n",
"| | Libya |\n",
"| | Qatar |\n",
"| Asia Pacific | APLNG |\n",
"| | Malaysia |\n",
"| | China |\n",
"| Canada | Montney |\n",
"| | Surmont |\n",
"parsed_text: Our Differentiated Portfolio: Deep; Durable and Diverse\n",
" 20 BBOE of Resource Diverse Production Base\n",
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3 $40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
" 8 $20\n",
" E Qatar Libya\n",
" Asia Pacific Canada\n",
" $10 Permian\n",
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
]
}
],
"source": [
"print(response.source_nodes[0].get_content(metadata_mode=\"all\"))"
"print(response.source_nodes[7].get_content(metadata_mode=\"all\"))"
]
},
{
@@ -855,26 +921,20 @@
"name": "stdout",
"output_type": "stream",
"text": [
"Added user message to memory: How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"Conoco Phillips capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20/BOE.\n",
"=== Calling Function ===\n",
"Calling function: vector_tool with args: {\"input\": \"competitors capex/EUR in the Delaware Basin\"}\n",
"=== Function Output ===\n",
"The average single well capex/EUR for competitors in the Delaware Basin is between $10 and $25 per BOE.\n",
"=== LLM Response ===\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20 per BOE. In comparison, the average capex/EUR for competitors in the Delaware Basin ranges between $10 and $25 per BOE. This places ConocoPhillips' capex/EUR towards the higher end of the competitive range.\n",
"ConocoPhillips' capex/EUR in the Delaware Basin is approximately $20 per BOE. In comparison, the average capex/EUR for competitors in the Delaware Basin ranges between $10 and $25 per BOE. This places ConocoPhillips' capex/EUR towards the higher end of the competitive range.\n"
"ConocoPhillips has a diverse production base spanning several key geographies:\n",
"\n",
"1. **Lower 48 (United States)**: This includes major production areas such as the Permian Basin, Bakken Formation, and Eagle Ford Shale.\n",
"2. **Alaska**: Significant operations in the North Slope region.\n",
"3. **Canada**: Operations in the Montney Formation and the Surmont oil sands project.\n",
"4. **EMENA (Europe, Middle East, and North Africa)**: Notable operations in Norway and Libya.\n",
"5. **Asia Pacific**: Includes operations in Malaysia, China, and the Australia Pacific LNG (APLNG) project.\n",
"6. **Qatar**: Involvement in the country's energy sector.\n",
"\n",
"These regions highlight the company's extensive and varied geographical footprint in the energy production industry.\n"
]
}
],
"source": [
"# base_response = base_agent.query(\"Tell me about the different regions and subregions where Conoco Phillips has a production base.\")\n",
"base_response = base_agent.query(\n",
" \"How does the Conoco Phillips capex/EUR in the delaware basin compare against other competitors?\"\n",
")\n",
"print(str(base_response))"
]
},
@@ -888,30 +948,31 @@
"name": "stdout",
"output_type": "stream",
"text": [
"Deep, Durable and Diverse Portfolio with Significant Growth Runway\n",
" 1,2002022 Lower 48 Unconventional Production' (MBOED S50 ~S32/BBL\n",
" 000 ConocoPhillips Cost of SupplyAverage\n",
" 00 S40\n",
" 500 3\n",
" 400 1 S30\n",
" 200\n",
" 5\n",
" 15,000ConocoPhillipsNet Remaining Well Inventory? 1 S20\n",
" 12,000 S10\n",
" 000\n",
" 0o0 SO\n",
" 3,000 10\n",
" Resource (BBOE)\n",
" Delaware Basin Midland Basin Eagle Ford Bakken Other\n",
" Largest Lower 48 Unconventional Producer; Growing into the Next Decade\n",
" onshore operated inventory that achieves 15% IRR at $SO/BBL WTI, Competitors include CVX, DVN, EOG, FANG, MRO, OXY, PXD,and XOM:\n",
" Source: Wood Mackenzie Lower 48 Unconventional Plays 2022 ProductionCompetitors include CVX, DVN; EOG, FANG, MRO, OXY, PXD and XOM; greaterthan50% liquids weight: ?Source: Wood Mackenzie (March 2023), Lower 48\n",
" ConocoPhillips\n"
"Our Differentiated Portfolio: Deep; Durable and Diverse\n",
" 20 BBOE of Resource Diverse Production Base\n",
" Under $40/BBL Cost of Supply 10-Year Plan Cumulative Production (BBOE)\n",
" S50 S32/BBL Lower 48 Alaska\n",
" Average Cost of Supply\n",
" 3 $40 GKA GWA\n",
" GPA WNS\n",
" $30 EMENA\n",
" 3 Norway\n",
" 8 $20\n",
" E Qatar Libya\n",
" Asia Pacific Canada\n",
" $10 Permian\n",
" APLNG Montney\n",
" S0\n",
" 10 15 20 Bakken\n",
" Resource (BBOE) Eagle Ford Other Malaysia ChinaSurmont\n",
" Lower 48 Canada Alaska EMENA Asia Pacific\n",
"Costs assumemid-cycle price environment of S60/BBL WTI:\n",
" ConocoPhillips\n"
]
}
],
"source": [
"print(base_response.source_nodes[0].get_content(metadata_mode=\"llm\"))"
"print(base_response.source_nodes[1].get_content(metadata_mode=\"all\"))"
]
}
],
@@ -11,6 +11,8 @@
"\n",
"In this cookbook we show you how to build a multimodal report generation agent from a bank of research reports. We use the a set of ICLR papers (which were also used as the dataset in our [DeepLearning.ai course](https://www.deeplearning.ai/short-courses/building-agentic-rag-with-llamaindex/?utm_campaign=llamaindexC2-launch&utm_medium=headband&utm_source=dlai-homepage).\n",
"\n",
"![](multimodal_report_generation_agent_img.png)\n",
"\n",
"We use our workflow abstraction to define an agentic system that contains two main phases: a research phase that pulls in relevant files through chunk-level or file-level retrieval, and then a blog generation phase that synthesizes the final report."
]
},
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 350 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 47 KiB

@@ -0,0 +1,602 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/parsing_instructions.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"# Parsing documents with Instructions\n",
"\n",
"Parsing instructions allow you to guide our parsing model in the same way you would instruct an LLM.\n",
"\n",
"These instructions can be useful for improving the parser's performance on complex document layouts, extracting data in a specific format, or transforming the document in other ways.\n",
"\n",
"### Why This Matters:\n",
"Traditional document parsing can be rigid and error-prone, often missing crucial context and nuances in complex layouts. Our instruction-based parsing allows you to:\n",
"\n",
"1. Extract specific information with pinpoint accuracy\n",
"2. Handle complex document layouts with ease\n",
"3. Transform unstructured data into structured formats effortlessly\n",
"4. Save hours of manual data entry and verification\n",
"5. Reduce errors in document processing workflows\n",
"\n",
"In this demonstration, we showcase how parsing instructions can be used to extract specific information from unstructured documents. Below are the documents we use for testing:\n",
"\n",
"1. McDonald's Receipt - Extracting the price of each order and the final amount to be paid.\n",
"\n",
"2. Expense Report Document - Extracting employee name, employee ID, position, department, date ranges, individual expense items with dates, categories, and amounts.\n",
"\n",
"3. Purchase Order Document - Identifying the PO number, vendor details, shipping terms, and an itemized list of products with quantities and unit prices.\n",
"\n",
"Let's jump into these real-world examples and see how parsing instructions can help us extract specific information."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Setup API Key"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### McDonald's Receipt\n",
"\n",
"Here we extract the price of each order and the final amount to be paid."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"mcdonalds_receipt.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 66643b81-e2f4-408b-890b-8e116472210b\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\"./mcdonalds_receipt.png\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Rate us HIGHLY SATISFIED\n",
"\n",
"Purchase any sandwich and receive a FREE ITEM\n",
"\n",
"Go to WWW.mcdvoice.com within 7 days of purchase of equal or lesser value and tell us about your visit.\n",
"\n",
"Validation Code: 31278-01121-21018-20481-00081-0\n",
"\n",
"Valid at participating US McDonald's\n",
"\n",
"Expires 30 days after receipt date\n",
"\n",
"# McDonald's Restaurant #312782378\n",
"\n",
"PINE RD NW\n",
"\n",
"RICE MN 56367-9740\n",
"\n",
"TEL# 320 393 4600\n",
"\n",
"KS# 12/08/2022 08:48 PM\n",
"\n",
"# Order\n",
"\n",
"|Happy Meal 6 Pc|$4.89|\n",
"|---|---|\n",
"|Creamy Ranch Cup| |\n",
"|Extra Kids Fry| |\n",
"|Wreck It Ralph 2 Snack| |\n",
"|Oreo McFlurry|$2.69|\n",
"\n",
"# Summary\n",
"\n",
"|Subtotal|$7.58|\n",
"|---|---|\n",
"|Tax|$0.52|\n",
"|Take-Out Total|$8.10|\n",
"|Cash Tendered|$10.00|\n",
"|Change|$1.90|\n",
"\n",
"### Not ACCEPTING APPLICATIONS *++ McDonald's Restaurant Rice\n",
"\n",
"Text to #36453 apply 31278\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 1a04fdbb-5415-4a36-a1bd-26bfb5d618fa\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"The provided document is a McDonald's receipt.\n",
" Provide the price of each order and final amount to be paid.\"\"\"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./mcdonalds_receipt.png\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Here are the prices for each order from the McDonald's receipt:\n",
"\n",
"1. Happy Meal 6 Pc: $4.89\n",
"2. Snack Oreo McFlurry: $2.69\n",
"\n",
"**Subtotal:** $7.58\n",
"**Tax:** $0.52\n",
"**Total Amount to be Paid:** $8.10\n",
"\n",
"The cash tendered was $10.00, and the change given was $1.90.\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Expense Report Document\n",
"\n",
"Here we extract employee name, employee ID, position, department, date ranges, individual expense items with dates, categories, and amounts."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"expense_report_document.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b6bcc6e1-7d30-4522-9abd-ace196781a70\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\n",
" \"./expense_report_document.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# QUANTUM DYNAMICS CORPORATION\n",
"\n",
"# EMPLOYEE EXPENSE REPORT\n",
"\n",
"# FISCAL YEAR 2024\n",
"\n",
"# EMPLOYEE INFORMATION:\n",
"\n",
"Name: Dr. Alexandra Chen-Martinez, PhD\n",
"\n",
"Employee ID: QD-2022-1457\n",
"\n",
"Department: Advanced Research & Development\n",
"\n",
"Cost Center: CC-ARD-NA-003\n",
"\n",
"Project Codes: QD-QUANTUM-2024-01, QD-AI-2024-03\n",
"\n",
"Position: Principal Research Scientist\n",
"\n",
"Reporting Manager: Dr. James Thompson\n",
"\n",
"# TRIP/EXPENSE PERIOD:\n",
"\n",
"Start Date: November 15, 2024\n",
"\n",
"End Date: December 10, 2024\n",
"\n",
"Purpose: International Conference Attendance & Client Meetings\n",
"\n",
"Locations: Tokyo, Japan → Singapore → Sydney, Australia\n",
"\n",
"# CURRENCY CONVERSION RATES APPLIED:\n",
"\n",
"JPY (¥) → USD: 0.0068 (as of 11/15/2024)\n",
"\n",
"SGD (S$) → USD: 0.74 (as of 11/28/2024)\n",
"\n",
"AUD (A$) → USD: 0.65 (as of 12/03/2024)\n",
"\n",
"# ITEMIZED EXPENSES:\n",
"\n",
"|Date|Category|Description|Original|Currency|USD|\n",
"|---|---|---|---|---|---|\n",
"|11/15/2024|Transportation|JFK → NRT Business Class|4,250.00|USD|4,250.00|\n",
"|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|\n",
"|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|\n",
"|11/16/2024|Accommodation|Hilton Tokyo - 5 nights|225,000|JPY|1,530.00|\n",
"|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 7b0d05bb-947b-4475-8d0f-f10386f7446e\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"You are provided with an expense report. \n",
"Extract employee name, employee id, position, department, date ranges, individual expense items with dates, categories, and amounts.\"\"\"\n",
"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./expense_report_document.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"**Employee Information:**\n",
"- **Name:** Dr. Alexandra Chen-Martinez, PhD\n",
"- **Employee ID:** QD-2022-1457\n",
"- **Position:** Principal Research Scientist\n",
"- **Department:** Advanced Research & Development\n",
"\n",
"**Trip/Expense Period:**\n",
"- **Start Date:** November 15, 2024\n",
"- **End Date:** December 10, 2024\n",
"\n",
"**Expense Items:**\n",
"1. **Date:** 11/15/2024\n",
"- **Category:** Transportation\n",
"- **Description:** JFK → NRT Business Class\n",
"- **Original Amount:** $4,250.00\n",
"- **Currency:** USD\n",
"- **USD Amount:** $4,250.00\n",
"- **Booking Reference:** QF78956 - Corporate Rate Applied\n",
"- **Project Code:** QD-QUANTUM-2024-01\n",
"\n",
"2. **Date:** 11/16/2024\n",
"- **Category:** Accommodation\n",
"- **Description:** Hilton Tokyo - 5 nights\n",
"- **Original Amount:** ¥225,000\n",
"- **Currency:** JPY\n",
"- **USD Amount:** $1,530.00\n",
"- **Confirmation:** HTK-2024-78956\n",
"\n",
"**Locations:**\n",
"- Tokyo, Japan\n",
"- Singapore\n",
"- Sydney, Australia\n",
"\n",
"**Currency Conversion Rates Applied:**\n",
"- JPY (¥) → USD: 0.0068 (as of 11/15/2024)\n",
"- SGD (S$) → USD: 0.74 (as of 11/28/2024)\n",
"- AUD (A$) → USD: 0.65 (as of 12/03/2024)\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Purchase Order Document \n",
"\n",
"Here we identify the PO number, vendor details, shipping terms, and an itemized list of products with quantities and unit prices."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"purchase_order_document.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b8cb11c3-7dce-4e6a-94bb-1a4e50e45e55\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\n",
" \"./purchase_order_document.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# GLOBAL TECH SOLUTIONS, INC.\n",
"\n",
"# PURCHASE ORDER\n",
"\n",
"Document Reference: PO-2024-GT-9876/REV.2\n",
"\n",
"[Original: PO-2024-GT-9876]\n",
"\n",
"Amendment Date: 12/10/2024\n",
"\n",
"# VENDOR INFORMATION:\n",
"\n",
"Quantum Electronics Manufacturing\n",
"\n",
"DUNS: 78-456-7890\n",
"\n",
"Tax ID: EU8976543210\n",
"\n",
"Hoofdorp, Netherlands\n",
"\n",
"Vendor #: QEM-EU-2024-001\n",
"\n",
"# SHIP TO:\n",
"\n",
"Global Tech Solutions, Inc.\n",
"\n",
"Building 7A, Innovation Park\n",
"\n",
"2100 Technology Drive\n",
"\n",
"Austin, TX 78701\n",
"\n",
"USA\n",
"\n",
"Attn: Sarah Martinez, Receiving Manager\n",
"\n",
"Tel: +1 (512) 555-0123\n",
"\n",
"# PAYMENT TERMS:\n",
"\n",
"Net 45\n",
"\n",
"2% discount if paid within 15 days\n",
"\n",
"# SHIPPING TERMS:\n",
"\n",
"DDP (Delivered Duty Paid) - Incoterms 2020\n",
"\n",
"Insurance Required: Yes\n",
"\n",
"Preferred Carrier: DHL/FedEx\n",
"\n",
"Required Delivery Date: 01/15/2025\n",
"\n",
"# SPECIAL INSTRUCTIONS:\n",
"\n",
"1. All shipments must include Certificate of Conformance\n",
"2. ESD-sensitive items must be properly packaged\n",
"3. Temperature logging required for items marked with *\n",
"4. Partial shipments accepted with prior approval\n",
"5. Quote PO number on all correspondence\n",
"\n",
"# ITEM DETAILS:\n",
"\n",
"|Line|Part Number|Description|Qty|UOM|Unit Price|Total|\n",
"|---|---|---|---|---|---|---|\n",
"|1|QE-MCU-5590|Microcontroller Unit|500|EA|$12.50|$6,250.00|\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id d2731305-984d-4633-8a52-0493748cf10b\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"You are provided with a purchase order. \n",
"Identify the PO number, vendor details, shipping terms, and itemized list of products with quantities and unit prices.\"\"\"\n",
"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./purchase_order_document.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Here are the details extracted from the purchase order:\n",
"\n",
"**PO Number:** PO-2024-GT-9876/REV.2\n",
"\n",
"**Vendor Details:**\n",
"- **Vendor Name:** Quantum Electronics Manufacturing\n",
"- **DUNS:** 78-456-7890\n",
"- **Tax ID:** EU8976543210\n",
"- **Address:** Hoofdorp, Netherlands\n",
"- **Vendor Number:** QEM-EU-2024-001\n",
"- **Contact Person:** Sarah Martinez, Receiving Manager\n",
"- **Phone:** +1 (512) 555-0123\n",
"\n",
"**Shipping Terms:**\n",
"- **Terms:** DDP (Delivered Duty Paid) - Incoterms 2020\n",
"- **Insurance Required:** Yes\n",
"- **Preferred Carrier:** DHL/FedEx\n",
"- **Required Delivery Date:** 01/15/2025\n",
"\n",
"**Itemized List of Products:**\n",
"1. **Part Number:** QE-MCU-5590\n",
"- **Description:** Microcontroller Unit\n",
"- **Quantity:** 500 EA\n",
"- **Unit Price:** $12.50\n",
"- **Total:** $6,250.00\n",
"\n",
"**Payment Terms:**\n",
"- Net 45\n",
"- 2% discount if paid within 15 days\n",
"\n",
"**Special Instructions:**\n",
"1. All shipments must include Certificate of Conformance\n",
"2. ESD-sensitive items must be properly packaged\n",
"3. Temperature logging required for items marked with *\n",
"4. Partial shipments accepted with prior approval\n",
"5. Quote PO number on all correspondence\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 344 KiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 100 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 464 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 410 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 444 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 610 KiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 986 KiB

@@ -46,7 +46,7 @@
"metadata": {},
"outputs": [],
"source": [
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<LLAMA_CLOUD_API_KEY>"
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<LLAMA_CLOUD_API_KEY>\""
]
},
{
+572 -144
View File
@@ -1,23 +1,26 @@
import os
import asyncio
from urllib.parse import urlparse
import httpx
import mimetypes
import time
from pathlib import Path
from typing import AsyncGenerator, List, Optional, Union
from pathlib import Path, PurePath, PurePosixPath
from typing import AsyncGenerator, Any, Dict, List, Optional, Union
from contextlib import asynccontextmanager
from io import BufferedIOBase
from llama_index.core.async_utils import run_jobs
from fsspec import AbstractFileSystem
from llama_index.core.async_utils import asyncio_run, run_jobs
from llama_index.core.bridge.pydantic import Field, field_validator
from llama_index.core.constants import DEFAULT_BASE_URL
from llama_index.core.readers.base import BasePydanticReader
from llama_index.core.readers.file.base import get_default_fs
from llama_index.core.schema import Document
from llama_parse.utils import (
nest_asyncio_err,
nest_asyncio_msg,
ResultType,
Language,
SUPPORTED_FILE_TYPES,
)
from copy import deepcopy
@@ -32,6 +35,7 @@ _DEFAULT_SEPARATOR = "\n---\n"
class LlamaParse(BasePydanticReader):
"""A smart-parser for files."""
# Library / access specific configurations
api_key: str = Field(
default="",
description="The API key for the LlamaParse API.",
@@ -41,8 +45,20 @@ class LlamaParse(BasePydanticReader):
default=DEFAULT_BASE_URL,
description="The base URL of the Llama Parsing API.",
)
result_type: ResultType = Field(
default=ResultType.TXT, description="The result type for the parser."
check_interval: int = Field(
default=1,
description="The interval in seconds to check if the parsing is done.",
)
custom_client: Optional[httpx.AsyncClient] = Field(
default=None, description="A custom HTTPX client to use for sending requests."
)
ignore_errors: bool = Field(
default=True,
description="Whether or not to ignore and skip errors raised during parsing.",
)
max_timeout: int = Field(
default=2000,
description="The maximum timeout in seconds to wait for the parsing to finish.",
)
num_workers: int = Field(
default=4,
@@ -50,59 +66,214 @@ class LlamaParse(BasePydanticReader):
lt=10,
description="The number of workers to use sending API requests for parsing.",
)
check_interval: int = Field(
default=1,
description="The interval in seconds to check if the parsing is done.",
)
max_timeout: int = Field(
default=2000,
description="The maximum timeout in seconds to wait for the parsing to finish.",
)
verbose: bool = Field(
default=True, description="Whether to print the progress of the parsing."
result_type: ResultType = Field(
default=ResultType.TXT, description="The result type for the parser."
)
show_progress: bool = Field(
default=True, description="Show progress when parsing multiple files."
)
language: Language = Field(
default=Language.ENGLISH, description="The language of the text to parse."
split_by_page: bool = Field(
default=True,
description="Whether to split by page using the page separator",
)
parsing_instruction: Optional[str] = Field(
default="", description="The parsing instruction for the parser."
verbose: bool = Field(
default=True, description="Whether to print the progress of the parsing."
)
skip_diagonal_text: Optional[bool] = Field(
# Parsing specific configurations (Alphabetical order)
annotate_links: Optional[bool] = Field(
default=False,
description="If set to true, the parser will ignore diagonal text (when the text rotation in degrees modulo 90 is not 0).",
description="Annotate links found in the document to extract their URL.",
)
invalidate_cache: Optional[bool] = Field(
auto_mode: Optional[bool] = Field(
default=False,
description="If set to true, the cache will be ignored and the document re-processes. All document are kept in cache for 48hours after the job was completed to avoid processing the same document twice.",
description="If set to true, the parser will automatically select the best mode to extract text from documents based on the rules provide. Will use the 'accurate' default mode by default and will upgrade page that match the rule to Premium mode.",
)
auto_mode_trigger_on_image_in_page: Optional[bool] = Field(
default=False,
description="If auto_mode is set to true, the parser will upgrade the page that contain an image to Premium mode.",
)
auto_mode_trigger_on_table_in_page: Optional[bool] = Field(
default=False,
description="If auto_mode is set to true, the parser will upgrade the page that contain a table to Premium mode.",
)
auto_mode_trigger_on_text_in_page: Optional[str] = Field(
default=None,
description="If auto_mode is set to true, the parser will upgrade the page that contain the text to Premium mode.",
)
auto_mode_trigger_on_regexp_in_page: Optional[str] = Field(
default=None,
description="If auto_mode is set to true, the parser will upgrade the page that match the regexp to Premium mode.",
)
azure_openai_api_version: Optional[str] = Field(
default=None, description="Azure Openai API Version"
)
azure_openai_deployment_name: Optional[str] = Field(
default=None, description="Azure Openai Deployment Name"
)
azure_openai_endpoint: Optional[str] = Field(
default=None, description="Azure Openai Endpoint"
)
azure_openai_key: Optional[str] = Field(
default=None, description="Azure Openai Key"
)
bbox_bottom: Optional[float] = Field(
default=None,
description="The bottom margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page height.",
)
bbox_left: Optional[float] = Field(
default=None,
description="The left margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page width.",
)
bbox_right: Optional[float] = Field(
default=None,
description="The right margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page width.",
)
bbox_top: Optional[float] = Field(
default=None,
description="The top margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page height.",
)
continuous_mode: Optional[bool] = Field(
default=False,
description="Parse documents continuously, leading to better results on documents where tables span across two pages.",
)
disable_ocr: Optional[bool] = Field(
default=False,
description="Disable the OCR on the document. LlamaParse will only extract the copyable text from the document.",
)
disable_image_extraction: Optional[bool] = Field(
default=False,
description="If set to true, the parser will not extract images from the document. Make the parser faster.",
)
do_not_cache: Optional[bool] = Field(
default=False,
description="If set to true, the document will not be cached. This mean that you will be re-charged it you reprocess them as they will not be cached.",
)
fast_mode: Optional[bool] = Field(
default=False,
description="Note: Non compatible with gpt-4o. If set to true, the parser will use a faster mode to extract text from documents. This mode will skip OCR of images, and table/heading reconstruction.",
)
do_not_unroll_columns: Optional[bool] = Field(
default=False,
description="If set to true, the parser will keep column in the text according to document layout. Reduce reconstruction accuracy, and LLM's/embedings performances in most case.",
)
page_separator: Optional[str] = Field(
extract_charts: Optional[bool] = Field(
default=False,
description="If set to true, the parser will extract/tag charts from the document.",
)
extract_layout: Optional[bool] = Field(
default=False,
description="If set to true, the parser will extract the layout information of the document. Cost 1 credit per page.",
)
fast_mode: Optional[bool] = Field(
default=False,
description="Note: Non compatible with gpt-4o. If set to true, the parser will use a faster mode to extract text from documents. This mode will skip OCR of images, and table/heading reconstruction.",
)
guess_xlsx_sheet_names: Optional[bool] = Field(
default=False,
description="Whether to guess the sheet names of the xlsx file.",
)
html_make_all_elements_visible: Optional[bool] = Field(
default=False,
description="If set to true, when parsing HTML the parser will consider all elements display not element as display block.",
)
html_remove_fixed_elements: Optional[bool] = Field(
default=False,
description="If set to true, when parsing HTML the parser will remove fixed elements. Useful to hide cookie banners.",
)
html_remove_navigation_elements: Optional[bool] = Field(
default=False,
description="If set to true, when parsing HTML the parser will remove navigation elements. Useful to hide menus, header, footer.",
)
http_proxy: Optional[str] = Field(
default=None,
description="A templated page separator to use to split the text. If it contain `{page_number}`,it will be replaced by the next page number. If not set will the default separator '\\n---\\n' will be used.",
description="(optional) If set with input_url will use the specified http proxy to download the file.",
)
invalidate_cache: Optional[bool] = Field(
default=False,
description="If set to true, the cache will be ignored and the document re-processes. All document are kept in cache for 48hours after the job was completed to avoid processing the same document twice.",
)
is_formatting_instruction: Optional[bool] = Field(
default=False,
description="Allow the parsing instruction to also format the output. Disable to have a cleaner markdown output.",
)
language: Optional[str] = Field(
default="en", description="The language of the text to parse."
)
max_pages: Optional[int] = Field(
default=None,
description="The maximum number of pages to extract text from documents. If set to 0 or not set, all pages will be that should be extracted will be extracted (can work in combination with targetPages).",
)
output_pdf_of_document: Optional[bool] = Field(
default=False,
description="If set to true, the parser will also output a PDF of the document. (except for spreadsheets)",
)
output_s3_path_prefix: Optional[str] = Field(
default=None,
description="An S3 path prefix to store the output of the parsing job. If set, the parser will upload the output to S3. The bucket need to be accessible from the LlamaIndex organization.",
)
page_prefix: Optional[str] = Field(
default=None,
description="A templated prefix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
)
page_separator: Optional[str] = Field(
default=None,
description="A templated page separator to use to split the text. If it contain `{page_number}`,it will be replaced by the next page number. If not set will the default separator '\\n---\\n' will be used.",
)
page_suffix: Optional[str] = Field(
default=None,
description="A templated suffix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
)
gpt4o_mode: bool = Field(
parsing_instruction: Optional[str] = Field(
default="", description="The parsing instruction for the parser."
)
premium_mode: Optional[bool] = Field(
default=False,
description="Use our best parser mode if set to True.",
)
skip_diagonal_text: Optional[bool] = Field(
default=False,
description="If set to true, the parser will ignore diagonal text (when the text rotation in degrees modulo 90 is not 0).",
)
structured_output: Optional[bool] = Field(
default=False,
description="If set to true, the parser will output structured data based on the provided JSON Schema.",
)
structured_output_json_schema: Optional[str] = Field(
default=None,
description="A JSON Schema to use to structure the output of the parsing job. If set, the parser will output structured data based on the provided JSON Schema.",
)
structured_output_json_schema_name: Optional[str] = Field(
default=None,
description="The named JSON Schema to use to structure the output of the parsing job. For convenience / testing, LlamaParse provides a few named JSON Schema that can be used directly. Use 'imFeelingLucky' to let llamaParse dream the schema.",
)
take_screenshot: Optional[bool] = Field(
default=False,
description="Whether to take screenshot of each page of the document.",
)
target_pages: Optional[str] = Field(
default=None,
description="The target pages to extract text from documents. Describe as a comma separated list of page numbers. The first page of the document is page 0",
)
use_vendor_multimodal_model: Optional[bool] = Field(
default=False,
description="Whether to use the vendor multimodal API.",
)
vendor_multimodal_api_key: Optional[str] = Field(
default=None,
description="The API key for the multimodal API.",
)
vendor_multimodal_model_name: Optional[str] = Field(
default=None,
description="The model name for the vendor multimodal API.",
)
webhook_url: Optional[str] = Field(
default=None,
description="A URL that needs to be called at the end of the parsing job.",
)
# Deprecated
bounding_box: Optional[str] = Field(
default=None,
description="The bounding box to use to extract text from documents describe as a string containing the bounding box margins",
)
gpt4o_mode: Optional[bool] = Field(
default=False,
description="Whether to use gpt-4o extract text from documents.",
)
@@ -110,41 +281,6 @@ class LlamaParse(BasePydanticReader):
default=None,
description="The API key for the GPT-4o API. Lowers the cost of parsing.",
)
bounding_box: Optional[str] = Field(
default=None,
description="The bounding box to use to extract text from documents describe as a string containing the bounding box margins",
)
target_pages: Optional[str] = Field(
default=None,
description="The target pages to extract text from documents. Describe as a comma separated list of page numbers. The first page of the document is page 0",
)
ignore_errors: bool = Field(
default=True,
description="Whether or not to ignore and skip errors raised during parsing.",
)
split_by_page: bool = Field(
default=True,
description="Whether to split by page using the page separator",
)
vendor_multimodal_api_key: Optional[str] = Field(
default=None,
description="The API key for the multimodal API.",
)
use_vendor_multimodal_model: bool = Field(
default=False,
description="Whether to use the vendor multimodal API.",
)
vendor_multimodal_model_name: Optional[str] = Field(
default=None,
description="The model name for the vendor multimodal API.",
)
take_screenshot: bool = Field(
default=False,
description="Whether to take screenshot of each page of the document.",
)
custom_client: Optional[httpx.AsyncClient] = Field(
default=None, description="A custom HTTPX client to use for sending requests."
)
@field_validator("api_key", mode="before", check_fields=True)
@classmethod
@@ -176,14 +312,51 @@ class LlamaParse(BasePydanticReader):
async with httpx.AsyncClient(timeout=self.max_timeout) as client:
yield client
def _is_input_url(self, file_path: FileInput) -> bool:
"""Check if the input is a valid URL.
This method checks for:
- Proper URL scheme (http/https)
- Valid URL structure
- Network location (domain)
"""
if not isinstance(file_path, str):
return False
try:
result = urlparse(file_path)
return all(
[
result.scheme in ("http", "https"),
result.netloc, # Has domain
result.scheme, # Has scheme
]
)
except Exception:
return False
def _is_s3_url(self, file_path: FileInput) -> bool:
"""Check if the input is a valid URL.
This method checks for:
- Proper S3 scheme (s3://)
"""
if isinstance(file_path, str):
return file_path.startswith("s3://")
return False
# upload a document and get back a job_id
async def _create_job(
self, file_input: FileInput, extra_info: Optional[dict] = None
self,
file_input: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> str:
headers = {"Authorization": f"Bearer {self.api_key}"}
url = f"{self.base_url}/api/parsing/upload"
files = None
file_handle = None
input_url = file_input if self._is_input_url(file_input) else None
input_s3_path = file_input if self._is_s3_url(file_input) else None
if isinstance(file_input, (bytes, BufferedIOBase)):
if not extra_info or "file_name" not in extra_info:
@@ -193,7 +366,11 @@ class LlamaParse(BasePydanticReader):
file_name = extra_info["file_name"]
mime_type = mimetypes.guess_type(file_name)[0]
files = {"file": (file_name, file_input, mime_type)}
elif isinstance(file_input, (str, Path)):
elif input_url is not None:
files = None
elif input_s3_path is not None:
files = None
elif isinstance(file_input, (str, Path, PurePosixPath, PurePath)):
file_path = str(file_input)
file_ext = os.path.splitext(file_path)[1].lower()
if file_ext not in SUPPORTED_FILE_TYPES:
@@ -203,46 +380,195 @@ class LlamaParse(BasePydanticReader):
)
mime_type = mimetypes.guess_type(file_path)[0]
# Open the file here for the duration of the async context
file_handle = open(file_path, "rb")
# load data, set the mime type
fs = fs or get_default_fs()
file_handle = fs.open(file_input, "rb")
files = {"file": (os.path.basename(file_path), file_handle, mime_type)}
else:
raise ValueError(
"file_input must be either a file path string, file bytes, or buffer object"
)
data = {
"language": self.language.value,
"parsing_instruction": self.parsing_instruction,
"invalidate_cache": self.invalidate_cache,
"skip_diagonal_text": self.skip_diagonal_text,
"do_not_cache": self.do_not_cache,
"fast_mode": self.fast_mode,
"do_not_unroll_columns": self.do_not_unroll_columns,
"gpt4o_mode": self.gpt4o_mode,
"gpt4o_api_key": self.gpt4o_api_key,
"vendor_multimodal_api_key": self.vendor_multimodal_api_key,
"use_vendor_multimodal_model": self.use_vendor_multimodal_model,
"vendor_multimodal_model_name": self.vendor_multimodal_model_name,
"take_screenshot": self.take_screenshot,
}
data: Dict[str, Any] = {}
data["from_python_package"] = True
if self.annotate_links:
data["annotate_links"] = self.annotate_links
if self.auto_mode:
data["auto_mode"] = self.auto_mode
if self.auto_mode_trigger_on_image_in_page:
data[
"auto_mode_trigger_on_image_in_page"
] = self.auto_mode_trigger_on_image_in_page
if self.auto_mode_trigger_on_table_in_page:
data[
"auto_mode_trigger_on_table_in_page"
] = self.auto_mode_trigger_on_table_in_page
if self.auto_mode_trigger_on_text_in_page is not None:
data[
"auto_mode_trigger_on_text_in_page"
] = self.auto_mode_trigger_on_text_in_page
if self.auto_mode_trigger_on_regexp_in_page is not None:
data[
"auto_mode_trigger_on_regexp_in_page"
] = self.auto_mode_trigger_on_regexp_in_page
if self.azure_openai_api_version is not None:
data["azure_openai_api_version"] = self.azure_openai_api_version
if self.azure_openai_deployment_name is not None:
data["azure_openai_deployment_name"] = self.azure_openai_deployment_name
if self.azure_openai_endpoint is not None:
data["azure_openai_endpoint"] = self.azure_openai_endpoint
if self.azure_openai_key is not None:
data["azure_openai_key"] = self.azure_openai_key
if self.bbox_bottom is not None:
data["bbox_bottom"] = self.bbox_bottom
if self.bbox_left is not None:
data["bbox_left"] = self.bbox_left
if self.bbox_right is not None:
data["bbox_right"] = self.bbox_right
if self.bbox_top is not None:
data["bbox_top"] = self.bbox_top
if self.continuous_mode:
data["continuous_mode"] = self.continuous_mode
if self.disable_ocr:
data["disable_ocr"] = self.disable_ocr
if self.disable_image_extraction:
data["disable_image_extraction"] = self.disable_image_extraction
if self.do_not_cache:
data["do_not_cache"] = self.do_not_cache
if self.do_not_unroll_columns:
data["do_not_unroll_columns"] = self.do_not_unroll_columns
if self.extract_charts:
data["extract_charts"] = self.extract_charts
if self.extract_layout:
data["extract_layout"] = self.extract_layout
if self.fast_mode:
data["fast_mode"] = self.fast_mode
if self.guess_xlsx_sheet_names:
data["guess_xlsx_sheet_names"] = self.guess_xlsx_sheet_names
if self.html_make_all_elements_visible:
data["html_make_all_elements_visible"] = self.html_make_all_elements_visible
if self.html_remove_fixed_elements:
data["html_remove_fixed_elements"] = self.html_remove_fixed_elements
if self.html_remove_navigation_elements:
data[
"html_remove_navigation_elements"
] = self.html_remove_navigation_elements
if self.http_proxy is not None:
data["http_proxy"] = self.http_proxy
if input_url is not None:
files = None
data["input_url"] = str(input_url)
if input_s3_path is not None:
files = None
data["input_s3_path"] = str(input_s3_path)
if self.invalidate_cache:
data["invalidate_cache"] = self.invalidate_cache
if self.is_formatting_instruction:
data["is_formatting_instruction"] = self.is_formatting_instruction
if self.language:
data["language"] = self.language
if self.max_pages is not None:
data["max_pages"] = self.max_pages
if self.output_pdf_of_document:
data["output_pdf_of_document"] = self.output_pdf_of_document
if self.output_s3_path_prefix is not None:
data["output_s3_path_prefix"] = self.output_s3_path_prefix
if self.page_prefix is not None:
data["page_prefix"] = self.page_prefix
# only send page separator to server if it is not None
# as if a null, "" string is sent the server will then ignore the page separator instead of using the default
if self.page_separator is not None:
data["page_separator"] = self.page_separator
if self.page_prefix is not None:
data["page_prefix"] = self.page_prefix
if self.page_suffix is not None:
data["page_suffix"] = self.page_suffix
if self.bounding_box is not None:
data["bounding_box"] = self.bounding_box
if self.parsing_instruction is not None:
data["parsing_instruction"] = self.parsing_instruction
if self.premium_mode:
data["premium_mode"] = self.premium_mode
if self.skip_diagonal_text:
data["skip_diagonal_text"] = self.skip_diagonal_text
if self.structured_output:
data["structured_output"] = self.structured_output
if self.structured_output_json_schema is not None:
data["structured_output_json_schema"] = self.structured_output_json_schema
if self.structured_output_json_schema_name is not None:
data[
"structured_output_json_schema_name"
] = self.structured_output_json_schema_name
if self.take_screenshot:
data["take_screenshot"] = self.take_screenshot
if self.target_pages is not None:
data["target_pages"] = self.target_pages
if self.use_vendor_multimodal_model:
data["use_vendor_multimodal_model"] = self.use_vendor_multimodal_model
if self.vendor_multimodal_api_key is not None:
data["vendor_multimodal_api_key"] = self.vendor_multimodal_api_key
if self.vendor_multimodal_model_name is not None:
data["vendor_multimodal_model_name"] = self.vendor_multimodal_model_name
if self.webhook_url is not None:
data["webhook_url"] = self.webhook_url
# Deprecated
if self.bounding_box is not None:
data["bounding_box"] = self.bounding_box
if self.gpt4o_mode:
data["gpt4o_mode"] = self.gpt4o_mode
if self.gpt4o_api_key is not None:
data["gpt4o_api_key"] = self.gpt4o_api_key
try:
async with self.client_context() as client:
response = await client.post(
@@ -261,7 +587,7 @@ class LlamaParse(BasePydanticReader):
async def _get_job_result(
self, job_id: str, result_type: str, verbose: bool = False
) -> dict:
) -> Dict[str, Any]:
result_url = f"{self.base_url}/api/parsing/job/{job_id}/result/{result_type}"
status_url = f"{self.base_url}/api/parsing/job/{job_id}"
headers = {"Authorization": f"Bearer {self.api_key}"}
@@ -287,7 +613,8 @@ class LlamaParse(BasePydanticReader):
continue
# Allowed values "PENDING", "SUCCESS", "ERROR", "CANCELED"
status = result.json()["status"]
result_json = result.json()
status = result_json["status"]
if status == "SUCCESS":
parsed_result = await client.get(result_url, headers=headers)
return parsed_result.json()
@@ -299,22 +626,25 @@ class LlamaParse(BasePydanticReader):
print(".", end="", flush=True)
await asyncio.sleep(self.check_interval)
continue
else:
raise Exception(
f"Failed to parse the file: {job_id}, status: {status}"
error_code = result_json.get("error_code", "No error code found")
error_message = result_json.get(
"error_message", "No error message found"
)
exception_str = f"Job ID: {job_id} failed with status: {status}, Error code: {error_code}, Error message: {error_message}"
raise Exception(exception_str)
async def _aload_data(
self,
file_path: FileInput,
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
verbose: bool = False,
) -> List[Document]:
"""Load data from the input path."""
try:
job_id = await self._create_job(file_path, extra_info=extra_info)
job_id = await self._create_job(file_path, extra_info=extra_info, fs=fs)
if verbose:
print("Started parsing the file under job_id %s" % job_id)
@@ -345,17 +675,19 @@ class LlamaParse(BasePydanticReader):
self,
file_path: Union[List[FileInput], FileInput],
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> List[Document]:
"""Load data from the input path."""
if isinstance(file_path, (str, Path, bytes, BufferedIOBase)):
if isinstance(file_path, (str, PurePosixPath, Path, bytes, BufferedIOBase)):
return await self._aload_data(
file_path, extra_info=extra_info, verbose=self.verbose
file_path, extra_info=extra_info, fs=fs, verbose=self.verbose
)
elif isinstance(file_path, list):
jobs = [
self._aload_data(
f,
extra_info=extra_info,
fs=fs,
verbose=self.verbose and not self.show_progress,
)
for f in file_path
@@ -384,10 +716,11 @@ class LlamaParse(BasePydanticReader):
self,
file_path: Union[List[FileInput], FileInput],
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> List[Document]:
"""Load data from the input path."""
try:
return asyncio.run(self.aload_data(file_path, extra_info))
return asyncio_run(self.aload_data(file_path, extra_info, fs=fs))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
@@ -402,12 +735,13 @@ class LlamaParse(BasePydanticReader):
job_id = await self._create_job(file_path, extra_info=extra_info)
if self.verbose:
print("Started parsing the file under job_id %s" % job_id)
result = await self._get_job_result(job_id, "json")
result["job_id"] = job_id
result["file_path"] = file_path
return [result]
if not isinstance(file_path, (bytes, BufferedIOBase)):
result["file_path"] = str(file_path)
return [result]
except Exception as e:
file_repr = file_path if isinstance(file_path, str) else "<bytes/buffer>"
print(f"Error while parsing the file '{file_repr}':", e)
@@ -453,59 +787,89 @@ class LlamaParse(BasePydanticReader):
) -> List[dict]:
"""Parse the input path."""
try:
return asyncio.run(self.aget_json(file_path, extra_info))
return asyncio_run(self.aget_json(file_path, extra_info))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
async def aget_assets(
self, json_result: List[dict], download_path: str, asset_key: str
) -> List[dict]:
"""Download assets (images or charts) from the parsed result."""
headers = {"Authorization": f"Bearer {self.api_key}"}
# Make the download path
if not os.path.exists(download_path):
os.makedirs(download_path)
try:
assets = []
for result in json_result:
job_id = result["job_id"]
for page in result["pages"]:
if self.verbose:
print(
f"> {asset_key.capitalize()} for page {page['page']}: {page[asset_key]}"
)
for asset in page[asset_key]:
asset_name = asset["name"]
# Get the full path
asset_path = os.path.join(
download_path, f"{job_id}-{asset_name}"
)
# Get a valid asset path
if not asset_path.endswith(".png"):
if not asset_path.endswith(".jpg"):
asset_path += ".png"
asset["path"] = asset_path
asset["job_id"] = job_id
asset["original_file_path"] = result.get("file_path", None)
asset["page_number"] = page["page"]
with open(asset_path, "wb") as f:
asset_url = f"{self.base_url}/api/parsing/job/{job_id}/result/image/{asset_name}"
async with self.client_context() as client:
res = await client.get(
asset_url, headers=headers, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
assets.append(asset)
return assets
except Exception as e:
print(f"Error while downloading {asset_key} from the parsed result:", e)
if self.ignore_errors:
return []
else:
raise e
async def aget_images(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download images from the parsed result."""
headers = {"Authorization": f"Bearer {self.api_key}"}
# make the download path
if not os.path.exists(download_path):
os.makedirs(download_path)
try:
images = []
for result in json_result:
job_id = result["job_id"]
for page in result["pages"]:
if self.verbose:
print(f"> Image for page {page['page']}: {page['images']}")
for image in page["images"]:
image_name = image["name"]
# get the full path
image_path = os.path.join(
download_path, f"{job_id}-{image_name}"
)
# get a valid image path
if not image_path.endswith(".png"):
if not image_path.endswith(".jpg"):
image_path += ".png"
image["path"] = image_path
image["job_id"] = job_id
image["original_pdf_path"] = result["file_path"]
image["page_number"] = page["page"]
with open(image_path, "wb") as f:
image_url = f"{self.base_url}/api/parsing/job/{job_id}/result/image/{image_name}"
async with self.client_context() as client:
res = await client.get(
image_url, headers=headers, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
images.append(image)
return images
return await self.aget_assets(json_result, download_path, "images")
except Exception as e:
print("Error while downloading images from the parsed result:", e)
print("Error while downloading images:", e)
if self.ignore_errors:
return []
else:
raise e
async def aget_charts(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download charts from the parsed result."""
try:
return await self.aget_assets(json_result, download_path, "charts")
except Exception as e:
print("Error while downloading charts:", e)
if self.ignore_errors:
return []
else:
@@ -514,7 +878,71 @@ class LlamaParse(BasePydanticReader):
def get_images(self, json_result: List[dict], download_path: str) -> List[dict]:
"""Download images from the parsed result."""
try:
return asyncio.run(self.aget_images(json_result, download_path))
return asyncio_run(self.aget_images(json_result, download_path))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
def get_charts(self, json_result: List[dict], download_path: str) -> List[dict]:
"""Download charts from the parsed result."""
try:
return asyncio_run(self.aget_charts(json_result, download_path))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
async def aget_xlsx(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download xlsx from the parsed result."""
headers = {"Authorization": f"Bearer {self.api_key}"}
# make the download path
if not os.path.exists(download_path):
os.makedirs(download_path)
try:
xlsx_list = []
for result in json_result:
job_id = result["job_id"]
if self.verbose:
print("> XLSX")
xlsx_path = os.path.join(download_path, f"{job_id}.xlsx")
xlsx = {}
xlsx["path"] = xlsx_path
xlsx["job_id"] = job_id
xlsx["original_file_path"] = result.get("file_path", None)
with open(xlsx_path, "wb") as f:
xlsx_url = (
f"{self.base_url}/api/parsing/job/{job_id}/result/raw/xlsx"
)
async with self.client_context() as client:
res = await client.get(
xlsx_url, headers=headers, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
xlsx_list.append(xlsx)
return xlsx_list
except Exception as e:
print("Error while downloading xlsx:", e)
if self.ignore_errors:
return []
else:
raise e
def get_xlsx(self, json_result: List[dict], download_path: str) -> List[dict]:
"""Download xlsx from the parsed result."""
try:
return asyncio_run(self.aget_xlsx(json_result, download_path))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
View File
+92
View File
@@ -0,0 +1,92 @@
import click
import json
from enum import Enum
from pathlib import Path
from pydantic.fields import FieldInfo
from typing import Any, Callable, List
from llama_parse.base import LlamaParse
def pydantic_field_to_click_option(name: str, field: FieldInfo) -> click.Option:
"""Convert a Pydantic field to a Click option."""
kwargs = {
"default": field.default if field.default else None,
"help": field.description,
}
if isinstance(kwargs["default"], Enum):
kwargs["default"] = kwargs["default"].value
if field.annotation is bool:
kwargs["is_flag"] = True
if field.default and field.default is True:
name = f"no-{name}"
return click.option(f'--{name.replace("_", "-")}', **kwargs)
def add_options(options: List[click.Option]) -> Callable:
def _add_options(func: Callable) -> Callable:
for option in reversed(options):
func = option(func)
return func
return _add_options
@click.command()
@click.argument("file_paths", nargs=-1, type=click.Path(exists=True, path_type=Path))
@click.option(
"--output-file", type=click.Path(path_type=Path), help="Path to save the output"
)
@click.option("--output-raw-json", is_flag=True, help="Output the raw JSON result")
@add_options(
[
pydantic_field_to_click_option(name, field)
for name, field in LlamaParse.model_fields.items()
if name not in ["custom_client"]
]
)
def parse(**kwargs: Any) -> None:
"""Parse files using LlamaParse and output the results."""
file_paths = kwargs.pop("file_paths")
output_file = kwargs.pop("output_file")
output_raw_json = kwargs.pop("output_raw_json")
# Remove None values to use LlamaParse defaults
kwargs = {k: v for k, v in kwargs.items() if v is not None}
# Remove no- prefix for boolean flags
kwargs = {k.replace("no_", ""): v for k, v in kwargs.items()}
parser = LlamaParse(**kwargs)
if output_raw_json:
results = parser.get_json_result(list(file_paths))
if output_file:
with output_file.open("w") as f:
json.dump(results, f)
click.echo(f"Results saved to {output_file}")
else:
click.echo(results)
else:
results = parser.load_data(list(file_paths))
if output_file:
with output_file.open("w") as f:
for i, doc in enumerate(results):
f.write(f"File: {doc.metadata.get('file_path', 'Unknown')}\n") # type: ignore
f.write(doc.text) # type: ignore
if i < len(results) - 1:
f.write("\n\n---\n\n")
click.echo(f"Results saved to {output_file}")
else:
for i, doc in enumerate(results):
click.echo(f"File: {doc.metadata.get('file_path', 'Unknown')}") # type: ignore
click.echo(doc.text) # type: ignore
if i < len(results) - 1:
click.echo("\n---\n")
if __name__ == "__main__":
parse()
+8
View File
@@ -11,6 +11,7 @@ class ResultType(str, Enum):
TXT = "text"
MD = "markdown"
JSON = "json"
STRUCTURED = "structured"
class Language(str, Enum):
@@ -190,4 +191,11 @@ SUPPORTED_FILE_TYPES = [
".xlr",
".eth",
".tsv",
".mp3",
".mp4",
".mpeg",
".mpga",
".m4a",
".wav",
".webm",
]
Generated
+1472 -1248
View File
File diff suppressed because it is too large Load Diff
+8 -2
View File
@@ -4,7 +4,7 @@ build-backend = "poetry.core.masonry.api"
[tool.poetry]
name = "llama-parse"
version = "0.5.2"
version = "0.5.18"
description = "Parse files into RAG-Optimized formats."
authors = ["Logan Markewich <logan@llamaindex.ai>"]
license = "MIT"
@@ -12,9 +12,15 @@ readme = "README.md"
packages = [{include = "llama_parse"}]
[tool.poetry.dependencies]
python = ">=3.8.1,<4.0"
python = ">=3.9,<4.0"
llama-index-core = ">=0.11.0"
pydantic = "!=2.10"
click = "^8.1.7"
[tool.poetry.group.dev.dependencies]
pytest = "^8.0.0"
pytest-asyncio = "*"
ipykernel = "^6.29.0"
[tool.poetry.scripts]
llama-parse = "llama_parse.cli.main:parse"
Binary file not shown.

After

Width:  |  Height:  |  Size: 347 KiB

+91 -5
View File
@@ -1,6 +1,9 @@
import os
import pytest
import shutil
from fsspec.implementations.local import LocalFileSystem
from httpx import AsyncClient
from llama_parse import LlamaParse
@@ -74,13 +77,29 @@ def test_simple_page_markdown_buffer(markdown_parser: LlamaParse) -> None:
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
def test_simple_page_progress_workers() -> None:
@pytest.mark.asyncio
async def test_simple_page_with_custom_fs() -> None:
parser = LlamaParse(result_type="markdown")
fs = LocalFileSystem()
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
result = await parser.aload_data(filepath, fs=fs)
assert len(result) == 1
@pytest.mark.skipif(
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_simple_page_progress_workers() -> None:
parser = LlamaParse(result_type="markdown", show_progress=True, verbose=True)
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
result = parser.load_data([filepath, filepath])
result = await parser.aload_data([filepath, filepath])
assert len(result) == 2
assert len(result[0].text) > 0
@@ -91,7 +110,7 @@ def test_simple_page_progress_workers() -> None:
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
result = parser.load_data([filepath, filepath])
result = await parser.aload_data([filepath, filepath])
assert len(result) == 2
assert len(result[0].text) > 0
@@ -100,12 +119,79 @@ def test_simple_page_progress_workers() -> None:
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
def test_custom_client() -> None:
@pytest.mark.asyncio
async def test_custom_client() -> None:
custom_client = AsyncClient(verify=False, timeout=10)
parser = LlamaParse(result_type="markdown", custom_client=custom_client)
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
result = parser.load_data(filepath)
result = await parser.aload_data(filepath)
assert len(result) == 1
assert len(result[0].text) > 0
@pytest.mark.skipif(
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_input_url() -> None:
parser = LlamaParse(result_type="markdown")
# links to a resume example
input_url = "https://cdn-blog.novoresume.com/articles/google-docs-resume-templates/basic-google-docs-resume.png"
result = await parser.aload_data(input_url)
assert len(result) == 1
assert "your name" in result[0].text.lower()
@pytest.mark.skipif(
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_input_url_with_website_input() -> None:
parser = LlamaParse(result_type="markdown")
input_url = "https://www.google.com"
result = await parser.aload_data(input_url)
assert len(result) == 1
assert "google" in result[0].text.lower()
@pytest.mark.skipif(
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_mixing_input_types() -> None:
parser = LlamaParse(result_type="markdown")
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
input_url = "https://cdn-blog.novoresume.com/articles/google-docs-resume-templates/basic-google-docs-resume.png"
result = await parser.aload_data([filepath, input_url])
assert len(result) == 2
@pytest.mark.skipif(
os.environ.get("LLAMA_CLOUD_API_KEY", "") == "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_download_images() -> None:
parser = LlamaParse(result_type="markdown", take_screenshot=True)
filepath = os.path.join(
os.path.dirname(__file__), "test_files/attention_is_all_you_need.pdf"
)
json_result = await parser.aget_json([filepath])
assert len(json_result) == 1
assert len(json_result[0]["pages"][0]["images"]) > 0
download_path = os.path.join(os.path.dirname(__file__), "test_files/images")
shutil.rmtree(download_path, ignore_errors=True)
await parser.aget_images(json_result, download_path)
assert len(os.listdir(download_path)) == len(json_result[0]["pages"][0]["images"])