Compare commits

...

14 Commits

Author SHA1 Message Date
Sacha Bron ae38f406fd fix release pipeline no-dev issue (#592) 2025-01-24 15:17:01 +01:00
Pierre-Loic Doulcet 4897d01cb0 add new formatting instruction parameters (#582)
* add new formatting instruction parameters

* bump version

* wip

* s3 region

* update test
2025-01-22 15:56:57 +01:00
Logan bd7b563463 v0.5.19 (#569) 2024-12-27 13:07:01 -06:00
apostoli 530241dd0b Stoli/feat/connection handling (#568) 2024-12-27 11:34:00 -06:00
Pierre-Loic Doulcet 6338641107 Extract layout, audio files (#557) 2024-12-18 16:29:17 +01:00
Bharath Lakshman Kumar 6d62fb89c3 Fix docstring for aget_xlsx method (#551)
Updated docstring to describe xlsx download instead of image download
2024-12-13 20:35:18 +05:30
Ravi Theja 7d4df3b6e5 Add cookbook for parsing instructions (#550) 2024-12-13 06:44:17 -08:00
Ravi Theja bc28db5b92 Update cache parameter (#548) 2024-12-11 16:15:59 +01:00
Jerry Liu f78186c0f7 update auto-mode (#545) 2024-12-09 16:13:09 -06:00
Laurie Voss e3292f5566 Expanding auto mode notebook with strings and regex triggers (#544) 2024-12-09 12:03:09 -08:00
Jerry Liu 58f980f411 auto-mode notebook (#540)
Co-authored-by: Laurie Voss <github@seldo.com>
2024-12-09 08:59:21 -08:00
Ravi Theja 4740d0611d Add get charts function (#542)
* Add get charts function

* code refactoring

* solve linting

* Add cookbook
2024-12-09 21:28:48 +05:30
Laurie Voss 3651a10e80 JSON mode tour notebook (#531) 2024-12-06 14:21:15 -08:00
Pierre-Loic Doulcet 483b51c51c Add support for html_remove_navigation_elements. (#532) 2024-12-06 12:05:46 +01:00
29 changed files with 3526 additions and 548 deletions
+2 -2
View File
@@ -31,10 +31,10 @@ jobs:
shell: bash
run: pip install -e .
- name: Build and publish to pypi
uses: JRubics/poetry-publish@v1.17
uses: JRubics/poetry-publish@v2.1
with:
pypi_token: ${{ secrets.LLAMA_PARSE_PYPI_TOKEN }}
ignore_dev_requirements: "yes"
poetry_install_options: "--without dev"
- name: Create GitHub Release
id: create_release
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
# You can use PyPy versions in python-version.
# For example, pypy-2.7 and pypy-3.8
matrix:
python-version: ["3.8", "3.10", "3.11"]
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v3
with:
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+354 -412
View File
@@ -1,415 +1,357 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "97c79c38-38a3-40f3-ba2e-250649347d63",
"metadata": {
"id": "97c79c38-38a3-40f3-ba2e-250649347d63"
},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_starter_multimodal.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
"cell_type": "markdown",
"id": "4e081457",
"metadata": {},
"source": [
"# Multimodal Parsing using LlamaParse\n",
"\n",
"This cookbook shows you how to use LlamaParse to parse any document with the multimodal capabilities of Multi-Modal LLMs from Anthropic/ OpenAI.\n",
"\n",
"LlamaParse allows you to plug in external, multimodal model vendors for parsing - we handle the error correction, validation, and scalability/reliability for you.\n"
]
},
{
"cell_type": "markdown",
"id": "qOdqBxCS51Ow",
"metadata": {
"id": "qOdqBxCS51Ow"
},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "H_Vqcylb50vm",
"metadata": {
"id": "H_Vqcylb50vm"
},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"id": "15e60ecf-519c-41fc-911b-765adaf8bad4",
"metadata": {
"id": "15e60ecf-519c-41fc-911b-765adaf8bad4"
},
"source": [
"### Setup\n",
"\n",
"Here we setup `LLAMA_CLOUD_API_KEY` for using `LlamaParse`."
]
},
{
"cell_type": "code",
"execution_count": 1,
"id": "91a9e532-1454-40e0-bbf0-fd442c350121",
"metadata": {
"id": "91a9e532-1454-40e0-bbf0-fd442c350121"
},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<YOUR LLAMACLOUD API KEY>\""
]
},
{
"cell_type": "markdown",
"id": "LGwBNPNotZRQ",
"metadata": {
"id": "LGwBNPNotZRQ"
},
"source": [
"## Download Data\n",
"\n",
"For this demonstration, we will use OpenAI's recent paper `Evaluation of OpenAI o1: Opportunities and Challenges of AGI`."
]
},
{
"cell_type": "code",
"execution_count": 2,
"id": "IjtKDQRLrylI",
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "IjtKDQRLrylI",
"outputId": "31df0fac-51f2-4697-f78b-0b7c0b8cd145"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 18:54:24-- https://arxiv.org/pdf/2409.18486\n",
"Resolving arxiv.org (arxiv.org)... 151.101.67.42, 151.101.131.42, 151.101.3.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.67.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 13986265 (13M) [application/pdf]\n",
"Saving to: o1.pdf\n",
"\n",
"o1.pdf 100%[===================>] 13.34M 11.8MB/s in 1.1s \n",
"\n",
"2024-12-05 18:54:26 (11.8 MB/s) - o1.pdf saved [13986265/13986265]\n",
"\n"
]
}
],
"source": [
"!wget \"https://arxiv.org/pdf/2409.18486\" -O \"o1.pdf\""
]
},
{
"cell_type": "markdown",
"id": "4e29a9d7-5bd9-4fb8-8ec1-4c128a748662",
"metadata": {
"id": "4e29a9d7-5bd9-4fb8-8ec1-4c128a748662"
},
"source": [
"## Initialize LlamaParse\n",
"\n",
"Initialize LlamaParse in multimodal mode, and specify the vendor.\n",
"\n",
"**NOTE**: optionally you can specify the Anthropic/ OpenAI API key. If you choose to do so LlamaParse will only charge you 1 credit (0.3c) per page. \n",
"\n",
"\n",
"Using your own API key may incur additional costs from your model provider and could result in failed pages or documents if you do not have sufficient usage limits."
]
},
{
"cell_type": "code",
"execution_count": 3,
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9",
"metadata": {
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9"
},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"md\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
"cell_type": "markdown",
"id": "1b5d6da6",
"metadata": {},
"source": [
"### With anthropic-sonnet-3.5"
]
},
{
"cell_type": "code",
"execution_count": 5,
"id": "f2e9d9cf-8189-4fcb-b34f-cde6cc0b59c8",
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "f2e9d9cf-8189-4fcb-b34f-cde6cc0b59c8",
"outputId": "a337cbdd-60db-4a73-b66b-2bd6159e81f2"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id dd9d5e0f-160e-486a-89a2-6005e5a1c2ac\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model_name=\"anthropic-sonnet-3.5\",\n",
" target_pages=\"24\"\n",
" # invalidate_cache=True\n",
")\n",
"json_objs = parser.get_json_result(\"o1.pdf\")\n",
"json_list = json_objs[0][\"pages\"]\n",
"docs = get_text_nodes(json_list)"
]
},
{
"cell_type": "markdown",
"id": "4f3c51b0-7878-48d7-9bc3-02b516500128",
"metadata": {
"id": "4f3c51b0-7878-48d7-9bc3-02b516500128"
},
"source": [
"### With GPT-4o\n",
"\n",
"For comparison, we will also parse the document using GPT-4o."
]
},
{
"cell_type": "code",
"execution_count": 6,
"id": "6fc3f258-50ae-4988-b904-c105463a498f",
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "6fc3f258-50ae-4988-b904-c105463a498f",
"outputId": "89c525c4-2b93-4909-9657-55646e034637"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 6a4dea44-4f90-406b-b290-9e98620b1232\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model=\"openai-gpt4o\",\n",
" target_pages=\"24\",\n",
" # invalidate_cache=True\n",
")\n",
"json_objs_gpt4o = parser_gpt4o.get_json_result(\"o1.pdf\")\n",
"json_list_gpt4o = json_objs_gpt4o[0][\"pages\"]\n",
"docs_gpt4o = get_text_nodes(json_list_gpt4o)"
]
},
{
"cell_type": "markdown",
"id": "44c20f7a-2901-4dd0-b635-a4b33c5664c1",
"metadata": {
"id": "44c20f7a-2901-4dd0-b635-a4b33c5664c1"
},
"source": [
"### View Results\n",
"\n",
"Let's visualize the results along with the original document page."
]
},
{
"cell_type": "code",
"execution_count": 7,
"id": "778698aa-da7e-4081-b3b5-0372f228536f",
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "778698aa-da7e-4081-b3b5-0372f228536f",
"outputId": "bb89e323-7041-4fc3-d835-95e373189d02"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|-----------------|----------------------------------|\n",
"| Attribute | Value | Basic Personal Information: Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia. |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White | Biomarker Measurements: The subject's genetic profile includes an ApoE4 status of 0.0... |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... | Cognitive and Neurofunctional Assessments: The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 | Volumetric Data: Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 54422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 2782.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0.... |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-preview's performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the model's logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-preview's solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-preview's solutions. By comparing o1-preview's solutions with the dataset's solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-preview's overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using Sonnet-3.5\n",
"print(docs[0].get_content(metadata_mode=\"all\"))"
]
},
{
"cell_type": "code",
"execution_count": 8,
"id": "1511a30f-3efc-4142-9668-7dc056a24d0c",
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "1511a30f-3efc-4142-9668-7dc056a24d0c",
"outputId": "2e5e8e20-2b41-4183-f21f-dff503a03089"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|----------------|--------------------------------|\n",
"| **Attribute** | **Value** |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"**Basic Personal Information:** Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia.\n",
"\n",
"**Biomarker Measurements:** The subject's genetic profile includes an ApoE4 status of 0.0...\n",
"\n",
"**Cognitive and Neurofunctional Assessments:** The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall...\n",
"\n",
"**Volumetric Data:** Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross-Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 84422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 27820.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0...\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"----\n",
"\n",
"Skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-previews performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the models logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-previews solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-previews solutions. By comparing o1-previews solutions with the datasets solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-previews overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using GPT-4o\n",
"print(docs_gpt4o[0].get_content(metadata_mode=\"all\"))"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1c75bb85",
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.12.4"
}
"cells": [
{
"cell_type": "markdown",
"id": "97c79c38-38a3-40f3-ba2e-250649347d63",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_starter_multimodal.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
"nbformat": 4,
"nbformat_minor": 5
{
"cell_type": "markdown",
"id": "4e081457",
"metadata": {},
"source": [
"# Multimodal Parsing using LlamaParse\n",
"\n",
"This cookbook shows you how to use LlamaParse to parse any document with the multimodal capabilities of Multi-Modal LLMs from Anthropic/ OpenAI.\n",
"\n",
"LlamaParse allows you to plug in external, multimodal model vendors for parsing - we handle the error correction, validation, and scalability/reliability for you.\n"
]
},
{
"cell_type": "markdown",
"id": "qOdqBxCS51Ow",
"metadata": {},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "H_Vqcylb50vm",
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"id": "15e60ecf-519c-41fc-911b-765adaf8bad4",
"metadata": {},
"source": [
"### Setup\n",
"\n",
"Here we setup `LLAMA_CLOUD_API_KEY` for using `LlamaParse`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "91a9e532-1454-40e0-bbf0-fd442c350121",
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"# API access to llama-cloud\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"<YOUR LLAMACLOUD API KEY>\""
]
},
{
"cell_type": "markdown",
"id": "LGwBNPNotZRQ",
"metadata": {},
"source": [
"## Download Data\n",
"\n",
"For this demonstration, we will use OpenAI's recent paper `Evaluation of OpenAI o1: Opportunities and Challenges of AGI`."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "IjtKDQRLrylI",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2024-12-05 18:54:24-- https://arxiv.org/pdf/2409.18486\n",
"Resolving arxiv.org (arxiv.org)... 151.101.67.42, 151.101.131.42, 151.101.3.42, ...\n",
"Connecting to arxiv.org (arxiv.org)|151.101.67.42|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 13986265 (13M) [application/pdf]\n",
"Saving to: o1.pdf\n",
"\n",
"o1.pdf 100%[===================>] 13.34M 11.8MB/s in 1.1s \n",
"\n",
"2024-12-05 18:54:26 (11.8 MB/s) - o1.pdf saved [13986265/13986265]\n",
"\n"
]
}
],
"source": [
"!wget \"https://arxiv.org/pdf/2409.18486\" -O \"o1.pdf\""
]
},
{
"cell_type": "markdown",
"id": "4e29a9d7-5bd9-4fb8-8ec1-4c128a748662",
"metadata": {},
"source": [
"## Initialize LlamaParse\n",
"\n",
"Initialize LlamaParse in multimodal mode, and specify the vendor.\n",
"\n",
"**NOTE**: optionally you can specify the Anthropic/ OpenAI API key. If you choose to do so LlamaParse will only charge you 1 credit (0.3c) per page. \n",
"\n",
"\n",
"Using your own API key may incur additional costs from your model provider and could result in failed pages or documents if you do not have sufficient usage limits."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "dc921729-3446-42ca-8e1b-a6fd26195ed9",
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.schema import TextNode\n",
"from typing import List\n",
"\n",
"\n",
"def get_text_nodes(json_list: List[dict]):\n",
" text_nodes = []\n",
" for idx, page in enumerate(json_list):\n",
" text_node = TextNode(text=page[\"md\"], metadata={\"page\": page[\"page\"]})\n",
" text_nodes.append(text_node)\n",
" return text_nodes"
]
},
{
"cell_type": "markdown",
"id": "1b5d6da6",
"metadata": {},
"source": [
"### With anthropic-sonnet-3.5"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f2e9d9cf-8189-4fcb-b34f-cde6cc0b59c8",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id dd9d5e0f-160e-486a-89a2-6005e5a1c2ac\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model_name=\"anthropic-sonnet-3.5\",\n",
" target_pages=\"24\"\n",
" # invalidate_cache=True\n",
")\n",
"json_objs = parser.get_json_result(\"o1.pdf\")\n",
"json_list = json_objs[0][\"pages\"]\n",
"docs = get_text_nodes(json_list)"
]
},
{
"cell_type": "markdown",
"id": "4f3c51b0-7878-48d7-9bc3-02b516500128",
"metadata": {},
"source": [
"### With GPT-4o\n",
"\n",
"For comparison, we will also parse the document using GPT-4o."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "6fc3f258-50ae-4988-b904-c105463a498f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 6a4dea44-4f90-406b-b290-9e98620b1232\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser_gpt4o = LlamaParse(\n",
" result_type=\"markdown\",\n",
" use_vendor_multimodal_model=True,\n",
" vendor_multimodal_model=\"openai-gpt4o\",\n",
" target_pages=\"24\",\n",
" # invalidate_cache=True\n",
")\n",
"json_objs_gpt4o = parser_gpt4o.get_json_result(\"o1.pdf\")\n",
"json_list_gpt4o = json_objs_gpt4o[0][\"pages\"]\n",
"docs_gpt4o = get_text_nodes(json_list_gpt4o)"
]
},
{
"cell_type": "markdown",
"id": "44c20f7a-2901-4dd0-b635-a4b33c5664c1",
"metadata": {},
"source": [
"### View Results\n",
"\n",
"Let's visualize the results along with the original document page."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "778698aa-da7e-4081-b3b5-0372f228536f",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|-----------------|----------------------------------|\n",
"| Attribute | Value | Basic Personal Information: Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia. |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White | Biomarker Measurements: The subject's genetic profile includes an ApoE4 status of 0.0... |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... | Cognitive and Neurofunctional Assessments: The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 | Volumetric Data: Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 54422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 2782.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0.... |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-preview's performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the model's logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-preview's solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-preview's solutions. By comparing o1-preview's solutions with the dataset's solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-preview's overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using Sonnet-3.5\n",
"print(docs[0].get_content(metadata_mode=\"all\"))"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1511a30f-3efc-4142-9668-7dc056a24d0c",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"page: 25\n",
"\n",
"\n",
"| Participant_ID | clinical Description Reference |\n",
"|----------------|--------------------------------|\n",
"| **Attribute** | **Value** |\n",
"| Age | 72.0 |\n",
"| Sex | Female |\n",
"| Education | 15 |\n",
"| Race | White |\n",
"| DX_bl | AD |\n",
"| DX | Dementia |\n",
"| ... | ... |\n",
"| APOE4 | 1.0 |\n",
"| TAU | 212.5 |\n",
"| ... | ... |\n",
"| MMSE | 29.0 |\n",
"| CDRSB | 0.0 |\n",
"| ... | ... |\n",
"| FLDSTRENG | 1.5 Tesla MRI |\n",
"| Ventricles | 84599 |\n",
"| Hippocampus | 5319 |\n",
"| ... | ... |\n",
"\n",
"**Basic Personal Information:** Subject 098_S_0896 is a 72.0-year-old Female who has completed 15 years of education. The ethnicity is Not Hisp/Latino and race is White. Marital status is Married. Initially diagnosed as AD, as of the date 2007-10-24, the final diagnosis was Dementia.\n",
"\n",
"**Biomarker Measurements:** The subject's genetic profile includes an ApoE4 status of 0.0...\n",
"\n",
"**Cognitive and Neurofunctional Assessments:** The Mini-Mental State Examination score stands at 29.0. The Clinical Dementia Rating, sum of boxes, is 1.0. ADAS 11 and 13 scores are 4.67 and 4.67 respectively, with a score of 1.0 in delayed word recall...\n",
"\n",
"**Volumetric Data:** Under MRI conditions at a field strength of 1.5 Tesla MRI Tesla, using Cross-Sectional FreeSurfer (FreeSurfer Version 4.3), the imaging data recorded includes ventricles volume at 84422.0, hippocampus volume at 6677.0, whole brain volume at 1147980.0, entorhinal cortex volume at 27820.0, fusiform gyrus volume at 19432.0, and middle temporal area volume at 24951.0. The intracranial volume measured is 1799580.0...\n",
"\n",
"Figure 2: An example of a patient table and its corresponding clinical description.\n",
"\n",
"----\n",
"\n",
"Skills. Mathematics, as a highly structured and logic-driven discipline, provides an ideal testing ground for evaluating this reasoning ability. To investigate o1-previews performance, we designed a series of tests covering various difficulty levels. We begin with high school-level math competition problems in this section, followed by college-level mathematics problems in the next section, allowing us to observe the models logical reasoning across varying levels of complexity.\n",
"\n",
"In this section, we selected two primary areas of mathematics: algebra and counting and probability in this section. We chose these two topics because of their heavy reliance on problem-solving skills and their frequent use in assessing logical and abstract thinking [46]. The dataset used in testing is from the MATH dataset [46]. The problems in the dataset cover a wide range of subjects, including Prealgebra, Intermediate Algebra, Algebra, Geometry, Counting and Probability, Number Theory, and Precalculus. Each problem is categorized based on difficulty, ranked from level 1 to 5, according to the Art of Problem Solving (AoPS). The dataset mainly comprises problems from various high school math competitions, including the American Mathematics Competitions (AMC) 10 and 12, as well as the American Invitational Mathematics Examination (AIME), and other similar contests. Each problem comes with detailed reference solutions, allowing for a comprehensive comparison of o1-previews solutions.\n",
"\n",
"In addition to evaluating the final answers produced by o1-preview, our analysis delves into the step-by-step reasoning process of the o1-previews solutions. By comparing o1-previews solutions with the datasets solutions, we assess its ability to engage in logical reasoning, handle abstract problem-solving tasks, and apply structured approaches to reach correct answers. This deeper analysis offers insights into o1-previews overall reasoning capabilities, using mathematics as a reliable indicator for logical and structured thought processes.\n"
]
}
],
"source": [
"# using GPT-4o\n",
"print(docs_gpt4o[0].get_content(metadata_mode=\"all\"))"
]
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
@@ -43,7 +43,7 @@
},
{
"cell_type": "code",
"execution_count": 1,
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
@@ -69,7 +69,7 @@
},
{
"cell_type": "code",
"execution_count": 2,
"execution_count": null,
"metadata": {},
"outputs": [
{
@@ -105,7 +105,7 @@
},
{
"cell_type": "code",
"execution_count": 3,
"execution_count": null,
"metadata": {},
"outputs": [
{
@@ -119,17 +119,14 @@
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" target_pages=\"0,1,2\",\n",
" result_type=\"markdown\"\n",
")\n",
"parser = LlamaParse(target_pages=\"0,1,2\", result_type=\"markdown\")\n",
"\n",
"documents = parser.load_data('./uber_2021.pdf')"
"documents = parser.load_data(\"./uber_2021.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": 4,
"execution_count": null,
"metadata": {},
"outputs": [
{
@@ -140,7 +137,7 @@
" Document(id_='ad988239-3ab5-498d-85ba-a29241db24d4', embedding=None, metadata={}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, metadata_template='{key}: {value}', metadata_separator='\\n', text='# UBER TECHNOLOGIES, INC.\\n\\n# TABLE OF CONTENTS\\n\\n|Special Note Regarding Forward-Looking Statements|2|\\n|---|---|\\n|PART I|PART I|\\n|Item 1. Business|4|\\n|Item 1A. Risk Factors|11|\\n|Item 1B. Unresolved Staff Comments|46|\\n|Item 2. Properties|46|\\n|Item 3. Legal Proceedings|46|\\n|Item 4. Mine Safety Disclosures|47|\\n|PART II|PART II|\\n|Item 5. Market for Registrants Common Equity, Related Stockholder Matters and Issuer Purchases of Equity Securities|47|\\n|Item 6. [Reserved]|48|\\n|Item 7. Managements Discussion and Analysis of Financial Condition and Results of Operations|48|\\n|Item 7A. Quantitative and Qualitative Disclosures About Market Risk|69|\\n|Item 8. Financial Statements and Supplementary Data|70|\\n|Item 9. Changes in and Disagreements with Accountants on Accounting and Financial Disclosure|146|\\n|Item 9A. Controls and Procedures|147|\\n|Item 9B. Other Information|147|\\n|Item 9C. Disclosure Regarding Foreign Jurisdictions that Prevent Inspections|147|\\n|PART III|PART III|\\n|Item 10. Directors, Executive Officers and Corporate Governance|147|\\n|Item 11. Executive Compensation|147|\\n|Item 12. Security Ownership of Certain Beneficial Owners and Management and Related Stockholder Matters|148|\\n|Item 13. Certain Relationships and Related Transactions, and Director Independence|148|\\n|Item 14. Principal Accounting Fees and Services|148|\\n|PART IV|PART IV|\\n|Item 15. Exhibits, Financial Statement Schedules|148|\\n|Item 16. Form 10-K Summary|148|\\n|Exhibit Index|149|\\n|Signatures|152|', mimetype='text/plain', start_char_idx=None, end_char_idx=None, metadata_seperator='\\n', text_template='{metadata_str}\\n\\n{content}')]"
]
},
"execution_count": 4,
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
@@ -148,13 +145,6 @@
"source": [
"documents"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
@@ -172,8 +162,7 @@
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.12.4"
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
Binary file not shown.

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 173 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 200 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 115 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 350 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 47 KiB

@@ -0,0 +1,602 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/parsing_instructions.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"# Parsing documents with Instructions\n",
"\n",
"Parsing instructions allow you to guide our parsing model in the same way you would instruct an LLM.\n",
"\n",
"These instructions can be useful for improving the parser's performance on complex document layouts, extracting data in a specific format, or transforming the document in other ways.\n",
"\n",
"### Why This Matters:\n",
"Traditional document parsing can be rigid and error-prone, often missing crucial context and nuances in complex layouts. Our instruction-based parsing allows you to:\n",
"\n",
"1. Extract specific information with pinpoint accuracy\n",
"2. Handle complex document layouts with ease\n",
"3. Transform unstructured data into structured formats effortlessly\n",
"4. Save hours of manual data entry and verification\n",
"5. Reduce errors in document processing workflows\n",
"\n",
"In this demonstration, we showcase how parsing instructions can be used to extract specific information from unstructured documents. Below are the documents we use for testing:\n",
"\n",
"1. McDonald's Receipt - Extracting the price of each order and the final amount to be paid.\n",
"\n",
"2. Expense Report Document - Extracting employee name, employee ID, position, department, date ranges, individual expense items with dates, categories, and amounts.\n",
"\n",
"3. Purchase Order Document - Identifying the PO number, vendor details, shipping terms, and an itemized list of products with quantities and unit prices.\n",
"\n",
"Let's jump into these real-world examples and see how parsing instructions can help us extract specific information."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Setup API Key"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### McDonald's Receipt\n",
"\n",
"Here we extract the price of each order and the final amount to be paid."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"mcdonalds_receipt.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 66643b81-e2f4-408b-890b-8e116472210b\n"
]
}
],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\"./mcdonalds_receipt.png\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# Rate us HIGHLY SATISFIED\n",
"\n",
"Purchase any sandwich and receive a FREE ITEM\n",
"\n",
"Go to WWW.mcdvoice.com within 7 days of purchase of equal or lesser value and tell us about your visit.\n",
"\n",
"Validation Code: 31278-01121-21018-20481-00081-0\n",
"\n",
"Valid at participating US McDonald's\n",
"\n",
"Expires 30 days after receipt date\n",
"\n",
"# McDonald's Restaurant #312782378\n",
"\n",
"PINE RD NW\n",
"\n",
"RICE MN 56367-9740\n",
"\n",
"TEL# 320 393 4600\n",
"\n",
"KS# 12/08/2022 08:48 PM\n",
"\n",
"# Order\n",
"\n",
"|Happy Meal 6 Pc|$4.89|\n",
"|---|---|\n",
"|Creamy Ranch Cup| |\n",
"|Extra Kids Fry| |\n",
"|Wreck It Ralph 2 Snack| |\n",
"|Oreo McFlurry|$2.69|\n",
"\n",
"# Summary\n",
"\n",
"|Subtotal|$7.58|\n",
"|---|---|\n",
"|Tax|$0.52|\n",
"|Take-Out Total|$8.10|\n",
"|Cash Tendered|$10.00|\n",
"|Change|$1.90|\n",
"\n",
"### Not ACCEPTING APPLICATIONS *++ McDonald's Restaurant Rice\n",
"\n",
"Text to #36453 apply 31278\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 1a04fdbb-5415-4a36-a1bd-26bfb5d618fa\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"The provided document is a McDonald's receipt.\n",
" Provide the price of each order and final amount to be paid.\"\"\"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./mcdonalds_receipt.png\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Here are the prices for each order from the McDonald's receipt:\n",
"\n",
"1. Happy Meal 6 Pc: $4.89\n",
"2. Snack Oreo McFlurry: $2.69\n",
"\n",
"**Subtotal:** $7.58\n",
"**Tax:** $0.52\n",
"**Total Amount to be Paid:** $8.10\n",
"\n",
"The cash tendered was $10.00, and the change given was $1.90.\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Expense Report Document\n",
"\n",
"Here we extract employee name, employee ID, position, department, date ranges, individual expense items with dates, categories, and amounts."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"expense_report_document.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b6bcc6e1-7d30-4522-9abd-ace196781a70\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\n",
" \"./expense_report_document.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# QUANTUM DYNAMICS CORPORATION\n",
"\n",
"# EMPLOYEE EXPENSE REPORT\n",
"\n",
"# FISCAL YEAR 2024\n",
"\n",
"# EMPLOYEE INFORMATION:\n",
"\n",
"Name: Dr. Alexandra Chen-Martinez, PhD\n",
"\n",
"Employee ID: QD-2022-1457\n",
"\n",
"Department: Advanced Research & Development\n",
"\n",
"Cost Center: CC-ARD-NA-003\n",
"\n",
"Project Codes: QD-QUANTUM-2024-01, QD-AI-2024-03\n",
"\n",
"Position: Principal Research Scientist\n",
"\n",
"Reporting Manager: Dr. James Thompson\n",
"\n",
"# TRIP/EXPENSE PERIOD:\n",
"\n",
"Start Date: November 15, 2024\n",
"\n",
"End Date: December 10, 2024\n",
"\n",
"Purpose: International Conference Attendance & Client Meetings\n",
"\n",
"Locations: Tokyo, Japan → Singapore → Sydney, Australia\n",
"\n",
"# CURRENCY CONVERSION RATES APPLIED:\n",
"\n",
"JPY (¥) → USD: 0.0068 (as of 11/15/2024)\n",
"\n",
"SGD (S$) → USD: 0.74 (as of 11/28/2024)\n",
"\n",
"AUD (A$) → USD: 0.65 (as of 12/03/2024)\n",
"\n",
"# ITEMIZED EXPENSES:\n",
"\n",
"|Date|Category|Description|Original|Currency|USD|\n",
"|---|---|---|---|---|---|\n",
"|11/15/2024|Transportation|JFK → NRT Business Class|4,250.00|USD|4,250.00|\n",
"|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|Booking Ref: QF78956 - Corporate Rate Applied|\n",
"|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|Project Code: QD-QUANTUM-2024-01|\n",
"|11/16/2024|Accommodation|Hilton Tokyo - 5 nights|225,000|JPY|1,530.00|\n",
"|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|Confirmation: HTK-2024-78956|\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 7b0d05bb-947b-4475-8d0f-f10386f7446e\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"You are provided with an expense report. \n",
"Extract employee name, employee id, position, department, date ranges, individual expense items with dates, categories, and amounts.\"\"\"\n",
"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./expense_report_document.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"**Employee Information:**\n",
"- **Name:** Dr. Alexandra Chen-Martinez, PhD\n",
"- **Employee ID:** QD-2022-1457\n",
"- **Position:** Principal Research Scientist\n",
"- **Department:** Advanced Research & Development\n",
"\n",
"**Trip/Expense Period:**\n",
"- **Start Date:** November 15, 2024\n",
"- **End Date:** December 10, 2024\n",
"\n",
"**Expense Items:**\n",
"1. **Date:** 11/15/2024\n",
"- **Category:** Transportation\n",
"- **Description:** JFK → NRT Business Class\n",
"- **Original Amount:** $4,250.00\n",
"- **Currency:** USD\n",
"- **USD Amount:** $4,250.00\n",
"- **Booking Reference:** QF78956 - Corporate Rate Applied\n",
"- **Project Code:** QD-QUANTUM-2024-01\n",
"\n",
"2. **Date:** 11/16/2024\n",
"- **Category:** Accommodation\n",
"- **Description:** Hilton Tokyo - 5 nights\n",
"- **Original Amount:** ¥225,000\n",
"- **Currency:** JPY\n",
"- **USD Amount:** $1,530.00\n",
"- **Confirmation:** HTK-2024-78956\n",
"\n",
"**Locations:**\n",
"- Tokyo, Japan\n",
"- Singapore\n",
"- Sydney, Australia\n",
"\n",
"**Currency Conversion Rates Applied:**\n",
"- JPY (¥) → USD: 0.0068 (as of 11/15/2024)\n",
"- SGD (S$) → USD: 0.74 (as of 11/28/2024)\n",
"- AUD (A$) → USD: 0.65 (as of 12/03/2024)\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Purchase Order Document \n",
"\n",
"Here we identify the PO number, vendor details, shipping terms, and an itemized list of products with quantities and unit prices."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<img src=\"purchase_order_document.png\" alt=\"Alt Text\" width=\"500\">"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id b8cb11c3-7dce-4e6a-94bb-1a4e50e45e55\n"
]
}
],
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\n",
" \"./purchase_order_document.pdf\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"# GLOBAL TECH SOLUTIONS, INC.\n",
"\n",
"# PURCHASE ORDER\n",
"\n",
"Document Reference: PO-2024-GT-9876/REV.2\n",
"\n",
"[Original: PO-2024-GT-9876]\n",
"\n",
"Amendment Date: 12/10/2024\n",
"\n",
"# VENDOR INFORMATION:\n",
"\n",
"Quantum Electronics Manufacturing\n",
"\n",
"DUNS: 78-456-7890\n",
"\n",
"Tax ID: EU8976543210\n",
"\n",
"Hoofdorp, Netherlands\n",
"\n",
"Vendor #: QEM-EU-2024-001\n",
"\n",
"# SHIP TO:\n",
"\n",
"Global Tech Solutions, Inc.\n",
"\n",
"Building 7A, Innovation Park\n",
"\n",
"2100 Technology Drive\n",
"\n",
"Austin, TX 78701\n",
"\n",
"USA\n",
"\n",
"Attn: Sarah Martinez, Receiving Manager\n",
"\n",
"Tel: +1 (512) 555-0123\n",
"\n",
"# PAYMENT TERMS:\n",
"\n",
"Net 45\n",
"\n",
"2% discount if paid within 15 days\n",
"\n",
"# SHIPPING TERMS:\n",
"\n",
"DDP (Delivered Duty Paid) - Incoterms 2020\n",
"\n",
"Insurance Required: Yes\n",
"\n",
"Preferred Carrier: DHL/FedEx\n",
"\n",
"Required Delivery Date: 01/15/2025\n",
"\n",
"# SPECIAL INSTRUCTIONS:\n",
"\n",
"1. All shipments must include Certificate of Conformance\n",
"2. ESD-sensitive items must be properly packaged\n",
"3. Temperature logging required for items marked with *\n",
"4. Partial shipments accepted with prior approval\n",
"5. Quote PO number on all correspondence\n",
"\n",
"# ITEM DETAILS:\n",
"\n",
"|Line|Part Number|Description|Qty|UOM|Unit Price|Total|\n",
"|---|---|---|---|---|---|---|\n",
"|1|QE-MCU-5590|Microcontroller Unit|500|EA|$12.50|$6,250.00|\n"
]
}
],
"source": [
"print(vanilaParsing[0].text)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id d2731305-984d-4633-8a52-0493748cf10b\n"
]
}
],
"source": [
"parsingInstruction = \"\"\"You are provided with a purchase order. \n",
"Identify the PO number, vendor details, shipping terms, and itemized list of products with quantities and unit prices.\"\"\"\n",
"\n",
"withInstructionParsing = LlamaParse(\n",
" result_type=\"markdown\", parsing_instruction=parsingInstruction\n",
").load_data(\"./purchase_order_document.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Here are the details extracted from the purchase order:\n",
"\n",
"**PO Number:** PO-2024-GT-9876/REV.2\n",
"\n",
"**Vendor Details:**\n",
"- **Vendor Name:** Quantum Electronics Manufacturing\n",
"- **DUNS:** 78-456-7890\n",
"- **Tax ID:** EU8976543210\n",
"- **Address:** Hoofdorp, Netherlands\n",
"- **Vendor Number:** QEM-EU-2024-001\n",
"- **Contact Person:** Sarah Martinez, Receiving Manager\n",
"- **Phone:** +1 (512) 555-0123\n",
"\n",
"**Shipping Terms:**\n",
"- **Terms:** DDP (Delivered Duty Paid) - Incoterms 2020\n",
"- **Insurance Required:** Yes\n",
"- **Preferred Carrier:** DHL/FedEx\n",
"- **Required Delivery Date:** 01/15/2025\n",
"\n",
"**Itemized List of Products:**\n",
"1. **Part Number:** QE-MCU-5590\n",
"- **Description:** Microcontroller Unit\n",
"- **Quantity:** 500 EA\n",
"- **Unit Price:** $12.50\n",
"- **Total:** $6,250.00\n",
"\n",
"**Payment Terms:**\n",
"- Net 45\n",
"- 2% discount if paid within 15 days\n",
"\n",
"**Special Instructions:**\n",
"1. All shipments must include Certificate of Conformance\n",
"2. ESD-sensitive items must be properly packaged\n",
"3. Temperature logging required for items marked with *\n",
"4. Partial shipments accepted with prior approval\n",
"5. Quote PO number on all correspondence\n"
]
}
],
"source": [
"print(withInstructionParsing[0].text)"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llamacloud",
"language": "python",
"name": "llamacloud"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 344 KiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 2.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 100 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 464 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 410 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 444 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 610 KiB

+286 -111
View File
@@ -1,29 +1,29 @@
import os
import asyncio
import mimetypes
import os
import time
from contextlib import asynccontextmanager
from copy import deepcopy
from io import BufferedIOBase
from pathlib import Path, PurePath, PurePosixPath
from typing import Any, AsyncGenerator, Dict, List, Optional, Union
from urllib.parse import urlparse
import httpx
import mimetypes
import time
from pathlib import Path, PurePath, PurePosixPath
from typing import AsyncGenerator, Any, Dict, List, Optional, Union
from contextlib import asynccontextmanager
from io import BufferedIOBase
from fsspec import AbstractFileSystem
from llama_index.core.async_utils import asyncio_run, run_jobs
from llama_index.core.bridge.pydantic import Field, field_validator
from llama_index.core.bridge.pydantic import Field, PrivateAttr, field_validator
from llama_index.core.constants import DEFAULT_BASE_URL
from llama_index.core.readers.base import BasePydanticReader
from llama_index.core.readers.file.base import get_default_fs
from llama_index.core.schema import Document
from llama_parse.utils import (
SUPPORTED_FILE_TYPES,
ResultType,
nest_asyncio_err,
nest_asyncio_msg,
ResultType,
SUPPORTED_FILE_TYPES,
)
from copy import deepcopy
# can put in a path to the file or the file bytes itself
# if passing as bytes or a buffer, must provide the file_name in extra_info
@@ -32,6 +32,11 @@ FileInput = Union[str, bytes, BufferedIOBase]
_DEFAULT_SEPARATOR = "\n---\n"
JOB_RESULT_URL = "/api/parsing/job/{job_id}/result/{result_type}"
JOB_STATUS_ROUTE = "/api/parsing/job/{job_id}"
JOB_UPLOAD_ROUTE = "/api/parsing/upload"
class LlamaParse(BasePydanticReader):
"""A smart-parser for files."""
@@ -49,9 +54,11 @@ class LlamaParse(BasePydanticReader):
default=1,
description="The interval in seconds to check if the parsing is done.",
)
custom_client: Optional[httpx.AsyncClient] = Field(
default=None, description="A custom HTTPX client to use for sending requests."
)
ignore_errors: bool = Field(
default=True,
description="Whether or not to ignore and skip errors raised during parsing.",
@@ -133,6 +140,14 @@ class LlamaParse(BasePydanticReader):
default=None,
description="The top margin of the bounding box to use to extract text from documents expressed as a float between 0 and 1 representing the percentage of the page height.",
)
complemental_formatting_instruction: Optional[str] = Field(
default=None,
description="The complemental formatting instruction for the parser. Tell llamaParse how some thing should to be formatted, while retaining the markdown output.",
)
content_guideline_instruction: Optional[str] = Field(
default=None,
description="The content guideline for the parser. Tell LlamaParse how the content should be changed / transformed.",
)
continuous_mode: Optional[bool] = Field(
default=False,
description="Parse documents continuously, leading to better results on documents where tables span across two pages.",
@@ -157,10 +172,18 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, the parser will extract/tag charts from the document.",
)
extract_layout: Optional[bool] = Field(
default=False,
description="If set to true, the parser will extract the layout information of the document. Cost 1 credit per page.",
)
fast_mode: Optional[bool] = Field(
default=False,
description="Note: Non compatible with gpt-4o. If set to true, the parser will use a faster mode to extract text from documents. This mode will skip OCR of images, and table/heading reconstruction.",
)
formatting_instruction: Optional[str] = Field(
default=None,
description="The Formatting instruction for the parser. Override default llamaParse behavior. In most case you want to use complemental_formatting_instruction instead.",
)
guess_xlsx_sheet_names: Optional[bool] = Field(
default=False,
description="Whether to guess the sheet names of the xlsx file.",
@@ -173,17 +196,33 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, when parsing HTML the parser will remove fixed elements. Useful to hide cookie banners.",
)
html_remove_navigation_elements: Optional[bool] = Field(
default=False,
description="If set to true, when parsing HTML the parser will remove navigation elements. Useful to hide menus, header, footer.",
)
http_proxy: Optional[str] = Field(
default=None,
description="(optional) If set with input_url will use the specified http proxy to download the file.",
)
ignore_document_elements_for_layout_detection: Optional[bool] = Field(
default=False,
description="If set to true, the parser will ignore document elements for layout detection and only rely on a vision model.",
)
input_s3_region: Optional[str] = Field(
default=None,
description="The region of the input S3 bucket if input_s3_path is specified.",
)
invalidate_cache: Optional[bool] = Field(
default=False,
description="If set to true, the cache will be ignored and the document re-processes. All document are kept in cache for 48hours after the job was completed to avoid processing the same document twice.",
)
is_formatting_instruction: Optional[bool] = Field(
default=False,
description="Allow the parsing instruction to also format the output. Disable to have a cleaner markdown output.",
job_timeout_extra_time_per_page_in_seconds: Optional[float] = Field(
default=None,
description="The extra time in seconds to wait for the parsing to finish per page. Get added to job_timeout_in_seconds.",
)
job_timeout_in_seconds: Optional[float] = Field(
default=None,
description="The maximum timeout in seconds to wait for the parsing to finish. Override default timeout of 30 minutes. Minimum is 120 seconds.",
)
language: Optional[str] = Field(
default="en", description="The language of the text to parse."
@@ -200,6 +239,14 @@ class LlamaParse(BasePydanticReader):
default=None,
description="An S3 path prefix to store the output of the parsing job. If set, the parser will upload the output to S3. The bucket need to be accessible from the LlamaIndex organization.",
)
output_s3_region: Optional[str] = Field(
default=None,
description="The AWS region of the output S3 bucket defined in output_s3_path_prefix.",
)
output_tables_as_HTML: Optional[bool] = Field(
default=False,
description="If set to true, the parser will output tables as HTML in the markdown.",
)
page_prefix: Optional[str] = Field(
default=None,
description="A templated prefix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
@@ -212,9 +259,6 @@ class LlamaParse(BasePydanticReader):
default=None,
description="A templated suffix to add to the beginning of each page. If it contain `{page_number}`, it will be replaced by the page number.",
)
parsing_instruction: Optional[str] = Field(
default="", description="The parsing instruction for the parser."
)
premium_mode: Optional[bool] = Field(
default=False,
description="Use our best parser mode if set to True.",
@@ -223,6 +267,31 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, the parser will ignore diagonal text (when the text rotation in degrees modulo 90 is not 0).",
)
spreadsheet_extract_sub_tables: Optional[bool] = Field(
default=False,
description="If set to true, the parser will extract sub-tables from the spreadsheet when possible (more than one table per sheet).",
)
strict_mode_buggy_font: Optional[bool] = Field(
default=False,
description="If set to true, the parser will fail if it can't extract text from a document because of a buggy font.",
)
strict_mode_image_extraction: Optional[bool] = Field(
default=False,
description="If set to true, the parser will fail if it can't extract an image from the document.",
)
strict_mode_image_ocr: Optional[bool] = Field(
default=False,
description="If set to true, the parser will fail if it can't OCR an image from the document.",
)
strict_mode_reconstruction: Optional[bool] = Field(
default=False,
description="If set to true, the parser will fail if it can't reconstruct a table or a heading from the document.",
)
structured_output: Optional[bool] = Field(
default=False,
description="If set to true, the parser will output structured data based on the provided JSON Schema.",
@@ -273,6 +342,13 @@ class LlamaParse(BasePydanticReader):
default=None,
description="The API key for the GPT-4o API. Lowers the cost of parsing.",
)
is_formatting_instruction: Optional[bool] = Field(
default=False,
description="Allow the parsing instruction to also format the output. Disable to have a cleaner markdown output.",
)
parsing_instruction: Optional[str] = Field(
default="", description="The parsing instruction for the parser."
)
@field_validator("api_key", mode="before", check_fields=True)
@classmethod
@@ -295,6 +371,25 @@ class LlamaParse(BasePydanticReader):
url = os.getenv("LLAMA_CLOUD_BASE_URL", None)
return url or v or DEFAULT_BASE_URL
_aclient: Union[httpx.AsyncClient, None] = PrivateAttr(default=None, init=False)
@property
def aclient(self) -> httpx.AsyncClient:
if not self._aclient:
self._aclient = self.custom_client or httpx.AsyncClient()
# need to do this outside instantiation in case user
# updates base_url, api_key, or max_timeout later
# ... you wouldn't usually expect that, except
# if someone does do it and it doesn't reflect on
# the client they'll end up pretty confused, so
# for the sake of ergonomics...
self._aclient.base_url = self.base_url
self._aclient.headers["Authorization"] = f"Bearer {self.api_key}"
self._aclient.timeout = self.max_timeout
return self._aclient
@asynccontextmanager
async def client_context(self) -> AsyncGenerator[httpx.AsyncClient, None]:
"""Create a context for the HTTPX client."""
@@ -343,8 +438,6 @@ class LlamaParse(BasePydanticReader):
extra_info: Optional[dict] = None,
fs: Optional[AbstractFileSystem] = None,
) -> str:
headers = {"Authorization": f"Bearer {self.api_key}"}
url = f"{self.base_url}/api/parsing/upload"
files = None
file_handle = None
input_url = file_input if self._is_input_url(file_input) else None
@@ -435,6 +528,14 @@ class LlamaParse(BasePydanticReader):
if self.bbox_top is not None:
data["bbox_top"] = self.bbox_top
if self.complemental_formatting_instruction:
data[
"complemental_formatting_instruction"
] = self.complemental_formatting_instruction
if self.content_guideline_instruction:
data["content_guideline_instruction"] = self.content_guideline_instruction
if self.continuous_mode:
data["continuous_mode"] = self.continuous_mode
@@ -453,9 +554,15 @@ class LlamaParse(BasePydanticReader):
if self.extract_charts:
data["extract_charts"] = self.extract_charts
if self.extract_layout:
data["extract_layout"] = self.extract_layout
if self.fast_mode:
data["fast_mode"] = self.fast_mode
if self.formatting_instruction:
data["formatting_instruction"] = self.formatting_instruction
if self.guess_xlsx_sheet_names:
data["guess_xlsx_sheet_names"] = self.guess_xlsx_sheet_names
@@ -465,9 +572,19 @@ class LlamaParse(BasePydanticReader):
if self.html_remove_fixed_elements:
data["html_remove_fixed_elements"] = self.html_remove_fixed_elements
if self.html_remove_navigation_elements:
data[
"html_remove_navigation_elements"
] = self.html_remove_navigation_elements
if self.http_proxy is not None:
data["http_proxy"] = self.http_proxy
if self.ignore_document_elements_for_layout_detection:
data[
"ignore_document_elements_for_layout_detection"
] = self.ignore_document_elements_for_layout_detection
if input_url is not None:
files = None
data["input_url"] = str(input_url)
@@ -476,12 +593,23 @@ class LlamaParse(BasePydanticReader):
files = None
data["input_s3_path"] = str(input_s3_path)
if self.input_s3_region is not None:
data["input_s3_region"] = self.input_s3_region
if self.invalidate_cache:
data["invalidate_cache"] = self.invalidate_cache
if self.is_formatting_instruction:
data["is_formatting_instruction"] = self.is_formatting_instruction
if self.job_timeout_extra_time_per_page_in_seconds is not None:
data[
"job_timeout_extra_time_per_page_in_seconds"
] = self.job_timeout_extra_time_per_page_in_seconds
if self.job_timeout_in_seconds is not None:
data["job_timeout_in_seconds"] = self.job_timeout_in_seconds
if self.language:
data["language"] = self.language
@@ -494,6 +622,12 @@ class LlamaParse(BasePydanticReader):
if self.output_s3_path_prefix is not None:
data["output_s3_path_prefix"] = self.output_s3_path_prefix
if self.output_s3_region is not None:
data["output_s3_region"] = self.output_s3_region
if self.output_tables_as_HTML:
data["output_tables_as_HTML"] = self.output_tables_as_HTML
if self.page_prefix is not None:
data["page_prefix"] = self.page_prefix
@@ -506,6 +640,9 @@ class LlamaParse(BasePydanticReader):
data["page_suffix"] = self.page_suffix
if self.parsing_instruction is not None:
print(
"WARNING: parsing_instruction is deprecated. Use complemental_formatting_instruction or content_guideline_instruction instead."
)
data["parsing_instruction"] = self.parsing_instruction
if self.premium_mode:
@@ -514,6 +651,21 @@ class LlamaParse(BasePydanticReader):
if self.skip_diagonal_text:
data["skip_diagonal_text"] = self.skip_diagonal_text
if self.spreadsheet_extract_sub_tables:
data["spreadsheet_extract_sub_tables"] = self.spreadsheet_extract_sub_tables
if self.strict_mode_buggy_font:
data["strict_mode_buggy_font"] = self.strict_mode_buggy_font
if self.strict_mode_image_extraction:
data["strict_mode_image_extraction"] = self.strict_mode_image_extraction
if self.strict_mode_image_ocr:
data["strict_mode_image_ocr"] = self.strict_mode_image_ocr
if self.strict_mode_reconstruction:
data["strict_mode_reconstruction"] = self.strict_mode_reconstruction
if self.structured_output:
data["structured_output"] = self.structured_output
@@ -554,17 +706,12 @@ class LlamaParse(BasePydanticReader):
data["gpt4o_api_key"] = self.gpt4o_api_key
try:
async with self.client_context() as client:
response = await client.post(
url,
files=files,
headers=headers,
data=data,
)
if not response.is_success:
raise Exception(f"Failed to parse the file: {response.text}")
job_id = response.json()["id"]
return job_id
resp = await self.aclient.post(JOB_UPLOAD_ROUTE, files=files, data=data) # type: ignore
resp.raise_for_status() # this raises if status is not 2xx
return resp.json()["id"]
except httpx.HTTPStatusError as err: # this catches it
msg = f"Failed to parse the file: {err.response.text}"
raise Exception(msg) from err # this preserves the exception context
finally:
if file_handle is not None:
file_handle.close()
@@ -572,52 +719,51 @@ class LlamaParse(BasePydanticReader):
async def _get_job_result(
self, job_id: str, result_type: str, verbose: bool = False
) -> Dict[str, Any]:
result_url = f"{self.base_url}/api/parsing/job/{job_id}/result/{result_type}"
status_url = f"{self.base_url}/api/parsing/job/{job_id}"
headers = {"Authorization": f"Bearer {self.api_key}"}
start = time.time()
tries = 0
# so we're not re-setting the headers & stuff on each
# usage... assume that there is not some other
# coro also modifying base_url and the other client related configs.
client = self.aclient
while True:
await asyncio.sleep(self.check_interval)
async with self.client_context() as client:
tries += 1
tries += 1
result = await client.get(JOB_STATUS_ROUTE.format(job_id=job_id))
if result.status_code != 200:
end = time.time()
if end - start > self.max_timeout:
raise Exception(f"Timeout while parsing the file: {job_id}")
if verbose and tries % 10 == 0:
print(".", end="", flush=True)
await asyncio.sleep(self.check_interval)
continue
result = await client.get(status_url, headers=headers)
# Allowed values "PENDING", "SUCCESS", "ERROR", "CANCELED"
result_json = result.json()
status = result_json["status"]
if status == "SUCCESS":
parsed_result = await client.get(
JOB_RESULT_URL.format(job_id=job_id, result_type=result_type),
)
return parsed_result.json()
if result.status_code != 200:
end = time.time()
if end - start > self.max_timeout:
raise Exception(f"Timeout while parsing the file: {job_id}")
if verbose and tries % 10 == 0:
print(".", end="", flush=True)
elif status == "PENDING":
end = time.time()
if end - start > self.max_timeout:
raise Exception(f"Timeout while parsing the file: {job_id}")
if verbose and tries % 10 == 0:
print(".", end="", flush=True)
await asyncio.sleep(self.check_interval)
await asyncio.sleep(self.check_interval)
else:
error_code = result_json.get("error_code", "No error code found")
error_message = result_json.get(
"error_message", "No error message found"
)
continue
# Allowed values "PENDING", "SUCCESS", "ERROR", "CANCELED"
result_json = result.json()
status = result_json["status"]
if status == "SUCCESS":
parsed_result = await client.get(result_url, headers=headers)
return parsed_result.json()
elif status == "PENDING":
end = time.time()
if end - start > self.max_timeout:
raise Exception(f"Timeout while parsing the file: {job_id}")
if verbose and tries % 10 == 0:
print(".", end="", flush=True)
await asyncio.sleep(self.check_interval)
else:
error_code = result_json.get("error_code", "No error code found")
error_message = result_json.get(
"error_message", "No error message found"
)
exception_str = f"Job ID: {job_id} failed with status: {status}, Error code: {error_code}, Error message: {error_message}"
raise Exception(exception_str)
exception_str = f"Job ID: {job_id} failed with status: {status}, Error code: {error_code}, Error message: {error_message}"
raise Exception(exception_str)
async def _aload_data(
self,
@@ -778,54 +924,77 @@ class LlamaParse(BasePydanticReader):
else:
raise e
async def aget_images(
self, json_result: List[dict], download_path: str
async def aget_assets(
self, json_result: List[dict], download_path: str, asset_key: str
) -> List[dict]:
"""Download images from the parsed result."""
headers = {"Authorization": f"Bearer {self.api_key}"}
# make the download path
"""Download assets (images or charts) from the parsed result."""
# Make the download path
if not os.path.exists(download_path):
os.makedirs(download_path)
client = self.aclient
try:
images = []
assets = []
for result in json_result:
job_id = result["job_id"]
for page in result["pages"]:
if self.verbose:
print(f"> Image for page {page['page']}: {page['images']}")
for image in page["images"]:
image_name = image["name"]
print(
f"> {asset_key.capitalize()} for page {page['page']}: {page[asset_key]}"
)
for asset in page[asset_key]:
asset_name = asset["name"]
# get the full path
image_path = os.path.join(
download_path, f"{job_id}-{image_name}"
# Get the full path
asset_path = os.path.join(
download_path, f"{job_id}-{asset_name}"
)
# get a valid image path
if not image_path.endswith(".png"):
if not image_path.endswith(".jpg"):
image_path += ".png"
# Get a valid asset path
if not asset_path.endswith(".png"):
if not asset_path.endswith(".jpg"):
asset_path += ".png"
image["path"] = image_path
image["job_id"] = job_id
asset["path"] = asset_path
asset["job_id"] = job_id
asset["original_file_path"] = result.get("file_path", None)
asset["page_number"] = page["page"]
image["original_file_path"] = result.get("file_path", None)
image["page_number"] = page["page"]
with open(image_path, "wb") as f:
image_url = f"{self.base_url}/api/parsing/job/{job_id}/result/image/{image_name}"
async with self.client_context() as client:
res = await client.get(
image_url, headers=headers, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
images.append(image)
return images
with open(asset_path, "wb") as f:
asset_url = f"{self.base_url}/api/parsing/job/{job_id}/result/image/{asset_name}"
resp = await client.get(asset_url)
resp.raise_for_status()
f.write(resp.content)
assets.append(asset)
return assets
except Exception as e:
print("Error while downloading images from the parsed result:", e)
print(f"Error while downloading {asset_key} from the parsed result:", e)
if self.ignore_errors:
return []
else:
raise e
async def aget_images(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download images from the parsed result."""
try:
return await self.aget_assets(json_result, download_path, "images")
except Exception as e:
print("Error while downloading images:", e)
if self.ignore_errors:
return []
else:
raise e
async def aget_charts(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download charts from the parsed result."""
try:
return await self.aget_assets(json_result, download_path, "charts")
except Exception as e:
print("Error while downloading charts:", e)
if self.ignore_errors:
return []
else:
@@ -841,15 +1010,24 @@ class LlamaParse(BasePydanticReader):
else:
raise e
def get_charts(self, json_result: List[dict], download_path: str) -> List[dict]:
"""Download charts from the parsed result."""
try:
return asyncio_run(self.aget_charts(json_result, download_path))
except RuntimeError as e:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
else:
raise e
async def aget_xlsx(
self, json_result: List[dict], download_path: str
) -> List[dict]:
"""Download images from the parsed result."""
headers = {"Authorization": f"Bearer {self.api_key}"}
"""Download xlsx from the parsed result."""
# make the download path
if not os.path.exists(download_path):
os.makedirs(download_path)
client = self.aclient
try:
xlsx_list = []
for result in json_result:
@@ -869,12 +1047,9 @@ class LlamaParse(BasePydanticReader):
xlsx_url = (
f"{self.base_url}/api/parsing/job/{job_id}/result/raw/xlsx"
)
async with self.client_context() as client:
res = await client.get(
xlsx_url, headers=headers, timeout=self.max_timeout
)
res.raise_for_status()
f.write(res.content)
res = await client.get(xlsx_url)
res.raise_for_status()
f.write(res.content)
xlsx_list.append(xlsx)
return xlsx_list
+7
View File
@@ -191,4 +191,11 @@ SUPPORTED_FILE_TYPES = [
".xlr",
".eth",
".tsv",
".mp3",
".mp4",
".mpeg",
".mpga",
".m4a",
".wav",
".webm",
]
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "poetry.core.masonry.api"
[tool.poetry]
name = "llama-parse"
version = "0.5.16"
version = "0.5.20"
description = "Parse files into RAG-Optimized formats."
authors = ["Logan Markewich <logan@llamaindex.ai>"]
license = "MIT"
+2 -2
View File
@@ -153,10 +153,10 @@ async def test_input_url() -> None:
@pytest.mark.asyncio
async def test_input_url_with_website_input() -> None:
parser = LlamaParse(result_type="markdown")
input_url = "https://www.google.com"
input_url = "https://www.example.com"
result = await parser.aload_data(input_url)
assert len(result) == 1
assert "google" in result[0].text.lower()
assert "example" in result[0].text.lower()
@pytest.mark.skipif(