Compare commits

..

13 Commits

Author SHA1 Message Date
Pierre-Loic Doulcet 2e6c064682 add more supported format 2024-04-23 11:18:38 +08:00
Jerry Liu 4252f6186b fixes to insurance demo (#97) 2024-03-21 00:02:47 -07:00
Jerry Liu 22148ade9f fix advanced RAG notebook title (#98) 2024-03-21 00:02:38 -07:00
Jerry Liu b8332fe8e1 nit: add colab badge to mongodb notebook (#109) 2024-03-21 00:02:29 -07:00
Ravi Theja e40e92a133 Add mongodb llamaparse example (#107) 2024-03-20 23:37:42 -07:00
Jerry Liu ba8f345f80 Revert "cr"
This reverts commit 2ddbf1ba0d.
2024-03-19 00:21:29 -07:00
Jerry Liu 2ddbf1ba0d cr 2024-03-19 00:20:42 -07:00
Haotian Zhang 23567c8f98 Init LlamaParseJsonNodeParser example (#93) 2024-03-18 15:15:36 -04:00
Logan 8d39ae7763 add agent demo (#88)
* add agent demo

* remove mention of react agent

* agents folder
2024-03-18 16:12:56 +01:00
Jerry Liu a2edc41fc7 nit: fix grammar in insurance cookbook (#89)
cr
2024-03-18 16:11:03 +01:00
Ikko Eltociear Ashimine 591b6fc44d Update demo_parsing_instructions.ipynb (#86)
usefull -> useful
2024-03-16 18:15:30 +01:00
Pierre-Loic Doulcet f8a3d92ce0 demo insurance + parsing instructions (#84) 2024-03-16 18:14:49 +01:00
Laurie Voss a1d18d83da Adding parsing instructions demo (#82) 2024-03-15 10:07:53 +01:00
8 changed files with 1912 additions and 156 deletions
@@ -0,0 +1,302 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# LlamaParse Agent\n",
"\n",
"This demo walks through using an OpenAI Agent with [LlamaParse](https://cloud.llamaindex.ai)."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install llama-parse llama-index llama-index-postprocessor-sbert-rerank"
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\"\n",
"os.environ[\"OPENAI_API_KEY\"] = \"sk-...\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import Settings\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.llms.openai import OpenAI\n",
"\n",
"Settings.embed_model = OpenAIEmbedding(model=\"text-embedding-3-small\")\n",
"Settings.llm = OpenAI(model=\"gpt-3.5-turbo\", temperature=0.2)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Parsing \n",
"\n",
"For parsing, lets use a [recent paper](https://huggingface.co/papers/2403.09611) on Multi-Modal pretraining"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!wget https://arxiv.org/pdf/2403.09611.pdf -O paper.pdf"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Below, we can tell the parser to skip content we don't want. In this case, the references section will just add noise to a RAG system."
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {},
"outputs": [],
"source": [
"from llama_parse import LlamaParse\n",
"\n",
"parser = LlamaParse(\n",
" result_type=\"markdown\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 81251f39-01be-434e-99e8-1c1b83b82098\n"
]
}
],
"source": [
"documents = await parser.aload_data(\"paper.pdf\")"
]
},
{
"cell_type": "code",
"execution_count": 5,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Embeddings have been explicitly disabled. Using MockEmbedding.\n"
]
},
{
"name": "stderr",
"output_type": "stream",
"text": [
"41it [00:00, 26765.21it/s]\n",
"100%|██████████| 41/41 [00:13<00:00, 2.98it/s]\n"
]
}
],
"source": [
"import nest_asyncio\n",
"nest_asyncio.apply()\n",
"\n",
"from llama_index.core.node_parser import MarkdownElementNodeParser, SentenceSplitter\n",
"\n",
"# explicitly extract tables with the MarkdownElementNodeParser\n",
"node_parser = MarkdownElementNodeParser(num_workers=8)\n",
"nodes = node_parser.get_nodes_from_documents(documents)\n",
"nodes, objects = node_parser.get_nodes_and_objects(nodes)\n",
"\n",
"# Chain splitters to ensure chunk size requirements are met\n",
"nodes = SentenceSplitter(chunk_size=512, chunk_overlap=20).get_nodes_from_documents(nodes)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Chat over the paper, lets find out what it is about!"
]
},
{
"cell_type": "code",
"execution_count": 6,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core import VectorStoreIndex, SummaryIndex\n",
"\n",
"vector_index = VectorStoreIndex(nodes=nodes)\n",
"summary_index = SummaryIndex(nodes=nodes)"
]
},
{
"cell_type": "code",
"execution_count": 7,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.agent.openai import OpenAIAgent\n",
"from llama_index.core.tools import QueryEngineTool, ToolMetadata\n",
"from llama_index.postprocessor.colbert_rerank import ColbertRerank\n",
"\n",
"tools = [\n",
" QueryEngineTool(\n",
" vector_index.as_query_engine(\n",
" similarity_top_k=8,\n",
" node_postprocessors=[ColbertRerank(top_n=3)]\n",
" ),\n",
" metadata=ToolMetadata(\n",
" name=\"search\",\n",
" description=\"Search the document, pass the entire user message in the query\",\n",
" ),\n",
" ),\n",
" QueryEngineTool(\n",
" summary_index.as_query_engine(),\n",
" metadata=ToolMetadata(\n",
" name=\"summarize\",\n",
" description=\"Summarize the document using the user message\",\n",
" ),\n",
" ),\n",
"]\n",
"\n",
"agent = OpenAIAgent.from_tools(\n",
" tools=tools, \n",
" verbose=True\n",
")"
]
},
{
"cell_type": "code",
"execution_count": 8,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Added user message to memory: What is the summary of the paper?\n",
"=== Calling Function ===\n",
"Calling function: summarize with args: {\"input\":\"summary\"}\n",
"Got output: The research focuses on developing Multimodal Large Language Models (MLLMs) by incorporating image-caption, interleaved image-text, and text-only data for pre-training. It highlights the importance of factors like the image encoder, resolution, and token count, while downplaying the design of the vision-language connector. With models scaling up to 30B parameters, the MM1 family demonstrates impressive performance in pre-training metrics and competitive outcomes on diverse multimodal benchmarks. It demonstrates abilities such as in-context learning and multi-image reasoning, aiming to provide valuable insights for creating MLLMs that benefit the research community.\n",
"========================\n",
"\n"
]
}
],
"source": [
"# note -- this will take a while with local LLMs, its sending every node in the document to the LLM\n",
"resp = agent.chat(\"What is the summary of the paper?\")"
]
},
{
"cell_type": "code",
"execution_count": 9,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The summary of the paper highlights the development of Multimodal Large Language Models (MLLMs) by incorporating image-caption, interleaved image-text, and text-only data for pre-training. The research emphasizes factors like the image encoder, resolution, and token count, while de-emphasizing the design of the vision-language connector. The MM1 family of models, scaling up to 30B parameters, shows impressive performance in pre-training metrics and competitive outcomes on various multimodal benchmarks. These models demonstrate capabilities such as in-context learning and multi-image reasoning, aiming to provide valuable insights for creating MLLMs that benefit the research community.\n"
]
}
],
"source": [
"print(str(resp))"
]
},
{
"cell_type": "code",
"execution_count": 10,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Added user message to memory: How do the authors evaluate their work?\n",
"=== Calling Function ===\n",
"Calling function: search with args: {\"input\":\"evaluation methods\"}\n",
"Got output: The evaluation methods involve synthesizing all benchmark results into a single meta-average number to simplify comparisons. This is achieved by normalizing the evaluation metrics with respect to a baseline configuration, standardizing the results for each task, adjusting every metric by dividing it by its respective baseline, and then averaging across all metrics.\n",
"========================\n",
"\n"
]
}
],
"source": [
"resp = agent.chat(\"How do the authors evaluate their work?\")"
]
},
{
"cell_type": "code",
"execution_count": 11,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The authors evaluate their work by synthesizing all benchmark results into a single meta-average number to simplify comparisons. They normalize the evaluation metrics with respect to a baseline configuration, standardize the results for each task, adjust every metric by dividing it by its respective baseline, and then average across all metrics for evaluation.\n"
]
}
],
"source": [
"print(str(resp))"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "llama-parse-aNC435Vv-py3.10",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.12"
},
"orig_nbformat": 4
},
"nbformat": 4,
"nbformat_minor": 2
}
+4 -2
View File
@@ -4,9 +4,11 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"# Llama Parser <> LlamaIndex\n",
"# Advanced RAG with LlamaParse\n",
"\n",
"This notebook is a complete walkthrough for using `LlamaParse` for RAG applications with `LlamaIndex`.\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you how to use LlamaParse with our advanced markdown ingestion and recursive retrieval algorithms to model tables/text within a document hierarchically. This lets you ask questions over both tables and text.\n",
"\n",
"Note for this example, we are using the `llama_index >=0.10.4` version"
]
+2 -3
View File
@@ -182,9 +182,8 @@
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.12"
},
"orig_nbformat": 4
"version": "3.10.10"
}
},
"nbformat": 4,
"nbformat_minor": 4
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+434
View File
@@ -0,0 +1,434 @@
{
"cells": [
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "W6SX9VAnximx"
},
"source": [
"# LlamaParse With MongoDB\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_mongodb.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook, we provide a straightforward example of using LlamaParse with MongoDB Atlas VectorSearch.\n",
"\n",
"We illustrate the process of using llama-parse to parse a PDF document, then index the document with a MongoDB vector store, and subsequently perform basic queries against this store.\n",
"\n",
"This notebook is structured similarly to quick start guides, aiming to introduce users to utilizing llama-parse in conjunction with a MongoDB Atlas VectorSearch."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "rUJKhWDHxr_k"
},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "U6ZkIeBnxfRb"
},
"outputs": [],
"source": [
"!pip install llama-index llama-parse pip install llama-index-vector-stores-mongodb llama-index-llms-openai"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "wh1eeFJe1gkY"
},
"source": [
"### Setup API Keys"
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {
"id": "I5slpdnyxwIB"
},
"outputs": [],
"source": [
"import os\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = '' # Get it from https://cloud.llamaindex.ai/api-key\n",
"os.environ['OPENAI_API_KEY'] = '' # Get it from https://platform.openai.com/api-keys"
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {
"id": "es2mz_OVyQw9"
},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import requests\n",
"import pymongo\n",
"\n",
"from llama_index.vector_stores.mongodb import MongoDBAtlasVectorSearch\n",
"from llama_parse import LlamaParse\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"from llama_index.core.node_parser import SimpleNodeParser"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Ou3bVdHQ10X5"
},
"source": [
"### Download Document\n",
"\n",
"We will use `Attention is all you need` paper."
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "YO9lAk6bybV3",
"outputId": "5cee588a-bec5-482e-e8ef-fbb78e8a5967"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Download complete.\n"
]
}
],
"source": [
"# The URL of the file you want to download\n",
"url = \"https://arxiv.org/pdf/1706.03762.pdf\"\n",
"# The local path where you want to save the file\n",
"file_path = \"./attention.pdf\"\n",
"\n",
"# Perform the HTTP request\n",
"response = requests.get(url)\n",
"\n",
"# Check if the request was successful\n",
"if response.status_code == 200:\n",
" # Open the file in binary write mode and save the content\n",
" with open(file_path, \"wb\") as file:\n",
" file.write(response.content)\n",
" print(\"Download complete.\")\n",
"else:\n",
" print(\"Error downloading the file.\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1NtR7PGo13Hh"
},
"source": [
"### Parse the document using `LlamaParse`."
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "reeJsblfyeSd",
"outputId": "bb569e9f-fe31-47b9-a059-d7da369b3f94"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 09a49745-9f21-4190-9de8-27e4e1a4bdf5\n"
]
}
],
"source": [
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
]
},
{
"cell_type": "code",
"execution_count": 5,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "-NIXtCBwyiPp",
"outputId": "ad4b3cec-2c23-4858-81f0-994ae2c96b8f"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"rmer - model architecture.\n",
"The Transformer follows this overall architecture using stacked self-attention and point-wise, fully\n",
"connected layers for both the encoder and decoder, shown in the left and right halves of Figure 1,\n",
"respectively.\n",
"3.1 Encoder and Decoder Stacks\n",
"Encoder: The encoder is composed of a stack of N = 6 identical layers. Each layer has two\n",
"sub-layers. The first is a multi-head self-attention mechanism, and the second is a simple, position-\n",
"wise fully connected feed-forward network. We employ a residual connection [11] around each of\n",
"the two sub-layers, followed by layer normalization [1]. That is, the output of each sub-layer is\n",
"LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer\n",
"itself. To facilitate these residual connections, all sub-layers in the model, as well as the embedding\n",
"layers, produce outputs of dimension dmodel = 512.\n",
"Decoder: The decoder is also composed of a stack of N = 6 identical layers. In addition \n"
]
}
],
"source": [
"# Take a quick look at some of the parsed text from the document:\n",
"print(documents[0].get_content()[10000:11000])"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "wP9I5dhB1-w1"
},
"source": [
"### Create `MongoDBAtlasVectorSearch`."
]
},
{
"cell_type": "code",
"execution_count": 6,
"metadata": {
"id": "-4Ek0oK-yp3L"
},
"outputs": [],
"source": [
"mongo_uri = os.environ[\"MONGO_URI\"]\n",
"\n",
"mongodb_client = pymongo.MongoClient(mongo_uri)\n",
"mongodb_vector_store = MongoDBAtlasVectorSearch(mongodb_client)\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "GYiVwFok2DNf"
},
"source": [
"### Create nodes."
]
},
{
"cell_type": "code",
"execution_count": 7,
"metadata": {
"id": "aqdF6ZonytHF"
},
"outputs": [],
"source": [
"node_parser = SimpleNodeParser()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "U5fMoGrA2GSH"
},
"source": [
"### Create Index and Query Engine."
]
},
{
"cell_type": "code",
"execution_count": 8,
"metadata": {
"id": "gQUieIrAywSC"
},
"outputs": [],
"source": [
"storage_context = StorageContext.from_defaults(vector_store=mongodb_vector_store)\n",
"\n",
"index = VectorStoreIndex(\n",
" nodes=nodes,\n",
" storage_context=storage_context,\n",
" embed_model=OpenAIEmbedding(),\n",
")"
]
},
{
"cell_type": "code",
"execution_count": 9,
"metadata": {
"id": "snkZZss-zKDb"
},
"outputs": [],
"source": [
"query_engine = index.as_query_engine(similarity_top_k=2)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "rTKT34XO2LYk"
},
"source": [
"### Test Query"
]
},
{
"cell_type": "code",
"execution_count": 10,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "r66ciuPkzNv1",
"outputId": "919218e3-0884-4992-802c-ab1c4622ec4b"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********New LlamaParse+ Basic Query Engine***********\n",
"The BLEU score on the WMT 2014 English-to-German translation task is 28.4.\n"
]
}
],
"source": [
"query = \"What is BLEU score on the WMT 2014 English-to-German translation task?\"\n",
"\n",
"response = query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Basic Query Engine***********\")\n",
"print(response)"
]
},
{
"cell_type": "code",
"execution_count": 11,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "K7RsivpwzQBo",
"outputId": "9bcbf62e-250c-46db-f247-e1f293c09bbe"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"We varied the learning\n",
"rate over the course of training, according to the formula:\n",
" lrate = d0.5 (3)\n",
" model · min(step_num0.5, step_num · warmup_steps1.5)\n",
"This corresponds to increasing the learning rate linearly for the first warmup_steps training steps,\n",
"and decreasing it thereafter proportionally to the inverse square root of the step number. We used\n",
"warmup_steps = 4000.\n",
"5.4 Regularization\n",
"We employ three types of regularization during training:\n",
" 7\n",
"---\n",
"Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the\n",
"English-to-German and English-to-French newstest2014 tests at a fraction of the training cost.\n",
" Model BLEU Training Cost (FLOPs)\n",
" EN-DE EN-FR EN-DE EN-FR\n",
" ByteNet [18] 23.75\n",
" Deep-Att + PosUnk [39] 39.2 1.0 · 1020\n",
" GNMT + RL [38] 24.6 39.92 2.3 · 1019 1.4 · 1020\n",
" ConvS2S [9] 25.16 40.46 9.6 · 1018 1.5 · 1020\n",
" MoE [32] 26.03 40.56 2.0 · 1019 1.2 · 1020\n",
" Deep-Att + PosUnk Ensemble [39] 40.4 8.0 · 1020\n",
" GNMT + RL Ensemble [38] 26.30 41.16 1.8 · 1020 1.1 · 1021\n",
" ConvS2S Ensemble [9] 26.36 41.29 7.7 · 1019 1.2 · 1021\n",
" Transformer (base model) 27.3 38.1 3.3 · 1018\n",
" Transformer (big) 28.4 41.8 2.3 · 1019\n",
"Residual Dropout We apply dropout [33] to the output of each sub-layer, before it is added to the\n",
"sub-layer input and normalized. In addition, we apply dropout to the sums of the embeddings and the\n",
"positional encodings in both the encoder and decoder stacks. For the base model, we use a rate of\n",
"Pdrop = 0.1.\n",
"Label Smoothing During training, we employed label smoothing of value ϵls = 0.1 [36]. This\n",
"hurts perplexity, as the model learns to be more unsure, but improves accuracy and BLEU score.\n",
"6 Results\n",
"6.1 Machine Translation\n",
"On the WMT 2014 English-to-German translation task, the big transformer model (Transformer (big)\n",
"in Table 2) outperforms the best previously reported models (including ensembles) by more than 2.0\n",
"BLEU, establishing a new state-of-the-art BLEU score of 28.4. The configuration of this model is\n",
"listed in the bottom line of Table 3. Training took 3.5 days on 8 P100 GPUs. Even our base model\n",
"surpasses all previously published models and ensembles, at a fraction of the training cost of any of\n",
"the competitive models.\n",
"On the WMT 2014 English-to-French translation task, our big model achieves a BLEU score of 41.0,\n",
"outperforming all of the previously published single models, at less than 1/4 the training cost of the\n",
"previous state-of-the-art model. The Transformer (big) model trained for English-to-French used\n",
"dropout rate Pdrop = 0.1, instead of 0.3.\n",
"For the base models, we used a single model obtained by averaging the last 5 checkpoints, which\n",
"were written at 10-minute intervals. For the big models, we averaged the last 20 checkpoints. We\n",
"used beam search with a beam size of 4 and length penalty α = 0.6 [38]. These hyperparameters\n",
"were chosen after experimentation on the development set. We set the maximum output length during\n",
"inference to input length + 50, but terminate early when possible [38].\n",
"Table 2 summarizes our results and compares our translation quality and training costs to other model\n",
"architectures from the literature.\n"
]
}
],
"source": [
"# Take a look at one of the source nodes from the response\n",
"print(response.source_nodes[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "anthropic_env",
"language": "python",
"name": "anthropic_env"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.3"
},
"vscode": {
"interpreter": {
"hash": "b0fa6594d8f4cbf19f97940f81e996739fb7646882a419484c72d19e05852a7e"
}
}
},
"nbformat": 4,
"nbformat_minor": 0
}
+615
View File
@@ -0,0 +1,615 @@
{
"nbformat": 4,
"nbformat_minor": 0,
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"name": "python3",
"display_name": "Python 3"
},
"language_info": {
"name": "python"
}
},
"cells": [
{
"cell_type": "markdown",
"source": [
"# LlamaParse - Parsing comic books with parsing intructions\n",
"Parsing intructions allow you to instruct our parsing model the same way you would instruct an LLM!\n",
"\n",
"They can be useful to help the parser get better results on complex document layouts, to extract data in a specific format, or to transform the document in other ways.\n",
"\n",
"Using Parsing Instruction you will get better results out of LlamaParse on complicated documents, and also be able to simplify your application code."
],
"metadata": {
"id": "eld1dKaN7P8B"
}
},
{
"cell_type": "markdown",
"source": [
"## Installation\n",
"\n",
"Parsing instructions are part of the llamaParse API. They can be accessed by directly specifying the parsing_instruction parameter in the API or by using the LlamaParse python module (which we will use for this tutorial).\n",
"\n",
"To install llama-parse, just get it from PIP:"
],
"metadata": {
"id": "goB1sV8zu_Xl"
}
},
{
"cell_type": "code",
"source": [
"!pip install llama-parse"
],
"metadata": {
"id": "7Y3_BwQLu-qK",
"colab": {
"base_uri": "https://localhost:8080/"
},
"outputId": "652f8957-ae7a-48e3-86ae-2e9e885c4e1e"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"Collecting llama-parse\n",
" Downloading llama_parse-0.3.8-py3-none-any.whl (6.7 kB)\n",
"Collecting llama-index-core>=0.10.7 (from llama-parse)\n",
" Downloading llama_index_core-0.10.19-py3-none-any.whl (15.3 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m15.3/15.3 MB\u001b[0m \u001b[31m31.9 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: PyYAML>=6.0.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (6.0.1)\n",
"Requirement already satisfied: SQLAlchemy[asyncio]>=1.4.49 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.0.28)\n",
"Requirement already satisfied: aiohttp<4.0.0,>=3.8.6 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.9.3)\n",
"Collecting dataclasses-json (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dataclasses_json-0.6.4-py3-none-any.whl (28 kB)\n",
"Collecting deprecated>=1.2.9.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading Deprecated-1.2.14-py2.py3-none-any.whl (9.6 kB)\n",
"Collecting dirtyjson<2.0.0,>=1.0.8 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading dirtyjson-1.0.8-py3-none-any.whl (25 kB)\n",
"Requirement already satisfied: fsspec>=2023.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2023.6.0)\n",
"Collecting httpx (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpx-0.27.0-py3-none-any.whl (75 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m75.6/75.6 kB\u001b[0m \u001b[31m6.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hCollecting llamaindex-py-client<0.2.0,>=0.1.13 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading llamaindex_py_client-0.1.13-py3-none-any.whl (107 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m108.0/108.0 kB\u001b[0m \u001b[31m10.0 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: nest-asyncio<2.0.0,>=1.5.8 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.6.0)\n",
"Requirement already satisfied: networkx>=3.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.2.1)\n",
"Requirement already satisfied: nltk<4.0.0,>=3.8.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (3.8.1)\n",
"Requirement already satisfied: numpy in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.25.2)\n",
"Collecting openai>=1.1.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading openai-1.13.3-py3-none-any.whl (227 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m227.4/227.4 kB\u001b[0m \u001b[31m16.3 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: pandas in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (1.5.3)\n",
"Requirement already satisfied: pillow>=9.0.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (9.4.0)\n",
"Requirement already satisfied: requests>=2.31.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (2.31.0)\n",
"Requirement already satisfied: tenacity<9.0.0,>=8.2.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (8.2.3)\n",
"Collecting tiktoken>=0.3.3 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading tiktoken-0.6.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.8 MB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m1.8/1.8 MB\u001b[0m \u001b[31m43.1 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: tqdm<5.0.0,>=4.66.1 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.66.2)\n",
"Requirement already satisfied: typing-extensions>=4.5.0 in /usr/local/lib/python3.10/dist-packages (from llama-index-core>=0.10.7->llama-parse) (4.10.0)\n",
"Collecting typing-inspect>=0.8.0 (from llama-index-core>=0.10.7->llama-parse)\n",
" Downloading typing_inspect-0.9.0-py3-none-any.whl (8.8 kB)\n",
"Requirement already satisfied: aiosignal>=1.1.2 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Requirement already satisfied: attrs>=17.3.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (23.2.0)\n",
"Requirement already satisfied: frozenlist>=1.1.1 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.4.1)\n",
"Requirement already satisfied: multidict<7.0,>=4.5 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (6.0.5)\n",
"Requirement already satisfied: yarl<2.0,>=1.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (1.9.4)\n",
"Requirement already satisfied: async-timeout<5.0,>=4.0 in /usr/local/lib/python3.10/dist-packages (from aiohttp<4.0.0,>=3.8.6->llama-index-core>=0.10.7->llama-parse) (4.0.3)\n",
"Requirement already satisfied: wrapt<2,>=1.10 in /usr/local/lib/python3.10/dist-packages (from deprecated>=1.2.9.3->llama-index-core>=0.10.7->llama-parse) (1.14.1)\n",
"Requirement already satisfied: pydantic>=1.10 in /usr/local/lib/python3.10/dist-packages (from llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.6.3)\n",
"Requirement already satisfied: anyio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.7.1)\n",
"Requirement already satisfied: certifi in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (2024.2.2)\n",
"Collecting httpcore==1.* (from httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading httpcore-1.0.4-py3-none-any.whl (77 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m77.8/77.8 kB\u001b[0m \u001b[31m8.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: idna in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (3.6)\n",
"Requirement already satisfied: sniffio in /usr/local/lib/python3.10/dist-packages (from httpx->llama-index-core>=0.10.7->llama-parse) (1.3.1)\n",
"Collecting h11<0.15,>=0.13 (from httpcore==1.*->httpx->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading h11-0.14.0-py3-none-any.whl (58 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m58.3/58.3 kB\u001b[0m \u001b[31m5.7 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: click in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (8.1.7)\n",
"Requirement already satisfied: joblib in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (1.3.2)\n",
"Requirement already satisfied: regex>=2021.8.3 in /usr/local/lib/python3.10/dist-packages (from nltk<4.0.0,>=3.8.1->llama-index-core>=0.10.7->llama-parse) (2023.12.25)\n",
"Requirement already satisfied: distro<2,>=1.7.0 in /usr/lib/python3/dist-packages (from openai>=1.1.0->llama-index-core>=0.10.7->llama-parse) (1.7.0)\n",
"Requirement already satisfied: charset-normalizer<4,>=2 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (3.3.2)\n",
"Requirement already satisfied: urllib3<3,>=1.21.1 in /usr/local/lib/python3.10/dist-packages (from requests>=2.31.0->llama-index-core>=0.10.7->llama-parse) (2.0.7)\n",
"Requirement already satisfied: greenlet!=0.4.17 in /usr/local/lib/python3.10/dist-packages (from SQLAlchemy[asyncio]>=1.4.49->llama-index-core>=0.10.7->llama-parse) (3.0.3)\n",
"Collecting mypy-extensions>=0.3.0 (from typing-inspect>=0.8.0->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading mypy_extensions-1.0.0-py3-none-any.whl (4.7 kB)\n",
"Collecting marshmallow<4.0.0,>=3.18.0 (from dataclasses-json->llama-index-core>=0.10.7->llama-parse)\n",
" Downloading marshmallow-3.21.1-py3-none-any.whl (49 kB)\n",
"\u001b[2K \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m49.4/49.4 kB\u001b[0m \u001b[31m4.5 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
"\u001b[?25hRequirement already satisfied: python-dateutil>=2.8.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2.8.2)\n",
"Requirement already satisfied: pytz>=2020.1 in /usr/local/lib/python3.10/dist-packages (from pandas->llama-index-core>=0.10.7->llama-parse) (2023.4)\n",
"Requirement already satisfied: exceptiongroup in /usr/local/lib/python3.10/dist-packages (from anyio->httpx->llama-index-core>=0.10.7->llama-parse) (1.2.0)\n",
"Requirement already satisfied: packaging>=17.0 in /usr/local/lib/python3.10/dist-packages (from marshmallow<4.0.0,>=3.18.0->dataclasses-json->llama-index-core>=0.10.7->llama-parse) (23.2)\n",
"Requirement already satisfied: annotated-types>=0.4.0 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (0.6.0)\n",
"Requirement already satisfied: pydantic-core==2.16.3 in /usr/local/lib/python3.10/dist-packages (from pydantic>=1.10->llamaindex-py-client<0.2.0,>=0.1.13->llama-index-core>=0.10.7->llama-parse) (2.16.3)\n",
"Requirement already satisfied: six>=1.5 in /usr/local/lib/python3.10/dist-packages (from python-dateutil>=2.8.1->pandas->llama-index-core>=0.10.7->llama-parse) (1.16.0)\n",
"Installing collected packages: dirtyjson, mypy-extensions, marshmallow, h11, deprecated, typing-inspect, tiktoken, httpcore, httpx, dataclasses-json, openai, llamaindex-py-client, llama-index-core, llama-parse\n",
"Successfully installed dataclasses-json-0.6.4 deprecated-1.2.14 dirtyjson-1.0.8 h11-0.14.0 httpcore-1.0.4 httpx-0.27.0 llama-index-core-0.10.19 llama-parse-0.3.8 llamaindex-py-client-0.1.13 marshmallow-3.21.1 mypy-extensions-1.0.0 openai-1.13.3 tiktoken-0.6.0 typing-inspect-0.9.0\n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"## API key\n",
"\n",
"The use of LlamaParse requires an API key which you can get here: https://cloud.llamaindex.ai/parse"
],
"metadata": {
"id": "i-Rg2D_Rvf2i"
}
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "af6i2P1vuU-U"
},
"outputs": [],
"source": [
"import os\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"llx-...\""
]
},
{
"cell_type": "markdown",
"source": [
"## Async (Notebook only)\n",
"llama-parse is async-first, so running the code in a notebook requires the use of nest_asyncio\n"
],
"metadata": {
"id": "p8Eq-aX-wAEo"
}
},
{
"cell_type": "code",
"source": [
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()"
],
"metadata": {
"id": "4OB0BkTqv_0l"
},
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"source": [
"## Import the package"
],
"metadata": {
"id": "dz927ecMyYo_"
}
},
{
"cell_type": "code",
"source": [
"from llama_parse import LlamaParse"
],
"metadata": {
"id": "nSW-6sEwyXwx"
},
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"source": [
"## Using llamaparse for getting better results (on Manga!)\n",
"\n",
"Sometimes the layout of a page is unusual and you will get sub-optimal reading order results with LlamaParse. For example, when parsing manga you expect the reading order to be right to left even if the content is in English!"
],
"metadata": {
"id": "l_D4YsAHwUSk"
}
},
{
"cell_type": "markdown",
"source": [
"Let's download an extract of a great manga \"The manga guide to calculus\", by Hiroyuki Kojima (https://www.amazon.com/Manga-Guide-Calculus-Hiroyuki-Kojima/dp/1593271948)\n",
"\n"
],
"metadata": {
"id": "SV4K2RivxzJG"
}
},
{
"cell_type": "code",
"source": [
"! wget \"https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\" -O ./manga.pdf"
],
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "d3qeuiyawT0U",
"outputId": "e6da0635-dea2-4f2b-ec03-d8db99a75e17"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/uc?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Resolving drive.usercontent.google.com (drive.usercontent.google.com)... 173.194.211.132, 2607:f8b0:400c:c10::84\n",
"Connecting to drive.usercontent.google.com (drive.usercontent.google.com)|173.194.211.132|:443... connected.\n",
"HTTP request sent, awaiting response... 303 See Other\n",
"Location: https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download [following]\n",
"--2024-03-13 13:57:19-- https://drive.usercontent.google.com/download?id=1tZJhcpepLRdQFJFCFX50QIqLyLgqzZsY&export=download\n",
"Reusing existing connection to drive.usercontent.google.com:443.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 3041634 (2.9M) [application/octet-stream]\n",
"Saving to: ./manga.pdf\n",
"\n",
"./manga.pdf 100%[===================>] 2.90M --.-KB/s in 0.04s \n",
"\n",
"2024-03-13 13:57:20 (78.6 MB/s) - ./manga.pdf saved [3041634/3041634]\n",
"\n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"### Without parsing instructions\n",
"For the sake of comparison, let's first parse without any instructions."
],
"metadata": {
"id": "Gbr8RiHEyF3-"
}
},
{
"cell_type": "code",
"source": [
"vanilaParsing = LlamaParse(result_type=\"markdown\").load_data(\"./manga.pdf\")"
],
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "3jKnXCuAyQ9_",
"outputId": "8ab58d56-8c51-44de-8d34-0d94b1bb440f"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"Started parsing the file under job_id 25bf4202-78d8-4705-88cf-c616ae7c82af\n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"As you can see below, LlamaParse is not doing a great job here. It is interpreting the grid of comic panels as a table, and trying to fit the dialogue into a table. It's very hard to follow."
],
"metadata": {
"id": "p4GVOdWzzvYg"
}
},
{
"cell_type": "code",
"source": [
"print(vanilaParsing[0].text[100:1000])"
],
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "ZMhWfKrzzhgQ",
"outputId": "5426c0cc-7e62-4836-9877-87258e1f0b6f"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"\n",
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"---\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = \n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"### Using parsing instructions\n",
"Let's try to parse the manga with custom instructions:\n",
"\n",
"\"The provided document is a manga comic book. Most pages do NOT have title. It does not contain tables. Try to reconstruct the dialogue happening in a cohesive way.\"\n",
"\n",
"To do so just pass the parsing instruction as a parameter to LlamaParse:"
],
"metadata": {
"id": "sUq6znUryiu0"
}
},
{
"cell_type": "code",
"source": [
"parsingInstructionManga = \"\"\"The provided document is a manga comic book, most page do NOT have title.\n",
"It does not contain table.\n",
"Try to reconstruct the dialog happening in a cohesive way.\"\"\"\n",
"withInstructionParsing = LlamaParse(result_type=\"markdown\", parsing_instruction=parsingInstructionManga).load_data(\"./manga.pdf\")"
],
"metadata": {
"id": "dEX7Mv9V0UvM",
"colab": {
"base_uri": "https://localhost:8080/"
},
"outputId": "184b77b7-7e3a-4991-f2c9-eb9105a35a7b"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"Started parsing the file under job_id 88ab273e-b2a7-4f84-8e72-e9367cf6b114\n",
"."
]
}
]
},
{
"cell_type": "markdown",
"source": [
"Let's see how it compare with page 3! We encourage you to play with the target page and explore other pages. As you will see, the parsing instruction allowed LlamaParse to make sense of the document!\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1M87rXTIZE8d5v7aHmVZVW6gW3eDGq6ks&authuser=0\" />\n",
"\n",
"\n",
"\n"
],
"metadata": {
"id": "-UQcA-YW2kjd"
}
},
{
"cell_type": "code",
"source": [
"target_page=1\n",
"print(vanilaParsing[0].text.split('\\n---\\n')[target_page])\n",
"print(\"\\n\\n------------------------------------------------------------\\n\\n\")\n",
"print(withInstructionParsing[0].text.split('\\n---\\n')[target_page])"
],
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "0oPHXg0F0yAS",
"outputId": "38b729bc-b0b7-42f2-97e0-b4e9d7d459a1"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"The Asagake Times Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor? do I have the wrong map?\n",
"\n",
"Youre looking Its next for the Sanda-cho door. branch office? Everybody mistakes us for the office because we are larger. What Is a Function? 3\n",
"\n",
"\n",
"------------------------------------------------------------\n",
"\n",
"\n",
"# The Asagake Times\n",
"\n",
"Sanda-Cho Distributor\n",
"\n",
"A newspaper distributor?\n",
"\n",
"Do I have the wrong map?\n",
"\n",
"You're looking for the Sanda-cho branch office?\n",
"\n",
"It's next door.\n",
"\n",
"Everybody mistakes us for the office because we are larger.\n",
"\n",
"What Is a Function? 3\n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"### Math - doing more with parsing instuction!\n",
"\n",
"But this manga is about math and full of equations, why not ask the parser to output them in **LaTeX**?\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1tze3xcQ7axVA-vC_iZeAj_GvYcyNuYDa&authuser=0\" />"
],
"metadata": {
"id": "yU_jyYWI5fMH"
}
},
{
"cell_type": "code",
"source": [
"parsingInstructionMangaLatex = \"\"\"The provided document is a manga comic book, most page do NOT have title.\n",
"It does not contain table. Do not output table.\n",
"Try to reconstruct the dialog happening in a cohesive way.\n",
"Output any math equation in LATEX markdown (between $$)\"\"\"\n",
"withLatex = LlamaParse(result_type=\"markdown\", parsing_instruction=parsingInstructionMangaLatex).load_data(\"./manga.pdf\")"
],
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "FP_YdO2y5e5o",
"outputId": "45171cdd-298a-4f1f-e67d-6be3cb1a1156"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"Started parsing the file under job_id 3a055e64-d91e-484e-b9b0-99a2e637c08d\n",
"."
]
}
]
},
{
"cell_type": "code",
"source": [
"target_page=2\n",
"print(\"\\n\\n[Without instruction]------------------------------------------------------------\\n\\n\")\n",
"print(vanilaParsing[0].text.split('\\n---\\n')[target_page])\n",
"print(\"\\n\\n[With instruction to output math in LATEX!]------------------------------------------------------------\\n\\n\")\n",
"print(withLatex[0].text.split('\\n---\\n')[target_page])\n"
],
"metadata": {
"id": "TntdRRGp6Rui",
"colab": {
"base_uri": "https://localhost:8080/"
},
"outputId": "e2e503fa-eb87-4f78-83d8-209fe7cebd9e"
},
"execution_count": null,
"outputs": [
{
"output_type": "stream",
"name": "stdout",
"text": [
"\n",
"\n",
"[Without instruction]------------------------------------------------------------\n",
"\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"|1.|Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 (f(a + ε) - f(a)) / ε = lim ε→0 (α - α) = lim ε→0 0 = 0|\n",
"| |Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.|\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"|2.|Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is|\n",
"|---|---|\n",
"| |lim ε→0 (f(α + ε) - f(a)) = lim ε→0 (α(a + ε) + β - (αa + β)) = lim ε→0 α = α|\n",
"| |Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.|\n",
"\n",
"|3.|Lets find the derivative of f(x) = x^2, which appeared in the story. The differential coefficient of f(x) at x = a is|\n",
"|---|---|\n",
"| |lim ε→0 ((a + ε)^2 - a^2) / ε = lim (a^2 + 2aε + ε^2 - a^2) / ε = lim (2aε + ε^2) = lim (2a + ε) = 2a|\n",
"| |Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.|\n",
"\n",
"## Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n",
"\n",
"## Chapter 1 Lets Differentiate a Function!\n",
"\n",
"\n",
"[With instruction to output math in LATEX!]------------------------------------------------------------\n",
"\n",
"\n",
"# Derivative of Constant, Linear, or Quadratic Function\n",
"\n",
"## Calculating the Derivative of a Constant, Linear, or Quadratic Function\n",
"\n",
"1. Lets find the derivative of constant function f(x) = α. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha - \\alpha}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} 0 = 0 \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = 0. This makes sense, since our function is constant—the rate of change is 0.\n",
"\n",
"Note: The differential coefficient of f(x) at x = a is often simply called the derivative of f(x) at x = a, or just f(a).\n",
"\n",
"2. Lets calculate the derivative of linear function f(x) = αx + β. The derivative of f(x) at x = α is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(\\alpha + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\frac{\\alpha(a + \\varepsilon) + \\beta - (\\alpha a + \\beta)}{\\varepsilon} = \\lim_{{\\varepsilon \\to 0}} \\alpha = \\alpha \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the derivative of f(x) is f(x) = α, a constant value. This result should also be intuitive—linear functions have a constant rate of change by definition.\n",
"\n",
"3. Lets find the derivative of f(x) = x2. The differential coefficient of f(x) at x = a is\n",
"\n",
"$$\n",
"\\begin{align*}\n",
"&\\lim_{{\\varepsilon \\to 0}} \\left( \\frac{f(a + \\varepsilon) - f(a)}{\\varepsilon} \\right) = \\lim_{{\\varepsilon \\to 0}} \\left( (a + \\varepsilon)^2 - a^2 \\right) = \\lim_{{\\varepsilon \\to 0}} 2a\\varepsilon + \\varepsilon = \\lim_{{\\varepsilon \\to 0}} (2a + \\varepsilon) = 2a \\\\\n",
"\\end{align*}\n",
"$$\n",
"Thus, the differential coefficient of f(x) at x = a is 2a, or f(a) = 2a. Therefore, the derivative of f(x) is f(x) = 2x.\n",
"\n",
"### Summary\n",
"\n",
"- The calculation of a limit that appears in calculus is simply a formula calculating an error.\n",
"- A limit is used to obtain a derivative.\n",
"- The derivative is the slope of the tangent line at a given point.\n",
"- The derivative is nothing but the rate of change.\n"
]
}
]
},
{
"cell_type": "markdown",
"source": [
"And here is the result as rendered by https://upmath.me/ .\n",
"\n",
"\n",
"<img src=\"https://drive.usercontent.google.com/download?id=1qGo5bMGYOiIC9MnprcgEByaYjU9YII2Q&authuser=0\" />\n",
"\n",
"\n",
"Over this short notebook we saw how to use parsing instructions to increase the quality and accuracy of parsing with LLamaParse!"
],
"metadata": {
"id": "rfFdeWZKmmLW"
}
}
]
}
+42 -6
View File
@@ -110,17 +110,53 @@ class Language(str, Enum):
SUPPORTED_FILE_TYPES = [
".pdf",
".xml"
".602",
".abw",
".cgm",
".cwk",
".doc",
".docx",
".pptx",
".rtf",
".pages",
".docm",
".dot",
".dotm",
".hwp",
".key",
".epub"
".lwp",
".mw",
".mcw",
".pages",
".pbd",
".ppt",
".pptm",
".pptx",
".pot",
".potm",
".potx",
".rtf",
".sda",
".sdd",
".sdp",
".sdw",
".sgl",
".sti",
".sxi",
".sxw",
".stw",
".sxg",
".txt",
".uof",
".uop",
".uot",
".vor",
".wpd",
".wps",
".xml",
".zabw",
".epub",
".htm",
".html"
]
class LlamaParse(BasePydanticReader):
"""A smart-parser for files."""