Compare commits

..

10 Commits

Author SHA1 Message Date
Pierre-Loic Doulcet 2e6c064682 add more supported format 2024-04-23 11:18:38 +08:00
Jerry Liu 4252f6186b fixes to insurance demo (#97) 2024-03-21 00:02:47 -07:00
Jerry Liu 22148ade9f fix advanced RAG notebook title (#98) 2024-03-21 00:02:38 -07:00
Jerry Liu b8332fe8e1 nit: add colab badge to mongodb notebook (#109) 2024-03-21 00:02:29 -07:00
Ravi Theja e40e92a133 Add mongodb llamaparse example (#107) 2024-03-20 23:37:42 -07:00
Jerry Liu ba8f345f80 Revert "cr"
This reverts commit 2ddbf1ba0d.
2024-03-19 00:21:29 -07:00
Jerry Liu 2ddbf1ba0d cr 2024-03-19 00:20:42 -07:00
Haotian Zhang 23567c8f98 Init LlamaParseJsonNodeParser example (#93) 2024-03-18 15:15:36 -04:00
Logan 8d39ae7763 add agent demo (#88)
* add agent demo

* remove mention of react agent

* agents folder
2024-03-18 16:12:56 +01:00
Jerry Liu a2edc41fc7 nit: fix grammar in insurance cookbook (#89)
cr
2024-03-18 16:11:03 +01:00
6 changed files with 995 additions and 156 deletions
+4 -2
View File
@@ -4,9 +4,11 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"# Llama Parser <> LlamaIndex\n",
"# Advanced RAG with LlamaParse\n",
"\n",
"This notebook is a complete walkthrough for using `LlamaParse` for RAG applications with `LlamaIndex`.\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you how to use LlamaParse with our advanced markdown ingestion and recursive retrieval algorithms to model tables/text within a document hierarchically. This lets you ask questions over both tables and text.\n",
"\n",
"Note for this example, we are using the `llama_index >=0.10.4` version"
]
+2 -3
View File
@@ -182,9 +182,8 @@
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.12"
},
"orig_nbformat": 4
"version": "3.10.10"
}
},
"nbformat": 4,
"nbformat_minor": 4
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+434
View File
@@ -0,0 +1,434 @@
{
"cells": [
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "W6SX9VAnximx"
},
"source": [
"# LlamaParse With MongoDB\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_mongodb.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook, we provide a straightforward example of using LlamaParse with MongoDB Atlas VectorSearch.\n",
"\n",
"We illustrate the process of using llama-parse to parse a PDF document, then index the document with a MongoDB vector store, and subsequently perform basic queries against this store.\n",
"\n",
"This notebook is structured similarly to quick start guides, aiming to introduce users to utilizing llama-parse in conjunction with a MongoDB Atlas VectorSearch."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "rUJKhWDHxr_k"
},
"source": [
"### Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "U6ZkIeBnxfRb"
},
"outputs": [],
"source": [
"!pip install llama-index llama-parse pip install llama-index-vector-stores-mongodb llama-index-llms-openai"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "wh1eeFJe1gkY"
},
"source": [
"### Setup API Keys"
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {
"id": "I5slpdnyxwIB"
},
"outputs": [],
"source": [
"import os\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = '' # Get it from https://cloud.llamaindex.ai/api-key\n",
"os.environ['OPENAI_API_KEY'] = '' # Get it from https://platform.openai.com/api-keys"
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {
"id": "es2mz_OVyQw9"
},
"outputs": [],
"source": [
"# llama-parse is async-first, running the sync code in a notebook requires the use of nest_asyncio\n",
"import nest_asyncio\n",
"\n",
"nest_asyncio.apply()\n",
"\n",
"import requests\n",
"import pymongo\n",
"\n",
"from llama_index.vector_stores.mongodb import MongoDBAtlasVectorSearch\n",
"from llama_parse import LlamaParse\n",
"from llama_index.embeddings.openai import OpenAIEmbedding\n",
"from llama_index.core import VectorStoreIndex, StorageContext\n",
"from llama_index.core.node_parser import SimpleNodeParser"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Ou3bVdHQ10X5"
},
"source": [
"### Download Document\n",
"\n",
"We will use `Attention is all you need` paper."
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "YO9lAk6bybV3",
"outputId": "5cee588a-bec5-482e-e8ef-fbb78e8a5967"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Download complete.\n"
]
}
],
"source": [
"# The URL of the file you want to download\n",
"url = \"https://arxiv.org/pdf/1706.03762.pdf\"\n",
"# The local path where you want to save the file\n",
"file_path = \"./attention.pdf\"\n",
"\n",
"# Perform the HTTP request\n",
"response = requests.get(url)\n",
"\n",
"# Check if the request was successful\n",
"if response.status_code == 200:\n",
" # Open the file in binary write mode and save the content\n",
" with open(file_path, \"wb\") as file:\n",
" file.write(response.content)\n",
" print(\"Download complete.\")\n",
"else:\n",
" print(\"Error downloading the file.\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1NtR7PGo13Hh"
},
"source": [
"### Parse the document using `LlamaParse`."
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "reeJsblfyeSd",
"outputId": "bb569e9f-fe31-47b9-a059-d7da369b3f94"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Started parsing the file under job_id 09a49745-9f21-4190-9de8-27e4e1a4bdf5\n"
]
}
],
"source": [
"documents = LlamaParse(result_type=\"text\").load_data(file_path)"
]
},
{
"cell_type": "code",
"execution_count": 5,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "-NIXtCBwyiPp",
"outputId": "ad4b3cec-2c23-4858-81f0-994ae2c96b8f"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"rmer - model architecture.\n",
"The Transformer follows this overall architecture using stacked self-attention and point-wise, fully\n",
"connected layers for both the encoder and decoder, shown in the left and right halves of Figure 1,\n",
"respectively.\n",
"3.1 Encoder and Decoder Stacks\n",
"Encoder: The encoder is composed of a stack of N = 6 identical layers. Each layer has two\n",
"sub-layers. The first is a multi-head self-attention mechanism, and the second is a simple, position-\n",
"wise fully connected feed-forward network. We employ a residual connection [11] around each of\n",
"the two sub-layers, followed by layer normalization [1]. That is, the output of each sub-layer is\n",
"LayerNorm(x + Sublayer(x)), where Sublayer(x) is the function implemented by the sub-layer\n",
"itself. To facilitate these residual connections, all sub-layers in the model, as well as the embedding\n",
"layers, produce outputs of dimension dmodel = 512.\n",
"Decoder: The decoder is also composed of a stack of N = 6 identical layers. In addition \n"
]
}
],
"source": [
"# Take a quick look at some of the parsed text from the document:\n",
"print(documents[0].get_content()[10000:11000])"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "wP9I5dhB1-w1"
},
"source": [
"### Create `MongoDBAtlasVectorSearch`."
]
},
{
"cell_type": "code",
"execution_count": 6,
"metadata": {
"id": "-4Ek0oK-yp3L"
},
"outputs": [],
"source": [
"mongo_uri = os.environ[\"MONGO_URI\"]\n",
"\n",
"mongodb_client = pymongo.MongoClient(mongo_uri)\n",
"mongodb_vector_store = MongoDBAtlasVectorSearch(mongodb_client)\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "GYiVwFok2DNf"
},
"source": [
"### Create nodes."
]
},
{
"cell_type": "code",
"execution_count": 7,
"metadata": {
"id": "aqdF6ZonytHF"
},
"outputs": [],
"source": [
"node_parser = SimpleNodeParser()\n",
"\n",
"nodes = node_parser.get_nodes_from_documents(documents)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "U5fMoGrA2GSH"
},
"source": [
"### Create Index and Query Engine."
]
},
{
"cell_type": "code",
"execution_count": 8,
"metadata": {
"id": "gQUieIrAywSC"
},
"outputs": [],
"source": [
"storage_context = StorageContext.from_defaults(vector_store=mongodb_vector_store)\n",
"\n",
"index = VectorStoreIndex(\n",
" nodes=nodes,\n",
" storage_context=storage_context,\n",
" embed_model=OpenAIEmbedding(),\n",
")"
]
},
{
"cell_type": "code",
"execution_count": 9,
"metadata": {
"id": "snkZZss-zKDb"
},
"outputs": [],
"source": [
"query_engine = index.as_query_engine(similarity_top_k=2)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "rTKT34XO2LYk"
},
"source": [
"### Test Query"
]
},
{
"cell_type": "code",
"execution_count": 10,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "r66ciuPkzNv1",
"outputId": "919218e3-0884-4992-802c-ab1c4622ec4b"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"***********New LlamaParse+ Basic Query Engine***********\n",
"The BLEU score on the WMT 2014 English-to-German translation task is 28.4.\n"
]
}
],
"source": [
"query = \"What is BLEU score on the WMT 2014 English-to-German translation task?\"\n",
"\n",
"response = query_engine.query(query)\n",
"print(\"\\n***********New LlamaParse+ Basic Query Engine***********\")\n",
"print(response)"
]
},
{
"cell_type": "code",
"execution_count": 11,
"metadata": {
"colab": {
"base_uri": "https://localhost:8080/"
},
"id": "K7RsivpwzQBo",
"outputId": "9bcbf62e-250c-46db-f247-e1f293c09bbe"
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"We varied the learning\n",
"rate over the course of training, according to the formula:\n",
" lrate = d0.5 (3)\n",
" model · min(step_num0.5, step_num · warmup_steps1.5)\n",
"This corresponds to increasing the learning rate linearly for the first warmup_steps training steps,\n",
"and decreasing it thereafter proportionally to the inverse square root of the step number. We used\n",
"warmup_steps = 4000.\n",
"5.4 Regularization\n",
"We employ three types of regularization during training:\n",
" 7\n",
"---\n",
"Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the\n",
"English-to-German and English-to-French newstest2014 tests at a fraction of the training cost.\n",
" Model BLEU Training Cost (FLOPs)\n",
" EN-DE EN-FR EN-DE EN-FR\n",
" ByteNet [18] 23.75\n",
" Deep-Att + PosUnk [39] 39.2 1.0 · 1020\n",
" GNMT + RL [38] 24.6 39.92 2.3 · 1019 1.4 · 1020\n",
" ConvS2S [9] 25.16 40.46 9.6 · 1018 1.5 · 1020\n",
" MoE [32] 26.03 40.56 2.0 · 1019 1.2 · 1020\n",
" Deep-Att + PosUnk Ensemble [39] 40.4 8.0 · 1020\n",
" GNMT + RL Ensemble [38] 26.30 41.16 1.8 · 1020 1.1 · 1021\n",
" ConvS2S Ensemble [9] 26.36 41.29 7.7 · 1019 1.2 · 1021\n",
" Transformer (base model) 27.3 38.1 3.3 · 1018\n",
" Transformer (big) 28.4 41.8 2.3 · 1019\n",
"Residual Dropout We apply dropout [33] to the output of each sub-layer, before it is added to the\n",
"sub-layer input and normalized. In addition, we apply dropout to the sums of the embeddings and the\n",
"positional encodings in both the encoder and decoder stacks. For the base model, we use a rate of\n",
"Pdrop = 0.1.\n",
"Label Smoothing During training, we employed label smoothing of value ϵls = 0.1 [36]. This\n",
"hurts perplexity, as the model learns to be more unsure, but improves accuracy and BLEU score.\n",
"6 Results\n",
"6.1 Machine Translation\n",
"On the WMT 2014 English-to-German translation task, the big transformer model (Transformer (big)\n",
"in Table 2) outperforms the best previously reported models (including ensembles) by more than 2.0\n",
"BLEU, establishing a new state-of-the-art BLEU score of 28.4. The configuration of this model is\n",
"listed in the bottom line of Table 3. Training took 3.5 days on 8 P100 GPUs. Even our base model\n",
"surpasses all previously published models and ensembles, at a fraction of the training cost of any of\n",
"the competitive models.\n",
"On the WMT 2014 English-to-French translation task, our big model achieves a BLEU score of 41.0,\n",
"outperforming all of the previously published single models, at less than 1/4 the training cost of the\n",
"previous state-of-the-art model. The Transformer (big) model trained for English-to-French used\n",
"dropout rate Pdrop = 0.1, instead of 0.3.\n",
"For the base models, we used a single model obtained by averaging the last 5 checkpoints, which\n",
"were written at 10-minute intervals. For the big models, we averaged the last 20 checkpoints. We\n",
"used beam search with a beam size of 4 and length penalty α = 0.6 [38]. These hyperparameters\n",
"were chosen after experimentation on the development set. We set the maximum output length during\n",
"inference to input length + 50, but terminate early when possible [38].\n",
"Table 2 summarizes our results and compares our translation quality and training costs to other model\n",
"architectures from the literature.\n"
]
}
],
"source": [
"# Take a look at one of the source nodes from the response\n",
"print(response.source_nodes[0].get_content())"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"colab": {
"provenance": []
},
"kernelspec": {
"display_name": "anthropic_env",
"language": "python",
"name": "anthropic_env"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.3"
},
"vscode": {
"interpreter": {
"hash": "b0fa6594d8f4cbf19f97940f81e996739fb7646882a419484c72d19e05852a7e"
}
}
},
"nbformat": 4,
"nbformat_minor": 0
}
+42 -6
View File
@@ -110,17 +110,53 @@ class Language(str, Enum):
SUPPORTED_FILE_TYPES = [
".pdf",
".xml"
".602",
".abw",
".cgm",
".cwk",
".doc",
".docx",
".pptx",
".rtf",
".pages",
".docm",
".dot",
".dotm",
".hwp",
".key",
".epub"
".lwp",
".mw",
".mcw",
".pages",
".pbd",
".ppt",
".pptm",
".pptx",
".pot",
".potm",
".potx",
".rtf",
".sda",
".sdd",
".sdp",
".sdw",
".sgl",
".sti",
".sxi",
".sxw",
".stw",
".sxg",
".txt",
".uof",
".uop",
".uot",
".vor",
".wpd",
".wps",
".xml",
".zabw",
".epub",
".htm",
".html"
]
class LlamaParse(BasePydanticReader):
"""A smart-parser for files."""