Compare commits

..

9 Commits

Author SHA1 Message Date
Emanuel Ferreira d71fa8d530 docs: parse -> classify -> extract 2025-09-22 19:02:21 -03:00
Adrian Lyjak 9d9b816644 Handle reasoning field conflict (#929)
* Handle reasoning field conflict

* update version to 0.6.69
2025-09-22 11:29:11 -04:00
Adrian Lyjak 83555f76e6 Handle validation errors for agent data retrieval (#928)
* feat: Add untyped agent data retrieval and handling

Introduces methods to retrieve agent data as untyped dictionaries,
handling validation errors gracefully. This allows for more flexible
data access when strict typing is not required or when data may be
malformed.

Co-authored-by: adrian <adrian@runllama.ai>

* Expose raw api result

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-09-22 11:28:49 -04:00
Adrian Lyjak 5edf5f914a Support creating indexes in a specified project_id (#924)
* Support creating indexes in a specified project_id

* Bump
2025-09-18 11:07:07 -04:00
Adrian Lyjak 22e4975cb2 Refactor agent fields in llama_cloud_services (#921) 2025-09-17 15:14:40 -04:00
Peter Rowlands (변기호) bc2f04379b py: bump version to v.0.6.66 (#920) 2025-09-16 19:34:18 +09:00
Peter Rowlands (변기호) f9f951d5d8 parse: expose spreadsheet_force_formula_computation option (#919) 2025-09-16 19:28:03 +09:00
Emmanuel Ferdman 355129fea5 Fix colab broken links (#750)
Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>
2025-09-14 23:10:21 +02:00
Adrian Lyjak d9aed80ded fix: v prefix goes deeper. Fix more (#899) 2025-09-08 17:45:06 -04:00
34 changed files with 8331 additions and 3846 deletions
+1 -1
View File
@@ -29,7 +29,7 @@ repos:
- id: black-jupyter
name: black-src
alias: black
exclude: ".*uv.lock"
exclude: ".*uv.lock|examples/extract/solar_panel_e2e_comparison.ipynb"
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.0.1
hooks:
@@ -7,7 +7,7 @@
"source": [
"# Extraction and Analysis over a Fidelity Multi-Fund Annual Report\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/extract/asset_manager_fund_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/asset_manager_fund_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook we show you how to create an agentic document workflow over a complex document that contains annual reports for multiple funds - each fund reports financials in a standardized reporting structure, and it's all consolidated in the same document.\n",
"\n",
@@ -7,7 +7,7 @@
"source": [
"# Automotive Equity Research: A Multi-Step Agentic Workflow\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/extract/automotive_sector_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/extract/automotive_sector_analysis.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates an endtoend agentic workflow using LlamaExtract and the LlamaIndex eventdriven workflow framework for automotive sector analysis.\n",
"\n",
+2 -2
View File
@@ -1035,7 +1035,7 @@
],
"metadata": {
"kernelspec": {
"display_name": ".venv",
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
@@ -1052,5 +1052,5 @@
}
},
"nbformat": 4,
"nbformat_minor": 2
"nbformat_minor": 4
}
@@ -0,0 +1,765 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Complete Parse → Classify → Extract Workflow with LlamaCloud Services\n",
"\n",
"This notebook demonstrates the complete workflow for processing documents using LlamaCloud services:\n",
"1. **Parse** - Extract and convert documents to markdown\n",
"2. **Classify** - Categorize documents based on their content\n",
"3. **Extract** - Extract structured data using the markdown as input via SourceText\n",
"\n",
"## Overview of the Workflow\n",
"\n",
"### 1. Parse Phase\n",
"- Use `LlamaParse` to convert documents (PDFs, Word docs, etc.) into structured formats\n",
"- Extract markdown content that preserves document structure\n",
"- Get both raw text and markdown representations\n",
"\n",
"### 2. Classify Phase\n",
"- Use `ClassifyClient` to categorize documents based on content\n",
"- Apply classification rules to route documents appropriately\n",
"- Handle different document types with specific processing logic\n",
"\n",
"### 3. Extract Phase\n",
"- Use `LlamaExtract` with `SourceText` to extract structured data\n",
"- Pass the markdown content as input for more accurate extraction\n",
"- Define custom schemas for structured data extraction\n",
"\n",
"Let's walk through each step with practical examples."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Setup and Installation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Install required packages\n",
"!pip install llama-cloud-services\n",
"!pip install python-dotenv"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"✅ API key configured\n"
]
}
],
"source": [
"import os\n",
"import nest_asyncio\n",
"from getpass import getpass\n",
"from dotenv import load_dotenv\n",
"\n",
"# Load environment variables\n",
"load_dotenv()\n",
"nest_asyncio.apply()\n",
"\n",
"# Set up API key\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"\" # edit it\n",
"\n",
"# Setup Base URL\n",
"# os.envrion[\"LLAMA_CLOUD_BASE_URL\"] = \"https://api.cloud.eu.llamaindex.ai/\" # update if necessay\n",
"\n",
"print(\"✅ API key configured\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Download Sample Documents\n",
"\n",
"Let's download some sample documents to work with:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📁 financial_report.pdf already exists\n",
"📁 technical_spec.pdf already exists\n",
"\n",
"📂 Sample documents ready!\n"
]
}
],
"source": [
"import requests\n",
"import os\n",
"\n",
"# Create directory for sample documents\n",
"os.makedirs(\"sample_docs\", exist_ok=True)\n",
"\n",
"# Download sample documents\n",
"docs_to_download = {\n",
" \"financial_report.pdf\": \"https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf\",\n",
" \"technical_spec.pdf\": \"https://www.ti.com/lit/ds/symlink/lm317.pdf\",\n",
"}\n",
"\n",
"for filename, url in docs_to_download.items():\n",
" filepath = f\"sample_docs/{filename}\"\n",
" if not os.path.exists(filepath):\n",
" print(f\"Downloading {filename}...\")\n",
" response = requests.get(url)\n",
" if response.status_code == 200:\n",
" with open(filepath, \"wb\") as f:\n",
" f.write(response.content)\n",
" print(f\"✅ Downloaded {filename}\")\n",
" else:\n",
" print(f\"❌ Failed to download {filename}\")\n",
" else:\n",
" print(f\"📁 {filename} already exists\")\n",
"\n",
"print(\"\\n📂 Sample documents ready!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Phase 1: Document Parsing\n",
"\n",
"First, let's parse our documents using LlamaParse to extract clean markdown content."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🔄 Parsing documents...\n",
"Started parsing the file under job_id 8a8c76f9-354d-4275-91d8-312ff1adc762\n",
"...✅ Parsed financial report (Job ID: 8a8c76f9-354d-4275-91d8-312ff1adc762)\n",
"Started parsing the file under job_id 7e603448-ed80-4d18-948b-6801ed51c41b\n",
"✅ Parsed technical spec (Job ID: 7e603448-ed80-4d18-948b-6801ed51c41b)\n",
"\n",
"📄 Parsing complete!\n"
]
}
],
"source": [
"from llama_cloud_services.parse.base import LlamaParse\n",
"from llama_cloud_services.parse.utils import ResultType\n",
"import asyncio\n",
"\n",
"# Initialize the parser\n",
"parser = LlamaParse(\n",
" result_type=ResultType.MD, # Get markdown output\n",
" verbose=True,\n",
" language=\"en\",\n",
" # Premium mode for better accuracy\n",
" premium_mode=True,\n",
" # Extract tables as HTML for better structure\n",
" output_tables_as_HTML=True,\n",
" # Parse only first few pages for demo\n",
")\n",
"\n",
"print(\"🔄 Parsing documents...\")\n",
"\n",
"# Parse the financial report\n",
"financial_result = await parser.aparse(\"sample_docs/financial_report.pdf\")\n",
"print(f\"✅ Parsed financial report (Job ID: {financial_result.job_id})\")\n",
"\n",
"# Parse the technical specification\n",
"technical_result = await parser.aparse(\"sample_docs/technical_spec.pdf\")\n",
"print(f\"✅ Parsed technical spec (Job ID: {technical_result.job_id})\")\n",
"\n",
"print(\"\\n📄 Parsing complete!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Extract Markdown Content\n",
"\n",
"Now let's get the markdown content from our parsed documents:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📋 Financial Report Markdown (first 500 chars):\n",
"\n",
"\n",
"# UNITED STATES\n",
"# SECURITIES AND EXCHANGE COMMISSION\n",
"Washington, D.C. 20549\n",
"\n",
"## FORM 10-K\n",
"\n",
"(Mark One)\n",
"\n",
"☒ ANNUAL REPORT PURSUANT TO SECTION 13 OR 15(d) OF THE SECURITIES EXCHANGE ACT OF 1934\n",
"For the fiscal year ended December 31, 2021\n",
"OR\n",
"☐ TRANSITION REPORT PURSUANT TO SECTION 13 OR 15(d) OF THE SECURITIES EXCHANGE ACT OF 1934\n",
"For the transition period from_____ to _____\n",
"Commission File Number: 001-38902\n",
"\n",
"# UBER TECHNOLOGIES, INC.\n",
"(Exact name of registrant as specified in its charter)\n",
"\n",
"Delaware\n",
"...\n",
"\n",
"📋 Technical Spec Markdown (first 500 chars):\n",
"\n",
"\n",
"LM317\n",
"SLVS044Z SEPTEMBER 1997 REVISED APRIL 2025\n",
"\n",
"# LM317 3-Pin Adjustable Regulator\n",
"\n",
"## 1 Features\n",
"\n",
"• Output voltage range:\n",
" Adjustable: 1.25V to 37V\n",
"• Output current: 1.5A\n",
"• Line regulation: 0.01%/V (typ)\n",
"• Load regulation: 0.1% (typ)\n",
"• Internal short-circuit current limiting\n",
"• Thermal overload protection\n",
"• Output safe-area compensation (new chip)\n",
"• PSRR: 80dB at 120Hz for CADJ = 10μF (new chip)\n",
"• Packages:\n",
" 4-pin, SOT-223 (DCY)\n",
" 3-pin, TO-263 (KTT)\n",
" 3-pin, TO-220 (KCS, KCT),\n",
"...\n",
"\n",
"📏 Financial report markdown length: 1348671 characters\n",
"📏 Technical spec markdown length: 90971 characters\n"
]
}
],
"source": [
"# Get markdown content from parsed documents\n",
"financial_markdown = await financial_result.aget_markdown()\n",
"technical_markdown = await technical_result.aget_markdown()\n",
"\n",
"print(\"📋 Financial Report Markdown (first 500 chars):\")\n",
"print(financial_markdown[:500])\n",
"print(\"...\\n\")\n",
"\n",
"print(\"📋 Technical Spec Markdown (first 500 chars):\")\n",
"print(technical_markdown[:500])\n",
"print(\"...\\n\")\n",
"\n",
"print(f\"📏 Financial report markdown length: {len(financial_markdown)} characters\")\n",
"print(f\"📏 Technical spec markdown length: {len(technical_markdown)} characters\")\n",
"\n",
"document_texts = [financial_markdown, technical_markdown]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Phase 2: Document Classification\n",
"\n",
"Next, let's classify our documents based on their content using the ClassifyClient."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🏷️ Setting up document classification...\n",
"📝 Created 3 classification rules\n"
]
}
],
"source": [
"from llama_cloud_services.beta.classifier.client import ClassifyClient\n",
"from llama_cloud.types import ClassifierRule\n",
"from llama_cloud_services.files.client import FileClient\n",
"from llama_cloud.client import AsyncLlamaCloud\n",
"\n",
"# Initialize the classify client\n",
"api_key = os.environ[\"LLAMA_CLOUD_API_KEY\"]\n",
"classify_client = ClassifyClient.from_api_key(api_key)\n",
"\n",
"print(\"🏷️ Setting up document classification...\")\n",
"\n",
"# Define classification rules\n",
"classification_rules = [\n",
" ClassifierRule(\n",
" type=\"financial_document\",\n",
" description=\"Documents containing financial data, revenue, expenses, SEC filings, or financial statements\",\n",
" ),\n",
" ClassifierRule(\n",
" type=\"technical_specification\",\n",
" description=\"Technical datasheets, component specifications, engineering documents, or technical manuals\",\n",
" ),\n",
" ClassifierRule(\n",
" type=\"general_document\",\n",
" description=\"General business documents, contracts, or other unspecified document types\",\n",
" ),\n",
"]\n",
"\n",
"print(f\"📝 Created {len(classification_rules)} classification rules\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Phase 3: Structured Data Extraction using SourceText\n",
"\n",
"Now comes the key part - using the markdown content as input for structured data extraction via SourceText."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"⚙️ LlamaExtract initialized\n"
]
}
],
"source": [
"from llama_cloud_services.extract.extract import LlamaExtract, SourceText\n",
"from llama_cloud.types import ExtractConfig, ExtractMode\n",
"from pydantic import BaseModel, Field\n",
"from typing import List, Optional\n",
"\n",
"# Initialize LlamaExtract\n",
"llama_extract = LlamaExtract(api_key=api_key, verbose=True)\n",
"\n",
"print(\"⚙️ LlamaExtract initialized\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Define Extraction Schemas\n",
"\n",
"Let's define different schemas for different document types:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📋 Extraction schemas defined\n"
]
}
],
"source": [
"# Schema for financial documents\n",
"class FinancialMetrics(BaseModel):\n",
" company_name: str = Field(description=\"Name of the company\")\n",
" document_type: str = Field(\n",
" description=\"Type of financial document (10-K, 10-Q, annual report, etc.)\"\n",
" )\n",
" fiscal_year: int = Field(description=\"Fiscal year of the report\")\n",
" revenue_2021: str = Field(description=\"Total revenue in 2021\")\n",
" net_income_2021: str = Field(description=\"Net income in 2021\")\n",
" key_business_segments: List[str] = Field(\n",
" default=[], description=\"Main business segments or divisions\"\n",
" )\n",
" risk_factors: List[str] = Field(\n",
" default=[], description=\"Key risk factors mentioned\"\n",
" )\n",
"\n",
"\n",
"# Schema for technical specifications\n",
"class VoltageRange(BaseModel):\n",
" min_voltage: Optional[float] = Field(description=\"Minimum voltage\")\n",
" max_voltage: Optional[float] = Field(description=\"Maximum voltage\")\n",
" unit: str = Field(default=\"V\", description=\"Voltage unit\")\n",
"\n",
"\n",
"class TechnicalSpec(BaseModel):\n",
" component_name: str = Field(description=\"Name of the technical component\")\n",
" manufacturer: Optional[str] = Field(description=\"Manufacturer name\")\n",
" part_number: Optional[str] = Field(description=\"Part or model number\")\n",
" description: str = Field(description=\"Brief description of the component\")\n",
" operating_voltage: Optional[VoltageRange] = Field(\n",
" description=\"Operating voltage range\"\n",
" )\n",
" maximum_current: Optional[float] = Field(\n",
" description=\"Maximum current rating in amperes\"\n",
" )\n",
" key_features: List[str] = Field(\n",
" default=[], description=\"Key features and capabilities\"\n",
" )\n",
" applications: List[str] = Field(default=[], description=\"Typical applications\")\n",
"\n",
"\n",
"print(\"📋 Extraction schemas defined\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Complete Workflow Summary\n",
"\n",
"Let's create a function that demonstrates the complete workflow:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🔧 Workflow function defined!\n"
]
}
],
"source": [
"import tempfile\n",
"from pathlib import Path\n",
"from llama_cloud import ExtractConfig\n",
"\n",
"\n",
"async def complete_document_workflow(markdown_content: str):\n",
" \"\"\"\n",
" Complete workflow: Parse → Classify → Extract\n",
" \"\"\"\n",
" print(f\"🚀 Starting complete workflow\")\n",
" print(\"=\" * 60)\n",
"\n",
" # Step 1: Classify\n",
" print(\"🏷️ Step 2: Classifying document...\")\n",
"\n",
" with tempfile.NamedTemporaryFile(\n",
" mode=\"w\", suffix=\".md\", delete=False, encoding=\"utf-8\"\n",
" ) as tmp:\n",
" tmp.write(markdown_content)\n",
" temp_path = Path(tmp.name)\n",
"\n",
" print(temp_path)\n",
"\n",
" classification = await classify_client.aclassify_file_path(\n",
" rules=classification_rules, file_input_path=str(temp_path)\n",
" )\n",
" doc_type = classification.items[0].result.type\n",
" confidence = classification.items[0].result.confidence\n",
" print(f\" ✅ Classified as: {doc_type} (confidence: {confidence:.2f})\")\n",
"\n",
" # Step 2: Extract based on classification\n",
" print(\"🔍 Step 3: Extracting structured data using SourceText...\")\n",
" source_text = SourceText(\n",
" text_content=markdown_content,\n",
" filename=f\"{os.path.basename(temp_path)}_markdown.md\",\n",
" )\n",
"\n",
" # Choose schema based on classification\n",
" if \"financial\" in doc_type.lower():\n",
" schema = FinancialMetrics\n",
" print(\" 📊 Using FinancialMetrics schema\")\n",
" elif \"technical\" in doc_type.lower():\n",
" schema = TechnicalSpec\n",
" print(\" 🔧 Using TechnicalSpec schema\")\n",
" else:\n",
" schema = FinancialMetrics # Default fallback\n",
" print(\" 📊 Using default FinancialMetrics schema\")\n",
"\n",
" extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
" )\n",
"\n",
" extraction_result = llama_extract.extract(\n",
" data_schema=schema, config=extract_config, files=source_text\n",
" )\n",
"\n",
" print(\" ✅ Extraction complete!\")\n",
"\n",
" return {\n",
" \"file_path\": temp_path,\n",
" \"markdown_length\": len(markdown_content),\n",
" \"classification\": doc_type,\n",
" \"confidence\": confidence,\n",
" \"extracted_data\": extraction_result.data,\n",
" \"markdown_sample\": markdown_content[:200] + \"...\"\n",
" if len(markdown_content) > 200\n",
" else markdown_content,\n",
" }\n",
"\n",
"\n",
"print(\"🔧 Workflow function defined!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Run Complete Workflow on Both Documents"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🚀 Starting complete workflow\n",
"============================================================\n",
"🏷️ Step 2: Classifying document...\n",
"/var/folders/g6/4b5lpp5974gcpr890ybhbw4r0000gn/T/tmpos3b62tm.md\n",
" ✅ Classified as: financial_document (confidence: 1.00)\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 📊 Using FinancialMetrics schema\n",
".. ✅ Extraction complete!\n",
"\n",
"============================================================\n",
"\n",
"🚀 Starting complete workflow\n",
"============================================================\n",
"🏷️ Step 2: Classifying document...\n",
"/var/folders/g6/4b5lpp5974gcpr890ybhbw4r0000gn/T/tmpppz9ub_m.md\n",
" ✅ Classified as: technical_specification (confidence: 1.00)\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 🔧 Using TechnicalSpec schema\n",
" ✅ Extraction complete!\n",
"\n",
"============================================================\n",
"\n",
"📋 Processed 2 documents successfully!\n"
]
}
],
"source": [
"# Process both documents through the complete workflow\n",
"results = []\n",
"\n",
"for doc_text in document_texts:\n",
" try:\n",
" result = await complete_document_workflow(doc_text)\n",
" results.append(result)\n",
" print(\"\\n\" + \"=\" * 60 + \"\\n\")\n",
" except Exception as e:\n",
" print(f\"❌ Error processing {doc_path}: {str(e)}\")\n",
" print(\"\\n\" + \"=\" * 60 + \"\\n\")\n",
"\n",
"print(f\"📋 Processed {len(results)} documents successfully!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Final Results Summary"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"📈 COMPLETE WORKFLOW RESULTS SUMMARY\n",
"======================================================================\n",
"\n",
"📄 Document 1: tmpos3b62tm.md\n",
" 📊 Classification: financial_document (confidence: 1.00)\n",
" 📝 Markdown length: 1,348,671 characters\n",
" 📋 Markdown sample: \n",
"\n",
"# UNITED STATES\n",
"# SECURITIES AND EXCHANGE COMMISSION\n",
"Washington, D.C. 20549\n",
"\n",
"## FORM 10-K\n",
"\n",
"(Mark O...\n",
" 🎯 Extracted fields: 7 fields\n",
" • company_name: Uber Technologies, Inc.\n",
" • document_type: Annual Report on Form 10-K\n",
" • fiscal_year: 2021\n",
" • revenue_2021: $21,764\n",
" • net_income_2021: $(496)\n",
" • key_business_segments: ['Mobility', 'Delivery', 'Freight', 'All Other (including former New Mobility, e-bikes, e-scooters, Advanced Technologies Group and other technology programs)']\n",
" • risk_factors: [\"The company faces numerous risk factors across its business operations and environment. The COVID-19 pandemic and related mitigation measures have adversely affected parts of the business, including reduced demand for Mobility offerings and creating ongoing uncertainties. The company's operational and financial performance is influenced by competitive pressure in the mobility, delivery, and logistics industries, characterized by well-established alternatives, low barriers to entry, and low switching costs. Driver classification risks exist if Drivers are deemed employees, workers, or quasi-employees rather than independent contractors, exposing the company to legal actions and financial liabilities globally. Competition challenges require the company to sometimes lower fares, offer incentives, and promotions, which impacts profitability. There are significant operating losses historically with substantial future operating expense increases anticipated, and the ability to achieve or maintain profitability is uncertain. Network value depends on maintaining critical mass among Drivers, consumers, merchants, shippers, and carriers, and failures to do so diminish platform attractiveness. Brand and reputation maintenance is critical, with exposure to negative publicity, media coverage, and risks from associated companies' brands or licensed brands in joint ventures.\\n\\nOperational risks include historical workplace culture and compliance challenges, management complexity due to rapid growth, technological infrastructure issues potentially causing disruptions or poor user experience, and security or data privacy breaches that could impact revenue and reputation. Platform users may engage in or be subjected to criminal, violent, or dangerous activity leading to safety incidents and legal actions. New offerings and technologies investments are inherently risky without guaranteed benefits. Economic conditions, inflation, and increased costs (fuel, food, labor, energy) may negatively impact results. Regulatory risks are extensive and global, involving payment and financial services compliance, licensing, anti-money laundering laws, data privacy (GDPR, CCPA, LGPD), and labor laws. Legal and regulatory investigations and inquiries, including antitrust, FCPA, labor classification, data protection, and intellectual property matters, pose risks of fines, penalties, operational changes, and increased costs.\\n\\nGeopolitical and jurisdictional risks include operating limitations or bans in some locations, currency exchange risk, and complex evolving regulations with the potential for fines and loss of licenses or permits. Insurance risks include potential inadequacy of reserves, liability exposure from accidents or impersonation, and insurer insolvency. Driver qualification requirements and background checks may increase costs or fail to expose all relevant information, with associated insurance cost risks and potential for courtroom or regulatory challenges to pricing models.\\n\\nFinancial risks comprise significant accumulated deficits, requirement for additional capital with uncertain availability, debt obligations, tax exposure including uncertain positions and observed changes in tax laws, and volatility in common stock price with no expected cash dividends. Accounting judgments and estimates involve critical assumptions affecting reported financial metrics related to goodwill, revenue recognition, incentive accruals, and stock-based compensation. Cybersecurity risks include exposures to malware, ransomware, phishing, and other cyberattacks. Climate change presents physical and transitional risks that may impact operations and costs, and failure to meet climate commitments may have operational and reputational consequences.\\n\\nOther risks include potential liability under anti-corruption and anti-terrorism laws, adverse effects from defaults under debt agreements, limitations in takeover actions due to corporate governance provisions, and the impact of non-GAAP financial measure limitations. Overall, these diverse and interconnected risk factors contribute to significant uncertainty regarding the company's future business prospects, operating results, and financial condition.\"]\n",
"\n",
"📄 Document 2: tmpppz9ub_m.md\n",
" 📊 Classification: technical_specification (confidence: 1.00)\n",
" 📝 Markdown length: 90,971 characters\n",
" 📋 Markdown sample: \n",
"\n",
"LM317\n",
"SLVS044Z SEPTEMBER 1997 REVISED APRIL 2025\n",
"\n",
"# LM317 3-Pin Adjustable Regulator\n",
"\n",
"## 1 Fea...\n",
" 🎯 Extracted fields: 8 fields\n",
" • component_name: LM317\n",
" • manufacturer: Texas Instruments\n",
" • part_number: LM317\n",
" • description: The LM317 is an adjustable three-pin, positive-voltage regulator capable of supplying up to 1.5A over an output voltage range of 1.25V to 37V. It features line and load regulation, internal current limiting, thermal overload protection, and safe operating area compensation.\n",
" • operating_voltage: {'min_voltage': 1.25, 'max_voltage': 37.0, 'unit': 'V'}\n",
" • maximum_current: 1.5\n",
" • key_features: ['Adjustable output voltage: 1.25V to 37V', 'Output current up to 1.5A', 'Line regulation: 0.01%/V (typical)', 'Load regulation: 0.1% (typical)', 'Internal short-circuit current limiting', 'Thermal overload protection', 'Output safe-area compensation', 'High power supply rejection ratio (PSRR): 80dB at 120Hz (new chip)', 'Available in SOT-223, TO-263, and TO-220 packages']\n",
" • applications: ['Multifunction printers', 'AC drive power stage modules', 'Electricity meters', 'Servo drive control modules', 'Merchant network and server power supply units']\n",
"\n",
"✨ Workflow completed successfully!\n",
"\n",
"📚 Key Learnings:\n",
" • Parse: Converted documents to clean markdown format\n",
" • Classify: Automatically categorized document types\n",
" • Extract: Used SourceText with markdown for structured data extraction\n",
" • The markdown content provides much better context for extraction than raw PDFs\n"
]
}
],
"source": [
"print(\"📈 COMPLETE WORKFLOW RESULTS SUMMARY\")\n",
"print(\"=\" * 70)\n",
"\n",
"for i, result in enumerate(results, 1):\n",
" print(f\"\\n📄 Document {i}: {os.path.basename(result['file_path'])}\")\n",
" print(\n",
" f\" 📊 Classification: {result['classification']} (confidence: {result['confidence']:.2f})\"\n",
" )\n",
" print(f\" 📝 Markdown length: {result['markdown_length']:,} characters\")\n",
" print(f\" 📋 Markdown sample: {result['markdown_sample'][:100]}...\")\n",
" print(f\" 🎯 Extracted fields: {len(result['extracted_data'])} fields\")\n",
"\n",
" # Print all keyvalue pairs\n",
" extracted = result[\"extracted_data\"]\n",
" for key, value in extracted.items():\n",
" print(f\" • {key}: {value}\")\n",
"\n",
"print(\"\\n✨ Workflow completed successfully!\")\n",
"print(\"\\n📚 Key Learnings:\")\n",
"print(\" • Parse: Converted documents to clean markdown format\")\n",
"print(\" • Classify: Automatically categorized document types\")\n",
"print(\" • Extract: Used SourceText with markdown for structured data extraction\")\n",
"print(\n",
" \" • The markdown content provides much better context for extraction than raw PDFs\"\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conclusion\n",
"\n",
"This notebook demonstrated the complete **Parse → Classify → Extract** workflow using LlamaCloud services:\n",
"\n",
"### Key Components:\n",
"\n",
"1. **LlamaParse** (`llama_cloud_services.parse.base.LlamaParse`):\n",
" - Converts documents to clean, structured markdown\n",
" - Preserves document structure and formatting\n",
" - Handles various file types (PDF, DOCX, etc.)\n",
"\n",
"2. **ClassifyClient** (`llama_cloud_services.beta.classifier.client.ClassifyClient`):\n",
" - Automatically categorizes documents based on content\n",
" - Uses customizable rules for classification\n",
" - Provides confidence scores for classifications\n",
"\n",
"3. **LlamaExtract with SourceText** (`llama_cloud_services.extract.extract.LlamaExtract`, `SourceText`):\n",
" - Extracts structured data using custom Pydantic schemas\n",
" - **SourceText** allows using markdown content as input instead of raw files\n",
" - Provides much better extraction accuracy when using processed markdown\n",
"\n",
"### Workflow Benefits:\n",
"\n",
"- **Better Accuracy**: Using markdown from parsing provides cleaner, more structured input for extraction\n",
"- **Automatic Routing**: Classification allows different processing logic for different document types\n",
"- **Structured Output**: Custom schemas ensure consistent, structured data extraction\n",
"- **Flexible Input**: SourceText supports text content, file paths, and bytes\n",
"\n",
"### Key Insights:\n",
"\n",
"1. **SourceText is the bridge**: It allows you to pass the clean markdown content from parsing directly to extraction\n",
"2. **Markdown improves extraction**: Pre-processed markdown provides much better context than raw PDFs\n",
"3. **Classification enables smart routing**: Different document types can use different extraction schemas\n",
"4. **End-to-end automation**: The entire workflow can be automated for production use\n",
"\n",
"This approach is ideal for production document processing pipelines where you need to:\n",
"- Process various document types automatically\n",
"- Extract structured data consistently\n",
"- Maintain high accuracy and reliability\n",
"- Handle documents at scale\n",
"\n",
"The combination of these three services provides a powerful, flexible document processing pipeline that can handle complex, real-world document processing requirements."
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
@@ -7,7 +7,7 @@
"source": [
"# Dynamic Section Retrieval with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services-demo/blob/main/examples/parse/advanced_rag/dynamic_section_retrieval.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/advanced_rag/dynamic_section_retrieval.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook showcases a concept called \"dynamic section retrieval\".\n",
"\n",
+1 -1
View File
@@ -6,7 +6,7 @@
"source": [
"# Advanced RAG with LlamaParse\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_parse/blob/main/examples/parse/demo_advanced.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook is a complete walkthrough for using LlamaParse with advanced indexing/retrieval techniques in LlamaIndex over the Apple 10K Filing. \n",
"\n",
+1 -1
View File
@@ -6,7 +6,7 @@
"source": [
"# RAG with Excel Spreadsheet using LlamaPrase\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_excel.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows you using LlamaParse with Excel Spreadsheet.\n",
"\n",
+1 -1
View File
@@ -7,7 +7,7 @@
"source": [
"# Download Charts\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_get_charts.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_get_charts.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook demonstrates how to download charts from a document using the result object.\n",
"\n",
+4 -4
View File
@@ -6,7 +6,7 @@
"source": [
"# LlamaParse - Fast checking Insurance Contract for Coverage\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_insurance.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"In this notebook we will look at how LlamaParse can be used to extract structured coverage information from an insurance policy.\n",
"\n",
@@ -36,7 +36,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"## Download an insurance policy fron IRDAI\n",
"## Download an insurance policy from IRDAI\n",
"\n",
"The Insurance Regulatory and Development Authority of India (IRDAI) maintains a great resource: https://policyholder.gov.in/web/guest/non-life-insurance-products where all insurance policies available in India are publicly available for download! Let's download a complex health insurance policy as an example."
]
@@ -228,11 +228,11 @@
" result_type=\"markdown\",\n",
" system_prompt_append=\"\"\"\n",
"This document is an insurance policy.\n",
"When a benefits/coverage/exlusion is describe in the document ammend to it add a text in the follwing benefits string format (where coverage could be an exclusion).\n",
"When a benefits/coverage/exlusion is describe in the document amend to it add a text in the following benefits string format (where coverage could be an exclusion).\n",
"\n",
"For {nameofrisk} and in this condition {whenDoesThecoverageApply} the coverage is {coverageDescription}. \n",
" \n",
"If the document contain a benefits TABLE that describe coverage amounts, do not ouput it as a table, but instead as a list of benefits string.\n",
"If the document contain a benefits TABLE that describe coverage amounts, do not output it as a table, but instead as a list of benefits string.\n",
" \n",
"\"\"\",\n",
").aparse(\"./policy.pdf\")\n",
+1 -1
View File
@@ -7,7 +7,7 @@
"source": [
"# LlamaParse `JobResult` Tour\n",
"\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/demo_json.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"The `JobResult` object is the main object returned by the LlamaParse API. It contains all the information about the job, including the parsed data, metadata, and any errors.\n",
"\n",
+1 -1
View File
@@ -9,7 +9,7 @@
"\n",
"LlamaParse supports users to specify a `language` parameter before uploading documents, giving users better OCR capabilities over non-English PDFs, parsing images into more accurate representations.\n",
"\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_cloud_services/blob/main/llama_parse/base.py.\n",
"You can specify 80+ different languages: see this file for a full list of supported languages: https://github.com/run-llama/llama_cloud_services/blob/main/py/llama_cloud_services/parse/base.py.\n",
"\n",
"This notebook shows a demo of this in action. \n",
"\n",
+1 -1
View File
@@ -4,7 +4,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/excel/o1_excel_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/parse/excel/o1_excel_rag.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>"
]
},
{
@@ -740,7 +740,7 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"In this example, these pages aren't going to be that different when parsed, but we can verify which pages triggered auto-made by looking at the [JSON output](https://github.com/run-llama/llama_cloud_services/blob/main/examples/demo_json_tour.ipynb) of LlamaParse:"
"In this example, these pages aren't going to be that different when parsed, but we can verify which pages triggered auto-made by looking at the [JSON output](https://github.com/run-llama/llama_cloud_services/blob/main/examples/parse/demo_json_tour.ipynb) of LlamaParse:"
]
},
{
@@ -1,6 +1,11 @@
import os
from typing import Any, Dict, Generic, List, Optional, Type
from llama_cloud import (
AgentData,
PaginatedResponseAgentData,
PaginatedResponseAggregateGroup,
)
from llama_cloud.client import AsyncLlamaCloud
from tenacity import (
WrappedFn,
@@ -86,7 +91,7 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
client=llama_client,
type=ExtractedPerson,
collection="extracted_people",
agent_url_id="person-extraction-agent"
deployment_name="person-extraction-agent"
)
# Create data
@@ -109,10 +114,12 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
self,
type: Type[AgentDataT],
collection: str = "default",
agent_url_id: Optional[str] = None,
deployment_name: Optional[str] = None,
client: Optional[AsyncLlamaCloud] = None,
token: Optional[str] = None,
base_url: Optional[str] = None,
# deprecated, use deployment_name instead
agent_url_id: Optional[str] = None,
):
"""
Initialize the AsyncAgentDataClient.
@@ -123,11 +130,11 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
collection: Named collection within the agent for organizing data.
Defaults to "default". Collections allow logical separation of
different data types or workflows within the same agent.
agent_url_id: Unique identifier for the agent. This normally appears in the
url of an agent within the llama cloud platform. If not provided,
will attempt to use the LLAMA_DEPLOY_DEPLOYMENT_NAME environment
variable. Data can only be added to an already existing agent in the
platform.
deployment_name: Unique identifier for the agent deployment. This normally
appears in the URL of an agent within the Llama Cloud platform. If not
provided, will attempt to use the LLAMA_DEPLOY_DEPLOYMENT_NAME
environment variable. Data can only be added to an already existing
agent in the platform.
client: AsyncLlamaCloud client instance for API communication. If not provided, will
construct one from the provided api token and base url
token: Llama Cloud API token. Reads from LLAMA_CLOUD_API_KEY if not provided
@@ -135,15 +142,14 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
defaults to https://api.cloud.llamaindex.ai
Raises:
ValueError: If agent_url_id is not provided and the
ValueError: If deployment_name is not provided and the
LLAMA_DEPLOY_DEPLOYMENT_NAME environment variable is not set
Note:
The client automatically applies retry logic to all API calls with
exponential backoff for timeout, connection, and HTTP status errors.
"""
self.agent_url_id = agent_url_id or get_default_agent_id()
self.deployment_name = deployment_name or agent_url_id or get_default_agent_id()
self.collection = collection
if not client:
@@ -156,15 +162,19 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
@agent_data_retry
async def get_item(self, item_id: str) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.get_agent_data(
raw_data = await self.untyped_get_item(item_id)
return TypedAgentData.from_raw(raw_data, self.type)
@agent_data_retry
async def untyped_get_item(self, item_id: str) -> AgentData:
return await self.client.beta.get_agent_data(
item_id=item_id,
)
return TypedAgentData.from_raw(raw_data, validator=self.type)
@agent_data_retry
async def create_item(self, data: AgentDataT) -> TypedAgentData[AgentDataT]:
raw_data = await self.client.beta.create_agent_data(
agent_slug=self.agent_url_id,
deployment_name=self.deployment_name,
collection=self.collection,
data=data.model_dump(),
)
@@ -210,9 +220,7 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
offset: Number of items to skip from the beginning. Defaults to 0.
include_total: Whether to include the total count in the response. Defaults to False to improve performance. It's recommended to only request on the first page.
"""
raw = await self.client.beta.search_agent_data_api_v_1_beta_agent_data_search_post(
agent_slug=self.agent_url_id,
collection=self.collection,
raw = await self.untyped_search(
filter=filter,
order_by=order_by,
offset=offset,
@@ -227,6 +235,25 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
total=raw.total_size,
)
@agent_data_retry
async def untyped_search(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
include_total: bool = False,
) -> PaginatedResponseAgentData:
return await self.client.beta.search_agent_data_api_v_1_beta_agent_data_search_post(
deployment_name=self.deployment_name,
collection=self.collection,
filter=filter,
order_by=order_by,
offset=offset,
page_size=page_size,
include_total=include_total,
)
@agent_data_retry
async def aggregate(
self,
@@ -253,8 +280,38 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
offset: Number of groups to skip from the beginning. Defaults to 0.
page_size: Maximum number of groups to return per page.
"""
raw = await self.client.beta.aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
agent_slug=self.agent_url_id,
raw = await self.untyped_aggregate(
filter=filter,
group_by=group_by,
count=count,
first=first,
order_by=order_by,
offset=offset,
page_size=page_size,
)
return TypedAggregateGroupItems(
items=[
TypedAggregateGroup.from_raw(grp, validator=self.type)
for grp in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@agent_data_retry
async def untyped_aggregate(
self,
filter: Optional[Dict[str, Dict[ComparisonOperator, Any]]] = None,
group_by: Optional[List[str]] = None,
count: Optional[bool] = None,
first: Optional[bool] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
) -> PaginatedResponseAggregateGroup:
return await self.client.beta.aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
deployment_name=self.deployment_name,
collection=self.collection,
page_size=page_size,
filter=filter,
@@ -264,11 +321,3 @@ class AsyncAgentDataClient(Generic[AgentDataT]):
first=first,
offset=offset,
)
return TypedAggregateGroupItems(
items=[
TypedAggregateGroup.from_raw(item, validator=self.type)
for item in raw.items
],
has_more=raw.next_page_token is not None,
total=raw.total_size,
)
@@ -10,7 +10,7 @@ CRUD operations, search capabilities, filtering, and aggregation functionality
for managing agent-generated data at scale.
Key Concepts:
- Agent Slug: Unique identifier for an agent instance
- Deployment Name: Unique identifier for an agent deployment
- Collection: Named grouping of data within an agent (defaults to "default"). Data within a collection should be of the same type.
- Agent Data: Individual structured data records with metadata and timestamps
@@ -26,7 +26,7 @@ Example Usage:
client=async_llama_cloud,
type=Person,
collection="people",
agent_url_id="my-extraction-agent-xyz"
deployment_name="my-extraction-agent-xyz"
)
# Create typed data
@@ -56,7 +56,6 @@ from typing import (
# Type variable for user-defined data models
AgentDataT = TypeVar("AgentDataT", bound=BaseModel)
# Type variable for extracted data (can be dict or Pydantic model)
ExtractedT = TypeVar("ExtractedT", bound=Union[BaseModel, dict])
@@ -78,7 +77,7 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
Attributes:
id: Unique identifier for this data record
agent_url_id: Identifier of the agent that created this data
deployment_name: Identifier of the agent deployment that created this data
collection: Named collection within the agent (used for organization)
data: The actual structured data payload (typed as AgentDataT)
created_at: Timestamp when the record was first created
@@ -94,8 +93,8 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
"""
id: Optional[str] = Field(description="Unique identifier for this data record")
agent_url_id: str = Field(
description="Identifier of the agent that created this data"
deployment_name: str = Field(
description="Identifier of the agent deployment that created this data"
)
collection: Optional[str] = Field(
description="Named collection within the agent for data organization"
@@ -116,15 +115,15 @@ class TypedAgentData(BaseModel, Generic[AgentDataT]):
Args:
raw_data: Raw agent data from the API
validator: Pydantic model class to validate the data field
Returns:
TypedAgentData instance with validated data
"""
data: AgentDataT = validator.model_validate(raw_data.data)
return cls(
id=raw_data.id,
agent_url_id=raw_data.agent_slug,
deployment_name=raw_data.deployment_name,
collection=raw_data.collection,
data=data,
created_at=raw_data.created_at,
@@ -222,12 +221,16 @@ def parse_extracted_field_metadata(
return {
k: _parse_extracted_field_metadata_recursive(v)
for k, v in field_metadata.items()
if k not in _METADATA_FIELDS_SIBLING_TO_LEAF
and k not in _ADDITIONAL_ROOT_METADATA_FIELDS
if not _is_reasoning_field(k, v) and k not in _ADDITIONAL_ROOT_METADATA_FIELDS
}
_METADATA_FIELDS_SIBLING_TO_LEAF = {"reasoning"}
def _is_reasoning_field(field_name: str, field_value: Any) -> bool:
# There can either be a user specified reasoning field (from the schema), or a reasoning metadata field for the
# dict of values
return field_name == "reasoning" and isinstance(field_value, str)
_ADDITIONAL_ROOT_METADATA_FIELDS = {"error"}
@@ -257,14 +260,12 @@ def _parse_extracted_field_metadata_recursive(
except ValidationError:
pass
additional_fields = {
k: v
for k, v in field_value.items()
if k in _METADATA_FIELDS_SIBLING_TO_LEAF
k: v for k, v in field_value.items() if _is_reasoning_field(k, v)
}
return {
k: _parse_extracted_field_metadata_recursive(v, additional_fields)
for k, v in field_value.items()
if k not in _METADATA_FIELDS_SIBLING_TO_LEAF
if not _is_reasoning_field(k, v)
}
elif isinstance(field_value, list):
return [_parse_extracted_field_metadata_recursive(item) for item in field_value]
+14 -12
View File
@@ -489,6 +489,7 @@ class LlamaCloudIndex(BaseManagedIndex):
name: str,
project_name: str = DEFAULT_PROJECT_NAME,
organization_id: Optional[str] = None,
project_id: Optional[str] = None,
api_key: Optional[str] = None,
base_url: Optional[str] = None,
app_url: Optional[str] = None,
@@ -504,15 +505,15 @@ class LlamaCloudIndex(BaseManagedIndex):
app_url = app_url or os.environ.get("LLAMA_CLOUD_APP_URL", DEFAULT_APP_URL)
client = get_client(api_key, base_url, app_url, timeout)
# create project if it doesn't exist
project = client.projects.upsert_project(
organization_id=organization_id, request=ProjectCreate(name=project_name)
)
if project.id is None:
raise ValueError(f"Failed to create/get project {project_name}")
if verbose:
print(f"Created project {project.id} with name {project.name}")
if project_id is None:
# create project if it doesn't exist
project = client.projects.upsert_project(
organization_id=organization_id,
request=ProjectCreate(name=project_name),
)
project_id = project.id
if verbose:
print(f"Created project {project_id} with name {project_name}")
# create pipeline
pipeline_create = PipelineCreate(
@@ -523,7 +524,7 @@ class LlamaCloudIndex(BaseManagedIndex):
llama_parse_parameters=llama_parse_parameters or LlamaParseParameters(),
)
pipeline = client.pipelines.upsert_pipeline(
project_id=project.id, request=pipeline_create
project_id=project_id, request=pipeline_create
)
if pipeline.id is None:
raise ValueError(f"Failed to create/get pipeline {name}")
@@ -532,8 +533,7 @@ class LlamaCloudIndex(BaseManagedIndex):
return cls(
name,
project_name=project.name,
organization_id=project.organization_id,
project_id=project_id,
api_key=api_key,
base_url=base_url,
app_url=app_url,
@@ -606,6 +606,7 @@ class LlamaCloudIndex(BaseManagedIndex):
name: str,
project_name: str = DEFAULT_PROJECT_NAME,
organization_id: Optional[str] = None,
project_id: Optional[str] = None,
api_key: Optional[str] = None,
base_url: Optional[str] = None,
app_url: Optional[str] = None,
@@ -631,6 +632,7 @@ class LlamaCloudIndex(BaseManagedIndex):
verbose=verbose,
embedding_config=embedding_config,
transform_config=transform_config,
project_id=project_id,
)
app_url = app_url or os.environ.get("LLAMA_CLOUD_APP_URL", DEFAULT_APP_URL)
+9
View File
@@ -420,6 +420,10 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, the parser will extract sub-tables from the spreadsheet when possible (more than one table per sheet).",
)
spreadsheet_force_formula_computation: Optional[bool] = Field(
default=False,
description="If set to true, the parser will re-compute values for all spreadsheet cells containing formulas.",
)
specialized_chart_parsing_agentic: Optional[bool] = Field(
default=False,
description="If set to true, the parser will use a specialized agentic chart parsing model to extract data from charts. This model is able to understand the chart type and extract the data accordingly.",
@@ -965,6 +969,11 @@ class LlamaParse(BasePydanticReader):
if self.spreadsheet_extract_sub_tables:
data["spreadsheet_extract_sub_tables"] = self.spreadsheet_extract_sub_tables
if self.spreadsheet_force_formula_computation:
data[
"spreadsheet_force_formula_computation"
] = self.spreadsheet_force_formula_computation
if self.specialized_chart_parsing_agentic:
data[
"specialized_chart_parsing_agentic"
+2 -2
View File
@@ -11,13 +11,13 @@ dev = [
[project]
name = "llama-parse"
version = "0.6.65"
version = "0.6.69"
description = "Parse files into RAG-Optimized formats."
authors = [{name = "Logan Markewich", email = "logan@llamaindex.ai"}]
requires-python = ">=3.9,<4.0"
readme = "README.md"
license = "MIT"
dependencies = ["llama-cloud-services>=0.6.64"]
dependencies = ["llama-cloud-services>=0.6.69"]
[project.scripts]
llama-parse = "llama_parse.cli.main:parse"
+2 -2
View File
@@ -19,7 +19,7 @@ dev = [
[project]
name = "llama-cloud-services"
version = "0.6.65"
version = "0.6.69"
description = "Tailored SDK clients for LlamaCloud services."
authors = [{name = "Logan Markewich", email = "logan@runllama.ai"}]
requires-python = ">=3.9,<4.0"
@@ -27,7 +27,7 @@ readme = "README.md"
license = "MIT"
dependencies = [
"llama-index-core>=0.12.0",
"llama-cloud==0.1.41",
"llama-cloud==0.1.42",
"pydantic>=2.8,!=2.10",
"click>=8.1.7,<9",
"python-dotenv>=1.0.1,<2",
@@ -68,7 +68,7 @@ async def test_agent_data_crud_operations():
client=client,
type=ExampleData,
collection=f"test-collection-{test_id[:8]}",
agent_url_id=LLAMA_DEPLOY_DEPLOYMENT_NAME,
deployment_name=LLAMA_DEPLOY_DEPLOYMENT_NAME,
)
# Create test data
@@ -0,0 +1,158 @@
import pytest
from typing import Any, Dict, List, Optional
from pydantic import BaseModel
from datetime import datetime
from llama_cloud.types.agent_data import AgentData
from llama_cloud.types.aggregate_group import AggregateGroup
from llama_cloud_services.beta.agent_data.client import AsyncAgentDataClient
class Person(BaseModel):
name: str
age: int
class FakeBeta:
def __init__(self) -> None:
self._get_item_response: Optional[AgentData] = None
self._search_items: List[AgentData] = []
self._aggregate_items: List[AggregateGroup] = []
self._total_size: Optional[int] = None
self._next_page_token: Optional[str] = None
# Single get
async def get_agent_data(self, item_id: str) -> AgentData:
assert self._get_item_response is not None, "_get_item_response not set"
return self._get_item_response
# Search
async def search_agent_data_api_v_1_beta_agent_data_search_post(
self,
*,
deployment_name: str,
collection: str,
filter: Optional[Dict[str, Any]] = None,
order_by: Optional[str] = None,
offset: Optional[int] = None,
page_size: Optional[int] = None,
include_total: bool = False,
) -> Any:
class Resp:
def __init__(
self,
items: List[AgentData],
total_size: Optional[int],
next_page_token: Optional[str],
) -> None:
self.items = items
self.total_size = total_size
self.next_page_token = next_page_token
return Resp(self._search_items, self._total_size, self._next_page_token)
# Aggregate
async def aggregate_agent_data_api_v_1_beta_agent_data_aggregate_post(
self,
*,
deployment_name: str,
collection: str,
page_size: Optional[int] = None,
filter: Optional[Dict[str, Any]] = None,
order_by: Optional[str] = None,
group_by: Optional[List[str]] = None,
count: Optional[bool] = None,
first: Optional[bool] = None,
offset: Optional[int] = None,
) -> Any:
class Resp:
def __init__(
self,
items: List[AggregateGroup],
total_size: Optional[int],
next_page_token: Optional[str],
) -> None:
self.items = items
self.total_size = total_size
self.next_page_token = next_page_token
return Resp(self._aggregate_items, self._total_size, self._next_page_token)
class FakeClient:
def __init__(self) -> None:
self.beta = FakeBeta()
def make_agent_data(data: Dict[str, Any]) -> AgentData:
return AgentData(
id="id-1",
deployment_name="dep",
collection="col",
data=data,
created_at=datetime.now(),
updated_at=datetime.now(),
)
def make_group(
group_key: Dict[str, Any],
first_item: Optional[Dict[str, Any]],
count: Optional[int] = None,
) -> AggregateGroup:
return AggregateGroup(group_key=group_key, count=count, first_item=first_item)
@pytest.mark.asyncio
async def test_untyped_get_item_valid_to_dict() -> None:
client = FakeClient()
client.beta._get_item_response = make_agent_data({"name": "Alice", "age": 30})
adc = AsyncAgentDataClient(type=Person, client=client, deployment_name="dep")
item = await adc.untyped_get_item("id-1")
assert item.data == {"name": "Alice", "age": 30}
@pytest.mark.asyncio
async def test_untyped_get_item_invalid_retains_dict() -> None:
client = FakeClient()
# age wrong type; will fail validation and should be returned as dict
client.beta._get_item_response = make_agent_data({"name": "Bob", "age": "x"})
adc = AsyncAgentDataClient(type=Person, client=client, deployment_name="dep")
item = await adc.untyped_get_item("id-1")
assert item.data == {"name": "Bob", "age": "x"}
@pytest.mark.asyncio
async def test_untyped_search_mixed_items() -> None:
client = FakeClient()
client.beta._search_items = [
make_agent_data({"name": "Carol", "age": 22}),
make_agent_data({"name": "Dave", "age": "bad"}),
]
client.beta._total_size = 2
adc = AsyncAgentDataClient(type=Person, client=client, deployment_name="dep")
results = await adc.untyped_search(include_total=True)
assert len(results.items) == 2
assert results.items[0].data == {"name": "Carol", "age": 22}
assert results.items[1].data == {"name": "Dave", "age": "bad"}
assert results.total_size == 2
@pytest.mark.asyncio
async def test_untyped_aggregate_first_item_dict() -> None:
client = FakeClient()
client.beta._aggregate_items = [
make_group({"k": 1}, {"name": "Eve", "age": 40}),
make_group({"k": 2}, {"name": "Frank", "age": "bad"}),
]
client.beta._total_size = 2
adc = AsyncAgentDataClient(type=Person, client=client, deployment_name="dep")
results = await adc.untyped_aggregate(group_by=["k"], first=True)
assert len(results.items) == 2
assert results.items[0].first_item == {"name": "Eve", "age": 40}
assert results.items[1].first_item == {"name": "Frank", "age": "bad"}
@@ -38,7 +38,7 @@ def test_typed_agent_data_from_raw():
"""Test TypedAgentData.from_raw class method."""
raw_data = AgentData(
id="456",
agent_slug="extraction-agent",
deployment_name="extraction-agent",
collection="employees",
data={"name": "Jane Smith", "age": 25, "email": "jane@company.com"},
created_at=datetime.now(),
@@ -48,7 +48,7 @@ def test_typed_agent_data_from_raw():
typed_data = TypedAgentData.from_raw(raw_data, Person)
assert typed_data.id == "456"
assert typed_data.agent_url_id == "extraction-agent"
assert typed_data.deployment_name == "extraction-agent"
assert typed_data.collection == "employees"
assert typed_data.data.name == "Jane Smith"
assert typed_data.data.age == 25
@@ -56,10 +56,10 @@ def test_typed_agent_data_from_raw():
def test_typed_agent_data_from_raw_validation_error():
"""Test TypedAgentData.from_raw with invalid data."""
"""Test TypedAgentData.from_raw with invalid data now raises InvalidTypedAgentData."""
raw_data = AgentData(
id="789",
agent_slug="test-agent",
deployment_name="test-agent",
collection="people",
data={"name": "Invalid Person", "age": "not_a_number"}, # Invalid age
created_at=datetime.now(),
@@ -613,3 +613,51 @@ def test_parses_field_metadata_with_error_field():
}
assert parsed.metadata.get("field_errors") == "This is an error"
assert parsed.metadata.get("job_id") == "job-123"
REASONING_IN_SCHEMA = {
"majority_opinion": {
"type": {
"citation": [
{
"page": 4,
"matching_text": "BARRETT, J., delivered the opinion for a unanimous Court.",
},
{"page": 11, "matching_text": "Opinion of the Court"},
],
"parsing_confidence": 1.0,
"extraction_confidence": 0.9999998919950147,
"confidence": 0.9999998919950147,
},
"reasoning": {
"citation": [
{
"page": 15,
"matching_text": "We hold that §5110(b)(1) is not subject to equitable tolling and affirm the judg...",
}
],
"parsing_confidence": 1.0,
"extraction_confidence": 0.414292785946868,
"confidence": 0.414292785946868,
},
},
"reasoning": {
"citation": [
{
"page": 15,
"matching_text": "We hold that §5110(b)(1) is not subject to equitable tolling and affirm the judg...",
}
],
"parsing_confidence": 1.0,
"extraction_confidence": 0.414292785946868,
"confidence": 0.414292785946868,
},
}
def test_field_conflict_in_schema():
extracted = parse_extracted_field_metadata(REASONING_IN_SCHEMA)
assert isinstance(extracted["reasoning"], ExtractedFieldMetadata)
assert isinstance(
extracted["majority_opinion"]["reasoning"], ExtractedFieldMetadata
)
+143 -4
View File
@@ -1,16 +1,155 @@
import pytest
from unittest.mock import MagicMock, patch
import llama_cloud_services.index.base as base
from llama_cloud import (
PipelineEmbeddingConfig_ManagedOpenaiEmbedding,
Project,
Pipeline,
CloudDocument,
)
from llama_index.core.constants import DEFAULT_PROJECT_NAME
from llama_index.core.indices.managed.base import BaseManagedIndex
from llama_cloud_services.index import (
LlamaCloudIndex,
from llama_index.core.schema import Document
from llama_cloud_services.index import LlamaCloudIndex
# Simple test data as values, not fixtures
TEST_PROJECT = Project(id="proj-123", name="test-project", organization_id="org-123")
EMBEDDING_CONFIG = PipelineEmbeddingConfig_ManagedOpenaiEmbedding(
type="MANAGED_OPENAI_EMBEDDING"
)
TEST_PIPELINE = Pipeline(
id="pipe-456",
name="test-pipeline",
project_id="proj-123",
embedding_config=PipelineEmbeddingConfig_ManagedOpenaiEmbedding(
type="MANAGED_OPENAI_EMBEDDING"
),
)
def test_class():
@pytest.fixture
def mock_client() -> MagicMock:
"""Mock client with sensible defaults."""
client = MagicMock()
client.projects.upsert_project.return_value = Project(
id="default-proj", name=DEFAULT_PROJECT_NAME, organization_id="default-org"
)
client.pipelines.upsert_pipeline.return_value = Pipeline(
id="default-pipe",
name="default",
project_id="default-proj",
embedding_config=EMBEDDING_CONFIG,
)
client.pipelines.upsert_batch_pipeline_documents.return_value = [
CloudDocument(id="doc-1", text="test", metadata={})
]
return client
@pytest.fixture(autouse=True)
def base_patches(mock_client: MagicMock) -> None:
"""Auto-applied patches for all tests."""
with (
patch.object(base, "get_client", return_value=mock_client),
patch.object(
base,
"resolve_project_and_pipeline",
return_value=(TEST_PROJECT, TEST_PIPELINE),
),
patch.object(base.LlamaCloudIndex, "wait_for_completion"),
):
yield
def test_class() -> None:
names_of_base_classes = [b.__name__ for b in LlamaCloudIndex.__mro__]
assert BaseManagedIndex.__name__ in names_of_base_classes
def test_conflicting_index_identifiers():
def test_conflicting_index_identifiers() -> None:
with pytest.raises(ValueError):
LlamaCloudIndex(name="test", pipeline_id="test", index_id="test")
def test_from_documents_uses_provided_project_id(mock_client: MagicMock) -> None:
provided_project_id = "proj-123"
organization_id = "org-abc"
index_name = "my_new_index"
# Override resolve to return project with provided ID
test_project = Project(
id=provided_project_id, name="my_project", organization_id=organization_id
)
test_pipeline = Pipeline(
id="pipe-xyz",
name=index_name,
project_id=provided_project_id,
embedding_config=EMBEDDING_CONFIG,
)
with patch.object(
base, "resolve_project_and_pipeline", return_value=(test_project, test_pipeline)
):
docs = [Document(text="hello")]
index = LlamaCloudIndex.from_documents(
documents=docs,
name=index_name,
project_id=provided_project_id,
)
# Assert - project upsert not called; pipeline uses provided project_id
mock_client.projects.upsert_project.assert_not_called()
assert mock_client.pipelines.upsert_pipeline.call_count == 1
assert (
mock_client.pipelines.upsert_pipeline.call_args.kwargs["project_id"]
== provided_project_id
)
assert index.project.id == provided_project_id
def test_from_documents_upserts_project_when_project_id_missing(
mock_client: MagicMock,
) -> None:
organization_id = "org-xyz"
index_name = "my_new_index"
# Project is created when project_id is not provided
upserted_project = Project(
id="proj-999", name=DEFAULT_PROJECT_NAME, organization_id=organization_id
)
mock_client.projects.upsert_project.return_value = upserted_project
test_pipeline = Pipeline(
id="pipe-xyz",
name=index_name,
project_id=upserted_project.id,
embedding_config=EMBEDDING_CONFIG,
)
with patch.object(
base,
"resolve_project_and_pipeline",
return_value=(upserted_project, test_pipeline),
):
docs = [Document(text="world")]
index = LlamaCloudIndex.from_documents(
documents=docs,
name=index_name,
organization_id=organization_id,
)
# Assert - project was upserted with org id and default project name
mock_client.projects.upsert_project.assert_called_once()
kwargs = mock_client.projects.upsert_project.call_args.kwargs
assert kwargs["organization_id"] == organization_id
assert kwargs["request"].name == DEFAULT_PROJECT_NAME
# Pipeline created under the upserted project id
assert (
mock_client.pipelines.upsert_pipeline.call_args.kwargs["project_id"]
== upserted_project.id
)
assert index.project.id == upserted_project.id
Generated
+5 -5
View File
@@ -1582,21 +1582,21 @@ wheels = [
[[package]]
name = "llama-cloud"
version = "0.1.41"
version = "0.1.42"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "certifi" },
{ name = "httpx" },
{ name = "pydantic" },
]
sdist = { url = "https://files.pythonhosted.org/packages/62/6c/b2e84eebed376aea34c446cab745da5fc4e9dc53309180672299083219d5/llama_cloud-0.1.41.tar.gz", hash = "sha256:dcb741b779e3e740cd64928cfffc8ef70ed0e9bae9ef26acbe1d7e32aa737bdc", size = 109854, upload-time = "2025-09-05T22:45:13.069Z" }
sdist = { url = "https://files.pythonhosted.org/packages/21/04/ae0694b582d6aab4d6e7957febb7bff048897ac231ad80ba1bd71547d944/llama_cloud-0.1.42.tar.gz", hash = "sha256:485aa0e364ea648e3aaa3b2c54af7bcb6f2242c50b4f86ec022e137413fff464", size = 112480, upload-time = "2025-09-16T20:25:42.631Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/1e/4d/f0af76b389310840ce3483a92560a152025b0eefe4eee0c81102bf3317e6/llama_cloud-0.1.41-py3-none-any.whl", hash = "sha256:c847f288f0d3f4b23f47345088006deae5f2cf3f223ac1819d4c1531e9aaa13e", size = 307646, upload-time = "2025-09-05T22:45:11.597Z" },
{ url = "https://files.pythonhosted.org/packages/6a/61/85d115699a59d03f0783e119aaf6d534fca95dbe1a4531a8056e6a4774ed/llama_cloud-0.1.42-py3-none-any.whl", hash = "sha256:4ed3edde4a277ff52eeb831188c8476eb079b5e4605ad3142157a0f054b27d96", size = 311857, upload-time = "2025-09-16T20:25:41.479Z" },
]
[[package]]
name = "llama-cloud-services"
version = "0.6.65"
version = "0.6.68"
source = { editable = "." }
dependencies = [
{ name = "click", version = "8.1.8", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.10'" },
@@ -1631,7 +1631,7 @@ dev = [
requires-dist = [
{ name = "click", specifier = ">=8.1.7,<9" },
{ name = "eval-type-backport", marker = "python_full_version < '3.10'", specifier = ">=0.2.0,<0.3" },
{ name = "llama-cloud", specifier = "==0.1.41" },
{ name = "llama-cloud", specifier = "==0.1.42" },
{ name = "llama-index-core", specifier = ">=0.12.0" },
{ name = "packaging", specifier = ">=25.0" },
{ name = "platformdirs", specifier = ">=4.3.7,<5" },
+7 -8
View File
@@ -120,7 +120,7 @@ def get_current_branch() -> str:
return result.stdout.strip()
def create_if_not_exists(version: str) -> None:
def create_if_not_exists(version: str) -> str:
"""Create a git tag and push it."""
current_branch = get_current_branch()
if current_branch != "main":
@@ -130,26 +130,25 @@ def create_if_not_exists(version: str) -> None:
sys.exit(1)
tag_name = f"v{version}" if version[0].isdigit() else version
if not tag_exists(version):
if not tag_exists(tag_name):
# Create tag
subprocess.run(["git", "tag", tag_name], check=True)
click.echo(f"Created tag {tag_name}")
else:
click.echo(f"Tag {tag_name} already exists")
return tag_name
def tag_exists(version: str) -> bool:
def tag_exists(tag_name: str) -> bool:
"""Check if a git tag exists."""
tag_name = f"v{version}"
result = subprocess.run(
["git", "tag", "-l", tag_name], capture_output=True, text=True, check=True
)
return tag_name in result.stdout.strip()
def push_tag(version: str) -> None:
def push_tag(tag_name: str) -> None:
"""Push a git tag."""
tag_name = f"v{version}"
subprocess.run(["git", "push", "origin", tag_name], check=True)
click.echo(f"Pushed tag {tag_name}")
@@ -218,9 +217,9 @@ def tag(version: str | None = None, push: bool = False, js: bool = False) -> Non
main_version, _, _, js_version = get_current_versions()
version = f"llama-cloud-services@{js_version}" if js else main_version
create_if_not_exists(version)
tag_name = create_if_not_exists(version)
if push:
push_tag(version)
push_tag(tag_name)
if __name__ == "__main__":
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "llama-cloud-services",
"version": "0.3.5",
"version": "0.3.6",
"type": "module",
"license": "MIT",
"scripts": {
@@ -25,20 +25,23 @@ import type {
export class AgentClient<T = unknown> {
private client: ReturnType<typeof createClient>;
private collection: string;
private agentUrlId: string;
private deploymentName: string;
constructor({
client = defaultClient,
collection = "default",
agentUrlId = "_public",
deploymentName = "_public",
agentUrlId,
}: {
client?: ReturnType<typeof createClient>;
collection?: string;
deploymentName?: string;
// deprecated, use deploymentName instead
agentUrlId?: string;
}) {
this.client = client;
this.collection = collection;
this.agentUrlId = agentUrlId;
this.deploymentName = agentUrlId || deploymentName;
}
/**
@@ -48,7 +51,7 @@ export class AgentClient<T = unknown> {
const response = await createAgentDataApiV1BetaAgentDataPost({
throwOnError: true,
body: {
agent_slug: this.agentUrlId,
deployment_name: this.deploymentName,
collection: this.collection,
data: data as Record<string, unknown>,
},
@@ -118,7 +121,7 @@ export class AgentClient<T = unknown> {
const response = await searchAgentDataApiV1BetaAgentDataSearchPost({
throwOnError: true,
body: {
agent_slug: this.agentUrlId,
deployment_name: this.deploymentName,
...(this.collection !== undefined && {
collection: this.collection,
}),
@@ -165,7 +168,7 @@ export class AgentClient<T = unknown> {
const response = await aggregateAgentDataApiV1BetaAgentDataAggregatePost({
throwOnError: true,
body: {
agent_slug: this.agentUrlId,
deployment_name: this.deploymentName,
...(this.collection !== undefined && {
collection: this.collection,
}),
@@ -209,7 +212,7 @@ export class AgentClient<T = unknown> {
private transformResponse(data: AgentData): TypedAgentData<T> {
const result: TypedAgentData<T> = {
id: data.id!,
agentUrlId: data.agent_slug,
deploymentName: data.deployment_name,
data: data.data as T,
createdAt: new Date(data.created_at!),
updatedAt: new Date(data.updated_at!),
@@ -250,10 +253,10 @@ export interface AgentDataClientOptions {
/** Base URL for the client */
/** Base URL of the llama cloud api */
baseUrl?: string;
/** If running in an agent runtime, optionally provide the window url to infer the agent url id */
/** If running in an agent runtime, optionally provide the window url to infer the deployment name */
windowUrl?: string;
/** Agent URL ID for the client, if not provided, it will be inferred from the window url, or fall back to "default" */
agentUrlId?: string;
/** Deployment name for the client, if not provided, it will be inferred from the window url, or fall back to "default" */
deploymentName?: string;
/** Collection name for the client, defaults to "default" */
collection?: string;
}
@@ -267,22 +270,25 @@ export function createAgentDataClient<T = unknown>({
client = defaultClient,
windowUrl,
env,
deploymentName,
agentUrlId,
collection = "default",
}: {
client?: ReturnType<typeof createClient>;
windowUrl?: string;
env?: Record<string, string>;
deploymentName?: string;
// deprecated, use deploymentName instead
agentUrlId?: string;
collection?: string;
} = {}): AgentClient<T> {
if (env && !agentUrlId) {
agentUrlId =
if (env && !deploymentName) {
deploymentName =
env.LLAMA_DEPLOY_DEPLOYMENT_NAME ||
env.NEXT_PUBLIC_LLAMA_DEPLOY_DEPLOYMENT_NAME ||
env.VITE_LLAMA_DEPLOY_DEPLOYMENT_NAME;
}
if (windowUrl && !agentUrlId) {
if (windowUrl && !deploymentName) {
try {
const url = new URL(windowUrl);
const path = url.pathname;
@@ -291,17 +297,18 @@ export function createAgentDataClient<T = unknown>({
url.hostname.includes("127.0.0.1");
if (path.startsWith("/deployments/") && !isLocalhost) {
// /deployments/<agent-url-id>/ui/ -> ["", "deployments", "<agent-url-id>", "ui"]
agentUrlId = path.split("/")[2];
deploymentName = path.split("/")[2];
}
} catch (error) {
console.warn(
"Failed to infer agent url id from window url, falling back to default",
"Failed to infer deployment name from window url, falling back to default",
error,
);
}
}
return new AgentClient({
...(deploymentName && { deploymentName }),
...(agentUrlId && { agentUrlId }),
collection,
client,
@@ -87,8 +87,8 @@ export interface ExtractedData<T = unknown> {
export interface TypedAgentData<T = unknown> {
/** The unique ID of the agent data record. */
id: string;
/** The ID of the agent that created the data. */
agentUrlId: string;
/** The deployment name of the agent that created the data. */
deploymentName: string;
/** The collection of the agent data. */
collection?: string;
/** The data of the agent data. Usually an ExtractedData&lt;SomeOtherType&gt; */
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff