Compare commits

...

59 Commits

Author SHA1 Message Date
Clelia (Astra) Bertelli e110334273 feat: add mocking for agent data (#1028)
* wip: add mocking for agent data

* chore: add tests; fix: various fixes

* chore: fix schema generation null handling
2025-11-25 21:59:17 +01:00
Adrian Lyjak 3b48ea78c0 fix tests 2025-11-24 08:47:08 -05:00
Adrian Lyjak 5a5986e59b Implement test plan and create mocked tests (#1023)
feat: Add FakeLlamaCloudServer for testing

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2025-11-24 08:05:29 -05:00
Adrian Lyjak 8aa5ab0756 plan 2025-11-24 00:05:01 -05:00
Adrian Lyjak 30e36f3cc3 docs_improved2 2025-11-23 21:10:58 -05:00
Adrian Lyjak 549efbd8b9 docs_improved 2025-11-23 20:50:56 -05:00
Adrian Lyjak d312dc9890 add docs.md 2025-11-23 19:51:37 -05:00
Adrian Lyjak 10ab5fd2b9 planish 2025-11-23 15:50:36 -05:00
Neeraj Pradhan ad38ef5cd7 Add notebook for tabular extraction (#1017) 2025-11-18 09:47:07 -08:00
Logan Markewich 4c4c6e6575 fix sheets test 2025-11-17 16:14:29 -06:00
github-actions[bot] 740b47d9dc chore: version packages (#1016)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-17 16:11:18 -06:00
Logan f3233deb2e propagate retrieval metadata to retrieved nodes (#1015) 2025-11-17 16:06:52 -06:00
github-actions[bot] fd45127678 chore: version packages (#1014)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-17 21:18:09 +01:00
Clelia (Astra) Bertelli 0506c88735 chore: rename classifyclient and keep it backward compatible (#1013)
* chore: rename classifyclient and keep it backward compatible

* chore: Replace ClassifyClient in notebooks

* chore: changesets
2025-11-17 21:16:23 +01:00
Logan 4bc9eb6c0d beta sheets API (#992) 2025-11-17 11:32:06 -06:00
Patricia 5a3dac655c Add support for custom metadata in file upload methods (#1012) 2025-11-17 11:18:11 -06:00
github-actions[bot] 519254efbe chore: version packages (#999)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-11-04 14:18:27 -05:00
Adrian Lyjak 6ab56b79f3 fix version breaking (#998) 2025-11-04 14:14:38 -05:00
Adrian Lyjak e020e3e2b1 Remove organization id from classify (#997) 2025-11-04 14:05:19 -05:00
Adrian Lyjak f293547910 destructured keyword params for classify (#996) 2025-11-04 14:04:41 -05:00
github-actions[bot] 662bc37462 chore: version packages (#995) 2025-11-03 20:15:50 -06:00
Neeraj Pradhan 9f1ef4ef1f Bump to version 0.6.78 (#994) 2025-11-03 20:11:18 -06:00
github-actions[bot] 1243573924 chore: version packages (#991) 2025-10-30 10:11:16 -06:00
Preston Carlson 407292b177 Fix: Return partial results on job failure (#990)
* Return partial result on failed job, especially job id

* Maintains NO_DATA_FOUND_IN_FILE throw behavior
2025-10-23 13:44:41 -07:00
Clelia (Astra) Bertelli a7df7c0912 docs: add llamaclassify demo (#989) 2025-10-23 17:38:57 +02:00
github-actions[bot] c758144bfe chore: version packages (#988)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-22 14:41:44 +02:00
Clelia (Astra) Bertelli fee516dd19 feat: add classify to ts sdk (#985)
* feat: add classify to ts sdk

* ci: changesets

* chore: camelCase for everyone; refactor: slimmer logic for fileContents/filePaths handling

* chore: implement claude suggestions
2025-10-22 14:39:20 +02:00
Neeraj Pradhan 032fbd5768 Add common SourceText class for classify/extract text inputs (#986) 2025-10-21 13:37:41 -07:00
Jerry Liu 970e864514 improve classify notebook (#983) 2025-10-20 10:07:35 -07:00
github-actions[bot] d0649ece6e chore: version packages (#982) 2025-10-16 16:58:29 -06:00
MartijnLeplae 5d4cabd843 Add ImageNode support in TypeScript (#969) 2025-10-16 16:56:28 -06:00
github-actions[bot] 9070a6ac16 chore: version packages (#981) 2025-10-15 12:01:34 -06:00
Bogdan Gheorghe 4f24f537f6 Add agressive table extraction argument (#980) 2025-10-15 11:57:34 -06:00
github-actions[bot] 8859a203e2 chore: version packages (#977) 2025-10-14 19:03:36 -06:00
dependabot[bot] b091364054 build(deps): bump astral-sh/setup-uv from 6 to 7 (#974) 2025-10-14 19:02:32 -06:00
dependabot[bot] 43b1a013ca build(deps): bump github/codeql-action from 3 to 4 (#973) 2025-10-14 19:02:20 -06:00
Logan f81532e7f2 safest types possible for parse (#976) 2025-10-14 19:02:07 -06:00
github-actions[bot] 986d3987d3 chore: version packages (#965)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-14 08:14:49 -06:00
Logan 1bf522311f fix default bbox values (#975) 2025-10-14 07:44:35 -06:00
Preston Carlson 24166dcfc8 Only escape single dollar sign in notebook md (#964)
* Limit escaping to lone dollar signs - preserve double dollar for latex equations

* Updated uv.lock via make lint

* Patch bump

* Unit test for _format_markdown_for_notebook

Test doesn't depend on getting real results/is just testing a string manipulation function, so inserting before other tests. Should move to its own file if we add additional formatting configurations
2025-10-07 08:06:03 -07:00
github-actions[bot] bfb7f3973f chore: version packages (#956)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-06 11:15:55 -04:00
dependabot[bot] 979f643c77 build(deps): bump actions/checkout from 4 to 5 (#961) 2025-10-06 09:12:38 -06:00
dependabot[bot] aefd89cf1b build(deps): bump actions/setup-python from 5 to 6 (#960) 2025-10-06 09:12:30 -06:00
dependabot[bot] 8ea2b2c64e build(deps): bump pnpm/action-setup from 3 to 4 (#959) 2025-10-06 09:12:20 -06:00
dependabot[bot] 4a9a2a21d8 build(deps): bump astral-sh/setup-uv from 3 to 6 (#958) 2025-10-06 09:12:08 -06:00
Logan e6a7939206 loosen packaging requirements (#962) 2025-10-06 09:11:57 -06:00
Adrian Lyjak 104a03e829 fix: re-enable js publishing (#963) 2025-10-06 11:10:46 -04:00
Terry Zhao 6e0f2f4ca0 citation can be null (#869)
* citation can be null

* Add changeset

---------

Co-authored-by: Terry Zhao <terryzhao@runllama.ai>
Co-authored-by: Adrian Lyjak <adrianlyjak@gmail.com>
2025-10-04 16:26:11 -04:00
dependabot[bot] 0708d11f8a Bump actions/setup-node from 4 to 5 (#909)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 4 to 5.
- [Release notes](https://github.com/actions/setup-node/releases)
- [Commits](https://github.com/actions/setup-node/compare/v4...v5)

---
updated-dependencies:
- dependency-name: actions/setup-node
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-10-04 16:21:50 -04:00
github-actions[bot] be19185503 chore: version packages (#954)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2025-10-03 20:14:04 -04:00
Adrian Lyjak 7571b0d6c4 Missed some things again with tag fixes (#955)
guh
2025-10-03 20:12:53 -04:00
Adrian Lyjak ad6734bf80 fixup tagging more better (#953)
* fix: correct private field type in py/package.json to be recognized by pnpm

* use packages more directly, make public

* add bump

* fix crash
2025-10-03 19:53:57 -04:00
github-actions[bot] 9ec2a8322e chore: version packages (#952) 2025-10-03 15:11:14 -06:00
Logan 51011b9f30 fix changeset harder (#951) 2025-10-03 15:09:58 -06:00
Logan 09805f9e15 swap changesets (#949) 2025-10-03 15:06:00 -06:00
Adrian Lyjak 8ced6f6eab fix: explicitly tag. I thought the action did this (#948) 2025-10-03 16:59:41 -04:00
Preston Carlson 081ddeca34 Escaping dollar signs in md output when running in a jupyter notebook (#945) 2025-10-03 14:52:26 -06:00
Adrian Lyjak 2460908789 Disable npm release (#946) 2025-10-03 16:13:16 -04:00
Adrian Lyjak c226d6a54c Fix more bugs in publishing (#944) 2025-10-03 11:16:43 -04:00
115 changed files with 12633 additions and 668 deletions
+1 -1
View File
@@ -27,7 +27,7 @@ jobs:
- uses: actions/checkout@v5
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+1 -1
View File
@@ -21,7 +21,7 @@ jobs:
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
+2 -2
View File
@@ -30,12 +30,12 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v3
uses: github/codeql-action/init@v4
with:
languages: python
dependency-caching: true
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v3
uses: github/codeql-action/analyze@v4
with:
category: "/language:python"
+2 -2
View File
@@ -22,7 +22,7 @@ jobs:
with:
fetch-depth: ${{ github.event_name == 'pull_request' && 2 || 0 }}
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
@@ -31,7 +31,7 @@ jobs:
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
+1 -1
View File
@@ -22,7 +22,7 @@ jobs:
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
with:
fetch-depth: 0
- name: Install uv
uses: astral-sh/setup-uv@v6
uses: astral-sh/setup-uv@v7
with:
version: ${{ env.UV_VERSION }}
+1 -1
View File
@@ -24,7 +24,7 @@ jobs:
- uses: actions/checkout@v5
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version-file: "ts/llama_cloud_services/.nvmrc"
- name: Install dependencies
@@ -15,23 +15,23 @@ jobs:
if: github.ref == 'refs/heads/main'
steps:
- name: Checkout Repo
uses: actions/checkout@v4
uses: actions/checkout@v5
- uses: pnpm/action-setup@v3
- uses: pnpm/action-setup@v4
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v5
with:
node-version: "22"
cache: "pnpm"
- name: Setup Python
uses: actions/setup-python@v5
uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Install uv
uses: astral-sh/setup-uv@v3
uses: astral-sh/setup-uv@v7
- name: Install dependencies
run: pnpm install
+1 -1
View File
@@ -34,7 +34,7 @@ repos:
rev: v1.0.1
hooks:
- id: mypy
exclude: ^py/tests|^py/unit_tests
exclude: ^py/tests|^py/unit_tests|^examples
additional_dependencies:
[
"types-requests",
+21
View File
@@ -0,0 +1,21 @@
node_modules
package-lock.json
yarn.lock
.DS_Store
.cache
.env
.vercel
.output
.nitro
/build/
/api/
/server/build
/public/build# Sentry Config File
.env.sentry-build-plugin
/test-results/
/playwright-report/
/blob-report/
/playwright/.cache/
.tanstack
.vscode
+4
View File
@@ -0,0 +1,4 @@
**/build
**/public
pnpm-lock.yaml
routeTree.gen.ts
+88
View File
@@ -0,0 +1,88 @@
# LlamaClassify Demo
A TypeScript demo application showcasing the power of **LlamaClassify** - an agentic documents classification service from [LlamaCloud](https://cloud.llamaindex.ai). This demo allows you to classify financial documents among three different types (Cash flow statement, Income Statement and Balance Sheet).
## Table of Contents
- [Features](#features)
- [Prerequisites](#prerequisites)
- [Installation](#installation)
- [Usage](#usage)
- [Start the Demo](#start-the-demo)
- [How It Works](#how-it-works)
- [Troubleshooting](#troubleshooting)
- [Common Issues](#common-issues)
- [License](#license)
- [Contributing](#contributing)
## Features
- 📄 **Documemt Classification**: Classify files based on well-defined rules you can customized and play around with.
- 🤖 **Reasoning-based Actionable Insights**: Get in-depth, reasoning based insights on the document classification, accompanied by confidence scores.
- 🎨 **Beautiful UI**: [DaisyUI](https://daisyui.com)-based interface powered by [TanStack](https://tanstack.com)
-**Fast Development**: Hot reload support with development mode
- 🛠️ **TypeScript**: Full TypeScript support with strict type checking
## Prerequisites
- Node.js (version 22 or higher)
- pnpm package manager
- LlamaCloud API key
## Installation
1. Clone the repository:
```bash
git clone https://github.com/run-llama/llama_cloud_services
cd lama_cloud_services/examples-ts/classify/
```
2. Install dependencies:
```bash
npm install
```
3. Set up your environment variables:
```bash
# Add your API key to your environment
export LLAMA_CLOUD_API_KEY="your-llamacloud-api-key"
```
## Usage
### Start the Demo
```bash
npm run dev
```
The application will be up and running on http://localhost:3000
## How It Works
1. **Document Input**: Enter the path to your document when prompted
2. **Parsing**: LlamaClassify, based on the rules you can find [here](./src/utils/classifier.ts), processes the document and classifies it
3. **Results**: The classification outcome, as well as the reasoning behind it and the confidence score, are displayed in the UI.
## Troubleshooting
### Common Issues
1. **Module Resolution Errors**: Ensure you're using Node.js 22+ and have all dependencies installed
2. **API Key Issues**: Verify your LlamaCloud API key is correctly set
3. **File Path Errors**: Use absolute paths or ensure relative paths are correct from the project root
## License
MIT License - see the [LICENSE](../../LICENSE) file for details.
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Run `npm run format` and `npm run lint`
5. Submit a pull request
+34
View File
@@ -0,0 +1,34 @@
{
"name": "tanstack-start-example-basic",
"private": true,
"sideEffects": false,
"type": "module",
"scripts": {
"dev": "vite dev",
"build": "vite build && tsc --noEmit",
"start": "node .output/server/index.mjs"
},
"dependencies": {
"@tanstack/react-router": "^1.133.22",
"@tanstack/react-router-devtools": "^1.133.22",
"@tanstack/react-start": "^1.133.22",
"llama-cloud-services": "file:../../ts/llama_cloud_services",
"react": "^19.0.0",
"react-dom": "^19.0.0",
"tailwind-merge": "^2.6.0",
"zod": "^3.24.2"
},
"devDependencies": {
"@tailwindcss/postcss": "^4.1.15",
"@types/node": "^22.5.4",
"@types/react": "^19.0.8",
"@types/react-dom": "^19.0.3",
"@vitejs/plugin-react": "^4.6.0",
"daisyui": "^5.3.7",
"postcss": "^8.5.1",
"tailwindcss": "^4.1.15",
"typescript": "^5.7.2",
"vite": "^7.1.7",
"vite-tsconfig-paths": "^5.1.4"
}
}
+5
View File
@@ -0,0 +1,5 @@
export default {
plugins: {
'@tailwindcss/postcss': {},
},
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 21 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.8 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 862 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.0 KiB

@@ -0,0 +1,19 @@
{
"name": "",
"short_name": "",
"icons": [
{
"src": "/android-chrome-192x192.png",
"sizes": "192x192",
"type": "image/png"
},
{
"src": "/android-chrome-512x512.png",
"sizes": "512x512",
"type": "image/png"
}
],
"theme_color": "#ffffff",
"background_color": "#ffffff",
"display": "standalone"
}
@@ -0,0 +1,53 @@
import {
ErrorComponent,
Link,
rootRouteId,
useMatch,
useRouter,
} from '@tanstack/react-router'
import type { ErrorComponentProps } from '@tanstack/react-router'
export function DefaultCatchBoundary({ error }: ErrorComponentProps) {
const router = useRouter()
const isRoot = useMatch({
strict: false,
select: (state) => state.id === rootRouteId,
})
console.error('DefaultCatchBoundary Error:', error)
return (
<div className="min-w-0 flex-1 p-4 flex flex-col items-center justify-center gap-6">
<ErrorComponent error={error} />
<div className="flex gap-2 items-center flex-wrap">
<button
onClick={() => {
router.invalidate()
}}
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
>
Try Again
</button>
{isRoot ? (
<Link
to="/"
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
>
Home
</Link>
) : (
<Link
to="/"
className={`px-2 py-1 bg-gray-600 dark:bg-gray-700 rounded-sm text-white uppercase font-extrabold`}
onClick={(e) => {
e.preventDefault()
window.history.back()
}}
>
Go Back
</Link>
)}
</div>
</div>
)
}
@@ -0,0 +1,25 @@
import { Link } from '@tanstack/react-router'
export function NotFound({ children }: { children?: any }) {
return (
<div className="space-y-2 p-2">
<div className="text-gray-600 dark:text-gray-400">
{children || <p>The page you are looking for does not exist.</p>}
</div>
<p className="flex items-center gap-2 flex-wrap">
<button
onClick={() => window.history.back()}
className="bg-emerald-500 text-white px-2 py-1 rounded-sm uppercase font-black text-sm"
>
Go back
</button>
<Link
to="/"
className="bg-cyan-600 text-white px-2 py-1 rounded-sm uppercase font-black text-sm"
>
Start Over
</Link>
</p>
</div>
)
}
+225
View File
@@ -0,0 +1,225 @@
/* eslint-disable */
// @ts-nocheck
// noinspection JSUnusedGlobalSymbols
// This file was automatically generated by TanStack Router.
// You should NOT make any changes in this file as it will be overwritten.
// Additionally, you should also exclude this file from your linter and/or formatter to prevent it from being checked or modified.
import { Route as rootRouteImport } from './routes/__root'
import { Route as UsersRouteImport } from './routes/users'
import { Route as IndexRouteImport } from './routes/index'
import { Route as UsersIndexRouteImport } from './routes/users.index'
import { Route as PostsIndexRouteImport } from './routes/posts.index'
import { Route as UsersUserIdRouteImport } from './routes/users.$userId'
import { Route as PostsPostIdRouteImport } from './routes/posts.$postId'
import { Route as ApiClassifyRouteImport } from './routes/api/classify'
import { Route as PostsPostIdDeepRouteImport } from './routes/posts_.$postId.deep'
const UsersRoute = UsersRouteImport.update({
id: '/users',
path: '/users',
getParentRoute: () => rootRouteImport,
} as any)
const IndexRoute = IndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => rootRouteImport,
} as any)
const UsersIndexRoute = UsersIndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => UsersRoute,
} as any)
const PostsIndexRoute = PostsIndexRouteImport.update({
id: '/posts/',
path: '/posts/',
getParentRoute: () => rootRouteImport,
} as any)
const UsersUserIdRoute = UsersUserIdRouteImport.update({
id: '/$userId',
path: '/$userId',
getParentRoute: () => UsersRoute,
} as any)
const PostsPostIdRoute = PostsPostIdRouteImport.update({
id: '/posts/$postId',
path: '/posts/$postId',
getParentRoute: () => rootRouteImport,
} as any)
const ApiClassifyRoute = ApiClassifyRouteImport.update({
id: '/api/classify',
path: '/api/classify',
getParentRoute: () => rootRouteImport,
} as any)
const PostsPostIdDeepRoute = PostsPostIdDeepRouteImport.update({
id: '/posts_/$postId/deep',
path: '/posts/$postId/deep',
getParentRoute: () => rootRouteImport,
} as any)
export interface FileRoutesByFullPath {
'/': typeof IndexRoute
'/users': typeof UsersRouteWithChildren
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts': typeof PostsIndexRoute
'/users/': typeof UsersIndexRoute
'/posts/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRoutesByTo {
'/': typeof IndexRoute
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts': typeof PostsIndexRoute
'/users': typeof UsersIndexRoute
'/posts/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRoutesById {
__root__: typeof rootRouteImport
'/': typeof IndexRoute
'/users': typeof UsersRouteWithChildren
'/api/classify': typeof ApiClassifyRoute
'/posts/$postId': typeof PostsPostIdRoute
'/users/$userId': typeof UsersUserIdRoute
'/posts/': typeof PostsIndexRoute
'/users/': typeof UsersIndexRoute
'/posts_/$postId/deep': typeof PostsPostIdDeepRoute
}
export interface FileRouteTypes {
fileRoutesByFullPath: FileRoutesByFullPath
fullPaths:
| '/'
| '/users'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts'
| '/users/'
| '/posts/$postId/deep'
fileRoutesByTo: FileRoutesByTo
to:
| '/'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts'
| '/users'
| '/posts/$postId/deep'
id:
| '__root__'
| '/'
| '/users'
| '/api/classify'
| '/posts/$postId'
| '/users/$userId'
| '/posts/'
| '/users/'
| '/posts_/$postId/deep'
fileRoutesById: FileRoutesById
}
export interface RootRouteChildren {
IndexRoute: typeof IndexRoute
UsersRoute: typeof UsersRouteWithChildren
ApiClassifyRoute: typeof ApiClassifyRoute
PostsPostIdRoute: typeof PostsPostIdRoute
PostsIndexRoute: typeof PostsIndexRoute
PostsPostIdDeepRoute: typeof PostsPostIdDeepRoute
}
declare module '@tanstack/react-router' {
interface FileRoutesByPath {
'/users': {
id: '/users'
path: '/users'
fullPath: '/users'
preLoaderRoute: typeof UsersRouteImport
parentRoute: typeof rootRouteImport
}
'/': {
id: '/'
path: '/'
fullPath: '/'
preLoaderRoute: typeof IndexRouteImport
parentRoute: typeof rootRouteImport
}
'/users/': {
id: '/users/'
path: '/'
fullPath: '/users/'
preLoaderRoute: typeof UsersIndexRouteImport
parentRoute: typeof UsersRoute
}
'/posts/': {
id: '/posts/'
path: '/posts'
fullPath: '/posts'
preLoaderRoute: typeof PostsIndexRouteImport
parentRoute: typeof rootRouteImport
}
'/users/$userId': {
id: '/users/$userId'
path: '/$userId'
fullPath: '/users/$userId'
preLoaderRoute: typeof UsersUserIdRouteImport
parentRoute: typeof UsersRoute
}
'/posts/$postId': {
id: '/posts/$postId'
path: '/posts/$postId'
fullPath: '/posts/$postId'
preLoaderRoute: typeof PostsPostIdRouteImport
parentRoute: typeof rootRouteImport
}
'/api/classify': {
id: '/api/classify'
path: '/api/classify'
fullPath: '/api/classify'
preLoaderRoute: typeof ApiClassifyRouteImport
parentRoute: typeof rootRouteImport
}
'/posts_/$postId/deep': {
id: '/posts_/$postId/deep'
path: '/posts/$postId/deep'
fullPath: '/posts/$postId/deep'
preLoaderRoute: typeof PostsPostIdDeepRouteImport
parentRoute: typeof rootRouteImport
}
}
}
interface UsersRouteChildren {
UsersUserIdRoute: typeof UsersUserIdRoute
UsersIndexRoute: typeof UsersIndexRoute
}
const UsersRouteChildren: UsersRouteChildren = {
UsersUserIdRoute: UsersUserIdRoute,
UsersIndexRoute: UsersIndexRoute,
}
const UsersRouteWithChildren = UsersRoute._addFileChildren(UsersRouteChildren)
const rootRouteChildren: RootRouteChildren = {
IndexRoute: IndexRoute,
UsersRoute: UsersRouteWithChildren,
ApiClassifyRoute: ApiClassifyRoute,
PostsPostIdRoute: PostsPostIdRoute,
PostsIndexRoute: PostsIndexRoute,
PostsPostIdDeepRoute: PostsPostIdDeepRoute,
}
export const routeTree = rootRouteImport
._addFileChildren(rootRouteChildren)
._addFileTypes<FileRouteTypes>()
import type { getRouter } from './router.tsx'
import type { createStart } from '@tanstack/react-start'
declare module '@tanstack/react-start' {
interface Register {
ssr: true
router: Awaited<ReturnType<typeof getRouter>>
}
}
+15
View File
@@ -0,0 +1,15 @@
import { createRouter } from '@tanstack/react-router'
import { routeTree } from './routeTree.gen'
import { DefaultCatchBoundary } from './components/DefaultCatchBoundary'
import { NotFound } from './components/NotFound'
export function getRouter() {
const router = createRouter({
routeTree,
defaultPreload: 'intent',
defaultErrorComponent: DefaultCatchBoundary,
defaultNotFoundComponent: () => <NotFound />,
scrollRestoration: true,
})
return router
}
+128
View File
@@ -0,0 +1,128 @@
/// <reference types="vite/client" />
import {
HeadContent,
Scripts,
createRootRoute,
} from '@tanstack/react-router'
import * as React from 'react'
import { DefaultCatchBoundary } from '~/components/DefaultCatchBoundary'
import { NotFound } from '~/components/NotFound'
import { seo } from '~/utils/seo'
export const Route = createRootRoute({
head: () => ({
meta: [
{
charSet: 'utf-8',
},
{
name: 'viewport',
content: 'width=device-width, initial-scale=1',
},
...seo({
title:
'Financial Documents Classification Agent',
description: `Classify financial documents as balance sheets, income statements and cash flow statemets. `,
}),
],
links: [
{ rel: 'stylesheet', href: "https://cdn.jsdelivr.net/npm/daisyui@5" },
{
rel: 'apple-touch-icon',
sizes: '180x180',
href: '/apple-touch-icon.png',
},
{
rel: 'icon',
type: 'image/png',
sizes: '32x32',
href: '/favicon-32x32.png',
},
{
rel: 'icon',
type: 'image/png',
sizes: '16x16',
href: '/favicon-16x16.png',
},
{ rel: 'manifest', href: '/site.webmanifest', color: '#fffff' },
{ rel: 'icon', href: '/favicon.ico' },
],
scripts: [
{
src: '/customScript.js',
type: 'text/javascript',
},
{
src: "https://cdn.jsdelivr.net/npm/@tailwindcss/browser@4",
type: "text/javascript",
}
],
}),
errorComponent: DefaultCatchBoundary,
notFoundComponent: () => <NotFound />,
shellComponent: RootDocument,
})
function RootDocument({ children }: { children: React.ReactNode }) {
return (
<html>
<head>
<HeadContent />
</head>
<body>
<div className="navbar bg-base-100 shadow-sm">
<div className="navbar-start">
<div className="dropdown">
<div tabIndex={0} role="button" className="btn btn-ghost btn-circle">
<svg
xmlns="http://www.w3.org/2000/svg"
className="h-5 w-5"
fill="none"
viewBox="0 0 24 24"
stroke="currentColor"
>
<path
strokeLinecap="round"
strokeLinejoin="round"
strokeWidth="2"
d="M4 6h16M4 12h16M4 18h7"
/>
</svg>
</div>
<ul
tabIndex={0}
className="menu menu-lg dropdown-content bg-base-100 rounded-box z-1 mt-3 w-80 p-2 shadow"
>
<li><a href="/">Home</a></li>
<li><a href="https://cloud.llamaindex.ai">Get Started with LlamaCloud</a></li>
<li><a href="https://developers.llamaindex.ai/python/cloud/llamaclassify/getting_started/">LlamaClassify Docs</a></li>
</ul>
</div>
</div>
<div className="navbar-center">
<a className="btn btn-ghost text-xl" href="/">Financial Documents Classification Agent</a>
</div>
<div className="navbar-end">
<a href="https://github.com/run-llama/llama_cloud_services/main/blob/examples-ts/classify">
<button className="btn btn-ghost btn-circle">
<div className="indicator">
<svg
xmlns="http://www.w3.org/2000/svg"
className="h-10 w-10"
fill="currentColor"
viewBox="0 0 640 512"
>
<path d="M237.9 461.4C237.9 463.4 235.6 465 232.7 465C229.4 465.3 227.1 463.7 227.1 461.4C227.1 459.4 229.4 457.8 232.3 457.8C235.3 457.5 237.9 459.1 237.9 461.4zM206.8 456.9C206.1 458.9 208.1 461.2 211.1 461.8C213.7 462.8 216.7 461.8 217.3 459.8C217.9 457.8 216 455.5 213 454.6C210.4 453.9 207.5 454.9 206.8 456.9zM251 455.2C248.1 455.9 246.1 457.8 246.4 460.1C246.7 462.1 249.3 463.4 252.3 462.7C255.2 462 257.2 460.1 256.9 458.1C256.6 456.2 253.9 454.9 251 455.2zM316.8 72C178.1 72 72 177.3 72 316C72 426.9 141.8 521.8 241.5 555.2C254.3 557.5 258.8 549.6 258.8 543.1C258.8 536.9 258.5 502.7 258.5 481.7C258.5 481.7 188.5 496.7 173.8 451.9C173.8 451.9 162.4 422.8 146 415.3C146 415.3 123.1 399.6 147.6 399.9C147.6 399.9 172.5 401.9 186.2 425.7C208.1 464.3 244.8 453.2 259.1 446.6C261.4 430.6 267.9 419.5 275.1 412.9C219.2 406.7 162.8 398.6 162.8 302.4C162.8 274.9 170.4 261.1 186.4 243.5C183.8 237 175.3 210.2 189 175.6C209.9 169.1 258 202.6 258 202.6C278 197 299.5 194.1 320.8 194.1C342.1 194.1 363.6 197 383.6 202.6C383.6 202.6 431.7 169 452.6 175.6C466.3 210.3 457.8 237 455.2 243.5C471.2 261.2 481 275 481 302.4C481 398.9 422.1 406.6 366.2 412.9C375.4 420.8 383.2 435.8 383.2 459.3C383.2 493 382.9 534.7 382.9 542.9C382.9 549.4 387.5 557.3 400.2 555C500.2 521.8 568 426.9 568 316C568 177.3 455.5 72 316.8 72zM169.2 416.9C167.9 417.9 168.2 420.2 169.9 422.1C171.5 423.7 173.8 424.4 175.1 423.1C176.4 422.1 176.1 419.8 174.4 417.9C172.8 416.3 170.5 415.6 169.2 416.9zM158.4 408.8C157.7 410.1 158.7 411.7 160.7 412.7C162.3 413.7 164.3 413.4 165 412C165.7 410.7 164.7 409.1 162.7 408.1C160.7 407.5 159.1 407.8 158.4 408.8zM190.8 444.4C189.2 445.7 189.8 448.7 192.1 450.6C194.4 452.9 197.3 453.2 198.6 451.6C199.9 450.3 199.3 447.3 197.3 445.4C195.1 443.1 192.1 442.8 190.8 444.4zM179.4 429.7C177.8 430.7 177.8 433.3 179.4 435.6C181 437.9 183.7 438.9 185 437.9C186.6 436.6 186.6 434 185 431.7C183.6 429.4 181 428.4 179.4 429.7z" />
</svg>
</div>
</button>
</a>
</div>
</div>
<hr />
{children}
<Scripts />
</body>
</html>
)
}
@@ -0,0 +1,45 @@
import { createFileRoute } from '@tanstack/react-router'
import { classifier, classificationRules, parsingConfig } from '~/utils/classifier'
export const Route = createFileRoute('/api/classify')({
component: RouteComponent,
server: {
handlers: {
POST: async ({ request }) => {
const body = await request.formData()
const fl = body.get("file") as File;
if (!fl) {
return new Response(JSON.stringify({"result": "you need to provide a file"}))
}
const buff = await fl.arrayBuffer()
const rawRes = await classifier.classify(
classificationRules,
parsingConfig,
{ fileContents: [new Uint8Array(buff)] },
)
const results = rawRes.items
let classification = ""
for (const result of results) {
if ("result" in result && result.result) {
classification += `
<div class="card bg-base-100 shadow-xl p-6 mb-4">
<div class="space-y-3">
<p><span class="font-semibold">📄 Document:</span> ${fl.name}</p>
<p><span class="font-semibold">🏷️ Type:</span> <span class="badge badge-primary">${result.result.type}</span></p>
<p><span class="font-semibold">📊 Confidence:</span> ${result.result.confidence*100}%</p>
<p><span class="font-semibold">💭 Reasoning:</span> ${result.result.reasoning}</p>
</div>
</div>
`
}
}
return new Response(JSON.stringify({"result": classification}))
},
},
},
})
function RouteComponent() {
return
}
+99
View File
@@ -0,0 +1,99 @@
import { createFileRoute } from '@tanstack/react-router'
import { useRef, useState } from 'react'
export const Route = createFileRoute('/')({
component: Home,
})
function Home() {
const [file, setFile] = useState<null | File>(null)
const fileInputRef = useRef<HTMLInputElement>(null)
const [reply, setReply] = useState<null | string>(null)
const [loading, setLoading] = useState<boolean>(false)
const handleFileChange = (event: React.ChangeEvent<HTMLInputElement>) => {
const selectedFile = event.target.files?.[0]
if (selectedFile) {
setFile(selectedFile)
}
}
const handleClearFile = () => {
if (file) {
setFile(null)
}
if (fileInputRef.current) {
fileInputRef.current.value = ''
}
if (reply) {
setReply(null)
}
}
const handleClassify = async () => {
if (!file) return
if (reply) {
setReply(null)
}
setLoading(true)
try {
const formData = new FormData()
formData.append('file', file)
const res = await fetch('/api/classify', {
method: 'POST',
body: formData,
})
const data = await res.json()
setReply(data.result)
} catch (error) {
console.error('Error:', error)
} finally {
setLoading(false)
}
}
return (
<div className="flex flex-col justify-center items-center gap-y-8">
<br />
<h1 className="text-xl font-bold text-gray-700">AI-Powered finacial document classification</h1>
<h2 className="text-lg font-semibold text-gray-500">Need help sorting out the financial documents jungle? Let our classification agent handle it!</h2>
<fieldset className="fieldset bg-base-100 border-base-300 rounded-box w-200 border p-4">
<legend className="fieldset-legend text-lg">Upload your financial document here</legend>
<label className="label flex justify-center">
<input type="file" className="file-input" onChange={handleFileChange} accept='application/pdf' ref={fileInputRef} />
</label>
</fieldset>
{file && (
<div className="flex flex-col justify-center items-center gap-y-8">
<p className="text-sm text-gray-600">Selected file: {file.name}</p>
<div className='grid grid-cols-2 gap-x-6'>
<button
type="button"
className='btn bg-gray-500 text-white shadow-lg hover:bg-gray-600 hover:shadow-xl rounded'
onClick={handleClassify}
>
Classify
</button>
<button
onClick={handleClearFile}
type="button"
className="px-4 py-2 bg-red-300 text-black rounded hover:bg-red-400 hover:shadow-xl shadow-lg"
>
Clear
</button>
</div>
</div>
)}
{loading && (
<span className="loading loading-spinner text-primary"></span>
)}
{reply && (
<div
className="max-w-2xl w-full"
dangerouslySetInnerHTML={{ __html: reply }}
/>
)}
</div>
)
}
@@ -0,0 +1,23 @@
import { LlamaClassify, ClassifierRule, ClassifyParsingConfiguration } from "llama-cloud-services"
export const classifier = new LlamaClassify(process.env.LLAMA_CLOUD_API_KEY);
export const classificationRules: ClassifierRule[] = [
{
description: "Shows a company's assets, liabilities, and shareholders' equity at a specific point in time, providing a snapshot of financial position.",
type: "balance_sheet"
},
{
description: "Reports cash inflows and outflows from operating, investing, and financing activities, highlighting liquidity and cash management.",
type: "cash_flow_statement"
},
{
description: "Summarizes revenues, expenses, and profits over a period, indicating financial performance and profitability.",
type: "income_statement"
},
];
export const parsingConfig: ClassifyParsingConfiguration = {
lang: "en",
max_pages: 20,
}
+33
View File
@@ -0,0 +1,33 @@
export const seo = ({
title,
description,
keywords,
image,
}: {
title: string
description?: string
image?: string
keywords?: string
}) => {
const tags = [
{ title },
{ name: 'description', content: description },
{ name: 'keywords', content: keywords },
{ name: 'twitter:title', content: title },
{ name: 'twitter:description', content: description },
{ name: 'twitter:creator', content: '@tannerlinsley' },
{ name: 'twitter:site', content: '@tannerlinsley' },
{ name: 'og:type', content: 'website' },
{ name: 'og:title', content: title },
{ name: 'og:description', content: description },
...(image
? [
{ name: 'twitter:image', content: image },
{ name: 'twitter:card', content: 'summary_large_image' },
{ name: 'og:image', content: image },
]
: []),
]
return tags
}
+22
View File
@@ -0,0 +1,22 @@
{
"include": ["**/*.ts", "**/*.tsx"],
"compilerOptions": {
"strict": true,
"esModuleInterop": true,
"jsx": "react-jsx",
"module": "ESNext",
"moduleResolution": "Bundler",
"lib": ["DOM", "DOM.Iterable", "ES2022"],
"isolatedModules": true,
"resolveJsonModule": true,
"skipLibCheck": true,
"target": "ES2022",
"allowJs": true,
"forceConsistentCasingInFileNames": true,
"baseUrl": ".",
"paths": {
"~/*": ["./src/*"]
},
"noEmit": true
}
}
+19
View File
@@ -0,0 +1,19 @@
import { tanstackStart } from '@tanstack/react-start/plugin/vite'
import { defineConfig } from 'vite'
import tsConfigPaths from 'vite-tsconfig-paths'
import viteReact from '@vitejs/plugin-react'
export default defineConfig({
server: {
port: 3000,
},
plugins: [
tsConfigPaths({
projects: ['./tsconfig.json'],
}),
tanstackStart({
srcDirectory: 'src',
}),
viteReact(),
],
})
Binary file not shown.

After

Width:  |  Height:  |  Size: 287 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 769 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 942 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.5 MiB

+508
View File
@@ -0,0 +1,508 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "a7oq3cfnync",
"metadata": {},
"source": [
"# Extracting Repeating Entities from Documents\n",
"\n",
"This notebook demonstrates how to use the `PER_TABLE_ROW` extraction target to extract structured data from documents containing repeating entities like tables, lists, or catalogs.\n",
"\n",
"## Why Use the Tabular Extraction Target?\n",
"\n",
"`PER_DOC` (refer to the table below for a quick overview of the different extraction targets) is the default extraction target in LlamaExtract, which looks at the entire document's context when doing an extraction. When extracting lists of entities, LLM-based extraction has a critical failure mode — it often **only extracts the first few tens of entries** from a long list. This happens because LLMs have limited attention spans for repetitive data. Document-level extraction doesn't guarantee exhaustive coverage, and long lists lead to incomplete extractions.\n",
"\n",
"**The Solution**: `PER_TABLE_ROW` solves this by processing each entity individually or in smaller batches, ensuring **exhaustive extraction** of all entries regardless of list length.\n",
"\n",
"### Entity-Level Extraction\n",
"\n",
"When using `extraction_target=ExtractTarget.PER_TABLE_ROW`, you define a schema for a **single entity** (e.g., one hospital, one product, one invoice line item), not the full document. LlamaExtract automatically:\n",
"- Detects the formatting patterns that distinguish individual entities (table rows, list items, section headers, etc.)\n",
"- Applies your schema to each identified entity\n",
"- Returns a `list[YourSchema]` with one object per entity\n",
"\n",
"This approach is ideal when each entity locally contains all the information needed for your schema.\n",
"\n",
"### Choosing the Right Extraction Target\n",
"\n",
"| Extraction Target | Best For | Returns |\n",
"|-------------------|----------|---------|\n",
"| `PER_DOC` | Single-entity documents, summaries, or short lists | One JSON object for entire document |\n",
"| `PER_PAGE` | Multi-page documents where each page is independent | One JSON object per page |\n",
"| `PER_TABLE_ROW` | **Long lists, tables, catalogs with repeating entities** | List of JSON objects (one per entity) |\n",
"\n",
"📖 For more details, see the [Extraction Target documentation](https://developers.llamaindex.ai/python/cloud/llamaextract/features/concepts/#extraction-target)."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "9427d1de",
"metadata": {},
"outputs": [],
"source": [
"from dotenv import load_dotenv\n",
"from llama_cloud_services import LlamaExtract\n",
"\n",
"\n",
"# Load environment variables (put LLAMA_CLOUD_API_KEY in your .env file)\n",
"load_dotenv(override=True)\n",
"\n",
"# Optionally, add your project id/organization id\n",
"llama_extract = LlamaExtract()"
]
},
{
"cell_type": "markdown",
"id": "4426b360",
"metadata": {},
"source": [
"## Table of Hospitals by County and Insurance Plans\n",
"\n",
"We have a PDF document with a list of hospitals by county and different insurance plans offered by Blue Shield of California. \n",
"\n",
"\n",
"![First few entries from the PDF](./data/tables/bsc_page1.png)"
]
},
{
"cell_type": "markdown",
"id": "c86sjymhn1r",
"metadata": {},
"source": [
"We want to extract each hospital from this table along with a list of applicable insurance plans. \n",
"\n",
"### Example 1: Structured Table\n",
"\n",
"This is an ideal use case for `PER_TABLE_ROW` extraction:\n",
"- **Clear structure**: The document has explicit table formatting with rows and columns\n",
"- **Repeating entities**: Each row represents one hospital with consistent attributes\n",
"- **Local information**: All data for each hospital (county, name, plans) is contained within its row\n",
"\n",
"Notice that our `Hospital` schema describes a **single hospital**, not the full document. LlamaExtract will return a `list[Hospital]` with one entry per table row."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "7c61a802",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"\n",
"\n",
"class Hospital(BaseModel):\n",
" \"\"\"List of hospitals by county available for different BSC plans\"\"\"\n",
"\n",
" county: str = Field(description=\"County name\")\n",
" hospital_name: str = Field(description=\"Name of the hospital\")\n",
" plan_names: list[str] = Field(\n",
" description=\"List of plans available at the hospital. One of: Trio HMO, SaveNet, Access+ HMO, BlueHPN PPO, Tandem PPO, PPO\"\n",
" )"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "b8a69b7a",
"metadata": {},
"outputs": [],
"source": [
"from llama_cloud_services.extract import ExtractConfig, ExtractMode, ExtractTarget\n",
"\n",
"\n",
"result = await llama_extract.aextract(\n",
" data_schema=Hospital,\n",
" files=\"./data/tables/BSC-Hospital-List-by-County.pdf\",\n",
" config=ExtractConfig(\n",
" extraction_mode=ExtractMode.PREMIUM,\n",
" extraction_target=ExtractTarget.PER_TABLE_ROW,\n",
" parse_model=\"anthropic-sonnet-4.5\",\n",
" ),\n",
")"
]
},
{
"cell_type": "markdown",
"id": "43722cda",
"metadata": {},
"source": [
"### Results"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "95b5aca6",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"380"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"len(result.data)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "1e355770",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[{'county': 'Alameda',\n",
" 'hospital_name': 'Alameda Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Med Ctr Herrick Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Med Ctr Alta Bates Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Med Ctr Summit Campus',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Alta Bates Summit Medical Center',\n",
" 'plan_names': ['Trio HMO',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'BHC Fremont Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Centre For Neuro Skills San Francisco',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Eden Medical Center',\n",
" 'plan_names': ['Trio HMO', 'Access+ HMO', 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Fairmont Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']},\n",
" {'county': 'Alameda',\n",
" 'hospital_name': 'Highland Hospital',\n",
" 'plan_names': ['Trio HMO',\n",
" 'SaveNet',\n",
" 'Access+ HMO',\n",
" 'BlueHPN PPO',\n",
" 'Tandem PPO',\n",
" 'PPO']}]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"result.data[:10]"
]
},
{
"cell_type": "markdown",
"id": "e28f0de8",
"metadata": {},
"source": [
"![](./data/tables/bsc_results.png)"
]
},
{
"cell_type": "markdown",
"id": "di156pb7s6j",
"metadata": {},
"source": [
"**Success!** We extracted all **380 hospitals** from the multi-page PDF. Each entity was correctly parsed with its county, hospital name, and applicable insurance plans. With `PER_DOC`, we would likely have only gotten the first 20-30 entries."
]
},
{
"cell_type": "markdown",
"id": "gelvl6db268",
"metadata": {},
"source": [
"## Extracting from a Toy Catalog\n",
"\n",
"### Example 2: Semi-Structured List\n",
"\n",
"The `PER_TABLE_ROW` extraction target also works well for documents that aren't explicit tables but have similar properties:\n",
"- **Ordered listing**: The toys are listed sequentially with visual separation (section headers, spacing)\n",
"- **Repeating pattern**: Each toy entry has a consistent structure (code, name, specs, description)\n",
"- **Local information**: All attributes for each toy are grouped together in its entry\n",
"\n",
"Even though this isn't a traditional table format, each toy entity locally contains all the information needed for our schema. LlamaExtract detects the formatting patterns that distinguish each toy and extracts them as separate entities.\n",
"\n",
"![](./data/tables/toy_catalog_page.png)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "8cf0b2db",
"metadata": {},
"outputs": [],
"source": [
"from pydantic import BaseModel, Field\n",
"\n",
"\n",
"class ToyCatalog(BaseModel):\n",
" \"\"\"Product information from a toy catalog.\"\"\"\n",
"\n",
" section_name: str = Field(\n",
" description=\"The name of the toy section (e.g. Table Toys, Active Toys).\"\n",
" )\n",
" product_code: str = Field(\n",
" description=\"The unique product code for the toy (e.g., GA457).\"\n",
" )\n",
" toy_name: str = Field(description=\"The name of the toy.\")\n",
" age_range: str = Field(\n",
" description=\"The recommended age range for the toy (e.g., 6 +, 4 +).\",\n",
" )\n",
" player_range: str = Field(\n",
" description=\"The number of players the toy is designed for (e.g., 2, 2-4, 1-6).\",\n",
" )\n",
" material: str = Field(\n",
" description=\"The primary material(s) the toy is made of (e.g., wood, cardboard).\",\n",
" )\n",
" description: str = Field(\n",
" description=\"A brief description of the toy and its components and dimensions.\",\n",
" )"
]
},
{
"cell_type": "markdown",
"id": "mysu1i2qo9e",
"metadata": {},
"source": [
"### Results\n",
"\n",
"Again, our schema represents a **single toy product**, not the entire catalog. The system will return a `list[ToyCatalog]` with one entry per toy."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "5b38b806",
"metadata": {},
"outputs": [],
"source": [
"result = await llama_extract.aextract(\n",
" data_schema=ToyCatalog,\n",
" files=\"./data/tables/Click-BS-Toys-Catalogue-2024.pdf\",\n",
" config=ExtractConfig(\n",
" extraction_mode=ExtractMode.PREMIUM,\n",
" extraction_target=ExtractTarget.PER_TABLE_ROW,\n",
" parse_model=\"anthropic-sonnet-4.5\",\n",
" ),\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "91aface0",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"153"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"len(result.data)"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "51278736",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"[{'section_name': 'Table Toys',\n",
" 'product_code': 'GA457',\n",
" 'toy_name': 'Dots and Boxes',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': 'base 17x17 cm\\n50 border pieces 4x1,2x0,3 cm\\n34 trees 2,6x1,4 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA456',\n",
" 'toy_name': '3 In a Row',\n",
" 'age_range': '8+',\n",
" 'player_range': '2',\n",
" 'material': 'wood, pine, cardboard',\n",
" 'description': 'base 24x22,5x2,5 cm\\n30 cards 5,5x5 cm\\n6 chips'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA467',\n",
" 'toy_name': 'Which Cow am i?',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood, beech',\n",
" 'description': '2 cow bases 56x4x4,5 cm\\n16 cards 4x5 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA460',\n",
" 'toy_name': 'Balance Bunnies',\n",
" 'age_range': '4+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': '1 base 35x12x25 cm\\n7 bunnies 7 foxes\\n1 dice 3 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA462',\n",
" 'toy_name': 'Color Combination Race',\n",
" 'age_range': '4+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood, cardboard',\n",
" 'description': 'base 6,5x6,5x15 cm, rings 5,5x5,5x0,5 mm\\ncardholder 6x6x2 cm, cards 5,5x5,5 cm\\ncolor cards Ø 15,5 cm - Ø 7 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA465',\n",
" 'toy_name': 'Plop It',\n",
" 'age_range': '6+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood, elastic, cardboard',\n",
" 'description': 'Catch the right balls and plop them in the net!\\n* 2 ploppers 8x5 cm\\n* 2 net holders Ø 5cm, length 55 cm\\n* 6 cards 1,5x2,5 cm, 30 balls Ø 2,5 cm\\n* 1 rope 120 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA466',\n",
" 'toy_name': 'Whack a Shape',\n",
" 'age_range': '4+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood',\n",
" 'description': '* base 38,5x15,5 cm\\n* 2 stands 36 half balls, 4 hammers\\n* 1 dice 2,5 cm\\n* 4 cards'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA458',\n",
" 'toy_name': 'Sling Puck | Table Hockey',\n",
" 'age_range': '6+',\n",
" 'player_range': '2',\n",
" 'material': 'wood',\n",
" 'description': '* double sides base 39x21x3 cm\\n* 10 chips Ø 2,5 cm\\n* 2 pushers 4x4x3 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA039',\n",
" 'toy_name': 'DIY Birdhouse',\n",
" 'age_range': '3+',\n",
" 'player_range': '1',\n",
" 'material': 'wood',\n",
" 'description': '* house 9x9x13 cm'},\n",
" {'section_name': 'Table Toys',\n",
" 'product_code': 'GA319',\n",
" 'toy_name': 'Triangle Domino',\n",
" 'age_range': '6+',\n",
" 'player_range': '2-4',\n",
" 'material': 'wood',\n",
" 'description': '* 35 triangles 10x10 x10 cm'}]"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"result.data[:10]"
]
},
{
"cell_type": "markdown",
"id": "d1810c0a",
"metadata": {},
"source": [
"![](./data/tables/toy_catalog_results.png)"
]
},
{
"cell_type": "markdown",
"id": "ezur9gnhmsb",
"metadata": {},
"source": [
"**Success!** Despite the semi-structured format, we extracted all **152 toy products** from the catalog (there's an extra repeated extracted toy from the Appendix section). LlamaExtract automatically detected the visual patterns separating each toy entry and applied our schema to each one."
]
},
{
"cell_type": "markdown",
"id": "aeyr3io29u",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"The `PER_TABLE_ROW` extraction target is powerful for extracting repeating structured entities from documents. Key takeaways:\n",
"\n",
"1. **Schema design**: Define your schema for a single entity, not the full document. The system returns `list[YourSchema]`.\n",
"\n",
"2. **Works with various formats**: Not just traditional tables—any document with distinguishable repeating entities (bullets, numbering, headers, visual separation, etc.). The common requirement is that each entity should contain all the necessary data for your schema within its local context.\n",
"\n",
"3. **Automatic pattern detection**: LlamaExtract identifies the formatting patterns that distinguish entities and applies your schema to each one."
]
}
],
"metadata": {
"kernelspec": {
"display_name": ".venv",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
@@ -4,31 +4,19 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"# Complete Parse → Classify → Extract Workflow with LlamaCloud Services\n",
"# Document Classification + Extraction Workflow with LlamaCloud + LlamaIndex Workflows\n",
"\n",
"This notebook demonstrates the complete workflow for processing documents using LlamaCloud services:\n",
"1. **Parse** - Extract and convert documents to markdown\n",
"<a href=\"https://colab.research.google.com/github/run-llama/llama_cloud_services/blob/main/examples/misc/parse_classify_extract_workflow.ipynb\" target=\"_parent\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" alt=\"Open In Colab\"/></a>\n",
"\n",
"This notebook shows a multi-step agentic document workflow that uses the **parsing**, **classification** and **extraction** modules in LlamaCloud, orchestrated through **LlamaIndex Workflows**. The workflow can take in a complex input document, parse it into clean markdown, classify it according to its subtype, and extract data according to a specified schema for that subtype. This allows you to automate document extraction of various types within the same workflow instead of having to manually separate the data beforehand. \n",
"\n",
"This notebook uses the following modules:\n",
"1. **Parse (LlamaParse)** - Extract and convert documents to markdown\n",
"2. **Classify** - Categorize documents based on their content\n",
"3. **Extract** - Extract structured data using the markdown as input via SourceText\n",
"3. **Extract (LlamaExtract)** - Extract structured data using the markdown as input via SourceText\n",
"4. **LlamaIndex Workflows** - Event-driven orchestration of the parse, classify and extract steps\n",
"\n",
"## Overview of the Workflow\n",
"\n",
"### 1. Parse Phase\n",
"- Use `LlamaParse` to convert documents (PDFs, Word docs, etc.) into structured formats\n",
"- Extract markdown content that preserves document structure\n",
"- Get both raw text and markdown representations\n",
"\n",
"### 2. Classify Phase\n",
"- Use `ClassifyClient` to categorize documents based on content\n",
"- Apply classification rules to route documents appropriately\n",
"- Handle different document types with specific processing logic\n",
"\n",
"### 3. Extract Phase\n",
"- Use `LlamaExtract` with `SourceText` to extract structured data\n",
"- Pass the markdown content as input for more accurate extraction\n",
"- Define custom schemas for structured data extraction\n",
"\n",
"Let's walk through each step with practical examples."
"The workflow is implemented as a proper LlamaIndex Workflow with separate steps for parsing, classification, and extraction, connected by typed events. This provides modularity, observability, and type safety."
]
},
{
@@ -45,8 +33,8 @@
"outputs": [],
"source": [
"# Install required packages\n",
"!pip install llama-cloud-services\n",
"!pip install python-dotenv"
"%pip install llama-cloud-services\n",
"%pip install python-dotenv"
]
},
{
@@ -73,7 +61,7 @@
"nest_asyncio.apply()\n",
"\n",
"# Set up API key\n",
"os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"\" # edit it\n",
"# os.environ[\"LLAMA_CLOUD_API_KEY\"] = \"\" # edit it\n",
"\n",
"# Setup Base URL\n",
"# os.envrion[\"LLAMA_CLOUD_BASE_URL\"] = \"https://api.cloud.eu.llamaindex.ai/\" # update if necessay\n",
@@ -99,7 +87,8 @@
"name": "stdout",
"output_type": "stream",
"text": [
"📁 financial_report.pdf already exists\n",
"Downloading financial_report.pdf...\n",
"✅ Downloaded financial_report.pdf\n",
"📁 technical_spec.pdf already exists\n",
"\n",
"📂 Sample documents ready!\n"
@@ -115,7 +104,7 @@
"\n",
"# Download sample documents\n",
"docs_to_download = {\n",
" \"financial_report.pdf\": \"https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/10k/uber_2021.pdf\",\n",
" \"financial_report.pdf\": \"https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/10k/uber_2021.pdf\",\n",
" \"technical_spec.pdf\": \"https://www.ti.com/lit/ds/symlink/lm317.pdf\",\n",
"}\n",
"\n",
@@ -155,10 +144,10 @@
"output_type": "stream",
"text": [
"🔄 Parsing documents...\n",
"Started parsing the file under job_id 8a8c76f9-354d-4275-91d8-312ff1adc762\n",
"...✅ Parsed financial report (Job ID: 8a8c76f9-354d-4275-91d8-312ff1adc762)\n",
"Started parsing the file under job_id 7e603448-ed80-4d18-948b-6801ed51c41b\n",
"✅ Parsed technical spec (Job ID: 7e603448-ed80-4d18-948b-6801ed51c41b)\n",
"Started parsing the file under job_id 530c187a-bd2d-4eea-b38d-9e5738eab465\n",
".✅ Parsed financial report (Job ID: 530c187a-bd2d-4eea-b38d-9e5738eab465)\n",
"Started parsing the file under job_id a6e27710-776b-4445-8b94-8d75959ff5db\n",
"✅ Parsed technical spec (Job ID: a6e27710-776b-4445-8b94-8d75959ff5db)\n",
"\n",
"📄 Parsing complete!\n"
]
@@ -246,23 +235,23 @@
"\n",
"## 1 Features\n",
"\n",
" Output voltage range:\n",
"- Output voltage range:\n",
" Adjustable: 1.25V to 37V\n",
" Output current: 1.5A\n",
" Line regulation: 0.01%/V (typ)\n",
" Load regulation: 0.1% (typ)\n",
" Internal short-circuit current limiting\n",
" Thermal overload protection\n",
" Output safe-area compensation (new chip)\n",
" PSRR: 80dB at 120Hz for CADJ = 10μF (new chip)\n",
" Packages:\n",
"- Output current: 1.5A\n",
"- Line regulation: 0.01%/V (typ)\n",
"- Load regulation: 0.1% (typ)\n",
"- Internal short-circuit current limiting\n",
"- Thermal overload protection\n",
"- Output safe-area compensation (new chip)\n",
"- PSRR: 80dB at 120Hz for CADJ = 10μF (new chip)\n",
"- Packages:\n",
" 4-pin, SOT-223 (DCY)\n",
" 3-pin, TO-263 (KTT)\n",
" 3-pin, TO-220 (KCS, KCT),\n",
"...\n",
"\n",
"📏 Financial report markdown length: 1348671 characters\n",
"📏 Technical spec markdown length: 90971 characters\n"
"📏 Financial report markdown length: 1338499 characters\n",
"📏 Technical spec markdown length: 92483 characters\n"
]
}
],
@@ -291,7 +280,7 @@
"source": [
"## Phase 2: Document Classification\n",
"\n",
"Next, let's classify our documents based on their content using the ClassifyClient."
"Next, let's classify our documents based on their content using `LlamaClassify`."
]
},
{
@@ -309,14 +298,14 @@
}
],
"source": [
"from llama_cloud_services.beta.classifier.client import ClassifyClient\n",
"from llama_cloud_services.beta.classifier.client import LlamaClassify\n",
"from llama_cloud.types import ClassifierRule\n",
"from llama_cloud_services.files.client import FileClient\n",
"from llama_cloud.client import AsyncLlamaCloud\n",
"\n",
"# Initialize the classify client\n",
"api_key = os.environ[\"LLAMA_CLOUD_API_KEY\"]\n",
"classify_client = ClassifyClient.from_api_key(api_key)\n",
"classify_client = LlamaClassify.from_api_key(api_key)\n",
"\n",
"print(\"🏷️ Setting up document classification...\")\n",
"\n",
@@ -339,6 +328,72 @@
"print(f\"📝 Created {len(classification_rules)} classification rules\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Try Classification Independently\n",
"\n",
"Let's test the classification on one of our parsed documents to see how it works:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"🔍 Classifying financial document...\n",
" Document length: 1,338,499 characters\n",
"\n",
"✅ Classification Result:\n",
" Type: financial_document\n",
" Confidence: 100.00%\n",
" Reasoning: This document is a Form 10-K, which is an annual report required by the U.S. Securities and Exchange Commission (SEC) for publicly traded companies. It contains financial data, information about the c...\n",
"\n",
"======================================================================\n"
]
}
],
"source": [
"# Let's classify the financial document\n",
"print(\"🔍 Classifying financial document...\")\n",
"print(f\" Document length: {len(financial_markdown):,} characters\\n\")\n",
"\n",
"# Write to temp file for classification\n",
"import tempfile\n",
"from pathlib import Path\n",
"\n",
"with tempfile.NamedTemporaryFile(\n",
" mode=\"w\", suffix=\".md\", delete=False, encoding=\"utf-8\"\n",
") as tmp:\n",
" tmp.write(financial_markdown)\n",
" temp_financial_path = Path(tmp.name)\n",
"\n",
"# Classify the document\n",
"financial_classification = await classify_client.aclassify_file_path(\n",
" rules=classification_rules, file_input_path=str(temp_financial_path)\n",
")\n",
"\n",
"doc_type = financial_classification.items[0].result.type\n",
"confidence = financial_classification.items[0].result.confidence\n",
"reasoning = financial_classification.items[0].result.reasoning\n",
"\n",
"print(f\"✅ Classification Result:\")\n",
"print(f\" Type: {doc_type}\")\n",
"print(f\" Confidence: {confidence:.2%}\")\n",
"print(\n",
" f\" Reasoning: {reasoning[:200]}...\"\n",
" if reasoning and len(reasoning) > 200\n",
" else f\" Reasoning: {reasoning}\"\n",
")\n",
"\n",
"print(\"\\n\" + \"=\" * 70)"
]
},
{
"cell_type": "markdown",
"metadata": {},
@@ -444,9 +499,31 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"## Complete Workflow Summary\n",
"## Building the Complete Workflow\n",
"\n",
"Let's create a function that demonstrates the complete workflow:"
"Now that we've seen how parsing works, let's build a complete 3-step workflow (Parse → Classify → Extract) using LlamaIndex Workflows. We'll define the workflow structure here, and you can see it in action below where we also demonstrate the classification and extraction modules independently.\n",
"\n",
"### Install Workflows Package\n",
"\n",
"First, let's install the LlamaIndex workflows package:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index-workflows llama-index-utils-workflow"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Define the Workflow\n",
"\n",
"Let's restructure the document processing into a proper LlamaIndex Workflow with separate classification and extraction steps:\n"
]
},
{
@@ -458,7 +535,7 @@
"name": "stdout",
"output_type": "stream",
"text": [
"🔧 Workflow function defined!\n"
"🔧 Workflow defined!\n"
]
}
],
@@ -466,81 +543,286 @@
"import tempfile\n",
"from pathlib import Path\n",
"from llama_cloud import ExtractConfig\n",
"from workflows import Workflow, step, Context\n",
"from workflows.events import Event, StartEvent, StopEvent\n",
"\n",
"\n",
"async def complete_document_workflow(markdown_content: str):\n",
"# Define workflow events\n",
"class ParseEvent(Event):\n",
" \"\"\"Event emitted after parsing\"\"\"\n",
"\n",
" file_path: str\n",
" markdown_content: str\n",
" job_id: str\n",
"\n",
"\n",
"class ClassifyEvent(Event):\n",
" \"\"\"Event emitted after classification\"\"\"\n",
"\n",
" markdown_content: str\n",
" temp_path: str\n",
" doc_type: str\n",
" confidence: float\n",
"\n",
"\n",
"class ExtractEvent(Event):\n",
" \"\"\"Event emitted after extraction\"\"\"\n",
"\n",
" doc_type: str\n",
" confidence: float\n",
" extracted_data: dict\n",
" markdown_length: int\n",
" temp_path: str\n",
" markdown_sample: str\n",
"\n",
"\n",
"class DocumentWorkflow(Workflow):\n",
" \"\"\"\n",
" Complete workflow: Parse → Classify → Extract\n",
" Complete document processing workflow: Parse → Classify → Extract\n",
" \"\"\"\n",
" print(f\"🚀 Starting complete workflow\")\n",
" print(\"=\" * 60)\n",
"\n",
" # Step 1: Classify\n",
" print(\"🏷️ Step 2: Classifying document...\")\n",
" def __init__(\n",
" self,\n",
" parser,\n",
" classify_client,\n",
" classification_rules,\n",
" llama_extract,\n",
" financial_schema,\n",
" technical_schema,\n",
" **kwargs,\n",
" ):\n",
" super().__init__(**kwargs)\n",
" self.parser = parser\n",
" self.classify_client = classify_client\n",
" self.classification_rules = classification_rules\n",
" self.llama_extract = llama_extract\n",
" self.financial_schema = financial_schema\n",
" self.technical_schema = technical_schema\n",
"\n",
" with tempfile.NamedTemporaryFile(\n",
" mode=\"w\", suffix=\".md\", delete=False, encoding=\"utf-8\"\n",
" ) as tmp:\n",
" tmp.write(markdown_content)\n",
" temp_path = Path(tmp.name)\n",
" @step\n",
" async def parse_document(self, ctx: Context, ev: StartEvent) -> ParseEvent:\n",
" \"\"\"\n",
" Step 1: Parse the document to extract markdown\n",
" \"\"\"\n",
" file_path = ev.file_path\n",
" print(f\"📄 Step 1: Parsing document: {file_path}...\")\n",
"\n",
" print(temp_path)\n",
" # Parse the document\n",
" parse_result = await self.parser.aparse(file_path)\n",
" markdown_content = await parse_result.aget_markdown()\n",
" job_id = parse_result.job_id\n",
"\n",
" classification = await classify_client.aclassify_file_path(\n",
" rules=classification_rules, file_input_path=str(temp_path)\n",
" )\n",
" doc_type = classification.items[0].result.type\n",
" confidence = classification.items[0].result.confidence\n",
" print(f\" ✅ Classified as: {doc_type} (confidence: {confidence:.2f})\")\n",
" print(f\" ✅ Parsed successfully (Job ID: {job_id})\")\n",
" print(f\" 📝 Extracted {len(markdown_content):,} characters\")\n",
"\n",
" # Step 2: Extract based on classification\n",
" print(\"🔍 Step 3: Extracting structured data using SourceText...\")\n",
" source_text = SourceText(\n",
" text_content=markdown_content,\n",
" filename=f\"{os.path.basename(temp_path)}_markdown.md\",\n",
" )\n",
" # Write event to stream for monitoring\n",
" parse_event = ParseEvent(\n",
" file_path=file_path,\n",
" markdown_content=markdown_content,\n",
" job_id=job_id,\n",
" )\n",
" ctx.write_event_to_stream(parse_event)\n",
"\n",
" # Choose schema based on classification\n",
" if \"financial\" in doc_type.lower():\n",
" schema = FinancialMetrics\n",
" print(\" 📊 Using FinancialMetrics schema\")\n",
" elif \"technical\" in doc_type.lower():\n",
" schema = TechnicalSpec\n",
" print(\" 🔧 Using TechnicalSpec schema\")\n",
" else:\n",
" schema = FinancialMetrics # Default fallback\n",
" print(\" 📊 Using default FinancialMetrics schema\")\n",
" return parse_event\n",
"\n",
" extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
" )\n",
" @step\n",
" async def classify_document(self, ctx: Context, ev: ParseEvent) -> ClassifyEvent:\n",
" \"\"\"\n",
" Step 2: Classify the document based on its content\n",
" \"\"\"\n",
" markdown_content = ev.markdown_content\n",
" print(\"🏷️ Step 2: Classifying document...\")\n",
"\n",
" extraction_result = llama_extract.extract(\n",
" data_schema=schema, config=extract_config, files=source_text\n",
" )\n",
" # Write markdown to temp file for classification\n",
" with tempfile.NamedTemporaryFile(\n",
" mode=\"w\", suffix=\".md\", delete=False, encoding=\"utf-8\"\n",
" ) as tmp:\n",
" tmp.write(markdown_content)\n",
" temp_path = Path(tmp.name)\n",
"\n",
" print(\" ✅ Extraction complete!\")\n",
" # Classify the document\n",
" classification = await self.classify_client.aclassify_file_path(\n",
" rules=self.classification_rules, file_input_path=str(temp_path)\n",
" )\n",
" doc_type = classification.items[0].result.type\n",
" confidence = classification.items[0].result.confidence\n",
"\n",
" return {\n",
" \"file_path\": temp_path,\n",
" \"markdown_length\": len(markdown_content),\n",
" \"classification\": doc_type,\n",
" \"confidence\": confidence,\n",
" \"extracted_data\": extraction_result.data,\n",
" \"markdown_sample\": markdown_content[:200] + \"...\"\n",
" if len(markdown_content) > 200\n",
" else markdown_content,\n",
" }\n",
" print(f\" ✅ Classified as: {doc_type} (confidence: {confidence:.2f})\")\n",
"\n",
" # Write event to stream for monitoring\n",
" classify_event = ClassifyEvent(\n",
" markdown_content=markdown_content,\n",
" temp_path=str(temp_path),\n",
" doc_type=doc_type,\n",
" confidence=confidence,\n",
" )\n",
" ctx.write_event_to_stream(classify_event)\n",
"\n",
" return classify_event\n",
"\n",
" @step\n",
" async def extract_data(self, ctx: Context, ev: ClassifyEvent) -> ExtractEvent:\n",
" \"\"\"\n",
" Step 3: Extract structured data based on classification\n",
" \"\"\"\n",
" print(\"🔍 Step 3: Extracting structured data using SourceText...\")\n",
"\n",
" # Choose schema based on classification\n",
" if \"financial\" in ev.doc_type.lower():\n",
" schema = self.financial_schema\n",
" print(\" 📊 Using FinancialMetrics schema\")\n",
" elif \"technical\" in ev.doc_type.lower():\n",
" schema = self.technical_schema\n",
" print(\" 🔧 Using TechnicalSpec schema\")\n",
" else:\n",
" schema = self.financial_schema # Default fallback\n",
" print(\" 📊 Using default FinancialMetrics schema\")\n",
"\n",
" # Create SourceText from markdown content\n",
" source_text = SourceText(\n",
" text_content=ev.markdown_content,\n",
" filename=f\"{os.path.basename(ev.temp_path)}_markdown.md\",\n",
" )\n",
"\n",
" # Configure extraction\n",
" extract_config = ExtractConfig(\n",
" extraction_mode=\"BALANCED\",\n",
" )\n",
"\n",
" # Perform extraction\n",
" extraction_result = self.llama_extract.extract(\n",
" data_schema=schema, config=extract_config, files=source_text\n",
" )\n",
"\n",
" print(\" ✅ Extraction complete!\")\n",
"\n",
" # Create markdown sample\n",
" markdown_sample = (\n",
" ev.markdown_content[:200] + \"...\"\n",
" if len(ev.markdown_content) > 200\n",
" else ev.markdown_content\n",
" )\n",
"\n",
" extract_event = ExtractEvent(\n",
" doc_type=ev.doc_type,\n",
" confidence=ev.confidence,\n",
" extracted_data=extraction_result.data,\n",
" markdown_length=len(ev.markdown_content),\n",
" temp_path=ev.temp_path,\n",
" markdown_sample=markdown_sample,\n",
" )\n",
" ctx.write_event_to_stream(extract_event)\n",
"\n",
" return extract_event\n",
"\n",
" @step\n",
" async def finalize_results(self, ctx: Context, ev: ExtractEvent) -> StopEvent:\n",
" \"\"\"\n",
" Step 4: Finalize and return results\n",
" \"\"\"\n",
" result = {\n",
" \"file_path\": ev.temp_path,\n",
" \"markdown_length\": ev.markdown_length,\n",
" \"classification\": ev.doc_type,\n",
" \"confidence\": ev.confidence,\n",
" \"extracted_data\": ev.extracted_data,\n",
" \"markdown_sample\": ev.markdown_sample,\n",
" }\n",
"\n",
" return StopEvent(result=result)\n",
"\n",
"\n",
"print(\"🔧 Workflow function defined!\")"
"print(\"🔧 Workflow defined!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Run Complete Workflow on Both Documents"
"### Workflow Structure\n",
"\n",
"The workflow consists of four steps connected by typed events:\n",
"\n",
"```\n",
"┌─────────────┐\n",
"│ StartEvent │ (file_path)\n",
"└──────┬──────┘\n",
" │\n",
" ▼\n",
"┌──────────────────┐\n",
"│ parse_document │ Step 1: Parse PDF to markdown\n",
"└──────┬───────────┘\n",
" │\n",
" ▼\n",
"┌─────────────┐\n",
"│ ParseEvent │ (markdown_content, job_id)\n",
"└──────┬──────┘\n",
" │\n",
" ▼\n",
"┌─────────────────────┐\n",
"│ classify_document │ Step 2: Classification\n",
"└──────┬──────────────┘\n",
" │\n",
" ▼\n",
"┌──────────────┐\n",
"│ ClassifyEvent│ (doc_type, confidence, markdown_content)\n",
"└──────┬───────┘\n",
" │\n",
" ▼\n",
"┌──────────────┐\n",
"│ extract_data │ Step 3: Extraction with schema selection\n",
"└──────┬───────┘\n",
" │\n",
" ▼\n",
"┌──────────────┐\n",
"│ ExtractEvent │ (extracted_data, doc_type, confidence)\n",
"└──────┬───────┘\n",
" │\n",
" ▼\n",
"┌──────────────────┐\n",
"│ finalize_results │ Step 4: Format and return results\n",
"└──────┬───────────┘\n",
" │\n",
" ▼\n",
"┌─────────────┐\n",
"│ StopEvent │ (final result dictionary)\n",
"└─────────────┘\n",
"```\n",
"\n",
"**Key Features:**\n",
"- **Step 1 (parse_document)**: Takes a file path and parses the document into clean markdown\n",
"- **Step 2 (classify_document)**: Takes markdown content and classifies it into document types\n",
"- **Step 3 (extract_data)**: Selects appropriate schema based on classification and extracts structured data\n",
"- **Step 4 (finalize_results)**: Packages all results into final output format\n",
"- Events are written to the stream for real-time monitoring\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Visualize the Workflow\n",
"\n",
"Let's visualize the workflow structure to see the flow of events:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Initialize the workflow\n",
"workflow = DocumentWorkflow(\n",
" parser=parser,\n",
" classify_client=classify_client,\n",
" classification_rules=classification_rules,\n",
" llama_extract=llama_extract,\n",
" financial_schema=FinancialMetrics,\n",
" technical_schema=TechnicalSpec,\n",
" timeout=300,\n",
" verbose=True,\n",
")"
]
},
{
@@ -552,53 +834,173 @@
"name": "stdout",
"output_type": "stream",
"text": [
"🚀 Starting complete workflow\n",
"============================================================\n",
"🏷️ Step 2: Classifying document...\n",
"/var/folders/g6/4b5lpp5974gcpr890ybhbw4r0000gn/T/tmpos3b62tm.md\n",
" ✅ Classified as: financial_document (confidence: 1.00)\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 📊 Using FinancialMetrics schema\n",
".. ✅ Extraction complete!\n",
"\n",
"============================================================\n",
"\n",
"🚀 Starting complete workflow\n",
"============================================================\n",
"🏷️ Step 2: Classifying document...\n",
"/var/folders/g6/4b5lpp5974gcpr890ybhbw4r0000gn/T/tmpppz9ub_m.md\n",
" ✅ Classified as: technical_specification (confidence: 1.00)\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 🔧 Using TechnicalSpec schema\n",
" ✅ Extraction complete!\n",
"\n",
"============================================================\n",
"\n",
"📋 Processed 2 documents successfully!\n"
"document_workflow.html\n"
]
}
],
"source": [
"# Process both documents through the complete workflow\n",
"results = []\n",
"# Draw the workflow visualization\n",
"from llama_index.utils.workflow import draw_all_possible_flows\n",
"\n",
"for doc_text in document_texts:\n",
" try:\n",
" result = await complete_document_workflow(doc_text)\n",
" results.append(result)\n",
" print(\"\\n\" + \"=\" * 60 + \"\\n\")\n",
" except Exception as e:\n",
" print(f\"❌ Error processing {doc_path}: {str(e)}\")\n",
" print(\"\\n\" + \"=\" * 60 + \"\\n\")\n",
"\n",
"print(f\"📋 Processed {len(results)} documents successfully!\")"
"draw_all_possible_flows(\n",
" workflow,\n",
" filename=\"document_workflow.html\",\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Final Results Summary"
"The workflow has been visualized and saved to `document_workflow.html`. You can open this file in a browser to see the interactive workflow diagram.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The workflow visualization shows:\n",
"1. **StartEvent** → **parse_document** step\n",
"2. **ParseEvent** → **classify_document** step\n",
"3. **ClassifyEvent** → **extract_data** step \n",
"4. **ExtractEvent** → **finalize_results** step\n",
"5. **StopEvent** (final output)\n",
"\n",
"Each step is connected by typed events, allowing for clean separation of concerns and easy monitoring of the workflow execution.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Run the Workflow on Both Documents\n",
"\n",
"Now let's run the workflow on both documents and monitor the events:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"======================================================================\n",
"🚀 Processing Document 1: sample_docs/financial_report.pdf\n",
"======================================================================\n",
"\n",
"Running step parse_document\n",
"📄 Step 1: Parsing document: sample_docs/financial_report.pdf...\n",
"Started parsing the file under job_id bb53c6bf-79cc-4f63-9c97-16983d59f29d\n",
". ✅ Parsed successfully (Job ID: bb53c6bf-79cc-4f63-9c97-16983d59f29d)\n",
" 📝 Extracted 1,338,499 characters\n",
"Step parse_document produced event ParseEvent\n",
"📄 Parse Event: Extracted 1,338,499 characters\n",
"Running step classify_document\n",
"🏷️ Step 2: Classifying document...\n",
" ✅ Classified as: financial_document (confidence: 1.00)\n",
"Step classify_document produced event ClassifyEvent\n",
"📊 Classification Event: financial_document (1.00)\n",
"Running step extract_data\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 📊 Using FinancialMetrics schema\n",
".. ✅ Extraction complete!\n",
"Step extract_data produced event ExtractEvent\n",
"Running step finalize_results\n",
"Step finalize_results produced event StopEvent\n",
"✅ Extraction Event: 7 fields extracted\n",
"\n",
"✅ Document 1 processed successfully!\n",
"\n",
"======================================================================\n",
"🚀 Processing Document 2: sample_docs/technical_spec.pdf\n",
"======================================================================\n",
"\n",
"Running step parse_document\n",
"📄 Step 1: Parsing document: sample_docs/technical_spec.pdf...\n",
"Started parsing the file under job_id 944905c1-3c49-431a-ad86-4436d16f3d1c\n",
" ✅ Parsed successfully (Job ID: 944905c1-3c49-431a-ad86-4436d16f3d1c)\n",
" 📝 Extracted 92,483 characters\n",
"Step parse_document produced event ParseEvent\n",
"📄 Parse Event: Extracted 92,483 characters\n",
"Running step classify_document\n",
"🏷️ Step 2: Classifying document...\n",
" ✅ Classified as: technical_specification (confidence: 1.00)\n",
"Step classify_document produced event ClassifyEvent\n",
"📊 Classification Event: technical_specification (1.00)\n",
"Running step extract_data\n",
"🔍 Step 3: Extracting structured data using SourceText...\n",
" 🔧 Using TechnicalSpec schema\n",
" ✅ Extraction complete!\n",
"Step extract_data produced event ExtractEvent\n",
"Running step finalize_results\n",
"Step finalize_results produced event StopEvent\n",
"✅ Extraction Event: 8 fields extracted\n",
"\n",
"✅ Document 2 processed successfully!\n",
"\n",
"\n",
"📋 Processed 2 documents successfully!\n"
]
}
],
"source": [
"# Process both documents through the workflow\n",
"results = []\n",
"\n",
"# Define the document files to process\n",
"document_files = [\n",
" \"sample_docs/financial_report.pdf\",\n",
" \"sample_docs/technical_spec.pdf\",\n",
"]\n",
"\n",
"for i, file_path in enumerate(document_files, 1):\n",
" print(f\"\\n{'='*70}\")\n",
" print(f\"🚀 Processing Document {i}: {file_path}\")\n",
" print(f\"{'='*70}\\n\")\n",
"\n",
" try:\n",
" # Run the workflow\n",
" handler = workflow.run(file_path=file_path)\n",
"\n",
" # Monitor events as they are emitted\n",
" async for event in handler.stream_events():\n",
" if isinstance(event, ParseEvent):\n",
" print(\n",
" f\"📄 Parse Event: Extracted {len(event.markdown_content):,} characters\"\n",
" )\n",
" elif isinstance(event, ClassifyEvent):\n",
" print(\n",
" f\"📊 Classification Event: {event.doc_type} ({event.confidence:.2f})\"\n",
" )\n",
" elif isinstance(event, ExtractEvent):\n",
" print(\n",
" f\"✅ Extraction Event: {len(event.extracted_data)} fields extracted\"\n",
" )\n",
"\n",
" # Get final result\n",
" result = await handler\n",
" results.append(result)\n",
"\n",
" print(f\"\\n✅ Document {i} processed successfully!\")\n",
"\n",
" except Exception as e:\n",
" print(f\"❌ Error processing document {i}: {str(e)}\")\n",
" import traceback\n",
"\n",
" traceback.print_exc()\n",
"\n",
"print(f\"\\n\\n📋 Processed {len(results)} documents successfully!\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Final Results Summary\n"
]
},
{
@@ -613,9 +1015,9 @@
"📈 COMPLETE WORKFLOW RESULTS SUMMARY\n",
"======================================================================\n",
"\n",
"📄 Document 1: tmpos3b62tm.md\n",
"📄 Document 1: tmpuyxzpd3x.md\n",
" 📊 Classification: financial_document (confidence: 1.00)\n",
" 📝 Markdown length: 1,348,671 characters\n",
" 📝 Markdown length: 1,338,499 characters\n",
" 📋 Markdown sample: \n",
"\n",
"# UNITED STATES\n",
@@ -629,14 +1031,14 @@
" • company_name: Uber Technologies, Inc.\n",
" • document_type: Annual Report on Form 10-K\n",
" • fiscal_year: 2021\n",
" • revenue_2021: $21,764\n",
" • net_income_2021: $(496)\n",
" • key_business_segments: ['Mobility', 'Delivery', 'Freight', 'All Other (including former New Mobility, e-bikes, e-scooters, Advanced Technologies Group and other technology programs)']\n",
" • risk_factors: [\"The company faces numerous risk factors across its business operations and environment. The COVID-19 pandemic and related mitigation measures have adversely affected parts of the business, including reduced demand for Mobility offerings and creating ongoing uncertainties. The company's operational and financial performance is influenced by competitive pressure in the mobility, delivery, and logistics industries, characterized by well-established alternatives, low barriers to entry, and low switching costs. Driver classification risks exist if Drivers are deemed employees, workers, or quasi-employees rather than independent contractors, exposing the company to legal actions and financial liabilities globally. Competition challenges require the company to sometimes lower fares, offer incentives, and promotions, which impacts profitability. There are significant operating losses historically with substantial future operating expense increases anticipated, and the ability to achieve or maintain profitability is uncertain. Network value depends on maintaining critical mass among Drivers, consumers, merchants, shippers, and carriers, and failures to do so diminish platform attractiveness. Brand and reputation maintenance is critical, with exposure to negative publicity, media coverage, and risks from associated companies' brands or licensed brands in joint ventures.\\n\\nOperational risks include historical workplace culture and compliance challenges, management complexity due to rapid growth, technological infrastructure issues potentially causing disruptions or poor user experience, and security or data privacy breaches that could impact revenue and reputation. Platform users may engage in or be subjected to criminal, violent, or dangerous activity leading to safety incidents and legal actions. New offerings and technologies investments are inherently risky without guaranteed benefits. Economic conditions, inflation, and increased costs (fuel, food, labor, energy) may negatively impact results. Regulatory risks are extensive and global, involving payment and financial services compliance, licensing, anti-money laundering laws, data privacy (GDPR, CCPA, LGPD), and labor laws. Legal and regulatory investigations and inquiries, including antitrust, FCPA, labor classification, data protection, and intellectual property matters, pose risks of fines, penalties, operational changes, and increased costs.\\n\\nGeopolitical and jurisdictional risks include operating limitations or bans in some locations, currency exchange risk, and complex evolving regulations with the potential for fines and loss of licenses or permits. Insurance risks include potential inadequacy of reserves, liability exposure from accidents or impersonation, and insurer insolvency. Driver qualification requirements and background checks may increase costs or fail to expose all relevant information, with associated insurance cost risks and potential for courtroom or regulatory challenges to pricing models.\\n\\nFinancial risks comprise significant accumulated deficits, requirement for additional capital with uncertain availability, debt obligations, tax exposure including uncertain positions and observed changes in tax laws, and volatility in common stock price with no expected cash dividends. Accounting judgments and estimates involve critical assumptions affecting reported financial metrics related to goodwill, revenue recognition, incentive accruals, and stock-based compensation. Cybersecurity risks include exposures to malware, ransomware, phishing, and other cyberattacks. Climate change presents physical and transitional risks that may impact operations and costs, and failure to meet climate commitments may have operational and reputational consequences.\\n\\nOther risks include potential liability under anti-corruption and anti-terrorism laws, adverse effects from defaults under debt agreements, limitations in takeover actions due to corporate governance provisions, and the impact of non-GAAP financial measure limitations. Overall, these diverse and interconnected risk factors contribute to significant uncertainty regarding the company's future business prospects, operating results, and financial condition.\"]\n",
" • revenue_2021: $17,455 and $21,764\n",
" • net_income_2021: $(496) to (700)\n",
" • key_business_segments: ['Borrower and the Restricted Subsidiaries', 'Holdings', 'Guarantors', 'Material Domestic Subsidiaries', 'Material Foreign Subsidiaries']\n",
" • risk_factors: ['Indemnification obligations of the borrower for losses, claims, damages, liabilities, and out-of-pocket expenses incurred by agents, lenders, arrangers, and related parties in connection with the agreement or loans, except in certain cases such as gross negligence, bad faith, willful misconduct, or material breach by the indemnitee.', \"Borrower not required to indemnify any indemnitee for settlements entered into without the borrower's consent.\", 'Limitation of liability for special, indirect, consequential, or punitive damages, and for damages from unauthorized use of information, except for direct damages resulting from gross negligence, bad faith, or willful misconduct.', 'Obligation of the borrower to indemnify the administrative agent for liabilities arising from performance of duties, except in cases of gross negligence, bad faith, or willful misconduct.', 'Limitations and conditions on assignments and participations of lender rights, including restrictions on assignments to disqualified institutions, loan parties, affiliates of loan parties, defaulting lenders, and natural persons.', 'Setoff rights for lenders and issuing banks after an event of default, allowing them to apply borrower deposits toward obligations under the agreement.', 'Potential for increased obligations under the agreement as a result of changes in law affecting payment terms.', 'Requirement for the borrower and guarantors to provide information to comply with anti-money laundering rules and the USA PATRIOT Act.']\n",
"\n",
"📄 Document 2: tmpppz9ub_m.md\n",
"📄 Document 2: tmp7ower2xm.md\n",
" 📊 Classification: technical_specification (confidence: 1.00)\n",
" 📝 Markdown length: 90,971 characters\n",
" 📝 Markdown length: 92,483 characters\n",
" 📋 Markdown sample: \n",
"\n",
"LM317\n",
@@ -648,20 +1050,14 @@
" 🎯 Extracted fields: 8 fields\n",
" • component_name: LM317\n",
" • manufacturer: Texas Instruments\n",
" • part_number: LM317\n",
" • description: The LM317 is an adjustable three-pin, positive-voltage regulator capable of supplying up to 1.5A over an output voltage range of 1.25V to 37V. It features line and load regulation, internal current limiting, thermal overload protection, and safe operating area compensation.\n",
" • part_number: LM317, SLVS044Z\n",
" • description: The LM317 is an adjustable three-pin, positive-voltage regulator capable of supplying more than 1.5A (typically up to 1.5A) over an output voltage range of 1.25V to 37V. The device requires only two external resistors to set the output voltage. It features a typical line regulation of 0.01% and typical load regulation of 0.1%. The LM317 includes current limiting, thermal overload protection, and safe operating area protection. Overload protection remains functional even if the ADJUST pin is disconnected. The regulator is used in applications such as constant-current battery-charger circuits, slow turn-on 15V regulator circuits, AC voltage-regulator circuits, current-limited charger circuits, and high-current and adjustable regulator circuits. It is available in packages including SOT-223 (DCY), TO-220 (KCS), and TO-263 (KTT).\n",
" • operating_voltage: {'min_voltage': 1.25, 'max_voltage': 37.0, 'unit': 'V'}\n",
" • maximum_current: 1.5\n",
" • key_features: ['Adjustable output voltage: 1.25V to 37V', 'Output current up to 1.5A', 'Line regulation: 0.01%/V (typical)', 'Load regulation: 0.1% (typical)', 'Internal short-circuit current limiting', 'Thermal overload protection', 'Output safe-area compensation', 'High power supply rejection ratio (PSRR): 80dB at 120Hz (new chip)', 'Available in SOT-223, TO-263, and TO-220 packages']\n",
" • applications: ['Multifunction printers', 'AC drive power stage modules', 'Electricity meters', 'Servo drive control modules', 'Merchant network and server power supply units']\n",
" • maximum_current: 4.0\n",
" • key_features: ['Adjustable output voltage range: 1.25V to 37V', 'Output current up to 1.5A (up to 4A with external pass elements)', 'Line regulation: typically 0.01%/V', 'Load regulation: typically 0.1%', 'Internal short-circuit current limiting / Current limiting', 'Thermal overload protection / Thermal shutdown', 'Output safe-area compensation / Safe operating area protection', 'PSRR: 80dB at 120Hz for CADJ = 10μF (new chip)', 'NPN Darlington output drive', 'Programmable feedback', 'Multiple package options (SOT-223, TO-220, TO-263)', 'Can be used in constant-current, battery-charging, and regulator applications']\n",
" • applications: ['Multifunction printers, AC drive power stage modules, Electricity meters, Servo drive control modules, Merchant network and server PSU, Adjustable voltage regulator, 0V to 30V regulator circuit, Regulator circuit with improved ripple rejection, Precision current-limiter, Tracking preregulator, 1.25V to 20V regulator, Battery charger circuit, Constant-current battery charger circuits, Slow turn-on regulator, AC voltage-regulator, Current-limited charger circuits, High-current adjustable regulator circuits, General-purpose adjustable power supply']\n",
"\n",
"✨ Workflow completed successfully!\n",
"\n",
"📚 Key Learnings:\n",
" • Parse: Converted documents to clean markdown format\n",
" • Classify: Automatically categorized document types\n",
" • Extract: Used SourceText with markdown for structured data extraction\n",
" • The markdown content provides much better context for extraction than raw PDFs\n"
"✨ Workflow completed successfully!\n"
]
}
],
@@ -683,14 +1079,7 @@
" for key, value in extracted.items():\n",
" print(f\" • {key}: {value}\")\n",
"\n",
"print(\"\\n✨ Workflow completed successfully!\")\n",
"print(\"\\n📚 Key Learnings:\")\n",
"print(\" • Parse: Converted documents to clean markdown format\")\n",
"print(\" • Classify: Automatically categorized document types\")\n",
"print(\" • Extract: Used SourceText with markdown for structured data extraction\")\n",
"print(\n",
" \" • The markdown content provides much better context for extraction than raw PDFs\"\n",
")"
"print(\"\\n✨ Workflow completed successfully!\")"
]
},
{
@@ -699,54 +1088,33 @@
"source": [
"## Conclusion\n",
"\n",
"This notebook demonstrated the complete **Parse → Classify → Extract** workflow using LlamaCloud services:\n",
"The notebook shows you how to build an e2e document **Classify → Extract** workflow using LlamaCloud. This uses some of our core building blocks around **classification** interleaved with **document extraction**.\n",
"\n",
"### Key Components:\n",
"### Main Components:\n",
"\n",
"1. **LlamaParse** (`llama_cloud_services.parse.base.LlamaParse`):\n",
" - Converts documents to clean, structured markdown\n",
" - Preserves document structure and formatting\n",
" - Handles various file types (PDF, DOCX, etc.)\n",
"\n",
"2. **ClassifyClient** (`llama_cloud_services.beta.classifier.client.ClassifyClient`):\n",
"2. **LlamaClassify** (`llama_cloud_services.beta.classifier.client.LlamaClassify`):\n",
" - Automatically categorizes documents based on content\n",
" - Uses customizable rules for classification\n",
" - Provides confidence scores for classifications\n",
"\n",
"3. **LlamaExtract with SourceText** (`llama_cloud_services.extract.extract.LlamaExtract`, `SourceText`):\n",
" - Extracts structured data using custom Pydantic schemas\n",
" - **SourceText** allows using markdown content as input instead of raw files\n",
" - Provides much better extraction accuracy when using processed markdown\n",
" - You can either feed in the file directly (in which case parsing will happen under the hood), or the parsed text through the **SourceText** object (which is the case in this example) \n",
"\n",
"### Workflow Benefits:\n",
"\n",
"- **Better Accuracy**: Using markdown from parsing provides cleaner, more structured input for extraction\n",
"- **Automatic Routing**: Classification allows different processing logic for different document types\n",
"- **Structured Output**: Custom schemas ensure consistent, structured data extraction\n",
"- **Flexible Input**: SourceText supports text content, file paths, and bytes\n",
"\n",
"### Key Insights:\n",
"\n",
"1. **SourceText is the bridge**: It allows you to pass the clean markdown content from parsing directly to extraction\n",
"2. **Markdown improves extraction**: Pre-processed markdown provides much better context than raw PDFs\n",
"3. **Classification enables smart routing**: Different document types can use different extraction schemas\n",
"4. **End-to-end automation**: The entire workflow can be automated for production use\n",
"\n",
"This approach is ideal for production document processing pipelines where you need to:\n",
"- Process various document types automatically\n",
"- Extract structured data consistently\n",
"- Maintain high accuracy and reliability\n",
"- Handle documents at scale\n",
"\n",
"The combination of these three services provides a powerful, flexible document processing pipeline that can handle complex, real-world document processing requirements."
"**Benefits of an e2e workflow**: The main benefit of doing Classify -> Extract, instead of only Extract, is the fact that you can handle documents of different types/different expected schemas within the same workflow, without having to separate out the data before and running separate extractions on each data subset. "
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"display_name": "llama_parse",
"language": "python",
"name": "python3"
"name": "llama_parse"
},
"language_info": {
"codemirror_mode": {
@@ -0,0 +1,73 @@
This project uses LlamaSheets to extract data from spreadsheets for analysis.
## Current Project Structure
- `data/` - Contains extracted parquet files from LlamaSheets
- `{name}_region_{N}.parquet` - Table data files
- `{name}_metadata_{N}.parquet` - Cell metadata files
- `{name}_job_metadata.json` - Extraction job information
- `scripts/` - Analysis and helper scripts
- `reports/` - Your generated reports and outputs
## Working with LlamaSheets Data
### Understanding the Files
When a spreadsheet is extracted, you'll find:
1. **Table parquet files** (`region_*.parquet`): The actual table data
- Columns correspond to spreadsheet columns
- Data types are preserved (dates, numbers, strings, booleans)
2. **Metadata parquet files** (`metadata_*.parquet`): Rich cell-level metadata
- Formatting: `font_bold`, `font_italic`, `font_size`, `background_color_rgb`
- Position: `row_number`, `column_number`, `coordinate` (e.g., "A1")
- Type detection: `data_type`, `is_date_like`, `is_percentage`, `is_currency`
- Layout: `is_in_first_row`, `is_merged_cell`, `horizontal_alignment`
- Content: `cell_value`, `raw_cell_value`
3. **Job metadata JSON** (`job_metadata.json`): Overall extraction results
- `regions[]`: List of extracted regions with IDs, locations, and titles/descriptions
- `worksheet_metadata[]`: Generated titles and descriptions
- `status`: Success/failure status
### Key Principles
1. **Use metadata to understand structure**: Bold cells often indicate headers, colors indicate groupings
2. **Validate before analysis**: Check data types, look for missing values
3. **Preserve formatting context**: The metadata tells you what the spreadsheet author emphasized
4. **Save intermediate results**: Store cleaned data as new parquet files
### Common Patterns
**Loading data:**
```python
import pandas as pd
df = pd.read_parquet("data/region_1_Sheet1.parquet")
meta_df = pd.read_parquet("data/metadata_1_Sheet1.parquet")
```
**Finding headers:**
```python
headers = meta_df[meta_df["font_bold"] == True]["cell_value"].tolist()
```
**Finding date columns:**
```python
date_cols = meta_df[meta_df["is_date_like"] == True]["column_number"].unique()
```
## Tools Available
- **Python 3.11+**: For data analysis
- **pandas**: DataFrame manipulation
- **pyarrow**: Parquet file reading
- **matplotlib**: Visualization (optional)
## Guidelines
- Always read the job_metadata.json first to understand what was extracted
- Check both table data and metadata before making assumptions
- Write reusable functions for common operations
- Document any data quality issues discovered
@@ -0,0 +1,278 @@
"""
Generate sample spreadsheets for LlamaSheets + Claude workflows.
This script creates example Excel files that demonstrate different use cases:
1. Simple data table (for Workflow 1)
2. Regional sales data (for Workflow 2)
3. Complex budget with formatting (for Workflow 3)
4. Weekly sales report (for Workflow 4)
Usage:
python generate_sample_data.py
"""
import random
from datetime import datetime, timedelta
from pathlib import Path
import pandas as pd
from openpyxl import Workbook
from openpyxl.styles import Font, PatternFill, Alignment
def generate_workflow_1_data(output_dir: Path) -> None:
"""Generate simple financial report for Workflow 1."""
print("📊 Generating Workflow 1: financial_report_q1.xlsx")
# Create sample quarterly data
months = ["January", "February", "March"]
categories = ["Revenue", "Cost of Goods Sold", "Operating Expenses", "Net Income"]
data = []
for category in categories:
row: dict[str, str | int] = {"Category": category}
for month in months:
if category == "Revenue":
value = random.randint(80000, 120000)
elif category == "Cost of Goods Sold":
value = random.randint(30000, 50000)
elif category == "Operating Expenses":
value = random.randint(20000, 35000)
else: # Net Income
value = int(
int(row.get("January", 0))
+ int(row.get("February", 0))
+ int(row.get("March", 0))
)
value = random.randint(15000, 40000)
row[month] = value
data.append(row)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / "financial_report_q1.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Q1 Summary", index=False)
# Format it nicely
worksheet = writer.sheets["Q1 Summary"]
for cell in worksheet[1]: # Header row
cell.font = Font(bold=True)
cell.fill = PatternFill(
start_color="4F81BD", end_color="4F81BD", fill_type="solid"
)
cell.font = Font(color="FFFFFF", bold=True)
print(f" ✅ Created {output_file}")
def generate_workflow_2_data(output_dir: Path) -> None:
"""Generate regional sales data for Workflow 2."""
print("\n📊 Generating Workflow 2: Regional sales data")
regions = ["northeast", "southeast", "west"]
products = ["Widget A", "Widget B", "Widget C", "Gadget X", "Gadget Y"]
for region in regions:
data = []
start_date = datetime(2024, 1, 1)
# Generate 90 days of sales data
for day in range(90):
date = start_date + timedelta(days=day)
# Random number of sales per day (3-8)
for _ in range(random.randint(3, 8)):
product = random.choice(products)
units_sold = random.randint(1, 20)
price_per_unit = random.randint(50, 200)
revenue = units_sold * price_per_unit
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units_Sold": units_sold,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / f"sales_{region}.xlsx"
df.to_excel(output_file, sheet_name="Sales", index=False)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def generate_workflow_3_data(output_dir: Path) -> None:
"""Generate complex budget spreadsheet with formatting for Workflow 3."""
print("\n📊 Generating Workflow 3: company_budget_2024.xlsx")
wb = Workbook()
ws = wb.active
ws.title = "Budget"
# Define departments with colors
departments = {
"Engineering": "C6E0B4",
"Marketing": "FFD966",
"Sales": "F4B084",
"Operations": "B4C7E7",
}
# Define categories
categories = {
"Personnel": ["Salaries", "Benefits", "Training"],
"Infrastructure": ["Office Rent", "Equipment", "Software Licenses"],
"Operations": ["Travel", "Supplies", "Miscellaneous"],
}
# Styles
header_font = Font(bold=True, size=12)
category_font = Font(bold=True, size=11)
row = 1
# Title
ws.merge_cells(f"A{row}:E{row}")
ws[f"A{row}"] = "2024 Annual Budget"
ws[f"A{row}"].font = Font(bold=True, size=14)
ws[f"A{row}"].alignment = Alignment(horizontal="center")
row += 2
# Headers
ws[f"A{row}"] = "Category"
ws[f"B{row}"] = "Item"
for i, dept in enumerate(departments.keys()):
ws.cell(row, 3 + i, dept)
ws.cell(row, 3 + i).font = header_font
for cell in ws[row]:
cell.font = header_font
row += 1
# Data
for category, items in categories.items():
# Category header (bold)
ws[f"A{row}"] = category
ws[f"A{row}"].font = category_font
row += 1
# Items with department budgets
for item in items:
ws[f"A{row}"] = ""
ws[f"B{row}"] = item
# Add budget amounts for each department (with color)
for i, (dept, color) in enumerate(departments.items()):
amount = random.randint(5000, 50000)
cell = ws.cell(row, 3 + i, amount)
cell.fill = PatternFill(
start_color=color, end_color=color, fill_type="solid"
)
cell.number_format = "$#,##0"
row += 1
row += 1 # Blank row between categories
# Adjust column widths
ws.column_dimensions["A"].width = 20
ws.column_dimensions["B"].width = 25
for i in range(len(departments)):
ws.column_dimensions[chr(67 + i)].width = 15 # C, D, E, F
output_file = output_dir / "company_budget_2024.xlsx"
wb.save(output_file)
print(f" ✅ Created {output_file}")
print(" • Bold categories, colored departments, merged title cell")
def generate_workflow_4_data(output_dir: Path) -> None:
"""Generate weekly sales report for Workflow 4."""
print("\n📊 Generating Workflow 4: sales_weekly.xlsx")
products = [
"Product A",
"Product B",
"Product C",
"Product D",
"Product E",
"Product F",
"Product G",
"Product H",
]
# Generate one week of data
data = []
start_date = datetime(2024, 11, 4) # Monday
for day in range(7):
date = start_date + timedelta(days=day)
# Each product has 3-10 transactions per day
for product in products:
for _ in range(random.randint(3, 10)):
units = random.randint(1, 15)
price = random.randint(20, 150)
revenue = units * price
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units": units,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel with some formatting
output_file = output_dir / "sales_weekly.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Weekly Sales", index=False)
# Format header
worksheet = writer.sheets["Weekly Sales"]
for cell in worksheet[1]:
cell.font = Font(bold=True)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def main() -> None:
"""Generate all sample data files."""
print("=" * 60)
print("Generating Sample Data for LlamaSheets + Coding Agent Workflows")
print("=" * 60)
# Create output directory
output_dir = Path("input_data")
output_dir.mkdir(exist_ok=True)
# Generate data for each workflow
generate_workflow_1_data(output_dir)
generate_workflow_2_data(output_dir)
generate_workflow_3_data(output_dir)
generate_workflow_4_data(output_dir)
print("\n" + "=" * 60)
print("✅ All sample data generated!")
print("=" * 60)
print(f"\nFiles created in {output_dir.absolute()}:")
print("\nWorkflow 1 (Understanding a New Spreadsheet):")
print(" • financial_report_q1.xlsx")
print("\nWorkflow 2 (Generating Analysis Scripts):")
print(" • sales_northeast.xlsx")
print(" • sales_southeast.xlsx")
print(" • sales_west.xlsx")
print("\nWorkflow 3 (Using Cell Metadata):")
print(" • company_budget_2024.xlsx")
print("\nWorkflow 4 (Complete Automation):")
print(" • sales_weekly.xlsx")
print("\nYou can now use these files with the workflows in the documentation!")
if __name__ == "__main__":
main()
@@ -0,0 +1,5 @@
llama-cloud-services # LlamaSheets SDK
pandas>=2.0.0
pyarrow>=12.0.0
openpyxl>=3.0.0 # For Excel file support
matplotlib>=3.7.0 # For visualizations (optional)
@@ -0,0 +1,100 @@
"""Helper script to extract spreadsheets using LlamaSheets."""
import asyncio
import json
import os
import dotenv
from pathlib import Path
from llama_cloud_services.beta.sheets import LlamaSheets
from llama_cloud_services.beta.sheets.types import (
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
dotenv.load_dotenv()
async def extract_spreadsheet(
file_path: str, output_dir: str = "data", generate_metadata: bool = True
) -> dict:
"""Extract a spreadsheet using LlamaSheets."""
client = LlamaSheets(
base_url="https://api.cloud.llamaindex.ai",
api_key=os.getenv("LLAMA_CLOUD_API_KEY"),
)
print(f"Extracting {file_path}...")
# Extract regions
config = SpreadsheetParsingConfig(
sheet_names=None, # Extract all sheets
generate_additional_metadata=generate_metadata,
)
job_result = await client.aextract_regions(file_path, config=config)
print(f"Extracted {len(job_result.regions)} region(s)")
# Create output directory
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
# Get base name for files
base_name = Path(file_path).stem
# Save job metadata
job_metadata_path = output_path / f"{base_name}_job_metadata.json"
with open(job_metadata_path, "w") as f:
json.dump(job_result.model_dump(mode="json"), f, indent=2)
print(f"Saved job metadata to {job_metadata_path}")
# Download each region
for idx, region in enumerate(job_result.regions, 1):
sheet_name = region.sheet_name.replace(" ", "_")
# Download region data
region_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=region.region_type,
)
region_path = output_path / f"{base_name}_region_{idx}_{sheet_name}.parquet"
with open(region_path, "wb") as f:
f.write(region_bytes)
print(f" Table {idx}: {region_path}")
# Download metadata
metadata_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=SpreadsheetResultType.CELL_METADATA,
)
metadata_path = output_path / f"{base_name}_metadata_{idx}_{sheet_name}.parquet"
with open(metadata_path, "wb") as f:
f.write(metadata_bytes)
print(f" Metadata {idx}: {metadata_path}")
print(f"\nAll files saved to {output_path}/")
return job_result.model_dump(mode="json")
if __name__ == "__main__":
import sys
if len(sys.argv) < 2:
print("Usage: python scripts/extract.py <spreadsheet_file>")
sys.exit(1)
file_path = sys.argv[1]
if not Path(file_path).exists():
print(f"❌ File not found: {file_path}")
sys.exit(1)
result = asyncio.run(extract_spreadsheet(file_path))
print(f"\n✅ Extraction complete! Job ID: {result['id']}")
@@ -0,0 +1,278 @@
"""
Generate sample spreadsheets for LlamaSheets + LlamaIndex Agent workflows.
This script creates example Excel files that demonstrate different use cases:
1. Simple data table (for Workflow 1)
2. Regional sales data (for Workflow 2)
3. Complex budget with formatting (for Workflow 3)
4. Weekly sales report (for Workflow 4)
Usage:
python generate_sample_data.py
"""
import random
from datetime import datetime, timedelta
from pathlib import Path
import pandas as pd
from openpyxl import Workbook
from openpyxl.styles import Font, PatternFill, Alignment
def generate_workflow_1_data(output_dir: Path) -> None:
"""Generate simple financial report for Workflow 1."""
print("📊 Generating Workflow 1: financial_report_q1.xlsx")
# Create sample quarterly data
months = ["January", "February", "March"]
categories = ["Revenue", "Cost of Goods Sold", "Operating Expenses", "Net Income"]
data = []
for category in categories:
row: dict[str, str | int] = {"Category": category}
for month in months:
if category == "Revenue":
value = random.randint(80000, 120000)
elif category == "Cost of Goods Sold":
value = random.randint(30000, 50000)
elif category == "Operating Expenses":
value = random.randint(20000, 35000)
else: # Net Income
value = int(
int(row.get("January", 0))
+ int(row.get("February", 0))
+ int(row.get("March", 0))
)
value = random.randint(15000, 40000)
row[month] = value
data.append(row)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / "financial_report_q1.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Q1 Summary", index=False)
# Format it nicely
worksheet = writer.sheets["Q1 Summary"]
for cell in worksheet[1]: # Header row
cell.font = Font(bold=True)
cell.fill = PatternFill(
start_color="4F81BD", end_color="4F81BD", fill_type="solid"
)
cell.font = Font(color="FFFFFF", bold=True)
print(f" ✅ Created {output_file}")
def generate_workflow_2_data(output_dir: Path) -> None:
"""Generate regional sales data for Workflow 2."""
print("\n📊 Generating Workflow 2: Regional sales data")
regions = ["northeast", "southeast", "west"]
products = ["Widget A", "Widget B", "Widget C", "Gadget X", "Gadget Y"]
for region in regions:
data = []
start_date = datetime(2024, 1, 1)
# Generate 90 days of sales data
for day in range(90):
date = start_date + timedelta(days=day)
# Random number of sales per day (3-8)
for _ in range(random.randint(3, 8)):
product = random.choice(products)
units_sold = random.randint(1, 20)
price_per_unit = random.randint(50, 200)
revenue = units_sold * price_per_unit
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units_Sold": units_sold,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel
output_file = output_dir / f"sales_{region}.xlsx"
df.to_excel(output_file, sheet_name="Sales", index=False)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def generate_workflow_3_data(output_dir: Path) -> None:
"""Generate complex budget spreadsheet with formatting for Workflow 3."""
print("\n📊 Generating Workflow 3: company_budget_2024.xlsx")
wb = Workbook()
ws = wb.active
ws.title = "Budget"
# Define departments with colors
departments = {
"Engineering": "C6E0B4",
"Marketing": "FFD966",
"Sales": "F4B084",
"Operations": "B4C7E7",
}
# Define categories
categories = {
"Personnel": ["Salaries", "Benefits", "Training"],
"Infrastructure": ["Office Rent", "Equipment", "Software Licenses"],
"Operations": ["Travel", "Supplies", "Miscellaneous"],
}
# Styles
header_font = Font(bold=True, size=12)
category_font = Font(bold=True, size=11)
row = 1
# Title
ws.merge_cells(f"A{row}:E{row}")
ws[f"A{row}"] = "2024 Annual Budget"
ws[f"A{row}"].font = Font(bold=True, size=14)
ws[f"A{row}"].alignment = Alignment(horizontal="center")
row += 2
# Headers
ws[f"A{row}"] = "Category"
ws[f"B{row}"] = "Item"
for i, dept in enumerate(departments.keys()):
ws.cell(row, 3 + i, dept)
ws.cell(row, 3 + i).font = header_font
for cell in ws[row]:
cell.font = header_font
row += 1
# Data
for category, items in categories.items():
# Category header (bold)
ws[f"A{row}"] = category
ws[f"A{row}"].font = category_font
row += 1
# Items with department budgets
for item in items:
ws[f"A{row}"] = ""
ws[f"B{row}"] = item
# Add budget amounts for each department (with color)
for i, (dept, color) in enumerate(departments.items()):
amount = random.randint(5000, 50000)
cell = ws.cell(row, 3 + i, amount)
cell.fill = PatternFill(
start_color=color, end_color=color, fill_type="solid"
)
cell.number_format = "$#,##0"
row += 1
row += 1 # Blank row between categories
# Adjust column widths
ws.column_dimensions["A"].width = 20
ws.column_dimensions["B"].width = 25
for i in range(len(departments)):
ws.column_dimensions[chr(67 + i)].width = 15 # C, D, E, F
output_file = output_dir / "company_budget_2024.xlsx"
wb.save(output_file)
print(f" ✅ Created {output_file}")
print(" • Bold categories, colored departments, merged title cell")
def generate_workflow_4_data(output_dir: Path) -> None:
"""Generate weekly sales report for Workflow 4."""
print("\n📊 Generating Workflow 4: sales_weekly.xlsx")
products = [
"Product A",
"Product B",
"Product C",
"Product D",
"Product E",
"Product F",
"Product G",
"Product H",
]
# Generate one week of data
data = []
start_date = datetime(2024, 11, 4) # Monday
for day in range(7):
date = start_date + timedelta(days=day)
# Each product has 3-10 transactions per day
for product in products:
for _ in range(random.randint(3, 10)):
units = random.randint(1, 15)
price = random.randint(20, 150)
revenue = units * price
data.append(
{
"Date": date.strftime("%Y-%m-%d"),
"Product": product,
"Units": units,
"Revenue": revenue,
}
)
df = pd.DataFrame(data)
# Write to Excel with some formatting
output_file = output_dir / "sales_weekly.xlsx"
with pd.ExcelWriter(output_file, engine="openpyxl") as writer:
df.to_excel(writer, sheet_name="Weekly Sales", index=False)
# Format header
worksheet = writer.sheets["Weekly Sales"]
for cell in worksheet[1]:
cell.font = Font(bold=True)
print(f" ✅ Created {output_file} ({len(df)} rows)")
def main() -> None:
"""Generate all sample data files."""
print("=" * 60)
print("Generating Sample Data for LlamaSheets + Coding Agent Workflows")
print("=" * 60)
# Create output directory
output_dir = Path("input_data")
output_dir.mkdir(exist_ok=True)
# Generate data for each workflow
generate_workflow_1_data(output_dir)
generate_workflow_2_data(output_dir)
generate_workflow_3_data(output_dir)
generate_workflow_4_data(output_dir)
print("\n" + "=" * 60)
print("✅ All sample data generated!")
print("=" * 60)
print(f"\nFiles created in {output_dir.absolute()}:")
print("\nWorkflow 1 (Understanding a New Spreadsheet):")
print(" • financial_report_q1.xlsx")
print("\nWorkflow 2 (Generating Analysis Scripts):")
print(" • sales_northeast.xlsx")
print(" • sales_southeast.xlsx")
print(" • sales_west.xlsx")
print("\nWorkflow 3 (Using Cell Metadata):")
print(" • company_budget_2024.xlsx")
print("\nWorkflow 4 (Complete Automation):")
print(" • sales_weekly.xlsx")
print("\nYou can now use these files with the workflows in the documentation!")
if __name__ == "__main__":
main()
@@ -0,0 +1,308 @@
"""
LlamaSheets Agent with LlamaIndex
This example shows how to build an agent that can work with spreadsheet data
extracted by LlamaSheets using Python code execution.
The agent has minimal tools but maximum flexibility - it can execute arbitrary
pandas code against the extracted data, similar to a coding agent.
NOTE: Code execution should be handled safely in a sandboxed environment for security.
"""
import io
import json
import sys
from pathlib import Path
from typing import Any, Dict, Optional
import dotenv
import pandas as pd
from llama_index.core.agent import FunctionAgent, ToolCall, ToolCallResult, AgentStream
from llama_index.llms.openai import OpenAI
from workflows import Context
dotenv.load_dotenv()
# Global context for loaded dataframes
_dataframe_context: Dict[str, Any] = {}
# Helper function for initial agent context
def list_extracted_data(data_dir: str = "data") -> str:
"""
List all regions and metadata files extracted by LlamaSheets.
This helps discover what data is available to work with.
Args:
data_dir: Directory containing extracted parquet files (default: "data")
Returns:
JSON string with information about available files
"""
data_path = Path(data_dir)
if not data_path.exists():
return json.dumps({"error": f"Data directory '{data_dir}' not found"})
# Find all parquet and metadata files
region_files = list(data_path.glob("*_region_*.parquet"))
job_metadata_files = list(data_path.glob("*_job_metadata.json"))
regions = []
for region_file in region_files:
# Quick peek at dimensions
df = pd.read_parquet(region_file)
# Find corresponding metadata file
base_name = region_file.stem.replace("_region_", "_metadata_")
metadata_path = region_file.parent / f"{base_name}.parquet"
regions.append(
{
"region_file": str(region_file),
"metadata_file": str(metadata_path) if metadata_path.exists() else None,
"shape": {"rows": len(df), "columns": len(df.columns)},
"columns": list(df.columns),
}
)
result = {
"data_directory": str(data_path.absolute()),
"num_regions": len(regions),
"regions": regions,
"job_metadata_files": [str(f) for f in job_metadata_files],
}
return json.dumps(result, indent=2)
# Agent tool for code execution against dataframes
def execute_dataframe_code(
code: str, load_files: Optional[Dict[str, str]] = None
) -> str:
"""
Execute Python pandas code against LlamaSheets extracted data.
This tool allows flexible data analysis by executing arbitrary pandas code.
You can load parquet files, manipulate dataframes, and return results.
The code executes in a context where:
- pandas is available as 'pd'
- json is available for formatting output
- Previously loaded dataframes are accessible by their variable names
Args:
code: Python code to execute. Any print() statements or stdout/stderr
will be captured and returned. Optionally set a 'result' variable
for structured output.
load_files: Optional dict mapping variable names to file paths to load
Example: {"df": "data/sales_region_1.parquet",
"meta": "data/sales_metadata_1.parquet"}
Returns:
String containing:
- Any stdout/stderr output from the code execution
- The 'result' variable if it was set (formatted appropriately)
- Error message if execution failed
Example usage:
code = '''
# Load and inspect data
df = pd.read_parquet("data/sales_region_1.parquet")
print(f"Loaded {len(df)} rows")
result = {
"shape": df.shape,
"columns": list(df.columns),
"sample": df.head(3).to_dict(orient="records")
}
'''
"""
global _dataframe_context
# Capture stdout and stderr
stdout_capture = io.StringIO()
stderr_capture = io.StringIO()
old_stdout = sys.stdout
old_stderr = sys.stderr
try:
# Redirect stdout/stderr
sys.stdout = stdout_capture
sys.stderr = stderr_capture
# Create execution context with pandas, json, and previously loaded dfs
exec_context = {
"pd": pd,
"json": json,
"Path": Path,
**_dataframe_context, # Include previously loaded dataframes
}
# Load any requested files into context
if load_files:
for var_name, file_path in load_files.items():
if file_path.endswith(".parquet"):
exec_context[var_name] = pd.read_parquet(file_path)
# Also save to global context for future calls
_dataframe_context[var_name] = exec_context[var_name]
elif file_path.endswith(".json"):
with open(file_path, "r") as f:
exec_context[var_name] = json.load(f)
_dataframe_context[var_name] = exec_context[var_name]
# Execute the code
exec(code, exec_context)
# Restore stdout/stderr
sys.stdout = old_stdout
sys.stderr = old_stderr
# Collect output
stdout_output = stdout_capture.getvalue()
stderr_output = stderr_capture.getvalue()
output_parts = []
# Add stdout if any
if stdout_output:
output_parts.append(f"<stdout>{stdout_output}</stdout>")
# Add stderr if any
if stderr_output:
output_parts.append(f"<stderr>{stderr_output}</stderr>")
# Try to get a result (if code set a 'result' variable)
if "result" in exec_context:
result = exec_context["result"]
result_str = None
if isinstance(result, pd.DataFrame):
# Convert DataFrame to readable format
result_str = result.to_string()
elif isinstance(result, (dict, list)):
result_str = json.dumps(result, indent=2, default=str)
else:
result_str = str(result)
if result_str:
output_parts.append(f"<result_var>{result_str}</result_var>")
# Return combined output or success message
if output_parts:
return "\n\n".join(output_parts)
else:
return "Code executed successfully (no output or result)"
except Exception as e:
# Restore stdout/stderr in case of error
sys.stdout = old_stdout
sys.stderr = old_stderr
# Get any partial output
stdout_output = stdout_capture.getvalue()
stderr_output = stderr_capture.getvalue()
error_parts = []
if stdout_output:
error_parts.append(f"=== STDOUT (before error) ===\n{stdout_output}")
if stderr_output:
error_parts.append(f"=== STDERR (before error) ===\n{stderr_output}")
error_parts.append(f"=== ERROR ===\n{str(e)}")
error_parts.append(f"\n=== CODE ===\n{code}")
return "\n\n".join(error_parts)
def create_llamasheets_agent(
llm_model: str = "gpt-4.1", api_key: Optional[str] = None
) -> FunctionAgent:
# Initialize LLM
llm = OpenAI(model=llm_model, api_key=api_key)
# Create tools - just 4 simple but powerful tools
tools = [execute_dataframe_code]
# System prompt to guide the agent
available_regions = list_extracted_data()
system_prompt = f"""You are an AI assistant that helps analyze spreadsheet data extracted by LlamaSheets.
LlamaSheets extracts messy spreadsheets into clean parquet files with two types of outputs:
1. Region files (*_region_*.parquet) - The actual data with columns and rows
2. Metadata files (*_metadata_*.parquet) - Rich cell-level metadata including:
- Formatting: font_bold, font_italic, font_size, background_color_rgb
- Position: row_number, column_number, coordinate
- Type detection: data_type, is_date_like, is_percentage, is_currency
- Layout: is_in_first_row, is_merged_cell, horizontal_alignment
Your approach:
1. Use list_extracted_data() to discover available files
2. Use execute_dataframe_code() to load and analyze data with pandas
3. Use metadata to understand structure (bold = headers, colors = groups)
4. Use save_dataframe() to export results
Key tips:
- Bold cells in metadata often indicate headers
- Background colors often indicate groupings or departments
- Load both region and metadata files for complete analysis
- Write clear pandas code - you have full pandas functionality available
- Store results in variables for reuse across multiple code executions
Existing Processed Regions:
{available_regions}
"""
# Configure agent
return FunctionAgent(tools=tools, llm=llm, system_prompt=system_prompt)
async def main():
"""Example of using the LlamaSheets agent."""
# Create the agent
agent = create_llamasheets_agent()
ctx = Context(agent)
# Example queries the agent can handle:
queries = [
# Discovery
"What spreadsheet data is available?",
# Simple analysis
"Load the sales data and show me the first few rows with column info",
# Using metadata
"Find all bold cells in the metadata - these are likely headers",
]
# Example: Run a query
for query in queries:
print(f"\n=== Query: {query} ===")
handler = agent.run(query, ctx=ctx)
async for ev in handler.stream_events():
if isinstance(ev, ToolCall):
tool_kwargs_str = (
str(ev.tool_kwargs)[:500] + " ..."
if len(str(ev.tool_kwargs)) > 500
else str(ev.tool_kwargs)
)
print(f"\n[Tool Call] {ev.tool_name} with args:\n{tool_kwargs_str}\n\n")
elif isinstance(ev, ToolCallResult):
result_str = (
str(ev.tool_output)[:500] + " ..."
if len(str(ev.tool_output)) > 500
else str(ev.tool_output)
)
print(f"\n[Tool Result] {ev.tool_name}:\n{result_str}\n\n")
elif isinstance(ev, AgentStream):
print(ev.delta, end="", flush=True)
_ = await handler
print("=== End Query ===\n")
if __name__ == "__main__":
import asyncio
asyncio.run(main())
@@ -0,0 +1,7 @@
llama-cloud-services # LlamaSheets SDK
llama-index-core
llama-index-llms-openai
pandas>=2.0.0
pyarrow>=12.0.0
openpyxl>=3.0.0 # For Excel file support
matplotlib>=3.7.0 # For visualizations (optional)
@@ -0,0 +1,100 @@
"""Helper script to extract spreadsheets using LlamaSheets."""
import asyncio
import json
import os
import dotenv
from pathlib import Path
from llama_cloud_services.beta.sheets import LlamaSheets
from llama_cloud_services.beta.sheets.types import (
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
dotenv.load_dotenv()
async def extract_spreadsheet(
file_path: str, output_dir: str = "data", generate_metadata: bool = True
) -> dict:
"""Extract a spreadsheet using LlamaSheets."""
client = LlamaSheets(
base_url="https://api.cloud.llamaindex.ai",
api_key=os.getenv("LLAMA_CLOUD_API_KEY"),
)
print(f"Extracting {file_path}...")
# Extract regions
config = SpreadsheetParsingConfig(
sheet_names=None, # Extract all sheets
generate_additional_metadata=generate_metadata,
)
job_result = await client.aextract_regions(file_path, config=config)
print(f"Extracted {len(job_result.regions)} region(s)")
# Create output directory
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
# Get base name for files
base_name = Path(file_path).stem
# Save job metadata
job_metadata_path = output_path / f"{base_name}_job_metadata.json"
with open(job_metadata_path, "w") as f:
json.dump(job_result.model_dump(mode="json"), f, indent=2)
print(f"Saved job metadata to {job_metadata_path}")
# Download each region
for idx, region in enumerate(job_result.regions, 1):
sheet_name = region.sheet_name.replace(" ", "_")
# Download region data
region_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=region.region_type,
)
region_path = output_path / f"{base_name}_region_{idx}_{sheet_name}.parquet"
with open(region_path, "wb") as f:
f.write(region_bytes)
print(f" Table {idx}: {region_path}")
# Download metadata
metadata_bytes = await client.adownload_region_result(
job_id=job_result.id,
region_id=region.region_id,
result_type=SpreadsheetResultType.CELL_METADATA,
)
metadata_path = output_path / f"{base_name}_metadata_{idx}_{sheet_name}.parquet"
with open(metadata_path, "wb") as f:
f.write(metadata_bytes)
print(f" Metadata {idx}: {metadata_path}")
print(f"\nAll files saved to {output_path}/")
return job_result.model_dump(mode="json")
if __name__ == "__main__":
import sys
if len(sys.argv) < 2:
print("Usage: python scripts/extract.py <spreadsheet_file>")
sys.exit(1)
file_path = sys.argv[1]
if not Path(file_path).exists():
print(f"❌ File not found: {file_path}")
sys.exit(1)
result = asyncio.run(extract_spreadsheet(file_path))
print(f"\n✅ Extraction complete! Job ID: {result['id']}")
+2 -2
View File
@@ -8,7 +8,7 @@
"scripts": {
"pre-commit-version": "pnpm changeset",
"version": "./scripts/changeset-version.py version",
"publish": "./scripts/changeset-version.py publish"
"publish": "./scripts/changeset-version.py publish --tag"
},
"devDependencies": {
"prettier": "^3.6.2",
@@ -19,7 +19,7 @@
"lint-staged": {
"ts/llama_cloud_services/src/**/*.{ts,tsx,js,jsx}": [
"pnpm --filter llama-cloud-services exec eslint --fix",
"pnpm --filter llama-cloud-services exec prettier --write"
"pnpm --filter llama-cloud-services exec prettier --write src/ tests/"
]
},
"packageManager": "pnpm@10.11.1+sha512.e519b9f7639869dc8d5c3c5dfef73b3f091094b0a006d7317353c72b124e80e1afd429732e28705ad6bfa1ee879c1fce46c128ccebd3192101f43dd67c667912"
+3383 -20
View File
File diff suppressed because it is too large Load Diff
+1
View File
@@ -1,3 +1,4 @@
packages:
- "ts/*"
- "py"
- "py/*"
+67
View File
@@ -1,5 +1,72 @@
# llama-cloud-services-py
## 0.6.81
### Patch Changes
- f3233de: Propagate retrieval metadata to retriever nodes
## 0.6.80
### Patch Changes
- 0506c88: Moved ClassifyClient to LlamaClassify (backward compatible)
## 0.6.79
### Patch Changes
- e020e3e: Remove unneeded organization_id param from beta classifier client
## 0.6.78
### Patch Changes
- 9f1ef4e: Fix extract
## 0.6.77
### Patch Changes
- 407292b: Now return partial results on job failure
## 0.6.76
### Patch Changes
- 4f24f53: Add aggressive_table_extraction flag in python sdk
## 0.6.75
### Patch Changes
- f81532e: Safest types possible for parse
## 0.6.74
### Patch Changes
- 1bf5223: Fix default bbox values
- 24166dc: Now only escape single dollar signs - preserve double for latex equations
## 0.6.73
### Patch Changes
- e6a7939: Loosen packaging dep requirement
## 0.6.72
### Patch Changes
- ad6734b: Fixup and test versioning
## 0.6.71
### Patch Changes
- 51011b9: Escape dollar signs in jupyter notebooks
## 0.6.70
### Patch Changes
+5 -1
View File
@@ -1,5 +1,7 @@
from llama_cloud_services.parse import LlamaParse
from llama_cloud_services.extract import LlamaExtract, ExtractionAgent, SourceText
from llama_cloud_services.extract import LlamaExtract, ExtractionAgent
from llama_cloud_services.testing_utils import FakeLlamaCloudServer
from llama_cloud_services.utils import SourceText, FileInput
from llama_cloud_services.constants import EU_BASE_URL
from llama_cloud_services.index import (
LlamaCloudCompositeRetriever,
@@ -12,8 +14,10 @@ __all__ = [
"LlamaExtract",
"ExtractionAgent",
"SourceText",
"FileInput",
"EU_BASE_URL",
"LlamaCloudIndex",
"LlamaCloudRetriever",
"LlamaCloudCompositeRetriever",
"FakeLlamaCloudServer",
]
@@ -0,0 +1,11 @@
from llama_cloud_services.beta.classifier.client import LlamaClassify, ClassifyClient
from llama_cloud_services.beta.classifier.types import ClassifyJobResultsWithFiles
from llama_cloud_services.utils import SourceText, FileInput
__all__ = [
"LlamaClassify",
"ClassifyClient",
"ClassifyJobResultsWithFiles",
"SourceText",
"FileInput",
]
+152 -38
View File
@@ -1,6 +1,7 @@
import asyncio
import time
from typing import Optional
import warnings
from typing import Optional, List, Union
from pydantic import BaseModel
from llama_cloud.client import AsyncLlamaCloud
from llama_cloud.types import (
@@ -14,7 +15,11 @@ from llama_cloud.types import (
from llama_cloud.resources.classifier.client import OMIT
from llama_cloud_services.files.client import FileClient
from llama_cloud_services.constants import POLLING_TIMEOUT_SECONDS
from llama_cloud_services.utils import is_terminal_status, augment_async_errors
from llama_cloud_services.utils import (
is_terminal_status,
augment_async_errors,
FileInput,
)
from llama_index.core.async_utils import DEFAULT_NUM_WORKERS, run_jobs
from llama_cloud_services.beta.classifier.types import (
ClassifyJobResultsWithFiles,
@@ -26,7 +31,7 @@ class ClassificationOutput(BaseModel):
classification: str
class ClassifyClient:
class LlamaClassify:
"""
Experimental - Client for interacting with the LlamaCloud Classifier API.
The Classification API is currently in beta and may change in the future without notice.
@@ -34,7 +39,6 @@ class ClassifyClient:
Args:
client: The LlamaCloud client to use.
project_id: The project ID to use.
organization_id: The organization ID to use.
polling_interval: The interval to poll for job completion in seconds.
polling_timeout: The timeout for the job to complete in seconds.
"""
@@ -43,15 +47,13 @@ class ClassifyClient:
self,
client: AsyncLlamaCloud,
project_id: Optional[str] = None,
organization_id: Optional[str] = None,
polling_interval: float = 1.0,
polling_timeout: float = POLLING_TIMEOUT_SECONDS,
):
self.client = client
self.project_id = project_id
self.organization_id = organization_id
self.polling_interval = polling_interval
self.file_client = FileClient(client, project_id, organization_id)
self.file_client = FileClient(client, project_id)
self.polling_timeout = polling_timeout
@classmethod
@@ -59,7 +61,6 @@ class ClassifyClient:
cls,
api_key: str,
project_id: Optional[str] = None,
organization_id: Optional[str] = None,
base_url: Optional[str] = None,
) -> "ClassifyClient":
"""
@@ -69,7 +70,6 @@ class ClassifyClient:
return cls(
client,
project_id,
organization_id,
)
async def acreate_classify_job(
@@ -96,7 +96,6 @@ class ClassifyClient:
file_ids=file_ids,
parsing_configuration=parsing_configuration or OMIT,
project_id=self.project_id,
organization_id=self.organization_id,
)
def create_classify_job(
@@ -147,7 +146,6 @@ class ClassifyClient:
results = await self.client.classifier.get_classification_job_results(
classify_job_with_status.id,
project_id=self.project_id,
organization_id=self.organization_id,
)
return results
@@ -166,6 +164,98 @@ class ClassifyClient:
)
)
async def aclassify(
self,
rules: list[ClassifierRule],
files: Union[FileInput, List[FileInput]],
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResultsWithFiles:
"""
Classify one or more files from various input types.
Args:
rules: The rules to use for classification.
files: The file(s) to classify. Can be a single file or list of files. Each can be:
- str/Path: File path
- SourceText: Text content or file with explicit filename
- File: Already uploaded file
- BufferedIOBase: File-like object
parsing_configuration: The parsing configuration to use for classification.
raise_on_error: Whether to raise an error if the classification job fails.
workers: Number of parallel workers for uploading files.
show_progress: Whether to show progress bars.
Returns:
The results of the classification job with file metadata.
"""
# Normalize to list
if not isinstance(files, list):
files = [files]
# Upload all files
coroutines = [
self.file_client.upload_content(file_input) for file_input in files
]
uploaded_files: List[File] = await run_jobs(
coroutines,
show_progress=show_progress,
workers=workers,
desc="Uploading files for classification",
)
# Classify
results = await self.aclassify_file_ids(
rules,
[file.id for file in uploaded_files],
parsing_configuration,
raise_on_error,
)
return ClassifyJobResultsWithFiles.from_classify_job_results(
results, uploaded_files
)
def classify(
self,
rules: list[ClassifierRule],
files: Union[FileInput, List[FileInput]],
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResultsWithFiles:
"""
Classify one or more files from various input types (synchronous version).
Args:
rules: The rules to use for classification.
files: The file(s) to classify. Can be a single file or list of files. Each can be:
- str/Path: File path
- SourceText: Text content or file with explicit filename
- File: Already uploaded file
- BufferedIOBase: File-like object
parsing_configuration: The parsing configuration to use for classification.
raise_on_error: Whether to raise an error if the classification job fails.
workers: Number of parallel workers for uploading files.
show_progress: Whether to show progress bars.
Returns:
The results of the classification job with file metadata.
"""
with augment_async_errors():
return asyncio.run(
self.aclassify(
rules,
files,
parsing_configuration,
raise_on_error,
workers,
show_progress,
)
)
async def aclassify_file_path(
self,
rules: list[ClassifierRule],
@@ -173,11 +263,17 @@ class ClassifyClient:
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResultsWithFiles:
file = await self.file_client.upload_file(file_input_path)
results = await self.aclassify_file_ids(
rules, [file.id], parsing_configuration, raise_on_error
"""
Deprecated: Use aclassify() instead.
"""
warnings.warn(
"aclassify_file_path is deprecated, use aclassify() instead",
DeprecationWarning,
stacklevel=2,
)
return await self.aclassify(
rules, file_input_path, parsing_configuration, raise_on_error
)
return ClassifyJobResultsWithFiles.from_classify_job_results(results, [file])
def classify_file_path(
self,
@@ -186,12 +282,17 @@ class ClassifyClient:
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResultsWithFiles:
with augment_async_errors():
return asyncio.run(
self.aclassify_file_path(
rules, file_input_path, parsing_configuration, raise_on_error
)
)
"""
Deprecated: Use classify() instead.
"""
warnings.warn(
"classify_file_path is deprecated, use classify() instead",
DeprecationWarning,
stacklevel=2,
)
return self.classify(
rules, file_input_path, parsing_configuration, raise_on_error
)
async def aclassify_file_paths(
self,
@@ -202,17 +303,22 @@ class ClassifyClient:
workers: int = DEFAULT_NUM_WORKERS,
show_progress: bool = False,
) -> ClassifyJobResultsWithFiles:
coroutines = [self.file_client.upload_file(path) for path in file_input_paths]
files: list[File] = await run_jobs(
coroutines,
show_progress=show_progress,
workers=workers,
desc="Uploading files for classification",
"""
Deprecated: Use aclassify() instead.
"""
warnings.warn(
"aclassify_file_paths is deprecated, use aclassify() instead",
DeprecationWarning,
stacklevel=2,
)
results = await self.aclassify_file_ids(
rules, [file.id for file in files], parsing_configuration, raise_on_error
return await self.aclassify(
rules,
file_input_paths,
parsing_configuration,
raise_on_error,
workers,
show_progress,
)
return ClassifyJobResultsWithFiles.from_classify_job_results(results, files)
def classify_file_paths(
self,
@@ -221,12 +327,17 @@ class ClassifyClient:
parsing_configuration: Optional[ClassifyParsingConfiguration] = None,
raise_on_error: bool = True,
) -> ClassifyJobResultsWithFiles:
with augment_async_errors():
return asyncio.run(
self.aclassify_file_paths(
rules, file_input_paths, parsing_configuration, raise_on_error
)
)
"""
Deprecated: Use classify() instead.
"""
warnings.warn(
"classify_file_paths is deprecated, use classify() instead",
DeprecationWarning,
stacklevel=2,
)
return self.classify(
rules, file_input_paths, parsing_configuration, raise_on_error
)
async def wait_for_job_completion(self, job_id: str) -> ClassifyJob:
"""
@@ -241,7 +352,7 @@ class ClassifyClient:
The classify job with status.
"""
job = await self.client.classifier.get_classify_job(
job_id, project_id=self.project_id, organization_id=self.organization_id
job_id, project_id=self.project_id
)
start_time = time.time()
while not is_terminal_status(job.status):
@@ -252,6 +363,9 @@ class ClassifyClient:
)
await asyncio.sleep(self.polling_interval)
job = await self.client.classifier.get_classify_job(
job_id, project_id=self.project_id, organization_id=self.organization_id
job_id, project_id=self.project_id
)
return job
ClassifyClient = LlamaClassify
@@ -0,0 +1,43 @@
"""LlamaCloud Spreadsheet API SDK
This module provides a Python SDK for the LlamaCloud Spreadsheet API.
"""
from llama_cloud_services.beta.sheets.client import (
LlamaSheets,
SpreadsheetAPIError,
SpreadsheetJobError,
SpreadsheetTimeoutError,
)
from llama_cloud_services.beta.sheets.types import (
ExtractedRegionSummary,
FileUploadResponse,
JobStatus,
PresignedUrlResponse,
SpreadsheetJob,
SpreadsheetJobResult,
SpreadsheetParseResult,
SpreadsheetParsingConfig,
SpreadsheetResultType,
WorksheetMetadata,
)
__all__ = [
# Client
"LlamaSheets",
# Exceptions
"SpreadsheetAPIError",
"SpreadsheetJobError",
"SpreadsheetTimeoutError",
# Types
"ExtractedRegionSummary",
"FileUploadResponse",
"JobStatus",
"PresignedUrlResponse",
"SpreadsheetJob",
"SpreadsheetJobResult",
"SpreadsheetParseResult",
"SpreadsheetParsingConfig",
"SpreadsheetResultType",
"WorksheetMetadata",
]
@@ -0,0 +1,520 @@
from __future__ import annotations
import asyncio
import io
import os
import time
from typing import TYPE_CHECKING
import httpx
from llama_cloud.client import AsyncLlamaCloud
from tenacity import (
AsyncRetrying,
retry_if_exception,
stop_after_attempt,
wait_exponential,
)
from llama_cloud_services.beta.sheets.types import (
FileUploadResponse,
JobStatus,
PresignedUrlResponse,
SpreadsheetJob,
SpreadsheetJobResult,
SpreadsheetParsingConfig,
SpreadsheetResultType,
)
from llama_cloud_services.constants import BASE_URL
from llama_cloud_services.files.client import FileClient
from llama_cloud_services.utils import (
augment_async_errors,
FileInput,
)
if TYPE_CHECKING:
import pandas as pd
def _should_retry_exception(exception: BaseException) -> bool:
"""Determine if an exception should be retried."""
if isinstance(exception, httpx.HTTPStatusError):
return exception.response.status_code in (429, 500, 502, 503, 504)
return False
class SpreadsheetAPIError(Exception):
"""Base exception for spreadsheet API errors"""
pass
class SpreadsheetJobError(SpreadsheetAPIError):
"""Exception raised when a spreadsheet job fails"""
pass
class SpreadsheetTimeoutError(SpreadsheetAPIError):
"""Exception raised when a job times out"""
pass
class LlamaSheets:
"""Client for the LlamaCloud Spreadsheet API"""
def __init__(
self,
api_key: str | None = None,
base_url: str | None = None,
max_timeout: int = 300,
poll_interval: int = 5,
max_retries: int = 3,
async_httpx_client: httpx.AsyncClient | None = None,
) -> None:
"""Initialize the LlamaSheets client.
Args:
api_key: API key for authentication. If not provided, will use LLAMA_CLOUD_API_KEY env var
base_url: Base URL for the API
max_timeout: Maximum time to wait for job completion in seconds
poll_interval: Interval between status checks in seconds
max_retries: Maximum number of retries for failed requests
async_httpx_client: Optional custom async httpx client
"""
self.api_key = api_key or os.environ.get("LLAMA_CLOUD_API_KEY")
if not self.api_key:
raise ValueError(
"An API key must be provided either as an argument or via the LLAMA_CLOUD_API_KEY environment variable."
)
base_url = base_url or os.environ.get("LLAMA_CLOUD_BASE_URL", BASE_URL)
self.base_url = str(base_url).rstrip("/")
self.max_timeout = max_timeout
self.poll_interval = poll_interval
self.max_retries = max_retries
self._async_client: httpx.AsyncClient | None = async_httpx_client
self._files_client = FileClient(
AsyncLlamaCloud(
token=self.api_key,
base_url=self.base_url,
httpx_client=async_httpx_client,
)
)
def _get_async_client(self) -> httpx.AsyncClient:
"""Get or create the async httpx client"""
if self._async_client is None:
self._async_client = httpx.AsyncClient(
timeout=httpx.Timeout(60.0),
follow_redirects=True,
)
return self._async_client
def _get_headers(self) -> dict[str, str]:
"""Get common headers for API requests"""
return {
"Authorization": f"Bearer {self.api_key}",
"Content-Type": "application/json",
}
# Sync methods
def upload_file(
self, file_obj: FileInput, file_name: str | None = None
) -> FileUploadResponse:
"""Upload a file to the Files API.
Args:
file_obj: File to upload (path, bytes, or file-like object)
file_name: Optional name for the uploaded filename
Returns:
FileUploadResponse with the uploaded file ID
"""
with augment_async_errors():
return asyncio.run(self.aupload_file(file_obj))
def create_job(
self,
file_id: str,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJob:
"""Create a new spreadsheet parsing job.
Args:
file_id: ID of the uploaded file
config: Parsing configuration
Returns:
SpreadsheetJob with job details
"""
with augment_async_errors():
return asyncio.run(self.acreate_job(file_id, config))
def get_job(
self, job_id: str, include_results_metadata: bool = True
) -> SpreadsheetJobResult:
"""Get the status of a spreadsheet parsing job.
Args:
job_id: ID of the job
include_results_metadata: Whether to include results metadata in the response
Returns:
SpreadsheetJobResult with job status and optionally results
"""
with augment_async_errors():
return asyncio.run(self.aget_job(job_id, include_results_metadata))
def wait_for_completion(self, job_id: str) -> SpreadsheetJobResult:
"""Wait for a job to complete by polling.
Args:
job_id: ID of the job to wait for
Returns:
SpreadsheetJobResult when job is complete
Raises:
SpreadsheetTimeoutError: If job doesn't complete within max_timeout
SpreadsheetJobError: If job fails
"""
with augment_async_errors():
return asyncio.run(self.await_for_completion(job_id))
def download_region_result(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> bytes:
"""Download a region result (either region data or cell metadata).
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
Raw bytes of the parquet file
"""
with augment_async_errors():
return asyncio.run(
self.adownload_region_result(job_id, region_id, result_type)
)
def download_region_as_dataframe(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> "pd.DataFrame":
"""Download a region result as a pandas DataFrame.
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
pandas DataFrame
"""
with augment_async_errors():
return asyncio.run(
self.adownload_region_as_dataframe(job_id, region_id, result_type)
)
def extract_regions(
self,
file_obj: FileInput,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJobResult:
"""High-level method to parse a spreadsheet file.
This method handles the entire workflow:
1. Upload the file
2. Create a parsing job
3. Wait for completion
4. Return results
Args:
file_obj: File to parse (path, bytes, or file-like object)
config: Parsing configuration
Returns:
SpreadsheetJobResult with parsing results
"""
with augment_async_errors():
return asyncio.run(self.aextract_regions(file_obj, config))
# Async methods
async def aupload_file(
self, file_obj: FileInput, file_name: str | None = None
) -> FileUploadResponse:
"""Upload a file to the Files API.
Args:
file_obj: File to upload (path, bytes, or file-like object)
file_name: Optional name for the uploaded filename
Returns:
FileUploadResponse with the uploaded file ID
"""
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
return await self._files_client.upload_content(
file_obj, external_file_id=file_name
)
except Exception as e:
raise SpreadsheetAPIError(f"Failed to upload file: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def acreate_job(
self,
file_id: str,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJob:
"""Create a new spreadsheet parsing job.
Args:
file_id: ID of the uploaded file
config: Parsing configuration
Returns:
SpreadsheetJob with job details
"""
if config is None:
config = SpreadsheetParsingConfig()
elif isinstance(config, dict):
config = SpreadsheetParsingConfig.model_validate(config)
if not isinstance(config, SpreadsheetParsingConfig):
raise ValueError(
"config must be a dict or SpreadsheetParsingConfig instance"
)
payload = {
"file_id": file_id,
"config": config.model_dump(mode="json", exclude_none=True),
}
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
response = await client.post(
f"{self.base_url}/api/v1/beta/sheets/jobs",
headers=self._get_headers(),
json=payload,
)
response.raise_for_status()
return SpreadsheetJob.model_validate(response.json())
except Exception as e:
raise SpreadsheetAPIError(f"Failed to create job: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def aget_job(
self, job_id: str, include_results_metadata: bool = True
) -> SpreadsheetJobResult:
"""Get the status of a spreadsheet parsing job.
Args:
job_id: ID of the job
include_results_metadata: Whether to include results in the response
Returns:
SpreadsheetJobResult with job status and optionally results
"""
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
response = await client.get(
f"{self.base_url}/api/v1/beta/sheets/jobs/{job_id}",
headers=self._get_headers(),
params={"include_results": include_results_metadata},
)
response.raise_for_status()
return SpreadsheetJobResult.model_validate(response.json())
except Exception as e:
raise SpreadsheetAPIError(f"Failed to get job status: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def await_for_completion(self, job_id: str) -> SpreadsheetJobResult:
"""Wait for a job to complete by polling.
Args:
job_id: ID of the job to wait for
Returns:
SpreadsheetJobResult when job is complete
Raises:
SpreadsheetTimeoutError: If job doesn't complete within max_timeout
SpreadsheetJobError: If job fails
"""
start_time = time.time()
while (time.time() - start_time) < self.max_timeout:
job_result = await self.aget_job(job_id, include_results_metadata=True)
if job_result.status in (
JobStatus.SUCCESS,
JobStatus.PARTIAL_SUCCESS,
JobStatus.ERROR,
JobStatus.FAILURE,
):
if job_result.status in (JobStatus.SUCCESS, JobStatus.PARTIAL_SUCCESS):
return job_result
else:
error_msg = f"Job failed with status: {job_result.status}"
if job_result.errors:
error_msg += f"\nErrors: {', '.join(job_result.errors)}"
raise SpreadsheetJobError(error_msg)
await asyncio.sleep(self.poll_interval)
raise SpreadsheetTimeoutError(
f"Job did not complete within {self.max_timeout} seconds"
)
async def adownload_region_result(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> bytes:
"""Download a region result (either region data or cell metadata).
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
Raw bytes of the parquet file
"""
# Get presigned URL
presigned_response = None
result_type_str = str(result_type)
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
client = self._get_async_client()
response = await client.get(
f"{self.base_url}/api/v1/beta/sheets/jobs/{job_id}/regions/{region_id}/result/{result_type_str}",
headers=self._get_headers(),
)
response.raise_for_status()
presigned_response = PresignedUrlResponse.model_validate(
response.json()
)
except Exception as e:
raise SpreadsheetAPIError(f"Failed to get presigned URL: {e}") from e
# Download using presigned URL
if presigned_response is None:
raise SpreadsheetAPIError("Failed to obtain presigned URL.")
try:
async for attempt in AsyncRetrying(
stop=stop_after_attempt(self.max_retries),
wait=wait_exponential(multiplier=1, min=1, max=32),
retry=retry_if_exception(_should_retry_exception),
reraise=True,
):
with attempt:
download_response = await client.get(presigned_response.url)
download_response.raise_for_status()
return download_response.content
except Exception as e:
raise SpreadsheetAPIError(f"Failed to download result: {e}") from e
raise RuntimeError("Tenacity did not execute")
async def adownload_region_as_dataframe(
self,
job_id: str,
region_id: str,
result_type: SpreadsheetResultType = SpreadsheetResultType.TABLE,
) -> "pd.DataFrame":
"""Download a region result as a pandas DataFrame.
Args:
job_id: ID of the job
region_id: ID of the region
result_type: Type of result to download (region or cell_metadata)
Returns:
pandas DataFrame
"""
import pandas as pd
parquet_bytes = await self.adownload_region_result(
job_id, region_id, result_type
)
return pd.read_parquet(io.BytesIO(parquet_bytes))
async def aextract_regions(
self,
file_obj: FileInput,
config: dict | SpreadsheetParsingConfig | None = None,
) -> SpreadsheetJobResult:
"""High-level method to parse a spreadsheet file.
This method handles the entire workflow:
1. Upload the file
2. Create a parsing job
3. Wait for completion
4. Return results
Args:
file_obj: File to parse (path, bytes, or file-like object)
config: Parsing configuration
Returns:
SpreadsheetJobResult with parsing results
"""
# Upload file
file_response = await self.aupload_file(file_obj)
# Create job
job = await self.acreate_job(file_response.id, config)
# Wait for completion
return await self.await_for_completion(job.id)
async def aclose(self) -> None:
"""Close all HTTP clients (async)"""
if self._async_client:
await self._async_client.aclose()
async def __aenter__(self) -> "LlamaSheets":
return self
async def __aexit__(self, _exc_type, _exc_val, _exc_tb) -> None: # type: ignore
await self.aclose()
@@ -0,0 +1,158 @@
from __future__ import annotations
from datetime import datetime
from enum import Enum
from pydantic import BaseModel, ConfigDict, Field, field_validator
class SpreadsheetResultType(str, Enum):
TABLE = "table"
EXTRA = "extra"
CELL_METADATA = "cell_metadata"
def __str__(self) -> str:
return self.value
class ExtractedRegionSummary(BaseModel):
"""A summary of a single extracted region from a spreadsheet"""
region_id: str = Field(
...,
description="Unique identifier for this region within the file",
)
sheet_name: str = Field(..., description="Worksheet name where region was found")
location: str = Field(..., description="Location of the region in the spreadsheet")
title: str | None = Field(None, description="Generated title for the region")
description: str | None = Field(
None, description="Generated description of the region"
)
region_type: SpreadsheetResultType = Field(
..., description="Type of the extracted region"
)
class WorksheetMetadata(BaseModel):
"""Metadata about a worksheet in a spreadsheet"""
sheet_name: str = Field(..., description="Name of the worksheet")
title: str | None = Field(None, description="Generated title for the worksheet")
description: str | None = Field(
None, description="Generated description of the worksheet"
)
class SpreadsheetParseResult(BaseModel):
"""Result of parsing a single spreadsheet file"""
success: bool = Field(..., description="Whether parsing was successful")
file_name: str = Field(..., description="Original filename")
regions: list[ExtractedRegionSummary] = Field(
default_factory=list, description="All successfully extracted regions"
)
worksheet_metadata: list[WorksheetMetadata] = Field(
default_factory=list, description="Metadata for each processed worksheet"
)
# Error information
errors: list[str] = Field(
default_factory=list, description="Any errors encountered during parsing"
)
class SpreadsheetParsingConfig(BaseModel):
"""Configuration for spreadsheet parsing and region extraction"""
model_config = ConfigDict(extra="forbid")
sheet_names: list[str] | None = Field(
default=None,
description="The names of the sheets to extract regions from. If empty, the default sheet is extracted.",
)
include_hidden_cells: bool = Field(
default=True,
description="Whether to include hidden cells when extracting regions from the spreadsheet.",
)
extraction_range: str | None = Field(
default=None,
description="A1 notation of the range to extract a single region from. If None, the entire sheet is used.",
)
generate_additional_metadata: bool = Field(
default=True,
description="Whether to generate additional metadata (title, description) for each extracted region.",
)
use_experimental_processing: bool = Field(
default=False,
description="Enables experimental processing. Accuracy may be impacted.",
)
class SpreadsheetJob(BaseModel):
"""A spreadsheet parsing job"""
id: str = Field(..., description="The ID of the job")
user_id: str = Field(..., description="The ID of the user")
project_id: str = Field(..., description="The ID of the project")
file: dict = Field(..., description="The file object being parsed")
config: SpreadsheetParsingConfig = Field(
..., description="Configuration for the parsing job"
)
status: str = Field(..., description="The status of the parsing job")
created_at: str = Field(..., description="When the job was created")
updated_at: str = Field(..., description="When the job was last updated")
@field_validator("created_at", "updated_at", mode="before")
def validate_dates(cls, v: str) -> str:
"""Validate that the dates are in the correct format"""
if isinstance(v, datetime):
return v.isoformat()
else:
return v
class SpreadsheetJobResult(SpreadsheetJob):
"""A spreadsheet parsing job result."""
# Results are included when the job is complete
success: bool | None = Field(
None, description="Whether the job completed successfully"
)
regions: list[ExtractedRegionSummary] = Field(
default_factory=list,
description="All extracted regions (populated when job is complete)",
)
worksheet_metadata: list[WorksheetMetadata] = Field(
default_factory=list,
description="Metadata for each processed worksheet (populated when job is complete)",
)
errors: list[str] = Field(
default_factory=list, description="Any errors encountered"
)
class JobStatus(str, Enum):
"""Status of a spreadsheet parsing job"""
PENDING = "PENDING"
IN_PROGRESS = "IN_PROGRESS"
SUCCESS = "SUCCESS"
PARTIAL_SUCCESS = "PARTIAL_SUCCESS"
ERROR = "ERROR"
FAILURE = "FAILURE"
class PresignedUrlResponse(BaseModel):
"""Response containing a presigned URL for downloading results"""
url: str = Field(..., description="The presigned URL for downloading")
class FileUploadResponse(BaseModel):
"""Response from uploading a file"""
id: str = Field(..., description="The ID of the uploaded file")
name: str = Field(..., description="The name of the file")
project_id: str = Field(..., description="The project ID")
user_id: str = Field(..., description="The user ID")
+1
View File
@@ -1,2 +1,3 @@
BASE_URL = "https://api.cloud.llamaindex.ai"
EU_BASE_URL = "https://api.cloud.eu.llamaindex.ai"
POLLING_TIMEOUT_SECONDS = 300.0
+2 -1
View File
@@ -2,15 +2,16 @@ from llama_cloud_services.extract.extract import (
LlamaExtract,
ExtractConfig,
ExtractionAgent,
SourceText,
ExtractTarget,
ExtractMode,
)
from llama_cloud_services.utils import SourceText, FileInput
__all__ = [
"LlamaExtract",
"ExtractionAgent",
"SourceText",
"FileInput",
"ExtractConfig",
"ExtractTarget",
"ExtractMode",
+7 -100
View File
@@ -2,10 +2,9 @@ import asyncio
import base64
import os
import time
from io import BufferedIOBase, BufferedReader, BytesIO, TextIOWrapper
from io import BufferedIOBase, TextIOWrapper
from pathlib import Path
from typing import List, Optional, Type, Union, Coroutine, Any, TypeVar
import secrets
import warnings
import httpx
from pydantic import BaseModel
@@ -33,7 +32,8 @@ from llama_cloud_services.extract.utils import (
JSONObjectType,
ExperimentalWarning,
)
from llama_cloud_services.utils import augment_async_errors
from llama_cloud_services.utils import augment_async_errors, SourceText, FileInput
from llama_cloud_services.files.client import FileClient
from llama_index.core.schema import BaseComponent
from llama_index.core.async_utils import run_jobs
from llama_index.core.bridge.pydantic import Field, PrivateAttr
@@ -188,46 +188,6 @@ async def _wait_for_job_result(
)
class SourceText:
def __init__(
self,
*,
file: Union[bytes, BufferedIOBase, TextIOWrapper, str, Path, None] = None,
text_content: Optional[str] = None,
filename: Optional[str] = None,
):
self.file = file
self.filename = filename
self.text_content = text_content
self._validate()
def _validate(self) -> None:
"""Ensure filename is provided when needed."""
if not ((self.file is None) ^ (self.text_content is None)):
raise ValueError("Either file or text_content must be provided.")
if self.text_content is not None:
if not self.filename:
random_hex = secrets.token_hex(4)
self.filename = f"text_input_{random_hex}.txt"
return
if isinstance(self.file, (bytes, BufferedIOBase, TextIOWrapper)):
if not self.filename and hasattr(self.file, "name"):
self.filename = os.path.basename(str(self.file.name))
elif not hasattr(self.file, "name") and self.filename is None:
raise ValueError(
"filename must be provided when file is bytes or a file-like object without a name"
)
elif isinstance(self.file, (str, Path)):
if not self.filename:
self.filename = os.path.basename(str(self.file))
else:
raise ValueError(f"Unsupported file type: {type(self.file)}")
FileInput = Union[str, Path, BufferedIOBase, SourceText, File]
def run_in_thread(
coro: Coroutine[Any, Any, T],
thread_pool: ThreadPoolExecutor,
@@ -320,6 +280,7 @@ class ExtractionAgent:
self._thread_pool = ThreadPoolExecutor(
max_workers=min(10, (os.cpu_count() or 1) + 4)
)
self._file_client = FileClient(client, project_id, organization_id)
@property
def id(self) -> str:
@@ -369,65 +330,11 @@ class ExtractionAgent:
ValueError: If filename is not provided for bytes input or for file-like objects
without a name attribute.
"""
file_contents: Optional[Union[BufferedIOBase, BytesIO]] = None
try:
if file_input.text_content is not None:
# Handle direct text content
file_contents = BytesIO(file_input.text_content.encode("utf-8"))
elif isinstance(file_input.file, TextIOWrapper):
# Handle text-based IO objects
file_contents = BytesIO(file_input.file.read().encode("utf-8"))
elif isinstance(file_input.file, (str, Path)):
# Handle file paths
file_contents = open(file_input.file, "rb")
elif isinstance(file_input.file, bytes):
# Handle bytes
file_contents = BytesIO(file_input.file)
elif isinstance(file_input.file, BufferedIOBase):
# Handle binary IO objects
file_contents = file_input.file
else:
raise ValueError(f"Unsupported file type: {type(file_input.file)}")
# Add name attribute to file object if needed
if not hasattr(file_contents, "name"):
file_contents.name = file_input.filename # type: ignore
return await self._client.files.upload_file(
project_id=self._project_id, upload_file=file_contents
)
finally:
if file_contents is not None and isinstance(
file_contents, (BufferedReader, BytesIO)
):
file_contents.close()
return await self._file_client.upload_content(file_input)
async def _upload_file(self, file_input: FileInput) -> File:
source_text = None
if isinstance(file_input, File):
return file_input
if isinstance(file_input, SourceText):
source_text = file_input
elif isinstance(file_input, (str, Path)):
path = Path(file_input)
source_text = SourceText(file=path, filename=path.name)
else:
# Try to get filename from the file object if not provided
filename = None
if hasattr(file_input, "name"):
filename = os.path.basename(str(file_input.name))
if filename is None:
raise ValueError(
"Use SourceText to provide filename when uploading bytes or file-like objects."
)
warnings.warn(
"Use SourceText instead of bytes or file-like objects",
DeprecationWarning,
)
source_text = SourceText(file=file_input, filename=filename)
return await self.upload_file(source_text)
"""Upload a file from various input types using FileClient."""
return await self._file_client.upload_content(file_input)
async def _wait_for_job_result(self, job_id: str) -> Optional[ExtractRun]:
"""Wait for and return the results of an extraction job."""
+82
View File
@@ -1,9 +1,11 @@
from io import BytesIO
from typing import BinaryIO
import os
from pathlib import Path
from llama_cloud.client import AsyncLlamaCloud
from llama_cloud.types import File, FileCreate
from typing import Optional
from llama_cloud_services.utils import SourceText, FileInput
class FileClient:
@@ -95,3 +97,83 @@ class FileClient:
project_id=self.project_id,
organization_id=self.organization_id,
)
async def upload_content(
self, file_input: FileInput, external_file_id: Optional[str] = None
) -> File:
"""
Upload content from various input types or fetch an already-uploaded file.
Args:
file_input: The content to upload. Can be:
- File: Already uploaded file (returned as-is)
- str/Path: Path to a file on disk
- SourceText: Text content, file, or file_id with explicit filename
- BufferedIOBase: File-like binary object
external_file_id: Optional external identifier for the file
Returns:
File: The uploaded (or fetched) file object
Raises:
ValueError: If the input type is not supported or required info is missing
"""
# If already a File object, return it
if isinstance(file_input, File):
return file_input
# Handle SourceText
if isinstance(file_input, SourceText):
# If file_id is provided, fetch the file object
if file_input.file_id is not None:
return await self.get_file(file_input.file_id)
elif file_input.text_content is not None:
# Handle direct text content
text_bytes = file_input.text_content.encode("utf-8")
return await self.upload_bytes(
text_bytes, external_file_id or file_input.filename or "file"
)
elif isinstance(file_input.file, (str, Path)):
# Handle file paths using the existing upload_file method
return await self.upload_file(
str(file_input.file), external_file_id or file_input.filename
)
elif isinstance(file_input.file, bytes):
# Handle bytes
return await self.upload_bytes(
file_input.file, external_file_id or file_input.filename or "file"
)
elif hasattr(file_input.file, "read"):
# Handle any file-like object (TextIOWrapper, BytesIO, BufferedReader, BufferedIOBase, etc.)
content = file_input.file.read() # type: ignore
if isinstance(content, str):
content = content.encode("utf-8")
return await self.upload_bytes(
content, external_file_id or file_input.filename or "file"
)
else:
raise ValueError(f"Unsupported file type: {type(file_input.file)}")
# Handle string/Path directly
elif isinstance(file_input, (str, Path)):
return await self.upload_file(str(file_input), external_file_id)
# Handle raw file-like objects
elif hasattr(file_input, "read"):
if hasattr(file_input, "name"):
filename = os.path.basename(str(file_input.name))
else:
filename = external_file_id or "file"
# Read content to determine size
content = file_input.read()
if isinstance(content, str):
content = content.encode("utf-8")
return await self.upload_bytes(content, external_file_id or filename)
else:
raise ValueError(
f"Unsupported file input type: {type(file_input)}. "
f"Supported types: str, Path, SourceText, BufferedIOBase, or File."
)
+18 -2
View File
@@ -258,6 +258,7 @@ def page_screenshot_nodes_to_node_with_score(
client: LlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
if not raw_image_nodes:
return []
@@ -273,6 +274,7 @@ def page_screenshot_nodes_to_node_with_score(
image_base64 = base64.b64encode(image_bytes).decode("utf-8")
image_node_metadata: Dict[str, Any] = {
**(raw_image_node.node.metadata or {}),
**(metadata or {}),
"file_id": raw_image_node.node.file_id,
"page_index": raw_image_node.node.page_index,
}
@@ -289,6 +291,7 @@ def image_nodes_to_node_with_score(
client: LlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
"""
Legacy method to alias page_screenshot_nodes_to_node_with_score.
@@ -297,7 +300,10 @@ def image_nodes_to_node_with_score(
return []
return page_screenshot_nodes_to_node_with_score(
client=client, raw_image_nodes=raw_image_nodes, project_id=project_id
client=client,
raw_image_nodes=raw_image_nodes,
project_id=project_id,
metadata=metadata,
)
@@ -305,6 +311,7 @@ def page_figure_nodes_to_node_with_score(
client: LlamaCloud,
raw_figure_nodes: Optional[List[PageFigureNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
if not raw_figure_nodes:
return []
@@ -321,6 +328,7 @@ def page_figure_nodes_to_node_with_score(
figure_base64 = base64.b64encode(figure_bytes).decode("utf-8")
figure_node_metadata: Dict[str, Any] = {
**(raw_figure_node.node.metadata or {}),
**(metadata or {}),
"file_id": raw_figure_node.node.file_id,
"page_index": raw_figure_node.node.page_index,
"figure_name": raw_figure_node.node.figure_name,
@@ -337,6 +345,7 @@ async def apage_screenshot_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
if not raw_image_nodes:
return []
@@ -357,6 +366,7 @@ async def apage_screenshot_nodes_to_node_with_score(
image_base64 = base64.b64encode(image_bytes).decode("utf-8")
image_node_metadata: Dict[str, Any] = {
**(raw_image_node.node.metadata or {}),
**(metadata or {}),
"file_id": raw_image_node.node.file_id,
"page_index": raw_image_node.node.page_index,
}
@@ -372,6 +382,7 @@ async def aimage_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_image_nodes: Optional[List[PageScreenshotNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
"""
Legacy method to alias apage_screenshot_nodes_to_node_with_score.
@@ -380,7 +391,10 @@ async def aimage_nodes_to_node_with_score(
return []
return await apage_screenshot_nodes_to_node_with_score(
client=client, raw_image_nodes=raw_image_nodes, project_id=project_id
client=client,
raw_image_nodes=raw_image_nodes,
project_id=project_id,
metadata=metadata,
)
@@ -388,6 +402,7 @@ async def apage_figure_nodes_to_node_with_score(
client: AsyncLlamaCloud,
raw_figure_nodes: Optional[List[PageFigureNodeWithScore]],
project_id: str,
metadata: Optional[dict] = None,
) -> List[NodeWithScore]:
if not raw_figure_nodes:
return []
@@ -409,6 +424,7 @@ async def apage_figure_nodes_to_node_with_score(
figure_base64 = base64.b64encode(figure_bytes).decode("utf-8")
figure_node_metadata: Dict[str, Any] = {
**(raw_figure_node.node.metadata or {}),
**(metadata or {}),
"file_id": raw_figure_node.node.file_id,
"page_index": raw_figure_node.node.page_index,
"figure_name": raw_figure_node.node.figure_name,
+31 -10
View File
@@ -19,6 +19,7 @@ from llama_cloud import (
PipelineCreate,
PipelineCreateEmbeddingConfig,
PipelineCreateTransformConfig,
PipelineFileCreateCustomMetadataValue,
PipelineType,
ProjectCreate,
ManagedIngestionStatus,
@@ -333,7 +334,7 @@ class LlamaCloudIndex(BaseManagedIndex):
if file_ids:
self._wait_for_resources(
file_ids,
lambda fid: self._client.pipelines.get_pipeline_file_status(
lambda fid: self._client.pipeline_files.get_pipeline_file_status(
pipeline_id=self.pipeline.id, file_id=fid
),
resource_name="file",
@@ -420,7 +421,7 @@ class LlamaCloudIndex(BaseManagedIndex):
if file_ids:
await self._await_for_resources(
file_ids,
lambda fid: self._aclient.pipelines.get_pipeline_file_status(
lambda fid: self._aclient.pipeline_files.get_pipeline_file_status(
pipeline_id=self.pipeline.id, file_id=fid
),
resource_name="file",
@@ -905,6 +906,9 @@ class LlamaCloudIndex(BaseManagedIndex):
def upload_file(
self,
file_path: str,
custom_metadata: Optional[
dict[str, Optional[PipelineFileCreateCustomMetadataValue]]
] = None,
verbose: bool = False,
wait_for_ingestion: bool = True,
raise_on_error: bool = False,
@@ -918,8 +922,10 @@ class LlamaCloudIndex(BaseManagedIndex):
print(f"Uploaded file {file.id} with name {file.name}")
# Add file to pipeline
pipeline_file_create = PipelineFileCreate(file_id=file.id)
self._client.pipelines.add_files_to_pipeline_api(
pipeline_file_create = PipelineFileCreate(
file_id=file.id, custom_metadata=custom_metadata
)
self._client.pipeline_files.add_files_to_pipeline_api(
pipeline_id=self.pipeline.id, request=[pipeline_file_create]
)
@@ -932,6 +938,9 @@ class LlamaCloudIndex(BaseManagedIndex):
async def aupload_file(
self,
file_path: str,
custom_metadata: Optional[
dict[str, Optional[PipelineFileCreateCustomMetadataValue]]
] = None,
verbose: bool = False,
wait_for_ingestion: bool = True,
raise_on_error: bool = False,
@@ -945,8 +954,10 @@ class LlamaCloudIndex(BaseManagedIndex):
print(f"Uploaded file {file.id} with name {file.name}")
# Add file to pipeline
pipeline_file_create = PipelineFileCreate(file_id=file.id)
await self._aclient.pipelines.add_files_to_pipeline_api(
pipeline_file_create = PipelineFileCreate(
file_id=file.id, custom_metadata=custom_metadata
)
await self._aclient.pipeline_files.add_files_to_pipeline_api(
pipeline_id=self.pipeline.id, request=[pipeline_file_create]
)
@@ -961,6 +972,9 @@ class LlamaCloudIndex(BaseManagedIndex):
self,
file_name: str,
url: str,
custom_metadata: Optional[
dict[str, Optional[PipelineFileCreateCustomMetadataValue]]
] = None,
proxy_url: Optional[str] = None,
request_headers: Optional[Dict[str, str]] = None,
verify_ssl: bool = True,
@@ -983,8 +997,10 @@ class LlamaCloudIndex(BaseManagedIndex):
print(f"Uploaded file {file.id} with ID {file.id}")
# Add file to pipeline
pipeline_file_create = PipelineFileCreate(file_id=file.id)
self._client.pipelines.add_files_to_pipeline_api(
pipeline_file_create = PipelineFileCreate(
file_id=file.id, custom_metadata=custom_metadata
)
self._client.pipeline_files.add_files_to_pipeline_api(
pipeline_id=self.pipeline.id, request=[pipeline_file_create]
)
@@ -998,6 +1014,9 @@ class LlamaCloudIndex(BaseManagedIndex):
self,
file_name: str,
url: str,
custom_metadata: Optional[
dict[str, Optional[PipelineFileCreateCustomMetadataValue]]
] = None,
proxy_url: Optional[str] = None,
request_headers: Optional[Dict[str, str]] = None,
verify_ssl: bool = True,
@@ -1020,8 +1039,10 @@ class LlamaCloudIndex(BaseManagedIndex):
print(f"Uploaded file {file.id} with ID {file.id}")
# Add file to pipeline
pipeline_file_create = PipelineFileCreate(file_id=file.id)
await self._aclient.pipelines.add_files_to_pipeline_api(
pipeline_file_create = PipelineFileCreate(
file_id=file.id, custom_metadata=custom_metadata
)
await self._aclient.pipeline_files.add_files_to_pipeline_api(
pipeline_id=self.pipeline.id, request=[pipeline_file_create]
)
+25 -8
View File
@@ -129,11 +129,12 @@ class LlamaCloudRetriever(BaseRetriever):
)
def _result_nodes_to_node_with_score(
self, result_nodes: List[TextNodeWithScore]
self, result_nodes: List[TextNodeWithScore], metadata: Optional[dict] = None
) -> List[NodeWithScore]:
nodes = []
for res in result_nodes:
text_node = TextNode.parse_obj(res.node.dict())
text_node = TextNode.model_validate(res.node.dict())
text_node.metadata.update(metadata or {})
nodes.append(NodeWithScore(node=text_node, score=res.score))
return nodes
@@ -161,17 +162,25 @@ class LlamaCloudRetriever(BaseRetriever):
search_filters_inference_schema=search_filters_inference_schema,
)
result_nodes = self._result_nodes_to_node_with_score(results.retrieval_nodes)
result_nodes = self._result_nodes_to_node_with_score(
results.retrieval_nodes, metadata=results.metadata
)
if self._retrieve_page_screenshot_nodes:
result_nodes.extend(
page_screenshot_nodes_to_node_with_score(
self._client, results.image_nodes, self.project.id
self._client,
results.image_nodes,
self.project.id,
metadata=results.metadata,
)
)
if self._retrieve_page_figure_nodes:
result_nodes.extend(
page_figure_nodes_to_node_with_score(
self._client, results.page_figure_nodes, self.project.id
self._client,
results.page_figure_nodes,
self.project.id,
metadata=results.metadata,
)
)
@@ -200,17 +209,25 @@ class LlamaCloudRetriever(BaseRetriever):
search_filters_inference_schema=search_filters_inference_schema,
)
result_nodes = self._result_nodes_to_node_with_score(results.retrieval_nodes)
result_nodes = self._result_nodes_to_node_with_score(
results.retrieval_nodes, metadata=results.metadata
)
if self._retrieve_page_screenshot_nodes:
result_nodes.extend(
await apage_screenshot_nodes_to_node_with_score(
self._aclient, results.image_nodes, self.project.id
self._aclient,
results.image_nodes,
self.project.id,
metadata=results.metadata,
)
)
if self._retrieve_page_figure_nodes:
result_nodes.extend(
await apage_figure_nodes_to_node_with_score(
self._aclient, results.page_figure_nodes, self.project.id
self._aclient,
results.page_figure_nodes,
self.project.id,
metadata=results.metadata,
)
)
+40 -3
View File
@@ -188,6 +188,10 @@ class LlamaParse(BasePydanticReader):
default=False,
description="If set to true, LlamaParse will try to detect long table and adapt the output.",
)
aggressive_table_extraction: Optional[bool] = Field(
default=False,
description="If set to true, LlamaParse will try to extract tables aggressively, may lead to false positives.",
)
annotate_links: Optional[bool] = Field(
default=False,
description="Annotate links found in the document to extract their URL.",
@@ -713,6 +717,9 @@ class LlamaParse(BasePydanticReader):
if self.adaptive_long_table:
data["adaptive_long_table"] = self.adaptive_long_table
if self.aggressive_table_extraction:
data["aggressive_table_extraction"] = self.aggressive_table_extraction
if self.annotate_links:
data["annotate_links"] = self.annotate_links
@@ -1139,6 +1146,25 @@ class LlamaParse(BasePydanticReader):
)
current_interval = self._calculate_backoff(current_interval)
async def _get_job_result_with_error_handling(
self, job_id: str, result_type: str, verbose: bool = False
) -> Dict[str, Any]:
"""Get job result with error handling based on ignore_errors setting."""
try:
return await self._get_job_result(job_id, result_type, verbose=verbose)
except JobFailedException as e:
if self.ignore_errors:
# Return error information when ignore_errors is True
return {
"pages": [],
"job_metadata": {},
"error": f"{e.status}: {e.error_message or 'No error message'}",
"error_code": e.error_code,
"status": e.status,
}
else:
raise e
async def _parse_one(
self,
file_path: FileInput,
@@ -1180,7 +1206,7 @@ class LlamaParse(BasePydanticReader):
)
if self.verbose:
print("Started parsing the file under job_id %s" % job_id)
result = await self._get_job_result(
result = await self._get_job_result_with_error_handling(
job_id, result_type or self.result_type.value, verbose=self.verbose
)
return job_id, result
@@ -1243,6 +1269,15 @@ class LlamaParse(BasePydanticReader):
result_type=ResultType.JSON.value,
partition_target_pages=f"{total}-{total + size - 1}",
)
# Check if the result is an error result (when ignore_errors=True)
if json_result.get("error_code") == "NO_DATA_FOUND_IN_FILE":
raise JobFailedException(
job_id=job_id,
status=json_result.get("status", "ERROR"),
error_code=json_result.get("error_code"),
error_message=json_result.get("error"),
)
result_type = result_type or self.result_type.value
if result_type == ResultType.JSON.value:
job_result = json_result
@@ -1768,7 +1803,7 @@ class LlamaParse(BasePydanticReader):
JobResult object or list of JobResult objects if multiple job IDs were provided.
"""
if isinstance(job_id, str):
result = await self._get_job_result(
result = await self._get_job_result_with_error_handling(
job_id, ResultType.JSON.value, verbose=self.verbose
)
return JobResult(
@@ -1783,7 +1818,9 @@ class LlamaParse(BasePydanticReader):
elif isinstance(job_id, list):
results = []
jobs = [
self._get_job_result(id_, ResultType.JSON.value, verbose=self.verbose)
self._get_job_result_with_error_handling(
id_, ResultType.JSON.value, verbose=self.verbose
)
for id_ in job_id
]
results = await run_jobs(
+153 -27
View File
@@ -1,17 +1,87 @@
import httpx
import os
import re
from pydantic import BaseModel, Field, SerializeAsAny
from typing import Dict, Any, List, Optional
from pydantic import BaseModel, ConfigDict, Field, SerializeAsAny, model_validator
from typing import Dict, Any, List, Optional, get_origin, get_args
from llama_cloud_services.parse.utils import make_api_request
from llama_cloud_services.parse.utils import (
make_api_request,
is_jupyter,
)
from llama_index.core.async_utils import asyncio_run
from llama_index.core.schema import Document, ImageDocument, ImageNode, TextNode
PAGE_REGEX = r"page[-_](\d+)\.jpg$"
SAFE_MODEL_CONFIGS = ConfigDict(
extra="allow",
validate_assignment=False,
arbitrary_types_allowed=True,
validate_default=False,
)
class JobMetadata(BaseModel):
class SafeBaseModel(BaseModel):
"""Base model that gracefully handles None values from unstable backend responses."""
model_config = SAFE_MODEL_CONFIGS
@model_validator(mode="before")
@classmethod
def coerce_none_to_defaults(cls, data: Any) -> Any:
"""
Replace None values with appropriate defaults based on field type annotations.
This prevents validation errors when the backend returns None for non-optional fields.
"""
if not isinstance(data, dict):
return data
# Process each field that has a None value
result = {}
for key, value in data.items():
if value is not None or key not in cls.model_fields:
result[key] = value
continue
# Value is None and field exists in model
field_info = cls.model_fields[key]
# If field has a default or default_factory, let Pydantic handle it
from pydantic_core import PydanticUndefined
if (
field_info.default is not PydanticUndefined
or field_info.default_factory is not None
):
continue
# Otherwise, provide a sensible default based on the type annotation
annotation = field_info.annotation
origin = get_origin(annotation)
# Handle List types
if origin is list:
result[key] = []
# Handle Dict types
elif origin is dict:
result[key] = {}
# Handle basic types
elif annotation == str or (origin and str in get_args(annotation)):
result[key] = ""
elif annotation == int or (origin and int in get_args(annotation)):
result[key] = 0
elif annotation == float or (origin and float in get_args(annotation)):
result[key] = 0.0
elif annotation == bool or (origin and bool in get_args(annotation)):
result[key] = False
# If we can't determine a safe default, skip (let Pydantic try)
else:
result[key] = value
return result
class JobMetadata(SafeBaseModel):
"""Metadata about the job."""
job_pages: int = Field(default=0, description="The number of pages in the job.")
@@ -24,19 +94,31 @@ class JobMetadata(BaseModel):
)
class BBox(BaseModel):
class BBox(SafeBaseModel):
"""A bounding box."""
x: float = Field(description="The x-coordinate of the bounding box.")
y: float = Field(description="The y-coordinate of the bounding box.")
w: float = Field(description="The width of the bounding box.")
h: float = Field(description="The height of the bounding box.")
x: Optional[float] = Field(
default=None,
description="The x-coordinate of the bounding box.",
)
y: Optional[float] = Field(
default=None,
description="The y-coordinate of the bounding box.",
)
w: Optional[float] = Field(
default=None,
description="The width of the bounding box.",
)
h: Optional[float] = Field(
default=None,
description="The height of the bounding box.",
)
class PageItem(BaseModel):
class PageItem(SafeBaseModel):
"""An item in a page."""
type: str = Field(description="The type of the item.")
type: str = Field(default="", description="The type of the item.")
lvl: Optional[int] = Field(
default=None, description="The level of indentation of the item."
)
@@ -58,10 +140,10 @@ class PageItem(BaseModel):
)
class ImageItem(BaseModel):
class ImageItem(SafeBaseModel):
"""An image in a page."""
name: str = Field(description="The name of the image.")
name: str = Field(default="", description="The name of the image.")
height: Optional[float] = Field(
default=None, description="The height of the image."
)
@@ -81,22 +163,28 @@ class ImageItem(BaseModel):
type: Optional[str] = Field(default=None, description="The type of the image.")
class LayoutItem(BaseModel):
class LayoutItem(SafeBaseModel):
"""The layout of a page."""
image: str = Field(description="The name of the image containing the layout item")
confidence: float = Field(description="The confidence of the layout item.")
label: str = Field(description="The label of the layout item.")
image: str = Field(
default="", description="The name of the image containing the layout item"
)
confidence: float = Field(
default=0.0, description="The confidence of the layout item."
)
label: str = Field(default="", description="The label of the layout item.")
bbox: Optional[BBox] = Field(
default=None, description="The bounding box of the layout item."
)
isLikelyNoise: bool = Field(description="Whether the layout item is likely noise.")
isLikelyNoise: bool = Field(
default=False, description="Whether the layout item is likely noise."
)
class ChartItem(BaseModel):
class ChartItem(SafeBaseModel):
"""A chart in a page."""
name: str = Field(description="The name of the chart.")
name: str = Field(default="", description="The name of the chart.")
x: Optional[float] = Field(
default=None, description="The x-coordinate of the chart."
)
@@ -109,7 +197,7 @@ class ChartItem(BaseModel):
)
class Page(BaseModel):
class Page(SafeBaseModel):
"""A page of the document."""
page: int = Field(default=0, description="The page number.")
@@ -164,7 +252,7 @@ class Page(BaseModel):
)
class JobResult(BaseModel):
class JobResult(SafeBaseModel):
"""The raw JSON result from the LlamaParse API."""
pages: List[Page] = Field(
@@ -181,6 +269,13 @@ class JobResult(BaseModel):
error: Optional[str] = Field(
default=None, description="The error message if the job failed."
)
error_code: Optional[str] = Field(
default=None, description="The error code if the job failed."
)
status: Optional[str] = Field(
default=None,
description="The job status (e.g., PENDING, SUCCESS, ERROR, CANCELED).",
)
def __init__(
self,
@@ -258,6 +353,29 @@ class JobResult(BaseModel):
documents = await self.aget_text_documents(split_by_page)
return [TextNode(text=doc.text, metadata=doc.metadata) for doc in documents]
def _format_markdown_for_notebook(self, text: Optional[str]) -> Optional[str]:
"""Format markdown text for Jupyter notebook display by escaping dollar signs."""
if text is None:
return None
def escape_single_dollar_signs(text: str) -> str:
"""Escape single dollar signs in text to prevent Jupyter from interpreting them as LaTeX.
Preserves all strings of dollar signs greater than length 1,
especially preserving double dollar signs ($$) which denote LaTeX equations.
Args:
text: The text to escape
Returns:
Text with single dollar signs escaped
"""
# Replace single $ with \$, but preserve $$
# Use negative lookahead and lookbehind to match $ not preceded or followed by $
return re.sub(r"(?<!\$)\$(?!\$)", r"\$", text)
return escape_single_dollar_signs(text)
def get_markdown_documents(self, split_by_page: bool = False) -> List[Document]:
"""
Get the markdown documents from the job.
@@ -268,17 +386,22 @@ class JobResult(BaseModel):
if split_by_page:
return [
Document(
text=page.md,
text=self._format_markdown_for_notebook(page.md)
if is_jupyter()
else page.md,
metadata={"page_number": page.page, "file_name": self.file_name},
)
for page in self.pages
]
else:
text = self._page_separator.join(
[page.md if page.md is not None else "" for page in self.pages]
)
return [
Document(
text=self._page_separator.join(
[page.md if page.md is not None else "" for page in self.pages]
),
text=self._format_markdown_for_notebook(text)
if is_jupyter()
else text,
metadata={"file_name": self.file_name},
)
]
@@ -328,7 +451,10 @@ class JobResult(BaseModel):
"""
url = f"{self._base_url}/api/v1/parsing/job/{self.job_id}/result/raw/markdown"
response = await make_api_request(self._client, "GET", url)
return response.content.decode("utf-8")
markdown = response.content.decode("utf-8")
return (
self._format_markdown_for_notebook(markdown) if is_jupyter() else markdown
)
def get_text(self) -> str:
"""
+12
View File
@@ -1,3 +1,4 @@
import functools
import httpx
import itertools
import logging
@@ -356,6 +357,17 @@ def partition_pages(
return
@functools.lru_cache(maxsize=1)
def is_jupyter() -> bool:
"""Check if we're running in a Jupyter environment."""
try:
from IPython import get_ipython
return get_ipython().__class__.__name__ == "ZMQInteractiveShell"
except (ImportError, AttributeError):
return False
def extract_tables_from_json_results(
json_results: List[dict], download_path: str
) -> List[str]:
@@ -0,0 +1,9 @@
from .matchers import FileMatcher, RequestMatcher, SchemaMatcher
from .server import FakeLlamaCloudServer
__all__ = [
"FakeLlamaCloudServer",
"FileMatcher",
"SchemaMatcher",
"RequestMatcher",
]
@@ -0,0 +1,192 @@
from __future__ import annotations
import hashlib
import json
import random
from datetime import datetime, timezone
from typing import Any, Iterable, Mapping, MutableMapping
def hash_chunks(chunks: Iterable[bytes]) -> str:
digest = hashlib.sha256()
for chunk in chunks:
digest.update(chunk)
return digest.hexdigest()
def fingerprint_file(content: bytes, filename: str | None = None) -> str:
name_bytes = filename.encode("utf-8") if filename else b""
return hash_chunks((content, name_bytes))
def hash_schema(schema: Any) -> str:
json_string = json.dumps(
_to_serializable(schema),
sort_keys=True,
separators=(",", ":"),
)
return hashlib.sha256(json_string.encode("utf-8")).hexdigest()
def combined_seed(*parts: str) -> int:
digest = hash_chunks(tuple(part.encode("utf-8") for part in parts))
return int(digest[:16], 16)
def generate_data_from_schema(schema: Any, seed: int) -> Any:
rng = random.Random(seed)
return _generate_value(schema, rng, depth=0)
def generate_text_blob(seed: int, sentences: int = 3) -> str:
rng = random.Random(seed)
words = [
"aurora",
"copper",
"delta",
"ember",
"fable",
"glyph",
"harbor",
"iris",
"juniper",
"kepler",
"lumen",
"monarch",
"nylon",
"onyx",
"paragon",
"quartz",
"raptor",
"solstice",
"topaz",
"umbra",
"verdant",
"willow",
"xenon",
"yonder",
"zephyr",
]
sentence_pieces = []
for _ in range(sentences):
length = rng.randint(6, 12)
chosen = rng.sample(words, k=length)
sentence = " ".join(chosen).capitalize() + "."
sentence_pieces.append(sentence)
return " ".join(sentence_pieces)
def utcnow() -> datetime:
return datetime.now(timezone.utc)
def _to_serializable(value: Any) -> Any:
if value is None:
return None
if isinstance(value, (str, int, float, bool)):
return value
if isinstance(value, bytes):
return value.decode("utf-8", errors="ignore")
if isinstance(value, Mapping):
return {key: _to_serializable(val) for key, val in value.items()}
if isinstance(value, MutableMapping):
return {key: _to_serializable(val) for key, val in value.items()}
if isinstance(value, (list, tuple, set)):
return [_to_serializable(item) for item in value]
if hasattr(value, "model_dump_json"):
return json.loads(value.model_dump_json())
if hasattr(value, "model_dump"):
return value.model_dump()
if hasattr(value, "dict"):
return value.dict() # type: ignore[call-arg]
if hasattr(value, "model_json_schema"):
return value.model_json_schema()
return str(value)
def _generate_value(schema: Any, rng: random.Random, depth: int) -> Any:
if depth > 8:
return rng.choice(
(
rng.randint(1, 999),
rng.random(),
generate_text_blob(rng.randint(0, 1_000_000), sentences=1),
)
)
if schema is None:
return generate_text_blob(rng.randint(0, 1_000_000), sentences=1)
if isinstance(schema, list):
return [_generate_value(item, rng, depth + 1) for item in schema]
if isinstance(schema, str):
return f"{schema}-{rng.randint(100, 999)}"
if isinstance(schema, Mapping):
if "enum" in schema:
options = schema["enum"]
if options:
index = rng.randint(0, len(options) - 1)
return options[index]
schema_type = schema.get("type")
if schema_type == "object":
properties = schema.get("properties", {})
result = {}
for key, subschema in properties.items():
result[key] = _generate_value(subschema, rng, depth + 1)
return result
if schema_type == "array":
items_schema = schema.get("items", {})
min_items = schema.get("minItems", 1)
max_items = schema.get("maxItems", max(3, min_items))
length = rng.randint(min_items, min(min_items + 2, max_items))
return [
_generate_value(items_schema, rng, depth + 1) for _ in range(length)
]
if schema_type == "integer":
minimum = schema.get("minimum", 0)
maximum = schema.get("maximum", minimum + 500)
return rng.randint(int(minimum), int(maximum))
if schema_type == "number":
minimum = schema.get("minimum", 0.0)
maximum = schema.get("maximum", minimum + 500.0)
value = rng.uniform(float(minimum), float(maximum))
return round(value, 2)
if schema_type == "boolean":
return rng.choice((True, False))
if schema_type == "string":
fmt = schema.get("format")
if fmt == "date-time":
timestamp = utcnow().isoformat()
return timestamp
if fmt == "email":
return f"user{rng.randint(1000, 9999)}@example.com"
if fmt == "uri":
return f"https://example.com/{rng.randint(1000, 9999)}"
min_length = schema.get("minLength", 5)
max_length = schema.get("maxLength", max(10, min_length))
length = rng.randint(min_length, min(min_length + 5, max_length))
return generate_text_blob(
rng.randint(0, 1_000_000), sentences=max(1, length // 5)
)
if schema_type == "null":
return None
if "oneOf" in schema:
option = rng.choice(schema["oneOf"])
return _generate_value(option, rng, depth + 1)
if "anyOf" in schema:
option = rng.choice(schema["anyOf"])
return _generate_value(option, rng, depth + 1)
return generate_text_blob(rng.randint(0, 1_000_000), sentences=1)
@@ -0,0 +1,335 @@
from __future__ import annotations
import httpx
import re
from dataclasses import dataclass
from typing import TYPE_CHECKING, Any, Dict
from ._deterministic import utcnow, hash_schema
if TYPE_CHECKING:
from .server import FakeLlamaCloudServer
@dataclass
class StoredAgentData:
data: dict[str, Any]
id: str
collection: str
deployment_name: str
def __getattr__(self, name: str) -> Any:
return self.data.get(name)
def __setattr__(self, name: str, value: Any) -> None:
if name in ("data", "id", "collection", "deployment_name"):
super().__setattr__(name, value)
else:
self.data[name] = value
@classmethod
def from_request_data(cls, data: dict[str, Any]) -> "StoredAgentData":
return cls(
data=data.get("data", {}),
collection=data.get("collection", "default"),
deployment_name=data.get("deployment_name", ""),
id=hash_schema(data.get("data", {}))[:7],
)
def apply_filter(data: dict, filters: dict) -> bool:
"""Check if data matches all filters"""
ops = {
"gt": lambda a, b: a > b,
"gte": lambda a, b: a >= b,
"lt": lambda a, b: a < b,
"lte": lambda a, b: a <= b,
"eq": lambda a, b: a == b,
"ne": lambda a, b: a != b,
"in": lambda a, b: a in b,
"nin": lambda a, b: a not in b,
}
for key, condition in filters.items():
if key not in data:
return False
if isinstance(condition, dict):
for op, value in condition.items():
if op in ops:
if not ops[op](data[key], value):
return False
else:
return False
else:
if data[key] != condition:
return False
return True
class FakeAgentDataNamespace:
def __init__(
self,
*,
server: "FakeLlamaCloudServer",
) -> None:
self._server = server
self.stored: list[StoredAgentData] = []
self.routes: Dict[str, Any] = {}
def _create_data(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request=request)
data = StoredAgentData.from_request_data(payload)
self.stored.append(data)
response = {
"data": data.data,
"collection": data.collection,
"deployment_name": data.deployment_name,
"created_at": utcnow().isoformat(),
"updated_at": None,
"id": data.id,
"project_id": None,
"organization_id": None,
}
return self._server.json_response(response, status_code=200)
def _delete_data_by_query(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request=request)
delete_count = 0
if (filters := payload.get("filter")) is not None:
to_keep = []
for data in self.stored:
if data.collection == payload.get(
"collection", "default"
) and data.deployment_name == payload.get("deployment_name"):
if not apply_filter(data.data, filters):
to_keep.append(data)
else:
delete_count += 1
self.stored = to_keep
return self._server.json_response(
{"deleted_count": delete_count}, status_code=200
)
def _delete_data_by_id(self, request: httpx.Request) -> httpx.Response:
item_id = self._find_item_id(request=request)
if not item_id:
return self._server.json_response(
{
"detail": "An item_id path parameter is required to perform this operation"
},
status_code=400,
)
if not any(data.id == item_id for data in self.stored):
return self._server.json_response(
{"detail": f"No data with ID: {item_id}"}, status_code=404
)
self.stored = [data for data in self.stored if data.id != item_id]
return self._server.json_response({}, status_code=200)
def _get_data_by_id(self, request: httpx.Request) -> httpx.Response:
item_id = self._find_item_id(request=request)
if not item_id:
return self._server.json_response(
{
"detail": "An item_id path parameter is required to perform this operation"
},
status_code=400,
)
data = [data for data in self.stored if data.id == item_id]
if data:
response = {
"data": data[0].data,
"collection": data[0].collection,
"deployment_name": data[0].deployment_name,
"created_at": utcnow().isoformat(),
"updated_at": None,
"id": data[0].id,
"project_id": None,
"organization_id": None,
}
return self._server.json_response(response, status_code=200)
else:
return self._server.json_response(
{"detail": f"No data with ID: {item_id}"}, status_code=404
)
def _search_data(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request=request)
found = []
if (filters := payload.get("filter")) is not None:
for data in self.stored:
if data.collection == payload.get(
"collection", "default"
) and data.deployment_name == payload.get("deployment_name"):
if apply_filter(data.data, filters):
found.append(
{
"data": data.data,
"collection": data.collection,
"deployment_name": data.deployment_name,
"created_at": utcnow().isoformat(),
"updated_at": None,
"id": data.id,
"project_id": None,
"organization_id": None,
}
)
else:
for data in self.stored:
if data.collection == payload.get(
"collection", "default"
) and data.deployment_name == payload.get("deployment_name"):
found.append(
{
"data": data.data,
"collection": data.collection,
"deployment_name": data.deployment_name,
"created_at": utcnow().isoformat(),
"updated_at": None,
"id": data.id,
"project_id": None,
"organization_id": None,
}
)
return self._server.json_response(
{"items": found, "next_page_token": None, "total_size": len(found)},
status_code=200,
)
def _update_data(self, request: httpx.Request) -> httpx.Response:
item_id = self._find_item_id(request=request)
payload = self._server.json(request=request)
if not item_id:
return self._server.json_response(
{
"detail": "An item_id path parameter is required to perform this operation"
},
status_code=400,
)
updated = None
for i, data in enumerate(self.stored):
if data.id == item_id:
updated = data
updated.data = payload.get("data", data.data)
self.stored[i] = updated
print(updated)
if updated is not None:
response = {
"data": updated.data,
"collection": updated.collection,
"deployment_name": updated.deployment_name,
"created_at": None,
"updated_at": utcnow().isoformat(),
"id": updated.id,
"project_id": None,
"organization_id": None,
}
status_code = 200
else:
response = {"detail": f"Record with id {item_id} not found"}
status_code = 404
return self._server.json_response(response, status_code=status_code)
def _aggregate_data(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request=request)
add_count = payload.get("count", False)
group_bys: list[str] = payload.get("group_by", [])
groups: dict[str, dict[str, list[dict]]] = {key: {} for key in group_bys}
if (filters := payload.get("filter")) is not None:
for data in self.stored:
if data.collection == payload.get(
"collection", "default"
) and data.deployment_name == payload.get("deployment_name"):
if apply_filter(data.data, filters):
for key in group_bys:
if key in data.data and data.data[key] in groups[key]:
groups[key][data.data[key]].append(data.data)
elif key in data.data and data.data[key] not in groups[key]:
groups[key][data.data[key]] = [data.data]
else:
for data in self.stored:
if data.collection == payload.get(
"collection", "default"
) and data.deployment_name == payload.get("deployment_name"):
for key in group_bys:
if key in data.data and data.data[key] in groups[key]:
groups[key][data.data[key]].append(data.data)
elif key in data.data and data.data[key] not in groups[key]:
groups[key][data.data[key]] = [data.data]
response: dict[str, Any] = {
"items": [],
"next_page_token": None,
"total_size": 0,
}
for k in groups:
if len(groups[k]) > 0:
for v in groups[k]:
if groups[k][v]:
first_element = groups[k][v][0]
else:
first_element = None
response["items"].append(
{
"first_item": first_element,
"count": len(groups[k][v]) if add_count else None,
"group_key": {k: v},
}
)
response["total_size"] = len(response["items"])
return self._server.json_response(response, status_code=200)
def _find_item_id(self, request: httpx.Request) -> str | None:
matchgroups = re.search(r"/agent-data\/([^\/]+)$", request.url.path)
return matchgroups.group(1) if matchgroups is not None else None
def register(self) -> None:
server = self._server
route = server.add_route(
"POST",
"/api/v1/beta/agent-data",
self._create_data,
namespace="create_item",
)
self.routes["stateless_run"] = route
self.stateless_run = route
server.add_route(
"POST",
"/api/v1/beta/agent-data/:aggregate",
self._aggregate_data,
namespace="untyped_aggregate",
alias="aggregate",
)
server.add_route(
"POST",
"/api/v1/beta/agent-data/:delete",
self._delete_data_by_query,
namespace="delete",
)
server.add_route(
"POST",
"/api/v1/beta/agent-data/:search",
self._search_data,
namespace="untyped_search",
alias="search",
)
server.add_route(
"DELETE",
"/api/v1/beta/agent-data/{item_id}",
self._delete_data_by_id,
namespace="delete_item",
)
server.add_route(
"GET",
"/api/v1/beta/agent-data/{item_id}",
self._get_data_by_id,
namespace="untyped_get_item",
alias="get_item",
)
server.add_route(
"PUT",
"/api/v1/beta/agent-data/{item_id}",
self._update_data,
namespace="update_item",
)
@@ -0,0 +1,153 @@
from __future__ import annotations
from dataclasses import dataclass
from typing import TYPE_CHECKING, Dict, List
import httpx
from llama_cloud.types import (
ClassifierRule,
ClassifyJob,
ClassifyJobResults,
ClassificationResult,
FileClassification,
StatusEnum,
)
from ._deterministic import combined_seed, utcnow
from .files import FakeFilesNamespace, StoredFile
if TYPE_CHECKING:
from .server import FakeLlamaCloudServer
@dataclass
class ClassificationJobRecord:
job: ClassifyJob
results: ClassifyJobResults
files: List[StoredFile]
class FakeClassifyNamespace:
def __init__(
self, *, server: "FakeLlamaCloudServer", files: FakeFilesNamespace
) -> None:
self._server = server
self._files = files
self._jobs: Dict[str, ClassificationJobRecord] = {}
def register(self) -> None:
server = self._server
server.add_route(
"POST",
"/api/v1/classifier/jobs",
self._handle_create_job,
namespace="classify",
)
server.add_route(
"GET",
"/api/v1/classifier/jobs",
self._handle_list_jobs,
namespace="classify",
)
server.add_route(
"GET",
"/api/v1/classifier/jobs/{job_id}",
self._handle_get_job,
namespace="classify",
)
server.add_route(
"GET",
"/api/v1/classifier/jobs/{job_id}/results",
self._handle_get_results,
namespace="classify",
)
def _handle_create_job(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
file_ids = payload.get("file_ids", [])
rules_payload = payload.get("rules", [])
rules = [ClassifierRule.parse_obj(rule) for rule in rules_payload]
stored_files = []
for file_id in file_ids:
stored = self._files.get(file_id)
if not stored:
return self._server.json_response(
{"detail": f"File {file_id} not found"}, status_code=404
)
stored_files.append(stored)
job_id = self._server.new_id("classify-job")
job = ClassifyJob(
id=job_id,
project_id=request.url.params.get(
"project_id", self._server.default_project_id
),
user_id="fake-user",
rules=rules,
parsing_configuration=None,
status=StatusEnum.SUCCESS,
created_at=utcnow(),
updated_at=utcnow(),
effective_at=utcnow(),
error_message=None,
job_record_id=None,
)
results = self._build_results(job_id, stored_files, rules)
record = ClassificationJobRecord(job=job, results=results, files=stored_files)
self._jobs[job_id] = record
return self._server.json_response(job.dict())
def _handle_list_jobs(self, request: httpx.Request) -> httpx.Response:
return self._server.json_response(
[record.job.dict() for record in self._jobs.values()]
)
def _handle_get_job(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-1]
record = self._jobs.get(job_id)
if not record:
return self._server.json_response(
{"detail": "Job not found"}, status_code=404
)
return self._server.json_response(record.job.dict())
def _handle_get_results(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-2]
record = self._jobs.get(job_id)
if not record:
return self._server.json_response(
{"detail": "Results not found"}, status_code=404
)
return self._server.json_response(record.results.dict())
def _build_results(
self,
job_id: str,
stored_files: List[StoredFile],
rules: List[ClassifierRule],
) -> ClassifyJobResults:
items: List[FileClassification] = []
for stored in stored_files:
seed = combined_seed(stored.sha256, job_id)
rule_index = seed % len(rules) if rules else 0
predicted_type = rules[rule_index].type if rules else "unlabeled"
confidence = 0.55 + (seed % 40) / 100
reasoning = (
f"Selected rule '{predicted_type}' using deterministic seed {seed}."
)
classification = FileClassification(
id=self._server.new_id("classification"),
file_id=stored.file.id,
classify_job_id=job_id,
created_at=utcnow(),
updated_at=utcnow(),
result=ClassificationResult(
type=predicted_type,
confidence=min(confidence, 0.95),
reasoning=reasoning,
),
)
items.append(classification)
return ClassifyJobResults(
items=items, next_page_token=None, total_size=len(items)
)
@@ -0,0 +1,673 @@
from __future__ import annotations
from dataclasses import dataclass
from typing import TYPE_CHECKING, Any, Dict, List, Optional
import httpx
from llama_cloud.types import (
ExtractAgent,
ExtractConfig,
ExtractJob,
ExtractRun,
ExtractState,
File as CloudFile,
PaginatedExtractRunsResponse,
StatusEnum,
)
from ._deterministic import (
combined_seed,
generate_data_from_schema,
hash_schema,
utcnow,
)
from ._deterministic import fingerprint_file
from .files import FakeFilesNamespace, StoredFile
from .matchers import RequestContext, RequestMatcher
if TYPE_CHECKING:
from .server import FakeLlamaCloudServer
@dataclass
class ExtractRunStub:
matcher: Optional[RequestMatcher]
data: Optional[Any]
status: Optional[str]
metadata: Optional[Dict[str, Any]]
error: Optional[str]
job_status: Optional[str]
once: bool
@dataclass
class AgentRunStub:
agent_id: str
matcher: Optional[RequestMatcher]
job_status: Optional[str]
run_status: Optional[str]
error: Optional[str]
once: bool
@dataclass
class StoredRun:
job: ExtractJob
run: ExtractRun
class FakeExtractNamespace:
def __init__(
self,
*,
server: "FakeLlamaCloudServer",
files: FakeFilesNamespace,
) -> None:
self._server = server
self._files = files
self._jobs: Dict[str, StoredRun] = {}
self._runs: Dict[str, ExtractRun] = {}
self._agents: Dict[str, ExtractAgent] = {}
self._agents_by_name: Dict[str, str] = {}
self._run_stubs: List[ExtractRunStub] = []
self._agent_run_stubs: List[AgentRunStub] = []
self.routes: Dict[str, Any] = {}
# Public APIs ----------------------------------------------------
def stub_run(
self,
matcher: Optional[RequestMatcher],
*,
data: Optional[Any] = None,
status: Optional[str] = None,
job_status: Optional[str] = None,
metadata: Optional[Dict[str, Any]] = None,
error: Optional[str] = None,
once: bool = True,
) -> None:
self._run_stubs.append(
ExtractRunStub(
matcher=matcher,
data=data,
status=status,
metadata=metadata,
error=error,
job_status=job_status,
once=once,
)
)
def stub_agent_run(
self,
*,
agent_id: str,
matcher: Optional[RequestMatcher],
job_status: Optional[str] = None,
run_status: Optional[str] = None,
error: Optional[str] = None,
once: bool = True,
) -> None:
self._agent_run_stubs.append(
AgentRunStub(
agent_id=agent_id,
matcher=matcher,
job_status=job_status,
run_status=run_status,
error=error,
once=once,
)
)
# Route registration ---------------------------------------------
def register(self) -> None:
server = self._server
route = server.add_route(
"POST",
"/api/v1/extraction/run",
self._handle_stateless_run,
namespace="extract",
alias="extract_run",
)
self.routes["stateless_run"] = route
self.stateless_run = route
server.add_route(
"POST",
"/api/v1/extraction/extraction-agents",
self._handle_create_agent,
namespace="extract",
)
server.add_route(
"PATCH",
"/api/v1/extraction/extraction-agents/{agent_id}",
self._handle_update_agent,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/extraction-agents/{agent_id}",
self._handle_get_agent,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/extraction-agents/by-name/{name}",
self._handle_get_agent_by_name,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/extraction-agents",
self._handle_list_agents,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/extraction-agents/default",
self._handle_get_default_agent,
namespace="extract",
)
server.add_route(
"DELETE",
"/api/v1/extraction/extraction-agents/{agent_id}",
self._handle_delete_agent,
namespace="extract",
)
server.add_route(
"POST",
"/api/v1/extraction/extraction-agents/schema/validation",
self._handle_validate_schema,
namespace="extract",
)
agent_job_route = server.add_route(
"POST",
"/api/v1/extraction/jobs",
self._handle_agent_job,
namespace="extract",
alias="agent_job",
)
self.routes["agent_job"] = agent_job_route
self.agent_job = agent_job_route
server.add_route(
"POST",
"/api/v1/extraction/jobs/batch",
self._handle_agent_job_batch,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/jobs",
self._handle_list_jobs,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/jobs/{job_id}",
self._handle_get_job,
namespace="extract",
)
agent_run_route = server.add_route(
"GET",
"/api/v1/extraction/runs/by-job/{job_id}",
self._handle_get_run_by_job,
namespace="extract",
alias="agent_run",
)
self.routes["agent_run"] = agent_run_route
self.agent_run = agent_run_route
server.add_route(
"GET",
"/api/v1/extraction/runs/{run_id}",
self._handle_get_run,
namespace="extract",
)
server.add_route(
"DELETE",
"/api/v1/extraction/runs/{run_id}",
self._handle_delete_run,
namespace="extract",
)
server.add_route(
"GET",
"/api/v1/extraction/runs",
self._handle_list_runs,
namespace="extract",
)
# Handlers -------------------------------------------------------
def _handle_stateless_run(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
config = ExtractConfig.parse_obj(payload["config"])
data_schema = payload["data_schema"]
schema_hash = hash_schema(data_schema)
file_info = self._extract_file_info(payload, request)
agent = self._build_ephemeral_agent(
config, data_schema, file_info.file.project_id
)
context = RequestContext(
request=request,
json=payload,
file_id=file_info.file.id,
filename=file_info.file.name,
file_sha256=file_info.sha256,
schema_hash=schema_hash,
project_id=file_info.file.project_id,
organization_id=self._server.default_organization_id,
)
stub = self._pop_stub(self._run_stubs, context)
job_status = StatusEnum.SUCCESS
run_status = ExtractState.SUCCESS
metadata = {"deterministic": {"value": True}}
error = None
run_data = self._generate_run_data(data_schema, file_info.sha256)
if stub:
if stub.job_status:
job_status = StatusEnum(stub.job_status)
if stub.status:
run_status = ExtractState(stub.status)
if stub.metadata:
metadata = stub.metadata
if stub.error:
error = stub.error
if stub.data is not None:
if callable(stub.data):
run_data = stub.data(payload) # type: ignore[assignment]
else:
run_data = stub.data
stored = self._create_job_and_run(
agent=agent,
config=config,
data_schema=data_schema,
file_info=file_info,
job_status=job_status,
run_status=run_status,
metadata=metadata,
data=run_data,
error=error,
project_id=file_info.file.project_id,
)
return self._server.json_response(stored.job.dict())
def _handle_create_agent(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
name = payload["name"]
config = ExtractConfig.parse_obj(payload["config"])
data_schema = payload["data_schema"]
agent_id = self._server.new_id("agent")
agent = ExtractAgent(
id=agent_id,
name=name,
config=config,
data_schema=data_schema,
project_id=request.url.params.get(
"project_id", self._server.default_project_id
),
created_at=utcnow(),
updated_at=utcnow(),
custom_configuration=None,
)
self._agents[agent_id] = agent
self._agents_by_name[name] = agent_id
return self._server.json_response(agent.dict())
def _handle_update_agent(self, request: httpx.Request) -> httpx.Response:
agent_id = request.url.path.split("/")[-1]
if agent_id not in self._agents:
return self._server.json_response(
{"detail": "Agent not found"}, status_code=404
)
payload = self._server.json(request)
agent = self._agents[agent_id]
config = payload.get("config", agent.config)
data_schema = payload.get("data_schema", agent.data_schema)
updated = agent.copy(
update={
"config": ExtractConfig.parse_obj(config)
if isinstance(config, dict)
else config,
"data_schema": data_schema,
"updated_at": utcnow(),
}
)
self._agents[agent_id] = updated
return self._server.json_response(updated.dict())
def _handle_get_agent(self, request: httpx.Request) -> httpx.Response:
agent_id = request.url.path.split("/")[-1]
agent = self._agents.get(agent_id)
if not agent:
return self._server.json_response(
{"detail": "Agent not found"}, status_code=404
)
return self._server.json_response(agent.dict())
def _handle_get_agent_by_name(self, request: httpx.Request) -> httpx.Response:
name = request.url.path.split("/")[-1]
agent_id = self._agents_by_name.get(name)
if not agent_id:
return self._server.json_response(
{"detail": "Agent not found"}, status_code=404
)
return self._server.json_response(self._agents[agent_id].dict())
def _handle_list_agents(self, request: httpx.Request) -> httpx.Response:
include_default = (
request.url.params.get("include_default", "false").lower() == "true"
)
agents = list(self._agents.values())
if include_default and not agents:
default_agent = self._build_ephemeral_agent(
ExtractConfig(),
{"type": "object", "properties": {}},
self._server.default_project_id,
)
agents.append(default_agent)
return self._server.json_response([agent.dict() for agent in agents])
def _handle_get_default_agent(self, request: httpx.Request) -> httpx.Response:
if self._agents:
agent = next(iter(self._agents.values()))
else:
agent = self._build_ephemeral_agent(
ExtractConfig(),
{"type": "object", "properties": {}},
self._server.default_project_id,
)
return self._server.json_response(agent.dict())
def _handle_delete_agent(self, request: httpx.Request) -> httpx.Response:
agent_id = request.url.path.split("/")[-1]
agent = self._agents.pop(agent_id, None)
if agent:
self._agents_by_name.pop(agent.name, None)
return self._server.json_response({}, status_code=200)
def _handle_validate_schema(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
return self._server.json_response({"data_schema": payload["data_schema"]})
def _handle_agent_job(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
agent_id = payload["extraction_agent_id"]
agent = self._agents.get(agent_id)
if not agent:
return self._server.json_response(
{"detail": "Agent not found"}, status_code=404
)
file_id = payload["file_id"]
stored_file = self._files._files.get(file_id)
if not stored_file:
return self._server.json_response(
{"detail": "File not found"}, status_code=404
)
schema = payload.get("data_schema_override", agent.data_schema)
config_payload = payload.get("config_override", agent.config)
config = (
ExtractConfig.parse_obj(config_payload)
if isinstance(config_payload, dict)
else config_payload
)
stub = self._pop_agent_stub(
agent_id, RequestContext(request=request, json=payload)
)
job_status = StatusEnum.SUCCESS
run_status = ExtractState.SUCCESS
error = None
if stub:
if stub.job_status:
job_status = StatusEnum(stub.job_status)
if stub.run_status:
run_status = ExtractState(stub.run_status)
if stub.error:
error = stub.error
stored = self._create_job_and_run(
agent=agent,
config=config,
data_schema=schema,
file_info=stored_file,
job_status=job_status,
run_status=run_status,
metadata={"agent": {"value": agent.id}},
data=self._generate_run_data(schema, stored_file.sha256),
error=error,
project_id=agent.project_id,
)
return self._server.json_response(stored.job.dict())
def _handle_agent_job_batch(self, request: httpx.Request) -> httpx.Response:
payload = self._server.json(request)
file_ids = payload.get("file_ids", [])
jobs = []
for file_id in file_ids:
request_body = payload.copy()
request_body["file_id"] = file_id
fake_request = request.copy()
fake_request._content = self._server.encode_json(request_body)
response = self._handle_agent_job(fake_request)
if response.status_code != 200:
return response
jobs.append(response.json())
return self._server.json_response(jobs)
def _handle_list_jobs(self, request: httpx.Request) -> httpx.Response:
agent_id = request.url.params.get("extraction_agent_id")
items = []
for stored in self._jobs.values():
if agent_id and stored.job.extraction_agent.id != agent_id:
continue
items.append(stored.job.dict())
return self._server.json_response(items)
def _handle_get_job(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-1]
stored = self._jobs.get(job_id)
if not stored:
return self._server.json_response(
{"detail": "Job not found"}, status_code=404
)
return self._server.json_response(stored.job.dict())
def _handle_get_run_by_job(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-1]
stored = self._jobs.get(job_id)
if not stored:
return self._server.json_response(
{"detail": "Run not found"}, status_code=404
)
return self._server.json_response(stored.run.dict())
def _handle_get_run(self, request: httpx.Request) -> httpx.Response:
run_id = request.url.path.split("/")[-1]
run = self._runs.get(run_id)
if not run:
return self._server.json_response(
{"detail": "Run not found"}, status_code=404
)
return self._server.json_response(run.dict())
def _handle_delete_run(self, request: httpx.Request) -> httpx.Response:
run_id = request.url.path.split("/")[-1]
self._runs.pop(run_id, None)
to_delete = [
job_id for job_id, stored in self._jobs.items() if stored.run.id == run_id
]
for job_id in to_delete:
self._jobs.pop(job_id, None)
return self._server.json_response({}, status_code=200)
def _handle_list_runs(self, request: httpx.Request) -> httpx.Response:
agent_id = request.url.params.get("extraction_agent_id")
skip = int(request.url.params.get("skip", "0"))
limit = int(request.url.params.get("limit", "50"))
filtered = [
stored.run
for stored in self._jobs.values()
if not agent_id or stored.job.extraction_agent.id == agent_id
]
page = filtered[skip : skip + limit]
response = PaginatedExtractRunsResponse(
items=page,
skip=skip,
limit=limit,
total=len(filtered),
)
return self._server.json_response(response.dict())
# Internal helpers -----------------------------------------------
def _extract_file_info(
self, payload: Dict[str, Any], request: httpx.Request
) -> StoredFile:
if "file_id" in payload:
file_id = payload["file_id"]
stored = self._files.get(file_id)
if not stored:
raise ValueError("file_id not found in fake store")
return stored
if "file" in payload:
content, filename = self._files.decode_file_data(payload)
file_id = self._server.new_id("file")
stored = StoredFile(
file=CloudFile(
id=file_id,
name=filename or f"inline-{file_id}",
project_id=request.url.params.get(
"project_id", self._server.default_project_id
),
external_file_id=None,
file_size=len(content),
file_type=None,
created_at=utcnow(),
updated_at=utcnow(),
data_source_id=None,
permission_info=None,
resource_info=None,
last_modified_at=utcnow(),
),
content=content,
sha256=fingerprint_file(content, filename),
)
return stored
if "text" in payload:
text_bytes = payload["text"].encode("utf-8")
file_id = self._server.new_id("file")
stored = StoredFile(
file=CloudFile(
id=file_id,
name=f"text-{file_id}.txt",
project_id=self._server.default_project_id,
external_file_id=None,
file_size=len(text_bytes),
file_type="text/plain",
created_at=utcnow(),
updated_at=utcnow(),
data_source_id=None,
permission_info=None,
resource_info=None,
last_modified_at=utcnow(),
),
content=text_bytes,
sha256=fingerprint_file(text_bytes, None),
)
return stored
raise ValueError("file payload missing")
def _build_ephemeral_agent(
self,
config: ExtractConfig,
data_schema: Dict[str, Any],
project_id: str,
) -> ExtractAgent:
return ExtractAgent(
id=self._server.new_id("agent"),
name="stateless-agent",
config=config,
data_schema=data_schema,
project_id=project_id,
created_at=utcnow(),
updated_at=utcnow(),
custom_configuration=None,
)
def _generate_run_data(self, schema: Dict[str, Any], file_hash: str) -> Any:
seed = combined_seed(file_hash, hash_schema(schema))
return generate_data_from_schema(schema, seed)
def _create_job_and_run(
self,
*,
agent: ExtractAgent,
config: ExtractConfig,
data_schema: Dict[str, Any],
file_info: StoredFile,
job_status: StatusEnum,
run_status: ExtractState,
metadata: Dict[str, Any],
data: Any,
error: Optional[str],
project_id: str,
) -> StoredRun:
job_id = self._server.new_id("job")
run_id = self._server.new_id("run")
now = utcnow()
job = ExtractJob(
id=job_id,
file=file_info.file,
extraction_agent=agent,
status=job_status,
error=error,
)
run = ExtractRun(
id=run_id,
job_id=job_id,
file=file_info.file,
extraction_agent_id=agent.id,
status=run_status,
config=config,
data_schema=data_schema,
data=data,
extraction_metadata=metadata,
created_at=now,
updated_at=now,
from_ui=False,
error=error,
project_id=project_id,
)
stored = StoredRun(job=job, run=run)
self._jobs[job_id] = stored
self._runs[run_id] = run
return stored
def _pop_stub(
self,
stubs: List[ExtractRunStub],
context: RequestContext,
) -> Optional[ExtractRunStub]:
for index, stub in enumerate(list(stubs)):
if context.matches(stub.matcher):
if stub.once:
stubs.pop(index)
return stub
return None
def _pop_agent_stub(
self,
agent_id: str,
context: RequestContext,
) -> Optional[AgentRunStub]:
for index, stub in enumerate(list(self._agent_run_stubs)):
if stub.agent_id != agent_id:
continue
if context.matches(stub.matcher):
if stub.once:
self._agent_run_stubs.pop(index)
return stub
return None
@@ -0,0 +1,324 @@
from __future__ import annotations
import base64
from dataclasses import dataclass
from pathlib import Path
from typing import TYPE_CHECKING, Any, Dict, List, Optional
from urllib.parse import urlencode
import httpx
import respx
from llama_cloud.types import File as CloudFile
from llama_cloud.types import FileIdPresignedUrl, PresignedUrl
from ._deterministic import fingerprint_file, hash_chunks, utcnow
from .matchers import RequestContext, RequestMatcher
if TYPE_CHECKING:
from .server import FakeLlamaCloudServer
@dataclass
class StoredFile:
file: CloudFile
content: bytes
sha256: str
@dataclass
class PendingUpload:
file_id: str
filename: str
project_id: str
organization_id: str
external_file_id: Optional[str]
expected_size: Optional[int]
class FakeFilesNamespace:
def __init__(
self,
*,
server: "FakeLlamaCloudServer",
upload_base_url: str,
download_base_url: str,
) -> None:
self._server = server
self._upload_base_url = upload_base_url.rstrip("/")
self._download_base_url = download_base_url.rstrip("/")
self._files: Dict[str, StoredFile] = {}
self._pending: Dict[str, PendingUpload] = {}
self._upload_stubs: List[
tuple[RequestMatcher | None, int, Dict[str, Any], bool]
] = []
self.routes: Dict[str, respx.Route] = {}
# Public helpers -------------------------------------------------
def preload(self, *, path: str | Path, filename: Optional[str] = None) -> str:
path = Path(path)
content = path.read_bytes()
file_id = self._server.new_id("file")
name = filename or path.name
stored = self._build_file(
file_id=file_id,
name=name,
project_id=self._server.default_project_id,
organization_id=self._server.default_organization_id,
content=content,
external_file_id=None,
)
self._files[file_id] = stored
return file_id
def read(self, file_id: str) -> bytes:
return self._files[file_id].content
def get(self, file_id: str) -> Optional[StoredFile]:
return self._files.get(file_id)
def stub_upload(
self,
matcher: Optional[RequestMatcher],
*,
status_code: int = 413,
json_body: Optional[Dict[str, Any]] = None,
once: bool = True,
) -> None:
body = json_body or {"detail": "upload rejected by fake server"}
self._upload_stubs.append((matcher, status_code, body, once))
# Route registration ---------------------------------------------
def register(self) -> None:
server = self._server
server.add_route(
"PUT",
"/api/v1/files",
self._handle_generate_presigned_url,
namespace="files",
alias="generate_presigned_url",
)
upload_route = server.add_route(
"POST",
"/api/v1/files",
self._handle_direct_upload,
namespace="files",
alias="upload",
)
self.routes["upload"] = upload_route
get_route = server.add_route(
"GET",
"/api/v1/files/{file_id}",
self._handle_get_metadata,
namespace="files",
alias="get",
)
self.routes["get"] = get_route
server.add_route(
"DELETE",
"/api/v1/files/{file_id}",
self._handle_delete,
namespace="files",
)
server.add_route(
"GET",
"/api/v1/files/{file_id}/content",
self._handle_read_content,
namespace="files",
alias="read_content",
)
server.add_route(
"PUT",
"/upload/{file_id}",
self._handle_presigned_upload,
namespace="files",
base_urls=[self._upload_base_url],
alias="presigned_upload",
)
server.add_route(
"GET",
"/files/{file_id}",
self._handle_presigned_download,
namespace="files",
base_urls=[self._download_base_url],
alias="download",
)
# Handlers -------------------------------------------------------
def _handle_generate_presigned_url(self, request: httpx.Request) -> httpx.Response:
data = self._server.json(request)
now = utcnow()
file_id = self._server.new_id("file")
name = data.get("name") or f"upload-{file_id}.bin"
pending = PendingUpload(
file_id=file_id,
filename=name,
project_id=request.url.params.get(
"project_id", self._server.default_project_id
),
organization_id=request.url.params.get(
"organization_id", self._server.default_organization_id
),
external_file_id=data.get("external_file_id"),
expected_size=data.get("file_size"),
)
self._pending[file_id] = pending
presigned = FileIdPresignedUrl(
file_id=file_id,
url=f"{self._upload_base_url}/upload/{file_id}",
expires_at=now,
form_fields=None,
)
return self._server.json_response(presigned.dict())
def _handle_direct_upload(self, request: httpx.Request) -> httpx.Response:
file_bytes, filename = self._extract_multipart_file(request)
file_id = self._server.new_id("file")
stored = self._build_file(
file_id=file_id,
name=filename or f"upload-{file_id}.bin",
project_id=request.url.params.get(
"project_id", self._server.default_project_id
),
organization_id=request.url.params.get(
"organization_id", self._server.default_organization_id
),
content=file_bytes,
external_file_id=request.url.params.get("external_file_id"),
)
self._files[file_id] = stored
return self._server.json_response(stored.file.dict())
def _handle_get_metadata(self, request: httpx.Request) -> httpx.Response:
file_id = request.url.path.split("/")[-1]
if file_id not in self._files:
return self._server.json_response(
{"detail": "File not found"}, status_code=404
)
return self._server.json_response(self._files[file_id].file.dict())
def _handle_delete(self, request: httpx.Request) -> httpx.Response:
file_id = request.url.path.split("/")[-1]
self._files.pop(file_id, None)
self._pending.pop(file_id, None)
return self._server.json_response({}, status_code=200)
def _handle_read_content(self, request: httpx.Request) -> httpx.Response:
file_id = request.url.path.split("/")[-2]
if file_id not in self._files:
return self._server.json_response(
{"detail": "File not found"}, status_code=404
)
presigned = PresignedUrl(
url=f"{self._download_base_url}/files/{file_id}?{urlencode({'token': 'fake'})}",
expires_at=utcnow(),
form_fields=None,
)
return self._server.json_response(presigned.dict())
def _handle_presigned_upload(self, request: httpx.Request) -> httpx.Response:
file_id = request.url.path.split("/")[-1]
pending = self._pending.get(file_id)
context = RequestContext(
request=request,
json=None,
file_id=file_id,
filename=pending.filename if pending else None,
file_sha256=hash_chunks([request.content]),
)
for index, (matcher, status, body, once) in enumerate(list(self._upload_stubs)):
if context.matches(matcher):
if once:
self._upload_stubs.pop(index)
return self._server.json_response(body, status_code=status)
if pending is None:
return self._server.json_response(
{"detail": "Unknown file"}, status_code=404
)
stored = self._build_file(
file_id=file_id,
name=pending.filename,
project_id=pending.project_id,
organization_id=pending.organization_id,
content=request.content,
external_file_id=pending.external_file_id,
)
self._files[file_id] = stored
self._pending.pop(file_id, None)
return httpx.Response(204)
def _handle_presigned_download(self, request: httpx.Request) -> httpx.Response:
file_id = request.url.path.split("/")[-1]
stored = self._files.get(file_id)
if not stored:
return httpx.Response(404, json={"detail": "File not found"})
return httpx.Response(200, content=stored.content)
# Internal helpers -----------------------------------------------
def _build_file(
self,
*,
file_id: str,
name: str,
project_id: str,
organization_id: str,
content: bytes,
external_file_id: Optional[str],
) -> StoredFile:
sha256 = fingerprint_file(content, name)
now = utcnow()
cloud_file = CloudFile(
id=file_id,
name=name,
project_id=project_id,
external_file_id=external_file_id,
file_size=len(content),
file_type=Path(name).suffix or "application/octet-stream",
created_at=now,
updated_at=now,
data_source_id=None,
permission_info=None,
resource_info=None,
last_modified_at=now,
)
return StoredFile(file=cloud_file, content=content, sha256=sha256)
def _extract_multipart_file(
self, request: httpx.Request
) -> tuple[bytes, Optional[str]]:
content_type = request.headers.get("content-type", "")
if "multipart/form-data" not in content_type:
raise ValueError("Expected multipart upload")
boundary = content_type.split("boundary=")[-1]
boundary_bytes = boundary.encode("utf-8")
body = request.content
delimiter = b"--" + boundary_bytes
parts = [
part
for part in body.split(delimiter)
if part.strip(b"\r\n") and part.strip(b"\r\n") != b"--"
]
for part in parts:
headers, _, payload = part.partition(b"\r\n\r\n")
header_text = headers.decode("utf-8", errors="ignore")
if 'name="upload_file"' in header_text or 'name="file"' in header_text:
filename = None
if "filename=" in header_text:
filename = (
header_text.split("filename=")[-1].strip().strip('"').strip("'")
)
return payload.rstrip(b"\r\n"), filename
raise ValueError("upload file part not found")
def decode_file_data(self, data: Dict[str, Any]) -> tuple[bytes, Optional[str]]:
if "file" not in data:
raise ValueError("file payload missing")
file_payload = data["file"]
encoded = file_payload["data"]
content = base64.b64decode(encoded)
filename = file_payload.get("filename")
return content, filename
@@ -0,0 +1,100 @@
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Callable, Optional
import httpx
MatcherPredicate = Callable[[httpx.Request], bool]
@dataclass
class FileMatcher:
filename: Optional[str] = None
sha256: Optional[str] = None
file_id: Optional[str] = None
@dataclass
class SchemaMatcher:
model: Optional[type] = None
schema_hash: Optional[str] = None
@dataclass
class RequestMatcher:
file: Optional[FileMatcher | MatcherPredicate] = None
schema: Optional[SchemaMatcher] = None
agent_id: Optional[str] = None
project_id: Optional[str] = None
organization_id: Optional[str] = None
predicate: Optional[MatcherPredicate] = None
@dataclass
class RequestContext:
request: httpx.Request
json: Optional[dict[str, Any]]
file_id: Optional[str] = None
filename: Optional[str] = None
file_sha256: Optional[str] = None
schema_hash: Optional[str] = None
agent_id: Optional[str] = None
project_id: Optional[str] = None
organization_id: Optional[str] = None
def matches(self, matcher: Optional[RequestMatcher]) -> bool:
if matcher is None:
return True
if matcher.project_id and matcher.project_id != self.project_id:
return False
if matcher.organization_id and matcher.organization_id != self.organization_id:
return False
if matcher.agent_id and matcher.agent_id != self.agent_id:
return False
if matcher.file:
if isinstance(matcher.file, FileMatcher):
if matcher.file.filename and matcher.file.filename != self.filename:
return False
if matcher.file.file_id and matcher.file.file_id != self.file_id:
return False
if matcher.file.sha256 and matcher.file.sha256 != self.file_sha256:
return False
else:
if not matcher.file(self.request):
return False
if matcher.schema:
if (
matcher.schema.schema_hash
and matcher.schema.schema_hash != self.schema_hash
):
return False
if matcher.schema.model and matcher.schema.schema_hash:
return matcher.schema.schema_hash == self.schema_hash
if matcher.schema.model and matcher.schema.schema_hash is None:
expected = _schema_hash_from_model(matcher.schema.model)
return expected == self.schema_hash
if matcher.predicate and not matcher.predicate(self.request):
return False
return True
def _schema_hash_from_model(model: type) -> Optional[str]:
if hasattr(model, "model_json_schema"):
schema = model.model_json_schema()
elif hasattr(model, "schema"):
schema = model.schema() # type: ignore[attr-defined]
else:
return None
from ._deterministic import hash_schema
return hash_schema(schema)
@@ -0,0 +1,153 @@
from __future__ import annotations
from dataclasses import dataclass
from typing import TYPE_CHECKING, Any, Dict
import httpx
from ._deterministic import generate_text_blob, hash_schema
if TYPE_CHECKING:
from .server import FakeLlamaCloudServer
@dataclass
class ParseJobRecord:
job_id: str
file_name: str
status: str
result: Dict[str, Any]
content: bytes
class FakeParseNamespace:
def __init__(self, *, server: "FakeLlamaCloudServer") -> None:
self._server = server
self._jobs: Dict[str, ParseJobRecord] = {}
self.routes: Dict[str, Any] = {}
def register(self) -> None:
server = self._server
server.add_route(
"POST",
"/api/parsing/upload",
self._handle_upload,
namespace="parse",
)
server.add_route(
"GET",
"/api/parsing/job/{job_id}",
self._handle_job_status,
namespace="parse",
)
server.add_route(
"GET",
"/api/parsing/job/{job_id}/result/{result_type}",
self._handle_job_result,
namespace="parse",
)
def _handle_upload(self, request: httpx.Request) -> httpx.Response:
file_bytes, filename, form_data = self._split_multipart(request)
job_id = self._server.new_id("parse-job")
seed_hash = hash_schema({"filename": filename, "form": form_data})
seed = int(seed_hash[:16], 16)
page_text = generate_text_blob(seed, sentences=3)
pages: list[Dict[str, Any]] = [
{
"page": index + 1,
"text": f"{page_text} (page {index + 1})",
"md": f"{page_text} (page {index + 1})",
"images": [],
"charts": [],
"tables": [],
"layout": [],
"items": [],
"status": "SUCCESS",
"links": [],
"width": 8.5,
"height": 11.0,
"parsingMode": "deterministic",
"structuredData": {},
"noStructuredContent": False,
"noTextContent": False,
"isAudioTranscript": False,
"durationInSeconds": None,
"slideSpeakerNotes": None,
}
for index in range(1)
]
result = {
"job_id": job_id,
"status": "SUCCESS",
"file_name": filename,
"is_done": True,
"pages": pages,
"job_metadata": {"job_pages": len(pages)},
"text": "\n\n".join(str(page["text"]) for page in pages),
"markdown": "\n\n".join(str(page["md"]) for page in pages),
"json": {"pages": pages},
}
record = ParseJobRecord(
job_id=job_id,
file_name=filename,
status="SUCCESS",
result=result,
content=file_bytes,
)
self._jobs[job_id] = record
return self._server.json_response({"id": job_id})
def _handle_job_status(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-1]
job = self._jobs.get(job_id)
if not job:
return self._server.json_response(
{"detail": "Job not found"}, status_code=404
)
return self._server.json_response({"id": job_id, "status": job.status})
def _handle_job_result(self, request: httpx.Request) -> httpx.Response:
job_id = request.url.path.split("/")[-3]
job = self._jobs.get(job_id)
if not job:
return self._server.json_response(
{"detail": "Result not found"}, status_code=404
)
return self._server.json_response(job.result)
def _split_multipart(
self, request: httpx.Request
) -> tuple[bytes, str, Dict[str, str]]:
content_type = request.headers.get("content-type", "")
if "multipart/form-data" not in content_type:
raise ValueError("Expected multipart form data for parse upload")
boundary = content_type.split("boundary=")[-1]
delimiter = f"--{boundary}".encode()
closing = f"--{boundary}--".encode()
parts = []
body = request.content
for chunk in body.split(delimiter):
chunk = chunk.strip()
if not chunk or chunk == closing:
continue
parts.append(chunk)
file_bytes = b""
filename = "upload.pdf"
form_data: Dict[str, str] = {}
for part in parts:
header_blob, _, payload = part.partition(b"\r\n\r\n")
payload = payload.rstrip(b"\r\n")
header_text = header_blob.decode("utf-8", errors="ignore")
if "filename=" in header_text:
filename = (
header_text.split("filename=")[-1].strip().strip('"').strip("'")
)
file_bytes = payload
else:
name = header_text.split('name="')[-1].split('"')[0].strip()
form_data[name] = payload.decode("utf-8", errors="ignore")
if not file_bytes:
raise ValueError("File part missing from multipart payload")
return file_bytes, filename, form_data
@@ -0,0 +1,173 @@
from __future__ import annotations
import json
import re
import uuid
from typing import Any, Callable, Dict, Optional, Sequence
import httpx
import respx
from .classify import FakeClassifyNamespace
from .extract import FakeExtractNamespace
from .files import FakeFilesNamespace
from .parse import FakeParseNamespace
from .agent_data import FakeAgentDataNamespace
Handler = Callable[[httpx.Request], httpx.Response]
class FakeLlamaCloudServer:
DEFAULT_BASE_URL = "https://api.cloud.llamaindex.ai"
DEFAULT_UPLOAD_BASE = "https://uploads.fake-llama.test"
DEFAULT_DOWNLOAD_BASE = "https://downloads.fake-llama.test"
def __init__(
self,
*,
base_urls: Optional[Sequence[str]] = None,
namespaces: Optional[Sequence[str]] = None,
upload_base_url: Optional[str] = None,
download_base_url: Optional[str] = None,
default_project_id: str = "proj-test",
default_organization_id: str = "org-test",
) -> None:
self.base_urls = tuple(base_urls or (self.DEFAULT_BASE_URL,))
selected = namespaces or ("files", "extract", "parse", "classify", "agent_data")
self._namespace_names = {name.lower() for name in selected}
self._upload_base_url = upload_base_url or self.DEFAULT_UPLOAD_BASE
self._download_base_url = download_base_url or self.DEFAULT_DOWNLOAD_BASE
self.default_project_id = default_project_id
self.default_organization_id = default_organization_id
self.router = respx.MockRouter(assert_all_called=False)
self._installed = False
self._registered = False
self.files = FakeFilesNamespace(
server=self,
upload_base_url=self._upload_base_url,
download_base_url=self._download_base_url,
)
self.extract = FakeExtractNamespace(server=self, files=self.files)
self.parse = FakeParseNamespace(server=self)
self.classify = FakeClassifyNamespace(server=self, files=self.files)
self.agent_data = FakeAgentDataNamespace(server=self)
# Context management ----------------------------------------------
def install(self) -> "FakeLlamaCloudServer":
if not self._registered:
self._register_namespaces()
if not self._installed:
self.router.__enter__()
self._installed = True
return self
def uninstall(self) -> None:
if self._installed:
self.router.__exit__(None, None, None)
self._installed = False
def __enter__(self) -> "FakeLlamaCloudServer":
return self.install()
def __exit__(self, exc_type: Any, exc: Any, tb: Any) -> None:
self.uninstall()
# Route utilities -------------------------------------------------
def add_route(
self,
method: str,
path: str,
handler: Handler,
*,
namespace: str,
alias: Optional[str] = None,
base_urls: Optional[Sequence[str]] = None,
) -> respx.Route:
urls = base_urls or self.base_urls
first_route: Optional[respx.Route] = None
for base in urls:
route = self._register_route(method, base, path, handler)
if first_route is None:
first_route = route
if alias and first_route:
setattr(self, alias, first_route)
return first_route # type: ignore[return-value]
def _register_route(
self,
method: str,
base: str,
path: str,
handler: Handler,
) -> respx.Route:
url = self._build_url(base, path)
if "{" in path:
regex = self._compile_regex(base, path)
route = self.router.route(method=method, url__regex=regex)
else:
route = self.router.route(method=method, url=url)
route.mock(side_effect=lambda request, func=handler: func(request))
return route
def _build_url(self, base: str, path: str) -> str:
base = base.rstrip("/")
if not path.startswith("/"):
path = "/" + path
return f"{base}{path}"
def _compile_regex(self, base: str, path: str) -> re.Pattern[str]:
escaped = re.escape(base.rstrip("/"))
regex_path = re.sub(r"\{[^/]+\}", r"[^/]+", path)
pattern = f"^{escaped}{regex_path}$"
return re.compile(pattern)
# Helpers ---------------------------------------------------------
def json(self, request: httpx.Request) -> Dict[str, Any]:
if not request.content:
return {}
return json.loads(request.content.decode("utf-8"))
def encode_json(self, payload: Dict[str, Any]) -> bytes:
return json.dumps(payload).encode("utf-8")
def json_response(self, payload: Any, *, status_code: int = 200) -> httpx.Response:
body = json.dumps(payload, default=self._json_default).encode("utf-8")
headers = {"content-type": "application/json"}
return httpx.Response(status_code=status_code, headers=headers, content=body)
def new_id(self, prefix: str) -> str:
return f"{prefix}_{uuid.uuid4().hex[:8]}"
# Internal --------------------------------------------------------
def _json_default(self, value: Any) -> Any:
if hasattr(value, "model_dump"):
return value.model_dump()
if hasattr(value, "dict"):
return value.dict()
if isinstance(value, (set, frozenset)):
return list(value)
if isinstance(value, (bytes, bytearray)):
return value.decode("utf-8")
if hasattr(value, "isoformat"):
try:
return value.isoformat() # datetime/date support
except Exception:
pass
raise TypeError(f"{value!r} is not JSON serializable")
def _register_namespaces(self) -> None:
if "files" in self._namespace_names:
self.files.register()
if "extract" in self._namespace_names:
self.extract.register()
if "parse" in self._namespace_names:
self.parse.register()
if "classify" in self._namespace_names:
self.classify.register()
if "agent_data" in self._namespace_names:
self.agent_data.register()
self._registered = True
__all__ = ["FakeLlamaCloudServer"]
+102 -2
View File
@@ -3,11 +3,14 @@ import importlib.metadata
from contextlib import contextmanager
from typing import Generator
import difflib
from llama_cloud.types import StatusEnum
from llama_cloud.types import StatusEnum, File
import httpx
import packaging.version
from pydantic import BaseModel
from typing import Any, Dict, List, Tuple, Type
from typing import Any, Dict, List, Tuple, Type, Union, Optional
from io import BufferedIOBase, TextIOWrapper
from pathlib import Path
import secrets
# Asyncio error messages
nest_asyncio_err = "cannot be called from a running event loop"
@@ -104,3 +107,100 @@ def augment_async_errors() -> Generator[None, None, None]:
if nest_asyncio_err in str(e):
raise RuntimeError(nest_asyncio_msg)
raise
class SourceText:
"""
A wrapper class for providing text or file input with optional filename specification.
This class allows you to provide input in multiple ways:
- Direct text content via text_content parameter
- File paths as strings or Path objects
- Raw bytes
- File-like objects (BufferedIOBase, TextIOWrapper)
- Already-uploaded file ID via file_id parameter
Args:
file: The file input (bytes, file-like object, str path, or Path).
Mutually exclusive with text_content and file_id.
text_content: Raw text content to process. Mutually exclusive with file and file_id.
file_id: ID of an already-uploaded file. Mutually exclusive with file and text_content.
filename: Optional filename. Required for bytes/file-like objects without names.
If not provided, will be auto-generated for text_content or inferred from paths.
Examples:
# Direct text input
source = SourceText(text_content="Hello world")
# File path
source = SourceText(file="document.pdf")
# Bytes with filename
source = SourceText(file=b"...", filename="document.pdf")
# File-like object (will read from current position)
with open("document.pdf", "rb") as f:
source = SourceText(file=f)
# Already-uploaded file
source = SourceText(file_id="file_abc123")
"""
def __init__(
self,
*,
file: Union[bytes, BufferedIOBase, TextIOWrapper, str, Path, None] = None,
text_content: Optional[str] = None,
file_id: Optional[str] = None,
filename: Optional[str] = None,
):
self.file = file
self.filename = filename
self.text_content = text_content
self.file_id = file_id
self._validate()
def _validate(self) -> None:
"""Ensure filename is provided when needed."""
# Check that exactly one of file, text_content, or file_id is provided
provided = sum(
[
self.file is not None,
self.text_content is not None,
self.file_id is not None,
]
)
if provided == 0:
raise ValueError("One of file, text_content, or file_id must be provided.")
elif provided > 1:
raise ValueError(
"Only one of file, text_content, or file_id can be provided."
)
# If file_id is provided, we don't need filename validation
if self.file_id is not None:
return
if self.text_content is not None:
if not self.filename:
random_hex = secrets.token_hex(4)
self.filename = f"text_input_{random_hex}.txt"
return
if isinstance(self.file, (bytes, BufferedIOBase, TextIOWrapper)):
if not self.filename and hasattr(self.file, "name"):
self.filename = os.path.basename(str(self.file.name))
elif self.filename is None and not hasattr(self.file, "name"):
raise ValueError(
"filename must be provided when file is bytes or a file-like object without a name"
)
elif isinstance(self.file, (str, Path)):
if not self.filename:
self.filename = os.path.basename(str(self.file))
else:
raise ValueError(f"Unsupported file type: {type(self.file)}")
# Type alias for file input that can be used across services
FileInput = Union[str, Path, BufferedIOBase, SourceText, File]
+73
View File
@@ -0,0 +1,73 @@
# llama_parse
## 0.6.81
### Patch Changes
- Updated dependencies [f3233de]
- llama-cloud-services-py@0.6.81
## 0.6.80
### Patch Changes
- Updated dependencies [0506c88]
- llama-cloud-services-py@0.6.80
## 0.6.79
### Patch Changes
- Updated dependencies [e020e3e]
- llama-cloud-services-py@0.6.79
## 0.6.78
### Patch Changes
- 9f1ef4e: Fix extract
- Updated dependencies [9f1ef4e]
- llama-cloud-services-py@0.6.78
## 0.6.77
### Patch Changes
- Updated dependencies [407292b]
- llama-cloud-services-py@0.6.77
## 0.6.76
### Patch Changes
- Updated dependencies [4f24f53]
- llama-cloud-services-py@0.6.76
## 0.6.75
### Patch Changes
- Updated dependencies [f81532e]
- llama-cloud-services-py@0.6.75
## 0.6.74
### Patch Changes
- Updated dependencies [1bf5223]
- Updated dependencies [24166dc]
- llama-cloud-services-py@0.6.74
## 0.6.73
### Patch Changes
- Updated dependencies [e6a7939]
- llama-cloud-services-py@0.6.73
## 0.6.72
### Patch Changes
- Updated dependencies [ad6734b]
- llama-cloud-services-py@0.6.72
+20
View File
@@ -0,0 +1,20 @@
{
"name": "llama_parse",
"version": "0.6.81",
"description": "",
"main": "index.js",
"private": false,
"scripts": {
"test": "echo \"Error: no test specified\" && exit 1"
},
"dependencies": {
"llama-cloud-services-py": "workspace:*"
},
"keywords": [],
"author": "",
"license": "ISC",
"packageManager": "pnpm@10.11.1",
"devDependencies": {
"changesets": "^1.0.2"
}
}
+2 -2
View File
@@ -11,13 +11,13 @@ dev = [
[project]
name = "llama-parse"
version = "0.6.70"
version = "0.6.81"
description = "Parse files into RAG-Optimized formats."
authors = [{name = "Logan Markewich", email = "logan@llamaindex.ai"}]
requires-python = ">=3.9,<4.0"
readme = "README.md"
license = "MIT"
dependencies = ["llama-cloud-services>=0.6.70"]
dependencies = ["llama-cloud-services>=0.6.81"]
[project.scripts]
llama-parse = "llama_parse.cli.main:parse"
+6 -3
View File
@@ -1,7 +1,10 @@
{
"name": "llama-cloud-services-py",
"version": "0.6.70",
"private": "true",
"version": "0.6.81",
"private": false,
"license": "MIT",
"scripts": {}
"scripts": {},
"devDependencies": {
"changesets": "^1.0.2"
}
}
+8 -4
View File
@@ -14,12 +14,15 @@ dev = [
"ipython>=8.12.3,<9",
"jupyter>=1.1.1,<2",
"mypy>=1.14.1,<2",
"pydantic-settings>=2.10.1"
"pydantic-settings>=2.10.1",
"pandas",
"openpyxl",
"pyarrow"
]
[project]
name = "llama-cloud-services"
version = "0.6.70"
version = "0.6.81"
description = "Tailored SDK clients for LlamaCloud services."
authors = [{name = "Logan Markewich", email = "logan@runllama.ai"}]
requires-python = ">=3.9,<4.0"
@@ -27,14 +30,15 @@ readme = "README.md"
license = "MIT"
dependencies = [
"llama-index-core>=0.12.0",
"llama-cloud==0.1.43",
"llama-cloud==0.1.44",
"pydantic>=2.8,!=2.10",
"click>=8.1.7,<9",
"python-dotenv>=1.0.1,<2",
"eval-type-backport>=0.2.0,<0.3 ; python_version < '3.10'",
"platformdirs>=4.3.7,<5",
"tenacity>=8.5.0, <10.0",
"packaging>=25.0"
"packaging>=23.0",
"respx[tests]>=0.22.0"
]
[project.scripts]
+101
View File
@@ -0,0 +1,101 @@
# Testing Utils Implementation Plan
## Goal
Build a `FakeLlamaCloudServer` that intercepts all SDK HTTP traffic (extract, parse, classify, files, etc.) and returns deterministic responses so offline tests behave like production, per `testing_utils_spec.md`.
## High-Level Phases
1. **Router + lifecycle scaffold**
- Implement `FakeLlamaCloudServer` with `respx.Router` that can be used as a context manager or via explicit `install()` / `uninstall()`.
- Support multiple base URLs and namespace filtering so only selected APIs are intercepted.
2. **State stores + deterministic generators**
- Create in-memory stores for files, jobs, runs, parse results, and classification predictions.
- Implement deterministic data generation seeded by (file hash, schema hash, namespace) as described in the spec.
3. **Namespace handlers**
- Extract: stub `/api/v1/extraction/extraction-agents/*`, `/api/v1/extraction/run`, `/api/v1/extraction/jobs*`, `/api/v1/extraction/runs/by-job/{id}`.
- Files: stub `/api/v1/files/**` plus presigned upload/download workflows.
- Parse: stub `/api/parsing/upload`, `/api/parsing/job/{id}`, `/api/parsing/job/{id}/result/{result_type}`.
- Classify: stub `/api/v1/classifier/*` (job creation, polling, results).
4. **Matcher + override system**
- Implement `RequestMatcher`, `FileMatcher`, `SchemaMatcher`, etc., and expose helper APIs like `fake.extract.stub_run`.
5. **Ergonomic utilities**
- Provide helper shortcuts (`fake.extract.stateless_run`) and spy APIs (call count assertions).
6. **Docs + tests**
- Document usage in `testing_utils_spec.md`.
- Add tests that demonstrate end-to-end flows using the fake.
## Detailed Steps & Considerations
### 1. Router & Lifecycle
- Create a single `FakeLlamaCloudServer` class that holds a `respx.Router` configured for each base URL.
- Provide `__enter__/__exit__` plus `install()/uninstall()` to attach/detach the router.
- Complexity: need to handle both sync and async clients since `LlamaParse` uses raw `httpx.AsyncClient` instances constructed on the fly. Ensure `respx.mock(assert_all_called=False)` works for both.
### 2. State Management & Determinism
- Implement a `FileStore` that tracks uploaded file bytes, metadata, generated IDs, and seeded RNG values.
- Implement `ExtractStore`, `ParseStore`, `ClassifyStore` to track job lifecycles and generated runs.
- Deterministic generator design:
- Compute SHA256 of (file bytes + filename) and of normalized schema JSON.
- Combine into a seed (e.g., `seed = sha256(file_hash + schema_hash)`).
- Use that seed for namespace-specific RNG (extract uses schema walk, parse uses layout heuristics, classify uses label sets).
- Complexity: schema normalization requires Pydantic `model_json_schema()` ordering; ensure we match production ordering to avoid drift.
### 3. Namespace Handlers
#### Extract
- Stub endpoints listed under `LlamaExtract` usage (`create_extraction_agent`, `run_job`, stateless run, poll job/run).
- Mirror response bodies (`ExtractJob`, `ExtractRun`, `PaginatedExtractRunsResponse`) so the SDKs type deserialization works.
- Manage transitions `PENDING → SUCCESS/FAILED` with realistic timestamps.
#### Files
- Implement both presigned workflow and direct upload fallback:
- `POST /api/v1/files` (or equivalent) returns a fake presigned URL (e.g., `https://fake-upload.local/{file_id}`) that our router also intercepts.
- The subsequent `PUT` should store the bytes and mark upload complete.
- `GET /api/v1/files/{id}` returns stored metadata; `read_file_content` returns presigned download URLs or raw bytes.
- Complexity: need to intercept arbitrary presigned hostnames (e.g., AWS S3). Spec does not clarify if presigned URLs live on the same base; we may need to whitelist custom domains or provide a fake S3 host.
#### Parse
- Because `LlamaParse` manually constructs `/api/parsing/*` URLs, ensure the fake registers these exact routes against every provided base URL.
- Store job configs, return deterministic `JobResult` payloads (text, markdown, JSON), and support partitioned jobs.
#### Classify
- Stub job creation/polling/responses, ensuring statuses transition according to `StatusEnum`.
- Return deterministically chosen labels based on input payload + rules (seed derived from contents).
### 4. Matcher / Override System
- Provide dataclasses from the spec (`FileMatcher`, `SchemaMatcher`, `RequestMatcher`).
- Implement matcher evaluation order with `once=True` behavior to remove one-time overrides.
- Expose helper APIs:
- `fake.extract.stub_run(...)`
- `fake.parse.stub_parse(...)`
- `fake.classify.stub_prediction(...)`
- `fake.files.stub_upload(...)`, etc.
- Complexity: Need to ensure matcher evaluation can inspect raw `httpx.Request` bodies/headers for both sync and async flows without consuming the stream twice.
### 5. Assertions & Spies
- Expose convenience attributes pointing to `respx.Route` objects for frequently used paths (e.g., `fake.extract.stateless_run`).
- Provide helper methods for call counts, captured requests, etc.
- Ensure naming stays stable to avoid brittle tests.
### 6. Testing Strategy
- Add pytest fixtures to install the fake server globally for integration tests.
- Cover scenarios:
- Stateless extract returns deterministic payload.
- Agent-backed extract polls job/runs.
- Files API handles presigned upload and retrieval.
- Parse job lifecycle for both success and failure.
- Classification job with deterministic label output.
- Matcher overrides injection & once-only behavior.
- Mixed namespace configurations (e.g., intercept extract only, let parse hit real network).
## Extra Complexity & Spec Concerns
- **Presigned URL scope**: Spec assumes presigned uploads can be intercepted the same way as SaaS APIs, but actual presigned URLs often point to AWS domains outside `base_urls`. Need a strategy (e.g., generate fake host names the SDK will call, or rewrite responses to use local URLs).
- **Async client coverage**: LlamaParse builds new `httpx.AsyncClient` objects; the specs install/uninstall story must ensure respx patches all clients, not just the global one.
- **Deterministic generators**: The spec outlines hashing inputs but doesnt define exact algorithms. Without mirroring production generator logic, fixtures might diverge. We may need to document any intentional differences.
- **Job state timelines**: The spec expects transitions (`PENDING → SUCCESS`) with realistic timestamps. Need to ensure we schedule async updates or respond with multi-step polling; otherwise, tests relying on delays may behave differently.
- **Namespace toggling**: Clarify behavior when a namespace is disabled—should unmatched routes fall through automatically or raise? Current spec implies fall-through to real network, but that could be surprising in CI.
- **Schema handling**: `_validate_schema` currently calls production for dict schemas. The fake must emulate validation; otherwise tests will still hit SaaS. Spec doesnt detail validation logic, so we must decide on a simplified validator or deterministic echo.
- **Parse partitioning**: `LlamaParse` can partition jobs and expects consistent pagination semantics. Need to ensure deterministic results respect `target_pages`, `partition_pages`, etc., or document limitations.
## Next Actions
1. Prototype router + lifecycle with namespace toggles.
2. Implement FileStore + presigned workflow since other namespaces depend on file IDs.
3. Build deterministic generators and stores for extract/parse/classify.
4. Layer matcher/override system on top of the stores.
5. Write initial tests per namespace and refine spec gaps (presigned host, validation behavior).
6. Update `testing_utils_spec.md` with any clarifications discovered above.
+199
View File
@@ -0,0 +1,199 @@
# Testing Utils Research
## Objective
The `FakeLlamaCloudServer` proposal aims to intercept raw HTTP traffic for extract, parse, classify, and files APIs so local tests behave like SaaS without per-test stubbing. The mock must respect base URLs, support context-manager or long-lived install modes, and deterministically synthesize payloads from file + schema hashes while still allowing targeted overrides.
```1:67:py/testing_utils_spec.md
## Local Testing Utilities 2.0 (Spec Draft)
- **Everything mocked by default** …
- **Context manager optional** …
- **Pydantic-first ergonomics** …
- **API-only contract** …
```
## Python Hand-Written SDK Surfaces
These modules sit on top of the generated `llama_cloud` client and are what most application tests exercise. `FakeLlamaCloudServer` must satisfy the HTTP contracts they rely on.
### Extract flow (`py/llama_cloud_services/extract/extract.py`)
- The class imports resource types from `llama_cloud` and owns an `AsyncLlamaCloud` client shared across both stateless and agent-backed flows.
```17:37:py/llama_cloud_services/extract/extract.py
from llama_cloud import (
ExtractAgent as CloudExtractAgent,
ExtractConfig,
)
from llama_cloud.client import AsyncLlamaCloud
```
- Schema validation always calls `POST /api/v1/extraction/extraction-agents/schema/validation` through `client.llama_extract.validate_extraction_schema`, so the fake must mimic that endpoint.
```65:82:py/llama_cloud_services/extract/extract.py
async def _validate_schema(...):
validated_schema = await client.llama_extract.validate_extraction_schema(
data_schema=processed_schema
)
```
- Agent creation, listing, and run management use the `llama_extract` namespace. These calls surface the same request/response bodies as the SaaS API, so tests that interact through agents expect consistent metadata (IDs, status enums, etc.).
```636:738:py/llama_cloud_services/extract/extract.py
def create_agent(...):
agent = self._run_in_thread(
self._async_client.llama_extract.create_extraction_agent(
project_id=self._project_id,
)
)
return ExtractionAgent(...)
```
- Stateless extraction queues work by converting file inputs into either `file_id`, inline text, or base64 payloads and forwarding them to `POST /api/v1/extraction/run`. Deterministic responses from the fake should be keyed off `processed_schema` + whichever file representation the SDK sent.
```921:1018:py/llama_cloud_services/extract/extract.py
async def queue_extraction(...):
processed_schema = await _validate_schema(...)
job = await self._async_client.llama_extract.extract_stateless(
project_id=self._project_id,
organization_id=self._organization_id,
data_schema=processed_schema,
config=config,
**file_args,
)
```
### File uploads (`py/llama_cloud_services/files/client.py`)
- All higher-level services route file uploads/downloads through `FileClient`. When `use_presigned_url` is enabled, the client first calls `POST /api/v1/files` (generate URL), then performs the PUT upload directly, then fetches metadata via `GET /api/v1/files/{id}`. The fake must intercept both the API calls and the presigned PUT hops to return consistent `file_id`s and stored bytes.
```63:140:py/llama_cloud_services/files/client.py
presigned_url = await self.client.files.generate_presigned_url(...)
upload_response = await httpx_client.put(presigned_url.url, data=buffer.read())
return await self.client.files.get_file(presigned_url.file_id, …)
```
### Parse reader (`py/llama_cloud_services/parse/base.py`)
- `LlamaParse` is a bespoke reader that talks straight to HTTP routes defined in the module (e.g., `/api/parsing/upload`, `/api/parsing/job/{id}`). Unlike `LlamaExtract`, it does not go through the generated client; instead it builds URLs manually and uses `httpx`/`make_api_request`. A fake server must therefore implement these exact paths.
```49:66:py/llama_cloud_services/parse/base.py
JOB_RESULT_URL = "/api/parsing/job/{job_id}/result/{result_type}"
JOB_STATUS_ROUTE = "/api/parsing/job/{job_id}"
JOB_UPLOAD_ROUTE = "/api/parsing/upload"
```
```1056:1070:py/llama_cloud_services/parse/base.py
url = build_url(JOB_UPLOAD_ROUTE, self.organization_id, self.project_id)
resp = await make_api_request(self.aclient, "POST", url, …, files=files, data=data)
```
### Classifier beta client (`py/llama_cloud_services/beta/classifier/client.py`)
- Classification flows reuse `FileClient` for uploads and then call `AsyncLlamaCloud.classifier` endpoints (`create_classify_job`, `get_classify_job`, `get_classification_job_results`). Long-running tests poll until status becomes terminal, so mocking needs to cover both the enqueue POST and the follow-up GETs.
```75:151:py/llama_cloud_services/beta/classifier/client.py
return await self.client.classifier.create_classify_job(...)
results = await self.client.classifier.get_classification_job_results(
classify_job_with_status.id,
project_id=self.project_id,
)
```
## Generated Python Client
- The repo depends on the published `llama_cloud` package (currently 0.1.44) which is itself generated from the OpenAPI spec. All hand-written modules import types and service clients from this package, so the fake server may need to mirror whatever transport settings `AsyncLlamaCloud` expects (headers, pagination, etc.).
```1605:1608:py/uv.lock
sdist = { … "llama_cloud-0.1.44.tar.gz", … }
wheels = [{ … "llama_cloud-0.1.44-py3-none-any.whl", … }]
```
- Since `AsyncLlamaCloud` handles auth headers and base URLs, integrating the fake server typically means pointing `LLAMA_CLOUD_BASE_URL` at the mock and letting the generated client continue to build resource paths.
## TypeScript + OpenAPI Assets
- The canonical OpenAPI document lives in `ts/llama_cloud_services/openapi.json`. It defines every path/operation used by both the generated TypeScript SDK and the Python client. For example, the stateless extract endpoint is captured as `POST /api/v1/extraction/run`.
```13825:13872:ts/llama_cloud_services/openapi.json
"/api/v1/extraction/run": {
"post": {
"summary": "Extract Stateless",
"description": "… Requires data_schema, config, and either file_id, text, or base64 encoded file data.",
}
}
```
- The OpenAPI document is downloaded from production via `scripts/download.mjs` and then fed into `@hey-api/openapi-ts` to regenerate the TypeScript client and schema wrappers. Keeping the fake server in sync with the spec means you can diff regenerated clients when the API evolves.
```3:21:ts/llama_cloud_services/scripts/download.mjs
const response = await fetch('https://api.cloud.llamaindex.ai/api/openapi.json');
fs.writeFileSync('openapi.json', JSON.stringify(data, null, 2));
```
```1:24:ts/llama_cloud_services/openapi-ts.config.ts
export default defineConfig({
input: "./openapi.json",
output: { path: "./src/client", format: "prettier", lint: "eslint" },
plugins: [ … "@hey-api/sdk", "@hey-api/typescript" ],
});
```
- The public TypeScript surface (`src/LlamaClassify.ts`, etc.) already consumes the generated client by injecting auth headers and delegating to `classify(...)`. If Python tests eventually need to reuse the same fake server, TypeScript examples provide another reference for how SDK consumers expect responses to look.
```12:74:ts/llama_cloud_services/src/LlamaClassify.ts
export class LlamaClassify {
constructor(apiKey?: string, baseUrl?: string, region?: string) {
this.client = createClient(createConfig({ baseUrl: url, headers: { Authorization: `Bearer ${key}` }}));
}
async classify(rules, configuration, { fileContents, filePaths, projectId, … }) {
const result = await classify(rules, configuration, {
fileContents,
filePaths,
projectId: projectId ?? undefined,
client: this.client,
});
return result;
}
}
```
## Endpoint Map to Stub First
Cross-referencing the Python call-sites with the OpenAPI spec yields the minimum set of HTTP routes the fake server must implement:
1. `/api/v1/extraction/run` for stateless jobs, plus the agent CRUD endpoints under `/api/v1/extraction/extraction-agents`, `/api/v1/extraction/jobs`, and `/api/v1/extraction/runs/by-job/{id}` (see `LlamaExtract` usage above).
2. `/api/v1/files/**` for upload/generate-presigned/list/get/delete, plus any presigned `PUT` destinations (`FileClient`).
3. `/api/parsing/upload`, `/api/parsing/job/{job_id}`, `/api/parsing/job/{job_id}/result/{result_type}` (direct HTTPX calls in `LlamaParse`).
4. `/api/v1/classifier/**` for job creation, polling, and result retrieval (`LlamaClassify` and `ClassifyClient`).
Having deterministic handlers for these routes unlocks end-to-end coverage of extract/parse/classify flows without touching live SaaS.
## Implementation Reminders from the Spec
- Namespace toggles let tests intercept a subset of APIs while letting others fall through—mirror this by allowing `FakeLlamaCloudServer(namespaces=[...])` to selectively register respx routes.
- Deterministic payloads should derive from uploaded file bytes + schema hashes for extract, layout characteristics for parse, and label sets for classify so that rerunning the same test yields identical responses (reducing fixture churn).
- Keep the matcher system (`RequestMatcher`, `FileMatcher`, etc.) flexible so individual tests can stub failures (e.g., presigned upload errors, job timeouts) without reconfiguring the entire fake.
```44:210:py/testing_utils_spec.md
with FakeLlamaCloudServer() as fake:
extractor = LlamaExtract(...)
parser = LlamaParse(...)
classifier = LlamaClassify(...)
fake.extract.stub_run(... RequestMatcher ...)
```
Armed with the file map above, a new developer can trace any SDK call from the hand-written layers down to the generated client and the authoritative OpenAPI route, making it clear where the fake server needs to hook in.
+329
View File
@@ -0,0 +1,329 @@
## Local Testing Utilities 2.0 (Spec Draft)
Offline testing should feel identical to calling the public LlamaCloud API. The new utilities center a single `FakeLlamaCloudServer` that intercepts HTTP traffic at the API boundary, deterministically generates responses from real files + schemas, and only requires overrides when a test needs to exercise edge cases.
### Design goals
- **Everything mocked by default**: instantiating `FakeLlamaCloudServer()` wires up every public LlamaCloud namespace (extract, parse, classify, files, etc.) so the SDK behaves as if it were talking to production. Deterministic responses are returned without any additional wiring.
- **Context manager optional**: `FakeLlamaCloudServer` still supports `with ...` for pytest isolation, but you can call `install()` / `uninstall()` to keep the mock server active inside a long-running process (e.g., a FastAPI dev server that proxies to the fake).
- **Pydantic-first ergonomics**: all documentation and helpers assume schemas are declared as `BaseModel` subclasses. JSON Schema dictionaries are still accepted for compatibility.
- **API-only contract**: handlers talk raw HTTP (request dicts, status codes, JSON payloads) so we can reuse the mock in future SDKs or other languages without depending on `LlamaExtract`.
### Quick start (pytest-friendly, deterministic by default)
```python
from pathlib import Path
from pydantic import BaseModel, Field
from llama_cloud import ExtractConfig, ExtractMode
from llama_cloud_services.extract import LlamaExtract
from llama_cloud_services.testing_utils import FakeLlamaCloudServer
class Receipt(BaseModel):
merchant: str = Field(description="Vendor name")
total: float = Field(description="Grand total in USD")
config = ExtractConfig(extraction_mode=ExtractMode.FAST)
pdf_path = Path("tests/fixtures/receipt.pdf")
with FakeLlamaCloudServer() as fake:
extractor = LlamaExtract(
api_key="test-key",
verify=False,
)
run = extractor.extract(Receipt, config, pdf_path)
assert run.status.value == "SUCCESS"
assert run.data["total"] > 0 # generated entirely from file + schema
```
Key points:
- No manual stubbing required. The fake server hashes the uploaded file bytes + schema JSON to derive a deterministic seed and walks the schema to produce stable mock data.
- `FakeLlamaCloudServer` automatically intercepts the default SaaS URL (`https://api.cloud.llamaindex.ai`). If your tests point at another host (e.g., BYOC), pass it via `FakeLlamaCloudServer(base_urls=["https://byoc.dev/api"])`; otherwise, keep using your normal SDK base URL.
### Works across extract, parse, classify
```python
from llama_cloud_services.testing_utils import FakeLlamaCloudServer
from llama_cloud_services.extract import LlamaExtract
from llama_cloud_services.parse import LlamaParse
from llama_cloud_services.classify import LlamaClassify
with FakeLlamaCloudServer() as fake:
extractor = LlamaExtract(api_key="test-key")
parser = LlamaParse(api_key="test-key")
classifier = LlamaClassify(api_key="test-key")
run = extractor.extract(
Receipt, config, "noisebridge.pdf"
) # reuse quick-start schema/config
parse_result = parser.parse("noisebridge.pdf")
classification = classifier.classify({"text": "foo"})
assert run.status.value == "SUCCESS"
assert parse_result.documents[0].text # deterministically generated
assert classification.prediction in {
"A",
"B",
} # stable RNG driven by payload
```
Every namespace uses its own deterministic generator (schema-driven for extract, layout-driven for parse, label-driven for classify) but shares the same matcher/override system described below.
### Limiting intercepted APIs
If you only need a subset of APIs (e.g., extract + files during early bring-up), pass `namespaces` explicitly. Anything omitted will fall through to the real network, which is handy for hybrid tests.
```python
fake = FakeLlamaCloudServer(
namespaces=["extract", "files"],
base_urls=["https://api.cloud.llamaindex.ai"], # optionally point at BYOC
)
with fake:
extractor = LlamaExtract(api_key="test-key")
extractor.extract(Receipt, config, "noisebridge.pdf")
```
### Long-lived install for iterative development
```python
import os
from contextlib import asynccontextmanager
from fastapi import FastAPI
fake = FakeLlamaCloudServer(
namespaces=["extract"],
base_urls=[
os.environ.get(
"LLAMA_CLOUD_BASE_URL", "https://api.cloud.llamaindex.ai"
)
],
)
@asynccontextmanager
async def lifespan(app):
fake.install()
app.state.extractor = LlamaExtract(
api_key="dev",
verify=False,
)
try:
yield
finally:
fake.uninstall()
app = FastAPI(lifespan=lifespan)
```
`install()`/`uninstall()` simply wrap the respx router lifecycle so you can keep the mock server hot for REPLs, background workers, or manual QA sessions without relying on a context manager.
### Files API behavior
`FakeLlamaCloudServer` ships with a first-class fake for `/api/v1/files/*` because uploads sit on the critical path for both extract and downstream workflows that pre-stage files.
- Every call to `POST /files/generate-presigned-url`, the subsequent `PUT` upload, and the follow-up `GET /files/{file_id}` is intercepted and stored in-memory. The response objects mirror the real API so `FileClient` keeps working unchanged.
- `fake.files.preload(path="tests/fixtures/plan.pdf", filename="plan.pdf")` ingests local fixtures ahead of time and returns a reusable `file_id`, which is useful when tests pass `SourceText(file_id=...)`.
- `fake.files.stub_upload(...)` lets you simulate storage failures (e.g., 413 "file too large") using the same matcher system as extract.
- You can download what the SDK uploaded via `fake.files.read(file_id)` to assert on the bytes or to feed downstream mocks.
```python
from llama_cloud_services.extract import SourceText
with FakeLlamaCloudServer() as fake:
file_id = fake.files.preload(path="tests/fixtures/noisebridge.pdf")
extractor = LlamaExtract(api_key="test-key")
run = extractor.extract(Receipt, config, SourceText(file_id=file_id))
assert fake.files.read(file_id).startswith(b"%PDF")
```
### Deterministic response generation
1. Files uploaded via `/api/v1/files` (or inlined via `extract_stateless`) are fingerprinted using SHA256 (file content bytes + filename).
2. Schemas are normalized (Pydantic `model_json_schema()` plus sorted keys) and hashed.
3. A seed derived from `sha256(file_fingerprint + schema_digest)` feeds a tiny RNG that walks the schema to synthesize values (numbers, strings, arrays) while respecting field metadata (descriptions hint names, numeric ranges, etc.).
4. Runs transition through the same states as production (`PENDING``SUCCESS`) and return realistic timestamps, metadata, and config echoes.
Because the seed is stable, rerunning the same schema/file pair yields identical mock payloads without stubbing.
### Stubbing, spying, and assertions
Most tests only need the deterministic defaults, but the fake server provides a layered set of helpers for overriding responses, asserting call counts, and finally dropping down to raw `respx` when necessary.
#### Matcher API
```python
from dataclasses import dataclass
from typing import Callable, Optional
from httpx import Request
@dataclass
class FileMatcher:
filename: str | None = None
sha256: str | None = None
file_id: str | None = None
@dataclass
class SchemaMatcher:
model: type[BaseModel] | None = None
schema_hash: str | None = None
@dataclass
class RequestMatcher:
file: FileMatcher | Callable[[Request], bool] | None = None
schema: SchemaMatcher | None = None
agent_id: str | None = None
project_id: str | None = None
organization_id: str | None = None
predicate: Callable[[Request], bool] | None = None
```
Every callable matcher receives the raw `httpx.Request` object that respx captured, so you can inspect headers, cookies, bodies, etc., without learning another wrapper type. The helper dataclasses (`FileMatcher`, `SchemaMatcher`) just cover the common cases; mix and match as needed. Stubs are evaluated in registration order, and `once=True` removes the stub after the first match.
#### Stateless extraction example
```python
fake.extract.stub_run(
matcher=RequestMatcher(file=FileMatcher(filename="noisebridge.pdf")),
data={"merchant": "Noisebridge", "total": 42.0},
status="SUCCESS", # defaults to deterministic timeline when omitted
metadata={"source": "unit-test"},
once=True,
)
run = extractor.extract(Receipt, config, "noisebridge.pdf")
assert run.data["merchant"] == "Noisebridge"
```
`data` accepts dictionaries, Pydantic models, or callables (`Callable[[Request], dict]`). If you omit `status`, the stub only replaces the payload while preserving the deterministic job/run lifecycle.
#### Agent extraction example
```python
agent = extractor.create_agent(
name="tests", data_schema=Receipt, config=config
)
fake.extract.stub_agent_run(
agent_id=agent.id,
matcher=RequestMatcher(file=FileMatcher(filename="bad.pdf")),
job_status="FAILED", # overrides POST /extraction/jobs
run_status="FAILED", # overrides GET /runs/by-job
error={"message": "Schema mismatch"},
)
with pytest.raises(ApiError):
agent.extract("bad.pdf")
```
`stub_agent_run` targets the stateful job pipeline (`/extraction/jobs`, `/extraction/jobs/{id}`, `/extraction/runs/by-job/{id}`) so you can mimic long-running failures, retries, or partial completions without hand-writing multiple HTTP handlers.
#### Parse, classify, and files stubs
- `fake.parse.stub_parse(...)` lets you override document splits, token counts, or even return structured HTML for specific file IDs.
- `fake.classify.stub_prediction(...)` accepts label sets and score distributions so you can test downstream logic that inspects confidences.
- `fake.files.stub_upload(...)` / `fake.files.stub_download(...)` simulate storage edge cases such as timeouts or corrupted content.
Because every namespace uses the same matcher primitives, you can coordinate multi-API scenarios (e.g., stub the file upload and the subsequent extract run) without duplicating predicates.
#### Assertions without extra abstractions
Since everything runs through the same `respx.MockRouter` you already use elsewhere, assertions stay lightweight:
```python
with FakeLlamaCloudServer() as fake:
route = fake.router["POST", "/api/v1/extraction/run"]
extractor = LlamaExtract(api_key="test-key")
extractor.extract(Receipt, config, "noisebridge.pdf")
assert route.called
assert route.call_count == 1
req = route.calls[0].request # this is httpx.Request
assert req.headers["authorization"].startswith("Bearer ")
```
For friendlier names, every frequently used route is also pinned to a stable attribute:
- `fake.extract.stateless_run` (alias: `fake.extract_run`) → `POST /api/v1/extraction/run`
- `fake.extract.agent_job``POST /api/v1/extraction/jobs`
- `fake.extract.agent_run``GET /api/v1/extraction/runs/by-job/{job_id}`
- `fake.files.upload``POST /api/v1/files/upload` (or the presigned PUT hop, depending on mode)
- `fake.files.get``GET /api/v1/files/{file_id}`
Each attribute is the underlying `respx.Route`, so assertions feel natural:
```python
assert fake.extract.stateless_run.called
fake.extract.stateless_run.assert_called_once()
assert fake.files.upload.call_count == 1
assert fake.extract_run.called # global alias for the same route
```
If you ever need the full mapping, `fake.extract.routes["stateless_run"] is fake.extract.stateless_run`.
#### Advanced (respx-level) overrides
When you need total control, drop straight into respx:
```python
route = fake.router["POST", "/api/v1/extraction/run"]
route.mock(side_effect=lambda request: (418, {"detail": "I'm a teapot"}))
```
Or use the attribute shortcuts:
```python
fake.extract.stateless_run.mock(
side_effect=lambda request: (500, {"detail": "boom"})
)
```
Either way you're dealing with the canonical respx objects, so regex paths, call assertions, and other ecosystem tools keep working. The only convention is that handlers should return `(status_code, json_body | bytes)` so logging and deterministic fallbacks remain consistent.
### API-layer implementation hints
- Route decorators such as `server.add_handler("POST", "/api/v1/extraction/run")` install handlers for **every** registered base URL declared in the constructor, keeping the mock independent of SDK client classes.
- Request objects passed to handlers expose method, URL, headers, query params, JSON body, and raw bytes—everything needed to mirror production without importing internal models.
- Namespaces self-register via the constructor (e.g., `namespaces=["extract"]`) so no additional attach helpers are required; future SDKs can opt into the same HTTP contracts by toggling the namespaces they care about.
## Research: Extract SDK surface map
The current Python SDK (`py/llama_cloud_services/extract/extract.py`) is a thin wrapper over the HTTP API exposed in `ts/llama_cloud_services/openapi.json`. Understanding this mapping helps ensure the fake server mirrors the real contract.
### Core classes
- `LlamaExtract`: factory that owns an `AsyncLlamaCloud` client, manages thread pools, and exposes both stateless extraction (`extract`, `aextract`, `queue_extraction`) and agent CRUD helpers.
- `ExtractionAgent`: wraps an existing agent returned by the API and provides methods for queuing files, polling jobs, listing runs, updating schemas/configs, and deleting runs.
- `FileClient`: abstracts the `/api/v1/files` upload + download flow, including presigned URL handling for uploads.
### Stateless extraction flow
1. `LlamaExtract.queue_extraction(data_schema, config, files)` validates schemas via `POST /api/v1/extraction/extraction-agents/schema/validation`, converts input files into either `file_id`, `file` (base64 body), or inline `text`.
2. For each file the SDK calls `POST /api/v1/extraction/run` with the processed schema + config + file payload. The API responds with an `ExtractJob`.
3. `LlamaExtract.aextract` waits for completion by polling `_wait_for_job_result`, which hits `GET /api/v1/extraction/jobs/{job_id}` until the job is `SUCCESS`/`FAILED`, then fetches the run via `GET /api/v1/extraction/runs/by-job/{job_id}`. The synchronous `extract` just wraps this coroutine in a worker thread.
### Agent-backed flow
1. `create_agent` issues `POST /api/v1/extraction/extraction-agents` with name, schema, and config; responses seed `ExtractionAgent`.
2. `ExtractionAgent.queue_extraction` uploads files via `FileClient`, then enqueues jobs with `POST /api/v1/extraction/jobs` (or `/jobs/file` for multipart uploads). Returned job IDs are polled via `_wait_for_job_result` just like the stateless path.
3. `ExtractionAgent.list_extraction_runs` and `delete_extraction_run` map to `GET /api/v1/extraction/runs` (with pagination) and `DELETE /api/v1/extraction/runs/{run_id}` respectively.
4. Manual inspection helpers (`get_extraction_job`, `get_extraction_run_for_job`, `get_extraction_run`) call `GET /api/v1/extraction/jobs/{job_id}` and `GET /api/v1/extraction/runs/by-job/{job_id}` / `GET /api/v1/extraction/runs/{run_id}`.
### Files API touch points
- Uploads default to presigned URLs: the SDK first calls `POST /api/v1/files/generate-presigned-url`, then performs an HTTP PUT to the returned URL, and finally fetches the file metadata via `GET /api/v1/files/{file_id}`.
- When BYOC deployments disable presigned uploads, `FileClient` falls back to `POST /api/v1/files/upload`.
### Implications for the fake server
- **API-level parity**: mocking should happen at the HTTP layer (matching the endpoints listed above) so new SDKs can reuse the fake by simply pointing their base URL at it.
- **State surfaces**: to emulate production, the fake needs in-memory stores for files, jobs, and runs keyed by UUIDs, plus schema validation stubs that mimic the `/schema/validation` endpoint.
- **Deterministic generators**: since `ExtractRun.data` is derived from schema + file, implementing the generator once at the API layer ensures consistency across SDKs.
- **Error simulation hooks**: overrides should let us short-circuit any endpoint (jobs, runs, schema validation) without changing SDK code, mirroring how the real API might fail.
This map should serve as the checklist when we implement the mock: if an SDK method calls a certain path, our fake server must expose the same path with compatible request/response bodies so we can eventually lift these utilities into a standalone package.
## Implementation status
- `FakeLlamaCloudServer` now lives in `llama_cloud_services.testing_utils` and registers the files, extract, parse, and classify namespaces by default. It can be used as `with FakeLlamaCloudServer(): ...` or by calling `install()` / `uninstall()` explicitly.
- Deterministic payloads derive from file fingerprints plus schema hashes for extraction, seeded layout information for parse, and rule/file hashes for classification. The helper exposes matcher-driven overrides (`stub_run`, `stub_agent_run`, `files.stub_upload`) so tests can simulate errors.
- Common `respx.Route` handles are exposed via namespaces (`fake.extract.stateless_run`, `fake.extract.agent_job`, `fake.files.routes["upload"]`, etc.) for assertions that mirror the examples in this spec.
- End-to-end usage is covered in `py/unit_tests/testing_utils/test_fake_server.py`, which exercises stateless extraction, agent-backed uploads, and LlamaParse readers entirely against the fake server without external network calls.
+171
View File
@@ -0,0 +1,171 @@
import os
import tempfile
import pytest
import pandas as pd
from llama_cloud_services.beta.sheets import LlamaSheets
from llama_cloud_services.beta.sheets.types import SpreadsheetParsingConfig
@pytest.fixture
def sheets_client():
"""Create a LlamaSheets client for testing."""
api_key = os.getenv(
"LLAMA_CLOUD_API_KEY", "llx-3AEorIw5v0lnJPzEOI9xSl0N8yFx3fguw0Zn8QJHzGWmwg5r"
)
base_url = os.getenv("LLAMA_CLOUD_BASE_URL", "https://api.staging.llamaindex.ai")
client = LlamaSheets(
api_key=api_key,
base_url=base_url,
max_timeout=300,
poll_interval=2,
)
return client
@pytest.fixture
def sample_excel_file():
"""Create a temporary Excel file with sample data."""
# Create a simple dataframe with various data types
data = {
"Name": ["Alice", "Bob", "Charlie", "David", "Eve"],
"Age": [25, 30, 35, 40, 45],
"City": ["New York", "Los Angeles", "Chicago", "Houston", "Phoenix"],
"Salary": [50000.50, 75000.75, 100000.00, 125000.25, 150000.50],
}
df = pd.DataFrame(data)
# Create a temporary file
with tempfile.NamedTemporaryFile(suffix=".xlsx", delete=False) as tmp:
tmp_path = tmp.name
df.to_excel(tmp_path, index=False, sheet_name="TestSheet")
yield tmp_path
# Cleanup
try:
os.unlink(tmp_path)
except Exception:
pass
@pytest.mark.skipif(
os.environ.get(
"LLAMA_CLOUD_API_KEY", "llx-3AEorIw5v0lnJPzEOI9xSl0N8yFx3fguw0Zn8QJHzGWmwg5r"
)
== "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_spreadsheet_extraction_e2e(
sheets_client: LlamaSheets, sample_excel_file: str
):
"""End-to-end test for spreadsheet extraction.
This test:
1. Creates a temporary Excel file with sample data
2. Uploads and extracts tables from the file
3. Downloads the extracted table as a DataFrame
4. Verifies the extracted data matches the original data
"""
# Extract tables from the spreadsheet
result = await sheets_client.aextract_regions(sample_excel_file)
# Verify job completed successfully
assert result.status in ("SUCCESS", "PARTIAL_SUCCESS")
assert result.success is True
# Verify we extracted at least one table
assert len(result.regions) > 0, "Expected at least one table to be extracted"
# Get the first table
first_table = result.regions[0]
assert first_table.sheet_name == "TestSheet"
# Download the table as a DataFrame
extracted_df = await sheets_client.adownload_region_as_dataframe(
job_id=result.id,
region_id=first_table.region_id,
result_type=first_table.region_type,
)
# Load the original dataframe for comparison
original_df = pd.read_excel(sample_excel_file)
# Verify the extracted DataFrame has the expected shape
assert extracted_df.shape[0] == original_df.shape[0], (
f"Row count mismatch: extracted {extracted_df.shape[0]}, "
f"original {original_df.shape[0]}"
)
assert extracted_df.shape[1] == original_df.shape[1], (
f"Column count mismatch: extracted {extracted_df.shape[1]}, "
f"original {original_df.shape[1]}"
)
# Verify column names match
assert list(extracted_df.columns) == list(original_df.columns), (
f"Column names mismatch: extracted {list(extracted_df.columns)}, "
f"original {list(original_df.columns)}"
)
# Verify data types are preserved (at least numerically)
for col in original_df.columns:
if original_df[col].dtype in ["int64", "float64"]:
assert extracted_df[col].dtype in ["int64", "float64"], (
f"Column {col} type mismatch: extracted {extracted_df[col].dtype}, "
f"original {original_df[col].dtype}"
)
# Verify the data values match (allowing for minor type conversions)
for col in original_df.columns:
original_values = original_df[col].tolist()
extracted_values = extracted_df[col].tolist()
# Convert both to strings for comparison to handle type differences
original_str = [str(v) for v in original_values]
extracted_str = [str(v) for v in extracted_values]
assert original_str == extracted_str, (
f"Column {col} values mismatch:\n"
f"Original: {original_str}\n"
f"Extracted: {extracted_str}"
)
@pytest.mark.skipif(
os.environ.get(
"LLAMA_CLOUD_API_KEY", "llx-3AEorIw5v0lnJPzEOI9xSl0N8yFx3fguw0Zn8QJHzGWmwg5r"
)
== "",
reason="LLAMA_CLOUD_API_KEY not set",
)
@pytest.mark.asyncio
async def test_spreadsheet_extraction_with_config(
sheets_client: LlamaSheets, sample_excel_file: str
):
"""Test spreadsheet extraction with custom configuration."""
# Create a config with specific settings
config = SpreadsheetParsingConfig(
sheet_names=["TestSheet"],
include_hidden_cells=True,
generate_additional_metadata=True,
)
# Extract tables with the config
result = await sheets_client.aextract_regions(sample_excel_file, config=config)
# Verify job completed successfully
assert result.status in ("SUCCESS", "PARTIAL_SUCCESS")
assert result.success is True
# Verify that additional metadata was generated
assert len(result.worksheet_metadata) > 0
assert result.worksheet_metadata[0].title is not None
assert result.worksheet_metadata[0].description is not None
# Verify we extracted at least one table
assert len(result.regions) > 0
# Verify the sheet name matches
assert result.regions[0].sheet_name == "TestSheet"
-3
View File
@@ -44,7 +44,6 @@ def classify_client(
return ClassifyClient(
async_llama_cloud_client,
project_id=project.id,
organization_id=project.organization_id,
polling_interval=1,
)
@@ -56,7 +55,6 @@ def file_client(
return FileClient(
async_llama_cloud_client,
project_id=project.id,
organization_id=project.organization_id,
use_presigned_url=False,
)
@@ -148,7 +146,6 @@ async def test_classify_file_ids_from_api_key(
api_key=e2e_test_settings.LLAMA_CLOUD_API_KEY.get_secret_value(),
base_url=e2e_test_settings.LLAMA_CLOUD_BASE_URL,
project_id=pdf_file.project_id,
organization_id=e2e_test_settings.LLAMA_CLOUD_ORGANIZATION_ID,
)
# Classify the uploaded files
+2
View File
@@ -1,3 +1,5 @@
from __future__ import annotations
import pytest
from llama_index.core.constants import DEFAULT_BASE_URL
from pydantic import Field, SecretStr
+2
View File
@@ -58,6 +58,8 @@ def get_test_cases():
settings = [
ExtractConfig(extraction_mode=ExtractMode.FAST),
ExtractConfig(extraction_mode=ExtractMode.BALANCED),
ExtractConfig(extraction_mode=ExtractMode.MULTIMODAL),
ExtractConfig(extraction_mode=ExtractMode.PREMIUM),
]
for input_file in sorted(input_files):
+121 -2
View File
@@ -44,7 +44,7 @@ def index_name() -> Generator[str, None, None]:
client = LlamaCloud(token=api_key, base_url=base_url)
pipeline = client.pipelines.search_pipelines(project_name=name)
if pipeline:
client.pipelines.delete(pipeline_id=pipeline[0].id)
client.pipelines.delete_pipeline(pipeline_id=pipeline[0].id)
@pytest.fixture()
@@ -83,7 +83,7 @@ def _setup_index_with_file(
# add file to pipeline
pipeline_file_create = PipelineFileCreate(file_id=file.id)
client.pipelines.add_files_to_pipeline_api(
client.pipeline_files.add_files_to_pipeline_api(
pipeline_id=pipeline.id, request=[pipeline_file_create]
)
@@ -170,6 +170,43 @@ def test_upload_file(index_name: str):
os.remove(temp_file_path)
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
def test_upload_file_with_custom_metadata(index_name: str):
index = LlamaCloudIndex.create_index(
name=index_name,
project_name=project_name,
organization_id=organization_id,
api_key=api_key,
base_url=base_url,
)
# Create a temporary file to upload
with tempfile.NamedTemporaryFile(delete=False, suffix=".txt") as temp_file:
temp_file.write(b"Sample content for testing upload.")
temp_file_path = temp_file.name
custom_metadata = {"foo": "bar"}
try:
# Upload the file
file_id = index.upload_file(
temp_file_path, custom_metadata=custom_metadata, verbose=True
)
assert file_id is not None
# Verify the file is part of the index
docs = index.ref_doc_info
temp_file_name = os.path.basename(temp_file_path)
assert any(
temp_file_name == doc.metadata.get("file_name") for doc in docs.values()
)
finally:
# Clean up the temporary file
os.remove(temp_file_path)
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
@@ -196,6 +233,38 @@ def test_upload_file_from_url(remote_file: Tuple[str, str], index_name: str):
assert any(test_file_name == doc.metadata.get("file_name") for doc in docs.values())
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
def test_upload_file_from_url_with_custom_metadata(
remote_file: Tuple[str, str], index_name: str
):
index = LlamaCloudIndex.create_index(
name=index_name,
project_name=project_name,
organization_id=organization_id,
api_key=api_key,
base_url=base_url,
)
# Define a URL to a file for testing
custom_metadata = {"foo": "bar"}
test_file_url, test_file_name = remote_file
# Upload the file from the URL
file_id = index.upload_file_from_url(
file_name=test_file_name,
url=test_file_url,
custom_metadata=custom_metadata,
verbose=True,
)
assert file_id is not None
# Verify the file is part of the index
docs = index.ref_doc_info
assert any(test_file_name == doc.metadata.get("file_name") for doc in docs.values())
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
@@ -507,6 +576,33 @@ async def test_async_upload_file_from_url(
await index.await_for_completion()
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
@pytest.mark.asyncio
async def test_async_upload_file_from_url_with_custom_metadata(
remote_file: Tuple[str, str], index_name: str
):
index = await LlamaCloudIndex.acreate_index(
name=index_name,
project_name=project_name,
api_key=api_key,
base_url=base_url,
)
custom_metadata = {"foo": "bar"}
test_file_url, test_file_name = remote_file
file_id = await index.aupload_file_from_url(
file_name=test_file_name,
url=test_file_url,
custom_metadata=custom_metadata,
verbose=True,
)
assert file_id is not None
await index.await_for_completion()
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
@@ -525,6 +621,29 @@ async def test_async_index_from_file(index_name: str, local_file: str):
await index.await_for_completion()
@pytest.mark.skipif(
not base_url or not api_key, reason="No platform base url or api key set"
)
@pytest.mark.asyncio
async def test_async_index_from_file_with_custom_metadata(
index_name: str, local_file: str
):
index = await LlamaCloudIndex.acreate_index(
name=index_name,
project_name=project_name,
api_key=api_key,
base_url=base_url,
)
custom_metadata = {"foo": "bar"}
file_id = await index.aupload_file(
file_path=local_file, custom_metadata=custom_metadata, verbose=True
)
assert file_id is not None
await index.await_for_completion()
class DummySchema(BaseModel):
source: str
+34
View File
@@ -6,6 +6,40 @@ from llama_cloud_services import LlamaParse
from llama_cloud_services.parse.types import JobResult
def test_format_parse_result_markdown_for_notebook():
"""Test the _format_markdown_for_notebook function.
Right now, the only work it does is escape single dollar signs."""
result = JobResult(job_id="test", file_name="test.pdf", job_result={})
# Test None input
assert result._format_markdown_for_notebook(None) is None
# Test single dollar sign gets escaped
assert result._format_markdown_for_notebook("This costs $5") == "This costs \\$5"
# Test double dollar signs are preserved (LaTeX equations)
assert (
result._format_markdown_for_notebook("$$x^2 + y^2 = z^2$$")
== "$$x^2 + y^2 = z^2$$"
)
# Test mixed single and double dollar signs
text = "This costs $5, but $$E = mc^2$$ is priceless"
expected = "This costs \\$5, but $$E = mc^2$$ is priceless"
assert result._format_markdown_for_notebook(text) == expected
# Test multiple single dollar signs
assert result._format_markdown_for_notebook("$10 and $20") == "\\$10 and \\$20"
# Test three or more consecutive dollar signs (preserve them)
assert result._format_markdown_for_notebook("$$$") == "$$$"
# Test adjacent dollar signs with text in between
text = "$$inline$$ and $separate"
expected = "$$inline$$ and \\$separate"
assert result._format_markdown_for_notebook(text) == expected
@pytest.fixture
def file_path() -> str:
return "tests/test_files/attention_is_all_you_need.pdf"
@@ -2,6 +2,7 @@ from datetime import datetime
import json
from pathlib import Path
from typing import Any, Dict, Optional
import uuid
import pytest
from llama_cloud import ExtractRun, File
@@ -434,6 +435,7 @@ def create_extract_run(
"extraction_agent_id": "extraction-agent-123",
"config": {},
"status": "SUCCESS",
"project_id": str(uuid.uuid4()),
"from_ui": False,
}
)
+1
View File
@@ -112,5 +112,6 @@
"num_output_tokens": 3440
}
},
"project_id": "77bdc79f-fb69-49ae-a783-fcc573eec7ce",
"from_ui": false
}
@@ -0,0 +1,311 @@
from __future__ import annotations
from pathlib import Path
import pytest
from llama_cloud import ExtractConfig
from llama_cloud.types import ExtractMode
from llama_cloud.core.api_error import ApiError
from llama_cloud_services.extract import LlamaExtract
from llama_cloud_services.parse import LlamaParse
from llama_cloud_services.beta.agent_data import AsyncAgentDataClient
from llama_cloud_services.testing_utils import FakeLlamaCloudServer
from llama_cloud_services.testing_utils._deterministic import hash_schema
from pydantic import BaseModel, Field
class Receipt(BaseModel):
merchant: str = Field(description="Vendor name")
total: float = Field(description="Grand total")
@pytest.fixture(autouse=True)
def _fake_env(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("LLAMA_CLOUD_API_KEY", "unit-test-key")
monkeypatch.setenv("LLAMA_CLOUD_BASE_URL", FakeLlamaCloudServer.DEFAULT_BASE_URL)
@pytest.fixture
def fake_server() -> FakeLlamaCloudServer:
with FakeLlamaCloudServer() as server:
yield server
def _write_sample_file(tmp_path: Path, name: str, content: str) -> Path:
target = tmp_path / name
target.write_text(content)
return target
def test_stateless_extract_is_deterministic(
fake_server: FakeLlamaCloudServer, tmp_path: Path
) -> None:
extractor = LlamaExtract(api_key="unit-test-key", verify=False)
config = ExtractConfig(extraction_mode=ExtractMode.FAST)
sample_path = _write_sample_file(
tmp_path, "receipt.txt", "Merchant: Lunar Bistro\nTotal: 123.45"
)
first_run = extractor.extract(Receipt, config, sample_path)
second_run = extractor.extract(Receipt, config, sample_path)
assert first_run.status.value == "SUCCESS"
assert second_run.data == first_run.data
assert "merchant" in first_run.data
assert fake_server.extract.stateless_run.called
def test_agent_flow_uploads_and_processes_files(
fake_server: FakeLlamaCloudServer, tmp_path: Path
) -> None:
extractor = LlamaExtract(api_key="unit-test-key", verify=False)
config = ExtractConfig(extraction_mode=ExtractMode.FAST)
agent = extractor.create_agent(
name="unit-test-agent", data_schema=Receipt, config=config
)
sample_path = _write_sample_file(
tmp_path, "contract.pdf", "Agreement between parties."
)
run = agent.extract(sample_path)
assert run.status.value == "SUCCESS"
assert "merchant" in run.data
uploaded_bytes = fake_server.files.read(run.file.id)
assert uploaded_bytes.startswith(b"Agreement")
assert fake_server.extract.agent_job.called
assert fake_server.extract.agent_run.called
def test_parse_load_data_returns_documents(
fake_server: FakeLlamaCloudServer, tmp_path: Path
) -> None:
parser = LlamaParse(
api_key="unit-test-key", base_url=FakeLlamaCloudServer.DEFAULT_BASE_URL
)
sample_path = _write_sample_file(
tmp_path, "report.pdf", "Executive summary of quarterly goals."
)
documents = parser.load_data(sample_path)
assert documents
assert "(page 1)" in documents[0].text
@pytest.mark.asyncio
async def test_agent_data_create(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data = Receipt(merchant="Test Inc", total=1000)
item = await client.create_item(data)
assert item.id == hash_schema(data)[:7]
assert item.data.merchant == data.merchant and item.data.total == data.total
assert item.collection == "extracted_data"
assert item.deployment_name == "extraction_agent"
@pytest.mark.asyncio
async def test_agent_data_update(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data = Receipt(merchant="Test Inc", total=1000)
item = await client.create_item(data)
assert item.id is not None
updated_data = Receipt(merchant="Testing Inc", total=1100)
updated_item = await client.update_item(item_id=item.id, data=updated_data)
# ensure that the data actually changed
assert (
updated_item.data.merchant == updated_data.merchant
and updated_item.data.total == updated_data.total
)
# make sure nothing else changed
assert updated_item.id == item.id
assert updated_item.collection == item.collection
assert updated_item.deployment_name == item.deployment_name
@pytest.mark.asyncio
async def test_agent_data_search(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data1 = Receipt(merchant="Test Inc", total=1000)
data2 = Receipt(merchant="Test Inc", total=1300)
data3 = Receipt(merchant="Testing Inc", total=1100)
item1 = await client.create_item(data1)
item2 = await client.create_item(data2)
item3 = await client.create_item(data3)
result = await client.search(filter={"merchant": {"eq": "Test Inc"}})
assert result.total == 2
assert any(item.id == item1.id for item in result.items) and any(
item.id == item2.id for item in result.items
)
assert all(item.data.merchant == "Test Inc" for item in result.items)
result1 = await client.search(filter={"total": {"lt": 1200}})
assert result.total == 2
assert any(item.id == item1.id for item in result1.items) and any(
item.id == item3.id for item in result1.items
)
assert all(item.data.total < 1200 for item in result1.items)
@pytest.mark.asyncio
async def test_agent_data_aggregate(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data1 = Receipt(merchant="Test Inc", total=1000)
data2 = Receipt(merchant="Test Inc", total=1300)
data3 = Receipt(merchant="Testing Inc", total=1100)
await client.create_item(data1)
await client.create_item(data2)
await client.create_item(data3)
result = await client.aggregate(
filter={"merchant": {"eq": "Test Inc"}},
group_by=["merchant"],
count=True,
)
# filtering for 'Test Inc' on merchant means that only data with 'Test Inc' are left, meaning that there is only one group of data for merchant, i.e. the 'Test Inc' group
assert len(result.items) == 1
assert result.items[0].count == 2
assert result.items[0].first_item is not None
assert result.items[0].first_item.merchant == data1.merchant
assert result.items[0].first_item.total == data1.total
assert result.items[0].group_key == {"merchant": "Test Inc"}
result = await client.aggregate(
group_by=["merchant"],
count=True,
)
assert len(result.items) == 2
assert result.items[0].count == 2
assert result.items[0].first_item is not None
assert result.items[0].first_item.merchant == data1.merchant
assert result.items[0].first_item.total == data1.total
assert result.items[0].group_key == {"merchant": "Test Inc"}
assert result.items[1].count == 1
assert result.items[1].first_item is not None
assert result.items[1].first_item.merchant == data3.merchant
assert result.items[1].first_item.total == data3.total
assert result.items[1].group_key == {"merchant": "Testing Inc"}
@pytest.mark.asyncio
async def test_agent_data_get(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data1 = Receipt(merchant="Test Inc", total=1000)
data2 = Receipt(merchant="Test Inc", total=1300)
item1 = await client.create_item(data1)
assert item1.id is not None
item2 = await client.create_item(data2)
assert item2.id is not None
item = await client.get_item(item1.id)
assert item.collection == item1.collection
assert item.deployment_name == item1.deployment_name
assert item.data.merchant == data1.merchant
assert item.data.total == data1.total
# using this pattern instead of `with pytest.raise` for more granual control over the error itself
try:
notitem = await client.get_item(item2.id + "thisdoesnotexist")
e = None
except ApiError as err:
e = err
notitem = None
assert notitem is None
assert e is not None
assert e.status_code == 404
assert e.body == {"detail": f"No data with ID: {item2.id+'thisdoesnotexist'}"}
@pytest.mark.asyncio
async def test_agent_data_delete_by_id(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data = Receipt(merchant="Test Inc", total=1300)
item = await client.create_item(data)
assert item.id is not None
await client.delete_item(item.id)
# using this pattern instead of `with pytest.raise` for more granual control over the error itself
try:
notitem = await client.get_item(item.id)
e = None
except ApiError as err:
e = err
notitem = None
assert notitem is None
assert e is not None
assert e.status_code == 404
assert e.body == {"detail": f"No data with ID: {item.id}"}
# using this pattern instead of `with pytest.raise` for more granual control over the error itself
try:
await client.delete_item(item.id)
e = None
except ApiError as err:
e = err
assert e is not None
assert e.status_code == 404
assert e.body == {"detail": f"No data with ID: {item.id}"}
@pytest.mark.asyncio
async def test_agent_data_delete_by_query(fake_server: FakeLlamaCloudServer):
with fake_server as _:
client = AsyncAgentDataClient(
Receipt,
collection="extracted_data",
deployment_name="extraction_agent",
token="fake-api-key",
)
data1 = Receipt(merchant="Test Inc", total=1000)
data2 = Receipt(merchant="Test Inc", total=1300)
data3 = Receipt(merchant="Testing Inc", total=1100)
item1 = await client.create_item(data1)
item2 = await client.create_item(data2)
item3 = await client.create_item(data3)
result = await client.delete(filter={"merchant": {"eq": "Test Inc"}})
assert result == 2
for item in (item1, item2):
assert item.id is not None
try:
notitem = await client.get_item(item.id)
e = None
except ApiError as err:
e = err
notitem = None
assert notitem is None
assert e is not None
assert e.status_code == 404
assert e.body == {"detail": f"No data with ID: {item.id}"}
assert item3.id is not None
itemfound = await client.get_item(item3.id)
assert itemfound.id == item3.id
Generated
+265 -13
View File
@@ -1,9 +1,10 @@
version = 1
revision = 2
revision = 3
requires-python = ">=3.9, <4.0"
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
"python_full_version == '3.10.*'",
"python_full_version < '3.10'",
]
@@ -220,7 +221,8 @@ name = "argon2-cffi-bindings"
version = "25.1.0"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
"python_full_version == '3.10.*'",
"python_full_version < '3.10'",
]
@@ -589,7 +591,8 @@ version = "8.2.1"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
"python_full_version == '3.10.*'",
]
dependencies = [
@@ -720,6 +723,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/33/6b/e0547afaf41bf2c42e52430072fa5658766e3d65bd4b03a563d1b6336f57/distlib-0.4.0-py2.py3-none-any.whl", hash = "sha256:9659f7d87e46584a30b5780e43ac7a2143098441670ff0a49d5f9034c54a6c16", size = 469047, upload-time = "2025-07-17T16:51:58.613Z" },
]
[[package]]
name = "et-xmlfile"
version = "2.0.0"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/d3/38/af70d7ab1ae9d4da450eeec1fa3918940a5fafb9055e934af8d6eb0c2313/et_xmlfile-2.0.0.tar.gz", hash = "sha256:dab3f4764309081ce75662649be815c4c9081e88f0837825f90fd28317d4da54", size = 17234, upload-time = "2024-10-25T17:25:40.039Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/c1/8b/5fe2cc11fee489817272089c4203e679c63b570a5aaeb18d852ae3cbba6a/et_xmlfile-2.0.0-py3-none-any.whl", hash = "sha256:7a91720bc756843502c3b7504c77b8fe44217c85c537d85037f0f536151b2caa", size = 18059, upload-time = "2024-10-25T17:25:39.051Z" },
]
[[package]]
name = "eval-type-backport"
version = "0.2.2"
@@ -1120,7 +1132,8 @@ version = "8.37.0"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
"python_full_version == '3.10.*'",
]
dependencies = [
@@ -1582,21 +1595,21 @@ wheels = [
[[package]]
name = "llama-cloud"
version = "0.1.43"
version = "0.1.44"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "certifi" },
{ name = "httpx" },
{ name = "pydantic" },
]
sdist = { url = "https://files.pythonhosted.org/packages/9b/33/33a8bd3a617c071caf450ca2627969f8b28272d0692f122997c10a32247e/llama_cloud-0.1.43.tar.gz", hash = "sha256:00429f05aea515449d90cde91ef3ed3687fcd93e46f6246d08cbea02f9b397a9", size = 112992, upload-time = "2025-10-02T21:55:38.355Z" }
sdist = { url = "https://files.pythonhosted.org/packages/54/eb/16e31fb0fc4df91b08fa19cc3f28ac6e3c7d4df0bcbb71dd2bf596e9586f/llama_cloud-0.1.44.tar.gz", hash = "sha256:276a2b4f94463da037431ca3063331b3b6be398bbfb003113ee76b7c2a873b53", size = 120502, upload-time = "2025-11-04T00:51:58.578Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2b/54/559a67542396d5660a71115b29e0160e9dd784e570e1f4ef55ad22bf5b39/llama_cloud-0.1.43-py3-none-any.whl", hash = "sha256:540605d4dd13c6536a3b75cd4d04b211f29b16d17faee9381e3793a651f1dec1", size = 311460, upload-time = "2025-10-02T21:55:37.282Z" },
{ url = "https://files.pythonhosted.org/packages/69/0a/fabe54c21d5927d626550cb9560a20e51e42468355f5f0fb300f84806e28/llama_cloud-0.1.44-py3-none-any.whl", hash = "sha256:dfdcc4932353711fc8639f14261cbb54a88139b7790ebdd3ed4fde29bbbc0b88", size = 332779, upload-time = "2025-11-04T00:51:57.371Z" },
]
[[package]]
name = "llama-cloud-services"
version = "0.6.69"
version = "0.6.81"
source = { editable = "." }
dependencies = [
{ name = "click", version = "8.1.8", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.10'" },
@@ -1608,6 +1621,7 @@ dependencies = [
{ name = "platformdirs" },
{ name = "pydantic" },
{ name = "python-dotenv" },
{ name = "respx" },
{ name = "tenacity" },
]
@@ -1620,7 +1634,11 @@ dev = [
{ name = "ipython", version = "8.37.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.10'" },
{ name = "jupyter" },
{ name = "mypy" },
{ name = "openpyxl" },
{ name = "pandas" },
{ name = "pre-commit" },
{ name = "pyarrow", version = "21.0.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.10'" },
{ name = "pyarrow", version = "22.0.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.10'" },
{ name = "pydantic-settings" },
{ name = "pytest" },
{ name = "pytest-asyncio" },
@@ -1631,12 +1649,13 @@ dev = [
requires-dist = [
{ name = "click", specifier = ">=8.1.7,<9" },
{ name = "eval-type-backport", marker = "python_full_version < '3.10'", specifier = ">=0.2.0,<0.3" },
{ name = "llama-cloud", specifier = "==0.1.43" },
{ name = "llama-cloud", specifier = "==0.1.44" },
{ name = "llama-index-core", specifier = ">=0.12.0" },
{ name = "packaging", specifier = ">=25.0" },
{ name = "packaging", specifier = ">=23.0" },
{ name = "platformdirs", specifier = ">=4.3.7,<5" },
{ name = "pydantic", specifier = ">=2.8,!=2.10" },
{ name = "python-dotenv", specifier = ">=1.0.1,<2" },
{ name = "respx", extras = ["tests"], specifier = ">=0.22.0" },
{ name = "tenacity", specifier = ">=8.5.0,<10.0" },
]
@@ -1648,7 +1667,10 @@ dev = [
{ name = "ipython", specifier = ">=8.12.3,<9" },
{ name = "jupyter", specifier = ">=1.1.1,<2" },
{ name = "mypy", specifier = ">=1.14.1,<2" },
{ name = "openpyxl" },
{ name = "pandas" },
{ name = "pre-commit", specifier = "==3.2.0" },
{ name = "pyarrow" },
{ name = "pydantic-settings", specifier = ">=2.10.1" },
{ name = "pytest", specifier = ">=8.0.0,<9" },
{ name = "pytest-asyncio" },
@@ -2098,7 +2120,8 @@ version = "3.5"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
]
sdist = { url = "https://files.pythonhosted.org/packages/6c/4f/ccdb8ad3a38e583f214547fd2f7ff1fc160c43a75af88e6aec213404b96a/networkx-3.5.tar.gz", hash = "sha256:d4c6f9cf81f52d69230866796b82afbccdec3db7ae4fbd1b65ea750feed50037", size = 2471065, upload-time = "2025-05-29T11:35:07.804Z" }
wheels = [
@@ -2284,7 +2307,8 @@ version = "2.3.2"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.11' and python_full_version < '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
]
sdist = { url = "https://files.pythonhosted.org/packages/37/7d/3fec4199c5ffb892bed55cff901e4f39a58c81df9c44c280499e92cad264/numpy-2.3.2.tar.gz", hash = "sha256:e0486a11ec30cdecb53f184d496d1c6a20786c81e55e41640270130056f8ee48", size = 20489306, upload-time = "2025-07-24T21:32:07.553Z" }
wheels = [
@@ -2363,6 +2387,18 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/78/e3/6690b3f85a05506733c7e90b577e4762517404ea78bab2ca3a5cb1aeb78d/numpy-2.3.2-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:6936aff90dda378c09bea075af0d9c675fe3a977a9d2402f95a87f440f59f619", size = 12977811, upload-time = "2025-07-24T21:29:18.234Z" },
]
[[package]]
name = "openpyxl"
version = "3.1.5"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "et-xmlfile" },
]
sdist = { url = "https://files.pythonhosted.org/packages/3d/f9/88d94a75de065ea32619465d2f77b29a0469500e99012523b91cc4141cd1/openpyxl-3.1.5.tar.gz", hash = "sha256:cf0e3cf56142039133628b5acffe8ef0c12bc902d2aadd3e0fe5878dc08d1050", size = 186464, upload-time = "2024-06-28T14:03:44.161Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/c0/da/977ded879c29cbd04de313843e76868e6e13408a94ed6b987245dc7c8506/openpyxl-3.1.5-py2.py3-none-any.whl", hash = "sha256:5282c12b107bffeef825f4617dc029afaf41d0ea60823bbb665ef3079dc79de2", size = 250910, upload-time = "2024-06-28T14:03:41.161Z" },
]
[[package]]
name = "orderly-set"
version = "5.5.0"
@@ -2390,6 +2426,76 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/20/12/38679034af332785aac8774540895e234f4d07f7545804097de4b666afd8/packaging-25.0-py3-none-any.whl", hash = "sha256:29572ef2b1f17581046b3a2227d5c611fb25ec70ca1ba8554b24b0e69331a484", size = 66469, upload-time = "2025-04-19T11:48:57.875Z" },
]
[[package]]
name = "pandas"
version = "2.3.3"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "numpy", version = "2.0.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.10'" },
{ name = "numpy", version = "2.2.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version == '3.10.*'" },
{ name = "numpy", version = "2.3.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" },
{ name = "python-dateutil" },
{ name = "pytz" },
{ name = "tzdata" },
]
sdist = { url = "https://files.pythonhosted.org/packages/33/01/d40b85317f86cf08d853a4f495195c73815fdf205eef3993821720274518/pandas-2.3.3.tar.gz", hash = "sha256:e05e1af93b977f7eafa636d043f9f94c7ee3ac81af99c13508215942e64c993b", size = 4495223, upload-time = "2025-09-29T23:34:51.853Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/3d/f7/f425a00df4fcc22b292c6895c6831c0c8ae1d9fac1e024d16f98a9ce8749/pandas-2.3.3-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:376c6446ae31770764215a6c937f72d917f214b43560603cd60da6408f183b6c", size = 11555763, upload-time = "2025-09-29T23:16:53.287Z" },
{ url = "https://files.pythonhosted.org/packages/13/4f/66d99628ff8ce7857aca52fed8f0066ce209f96be2fede6cef9f84e8d04f/pandas-2.3.3-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:e19d192383eab2f4ceb30b412b22ea30690c9e618f78870357ae1d682912015a", size = 10801217, upload-time = "2025-09-29T23:17:04.522Z" },
{ url = "https://files.pythonhosted.org/packages/1d/03/3fc4a529a7710f890a239cc496fc6d50ad4a0995657dccc1d64695adb9f4/pandas-2.3.3-cp310-cp310-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5caf26f64126b6c7aec964f74266f435afef1c1b13da3b0636c7518a1fa3e2b1", size = 12148791, upload-time = "2025-09-29T23:17:18.444Z" },
{ url = "https://files.pythonhosted.org/packages/40/a8/4dac1f8f8235e5d25b9955d02ff6f29396191d4e665d71122c3722ca83c5/pandas-2.3.3-cp310-cp310-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:dd7478f1463441ae4ca7308a70e90b33470fa593429f9d4c578dd00d1fa78838", size = 12769373, upload-time = "2025-09-29T23:17:35.846Z" },
{ url = "https://files.pythonhosted.org/packages/df/91/82cc5169b6b25440a7fc0ef3a694582418d875c8e3ebf796a6d6470aa578/pandas-2.3.3-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:4793891684806ae50d1288c9bae9330293ab4e083ccd1c5e383c34549c6e4250", size = 13200444, upload-time = "2025-09-29T23:17:49.341Z" },
{ url = "https://files.pythonhosted.org/packages/10/ae/89b3283800ab58f7af2952704078555fa60c807fff764395bb57ea0b0dbd/pandas-2.3.3-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:28083c648d9a99a5dd035ec125d42439c6c1c525098c58af0fc38dd1a7a1b3d4", size = 13858459, upload-time = "2025-09-29T23:18:03.722Z" },
{ url = "https://files.pythonhosted.org/packages/85/72/530900610650f54a35a19476eca5104f38555afccda1aa11a92ee14cb21d/pandas-2.3.3-cp310-cp310-win_amd64.whl", hash = "sha256:503cf027cf9940d2ceaa1a93cfb5f8c8c7e6e90720a2850378f0b3f3b1e06826", size = 11346086, upload-time = "2025-09-29T23:18:18.505Z" },
{ url = "https://files.pythonhosted.org/packages/c1/fa/7ac648108144a095b4fb6aa3de1954689f7af60a14cf25583f4960ecb878/pandas-2.3.3-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:602b8615ebcc4a0c1751e71840428ddebeb142ec02c786e8ad6b1ce3c8dec523", size = 11578790, upload-time = "2025-09-29T23:18:30.065Z" },
{ url = "https://files.pythonhosted.org/packages/9b/35/74442388c6cf008882d4d4bdfc4109be87e9b8b7ccd097ad1e7f006e2e95/pandas-2.3.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:8fe25fc7b623b0ef6b5009149627e34d2a4657e880948ec3c840e9402e5c1b45", size = 10833831, upload-time = "2025-09-29T23:38:56.071Z" },
{ url = "https://files.pythonhosted.org/packages/fe/e4/de154cbfeee13383ad58d23017da99390b91d73f8c11856f2095e813201b/pandas-2.3.3-cp311-cp311-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b468d3dad6ff947df92dcb32ede5b7bd41a9b3cceef0a30ed925f6d01fb8fa66", size = 12199267, upload-time = "2025-09-29T23:18:41.627Z" },
{ url = "https://files.pythonhosted.org/packages/bf/c9/63f8d545568d9ab91476b1818b4741f521646cbdd151c6efebf40d6de6f7/pandas-2.3.3-cp311-cp311-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b98560e98cb334799c0b07ca7967ac361a47326e9b4e5a7dfb5ab2b1c9d35a1b", size = 12789281, upload-time = "2025-09-29T23:18:56.834Z" },
{ url = "https://files.pythonhosted.org/packages/f2/00/a5ac8c7a0e67fd1a6059e40aa08fa1c52cc00709077d2300e210c3ce0322/pandas-2.3.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:1d37b5848ba49824e5c30bedb9c830ab9b7751fd049bc7914533e01c65f79791", size = 13240453, upload-time = "2025-09-29T23:19:09.247Z" },
{ url = "https://files.pythonhosted.org/packages/27/4d/5c23a5bc7bd209231618dd9e606ce076272c9bc4f12023a70e03a86b4067/pandas-2.3.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:db4301b2d1f926ae677a751eb2bd0e8c5f5319c9cb3f88b0becbbb0b07b34151", size = 13890361, upload-time = "2025-09-29T23:19:25.342Z" },
{ url = "https://files.pythonhosted.org/packages/8e/59/712db1d7040520de7a4965df15b774348980e6df45c129b8c64d0dbe74ef/pandas-2.3.3-cp311-cp311-win_amd64.whl", hash = "sha256:f086f6fe114e19d92014a1966f43a3e62285109afe874f067f5abbdcbb10e59c", size = 11348702, upload-time = "2025-09-29T23:19:38.296Z" },
{ url = "https://files.pythonhosted.org/packages/9c/fb/231d89e8637c808b997d172b18e9d4a4bc7bf31296196c260526055d1ea0/pandas-2.3.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:6d21f6d74eb1725c2efaa71a2bfc661a0689579b58e9c0ca58a739ff0b002b53", size = 11597846, upload-time = "2025-09-29T23:19:48.856Z" },
{ url = "https://files.pythonhosted.org/packages/5c/bd/bf8064d9cfa214294356c2d6702b716d3cf3bb24be59287a6a21e24cae6b/pandas-2.3.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:3fd2f887589c7aa868e02632612ba39acb0b8948faf5cc58f0850e165bd46f35", size = 10729618, upload-time = "2025-09-29T23:39:08.659Z" },
{ url = "https://files.pythonhosted.org/packages/57/56/cf2dbe1a3f5271370669475ead12ce77c61726ffd19a35546e31aa8edf4e/pandas-2.3.3-cp312-cp312-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:ecaf1e12bdc03c86ad4a7ea848d66c685cb6851d807a26aa245ca3d2017a1908", size = 11737212, upload-time = "2025-09-29T23:19:59.765Z" },
{ url = "https://files.pythonhosted.org/packages/e5/63/cd7d615331b328e287d8233ba9fdf191a9c2d11b6af0c7a59cfcec23de68/pandas-2.3.3-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b3d11d2fda7eb164ef27ffc14b4fcab16a80e1ce67e9f57e19ec0afaf715ba89", size = 12362693, upload-time = "2025-09-29T23:20:14.098Z" },
{ url = "https://files.pythonhosted.org/packages/a6/de/8b1895b107277d52f2b42d3a6806e69cfef0d5cf1d0ba343470b9d8e0a04/pandas-2.3.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:a68e15f780eddf2b07d242e17a04aa187a7ee12b40b930bfdd78070556550e98", size = 12771002, upload-time = "2025-09-29T23:20:26.76Z" },
{ url = "https://files.pythonhosted.org/packages/87/21/84072af3187a677c5893b170ba2c8fbe450a6ff911234916da889b698220/pandas-2.3.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:371a4ab48e950033bcf52b6527eccb564f52dc826c02afd9a1bc0ab731bba084", size = 13450971, upload-time = "2025-09-29T23:20:41.344Z" },
{ url = "https://files.pythonhosted.org/packages/86/41/585a168330ff063014880a80d744219dbf1dd7a1c706e75ab3425a987384/pandas-2.3.3-cp312-cp312-win_amd64.whl", hash = "sha256:a16dcec078a01eeef8ee61bf64074b4e524a2a3f4b3be9326420cabe59c4778b", size = 10992722, upload-time = "2025-09-29T23:20:54.139Z" },
{ url = "https://files.pythonhosted.org/packages/cd/4b/18b035ee18f97c1040d94debd8f2e737000ad70ccc8f5513f4eefad75f4b/pandas-2.3.3-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:56851a737e3470de7fa88e6131f41281ed440d29a9268dcbf0002da5ac366713", size = 11544671, upload-time = "2025-09-29T23:21:05.024Z" },
{ url = "https://files.pythonhosted.org/packages/31/94/72fac03573102779920099bcac1c3b05975c2cb5f01eac609faf34bed1ca/pandas-2.3.3-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:bdcd9d1167f4885211e401b3036c0c8d9e274eee67ea8d0758a256d60704cfe8", size = 10680807, upload-time = "2025-09-29T23:21:15.979Z" },
{ url = "https://files.pythonhosted.org/packages/16/87/9472cf4a487d848476865321de18cc8c920b8cab98453ab79dbbc98db63a/pandas-2.3.3-cp313-cp313-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:e32e7cc9af0f1cc15548288a51a3b681cc2a219faa838e995f7dc53dbab1062d", size = 11709872, upload-time = "2025-09-29T23:21:27.165Z" },
{ url = "https://files.pythonhosted.org/packages/15/07/284f757f63f8a8d69ed4472bfd85122bd086e637bf4ed09de572d575a693/pandas-2.3.3-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:318d77e0e42a628c04dc56bcef4b40de67918f7041c2b061af1da41dcff670ac", size = 12306371, upload-time = "2025-09-29T23:21:40.532Z" },
{ url = "https://files.pythonhosted.org/packages/33/81/a3afc88fca4aa925804a27d2676d22dcd2031c2ebe08aabd0ae55b9ff282/pandas-2.3.3-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:4e0a175408804d566144e170d0476b15d78458795bb18f1304fb94160cabf40c", size = 12765333, upload-time = "2025-09-29T23:21:55.77Z" },
{ url = "https://files.pythonhosted.org/packages/8d/0f/b4d4ae743a83742f1153464cf1a8ecfafc3ac59722a0b5c8602310cb7158/pandas-2.3.3-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:93c2d9ab0fc11822b5eece72ec9587e172f63cff87c00b062f6e37448ced4493", size = 13418120, upload-time = "2025-09-29T23:22:10.109Z" },
{ url = "https://files.pythonhosted.org/packages/4f/c7/e54682c96a895d0c808453269e0b5928a07a127a15704fedb643e9b0a4c8/pandas-2.3.3-cp313-cp313-win_amd64.whl", hash = "sha256:f8bfc0e12dc78f777f323f55c58649591b2cd0c43534e8355c51d3fede5f4dee", size = 10993991, upload-time = "2025-09-29T23:25:04.889Z" },
{ url = "https://files.pythonhosted.org/packages/f9/ca/3f8d4f49740799189e1395812f3bf23b5e8fc7c190827d55a610da72ce55/pandas-2.3.3-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:75ea25f9529fdec2d2e93a42c523962261e567d250b0013b16210e1d40d7c2e5", size = 12048227, upload-time = "2025-09-29T23:22:24.343Z" },
{ url = "https://files.pythonhosted.org/packages/0e/5a/f43efec3e8c0cc92c4663ccad372dbdff72b60bdb56b2749f04aa1d07d7e/pandas-2.3.3-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:74ecdf1d301e812db96a465a525952f4dde225fdb6d8e5a521d47e1f42041e21", size = 11411056, upload-time = "2025-09-29T23:22:37.762Z" },
{ url = "https://files.pythonhosted.org/packages/46/b1/85331edfc591208c9d1a63a06baa67b21d332e63b7a591a5ba42a10bb507/pandas-2.3.3-cp313-cp313t-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6435cb949cb34ec11cc9860246ccb2fdc9ecd742c12d3304989017d53f039a78", size = 11645189, upload-time = "2025-09-29T23:22:51.688Z" },
{ url = "https://files.pythonhosted.org/packages/44/23/78d645adc35d94d1ac4f2a3c4112ab6f5b8999f4898b8cdf01252f8df4a9/pandas-2.3.3-cp313-cp313t-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:900f47d8f20860de523a1ac881c4c36d65efcb2eb850e6948140fa781736e110", size = 12121912, upload-time = "2025-09-29T23:23:05.042Z" },
{ url = "https://files.pythonhosted.org/packages/53/da/d10013df5e6aaef6b425aa0c32e1fc1f3e431e4bcabd420517dceadce354/pandas-2.3.3-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:a45c765238e2ed7d7c608fc5bc4a6f88b642f2f01e70c0c23d2224dd21829d86", size = 12712160, upload-time = "2025-09-29T23:23:28.57Z" },
{ url = "https://files.pythonhosted.org/packages/bd/17/e756653095a083d8a37cbd816cb87148debcfcd920129b25f99dd8d04271/pandas-2.3.3-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:c4fc4c21971a1a9f4bdb4c73978c7f7256caa3e62b323f70d6cb80db583350bc", size = 13199233, upload-time = "2025-09-29T23:24:24.876Z" },
{ url = "https://files.pythonhosted.org/packages/04/fd/74903979833db8390b73b3a8a7d30d146d710bd32703724dd9083950386f/pandas-2.3.3-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:ee15f284898e7b246df8087fc82b87b01686f98ee67d85a17b7ab44143a3a9a0", size = 11540635, upload-time = "2025-09-29T23:25:52.486Z" },
{ url = "https://files.pythonhosted.org/packages/21/00/266d6b357ad5e6d3ad55093a7e8efc7dd245f5a842b584db9f30b0f0a287/pandas-2.3.3-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:1611aedd912e1ff81ff41c745822980c49ce4a7907537be8692c8dbc31924593", size = 10759079, upload-time = "2025-09-29T23:26:33.204Z" },
{ url = "https://files.pythonhosted.org/packages/ca/05/d01ef80a7a3a12b2f8bbf16daba1e17c98a2f039cbc8e2f77a2c5a63d382/pandas-2.3.3-cp314-cp314-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6d2cefc361461662ac48810cb14365a365ce864afe85ef1f447ff5a1e99ea81c", size = 11814049, upload-time = "2025-09-29T23:27:15.384Z" },
{ url = "https://files.pythonhosted.org/packages/15/b2/0e62f78c0c5ba7e3d2c5945a82456f4fac76c480940f805e0b97fcbc2f65/pandas-2.3.3-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ee67acbbf05014ea6c763beb097e03cd629961c8a632075eeb34247120abcb4b", size = 12332638, upload-time = "2025-09-29T23:27:51.625Z" },
{ url = "https://files.pythonhosted.org/packages/c5/33/dd70400631b62b9b29c3c93d2feee1d0964dc2bae2e5ad7a6c73a7f25325/pandas-2.3.3-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:c46467899aaa4da076d5abc11084634e2d197e9460643dd455ac3db5856b24d6", size = 12886834, upload-time = "2025-09-29T23:28:21.289Z" },
{ url = "https://files.pythonhosted.org/packages/d3/18/b5d48f55821228d0d2692b34fd5034bb185e854bdb592e9c640f6290e012/pandas-2.3.3-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:6253c72c6a1d990a410bc7de641d34053364ef8bcd3126f7e7450125887dffe3", size = 13409925, upload-time = "2025-09-29T23:28:58.261Z" },
{ url = "https://files.pythonhosted.org/packages/a6/3d/124ac75fcd0ecc09b8fdccb0246ef65e35b012030defb0e0eba2cbbbe948/pandas-2.3.3-cp314-cp314-win_amd64.whl", hash = "sha256:1b07204a219b3b7350abaae088f451860223a52cfb8a6c53358e7948735158e5", size = 11109071, upload-time = "2025-09-29T23:32:27.484Z" },
{ url = "https://files.pythonhosted.org/packages/89/9c/0e21c895c38a157e0faa1fb64587a9226d6dd46452cac4532d80c3c4a244/pandas-2.3.3-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:2462b1a365b6109d275250baaae7b760fd25c726aaca0054649286bcfbb3e8ec", size = 12048504, upload-time = "2025-09-29T23:29:31.47Z" },
{ url = "https://files.pythonhosted.org/packages/d7/82/b69a1c95df796858777b68fbe6a81d37443a33319761d7c652ce77797475/pandas-2.3.3-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:0242fe9a49aa8b4d78a4fa03acb397a58833ef6199e9aa40a95f027bb3a1b6e7", size = 11410702, upload-time = "2025-09-29T23:29:54.591Z" },
{ url = "https://files.pythonhosted.org/packages/f9/88/702bde3ba0a94b8c73a0181e05144b10f13f29ebfc2150c3a79062a8195d/pandas-2.3.3-cp314-cp314t-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:a21d830e78df0a515db2b3d2f5570610f5e6bd2e27749770e8bb7b524b89b450", size = 11634535, upload-time = "2025-09-29T23:30:21.003Z" },
{ url = "https://files.pythonhosted.org/packages/a4/1e/1bac1a839d12e6a82ec6cb40cda2edde64a2013a66963293696bbf31fbbb/pandas-2.3.3-cp314-cp314t-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:2e3ebdb170b5ef78f19bfb71b0dc5dc58775032361fa188e814959b74d726dd5", size = 12121582, upload-time = "2025-09-29T23:30:43.391Z" },
{ url = "https://files.pythonhosted.org/packages/44/91/483de934193e12a3b1d6ae7c8645d083ff88dec75f46e827562f1e4b4da6/pandas-2.3.3-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:d051c0e065b94b7a3cea50eb1ec32e912cd96dba41647eb24104b6c6c14c5788", size = 12699963, upload-time = "2025-09-29T23:31:10.009Z" },
{ url = "https://files.pythonhosted.org/packages/70/44/5191d2e4026f86a2a109053e194d3ba7a31a2d10a9c2348368c63ed4e85a/pandas-2.3.3-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:3869faf4bd07b3b66a9f462417d0ca3a9df29a9f6abd5d0d0dbab15dac7abe87", size = 13202175, upload-time = "2025-09-29T23:31:59.173Z" },
{ url = "https://files.pythonhosted.org/packages/56/b4/52eeb530a99e2a4c55ffcd352772b599ed4473a0f892d127f4147cf0f88e/pandas-2.3.3-cp39-cp39-macosx_10_9_x86_64.whl", hash = "sha256:c503ba5216814e295f40711470446bc3fd00f0faea8a086cbc688808e26f92a2", size = 11567720, upload-time = "2025-09-29T23:33:06.209Z" },
{ url = "https://files.pythonhosted.org/packages/48/4a/2d8b67632a021bced649ba940455ed441ca854e57d6e7658a6024587b083/pandas-2.3.3-cp39-cp39-macosx_11_0_arm64.whl", hash = "sha256:a637c5cdfa04b6d6e2ecedcb81fc52ffb0fd78ce2ebccc9ea964df9f658de8c8", size = 10810302, upload-time = "2025-09-29T23:33:35.846Z" },
{ url = "https://files.pythonhosted.org/packages/13/e6/d2465010ee0569a245c975dc6967b801887068bc893e908239b1f4b6c1ac/pandas-2.3.3-cp39-cp39-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:854d00d556406bffe66a4c0802f334c9ad5a96b4f1f868adf036a21b11ef13ff", size = 12154874, upload-time = "2025-09-29T23:33:49.939Z" },
{ url = "https://files.pythonhosted.org/packages/1f/18/aae8c0aa69a386a3255940e9317f793808ea79d0a525a97a903366bb2569/pandas-2.3.3-cp39-cp39-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:bf1f8a81d04ca90e32a0aceb819d34dbd378a98bf923b6398b9a3ec0bf44de29", size = 12790141, upload-time = "2025-09-29T23:34:05.655Z" },
{ url = "https://files.pythonhosted.org/packages/f7/26/617f98de789de00c2a444fbe6301bb19e66556ac78cff933d2c98f62f2b4/pandas-2.3.3-cp39-cp39-musllinux_1_2_aarch64.whl", hash = "sha256:23ebd657a4d38268c7dfbdf089fbc31ea709d82e4923c5ffd4fbd5747133ce73", size = 13208697, upload-time = "2025-09-29T23:34:21.835Z" },
{ url = "https://files.pythonhosted.org/packages/b9/fb/25709afa4552042bd0e15717c75e9b4a2294c3dc4f7e6ea50f03c5136600/pandas-2.3.3-cp39-cp39-musllinux_1_2_x86_64.whl", hash = "sha256:5554c929ccc317d41a5e3d1234f3be588248e61f08a74dd17c9eabb535777dc9", size = 13879233, upload-time = "2025-09-29T23:34:35.079Z" },
{ url = "https://files.pythonhosted.org/packages/98/af/7be05277859a7bc399da8ba68b88c96b27b48740b6cf49688899c6eb4176/pandas-2.3.3-cp39-cp39-win_amd64.whl", hash = "sha256:d3e28b3e83862ccf4d85ff19cf8c20b2ae7e503881711ff2d534dc8f761131aa", size = 11359119, upload-time = "2025-09-29T23:34:46.339Z" },
]
[[package]]
name = "pandocfilters"
version = "1.5.1"
@@ -2735,6 +2841,122 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/8e/37/efad0257dc6e593a18957422533ff0f87ede7c9c6ea010a2177d738fb82f/pure_eval-0.2.3-py3-none-any.whl", hash = "sha256:1db8e35b67b3d218d818ae653e27f06c3aa420901fa7b081ca98cbedc874e0d0", size = 11842, upload-time = "2024-07-21T12:58:20.04Z" },
]
[[package]]
name = "pyarrow"
version = "21.0.0"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version < '3.10'",
]
sdist = { url = "https://files.pythonhosted.org/packages/ef/c2/ea068b8f00905c06329a3dfcd40d0fcc2b7d0f2e355bdb25b65e0a0e4cd4/pyarrow-21.0.0.tar.gz", hash = "sha256:5051f2dccf0e283ff56335760cbc8622cf52264d67e359d5569541ac11b6d5bc", size = 1133487, upload-time = "2025-07-18T00:57:31.761Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/17/d9/110de31880016e2afc52d8580b397dbe47615defbf09ca8cf55f56c62165/pyarrow-21.0.0-cp310-cp310-macosx_12_0_arm64.whl", hash = "sha256:e563271e2c5ff4d4a4cbeb2c83d5cf0d4938b891518e676025f7268c6fe5fe26", size = 31196837, upload-time = "2025-07-18T00:54:34.755Z" },
{ url = "https://files.pythonhosted.org/packages/df/5f/c1c1997613abf24fceb087e79432d24c19bc6f7259cab57c2c8e5e545fab/pyarrow-21.0.0-cp310-cp310-macosx_12_0_x86_64.whl", hash = "sha256:fee33b0ca46f4c85443d6c450357101e47d53e6c3f008d658c27a2d020d44c79", size = 32659470, upload-time = "2025-07-18T00:54:38.329Z" },
{ url = "https://files.pythonhosted.org/packages/3e/ed/b1589a777816ee33ba123ba1e4f8f02243a844fed0deec97bde9fb21a5cf/pyarrow-21.0.0-cp310-cp310-manylinux_2_28_aarch64.whl", hash = "sha256:7be45519b830f7c24b21d630a31d48bcebfd5d4d7f9d3bdb49da9cdf6d764edb", size = 41055619, upload-time = "2025-07-18T00:54:42.172Z" },
{ url = "https://files.pythonhosted.org/packages/44/28/b6672962639e85dc0ac36f71ab3a8f5f38e01b51343d7aa372a6b56fa3f3/pyarrow-21.0.0-cp310-cp310-manylinux_2_28_x86_64.whl", hash = "sha256:26bfd95f6bff443ceae63c65dc7e048670b7e98bc892210acba7e4995d3d4b51", size = 42733488, upload-time = "2025-07-18T00:54:47.132Z" },
{ url = "https://files.pythonhosted.org/packages/f8/cc/de02c3614874b9089c94eac093f90ca5dfa6d5afe45de3ba847fd950fdf1/pyarrow-21.0.0-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:bd04ec08f7f8bd113c55868bd3fc442a9db67c27af098c5f814a3091e71cc61a", size = 43329159, upload-time = "2025-07-18T00:54:51.686Z" },
{ url = "https://files.pythonhosted.org/packages/a6/3e/99473332ac40278f196e105ce30b79ab8affab12f6194802f2593d6b0be2/pyarrow-21.0.0-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:9b0b14b49ac10654332a805aedfc0147fb3469cbf8ea951b3d040dab12372594", size = 45050567, upload-time = "2025-07-18T00:54:56.679Z" },
{ url = "https://files.pythonhosted.org/packages/7b/f5/c372ef60593d713e8bfbb7e0c743501605f0ad00719146dc075faf11172b/pyarrow-21.0.0-cp310-cp310-win_amd64.whl", hash = "sha256:9d9f8bcb4c3be7738add259738abdeddc363de1b80e3310e04067aa1ca596634", size = 26217959, upload-time = "2025-07-18T00:55:00.482Z" },
{ url = "https://files.pythonhosted.org/packages/94/dc/80564a3071a57c20b7c32575e4a0120e8a330ef487c319b122942d665960/pyarrow-21.0.0-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:c077f48aab61738c237802836fc3844f85409a46015635198761b0d6a688f87b", size = 31243234, upload-time = "2025-07-18T00:55:03.812Z" },
{ url = "https://files.pythonhosted.org/packages/ea/cc/3b51cb2db26fe535d14f74cab4c79b191ed9a8cd4cbba45e2379b5ca2746/pyarrow-21.0.0-cp311-cp311-macosx_12_0_x86_64.whl", hash = "sha256:689f448066781856237eca8d1975b98cace19b8dd2ab6145bf49475478bcaa10", size = 32714370, upload-time = "2025-07-18T00:55:07.495Z" },
{ url = "https://files.pythonhosted.org/packages/24/11/a4431f36d5ad7d83b87146f515c063e4d07ef0b7240876ddb885e6b44f2e/pyarrow-21.0.0-cp311-cp311-manylinux_2_28_aarch64.whl", hash = "sha256:479ee41399fcddc46159a551705b89c05f11e8b8cb8e968f7fec64f62d91985e", size = 41135424, upload-time = "2025-07-18T00:55:11.461Z" },
{ url = "https://files.pythonhosted.org/packages/74/dc/035d54638fc5d2971cbf1e987ccd45f1091c83bcf747281cf6cc25e72c88/pyarrow-21.0.0-cp311-cp311-manylinux_2_28_x86_64.whl", hash = "sha256:40ebfcb54a4f11bcde86bc586cbd0272bac0d516cfa539c799c2453768477569", size = 42823810, upload-time = "2025-07-18T00:55:16.301Z" },
{ url = "https://files.pythonhosted.org/packages/2e/3b/89fced102448a9e3e0d4dded1f37fa3ce4700f02cdb8665457fcc8015f5b/pyarrow-21.0.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:8d58d8497814274d3d20214fbb24abcad2f7e351474357d552a8d53bce70c70e", size = 43391538, upload-time = "2025-07-18T00:55:23.82Z" },
{ url = "https://files.pythonhosted.org/packages/fb/bb/ea7f1bd08978d39debd3b23611c293f64a642557e8141c80635d501e6d53/pyarrow-21.0.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:585e7224f21124dd57836b1530ac8f2df2afc43c861d7bf3d58a4870c42ae36c", size = 45120056, upload-time = "2025-07-18T00:55:28.231Z" },
{ url = "https://files.pythonhosted.org/packages/6e/0b/77ea0600009842b30ceebc3337639a7380cd946061b620ac1a2f3cb541e2/pyarrow-21.0.0-cp311-cp311-win_amd64.whl", hash = "sha256:555ca6935b2cbca2c0e932bedd853e9bc523098c39636de9ad4693b5b1df86d6", size = 26220568, upload-time = "2025-07-18T00:55:32.122Z" },
{ url = "https://files.pythonhosted.org/packages/ca/d4/d4f817b21aacc30195cf6a46ba041dd1be827efa4a623cc8bf39a1c2a0c0/pyarrow-21.0.0-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:3a302f0e0963db37e0a24a70c56cf91a4faa0bca51c23812279ca2e23481fccd", size = 31160305, upload-time = "2025-07-18T00:55:35.373Z" },
{ url = "https://files.pythonhosted.org/packages/a2/9c/dcd38ce6e4b4d9a19e1d36914cb8e2b1da4e6003dd075474c4cfcdfe0601/pyarrow-21.0.0-cp312-cp312-macosx_12_0_x86_64.whl", hash = "sha256:b6b27cf01e243871390474a211a7922bfbe3bda21e39bc9160daf0da3fe48876", size = 32684264, upload-time = "2025-07-18T00:55:39.303Z" },
{ url = "https://files.pythonhosted.org/packages/4f/74/2a2d9f8d7a59b639523454bec12dba35ae3d0a07d8ab529dc0809f74b23c/pyarrow-21.0.0-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:e72a8ec6b868e258a2cd2672d91f2860ad532d590ce94cdf7d5e7ec674ccf03d", size = 41108099, upload-time = "2025-07-18T00:55:42.889Z" },
{ url = "https://files.pythonhosted.org/packages/ad/90/2660332eeb31303c13b653ea566a9918484b6e4d6b9d2d46879a33ab0622/pyarrow-21.0.0-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:b7ae0bbdc8c6674259b25bef5d2a1d6af5d39d7200c819cf99e07f7dfef1c51e", size = 42829529, upload-time = "2025-07-18T00:55:47.069Z" },
{ url = "https://files.pythonhosted.org/packages/33/27/1a93a25c92717f6aa0fca06eb4700860577d016cd3ae51aad0e0488ac899/pyarrow-21.0.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:58c30a1729f82d201627c173d91bd431db88ea74dcaa3885855bc6203e433b82", size = 43367883, upload-time = "2025-07-18T00:55:53.069Z" },
{ url = "https://files.pythonhosted.org/packages/05/d9/4d09d919f35d599bc05c6950095e358c3e15148ead26292dfca1fb659b0c/pyarrow-21.0.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:072116f65604b822a7f22945a7a6e581cfa28e3454fdcc6939d4ff6090126623", size = 45133802, upload-time = "2025-07-18T00:55:57.714Z" },
{ url = "https://files.pythonhosted.org/packages/71/30/f3795b6e192c3ab881325ffe172e526499eb3780e306a15103a2764916a2/pyarrow-21.0.0-cp312-cp312-win_amd64.whl", hash = "sha256:cf56ec8b0a5c8c9d7021d6fd754e688104f9ebebf1bf4449613c9531f5346a18", size = 26203175, upload-time = "2025-07-18T00:56:01.364Z" },
{ url = "https://files.pythonhosted.org/packages/16/ca/c7eaa8e62db8fb37ce942b1ea0c6d7abfe3786ca193957afa25e71b81b66/pyarrow-21.0.0-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:e99310a4ebd4479bcd1964dff9e14af33746300cb014aa4a3781738ac63baf4a", size = 31154306, upload-time = "2025-07-18T00:56:04.42Z" },
{ url = "https://files.pythonhosted.org/packages/ce/e8/e87d9e3b2489302b3a1aea709aaca4b781c5252fcb812a17ab6275a9a484/pyarrow-21.0.0-cp313-cp313-macosx_12_0_x86_64.whl", hash = "sha256:d2fe8e7f3ce329a71b7ddd7498b3cfac0eeb200c2789bd840234f0dc271a8efe", size = 32680622, upload-time = "2025-07-18T00:56:07.505Z" },
{ url = "https://files.pythonhosted.org/packages/84/52/79095d73a742aa0aba370c7942b1b655f598069489ab387fe47261a849e1/pyarrow-21.0.0-cp313-cp313-manylinux_2_28_aarch64.whl", hash = "sha256:f522e5709379d72fb3da7785aa489ff0bb87448a9dc5a75f45763a795a089ebd", size = 41104094, upload-time = "2025-07-18T00:56:10.994Z" },
{ url = "https://files.pythonhosted.org/packages/89/4b/7782438b551dbb0468892a276b8c789b8bbdb25ea5c5eb27faadd753e037/pyarrow-21.0.0-cp313-cp313-manylinux_2_28_x86_64.whl", hash = "sha256:69cbbdf0631396e9925e048cfa5bce4e8c3d3b41562bbd70c685a8eb53a91e61", size = 42825576, upload-time = "2025-07-18T00:56:15.569Z" },
{ url = "https://files.pythonhosted.org/packages/b3/62/0f29de6e0a1e33518dec92c65be0351d32d7ca351e51ec5f4f837a9aab91/pyarrow-21.0.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:731c7022587006b755d0bdb27626a1a3bb004bb56b11fb30d98b6c1b4718579d", size = 43368342, upload-time = "2025-07-18T00:56:19.531Z" },
{ url = "https://files.pythonhosted.org/packages/90/c7/0fa1f3f29cf75f339768cc698c8ad4ddd2481c1742e9741459911c9ac477/pyarrow-21.0.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:dc56bc708f2d8ac71bd1dcb927e458c93cec10b98eb4120206a4091db7b67b99", size = 45131218, upload-time = "2025-07-18T00:56:23.347Z" },
{ url = "https://files.pythonhosted.org/packages/01/63/581f2076465e67b23bc5a37d4a2abff8362d389d29d8105832e82c9c811c/pyarrow-21.0.0-cp313-cp313-win_amd64.whl", hash = "sha256:186aa00bca62139f75b7de8420f745f2af12941595bbbfa7ed3870ff63e25636", size = 26087551, upload-time = "2025-07-18T00:56:26.758Z" },
{ url = "https://files.pythonhosted.org/packages/c9/ab/357d0d9648bb8241ee7348e564f2479d206ebe6e1c47ac5027c2e31ecd39/pyarrow-21.0.0-cp313-cp313t-macosx_12_0_arm64.whl", hash = "sha256:a7a102574faa3f421141a64c10216e078df467ab9576684d5cd696952546e2da", size = 31290064, upload-time = "2025-07-18T00:56:30.214Z" },
{ url = "https://files.pythonhosted.org/packages/3f/8a/5685d62a990e4cac2043fc76b4661bf38d06efed55cf45a334b455bd2759/pyarrow-21.0.0-cp313-cp313t-macosx_12_0_x86_64.whl", hash = "sha256:1e005378c4a2c6db3ada3ad4c217b381f6c886f0a80d6a316fe586b90f77efd7", size = 32727837, upload-time = "2025-07-18T00:56:33.935Z" },
{ url = "https://files.pythonhosted.org/packages/fc/de/c0828ee09525c2bafefd3e736a248ebe764d07d0fd762d4f0929dbc516c9/pyarrow-21.0.0-cp313-cp313t-manylinux_2_28_aarch64.whl", hash = "sha256:65f8e85f79031449ec8706b74504a316805217b35b6099155dd7e227eef0d4b6", size = 41014158, upload-time = "2025-07-18T00:56:37.528Z" },
{ url = "https://files.pythonhosted.org/packages/6e/26/a2865c420c50b7a3748320b614f3484bfcde8347b2639b2b903b21ce6a72/pyarrow-21.0.0-cp313-cp313t-manylinux_2_28_x86_64.whl", hash = "sha256:3a81486adc665c7eb1a2bde0224cfca6ceaba344a82a971ef059678417880eb8", size = 42667885, upload-time = "2025-07-18T00:56:41.483Z" },
{ url = "https://files.pythonhosted.org/packages/0a/f9/4ee798dc902533159250fb4321267730bc0a107d8c6889e07c3add4fe3a5/pyarrow-21.0.0-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:fc0d2f88b81dcf3ccf9a6ae17f89183762c8a94a5bdcfa09e05cfe413acf0503", size = 43276625, upload-time = "2025-07-18T00:56:48.002Z" },
{ url = "https://files.pythonhosted.org/packages/5a/da/e02544d6997037a4b0d22d8e5f66bc9315c3671371a8b18c79ade1cefe14/pyarrow-21.0.0-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:6299449adf89df38537837487a4f8d3bd91ec94354fdd2a7d30bc11c48ef6e79", size = 44951890, upload-time = "2025-07-18T00:56:52.568Z" },
{ url = "https://files.pythonhosted.org/packages/e5/4e/519c1bc1876625fe6b71e9a28287c43ec2f20f73c658b9ae1d485c0c206e/pyarrow-21.0.0-cp313-cp313t-win_amd64.whl", hash = "sha256:222c39e2c70113543982c6b34f3077962b44fca38c0bd9e68bb6781534425c10", size = 26371006, upload-time = "2025-07-18T00:56:56.379Z" },
{ url = "https://files.pythonhosted.org/packages/3e/cc/ce4939f4b316457a083dc5718b3982801e8c33f921b3c98e7a93b7c7491f/pyarrow-21.0.0-cp39-cp39-macosx_12_0_arm64.whl", hash = "sha256:a7f6524e3747e35f80744537c78e7302cd41deee8baa668d56d55f77d9c464b3", size = 31211248, upload-time = "2025-07-18T00:56:59.7Z" },
{ url = "https://files.pythonhosted.org/packages/1f/c2/7a860931420d73985e2f340f06516b21740c15b28d24a0e99a900bb27d2b/pyarrow-21.0.0-cp39-cp39-macosx_12_0_x86_64.whl", hash = "sha256:203003786c9fd253ebcafa44b03c06983c9c8d06c3145e37f1b76a1f317aeae1", size = 32676896, upload-time = "2025-07-18T00:57:03.884Z" },
{ url = "https://files.pythonhosted.org/packages/68/a8/197f989b9a75e59b4ca0db6a13c56f19a0ad8a298c68da9cc28145e0bb97/pyarrow-21.0.0-cp39-cp39-manylinux_2_28_aarch64.whl", hash = "sha256:3b4d97e297741796fead24867a8dabf86c87e4584ccc03167e4a811f50fdf74d", size = 41067862, upload-time = "2025-07-18T00:57:07.587Z" },
{ url = "https://files.pythonhosted.org/packages/fa/82/6ecfa89487b35aa21accb014b64e0a6b814cc860d5e3170287bf5135c7d8/pyarrow-21.0.0-cp39-cp39-manylinux_2_28_x86_64.whl", hash = "sha256:898afce396b80fdda05e3086b4256f8677c671f7b1d27a6976fa011d3fd0a86e", size = 42747508, upload-time = "2025-07-18T00:57:13.917Z" },
{ url = "https://files.pythonhosted.org/packages/3b/b7/ba252f399bbf3addc731e8643c05532cf32e74cebb5e32f8f7409bc243cf/pyarrow-21.0.0-cp39-cp39-musllinux_1_2_aarch64.whl", hash = "sha256:067c66ca29aaedae08218569a114e413b26e742171f526e828e1064fcdec13f4", size = 43345293, upload-time = "2025-07-18T00:57:19.828Z" },
{ url = "https://files.pythonhosted.org/packages/ff/0a/a20819795bd702b9486f536a8eeb70a6aa64046fce32071c19ec8230dbaa/pyarrow-21.0.0-cp39-cp39-musllinux_1_2_x86_64.whl", hash = "sha256:0c4e75d13eb76295a49e0ea056eb18dbd87d81450bfeb8afa19a7e5a75ae2ad7", size = 45060670, upload-time = "2025-07-18T00:57:24.477Z" },
{ url = "https://files.pythonhosted.org/packages/10/15/6b30e77872012bbfe8265d42a01d5b3c17ef0ac0f2fae531ad91b6a6c02e/pyarrow-21.0.0-cp39-cp39-win_amd64.whl", hash = "sha256:cdc4c17afda4dab2a9c0b79148a43a7f4e1094916b3e18d8975bfd6d6d52241f", size = 26227521, upload-time = "2025-07-18T00:57:29.119Z" },
]
[[package]]
name = "pyarrow"
version = "22.0.0"
source = { registry = "https://pypi.org/simple" }
resolution-markers = [
"python_full_version >= '3.14'",
"python_full_version >= '3.12' and python_full_version < '3.14'",
"python_full_version == '3.11.*'",
"python_full_version == '3.10.*'",
]
sdist = { url = "https://files.pythonhosted.org/packages/30/53/04a7fdc63e6056116c9ddc8b43bc28c12cdd181b85cbeadb79278475f3ae/pyarrow-22.0.0.tar.gz", hash = "sha256:3d600dc583260d845c7d8a6db540339dd883081925da2bd1c5cb808f720b3cd9", size = 1151151, upload-time = "2025-10-24T12:30:00.762Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/d9/9b/cb3f7e0a345353def531ca879053e9ef6b9f38ed91aebcf68b09ba54dec0/pyarrow-22.0.0-cp310-cp310-macosx_12_0_arm64.whl", hash = "sha256:77718810bd3066158db1e95a63c160ad7ce08c6b0710bc656055033e39cdad88", size = 34223968, upload-time = "2025-10-24T10:03:31.21Z" },
{ url = "https://files.pythonhosted.org/packages/6c/41/3184b8192a120306270c5307f105b70320fdaa592c99843c5ef78aaefdcf/pyarrow-22.0.0-cp310-cp310-macosx_12_0_x86_64.whl", hash = "sha256:44d2d26cda26d18f7af7db71453b7b783788322d756e81730acb98f24eb90ace", size = 35942085, upload-time = "2025-10-24T10:03:38.146Z" },
{ url = "https://files.pythonhosted.org/packages/d9/3d/a1eab2f6f08001f9fb714b8ed5cfb045e2fe3e3e3c0c221f2c9ed1e6d67d/pyarrow-22.0.0-cp310-cp310-manylinux_2_28_aarch64.whl", hash = "sha256:b9d71701ce97c95480fecb0039ec5bb889e75f110da72005743451339262f4ce", size = 44964613, upload-time = "2025-10-24T10:03:46.516Z" },
{ url = "https://files.pythonhosted.org/packages/46/46/a1d9c24baf21cfd9ce994ac820a24608decf2710521b29223d4334985127/pyarrow-22.0.0-cp310-cp310-manylinux_2_28_x86_64.whl", hash = "sha256:710624ab925dc2b05a6229d47f6f0dac1c1155e6ed559be7109f684eba048a48", size = 47627059, upload-time = "2025-10-24T10:03:55.353Z" },
{ url = "https://files.pythonhosted.org/packages/3a/4c/f711acb13075c1391fd54bc17e078587672c575f8de2a6e62509af026dcf/pyarrow-22.0.0-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:f963ba8c3b0199f9d6b794c90ec77545e05eadc83973897a4523c9e8d84e9340", size = 47947043, upload-time = "2025-10-24T10:04:05.408Z" },
{ url = "https://files.pythonhosted.org/packages/4e/70/1f3180dd7c2eab35c2aca2b29ace6c519f827dcd4cfeb8e0dca41612cf7a/pyarrow-22.0.0-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:bd0d42297ace400d8febe55f13fdf46e86754842b860c978dfec16f081e5c653", size = 50206505, upload-time = "2025-10-24T10:04:15.786Z" },
{ url = "https://files.pythonhosted.org/packages/80/07/fea6578112c8c60ffde55883a571e4c4c6bc7049f119d6b09333b5cc6f73/pyarrow-22.0.0-cp310-cp310-win_amd64.whl", hash = "sha256:00626d9dc0f5ef3a75fe63fd68b9c7c8302d2b5bbc7f74ecaedba83447a24f84", size = 28101641, upload-time = "2025-10-24T10:04:22.57Z" },
{ url = "https://files.pythonhosted.org/packages/2e/b7/18f611a8cdc43417f9394a3ccd3eace2f32183c08b9eddc3d17681819f37/pyarrow-22.0.0-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:3e294c5eadfb93d78b0763e859a0c16d4051fc1c5231ae8956d61cb0b5666f5a", size = 34272022, upload-time = "2025-10-24T10:04:28.973Z" },
{ url = "https://files.pythonhosted.org/packages/26/5c/f259e2526c67eb4b9e511741b19870a02363a47a35edbebc55c3178db22d/pyarrow-22.0.0-cp311-cp311-macosx_12_0_x86_64.whl", hash = "sha256:69763ab2445f632d90b504a815a2a033f74332997052b721002298ed6de40f2e", size = 35995834, upload-time = "2025-10-24T10:04:35.467Z" },
{ url = "https://files.pythonhosted.org/packages/50/8d/281f0f9b9376d4b7f146913b26fac0aa2829cd1ee7e997f53a27411bbb92/pyarrow-22.0.0-cp311-cp311-manylinux_2_28_aarch64.whl", hash = "sha256:b41f37cabfe2463232684de44bad753d6be08a7a072f6a83447eeaf0e4d2a215", size = 45030348, upload-time = "2025-10-24T10:04:43.366Z" },
{ url = "https://files.pythonhosted.org/packages/f5/e5/53c0a1c428f0976bf22f513d79c73000926cb00b9c138d8e02daf2102e18/pyarrow-22.0.0-cp311-cp311-manylinux_2_28_x86_64.whl", hash = "sha256:35ad0f0378c9359b3f297299c3309778bb03b8612f987399a0333a560b43862d", size = 47699480, upload-time = "2025-10-24T10:04:51.486Z" },
{ url = "https://files.pythonhosted.org/packages/95/e1/9dbe4c465c3365959d183e6345d0a8d1dc5b02ca3f8db4760b3bc834cf25/pyarrow-22.0.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:8382ad21458075c2e66a82a29d650f963ce51c7708c7c0ff313a8c206c4fd5e8", size = 48011148, upload-time = "2025-10-24T10:04:59.585Z" },
{ url = "https://files.pythonhosted.org/packages/c5/b4/7caf5d21930061444c3cf4fa7535c82faf5263e22ce43af7c2759ceb5b8b/pyarrow-22.0.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:1a812a5b727bc09c3d7ea072c4eebf657c2f7066155506ba31ebf4792f88f016", size = 50276964, upload-time = "2025-10-24T10:05:08.175Z" },
{ url = "https://files.pythonhosted.org/packages/ae/f3/cec89bd99fa3abf826f14d4e53d3d11340ce6f6af4d14bdcd54cd83b6576/pyarrow-22.0.0-cp311-cp311-win_amd64.whl", hash = "sha256:ec5d40dd494882704fb876c16fa7261a69791e784ae34e6b5992e977bd2e238c", size = 28106517, upload-time = "2025-10-24T10:05:14.314Z" },
{ url = "https://files.pythonhosted.org/packages/af/63/ba23862d69652f85b615ca14ad14f3bcfc5bf1b99ef3f0cd04ff93fdad5a/pyarrow-22.0.0-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:bea79263d55c24a32b0d79c00a1c58bb2ee5f0757ed95656b01c0fb310c5af3d", size = 34211578, upload-time = "2025-10-24T10:05:21.583Z" },
{ url = "https://files.pythonhosted.org/packages/b1/d0/f9ad86fe809efd2bcc8be32032fa72e8b0d112b01ae56a053006376c5930/pyarrow-22.0.0-cp312-cp312-macosx_12_0_x86_64.whl", hash = "sha256:12fe549c9b10ac98c91cf791d2945e878875d95508e1a5d14091a7aaa66d9cf8", size = 35989906, upload-time = "2025-10-24T10:05:29.485Z" },
{ url = "https://files.pythonhosted.org/packages/b4/a8/f910afcb14630e64d673f15904ec27dd31f1e009b77033c365c84e8c1e1d/pyarrow-22.0.0-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:334f900ff08ce0423407af97e6c26ad5d4e3b0763645559ece6fbf3747d6a8f5", size = 45021677, upload-time = "2025-10-24T10:05:38.274Z" },
{ url = "https://files.pythonhosted.org/packages/13/95/aec81f781c75cd10554dc17a25849c720d54feafb6f7847690478dcf5ef8/pyarrow-22.0.0-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:c6c791b09c57ed76a18b03f2631753a4960eefbbca80f846da8baefc6491fcfe", size = 47726315, upload-time = "2025-10-24T10:05:47.314Z" },
{ url = "https://files.pythonhosted.org/packages/bb/d4/74ac9f7a54cfde12ee42734ea25d5a3c9a45db78f9def949307a92720d37/pyarrow-22.0.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:c3200cb41cdbc65156e5f8c908d739b0dfed57e890329413da2748d1a2cd1a4e", size = 47990906, upload-time = "2025-10-24T10:05:58.254Z" },
{ url = "https://files.pythonhosted.org/packages/2e/71/fedf2499bf7a95062eafc989ace56572f3343432570e1c54e6599d5b88da/pyarrow-22.0.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:ac93252226cf288753d8b46280f4edf3433bf9508b6977f8dd8526b521a1bbb9", size = 50306783, upload-time = "2025-10-24T10:06:08.08Z" },
{ url = "https://files.pythonhosted.org/packages/68/ed/b202abd5a5b78f519722f3d29063dda03c114711093c1995a33b8e2e0f4b/pyarrow-22.0.0-cp312-cp312-win_amd64.whl", hash = "sha256:44729980b6c50a5f2bfcc2668d36c569ce17f8b17bccaf470c4313dcbbf13c9d", size = 27972883, upload-time = "2025-10-24T10:06:14.204Z" },
{ url = "https://files.pythonhosted.org/packages/a6/d6/d0fac16a2963002fc22c8fa75180a838737203d558f0ed3b564c4a54eef5/pyarrow-22.0.0-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:e6e95176209257803a8b3d0394f21604e796dadb643d2f7ca21b66c9c0b30c9a", size = 34204629, upload-time = "2025-10-24T10:06:20.274Z" },
{ url = "https://files.pythonhosted.org/packages/c6/9c/1d6357347fbae062ad3f17082f9ebc29cc733321e892c0d2085f42a2212b/pyarrow-22.0.0-cp313-cp313-macosx_12_0_x86_64.whl", hash = "sha256:001ea83a58024818826a9e3f89bf9310a114f7e26dfe404a4c32686f97bd7901", size = 35985783, upload-time = "2025-10-24T10:06:27.301Z" },
{ url = "https://files.pythonhosted.org/packages/ff/c0/782344c2ce58afbea010150df07e3a2f5fdad299cd631697ae7bd3bac6e3/pyarrow-22.0.0-cp313-cp313-manylinux_2_28_aarch64.whl", hash = "sha256:ce20fe000754f477c8a9125543f1936ea5b8867c5406757c224d745ed033e691", size = 45020999, upload-time = "2025-10-24T10:06:35.387Z" },
{ url = "https://files.pythonhosted.org/packages/1b/8b/5362443737a5307a7b67c1017c42cd104213189b4970bf607e05faf9c525/pyarrow-22.0.0-cp313-cp313-manylinux_2_28_x86_64.whl", hash = "sha256:e0a15757fccb38c410947df156f9749ae4a3c89b2393741a50521f39a8cf202a", size = 47724601, upload-time = "2025-10-24T10:06:43.551Z" },
{ url = "https://files.pythonhosted.org/packages/69/4d/76e567a4fc2e190ee6072967cb4672b7d9249ac59ae65af2d7e3047afa3b/pyarrow-22.0.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:cedb9dd9358e4ea1d9bce3665ce0797f6adf97ff142c8e25b46ba9cdd508e9b6", size = 48001050, upload-time = "2025-10-24T10:06:52.284Z" },
{ url = "https://files.pythonhosted.org/packages/01/5e/5653f0535d2a1aef8223cee9d92944cb6bccfee5cf1cd3f462d7cb022790/pyarrow-22.0.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:252be4a05f9d9185bb8c18e83764ebcfea7185076c07a7a662253af3a8c07941", size = 50307877, upload-time = "2025-10-24T10:07:02.405Z" },
{ url = "https://files.pythonhosted.org/packages/2d/f8/1d0bd75bf9328a3b826e24a16e5517cd7f9fbf8d34a3184a4566ef5a7f29/pyarrow-22.0.0-cp313-cp313-win_amd64.whl", hash = "sha256:a4893d31e5ef780b6edcaf63122df0f8d321088bb0dee4c8c06eccb1ca28d145", size = 27977099, upload-time = "2025-10-24T10:08:07.259Z" },
{ url = "https://files.pythonhosted.org/packages/90/81/db56870c997805bf2b0f6eeeb2d68458bf4654652dccdcf1bf7a42d80903/pyarrow-22.0.0-cp313-cp313t-macosx_12_0_arm64.whl", hash = "sha256:f7fe3dbe871294ba70d789be16b6e7e52b418311e166e0e3cba9522f0f437fb1", size = 34336685, upload-time = "2025-10-24T10:07:11.47Z" },
{ url = "https://files.pythonhosted.org/packages/1c/98/0727947f199aba8a120f47dfc229eeb05df15bcd7a6f1b669e9f882afc58/pyarrow-22.0.0-cp313-cp313t-macosx_12_0_x86_64.whl", hash = "sha256:ba95112d15fd4f1105fb2402c4eab9068f0554435e9b7085924bcfaac2cc306f", size = 36032158, upload-time = "2025-10-24T10:07:18.626Z" },
{ url = "https://files.pythonhosted.org/packages/96/b4/9babdef9c01720a0785945c7cf550e4acd0ebcd7bdd2e6f0aa7981fa85e2/pyarrow-22.0.0-cp313-cp313t-manylinux_2_28_aarch64.whl", hash = "sha256:c064e28361c05d72eed8e744c9605cbd6d2bb7481a511c74071fd9b24bc65d7d", size = 44892060, upload-time = "2025-10-24T10:07:26.002Z" },
{ url = "https://files.pythonhosted.org/packages/f8/ca/2f8804edd6279f78a37062d813de3f16f29183874447ef6d1aadbb4efa0f/pyarrow-22.0.0-cp313-cp313t-manylinux_2_28_x86_64.whl", hash = "sha256:6f9762274496c244d951c819348afbcf212714902742225f649cf02823a6a10f", size = 47504395, upload-time = "2025-10-24T10:07:34.09Z" },
{ url = "https://files.pythonhosted.org/packages/b9/f0/77aa5198fd3943682b2e4faaf179a674f0edea0d55d326d83cb2277d9363/pyarrow-22.0.0-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:a9d9ffdc2ab696f6b15b4d1f7cec6658e1d788124418cb30030afbae31c64746", size = 48066216, upload-time = "2025-10-24T10:07:43.528Z" },
{ url = "https://files.pythonhosted.org/packages/79/87/a1937b6e78b2aff18b706d738c9e46ade5bfcf11b294e39c87706a0089ac/pyarrow-22.0.0-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:ec1a15968a9d80da01e1d30349b2b0d7cc91e96588ee324ce1b5228175043e95", size = 50288552, upload-time = "2025-10-24T10:07:53.519Z" },
{ url = "https://files.pythonhosted.org/packages/60/ae/b5a5811e11f25788ccfdaa8f26b6791c9807119dffcf80514505527c384c/pyarrow-22.0.0-cp313-cp313t-win_amd64.whl", hash = "sha256:bba208d9c7decf9961998edf5c65e3ea4355d5818dd6cd0f6809bec1afb951cc", size = 28262504, upload-time = "2025-10-24T10:08:00.932Z" },
{ url = "https://files.pythonhosted.org/packages/bd/b0/0fa4d28a8edb42b0a7144edd20befd04173ac79819547216f8a9f36f9e50/pyarrow-22.0.0-cp314-cp314-macosx_12_0_arm64.whl", hash = "sha256:9bddc2cade6561f6820d4cd73f99a0243532ad506bc510a75a5a65a522b2d74d", size = 34224062, upload-time = "2025-10-24T10:08:14.101Z" },
{ url = "https://files.pythonhosted.org/packages/0f/a8/7a719076b3c1be0acef56a07220c586f25cd24de0e3f3102b438d18ae5df/pyarrow-22.0.0-cp314-cp314-macosx_12_0_x86_64.whl", hash = "sha256:e70ff90c64419709d38c8932ea9fe1cc98415c4f87ea8da81719e43f02534bc9", size = 35990057, upload-time = "2025-10-24T10:08:21.842Z" },
{ url = "https://files.pythonhosted.org/packages/89/3c/359ed54c93b47fb6fe30ed16cdf50e3f0e8b9ccfb11b86218c3619ae50a8/pyarrow-22.0.0-cp314-cp314-manylinux_2_28_aarch64.whl", hash = "sha256:92843c305330aa94a36e706c16209cd4df274693e777ca47112617db7d0ef3d7", size = 45068002, upload-time = "2025-10-24T10:08:29.034Z" },
{ url = "https://files.pythonhosted.org/packages/55/fc/4945896cc8638536ee787a3bd6ce7cec8ec9acf452d78ec39ab328efa0a1/pyarrow-22.0.0-cp314-cp314-manylinux_2_28_x86_64.whl", hash = "sha256:6dda1ddac033d27421c20d7a7943eec60be44e0db4e079f33cc5af3b8280ccde", size = 47737765, upload-time = "2025-10-24T10:08:38.559Z" },
{ url = "https://files.pythonhosted.org/packages/cd/5e/7cb7edeb2abfaa1f79b5d5eb89432356155c8426f75d3753cbcb9592c0fd/pyarrow-22.0.0-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:84378110dd9a6c06323b41b56e129c504d157d1a983ce8f5443761eb5256bafc", size = 48048139, upload-time = "2025-10-24T10:08:46.784Z" },
{ url = "https://files.pythonhosted.org/packages/88/c6/546baa7c48185f5e9d6e59277c4b19f30f48c94d9dd938c2a80d4d6b067c/pyarrow-22.0.0-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:854794239111d2b88b40b6ef92aa478024d1e5074f364033e73e21e3f76b25e0", size = 50314244, upload-time = "2025-10-24T10:08:55.771Z" },
{ url = "https://files.pythonhosted.org/packages/3c/79/755ff2d145aafec8d347bf18f95e4e81c00127f06d080135dfc86aea417c/pyarrow-22.0.0-cp314-cp314-win_amd64.whl", hash = "sha256:b883fe6fd85adad7932b3271c38ac289c65b7337c2c132e9569f9d3940620730", size = 28757501, upload-time = "2025-10-24T10:09:59.891Z" },
{ url = "https://files.pythonhosted.org/packages/0e/d2/237d75ac28ced3147912954e3c1a174df43a95f4f88e467809118a8165e0/pyarrow-22.0.0-cp314-cp314t-macosx_12_0_arm64.whl", hash = "sha256:7a820d8ae11facf32585507c11f04e3f38343c1e784c9b5a8b1da5c930547fe2", size = 34355506, upload-time = "2025-10-24T10:09:02.953Z" },
{ url = "https://files.pythonhosted.org/packages/1e/2c/733dfffe6d3069740f98e57ff81007809067d68626c5faef293434d11bd6/pyarrow-22.0.0-cp314-cp314t-macosx_12_0_x86_64.whl", hash = "sha256:c6ec3675d98915bf1ec8b3c7986422682f7232ea76cad276f4c8abd5b7319b70", size = 36047312, upload-time = "2025-10-24T10:09:10.334Z" },
{ url = "https://files.pythonhosted.org/packages/7c/2b/29d6e3782dc1f299727462c1543af357a0f2c1d3c160ce199950d9ca51eb/pyarrow-22.0.0-cp314-cp314t-manylinux_2_28_aarch64.whl", hash = "sha256:3e739edd001b04f654b166204fc7a9de896cf6007eaff33409ee9e50ceaff754", size = 45081609, upload-time = "2025-10-24T10:09:18.61Z" },
{ url = "https://files.pythonhosted.org/packages/8d/42/aa9355ecc05997915af1b7b947a7f66c02dcaa927f3203b87871c114ba10/pyarrow-22.0.0-cp314-cp314t-manylinux_2_28_x86_64.whl", hash = "sha256:7388ac685cab5b279a41dfe0a6ccd99e4dbf322edfb63e02fc0443bf24134e91", size = 47703663, upload-time = "2025-10-24T10:09:27.369Z" },
{ url = "https://files.pythonhosted.org/packages/ee/62/45abedde480168e83a1de005b7b7043fd553321c1e8c5a9a114425f64842/pyarrow-22.0.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:f633074f36dbc33d5c05b5dc75371e5660f1dbf9c8b1d95669def05e5425989c", size = 48066543, upload-time = "2025-10-24T10:09:34.908Z" },
{ url = "https://files.pythonhosted.org/packages/84/e9/7878940a5b072e4f3bf998770acafeae13b267f9893af5f6d4ab3904b67e/pyarrow-22.0.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:4c19236ae2402a8663a2c8f21f1870a03cc57f0bef7e4b6eb3238cc82944de80", size = 50288838, upload-time = "2025-10-24T10:09:44.394Z" },
{ url = "https://files.pythonhosted.org/packages/7b/03/f335d6c52b4a4761bcc83499789a1e2e16d9d201a58c327a9b5cc9a41bd9/pyarrow-22.0.0-cp314-cp314t-win_amd64.whl", hash = "sha256:0c34fe18094686194f204a3b1787a27456897d8a2d62caf84b61e8dfbc0252ae", size = 29185594, upload-time = "2025-10-24T10:09:53.111Z" },
]
[[package]]
name = "pycparser"
version = "2.22"
@@ -2969,6 +3191,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/08/20/0f2523b9e50a8052bc6a8b732dfc8568abbdc42010aef03a2d750bdab3b2/python_json_logger-3.3.0-py3-none-any.whl", hash = "sha256:dd980fae8cffb24c13caf6e158d3d61c0d6d22342f932cb6e9deedab3d35eec7", size = 15163, upload-time = "2025-03-07T07:08:25.627Z" },
]
[[package]]
name = "pytz"
version = "2025.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/f8/bf/abbd3cdfb8fbc7fb3d4d38d320f2441b1e7cbe29be4f23797b4a2b5d8aac/pytz-2025.2.tar.gz", hash = "sha256:360b9e3dbb49a209c21ad61809c7fb453643e048b38924c765813546746e81c3", size = 320884, upload-time = "2025-03-25T02:25:00.538Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/81/c4/34e93fe5f5429d7570ec1fa436f1986fb1f00c3e0f43a589fe2bbcd22c3f/pytz-2025.2-py2.py3-none-any.whl", hash = "sha256:5ddf76296dd8c44c26eb8f4b6f35488f3ccbf6fbbd7adee0b7262d43f0ec2f00", size = 509225, upload-time = "2025-03-25T02:24:58.468Z" },
]
[[package]]
name = "pywin32"
version = "311"
@@ -3360,6 +3591,18 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/7c/e4/56027c4a6b4ae70ca9de302488c5ca95ad4a39e190093d6c1a8ace08341b/requests-2.32.4-py3-none-any.whl", hash = "sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c", size = 64847, upload-time = "2025-06-09T16:43:05.728Z" },
]
[[package]]
name = "respx"
version = "0.22.0"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "httpx" },
]
sdist = { url = "https://files.pythonhosted.org/packages/f4/7c/96bd0bc759cf009675ad1ee1f96535edcb11e9666b985717eb8c87192a95/respx-0.22.0.tar.gz", hash = "sha256:3c8924caa2a50bd71aefc07aa812f2466ff489f1848c96e954a5362d17095d91", size = 28439, upload-time = "2024-12-19T22:33:59.374Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/8e/67/afbb0978d5399bc9ea200f1d4489a23c9a1dad4eee6376242b8182389c79/respx-0.22.0-py2.py3-none-any.whl", hash = "sha256:631128d4c9aba15e56903fb5f66fb1eff412ce28dd387ca3a81339e52dbd3ad0", size = 25127, upload-time = "2024-12-19T22:33:57.837Z" },
]
[[package]]
name = "rfc3339-validator"
version = "0.1.4"
@@ -3860,6 +4103,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/17/69/cd203477f944c353c31bade965f880aa1061fd6bf05ded0726ca845b6ff7/typing_inspection-0.4.1-py3-none-any.whl", hash = "sha256:389055682238f53b04f7badcb49b989835495a96700ced5dab2d8feae4b26f51", size = 14552, upload-time = "2025-05-21T18:55:22.152Z" },
]
[[package]]
name = "tzdata"
version = "2025.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/95/32/1a225d6164441be760d75c2c42e2780dc0873fe382da3e98a2e1e48361e5/tzdata-2025.2.tar.gz", hash = "sha256:b60a638fcc0daffadf82fe0f57e53d06bdec2f36c4df66280ae79bce6bd6f2b9", size = 196380, upload-time = "2025-03-23T13:54:43.652Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/5c/23/c7abc0ca0a1526a0774eca151daeb8de62ec457e77262b66b359c3c7679e/tzdata-2025.2-py2.py3-none-any.whl", hash = "sha256:1a403fada01ff9221ca8044d701868fa132215d84beb92242d9acd2147f667a8", size = 347839, upload-time = "2025-03-23T13:54:41.845Z" },
]
[[package]]
name = "uri-template"
version = "1.3.0"

Some files were not shown because too many files have changed in this diff Show More