> ## Documentation Index
> Fetch the complete documentation index at: https://daily-main.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# File Resolution and Storage

> Resolve send-file URLs into bytes with FileResolver and store client uploads behind pluggable FileStorage backends like LocalFileStorage.

## Overview

A file sent to the bot in an RTVI [`send-file`](/api-reference/server/rtvi/rtvi-processor#send-file) message lands in the LLM context as a URL. That URL can take several forms: `pipecat:<id>` for a file uploaded through the [development runner's upload endpoints](/api-reference/server/utilities/runner/guide), `s3://` or `gs://` for a file in cloud storage, or a plain `http(s)` link. Some of these URLs can be passed to the LLM provider as-is; others must first be fetched by the server and inlined as bytes — a decision that depends on the provider and on who can reach the URL. Two classes in `pipecat.utils` split that work:

* **`FileStorage`** (`pipecat.utils.file_storage`) is the storage half: the interface behind the runner's upload endpoints, which saves client uploads and mints the URLs that come back in `send-file` messages.
* **`FileResolver`** (`pipecat.utils.file_resolver`) is the fetching half: configured on an LLM service via its [`file_resolver`](/pipecat/learn/llm#base-class-configuration) parameter, it downloads the bytes right before a completion when the provider can't fetch the URL itself — loading stored uploads back through the storage backend.

Passing the same storage instance to both sides — which the development runner does for you via `runner_args.file_storage` — means the upload endpoint and the resolver agree on where uploads live.

## FileStorage

The abstract interface between an upload endpoint, which saves uploads and hands the resulting URL back to the client, and the `FileResolver`, which loads the bytes back when the client references that URL in a `send-file` message.

The URL is the contract: `save()` returns one, callers pass it back unmodified, and each implementation mints whatever form it can later resolve — `pipecat:<id>` for `LocalFileStorage`, or e.g. `gs://`/`s3://` for a custom cloud backend (which the LLM provider may then even fetch itself, without `load()` being involved at all).

### Interface

<ParamField path="save(filename, contents)" type="async → str">
  Store a file and return the URL later passed to `load()` or `delete()`.
  Implementations should not derive the returned URL from `filename`.
</ParamField>

<ParamField path="load(file_url)" type="async → bytes">
  Return the stored contents for a URL previously returned by `save()`. Raises
  `FileNotFoundError` if the URL is invalid or no longer stored.
</ParamField>

<ParamField path="delete(file_url)" type="async → None">
  Remove the stored file, if present.
</ParamField>

<ParamField path="delete_after_load" type="class attribute, bool" default="False">
  Whether a file should be deleted once its bytes have been consumed (loaded and
  inlined into the LLM context by a `FileResolver`). Safe only when every URL
  the backend resolves is an upload it owns; leave `False` for backends whose
  URLs may point at shared data (e.g. arbitrary objects in a cloud bucket),
  where "read into a conversation" must not mean "destroy".
</ParamField>

### LocalFileStorage

The default backend: files on local disk, addressed as `pipecat:<hex>`. Suitable for local development and single-process deployments. The random hex suffix can't be used for path traversal and doesn't leak the original filename. Sets `delete_after_load = True`, since every `pipecat:` URL is an upload it owns; deleting an already-missing file is a no-op, so a consumed upload may safely be deleted again by session-end cleanup.

<ParamField path="folder" type="str">
  Directory to store files in. Created on first save if missing.
</ParamField>

<ParamField path="max_files" type="int" default="10">
  Maximum number of files to retain; oldest files (by mtime) are deleted once
  this is exceeded. Set to `0` to disable trimming.
</ParamField>

The development runner constructs one for you from `-u/--uploads-folder` and `--uploads-folder-max-files`.

### Custom backends

To store uploads somewhere other than local disk, implement the interface and — when using the development runner — return it from a synchronous `create_file_storage()` function defined next to your `bot()`. The runner installs it behind the upload endpoints and injects it into the bot as `runner_args.file_storage`:

```python theme={null}
class BucketFileStorage(FileStorage):
    async def save(self, filename: str, contents: bytes) -> str:
        blob_name = f"uploads/{uuid.uuid4().hex}"
        await upload_to_bucket(MY_BUCKET, blob_name, contents)
        return f"gs://{MY_BUCKET}/{blob_name}"

    async def load(self, file_url: str) -> bytes:
        return await download_from_bucket(file_url)

    async def delete(self, file_url: str) -> None:
        await delete_from_bucket(file_url)


def create_file_storage():
    return BucketFileStorage()


async def bot(runner_args: RunnerArguments):
    llm = OpenAILLMService(
        api_key=os.getenv("OPENAI_API_KEY"),
        file_resolver=FileResolver(file_storage=runner_args.file_storage),
    )
    ...
```

See the [runner guide](/api-reference/server/utilities/runner/guide) for the upload endpoints, CLI flags, and session-end cleanup behavior.

## FileResolver

Fetches and caches the bytes behind file URLs the LLM provider can't fetch itself.

```python theme={null}
from pipecat.utils.file_resolver import FileResolver

llm = OpenAILLMService(
    api_key=os.getenv("OPENAI_API_KEY"),
    file_resolver=FileResolver(file_storage=runner_args.file_storage),
)
```

### Constructor

<ParamField path="file_storage" type="FileStorage" default="None">
  Storage backend used to resolve URLs it minted (e.g. `pipecat:<id>` from the
  development runner's upload endpoints — pass `runner_args.file_storage`
  there). Without it, only `data:` and `http(s)` URLs can be resolved.
</ParamField>

<ParamField path="allowed_url_networks" type="list[str]" default="None">
  CIDR ranges (e.g. `["10.0.0.0/8"]`) this server is trusted to fetch from, in
  addition to the public internet. An `http(s)` URL that resolves outside both
  is refused. Defaults to the `PIPECAT_ALLOWED_FILE_URL_NETWORKS` environment
  variable (comma-separated).
</ParamField>

<ParamField path="max_fetch_bytes" type="int" default="52428800">
  Maximum size of a fetched file. Defaults to 50 MB.
</ParamField>

<ParamField path="fetch_timeout_secs" type="float" default="30">
  Total timeout for an `http(s)` fetch.
</ParamField>

### URL handling

The resolver handles three kinds of URL:

* **`data:` URLs** are decoded directly.
* **`http(s)` URLs** are fetched by this server, but only when the resolved address is publicly routable or falls within `allowed_url_networks` — a client-supplied URL must not be able to point your server at arbitrary private address space (SSRF). Redirects are refused rather than followed, and downloads are capped at `max_fetch_bytes`.
* **Any other scheme** is delegated to `file_storage`, which resolves the URLs it minted (`pipecat:<id>` for `LocalFileStorage`; a custom backend resolves whatever its `save()` returned, e.g. `gs://` with the deployment's own credentials).

A URL that can't be resolved — refused by the reachability policy, over the size limit, failed to download, or a scheme nothing is configured for — raises `FileResolverError`. During context conversion this surfaces as `LLMContextConversionError`, and the offending file message is removed from the context so later completions aren't blocked.

<Note>
  URLs the provider consumes directly — public `http(s)` URLs, `s3://` on AWS
  Bedrock, `gs://` on Google Vertex — never reach the resolver's fetch path: the
  adapter passes them through and the provider fetches them itself.
</Note>

### Caching and sharing

Each URL is fetched once and served from an in-memory cache afterwards: a URL is assumed to identify one immutable piece of content. Content that changes should live at a new URL (uploads mint a fresh URL per upload; a cache-busting query string works for external URLs), or call `forget(url)` when your application knows a URL's content changed.

The caches make resolved files provider-neutral, which drives two scoping rules:

* **Share one instance across LLM services** that may consume the same files (e.g. services switched between mid-session), so they reuse each other's downloads. This matters doubly for `delete_after_load` backends: the stored copy is deleted once the bytes are cached, so the cache entry becomes the surviving copy.
* **Scope a resolver to one session.** An instance shared across sessions would hold every session's files in memory and stretch the one-fetch-per-URL assumption across all of them.

### Methods

<ParamField path="fetch(url)" type="async → bytes">
  Return the bytes behind `url`, fetching and caching them on first use. A
  storage-resolved file is deleted after a successful load when the backend sets
  `delete_after_load`.
</ParamField>

<ParamField path="ensure_data_url(url, mime_type)" type="async → str">
  Return the `data:` form of `url`'s bytes, encoding and caching it on first use
  — for adapters whose provider consumes base64 rather than raw bytes.
</ParamField>

<ParamField path="classify(url)" type="async → UrlReachability">
  Classify who can reach an `http(s)` URL, honoring `allowed_url_networks`.
  Classified once per URL; later calls return the cached verdict.
</ParamField>

<ParamField path="forget(url)" type="None">
  Drop everything cached for `url`, forcing a re-fetch on next use. A forgotten
  `delete_after_load` upload can't be re-fetched — its stored copy was deleted
  when first cached.
</ParamField>

<ParamField path="cached_bytes(url) / cached_data_url(url) / cached_reachability(url)" type="… | None">
  Synchronous cache reads, returning `None` when nothing is cached for `url`.
</ParamField>

<ParamField path="file_storage" type="property → FileStorage | None">
  The storage backend used to resolve URLs it minted, if any.
</ParamField>

Subclass and override `fetch()` to support additional schemes or credentialed fetches beyond what the storage backend provides.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.