# Payload Limits, Rate Limits & Timeouts

> Operational ceilings for the five benchmarked platforms: request/batch payload
> caps, rate-limit enforcement and timeout budgets, with per-vendor backoff guidance.
Last verified: 2026-10-09

Figures are consolidated from vendor documentation during automated extraction and
are **indicative planning values** — always confirm against the vendor limits page
linked per row before capacity planning. The machine-readable summary also lives in
the [comparison matrix](/matrix.md) (section 4).

## 1. Request payload & batch limits

| Vendor | Request payload | Batch / ingest guidance | Metadata ceiling | Limits doc |
| :-- | :-- | :-- | :-- | :-- |
| [Pinecone](/go/pinecone) | ~100 MB hard ceiling per upsert; keep queries well under 2 MB | Upsert in batches of ~100–1,000 vectors; one vector per `(id)` tuple | ~40 KB total metadata per vector | [docs](https://docs.pinecone.io/) |
| [Qdrant Cloud](/go/qdrant) | Large JSON point payloads supported; chunk uploads above ~32 MB per point | gRPC/REST `upsert` with arbitrary batch sizes; shard-aware parallel writes | Payload size limited by max request size, not per-field caps | [docs](https://qdrant.tech/documentation/) |
| [Weaviate Cloud](/go/weaviate) | No fixed per-request ceiling documented; batch endpoint auto-chunks | `batch` API auto-tunes concurrency; enable `dynamic batching` for bulk loads | Object properties sized by request body limit | [docs](https://docs.weaviate.io/) |
| [Milvus (Zilliz Cloud)](/go/milvus) | ~64 MB per gRPC/REST message (server default, configurable) | `bulk_insert` in segments; avoid single messages beyond the gRPC limit | Scalar fields stored per-entity; JSON field limits apply | [docs](https://docs.zilliz.com/) |
| [Supabase pgvector](/go/supabase) | Standard Postgres/PostgREST limits — size batches by statement memory | `INSERT ... vectors` batches of 1k–10k rows; use `COPY` for large loads | Row/JSON limits of Postgres (~2 KB/page per TOAST row practical) | [docs](https://supabase.com/docs) |

## 2. Rate limits & backoff

| Vendor | Enforcement signal | Typical trigger | Client strategy |
| :-- | :-- | :-- | :-- |
| [Pinecone](/go/pinecone) | HTTP `429` + `Retry-After` | Per-second request ceilings of the pod/scale unit | Exponential backoff on 429; scale replicas for sustained RPS |
| [Qdrant Cloud](/go/qdrant) | HTTP `429` + `Retry-After` | Plan-based requests/second with burst headroom | Jittered backoff; prefer connection reuse over new TLS handshakes |
| [Weaviate Cloud](/go/weaviate) | HTTP `429` + `Retry-After` | Plan-based request quotas (queries per minute) | Back off on 429; batch writes instead of N single inserts |
| [Milvus (Zilliz Cloud)](/go/milvus) | HTTP `429` / resource-unit exhaustion | Serverless request-unit (RU) quotas | Pace ingestion to RU budget; exponential backoff with jitter |
| [Supabase pgvector](/go/supabase) | HTTP `429` / connection queueing | Per-plan API request throttling and Postgres connection limits | Pool via PgBouncer; queue heavy jobs (pgmq/graphile-worker) |

Universal pattern: treat `429` as the only throttle signal, honour `Retry-After`,
apply exponential backoff with full jitter, and cap in-flight requests per worker.

## 3. Timeouts

| Vendor | API / query timeout | Long-running operations | Guidance |
| :-- | :-- | :-- | :-- |
| [Pinecone](/go/pinecone) | ~60 s HTTP request timeout | Upserts stream; search is synchronous | Split large imports; keep top-k queries in low tens of ms by design |
| [Qdrant Cloud](/go/qdrant) | Per-request timeout parameter; no fixed server cap | Scroll/pagination for large scans | Set explicit client deadlines; page through `scroll` cursors |
| [Weaviate Cloud](/go/weaviate) | Configurable server-side query timeout | Batch imports are async from the client's view | Bound GraphQL queries with `timeout` argument |
| [Milvus (Zilliz Cloud)](/go/milvus) | No fixed server timeout; ~60 s client deadline recommended | Index building is asynchronous (task polling) | Poll `index state` rather than blocking on DDL calls |
| [Supabase pgvector](/go/supabase) | PostgREST statements bounded by `statement_timeout` | Index builds (`CREATE INDEX CONCURRENTLY`) run outside transactions | Move embeddings generation to Edge Functions/queue; use iterative scans for filtered ANN |

## 4. Load-test checklist

1. Warm the index (populate vectors, run a few queries) before measuring p50/p95.
2. Measure with **and** without metadata filters — filtering stage dominates tail latency.
3. Push one dimension at a time: QPS ramp, batch size ramp, payload size ramp.
4. Record the `429` threshold per plan; it is part of the effective throughput spec.
5. Keep the client deadline below the platform's HTTP timeout to avoid retry storms.

See also: [comparison matrix](/matrix.md) · [payload/limits source JSON](/content/index.md) · [llms.txt](/llms.txt)
