The model bill is not what breaks a multi-tenant AI product
When people ask what it costs to run an AI product for many customers, they mean the model bill. We wrote that post, and the bill matters.
It is not what took our features down, though.
These are field notes from a multi-tenant AI SaaS we built and still operate: a retrieval-grounded assistant that many businesses embed on their own websites, with crawling, document ingestion, lead capture and a dashboard behind it. Each customer is a tenant. They share one API, one database, one task broker and one model-provider account.
The failures that hurt us came from three places, and none of them shows up on the provider's invoice:
- The connection ceiling. A managed database with 20 connections, and LLM calls that last seconds.
- Head-of-line blocking. One task queue, where a five-minute crawl could sit in front of a login code.
- The tenant key we chose in week one. Four tables stored the tenant ID as text, and we only found out when production queries started failing.
This post covers what each one cost us and how we fixed it, along with the isolation, rate-limiting and traceability design we run today. Every mechanism described here exists in the codebase as written, and we checked each one against the code before publishing. Where a fix is partial, we say so.
What makes multi-tenant AI different from ordinary SaaS
Multi-tenant SaaS architecture is a solved problem in the textbooks: shared infrastructure, a tenant ID on every row, and per-tenant authorisation. A multi-tenant AI agent SaaS architecture breaks three assumptions behind that playbook.
Requests are slow. A normal API request holds resources for tens of milliseconds. An LLM completion holds them for seconds, and an agentic loop or a crawl can hold them for minutes. Any resource you hold during a request (a database connection, a worker process, a broker socket) gets used up about a hundred times faster than your capacity planning assumed.
Every tenant shares one upstream limit. OpenAI's documentation states it plainly: "Rate limits are defined at the organization level and at the project level, not user level." (OpenAI rate limits guide, accessed October 2026). The provider cannot tell your customers apart. Unless you add controls of your own, one tenant's import job spends the token budget that everyone else's live chats depend on.
Consumption is a security category now. The OWASP Top 10 for LLM Applications 2026, published in early August 2026, lists Unbounded Consumption at #6. Excessive Agency moved from #6 to #3, the biggest single change on the list. According to the Cloud Security Alliance's analysis (15 August 2026), this edition is the first to weight real incidents: 6,639 classifiable incidents out of 7,714 reported, counted at 25% of the ranking against 75% for practitioner voting. In a multi-tenant product, unbounded consumption shows up as one customer degrading the service for all the others.
Constraint 1: the connection ceiling
Our database plan allows 20 connections in total. That is the published limit for Heroku Postgres Essential-0 and Essential-1 (Heroku plan documentation, accessed October 2026), and most managed providers have a similar cap on their entry tiers. For comparison, self-hosted PostgreSQL ships with a default of "typically 100 connections" (PostgreSQL docs).
Twenty connections looks like a lot until you multiply it out. Every process has its own pool, and the plan cap has to cover the sum of all of them. This is the budget written into our settings file, in a comment that sits next to the numbers so nobody "tidies" them later:
web = pool_size 5 + max_overflow 5 = 10
worker child = 1 task conn + 1 sync pool + 1 overflow = 3
worker = 3 x 2 prefork children = 6
beat = 0 (schedules only; pools never open)
------------------------------------------------------------
16
headroom for an interactive psql session + release-phase migration = 4
The key comment in that file says: a pool is a reuse cache, not a multiplexer. One pooled connection serves one transaction at a time. Ten connections on the web process means at most ten requests can be in a transaction at once.
Now add an LLM call. An ORM session checks out a connection on the first query and keeps it until the transaction ends. If a chat request reads the conversation, waits four seconds for the model and then writes the reply, it holds a connection for the whole four seconds without running any SQL. With ten connections, that caps you at ten concurrent conversations across the entire product, for every tenant combined.
The fix is a small function that we call before every slow external call on the request path:
async def release_connection(db: AsyncSession) -> None:
"""Hand the pooled connection back before a slow external call."""
if db.in_transaction():
await db.commit()
Committing ends the transaction and returns the connection to the pool. The next query starts a fresh transaction. By our own estimate, a chat request now holds a connection for the roughly 50 ms it actually spends querying, instead of the roughly 4 s it spends waiting on the model. We call it before provider calls in chat, chatbot administration, dashboard insights, listings and the install assistant.
There are two conditions. First, the session must be built with expire_on_commit=False. Otherwise every ORM object loaded before the commit expires, and the first attribute access afterwards triggers a lazy load, which raises an error under asyncio. Second, a commit in the middle of a request makes earlier writes durable. Only release where a partial commit is acceptable.
To keep it honest, the request engine runs with a ten-minute idle_in_transaction_session_timeout. A request path that waits on the model while holding a transaction is a leak, and the database now kills it instead of letting it pin a connection until the process restarts.
PgBouncer is not the fix by itself. We have transaction-mode PgBouncer built and configured, and it is switched off. A pooler lets many app-side connections share a few server-side ones, but only between transactions. A transaction held open across an LLM call pins a PgBouncer server connection exactly as it would pin a direct one. Releasing the connection comes first, and the pooler only helps after that.
Constraint 2: head-of-line blocking
For a long time, all background work in this system shared one Celery queue. That included login one-time codes, chat analysis, website crawls and nightly digests. A crawl takes minutes. Someone waiting for a login code at the wrong moment waited behind it.
We now run four queues, split by how long a person is prepared to wait:
criticallogin codes, invitesprefork workerpriority pollingrealtimechat completion, prompt rebuildsingestioncrawl, embed, parsemaintenancebilling sync, digests, sweeps
Three decisions make this work in practice.
Routing lives in exactly one place. A single task-to-queue map, plus module-level wildcard patterns as a fallback, so a new task cannot silently end up on the wrong queue. Passing queue= to a task decorator or to apply_async is banned. Our enqueue helper raises an error if it sees one. A hard-coded queue argument silently overrides the routing table, and that is how everything ended up on one queue in the first place. Unrouted tasks default to maintenance, not critical: an unclassified task is far more likely to be a cron job than a login code, and filing a cron job as urgent is the more expensive mistake.
Priority polling, with the limits written down. The Redis transport polls queues in a fixed order (queue_order_strategy="priority"), so a free worker always checks critical first. The code comment is explicit about what this does not do: it does not preempt a running task. The worker runs two prefork children. One long crawl still leaves a child free for login codes. Two concurrent crawls can still make a login code wait.
The real guarantee needs a second worker process pinned to -Q critical,realtime. Celery's own optimisation guide says the same thing: "If you have a combination of long- and short-running tasks, the best option is to use two worker nodes that are configured separately." We have not done this yet, because a second worker's broker connections do not fit within the roughly 20-connection cap of our entry-tier Redis plan. That is the connection ceiling again, this time at the broker. It is a known gap and it is written down next to the queue definitions.
Each queue gets its own failure semantics. The global config uses acks_late=True with task_reject_on_worker_lost, so a crash re-delivers the task, plus worker_prefetch_multiplier=1. Ingestion tasks are the exception: they acknowledge on receipt. A crawl that hangs inside a C extension gets killed by the hard time limit. Under acks_late it would be re-delivered, killed again, and occupy a child indefinitely without ever recording a failure. Turning acks_late off is safe for ingestion only because ingestion already has recovery sweeps that re-queue stuck jobs every 10 to 15 minutes. Login codes and billing writes have no sweep behind them, so they keep at-least-once delivery.
Constraint 3: the tenant key you chose in week one
This is the one we most wish we had got right on day one.
Every tenant-scoped table had a proper BIGINT foreign key to tenants.id, except four: conversations, conversation_insights, lead_notes and email_logs. Those stored the tenant as VARCHAR, holding str(tenants.id). Nobody decided this. It drifted in because one code path stringified the ID and another did not.
It cost us in three ways.
Joins failed outright instead of degrading. Any query that joined one of those tables to tenants needed an explicit cast. The queries that missed the cast did not get slower. They failed with:
operator does not exist: character varying = bigint
That error took down two super-admin analytics endpoints, and the same bug was waiting in a third. In the migration's words: "A cast convention that must be re-derived at every call site is not a convention, it is a trap; the type is the fix."
The missing foreign key hid problems. Nothing stopped a conversation from referencing a tenant that did not exist. Nothing cascaded when a tenant was deleted. The query planner had no join selectivity to work with.
Fixing it was not a rolling deploy. ALTER COLUMN ... TYPE rewrites the table and every index on it under an ACCESS EXCLUSIVE lock, and conversations is the largest table on the platform. Worse, the schema and the application were incompatible in both directions. Old code binds a string, which a BIGINT column rejects. New code binds an integer, which the old VARCHAR column rejects just as firmly. The platform runs migrations while the old processes are still serving traffic, so there is a window in which every conversation query fails. The only safe procedure is to turn on maintenance mode, snapshot, migrate, deploy, and turn maintenance off. The migration's own docstring warns that running downgrade() alone does not restore service, because the old code has to be redeployed with it.
The migration itself is the part worth copying. It refuses to run rather than damage data it does not understand. Before the first ALTER, it scans all four tables for non-numeric tenant values and for references to tenants that no longer exist. If it finds any, it aborts with the actual bad values in the error message. There is one deliberate exception. email_logs is a delivery record that has to outlive the tenant it was sent for, so its foreign key is ON DELETE SET NULL and orphaned rows there are detached and counted instead of blocking the deploy. The other three use CASCADE.
What we would tell any team on day one:
- Use one tenant key type everywhere, enforced by a foreign key. The chat code now carries a comment saying so at the point where the tenant ID is resolved.
- Use two identifiers, not one. Our tenants have an internal
BIGSERIALused for every join and never exposed by any API, and a public UUIDv7 that appears in URLs and responses. Joins stay cheap, and an outsider cannot enumerate tenants. - Put the tenant ID on every row, including derived rows. Our document chunks carry
tenant_iddirectly, even though it can be derived through document → knowledge source → chatbot. Retrieval then never depends on a chain of joins being right.
Tenant isolation is a data-model decision
The most important isolation property in an AI product is that retrieval cannot return another tenant's knowledge. If a prompt contains another customer's documents, no guardrail on the output will catch it.
We use the most common pattern: one shared pgvector table and one HNSW index, with every query filtered by tenant. What makes this safe is not the filter. It is that there is exactly one retrieval function, and its signature requires both tenant_id and chatbot_config_id. The scope clause is built once and applied to both retrieval legs (dense vector and Postgres full-text), which are then merged with reciprocal rank fusion. Nobody writes their own vector query elsewhere in the codebase.
We do not use row-level security. Isolation is enforced in the application, through tenant-scoped dependencies that resolve the tenant and the caller's membership before any handler runs. That is a legitimate choice. It does mean that a handler which skips the dependency can expose data across tenants, so authorisation coverage is something we audit, not something we assume. If you need a hard guarantee for a regulated tier, Postgres RLS with a transaction-local tenant setting moves the check into the database.
A trap that is about relevance, not isolation. pgvector's documentation notes that "with approximate indexes, filtering is applied after the index is scanned." With the default hnsw.ef_search of 40 and a filter matching 10% of rows, you get about four results on average (pgvector README). In a shared index, a small tenant is a highly selective filter. The answer is still correctly isolated, but it can be badly incomplete. pgvector 0.8.0 added iterative index scans (SET hnsw.iterative_scan = strict_order) to fix exactly this, and partial indexes or partitioning are the alternatives. We have not enabled iterative scans yet. Our full-text leg hides some of the recall loss, which is not the same as having no recall loss. It is on our list.
Noisy neighbours: per-tenant rate limiting for LLM calls
Since the provider rate-limits the whole organisation, fairness between tenants is our responsibility. We built an admission-control layer that decides whether to spend tokens before calling the provider, instead of learning the answer from a 429.
It has three controls:
- Lanes. Interactive traffic (a visitor waiting on a widget reply) and background traffic (crawl, OCR, insights, digests) each get their own tokens-per-minute budget, carved from the account ceiling. A crawl storm cannot starve a live conversation.
- A per-tenant budget, set as a share of the interactive lane rather than as an absolute number. With the share at one half, no single customer can hold more than half of the platform's live capacity. Because the number scales with the account tier, signing your eleventh tenant does not mean retuning it.
- A concurrency ceiling on in-flight completions. Before this, chat completions had no limit other than an HTTP pool of 1,000, so extra load turned straight into 429s.
Each call reserves an estimate (prompt characters ÷ 4, plus the expected completion) and settles to the real usage when the response returns. The expected completion matters. Our widget replies are measured at 100 to 200 tokens against a truncation cap of 800, so reserving the cap was wasting a large share of every tenant's per-minute budget.
The war story behind the per-tenant budget: the first version was a fixed 15,000 TPM per tenant. A turn deep into a conversation reserves about 6,000 tokens once you include history, summary and retrieved chunks. That works out to roughly two turns per minute for the entire tenant, and a single visitor typing at a normal pace broke it. The settings file now keeps that table in a comment, and a test enforces that one visitor's own rate limit always stays a small fraction of the tenant budget.
Some honest notes on the design:
- We use a fixed one-minute window, keyed on the wall-clock minute, because that is how the provider accounts for TPM. A burst that straddles two windows can briefly exceed the limit. We accept that because our limits sit below the provider's real ceiling on purpose.
- Everything fails open. If Redis is down, calls are admitted. The provider's own 429 is the backstop, and a metering glitch must never take a widget down.
- Quotas gate the start of a conversation, never a message within one. Monthly conversation and token quotas are checked before the first turn. A quota can refuse the next visitor, but it never cuts off one who is mid-sentence. The overrun this allows is bounded by a per-conversation message cap.
- The platform-wide daily budget alerts but does not block live chat. It used to block, and one heavy day switched off every widget, including those of customers nowhere near their own limits. Now background work sheds at a fraction of the daily budget, interactive traffic keeps being served, and a human gets paged.
Failure isolation: two circuit breakers, on purpose
Around the model provider, a circuit breaker stores its state in Redis, so the web process and every worker share one trip decision. Five failures open it, and it probes again after 60 seconds. An exhausted billing quota on the provider side is recorded as a failure straight away. Every request after that would fail the same way until a human acts, so failing fast is kinder than making each visitor wait out a timeout. If Redis is unreachable, the breaker fails open.
Around the task broker, that design cannot work. A Redis-backed breaker cannot record an outage of Redis itself. So enqueueing on the request path goes through a safe_enqueue helper with a separate, in-process breaker. It publishes from a worker thread so a hung broker cannot block the event loop. It never raises on infrastructure failure, because the database write it follows has already been committed. It stops trying after three failures and probes with exactly one request after the recovery window. It is used only for work that something else will pick up, such as a recovery sweep or the next turn of a conversation.
Before this existed, a degraded Redis could take down the API. The web process inherited the worker's 120-second socket timeout and had a single broker connection, so one hung publish stalled every request on that process. The web process now has its own three-second publish timeout and retries a publish once.
Following one request across tenants and queues
You cannot debug a multi-tenant async system from logs that only say "task failed".
When a request resolves its tenant, the tenant's public ID is written into a logging context variable, the request state and the Sentry scope. From that point, every log line in the request carries it. When the request enqueues a task, a before_task_publish signal copies the request ID, tenant ID and user ID into the message headers. In the worker, task_prerun restores them, so the task's log lines share the HTTP request's correlation ID. Cron tasks get a minted task-<id> instead. task_postrun resets the context so an idle worker never logs a stale tenant.
Permanent failures write a dead-letter row containing the task name, arguments, traceback, the originating request ID and the tenant. The super-admin console can then answer "which tenant lost which job, triggered by what" and replay it. The request ID is also exposed to browsers as a CORS header, so a support ticket can quote it.
Two smaller decisions that pay off
Tenant-versioned caching. Dashboard aggregates are cached under keys that include a per-tenant version counter. Any CRM or campaign write runs a single INCR on that counter, which makes all of the tenant's cached aggregates unreachable immediately, with no key scanning. Old entries age out through their TTL. On a cache miss, a short per-key Redis lock means only one request recomputes while the others wait briefly, which prevents a stampede when a hot key expires. If the lock fails, callers compute live, so cache errors never break a request.
Two-layer CORS. The dashboard is cookie-authenticated, so it needs credentialed CORS. That only works with an explicit allowlist. The embedded widget runs on every customer's own domain, so those origins cannot be listed in advance. We run two middlewares. A public layer reflects any origin, but only on a fixed list of public widget routes and never with Allow-Credentials. The authorisation decision stays in the service layer, which checks each chatbot's allowed domains and returns a readable 403. A credentialed layer handles everything else. One subtle detail: Starlette's CORS middleware adds Allow-Credentials even to responses whose origin it rejected, so the public layer strips that header before reflecting the origin. Otherwise you end up with any-origin plus credentials, which this split exists to prevent.
Where agents and tools change the picture
Our assistant mostly retrieves and converses. Products that give an agent tools inherit the risk OWASP now ranks #3. OWASP splits Excessive Agency into excessive functionality, excessive permissions and excessive autonomy. In a multi-tenant system, the dangerous form is a tool that runs with platform credentials when it should run with this tenant's credentials.
The Model Context Protocol made this explicit in its 2026-07-28 authorization specification: servers "MUST validate that access tokens were issued specifically for them as the intended audience" and "MUST NOT accept or transit any other tokens." Token passthrough is the confused-deputy problem, and in a multi-tenant product the confused deputy acts across tenant boundaries. Apply the same discipline we apply to retrieval: one entry point per tool, a tenant scope required by its signature, and credentials resolved per tenant at call time, never from a shared environment variable.
The checklist
Run through this before your second customer, not your hundredth:
- [ ] Write the connection budget down, per process type, summed against the plan cap, with headroom for migrations and a psql session.
- [ ] Release the database connection before every LLM call on the request path, and add an idle-in-transaction timeout to catch the ones you missed.
- [ ] Split queues by latency budget, route them in one place, and ban hard-coded
queue=arguments. - [ ] Write down what priority polling does not guarantee, and plan the dedicated worker for urgent queues.
- [ ] Use one tenant key type, enforced by a foreign key, on every tenant-scoped table, including derived ones like chunks.
- [ ] Keep an internal integer ID for joins and a public opaque ID for the API.
- [ ] Have one retrieval function whose signature requires the tenant scope.
- [ ] Check vector recall for small tenants in a shared ANN index (iterative scans, partial indexes or partitions).
- [ ] Separate lanes for interactive and background LLM traffic, with a per-tenant share of the interactive lane.
- [ ] Reserve tokens before the call and settle after it. Estimate the expected completion, not the cap.
- [ ] Gate quotas at the start of a conversation, not mid-conversation.
- [ ] Use a shared circuit breaker for the provider and an in-process one for the broker.
- [ ] Propagate request, tenant and user IDs into task headers, and write them onto dead-letter rows.
- [ ] Scope cache invalidation per tenant with a version counter, plus single-flight on misses.
- [ ] Keep the CORS layers separate for credentialed first-party routes and any-origin public routes.
- [ ] Resolve tool credentials per tenant and never pass a token through.
If you are designing a multi-tenant AI product and would rather learn these lessons from someone else's migration window, talk to our engineering team. Everything above comes from a system we run ourselves.
Frequently asked questions
What is a multi-tenant AI SaaS architecture? It is a single deployment of an AI product (API, database, task queues, vector store and model-provider account) shared by many customer organisations, with each customer's data, capacity and failures isolated from the others. The AI-specific parts are isolating retrieval and context per tenant and sharing one upstream model rate limit fairly.
How do you isolate tenant data in a RAG system? Put the tenant ID on every stored chunk, and route all retrieval through one function whose signature requires the tenant scope, so no query can be written without it. Shared indexes with a tenant filter are the most common and cheapest pattern. Postgres row-level security or per-tenant partitions give stronger guarantees. Also check recall for small tenants, because approximate indexes filter after the scan.
How do you stop one tenant's AI usage from affecting others? Enforce admission control before the provider call. Use separate token budgets for interactive and background traffic, cap each tenant at a share of the interactive budget, and add a concurrency ceiling. Reserve an estimate up front and settle to actual usage afterwards. Provider rate limits apply to your whole organisation, so the provider will not do this for you.
Why do database connections run out in AI applications? Because an LLM call holds a request open for seconds, and an ORM session holds its database connection for the whole transaction. With a small pool, concurrency collapses to the pool size. Commit or release the connection before the model call. A transaction-mode pooler such as PgBouncer only helps after you have done that.
Should AI background jobs share a queue with user-facing tasks? No. Split queues by how long a person is prepared to wait, and route them in one place. Priority polling helps, but it does not preempt a running task, so a guarantee for urgent work needs a dedicated worker process that long tasks cannot occupy.
How do you trace a request across Celery workers in a multi-tenant app? Copy the request ID, tenant ID and user ID into task message headers when publishing, restore them into the logging context when the task starts, and reset them afterwards. Record the same IDs on permanent-failure (dead-letter) rows, so a lost job can be traced back to the tenant and request that caused it.
Is it expensive to change the tenant ID type later? Yes. Changing a column type rewrites the table under an exclusive lock, and old and new code are usually incompatible with the new type in both directions, so it requires a maintenance window rather than a rolling deploy. Pick an integer foreign key for joins and a separate public identifier on day one.
Does OWASP cover multi-tenant AI risks? Indirectly. The OWASP Top 10 for LLM Applications 2026 places Excessive Agency at #3, Unbounded Consumption at #6 and Vector and Embedding Weaknesses at #9. In a multi-tenant product, each of these can turn into one tenant affecting another, through over-privileged tools, shared capacity or shared embedding stores.
Sources & further reading
- OWASP Top 10 for LLM Applications 2026: OWASP GenAI Security Project, August 2026
- OWASP's 2026 LLM Top 10: incident data meets judgment: Cloud Security Alliance, August 2026
- MCP authorization specification, revision 2026-07-28: Model Context Protocol, July 2026
- Rate limits: OpenAI API documentation, accessed October 2026
- Heroku Postgres plans: Heroku Dev Center, accessed October 2026
- Connections and authentication: PostgreSQL documentation
- Optimizing: prefetch limits: Celery documentation
- pgvector: filtering and iterative index scans: pgvector
- Multi-tenant generative AI platform scenario: AWS Well-Architected Generative AI Lens
From us
- What an AI Agent Actually Costs to Run in Production: the token-economics side of the same system
- Why 40% of AI Agent Projects Get Cancelled, and What We Did Differently: the broader engineering disciplines
- Talk to our engineering team
