The two performance tiers
Reads and writes are not in the same band. Reads run about 400–1,000 RPS because databases scale well with concurrent readers. Writes fall to roughly 150–400 RPS because each one awaits lock constraints, a transaction commit, and index updates. A mixed 50/50 workload sits closer to 300–600 than to either extreme, and if your traffic is write-heavy you should budget for the low end.
Two questions decide which number is yours: which database engine (PostgreSQL, MySQL, SQLite) and whether you’re on an ORM like SQLAlchemy or SQLModel. An ORM plus a write-heavy workload plus a 1-core pod is the worst combination in the table above, and it lands you near 150 RPS.
What actually runs your request
A singleuvicorn worker is one OS thread running one event loop on one core. That one sentence explains most of the tuning:
- The loop is single-threaded. A handler that returns control quickly costs almost nothing in wall-clock terms; a handler that holds it is a stop-the-world event for every other in-flight request.
- The GIL is not your bottleneck at this scale — the loop is. With one core, multiple processes buy you nothing.
uvloop(installed withuvicorn[standard]) swaps the stdlib selector loop for a libuv-based one and typically buys 10–30 percent more loop iterations per second. It’s free, so take it, but it won’t rescue a serialization-bound app.- A sync
defendpoint does not run on the loop. FastAPI hands it to the Starlette/anyio threadpool, default capacity 40 worker threads — a ceiling you hit long before the core’s. - Worker count equals core count. On a 1-core instance that means exactly one Uvicorn worker; more workers on one core force aggressive context-switching and measurably degrade throughput.
min(CPU-time-per-request, event-loop overhead, pool concurrency, DB latency), and on a 1-core box the CPU term almost always wins.
Sync vs async under load: I measured it
I built a route that simulates 50 ms of blocking I/O twice — once asasync def with await asyncio.sleep(0.05), once as plain def with time.sleep(0.05) — and offered it 64 concurrent requests.
The arithmetic predicts this almost exactly. Async: 64 concurrent slots at 52.5 ms each is about 1,217 RPS — the loop never waits, so throughput scales linearly with offered concurrency. Sync: the threadpool only has 40 slots, so 64 in-flight requests means each thread runs 1.6 requests in series, which is precisely the p50 jump from 50 ms to 86 ms, and precisely the 40 / 0.05 s = 800 RPS theoretical cap the measured 741 sits just under.
So the honest framing is this: async does not make a single request faster. It makes the server’s concurrency the thing you get to choose. At low offered load (8 concurrent per client, three clients) sync and async were indistinguishable at 145–150 RPS per client, and the sync server cost slightly more memory — 32.5 MB resident versus 29.0 MB — because those threadpool threads have stacks. The gap only opens when you push past 40 simultaneous requests, and then it opens hard.
The corollary is the trap: if you call a sync driver (
psycopg2, sqlite3, requests) from an async def handler, you block the one loop and every other request queues behind it. My /mixed endpoint does 2,000,000 integer adds inside an async handler for exactly this reason — pure CPU work in a coroutine is the same footgun as blocking I/O there.
The serialization tax
In FastAPI the database query is usually not the slow part — it’s non-blocking network I/O, and the loop happily runs other requests while it’s outstanding. The hidden cost is Python converting response rows into models and then into JSON, on a single pegged core. I measured the three serialization paths directly, on a 500-row result set, as pure CPU time per response:
At 20 rows the difference is below my noise floor — my A/B of the small-payload hot path versus the validated path gave 540 vs 502 RPS in one run and 429 vs 478 in another, i.e. no signal. At 500 rows it dominates.
This is where I have to be careful with the standard advice, “drop
response_model to save CPU.” The direction is right — do not validate large payloads through Pydantic on a hot path — but the reason is subtler than “Pydantic is slow.” With no response_model, FastAPI still walks your return value through jsonable_encoder, and that pure-Python recursive walk is slower than Pydantic v2’s compiled Rust serializer. Returning raw dicts doesn’t remove the tax; it changes who charges it. The only path that skips conversion is returning a Response you built yourself — row 3 of that table, 8x cheaper than row 1. Use ORJSONResponse on genuinely hot routes, and response_model everywhere a contract and an OpenAPI schema are worth more than 0.6 ms.
The pool, the connection count, and the write path
Pool sizing is a concurrency budget, not a memory dial. Set it statically at startup (in thelifespan handler) so you never pay connection establishment mid-request: something like min_size=10, max_size=20 for asyncpg is a sane starting point for a 1-core service, and the corresponding SQLAlchemy/SQLModel knobs are pool_size plus max_overflow on create_async_engine with async_sessionmaker.
The relationship to hold: worker processes × pool max size = database connections. One worker with max_size=20 uses 20 connections. Six workers × 20 = 120. The Postgres default max_connections=100 is a wall you will hit by scaling horizontally without shrinking pools.
Using raw SQL with asyncpg is up to 3x faster than traditional sync drivers, because you remove both the threadpool hop and the ORM’s row-instantiation layer. That is the single biggest lever in the “Optimized” tier. On the write side, a normal single-row insert through FastAPI costs 1–2 ms per insert, which is expected rather than pathological — it’s the round trip and commit, not your Python.
One warning from the field that matches my own curve: a team reported their database topping out at 80% CPU and 80% RAM with roughly 300 connections, with connection pooling already in place. Beyond a few hundred open connections, the DB’s per-connection memory and scheduling cost becomes the limiter, and adding pool size makes it worse instead of better. If your engine can’t serve more concurrent readers, a bigger pool just moves the queue.
How to measure this without lying to yourself
1
Warm up, then time
Run 50–100 throwaway requests first. Bytecode caches, Pydantic core-schema construction, DNS, TCP handshakes and pool fills all land on the first requests and otherwise show up as p99.
2
Know that a closed loop hides latency
A closed-loop generator (N workers, each firing the next request only when the last completes) measures capacity. As the server slows, offered load drops with it and latency damage disappears — coordinated omission. Add a fixed-rate run if you care about p99 at peak.
3
Ramp concurrency, don't set it once
Sweep 1, 8, 24, 64 concurrent and record RPS plus p50, p95, p99 for each. You’re looking for the knee: where RPS stops rising and p99 starts climbing. Serve at 60–70 percent of that RPS.
4
Prove the client isn't the ceiling
This is the mistake I made first. My single
httpx generator process topped out near 1,000 RPS against an idle server, and its measured throughput fell as I raised its own concurrency (1,045 at 8 concurrent, 502 at 50, 290 at 100). The client is a Python process too. Run several — my 3-client test reached 2,172 RPS where one client managed 1,045 — or use a native generator (wrk, hey, vegeta, k6).5
Attribute the limit
Sample server CPU during the run. A worker pinned near 100 percent means CPU-bound: fix serialization or add cores. Low CPU with flat RPS means you’re waiting on the database, the pool, or the threadpool’s 40 slots.
The code
Two files. The app has six routes that isolate the variables discussed above; the generator reports throughput and percentiles. I ran these on Python 3.14, FastAPI 0.142, and Uvicorn 0.54.--loop uvloop, one client process at 8 concurrent:
Three things fall out of it. Payload size, not request count, sets your ceiling: the same query at 500 rows instead of 20 costs 6–8x throughput, so a “what’s my RPS” question is unanswerable without a response shape.
JSONResponse beat the Pydantic path by 24 percent on the identical 500-row result, confirming that the win comes from skipping conversion, not from skipping one converter in particular. And /mixed is the loop-blocking case made visible: 23 RPS means the eight concurrent clients were served strictly one at a time, 43 ms of serialized CPU each, which is what happens when you put CPU work (or a sync driver call) inside async def.
Scale the per-request CPU numbers to your core, not the absolute RPS.
Turning RPS into users
Active users is two different questions. Concurrent users are the ones clicking at the same second; DAU/MAU is the pool that generates them. Little’s Law adapted for web traffic gives you the first:
Worked example with page load cost included: 2 API requests per page and a 6-second pause between pages gives (400 × 6) / 2 = 1,200 concurrent users at peak.
To reach DAU, spread those requests over a day. Traffic is never flat — a peak hour typically carries 10–15% of the day’s total — so the 400 RPS headroom is only available for a slice of the day:
The uniform case is the arithmetic to trust: 400 RPS × 86,400 s = 34.56 million requests/day, and at ~30 requests per user-day that is about 1.1 million DAU. The spiky rows apply a 10–15% peak-hour concentration factor on top, which is why they collapse to a tenth. At the top of the band, 400 RPS of measured capacity supports 4,000 to 40,000+ concurrent active users depending on how aggressively their clients poll — before you buy anything, check whether your mobile app’s heartbeat is the thing setting that rate.
Sizing the pod from these numbers
For the 1-core/4 GB envelope: setrequests.cpu and limits.cpu to 1 core and run one Uvicorn worker. Worker RSS in my runs was about 30 MB, but a real SQLAlchemy-plus-ORM process is more like 150–400 MB, and each worker is a separate copy of that, so 4 GB divided by worker RSS is your worker ceiling regardless of anything else in this article. Keep requests.memory equal to limits.memory so a serialization-heavy response can’t get the pod evicted, set --limit-concurrency just above your measured knee so excess load sheds as a 503 instead of becoming 3-second p99s, and give terminationGracePeriodSeconds enough room for the loop to drain in-flight requests on SIGTERM.
If you autoscale on CPU, remember a CPU-bound worker pins one core well before it reaches its RPS ceiling — a 60 percent CPU target on a 1-core pod is roughly 240 RPS of the 400 you measured. Autoscale on requests-per-second-per-pod from your own knee test instead.
Sources I read while building this
The pages scraped into this note kept their titles but lost their URLs, so these are citations rather than links. Search the titles:- Medium — Asyncio Mastery: Scaling FastAPI to 10k+ RPS in 2025: connection pooling as the first bottleneck at scale
- Stack Overflow — FastAPI Insert Performance: 1-2 ms Per Insert — Is This Normal?: the single-insert latency baseline
- Medium — FastAPI: The Surprising Performance Workhorse That Gives Rust a Run
- Medium (Kenan Can) — FastAPI Async vs Sync: Benchmark Results: an I/O-bound point lookup (
db_user), a pattern search (db_search), and a mixed database-plus-CPU workload (mixed_db_cpu) at 1 and 4 Uvicorn workers, concluding that I/O-bound routes should useasync defwith an async driver - Reddit r/FastAPI — Not so much performance improvement when using async: the database topping out at 80% CPU and 80% RAM with about 300 connections under connection pooling
- Reddit r/FastAPI — How does FastAPI utilize the CPU? (single-worker Uvicorn) and Medium — What Actually Happens When FastAPI Handles a Request? (one worker, one OS thread, one core)
- Medium (Afolabi Ifeoluwa James) — FastAPI vs Flask vs Django: Uvicorn-served FastAPI taking the highest throughput on the
/jsonendpoint - GitHub — Understanding benchmarking and setting up async Postgres:
create_async_engineandasync_sessionmakerfromsqlmodel.ext.asyncio.session - Reddit — Implementing Async REST APIs in FastAPI with PostgreSQL CRUD: 6 Uvicorn workers and an async client against 20k POSTs