Atomicity without transactions
The alternatives all fall short in a specific way I have hit repeatedly:MULTI/EXECbatches commands but cannot branch on a value it reads. A read-modify-write like “increment, and only set the TTL if this is the first hit in the window” needs the result of the read to choose the next command, which is impossible inside a queue of pre-declared commands.WATCH/MULTI/EXECgives optimistic concurrency, so a contended key means the client re-runs the whole transaction. Under load the retry rate is the throughput limit, and you pay a round trip per attempt.- A client-side read-then-write has a race window measured in round trips. Two clients can both read
current = 4, both decide they are under the limit of 5, and both increment.
redis.call errors halfway through, the writes already applied stay applied. I verified this by running a script that wrote a key and then spun in a long loop; the key was there afterwards (GET during -> written), and no rollback occurred. Design your scripts so the last thing they do is the thing you care about.
The script I actually ship
The canonical example is a fixed-window rate limiter. It increments a counter for a given key, sets an expiration time on the first access, and returns whether the request is allowed or blocked:This reads then writes, so it costs three command executions inside the script on the first hit. The tightened version drops the
GET entirely: INCR first, set EXPIRE when the result is 1, then compare against the limit. It answers the same question with two server commands instead of three, and it is still correct because nothing else can observe the intermediate state.KEYSandARGV: Redis passes parameters into two separate arrays. By convention, all database keys must go inKEYS, and plain values (like timeouts, counts, or IDs) go inARGV. Lua arrays are 1-indexed (Redis EVAL intro, CodeSignal).redis.call(): this function executes any native Redis command within the script environment (freeCodeCamp guide).- Atomicity: the script blocks everything else while it executes. Keep your scripts short and fast to avoid triggering a
BUSYerror (Redis programmability, ScaleGrid on long-running scripts, Redis EVAL intro). - Local variables: always declare variables with the
localkeyword to ensure they remain safely sandboxed within your script context (Redis Lua API). A stray global is not confined to your script — Redis refuses scripts that create globals for exactly this reason.
redis-cli:
Why EVAL reduces round trips, with the arithmetic
Count the trips a client must make for one rate-limit decision:
The round-trip reduction is the whole performance story, because round trips do not overlap when the decision depends on the previous answer. At an intra-AZ Redis RTT of 0.25 ms, a three-trip client-side decision cannot exceed
1 / 0.00075 ≈ 1,333 decisions per second per connection, no matter how fast Redis itself is. One trip gives you 4,000/s per connection. Three times the headroom for the same server.
I measured it on a loopback Redis 8.8.0 with 20,000 rate-limit decisions per pass, counting each client-side request as one round trip (one fresh key per decision, so the EXPIRE always fires — the worst case for the naive path):
CONFIG RESETSTAT, 10,000 calls of the original script executed {'incr': 10000, 'evalsha': 10000, 'expire': 10000, 'get': 10000} inside itself, while 10,000 calls of the INCR-first version executed {'incr': 10000, 'evalsha': 10000, 'expire': 10000} — dropping GET removes a third of the per-decision command executions on the single thread that matters. The loopback RTT here is tens of microseconds, and I still got 2.88x — on a real network where RTT is the dominant term, the ratio approaches the trip ratio itself.
Script caching: EVALSHA instead of shipping source
Sending large scripts over the network repeatedly wastes bandwidth. Instead, production applications useSCRIPT LOAD to store the script on the server once, then execute it via its unique SHA1 hash using EVALSHA (DevGenius on Lua performance, Bullet-proofing Lua scripts in redis-py, Redis EVAL intro).
The hash is literally sha1(script_source), which I verified: my 576-byte script’s local SHA-1 and the digest Redis held for those exact bytes were the same 40-character value, f51d4c585273ad9be8595b474f9c31224b3c86f7 (SCRIPT EXISTS on it returned 1 right after redis-cli --eval loaded the file). Beware the off-by-one-byte trap here: passing the source through $(cat rate_limit.lua) strips the trailing newline, so the server hashes 575 bytes and reports a different digest. The byte math on the wire:
EVAL: 576 B of source + key/arg frames, every call.EVALSHA: 40 B of digest + the same frames, every call.- Saved per call: 536 B. At 1,000 calls/s that is 0.5 MB/s; at 10,000 calls/s it is 5.36 MB/s, about 463 GB/day of script source you stopped sending.
redis-cli --eval form: separate your KEYS and ARGV with a comma surrounded by spaces, so the CLI knows where keys end and arguments begin. And treat the script cache as volatile: SCRIPT FLUSH, a restart, or a failover clears it, so any client that uses EVALSHA must handle -NOSCRIPT and fall back to EVAL to re-register. I verified the reply text: NOSCRIPT No matching script. Please use EVAL. That retry loop is why you should use register_script / the Script object in redis-py rather than hand-rolling EVALSHA calls — the library manages the cache miss for you.
For Redis 7 and later I prefer functions (FUNCTION LOAD, FCALL) over ad-hoc scripts where the logic is part of the application: libraries are named, persisted, and replicated with the dataset, so a replica promotion does not leave you with a cold script cache and a fleet of clients hitting NOSCRIPT.
The single-threaded consequence of a slow script
Redis executes commands on one thread. A script is one command, so a slow script is not “one slow request” — it is a stop-the-world event for every other client on that node. I measured the effect directly against a loopback Redis 8.8.0:lua-time-limit (5000 ms by default), Redis stops queueing other clients and starts refusing them. With the limit lowered to 200 ms and a runaway read-only script, I watched 35,091 of 35,205 probe PINGs get rejected:
SCRIPT KILL answered -NOTBUSY No scripts in execution right now. The escape hatch depends on whether the script has written anything yet, which I caught correctly in a second run:
ARGV, never over the keyspace — KEYS * inside a script is how you take down a node), no unbatched fan-out, work proportional to a client-supplied argument count so a hostile caller cannot choose your runtime, and a slowlog/latency monitor alert on any script exceeding a millisecond or two.
Cluster and key-hygiene constraints
The corollaries I enforce in code review:- Never derive a key name from
ARGVinside the script (redis.call('GET', 'prefix:' .. ARGV[1])). Cluster routing is decided fromKEYSbefore the script runs, so an undeclared key is at best unroutable and at worst silently reads another node’s data. - Pass
numkeyscorrectly: the argument between the script and the values is the count of keys, and getting it wrong shifts every key intoARGV. - Return values cross the RESP boundary: return numbers, strings, and flat tables; a table with holes or a nested mixed table comes back truncated or as an error.
- Do not use
redis.call('TIME')for window boundaries when you want the logic reproducible — pass the timestamp inARGV, so the same inputs produce the same outputs and replicas are not asked to re-derive time. - Version the script by digest, not by filename.
SCRIPT LOADis idempotent, so deploys that re-load are free; but rolling restarts of clients and servers can briefly disagree on which body a SHA stands for, so deploy the script first, then the callers.
Which problem shape gets which tool
If you are picking this up cold, the two questions that determine the design are the same two I ask myself: what is the atomic decision (conditional update, atomic counter, multi-step workflow), and which client library and language are you in — because that decides whether you get automatic
NOSCRIPT retry, and whether your team can even read the script. Related: Sequential database scan for the same round-trip-versus-work tradeoff one layer down, and Hybrid logical clocks when the timestamp inside your script has to mean something across nodes.