Handling load spikes and resuming after idle: DSQL and existing services
How do DSQL and existing serverless and fixed-capacity services compare in latency, manual intervention, and cost?
Results summary
This experiment compared DSQL and Aurora Serverless v2 (A2) on whether they keep their latency targets without anyone adjusting capacity when request volume suddenly rises, or when requests resume after a long quiet period. As a minimum scope (MVP), we sent one spike load that changed the arrival rate from 20% (2 minutes) to 200% (1 minute) to 100% (2 minutes) to 5% (1 minute) of a 1,000 TPS baseline. After that, we closed all connections, waited 15 minutes, and sent a first request; this idle test was run twice.
Under the spike load, both services processed every stage with zero failures and no manual intervention. DSQL’s order-creation p95 was 29.2–36.0 ms and barely changed even when the arrival rate increased tenfold; p99 in the 2,000 TPS stage was 51.5 ms. A2, configured with a 4–32 ACU range, saw its writer grow from 10.5 ACU to a maximum of 21 ACU; its order-creation p95 was 8.2–12.1 ms, and p99 in the 2,000 TPS stage was 37.9 ms. This spike (up to 2,000 TPS) was smaller than the capacity of either service as measured in E002, so these results cannot tell us how the services handle spikes near their capacity limits.
The first request after idle showed a large difference. Even after 15 minutes with no connections, DSQL responded with a first connection of 112–287 ms and a first query of 115–123 ms, after which queries returned to about 2 ms. For this test, A2 was configured with a minimum capacity of 0 ACU and auto-pause after 300 seconds, and CloudWatch confirmed that it actually paused down to 0 ACU. In that state, the first connection failed both times because it did not complete within the driver’s 15-second connection timeout; that connection attempt triggered the resume, and capacity came back within about 1 minute.
Production readiness
For services whose requests arrive sparsely or stop overnight, DSQL had a clear advantage. DSQL handled the first request within 0.3 seconds even after 15 minutes idle, and no compute charges accrue while it is idle. Aurora Serverless v2 can also eliminate idle cost with 0 ACU auto-pause, but the first connection failed after exceeding 15 seconds during the resume, so connection timeouts and retries must be set long. At this spike size, both services handled the load without manual intervention, so spike handling near capacity limits must be judged together with the E002 capacity results. These results come from a single run on small data.
Question
How do DSQL and existing serverless and fixed-capacity services compare in latency, manual intervention, and cost?
Test conditions
- Region and time: Seoul (ap-northeast-2), 2026-09-29. Spike 08:43–08:51 UTC, idle 08:51–09:21 UTC.
- Configurations: D1 Aurora DSQL single-Region, A2 Aurora PostgreSQL 16.15 Serverless v2 (writer+reader, 4–32 ACU for the spike test). Data was 2% of the E002 scale. R1 and A1 (fixed capacity) were excluded from this MVP.
- Spike load: The same workload mix as E002 was sent at fixed arrival rates. Each stage ran as a new cell, with no warm-up. There was cell preparation time between stages (D1 about 20–45 seconds, A2 about 10 seconds), so the load was not continuous as planned.
- Idle test: We closed all of the runner’s connections, waited 900 seconds, and then timed a new connection, a single product lookup, and 20 further lookups. Repeated 2 times. During this test A2 was changed to
MinCapacity 0andSecondsUntilAutoPause 300, then set back to 4–32 ACU afterward. The connection timeout was the driver setting of 15 seconds. - Load generator: Spot runners, D1
c6g.4xlargeand A2m7g.4xlarge. - Deviations from the plan: Instead of ratios of Qref (20%→200%→100%→5%), we used a 1,000 TPS baseline; we shortened the stage durations (10/2/20/10 minutes → 2/1/2/1 minutes); and instead of 3 repetitions and 5 idle tests, we ran 1 and 2.
Performance results
Spike load
Order-creation p95 / p99 (ms). The failure rate was 0 in every stage (0.0008% in the D1 1,000 TPS stage).
| Stage | Arrival rate | D1 successful TPS | D1 order creation | A2 successful TPS | A2 order creation | A2 writer ACU (max) |
|---|---|---|---|---|---|---|
| 20% | 200 | 198 | 29.2 / 38.6 | 198 | 8.5 / 12.2 | 15 |
| 200% | 2,000 | 1,960 | 30.1 / 51.5 | 1,960 | 9.3 / 37.9 | 19–21 |
| 100% | 1,000 | 965 | 29.2 / 40.2 | 965 | 8.2 / 9.9 | 10.5–19 |
| 5% | 50 | 49 | 36.0 / 67.4 | 49 | 12.1 / 14.9 | 10.5 |
- We attribute the larger p99 in D1’s 5% stage to the small number of requests (about 3,000 in 1 minute), which let a few slow requests dominate p99.
First request after 15 minutes idle
| Configuration | Run | First connection | First query | p50 of next 20 | Result |
|---|---|---|---|---|---|
| D1 | 1 | 287 ms | 123 ms | 2.2 ms | Success |
| D1 | 2 | 112 ms | 115 ms | 1.9 ms | Success |
| A2 | 1 | Failed after 15.2 s | - | - | Connection timeout exceeded |
| A2 | 2 | Failed after 15.1 s | - | - | Connection timeout exceeded |
- The A2 writer’s ACU was 0 during 08:57–09:05 UTC and 09:12–09:20 UTC, and rose to 2–6 ACU within 1 minute after the first connection attempt. We could not measure the exact resume time with this tooling (15 seconds or more).
Development and operations
- During the spike test, neither service needed human intervention. A2 requires an ACU range (minimum and maximum) to be set in advance; DSQL has no capacity value to configure.
- To use A2’s auto-pause, the minimum capacity must be changed to 0, and the application’s connection timeout and retries must be set longer than the resume time. DSQL needed IAM token signing and a TLS connection for the first connection after idle, but this finished within 0.3 seconds.
Cost
E005, E006, E007, and E009 shared resources in the same run B (e002-20260929t082035z-afbe, 2026-09-29 08:20–10:20 UTC), so cost is reported per run. The estimates multiply CloudWatch usage by Seoul Region On-Demand prices; the actual charges were confirmed in Cost Explorer (2026-09-29 UTC, by usage type) on 2026-09-30. DSQL DPU and Spot runners were billed on the same lines as the second E002 run on the same day, so DPU was split by per-cluster CloudWatch values and Spot runners by running hours per instance type. Account charges that already existed before the experiments are excluded.
| Item | Usage | Estimated cost | Actual charge |
|---|---|---|---|
| A2 compute | estimated 19.05 ACU-hours (writer 8.53, reader 9.26, E007 restored copy 1.26), billed 18.77 ACU-hours × $0.20 | about $3.81 | $3.75 |
| A2 I/O | 1.198 million I/Os billed | about $0.29 | $0.29 |
| D1 DPU | 34,055 DPU + 33 DPU for the E007 restored copy | about $0.23 | $0.34 |
| E007 DSQL restore | 0.095 GB restored | - | $0.002 |
| Two Spot runners | c6g.4xlarge about 1.8 hours, m7g.4xlarge about 1.8 hours |
about $0.8 | $0.68 |
| Total | about $5.1 | $5.07 |
- About $0.65 for cross-AZ transfer, EBS, and public IPv4 shared by the two runs was not assigned to either run.
-
The D1 DPU estimate was lower than the charge because usage after the harness’s last measurement was missing from it.
- The harness’s cost guard counted unmeasured A2 intervals at the maximum of 32 ACU and estimated about $9.5; because of this, the guard limit was raised from $14 to $16 before the E006 re-measurement.
- Most of A2’s cost came from the writer and reader, each kept running at a minimum of 4 ACU even while there were almost no requests. Over the same period, DSQL was billed only for the DPUs used to process requests ($0.34).
Conclusions and limitations
- At this spike size, both services had zero failures without intervention, and DSQL’s latency stayed constant regardless of load.
- For the first request after idle, DSQL responded immediately, whereas A2, paused at 0 ACU, took longer than the 15-second connection timeout to resume.
- Limitations: a single run, small data, and a spike load with gaps between stages. The spike was smaller than the capacity limits, and we did not compare against the fixed-capacity controls (R1, A1) or compare idle-period costs.
Cleanup record
- Resources created in run B: BATCH (VPC, 2 subnets, IGW, 2 security groups, DB subnet group, runner IAM role and instance profile), the D1 DSQL cluster, the A2 cluster with writer and reader, an RDS-managed secret, and 2 Spot runners (D1
c6g.4xlarge, A2m7g.4xlarge). E007 additionally created an AWS Backup vault, a backup service role, 2 recovery points, 1 restored DSQL cluster, and a point-in-time-restored A2 cluster and instance. - Deletion: The backup vault and 2 recovery points were deleted at 09:35 UTC, and the restored DSQL cluster and A2 restored copy were deleted from 09:48 UTC to about 10:05 UTC after checking their ownership tags. Everything else was deleted with
e002.py batch-downat 10:04–10:20 UTC. No resource failed to delete. - Verification (10:20:27 UTC):
e002.py verifyreportedremaining_count=0. A manual cross-check confirmed 0 AWS Backup vaults, 0 RDS clusters and IAM roles with this run’s prefix, and 0 open Spot requests. The 1 DSQL cluster that still existed at that time belonged to run A (E002, E003, E008, E012), which was running concurrently.