Back to the experiment list
E002CompletedP0

Limits of OLTP throughput and latency

How many times the SLO capacity and latency of each control does DSQL have?

Results summary

This experiment checked how many more requests DSQL handles than the controls, and how its latency differs, when the same order workload runs under the same latency objective (SLO). The workload mixed 40% product lookups, 30% order history lookups, 20% order creation, and 10% order cancellation, and the SLO was read p95 50 ms and p99 100 ms, write p95 100 ms and p99 200 ms, and a technical failure rate of 0.1% or less. The experiment ran in two parts. On 2026-09-28 we brought up all four services together for a pilot and a concurrency search, but the DSQL search stopped because of slow resets and the budget; on 2026-09-29 we brought up DSQL alone and measured again without deleting data between cells.

With a fixed number of connections sending requests without pause (closed-loop), DSQL handled 6,680 TPS with 64 connections and 24,541 TPS with 256 connections, meeting the SLO in both. Even at 256 connections, order creation p99 was 35.7 ms, almost the same as with 64 connections (30.4 ms). In the control measurements with the same method (2026-09-28), write latency exceeded the SLO for all three controls at 256 connections, and the highest throughput that met the SLO was 5,839 TPS for RDS Multi-AZ (R1), 3,375 TPS for Aurora Provisioned (A1), and 11,559 TPS for Aurora Serverless v2 (A2). At 256 connections, DSQL handled about 2.1 times A2, about 4.2 times R1, and about 7.3 times A1. At that point the load generator CPU was at 80.8%, close to the saturation threshold (85%), so the DSQL value is a lower bound for this measurement, not a service limit.

In absolute terms, DSQL latency was longer. At the same 1,600 TPS, order creation p95 was 41.0 ms on DSQL (2026-09-28) and 6.8 ms on A2, and in the 2026-09-29 measurement DSQL write p95 was about 28–30 ms. When requests were sent at a fixed arrival rate (open-loop), DSQL met the SLO up to 6,750 TPS, and p99 exceeded 200 ms at 8,100 TPS. However, this failure came not from DSQL processing latency (p95 unchanged at about 28 ms) but from time the load generator spent waiting for the connections that each process shares (wait p99 about 190 ms), so we interpret the open-loop limit as a connection pool limit of the measurement tool. There were zero business invariant violations in every cell. The planned main measurement at common arrival rates and the repetitions were not done, so the differences between services are single-measurement values.

Production readiness

In terms of throughput, DSQL scaled the most in this experiment. At 256 connections it handled more than 24,000 mixed order operations per second without a rise in latency, while all the existing services exceeded the latency objective under the same conditions. Holding up as load grows without capacity planning is an advantage for services with large traffic swings. On the other hand, the latency of a single request was 3–6 times longer than Aurora (write p95 about 30–40 ms), so screens that run several queries in sequence per request will see longer response times. Reaching the same throughput requires more connections and more concurrency, and sustaining high throughput increases DPU charges (E010).

Question

How many times the SLO capacity and latency of each control does DSQL have?

We can answer on the basis of a fixed concurrency of 256 connections: DSQL’s throughput while passing the SLO is at least about 2.1 times A2, at least about 4.2 times R1, and at least about 7.3 times A1. Latency is the reverse, with DSQL 3–6 times longer. The main measurement at common arrival rates (0.5/1.0/1.2 × Qref) was not done.

Test conditions

  • Region and time: Seoul (ap-northeast-2). First run (four configurations) 2026-09-28 03:34–12:07 UTC, second run (DSQL only) 2026-09-29 08:11–11:10 UTC.
  • Configurations:
    • D1: Aurora DSQL single-Region, IAM token authentication, public endpoint.
    • R1: RDS for PostgreSQL 16, db.r6g.xlarge Multi-AZ, gp3 400 GiB.
    • A1: Aurora PostgreSQL 16 Provisioned, one db.r6g.xlarge writer + one reader in another AZ, Aurora Standard.
    • A2: Aurora PostgreSQL 16 Serverless v2, one writer + one reader, 4–32 ACU each, Aurora Standard.
  • Data: 1 million customers, 200,000 products, 11 million orders, 27.5 million order items, and 11 million ledger rows (about 5 GiB). For D1, we counted the loaded rows directly in the database and confirmed that they matched the design values.
  • Workload and verdict: the mixed workload above, explicit REPEATABLE READ, and a common retry policy. A cell passes only if it meets both the technical failure rate of 0.1% or less and the SLO. Inventory, order, and ledger invariants were checked in each cell.
  • Load generator: one Spot runner per configuration. The pilot used c7g.4xlarge; for the search, all four configurations used c6g.4xlarge because of Spot capacity shortages and interruptions. Generator CPU peaked at 69.7% (A2, 256 connections), below the saturation threshold (85%).
  • Cell duration: pilot warm-up 60 s + measurement 120 s; search and second-run warm-up 120 s + measurement 300 s; once each.
  • Second run (DSQL only): we loaded the same data into a new DSQL cluster and let the rows created by each cell accumulate instead of deleting them. Consistency was checked in each cell using only the rows that cell created (a separately assigned ID range and a business ID prefix per cell). Orders grow by about 1% per cell. The runner started as c6g.4xlarge and was replaced by m7g.4xlarge after a Spot interruption; the 64- and 256-connection cells and the cells at 5,400 TPS and above ran on m7g.4xlarge. For the controls, the 2026-09-28 results were used as is for comparison.
  • Deviations from the plan:
    • The main measurement (0.5/1.0/1.2 × Qref) and the boundary search could not be run.
    • The cost cap was raised from USD 50 to 60 (user decision on 2026-09-28).
    • The searches for the controls and for DSQL ran on different days, with different runner specifications (c6g.4xlarge vs m7g.4xlarge) and different runner locations. DSQL RTT was 2.7–5.1 ms in the first run and 1.6–1.8 ms in the second, and write latency was also about 30% shorter in the second run.
    • Cost caps: USD 60 for the first run (user decision on 2026-09-28) and USD 20 for the second run (user decision on 2026-09-29).

Performance results

Fixed arrival rate pilot (open-loop)

p95 / p99, in ms. In every cell the failure rate was 0, there were 0 invariant violations, and the SLO was met.

Configuration Arrival rate Product lookup Order history Order creation Order cancellation RTT
D1 100 10.5 / 12.3 14.3 / 16.1 47.6 / 55.2 50.0 / 59.2 2.8
D1 400 10.3 / 11.5 13.2 / 14.3 47.6 / 53.6 48.1 / 54.2 3.0
D1 1600 9.1 / 10.6 11.7 / 13.3 41.0 / 46.2 41.4 / 47.6 5.1
A2 100 3.6 / 4.5 9.9 / 11.8 13.1 / 16.2 7.6 / 9.4 0.3
A2 400 2.4 / 2.8 3.9 / 5.2 9.3 / 11.4 9.0 / 10.9 0.1
A2 1600 2.2 / 2.7 2.6 / 3.2 6.8 / 7.9 6.9 / 8.1 0.1

The A2 writer used 12 ACU at 1600 TPS. D1 consumed about 0.0375 DPU per request attempt.

Fixed concurrency search (closed-loop)

Successful throughput (TPS) and SLO verdict. In every cell the failure rate was 0 and there were 0 invariant violations. The controls were measured on 2026-09-28, and D1 at 64 and 256 connections on 2026-09-29.

Configuration 16 connections 64 connections 256 connections Highest throughput meeting the SLO
D1 1,006 pass (first run) 6,680 pass 24,541 pass 24,541 or more (lower bound)
R1 3,065 pass 5,839 pass 5,347 exceeded 5,839
A1 3,375 pass 3,322 exceeded 4,254 exceeded 3,375
A2 5,657 pass 11,559 pass 15,115 exceeded 11,559
  • D1 at 256 connections, p95 / p99 (ms): product lookup 5.6 / 6.4, order history 8.3 / 9.5, order creation 30.4 / 35.7, order cancellation 28.9 / 32.9. These are almost the same as at 64 connections (4.9 / 5.5, 7.6 / 8.4, 27.8 / 30.4, 27.0 / 29.8), so quadrupling the connections barely increased latency. Runner CPU was 80.8% at that point.
  • D1’s attempted throughput at 256 connections was 26,285 TPS, about 7% more than its successful throughput. The difference consists of attempts retried after commit-time conflicts (40001).
  • A1 at 64 connections had an order history lookup p99 of 178.7 ms, exceeding the read SLO (100 ms).
  • At 256 connections, write p95 exceeded the write SLO (100 ms) for all three controls (order creation p95: R1 163.4 ms, A1 250.7 ms, A2 122.5 ms).
  • D1 at 16 connections (first run) was low at 1,006 TPS, because D1 latency in the first run was longer than in the second (RTT 2.7–5.1 ms) and, with few connections, per-request latency limited throughput.

Stepping up a fixed arrival rate (open-loop, DSQL only)

Arrival rate Successful TPS Verdict Order creation p95 / p99 Connection wait p99
2,400 2,329 Pass - -
3,600 3,493 Pass - -
5,400 5,202 Pass - -
6,750 6,416 Pass 27.8 / 101.4 26.5 ms
8,100 7,758 Exceeded 28.1 / 255.7 189.7 ms
  • At 8,100 TPS, p99 exceeded 200 ms for every operation, but processing latency excluding wait time (p95 about 28 ms) did not change. The runner splits its 256 connections into 16 per process across 16 processes, so in processes where arrivals bunched up, requests waited for a connection. We therefore interpret open-loop 6,750 TPS not as a DSQL limit but as a connection pool limit of this measurement tool.
  • DPU per request attempt was 0.029–0.032 (0.032 at 2,400 TPS, 0.029 at 26,285 TPS), slightly lower than the pilot estimate of 0.0375.

Development and operations

  • Bulk deletes are slow: in the reset that reverts the rows created by a cell, D1 took 6–7 minutes to process about 130,000 receipts and related rows (about 600,000 row changes in total across orders, items, ledger, and inventory restoration). The first implementation re-ran the target query for every batch and exceeded the 27-minute limit, and even after we changed it to query the targets only once, the reset for the 64-connection cell did not finish within the time limit. For the controls, the same reset did not cause problems for cell timing.
  • Reverse index for a foreign key: when deleting orders, there was no index on the ledger’s order_id, so every delete scanned the entire ledger, and D1 hit the 300-second transaction limit (ProgramLimitExceeded). The same problem appeared on A2 as a delete that took 24 minutes. We fixed it by adding an index.
  • Asynchronous indexes: indexes created with DSQL’s CREATE INDEX ASYNC appeared in pg_indexes before the build finished. We changed the tool to wait until pg_index.indisvalid became true.
  • Transient errors: XX000 server unavailable occurred during loading, so we added it to the retry targets. We split loading and deletion to fit the 3,000-row limit per transaction.
  • Connection management: DSQL connections have a one-hour limit, so we reconnected every 50 minutes and cached the IAM token for 10 minutes.
  • Concurrent token signing: in the second run, 256 connections tried to sign IAM tokens at the same time and the instance credential lookup failed (NoCredentialsError). We added a lock so that signing happens only once per process, and made transient credential errors retry.
  • Accumulating instead of resetting: to avoid slow DSQL bulk deletes, the second run did not delete rows created by cells. At first, retries counted rows left by attempts interrupted by Spot reclamation as their own rows, so we changed each attempt to use a unique ID range.

Cost

First run (four configurations, 2026-09-28): the estimated cost is what the harness calculated from usage and unit prices during the run, and the actual charge was checked on 2026-09-29 in Cost Explorer (Seoul Region, 2026-09-28 UTC, by usage type). All of the experiment ran within 2026-09-28 UTC. Standing account charges that also occurred the day before the experiment (load balancers, S3, existing public IPv4 and so on, about USD 2.9 per day) were excluded. The Cost Explorer values at the time of the query are estimates before month-end close.

Item Estimated cost Actual charge Usage and difference
D1 DPU About $10.7 $11.49 1,148,488 DPU. Usage after the harness’s last measurement (the aborted 64-connection cell, before and after deletion) was missing from the estimate.
A2 compute About $31.3 $23.99 119.9 ACU-hours. The estimate was high because unmeasured periods were calculated at the maximum ACU.
Aurora I/O (A1+A2) About $4.7 $6.53 27.2 million I/Os (Aurora Standard). I/O after the last measurement was missing from the estimate.
A1 instances About $2.9 $2.43 3.90 instance-hours of db.r6g.xlarge (writer+reader)
R1 About $2.7 $2.12 1.72 hours of Multi-AZ + gp3 storage
Runners and other About $0.9 $2.03 13.95 hours of Spot runners ($1.68), inter-AZ transfer, public IPv4, EBS
Total About $53 $48.57 About 92% of the estimate. Within the USD 60 cap and the USD 54 guard

Second run (DSQL only, 2026-09-29): the CloudWatch TotalDPU total of 1,784,023 DPU (about $17.84) plus about $0.73 for the Spot runner came to an estimated $18.6; the actual charge confirmed in Cost Explorer on 2026-09-30 was $18.54 (DSQL DPU $17.84, Spot runners $0.70). DSQL DPU for this run and the E005, E006, E007, and E009 run on the same day was billed on one line, so it was split by per-cluster CloudWatch DPU; the CloudWatch total for both runs (1,818,111 DPU) matched the billed DPU (1,818,110). Of this, data loading was 517,021 DPU (about $5.17), the step-up and closed-loop measurements were 1,093,524 DPU (about $10.94), and the rest was E003, E008, and E012, which ran on the same cluster. About $0.65 for cross-AZ transfer, EBS, and public IPv4 shared by the two runs was not assigned to either run.

  • This month’s DSQL free usage (100,000 DPU) was used up in E004, so D1 DPU is billed with no free usage deduction.
  • A2 scaled up its ACU significantly while taking 15,000 TPS of load in the search, and the idle time while waiting for a decision after the pilot is also included, making it the largest cost (about 49% of the total).
  • The Aurora Standard I/O charge was $6.53, larger than the A1 instance charge. For heavy-load experiments, I/O-Optimized may be cheaper (not compared this time).

Conclusions and limitations

  • Throughput: at a fixed concurrency of 256 connections, DSQL handled 24,541 TPS or more within the SLO, 2.1–7.3 times the controls’ highest throughput passing the SLO (R1 5,839, A1 3,375, A2 11,559 TPS). The DSQL value is a lower bound close to the load generator’s limit.
  • Latency: DSQL write p95 was about 28–41 ms, 3–6 times longer than Aurora Serverless v2 (about 7–9 ms), but it barely rose as load increased.
  • Not measured: the main measurement at common arrival rates (0.5/1.0/1.2 × Qref), the boundary search, variation across repetitions, and DSQL’s actual upper limit using multiple runners.
  • Limitations: every cell was run once. The searches for DSQL and the controls ran on different days with different runner specifications. DSQL used a public endpoint, while the controls used a path inside the VPC. In the second run, data grew by about 1% per cell.

Cleanup record

  • Resources created:
    • BATCH: VPC, two subnets, IGW, two security groups, DB subnet group, IAM role and instance profile.
    • Per configuration: the D1 cluster, the R1 instance, the A1 and A2 clusters with writers and readers, and three RDS-managed secrets.
    • Spot runners: several, including replacements. They were replaced because of two Spot interruptions (two runners at 06:45 UTC, and three started at 08:54–08:59 UTC) and capacity shortages.
  • A2 recreation: to reduce cost after the pilot, we deleted A2, then recreated it and loaded data for the main phase.
  • Deletion completion times (UTC): R1, A1, and A2 10:57:45, D1 12:06:31, BATCH 12:06:45. No resource failed to delete.
  • Verification (12:07:44 UTC): remaining_count=0 from e002.py verify. The first verification (12:06:55) returned 1 because one Spot request for a terminated runner remained in the active state; we cancelled the request and verified again. The fact that the cleanup code does not cancel Spot requests directly is an item to fix.
  • Second run (2026-09-29, e002-20260929t081139z-4dce): we created the BATCH network and IAM, one DSQL cluster, and two Spot runners (c6g.4xlarge interrupted by Spot at 08:33 UTC, replaced by m7g.4xlarge). After running E003, E008, and E012 on the same cluster, we deleted everything at 11:08–11:10 UTC and confirmed remaining_count=0 from e002.py verify (11:10:30 UTC). This was after we changed the tool to cancel Spot requests together with runner termination, and no Spot requests remained. A manual cross-check confirmed 0 DSQL clusters, 0 VPCs with experiment tags, 0 experiment IAM roles, and 0 open Spot requests.