Back to the experiment list
E004CompletedP0

Transaction contention and data consistency

How do contention and retries on DSQL affect consistency, latency, and implementation effort?

Results summary

This experiment checked whether DSQL and the controls preserve consistency when many requests update the same rows at the same time, and how conflicts and retries affect latency and failure rates. We ran order creation and account transfer workloads at concurrency 16/64/256, with uniform access and access concentrated on popular keys, and checked inventory, balance, and duplicate-processing invariants in every load cell. There were zero invariant violations across all 66 load cells. All four services worked without double processing or lost commits once conditional updates, business ID receipts, retries, and commit-status checks were implemented.

The way contention was handled differed greatly. DSQL reports conflicts (40001) at commit time without lock waits. As a result, under concentrated contention where 80% of requests go to the top 1% of keys, the p99 of requests that committed successfully was short, at 110 ms or less, but the conflict rate at concurrency 256 was 67–71%, and 39% still failed in the end even after up to 3 retries. The controls queued requests with row locks, and under concentrated contention at concurrency 64 and 256, waits exceeded the 2-second deadline, successful throughput dropped to 24–512 TPS, and final failure rates reached 18–63%. Because DSQL does not support READ COMMITTED, implementing retries is mandatory, and for workloads where contention concentrates, such as popular products, the retry budget and a schema design that spreads contention should be considered together.

Because of the cost cap (USD 5 at the time, later raised to 50), we scaled the controls down to small burstable instances and ran each cell only once for 25 seconds. The throughput differences are therefore not used as a basis for comparing services, and because the method uses fixed concurrency, latency while requests back up may be missed. A capacity and latency comparison at the planned scale will be measured again in E002 with a fixed request rate. The estimated cost calculated from usage was about USD 3.4, and the actual charge confirmed in Cost Explorer was about USD 2.07 (the main difference being the deduction of DSQL’s monthly free usage).

Production readiness

With retries and business ID receipts (idempotent processing) implemented, DSQL preserved consistency under contention without double processing or lost commits, so consistency itself was not an obstacle to production adoption. For workloads where contention is spread out, the conflict rate was similar to the existing services and retries resolved most conflicts. However, when writes concentrate on a few rows, as with inventory for popular products, a large share of failures remains even after retries (39% at concurrency 256), so such workloads are hard to move as they are without a schema redesign such as splitting inventory rows. The 3,000-row limit per transaction meant that bulk deletes and cleanup jobs had to be split, and a server error was observed once, so retries are mandatory. These results come from small controls and a single repetition, so capacity is judged in E002.

Question

How do contention and retries on DSQL affect consistency, latency, and implementation effort? We ran the same order and transfer workloads on D1 (Aurora DSQL) and three controls (R1 RDS PostgreSQL Multi-AZ, A1 Aurora Provisioned, A2 Aurora Serverless v2). We checked two things.

  • Isolation level scenarios: we fixed the execution order of two connections with a barrier to reproduce five kinds of anomalies, and looked at whether each service prevented them.
  • Contention load: we applied load while varying concurrency and access distribution, and checked business invariants in each cell.

Test conditions

  • Measurement time: 2026-09-25 22:50 UTC – 2026-09-26 01:57 UTC, region ap-northeast-2, run prefix e004-20260925t221414z-e5fd.
  • Code commit: the load cells ran on 9474ed6. Five D1 cells run earlier on e3b71f9 are also included in the results. The runner code did not change between the two commits; only the way the operations tool records results changed. The reproduction procedure is in the experiment README.
  • Compared configurations: scaled down from the plan because of the cost cap.
ID Actual configuration Difference from the planned configuration
D1 DSQL single-Region cluster, IAM token, verify-full TLS, public endpoint None
R1 RDS PostgreSQL 16.15, db.t4g.medium Multi-AZ, gp3 20 GiB, private endpoint Smaller class (burstable 2 vCPU) and storage
A1 Aurora PostgreSQL 16.15, one db.t4g.medium writer, I/O-Optimized No reader, small class, I/O-Optimized instead of Standard
A2 Aurora Serverless v2 16.15, one writer, 0.5–4 ACU, I/O-Optimized No reader, narrower ACU range
  • Load generator: an EC2 c7g.4xlarge Spot instance (16 vCPU) in the same VPC, accessed only through SSM. Tool versions were Python 3.11.16, psycopg 3.3.6 (libpq 18), and boto3 1.43. Generator CPU utilization during measurement peaked at 75%, and no cell was invalidated for saturation.
  • Median RTT: D1 2.6 ms, R1 0.7 ms, A1 0.2 ms, A2 0.16 ms. Only D1 goes through a public endpoint, so its RTT is the longest.
  • Workload: 70% order creation (conditional inventory decrement, storing the order, items, and receipt) and 30% account transfers (update two accounts after checking the balance). The SQL is identical on all four services. Business ID receipts prevent duplicate processing, and when it is uncertain whether a commit happened, the receipt is looked up to decide.
  • Condition matrix:
    • REPEATABLE READ (common to all four services): uniform and concentrated distributions × concurrency 16/64/256 × no retries / up to 3 retries.
    • READ COMMITTED: run on the controls only (the controls’ default isolation level, up to 3 retries). D1’s default isolation level is REPEATABLE READ, so the results from the common matrix are used as is.
    • The concentrated distribution sends 80% of requests to the top 1% of keys.
    • Retries apply only to 40001/40P01, with a total deadline of 2 seconds, exponential backoff (10 ms base), and full jitter.
  • Cells and repetitions: each cell is 5 seconds of warm-up + 20 seconds of measurement, with one repetition. Concurrency is fixed and the next request is sent as soon as a response arrives (closed-loop).
  • Deviations from the plan:
    • The plan was 10 minutes of warm-up + 20 minutes of collection, repeated 3 times.
    • In the D1 pilot, DSQL consumed about 0.095 DPU per attempt. Running under the planned conditions was estimated to cost about USD 16.5 in D1 DPU alone, so by user decision the number of repetitions and the cell duration were reduced.
    • In the pilot, c7g.2xlarge saturated at concurrency 256, so the runner was changed to c7g.4xlarge.
    • During the run, the Spot instance was reclaimed (23:19 UTC, insufficient capacity), so we replaced the runner and resumed from D1’s remaining cells.
    • One D1 cell (D1-RR-hot-c64-retry3) failed because of a DSQL server error (InternalError_: server unavailable, 22:57 UTC) and was run once more. The record of the first attempt was also kept.

Performance results

Isolation level scenarios

We fixed the execution order of two transactions and recorded the result. Each result is one of “anomaly occurred”, “blocked by error (SQLSTATE)”, or “blocked by waiting”. Hermitage-style notation is in parentheses. D1 rejected BEGIN ISOLATION LEVEL READ COMMITTED and SERIALIZABLE (0A000), so those rows are not applicable.

Scenario Isolation level D1 R1 · A1 · A2 (same for all three)
lost update (P4) READ COMMITTED Not applicable Anomaly occurred
lost update (P4) REPEATABLE READ Blocked by error (40001) Blocked by error (40001)
lost update (P4) SERIALIZABLE Not applicable Blocked by error (40001)
write skew (G2-item) READ COMMITTED Not applicable Anomaly occurred
write skew (G2-item) REPEATABLE READ Blocked by error (40001) Anomaly occurred
write skew (G2-item) SERIALIZABLE Not applicable Blocked by error (40001)
Duplicate order with the same business ID READ COMMITTED Not applicable Blocked after waiting (23505)
Duplicate order with the same business ID REPEATABLE READ Blocked by error (40001) Blocked after waiting (23505)
Duplicate order with the same business ID SERIALIZABLE Not applicable Blocked after waiting (40001)
Crossed updates (deadlock) All three levels Blocked at commit (40001, RR) Deadlock detected (40P01)
SELECT FOR UPDATE conditional decrement READ COMMITTED Not applicable Processed sequentially after waiting
SELECT FOR UPDATE conditional decrement REPEATABLE READ / SERIALIZABLE Blocked at commit (40001, RR) Blocked after waiting (40001)
  • The three control configurations: all matched the per-isolation-level behavior described in the PostgreSQL documentation. For example, REPEATABLE READ allowed write skew, and deadlocks ended with 40P01.
  • Two differences in D1:
    1. There is no lock waiting. When two transactions updated the same row, the one that committed first succeeded and the one that committed later failed with 40001. Even with FOR UPDATE, the other transaction did not wait.
    2. It blocked this write skew scenario with 40001 under REPEATABLE READ. PostgreSQL’s REPEATABLE READ allowed the same scenario.
  • Limits of interpretation: item 2 is a result observed once. We did not confirm in the official documentation whether this is behavior that DSQL’s concurrency control always guarantees. We therefore do not generalize it as “DSQL prevents write skew”.

Contention load

The metrics used in the tables below are defined as follows.

  • Successful TPS: the number of operations that committed successfully during the 20 seconds of measurement ÷ 20.
  • Raw conflict rate: the share of all attempts that failed with 40001/40P01.
  • Final failure rate: the share of operations that ultimately failed after retries and the deadline are taken into account. There were zero business rejections such as out of stock or insufficient balance.
  • p99 overall: the time until one operation finished, including retries and waiting. Failed operations are included. For operations that hit the 2-second deadline, time for replacing the connection and looking up the receipt is added until they finish, so values above 2 seconds can appear.
  • p99 committed: the p99 over operations that committed successfully only.

Every cell has one repetition, so variation is unknown. Each value is the result of a single measurement.

REPEATABLE READ · no retries

Distribution Concurrency Configuration Successful TPS Raw conflict rate Final failure rate p50 (ms) p99 overall (ms) p99 committed (ms)
Uniform 16 D1 841 0.6% 0.6% 19 28 28
Uniform 16 R1 1,500 0.5% 0.5% 10 23 23
Uniform 16 A1 855 0.6% 0.6% 18 44 44
Uniform 16 A2 990 0.5% 0.5% 11 59 59
Uniform 64 D1 2,363 2.5% 2.5% 26 39 39
Uniform 64 R1 1,745 2.6% 2.6% 34 80 81
Uniform 64 A1 1,050 2.1% 2.1% 56 152 152
Uniform 64 A2 922 2.4% 2.4% 69 186 186
Uniform 256 D1 8,494 8.6% 8.6% 27 41 41
Uniform 256 R1 1,412 9.1% 9.1% 159 341 345
Uniform 256 A1 1,165 8.4% 8.4% 192 400 400
Uniform 256 A2 0 0.0% 100.0% 7,098 8,449 -
Concentrated 16 D1 711 20.8% 20.8% 18 25 25
Concentrated 16 R1 890 21.6% 21.6% 10 25 22
Concentrated 16 A1 518 19.8% 19.8% 15 81 44
Concentrated 16 A2 608 20.1% 20.1% 9 73 65
Concentrated 64 D1 1,402 43.7% 43.7% 25 36 35
Concentrated 64 R1 423 43.5% 44.0% 27 1,020 208
Concentrated 64 A1 274 39.6% 41.0% 43 4,536 335
Concentrated 64 A2 302 42.1% 42.2% 43 1,020 678
Concentrated 256 D1 4,479 66.5% 66.5% 19 31 28
Concentrated 256 R1 512 62.1% 63.1% 81 1,852 639
Concentrated 256 A1 56 17.3% 42.5% 109 16,211 373
Concentrated 256 A2 37 1.5% 30.7% 64 18,267 208

REPEATABLE READ · up to 3 retries

Distribution Concurrency Configuration Successful TPS Raw conflict rate Final failure rate p50 (ms) p99 overall (ms) p99 committed (ms)
Uniform 16 D1 841 0.6% 0.0% 19 39 39
Uniform 16 R1 1,447 0.6% 0.0% 10 30 30
Uniform 16 A1 869 0.5% 0.0% 16 54 54
Uniform 16 A2 989 0.5% 0.0% 10 62 62
Uniform 64 D1 2,268 2.5% 0.0% 28 65 65
Uniform 64 R1 1,760 2.5% 0.0% 33 94 94
Uniform 64 A1 1,082 2.1% 0.0% 54 162 162
Uniform 64 A2 933 2.6% 0.0% 70 197 197
Uniform 256 D1 8,230 9.8% 0.4% 28 96 92
Uniform 256 R1 1,404 10.2% 0.4% 165 442 433
Uniform 256 A1 700 7.6% 2.3% 173 12,641 545
Uniform 256 A2 120 1.8% 0.0% 98 325 325
Concentrated 16 D1 421 27.2% 4.9% 27 109 104
Concentrated 16 R1 534 27.1% 5.0% 10 970 62
Concentrated 16 A1 343 25.9% 5.2% 15 990 120
Concentrated 16 A2 374 23.9% 4.5% 11 990 144
Concentrated 64 D1 1,102 50.4% 18.0% 29 110 105
Concentrated 64 R1 187 48.8% 19.9% 34 3,502 1,172
Concentrated 64 A1 190 47.0% 18.0% 51 4,581 1,104
Concentrated 64 A2 96 40.4% 21.9% 68 5,011 1,533
Concentrated 256 D1 3,273 71.0% 38.9% 49 94 86
Concentrated 256 R1 56 4.3% 27.7% 92 11,218 239
Concentrated 256 A1 39 3.9% 30.6% 112 13,286 325
Concentrated 256 A2 24 1.3% 38.6% 86 16,537 322

READ COMMITTED · up to 3 retries

Distribution Concurrency Configuration Successful TPS Raw conflict rate Final failure rate p50 (ms) p99 overall (ms) p99 committed (ms)
Uniform 16 R1 1,488 0.0% 0.0% 10 25 25
Uniform 16 A1 872 0.0% 0.0% 17 42 42
Uniform 16 A2 1,003 0.0% 0.0% 10 61 61
Uniform 64 R1 1,893 0.0% 0.0% 32 86 86
Uniform 64 A1 1,173 0.0% 0.0% 53 119 119
Uniform 64 A2 985 0.0% 0.0% 68 181 181
Uniform 256 R1 1,543 0.0% 0.0% 160 355 355
Uniform 256 A1 1,213 0.0% 0.0% 203 442 442
Uniform 256 A2 595 0.0% 0.0% 425 852 852
Concentrated 16 R1 177 0.6% 1.1% 11 2,129 1,040
Concentrated 16 A1 236 0.5% 0.9% 16 1,269 1,040
Concentrated 16 A2 152 0.7% 0.9% 10 2,006 1,728
Concentrated 64 R1 60 2.3% 18.1% 41 4,491 1,798
Concentrated 64 A1 61 2.7% 23.3% 52 3,868 1,890
Concentrated 64 A2 52 2.5% 24.5% 62 4,026 1,834
Concentrated 256 R1 60 0.0% 28.5% 92 11,674 199
Concentrated 256 A1 38 0.0% 31.0% 100 13,553 318
Concentrated 256 A2 0 0.0% 100.0% 9,286 9,693 -

Summary of observations

  • Consistency: there were zero invariant violations in all 66 cells. The invariants checked were no negative inventory, inventory changes matching quantities sold, conservation of the total balance, exactly one effect per business ID, no partial application, and no confirmed lost commits. All four services preserved the business invariants with this implementation (conditional updates, receipts, retries, commit-status checks).
  • Uniform distribution: regardless of service, raw conflict rates were similar: about 0.5% at concurrency 16, about 2–3% at 64, and about 8–10% at 256 (A2 at concurrency 256 excluded, as explained below). This is because all four services reject same-row conflicts under REPEATABLE READ. With up to 3 retries, final failure rates dropped to 0–2.3%. The highest, 2.3%, was A1 at concurrency 256, caused by transfers that hit the 2-second deadline.
  • D1 under the concentrated distribution:
    • Raw conflict rates rose to 21%, 44%, and 67% at concurrency 16/64/256 (no retries).
    • With up to 3 retries, final failure rates fell to 4.9%, 18%, and 39%, but not to 0.
    • There was no lock waiting, so the p99 of requests that committed successfully was 110 ms or less.
  • The controls under the concentrated distribution:
    • Row lock queues formed. On R1 at concurrency 256, up to 206 sessions were waiting for locks.
    • Under REPEATABLE READ at concurrency 64 and 256, more transfers hit the 2-second deadline, successful TPS dropped to 24–512, and final failure rates reached 18–63%.
    • Under READ COMMITTED there were almost no conflict errors, but because of the same lock waits, final failure rates at concurrency 64/256 were 18–31% (A2 at concurrency 256 excluded).
  • Two A2 cells at concurrency 256: there were zero successes (REPEATABLE READ, uniform, no retries; and READ COMMITTED, concentrated). All failures were connection timeouts (ConnectionTimeout). With only 5 seconds of warm-up, 256 new connections arrived all at once while ACU was low. From this result alone we cannot tell whether it is an A2 capacity limit or an effect of the short warm-up, so it is not used in the conclusions.
  • Why throughput is not compared: at concurrency 256 with the uniform distribution, D1 delivered about 8,500 TPS. The controls were t4g.medium (burstable 2 vCPU) and delivered about 1,200–1,500 TPS. This difference comes from the scaled-down configurations, not from the services, so it is not interpreted as a throughput ratio. A capacity comparison at the planned scale is done in E002.

Development and operations

  • Provisioning time (from the create request until usable, measured once): D1 42 seconds, A1 5 minutes 30 seconds, A2 6 minutes 37 seconds, R1 12 minutes 8 seconds (Multi-AZ).
  • Deletion time (from the end of the run until deletion was confirmed): D1 about 2 minutes, R1 about 3 minutes, A1 about 11 minutes, A2 about 11 minutes.
  • Impact of DSQL’s row modification limit: DSQL limits the number of modified rows per transaction to 3,000 (official quota, checked on 2026-09-26). The reset between cells (deleting the previous cell’s orders and receipts) therefore had to be split into batches of 1,000 rows.
    • As a result, D1 took about 3–8 minutes per cell, about 58 minutes for 12 cells.
    • The controls took about 1 minute per cell with the same code. R1 took 16 minutes for 18 cells.
    • This is an observation that bulk delete and cleanup jobs become an operational burden on DSQL.
  • What additionally had to be handled on DSQL:
    • A table created after a transaction started was not visible in that transaction (42P01). We therefore changed the scenario tool to create tables first.
    • SHOW max_connections returned 20. The official quota is 10,000 connections per cluster, so we changed the tool not to limit concurrency by this value.
    • A DSQL server error occurred once, and the cell was run again.
  • Retry implementation effort: under REPEATABLE READ, all four services needed retries and used the same retry code (run_op in load.py, about 60 lines including the commit-status check). In contrast, when the controls use their default READ COMMITTED, there are almost no conflict errors, so they work without retries. DSQL does not support READ COMMITTED, so implementing retries is mandatory. The working time spent on the implementation was not measured.

Cost

Unit prices are Seoul Region On-Demand prices (USD) checked with the AWS Price List API on 2026-09-25–26. DSQL is USD 10 per million DPU (version published 2026-09-11). The estimated costs in the table below were calculated from usage and unit prices at the time of the run, and the actual charges were checked on 2026-09-28 in Cost Explorer (Seoul Region, 2026-09-25–26 UTC, by usage type). Standing account charges unrelated to the experiment (ELB, S3, VPC and so on) were excluded.

Item Usage Estimated cost Actual charge Basis and difference
D1 DPU 222,363 DPU (including pilot, scenarios, and resets) About $2.22 $1.22 The billed DPU (222,362) matched the CloudWatch TotalDPU total. The difference in amount is because the monthly free usage of 100,000 DPU was deducted.
R1 0.52 hours About $0.11 $0.06 The billed Multi-AZ usage time was 0.29 hours, shorter than the estimated time measured from the create request to deletion completion.
A1 0.50 hours About $0.07 $0.05 Billed time 0.42 hours. See the note on rates below
A2 1.91 ACU-hours About $0.50 $0.39 Billed 1.86 ACU-hours (3% off from the integral of CloudWatch ServerlessDatabaseCapacity). See the note on rates below
Load generator and network 4 Spot runners, about 3.9 hours in total About $0.45 About $0.34 Billed Spot usage 2.77 hours ($0.28), EBS, inter-AZ transfer, and public IPv4 about $0.06
Total   About $3.4 About $2.07 About 61% of the estimate
  • Why the estimate was higher than the actual charge: of the difference of about $1.3, about $1.0 is due to DSQL’s monthly free usage. The rest is because the running time used for the estimate (from the create request to deletion completion) was longer than the actual billed time. This month’s DSQL free usage was entirely used up in this experiment.
  • I/O-Optimized rates: according to CloudTrail records, the A1 and A2 clusters were created as aurora-iopt1 (I/O-Optimized) and not changed afterward. However, the billing records show all of A1 and 1.54 ACU-hours of A2 under Standard usage types and rates ($0.113/hour, $0.20/ACU-hour), and only 0.32 ACU-hours of A2 were billed at the I/O-Optimized rate. The I/O charges that would apply under Standard were not billed. The configuration was indeed I/O-Optimized, and we could not determine why the billing was classified this way (the Cost Explorer values at the time of the query are estimates and may change at month-end close).

  • DSQL DPU consumption: about 0.095 DPU per attempt in the D1 pilot. At the Seoul unit price this comes to about USD 0.95 per million attempts.
  • Cost of maximum-throughput load: because DSQL throughput was high, DPU cost grew quickly even over short periods under a closed-loop load that sends requests without pause.
  • Why unit costs are not compared: the controls are billed by running time, and the cells here are short at 25 seconds. Comparing cost per successful operation between services would therefore be meaningless. This comparison is calculated in E010 at the same arrival rate.

Conclusions and limitations

  • Consistency (question 1): all four services preserved the business invariants under contention and retries. DSQL also worked without double processing or lost commits once retries and commit-status checks were implemented.
  • Differences in contention behavior (question 2):
    • DSQL reports conflicts at commit time without lock waits. As a result, under concentrated contention latency stays short, but the conflict rate is high and failures remain even after retries (final failure rate 39% at concentrated, concurrency 256).
    • The controls try to absorb contention with lock waits, so under high contention the waits exceeded the deadline and throughput collapsed.
    • Therefore, when moving workloads with concentrated contention, such as popular products, to DSQL, two things should be considered first: the retry budget (count and deadline) and a schema design that spreads contention. One example is splitting inventory across multiple rows.
  • Implementation effort (question 3): DSQL has no READ COMMITTED, so retry handling is mandatory. The controls can skip retries by using READ COMMITTED.
  • Limitations:
    • With one repetition and short 25-second cells, variation cannot be reported. The measurement window may include states before caches and ACU had stabilized.
    • The controls were small burstable configurations, so absolute throughput and latency differ from the planned configurations.
    • Because the load was closed-loop, latency during periods when request intervals widen is missing from the measurement (coordinated omission). The p99 should therefore be interpreted only as the response time at that concurrency.
    • A2 results at concurrency 256 are not interpreted because of the connection timeouts.
    • The scenarios were observed once per isolation level.
    • As an external reference, there is a report that DSQL had high failure rates on a hot-key schema (Marc Bowes’s TPC-B post). It points in the same direction as these results, but we did not verify its numbers.
  • Next experiments:
    • E002: measures SLO capacity and latency at the planned scale with an open-loop method that fixes the request rate.
    • E003: covers connection surges, including A2’s connection timeouts and DSQL’s connection rate limits.

Cleanup record

  • Resources created:
    • BATCH: VPC, two subnets, IGW, two security groups, DB subnet group, IAM role and instance profile, 4 Spot runners (one c7g.2xlarge for the pilot and three c7g.4xlarge).
    • Per configuration: the D1 cluster, the R1 instance, the A1 and A2 clusters with writers, and three RDS-managed secrets.
  • Deletion completion times (UTC): D1 00:26:37, R1 00:58:04, A1 01:31:21, A2 02:07:38, BATCH 02:08:14. No resource failed to delete.
  • One runner missing from the manifest: because of a bug in the runner replacement code, one c7g.4xlarge runner (created at 22:45:53 UTC) was created without being recorded in the manifest. We confirmed its ownership through the experiment tags and terminated it manually, and termination was confirmed at 22:48 UTC. We fixed the code so that the same problem does not recur.
  • Verification (02:09:45 UTC): remaining_count=0 from e004.py verify. The first verification (02:08:19) returned 1 because one Spot request for a terminated runner was still in the active state; we verified again after the request changed to closed.
  • Manual cross-check: 0 DSQL clusters, 0 RDS instances, clusters, and cluster snapshots with the e004 prefix, IAM role NoSuchEntity, and 0 EC2 instances, volumes, ENIs, and Spot requests (open/active) with experiment tags.