Back to the experiment list
E006CompletedP0

Connection failures, per-service recovery, and preservation of successful commits

After the same connection failure, how do business recovery, commit preservation, and reconnection effort compare?

Results summary

This experiment checked whether, after the connection between the application and the database is briefly lost, processing returns to normal once the connection is restored, and whether commits that already succeeded are neither lost nor applied twice. At a minimum scope (MVP), while sending order transactions at 300 per second, we blocked the path so that packets from the load generator to the database were dropped for 30 seconds (a blackhole route). This is a client-side network failure, not a failure or failover of the database itself.

For both DSQL (D1) and Aurora Serverless v2 (A2), requests during the 30-second block failed with connection errors and timeouts, so the technical failure rate over the whole 150-second measurement window was 20.6% for D1 and 20.7% for A2. Because the block covered 20% of the measurement window, we interpret this as most requests during the block failing while the rest of the window was processed normally. After the block was lifted, the first query on a new connection succeeded after 0.02 seconds on D1 and 0.07 seconds on A2.

The most important result is consistency. We implemented requests whose success was unclear because no commit response was received so that they check the business ID receipt before retrying, and after each cell we reconciled receipts, orders, ledger, and inventory. Both services had 0 lost commits, 0 double-applied changes, and 0 inventory mismatches. DSQL could also preserve consistency during a lost connection when the same approach as the baseline (business ID receipts and checking whether a commit happened) was implemented.

Production readiness

Within the scope of this test, DSQL neither lost nor double-applied successful commits when the connection was cut for 30 seconds, and it reconnected immediately once the connection came back. However, this is because the application implemented business ID receipts and “check, then retry when the commit outcome is unclear,” and this implementation is required for both DSQL and the existing services. Requests during the block failed on both services, so recovery time depends on how long it takes for the network to come back. How DSQL behaves under internal failures or AZ failures could not be checked this time because there is no public means of fault injection, and Aurora’s failover time was not measured either.

Question

After the same connection failure, how do business recovery, commit preservation, and reconnection effort compare? This MVP covers only a common client connection failure.

Test conditions

  • Region and time: Seoul (ap-northeast-2), 2026-09-29. A2 08:40–08:43 UTC, D1 10:00–10:04 UTC (the first D1 attempt was rerun because of a defect in the measurement tool).
  • Configurations: D1 Aurora DSQL single-Region, A2 Aurora PostgreSQL 16.15 Serverless v2 (writer+reader, 4–32 ACU). Data was 2% of the E002 size.
  • Load: The same business mix as E002 (product lookup 40%, order history 30%, order creation 20%, cancellation 10%) at a fixed arrival rate of 300 TPS. 30 seconds of warm-up + 150 seconds of measurement, with the block applied for 30 seconds starting 40 seconds after measurement began.
  • Fault injection: On the runner, the route to the IPs that the database endpoint resolves to was blocked with ip route add blackhole and removed after 30 seconds (2 IPs for D1, 1 for A2).
  • Retries: Up to 3 attempts per request, with a total deadline of 2 seconds. A request that did not receive a commit response looks up the business ID receipt to check whether it committed before retrying.
  • Consistency check: Selecting only rows created by the cell, we checked that the number of receipts was within the range of commits confirmed by the client plus unclear cases, and that changes to orders, ledger, and inventory matched the receipts.
  • Load generator: Spot runners, D1 c6g.4xlarge, A2 m7g.4xlarge.
  • Deviations from the plan: We did not run the two load levels of 30%/80% of Q, 3 repetitions per condition, 10 minutes of observation before and after the block, or service-specific failover (RDS reboot with failover, Aurora cluster failover).

Performance results

Configuration Attempted TPS Successful TPS Technical failure rate First successful connection after block lifted Invariant violations
D1 289.6 225.6 20.6% 0.02 s 0
A2 292.4 237.6 20.7% 0.07 s 0
  • Most failures were connection errors (for order creation, 1,572 of 1,614 failures on D1 and 1,619 of 1,663 on A2). The rest were deadline overruns of 2 seconds and a small number of serialization conflicts.
  • Because the cells included the blocked period, p99 increased to 1.5–1.9 seconds for both services. Latency outside the blocked period was at the same level as the E002 results (D1 write p95 about 29–30 ms, A2 about 9 ms).
  • On D1, 1,693 cancellation requests were rejected as “order already cancelled.” This is because earlier E009 tests had accumulated cancellations on the same data, and it does not affect the consistency verdict.

Development and operations

  • Both services recovered simply by the connection pool discarding broken connections and opening new ones, and no additional work specific to DSQL was needed. DSQL requires an IAM token for each new connection, but tokens were cached for 10 minutes, so they had almost no effect on reconnection latency.
  • Measurement tool defect: in the first D1 attempt, an administrative connection opened before the cell started was cut by the block, so the post-cell consistency check failed. We changed the tool to run the check on a new connection after the cell ended and measured again.

Cost

The total cost of run B is summarized in the cost section of the E009 report. This test applied about 3 minutes of load per configuration.

Conclusions and limitations

  • Under a 30-second client-side connection failure, DSQL and A2 showed the same results: requests failed during the block, reconnection happened right after it was lifted, and there were 0 lost or duplicated commits.
  • Limitations: one run, small data, and low load (300 TPS). We do not generalize the absence of loss and duplication in this single run into an absolute guarantee. Database failover and DSQL internal failures were not covered.

Cleanup record

  • Resources created in run B: BATCH (VPC, 2 subnets, IGW, 2 security groups, DB subnet group, runner IAM role and instance profile), the D1 DSQL cluster, the A2 cluster with writer and reader, an RDS-managed secret, and 2 Spot runners (D1 c6g.4xlarge, A2 m7g.4xlarge). E007 additionally created an AWS Backup vault, a backup service role, 2 recovery points, 1 restored DSQL cluster, and a point-in-time restored A2 cluster and instance.
  • Deletion: The backup vault and the 2 recovery points were deleted at 09:35 UTC, and the restored DSQL cluster and the A2 restore were deleted from 09:48 UTC to about 10:05 UTC after checking their ownership tags. The rest was deleted with e002.py batch-down at 10:04–10:20 UTC. No resource failed to delete.
  • Verification (10:20:27 UTC): e002.py verify reported remaining_count=0. A manual cross-check confirmed 0 AWS Backup vaults, 0 RDS clusters and IAM roles with this run’s prefix, and 0 open Spot requests. The one DSQL cluster still present at that time belonged to run A (E002, E003, E008, E012), which was running concurrently.