Back to the experiment list
E011CompletedP0

DSQL development and operations convenience and adoption burden

How much does DSQL reduce or increase the working time and amount of change for initial adoption and recurring operations?

Results summary

This experiment summarized how much infrastructure operations work decreases when adopting DSQL, and what burden is added to the application and operating procedures in exchange. As a minimum scope (MVP), without new AWS runs, we collected the work we actually went through from E001 to E012 in creating, loading, putting load on, recovering, and deleting four services (DSQL, RDS Multi-AZ, Aurora Provisioned, Aurora Serverless v2). We did not perform the planned per-task time measurements or three repeated runs, so we did not calculate a working-time reduction rate.

DSQL clearly required less infrastructure work. A DSQL cluster was usable 32 seconds after the create request, and deletion took about 2 minutes. Aurora Serverless v2 (writer+reader) took about 11 minutes to create and about 15 minutes to delete. DSQL had no step for choosing instance size or minimum and maximum capacity, no password to store, no maintenance settings such as vacuum, and no read/write routing, and capacity did not need adjusting even as load rose to 24,000 TPS.

On the other hand, the application-side burden increased with DSQL. 16 of 35 workload SQL statements had to be changed (E001), and code to retry commit-time conflicts was mandatory (E004). Transactions cannot exceed 3,000 rows and 300 seconds, so loading, deletion, and batch code had to be split (E002, E008); asynchronous indexes needed a separate completion check, and indexes on the foreign-key side had to be created manually. Connections had to be rotated every hour and IAM tokens signed (E003), and because there is no point-in-time recovery, recovery from mistakes was tied to the backup interval (E007).

Production readiness

DSQL greatly reduces “the work of operating a database server.” There was no need to think about capacity planning, instance replacement, password rotation, vacuum, or reader routing, and cluster creation took just over 30 seconds. In exchange, a substantial part of that burden moves to the application. Existing PostgreSQL code needs SQL changes, retries, transaction splitting, and connection rotation to be redesigned, and because there is no point-in-time recovery, the backup interval and recovery procedure must be defined separately. The benefit is large for a new service designed around these constraints from the start, and the burden is large when migrating an existing service that relies heavily on stored procedures and large batches. Working time was not measured.

Question

How much does DSQL reduce or increase the working time and amount of change for initial adoption and recurring operations?

Test conditions

  • Additional AWS runs: None. Compiled from the run records of E001–E012, the change history of the tooling code, and the published reports.
  • Subjects: D1 Aurora DSQL, R1 RDS PostgreSQL 16 Multi-AZ, A1 Aurora PostgreSQL 16 Provisioned, A2 Aurora PostgreSQL 16 Serverless v2 (Seoul).
  • Deviations from the plan: We did not measure actual working time and waiting time per task, perform a first implementation plus three independent re-runs, test permission-error diagnosis, or manage per-service diffs against a baseline PostgreSQL app. The times in the table below are service waiting times read from automation tool logs.

Performance results

Infrastructure work (based on automation logs)

Task DSQL Aurora Serverless v2 (writer+reader) Source
Create request → available 32 seconds about 11 minutes E005–E009 run B, 2026-09-29
Delete request → deletion complete about 2 minutes about 15 minutes E002 second measurement, run B
Capacity selection None ACU minimum and maximum, reader promotion tier E002
Authentication setup IAM permission (dsql:DbConnectAdmin) Managed password (Secrets Manager) E002
Read routing None (single endpoint, immediately up to date) Reader endpoint; reads right after writes go to the writer E005
Responding to load increases None (up to 24,541 TPS) Automatic within the maximum ACU, upper limit 32 ACU E002, E009
Recovering from mistakes Prepare backup vault and role, full backup → new cluster (129 seconds) Point-in-time restore → new cluster + instance (587 seconds) E007

Work added to the application and procedures (DSQL)

Item Required change Source
SQL compatibility 16 of 35 changed: sequence and identity, serial, index creation, temporary tables and partitions, PL/pgSQL and triggers, isolation levels, statement_timeout, server-side cancellation E001
Concurrency Retrying commit-time conflicts (40001) and checking whether a commit happened by business ID are mandatory. No READ COMMITTED E004
Transaction size and time Splitting loading, deletion, and batches to fit the 3,000-row and 300-second limits E002, E008
Indexes Check indisvalid after CREATE INDEX ASYNC. Create indexes on the foreign-key side manually E002, E008
Connections Rotation matched to the 1-hour connection lifetime, IAM token signing and caching, preventing signing contention on simultaneous first connections E002, E003
Transient errors Retrying XX000 server unavailable during loading E002
Bulk deletion Slow enough that the measurement tool’s reset method had to be changed E002
Recovery No point-in-time recovery. RPO is determined by the backup interval; after restore, switch to the new endpoint and permissions E007
Diagnostics SHOW max_connections returns 20 (the actual quota is 10,000) E004

Development and operations

  • On the infrastructure operations side, the work DSQL reduced was waiting for creation and deletion, capacity planning, password management, maintenance settings, and read routing. These tasks recur in every round of operations, so the effect grows with the number of services.
  • On the application side, most of the added work is code that can be reused once implemented (retries, transaction splitting, token signing, connection rotation), but an existing codebase needs changes across all areas.
  • The load tool for this study was also modified several times to support DSQL: server unavailable retries, resuming loading per chunk, waiting for asynchronous index completion, adding foreign-key indexes, transaction splitting, a lock around token signing, and a change to the reset method.

Cost

This report itself did not create any AWS resources. The costs of the cited runs are summarized in each report and in E010.

Conclusions and limitations

  • DSQL reduces infrastructure operations work and increases the application design burden. The benefit is large for newly designed services, and the migration cost is large for services that rely on existing PostgreSQL features.
  • Limitations: working time was not measured, and the time people actually spent and the learning cost were not compared in monetary terms. Recurring operations work on the control side (patching, instance replacement, storage expansion) was not performed. The observations are cases encountered in this study and do not generalize to all workloads.

Cleanup record

  • This report did not create new AWS resources. For the three cited measurement runs, each report confirmed 0 remaining resources (2026-09-28 12:07:44, 2026-09-29 11:10:30, and 10:20:27 UTC).