Same workload,
how different is DSQL?

From throughput and latency to the hours spent on development and operations.
We compare DSQL directly with established databases and publish the evidence.

Read the adoption decision

DSQL usage guide · Browse the experiments

Experiment progress

Planned
0
Running
0
Completed
12

Published results

Measured and reviewed records
E001

SQL compatibility and executability

Of 35 SQL checks, DSQL passed 17 and rejected 16 as unsupported (0A000). The three control services passed 34, all except the DSQL-only CREATE INDEX ASYNC. The order transaction (including FKs) passed on DSQL without modification.

Measured 2026-09-24
E002

Limits of OLTP throughput and latency

In a DSQL-only re-measurement (2026-09-29), DSQL handled 24,541 TPS at a fixed concurrency of 256 connections while meeting the SLO (write p99 35.7 ms). Under the same conditions, the highest throughput at which the controls met the SLO was 5,839 TPS for RDS Multi-AZ, 3,375 for Aurora Provisioned, and 11,559 for Aurora Serverless v2. DSQL latency was almost independent of load, but in absolute terms write p95 was about 28–41 ms, longer than Serverless v2 (7–9 ms). The load generator was close to saturation, so the DSQL value is a lower bound, and the main measurement and repetitions were not done.

Measured 2026-09-29
E003

Connection count, connection surges, and pooling

MVP (DSQL only). Signing an IAM token took p50 0.2 ms, a small burden, and a new connection (TLS, authentication, first query) took p50 15.5 ms and p99 120 ms when run sequentially. Requesting 500 and 1,000 connections at once succeeded for all of them with no rejections. However, opening a new connection per request brought the read throughput of 16 workers to 121 TPS, about 1/70 of the rate with persistent connections (8,370 TPS). Using a connection pool is effectively mandatory; the control comparison and RDS Proxy were not measured.

Measured 2026-09-29
E004

Transaction contention and data consistency

There were zero business invariant violations across all 66 load cells and all isolation level scenarios. Under concentrated (hot) contention, DSQL failed with commit-time conflicts (40001) without waiting and kept p99 at or below 110 ms, but its raw conflict rate rose to 67–71% at concurrency 256. The controls (small configurations) saw lock waits exceed the 2-second deadline, and throughput dropped to tens of TPS. These are results from one repetition, 25-second cells, and small controls, and are not a basis for capacity comparison.

Measured 2026-09-26
E005

Read scaling and visibility of the latest data

MVP. In 200 trials that read the same row from another connection immediately after a commit was acknowledged, DSQL returned the new value on the first read in all 200 trials (p99 5.5 ms until visible). With the Aurora Serverless v2 reader endpoint, 199 of 200 first reads returned the old value, and the new value took p50 21.8 ms and p99 32.6 ms to become visible. Another connection on the writer endpoint saw it immediately. Read throughput scaling was not measured.

Measured 2026-09-29
E006

Connection failures, per-service recovery, and preservation of successful commits

MVP. During an order load of 300 transactions per second, we created a client-side network failure by dropping packets from the load generator to the database for 30 seconds. For both DSQL and Aurora Serverless v2, requests during the blocked period failed with connection errors, giving an overall cell failure rate of about 20.6%, but a new connection succeeded 0.02–0.07 seconds after the block was lifted, and both services had 0 lost commits and 0 duplicated effects when reconciled against business receipts. Database-side failover and DSQL internal failures were not tested.

Measured 2026-09-29
E007

Backup, point-in-time restore, and recovery from mistakes

MVP. On about 100 MB of data, we wrote 30 markers one second apart, then made a mistake that overwrote all the markers, and recovered. DSQL has no point-in-time restore, so we restored to a new cluster from an AWS Backup full backup taken before the mistake (481 seconds); the restore took 129 seconds. Aurora Serverless v2 was restored to a point in time just before the mistake; the cluster was ready in 223 seconds and the instance in 587 seconds. Both restores had all 30 correct markers intact, and 0 overwritten rows and 0 rows added after the mistake.

Measured 2026-09-29
E008

Data growth and the impact of operational tasks

MVP (DSQL only). Under an order load of 1,000 transactions per second, we built an asynchronous index on 11 million order rows. The command returned in 0.1 seconds, and it took 736 seconds until the index was actually usable. During the build, order creation p95 rose from 27.5 ms to 52.0 ms and p99 from 30.7 ms to 79.8 ms, but stayed within the SLO with 0 failures. Adding and dropping a column and dropping the index each finished within 0.1 seconds. A transaction changing 5,000 rows and a transaction held open for 310 seconds were rejected by the row-count limit and the 300-second limit, respectively.

Measured 2026-09-29
E009

Handling load spikes and resuming after idle: DSQL and existing services

MVP. Under a spike load that changed from 200 to 2,000 to 1,000 to 50 requests per second, both DSQL and Aurora Serverless v2 handled every request with zero failures and no manual intervention. DSQL's write p95 stayed at 29–36 ms regardless of load, while A2 was at 8–12 ms. After 15 minutes with no connections at all, DSQL answered the first request with a 112–287 ms connection and a 115–123 ms first query, but A2, which had auto-paused down to 0 ACU, resumed slowly and its first connection failed both times by exceeding the 15-second timeout.

Measured 2026-09-29
E010

Cost of the same workload and selection criteria

MVP (no additional AWS runs). Using measurements from E002 and E009, we calculated the hourly cost of the same order workload. DSQL costs about USD 0.31 per million requests, proportional to the number of requests (0.029–0.032 DPU per request), while RDS Multi-AZ cost about USD 1.22 per hour regardless of load. When average load is below about 1,100 TPS or idle time is long, DSQL is cheaper; when load above that level continues, a fixed instance is cheaper. Aurora Serverless v2 with a read replica was more expensive than or similar to DSQL at every measured load.

Measured 2026-09-29
E011

DSQL development and operations convenience and adoption burden

MVP (no additional AWS runs). We collected the work involved in actually creating, using, and deleting four services across E001–E012. DSQL was faster than Aurora Serverless v2 (about 11 minutes to create, about 15 minutes to delete), with 32 seconds to create and about 2 minutes to delete, and it needed no capacity selection, password management, vacuum, or reader routing. On the other hand, 16 of 35 workload SQL statements had to be changed, and additional code and procedures were needed for conflict retries, splitting transactions at 3,000 rows and 300 seconds, confirming completion of asynchronous indexes, manually creating indexes on the foreign-key side, rotating connections every hour, IAM token signing, and the absence of point-in-time recovery. Working time was not measured.

Measured 2026-09-29
E012

Large queries and aggregations on DSQL and interference with OLTP

MVP (DSQL only, S data). A JOIN for a customer's recent orders had a p50 of 4.4 ms; a daily sales aggregation over 10% of orders (1.1 million rows) took 2.4–2.6 seconds; a product ranking over about 270,000 order-item rows took 0.7 seconds; and a full aggregation over all 11 million order rows took about 36 seconds. One full aggregation used about 1,700 DPU (about USD 0.017). Even with the 10% aggregation repeated continuously during an order load of 1,000 per second, OLTP met the SLO with zero failures, and order-creation p99 rose from 30.7 ms to 41.8 ms. Comparison with control configurations and L-scale data were not measured.

Measured 2026-09-29

The order is OLTP, then bursts and idle periods, then large queries. Ease of use and cost are recorded from first creation to deletion. Experiment numbers are identifiers, not the run order.

E001

SQL compatibility and executability

What must change to implement the same workload on DSQL?

Result DSQL passed 17/35 unmodified, 16 unsupported (0A000). The three controls passed 34/35

Show the summary and production readiness

This experiment measured how much SQL has to be changed to move an existing PostgreSQL workload to Aurora DSQL. We ran 35 workload SQL items unchanged on DSQL (D1) and on three control services (RDS PostgreSQL, Aurora Provisioned, Aurora Serverless v2). DSQL passed 17 items without modification and rejected 16 items as “feature not supported” (SQLSTATE 0A000). The three control services passed all 34 items other than CREATE INDEX ASYNC, which is DSQL-only syntax.

The areas that needed changes on DSQL were sequence and identity declarations, the serial type, how indexes are created, temporary tables and partitions, PL/pgSQL functions and triggers, isolation level settings, statement_timeout, and server-side statement cancellation. In contrast, constraints such as PK, UNIQUE, CHECK and foreign keys, JOINs, CTEs, window functions, JSONB, and the core order transaction worked as written. SELECT FOR UPDATE was accepted syntactically but behaved differently. The controls made the other write wait, whereas DSQL did not make it wait and instead failed one side with a conflict (40001) at commit time. Applications moved to DSQL therefore need logic that retries failed transactions.

These results come from SQL feature checks only, on small configurations with a local client, and each configuration was run once. Performance, latency, and high availability were not evaluated, and because DSQL’s supported feature set keeps changing, the results should be read as observations as of 2026-09-24.

Production readiness

DSQL ran core OLTP SQL, such as constraints, foreign keys, JOINs, JSONB, and the order transaction, without modification, so for a new service designed around DSQL’s constraints from the start, the SQL-level barrier to adoption is low. Moving an existing PostgreSQL service, however, requires changing sequence and serial declarations, how indexes are created, PL/pgSQL and triggers, and temporary tables and partitions. In addition, REPEATABLE READ is the only isolation level, and statement_timeout and server-side cancellation do not work, so conflict retries and request deadlines must be implemented in the application. The more logic a system keeps in stored procedures and triggers, the higher the migration cost. This experiment is only a feature check; performance and availability are judged in E002 and later experiments.

Read the full report
CompletedP0Issue #1
E002

Limits of OLTP throughput and latency

How many times the SLO capacity and latency of each control does DSQL have?

Result Highest throughput passing the SLO at 256 connections: DSQL 24,541 TPS or more (lower bound), A2 11,559, R1 5,839, A1 3,375 TPS. DSQL write p95 about 28–41 ms, A2 about 7–9 ms

Show the summary and production readiness

This experiment checked how many more requests DSQL handles than the controls, and how its latency differs, when the same order workload runs under the same latency objective (SLO). The workload mixed 40% product lookups, 30% order history lookups, 20% order creation, and 10% order cancellation, and the SLO was read p95 50 ms and p99 100 ms, write p95 100 ms and p99 200 ms, and a technical failure rate of 0.1% or less. The experiment ran in two parts. On 2026-09-28 we brought up all four services together for a pilot and a concurrency search, but the DSQL search stopped because of slow resets and the budget; on 2026-09-29 we brought up DSQL alone and measured again without deleting data between cells.

With a fixed number of connections sending requests without pause (closed-loop), DSQL handled 6,680 TPS with 64 connections and 24,541 TPS with 256 connections, meeting the SLO in both. Even at 256 connections, order creation p99 was 35.7 ms, almost the same as with 64 connections (30.4 ms). In the control measurements with the same method (2026-09-28), write latency exceeded the SLO for all three controls at 256 connections, and the highest throughput that met the SLO was 5,839 TPS for RDS Multi-AZ (R1), 3,375 TPS for Aurora Provisioned (A1), and 11,559 TPS for Aurora Serverless v2 (A2). At 256 connections, DSQL handled about 2.1 times A2, about 4.2 times R1, and about 7.3 times A1. At that point the load generator CPU was at 80.8%, close to the saturation threshold (85%), so the DSQL value is a lower bound for this measurement, not a service limit.

In absolute terms, DSQL latency was longer. At the same 1,600 TPS, order creation p95 was 41.0 ms on DSQL (2026-09-28) and 6.8 ms on A2, and in the 2026-09-29 measurement DSQL write p95 was about 28–30 ms. When requests were sent at a fixed arrival rate (open-loop), DSQL met the SLO up to 6,750 TPS, and p99 exceeded 200 ms at 8,100 TPS. However, this failure came not from DSQL processing latency (p95 unchanged at about 28 ms) but from time the load generator spent waiting for the connections that each process shares (wait p99 about 190 ms), so we interpret the open-loop limit as a connection pool limit of the measurement tool. There were zero business invariant violations in every cell. The planned main measurement at common arrival rates and the repetitions were not done, so the differences between services are single-measurement values.

Production readiness

In terms of throughput, DSQL scaled the most in this experiment. At 256 connections it handled more than 24,000 mixed order operations per second without a rise in latency, while all the existing services exceeded the latency objective under the same conditions. Holding up as load grows without capacity planning is an advantage for services with large traffic swings. On the other hand, the latency of a single request was 3–6 times longer than Aurora (write p95 about 30–40 ms), so screens that run several queries in sequence per request will see longer response times. Reaching the same throughput requires more connections and more concurrency, and sustaining high throughput increases DPU charges (E010).

Read the full report
CompletedP0Issue #2
E003

Connection count, connection surges, and pooling

What burden do authentication, connection renewal, and surge handling impose on DSQL?

Result MVP: new connection p50 15.5 ms, 0 rejections for 1,000 concurrent connections. New connection per request 121 TPS vs persistent connections 8,370 TPS (about 1/70)

Show the summary and production readiness

This experiment checked how much of a burden IAM authentication and connection setup impose when an application connects to DSQL, and whether rejections or delays occur when connections arrive all at once. As a minimum scope (MVP), we measured token signing, sequential connections, persistent connections compared with a new connection per request, and simultaneous requests for 500 and 1,000 connections, from one load generator against one DSQL cluster holding E002-scale data. No controls were measured this time.

DSQL authenticates with a token signed through IAM instead of a password. Token signing happens inside the runner without a network call and took p50 0.2 ms; only the first call took 131 ms because of client initialization. Opening a new connection and completing the first query took p50 15.5 ms, p95 109 ms, and p99 120 ms over 100 sequential runs.

Sixteen workers with persistent connections handled 8,370 product lookups per second (p50 1.9 ms), but when they opened a new connection per request this fell to 121 per second (p50 131 ms), about one seventieth. When 500 and 1,000 connections were requested at once, all succeeded without rejection, but establishing all of them took 3.7 seconds and 8.2 seconds respectively. This rate (about 120–135 connections per second) is similar to the throughput with a new connection per request, so this measurement could not distinguish whether it is a DSQL limit on the connection establishment rate or the limit of the runner’s single Python process in handling TLS connections. New connections right after the surge took p50 16.2 ms, the same as usual.

Production readiness

DSQL’s IAM token authentication itself imposed little burden. Signing took 0.2 ms, and reusing a token for several minutes removes the need to sign for every connection. Requesting 1,000 connections at once produced no rejections. However, each new connection took 15–120 ms, so a design that opens a new connection for every request (for example, short-lived functions that run without a connection pool) dropped to about 1/70 of the throughput with persistent connections. Use a connection pool, and because DSQL closes connections older than one hour, configure the pool to replace connections periodically. Controls and RDS Proxy were not compared.

Read the full report
CompletedP0Issue #3
E004

Transaction contention and data consistency

How do contention and retries on DSQL affect consistency, latency, and implementation effort?

Result Invariant violations 0/66 cells. Under concentrated contention DSQL failed with 40001 without waiting (conflict rate 67–71% at concurrency 256, p99 of successful requests 110 ms or less), while the small controls saw throughput collapse from lock waits

Show the summary and production readiness

This experiment checked whether DSQL and the controls preserve consistency when many requests update the same rows at the same time, and how conflicts and retries affect latency and failure rates. We ran order creation and account transfer workloads at concurrency 16/64/256, with uniform access and access concentrated on popular keys, and checked inventory, balance, and duplicate-processing invariants in every load cell. There were zero invariant violations across all 66 load cells. All four services worked without double processing or lost commits once conditional updates, business ID receipts, retries, and commit-status checks were implemented.

The way contention was handled differed greatly. DSQL reports conflicts (40001) at commit time without lock waits. As a result, under concentrated contention where 80% of requests go to the top 1% of keys, the p99 of requests that committed successfully was short, at 110 ms or less, but the conflict rate at concurrency 256 was 67–71%, and 39% still failed in the end even after up to 3 retries. The controls queued requests with row locks, and under concentrated contention at concurrency 64 and 256, waits exceeded the 2-second deadline, successful throughput dropped to 24–512 TPS, and final failure rates reached 18–63%. Because DSQL does not support READ COMMITTED, implementing retries is mandatory, and for workloads where contention concentrates, such as popular products, the retry budget and a schema design that spreads contention should be considered together.

Because of the cost cap (USD 5 at the time, later raised to 50), we scaled the controls down to small burstable instances and ran each cell only once for 25 seconds. The throughput differences are therefore not used as a basis for comparing services, and because the method uses fixed concurrency, latency while requests back up may be missed. A capacity and latency comparison at the planned scale will be measured again in E002 with a fixed request rate. The estimated cost calculated from usage was about USD 3.4, and the actual charge confirmed in Cost Explorer was about USD 2.07 (the main difference being the deduction of DSQL’s monthly free usage).

Production readiness

With retries and business ID receipts (idempotent processing) implemented, DSQL preserved consistency under contention without double processing or lost commits, so consistency itself was not an obstacle to production adoption. For workloads where contention is spread out, the conflict rate was similar to the existing services and retries resolved most conflicts. However, when writes concentrate on a few rows, as with inventory for popular products, a large share of failures remains even after retries (39% at concurrency 256), so such workloads are hard to move as they are without a schema redesign such as splitting inventory rows. The 3,000-row limit per transaction meant that bulk deletes and cleanup jobs had to be split, and a server error was observed once, so retries are mandatory. These results come from small controls and a single repetition, so capacity is judged in E002.

Read the full report
CompletedP0Issue #4
E005

Read scaling and visibility of the latest data

How do performance, freshness, and routing work in DSQL compare with distributing reads to readers?

Result MVP: on reads from another connection right after a write, DSQL stale reads 0/200, A2 reader stale reads 199/200 (p99 32.6 ms until visible)

Show the summary and production readiness

This experiment compared the visibility of the latest data in DSQL and Aurora Serverless v2 (A2), that is, whether a value can be read from another connection immediately after its write is committed. Instead of the full planned scope, we reduced it to a minimum scope (MVP): one connection inserted a row, and as soon as the commit response arrived, another connection queried that row every 10 ms. We repeated this trial 200 times.

On DSQL (D1), the new row was visible on the first query in all 200 trials. The time until the second connection saw the row was p50 2.1 ms and p99 5.5 ms, which is about the round-trip time of a single query. DSQL has only one connection endpoint, and committed data was readable immediately from any connection, so the application did not need a separate routing rule that sends reads right after a write to the writer.

On A2, another connection on the writer endpoint saw the row immediately in all 200 trials (p99 0.5 ms), but when querying through the reader endpoint, 199 of 200 trials did not find the row on the first query. The new row took p50 21.8 ms, p95 22.1 ms, p99 32.6 ms, and at most 65.5 ms to become visible on the reader. In a configuration that distributes reads to readers, requests that need “read your own write right after writing” must either be sent to the writer or tolerate a delay of tens of milliseconds. This time we did not measure read throughput scaling or performance as a function of the number of readers.

Production readiness

In terms of freshness, DSQL had the advantage. All 200 trials read the new value from another connection right after the commit, so there is no need for routing code that sends post-write reads to the writer to avoid reader replication lag, as with existing Aurora (p99 about 33 ms in this measurement). This simplifies the implementation of workloads that must read their own writes immediately, such as a screen that shows order details right after an order is placed. However, we did not measure how far read throughput can scale this time, and as seen in E002, DSQL’s read latency itself was about 4 times longer than that of the Aurora writer. Services with a very high share of reads should be evaluated together with the E002 capacity results.

Read the full report
CompletedP1Issue #5
E006

Connection failures, per-service recovery, and preservation of successful commits

After the same connection failure, how do business recovery, commit preservation, and reconnection effort compare?

Result MVP: after a 30-second connection block, DSQL and A2 both had 0 lost or duplicated commits, reconnection 0.02–0.07 s after the block was lifted (database failover not measured)

Show the summary and production readiness

This experiment checked whether, after the connection between the application and the database is briefly lost, processing returns to normal once the connection is restored, and whether commits that already succeeded are neither lost nor applied twice. At a minimum scope (MVP), while sending order transactions at 300 per second, we blocked the path so that packets from the load generator to the database were dropped for 30 seconds (a blackhole route). This is a client-side network failure, not a failure or failover of the database itself.

For both DSQL (D1) and Aurora Serverless v2 (A2), requests during the 30-second block failed with connection errors and timeouts, so the technical failure rate over the whole 150-second measurement window was 20.6% for D1 and 20.7% for A2. Because the block covered 20% of the measurement window, we interpret this as most requests during the block failing while the rest of the window was processed normally. After the block was lifted, the first query on a new connection succeeded after 0.02 seconds on D1 and 0.07 seconds on A2.

The most important result is consistency. We implemented requests whose success was unclear because no commit response was received so that they check the business ID receipt before retrying, and after each cell we reconciled receipts, orders, ledger, and inventory. Both services had 0 lost commits, 0 double-applied changes, and 0 inventory mismatches. DSQL could also preserve consistency during a lost connection when the same approach as the baseline (business ID receipts and checking whether a commit happened) was implemented.

Production readiness

Within the scope of this test, DSQL neither lost nor double-applied successful commits when the connection was cut for 30 seconds, and it reconnected immediately once the connection came back. However, this is because the application implemented business ID receipts and “check, then retry when the commit outcome is unclear,” and this implementation is required for both DSQL and the existing services. Requests during the block failed on both services, so recovery time depends on how long it takes for the network to come back. How DSQL behaves under internal failures or AZ failures could not be checked this time because there is no public means of fault injection, and Aurora’s failover time was not measured either.

Read the full report
CompletedP0Issue #6
E007

Backup, point-in-time restore, and recovery from mistakes

How do DSQL's recovery features, time, and effort differ from existing services?

Result MVP: DSQL has no PITR, recovered the pre-mistake state with a 481 s backup + 129 s restore. A2 PITR 587 s. Markers 30/30 correct in both restores

Show the summary and production readiness

This experiment compared, on DSQL and Aurora Serverless v2 (A2), the procedure and time needed after an operator changes data by mistake to bring back the pre-mistake data so that the application can read it again. At a minimum scope (MVP), we wrote 30 markers at one-second intervals on top of about 100 MB of data, then created a “mistake” that overwrote all markers with -1 and inserted one new row, and recovered using each service’s method.

DSQL does not support point-in-time restore (PITR) to an arbitrary time; the only method is to restore a full AWS Backup backup into a new cluster. We therefore took an on-demand backup after writing the markers and before the mistake. The backup (about 102 MB) took 481 seconds, and the restore job that created a new cluster from this backup took 129 seconds. A2 was restored to a specified point in time just before the mistake; the restored cluster was ready after 223 seconds, and it became connectable after attaching an instance at 587 seconds.

In both restores, the 30 correct markers and their sum matched the original, and there were 0 overwritten rows and 0 rows inserted after the mistake. The first connection to the restore took 343 ms on D1 and 291 ms on A2. DSQL’s restore was fast, but the point that can be recovered is limited to “the time the last backup was taken,” so changes made between backups are lost.

Production readiness

Mistakes can be recovered with DSQL, but the biggest difference is that there is no point-in-time restore to “1 second before the mistake” as with existing Aurora. The most recent recoverable state is the last AWS Backup recovery point, so the backup schedule must be set to match the acceptable data loss window (RPO), and its cost must be accepted. With the small data in this test, the backup took 8 minutes and the restore 2 minutes, but every backup is a full backup, so time and cost may grow as data grows. A restore always creates a new cluster, so a procedure for switching the application’s connection address and IAM permissions to the new cluster must also be prepared.

Read the full report
CompletedP1Issue #7
E008

Data growth and the impact of operational tasks

How easy are DDL, diagnostics, and handling growth in DSQL, and what are the constraints?

Result MVP: 11-million-row async index under load in 736 s, write p99 during build 30.7→79.8 ms (within SLO, 0 failures). 5,000-row and 310-second transactions rejected by limits

Show the summary and production readiness

This experiment checked how easy it is in DSQL to change the schema or create indexes while a service is running, what impact this has on business transactions, and how DSQL’s transaction constraints actually show up. At a minimum scope (MVP), we ran index creation and column addition and removal on a DSQL cluster holding E002-scale data while sending order transactions at 1,000 per second, and compared the result with a cell that sent the same load only. The baseline was not measured this time.

When we created an index on the status column of 11 million order rows with CREATE INDEX ASYNC, the command returned in 0.1 seconds, and it took 736 seconds (about 12 minutes) until the index was actually usable (indisvalid). In the 5-minute measurement window during the build, order creation p95 rose from 27.5 ms to 52.0 ms and p99 from 30.7 ms to 79.8 ms, and order history lookup p99 also rose from 9.8 ms to 45.7 ms. Even so, all transactions stayed within the SLO, with 0 failures and 0 business invariant violations. The column addition, column drop, and index drop that followed each finished within 0.1 seconds.

DSQL’s constraints showed up clearly. An attempt to change 5,000 rows in one transaction was rejected with “transaction row limit exceeded” (54000), and a transaction held open for 310 seconds was rejected with “transaction age limit of 300s exceeded.” Bulk changes must be split into 3,000 rows or fewer, and work longer than 5 minutes must be split into multiple transactions. In E002 as well, these constraints forced us to split the loading and cleanup code, and bulk deletion was slow enough that we had to change the measurement method.

Production readiness

In DSQL, adding indexes and changing columns during operation could be done without stopping the service. During the 11-million-row index build, write latency rose 2–3 times but stayed within the target, with no failures. However, the index remains unusable for about 12 minutes after the command returns, so the deployment procedure must check indisvalid before enabling code that depends on the new index. The 3,000-row and 300-second transaction limits force bulk updates, bulk deletes, and long batch jobs to be split across multiple transactions, so existing batch and data migration scripts must be rewritten. Users have nothing to do for maintenance such as vacuum, but we did not compare the operational burden with the baseline.

Read the full report
CompletedP1Issue #8
E009

Handling load spikes and resuming after idle: DSQL and existing services

How do DSQL and existing serverless and fixed-capacity services compare in latency, manual intervention, and cost?

Result MVP: failures during the spike (up to 2,000 TPS) were 0 for both DSQL and A2. First connection after 15 minutes idle: DSQL 112–287 ms; auto-paused A2 exceeded the 15-second connection timeout 2 of 2 times

Show the summary and production readiness

This experiment compared DSQL and Aurora Serverless v2 (A2) on whether they keep their latency targets without anyone adjusting capacity when request volume suddenly rises, or when requests resume after a long quiet period. As a minimum scope (MVP), we sent one spike load that changed the arrival rate from 20% (2 minutes) to 200% (1 minute) to 100% (2 minutes) to 5% (1 minute) of a 1,000 TPS baseline. After that, we closed all connections, waited 15 minutes, and sent a first request; this idle test was run twice.

Under the spike load, both services processed every stage with zero failures and no manual intervention. DSQL’s order-creation p95 was 29.2–36.0 ms and barely changed even when the arrival rate increased tenfold; p99 in the 2,000 TPS stage was 51.5 ms. A2, configured with a 4–32 ACU range, saw its writer grow from 10.5 ACU to a maximum of 21 ACU; its order-creation p95 was 8.2–12.1 ms, and p99 in the 2,000 TPS stage was 37.9 ms. This spike (up to 2,000 TPS) was smaller than the capacity of either service as measured in E002, so these results cannot tell us how the services handle spikes near their capacity limits.

The first request after idle showed a large difference. Even after 15 minutes with no connections, DSQL responded with a first connection of 112–287 ms and a first query of 115–123 ms, after which queries returned to about 2 ms. For this test, A2 was configured with a minimum capacity of 0 ACU and auto-pause after 300 seconds, and CloudWatch confirmed that it actually paused down to 0 ACU. In that state, the first connection failed both times because it did not complete within the driver’s 15-second connection timeout; that connection attempt triggered the resume, and capacity came back within about 1 minute.

Production readiness

For services whose requests arrive sparsely or stop overnight, DSQL had a clear advantage. DSQL handled the first request within 0.3 seconds even after 15 minutes idle, and no compute charges accrue while it is idle. Aurora Serverless v2 can also eliminate idle cost with 0 ACU auto-pause, but the first connection failed after exceeding 15 seconds during the resume, so connection timeouts and retries must be set long. At this spike size, both services handled the load without manual intervention, so spike handling near capacity limits must be judged together with the E002 capacity results. These results come from a single run on small data.

Read the full report
CompletedP0Issue #9
E010

Cost of the same workload and selection criteria

What costs and constraints come with DSQL's performance and convenience advantages?

Result MVP: DSQL about USD 0.31 per million requests. Break-even with RDS Multi-AZ (about USD 1.22 per hour) at an average of about 1,100 TPS. Daily cost for 8 hours active + 16 hours idle: DSQL about USD 9 vs RDS about USD 29

Show the summary and production readiness

This experiment calculated how the cost of DSQL and existing services changes with the shape of the load, for the same workload and the same latency target. As a minimum scope (MVP), we did not launch any new AWS resources; instead, we multiplied values measured in E002 (throughput, DPU, ACU) and E009 (idle behavior) by Seoul Region On-Demand prices. The workload is the E002 order mix (40% product lookup, 30% order history, 20% order creation, 10% cancellation).

DSQL is billed in DPUs for the amount of work it processes, and this workload used 0.029–0.032 DPU per request. At the Seoul price ($10 per million DPUs), that is about $0.31 per million requests, or about $1.12 to sustain 1,000 TPS for one hour. By contrast, RDS Multi-AZ (db.r6g.xlarge, gp3 400 GiB) costs about $1.22 per hour regardless of load and handled up to 5,839 TPS within the SLO. Therefore, DSQL is cheaper when average load is below about 1,100 TPS, and a fixed instance is cheaper when load above that level continues. Aurora Provisioned (writer+reader, about $1.25 per hour + I/O) shows a similar break-even point.

The difference grows depending on the shape of the load. For a constant 1,000 TPS load over 24 hours, the daily cost was similar: DSQL about $27 and RDS Multi-AZ about $29. For a service that works at 1,000 TPS for only 8 hours and is idle for 16, DSQL cost about $9 and RDS about $29, so DSQL was about one third. Aurora Serverless v2 (A2) with a read replica scales the writer and reader together, costing about $3.4 per hour at 1,000 TPS, and because of its minimum capacity it is billed during idle time too, for about $53 per day. Using 0 ACU auto-pause reduces this to about $27, but in E009 the first connection failed after exceeding 15 seconds.

Production readiness

On cost alone, DSQL favors “services with few or irregular requests” and is at a disadvantage for “services with high load all day.” For this workload the break-even point was an average of about 1,100 TPS. Below that, DSQL is cheaper because the cost of keeping a fixed instance running disappears and high availability is included by default. Conversely, if thousands of TPS are processed continuously, the DPU charges that scale with request count exceed the cost of a fixed instance. Large aggregations are also billed by the amount processed (E012), so services with heavy analytical workloads need a separate cost calculation. This calculation is an estimate based on one workload mix, On-Demand prices, and measurements from a single run.

Read the full report
CompletedP0Issue #10
E011

DSQL development and operations convenience and adoption burden

How much does DSQL reduce or increase the working time and amount of change for initial adoption and recurring operations?

Result MVP: infrastructure work decreased (32-second creation; no capacity, password, vacuum, or routing work), but the application-side burden increased (16/35 SQL statements changed; handling retries, transaction splitting, asynchronous indexes, connection rotation, and the absence of PITR)

Show the summary and production readiness

This experiment summarized how much infrastructure operations work decreases when adopting DSQL, and what burden is added to the application and operating procedures in exchange. As a minimum scope (MVP), without new AWS runs, we collected the work we actually went through from E001 to E012 in creating, loading, putting load on, recovering, and deleting four services (DSQL, RDS Multi-AZ, Aurora Provisioned, Aurora Serverless v2). We did not perform the planned per-task time measurements or three repeated runs, so we did not calculate a working-time reduction rate.

DSQL clearly required less infrastructure work. A DSQL cluster was usable 32 seconds after the create request, and deletion took about 2 minutes. Aurora Serverless v2 (writer+reader) took about 11 minutes to create and about 15 minutes to delete. DSQL had no step for choosing instance size or minimum and maximum capacity, no password to store, no maintenance settings such as vacuum, and no read/write routing, and capacity did not need adjusting even as load rose to 24,000 TPS.

On the other hand, the application-side burden increased with DSQL. 16 of 35 workload SQL statements had to be changed (E001), and code to retry commit-time conflicts was mandatory (E004). Transactions cannot exceed 3,000 rows and 300 seconds, so loading, deletion, and batch code had to be split (E002, E008); asynchronous indexes needed a separate completion check, and indexes on the foreign-key side had to be created manually. Connections had to be rotated every hour and IAM tokens signed (E003), and because there is no point-in-time recovery, recovery from mistakes was tied to the backup interval (E007).

Production readiness

DSQL greatly reduces “the work of operating a database server.” There was no need to think about capacity planning, instance replacement, password rotation, vacuum, or reader routing, and cluster creation took just over 30 seconds. In exchange, a substantial part of that burden moves to the application. Existing PostgreSQL code needs SQL changes, retries, transaction splitting, and connection rotation to be redesigned, and because there is no point-in-time recovery, the backup interval and recovery procedure must be defined separately. The benefit is large for a new service designed around these constraints from the start, and the burden is large when migrating an existing service that relies heavily on stored procedures and large batches. Working time was not measured.

Read the full report
CompletedP0Issue #11
E012

Large queries and aggregations on DSQL and interference with OLTP

What are DSQL's large-query performance, OLTP interference, and tuning burden?

Result MVP: full aggregation over 11 million rows about 36 seconds (about 1,700 DPU), aggregation over 1.1 million rows 2.4–2.6 seconds. OLTP failures 0 while repeating aggregations; order-creation p99 30.7→41.8 ms

Show the summary and production readiness

This experiment checked how long large JOIN and aggregation queries take on DSQL and how much they cost, and how much OLTP (the order workload) slows down when such queries run at the same time. As a minimum scope (MVP), instead of the planned L scale, we ran four queries with different selectivity on the E002 S-scale data (11 million orders, 27.5 million order-item rows), and ran an interference test that repeated an aggregation during an OLTP load of 1,000 per second. Control configurations were not measured this time.

The JOIN for a customer’s 20 most recent orders, which uses an index, had a p50 of 4.4 ms and a p95 of 32.9 ms over 30 runs. The query that aggregates 10% of orders (about 1.1 million rows) by date took 2.4–2.6 seconds; the window query that ranks products using order items in a 1% range of orders (about 270,000 rows) took 0.7 seconds; and the query that aggregates all 11 million order rows by status took about 36 seconds. Running the same query twice produced the same row count and result hash, and the second run took almost the same time as the first, so there was no visible speed-up from repeated execution.

DSQL charges in DPUs for the amount a query processes. Estimated from per-minute usage, one full aggregation used about 1,700 DPU (about $0.017), and one 10% aggregation used about 100–150 DPU. When the 10% aggregation was repeated continuously (124 times in 5 minutes, p50 2.4 seconds) while sending an OLTP load of 1,000 per second, OLTP met the SLO with zero failures; order-creation p95 rose from 27.5 ms to 32.0 ms, p99 from 30.7 ms to 41.8 ms, and product-lookup p99 from 6.9 ms to 16.4 ms.

Production readiness

On DSQL, index-based lookups are fast, but aggregations that scan millions of rows or more took from seconds to tens of seconds (about 36 seconds for 11 million rows), and because of the 300-second transaction limit, larger aggregations must be split. Server-side statement_timeout and statement cancellation do not work (E001), so the means to stop a long-running query midway are also limited. Even with repeated aggregations, OLTP stayed within its targets, so interference was small. However, aggregations incur DPU charges proportional to the data processed, so for workloads with frequent large aggregations, such as periodic reports or dashboards, separating them into an analytics store is the safer choice in terms of cost. Speed was not compared against control configurations.

Read the full report
CompletedP1Issue #12

What do we compare, and how?

We do not assume that DSQL wins. We judge only results that meet the same workload, consistency, and SLO.

Performance
Successful throughput within deadlines and p95/p99 latency, including connection and retry time, with the variation across repeated runs.
Ease of use
Actual work time, manual steps, and code changes for first connection, migration, scaling, diagnosis, and recovery, separating first-time learning from repeated operations.
Cost and constraints
Cost per successful task under the same requirements. We state unsupported features and unmeasured items and delete every experiment resource when done.