Read scaling and visibility of the latest data
How do performance, freshness, and routing work in DSQL compare with distributing reads to readers?
Results summary
This experiment compared the visibility of the latest data in DSQL and Aurora Serverless v2 (A2), that is, whether a value can be read from another connection immediately after its write is committed. Instead of the full planned scope, we reduced it to a minimum scope (MVP): one connection inserted a row, and as soon as the commit response arrived, another connection queried that row every 10 ms. We repeated this trial 200 times.
On DSQL (D1), the new row was visible on the first query in all 200 trials. The time until the second connection saw the row was p50 2.1 ms and p99 5.5 ms, which is about the round-trip time of a single query. DSQL has only one connection endpoint, and committed data was readable immediately from any connection, so the application did not need a separate routing rule that sends reads right after a write to the writer.
On A2, another connection on the writer endpoint saw the row immediately in all 200 trials (p99 0.5 ms), but when querying through the reader endpoint, 199 of 200 trials did not find the row on the first query. The new row took p50 21.8 ms, p95 22.1 ms, p99 32.6 ms, and at most 65.5 ms to become visible on the reader. In a configuration that distributes reads to readers, requests that need “read your own write right after writing” must either be sent to the writer or tolerate a delay of tens of milliseconds. This time we did not measure read throughput scaling or performance as a function of the number of readers.
Production readiness
In terms of freshness, DSQL had the advantage. All 200 trials read the new value from another connection right after the commit, so there is no need for routing code that sends post-write reads to the writer to avoid reader replication lag, as with existing Aurora (p99 about 33 ms in this measurement). This simplifies the implementation of workloads that must read their own writes immediately, such as a screen that shows order details right after an order is placed. However, we did not measure how far read throughput can scale this time, and as seen in E002, DSQL’s read latency itself was about 4 times longer than that of the Aurora writer. Services with a very high share of reads should be evaluated together with the E002 capacity results.
Question
How do performance, freshness, and routing work in DSQL compare with distributing reads to readers? This MVP answers only the freshness part and the resulting need for routing.
Test conditions
- Region and time: Seoul (ap-northeast-2), 2026-09-29 08:38–08:40 UTC.
- Configurations:
- D1: Aurora DSQL single-Region cluster.
- A2: Aurora PostgreSQL 16.15 Serverless v2, 1 writer + 1 reader in another AZ (promotion tier 1), 4–32 ACU each, Aurora Standard.
- Data: 2% of the E002 data size (20,000 customers, 220,000 order rows, about 100 MB). We created a separate table
mvp_visibilityfor the visibility trials and dropped it afterwards. - Method: The write connection ran an
INSERTin autocommit mode, and from the moment the response arrived, the observing connection queried that row every 10 ms in a new transaction. If the row was not visible within 5 seconds, the trial was counted as censored. Each trial set had 200 repetitions. - Observing connections: For D1, a different connection to the same endpoint. For A2, two paths: a different connection to the writer endpoint, and the reader endpoint (confirmed to be a replica by
pg_is_in_recovery() = true). - Load generator: One Spot runner per configuration (D1
c6g.4xlarge, A2m7g.4xlarge). The visibility trials use only two connections, so the difference in runner specifications does not affect the results. - Deviations from the plan: We did not run the throughput and p99 comparison under read-heavy load, changes in the number of readers, the R2 configuration, or collection of engine replication metrics.
Performance results
| Configuration | Observation path | Old value on first query | Over 5 s | p50 until visible | p95 | p99 | Max |
|---|---|---|---|---|---|---|---|
| D1 | Same endpoint, different connection | 0/200 | 0 | 2.1 ms | 2.8 ms | 5.5 ms | 41.8 ms |
| A2 | Writer endpoint, different connection | 0/200 | 0 | 0.2 ms | 0.3 ms | 0.5 ms | 4.1 ms |
| A2 | Reader endpoint | 199/200 | 0 | 21.8 ms | 22.1 ms | 32.6 ms | 65.5 ms |
- The time until visible includes the round-trip time of the observing query itself. D1’s values are larger than those of the A2 writer not because of replication lag but because of the latency of a single DSQL query (reads of about 2–10 ms in E002).
- The A2 reader lag is measured in units of the 10 ms polling interval, so it has an error of about ±10 ms.
Development and operations
- DSQL has a single connection endpoint, so no read/write routing configuration was needed. A2 has a reader endpoint separate from the cluster endpoint (
.cluster-ro-), and the application must decide where to send reads that follow a write. - DDL to create and drop the test table on DSQL ran without problems.
Cost
This trial was part of run B (e002-20260929t082035z-afbe), which ran E005, E006, E007, and E009 together; the visibility trials themselves took about 2 minutes. The total cost of run B is summarized in the cost section of the E009 report.
Conclusions and limitations
- DSQL returned the new value from another connection right after the commit, while the A2 reader showed it tens of milliseconds late almost every time. This difference determines whether routing of post-write reads is needed.
- Limitations: 200 repetitions per trial set, one run. The measurements used small data with no load, so replication lag under heavy load may be larger. Read throughput scaling was not measured.
Cleanup record
- Resources created in run B: BATCH (VPC, 2 subnets, IGW, 2 security groups, DB subnet group, runner IAM role and instance profile), the D1 DSQL cluster, the A2 cluster with writer and reader, an RDS-managed secret, and 2 Spot runners (D1
c6g.4xlarge, A2m7g.4xlarge). E007 additionally created an AWS Backup vault, a backup service role, 2 recovery points, 1 restored DSQL cluster, and a point-in-time restored A2 cluster and instance. - Deletion: The backup vault and the 2 recovery points were deleted at 09:35 UTC, and the restored DSQL cluster and the A2 restore were deleted from 09:48 UTC to about 10:05 UTC after checking their ownership tags. The rest was deleted with
e002.py batch-downat 10:04–10:20 UTC. No resource failed to delete. - Verification (10:20:27 UTC):
e002.py verifyreportedremaining_count=0. A manual cross-check confirmed 0 AWS Backup vaults, 0 RDS clusters and IAM roles with this run’s prefix, and 0 open Spot requests. The one DSQL cluster still present at that time belonged to run A (E002, E003, E008, E012), which was running concurrently.