Back to the experiment list
E007CompletedP1

Backup, point-in-time restore, and recovery from mistakes

How do DSQL's recovery features, time, and effort differ from existing services?

Results summary

This experiment compared, on DSQL and Aurora Serverless v2 (A2), the procedure and time needed after an operator changes data by mistake to bring back the pre-mistake data so that the application can read it again. At a minimum scope (MVP), we wrote 30 markers at one-second intervals on top of about 100 MB of data, then created a “mistake” that overwrote all markers with -1 and inserted one new row, and recovered using each service’s method.

DSQL does not support point-in-time restore (PITR) to an arbitrary time; the only method is to restore a full AWS Backup backup into a new cluster. We therefore took an on-demand backup after writing the markers and before the mistake. The backup (about 102 MB) took 481 seconds, and the restore job that created a new cluster from this backup took 129 seconds. A2 was restored to a specified point in time just before the mistake; the restored cluster was ready after 223 seconds, and it became connectable after attaching an instance at 587 seconds.

In both restores, the 30 correct markers and their sum matched the original, and there were 0 overwritten rows and 0 rows inserted after the mistake. The first connection to the restore took 343 ms on D1 and 291 ms on A2. DSQL’s restore was fast, but the point that can be recovered is limited to “the time the last backup was taken,” so changes made between backups are lost.

Production readiness

Mistakes can be recovered with DSQL, but the biggest difference is that there is no point-in-time restore to “1 second before the mistake” as with existing Aurora. The most recent recoverable state is the last AWS Backup recovery point, so the backup schedule must be set to match the acceptable data loss window (RPO), and its cost must be accepted. With the small data in this test, the backup took 8 minutes and the restore 2 minutes, but every backup is a full backup, so time and cost may grow as data grows. A restore always creates a new cluster, so a procedure for switching the application’s connection address and IAM permissions to the new cluster must also be prepared.

Question

How do DSQL’s recovery features, time, and effort differ from existing services?

Test conditions

  • Region and time: Seoul (ap-northeast-2), 2026-09-29 09:24–09:47 UTC.
  • Configurations: D1 Aurora DSQL single-Region, A2 Aurora PostgreSQL 16.15 Serverless v2 (writer+reader, 4–32 ACU). Data was 2% of the E002 size (D1 backup size 101,940,777 bytes).
  • Markers and mistake: We wrote 30 rows at one-second intervals to the marker table mvp_marker (09:24:12–09:24:42 UTC), and at 09:32:54 changed all rows to v = -1 and inserted a row with id 100000.
  • D1 procedure: We created a dedicated backup vault and an AWS Backup service role (AWSBackupServiceRolePolicyForBackup, ...ForRestores), then proceeded in this order: on-demand backup after the markers → mistake → start-restore-job from the recovery point (deletion protection off) → new cluster ACTIVE → grant the runner access to the new cluster → verification. The AWS Backup opt-in setting for DSQL was already enabled in the account.
  • A2 procedure: After waiting until the latest restorable time passed the target time (0.2 seconds), restore-db-cluster-to-point-in-time to the midpoint between the last correct marker and the mistake → cluster available → create a db.serverless instance → available → verification.
  • Verification: From the runner, we opened a new connection to each restore and checked the number of markers, the number of overwritten rows, the number of rows added after the mistake, the marker sum, and the number of orders.
  • Deviations from the plan: We did not run 3 repetitions per method and size, a 7-day retention policy, restore under normal load with measurement of SLO recovery, a snapshot restore comparison, or R1 and A1.

Performance results

Configuration Recovery method Backup Restore (request→ready) First connection Markers Overwritten rows Rows after mistake
D1 On-demand full backup → new cluster 481 s 129 s 343 ms 30/30 0 0
A2 Point-in-time restore → new cluster + instance Not needed (continuous backup) 587 s (cluster 223 s) 291 ms 30/30 0 0
  • The marker sum (435) in both restores matched the original. The order counts were 282,772 for D1 and 282,773 for A2, values that reflect the rows accumulated by earlier tests on each original.
  • D1’s restore time does not include the backup time. In a real incident the backup must already exist, so the most recent recoverable point is the time of the last backup.

Development and operations

  • Preparing DSQL backups: An AWS Backup vault and a service role are required. With the CLI, a default vault is not created automatically. Querying a vault that does not exist returns AccessDeniedException rather than ResourceNotFound, so the automation code had to treat this as “no vault.”
  • Result of a DSQL restore: A restore creates a new cluster with deletion protection enabled by default. We turned off deletion protection through the regionalConfig metadata and copied the original tags. The new cluster has a different identifier and endpoint, so IAM access permissions had to be granted again.
  • Aurora point-in-time restore: The restore API does not accept the new managed password option and keeps the original’s master password. After the cluster is restored, an instance must be created separately before it can be connected to.
  • Cleanup: The vault cannot be deleted before backup jobs finish, so we waited for the running backup job to finish and then deleted the recovery points and the vault.

Cost

The total cost of run B is summarized in the cost section of the E009 report. The additional resources created by this test were 2 AWS Backup recovery points (about 102 MB, kept for a few minutes), the restored D1 cluster (about 10 minutes), and the restored A2 cluster and instance (about 15 minutes).

Conclusions and limitations

  • Both services accurately brought back the pre-mistake data. DSQL restored quickly, but without point-in-time restore its recovery point is tied to the backup schedule.
  • Limitations: one run, about 100 MB of data. Backup and restore times can vary with data size and service state. The full recovery time including restore under load and switching the application over was not measured.

Cleanup record

  • Resources created in run B: BATCH (VPC, 2 subnets, IGW, 2 security groups, DB subnet group, runner IAM role and instance profile), the D1 DSQL cluster, the A2 cluster with writer and reader, an RDS-managed secret, and 2 Spot runners (D1 c6g.4xlarge, A2 m7g.4xlarge). E007 additionally created an AWS Backup vault, a backup service role, 2 recovery points, 1 restored DSQL cluster, and a point-in-time restored A2 cluster and instance.
  • Deletion: The backup vault and the 2 recovery points were deleted at 09:35 UTC, and the restored DSQL cluster and the A2 restore were deleted from 09:48 UTC to about 10:05 UTC after checking their ownership tags. The rest was deleted with e002.py batch-down at 10:04–10:20 UTC. No resource failed to delete.
  • Verification (10:20:27 UTC): e002.py verify reported remaining_count=0. A manual cross-check confirmed 0 AWS Backup vaults, 0 RDS clusters and IAM roles with this run’s prefix, and 0 open Spot requests. The one DSQL cluster still present at that time belonged to run A (E002, E003, E008, E012), which was running concurrently.