Synchronization Is a Reconciliation Problem
Persistent local state, account-scoped idempotency, and an Android failure that showed why bounded workers must make durable progress.
The difficult question in photo backup is what the next execution should do after the previous one stopped.
A phone may know that an upload started. Its native transfer service may know that it finished. S3 may already hold the original. The cloud catalogue may still be waiting for the preview worker.
Each observation describes a different part of the operation. Recovery consists of reconciling them.
Give each record one job
The local database separates three concerns:
| Table | Responsibility |
|---|---|
backup_assets |
Discovered library assets, versions, exclusions and preparation errors |
upload_tasks |
Durable transfer intent, request identity, working file and upload session |
media_items |
Local projection of the cloud catalogue for the UI |
Settings hold additional coordination state: whether backup was requested, scan position, network policy and leases. Each account has its own database file.
This structure avoids pretending that a library asset and a cloud media object are the same thing. Local identifiers describe the phone’s library. Content identity can connect different local assets to the same cloud object.
There are several state machines
Discovery records use these states:
DISCOVERED → QUEUED
│
├────→ SAVED
│
└────→ ERROR → later attempt
QUEUED means preparation produced or found a transfer task. SAVED is a discovery bookkeeping state: it can be set when a corresponding non-cancelled local task exists, as well as when cloud content is recognized. It should not be interpreted by itself as proof that the original is remotely stored.
Transfer records distinguish PENDING_UPLOAD, UPLOADING, PAUSED, UPLOAD_FAILED, CANCELLED, PROCESSING, READY and PROCESSING_FAILED.
The cloud model is smaller:
PENDING_UPLOAD → PROCESSING → READY
│
└────→ FAILED
These states describe storage and processing rather than every mobile interaction. A percentage reaching 100% does not establish that a preview is ready.
Persist intent before depending on a process
Preparation copies the original into application storage and hashes that stable working file. It then persists the task and its identifiers.
The task contains the local asset ID, working-file path, size, MIME type, content hash, request ID, media ID, retry schedule and native task identity. Multipart uploads add session information and the current part or frame.
On restart, the application compares these records with active native transfers and persisted native results. The database is essential, but it is not the authority for bytes already accepted by S3.
Nor is it an infallible journal: a working file can disappear. The recovery path needs to report that condition rather than continually retry an unavailable source.
Idempotency and deduplication answer different questions
The backend stores records under an account partition:
PK = USER#<authenticated-user>
SK = MEDIA#<media-id>
SK = REQ#<client-request-id>
SK = CONTENT#<sha256>
A DynamoDB transaction reserves the request and, for new content, the media and content-hash mapping. A replay with matching metadata returns the existing object. Reusing the request ID with conflicting metadata produces a conflict.
Two different requests containing the same hash can converge on a canonical media record. The client must therefore accept the media ID returned by the server rather than assuming its proposed ID always wins.
This deduplication is scoped to an account. It is based on identical bytes, not visual similarity. It also depends on the supplied hash; the presence of a hash field alone is not evidence that every upload path independently recomputes it on the server.
A worker can run repeatedly and accomplish nothing
An Android backup failure exposed an architectural mistake in work ordering.
The application discovered more than two thousand assets. Before preparing transfers, the worker attempted to recognize already-saved content by exporting and hashing discovered originals across the library.
Each run had a time budget. For a large unsaved library, the recognition pass consumed that budget before queue preparation. A later run could perform the same expensive work again without creating a transfer.
The service was active, but durable progress was absent.
The correction keeps recognition and preparation together:
Check available preparation slots
↓
Choose a small number of candidates
↓
For each candidate:
recognize existing content
or prepare its upload immediately
↓
Persist the resulting state
Preparation is limited to two occupied queue slots. Discovery is paginated and persists its cursor. The deadline is checked between candidates; it is a cooperative budget, not a hard cancellation boundary around a slow export.
One optimization remains possible: recognition and preparation can hash the same unsaved original separately. Bounding that work fixed starvation without eliminating every repeated read.
Coordination does not guarantee execution
Foreground and headless workers use database leases to coordinate discovery and transfer pumping. A lease has an owner, an expiry and a heartbeat, allowing a later process to recover after an owner disappears.
Background work is scheduled through the mobile operating system. Its requested frequency does not guarantee exact execution times. Permissions, connectivity, account state and platform scheduling still constrain progress.
The practical design target is that each allowed execution can recover and move a bounded amount of work forward.
That same reasoning applies inside an individual large transfer. The next article follows a multipart session through interruptions and uncertain outcomes.