← All writing
awsmobiledistributed-systemsmultipart

Resuming Uploads Without Starting Over

How the mobile client and S3 reconcile multipart sessions, refresh upload permissions, and recover when completion or cleanup changes remote state.

A transfer interrupted near the end of a large video makes the cost of a retry visible.

For a simple PUT, restarting means sending the file again. For multipart, a retry can preserve completed parts and repeat only missing work. S3 assembles the final object when the multipart upload is completed. AWS describes the protocol here.

The application still needs to decide what survived. That is where the implementation becomes interesting.

Choose the transfer mode explicitly

The backend offers multipart for videos. It also offers it for photos larger than 5 MiB when the caller opts in; the mobile client does.

The default part size is 5 MiB. Smaller photos use a single signed PUT.

The distinction affects pause and recovery. The native adapter cancels an individual PUT when paused. Resuming a single-PUT upload starts from zero. Resuming multipart preserves parts S3 has already accepted, but the interrupted part itself may need retransmission.

There is no promise of arbitrary byte-offset resume.

Keep authorization separate from transport

The client initiates an upload through the API using its persisted request identity. The server creates or retrieves the canonical media record, then returns a multipart offer.

The offer includes the upload ID, part size, part count, completed parts and signed URLs for missing parts. Video offers also include signed destinations for preview frames.

Mobile → API: initiate or recover
API → S3: inspect multipart state
API → Mobile: completed parts + authorized missing parts
Mobile → S3: upload bytes
Mobile → API: complete

The API coordinates the operation while S3 receives the original bytes.

Recover remote facts rather than trusting a checkpoint

Consider this interruption:

S3 accepts part 3
    ↓
The app stops before saving the result locally

The local checkpoint is behind reality. If recovery trusted it exclusively, part 3 would appear missing.

The backend calls ListParts when producing another offer. S3’s accepted parts determine which upload work remains. The client stores the offer locally and uploads the next missing part.

The implementation creates one temporary slice for the current part and sends one part at a time per media task. Two media transfers can be active; it does not launch every part of a video concurrently.

This keeps temporary storage and concurrency bounded. Multipart is useful for recovery even without aggressive parallelism.

ETags are part of the protocol

A successful native part transfer can report an ETag. The client records the part number and ETag in its session, then clears the active native task and schedules more work.

A simplified checkpoint looks like this:

{
  "uploadId": "opaque-session-id",
  "partSize": 5242880,
  "partCount": 3,
  "completed": [
    { "partNumber": 1, "etag": "opaque-part-etag" }
  ],
  "framesDone": []
}

This is an illustration of the persisted shape, not a complete API response. An ETag here is a completion token; it should not be treated as the application’s SHA-256 content identity.

When no parts or frames remain, the client requests completion. The backend checks the completion list against remote part information before assembling the object.

An absent session has more than one explanation

NoSuchUpload is ambiguous from the client’s perspective.

The upload may have completed already. It may have been aborted. A lifecycle rule may have removed an abandoned session.

The backend first checks whether the original exists with the expected size and content type. If it does, recovery returns a processing or later status instead of starting another upload.

If the original is absent, it conditionally clears the old upload ID and creates a new session. The conditional update prevents stale recovery from clearing a session another caller has installed.

Session creation also handles competing callers: attaching an upload ID is conditional, and a losing caller aborts the extra session it created.

These checks address specific races. They do not turn S3 and DynamoDB into one transaction.

Permissions expire while work waits

A native upload can wait for Wi-Fi longer than its URL remains valid. Presigned URL validity depends on both its configured lifetime and the credentials used to sign it. AWS explains those expiration rules.

Before starting a native task, the adapter checks the signature timestamp and lifetime with a fifteen-minute margin. A stale offer is rejected so the persistent queue can obtain fresh authorization.

The content identity and request identity stay stable while the temporary permission changes.

Cleanup limits how long resume remains possible

Terraform configures this rule on originals:

abort_incomplete_multipart_upload {
  days_after_initiation = 7
}

S3 can abort incomplete uploads through lifecycle configuration, removing the need for a separate cleanup worker. AWS documents that mechanism.

The tradeoff is explicit: parts from a sufficiently old abandoned session may be gone. Recovery can establish a new session, but it cannot recover deleted parts.

The useful guarantee is conditional. While the session and source file survive, accepted parts can be reused. After cleanup, the operation can restart under the same logical media identity.

Thanks for reading.Back to the notebook →