Building a Private Family Photo Vault Without Overengineering It
How I designed a private photo synchronization system by adding complexity only when a requirement justified it.
It started with a very ordinary problem.
I was moving from Android to iOS, and keeping my photos synchronized became another ecosystem decision. Google Photos and iCloud already solve much of this problem. NAS products and other storage services offer alternatives.
But switching devices made me think about where our family photos lived, who controlled access, and what preserving them would cost over time. Family members were also asking whether I could build a private place for our photos and videos.
The ambition was modest: a digital family vault that remained understandable and inexpensive to operate.
For most people, buying an existing solution is a sensible answer. I chose to build because I wanted control over the storage and wanted to understand the system myself.
That choice came with a rule:
Start with the simplest design, then add a mechanism when a requirement gives it a reason to exist.
Begin with the bytes
Remove the constraints and the architecture looks almost trivial:
Mobile app → S3
A user selects a photo. The application uploads it. There is little to explain.
This is a useful starting point for reasoning, even though it is insufficient for the application I wanted. Each additional constraint should explain a specific architectural change.
The application should not contain AWS credentials
A distributed mobile binary cannot safely hold long-lived AWS secrets. Each account also needs its own authorized view of the archive.
I use Cognito for identity and an HTTP API backed by Lambda for application operations. The API uses the authenticated identity to access metadata and derive object paths. The client does not choose another user’s storage prefix.
The backend is the control plane. It decides which operation is allowed.
The file bytes take a shorter path:
Mobile app → API → permission to upload
Mobile app ─────→ S3 → original bytes
The API issues temporary presigned URLs. This avoids routing every video through the application backend. Those URLs still grant access to their holder and must be treated as sensitive capabilities. AWS documents the access and expiration model.
Account isolation is the current boundary. A private application used by a family does not automatically imply shared albums or permissions between relatives; those would require another authorization model.
Uploading a picture is easier than synchronizing a library
A manual upload has a clear starting point. Automatic backup must discover work, remember it and recover after interruption.
The phone and the cloud have different knowledge. A photo can exist locally, have a prepared upload, already be stored remotely, or be waiting for preview generation.
A single isSynced flag cannot explain these situations.
The mobile application therefore keeps persistent metadata in SQLite through Drift. Discovery, transfers and the local cloud catalogue have separate records. An application restart changes the running process; it should not erase the user’s intent.
This is the first substantial increase in complexity. The requirement is recovery, and the mechanism is durable local state.
Retries need an identity
An upload may succeed while its response disappears. Retrying with a new identity would turn an uncertain outcome into a duplicate.
The application persists both a media identifier and a request identifier. The backend reserves metadata transactionally in DynamoDB and returns the previous result when the same valid request is replayed.
A SHA-256 content hash serves another purpose: recognizing identical bytes within an account. It does not replace request identity, and it does not make a multi-step workflow atomic.
The distinction matters. Idempotency handles repeated operations. Deduplication recognizes content. Recovery coordinates the remaining steps.
Large files change the cost of failure
Restarting a small photo upload may be acceptable. Restarting a large video after transferring most of it is wasteful.
The implementation uses multipart uploads for videos and for photos larger than 5 MiB when the client requests that mode. The default part size is 5 MiB. Completed parts can be recovered from S3, allowing the next attempt to send the missing work. The S3 multipart documentation describes that lifecycle.
This introduces sessions, checkpoints and completion logic. It also introduces abandoned parts. A bucket lifecycle rule aborts incomplete uploads after seven days, so a custom cleanup scheduler is unnecessary.
Browsing should use smaller representations
A gallery needs a small visual representation, not the full original for every tile.
Original creation events go through SQS to a Lambda worker. For photos, the worker generates a JPEG preview. For videos, the mobile application supplies extracted images that the worker uses for the preview.
Previews live in a separate private bucket and are cached locally. Processing has its own status because successful storage and successful preview generation are different outcomes.
The resulting architecture
Cognito
│
▼
Mobile app ───────→ HTTP API / Lambda ─────→ DynamoDB
│
├── local state: SQLite + working files
│
└── authorized uploads → S3 originals
│
▼
SQS
│
▼
Preview worker
│ │
▼ ▼
S3 previews DynamoDB
| Constraint | Mechanism |
|---|---|
| Keep AWS secrets out of the app | Managed identity and API authorization |
| Avoid proxying large files | Presigned direct transfers |
| Recover after process termination | Persistent local records |
| Make retries safe | Stable request identity and conditional writes |
| Recognize identical content | Account-scoped content hash reservation |
| Resume large transfers | Multipart uploads and remote part reconciliation |
| Clean abandoned parts | S3 lifecycle rule |
| Browse efficiently | Derived previews and local caching |
Completed originals currently have no configured transition to colder storage. That remains a possible cost decision, rather than a feature I can claim already exists.
I also have no requirement that justifies a permanently running server, Kubernetes or a design for millions of users.
The interesting work is in the boundaries: what survives a crash, which component knows the truth, and how the next attempt makes progress. The next article examines those questions through the actual synchronization model.