Writing

STRIPE: A Scan-Tag-Recall Image Pipeline for the Family Archive

Original writing, AI-assisted editing.

The homelab exists to serve one workload: scanning thousands of physical family photos and storing them in a searchable archive. The Scan-Tag-Recall Image Pipeline Engine (STRIPE) is that workload. It takes a scan, describes what’s in the frame, pulls out any text on it, works out who’s in it, and files all of that where a natural-language query can find it later.

This piece walks the pipeline stage by stage and records where each one actually stands today — which, right now, is mostly “designed, not yet built.”

Why a Pipeline

The sheer volume of physical photos to process alone demands automation. While my personal knowledge related to the photos is irreplaceable, I can’t be the human in the loop for all aspects of this process for all photos. The importance of some images will justify my personal attention, like the lone Civil War era photo of my 2nd great-grandfather, but for a majority of the photos a standard processing will be sufficient. A single script might be simpler. However, it would also be impossible to debug, impossible to re-run one stage in isolation, and impossible to spread across the nodes that suit each job. Splitting the work into discrete stages behind a shared metadata queue lets each stage fail, retry, and scale on its own — which is also why the queue was the first thing actually built, before any of the stages that feed it.

Pipeline Overview

graph LR
ingest["Scan & Ingest<br/>flex2<br/>watcher is a stub"]
queue[("Metadata Queue<br/>flex2<br/>archive-db-mcp / Postgres")]
ocr["OCR<br/>inspiron<br/>Tesseract"]
caption["Vision Captioning<br/>& Restoration<br/>g11cd — Ollama"]
faces["Face Detect<br/>& Cluster<br/>inspiron — Coral TPU"]
index["Embed & Index<br/>allinone<br/>Qdrant / Chroma"]
rag["RAG Search<br/>& Chat<br/>allinone — Open WebUI"]

ingest --> queue --> ocr --> caption --> faces --> index --> rag

classDef live fill:#eafaf1,stroke:#1a7a4c,color:#1a1d21
classDef wip fill:#fdf0e1,stroke:#c2410c,color:#1a1d21
classDef planned fill:#f7f8fa,stroke:#9aa3af,color:#5a6270

class queue live
class ingest wip
class ocr,caption,faces,index,rag planned
Stage Status Host Notes
Scan & Ingest In progress Windows Laptop, flex2 Images are scanned to a shared Windows directory. Flex2 detects scans and queues them for further processing.
Metadata Queue Live flex2 archive-db-mcp (Postgres) — deployed and verified. Also where household/person naming gets standardized (e.g. “Betty” vs. “Elizabeth”) before it’s written.
OCR Planned inspiron Tesseract
Vision Captioning & Restoration Planned g11cd Ollama (RTX 3060), orchestrated by Hermes Agent on inspiron
Face Detect & Cluster Planned inspiron Coral USB TPU
Embed & Index Planned allinone Qdrant / Chroma
RAG Search & Chat Planned allinone Open WebUI

Stage: Scan & Ingest

Photos are scanned via a ScanSnap iX1600 to a Windows network share, CIFS-mounted by a polling watcher on flex2. The watcher is designed to hash each file for dedup, copy it to the primary archive on g11cd, and write a scans row with status pending for every stage — today it’s a running stub that confirms the service is wired up; that logic isn’t implemented yet.

Stage: Metadata Queue

As I scan the images in batches, I can supply some metadata like: whose photo collection is it from; what time period is it; what’s the location; who are known individuals in the image; etc. One challenge here is standardizing the metadata. For example when tagging someone within an image, if a person’s given name was “Elizabeth”, but they were generally called “Betty”, all of the images with her in them should be tagged consistently. In another case, I have 3 Joshua Hoiles in my family tree – two of which were father and son. How should I handle the metadata so that I know which unique person is being referred to?

Stage: OCR

Some people, like my mom, were very good about writing names and dates on the margins or backs of photos. This information is critical for older photos where I may not know or recognize who is in them. This information cannot be lost.

Stage: Vision Captioning & Restoration

While it may be valuable for me to manually identify who is in a photo, I don’t want to manually describe photos of a “farm house”, a “dog in the back of a pickup”, a “pretty sunset”. I need some automation that can not only provide a general description of the image, but also help classify them so that they can be further analyzed, for example, when people are present.

Additionally there will be many images that will require some level of restoration. Some older photos are little more than the size of a postage stamp. Many color photos from the 60s through the 80s are badly faded and need color restoration. Other photos may have cracks or scratches that need to be repaired.

Stage: Face Detect & Cluster

For those images that have been identified as having one or more people in them, I will want to take the processing further to detect faces and cluster them together with specific individuals.

Stage: Embed & Index

Once a scan has a caption, OCR text, and any detected faces attached, it needs to become something searchable rather than just a row in a database. That’s this stage’s job: embed the image, its caption, and its OCR text into vector space, and write those alongside the structured metadata — dates, locations, people, the household and source box it came from — into a vector database on allinone.

I haven’t settled between Qdrant and Chroma yet. Both run comfortably in a container on modest hardware, so the real decision comes down to how well collection layout and metadata filtering hold up once the archive is thousands of records deep rather than a handful of test scans — that trial still needs to happen. allinone is the natural home for it either way: it’s already the cluster’s spot for network-facing, steady-state services, as opposed to the bursty, thread-heavy work parked on inspiron.

Nothing here is built yet, but the schema is ready for it. index_status is one of the four per-stage columns tracked on every scan in the metadata queue, and it’s written to explicitly require caption, OCR, and face-clustering to all be done first. The service itself is still an empty folder with a one-line README.

Stage: RAG Search & Chat

The last stage is the one that actually matters to the rest of the family: a place to ask a question in plain language and get photos back. Open WebUI, also on allinone, sits on top of the vector index and is meant to answer something like “find photos of Grandma at the lake in the 1970s” by combining semantic search over captions and OCR text with the structured metadata already attached to each scan — date range, location, household.

Access is the easy part, at least on paper: everything runs behind Tailscale, so anyone in the family can reach the search interface from their own device without me exposing a port to the public internet or handing the photos to a third party.

This is the stage I’m least worried about technically — Open WebUI’s retrieval support is well-trodden ground — and most worried about in terms of sequencing. It’s the last domino. Every stage before it has to actually be producing real data before there’s anything here worth chatting with.

Tech Stack

Component Role
Proxmox Cluster + standalone virtualization
Postgres Metadata queue (the scans table)
Ollama Local vision-language model — captioning + restoration
Coral TPU Face detection acceleration
Qdrant / Chroma Vector store (under evaluation)
Tesseract OCR
Open WebUI Retrieval front end
Hermes Agent Pipeline orchestration
Tailscale Flat network across nodes; remote access without a public endpoint

Where It Stands

While the first couple of steps have an initial, minimal implementation, I already have some foundational changes that are needed for this to truly serve as an archive. Primarily, I want to break any dependencies on an actual database for storing the images and the image metadata. I don’t want to get sidetracked keeping the underlying database patched and up to date. While I’m capable of performing those tasks, it’s very unlikely that the next person who becomes the caretaker of this family history will have a IT background. Along those same lines, I don’t want to rely on XML and schemas or JSON that are more difficult for non-technologists to work with. I think Markdown is the form I want to use. It’s easily readable by humans and AI with a simple structure and the flexibility to add additional metadata fields as needed. The archive needs to be contained and described by the directory structure.

As I continue to learn about hosting my own LLMs and building out the necessary hardware, I need to identify the supporting MCP servers and hosted applications that will meet the demands of this effort.