STRIPE: A Scan-Tag-Recall Image Pipeline for the Family Archive
Original writing, AI-assisted editing.
The homelab exists to serve one workload: scanning thousands of physical family photos and storing them in a searchable archive. The Scan-Tag-Recall Image Pipeline Engine (STRIPE) is that workload. It takes a scan, describes what’s in the frame, pulls out any text on it, works out who’s in it, and files all of that where a natural-language query can find it later.
This piece walks the pipeline stage by stage and records where each one actually stands today — which, right now, is mostly “designed, not yet built.”
Why a Pipeline
The sheer volume of physical photos to process alone demands automation. While my personal knowledge related to the photos is irreplaceable, I can’t be the human in the loop for all aspects of this process for all photos. The importance of some images will justify my personal attention, like the lone Civil War era photo of my 2nd great-grandfather, but for a majority of the photos a standard processing will be sufficient. A single script might be simpler. However, it would also be impossible to debug, impossible to re-run one stage in isolation, and impossible to spread across the nodes that suit each job. Splitting the work into discrete stages behind a shared metadata queue lets each stage fail, retry, and scale on its own — which is also why the queue was the first thing actually built, before any of the stages that feed it.
Pipeline Overview
graph LR
ingest["Scan & Ingest<br/>flex2<br/>watcher is a stub"]
queue[("Metadata Queue<br/>flex2<br/>archive-db-mcp / Postgres")]
ocr["OCR<br/>inspiron<br/>Tesseract"]
caption["Vision Captioning<br/>& Restoration<br/>g11cd — Ollama"]
faces["Face Detect<br/>& Cluster<br/>inspiron — Coral TPU"]
index["Embed & Index<br/>allinone<br/>Qdrant / Chroma"]
rag["RAG Search<br/>& Chat<br/>allinone — Open WebUI"]
ingest --> queue --> ocr --> caption --> faces --> index --> rag
classDef live fill:#eafaf1,stroke:#1a7a4c,color:#1a1d21
classDef wip fill:#fdf0e1,stroke:#c2410c,color:#1a1d21
classDef planned fill:#f7f8fa,stroke:#9aa3af,color:#5a6270
class queue live
class ingest wip
class ocr,caption,faces,index,rag planned
| Stage | Status | Host | Notes |
|---|---|---|---|
| Scan & Ingest | In progress | Windows Laptop, flex2 | Images are scanned to a shared Windows directory. Flex2 detects scans and queues them for further processing. |
| Metadata Queue | Live | flex2 | archive-db-mcp (Postgres) — deployed and verified. Also where household/person naming gets standardized (e.g. “Betty” vs. “Elizabeth”) before it’s written. |
| OCR | Planned | inspiron | Tesseract |
| Vision Captioning & Restoration | Planned | g11cd | Ollama (RTX 3060), orchestrated by Hermes Agent on inspiron |
| Face Detect & Cluster | Planned | inspiron | Coral USB TPU |
| Embed & Index | Planned | allinone | Qdrant / Chroma |
| RAG Search & Chat | Planned | allinone | Open WebUI |
Stage: Scan & Ingest
Photos are scanned via a ScanSnap iX1600 to a Windows network share, CIFS-mounted by a polling watcher on flex2. The watcher is designed to hash each file for dedup, copy it to the primary archive on g11cd, and write a scans row with status pending for every stage — today it’s a running stub that confirms the service is wired up; that logic isn’t implemented yet.
Stage: Metadata Queue
As I scan the images in batches, I can supply some metadata like: whose photo collection is it from; what time period is it; what’s the location; who are known individuals in the image; etc. One challenge here is standardizing the metadata. For example when tagging someone within an image, if a person’s given name was “Elizabeth”, but they were generally called “Betty”, all of the images with her in them should be tagged consistently. In another case, I have 3 Joshua Hoiles in my family tree – two of which were father and son. How should I handle the metadata so that I know which unique person is being referred to?
Stage: OCR
Some people, like my mom, were very good about writing names and dates on the margins or backs of photos. This information is critical for older photos where I may not know or recognize who is in them. This information cannot be lost.
Stage: Vision Captioning & Restoration
While it may be valuable for me to manually identify who is in a photo, I don’t want to manually describe photos of a “farm house”, a “dog in the back of a pickup”, a “pretty sunset”. I need some automation that can not only provide a general description of the image, but also help classify them so that they can be further analyzed, for example, when people are present.
Additionally there will be many images that will require some level of restoration. Some older photos are little more than the size of a postage stamp. Many color photos from the 60s through the 80s are badly faded and need color restoration. Other photos may have cracks or scratches that need to be repaired.
Stage: Face Detect & Cluster
For those images that have been identified as having one or more people in them, I will want to take the processing further to detect faces and cluster them together with specific individuals.
Stage: Embed & Index
Once a scan has a caption, OCR text, and any detected faces attached, it needs to become something searchable rather than just a row in a database. That’s this stage’s job: embed the image, its caption, and its OCR text into vector space, and write those alongside the structured metadata — dates, locations, people, the household and source box it came from — into a vector database on allinone.
I haven’t settled between Qdrant and Chroma yet. Both run comfortably in a container on modest hardware, so the real decision comes down to how well collection layout and metadata filtering hold up once the archive is thousands of records deep rather than a handful of test scans — that trial still needs to happen. allinone is the natural home for it either way: it’s already the cluster’s spot for network-facing, steady-state services, as opposed to the bursty, thread-heavy work parked on inspiron.
Nothing here is built yet, but the schema is ready for it. index_status is one of the four per-stage columns tracked on every scan in the metadata queue, and it’s written to explicitly require caption, OCR, and face-clustering to all be done first. The service itself is still an empty folder with a one-line README.
Stage: RAG Search & Chat
The last stage is the one that actually matters to the rest of the family: a place to ask a question in plain language and get photos back. Open WebUI, also on allinone, sits on top of the vector index and is meant to answer something like “find photos of Grandma at the lake in the 1970s” by combining semantic search over captions and OCR text with the structured metadata already attached to each scan — date range, location, household.
Access is the easy part, at least on paper: everything runs behind Tailscale, so anyone in the family can reach the search interface from their own device without me exposing a port to the public internet or handing the photos to a third party.
This is the stage I’m least worried about technically — Open WebUI’s retrieval support is well-trodden ground — and most worried about in terms of sequencing. It’s the last domino. Every stage before it has to actually be producing real data before there’s anything here worth chatting with.
Tech Stack
| Component | Role |
|---|---|
| Proxmox | Cluster + standalone virtualization |
| Postgres | Metadata queue (the scans table) |
| Ollama | Local vision-language model — captioning + restoration |
| Coral TPU | Face detection acceleration |
| Qdrant / Chroma | Vector store (under evaluation) |
| Tesseract | OCR |
| Open WebUI | Retrieval front end |
| Hermes Agent | Pipeline orchestration |
| Tailscale | Flat network across nodes; remote access without a public endpoint |
Where It Stands
While the first couple of steps have an initial, minimal implementation, I already have some foundational changes that are needed for this to truly serve as an archive. Primarily, I want to break any dependencies on an actual database for storing the images and the image metadata. I don’t want to get sidetracked keeping the underlying database patched and up to date. While I’m capable of performing those tasks, it’s very unlikely that the next person who becomes the caretaker of this family history will have a IT background. Along those same lines, I don’t want to rely on XML and schemas or JSON that are more difficult for non-technologists to work with. I think Markdown is the form I want to use. It’s easily readable by humans and AI with a simple structure and the flexibility to add additional metadata fields as needed. The archive needs to be contained and described by the directory structure.
As I continue to learn about hosting my own LLMs and building out the necessary hardware, I need to identify the supporting MCP servers and hosted applications that will meet the demands of this effort.