Ali FarooqBook a call

Case study · Document AI · 2026

Screenshot-to-CRM pipeline

A vision-LLM pipeline that reads casting screenshots, drafts structured records, and publishes them to a production CRM once a person signs off.

  1. Screenshot or page capture
  2. Classify
  3. Extract
  4. Verify
  5. Human review
  6. Production CRM
Five queued stages per document. Nothing reaches the CRM until a person approves the draft.
Client
Casting-industry CRM
Engagement
Independent consulting
Timeline
Jul – Aug 2026
My role
Sole engineer
Stack
TypeScript · Next.js · BullMQ · Postgres
Outcome
Reviewed records in production
On this page
  1. 01Problem
  2. 02Constraints
  3. 03What I owned
  4. 04Approach
  5. 05Results
  6. 06Skills this shows
  7. 07For your project

Problem

Casting submissions and talent reports arrive as screenshots. Getting them into the client’s CRM as structured projects, roles and talent records was manual work. The two document types carry the same information in inverted layouts: one project with many people, or one person across many projects.

Constraints

  • A wrong contact or a wrong merge is far worse than a paraphrased title, and hard to undo once it reaches production.
  • Screenshots are noisy, rotated and full of small text. Generic vision-to-JSON invents fields.
  • The CRM’s pages are built from frames, so a normal browser screenshot captures a blank shell.
  • Model confidence isn’t a calibrated probability.

What I owned

I was the only engineer, from the first commit to production.

  • Vision pipeline. Classify, extract, verify and enrich as queued jobs, over a closed Zod schema of entities and relationships.
  • Review interface. Two review forms for the two document types, with versioned drafts and conflict detection.
  • CRM write-back. Server-side publishing of projects, roles and roster links. The browser never holds the CRM token.
  • Evaluation. A multi-model harness (Claude, Gemini, GPT) scoring classification and extraction against golden screenshots, with per-stage token, cost and latency numbers.
  • Capture extension. A Chrome MV3 side panel that captures full framed pages over the DevTools protocol and sends them for processing.

Approach

The model reads the image and emits typed records in one call, validated against a closed vocabulary. There is no separate OCR step. Classification is a hard gate: a low-confidence or unknown document stops before extraction or any database write.

Trust thresholds follow the cost of an error. Contact details and merges need more confidence than titles, and anything uncertain goes to review instead of the CRM. A cheaper model classifies and a stronger one extracts, chosen after measuring that extraction dominated cost and latency.

Fuzzy cross-document merging was built, then taken off the live path once it proved to be the costliest kind of mistake.

Results

  • Structured records published to the client’s production CRM, each approved by a person first.
  • Models chosen from a benchmark against golden screenshots, on accuracy, cost and latency, not reputation.
  • Full-page capture of the client’s framed pages, which standard browser capture couldn’t do.

Skills this shows

If your role needs Evidence in this project
Document AI and extraction Schema-constrained vision extraction with a hard classification gate.
LLM evaluation A multi-model harness with golden screenshots and per-stage cost and latency.
Human-in-the-loop design A review UI with versioned drafts; nothing reaches production unreviewed.
Integrating with existing systems Server-side CRM write-back with typed errors for partial failures.
Browser tooling An MV3 extension capturing framed pages over CDP.

What this means for your project

  • Have documents or screenshots people retype? I’ll get them into your system as structured, reviewed records, and measure accuracy and cost before launch.
  • Worried about AI writing bad data? I design where automation stops and a person approves, based on what each kind of mistake costs.
  • Need one engineer to own it end to end? This was a solo build, from browser capture to production write-back.

© 2026 Ali Farooq