All case studies
Agentic AIPersonal Tooling

AI Job Search — Pipeline & Status Automation

Extended an open-source job-search framework with Gmail recommendation-lead detection, repost-dedup hardening, a live application dashboard, and a shared posting-fetch cache — automation layered on a forked base, not authored from scratch.

The takeaway: Confirmed clean of PII or client data via a full git-history audit — and a real repost bug (one already-rejected role scraped three times under three different LinkedIn IDs, nearly reapplied to) closed for good once caught (commit 992a944).

Role

Sole developer & end user — directed Claude Code for all implementation

Timeline

Aug 2026 · ongoing

Stack

JavaScript (Bun) · Claude Code · Gmail API

Status

Active — used daily for my own job search

Run in Fast-Iterate mode — lighter process ceremony by design for solo personal tooling, not a lower bar. See the governance comparison against OpenCAM →

Problem

What was missing

The framework's default output was a static snapshot and application outcomes lived entirely in employer emails I had to notice, reread, and hand-transcribe — with dozens of applications in flight, status updates silently lagged reality.

Portal-level dedup only matched on exact URL or job ID, which missed an employer relisting an unfilled role under a brand-new ID. That wasn't hypothetical: a role I'd already been rejected from was scraped three times under three different LinkedIn IDs, and the surviving copy nearly went out again in a fresh application batch before I caught it (commit 992a944).

Scope

What's mine, and what came with the fork

Inherited from upstream

MadsLorentzen/ai-job-search

  • Core /apply workflow: fit evaluation → CV/cover-letter drafting → compile → verify
  • Job-scraping CLI framework and the portal-skill pattern
  • Gmail sync's core mechanism: propose-then-approve tracker status updates
  • Base /rank and /outcome command structure

Built on top

57+ fork-specific commits

  • Extended Gmail sync to also detect and register job-recommendation digest emails as new scrape leads — the base sync only covered existing applications
  • Live, filterable application dashboard — iterated from an early CSV export through a static version to today's live one; job_scraper/ didn't exist in upstream at all
  • Cross-portal repost-dedup hardening — two real incidents, two fixes (PR #1, then commit 992a944)
  • Shared posting-fetch cache between /rank and /apply
  • Security guards enforcing gitignore/permissions/hooks discipline on a public fork

Process

How I worked it

1

Used my own job search as the test bed

Forked the framework in August 2026 to run my active search across the UK, Germany, and Ireland, and used the fork itself as the vehicle to close real gaps I hit using it daily.

2

Ran it direct-to-master, deliberately

57 fork-specific commits with no PR-per-change or issue-tracked backlog — a live scraper hitting real job portals needed a tight edit-run-observe loop, and review overhead has no payoff when I'm the only contributor and the only person affected by a regression.

3

Closed the repost gap after it actually cost me

The three-LinkedIn-ID incident directly drove a same-company-title check against the entire scrape history, not just the current run's pool — catching the exact failure mode that had already happened once.

4

Held data governance to a fixed bar regardless of iteration speed

Every piece of sensitive data — the tracker, scraped-job cache, Gmail sync state, cached research — stayed git-ignored throughout, verified afterward with a full git log --full-history audit rather than just assumed.

Decisions

Calls I made, and why

Why direct-to-master here

Contrast this with OpenCAM, a multi-user tool that has to be defensible to someone else. This is solo personal tooling — direct-to-master rapid feedback loops win when I'm the only stakeholder.

Checks the whole scrape history

ID-based checks only catch a repost when the ID matches. Comparing normalized company + title against every existing entry, regardless of prior scrape date, catches the case ID checks structurally can't (commit 992a944).

Cache once, reuse everywhere

/rank and /apply were each independently re-fetching the same posting. A shared, normalized-filename cache means the common rank-then-apply path costs one fetch instead of two.

See how this compares to the other project's delivery model

Outcomes

What changed

0

PII or client data currently exposed — confirmed via full git-history audit

150+

job postings processed & auto-deduped

57

fork-specific commits shipped, direct-to-master

  • The repost-hardening fix closes a failure mode that had already cost one wasted application before it shipped — zero repeats since.
  • One representative Gmail-sync run resolved 7 outcome updates and surfaced 2 new leads from an 11-thread window in a single pass, replacing unbounded manual inbox rereading.

Supporting visuals

Inside the build

57 fork-specific commits on GitHub, diverged from the upstream MadsLorentzen/ai-job-search base.

The live dashboard, filtered to postings not yet applied to — fit scoring, gates, and dedup status per row.