Development•Advanced

Create a Lead Database Using Code Assistants

Stop asking AI to know the data. Ask AI to build the tools to get the data.

•
AI AgentsData EngineeringWeb ScrapingSales ProspectingAutomation

Last week, I needed a database of CrossFit gyms for a project I'm building. Not a generic list, I needed names, owner contacts, emails, class schedules, and everything formatted for my specific database.

I tried the obvious approaches: asking an LLM to research it and also considered using no-code tools like n8n. Both fell short. What actually worked was treating my code assistant as a data engineering partner. In this article, I'll walk you through exactly how I did it and share a framework you can apply to your own lead collection challenges.

  • Create lead database with code assistants flow

1. Why "Automated Research" Fails

Here's what most people try first: open ChatGPT, type something like "Give me a list of CrossFit gym owners in the Vaud region of Switzerland with their emails and websites", and expect magic.

The result? A mix of hallucinated business names, outdated information, and generic responses like "I don't have access to real-time data."

This isn't the AI's fault. LLMs are reasoning engines, not databases. They're trained on static snapshots of the internet, not live directories. When you ask them to "research" something, you're really asking them to guess based on patterns they've seen before.

The shift in thinking: Instead of asking the AI to know the data, ask it to build the tools to get the data.

This is the difference between hiring someone who claims to have all the answers versus hiring someone who knows how to find them. The second person is far more useful.


2. The Problem

I'm currently building a tool for CrossFit gyms. Before I can do anything useful, I need something basic: a database of gyms that actually exist in my target region.

I need:

  • A list of gyms (names, locations, websites)
  • Contact information (owner names, emails, phone numbers)
  • Enriched data (class schedules, types of training offered)
  • Everything formatted to match my database schema

Doing this manually? Hours of copy-pasting. Hiring a virtual assistant? Expensive and slow. Using a pre-built lead database? Often outdated, generic, or missing the specific fields I need.

What I actually wanted was a custom pipeline that:

  1. Pulls from authoritative sources (Google Maps, directories, etc.)
  2. Visits each gym's website to extract specific data
  3. Handles edge cases (like information trapped in images)
  4. Outputs everything in my exact format

This is exactly what I built with a code assistant in an afternoon.


3. The Workflow

Here's the high-level flow I built:

Data sources I identified:

  • Primary source: Google Maps, the most complete list of any business in any region
  • Secondary sources: Each gym's own website for owner info, contacts, and schedules
  • Edge cases: Schedule pages that are images, not text

The key insight: I wasn't building one monolithic scraper. I was building a sequence of specialized scripts, each solving one micro-problem. Think of them as micro-agents, each one specialized in a specific task, working together in a pipeline.


4. The Tech Stack

Code Assistants: Your Data Engineering Co-Pilot

Tools like Cursor or Claude Code aren't just chat interfaces, they're agents with context. They can:

  • Read and write files in your project
  • Execute commands in your terminal
  • Iterate on code based on error messages
  • Remember what you've built so far

This is fundamentally different from pasting code snippets from ChatGPT or configuring blocks individually in n8n. The assistant sees your environment and can debug alongside you.

Code vs. No-Code

I initially considered using n8n for this workflow. It's powerful, visual, and doesn't require writing code. But here's the trade-off:

ApproachProsCons
No-code (n8n, Make, Zapier)Visual, no coding requiredMust build every block manually, limited flexibility for edge cases
Pure LLM promptingFast for simple tasksHallucinations, no real data access
Code assistantFlexible, handles edge cases, fast iterationRequires ability to run/test code

The code assistant sits in a sweet spot: you describe what you want in plain language, and it handles the implementation. You don't need to be a senior engineer, you just need to be able to run a script, read an error message, and say "this didn't work, try something else."

A productivity note: I no longer type my instructions. I use voice transcription (I use

Wispr Flow

, but many alternatives exist) to dictate what I want the assistant to do.

With voice, you can be far more detailed and comprehensive. You're not subconsciously saving characters because typing is slow. You think out loud, describe edge cases naturally, give precise context, and iterate much faster. When you're doing back-and-forth with a code assistant, this speed and depth matters enormously.


5. The Methodology: The "Micro-Problem" Framework

Here's the core principle that made this work: never try to solve the whole problem in one prompt.

LLMs get confused when you ask for too much at once. Instead, break your pipeline into discrete, testable steps. Each step should:

  • Have a clear input and output
  • Be independently verifiable
  • Fail in obvious ways if something goes wrong

Step 1: The Discovery Script

Goal: Get the root list of leads from your primary source.

I asked the assistant to write a script that queries the Google Maps API for "CrossFit gyms in Canton of Vaud, Switzerland."

The conversation went something like:

"I need a script that uses the Google Maps Places API to find all CrossFit gyms in a specific region and outputs the results as JSON."

The assistant told me I needed an API key, walked me through setting it up, then generated a script in TypeScript (the coding language of my project) that returned 60 results with: name, address, phone, website URL, and Google Maps rating.

Output: gyms_raw.json with a list of leads with basic info.

Step 2: The Deep Diver

Goal: Visit each lead's website and extract specific fields.

The Google Maps data was a starting point, but I needed more: owner names, contact emails, class types offered, and the URL to their schedule page.

"Now I need a script that takes my JSON file, visits each gym's website, and extracts: owner or coach names, contact emails, types of classes offered, and the URL to their schedule or timetable page."

The assistant wrote a scraper using Playwright, which spins up an actual browser on my machine. This was important because many gym websites load content dynamically with JavaScript, so simply fetching the HTML wouldn't work. Playwright renders the page fully before extracting data, just like a real user would see it.

It wasn't perfect on every site, but it worked on about 80% of them, which was far better than doing it manually.

Output: gyms_enriched.json containing the same list, now with contacts and metadata.

Step 3: The Vision Layer

Goal: Extract data that exists only in images.

Here's where it got interesting. Most CrossFit gyms publish their weekly schedule as an image (a designed graphic, not HTML text). No traditional scraper can read this.

My solution:

"Create a script that takes a URL, captures a screenshot of the page, sends it to OpenAI's Vision API, and extracts the class schedule in a structured format: day, time, class name."

I also gave it my target schema so the output would match my database model directly. The assistant handled the screenshot capture (using Playwright again), the API call, and the response parsing.

Output: schedules.json with structured timetable data extracted from images.

Step 4: The Harmonizer

Goal: Merge everything into your final database format.

Now I had three data sources: raw Google Maps data, enriched website data, and vision-extracted schedules. The final step was combining them.

"Write a script that merges gyms_raw.json, gyms_enriched.json, and schedules.json into a single file matching this schema: [provided my database model]"

The assistant wrote a merge script that handled duplicates, missing fields, and data type conversions.

Output: gyms_final.json ready to import into my database.

  • Create lead database with code assistants

6. When to Use This Approach

This approach shines when you're in the messy middle: the problem is too complex for a simple scraper, but too specific to justify building production infrastructure.

Use this when:

  • ✅

    You need custom fields that don't exist in standard databases (e.g., "gyms that offer Olympic lifting classes")

  • ✅

    Your data sources are scattered across multiple public websites, APIs, or directories

  • ✅

    You're dealing with edge cases like data trapped in images, PDFs, or JavaScript-heavy sites

  • ✅

    Speed matters more than perfection (you need results in days, not months)

  • ✅

    You can run basic commands (if you can install Node.js and run a script, you're good)

Skip this if:

  • ❌

    You need daily automation (this is for one-time or occasional data collection, not recurring pipelines)

  • ❌

    The data requires authentication (dealing with logins adds significant complexity)

  • ❌

    You're completely non-technical and have no one to help with basic setup

  • ❌

    An off-the-shelf database already has what you need (don't reinvent the wheel)


7. Other Use Cases

The "Micro-Problem" framework (building micro-agents for each step) applies far beyond gym databases:

  • Sales prospecting: Find companies in a specific niche → enrich with LinkedIn data → extract decision-maker emails
  • Competitive analysis: List competitors from a directory → scrape their pricing pages → extract features from screenshots
  • Real estate: Pull listings from multiple platforms → visit each listing → extract details not in the API
  • Recruiting: Find candidates from GitHub/portfolios → scrape their project pages → summarize their tech stack
  • Content research: Gather articles on a topic → extract key quotes → compile into a research database

The pattern is always the same:

  1. Primary source for the root list
  2. Secondary sources for enrichment
  3. Vision layer for non-text data
  4. Harmonizer for your target format

8. The Result

After running the pipeline, I had a complete database of CrossFit gyms in the Vaud region. But the real value came from turning that data into something useful: a gym explorer application.

Gym Explorer application showing the results of scraped gyms

The application displays all the gyms I collected, complete with:

  • Names, locations, and contact information
  • Owner details and email addresses
  • Class schedules extracted from images
  • Interactive map showing gym locations across the region

This is the power of the micro-agent approach: what started as a vague idea ("I need a list of gyms") became a fully functional database and application in a single afternoon. The same pattern: discovery → enrichment → vision layer → harmonizer, can be applied to any data collection challenge.


9. Try It This Week

Here's your challenge: pick one small data collection task you've been putting off. Maybe it's:

  • Finding 20 potential podcast guests in your niche
  • Building a list of local businesses for outreach
  • Gathering competitor pricing for a market analysis

Open Cursor (or your code assistant of choice), and start with step one:

"I need a script that [gets the root list] from [your primary source] and outputs JSON."

See what happens. Let the assistant guide you through API keys, dependencies, and errors. When the first script works, move to step two.

You'll be surprised how quickly a "custom database" goes from overwhelming to done.


If you need help setting up your environment or guidelines on how to start with your code assistant, don't hesitate to reach out, and we are happy to help.