Updated September 28, 2026 | 13 min read

I Turned Any Website Into a Lead List With Claude Skills (Step by Step)

I’ve watched sales teams burn hours copy-pasting names from LinkedIn tabs and conference directories into spreadsheets. One SDR I spoke with tracked it: 11 hours a week, just on manual research. That’s more than a full workday, gone.
In this post, I’ll show you exactly how to turn any webpage into a clean, importable lead list using the lemlist Website Scraper Claude Skill. We’ll cover how scraping works, which tool fits which job, and how to go from raw URL to a running outbound campaign. lemlist’s own MCP benchmark puts the full workflow at under two minutes (Best Sales Intelligence Platforms in 2026 Ranked and Reviewed), so the time savings are real.
If you’re here to start scraping right now: get the Website Scraper Claude Skill on GitHub → and skip straight to the step-by-step section below.

What a website scraper does and why sales teams need one

A website scraper automatically extracts data from a webpage and turns it into a structured file, like a spreadsheet. It reads the page’s underlying code, pulls out specific fields you define (names, titles, emails, company info), and exports them into rows and columns you can actually use.
For sales teams, that’s the whole appeal. You’ve probably spent hours copying names and titles from LinkedIn tabs into a spreadsheet by hand. A scraper compresses that into one automated step, usually landing in a CSV you can import straight into your CRM or outreach tool.
Here’s what that looks like in practice:
  • The scraper identifies and pulls the fields you define, like job titles or email addresses
  • Raw page content becomes something you can sort, filter, and search
  • Once set up, the same scrape runs across hundreds of pages without you touching each one
  • It normalizes messy HTML into consistent columns, so you spend less time cleaning data afterward
And sure, a scraper won’t solve bad targeting. Garbage URLs in, garbage data out. But when you’re pointed at the right sources, the speed difference is hard to overstate. lemlist’s own database covers 600M+ contacts (The only platform you need for outbound | lemlist), and the platform holds a 4.6/5 rating on G2 (Lemlist Review 2026: My Honest Opinion After Using the Platform), so when we talk about scaling prospecting, we’ve done the homework.

How website scraping works

Scraping breaks down into two steps: fetching the page, then pulling out what you need. Understanding both helps you troubleshoot when a scrape doesn’t return what you expected.

Extraction and parsing

First, the scraper sends a request to the target page, similar to what happens when you type a URL into your browser. The server responds with the page’s HTML, the raw code containing all the visible text, images, and structure.
The scraper then parses that HTML into a structure it can navigate, sometimes called a DOM tree. Parsing just means mapping out the page (headings, tables, links) so the tool knows exactly where to look for the data you asked for. Skip this step, and you’re left with a wall of code and no way to locate anything useful.
That said, a huge number of modern sites render content with JavaScript after the initial HTML loads. If the scraper only reads the static HTML response, it’ll miss anything built dynamically by React, Angular, or similar frameworks. On top of that, many sites deploy anti-bot detection (CAPTCHAs, fingerprinting, IP rate limits) that will block repeated automated requests. Basic scrapers break on both of these walls, which is why the tool category has shifted toward solutions that handle rendering and detection out of the box.

Output formats and data structuring

Once the scraper has pulled the data, it cleans it up: stripping HTML tags, fixing text formatting, and organizing everything into a usable file. Which format makes sense depends on what you’re doing with the data next.
Format
Best for
Sales use case
CSV
Spreadsheet review, CRM imports
Upload directly to HubSpot, Salesforce, or lemlist
JSON
API integrations, developer workflows
Pass data between tools programmatically
Markdown
Documentation, AI context
Feed content into Claude or another LLM for analysis
For most sales workflows, CSV is the format you’ll reach for. It’s what your CRM expects, and it’s the default output for the lemlist Website Scraper Claude Skill (more on that shortly).

Types of website scraper tools

Not every scraper fits every job. The right pick depends on how technical you are, how much data you need, and how often you’ll run the task.

Browser extensions and no-code scrapers

Tools like Web Scraper (Chrome), Octoparse, or Browse AI run right in your browser. You click on the elements you want, and the extension builds the scraper for you.
They’re great for quick, small jobs. But once you’re past a few hundred records, or the site uses heavy JavaScript, you’ll likely hit walls that require a paid upgrade or manual workarounds.

Scraping APIs and developer frameworks

Frameworks like Python’s BeautifulSoup and Scrapy, or JavaScript’s Puppeteer, give developers full control over how a page gets fetched and parsed. They handle things like proxy rotation and JavaScript-heavy pages well.
The catch: someone has to write and maintain the code. Without engineering support, this route often creates more work than it saves.

Managed scraping APIs

A newer category sits between raw frameworks and no-code tools. Services like Firecrawl and ScrapingBee offer managed infrastructure that handles JavaScript rendering, proxy rotation, and anti-bot bypass through a single API call. You send a URL, get structured data back. No browser to manage, no proxies to configure.
For teams that need reliable extraction at scale but don’t want to maintain a scraping codebase, this tier fills the gap. The tradeoff is cost: managed APIs charge per request, and that adds up at high volume.

AI-powered scrapers and Claude Skills

This category skips both the clicking and the coding. Instead, you describe what you want in plain English, and an AI agent handles the rest.
A Claude Skill is a plain-text instruction file, called SKILL.md, that tells Claude how to carry out a specific task, in this case, scraping. The lemlist Website Scraper Claude Skill lets you type something like “Scrape all contacts from this URL into a CSV with name, title, company, and email.” Claude reads the page, identifies the fields, and builds the file.
For teams without a developer on standby, this is the most direct path from a URL to a usable lead list.

How to pick the right scraper for your workflow

If you need a one-off, small list, a browser extension gets it done, free and fast. For recurring, large-scale extraction, a scraping API, managed service, or framework makes sense, assuming you have dev support or budget. If you’re prospecting without code, an AI-powered scraper or Claude Skill removes the technical barrier entirely. And for full outbound automation, a Claude Skill paired with lemlist handles scraping, enrichment, and campaign launch in one flow.

How to scrape a website and export data to CSV

Here’s the workflow using the lemlist Website Scraper Claude Skill, step by step.
1. Choose your target URL and define your fields. Pick your source, a directory, a team page, an industry listing, and decide up front which columns you want (name, title, email, company, location). Skipping this leads to messy output with irrelevant data mixed in.
2. Install the skill. Head to the GitHub repo, clone it, and point Claude Desktop at the SKILL.md file through your MCP config. Setup usually takes under five minutes.
3. Run the scrape and check your CSV. Type a prompt specifying the URL and the columns you defined. Claude reads the page, pulls the data, and formats a CSV, one row per record, so spot-check a few entries against the original page before you trust it fully.
4. Handle pagination. Many directories spread results across dozens of pages. The skill follows “next page” links automatically, and for larger sites, sessions are resumable, so a timeout doesn’t mean starting over.

Using a website scraper for sales lead generation

Scraping only earns its place in your workflow when it’s tied to something specific you’re trying to find. A few use cases worth trying:
  • Conference speaker lists, association pages, and award announcements often contain prospects you can filter by role or geography
  • About pages and leadership bios give you names, titles, and context on what the company does, which is what makes your first message read like research instead of a template
  • A company hiring for “Head of Revenue Operations” is signaling it’s building out a team your product might support
  • New customer logos or pricing updates on a prospect’s site often mean they’re more likely to respond right now
That last point is worth dwelling on. Timing matters as much as targeting, and manually checking competitor sites for changes isn’t sustainable at scale. lemlist’s intent signal agents can automate that monitoring, surfacing these moments and adding leads to campaigns with messaging tied to the trigger.

How to clean and enrich scraped data before outreach

A scraped CSV is a starting point, not something you send from directly. Raw scrapes typically carry duplicates, inconsistent formatting, and contact info nobody’s verified yet.
You’ll want to remove repeated records from overlapping pages first. Then normalize job titles (“VP Sales” vs. “Vice President of Sales”) and phone formats so your segmentation actually works. After that, check email addresses at the SMTP level, since syntax checks alone miss the real problems. Finally, link contacts to the right CRM record so nothing lands in the wrong account.
lemlist’s data enrichment agents handle this automatically, pulling verified emails and LinkedIn context from multiple providers and matching contacts to company records.
Tip: if you import scraped leads into lemlist, the platform flags unverified emails before you send, protecting your sender reputation. You can enrich missing fields right inside the campaign builder.
The legality question comes up almost immediately, and the honest answer has some nuance to it.

What the law says about scraping public data

The Ninth Circuit’s ruling in hiQ Labs v. LinkedIn held that the Computer Fraud and Abuse Act’s “without authorization” clause doesn’t apply to publicly available data, since there’s no access permission to violate in the first place (California Lawyers Association). That said, creating an account and agreeing to a site’s terms of service is a different situation entirely. Scraping in violation of those terms is a contract issue, separate from the CFAA question (LegalClarity).
The full picture matters here. The dispute between hiQ Labs and LinkedIn actually ended in December 2022, when the parties reached a private settlement. hiQ agreed to a permanent injunction requiring it to cease web scraping, delete all source code and data, and pay $500,000 in damages to LinkedIn (hiQ v. LinkedIn Wrapped Up: Web Scraping Lessons Learned) (ZwillGen). The various decisions leading up to this point show that, under certain circumstances, data scraping publicly available websites may be legal under the CFAA but can still create liability risk under a breach of contract claim or common law torts claims (LinkedIn v. hiQ: Landmark Data Scraping Suit Provides Guidance to Data Scrapers and Web Operators – Tech & Sourcing @ Morgan Lewis) (Morgan Lewis). In other words: the CFAA question went one way, but the contract question went the other.

Robots.txt and rate limits

Robots.txt is a file that tells automated tools which pages a site allows and disallows. Checking it before scraping is a baseline practice, and throttling your request rate matters too, since hammering a server with requests can degrade the site for actual visitors.

GDPR considerations

Scraping personal data belonging to EU residents brings GDPR into play. For B2B prospecting, the legitimate interest basis under Article 6(1)(f), supported by Recital 47’s explicit mention of direct marketing, provides a legal foundation for B2B cold outreach in the EU, so long as you can demonstrate a genuine business reason and respect the right to object (Is Cold Email Legal Under GDPR? Yes, for B2B — Here’s How | Is Cold Email Legal?) (Is Cold Email Legal). Legitimate interest is not a blank check, though. Organizations must conduct a Legitimate Interest Assessment documenting their reasoning (GDPR Compliance for B2B Sales: Guide | Cleanlist).
lemlist is GDPR compliant on the outreach side, but the data collection step is worth a conversation with your legal team if you’re targeting the EU.

Turn scraped prospect data into outbound campaigns

A clean CSV only matters once it turns into real conversations. That’s the last mile: getting leads from spreadsheet to sequence.
With lemlist, you upload your CSV and launch multichannel sequences across email, LinkedIn, calls, SMS, and WhatsApp from one workflow. It syncs with HubSpot and Salesforce so records stay current, and a unified inbox lets you manage replies no matter which channel or mailbox they came through. lemwarm, included with every seat, keeps your deliverability steady even when you’re sending at scale.
The full path looks like this: scrape with the Claude Skill, enrich in lemlist, launch the campaign, manage replies in one inbox. What used to eat 45 minutes of tab-hopping now takes under two minutes (Best Sales Intelligence Platforms in 2026 Ranked and Reviewed).

FAQs about website scrapers

Is web scraping detectable by target websites? Yes. Servers can spot patterns like unusually fast page loads or repeated requests from a single IP. Throttling requests and rotating IPs reduces visibility, though detection methods keep evolving.
Can ChatGPT or Claude scrape a website in real time? Not on their own. Claude can run real-time scraping when connected to a Claude Skill (like the lemlist Website Scraper) or an MCP server that provides web fetching tools.
What’s the difference between web scraping and web crawling? Crawling discovers and indexes pages across a site, which is what search engines do. Scraping pulls specific data from those pages into a structured file like CSV or JSON.
How many pages can a website scraper handle in one session? It varies by tool. Browser extensions typically manage tens of pages, scraping APIs handle thousands, and AI skills like the lemlist Website Scraper support resumable sessions that work through large paginated sites over multiple runs.
What’s the difference between a Claude Skill and lemlist MCP? A Claude Skill is a standalone instruction file you drop into Claude Desktop for a single task, like scraping a website. No API keys, no configuration beyond pointing Claude at the SKILL.md file. lemlist MCP goes further: it connects Claude to your full lemlist account, so the AI can search your lead database, launch campaigns, manage sequences, and pull analytics, all through natural-language prompts. Use the Claude Skill when you want quick, no-setup scraping. Use lemlist MCP when you want Claude to run your entire outbound workflow.

Over to you

Here’s the short version: pick a URL, install the Claude Skill, run the scrape, clean the data in lemlist, and launch. The whole loop, from a raw webpage to a live multichannel campaign, fits inside one workflow now.
I’d start with a single directory page you already know has good prospects. Run the skill, spot-check the CSV, and import it into lemlist to see the enrichment and verification layer in action. You’ll know within one test run whether this replaces your current process.
Start a 14-day free trial, no credit card required, cancel anytime.
Share this post