Install it
Read the ground rules before you point this at anything
Scrapling walks past blocks that exist for a reason. There's a section further down called Before you scrape anything — robots.txt, rate limits, terms of service, personal data. It is the most important part of this page and it takes two minutes. Don't skip it because the install worked.
Three commands. You need Python already on your machine — if python3 --version returns a number in your terminal, you're set.
Install Scrapling with its browsers
The first line gets the library plus the fetchers. The second downloads the browser engines it drives — that's the part that gets past Cloudflare, and it takes a few minutes.
Add the AI extras
This is what gives you the MCP server, so Claude Code or Codex can drive the scraping itself instead of you writing Python.
Hook it up to Claude Code
One line. Then check it landed with claude mcp list, or type /mcp inside a session — you want scrapling · connected.
If that errors with "command not found"
Your shell can't see where pip put it. Run which scrapling-mcp, copy the full path it prints, and use that instead: claude mcp add scrapling -- /full/path/to/scrapling-mcp.
Codex instead of Claude Code
Same server. Run codex mcp add and point it at the scrapling-mcp command, or add it to ~/.codex/config.toml yourself. Confirm with /mcp in a session.
Undo it
claude mcp remove scrapling unhooks it. pip uninstall scrapling removes the library. The downloaded browsers live in your browser cache folder and can stay — other tools use them too.
The prompt you hand Claude Code
You don't write Python. You describe the job and the agent drives Scrapling through the MCP tools. This is the one I'd start with — fill in the two brackets and send.
Why the numbered rules are in there
Rule 1 makes it check permission before it acts. Rule 2 stops it reaching for the stealth browser on a site that would have answered a normal request — which is both politer and about ten times faster. Rule 3 keeps personal data out of your spreadsheet by default. Rule 5 is the one that stops a confident agent quietly making things up.
Once one page works, this is how you scale it without being a menace:
A one-off, no-code version
For a single page you don't even need the agent. Install the shell extra with pip install "scrapling[shell]", then run the command below — it turns any page into clean markdown you can read or feed to an AI. End the filename in .txt for plain text or .html for the raw HTML instead.
If that page is behind Cloudflare
Swap get for stealthy-fetch and add --solve-cloudflare. Same command shape, slower, and it opens a real browser to do it.
Before you scrape anything
Scrapling's own README says it's "provided for educational and research purposes only" and tells you to respect terms of service and robots.txt. That's the author covering himself — but it's also genuinely the line. This tool removes the technical barrier. It does not remove the legal or ethical one, and the two were never the same thing.
Here's the honest version, for someone running a business rather than a research lab.
Generally fine
- Public pages with no login, at a human-ish pace
- Facts: prices, product names, stock status, published dates, public job titles
- Your own sites, listings and profiles
- Sites whose robots.txt allows the path you're pulling
- Anything you'd be comfortable telling the site owner you did
Don't
- Anything behind a login — that's a contract you agreed to, and breaking it is a different category of problem
- Personal data: names, emails, phone numbers, addresses. Scraped email lists are a fast route to a privacy complaint
- Republishing someone's content as your own. Facts aren't copyright; their words and photos are
- Hammering a small business's site. Hundreds of requests a minute is an attack, whatever you meant by it
- Ignoring a 429 or a captcha and forcing your way through anyway
The four checks, every single time
1. robots.txt. Type the site's address then /robots.txt in your browser. It's a plain text file saying which paths the owner is happy for bots to touch. It isn't law, but ignoring it after reading it is the difference between an accident and a decision.
2. Terms of service. Search their terms for "scrape", "crawl", "automated" or "bot". If it's forbidden and you're logged in, you're breaching a contract you signed. Logged out, it's murkier — but you now know their position.
3. Rate limit yourself. A few seconds between requests. If you get a 429, a 403 or a captcha, that's the site saying stop. Stop. The whole reason blocks exist is that someone's server is being hurt.
4. Personal data. In Australia the Privacy Act and the Australian Privacy Principles cover personal information regardless of how you got it — "it was on a public website" is not a defence. EU residents bring GDPR with them. Collect facts about businesses, not details about people, and you sidestep nearly all of it.
The captcha question, answered straight
Scrapling can solve Cloudflare challenges. A captcha is the site's owner saying "I don't want automated traffic here." Walking past it isn't a hack — it's you deciding your convenience beats their stated wish. Sometimes that's genuinely defensible: your own site, a client's site with permission, a public price you're allowed to see anyway. Often it isn't. Ask yourself whether you'd send the site owner a screenshot of what you're doing. If the answer's no, that's your answer.
Three things that make this a non-issue
Check for an official API first — most big sites have one, and it's faster, allowed, and doesn't break when they change their layout. Scrape facts about companies rather than details about people. And keep a note of what you pulled, from where, on what date, so if anyone ever asks you have an answer ready.
Not legal advice
I'm not a lawyer and this isn't advice. If you're building scraping into a product, or pulling anything at scale, or touching personal data at all, that's a conversation with an actual solicitor — not a page on my website.
Four things worth using it for
All four stay on the right side of the section above — public facts about businesses, not details about people.
Competitor pricing, tracked over time
Pull the price and what's included from five competitors' pages, once a week, into one sheet. Three months in you can see who's discounting, who's quietly raised, and whether you're the cheapest without meaning to be. Prices on a public page are facts — this is the cleanest use there is.
Reviews, read at volume
Pull the review text (not the reviewers' names) for a product or category, hand the file to Claude, and ask what people complain about most. It's the fastest market research going — you're reading your customers' actual words instead of guessing at their objections.
Watching your own listings
Your Google Business listing, your product pages, your directory entries. Check weekly that hours, prices and links are still right and nothing's gone stale or broken. Your own data, no ethics question at all, and it catches the thing you'd otherwise find out about from a customer.
Turning a site into something AI can read
Scrapling converts any page to clean markdown with the prompt-injection junk stripped out. Point it at a documentation site or a long report and you get a file you can hand to Claude and ask questions of. This is the use that has nothing to do with bypassing anything.
The pattern under all four
Scrapling gets the data. Claude does the thinking. Neither is much use alone — a spreadsheet of scraped prices sits there doing nothing, and an agent with no data just makes plausible-sounding numbers up.
What "700 times faster" actually measures
I went and read the benchmark
The repo doesn't publish a 700× number. What it publishes is a parsing benchmark: pulling text out of 5,000 nested HTML elements that are already sitting in memory. Scrapling does it in 1.99ms. BeautifulSoup with lxml takes 1,562ms. That's the ~785× that gets rounded down to "700 times faster" on the internet.
Which is a real result, and also not the result most people think they're hearing. Here's the same table with the comparison that matters:
The published numbers
- Scrapling — 1.99ms
- Parsel / Scrapy — 2.06ms (1.03×)
- Raw lxml — 2.56ms (1.29×)
- PyQuery — 23.98ms (12×)
- BeautifulSoup + lxml — 1,562ms (785×)
- BeautifulSoup + html5lib — 3,413ms (1,715×)
What that means for you
- Against the nearest serious competitor it's 3% faster, not 700×
- The giant multiplier is against BeautifulSoup, the slow beginner-friendly one
- It measures parsing, not fetching — the network is what actually makes scraping slow
- In a real job the wait is the website answering, and no library changes that
- Parsing 5,000 elements is a stress test, not a normal page
So is it still worth installing?
Yes — just not for the number. The reasons that hold up: it gets past Cloudflare where a plain request gets a wall, its selectors survive a site redesign instead of silently returning nothing, and it ships an MCP server so your agent can use it without you writing Python. Speed is the headline. Those three are the reason.
While we're checking numbers
The star count is 78,900 as of 07/09/2026, not 77,000 — it's gone up, so if you saw the lower figure it was just out of date. Licence is BSD-3-Clause, so commercial use is fine.
Ready to go deeper?
Where to go from here.
Wright Mode Membership
Join a community of women entrepreneurs implementing AI and automation in their businesses. Live calls, templates, and ongoing support.
Claude Masterclass
Learn how to use Claude like a pro. From prompting fundamentals to building real workflows that save hours every week.
Claude Code Masterclass
Go beyond the chat interface. Build automations, process data, and create tools with Claude Code — no developer experience needed.