Is there a simple way to severly impede webscraping and LLM data collection of my website?

Maroon@lemmy.world · edit-2 1 year ago

Is there a simple way to severly impede webscraping and LLM data collection of my website?

corroded@lemmy.world · 1 year ago

Speaking from experience, be careful you don’t become over-zealous in your anti-scraping efforts.

I often buy parts and equipment from a particular online supplier. I also use custom inventory software to catalog my parts. In the past, I could use cURL to pull from their website, and my software would parse the website and save part specifications to my local database.

They have since enacted intense anti-scraping measures, to the point that cURL no longer works. I’ve had to resort to having the software launch Firefox to load the web page, then the software extracts the HTML from Firefox.

I doubt that their goal was to block customers from accessing data for items they purchased, but that’s exactly what they did in my case. I’ve bought thousands of dollars of product from them in the past, but this was enough of a pain in the ass to make me consider switching to a supplier with a decent API or at least a less restrictive website.

Simple rate limiting may have been a better choice.

IphtashuFitz@lemmy.world · 1 year ago

Try using “curl -A” to specify a User-Agent string that matches Chrome or Firefox.

corroded@lemmy.world · 1 year ago

I probably should have specified I’m using libcurl, but I did try the equivalent of what you suggested. I even tried setting a list of user agents and having it cycle through. None of them work. A lot of anti-scraping methods use much more complex schemes than just validating the user agent. In some cases, even a headless browser will be blocked.

bloodfart@lemmy.ml · 1 year ago

Mouser?

Deckweiss@lemmy.world · 1 year ago

I did this a while back for blocking LLMs and there are more methods discussed in that threads comments.

https://lemmy.world/post/14767952

TheAnonymouseJoker@lemmy.ml · edit-2 1 year ago

Removed by mod

RvTV95XBeo@sh.itjust.works · 1 year ago

Are you perhaps an LLM in disguise?

TheAnonymouseJoker@lemmy.ml · edit-2 1 year ago

Removed by mod

otp@sh.itjust.works · 1 year ago

That’s exactly what an LLM would say…

TheAnonymouseJoker@lemmy.ml · edit-2 1 year ago

Removed by mod

habitualTartare@lemmy.world · 1 year ago

https://en.wikipedia.org/wiki/Robots.txt

Should cover any polite web crawlers but it is voluntary.

https://platform.openai.com/docs/gptbot

Might have to put it behind a captcha or other type to severely limit automated access.

It’s not realistic to assume it won’t get scraped eventually. Such as someone paying people to bypass capatcha or web crawlers that don’t respect robots.txt. I also don’t know if Google and Microsoft bundle their AI data collection that doesn’t also remove your site from web search.

fubarx@lemmy.ml · 1 year ago

Scrape a bunch of Onion articles, link them together in an index, then post an invsible link from your home page that spiders will follow but humans can’t see.

Write a script to randomize the words on all the articles and link them in too. Then change the image tags to point to random wikimedia files.

If there’s one thing we’ve learned, it’s that there’s very little quality control. Channel your inner Ken Kesey / Merry Prankster. Have fun.

FactualPerson@lemmy.world · 1 year ago

Why not add a basic http Auth to the site? Most web hosts provide a simple way to protect a site or directory.

You can have a simple username and pass for humans, but it will stop scrapers as they won’t get past the Auth challenge unless they know the details. I’m pretty sure you can even show login details in the Auth dialog, if you wanted to, rather than pre sharing them.

Maroon@lemmy.world · 1 year ago

deleted by creator

quafeinum@lemmy.world · 1 year ago

With htacces everyone can use the same credentials and you can have a message in the popup like ‚use username admin, passeword= what’s a duck? as the login‘ The other option would be an actual captcha

GBU_28@lemm.ee · 1 year ago

Any attempts to mangle the body of the pages or obscure in in JS are moot. Most competent stealthy scrapers have visionai as a fallback, so even if you reduce the ability to programmatically parse the page body, the bot can just snatch an image of the page and OCR the contents.

Flyswat@lemmy.ml · 1 year ago

Use a special/custom font where the letter shown differs from the character used.

utopiah@lemmy.ml · 1 year ago

Possibly, see https://github.com/ai-robots-txt/ai.robots.txt but I just discovered it myself while looking for a Robots.txt a la CrowdSec/AdBlocking lists, so feedback appreciated!

utopiah@lemmy.ml · 1 year ago

Actually https://darkvisitors.com/docs/robots-txt might be more direct.

IphtashuFitz@lemmy.world · 1 year ago

deleted by creator

istanbullu@lemmy.ml · 1 year ago

use javascript 🤣