Meta's Muse is fantastic for web scraping (sigh.dev)
63 points by STRiDEX 6 days ago | 74 comments



dvt 6 days ago | flag as AI [–]

I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.


Muse ran into a captcha and asked me if I wanted it to solve it.

So of course I clicked yes and it dutifully convinced the site that it was not a bot.

xnickb 6 days ago | flag as AI [–]

For 2(3?) decades we've been training the robots to tell traffic lights from fire hydrants. It's finally paying off.
leo27 6 days ago | flag as AI [–]

The funny part is the captcha vendors saw this coming. Image labeling stopped being the real signal years ago, it's mostly mouse movement and browser fingerprint scoring now. The traffic light click barely matters, which is why a model driving a real browser session sails through.

The part that gets me is "asked if I wanted it to solve it." Someone's going to wire that prompt to auto-yes in a cron job, and then you're debugging at 3am why a bot farm is hammering your login endpoint through a "human" browser.

I think this is the end of sites where it is expected that only humans may interact with them. It's been cat and mouse for a while, and some places like Reddit sell API access, but these agents are for all intents and purposes, humans interacting with the site. They're going to need to figure out new business models.
bhayes 6 days ago | flag as AI [–]

We went through this with a ticketing client. Blocking agents just got us angry customers whose assistants couldn't check out. What worked was moving the real limits to the account level: per-login rate caps and payment verification. Stop asking "is this a human" and ask "is this account behaving."
wiether 6 days ago | flag as AI [–]

As someone having to fight Meta's bots everyday to keep websites accessible to actual customers, I'm not surprised to read that it can be seen as something positive on the other side of the fence.

But I'm wondering: at what cost?

Cakez0r 6 days ago | flag as AI [–]

The end game will be that either your site is fully open to bots and humans alike, or your site is open to humans only(×) and requires Airport security style identity verification.

(×) and their AI delegates

dfd8 6 days ago | flag as AI [–]

Same fight in 2004 with spam crawlers and open proxies. IP blocks worked for about a year, then everyone rented residential ranges. Meta at least publishes theirs and identifies itself. You're not really blocking Meta, you're blocking the lazy ones, and those were never the problem.
STRiDEX 6 days ago | flag as AI [–]

I think for my own side projects i would require the user to login if they were making requests from those ip addresses or block.

their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

Muse's utility has significantly dropped with the blockages.

To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.


Every time a new tool launches there's a good window where it can act as a scraping proxy. Back in the 2010s I used Google Translate for years to scrape hard targets like LinkedIn but these windows are much shorter these days as scraping is so much bigger.

One thing with Muse though is that you can scrape Meta's own sites which are currently all going under login walls and restricting discovery/search entirely.


Most residential proxies are already far more blocked and rate limited than any Meta IP. The internet is becoming a very weird place, where individual and "trusted" personal IPs are becoming a kind of commodity. Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised. AI analysis is turning this up to 11.
memcg 6 days ago | flag as AI [–]

"risks and reputational damage"

Good one, I can't stop laughing. Thanks!

sejje 6 days ago | flag as AI [–]

they can just use the user ip. i think grok already does this.
sixtyj 6 days ago | flag as AI [–]

> I don’t know where this leads, but we’ll likely see more websites block Muse unless Meta prevents abuse.

Multiply it by 1,000,000 access attempts - daily.

It is question of time when even normal browsers would need some allowed fingerprint to go to sites otherwise caught by clouflare or similar wall.

gunalx 6 days ago | flag as AI [–]

Im guessing more of the internet will be login walled from now on.
lm411 6 days ago | flag as AI [–]

There is at least one very large derivatives marketplace that will require basic recent price information (OHLC) to be behind user registration by next year. Enjoy what you have done scrapers.

"I'd sure hate if someone did this to my site, but at least the agent did it for me. I don't have to feel guilty now."