Tag

Ai Scraping

All articles tagged with #ai scraping

Reddit Signals End for Old Reddit as API and Bot Controls Tighten
technology21 days ago

Reddit Signals End for Old Reddit as API and Bot Controls Tighten

Reddit is moving to curb AI scraping by planning changes to Old Reddit, migrating bots and moderator workflows to its modern stack, and gradually restricting access—potentially blocking logged-out users. The company is phasing out its public API in favor of the Developer Platform, while expanding a Rules Hub moderation system to replace AutoMod in the future, with no hard timeline but clear intent to push developers and communities toward the new stack.

ShieldFont uses word-swapping fonts to slow AI scrapers
technology26 days ago

ShieldFont uses word-swapping fonts to slow AI scrapers

A new open-source project called ShieldFont aims to deter AI scrapers by using extended GSUB-based substitutions in OpenType fonts, making page text appear normal to human readers while delivering gibberish to automated crawlers. The font maps words through predefined pools (about 250) so about a quarter of words are swapped into semantically related, unreadable forms, complicating data ingestion without completely blocking access. While it can slow or confuse scrapers, it’s not foolproof—OCR, targeted AI that downloads the font, or other scraping methods can bypass it, and there could be SEO and accessibility drawbacks. ShieldFont is available as a demo, a React component, and CSS/CDN-ready options, with the project described as alpha and evolving.

Reddit Sues Data Scrapers in Battle Over AI and Internet Privacy
technology10 months ago

Reddit Sues Data Scrapers in Battle Over AI and Internet Privacy

Reddit's lawsuit against AI and data scraping companies is portrayed as an attack on the open internet, but it primarily targets companies that scrape Google search results to link Reddit content, using a twisted interpretation of copyright law and the DMCA's anti-circumvention clause. The case could threaten the fundamental workings of search engines and the open web, as it claims that bypassing technological measures to access publicly available data constitutes copyright infringement, which critics argue is a dangerous overreach.

Reddit to Block Internet Archive Access
technology1 year ago

Reddit to Block Internet Archive Access

Reddit will restrict the Internet Archive's Wayback Machine from crawling most of its content after discovering AI companies scraping data, citing concerns over privacy and policy violations. The move limits the archive to only indexing Reddit's homepage, aiming to protect user data and enforce platform policies. Reddit has previously restricted access to data for AI training and has ongoing disputes with AI companies over data scraping practices.