Security & data research
I probe how sites defend themselves — then extract the data behind those defenses, or disclose the flaws in them.
The same defenses a scraper has to pass are the ones a security test pokes at. One is 227 production scrapers on Apify; the other is vulnerability research under coordinated disclosure. Same craft, both sides of the boundary.
Most-run scrapers
Bot protection, TLS fingerprinting, rate limits — passed reliably. Ranked by total runs.
GMGN Scraper — Trending Memecoins, Smart Money & Wallets
Scrape GMGN.ai with no API key: trending memecoins, new launches, smart-money buys, top-trader wallets, token rugcheck, holders, wallet PnL and KOL signals across 6 chains.
Website Contact Scraper – Email, Phone & Social Extractor
Scrape website contact details: emails, phones and social profiles (LinkedIn, Instagram, X/Twitter, Facebook, YouTube). Auto-crawls Contact/About pages. Export JSON for B2B leads and enrichment.
Welcome to the Jungle Jobs Scraper – WTTJ Jobs & Salary Data
Scrape WTTJ job listings by keyword, contract type, remote and salary. Extract title, company, salary, offices and profession. Algolia API — no proxy or login needed.
Pinterest Scraper - Scrape Pins, Boards & Profiles
Scrape Pinterest pins, boards and profiles by keyword, board, user or pin URL: images, videos, links, save counts. No login or API key. Export CSV, JSON.
PropertyFinder Scraper — Listings, Prices & Agent Leads
Scrape PropertyFinder real estate in Dubai, UAE & the Gulf: price, beds, size, geo and agent/broker email & phone. No login or API key. Export JSON, CSV, Excel.
DexScreener Scraper - Boosted & Trending Tokens API
DexScreener scraper, no API key: boosted/trending tokens + marketing spend, search pairs, price/volume/liquidity, plus token security audits. Export CSV, JSON.
A crawler a marketplace trusts is a probe of the very defenses a bug bounty tests. The recon that maps a target and the engineering that extracts from one are the same engineering — only the output schema changes.
Vulnerability research
On web applications, inside published scope, under coordinated disclosure.
Attack surface, mapped at scale
Certificate Transparency, DNS, TLS, archived paths and sitemaps — resolved into an inventory of hosts and endpoints before anything is tested by hand.
Read by hand where it counts
Authorisation boundaries, OAuth and session flows, server-side fetches and business logic. The classes a scanner has no signature for are the ones worth the hours.
Disclosed, never published
Testing stays inside a published scope, proof stops at the shallowest depth that shows impact, and nothing goes public before a fix ships.
Browse scrapers by category
16 categories, every actor tagged for fast discovery.
B2B contact, registry & prospect data for sales pipelines.
Competitor intel, market research, registries.
APIs, datasets and dev infrastructure scrapers.
Headless workflows, data pipelines, API replacements.
Niche datasets and utilities.
Reddit, LinkedIn, podcast & content platforms.
Ad libraries, campaign data, brand monitoring.
Property listings, prices, agents — EU & global.
Product catalogs, prices, merchant intelligence.
Job board scrapers across geos and verticals.
Sitemap, schema, broken link, technical SEO data.
News aggregation, RSS, content monitoring.
Video & podcast platform data.
AI training data, models, datasets, RAG inputs.
Hotel prices, OTA, destination data.
Routes, scores, live event data.
Latest guides
Deep dives on getting real-world sites to give up their data.
How to Build a Deep-Research Retrieval Layer for AI Agents
Give an agent a topic and get ranked web sources, full-page Markdown, and recent news with citations — keyless multi-source retrieval for RAG and grounding.
How to Extract Structured Data from Any URL in 2026
Turn any web page into clean JSON without an LLM: parse schema.org JSON-LD, OpenGraph, tables, prices and contacts deterministically — no API key, no browser.
How to Add Live Web Search to AI Agents (No API Key)
Give your LLM fresh, citable search results without a Tavily or SerpAPI key. How live SERP extraction works: ranked results plus Markdown page content for RAG.
How to Scrape Allabolag.se Sweden Company Leads in 2026
Extract Swedish company data from allabolag.se — org number, revenue, CEO, phone and email — without an API key. A guide to bulk firmographics and B2B leads.
How to Scrape arXiv Papers, Abstracts & Metadata in 2026
Build a research-paper dataset from arXiv's public API: titles, full abstracts, author lists, categories, PDF links and DOIs. No API key, no browser, at scale.
How to Scrape the Australia Business Register (ABN/ABR)
Bulk-export Australian business names, ABNs, status and registration dates from the official data.gov.au CKAN register — no API key, no browser, no captcha.
Need a scraper that doesn't exist yet — or found a bug in one?
If the target is publicly accessible and the data justifies a recurring pipeline, say so and it gets built on the same infrastructure as everything listed here. Security reports go to the same address with [SECURITY] in the subject and are answered first.