This guide puts CAPTCHA handling where it belongs in a training-corpus or RAG pipeline: in the fetch layer, behind an authorization gate, with every document carrying a record of how it was obtained. You get a pre-solve gate for robots.txt AI tokens and TDM reservations, a session-level solver built on CaptchaAI, and drop-in ingestion code for a LlamaIndex reader, an Airbyte Python CDK source and a Chroma collection. One constraint shapes all of it: being able to solve a CAPTCHA is not permission to collect, so the gate runs before the solver, never after.
The code is Python 3.10+ (pip install requests beautifulsoup4, plus llama-index-core, airbyte-cdk or chromadb for the section you use), with a Node.js version of the solve helper (tested with undici 7.30), and was exercised against local mock servers with llama-index-core 0.14.25, airbyte-cdk 7.32.0 and chromadb 1.5.9. Export your key as CAPTCHAAI_API_KEY (see the quickstart). Agents that browse at run time are a separate problem, covered in solving CAPTCHAs in AI browser agents.
When a CAPTCHA is a stop sign: the authorization checklist
A CAPTCHA is the site operator's control over automated access. Before any solve, settle whether you may collect this content and use it for this purpose. A defensible yes rests on things you can write down, per host:
- A written basis. Your own property, an explicit licence (open-data, Creative Commons), a contract or data partnership, or the operator's written permission. "It is publicly reachable" is not, on its own, a basis for training use.
- Nothing auth-walled or private that you do not own. Login walls, paywalls, member areas and other people's accounts stay out, CAPTCHA or not. Leave personal data out too; it needs its own legal basis.
- robots.txt, including the AI tokens. Obey the groups for your crawler's own token and for
*, then read those aimed at AI crawlers:GPTBot(OpenAI says it crawls content that may be used to train its foundation models),CCBot(Common Crawl's crawler) andGoogle-Extended, a control-only token with no user agent of its own that decides whether content Google crawls may train Gemini models and be used for grounding. None of these groups addresses your crawler, but each states the operator's position on AI use, and a conservative gate treats a Disallow there as a no. - TDM reservations. In the EU, the general text-and-data-mining exception in Article 4 of Directive (EU) 2019/790, the one open to commercial users (Article 3 covers research organisations), applies only where rightholders have not reserved their rights "in an appropriate manner, such as machine-readable means" for content made public online. Recital 18 counts metadata and a website's terms and conditions as such means. If you provide a general-purpose AI model in the EU, Article 53(1)(c) of the AI Act also requires a copyright policy that identifies and complies with these reservations. The W3C community's TDM Reservation Protocol is one machine-readable form: a
/.well-known/tdmrep.jsonfile, atdm-reservationHTTP header and a<meta name="tdm-reservation" content="1">tag. What counts as a valid reservation in your jurisdiction is a question for counsel, not code. - Licence conditions. Record the licence identifier and any attribution or non-commercial condition; they travel with the document into the corpus.
- A logged decision. Who approved the source, on what basis, and when. The gate below writes one JSON line per URL, which pairs well with an audit log for CAPTCHA solving.
Even with a yes, the cleanest fix is often not a solver: ask the operator to exempt your crawler's IPs or user agent from the challenge, or for a feed or dump. CaptchaAI fits the remaining case, where you are authorized but the challenge stays, such as a public data portal that challenges all automated traffic or a partner whose security settings you cannot change. The guidelines on authorized CAPTCHA solving cover the wider picture.
Pre-solve gate in code: robots.txt AI tokens, TDM reservations, source allowlist
The gate answers one question per URL before any fetch or solve: may this enter the corpus for this use (rag or training)? Three choices in it are deliberate:
- robots.txt is fetched with your own session, so the request carries your crawler's User-Agent, and parsed with the standard library's
RobotFileParser. A 401, a 403, a 5xx or a challenged robots.txt means no permission is established. That is stricter than RFC 9309, which lets crawlers treat a 4xx as "no rules" but requires them to assume complete disallow on a 5xx. A 404 means no rules. - TDMRep order follows the protocol. The first matching rule in
tdmrep.jsoncounts as the most specific, atdm-reservationresponse header supersedes it, and the page's meta tag supersedes both. The meta tag can only be read after the fetch, so the reader checks it later, and only to add a reservation: a URL the gate refused is never fetched to look for a0. Atdmrep.jsonthat answers 401, 403, 5xx or a challenge is treated like an unreadable robots.txt: no permission is established. - Every decision is logged, allowed or not, with the licence that will be stamped on each document.
"""source_gate.py: decide whether a URL may enter the corpus before anything is fetched or solved."""
import json
import re
import time
from dataclasses import asdict, dataclass
from functools import lru_cache
from urllib.parse import urlsplit
from urllib.robotparser import RobotFileParser
import requests
CRAWLER_TOKEN = "AcmeCorpusBot" # your crawler's own robots.txt product token
AI_TOKENS = ("GPTBot", "CCBot", "Google-Extended") # the operator's stated position on AI use
# A host enters the corpus only with a written basis. The licence travels with every document.
ALLOWLIST = {
"data.example.gov": {"licence": "CC-BY-4.0", "basis": "open-data licence", "uses": {"rag", "training"}},
"docs.partner.example": {"licence": "contract-2026-014", "basis": "partner agreement", "uses": {"rag"}},
}
@dataclass
class GateDecision:
url: str
use: str
allowed: bool
reason: str
licence: str | None
checked_at: str
@lru_cache(maxsize=256) # once per host per run; re-checked on the next crawl
def _robots(origin: str, http: requests.Session) -> RobotFileParser | None:
resp = http.get(f"{origin}/robots.txt", timeout=20)
if resp.status_code in (401, 403) or resp.status_code >= 500 or "cf-mitigated" in resp.headers:
return None # the rules cannot be read, so no permission is established
parser = RobotFileParser()
parser.parse(resp.text.splitlines() if resp.ok else []) # a 404 means no rules
return parser
@lru_cache(maxsize=256)
def _tdm_rules(origin: str, http: requests.Session) -> tuple | None:
resp = http.get(f"{origin}/.well-known/tdmrep.json", timeout=20)
if resp.status_code in (401, 403) or resp.status_code >= 500 or "cf-mitigated" in resp.headers:
return None # unreadable, as with robots.txt: no permission is established
try:
rules = resp.json() if resp.ok else [] # a 404 means no file, so no reservation from it
except ValueError:
return ()
return tuple(r for r in rules if isinstance(r, dict)) if isinstance(rules, list) else ()
def _tdm_reserved(origin: str, path: str, http: requests.Session) -> bool:
"""TDMRep: the /.well-known/tdmrep.json rule first, then a tdm-reservation header supersedes it."""
rules = _tdm_rules(origin, http)
reserved = rules is None
for rule in rules or (): # the first matching rule is the most specific one
pattern = re.escape(rule.get("location", "")).replace(r"\*", ".*").replace(r"\$", "$")
if re.match(pattern, path):
reserved = rule.get("tdm-reservation") == 1
break
header = http.head(f"{origin}{path}", allow_redirects=True, timeout=20).headers.get("tdm-reservation")
return header.strip() == "1" if header is not None else reserved
def check_source(url: str, use: str, http: requests.Session, log_path: str = "gate-log.jsonl") -> GateDecision:
parts = urlsplit(url)
origin, entry = f"{parts.scheme}://{parts.netloc}", ALLOWLIST.get(parts.hostname or "")
if entry is None or use not in entry["uses"]:
allowed, reason = False, f"no written basis for '{use}' use"
elif (robots := _robots(origin, http)) is None:
allowed, reason = False, "robots.txt unreadable (401, 403, 5xx or challenged)"
elif not robots.can_fetch(CRAWLER_TOKEN, url):
allowed, reason = False, f"robots.txt disallows {CRAWLER_TOKEN}"
elif blocked := [t for t in AI_TOKENS if not robots.can_fetch(t, url)]:
allowed, reason = False, f"robots.txt disallows AI tokens {blocked}"
elif _tdm_reserved(origin, parts.path or "/", http):
allowed, reason = False, "TDM rights reserved, or tdmrep.json unreadable (TDMRep)"
else:
allowed, reason = True, entry["basis"]
decision = GateDecision(url, use, allowed, reason, entry["licence"] if entry else None,
time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()))
with open(log_path, "a", encoding="utf-8") as log:
log.write(json.dumps(asdict(decision)) + "\n")
return decision
Run the gate again on every crawl. Operators add AI-token rules and TDM reservations over time, and a source that was open last quarter may not be now.
Architecture: fetch, session-level CAPTCHA gate, extract, provenance, store or embed
source list ─► compliance gate ─► fetch (one session per host, one proxy, one UA)
│ no │
▼ ▼
gate-log.jsonl challenge detection ── unsupported ──► quarantine.jsonl
│ supported
▼
CaptchaAI solve (once per session)
│
▼
extract text ─► provenance record ─► chunk ─► embed ─► upsert ─► refresh schedule (re-gate, re-fetch)
Three rules keep this sane:
- The model never sees a CAPTCHA. The fetch layer returns either a document or a quarantine record. Nothing about challenges or tokens reaches a prompt, a tool call or an embedding.
- Solve per session, not per document. A Cloudflare
cf_clearancecookie covers every page on that host for as long as the same IP and User-Agent are used; its default lifetime is 30 minutes, and the site owner can change it. A widget-gated search form, once submitted, usually leaves the site's own session cookie behind. Either way, one solve unlocks many documents. - Downstream stages are CAPTCHA-agnostic. Chunking, embedding and storage only see a
captcha_type_seenfield in the provenance record.
Detecting the challenge: supported vs unsupported types
Detection runs on every response, in this order:
| Signal in the response | What it is | CaptchaAI | Pipeline action |
|---|---|---|---|
cf-mitigated: challenge header |
Cloudflare Challenge interstitial | ✅ cloudflare_challenge (proxy required) |
Solve once, then reuse the cookie, User-Agent and proxy |
hcaptcha.com script or h-captcha element |
hCaptcha | ❌ Not supported | Quarantine |
arkoselabs.com or funcaptcha.com script |
FunCaptcha (Arkose Labs) | ❌ Not supported | Quarantine |
recaptcha/enterprise.js |
reCAPTCHA Enterprise | ✅ userrecaptcha with enterprise=1; answer in result with a user_agent to reuse |
Not wired in this helper; quarantine for manual review |
geetest, frc-captcha, captchafox or leminnow in the page |
GeeTest, Friendly Captcha, CaptchaFox, Lemin | ✅ GeeTest v3; ✅ Beta for Friendly Captcha, CaptchaFox and Lemin; ❌ GeeTest v4 | Not wired in this helper; quarantine for manual review |
.cf-turnstile[data-sitekey] element |
Cloudflare Turnstile widget | ✅ turnstile |
Solve when submitting the gated form; token goes in cf-turnstile-response |
recaptcha/api.js?render=<sitekey> plus grecaptcha.execute(..., {action}) |
reCAPTCHA v3 | ✅ userrecaptcha with version=v3 and action |
Same; token goes in g-recaptcha-response |
.g-recaptcha[data-sitekey] element |
reCAPTCHA v2 | ✅ userrecaptcha |
Same |
The header check comes first for a reason. Cloudflare sets cf-mitigated on every Challenge Page response and challenge is its only value, which makes it the reliable signal (Cloudflare's detection guide). An interstitial can also contain a Turnstile widget, and it must be handled as a Challenge (a clearance cookie), not as a form token. The not-wired rows matter in a corpus pipeline: without them, a page gated by one of those widgets comes back as none and is ingested as a document. Image and grid CAPTCHAs and BLS are supported by CaptchaAI too, but they have no reliable page signature, so this detector does not look for them. GeeTest v4 support is coming soon; it is not available yet.
Solving once per session: cookies, UA/proxy consistency, rate limits
The helper calls the CaptchaAI API as documented:
- Submit to
https://ocr.captchaai.com/in.phpwithkey,methodandjson=1. Turnstile addssitekeyandpageurl(andactionwhen the widget declares one); reCAPTCHA v2 addsgooglekeyandpageurl, and v3 alsoversion=v3andaction, with no proxy. Cloudflare Challenge requirespageurl,proxyandproxytype. On success,requestholds the task ID. - Poll
https://ocr.captchaai.com/res.phpwithaction=get, the task ID andjson=1: first after 15 seconds for Turnstile and 20 for reCAPTCHA and Cloudflare Challenge, then every 5 seconds while the answer isCAPCHA_NOT_READY, for at most 120 seconds. - Read the answer. Tokens arrive in
request. A Cloudflare Challenge answer arrives inresult(thecf_clearancevalue) with auser_agent, which onlyjson=1returns. - Classify errors.
ERROR_WRONG_USER_KEY,ERROR_KEY_DOES_NOT_EXISTandIP_BANNEDstop the crawl.ERROR_ZERO_BALANCEmeans no free thread (or no active plan), so the helper backs off a bounded number of times. Server errors are retried after about 10 seconds, andERROR_CAPTCHA_UNSOLVABLEquarantines the source. Plain-text codes can arrive even withjson=1, and_read()handles them.
"""captcha_fetch.py: detect a challenge in the fetch layer, solve it once with CaptchaAI, keep the session."""
import os
import re
import time
from dataclasses import dataclass
from urllib.parse import urlsplit
import requests
from bs4 import BeautifulSoup
API = "https://ocr.captchaai.com"
API_KEY = os.environ["CAPTCHAAI_API_KEY"] # YOUR_API_KEY, 32 characters
FIRST_WAIT = {"turnstile": 15, "userrecaptcha": 20, "cloudflare_challenge": 20}
FATAL = {"ERROR_WRONG_USER_KEY", "ERROR_KEY_DOES_NOT_EXIST", "IP_BANNED"}
TRANSIENT = {"ERROR_SERVER_ERROR", "ERROR_INTERNAL_SERVER_ERROR"}
class CaptchaStop(RuntimeError):
"""Key, plan or request problem: stop the crawl and fix it."""
class SourceUnsolvable(RuntimeError):
"""Quarantine this source for review and stop fetching it."""
@dataclass
class Challenge:
kind: str # none | cloudflare_challenge | turnstile | recaptcha_v2 | recaptcha_v3 | unsupported:<name>
sitekey: str | None = None
action: str | None = None
def detect(resp: requests.Response) -> Challenge:
if resp.headers.get("cf-mitigated") == "challenge":
return Challenge("cloudflare_challenge")
html = resp.text
if "hcaptcha.com" in html or 'class="h-captcha' in html:
return Challenge("unsupported:hcaptcha")
if "arkoselabs.com" in html or "funcaptcha.com" in html:
return Challenge("unsupported:funcaptcha")
if "recaptcha/enterprise.js" in html:
return Challenge("unsupported:recaptcha_enterprise_not_wired")
if other := re.search(r"geetest|frc-captcha|captchafox|leminnow", html):
return Challenge(f"unsupported:{other.group(0)}_not_wired") # see the detection table: not wired here
soup = BeautifulSoup(html, "html.parser")
if widget := soup.select_one(".cf-turnstile[data-sitekey]"):
return Challenge("turnstile", widget["data-sitekey"], widget.get("data-action"))
if v3 := re.search(r"recaptcha/api\.js\?[^\"']*render=([\w-]{20,})", html):
action = re.search(r"grecaptcha\.execute\([^)]*action:\s*['\"]([\w/]+)", html)
return Challenge("recaptcha_v3", v3.group(1), action.group(1) if action else "verify")
if widget := soup.select_one(".g-recaptcha[data-sitekey]"):
return Challenge("recaptcha_v2", widget["data-sitekey"])
return Challenge("none")
def _read(resp: requests.Response) -> dict:
try:
return resp.json()
except ValueError: # plain-text codes can arrive even with json=1; HTML 5xx pages too
text = resp.text.strip()
return {"status": 0, "request": text if re.fullmatch(r"[A-Z_]+", text) else "UNREADABLE"}
def solve(ch: Challenge, pageurl: str, proxy: str | None = None) -> dict:
if ch.kind == "turnstile":
params = {"method": "turnstile", "sitekey": ch.sitekey, "pageurl": pageurl}
if ch.action:
params["action"] = ch.action
elif ch.kind == "recaptcha_v2":
params = {"method": "userrecaptcha", "googlekey": ch.sitekey, "pageurl": pageurl}
elif ch.kind == "recaptcha_v3": # no proxy for standard v3
params = {"method": "userrecaptcha", "version": "v3", "googlekey": ch.sitekey,
"pageurl": pageurl, "action": ch.action or "verify"}
elif ch.kind == "cloudflare_challenge" and proxy:
params = {"method": "cloudflare_challenge", "pageurl": pageurl, "proxy": proxy, "proxytype": "HTTP"}
else:
raise SourceUnsolvable(f"{ch.kind} is not solved in this pipeline (a Challenge needs a proxy)")
params.update(key=API_KEY, json=1)
for attempt in range(4): # submit; back off while every plan thread is busy
sub = _read(requests.post(f"{API}/in.php", data=params, timeout=30))
if sub.get("status") == 1:
break
if sub["request"] in TRANSIENT | {"ERROR_ZERO_BALANCE", "UNREADABLE"} and attempt < 3:
time.sleep(10 * 2 ** attempt)
continue
raise CaptchaStop(f"in.php: {sub['request']}") # FATAL codes, bad parameters, bad proxy, no plan
time.sleep(FIRST_WAIT[params["method"]])
deadline = time.monotonic() + 120
while time.monotonic() < deadline:
res = _read(requests.get(f"{API}/res.php", params={
"key": API_KEY, "action": "get", "id": sub["request"], "json": 1}, timeout=30))
if res.get("status") == 1:
return res # token in "request"; Cloudflare Challenge: "result" + "user_agent"
code = res.get("request")
if code in ("CAPCHA_NOT_READY", "UNREADABLE"):
time.sleep(5)
elif code in TRANSIENT:
time.sleep(10)
elif code == "ERROR_CAPTCHA_UNSOLVABLE":
raise SourceUnsolvable(f"{pageurl}: unsolvable, check the challenge type and parameters")
else:
raise CaptchaStop(f"res.php: {code}")
raise SourceUnsolvable(f"{pageurl}: no answer within 120 s")
class GatedSession:
"""One per source host: cookies, User-Agent and proxy stay fixed for the session's lifetime."""
def __init__(self, user_agent: str, proxy: str | None = None, max_solves: int = 3):
self.http = requests.Session()
self.http.headers["User-Agent"] = user_agent
self.proxy = proxy # "user:pass@host:port", the format CaptchaAI expects
if proxy:
self.http.proxies = {"http": f"http://{proxy}", "https": f"http://{proxy}"}
self.max_solves, self.solves, self.seen = max_solves, 0, "none"
def _budget(self) -> None:
if self.solves >= self.max_solves:
raise SourceUnsolvable("solve budget spent: the source keeps challenging, stop and review")
self.solves += 1
def get(self, url: str) -> requests.Response:
resp = self.http.get(url, timeout=30)
ch = detect(resp)
if ch.kind != "cloudflare_challenge":
return resp # widget pages go through unlock_form(); the caller checks detect()
self._budget()
answer = solve(ch, url, self.proxy)
self.seen = ch.kind # what this session cleared, stamped on every document it fetches
self.http.headers["User-Agent"] = answer["user_agent"] # cf_clearance is bound to this UA and IP
self.http.cookies.set("cf_clearance", answer["result"], domain=urlsplit(url).hostname)
resp = self.http.get(url, timeout=30)
if detect(resp).kind == "cloudflare_challenge":
raise SourceUnsolvable(f"{url}: still challenged with a fresh cf_clearance")
return resp
def unlock_form(self, page_url: str, post_url: str, fields: dict) -> requests.Response:
"""Submit a widget-gated form once; the site's own session cookie then covers later fetches."""
ch = detect(self.get(page_url))
if ch.kind not in ("turnstile", "recaptcha_v2", "recaptcha_v3"):
raise SourceUnsolvable(f"{page_url}: expected a widget, found {ch.kind}")
self._budget()
token = solve(ch, page_url)["request"] # single use: submit it straight away
self.seen = ch.kind
field = "cf-turnstile-response" if ch.kind == "turnstile" else "g-recaptcha-response"
return self.http.post(post_url, data={**fields, field: token}, timeout=30)
The session rules matter more than the API calls:
- One proxy, one User-Agent, one cookie jar per host. The
cf_clearancecookie is bound to the solver's User-Agent and IP, so the session fetches through the proxy it gave CaptchaAI and adopts the returneduser_agent. Use a sticky exit IP, not a rotating one, passed asuser:pass@host:portwith no scheme. Proxy use is disabled on CaptchaAI accounts by default, so ask support to enable it before you rely on Challenge solves. The cf_clearance cookie guide goes deeper. - Tokens are single-use and short-lived. A Turnstile token is valid for 300 seconds, a reCAPTCHA token for two minutes, and each can be verified once. So
unlock_form()solves immediately before the POST and never caches a token; the Turnstile API walkthrough covers the widget side. - Identification changes after a Challenge solve. The session now presents the solver's User-Agent, not your crawler's token. If your agreement requires your crawler to identify itself, ask for an exemption instead.
- The solve does not license volume. Honour
Crawl-delay, keep one or two connections per host, and stop on HTTP 429.max_solvescaps re-solves before the source is quarantined for review.
The same helper in Node.js
For Node crawlers, detection, the solve and Challenge clearance carry over; form unlocking and the solve budget are left to the caller. npm install undici now installs undici 8, which needs Node 22.19 or later; on Node 20.18+, install undici@7. In both, ProxyAgent reads the credentials from the proxy URI and tunnels through it, so fetches leave from the IP the solver used.
// captcha-fetch.mjs: the same fetch-layer helper for Node crawlers (Node 22.19+: npm install undici; Node 20.18+: npm install undici@7)
import { pathToFileURL } from "node:url";
import { fetch, ProxyAgent } from "undici";
const API = "https://ocr.captchaai.com";
const KEY = process.env.CAPTCHAAI_API_KEY; // YOUR_API_KEY
const FIRST_WAIT = { turnstile: 15, userrecaptcha: 20, cloudflare_challenge: 20 };
const sleep = (s) => new Promise((resolve) => setTimeout(resolve, s * 1000));
const attr = (tag, name) => tag?.match(new RegExp(`${name}="([^"]+)"`))?.[1];
export class SourceUnsolvable extends Error {}
export function detect(res, html) {
if (res.headers.get("cf-mitigated") === "challenge") return { kind: "cloudflare_challenge" };
if (html.includes("hcaptcha.com")) return { kind: "unsupported:hcaptcha" };
if (html.includes("arkoselabs.com") || html.includes("funcaptcha.com")) return { kind: "unsupported:funcaptcha" };
if (html.includes("recaptcha/enterprise.js")) return { kind: "unsupported:recaptcha_enterprise_not_wired" };
const other = html.match(/geetest|frc-captcha|captchafox|leminnow/);
if (other) return { kind: `unsupported:${other[0]}_not_wired` };
const tags = html.match(/<[^>]+>/g) ?? [];
const widget = (cls) => tags.find((t) => t.includes(cls) && t.includes("data-sitekey="));
const turnstile = widget("cf-turnstile");
if (turnstile) return { kind: "turnstile", sitekey: attr(turnstile, "data-sitekey"), action: attr(turnstile, "data-action") };
const v3 = html.match(/recaptcha\/api\.js\?[^"']*render=([\w-]{20,})/);
if (v3) {
const action = html.match(/grecaptcha\.execute\([^)]*action:\s*['"]([\w/]+)/);
return { kind: "recaptcha_v3", sitekey: v3[1], action: action ? action[1] : "verify" };
}
const v2 = widget("g-recaptcha");
return v2 ? { kind: "recaptcha_v2", sitekey: attr(v2, "data-sitekey") } : { kind: "none" };
}
async function call(path, params) {
const res = await fetch(`${API}/${path}`, { method: "POST", body: new URLSearchParams(params) });
const text = (await res.text()).trim();
try {
return JSON.parse(text);
} catch {
return { status: 0, request: /^[A-Z_]+$/.test(text) ? text : "UNREADABLE" };
}
}
export async function solve(ch, pageurl, proxy) {
const params = {
turnstile: { method: "turnstile", sitekey: ch.sitekey, pageurl, ...(ch.action && { action: ch.action }) },
recaptcha_v2: { method: "userrecaptcha", googlekey: ch.sitekey, pageurl },
recaptcha_v3: { method: "userrecaptcha", version: "v3", googlekey: ch.sitekey, pageurl, action: ch.action },
cloudflare_challenge: proxy && { method: "cloudflare_challenge", pageurl, proxy, proxytype: "HTTP" },
}[ch.kind];
if (!params) throw new SourceUnsolvable(`${ch.kind} is not solved in this pipeline`);
let sub;
for (let attempt = 0; attempt < 4; attempt += 1) {
sub = await call("in.php", { ...params, key: KEY, json: 1 });
if (sub.status === 1) break;
const retryable = ["ERROR_ZERO_BALANCE", "ERROR_SERVER_ERROR", "ERROR_INTERNAL_SERVER_ERROR", "UNREADABLE"];
if (!retryable.includes(sub.request) || attempt === 3) throw new Error(`in.php: ${sub.request}`);
await sleep(10 * 2 ** attempt);
}
await sleep(FIRST_WAIT[params.method]);
for (const end = Date.now() + 120_000; Date.now() < end; ) {
const res = await call("res.php", { key: KEY, action: "get", id: sub.request, json: 1 });
if (res.status === 1) return res; // token in request; Cloudflare Challenge: result + user_agent
if (res.request === "ERROR_CAPTCHA_UNSOLVABLE") throw new SourceUnsolvable(`${pageurl}: unsolvable`);
if (!["CAPCHA_NOT_READY", "UNREADABLE", "ERROR_SERVER_ERROR", "ERROR_INTERNAL_SERVER_ERROR"].includes(res.request)) {
throw new Error(`res.php: ${res.request}`);
}
await sleep(res.request.startsWith("ERROR_") ? 10 : 5);
}
throw new SourceUnsolvable(`${pageurl}: no answer within 120 s`);
}
// Fetch a page through one proxy; clear a Cloudflare Challenge once and keep its cookie and UA.
export async function gatedGet(url, session) {
const dispatcher = session.proxy ? new ProxyAgent(`http://${session.proxy}`) : undefined;
const headers = { "User-Agent": session.userAgent, ...(session.cookie && { Cookie: session.cookie }) };
let res = await fetch(url, { headers, dispatcher });
let html = await res.text();
let kind = detect(res, html).kind;
if (kind === "cloudflare_challenge") {
const answer = await solve({ kind }, url, session.proxy);
Object.assign(session, { userAgent: answer.user_agent, cookie: `cf_clearance=${answer.result}`, seen: kind });
res = await fetch(url, { headers: { "User-Agent": session.userAgent, Cookie: session.cookie }, dispatcher });
html = await res.text();
kind = detect(res, html).kind;
if (kind === "cloudflare_challenge") throw new SourceUnsolvable(`${url}: still challenged`);
}
return { res, html, kind }; // "none", or a widget or unsupported kind for the caller; session.seen = what was cleared
}
if (import.meta.url === pathToFileURL(process.argv[1]).href && process.argv[2]) {
const session = { userAgent: "AcmeCorpusBot/1.0", proxy: process.env.CRAWL_PROXY };
const { res, kind } = await gatedGet(process.argv[2], session);
console.log(res.status, kind, session.seen ?? "none");
}
Provenance records for every document
Every document that leaves the fetch layer carries enough metadata to answer three questions later: where did this come from, were we allowed to take it, and has it changed? The same record drives deletion on opt-out and skipping unchanged pages. The samples below emit the core fields; a fuller record stored next to each chunk looks like this:
{
"source_url": "https://data.example.gov/reports/2026-q2",
"licence": "CC-BY-4.0",
"gate_reason": "open-data licence",
"gate_checked_at": "2026-09-29T08:14:01Z",
"fetched_at": "2026-09-29T08:14:03+00:00",
"http_status": 200,
"captcha_type_seen": "cloudflare_challenge",
"content_sha256": "babc0b6969482c20da239e0993f8dd2797595f9919729672ddfcfa2bf9f9a4d3",
"extractor": "bs4 get_text",
"chunk_index": 0
}
content_sha256is computed over the extracted text, not the raw HTML. Pages embed rotating tokens and timestamps, and a raw-HTML hash would change on every fetch.captcha_type_seenrecords what the fetch layer had to clear. When a host's value changes between crawls, its protection changed, and the operator may be sending you a message.licenceandgate_reasoncome straight from the gate decision, so a training-set build can filter by licence without re-deriving anything.
LlamaIndex: a custom reader that fetches through the solved session
LlamaIndex readers subclass BaseReader from llama_index.core.readers.base. Implement lazy_load_data() and the base class supplies load_data(), which returns list(self.lazy_load_data(...)), plus async variants that run it in a thread. The reader keeps one GatedSession per host, runs the gate per URL, quarantines unsolvable and unsupported sources and yields Document objects with the provenance metadata. Hash-like fields are excluded from what the embedding model and the LLM see; the source URL and licence stay visible for citations.
"""gated_reader.py: a LlamaIndex reader that fetches through the source gate and the solved session."""
import hashlib
import json
from datetime import datetime, timezone
from typing import Iterable
from urllib.parse import urlsplit
from bs4 import BeautifulSoup
from llama_index.core import Document, VectorStoreIndex
from llama_index.core.readers.base import BaseReader
from captcha_fetch import GatedSession, SourceUnsolvable, detect
from source_gate import check_source
PROVENANCE_ONLY = ["content_sha256", "fetched_at", "captcha_type_seen", "gate_reason"]
def quarantine(url: str, reason: str) -> None:
with open("quarantine.jsonl", "a", encoding="utf-8") as q:
q.write(json.dumps({"url": url, "reason": reason}) + "\n")
class GatedWebReader(BaseReader):
def __init__(self, user_agent: str, use: str = "rag", proxy: str | None = None):
self.user_agent, self.use, self.proxy = user_agent, use, proxy
self.sessions: dict[str, GatedSession] = {} # one solved session per host
def lazy_load_data(self, urls: list[str]) -> Iterable[Document]:
for url in urls:
host = urlsplit(url).hostname or ""
if host not in self.sessions:
self.sessions[host] = GatedSession(self.user_agent, self.proxy)
session = self.sessions[host]
decision = check_source(url, self.use, session.http)
if not decision.allowed:
continue
try:
resp = session.get(url)
except SourceUnsolvable as exc:
quarantine(url, str(exc))
continue
kind = detect(resp).kind
if kind.startswith("unsupported:"):
quarantine(url, kind) # hCaptcha, FunCaptcha, or a type left to manual review
continue
if kind != "none" or not resp.ok:
continue # a widget-gated page needs session.unlock_form() first
soup = BeautifulSoup(resp.text, "html.parser")
if soup.find("meta", attrs={"name": "tdm-reservation", "content": "1"}):
continue # page-level TDMRep reservation overrides the site-wide file
text = soup.get_text(" ", strip=True)
doc = Document(
id_=url,
text=text,
metadata={
"source_url": url,
"licence": decision.licence,
"gate_reason": decision.reason,
"fetched_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"content_sha256": hashlib.sha256(text.encode()).hexdigest(),
"captcha_type_seen": session.seen,
},
)
doc.excluded_embed_metadata_keys = PROVENANCE_ONLY # keep hashes out of the vectors
doc.excluded_llm_metadata_keys = PROVENANCE_ONLY
yield doc
if __name__ == "__main__":
reader = GatedWebReader(user_agent="AcmeCorpusBot/1.0 (+https://acme.example/bot)")
documents = reader.load_data(urls=["https://data.example.gov/reports/2026-q2"])
index = VectorStoreIndex.from_documents(documents) # uses Settings.embed_model
print(index.as_query_engine().query("What changed in Q2?"))
VectorStoreIndex.from_documents() embeds with Settings.embed_model. Left at its default, that is OpenAI's embedding model, which needs llama-index-embeddings-openai and an OpenAI key, so set it explicitly if you embed elsewhere. The query step likewise answers with Settings.llm, which defaults to OpenAI through llama-index-llms-openai.
When should you wrap an existing reader instead? llama-index-readers-web ships readers such as SimpleWebPageReader, TrafilaturaWebReader and BeautifulSoupWebReader. They are the right choice for sources that never challenge. SimpleWebPageReader, for example, fetches each URL with a plain requests.get() and no headers, cookies or session, so it cannot carry a cf_clearance cookie or the User-Agent it is bound to. For challenged sources, keep the fetch in your own session as above and reuse only the extraction logic you like.
Airbyte: a Python CDK source with a solved session
Airbyte recommends its Connector Builder for most new API sources; the Python CDK is the most flexible option, at the cost of more code. A probe, a solve, and a cookie and User-Agent reused through one proxy is custom, stateful logic, so this source subclasses the CDK's HttpStream (see the HTTP streams documentation). The hooks:
stream_slices()yields one slice per document path and runs the gate on each, so a disallowed path never produces a request.request_headers()solves lazily on its first call and returns the solver's User-Agent and thecf_clearancecookie. The CDK calls it for every page request it builds, but a retry resends the same prepared request, so a fresh solve cannot be slipped into a retry.request_kwargs()returns the session's proxies, so the sync leaves from the IP the solve used.get_error_handler()is the current hook. CDK 3.0.0 removedshould_retryandbackoff_timefromHttpStream(a stream that still defines them is wrapped in deprecated adapters) and deprecatedraise_on_http_errors;get_error_handler()andget_backoff_strategy()replace them. The default mapping fails a 403 as aconfig_error("HTTP Status Code: 403. Error: Forbidden. You don't have permission to access this resource."), which misdescribes an expired clearance. The custom handler fails the stream with atransient_errorinstead; the next sync attempt starts in a new process with no clearance and solves again.parse_response()emits one record per document with the provenance fields, and logs and skips a page that still shows a widget or a type this helper does not solve, so a challenge page never becomes a record.
"""source_gated_docs.py: an Airbyte Python CDK source that reads documents through a solved session."""
import hashlib
import logging
import sys
from datetime import datetime, timezone
from typing import Any, Iterable, List, Mapping, Optional, Tuple
import requests
from airbyte_cdk.entrypoint import launch
from airbyte_cdk.models import ConnectorSpecification, FailureType
from airbyte_cdk.sources import AbstractSource
from airbyte_cdk.sources.streams import Stream
from airbyte_cdk.sources.streams.http import HttpStream
from airbyte_cdk.sources.streams.http.error_handlers import (
ErrorHandler, ErrorResolution, HttpStatusErrorHandler, ResponseAction)
from bs4 import BeautifulSoup
from captcha_fetch import GatedSession, detect
from source_gate import check_source
class ChallengeAware(HttpStatusErrorHandler):
"""A Challenge served mid-sync means the clearance expired: fail as transient and re-solve next attempt."""
def __init__(self, stream: "GatedDocuments"):
super().__init__(logger=logging.getLogger("airbyte"))
self.stream = stream
def interpret_response(self, response_or_exception=None) -> ErrorResolution:
if isinstance(response_or_exception, requests.Response) and \
detect(response_or_exception).kind == "cloudflare_challenge":
self.stream.clearance = None
return ErrorResolution(ResponseAction.FAIL, FailureType.transient_error,
"Cloudflare Challenge returned mid-sync; the next attempt solves again")
return super().interpret_response(response_or_exception)
class GatedDocuments(HttpStream):
primary_key = "source_url"
def __init__(self, config: Mapping[str, Any]):
self.config = config
self.session = GatedSession(config["user_agent"], config.get("proxy"))
self.clearance: Optional[dict] = None
super().__init__()
@property
def url_base(self) -> str:
return self.config["base_url"]
def get_error_handler(self) -> Optional[ErrorHandler]:
return ChallengeAware(self)
def stream_slices(self, **kwargs) -> Iterable[Optional[Mapping[str, Any]]]:
for path in self.config["paths"]: # the gate runs per document, not once per source
if check_source(self.url_base + path, "rag", self.session.http).allowed:
yield {"path": path}
def path(self, *, stream_state=None, stream_slice=None, next_page_token=None) -> str:
return stream_slice["path"]
def next_page_token(self, response: requests.Response) -> Optional[Mapping[str, Any]]:
return None
def request_headers(self, stream_state, stream_slice=None, next_page_token=None) -> Mapping[str, Any]:
if self.clearance is None: # probe once; GatedSession solves a Challenge if one is served
self.session.get(self.url_base)
cookies = "; ".join(f"{c.name}={c.value}" for c in self.session.http.cookies)
self.clearance = {"User-Agent": self.session.http.headers["User-Agent"], "Cookie": cookies}
return self.clearance
def request_kwargs(self, stream_state, stream_slice=None, next_page_token=None) -> Mapping[str, Any]:
return {"proxies": self.session.http.proxies, "timeout": 30} # same IP as the solve
def parse_response(self, response, *, stream_state, stream_slice=None, next_page_token=None):
if (kind := detect(response).kind) != "none": # a widget or an unwired type: review, not a record
self.logger.warning("%s skipped: %s", response.url, kind)
return
text = BeautifulSoup(response.text, "html.parser").get_text(" ", strip=True)
yield {
"source_url": response.url,
"text": text,
"content_sha256": hashlib.sha256(text.encode()).hexdigest(),
"fetched_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"captcha_type_seen": self.session.seen,
"licence": self.config["licence"],
}
def get_json_schema(self) -> Mapping[str, Any]:
fields = ("source_url", "text", "content_sha256", "fetched_at", "captcha_type_seen", "licence")
return {"type": "object", "properties": {f: {"type": "string"} for f in fields}}
class SourceGatedDocs(AbstractSource):
def spec(self, logger: logging.Logger) -> ConnectorSpecification:
required = ["base_url", "paths", "licence", "user_agent"]
return ConnectorSpecification(connectionSpecification={
"type": "object", "required": required,
"properties": {**{k: {"type": "string"} for k in required if k != "paths"},
"paths": {"type": "array", "items": {"type": "string"}},
"proxy": {"type": "string", "airbyte_secret": True}}})
def check_connection(self, logger, config) -> Tuple[bool, Optional[Any]]:
decision = check_source(config["base_url"], "rag", GatedSession(config["user_agent"]).http)
return decision.allowed, None if decision.allowed else decision.reason
def streams(self, config: Mapping[str, Any]) -> List[Stream]:
return [GatedDocuments(config)]
if __name__ == "__main__":
launch(SourceGatedDocs(), sys.argv[1:])
Run it like any CDK source: python source_gated_docs.py check --config config.json, then discover, then read with a configured catalog. In a local run the CDK emitted a state message after each slice, so a sync that fails mid-way records which documents were already read. For a true incremental stream, add a cursor such as the site's last-modified date instead of refetching every path.
Vector-DB ingestion: dedupe, re-crawl cadence and deleting on opt-out
The ingestion side needs three behaviours: do not re-embed unchanged pages, do not leave orphaned chunks when a page shrinks, and remove a whole source cleanly when permission ends. With Chroma, collection.upsert() updates records whose IDs exist and adds the rest. Deterministic IDs, sha256(url + "#" + chunk_index), make re-ingestion idempotent.
"""ingest_chroma.py: chunk, dedupe and upsert documents with provenance; forget a source on opt-out."""
import hashlib
import chromadb
client = chromadb.PersistentClient(path="./corpus-db")
corpus = client.get_or_create_collection("corpus")
def sha256(value: str) -> str:
return hashlib.sha256(value.encode()).hexdigest()
def chunks(text: str, size: int = 1500, overlap: int = 200) -> list[str]:
step = size - overlap
return [text[i:i + size] for i in range(0, max(len(text) - overlap, 1), step)]
def ingest(url: str, text: str, provenance: dict) -> str:
"""provenance: licence, fetched_at, captcha_type_seen, gate_reason (str/int/float/bool only)."""
digest = sha256(text)
current = corpus.get(where={"source_url": url}, include=["metadatas"])
if current["ids"] and current["metadatas"][0]["content_sha256"] == digest:
return "unchanged" # same extracted text: no re-embedding
pieces = chunks(text)
ids = [sha256(f"{url}#{i}") for i in range(len(pieces))]
if stale := set(current["ids"]) - set(ids):
corpus.delete(ids=list(stale)) # the page got shorter
corpus.upsert(
ids=ids,
documents=pieces, # embedded with the collection's embedding function
metadatas=[{**provenance, "source_url": url, "source_host": url.split("/")[2],
"content_sha256": digest, "chunk_index": i} for i in range(len(pieces))],
)
return "upserted"
def forget_host(host: str) -> None:
"""Opt-out, licence ended or robots.txt now disallows: remove every chunk from that host."""
corpus.delete(where={"source_host": host})
if __name__ == "__main__":
print(ingest("https://data.example.gov/reports/2026-q2", "Quarterly report text ... " * 200,
{"licence": "CC-BY-4.0", "fetched_at": "2026-09-29T08:14:03+00:00",
"captcha_type_seen": "cloudflare_challenge", "gate_reason": "open-data licence"}))
- Skip unchanged pages. If the stored
content_sha256matches the new text, nothing is re-embedded. That covers most pages on most re-crawls and keeps embedding cost proportional to change. - Delete on opt-out. When a host adds an AI-token Disallow or a TDM reservation, or its licence ends,
forget_host()removes its chunks with a metadata filter. Log it next to the gate decision that triggered it. - Quarantine, do not retry forever. Have the next crawl skip sources listed in
quarantine.jsonluntil someone reviews them; re-solving a source that keeps returningERROR_CAPTCHA_UNSOLVABLEwastes threads. - Re-crawl cadence per source. Set it from how often the source changes, and re-run the gate each time. Chroma's default embedding function is a local all-MiniLM-L6-v2 model; with pgvector, the same columns and an
INSERT ... ON CONFLICTupsert do the job.
Training corpora: dedup, licences and refresh
A training corpus has stricter needs than a RAG index, because a mistake is baked into a model instead of being fixed by deleting a row.
- Dedup in two passes. Exact duplicates fall out of
content_sha256. Near-duplicates (the same article on a mirror, or a page with a changed footer) need MinHash signatures with locality-sensitive hashing, run per snapshot before tokenization. - Licence as a first-class column. Build each training set from a query over
licence, not from a directory of files. A licence that excludes commercial use excludes the document from a commercial model's corpus, and attribution licences need thesource_urlkept. - Snapshots and refresh. Freeze a dated snapshot per training run and keep the gate log with it. When a refresh finds a new reservation, drop the source from the next snapshot and record why; whether earlier snapshots are affected is a legal question.
- The CAPTCHA does not change the terms. A challenge controls access to the transport; a licence governs use of the content. Solving one gets you the bytes and grants no rights. If a partner adds a challenge after your agreement was signed, tell them: it may be deliberate.
Throughput and cost: sizing threads for crawls
CaptchaAI bills per concurrent thread with unlimited solves per thread, and a thread is one in-flight CAPTCHA. Because this pipeline solves per session, not per page, the number that sets your plan is peak concurrent solves.
Take 8 crawl workers over 20 Cloudflare-protected hosts. If each worker keeps its own session per host, that is 160 sessions, each re-solving when its clearance lapses (every 30 minutes by default): about 320 solves an hour, which is small. The peak is what counts. A worker blocks on its own solve, so at most 8 solves are in flight, all 8 at a cold start. BASIC ($15/month, 5 threads) makes three workers wait at start-up; STANDARD ($30/month, 15 threads) covers it with room for retries; ADVANCE ($90/month, 50 threads) suits fleets of several dozen workers. Sharing one clearance per host across workers on the same sticky proxy IP cuts 160 sessions to 20. Current plans are on the pricing page.
Cloudflare Challenge typically clears in under 15 seconds, Turnstile in under 10, reCAPTCHA v3 in under 4 and reCAPTCHA v2 in under 60, plus your first-wait and polling overhead. Cap concurrent solves with a semaphore set to your plan's threads, and check live usage with the res.php action threadsinfo, which returns threads and working_threads. The thread count guide covers sizing for other workloads.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Gate rejects a whole host with "robots.txt unreadable" | robots.txt returns 403 or sits behind the challenge | Get the rules from the operator in writing; for a host allowlisted by contract, read robots.txt through the solved session |
| Form rejects the solved token | Turnstile token older than 300 seconds or reused; reCAPTCHA token older than two minutes, or a v3 action that does not match the page |
Solve right before the POST, send the action from grecaptcha.execute, never share a token between workers |
Challenge returns straight after setting cf_clearance |
IP or User-Agent mismatch: rotating proxy, or a default UA header overwriting the solver's | Sticky proxy session; use user_agent from the answer verbatim; the same proxy for solve and fetch |
| Challenge returns after about 30 minutes | Clearance expired (default lifetime, set by the site owner) | Re-solve within the session budget; the Airbyte handler fails the attempt as transient and re-solves next time |
ERROR_ZERO_BALANCE during crawl bursts |
All plan threads busy, or no active plan | Semaphore sized to your threads, staggered worker start; check threadsinfo |
Repeated ERROR_CAPTCHA_UNSOLVABLE on one source |
Wrong type or parameters, or a variant the helper does not handle | Mark the source and stop; see fixing ERROR_CAPTCHA_UNSOLVABLE |
ERROR_BAD_PROXY or ERROR_PROXY_CONNECTION_FAILED on cloudflare_challenge |
Proxy marked bad, or unreachable from the solver | Switch to another sticky proxy and submit a new task; confirm proxy use is enabled on your account |
| Chroma re-embeds every page on every crawl | Hash taken over raw HTML with rotating tokens | Hash the extracted text, as ingest() does |
| Airbyte sync fails with "HTTP Status Code: 403. Error: Forbidden. You don't have permission to access this resource." | Default 403 mapping hit a challenge page | Use the custom error handler so challenges surface as transient errors |
FAQ
Does solving the CAPTCHA mean we may use the content for training?
No. The challenge controls access; the licence, the site's terms and any TDM reservation govern use. Decide eligibility in the gate first, and solve only for sources that pass.
Should the LLM or a retrieval agent handle CAPTCHAs at query time?
Not in an ingestion pipeline. Fetching, solving and provenance belong in the batch fetch layer, where decisions are logged and repeatable. Live browsing agents are a separate design; see solving CAPTCHAs in AI browser agents.
What happens to sources behind hCaptcha or FunCaptcha?
CaptchaAI does not solve hCaptcha or FunCaptcha (Arkose Labs), so the detector quarantines those sources. Ask the operator for a feed, an export or an allowlist entry, or drop the source.
Do we need a proxy for every source?
No. Cloudflare Challenge requires one, and it must be the proxy you then fetch through. Turnstile and reCAPTCHA v2 work without one, and CaptchaAI's proxy guide advises against sending one for standard reCAPTCHA v3. Proxy use is disabled on accounts by default, so request it before your first Challenge solve.
Next step: gate first, then solve
Put source_gate.py in front of your crawler, add the fetch helper for the challenge types your sources actually use, and wire the provenance fields into whichever sink you run: a LlamaIndex index, an Airbyte connection or a Chroma collection. Start with one allowlisted host, check the gate log and the captcha_type_seen values, then widen the source list.
Get a CaptchaAI key for your ingestion pipeline's fetch layer