gurl
Fetch any URL. Get clean markdown. Save 90%+ tokens.
When an AI agent fetches a web page with curl or web_fetch, it gets raw HTML. A typical docs page is 300K+ tokens of tags, scripts, and navigation. The agent wastes context window, money, and reasoning on noise.
gurl fetches the same URL and returns clean markdown in a structured JSON envelope. Same content, 90-100% fewer tokens.
curl https://docs.stripe.com/api → 279,363 tokens (raw HTML)
gurl https://docs.stripe.com/api → 2,814 tokens (clean markdown) — 99% saved
gurl get https://docs.stripe.com/api{
"url": "https://docs.stripe.com/api",
"response": { "status": 200, "headers": { "content-type": "text/html; charset=utf-8" } },
"content": {
"type": "markdown",
"original_type": "text/html; charset=utf-8",
"title": "Stripe API Reference",
"body": "# Stripe API Reference\n\nThe Stripe API is organized around REST..."
},
"timing": { "total_ms": 1139 }
}Every response is a JSON envelope with metadata + clean content. HTML becomes markdown. PDFs become text. JSON stays structured. Agents parse one format, always.
# Shell installer (macOS / Linux)
curl -fsSL https://raw.githubusercontent.com/HamedMP/gurl/main/install.sh | sh
# Homebrew
brew install HamedMP/gurl/gurl
# Cargo
cargo install gurl-cli
# Debian / Ubuntu
curl -LO https://github.com/HamedMP/gurl/releases/latest/download/gurl_0.1.0_amd64.deb
sudo dpkg -i gurl_0.1.0_amd64.deb
# Arch Linux (AUR)
# PKGBUILD available at dist/PKGBUILD
# Pre-built binaries
# Download from https://github.com/HamedMP/gurl/releasesOr build from source:
git clone https://github.com/HamedMP/gurl.git
cd gurl
cargo build --release
# Binary at ./target/release/gurl# Fetch a page — returns structured JSON envelope with markdown content
gurl get https://docs.stripe.com/api
# Body only (for piping to agents)
gurl get https://react.dev/reference/react --quiet
# Raw response (no conversion, actual HTML/bytes)
gurl get https://example.com --raw
# Extract specific fields with dot notation
gurl get https://httpbin.org/json --select content.body
# POST with JSON body
gurl post https://httpbin.org/post --json '{"key": "value"}'
# Form data
gurl post https://httpbin.org/post -F "name=gurl" -F "type=cli"
# Custom headers
gurl get https://api.example.com -H "Authorization: Bearer token"
# Save to file
gurl get https://arxiv.org/pdf/1706.03762 -o paper.md
# Timeout
gurl get https://slow-site.com --timeout 5| Flag | Output |
|---|---|
| (default) | Full JSON envelope (pretty in terminal, compact when piped) |
--quiet |
Body content only |
--raw |
Original response bytes, no conversion |
--select path |
Extract field via dot notation (content.body, response.status) |
-o file |
Save body to file |
gurl auto-detects content type and converts to the optimal format for AI consumption:
| Input | Output | Method |
|---|---|---|
| HTML pages | Markdown | Readability + noise stripping |
| Next.js RSC sites | Markdown | RSC payload extraction |
| Cookie-wall pages | Markdown | Banner removal + article extraction |
| PDF documents | Text with page breaks | pdf-extract |
| JSON / API responses | Structured JSON | Pass-through with formatting |
| CSV / TSV | Markdown tables | Column detection |
| DOCX | Markdown | XML parsing |
| XLSX | Markdown tables | Sheet extraction |
| EPUB | Markdown | Chapter extraction |
| PPTX | Markdown | Slide-by-slide text |
| RSS / Atom | Markdown | Feed item extraction |
| Images (EXIF) | Metadata markdown | EXIF tag extraction |
| Outlook .msg | Markdown | CFB parsing |
| Jupyter notebooks | Markdown | Cell extraction |
| Wikipedia | Markdown | Infobox + article extraction |
| ZIP archives | Content listing | File enumeration |
| Plain text | Pass-through | Charset detection |
Modern Next.js sites (Vercel docs, etc.) render content client-side via React Server Components. The actual text lives in self.__next_f.push() script chunks, not the DOM. gurl extracts content directly from the RSC payload — no headless browser needed.
Before extraction, gurl removes cookie banners, consent dialogs, navigation, sidebars, footers, modals, and inline scripts. A quality gate rejects results that are predominantly navigation links.
SITE gurl_ms curl_tok gurl_tok saved STATUS
──────────────────────────────────────────────────────────────────────
python-stdlib 271ms 19531 5558 72% OK
rust-std 178ms 13240 7818 41% OK
mdn-js 77ms 54735 5976 90% OK
react-ref 176ms 39001 1175 97% OK
nextjs-docs 190ms 191117 8184 96% OK
fastapi 119ms 40131 8906 78% OK
stripe-api 1139ms 279363 2814 99% OK
github-rest 55ms 103150 5106 96% OK
anthropic-api 1070ms 122772 2212 99% OK
vercel-nextjs 278ms 467820 3935 100% OK
claude-tools 1483ms 169138 4124 98% OK
cloudflare-workers 81ms 61316 1382 98% OK
neon-docs 247ms 97939 1608 99% OK
supabase-docs 152ms 69433 1689 98% OK
redis-cmds 140ms 172909 18101 90% OK
docker-ref 109ms 73236 16024 79% OK
k8s-concepts 147ms 120643 752 100% OK
wikipedia-rust 322ms 146789 24417 84% OK
arxiv-pdf 316ms 541536 9903 99% OK
...
Results: 32/34 passed (>200 chars extracted)
curl_tok is what curl or web_fetch costs your agent in context window tokens. gurl_tok is what gurl costs. The difference is wasted money.
Run yourself: bash bench/sites.sh
gurl-cli # CLI binary (clap)
|
+-- gurl-core # HTTP client + content pipeline
|
+-- markitdown-rs # Document conversion (Rust port of Microsoft MarkItDown)
|
+-- 15 converters (HTML, PDF, DOCX, XLSX, CSV, EPUB, RSS, ...)
+-- MIME detection (magic bytes + extension)
+-- Charset detection (encoding_rs)
Three crates in a Cargo workspace:
| Crate | Purpose |
|---|---|
gurl-cli |
CLI interface, output formatting, TTY detection |
gurl-core |
HTTP client (reqwest), response envelope, content routing |
markitdown-rs |
15 document converters, content detection, noise stripping |
A standalone Rust port of Microsoft MarkItDown with feature flags for each converter:
[dependencies]
markitdown-rs = { version = "0.1", features = ["html", "pdf"] }
# or everything:
markitdown-rs = { version = "0.1", features = ["all"] }Available features: html, pdf, docx, xlsx, csv-convert, epub, rss, image, outlook, ipynb, pptx, wikipedia, zip-convert
gurl is designed as a drop-in tool for AI agents. The structured JSON envelope means your agent always gets:
# Pseudo-code for any agent framework
result = run("gurl get https://docs.stripe.com/api --quiet")
# result is clean markdown, ready for the context windowCompared to alternatives:
| Tool | Output | Tokens for Stripe API docs |
|---|---|---|
curl |
Raw HTML | ~279K |
Claude web_fetch |
HTML with cookie noise | ~169K |
gurl |
Clean markdown | ~2.8K |
The JSON envelope also includes timing, headers, and content metadata — useful for agents that need to make decisions based on response characteristics.
# Run tests
cargo test --all-features # 52 tests
# Build release
cargo build --release
# Run benchmarks
bash bench/compare.sh # gurl vs curl comparison
bash bench/sites.sh # 34-site extraction quality testMIT
gurl. Fetch anything. Read everything. Waste nothing.